Skip to content

Evaluate your own model

There are three ways to evaluate a voice cloning model with RVCBench.

Route Use it when You change
Score your own outputs You can run the model yourself, or only through an API Nothing; you generate WAV files
External adapter You want RVCBench to drive generation One Python file of your own
Built-in integration You want the model listed in this repository The package, via a pull request

All routes use the same inputs, run records and metrics.

Route 0: score audio you generated anywhere

Your model runs in its own environment, with your own code. RVCBench only needs the resulting WAV files.

python -m pip install 'rvcbench[eval]'
rvcbench setup-scorers
rvcbench prompts --suite core-v1 --output prompts/

prompts/ lists each utterance in three files; use the one your code already reads.

File Line format Read by
prompts.tsv id<TAB>reference_text<TAB>reference_audio<TAB>text ZipVoice --test-list
prompts.lst id\|reference_text\|reference_audio\|text Seed-TTS-eval scripts (F5-TTS, CosyVoice, ...)
prompts.jsonl one JSON object per line (below) your own script
{"id": "english-libritts__LibriTTS-4992-000033", "task": "english-libritts", "pair_id": "LibriTTS-4992-000033",
 "speaker_id": "4992", "reference_audio": "references/english-libritts/LibriTTS-4992-000033.wav",
 "reference_sha256": "3849cd34...", "reference_text": "I wish to prove my friendship to Miss Milner, ...",
 "text": "The Carey children had only found it by accident.", "language": "EN",
 "output_file": "english-libritts/LibriTTS-4992-000033.wav"}

Synthesize each text in the voice of reference_audio and save a mono WAV as <your directory>/<id>.wav (one flat directory, as batch scripts write it) or as <your directory>/<output_file>. Reference paths in prompts.tsv and prompts.lst are absolute, so the lists work from any directory; export again if you move prompts/. Then score:

rvcbench score --suite core-v1 --generated my_outputs/ --output results/my-model/ --device cuda

The model name defaults to the directory name (my_outputs here); set it with --model.

Batch inference scripts

With ZipVoice, for example, the whole loop is its own batch command:

python3 -m zipvoice.bin.infer_zipvoice --model-name zipvoice --test-list prompts/prompts.tsv --res-dir outputs/zipvoice
rvcbench score --suite core-v1 --generated outputs/zipvoice --output results/zipvoice --device cuda

Scripts that read Seed-TTS-eval lists take prompts/prompts.lst and write <utt>.wav, which is the same layout. The lists carry no language column; language is in prompts.jsonl for models that need it.

Several models at once

rvcbench score --suite core-v1 --generated outputs/zipvoice outputs/zipvoice_distill --output results/ --device cuda

Each directory is scored as one model, named after the directory (or --model a b), into results/<model>/. Each metric model is loaded once for all of them. results/comparison.md shows every task and metric with one column per model, the best value in bold and each perturbed condition's change against its clean counterpart; comparison.csv and comparison.json hold the same data.

To compare results scored at different times, for example a new checkpoint against earlier ones:

rvcbench compare results/*/submission.json --output results/

Comparison requires the same suite and matching per-task scoring fingerprints, including scorer code, assets, dependencies and settings. Rescore models together after an environment upgrade. --allow-incompatible explicitly produces an unranked inspection with the differences listed.

Resume a stopped evaluation

Repeat the original scoring command with --resume:

rvcbench score --suite core-v1 --generated my_outputs/ --output results/my-model/ --device cuda --resume

The output directory must belong to the same suite, model name and generated-audio directory. Inputs are checked again. Successful scores with matching audio and scorer fingerprints are reused; missing outputs and failed scores are retried. After replacing a bad WAV, the changed sample is rescored. The same flag works with several model directories. Results from releases before the resume journal was introduced need a new output directory. Avoid simultaneous writers to the same results directory.

Your own inference loop

Read prompts.jsonl and call the inference function from your model's environment:

import json
from pathlib import Path
import soundfile as sf

prompts = Path("prompts")
outputs = Path("my_outputs")
outputs.mkdir(exist_ok=True)
for line in (prompts / "prompts.jsonl").read_text().splitlines():
    item = json.loads(line)
    waveform, sample_rate = your_model.synthesize(  # replace with your model's API
        text=item["text"],
        reference_audio=str(prompts / item["reference_audio"]),
        reference_text=item["reference_text"],
        language=item["language"],
    )
    sf.write(outputs / f"{item['id']}.wav", waveform, sample_rate)

Then run rvcbench score in the scoring environment. A model's native batch script is equally valid.

Results

  • results/my-model/submission.json holds the per-task metric means, coverage, the hash of every scored file and the suite version. Each task also gets a full run directory, so rvcbench status and rvcbench report work on results/my-model/<task>/.
  • A missing or invalid file fails that sample, and a task with a failed sample is partial. A file identical to the dataset's target recording is rejected.
  • The suite pins the dataset revision and the hash of every input. Both commands download only the files the suite uses, and refuse to run if any input differs from the frozen suite. Pass --data-root to use a local copy laid out like the Hub dataset.
  • Target recordings are never exported; they are read only while scoring.
  • core-v1 has 480 utterances across the paper's robustness evaluations; see the Core suite. It is not yet a leaderboard suite.
  • onboarding-v1 (52 utterances) is a smaller suite for checking the workflow.

Route 1: external adapter

1. Install

python -m pip install "rvcbench[eval]"
rvcbench setup-scorers
rvcbench doctor

Install into the environment that already runs your model.

2. Write the adapter

Subclass rvcbench.VoiceCloningAdapter and implement clone:

# my_model_adapter.py
from rvcbench import VoiceCloningAdapter


class MyModelAdapter(VoiceCloningAdapter):
    def load(self):
        # Called once per run. self.config is the `adversary` block of the run config.
        self.model = load_my_model(self.config.checkpoint, device=self.device)

    def clone(self, *, text, reference_audio, reference_text, language):
        # Speak `text` in the voice of the audio file `reference_audio`.
        waveform = self.model.synthesize(text, prompt_wav=str(reference_audio), prompt_text=reference_text)
        return waveform, self.model.sample_rate   # mono float waveform in [-1, 1], sample rate in Hz

The contract:

  • clone is called once per benchmark sample. reference_audio is a pathlib.Path; reference_text is its transcript (may be empty); language is the dataset's language tag or None.
  • Return (waveform, sample_rate). The runner writes the WAV file, names it, and checks that it is nonempty and finite.
  • Raise an exception when an utterance cannot be generated. The runner records the failure for that sample and continues; it does not skip silently.
  • The runner seeds Python, NumPy and PyTorch before each call (run seed plus the sample's source index). Do not reseed inside clone.
  • load and unload are optional and run once per run.
  • The adapter must be importable from a .py file, because each run records a hash of the adapter's source and of the modules it imports from the same package.

examples/echo_adapter.py is a complete adapter that returns the reference audio; it is run by the test suite.

3. Run

PYTHONPATH=. rvcbench run --config-name ots_vc/clean/libritts/custom_ots \
    vc.adapter=my_model_adapter:MyModelAdapter \
    vc.model=my_model run_name=my_model_on_libritts \
    +adversary.checkpoint=/path/to/checkpoint.pt \
    adversary.seed=42 adversary.max_samples=20
  • vc.adapter is package.module:ClassName. The module has to be importable: install your package, or put its directory on PYTHONPATH as above.
  • vc.model is the name recorded in the run and shown in reports. It must differ from the built-in model names (for example xtts), which select model-specific backends and checks.
  • custom_ots is a template config. Keys under adversary reach your adapter as self.config; prefix a key with + when the template does not define it.
  • Add +vc.generate_only=true to generate without scoring.
  • The first run downloads the selected dataset from the Hugging Face Hub. To use a local copy, pass dataset.use_hf_dataset=false dataset.root_path=/path/to/dataset.
  • Other datasets: copy the template next to the other configs of that dataset, or keep your own configs in a directory and pass --config-dir /path/to/your/configs.

4. Read the results

Each run writes results/<run_name>/<timestamp>/:

File Content
run_manifest.json Per-sample status, input and output hashes, seeds, resolved config, environment and source provenance
generated_audio/ One WAV per sample
metrics.json Aggregate metrics with coverage over all requested samples
rvcbench status results/my_model_on_libritts/<timestamp>
rvcbench report results/my_model_on_libritts/<timestamp> --output my_model_report.json

rvcbench report refuses runs with missing outputs or missing required metrics. See the run guide for resume, retries, evaluation-only scoring and the conditions under which two runs may be compared.

Route 2: built-in integration

A pull request that adds a model to this repository contains:

Piece Location Notes
Adapter src/rvcbench/adversary/<model>_ots.py Subclass BaseAdversary. Keep model imports inside methods so that importing the module needs no model dependency.
Model wrapper src/rvcbench/models/<model>/ Loading and inference, with explicit checkpoint paths or pinned Hub revisions.
Registry entry src/rvcbench/benchmark/registry.py "<model>": "rvcbench.adversary.<model>_ots:<ClassName>"
Configs src/rvcbench/configs/ots_vc/clean/<dataset>/<model>_ots.yaml Start from an existing config of the same dataset.
Environment envs/<model>.yml Models have incompatible dependency stacks; each gets its own environment.
Catalog entry src/rvcbench/benchmark/model_catalog.json Use experimental_adapter until a subset validation is recorded.
Tests tests/test_<model>_runtime.py Use fake modules or stub processes. Tests must not download or load a model.
Documentation docs/validation.md, README.md Add a validation row; state what was and was not run.

Requirements for the adapter:

  • Fail loudly. Missing references, empty text, incomplete checkpoints and invalid audio raise an error; they are not skipped or patched.
  • Use the sample's original source index for seeding (self._sample_seed(sample)), not its position in a filtered list.
  • Release owned resources in close, including worker processes.
  • Do not write absolute paths of your machine into tracked files; tests/test_public_paths.py rejects them.

A model may additionally get a direct backend in src/rvcbench/benchmark/backends.py (an adapter generate_sample method plus a backend class with its own timing scope). tests/test_direct_backends.py shows the seed and failure checks those backends must pass.

Before opening the pull request, run the model on a fixed subset (for example reproduction/subsets/libritts16_v1) and attach the rvcbench report output. Describe the result as subset validation; it does not establish paper-table reproduction.