Evaluate your own model¶
There are three ways to evaluate a voice cloning model with RVCBench.
| Route | Use it when | You change |
|---|---|---|
| Score your own outputs | You can run the model yourself, or only through an API | Nothing; you generate WAV files |
| External adapter | You want RVCBench to drive generation | One Python file of your own |
| Built-in integration | You want the model listed in this repository | The package, via a pull request |
All routes use the same inputs, run records and metrics.
Route 0: score audio you generated anywhere¶
Your model runs in its own environment, with your own code. RVCBench only needs the resulting WAV files.
python -m pip install 'rvcbench[eval]'
rvcbench setup-scorers
rvcbench prompts --suite core-v1 --output prompts/
prompts/ lists each utterance in three files; use the one your code already reads.
| File | Line format | Read by |
|---|---|---|
prompts.tsv |
id<TAB>reference_text<TAB>reference_audio<TAB>text |
ZipVoice --test-list |
prompts.lst |
id\|reference_text\|reference_audio\|text |
Seed-TTS-eval scripts (F5-TTS, CosyVoice, ...) |
prompts.jsonl |
one JSON object per line (below) | your own script |
{"id": "english-libritts__LibriTTS-4992-000033", "task": "english-libritts", "pair_id": "LibriTTS-4992-000033",
"speaker_id": "4992", "reference_audio": "references/english-libritts/LibriTTS-4992-000033.wav",
"reference_sha256": "3849cd34...", "reference_text": "I wish to prove my friendship to Miss Milner, ...",
"text": "The Carey children had only found it by accident.", "language": "EN",
"output_file": "english-libritts/LibriTTS-4992-000033.wav"}
Synthesize each text in the voice of reference_audio and save a mono WAV as <your directory>/<id>.wav
(one flat directory, as batch scripts write it) or as <your directory>/<output_file>. Reference paths in
prompts.tsv and prompts.lst are absolute, so the lists work from any directory; export again if you
move prompts/. Then score:
The model name defaults to the directory name (my_outputs here); set it with --model.
Batch inference scripts¶
With ZipVoice, for example, the whole loop is its own batch command:
python3 -m zipvoice.bin.infer_zipvoice --model-name zipvoice --test-list prompts/prompts.tsv --res-dir outputs/zipvoice
rvcbench score --suite core-v1 --generated outputs/zipvoice --output results/zipvoice --device cuda
Scripts that read Seed-TTS-eval lists take prompts/prompts.lst and write <utt>.wav, which is the same
layout. The lists carry no language column; language is in prompts.jsonl for models that need it.
Several models at once¶
rvcbench score --suite core-v1 --generated outputs/zipvoice outputs/zipvoice_distill --output results/ --device cuda
Each directory is scored as one model, named after the directory (or --model a b), into
results/<model>/. Each metric model is loaded once for all of them. results/comparison.md shows every
task and metric with one column per model, the best value in bold and each perturbed condition's change
against its clean counterpart; comparison.csv and comparison.json hold the same data.
To compare results scored at different times, for example a new checkpoint against earlier ones:
Comparison requires the same suite and matching per-task scoring fingerprints, including scorer code,
assets, dependencies and settings. Rescore models together after an environment upgrade.
--allow-incompatible explicitly produces an unranked inspection with the differences listed.
Resume a stopped evaluation¶
Repeat the original scoring command with --resume:
rvcbench score --suite core-v1 --generated my_outputs/ --output results/my-model/ --device cuda --resume
The output directory must belong to the same suite, model name and generated-audio directory. Inputs are checked again. Successful scores with matching audio and scorer fingerprints are reused; missing outputs and failed scores are retried. After replacing a bad WAV, the changed sample is rescored. The same flag works with several model directories. Results from releases before the resume journal was introduced need a new output directory. Avoid simultaneous writers to the same results directory.
Your own inference loop¶
Read prompts.jsonl and call the inference function from your model's environment:
import json
from pathlib import Path
import soundfile as sf
prompts = Path("prompts")
outputs = Path("my_outputs")
outputs.mkdir(exist_ok=True)
for line in (prompts / "prompts.jsonl").read_text().splitlines():
item = json.loads(line)
waveform, sample_rate = your_model.synthesize( # replace with your model's API
text=item["text"],
reference_audio=str(prompts / item["reference_audio"]),
reference_text=item["reference_text"],
language=item["language"],
)
sf.write(outputs / f"{item['id']}.wav", waveform, sample_rate)
Then run rvcbench score in the scoring environment. A model's native batch script is equally valid.
Results¶
results/my-model/submission.jsonholds the per-task metric means, coverage, the hash of every scored file and the suite version. Each task also gets a full run directory, sorvcbench statusandrvcbench reportwork onresults/my-model/<task>/.- A missing or invalid file fails that sample, and a task with a failed sample is
partial. A file identical to the dataset's target recording is rejected. - The suite pins the dataset revision and the hash of every input. Both commands download only the files the suite uses, and refuse to run if any input differs from the frozen suite. Pass
--data-rootto use a local copy laid out like the Hub dataset. - Target recordings are never exported; they are read only while scoring.
core-v1has 480 utterances across the paper's robustness evaluations; see the Core suite. It is not yet a leaderboard suite.onboarding-v1(52 utterances) is a smaller suite for checking the workflow.
Route 1: external adapter¶
1. Install¶
Install into the environment that already runs your model.
2. Write the adapter¶
Subclass rvcbench.VoiceCloningAdapter and implement clone:
# my_model_adapter.py
from rvcbench import VoiceCloningAdapter
class MyModelAdapter(VoiceCloningAdapter):
def load(self):
# Called once per run. self.config is the `adversary` block of the run config.
self.model = load_my_model(self.config.checkpoint, device=self.device)
def clone(self, *, text, reference_audio, reference_text, language):
# Speak `text` in the voice of the audio file `reference_audio`.
waveform = self.model.synthesize(text, prompt_wav=str(reference_audio), prompt_text=reference_text)
return waveform, self.model.sample_rate # mono float waveform in [-1, 1], sample rate in Hz
The contract:
cloneis called once per benchmark sample.reference_audiois apathlib.Path;reference_textis its transcript (may be empty);languageis the dataset's language tag orNone.- Return
(waveform, sample_rate). The runner writes the WAV file, names it, and checks that it is nonempty and finite. - Raise an exception when an utterance cannot be generated. The runner records the failure for that sample and continues; it does not skip silently.
- The runner seeds Python, NumPy and PyTorch before each call (run seed plus the sample's source index). Do not reseed inside
clone. loadandunloadare optional and run once per run.- The adapter must be importable from a
.pyfile, because each run records a hash of the adapter's source and of the modules it imports from the same package.
examples/echo_adapter.py is a complete adapter that returns the reference audio; it is run by the test suite.
3. Run¶
PYTHONPATH=. rvcbench run --config-name ots_vc/clean/libritts/custom_ots \
vc.adapter=my_model_adapter:MyModelAdapter \
vc.model=my_model run_name=my_model_on_libritts \
+adversary.checkpoint=/path/to/checkpoint.pt \
adversary.seed=42 adversary.max_samples=20
vc.adapterispackage.module:ClassName. The module has to be importable: install your package, or put its directory onPYTHONPATHas above.vc.modelis the name recorded in the run and shown in reports. It must differ from the built-in model names (for examplextts), which select model-specific backends and checks.custom_otsis a template config. Keys underadversaryreach your adapter asself.config; prefix a key with+when the template does not define it.- Add
+vc.generate_only=trueto generate without scoring. - The first run downloads the selected dataset from the Hugging Face Hub. To use a local copy, pass
dataset.use_hf_dataset=false dataset.root_path=/path/to/dataset. - Other datasets: copy the template next to the other configs of that dataset, or keep your own configs in a directory and pass
--config-dir /path/to/your/configs.
4. Read the results¶
Each run writes results/<run_name>/<timestamp>/:
| File | Content |
|---|---|
run_manifest.json |
Per-sample status, input and output hashes, seeds, resolved config, environment and source provenance |
generated_audio/ |
One WAV per sample |
metrics.json |
Aggregate metrics with coverage over all requested samples |
rvcbench status results/my_model_on_libritts/<timestamp>
rvcbench report results/my_model_on_libritts/<timestamp> --output my_model_report.json
rvcbench report refuses runs with missing outputs or missing required metrics. See the run guide for resume, retries, evaluation-only scoring and the conditions under which two runs may be compared.
Route 2: built-in integration¶
A pull request that adds a model to this repository contains:
| Piece | Location | Notes |
|---|---|---|
| Adapter | src/rvcbench/adversary/<model>_ots.py |
Subclass BaseAdversary. Keep model imports inside methods so that importing the module needs no model dependency. |
| Model wrapper | src/rvcbench/models/<model>/ |
Loading and inference, with explicit checkpoint paths or pinned Hub revisions. |
| Registry entry | src/rvcbench/benchmark/registry.py |
"<model>": "rvcbench.adversary.<model>_ots:<ClassName>" |
| Configs | src/rvcbench/configs/ots_vc/clean/<dataset>/<model>_ots.yaml |
Start from an existing config of the same dataset. |
| Environment | envs/<model>.yml |
Models have incompatible dependency stacks; each gets its own environment. |
| Catalog entry | src/rvcbench/benchmark/model_catalog.json |
Use experimental_adapter until a subset validation is recorded. |
| Tests | tests/test_<model>_runtime.py |
Use fake modules or stub processes. Tests must not download or load a model. |
| Documentation | docs/validation.md, README.md |
Add a validation row; state what was and was not run. |
Requirements for the adapter:
- Fail loudly. Missing references, empty text, incomplete checkpoints and invalid audio raise an error; they are not skipped or patched.
- Use the sample's original source index for seeding (
self._sample_seed(sample)), not its position in a filtered list. - Release owned resources in
close, including worker processes. - Do not write absolute paths of your machine into tracked files;
tests/test_public_paths.pyrejects them.
A model may additionally get a direct backend in src/rvcbench/benchmark/backends.py (an adapter generate_sample method plus a backend class with its own timing scope). tests/test_direct_backends.py shows the seed and failure checks those backends must pass.
Before opening the pull request, run the model on a fixed subset (for example reproduction/subsets/libritts16_v1) and attach the rvcbench report output. Describe the result as subset validation; it does not establish paper-table reproduction.