# RVCBench documentation A plain-text export of the public English guides. --- Source: https://nanboy-ronan.github.io/RVCBench/docs/ # Using RVCBench [Read the documentation website](https://nanboy-ronan.github.io/RVCBench/docs/) · [Project homepage](https://nanboy-ronan.github.io/RVCBench/) · [PyPI](https://pypi.org/project/rvcbench/) Score generated speech, or evaluate a model on RVCBench data. | I want to… | Open | | --- | --- | | Install and run my first evaluation | [Quickstart](https://nanboy-ronan.github.io/RVCBench/docs/quickstart/) | | Look up a command argument, its values or its default | [CLI reference](https://nanboy-ronan.github.io/RVCBench/docs/cli/) | | Choose a paper scenario | [Scenarios and task IDs](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/) | | Score audio from my own dataset | [Python API arguments](https://nanboy-ronan.github.io/RVCBench/docs/api/) and [metric examples](https://nanboy-ronan.github.io/RVCBench/docs/metrics/) | | Connect my model or use a batch inference script | [Model evaluation](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/) | | Set up GPU scoring or fix installation | [Installation](https://nanboy-ronan.github.io/RVCBench/docs/installation/) | | Find dataset names and file formats | [Datasets](https://nanboy-ronan.github.io/RVCBench/docs/datasets/) | | Run a built-in model | [Models](https://nanboy-ronan.github.io/RVCBench/docs/models/) and [environments](https://nanboy-ronan.github.io/RVCBench/docs/model_environments/) | | Inspect or resume a run | [Run guide](https://nanboy-ronan.github.io/RVCBench/docs/run_protocol/) | | Reproduce the paper's results | [Frozen v1 codebase](https://nanboy-ronan.github.io/RVCBench/docs/versions/) | ## Run one scenario After [installation](https://nanboy-ronan.github.io/RVCBench/docs/installation/): ```bash rvcbench tasks --suite core-v1 rvcbench prompts --suite core-v1 --tasks chinese --output prompts/ # Generate every prompt with your model in outputs/my-model/. rvcbench score --suite core-v1 --tasks chinese \ --generated outputs/my-model --output results/my-model --device cpu ``` Read `results/my-model/submission.json` for per-task scores. Replace `chinese` with [task IDs](https://nanboy-ronan.github.io/RVCBench/docs/cli/#task-values), separated by spaces. ## Contribute See [CONTRIBUTING](https://github.com/Nanboy-Ronan/RVCBench/blob/main/CONTRIBUTING.md) for development checks and release validation, [model integration](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#route-2-built-in-integration) for adding an adapter, and [website maintenance](https://github.com/Nanboy-Ronan/RVCBench/blob/main/docs/site-src/README.md) for rebuilding the homepage. --- Source: https://nanboy-ronan.github.io/RVCBench/docs/quickstart/ # Quickstart | You have | Use | | --- | --- | | Generated audio and your own reference data | [Score your own audio](https://nanboy-ronan.github.io/RVCBench/docs/quickstart/#2a-score-your-own-audio) | | A model to evaluate on RVCBench data | [Export, generate, score](https://nanboy-ronan.github.io/RVCBench/docs/quickstart/#2b-evaluate-a-model-with-our-data) | | A parameter to look up | [CLI values and defaults](https://nanboy-ronan.github.io/RVCBench/docs/cli/) or [Python API arguments](https://nanboy-ronan.github.io/RVCBench/docs/api/) | ## 1. Install and prepare the scorers Linux, Python 3.10–3.13, FFmpeg. CPU installation: ```bash python -m venv .venv source .venv/bin/activate python -m pip install --upgrade pip python -m pip install torch==2.6.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cpu python -m pip install "rvcbench[eval]==2.2.1" rvcbench setup-scorers ``` Install FFmpeg with your system package manager, e.g. `sudo apt-get install ffmpeg` on Ubuntu. Scorer downloads need several GB of cache space. [GPU installation](https://nanboy-ronan.github.io/RVCBench/docs/installation/). ## 2A. Score your own audio Replace the file paths and text, then run this Python code: ```python from rvcbench import metrics with metrics.Evaluator(["sim", "wer", "speechmos"], device="cpu") as evaluator: scores = evaluator.score( "generated.wav", reference="speaker_reference.wav", text="Hello there.", language="en", ) print(scores) ``` | Argument | Values | | --- | --- | | `metrics` | `sim`, `sva`, `wer`, `speechmos`, `mcd`, `stoi`, `emotion`; pass one name, a list, or `"all"` | | `device` | `cpu`, `cuda`, `cuda:N` | | `language` | `en` (English), `zh` (Chinese), `fr` (French), or `None` for detection; other Whisper languages also work | `sim`: higher speaker similarity is better. `wer`: lower error rate is better. `speechmos`: higher predicted naturalness is better. MCD/STOI also need a same-text recording via `target=`. [All inputs and defaults](https://nanboy-ronan.github.io/RVCBench/docs/api/). ## 2B. Evaluate a model with our data ### Export prompts ```bash rvcbench tasks --suite core-v1 rvcbench prompts --suite core-v1 --tasks chinese --output prompts/ ``` | Argument | Values | | --- | --- | | `--suite` | `onboarding-v1` (52 prompts), `core-v1` (480), `full-v1` (12,724); counts are for whole suites | | `--tasks` | Space-separated task IDs, e.g. `chinese french background`; omit for all tasks. [Complete list](https://nanboy-ronan.github.io/RVCBench/docs/cli/#task-values) | | `--output` | New or empty directory for prompt lists and reference audio | The example exports 24 Chinese prompts. Some tasks automatically add clean controls; generate all exported prompts. [Task counts and paper scenarios](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/#tasks). ### Generate with your model Read `prompts/prompts.jsonl`, synthesize each entry, and save `outputs/my-model/.wav`. Replace `model.synthesize` below with your model's API: ```python import json from pathlib import Path import soundfile as sf # Load your voice cloning model here as `model`. # The synthesize call below is an integration example: replace it with your model's API. prompts = Path("prompts") outputs = Path("outputs/my-model") outputs.mkdir(parents=True, exist_ok=True) for line in (prompts / "prompts.jsonl").read_text().splitlines(): item = json.loads(line) waveform, sample_rate = model.synthesize( text=item["text"], reference_audio=str(prompts / item["reference_audio"]), reference_text=item["reference_text"], language=item["language"], ) sf.write(outputs / f"{item['id']}.wav", waveform, sample_rate) ``` Use the exported `id` unchanged. Write mono audio at the model's actual sample rate. Native batch scripts can use `prompts.tsv` or `prompts.lst`; see [batch formats](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#batch-inference-scripts). ### Score and inspect results ```bash rvcbench score --suite core-v1 --tasks chinese \ --generated outputs/my-model --output results/my-model --device cpu ``` Use the same `--suite` and `--tasks` as export. Set `--device cuda` or `--device cuda:1` to score on a GPU. | Read | Contains | | --- | --- | | `results/my-model/submission.json` | Per-task metrics, completion status and coverage | | `results/my-model/chinese/run_manifest.json` | Per-sample scores and errors | `complete`: all selected samples and metrics succeeded. `partial`: some failed; incomplete tasks have no mean score. Fix the failed outputs, then repeat the score command with `--resume`. ### Compare models and expand coverage Score several models with the same suite and tasks: ```bash rvcbench score --suite core-v1 --tasks chinese \ --generated outputs/model-a outputs/model-b --output results/comparison --device cpu ``` Read `results/comparison/comparison.md` (also `.csv` and `.json`). To compare existing results: ```bash rvcbench compare results/model-a results/model-b --output results/compare ``` To change tasks or suites, export new prompts and use new output directories. ## Next steps - [CLI reference](https://nanboy-ronan.github.io/RVCBench/docs/cli/): every argument, accepted value and default. - [Scenarios](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/): task IDs, counts, controls and metrics. - [Metrics API](https://nanboy-ronan.github.io/RVCBench/docs/api/): Python arguments and input requirements. --- Source: https://nanboy-ronan.github.io/RVCBench/docs/installation/ # Install RVCBench The pip package supports two workflows: automatic metrics for your own audio, and dataset-backed model evaluation. Neither requires cloning this repository. ## Recommended installation Use Python 3.10–3.13 on Linux. Start in a separate environment: ```bash python -m venv .venv source .venv/bin/activate python -m pip install --upgrade pip python -m pip install "rvcbench[eval]" ``` Install FFmpeg through your operating system if it is missing. On Ubuntu/Debian: ```bash sudo apt-get install ffmpeg ``` Download the scoring models once: ```bash rvcbench setup-scorers ``` This prepares all seven public metrics. Model downloads need several GB of disk space; Whisper medium is the largest download. For only speaker similarity, WER and MOS: ```bash rvcbench setup-scorers --metrics sim wer speechmos ``` - **Your own audio:** follow the [metrics API](https://nanboy-ronan.github.io/RVCBench/docs/metrics/). - **Our datasets:** follow [Evaluate your model](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/). - **Check the installation:** `rvcbench doctor --eval --imports` checks imports without model downloads. - **Check downloaded models:** `rvcbench setup-scorers --check-only` verifies and loads them without downloads. `pip install rvcbench` alone provides the runner, suite definitions and command line. Add `[eval]` to install the speech-scoring dependencies. Built-in model runtimes are separate; the two workflows above can score audio generated by any model, without installing that model inside RVCBench's environment. ## CPU and GPU CPU scoring works with `device="cpu"` in Python and `--device cpu` on the command line. To avoid installing CUDA libraries in a CPU environment, install CPU PyTorch first: ```bash python -m pip install torch==2.6.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cpu python -m pip install "rvcbench[eval]" ``` For GPU scoring, use matching PyTorch and torchaudio versions compatible with your NVIDIA driver, then install `rvcbench[eval]` and select `cuda` or `cuda:0`. The scoring extra currently supports PyTorch 2.3–2.9; see [model environments](https://nanboy-ronan.github.io/RVCBench/docs/model_environments/) for the validated scoring environments. ## Downloads and offline use Scorer assets are checked against pinned source revisions and SHA-256 hashes. Speaker, emotion and SpeechMOS assets use `RVCBENCH_ASSET_DIR` when set, otherwise a per-user cache under `~/.cache/rvcbench/` (existing checkout assets can also be used). Whisper uses `~/.cache/whisper/`. Set `XDG_CACHE_HOME` to relocate both default caches. ```bash export RVCBENCH_ASSET_DIR="$HOME/.cache/rvcbench" rvcbench setup-scorers rvcbench setup-scorers --check-only ``` Legacy SpeechMOS Torch Hub files are reused during setup only if their hashes match the release pins. A mismatched file is reported and preserved; move it aside and rerun setup to fetch the pinned version. Dataset audio is downloaded from Hugging Face when you run `rvcbench prompts` or `rvcbench score`. Only the suite's files are requested. To use an existing dataset copy, pass `--data-root` to both commands; see the [suite guide](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/). Log in with `hf auth login` if downloads are rate-limited. After preparing the assets and data, scoring can run offline. ## Upgrade from 2.0.0 ```bash python -m pip install --upgrade "rvcbench[eval]==2.2.1" rvcbench setup-scorers rvcbench doctor --eval --imports ``` Use a new results directory for the first 2.1.0 evaluation. Earlier reports remain readable, but scorer fingerprints have changed: rescore all models in the same environment before comparing them. `--resume` requires the request journal written by 2.1.0. Keep your existing generated WAV files; they can be scored again. ## Troubleshooting | Symptom | Action | | --- | --- | | Missing Whisper, SpeechBrain or another scorer dependency | Install `rvcbench[eval]` in the interpreter that runs the command. | | PyTorch/torchaudio binary import error | Install matching versions of torch and torchaudio from the same CPU/CUDA wheel index. | | NumPy/Numba conflict | Reinstall the evaluation extra in a fresh environment; it requires NumPy below 2.3. | | `pysptk` cannot build | See [building evaluation extras](https://nanboy-ronan.github.io/RVCBench/docs/model_environments/#building-the-evaluation-extras-from-source) for compiler prerequisites. | | Scorer asset missing | Run `rvcbench setup-scorers --metrics ` in the same cache environment. | | `ffmpeg` not found | Install the operating-system FFmpeg package and ensure it is on `PATH`. | | An evaluation stopped halfway | Repeat `rvcbench score` with the same arguments and `--resume`. | | Existing reports have incompatible scoring fingerprints | Rescore both models in one environment; use `compare --allow-incompatible` only for unranked inspection. | --- Source: https://nanboy-ronan.github.io/RVCBench/docs/metrics/ # Automatic speech metrics Use `rvcbench.metrics` to score your own voice-cloning outputs, with your own data. No benchmark suite or model adapter is required. Install `rvcbench[eval]` and prepare the scorers: ```bash python -m pip install "rvcbench[eval]" rvcbench setup-scorers ``` See [installation](https://nanboy-ronan.github.io/RVCBench/docs/installation/) for FFmpeg, CPU/GPU environments and downloads. ## Score one file ```python from rvcbench import metrics with metrics.Evaluator(["sim", "wer", "speechmos"], device="cpu") as evaluator: scores = evaluator.score( "generated.wav", reference="speaker_reference.wav", text="Hello there.", language="en", ) print(scores) # {'sim': ..., 'wer': ..., 'speechmos': ...} ``` Choose `device="cuda"` to use a GPU. `metrics.available()` lists the seven supported names. ## Score every metric ```python with metrics.Evaluator("all", device="cpu") as evaluator: scores = evaluator.score( "generated.wav", reference="speaker_reference.wav", target="same_text_recording.wav", text="Hello there.", language="en", ) ``` `reference` is a recording of the intended speaker. `target` is a recording of the **same text** as the generated audio; it is used for MCD and STOI. These can be different recordings. When `target` is omitted, MCD/STOI use `reference` for backwards compatibility, so only omit it when that recording has the same text. If no same-text recording exists, select the other metrics rather than calculating MCD/STOI on unrelated words. Emotion agreement uses `reference`; choose a reference with the intended expression. ## Score a dataset Reuse one evaluator to avoid reloading models for each file. For example, create `audio.jsonl`: ```json {"id": "sample-1", "generated": "outputs/1.wav", "reference": "refs/1.wav", "text": "Hello there.", "language": "en"} {"id": "sample-2", "generated": "outputs/2.wav", "reference": "refs/2.wav", "text": "Good morning.", "language": "en"} ``` Then run: ```python import json from rvcbench import metrics with metrics.Evaluator(["sim", "wer", "speechmos"], device="cpu") as evaluator: with open("audio.jsonl") as source, open("scores.jsonl", "w") as output: for line in source: item = json.loads(line) sample_id = item.pop("id") scores = evaluator.score(**item) output.write(json.dumps({"id": sample_id, **scores}) + "\n") ``` Paths in this example are relative to the directory where the script runs. Each scorer loads on first use and stays loaded until the context closes. Errors are raised with their cause; invalid scores are not silently included in an average. Scoring restores the caller's Python, NumPy, CPU and selected CUDA RNG states and cuDNN flags, including after an error. Run concurrent training/scoring in separate processes, as RNG state is temporarily changed during a call. ## One-line functions ```python metrics.speaker_similarity("generated.wav", "speaker_reference.wav") metrics.word_error_rate("generated.wav", "Hello there.", language="en") metrics.mos("generated.wav") metrics.mel_cepstral_distortion("generated.wav", "same_text_recording.wav") metrics.stoi("generated.wav", "same_text_recording.wav") ``` These functions load and release the required model per call. Use an `Evaluator` for repeated scoring. ## Metric definitions | Name | Definition | Compared with | | --- | --- | --- | | `sim` | Cosine similarity of ECAPA-TDNN speaker embeddings (SpeechBrain/VoxCeleb). Higher is better. | `reference` | | `sva` | Boolean speaker verification decision from the same model. | `reference` | | `wer` | Whisper medium transcription error rate, lowercased with ASCII punctuation removed; Chinese is segmented with jieba. Lower is better. | `text` | | `speechmos` | UTMOS22 strong predicted naturalness, nominally 1–5. Higher is better. | No reference | | `mcd` | DTW-aligned mel-cepstral distortion (pymcd). Lower is better. | `target`, or same-text `reference` | | `stoi` | Short-time objective intelligibility, resampled to 16 kHz and trimmed to the shorter recording. Higher is better. | `target`, or same-text `reference` | | `emotion` | Boolean equality of emotion labels predicted by SpeechBrain wav2vec2/IEMOCAP. | `reference` | `language` is passed to Whisper as a hint (`"en"`, `"zh"`, `"fr"`, etc.); without it Whisper detects the language. Automatic MOS is a model prediction, and emotion is agreement within the recognizer's four labels. These are the seven public API metrics; the research workflows contain additional specialized measurements. ## Relationship to the benchmark The API and `rvcbench score` use the same scorer implementations and pinned models. A suite supplies its target recordings and expected text automatically. Standalone calls let you choose your own references. Matching files, normalization and seeds give matching metric definitions; a suite seeds Whisper per sample, so fallback sampling can differ from the standalone API's default seed of 42. You can set `seed=` on an `Evaluator`. Floating-point differences between hardware and library versions can also occur. Suite reports additionally record scoring fingerprints and coverage. Use the [dataset workflow](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/) for comparable task reports, interruption recovery and multi-model tables. --- Source: https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/ # Evaluate your own model There are three ways to evaluate a voice cloning model with RVCBench. | Route | Use it when | You change | | --- | --- | --- | | [Score your own outputs](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#route-0-score-audio-you-generated-anywhere) | You can run the model yourself, or only through an API | Nothing; you generate WAV files | | [External adapter](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#route-1-external-adapter) | You want RVCBench to drive generation | One Python file of your own | | [Built-in integration](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#route-2-built-in-integration) | You want the model listed in this repository | The package, via a pull request | All routes use the same inputs, run records and metrics. ## Route 0: score audio you generated anywhere Your model runs in its own environment, with your own code. RVCBench only needs the resulting WAV files. ```bash python -m pip install 'rvcbench[eval]' rvcbench setup-scorers rvcbench prompts --suite core-v1 --output prompts/ ``` `prompts/` lists each utterance in three files; use the one your code already reads. | File | Line format | Read by | | --- | --- | --- | | `prompts.tsv` | `idreference_textreference_audiotext` | ZipVoice `--test-list` | | `prompts.lst` | `id\|reference_text\|reference_audio\|text` | Seed-TTS-eval scripts (F5-TTS, CosyVoice, ...) | | `prompts.jsonl` | one JSON object per line (below) | your own script | ```json {"id": "english-libritts__LibriTTS-4992-000033", "task": "english-libritts", "pair_id": "LibriTTS-4992-000033", "speaker_id": "4992", "reference_audio": "references/english-libritts/LibriTTS-4992-000033.wav", "reference_sha256": "3849cd34...", "reference_text": "I wish to prove my friendship to Miss Milner, ...", "text": "The Carey children had only found it by accident.", "language": "EN", "output_file": "english-libritts/LibriTTS-4992-000033.wav"} ``` Synthesize each `text` in the voice of `reference_audio` and save a mono WAV as `/.wav` (one flat directory, as batch scripts write it) or as `/`. Reference paths in `prompts.tsv` and `prompts.lst` are absolute, so the lists work from any directory; export again if you move `prompts/`. Then score: ```bash rvcbench score --suite core-v1 --generated my_outputs/ --output results/my-model/ --device cuda ``` The model name defaults to the directory name (`my_outputs` here); set it with `--model`. ### Batch inference scripts With ZipVoice, for example, the whole loop is its own batch command: ```bash python3 -m zipvoice.bin.infer_zipvoice --model-name zipvoice --test-list prompts/prompts.tsv --res-dir outputs/zipvoice rvcbench score --suite core-v1 --generated outputs/zipvoice --output results/zipvoice --device cuda ``` Scripts that read Seed-TTS-eval lists take `prompts/prompts.lst` and write `.wav`, which is the same layout. The lists carry no language column; `language` is in `prompts.jsonl` for models that need it. ### Several models at once ```bash rvcbench score --suite core-v1 --generated outputs/zipvoice outputs/zipvoice_distill --output results/ --device cuda ``` Each directory is scored as one model, named after the directory (or `--model a b`), into `results//`. Each metric model is loaded once for all of them. `results/comparison.md` shows every task and metric with one column per model, the best value in bold and each perturbed condition's change against its clean counterpart; `comparison.csv` and `comparison.json` hold the same data. To compare results scored at different times, for example a new checkpoint against earlier ones: ```bash rvcbench compare results/*/submission.json --output results/ ``` Comparison requires the same suite and matching per-task scoring fingerprints, including scorer code, assets, dependencies and settings. Rescore models together after an environment upgrade. `--allow-incompatible` explicitly produces an unranked inspection with the differences listed. ### Resume a stopped evaluation Repeat the original scoring command with `--resume`: ```bash rvcbench score --suite core-v1 --generated my_outputs/ --output results/my-model/ --device cuda --resume ``` The output directory must belong to the same suite, model name and generated-audio directory. Inputs are checked again. Successful scores with matching audio and scorer fingerprints are reused; missing outputs and failed scores are retried. After replacing a bad WAV, the changed sample is rescored. The same flag works with several model directories. Results from releases before the resume journal was introduced need a new output directory. Avoid simultaneous writers to the same results directory. ### Your own inference loop Read `prompts.jsonl` and call the inference function from your model's environment: ```python import json from pathlib import Path import soundfile as sf prompts = Path("prompts") outputs = Path("my_outputs") outputs.mkdir(exist_ok=True) for line in (prompts / "prompts.jsonl").read_text().splitlines(): item = json.loads(line) waveform, sample_rate = your_model.synthesize( # replace with your model's API text=item["text"], reference_audio=str(prompts / item["reference_audio"]), reference_text=item["reference_text"], language=item["language"], ) sf.write(outputs / f"{item['id']}.wav", waveform, sample_rate) ``` Then run `rvcbench score` in the scoring environment. A model's native batch script is equally valid. ### Results - `results/my-model/submission.json` holds the per-task metric means, coverage, the hash of every scored file and the suite version. Each task also gets a full run directory, so `rvcbench status` and `rvcbench report` work on `results/my-model//`. - A missing or invalid file fails that sample, and a task with a failed sample is `partial`. A file identical to the dataset's target recording is rejected. - The suite pins the dataset revision and the hash of every input. Both commands download only the files the suite uses, and refuse to run if any input differs from the frozen suite. Pass `--data-root` to use a local copy laid out like the Hub dataset. - Target recordings are never exported; they are read only while scoring. - `core-v1` has 480 utterances across the paper's robustness evaluations; see the [Core suite](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/). It is not yet a leaderboard suite. - `onboarding-v1` (52 utterances) is a smaller suite for checking the workflow. ## Route 1: external adapter ### 1. Install ```bash python -m pip install "rvcbench[eval]" rvcbench setup-scorers rvcbench doctor ``` Install into the environment that already runs your model. ### 2. Write the adapter Subclass `rvcbench.VoiceCloningAdapter` and implement `clone`: ```python # my_model_adapter.py from rvcbench import VoiceCloningAdapter class MyModelAdapter(VoiceCloningAdapter): def load(self): # Called once per run. self.config is the `adversary` block of the run config. self.model = load_my_model(self.config.checkpoint, device=self.device) def clone(self, *, text, reference_audio, reference_text, language): # Speak `text` in the voice of the audio file `reference_audio`. waveform = self.model.synthesize(text, prompt_wav=str(reference_audio), prompt_text=reference_text) return waveform, self.model.sample_rate # mono float waveform in [-1, 1], sample rate in Hz ``` The contract: - `clone` is called once per benchmark sample. `reference_audio` is a `pathlib.Path`; `reference_text` is its transcript (may be empty); `language` is the dataset's language tag or `None`. - Return `(waveform, sample_rate)`. The runner writes the WAV file, names it, and checks that it is nonempty and finite. - Raise an exception when an utterance cannot be generated. The runner records the failure for that sample and continues; it does not skip silently. - The runner seeds Python, NumPy and PyTorch before each call (run seed plus the sample's source index). Do not reseed inside `clone`. - `load` and `unload` are optional and run once per run. - The adapter must be importable from a `.py` file, because each run records a hash of the adapter's source and of the modules it imports from the same package. [`examples/echo_adapter.py`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/examples/echo_adapter.py) is a complete adapter that returns the reference audio; it is run by the test suite. ### 3. Run ```bash PYTHONPATH=. rvcbench run --config-name ots_vc/clean/libritts/custom_ots \ vc.adapter=my_model_adapter:MyModelAdapter \ vc.model=my_model run_name=my_model_on_libritts \ +adversary.checkpoint=/path/to/checkpoint.pt \ adversary.seed=42 adversary.max_samples=20 ``` - `vc.adapter` is `package.module:ClassName`. The module has to be importable: install your package, or put its directory on `PYTHONPATH` as above. - `vc.model` is the name recorded in the run and shown in reports. It must differ from the built-in model names (for example `xtts`), which select model-specific backends and checks. - `custom_ots` is a template config. Keys under `adversary` reach your adapter as `self.config`; prefix a key with `+` when the template does not define it. - Add `+vc.generate_only=true` to generate without scoring. - The first run downloads the selected dataset from the Hugging Face Hub. To use a local copy, pass `dataset.use_hf_dataset=false dataset.root_path=/path/to/dataset`. - Other datasets: copy the template next to the other configs of that dataset, or keep your own configs in a directory and pass `--config-dir /path/to/your/configs`. ### 4. Read the results Each run writes `results///`: | File | Content | | --- | --- | | `run_manifest.json` | Per-sample status, input and output hashes, seeds, resolved config, environment and source provenance | | `generated_audio/` | One WAV per sample | | `metrics.json` | Aggregate metrics with coverage over all requested samples | ```bash rvcbench status results/my_model_on_libritts/ rvcbench report results/my_model_on_libritts/ --output my_model_report.json ``` `rvcbench report` refuses runs with missing outputs or missing required metrics. See the [run guide](https://nanboy-ronan.github.io/RVCBench/docs/run_protocol/) for resume, retries, evaluation-only scoring and the conditions under which two runs may be compared. ## Route 2: built-in integration A pull request that adds a model to this repository contains: | Piece | Location | Notes | | --- | --- | --- | | Adapter | `src/rvcbench/adversary/_ots.py` | Subclass `BaseAdversary`. Keep model imports inside methods so that importing the module needs no model dependency. | | Model wrapper | `src/rvcbench/models//` | Loading and inference, with explicit checkpoint paths or pinned Hub revisions. | | Registry entry | `src/rvcbench/benchmark/registry.py` | `"": "rvcbench.adversary._ots:"` | | Configs | `src/rvcbench/configs/ots_vc/clean//_ots.yaml` | Start from an existing config of the same dataset. | | Environment | `envs/.yml` | Models have incompatible dependency stacks; each gets its own environment. | | Catalog entry | `src/rvcbench/benchmark/model_catalog.json` | Use `experimental_adapter` until a subset validation is recorded. | | Tests | `tests/test__runtime.py` | Use fake modules or stub processes. Tests must not download or load a model. | | Documentation | `docs/validation.md`, `README.md` | Add a validation row; state what was and was not run. | Requirements for the adapter: - Fail loudly. Missing references, empty text, incomplete checkpoints and invalid audio raise an error; they are not skipped or patched. - Use the sample's original source index for seeding (`self._sample_seed(sample)`), not its position in a filtered list. - Release owned resources in `close`, including worker processes. - Do not write absolute paths of your machine into tracked files; `tests/test_public_paths.py` rejects them. A model may additionally get a direct backend in `src/rvcbench/benchmark/backends.py` (an adapter `generate_sample` method plus a backend class with its own timing scope). `tests/test_direct_backends.py` shows the seed and failure checks those backends must pass. Before opening the pull request, run the model on a fixed subset (for example `reproduction/subsets/libritts16_v1`) and attach the `rvcbench report` output. Describe the result as subset validation; it does not establish paper-table reproduction. --- Source: https://nanboy-ronan.github.io/RVCBench/docs/core_suite/ # Scenarios and tasks ## Run specific scenarios ```bash # List available task IDs. rvcbench tasks --suite core-v1 # Export Chinese voice-cloning inputs. rvcbench prompts --suite core-v1 --tasks chinese --output prompts-chinese/ # Generate one .wav per prompt with your model in outputs/chinese/. # Score the generated audio. rvcbench score --suite core-v1 --tasks chinese \ --generated outputs/chinese --output results/chinese --device cuda ``` Read `results/chinese/submission.json` for per-task scores. | Argument | What to pass | | --- | --- | | `--suite` | `onboarding-v1`, `core-v1` or `full-v1` | | `--tasks` | IDs from the table below, separated by spaces; e.g. `chinese french background` | | `--output` | A new output directory | | `--generated` | Directory containing your model's WAV files | | `--device` | `cpu`, `cuda` or `cuda:N`, e.g. `cuda:1`; default `cpu` | Use the same `--suite` and `--tasks` for export and scoring. Omit `--tasks` to run the whole suite. For all arguments, see the [CLI reference](https://nanboy-ronan.github.io/RVCBench/docs/cli/). ## Choose a suite | `--suite` | Audio files to generate | Use | | --- | ---: | --- | | `onboarding-v1` | 52 | Quick integration check: `libritts` (16), `vctk` (16), `robotcall` (20) | | `core-v1` | 480 | Small evaluation across the scenarios below | | `full-v1` | 12,724 | Larger evaluation; excludes AdvNoise and AntiProtect | Counts above are for a whole suite. Selecting tasks reduces the workload. These suites are evaluation previews. To reproduce the paper's reported results, use [v1](https://nanboy-ronan.github.io/RVCBench/docs/versions/). ## Tasks Each row is one accepted `--tasks` value for `core-v1`. The Full column shows availability in `full-v1`. Counts are WAV files **you generate**, excluding automatically added tasks. `Auto` means RVCBench creates that task's audio during scoring; ffmpeg is required. | `--tasks` value | Paper scenario | Core | Full | Automatically added | | --- | --- | ---: | ---: | --- | | `audioshift` | AudioShift: accent, gender and age | 24 | 2000 | — | | `textshift-standard` | TextShift: standard-text control | 24 | 200 | — | | `textshift-hallucination` | TextShift: hallucination prompts | 24 | 200 | `textshift-standard` | | `textshift-scam` | TextShift / Expression: scam scripts | 20 | 200 | `textshift-scam-standard` | | `textshift-scam-standard` | Expression: normal-text control | 20 | 100 | — | | `english-libritts` | Multilingual: English | 24 | 2000 | — | | `chinese` | Multilingual: Chinese | 24 | 1998 | — | | `crosslingual` | Multilingual: English ↔ Chinese | 24 | 650 | — | | `french` | Multilingual: French | 24 | 2000 | — | | `longtext` | LongContext: long target text | 20 | 20 | — | | `longaudio` | LongContext: long reference audio | 24 | 156 | — | | `background-clean` | PassiveNoise: clean background control | 20 | 800 | — | | `background` | PassiveNoise: background noise | 20 | 800 | `background-clean` | | `multispeaker-clean` | PassiveNoise: single-speaker control | 24 | 800 | — | | `multispeaker` | PassiveNoise: competing speakers | 24 | 800 | `multispeaker-clean` | | `adv-clean` | AdvNoise / AntiProtect: clean control | 20 | Unavailable | — | | `adv-gaussian` | AdvNoise: Gaussian noise | 20 | Unavailable | `adv-clean` | | `adv-spec` | AdvNoise: SPEC | 20 | Unavailable | `adv-clean` | | `adv-safespeech` | AdvNoise: SafeSpeech | 20 | Unavailable | `adv-clean` | | `adv-pop` | AdvNoise: POP | 20 | Unavailable | `adv-clean` | | `adv-enkidu` | AdvNoise: Enkidu | 20 | Unavailable | `adv-clean` | | `antiprotect-spec` | AntiProtect: DEMUCS on SPEC | 20 | Unavailable | `adv-clean` | | `compression-mp3-64k` | Compression: MP3 64 kbps | Auto | Auto | `audioshift` | | `compression-aac-64k` | Compression: AAC 64 kbps | Auto | Auto | `audioshift` | | `compression-opus-24k` | Compression: OPUS 24 kbps | Auto | Auto | `audioshift` | | `compression-mp3-32k` | Compression: MP3 32 kbps | Auto | Auto | `audioshift` | | `compression-aac-32k` | Compression: AAC 32 kbps | Auto | Auto | `audioshift` | | `compression-opus-16k` | Compression: OPUS 16 kbps | Auto | Auto | `audioshift` | | `compression-narrowband` | Compression: telephone band | Auto | Auto | `audioshift` | Examples: ```bash # Chinese and French: 24 + 24 = 48 generations. rvcbench prompts --suite core-v1 --tasks chinese french --output prompts-languages/ # Background noise: 20 noisy + 20 clean = 40 generations. rvcbench prompts --suite core-v1 --tasks background --output prompts-noise/ # MP3 compression: generate 24 audioshift prompts; scoring creates the MP3 versions. rvcbench prompts --suite core-v1 --tasks compression-mp3-64k --output prompts-mp3/ ``` Generate every exported prompt, including automatically added controls. Selecting `audioshift` alone does not run compression tasks. `onboarding-v1` accepts only `libritts`, `vctk`, and `robotcall`. ## Metrics and results | Tasks | Reported metrics | | --- | --- | | Most tasks | `sim`, `speechmos`, `wer`, `mcd` | | `textshift-scam`, `textshift-scam-standard` | `sim`, `speechmos`, `wer`, `sva`, `emotion` | | `compression-*` | `stoi`, `mcd`, `sim`, `wer` | Metric definitions: [SIM, SVA, WER, MOS, MCD, STOI and emotion](https://nanboy-ronan.github.io/RVCBench/docs/metrics/#metric-definitions). The suite chooses the metrics; `score` has no `--metrics` argument. | Result | Where to find it | | --- | --- | | Per-task means and completion status | `submission.json` → `tasks` → task ID | | Change from the clean control | Task's `relative_change_percent` | | Accent, gender and age breakdown | `audioshift` → `group_means` | | Other group breakdowns | `english-libritts`: gender; `crosslingual`: direction; `longaudio`: reference duration; `background`: noise; `multispeaker`: SNR/interferer; `textshift-scam`: scam type | | Per-sample scores, errors and coverage | `/run_manifest.json` | | 95% bootstrap intervals | Metric reports in the task directory | `complete` means every selected task and automatically added dependency succeeded. Incomplete tasks have no `means`. Resume with the original command plus `--resume`. Compare models using the same suite, task selection and scoring setup. ## Coverage limits - Deepfake detection and the audio-LLM emotion-alignment judge are not included. - `core-v1` includes AdvNoise and AntiProtect; `full-v1` does not. - For shared tasks, core pairs are included in full with the same IDs and inputs. ## Full suite ```bash rvcbench prompts --suite full-v1 --output prompts-full/ # Generate all exported prompts in outputs/my-model/. rvcbench score --suite full-v1 --generated outputs/my-model \ --output results-full/my-model --device cuda ``` The full suite needs 12,724 generated WAVs and scores 26,724 files including compression variants. You can still select a smaller set with `--tasks`. To use a local dataset copy: ```bash hf download Nanboy/RVCBench --repo-type dataset --revision a932bd08d6858f14bdda52356dde5a7b771f0245 \ --local-dir rvcbench-data rvcbench prompts --suite full-v1 --tasks chinese --output prompts-local/ --data-root rvcbench-data ``` Pass the same `--data-root rvcbench-data` to `score`. Downloads require about 12.6 GB for the whole snapshot; use `hf auth login` if the Hub rate-limits requests. ## How it is built
Sampling, input verification and comparison details - Core selections are fixed before looking at model outputs, using seeded SHA-256 ranking (seed 20261002). Audio and metadata are checked against the suite's recorded hashes. - Noisy and protected tasks use matched clean controls. Compression tasks compare processed generated audio with its unprocessed version. - Selected runs record `selected_tasks`, `parent_suite_sha256` and `suite_sha256`. Different task selections cannot share a resumed run or a ranked comparison. Selecting all tasks keeps the whole-suite identity. - Protected references come from `Protected_LibriTTS/`. POP is called `em` in the protection code.
## Things to know when reading results
Task-specific interpretation - Enkidu and DEMUCS references are 16 kHz; clean LibriTTS references are 24 kHz. Their comparison includes the bandwidth difference. - Scam scripts have no same-text target recording. SIM and emotion compare with the speaker reference; MCD is omitted. WER checks whether the generated speech follows the script. - In full English-LibriTTS, 19 speakers have no gender annotation and appear as `unknown`. - Multi-line text is preserved in `prompts.jsonl`; TSV/LST exports replace line breaks with spaces. WER treats these line breaks as spaces.
--- Source: https://nanboy-ronan.github.io/RVCBench/docs/datasets/ # Datasets The benchmark data is on Hugging Face at [Nanboy/RVCBench](https://huggingface.co/datasets/Nanboy/RVCBench). Dataset configs download it automatically (`use_hf_dataset: true`, the default); `rvcbench prompts` and `rvcbench score` download only the files a suite uses. An offline snapshot is also available on [Google Drive](https://drive.google.com/file/d/1ZDOMorDGV8i5oVNtA5BaJLbFj2dVo5AU/view?usp=drive_link). How each source corpus was selected and preprocessed is described in [dataset preprocessing](https://nanboy-ronan.github.io/RVCBench/docs/dataset_preprocessing/). ## Folders and the paper's evaluations | Hub folder | Source | Used for (paper evaluation) | Dataset config | | --- | --- | --- | --- | | `Libritts` | LibriTTS test/dev-clean, 40 speakers | English-VC, LongAudio, AdvNoise and AntiProtect references | `libritts` | | `VCTK` | VCTK, 40 speakers, 12 accents | AudioShift (accent, gender, age), English-VC | `vctk` | | `vctk_text_robust` | VCTK voices, hallucination-style prompts | TextShift: Hallucination | `vctk_text_robust` | | `robotcall` | Robocall scam scripts, VCTK voices | TextShift: Scam, Expression | `robotcall` | | `AISHELL1_dev` | AISHELL-1, 40 speakers | Multilingual: Chinese-VC | `aishell` | | `Bilingual_uedin` | EMIME English-Mandarin, 13 speakers | Multilingual: CrossLingual | `bilingual_uedin` | | `CommonVoiceFR_dev` | Common Voice French, 40 speakers | Multilingual: French (appendix) | `french` | | `Long_context` | LibriSpeech-Long, 10 speakers | LongContext: LongText | `long_librispeech` | | `Background_noise` | VoiceBank+DEMAND, 10 noise types at 10 dB | PassiveNoise: Background | `background_noise`, `background_clean` | | `Multispeaker_libri` | LibriTTS mixed with two interferers at four SNRs | PassiveNoise: MultiSpeaker | `multispeaker_libri` | | `Protected_LibriTTS` | Protected references of the `core-v1` suite | AdvNoise, AntiProtect | used by `core-v1` | Dataset configs live in [`src/rvcbench/configs/dataset/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/configs/dataset/). To use a local copy, pass `dataset.use_hf_dataset=false dataset.root_path=/path/to/`. ## Format Each folder follows one layout: ```text / ├── audios//*.wav ├── filelists/ # legacy per-speaker JSON manifests, kept for compatibility └── metadata.parquet # canonical manifest read by all loaders ``` `metadata.parquet` has one row per evaluation pair: | Column | Description | | --- | --- | | `speaker_id` | Target speaker | | `prompt_file_name` | Reference (prompt) audio | | `prompt_text`, `prompt_language` | Transcript and language of the reference | | `target_file_name` | Ground-truth target audio | | `target_text`, `target_language` | Text to synthesize and its language | | `pair_id`, `dataset_name`, `split` | Provenance | Phoneme and alignment annotations (`prompt_phonemes`, `prompt_tone`, `prompt_word2ph` and their `target_*` counterparts) are kept when available, and dataset-specific fields such as `spam_type` in `robotcall` are extra columns. LibriTTS has two annotation exports (`speaker` and `speaker_text`); they stay distinct, see the [run guide](https://nanboy-ronan.github.io/RVCBench/docs/run_protocol/#libritts-manifest-variants). To rebuild canonical manifests from legacy per-speaker JSON files: ```bash python src/rvcbench/datasets/build_canonical_manifests.py --force ``` --- Source: https://nanboy-ronan.github.io/RVCBench/docs/api/ # Python API reference Import the metrics API: ```python from rvcbench import metrics ``` ## Available metrics ```python metrics.available() # ['sim', 'sva', 'wer', 'speechmos', 'mcd', 'stoi', 'emotion'] ``` See [metric definitions](https://nanboy-ronan.github.io/RVCBench/docs/metrics/#metric-definitions) for scoring directions and required inputs. ## Evaluator ```text metrics.Evaluator( metrics=("sim", "wer", "speechmos"), *, device="cpu", seed=42, logger=None, ) ``` | Argument | Accepted values | Default | | --- | --- | --- | | `metrics` | `"sim"`, `"sva"`, `"wer"`, `"speechmos"`, `"mcd"`, `"stoi"`, `"emotion"`; a list/tuple of these; or `"all"` | `("sim", "wer", "speechmos")` | | `device` | `"cpu"`, `"cuda"`, `"cuda:N"` (e.g. `"cuda:1"`), or a compatible `torch.device` | `"cpu"` | | `seed` | Integer | `42` | | `logger` | `logging.Logger` or `None` | `None` | Example: `metrics.Evaluator(["sim", "wer"], device="cuda")`. Unknown metric names and empty lists raise `ValueError`. Scorer models load on first use and are reused. Use the evaluator as a context manager so its resources are released, including when a call raises an error: ```python with metrics.Evaluator("speechmos", device="cpu") as evaluator: scores = evaluator.score("generated.wav") ``` ### score ```text evaluator.score(generated, *, reference=None, target=None, text=None, language=None) ``` | Argument | Accepted values | Required / default | | --- | --- | --- | | `generated` | Audio file path: `str` or `pathlib.Path`, e.g. `"generated.wav"` | **Required** | | `reference` | Speaker reference audio path: `str`, `Path` or `None` | Required for `sim`, `sva`, `emotion`; default `None` | | `target` | Same-text target audio path: `str`, `Path` or `None` | Used by `mcd` and `stoi`; falls back to `reference` | | `text` | Nonempty transcript string, e.g. `"Hello there."`, or `None` | Required for `wer`; default `None` | | `language` | Whisper language code/name; benchmark languages are `"en"`, `"zh"`, `"fr"`; `None` or `"auto"` for detection | `None` | `target` must contain the same words as `generated`. `reference` may contain different words. For MCD/STOI, provide either `target` or a same-text `reference`. ```python with metrics.Evaluator(["sim", "wer"], device="cpu") as evaluator: scores = evaluator.score( "generated.wav", reference="speaker.wav", text="Hello there.", language="en" ) ``` Returns a dictionary with exactly the selected metric names. Values are Python floats; `sva` and `emotion` are booleans. Missing required arguments raise `ValueError`; missing files raise `FileNotFoundError`. Scorer/dependency errors propagate to the caller, and invalid metric values raise an error rather than silently entering an average. ### close `evaluator.close()` releases the loaded scorers. The context manager calls it automatically. Python, NumPy and relevant PyTorch RNG states are restored after scoring; run concurrent training and scoring in separate processes because global RNG state is temporarily changed during a call. ## Convenience functions Each function scores one file and then releases its scorer. For repeated calls, prefer `Evaluator`. ```text metrics.speaker_similarity(generated, reference, *, device="cpu") metrics.word_error_rate(generated, text, *, language=None, device="cpu") metrics.mos(generated, *, device="cpu") metrics.mel_cepstral_distortion(generated, reference) metrics.stoi(generated, reference) ``` These are signature summaries: `*` marks keyword-only parameters. For MCD/STOI, the convenience function's `reference` must be a recording of the same text. Use `Evaluator` for SVA and emotion. ## Model adapters `from rvcbench import VoiceCloningAdapter` exposes the external adapter base class. Implement `clone(self, *, text, reference_audio, reference_text, language)` and return a mono waveform and sample rate. Optional `load()` and `unload()` methods control model lifetime. See the [adapter contract and full example](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#route-1-external-adapter). --- Source: https://nanboy-ronan.github.io/RVCBench/docs/cli/ # Command-line reference ```bash rvcbench --version rvcbench --help rvcbench score --help ``` Flags such as `--resume` take no value. Lists use spaces: `--tasks chinese french`, not commas. ## List scenarios ```bash rvcbench tasks --suite core-v1 rvcbench tasks --suite core-v1 --json ``` | Argument | Accepted values | Default | Result | | --- | --- | --- | --- | | `--suite` | `onboarding-v1`, `core-v1`, `full-v1`, or a custom suite `.json` path | `core-v1` | Tasks to list | | `--json` | Flag; no value | Off | Print JSON instead of text | ## Export evaluation inputs ```bash rvcbench prompts --suite core-v1 --tasks chinese --output prompts/ ``` | Argument | Accepted values | Default | Result | | --- | --- | --- | --- | | `--suite` | `onboarding-v1`, `core-v1`, `full-v1`, or a custom suite `.json` path | `onboarding-v1` | Evaluation suite | | `--tasks` | One or more IDs from the [task list below](https://nanboy-ronan.github.io/RVCBench/docs/cli/#task-values) | All tasks in the suite | Export only selected tasks and their dependencies | | `--output` | Directory path, e.g. `prompts/`; must be new or empty | **Required** | Prompt lists and reference WAV files | | `--data-root` | Dataset directory, e.g. `/data/RVCBench` containing `Libritts/`, `VCTK/`, etc. | Download from Hugging Face | Read local data | Output: `prompts.jsonl`, `prompts.tsv`, `prompts.lst`, reference audio, `suite.json` and `README.md`. Generate one WAV per prompt with your model. Save it as `.wav` or `/.wav`. See [inference examples](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#route-0-score-audio-you-generated-anywhere). ## Score generated outputs ```bash rvcbench score --suite core-v1 --tasks chinese \ --generated outputs/my-model --output results/my-model --device cuda ``` | Argument | Accepted values | Default | Result | | --- | --- | --- | --- | | `--suite` | `onboarding-v1`, `core-v1`, `full-v1`, or a custom suite `.json` path | `onboarding-v1` | Must match prompt export | | `--tasks` | One or more IDs from the [task list below](https://nanboy-ronan.github.io/RVCBench/docs/cli/#task-values) | All tasks in the suite | Must match prompt export | | `--generated` | One or more directories, e.g. `outputs/model-a outputs/model-b` | **Required** | Generated WAV files to score | | `--model` | One name per generated directory, e.g. `model-a model-b` | Directory names | Labels in reports; batch names must start with a letter/digit, contain only letters, digits, `_`, `-`, `.`, and exclude `__` | | `--output` | Directory path, e.g. `results/my-model` | **Required** | Write results here; use a new directory unless resuming | | `--device` | `cpu`, `cuda`, or `cuda:N` where N is a GPU index, e.g. `cuda:1` | `cpu` | Scoring device | | `--resume` | Flag; no value | Off | Reuse matching scores in an existing output directory | | `--data-root` | Dataset directory, same as prompt export | Download from Hugging Face | Read local data | One model: read `OUTPUT/submission.json`. Several models: read `OUTPUT/comparison.md`, `.csv` or `.json`; each model also gets `OUTPUT/MODEL/submission.json`. Scores are reported per task. Resume: repeat the same command with `--resume`. Keep the suite, tasks, model names and directories the same. ## Suite values | `--suite` | Audio files to generate per model | Use | | --- | ---: | --- | | `onboarding-v1` | 52 | Check your integration | | `core-v1` | 480 | Small evaluation across paper scenarios | | `full-v1` | 12,724 | Larger evaluation; excludes AdvNoise and AntiProtect | | `/path/to/suite.json` | Defined in the file | Custom evaluation; manifest paths are relative to that file | These counts are for whole suites. `--tasks` reduces the work to your selection plus its dependencies. ## Task values | Suite | Scenario | Accepted `--tasks` values | | --- | --- | --- | | onboarding-v1 | Onboarding | `libritts`, `vctk`, `robotcall` | | core-v1, full-v1 | AudioShift | `audioshift` | | core-v1, full-v1 | TextShift / Expression | `textshift-standard`, `textshift-hallucination`, `textshift-scam`, `textshift-scam-standard` | | core-v1, full-v1 | Multilingual | `english-libritts`, `chinese`, `french`, `crosslingual` | | core-v1, full-v1 | LongContext | `longtext`, `longaudio` | | core-v1, full-v1 | PassiveNoise | `background`, `background-clean`, `multispeaker`, `multispeaker-clean` | | core-v1 | AdvNoise / AntiProtect | `adv-clean`, `adv-gaussian`, `adv-spec`, `adv-safespeech`, `adv-pop`, `adv-enkidu`, `antiprotect-spec` | | core-v1, full-v1 | Compression | `compression-mp3-64k`, `compression-aac-64k`, `compression-opus-24k`, `compression-mp3-32k`, `compression-aac-32k`, `compression-opus-16k`, `compression-narrowband` | ```bash # Multiple tasks: separate IDs with spaces. rvcbench prompts --suite core-v1 --tasks chinese crosslingual background --output prompts/ ``` Omit `--tasks` to select all tasks. `--tasks all` is not supported. Clean controls and source tasks are added automatically: `background` adds `background-clean`; `compression-mp3-64k` adds `audioshift`. Generate every exported prompt. See [task meanings, sample counts and dependencies](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/#tasks). ## Check the environment ```bash rvcbench doctor --eval --imports rvcbench setup-scorers --metrics sim wer speechmos ``` ### doctor | Argument | Accepted values | Default | Result | | --- | --- | --- | --- | | `--model` | `qwen3`, `qwen3_omni`, `sparktts` | No model check | Include that model's dependencies | | `--eval` | Flag; no value | Off | Include scoring dependencies | | `--imports` | Flag; no value | Off | Import dependencies to detect broken installations | ### setup-scorers | Argument | Accepted values | Default | Result | | --- | --- | --- | --- | | `--metrics` | One or more of `sim`, `sva`, `wer`, `speechmos`, `mcd`, `stoi`, `emotion` | `sim speechmos wer mcd emotion stoi` | Download and verify these scorers | | `--check-only` | Flag; no value | Off | Check cached files and load models; download nothing | `sim` and `sva` share a model. `mcd` and `stoi` need no model files. Omit `--metrics` to set up all public scorers; `--metrics all` is not supported. This command prepares assets; it does not change the metrics used by `rvcbench score`, which are fixed by the suite. ## Compare models ```bash rvcbench compare results/model-a results/model-b --output results/compare ``` | Argument | Accepted values | Default | Result | | --- | --- | --- | --- | | `submissions` | One or more result directories or `submission.json` paths | **Required** | Reports to compare | | `--output` | Directory path, e.g. `results/compare` | Print only | Write `comparison.md`, `.csv` and `.json` | | `--allow-incompatible` | Flag; no value | Off | Show differing scorer protocols without ranking | All reports must use the same suite and task selection, even with `--allow-incompatible`. ## Inspect a task run ```bash rvcbench status results/my-model/chinese rvcbench report results/my-model/chinese --output chinese-report.json ``` | Command | Argument | Accepted values | Default | | --- | --- | --- | --- | | `status` | `run_dir` | Task directory containing `run_manifest.json` | **Required** | | `report` | `run_dir` | Task directory containing `run_manifest.json`; run must be complete | **Required** | | `report` | `--output` | JSON file path, e.g. `chinese-report.json` | **Required** | | `smoke` | `--output` | New directory path for a synthetic CPU check | `results/smoke` | ## Built-in models and research workflows `run`, `run-protected`, `protect` and `denoise` use Hydra `key=value` arguments: ```bash rvcbench run --config-name ots_vc/clean/libritts/qwen3_tts_ots --cfg job ``` | Argument | Accepted values | Result | | --- | --- | --- | | `--config-name` | Config path without `.yaml`, e.g. `ots_vc/clean/libritts/qwen3_tts_ots` | Choose a run configuration | | `--cfg` | `job`, `hydra`, `all` | Print the selected configuration without running | | `key=value` | A key and value from that configuration, e.g. `device=cuda:0` | Override a setting | | `+key=value` | A new configuration key, e.g. `+vc.generate_only=true` | Add a setting | Config names and model-specific settings are listed in [built-in models](https://nanboy-ronan.github.io/RVCBench/docs/models/) and [model setup](https://nanboy-ronan.github.io/RVCBench/docs/quickstart_model_setup/). `rvcbench run --help` lists available config groups. ### Audit and compare run records | Command | Argument | Accepted values | Default | | --- | --- | --- | --- | | `audit-source` | `run_dir` | Task run directory | **Required** | | `audit-source` | `--root` | Source checkout or installed package root | Current installation | | `audit-source` | `--output` | JSON file path | Print only | | `compare-check`, `compare-timing` | `left`, `right` | Two task run directories | **Required** | | `compare-check` | `--metrics` | Space-separated metric IDs: `sim`, `sva`, `wer`, `speechmos`, `mcd`, `stoi`, `emotion`; legacy `dnsmos` | `mcd wer sim` | | `compare-check`, `compare-timing` | `--output` | JSON file path | Print only | ### replay-gr Replay a recorded Gaussian-noise stage. [Input file formats](https://nanboy-ronan.github.io/RVCBench/docs/gr_noise_replay/). | Argument | Accepted values | Default | | --- | --- | --- | | `--dataset-root` | Local dataset directory | **Required** | | `--subset-manifest` | Frozen selection JSON file | **Required** | | `--noise-archive` | Recorded noise archive path | **Required** | | `--historical-directory` | Directory of historical WAV files to verify against | **Required** | | `--output` | Output directory | **Required** | | `--batch-size` | Integer | `8` | | `--sample-rate` | Integer, Hz | `24000` | | `--hop-length` | Integer, samples | `512` | | `--regenerate-rng` | Flag; no value | Off | | `--seed` | Integer; used with `--regenerate-rng` | `42` | | `--epsilon` | Float; perturbation magnitude | `0.03137255` | | `--device` | `cpu`, `cuda`, `cuda:N` | `cpu` | ### denoise-dns64 Denoise protected references. [Input file formats](https://nanboy-ronan.github.io/RVCBench/docs/dns64_stage/). | Argument | Accepted values | Default | | --- | --- | --- | | `--dataset-root` | Local dataset directory | **Required** | | `--subset-manifest` | Frozen selection JSON file | **Required** | | `--reference-directory` | Protected reference WAV directory | **Required** | | `--weights` | Local DNS64 weights file | **Required** | | `--output` | Output directory | **Required** | | `--device` | `cpu`, `cuda`, `cuda:N` | `cpu` | | `--dry` | Float from `0` to `1`; fraction of original audio to mix in | `0.0` | | `--dataset-rate` | Positive integer, Hz | `16000` | | `--runtime-python` | Python executable path | Current interpreter | | `--timeout-seconds` | Number of seconds for the external runtime | `600` | ### protect-enkidu Generate Enkidu-protected references. [Input file formats](https://nanboy-ronan.github.io/RVCBench/docs/enkidu_stage/). | Argument | Accepted values | Default | | --- | --- | --- | | `--dataset-root` | Local dataset directory | **Required** | | `--subset-manifest` | Frozen selection JSON file | **Required** | | `--model-directory` | Local Enkidu model directory | **Required** | | `--output` | Output directory | **Required** | | `--device` | `cpu`, `cuda`, `cuda:N` | `cpu` | | `--seed` | Integer | `42` | | `--epochs` | Positive integer | `10` | --- Source: https://nanboy-ronan.github.io/RVCBench/docs/faq/ # Voice cloning evaluation FAQ RVCBench is a Python package for comprehensive voice cloning evaluation. It combines automatic speech metrics with ready-to-use evaluation datasets. You can score your own audio or use benchmark prompts to compare models across languages, speakers and recording conditions. ## Can I use the metrics without downloading benchmark data? Yes. Install `rvcbench[eval]`, run `rvcbench setup-scorers`, and use `rvcbench.metrics.Evaluator` with your own generated audio, references and text. The [metrics guide](https://nanboy-ronan.github.io/RVCBench/docs/metrics/) includes single-file and batch examples. Metric models are downloaded once and reused locally. ## Which aspects of voice cloning does RVCBench measure? The public API exposes seven metrics across five aspects: | Aspect | Metrics | Inputs beyond generated audio | | --- | --- | --- | | Speaker identity | `sim`, `sva` | Recording of the intended speaker | | Content accuracy | `wer` | Expected text; optional language hint | | Predicted naturalness | `speechmos` (UTMOS) | None | | Acoustic fidelity and intelligibility | `mcd`, `stoi` | Recording of the same text | | Emotion consistency | `emotion` | Reference recording with the intended expression | These scores describe complementary properties. They are not a single universal quality score and do not replace human listening tests. See [metric definitions](https://nanboy-ronan.github.io/RVCBench/docs/metrics/#metric-definitions). ## How do I benchmark my model with the provided datasets? Use `rvcbench prompts` to download selected data and export reference recordings and text. Generate speech in your model's own environment or through its API, then run `rvcbench score` on the WAV files. RVCBench handles data preparation, scoring and reports; you supply model inference. See the [complete quickstart](https://nanboy-ronan.github.io/RVCBench/docs/quickstart/#2b-evaluate-a-model-with-our-data). ## Does a new model need a built-in adapter? No. Any model that can generate WAV files from the exported prompts can use the dataset scoring workflow. Save outputs using the exported IDs. JSONL, ZipVoice-style TSV and Seed-TTS-style lists are provided. Built-in adapters are an additional option, not the limit on supported model submissions. See [model integration](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/). ## Which benchmark suite should I start with? Use `onboarding-v1` with 52 outputs to check your integration. Use `core-v1` with 480 outputs for broader coverage, including protected references. Use `full-v1` with 12,724 outputs for larger datasets; protection tasks currently require `core-v1` separately. These are preview suite protocols, distinct from paper reproduction. See [suite coverage](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/). ## What if I do not have a recording of the same text? Use the metrics supported by your inputs. Speaker similarity only needs a recording of the intended speaker, WER needs expected text, and SpeechMOS needs only generated audio. MCD and STOI require a same-text recording; do not compare unrelated sentences. See [audio input requirements](https://nanboy-ronan.github.io/RVCBench/docs/api/#score). ## Can I score on CPU and resume an interrupted evaluation? Yes. CPU is the default device, and you can select CUDA for GPU scoring. Reuse an `Evaluator` for batch API scoring; for dataset scoring, repeat the same command with `--resume` to reuse matching successful scores and retry failed or changed samples. See [installation](https://nanboy-ronan.github.io/RVCBench/docs/installation/) and [resume behavior](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#resume-a-stopped-evaluation). ## How can I compare models fairly? Use the same suite, inputs and scoring environment for every model. RVCBench checks per-task scoring fingerprints and reports coverage and failures. `rvcbench compare` refuses incompatible results by default. Its explicit incompatibility override produces an unranked inspection. See the [comparison workflow](https://nanboy-ronan.github.io/RVCBench/docs/adding_a_model/#several-models-at-once). ## Does comprehensive evaluation mean every possible metric is included? It means the package covers the five aspects above through one API, with datasets spanning multiple languages, speakers and recording conditions. The paper additionally studies deepfake detectability and an audio-LLM expression judge; those two components are not yet in the public API or packaged suites. See [coverage](https://nanboy-ronan.github.io/RVCBench/docs/core_suite/) before selecting a protocol for your model. --- Source: https://nanboy-ronan.github.io/RVCBench/docs/versions/ # Codebase versions v1 and v2 are two versions of the code for the same benchmark: the datasets, metrics and paper results are the same. **To reproduce the paper, use v1.** To evaluate a new model, use v2. ```text tag v1.0 (2026-09-30): the code released with the paper ├── branch v1 v1.0 plus README changes only; frozen └── branch main v1.0 plus the v2 refactor (2026-09-30 to 2026-10-02); under development ``` | | v1 (branch [`v1`](https://github.com/Nanboy-Ronan/RVCBench/tree/v1), tag [`v1.0`](https://github.com/Nanboy-Ronan/RVCBench/tree/v1.0)) | v2 (branch [`main`](https://github.com/Nanboy-Ronan/RVCBench/tree/main)) | | --- | --- | --- | | Use it to | **Reproduce the paper** | Evaluate a new model | | Status | Frozen at the paper release | Under active development | | Install | Clone, then `pip install` a list of packages | `pip install` the `rvcbench` package from GitHub | | Run a built-in model | `python run_vc.py --config-name ...` | `rvcbench run --config-name ...` (same config names) | | Evaluate your own model | Add an adapter to the codebase | Score audio generated anywhere (`rvcbench prompts`, `rvcbench score`), or a one-file adapter | | Evaluation data | Full datasets | Full datasets, plus the `core-v1` suite: 480 pinned utterances covering 16 of the paper's 18 evaluations | | Run output | `metrics.json` per run | Per-sample `run_manifest.json` with input and output hashes, failures and metric coverage | ```bash git clone --branch v1 https://github.com/Nanboy-Ronan/RVCBench.git RVCBench-v1 # reproduce the paper git clone https://github.com/Nanboy-Ronan/RVCBench.git RVCBench # v2 ``` v1 and v2 name versions of this code. They are unrelated to the paper's arXiv versions and to suite names such as `core-v1`. ## What changed in v2 | Area | Main code | Change | | --- | --- | --- | | Package and commands | [`src/rvcbench/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/), [`benchmark/cli.py`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/benchmark/cli.py) | Installable `rvcbench` package with configs inside it; `rvcbench run`, `prompts`, `score`, `setup-scorers` and the run-record commands. | | Suites | [`benchmark/submission.py`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/benchmark/submission.py), [`suites/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/suites/) | Versioned suites with pinned inputs; score audio generated anywhere. | | Execution and reporting | [`benchmark/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/benchmark/), [`workflows/vc.py`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/workflows/vc.py) | Per-sample records of inputs, outputs, failures, metric coverage and run state; generation and scoring can run in separate environments. Completion and comparability are checked separately. | | Model adapters | [`adversary/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/adversary/), [`models/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/models/) | Explicit per-sample requests for 13 models, seeds from the original sample index, invalid conditioning or audio rejected, owned resources released. Other models still use the compatibility path. | | Model workers | [`models/worker_protocol.py`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/models/worker_protocol.py), [`scripts/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/scripts/) | Bounded waits, request/response matching and process cleanup. | | Checkpoints and dependencies | [`benchmark/model_assets.py`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/benchmark/model_assets.py), [Hub revisions](https://nanboy-ronan.github.io/RVCBench/docs/hub_revisions/) | Explicit paths or pinned revisions, recorded hashes, strict state loading for selected models. | | Data and reproduction | [`datasets/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/datasets/), [`reproduction/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/reproduction/) | Annotation variants kept in sample identity, validated source indices, frozen subsets with input hashes. | | Protection, scoring and timing | [`benchmark/`](https://github.com/Nanboy-Ronan/RVCBench/blob/main/src/rvcbench/benchmark/), [run guide](https://nanboy-ronan.github.io/RVCBench/docs/run_protocol/) | Traced reference stages and scorer provenance; timing scopes declared so incompatible runs are not compared. | **Behaviour changes.** Malformed inputs or incomplete checkpoints that v1 skipped or patched now fail. Model versions, seed policies and retry or conditioning variants must be recorded when comparing v1 and v2 results. v2 runs do not replace the paper's tables. To reproduce those, use v1 together with its environment files (`envs/`) and the checkpoint bundle linked from its README; the code alone does not pin model weights. ## Validation status Real generation and core scoring on a fixed subset have been recorded for 27 of the 32 integrations, including the paper's 18 models. These runs span different stages of the refactor and do not establish reproduction on the current combined source or of the paper's tables. The other five integrations lack assets or a verified cloning protocol. See [validation coverage](https://nanboy-ronan.github.io/RVCBench/docs/validation/) and the [reproduction plan](https://github.com/Nanboy-Ronan/RVCBench/blob/main/reproduction/plan.json).