Skip to content

Scenarios and tasks

Run specific scenarios

# List available task IDs.
rvcbench tasks --suite core-v1

# Export Chinese voice-cloning inputs.
rvcbench prompts --suite core-v1 --tasks chinese --output prompts-chinese/

# Generate one <id>.wav per prompt with your model in outputs/chinese/.

# Score the generated audio.
rvcbench score --suite core-v1 --tasks chinese \
  --generated outputs/chinese --output results/chinese --device cuda

Read results/chinese/submission.json for per-task scores.

Argument What to pass
--suite onboarding-v1, core-v1 or full-v1
--tasks IDs from the table below, separated by spaces; e.g. chinese french background
--output A new output directory
--generated Directory containing your model's WAV files
--device cpu, cuda or cuda:N, e.g. cuda:1; default cpu

Use the same --suite and --tasks for export and scoring. Omit --tasks to run the whole suite. For all arguments, see the CLI reference.

Choose a suite

--suite Audio files to generate Use
onboarding-v1 52 Quick integration check: libritts (16), vctk (16), robotcall (20)
core-v1 480 Small evaluation across the scenarios below
full-v1 12,724 Larger evaluation; excludes AdvNoise and AntiProtect

Counts above are for a whole suite. Selecting tasks reduces the workload. These suites are evaluation previews. To reproduce the paper's reported results, use v1.

Tasks

Each row is one accepted --tasks value for core-v1. The Full column shows availability in full-v1. Counts are WAV files you generate, excluding automatically added tasks. Auto means RVCBench creates that task's audio during scoring; ffmpeg is required.

--tasks value Paper scenario Core Full Automatically added
audioshift AudioShift: accent, gender and age 24 2000 —
textshift-standard TextShift: standard-text control 24 200 —
textshift-hallucination TextShift: hallucination prompts 24 200 textshift-standard
textshift-scam TextShift / Expression: scam scripts 20 200 textshift-scam-standard
textshift-scam-standard Expression: normal-text control 20 100 —
english-libritts Multilingual: English 24 2000 —
chinese Multilingual: Chinese 24 1998 —
crosslingual Multilingual: English ↔ Chinese 24 650 —
french Multilingual: French 24 2000 —
longtext LongContext: long target text 20 20 —
longaudio LongContext: long reference audio 24 156 —
background-clean PassiveNoise: clean background control 20 800 —
background PassiveNoise: background noise 20 800 background-clean
multispeaker-clean PassiveNoise: single-speaker control 24 800 —
multispeaker PassiveNoise: competing speakers 24 800 multispeaker-clean
adv-clean AdvNoise / AntiProtect: clean control 20 Unavailable —
adv-gaussian AdvNoise: Gaussian noise 20 Unavailable adv-clean
adv-spec AdvNoise: SPEC 20 Unavailable adv-clean
adv-safespeech AdvNoise: SafeSpeech 20 Unavailable adv-clean
adv-pop AdvNoise: POP 20 Unavailable adv-clean
adv-enkidu AdvNoise: Enkidu 20 Unavailable adv-clean
antiprotect-spec AntiProtect: DEMUCS on SPEC 20 Unavailable adv-clean
compression-mp3-64k Compression: MP3 64 kbps Auto Auto audioshift
compression-aac-64k Compression: AAC 64 kbps Auto Auto audioshift
compression-opus-24k Compression: OPUS 24 kbps Auto Auto audioshift
compression-mp3-32k Compression: MP3 32 kbps Auto Auto audioshift
compression-aac-32k Compression: AAC 32 kbps Auto Auto audioshift
compression-opus-16k Compression: OPUS 16 kbps Auto Auto audioshift
compression-narrowband Compression: telephone band Auto Auto audioshift

Examples:

# Chinese and French: 24 + 24 = 48 generations.
rvcbench prompts --suite core-v1 --tasks chinese french --output prompts-languages/

# Background noise: 20 noisy + 20 clean = 40 generations.
rvcbench prompts --suite core-v1 --tasks background --output prompts-noise/

# MP3 compression: generate 24 audioshift prompts; scoring creates the MP3 versions.
rvcbench prompts --suite core-v1 --tasks compression-mp3-64k --output prompts-mp3/

Generate every exported prompt, including automatically added controls. Selecting audioshift alone does not run compression tasks. onboarding-v1 accepts only libritts, vctk, and robotcall.

Metrics and results

Tasks Reported metrics
Most tasks sim, speechmos, wer, mcd
textshift-scam, textshift-scam-standard sim, speechmos, wer, sva, emotion
compression-* stoi, mcd, sim, wer

Metric definitions: SIM, SVA, WER, MOS, MCD, STOI and emotion. The suite chooses the metrics; score has no --metrics argument.

Result Where to find it
Per-task means and completion status submission.json → tasks → task ID
Change from the clean control Task's relative_change_percent
Accent, gender and age breakdown audioshift → group_means
Other group breakdowns english-libritts: gender; crosslingual: direction; longaudio: reference duration; background: noise; multispeaker: SNR/interferer; textshift-scam: scam type
Per-sample scores, errors and coverage <task>/run_manifest.json
95% bootstrap intervals Metric reports in the task directory

complete means every selected task and automatically added dependency succeeded. Incomplete tasks have no means. Resume with the original command plus --resume. Compare models using the same suite, task selection and scoring setup.

Coverage limits

  • Deepfake detection and the audio-LLM emotion-alignment judge are not included.
  • core-v1 includes AdvNoise and AntiProtect; full-v1 does not.
  • For shared tasks, core pairs are included in full with the same IDs and inputs.

Full suite

rvcbench prompts --suite full-v1 --output prompts-full/
# Generate all exported prompts in outputs/my-model/.
rvcbench score --suite full-v1 --generated outputs/my-model \
  --output results-full/my-model --device cuda

The full suite needs 12,724 generated WAVs and scores 26,724 files including compression variants. You can still select a smaller set with --tasks.

To use a local dataset copy:

hf download Nanboy/RVCBench --repo-type dataset --revision a932bd08d6858f14bdda52356dde5a7b771f0245 \
  --local-dir rvcbench-data
rvcbench prompts --suite full-v1 --tasks chinese --output prompts-local/ --data-root rvcbench-data

Pass the same --data-root rvcbench-data to score. Downloads require about 12.6 GB for the whole snapshot; use hf auth login if the Hub rate-limits requests.

How it is built

Sampling, input verification and comparison details
  • Core selections are fixed before looking at model outputs, using seeded SHA-256 ranking (seed 20261002). Audio and metadata are checked against the suite's recorded hashes.
  • Noisy and protected tasks use matched clean controls. Compression tasks compare processed generated audio with its unprocessed version.
  • Selected runs record selected_tasks, parent_suite_sha256 and suite_sha256. Different task selections cannot share a resumed run or a ranked comparison. Selecting all tasks keeps the whole-suite identity.
  • Protected references come from Protected_LibriTTS/. POP is called em in the protection code.

Things to know when reading results

Task-specific interpretation
  • Enkidu and DEMUCS references are 16 kHz; clean LibriTTS references are 24 kHz. Their comparison includes the bandwidth difference.
  • Scam scripts have no same-text target recording. SIM and emotion compare with the speaker reference; MCD is omitted. WER checks whether the generated speech follows the script.
  • In full English-LibriTTS, 19 speakers have no gender annotation and appear as unknown.
  • Multi-line text is preserved in prompts.jsonl; TSV/LST exports replace line breaks with spaces. WER treats these line breaks as spaces.