Evaluate voice cloning.
Metrics and datasets in one package.
RVCBench brings automatic speech metrics and ready-to-use evaluation datasets together. Score your own audio, or evaluate a model across speaker identity, speech quality, intelligibility, languages, expression and recording conditions — with one pip package.
Paper vs. codebase. The paper (arXiv v3) reports 18 models across 18 robustness evaluations, 204 speakers, and 14,370 utterances, plus four more models in its appendix. This repository currently ships 32 integration entries.
One package, two ways to evaluate.
Use the metrics in an existing evaluation script, or use the packaged datasets to evaluate a new model. Your model keeps its own inference code and environment.
| You provide | RVCBench | |
|---|---|---|
| Score your audio | Generated audio, references and text | 7 automatic speech metrics through one Python API |
| Evaluate with our data | A model that generates WAV files | 3 versioned suites, prepared prompts and automatic scoring |
| Compare models | Output folders or existing reports | Checked scoring protocols, per-task comparisons and coverage |
| Reproduce a run | The original data and model settings | Input hashes, scoring provenance and interruption recovery |
“Bring your model or your audio. RVCBench provides the scoring and the evaluation data.”
Prepare prompts → generate speech → get your report.
Use your own files directly with the metrics API, or follow the dataset workflow below. Model inference runs with your code; RVCBench prepares inputs and computes the scores.
Choose data
Use your own recordings, or select onboarding-v1, core-v1 or full-v1.
3 packaged suitesPrepare prompts
Download verified reference clips and export texts in JSONL or native batch-list formats.
rvcbench promptsGenerate speech
Run your model locally or through an API. Save the output WAV files using the prompt identifiers.
your inference codeScore automatically
Compute the task metrics with pinned scoring models. Resume interrupted scoring.
rvcbench scoreCompare results
Read per-task means, sample scores and coverage; compare models with checked scoring protocols.
rvcbench compareFour dimensions, one benchmark.
Evaluate speaker identity, content and quality across input conditions, languages, output processing and protected references. The tables below show published paper results; packaged-suite results use their own versioned protocol.
Inputs and speakers
Does it still work when the reference audio or text prompt isn't clean studio speech?
- Reference-audio shifts — accents, ages, multi-speaker clips, café/station/train noise
- Text-prompt shifts — unusual, robocall-style, or hallucination-inducing prompts
Languages and expression
Does cloning quality hold up across model architectures, languages, and utterance length?
- 32 integration entries spanning codec-LM, diffusion, and hybrid architectures
- Multilingual (EN/ZH/FR), long-form generation, and emotion preservation
Output processing
Does the cloned output survive real-world post-processing, and can it be told apart from the real speaker?
- Post-processing resilience — MP3/AAC/Opus compression, phone-narrowband simulation
- Deepfake detectability — ground-truth vs. cloned speech classification
Noise and protection
Can a protection method actually stop a clone — and survive an attacker trying to denoise it back out?
- Passive perturbation — natural multi-speaker interference and environmental noise
- Proactive perturbation (5 methods) and counteract perturbation (adaptive denoising)
Which model clones a voice best?
Ranked by speaker similarity (SIM) on unprotected prompts, averaged across the full LibriTTS speaker set, as reported in the paper (arXiv v3); † marks models reported in its appendix. Click a quality column to sort. Historical RTF is raw and incomparable: timing boundaries and execution conditions are unverified. Speed ranking is disabled.
| # | Model | SIM ↑▾ | WER ↓ | MOS ↑ | MCD ↓ | RTF (raw, incomparable) | SVA ↑ | Emo ↑ |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3-TTS | 0.05 | 4.39 | 5.79 | 2.02 | 0.97 | 0.73 | |
| 2 | IndexTTS | 0.05 | 4.06 | 6.61 | 2.23 | 0.97 | 0.69 | |
| 3 | dots.tts † | 0.06 | 4.17 | 6.11 | 0.67 | 0.96 | 0.71 | |
| 4 | CosyVoice 2 | 0.05 | 4.37 | 6.02 | 4.81 | 0.97 | 0.70 | |
| 5 | ZipVoice | 0.05 | 4.13 | 7.09 | 1.46 | 0.95 | 0.68 | |
| 6 | MOSS-TTS v1.5 † | 0.06 | 4.32 | 6.66 | 0.76 | 0.95 | 0.70 | |
| 7 | GLM-TTS | 0.09 | 4.08 | 6.41 | 1.74 | 0.95 | 0.68 | |
| 8 | MaskGCT | 0.09 | 3.93 | 6.91 | 1.36 | 0.94 | 0.68 | |
| 9 | Higgs TTS 3 † | 0.05 | 4.23 | 6.32 | 0.64 | 0.96 | 0.71 | |
| 10 | F5-TTS | 0.12 | 3.99 | 6.96 | 0.61 | 0.94 | 0.68 | |
| 11 | Higgs Audio | 0.25 | 4.30 | 6.06 | 1.42 | 0.94 | 0.72 | |
| 12 | Fish Audio S2 † | 0.04 | 4.37 | 6.16 | 4.67 | 0.95 | 0.71 | |
| 13 | MGM-Omni | 0.09 | 4.28 | 5.82 | 0.84 | 0.93 | 0.68 | |
| 14 | PlayDiffusion | 0.05 | 4.15 | 8.06 | 0.73 | 0.94 | 0.68 | |
| 15 | MOSS-TTSD | 0.38 | 4.10 | 7.09 | 0.62 | 0.88 | 0.67 | |
| 16 | VibeVoice | 0.23 | 3.83 | 6.76 | 1.86 | 0.85 | 0.62 | |
| 17 | FishSpeech | 0.17 | 4.37 | 6.47 | 3.61 | 0.91 | 0.68 | |
| 18 | XTTS-v2 | 0.07 | 3.81 | 8.62 | 0.62 | 0.91 | 0.64 | |
| 19 | Spark-TTS | 0.33 | 4.06 | 5.83 | 1.56 | 0.76 | 0.67 | |
| 20 | OZSpeech | 0.06 | 3.21 | 6.87 | 8.75 | 0.84 | 0.64 | |
| 21 | OpenVoice V2 | 0.07 | 4.30 | 7.06 | 0.08 | 0.47 | 0.60 | |
| 22 | StyleTTS 2 | 0.05 | 4.30 | 6.81 | 0.11 | 0.39 | 0.59 |
SIM: speaker cosine similarity · WER: word error rate · MOS: SpeechMOS perceptual score · MCD: mel-cepstral distortion · RTF: raw real-time factor; incomparable historical scopes, no speed ranking · SVA: speaker-verification accuracy · Emo: emotion match rate. All values on clean prompts.
How far can protection push similarity down?
Each row spans a model's clean SIM down to its lowest SIM under any of the 5 proactive-perturbation methods below — the larger the gap, the more that model's voice can be shielded. The pipeline's optional Denoise step is the counteract-perturbation half of this dimension: an adaptive attacker trying to undo the protection before cloning.
Show full data — every protection method, every model
| Model | Clean | SafeSpeech | Enkidu | Spectral | GR-Noise | EM |
|---|---|---|---|---|---|---|
| Qwen3-TTS | 0.61 | 0.38 | 0.50 | 0.36 | 0.41 | 0.58 |
| IndexTTS | 0.61 | 0.35 | 0.47 | 0.32 | 0.39 | 0.57 |
| dots.tts † | 0.60 | 0.41 | 0.49 | 0.39 | 0.44 | 0.57 |
| CosyVoice 2 | 0.58 | 0.32 | 0.45 | 0.30 | 0.38 | 0.55 |
| ZipVoice | 0.58 | 0.29 | 0.44 | 0.26 | 0.26 | 0.54 |
| MOSS-TTS v1.5 † | 0.57 | 0.33 | 0.43 | 0.31 | 0.33 | 0.53 |
| GLM-TTS | 0.57 | 0.33 | 0.44 | 0.31 | 0.39 | 0.53 |
| MaskGCT | 0.57 | 0.30 | 0.41 | 0.28 | 0.31 | 0.53 |
| Higgs TTS 3 † | 0.56 | 0.48 | 0.49 | 0.48 | 0.34 | 0.53 |
| F5-TTS | 0.56 | 0.21 | 0.43 | 0.18 | 0.14 | 0.52 |
| Higgs Audio | 0.56 | 0.26 | 0.43 | 0.24 | 0.27 | 0.52 |
| Fish Audio S2 † | 0.54 | 0.32 | 0.43 | 0.30 | 0.34 | 0.52 |
| MGM-Omni | 0.54 | 0.18 | 0.32 | 0.17 | 0.23 | 0.49 |
| PlayDiffusion | 0.51 | 0.17 | 0.34 | 0.15 | 0.16 | 0.47 |
| MOSS-TTSD | 0.49 | 0.24 | 0.34 | 0.22 | 0.25 | 0.08 |
| VibeVoice | 0.48 | 0.27 | 0.37 | 0.25 | 0.28 | 0.45 |
| FishSpeech | 0.47 | 0.24 | 0.33 | 0.21 | 0.23 | 0.01 |
| XTTS-v2 | 0.45 | 0.26 | 0.31 | 0.24 | 0.24 | 0.41 |
| Spark-TTS | 0.41 | 0.13 | 0.14 | 0.11 | 0.06 | 0.36 |
| OZSpeech | 0.39 | 0.16 | 0.19 | 0.15 | 0.15 | 0.34 |
| OpenVoice V2 | 0.24 | 0.18 | 0.19 | 0.18 | 0.18 | 0.24 |
| StyleTTS 2 | 0.23 | 0.09 | 0.12 | 0.08 | 0.03 | 0.21 |
Does it still work outside LibriTTS?
Speaker similarity across 9 dataset conditions from the paper (arXiv v3), grouped by which dimension they test: Input Robustness shifts the reference audio or the text, Generation Robustness shifts the language or utterance length. Multi-speaker results are reported per SNR level in the paper. Hover a cell for the exact value.
| Model | Baseline | Input Robustness | Generation Robustness | ||||||
|---|---|---|---|---|---|---|---|---|---|
| LibriTTS | VCTK | BG-clean | BG-noise | Halluc. | Long | AISHELL | French | Bilingual | |
| Qwen3-TTS | 0.61 | 0.62 | 0.69 | 0.57 | 0.51 | 0.56 | 0.72 | 0.54 | 0.67 |
| IndexTTS | 0.61 | 0.57 | 0.59 | 0.53 | 0.53 | 0.78 | 0.72 | 0.40 | 0.67 |
| dots.tts † | 0.60 | 0.57 | 0.65 | 0.56 | 0.54 | 0.72 | 0.68 | 0.46 | 0.61 |
| CosyVoice 2 | 0.58 | 0.58 | 0.63 | 0.51 | 0.52 | 0.53 | 0.72 | 0.38 | 0.65 |
| ZipVoice | 0.58 | 0.55 | 0.62 | 0.46 | 0.51 | 0.73 | 0.71 | 0.36 | 0.63 |
| MOSS-TTS v1.5 † | 0.57 | 0.51 | 0.58 | 0.46 | 0.46 | 0.66 | 0.68 | 0.46 | 0.65 |
| GLM-TTS | 0.57 | 0.57 | 0.62 | 0.53 | 0.53 | 0.76 | 0.69 | 0.40 | 0.66 |
| MaskGCT | 0.57 | 0.56 | 0.61 | 0.49 | 0.50 | 0.76 | 0.67 | 0.49 | 0.63 |
| Higgs TTS 3 † | 0.56 | 0.45 | 0.49 | 0.31 | 0.40 | 0.74 | 0.63 | 0.47 | 0.30 |
| F5-TTS | 0.56 | 0.54 | 0.58 | 0.41 | 0.46 | 0.61 | 0.70 | 0.30 | 0.65 |
| Higgs Audio | 0.56 | 0.52 | 0.59 | 0.42 | 0.42 | 0.52 | 0.58 | 0.35 | 0.54 |
| Fish Audio S2 † | 0.54 | 0.51 | 0.58 | 0.48 | 0.41 | 0.50 | 0.66 | 0.46 | 0.62 |
| MGM-Omni | 0.54 | 0.45 | 0.52 | 0.33 | 0.40 | 0.44 | 0.71 | 0.23 | 0.63 |
| PlayDiffusion | 0.51 | 0.43 | 0.43 | 0.31 | 0.41 | 0.64 | 0.44 | 0.28 | 0.46 |
| MOSS-TTSD | 0.49 | 0.44 | 0.49 | 0.04 | 0.42 | 0.64 | 0.44 | 0.33 | 0.44 |
| VibeVoice | 0.48 | 0.44 | 0.51 | 0.36 | 0.41 | 0.62 | 0.56 | 0.34 | 0.53 |
| FishSpeech | 0.47 | 0.43 | 0.50 | 0.39 | 0.35 | 0.57 | 0.61 | 0.37 | 0.57 |
| XTTS-v2 | 0.45 | 0.45 | 0.55 | 0.39 | 0.49 | 0.61 | 0.57 | 0.45 | 0.51 |
| Spark-TTS | 0.41 | 0.53 | 0.59 | 0.33 | 0.34 | 0.35 | 0.57 | 0.16 | 0.48 |
| OZSpeech | 0.39 | 0.25 | 0.27 | 0.16 | 0.28 | 0.42 | 0.20 | 0.11 | 0.17 |
| OpenVoice V2 | 0.24 | 0.39 | 0.48 | 0.36 | 0.36 | 0.28 | 0.43 | 0.27 | 0.30 |
| StyleTTS 2 | 0.23 | 0.24 | 0.20 | 0.17 | 0.18 | 0.20 | 0.11 | 0.11 | 0.21 |
Results with sample-level coverage
The tables above preserve the historical release. New reports below are generated from complete run manifests with explicit sample counts and metric coverage; historical rows have not been retroactively certified.
No reports have been published under the new manifest protocol yet.
32 integration entries, one interface.
Adapters with published historical results are distinguished from experimental integrations. Environment templates require the corresponding upstream runtime; an adapter is not a guarantee of a tested installation.
Five ways to make a voice harder to clone.
SafeSpeech
Adversarial perturbation optimised against a surrogate VC model.
Enkidu
Perceptual-loss adversarial perturbation.
EM
Expectation–Maximisation perturbation.
GRNoise
Gaussian random noise — no surrogate model required.
Spectral
SafeSpeech's spectral perturbation mode.
10 benchmark conditions, publicly hosted.
Every subset is a Hugging Face dataset config under Nanboy/RVCBench, fetched automatically by the data loader.
| Config | Language | Typical use |
|---|---|---|
| Libritts | EN | English zero-shot VC/TTS benchmark prompts. |
| VCTK | EN | Multi-speaker English voice cloning. |
| Multispeaker_libri | EN | Multi-speaker LibriSpeech-style evaluation. |
| Long_context | EN | Longer-context voice-cloning prompts. |
| AISHELL1_dev | ZH | Mandarin speech evaluation. |
| CommonVoiceFR_dev | FR | French speech evaluation. |
| Bilingual_uedin | EN/ZH | Bilingual speech evaluation. |
| Background_noise | EN | Noisy-prompt robustness. |
| robotcall | EN | Robocall-style speech robustness. |
| vctk_text_robust | EN | Text robustness on VCTK-style prompts. |
The referee's scorecard.
Speaker cosine similarity between cloned and target voice.
Word error rate — how intelligible the generated speech is.
SpeechMOS perceptual quality score.
Mel-cepstral distortion versus the reference.
Raw real-time factor. Historical timing scopes are unverified; cross-model speed ranking is not supported.
Speaker-verification accuracy against the target identity.
Emotion match rate between clone and target.
Fidelity metrics reported for protection and denoising runs.
Questions people ask about RVCBench.
Can I score my own audio without using the datasets?
Yes. Install rvcbench[eval], prepare the scorers with rvcbench setup-scorers, and use rvcbench.metrics.Evaluator. Choose from SIM, SVA, WER, MOS, MCD, STOI and emotion consistency. MCD and STOI require a recording of the same text; other metrics use the reference voice or expected text.
How do I evaluate a model using RVCBench data?
Run rvcbench prompts to prepare a versioned suite, synthesize the listed texts with your model, and run rvcbench score on its WAV files. Start with onboarding-v1 (52 utterances), then use core-v1 (480) or full-v1 (12,724). Scoring, reports and --resume are included in the pip package. The packaged suites are previews; paper results use a separate protocol.
What is RVCBench?
RVCBench is a general-purpose package for voice cloning evaluation, with automatic speech metrics and ready-to-use datasets. It provides 32 integration entries and paper results for 22 models (18 in the main results, 4 in the appendix), with 5 audio-protection methods across 10 dataset configurations, scoring speaker similarity, intelligibility, perceptual quality, and runtime.
How many voice-cloning models does RVCBench evaluate?
The RVCBench codebase includes 32 TTS/VC integration entries. The paper (arXiv v3) reports results for 18 of those models across 18 robustness evaluations, 204 speakers, and 14,370 utterances.
What audio-protection methods does RVCBench compare?
Five methods on equal footing: SafeSpeech (adversarial perturbation against a surrogate VC model), Enkidu (perceptual-loss adversarial perturbation), POP (error-minimizing perturbation, named em in the code), Spectral (SafeSpeech's spectral perturbation mode), and GR-Noise (Gaussian random noise).
Which model is hardest to clone under protection, according to RVCBench?
Across the LibriTTS leaderboard, StyleTTS 2 and OpenVoice V2 have the lowest clean speaker similarity and drop furthest under protection — GR-Noise pushes StyleTTS 2's similarity from 0.23 down to 0.03.
Is the RVCBench dataset public?
Yes. The benchmark dataset is hosted on Hugging Face at huggingface.co/datasets/Nanboy/RVCBench under a CC0-1.0 license, with 10 dataset configurations spanning English, Mandarin, and French.
How do I cite RVCBench?
Cite the NeurIPS 2026 paper: Jin, Ruinan; Liao, Xinting; Yu, Hanlin; Pandya, Deval; Li, Xiaoxiao. “RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models.” Advances in Neural Information Processing Systems (NeurIPS), 2026. arXiv:2602.00443.
Install once. Score your audio or evaluate with our data.
Both workflows work from the pip package. Install with scoring dependencies, then download the metric models once. Linux, Python 3.10+ and FFmpeg are required.
python -m pip install "rvcbench[eval]" rvcbench setup-scorers
1. Automatic metrics for your audio
from rvcbench import metrics
with metrics.Evaluator(["sim", "wer", "speechmos"]) as e:
scores = e.score("generated.wav",
reference="reference.wav", text="Hello there.",
language="en")
print(scores)2. Evaluate with our datasets
rvcbench prompts --suite onboarding-v1 --output prompts/ # Generate each prompt with your model into outputs/my-model/. rvcbench score --suite onboarding-v1 \ --generated outputs/my-model --output results/my-model # Add --resume to continue an interrupted evaluation.
If RVCBench is useful, cite it.
Contributions are welcome — new protection methods, adversary wrappers, dataset adapters, and evaluation metrics. Open an issue or PR on GitHub.
@inproceedings{jin2026rvcbench,
title = {RVCBench: Benchmarking the Robustness of
Voice Cloning Across Modern Audio
Generation Models},
author = {Jin, Ruinan and Liao, Xinting and Yu, Hanlin
and Pandya, Deval and Li, Xiaoxiao},
booktitle = {Advances in Neural Information Processing Systems},
url = {https://arxiv.org/abs/2602.00443},
year = {2026}
}
RVCBench