Comprehensive voice cloning evaluation

Evaluate voice cloning.
Metrics and datasets in one package.

RVCBench brings automatic speech metrics and ready-to-use evaluation datasets together. Score your own audio, or evaluate a model across speaker identity, speech quality, intelligibility, languages, expression and recording conditions — with one pip package.

Accepted to NeurIPS 2026 · arXiv:2602.00443 · CC0-1.0 · Jin, Liao, Yu, Pandya, Li

signal.monitor — reference → generate → scorelive
■ clean signal (SIM ↑) ■ processed speech
32
VC / TTS models
5
protection methods
10
dataset configs
204
speakers (paper)
14,370
utterances (paper)
3
languages — EN / ZH / FR

Paper vs. codebase. The paper (arXiv v3) reports 18 models across 18 robustness evaluations, 204 speakers, and 14,370 utterances, plus four more models in its appendix. This repository currently ships 32 integration entries.

Why RVCBench

One package, two ways to evaluate.

Use the metrics in an existing evaluation script, or use the packaged datasets to evaluate a new model. Your model keeps its own inference code and environment.

You provideRVCBench
Score your audioGenerated audio, references and text7 automatic speech metrics through one Python API
Evaluate with our dataA model that generates WAV files3 versioned suites, prepared prompts and automatic scoring
Compare modelsOutput folders or existing reportsChecked scoring protocols, per-task comparisons and coverage
Reproduce a runThe original data and model settingsInput hashes, scoring provenance and interruption recovery

“Bring your model or your audio. RVCBench provides the scoring and the evaluation data.”

How it works

Prepare prompts → generate speech → get your report.

Use your own files directly with the metrics API, or follow the dataset workflow below. Model inference runs with your code; RVCBench prepares inputs and computes the scores.

1

Choose data

Use your own recordings, or select onboarding-v1, core-v1 or full-v1.

3 packaged suites
2

Prepare prompts

Download verified reference clips and export texts in JSONL or native batch-list formats.

rvcbench prompts
3

Generate speech

Run your model locally or through an API. Save the output WAV files using the prompt identifiers.

your inference code
4

Score automatically

Compute the task metrics with pinned scoring models. Resume interrupted scoring.

rvcbench score
5

Compare results

Read per-task means, sample scores and coverage; compare models with checked scoring protocols.

rvcbench compare
The framework

Four dimensions, one benchmark.

Evaluate speaker identity, content and quality across input conditions, languages, output processing and protected references. The tables below show published paper results; packaged-suite results use their own versioned protocol.

Inputs and speakers

Does it still work when the reference audio or text prompt isn't clean studio speech?

  • Reference-audio shifts — accents, ages, multi-speaker clips, café/station/train noise
  • Text-prompt shifts — unusual, robocall-style, or hallucination-inducing prompts
See it in the cross-dataset heatmap →

Languages and expression

Does cloning quality hold up across model architectures, languages, and utterance length?

  • 32 integration entries spanning codec-LM, diffusion, and hybrid architectures
  • Multilingual (EN/ZH/FR), long-form generation, and emotion preservation
See it in the leaderboard →

Output processing

Does the cloned output survive real-world post-processing, and can it be told apart from the real speaker?

  • Post-processing resilience — MP3/AAC/Opus compression, phone-narrowband simulation
  • Deepfake detectability — ground-truth vs. cloned speech classification
In the codebase (data/compression/, src/rvcbench/datasets/deepfake_preprocess.py) — not yet on the public leaderboard

Noise and protection

Can a protection method actually stop a clone — and survive an attacker trying to denoise it back out?

  • Passive perturbation — natural multi-speaker interference and environmental noise
  • Proactive perturbation (5 methods) and counteract perturbation (adaptive denoising)
See it in the protection-robustness chart →
Generation Robustness Leaderboard — LibriTTS, clean prompts

Which model clones a voice best?

Ranked by speaker similarity (SIM) on unprotected prompts, averaged across the full LibriTTS speaker set, as reported in the paper (arXiv v3); † marks models reported in its appendix. Click a quality column to sort. Historical RTF is raw and incomparable: timing boundaries and execution conditions are unverified. Speed ranking is disabled.

# Model SIM ↑▾ WER ↓ MOS ↑ MCD ↓ RTF (raw, incomparable) SVA ↑ Emo ↑
1Qwen3-TTS0.610.054.395.792.020.970.73
2IndexTTS0.610.054.066.612.230.970.69
3dots.tts †0.600.064.176.110.670.960.71
4CosyVoice 20.580.054.376.024.810.970.70
5ZipVoice0.580.054.137.091.460.950.68
6MOSS-TTS v1.5 †0.570.064.326.660.760.950.70
7GLM-TTS0.570.094.086.411.740.950.68
8MaskGCT0.570.093.936.911.360.940.68
9Higgs TTS 3 †0.560.054.236.320.640.960.71
10F5-TTS0.560.123.996.960.610.940.68
11Higgs Audio0.560.254.306.061.420.940.72
12Fish Audio S2 †0.540.044.376.164.670.950.71
13MGM-Omni0.540.094.285.820.840.930.68
14PlayDiffusion0.510.054.158.060.730.940.68
15MOSS-TTSD0.490.384.107.090.620.880.67
16VibeVoice0.480.233.836.761.860.850.62
17FishSpeech0.470.174.376.473.610.910.68
18XTTS-v20.450.073.818.620.620.910.64
19Spark-TTS0.410.334.065.831.560.760.67
20OZSpeech0.390.063.216.878.750.840.64
21OpenVoice V20.240.074.307.060.080.470.60
22StyleTTS 20.230.054.306.810.110.390.59

SIM: speaker cosine similarity · WER: word error rate · MOS: SpeechMOS perceptual score · MCD: mel-cepstral distortion · RTF: raw real-time factor; incomparable historical scopes, no speed ranking · SVA: speaker-verification accuracy · Emo: emotion match rate. All values on clean prompts.

Audio Perturbation Robustness Protection robustness — SIM, LibriTTS

How far can protection push similarity down?

Each row spans a model's clean SIM down to its lowest SIM under any of the 5 proactive-perturbation methods below — the larger the gap, the more that model's voice can be shielded. The pipeline's optional Denoise step is the counteract-perturbation half of this dimension: an adaptive attacker trying to undo the protection before cloning.

Clean prompt Best-case protection (lowest SIM reached, method labeled)
0.00.10.20.30.40.50.6Qwen3-TTSQwen3-TTS · Spectral: 0.360Qwen3-TTS · Clean: 0.610SpectralIndexTTSIndexTTS · Spectral: 0.320IndexTTS · Clean: 0.610Spectraldots.tts †dots.tts † · Spectral: 0.390dots.tts † · Clean: 0.600SpectralCosyVoice 2CosyVoice 2 · Spectral: 0.300CosyVoice 2 · Clean: 0.580SpectralZipVoiceZipVoice · Spectral: 0.260ZipVoice · Clean: 0.580SpectralMOSS-TTS v1.5 †MOSS-TTS v1.5 † · Spectral: 0.310MOSS-TTS v1.5 † · Clean: 0.570SpectralGLM-TTSGLM-TTS · Spectral: 0.310GLM-TTS · Clean: 0.570SpectralMaskGCTMaskGCT · Spectral: 0.280MaskGCT · Clean: 0.570SpectralHiggs TTS 3 †Higgs TTS 3 † · GR-Noise: 0.340Higgs TTS 3 † · Clean: 0.560GR-NoiseF5-TTSF5-TTS · GR-Noise: 0.140F5-TTS · Clean: 0.560GR-NoiseHiggs AudioHiggs Audio · Spectral: 0.240Higgs Audio · Clean: 0.560SpectralFish Audio S2 †Fish Audio S2 † · Spectral: 0.300Fish Audio S2 † · Clean: 0.540SpectralMGM-OmniMGM-Omni · Spectral: 0.170MGM-Omni · Clean: 0.540SpectralPlayDiffusionPlayDiffusion · Spectral: 0.150PlayDiffusion · Clean: 0.510SpectralMOSS-TTSDMOSS-TTSD · EM: 0.080MOSS-TTSD · Clean: 0.490EMVibeVoiceVibeVoice · Spectral: 0.250VibeVoice · Clean: 0.480SpectralFishSpeechFishSpeech · EM: 0.010FishSpeech · Clean: 0.470EMXTTS-v2XTTS-v2 · Spectral: 0.240XTTS-v2 · Clean: 0.450SpectralSpark-TTSSpark-TTS · GR-Noise: 0.060Spark-TTS · Clean: 0.410GR-NoiseOZSpeechOZSpeech · Spectral: 0.150OZSpeech · Clean: 0.390SpectralOpenVoice V2OpenVoice V2 · SafeSpeech: 0.180OpenVoice V2 · Clean: 0.240SafeSpeechStyleTTS 2StyleTTS 2 · GR-Noise: 0.030StyleTTS 2 · Clean: 0.230GR-Noise
Show full data — every protection method, every model
ModelCleanSafeSpeechEnkiduSpectralGR-NoiseEM
Qwen3-TTS0.610.380.500.360.410.58
IndexTTS0.610.350.470.320.390.57
dots.tts †0.600.410.490.390.440.57
CosyVoice 20.580.320.450.300.380.55
ZipVoice0.580.290.440.260.260.54
MOSS-TTS v1.5 †0.570.330.430.310.330.53
GLM-TTS0.570.330.440.310.390.53
MaskGCT0.570.300.410.280.310.53
Higgs TTS 3 †0.560.480.490.480.340.53
F5-TTS0.560.210.430.180.140.52
Higgs Audio0.560.260.430.240.270.52
Fish Audio S2 †0.540.320.430.300.340.52
MGM-Omni0.540.180.320.170.230.49
PlayDiffusion0.510.170.340.150.160.47
MOSS-TTSD0.490.240.340.220.250.08
VibeVoice0.480.270.370.250.280.45
FishSpeech0.470.240.330.210.230.01
XTTS-v20.450.260.310.240.240.41
Spark-TTS0.410.130.140.110.060.36
OZSpeech0.390.160.190.150.150.34
OpenVoice V20.240.180.190.180.180.24
StyleTTS 20.230.090.120.080.030.21
InputGeneration Cross-dataset generalisation — SIM, clean prompts

Does it still work outside LibriTTS?

Speaker similarity across 9 dataset conditions from the paper (arXiv v3), grouped by which dimension they test: Input Robustness shifts the reference audio or the text, Generation Robustness shifts the language or utterance length. Multi-speaker results are reported per SNR level in the paper. Hover a cell for the exact value.

ModelBaselineInput RobustnessGeneration Robustness
LibriTTSVCTKBG-cleanBG-noiseHalluc.LongAISHELLFrenchBilingual
Qwen3-TTS0.610.620.690.570.510.560.720.540.67
IndexTTS0.610.570.590.530.530.780.720.400.67
dots.tts †0.600.570.650.560.540.720.680.460.61
CosyVoice 20.580.580.630.510.520.530.720.380.65
ZipVoice0.580.550.620.460.510.730.710.360.63
MOSS-TTS v1.5 †0.570.510.580.460.460.660.680.460.65
GLM-TTS0.570.570.620.530.530.760.690.400.66
MaskGCT0.570.560.610.490.500.760.670.490.63
Higgs TTS 3 †0.560.450.490.310.400.740.630.470.30
F5-TTS0.560.540.580.410.460.610.700.300.65
Higgs Audio0.560.520.590.420.420.520.580.350.54
Fish Audio S2 †0.540.510.580.480.410.500.660.460.62
MGM-Omni0.540.450.520.330.400.440.710.230.63
PlayDiffusion0.510.430.430.310.410.640.440.280.46
MOSS-TTSD0.490.440.490.040.420.640.440.330.44
VibeVoice0.480.440.510.360.410.620.560.340.53
FishSpeech0.470.430.500.390.350.570.610.370.57
XTTS-v20.450.450.550.390.490.610.570.450.51
Spark-TTS0.410.530.590.330.340.350.570.160.48
OZSpeech0.390.250.270.160.280.420.200.110.17
OpenVoice V20.240.390.480.360.360.280.430.270.30
StyleTTS 20.230.240.200.170.180.200.110.110.21
0.00.78 SIM
Reproducible runs

Results with sample-level coverage

The tables above preserve the historical release. New reports below are generated from complete run manifests with explicit sample counts and metric coverage; historical rows have not been retroactively certified.

No reports have been published under the new manifest protocol yet.

Supported adversaries

32 integration entries, one interface.

Adapters with published historical results are distinguished from experimental integrations. Environment templates require the corresponding upstream runtime; an adapter is not a guarantee of a tested installation.

On the LibriTTS leaderboard (18) Experimental adapters without leaderboard results (14)
BertVITS2bert
Qwen3-TTSqwen3_tts
Qwen3-Omniqwen3_omni
FireRedTTS-2fireredtts2
VoxCPMvoxcpm
F5-TTSf5_tts
MaskGCTmaskgct
OpenVoice V2openvoice
Coqui XTTS-v2xtts
IndexTTSindex_tts
ZipVoicezipvoice
FishSpeechfishspeech
Fish Audio S2 (in-proc)fishspeech_s2
Fish Audio S2 (server)fish_audio_s2
CosyVoice / 2cosyvoice
Higgs Audiohiggs_audio
Higgs TTS 3higgs_tts_3
SparkTTSsparktts
VALL-Evall_e
StyleTTS 2styletts2
GLM-TTSglm_tts
GlowTTSglowtts
Kimi Audiokimi_audio
MGM-Omnimgm_omni
MOSS TTSDmoss_ttsd
MOSS-TTSmoss_tts
dots.ttsdots_tts
ZONOS2zonos2
PlayDiffusionplaydiffusion
Bark Voice Clonebark_voice_clone
OZSpeechozspeech
VibeVoicevibevoice
Protection methods

Five ways to make a voice harder to clone.

SafeSpeech

Adversarial perturbation optimised against a surrogate VC model.

Enkidu

Perceptual-loss adversarial perturbation.

EM

Expectation–Maximisation perturbation.

GRNoise

Gaussian random noise — no surrogate model required.

Spectral

SafeSpeech's spectral perturbation mode.

Datasets

10 benchmark conditions, publicly hosted.

Every subset is a Hugging Face dataset config under Nanboy/RVCBench, fetched automatically by the data loader.

ConfigLanguageTypical use
LibrittsENEnglish zero-shot VC/TTS benchmark prompts.
VCTKENMulti-speaker English voice cloning.
Multispeaker_libriENMulti-speaker LibriSpeech-style evaluation.
Long_contextENLonger-context voice-cloning prompts.
AISHELL1_devZHMandarin speech evaluation.
CommonVoiceFR_devFRFrench speech evaluation.
Bilingual_uedinEN/ZHBilingual speech evaluation.
Background_noiseENNoisy-prompt robustness.
robotcallENRobocall-style speech robustness.
vctk_text_robustENText robustness on VCTK-style prompts.
Evaluated metrics

The referee's scorecard.

SIM↑

Speaker cosine similarity between cloned and target voice.

WER↓

Word error rate — how intelligible the generated speech is.

MOS↑

SpeechMOS perceptual quality score.

MCD↓

Mel-cepstral distortion versus the reference.

RTF↓

Raw real-time factor. Historical timing scopes are unverified; cross-model speed ranking is not supported.

SVA↑

Speaker-verification accuracy against the target identity.

Emo↑

Emotion match rate between clone and target.

SNR / STOI / DNSMOS

Fidelity metrics reported for protection and denoising runs.

FAQ

Questions people ask about RVCBench.

Can I score my own audio without using the datasets?

Yes. Install rvcbench[eval], prepare the scorers with rvcbench setup-scorers, and use rvcbench.metrics.Evaluator. Choose from SIM, SVA, WER, MOS, MCD, STOI and emotion consistency. MCD and STOI require a recording of the same text; other metrics use the reference voice or expected text.

How do I evaluate a model using RVCBench data?

Run rvcbench prompts to prepare a versioned suite, synthesize the listed texts with your model, and run rvcbench score on its WAV files. Start with onboarding-v1 (52 utterances), then use core-v1 (480) or full-v1 (12,724). Scoring, reports and --resume are included in the pip package. The packaged suites are previews; paper results use a separate protocol.

What is RVCBench?

RVCBench is a general-purpose package for voice cloning evaluation, with automatic speech metrics and ready-to-use datasets. It provides 32 integration entries and paper results for 22 models (18 in the main results, 4 in the appendix), with 5 audio-protection methods across 10 dataset configurations, scoring speaker similarity, intelligibility, perceptual quality, and runtime.

How many voice-cloning models does RVCBench evaluate?

The RVCBench codebase includes 32 TTS/VC integration entries. The paper (arXiv v3) reports results for 18 of those models across 18 robustness evaluations, 204 speakers, and 14,370 utterances.

What audio-protection methods does RVCBench compare?

Five methods on equal footing: SafeSpeech (adversarial perturbation against a surrogate VC model), Enkidu (perceptual-loss adversarial perturbation), POP (error-minimizing perturbation, named em in the code), Spectral (SafeSpeech's spectral perturbation mode), and GR-Noise (Gaussian random noise).

Which model is hardest to clone under protection, according to RVCBench?

Across the LibriTTS leaderboard, StyleTTS 2 and OpenVoice V2 have the lowest clean speaker similarity and drop furthest under protection — GR-Noise pushes StyleTTS 2's similarity from 0.23 down to 0.03.

Is the RVCBench dataset public?

Yes. The benchmark dataset is hosted on Hugging Face at huggingface.co/datasets/Nanboy/RVCBench under a CC0-1.0 license, with 10 dataset configurations spanning English, Mandarin, and French.

How do I cite RVCBench?

Cite the NeurIPS 2026 paper: Jin, Ruinan; Liao, Xinting; Yu, Hanlin; Pandya, Deval; Li, Xiaoxiao. “RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models.” Advances in Neural Information Processing Systems (NeurIPS), 2026. arXiv:2602.00443.

Getting started

Install once. Score your audio or evaluate with our data.

Both workflows work from the pip package. Install with scoring dependencies, then download the metric models once. Linux, Python 3.10+ and FFmpeg are required.

python -m pip install "rvcbench[eval]"
rvcbench setup-scorers

1. Automatic metrics for your audio

from rvcbench import metrics

with metrics.Evaluator(["sim", "wer", "speechmos"]) as e:
    scores = e.score("generated.wav",
        reference="reference.wav", text="Hello there.",
        language="en")
    print(scores)

2. Evaluate with our datasets

rvcbench prompts --suite onboarding-v1 --output prompts/
# Generate each prompt with your model into outputs/my-model/.
rvcbench score --suite onboarding-v1 \
  --generated outputs/my-model --output results/my-model
# Add --resume to continue an interrupted evaluation.
Citation

If RVCBench is useful, cite it.

Contributions are welcome — new protection methods, adversary wrappers, dataset adapters, and evaluation metrics. Open an issue or PR on GitHub.

@inproceedings{jin2026rvcbench,
  title   = {RVCBench: Benchmarking the Robustness of
             Voice Cloning Across Modern Audio
             Generation Models},
  author  = {Jin, Ruinan and Liao, Xinting and Yu, Hanlin
             and Pandya, Deval and Li, Xiaoxiao},
  booktitle = {Advances in Neural Information Processing Systems},
  url     = {https://arxiv.org/abs/2602.00443},
  year    = {2026}
}