RVCBench¶
Accepted to NeurIPS 2026.
RVCBench is a benchmark dataset for studying robustness in voice cloning, text-to-speech, speaker privacy, audio protection, adversarial audio perturbations, and related audio generation pipelines.
Dataset page: https://huggingface.co/datasets/Nanboy/RVCBench Code repository: https://github.com/Nanboy-Ronan/RVCBench Paper: https://arxiv.org/abs/2602.00443
RVCBench is designed for evaluating how modern voice cloning (VC), TTS, and audio generation systems behave under clean prompts, protected prompts, and denoised protected prompts. It supports research on audio deepfake robustness, anti-spoofing, speaker verification resilience, privacy-preserving speech generation, and standardized benchmark evaluation.
Each subset is exposed as its own Hugging Face dataset configuration. Most subsets contain:
metadata.parquetaudios/
The canonical metadata stores one row per benchmark pair with columns such as:
speaker_idprompt_file_nameprompt_textprompt_languagetarget_file_nametarget_texttarget_languagepair_iddataset_namedata_split
When available, training-oriented annotations are also preserved:
prompt_phonemes,prompt_tone,prompt_word2phtarget_phonemes,target_tone,target_word2ph
Some subsets include additional task-specific metadata, for example spam_type in robotcall.
Available Configs¶
AISHELL1_devBackground_noiseBilingual_uedinCommonVoiceFR_devLibrittsLong_contextMultispeaker_libriVCTKrobotcallvctk_text_robust
Getting started¶
Follow the public quickstart. The Qwen quickstart downloads the selected speaker audio and generates a small sample without loading the full evaluation stack. Use the run protocol to evaluate, inspect coverage, or resume a run. Keep the dataset revision or input hashes with any reported result; do not compare scores computed on different successful subsets without reporting coverage.
Intended Use¶
Use this dataset with the RVCBench codebase to run reproducible voice cloning robustness experiments across source audio, protected audio, denoised audio, and generated audio. Typical tasks include:
- benchmarking zero-shot voice cloning and TTS models;
- comparing audio protection methods such as SafeSpeech, Enkidu, EM, spectral perturbation, and Gaussian noise;
- measuring speaker similarity, intelligibility, perceptual quality, word error rate, and runtime;
- studying speaker privacy and audio deepfake robustness under adversarial perturbations.
Citation¶
If you use RVCBench in your research, please cite:
@inproceedings{jin2026rvcbench,
title = {RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models},
author = {Jin, Ruinan and Liao, Xinting and Yu, Hanlin and Pandya, Deval and Li, Xiaoxiao},
booktitle = {Advances in Neural Information Processing Systems},
url = {https://arxiv.org/abs/2602.00443},
year = {2026}
}
Notes¶
- The
compressiondirectory is intentionally excluded from this dataset release. - This repository is organized for direct browsing in the Hugging Face dataset viewer.