Skip to content

RVCBench

NeurIPS 2026

Accepted to NeurIPS 2026.

RVCBench is a benchmark dataset for studying robustness in voice cloning, text-to-speech, speaker privacy, audio protection, adversarial audio perturbations, and related audio generation pipelines.

Dataset page: https://huggingface.co/datasets/Nanboy/RVCBench Code repository: https://github.com/Nanboy-Ronan/RVCBench Paper: https://arxiv.org/abs/2602.00443

RVCBench is designed for evaluating how modern voice cloning (VC), TTS, and audio generation systems behave under clean prompts, protected prompts, and denoised protected prompts. It supports research on audio deepfake robustness, anti-spoofing, speaker verification resilience, privacy-preserving speech generation, and standardized benchmark evaluation.

Each subset is exposed as its own Hugging Face dataset configuration. Most subsets contain:

  • metadata.parquet
  • audios/

The canonical metadata stores one row per benchmark pair with columns such as:

  • speaker_id
  • prompt_file_name
  • prompt_text
  • prompt_language
  • target_file_name
  • target_text
  • target_language
  • pair_id
  • dataset_name
  • data_split

When available, training-oriented annotations are also preserved:

  • prompt_phonemes, prompt_tone, prompt_word2ph
  • target_phonemes, target_tone, target_word2ph

Some subsets include additional task-specific metadata, for example spam_type in robotcall.

Available Configs

  • AISHELL1_dev
  • Background_noise
  • Bilingual_uedin
  • CommonVoiceFR_dev
  • Libritts
  • Long_context
  • Multispeaker_libri
  • VCTK
  • robotcall
  • vctk_text_robust

Getting started

Follow the public quickstart. The Qwen quickstart downloads the selected speaker audio and generates a small sample without loading the full evaluation stack. Use the run protocol to evaluate, inspect coverage, or resume a run. Keep the dataset revision or input hashes with any reported result; do not compare scores computed on different successful subsets without reporting coverage.

Intended Use

Use this dataset with the RVCBench codebase to run reproducible voice cloning robustness experiments across source audio, protected audio, denoised audio, and generated audio. Typical tasks include:

  • benchmarking zero-shot voice cloning and TTS models;
  • comparing audio protection methods such as SafeSpeech, Enkidu, EM, spectral perturbation, and Gaussian noise;
  • measuring speaker similarity, intelligibility, perceptual quality, word error rate, and runtime;
  • studying speaker privacy and audio deepfake robustness under adversarial perturbations.

Citation

If you use RVCBench in your research, please cite:

@inproceedings{jin2026rvcbench,
  title   = {RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models},
  author  = {Jin, Ruinan and Liao, Xinting and Yu, Hanlin and Pandya, Deval and Li, Xiaoxiao},
  booktitle = {Advances in Neural Information Processing Systems},
  url     = {https://arxiv.org/abs/2602.00443},
  year    = {2026}
}

Notes

  • The compression directory is intentionally excluded from this dataset release.
  • This repository is organized for direct browsing in the Hugging Face dataset viewer.