Running the built-in models¶
RVCBench includes adapters for 32 voice cloning integrations. Each one needs its own Python environment, the upstream inference code and the model checkpoints; none of these are bundled with the package. If you only want scores for a model you can already run, you do not need any of this: score your own outputs instead.
Installation¶
Python 3.10 or newer on Linux. A GPU is recommended for scoring.
# package only
pip install "rvcbench[eval]"
# or a source checkout, needed to run the built-in models
git clone https://github.com/Nanboy-Ronan/RVCBench.git
cd RVCBench
python -m pip install -e '.[eval]'
rvcbench doctor # check dependencies
rvcbench smoke --output results/smoke # synthetic CPU pipeline check, no downloads
| Extra | Adds |
|---|---|
| (none) | Runner, run records, suites and the rvcbench command |
eval |
Metrics: Whisper, SpeechBrain ECAPA and emotion, SpeechMOS, MCD, STOI |
qwen3 |
Qwen3-TTS runtime for the quickstart |
http |
Clients for server-backed models |
enkidu |
Enkidu protection |
dev |
Tests, lint, pre-commit and build tools |
- FFmpeg is needed for the compression tasks and by several models.
- Hugging Face login (
hf auth login) is recommended: anonymous downloads are rate-limited. - Scorer models are stored in
$RVCBENCH_ASSET_DIR, else./checkpoints/when it exists, else~/.cache/rvcbench/; see model environments. - If installing
[eval]reports that nopysptkversion matches, see building the evaluation extras.
Supported models¶
"Paper" marks the 18 models of the paper's main results and the four reported in its appendix (arXiv v3). "v2 subset run" means real generation and core scoring were recorded on a fixed subset during the v2 refactor; see validation coverage.
| Model | vc.model |
Paper | v2 subset run |
|---|---|---|---|
| BertVITS2 | bertvits2 |
pending | |
| Qwen3-TTS | qwen3_tts |
main | ✓ |
| Qwen3-Omni | qwen3_omni |
pending | |
| FireRedTTS-2 | fireredtts2 |
✓ | |
| VoxCPM | voxcpm |
✓ | |
| F5-TTS | f5_tts |
main | ✓ |
| MaskGCT | maskgct |
main | ✓ |
| OpenVoice V2 | openvoice |
main | ✓ |
| Coqui XTTS-v2 | xtts |
main | ✓ |
| IndexTTS | index_tts |
main | ✓ |
| ZipVoice | zipvoice |
main | ✓ |
| FishSpeech | fishspeech |
main | ✓ |
| Fish Audio S2 (in-proc) | fishspeech_s2 |
appendix | ✓ |
| Fish Audio S2 (server) | fish_audio_s2 |
appendix | ✓ |
| CosyVoice / 2 | cosyvoice |
main | ✓ |
| Higgs Audio | higgs_audio |
main | ✓ |
| Higgs TTS 3 | higgs_tts_3 |
appendix | pending |
| SparkTTS | sparktts |
main | ✓ |
| VALL-E | vall_e |
pending | |
| StyleTTS 2 | styletts2 |
main | ✓ |
| GLM-TTS | glm_tts |
main | ✓ |
| GlowTTS | glowtts |
pending | |
| Kimi Audio | kimi_audio |
✓ | |
| MGM-Omni | mgm_omni |
main | ✓ |
| MOSS TTSD | moss_ttsd |
main | ✓ |
| MOSS-TTS | moss_tts |
appendix | ✓ |
| dots.tts | dots_tts |
appendix | ✓ |
| ZONOS2 | zonos2 |
✓ | |
| PlayDiffusion | playdiffusion |
main | ✓ |
| Bark Voice Clone | bark_voice_clone |
✓ | |
| OZSpeech | ozspeech |
main | ✓ |
| VibeVoice | vibevoice |
main | ✓ |
One model, step by step¶
- Pick the model's config under
src/rvcbench/configs/ots_vc/clean/, for exampleots_vc/clean/libritts/qwen3_tts_ots. - Create and activate the model's environment from
envs/; see model environments for the map. - Install the model's runtime and download its checkpoints (notes below).
- Point local paths at them with Hydra overrides such as
adversary.code_path=...oradversary.checkpoint_path=.... - Run it:
rvcbench run --config-name <model_config> \
dataset.speaker_id=<speaker_id> \
adversary.max_samples=<n>
rvcbench run is the same as python run_vc.py in a source checkout; both take Hydra arguments.
Results go to results/<run_name>/<timestamp>/ (see run guide).
Examples¶
# Qwen3-TTS
python -m pip install -e '.[qwen3]'
huggingface-cli download Qwen/Qwen3-TTS-12Hz-1.7B-Base --local-dir checkpoints/Qwen3-TTS-12Hz-1.7B-Base
rvcbench run --config-name ots_vc/clean/libritts/qwen3_tts_ots \
dataset.speaker_id=1089 \
adversary.checkpoint_path=checkpoints/Qwen3-TTS-12Hz-1.7B-Base
# FishSpeech
git clone https://github.com/fishaudio/fish-speech.git checkpoints/fish_speech_s1
git -C checkpoints/fish_speech_s1 checkout d3df50503b36314a964f66cac1af1e19e95bcfa3
python -m pip install -e checkpoints/fish_speech_s1
huggingface-cli download fishaudio/s1-mini --local-dir checkpoints/fish_speech/openaudio-s1-mini
rvcbench run --config-name ots_vc/clean/vctk/fishspeech_ots \
dataset.speaker_id=p226 \
adversary.code_path=checkpoints/fish_speech_s1 \
adversary.llama_checkpoint_path=checkpoints/fish_speech/openaudio-s1-mini \
adversary.decoder_checkpoint_path=checkpoints/fish_speech/openaudio-s1-mini/codec.pth
# FireRedTTS-2 and VoxCPM
rvcbench run --config-name ots_vc/clean/libritts/fireredtts2_ots
rvcbench run --config-name ots_vc/clean/libritts/voxcpm_ots
# dots.tts (env: dots-tts) and MOSS-TTS (env: moss-tts)
rvcbench run --config-name ots_vc/clean/libritts/dots_tts_ots device=cuda:<gpu>
rvcbench run --config-name ots_vc/clean/libritts/moss_tts_ots device=cuda:<gpu>
# ZONOS2 (uv-managed environment; run with its own interpreter, GPU pinned with CUDA_VISIBLE_DEVICES)
CUDA_VISIBLE_DEVICES=<gpu> checkpoints/ZONOS2-repo/.venv/bin/python run_vc.py \
--config-name ots_vc/clean/libritts/zonos2_ots
More overrides:
rvcbench run --config-name ots_vc/clean/libritts/fireredtts2_ots adversary.max_samples=20 dataset.speaker_id=1089
rvcbench run --config-name ots_vc/clean/libritts/voxcpm_ots adversary.local_files_only=true adversary.cache_dir=/path/to/hf-cache
Quickstart model setup has the exact commands for the quickstart models, including gated downloads.
Model-specific notes¶
- Qwen3-TTS accepts the Hugging Face ID
Qwen/Qwen3-TTS-12Hz-1.7B-Baseor a local directory viaadversary.checkpoint_path, and needs theqwen-ttspackage (.[qwen3]extra).envs/qwen3-tts.ymlpins the exact versions used for the published runs. - FishSpeech needs a checkout of
fishaudio/fish-speechand thefishaudio/s1-minicheckpoint, passed withadversary.code_path,adversary.llama_checkpoint_pathandadversary.decoder_checkpoint_path. - Fish Audio S2 (paper) needs a separate S2-compatible checkout and
fishaudio/s2-pro. Keep the pinned S1 checkout separate: newer upstream tokenizer code is incompatible with the released S1-mini tiktoken files. The opt-in stable HTTP codec variant produced identical waveforms across two service instances; the original HTTP path has a cold/warm numerical difference. See native setup and HTTP setup. - FireRedTTS-2 expects the upstream checkout at
checkpoints/FireRedTTS2and weights undercheckpoints/FireRedTTS2/pretrained_models/FireRedTTS2by default. - VoxCPM defaults to
openbmb/VoxCPM2; setadversary.local_files_only=true(and optionallyadversary.cache_dir) to load offline. - dots.tts (
rednote-hilab/dots.tts-soar) needs its own environment (envs/dots-tts.yml, which lists the install order; it requirestorch>=2.8.0).device: cuda:Nis honoured. - MOSS-TTS (
OpenMOSS-Team/MOSS-TTS-v1.5,transformerswithtrust_remote_code) needsenvs/moss-tts.yml; install its packages in the order listed there. - ZONOS2 (
Zyphra/ZONOS2) is auvproject (envs/zonos2.ymllists the clone anduv syncsteps). Its scheduler ignoresdevice: cuda:N, so pin the GPU withCUDA_VISIBLE_DEVICES; it pre-allocates a large KV cache (about 55 GB on an 80 GB A100), so run it alone on its GPU. - dots.tts, ZONOS2 and MOSS-TTS pass the sample's
target_languageto the model, which matters for the AISHELL, French and cross-lingual configs. - Fish Audio S2 (server) and Higgs TTS 3 call a local inference server instead of loading weights; see below.
Server-backed models¶
These adapters send requests to a local HTTP server (adversary.endpoint_url). Start the server, then run
the config from an environment with python -m pip install -e '.[http]'.
Fish Audio S2 (fishaudio/s2-pro, Fish Speech /v1/tts API):
git clone https://github.com/fishaudio/fish-speech.git checkpoints/fish_speech
conda env create -f envs/fish-speech-s2.yml
conda activate fish-speech-s2
uv pip install 'numba==0.63.1' 'llvmlite==0.46.0'
cd checkpoints/fish_speech && uv pip install -e '.[cu126]' && uv pip install 'protobuf>=6.31.1,<7' && cd ../..
hf download fishaudio/s2-pro --local-dir checkpoints/s2-pro
CUDA_VISIBLE_DEVICES=<gpu> python checkpoints/fish_speech/tools/api_server.py \
--llama-checkpoint-path checkpoints/s2-pro \
--decoder-checkpoint-path checkpoints/s2-pro/codec.pth \
--listen 0.0.0.0:8001 --half
# in another shell:
rvcbench run --config-name ots_vc/clean/libritts/fish_audio_s2_ots
Higgs TTS 3 (bosonai/higgs-tts-3-4b, vLLM-Omni /v1/audio/speech API):
conda env create -f envs/vllm-omni-cu129.yml
conda activate vllm-omni-cu129
uv pip install --torch-backend=cu129 --extra-index-url https://wheels.vllm.ai/0.24.0/cu129 vllm==0.24.0
uv pip install --torch-backend=cu129 vllm-omni==0.24.0
hf download bosonai/higgs-tts-3-4b
CUDA_VISIBLE_DEVICES=<gpu> vllm-omni serve bosonai/higgs-tts-3-4b \
--host 0.0.0.0 --port 8000 --trust-remote-code --omni --allowed-local-media-path "$(pwd)"
# in another shell:
rvcbench run --config-name ots_vc/clean/libritts/higgs_tts3_ots
Both configs set device: cpu for the client process and evaluation.device: cuda:0 for scoring; the
server's GPU is chosen with CUDA_VISIBLE_DEVICES.
Known metric gaps in some model environments¶
These come from package versions in the model environments, not from the models (see the
emotion_pairs and speechmos_pairs fields of a run's metrics.json):
- dots-tts, moss-tts: no emotion scores; the emotion recognizer needs
AutoModelWithLMHead, which thetransformersversion these models require has removed. - zonos2: no SpeechMOS or emotion scores;
torchcodecfails to load its native library with this environment's torch and FFmpeg. WER, MCD, SIM and DNSMOS are unaffected.
Scoring the saved audio from a separate evaluation environment (+vc.evaluate_only=true, see the
run guide) avoids these gaps.
Checkpoints¶
A bundle of the supported models' code and checkpoints is available here (about 58 GB). For a single model, cloning only that model's repository is faster. Check the paths in the model's config before running.
Protection and denoising¶
The protection pipeline perturbs reference audio, clones from the protected references, and optionally denoises them first:
rvcbench protect --config-name safespeech_on_libritts # protect and measure fidelity
rvcbench run-protected --config-name ots_vc/protection/safespeech/ozspeech_ots \
protected_audio_dir=results/safespeech_on_libritts/<timestamp>/protected_audio
rvcbench denoise --config-name denoise/denoiser_dns64_on_protected_libritts_spec
Protection configs: grnoise_on_libritts (Gaussian noise), spec_on_libritts (SPEC),
safespeech_on_libritts (SafeSpeech), em_on_libritts (the paper's POP) and enkidu_on_libritts (Enkidu).
SPEC and SafeSpeech need surrogate-model checkpoints; see quickstart model setup.
Quickstart scripts and notebooks¶
| Example | Notebook | Script |
|---|---|---|
| Qwen3-TTS on LibriTTS | notebooks/rvcbench_qwen3tts_quickstart.ipynb |
scripts/run_qwen3tts_quickstart.py |
| FishSpeech on VCTK | notebooks/rvcbench_fishspeech_quickstart.ipynb |
scripts/run_fishspeech_quickstart.py |
| Fish Audio S2 on VCTK | — | scripts/run_fishspeech_s2_quickstart.py |
| Gaussian noise or SafeSpeech protection, then Qwen3-TTS | notebooks/rvcbench_safespeech_qwen3tts_quickstart.ipynb |
scripts/run_protect_qwen3tts_quickstart.py |
The scripts download the selected speakers from the Hugging Face dataset unless --no-hf-download is
passed; --help lists their options. python scripts/validate_quickstarts.py checks their commands and
layouts without running a model.