Skip to content

Model Environments

RVCBench integrates many third-party voice cloning and TTS models. These projects often pin incompatible versions of PyTorch, Transformers, ONNX Runtime, tokenizers, or model-specific helper packages, so one global Python environment is not expected to run every model.

The model environment specs are maintained in this repository under envs/. Use the environment that matches the model you are launching.

Creating an Environment

The Qwen environment installs only the core package and Qwen runtime. Other base templates install the core package and may require a separate upstream checkout or package. They are installation starting points, not tested lock files. Run from envs/ so editable package paths resolve correctly:

cd envs
conda env create -f qwen3-tts.yml
conda activate qwen3
cd ..

Then run the corresponding benchmark command, for example:

python scripts/run_qwen3tts_quickstart.py --max-samples 5

To update an existing environment:

cd envs
conda env update -f qwen3-tts.yml --prune
cd ..

Model-to-Environment Map

Model family Environment file Conda environment name
General benchmark / fallback envs/audiobench.yml audiobench
VoxCPM2 envs/voxcpm.yml (base recipe; requires the pinned upstream checkout) voxcpm
FireRedTTS2 envs/fireredtts2.yml (native monologue canary validated; requires pinned upstream checkout) fireredtts2
BertVITS2 / SafeSpeech surrogate envs/bertvits2.yml bertvits2
CosyVoice envs/cosyvoice.yml cosyvoice
dots.tts envs/dots-tts.yml dots-tts
Fish Audio S2 native envs/fish-s2-native.yml (setup; base recipe) fish-s2-native
F5-TTS envs/f5-tts.yml (validated Python 3.11 / Torch 2.8 / SDK 1.1.18 pins; base recipe) f5-tts
Fish Audio S2 (local API server) envs/fish-speech-s2.yml (base env; serves the model, see README) fish-speech-s2
FishSpeech envs/fishspeech.yml fishspeech
GLM-TTS envs/glm-tts.yml glm-tts
GlowTTS envs/glowtts.yml glowtts
Higgs Audio envs/higgs-audio.yml higgs-audio
Higgs TTS 3 envs/vllm-omni-cu129.yml (base env; serves the model, see README) vllm-omni-cu129
IndexTTS envs/indextts.yml indextts
Kimi Audio envs/kimi-audio.yml kimi-audio
MaskGCT envs/maskgct.yml maskgct
MGM-Omni envs/mgm-omni.yml mgm-omni
MOSS-TTS envs/moss-tts.yml moss-tts
MOSS-TTSD (native MossTTSD checkpoint) envs/moss-ttsd.yml moss-ttsd
MOSS-TTSD (legacy Asteroid checkpoint) envs/moss.yml moss
OpenVoice envs/openvoice.yml openvoice
OZSpeech envs/ozspeech.yml ozspeech
PlayDiffusion envs/playdiffusion.yml playdiffusion
Qwen3-Omni envs/qwen3-omni.yml qwen3-omni
Qwen3-TTS envs/qwen3-tts.yml qwen3
Spark-TTS envs/sparktts.yml sparktts
StyleTTS2 envs/styletts2.yml styletts2
VALL-E envs/vall-e.yml vall-e
Bark Voice Clone envs/bark-native.yml (native LibriTTS16 generated/scored; base recipe; setup) bark-native
Amphion VALL-E (separate implementation) envs/amphion-valle.yml (native LibriTTS16 generated/scored; base recipe) amphion-valle
VibeVoice envs/vibevoice.yml vibevoice
XTTS-v2 envs/xtts-v2.yml xtts-v2
ZipVoice envs/zipvoice.yml zipvoice
ZONOS2 not a conda env — see envs/zonos2.yml (uv-managed checkpoints/ZONOS2-repo/.venv) —

Reproducibility Options

The Qwen quickstart and CPU smoke path are covered by offline CI. The other model families have adapter integrations and base environment templates; they have not all been revalidated with current upstream packages. Dependencies are declared in pyproject.toml; follow each model runtime setup before using its adapter.

Generation-only usage needs the core plus model runtime. Install .[eval] explicitly for full evaluation; run rvcbench doctor --model qwen3 --eval --imports to check dependency availability before launching. This check does not load models or certify CUDA compatibility. Training/protection runtimes need their own upstream dependencies.

For stronger reproducibility, generate platform-specific lock files from these YAML files with conda-lock, or publish prebuilt Docker/Apptainer images per model family. Containers are usually the most reliable option for CUDA-heavy third-party inference stacks, while Conda or Micromamba specs are easier for users who need to adapt paths, CUDA versions, or local checkpoint locations.

For models requiring a different Torch version, generate with +vc.generate_only=true in that model's environment and evaluate in a separate environment containing .[eval]. The Qwen, Enkidu and evaluation extras accept PyTorch 2.3–2.9 and keep an existing build; envs/qwen3-tts.yml and envs/enkidu.yml pin the exact versions used for the published runs. Other model base templates do not install the evaluation stack automatically.

Generation provenance records statically discovered unmapped_imports separately from distribution versions. These names may be authored upstream namespaces, missing modules or packages without distribution metadata. The inventory includes literal __import__() and import_module() calls; it does not resolve nonliteral imports or certify a complete environment lock.

For SparkTTS, run rvcbench doctor --model sparktts --imports in the chosen runtime before generation. It checks the main native dependencies, including einx, without loading weights. Dependency availability does not validate the upstream source tree, checkpoint compatibility, CUDA or full model inference.

OpenVoice's converter-only CPU check uses native configuration and weights, but its full audio API also needs a compatible NumPy/Numba/Librosa combination. Import the audio stack in the intended runtime before loading checkpoints. Read-only environment installations may require NUMBA_CACHE_DIR pointing to a writable local directory. This addresses cache placement; it does not repair NumPy/Numba version incompatibility or establish a clean OpenVoice/Melo lock.

Building the evaluation extras from source

pysptk (needed for MCD) is published only as a source package. Its build appends the commit of any Git repository that contains the build directory to its version number, which makes pip discard the build. If pip install -e '.[eval]' reports that no pysptk version matches, build it outside a Git checkout or pin the version explicitly:

PYSPTK_BUILD_VERSION=1.0.1 python -m pip install 'pysptk==1.0.1'

Where scorer models are stored

The speaker, emotion and DNSMOS scorers keep their model files in one directory per scorer: $RVCBENCH_ASSET_DIR/<name> when that variable is set, otherwise ./checkpoints/<name> when it exists (as in a source checkout), otherwise ~/.cache/rvcbench/<name>. A speaker model fetched from the Hugging Face Hub is copied there rather than linked into the Hub cache, so moving or clearing that cache does not break it.