Install RVCBench¶
The pip package supports two workflows: automatic metrics for your own audio, and dataset-backed model evaluation. Neither requires cloning this repository.
Recommended installation¶
Use Python 3.10–3.13 on Linux. Start in a separate environment:
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "rvcbench[eval]"
Install FFmpeg through your operating system if it is missing. On Ubuntu/Debian:
Download the scoring models once:
This prepares all seven public metrics. Model downloads need several GB of disk space; Whisper medium is the largest download. For only speaker similarity, WER and MOS:
- Your own audio: follow the metrics API.
- Our datasets: follow Evaluate your model.
- Check the installation:
rvcbench doctor --eval --importschecks imports without model downloads. - Check downloaded models:
rvcbench setup-scorers --check-onlyverifies and loads them without downloads.
pip install rvcbench alone provides the runner, suite definitions and command line. Add [eval]
to install the speech-scoring dependencies. Built-in model runtimes are separate; the two workflows above
can score audio generated by any model, without installing that model inside RVCBench's environment.
CPU and GPU¶
CPU scoring works with device="cpu" in Python and --device cpu on the command line. To avoid
installing CUDA libraries in a CPU environment, install CPU PyTorch first:
python -m pip install torch==2.6.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cpu
python -m pip install "rvcbench[eval]"
For GPU scoring, use matching PyTorch and torchaudio versions compatible with your NVIDIA driver,
then install rvcbench[eval] and select cuda or cuda:0. The scoring extra currently supports
PyTorch 2.3–2.9; see model environments for the validated scoring environments.
Downloads and offline use¶
Scorer assets are checked against pinned source revisions and SHA-256 hashes. Speaker, emotion and
SpeechMOS assets use RVCBENCH_ASSET_DIR when set, otherwise a per-user cache under
~/.cache/rvcbench/ (existing checkout assets can also be used). Whisper uses ~/.cache/whisper/.
Set XDG_CACHE_HOME to relocate both default caches.
export RVCBENCH_ASSET_DIR="$HOME/.cache/rvcbench"
rvcbench setup-scorers
rvcbench setup-scorers --check-only
Legacy SpeechMOS Torch Hub files are reused during setup only if their hashes match the release pins. A mismatched file is reported and preserved; move it aside and rerun setup to fetch the pinned version.
Dataset audio is downloaded from Hugging Face when you run rvcbench prompts or rvcbench score.
Only the suite's files are requested. To use an existing dataset copy, pass --data-root to both
commands; see the suite guide. Log in with hf auth login if downloads are rate-limited.
After preparing the assets and data, scoring can run offline.
Upgrade from 2.0.0¶
python -m pip install --upgrade "rvcbench[eval]==2.2.1"
rvcbench setup-scorers
rvcbench doctor --eval --imports
Use a new results directory for the first 2.1.0 evaluation. Earlier reports remain readable, but scorer
fingerprints have changed: rescore all models in the same environment before comparing them. --resume
requires the request journal written by 2.1.0. Keep your existing generated WAV files; they can be scored again.
Troubleshooting¶
| Symptom | Action |
|---|---|
| Missing Whisper, SpeechBrain or another scorer dependency | Install rvcbench[eval] in the interpreter that runs the command. |
| PyTorch/torchaudio binary import error | Install matching versions of torch and torchaudio from the same CPU/CUDA wheel index. |
| NumPy/Numba conflict | Reinstall the evaluation extra in a fresh environment; it requires NumPy below 2.3. |
pysptk cannot build |
See building evaluation extras for compiler prerequisites. |
| Scorer asset missing | Run rvcbench setup-scorers --metrics <metric> in the same cache environment. |
ffmpeg not found |
Install the operating-system FFmpeg package and ensure it is on PATH. |
| An evaluation stopped halfway | Repeat rvcbench score with the same arguments and --resume. |
| Existing reports have incompatible scoring fingerprints | Rescore both models in one environment; use compare --allow-incompatible only for unranked inspection. |