Codebase versions¶
v1 and v2 are two versions of the code for the same benchmark: the datasets, metrics and paper results are the same. To reproduce the paper, use v1. To evaluate a new model, use v2.
tag v1.0 (2026-09-30): the code released with the paper
├── branch v1 v1.0 plus README changes only; frozen
└── branch main v1.0 plus the v2 refactor (2026-09-30 to 2026-10-02); under development
v1 (branch v1, tag v1.0) |
v2 (branch main) |
|
|---|---|---|
| Use it to | Reproduce the paper | Evaluate a new model |
| Status | Frozen at the paper release | Under active development |
| Install | Clone, then pip install a list of packages |
pip install the rvcbench package from GitHub |
| Run a built-in model | python run_vc.py --config-name ... |
rvcbench run --config-name ... (same config names) |
| Evaluate your own model | Add an adapter to the codebase | Score audio generated anywhere (rvcbench prompts, rvcbench score), or a one-file adapter |
| Evaluation data | Full datasets | Full datasets, plus the core-v1 suite: 480 pinned utterances covering 16 of the paper's 18 evaluations |
| Run output | metrics.json per run |
Per-sample run_manifest.json with input and output hashes, failures and metric coverage |
git clone --branch v1 https://github.com/Nanboy-Ronan/RVCBench.git RVCBench-v1 # reproduce the paper
git clone https://github.com/Nanboy-Ronan/RVCBench.git RVCBench # v2
v1 and v2 name versions of this code. They are unrelated to the paper's arXiv versions and to suite names such
as core-v1.
What changed in v2¶
| Area | Main code | Change |
|---|---|---|
| Package and commands | src/rvcbench/, benchmark/cli.py |
Installable rvcbench package with configs inside it; rvcbench run, prompts, score, setup-scorers and the run-record commands. |
| Suites | benchmark/submission.py, suites/ |
Versioned suites with pinned inputs; score audio generated anywhere. |
| Execution and reporting | benchmark/, workflows/vc.py |
Per-sample records of inputs, outputs, failures, metric coverage and run state; generation and scoring can run in separate environments. Completion and comparability are checked separately. |
| Model adapters | adversary/, models/ |
Explicit per-sample requests for 13 models, seeds from the original sample index, invalid conditioning or audio rejected, owned resources released. Other models still use the compatibility path. |
| Model workers | models/worker_protocol.py, scripts/ |
Bounded waits, request/response matching and process cleanup. |
| Checkpoints and dependencies | benchmark/model_assets.py, Hub revisions |
Explicit paths or pinned revisions, recorded hashes, strict state loading for selected models. |
| Data and reproduction | datasets/, reproduction/ |
Annotation variants kept in sample identity, validated source indices, frozen subsets with input hashes. |
| Protection, scoring and timing | benchmark/, run guide |
Traced reference stages and scorer provenance; timing scopes declared so incompatible runs are not compared. |
Behaviour changes. Malformed inputs or incomplete checkpoints that v1 skipped or patched now fail.
Model versions, seed policies and retry or conditioning variants must be recorded when comparing v1 and v2
results. v2 runs do not replace the paper's tables. To reproduce those, use v1 together with its environment
files (envs/) and the checkpoint bundle linked from its README; the code alone does not pin model weights.
Validation status¶
Real generation and core scoring on a fixed subset have been recorded for 27 of the 32 integrations, including the paper's 18 models. These runs span different stages of the refactor and do not establish reproduction on the current combined source or of the paper's tables. The other five integrations lack assets or a verified cloning protocol. See validation coverage and the reproduction plan.