Validation coverage¶
The repository exposes 32 model integrations. Currently, 27 have completed real generation and scoring on a fixed subset. This is integration validation; full paper-table and bitwise generation reproduction remain unestablished.
Paper membership refers to the 18 models of the paper's main results (arXiv v3). Experimental integrations remain in scope even when assets or compatible protocols are missing.
| Model | Paper | Subset | Generation / scoring |
|---|---|---|---|
| BertVITS2 | No | libritts16_v1 | Pending |
| Qwen3-TTS | Yes | libritts16_v1 | Complete |
| Qwen3-Omni | No | libritts16_v1 | Pending |
| FireRedTTS-2 | No | libritts_short10_legacy_v1 | Complete |
| VoxCPM | No | libritts16_v1 | Complete |
| F5-TTS | Yes | libritts16_v1 | Complete |
| MaskGCT | Yes | libritts16_v1 | Complete |
| OpenVoice V2 | Yes | libritts16_v1 | Complete |
| Coqui XTTS-v2 | Yes | libritts16_v1 | Complete |
| IndexTTS | Yes | libritts16_v1 | Complete |
| ZipVoice | Yes | libritts16_v1 | Complete |
| FishSpeech | Yes | libritts16_v1 | Complete |
| Fish Audio S2 (in-proc) | No | libritts16_v1 | Complete |
| Fish Audio S2 (server) | No | libritts16_v1 | Complete |
| CosyVoice / 2 | Yes | libritts16_v1 | Complete |
| Higgs Audio | Yes | libritts16_v1 | Complete |
| Higgs TTS 3 | No | libritts16_v1 | Pending |
| SparkTTS | Yes | libritts16_v1 | Complete |
| VALL-E | No | libritts16_v1 | Pending |
| StyleTTS 2 | Yes | libritts16_v1 | Complete |
| GLM-TTS | Yes | libritts16_v1 | Complete |
| GlowTTS | No | libritts16_v1 | Pending |
| Kimi Audio | No | libritts16_v1 | Complete |
| MGM-Omni | Yes | libritts16_v1 | Complete |
| MOSS TTSD | Yes | libritts16_v1 | Complete |
| MOSS-TTS | No | libritts16_v1 | Complete |
| dots.tts | No | libritts16_v1 | Complete |
| ZONOS2 | No | vctk16_v1 | Complete |
| PlayDiffusion | Yes | libritts16_v1 | Complete |
| Bark Voice Clone | No | libritts16_v1 | Complete |
| OZSpeech | Yes | libritts16_v1 | Complete |
| VibeVoice | Yes | libritts16_v1 | Complete |
Pending integrations¶
| Integration | Required before real subset validation |
|---|---|
| BERT-VITS2 | A trained reference-conditioned checkpoint and compatible native inference. The current closed-set speaker model cannot substantiate zero-shot cloning. |
| Qwen3-Omni | A complete checkpoint and a verified reference-speaker cloning mechanism. |
| Higgs TTS 3 | A compatible speech-generation service with the configured model and reference-audio support. |
| Native VALL-E | A checkpoint compatible with lifeiteng/vall-e. The validated Amphion variant has a separate configuration. |
| GlowTTS | Model assets and a validated wrapper implementing the intended speaker-conditioning protocol. |
GR seeded-noise reconstruction, DNS64 and Enkidu cohort production have been validated on the fixed LibriTTS subset. Enkidu completed the full 2,000-pair training cohort and Qwen3 generation/core scoring for 16 selected outputs. Its audio-only protocol differs from historical protection output and does not establish historical equivalence.
The Bark, FireRedTTS2, VoxCPM, IndexTTS, MaskGCT, XTTS, ZipVoice, SparkTTS, CosyVoice, OpenVoice and StyleTTS2 direct-backend migrations have contract checks, including explicit seed validation and failure propagation. IndexTTS and MaskGCT also have worker timeout, request matching and cleanup checks. Native generation of the other migrated frozen subsets remains pending. OpenVoice additionally completed all 16 current native CPU generation requests with matching frozen inputs, valid WAV/hash/seed receipts and matching recorded source hashes. Its MCD/WER/SIM scoring is complete for all 16 pairs, and a second fresh CPU generation matches every WAV byte. Historical comparison fails the scorer/runtime consistency check; GPU/full-paper equivalence remains unproven. StyleTTS2 also completed two fresh CPU generations of all 16 frozen samples with byte-identical WAVs and full MCD/WER/SIM scoring. All 13 constructed native modules load complete finite checkpoint state, and corrupted learned-weight checks fail explicitly. Its historical GPU scorer/runtime fingerprints also differ, so historical/full-paper equivalence remains unproven. FireRedTTS2 completed native CPU generation and MCD/WER/SIM scoring for all 10 frozen short10/retry20 pairs with row/CSV/aggregate checks. This explicit variant preserves its historical input and asset hashes; historical GPU scorer/runtime fingerprints differ, and native CPU repeatability has not yet been measured. These validations apply to each migration's recorded source revision; final combined-source native revalidation remains outstanding. The table above retains the earlier generation/scoring campaigns.
Other outstanding work includes remaining protection-method production, historical auxiliary metric replay, controlled timing campaigns across all runtimes, and final website/Hugging Face integration. These are retained in the reproduction plan.
Run records store sample hashes, assets, runtime provenance and metric coverage. Local debug runs and detailed audit snapshots are excluded from Git. Use the run guide to generate your own reports.
Check the source version of a retained run¶
This read-only command compares each recorded generation source hash with the
current file. Use --root /path/to/checkout for another checkout and
--output results/source_audit.json to save the report. A changed or missing
file, unavailable provenance or invalid recorded hash fails the check. Absolute
upstream paths are checked at their recorded locations.
A match covers only recorded files; it does not verify complete current import coverage, dependencies, weights, audio or metrics. A mismatch marks a source version difference and does not invalidate historical results. The table above records completed subset campaigns, rather than a guarantee for every later source revision. Qwen3 and F5 direct campaigns were run before subsequent shared runner changes; native revalidation of the final combined refactor remains outstanding.