Fixed Hub revisions and offline reproduction¶
Revision resolution preserves an explicit full 40-character lowercase Git
commit without a separate model_info request. Branches, tags, abbreviated commits
and unspecified revisions are resolved online. Offline use of such mutable
references fails with an instruction to pin a full commit first. This applies
to Qwen3-TTS, F5-TTS checkpoints/vocoders, PlayDiffusion presets and
OzSpeech's downloaded codec assets, MaskGCT's primary and semantic checkpoints, XTTS-v2 and ZipVoice.
Snapshot downloads may still contact the Hub when online.
The pin identifies a version; it does not prove that the required files exist in the local cache. A missing file still fails during snapshot or model loading. The Hub's download documentation describes full commit revisions and versioned snapshots.
For Qwen3-TTS Hub checkpoints, the runner resolves the pinned snapshot to a local
directory before loading the upstream wrapper. The installed upstream wrapper
forwards loading options to its model, but loads its processor separately
without forwarding revision. Using the local snapshot keeps both components
on the selected version and avoids the processor's repo metadata request in
offline mode. No shared upstream package is modified.
The runner records the Hub repo and commit under model_reference.assets,
alongside hashes of supported asset files in the snapshot. Its effective
adversary.checkpoint_path is the local snapshot directory; the original
configured repo remains in the saved run configuration. Explicit local
checkpoints continue to use their actual file hashes.
With the model already cached, a fixed LibriTTS subset can be generated using:
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python run_vc.py \
--config-name ots_vc/clean/libritts/qwen3_tts_ots \
dataset.use_hf_dataset=false \
+dataset.manifest_filename=/absolute/path/to/reproduction/subsets/libritts16_v1/metadata.json \
adversary.revision=fd4b254389122332181a7c3db7f27e918eec64e3 \
adversary.max_samples=16 +vc.generate_only=true +seed=42
This is an example revision used in the retained reproduction artifacts, not a claim that every upstream revision supports the same cloning interface.
F5-TTS assets¶
F5-TTS resolves checkpoint and vocoder revisions independently. The runner
supplies local versioned assets to the upstream API, records their hashes,
and hashes the installed preset configuration and default vocabulary.
Explicit ckpt_file, vocab_file and vocoder_local_path overrides are retained.
Checkpoint paths preserve the logical .safetensors filename in Hub caches;
resolving the symlink to a bare blob filename can select the wrong loader.
With both snapshots cached, the validated default preset can run offline:
HF_HUB_OFFLINE=1 python run_vc.py \
--config-name ots_vc/clean/libritts/f5_tts_ots \
dataset.use_hf_dataset=false \
+dataset.manifest_filename=reproduction/subsets/libritts16_v1/metadata.json \
adversary.revision=84e5a410d9cead4de2f847e7c9369a6440bdfaca \
adversary.vocoder_revision=0feb3fdd929bcd6649e0e7c5a688cf7dd012ef21 \
+vc.generate_only=true +seed=42
These pins apply to F5TTS_v1_Base and its Vocos vocoder. Another preset or
vocoder needs compatible revisions or explicit local assets. Current asset
provenance does not reconstruct missing checkpoint identities from legacy runs.
MaskGCT checkpoints¶
The runner resolves adversary.repo_id and adversary.revision into an explicit
adversary.checkpoint_dir containing all six native checkpoint files. It records
their hashes before starting the worker; missing or empty files fail asset
resolution. The worker loads these local paths without downloading replacement
primary weights. Explicit local checkpoint directories bypass Hub resolution.
Offline resolution requires a full commit and all six files in the cache. The
cached revision 265c6cef07625665d0c28d2faafb1415562379dc has been resolved and
hashed locally. This verifies asset availability, not native generation after
the backend migration or equivalence to historical weights.
The runner also resolves adversary.semantic_revision for
facebook/w2v-bert-2.0 to adversary.semantic_model_path, containing its model
configuration, feature-extractor configuration and safetensors weights. Both
upstream semantic loaders use this same snapshot with local_files_only=True.
The redirection is scoped to startup inside the isolated worker and restores
the original loading methods afterward. Upstream files are not modified.
Explicit local semantic snapshots bypass Hub resolution and must contain all
three required files.
The cached semantic revision da985ba0987f70aaeb84a80f2851cfac8c697a7b has been
hashed and loaded with the native SDK on CPU without missing, unexpected or
mismatched weights. Its feature extractor matches the currently cached default
processor on all 16 frozen references. These checks do not establish complete
MaskGCT generation after migration or historical weight equivalence.
For offline runs, provide both full revisions after caching their required files:
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python run_vc.py \
--config-name ots_vc/clean/libritts/maskgct_ots \
dataset.use_hf_dataset=false \
+dataset.manifest_filename=reproduction/subsets/libritts16_v1/metadata.json \
adversary.revision=265c6cef07625665d0c28d2faafb1415562379dc \
adversary.semantic_revision=da985ba0987f70aaeb84a80f2851cfac8c697a7b \
+vc.generate_only=true +seed=42
The local dataset root must contain the referenced audio. Configure the
MaskGCT worker environment through MASKGCT_RUNTIME_PYTHON or
adversary.runtime_python before running. The runner hashes local semantic normalization statistics when
code_path is configured. It also hashes the G2P source/data files, including
the eagerly imported Chinese ONNX model and dictionaries. An eSpeak probe uses
the worker interpreter and native-library search path; its version, selected
shared library and complete data directory are recorded, including compiled
dictionary files without extensions. Before model loading, the worker recomputes
the eSpeak binary/data hashes and checks them against the resolved configuration.
The parent also checks the startup receipt; changed resources or a missing
receipt reject startup and reap the worker. Native English tokenization has been
checked on the 32 prompt/target texts in the frozen 16-pair subset. Other-language
package resources and a clean environment lock remain incomplete; these checks
do not establish full generation or historical equivalence.
XTTS-v2 checkpoints¶
The runner resolves adversary.checkpoint and adversary.revision to a local
snapshot before creating XTTS. It validates and hashes the native model, JSON
configuration and vocabulary; a preset speaker bank is also captured when
present. The native cloning path supports checkpoints without a preset speaker
bank. Explicit file overrides are validated, hashed and passed to the loader.
Missing or empty required files fail before loading weights.
The cached revision below resolves offline, and its four consumed file hashes match the retained LibriTTS16 run. Migrated native generation and scoring remain pending; these asset checks do not establish output equivalence or a complete runtime environment lock.
HF_HUB_OFFLINE=1 python run_vc.py \
--config-name ots_vc/clean/libritts/xtts_ots \
dataset.use_hf_dataset=false \
+dataset.manifest_filename=reproduction/subsets/libritts16_v1/metadata.json \
adversary.revision=6c2b0d75eae4b7047358e3b6bd9325f857d43f77 \
adversary.local_files_only=true +vc.generate_only=true +seed=42
Standalone XttsGeneratorConfig accepts the same revision and
local_files_only options. Explicit local checkpoint directories bypass Hub
resolution.
ZipVoice assets¶
The runner resolves the ZipVoice model and Vocos vocoder independently, using
adversary.revision and adversary.vocoder_revision. The selected
zipvoice or zipvoice_distill directory contains the configured checkpoint,
model.json and tokens.txt. The vocoder contains config.yaml and
pytorch_model.bin. Missing or empty required files fail before loading models.
Their local paths and hashes are recorded, and both native and CLI inference
receive those explicit directories. Local model_dir, checkpoint_name and
vocoder_path overrides are retained.
The following cached model and vocoder snapshots resolve offline; all five
consumed file hashes match the retained Emilia-tokenizer LibriTTS16 run. This
does not establish migrated native generation equivalence or freeze all
tokenizer/package resources. The model pin below has been checked for
zipvoice; another model variant needs its own compatible cached files.
HF_HUB_OFFLINE=1 python run_vc.py \
--config-name ots_vc/clean/libritts/zipvoice_ots \
dataset.use_hf_dataset=false \
+dataset.manifest_filename=reproduction/subsets/libritts16_v1/metadata.json \
adversary.revision=4ed45fb6e7e9527b780bef9e097a04bf13fe4e6b \
adversary.vocoder_revision=0feb3fdd929bcd6649e0e7c5a688cf7dd012ef21 \
adversary.local_files_only=true +vc.generate_only=true +seed=42
Provide the native ZipVoice source via adversary.code_path. Pin resolution
applies to the benchmark runner. Standalone ZipVoiceGenerator callers need
explicit local model_dir and vocoder_path to avoid implicit upstream
downloads; both runtimes recheck configured local required files before model
startup.
SparkTTS local assets¶
SparkTTS's relative model_dir resolves under its code_path, matching the
native loader even if a different directory with the same name exists in the
current working directory. The runner records the resolved absolute directory
and hashes its assets. Before native imports, both runner and generator check
configuration, BiCodec weights, LLM weights/tokenizer and wav2vec2
weights/feature-extractor files. Standard single-file or indexed Hugging Face
weight layouts are accepted; every referenced shard must be present and nonempty.
Native BiCodec loading rejects missing learned state, unexpected checkpoint
entries and shape mismatches. Only the two configuration-derived registered
mel buffers, mel_transformer.spectrogram.window and
mel_transformer.mel_scale.fb, may be absent. A missing entry is accepted only as a registered non-trainable
finite buffer. Their shapes, dtypes and value hashes are recorded in each
successful sample's native_initialization receipt and retained during reuse.
The scoped check restores the native class method after startup or failure;
upstream source files remain unchanged.
The current local bundle's 14 asset hashes match the retained LibriTTS16 run.
Actual native BiCodec CPU loading accepts its 840 checkpoint entries with only
those two derived buffers absent, and rejects removal of a learned encoder
tensor in memory. This component check used the existing audiobench runtime;
the default controller environment lacks einx. It does not establish full
SparkTTS generation after migration, LLM/wav2vec loading completeness, a clean
runtime lock or immutable historical weights. Use a compatible SparkTTS runtime
with a complete explicit local model bundle for full generation.
CosyVoice local assets¶
CosyVoice requires an explicit local model_dir. Discovery stays inside that
directory, and the runner binds and hashes the effective payload directory.
Runner and generator reject missing or empty core checkpoints, the selected
variant's YAML configuration and speech tokenizer before importing the native
SDK. A configured zero_shot_spk_id also requires a nonempty spk2info.pt bank;
its actual speaker membership is still checked by native inference.
CosyVoice2's native constructor overrides its Qwen pretrained path to
CosyVoice-BlankEN under the payload directory. The preflight therefore also
requires its configuration, vocabulary, merges and weights. Standard single-file
or indexed Transformers weights are accepted, with every declared shard present
and nonempty. The generator rechecks these files at model loading. CosyVoice1
receives only the arguments accepted by its constructor; load_vllm is restricted
to CosyVoice2.
These checks validate local file completeness, not tensor loading or speech quality. CosyVoice1's YAML-selected text assets, custom YAML dependencies, JIT/TensorRT/vLLM assets and text-frontend resources need further runtime evidence. The previously measured native subset uses CosyVoice2 with accelerators disabled; post-migration native subset generation remains pending.
OpenVoice converter loading¶
OpenVoice's native converter uses strict=False. The wrapper requires a complete
finite checkpoint with matching tensor shapes instead. Missing or unexpected
keys, shape errors and nonfinite tensors fail startup. The check is scoped to
the converter's owned Torch model; other Torch modules retain their behavior,
and the original method is restored after success or failure. Startup must
perform exactly one checked load. Successful rows record the converter's key
count, strictness and compatibility result in native_initialization, retained
on resume and saved-audio scoring.
Actual native converter CPU startup accepts the local 486-entry checkpoint with
no missing or unexpected state. Its 32,792,226 parameters load without changing
any checkpoint tensor, and removing a learned decoder parameter in memory is
rejected. This check used the existing audiobench environment with a writable
Numba cache and bypassed the Melo import stage to isolate the converter.
This verifies the converter component only. Melo checkpoints, BERT/text assets, speaker embeddings and post-migration native subset generation still require separate evidence. A compatible full OpenVoice/Melo runtime is required; loading the standalone converter architecture does not validate its audio API imports.
OpenVoice Melo language models¶
The LibriTTS OpenVoice configuration pins the English Melo model to
myshell-ai/MeloTTS-English revision
bb4fb7346d566d277ba8c8c7dbfdf6786139b8ef. The runner resolves its configuration
and checkpoint from that same snapshot, checks both files are nonempty, records
their hashes, and supplies their explicit local paths to the native constructor.
adversary.local_files_only=true uses an already cached pinned snapshot.
Configure additional uppercase language entries in adversary.melo_models
with repo_id and revision, or both config_path and checkpoint_path for
explicit local files. Once this mapping is configured, an absent language fails
rather than falling back to an implicit download. Standalone generator calls
require resolved local paths in the mapping. Older configurations without a
mapping retain the native implicit asset behavior as a separate unpinned path.
The actual pinned English architecture accepts all 1,051 checkpoint entries
with no missing or unexpected state; its 51,874,097 parameters match every
checkpoint tensor and all tensors are finite. This CPU component check excludes
Melo's text API and speech inference. The earlier native subset did not record
Melo asset hashes, so it does not establish historical checkpoint identity.
Transformer text assets use the separate text_models mapping described below.
Full native text and speech validation remains pending.
OpenVoice English pronunciation resources¶
The LibriTTS configuration enables
adversary.capture_english_text_resources=true. Before model loading, the runner
records Melo's CMU dictionary and its existing pickle cache, the trained
g2p_en parameter archive, and NLTK's CMU/tagger resources. It includes both
archives checked by g2p_en during import and the active corpus/tagger used by
the installed NLTK version. Directory resources include extensionless files.
Missing resources fail preflight; this path does not call a downloader.
These are resources discovered in the generation interpreter. Configure
melo_code_path for the intended checkout and install the needed NLTK data in
that runtime before running. Matching hashes identify the observed resource
files; they do not pin their installation source or guarantee a clean runtime.
The local inventory records seven resource entries and ten files. BERT and
its tokenizers, non-English resources and migrated speech generation remain
outside this check.
OpenVoice Transformer text assets¶
The LibriTTS configuration lists fixed snapshots in adversary.text_models,
keyed by the upstream repository IDs used by Melo. Melo eagerly imports six
language tokenizers even for English generation. The mapping provides these
tokenizers and enables weights only for English bert-base-uncased with
load_model: true. The runner validates and hashes the selected files; cached
weights for other languages are not included unless enabled.
During native imports and text inference, scoped loading redirects configured
repository IDs to their resolved local snapshot paths with local_files_only
and remote custom code disabled. Unlisted IDs and weight loading on a
tokenizer-only entry fail explicitly. Loader methods are restored after success
or failure. Configurations without text_models retain the separate legacy
implicit-loading path. Standalone generator calls require resolved local paths.
This fixes asset selection, not Transformer version compatibility or all text resources. The observed runtime also emits a tokenizer regex warning for the multilingual snapshot. Actual native English G2P, phone/tone/language tensors and expanded BERT features have repeated identically for the 16 selected target texts on CPU. Historical repo-ID tokenization equivalence and speech reproduction remain unproven.
OpenVoice tracks Melo text module objects imported by its own serial operations. On close, it releases their native global Torch model caches as well as its language models and reference embeddings. A replacement module or a previously imported foreign module is not selected for cleanup. Partial import failures still record owned modules for cleanup. The native LibriTTS16 text check verifies that closing the generator releases the actual English BERT cache without changing any of the checked text feature hashes.