Skip to content

Binding protected and denoised references

The zero-shot runner accepts vc.reference_audio_dir as an explicit directory of replacement reference waveforms. It keeps the selected sample IDs, target waveforms and transcripts from the clean manifest. vc.reference_stage names the protocol variant; the name does not attest how the audio was produced.

For example, append these overrides to a run_vc.py zero-shot command:

+vc.reference_audio_dir=/absolute/path/to/protected_audio
+vc.reference_stage=protected_grnoise_legacy

Each selected reference must resolve uniquely under that directory as <speaker>/<filename>, <filename>, or its manifest-relative path. Missing references, ambiguous paths and flat files shared by different clean identities fail before model loading. The dataset object is not modified. When a legacy wrapper also supplies a protected directory, it must agree with the explicit directory.

run_manifest.json records a reference_stage section with the clean reference and replacement reference paths and SHA-256 hashes for every sample. Its fingerprint enters the generation configuration. Resume and evaluation reject changed lineage, including changes to the clean reference when the replacement and target bytes remain unchanged.

An adjacent stage_manifest.json is hashed and recorded when present. This is recorded producer metadata; unknown producer formats do not attest their claims. A producer declaring an unfinished or failed status is rejected. For the known gr_archived_noise_replay_v1, gr_seeded_batch_rng_v1, dns64_dataset_rate_v1 and enkidu_audio_only_cohort_v1 formats, the runner additionally requires complete verification counts and checks selected sample identities and clean/output hashes against the producer rows. GR replay also requires historical output hashes; Enkidu requires complete cohort training. Such bindings are labeled verified_selected_output_hashes. Historical directories without a producer manifest are labeled legacy_directory_content_only; binding them does not reproduce the protection or denoising algorithm.

Produce reference audio

Use Enkidu cohort production, archived Gaussian-noise replay when the historical noise archive is available, or DNS64 production to denoise an explicit reference directory. These commands write a stage manifest that the cloning runner verifies against the selected inputs and output bytes.

Legacy protectors store their noise archive at the protection run root and write WAVs under <speaker>/<filename>. Enkidu requires batch size 1 and 16 kHz processing; its output WAVs retain that rate even when source audio is 24 kHz. Its surrogate optimization retains the historical gradient accumulation behavior. Validate full production separately from binding an existing protected directory.

SafeSpeech/SPEC/EM surrogate loading requires the canonical 112-symbol table and matching complete checkpoints. The default model.checkpoint_load_policy=strict rejects missing or differently shaped weights before applying them. An explicit model.checkpoint_load_policy=legacy_partial retains legacy initialization fallbacks and records the affected tensors; this is a distinct diagnostic protocol and does not establish trained-surrogate equivalence. Text symbol imports do not prepare language models or download checkpoints.

The training loader uses dataset.text_feature_policy=strict_cached by default: current-language .bert.pt features must be finite floating tensors with 1024 channels and the expected phoneme length. Missing or invalid caches fail instead of fabricating features. legacy_random explicitly restores the old fallback; non-current-language random channels retain the historical construction in both policies. A shape check does not establish a cache's tokenizer/model provenance.