Skip to content

Native Bark voice cloning

The measured implementation uses the serp-ai Bark fork at commit 3b567365f650ee481c52dc5c32e55d4fddf5b6d6, native Fairseq HuBERT layer 9, the released learned semantic tokenizer, and Encodec reference codes. All assets must exist locally before model preparation. Missing files fail explicitly; the adapter does not substitute reference tokens or download weights.

Create the environment from envs/bark-native.yml, then install Fairseq:

cd envs
conda env create -f bark-native.yml
conda activate bark-native
cd ..
python -m pip install fairseq==0.12.2 --no-deps
git clone https://github.com/serp-ai/bark-with-voice-clone.git checkpoints/bark-with-voice-clone
git -C checkpoints/bark-with-voice-clone checkout 3b567365f650ee481c52dc5c32e55d4fddf5b6d6

Fairseq declares older Hydra/OmegaConf constraints. The measured runtime uses Hydra 1.3.2 and OmegaConf 2.3.0 through native legacy task/model factories. The recipe is an installation starting point, not a tested clean installation lock. The measured environment was an isolated overlay on an existing Python 3.10 environment. The factory control checks all 214 state tensors and layer-9 features on one second of reference audio; it does not establish paper scores.

Supply these assets through the YAML defaults or explicit adversary.* overrides:

Option Required local asset
models_dir Directory containing text_2.pt, coarse_2.pt, fine_2.pt from suno/bark
hubert_checkpoint hubert_base_ls960.pt from https://dl.fbaipublicfiles.com/hubert/hubert_base_ls960.pt
hubert_tokenizer quantifier_hubert_base_ls960_14.pth from GitMylo/bark-voice-cloning revision c26e70f3311c6973ca86511dd18b6a8ee073e830
text_tokenizer_path Local tokenizer directory for google-bert/bert-base-multilingual-cased, revision 3f076fdb1ab68d5b2880cb87a0886f315b8146f8

The multilingual tokenizer needs vocab.txt, tokenizer.json, tokenizer_config.json, and config.json. Provision Encodec's encodec_24khz-d7cc33bc.th in the Torch Hub checkpoint cache before an offline run. The audit records actual asset hashes; the three Bark weights were reused read-only from existing local assets, without a claimed immutable Hub revision. The quantizer filename's 14 does not select HuBERT layer 14: native reference encoding uses layer 9. LoRA weights require a separate runtime and are rejected. Duplicate compiled prefix keys also fail before native key normalization. Text/coarse/fine checkpoints must contain complete finite state with matching tensor shapes. Extra or missing .attn.bias tensors are accepted only for a known native attention module and an exact lower-triangular causal mask of the configured size. Each component records strict loading and validated mask hashes in its initialization receipt.

The runner uses explicit single-sample requests with original-index seeds. Missing reference audio, empty target text and invalid mono audio fail the sample; the runner records the error and continues. Reference transcripts are log metadata; native conditioning uses HuBERT/quantizer and Encodec reference tokens. Timing bark_reference_encoding_semantic_coarse_fine_and_codec_excluding_output_write_v1 includes reference encoding and all synthesis stages, excluding final WAV writing. Direct-request/loading contracts are checked. Current native CPU generation and core scoring of the frozen subset require separate validation.

Run the frozen subset after placing assets at the default paths:

python run_vc.py --config-name ots_vc/clean/libritts/bark_voice_clone_ots \
  run_name=bark_libritts16 dataset.use_hf_dataset=false \
  +dataset.manifest_filename=reproduction/subsets/libritts16_v1/metadata.json \
  +vc.generate_only=true +seed=42

Score saved outputs using the evaluation-only command in a separate evaluation environment. Subset validation does not establish historical paper equivalence.