Skip to content

Native Fish Audio S2 runtime

The S2 integration uses its own native checkout, separate from the historical S1 adapter. The source revision is 214da3cd841bda85da2496b96cd3c4d7edb1337e. Its native conversion maps the released fish_qwen3_omni sharded checkpoint into DualARTransformer weights. Both the main model and codec load strictly. The six legacy codec RoPE/causal-mask caches are accepted only when they are registered nonpersistent buffers and equal the corresponding region of the native reconstruction, including dtype. Missing trained weights remain errors.

git clone https://github.com/fishaudio/fish-speech.git checkpoints/fish-speech-s2-native
git -C checkpoints/fish-speech-s2-native checkout 214da3cd841bda85da2496b96cd3c4d7edb1337e
cd envs
conda env create -f fish-s2-native.yml
conda activate fish-s2-native
cd ..

Provision the released S2-Pro checkpoint in checkpoints/s2-pro, including the tokenizer, config, shard index, both safetensors shards, and codec.pth. Alternatively override adversary.llama_checkpoint_path and adversary.decoder_checkpoint_path to existing local assets. The environment recipe is a starting point, not a tested clean-install lock. The measured environment is a private overlay on an existing Python 3.12/Torch 2.8 install.

python run_vc.py --config-name ots_vc/clean/libritts/fishspeech_s2_ots \
  run_name=fish_s2_libritts16 dataset.use_hf_dataset=false \
  +dataset.manifest_filename=reproduction/subsets/libritts16_v1/metadata.json \
  +vc.generate_only=true +seed=42

Native reference loading, semantic generation and codec decoding run serially on the calling thread. This preserves native generation functions while avoiding the upstream worker's unbounded initialization wait and global Hydra reset. Resources belong to the adapter and are released on close; no shared endpoint is restarted. Use compile=false and half=false (BF16 main model). The native sampler's top_k default is 30. Its repetition-penalty request argument is accepted upstream but has no use in the current generation function; the adapter does not claim that changing it alters generation. The native engine also accepts normalize without applying a normalization transform.

Actual reference and target transcripts are required. Empty target text is an error, rather than a request to synthesize the reference sentence. Scoring runs separately in the evaluation environment, after generation resources close. The HTTP S2 integration remains a separate entry requiring a verified S2 server; see its measured HTTP protocol and repeatability limits.