Fish S2 HTTP validation¶
The fish_audio_s2 integration sends requests to an independently managed
Fish Speech S2 server. Install the client with pip install -e '.[http]'.
For server setup and checkpoint provisioning, use the S2 environment and
source revision. The measured source revision is
214da3cd841bda85da2496b96cd3c4d7edb1337e.
Run the official service from that checkout in its own environment:
CUDA_VISIBLE_DEVICES=0 python -m tools.api_server \
--listen 127.0.0.1:18011 --workers 1 --device cuda:0 \
--llama-checkpoint-path /absolute/path/to/s2-pro \
--decoder-checkpoint-path /absolute/path/to/s2-pro/codec.pth \
--decoder-config-name modded_dac_vq
Use a free port and retain ownership of the server process. A successful
GET /v1/health confirms readiness; it does not attest server weights or code.
The benchmark checks readiness before generation, requires actual reference
and target transcripts, sends seed plus original sample index, and rejects
non-WAV success responses and permanent HTTP errors. Transient retries reuse
the exact serialized request. Remote inference failure after acceptance can
still leave an unknown server-side outcome.
python run_vc.py --config-name ots_vc/clean/libritts/fish_audio_s2_ots \
run_name=fish_s2_http_libritts16 dataset.use_hf_dataset=false \
+dataset.manifest_filename=/absolute/path/to/reproduction/subsets/libritts16_v1/metadata.json \
adversary.endpoint_url=http://127.0.0.1:18011/v1/tts \
+adversary.service_code_path=/absolute/path/to/fish-speech \
+adversary.service_checkpoint_path=/absolute/path/to/s2-pro \
+adversary.service_codec_checkpoint_path=/absolute/path/to/s2-pro/codec.pth \
+vc.generate_only=true +seed=42
Those optional local paths record source and asset hashes; they must refer to the assets actually used by the owned server. They do not verify an arbitrary remote server. The official server's permissive loader was independently checked against strict native loading of the same local assets for this run.
Score saved audio in the evaluation environment with +vc.evaluate_only=true
and +vc.evaluation.generated_audio_dir=/absolute/path/to/generated_audio.
vc.evaluation.device takes precedence over evaluation.device, followed by
the generation device. HTTP generation uses CPU on the client while the
default scoring device is CUDA. Scoring provenance includes execution device,
Torch/CUDA versions, CPU thread count, and CUDA device properties, preventing
cross-device cache reuse.
All 16 fixed LibriTTS pairs were generated and scored. Their historical inputs are matched by audio hashes and transcripts. Historical server weights, runtime, and RNG state are unavailable, so the comparison remains diagnostic. The historical S2 entry is experimental and outside the paper's 18-model set.
For diagnosis, scripts/serve_fish_s2.py invokes the official single-worker
API after native startup seed setup. It rejects multiple workers because
spawned workers would not inherit that initialization. This is a diagnostic
launcher, not a guarantee of deterministic inference. An optional
--diagnostic-trace /absolute/path/to/new_trace.jsonl records reference and
generated token hashes on the owned engine instance. It synchronizes CUDA,
so traced runs must not be used for timing comparisons. The original upstream
checkout is unchanged. Example from the repository root:
python scripts/serve_fish_s2.py \
--code-path checkpoints/fish-speech-s2-native --startup-seed 42 \
--diagnostic-trace /absolute/path/to/new_trace.jsonl -- \
--listen 127.0.0.1:18011 --workers 1 --device cuda:0 \
--llama-checkpoint-path /absolute/path/to/s2-pro \
--decoder-checkpoint-path /absolute/path/to/s2-pro/codec.pth \
--decoder-config-name modded_dac_vq
All owned validation services have been stopped; existing services were preserved.