ZONOS2#
ZONOS2 is a mixture-of-experts (MoE)
text-to-speech model from Zyphra. A MoE autoregressive decoder predicts
9 DAC audio codebooks scheduled in a delay pattern; the codes are then
decoded back to 44.1 kHz speech by a DAC vocoder. It clones a voice from a
short reference clip and supports optional language-specific text normalization. In
SGLang-Omni it runs as a preprocessing β speaker_encode β tts_engine β vocoder pipeline and is served through the OpenAI-compatible
/v1/audio/speech endpoint.
Prerequisites#
Install sglang-omni by following Installation, then
download the model:
hf download Zyphra/zonos2
The processor ships with the checkpoint, so no extra TTS package is needed. Voice cloning
transcodes reference audio (file, URL, or base64 data-URI) with ffmpeg, so ffmpeg must
be on the serverβs PATH (e.g. apt-get install ffmpeg).
Server Configuration#
The pipeline is preprocessing β speaker_encode β tts_engine β vocoder.
ZONOS2 ships a params.json whose model_type (zonos2) auto-selects the
Zonos2ForCausalLM architecture, so serve needs only --model-path β no
--config (mirrors Higgs).
sgl-omni serve \
--model-path Zyphra/zonos2 \
--port 8000
Synthesizing Speech#
Voice Cloning#
ZONOS2 clones a voice from a reference clip. The references field accepts audio_path
(a local path, HTTP URL, or base64 data URI) and text (the transcript of that clip).
The shared speech API accepts the transcript, but ZONOS2 currently conditions only on
the reference audio.
curl -X POST http://localhost:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "SGLang-Omni is a great project!",
"references": [{
"audio_path": "https://huggingface.co/datasets/zhaochenyang20/seed-tts-eval-mini/resolve/main/en/prompt-wavs/common_voice_en_10119832.wav",
"text": "We asked over twenty different people, and they all said it was his."
}]
}' \
--output output.wav
ref_audio and ref_text are accepted as shorthand for references[0].audio_path and
references[0].text. ref_text is currently unused by ZONOS2.
Python#
import requests
resp = requests.post(
"http://localhost:8000/v1/audio/speech",
json={
"input": "Get the trust fund to the bank early.",
"ref_audio": "https://huggingface.co/datasets/zhaochenyang20/seed-tts-eval-mini/resolve/main/en/prompt-wavs/common_voice_en_10119832.wav",
"ref_text": "We asked over twenty different people, and they all said it was his.",
},
)
resp.raise_for_status()
with open("output.wav", "wb") as f:
f.write(resp.content)
Reference Audio Sources#
audio_path / ref_audio may be a local filesystem path readable by the server, an HTTP(S)
URL, or a base64 data URI (data:audio/wav;base64,<...>, transcoded via ffmpeg):
{"ref_audio": "data:audio/wav;base64,UklGR.....", "ref_text": "Transcript of the clip."}
Streaming#
Set "stream": true with "response_format": "pcm" to receive raw signed 16-bit
little-endian PCM bytes. The response is not SSE or base64. ZONOS2 returns mono 44.1 kHz
audio with Content-Type: audio/pcm and the X-Sample-Rate, X-Channels, and
X-Bit-Depth headers.
curl -sS -D output.headers -X POST http://localhost:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"input": "Get the trust fund to the bank early.",
"ref_audio": "https://huggingface.co/datasets/zhaochenyang20/seed-tts-eval-mini/resolve/main/en/prompt-wavs/common_voice_en_10119832.wav",
"ref_text": "We asked over twenty different people, and they all said it was his.",
"response_format": "pcm",
"stream": true
}' \
--output output.pcm
ffmpeg -f s16le -ar 44100 -ac 1 -i output.pcm output.wav
Language#
language selects the NeMo written-to-spoken normalizer before the prompt is
tokenized. English, Chinese, Japanese, Korean, German, French,
Portuguese, Spanish, and Italian are normalized. Omit it or use Auto to
keep the input unchanged; accepted languages without a NeMo package, such as
Russian, also pass through unchanged. The model infers the spoken language
from the resulting text.
{
"input": "δ»ε€©ε€©ζ°δΈιοΌε°±θ―₯εΊε»ζζε€ͺι³γ",
"ref_audio": "...", "ref_text": "...",
"language": "Chinese"
}
Generation Parameters#
Parameter |
Default |
Notes |
|---|---|---|
|
(required) |
Text to synthesize |
|
|
Reference clip for cloning; |
|
|
Shorthand fields; |
|
|
Use |
|
|
Stream raw mono 44.1 kHz PCM16 bytes; requires |
|
|
Optional NeMo text-normalization language; |
|
(model default) |
Maximum generated frames; an explicit value must be |
|
(model default) |
Sampling temperature |
|
(model default) |
Top-p sampling |
|
(model default) |
Top-k sampling |
|
(model default) |
Min-p sampling |
|
(model default) |
Audio repetition penalty |
Benchmarking#
ZONOS2 clones from each prompt (--ref-format references). Run the seed-tts-eval voice-clone
benchmark against a running server:
python -m benchmarks.eval.benchmark_tts_seedtts \
--meta zhaochenyang20/seed-tts-eval-arrow \
--model Zyphra/zonos2 --port 8000 \
--ref-format references \
--output-dir results/zonos2_en --lang en --max-concurrency 16
Use --lang zh for the Chinese split. See benchmarks/README.md for the full workflow.
Benchmark Results#
Seed-TTS-Eval full sets (EN 1088 / ZH 2020), 1Γ H100, concurrency 16, cold full clock,
--ref-format references, fp8 + frame_graph + async_decode + compile_sampler. Scored with
Qwen3-ASR-1.7B (the CI scorer; EN word-WER, ZH char-CER). Median WER/CER is 0% for every
config β the corpus numbers below are inflated by a few catastrophic-collapse samples from
unseeded sampling (EN β€3/1088, ZH 9β16/2020), so cross-config corpus deltas are run noise, not
config differences (excluding those, corpus is EN ~1.3% / ZH ~1.4β1.7%). All EN configs clear the
<2% gate.
Config |
WER (corpus) |
RTF mean |
Throughput (qps) |
|---|---|---|---|
non-streaming |
1.7% |
0.69 |
6.2 |
streaming, |
1.4% |
0.92 |
4.6 |
streaming, adaptive |
1.3% |
0.77 |
5.6 |
streaming, 2-GPU ( |
1.5% |
0.69 |
6.2 |
Streaming throughput (stream_emit_chunk_frames)#
In streaming mode the AR engine pushes sampled frames to the vocoder over an in-process queue.
By default the engine coalesces stream_emit_chunk_frames=32 frames into one message instead
of one put() per frame; the per-frame puts run on the resolve host loop and serialize against
the next decode launch, so batching them gives β17% streaming RTF and +21% throughput,
WER-neutral, vs the per-frame path (inter-chunk latency also improves, 0.37 s β 0.28 s). To keep
first-audio latency low, the default also emits a smaller first chunk
(stream_emit_first_chunk_frames=24), so time-to-first-chunk stays at the per-frame level
(~0.30 s) instead of waiting for a full 32-frame batch (which would add ~0.09 s). Set
stream_emit_first_chunk_frames=0 to disable the adaptive first chunk, or
stream_emit_chunk_frames=1 for the lowest-latency per-frame streaming. The 2-GPU multi_gpu
pipeline (codec + speaker encoder on cuda:1) stacks on top for the best streaming RTF (~0.69).
ZH (
--lang zh, full 2020) β Qwen3-ASR corpus CER 1.8β3.0% (median 0%; the corpus reflects a 9β16/2020 catastrophic-sample tail, not config differences β excluding those, ~1.4β1.7%). Non-streaming RTF ~0.62β0.64, streaming ~0.65β0.72; coalescing is audio-neutral after theon_stream_donetail fix.
Known Limitations#
Reference transcripts are not used. The shared API accepts
text/ref_text, but ZONOS2 currently conditions only on reference audio.Language controls text normalization only. It does not add a separate model-side language condition.
Per-request seeds are unsupported. Requests with
seedare rejected until sampling has an isolated request RNG.Rare runaway generation. A small fraction of utterances can loop and generate up to
max_new_tokens; loweringmax_new_tokensbounds the output.