PersonaPlex#
nvidia/personaplex-7b-v1 is a 7B full-duplex speech-to-speech model built on Moshi: every 80 ms it reads one text token and 16 Mimi codes (8 for what it says, 8 for what it hears) and writes the next frame. A voice prompt and a <system> role prompt set the persona, and the model decides for itself when to speak; there is no VAD.
SGLang-Omni serves it as an offline pipeline: one recording of the caller’s side in, the agent’s reply as text and 24 kHz audio of the same length out. Live duplex sessions over /v1/realtime are tracked in #1909.
Prerequisites#
The checkpoint is gated; accept the license on the model page and log in (hf auth login), then:
hf download nvidia/personaplex-7b-v1
It ships the 7B weights, the Mimi codec, the SentencePiece text model and voices.tgz, which is unpacked into voices/ next to the checkpoint on first use, or into the temp directory when that folder cannot be written. Recorded voice prompts (--voice some.wav) also need pip install pyloudnorm; the packaged .pt voices do not.
Everything runs on one GPU. On an H200 the LM engine reserves mem_fraction_static=0.3 and the two Mimi instances stay under 1 GB each; lower --lm.engine.mem_fraction_static on smaller cards.
Running the offline example#
python examples/run_personaplex.py \
--model-path nvidia/personaplex-7b-v1 \
--audio /path/to/caller.wav \
--voice NATF2 \
--text-prompt "You are a wise and friendly teacher. Answer questions or provide advice in a clear and engaging way." \
--out reply.wav
Five stages run under MultiProcessPipelineRunner (preprocessing, Mimi encode, the LM engine, text decode, streaming code2wav). Pick the GPU with CUDA_VISIBLE_DEVICES; stage settings take the same dotted flags as serve, such as --lm.engine.mem_fraction_static 0.25. Input conventions:
The recording is resampled to 24 kHz and channel 0 is used.
The reply is exactly as long as the input, offset by one frame: the model answers while it listens, so leave silence after the caller’s last words if you want a full answer.
Prompt plus reply must fit the LM context, 8192 positions by default (about 10.8 minutes); a longer recording is rejected with the limit in the message and needs
--lm.engine.context_length.The reply text is the model’s inner monologue with the frame markers (
PAD,EPAD,BOS,EOS) removed.
Serving over HTTP#
python -m sglang_omni.cli serve --model-path nvidia/personaplex-7b-v1 --port 8000
Send the recording as the prompt of POST /generate (a path, URL or data: URI, with "return_logprob": false), or as the single entry of audios in POST /v1/chat/completions with "modalities": ["text", "audio"]; the reply audio comes back base64-encoded as WAV. Options with no request field of their own go under stage_params: voice and text_prompt under preprocessing, audio_temperature, audio_top_k and seed under lm.
curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "nvidia/personaplex-7b-v1",
"messages": [{"role": "user", "content": ""}],
"audios": ["/path/to/caller.wav"],
"modalities": ["text", "audio"],
"stage_params": {"preprocessing": {"voice": "NATM1"}, "lm": {"audio_temperature": 0.8}}
}'
Request parameters#
Parameter |
Effect |
|---|---|
|
A packaged voice name ( |
|
The role prompt, wrapped in |
|
Text sampling; defaults 0.7 / 25, which the client’s filler values (1.0 / -1) do not override unless set explicitly. |
|
Code sampling in the depformer; defaults 0.8 / 250. |
|
Makes both draws reproducible (a child seed each for text and audio). |
|
Ignored; the frame count is fixed by the input length. |
Explicit stage sampling#
HTTP chat and rollout requests preserve which fields were provided in
stage_sampling.lm. For example, {"seed": 7} keeps the PersonaPlex text
defaults (temperature 0.7, top-k 25), while
{"temperature": 1.0, "top_k": -1} explicitly selects temperature 1.0 and
disables top-k filtering. In-process stage SamplingParams objects are complete configurations:
all serialized fields are explicit. Thus SamplingParams(seed=7) in
stage_sampling.lm selects temperature 1.0 and top-k -1 as well as seed 7.
To keep PersonaPlex text defaults, omit stage_sampling and set the seed in
the top-level sampling=SamplingParams(seed=7) instead. Top-level client
filler handling is unchanged.
Text sampling uses stage_sampling.lm, then stage_params.lm, then top-level
request values. Stage sampling also takes precedence for seed.
top_p, min_p, and repetition_penalty apply to the text sampler only;
audio_temperature and audio_top_k continue to control the audio sampler.
The fixed caller-frame budget still determines the number of generated frames.
Known limitations#
Offline, one request at a time (
max_running_requests=1).CUDA graphs are off; a 7B decode step plus 8 depformer steps runs close to the 80 ms frame budget rather than well inside it.
The temporal attention window follows the streaming ring, including the masked oldest slot once its 3000-position cache fills. Boundary tests check this rule; they do not measure long-input audio quality.
Tests#
tests/unit_test/personaplex/ runs on CPU without weights: the delayed timeline, chunked Mimi against whole-sequence Mimi, the depformer, checkpoint weight routing, the model-runner hooks, the streaming codec stage, the checkpoint shim, preprocessing and voice unpacking, and request lowering.
pytest tests/unit_test/personaplex/ -q
Reference comparisons and reproducibility checks are evaluation scripts under
benchmarks/eval/, separate from the unit suite. See the
PersonaPlex evaluation guide for checkpoint,
reference-environment and output setup. The greedy comparison checks a matching
prefix, not full-output equality; the component comparison includes a diagnostic
for the reference depformer ring behavior.