MiniCPM-o Reference Audio#
On the speech pipeline, pass an explicit speaker reference in
audio.ref_audio on /v1/chat/completions:
import base64
from pathlib import Path
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="unused")
reference = base64.b64encode(Path("reference.wav").read_bytes()).decode("ascii")
response = client.chat.completions.create(
model="MiniCPM-o-4_5",
messages=[{"role": "user", "content": "Please say hello."}],
modalities=["text", "audio"],
audio={
"format": "wav",
"ref_audio": f"data:audio/wav;base64,{reference}",
},
)
stage_params.code2wav.ref_audio is an alternative, with higher priority than
audio.ref_audio. Both accept prompt_wav as an alias. The Python pipeline
client can also supply ref_audio through extra_params. References must be
base64 audio data URIs, inline {data, media_type} descriptors, or encoded audio
bytes for the Python client. Paths and HTTP URLs are not fetched by this stage;
read or download the file on the client before sending it.
The reference conditions Token2wav’s speaker embedding, prompt tokens, and mel
features. Audio supplied in chat messages remains understanding input and is not
automatically used as the speaker reference. Without an explicit reference,
Token2wav uses the checkpoint’s assets/HT_ref_audio.wav when available.
The vocoder caches only the most recently used reference by audio content. A different reference, including switching back to the default, rebuilds the conditioning. Invalid references fail instead of silently using the default. Audio output remains non-streaming.