MiniCPM-o Reference Audio

MiniCPM-o Reference Audio#

On the speech pipeline, pass an explicit speaker reference in audio.ref_audio on /v1/chat/completions:

import base64
from pathlib import Path

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="unused")
reference = base64.b64encode(Path("reference.wav").read_bytes()).decode("ascii")
response = client.chat.completions.create(
    model="MiniCPM-o-4_5",
    messages=[{"role": "user", "content": "Please say hello."}],
    modalities=["text", "audio"],
    audio={
        "format": "wav",
        "ref_audio": f"data:audio/wav;base64,{reference}",
    },
)

stage_params.code2wav.ref_audio is an alternative, with higher priority than audio.ref_audio. Both accept prompt_wav as an alias. The Python pipeline client can also supply ref_audio through extra_params. References must be base64 audio data URIs, inline {data, media_type} descriptors, or encoded audio bytes for the Python client. Paths and HTTP URLs are not fetched by this stage; read or download the file on the client before sending it.

The reference conditions Token2wav’s speaker embedding, prompt tokens, and mel features. Audio supplied in chat messages remains understanding input and is not automatically used as the speaker reference. Without an explicit reference, Token2wav uses the checkpoint’s assets/HT_ref_audio.wav when available.

The vocoder caches only the most recently used reference by audio content. A different reference, including switching back to the default, rebuilds the conditioning. Invalid references fail instead of silently using the default. Audio output remains non-streaming.