Narrators & voices
A narrator is a named profile: engine, model, voice, and style. One narrator renders a whole episode; a failure never falls back silently to another narrator.
Built-in narrators
| Narrator | Engine | Model | Voice |
|---|---|---|---|
gemini |
gemini | gemini-3.8-flash-tts |
Charon |
gemini-lite |
gemini | gemini-3.8-flash-lite-tts |
— |
openai |
openai | gpt-4o-mini-tts-2025-12-15 |
marin |
tone |
tone | — | — |
gemini is the default. gemini-lite falls back to the Charon voice
when its profile names none. tone is the offline deterministic engine
for tests, CI smoke, and demos: it costs nothing and renders at about
150 wpm.
Engines
- gemini calls
client.interactions.createwith turn-levelspeech_metadata.stylestyling and a prebuilt voice. The transcript is sent verbatim; style is never inlined into the text. - openai POSTs to
{base_url}/audio/speechwith{model, input, voice, response_format: "pcm"}, plusinstructionswhen the narrator has a style andspeedwhen it differs from 1.0. The defaultbase_urlishttps://api.openai.com/v1. - tone renders deterministic speech-paced tone bursts locally.
The style prompt
The default narrator speaks with a fixed calm technical-briefing style:
Calm, clear technical-briefing narrator. Moderate pace. Slight emphasis on numbers. Read the text exactly as written.
For Gemini it travels as a structured speech_metadata.style annotation;
for OpenAI-compatible endpoints as instructions. The style applies to
spoken audio only, never to ID3 or feed chapter titles.
Adding a Kokoro profile
Point a narrator at a local Kokoro-FastAPI server through the OpenAI
adapter with a custom base_url:
narrators:
kokoro:
engine: openai
model: kokoro
voice: af_heart
base_url: http://localhost:8880/v1
Kokoro voices may prefer shorter inputs; the adapter's per-call
max_chars override (default 3,500) exists for that. Costs nothing:
unknown models estimate at $0.00.