Events and transcripts
VoiceSession.run() yields typed events from timbal.voice. The server converts them to transport JSON; WebRTC and LiveKit carry speech on audio tracks instead of JSON audio frames.
session.transcript contains committed user and assistant entries with timestamps. Fillers are marked filler=True. At session end, the server sends session_transcript followed by session_ended. Transcript text reflects interruption handling and may contain only the heard prefix of a cancelled assistant reply.
Latency
TurnMetricsEvent reports millisecond durations and a run_id for joining a voice turn to the agent trace. Unobserved stages have None, not zero.
speech_end_to_transcript_ms: speech-end evidence to the committed transcript.speech_end_sourceidentifies VAD or the coarser partial-transcript estimate.eou_to_llm_first_token_ms: committed transcript to the first model text delta.eou_to_first_audio_ms: committed transcript to the first emitted speech bytes. This excludes transcription delay and network/client playback delay.turn_total_ms,llm_total_ms, andtts_total_ms: turn and stage durations.interrupted,heard_bytes, andplayback_acks_received: interruption and playback evidence.filler_spokenandfiller_count: whether a latency-masking phrase was spoken. When present, first-audio latency can measure the filler rather than the substantive answer.
Provider-reported speech usage
OpenAI transcription and mini-TTS adapters emitVoiceUsageEvent independently of transcripts and audio. These events carry provider quantities, not dollar amounts. Other providers and custom adapters emit nothing unless they implement the usage capability.
For billing, register a non-blocking listener on the adapters before the session starts. Delivery through the client connection alone can lose teardown events:
provider, operation (stt or tts), and model identify the resolved operation. Deduplicate by usage_id; prefer a complete report over an earlier incomplete report with the same ID. request_id and item_id, when present, help reconcile provider operations.
status="complete" means terminal provider counters arrived. status="incomplete" means quantities are missing, invalid, or unavailable after cancellation, error, or closure. Incomplete is not zero cost. usage preserves the provider object, including any partial details.
OpenAI transcription can report token usage with separate input audio/text details, or duration usage with seconds. Mini-TTS reports input text and output audio tokens through the Speech API’s SSE completion event. Legacy tts-1 and tts-1-hd use PCM streaming and do not supply these SSE usage reports. Do not apply one text rate to total_tokens or charge both an estimate and the reported quantities. Generated or cancelled audio can be billed even when the caller never hears it.
Apply the resolved speech model’s published rates separately from the text agent’s OutputEvent.usage. The text model catalog and cost SQL generator do not provide a complete speech rate card. Use an explicit estimate policy for incomplete operations, and label those estimates.