[ ABORT TO HUD ]
SEQ. 1
SEQ. 2
SEQ. 3
SEQ. 4
SEQ. 5
Speech Services & Voice Live
Speech-to-Text & Real-Time Voice
Foundry's Speech services enable voice-powered AI applications with high-quality transcription and synthesis.
Speech Capabilities
| Service | Function | Key Features |
|---|---|---|
| Speech-to-Text | Transcribe audio to text | Real-time & batch, 100+ languages, custom models |
| Text-to-Speech | Convert text to natural speech | 400+ neural voices, custom voice cloning |
| Voice Live | Real-time speech-to-speech | Fully managed runtime, noise suppression, barge-in (New in 2026) |
| Speaker Recognition | Identify speakers by voice | Verification and identification modes |
Building Voice-Enabled Agents
Combine Speech services with the Agent Service to build voice-controlled AI assistants. With the 2026 Voice Live integration, this is easier than ever:
- User speaks → Voice Live captures audio, handling noise suppression natively
- Direct integration → Sent to Foundry Agent (e.g. GPT-4o Audio) for processing
- Agent response → Voice Live streams synthesis immediately
- User can interrupt ("barge-in") smoothly
🎯 Pro Tip: Use the fully managed Voice Live runtime for interactive conversational agents rather than building custom STT/TTS pipelines. This natively handles complex edge cases like user interruptions ("barge-in") and echo cancellation.
FOUNDRY VERIFICATION
QUERY 1 // 2
What is the primary advantage of the new Voice Live runtime for agents?
It provides video streaming
It natively handles real-time speech-to-speech with features like barge-in and noise suppression
It translates text faster
It reduces token costs by 50%