Skip to content

← Back to the Labs

L3:voice-ai-contact-center · AI Wing · Hero

Realtime Voice AI Contact Center

Callers interrupt the agent naturally (~2 s to ~0.2 s), get grounded answers and reach a human mid-call.

Role
Owner and lead engineer: designed, directed, reviewed and shipped it with AI coding agents
Period
May 2026 to Oct 2026

Results

  • ~2 s to ~0.2 s

    Caller interruption during knowledge lookups

    Controlled A/B test on the same deployment, index and audio, n=4 calls per arm (1,875 to 2,140 ms vs 93 to 220 ms).

    Verified
  • 2 of 2

    Node-loss drills passed with live calls surviving

    Hard VM-stop failover and caller-reconnect drills, run after four root causes were found and fixed.

    Verified
  • 6 vs 18

    Hallucinated answers: the quality gate kept the stronger realtime model

    Offline evaluation with a calibrated LLM judge (mean score 4.71 vs 4.28).

    Verified
  • Migrated

    Voice channel moved off a retiring service, old path removed

    Live cut-over well ahead of the announced 2028 retirement; about 20.5k lines of the old path deleted.

    Verified

Problem

Phone menus frustrate callers. The goal was a voice agent that answers from a company knowledge base, can be interrupted while it talks or searches, hands the caller to a human mid-call and back, and gives supervisors a live console.

What I did

  • Merged two prototypes into one product and ported the realtime backend to Python and FastAPI.
  • Made barge-in playout-aware and moved knowledge-base tool calls off the realtime loop into cancellable tasks with a single time budget.
  • Ran a production-readiness program: atomic writes, crash recovery, a replayable dead-letter queue, drain-gated releases, managed identity and per-environment infrastructure as code.
  • Wrote the decision record to move the voice channel off a service announced for retirement onto self-hosted WebRTC media on zonal VMs, then proved failover with hard VM-stop drills.
  • Built a voice evaluation harness (media emulator, golden questions, calibrated LLM judge) and kept changes the data did not support switched off.

Architecture

  1. ClientCallerBrowser over WebRTC
  2. HumanSupervisorNext.js operator console
  3. ServiceLiveKit mediaSelf-hosted on zonal VMs
  4. ServiceFastAPI realtime backendCancellable tools, 2.5 s budget
  5. AIAzure OpenAI RealtimeSpeech to speech, barge-in
  6. DataAzure AI SearchHybrid and semantic
  7. DataCosmos DBAtomic writes, dead-letter replay
  8. ServiceKnowledge ingestionDurable Functions, Document Intelligence

Data flow

  • Caller to LiveKit media (WebRTC audio)
  • LiveKit media to FastAPI realtime backend (audio stream)
  • FastAPI realtime backend to Azure OpenAI Realtime (speech to speech)
  • FastAPI realtime backend to Azure AI Search (grounding tools)
  • FastAPI realtime backend to Cosmos DB (calls, transcripts)
  • Supervisor to FastAPI realtime backend (takeover, hand-back)
  • Knowledge ingestion sends asynchronously to Azure AI Search (index)

Stack

  • Azure OpenAI Realtime
  • Azure AI Search
  • Azure Speech
  • Document Intelligence
  • Durable Functions
  • Cosmos DB
  • App Service
  • LiveKit
  • Azure VMs
  • Python
  • FastAPI
  • Next.js
  • Entra ID
  • OpenTelemetry
  • Application Insights
  • Bicep
  • GitHub Actions

L2 · On the journeyTechnical Consultant, Software & AI · Metrodata (PT Mitra Integrasi Informatika)