RMT Engineering Logo
Real-time voice agents

Voice AI Agents

Real-time phone conversations over an open WebRTC media path. The caller cuts in mid-sentence and the agent stops talking. Three turns later they slip from English into Hindi, and the agent goes with them. A speech provider degrades. The call carries on.

  • Barge-in aborted mid-word
  • 20+ languages, switched mid-call
  • Provider failover without dropping
We reply within one business day. No newsletter.
20+ Languages, including 22 Indian
8 + 8 Speech and voice engines
11 LLM providers with failover
Hear a real call

From understanding intent to completing the action — pick an industry and listen.

Now playing E-Commerce Order Tracking Punjabi
300ms Eager endpointing, set per agent

A phone call is a contest over whose turn it is

Most voice bots lose it. They talk over the caller, or they leave such a long gap that the caller gives up and repeats themselves. Both failures come from the same thing: one fixed guess at when a sentence has ended, applied to everyone who rings in. Endpointing is the setting most voice platforms bury. Here it is a number you set per agent, anywhere from 300ms up to 1200ms, because a caller reading out an order number and a caller hunting through a drawer for a policy document do not pause the same way.

How it fits together

The path a conversation takes through Voice AI Agents

The OptiML voice media pipeline, and the path a barge-in takes back through it Caller SIP trunk / WebRTC Voice activity + endpointing Speech-to-text 8 engines Agent runtime knowledge + tools Text-to- speech Synthesised audio back down the same media path Barge-in — the caller speaks: synthesis aborted, in-flight model call cancelled Every stage traced separately: speech, model and synthesis latency per turn.

The forward path carries the caller to the agent; the return path carries synthesised speech back. The dashed line is barge-in, which runs the other way and stops the agent talking.

By the numbers

Few Facts - at a glance

  • 20+

    languages, including 22 Indian languages through Bhashini

  • 8 + 8

    speech-to-text and text-to-speech engines behind one adapter layer

  • 11

    LLM providers with automatic failover, so a provider outage never drops a call

The process

How It Works

The four stages, end to end

The media pipeline is WebRTC. Open and standards-based, not a proprietary tunnel you never get to see inside. Four stages run between the caller and the answer, and each one is traced on its own, so a slow turn points at speech or model or synthesis instead of leaving you to guess. Barge-in runs the same four stages backwards.

  1. The call lands

    A call arrives over your SIP trunk or a WebRTC endpoint. Dispatch rules decide which agent picks it up, and the caller ID lookup pulls the customer record in before anyone has said hello.

  2. Listening and endpointing

    Silero voice activity detection and end-of-utterance detection work out when the caller has actually finished, at the eagerness you set. Not when they paused for breath. The audio then goes to whichever of the eight speech-to-text engines the router picked for that language.

  3. The agent works out the answer

    The runtime reads the transcript against your knowledge base, the customer record and whatever tools this agent is allowed to call, and it does all of that inside the turn rather than parking the caller on hold music while it thinks. Guardrails apply here. Nothing reaches the caller before they have run.

  4. It speaks, and it stops

    Text-to-speech sends synthesised audio back down the same media path. If the caller speaks over it, synthesis aborts and the in-flight model call is cancelled. Either way every stage is timed on its own, so the trace tells you where the second went.

Features

Every capability you need in one module

1. Natural turn-taking

Barge-in means the caller cuts in and the agent stops speaking and starts listening. The in-flight speech and the model call are both killed on the spot. Voice activity detection, end-of-utterance detection and endpointing eagerness are tunable per agent: eager for a fast IVR replacement, patient for a helpline where most of the callers are over seventy. A cough is not an interruption. The false-interruption guard catches that one, and the agent carries on instead of derailing.

  • Silero VAD
  • End-of-utterance detection, so a pause for breath is not read as a full stop
  • Gated barge-in, aborted mid-word
  • Cough guard
Illustrative interface. Counts and settings are sample data, not a performance claim. The screen shows the turn-taking settings for a voice agent: endpointing eagerness set to eager at 300 milliseconds with normal and patient alternatives, sliders for voice activity detection, barge-in sensitivity and the false-interruption guard, and a summary of how many interruptions were honoured and how many coughs were ignored across recent calls.

2. Multilingual, including mid-call

20+ languages, with 22 Indian languages available through Bhashini. Language is not a setting you choose at the start of the call and then hold for the duration. The agent detects a change inside the conversation and switches with the caller. Someone who opens in English and slips into Hindi three turns later is followed, not restarted, and the customer record stays exactly as it was. Same call, same context.

  • Detection inside the conversation, not only at the start
  • Bhashini and Sarvam for Indian-language speech
  • 22 Indian languages
Illustrative interface. The exchange is sample data, not a transcript of a real call. The screen shows a call transcript in which the caller asks two questions in English and then switches to Hindi. The agent answers in Hindi from the same conversation, with a panel confirming that the language change was detected, the voice swapped, and neither the conversation nor the customer record was restarted.

3. Every engine, no lock-in

Speech-to-text across Deepgram, Whisper, AssemblyAI, Google, Azure and AWS Transcribe. Text-to-speech across ElevenLabs, OpenAI, Azure, Google, PlayHT and Cartesia. Bhashini and Sarvam handle Indian regional voice. A quality/cost router picks the cheapest engine that still clears the quality floor you set, and if that engine starts degrading mid-call it fails over rather than dropping the caller. Bring your own provider keys where you already have commercial terms.

8 speech-to-text engines 8 text-to-speech engines Quality/cost router with a quality floor Automatic failover mid-call

See how model and provider access is governed

4. Voice that's yours

Neural voices with SSML, emotion and tone control, plus per-agent speed and stability settings, because the pace that works on a collections call is not the pace you want on a bereavement line. Voice cloning turns a sample into a branded voice, so the same voice answers the phone in every market you run. Speaker biometrics identify the caller by voice.

SSML, emotion and tone control Voice cloning from a sample Speaker biometrics Per-agent speed and stability

5. Inbound, outbound and IVR

One voice agent, three directions. Answer inbound with caller-ID capture and autonomous resolution wherever the agent can manage it. Run outbound campaigns on predictive dialing. Or pull the menu tree out altogether: the visual IVR builder combines DTMF keypad input with conversational routing, and data-dip nodes let a menu check a balance or an order status before it decides where the call goes.

  • Inbound hands over the transcript, the detected intent and the customer record
  • Outbound uses the same agent config
  • No second IVR product to licence
Inbound

Inbound calls are answered with caller-ID capture and resolved autonomously where the agent can. Where it cannot, the handover carries the transcript, the intent and the customer record with it, so the human picks the call up mid-sentence rather than at the beginning.

See what the human agent receives

Outbound

Outbound campaigns run on predictive dialing, and the same voice agent that answers the phone is the one that places the call. One configuration, one set of guardrails, one record of what was said.

See dialing modes and compliance

IVR replacement

The visual IVR builder combines DTMF keypad input with conversational routing, and data-dip nodes let a menu look up a balance or an order before it routes. There is no separate menu product to licence and no second place to keep the routing logic.

Build the call flow on the canvas

6. PSTN and SIP, connected

Inbound and outbound SIP trunk provisioning, with dispatch rules, number management and porting. Bring your own carrier — LiveKit, Twilio, Telnyx, Vonage. Call forwarding knows about timezones and public holidays, so an out-of-hours rule does not send Monday morning calls to voicemail because it happens to be a holiday somewhere else. Call masking keeps the raw customer number out of the agent's view, which reads like a compliance line until the first time somebody leaves and takes a call list with them.

  • SIP ingress and egress, with dispatch rules
  • Number management and porting
  • Timezone- and holiday-aware forwarding
  • Call masking, so agents never see the raw number

Recording, kept properly

Call recordings are held under AES-256-GCM envelope encryption, with retention policies and consent capture, in the region you chose. Which region, how long recordings live and who can reach them are decisions you make at deployment.

See hosting, residency and retention

Use Cases

Where Voice AI Agents delivers value

Natural barge-in

The caller corrects themselves mid-sentence

A mid-sized general insurer (illustrative)

Scenario

A policyholder starts asking about a renewal date, then changes their mind halfway through the agent's reply. "Sorry, just the premium amount." The caller speaks over the agent. Synthesis stops, the in-flight model call is cancelled, and the corrected question is answered in the same turn.

Outcome

Callers who hesitate, back up and correct themselves get to the end of the call instead of hanging up, dialing in again and starting the whole thing over with a different agent. A cough no longer derails the answer halfway through.

Mid-call language switching

One call that begins in English and ends in Hindi

A regional retail bank (illustrative)

Scenario

A customer opens in English to check a balance. Three turns in, they switch to Hindi to explain a problem at the branch, because that is the language the complaint actually lives in. The change is detected inside the conversation. The voice swaps through Bhashini, and the customer record and the context carry straight over.

Outcome

There is no separate Hindi queue to staff. Nobody has to say "please call our Hindi line", and the same agent, knowledge base and customer record serve both halves of the call.

IVR replacement with data dips

The menu tree comes out, the agent goes in

A utility contact centre (illustrative)

Scenario

Press one for billing, press two for a new connection. Callers who are not sure which one they need pick wrong, and the queue behind the wrong option fills up with transfers. So the flow gets rebuilt in the visual IVR builder. Keypad input still works for the people who prefer it, conversational routing handles everyone else, and a data-dip node reads the account balance before the call is sent anywhere.

Outcome

Routing logic sits where the prompts and guardrails already sit. One change, one place. No pair of systems drifting out of step for six months before anyone notices.

At a glance

Specification

The numbers and limits, without the sales copy

Specification for Voice AI Agents
Specification Detail
Media pipeline WebRTC, open and standards-based
Speech-to-text Deepgram Whisper AssemblyAI Google Azure AWS Transcribe Bhashini Sarvam
Text-to-speech ElevenLabs OpenAI Azure Google PlayHT Cartesia Bhashini Sarvam
Speech-to-speech OpenAI Realtime low-latency path
Turn-taking Silero VAD, end-of-utterance detection, gated barge-in, false-interruption resume
Endpointing Eager 300ms / normal 700ms / patient 1200ms, per agent
Languages 20+, including 22 Indian languages
Telephony LiveKit, Twilio, Telnyx, Vonage — SIP ingress and egress
Recording AES-256-GCM envelope encryption, retention policies, consent capture
Observability Per-stage latency tracing across speech, model and synthesis
FAQ

Questions,
answered

What teams ask us before they roll out Voice AI Agents — how it works, what it needs from your side, and what happens when it gets something wrong

Still not sure?

Talk to a specialist and get a straight answer.

Ask our team

Voice AI Agents connect over SIP ingress and egress, with trunk provisioning and dispatch rules deciding which agent answers which number. Bring your own carrier (LiveKit, Twilio, Telnyx, Vonage), or port your existing numbers across and keep them. If the calls are in-app or in-browser, a WebRTC endpoint works with no PSTN leg at all, so a pilot can run before anyone goes near a carrier contract.

Yes. Voice AI Agents put eight speech-to-text and eight text-to-speech engines behind one adapter layer, so changing provider is configuration rather than a rebuild. Bring your own provider accounts and keys where you already have commercial terms. You also set the quality floor the cost router has to clear before it is allowed to pick the cheaper engine, so cost optimisation does not quietly turn into a transcription accuracy problem that surfaces in a QA review three weeks later.

The quality/cost router in Voice AI Agents watches provider health and moves the call to a standby engine instead of hanging up. The same holds at the model layer, where eleven LLM providers are available for failover. The caller stays on the line. The conversation picks up from where it was.

Voice AI Agents resolve what they can on their own and hand over the rest. The trigger is a request outside what the agent is allowed to do, a guardrail that blocks the answer, or a caller who simply asks for a person. Whichever it is, the handover carries the transcript, the detected intent and the customer record with it, so the human picks the call up mid-sentence rather than at the beginning. The per-stage latency trace and the recording stay attached to the call, so a handover that should not have been needed can be traced back to speech, to retrieval or to the prompt.

Call recordings from Voice AI Agents are held under AES-256-GCM envelope encryption, with retention policies and consent capture, in the region you pick at deployment. Live audio only reaches the speech and model providers you have switched on. The list of third parties in the call path is one you set, not one you inherit.

Connected solutions

Where Voice AI Agents is used

The Solutions pages that lean on this module, and what it looks like once it is configured for a particular floor, job title or job to be done.

16 solutions built on this module
Talk to a specialist

Hear it interrupt, and stop mid-word

Bring a call recording, a policy document or a WhatsApp thread from your own operation. We ground an agent in it and put it on a live call while you watch. Not a canned demo.

  • 30 minutes
  • A working agent grounded in your content
  • No slide deck unless you want one

Book your slot

Leave your email and our team will come back to you within one business day.

or reach us directly

Your details stay private. We never share them.

This website uses cookies.

Cookies are small text files that allow us to create the best browsing experience for you on our site. By continuing to use this website or clicking "Accept & Close", you are agreeing to our use of cookies. To understand how we use cookies or how to manage them, please see our cookies policy.

Ask OptiML

Powered by RMT Engineering