House of MohnyTest protocol · v1

Voice Platform Test Protocol

How to test AI voice platforms on an Indian phone number in an afternoon, and the scorecard to run them through. This is the method — the results are yours, because the only results that matter are on your number, your accent, your callers.

Why this isn't a ranked list

Every voice-platform comparison you'll read was run in a different country, on a different codec, in a different accent, on a version that shipped months ago. A benchmark you didn't run is a marketing claim. This protocol takes an afternoon and gives you a number you can actually defend.

The criteria

Three things, nothing else


Feature lists are a distraction. Three things decide whether a voice agent survives contact with a real Indian caller.

Latency to first audio

Time from the caller finishing a sentence to the first syllable coming back. Measure it on a real call, over the actual phone network — not in the browser demo on the vendor's site, which skips the part that's slow.

Fail thresholdOver 1.5 seconds. Callers begin talking over it or assume the line dropped.

Code-switch survival

Say one sentence that changes language halfway — "Tuesday ko appointment mil jayega?" — and see whether both halves survive transcription. This is the single most common failure and almost no vendor tests for it.

Fail thresholdEither half dropped, or the reply answers a question you didn't ask.

Calendar write

Not "can it discuss availability" — can it put a correctly-dated booking into a real calendar, unattended? Check the date, not just that something appeared.

Fail thresholdWrong day, wrong timezone, or a booking it claims to have made and didn't.
Method

How to run it


  • One script, read identically every time. Write six lines including one code-switch and one ambiguous date ("next Tuesday"). Same words, same pace, every platform.
  • Call from a real phone on mobile data, not a laptop on office wifi. The network is part of what you're testing.
  • Three calls per platform. One is noise. Take the median, not the best — you'll be tempted to keep the good one.
  • Record every call. You will disagree with your own notes later.
  • Read the latency off the waveform, not off a stopwatch. Open the recording in any free audio editor — Audacity will do — select from the end of your last word to the first syllable of the reply, and read the selection length. That is accurate to a few milliseconds. A stopwatch carries your reaction time at both ends, roughly a third of a second, and the platforms sit closer together than that, so a stopwatch tells you which ones are broken and nothing more.
  • Test one thing you'd never ask. Interrupt it mid-sentence. How it recovers tells you more than the happy path.
Scorecard

Fill this in


Print it or copy it. Median of three calls per row.

The Platform column is blank because the shortlist is yours. As of the field splits three ways, and a useful afternoon covers at least two of them. End-to-end voice agent platforms — Vapi, Retell AI, Bland AI, ElevenLabs Agents — hand you telephony, transcription, model and speech in one product. Assembled stacks put a speech-to-text service (Deepgram, AssemblyAI, Sarvam AI) and a text-to-speech service (ElevenLabs, Cartesia, Sarvam AI) either side of a model you pick yourself: more work, more control. Indian telephony — Exotel, Plivo, Knowlarity, Ozonetel, or Twilio on an Indian number — sits underneath both, and it is where the latency and the code-switch survive or don't. Naming them is not ranking them, and this paragraph ages faster than the method does. Check what each one supports on the day you test.

PlatformLatencyCode-switchCalendarInterruptVerdict
1.___ mspass / failpass / failpass / fail___
2.___ mspass / failpass / failpass / fail___
3.___ mspass / failpass / failpass / fail___
4.___ mspass / failpass / failpass / fail___
5.___ mspass / failpass / failpass / fail___
Send me your scorecard →

Any platform failing code-switch is out, regardless of latency. A fast agent that mishears half your callers is worse than a slow one that doesn't.

What usually happens

Most people run this expecting latency to be the differentiator and find that code-switching is. The platforms cluster within a few hundred milliseconds of each other and separate hard on whether they can follow a sentence that changes language halfway through.

Which is the whole reason a comparison run in California doesn't tell you anything useful here.

Run it, then send me your scorecard. I'll tell you what I'd deploy and where I'd expect it to break — and if your numbers contradict mine, I want to know that more than you do. I'm in Bengaluru, testing on the same networks you are.

@houseofmohny on Instagram →