House of MohnyTest protocol ยท v1

Voice Platform Test Protocol

How to test AI voice platforms on an Indian phone number in an afternoon, and the scorecard to run them through. This is the method โ€” the results are yours, because the only results that matter are on your number, your accent, your callers.

Why this isn't a ranked list

Every voice-platform comparison you'll read was run in a different country, on a different codec, in a different accent, on a version that shipped months ago. A benchmark you didn't run is a marketing claim. This protocol takes an afternoon and gives you a number you can actually defend.

The criteria

Three things, nothing else


Feature lists are a distraction. Three things decide whether a voice agent survives contact with a real Indian caller.

Latency to first audio

Time from the caller finishing a sentence to the first syllable coming back. Measure it on a real call, over the actual phone network โ€” not in the browser demo on the vendor's site, which skips the part that's slow.

Fail thresholdOver 1.5 seconds. Callers begin talking over it or assume the line dropped.

Code-switch survival

Say one sentence that changes language halfway โ€” "Tuesday ko appointment mil jayega?" โ€” and see whether both halves survive transcription. This is the single most common failure and almost no vendor tests for it.

Fail thresholdEither half dropped, or the reply answers a question you didn't ask.

Calendar write

Not "can it discuss availability" โ€” can it put a correctly-dated booking into a real calendar, unattended? Check the date, not just that something appeared.

Fail thresholdWrong day, wrong timezone, or a booking it claims to have made and didn't.
Method

How to run it


Scorecard

Fill this in


Print it or copy it. Median of three calls per row.

PlatformLatencyCode-switchCalendarInterruptVerdict
1.___ mspass / failpass / failpass / fail___
2.___ mspass / failpass / failpass / fail___
3.___ mspass / failpass / failpass / fail___
4.___ mspass / failpass / failpass / fail___
5.___ mspass / failpass / failpass / fail___

Any platform failing code-switch is out, regardless of latency. A fast agent that mishears half your callers is worse than a slow one that doesn't.

What usually happens

Most people run this expecting latency to be the differentiator and find that code-switching is. The platforms cluster within a few hundred milliseconds of each other and separate hard on whether they can follow a sentence that changes language halfway through.

Which is the whole reason a comparison run in California doesn't tell you anything useful here.

Run it and reply to the DM with your scorecard. I'll tell you what I'd deploy and where I'd expect it to break โ€” and if your numbers contradict mine, I want to know that more than you do.

@houseofmohny ยท Bengaluru