Bluejay reposted this
Earlier this week at Bluejay we launched MIVAS: the most comprehensive speech-to-speech benchmark for production voice AI. Faraz Siddiqi and Yash Savalia built production-grade, multi-agent simulation environments across healthcare, legal, and customer support. Hyper-realistic scenarios, 72+ detailed tasks, deterministic verifiers scoring every tool call, handoff, and final database state. Then they ran 8 speech-to-speech models through all of it. Three things stood out to me in the data: 1/ There's no single leaderboard winner. Grok Voice 2 leads healthcare at 50.8% pass^5, then drops to 3rd in legal at less than half that score. OpenAI Realtime 2.1 takes legal and customer support instead. 2/ Legal separates the field like nothing else. OpenAI Realtime 2.1 hits 56.0% pass^5 on legal tasks, more than double the runner-up at 25.5%. The bottom of the table passes 5-of-5 runs on just 1.5% of tasks. Healthcare is the opposite story: five harnesses packed within about 11 points. Industry difficulty profiles are wildly different, and averages hide that. 3/ Fast and accurate turned out to be the same thing. We expected a speed/quality trade-off. Instead, the most reliable harness is also the fastest (2.05s latency, 46.8% pass^5), and the slowest (3.89s) is the least reliable (15.0%). Latency and reliability look less like a dial you tune and more like two symptoms of the same underlying capability. Full leaderboard, methodology, and task suites: https://lnkd.in/gv-jHdMb