Bluejay’s cover photo
Bluejay

Bluejay

Software Development

San Francisco, CA 5,644 followers

Test, monitor, and improve voice and chat AI agents

About us

Bluejay helps teams test, monitor, and improve their conversational AI agents across voice, chat, and IVR with powerful simulation and observability tools. We serve everyone from solo developers and fast-growing AI startups to mid-market companies and Fortune 500 enterprises. Conversational AI operates differently than traditional software—agents make autonomous decisions in unpredictable, high-stakes customer interactions. Bluejay gives you the infrastructure to build with confidence throughout your agent's entire lifecycle. Run full simulations with thousands of realistic conversation scenarios to catch issues before launch. Load-test your system under high-traffic conditions to ensure reliability at scale. Monitor production performance in real-time with deep observability into conversation flows, agent decision-making, and failure patterns. Automatically detect regressions, compliance violations, and unexpected behaviors before they impact customers. We help you ship accurate, compliant, and reliable AI experiences—and continuously improve them with data-driven insights. Whether you're prototyping your first voice assistant or managing mission-critical conversational systems at global scale, Bluejay gives you the confidence to build better AI, faster.

Website
https://getbluejay.ai/
Industry
Software Development
Company size
2-10 employees
Headquarters
San Francisco, CA
Type
Privately Held

Locations

Employees at Bluejay

Updates

  • Bluejay reposted this

    Earlier this week at Bluejay we launched MIVAS: the most comprehensive speech-to-speech benchmark for production voice AI. Faraz Siddiqi and Yash Savalia built production-grade, multi-agent simulation environments across healthcare, legal, and customer support. Hyper-realistic scenarios, 72+ detailed tasks, deterministic verifiers scoring every tool call, handoff, and final database state. Then they ran 8 speech-to-speech models through all of it. Three things stood out to me in the data: 1/ There's no single leaderboard winner. Grok Voice 2 leads healthcare at 50.8% pass^5, then drops to 3rd in legal at less than half that score. OpenAI Realtime 2.1 takes legal and customer support instead. 2/ Legal separates the field like nothing else. OpenAI Realtime 2.1 hits 56.0% pass^5 on legal tasks, more than double the runner-up at 25.5%. The bottom of the table passes 5-of-5 runs on just 1.5% of tasks. Healthcare is the opposite story: five harnesses packed within about 11 points. Industry difficulty profiles are wildly different, and averages hide that. 3/ Fast and accurate turned out to be the same thing. We expected a speed/quality trade-off. Instead, the most reliable harness is also the fastest (2.05s latency, 46.8% pass^5), and the slowest (3.89s) is the least reliable (15.0%). Latency and reliability look less like a dial you tune and more like two symptoms of the same underlying capability. Full leaderboard, methodology, and task suites: https://lnkd.in/gv-jHdMb

  • Bluejay reposted this

    I'm excited to unveil Bluejay Labs, and our first open contribution: MIVAS Bench. Bluejay Labs is an audio data lab with a clear mission: Accelerate the adoption of safe, trustworthy, and reliable conversational intelligence systems. Over the next decade, voice will become the default medium of interaction between humanity and technology. We believe speech is more intuitive than keyboards, touchscreens, controllers, or any other interface built to command a machine. As artificial systems surpass human intelligence, human control becomes an ethical priority. Our goal is to build training and evaluation instruments that encourage conversational intelligence models to prioritize human interests with superhuman intelligence. In our time working with frontier labs and enterprises, we have seen the gaps between model development and production, and are committed to building evaluation and training instruments to bridge the gap. Multi-Industry Voice Agent Simulation (MIVAS) Bench is an open source, open data indicator of voice AI performance across economic sectors with over 8600 conversations across industries. MIVAS is comprised of tasks and verifiers with granular rewards executed in production-grade multi-agent RL environments across industries where voice AI is being adopted the fastest. Our primary contributions with MIVAS are the multi-agent topologies we've built from our experience working with voice AI companies in each industry, and in our verification methodology. In our blog, we propose a conjunctive verifier, surpassing traditional limitations Tau Bench and other "DB State Diff" benchmarks encounter. Today, we released mivas-{healthcare, legal, customer-support}. We will be expanding support for industries in the coming weeks. Beginning with MIVAS, we will release the benchmarks and datasets that make trustworthy voice AI development public, mainstream, and easy to do well. Couldn't have shipped this without the brilliant Yash Savalia :) Benchmark link in comments! #ai #ml #aiml #machinelearning #artificialintelligence #startup #yc #voiceai #voice #voiceagents #aiagents #agentic #conversationalai #bluejay #bluejaylabs #audio #audiodata #rl #rlenvironments #selfimproving #rlvr

    • No alternative text description for this image
  • Bluejay reposted this

    "Bluejay is not a tool we use. They are partners." That's Muizz Matemilola at 11x. Our case study with them is out today. Some background on 11x: they've raised over $70M from Andreessen Horowitz and Benchmark, and their agent Julian handles inbound demand for go-to-market teams at companies like Checkr and Canibuild. Julian talks to real prospects, in a customer's name, in the wild. Before Bluejay, validating a single change meant calling the agent over and over, one conversation at a time, and edge case coverage stopped at what the team could imagine. Every gap was a risk in front of a real prospect. Here's why 11x chose to #GoBlue: 1/ Thousands of realistic simulations before an agent ever takes a live call 2/ Automated red teaming graded against OWASP, MLCommons AILuminate, and MITRE ATLAS 3/ Living eval suites that evolve with each customer instead of a checklist that only grows 4/ A self-healing loop where agents improve agents until quality clears the bar 5/ A true partner behind it all. They came to us with one problem, we solved it together, they found more, we solved those too So far that's saved them over 1,100 hours of manual QA. More importantly, their customers deploy faster because they can see the agent already survived scenarios they hadn't thought of themselves. Thank you to Sachi Angle, Muizz Matemilola, Francisco Izaguirre, Aaron Anantharajah, Santiago Gutierrez, and the entire 11x team for sharing their story. Full case study linked in the comments.

  • Bluejay reposted this

    A Fortune 100 company became a paying Bluejay customer overnight. No sales process. No procurement cycle. No enterprise negotiation. They signed up for pay-as-you-go, upgraded to Growth, then Scale, then exceeded Scale. All within days. That's the power of self-serve. But launching PLG taught us more than we expected: 1/ Being closed at first was the right call Hot take: maybe companies shouldn't be self-serve from day one. Bluejay started as a closed, sales-led product. That let us hand-hold every user, hear their concerns directly, and hill-climb fast. Those conversations surfaced insights analytics alone never would have. Self-serve came later, built on over a year's worth of customer conversations. 2/ Learning changes shape When every customer came through a sales call, we learned by talking. Now most users never get on a call, so learning comes from screen sessions, clicks, and behavior. Internal observability became as important as the product itself. We lean heavily on internal agents to diagnose what's happening, spearheaded by Faraz Siddiqi. 3/ PLG still requires proactive and curated outreach Opening up the product doesn't mean pipeline runs itself. When a new customer signs up, we get genuinely excited and want to reach out right away. The line we're always trying to walk: curated and welcoming outreach instead of spam. If someone's stumbling, we reach out to help, not to pitch. That kind of outreach doesn't happen on its own. A signup notification isn't enough. You need to know what they did, what they couldn't do, and what drove them to sign up. Design the system so the right outreach actually happens. 4/ We didn't restrict signups to company emails We debated this internally. Restricting to work emails felt safer. I was also connected with Tommy Fink to advise on PLG best practices, and his take leaned the same direction. We agreed, because we wanted people to be able to try Bluejay freely. The tradeoff: enrichment is critical. We enrich every signup with signals and flags so our team knows who's coming in and how to help. 5/ Pricing: make the barrier to entry as low as possible We optimized our plans around usage. No seat limits. No project limits. No agent limits. The goal was to make it dead simple to start small, see value, and grow with usage, without a big commitment upfront. It's a bet we're still testing. But we're confident because it aligns our incentives with our customers': Bluejay profits when usage grows, and usage only grows when people get real value. Early traction since launch supports this. 6/ People will try to break your product They will spam you. They will probe your limits. Have guardrails and a safety net ready before launch, not after. — Going self-serve touched everything: design, engineering, marketing, sales, pricing. A true team-wide effort, and I couldn't be prouder of the Bluejay team. More to share soon.

  • Bluejay reposted this

    We're giving Bluejay to current YC companies for free. For a full year. 18 months ago, Faraz Siddiqi and I got into YC's P25 batch. The funding mattered. But the thing that carried equal weight: the credits. Being able to use real tools for free while we were figuring everything out changed how fast we could move. Today we're paying that forward. Every company in the current batch gets: - Bluejay's growth plan free for 12 months ($6,000+ in value) - a dedicated support channel to get you running - weekly check-ins with Bluejay engineers And because shipping should feel good, we're throwing in The Bluejay Call Bell. Ring it every time you ship during the batch. The YC community unlocked a lot for us. Excited to give some of that back. If you're in the current batch, come find us. More details in the comments. Sidekick (YC S26) Dialogus (YC S26) Marble (YC S26) Peer Derya (YC S26) Illume Labs (YC S26) talentpluto

    • No alternative text description for this image
  • Bluejay reposted this

    Many teams manually write 5-10 test cases for a knowledge base with 200 pages in it. This doesn’t work. We built a better system that does! We built a RAG system that understands your knowledge base, clusters it into key items to test, and generates test cases from those topics to comprehensively cover your agent’s capability space. This way, we test all of your agent’s knowledge, not just the parts you remember. 🧠 This is the full coverage approach to testing your agent, only on Bluejay. You can check it out by signing up at getbluejay.ai! #ai #ml #aiml #machinelearning #artificialintelligence #startup #yc #voiceai #voice #voiceagents #aiagents #agentic #conversationalai #bluejay

  • Bluejay reposted this

    When Sanas launched their first version of accent translation, Shawn Zhang did not watch the metrics from San Francisco. He flew to the Philippines and lived on a call center floor for a month. This week on Skywatch, Bluejay's car podcast, I sat down with Shawn, co-founder and CTO of Sanas, and we talked about what it actually takes to build speech AI that works in the real world and why the problem they are solving is bigger than most people realize. Shawn came out of Stanford's ML group where he worked under Andrew Ng on computer vision and NLP. When it came time to build Sanas, he knew the hardest problem would not be the model. It would be the data. Their first few commercial deals were not about revenue. They traded their product for access to real call center speech. A few things from this conversation I have not stopped thinking about: 👉 Sanas has collected over eighteen million hours of speech data across six hundred accents and dialects. That is what makes the difference between a good demo and something that actually works in production. 👉 Shawn's ten year vision for Sanas is not a product. It is a utility. Wherever there is a conduit of speech, Sanas sits in and makes it more understood. If you are building in voice AI, speech infrastructure, or customer experience, this one is worth your time. Full episode on YouTube and Spotify in the comments 👇

  • Bluejay reposted this

    Congratulations to Athin Shetty, our first winner of Bluejay's Bug Bounty! Athin was the first to find all three bugs in the Case Western Health agent, and used Bluejay to discover all three failure modes! There are two agents remaining to debug, and $4,000 left in the prize pool. The Bug Bounty closes on August 10th! Happy cracking! #ai #ml #aiml #machinelearning #artificialintelligence #startup #yc #voiceai #voice #voiceagents #aiagents #agentic #conversationalai #bluejay

    • No alternative text description for this image
  • Today, we're giving you a closer look at Bluejay AI, aka Jay. You can tell it to get your agent production-ready, draft a test plan, run simulations against your Digital Humans, and read through the results. You can also point it at production, and it watches real conversations, catching failures the same way. From there, it proposes the fix: concrete changes to your agent, with the reasoning behind them. You review it, approve it, and Bluejay AI pushes the change and reruns to prove it worked. Have Jay start getting your agent production-ready today at getbluejay.ai!

Similar pages