Startup Spotlight: Roark turns failed voice calls into regression tests

James Zammit and Daniel Gauci Mizzi built a QA layer for voice agents; YC says Roark processed more than 10 million call minutes in a six-month period.

By · Published

Primary source: Y Combinator

Why it matters

Roark is betting that voice-agent reliability needs a continuous QA loop, not a one-time demo. Its reported call volume suggests usage; evidence of fewer failures and better customer outcomes would show whether that loop works.

Startup Spotlight: Roark turns failed voice calls into regression tests — James Zammit and Daniel Gauci Mizzi built a QA layer for voice agents; YC says Roark processed more than 10 million call minutes in a six-month period.

Roark was built by co-founders James Zammit and Daniel Gauci Mizzi around a problem they encountered in different industries: a voice agent can sound convincing in a demo and still fail when a real caller interrupts, speaks with an unfamiliar accent, or waits through dead air. Zammit, Roark's CEO, saw the stakes in financial services; Gauci Mizzi, its CTO, worked on systems serving users who would not file a bug report when something went wrong. Their product turns call testing from a manual spot-check into a repeatable process that runs before and after deployment.

That is the case for Roark, a San Francisco-based startup in Y Combinator's Winter 2025 batch. YC's company profile says Roark processed more than 10 million minutes of calls in a six-month period, supporting monitoring and simulation for voice-AI teams, including YC companies. The profile does not give the period's exact start and end dates or a methodology for the minute count, so it is best read as a Roark-reported usage measure, not an independently verified measure of reliability or customer adoption.

YC listed Roark's public launch on August 26th, 2025. The open question is whether the founders can make a systematic quality loop part of how voice agents are built and operated.

Two careers, one stubborn failure mode

Zammit came to the problem after building infrastructure and AI systems at AngelList. YC says he was a senior engineer there as AngelList's assets under management grew from $10 billion to $124 billion, and led Relay, an AI-powered portfolio manager that processed thousands of financial documents each month. Roark's own account gives the work a practical edge: when a financial system reads a number incorrectly aloud, a fluent conversation can still be consequentially wrong.

Gauci Mizzi approached the same issue from consumer software and gaming. YC says he spent seven years at Casumo, leading development of a mobile app used by millions of players, and later worked on Akiflow's mobile-development team. Roark says his time at Akiflow coincided with Akiflow reaching $1.5 million in annual recurring revenue and more than 10,000 customers. Those are background claims reported by Roark and YC; they help explain the founders' operating experience, but do not establish Roark's own revenue or customer count.

The founders describe a shared lesson: transcripts and dashboards can flatten away the details that make a call feel broken. A transcript may show a correct answer while missing a mispronounced name, a long pause, a caller being talked over, or a disclosure delivered unclearly. In Roark's account of how it began, the founders say they first tried transcript review, LLM grading and dashboards. Their conclusion was that the sound and timing of a conversation required its own testing discipline.

Roark says an early attempt to build a dental-clinic voice agent helped make the problem concrete. The agent looped, mishandled insurance confirmation and gave irrelevant answers, while manual calls and transcript review made the faults difficult to isolate. The story is a useful origin because it is about the founders encountering their own product failure, not spotting an abstract market category from a distance.

QA for agents built elsewhere

Roark is an evaluation and monitoring layer for existing voice agents, rather than a platform for building the agents themselves. Its YC profile describes monitoring and evaluation, custom dashboards and alerts, phone and WebSocket simulations, persona-based callers, and graph-based conversation tests that branch into edge cases. It also says the product had more than 40 built-in metrics and automatic identification of up to 15 speakers.

Roark's current product and pricing pages present a wider set of capabilities, including 500-plus metrics, simulation over phone and WebRTC, production-call replay, integrations with tools such as Vapi, Retell, LiveKit and Pipecat, and OpenTelemetry traces. The gap between the 40-plus metrics in the YC profile and 500-plus on the current site likely reflects how the product is now packaged or described; a metric count alone does not show which metrics teams use or how well they predict customer outcomes.

The operating loop is more important than the count. A team can simulate a call using a chosen persona and scenario, score its behavior, monitor production conversations, and convert a failed call into a test to rerun against a revised agent. That offers a practical answer to a recurring engineering problem: every prompt, model or workflow change can alter a system that previously appeared to work. Roark's claim is that failures can become test cases instead of remaining isolated incidents found through a caller complaint or a manual review.

Voice agents can return the right words and still be difficult to use if they speak over the customer, pause too long, or handle a name poorly. Roark says its evaluations include audio-native measures alongside conversational, compliance and performance checks. Buyers will still need to decide whether those measures align with their own definitions of a successful call, and whether the simulations reproduce the conditions their customers actually encounter.

Usage is a start, not proof of outcomes

The 10-million-minute figure is the strongest quantitative traction claim in the YC profile, and it indicates that Roark has been used on meaningful volumes of calls. But call minutes can include production monitoring, simulations, or a mix; the profile does not break that down. It also does not identify the teams behind the number or report how many failures Roark caught, how often a fix passed a regression test, or whether customer outcomes improved. Those missing denominators matter when evaluating a QA product: activity on the platform is not the same as a demonstrated reduction in failed calls.

Roark's pricing page lays out a usage-based model: a free starting tier with $50 in credit, a $500-per-month Team commitment that is described as included usage, and enterprise plans starting at $4,000 per month. Simulation and production evaluation are billed differently, and Roark says telephony and speech-provider costs for simulated calls are passed through at cost. That structure makes it possible for small teams to try the product, while tying costs to call volume and evaluation rather than solely to software seats. The larger commitment also puts the tool into enterprise procurement territory, where security, retention, data residency and service-level requirements can be as important as the test results.

The category is drawing funded competitors. Coval announced a $28 million Series A on June 24th, 2026, led by Norwest. Coval also markets simulation, evaluation and production QA for voice agents. That does not establish Roark's fundraising or market position, but it shows investors are backing the broader premise that deploying voice agents creates a distinct testing and reliability need. Roark's bet is to serve teams across different agent stacks rather than own the agent runtime itself.

The test is whether the loop sticks

Roark lists Y Combinator, F-Prime Capital, True Ventures, Liquid 2 Ventures and MTV among its backers. The public materials reviewed here do not tie those names to a disclosed Roark round size or valuation, so they indicate investor support without establishing the terms or total capital raised.

The founders have a credible reason to focus on the unglamorous part of voice AI. Zammit worked on systems where a spoken number could carry financial consequences; Gauci Mizzi helped build products whose users would simply move on when the experience failed. Roark turns those lessons into a concrete engineering product: make a call repeatable, measure the failure, and check the correction against the same scenario.

Roark's harder task is proving that teams will keep that loop running once the first agent is live. The 10 million minutes show usage, while the test-to-fix workflow explains the product thesis. To validate that thesis, Roark needs customer-level evidence that production failures become tests, regressions get caught before release, and calls improve as a result.

Reader comments

Conversation for this story loads after sign-in.