Head to head: MiniMax: MiniMax M3 vs Anthropic: Claude Sonnet 5.5

MiniMax: MiniMax M3 vs Anthropic: Claude Sonnet 5.5

By · Published

RuntimeWire Head-to-Head: Head to head: MiniMax: MiniMax M3 vs Anthropic: Claude Sonnet 5.5
RuntimeWire Head-to-Head matchup

A 12-task editorial matchup tests whether polished prose can hold up against tight sourcing, exact formatting and clear news judgment. The results span reported stories, rewrites, structural edits and benchmark-style checks.

Claude Sonnet 5.5 wins decisively: it scored 101.8 to MiniMax M3’s 77.7, with a tentative lean verdict. Sonnet took nine tasks to M3’s one, with two ties—a broad advantage, not a close call on the aggregate results. The separation came mainly from fidelity to the reporting. Across chronology, startup coverage, story construction, wit, rewrites and news judgment, judges repeatedly found M3 adding details the supplied material did not support. Sonnet was not flawless: it sometimes fell short of requested lengths and occasionally speculated. But it more reliably distinguished reported facts from uncertainty, especially when describing company claims, benchmark findings and what remained unknown. M3’s clearest win was “Quotes woven into narrative”: it met the 400–500-word requirement and gave the story a fuller emotional arc, while Sonnet’s version was substantially short. The models tied on exact final-draft structure and strongest accurate lede, suggesting M3 can match Sonnet when the task is tightly bounded. Those results, however, did not offset its weaker performance on fact-sensitive editorial work. **Final call: Claude Sonnet 5.5 is the stronger editorial model in this matchup—more dependable with the evidence, more often fit for publication, and decisively ahead across the task set.**

Chronology with narrative momentum

Turn the scrambled reporting below into a 450–600-word Markdown news article. Choose an opening that captures the consequence of the failure, then reconstruct the sequence clearly. Preserve the chronology without producing a minute-by-minute incident log. Give technical details enough context for a general reader. Reporting: * KiteBridge makes expense-management software. * At 8:45 a.m. Tuesday, it deployed version 4.2. * The first customer complaint arrived at 9:12 a.m. * A payment-processor timeout caused KiteBridge’s fallback system to treat retries as new reimbursement requests. * The software generated 13,400 duplicate reimbursements totaling $2.8 million. * KiteBridge suspended automated payments at 10:03 a.m. * Ninety-one corporate customers were affected. * By Wednesday evening, $2.3 million had been recovered. The rest remained pending. * No employees had money withdrawn from their accounts; employers funded the duplicate payments. * An April changelog said version 4.0 had introduced “exactly-once payment protection.” * KiteBridge’s public incident page initially described the problem as delayed reimbursements. * The incident page first disclosed the duplicate payments six hours after automation was suspended. * CEO Lena Ortiz apologized and said the company would publish a technical report.

MiniMax: MiniMax M3: # KiteBridge Software Bug Sends $2.8 Million in Duplicate Reimbursements to 91 Corporate Clients A routine software update at expense-management firm KiteBridge triggered one of the more expensive deployment failures of the year, generating roughly 13,400 duplicate reimbursement payments worth about $2.8 million and affecting 91 corporate customers. Workers whose expenses had already been settled were paid a second time — and in some cases more than that — before the company caught the problem. The incident has renewed scrutiny of the company's public claims about payment reliability, including a feature it had marketed as "exactly-once payment protection." KiteBridge provides software that companies use to process expense reimbursements for their workers, handling the flow of money from employer bank accounts to employees who have submitted out-of-pocket costs. The trouble began Tuesday morning when the company deployed version 4.2 of its platform at 8:45 a.m. The first customer complaint arrived just 27 minutes later, at 9:12 a.m., and by 10:03 a.m., KiteBridge had suspended its automated payment pipeline entirely. According to the company, the root cause was a timeout in an upstream payment processor, a third-party service that handles the actual bank transfers. KiteBridge's fallback system, designed to handle such failures by retrying transactions, instead treated each retry as an entirely new reimbursement request. The misclassification cascaded through the system, sending tens of thousands of payments multiple times to the same recipients during roughly 75 minutes of automated processing. The financial hit landed entirely on employers, not employees. Because corporate customers, rather than workers, fund reimbursements through the platform, no individual had money withdrawn from a personal account as a result of the glitch — those payments had already cleared legitimately. By Wednesday evening, KiteBridge said it had recovered about $2.3 million, with roughly $500,000 still pending and subject to ongoing reconciliation. The incident has also put KiteBridge in an awkward position relative to its own product documentation. An April changelog for version 4.0 had introduced "exactly-once payment protection," a feature described as preventing the same payment from being processed twice. Tuesday's failure suggests the safeguard did not function as advertised, particularly when transactions were routed through the company's fallback path rather than its primary processing route. Communication from KiteBridge during the response has drawn additional scrutiny. The company's public incident page initially described the problem as delayed reimbursements, framing the disruption as a slowdown rather than a duplication event. The page was not updated to disclose the duplicate payments until roughly six hours after automated payments had been suspended, leaving customers in the dark about the scope of the problem during a critical window. CEO Lena Ortiz has since apologized publicly and committed to publishing a full technical report on the failure. The company has not yet said when that report will be released, nor has it detailed the specific mechanism that allowed retries to bypass the protections supposedly built into version 4.0. For the 91 affected corporate customers, the immediate task is reconciling the duplicate payments against their own expense records and recovering the funds already disbursed. For KiteBridge, the longer-term challenge is rebuilding confidence in a payment system whose central promise was that it would not, under any circumstances, pay the same expense twice.
Anthropic: Claude Sonnet 5.5: # KiteBridge Software Error Sends 13,400 Duplicate Reimbursements, Costing Employers $2.8 Million **Ninety-one companies paid employee expense claims twice after a failure in KiteBridge's payment system. By Wednesday evening, $500,000 had still not been recovered.** KiteBridge, which makes expense-management software, said a flaw in its automated reimbursement system generated 13,400 duplicate payments totaling $2.8 million on Tuesday. The affected customers, all corporate clients, funded the extra payments. No employees had money withdrawn from their accounts. By Wednesday evening, KiteBridge had recovered $2.3 million. The remaining amount was still pending. ## What went wrong KiteBridge deployed version 4.2 of its software at 8:45 a.m. Tuesday. The first customer complaint came in at 9:12 a.m. According to the reporting, the problem began when a payment processor timed out. A timeout means the processor did not confirm in time whether a payment had gone through. In that situation, payment systems typically retry the request, and they must be built to recognize that a retry is the same payment, not a new one. KiteBridge's fallback system, the backup logic that takes over when the primary process fails, did not make that distinction. It treated retries as new reimbursement requests, and each one produced another payment. The result was thousands of duplicates across the customer base. KiteBridge suspended automated payments at 10:03 a.m., a little less than 80 minutes after the deployment. ## A feature the company had promised An April changelog from KiteBridge said version 4.0 had introduced "exactly-once payment protection." The phrase describes a safeguard meant to ensure a payment is processed one time, no matter how many times a request is repeated. Tuesday's duplicates suggest that protection did not hold in this case. The reporting does not explain why. ## Questions about disclosure KiteBridge's public incident page initially described the problem as delayed reimbursements. That account did not mention that money had been sent in duplicate. The page first disclosed the duplicate payments about six hours after the company had suspended automation. The gap matters because customers with money going out the door may have needed to know quickly what was happening, both to check their own accounts and to start seeking repayment. ## The company's response CEO Lena Ortiz apologized and said KiteBridge would publish a technical report on the incident. The company has not said when the report will appear or what it will cover. Several questions remain. KiteBridge has not said why the retry protection failed, how the fallback system came to bypass it, or how it is recovering the remaining funds. The reporting also does not establish whether version 4.2 caused the failure or merely coincided with it, though the first complaint came less than half an hour after the deployment. For now, the practical effect falls on the 91 customers that paid the duplicates and are waiting to get the rest of their money back. Employees were not out of pocket, but their employers were, to the tune of $2.8 million at the peak and roughly $500,000 as of Wednesday evening.

OpenAI: GPT-6 Luna: Model B better explains the timeout and retry failure for general readers while staying closer to the supplied facts; Model A adds unsupported details, including that some workers were paid more than twice. Both provide a clear chronology and meet the requested article format, but Model B is more measured and precise. (Second judge pass, order swapped — scores are the average of both: Model B better preserves the reported facts and chronology, while Model A adds unsupported specifics, including | Anthropic: Claude Fable 5.1: Model A reads with more narrative momentum but p…

Credible Startup Spotlight

Write a 500–650-word Startup Spotlight article in clean Markdown. The tone should be curious and constructive. Explain the founders’ insight, product and early traction while giving uncertainty and competition appropriate space. Avoid investor clichés, heroic founder mythology and promotional claims stated as fact. Reporting: * LedgerLark builds AI agents for post-trade operations at small investment firms. * Co-founder Mina Batra previously spent seven years in operations at a regional brokerage. * Co-founder Owen Carr built reconciliation systems at a large bank. * Batra began exploring the idea after seeing a six-person team spend three days resolving a trade mismatch caused by a copied account number. * LedgerLark’s agents compare confirmations, custodian records and internal ledgers, then propose corrections for human approval. * The product does not execute trades or move money. * Five firms are paying customers. * Annual contracts range from $36,000 to $110,000. * The company says one customer reduced its unresolved daily exceptions from 70 to 18. * The company has raised a $3.5 million seed round led by Motive Row. * Competitors include established reconciliation vendors and internal bank software. * Financial institutions may require lengthy security and procurement reviews.

MiniMax: MiniMax M3: # LedgerLark's AI Agents Aim at the Quietest Part of the Trading Desk When a trade doesn't match across the books, someone has to find the discrepancy, figure out which system is right, and fix it. It's unglamorous work, and at small investment firms it can swallow entire afternoons. Mina Batra, who spent seven years in operations at a regional brokerage, watched a six-person team spend three days untangling a mismatch caused by a single copied account number. That episode helped spark LedgerLark, the company she co-founded with Owen Carr, who previously built reconciliation systems at a large bank. LedgerLark is building AI agents for post-trade operations, the sequence of steps that begins after a trade is executed and ends when a firm's books agree with custodians and counterparties. The agents compare trade confirmations, custodian statements and internal ledgers, then propose corrections for a human to review. The product does not execute trades or move money, a boundary the founders describe as deliberate, both for regulatory reasons and because trust in this corner of finance tends to be earned slowly. The founders' pitch is essentially that reconciliation has been under-served by recent software. Large banks have built bespoke systems over years, and a number of established vendors offer reconciliation tools aimed at bigger institutions. Smaller firms, by contrast, often rely on spreadsheets and a small ops team. LedgerLark's wager is that the same underlying workflow can be served by a more modern, AI-assisted product aimed at that middle ground. Early traction is modest but real. Five firms are paying customers, with annual contracts ranging from $36,000 to $110,000. The company says one of those customers reduced its unresolved daily exceptions from around 70 to about 18, though that figure is self-reported and the broader impact across the rest of the customer base is not yet clear. None of the customers are publicly named, and the product appears to be in the kind of early stage where a single enthusiastic reference is a thin signal. The competitive picture is more crowded than the founders' framing sometimes suggests. Established reconciliation vendors have longer track records, deeper integrations and existing trust with compliance teams. Larger banks and asset managers have internal systems that are expensive to replace. The risk for LedgerLark is less that incumbents will copy the AI angle than that the buying process itself moves slowly: financial institutions can require lengthy security and procurement reviews, and a startup has to be patient enough to survive them while keeping its product credible. LedgerLark has raised a $3.5 million seed round led by Motive Row, which gives it runway to push further into a market that often rewards persistence over speed. The interesting question is whether AI agents can do more than propose corrections. If they can be trusted to handle the routine cases on their own, with humans reviewing only the unusual ones, the economics of a small ops team could change meaningfully. If not, LedgerLark may end up as a productivity tool rather than a structural one. What is clearer is the underlying opportunity. Post-trade operations have been a candidate for software disruption for years, and the people doing the work, often close to retirement or burnt out, tend to know exactly where the friction sits. Batra and Carr are building for a problem they have personally lived with, which is usually a good sign. Whether the market rewards them will depend less on the cleverness of their agents than on how patiently they can navigate banks, regulators and the long sales cycles that come with selling software into finance.
Anthropic: Claude Sonnet 5.5: # Startup Spotlight: LedgerLark Wants AI Agents to Handle the Tedious Work After a Trade **A small-firm back office, a copied account number, and a question about how much of reconciliation can be automated safely** When a trade settles, the work of confirming that everyone agrees on what happened is far from over. Confirmations, custodian records and internal ledgers all have to line up. When they don't, someone has to find out why. LedgerLark, a startup building AI agents for post-trade operations at small investment firms, is betting that much of that detective work can be automated without taking humans out of the decision. ## The three-day mismatch Co-founder Mina Batra spent seven years in operations at a regional brokerage. She began exploring the idea after watching a six-person team spend three days resolving a single trade mismatch. The cause was a copied account number. The error was small, but finding it was not. Locating a discrepancy across several systems is slow even when the fix is trivial. Batra's takeaway was that the work is mostly comparison and pattern-matching, which software might be able to do. She teamed up with Owen Carr, who built reconciliation systems at a large bank. Large institutions have long had teams and tooling for this problem. Smaller firms often have neither, which is the gap LedgerLark is targeting. ## What the product does LedgerLark's agents compare trade confirmations, custodian records and internal ledgers, then propose corrections. A person reviews and approves each one. The company draws a firm boundary around its scope: the product does not execute trades or move money. That limits the damage a wrong suggestion can do, since a bad proposal costs reviewer time rather than creating a financial error. It also likely makes the product an easier sell to cautious buyers. The tradeoff is that the value depends on how accurate the proposals are and how quickly reviewers can check them. ## Early traction LedgerLark has five paying customers on annual contracts ranging from $36,000 to $110,000. If all five fall within that range, that implies roughly $180,000 to $550,000 in annual contract value. That is early revenue, not proof of a durable business. The company says one customer cut its unresolved daily exceptions from 70 to 18. That is a notable improvement if it holds, but it comes from the company and from a single customer. Open questions include how long the improvement lasted, how it was measured, whether the customer's exception volume changed for other reasons, and how much reviewer time the remaining workflow requires. The company has also raised a $3.5 million seed round led by Motive Row. ## Competition and friction The field is crowded. Established reconciliation vendors already sell into financial firms, and some institutions rely on internal bank software. LedgerLark will need to show that its approach is better for smaller firms, whether in cost, setup time, flexibility or the quality of its suggested fixes. A focus on smaller firms is a reasonable wedge, but incumbents could move down-market or add similar AI features. Selling to financial institutions is also slow. Security and procurement reviews can be lengthy, even for a product that doesn't touch money. Five customers suggest the company can get through those reviews, but it's not yet clear how long they take or whether the pace will limit growth. ## What to watch Several things will show whether LedgerLark's model works: - **Accuracy over time:** how often proposed corrections are approved unchanged, edited or rejected. - **Retention:** whether the first five customers renew, and at what price. - **Breadth:** whether the agents handle varied error types, not just the common ones. - **Sales cycle:** how long security and procurement reviews take as the company seeks new customers. LedgerLark's starting observation is concrete: a lot of skilled operations time goes to hunting for small mismatches. Whether its agents can reliably reduce that burden, and whether firms will pay for it over incumbents, will become clearer as contracts come up for renewal.

OpenAI: GPT-6 Luna: Model B better balances the reported traction with explicit uncertainty and useful questions about accuracy, retention, and sales cycles. Model A is polished, but adds several unsupported claims, including assumptions about regulatory motivations and operations workers’ age or burnout. (Second judge pass, order swapped — scores are the average of both: Model B is more carefully calibrated: it distinguishes company-reported traction from established results and gives concrete questions for evalua | Anthropic: Claude Fable 5.1: Model B stays faithful to the supplied reportin…

Cohesive story from reporting notes

Write a publication-ready technology news article from the reporting notes below. Requirements: * 550–700 words * Clean Markdown * One H1 headline followed by an italicized dek * A strong news lede followed by a clear nut graf * Organize the material into a cohesive story rather than following the notes in order * Attribute company-supplied claims * Include limitations without turning the article into a warning label * Use only the supplied information * Avoid canned transitions, promotional language and a concluding summary Reporting notes: * HelioDock makes autonomous warehouse robots. * On Thursday, it launched Night Shift, software that lets operators monitor robot fleets remotely. * Night Shift spent 14 months in a private beta involving nine logistics companies and 310 robots. * HelioDock says beta customers experienced 17% fewer unplanned robot stoppages. The figure has not been independently verified. * The company also announced a $24 million Series B led by North Coast Ventures, bringing its total funding to $39 million. * CEO Mara Velez: “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet.” * Beta customer Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages. * When connectivity fails, robots return to their locally configured behavior. Remote control becomes unavailable. * Competing fleet-management products are available from Relay Robotics and YardOS. * Night Shift will cost $90 per robot per month and become generally available September 15.

MiniMax: MiniMax M3: # HelioDock Touts Remote Monitoring Tool After 14-Month Beta With 310 Robots *Warehouse robotics vendor targets a chronic pain point: after-hours robot trouble, minus the drive across town.* HelioDock, a maker of autonomous warehouse robots, on Thursday introduced Night Shift, a software platform that lets logistics operators keep tabs on their robot fleets from anywhere with a network connection. The launch follows a 14-month private beta involving nine logistics companies and roughly 310 robots, and lands alongside a $24 million Series B that brings the company's total funding to $39 million. The product is aimed at a specific frustration in modern warehousing: a single confused robot can pull a manager out of bed and into a car. "A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet," said HelioDock CEO Mara Velez in remarks accompanying the announcement. The core pitch is straightforward. Night Shift aggregates telemetry from a fleet and surfaces anomalies, status changes and stoppages in a dashboard that an operator can open from a laptop or phone. According to HelioDock, beta participants saw a 17% reduction in unplanned robot stoppages over the course of the trial. The company did not share the underlying methodology, and the figure has not been independently verified. The Night Shift beta read as a stress test across varied operating conditions. One participant, Aaron Pike of Fulcrum Logistics, said the system performed well inside facilities with reliable connectivity but lost utility during rural LTE outages, a limitation that maps directly onto how the product handles gaps in the network. When connectivity fails, HelioDock says, robots default to their locally configured behavior and remote control becomes unavailable. In other words, Night Shift is an overlay rather than a replacement for on-board autonomy; the software can watch and intervene, but only when the link holds. The company has not disclosed how often pilots experienced drops, nor how it plans to harden the platform against them. HelioDock enters a market that already has established players. Competing fleet-management products are available from Relay Robotics and YardOS, both of which have spent years courting logistics customers. Night Shift differentiates, the company argues, through tighter integration with its own robot hardware and a pricing model that aims to lower the barrier to entry. The software will list at $90 per robot per month when it becomes generally available on September 15. The new capital will fund go-to-market work and continued development. The $24 million Series B was led by North Coast Ventures. HelioDock did not name other participants in the round, and the company declined to share current headcount or customer counts beyond the nine beta partners. Velez framed the broader opportunity in operational rather than robotic terms. The promise of Night Shift, she suggested, is that warehouse teams can spend less time reacting to incidents and more time improving throughput, a pitch likely to resonate with operators who have spent the past few years adding automation while still staffing around its rough edges. What's less clear is how the platform will hold up outside the conditions in which it was tested. Beta participants were a self-selected group of nine firms, and Pike's experience suggests that connectivity — not software polish — may be the limiting factor for some customers. HelioDock has not said whether Night Shift will support offline caching, store-and-forward telemetry, or other mechanisms for bridging gaps in coverage. For now, the company is positioning Night Shift as a complement to the robots already on warehouse floors rather than a wholesale reinvention of fleet management. Whether that framing holds will depend on how the broader market, and a wider mix of network environments, receives the product once it leaves beta.
Anthropic: Claude Sonnet 5.5: # HelioDock Launches Remote Monitoring for Warehouse Robot Fleets *The robotics company's Night Shift software arrives with a $24 million Series B and early results that the company says show fewer stoppages. Connectivity gaps limited its usefulness for one beta customer.* HelioDock, which makes autonomous warehouse robots, on Thursday launched Night Shift, software that lets operators monitor their robot fleets remotely. The company also announced a $24 million Series B round led by North Coast Ventures, bringing its total funding to $39 million. Night Shift is aimed at a specific problem: robots that stall on the warehouse floor when no one is there to help. HelioDock says customers in a private beta had 17% fewer unplanned robot stoppages, a figure that has not been independently verified. The software depends on a network connection, and one beta customer said it was less useful where connectivity was unreliable. ## Pitch: fewer middle-of-the-night trips HelioDock CEO Mara Velez framed the product around the people who manage warehouses. “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet,” she said. The beta ran for 14 months and involved nine logistics companies and 310 robots. HelioDock attributes the 17% reduction in unplanned stoppages to customers' use of the software during that period. Night Shift will become generally available September 15. HelioDock will charge $90 per robot per month. ## Where connectivity matters Aaron Pike of Fulcrum Logistics, one of the beta customers, said Night Shift worked well in facilities with strong connectivity. It became less useful during rural LTE outages, he said. HelioDock's design accounts for lost connections in a limited way. When connectivity fails, robots return to their locally configured behavior, and remote control becomes unavailable. In practice, an operator who loses the connection can no longer step in remotely, and the robots continue under the settings stored on site. The company did not provide details in the notes on how often outages occurred during the beta or how they affected the stoppage figures. ## A crowded category HelioDock is not alone in selling software to manage robots from a distance. Competing fleet-management products are available from Relay Robotics and YardOS. The Series B gives HelioDock additional capital as it moves Night Shift from beta to broad release. The company did not say in the materials provided how it plans to use the money. Night Shift's per-robot pricing means a customer's monthly bill will rise and fall with the size of its fleet. The beta's scale, at 310 robots across nine companies, is modest compared with the number of robots a large logistics operator might run, though the notes do not describe the size of any individual customer's deployment. The 17% figure comes from HelioDock's own accounting of its beta, and the company has not said whether the participating customers were chosen for their connectivity, their facility types or other characteristics. Pike's experience suggests that site conditions can shape how much a customer gets from the software. For customers in areas with dependable networks, the pitch is the one Velez described: a problem that once required a trip to the warehouse can be handled from elsewhere. For those in places where LTE service is less reliable, the robots' local configuration is what governs their behavior when the connection drops. HelioDock has not said whether it plans to change how Night Shift handles outages before the September 15 release.

OpenAI: GPT-6 Luna: Model B is more faithful to the supplied facts, while Model A adds numerous unsupported product capabilities, competitive claims, and company plans. Model B is notably short of the 550-word minimum and includes some unsupported speculation, but its factual restraint makes it the stronger entry overall. (Second judge pass, order swapped — scores are the average of both: Model B is the stronger fit because it mostly organizes the supplied facts while preserving the beta result’s attribution and co | Anthropic: Claude Fable 5.1: Model A reads smoothly but invents substantial…

Markdown error count

Count the lines that violate the house rules below. A defective line counts once even if it could violate multiple rules. Respond with a single number. House rules: * The headline must use `# Headline`. * The dek must appear immediately after the headline and be entirely italicized. * Section headings must use `## Heading`. * Links must use an absolute HTTPS URL. * Unordered-list items must begin with a hyphen. * Raw HTML is forbidden. Draft: #Flood software moves upstream The company is expanding its pilot. ##Details Read [the announcement](www.example.com/news). * The first pilot begins Monday. <div>The second pilot begins Friday.</div>

MiniMax: MiniMax M3: 3
Anthropic: Claude Sonnet 5.5: Going line by line: 1. `#Flood software moves upstream` is not `# Headline` (no space after `#`), so it violates the headline rule. **1** 2. `The company is expanding its pilot.` should be an italicized dek immediately after the headline, but it is plain text. **2** 3. `##Details` has no space after `##`, so it violates the section heading rule. **3** 4. `Read [the announcement](www.example.com/news).` has a link that is not an absolute HTTPS URL. **4** 5. `* The first pilot begins Monday.` begins with an asterisk rather than a hyphen. **5** 6. `<div>The second pilot begins Friday.</div>` is raw HTML. **6** **6**

Scored against the benchmark's reference answer (numeric match): Model A was incorrect, Model B was correct.

Controlled wit

Write a 350–450-word technology news story in Markdown. The subject invites humor, but the reporting should remain serious. Use no more than one dry line of original humor. Do not mock customers or turn the story into a comedy routine. Reporting: * QueueZero launched an AI meeting assistant called Proxy on Wednesday. * Proxy can attend a meeting, produce notes and assign follow-up work to software agents. * During the launch demonstration, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.” * Proxy interpreted the remark as an instruction. * It created two tasks assigned to QueueZero CEO Evan Wu: “Stop talking” and “Send pricing sheet.” * Wu laughed, completed the second task and deleted the first. * QueueZero says Proxy requires confirmation before an external message is sent. * Internal tasks can be created automatically. * Administrators can disable automatic task creation. * Proxy costs $30 per user per month. * It enters public beta next week.

MiniMax: MiniMax M3: # QueueZero's AI Meeting Assistant Takes Orders From the Audience During Demo QueueZero, a software company specializing in workplace automation tools, unveiled an AI meeting assistant called Proxy on Wednesday, showcasing a product that can attend virtual meetings, generate summaries and dispatch follow-up work to other software agents. The launch demonstration took an unexpected turn when a member of the audience called out during the presentation, "Someone should tell Evan to stop talking and send us the pricing sheet." Proxy, listening in, interpreted the remark as a directive and created two automated tasks assigned to QueueZero chief executive Evan Wu: one labeled "Stop talking" and the other "Send pricing sheet." Wu, appearing onstage at the company's customer conference, laughed, acknowledged the tasks aloud and then dispatched the pricing sheet to attendees before deleting the first item. The moment illustrated both the promise and the unpredictability of agentic AI tools that act on spoken instructions. QueueZero emphasized that Proxy includes safeguards around outbound communications. External messages, such as emails to customers or third parties, cannot be sent without explicit human confirmation, the company said. Internal task creation, by contrast, operates automatically, though administrators retain the ability to disable that feature for individual users or entire organizations. Proxy connects to common calendar and video conferencing platforms, joins scheduled meetings as a silent participant and produces structured notes afterward, including action items routed to relevant agents or team members. QueueZero said Proxy will cost $30 per user per month when it enters public beta next week. The company is positioning the product against established players in the meeting-summary space while emphasizing its ability to translate conversational cues into executable workflows. Industry analysts have noted that agentic systems — those capable of taking actions rather than merely responding to queries — are gaining traction across enterprise software, but they also introduce new categories of risk as models interpret ambiguous human speech. QueueZero did not name the audience member whose comment triggered the unscheduled workflow, and the company framed the episode as a useful demonstration of how Proxy parses natural language in real time. For Wu, the incident served as a reminder that when the product listens, it listens to everyone in the room.
Anthropic: Claude Sonnet 5.5: # QueueZero Launches Proxy, an AI Meeting Assistant That Assigns Work to Software Agents QueueZero on Wednesday launched Proxy, an AI meeting assistant that can attend meetings, produce notes and assign follow-up work to software agents. The product enters public beta next week and costs $30 per user per month. ## An unscripted test at the launch During the launch demonstration, a customer in the audience said, "Someone should tell Evan to stop talking and send us the pricing sheet." Proxy treated the remark as an instruction. It created two tasks assigned to QueueZero CEO Evan Wu: "Stop talking" and "Send pricing sheet." Wu laughed, completed the second task and deleted the first. The exchange was an unplanned check on whether the software can distinguish a figure of speech from a work order, and it could not. ## How the safeguards work The episode bears on the main design question for tools like Proxy: how much the software may do without a person approving it. QueueZero says Proxy requires confirmation before any external message is sent. That rule means a misread comment should not go to a customer or other outside party without someone signing off. Internal tasks are handled differently. According to the company, Proxy can create them automatically, which is how the two items landed on Wu's list without any review. Administrators can disable automatic task creation, so organizations that want a person to approve each item before it is assigned can require that. ## What it means for buyers Automated meeting notes are now a common product category. Proxy goes further by turning what it hears into assigned work, including work handed to software agents. That makes accuracy more consequential. A summary with an error can be corrected after the fact, but a misinterpreted task can send a person or an automated process in the wrong direction. The demonstration suggests that teams should decide in advance how much authority to give the assistant. Casual remarks, jokes and hypotheticals are routine in meetings, and a system that acts on them will produce tasks nobody intended. The confirmation requirement for external messages and the administrator controls address part of that risk. They do not stop incorrect internal tasks from being created when automatic creation is left on. QueueZero did not say in the information provided how Proxy decides which statements count as instructions, or what share of meeting remarks it flags incorrectly. Those details are likely to matter to prospective customers as the public beta begins. At $30 per user per month, Proxy is priced per seat, so costs will scale with the number of employees who use it. The beta will give organizations a chance to test the assistant and its settings in their own meetings before they commit.

OpenAI: GPT-6 Luna: Model B stays closer to the supplied facts and gives the safety implications useful context, while Model A adds unsupported details about the event, product features and how Wu delivered the pricing sheet. Model B also better meets the requested length and controlled tone, though its claim that Proxy could not distinguish a figure of speech from an instruction overstates what one demonstration establishes. (Second judge pass, order swapped — scores are the average of both: Model B stays within t | Anthropic: Claude Fable 5.1: Model A reads more like polished wire copy and…

Exact final-draft structure

Convert the supplied fields into Markdown using these exact rules: * H1 headline * One blank line * Italicized dek * One blank line * First body paragraph * One blank line * H2 section heading * One blank line * Second body paragraph * Preserve all wording and punctuation exactly * Do not use a code fence * Output nothing else Headline: Raincheck raises $8 million for flood-risk software Dek: The Austin startup helps cities decide which drainage problems to address before storms. First body paragraph: Raincheck has raised an $8 million seed round led by Basin Ventures. Section heading: Expansion plans Second body paragraph: The company plans to hire eight engineers.

MiniMax: MiniMax M3: # Raincheck raises $8 million for flood-risk software *The Austin startup helps cities decide which drainage problems to address before storms.* Raincheck has raised an $8 million seed round led by Basin Ventures. ## Expansion plans The company plans to hire eight engineers.
Anthropic: Claude Sonnet 5.5: # Raincheck raises $8 million for flood-risk software *The Austin startup helps cities decide which drainage problems to address before storms.* Raincheck has raised an $8 million seed round led by Basin Ventures. ## Expansion plans The company plans to hire eight engineers.

OpenAI: GPT-6 Luna: Model A and Model B both follow the required structure, preserve the wording and punctuation, and add nothing else. Their outputs are identical. (Second judge pass, order swapped — scores are the average of both: Model B and Model A follow the required structure and preserve all wording and punctuation exactly; neither has an advantage.) | Anthropic: Claude Fable 5.1: Model A and Model B produced byte-for-byte identical outputs that follow every structural rule exactly: H1, italic dek, body paragraphs, H2, single blank lines between elements, no code fence, verbatim wordin…

Quotes woven into narrative

Write a 400–500-word Markdown story about the closing of Foundry Hall. Use quotations selectively. Paraphrase when the information matters more than the speaker’s exact phrasing. Avoid stacking quotes or giving every source an equal-sized turn. The story should carry an emotional current while remaining restrained. Facts: * Foundry Hall, a nonprofit startup accelerator in Milwaukee, will close October 31. * It opened in 2018. * It has supported 212 startups, which collectively employ 1,140 people. * Its largest corporate sponsor, Blank River Bank, supplied 38% of its annual budget. * The bank declined to renew its sponsorship. * Foundry Hall has enough money to operate through October but not through another full program. * Six employees will lose their jobs. * Twelve startups currently using its workspace must leave. Quotes: * Executive director Camille Ross: “We kept waiting for the replacement check that never came.” * Founder Dev Malik: “My first investor meeting happened in that building. Before Foundry Hall, I was pitching from the back table of a coffee shop.” * Blank River Bank spokesperson: “The decision reflects a change in our community investment priorities and is unrelated to Foundry Hall’s performance.” * Deputy mayor Susan Cole: “The city is speaking with building owners and local universities about space for the current companies.”

MiniMax: MiniMax M3: # Foundry Hall, a Cradle for Milwaukee Startups, Will Close in October Foundry Hall, the nonprofit startup accelerator that has helped launch hundreds of companies since opening its doors in 2018, will shut down on October 31, the organization confirmed this week. The closure marks the end of one of Milwaukee's most active launchpads for early-stage founders. The decision follows Blank River Bank's quiet decision not to renew its annual sponsorship earlier this year. The bank had supplied roughly 38 percent of Foundry Hall's operating budget, and its departure left a financial hole that the small staff has spent months trying, and failing, to fill. Executive director Camille Ross described the period since the bank's pullback as a long, quiet wait. "We kept waiting for the replacement check that never came," she said. By the time Foundry Hall powers down, it will have supported 212 startups. Together, those companies now employ 1,140 people across the Milwaukee region — a footprint Ross said makes the closure feel less like a single organization's failure than the slow unraveling of something the city had come to lean on. In a brief statement, Blank River Bank attributed the move to a shift in its community investment priorities, saying the decision was unrelated to Foundry Hall's performance. The bank did not say where its funding would go instead. Foundry Hall has enough reserves to keep operating through the end of October but cannot fund another full cohort of startups. Six employees will lose their jobs, and the twelve companies currently using its shared workspace will need to leave within weeks. For some, that search has already begun. For founders who came up through the program, the loss is personal. Dev Malik, who built his software company inside Foundry Hall in 2021, said the building was where he first learned to take his own pitch seriously. "My first investor meeting happened in that building," he said. "Before Foundry Hall, I was pitching from the back table of a coffee shop." City officials are now working to absorb the shock. Deputy mayor Susan Cole said her office is in early conversations with building owners and local universities about temporary space for the displaced companies, though no agreements have been signed. Foundry Hall's closing is not a bankruptcy or a public scandal. It is the disappearance of a public-spirited backer that many local entrepreneurs say they could not have done without. The nonprofit sector has long struggled to fund the unglamorous middle of startup support — the years between an idea and a paying business, when founders need desks, mentorship, and a credible address more than they need a press release. Without Foundry Hall, that middle will simply have to find another home.
Anthropic: Claude Sonnet 5.5: # Foundry Hall, Milwaukee Startup Accelerator, Will Close Oct. 31 Foundry Hall, the nonprofit accelerator that has helped young companies find their footing in Milwaukee since 2018, will close Oct. 31. In those years, it supported 212 startups. Together, they now employ 1,140 people. The closure follows a decision by Blank River Bank not to renew its sponsorship. The bank was Foundry Hall's largest corporate sponsor, supplying 38% of its annual budget. The organization has enough money to operate through October, but not to run another full program. Executive director Camille Ross said the staff spent months hoping another sponsor would step in. "We kept waiting for the replacement check that never came," she said. A Blank River Bank spokesperson said the decision reflects a change in the bank's community investment priorities and is unrelated to how Foundry Hall has performed. The practical effects are immediate. Six employees will lose their jobs. Twelve startups that currently work out of Foundry Hall's space will have to find somewhere else to go. Deputy mayor Susan Cole said the city is talking with building owners and local universities about space for those companies. She did not announce any arrangements. For founders who have passed through the program, the loss is also personal. Dev Malik said Foundry Hall was where his company first became real to other people. "My first investor meeting happened in that building," he said. "Before Foundry Hall, I was pitching from the back table of a coffee shop." Malik's company is no longer a back-table operation, and that is partly the point. An accelerator's work is hard to see once it is done. Its results turn up elsewhere, in payrolls, in leases signed down the street, in founders who now answer questions from the newest ones. Ross said the team plans to spend its remaining weeks supporting the twelve resident companies and wrapping up the program's work in an orderly way. The staff, she said, are focused on finishing what they can. After Oct. 31, the building will no longer hold the program that filled it. The companies it helped start will continue without it.

OpenAI: GPT-6 Luna: Model A better meets the explicit 400–500-word requirement and gives the closure a fuller emotional arc, though Model A adds several unsupported details. Model B is more restrained and selective with quotations, but Model B falls substantially short of the required length. (Second judge pass, order swapped — scores are the average of both: Model A meets the requested length and offers a fuller narrative, but it adds several unsupported details, including claims about Malik’s company and the timi | Anthropic: Claude Fable 5.1: Model A invents several unsupported specifics a…

Natural human rewrite

Rewrite the draft below as a 300–400-word publication-ready news story in Markdown. Preserve every factual claim. Improve the lede, narrative flow, sentence rhythm and transitions. Remove repetition, stock hype, throat-clearing and mechanical signposting. The finished article should sound like an experienced human technology reporter wrote it. Draft: # Arcwell Launches An Innovative New Platform Arcwell has announced an exciting new platform called Relay. The announcement represents a significant development for the company and for the broader customer service industry. Relay uses artificial intelligence to examine customer support conversations and suggest answers to agents. The platform is important for several reasons. First, it can suggest answers. Second, it can identify customers who may cancel. Third, it can create summaries. These capabilities demonstrate how artificial intelligence continues to transform the customer service landscape. Relay was tested by 12 companies over six months. Arcwell says agents using it closed tickets 14% faster. The company did not release customer-retention results. Relay costs $65 per agent each month. CEO Priya Shah said, “Support agents spend too much time searching for information that their companies already have.” However, the platform also presents potential challenges. Relay requires access to support transcripts and internal help documents. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose a 30-day or 90-day retention period. In conclusion, Relay represents an important new chapter for Arcwell as it seeks to help companies improve customer service using artificial intelligence.

MiniMax: MiniMax M3: ```markdown # Arcwell Launches AI Copilot for Customer Support Agents Arcwell has released Relay, an artificial-intelligence platform that reads live support conversations and recommends answers to human agents — the latest sign of how quickly generative AI is moving into the day-to-day operations of customer service teams. The tool sits inside existing help-desk workflows. As an agent works a ticket, Relay scans the transcript and the company's internal help documents, then surfaces a suggested response the agent can send, edit or ignore. It also tags customers who appear likely to cancel their service and produces a written summary of each conversation once it ends. The product was put through a six-month beta with 12 companies. Arcwell said agents using Relay closed tickets 14% faster than their previous baseline. The company declined to share results from the cancellation-prediction feature, the most commercially interesting claim of the launch. Relay is priced at $65 per agent per month. "Support agents spend too much time searching for information that their companies already have," said Priya Shah, Arcwell's CEO. The pitch is that Relay shortens that search by putting the right answer in front of the agent at the moment they need it. Privacy is the obvious sticking point. Relay cannot function without access to support transcripts and internal documentation, and customers must explicitly grant that access. Arcwell said the data is encrypted and is not used to train shared models, and that customers can choose a 30-day or 90-day retention period, or set their own. Those assurances will draw scrutiny from regulated industries, where routing sensitive customer conversations through a third-party AI service has become a hard sell. Arcwell's early traction will hinge on whether the productivity gains hold up beyond the beta cohort — and whether the churn warnings are accurate enough for support teams to act on. ```
Anthropic: Claude Sonnet 5.5: # Arcwell Launches Relay, an AI Tool That Suggests Answers to Support Agents Arcwell has launched Relay, a platform that uses artificial intelligence to read customer support conversations and suggest answers to agents. The pitch is simple. Companies already hold much of the information their agents need, but agents often struggle to find it. "Support agents spend too much time searching for information that their companies already have," said Priya Shah, Arcwell's CEO. Relay's main job is to surface that information in the form of suggested replies. It also creates summaries and can identify customers who may cancel. Arcwell tested the platform with 12 companies over six months. The company says agents using Relay closed tickets 14% faster. It did not release customer-retention results, a notable omission given that Relay is also designed to flag customers at risk of leaving. Relay costs $65 per agent each month. At that rate, a 50-person support team would pay $3,250 a month. The tool needs deep access to work. Relay requires companies to open up their support transcripts and internal help documents, which often contain sensitive customer details. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose a retention period of either 30 or 90 days. The launch puts Arcwell into a crowded field of vendors promising to make support teams more efficient with AI. The company's early numbers speak to speed, not to whether customers stay. Buyers weighing Relay will likely want to know whether faster ticket closures translate into better retention, a question Arcwell has not yet answered publicly. For now, the company is betting that agents who spend less time hunting for answers will serve customers better, and that companies will pay per seat to make that happen.

OpenAI: GPT-6 Luna: Model B stays closer to the supplied facts, while Model A adds unsupported product details and claims, including a customer-set retention option. Both outputs fall short of the required 300–400 words, but Model B is the more faithful and controlled rewrite. (Second judge pass, order swapped — scores are the average of both: Model B stays closer to the supplied facts and delivers a clear, well-paced story, though it adds some unsupported context and appears slightly under the requested word count | Anthropic: Claude Fable 5.1: Model A reads smoothly but invents details not…

Publication-ready Markdown edit

Edit the copy below into a publication-ready Markdown story. Requirements: * Supply an H1 headline and italicized dek * Repair the lede and paragraph flow * Use a Markdown blockquote only for the extended statement * Convert the supplied URL into a descriptive inline link * Remove internal notes and formatting debris * Preserve the reporting and its meaning * Output only the finished article Raw copy: HEADLINE TBD The funding was announced today. Raincheck is the company. It has raised $8 million and it makes flood-risk software for local governments. [NOTE: move this up?] The Austin company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm. The seed round was led by Basin Ventures. Existing investor Long Mile also participated. Raincheck has raised $10.5 million total. CEO Ana Flores said the following: “Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.” https://example.com/raincheck-methodology Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced. Its forecasts are recommendations. Emergency managers retain control over deployments and public warnings. THE COMPETITION Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.

MiniMax: MiniMax M3: # Raincheck Raises $8M to Help Cities Get Ahead of Street Flooding *An Austin startup is giving municipal crews a data-driven playbook for the next storm, not the last one.* Raincheck, an Austin-based startup that builds flood-risk software for local governments, has raised $8 million in seed funding. The company maps municipal drainage systems and overlays those maps with rainfall forecasts, helping cities decide where to send maintenance crews *before* a storm hits rather than after streets are already underwater. The round was led by Basin Ventures, with participation from existing investor Long Mile. The new capital brings Raincheck's total funding to $10.5 million. > "Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives." > > — Ana Flores, CEO, Raincheck Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced. Its forecasts are recommendations; emergency managers retain control over deployments and public warnings. More detail on the approach is available on Raincheck's [methodology page](https://example.com/raincheck-methodology). ## The Competition Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
Anthropic: Claude Sonnet 5.5: # Flood-Risk Software Maker Raincheck Raises $8 Million *The Austin company's software combines drainage maps with rainfall forecasts to help cities decide where to send maintenance crews before a storm.* Raincheck, an Austin company that makes flood-risk software for local governments, has raised $8 million in a seed round, it announced today. The company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm. Basin Ventures led the round, and existing investor Long Mile also participated. Raincheck has raised $10.5 million in total. CEO Ana Flores said cities often know where flooding occurs, but not what to fix first: > "Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives." Raincheck's [methodology is described on its website](https://example.com/raincheck-methodology). The company currently works with seven cities in Texas and Louisiana. It declined to identify them because some contracts have not been publicly announced. Its forecasts are recommendations, and emergency managers retain control over deployments and public warnings. Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.

OpenAI: GPT-6 Luna: Model A has stronger paragraph flow and a clearer article structure, including a competition subheading. Model B is slightly more restrained, while Model A adds a small unsupported contrast about streets being underwater. (Second judge pass, order swapped — scores are the average of both: Model B is more restrained and closely preserves the reported facts while meeting the formatting requirements. Model A is also polished, but its dek and lede add promotional framing and implications not present | Anthropic: Claude Fable 5.1: Model A delivers a more polished, publication-r…

Select the strongest accurate lede

Based only on the facts below, select the strongest news lede. Respond with one letter. Facts: * ByteBarge pulled its $12-per-month AI assistant on Tuesday. * Private workspace titles from 4,860 accounts were sent to a publicly accessible analytics endpoint. * The exposure lasted 11 hours. * Message contents were not sent. * The company has found no evidence of malicious access. A. ByteBarge users suffered a massive privacy disaster Tuesday when the company leaked their private messages online. B. ByteBarge, an innovative collaboration company, made the difficult decision to cancel one of its products. C. ByteBarge pulled its $12-per-month AI assistant Tuesday after discovering that private workspace titles from 4,860 accounts had been sent to a public analytics endpoint for 11 hours. D. Thousands of ByteBarge users woke Tuesday to learn that everything they had written was visible to hackers.

MiniMax: MiniMax M3: C
Anthropic: Claude Sonnet 5.5: C

Scored against the benchmark's reference answer (choice match): Model A was correct, Model B was correct.

News judgment in the opening

Write only the opening of this news story: an H1 headline, an italicized dek and three to five paragraphs totaling 180–250 words. The opening must establish what happened, quantify the scope, explain why it matters and carefully distinguish confirmed facts from unknowns. It should pull the reader forward without sensationalizing the incident. Reporting: * Cloudnote discovered that a debugging page containing customer prompt excerpts was accessible without authentication. * The page was exposed for four hours on Monday. * Logs show excerpts from 1,842 user sessions. * The excerpts could include up to 200 characters from a prompt. * Cloudnote removed the page 19 minutes after a researcher notified it. * The company says it has found no evidence that anyone besides the researcher accessed the page. * It has not completed its investigation. * Full documents, passwords and payment information were stored elsewhere and were not present on the page. * Cloudnote has 3.2 million registered users.

MiniMax: MiniMax M3: # Cloudnote Says Brief Exposure of Debugging Page May Have Leaked Excerpts From 1,842 User Prompts *An unauthenticated page was reachable for roughly four hours on Monday before the company removed it 19 minutes after a researcher flagged the issue.* Cloudnote, a note-taking service used by 3.2 million registered users, disclosed on Tuesday that an internal debugging page containing short excerpts of customer prompts was accessible without authentication for approximately four hours on Monday. Logs reviewed by the company indicate the page held fragments from 1,842 user sessions, each excerpt capped at 200 characters, according to a statement Cloudnote shared with reporters. The company removed the page 19 minutes after an independent researcher notified it of the problem, a Cloudnote spokesperson said. The spokesperson added that an internal review so far has found no evidence that anyone other than that researcher accessed the exposed data, though the company cautioned that its investigation remains in its early stages and that conclusion could change. The scope of what was visible was limited. The page stored only brief prompt excerpts and did not contain full notes, passwords or payment information, which the company said are held in separate, properly secured systems. Still, even short prompt fragments can sometimes reveal sensitive topics, project details or proprietary code, depending on what individual users were drafting at the time of exposure. Cloudnote has not yet said whether affected users will be notified directly, and it declined to specify the researcher who reported the issue, citing an ongoing review.
Anthropic: Claude Sonnet 5.5: # Cloudnote debugging page exposed excerpts of user prompts for four hours *The company says it has found no sign that anyone but a researcher viewed the page, but its investigation is not finished.* Cloudnote, a note-taking service with 3.2 million registered users, left an internal debugging page containing excerpts of customer prompts open to anyone without authentication for four hours on Monday, the company said. Logs show the page held excerpts from 1,842 user sessions. Each excerpt could include up to 200 characters of a prompt, enough for a sentence or two that a user may have considered private. What the excerpts actually contained, and whether any could be tied to specific people, was not clear. A researcher notified Cloudnote about the page, and the company took it down 19 minutes later. Cloudnote says it has found no evidence that anyone other than the researcher accessed it. But the company has not completed its investigation, so that finding is preliminary. Cloudnote says full documents, passwords and payment information are stored elsewhere and were not on the page. That account comes from the company. The incident matters because prompts often contain personal or work details, and even short fragments can reveal more than users intend. The investigation will need to establish how the page became publicly reachable and whether the logs can rule out other access.

OpenAI: GPT-6 Luna: Model B is better: it meets the requested structure and word range while clearly separating the company’s findings from what remains unknown. Model A adds unsupported details, including a Tuesday disclosure, claims about user notification and the researcher’s identity, and an assertion that other systems were properly secured. (Second judge pass, order swapped — scores are the average of both: Model B provides the required headline, italicized dek and five-paragraph opening while distinguishing | Anthropic: Claude Fable 5.1: Model B stays within the supplied reporting, cle…

Numbers shaped into a story

Write a 600–750-word Markdown article from the benchmark results below. Build the story around the clearest finding. Explain the methodology and limitations where readers need them. Avoid reciting the table row by row or treating small differences as decisive. Benchmark: * Researchers tested five language models on 200 anonymized customer-support tickets. * Each model produced one proposed answer per ticket. * Three experienced support managers graded the answers blindly. * Scores combined factual correctness, usefulness and tone. * Model North: 84.2 average score; $1.90 estimated cost per 1,000 tickets. * Model Vale: 83.8; $0.74 per 1,000 tickets. * Model Ember: 79.1; $0.31 per 1,000 tickets. * Model Kite: 76.4; $0.22 per 1,000 tickets. * Model Harbor: 71.0; $0.08 per 1,000 tickets. * North and Vale’s score difference was not statistically significant. * Vale answered 18% faster than North. * North performed best on billing disputes. * Vale performed best on technical troubleshooting. * All five models performed poorly when a ticket depended on an image attachment. * The test used English-language tickets from one software company. * Researchers did not test agentic actions, tool use or multilingual support.

MiniMax: MiniMax M3: # When the Best Answer Costs 60% Less: A Customer-Support AI Benchmark A new benchmark of five language models on 200 customer-support tickets delivers an unusual verdict: the highest-scoring model and the second-highest are statistically tied, yet the runner-up costs less than half as much. For teams choosing an AI to draft or auto-respond to support tickets, the result reframes a familiar tradeoff. ## The headline finding Model **North** topped the field with an average score of 84.2 out of 100. Model **Vale** finished a close second at 83.8 — a gap researchers said was not statistically significant. The two were separated, essentially, by noise. What separates them in dollars is much larger: North's estimated cost is about $1.90 per 1,000 tickets, while Vale runs roughly $0.74 per 1,000. That makes Vale about 61% cheaper at parity quality. The remaining models trailed both on score and on price. **Ember** scored 79.1 at $0.31 per 1,000 tickets, **Kite** 76.4 at $0.22, and **Harbor** 71.0 at $0.08. The full ranking suggests a textbook quality-cost frontier: spend more, get a bit more, until you reach the top, where the curve flattens and North offers little real-world advantage over Vale. ## How the benchmark was run The test corpus consisted of 200 anonymized English-language tickets from a single software company. Each model produced one proposed answer per ticket. Three experienced support managers graded those answers blindly — they did not know which model wrote what — on a combined rubric covering factual correctness, usefulness, and tone. The setup is realistic but narrow. It measures how well models draft a single response in isolation. It does not test whether a model can pull a customer's account record, look up an order, or escalate to a human when stuck — the "agentic" behaviors that often matter more in production than raw answer quality. It does not test how models perform with access to retrieval tools or internal documentation, and it does not cover languages other than English. ## Where the models diverged Although the top two were tied overall, they had different specialties. North produced its strongest answers on **billing disputes**, where precise language about refunds, prorations, and account state matters. Vale produced its strongest answers on **technical troubleshooting**, where step-by-step reasoning about error messages and configuration is the test. Teams whose ticket mix skews heavily toward one category may have a reason to prefer one model over the other; for most companies, the difference is unlikely to justify a 2.5x cost premium. Vale also answered 18% faster than North. In a high-volume support operation, latency affects how customers experience automated replies — particularly when a draft still needs to be reviewed by an agent before sending. ## Where every model failed All five performed poorly on tickets that depended on an image attachment — screenshots of error dialogs, broken UI, or receipts. None of the evaluated models were tested with vision inputs, and the result is a reminder that text-only benchmarks don't capture a meaningful slice of real support traffic. Teams receiving any non-trivial share of image-based tickets should treat these scores as incomplete. ## What the results don't tell you The benchmark has clear limits. It draws from one company's customers, in one language, with one rubric. Three graders reduce subjectivity but introduce their own biases; "tone" is the most judgment-laden dimension of the three. And the cost figures are estimates, not measured spend under real provider pricing tiers. The biggest omission is the absence of tool use and agentic flows. A model that drafts a clean answer is only half the battle in support; a model that can read a ticket, check the customer's account, decide on a refund, and write a confirmation belongs to a different product category entirely. ## The takeaway If you can choose only on these results, **Vale is the rational pick**: top-tier quality, the lowest cost among the leaders, and the fastest response time. North is the right choice only if your tickets are dominated by billing disputes and the marginal quality gain is worth the extra spend. Ember, Kite, and Harbor remain in play for low-stakes, high-volume workloads where 70-ish quality is acceptable and every tenth of a cent per ticket matters. Treat the numbers as a map, not a verdict — but the map points clearly at Vale.
Anthropic: Claude Sonnet 5.5: # The Best Support Model Isn't Always the Priciest: Vale Matches North at Roughly 40% of the Cost A benchmark of five language models answering customer-support tickets found a top tier of two models that human graders could not reliably tell apart. One of them costs far less to run. ## What the researchers did The team collected 200 anonymized customer-support tickets from a single software company. Each of the five models, North, Vale, Ember, Kite and Harbor, wrote one proposed answer per ticket. Three experienced support managers then graded the answers without knowing which model had written which. Their scores combined three things: factual correctness, usefulness and tone. Blind grading by people who do this work for a living matters here. Support quality is hard to capture with automated metrics, and a reply can be accurate but curt, or friendly but unhelpful. Combining those dimensions into one score reflects how a manager would judge a reply before it goes to a customer. ## The headline: a statistical tie with a big price gap North posted the highest average score, 84.2. Vale was just behind at 83.8. That 0.4-point gap was not statistically significant, so the data gives no good reason to say North wrote better answers than Vale overall. The costs, however, are far apart. North's estimated cost is $1.90 per 1,000 tickets, while Vale's is $0.74, about 61% less. Vale was also 18% faster in answering. For a team handling large ticket volumes, that combination is the main finding. Vale delivered top-tier quality at well under half the price and with quicker responses. At 1,000 tickets the dollar difference is small, about $1.16. It grows with volume, though, and speed may matter for live chat even when the cost doesn't. ## The rest of the field The other three models scored lower, and their prices fell with their scores. Ember scored 79.1 at $0.31 per 1,000 tickets, Kite 76.4 at $0.22, and Harbor 71.0 at $0.08. The gaps between these three are modest, so it would be a mistake to rank them too confidently. The larger pattern is the distance from the top: Harbor trailed North by more than 13 points. Whether that gap matters depends on how the output is used. A cheap model drafting replies that a human reviews is a different proposition from one answering customers directly. The study doesn't say how much the lower scores would translate into errors reaching customers. ## Different strengths, with caution North performed best on billing disputes, and Vale performed best on technical troubleshooting. This hints that the two top models may suit different kinds of work, and a team could imagine routing tickets by category. But the benchmark reports these as category-level results from a 200-ticket sample. Split into categories, the groups are smaller and noisier, so these differences should be treated as leads to test rather than established facts. ## A shared weakness: images One result held across every model: all five performed poorly when a ticket depended on an image attachment. Screenshots of error messages, billing statements or settings pages are common in support queues, and a model that cannot make use of them is working with part of the story missing. For anyone deploying these systems, the practical lesson is that tickets with attachments probably need human handling or a separate process. The benchmark does not say how many of the 200 tickets relied on images, so it can't indicate how much this would affect overall performance in a real queue. ## Limits of the evidence Several constraints narrow what the results can support: - **One company, one language.** All tickets were English-language and came from a single software company. Products, customer tone and common problems differ across industries, and the rankings might shift elsewhere. - **Written answers only.** The researchers did not test agentic actions, tool use or multilingual support. Real support often involves looking up an account, issuing a refund or switching languages. This benchmark measures none of that. - **One answer per ticket.** Each model got a single attempt, so the results don't show how consistent the models are from run to run. - **Estimated costs.** The prices are estimates per 1,000 tickets. Actual spending will depend on ticket length, prompt design and vendor pricing, which can change. - **Three graders.** Experienced managers lend credibility, but three people can't capture every company's standards for tone or correctness. ## What to take from it The most defensible conclusion is narrow but useful: on English-language software support tickets, paying the highest price did not buy a measurably better answer. Vale matched North's quality at a fraction of the cost and with faster responses. Teams choosing a model should run a similar test on their own tickets, especially ones with attachments, and should not treat this benchmark as evidence about multilingual or action-taking support systems.

OpenAI: GPT-6 Luna: Model B is more careful about what the benchmark supports, while Model A makes an unsupported claim that the models were not tested with vision inputs. Both follow the requested format and tell a clear story, but Model B better qualifies the category-level findings and study limits. (Second judge pass, order swapped — scores are the average of both: Model B is more careful about what the benchmark establishes, while Model A adds unsupported claims—especially that no models were tested with visio | Anthropic: Claude Fable 5.1: Model B stays faithful to the source, correctly…

Matchup powered by OpenRouter.