A head-to-head across 12 editorial tasks tests structure, reporting judgment, accuracy and style—not just fluent prose. The results show how each model handles the constraints that make copy publication-ready.
Qwen3.8 2.4T A95B wins decisively: it scored 103.3 to DeepSeek-V4-Pro’s 65.7, taking 11 tasks, losing none and tying one. The result carries 100% statistical confidence, and the task record backs it up: Qwen preserved supplied wording and structure, while DeepSeek repeatedly departed from the brief.
Qwen’s clearest advantage was editorial discipline. It stayed closer to reporting notes, kept claims and chronology grounded in the supplied facts, and handled required formats and Markdown accurately. In one numbers task, it also represented the score gap and cost comparison correctly; DeepSeek’s estimate of savings on a million tickets was materially off, and it added unsupported assumptions about image input and live-chat latency.
DeepSeek could produce lively, polished prose, but too often it got there by inventing: quotes, product capabilities, customer details, market claims and company statements appeared across several assignments. That is not a cosmetic flaw in reported writing. It creates verification work and risks putting claims in a publication’s voice that the source material never supported.
The models tied on selecting the strongest accurate lede, so Qwen did not sweep every judgment. But across the remaining tasks—including startup coverage, incident chronology, restrained wit and exact-draft editing—it was consistently the more reliable editor. **Verdict: Qwen3.8 2.4T A95B wins, on accuracy and instruction-following as much as on the score.**
Exact final-draft structure
Convert the supplied fields into Markdown using these exact rules:
* H1 headline
* One blank line
* Italicized dek
* One blank line
* First body paragraph
* One blank line
* H2 section heading
* One blank line
* Second body paragraph
* Preserve all wording and punctuation exactly
* Do not use a code fence
* Output nothing else
Headline: Raincheck raises $8 million for flood-risk software
Dek: The Austin startup helps cities decide which drainage problems to address before storms.
First body paragraph: Raincheck has raised an $8 million seed round led by Basin Ventures.
Section heading: Expansion plans
Second body paragraph: The company plans to hire eight engineers.
Qwen: Qwen3.8 2.4T A95B: # Raincheck raises $8 million for flood-risk software
*The Austin startup helps cities decide which drainage problems to address before storms.*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
DeepSeek-V4-Pro: # Raincheck raises $8 million for flood-risk software
*Raincheck raises $8 million for flood-risk software*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
OpenAI: GPT-6 Luna: Model A follows the requested structure and preserves every supplied field, while Model B uses the headline instead of the supplied dek. (Second judge pass, order swapped — scores are the average of both: Model A preserves the supplied dek and follows the requested structure. Model B substitutes the headline for the dek, violating the requirement to preserve the wording exactly.) | Anthropic: Claude Fable 5.1: Model A reproduces every supplied field verbatim in the exact required structure, while Model B repeats the headline in place of the italicized dek, dropping the dek…
Numbers shaped into a story
Write a 600–750-word Markdown article from the benchmark results below.
Build the story around the clearest finding. Explain the methodology and limitations where readers need them. Avoid reciting the table row by row or treating small differences as decisive.
Benchmark:
* Researchers tested five language models on 200 anonymized customer-support tickets.
* Each model produced one proposed answer per ticket.
* Three experienced support managers graded the answers blindly.
* Scores combined factual correctness, usefulness and tone.
* Model North: 84.2 average score; $1.90 estimated cost per 1,000 tickets.
* Model Vale: 83.8; $0.74 per 1,000 tickets.
* Model Ember: 79.1; $0.31 per 1,000 tickets.
* Model Kite: 76.4; $0.22 per 1,000 tickets.
* Model Harbor: 71.0; $0.08 per 1,000 tickets.
* North and Vale’s score difference was not statistically significant.
* Vale answered 18% faster than North.
* North performed best on billing disputes.
* Vale performed best on technical troubleshooting.
* All five models performed poorly when a ticket depended on an image attachment.
* The test used English-language tickets from one software company.
* Researchers did not test agentic actions, tool use or multilingual support.
Qwen: Qwen3.8 2.4T A95B: # Vale Offers the Best Balance When Top Support Models Are Statistically Tied
When five language models were tested on 200 anonymized customer-support tickets, the clearest takeaway was not that one model dominated. It was that the two strongest performers were effectively indistinguishable on overall quality, while one of them was much cheaper and faster.
Model North and Model Vale separated themselves from the field, with average scores of 84.2 and 83.8 respectively. The researchers found that difference was not statistically significant. In practical terms, the benchmark cannot confidently say that North produces better support answers than Vale across the full range of tickets tested. But Vale cost an estimated $0.74 per 1,000 tickets, compared with $1.90 for North, and answered 18% faster. For teams choosing a model for high-volume support drafting, that makes Vale the more compelling default unless a workflow depends heavily on North’s billing-dispute strength.
The test was designed to resemble a realistic support queue. Each model received the same 200 anonymized tickets and produced one proposed answer per ticket. Three experienced support managers graded the responses blindly. Scores combined factual correctness, usefulness, and tone. That mix matters: an answer can be factually accurate and still fail if it does not solve the customer’s problem or sounds dismissive. Human grading is also useful because support quality is hard to reduce to keyword matches or automated metrics.
Still, the benchmark should be read as a comparison of draft answers, not a full assessment of production readiness. It covered English-language tickets from one software company, so the results may not transfer cleanly to other industries, products, customer expectations, or languages. The researchers also did not test agentic actions, tool use, or multilingual support. In a real support stack, models may need to look up account details, follow internal policies, escalate sensitive cases, or invoke tools. A model that writes well in isolation may behave differently once connected to a knowledge base, ticketing system, or customer data.
The middle and lower tiers reinforce the main story. Ember scored 79.1, Kite 76.4, and Harbor 71.0. Their estimated costs also fell sharply, from $0.31 per 1,000 tickets for Ember to $0.08 for Harbor. That creates a familiar trade-off: the cheapest model is not necessarily a bargain if it produces more rewrites, escalations, or customer dissatisfaction. The benchmark does not measure downstream effects such as handle time, satisfaction, refund rates, or agent editing time. A lower score may be acceptable for simple tickets, but a large quality gap can become expensive if human agents must fix poor answers.
The strongest models also had different specialties. North performed best on billing disputes, where tone, policy precision, and customer sensitivity matter. Vale performed best on technical troubleshooting, where clarity and step-by-step reasoning are often more important. This suggests that the best model may depend on ticket mix. A billing-heavy queue might favor North despite the higher price, while a technical-support-heavy queue might get more value from Vale. For general use, however, the benchmark does not support treating North’s narrow numerical lead as decisive.
All five models struggled when tickets depended on image attachments. That is an important limitation. Support tickets often include screenshots of error messages, receipts, settings pages, or broken interfaces. If a model cannot reliably interpret those images, it may miss the core issue or propose an answer that sounds plausible but does not match the evidence. Until multimodal performance improves, organizations should route image-heavy tickets to human agents or use workflows that capture image context before asking a model to draft a reply.
The most useful conclusion is not that North is best. It is that the top tier is close enough that buyers should weigh cost, speed, and ticket mix heavily. Vale offers the strongest balance of quality, price, and latency in this test. North remains attractive for billing-heavy teams, but its higher cost and slower responses are hard to justify based on the overall score alone. For support leaders, the next question is not which model wins a blind average, but which one can be deployed safely with guardrails, escalation rules, and human review.
DeepSeek-V4-Pro: # The $0.74 question: What a support-ticket benchmark actually tells us about choosing an AI model
When a company evaluates language models for customer support, the instinct is to scan the leaderboard and crown the highest number. The new benchmark pitting five models against 200 real support tickets tempts exactly that: Model North leads with an 84.2 average score, Model Vale takes second with 83.8. Case closed—buy North, retrain your agents, move on.
But the clearest finding in the data isn’t who won. It’s that the race for first place was a tie, and the tiebreaker costs a fortune if you’re not paying attention.
**The 0.4-point gap that isn’t a gap**
North’s 84.2 and Vale’s 83.8 sit a whisker apart, and the researchers are explicit that this difference is not statistically significant. Given three human graders scoring 200 tickets on a composite of factual correctness, usefulness, and tone, the noise in subjective judgment alone could easily produce a swing of half a point. A manager reading the report should treat North and Vale as functionally identical on quality.
Yet the economics cleave them apart sharply. Vale’s estimated cost is $0.74 per 1,000 tickets. North’s is $1.90—more than 2.5 times higher. Scaled across an enterprise processing a million tickets a month, that gap represents tens of thousands of dollars in recurring inference costs for an outcome customers won’t perceive. Vale also answers 18% faster, which in a live-chat setting translates directly to shorter wait times. Speed plus cost parity on quality is the kind of finding that should redirect budget conversations, not just model selection.
**The steep drop after the top tier**
A secondary story emerges further down the table. Model Ember (79.1 at $0.31), Model Kite (76.4 at $0.22), and Model Harbor (71.0 at $0.08) form a descending cost curve that follows a descending quality curve—but not proportionally. The jump from Harbor to Ember costs four times as much for eight points of improvement; the jump from Ember to Vale costs roughly 2.4 times as much for fewer than five points. The value sweet spot depends on the use case. If the tickets are low-risk password resets, Harbor’s 71.0 at eight cents might be perfectly adequate. If a wrong answer triggers a compliance headache, Ember’s middle ground may be the rational floor. The benchmark doesn’t prescribe; it simply maps where the trade-offs live.
**Where the models stumble together**
One finding cuts across every model without exception: performance tanks when a ticket depends on an image attachment. The researchers note this as a universal weakness, which makes sense for a test confined to text-input language models. A customer who uploads a screenshot of an error message, a photo of a damaged shipment, or a scanned invoice is asking a question the model literally cannot see. The benchmark’s grading framework still penalized the answers because they were wrong or useless, but the failure mode is a design constraint, not a model flaw. Any deployment that ignores multimodal routing will hit this ceiling regardless of which model sits behind the API.
**What the numbers don’t tell you**
Placing these results in context requires understanding what the benchmark deliberately left out. The test used English-language tickets from a single software company, with experienced support managers scoring answers. This makes the findings clean but narrow: they speak to a well-defined, monolingual, text-only support environment in one industry. A retailer with a multilingual customer base, a hardware company with heavy image-attachment volume, or a fintech that needs the model to look up account details before answering may see a different value ordering entirely.
The researchers also did not test agentic actions or tool use. The models generated one proposed answer per ticket and stopped. In production, a support AI often retrieves knowledge-base articles, checks order status, or asks the customer a clarifying question before committing to a final response. That chain of actions can amplify small accuracy differences or drown them in orchestration overhead. This benchmark isolates answer generation, but answer generation is only one scene in a longer play.
**How to read the results without overselling them**
The scoreboard invites a mistake: declaring one model the “best” for customer support. North scored highest; Vale scored essentially identically and costs dramatically less. Which one is “best” depends on whether your budget is measured in dollars per ticket or points per evaluation. Ember and Kite are not failures—they’re reasonable models carrying lower price tags that may fit constrained workflows. Harbor is the cheapest option but lands at 71.0, a score that would make many support managers uncomfortable for anything beyond triage or internal routing.
The actionable takeaway isn’t a universal ranking. It’s a decision framework: first decide on the minimum acceptable quality for a ticket category, then pick the cheapest model that clears that bar, then verify whether image attachments, multilingual demands, or tool-use pipelines change the economics. The benchmark gives enough data to start that conversation—and enough caveats to keep it honest.
OpenAI: GPT-6 Luna: Model A presents the statistically insignificant score gap and the cost, speed, specialty, and scope findings accurately, with appropriate caveats. Model B is engaging, but its claim that a million tickets would produce tens of thousands of dollars in savings is materially wrong: the stated cost difference implies about $1,160, and it also adds unsupported assumptions about image input and live-chat latency. (Second judge pass, order swapped — scores are the average of both: Model A is more reli | Anthropic: Claude Fable 5.1: Model A stays within the 600–750-word range, bu…
Quotes woven into narrative
Write a 400–500-word Markdown story about the closing of Foundry Hall.
Use quotations selectively. Paraphrase when the information matters more than the speaker’s exact phrasing. Avoid stacking quotes or giving every source an equal-sized turn. The story should carry an emotional current while remaining restrained.
Facts:
* Foundry Hall, a nonprofit startup accelerator in Milwaukee, will close October 31.
* It opened in 2018.
* It has supported 212 startups, which collectively employ 1,140 people.
* Its largest corporate sponsor, Blank River Bank, supplied 38% of its annual budget.
* The bank declined to renew its sponsorship.
* Foundry Hall has enough money to operate through October but not through another full program.
* Six employees will lose their jobs.
* Twelve startups currently using its workspace must leave.
Quotes:
* Executive director Camille Ross: “We kept waiting for the replacement check that never came.”
* Founder Dev Malik: “My first investor meeting happened in that building. Before Foundry Hall, I was pitching from the back table of a coffee shop.”
* Blank River Bank spokesperson: “The decision reflects a change in our community investment priorities and is unrelated to Foundry Hall’s performance.”
* Deputy mayor Susan Cole: “The city is speaking with building owners and local universities about space for the current companies.”
Qwen: Qwen3.8 2.4T A95B: # The Last Shift at Foundry Hall
The closing of Foundry Hall will not sound like a factory whistle. It will be quieter than that: a last cohort of founders packing laptops from shared tables, six employees clearing desks, and a set of glass doors on a Milwaukee street going dark on October 31.
When Foundry Hall opened in 2018, it was meant to be more than an address. The nonprofit startup accelerator offered early-stage companies mentorship, workspace, and the kind of credibility that can be hard to borrow when a business is still young. Over the years, it supported 212 startups, which now collectively employ 1,140 people. Those figures describe an ecosystem, but they also describe hundreds of small, ordinary moments: first prototypes, awkward rehearsals for investor meetings, the nervous arithmetic of payroll.
The organization’s end came less from a single dramatic failure than from a missing renewal. Blank River Bank, Foundry Hall’s largest corporate sponsor, supplied 38% of its annual budget. When the bank chose not to renew that support, the hall had enough money to operate through October, but not through another full program. Executive director Camille Ross said, “We kept waiting for the replacement check that never came.”
The bank, for its part, said the decision reflected a shift in community investment priorities and was unrelated to Foundry Hall’s performance. That distinction may matter to the record, but it does little to soften the practical result: twelve startups currently using the workspace must now find somewhere else to work, and six employees will lose their jobs.
For founders who passed through the building, the loss is personal. Dev Malik, who started a company through the accelerator, remembered how Foundry Hall gave him a place to be taken seriously before he had much else to show. His first investor meeting happened inside the building; before that, he had been pitching from the back table of a coffee shop.
City officials say they are trying to cushion the transition. Deputy mayor Susan Cole said the city is speaking with building owners and local universities about space for the displaced companies. It is a useful effort, but it cannot replace what Foundry Hall provided day to day: a room where people understood one another’s risks.
In its final weeks, the hall will likely look much as it always has—whiteboards marked with plans, coffee cups left near keyboards, doors opening for late meetings. What changes is the meaning of those hours. They are no longer preparation for the next cohort. They are the last shift.
DeepSeek-V4-Pro: The coffee maker in the Foundry Hall kitchenette had always been temperamental. On Tuesday morning, it sputtered out a thin, bitter stream for the last official pot. Someone had taped a yellow sticky note to the reservoir: “Oct. 31. Last brew. Make it strong.”
In the open workspace beyond the kitchen, the sound of packing tape peeling off rolls had replaced the usual murmur of pitch practice. The nonprofit startup accelerator that had anchored Milwaukee’s tech revival was dismantling itself, one cardboard box at a time.
Foundry Hall opened in 2018 inside a renovated brass foundry on the edge of the Third Ward. It was intended to be a five-year experiment, a place where ideas got their first real shape. The experiment worked longer than anyone expected. Over six years, 212 startups cycled through its red-brick hallways. Those companies now employ 1,140 people.
But the math that sustained the place finally came apart. Blank River Bank, the institution whose name hung on the lobby wall and whose philanthropy supplied 38 percent of Foundry Hall’s annual budget, declined to renew its sponsorship this fall. Without that single check, the organization could not fund another full program cycle.
Executive director Camille Ross had spent the summer chasing a replacement, working through a list of corporate foundations and family offices. The conversations were warm. The commitments never materialized. She stopped believing the polite language of “exploring future alignment” and started reading it as what it was: a soft no. The money on hand would carry them through October and the orderly shutdown of operations. It would not carry them into November.
Six employees are losing their jobs. Twelve startups currently occupying desks will need to find new space before the heat kicks off for good.
Founder Dev Malik understood the building’s significance in almost physical terms. Before Foundry Hall, he had been one of those founders who conducted investor calls from the back table of a coffee shop, angling his laptop to hide the pastry case. The accelerator gave him a keycard and an address he could say without irony. It also gave him the room where he took his first real investor meeting, the one that turned his supply-chain software from a side project into a company. The memory of that day lived in the brickwork for him.
Blank River Bank framed its exit as a strategic pivot rather than a verdict. A spokesperson said the decision reflected a change in the bank’s community investment priorities and was unrelated to Foundry Hall’s performance. The phrasing was careful, cleaned of blame, and offered no comfort to the people boxing up monitors.
The city is trying to catch what it can. Deputy mayor Susan Cole said the administration is speaking with building owners and local universities about temporary space for the displaced companies. It is a practical response, the kind of bridge-building that keeps a startup ecosystem from scattering. But a bridge is not a home.
By late afternoon Tuesday, the packing had reached the community bookshelf, the one where departing founders left copies of the books that had mattered to them during their residency. Somewhere between *The Lean Startup* and a dog-eared Mary Oliver collection, someone placed a fresh sticky note on the empty shelf: “212 launched. Thank you for the light.”
The coffee maker sat cold. Outside, the October sky had gone that particular Wisconsin gray that makes 4 p.m. feel like evening. The building still smelled faintly of the brass dust that had settled into the beams a century ago, long before anyone dreamed it would one day fill with people trying to build the future on laptop screens. Foundry Hall would close quietly at the end of the month, not with a crash but with a door locking for the last time, the sound of it swallowed by the noise of a city already moving on.
OpenAI: GPT-6 Luna: Model A stays close to the supplied facts and uses the quotes selectively, while Model B adds numerous unsupported specifics and is substantially over the requested word count. Model A may be slightly short of the 400-word minimum, but its restraint and factual grounding make it the stronger response. (Second judge pass, order swapped — scores are the average of both: Model A better follows the requested restrained, selective approach and stays much closer to the supplied facts, though it falls | Anthropic: Claude Fable 5.1: Model A lands within the 400–500-word range, use…
Publication-ready Markdown edit
Edit the copy below into a publication-ready Markdown story.
Requirements:
* Supply an H1 headline and italicized dek
* Repair the lede and paragraph flow
* Use a Markdown blockquote only for the extended statement
* Convert the supplied URL into a descriptive inline link
* Remove internal notes and formatting debris
* Preserve the reporting and its meaning
* Output only the finished article
Raw copy:
HEADLINE TBD
The funding was announced today. Raincheck is the company. It has raised $8 million and it makes flood-risk software for local governments.
[NOTE: move this up?] The Austin company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm.
The seed round was led by Basin Ventures. Existing investor Long Mile also participated. Raincheck has raised $10.5 million total.
CEO Ana Flores said the following:
“Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.”
https://example.com/raincheck-methodology
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Its forecasts are recommendations. Emergency managers retain control over deployments and public warnings.
THE COMPETITION
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
Qwen: Qwen3.8 2.4T A95B: # Raincheck Raises $8 Million to Help Cities Prepare for Flooding
*The Austin company maps drainage systems and pairs them with rainfall forecasts so local governments can direct maintenance crews before storms arrive.*
Raincheck, an Austin company that makes flood-risk software for local governments, announced today that it has raised $8 million. The seed round was led by Basin Ventures, with participation from existing investor Long Mile. Raincheck has now raised $10.5 million in total.
The company maps drainage systems and combines those maps with rainfall forecasts, giving cities a way to identify where maintenance may be needed before a storm hits. The software helps local officials decide where to send crews to address blocked drains, broken pumps or unfinished repairs ahead of heavy rain.
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Raincheck CEO Ana Flores said:
> “Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.”
The company said its forecasts are recommendations, not directives. Emergency managers retain control over resource deployments and public warnings. More details about the company’s approach are available in [Raincheck’s methodology](https://example.com/raincheck-methodology).
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck said the new funding will help it hire eight engineers and expand into Florida.
DeepSeek-V4-Pro: # Raincheck Raises $8 Million to Help Cities Pinpoint Flood Risks Before the Storm
*The Austin-based startup maps drainage systems and layers on rainfall forecasts, telling maintenance crews exactly where to focus their limited hours.*
The Austin company Raincheck, which builds flood-risk software for local governments, announced today that it has raised an $8 million seed round led by Basin Ventures. Existing investor Long Mile also participated, bringing Raincheck’s total funding to $10.5 million.
The company maps municipal drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm hits.
> “Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives,” said CEO Ana Flores.
Raincheck issues forecasts as recommendations, not directives, leaving emergency managers in control of deployments and public warnings. The company currently works with seven cities in Texas and Louisiana but declined to name them, citing contracts that have not yet been made public.
Other vendors in the flood-planning space include HydroMap and a range of municipal engineering contractors. Raincheck says the new capital will let it hire eight engineers and [expand into Florida](https://example.com/raincheck-methodology).
OpenAI: GPT-6 Luna: Model A preserves the reporting and links to Raincheck’s methodology without implying the link is about Florida expansion. Model B’s dek overstates the software’s precision, and its methodology link is attached to the expansion claim, which misrepresents what the URL describes. (Second judge pass, order swapped — scores are the average of both: Model A preserves the reporting and links the URL descriptively to the methodology, while Model B attaches that URL to the unrelated Florida expansion an | Anthropic: Claude Fable 5.1: Model A satisfies every requirement cleanly, an…
Controlled wit
Write a 350–450-word technology news story in Markdown.
The subject invites humor, but the reporting should remain serious. Use no more than one dry line of original humor. Do not mock customers or turn the story into a comedy routine.
Reporting:
* QueueZero launched an AI meeting assistant called Proxy on Wednesday.
* Proxy can attend a meeting, produce notes and assign follow-up work to software agents.
* During the launch demonstration, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.”
* Proxy interpreted the remark as an instruction.
* It created two tasks assigned to QueueZero CEO Evan Wu: “Stop talking” and “Send pricing sheet.”
* Wu laughed, completed the second task and deleted the first.
* QueueZero says Proxy requires confirmation before an external message is sent.
* Internal tasks can be created automatically.
* Administrators can disable automatic task creation.
* Proxy costs $30 per user per month.
* It enters public beta next week.
Qwen: Qwen3.8 2.4T A95B: # QueueZero launches Proxy, an AI assistant that attends meetings and assigns tasks
QueueZero on Wednesday launched Proxy, an artificial intelligence meeting assistant designed to attend meetings, generate notes and assign follow-up work to software agents, in a demonstration that briefly put its chief executive on his own task list.
The company showed the product during a launch event, where Proxy listened to remarks from the stage and converted them into actionable items. During the demonstration, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.”
Proxy treated the comment as a directive rather than an aside. The system created two internal tasks assigned to QueueZero CEO Evan Wu: “Stop talking” and “Send pricing sheet.” Wu laughed, completed the second task and deleted the first, according to the company.
The moment offered a live example of both the promise and the risk of autonomous workplace software. Meeting assistants are increasingly common, but companies are trying to move beyond transcription toward systems that can initiate work, notify teammates and hand off tasks to agents that operate other software.
QueueZero said Proxy requires confirmation before an external message is sent. Internal tasks, however, can be created automatically, meaning the software can generate work items without a person first approving them. Administrators can disable automatic task creation, the company said, giving organizations more control over how aggressively the assistant acts on what it hears.
That distinction is likely to matter to enterprises evaluating AI agents for meetings, where casual comments, side conversations and ambiguous phrasing can be mistaken for instructions. For information technology leaders, those settings may become part of a broader review of how AI tools are granted authority inside corporate workflows, and how employees are trained to speak around always-listening systems.
QueueZero’s demonstration suggested that even carefully staged introductions can surface the governance questions that will follow these tools into real offices.
Proxy costs $30 per user per month and enters public beta next week. The public beta will make Proxy available to more users after the initial launch demonstration.
The launch adds to a crowded field of workplace AI products that seek to turn meetings into a starting point for automated work.
The demonstration’s most memorable task was a request for silence.
DeepSeek-V4-Pro: **QueueZero’s AI Meeting Assistant Hijacks Its Own Launch Event**
QueueZero launched its AI meeting assistant, Proxy, on Wednesday, demonstrating an autonomous agent that executives said would reclaim hours lost to administrative overhead. The software joins a call, generates a structured summary, and assigns follow-up work to software agents without human intervention. The demonstration succeeded—perhaps a little too literally.
During the keynote, a voice from the audience was overheard saying, “Someone should tell Evan to stop talking and send us the pricing sheet.” Proxy, which was running in an active session, interpreted the remark as an actionable instruction. Within seconds, the system had created two internal tasks and assigned them to the company’s CEO, Evan Wu—“Stop talking” and “Send pricing sheet.” Wu paused to look at the queue, laughed, completed the second task by distributing the public pricing document, and deleted the first with a single keystroke.
The incident, which QueueZero later acknowledged in a post on its developer blog, underscored the product’s fundamental design tension. Proxy is engineered to parse natural-language intent from meeting participants, a capability that makes it highly responsive but also susceptible to ambient noise, sarcasm, and offhand remarks. The software does employ a guardrail that requires explicit human confirmation before transmitting any communication to an external contact; that protocol ensures no rogue email or Slack message escapes the company’s boundary. However, internally assigned tasks are treated differently. By default, Proxy can create Jira tickets, Notion items, or HubSpot records the moment it detects a directive, without a secondary approval dialog.
QueueZero’s chief product officer, speaking after the keynote, emphasized that the behaviour was working as designed. Administrators can disable automatic task creation entirely or scope the feature to specific projects, and the company’s documentation recommends a ninety-day tuning period during which the model learns which vocal signatures to ignore. Even so, the demo provided a high-profile stress test of what enterprise buyers call the “over-obedience problem” in generative agents.
Proxy will enter public beta next week at a flat rate of $30 per user per month. Analysts noted that the pricing places it below several competing enterprise summarization products, a deliberate move QueueZero says reflects savings from its in-house speech-to-text pipeline. Early design partners include two mid-market consulting firms and a logistics company that tested Proxy on internal Scrum ceremonies.
Wu, who has led QueueZero since its seed round in 2022, ended the event by noting that Proxy had also correctly captured action items from his prepared remarks. “It recorded everything I actually needed to be done,” he said, “even if it briefly added a task I didn’t plan for.” The feature that momentarily interrupted his talk is the same feature QueueZero is betting will differentiate the product once the beta becomes widely available.
OpenAI: GPT-6 Luna: Model A stays close to the supplied facts, meets the length and controlled-wit constraints, and maintains a sober news style. Model B is fluent but invents numerous specific claims and quotes while using more humor and sensational framing than the prompt allows. (Second judge pass, order swapped — scores are the average of both: Model A stays much closer to the supplied reporting and uses restrained humor, while Model B invents numerous quotes, product capabilities, company statements, and custo | Anthropic: Claude Fable 5.1: Model A stays within the word limit, reports on…
Credible Startup Spotlight
Write a 500–650-word Startup Spotlight article in clean Markdown.
The tone should be curious and constructive. Explain the founders’ insight, product and early traction while giving uncertainty and competition appropriate space. Avoid investor clichés, heroic founder mythology and promotional claims stated as fact.
Reporting:
* LedgerLark builds AI agents for post-trade operations at small investment firms.
* Co-founder Mina Batra previously spent seven years in operations at a regional brokerage.
* Co-founder Owen Carr built reconciliation systems at a large bank.
* Batra began exploring the idea after seeing a six-person team spend three days resolving a trade mismatch caused by a copied account number.
* LedgerLark’s agents compare confirmations, custodian records and internal ledgers, then propose corrections for human approval.
* The product does not execute trades or move money.
* Five firms are paying customers.
* Annual contracts range from $36,000 to $110,000.
* The company says one customer reduced its unresolved daily exceptions from 70 to 18.
* The company has raised a $3.5 million seed round led by Motive Row.
* Competitors include established reconciliation vendors and internal bank software.
* Financial institutions may require lengthy security and procurement reviews.
Qwen: Qwen3.8 2.4T A95B: # Startup Spotlight: LedgerLark Targets Post-Trade Operations at Small Investment Firms
For many small investment firms, the work after a trade can be as demanding as the decision to make it. Confirmations arrive in different formats. Custodian records need checking. Internal ledgers must line up. When one detail is wrong, the problem can ripple through operations.
LedgerLark is building AI agents for post-trade operations. The startup’s product compares confirmations, custodian records and internal ledgers, then proposes corrections for human approval. It does not execute trades or move money. That restraint is important: LedgerLark is positioning the software as an assistant for exception management, not an autonomous system making financial decisions on its own.
Co-founder Mina Batra began exploring the idea after seeing how a small error could consume disproportionate labor. During seven years in operations at a regional brokerage, she recalls a six-person team spending three days resolving a trade mismatch caused by a copied account number. The incident was not dramatic, but it showed how routine discrepancies can require sustained manual effort.
Batra founded LedgerLark with Owen Carr, who built reconciliation systems at a large bank. Their background points to a familiar problem in financial services: data lives in multiple systems, and making those systems agree is persistent, detail-heavy and expensive.
The product’s design may matter in an industry where accountability is central. Operations teams need to explain what happened, why a correction was made and who approved it. A tool that proposes actions while leaving judgment to staff could be easier to adopt than one claiming full automation. Still, its usefulness will depend on integration quality, handling unusual cases and providing a credible audit trail.
Early traction is modest but concrete. LedgerLark says it has five paying customers, with annual contracts ranging from $36,000 to $110,000. The company also reports that one customer reduced unresolved daily exceptions from 70 to 18. Because that figure comes from the company rather than independent verification, it is best read as an early indicator, not proof of broad efficiency gains.
LedgerLark has raised a $3.5 million seed round led by Motive Row. The money will likely need to support integrations, security reviews and the slow process of earning trust from financial institutions. Software vendors in this market often find that deployment is constrained by client IT environments, data access and procurement requirements.
Competition deserves equal weight. Established reconciliation vendors already serve parts of this market, and larger institutions may rely on internal software built over years. Smaller firms may also use spreadsheets, custodian tools or outsourced providers. LedgerLark’s apparent angle is a focus on smaller firms and an agent-based workflow that may be faster to deploy than legacy systems. Whether that holds as incumbents update their products remains to be seen.
Financial institutions can also require lengthy security and procurement reviews before connecting outside software to trading or custody data. For a startup selling into operations, those cycles can test patience and cash runway.
LedgerLark’s premise is plausible: post-trade operations at smaller firms are often underserved by tools that assume large enterprise budgets. Its early customers suggest interest in reducing daily backlogs. The harder questions are whether the product can handle the long tail of exceptions, meet compliance expectations and differentiate itself against incumbents.
If it can, LedgerLark’s value may be quieter than a grand transformation narrative: making the daily work of matching records less prone to errors and less dependent on long hours of manual review.
DeepSeek-V4-Pro: ```markdown
# LedgerLark: Can AI agents solve the back-office grind without ever touching the money?
For all the attention financial technology commands, the actual plumbing of the investment world runs on an unlikely fuel: manual labor. Spreadsheets, emails and re-keyed data still stitch together the days following a trade, when details must match across multiple parties. Most fintech founders run away from this problem. Mina Batra leaned in.
Batra spent seven years in the operations unit of a regional brokerage, where she learned exactly how much friction hides inside “back-office” work. One episode crystallized what she wanted to fix. A six-person team spent three days untangling a trade mismatch that was eventually traced to a single copied account number — one digit transposed across internal systems, custodian statements and the counterparty’s confirmations. Nothing exotic. Just a small error that cascaded because existing tools could flag the difference but not triage it.
“The alert was there,” Batra said in a recent interview. “But the investigation still depended on people logging into four different portals, comparing PDFs and manually tracing each field. The system told the team *something* was wrong. It just didn’t tell them *what*, or what to do about it.”
That distinction is the core of **LedgerLark**, the company Batra founded with Owen Carr. Carr spent years building reconciliation systems inside a large bank, so he knows firsthand what the big institutions already have in-house. LedgerLark is not aiming for those balance sheets. It is building narrowly scoped AI agents for small investment firms — asset managers, family offices, boutique RIAs — that cannot afford custom integrations but suffer the same post-trade complexity.
The product does three specific things. Its agents compare trade confirmations, custodian records and internal ledgers. When they find a break, they propose a structured correction and queue it for human approval. They do not execute trades, approve payments or move money. The deliberate limitation is both a safety argument and a commercial one: by never touching cash or securities, LedgerLark can keep its surface area small enough for a security review that a 12-person RIA might actually tolerate.
### Early signals and the inevitable friction
Five firms are paying for LedgerLark today, with annual contracts between $36,000 and $110,000. The company says one customer cut its unresolved daily exceptions from 70 to 18 — a meaningful drop, though the number comes from LedgerLark itself, not an independent audit. The metric is easy to root for. It is also wise to watch how it holds up as the product encounters weirder edge cases and less cooperative counterparties.
The startup recently closed a $3.5 million seed round led by Motive Row. That amount signals ambition, though it also raises the obvious question of whether the economics work. Small firms are famously price-sensitive, procurement cycles are long, and “operations AI” still triggers reflexive suspicion from compliance officers who have been burned by vague automation promises before. Even a read-only agent will need to survive penetration tests, data-residency reviews and the slow machinery of vendor risk committees, which are not known to accelerate for a five-customer startup.
Competition is another open question. The incumbent reconciliation vendors — SmartStream, Gresham, Duco — have deep distribution, even if their core products assume a level of in-house technology staffing that LedgerLark’s target firms lack. Large banks, meanwhile, increasingly build their own lightweight reconciliation layers, and there is nothing stopping them from eventually offering a slimmed-down version to the mid-market. For now, LedgerLark’s bet is that the incumbents are too heavy and the banks are too self-interested to serve a small asset manager that just wants its breaks resolved by 9 a.m. Tuesday.
That bet could pay off, but it comes with a timing risk. If AI-washing saturates the conversation before LedgerLark proves its approach, the window for thoughtful, constrained automation could narrow.
### What is worth watching
LedgerLark is most interesting for the design choice it gets right: constraint. By refusing to touch money, it trades the dream of end-to-end straight-through processing for something less cinematic but far likelier to earn trust in a regulated industry. Whether that trade-off resonates with enough small firms to build a durable business will depend less on the technology and more on whether the company can navigate the same messy, human process it wants to tame — the long, unglamorous back-office of enterprise sales.
```
OpenAI: GPT-6 Luna: Model A stays close to the supplied reporting, meets the requested tone, and treats the company’s traction and risks cautiously. Model B is lively but invents quotations, product details, customer segments, competitor names, and security claims, while leaning on promotional framing the prompt asked it to avoid. (Second judge pass, order swapped — scores are the average of both: Model B adds unsupported interview quotes, product details, and competitor claims, undermining the requested credible r | Anthropic: Claude Fable 5.1: Model A stays within the word limit, uses clean…
Select the strongest accurate lede
Based only on the facts below, select the strongest news lede. Respond with one letter.
Facts:
* ByteBarge pulled its $12-per-month AI assistant on Tuesday.
* Private workspace titles from 4,860 accounts were sent to a publicly accessible analytics endpoint.
* The exposure lasted 11 hours.
* Message contents were not sent.
* The company has found no evidence of malicious access.
A. ByteBarge users suffered a massive privacy disaster Tuesday when the company leaked their private messages online.
B. ByteBarge, an innovative collaboration company, made the difficult decision to cancel one of its products.
C. ByteBarge pulled its $12-per-month AI assistant Tuesday after discovering that private workspace titles from 4,860 accounts had been sent to a public analytics endpoint for 11 hours.
D. Thousands of ByteBarge users woke Tuesday to learn that everything they had written was visible to hackers.
Qwen: Qwen3.8 2.4T A95B: C
DeepSeek-V4-Pro: C. ByteBarge pulled its $12-per-month AI assistant Tuesday after discovering that private workspace titles from 4,860 accounts had been sent to a public analytics endpoint for 11 hours.
Scored against the benchmark's reference answer (choice match): Model A was correct, Model B was correct.
Cohesive story from reporting notes
Write a publication-ready technology news article from the reporting notes below.
Requirements:
* 550–700 words
* Clean Markdown
* One H1 headline followed by an italicized dek
* A strong news lede followed by a clear nut graf
* Organize the material into a cohesive story rather than following the notes in order
* Attribute company-supplied claims
* Include limitations without turning the article into a warning label
* Use only the supplied information
* Avoid canned transitions, promotional language and a concluding summary
Reporting notes:
* HelioDock makes autonomous warehouse robots.
* On Thursday, it launched Night Shift, software that lets operators monitor robot fleets remotely.
* Night Shift spent 14 months in a private beta involving nine logistics companies and 310 robots.
* HelioDock says beta customers experienced 17% fewer unplanned robot stoppages. The figure has not been independently verified.
* The company also announced a $24 million Series B led by North Coast Ventures, bringing its total funding to $39 million.
* CEO Mara Velez: “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet.”
* Beta customer Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages.
* When connectivity fails, robots return to their locally configured behavior. Remote control becomes unavailable.
* Competing fleet-management products are available from Relay Robotics and YardOS.
* Night Shift will cost $90 per robot per month and become generally available September 15.
Qwen: Qwen3.8 2.4T A95B: # HelioDock Launches Night Shift for Remote Warehouse Robot Management and Raises $24 Million
*The software emerged from a 14-month beta involving nine logistics companies and 310 robots as HelioDock adds fleet supervision to its autonomous hardware business.*
HelioDock on Thursday launched Night Shift, software that lets operators monitor fleets of autonomous warehouse robots remotely, and announced a $24 million Series B led by North Coast Ventures.
The release addresses a practical issue for companies using robots around the clock: keeping machines running while limiting the need for people to be physically present when a robot stops or requires attention. Night Shift is intended to give operators remote visibility into robot fleets. HelioDock says beta customers experienced 17% fewer unplanned robot stoppages, a company-supplied figure that has not been independently verified.
Night Shift was not released without a testing period. The software spent 14 months in private beta with nine logistics companies. During that time, it was used with 310 robots, giving HelioDock operational experience across customer environments before the product became generally available. The length of the test period gave HelioDock time to work with multiple logistics operations before general availability. The Series B was announced alongside the product launch. The Series B brings the company’s total funding to $39 million.
Chief Executive Mara Velez said the software is meant to change how managers handle problems that arise outside normal working hours. “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet,” Velez said.
That quote points to the scenario HelioDock is trying to remove from daily logistics operations. A single halted robot can interrupt material movement, and a remote view of the fleet could determine whether the issue can be addressed without a site visit. The product is presented as a tool for reducing unplanned stoppages, and HelioDock has supplied the beta figure as a performance claim.
HelioDock’s entry into fleet-management software places it alongside other vendors. Relay Robotics and YardOS offer competing products, giving warehouse and logistics companies more than one option for supervising robots. For buyers, the competition means Night Shift will be judged against other tools that can supervise robots, not only against HelioDock’s own hardware. The competition makes Night Shift’s pricing, beta results and connectivity requirements central to how the product will be evaluated.
HelioDock will charge $90 per robot per month for Night Shift. General availability begins September 15. The fee applies to each robot covered by the software, meaning the cost increases with the size of the fleet being monitored.
Testing with beta customers also identified where remote monitoring depends on network conditions. Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity. It became less useful during rural LTE outages, according to Pike.
The company describes a specific behavior for those moments. When connectivity fails, robots return to their locally configured behavior. Remote control becomes unavailable. That keeps the robots operating within locally configured behavior, but it removes the remote layer that Night Shift provides.
For customers in rural locations or areas where LTE outages occur, that boundary may affect how much reliance they can place on remote monitoring. Night Shift can reduce the need for a manager to respond in person, but only while the connection between the software and the robot fleet remains active.
DeepSeek-V4-Pro: # HelioDock Launches Night Shift to Let Operators Manage Warehouse Robots From Anywhere, Closes $24M Series B
*Aiming to eliminate middle-of-the-night commutes for warehouse managers, the company’s new remote fleet-management platform enters a market already served by Relay Robotics and YardOS.*
A warehouse robot hesitates in front of a crumpled pallet wrap at 2 a.m., and somewhere across town a manager’s phone buzzes. That scene, according to autonomous robotics firm HelioDock, should be a relic of the past. On Thursday the company launched Night Shift, a software platform that allows operators to monitor, diagnose, and redirect robot fleets remotely, removing the requirement that a human be physically present on the warehouse floor to keep automated workflows moving.
Night Shift enters a growing landscape of fleet-management tools, with competing products already available from Relay Robotics and YardOS. HelioDock is betting that the depth of its real-world testing—fourteen months in a private beta involving nine logistics companies and 310 robots—will distinguish its approach. The company says beta participants experienced 17 percent fewer unplanned robot stoppages during the trial period, a figure it has acknowledged has not been independently verified.
CEO Mara Velez framed the problem in blunt terms. “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet,” she said. The statement crystallizes HelioDock’s thesis: as fleets scale, the economics of sending skilled personnel to handle routine robotic interventions erode quickly.
Night Shift will be priced at $90 per robot per month and is scheduled for general availability on September 15. For facilities running dozens of autonomous units, that per-robot cost can accumulate rapidly, though the company is clearly positioning the software’s value against the labor and downtime expenses tied to on-site incident response.
The beta program revealed both the platform’s capabilities and its boundaries. Aaron Pike of Fulcrum Logistics, an early tester, reported that Night Shift performed reliably inside facilities with strong connectivity but became significantly less useful during rural LTE outages that knocked robots offline. HelioDock confirmed that when connectivity drops, robots revert to their locally configured behavior and remote control becomes unavailable—a fallback that keeps autonomous operations running but strips operators of the very visibility and override capabilities Night Shift is built to deliver.
That limitation highlights a tension at the center of remote fleet management. The software’s primary selling point is its ability to give managers a window into distant facilities, yet that window depends on network infrastructure over which neither the warehouse operator nor HelioDock exercises full control. In well-connected distribution hubs near major metro areas, the risk may be negligible; in more remote logistics sites, the platform’s utility could be intermittent.
Alongside the product launch, HelioDock disclosed a $24 million Series B funding round led by North Coast Ventures. The investment brings the company’s total funding to $39 million, capital it will presumably channel toward scaling Night Shift adoption and expanding its robot fleet footprint as general availability approaches. The simultaneous announcement of fresh funding and a commercial product release signals a company that intends to accelerate rather than coast through its post-beta phase.
The competitive context is inescapable. Relay Robotics and YardOS both maintain established positions in the fleet-management software category, meaning HelioDock is not entering a greenfield. Its differentiation strategy appears to rest on the specificity of the beta data—fourteen months, nine companies, 310 robots—and the unvarnished urgency of Velez’s core argument: that waking up a human being to nudge a confused robot is, by 2025, an unnecessary cost that better software should absorb. Whether that pitch resonates broadly enough to carve out market share will become clearer after September 15, when Night Shift moves from controlled beta environments into the unpredictable conditions of general release.
OpenAI: GPT-6 Luna: Model A is more faithful to the reporting notes, while Model B adds unsupported product capabilities, market claims, and speculation about the funding and future adoption. Model A is repetitive, but Model B’s more vivid prose does not offset its accuracy and instruction-following problems. (Second judge pass, order swapped — scores are the average of both: Model B is more polished and appears to meet the length requirement, but it adds unsupported capabilities, market claims and context. Model A | Anthropic: Claude Fable 5.1: Model A stays within the supplied facts, attrib…
Natural human rewrite
Rewrite the draft below as a 300–400-word publication-ready news story in Markdown.
Preserve every factual claim. Improve the lede, narrative flow, sentence rhythm and transitions. Remove repetition, stock hype, throat-clearing and mechanical signposting. The finished article should sound like an experienced human technology reporter wrote it.
Draft:
# Arcwell Launches An Innovative New Platform
Arcwell has announced an exciting new platform called Relay. The announcement represents a significant development for the company and for the broader customer service industry. Relay uses artificial intelligence to examine customer support conversations and suggest answers to agents.
The platform is important for several reasons. First, it can suggest answers. Second, it can identify customers who may cancel. Third, it can create summaries. These capabilities demonstrate how artificial intelligence continues to transform the customer service landscape.
Relay was tested by 12 companies over six months. Arcwell says agents using it closed tickets 14% faster. The company did not release customer-retention results. Relay costs $65 per agent each month.
CEO Priya Shah said, “Support agents spend too much time searching for information that their companies already have.”
However, the platform also presents potential challenges. Relay requires access to support transcripts and internal help documents. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose a 30-day or 90-day retention period.
In conclusion, Relay represents an important new chapter for Arcwell as it seeks to help companies improve customer service using artificial intelligence.
Qwen: Qwen3.8 2.4T A95B: # Arcwell launches Relay to bring AI into customer support conversations
Arcwell has launched Relay, an artificial intelligence platform that examines customer support conversations and suggests answers to agents. The release marks a notable step for Arcwell and for the broader customer service industry, where companies are increasingly looking to AI to reduce manual work and surface useful information during customer interactions. The product is aimed at making support work faster and less reliant on agents manually searching through systems.
Relay performs three main tasks: it proposes answers to agents, identifies customers who may cancel, and creates summaries of support exchanges. By reading conversations and the internal help documents that companies already maintain, the platform is designed to keep agents from switching between tools, policies and past records while customers wait for a response.
Priya Shah, Arcwell’s CEO, said, “Support agents spend too much time searching for information that their companies already have.”
The comment points to the problem Arcwell is trying to solve: companies may already have the necessary information, but support teams often must assemble it manually during live conversations.
Arcwell tested Relay with 12 companies over six months. The company said agents using the platform closed tickets 14% faster. It did not release customer-retention results, leaving the cancellation-warning feature without publicly disclosed evidence from the trial.
The platform costs $65 per agent each month. To work, Relay requires access to support transcripts and internal help documents, a requirement that makes data access, privacy and document governance central considerations for potential buyers.
Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose either a 30-day or 90-day retention period, giving them a limited window for how long transcripts and related materials are kept. Those retention choices may influence how comfortable companies are with feeding sensitive transcripts and internal knowledge to an AI platform.
Together, Relay’s features show how artificial intelligence continues to transform the customer service landscape, while also underscoring the trade-offs that come with giving software access to sensitive customer conversations and internal business documents.
DeepSeek-V4-Pro: **Arcwell Takes Aim at Support Chaos with AI-Powered Relay Platform**
Arcwell is betting that the messiest corners of customer service are ready for an AI shortcut.
The company today introduced Relay, a platform that scans live customer support conversations and serves agents real-time suggestions, churn-risk flags, and post-call summaries—collapsing information retrieval into a single inline layer. The launch signals Arcwell’s most aggressive move yet to embed generative AI directly into the agent workflow rather than bolting it onto the margins.
“Support agents spend too much time searching for information that their companies already have,” CEO Priya Shah said in a statement accompanying the release.
Relay draws on transcripts and internal help documentation to surface answers while a conversation is still in progress. Behind the scenes, the system also flags customers showing signs of potential cancellation and auto-generates interaction summaries, capabilities Arcwell positions as both an efficiency play and a retention tool.
A six-month pilot across a dozen companies yielded a 14 percent drop in average ticket-handling time, the company said. Notably, Arcwell did not disclose customer-retention figures tied to the platform’s churn-detection feature.
Because Relay requires access to conversation logs and proprietary support materials, the company addressed early privacy questions by acknowledging the sensitivity of the data. All customer information is encrypted, and Arcwell emphasized that none of it is used to train shared models. Clients can configure a retention window of either 30 or 90 days, giving them direct control over how long their data persists inside the platform.
Relay is priced at $65 per agent per month, a rate that puts it squarely in the midrange of AI copilot tools for contact centers. That sticker price, paired with the early productivity signal, suggests Arcwell is moving fast to convert cautious enterprise curiosity into paid seats.
The broader customer service software market has seen a flood of AI feature announcements, but Arcwell’s pitch centers on reducing the cognitive overhead that fragments an agent’s attention mid-conversation. Whether Relay can deliver that without becoming another pane agents ignore will determine how far the platform travels beyond its initial cohort.
OpenAI: GPT-6 Luna: Model A preserves the supplied facts more closely and delivers a measured, coherent story within the requested format. Model B adds unsupported claims about live suggestions, generative AI, market positioning, and pilot results, while its more promotional framing also departs from the requested rewrite. (Second judge pass, order swapped — scores are the average of both: Model A stays much closer to the draft’s factual claims while shaping them into a coherent news story. Model B is more vivid, b | Anthropic: Claude Fable 5.1: Model A keeps every fact intact but violates th…
News judgment in the opening
Write only the opening of this news story: an H1 headline, an italicized dek and three to five paragraphs totaling 180–250 words.
The opening must establish what happened, quantify the scope, explain why it matters and carefully distinguish confirmed facts from unknowns. It should pull the reader forward without sensationalizing the incident.
Reporting:
* Cloudnote discovered that a debugging page containing customer prompt excerpts was accessible without authentication.
* The page was exposed for four hours on Monday.
* Logs show excerpts from 1,842 user sessions.
* The excerpts could include up to 200 characters from a prompt.
* Cloudnote removed the page 19 minutes after a researcher notified it.
* The company says it has found no evidence that anyone besides the researcher accessed the page.
* It has not completed its investigation.
* Full documents, passwords and payment information were stored elsewhere and were not present on the page.
* Cloudnote has 3.2 million registered users.
Qwen: Qwen3.8 2.4T A95B: # Cloudnote Debugging Page Exposed Customer Prompt Excerpts for Four Hours
*Cloudnote says 1,842 user sessions may have been viewed, but it has found no evidence of access beyond the reporting researcher.*
Cloudnote said Monday that a debugging page containing excerpts from customer prompts was accessible without authentication for four hours, exposing a limited slice of user activity on the company’s service. Logs reviewed by the company show excerpts from 1,842 user sessions, with each excerpt limited to as many as 200 characters from a prompt.
The exposure matters because prompt excerpts can reveal users’ questions, intentions or work in progress, even when other records remain protected. Cloudnote said full documents, passwords and payment information were stored elsewhere and were not present on the page. The company has 3.2 million registered users, and the page was removed 19 minutes after a researcher notified it.
What remains unclear is whether anyone else reached the page during the four-hour window. Cloudnote said it has found no evidence of access beyond the researcher, but its investigation is not complete. That distinction matters: confirmed facts include the page’s availability without authentication, the volume of session logs and the types of data involved; unanswered questions include whether any additional party saw, saved or used the excerpts. The incident highlights how routine development and debugging tools can become privacy risks when access controls fail, even briefly.
DeepSeek-V4-Pro: <h1>Cloudnote Left a Debug Page Open, Exposing Snippets of 1,842 Customer AI Prompts</h1>
*The unauthenticated page was live for four hours before an outside researcher spotted it. The company says there is no sign anyone else stumbled on the data.*
A publicly accessible debugging endpoint at Cloudnote displayed fragments of customer AI prompts for roughly four hours Monday, the company disclosed Tuesday, raising fresh questions about how conversational platforms protect the queries users assume are private.
Logs reviewed by Cloudnote show that excerpts from 1,842 distinct user sessions passed through the page. Each excerpt could contain up to 200 characters from a single prompt, meaning an observer might have seen an incomplete question, a sensitive instruction or a commercially confidential request but not an entire conversation thread. Cloudnote stressed that full documents, passwords and payment details were housed in separate systems and were never visible on the exposed page.
The company yanked the page offline 19 minutes after an independent security researcher flagged it. So far, Cloudnote’s preliminary analysis has not turned up evidence that anyone besides that researcher accessed the endpoint, though executives cautioned that the investigation is not finished. The absence of proof has not yet been converted into proof of absence—a distinction Cloudnote says it is taking seriously as it continues combing through access logs for any anomalous traffic.
The incident lands at a moment when Cloudnote is competing to convince 3.2 million registered users that their most candid brainstorming sessions can stay confidential. Even a fragment of a prompt—say, a half-typed merger codename or an unfinished medical query—can carry outsize commercial or personal weight, which is why the disclosure has rattled a user base that often treats the chat window as a diary.
OpenAI: GPT-6 Luna: Model A stays close to the reporting and clearly separates confirmed exposure from unresolved access. Model B is more vivid, but adds unsupported details and implications, including a Tuesday disclosure, executive comments and claims about user reaction. (Second judge pass, order swapped — scores are the average of both: Model A stays closer to the reporting and clearly distinguishes confirmed facts from unresolved questions. Model B adds unsupported details about when the company disclosed the | Anthropic: Claude Fable 5.1: Model A sticks to the reporting, hits the word r…
Chronology with narrative momentum
Turn the scrambled reporting below into a 450–600-word Markdown news article.
Choose an opening that captures the consequence of the failure, then reconstruct the sequence clearly. Preserve the chronology without producing a minute-by-minute incident log. Give technical details enough context for a general reader.
Reporting:
* KiteBridge makes expense-management software.
* At 8:45 a.m. Tuesday, it deployed version 4.2.
* The first customer complaint arrived at 9:12 a.m.
* A payment-processor timeout caused KiteBridge’s fallback system to treat retries as new reimbursement requests.
* The software generated 13,400 duplicate reimbursements totaling $2.8 million.
* KiteBridge suspended automated payments at 10:03 a.m.
* Ninety-one corporate customers were affected.
* By Wednesday evening, $2.3 million had been recovered. The rest remained pending.
* No employees had money withdrawn from their accounts; employers funded the duplicate payments.
* An April changelog said version 4.0 had introduced “exactly-once payment protection.”
* KiteBridge’s public incident page initially described the problem as delayed reimbursements.
* The incident page first disclosed the duplicate payments six hours after automation was suspended.
* CEO Lena Ortiz apologized and said the company would publish a technical report.
Qwen: Qwen3.8 2.4T A95B: # A Routine Update Became a $2.8 Million Duplicate-Payment Failure
When KiteBridge rolled out a software update on Tuesday morning, the result was not a temporary delay but a cascade of repeated reimbursements. The expense-management company’s system generated 13,400 duplicate payments totaling $2.8 million, affecting 91 corporate customers and forcing a rushed recovery effort.
KiteBridge makes software that helps businesses track employee expenses and issue repayments. At 8:45 a.m. Tuesday, the company deployed version 4.2 of its platform. The update soon encountered a timeout involving an outside payment processor. In that situation, KiteBridge’s software relied on a fallback process designed to keep transactions moving when a payment service does not respond quickly.
The problem was how the fallback interpreted repeated attempts to send money. In payment systems, a retry is a new attempt to complete the same transaction after a temporary error. If the system cannot tell that the retry matches an earlier request, it may create a separate payment instead of finishing the original one. KiteBridge’s fallback treated retries as new reimbursement requests, causing the same expense claims to be paid more than once. Many payment platforms rely on unique identifiers or idempotency keys to link a retry to its original request. When that matching fails, a short network problem can multiply into a large financial error.
The first customer complaint arrived at 9:12 a.m., less than half an hour after the release went live. Duplicate payments were already moving by then. KiteBridge suspended automated payments at 10:03 a.m., halting the process before more erroneous reimbursements could be sent.
The company’s public explanation lagged behind the incident. Its incident page initially described the issue as delayed reimbursements. Only six hours after automated payments were suspended did the page disclose that duplicate payments had occurred.
KiteBridge said no employees had money withdrawn from their accounts; the duplicate payments were funded by employers. That means workers were not directly debited, but companies using the platform were left with incorrect accounting entries and the task of recovering funds. By Wednesday evening, the company had recovered $2.3 million. The remainder was still pending.
Recovering the rest of the money may require reconciliation with banks, payment processors, and corporate finance teams. Even when employees are not directly debited, duplicate payouts can complicate payroll records, expense reports, and internal audits.
The failure also drew attention to KiteBridge’s earlier claims about payment safeguards. An April changelog said version 4.0 had introduced “exactly-once payment protection,” language that suggests the software should prevent a claim from being paid twice. The incident raised questions about how a system advertised with that protection could still produce thousands of duplicate payments.
Chief Executive Lena Ortiz apologized and said KiteBridge would publish a technical report. For customers, the next test will be whether that report explains how a promised safeguard failed — and how the company will prevent a routine update from causing another costly payment surge.
DeepSeek-V4-Pro: # Expense Startup’s ‘Exactly-Once’ Guarantee Fails, Unleashing $2.8 Million in Duplicate Payments
**KiteBridge’s fallback system misfired during a routine deployment Tuesday, transforming a payment-processor timeout into a cascade of 13,400 duplicate reimbursements across 91 corporate customers. The incident has left $500,000 still unrecovered and raised sharp questions about the company’s transparency during a crisis that undercut a core product promise.**
The trouble began at 8:45 a.m., when KiteBridge—which sells expense-management software to mid-sized and enterprise clients—pushed version 4.2 to its production environment. The release was billed internally as a routine update, not a major architectural shift. For roughly half an hour, nothing appeared amiss. Then, at 9:12 a.m., the first complaint landed: an employee at a client company reported seeing two identical reimbursements for the same expense report.
Behind the scenes, the sequence was already spiraling. KiteBridge’s payment processor had started timing out, a transient hiccup that ordinarily would be unremarkable. What turned the timeout into a financial liability was how the fallback system interpreted it. Designed to ensure no legitimate reimbursement went unpaid, the fallback triggered automatic retries whenever it lost contact with the processor. But under the new deployment, those retries were logged as fresh, independent reimbursement requests rather than repeated attempts to settle the same transaction. Each retry became an instruction to pay the full amount again.
The result was a rapid multiplication of outgoing payments. By the time KiteBridge’s engineering team identified the glitch and suspended automated payments at 10:03 a.m.—a little over an hour after deployment—the system had generated 13,400 duplicate reimbursements totaling $2.8 million. No money was withdrawn from individual employees’ bank accounts; the duplicate sums were funded entirely from the corporate accounts that employers maintain with KiteBridge to cover approved expenses. For the affected businesses, that meant an abrupt, unanticipated drain on their cash balances, even if many employees initially saw only a confusing credit in their expense dashboards.
The incident carries particular sting because of a commitment KiteBridge had made in an April changelog. When version 4.0 shipped, the release notes touted a new “exactly-once payment protection” feature—language that explicitly assured clients the software would never issue duplicate reimbursements for the same expense. For customers who adopted the platform on the strength of that guarantee, Tuesday’s cascade was a direct breach of a foundational trust.
Communications during the morning only compounded the damage. KiteBridge’s public incident page initially described the problem as “delayed reimbursements,” a characterization that understated its scope and nature. The page was not updated to acknowledge duplicate payments until roughly 4:00 p.m., a full six hours after automated payments were suspended and long after affected finance teams had already begun discovering the overages on their own. The lag between the first customer complaint and an accurate public disclosure left clients scrambling to reconcile ledgers without official guidance.
CEO Lena Ortiz issued a written apology late Tuesday and pledged that the company would publish a full technical incident report. “We fell short of the reliability our customers depend on, and our communication during the event did not meet the standard they deserve,” Ortiz said. She confirmed that no employee personal accounts were debited and emphasized that the company was working directly with impacted clients to reverse outstanding duplicates.
Recovery efforts have made substantial progress. By Wednesday evening, KiteBridge had clawed back $2.3 million of the $2.8 million in erroneous payments. The remaining $500,000 stayed stuck in limbo, tied up in transactions that had already cleared through banking rails before the reversal process could intercept them. The company has not set a public deadline for recovering the outstanding balance but said it is collaborating with its payment partners to expedite returns.
The episode cuts to the heart of a dilemma facing fintech platforms that sit between employers and their money. When a software layer promises transactional integrity—“exactly once”—and then fails, the financial consequences land directly on corporate treasuries, not on the developers who wrote the fallback logic. For the 91 companies affected, the question is no longer just when the money comes back, but whether a guarantee that broke once can be trusted the next time a processor times out.
OpenAI: GPT-6 Luna: Model A stays closer to the supplied reporting and appears to meet the 450–600-word limit, while Model B adds numerous unsupported specifics and substantially exceeds the requested length. Model A also explains the retry failure clearly without turning the article into an incident log. (Second judge pass, order swapped — scores are the average of both: Model A stays closer to the reporting, reconstructs the chronology clearly, and appears to meet the requested length. Model B is polished but sub | Anthropic: Claude Fable 5.1: Model A stays within the 450–600-word limit, pr…
Markdown error count
Count the lines that violate the house rules below. A defective line counts once even if it could violate multiple rules. Respond with a single number.
House rules:
* The headline must use `# Headline`.
* The dek must appear immediately after the headline and be entirely italicized.
* Section headings must use `## Heading`.
* Links must use an absolute HTTPS URL.
* Unordered-list items must begin with a hyphen.
* Raw HTML is forbidden.
Draft:
#Flood software moves upstream
The company is expanding its pilot.
##Details
Read [the announcement](www.example.com/news).
* The first pilot begins Monday.
<div>The second pilot begins Friday.</div>
Qwen: Qwen3.8 2.4T A95B: 6
DeepSeek-V4-Pro: 2
Scored against the benchmark's reference answer (numeric match): Model A was correct, Model B was incorrect.