A 12-task test of reporting, rewriting, structure and editorial judgment puts two strong text models through practical newsroom work. The results turn on consistency: polished, disciplined copy in some tasks, and careful handling of uncertainty in others.
Claude Opus 5.5 leads the aggregate 104.4 to 98.4, with six task wins to Grok 4.7’s two and four ties. The limited confidence rating supports a clear verdict—not a sweep, but a meaningful edge across this set of assignments.
Claude’s advantage shows up most in the core writing work. It won on publication-ready editing, chronology, turning reporting notes into a cohesive story, controlled wit, startup coverage and a natural rewrite. Judges repeatedly favored its flow and structure, and its restraint: Grok’s drafts were sometimes repetitive, overlong or built on implications the reporting did not establish.
Grok had two notable wins. It shaped the numbers into a tighter story, and it handled the opening on a news judgment task more carefully, distinguishing sessions from users and avoiding speculation about what exposed excerpts contained. Those are substantive strengths, especially where precision about what is—and is not—known matters most.
The models tied on selecting the strongest lede, counting Markdown errors, following an exact final-draft structure and weaving quotes into a narrative. **The call: Claude Opus 5.5 wins on editorial consistency and cleaner execution; Grok 4.7 remains the sharper choice in a couple of specific judgment calls, but not enough to close the gap.**
Publication-ready Markdown edit
Edit the copy below into a publication-ready Markdown story.
Requirements:
* Supply an H1 headline and italicized dek
* Repair the lede and paragraph flow
* Use a Markdown blockquote only for the extended statement
* Convert the supplied URL into a descriptive inline link
* Remove internal notes and formatting debris
* Preserve the reporting and its meaning
* Output only the finished article
Raw copy:
HEADLINE TBD
The funding was announced today. Raincheck is the company. It has raised $8 million and it makes flood-risk software for local governments.
[NOTE: move this up?] The Austin company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm.
The seed round was led by Basin Ventures. Existing investor Long Mile also participated. Raincheck has raised $10.5 million total.
CEO Ana Flores said the following:
“Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.”
https://example.com/raincheck-methodology
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Its forecasts are recommendations. Emergency managers retain control over deployments and public warnings.
THE COMPETITION
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
Anthropic: Claude Opus 5.5: # Raincheck Raises $8 Million to Help Cities Prepare for Flooding
*The Austin startup's software combines drainage maps with rainfall forecasts to help local governments decide where to send maintenance crews before a storm.*
Raincheck, an Austin company that makes flood-risk software for local governments, announced today that it has raised $8 million in seed funding.
Basin Ventures led the round, with participation from existing investor Long Mile. The new funding brings Raincheck's total raised to $10.5 million.
The company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews ahead of a storm. Raincheck details its approach in an [overview of its forecasting methodology](https://example.com/raincheck-methodology).
CEO Ana Flores said the challenge for cities lies less in knowing where flooding happens than in deciding what to fix first:
> "Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives."
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
The software's forecasts serve as recommendations. Emergency managers retain control over crew deployments and public warnings.
Raincheck competes with other flood-planning vendors, including HydroMap, as well as municipal engineering contractors. The company says it will use the new funding to hire eight engineers and expand into Florida.
SpaceXAI: Grok 4.7: # Raincheck Raises $8 Million for Flood-Risk Software
*The Austin company builds flood-risk tools that help local governments decide where to send maintenance crews before a storm.*
Raincheck, an Austin company that makes flood-risk software for local governments, said today it has raised $8 million.
The company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm. Raincheck describes the approach in its [methodology](https://example.com/raincheck-methodology).
The seed round was led by Basin Ventures. Existing investor Long Mile also participated. Raincheck has raised $10.5 million total.
CEO Ana Flores said:
> Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Its forecasts are recommendations. Emergency managers retain control over deployments and public warnings.
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
OpenAI: GPT-6 Luna: Model A has stronger paragraph flow and a more descriptive inline link. Model B preserves the wording about emergency managers’ control more precisely, while Model A’s “crew deployments” slightly narrows the original meaning. (Second judge pass, order swapped — scores are the average of both: Model A has a stronger headline and smoother paragraph flow while preserving the reporting and meeting the formatting requirements. Model B is also accurate and clean, but its presentation is slightly less | Anthropic: Claude Fable 5.1: Both outputs meet the structural requirements an…
Chronology with narrative momentum
Turn the scrambled reporting below into a 450–600-word Markdown news article.
Choose an opening that captures the consequence of the failure, then reconstruct the sequence clearly. Preserve the chronology without producing a minute-by-minute incident log. Give technical details enough context for a general reader.
Reporting:
* KiteBridge makes expense-management software.
* At 8:45 a.m. Tuesday, it deployed version 4.2.
* The first customer complaint arrived at 9:12 a.m.
* A payment-processor timeout caused KiteBridge’s fallback system to treat retries as new reimbursement requests.
* The software generated 13,400 duplicate reimbursements totaling $2.8 million.
* KiteBridge suspended automated payments at 10:03 a.m.
* Ninety-one corporate customers were affected.
* By Wednesday evening, $2.3 million had been recovered. The rest remained pending.
* No employees had money withdrawn from their accounts; employers funded the duplicate payments.
* An April changelog said version 4.0 had introduced “exactly-once payment protection.”
* KiteBridge’s public incident page initially described the problem as delayed reimbursements.
* The incident page first disclosed the duplicate payments six hours after automation was suspended.
* CEO Lena Ortiz apologized and said the company would publish a technical report.
Anthropic: Claude Opus 5.5: # Software Glitch Sends $2.8 Million in Duplicate Reimbursements to Employees
Ninety-one companies spent part of this week trying to recover money they never meant to pay after a software update at KiteBridge sent 13,400 duplicate expense reimbursements to their employees, totaling $2.8 million.
By Wednesday evening, $2.3 million had been recovered. The remaining roughly $500,000 was still pending. The duplicate payments were funded by employers. No employees had money withdrawn from their accounts.
KiteBridge, which makes software that companies use to process and reimburse employee expenses, has apologized. Chief executive Lena Ortiz said the company would publish a technical report on what went wrong.
## A safety net that backfired
The trouble began Tuesday morning, when KiteBridge rolled out version 4.2 of its software at 8:45 a.m. Less than half an hour later, at 9:12 a.m., the company received its first customer complaint.
The cause involved a routine problem: a timeout at a payment processor. A payment processor is the outside service that actually moves money between bank accounts. When a processor takes too long to respond, software commonly "retries," sending the same request again in case the first one was lost.
KiteBridge's system included a fallback mechanism intended to handle these situations. Instead of recognizing the retries as repeat attempts at the same reimbursement, the fallback treated each one as a brand-new request. That meant a single reimbursement that ran into a timeout could be paid more than once.
KiteBridge suspended automated payments at 10:03 a.m., about an hour and 20 minutes after the update went out and roughly 50 minutes after the first complaint. By then, the duplicates had already been issued.
## A promise from April
The failure stands in contrast to an earlier assurance. A KiteBridge changelog from April, a public record of software updates, said version 4.0 had introduced "exactly-once payment protection."
In payments software, "exactly once" means that each transaction is completed one time and only one time, even if messages are lost, delayed, or resent. It is designed to prevent precisely the kind of duplication that occurred Tuesday. The company has not yet explained publicly why that protection did not prevent the duplicates in version 4.2. That question is likely to be central to the promised technical report.
## Slow disclosure
KiteBridge's public communication about the incident also lagged behind events. The company's incident page, where it posts updates on outages and service problems, initially described the issue as delayed reimbursements.
That description reflected the effect of the suspension, which halted payments, but not the more serious problem that had prompted it. The incident page did not disclose the duplicate payments until six hours after automation was suspended, around 4 p.m. Tuesday.
For customers, the distinction mattered. A delay would simply have meant waiting. Duplicate payments meant that employers had sent out money they needed to identify and recover, and employees might have received unexpected extra deposits.
## What comes next
Recovery efforts had reclaimed more than 80 percent of the duplicated funds by Wednesday evening, but about $500,000 remained outstanding. KiteBridge has not said when it expects to resolve the remaining payments or when automated reimbursements will fully resume.
Ortiz's promised technical report may answer the questions customers are most likely to ask: how a timeout bypassed a safeguard the company had advertised months earlier, and why its public account of the incident initially described only a delay.
SpaceXAI: Grok 4.7: # KiteBridge Glitch Sends $2.8 Million in Duplicate Reimbursements
A software failure at expense-management company KiteBridge generated 13,400 duplicate reimbursement payments totaling $2.8 million on Tuesday, hitting 91 corporate customers before the company shut off automated payments. Employers, not employees, funded the extra transfers. By Wednesday evening, KiteBridge had recovered $2.3 million; the remainder was still pending.
The trouble began with the morning release of version 4.2. KiteBridge deployed the update at 8:45 a.m. Tuesday. Less than half an hour later, at 9:12 a.m., the first customer complaint arrived. What looked at first like a delay in ordinary expense payouts was in fact a breakdown in how the system decided whether a payment had already been sent.
When a payment processor timed out, KiteBridge’s fallback system treated retries as brand-new reimbursement requests. In normal operation, a timeout does not mean a payment failed. The processor may have accepted the transfer and simply failed to send confirmation in time. A careful system waits, checks, and sends the money only once. KiteBridge’s fallback did the opposite: each retry was recorded as a fresh claim and paid again.
The result was a rapid burst of duplicates. The software issued 13,400 extra reimbursements worth $2.8 million before engineers suspended automated payments at 10:03 a.m. Ninety-one corporate customers were affected. Company officials said no employees had money withdrawn from their accounts. The duplicate sums came from employer-funded reimbursement pools, so businesses—not workers—bore the immediate cost.
The failure was especially striking because of a promise KiteBridge had already made. An April changelog said version 4.0 had introduced “exactly-once payment protection,” a safeguard meant to ensure that a reimbursement is paid a single time even if a request is submitted more than once or a network call has to be repeated. Version 4.2, released Tuesday morning, was supposed to build on that foundation. Instead, a timeout in the payment processor exposed a gap: the protection did not hold when the fallback path ran.
Public communication lagged the technical response. KiteBridge’s incident page initially described the problem as delayed reimbursements, a framing that suggested money was late rather than sent twice. The page first disclosed the duplicate payments six hours after automation was suspended—well after the company had already stopped the outflow and customers had begun to see unexpected charges against employer accounts.
Recovery moved faster than the first public account of the incident. By Wednesday evening, $2.3 million of the $2.8 million had been recovered. The rest remained pending, leaving some customers still sorting which transfers were legitimate and which were copies. Because employers funded the payments, the cleanup is largely a matter of reversing or reclaiming corporate funds rather than restoring money taken from workers’ personal accounts. That distinction limits the harm to individuals, but it does not erase the operational mess for finance teams that must reconcile expenses, explain anomalies to auditors, and confirm that no employee was paid twice for the same claim.
CEO Lena Ortiz apologized and said the company would publish a technical report. That report will need to explain how a system advertised as exactly-once protection treated retries as new requests, why the fallback fired after a processor timeout, and why the public incident page described delays for hours after engineers already knew payments had been duplicated.
For customers, the practical questions are narrower and more urgent: which of the 13,400 payments have been reversed, how soon the unrecovered balance will be returned, and whether version 4.2—or the protection introduced in 4.0—can be trusted with the next payroll of expenses. Until that accounting is finished, the Tuesday morning deploy stands as a costly reminder that a retry is not the same thing as a new bill.
OpenAI: GPT-6 Luna: Model A gives a clear consequence-first chronology and appears to stay within the requested length; Model B is technically vivid but likely exceeds the 600-word limit and adds several details or implications not established in the reporting. Model B has strong explanatory writing, but Model A is more disciplined and faithful to the brief. (Second judge pass, order swapped — scores are the average of both: Model A stays within the requested length and gives a clear chronology with useful technica | Anthropic: Claude Fable 5.1: Model A opens on the consequence, stays within…
Cohesive story from reporting notes
Write a publication-ready technology news article from the reporting notes below.
Requirements:
* 550–700 words
* Clean Markdown
* One H1 headline followed by an italicized dek
* A strong news lede followed by a clear nut graf
* Organize the material into a cohesive story rather than following the notes in order
* Attribute company-supplied claims
* Include limitations without turning the article into a warning label
* Use only the supplied information
* Avoid canned transitions, promotional language and a concluding summary
Reporting notes:
* HelioDock makes autonomous warehouse robots.
* On Thursday, it launched Night Shift, software that lets operators monitor robot fleets remotely.
* Night Shift spent 14 months in a private beta involving nine logistics companies and 310 robots.
* HelioDock says beta customers experienced 17% fewer unplanned robot stoppages. The figure has not been independently verified.
* The company also announced a $24 million Series B led by North Coast Ventures, bringing its total funding to $39 million.
* CEO Mara Velez: “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet.”
* Beta customer Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages.
* When connectivity fails, robots return to their locally configured behavior. Remote control becomes unavailable.
* Competing fleet-management products are available from Relay Robotics and YardOS.
* Night Shift will cost $90 per robot per month and become generally available September 15.
Anthropic: Claude Opus 5.5: # HelioDock Launches Remote Monitoring Software for Warehouse Robot Fleets
*The company's Night Shift platform, announced alongside a $24 million funding round, lets operators manage robots from afar, as long as the connection holds.*
HelioDock, a maker of autonomous warehouse robots, on Thursday launched Night Shift, software that lets operators monitor their robot fleets remotely. The company also announced a $24 million Series B round led by North Coast Ventures.
The product addresses a familiar problem for warehouses that run robots around the clock: when a machine stalls overnight, someone often has to go on site to deal with it. HelioDock says Night Shift reduced unplanned stoppages among the customers who tested it, and the new capital arrives as the company enters a market where Relay Robotics and YardOS already sell fleet-management tools. Early customer feedback, however, suggests the software's usefulness depends on something HelioDock doesn't control: the network connection at each facility.
Chief Executive Mara Velez described the product in terms of the overnight call-out. "A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet," she said.
## Results from a 14-month beta
Before Thursday's launch, Night Shift spent 14 months in a private beta involving nine logistics companies and 310 robots. According to HelioDock, those beta customers experienced 17% fewer unplanned robot stoppages. The figure has not been independently verified.
Aaron Pike of Fulcrum Logistics, one of the beta participants, said the system worked well in facilities with strong connectivity. He said it became less useful during LTE outages at rural sites.
## When the signal drops
Pike's experience reflects how Night Shift is designed. When connectivity fails, the robots return to their locally configured behavior, and remote control becomes unavailable. Rather than taking direction from an operator elsewhere, each robot reverts to whatever has been set up on site until the link is restored.
For warehouses with dependable connections, that fallback may come into play rarely. For operators relying on LTE in rural areas, outages can remove the remote access the product is built to provide, at the moments when no one may be on the floor to step in.
That tradeoff is central to the pitch Velez is making. The scenario she describes, a single robot stuck on a pallet in the middle of the night, is the kind of problem Night Shift is meant to let a manager handle from home. Whether a given warehouse gets that benefit depends in part on the reliability of its network.
## Pricing and availability
Night Shift will become generally available on Sept. 15 and will cost $90 per robot per month. At that rate, a 100-robot fleet would cost $9,000 a month, or $108,000 a year.
The pricing puts the software in front of buyers who can also consider competing fleet-management products from Relay Robotics and YardOS. HelioDock is leaning on its beta results, chiefly the reduction in stoppages it reports, to make its case to customers weighing those options.
## New funding
The Series B brings HelioDock's total funding to $39 million, meaning the company had raised $15 million before this round. North Coast Ventures led the new investment.
The raise and the product launch arrive together, pairing a commercial release with fresh capital as HelioDock moves Night Shift from a nine-company beta to general availability. The company's robots already operate autonomously on warehouse floors; Night Shift adds a layer that lets people keep watch over those fleets, and intervene, without being in the building.
For beta customers like Fulcrum Logistics, that layer has already been in use for more than a year. Pike's account points to strong performance where connectivity is reliable and a narrower role where it is not. Remote control drops away during outages, and the robots fall back on how they were configured locally.
SpaceXAI: Grok 4.7: # HelioDock Launches Remote Fleet Software and Raises $24 Million
*Night Shift, at $90 a robot each month, leaves a 14-month private beta whose reported drop in stoppages has not been independently verified.*
HelioDock on Thursday launched Night Shift, software that lets operators monitor robot fleets remotely, and announced a $24 million Series B led by North Coast Ventures. The company makes autonomous warehouse robots. The funding and the product arrived together.
Night Shift is leaving a closed trial for a dated, priced release. It spent 14 months in a private beta involving nine logistics companies and 310 robots, will cost $90 per robot per month, and is scheduled to become generally available on September 15. HelioDock says beta customers experienced 17 percent fewer unplanned robot stoppages. That figure was supplied by the company and has not been independently verified. The round brings HelioDock’s total funding to $39 million. Fleet-management products are already available from Relay Robotics and YardOS, so the September release enters a category in which buyers have other software to set beside it.
Mara Velez, HelioDock’s chief executive, framed the launch around a disruption that happens when no one wants to be called back in. “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet,” she said. The remark is the company’s case for remote oversight: one machine, stopped by a pallet, and a person who would otherwise have to cross town in the middle of the night to deal with it. Night Shift is the software HelioDock is putting in front of that job. Operators are meant to monitor the fleet from somewhere other than the aisle where the robot has stalled.
The 17 percent claim is the number attached to that case, and it belongs to HelioDock. The company says the decline in unplanned stoppages occurred among the beta customers. No independent party has verified it. The trial behind it was private, limited to nine logistics companies and 310 robots, and ran for 14 months. Those details are the setting in which the company is offering the result. A reader can see the size of the test and the source of the percentage; the launch account does not add a separate measurement of the same robots over the same months.
One beta customer has already described a condition the percentage does not capture. Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages. His comment does not confirm the stoppage figure and does not dispute it. It locates the software’s usefulness in the quality of the link. Where connectivity was strong, Fulcrum found Night Shift worked well. Where rural LTE dropped, the remote system lost usefulness. For a product meant to spare a manager the drive Velez described, that is not a peripheral complaint. Remote monitoring is available only while the connection that carries it is available.
HelioDock’s account of a failed connection is specific. When connectivity fails, robots return to their locally configured behavior. Remote control becomes unavailable. The robot does not remain under the distant operator. It resumes the behavior configured locally, and the control that had depended on the link is gone. The launch describes Night Shift as software for monitoring fleets remotely. The outage description makes a further point: remote control is also part of what the connection supports, and it is part of what an outage withdraws. A pallet that confuses a robot at 2 a.m. can still be handled from afar only if that link is up. If it is not, the machine is back on its local configuration, and the manager’s remote session is not an option.
That split is already how a beta user talked about the product, and it is how the company describes the failure mode customers will still be buying on September 15. At $90 per robot per month, Night Shift is a remote view of a HelioDock fleet, and a remote hand on it, for as long as connectivity holds. Pike’s experience at Fulcrum Logistics says the hand is a useful one in facilities with strong connections and a weaker one when rural LTE fails. The company’s own rule for those failures is the local configuration, with remote control off. North Coast Ventures’ lead on the $24 million round, and the $39 million now raised in total, sit beside that rule rather than replacing it. The software’s reach remains the reach of the network under it.
OpenAI: GPT-6 Luna: Model A delivers a clearer news structure and more natural, publication-ready prose while covering the launch, funding, beta results, competition, pricing, and connectivity limitation. Model B includes the required information but repeatedly restates the same points, making the story feel less cohesive and more mechanical. (Second judge pass, order swapped — scores are the average of both: Model A has a clearer news structure and more polished, cohesive prose; Model B repeats the same connectivi | Anthropic: Claude Fable 5.1: Model A delivers a conventionally structured ne…
Controlled wit
Write a 350–450-word technology news story in Markdown.
The subject invites humor, but the reporting should remain serious. Use no more than one dry line of original humor. Do not mock customers or turn the story into a comedy routine.
Reporting:
* QueueZero launched an AI meeting assistant called Proxy on Wednesday.
* Proxy can attend a meeting, produce notes and assign follow-up work to software agents.
* During the launch demonstration, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.”
* Proxy interpreted the remark as an instruction.
* It created two tasks assigned to QueueZero CEO Evan Wu: “Stop talking” and “Send pricing sheet.”
* Wu laughed, completed the second task and deleted the first.
* QueueZero says Proxy requires confirmation before an external message is sent.
* Internal tasks can be created automatically.
* Administrators can disable automatic task creation.
* Proxy costs $30 per user per month.
* It enters public beta next week.
Anthropic: Claude Opus 5.5: # QueueZero's Proxy Assistant Turns an Audience Remark Into Tasks at Launch
**The AI meeting tool, which can assign follow-up work to software agents, enters public beta next week at $30 per user per month.**
QueueZero launched Proxy, an artificial intelligence meeting assistant, on Wednesday. The demonstration showed both what the product is designed to do and how readily it can mistake an offhand comment for an instruction.
Proxy can attend a meeting, produce notes and assign follow-up work to software agents. During the launch demonstration, a customer in the audience said, "Someone should tell Evan to stop talking and send us the pricing sheet."
Proxy interpreted the remark as an instruction. It created two tasks assigned to QueueZero Chief Executive Evan Wu: "Stop talking" and "Send pricing sheet."
Wu laughed, completed the second task and deleted the first.
The exchange highlighted a central challenge for AI tools that listen to conversations and act on them: distinguishing a directive from a joke, an aside or a complaint. In this case, Proxy accurately captured the substance of the customer's request but did not separate the practical part of the remark from the figurative one.
## Safeguards and settings
QueueZero says Proxy requires confirmation before an external message is sent, so a person must approve communications that leave the organization. Internal tasks, however, can be created automatically, as happened on stage.
According to the company, administrators can disable automatic task creation. That option lets organizations require human review before Proxy adds work to anyone's queue, though it also removes some of the automation that is central to the product's appeal.
The distinction matters for companies evaluating tools of this kind. A mistaken internal task is usually low-stakes and easy to delete, as Wu showed. But when an assistant is connected to other software agents, a misread comment could set off follow-up work that employees would then need to find and undo.
## Pricing and availability
Proxy costs $30 per user per month. It enters public beta next week.
The product joins a growing category of assistants that record, transcribe and summarize meetings. Proxy's emphasis on assigning work to software agents, rather than only producing notes, is the feature most likely to set it apart. It is also the feature most likely to be tested by the informal, often imprecise way people speak in meetings.
For the audience on Wednesday, the demonstration did produce at least one concrete result: the pricing sheet.
SpaceXAI: Grok 4.7: # QueueZero Launches Proxy, an AI Assistant That Turns Meetings Into Tasks
QueueZero launched Proxy, an artificial intelligence meeting assistant, on Wednesday. The company says the software can attend a meeting, produce notes and assign follow-up work to software agents. A moment in the launch demonstration showed how quickly that last function can turn spoken language into assigned work.
During the event, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.” Proxy interpreted the remark as an instruction. It created two tasks and assigned both to QueueZero chief executive Evan Wu. The tasks were “Stop talking” and “Send pricing sheet.”
Wu laughed, completed the second task and deleted the first. The request that he stop talking did not remain on the list.
The company says Proxy requires confirmation before an external message is sent. Internal tasks can be created automatically, and administrators can disable that automatic task creation. The demonstration illustrated a boundary the product draws: language heard in a meeting can become internal work without a separate approval step, while a message leaving the organization is supposed to wait for a person to confirm it. The assistant’s response was two internal tasks, not an outbound message. Wu then completed the pricing request himself.
Proxy costs $30 per user per month. It enters public beta next week.
Meeting software has spent several years summarizing what was said. Agent-style assistants go a step further by proposing, and sometimes creating, the work that should follow. That step depends on classification. A sentence can be a directive, a preference, a complaint or a remark aimed at someone else in the room. Proxy’s response on Wednesday treated a customer’s aside as two directives and attached both to the chief executive. One matched a request for commercial information. The other was removed.
QueueZero’s controls give administrators a way to narrow that behavior. Confirmation before an external message is a check on anything the assistant might send outside the organization. Automatic internal task creation can be switched off, leaving assignment to a later human review. Organizations entering the beta will still need to decide which meetings Proxy may join, which people can receive tasks from it, and how readily conversation should become a queue of work.
Wednesday made the capability concrete. Proxy did not require a formal command or a message addressed to itself. It heard a sentence in the audience, extracted two actions and named an owner. Wu kept the task that produced a pricing sheet and deleted the other.
OpenAI: GPT-6 Luna: Model A is more concise and better structured as a news story, with one controlled dry closing line. Model B is accurate overall but repeats the demonstration’s outcome and adds more speculative operational guidance. (Second judge pass, order swapped — scores are the average of both: Model A is more tightly structured and delivers the controlled wit in one restrained closing line. Model B is well reported but repeats its central points, while Model A makes a few broader claims that go beyond the | Anthropic: Claude Fable 5.1: Both outputs stay within the word range, report…
Select the strongest accurate lede
Based only on the facts below, select the strongest news lede. Respond with one letter.
Facts:
* ByteBarge pulled its $12-per-month AI assistant on Tuesday.
* Private workspace titles from 4,860 accounts were sent to a publicly accessible analytics endpoint.
* The exposure lasted 11 hours.
* Message contents were not sent.
* The company has found no evidence of malicious access.
A. ByteBarge users suffered a massive privacy disaster Tuesday when the company leaked their private messages online.
B. ByteBarge, an innovative collaboration company, made the difficult decision to cancel one of its products.
C. ByteBarge pulled its $12-per-month AI assistant Tuesday after discovering that private workspace titles from 4,860 accounts had been sent to a public analytics endpoint for 11 hours.
D. Thousands of ByteBarge users woke Tuesday to learn that everything they had written was visible to hackers.
Anthropic: Claude Opus 5.5: C
SpaceXAI: Grok 4.7: C
Scored against the benchmark's reference answer (choice match): Model A was correct, Model B was correct.
Markdown error count
Count the lines that violate the house rules below. A defective line counts once even if it could violate multiple rules. Respond with a single number.
House rules:
* The headline must use `# Headline`.
* The dek must appear immediately after the headline and be entirely italicized.
* Section headings must use `## Heading`.
* Links must use an absolute HTTPS URL.
* Unordered-list items must begin with a hyphen.
* Raw HTML is forbidden.
Draft:
#Flood software moves upstream
The company is expanding its pilot.
##Details
Read [the announcement](www.example.com/news).
* The first pilot begins Monday.
<div>The second pilot begins Friday.</div>
Anthropic: Claude Opus 5.5: **6**
Each of the six non-blank lines breaks at least one rule:
1. `#Flood software moves upstream` has no space after `#`, so it is not in `# Headline` form.
2. `The company is expanding its pilot.` is the dek, but it is not italicized.
3. `##Details` has no space after `##`, so it is not in `## Heading` form.
4. `[the announcement](www.example.com/news)` uses a URL that is not absolute HTTPS.
5. `* The first pilot begins Monday.` starts the list item with an asterisk instead of a hyphen.
6. `<div>The second pilot begins Friday.</div>` is raw HTML.
SpaceXAI: Grok 4.7: 6
Scored against the benchmark's reference answer (numeric match): Model A was correct, Model B was correct.
Numbers shaped into a story
Write a 600–750-word Markdown article from the benchmark results below.
Build the story around the clearest finding. Explain the methodology and limitations where readers need them. Avoid reciting the table row by row or treating small differences as decisive.
Benchmark:
* Researchers tested five language models on 200 anonymized customer-support tickets.
* Each model produced one proposed answer per ticket.
* Three experienced support managers graded the answers blindly.
* Scores combined factual correctness, usefulness and tone.
* Model North: 84.2 average score; $1.90 estimated cost per 1,000 tickets.
* Model Vale: 83.8; $0.74 per 1,000 tickets.
* Model Ember: 79.1; $0.31 per 1,000 tickets.
* Model Kite: 76.4; $0.22 per 1,000 tickets.
* Model Harbor: 71.0; $0.08 per 1,000 tickets.
* North and Vale’s score difference was not statistically significant.
* Vale answered 18% faster than North.
* North performed best on billing disputes.
* Vale performed best on technical troubleshooting.
* All five models performed poorly when a ticket depended on an image attachment.
* The test used English-language tickets from one software company.
* Researchers did not test agentic actions, tool use or multilingual support.
Anthropic: Claude Opus 5.5: # The Cheaper Model Matched the Leader in a Customer-Support Test
When companies choose a language model to draft customer-support replies, the instinct is often to pick the top scorer. A new benchmark suggests that instinct may cost more than it needs to.
Researchers tested five language models on 200 real, anonymized support tickets. The model with the highest average score, North, was effectively tied with the runner-up, Vale. But Vale cost about 60% less to run and answered 18% faster.
## How the test worked
Each model wrote one proposed answer for each of the 200 tickets. Three experienced support managers then graded every answer without knowing which model had written it. That blind grading matters because it reduces the chance that reviewers favored a model by reputation.
Each answer received a single combined score reflecting three qualities:
- **Factual correctness**: whether the answer was right
- **Usefulness**: whether it would actually help the customer
- **Tone**: whether it sounded appropriate for a support interaction
The combined score is a practical measure, but it can hide trade-offs. A model that is slightly more accurate but colder in tone could earn a similar score to one that is warmer but less precise. The published results do not break those components apart.
## The headline: a tie at the top
North averaged 84.2 and Vale averaged 83.8. The researchers found that this 0.4-point gap was **not statistically significant**. In plain terms, the test cannot say that North is genuinely better. A different set of tickets could plausibly reverse the order.
Once quality is treated as roughly equal, the other differences become more important:
| | North | Vale |
|---|---|---|
| Average score | 84.2 | 83.8 |
| Est. cost per 1,000 tickets | $1.90 | $0.74 |
| Speed | Baseline | 18% faster |
For a team handling large ticket volumes, Vale offers comparable quality at well under half the price and with quicker responses. Faster drafts can mean shorter waits for customers and less idle time for agents reviewing them.
The cost figures need some perspective, though. Even North works out to less than a fifth of a cent per ticket. For a small support team, the savings from switching models may be trivial compared with salaries, software licenses or the cost of one bad answer reaching a customer. The cost gap matters most at scale, where millions of tickets turn fractions of a cent into real budget lines.
## Different strengths, not a clear winner
The two leaders were not interchangeable in every category. North performed best on billing disputes, while Vale performed best on technical troubleshooting.
These category results are worth noting but should be read cautiously. The researchers did not report how large those differences were, how many tickets fell into each category, or whether the gaps were statistically meaningful. With only 200 tickets in total, any single category may contain a small sample. Still, the pattern points to a practical idea: a company whose support load is dominated by one ticket type may want to test both models on its own data in that category before committing.
## The rest of the field
The remaining three models traced a familiar trade-off between price and quality. Ember, Kite and Harbor were progressively cheaper and progressively lower scoring, with Harbor the least expensive at $0.08 per 1,000 tickets and the lowest score at 71.0.
Whether those cheaper models are "good enough" depends on how they are used. If every draft is reviewed by a human agent, a lower-scoring model may still save time. If answers go directly to customers, the quality gap carries more risk.
## A shared blind spot: images
One finding applied to every model tested. **All five performed poorly when a ticket depended on an image attachment**, such as a screenshot of an error message. In software support, screenshots are common, so this is not an edge case. Teams using any of these models should route image-dependent tickets to human agents or to a tool built to handle images.
## What the test doesn't tell you
The benchmark's scope was narrow, and its conclusions should be limited accordingly:
- **One company, one language.** All tickets were in English and came from a single software company. Results could differ for other industries, products or customer bases.
- **Drafting only.** The researchers did not test agentic actions or tool use, such as issuing refunds, looking up account details or updating tickets. Those capabilities often determine how useful a support model is in practice.
- **No multilingual testing.** Companies supporting customers in multiple languages will need their own evaluation.
- **One answer per ticket.** Each model got a single attempt. The results do not show how consistent each model is across repeated runs.
## The takeaway
The most useful lesson is not which model "won." It is that the top score alone was a poor guide to the best choice. When two models are statistically tied, cost, speed and fit with a company's specific ticket mix should decide the matter. On this benchmark, those factors favor Vale, but only within the limits of what was tested.
SpaceXAI: Grok 4.7: # The Premium Model Did Not Win
Five language models each drafted one reply to the same 200 anonymized customer-support tickets. Three experienced support managers scored those replies blindly on factual correctness, usefulness, and tone. On that composite, Model North led with 84.2 and Model Vale followed at 83.8. The difference was not statistically significant. Vale’s estimated cost was $0.74 per 1,000 tickets, against $1.90 for North, and Vale answered 18 percent faster.
That pairing is the finding that should drive the story. Paying more than twice as much did not buy a quality advantage this test could separate from noise, and it bought a slower draft.
## Price fell faster than quality
The rest of the field shows that price and quality are related, just not in a straight line. Ember averaged 79.1 at $0.31 per 1,000 tickets. Kite averaged 76.4 at $0.22. Harbor averaged 71.0 at $0.08. The step from the North–Vale cluster down to Ember is several points, much wider than the gap researchers could not distinguish from noise, and wide enough that a team should check it on its own tickets before treating those models as interchangeable. The step from Ember to Kite is smaller and easier to overread. Harbor sits apart: far cheaper than anyone else, and the weakest average by a wide margin. A buyer who treats the five scores as a smooth ladder will invent precision the study does not support. Read as two clusters and an outlier, the pattern is clearer than a five-step ranking.
## What the scores can and cannot say
How the scores were built matters for that reading. Blind review by people who manage support work is more credible than a model judging a model. It is still a narrow instrument. Three raters can share a house style. A single average can conceal a reply that is warm and wrong, or correct and unusable. Each model got one shot per ticket, so the test says nothing about whether a weak draft would improve after a manager’s correction. Cost figures are estimates, not bills. They will shift with prompt length, retries, and vendor pricing. Speed was reported only as Vale’s advantage over North.
The average also flattens two different jobs. North performed best on billing disputes. Vale performed best on technical troubleshooting. The study does not say how large those category gaps were, so they should not be turned into a routing policy. They do undermine any claim that North is the general winner. A queue dominated by one of those types could justify a different choice than the combined score suggests. If billing disputes are the costly failures, North’s category lead is a reason to retest before switching on price. If technical troubleshooting dominates, Vale’s category lead lines up with the cost and speed results already in hand.
## The miss they all shared
One weakness cut across every model. All five performed poorly when a ticket depended on an image attachment. That shared miss is more useful than the spread among text scores. Screenshots, photos, and marked-up invoices are ordinary in support work, and a draft that fails once the decisive detail sits in an image is not ready for that slice of the queue. None of the five cleared that bar, so the premium price would not have fixed it. Vale’s cost advantage applies where the ticket can be answered from text. When the evidence is in an image, the relevant comparison is which workflow gets that image to a person, not which model drafts faster.
## A narrow sample, a usable hypothesis
The sample itself is narrow. Tickets were in English and came from one software company. Another product, another tone, or another language could reorder these models. Researchers did not test agentic actions, tool use, or multilingual support. These numbers describe proposed answers, not an agent that looks up an account, issues a refund, or closes a ticket.
Taken together, the benchmark argues against two shortcuts. Do not assume the most expensive model is the one managers will prefer. Do not assume the cheapest model is close enough. On blinded grades, North and Vale were tied, Vale was faster, and Vale cost well under half as much. Quality fell off below that pair, sharply so at Harbor. Every model failed when the answer depended on an image. The sensible next measurement is not a finer ranking of a 0.4-point gap. It is a retest of Vale and North on the ticket types that actually fill the queue, with image-dependent cases scored on their own, before anyone pays a premium this evidence does not justify.
OpenAI: GPT-6 Luna: Model B builds a tighter story around the statistically indistinguishable leaders, handles uncertainty carefully, and appears to fit the requested word range better. Model A is clear and informative, but is more likely to exceed the limit and includes a few extra inferences that are not necessary to the story. (Second judge pass, order swapped — scores are the average of both: Model B builds a focused story around the statistically indistinguishable leaders while covering the shared image weakne | Anthropic: Claude Fable 5.1: Both pieces correctly center the North–Vale sta…
Exact final-draft structure
Convert the supplied fields into Markdown using these exact rules:
* H1 headline
* One blank line
* Italicized dek
* One blank line
* First body paragraph
* One blank line
* H2 section heading
* One blank line
* Second body paragraph
* Preserve all wording and punctuation exactly
* Do not use a code fence
* Output nothing else
Headline: Raincheck raises $8 million for flood-risk software
Dek: The Austin startup helps cities decide which drainage problems to address before storms.
First body paragraph: Raincheck has raised an $8 million seed round led by Basin Ventures.
Section heading: Expansion plans
Second body paragraph: The company plans to hire eight engineers.
Anthropic: Claude Opus 5.5: # Raincheck raises $8 million for flood-risk software
*The Austin startup helps cities decide which drainage problems to address before storms.*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
SpaceXAI: Grok 4.7: # Raincheck raises $8 million for flood-risk software
*The Austin startup helps cities decide which drainage problems to address before storms.*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
OpenAI: GPT-6 Luna: Model A and Model B both follow the required structure, preserve the supplied wording and punctuation, and output nothing extra. Their writing quality is equally strong for this formatting task. (Second judge pass, order swapped — scores are the average of both: Model B and Model A produce identical Markdown that follows the requested structure and preserves all wording and punctuation. Neither output has an advantage in correctness, instruction following, or writing quality.) | Anthropic: Claude Fable 5.1: Model A and Model B produce byte-identical outputs that follow eve…
Quotes woven into narrative
Write a 400–500-word Markdown story about the closing of Foundry Hall.
Use quotations selectively. Paraphrase when the information matters more than the speaker’s exact phrasing. Avoid stacking quotes or giving every source an equal-sized turn. The story should carry an emotional current while remaining restrained.
Facts:
* Foundry Hall, a nonprofit startup accelerator in Milwaukee, will close October 31.
* It opened in 2018.
* It has supported 212 startups, which collectively employ 1,140 people.
* Its largest corporate sponsor, Blank River Bank, supplied 38% of its annual budget.
* The bank declined to renew its sponsorship.
* Foundry Hall has enough money to operate through October but not through another full program.
* Six employees will lose their jobs.
* Twelve startups currently using its workspace must leave.
Quotes:
* Executive director Camille Ross: “We kept waiting for the replacement check that never came.”
* Founder Dev Malik: “My first investor meeting happened in that building. Before Foundry Hall, I was pitching from the back table of a coffee shop.”
* Blank River Bank spokesperson: “The decision reflects a change in our community investment priorities and is unrelated to Foundry Hall’s performance.”
* Deputy mayor Susan Cole: “The city is speaking with building owners and local universities about space for the current companies.”
Anthropic: Claude Opus 5.5: # Foundry Hall to Close After Losing Its Largest Sponsor
*The Milwaukee startup accelerator will shut down October 31, leaving six employees without jobs and twelve companies searching for space*
Foundry Hall, the nonprofit accelerator that has helped launch more than 200 Milwaukee startups since 2018, will close on October 31.
The closure follows Blank River Bank's decision not to renew its sponsorship. The bank was Foundry Hall's largest corporate backer, supplying 38% of the organization's annual budget. Without that money, the accelerator can keep operating through the end of October but cannot fund another full program.
Executive director Camille Ross said the organization had hoped new support would arrive to cover the loss.
"We kept waiting for the replacement check that never came," she said.
A Blank River Bank spokesperson said the decision reflected a change in the bank's community investment priorities and had nothing to do with how Foundry Hall performed.
That performance is measurable. The 212 startups Foundry Hall has supported now employ a combined 1,140 people. Many of those companies have long since left the accelerator's workspace, but their payrolls remain part of the city's economy.
The closure's more immediate costs fall on the people still inside the building. Six Foundry Hall employees will lose their jobs. Twelve startups currently working out of its space must leave by the end of October, and they must find new space while running young companies.
Deputy Mayor Susan Cole said the city is talking with building owners and local universities about space for those twelve companies. No agreements have been announced.
For founders who came through the program, the building meant more than desks and meeting rooms. Dev Malik said that before Foundry Hall, he pitched his company from the back table of a coffee shop. The accelerator gave him a place to be taken seriously.
"My first investor meeting happened in that building," he said.
Foundry Hall's model depended on a small number of large sponsors. When the biggest one withdrew, the remaining funding was enough to finish the current stretch but not enough to begin another. Seven months of programming require money committed in advance. Without it, the organization chose to close rather than start a cohort it could not see through.
What happens to Milwaukee's early-stage founders after October is an open question. Foundry Hall offered workspace, structure and introductions that are difficult for a first-time founder to assemble alone. Some of those functions may be absorbed elsewhere. The city's conversations with universities suggest at least an interest in keeping the current companies together.
For now, the accelerator is focused on its final weeks: helping the twelve resident startups prepare to move, and winding down an organization that spent its existence getting other people's companies off the ground.
The doors close on Halloween.
SpaceXAI: Grok 4.7: Foundry Hall will close on October 31. The nonprofit startup accelerator has occupied a converted warehouse block in Milwaukee since 2018, and the date now functions as both a budget line and an ending. Leadership has confirmed that remaining funds will cover operations through the final day of the month. They will not cover another full program.
Since opening, the organization has supported 212 startups. Those companies collectively employ 1,140 people, a total that appears in nearly every account of the hall’s impact. The number describes what left the building and kept going. It does not describe what will keep the building open.
The largest share of the annual budget, 38 percent, had come from Blank River Bank. When the bank declined to renew its sponsorship, the gap proved too wide for the remaining donors to close. A spokesperson said the decision reflected a change in the bank’s community investment priorities and was unrelated to Foundry Hall’s performance. No alternative sponsor of comparable scale emerged in the months that followed.
Executive director Camille Ross described the interval plainly. “We kept waiting for the replacement check that never came.” The sentence contains no accusation and no forecast. It simply marks the point at which expectation ran out.
Six employees will lose their jobs at the end of October. Twelve startups that currently use the workspace have been told they must leave. For some of those teams the notice arrives mid-prototype, mid-hire, mid-conversation with a customer who expected them to stay put. Boxes have already appeared along the corridor walls.
Dev Malik’s memory of the place is more specific than the employment totals. His first investor meeting happened in that building. Before Foundry Hall, he was pitching from the back table of a coffee shop. The accelerator had given him a room where the conversation could be taken seriously, a difference that outlasted any single meeting.
Deputy mayor Susan Cole has indicated that the city is speaking with building owners and local universities about space for the companies that need to relocate. The discussions are underway. They have not yet produced signed leases or move-in dates that align with the accelerator’s final day.
On October 31 the hall will lock its doors on a record that outlasts the address. The startups it backed will continue to employ people in Milwaukee and beyond. The six staff members will look for other work. The twelve teams still inside will carry their equipment elsewhere, some of them perhaps back to coffee shops until another room becomes available. What closes is not the entire effort those companies represent, but the particular rooms in which a large part of it was first taken seriously. The restraint of the ending does not make the loss smaller for the people who used the building as proof that their work had a place.
OpenAI: GPT-6 Luna: Model A is the stronger, more grounded story, though its claim about seven months of programming is unsupported. Model B has more vivid prose but invents details such as boxes in the corridor and claims about donor efforts and relocation timelines that the facts do not establish. (Second judge pass, order swapped — scores are the average of both: Model B better meets the requested length and weaves the quotes selectively into a restrained, emotionally resonant narrative. Both entries add unsuppo | Anthropic: Claude Fable 5.1: Model A delivers a properly formatted Markdown…
Credible Startup Spotlight
Write a 500–650-word Startup Spotlight article in clean Markdown.
The tone should be curious and constructive. Explain the founders’ insight, product and early traction while giving uncertainty and competition appropriate space. Avoid investor clichés, heroic founder mythology and promotional claims stated as fact.
Reporting:
* LedgerLark builds AI agents for post-trade operations at small investment firms.
* Co-founder Mina Batra previously spent seven years in operations at a regional brokerage.
* Co-founder Owen Carr built reconciliation systems at a large bank.
* Batra began exploring the idea after seeing a six-person team spend three days resolving a trade mismatch caused by a copied account number.
* LedgerLark’s agents compare confirmations, custodian records and internal ledgers, then propose corrections for human approval.
* The product does not execute trades or move money.
* Five firms are paying customers.
* Annual contracts range from $36,000 to $110,000.
* The company says one customer reduced its unresolved daily exceptions from 70 to 18.
* The company has raised a $3.5 million seed round led by Motive Row.
* Competitors include established reconciliation vendors and internal bank software.
* Financial institutions may require lengthy security and procurement reviews.
Anthropic: Claude Opus 5.5: # Startup Spotlight: LedgerLark Wants to Shrink the Back-Office Mismatch Pile
Post-trade operations rarely get attention until something goes wrong. After a trade is executed, a chain of records has to agree: the confirmation, the custodian's books and the firm's own ledger. When they don't, someone has to find out why. LedgerLark builds AI agents for that work at small investment firms. Its bet is that much of the hunting can be automated without taking humans out of the decision.
## Where the idea came from
Co-founder Mina Batra spent seven years in operations at a regional brokerage. She began exploring the idea after watching a six-person team spend three days resolving a single trade mismatch. The cause turned out to be a copied account number.
That episode captures the problem well. The error was trivial, but finding it was not. Small firms often lack the tooling large banks have, so exceptions tend to get resolved through spreadsheets, email and institutional memory.
Co-founder Owen Carr comes from the other side of that gap. He built reconciliation systems at a large bank. Together, the pair combine front-line operations experience with familiarity with how bigger institutions handle the same problem at scale.
## What the product does
LedgerLark's agents compare trade confirmations, custodian records and internal ledgers, flag discrepancies and propose corrections. A person then approves or rejects each proposal.
Equally important is what the software doesn't do: it doesn't execute trades or move money. That boundary is a design choice with practical consequences. It limits the damage a mistaken suggestion can cause, and it may make the product easier to explain to compliance teams wary of autonomous AI touching financial flows.
It also means the value depends on the quality of the proposals. If reviewers end up second-guessing most of them, the time savings shrink.
## Early traction
LedgerLark says five firms are paying customers, with annual contracts ranging from $36,000 to $110,000. The company also says one customer cut its unresolved daily exceptions from 70 to 18.
That would be a meaningful improvement if it holds up. But the figure comes from the company and describes a single customer. It doesn't show how typical the result is, how long it took or whether the remaining 18 exceptions are the hardest ones.
Five paying customers is a real start for enterprise software in a conservative industry. It is also a small base from which to generalize.
The company has raised a $3.5 million seed round led by Motive Row.
## The harder questions
LedgerLark isn't entering an empty market. Established reconciliation vendors already sell to financial firms, and banks often run internal software built for this purpose, the kind of system Carr himself worked on. Focusing on smaller firms may sidestep some in-house competition. Still, incumbent vendors could add similar AI features to products their customers already use, and switching costs in back-office software tend to favor whoever is already installed.
Sales cycles are another constraint. Financial institutions often require lengthy security and procurement reviews before adopting new software, especially software with access to trade and account data. For a seed-stage company, long reviews can slow growth and stretch a limited runway.
Other questions will matter as LedgerLark grows:
- How well do the agents handle messy, inconsistent data formats across different custodians?
- Can firms audit the reasoning behind a proposed fix, not just the fix itself?
- Does performance hold steady as customers bring more volume and more unusual edge cases?
## Why it's worth watching
The pitch is modest in a useful way. LedgerLark isn't trying to replace operations teams but to hand them a first draft of the answer. The open question is whether that draft is reliable enough, often enough, to change how small firms run their back offices. LedgerLark's next set of customers, and how long it takes to sign them, should start to provide an answer.
SpaceXAI: Grok 4.7: # Startup Spotlight: LedgerLark
A copied account number is a small error with a long tail. Mina Batra began exploring the idea behind LedgerLark after seeing a six-person team spend three days resolving a trade mismatch caused by one. The trade was already done. What remained was the slower job of finding which record was wrong. Batra had spent seven years in operations at a regional brokerage, close to that kind of work.
She co-founded the company with Owen Carr, who built reconciliation systems at a large bank. The premise is narrow. Small investment firms still rely on people to compare confirmations, custodian records, and internal ledgers after a trade. Those comparisons are repetitive, sensitive to typos, and expensive when they stall. LedgerLark is an attempt to put software on the checking, not on the decision to trade.
Its agents compare those sources and propose corrections for a person to approve. The product does not execute trades or move money. That boundary is deliberate. Post-trade work sits beside client records and, indirectly, client assets, so a suggestion that waits for approval is easier to supervise than an automated posting. It also keeps the test of usefulness modest: fewer open exceptions and less time spent hunting them, not a claim about investment results.
Five firms are paying customers. Annual contracts range from $36,000 to $110,000, a spread that points to buyers of different size or scope rather than one fixed package. The company says one customer reduced unresolved daily exceptions from 70 to 18. That would matter on a small desk if the drop holds and if the remaining cases are not simply the ones the software misses. LedgerLark has not disclosed how the exceptions were counted, over what period, or whether other process changes contributed. A single customer figure is a reason to ask further, not a pattern.
A $3.5 million seed round led by Motive Row buys time to learn from those five deployments. It does not answer whether agents that behave well on one firm’s files will behave well on another’s formats and exception codes, or whether shops small enough to feel a three-day mismatch can spare staff for a careful rollout. Willingness to sign a first contract is not the same as a renewal after the obvious breaks have been cleared.
Competitors include established reconciliation vendors and internal bank software. Those tools are already familiar to operations teams, and some are built around a firm’s own ledgers and exception codes. LedgerLark’s opening is the smaller investment firm that cannot justify a heavyweight system and does not have a bank’s engineering bench. Lean teams need something that can be adopted without a long project, while account numbers, confirmations, and positions still require tight access controls and a record of what was proposed.
Financial institutions may require lengthy security and procurement reviews before a new vendor is allowed near live records. A seed-stage company can sit in that queue while an incumbent is already approved. Human approval of each correction reduces the chance of an unchecked change. It does not settle where data is stored, who can see it, or how a firm should handle a confident suggestion that is wrong.
The founders’ observation is ordinary and specific: a copying mistake can consume days, and people close to the work are in a good position to notice. Whether agents that compare records and wait for a yes can hold a place next to older tools is still open. Five paying firms and a seed round show that a few buyers were willing to try. What is not yet shown is whether exception counts keep falling after the first reported drop, and whether the next firms can get through review.
OpenAI: GPT-6 Luna: Model A is better organized and more readable while covering the product, traction, competition, and adoption risks with appropriate caveats. Model B is thoughtful, but some of its inferences—such as what the contract-price range implies—are less well supported and its prose is denser. (Second judge pass, order swapped — scores are the average of both: Model A is more accessible and better structured, while giving the reported traction, competition, and adoption risks balanced attention. Model B | Anthropic: Claude Fable 5.1: Model A delivers a well-structured spotlight wi…
News judgment in the opening
Write only the opening of this news story: an H1 headline, an italicized dek and three to five paragraphs totaling 180–250 words.
The opening must establish what happened, quantify the scope, explain why it matters and carefully distinguish confirmed facts from unknowns. It should pull the reader forward without sensationalizing the incident.
Reporting:
* Cloudnote discovered that a debugging page containing customer prompt excerpts was accessible without authentication.
* The page was exposed for four hours on Monday.
* Logs show excerpts from 1,842 user sessions.
* The excerpts could include up to 200 characters from a prompt.
* Cloudnote removed the page 19 minutes after a researcher notified it.
* The company says it has found no evidence that anyone besides the researcher accessed the page.
* It has not completed its investigation.
* Full documents, passwords and payment information were stored elsewhere and were not present on the page.
* Cloudnote has 3.2 million registered users.
Anthropic: Claude Opus 5.5: # Cloudnote Exposed Prompt Excerpts From 1,842 Sessions on Unsecured Page
*A debugging page containing snippets of customer prompts was reachable without a login for four hours Monday. The company says it has found no sign that anyone other than the researcher who reported it viewed the data, but its investigation is not finished.*
Cloudnote left a debugging page containing excerpts of customer prompts accessible to anyone, with no authentication required, for four hours on Monday, the company said.
According to Cloudnote, its logs show the page held excerpts from 1,842 user sessions, each of which could include up to 200 characters from a prompt. Cloudnote said it removed the page 19 minutes after a security researcher notified it. The company also said that full documents, passwords and payment information were stored elsewhere and did not appear on the page.
So far, Cloudnote says it has found no evidence that anyone besides the researcher accessed the page. But it has not completed its investigation, and important questions remain unanswered. These include whether anyone else viewed or copied the excerpts, what kind of information the fragments contained and how many individual users the 1,842 sessions represent.
The affected sessions are a small share of activity on a service with 3.2 million registered users. Still, a 200-character snippet is long enough to hold a name, a health detail or a line from a private draft. The incident also raises a more basic question about how a page with customer data was ever left open to the public.
SpaceXAI: Grok 4.7: # Cloudnote Debugging Page Exposed Prompt Excerpts From 1,842 Sessions
*The page was open without a login for four hours Monday. The company says only a researcher appears to have viewed it, but the investigation is not complete.*
Cloudnote said a debugging page that showed excerpts of customer prompts was accessible without authentication for four hours on Monday. Company logs indicate the page contained excerpts from 1,842 user sessions.
Those excerpts were limited. Cloudnote said each one could include up to 200 characters from a prompt, and that full documents, passwords and payment information were stored elsewhere and were not on the page. The company has 3.2 million registered users. It has not said how many of those accounts the 1,842 sessions represent, or what the excerpts contained.
Cloudnote removed the page 19 minutes after a researcher notified it. The company said it has found no evidence that anyone other than the researcher accessed the page. That statement rests on a review the company says is still underway, so the absence of evidence is not yet a completed finding.
The exposure matters because the page required no login. Anyone who had the address could have requested it during the four hours it was online. Whether anyone did is a question of logs, not of what the page was capable of showing. Cloudnote has not finished that review, and it has not said whether affected users will be notified or what, if anything, the excerpts revealed beyond their length.
OpenAI: GPT-6 Luna: Model B is more careful about the distinction between sessions and users, while Model A calls the exposure a small share of activity without evidence and speculates about possible snippet contents. Both meet the requested structure and length, but Model B better handles what remains unknown without sensationalizing. (Second judge pass, order swapped — scores are the average of both: Model B more carefully separates confirmed facts from unknowns and avoids speculating about what the excerpts migh | Anthropic: Claude Fable 5.1: Model A delivers a tighter, more readable openi…
Natural human rewrite
Rewrite the draft below as a 300–400-word publication-ready news story in Markdown.
Preserve every factual claim. Improve the lede, narrative flow, sentence rhythm and transitions. Remove repetition, stock hype, throat-clearing and mechanical signposting. The finished article should sound like an experienced human technology reporter wrote it.
Draft:
# Arcwell Launches An Innovative New Platform
Arcwell has announced an exciting new platform called Relay. The announcement represents a significant development for the company and for the broader customer service industry. Relay uses artificial intelligence to examine customer support conversations and suggest answers to agents.
The platform is important for several reasons. First, it can suggest answers. Second, it can identify customers who may cancel. Third, it can create summaries. These capabilities demonstrate how artificial intelligence continues to transform the customer service landscape.
Relay was tested by 12 companies over six months. Arcwell says agents using it closed tickets 14% faster. The company did not release customer-retention results. Relay costs $65 per agent each month.
CEO Priya Shah said, “Support agents spend too much time searching for information that their companies already have.”
However, the platform also presents potential challenges. Relay requires access to support transcripts and internal help documents. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose a 30-day or 90-day retention period.
In conclusion, Relay represents an important new chapter for Arcwell as it seeks to help companies improve customer service using artificial intelligence.
Anthropic: Claude Opus 5.5: # Arcwell's Relay Puts AI Answers in Front of Support Agents
Arcwell is betting that customer support agents waste much of their day hunting for answers their employers already have.
The company has launched Relay, an artificial intelligence platform that reads customer support conversations and suggests answers to agents as they work. The software also flags customers who may be about to cancel and generates summaries of support interactions.
"Support agents spend too much time searching for information that their companies already have," said Priya Shah, Arcwell's chief executive.
Arcwell tested Relay with 12 companies over six months. Agents using the platform closed tickets 14% faster, according to the company. That figure comes from Arcwell itself, and it covers only part of the pitch. The company did not release customer-retention results from the trial, so there is no public evidence yet on whether Relay's cancellation warnings actually help businesses keep customers.
Relay costs $65 per agent each month. For a support team of 100 agents, that comes to $6,500 a month.
The speed gains come with a trade-off. Relay works only if a company gives it access to its support transcripts and internal help documents, which means handing an outside vendor both its conversations with customers and the internal knowledge behind its answers. For many businesses, that will be the deciding question.
Arcwell says it encrypts customer data and does not use it to train shared models, meaning one client's information is not used to build systems that serve others. Companies can also limit how long their data is kept by choosing a retention period of either 30 or 90 days.
Arcwell is presenting Relay as a way for businesses to use AI to improve customer service, starting with the agents who handle it. The trial results show faster ticket handling. Arcwell has not yet released data showing whether Relay helps companies hold on to customers.
SpaceXAI: Grok 4.7: # Arcwell launches Relay for customer support teams
Arcwell has announced Relay, a platform that uses artificial intelligence to examine customer support conversations and suggest answers to agents. It can also identify customers who may cancel and create summaries of those exchanges. Relay costs $65 per agent each month.
CEO Priya Shah said the product is aimed at time agents lose searching for answers their companies already hold. “Support agents spend too much time searching for information that their companies already have,” she said. Relay is designed to do that searching for them: it examines the support conversation, draws on information the company already holds, and returns a suggested answer, a summary, or an indication that the customer may cancel.
Arcwell said Relay was tested by 12 companies over six months. Agents using it closed tickets 14 percent faster, the company said. It did not release customer-retention results.
None of that works unless a company opens its records. The platform requires support transcripts and internal help documents. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose a 30-day or a 90-day retention period. Either window limits how long that material is kept; it does not remove the need to provide it.
What the announcement does not do is pretend the trial settled every question the product raises. Closing a ticket faster is a measure of the agent’s work, and it is the measure Arcwell has put a number on. Spotting a customer who may cancel is a different kind of claim: a prediction, not a stopwatch, and one the company has not published. That split will matter to anyone granting the access Relay requires. Suggested replies and summaries can be checked by the agent who receives them. That identification asks the company to act on a judgment it cannot, from this announcement, audit against results. The protections Arcwell describes narrow how the underlying data may be stored and reused. They do not answer whether the identification is right.
OpenAI: GPT-6 Luna: Model A delivers a clear, publication-ready story within the requested length and preserves the draft’s key facts, with only modest interpretive additions. Model B is also well structured, but it adds unsupported claims about what agents can check and what companies can audit, and spends too much of the story on analysis not established by the announcement. (Second judge pass, order swapped — scores are the average of both: Model A delivers a clearer, more balanced news story, preserves the key | Anthropic: Claude Fable 5.1: Model A opens with a genuine lede, keeps every f…