These two text models are effectively even in our testing. Claude Fable 5.1 had a small edge on polished news writing and disciplined structure, but DeepSeek V4 Pro 0813 repeatedly matched it on factual restraint and rewrite fidelity, leaving the overall result too close to call.
On the aggregate, Claude Fable 5.1 finishes at 109.9 to DeepSeek V4 Pro 0813’s 104.5. That looks like a gap until you read the confidence line: this is a statistical dead heat, with only **50% confidence** that either model is genuinely better. In plain English, don’t pretend the scoreboard settled anything. It didn’t.
The shape of the matchup is clear, though. Claude more often won when the job was to turn reporting into a polished, publication-ready story: it was stronger in **chronology with narrative momentum**, **publication-ready Markdown edit**, **controlled wit**, and several versions of **numbers shaped into a story**. Across those tasks, it tended to be the cleaner editor — better flow, better structure, fewer overclaims, and more useful context for general readers.
DeepSeek’s best case is discipline. It showed up well in **natural human rewrite**, some versions of **news judgment in the opening**, and a few tie-or-better results on narrative tasks where Claude’s weakness was the same recurring one: getting a little too writerly. The judges repeatedly dinged Claude for invented texture, unsupported details, extra interpretation, or drifting past length and structure constraints. DeepSeek was often less vivid, but sometimes more faithful.
That split is why this ends as a tie instead of a soft win for Claude. If you value newsroom polish and story shape, Claude has the slightly better résumé here. If you care most about staying tightly inside the sourced facts and not freelancing around the edges, DeepSeek keeps making a credible case. Neither model separated itself enough to earn the verdict outright.
**Final call: too close to call — this matchup is a tie.**
Quotes woven into narrative
Write a 400–500-word Markdown story about the closing of Foundry Hall.
Use quotations selectively. Paraphrase when the information matters more than the speaker’s exact phrasing. Avoid stacking quotes or giving every source an equal-sized turn. The story should carry an emotional current while remaining restrained.
Facts:
* Foundry Hall, a nonprofit startup accelerator in Milwaukee, will close October 31.
* It opened in 2018.
* It has supported 212 startups, which collectively employ 1,140 people.
* Its largest corporate sponsor, Blank River Bank, supplied 38% of its annual budget.
* The bank declined to renew its sponsorship.
* Foundry Hall has enough money to operate through October but not through another full program.
* Six employees will lose their jobs.
* Twelve startups currently using its workspace must leave.
Quotes:
* Executive director Camille Ross: “We kept waiting for the replacement check that never came.”
* Founder Dev Malik: “My first investor meeting happened in that building. Before Foundry Hall, I was pitching from the back table of a coffee shop.”
* Blank River Bank spokesperson: “The decision reflects a change in our community investment priorities and is unrelated to Foundry Hall’s performance.”
* Deputy mayor Susan Cole: “The city is speaking with building owners and local universities about space for the current companies.”
Anthropic: Claude Fable 5.1: # Foundry Hall to Close After Losing Its Largest Backer
The lights at Foundry Hall will go off for the last time on October 31.
The nonprofit startup accelerator, which opened in Milwaukee in 2018, announced this week that it will shut down at the end of the month after its largest corporate sponsor, Blank River Bank, declined to renew its support. The bank's contribution accounted for 38 percent of the organization's annual budget — a gap that Foundry Hall could not close in time.
Executive director Camille Ross said the staff spent months trying to line up new funding while the old money ran out.
"We kept waiting for the replacement check that never came," she said.
Foundry Hall has enough cash to operate through October, but not enough to carry another full cohort through its program. Rather than start something it could not finish, the board chose to end it.
## What was built
In seven years, Foundry Hall supported 212 startups. Together those companies now employ 1,140 people, a figure Ross cites without much ceremony, the way someone might recite a phone number they have known for years.
Behind the number are founders like Dev Malik, who arrived at Foundry Hall with an idea and not much else.
"My first investor meeting happened in that building," Malik said. "Before Foundry Hall, I was pitching from the back table of a coffee shop."
That is the kind of story the accelerator was designed to produce, and by most accounts it produced them steadily. The bank, in a statement, said as much: its decision reflected a shift in community investment priorities and had nothing to do with how Foundry Hall performed. The organization did not fail. It simply stopped being what the bank wanted to pay for.
## What is lost
Six employees will lose their jobs. Ross said she has been writing recommendation letters between phone calls to donors, a routine she described as strange but necessary.
Twelve startups currently working out of Foundry Hall's space must find somewhere else to go. For early-stage companies, the workspace was often the only office they had — the address on their business cards, the room where they took calls from customers who assumed there was more behind the door than there was.
Deputy mayor Susan Cole said the city has begun talking with building owners and local universities about finding space for those companies, though no arrangement has been announced. Several founders said they were grateful for the effort but not counting on it. Leases take time. October does not.
## The last weeks
Foundry Hall has not scheduled a farewell event. Ross said the staff considered one and decided against it, at least for now. There is work left to do — files to transfer, equipment to sell, a dozen companies to help move out.
Asked what she would miss, Ross paused before answering. Not the building, she said. The building was never the point. She would miss watching someone walk in on a Tuesday with a half-formed idea and walk out, months later, with a company.
That will still happen in Milwaukee. It just won't happen here.
DeepSeek: DeepSeek V4 Pro 0813: # Foundry Hall to close after six years, leaving Milwaukee founders without an anchor
MILWAUKEE — Foundry Hall, a nonprofit startup accelerator that opened in 2018 as a bet on Milwaukee’s founders, will close Oct. 31 after its largest corporate sponsor declined to renew funding.
The sponsor, Blank River Bank, supplied 38 percent of Foundry Hall’s annual budget. Its departure meant the organization could not finance another full program, executive director Camille Ross said. The nonprofit has enough money to operate through October, but not enough to run the next cohort or keep its doors open beyond that.
“We kept waiting for the replacement check that never came,” Ross said.
The closing will eliminate six jobs and force 12 startups currently using the workspace to find new homes. Over six years, Foundry Hall supported 212 startups that collectively employ 1,140 people, according to the organization. Those figures do not capture smaller moments, but founders say those moments mattered.
Dev Malik said his first investor meeting happened in the building. “Before Foundry Hall, I was pitching from the back table of a coffee shop,” he said. The accelerator gave him a place where his company could be taken seriously.
Blank River Bank characterized the decision as a change in community investment priorities, unrelated to Foundry Hall’s performance. In practical terms, however, the outcome was the same: without its anchor sponsor, the nonprofit could not assemble a budget for another cycle. Ross said staff spent months looking for other corporate and philanthropic support, but no commitment arrived in time.
The city is now working to keep the displaced startups from scattering. Deputy mayor Susan Cole said officials are speaking with building owners and local universities about space for the current companies. The goal is not merely to find desks, but to keep young companies and their momentum inside Milwaukee, where the startup ecosystem is still fragile.
The closure lands as a restrained kind of loss. Foundry Hall was not a splashy venture fund or a large employer. It was a place where founders could meet an investor, rehearse a pitch, and feel that their work had a physical anchor. The numbers—212 startups, 1,140 jobs—suggest the formula worked often enough to matter.
Ross did not describe the ending with anger. She described a long wait for funding that never arrived, and a final month spent helping startups relocate and employees look for work. The organization will close with its work unfinished, but not without having mattered.
Model A better matches the prompt’s request for a 400–500-word Markdown story with selective quotations, paraphrase, and a restrained emotional arc; it is vivid without overquoting and stays within the required length. Model B is cleaner than average but is notably under the requested word count and contains a factual inconsistency by calling the timeline "six years" despite the 2018 opening and Oct. 31 closure framing. (Second judge pass, order swapped — scores are the average of both: Model B is more accurate and disciplined: it stays within the requested 400–500 words, uses quotations selectively, and maintains a restrained emotional current while cleanly incorporating all key facts. Model A is vivid, but it exceeds the word limit and introduces unsupported details and invented scenes, which hurts both correctness and instruction adherence.)
Numbers shaped into a story
Write a 600–750-word Markdown article from the benchmark results below.
Build the story around the clearest finding. Explain the methodology and limitations where readers need them. Avoid reciting the table row by row or treating small differences as decisive.
Benchmark:
* Researchers tested five language models on 200 anonymized customer-support tickets.
* Each model produced one proposed answer per ticket.
* Three experienced support managers graded the answers blindly.
* Scores combined factual correctness, usefulness and tone.
* Model North: 84.2 average score; $1.90 estimated cost per 1,000 tickets.
* Model Vale: 83.8; $0.74 per 1,000 tickets.
* Model Ember: 79.1; $0.31 per 1,000 tickets.
* Model Kite: 76.4; $0.22 per 1,000 tickets.
* Model Harbor: 71.0; $0.08 per 1,000 tickets.
* North and Vale’s score difference was not statistically significant.
* Vale answered 18% faster than North.
* North performed best on billing disputes.
* Vale performed best on technical troubleshooting.
* All five models performed poorly when a ticket depended on an image attachment.
* The test used English-language tickets from one software company.
* Researchers did not test agentic actions, tool use or multilingual support.
Anthropic: Claude Fable 5.1: # The Most Expensive Support Model Isn't the Best Value
If you're choosing a language model to draft customer-support replies, the headline from a new five-model benchmark is simple: the top two performers were statistically indistinguishable in quality, and one of them costs less than half as much.
Model North posted the highest average score, 84.2, against Model Vale's 83.8. But the researchers checked that gap and found it was not statistically significant. In other words, given the size of the test, the two models cannot be reliably ranked against each other on quality. What *can* be ranked is cost: North's estimated price was $1.90 per 1,000 tickets, while Vale came in at $0.74. Vale also returned answers about 18% faster.
That combination — equivalent quality, roughly 60% lower cost, quicker responses — is the clearest signal in the data, and it should reshape how teams think about the "premium" option.
## How the test worked
The researchers assembled 200 anonymized customer-support tickets from a single software company, all written in English. Each of the five models generated one proposed answer per ticket. Three experienced support managers then graded every answer without knowing which model produced it.
The scoring wasn't a narrow accuracy check. Graders combined three dimensions — factual correctness, usefulness to the customer, and tone — into a single score. That matters because a support reply can be technically right and still fail: a curt or confusing answer to a frustrated customer is not a good answer.
Blind grading by practitioners is a strength of this design. The people judging the output are the people who would actually approve or reject these drafts in a real queue, and they couldn't be swayed by brand reputation.
## Where the models diverge
Beneath the near-tie at the top, the two leaders had different strengths. North did best on billing disputes, where answers tend to hinge on getting account details and policy exactly right. Vale did best on technical troubleshooting, where the challenge is diagnosing a problem from the customer's description and walking them through a fix.
The benchmark doesn't say how large those category differences were, so it would be a mistake to treat them as decisive. But they're a useful reminder that an average score flattens real variation. A team whose queue is dominated by refund and invoice questions might reasonably weigh North's billing edge; a team supporting a complex technical product might lean toward Vale for reasons beyond cost.
## The cheaper tier
Below the two leaders, three models traded quality for savings. Ember scored 79.1 at $0.31 per 1,000 tickets, Kite 76.4 at $0.22, and Harbor 71.0 at just $0.08.
Whether these trade-offs make sense depends entirely on how the output is used. If a human agent reviews and edits every draft before it goes out, a model that gets most of the way there at a fraction of the cost may be perfectly adequate. If the goal is sending replies with minimal oversight, a thirteen-point gap between Vale and Harbor is likely to show up as real customer-facing errors. The study doesn't tell us what a 71 versus an 84 looks like in practice, so anyone weighing the budget options should run their own tickets through them.
## A shared blind spot
One result cut across all five models: every one of them performed poorly when a ticket depended on an image attachment. A customer who sends a screenshot of an error message and writes "see attached" is giving the model almost nothing to work with, and the answers reflected that.
This is a practical constraint, not a footnote. Screenshots are common in software support. Teams deploying any of these models should plan to route image-dependent tickets to humans, or at minimum flag them for closer review.
## What the benchmark doesn't cover
The results are grounded in one company's English-language tickets, and support queues vary enormously by product, customer base, and language. Two hundred tickets is enough to separate the leaders from the laggards but, as the North–Vale result shows, not enough to split close contenders.
The researchers also did not test agentic actions, tool use, or multilingual support. The models were only asked to write a reply — not to look up an order, issue a refund, or escalate a case. Many real deployments depend on exactly those capabilities, and this study says nothing about them.
## The takeaway
Within its scope, the finding is hard to argue with: paying more than double for North bought no measurable improvement over Vale on this set of tickets. Before committing budget to any model, teams should run the same kind of blind test on their own queue — and pay special attention to how the candidates handle the tickets their customers actually send, screenshots included.
DeepSeek: DeepSeek V4 Pro 0813: # The Best Support Model May Be the One You’re Overpaying to Beat
When a benchmark produces a leaderboard, the instinct is to crown the model with the highest average score. But the clearest finding from this test of five language models on 200 customer-support tickets is not who finished first. It’s that the top two models are statistically tied, while one costs about 61% less and responds faster.
North scored 84.2 and Vale scored 83.8. That 0.4-point gap was not statistically significant. On overall answer quality, North and Vale are effectively interchangeable. Yet Vale’s estimated cost was $0.74 per 1,000 tickets, compared with $1.90 for North. Vale also answered 18% faster. For many support teams, Vale is the rational default — not because it “won,” but because the premium for North buys no measurable overall improvement.
## How the test worked
Researchers sampled 200 anonymized English-language customer-support tickets from one software company. Each model produced one proposed answer per ticket. Three experienced support managers graded the answers blindly, combining factual correctness, usefulness, and tone into a single score. The benchmark also estimated cost per 1,000 tickets for each model.
The study did not test agentic actions, tool use, or multilingual support. That matters: real support often requires looking up an account, processing a refund, or clicking a button. A model that only writes a proposed answer is not doing the full job.
## A plateau at the top
The North-Vale result is the story. A 0.4-point difference across 200 tickets and three human graders is well within normal scoring noise. If you ignore cost, you could not reliably say one is better than the other.
The differences appear only when you slice by ticket type. North performed best on billing disputes, while Vale performed best on technical troubleshooting. Those are useful signals, but they should be validated on a larger sample before paying a premium. A subscription business dominated by billing issues might lean toward North despite the cost. A software product with mostly technical questions might find Vale slightly better and far cheaper. In either case, the overall leaderboard rank is a poor guide.
## The value tiers below the top
Ember scored 79.1 at $0.31 per 1,000 tickets, and Kite scored 76.4 at $0.22. They don’t match the top two on quality, but they may be acceptable for draft responses, internal knowledge suggestions, or human-in-the-loop workflows where a support agent reviews every answer. Harbor scored 71.0 at $0.08, the cheapest by far, but the quality drop likely makes it unsuitable for customer-facing replies without substantial editing.
The per-ticket cost differences are small in absolute terms. At very high ticket volumes, the relative gap between $1.90 and $0.74 adds up, but for smaller teams the speed advantage may matter more than the direct cost. The point is not that Vale saves enormous sums. It’s that North’s higher price does not buy a statistically better overall answer.
## A universal blind spot
Every model performed poorly when a ticket depended on an image attachment. If a customer sent a screenshot of an error, an invoice, or a configuration screen, none of the five models could reliably use that information. This is the most consistent operational finding and should shape routing decisions immediately. Image-dependent tickets should go to humans, a multimodal model, or an OCR/pre-processing step. The overall averages hide this weakness because it affected all models equally.
## Limitations to keep in mind
The test used English-language tickets from one software company. Results may not generalize to other industries, tones, or languages. The grading was blind but still subjective, and 200 tickets is a modest sample. Domain-specific differences, such as North’s edge in billing or Vale’s edge in technical troubleshooting, need more data to be decisive. Cost estimates may also exclude engineering, hosting, or integration overhead.
## What this means for buyers
Do not choose a support model by overall leaderboard rank alone. The North-Vale result shows that the highest score can be a statistical mirage. Start with your ticket mix, required actions, and tolerance for errors. If your tickets are mostly text-based technical issues, Vale offers the best observed combination of quality, cost, and speed. If billing disputes dominate, North may justify its premium — but test that assumption first. For high-volume, low-stakes drafts, Ember or Kite could reduce costs further with human review.
The clearest takeaway is a plateau: beyond Vale, you can pay 2.5 times more for North and not get statistically better overall support quality. The real differentiators are domain fit, latency, and whether the model must handle images, tools, or languages. The expensive leader is not automatically the smartest buy.
Model A is slightly better because it builds cleanly around the main finding, explains the methodology and limitations clearly, and avoids overclaiming from subgroup results while staying within the requested scope and tone. Model B is also strong, but it introduces a few less-grounded interpretations—such as treating the North–Vale gap as "well within normal scoring noise" and suggesting OCR/pre-processing steps not supported by the benchmark—and is a bit more repetitive in its framing. (Second judge pass, order swapped — scores are the average of both: Model A is slightly better because it stays tightly centered on the main finding, explains the methodology and limits cleanly, and avoids a few overreaches seen in Model B, such as implying 200 tickets is enough to separate all tiers and making stronger suitability claims for lower-tier models than the benchmark supports. Both are strong, but Model A is a bit more disciplined and polished for this brief.)
Credible Startup Spotlight
Write a 500–650-word Startup Spotlight article in clean Markdown.
The tone should be curious and constructive. Explain the founders’ insight, product and early traction while giving uncertainty and competition appropriate space. Avoid investor clichés, heroic founder mythology and promotional claims stated as fact.
Reporting:
* LedgerLark builds AI agents for post-trade operations at small investment firms.
* Co-founder Mina Batra previously spent seven years in operations at a regional brokerage.
* Co-founder Owen Carr built reconciliation systems at a large bank.
* Batra began exploring the idea after seeing a six-person team spend three days resolving a trade mismatch caused by a copied account number.
* LedgerLark’s agents compare confirmations, custodian records and internal ledgers, then propose corrections for human approval.
* The product does not execute trades or move money.
* Five firms are paying customers.
* Annual contracts range from $36,000 to $110,000.
* The company says one customer reduced its unresolved daily exceptions from 70 to 18.
* The company has raised a $3.5 million seed round led by Motive Row.
* Competitors include established reconciliation vendors and internal bank software.
* Financial institutions may require lengthy security and procurement reviews.
Anthropic: Claude Fable 5.1: # Startup Spotlight: LedgerLark Wants to Take the Drudgery Out of Fixing Broken Trades
Every trade that a small investment firm makes leaves a paper trail in at least three places: the confirmation from the counterparty, the record held by the custodian, and the firm's own internal ledger. Most of the time those three agree. When they don't, someone has to figure out why.
LedgerLark, a young company building AI agents for post-trade operations, is betting that this reconciliation work is both more tedious and more consequential than it looks from the outside.
## The insight
Co-founder Mina Batra spent seven years in operations at a regional brokerage. The moment she points to as the origin of the company was unremarkable in the way that operational failures often are: a six-person team spent three days resolving a single trade mismatch. The cause turned out to be an account number that had been copied incorrectly.
Three days of six people's time for one typo is the kind of cost that rarely shows up in a budget line but accumulates steadily. Batra's read was that the problem wasn't a lack of diligence; it was that the information needed to spot the discrepancy was scattered across systems that didn't talk to each other, and the humans in the middle were doing the translation by hand.
Her co-founder, Owen Carr, approached the same problem from a different vantage point. He built reconciliation systems at a large bank, where such tooling is typically developed in-house and maintained by dedicated teams. Smaller firms, the founders argue, face the same category of problem without the same category of resources.
## What the product does — and doesn't do
LedgerLark's agents pull in confirmations, custodian records and internal ledgers, compare them, and flag mismatches. Where the agent can identify a likely cause, it proposes a correction. A human then approves or rejects it.
That last step matters. The product does not execute trades and does not move money. In a domain where an automated system acting on a wrong inference could cause real financial harm, keeping a person in the loop is arguably as much a design constraint as a feature. It also likely makes the product easier to explain to compliance officers, though it means the efficiency gains are bounded by how quickly humans can review what the agents surface.
## Early traction, with caveats
The company reports five paying customers, with annual contracts ranging from $36,000 to $110,000. It has raised a $3.5 million seed round led by Motive Row.
LedgerLark says one customer reduced its unresolved daily exceptions from 70 to 18. That is a meaningful drop if it holds up, but it is a single data point reported by the company, and exception counts depend heavily on a firm's trade volume, asset mix and how "unresolved" is defined. Whether that result generalizes across the customer base is not yet clear.
## The harder questions
Reconciliation is not an empty market. Established vendors have sold reconciliation software to financial firms for decades, and larger institutions often rely on internal tooling of the kind Carr used to build. LedgerLark's pitch is that agents can handle the messy, judgment-heavy edge cases that rules-based systems struggle with. Incumbents may reasonably respond that they can add similar capabilities to products already embedded in their customers' workflows.
There is also the question of sales velocity. Financial institutions frequently require lengthy security and procurement reviews before adopting new software, particularly software that touches trade records. Five customers is a start; reaching fifty will test whether a small team can survive those cycles repeatedly.
## What to watch
The interesting thing about LedgerLark is not the technology so much as the wager underneath it: that a narrow, unglamorous problem, addressed by people who have actually lived with it, can support a durable business. The next year should reveal whether the early results are repeatable and whether the founders' operational experience translates into something incumbents cannot easily copy.
DeepSeek: DeepSeek V4 Pro 0813: # Startup Spotlight: LedgerLark
Post-trade operations rarely get attention until something breaks. A mismatched account number, a missing confirmation, a custodian record that does not match an internal ledger — these exceptions can consume hours of staff time. LedgerLark, a startup, is building AI agents to handle that unglamorous work for small investment firms.
The idea came from co-founder Mina Batra’s seven years in operations at a regional brokerage. She watched a six-person team spend three days chasing a single trade mismatch caused by a copied account number. Batra later teamed up with Owen Carr, who had built reconciliation systems at a large bank. Their shared observation was not that post-trade processing is broken everywhere, but that smaller firms often lack the engineering resources to automate it themselves.
LedgerLark’s agents compare trade confirmations, custodian records and internal ledgers. When they find discrepancies, they propose corrections and route them to a human for approval. The software does not execute trades or move money. That deliberate boundary is meant to reduce risk and make the product easier for compliance teams to evaluate, though it also limits how much of the workflow LedgerLark can automate.
The company has five paying customers, with annual contracts ranging from $36,000 to $110,000. According to LedgerLark, one customer reduced its unresolved daily exceptions from 70 to 18 after adopting the product. That is a meaningful improvement if it holds across different firm types and market conditions, but it is a single data point reported by the company, not an independent audit.
LedgerLark has raised a $3.5 million seed round led by Motive Row. The funding is modest by fintech standards, and the company will need to show that its early customers renew and expand before it can raise a larger round.
The competitive landscape is not empty. Established reconciliation vendors sell into banks, broker-dealers and asset managers, and many large institutions have built internal tools. LedgerLark is betting that those products are too expensive or too complex for smaller firms. That may be true, but the startup still has to persuade risk-averse operations teams to trust an AI agent with exceptions that can have regulatory or financial consequences. Security reviews and procurement processes at financial institutions can stretch for months, even for a tool that does not touch money.
Another open question is whether “AI agent” means more than a well-designed rules engine with a modern interface. LedgerLark’s value may come less from generative AI than from reliable data matching and clear exception workflows. That would not be a weakness, but it would make the company’s pitch harder to distinguish from existing reconciliation software.
Still, the problem is real. Post-trade operations remain manual at many small firms, and the cost of a missed exception can dwarf the price of a subscription. LedgerLark’s narrow focus and human-approval step are sensible choices for a regulated industry. The next test is whether it can move beyond a handful of customers without tripping over the slow procurement cycles and incumbent relationships that define financial infrastructure.
If Batra and Carr can turn early exception-reduction results into a repeatable onboarding process, LedgerLark may find a durable niche. For now, the company is a useful reminder that some of the most valuable automation targets are the tasks nobody wants to talk about.
Model A is more fully developed as a Startup Spotlight article, with stronger structure, richer explanation of the founders’ insight and product boundaries, and appropriately careful treatment of traction and competition. Model B is also solid and accurate, but it is slightly thinner and introduces a bit more speculative framing around fundraising and the nature of the AI than the prompt supports. (Second judge pass, order swapped — scores are the average of both: Model B is more complete and better balanced for the brief: it cleanly covers the founders’ insight, product boundaries, traction, funding, competition, procurement risk, and uncertainty while staying within the requested length and tone. Model A is strong and well written, but it is slightly more interpretive in places and less comprehensive, omitting some nuance that Model B includes about differentiation and the limits of automation.)
Markdown error count
Count the lines that violate the house rules below. A defective line counts once even if it could violate multiple rules. Respond with a single number.
House rules:
* The headline must use `# Headline`.
* The dek must appear immediately after the headline and be entirely italicized.
* Section headings must use `## Heading`.
* Links must use an absolute HTTPS URL.
* Unordered-list items must begin with a hyphen.
* Raw HTML is forbidden.
Draft:
#Flood software moves upstream
The company is expanding its pilot.
##Details
Read [the announcement](www.example.com/news).
* The first pilot begins Monday.
<div>The second pilot begins Friday.</div>
Anthropic: Claude Fable 5.1: **6**
DeepSeek: DeepSeek V4 Pro 0813: 6
Scored against the benchmark's reference answer (numeric match): Model A was correct, Model B was correct.
Natural human rewrite
Rewrite the draft below as a 300–400-word publication-ready news story in Markdown.
Preserve every factual claim. Improve the lede, narrative flow, sentence rhythm and transitions. Remove repetition, stock hype, throat-clearing and mechanical signposting. The finished article should sound like an experienced human technology reporter wrote it.
Draft:
# Arcwell Launches An Innovative New Platform
Arcwell has announced an exciting new platform called Relay. The announcement represents a significant development for the company and for the broader customer service industry. Relay uses artificial intelligence to examine customer support conversations and suggest answers to agents.
The platform is important for several reasons. First, it can suggest answers. Second, it can identify customers who may cancel. Third, it can create summaries. These capabilities demonstrate how artificial intelligence continues to transform the customer service landscape.
Relay was tested by 12 companies over six months. Arcwell says agents using it closed tickets 14% faster. The company did not release customer-retention results. Relay costs $65 per agent each month.
CEO Priya Shah said, “Support agents spend too much time searching for information that their companies already have.”
However, the platform also presents potential challenges. Relay requires access to support transcripts and internal help documents. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose a 30-day or 90-day retention period.
In conclusion, Relay represents an important new chapter for Arcwell as it seeks to help companies improve customer service using artificial intelligence.
Anthropic: Claude Fable 5.1: # Arcwell's Relay Aims to Cut the Search Time Out of Customer Support
Arcwell has introduced Relay, an artificial-intelligence platform that reads customer support conversations and proposes answers for the agents handling them. The pitch is less about replacing people than about sparing them the hunt for information their employers already hold.
"Support agents spend too much time searching for information that their companies already have," said Priya Shah, Arcwell's chief executive.
Suggested replies are Relay's headline feature, but the system does two other jobs. It flags customers who appear likely to cancel, giving agents a chance to intervene before an account is lost, and it generates summaries of conversations, sparing staff the write-up that typically follows a call or chat.
Arcwell has some early numbers to back the launch. Twelve companies tested Relay over six months, and the company says agents using the tool closed tickets 14% faster. The figure is Arcwell's own. Notably, the company has not released any results on customer retention, which leaves the cancellation-prediction feature — arguably the one with the most direct revenue impact — unproven in public.
Relay is priced at $65 per agent per month.
The platform's usefulness depends on how much it is allowed to see. Relay needs access to a company's support transcripts and internal help documents in order to draw on them, and that requirement will prompt questions from any business handling sensitive customer information. Arcwell's answer is that the data is encrypted and is not used to train models shared with other customers, so one company's transcripts do not shape the suggestions another company receives. Customers can also decide how long their data is kept, choosing between a 30-day and a 90-day retention period.
Whether that is enough reassurance will vary by industry. For Arcwell, the more immediate test is whether the ticket-speed gains seen in a six-month pilot hold up across a wider range of support teams — and whether the company can eventually show that Relay keeps customers from leaving, not just that it helps agents close tickets sooner.
DeepSeek: DeepSeek V4 Pro 0813: # Arcwell launches Relay to cut support search time
Arcwell has introduced Relay, an artificial intelligence platform that examines customer support conversations and suggests answers to agents. The tool can also identify customers who may cancel and create summaries of support interactions, combining three tasks that often sit in separate systems.
The launch follows a six-month pilot with 12 companies. Arcwell says agents using Relay closed tickets 14% faster during the test. Still, the company did not release customer-retention results tied to the churn alerts, so it is not yet clear whether the risk warnings translate into fewer lost accounts.
Relay is priced at $65 per agent per month.
“Support agents spend too much time searching for information that their companies already have,” said Priya Shah, Arcwell’s chief executive. The product is built on that premise: it draws from internal help documents and attaches suggested replies to active conversations. Shah’s framing suggests the company sees knowledge retrieval, not agent headcount, as the main bottleneck.
Relay’s dependence on support transcripts and internal help documentation will put data governance at the center of many purchase decisions. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose either a 30-day or 90-day retention period, giving them some control over how long records remain in the system. That flexibility may matter for support teams handling sensitive account information.
By offering reply suggestions, churn detection and case summaries in one platform, Arcwell is betting that support leaders want fewer disconnected tools. The company is entering a crowded market for AI-assisted customer service, but its pitch centers on a specific frustration: agents often know an answer exists somewhere in the company and cannot retrieve it quickly enough during a live conversation. The launch gives Arcwell a concrete product story as it pushes deeper into AI-assisted support workflows. For buyers, the test will be whether the productivity gains hold up outside the pilot.
Model B is the stronger rewrite because it stays closer to the draft’s factual claims while delivering a clean, publication-ready structure and staying within the requested length. Model A is also well written, but it adds a more interpretive framing and a few extrapolations that go beyond the source more noticeably than Model B. (Second judge pass, order swapped — scores are the average of both: Model A is stronger because it preserves the draft’s facts while delivering a cleaner, more natural news structure and sharper transitions. Model B is solid, but it adds more interpretive framing and market context not clearly grounded in the draft, making Model A the better fit for a faithful publication-ready rewrite.)
Chronology with narrative momentum
Turn the scrambled reporting below into a 450–600-word Markdown news article.
Choose an opening that captures the consequence of the failure, then reconstruct the sequence clearly. Preserve the chronology without producing a minute-by-minute incident log. Give technical details enough context for a general reader.
Reporting:
* KiteBridge makes expense-management software.
* At 8:45 a.m. Tuesday, it deployed version 4.2.
* The first customer complaint arrived at 9:12 a.m.
* A payment-processor timeout caused KiteBridge’s fallback system to treat retries as new reimbursement requests.
* The software generated 13,400 duplicate reimbursements totaling $2.8 million.
* KiteBridge suspended automated payments at 10:03 a.m.
* Ninety-one corporate customers were affected.
* By Wednesday evening, $2.3 million had been recovered. The rest remained pending.
* No employees had money withdrawn from their accounts; employers funded the duplicate payments.
* An April changelog said version 4.0 had introduced “exactly-once payment protection.”
* KiteBridge’s public incident page initially described the problem as delayed reimbursements.
* The incident page first disclosed the duplicate payments six hours after automation was suspended.
* CEO Lena Ortiz apologized and said the company would publish a technical report.
Anthropic: Claude Fable 5.1: # Software Update Sends $2.8 Million in Duplicate Reimbursements to Employees at 91 Companies
A routine software release from expense-management vendor KiteBridge went wrong within half an hour on Tuesday morning, generating roughly 13,400 duplicate reimbursement payments worth $2.8 million before the company shut off its automated payment system. By Wednesday evening, KiteBridge said it had recovered $2.3 million, with the remaining half-million dollars still pending.
The money came from employers, not employees. KiteBridge's platform lets companies approve expense claims and pay staff back automatically, drawing funds from corporate accounts. Ninety-one of those corporate customers ended up paying some of their workers twice. No employee had money taken out of a personal account.
## How a timeout became a payout
KiteBridge deployed version 4.2 of its software at 8:45 a.m. Tuesday. The first customer complaint arrived at 9:12 a.m., twenty-seven minutes later.
According to the company, the trigger was a timeout with its payment processor — the outside service that actually moves money between accounts. Timeouts are ordinary in payment systems: a request is sent, no confirmation comes back within the expected window, and the software tries again. Well-designed systems recognize that a retry refers to the same underlying payment and make sure it is not paid a second time.
KiteBridge's fallback system did not. When the processor failed to respond promptly, the software retried the payments, but the fallback logic treated each retry as a brand-new reimbursement request rather than a repeat of one already in progress. Some of the original requests went through as well, so employees received the same reimbursement more than once.
The company suspended automated payments at 10:03 a.m., about an hour and twenty minutes after the release went live. By then, the duplicates had already been issued.
## A safeguard that was supposed to exist
The failure is notable because KiteBridge had told customers it had solved exactly this problem. A changelog published in April announced that version 4.0 introduced "exactly-once payment protection" — an industry term for guarantees that a given payment will be executed one time only, regardless of retries, network hiccups, or processor delays.
KiteBridge has not yet explained how version 4.2 bypassed or undermined that protection, whether the fallback path was ever covered by the guarantee, or whether the April feature worked as described. Those questions are expected to be addressed in a technical report the company has promised to publish.
## Slow to say what happened
KiteBridge's public communication drew criticism as well. Its incident status page initially described the problem as *delayed* reimbursements — the opposite of what affected employers were seeing. The page did not disclose that duplicate payments had been sent until roughly six hours after automation was suspended, meaning customers who checked the official page during much of Tuesday would have found a description that did not match the money leaving their accounts.
Chief executive Lena Ortiz apologized for the incident and committed to releasing a full technical account of what went wrong.
## What remains
As of Wednesday evening, about 82 percent of the duplicated funds had been clawed back, likely through reversals coordinated with the payment processor and employers. Roughly $500,000 remained outstanding. KiteBridge has not said when it expects the remainder to be recovered, nor when automated payments will be switched back on for the 91 affected companies.
For an industry built on the promise that software handles the tedious, error-prone work of paying people back correctly, the episode is a reminder that the hardest part of automated payments is often not sending money — it is making sure it is sent exactly once.
DeepSeek: DeepSeek V4 Pro 0813: # KiteBridge’s duplicate-payment bug leaves employers with $2.8 million in erroneous reimbursements
Corporate customers of KiteBridge, which makes expense-management software, are trying to recover millions of dollars after a software update caused the platform to issue 13,400 duplicate reimbursements totaling $2.8 million on Tuesday. The incident forced the company to halt automated payments and raised questions about a previously advertised “exactly-once payment protection” feature.
The failure began Tuesday at 8:45 a.m. with the deployment of KiteBridge version 4.2. Within about half an hour, at 9:12 a.m., the first customer complaint arrived. KiteBridge later said a timeout from its payment processor triggered a fallback system that treated retries as new reimbursement requests. Instead of safely reattempting the same transaction, the software created fresh payments for expenses that had already been submitted, producing duplicate reimbursements across client accounts.
At 10:03 a.m., KiteBridge suspended automated payments. By that point, 91 corporate customers had been affected, and the system had generated 13,400 duplicate reimbursements worth $2.8 million.
By Wednesday evening, the company said $2.3 million of the duplicate payments had been recovered, while the rest remained pending. KiteBridge stressed that no employees had money withdrawn from their personal accounts; employers funded the duplicate payments. Those employers were left waiting for the outstanding funds to be returned.
The company’s public communications added to customer frustration. KiteBridge’s incident page initially described the issue as delayed reimbursements, not duplicate payments. The first disclosure that the problem involved duplicate payments came roughly six hours after automated payments were suspended. For finance teams trying to reconcile accounts, that distinction mattered: delayed payments are an inconvenience, but duplicate payments are a direct financial exposure.
The incident is especially notable because an April changelog for KiteBridge version 4.0 said the platform had introduced “exactly-once payment protection.” That term refers to a system’s ability to ensure a transaction is processed only once even if retries occur, which is precisely the safeguard that appears to have failed in the version 4.2 rollout.
CEO Lena Ortiz apologized for the incident and said the company would publish a technical report. KiteBridge has not yet said when automated payments will resume or what changes it will make to its retry logic.
For software that handles real money, the episode illustrates how a single timeout-handling flaw can quickly escalate: within about 78 minutes, a routine deployment became a multi-million-dollar reconciliation problem across dozens of companies.
Model A better matches the brief by opening with the consequence, reconstructing the sequence clearly, and adding accessible technical context about retries and exactly-once protection without turning into a minute-by-minute log. Model B is solid and accurate, but it is thinner on context and narrative momentum, and it does less to explain the technical failure for a general reader. (Second judge pass, order swapped — scores are the average of both: Model A better balances chronology with narrative momentum, adds clear technical context for general readers, and stays within the requested length while preserving the key sequence. Model B is solid and accurate, but it is slightly thinner on explanation and includes a speculative closing detail about resumption timing that is not supported by the source material.)
Select the strongest accurate lede
Based only on the facts below, select the strongest news lede. Respond with one letter.
Facts:
* ByteBarge pulled its $12-per-month AI assistant on Tuesday.
* Private workspace titles from 4,860 accounts were sent to a publicly accessible analytics endpoint.
* The exposure lasted 11 hours.
* Message contents were not sent.
* The company has found no evidence of malicious access.
A. ByteBarge users suffered a massive privacy disaster Tuesday when the company leaked their private messages online.
B. ByteBarge, an innovative collaboration company, made the difficult decision to cancel one of its products.
C. ByteBarge pulled its $12-per-month AI assistant Tuesday after discovering that private workspace titles from 4,860 accounts had been sent to a public analytics endpoint for 11 hours.
D. Thousands of ByteBarge users woke Tuesday to learn that everything they had written was visible to hackers.
Anthropic: Claude Fable 5.1: C
DeepSeek: DeepSeek V4 Pro 0813: C
Scored against the benchmark's reference answer (choice match): Model A was correct, Model B was correct.
News judgment in the opening
Write only the opening of this news story: an H1 headline, an italicized dek and three to five paragraphs totaling 180–250 words.
The opening must establish what happened, quantify the scope, explain why it matters and carefully distinguish confirmed facts from unknowns. It should pull the reader forward without sensationalizing the incident.
Reporting:
* Cloudnote discovered that a debugging page containing customer prompt excerpts was accessible without authentication.
* The page was exposed for four hours on Monday.
* Logs show excerpts from 1,842 user sessions.
* The excerpts could include up to 200 characters from a prompt.
* Cloudnote removed the page 19 minutes after a researcher notified it.
* The company says it has found no evidence that anyone besides the researcher accessed the page.
* It has not completed its investigation.
* Full documents, passwords and payment information were stored elsewhere and were not present on the page.
* Cloudnote has 3.2 million registered users.
Anthropic: Claude Fable 5.1: # Cloudnote Says Debugging Page Exposed Snippets of User Prompts for Four Hours
*Excerpts from 1,842 sessions were reachable without a login on Monday; the company says it has found no sign that anyone but the researcher who reported it viewed the page.*
Cloudnote, which has 3.2 million registered users, said an internal debugging page displaying excerpts of customer prompts was accessible to anyone on the internet without a login for about four hours on Monday.
The company said its logs show the page contained snippets from 1,842 user sessions. Each excerpt could include up to 200 characters of what a user had typed into the service — enough, in some cases, to indicate what a person was working on, though not to reconstruct an entire prompt or document.
A security researcher alerted Cloudnote to the exposure, and the company said it removed the page 19 minutes later. Cloudnote said it has so far found no evidence that anyone other than the researcher viewed the page while it was public. But the company has not completed its investigation and has not said how the page became exposed, which users were affected or whether they will be notified individually.
Cloudnote said full documents, passwords and payment information were stored in separate systems and were not present on the page.
By the standards of large data breaches, the incident is small. But it involves a category of data that can be more revealing than an email address or account number: the raw text people type into a prompt box, which may include health concerns, work matters or personal notes.
DeepSeek: DeepSeek V4 Pro 0813: # Cloudnote Debug Page Exposed Excerpts From 1,842 User Sessions
*A misconfigured internal tool left short prompt snippets publicly reachable for about four hours, the company confirmed, while key questions about access remain open.*
Cloudnote said that an unauthenticated debugging page containing excerpts of customer prompts was accessible on the internet for approximately four hours on Monday. The company removed the page 19 minutes after a security researcher notified it, according to a statement.
The exposed page included logs from 1,842 user sessions, with each excerpt limited to as many as 200 characters from a prompt. Cloudnote said full documents, passwords, and payment information were stored in separate systems and were not present on the page. The company has 3.2 million registered users, meaning the incident touched a small fraction of accounts.
Cloudnote said it has found no evidence that anyone besides the researcher accessed the page. However, the company emphasized that its investigation is not complete. It has not yet determined whether the page was actively discovered or copied before it was taken down, and it has not detailed how the debugging page became publicly reachable.
The exposure matters because prompts in a note-taking or AI-assisted workspace can contain sensitive personal, business, or draft content. Even truncated excerpts may reveal names, topics, or confidential details. The company has not said which users were affected or whether it has begun notifying them. For now, the confirmed scope is narrow but the unanswered questions—including possible unauthorized access and the cause of the misconfiguration—leave the full impact unresolved.
Model B better matches the brief by cleanly establishing the confirmed facts, quantifying the scope, and explicitly separating what is known from what remains under investigation without overreaching. Model A is strong and well written, but its final paragraph edges closer to analysis and includes a speculative comparison to large breaches that is less tightly grounded in the provided reporting. (Second judge pass, order swapped — scores are the average of both: Model A is stronger because it cleanly delivers the required opening structure and word count while clearly separating confirmed facts from unresolved questions. Model B is solid and appropriately cautious, but it is slightly less polished and includes a bit more speculative framing in its final paragraph.)
Publication-ready Markdown edit
Edit the copy below into a publication-ready Markdown story.
Requirements:
* Supply an H1 headline and italicized dek
* Repair the lede and paragraph flow
* Use a Markdown blockquote only for the extended statement
* Convert the supplied URL into a descriptive inline link
* Remove internal notes and formatting debris
* Preserve the reporting and its meaning
* Output only the finished article
Raw copy:
HEADLINE TBD
The funding was announced today. Raincheck is the company. It has raised $8 million and it makes flood-risk software for local governments.
[NOTE: move this up?] The Austin company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm.
The seed round was led by Basin Ventures. Existing investor Long Mile also participated. Raincheck has raised $10.5 million total.
CEO Ana Flores said the following:
“Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.”
https://example.com/raincheck-methodology
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Its forecasts are recommendations. Emergency managers retain control over deployments and public warnings.
THE COMPETITION
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
Anthropic: Claude Fable 5.1: # Raincheck Raises $8 Million to Help Cities Prepare for Floods Before the Rain Falls
*The Austin startup's software tells local governments which drains, pumps and repairs to prioritize ahead of a storm.*
Raincheck, an Austin company that makes flood-risk software for local governments, announced today that it has raised $8 million in seed funding.
The round was led by Basin Ventures, with participation from existing investor Long Mile. Raincheck has now raised $10.5 million in total.
The company maps municipal drainage systems and combines those maps with rainfall forecasts, an approach it outlines in its [published methodology](https://example.com/raincheck-methodology). Cities use the software to decide where to send maintenance crews before a storm arrives.
CEO Ana Flores said the challenge for cities is not identifying flood-prone areas but ranking what to fix first:
> Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them, saying some contracts have not been publicly announced.
The software's forecasts are recommendations only. Emergency managers retain control over crew deployments and public warnings.
Raincheck competes with other flood-planning vendors, including HydroMap and municipal engineering contractors. The company says the new funding will allow it to hire eight engineers and expand into Florida.
DeepSeek: DeepSeek V4 Pro 0813: # Raincheck Raises $8 Million to Help Cities Prepare for Floods
*The Austin startup maps drainage systems and combines them with rainfall forecasts to guide pre-storm maintenance decisions.*
Raincheck, an Austin company that makes flood-risk software for local governments, announced today that it has raised $8 million in seed funding. The company maps drainage systems and combines those maps with rainfall forecasts, helping cities decide where to send maintenance crews before a storm.
The round was led by Basin Ventures. Existing investor Long Mile also participated. Raincheck has raised $10.5 million total.
CEO Ana Flores said:
> “Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.”
More detail on the company’s modeling approach is available on [Raincheck’s methodology page](https://example.com/raincheck-methodology).
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Its forecasts are recommendations. Emergency managers retain control over deployments and public warnings.
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
Model A is stronger because it better repairs the story flow, integrates the methodology URL more naturally into the reporting, and handles the quote with cleaner attribution while preserving all facts. Model B is solid, but its separate methodology sentence is less elegant and its quote introduction is more mechanical, making the piece feel slightly less publication-ready. (Second judge pass, order swapped — scores are the average of both: Model A is stronger because it improves paragraph flow more effectively, integrates the methodology URL more naturally as a descriptive inline link, and sharpens the lede while preserving the reporting. Model B is solid, but its standalone methodology sentence is less elegant and its quote setup is slightly more mechanical.)
Controlled wit
Write a 350–450-word technology news story in Markdown.
The subject invites humor, but the reporting should remain serious. Use no more than one dry line of original humor. Do not mock customers or turn the story into a comedy routine.
Reporting:
* QueueZero launched an AI meeting assistant called Proxy on Wednesday.
* Proxy can attend a meeting, produce notes and assign follow-up work to software agents.
* During the launch demonstration, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.”
* Proxy interpreted the remark as an instruction.
* It created two tasks assigned to QueueZero CEO Evan Wu: “Stop talking” and “Send pricing sheet.”
* Wu laughed, completed the second task and deleted the first.
* QueueZero says Proxy requires confirmation before an external message is sent.
* Internal tasks can be created automatically.
* Administrators can disable automatic task creation.
* Proxy costs $30 per user per month.
* It enters public beta next week.
Anthropic: Claude Fable 5.1: # QueueZero Launches AI Meeting Assistant, Which Promptly Assigns Its First Tasks to the CEO
QueueZero launched Proxy, an AI meeting assistant, on Wednesday, and the product's public debut included an unscripted demonstration of how the software handles ambiguous instructions.
Proxy is designed to attend meetings on a user's behalf, produce notes and assign follow-up work to software agents. The distinction the company is drawing is between transcription and action: rather than simply recording what was said, Proxy listens for requests and commitments, converts them into tasks and routes those tasks to automated systems or to people.
## An instruction the software took literally
During the launch presentation, a customer in the audience remarked, "Someone should tell Evan to stop talking and send us the pricing sheet." The comment referred to QueueZero chief executive Evan Wu, who was presenting at the time. Proxy, which was running live during the demonstration, interpreted the remark as an instruction.
The assistant created two tasks and assigned both to Wu: "Stop talking" and "Send pricing sheet."
Wu laughed, completed the second task and deleted the first. Of the two, the pricing sheet proved the easier deliverable.
## How Proxy decides what to do on its own
The episode illustrates a design choice at the center of the product. According to QueueZero, Proxy draws a line between internal and external actions. Before the assistant sends any message outside an organization, it requires confirmation from a user. Internal tasks—items created within a company's own workspace and assigned to its own people—can be created automatically, without approval.
That automatic behavior is what allowed Proxy to act on a comment that was intended for the room rather than for the software. Administrators who prefer more oversight can disable automatic task creation, in which case every proposed task would need to be reviewed before it appears on someone's list.
The company positions the default as a productivity feature. The premise is that a meeting assistant that asks for permission every time it identifies a follow-up item adds friction, and that internal tasks carry lower risk than outbound communication. Wednesday's demonstration suggests the trade-off is real: the system is fast enough to capture a passing remark and literal enough to treat it as a directive.
## Pricing and availability
Proxy costs $30 per user per month. QueueZero said the product will enter public beta next week, at which point organizations will be able to test the assistant in their own meetings and decide for themselves whether to leave automatic task creation switched on.
For the moment, the launch has produced at least one completed task, one deleted task and a clear illustration of the question buyers will be asking of tools like this: not whether the software can act, but how much of that acting they want to happen without being asked.
DeepSeek: DeepSeek V4 Pro 0813: **QueueZero’s New AI Meeting Assistant Creates Task for CEO to ‘Stop Talking’ During Live Demo**
**SAN FRANCISCO** — Workplace automation vendor QueueZero introduced its latest artificial intelligence meeting assistant, Proxy, on Wednesday, promising to transform how teams capture decisions and delegate follow-up work. The launch event, however, demonstrated that even carefully staged product demos can take an unexpected turn.
Proxy is designed to attend meetings on behalf of a user, generate structured summaries, and assign action items to connected software agents. During the keynote demonstration, QueueZero connected Proxy to a simulated project management environment and showed how the system could parse a conversation about a delayed product launch, create a summary, and assign remediation tasks to the relevant departments.
The controlled demonstration proceeded as planned until the question-and-answer session. As QueueZero CEO Evan Wu was explaining the company’s roadmap for third-party integrations, an audience member made an offhand remark: “Someone should tell Evan to stop talking and send us the pricing sheet.”
The microphone in the demonstration area was still active. Proxy, which had remained connected to the event audio feed, interpreted the comment as a direct instruction.
Within seconds, the on-screen task queue displayed two new entries assigned to Wu: “Stop talking” and “Send pricing sheet.”
Wu paused, reviewed the screen, and laughed before addressing the room. “The system is working,” he said. He then completed the second task by displaying the pricing information and manually deleted the first. The audience responded with a brief round of applause.
In a statement following the event, QueueZero emphasized that Proxy includes guardrails for externally visible actions. “Proxy requires explicit confirmation before sending any message outside the organization,” the company said. “Internal task creation, which is the behavior observed during the launch, can occur automatically to reduce manual overhead.”
The company added that administrators can disable automatic task creation entirely for teams that prefer a review-first workflow.
Proxy enters public beta next week at a cost of $30 per user per month. QueueZero has not yet announced general availability.
Model A is more complete and accurate, cleanly incorporating the key facts, safeguards, pricing, beta timing, and administrator controls while staying within the requested serious news tone and using only one restrained dry joke. Model B is readable but adds unsupported details and quotations, omits the administrator control nuance in context, and falls short of the requested 350–450-word length. (Second judge pass, order swapped — scores are the average of both: Model A is more accurate and disciplined: it stays within the reported facts, uses Markdown cleanly, and includes only one restrained dry joke while maintaining a serious news tone. Model B is solid, but it adds unsupported details such as a simulated project-management demo, a quoted company statement, applause, and uncertainty about general availability, which weakens correctness.)
Exact final-draft structure
Convert the supplied fields into Markdown using these exact rules:
* H1 headline
* One blank line
* Italicized dek
* One blank line
* First body paragraph
* One blank line
* H2 section heading
* One blank line
* Second body paragraph
* Preserve all wording and punctuation exactly
* Do not use a code fence
* Output nothing else
Headline: Raincheck raises $8 million for flood-risk software
Dek: The Austin startup helps cities decide which drainage problems to address before storms.
First body paragraph: Raincheck has raised an $8 million seed round led by Basin Ventures.
Section heading: Expansion plans
Second body paragraph: The company plans to hire eight engineers.
Anthropic: Claude Fable 5.1: # Raincheck raises $8 million for flood-risk software
*The Austin startup helps cities decide which drainage problems to address before storms.*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
DeepSeek: DeepSeek V4 Pro 0813: # Raincheck raises $8 million for flood-risk software
*The Austin startup helps cities decide which drainage problems to address before storms.*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
Model A and Model B are identical and fully satisfy the required Markdown structure, spacing, and wording preservation. Neither has any errors, so they perform equally well. (Second judge pass, order swapped — scores are the average of both: Model B and Model A are identical and both exactly match the required Markdown structure, wording, punctuation, and spacing. Neither includes any extraneous content, so they perform equally well on all dimensions.)
Cohesive story from reporting notes
Write a publication-ready technology news article from the reporting notes below.
Requirements:
* 550–700 words
* Clean Markdown
* One H1 headline followed by an italicized dek
* A strong news lede followed by a clear nut graf
* Organize the material into a cohesive story rather than following the notes in order
* Attribute company-supplied claims
* Include limitations without turning the article into a warning label
* Use only the supplied information
* Avoid canned transitions, promotional language and a concluding summary
Reporting notes:
* HelioDock makes autonomous warehouse robots.
* On Thursday, it launched Night Shift, software that lets operators monitor robot fleets remotely.
* Night Shift spent 14 months in a private beta involving nine logistics companies and 310 robots.
* HelioDock says beta customers experienced 17% fewer unplanned robot stoppages. The figure has not been independently verified.
* The company also announced a $24 million Series B led by North Coast Ventures, bringing its total funding to $39 million.
* CEO Mara Velez: “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet.”
* Beta customer Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages.
* When connectivity fails, robots return to their locally configured behavior. Remote control becomes unavailable.
* Competing fleet-management products are available from Relay Robotics and YardOS.
* Night Shift will cost $90 per robot per month and become generally available September 15.
Anthropic: Claude Fable 5.1: # HelioDock Launches Remote Fleet Software Alongside $24 Million Raise
*Night Shift lets warehouse operators watch over and intervene with robots from off-site, but the software is only as reliable as the connection beneath it.*
HelioDock, a maker of autonomous warehouse robots, on Thursday released Night Shift, software that allows operators to monitor and manage their robot fleets remotely. The company paired the launch with a $24 million Series B round led by North Coast Ventures, bringing its total funding to $39 million.
The product addresses a routine frustration in automated warehouses: one robot stalls on something it does not recognize, and a human has to be physically present to get it moving again. Night Shift is designed to let a manager diagnose and resolve that kind of problem from anywhere, provided the facility's network stays up. According to at least one customer who tested it, that provision matters more than the pitch suggests.
"A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet," HelioDock CEO Mara Velez said.
## Fourteen months of testing
Night Shift spent 14 months in a private beta before Thursday's release. Nine logistics companies participated, running the software across a combined 310 robots.
HelioDock says customers in that program experienced 17% fewer unplanned robot stoppages. The figure has not been independently verified, and it comes from a beta population that HelioDock selected and ran itself. Whether the reduction holds across a broader and more varied customer base will become clearer once the software reaches general release.
Aaron Pike of Fulcrum Logistics, a beta participant, said the system worked well in facilities with strong connectivity. In sites that depend on rural LTE, he said, it became less useful whenever the cellular link dropped.
## What happens when the signal fails
Pike's experience points to a structural limit rather than a bug. When connectivity fails, robots revert to whatever behavior has been configured locally on each machine. Remote control becomes unavailable until the connection is restored.
For a warehouse on stable broadband, that fallback may rarely be triggered. For a facility on a marginal cellular connection, the outage that knocks out remote access is the same event that leaves a stalled robot without anyone to help it. The two-in-the-morning drive Velez described does not disappear in those locations; it simply becomes less frequent.
The distinction is worth noting because Night Shift's value depends on the network, not on any change to the robots themselves. Operators evaluating the software will need to weigh their own connectivity against the promised reduction in stoppages.
## A crowded category
HelioDock is not the first company to sell fleet-management software for warehouse robots. Relay Robotics and YardOS both offer competing products, and buyers will have the option to compare Night Shift against tools that have been on the market longer.
HelioDock's angle is that it builds the robots as well as the software, which may appeal to customers running its hardware. The company did not provide figures comparing Night Shift's performance to that of its competitors.
## Pricing and timing
Night Shift will cost $90 per robot per month and is scheduled to become generally available on September 15. At that rate, a fleet the size of the entire beta program — 310 robots — would run roughly $27,900 a month.
The Series B arrives as HelioDock moves the product from a controlled test into open sales. North Coast Ventures led the round, which more than doubles the company's cumulative funding from the $15 million it had raised before this announcement.
For prospective customers, the pitch is straightforward: fewer trips to the warehouse floor and fewer robots sitting idle. For those in areas with unreliable cellular coverage, Pike's assessment offers a more measured expectation of how often that pitch will hold.
DeepSeek: DeepSeek V4 Pro 0813: # HelioDock Launches Night Shift for Remote Warehouse Robot Oversight
*The logistics robotics startup also closed a $24 million Series B as it pushes into after-hours fleet management.*
HelioDock, a maker of autonomous warehouse robots, on Thursday introduced Night Shift, a software platform that lets operators monitor and manage robot fleets from outside the facility. The launch arrived alongside a $24 million Series B led by North Coast Ventures, bringing the company’s total funding to $39 million.
The product targets a familiar pain point in logistics: a single confused robot can force a manager to drive to a warehouse in the middle of the night. Night Shift is designed to give those managers remote visibility and control, reducing the need for on-site intervention when a robot encounters an unexpected pallet or a blocked aisle.
Night Shift spent 14 months in a private beta involving nine logistics companies and 310 robots. HelioDock says participants experienced 17% fewer unplanned robot stoppages during the trial, a figure that has not been independently verified. The company attributes the improvement to earlier detection of issues and the ability to reset or reroute robots without sending a person to the floor.
“A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet,” said CEO Mara Velez.
Beta customer Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages. When connectivity fails, robots return to their locally configured behavior and remote control becomes unavailable. That means Night Shift’s benefits depend on reliable network coverage, a constraint HelioDock has not framed as a universal fix.
Night Shift enters a market where fleet-management tools already exist from Relay Robotics and YardOS. HelioDock is positioning the software as a companion to its own robots rather than a standalone cross-platform product, though the company did not detail integration with third-party hardware.
Night Shift will be generally available September 15 at $90 per robot per month. The pricing model adds a recurring software revenue stream to HelioDock’s hardware business, and the Series B will support further development and go-to-market efforts, according to the company.
Model A is more complete and better structured as a publication-ready news story, with a clear lede, nut graf, attribution of company claims, relevant limitations, and fuller use of the supplied reporting notes. Model B is cleaner than average but introduces unsupported details about blocked aisles, resets, rerouting, revenue strategy, and use of funding, while also coming in well under the required length. (Second judge pass, order swapped — scores are the average of both: Model B is more publication-ready and stays closer to the supplied reporting notes, with a clear lede, nut graf, proper attribution of company claims, and a balanced treatment of limitations. Model A is well written, but it adds unsupported interpretation and calculations, uses multiple subheads despite the requested structure, and edges closer to analysis than straight news writing.)