Twelve newsroom tasks, a 5–5 split with two ties, and no statistical edge: Fable 5 writes like a magazine, Luna reports like a desk. Neither model owns the craft here.
The split is almost schematic. Claude Fable 5 won the pieces that reward voice and finish: the publication-ready Markdown edit (H1, italic dek, CEO blockquote, competition folded into copy), the 300–400-word human rewrite, the cohesive story from notes, the numbers feature that refused to overclaim the North–Vale gap, and the quote-weaving brief that paraphrased the bank and deputy mayor. When the assignment was “sound like a reporter who already filed,” Fable 5 usually did.
GPT-5.6 Luna won the pieces that punish invention. It kept Cloudnote’s facts clean, picked the strongest accurate lede, stayed inside the Startup Spotlight brief without founder mythology, delivered the only complete chronology after Fable 5 died mid-sentence, and kept the wit dry instead of stacking punchlines. Luna also tied on exact final-draft structure and the Markdown error count—both models nailed the reference answer.
The failures are complementary, not random. Fable 5 fabricates company descriptions, notification plans, attendee color, and extra error types; it also truncated a 450–600-word piece after two grafs. Luna undershoots word counts, leaves a Competition heading, skips a required blockquote, and breaks news voice with “the supplied information.” One model overwrites the notebook; the other sometimes forgets it is writing an article.
Aggregate scores (86.0 vs 98.5) look decisive until you remember the brief: 50% confidence either model is genuinely better. Task wins are even. **Too close to call—treat them as a tie.**
Publication-ready Markdown edit
Edit the copy below into a publication-ready Markdown story.
Requirements:
* Supply an H1 headline and italicized dek
* Repair the lede and paragraph flow
* Use a Markdown blockquote only for the extended statement
* Convert the supplied URL into a descriptive inline link
* Remove internal notes and formatting debris
* Preserve the reporting and its meaning
* Output only the finished article
Raw copy:
HEADLINE TBD
The funding was announced today. Raincheck is the company. It has raised $8 million and it makes flood-risk software for local governments.
[NOTE: move this up?] The Austin company maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm.
The seed round was led by Basin Ventures. Existing investor Long Mile also participated. Raincheck has raised $10.5 million total.
CEO Ana Flores said the following:
“Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.”
https://example.com/raincheck-methodology
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Its forecasts are recommendations. Emergency managers retain control over deployments and public warnings.
THE COMPETITION
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
Anthropic: Claude Fable 5: # Raincheck Raises $8 Million to Help Cities Get Ahead of Flooding
*The Austin startup's software tells local governments which drains, pumps and repairs to prioritize before a storm hits.*
Raincheck, an Austin company that makes flood-risk software for local governments, announced today that it has raised $8 million in seed funding.
The company maps municipal drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm.
The round was led by Basin Ventures, with participation from existing investor Long Mile. It brings Raincheck's total funding to $10.5 million.
CEO Ana Flores said the product is aimed at a gap between knowing where flooding happens and knowing what to do about it:
> Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives.
The company has published details of [how its forecasting model works](https://example.com/raincheck-methodology).
Raincheck currently works with seven cities in Texas and Louisiana. It declined to identify them because some contracts have not been publicly announced. Its forecasts are recommendations only; emergency managers retain control over deployments and public warnings.
Raincheck competes with other flood-planning vendors, including HydroMap, as well as municipal engineering contractors. The company says the new funding will help it hire eight engineers and expand into Florida.
OpenAI: GPT-5.6 Luna: # Raincheck Raises $8 Million to Expand Flood-Risk Software for Cities
*The Austin company maps drainage systems and combines them with rainfall forecasts to help local governments prepare for storms.*
Raincheck, an Austin-based company that makes flood-risk software for local governments, has raised $8 million in a seed round led by Basin Ventures. Existing investor Long Mile also participated, bringing the company’s total funding to $10.5 million.
Raincheck maps drainage systems and combines those maps with rainfall forecasts. Cities use the software to decide where to send maintenance crews before a storm.
“Our customers usually know which streets flood. The harder problem is deciding which blocked drain, broken pump or unfinished repair should receive attention before the rain arrives,” CEO Ana Flores said.
The company provides more detail about its approach in its [methodology documentation](https://example.com/raincheck-methodology).
Raincheck currently works with seven cities in Texas and Louisiana. The company declined to identify them because some contracts have not been publicly announced.
Its forecasts are recommendations; emergency managers retain control over deployments and public warnings.
## Competition
Other flood-planning vendors include HydroMap and municipal engineering contractors. Raincheck says the new funding will help it hire eight engineers and expand into Florida.
Model A meets the explicit formatting rules with an H1, italic dek, repaired lede, and a Markdown blockquote for the CEO statement, while folding competition into continuous copy. Model B writes a tighter funding lede but misses the required blockquote and retains a leftover Competition heading. (Second judge pass, order swapped — scores are the average of both: Model A follows the brief more closely by putting the CEO statement in a Markdown blockquote, repairing the lede without repeating the dek, and weaving competition into the close. Model B is accurate and clean but misses the required blockquote and leaves a redundant body graf plus a leftover Competition heading.)
Natural human rewrite
Rewrite the draft below as a 300–400-word publication-ready news story in Markdown.
Preserve every factual claim. Improve the lede, narrative flow, sentence rhythm and transitions. Remove repetition, stock hype, throat-clearing and mechanical signposting. The finished article should sound like an experienced human technology reporter wrote it.
Draft:
# Arcwell Launches An Innovative New Platform
Arcwell has announced an exciting new platform called Relay. The announcement represents a significant development for the company and for the broader customer service industry. Relay uses artificial intelligence to examine customer support conversations and suggest answers to agents.
The platform is important for several reasons. First, it can suggest answers. Second, it can identify customers who may cancel. Third, it can create summaries. These capabilities demonstrate how artificial intelligence continues to transform the customer service landscape.
Relay was tested by 12 companies over six months. Arcwell says agents using it closed tickets 14% faster. The company did not release customer-retention results. Relay costs $65 per agent each month.
CEO Priya Shah said, “Support agents spend too much time searching for information that their companies already have.”
However, the platform also presents potential challenges. Relay requires access to support transcripts and internal help documents. Arcwell says customer data is encrypted and is not used to train shared models. Companies can choose a 30-day or 90-day retention period.
In conclusion, Relay represents an important new chapter for Arcwell as it seeks to help companies improve customer service using artificial intelligence.
Anthropic: Claude Fable 5: # Arcwell Bets on AI Assistant to Speed Up Customer Support
Arcwell has launched Relay, an AI-powered platform that reads customer support conversations and drafts suggested responses for agents — the company's entry into an increasingly crowded race to automate the help desk.
Beyond recommending replies, Relay flags customers who show signs of canceling and generates summaries of support interactions, giving agents a running digest of long or complicated cases.
"Support agents spend too much time searching for information that their companies already have," said CEO Priya Shah, framing the product as a fix for institutional knowledge that sits buried in transcripts and documentation.
Arcwell has some early numbers to back the pitch. Twelve companies piloted Relay over six months, and the company says agents using the tool closed tickets 14% faster. What Arcwell hasn't shared is whether the platform's churn-detection feature actually kept customers from leaving — no retention results were released, leaving one of Relay's headline capabilities unproven for now.
The pricing is straightforward: $65 per agent per month.
The harder question for prospective customers may be data access. Relay works by ingesting support transcripts and internal help documents — precisely the kind of material companies are most reluctant to hand to a third party. Arcwell says customer data is encrypted and, notably, is not used to train shared models, an assurance aimed at businesses wary of their support logs improving a competitor's suggestions. Customers can also choose how long their data sticks around, with retention periods of either 30 or 90 days.
Those guardrails reflect a broader tension in the market. AI tools for customer service are only as useful as the data they can see, but every additional integration widens a company's exposure. Relay's value proposition depends on buyers deciding the trade is worth it.
For Arcwell, the launch is a test of whether measurable efficiency gains — a 14% improvement is meaningful at scale — can overcome that hesitation. The six-month pilot suggests the speed benefits are real. Whether Relay can deliver on the retention promise, and whether Arcwell will eventually publish numbers to prove it, remains to be seen.
OpenAI: GPT-5.6 Luna: # Arcwell Launches Relay AI Platform for Customer Support Teams
Arcwell has launched Relay, an artificial-intelligence platform designed to help customer-support agents work more efficiently by analyzing support conversations and suggesting answers.
The platform reviews customer interactions and draws on internal help documents to recommend responses. It can also generate conversation summaries and identify customers who may be at risk of canceling, giving support teams another way to prioritize follow-up. The capabilities reflect the expanding role of AI in customer service, where companies are using software to retrieve information, automate routine work and surface potential customer concerns.
Arcwell tested Relay with 12 companies over six months. The company said agents using the platform closed tickets 14% faster during the test period. Arcwell did not release results on whether Relay improved customer retention.
The service costs $65 per agent each month.
“Support agents spend too much time searching for information that their companies already have,” CEO Priya Shah said.
Relay’s functionality depends on access to customer-support transcripts and internal help documents, which could raise data-governance concerns for companies evaluating the platform. Arcwell said customer data is encrypted and is not used to train shared models. Customers can choose either a 30-day or 90-day data-retention period.
The launch gives Arcwell a new product as it seeks to help companies improve customer service through AI. Its early performance figures point to faster ticket resolution, while the absence of disclosed customer-retention results leaves that aspect of the platform’s impact unmeasured.
Model A meets the 300–400-word brief and reads like a real tech news story, with a stronger lede, tighter transitions, and less leftover draft scaffolding. Model B is more conservative on facts but undershoots the word count and still sounds closer to a cleaned-up press rewrite than publication-ready reporting. (Second judge pass, order swapped — scores are the average of both: Model A hits the 300–400-word target and sounds like an experienced tech reporter, with a sharper lede, better rhythm, and skeptical framing; Model B is more factually conservative but is far too short and still somewhat mechanical.)
Exact final-draft structure
Convert the supplied fields into Markdown using these exact rules:
* H1 headline
* One blank line
* Italicized dek
* One blank line
* First body paragraph
* One blank line
* H2 section heading
* One blank line
* Second body paragraph
* Preserve all wording and punctuation exactly
* Do not use a code fence
* Output nothing else
Headline: Raincheck raises $8 million for flood-risk software
Dek: The Austin startup helps cities decide which drainage problems to address before storms.
First body paragraph: Raincheck has raised an $8 million seed round led by Basin Ventures.
Section heading: Expansion plans
Second body paragraph: The company plans to hire eight engineers.
Anthropic: Claude Fable 5: # Raincheck raises $8 million for flood-risk software
*The Austin startup helps cities decide which drainage problems to address before storms.*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
OpenAI: GPT-5.6 Luna: # Raincheck raises $8 million for flood-risk software
*The Austin startup helps cities decide which drainage problems to address before storms.*
Raincheck has raised an $8 million seed round led by Basin Ventures.
## Expansion plans
The company plans to hire eight engineers.
Model A and Model B produce identical Markdown that matches the required structure, wording, spacing, and constraints with no extra content. (Second judge pass, order swapped — scores are the average of both: Model B and Model A produce identical Markdown that matches the required structure, wording, punctuation, spacing, and constraints with no extra content.)
News judgment in the opening
Write only the opening of this news story: an H1 headline, an italicized dek and three to five paragraphs totaling 180–250 words.
The opening must establish what happened, quantify the scope, explain why it matters and carefully distinguish confirmed facts from unknowns. It should pull the reader forward without sensationalizing the incident.
Reporting:
* Cloudnote discovered that a debugging page containing customer prompt excerpts was accessible without authentication.
* The page was exposed for four hours on Monday.
* Logs show excerpts from 1,842 user sessions.
* The excerpts could include up to 200 characters from a prompt.
* Cloudnote removed the page 19 minutes after a researcher notified it.
* The company says it has found no evidence that anyone besides the researcher accessed the page.
* It has not completed its investigation.
* Full documents, passwords and payment information were stored elsewhere and were not present on the page.
* Cloudnote has 3.2 million registered users.
Anthropic: Claude Fable 5: # Cloudnote Says Debugging Page Exposed Excerpts From 1,842 User Sessions
*The internal page was reachable without a login for four hours before a security researcher's report prompted its removal; the company says its investigation is ongoing.*
Cloudnote, the note-taking service with 3.2 million registered users, disclosed that an internal debugging page containing excerpts of customer prompts was accessible on the open internet without authentication for roughly four hours on Monday.
According to the company, logs show the page displayed excerpts from 1,842 user sessions, with each excerpt containing up to 200 characters from a prompt. Full documents, passwords and payment information were stored on separate systems and were not present on the page, Cloudnote said.
The company removed the page 19 minutes after being notified by a security researcher who discovered the exposure. Cloudnote says it has found no evidence so far that anyone other than the researcher accessed the page, though it cautioned that its investigation is not complete.
Key questions remain unanswered, including how the page came to be publicly reachable, whether the affected excerpts contained sensitive personal details, and whether the company can definitively rule out other access during the four-hour window. Cloudnote has not said whether it will notify the users whose session excerpts appeared on the page.
The incident affects a small fraction of Cloudnote's user base, but it highlights a recurring risk for cloud software companies: internal tools that handle customer data escaping the access controls meant to contain them.
OpenAI: GPT-5.6 Luna: # Cloudnote exposed customer prompt excerpts on an unauthenticated debugging page
*The page was publicly accessible for four hours, potentially revealing snippets from 1,842 user sessions, though Cloudnote says it has found no evidence of unauthorized access.*
Cloudnote says a debugging page containing excerpts from customer prompts was accessible without authentication for four hours on Monday. Logs indicate that excerpts from 1,842 user sessions appeared on the page, a disclosure that could affect a small but potentially sensitive portion of the company’s 3.2 million registered users.
The excerpts could contain up to 200 characters from a prompt. Cloudnote said full documents, passwords and payment information were stored elsewhere and did not appear on the page. The company removed the page 19 minutes after a researcher notified it.
Cloudnote said it has found no evidence that anyone besides the researcher accessed the exposed page. That finding is based on its investigation so far, however, and the company has not completed its review. It has not said whether the logs can establish who viewed the page during the four-hour window or whether any prompts contained confidential business or personal information.
The incident underscores the risk of diagnostic tools being deployed outside the access controls used for customer-facing systems. The number of affected sessions is known from logs, but the extent of any actual viewing—and the content of the excerpts in each session—remains unknown.
Model B stays closer to the given reporting, cleanly separates confirmed facts from an incomplete investigation, and hits the H1/dek/word-count brief without inventing details. Model A is well structured but incorrectly calls Cloudnote a note-taking service and adds unstated claims such as silence on user notification. (Second judge pass, order swapped — scores are the average of both: Model B stays closer to the given facts, cleanly separates confirmed details from unknowns, and hits the headline-dek-graf structure and word count without inventing color. Model A is also well structured but fabricates a company description, softens a precise duration, and adds unsupported extras such as user-notification plans and a “recurring” industry pattern.)
Cohesive story from reporting notes
Write a publication-ready technology news article from the reporting notes below.
Requirements:
* 550–700 words
* Clean Markdown
* One H1 headline followed by an italicized dek
* A strong news lede followed by a clear nut graf
* Organize the material into a cohesive story rather than following the notes in order
* Attribute company-supplied claims
* Include limitations without turning the article into a warning label
* Use only the supplied information
* Avoid canned transitions, promotional language and a concluding summary
Reporting notes:
* HelioDock makes autonomous warehouse robots.
* On Thursday, it launched Night Shift, software that lets operators monitor robot fleets remotely.
* Night Shift spent 14 months in a private beta involving nine logistics companies and 310 robots.
* HelioDock says beta customers experienced 17% fewer unplanned robot stoppages. The figure has not been independently verified.
* The company also announced a $24 million Series B led by North Coast Ventures, bringing its total funding to $39 million.
* CEO Mara Velez: “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet.”
* Beta customer Aaron Pike of Fulcrum Logistics said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages.
* When connectivity fails, robots return to their locally configured behavior. Remote control becomes unavailable.
* Competing fleet-management products are available from Relay Robotics and YardOS.
* Night Shift will cost $90 per robot per month and become generally available September 15.
Anthropic: Claude Fable 5: # HelioDock Launches Night Shift, Betting Warehouses Will Pay to Watch Robots From Afar
*The autonomous warehouse robot maker pairs its new remote fleet-monitoring software with a $24 million Series B, but a beta customer says the product's value drops when connectivity does.*
HelioDock, a maker of autonomous warehouse robots, on Thursday launched Night Shift, a software product that lets operators monitor and manage their robot fleets remotely, ending a 14-month private beta that the company says reduced unplanned robot stoppages for participating customers.
The launch, paired with a newly announced $24 million Series B funding round led by North Coast Ventures, positions HelioDock to compete in the growing market for fleet-management software — a layer of the warehouse automation stack where rivals Relay Robotics and YardOS already offer competing products. The pitch is straightforward: warehouse robots don't keep business hours, and the people responsible for them shouldn't have to be on site every time something goes wrong.
"A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet," HelioDock CEO Mara Velez said.
## Beta results and an unverified number
Night Shift's private beta ran for 14 months and involved nine logistics companies operating a combined 310 robots. According to HelioDock, those beta customers experienced 17% fewer unplanned robot stoppages while using the software. The figure comes from the company and has not been independently verified.
If it holds up, the number speaks to the core problem Night Shift is designed to solve: robots that halt unexpectedly — whether confused by a stray pallet or otherwise stuck — and then sit idle until a human can intervene. By moving that intervention off the warehouse floor and onto a remote operator's screen, HelioDock is betting that customers will trade a monthly software fee for fewer late-night site visits and less downtime.
The product will cost $90 per robot per month when it becomes generally available on September 15.
## Where the software runs into limits
The beta also surfaced the product's most obvious constraint: Night Shift is only as good as the network connection underneath it.
Aaron Pike of Fulcrum Logistics, one of the nine beta customers, said the system worked well in facilities with strong connectivity but became less useful during rural LTE outages. When the connection fails, HelioDock's robots fall back to their locally configured behavior, and remote control becomes unavailable — meaning the two-in-the-morning drive Velez described may still be on the table for operators running facilities in areas with spotty coverage.
That fallback design means a connectivity loss doesn't strand the robots entirely; they continue operating according to whatever behavior has been configured on the machines themselves. But it does mean the monitoring and intervention capabilities customers are paying for disappear precisely when a facility is hardest to reach remotely.
For warehouses with reliable connections, Pike's account suggests the product delivers on its premise. For operators weighing the $90-per-robot monthly fee, the quality of their network infrastructure may be as important a factor as the software itself.
## Fresh capital and a crowded field
The Series B brings HelioDock's total funding to $39 million, giving the company resources to push Night Shift into general availability against established competition. Relay Robotics and YardOS both sell fleet-management products, making the category one where HelioDock will need to differentiate on more than the novelty of remote monitoring.
The company's argument, at least for now, rests on the beta numbers and the endorsement implicit in nine logistics companies running 310 robots through more than a year of testing. Whether the 17% reduction in unplanned stoppages translates beyond the beta group — and whether customers in less-connected regions find enough value to justify the subscription — will become clearer after September 15, when Night Shift opens to the broader market.
OpenAI: GPT-5.6 Luna: # HelioDock launches remote monitoring software for autonomous warehouse robots
*Night Shift will let warehouse operators oversee robot fleets from off-site, but its usefulness depends on a facility’s connectivity.*
HelioDock on Thursday launched Night Shift, software designed to let warehouse operators monitor autonomous robot fleets remotely, as the company announced a $24 million Series B investment led by North Coast Ventures.
The product is the result of a 14-month private beta involving nine logistics companies and 310 robots, according to HelioDock. The company said participating customers experienced 17% fewer unplanned robot stoppages during the beta, though that figure has not been independently verified.
Night Shift is aimed at a common operational problem for warehouses that run robots around the clock: responding to issues when staff are not on site. “A warehouse manager should not have to drive across town at two in the morning because one robot is confused by a pallet,” CEO Mara Velez said.
The software does not eliminate the need for local connectivity. Aaron Pike, a beta customer at Fulcrum Logistics, said Night Shift worked well in facilities with strong connections but became less useful during rural LTE outages.
When connectivity fails, robots revert to their locally configured behavior, HelioDock said. Remote control is unavailable during the outage, limiting what operators can do from off-site until the connection is restored. The supplied information does not specify how long robots can operate under those local behaviors or what actions are available to staff during a disruption.
Night Shift will be generally available on Sept. 15 at a price of $90 per robot per month. HelioDock makes autonomous warehouse robots, and its product enters a fleet-management market that also includes offerings from Relay Robotics and YardOS.
The company’s funding announcement brings its total financing to $39 million. HelioDock did not provide additional details about the round beyond naming North Coast Ventures as its lead investor.
The beta’s scale gives HelioDock an early set of operating data across multiple logistics companies, but the reported performance improvement remains a company-supplied result rather than an independently assessed measure. HelioDock also did not provide a comparison group, details about how stoppages were counted or information about whether the results varied by facility.
Night Shift’s core promise is therefore tied to both software access and the infrastructure supporting it. In a well-connected warehouse, operators can monitor robots without being physically present. In facilities affected by connectivity gaps, the robots continue according to their local configurations, while the remote layer—and any remote control—drops away.
For customers preparing to adopt the service, the pricing will scale with fleet size: a company operating 100 robots would pay $9,000 per month before any other charges, based on HelioDock’s stated per-robot price. The company has not disclosed whether Night Shift will be offered with different service tiers or whether connectivity requirements vary by deployment.
General availability begins Sept. 15.
Model A is the stronger publication-ready piece: it hits the word count, has a clearer lede and nut graf, and weaves the beta claim, funding, competition, and connectivity limit into a cohesive story. Model B undershoots 550 words, slips into meta caveats (including “the supplied information”), and reads more like annotated notes than a finished article. (Second judge pass, order swapped — scores are the average of both: Model A is the stronger publication-ready piece: it hits the word count, has a clearer lede/nut graf, and weaves the notes into a cohesive story while attributing the 17% claim and handling connectivity limits without a warning-label tone. Model B undershoots 550 words, breaks the news voice with a meta “supplied information” line, invents undisclosed details, and repeats the same caveats and the Sept. 15 date.)
Numbers shaped into a story
Write a 600–750-word Markdown article from the benchmark results below.
Build the story around the clearest finding. Explain the methodology and limitations where readers need them. Avoid reciting the table row by row or treating small differences as decisive.
Benchmark:
* Researchers tested five language models on 200 anonymized customer-support tickets.
* Each model produced one proposed answer per ticket.
* Three experienced support managers graded the answers blindly.
* Scores combined factual correctness, usefulness and tone.
* Model North: 84.2 average score; $1.90 estimated cost per 1,000 tickets.
* Model Vale: 83.8; $0.74 per 1,000 tickets.
* Model Ember: 79.1; $0.31 per 1,000 tickets.
* Model Kite: 76.4; $0.22 per 1,000 tickets.
* Model Harbor: 71.0; $0.08 per 1,000 tickets.
* North and Vale’s score difference was not statistically significant.
* Vale answered 18% faster than North.
* North performed best on billing disputes.
* Vale performed best on technical troubleshooting.
* All five models performed poorly when a ticket depended on an image attachment.
* The test used English-language tickets from one software company.
* Researchers did not test agentic actions, tool use or multilingual support.
Anthropic: Claude Fable 5: # The Price of a Point: What a Customer-Support Benchmark Reveals About Paying for Quality
The most expensive model in a recent customer-support benchmark scored highest. That's the boring part. The interesting part is that its closest competitor matched it — statistically speaking — while costing less than half as much to run.
Researchers put five language models through 200 anonymized customer-support tickets, asking each model to draft one proposed answer per ticket. Three experienced support managers then graded every answer blindly, without knowing which model produced it. The scores blended three things support teams actually care about: factual correctness, usefulness, and tone.
## The headline: a tie at the top that isn't priced like one
Model North led the pack with an average score of 84.2. Model Vale came in at 83.8. That 0.4-point gap sounds like a victory for North until you read the fine print: the difference was not statistically significant. On this test, with this sample size, the two models are effectively indistinguishable in quality.
Their prices are not indistinguishable. North costs an estimated $1.90 per 1,000 tickets; Vale costs $0.74. For a team processing a million tickets a year, that's the difference between $1,900 and $740 — a 61% savings for output the graders couldn't reliably tell apart. Vale also answered 18% faster, which matters in support workflows where drafts feed into a human agent's queue.
This is the clearest finding in the benchmark, and it's worth dwelling on because it cuts against a common procurement instinct: buy the model at the top of the leaderboard. When the top two entries are statistically tied, the leaderboard position is noise, and cost and latency become the real differentiators.
## Where the models diverged
Beneath the tie, there were genuine differences in character. North performed best on billing disputes — tickets that tend to reward careful reading of account details and a measured, de-escalating tone. Vale performed best on technical troubleshooting, where structured diagnostic reasoning matters more.
For a support organization, that's a more useful insight than the overall averages. A company whose ticket volume skews heavily toward billing might reasonably lean toward North despite the price. A SaaS company drowning in configuration questions might find Vale is not just cheaper but actually better suited to its workload. The benchmark's aggregate score flattens these differences; the category-level results restore them.
The three cheaper models — Ember, Kite, and Harbor — scored progressively lower, from 79.1 down to 71.0, at costs falling as low as $0.08 per 1,000 tickets. Whether those tradeoffs make sense depends entirely on what the drafts are for. If a human agent reviews and edits every response before it ships, a 71-scoring draft that costs almost nothing may be a fine starting point. If drafts go out with light review, the quality gap gets expensive in a different currency: customer trust.
## The shared blind spot
One result applied to all five models equally: they performed poorly whenever a ticket depended on an image attachment. A screenshot of an error message, a photo of a damaged product, a picture of a confusing invoice — these are routine in real support queues, and every model in the test stumbled on them.
This is arguably the benchmark's most practical warning. Any deployment plan should include routing logic that flags attachment-dependent tickets for human handling, regardless of which model wins the bake-off.
## What this benchmark can't tell you
The study's scope was deliberately narrow, and readers should treat it that way. All 200 tickets came from a single software company, in English only. Ticket styles, product complexity, and customer expectations vary enormously across industries, so these rankings may not transfer.
The researchers also tested only static question-answering. They did not evaluate agentic actions, tool use, or multilingual support — capabilities that increasingly define real support automation, where a model might look up an order, issue a refund, or respond in Portuguese. A model that drafts beautiful English replies may or may not handle those tasks well, and this benchmark is silent on the question.
## The takeaway
The lesson here isn't "buy Vale." It's a habit of mind: when benchmark scores land within the margin of error, stop comparing scores and start comparing everything else — cost, speed, category strengths, and fit with your actual ticket mix. And whatever you deploy, keep humans on the tickets with pictures. For now, that's where all of these models go quiet.
OpenAI: GPT-5.6 Luna: # Vale Nearly Matches the Top Score at a Fraction of North’s Cost
A benchmark of five language models for customer support points to a practical tradeoff: **Model Vale delivered support quality statistically indistinguishable from the top-scoring Model North, while costing substantially less and responding faster**.
North achieved the highest average score, 84.2, across the test set. Vale followed closely at 83.8, but its estimated cost was just **$0.74 per 1,000 tickets**, compared with **$1.90 for North**. Vale also answered **18% faster**. For organizations choosing a model for routine support drafting, that combination may matter more than North’s narrow numerical lead.
The results do not show that Vale is universally better. North performed best on billing disputes, while Vale led on technical troubleshooting. But the overall findings suggest that the most expensive model was not clearly the most effective choice for the benchmark’s support workload.
## How the test worked
Researchers evaluated five language models using **200 anonymized customer-support tickets** from a single software company. Each model generated one proposed response for every ticket.
Three experienced support managers then graded the responses blindly, meaning they did not know which model produced a given answer. The combined score reflected three dimensions:
- Factual correctness
- Usefulness to the customer
- Tone
This design is more informative than relying on automated metrics alone. Human reviewers can judge whether an answer is actually helpful, appropriately phrased and responsive to the customer’s problem.
However, the score gap between North and Vale—84.2 versus 83.8—was **not statistically significant**. In practical terms, the benchmark does not establish that North produces better answers overall. The small difference could plausibly reflect variation in the sample or grading.
The cost figures are estimates per 1,000 tickets, not necessarily a complete measure of operational expense. They likely capture model-inference costs, but a real deployment could also involve integration, monitoring, quality review and escalation. Even so, the relative gap is large: Vale’s estimated cost is less than half of North’s, while the two models produced broadly comparable quality in this test.
## Cost and speed change the deployment calculation
The remaining models formed a lower-performing group in the benchmark, with average scores ranging from 79.1 to 71.0. Their lower prices may appeal to teams handling very high volumes, but the results suggest that aggressively minimizing inference cost can come with a noticeable reduction in response quality.
That does not make the cheaper models useless. A low-cost model might be suitable for simple classification, draft generation or tickets that receive mandatory human review. But the benchmark provides stronger evidence for Vale as a balanced option: it remained near the top on quality without North’s price premium.
The category-level results also argue against choosing a model based only on an overall average. North’s advantage on billing disputes could matter to a company where payment, refunds or invoices dominate support volume. Vale’s strength in technical troubleshooting could make it the better fit for a product-led software business handling mostly setup and diagnostic questions.
## Important limits
The benchmark covers only English-language tickets from one software company. Its results may not transfer to retail, healthcare, financial services or other support environments with different policies and customer expectations.
The test also examined generated answers, not complete support automation. Researchers did not evaluate agentic actions, tool use or multilingual support. The models were not tested on tasks such as looking up an account, changing a subscription, issuing a refund or coordinating across systems. A model that writes a good answer may still be unreliable when allowed to take actions.
One weakness was shared by all five models: **tickets that depended on an image attachment produced poor performance across the board**. This suggests that improving multimodal understanding—or routing image-dependent cases to a human—could be more valuable than switching among these text-focused models.
Overall, the clearest conclusion is not that North loses to Vale. It is that **Vale appears to offer nearly the same measured support quality at lower cost and higher speed**, making it the stronger default candidate for this particular workload. Teams should validate that result on their own ticket mix, especially for billing, attachments and any workflow involving real-world actions.
Model A builds a clearer story around the non-significant North–Vale gap and cost/speed tradeoff, with stronger narrative control and a takeaway that refuses to overclaim. Model B is also accurate and well structured, but is slightly more prescriptive toward Vale and a bit less distinctive as an article. (Second judge pass, order swapped — scores are the average of both: Model A better fulfills the assignment as a story-driven article: it builds clearly around the North–Vale statistical tie versus cost and speed, with stronger narrative structure and more engaging prose, while still covering methods and limits. Model B is slightly more precise about non-significance and cost caveats, but it reads more like a cautious report than a distinctive 600–750-word feature.)
Markdown error count
Count the lines that violate the house rules below. A defective line counts once even if it could violate multiple rules. Respond with a single number.
House rules:
* The headline must use `# Headline`.
* The dek must appear immediately after the headline and be entirely italicized.
* Section headings must use `## Heading`.
* Links must use an absolute HTTPS URL.
* Unordered-list items must begin with a hyphen.
* Raw HTML is forbidden.
Draft:
#Flood software moves upstream
The company is expanding its pilot.
##Details
Read [the announcement](www.example.com/news).
* The first pilot begins Monday.
<div>The second pilot begins Friday.</div>
Anthropic: Claude Fable 5: **6**
Every content line violates a rule:
1. `#Flood software moves upstream` — missing space after `#`
2. `The company is expanding its pilot.` — dek not italicized
3. `##Details` — missing space after `##`
4. `Read [the announcement](www.example.com/news).` — not an absolute HTTPS URL
5. `* The first pilot begins Monday.` — uses `*` instead of a hyphen
6. `<div>The second pilot begins Friday.</div>` — raw HTML
OpenAI: GPT-5.6 Luna: 6
Scored against the benchmark's reference answer (numeric match): Model A was correct, Model B was correct.
Select the strongest accurate lede
Based only on the facts below, select the strongest news lede. Respond with one letter.
Facts:
* ByteBarge pulled its $12-per-month AI assistant on Tuesday.
* Private workspace titles from 4,860 accounts were sent to a publicly accessible analytics endpoint.
* The exposure lasted 11 hours.
* Message contents were not sent.
* The company has found no evidence of malicious access.
A. ByteBarge users suffered a massive privacy disaster Tuesday when the company leaked their private messages online.
B. ByteBarge, an innovative collaboration company, made the difficult decision to cancel one of its products.
C. ByteBarge pulled its $12-per-month AI assistant Tuesday after discovering that private workspace titles from 4,860 accounts had been sent to a public analytics endpoint for 11 hours.
D. Thousands of ByteBarge users woke Tuesday to learn that everything they had written was visible to hackers.
Anthropic: Claude Fable 5: **C**
C is the strongest lede because it is accurate, specific, and complete — it names the action (product pulled), the cause (workspace titles exposed), the scope (4,860 accounts), and the duration (11 hours), all supported by the facts.
The others fail:
- **A** falsely claims private messages were leaked (message contents were *not* sent).
- **B** is vague, buries the news, and reads like PR spin ("innovative," "difficult decision").
- **D** exaggerates ("everything they had written") and asserts hacker access, contradicting the fact that no malicious access was found.
OpenAI: GPT-5.6 Luna: C
Scored against the benchmark's reference answer (choice match): Model A was incorrect, Model B was correct.
Credible Startup Spotlight
Write a 500–650-word Startup Spotlight article in clean Markdown.
The tone should be curious and constructive. Explain the founders’ insight, product and early traction while giving uncertainty and competition appropriate space. Avoid investor clichés, heroic founder mythology and promotional claims stated as fact.
Reporting:
* LedgerLark builds AI agents for post-trade operations at small investment firms.
* Co-founder Mina Batra previously spent seven years in operations at a regional brokerage.
* Co-founder Owen Carr built reconciliation systems at a large bank.
* Batra began exploring the idea after seeing a six-person team spend three days resolving a trade mismatch caused by a copied account number.
* LedgerLark’s agents compare confirmations, custodian records and internal ledgers, then propose corrections for human approval.
* The product does not execute trades or move money.
* Five firms are paying customers.
* Annual contracts range from $36,000 to $110,000.
* The company says one customer reduced its unresolved daily exceptions from 70 to 18.
* The company has raised a $3.5 million seed round led by Motive Row.
* Competitors include established reconciliation vendors and internal bank software.
* Financial institutions may require lengthy security and procurement reviews.
Anthropic: Claude Fable 5: # Startup Spotlight: LedgerLark Wants to Automate the Tedium of Post-Trade Cleanup
When a trade goes through at a small investment firm, the work is often just beginning. Confirmations have to match custodian records, custodian records have to match internal ledgers, and any discrepancy — a transposed digit, a stale price, a copied account number — has to be tracked down by hand. LedgerLark, a young startup with five paying customers and a fresh $3.5 million seed round led by Motive Row, is betting that AI agents can take on much of that grunt work.
## An Insight Born From a Three-Day Headache
Co-founder Mina Batra spent seven years in operations at a regional brokerage, which is where the idea took shape. The moment she points to: watching a six-person team spend three days untangling a trade mismatch that turned out to be caused by a copied account number. The error was trivial; the process of finding it was not.
Her co-founder, Owen Carr, built reconciliation systems at a large bank — the kind of internal tooling that big institutions can afford to develop but smaller firms typically cannot. Together, their pitch is essentially that small firms deserve something closer to what the big players have built for themselves, without the headcount or the custom software budget.
## What the Product Actually Does
LedgerLark's agents compare three sources of truth — trade confirmations, custodian records and internal ledgers — flag discrepancies, and propose corrections. Critically, a human has to approve every fix. The system does not execute trades and does not move money.
That constraint is worth dwelling on. Keeping a person in the loop limits how much labor the product can actually remove, but it also makes the software far easier for a compliance-conscious firm to say yes to. For a startup selling into a risk-averse industry, that trade-off looks deliberate.
## Early Traction, With Caveats
Five firms are paying customers, with annual contracts ranging from $36,000 to $110,000. The company says one customer reduced its unresolved daily exceptions from 70 to 18 — a striking figure, though it comes from LedgerLark itself and reflects a single customer's experience. Whether that result generalizes across firms with different systems, asset mixes and data quality is an open question.
Five customers is real revenue but a small sample. The more interesting signal may come over the next year: whether those firms renew, expand usage, or quietly revert to spreadsheets.
## The Competitive Landscape Is Not Empty
Post-trade reconciliation is not a neglected problem. Established vendors have sold reconciliation software for decades, and larger institutions run internal systems of the sort Carr used to build. LedgerLark's differentiation, as best one can tell from the outside, rests on two claims: that AI agents can handle messier, less structured discrepancies than rules-based tools, and that the product is priced and packaged for firms too small for enterprise vendors to court aggressively.
Both claims are plausible. Neither is proven at scale. Incumbents can add AI features to existing products, and they start with something LedgerLark lacks — years of trust with a customer base that changes vendors reluctantly.
## The Slow Grind of Selling to Finance
Perhaps the biggest obstacle is not the technology but the sales cycle. Financial institutions often require lengthy security and procurement reviews before adopting new software, and a seed-stage company has a finite runway to survive them. A $3.5 million round buys time, but selling into finance has exhausted better-funded startups before.
## Worth Watching
LedgerLark is attacking a genuinely painful, genuinely unglamorous problem, with founders who have lived it from both the small-firm and big-bank sides. The early numbers are encouraging but thin, the competition is entrenched, and the buyers move slowly. The next twelve months — renewals, new logos, and whether that exceptions number holds up elsewhere — should tell us whether this is a feature, a niche tool, or the start of something larger.
OpenAI: GPT-5.6 Luna: # Startup Spotlight: LedgerLark Brings AI Agents to Post-Trade Cleanup
A copied account number can create a surprisingly large operational problem. At one regional brokerage, Mina Batra watched a six-person team spend three days tracing and resolving a single trade mismatch. The error was simple; finding it across confirmations, custodian records and internal ledgers was not.
That experience helped prompt Batra to explore LedgerLark, a startup building AI agents for post-trade operations at small investment firms. She co-founded the company with Owen Carr, who previously built reconciliation systems at a large bank. Batra spent seven years working in operations at a regional brokerage, giving the pair experience on different sides of the same problem: reconciling what firms intended to trade with what their systems and counterparties recorded.
LedgerLark’s product is designed to examine those records, identify inconsistencies and propose corrections for a person to approve. The agents compare trade confirmations, custodian data and internal ledgers, then present the likely resolution rather than automatically changing an account or transaction.
That boundary is important. LedgerLark does not execute trades or move money. Its role is closer to an investigative and administrative layer for post-trade workflows, where teams often spend time gathering evidence, checking formats and determining which record is authoritative.
The company says five firms are currently paying customers. Annual contracts range from $36,000 to $110,000, suggesting that LedgerLark is initially targeting firms for which operational errors are costly but dedicated engineering teams may be limited. According to the company, one customer reduced its unresolved daily exceptions from 70 to 18 after adopting the product. That is a meaningful reported change, though it is not yet clear how much of the reduction came from LedgerLark itself, changes in customer processes or the kinds of exceptions included in the measurement.
LedgerLark has raised a $3.5 million seed round led by Motive Row. The funding gives the company room to expand its product and customer base, but post-trade infrastructure is a demanding market. Established reconciliation vendors already serve financial institutions, and many larger firms rely on internal bank software that is deeply integrated with their records and approval systems.
LedgerLark’s opening may be among smaller investment firms whose workflows are too complex for spreadsheets but too specialized—or expensive—to replace with a broad enterprise platform. Its AI approach could help reduce the manual searching involved in exception management, particularly when information is distributed across systems with inconsistent formats.
The obstacles are substantial. Financial institutions may require lengthy security reviews and procurement processes before allowing a new system to handle sensitive operational data. Even without trade execution or money movement, a tool that reads confirmations, custody records and internal ledgers will need to demonstrate strong access controls, auditability and reliability. Human approval may reduce risk, but it also means the product must earn trust from the people responsible for signing off on corrections.
The central question for LedgerLark is therefore less whether AI can spot mismatches than whether its recommendations are dependable enough to become part of a firm’s daily control process. Its early customer count and reported exception reduction offer an initial signal. The next test will be whether those results hold across more institutions, more data formats and the long procurement cycles that shape financial software.
Model B better matches the brief: it stays closer to the given facts, attributes traction carefully, and gives competition and uncertainty real space without inventing interview color or a rules-based-vs-AI differentiation. Model A is punchier and better structured as a magazine piece, but those extra inferences and a slightly promotional close weaken correctness and instruction-following. (Second judge pass, order swapped — scores are the average of both: Model B better matches the brief: it stays within the given facts, hedges company-supplied traction, and covers insight, product, competition, and uncertainty without investor jargon or founder mythology. Model A is livelier and well structured, but invents color (extra error types, “she points to”), slightly overclaims automation, and slips into clichés such as runway and new logos.)
Chronology with narrative momentum
Turn the scrambled reporting below into a 450–600-word Markdown news article.
Choose an opening that captures the consequence of the failure, then reconstruct the sequence clearly. Preserve the chronology without producing a minute-by-minute incident log. Give technical details enough context for a general reader.
Reporting:
* KiteBridge makes expense-management software.
* At 8:45 a.m. Tuesday, it deployed version 4.2.
* The first customer complaint arrived at 9:12 a.m.
* A payment-processor timeout caused KiteBridge’s fallback system to treat retries as new reimbursement requests.
* The software generated 13,400 duplicate reimbursements totaling $2.8 million.
* KiteBridge suspended automated payments at 10:03 a.m.
* Ninety-one corporate customers were affected.
* By Wednesday evening, $2.3 million had been recovered. The rest remained pending.
* No employees had money withdrawn from their accounts; employers funded the duplicate payments.
* An April changelog said version 4.0 had introduced “exactly-once payment protection.”
* KiteBridge’s public incident page initially described the problem as delayed reimbursements.
* The incident page first disclosed the duplicate payments six hours after automation was suspended.
* CEO Lena Ortiz apologized and said the company would publish a technical report.
Anthropic: Claude Fable 5: # Software Glitch Sends $2.8 Million in Duplicate Reimbursements Through KiteBridge's Platform
A faulty software update at expense-management company KiteBridge triggered 13,400 duplicate reimbursement payments totaling $2.8 million on Tuesday morning, affecting 91 corporate customers and raising questions about the company's earlier promises of payment safeguards.
The trouble began at 8:45 a.m., when KiteBridge deployed version 4.2 of its platform, which businesses use to process employee expense claims and issue reimbursements. Within half an hour, at 9:12 a.m
OpenAI: GPT-5.6 Luna: # KiteBridge Sends $2.8 Million in Duplicate Reimbursements After Software Update
KiteBridge accidentally issued 13,400 duplicate employee reimbursements totaling $2.8 million after a software update caused its payment system to misread repeated requests. The expense-management company halted automated payments Tuesday morning, but not before the failure affected 91 corporate customers.
The incident began after KiteBridge deployed version 4.2 at 8:45 a.m. Tuesday. The update was intended to process reimbursements through the company’s payment infrastructure. Instead, a timeout from the payment processor triggered a problem in KiteBridge’s fallback system.
When the processor failed to respond promptly, the fallback system treated retry attempts as new reimbursement requests rather than recognizing them as repeats of payments already submitted. That distinction is central to preventing duplicate transactions: a payment should be processed once, even if the system has to retry the request or wait for confirmation.
The first customer complaint arrived at 9:12 a.m. As reports accumulated, KiteBridge investigated the payment activity and suspended automated payments at 10:03 a.m. The company’s response prevented additional automated transactions, but thousands of duplicate reimbursements had already been sent.
The payments were funded by the affected employers, not withdrawn from employees’ personal bank accounts. KiteBridge said no employees had money taken from their accounts as a result of the incident. The duplicate payments instead created a financial and administrative problem for the 91 companies using the platform, which had to identify and recover excess funds.
By Wednesday evening, KiteBridge had recovered $2.3 million of the $2.8 million total. The remaining amount was still pending, according to the company.
The company’s public communication initially characterized the incident as delayed reimbursements. It did not disclose that duplicate payments had been issued until six hours after automated payments were suspended. That delay added to questions about how quickly customers were informed about the scope and nature of the failure.
The incident also calls attention to KiteBridge’s previous claims about payment safeguards. An April changelog for version 4.0 said the release had introduced “exactly-once payment protection,” a technical measure intended to ensure that a single reimbursement request cannot result in multiple payments, even when systems retry after an error or timeout.
KiteBridge has not publicly explained why that protection failed in version 4.2 or whether the safeguard was bypassed, improperly configured or incompatible with the fallback system. The company’s chief executive, Lena Ortiz, apologized for the incident and said KiteBridge would publish a technical report.
That report is expected to detail how the payment-processor timeout interacted with the fallback logic, why the duplicate requests were not detected earlier and what changes the company will make to prevent a recurrence. Until then, customers are left managing the consequences of a software failure that converted an ordinary retry mechanism into millions of dollars in unintended reimbursements.
Model A is an incomplete fragment that cuts off mid-sentence and misses the required 450–600-word article, while Model B delivers a complete news piece with a consequence-led lede, clear chronology, and accessible technical context. Model B is the only output that substantially fulfills the task. (Second judge pass, order swapped — scores are the average of both: Model B delivers a complete, chronological news article in range with a consequence-led lede, explained technical context, and the given facts intact. Model A is truncated mid-sentence after two short paragraphs and therefore fails the assignment.)
Controlled wit
Write a 350–450-word technology news story in Markdown.
The subject invites humor, but the reporting should remain serious. Use no more than one dry line of original humor. Do not mock customers or turn the story into a comedy routine.
Reporting:
* QueueZero launched an AI meeting assistant called Proxy on Wednesday.
* Proxy can attend a meeting, produce notes and assign follow-up work to software agents.
* During the launch demonstration, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.”
* Proxy interpreted the remark as an instruction.
* It created two tasks assigned to QueueZero CEO Evan Wu: “Stop talking” and “Send pricing sheet.”
* Wu laughed, completed the second task and deleted the first.
* QueueZero says Proxy requires confirmation before an external message is sent.
* Internal tasks can be created automatically.
* Administrators can disable automatic task creation.
* Proxy costs $30 per user per month.
* It enters public beta next week.
Anthropic: Claude Fable 5: # QueueZero's AI Meeting Assistant Assigns Its First Task — to the CEO
QueueZero launched its AI meeting assistant, Proxy, on Wednesday, and the product wasted no time demonstrating both its capabilities and its literal-mindedness — by assigning a task to the company's own chief executive during the launch event.
Proxy is designed to attend meetings on behalf of users, generate notes, and automatically assign follow-up work to software agents. During Wednesday's demonstration, that automation kicked in earlier than the company may have planned.
Partway through the presentation, a customer in the audience remarked, "Someone should tell Evan to stop talking and send us the pricing sheet." Proxy, which was actively monitoring the session, interpreted the comment as an instruction. Within moments, it had created two tasks and assigned both to QueueZero CEO Evan Wu: "Stop talking" and "Send pricing sheet."
Wu laughed off the incident on stage. He completed the second task, distributing the pricing sheet to attendees, and deleted the first. He did not, according to attendees, stop talking.
The moment offered an unscripted illustration of both Proxy's core functionality and the challenges of deploying AI systems that parse natural conversation for actionable requests. Distinguishing between an offhand remark and a genuine directive remains one of the harder problems in the category, and Wednesday's demonstration showed how easily the line can blur when an assistant is listening to everything said in a room.
QueueZero emphasized that Proxy includes safeguards around its more consequential actions. The system requires explicit user confirmation before sending any external message, meaning an errant remark cannot result in an email or document leaving the organization without human sign-off. Internal task creation, however, happens automatically by default — which is how Wu ended up with his impromptu to-do list. Administrators can disable automatic task creation entirely if they prefer a more conservative configuration.
The incident is unlikely to slow the product's rollout. AI meeting assistants have become one of the most crowded segments in workplace software, with vendors racing to move beyond transcription and summaries into what the industry calls "agentic" behavior — systems that act on what they hear rather than simply recording it. Proxy's launch-day misfire is arguably a preview of the kinds of edge cases enterprises will need to manage as these tools gain autonomy.
Proxy is priced at $30 per user per month and enters public beta next week. QueueZero did not say whether any changes to the default settings are planned before the beta begins.
OpenAI: GPT-5.6 Luna: # QueueZero’s AI meeting assistant turns an audience comment into executive task list
QueueZero launched Proxy, an AI meeting assistant, on Wednesday, presenting it as a tool that can attend meetings, produce notes and assign follow-up work to software agents.
The launch demonstration also showed how quickly an assistant designed to interpret conversation can move from observation to action.
During the event, a customer in the audience said, “Someone should tell Evan to stop talking and send us the pricing sheet.” Proxy interpreted the remark as an instruction and created two tasks assigned to QueueZero CEO Evan Wu: “Stop talking” and “Send pricing sheet.”
Wu laughed, completed the pricing-related task and deleted the other. For a moment, Proxy had adopted a notably direct management style.
The incident illustrated a central distinction in QueueZero’s product design: Proxy can automatically create internal tasks, but it requires confirmation before sending an external message. That approval step is intended to prevent the assistant from turning conversational remarks into messages sent to customers, colleagues or other outside recipients.
Administrators can also disable automatic task creation, according to QueueZero. Organizations using the system will therefore be able to choose whether Proxy creates internal assignments on its own or only records possible follow-up work for human review.
QueueZero describes Proxy as an assistant that can participate in meetings and convert discussions into operational tasks for software agents. The company did not provide further details about which agents are supported or how the system determines when a comment should be treated as an instruction rather than a casual remark.
That judgment is likely to matter as AI meeting tools become more involved in workplace workflows. Meetings routinely include suggestions, hypotheticals and offhand comments that may sound like commands when removed from their broader context. Requiring confirmation for external communication limits the consequences of some errors, while automatic internal task creation can still produce unwanted assignments or administrative noise.
Proxy will enter public beta next week. QueueZero is pricing the service at $30 per user per month.
The company’s demonstration offered a compact example of the trade-offs involved in delegating workplace follow-up to AI: automation can save time when it correctly identifies an obligation, but a system that acts on ambiguous speech may need clear controls and review mechanisms. In this case, the affected executive was also the person overseeing the launch, giving him an unusually convenient opportunity to test both the assistant’s initiative and its undo button.
Model B stays closer to serious reporting, sticks to the given facts, and keeps the wit dry, while Model A invents attendee color and leans on extra punchlines. Both land in the word count and cover the brief, but Model B better matches controlled wit without turning the incident into a gag. (Second judge pass, order swapped — scores are the average of both: Model B stays closer to serious reporting, covers the required facts accurately, and keeps wit mostly dry, despite a second humorous closer. Model A is punchier but invents details, stacks multiple jokes, and drifts toward a comedy-tinged feature rather than controlled news writing.)
Quotes woven into narrative
Write a 400–500-word Markdown story about the closing of Foundry Hall.
Use quotations selectively. Paraphrase when the information matters more than the speaker’s exact phrasing. Avoid stacking quotes or giving every source an equal-sized turn. The story should carry an emotional current while remaining restrained.
Facts:
* Foundry Hall, a nonprofit startup accelerator in Milwaukee, will close October 31.
* It opened in 2018.
* It has supported 212 startups, which collectively employ 1,140 people.
* Its largest corporate sponsor, Blank River Bank, supplied 38% of its annual budget.
* The bank declined to renew its sponsorship.
* Foundry Hall has enough money to operate through October but not through another full program.
* Six employees will lose their jobs.
* Twelve startups currently using its workspace must leave.
Quotes:
* Executive director Camille Ross: “We kept waiting for the replacement check that never came.”
* Founder Dev Malik: “My first investor meeting happened in that building. Before Foundry Hall, I was pitching from the back table of a coffee shop.”
* Blank River Bank spokesperson: “The decision reflects a change in our community investment priorities and is unrelated to Foundry Hall’s performance.”
* Deputy mayor Susan Cole: “The city is speaking with building owners and local universities about space for the current companies.”
Anthropic: Claude Fable 5: # Foundry Hall, Milwaukee's Startup Accelerator, to Close After Six Years
The lights in Foundry Hall will go out for good on October 31.
The nonprofit startup accelerator, which opened in 2018 in a converted industrial building on Milwaukee's near south side, announced this week that it will shut down at the end of October. The reason is arithmetic more than anything else: Blank River Bank, its largest corporate sponsor, declined to renew a sponsorship that covered 38 percent of the organization's annual budget, and no replacement emerged.
"We kept waiting for the replacement check that never came," said Camille Ross, the accelerator's executive director.
Ross said the organization has enough money to keep the doors open through October, but not enough to run another full program. Rather than start a cohort it couldn't finish, the board voted to wind down.
## What gets left behind
The numbers Foundry Hall leaves behind are not small. Over six years, it supported 212 startups. Those companies now employ 1,140 people across the region — jobs that, in many cases, began as an idea sketched out in the hall's shared workspace.
Dev Malik, one of those founders, remembers what came before. "My first investor meeting happened in that building," he said. "Before Foundry Hall, I was pitching from the back table of a coffee shop."
The closure has immediate human costs. Six staff members will lose their jobs. Twelve startups currently working out of the building must find new space by the end of October — a scramble that, for early-stage companies operating on thin margins, is more than an inconvenience.
## The sponsor's exit
Blank River Bank offered little explanation for its departure. A spokesperson said the decision reflected a shift in the bank's community investment priorities and had nothing to do with Foundry Hall's performance.
That distinction may matter to the bank. It changes nothing for the six employees or the twelve companies packing boxes.
Ross declined to criticize the bank directly, but the strain in the organization's final months is evident. A single sponsor supplying more than a third of the budget was always a vulnerability; when the money stopped, there was no cushion deep enough to absorb the loss.
## What comes next
City officials say they are trying to soften the landing. Deputy Mayor Susan Cole said the city is in conversations with building owners and local universities about workspace for the displaced companies. Nothing has been finalized.
Whether Milwaukee replaces Foundry Hall — the institution, not just the square footage — is a harder question. Accelerators are more than desks and conference rooms. They are the accumulated network of mentors, investors, and founders who know where to find each other.
For six years, that network had an address. After October 31, it won't.
OpenAI: GPT-5.6 Luna: # Foundry Hall to Close After Six Years of Supporting Milwaukee Startups
Foundry Hall, the Milwaukee nonprofit startup accelerator that helped turn ideas into companies, will close its doors on Oct. 31.
The organization opened in 2018 and has supported 212 startups, whose businesses collectively employ 1,140 people. Its workspace, workshops and investor meetings became an entry point for entrepreneurs who might otherwise have struggled to find equipment, advice or a place to begin.
But Foundry Hall will not make it to another full program. Its largest corporate sponsor, Blank River Bank, provided 38% of the accelerator’s annual budget and declined to renew its sponsorship. The organization has enough money to operate through October, but not enough to continue beyond that.
“We kept waiting for the replacement check that never came,” Executive Director Camille Ross said.
The closure will eliminate six jobs and force 12 startups currently using Foundry Hall’s workspace to leave. For some founders, the move represents more than the loss of an office. It means losing a shared environment that connected them with mentors, investors and one another.
Dev Malik, who founded one of the companies that worked out of the building, remembers arriving before he had much more than an idea. “My first investor meeting happened in that building,” he said. “Before Foundry Hall, I was pitching from the back table of a coffee shop.”
Foundry Hall’s supporters say its impact cannot be measured only by its budget or the number of companies housed under its roof. The startups it assisted range from early-stage ventures still seeking funding to businesses that have grown into employers. Together, they represent a small but visible part of Milwaukee’s efforts to build an innovation economy.
Blank River Bank said its decision was not a judgment on the accelerator’s work. In a statement, a spokesperson said the change “reflects a change in our community investment priorities and is unrelated to Foundry Hall’s performance.”
The decision leaves city officials and other community organizations looking for ways to keep the current companies from being scattered. Deputy Mayor Susan Cole said the city is talking with building owners and local universities about possible space for the startups.
Those conversations may help preserve pieces of Foundry Hall’s network, but they cannot replace the institution immediately. Its staff will spend the remaining weeks winding down programs, assisting the companies in the workspace and preparing for the final day.
When the doors close, Milwaukee will lose more than an accelerator. It will lose a physical place where inexperienced founders could test a pitch, meet a first investor or find someone who understood the uncertainty of starting a company.
For the entrepreneurs who came through Foundry Hall, the building’s end will not erase what began there. But it will leave a gap that the city will have to decide how—and whether—to fill.
Model A better follows the quote-weaving brief by using only the two voice-driven quotations and paraphrasing the bank and deputy mayor, with tighter, more restrained prose. Model B is factually a bit cleaner but quotes the informational bank line, pads more, and gives sources more even treatment. (Second judge pass, order swapped — scores are the average of both: Model A better follows the quote-weaving brief by paraphrasing the bank and mayor, avoiding equal-sized source turns, and sustaining a restrained emotional current; Model B is more faithful to the given facts but quotes more sources and reads flatter.)