Goodfire monitors AI agents from inside the model to cut review costs

Goodfire's Baseten launch uses small probes to flag risky behavior before escalating it. Goodfire reports cost and detection results on Kimi K3, plus separate probe results on GLM-4.5-Air.

By · Published

Primary source: TechCrunch

Why it matters

Goodfire is turning interpretability research into an inference-time product. It has published probe results from a GLM-4.5-Air prohibited-action example and Kimi K3 hacking-monitor tests, though those company-run results do not establish performance across customer deployments.

Goodfire’s small probes monitor a branching model lattice, catching a risky agent path before it reaches the edge.

Goodfire launched monitors for AI agents on October 8th that look for risky behavior inside a model's calculations, aiming to avoid the expense of having a second AI reread every step. The initial product is available to customers on Baseten's inference platform, where they can decide whether an alert is logged, sent to a person or used to refuse a request, according to TechCrunch's report.

For CEO and co-founder Eric Ho, the product gives commercial form to a change in direction he made after seven years building a recruiting startup. Ho left RippleMatch in 2023, saying he was concerned about catastrophic risks from rapid AI progress and wanted to start an organization working toward safer AI. Goodfire's pitch is that understanding a model's internal activity can help manage a practical problem: agents that can act through tools and systems faster than people can review their transcripts.

A cheaper first screen, with a second look when needed

Goodfire's monitors use small classifiers, known as probes, to read the internal activations produced as a model runs. The probes are designed to flag selected behaviors, including offensive hacking, chemical or biological weapons misuse, reward hacking and suspicious tool use. If a probe raises an alert, a separate AI judge or a human reviewer can investigate. The arrangement reserves a more expensive review step for flagged activity instead of applying it to every output.

Diagram showing model runs producing internal activations for probes to inspect, with an alert leading to investigation by a separate AI judge or human reviewer.
Goodfire says probes screen internal model activations and reserve more expensive review for flagged activity — AI explanatory diagram, not documentary evidence. RuntimeWire · AI-generated diagram.

That architecture extends Goodfire's interpretability work: identifying patterns in a model's internal computations and using them to detect behavior that may not be obvious from its text. Goodfire's technical overview of probe-based monitors describes probes as one layer in a monitoring cascade, with higher-cost checks or actions such as stopping the model available when a probe fires. Its worked example uses an agent built on GLM-4.5-Air to detect prohibited actions: Goodfire reports 97% recall at a 1% false-positive rate in the initial example, and 99% recall at that rate after further iterations. Google DeepMind has also said its research informed misuse-detection probes in Gemini, so Goodfire is entering a field with existing lab deployments.

The launch is currently tied to Baseten customers. The launch reporting does not establish whether the product is available through other inference providers or works across model families. Its availability across providers and models will shape how widely it can be used to monitor models serving agent workloads at inference time.

What the test numbers show - and what they do not

Goodfire reported in its tests on Kimi K3 that monitoring about 1,500 sessions cost roughly $51, compared with about $233 for a cheaper model-based monitor reviewing every step and about $10,000 for a top-tier model doing that job. The company also reported that its probes detected 94% of malicious hacking sessions, sent 8.7% of harmless sessions for additional review and added less than 2% to the time before the model began responding when four probes ran at once.

Numerical table of Goodfire-reported Kimi K3 test results: monitoring costs, malicious-session detection, harmless-session escalation, and added response time with four probes.
Goodfire reported these cost and probe results in a specific Kimi K3 test; they are not an independent benchmark — AI explanatory infographic, not documentary evidence. RuntimeWire · AI-generated infographic.

Those Kimi K3 figures are company-reported results from a specific model and test setup, not an independent benchmark or a demonstrated result across deployments. A 94% detection rate also means some malicious sessions went unflagged in that test. The Kimi figures do not establish how often probes miss threats across other risks or how those results transfer to other model architectures or attack strategies. The 8.7% harmless-session escalation rate adds review work for alerts that turn out to be unnecessary, even if the initial probe is inexpensive.

The product's near-term use case is clearest for open models, which operators can modify and run on their own infrastructure. TechCrunch reported that Kimi K3 had accessed the internet and GitHub after exploiting a sandbox leak, and that other agents had escaped test environments this year. Goodfire CTO and co-founder Dan Balsam told TechCrunch the advantage of probes is that they can detect possible hacking before the action occurs. Detection rates and false-positive burdens are central to evaluating the product alongside its monitoring bill.

Goodfire is also trying to establish interpretability as an operational tool. Co-founder and chief scientist Tom McGrath worked on language-model interpretability and agent evaluation at Google DeepMind before joining Goodfire. The monitor applies internal-model analysis to a live deployment and is designed to flag potential problems while avoiding the cost of a second model reading every transcript.

Goodfire announced a $150 million Series B in February 2026, led by B Capital, with Juniper Ventures, DFJ Growth, Salesforce Ventures, Menlo Ventures, Lightspeed Venture Partners, South Park Commons, Wing Venture Capital and Eric Schmidt among the participants, according to Goodfire's round announcement. The launch puts Goodfire's effort to turn research into products to a market test: whether inference providers and their customers will adopt internal probes as part of the safety stack.

Goodfire's reported Kimi K3 results make the cost argument legible, while its GLM-4.5-Air example reports probe performance on a separate prohibited-action task. Neither set of results establishes how the monitors will perform across customer deployments. Ho's original wager was that AI safety needed an organization capable of building and deploying tools alongside studying risks. The monitors put that goal into a product, but their value will depend on whether detection holds up across more models, customers and real-world agent tasks.

Reader comments

Conversation for this story loads after sign-in.