Context.dev says its agent web extractor uses 97% fewer tokens than Codex

The September 28th launch pairs query-based passage extraction with company-run tests against Codex, Exa, Claude Code and Zilliz.

By · Published

Primary source: X

Why it matters

Agent builders pay to process the context they send to models, but aggressive extraction can discard evidence an answer needs. Context.dev's benchmark makes a credible efficiency case against a weak one-call Codex baseline and a closer comparison with Exa; independent testing will determine how well the savings hold up in real agent workflows.

An unbranded laptop shows a long web article with one short passage highlighted, beside a notepad and a ceramic mug.

Context.dev launched Context Highlights on September 28th, a tool that takes a webpage and a query and returns selected passages with nearby context for an AI agent. In an October 1st thread on X, the company said its extractor used 97% fewer tokens than Codex in a benchmark. That figure comes from Context.dev's own test, and the Codex comparison used a limited baseline.

https://x.com/getcontextdev/status/2105736079521481182

poster=/api/storage/public-objects/tweet-videos/context-dev-highlights-web-extraction-token-benchmark-poster-5bb6e430.jpg|Video from @getcontextdev on X
Video from the original post on X.

Founder and CEO Yahia Bakour (@mynameisyahia) built Context.dev around the web-data plumbing teams need when agents have to work with current online information. Before starting it, he worked at Amazon as a software engineer and checkout-platform tech lead, led engineering at Sunrun, and co-founded Stock Alarm, which was acquired in 2023. Y Combinator lists Context.dev in its Summer 2026 batch. Bakour says he started the company after seeing teams repeatedly build and maintain browsers, crawlers, parsers, queues and other web-data infrastructure themselves.

Highlights applies that infrastructure to a narrower problem: deciding what from a fetched page belongs in an agent's context window. Context.dev describes the output as extractive rather than a model-generated summary. Its documentation shows a request supplying a URL and a question, with controls for the maximum length of returned passages. The product aims to keep relevant evidence and surrounding text while leaving unrelated page content out.

The headline 97% figure is based on a 250-question WebCode benchmark adapted from Exa's test. Context.dev reported 76.0% groundedness and an average of 229 output tokens for Highlights, against 57.6% and 6,889 tokens for Codex's native page-opening tool. Groundedness measured whether the returned text supported every substantive fact in a reference answer. Context.dev also reported 75.6% groundedness and 286 tokens for Exa Highlights, putting the two products close on evidence coverage while Context returned about 20% fewer tokens.

The Codex result needs its setup alongside the percentage. Context.dev used one native page-open request, without passing the question or making follow-up calls; the company says an agent could retrieve relevant passages with additional calls. Its token count also included line and citation markers. The benchmark therefore measures a particular one-call Codex setup, not every way an agent could use Codex to find page evidence. Context.dev counted failed requests against groundedness but excluded them from token averages.

The same test illustrates why token reduction alone is an incomplete scorecard. Claude Code's WebFetch returned 231 tokens on average, close to Context Highlights' 229, but scored 50.4% groundedness. Firecrawl Highlights returned only 52 tokens and scored 48.8%. Exa and Context were nearly level on groundedness, so the company's results support a claim of lower output volume relative to that competitor, not a clear quality win.

A second company-run evaluation used all 1,000 questions in SimpleQA Verified. Context.dev said an LLM answering from Highlights achieved 76.7% accuracy with 183 tokens on average. Against Zilliz's semantic-highlighting model at one threshold, Context reported 74.0% accuracy and 577 tokens for Zilliz. At another Zilliz threshold, output fell to 133 tokens, while accuracy dropped to 60.0%. The test measures answers generated from supplied highlights; it does not show that every relevant fact survived extraction or establish general performance across agent tasks.

Context.dev says the benchmark runs took place on September 25th. For WebCode, the company used a local version of its extractor rather than the deployed endpoint, and a model judge to score the results. On SimpleQA Verified, it included 92 source-fetch failures as zero-token outputs. Those details make the results more interpretable, while leaving replication by an independent evaluator as the next test of the product's claims.

Highlights is available through Context.dev's scrape API and through its MCP setup for coding agents, including Claude Code and Codex. The move fits the company's broader web-data API business: Bakour has described the recurring burden of maintaining scrapers and related systems as the problem Context.dev was built to take on. Highlights puts a specific cost lever on that pitch - send less irrelevant text into the model - while its tests show the tradeoff the product must manage: keeping enough evidence for an agent to answer accurately.

Reader comments

Conversation for this story loads after sign-in.