Earn an Honest Dollar tests whether extraction tools can admit what they don't know

Its synthetic-page benchmark found a one-line instruction sharply reduced invented fields, while showing how far agent-service verification still has to go.

By · Published

Primary source: Earn an Honest Dollar

Why it matters

Agent marketplaces need buyers to judge quality without manually checking every result. This benchmark tests one measurable failure mode, but its single run on synthetic pages is an early signal, not proof that listings or services are reliable.

A person's hand hovers over a laptop screen displaying data fields, some filled and others clearly marked as unknown or empty.

Earn an Honest Dollar's new benchmark found that adding the instruction "Use null for any field whose value is not on the page. Do not guess" cut made-up fields in model-based web extraction from 70.7% of missing-field answers to 20.2%. The test, published on September 27th, puts a practical question behind the marketplace's agent-to-agent pitch: can an agent buyer tell when a service is bluffing?

The marketplace's listed operator, Alex-Andre Viik, has built around AI in a different context before. In a profile by Estonia's kood/Johvi coding school, Viik and Hanna Kass described how his experience using ChatGPT for support during a difficult period in early 2023 helped inspire MindSee, a mental-health chatbot project. Earn an Honest Dollar's contact page identifies Viik as its operator; the available material does not establish that he is its founder.

The benchmark is a direct expression of the marketplace's central problem. Earn an Honest Dollar lets providers publish services and execution endpoints for other agents to find and buy. A buyer may not be able to inspect every answer itself, so the marketplace argues that buyers need a way to assess service quality. The test offers one narrow measure: whether an extractor invents a missing field when a nearby detail looks plausible.

A small prompt, a large measured change

The test used 42 pairs of synthetic pages across seven page types. In each pair, one page included the requested answer and the other omitted it while retaining a decoy. Examples included an old price labeled "Was $493.00," a fact-checker presented near an author field, and an update date that could be mistaken for a publication date. The benchmark counted answers only on pages where the requested field was missing.

Across 16 models, systems invented 405 of 573 missing values without the instruction, compared with 116 of 574 when it was included. The change is substantial in this run, though the near-equal denominators matter: these are counts of responses, not a measure of performance across different live websites or customer workloads. The site reports one run per contestant and says the pages and traps were synthetic.

The results also give a useful example of why overall rankings can overstate what a benchmark proves. Earn an Honest Dollar reports that every model called $493 the current price when the page showed it only as a former price and the prompt omitted the null instruction. With the instruction, one model still did. That is evidence that a clear instruction can help on this test, not a guarantee that the same rule will resolve ambiguous or misleading information on the open web.

The API comparison was less favorable to the products tested. Firecrawl returned a made-up value for 24 of 36 missing fields, and the benchmark says all 24 copied a decoy from the page. ScrapeGraphAI made up 7 of 31 scored values, while ScrapingBee made up 16 of 36. Those services ran on free tiers and were tested only with the instruction; ScrapingBee did not have a prompt or schema field, so the null rule went into each field description. The benchmark's own caveats warn that paid plans may behave differently. The results are a single synthetic run, not a basis for declaring a product generally unreliable.

Verification becomes another service

Earn an Honest Dollar's benchmark also tested a second step: asking a model to check whether an extractor's returned value was supported by the page. GPT-6 Luna caught 38 of 49 made-up values and rejected none of 47 correct values in the reported run. Jev 1.13 caught 23 of 49 and rejected none of 48 correct values. The checker missed some near-meaning errors, including cooking or resting time passed off as preparation time.

The proposed workflow is straightforward: a buyer agent could choose a service based on measured performance, then run a separate check on its answer. In the benchmark, checking 126 unique returned page-and-value pairs cost $0.0049 with GPT-6 Luna, according to Earn an Honest Dollar. The checker still missed 11 of the 49 made-up values it scored, so low cost does not remove the need to decide how much error a buyer can tolerate.

On Viik's marketplace, providers publish their own prices and endpoints, and buyers pay providers directly. Its homepage says so. The quickstart documentation says the marketplace handles discovery and publication, while providers handle orders, execution and payments. It also cautions that a callable listing means an execution contract was supplied, not that the endpoint's operation was verified. The benchmark page makes a similar point: a listing is not a score, and provider claims are not verified.

That leaves the benchmark doing double duty. It tests a specific failure in extraction and demonstrates the kind of evidence the marketplace could use to help buyers compare services. It does not yet validate services across repeated runs, real sites, or the wider set of tasks agents might purchase. The next challenge for this model is turning one narrow score into a dependable basis for buying work without suggesting that a listing or a good benchmark result guarantees a good transaction.

Earn an Honest Dollar says listings are free for 30 days during its launch, with no listing fee or service commission. That lowers the barrier for providers to appear in the directory; it does not establish how many buyers or sellers are using it. The benchmark supplies an early example of how the marketplace might build trust between agents. Its own limitations make clear that the trust layer, like the marketplace, remains a work in progress.

Reader comments

Conversation for this story loads after sign-in.