OpenAI pitches Codex for tax prep after a 7,000-return pilot

The 2025 tax-year pilot saved accountants 31% of preparation time, according to Current, while keeping practitioners responsible for review.

By · Published

Primary source: X

Why it matters

OpenAI is positioning Codex as infrastructure for vertical agents, with its Thrive stake providing the real workflows and expert feedback that generic model access cannot supply.

An illustrated scene shows a small accountant figure reviewing documents amidst a vast landscape of interconnected digital tax forms and data, depicting OpenAI's Codex in action.

Greg Brockman (@gdb), OpenAI's president and co-founder, pitched Codex on August 20th as infrastructure for products outside software development, pointing to a tax-preparation system that processed 7,000 returns and reduced accountants' preparation time by about a third.

The figures describe a pilot from the 2025 tax year, rather than a new deployment. OpenAI first detailed the project on May 27th, after six months of work with Thrive Holdings and Crete Professionals Alliance, the accounting network that rebranded as Current on June 2nd. Brockman, who was Stripe's chief technology officer before helping found OpenAI, resurfaced the results to argue that developers can use Codex's open-source harness as the operating layer for specialized agents.

That is a broader product pitch than the coding assistant OpenAI originally put in terminals and development environments. The Codex repository is published under the Apache 2.0 license, and its App Server exposes the agent harness through a bidirectional interface that developers can embed in other products. The open-source code handles agent threads, tool execution, configuration and approvals; access to OpenAI's models still requires a ChatGPT account or API setup.

What the tax system handled

Tax AI was built for Current's network of accounting firms. Participating accountants uploaded source documents and client notes, and the system extracted information and prepared submissions for tax-engine review. The pilot covered 1040 individual returns and 1041 returns for estates and trusts.

OpenAI said data entry for medium- and high-complexity filings can consume as much as eight hours per return. The work includes pulling information from prior-year filings, spreadsheets and other inconsistent client documents, then mapping it to the correct tax fields.

Current reported an average 31% reduction in preparation time across the 7,000 returns. OpenAI said the system increased throughput by about 50% and produced drafts with up to 97% accuracy. Current later described accuracy as high as 98%. Those figures are reported by the organizations that built and deployed Tax AI, and the top-line accuracy number does not describe how results varied by return complexity.

OpenAI supplied a more useful measure of improvement over time. When Tax AI launched, one-quarter of evaluated returns reached at least 75% correct field completion. Six weeks later, 86% reached that threshold, even as Tax AI moved from relatively straightforward W-2 and 1099 inputs into K-1 forms, rental-property schedules and other complicated filings.

Practitioners remained responsible for reviewing the work and approving final returns. OpenAI limited Codex's automated engineering work to the extraction and mapping layer, while engineers retained control over architecture, product decisions and production releases.

Corrections became engineering tasks

The central mechanism was a feedback loop built around accountants' corrections. Tax AI recorded what the system proposed, what a practitioner changed and what ultimately entered the filed return. Repeated errors could then be grouped into a finding, converted into a targeted evaluation and assigned to Codex as a bounded engineering task.

Codex received the relevant production trace, source documents, expected tax-engine output, code and test commands. It could inspect a failure, propose a change and run targeted and regression evaluations. Ambiguous cases went back to engineers rather than being turned automatically into code changes.

That distinction matters in tax preparation, where a mismatch can reflect an extraction error, an accountant's judgment, a value carried over from a previous return or a change introduced elsewhere in the filing process. A raw correction is not necessarily evidence that the agent was wrong.

OpenAI said rental-property support took roughly six weeks and substantial engineering oversight to reach 90% precision and recall. The resulting evaluation and review patterns were then reused for other schedules.

OpenAI has a stake in the rollout

The pilot also reflects OpenAI's strategy for moving agents into established service businesses. OpenAI took an ownership stake in Thrive Holdings in December 2025, with accounting and IT services named as the partnership's first targets. OpenAI agreed to embed research, product and engineering staff inside Thrive-owned operations.

That structure gives OpenAI direct access to production workflows, expert corrections and a distribution channel across acquired businesses. Thrive gets purpose-built automation for labor-intensive services, while OpenAI gets a controlled proving ground for agent systems that need domain feedback and repeated evaluation.

Current says its network has since grown to 48 firms with more than 2,000 employees across 39 states. The accounting group plans to expand Tax AI across additional firms, while Thrive and OpenAI are applying the same design to bookkeeping, audit and IT help-desk workflows.

Brockman's post turns the pilot into a developer pitch: Codex can supply the agent loop beneath a vertical product, while domain experts, production traces and evaluations determine whether that product becomes reliable. The harness is available to copy. The difficult part remains securing the workflow, review process and proprietary feedback that made the tax pilot improve.

Reader comments

Conversation for this story loads after sign-in.