Zheqing (Bill) Zhu (@ZheqingZhu), founder and CEO of Pokee AI, launched Pokee-Isaac 28B on Tuesday, August 4, with vendor-reported results suggesting the proprietary agent model can handle 10 million tokens of context on a single Nvidia RTX 4090-class GPU.
In its announcement on X, Pokee said the 28-billion-parameter model scored 93.3% on RULER at the full 10-million-token length and reached prefill speeds of as much as 137,000 tokens per second on one GPU. The company described Pokee-Isaac as a "frontier-class agentic model" built on a proprietary "non-decoder-only" architecture.
Those performance figures come from Pokee's internal testing. The company has not released the model's weights, detailed test configurations or enough architecture information for outsiders to assess how it reduces the memory and compute demands associated with very long prompts. Pokee said in replies to the announcement that it intends to open source a version in the future, without specifying whether that release would use the same weights or setting a date.
The claimed context length is the center of the launch. If independently reproduced, a 10-million-token window could allow an agent to inspect large document collections, code repositories, tool catalogs and conversation histories in one session without relying as heavily on retrieval systems that split data into smaller chunks. Pokee's website says the model is designed to plan, execute and review long-running tasks while operating inside a customer's infrastructure.
RULER, the benchmark behind Pokee's headline score, is a synthetic evaluation that tests retrieval, multi-hop tracing and aggregation across long inputs. Its original research paper was designed to distinguish an advertised context window from the amount of context a model can use effectively. A strong RULER result would support Pokee's long-context claim under the tested conditions, but it would not establish general reasoning quality or reliability in multi-step enterprise workflows.
Pokee's launch graphic also shows its model trailing a system labeled GPT-5.6-Luna on Terminal-Bench 2.1 and MCP-Atlas while narrowly leading it on the company's reported BFCL v4 and tau3 averages. The graphic marks the comparisons as internal or sourced from external vendors. Pokee has not provided independent replication of the claimed 10-million-token performance or single-GPU throughput.
Pricing is inconsistent across Pokee's launch materials. The benchmark graphic advertises $0.15 per million input tokens and $1 per million output tokens. Pokee's API documentation, accessed on August 4, lists the standard Pokee-Isaac tier at $0.30 per million input tokens and $5 per million output tokens. A high-reasoning tier is listed at $3 and $15, respectively.
Zhu's enterprise deployment bet
The model extends the thesis Zhu brought to Pokee after leading applied reinforcement learning at Meta AI. According to Zhu's public biography, he led Meta's Pearl production reinforcement-learning project and completed a Stanford PhD in reinforcement learning while working full time. He founded Pokee in 2024 around agents that plan and use tools across applications rather than stopping at generated text.
That background shaped Pokee's initial product. The Seattle-area company applies reinforcement learning to tool selection and task sequencing across services including Google Workspace, Slack, GitHub and Notion. In July 2025, GeekWire reported that Pokee had raised a $12 million seed round led by Point72 Ventures, with participation from Qualcomm Ventures, Samsung NEXT and other investors. Pokee was pre-revenue at the time and operating a public beta, according to the report.
The single-GPU claim gives Zhu a potential enterprise sales argument, but its significance depends on independent validation. Pokee offers hosted access alongside private deployments in a customer's cloud account, and says its agent infrastructure can run on-premise or in an air-gapped environment. A model that can genuinely handle 10 million tokens on the hardware Pokee names could reduce dependence on large inference clusters and give companies tighter control over sensitive documents and workflow data.
Pokee is offering access through its API and custom deployments rather than releasing weights that developers can inspect and run independently. The company therefore retains control over the model and its distribution, even when customers keep workloads within their own network boundaries.
The commercial case now depends in part on whether outside developers can reproduce the headline performance under disclosed conditions. Until then, the 10-million-token RULER score, 137,000-token-per-second prefill rate and RTX 4090-class deployment remain vendor-reported results.