past.dev turns its agents' memory system into an API

Mehdi Djabri and Alex Ronse say past.dev's API grew out of agents they built for companies. Its BEAM scores lead published results, though competitors' runs used different models, judges and configurations.

By · Published

Primary source: PR Newswire

Why it matters

past.dev is betting that companies will buy a dedicated memory layer to track changing facts, evidence and access rights across AI agents. Its benchmark lead supports the pitch, while uneven comparison setups and the absence of named customers or quantified production deployments leave commercial scale unclear.

A laptop connected by cable to an organized archive of memory cards illustrates past.dev turning its agents' memory system into an API.

Mehdi Djabri and Alex Ronse built agents for product work, engineering, email and meetings before turning the memory infrastructure behind those systems into past.dev, an API for AI agents. past.dev described the product in an October 6th PR Newswire release, but the API was already presented as available in a past.dev blog post dated September 30th. The six-day gap means the October 6th release formally introduced an API past.dev had already presented as launched.

The pitch grew out of an operational problem the founders say they encountered while building agents for teams: business information changes, arrives late, comes from multiple sources and carries different access rights. A conventional search can surface a relevant message without establishing whether its contents are still current or whether the person asking should see it. past.dev aims to track the fact, its history, its source and its audience together.

Djabri's background spans product design and company-building: he worked at Lyft and Blackbird and co-founded Iteration X before past.dev. He and Ronse say they have worked on model context, memory systems and agents since the GPT-2 era. On past.dev's about page, they describe the API as an outgrowth of agents they had built to handle work across product, engineering, email and meetings. Their argument is grounded in the systems they say they needed themselves: once an agent follows a team's work over time, it must distinguish a current decision from an old one and retain evidence for the distinction.

Memory as a product layer

Developers send past.dev timestamped text, such as emails, meeting transcripts, messages or documents, then ask questions through a recall request that includes the requester's identity. The API is designed to return current facts, what those facts replaced, and dated source material, filtered by what the requester may see. It can also answer questions about an earlier date, according to the product materials.

Diagram of timestamped text entering past.dev memory, followed by an identity-bearing recall request and a response filtered by permissions, with a separate path for questions about an earlier date.
past.dev says its API connects timestamped text and a record of facts, history, sources and permissions to recall responses - AI explanatory diagram, not documentary evidence. RuntimeWire · AI-generated diagram.

past.dev packages retrieval with a persistent record of changing facts and permissions that can sit alongside different AI models. The API supports MCP, and past.dev says enterprise customers can self-host or use dedicated regions. The product is available through self-serve signup, according to the launch materials.

The founders' commercial bet is that teams deploying agents across internal systems will need an auditable memory layer rather than asking each application to assemble and maintain its own history. That need is plausible, especially where a stale answer or an answer shown to the wrong person carries a cost. The release says the system grew out of agents built for companies and had run in production for large teams across very large volumes of data. past.dev's About page says early customers receive help with evaluation, deployment and production testing. past.dev names no production customers or quantifies deployments, data volumes, revenue or usage, leaving the scale of demand unclear.

What the benchmark shows

past.dev reports scores of 92.08% at 100,000 tokens, 89.63% at 500,000, 90.65% at 1 million and 85.03% at 10 million on BEAM, a benchmark of long conversations and memory tasks published at ICLR 2026. Its benchmark page says its results came from a September 29th, 2026 run covering complete splits: 400 questions at 100,000 tokens, 700 each at 500,000 and 1 million, and 200 at 10 million.

The scores lead the other published results shown on past.dev's benchmark page. At 10 million tokens, the next listed result is 68.0%, from Exabase M-1. The smaller-history comparisons also show sizeable gaps. Those figures are company-published results, however, and the same page says the competing systems' results come from separate runs with different models, judges and configurations. They establish a lead among published numbers, not a controlled head-to-head test under identical conditions.

There is a discrepancy inside the launch materials. The release's table lists the previous best scores at 100,000, 500,000 and 1 million tokens as 86.2%, 80.1% and 79.1%. Its accompanying graphic instead shows 76.9%, 71.1% and 75.0%. The graphic's figures match the current benchmark page, while both versions give 68.0% at 10 million. That difference changes how large the claimed margin appears at the shorter history sizes. The benchmark page provides a reproducible evaluation harness, and Ronse said in the release the team opened it so others can check the results; readers comparing systems still have to account for the differences in published evaluation setups.

past.dev names General Catalyst, Varsity and Connect Ventures as backers, without giving a round size or valuation in the release. The release and company materials describe production work with teams, but disclose no customer names or quantified production deployments.

Reader comments

Conversation for this story loads after sign-in.