LiteLLM launches Lens to investigate production agent traces at scale
Co-founder Ishaan Jaffer is extending LiteLLM's gateway into trace analysis, with Lens in early access and 200,000-plus traces described as a future target.
By Ryan Merket · Published
Primary source: X
Why it matters
Lens is LiteLLM's bid to turn its gateway position into a foothold in agent observability. Its early-access release outlines the strategy; customer adoption and performance at scale remain unproven.

LiteLLM co-founder and CTO Ishaan Jaffer (@ishaan_jaff) announced Lens on September 30th, a tool that uses AI agents to investigate production agent traces. The launch extends LiteLLM's role from routing model requests toward helping teams inspect what their agents do after those requests reach tools and applications.
Jaffer has worked on LLM behavior monitoring before. Before LiteLLM, he co-founded Berri, a Carnegie Mellon alumni startup whose product monitored LLM applications for errors such as hallucinations and refusals. Carnegie Mellon listed Berri in its 2023 VentureBridge cohort. That earlier work gives Lens a clear throughline in Jaffer's career: making failures in AI applications easier to detect. The sources do not establish that Berri directly prompted Lens, but the technical problem is familiar.
LiteLLM's announcement describes its gateway as a chokepoint for all enterprise AI traffic. That is a company claim, not an independently established measure of its share of enterprise requests. LiteLLM's existing product provides a unified interface for routing calls to multiple model providers; Lens applies that gateway position to observability, where teams need to find patterns across many agent runs.
From individual traces to recurring failures
LiteLLM's Lens documentation describes a workflow in which a user specifies expected agent behavior and asks Lens to investigate a set of runs. Lens identifies failures, groups similar problems and links findings back to the original traces. Investigations can run on demand or on a schedule, while individual traces remain available for manual inspection in the proxy's logs interface.
The distinction is operational. A trace records what happened in one run; an investigation is intended to find repeated failure patterns across many runs. As agents make more tool calls and handle longer tasks, manually reviewing every execution becomes difficult. Lens is designed to analyze that growing record automatically, so engineers do not have to scan traces one by one.
The product's setup also makes the infrastructure tradeoff explicit. LiteLLM says Lens stores traces in ClickHouse and investigation results in PostgreSQL. Agents send OpenTelemetry traces to a LiteLLM proxy, and the Lens analyzer runs on the customer's infrastructure, using a model selected by the customer through LiteLLM. That approach fits the gateway's existing self-hosting pitch, while asking customers to operate the supporting data stores themselves.
Jaffer says the product is being built for a future in which agent swarms generate "200K+ traces." That figure describes a target scenario, not a reported Lens benchmark or a volume the product has already processed. The launch image accompanying the post illustrates hundreds of agents producing thousands of model and tool calls, with each run represented by a trace containing steps, tokens and cost.
The announcement also pitches agent-first APIs for querying traces with SQL and having coding tools such as Codex and Claude Code analyze them. The proposed benefit is direct access to customer-controlled trace data without the rate limits of hosted tracing platforms. LiteLLM's documentation provides setup instructions and a waitlist for early access; it does not establish general availability, Lens pricing or measured performance at the announced scale.
The gateway is the distribution bet
LiteLLM is competing in a field that already includes LangSmith, Langfuse, Arize Phoenix and Braintrust, products built around combinations of tracing, monitoring and evaluation. Lens's proposed advantage is that LiteLLM already sits in the request path for customers using its gateway. If those customers also send agent traces through LiteLLM, LiteLLM can connect routing and spend data with records of agent behavior in one system.
The launch offers a distribution hypothesis; customer adoption and performance data would test it. Lens is still in early access, and the announcement supplies no customer adoption, pricing or evaluation results. The practical test is whether teams already using LiteLLM will route enough trace data through it to make these investigations useful, and whether they prefer that workflow to their existing observability tools.
Jaffer's previous company focused on identifying errors in LLM applications. With Lens, he is applying that concern to systems that take multiple steps and call tools, where an isolated bad answer can be harder to diagnose than a failure in a conventional prompt-and-response flow. LiteLLM is betting that control of the gateway can make the trace data needed for that diagnosis easier to collect and query. The product's early-access stage leaves that bet to be demonstrated.