LiteLLM launches Lens to analyze agent traces inside its gateway

CTO Ishaan Jaffer says Lens will group failures across large trace volumes; its early-access setup stores traces on customer infrastructure.

By · Published

Primary source: X

Why it matters

Lens gives LiteLLM a route from model traffic management into agent debugging and evaluation. Its customer-hosted analyzer matches the gateway's self-hosted model, but the 200,000-trace figure is a future scenario, not a reported usage metric.

LiteLLM launches Lens to analyze agent traces inside its gateway — CTO Ishaan Jaffer says Lens will group failures across large trace volumes; its early-access setup stores traces on customer infrastructure.

LiteLLM launched Lens on September 30th, a tracing and analysis product that uses AI agents to find recurring problems across agent runs. The launch puts the gateway startup into a crowded observability category while extending its pitch: the infrastructure routing model calls can also help teams investigate what their agents did with them.

https://x.com/ishaan_jaff/status/2105495463638270115

poster=/api/storage/public-objects/tweet-videos/litellm-lens-agent-trace-analysis-launch-poster-e09bb912.jpg|Video from @ishaan_jaff on X
Video from the original post on X.

The news came from Ishaan Jaffer (@ishaan_jaff), LiteLLM's co-founder and CTO, who described the gateway as a chokepoint for enterprise AI traffic. That is the company's positioning, not an independently measured share of enterprise traffic. Jaffer said Lens is meant to make sense of trace volumes that could exceed 200,000, framing that figure as a future operating scenario rather than a current customer or usage metric.

Jaffer studied at Carnegie Mellon University and worked on parallel-computing projects there, including a graph simulator that he says in his public profile achieved a sixfold speedup through work partitioning and parallel execution. He and co-founder Krrish Dholakia built LiteLLM as an open-source interface for routing calls across model providers. Y Combinator lists the company in its Winter 2023 batch and says LiteLLM raised a $1.6 million seed round from Y Combinator, Gravity Fund and Pioneer Fund.

From routing calls to inspecting runs

LiteLLM's Lens documentation describes a workflow for teams that cannot manually review every agent trace. Developers send traces to the LiteLLM proxy using OpenTelemetry. Lens then analyzes selected runs against criteria the team defines, such as whether an agent followed a process or recovered from a tool failure. Its findings group similar issues and link back to the underlying traces.

The product is designed to run within a customer's deployment. Lens stores trace data in ClickHouse and investigation results in PostgreSQL; a separate analyzer runs on the customer's infrastructure and calls a model through LiteLLM. The documented setup requires both databases and a worker connected to the proxy. That arrangement fits LiteLLM's self-hosted pitch and gives security teams a way to inspect agent behavior without sending trace data to a separate hosted observability service.

Diagram of Lens in a customer deployment: OpenTelemetry traces enter the LiteLLM proxy, a worker connects to the proxy, and Lens stores trace data in ClickHouse and investigation results in PostgreSQL. A customer-infrastructure analyzer calls a model through the proxy.
LiteLLM’s documentation describes customer-deployed Lens with ClickHouse for trace data, PostgreSQL for investigation results, and an analyzer that calls a model through LiteLLM — AI explanatory diagram, not documentary evidence. RuntimeWire · AI-generated diagram.

In Jaffer's launch thread on X, he also described agent-first APIs, SQL queries and letting Codex or Claude Code analyze traces without platform rate limits. Those details extend the documentation's self-service investigation flow: users can define a good run, write checks, select executions and request analysis manually or on a schedule. LiteLLM's announcement offers early access through a waitlist.

The distinction from ordinary logging is practical. A trace can contain a sequence of model calls, tool use and intermediate steps, so a single failed outcome may require examining a long execution rather than one request-response record. Lens aims to identify repeated patterns in those runs and send engineers to examples they can inspect. The analysis depends on teams instrumenting their agents and writing useful expectations and checks; the product does not establish that an agent can improve itself automatically.

A crowded observability market

Lens enters territory already served by products such as LangSmith, which offers trace monitoring and automatic clustering and failure analysis, and Langfuse, an open-source platform that can be self-hosted. LangChain introduced a chronological agent-session view in LangSmith on September 24th, six days before LiteLLM's Lens post. The launch window puts agent-trace analysis squarely among active product priorities for infrastructure companies.

LiteLLM's proposed advantage is distribution through the gateway it already sells and maintains. Instead of asking platform teams to route traces to another service, LiteLLM wants those teams to send them through the same proxy handling model traffic, then analyze the resulting data there. That can reduce another integration for existing users, while giving LiteLLM a reason to expand from model routing into debugging and evaluation.

There is a commercial incentive in that expansion. LiteLLM's open-source gateway is free to self-host, while its current enterprise offering adds governance, security and support on an annual, usage-sized basis. Trace analysis gives the product a new operational job close to core gateway workloads. Whether customers adopt Lens as part of their existing deployment or evaluate it against dedicated observability vendors will determine how much that adjacency helps LiteLLM extend its position.

For now, LiteLLM is pitching Lens as a tool for large trace volumes and controlled, customer-run analysis. Its own 200,000-trace figure remains a forward-looking scenario. The immediate product bet is simpler: if a gateway already sees the model calls, it may be a convenient place to help teams find where agent runs go wrong.

Reader comments

Conversation for this story loads after sign-in.