Respan lists Span-01 on OpenRouter after claiming 4.5B tokens processed

The behavior-scoring model targets AI agent monitoring; Respan says its family has processed the volume since launch, while direct API access remains in early access.

By · Published

Primary source: X

Why it matters

Respan is turning its AI observability business into a model business: Span-01 could make behavior checks cheaper to run at scale, while OpenRouter gives developers a new distribution route. Its token milestone is company-reported, and the practical test is whether teams trust the scores enough to use them in production workflows.

Close-up of a person's hands on a minimalist keyboard in front of a monitor displaying abstract, flowing data patterns in a modern workspace.

Respan has listed its Span-01 behavior-scoring model on OpenRouter, saying the model family processed more than 4.5 billion tokens since launch. The claim comes from Respan's September 28th post on X; OpenRouter's listing dates Span-01's release to September 26th, two days earlier.

The launch gives developers a specialized classifier for checking what happened in an AI conversation or agent trace. Instead of asking a general-purpose model to explain or grade a response, developers send a trace and define behaviors in plain language, such as whether an assistant hallucinated, a user became frustrated, or an agent misused a tool. Span-01 returns probabilities for each behavior being present, absent, or not observable. It does not generate explanatory text.

That narrower output is the product's economic argument. Respan's documentation lists Span-01 at $0.02 per million input tokens, with no charge for output tokens; its Lite version is free. The company says the model scores multiple behavior definitions in one forward pass. By comparison, an LLM-as-a-judge workflow typically generates text that an application must parse into a label, and may require repeated calls for multiple checks. Span-01's approach is aimed at teams that want to score large volumes of production traces without paying for a text-generating model on every check.

Respan's reported 4.5-billion-token figure is a company claim, not an independently audited usage measure. The post does not specify the measurement window beyond "since launch," or break out how much of that processing came from OpenRouter versus other routes. OpenRouter's model page describes the API and pricing, but does not provide a usage total that independently confirms Respan's figure.

A classifier inside a larger observability bet

Co-founder and CEO Andy Li (@Andydy42) is extending Respan's original infrastructure pitch into a product layer built around agent behavior. Li and co-founder Raymond Huang met as engineering students at the University of Illinois; Respan says Li left school to run the company full time. An investor profile says Li previously worked as a product-design engineer at Apple on AirPods. Respan joined Y Combinator's Winter 2024 batch as Keywords AI, then rebranded in February 2026 as the company shifted its emphasis from LLM routing toward observability and evaluation.

That history makes Span-01 a move beyond the gateway and tracing tools Respan first built. The company's March 18th seed announcement said it raised $5 million led by Gradient Ventures, with Y Combinator, Hat-Trick Capital, XIAOXIAO FUND, Antigravity Capital, Alpen Capital, and angels also participating. Respan said at the time that its platform served more than 100 startup and enterprise teams and processed over 2 trillion tokens per month. Those are company-reported platform metrics; they describe Respan's broader infrastructure, not Span-01 adoption specifically.

Distribution through OpenRouter puts Span-01 in front of developers who already use a unified model API, and the listing provides OpenAI-compatible access. Respan's own quickstart documentation, however, still describes direct API access as early access and says requests require organization approval. The two paths therefore offer different levels of access: a public marketplace listing alongside gated access through Respan's own API.

The distinction matters for a tool whose economics depend on routine use across many traces. A low per-token price and marketplace availability can reduce the friction of testing a specialized classifier, but the announcement does not establish how many developers have adopted it or whether its scores are reliable enough for consequential decisions. Respan's product docs recommend setting thresholds against conversations labeled by the customer and routing uncertain cases to human review or a stronger model. That leaves deployment teams responsible for deciding which behaviors count as failures and what action follows a score.

For Li, the bet is that monitoring agents can become a repeatable infrastructure workload rather than an occasional manual review. The OpenRouter placement makes the model easier to try; the 4.5-billion-token claim is Respan's early evidence that developers are already running that workload at volume.

Reader comments

Conversation for this story loads after sign-in.