OpenRelay launches one inference API across GPUs, TPUs and Trainium

OpenRelay says its network handles 100B tokens weekly across 22 locations while routing jobs by cost, latency and availability.

By · Published

Primary source: Y Combinator

Why it matters

OpenRelay is betting that AI accelerators become interchangeable supply. Reliable routing could reduce cloud lock-in while giving hardware owners a market for unused capacity.

A central glowing processor chip radiating orange and yellow heat, connected by numerous network cables to diverse thermal forms of hardware in a distributed network.

Earlier in August, Jaden Wang and Prashant Patel launched OpenRelay, a hardware-agnostic inference network that routes AI workloads across accelerators from Nvidia, AMD, Google and Amazon through a single endpoint.

The two-person, Seattle-based company is part of Y Combinator's Summer 2026 batch. In its YC launch post, OpenRelay said the network is processing 100 billion tokens a week across 22 physical locations in North America, Europe, Asia-Pacific and the Middle East, using eight accelerator configurations. Those operating figures are self-reported.

OpenRelay's pitch is straightforward: developers send a model or container to OpenRelay, specify latency and throughput requirements, and let its scheduler select the cloud, region and chip. OpenRelay says the scheduler continuously benchmarks available capacity and routes each workload to the lowest-cost accelerator that meets the requested performance targets.

OpenRelay claims the approach can reduce inference costs by as much as 20%. Its website gives a narrower range of 10% to 20% for high-throughput workloads. OpenRelay has not published the underlying workload mix, comparison period or benchmark methodology behind those percentages, so the savings claim remains dependent on each customer's model, traffic pattern and available hardware.

A broker for fragmented compute

OpenRelay calls the architecture an "inference delivery network," borrowing the basic structure of a content delivery network. One endpoint sits in front of capacity distributed across multiple operators. Instead of routing static files to nearby servers, OpenRelay routes computation according to model availability, price, latency and hardware performance.

The term predates OpenRelay. A 2021 research paper described inference delivery networks as groups of computing nodes that coordinate to place models and satisfy inference requests across devices, edge infrastructure and clouds. OpenRelay is applying that concept to a commercial pool of AI accelerators.

The supply side matters as much as the API. Data centers, GPU clouds and businesses with reserved capacity can connect their machines to OpenRelay. OpenRelay handles workload isolation, orchestration, metering and billing, then shares usage revenue with the hardware operator. That turns the platform into a two-sided market: developers want cheap, dependable inference, while infrastructure owners want utilization.

OpenRelay is also offering dedicated GPU virtual machines and batch jobs from the same capacity pool. Its early-access documentation says the inference API supports OpenAI and Anthropic wire formats, allowing developers to keep existing SDKs and change the base URL and API key. The /v1 API is listed as stable, although the broader service remains in early access.

That compatibility reduces the work required to test OpenRelay. It does not remove the harder infrastructure problem: keeping models warm, routing around failures and delivering consistent performance when the underlying machines span different chips, clouds and operators. OpenRelay's ability to enforce service levels across that mixed supply will determine whether the network behaves like dependable infrastructure or another spot market for GPUs.

Founders from both sides of the GPU market

Wang and Patel arrived at the problem from different parts of the compute stack. Wang says he built a half-megawatt data center from a warehouse shell at age 20 before working on virtualization and high-performance computing at Voltage Park. When Voltage Park acquired GPU marketplace TensorDock in March 2025, Voltage Park named Wang as TensorDock's lead engineer.

Patel worked on model deployment from the cloud-provider side. He was a senior software development engineer on Amazon Bedrock, where his published work included optimizing and deploying custom models through OpenAI-compatible interfaces. AWS says Patel previously worked at IBM on running large AI and machine-learning workloads on Kubernetes and earned a master's degree from NYU Tandon.

The founders later overlapped at Voltage Park, according to their YC profiles. OpenRelay follows directly from that experience: Wang worked on aggregating distributed hardware, while Patel worked on turning accelerators into managed inference services.

The founders' larger bet is that AI accelerators will become interchangeable supply behind a routing layer. That abstraction is harder than it looks. A workload that runs cheaply on one chip may require different kernels, runtimes or model formats on another, and moving traffic can create cold starts or uneven latency. OpenRelay says it absorbs those differences through continuous benchmarking and automated routing.

If OpenRelay can make that machinery reliable, customers gain bargaining power across clouds and chip vendors without operating their own multi-provider stack. OpenRelay also gains control of the transaction between inference demand and unused hardware. The API is the front door; liquidity and predictable performance are the product.

Reader comments

Conversation for this story loads after sign-in.