Interfaze releases open-weight model for OCR, speech and structured extraction

Founded by repeat entrepreneur Yoeven Khemlani and computer-vision researcher Harsha Khurdula, the model is available under Apache 2.0 and is designed to run on one 80GB GPU.

By · Published

Primary source: X

Why it matters

Interfaze is selling a model designed to expose evidence alongside answers, giving developers a way to route uncertain document and audio outputs for review instead of treating fluent output as verified data.

A generic GPU accelerator sits beside a page in a document scanner and a small microphone, representing Interfaze’s OCR, speech, and extraction model.

Interfaze released interfaze-1-lite on October 5th, an open-weight model for document OCR, speech-to-text, structured extraction and other tasks where software needs outputs it can check. The San Francisco startup says the model returns confidence scores and location metadata, including bounding boxes, alongside its answers. The weights are available on Hugging Face under the Apache 2.0 license.

https://x.com/interfaze_ai/status/2107144505711595575

Video from the original post on X.

The launch reflects the thesis of co-founder and CEO Yoeven D. Khemlani (@yoeven), a repeat founder who previously built JigsawStack, the company now known as Interfaze. Khemlani has described the problem behind that work as a gap between AI that can converse and AI reliable enough to operate inside software. In a post about his path from an earlier startup to JigsawStack, he said the experience of selling a travel startup and stepping away from company-building did not last: he returned to building after running into the limits of general-purpose models on developer tasks. Interfaze's co-founder and CTO, Harsha Vardhan Khurdula (@khurdula), brings a research background in computer vision and AI, according to Y Combinator's company profile.

Interfaze calls the design Mixture of Architecture, or MoA. Rather than asking one general model to handle every input, the system pairs a reasoning core with task-specific specialist models. The core interprets the request and coordinates specialists; those models handle jobs such as reading documents, recognizing speech and locating objects. The combination is intended to provide schema-following, contextual reasoning and evidence metadata of the kind specialized vision and audio systems can return.

For a document workflow, that can mean extracting transactions into a fixed JSON schema while returning the OCR text, its position on the page and a confidence score. Interfaze's own example uses a confidence threshold to route uncertain rows to a human reviewer instead of automatically posting them. A confidence score alone does not establish that a system is reliable: developers still need to test how scores behave on their own documents and decide where the review threshold belongs.

The model accepts text, images, audio and files, including PDFs and Word documents, and can return text or JSON. Interfaze lists OCR, speaker diarization, classification, structured extraction, object detection, GUI detection, translation, forecasting and guardrails among its capabilities. Its company announcement says the full model runs on a single 80GB GPU, such as an H100; its setup instructions specify a GPU with compute capability 8.9 or newer. Teams need suitable hardware to self-host it.

Interfaze also sells API access, priced at $0.85 per million input tokens and $1.50 per million output tokens, according to the launch post. The model supports a 128,000-token context window and up to 32,000 output tokens. Developers can use the hosted API or load the weights through Transformers; the GitHub repository and model card provide setup details.

Interfaze's published benchmark table reports wins and losses across the compared models. The company reports interfaze-1-lite ahead of the four compared commercial models on MMMU-Pro multimodal reasoning, RefCOCO visual grounding, structured-output value accuracy, document processing on olmOCR and speech recognition on the VoxPopuli subset. It also reports weaker scores than its own larger interfaze-1 model on OCRBench V2, text-to-SQL, science questions and multilingual Q&A. On some tasks, its reported advantage over competitors is narrow: its 83.8% olmOCR score is just above Claude-Sonnet-5's 83.5%, while its 81.5% structured-output score is above Gemini-3.7-Flash's 80.2%. The company's leaderboard says results combine independently run evaluations and model-provider data; some scores are self-reported. These comparisons show performance on the listed tests, but do not establish that the model will outperform alternatives on every production workload.

The product bet is that developers will pay for reliable, inspectable outputs in workflows where a wrong character can break a downstream process. Interfaze offers the model as both a download and a metered API, giving teams a choice between self-hosting and a managed service. Its October 5th release makes the weights public. Whether confidence and location data reduce the human checking and model-stitching its target users already do remains to be seen.

Reader comments

Conversation for this story loads after sign-in.