Lamb Labsが毎秒20,000トークンを目標とするAIチップ計画を発表
Lamb Labsは、8B FPGAプロトタイプが10ワット未満で動作すると述べるが、その63倍の効率という主張には再現可能なベンチマークが欠けている。
By Ryan Merket · Published
Primary source: X
Why it matters
Sub-watt inference could move capable language models into robots and consumer devices, but Lamb Labs must translate FPGA experiments and internal metrics into tested silicon.

Lamb Labs, founded by Niki Kotecha and Thomas Lanning, said in Xでのスレッド on August 4th that its planned custom inference chips could generate more than 20,000 tokens per second and deliver 63 times the "Intelligence per Watt" of traditional GPUs.
https://x.com/LambLabs/status/2084692794388418746
Those figures are Lamb Labs targets, rather than independently tested results. Lamb Labs has not published the model, batch size, precision, latency methodology or GPU baseline behind the comparison. It has also not detailed the chip architecture, fabrication process or production schedule required to reproduce the result.
その数値は独立した試験結果ではなくLamb Labsの目標値だ。Lamb Labsは比較に使ったモデル、バッチサイズ、量子化精度、レイテンシの計測方法、GPUのベースラインを公表していない。また、結果を再現するために必要なチップアーキテクチャ、製造プロセス、量産スケジュールについても詳細を示していない。
The benchmark claims sharpen the hardware thesis that Kotecha and Lanning outlined when they brought Lamb Labs out of stealth in July. Lamb Labs is part of Y Combinator's Summer 2026 batch and is designing chips for local inference in robots, humanoids, phones, wearables and industrial equipment. The stated goal is to run capable models without a cloud connection or datacenter-scale power budget.
これらのベンチマーク主張は、KotechaとLanningが7月にLamb Labsをステルス状態から公表したときに示したハードウェア論を際立たせる。Lamb LabsはY CombinatorのSummer 2026バッチの一員であり、ロボット、ヒューマノイド、携帯電話、ウェアラブル、産業機器向けのローカル推論用チップを設計している。公表されている目標は、クラウド接続やデータセンター規模の電力予算なしで有能なモデルを動作させることだ。
Kotecha arrived at the chip project through AI research rather than semiconductor manufacturing. She earned BA and MEng degrees in chemical engineering from the University of Cambridge before pursuing a PhD at Imperial College London. Her work applied multi-agent reinforcement learning and graph neural networks to decentralized supply-chain decisions, including research on inventory control. She previously founded Transpexa, which used related methods for inventory, logistics and demand forecasting, according to an Imperial College profile.
Kotechaは半導体製造ではなくAI研究を通じてチッププロジェクトに関わるようになった。彼女はImperial College Londonで博士号を追求する前に、University of Cambridgeで化学工学のBAおよびMEngの学位を取得している。彼女の研究は、分散型のサプライチェーン意思決定に対してマルチエージェント強化学習とグラフニューラルネットワークを適用しており、在庫管理に関する研究も含まれる。Imperial Collegeのプロフィールによれば、彼女は以前にTranspexaを創業しており、在庫、物流、需要予測に関連する手法を用いていた。
Lanning's public profile says he announced plans to work as a student researcher at Google DeepMind. More recently, his updates have documented Lamb Labs' FPGA experiments, including a 7-billion-parameter model drawing about six watts and a quantized 27-billion-parameter model that fit on the same class of development board but ran slowly.
Lanningの公開プロフィールによれば、彼はGoogle DeepMindで学生研究員として働く計画を発表したという。最近のアップデートでは、Lamb LabsのFPGA実験が記録されており、約6ワットを消費する70億パラメータのモデルや、同じクラスの開発ボードに収まったが動作が遅かった量子化された270億パラメータのモデルなどが含まれている。
That prototype work shows the gap Lamb Labs still has to close. A Y Combinator post reproduced on Lamb Labs' LinkedIn page says the current 8-billion-parameter FPGA prototype operates below 10 watts. Lamb Labs is targeting deployments below one watt, a substantially tighter power envelope for battery-operated devices.
そのプロトタイプの作業は、Lamb Labsがまだ埋めるべきギャップを示している。Lamb LabsのLinkedInページに転載されたY Combinatorの投稿によれば、現在の80億パラメータのFPGAプロトタイプは10ワット未満で動作している。Lamb Labsが目指すのは1ワット未満での展開であり、これはバッテリー駆動機器にとってははるかに厳しい電力枠だ。
Y Combinator said Lamb Labs' approach combines model-specific hardware, aggressive quantization and diffusion-based inference. The accelerator described the long-term thesis as giving each AI model its own custom silicon. That would trade some of the flexibility of general-purpose GPUs for hardware tuned around a narrower model and workload.
Y Combinatorは、Lamb Labsのアプローチがモデル固有のハードウェア、積極的な量子化、拡散ベースの推論を組み合わせたものであると述べた。同アクセラレータは長期的な仮説を「各AIモデルに専用のシリコンを与えること」と説明した。それは汎用GPUの柔軟性の一部を放棄し、より狭いモデルとワークロードに合わせて調整されたハードウェアと引き換えることになる。
Lamb Labs calls its central metric Intelligence per Watt, or IPW. That framing attempts to account for model capability as well as speed and electricity use, but Lamb Labs has not published the formula used to score intelligence. Without a defined evaluation and common hardware baseline, the 63x figure cannot be compared with conventional performance-per-watt measurements.
Lamb Labsは主要な指標をIntelligence per Watt、略してIPWと呼んでいる。その枠組みはモデルの能力に加えて速度と電力使用量も考慮しようとするものだが、Lamb Labsは知能をスコア化する際に用いる計算式を公表していない。評価方法と共通のハードウェアベースラインが定義されていなければ、63倍という数値は従来のワット当たり性能の測定値と比較することはできない。
The market for specialized inference processors already spans datacenter systems and low-power devices. Groq built its LPU around deterministic execution and on-chip memory, while Cerebras sells high-speed inference on wafer-scale systems. Hailo targets edge equipment with lower-power AI accelerators. Lamb Labs is making a narrower wager around model-specific chips and power budgets small enough for embedded generative AI.
特殊化された推論プロセッサの市場はすでにデータセンターシステムと低消費電力デバイスにまたがっている。Groqは決定的実行とオンチップメモリを中心にLPUを構築し、Cerebrasはウェハースケールシステムでの高速推論を提供している。Hailoはより低消費電力のAIアクセラレータでエッジ機器をターゲットにしている。Lamb Labsはモデル固有のチップと組み込み型生成AI向けに十分小さい電力予算に賭ける、より狭い戦略を取っている。
Lamb Labs' place in YC gives Kotecha and Lanning initial financing while they work through that hardware cycle. YC's standard deal invests $500,000 through two safes: $125,000 for 7% and $375,000 on an uncapped most-favored-nation safe.
Lamb LabsがYCに参加していることは、KotechaとLanningがそのハードウェアサイクルを進める間の初期資金をもたらす。YCの標準的な契約は、2つのSAFEを通じて50万ドルを投資する:7%に対する12万5千ドルと、上限のないMost-Favored-Nation SAFEとして37万5千ドルだ。
A taped-out chip running a named model under specified conditions will determine whether Lamb Labs' figures describe a product or a design objective. The FPGA results establish that Kotecha and Lanning are running quantized language models on programmable hardware. Reproducible speed, model quality and wall-power measurements remain the proof required for the 20,000-token and 63x claims.
指定された条件下で既知のモデルを実行するテープアウト済みチップが、Lamb Labsの数値が製品を示すのか設計目標に過ぎないのかを判断するだろう。FPGAでの結果は、KotechaとLanningが量子化された言語モデルをプログラム可能なハードウェア上で動かしていることを示している。再現可能な速度、モデル品質、実際の消費電力(wall-power)の測定が、1秒あたり20,000トークンと63倍の主張を裏付けるために依然として必要な証拠である。