MiniMaxは、性能が未検証のままH3の消費者向けハードウェアのデモを予定している。

8月7日のセッションはオープンなH3ウェイト、ステレオ音声、既製のワークフローを約束しているが、公式のベンチマークはまだハードウェア、メモリ使用量、生成速度を特定していない。

By · Published

Primary source: MiniMax

Why it matters

H3 could put synchronized audio-video generation on personal machines, but developers still need reproducible GPU, memory and runtime measurements before they can assess its cost or compare it with Wan, HunyuanVideo and LTX-2.

An embroidered rendition of a graphical processing 'node' or 'module' (Embroidered textile patch - visible thread stitches, felt layers, slightly puffy dimensional fabric)

MiniMax, founded and led by Junjie Yan, scheduled an August 7, 2026 livestream with ComfyUI to demonstrate its H3 video model on consumer hardware. The session was set for 10 a.m. Pacific Time.

MiniMax の X上の投稿

ComfyUI の X上の投稿

The supplied announcement says MiniMax will show "how H3 got small enough to run on consumer hardware" and provide ready-to-use ComfyUI workflow templates. That establishes a promised demonstration, not verified performance. No official recording or logs available in the reviewed sources identify the GPU, VRAM, system memory, quantization settings, offloading configuration, generation time or whether a 2K clip was generated locally.

The event extends Yan's push to make advanced models cheaper and easier to distribute. Yan, MiniMax's chairman, CEO and CTO, joined SenseTime in 2015 and later became a vice president before founding MiniMax. He studied mathematics at Southeast University, earned an artificial-intelligence doctorate from the Chinese Academy of Sciences and conducted postdoctoral research at Tsinghua University, according to his MiniMax management biography.

In a 2025年基調講演, Yan said open-source innovation was catching up, inference costs were falling and useful AI needed to become affordable and accessible. MiniMax, founded in December 2021 and based in Shanghai, develops models and products spanning language, video, audio and music.

ハードウェアの詳細が「消費者向け」の定義を決める

The research brief supports this narrower description: H3 is an open-weight multimodal video-generation model with native stereo audio, output options up to 2K, and ComfyUI workflow support. MiniMax's プラットフォームのドキュメント lists text-to-video, image-to-video, first-and-last-frame generation and multimodal reference inputs, with clips ranging from four to 15 seconds at 24 frames per second.

ComfyUI's official テキスト→ビデオのワークフロ資料 give users ready-made H3 workflows for local generation. The available materials do not provide a completed local 2K benchmark or the hardware used to run one.

Those workflow files make the model inspectable and configurable. They do not show which personal computers can produce clips at useful speeds. "Consumer hardware" could cover systems ranging from high-memory desktop GPUs to laptops using aggressive quantization and CPU offloading. Without a named machine, visible memory consumption and measured runtime, developers cannot reproduce MiniMax's performance claim or estimate the cost of running H3 themselves.

ComfyUI's involvement gives the test practical weight. Its interface represents generation pipelines as connected nodes, exposing model loading, prompts, reference media, resolution and output handling. Comfy Org co-founder and CEO Yoland Yan, a former Google Search machine-learning engineer and Chromium committer, has helped develop the project beyond its original image-generation use case into video, audio and 3D workflows.

Comfy has also attracted institutional backing around that approach. Comfy Org raised $17 million in 2025, followed by a 2026年4月の$30 millionのラウンド at a reported $500 million valuation. Reported investors include Craft Ventures, Pace Capital, Chemistry, Abstract Ventures, Cursor Capital, Essence VC, Stratus Ventures, TruArrow and Vercel founder Guillermo Rauch.

H3が混戦のローカル動画競争に参入

MiniMax is competing with several open or downloadable video models whose developers publish more concrete memory targets. Alibaba's Wan2.2 includes a 5 billion-parameter text-and-image-to-video model targeting GPUs with at least 24GB of VRAM. Tencent's HunyuanVideo-1.5 uses an 8.3 billion-parameter architecture and documents a 14GB VRAM minimum when offloading is enabled. Lightricks' LTX-2 combines separate 14 billion-parameter video and 5 billion-parameter audio streams.

H3's announced combination of open weights, stereo sound and clips up to 2K gives MiniMax a distinct technical pitch. The comparison remains incomplete until MiniMax publishes enough runtime detail to measure H3 against those alternatives. Parameter count alone does not determine memory use or generation speed; model precision, attention implementation, encoder loading and offloading choices can materially change the result.

MiniMax has ample capital to support the model program. Alibaba led a financing round of more than $600 million in March 2024 at an approximate $2.5 billion valuation. MiniMax's 上場目論見書 identify additional backers including Tencent-related entities, MiHoYo, Hillhouse, HongShan, IDG Capital, Yunqi Capital, GL Ventures and Shanghai state-backed capital. The company later raised approximately $619 million in its January 2026 Hong Kong initial public offering, according to its listing prospectus.

That funding gives Yan a distribution base for H3. The unresolved issue is execution on personal machines. A reproducible benchmark needs the demonstrated GPU, VRAM and system-memory use, model precision, offloading settings, generation time, resolution and clip length. The reviewed sources establish H3's announced capabilities and downloadable ComfyUI workflow. They do not establish its actual consumer-hardware performance.

Reader comments

Conversation for this story loads after sign-in.