MiniMax schedules H3 consumer-hardware demo as performance remains unverified
The August 7 session promises open H3 weights, stereo audio and ready-made workflows, but no official benchmark yet identifies the hardware, memory use or generation speed.
By Ryan Merket · Published
Why it matters
H3 could put synchronized audio-video generation on personal machines, but developers still need reproducible GPU, memory and runtime measurements before they can assess its cost or compare it with Wan, HunyuanVideo and LTX-2.

MiniMax, founded and led by Junjie Yan, scheduled an August 7, 2026 livestream with ComfyUI to demonstrate its H3 video model on consumer hardware. The session was set for 10 a.m. Pacific Time.
The supplied announcement says MiniMax will show "how H3 got small enough to run on consumer hardware" and provide ready-to-use ComfyUI workflow templates. That establishes a promised demonstration, not verified performance. No official recording or logs available in the reviewed sources identify the GPU, VRAM, system memory, quantization settings, offloading configuration, generation time or whether a 2K clip was generated locally.
The event extends Yan's push to make advanced models cheaper and easier to distribute. Yan, MiniMax's chairman, CEO and CTO, joined SenseTime in 2015 and later became a vice president before founding MiniMax. He studied mathematics at Southeast University, earned an artificial-intelligence doctorate from the Chinese Academy of Sciences and conducted postdoctoral research at Tsinghua University, according to his MiniMax management biography.
In a 2025 keynote, Yan said open-source innovation was catching up, inference costs were falling and useful AI needed to become affordable and accessible. MiniMax, founded in December 2021 and based in Shanghai, develops models and products spanning language, video, audio and music.
The hardware details will decide what "consumer" means
The research brief supports this narrower description: H3 is an open-weight multimodal video-generation model with native stereo audio, output options up to 2K, and ComfyUI workflow support. MiniMax's platform documentation lists text-to-video, image-to-video, first-and-last-frame generation and multimodal reference inputs, with clips ranging from four to 15 seconds at 24 frames per second.
ComfyUI's official text-to-video workflow materials give users ready-made H3 workflows for local generation. The available materials do not provide a completed local 2K benchmark or the hardware used to run one.
Those workflow files make the model inspectable and configurable. They do not show which personal computers can produce clips at useful speeds. "Consumer hardware" could cover systems ranging from high-memory desktop GPUs to laptops using aggressive quantization and CPU offloading. Without a named machine, visible memory consumption and measured runtime, developers cannot reproduce MiniMax's performance claim or estimate the cost of running H3 themselves.
ComfyUI's involvement gives the test practical weight. Its interface represents generation pipelines as connected nodes, exposing model loading, prompts, reference media, resolution and output handling. Comfy Org co-founder and CEO Yoland Yan, a former Google Search machine-learning engineer and Chromium committer, has helped develop the project beyond its original image-generation use case into video, audio and 3D workflows.
Comfy has also attracted institutional backing around that approach. Comfy Org raised $17 million in 2025, followed by a $30 million round in April 2026 at a reported $500 million valuation. Reported investors include Craft Ventures, Pace Capital, Chemistry, Abstract Ventures, Cursor Capital, Essence VC, Stratus Ventures, TruArrow and Vercel founder Guillermo Rauch.
H3 enters a crowded local-video contest
MiniMax is competing with several open or downloadable video models whose developers publish more concrete memory targets. Alibaba's Wan2.2 includes a 5 billion-parameter text-and-image-to-video model targeting GPUs with at least 24GB of VRAM. Tencent's HunyuanVideo-1.5 uses an 8.3 billion-parameter architecture and documents a 14GB VRAM minimum when offloading is enabled. Lightricks' LTX-2 combines separate 14 billion-parameter video and 5 billion-parameter audio streams.
H3's announced combination of open weights, stereo sound and clips up to 2K gives MiniMax a distinct technical pitch. The comparison remains incomplete until MiniMax publishes enough runtime detail to measure H3 against those alternatives. Parameter count alone does not determine memory use or generation speed; model precision, attention implementation, encoder loading and offloading choices can materially change the result.
MiniMax has ample capital to support the model program. Alibaba led a financing round of more than $600 million in March 2024 at an approximate $2.5 billion valuation. MiniMax's listing materials identify additional backers including Tencent-related entities, MiHoYo, Hillhouse, HongShan, IDG Capital, Yunqi Capital, GL Ventures and Shanghai state-backed capital. The company later raised approximately $619 million in its January 2026 Hong Kong initial public offering, according to its listing prospectus.
That funding gives Yan a distribution base for H3. The unresolved issue is execution on personal machines. A reproducible benchmark needs the demonstrated GPU, VRAM and system-memory use, model precision, offloading settings, generation time, resolution and clip length. The reviewed sources establish H3's announced capabilities and downloadable ComfyUI workflow. They do not establish its actual consumer-hardware performance.