Head to head: Google: Gemini 3.7 Flash vs DeepSeek-V4-Flash
Google: Gemini 3.7 Flash vs DeepSeek-V4-Flash
By Ryan Merket · Published
This matchup turned on editorial discipline: source attribution, unsupported claims, format compliance, and the handling of shaky evidence. One model proved more reliable overall, but the result remained close enough to earn only a lean.
Gemini 3.7 Flash finished ahead, 50.8 to 44.0, winning four of six tasks. The statistical verdict gives it 77% confidence—a meaningful advantage, but not enough to call this a rout. Google’s model was notably better when the source packet demanded restraint. On the Gemini release, DeepSeek MI300X repository, Latitude Health funding, and AGent Energy stories, it more consistently separated reported facts from company claims, avoided broken or unverified material, and resisted inventing context. DeepSeek-V4-Flash repeatedly weakened otherwise polished copy with unsupported datelines, architecture or methodology claims, timeline errors, and unnecessary editorial notes. DeepSeek earned its two wins where completeness and format mattered more. Its MiniMax H3 response reconciled the dates more effectively, while its DeepGrove Maple-Preview story met the requested length and attributed specifications more consistently. Those wins also expose Gemini’s main weakness: it can become too terse, occasionally missing word-count requirements or sharpening qualified source language beyond what the evidence supports. **Final call: Gemini 3.7 Flash wins on a 77%-confidence lean. Its 4–2 task edge reflects stronger evidentiary discipline and cleaner news judgment, but DeepSeek-V4-Flash remains competitive when fuller treatment and strict length compliance are decisive.**
Google Gemini 3.7 Flash release
Write a 250–400-word RuntimeWire news story about Google's release of Gemini 3.7 Flash. Include a headline and dek. Explain what is changing, why the rapid replacement of Gemini 3.6 Flash matters, and summarize the performance results supplied by Ars Technica. Attribute company-reported benchmark figures clearly, preserve the distinction between reported results and independently verified performance, and do not infer technical specifications or product capabilities that the supplied materials do not establish. Do not rely on the broken Hacker News link as evidence. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 13, 2026, 3:26 PM CT) === --- Primary source: Ars Technica — "Google announces Gemini 3.7 Flash just three weeks after previous release" (published Aug 13, 2026, 12:00 PM CT) --- Google is announcing a new Gemini model today, but it's not the long-awaited 3.5 Pro. Gemini 3.7 Flash is now rolling out to replace 3.6 Flash, which itself was released only three weeks ago . This new "workhorse" model is supposedly the product of core optimizations and developer feedback, offering improved coding and agentic performance. And Google is hoping to counter the lower cost of some competing models with a lower "introductory price" for 3.7 Flash. According to Senior Director Tulsee Doshi, Gemini 3.7 Flash is noticeably better at coding than the previous Flash release. She cites a jump in the FrontierCode 1.1 Main test from 34.4 to 43.6 percent and DeepSWE v1.1 going from 49 to 65.3 percent. As for the vibes, Gemini 3.7 Flash's WebDev Arena score has risen to 1,588 from 1,538. People turning to Gemini and hoping it will "know" things may also see modest improvements in Gemini 3.7 Flash. The GDP.pdf benchmark, which measures how well a model can process complex documents, has gone up to 34 percent versus 22 percent with 3.6 Flash. AutomationBench tests how well models can execute common business workflows, and Gemini 3.7 Flash rose to 30.4 percent from 3.6's 17 percent score. Read full article Comments --- Additional source: Hacker News — "Gemini last models: temperature, top_p, and top_k are deprecated and ignored" (published Jul 21, 2026, 4:27 PM CT) --- [Skip to main content](https://ai.google.dev/gemini-api/docs/latest-model#main-content) [](https://ai.google.dev/) `/` Language - [English](https://ai.google.dev/gemini-api/docs/latest-model) - [Deutsch](https://ai.google.dev/gemini-api/docs/latest-model?hl=de) - [Español – América Latina](https://ai.google.dev/gemini-api/docs/latest-model?hl=es-419) - [Français](https://ai.google.dev/gemini-api/docs/latest-model?hl=fr) - [Indonesia](https://ai.google.dev/gemini-api/docs/latest-model?hl=id) - [Italiano](https://ai.google.dev/gemini-api/docs/latest-model?hl=it) - [Polski](https://ai.google.dev/gemini-api/docs/latest-model?hl=pl) - [Português – Brasil](https://ai.google.dev/gemini-api/docs/latest-model?hl=pt-br) - [Shqip](https://ai.google.dev/gemini-api/docs/latest-model?hl=sq) - [Tiếng Việt](https://ai.google.dev/gemini-api/docs/latest-model?hl=vi) - [Türkçe](https://ai.google.dev/gemini-api/docs/latest-model?hl=tr) - [Русский](https://ai.google.dev/gemini-api/docs/latest-model?hl=ru) - [עברית](https://ai.google.dev/gemini-api/docs/latest-model?hl=he) - [العربيّة](https://ai.google.dev/gemini-api/docs/latest-model?hl=ar) - [فارسی](https://ai.google.dev/gemini-api/docs/latest-model?hl=fa) - [हिंदी](https://ai.google.dev/gemini-api/docs/latest-model?hl=hi) - [বাংলা](https://ai.google.dev/gemini-api/docs/latest-model?hl=bn) - … --- Additional source: Hacker News — "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber" (published Jul 21, 2026, 10:17 AM CT) --- # This page doesn't exist. Let's get you back on track! Try using the search bar or [visiting our homepage](https://blog.google/). ## All stories - [**We’re announcing the Alliance for America’s Skilled Trades.**\\ ](https://blog.google/company-news/outreach-and-initiatives/creating-opportunity/alliance-america-skilled-trades/) - [**5 ways to build a side hustle with Gemini**\\ ](https://blog.google/products-and-platforms/products/gemini/launch-business-with-gemini/) - [**Designing emoji for the way we communicate today**\\ ](https://blog.google/products-and-platforms/platforms/android/world-emoji-day-noto-3d/) - [**Experience the legacy of Estadio Azteca on Google Earth.**\\ ](https://blog.google/products-and-platforms/products/earth/estadio-azteca/) - [**Beach vibes and temporary wallpaper are trending for back-to-school season.**](https://blog.google/products-and-platforms/products/shopping/back-to-school-trends/) - [**6 back-to-school shopping tricks every student … --- Prior RuntimeWire coverage --- - "Google launches Pixel 11 Pro Fold with gearless hinge and $1,899 price" (Aug 12, 2026, 11:48 AM CT): Google's lighter foldable pairs a gearless hinge with a company-claimed three-times durability gain. At $1,899, buyers will have to weigh that unverified claim against the device's premium price.
Model A clearly attributes the benchmark figures, distinguishes company claims from independent verification, and stays largely within the supplied evidence. Model B is polished but improperly incorporates the broken Hacker News material and adds unsupported claims about architecture, model stability, developer churn, benchmark scope, and the Gemini 3.5 Pro timeline. (Second judge pass, order swapped — scores are the average of both: Model A stays within the requested length, clearly labels the benchmark figures as company-reported, and avoids relying on the Hacker News material. Model B is more expansive but introduces unsupported benchmark descriptions, speculative claims about stability and developer churn, and an unnecessary API-parameter assertion drawn from the additional Hacker News material.)
minimax_h3_insufficient_support
Write a RuntimeWire news story about MiniMax's H3 model, including a headline and dek, in 250–350 words. Use only claims supported by the supplied materials, attribute company claims clearly, and do not infer details that are absent from the packet. Reconcile or flag the July 30 and Aug. 1 publication dates before publication, and verify the underlying announcement and the model's availability and open-model status. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 1, 2026, 1:02 AM CT) === --- Primary source: MiniMax (official) — "MiniMax H3: An open model breaking the boundaries between tasks and modalities https://t.co/pZxdX7ly4J" (published Aug 1, 2026, 12:50 AM CT) --- @MiniMax_AI (MiniMax (official)): MiniMax H3: An open model breaking the boundaries between tasks and modalities https://t.co/pZxdX7ly4J Links: https://twitter.com/renleanna/status/2083428921945780612 4 likes · 0 RT · 0 replies · 632 views --- Additional source: MiniMax Research — "MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities" (published Jul 30, 2026, 7:00 PM CT) --- Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length.
Model A is more cautious and avoids unsupported background details, while clearly flagging the unresolved availability and open-model claims, though it appears to fall short of the required word count. Model B better reconciles the dates and fits the requested format, but improperly adds that MiniMax is Chinese and uses an unqualified headline asserting launch and openness despite concluding those points require verification. (Second judge pass, order swapped — scores are the average of both: Model B better reconciles the dates as a research announcement followed by an official-account amplification, clearly attributes substantive claims, and explicitly concludes that availability and open-model status require verification before publication. Model B does introduce the unsupported description “Chinese AI startup,” while Model A misleadingly frames the dates as a discrepancy and may fall short of the required word count.)
DeepGrove Maple-Preview launch
Write a RuntimeWire news story about DeepGrove's launch of Maple-Preview, including a headline and dek, in 250–400 words. Use only claims supported by the supplied materials, clearly attribute company-reported figures, and preserve meaningful qualifications about the model's preview status and evaluation limits. Do not present promotional benchmark or device-performance claims as independently established facts. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 4, 2026, 6:18 PM CT) === --- Primary source: Hacker News — "Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone" (published Aug 4, 2026, 2:44 PM CT) --- Today we introduce **Maple-Preview**, an **open-source 20B-A1B ternary-weight** reasoning LLM. Maple-Preview is SOTA in its weight class and is even competitive with larger models. It solves IMO-level problems and runs at **200+ tokens/s** on a **Mac mini M4**, **5–16×** faster than efficient models like Gemma 4, Qwen3.5, and gpt-oss. [Chat with Maple-Preview →](https://chat.deepgrove.ai/) - **20B-A1B** Model - **218 tokens/s** M4 Mac mini - **5.31 GB** Checkpoint - **131,072** Token context  ### Maple-Preview solves [IMO 2024 P1](https://web.evanchen.cc/exams/IMO-2024-notes.pdf) (7/7) at 281.5 tokens/s on a MacBook Pro (M5 Pro). Maple-Preview earning 7/7 on IMO 2024 Problem 1 on a MacBook Pro (M5 Pro), running at 281.5 tokens/s. ### Maple-Preview runs at 127 tokens/s—13× faster than 1-bit Bonsai 27B (Qwen3.6 27B) on an iPhone **Maple-Preview** ≈ 00:09 **Bonsai 27B** \> 05:58 Maple Maple-Preview (127 tokens/s) vs. Bonsai 27B (9.6 tokens/s) on iPhone. Both receive the same prompt, “Make me a carrot cake.” Maple-Preview completes its response in about 10 seconds; Bonsai is still generating at the six-minute mark. ## Our Bet We envision a shift from the age of monolithic LLMs to **always-active on-device assistants, which continuously shape themselves to improve user experience.** We believe that this necessitates a move towards more efficient and performant architectures, such that these assistants can both be trained and run on everyday devices (e.g. laptops and phones). As exemplified in Maple-Preview, we view ultra-low precision as being crucial to this new era of efficient, performant modeling. At ultra-low bitwidths, matrix multiplication can be effectively replaced with additions, lowering the total arithmetic workload needed to infer through a model. With native support for low bitwidths and custom hardware that takes advantage of both the memory and arithmetic efficiency of such architectures, we find it easy to imagine a world where most inference is done on smaller, personalized models and infrequent, exceedingly difficult tasks are offloaded to cloud models. Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach; we think that this unnecessarily limits both performance (to being only some percentage of the full-precision model) and efficiency (by enforcing architectural constraints that may not be beneficial). We posit that in many ways, creating a high-performing, efficient model should be similar to creating simply a high-performing model. As such, instead of focusing on how to make a performant model efficient, we believe it is most fruitful to dedicate substantial effort toward working on optimization, data, and more to improve model performance while simultaneously enforcing efficiency through an inference-aware architectural design loop, pushing the frontier in both directions. **We believe that the precision a model runs at should be the precision it learns at.** At DeepGrove, we study how efficient models learn, rethinking architecture, training infrastructure, optimization, and hardware design to treat low precision as a first-class citizen. Maple-Preview is a natively trained ternary-weight network showing that low precision does not have to mean compromise. ## Architecture Maple-Preview is a 20B-A1B reasoning model designed from the start for efficient on-device reasoning. We designed the Maple architecture in a hardware-aware manner, testing all considered configurations on our Mac mini for inference speed. Starting with a 30-layer, 224-expert configuration, we optimized to a 24-layer, 256-expert configuration as a compromise between model performance and inference efficiency. We additionally chose hybrid sliding-window and global attention to bound KV-cache growth.   ## Evaluation On benchmarks, **Maple-Preview sets a new point on the Pareto frontier for both memory-to-performance and speed-to-performance,** demonstrating its strong reasoning capabilities. However, we note that this preview is focused primarily on raw reasoning and, as such, may underperform on agentic benchmarks. We intend to continue improving general performance through extended training before Maple's full release.
Model B better satisfies the 250–400-word requirement, consistently attributes performance claims, and preserves the preview and evaluation caveats, though its unsupported San Francisco dateline should be removed. Model A is concise and generally accurate but appears to fall short of the requested word count and overstates the source slightly by saying the preview is focused “strictly” rather than “primarily” on raw reasoning. (Second judge pass, order swapped — scores are the average of both: Model B more consistently attributes reported specifications and performance figures while preserving the preview and agentic-benchmark qualifications, though its unsupported San Francisco dateline and one ambiguous sentence are flaws. Model A is polished but presents the checkpoint size and context length without attribution and overstates the model’s focus as “strictly” raw reasoning when the source says “primarily.”)
Latitude Health Form D funding disclosure
Write a 250–350-word RuntimeWire news story about Latitude Health's newly disclosed financing. Include a headline and dek. Base the story on the SEC Form D and clearly distinguish filing facts from company-reported product claims. State the amount sold, total offering size, number of participating investors, filing date, company location and leadership where relevant. Do not identify investors or characterize the securities beyond what the filing establishes. Explain what the company does and note that its automation claims are claims by the company, not independently verified. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 7, 2026, 4:53 PM CT) === --- Primary source: SEC Form D — "Latitude Health raises $2M for AI-native utilization management" (published Aug 7, 2026, 2:50 PM CT) --- Latitude Health, Inc. filed SEC Form D on August 7, 2026, disclosing $2,000,000 in securities sold toward a $3,000,000 total offering. The Delaware corporation is based in San Francisco, CA. Latitude Health operates an AI-native platform for utilization management (UM) and prior authorization, claiming to automate 75% of manual work, double review throughput, and reduce clinician burnout. The company appears to have emerged from stealth around early 2025 and is led by Charles Feerick (Executive Officer and Director), with Jarred Bressner and Chris Palmieri also serving as directors. The filing shows only 2 investors have participated so far. The company's website is latitudehealth.com. This is the first concrete funding disclosure for the company. Source: SEC EDGAR Form D filing 0002148343-26-000001
Model A is more accurate, includes both a headline and dek, and cleanly separates filing facts from unverified company claims, though it falls short of the required word count. Model B meets the length requirement but lacks a dek, adds the unsupported assertion that the securities were sold on August 7, and uses a repetitive editor’s note that weakens the news style. (Second judge pass, order swapped — scores are the average of both: Model A better satisfies the requested format with a clear headline and dek, stays within the target length, and carefully separates SEC filing facts from unverified company claims. Model B lacks a dek and adds the unsupported assertion that the securities were sold on the filing date, along with an unnecessary editor’s note.)
deepseek-v4-flash-mi300x-unverified-repository
Write a 250–350-word RuntimeWire news story about the newly published GitHub repository and what it claims about running DeepSeek V4 Flash on an AMD MI300X. Include a headline and dek. Attribute claims clearly to the repository, distinguish reported results from independently verified facts, and use only details supported by the supplied materials. Assess whether the available evidence is sufficient for publication; do not infer the author's employment, affiliations, model specifications, deployment status, or benchmark validity beyond the sources. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 4, 2026, 6:08 AM CT) === --- Primary source: Hacker News — "DeepSeek V4 Flash on a Single AMD MI300X" (published Aug 4, 2026, 5:00 AM CT) --- [Skip to content](https://github.com/ryanzhou/deepseek-v4-flash-mi300x#start-of-content) You signed in with another tab or window. [Reload](https://github.com/ryanzhou/deepseek-v4-flash-mi300x) to refresh your session.You signed out in another tab or window. [Reload](https://github.com/ryanzhou/deepseek-v4-flash-mi300x) to refresh your session.You switched accounts on another tab or window. [Reload](https://github.com/ryanzhou/deepseek-v4-flash-mi300x) to refresh your session.Dismiss alert {{ message }} [ryanzhou](https://github.com/ryanzhou)/ **[deepseek-v4-flash-mi300x](https://github.com/ryanzhou/deepseek-v4-flash-mi300x)** Public - [Notifications](https://github.com/login?return_to=%2Fryanzhou%2Fdeepseek-v4-flash-mi300x) You must be signed in to change notification settings - [Fork\\ 0](https://github.com/login?return_to=%2Fryanzhou%2Fdeepseek-v4-flash-mi300x) - [Star\\ 0](https://github.com/login?return_to=%2Fryanzhou%2Fdeepseek-v4-flash-mi300x) main [**1** Branch](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/branches) [**0** Tags](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tags) [Go to Branches page](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/branches)[Go to Tags page](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tags) Go to file Code Open more actions menu ## Folders and files | Name | Name | Last commit message | Last commit date | | --- | --- | --- | --- | | ## Latest commit<br>[](https://github.com/ryanzhou)[ryanzhou](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commits?author=ryanzhou)<br>[Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0)<br>Open commit details<br>22 minutes agoAug 4, 2026<br>[7c06e57](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0) · 22 minutes agoAug 4, 2026<br>## History<br>[1 Commit](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commits/main/) <br>Open commit details<br>[View commit history for this file.](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commits/main/) 1 Commit | | [patches](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/patches "patches") | [patches](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/patches "patches") | [Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0 "Open-source DeepSeek V4 Flash on a single MI300X production stack Production-validated stack for deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X (gfx942): digest-pinned vLLM ROCm compose deployment, ten byte-identical correctness/performance overlays with SHA-256 pins, reference unified diffs against upstream vLLM/ROCm-triton/AITER, and AITER GEMM tuning tables. Includes blog-style README with measured results (168.6 tok/s median single-stream decode; 7.9-8.5K tok/s tuned prefill).") | 22 minutes agoAug 4, 2026 | | [tuning](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/tuning "tuning") | [tuning](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/tuning "tuning") | [Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0 "Open-source DeepSeek V4 Flash on a single MI300X production stack Production-validated stack for deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X (gfx942): digest-pinned vLLM ROCm compose deployment, ten byte-identical correctness/performance overlays with SHA-256 pins, reference unified diffs against upstream vLLM/ROCm-triton/AITER, and AITER GEMM tuning tables. Includes blog-style README with measured results (168.6 tok/s median single-stream decode; 7.9-8.5K tok/s tuned prefill).") | 22 minutes agoAug 4, 2026 | | [.gitignore](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/blob/main/.gitignore ".gitignore") | [.gitignore](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/blob/main/.gitignore ".gitignore") | [Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0 "Open-source DeepSeek V4 Flash on a single MI300X production stack Production-validated stack for deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X (gfx942): digest-pinned vLLM ROCm compose deployment, ten byte-identical correctness/performance overlays with SHA-256 pins, reference unified diffs against upstream vLLM/ROCm-triton/AITER, and AITER GEMM tuning tables. --- Additional source: Hacker News — "Bringing Up DeepSeek-V4-Flash on AMD MI300X" (published Jun 2, 2026, 12:52 PM CT) --- # Bringing up DeepSeek-V4-Flash on AMD MI300X 1 Jun 2026 9 min read At [Doubleword](https://app.doubleword.ai/) we are building an inference cloud designed for volume. To do that we have to reckon with the enveloping compute shortage. AMD’s MI300X launched in December 202311.At AMD’s [“Advancing AI” event](https://www.amd.com/en/newsroom/press-releases/2023-11-15-amd-announces-amd-instinct-mi300-accelerator-launc.html), 6 December 2023. as AMD’s response to NVIDIA’s H100, arriving alongside H200 in the same generation. It is an odd duck in the world of high-end AI accelerators. While H100 prices are climbing (up 40% in five months on one-year rentals, with on-demand capacity sold out across every major NVIDIA part22.SemiAnalysis, [The Great GPU Shortage: Rental Capacity](https://newsletter.semianalysis.com/p/the-great-gpu-shortage-rental-capacity), April 2026.), MI300X is perhaps still underappreciated. 192GB of HBM3 per card against the H100’s 80GB, comparable FP8 compute, list price roughly half. Yet you can rent one on-demand today (from [Hotaisle](https://www.hotaisle.ai/), for instance) for noticeably less than the equivalent NVIDIA capacity. The reason is software.
Model A stays within the requested format and length, attributes the performance claims carefully, and gives a sound verification assessment, though its dek slightly mischaracterizes the 168.6 tokens-per-second result as an “up to” figure. Model B exceeds the word limit and adds unsupported details or inferences, including a San Francisco dateline, “latest model,” lack of affiliation between sources, and claims about missing repository methodology that cannot be established from the supplied excerpt. (Second judge pass, order swapped — scores are the average of both: Model A stays within the requested length, clearly attributes the repository’s claims, and appropriately concludes that independent reproduction and technical review are needed. Model B exceeds the word limit and introduces unsupported or improper assertions, including a San Francisco dateline, claimed non-affiliation, and a purported RuntimeWire verification effort not established by the materials.)
AGent Energy raises $11 million Series Seed
Write a 250–350-word RuntimeWire news story about AGent Energy’s newly announced financing, including a clear headline and dek. Report the $11 million Series Seed and the company’s stated $17 million in total funding, identify the co-lead investors and other participants, and explain what AGent says it plans to do with the capital. Attribute company claims about behind-the-meter generation, grid conditions, market participation and emergency-event dispatches rather than presenting promotional assertions as independently verified facts. Make clear that the 200+ GW figure is AGent’s target or estimate of behind-the-meter generation, not disclosed capacity already controlled by the company. Do not add biographical details, customer figures, revenue figures, site counts or regulatory claims not supported by the packet. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 13, 2026, 5:29 PM CT) === --- Primary source: PR Newswire Business Technology — "AGent Energy Closes Series Seed to Unlock 200+ GW of Behind-the-Meter Generation Across Commercial, Industrial, and Institutional Sectors" (published Aug 13, 2026, 12:00 PM CT) --- [Accessibility Statement](https://www.cision.com/about/accessibility/) [Skip Navigation](https://www.prnewswire.com/news-releases/agent-energy-closes-series-seed-to-unlock-200-gw-of-behind-the-meter-generation-across-commercial-industrial-and-institutional-sectors-302851086.html#main) _Round Co-Led by Spero Ventures and MassMutual Ventures with Participation from Intrepid Investment Management and Existing Investors Zero Infinity Partners (ZIP) and CIV; Brings Total Funding to $17 Million in Just 12 Months, Making It One of the Fastest-Funded Distributed Energy Resource Companies to Date_ HOUSTON, Aug. 13, 2026 /PRNewswire/ -- AGent Energy, a trailblazing developer of AI-driven distributed power plants, today announced it has closed an $11 million Series Seed financing co-led by Spero Ventures and MassMutual Ventures, with participation from Intrepid Investment Management and existing investors CIV and Zero Infinity Partners (ZIP). The round follows a $6 million financing from CIV and ZIP, which closed within two months of founding, bringing AGent's total funding to $17 million in its first 12 months. It's a striking vote of confidence in behind-the-meter generation as the next great frontier of U.S. energy infrastructure. America's grid is under mounting strain. PJM's most recent capacity auction cleared at the price cap without enough capacity to meet demand, and data center load growth is outpacing new supply across every major market. AGent is unlocking a faster, smarter way to keep the power flowing: the backup generation that already sits at commercial, industrial, and mission-critical facilities, including AI data centers. AGent's AI-based platform aggregates, orchestrates, and monetizes these assets, turning them into rapidly dispatchable, highly reliable distributed power plants. Because the equipment is already built, already paid for, and idle most of the year, AGent delivers capacity at the lowest cost of any new grid resource, at zero cost to the asset owner, who earns new revenue instead. AGent is already dispatching in three of the largest wholesale markets in North America, having successfully delivered capacity during grid emergency events in PJM, MISO, and ERCOT. AGent will use the new capital to aggressively scale its team and accelerate its push to unlock 200+ GW of behind-the-meter generation across commercial and industrial facilities and the MUSH sector (municipalities, universities, schools, and hospitals), the properties with the highest concentration of on-site generation and the ones AGent's team knows best. "Twelve months ago, AGent was an idea backed by two investors who saw what we saw: 200+ GW of the most reliable generation in America sitting idle behind the meter," said Stephanie Hendricks, CEO and Co-Founder of AGent. "Closing $17 million in our first year reflects both the urgency of the grid reliability challenge and the speed at which our team executes. With Spero, MassMutual Ventures, and Intrepid joining CIV and ZIP, we now have the partners and capital to bring this dispatchable capacity to the markets that need it most." "The grid doesn't need to wait five years for new steel in the ground. The capacity is already there, and AGent has built the intelligence layer to unlock it. Stephanie and her team have moved faster than any company we've seen in this space, and we're proud to co-lead their Series Seed round," said Stephen Wemple of Spero Ventures. "When the grid is stressed, the difference between a rolling blackout and an ordinary afternoon is how fast dispatchable capacity shows up. AGent's AI platform turns generation already sitting at hospitals, universities, and industrial sites into exactly that: capacity that responds in minutes, with no new construction and no cost to the owner. It's a rare combination of climate resilience and hard economics, built by a team that has done this at scale before. That's the sweet spot for our Climate Technology Fund: AI applied to real assets, with economics that pencil from day one," said Aram Ouligian, Senior Associate at MassMutual Ventures. In connection with the financing, Stephen Wemple of Spero Ventures will join AGent's board of directors, and Aram Ouligian of MassMutual Ventures will join as a board observer. **About AGent, Inc.** Using its AI-based technology platform, AGent aggregates, orchestrates, and monetizes distributed generation resources in power markets, delivering significant reliability, economic, and sustainability benefits to large energy users, utilities, and grid operators. To learn more visit agentenergy.com. **About Spero Ventures** Spero Ventures is an early-stage venture capital firm that invests in founders building a future that belongs to everyone, with a focus on sustainable systems, health and longevity, and human potential. Learn more at spero.vc.
Model A is accurate, polished and properly attributes AGent’s claims while clearly presenting 200-plus GW as a target rather than controlled capacity. Model B incorrectly says the announcement came Wednesday, uses a more promotional headline, and ends with an unnecessary editorial note that weakens the news-story format. (Second judge pass, order swapped — scores are the average of both: Model A is more polished, stays within the requested news format, clearly frames the 200-plus GW figure as a target rather than controlled capacity, and consistently attributes operational and grid claims. Model B is largely compliant but incorrectly says the Thursday announcement occurred Wednesday, uses awkward phrasing in places, and adds an unnecessary editorial disclaimer.)
Matchup powered by OpenRouter.