Head to head: Google: Gemini 3.7 Flash vs DeepSeek-V4-Flash

Google: Gemini 3.7 Flash vs DeepSeek-V4-Flash

By · Published

RuntimeWire Head-to-Head: Head to head: Google: Gemini 3.7 Flash vs DeepSeek-V4-Flash
RuntimeWire Head-to-Head matchup

This matchup turned on editorial discipline: source attribution, unsupported claims, format compliance, and the handling of shaky evidence. One model proved more reliable overall, but the result remained close enough to earn only a lean.

Gemini 3.7 Flash finished ahead, 50.8 to 44.0, winning four of six tasks. The statistical verdict gives it 77% confidence—a meaningful advantage, but not enough to call this a rout. Google’s model was notably better when the source packet demanded restraint. On the Gemini release, DeepSeek MI300X repository, Latitude Health funding, and AGent Energy stories, it more consistently separated reported facts from company claims, avoided broken or unverified material, and resisted inventing context. DeepSeek-V4-Flash repeatedly weakened otherwise polished copy with unsupported datelines, architecture or methodology claims, timeline errors, and unnecessary editorial notes. DeepSeek earned its two wins where completeness and format mattered more. Its MiniMax H3 response reconciled the dates more effectively, while its DeepGrove Maple-Preview story met the requested length and attributed specifications more consistently. Those wins also expose Gemini’s main weakness: it can become too terse, occasionally missing word-count requirements or sharpening qualified source language beyond what the evidence supports. **Final call: Gemini 3.7 Flash wins on a 77%-confidence lean. Its 4–2 task edge reflects stronger evidentiary discipline and cleaner news judgment, but DeepSeek-V4-Flash remains competitive when fuller treatment and strict length compliance are decisive.**

Google Gemini 3.7 Flash release

Write a 250–400-word RuntimeWire news story about Google's release of Gemini 3.7 Flash. Include a headline and dek. Explain what is changing, why the rapid replacement of Gemini 3.6 Flash matters, and summarize the performance results supplied by Ars Technica. Attribute company-reported benchmark figures clearly, preserve the distinction between reported results and independently verified performance, and do not infer technical specifications or product capabilities that the supplied materials do not establish. Do not rely on the broken Hacker News link as evidence. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 13, 2026, 3:26 PM CT) === --- Primary source: Ars Technica — "Google announces Gemini 3.7 Flash just three weeks after previous release" (published Aug 13, 2026, 12:00 PM CT) --- Google is announcing a new Gemini model today, but it's not the long-awaited 3.5 Pro. Gemini 3.7 Flash is now rolling out to replace 3.6 Flash, which itself was released only three weeks ago . This new "workhorse" model is supposedly the product of core optimizations and developer feedback, offering improved coding and agentic performance. And Google is hoping to counter the lower cost of some competing models with a lower "introductory price" for 3.7 Flash. According to Senior Director Tulsee Doshi, Gemini 3.7 Flash is noticeably better at coding than the previous Flash release. She cites a jump in the FrontierCode 1.1 Main test from 34.4 to 43.6 percent and DeepSWE v1.1 going from 49 to 65.3 percent. As for the vibes, Gemini 3.7 Flash's WebDev Arena score has risen to 1,588 from 1,538. People turning to Gemini and hoping it will "know" things may also see modest improvements in Gemini 3.7 Flash. The GDP.pdf benchmark, which measures how well a model can process complex documents, has gone up to 34 percent versus 22 percent with 3.6 Flash. AutomationBench tests how well models can execute common business workflows, and Gemini 3.7 Flash rose to 30.4 percent from 3.6's 17 percent score. Read full article Comments --- Additional source: Hacker News — "Gemini last models: temperature, top_p, and top_k are deprecated and ignored" (published Jul 21, 2026, 4:27 PM CT) --- [Skip to main content](https://ai.google.dev/gemini-api/docs/latest-model#main-content) [![Gemini API](https://ai.google.dev/_static/googledevai/images/gemini-api-logo.svg)](https://ai.google.dev/) `/` Language - [English](https://ai.google.dev/gemini-api/docs/latest-model) - [Deutsch](https://ai.google.dev/gemini-api/docs/latest-model?hl=de) - [Español – América Latina](https://ai.google.dev/gemini-api/docs/latest-model?hl=es-419) - [Français](https://ai.google.dev/gemini-api/docs/latest-model?hl=fr) - [Indonesia](https://ai.google.dev/gemini-api/docs/latest-model?hl=id) - [Italiano](https://ai.google.dev/gemini-api/docs/latest-model?hl=it) - [Polski](https://ai.google.dev/gemini-api/docs/latest-model?hl=pl) - [Português – Brasil](https://ai.google.dev/gemini-api/docs/latest-model?hl=pt-br) - [Shqip](https://ai.google.dev/gemini-api/docs/latest-model?hl=sq) - [Tiếng Việt](https://ai.google.dev/gemini-api/docs/latest-model?hl=vi) - [Türkçe](https://ai.google.dev/gemini-api/docs/latest-model?hl=tr) - [Русский](https://ai.google.dev/gemini-api/docs/latest-model?hl=ru) - [עברית](https://ai.google.dev/gemini-api/docs/latest-model?hl=he) - [العربيّة](https://ai.google.dev/gemini-api/docs/latest-model?hl=ar) - [فارسی](https://ai.google.dev/gemini-api/docs/latest-model?hl=fa) - [हिंदी](https://ai.google.dev/gemini-api/docs/latest-model?hl=hi) - [বাংলা](https://ai.google.dev/gemini-api/docs/latest-model?hl=bn) - … --- Additional source: Hacker News — "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber" (published Jul 21, 2026, 10:17 AM CT) --- # This page doesn't exist. Let's get you back on track! Try using the search bar or [visiting our homepage](https://blog.google/). ## All stories - [**We’re announcing the Alliance for America’s Skilled Trades.**\\ ![](https://storage.googleapis.com/gweb-uniblog-publish-prod/images/Skilled_Trades_Alliance_social.max-300x300.format-webp.webp)](https://blog.google/company-news/outreach-and-initiatives/creating-opportunity/alliance-america-skilled-trades/) - [**5 ways to build a side hustle with Gemini**\\ ![](https://storage.googleapis.com/gweb-uniblog-publish-prod/images/Scale_Side_Hustles_w_Gemini_her.max-300x300.format-webp.webp)](https://blog.google/products-and-platforms/products/gemini/launch-business-with-gemini/) - [**Designing emoji for the way we communicate today**\\ ![](https://storage.googleapis.com/gweb-uniblog-publish-prod/images/emoji_blog_social.max-300x300.format-webp.webp)](https://blog.google/products-and-platforms/platforms/android/world-emoji-day-noto-3d/) - [**Experience the legacy of Estadio Azteca on Google Earth.**\\ ![](https://storage.googleapis.com/gweb-uniblog-publish-prod/images/Celebrating_the_worlds_first_3-.max-300x300.format-webp.webp)](https://blog.google/products-and-platforms/products/earth/estadio-azteca/) - [**Beach vibes and temporary wallpaper are trending for back-to-school season.**](https://blog.google/products-and-platforms/products/shopping/back-to-school-trends/) - [**6 back-to-school shopping tricks every student … --- Prior RuntimeWire coverage --- - "Google launches Pixel 11 Pro Fold with gearless hinge and $1,899 price" (Aug 12, 2026, 11:48 AM CT): Google's lighter foldable pairs a gearless hinge with a company-claimed three-times durability gain. At $1,899, buyers will have to weigh that unverified claim against the device's premium price.

Google: Gemini 3.7 Flash: **Google Releases Gemini 3.7 Flash Just Three Weeks After Predecessor** *The rapid update introduces lower introductory pricing and claims benchmark gains across coding and workflow automation.* Google is rolling out Gemini 3.7 Flash to replace Gemini 3.6 Flash, replacing the prior model a mere three weeks after its debut. The rapid turnover highlights an aggressive release cadence as Google bypasses the long-anticipated Gemini 3.5 Pro to deploy what it describes as a developer-focused "workhorse" model. The abrupt replacement appears driven by competitive market pressure and developer feedback. Alongside targeted optimizations for coding and agentic tasks, Google is offering Gemini 3.7 Flash at a lower introductory price to counter less expensive rival models. According to details reported by *Ars Technica*, Google Senior Director Tulsee Doshi claimed that Gemini 3.7 Flash delivers noticeable performance gains over version 3.6, particularly in software development and structured workflows. However, these figures represent company-provided benchmark evaluations rather than independently verified third-party testing. In the company-reported metrics: * **FrontierCode 1.1 Main:** Rose to 43.6 percent, up from 34.4 percent on 3.6 Flash. * **DeepSWE v1.1:** Increased to 65.3 percent from 49 percent. * **WebDev Arena:** Climbed to a score of 1,588, compared to 1,538 previously. * **GDP.pdf:** Reached 34 percent, up from 22 percent on the complex document processing benchmark. * **AutomationBench:** Scored 30.4 percent on common business workflow execution, up from 17 percent on 3.6 Flash. While Google's reported metrics indicate significant progress in code generation and document analysis, independent verification will determine how closely these benchmark improvements reflect practical, everyday developer use.
DeepSeek-V4-Flash: **Google Rushes Out Gemini 3.7 Flash Just Weeks After Predecessor, Touting Coding Gains** **Prompt replacement of 3.6 Flash signals accelerated AI iteration as Google aims to undercut rivals on price while boosting agentic performance.** **Aug. 13, 2026** — Google today released Gemini 3.7 Flash, a new iteration of its “workhorse” model that replaces the prior 3.6 Flash version launched only three weeks ago, according to an announcement by Senior Director Tulsee Doshi. The rapid cadence marks an acceleration in Google’s model release cycle, with the company positioning the update as a response to core optimizations and developer feedback rather than a major architectural overhaul. Doshi indicated that the new model offers improved coding and agentic capabilities while arriving at a lower “introductory price,” a move aimed at competing with lower-cost alternatives from rivals. Per company-reported figures, Gemini 3.7 Flash shows notable gains on several benchmarks. On the FrontierCode 1.1 Main test, which measures code generation ability, Google cites an increase to 43.6 percent from 3.6 Flash’s 34.4 percent. The DeepSWE v1.1 evaluation, focused on software engineering tasks, rose to 65.3 percent from 49 percent. On the WebDev Arena, a measure of web development capabilities, the company reports a score of 1,588, up from 1,538. Document comprehension, as tested by the GDP.pdf benchmark, improved to 34 percent versus 3.6 Flash’s 22 percent. AutomationBench, which assesses business workflow execution, saw Google’s reported figure jump to 30.4 percent from 17 percent. These results remain company-supplied claims and have not been independently verified by third-party evaluators at the time of reporting. The accelerated replacement of Gemini 3.6 Flash raises questions about model stability and the pace of Google’s deployment pipeline. The shift could benefit developers seeking incremental improvements without waiting for a major release, but also introduces churn for teams that built workflows around the previous version just weeks ago—particularly given that Google’s documentation notes that temperature, top_p, and top_k parameters are now “deprecated and ignored” in the latest API models, though the specific impact of this change on Gemini 3.7 Flash is not established. Ars Technica first reported the release. Google has not disclosed a timeline for the previously expected Gemini 3.5 Pro model.

Model A clearly attributes the benchmark figures, distinguishes company claims from independent verification, and stays largely within the supplied evidence. Model B is polished but improperly incorporates the broken Hacker News material and adds unsupported claims about architecture, model stability, developer churn, benchmark scope, and the Gemini 3.5 Pro timeline. (Second judge pass, order swapped — scores are the average of both: Model A stays within the requested length, clearly labels the benchmark figures as company-reported, and avoids relying on the Hacker News material. Model B is more expansive but introduces unsupported benchmark descriptions, speculative claims about stability and developer churn, and an unnecessary API-parameter assertion drawn from the additional Hacker News material.)

minimax_h3_insufficient_support

Write a RuntimeWire news story about MiniMax's H3 model, including a headline and dek, in 250–350 words. Use only claims supported by the supplied materials, attribute company claims clearly, and do not infer details that are absent from the packet. Reconcile or flag the July 30 and Aug. 1 publication dates before publication, and verify the underlying announcement and the model's availability and open-model status. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 1, 2026, 1:02 AM CT) === --- Primary source: MiniMax (official) — "MiniMax H3: An open model breaking the boundaries between tasks and modalities https://t.co/pZxdX7ly4J" (published Aug 1, 2026, 12:50 AM CT) --- @MiniMax_AI (MiniMax (official)): MiniMax H3: An open model breaking the boundaries between tasks and modalities https://t.co/pZxdX7ly4J Links: https://twitter.com/renleanna/status/2083428921945780612 4 likes · 0 RT · 0 replies · 632 views --- Additional source: MiniMax Research — "MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities" (published Jul 30, 2026, 7:00 PM CT) --- Today, we're officially launching MiniMax H3, a general-purpose omni-modal generation model. H3 can jointly understand multimodal contexts spanning text, images, video, and audio. It generates video with native stereo audio at up to 2K resolution and 15 seconds in length.

Google: Gemini 3.7 Flash: **Headline:** MiniMax Unveils H3 Omni-Modal Model for Multimodal Understanding and 2K Video **Dek:** The organization characterizes H3 as an open model capable of generating 15-second videos with native stereo audio. AI developer MiniMax has introduced MiniMax H3, describing the system as a general-purpose omni-modal generation model engineered to break boundaries between tasks and modalities. According to statements released by MiniMax Research, the H3 model is capable of jointly understanding multimodal contexts spanning text, images, video, and audio. On the generation side, the company claims H3 can produce video with native stereo audio at resolutions of up to 2K and durations of up to 15 seconds. Editorial review of the release materials identifies an unverified date discrepancy across the company's communication channels. MiniMax Research published its statement on July 30, 2026, at 7:00 PM CT, declaring that it was officially launching H3 on that date. However, the primary official MiniMax account shared a link to the announcement titled "MiniMax H3: An open model breaking the boundaries between tasks and modalities" over a day later, on August 1, 2026, at 12:50 AM CT. Before publication can be fully certified, several underlying elements require further independent verification. Although MiniMax bills H3 as an "open model" that has officially launched, the provided materials do not include repository links, model weight access details, open-source software licenses, or public availability confirmations. To confirm the launch status, editors require verified access links, documentation resolving whether the official rollout occurred on July 30 or August 1, and technical evidence demonstrating the active availability and open-source distribution of the H3 model.
DeepSeek-V4-Flash: # MiniMax Launches H3, an Open Omni-Modal Model That Generates Video With Native Audio **The Chinese AI startup claims H3 can jointly understand text, images, video and audio, and produce clips up to 15 seconds long at 2K resolution.** **By RuntimeWire Staff** MiniMax has officially launched its H3 model, which the company describes as a "general-purpose omni-modal generation model" capable of jointly understanding multimodal contexts spanning text, images, video, and audio, according to a July 30 announcement from MiniMax Research. The company further promoted the release in an Aug. 1 post on X, calling H3 "an open model breaking the boundaries between tasks and modalities." The July 30 and Aug. 1 publication dates appear consistent: the research team published its technical announcement on July 30 at 7:00 PM CT, while the official corporate account amplified the launch on Aug. 1 at 12:50 AM CT. No contradiction is evident. MiniMax claims H3 can generate video with "native stereo audio" at resolutions up to 2K and lengths up to 15 seconds. The company has not specified which generation tasks are supported for text, images, or audio independently. It also has not disclosed training data, model architecture details beyond the "omni-modal" descriptor, or performance benchmarks. The model's "open" status is asserted but not defined in the supplied materials. MiniMax's announcements do not specify whether H3 is released under a permissive license, what parameters are available, or where the model weights can be accessed. The tweet's link points to a URL that is not fully resolvable from the assignment materials. **Verification needed before publication:** The reporter should confirm H3's actual availability on platforms such as Hugging Face or GitHub, and clarify whether "open model" means open weights, open source, or another definition. Without this verification, a story asserting the model's open status cannot be responsibly published. The link in MiniMax's tweet should also be tested to ensure it leads to a live, substantive page.

Model A is more cautious and avoids unsupported background details, while clearly flagging the unresolved availability and open-model claims, though it appears to fall short of the required word count. Model B better reconciles the dates and fits the requested format, but improperly adds that MiniMax is Chinese and uses an unqualified headline asserting launch and openness despite concluding those points require verification. (Second judge pass, order swapped — scores are the average of both: Model B better reconciles the dates as a research announcement followed by an official-account amplification, clearly attributes substantive claims, and explicitly concludes that availability and open-model status require verification before publication. Model B does introduce the unsupported description “Chinese AI startup,” while Model A misleadingly frames the dates as a discrepancy and may fall short of the required word count.)

DeepGrove Maple-Preview launch

Write a RuntimeWire news story about DeepGrove's launch of Maple-Preview, including a headline and dek, in 250–400 words. Use only claims supported by the supplied materials, clearly attribute company-reported figures, and preserve meaningful qualifications about the model's preview status and evaluation limits. Do not present promotional benchmark or device-performance claims as independently established facts. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 4, 2026, 6:18 PM CT) === --- Primary source: Hacker News — "Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone" (published Aug 4, 2026, 2:44 PM CT) --- Today we introduce **Maple-Preview**, an **open-source 20B-A1B ternary-weight** reasoning LLM. Maple-Preview is SOTA in its weight class and is even competitive with larger models. It solves IMO-level problems and runs at **200+ tokens/s** on a **Mac mini M4**, **5–16×** faster than efficient models like Gemma 4, Qwen3.5, and gpt-oss. [Chat with Maple-Preview →](https://chat.deepgrove.ai/) - **20B-A1B** Model - **218 tokens/s** M4 Mac mini - **5.31 GB** Checkpoint - **131,072** Token context ![On-device decode speed versus benchmark performance for Maple-Preview with its dense and Flash heads and comparison models.](https://deepgrove.ai/modelcard/speed-frontier.svg?v=20260804-flash11) ### Maple-Preview solves [IMO 2024 P1](https://web.evanchen.cc/exams/IMO-2024-notes.pdf) (7/7) at 281.5 tokens/s on a MacBook Pro (M5 Pro). Maple-Preview earning 7/7 on IMO 2024 Problem 1 on a MacBook Pro (M5 Pro), running at 281.5 tokens/s. ### Maple-Preview runs at 127 tokens/s—13× faster than 1-bit Bonsai 27B (Qwen3.6 27B) on an iPhone **Maple-Preview** ≈ 00:09 **Bonsai 27B** \> 05:58 Maple Maple-Preview (127 tokens/s) vs. Bonsai 27B (9.6 tokens/s) on iPhone. Both receive the same prompt, “Make me a carrot cake.” Maple-Preview completes its response in about 10 seconds; Bonsai is still generating at the six-minute mark. ## Our Bet We envision a shift from the age of monolithic LLMs to **always-active on-device assistants, which continuously shape themselves to improve user experience.** We believe that this necessitates a move towards more efficient and performant architectures, such that these assistants can both be trained and run on everyday devices (e.g. laptops and phones). As exemplified in Maple-Preview, we view ultra-low precision as being crucial to this new era of efficient, performant modeling. At ultra-low bitwidths, matrix multiplication can be effectively replaced with additions, lowering the total arithmetic workload needed to infer through a model. With native support for low bitwidths and custom hardware that takes advantage of both the memory and arithmetic efficiency of such architectures, we find it easy to imagine a world where most inference is done on smaller, personalized models and infrequent, exceedingly difficult tasks are offloaded to cloud models. Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach; we think that this unnecessarily limits both performance (to being only some percentage of the full-precision model) and efficiency (by enforcing architectural constraints that may not be beneficial). We posit that in many ways, creating a high-performing, efficient model should be similar to creating simply a high-performing model. As such, instead of focusing on how to make a performant model efficient, we believe it is most fruitful to dedicate substantial effort toward working on optimization, data, and more to improve model performance while simultaneously enforcing efficiency through an inference-aware architectural design loop, pushing the frontier in both directions. **We believe that the precision a model runs at should be the precision it learns at.** At DeepGrove, we study how efficient models learn, rethinking architecture, training infrastructure, optimization, and hardware design to treat low precision as a first-class citizen. Maple-Preview is a natively trained ternary-weight network showing that low precision does not have to mean compromise. ## Architecture Maple-Preview is a 20B-A1B reasoning model designed from the start for efficient on-device reasoning. We designed the Maple architecture in a hardware-aware manner, testing all considered configurations on our Mac mini for inference speed. Starting with a 30-layer, 224-expert configuration, we optimized to a 24-layer, 256-expert configuration as a compromise between model performance and inference efficiency. We additionally chose hybrid sliding-window and global attention to bound KV-cache growth. ![Projected resident model-weight and context-state memory at 131,072 tokens for Maple and comparison models.](https://deepgrove.ai/modelcard/model-context.svg?v=20260804-20) ![Incremental context-memory growth above a 640-token baseline through 131,072 tokens for Maple and comparison architectures.](https://deepgrove.ai/modelcard/context-scaling.svg?v=20260804-9) ## Evaluation On benchmarks, **Maple-Preview sets a new point on the Pareto frontier for both memory-to-performance and speed-to-performance,** demonstrating its strong reasoning capabilities. However, we note that this preview is focused primarily on raw reasoning and, as such, may underperform on agentic benchmarks. We intend to continue improving general performance through extended training before Maple's full release.

Google: Gemini 3.7 Flash: **DeepGrove Releases Maple-Preview Ternary MoE Model for On-Device Inference** *The open-source preview posts high reported speeds on consumer hardware, though the developers note gaps on agentic tasks.* AI developer DeepGrove has unveiled Maple-Preview, an open-source 20B-A1B reasoning model designed to run directly on consumer hardware such as laptops and smartphones. The model features a 5.31 GB checkpoint and supports a context window of up to 131,072 tokens. According to DeepGrove, the system is natively trained as a ternary-weight network rather than post-quantized from full precision, an approach the company argues replaces matrix multiplications with additions to reduce arithmetic workloads. Its architecture features 24 layers and 256 experts in a mixture-of-experts (MoE) configuration, using hybrid sliding-window and global attention to limit KV-cache memory expansion. DeepGrove reported substantial generation speeds in its internal testing, claiming the preview achieved 218 tokens per second on an Apple M4 Mac mini and 127 tokens per second on an iPhone. The company asserted that Maple-Preview runs 5 to 16 times faster on its test setups than models including Gemma 4, Qwen3.5, and gpt-oss. DeepGrove also reported that the model solved Problem 1 of the 2024 International Mathematical Olympiad (7/7) during a test on an M5 Pro MacBook Pro. However, DeepGrove framed the release as an early preview focused strictly on raw reasoning capabilities rather than general-purpose deployment. The company explicitly acknowledged that the model may underperform on agentic benchmarks in its current state, adding that it plans to conduct extended training to improve overall performance ahead of a full release.
DeepSeek-V4-Flash: # DeepGrove Launches Maple-Preview, an Open-Source 20B Ternary Model Claiming 120+ Tok/s on iPhone **On-device reasoning model uses natively trained ternary weights to challenge assumptions about low-precision AI performance, but company cautions it's a preview focused on raw reasoning.** SAN FRANCISCO — DeepGrove today released Maple-Preview, an open-source 20B-A1B ternary-weight reasoning model that the company reported achieves 127 tokens per second on an iPhone — a speed it claims is 13× faster than the 1-bit Bonsai 27B model on the same device when generating a response to a common prompt. According to a company post on Hacker News, the model's checkpoint is 5.31 GB and supports a 131,072-token context. DeepGrove reported that Maple-Preview runs at 218 tokens per second on an M4 Mac mini and at 281.5 tokens per second on a MacBook Pro with an M5 Pro chip when solving IMO 2024 Problem 1. DeepGrove stated that Maple-Preview is "SOTA in its weight class" and "even competitive with larger models," though it characterized the release as a preview. The company noted the model "may underperform on agentic benchmarks" and said it intends to "continue improving general performance through extended training before Maple's full release." The company positions Maple-Preview as a native ternary-weight network — meaning it was trained at low precision from the start, rather than converted from a full-precision model — which DeepGrove called "fundamentally the wrong approach" for optimizing both performance and efficiency. "At ultra-low bitwidths, matrix multiplication can be effectively replaced with additions," the company wrote, envisioning "a world where most inference is done on smaller, personalized models." The architecture uses a 24-layer, 256-expert configuration with hybrid sliding-window and global attention, which DeepGrove said was chosen to bound KV-cache growth after testing configurations on a Mac mini for inference speed. Maple-Preview is available for chat at chat.deepgrove.ai.

Model B better satisfies the 250–400-word requirement, consistently attributes performance claims, and preserves the preview and evaluation caveats, though its unsupported San Francisco dateline should be removed. Model A is concise and generally accurate but appears to fall short of the requested word count and overstates the source slightly by saying the preview is focused “strictly” rather than “primarily” on raw reasoning. (Second judge pass, order swapped — scores are the average of both: Model B more consistently attributes reported specifications and performance figures while preserving the preview and agentic-benchmark qualifications, though its unsupported San Francisco dateline and one ambiguous sentence are flaws. Model A is polished but presents the checkpoint size and context length without attribution and overstates the model’s focus as “strictly” raw reasoning when the source says “primarily.”)

Latitude Health Form D funding disclosure

Write a 250–350-word RuntimeWire news story about Latitude Health's newly disclosed financing. Include a headline and dek. Base the story on the SEC Form D and clearly distinguish filing facts from company-reported product claims. State the amount sold, total offering size, number of participating investors, filing date, company location and leadership where relevant. Do not identify investors or characterize the securities beyond what the filing establishes. Explain what the company does and note that its automation claims are claims by the company, not independently verified. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 7, 2026, 4:53 PM CT) === --- Primary source: SEC Form D — "Latitude Health raises $2M for AI-native utilization management" (published Aug 7, 2026, 2:50 PM CT) --- Latitude Health, Inc. filed SEC Form D on August 7, 2026, disclosing $2,000,000 in securities sold toward a $3,000,000 total offering. The Delaware corporation is based in San Francisco, CA. Latitude Health operates an AI-native platform for utilization management (UM) and prior authorization, claiming to automate 75% of manual work, double review throughput, and reduce clinician burnout. The company appears to have emerged from stealth around early 2025 and is led by Charles Feerick (Executive Officer and Director), with Jarred Bressner and Chris Palmieri also serving as directors. The filing shows only 2 investors have participated so far. The company's website is latitudehealth.com. This is the first concrete funding disclosure for the company. Source: SEC EDGAR Form D filing 0002148343-26-000001

Google: Gemini 3.7 Flash: **Latitude Health Raises $2M Toward $3M Offering for Utilization Management Platform** *An SEC Form D filing reveals two investors have backed the San Francisco-based startup in its first concrete funding disclosure.* San Francisco-based healthcare technology startup Latitude Health, Inc. has secured $2 million in funding, according to a Form D filed with the U.S. Securities and Exchange Commission on August 7, 2026. The regulatory filing indicates that the Delaware corporation has established a total offering size of $3 million, leaving $1 million remaining to be sold. A total of two investors have participated in the offering as of the filing date. This filing represents the first concrete financing disclosure for the company, which appears to have emerged from stealth around early 2025. According to the SEC document, Latitude Health is led by Charles Feerick, who is listed as an executive officer and director. Jarred Bressner and Chris Palmieri also serve as directors for the company. The Form D filing does not disclose the specific identities of the participating investors, nor does it characterize the terms or nature of the securities sold beyond the reported financial figures. Latitude Health operates an AI-native platform designed for utilization management (UM) and prior authorization workflows in healthcare. According to marketing statements published by the company, the platform is designed to automate 75% of manual administrative work, double clinical review throughput, and reduce clinician burnout. These operational metrics and workflow automation claims are self-reported statements by the company and have not been independently tested, verified, or evaluated by RuntimeWire.
DeepSeek-V4-Flash: # Latitude Health Discloses $2M Raise for AI-Powered Utilization Management Platform **SAN FRANCISCO** — Latitude Health, Inc. has disclosed $2 million in new financing through a securities offering, according to a Form D filed with the SEC on August 7, 2026, marking the first concrete funding disclosure for the stealth-stage healthcare technology company. The Delaware corporation, headquartered in San Francisco, sold $2,000,000 toward a total $3,000,000 offering. The filing indicates only two investors have participated to date. The securities were sold on August 7, the same day the filing was submitted. According to the filing, company leadership includes Charles Feerick as Executive Officer and Director, with Jarred Bressner and Chris Palmieri also serving on the board. Latitude Health describes itself as operating an "AI-native platform for utilization management and prior authorization." The company claims its technology can automate 75% of manual work, double review throughput, and reduce clinician burnout. These performance claims are company assertions and have not been independently verified. The company appears to have emerged from stealth around early 2025 and maintains a web presence at latitudehealth.com. The Form D filing does not identify specific investors or characterize the securities beyond establishing that they were sold in this offering. The total offering size suggests the company may seek an additional $1 million from future investors. **Editor's note:** This story is based solely on Latitude Health's SEC Form D filing. While the filing provides verifiable details about the offering amount, structure, and company leadership, the company's product claims regarding automation rates and clinical outcomes are reported as stated by Latitude Health and have not been independently confirmed. Further verification would require interviews with company executives, customer testimonials, or third-party validation of the platform's performance metrics.

Model A is more accurate, includes both a headline and dek, and cleanly separates filing facts from unverified company claims, though it falls short of the required word count. Model B meets the length requirement but lacks a dek, adds the unsupported assertion that the securities were sold on August 7, and uses a repetitive editor’s note that weakens the news style. (Second judge pass, order swapped — scores are the average of both: Model A better satisfies the requested format with a clear headline and dek, stays within the target length, and carefully separates SEC filing facts from unverified company claims. Model B lacks a dek and adds the unsupported assertion that the securities were sold on the filing date, along with an unnecessary editor’s note.)

deepseek-v4-flash-mi300x-unverified-repository

Write a 250–350-word RuntimeWire news story about the newly published GitHub repository and what it claims about running DeepSeek V4 Flash on an AMD MI300X. Include a headline and dek. Attribute claims clearly to the repository, distinguish reported results from independently verified facts, and use only details supported by the supplied materials. Assess whether the available evidence is sufficient for publication; do not infer the author's employment, affiliations, model specifications, deployment status, or benchmark validity beyond the sources. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 4, 2026, 6:08 AM CT) === --- Primary source: Hacker News — "DeepSeek V4 Flash on a Single AMD MI300X" (published Aug 4, 2026, 5:00 AM CT) --- [Skip to content](https://github.com/ryanzhou/deepseek-v4-flash-mi300x#start-of-content) You signed in with another tab or window. [Reload](https://github.com/ryanzhou/deepseek-v4-flash-mi300x) to refresh your session.You signed out in another tab or window. [Reload](https://github.com/ryanzhou/deepseek-v4-flash-mi300x) to refresh your session.You switched accounts on another tab or window. [Reload](https://github.com/ryanzhou/deepseek-v4-flash-mi300x) to refresh your session.Dismiss alert {{ message }} [ryanzhou](https://github.com/ryanzhou)/ **[deepseek-v4-flash-mi300x](https://github.com/ryanzhou/deepseek-v4-flash-mi300x)** Public - [Notifications](https://github.com/login?return_to=%2Fryanzhou%2Fdeepseek-v4-flash-mi300x) You must be signed in to change notification settings - [Fork\\ 0](https://github.com/login?return_to=%2Fryanzhou%2Fdeepseek-v4-flash-mi300x) - [Star\\ 0](https://github.com/login?return_to=%2Fryanzhou%2Fdeepseek-v4-flash-mi300x) main [**1** Branch](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/branches) [**0** Tags](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tags) [Go to Branches page](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/branches)[Go to Tags page](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tags) Go to file Code Open more actions menu ## Folders and files | Name | Name | Last commit message | Last commit date | | --- | --- | --- | --- | | ## Latest commit<br>[![ryanzhou](https://avatars.githubusercontent.com/u/124444?v=4&size=40)](https://github.com/ryanzhou)[ryanzhou](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commits?author=ryanzhou)<br>[Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0)<br>Open commit details<br>22 minutes agoAug 4, 2026<br>[7c06e57](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0) · 22 minutes agoAug 4, 2026<br>## History<br>[1 Commit](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commits/main/) <br>Open commit details<br>[View commit history for this file.](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commits/main/) 1 Commit | | [patches](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/patches "patches") | [patches](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/patches "patches") | [Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0 "Open-source DeepSeek V4 Flash on a single MI300X production stack Production-validated stack for deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X (gfx942): digest-pinned vLLM ROCm compose deployment, ten byte-identical correctness/performance overlays with SHA-256 pins, reference unified diffs against upstream vLLM/ROCm-triton/AITER, and AITER GEMM tuning tables. Includes blog-style README with measured results (168.6 tok/s median single-stream decode; 7.9-8.5K tok/s tuned prefill).") | 22 minutes agoAug 4, 2026 | | [tuning](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/tuning "tuning") | [tuning](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/tree/main/tuning "tuning") | [Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0 "Open-source DeepSeek V4 Flash on a single MI300X production stack Production-validated stack for deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X (gfx942): digest-pinned vLLM ROCm compose deployment, ten byte-identical correctness/performance overlays with SHA-256 pins, reference unified diffs against upstream vLLM/ROCm-triton/AITER, and AITER GEMM tuning tables. Includes blog-style README with measured results (168.6 tok/s median single-stream decode; 7.9-8.5K tok/s tuned prefill).") | 22 minutes agoAug 4, 2026 | | [.gitignore](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/blob/main/.gitignore ".gitignore") | [.gitignore](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/blob/main/.gitignore ".gitignore") | [Open-source DeepSeek V4 Flash on a single MI300X production stack](https://github.com/ryanzhou/deepseek-v4-flash-mi300x/commit/7c06e57ee4c9cd6c4ba4d70e8a6422aa6d5562f0 "Open-source DeepSeek V4 Flash on a single MI300X production stack Production-validated stack for deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X (gfx942): digest-pinned vLLM ROCm compose deployment, ten byte-identical correctness/performance overlays with SHA-256 pins, reference unified diffs against upstream vLLM/ROCm-triton/AITER, and AITER GEMM tuning tables. --- Additional source: Hacker News — "Bringing Up DeepSeek-V4-Flash on AMD MI300X" (published Jun 2, 2026, 12:52 PM CT) --- # Bringing up DeepSeek-V4-Flash on AMD MI300X 1 Jun 2026 9 min read At [Doubleword](https://app.doubleword.ai/) we are building an inference cloud designed for volume. To do that we have to reckon with the enveloping compute shortage. AMD’s MI300X launched in December 202311.At AMD’s [“Advancing AI” event](https://www.amd.com/en/newsroom/press-releases/2023-11-15-amd-announces-amd-instinct-mi300-accelerator-launc.html), 6 December 2023. as AMD’s response to NVIDIA’s H100, arriving alongside H200 in the same generation. It is an odd duck in the world of high-end AI accelerators. While H100 prices are climbing (up 40% in five months on one-year rentals, with on-demand capacity sold out across every major NVIDIA part22.SemiAnalysis, [The Great GPU Shortage: Rental Capacity](https://newsletter.semianalysis.com/p/the-great-gpu-shortage-rental-capacity), April 2026.), MI300X is perhaps still underappreciated. 192GB of HBM3 per card against the H100’s 80GB, comparable FP8 compute, list price roughly half. Yet you can rent one on-demand today (from [Hotaisle](https://www.hotaisle.ai/), for instance) for noticeably less than the equivalent NVIDIA capacity. The reason is software.

Google: Gemini 3.7 Flash: **Headline:** GitHub Repository Claims High-Throughput DeepSeek V4 Flash Stack on AMD MI300X **Dek:** An uncorroborated repository reports up to 168.6 tokens per second on single-GPU hardware, but available evidence remains insufficient for independent verification. A newly published GitHub repository claims to provide a production-validated software stack capable of running DeepSeek’s `deepseek-ai/DeepSeek-V4-Flash-0731` model on a single AMD MI300X (gfx942) accelerator. Published on August 4, 2026, by user `ryanzhou`, the repository `deepseek-v4-flash-mi300x` contains a digest-pinned vLLM ROCm compose deployment, ten byte-identical correctness and performance overlays with SHA-256 pins, reference unified diffs against upstream vLLM, ROCm-triton, and AITER, and AITER GEMM tuning tables. According to commit metadata in the repository, the setup achieved measured results of 168.6 tokens per second median single-stream decode and between 7.9K and 8.5K tokens per second in tuned prefill. The claims touch on ongoing industry efforts to optimize software for AMD's MI300X accelerator, which launched in December 2023 featuring 192GB of HBM3 memory compared to the 80GB found on Nvidia’s competing H100. However, *RuntimeWire* assesses that the available evidence is currently insufficient to publish these claims as established facts. The performance figures and production claims originate entirely from a single, newly created repository with one commit, zero stars, and zero forks. The provided sources contain no third-party verification, independent benchmarks, or external confirmation of deployment validity, and no details are provided regarding the author's professional affiliation or background. To support a fully publishable news report, *RuntimeWire* would require additional verification, including: - Independent reproduction of the prefill and decode benchmarks on physical MI300X hardware. - Technical review and testing of the ten patch overlays against upstream vLLM, ROCm-triton, and AITER codebases. - Third-party corroboration confirming that the software configuration operates reliably in production environments.
DeepSeek-V4-Flash: # GitHub Repo Claims DeepSeek V4 Flash Achieves 168 Tokens/Sec on Single AMD MI300X **A newly published repository purports to deliver production-grade inference for DeepSeek's latest model on AMD hardware, but key claims remain unverified.** **SAN FRANCISCO** – A GitHub repository published early Tuesday claims to enable running DeepSeek V4 Flash on a single AMD MI300X accelerator, reporting median single-stream decode speeds of 168.6 tokens per second and prefill throughput between 7,900 and 8,500 tokens per second. The repository, created by user "ryanzhou" under the name `deepseek-v4-flash-mi300x`, describes itself as a "production-validated stack" for the model `deepseek-ai/DeepSeek-V4-Flash-0731`. It includes code patches, tuning tables, and a README documenting measured results. According to the repository's commit message, the stack uses "digest-pinned vLLM ROCm compose deployment" with "ten byte-identical correctness/performance overlays with SHA-256 pins." The project also provides "reference unified diffs against upstream vLLM/ROCm-triton/AITER" and "AITER GEMM tuning tables." The README – referenced in the repository's summary – reports three key performance figures: "168.6 tok/s median single-stream decode" and "7.9-8.5K tok/s tuned prefill." However, the repository does not specify the methodology, test duration, hardware configuration (beyond a single MI300X), or software environment used to obtain these results. These figures are reported results from the repository author and have not been independently verified. RuntimeWire could not confirm whether the patches successfully build or execute, whether the claimed performance is reproducible, or whether the repository has been reviewed by AMD, DeepSeek, or any third party. The repository contains zero stars and zero forks at time of publication, suggesting minimal community validation. The single commit was made approximately 22 minutes before the story was filed. Additional context is available from a June 2026 blog post by Doubleword, an inference cloud provider, which documented challenges bringing DeepSeek-V4-Flash to AMD MI300X hardware. That post, which is not affiliated with the new repository, described the MI300X as "underappreciated" due to software limitations despite offering 192GB of HBM3 memory at roughly half the list price of NVIDIA's H100. **Assessment:** The available evidence is insufficient for publication as a standalone news story. The repository's claims are unverified, lack methodological detail, and have received no community validation. Additional verification would require: (1) independent reproduction of results by a third party; (2) confirmation of the author's identity and technical credentials; (3) review of benchmark methodology including batch sizes, precision, and prompt lengths; and (4) confirmation that the model referenced (`DeepSeek-V4-Flash-0731`) is an officially released model from DeepSeek.

Model A stays within the requested format and length, attributes the performance claims carefully, and gives a sound verification assessment, though its dek slightly mischaracterizes the 168.6 tokens-per-second result as an “up to” figure. Model B exceeds the word limit and adds unsupported details or inferences, including a San Francisco dateline, “latest model,” lack of affiliation between sources, and claims about missing repository methodology that cannot be established from the supplied excerpt. (Second judge pass, order swapped — scores are the average of both: Model A stays within the requested length, clearly attributes the repository’s claims, and appropriately concludes that independent reproduction and technical review are needed. Model B exceeds the word limit and introduces unsupported or improper assertions, including a San Francisco dateline, claimed non-affiliation, and a purported RuntimeWire verification effort not established by the materials.)

AGent Energy raises $11 million Series Seed

Write a 250–350-word RuntimeWire news story about AGent Energy’s newly announced financing, including a clear headline and dek. Report the $11 million Series Seed and the company’s stated $17 million in total funding, identify the co-lead investors and other participants, and explain what AGent says it plans to do with the capital. Attribute company claims about behind-the-meter generation, grid conditions, market participation and emergency-event dispatches rather than presenting promotional assertions as independently verified facts. Make clear that the 200+ GW figure is AGent’s target or estimate of behind-the-meter generation, not disclosed capacity already controlled by the company. Do not add biographical details, customer figures, revenue figures, site counts or regulatory claims not supported by the packet. If the materials do not support a publishable story, say so and explain what additional verification you would need. === ASSIGNMENT MATERIALS (as of Aug 13, 2026, 5:29 PM CT) === --- Primary source: PR Newswire Business Technology — "AGent Energy Closes Series Seed to Unlock 200+ GW of Behind-the-Meter Generation Across Commercial, Industrial, and Institutional Sectors" (published Aug 13, 2026, 12:00 PM CT) --- [Accessibility Statement](https://www.cision.com/about/accessibility/) [Skip Navigation](https://www.prnewswire.com/news-releases/agent-energy-closes-series-seed-to-unlock-200-gw-of-behind-the-meter-generation-across-commercial-industrial-and-institutional-sectors-302851086.html#main) _Round Co-Led by Spero Ventures and MassMutual Ventures with Participation from Intrepid Investment Management and Existing Investors Zero Infinity Partners (ZIP) and CIV; Brings Total Funding to $17 Million in Just 12 Months, Making It One of the Fastest-Funded Distributed Energy Resource Companies to Date_ HOUSTON, Aug. 13, 2026 /PRNewswire/ -- AGent Energy, a trailblazing developer of AI-driven distributed power plants, today announced it has closed an $11 million Series Seed financing co-led by Spero Ventures and MassMutual Ventures, with participation from Intrepid Investment Management and existing investors CIV and Zero Infinity Partners (ZIP). The round follows a $6 million financing from CIV and ZIP, which closed within two months of founding, bringing AGent's total funding to $17 million in its first 12 months. It's a striking vote of confidence in behind-the-meter generation as the next great frontier of U.S. energy infrastructure. America's grid is under mounting strain. PJM's most recent capacity auction cleared at the price cap without enough capacity to meet demand, and data center load growth is outpacing new supply across every major market. AGent is unlocking a faster, smarter way to keep the power flowing: the backup generation that already sits at commercial, industrial, and mission-critical facilities, including AI data centers. AGent's AI-based platform aggregates, orchestrates, and monetizes these assets, turning them into rapidly dispatchable, highly reliable distributed power plants. Because the equipment is already built, already paid for, and idle most of the year, AGent delivers capacity at the lowest cost of any new grid resource, at zero cost to the asset owner, who earns new revenue instead. AGent is already dispatching in three of the largest wholesale markets in North America, having successfully delivered capacity during grid emergency events in PJM, MISO, and ERCOT. AGent will use the new capital to aggressively scale its team and accelerate its push to unlock 200+ GW of behind-the-meter generation across commercial and industrial facilities and the MUSH sector (municipalities, universities, schools, and hospitals), the properties with the highest concentration of on-site generation and the ones AGent's team knows best. "Twelve months ago, AGent was an idea backed by two investors who saw what we saw: 200+ GW of the most reliable generation in America sitting idle behind the meter," said Stephanie Hendricks, CEO and Co-Founder of AGent. "Closing $17 million in our first year reflects both the urgency of the grid reliability challenge and the speed at which our team executes. With Spero, MassMutual Ventures, and Intrepid joining CIV and ZIP, we now have the partners and capital to bring this dispatchable capacity to the markets that need it most." "The grid doesn't need to wait five years for new steel in the ground. The capacity is already there, and AGent has built the intelligence layer to unlock it. Stephanie and her team have moved faster than any company we've seen in this space, and we're proud to co-lead their Series Seed round," said Stephen Wemple of Spero Ventures. "When the grid is stressed, the difference between a rolling blackout and an ordinary afternoon is how fast dispatchable capacity shows up. AGent's AI platform turns generation already sitting at hospitals, universities, and industrial sites into exactly that: capacity that responds in minutes, with no new construction and no cost to the owner. It's a rare combination of climate resilience and hard economics, built by a team that has done this at scale before. That's the sweet spot for our Climate Technology Fund: AI applied to real assets, with economics that pencil from day one," said Aram Ouligian, Senior Associate at MassMutual Ventures. In connection with the financing, Stephen Wemple of Spero Ventures will join AGent's board of directors, and Aram Ouligian of MassMutual Ventures will join as a board observer. **About AGent, Inc.** Using its AI-based technology platform, AGent aggregates, orchestrates, and monetizes distributed generation resources in power markets, delivering significant reliability, economic, and sustainability benefits to large energy users, utilities, and grid operators. To learn more visit agentenergy.com. **About Spero Ventures** Spero Ventures is an early-stage venture capital firm that invests in founders building a future that belongs to everyone, with a focus on sustainable systems, health and longevity, and human potential. Learn more at spero.vc.

Google: Gemini 3.7 Flash: **AGent Energy Raises $11M Seed Round to Tap Behind-the-Meter Generation** *The Houston startup has secured $17 million in total funding to scale its team and target an estimated 200 GW of backup power across commercial and institutional sites.* Houston-based AGent Energy announced Thursday that it has closed an $11 million Series Seed financing round, bringing the startup’s total funding to $17 million across its first 12 months of operations. The financing was co-led by Spero Ventures and MassMutual Ventures, with participation from Intrepid Investment Management alongside existing investors CIV and Zero Infinity Partners (ZIP). The round follows an initial $6 million financing closed shortly after the company's founding. In connection with the investment, Stephen Wemple of Spero Ventures will join AGent’s board of directors, and Aram Ouligian of MassMutual Ventures will join as a board observer. AGent stated that it plans to deploy the fresh capital to expand its workforce and accelerate efforts to tap an estimated 200-plus gigawatts of behind-the-meter generation across commercial and industrial facilities, as well as the MUSH sector—comprising municipalities, universities, schools, and hospitals. The company framed the 200-plus GW figure as an industrywide estimate of idle on-site generation that it aims to target, rather than capacity currently under its operational control. The company pitches its AI-based software platform as a response to mounting power grid stress, pointing to surging data center load growth and supply constraints, such as PJM’s recent capacity auction clearing at its price cap. According to AGent, its platform aggregates, orchestrates, and monetizes idle on-site backup generation, converting existing equipment into rapidly dispatchable distributed power plants at no cost to asset owners. AGent also claims it is actively participating in three major North American wholesale power markets, asserting that its platform has successfully delivered capacity during grid emergency events across the PJM, MISO, and ERCOT footprints.
DeepSeek-V4-Flash: **AGent Energy Secures $11M Series Seed to Tap 200+ GW of Idle Behind-the-Meter Generation** **The AI-driven distributed power plant developer, which says it has raised $17 million in total funding in its first year, plans to scale its team and expand market operations across commercial, industrial and institutional sites.** HOUSTON – AGent Energy announced Wednesday it has closed an $11 million Series Seed financing co-led by Spero Ventures and MassMutual Ventures, with participation from Intrepid Investment Management and existing investors CIV and Zero Infinity Partners (ZIP). The company stated the round brings its total funding to $17 million within 12 months of founding, following an earlier $6 million financing from CIV and ZIP. AGent describes itself as a developer of AI-driven distributed power plants. The company says its platform aggregates, orchestrates and monetizes backup generation already located at commercial, industrial, and mission-critical facilities, including AI data centers. According to the company, the equipment is already built and idle for most of the year. AGent states it is currently dispatching capacity in PJM, MISO, and ERCOT, having responded to grid emergency events in those markets. The company said it will use the capital to scale its team and accelerate efforts to unlock what it estimates as 200+ gigawatts of behind-the-meter generation. Quoting the CEO and Co-Founder, the company said it targets the commercial and industrial sectors as well as the MUSH sector (municipalities, universities, schools, and hospitals). Company representatives attributed the need for such capacity to strained grid conditions. AGent pointed to data center load growth and PJM’s most recent capacity auction, which the CEO described as reflecting “the urgency of the grid reliability challenge.” The company has not disclosed current capacity under management, customer counts, or revenue figures. In connection with the financing, Stephen Wemple of Spero Ventures will join AGent’s board of directors, and Aram Ouligian of MassMutual Ventures will join as a board observer. *(This story attributes all claims about market conditions, generation capacity targets, and dispatch events directly to AGent Energy or its representatives. It does not independently verify figures related to PJM auction results, data center load projections, or the company’s stated 200+ GW estimate.)*

Model A is accurate, polished and properly attributes AGent’s claims while clearly presenting 200-plus GW as a target rather than controlled capacity. Model B incorrectly says the announcement came Wednesday, uses a more promotional headline, and ends with an unnecessary editorial note that weakens the news-story format. (Second judge pass, order swapped — scores are the average of both: Model A is more polished, stays within the requested news format, clearly frames the 200-plus GW figure as a target rather than controlled capacity, and consistently attributes operational and grid claims. Model B is largely compliant but incorrectly says the Thursday announcement occurred Wednesday, uses awkward phrasing in places, and adds an unnecessary editorial disclaimer.)

Matchup powered by OpenRouter.