Myrtle.ai's VOLLO hits 50M tree inferences per second in STAC test
On an AMD Alveo V80LL and Blackcore server, VOLLO delivered sub-2-microsecond p99 latency across three tree models; STAC lists it as ML-20260925, while myrtle.ai's release cites MRTL2026905.
By RuntimeWire Staff · Published
Primary source: PR Newswire
Why it matters
The benchmark quantifies VOLLO's performance on three tested models and identifies the prior system used for comparison, giving Myrtle.ai a concrete basis for its FPGA-inference pitch in latency-sensitive trading. The results remain specific to a defined workload and hardware configuration.
myrtle.ai founder and CEO Peter Baldwin says trading firms can run more powerful models without sacrificing speed. On October 6th, myrtle.ai reported STAC-audited records for its VOLLO inference accelerator on gradient-boosted-tree workloads: p99 latency below two microseconds across three models and 50 million inferences per second on the smallest one.
The result was unveiled at the STAC Summit in London. According to myrtle.ai's announcement, VOLLO cut 99th-percentile latency by more than 30% and delivered at least five times the throughput of the previous best results. On the smallest model, myrtle.ai reported 1.77-microsecond p99 latency at 50 million inferences per second. The tests ran on an AMD Alveo V80LL Compute Accelerator in a Blackcore ICON 3132-SM+ server.
The figures cover model inference under a defined benchmark. End-to-end trading latency also depends on receiving and preparing data and acting on a model's output. STAC's working group describes its Markets (Inference) suite as a benchmark for real-time financial data. STAC's description of the gradient-boosted-tree benchmark says it focuses on inference latency and performance consistency across models of varying complexity. The test-harness documentation describes the testing framework. STAC introduced a dedicated benchmark suite for gradient-boosted trees in 2025.
The release identifies the detailed results as STAC report SUT ID MRTL2026905, but that page has no accessible report text. STAC's working-group listing and its accessible ML-20260925 report list the matching myrtle.ai VOLLO, AMD Alveo V80LL and Blackcore ICON 3132-SM+ configuration under SUT ID ML-20260925. The accessible report verifies three model results: GBT_A delivered less than 1.5-microsecond p99 latency through 64 model instances and 50 million inferences per second; GBT_B delivered less than 1.5-microsecond p99 latency through 32 instances and 28 million inferences per second; GBT_C delivered less than 1.8-microsecond p99 latency and 1.9 million inferences per second. The report lists model error as 0.00 for all three. The release's 1.77-microsecond figure is for the smallest model at 50 million inferences per second; it is a reported result at that throughput, distinct from the report's latency figures at the stated instance counts.
The report compares VOLLO with the prior SUT XLRA260312, Xelera Silva on an AMD Alveo V80 in an HPE ProLiant DL385 Gen10 Plus v2 server. STAC identifies that system as its first audited submission for the El Popo gradient-boosted-tree suite. Against XLRA260312, STAC reports up to 42% lower p99 latency and up to 71% higher throughput at directly comparable model-instance counts. The release separately claims at least five times the throughput over previous best results; the available materials do not reconcile that figure with the named comparison. The standardized workload allows comparisons, but does not establish that every trading strategy, model or server will see the same advantage.
A founder's bet on hardware and software together
Baldwin has worked on software and specialized computing hardware for years. Myrtle.ai's team profile says he founded the business around hardware-software co-design for efficient machine-learning inference in 2017. He holds a PhD in pure mathematics from Cambridge, helped establish the speech-transcription benchmark for MLPerf, previously founded a software company focused on distributed data-center software and worked as a senior technical director at visual-effects companies.
The founding thesis was to make specialized acceleration usable by software developers who do not want to become FPGA engineers. The October announcement applies that idea to machine learning on structured financial data, where model outputs must arrive quickly enough for real-time decisions. Gradient-boosted trees are used for structured data, and STAC introduced a dedicated benchmark suite for them in 2025.
Myrtle.ai's product materials describe VOLLO as an accelerator for running machine-learning models on FPGA platforms. Myrtle.ai now offers a VOLLO Sandbox and SDK for developers to evaluate models without owning an FPGA. Baldwin says prospective users can try their own models before committing to specialized hardware. The sandbox lowers one barrier to evaluation, but does not answer questions about integration effort, deployment costs or production performance for a particular trading stack.
Myrtle.ai's earlier financing offers a glimpse of the people who backed that technical approach. In a Cambridge Angels funding account dated March 30th, 2017, the group said Myrtle Software had secured seed financing led by Robert Sansom, with participation from IQ Capital, Robert Swann, William Tunstall-Pedoe and Adrian Weller. The account did not give a round size. It described a team building hardware-based deep-learning technology and said the investment would support development and new products. That early backing predates VOLLO's current benchmark push and does not show myrtle.ai's present funding position.
What the record shows and its limits
Myrtle.ai says VOLLO has supported hundreds of thousands of hours of live trading that generated alpha for trading firms. Myrtle.ai does not name those customers or provide independently reported production results in the announcement. The claim concerns deployments in customer environments; the STAC test evaluates a specified technology stack.
On the named AMD accelerator and server, VOLLO handled the benchmark's tree-inference workload at the reported latency and throughput. Myrtle.ai says it can support other FPGA platforms too, but this result is tied to the hardware configuration disclosed for the test.
Baldwin said in the release that firms want to run more powerful models without giving up speed, and that developers can test their own models on VOLLO without FPGA expertise. The benchmark quantifies that pitch on the tested setup. The release points to SUT ID MRTL2026905, while the accessible STAC report and working-group listing give the matching benchmark configuration as ML-20260925. Myrtle.ai is inviting developers to run their own evaluations; the reported 50 million inferences per second applies to the smallest benchmark model and does not predict results for every model or trading system.