Build1 publisher2 min readPublished
STAC audit puts Myrtle.ai's FPGA tree inference under two microseconds at p99
Myrtle.ai's VOLLO ran three gradient-boosted-tree models under two microseconds p99 in a STAC audit, topping out at 50 million inferences a second. The audit times model inference alone, so trading desks still have to measure their own data and order path before budgeting for cards.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The tests ran on an AMD Alveo V80LL accelerator card installed in a Blackcore ICON 3132-SM+ server.
- Myrtle.ai's release cites STAC report MRTL2026905, which has no accessible text, while STAC report ML-20260925 lists the matching VOLLO configuration.
- In STAC's comparison with Xelera Silva on an AMD Alveo V80, VOLLO shows as much as 42% lower p99 latency and as much as 71% more throughput when the two run directly comparable numbers of model instances.
- The release also claims at least five times the throughput of previous best results, a figure the available materials do not reconcile with STAC's comparison.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- contradiction A desk that budgets on the release's fivefold throughput would be planning on roughly 2.9 times the gain that STAC's own comparison with Xelera supports.
- capability The VOLLO Sandbox and SDK give a team a way to evaluate its own trees against the tested complexity range ahead of any hardware decision.
- decision STAC's standardized workload lets a trading team shortlist FPGA vendors on audited figures and spend its own test time on the strategy and server fit the benchmark cannot settle.
The 50 million figure is a total across many copies of the model running in parallel. STAC's report reaches it on GBT_A with 64 model instances, at under 1.5 microseconds p99 [7]. That works out to about 780,000 inferences a second per instance [19]. GBT_B runs 32 instances at 28 million a second [8], or about 875,000 per instance [20]. Per copy, the two smaller models land close together. Most of the gap in total throughput between them comes from the instance count.
Baldwin says trading firms can run more powerful models without sacrificing speed [14]. On latency, the report backs him: GBT_C, the slowest of the three, stays under 1.8 microseconds p99 [9]. Its throughput is 1.9 million a second [9], about 26 times lower than GBT_A [21]. STAC built the suite around inference latency and how consistent it stays across models of varying complexity [4]. A desk whose production trees look like GBT_C should plan against 1.9 million a second.
The smallest model has two latency numbers. The release gives 1.77 microseconds p99 at 50 million inferences a second [2]. The runtimewire account describes that as a separate reported result at that throughput, distinct from the report's latency figures at stated instance counts [11]. A procurement memo should cite the report figure along with its instance count.
That Xelera system ran in an HPE ProLiant DL385 Gen10 Plus v2. STAC identifies it as the first audited submission to its El Popo gradient-boosted-tree suite [12]. Beating the first audited entry in a suite introduced in 2025 [5] is a record in the most literal sense. Both systems use V80-family cards in different servers [3][12]. The comparison is mostly between two inference stacks on related AMD silicon.
The part I would credit is the error column. Baldwin founded the company in 2017 on hardware-software co-design. The aim was to let software developers use FPGA acceleration without becoming FPGA engineers [15]. The report lists model error as 0.00 for all three models [10]. By STAC's error measure, then, the speed did not come from approximating the trees.
The audit stops at the model. End-to-end trading latency also depends on receiving and preparing data and on acting on the model's output [17]. Three conditions have to hold for the audited numbers to carry over to a desk. Its trees have to be close to the tested models in complexity. Its strategy has to make use of dozens of parallel instances. And the stages before and after inference have to be small next to two microseconds.
What to watch
- Accessible report text for STAC SUT MRTL2026905, and whether it supports the fivefold throughput claim.
- New audited submissions to STAC's El Popo suite that would give VOLLO a second baseline to beat.
- Trading firms publishing VOLLO latency measured across a full data-to-order path.