Product1 publisher3 min readPublished
Iterate.ai's own test shows Lifeboat doubling agent sessions on a single GPU
Iterate.ai launched Lifeboat, an inference engine it says fits two to six times more AI agent sessions on each GPU. Its own published test shows a doubling on one Nvidia card, still enough to make regulated firms measure the cards they own before buying more.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- According to Iterate.ai, four or five simultaneous long-context requests can be enough to stall a standard inference engine, the limit Lifeboat is built to lift.
- The company says its key-value cache changes double effective cache capacity while model weights stay at full precision.
- Iterate.ai's test ran a Qwen 30B-A3B model on a single Nvidia RTX PRO 6000 Blackwell GPU, where Lifeboat held 2,048 concurrent sessions with every request completing.
- A Confidential Computing edition refuses to serve requests until hardware attestation passes on supported AMD, Intel and Nvidia hardware, including H100, B200 and GB300 GPUs.
- Paid tiers cost $49.99 a month for professional support and $499.99 a month for the Confidential Computing edition, each with a seven-day trial that needs no credit card.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision Regulated firms choosing between more GPUs and a per-token cloud now have a step to take before either: measuring what scheduling and cache changes recover on cards they already own.
- exposure A team that sizes its hardware plan on the six-times headline would be counting on three times the session gain the only published test shows.
- cost Firms that need hardware attestation pay the Confidential Computing list price of $5,999.88 a year, a cost to set against the extra GPUs it is meant to make unnecessary.
A document-processing agent picks up a long file. Over a single task it can make dozens of model calls, and its context window grows with each one [2]. On a shared GPU that agent can crowd out the interactive sessions other staff are using, and Lifeboat's fair scheduling and admission control are meant to give each session its share of the card [4].
When teams hit that limit, SiliconANGLE reports, they tend to buy more GPUs or move the work to per-token cloud services, handing their data to a third party [10]. Both moves assume the cards are full. Iterate.ai's argument is that agent memory in the key-value cache is what runs out [2]. "Before any enterprise buys more GPUs for its agents, it should find out what the ones it already owns can do," Jon Nordmark, the chief executive, said [12]. According to Brian Sathianathan, the chief technology officer, banks, insurers and health systems want agents working on their own data "inside their own walls" without doubling GPU spending [11].
Against a pitch of two to six times [24], the published session result sits at the floor [19]. Throughput rose from 4,965 to 8,714 tokens per second, about 1.76 times [8][20]. Latency under memory pressure shows the widest gap. With 128 sessions sending 18,000-token requests, Lifeboat's 99th-percentile time to first token was 1.5 seconds against 189 for the baseline, a factor of 126 [9][21]. Each baseline was Lifeboat's own engine with its optimizations switched off [7]. The company did not publish a run against another vendor's engine, or one that reaches six times.
The company also said early access drew thousands of downloads within days [16]. A download counts someone who fetched a free engine. A buyer needs to know how many sessions its own agents hold per card in production.
For a bank or a health system, the tier that matters is the Confidential Computing edition, because the attestation features sit there [18]. That edition seals the model weights inside the trusted execution environment, where they remain encrypted while in use, on a cloud confidential virtual machine or on hardware the customer owns [15]. Nordmark pitched the same capacity argument to providers. "A data center or neo-cloud that doubles concurrent sessions per card gets that capacity back without adding racks or power," he said [13].
I'd sort the decision on two axes. One is context pressure: whether agents hold long contexts at high concurrency, the condition behind the 126-fold latency gap [21]. The other is residency: whether the data may leave hardware the firm controls. Long contexts with data that must stay on owned hardware is the quadrant Lifeboat is built for. There, the free Developer License covers evaluation on up to two inference servers on one node [17]. That is enough to replay a team's own agent traces on a card it already owns before a GPU order goes out. The tradeoff is that every published figure comes from Iterate.ai's own testing on one card and one model [6], so the planning number belongs nearer the doubling than the sixfold. Where data may leave, the same long-context workload puts a card running twice the sessions, plus a license, against the per-token cloud bill. With short contexts and data kept in-house, the latency result matters less, and the session and throughput gains are the ones to measure. Short contexts with data free to leave is the weakest case for changing anything.
What to watch
- An independent benchmark of Lifeboat against another vendor's inference engine on the same card and model.
- A bank, insurer or health system naming Lifeboat in production and reporting sessions held per GPU.
- Published results on the H100, B200 or GB300 GPUs that the Confidential Computing edition supports.