Iterate Studio Inc. launched Lifeboat, an inference engine for large language models with built-in confidential computing that fits two to six times more concurrent AI agent sessions on each graphics processing unit. In company testing on a single Nvidia RTX PRO 6000 Blackwell GPU running a Qwen 30B-A3B model, the engine held 2,048 concurrent sessions with every request completing. For banks, insurers and health systems that want agents inside their own infrastructure, the claim matters because it promises lower GPU spending without sending data to per-token cloud services.

Iterate. ai launches Lifeboat to fit up to six times more agents per GPU

How Lifeboat handles agent memory load

Lifeboat targets a memory bottleneck created by agents at scale. A chatbot question usually triggers one model call, while an agent on a single task can make dozens of calls and its context window grows each time, filling the key-value cache quickly. Iterate. ai says standard inference engines can stall at four or five long-context requests running at once. Teams that reach that ceiling typically buy more GPUs or shift work to outside services and hand data to a third party. Co-founder and chief technology officer Brian Sathianathan positioned Lifeboat as a way to run agents on proprietary data inside company walls.

The engine combines fair scheduling with admission control so each session receives its share of GPU capacity and a heavy document-processing agent cannot crowd out interactive work. Optimizations to the key-value cache double its effective capacity while model weights remain at full precision, according to the company. On mixture-of-experts models, only the experts needed by a given agent are loaded rather than the full set. Each session also runs in its own security capsule with filtering, token budgets and sandboxed execution. Together these controls address throughput, isolation and cost in one deployment.

The context for the release is rising enterprise demand for agents that work across long documents and multi-step tasks. That workload pattern differs from single-turn chat and puts sustained pressure on memory, scheduling and security rather than raw model quality. Data centers and so-called neo-clouds face the same constraint, since doubling concurrent sessions per card returns capacity without adding racks or power. Co-founder and chief executive Jon Nordmark said buyers should test what their existing GPUs can handle before purchasing more hardware. Early access to Lifeboat drew thousands of downloads within days, and the engine is generally available now.

What this means for companies running AI agents

For operating teams, the practical effect would be higher agent density on hardware already owned. The company reported 8,714 tokens per second with optimizations on versus 4,965 with them off, alongside twice the concurrent sessions in that comparison. In a memory-pressure run of 128 sessions sending 18,000-token requests, Lifeboat kept 99th-percentile time to first token at 1.5 seconds against 189 seconds for the baseline. A small firm could therefore pilot long-context agents on one or two servers, while a large organization could defer purchases and keep sensitive workloads on premises. The Developer License covers noncommercial and evaluation use on up to two inference servers on a single node.

The second consequence concerns regulated data and vendor selection. The top-tier Confidential Computing edition will not serve a request until hardware attestation passes, covering trusted execution features in Advanced Micro Devices and Intel processors plus Nvidia confidential computing mode on H100, B200, GB300 and other supported GPUs. Model weights are sealed inside the trusted execution environment and stay encrypted in use, whether in a cloud confidential virtual machine or on customer-owned hardware. Buyers should still verify attestation coverage for their exact chips, cloud configuration and compliance regime. The paid tiers can be tried free for seven days without a credit card, with Standard at $49.99 per month and Confidential Computing at $499.99 per month.

The marker to watch is independent confirmation of session density on production agent mixes rather than controlled tests. If enterprises or cloud operators report sustained twofold or greater concurrency on long-context workloads with stable latency, the approach will have moved from benchmark to operating practice. The company also discussed its Generate platform and partnership with NetApp on theCUBE at the NetApp Insight conference. Those deployments will show whether memory efficiency plus attestation shortens procurement for regulated industries.