CoreWeave has moved from raw GPU capacity to full-stack optimization for AI inference, launching the Forge platform and the RL Rollouts capability at its Fully Connected event. Testing of RL Rollouts showed a 15x improvement in model reload latency over a baseline configuration. The move matters because serving models faster and at lower cost is becoming the factor that decides the economics of AI deployments.
Forge platform and RL Rollouts in detail
Vice president of product and AI services Urvashi Chowdhary described the strategy in an interview with theCUBE Research's Dave Vellante and John Furrier. She said developers want to solve a problem as quickly as possible with strong performance and scalable cost. CoreWeave is therefore adding managed services for training, post-training and inference on top of its infrastructure. The comments were made during an exclusive broadcast on theCUBE, SiliconANGLE Media's livestreaming studio.
RL Rollouts is a preview capability built on Nvidia's Dynamo framework and aimed at reinforcement learning for agentic models. When customers train such models with rewards and verifiers, inference becomes the bottleneck during rollouts, Chowdhary said. The capability loads new checkpoints into a live deployment without stopping the scaled inference setup. As a result, training runs continue while inference scales, and iteration cycles become shorter.
CoreWeave Forge connects serving, observability, post-training and evaluation in one platform introduced at the event. It is free to start, with paid tiers adding further capabilities for developers and teams. Chowdhary said even an individual developer signing up alone receives the same performance and reliability without compromise. The design relies on open-source tools, with contributions back to open systems so customers retain flexibility.
What faster inference means for AI budgets
For companies running chatbots, assistants and agents, the tuning of layers above hardware directly affects latency and unit cost per request. CoreWeave cites work from the vLLM engine to quantized models and custom speculative decoders as part of that optimization. A small team can start on the free tier of Forge and access managed serving and evaluation without building its own stack. A large organization can standardize post-training and observability while scaling inference separately from training.
The approach has limits that buyers should verify before committing production workloads. The 15x reload improvement comes from testing against a baseline configuration, and actual gains will depend on model size, checkpoint frequency and deployment setup. Open-source flexibility does not remove dependence on CoreWeave infrastructure and Nvidia's Dynamo framework. Customers should ask about supported models, quantization trade-offs, service-level terms and how paid tiers differ from free access.
The marker to watch is adoption of Forge beyond the launch and conversion from free use to paid tiers over the next 12 months. A parallel signal is whether inference share keeps rising in customer workloads, as in the healthcare case cited by theCUBE Research. If both trends hold, full-stack inference optimization will become a standard requirement for AI cloud contracts.
