CoreWeave has introduced Forge, a platform workflow designed to move teams from a first agent to their best agent as quickly as possible. The loop runs through five stages — run, observe, curate, improve, evaluate — and then repeats, with each production run feeding the next model and agent iteration. The launch includes Agent Lens, Registry, RL Rollouts, model distillation and the generally available ARIA assistant. For businesses, the signal is that agent work is framed as continuous improvement rather than one-time deployment.
Forge loop, tools and partner network
Forge is positioned as a shared workflow for business units and machine learning teams that previously worked with disconnected tools. Susanne Seitinger, vice president of product marketing at CoreWeave, described the goal in an interview with John Furrier and Dave Vellante at the Fully Connected event on theCUBE. The company presents speed of learning as the main measure: faster learning leads to faster releases and faster value. The partner network is part of the same plan.
The improvement cycle is built around five linked operations that turn operating data into a stronger agent. Teams run agents in production, observe behavior, curate useful examples, improve models and configurations, then evaluate results before starting again. CoreWeave says each pass should build on the previous one. The design assumes post-training never stops, including work done during inference. That structure makes production traffic a training asset instead of only a cost center.
Several capabilities address specific bottlenecks inside that cycle. Agent Lens is for inspecting what agents are doing in practice, while Registry stores checkpoints and agent configurations for reuse and comparison. RL Rollouts and model distillation cover the improvement step, including compressing knowledge from a large model into a version tuned for a particular use case. ARIA, the AI Research and Iteration Agent, is now generally available to help users find patterns in runs. Together they connect observation, versioning and retraining.
The launch reflects a shift in enterprise workloads toward agents that need constant supervision and revision. Seitinger pointed to Cognition AI as an example, already running production workloads on CoreWeave first Nvidia Vera Rubin NVL72 racks. The company argues its expertise extends beyond initial training into later stages of the loop and inference operations. A partner network with tested, co-engineered integrations adds deployment recipes and playbooks. The aim is to reduce repeated trial and error across customers.
What faster agent iteration means for business
For companies putting agents into production, the practical effect is a shorter path from pilot to a stable version that handles real requests. A shared run-to-evaluation workflow can cut handoffs between product teams and machine learning specialists, while checkpoints make rollbacks and comparisons easier. Small firms gain ready-made integration patterns instead of building evaluation pipelines from scratch. Large organizations can standardize how multiple departments test, approve and update agents.
The limits sit in fit, maturity and operational discipline. Distillation and reinforcement-style rollouts help only when teams have representative production data and clear quality criteria for agents. Agent Lens and ARIA still require staff who can interpret traces, define failure modes and decide what to curate for retraining. The announcement by itself does not prove cost, latency or accuracy gains for a specific workload. Buyers should ask which integrations are tested, how versions are governed and what post-training support covers during inference.
A useful marker will be whether Cognition AI and other early production users report repeated measurable gains from successive Forge cycles on Vera Rubin capacity. Further evidence would be published playbooks, named partners and case data linking curation and evaluation to fewer agent errors. If those appear in coming quarters, continuous agent improvement will look like standard cloud practice. If not, Forge will remain a promising framework awaiting proof.
