CoreWeave brings Nvidia Vera Rubin systems into production with Cognition
CoreWeave announced availability of Nvidia’s Vera Rubin NVL72 systems on its cloud this week, naming Cognition, maker of the Devin AI coding agent, as the first customer running production workloads on the platform. The disclosure is the clearest public marker yet of how fast Nvidia’s newest chip generation is moving from lab benchmark to paying customer.
- Cognition reported up to 4.8x higher total token throughput on Vera Rubin NVL72 versus a GB200 NVL72 baseline
- CoreWeave’s Vera CPU deployment packs 128 CPUs and 11,264 cores into a single rack
- CoreWeave Forge, Agent Lens and Sandboxes moved to general availability alongside the hardware launch
- 4.8x token throughput gain, Vera Rubin vs GB200 baseline
- 3x faster agent sandbox startup on Vera CPU
- 20% improvement in agent failure detection, Agent Lens
- 40% lower cost for serverless RL vs self-managed setup
CoreWeave said in the release, published on Nvidia’s newsroom during CoreWeave’s Fully Connected conference in San Francisco this week, that it is “one of the first cloud providers to deliver the platform in customers’ hands.” The company said it stood up a production Vera Rubin cluster for Cognition “in days” and that Cognition scaled to thousands of GPUs on CoreWeave over nine months to power its Devin inference workloads. CoreWeave also said it will offer Nvidia’s Vera CPU, which Nvidia’s release describes as “the first CPU built for AI agents.”
Cognition’s benchmark and the Vera CPU numbers
Cognition ran the comparison itself, using a subset of tasks from the FrontierCode benchmark deployed as AI-agent workloads, and reported the 4.8x token-throughput figure against a GB200 NVL72 baseline for what it calls SWE-2 inference workloads. Silas Alberti of Cognition’s founding team said in the release: “Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship.”
On the CPU side, CoreWeave said its Vera deployment supports more than 11,000 concurrent agent environments at one core each, and that testing showed sandbox startup times more than 3x faster on Vera CPUs, with a 1.7x gain across all passing Terminal-Bench tasks.
What the release does not put a number on
Nvidia’s release does not disclose pricing for Vera Rubin NVL72 or Vera CPU capacity on CoreWeave Cloud, nor a timeline for moving beyond “early-access customers” to broader general availability. It also does not say how many GPUs Cognition is running in production today, only that it “scaled to thousands” over nine months, and it does not disclose Vera Rubin production volumes at CoreWeave beyond the first racks received “earlier this month.”
CoreWeave is a publicly traded cloud infrastructure provider that has built its business on long-term Nvidia GPU capacity, a relationship the companies describe as nearly a decade old; Cognition is a private AI lab whose Devin product markets itself as an autonomous software engineer. Both are well known to readers tracking AI infrastructure spending, but neither company’s release supplies a dollar figure for this deployment.
The economics behind the loop
The commercial argument in the release rests on cost per token and cost per fix rather than raw speed. CoreWeave said its Agent Lens service “improves failure detection by 20% and fixes issues at half of the cost,” and that serverless reinforcement learning on Forge “trains 1.4x faster at 40% lower cost than a self-managed setup.” Those are the numbers that matter to allocators comparing CoreWeave against hyperscaler alternatives, more than the headline GPU throughput figure.
Ian Buck, Nvidia’s vice president of hyperscale and high-performance computing, framed the pitch around depreciation rather than peak performance. “CoreWeave’s NVIDIA V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production,” Buck said in the release. “That’s the strength of the NVIDIA platform: infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload.”
“Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship.”
Silas Alberti, founding team, Cognition
The BlockWest read. A single-customer, vendor-run benchmark against last generation’s chip is not proof of market-wide economics, and CoreWeave has every incentive to publish the strongest number it has. What we would watch is whether the 40% cost reduction on serverless RL and the sandbox-density claims on Vera CPU hold up once enterprise customers like Capital One and Canva, both named as early Forge users, disclose their own usage costs rather than relying on CoreWeave’s framing.
CoreWeave’s Fully Connected conference runs through this week in San Francisco, where Nvidia says additional sessions and demos on Vera Rubin deployment are scheduled; neither company has given a date for Vera Rubin NVL72 to move from early access to general availability on CoreWeave Cloud.
BlockWest is a news publication. Nothing here is investment advice. Read our disclaimer and editorial policy.
