CoreWeave opens limited Vera Rubin access as Cognition runs production AI workloads
CoreWeave says Cognition is the first customer running production workloads on NVIDIA’s Vera Rubin NVL72 platform. An early test reports higher token throughput, but access remains limited.
CoreWeave said on September 30 that NVIDIA’s Vera Rubin NVL72 platform is in limited availability on its cloud, with Cognition running the first production workloads on the system. Announced during CoreWeave Fully Connected in San Francisco, the deployment gives selected customers access to new infrastructure for training and operating AI agents, while leaving wider availability and pricing unspecified.
NVIDIA identified Cognition, the company behind the Devin AI software engineer, as the first production customer. CoreWeave said it had deployed hundreds of Rubin GPUs across multiple regions and had begun onboarding customer workloads. Its description of limited availability matters: the announcement establishes a working deployment, but does not establish that any customer can immediately obtain capacity.
What Cognition is running on Vera Rubin
Cognition uses CoreWeave for both model training and production inference for Devin, according to the companies. Inference is the computing used when a deployed model responds to a request; training and reinforcement learning are part of improving that model. CoreWeave said Cognition began running production workloads within days of receiving the new racks, without having to build the underlying environment itself.
Devin’s work can involve reading code repositories, generating and testing code, and repeating those steps as a task develops, CoreWeave said. That makes throughput relevant beyond a single response: an agent may have to process substantial context and produce more output before it can complete an assignment. CoreWeave describes the new platform as supporting those workloads alongside the training that improves its models.
Cognition introduced SWE-2, its newer coding model, in a September 10 post. The company said then that SWE-2 was available in Devin Desktop and its command-line interface, with a rollout to Devin Web and Fusion under way. Those earlier product details explain the coding workload used in the new comparison; Cognition’s model benchmarks do not independently establish the speed of NVIDIA’s new hardware.
How to read the 4.8-times throughput figure
NVIDIA said Cognition tested Vera Rubin NVL72 against a GB200 NVL72 baseline after CoreWeave received its first production racks earlier in September. It said the test used AI agents working on a sampled subset of FrontierCode software engineering tasks. In those early tests, Cognition measured up to 4.8 times the total token throughput for SWE-2 inference workloads on Vera Rubin NVL72, NVIDIA reported.
CoreWeave also reports a 4.8-times inference result and separately says Cognition saw 3.8 times the output-token throughput for reinforcement-learning workloads at matched interactivity. The second figure concerns model improvement work, while the first concerns inference. They are different measures and should not be read as one overall speed increase for every task Devin performs.
The figures come from a customer test described by the companies, rather than an independently reproduced evaluation. The published accounts do not show whether the sampled tasks represent the full FrontierCode benchmark or the range of work customers give coding agents. Token throughput measures how much model text a system processes or generates over time; the reported comparison does not by itself measure completed software assignments, coding quality or savings for other customers.
The comparison also applies to the stated systems and workloads. A customer using another model, a different mix of tasks or another infrastructure configuration could see a different result. Neither announcement supplies a public price comparison that would let readers calculate the cost of producing those tokens on the two platforms.
What limited availability means for customers
CoreWeave calls the launch limited availability and directs prospective customers to request a briefing about capacity planning, onboarding timing and workload fit. It says customers can use Vera Rubin capacity through services including CoreWeave Kubernetes Service, CoreWeave Inference and Mission Control. The announcement gives no general release date, regional capacity breakdown or price list.
NVIDIA’s announcement says the CoreWeave systems use Spectrum-X 102.4T Ethernet networking. It also introduces CoreWeave Forge, an environment intended to connect model training, evaluation and improvement, drawing on Weights & Biases, OpenPipe expertise and the marimo notebook project. Forge is a separate part of the companies’ broader agent software announcement; Cognition’s production deployment is the concrete example they provide for the new Vera Rubin capacity.
For now, the established change is that CoreWeave has begun putting Vera Rubin NVL72 into selected customers’ production use and Cognition says it is already running workloads there. The reported throughput gains describe an early comparison for Cognition’s coding workloads. Broader access, independently reproduced performance and customer costs remain open questions.
Sources and context
- From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AINVIDIA
- Cognition Becomes First Customer for NVIDIA Vera Rubin NVL72 on CoreWeave CloudCoreWeave
- Introducing SWE-2: Pushing the Pareto FrontierCognition
AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.
About NewsJaws Desk
AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.