By Srabanti Chakraborty
Publication Date: 2026-10-01 11:24:00
Cognition reports up to 4.8x higher total token throughput for SWE-2 inference workloads versus a GB200 NVL72 baseline.
CoreWeave Inc. announced the availability of NVIDIA Vera Rubin NVL72 on CoreWeave, with Cognition as the first customer anywhere running production workloads on the system. Customers, like Cognition, run the system under the same operating model and tooling as their existing NVIDIA GB200 NVL72 and GB300 NVL72 fleets, with performance engineering from CoreWeave’s team. The news was shared during Fully Connected, CoreWeave’s AI cloud conference, which brings together more than 4,500 customers, partners, developers and AI leaders to share how they are building and running AI in production.
Cognition is the first customer in production with Vera Rubin NVL72
Cognition, the applied AI lab behind Devin, runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled from bridge capacity to thousands of GPUs for training and inference in less than nine months. Cognition worked with CoreWeave to stand up a Vera Rubin NVL72 cluster in early September, and Cognition’s own engineers ran the first customer-executed Vera Rubin inference benchmark, measured against a GB200 NVL72 cluster baseline.
Cognition’s engineers benchmarked Vera Rubin
In independent benchmarks run on CoreWeave Cloud, Cognition measured a…

