Site icon VMVirtualMachine.com

Nvidia is chasing China on open weights, by its own account

Nvidia is chasing China on open weights, by its own account

By Ana Maria Constantin
Publication Date: 2026-08-12 15:03:00

Nvidia released Nemotron 3.5 Lightning on Tuesday. It is a 30 billion parameter mixture-of-experts model with three billion active at any moment. The architecture is a hybrid of Mamba-2, MoE and attention layers, with a one million token context window. The weights are on Hugging Face and ModelScope.

The licence is the part that matters. It ships under OpenMDW-1.1, which permits commercial use, and Nvidia published the training data and the recipes alongside the weights. It is free for companies to download, use and modify without asking permission or paying Nvidia, CNBC notes.

The speed claim has two numbers

Nvidia leads on output speed of up to four times that of similar-sized models. Read further and the agentic figure is more modest. On PinchBench it hit 86% accuracy while finishing 10,000 tasks 30% faster than Alibaba’s Qwen 3.6 35B at similar accuracy.

Both numbers are in the same release. Four times faster describes token generation in the lab. Thirty percent is what happens when the model actually does a job. That gap is worth carrying whenever a vendor quotes throughput.

The benchmark scores are respectable rather than startling. It posts 81.94 on MMLU Pro, 75.44 on GPQA Diamond and 51.56 on SWE-bench Verified. Pre-training ran to more than 20 trillion tokens.

The router is the real product

Nvidia also released NeMo Switchyard, an open-source library. It sends each step of an agent workflow to whichever model handles it best. Plans…

Exit mobile version