Nvidia introduces Nemotron 3 Nano Omni with vision and speech for powerful agentic AI use  – SiliconANGLE

Nvidia introduces Nemotron 3 Nano Omni with vision and speech for powerful agentic AI use  – SiliconANGLE

By @SiliconANGLE
Publication Date: 2026-04-28 16:00:00

Nvidia Corp. today launched a powerful reasoning artificial intelligence model that unifies text, vision and speech, capable of acting as the “brains” of faster, smarter agentic AI applications. 

Dubbed Nemotron 3 Nano Omni, and weighing in at about 30 billion parameters, the new state-of-the-art model uses mixture-of-experts architecture to deliver extremely low latency and provides high flexibility and control. 

Nvidia combined vision and audio encoders with its 30B-AD3B hybrid MoE architecture to eliminate the need for separate perception modules, allowing its AI model to unify everything into one. The company said this allowed the model to improve efficiency at scale and provide up to nine times faster throughput than other open omni models on the market. 

“To build useful agents, you can’t wait seconds for a model to interpret a screen,” said Gautier Cloix, chief executive of H Company. “By building on Nemotron 3 Nano Omni, our agents can rapidly interpret full HD screen recordings — something that wasn’t practical before.” 

The result is a lower cost and higher scalability. With its smaller size, it can also be compressed enough to run on higher-end consumer hardware and execute efficiently on enterprise cloud deployments. 

The company said it is designed to run alongside other proprietary cloud models or other Nvidia Nemotron open models, such as Nemotron 3 Super for high-frequency execution or Super for…