By Frederic Lardinois
Publication Date: 2026-08-25 19:13:00
On Tuesday, IBM launched the latest family of its open-weight Granite large language models (LLMs). Weighing in at 3 billion, 8 billion, and 30 billion parameters, IBM is taking a very different approach to model building from some of the competition here, with dense, decoder-only reasoning models that it pre-trained from scratch.
Many recent models have moved from all-attention Transformers toward hybrid Mamba/attention architectures, including Nvidia’s Nemotron 3 family. But IBM tried that with its Granite 4.0 models. That generation included conventional dense, dense-hybrid, and hybrid MoE models.
With Granite 4.1, IBM returned the main family to an all-attention, dense Transformer architecture. At the time, IBM said that these new models outperformed the older generation, “while using a simpler — and therefore more flexible — architecture for fine-tuning for downstream tasks.”
Reasoning with Granite
IBM describes the 4.2 family as a…



