Speaker Diarization on 95MB: NVIDIA Tracks 8 Speakers On-Device

Speaker Diarization on 95MB: NVIDIA Tracks 8 Speakers On-Device

By Aaron Jackson
Publication Date: 2026-09-28 05:53:00

Ask a speech-to-text model what was said in a meeting and it will hand you back a clean transcript. Ask it who said it and it goes quiet. That gap is speaker diarization, and until recently it was the most expensive part of any voice pipeline: the models capable of it were large, ran on servers, and mostly topped out at two or four speakers.

NVIDIA’s Nemotron 3 Diarization, released on 23 September 2026, changes that. It is a 100-million-parameter model that separates up to eight overlapping speakers, runs live or on recordings from the same checkpoint, and has already been converted by independent developers into builds small enough for a phone, a Mac, and a web browser. The largest build is 397MB. The smallest is 95MB.

Here is what speaker diarization actually does, how the new model performs, and what it takes to run it yourself.

What Speaker Diarization Actually Does (and Why Whisper Cannot)

Speech recognition and speaker diarization answer different questions. Speech recognition answers “what was said.” Speaker diarization answers “who spoke when.” Neither one does the other job, which is why pairing them is standard practice rather than a luxury. Nemotron 3 Diarization sits firmly on the second side of that line. It has no vocabulary and cannot write a word: its output is a grid of probabilities showing who is talking and when, so the text in any finished transcript always comes from a separate speech recognition model.

Consider a transcript where every…