YFarmX logoYFarmX

Nvidia

Nemotron 3 Diarization

Nvidia’s Nemotron 3 Diarization marks when each speaker is active in recorded or live audio

Released 23 September 20261 min readAudio and VoiceLast updated:

Nemotron 3 Diarization editorial illustration

Key facts

23 Sep 2026
Released
100M parameters
Size
Up to 8
Speaker channels
OpenMDW 1.1
Licence

Nvidia’s Nemotron 3 Diarization marks when each speaker is active in recorded or live audio. Released on 23 September 2026, it supports up to eight speakers and overlapping speech.

Nemotron 3 Diarization is an AI model that marks when each speaker talks in an audio recording. It can also work on live audio and track overlapping speech. Nvidia released it on 23 September 2026, with support for up to eight speaker channels.

Speaker labels can be joined to a transcript

The model returns time intervals for anonymous speakers. A speech-recognition system can supply the words, and an application can combine those words with the intervals to produce a transcript attributed to Speaker 1, Speaker 2 and so on. Identifying a person’s real name requires separate information.

The model has 100 million parameters. Its model card specifies 16 kHz mono audio and a configurable streaming buffer. The lowest recommended buffer is 0.32 seconds; that describes one component of the delay, while total application latency also includes processing and delivery.

The same weights support live and recorded audio

Chunked processing lets the model work through long recordings. The weights are released under OpenMDW 1.1 for commercial and non-commercial use under the licence terms.

Nvidia reports a 14.72% diarisation error rate on VoiceArena’s benchmark at release. That figure measures errors in assigning speech time to speakers in the evaluation. A meeting system should also be checked on its own microphones, background sound and overlapping conversations, because these conditions affect the result.