Nvidia
Nemotron 3 Diarization
Nvidia’s Nemotron 3 Diarization marks when each speaker is active in recorded or live audio

Key facts
- 23 Sep 2026
- Released
- 100M parameters
- Size
- Up to 8
- Speaker channels
- OpenMDW 1.1
- Licence
Nvidia’s Nemotron 3 Diarization marks when each speaker is active in recorded or live audio. Released on 23 September 2026, it supports up to eight speakers and overlapping speech.
Nemotron 3 Diarization is an AI model that marks when each speaker talks in an audio recording. It can also work on live audio and track overlapping speech. Nvidia released it on 23 September 2026, with support for up to eight speaker channels.
Speaker labels can be joined to a transcript
The model returns time intervals for anonymous speakers. A speech-recognition system can supply the words, and an application can combine those words with the intervals to produce a transcript attributed to Speaker 1, Speaker 2 and so on. Identifying a person’s real name requires separate information.
The model has 100 million parameters. Its model card specifies 16 kHz mono audio and a configurable streaming buffer. The lowest recommended buffer is 0.32 seconds; that describes one component of the delay, while total application latency also includes processing and delivery.
The same weights support live and recorded audio
Chunked processing lets the model work through long recordings. The weights are released under OpenMDW 1.1 for commercial and non-commercial use under the licence terms.
Nvidia reports a 14.72% diarisation error rate on VoiceArena’s benchmark at release. That figure measures errors in assigning speech time to speakers in the evaluation. A meeting system should also be checked on its own microphones, background sound and overlapping conversations, because these conditions affect the result.
More in Audio and Voice
All Audio →- Alibaba QwenQwen-Audio-3.1Alibaba's upgraded audio stack, announced 23 September 2026
- Microsoft AIMAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flashlive transcription that tops the streaming accuracy board, and two new voices
- AlibabaHappyShrimp 1.0the name is a eulogy
- Microsoft AIMAI-Transcribe-2speech recognition at ten cents an hour
- SunoSuno v6a flagship, an exploratory variant and a free mini, released together
- SunoSuno v5.5the market leader