On September 23, 2026, NVIDIA released Nemotron 3 Diarization, an AI model that labels who spoke and when in a conversation, with timestamps. The earlier Streaming Sortformer handled up to four speakers. The new model supports up to eight and works on both recorded and live audio. That should make it easier to follow who said what in meeting transcripts, but "supports eight speakers" does not mean it can always tell eight people apart accurately. An independent evaluation shows that its strengths and weaknesses depend on the number of speakers.

AD

What changed with the release

NVIDIA has published the weights of a pretrained model with roughly 100 million parameters. Given audio, it returns, for each moment in time, the probability that speech is occurring on each of up to eight speaker channels. Each channel is labeled in the order speakers first appear in the conversation. These are anonymous labels such as "speaker 0" and "speaker 1," not identity verification that works out a person's name from their voice. When several people talk over each other, multiple channels can be active at the same time.

The older model also processed audio in small chunks and carried information about speakers found in earlier segments forward to the next. The new model keeps that arrival-order speaker cache and a memory of recent audio, while raising the number of speakers it can output from four to eight. So describing this release as "the first time speakers can be separated in real time" would overlook what the previous model already did. What is new is that a single model can now handle a wider range of larger conversations.

The model does not transcribe what is said. NVIDIA's model card describes how to obtain start and end times for each speaker and gives an example of combining it with a separate speech recognition model. To produce meeting minutes, you need to match the output of word recognition and speaker separation by timestamp. To display names, you also need additional information that ties the anonymous labels to actual participants.

Strengths and weaknesses in the 8-speaker evaluation

In early results from Voice Arena, an independent evaluator, Nemotron 3 Diarization ranked first among 12 systems, with an overall diarization error rate of 14.72%. The runner-up, DiariZen WavLM Large s80 MD v2, scored 19.34%. Diarization error rate combines the time speech was missed, the time silence was wrongly labeled as speech, and the time speech was assigned to the wrong speaker. Lower is better. The overall ranking alone, however, does not reveal what "up to eight speakers" means in practice.

In the same Voice Arena evaluation, Nemotron 3's error rate was 3.02 percentage points higher than the runner-up's in two-person conversations. In conversations with five to eight people, Nemotron 3's error rate was 11.57 points lower than the runner-up's.

Speakers in conversation Conversations Nemotron 3 Runner-up How to read it
2 37 10.31% 7.29% Runner-up has the lower error rate
5–8 17 22.18% 33.75% Nemotron 3 has the lower error rate

The figures come from Voice Arena's comparison by number of speakers: the two-speaker gap is 10.31 minus 7.29, and the five-to-eight-speaker gap is 33.75 minus 22.18. The evaluation used 139 English conversations totaling about 22 hours, checked against ground truth in which the final speaker timing was verified by humans after recording. Overlapping speech is scored, no tolerance margin is allowed at speech boundaries, and systems are not told the number of participants in advance. The 5–8 row combines 17 conversations and does not show accuracy for exactly eight speakers.

This breakdown shows why you need to check the number of speakers and the recording conditions when choosing what to compare against. For two-person conversations, the overall first-place ranking alone is not a reason to choose it. In this evaluation's 5–8 speaker group, the new model came out ahead, but an error rate of 22.18% remains. According to the breakdown of the evaluation audio, conversations with six to eight people are limited to phone calls, so the same gap cannot be assumed for in-person meetings with many participants. Because a wrong speaker attribution can change who appears to have made a commitment in the minutes, a design that lets users check important statements against the original audio is essential.

The "eight" limit means the output has eight speaker slots. It is not a guarantee of accuracy when eight people talk at once or when participants have similar voices. Voice Arena's breakdown by speaker count shows why specification values and measured results for a given use case should be read separately.

NVIDIA's own like-for-like comparison also shows improvement. On the multilingual DIHARD III evaluation overall, with the input wait time set to 1.04 seconds, the new model's error rate was 13.18%, versus 19.60% for the old four-speaker model. However, this figure comes from different audio and scoring conditions than Voice Arena's 139 English conversations. The DIHARD III "5–9 speakers" group also includes nine-speaker audio, which exceeds the new model's stated limit. Percentages from different evaluations cannot be lined up side by side and treated as a single performance gap.

Another evaluation, of phone conversations, also shows that speaker counts cannot be lumped together. In the CALLHOME-Part2 results NVIDIA published for the 30.4-second setting, the old model had the lower error rate for two-speaker conversations: 5.68% versus 5.98% for the new one. Across the whole dataset, though, the old model scored 10.32% and the new one 9.10%. An improvement seen with more speakers should not be mechanically applied to smaller conversations.

AD

What the "real-time" latency figures measure

The model card lists input wait time settings of 30.4, 1.04, 0.64, and 0.32 seconds. Shorter settings reduce the wait between audio arriving and a speaker label being output, but this figure is the time until the model has received a chunk of audio plus a little audio ahead of it. It does not include computation itself, speech recognition, network transmission, or on-screen display. It cannot be read as "meeting minutes are complete in 0.32 seconds."

Changing how much audio is awaited also changes the context available for recognition. In NVIDIA's overall DIHARD III evaluation, the new model's diarization error rate was 12.73% at the 30.4-second setting and 13.18% at 1.04 seconds. That difference applies to these evaluation conditions and does not predict the same degradation in every meeting. For live captions on a device, you would test whether speaker mix-ups are tolerable with a short wait time. For minutes produced after a recording, longer settings are also an option.

The model card says it was trained on 82,611 hours of synthesized multi-speaker audio and about 10,000 hours of real conversations. There is no basis, however, for extending the published evaluation tables directly to the accuracy of Japanese-language meetings. Audio varies not only with the number of speakers but also with overlapping speech, room reverberation, and microphone placement. When adopting it, verify per-speaker errors and total processing time on the recordings you actually use.

What to check before building it into meeting transcripts

NVIDIA has released the model weights and says they can be used commercially and non-commercially. The applicable OpenMDW 1.1 license permits free use but requires the license text and applicable rights notices to be retained when redistributing the model materials. "Freely available" does not mean speech recognition and operations come free and complete.

A real product needs processing that synchronizes the speaker separation and transcription results. NVIDIA's official explainer also shows a flow that combines speaker timing information with speech recognition. When voices overlap, deciding which speaker a word belongs to also becomes harder. Even if anonymous labels stay stable during a conversation, that alone does not guarantee the real names of participants.

The first thing to check is how many people appear in your audio and how much they overlap. Then measure how often speakers are confused in your Japanese meetings and with your microphones, and time the entire pipeline, including transcription. The eight-speaker specification widens the candidates, but measured results will decide whether to adopt it.