Model / API profile · Checked September 27, 2026

NVIDIA Nemotron 3 Diarization

NVIDIA Nemotron 3 Diarization identifies who spoke when in recorded or streaming audio. It produces speaker activity and timestamps, not a transcript: ASR is the separate task of recognizing what was said.

Release: 2026-09-23. Source-verified; not independently benchmarked.

Video Audio · Official starting point

Inputs, outputs and supported speakers

The approximately 100-million-parameter model accepts 16 kHz mono audio and produces activity probabilities for up to eight anonymous speaker channels. Postprocessing turns those probabilities into time segments. Multiple channels can be active during overlapping speech. Labels follow speaker arrival order; they do not identify a person by name.

Who it is for

Developers adding speaker attribution to meetings, interviews or call-processing pipelines. Pair the output with a separate ASR model and align its words with the speaker segments. A meeting app is a more direct choice if you need uploading, editing and export without maintaining inference code.

Availability, licence and cost

Weights and NVIDIA NeMo examples are public on Hugging Face under OpenMDW 1.1. The model card permits commercial and non-commercial use under that licence. Open weights do not include free hosting: plan for inference hardware, installation and integration. The card does not list a deployed Hugging Face Inference Provider at this check.

Streaming and long-audio limitations

Offline and streaming modes use the same model. Chunking, a speaker cache and recent-frame context remove a fixed model-imposed recording-duration limit. That is not a guarantee of reliable unlimited sessions. Noise, reverberation, domain changes and unusually long recordings can degrade results. Eight channels are the supported speaker limit, not unlimited participants.

What to validate before adoption

Use an authorized sample with known speaker turns and some overlap. Review missed speech, false detections, speaker swaps and boundaries separately from word errors in the ASR output. Check the combined pipeline before using its attribution for summaries or decisions. No local inference or independent benchmark was performed for this profile.

Official sources

Checked September 27, 2026. Access and prices can change; follow the linked provider documentation before use.

Related workflows

Compare hosted and self-hosted diarization routes.

Free transcription workflows · Meeting note tools