Who said what, and will it be remembered?
14 min read
We benchmarked eight speech systems, built on OpenAI, ElevenLabs, Deepgram, AssemblyAI, pyannoteAI and two open-source projects, on whether “who said what” holds up from one conversation to the next. Thine's research system made the fewest errors.
Our short paper on persistent speaker attribution was accepted to the IEEE SLT 2026 Demo Track.

Averaged across three settings, our research system's error rate across conversations (SI-cpWER) was 47.1%, against 54.8% for the next system, about 14% fewer errors. The technical report calls this system ThyVoice. In the charts and tables below it appears as Thine.
Every main table from the report is in Full results at the end, each with a short note on how to read it.
Why who said what has to hold up
AI notetakers and ambient assistants are only as good as their record of who said what. Summaries, action items, search and answers are all built on that record.
Maya says, “I'll send the revised budget on Friday.” In a later conversation she says, “The budget is ready, I've sent it.” Both transcripts can get her words exactly right. If the system stores Maya as two different people, though, the promise stays open, and a search for her outstanding commitments finds the promise without its resolution.
That is the problem Thine is built around. Thine learns from the in-person conversations of your day and keeps track of the open loops in them, and every open loop belongs to someone.
Standard accuracy metrics can't see this failure. Word accuracy (WER) ignores speakers entirely. Per-conversation scoring (cpWER) matches speakers afresh in every recording, so “Speaker 1” on Monday and “Speaker 1” on Friday never have to be the same person. We evaluated that directly, by checking whether each person keeps one identity across every recording (SI-cpWER).
Everyday conditions make this harder. People sit at different distances, rooms are noisy, people talk over each other, and the same person comes back later in a different room, on a different device, through a different microphone.
How we evaluated
We benchmarked eight speech systems. Five are commercial pipelines built on OpenAI (GPT-4o Transcribe Diarize), ElevenLabs (Scribe v2), Deepgram (Nova-3), AssemblyAI (Universal-3 Pro) and pyannoteAI (Precision-2). Two are open-source research systems, SE-DiCoW and WhisperX. The eighth is our research system.
Each comparison system produces a transcript with its own speaker labels for every recording. To link speakers across recordings, all seven use the same pyannoteAI voice-matching service; our research system uses its own. These results describe the configured systems above. They are not scores for each vendor's own identity product or for finished apps.
We used three settings:
- Office meetings. All 129 meetings in the NOTSOFAR evaluation set: 12 people, about 13 hours of audio, and roughly 30% overlapping speech. Each meeting is one recording from one microphone.
- The same meetings with added noise. Room noise, background chatter and short sounds such as keyboards and doors, mixed in at about 10 dB signal-to-noise on average.
- Dinner parties. Two real dinner-party recordings from CHiME-6, captured on far-field devices in homes, with eight people in total. Each recording is more than two hours long, so systems process it in roughly five-minute pieces, one after another.
Every setting starts with no known voices, and recordings arrive one after another in a fixed order. Nobody enrolls their voice in advance, so each system has to decide on its own when it is hearing someone new and when it is hearing someone again.
What we found
Strong results within one conversation did not carry over

Once the same person had to be recognized across conversations, our research system's error barely changed. Every other system's error rose, several of them sharply, including the system with the best word accuracy and the fewest errors within a conversation (cpWER).
Lower error than every commercial system in every setting

Across office meetings, the same meetings with added noise, and real dinner parties, our research system had lower error than every commercial system. The open-source SE-DiCoW led both office settings. At the dinner parties, our research system led all eight by 10 points.
The fewest words put in someone's mouth
A transcript can go wrong in two ways. It can say something that wasn't said: a wrong word, an extra word, or a real word given to the wrong person. Or it can leave words out.

Our research system made the fewest errors of the first kind (substitutions plus insertions), averaged across the three settings, and the fewest extra words in every setting. The trade-off is that it left out more words than the ElevenLabs, SE-DiCoW and OpenAI systems.
Overlapping audio is bad evidence for recognizing someone
When people talk over each other, the mixed audio can still be transcribed, but it is the wrong audio to trust when deciding who someone is.
We evaluated four ways of handling overlap on 20 office meetings. Two of them keep the overlap for transcribing and differ only in the audio used to recognize voices, clean or overlapping. Their error within a conversation is nearly identical (42.5 vs 42.9 cpWER), but recognizing from overlapping audio raised the error across conversations by more than half, from 41.4 to 67.8 (SI-cpWER). Separating overlapping voices first scored best.

This is how our research system handles overlap. It works out who is speaking when, separates two overlapping voices only where it can tell them apart with confidence, transcribes each person's audio on its own, and sets a higher bar for updating someone's voice profile than for recognizing them.
Full results
This section has every main table from the report. All numbers are error rates in percent, where lower is better, unless a column says otherwise. Mean is the simple average of the three settings. Clean is the office meetings, Noisy is the same meetings with added noise, and Dinner is CHiME-6. Thine is our research system. ElevenLabs results on the office settings are the average of three full runs. Per-setting breakdowns and the appendices are in the technical report.
What the metrics measure
- WER (word error rate): wrong, missing and extra words, ignoring who said them.
- cpWER: the same errors, but each person's words are scored separately, with speakers matched afresh in every recording.
- SI-cpWER: cpWER with one speaker mapping held across every recording, so each person has to keep the same identity throughout.
- Substitutions, deletions and insertions: wrong words, missing words and extra words. In cpWER, a real word given to the wrong person also counts here.
- DER and JER: how well a system tracks who is speaking when, before any words are involved.
- Purity and coverage: whether each speaker group a system creates holds one person (purity), and how much single-speaker speech gets any speaker label at all (coverage).
Errors across conversations (SI-cpWER)
| System | Clean | Noisy | Dinner | Mean |
|---|---|---|---|---|
| Thine | 34.89 | 51.27 | 55.24 | 47.13 |
| ElevenLabs | 36.13 | 53.89 | 74.22 | 54.75 |
| pyannoteAI | 44.80 | 69.35 | 65.30 | 59.82 |
| AssemblyAI | 48.26 | 58.30 | 77.41 | 61.32 |
| Deepgram | 58.71 | 79.16 | 90.03 | 75.97 |
| OpenAI | 61.09 | 76.01 | 80.47 | 72.52 |
| SE-DiCoW | 32.24 | 49.54 | 83.88 | 55.22 |
| WhisperX | 50.44 | 65.25 | 86.13 | 67.27 |
Our research system has the lowest mean and the lowest error of the commercial systems in every setting. SE-DiCoW has the lowest error on the two office settings.
Errors within a conversation (cpWER)
| System | Clean | Noisy | Dinner | Mean | Mean change across conversations |
|---|---|---|---|---|---|
| Thine | 35.53 | 51.42 | 53.76 | 46.90 | +0.23 |
| ElevenLabs | 33.57 | 43.58 | 44.25 | 40.47 | +14.28 |
| pyannoteAI | 39.51 | 52.71 | 56.43 | 49.55 | +10.27 |
| AssemblyAI | 37.24 | 51.36 | 63.77 | 50.79 | +10.53 |
| Deepgram | 55.14 | 74.08 | 88.87 | 72.70 | +3.27 |
| OpenAI | 47.22 | 63.86 | 71.45 | 60.84 | +11.68 |
| SE-DiCoW | 25.62 | 43.12 | 56.88 | 41.87 | +13.35 |
| WhisperX | 50.00 | 60.29 | 67.35 | 59.21 | +8.06 |
The last column is how much each system's mean changes from cpWER to SI-cpWER. It is the net effect of linking people across recordings, which can add errors or fix speakers that were split within a recording.
How many people each system thinks there are
| System | Clean (true: 12) | Noisy (true: 12) | Dinner (true: 8) |
|---|---|---|---|
| Thine | 14 | 21 | 17 |
| ElevenLabs | 14–16 | 23–28 | 27 |
| pyannoteAI | 75 | 101 | 80 |
| AssemblyAI | 14 | 22 | 20 |
| Deepgram | 26 | 26 | 16 |
| OpenAI | 71 | 124 | 62 |
| SE-DiCoW | 24 | 32 | 34 |
| WhisperX | 20 | 22 | 23 |
These are the identities each system ends up with after all recordings, against the true number of people. A count close to the true number is a good sign, but it does not prove the right words went to the right person. For the comparison systems, the shared pyannoteAI service did the matching. ElevenLabs' office counts are a range across three runs, and unknown-speaker labels are not counted.
What kind of errors each system makes
| System | Put in someone's mouth (substitutions + insertions) | Left out (deletions) |
|---|---|---|
| Thine | 13.21 | 33.69 |
| pyannoteAI | 14.43 | 35.12 |
| AssemblyAI | 16.55 | 34.24 |
| Deepgram | 18.43 | 54.27 |
| WhisperX | 19.65 | 39.56 |
| ElevenLabs | 20.85 | 19.62 |
| SE-DiCoW | 24.49 | 17.39 |
| OpenAI | 28.24 | 32.61 |
Means across the three settings, from the within-conversation scoring (cpWER). Substitutions plus insertions are words a transcript asserts wrongly: wrong words, extra words, and real words given to the wrong person. Deletions are words it leaves out. The report breaks these down by setting and by error type.
Word accuracy (WER)
| Words from | What was scored | Clean | Noisy | Dinner | Mean |
|---|---|---|---|---|---|
| ElevenLabs Scribe v2 | speaker-labelled transcript | 30.37 | 35.80 | 35.03 | 33.73 |
| AssemblyAI Universal-3 Pro | speaker-labelled transcript | 31.94 | 39.69 | 39.88 | 37.17 |
| Deepgram Nova-3 | speaker-labelled transcript | 31.68 | 50.31 | 82.60 | 54.86 |
| Parakeet TDT 0.6B v3 | recognizer on the original mixed audio | 32.69 | 42.26 | 48.43 | 41.12 |
| Qwen3-ASR-1.7B | recognizer on the original mixed audio | 33.11 | 43.44 | 52.27 | 42.94 |
| OpenAI GPT-4o Transcribe | standalone transcription API | 39.30 | 46.96 | 53.09 | 46.45 |
| WhisperX large-v2 | WhisperX recognizer output | 37.15 | 45.91 | 82.88 | 55.31 |
ElevenLabs had the best word accuracy in every setting. Two rows are recognizers run on the original mixed audio: Parakeet is the recognizer in pyannoteAI's pipeline, and Qwen3-ASR is the recognizer inside our research system. That Qwen3-ASR number is not our system's end-to-end word accuracy, because our system transcribes each person's own audio rather than the mix.
Who speaks when (DER and JER)
| System · component | Output | Clean DER / JER | Noisy DER / JER | Dinner DER / JER | Mean DER / JER |
|---|---|---|---|---|---|
| Thine · DiariZen md-v2 | native timeline | 18.32 / 23.65 | 29.51 / 37.14 | 68.24 / 68.75 | 38.69 / 43.18 |
| SE-DiCoW · DiariZen md | native timeline | 19.55 / 25.10 | 29.82 / 37.64 | 68.99 / 69.19 | 39.45 / 43.98 |
| pyannoteAI · Precision-2 | native timeline | 25.36 / 33.50 | 32.99 / 43.51 | 68.86 / 69.12 | 42.40 / 48.71 |
| WhisperX · pyannote Community-1 | native timeline | 30.36 / 38.42 | 37.50 / 48.08 | 72.35 / 73.81 | 46.74 / 53.44 |
| OpenAI · GPT-4o Transcribe Diarize | segments | 43.09 / 51.43 | 46.87 / 55.02 | 82.63 / 85.02 | 57.53 / 63.82 |
| ElevenLabs · Scribe v2 | rebuilt from word labels | 44.13 / 49.22 | 48.06 / 53.55 | 81.18 / 79.91 | 57.79 / 60.90 |
| AssemblyAI · Universal-3 Pro | rebuilt from word labels | 42.62 / 46.80 | 48.15 / 54.05 | 74.76 / 76.05 | 55.18 / 58.97 |
| Deepgram · Nova-3 | rebuilt from word labels | 41.97 / 48.07 | 59.78 / 68.14 | 92.55 / 92.44 | 64.77 / 69.55 |
The DiariZen md-v2 component in our research system had the lowest DER and JER in every setting. Outputs differ: some systems return a native timeline of who speaks when, OpenAI returns segments, and ElevenLabs, AssemblyAI and Deepgram label each word with one speaker. For those three we rebuilt speaker turns from the word labels, joining gaps of up to 0.25 seconds, so their scores partly reflect that rebuild.
Speaker evidence before voice profiles are made
| System · component | Output | Speaker-count error | Merged groups (%) | Purity (%) | Coverage (%) | Overlap found (%) |
|---|---|---|---|---|---|---|
| Thine · DiariZen md-v2 | native timeline | 0.36 / 0.50 | 1.7 / 11.2 | 96.2 / 91.2 | 95.2 / 92.0 | 78.5 / 53.0 |
| SE-DiCoW · DiariZen md | native timeline | 0.33 / 0.48 | 1.9 / 11.2 | 95.6 / 90.8 | 95.1 / 91.4 | 79.9 / 55.1 |
| pyannoteAI · Precision-2 | native timeline | 0.81 / 1.12 | 4.1 / 11.2 | 93.9 / 87.5 | 95.0 / 93.7 | 68.0 / 53.6 |
| WhisperX · pyannote Community-1 | native timeline | 0.33 / 0.87 | 6.6 / 19.2 | 91.2 / 84.4 | 92.8 / 92.0 | 51.7 / 40.0 |
| OpenAI · GPT-4o Transcribe Diarize | segments | 1.43 / 2.90 | 11.6 / 10.0 | 84.2 / 81.2 | 93.8 / 93.0 | 0.1 / 0.2 |
| ElevenLabs · Scribe v2 | rebuilt from word labels | 0.49 / 0.62 | 10.9 / 15.5 | 92.7 / 89.7 | 84.4 / 81.5 | 0.0 / 0.0 |
| AssemblyAI · Universal-3 Pro | rebuilt from word labels | 0.32 / 0.82 | 6.1 / 14.5 | 94.1 / 87.9 | 83.4 / 81.2 | 0.0 / 0.0 |
| Deepgram · Nova-3 | rebuilt from word labels | 0.71 / 1.78 | 13.1 / 37.5 | 89.6 / 79.2 | 90.1 / 71.2 | 0.0 / 0.0 |
Office settings only, shown as Clean / Noisy. These numbers describe the material a system could use to build voice profiles, not the quality of the profiles it ends up storing. Purity and coverage need to be read together, since high purity alone can hide speech that never got a label. Overlap found is the share of overlapping speech a system marks as overlapping. For outputs that label each word with one speaker, 0% means the output has no way to show two people at once; it says nothing about what the provider detects internally.
The overlap experiment
| How overlap is handled | Within one conversation (cpWER) | Across conversations (SI-cpWER) |
|---|---|---|
| Separate overlapping voices first | 34.39 | 32.73 |
| Leave overlap out | 55.39 | 54.12 |
| Keep overlap, recognize from clean audio | 42.53 | 41.42 |
| Keep overlap, recognize from overlapping audio | 42.86 | 67.76 |
Twenty office meetings, using our research system.
How we set up voice profiles for the comparison systems
| System | Clean: longest turn | Clean: pooled | Clean: change | Noisy: longest turn | Noisy: pooled | Noisy: change |
|---|---|---|---|---|---|---|
| pyannoteAI | 49.30 | 39.87 | −9.43 | 59.92 | 62.22 | +2.30 |
| ElevenLabs | 44.11 | 36.49 | −7.62 | 64.65 | 49.30 | −15.35 |
| AssemblyAI | 43.40 | 38.47 | −4.93 | 83.21 | 56.90 | −26.31 |
| OpenAI | 62.92 | 49.17 | −13.75 | 94.28 | 68.48 | −25.80 |
| Deepgram | 72.21 | 61.39 | −10.82 | 88.60 | 73.56 | −15.04 |
Errors across conversations (SI-cpWER) on 20 office meetings, one run each. When a comparison system meets a new speaker, the shared pyannoteAI service builds a voice profile from that speaker's audio. We compared building it from the speaker's single longest turn against joining all of their segments together (pooled). Pooling cut errors for every commercial system in the clean setting, by 4.9 to 13.8 points, and for four of the five in the noisy setting, by 15 to 26 points. pyannoteAI was the exception there. Our main results give every comparison system the pooled style of setup, up to 30 seconds of each speaker's audio joined together, rather than a single turn.
Scope
- English conversations, one microphone per recording.
- Two datasets in three settings. The dinner-party setting is two long recordings, and the noisy setting adds noise to the office recordings rather than capturing new audio.
- Recognition across conversations is evaluated by processing recordings in one fixed order, not by following the same people over weeks.
- We evaluated the speech technology inside these systems, not complete products.
Make this layer visible
This is the layer every AI notetaker and ambient assistant is built on. It should be reported as openly and comparably as word accuracy is today. We evaluated the speech technology rather than finished products, and we hold the Thine app to the same standard.
The full methodology is in the technical report, which expands on the accepted short paper.
We're continuing to improve this layer, and we'll keep sharing what we learn.
Shantanu Vispute and Siddhartha Saxena