Insight 7 min read Grafite Team
Why 'Who Said What' Is the Hardest Part of a Meeting Transcript
The Feature Everyone Promises and Almost Nobody Explains
Ask any meeting transcription tool if it can tell you who said what, and the answer is always yes. Ask how, and the answers get vague fast. That gap is not an accident. Speaker labeling (the technical term is diarization) is one of the hardest problems in the whole meeting-notes pipeline, and it is a different kind of hard than transcription itself.
Transcribing speech to text is a mature, well-understood problem. Figuring out which voice belongs to which person, especially in a live conversation with no visual cues, no name tags, and often no separate audio feed per speaker, is not. We spent real time this year building a version of it, measuring what it actually produced, and deciding what belonged in the product and what did not. This post is about what we learned, in plain terms, without pretending it is simpler than it is.
Why a Bot Makes This Easier, and Why We Don't Use One
The fastest way to solve speaker labels is to put a bot in the meeting. A visible participant named something like "Notetaker" joins the call, and because it is inside the meeting platform, it can often read participant names straight off the window, or in some cases request a separate audio stream for each person. That is a real shortcut, and it is why bot-based tools can advertise clean speaker labels with more confidence.
Grafite does not put a bot in your calls. We record from your device, browser or desktop, the same way you would use a voice memo app, so the meeting itself never changes and no extra participant joins. That choice is deliberate: it keeps the conversation natural and works on conversations that never go through a video platform at all: a phone call, an in-person meeting, a hallway chat. But it means we do not get the shortcut. We cannot read names off a participant list that does not include us, and we do not get a separate, clean audio stream for each person. We are listening to one mixed signal, the same way a human in the room would, and trying to work out who is talking.
There is a workaround some tools use: pull the recording from the meeting platform after the call. Zoom and similar platforms sometimes offer that, but it has its own friction. It usually requires the meeting host's permission, sometimes admin-level account access, and the recording is not available until well after the call ends, which defeats the point of live notes.
What Actually Makes This Hard
A few specific problems come up over and over once you try to build this without a bot.
Your own speakers bleed into your microphone. If you are not wearing headphones, the other person's voice comes out of your speakers and back into your mic, faintly, but the recording software cannot always tell that apart from you actually talking. This muddies the simplest possible split, you versus everyone else.
Overlapping speech is the worst case. People talk over each other constantly in real conversations, and untangling two voices in the same half-second of audio is far harder than labeling clean turns. On the hardest calls we measured, accuracy at separating the different far-end voices from each other topped out around one in three.
A mixed call sounds like one person. When several remote people are on a single audio feed, without separate streams for each of them, a system trying to cluster voices by sound alone often hears them as one blob, "the other side," rather than as distinct individuals.
Timing has to line up with the words. A label is only useful if it lands on the exact sentence someone said. Small differences in clock timing between the audio and the transcript can split a single sentence across two speakers, which reads as nonsense.
Voice clustering can badly over-split one person. We saw a real hour-long call where an early version of this pipeline heard one person as eighteen different "speakers," each a slightly different slice of the same voice. That is a failure mode that looks technical but produces something a reader immediately distrusts.
The compute cost is real. Processing steps like this add real overhead on a device: our own measurements put some of these steps at roughly a minute of processing per hour of audio, plus close to a gigabyte of memory on hour-long calls. That is workable, but it is not free, and it is one more reason this cannot be an afterthought bolted onto transcription.
In September we built a version of this that ran entirely on your own computer, tested it against real calls, and decided not to ship it. The results a person would actually read were not good enough yet, even where the underlying numbers looked promising on paper. That is the lesson we keep coming back to: measure what the user reads, not the lab score.
It Also Costs More
Speaker labels are not a free add-on to transcription. Plain speech-to-text is cheap and getting cheaper. Speaker-aware transcription runs a heavier engine, and the services that do it well charge more for every hour of audio they process. If you want truly separate audio for each participant, you pay for each of those streams, and every extra stream multiplies the bill. Bots carry their own costs too: someone has to run a virtual participant in every meeting, for its full length, whether anyone is speaking or not.
Doing this live makes it harder again. A system that labels speakers after the call can take its time and look at the whole recording at once. A live one has to decide who is talking a few seconds after they start, with no idea what comes next, and keep a connection open for the entire meeting. That is why speaker labels sit on our Professional plan rather than the free one. The cost is real, and we would rather be clear about that than quietly lose money on it.
What Actually Ships
On the Professional plan, Grafite's cloud transcription labels speakers live, as the call happens. When a meeting has two or more speakers, the AI suggests likely names, pulled from the meeting's attendees, so you are not stuck naming voices "A" and "B" the whole way through.
Because we do not claim this is perfect, we built the correction tools to match. Click a speaker label and you can rename the voice everywhere it appears in that transcript, merge two labels that turned out to be the same person, or reassign just one turn, or even a single highlighted phrase, to the right person without touching the rest. A correction made while you are still on the call is worth more than one made an hour later, because you still remember who was talking.
One more honest detail: streaming speaker engines like the one behind this feature track a limited number of concurrent voices at once, and the letters they assign can get reused as a call goes on. That is exactly why the correction tools matter as much as the initial guess does.
Be Skeptical of "Perfect"
If a transcription tool tells you its speaker labels are always accurate, ask how. If the answer involves a bot in your meetings, that helps, but it is still not perfect: crosstalk and shared room audio trip up even name-aware systems. If there is no bot and no separate audio stream per person, be more skeptical still. This is a genuinely hard problem, and any tool worth using should give you a fast way to fix what it gets wrong rather than asking you to trust it blindly.
A Few Things That Actually Help
If you want cleaner speaker labels on your own calls, a handful of habits go a long way. Headphones stop your speakers from bleeding into your mic, which removes one whole category of confusion. In a room with several people, one microphone per person, or at least per side of the table, beats a single laptop mic trying to hear everyone. Saying names out loud early in a call, "thanks, Priya, and Marcus, what do you think," gives the system, and anyone reading the transcript later, something concrete to anchor on. And when a label is wrong, correct it once. On Grafite, that correction applies across the whole note, not just the one place you noticed it.
Speaker labels are useful even when they are imperfect, as long as you can trust that "imperfect" is honestly disclosed and quickly fixable. That is the bar we are building to, not a marketing claim about never getting it wrong.
Read more about how the pieces of your meeting notes connect in building a meeting-powered knowledge base, or see how conversational search works once your transcripts are captured. Try Grafite, Grafite Core is free forever.
Share this article
