Improved Speaker Diarization & Persistent Speaker Tagging
Speaker diarization currently struggles with accurately identifying speakers, especially in single-speaker conversations where the system should confidently recognize one speaker at all times. Additionally, users should have the ability to tag speakers manually, and those tags should persist across conversations, reducing the need for repeated identification.
Challenges
Inconsistent speaker recognition: A user speaking solo should be identified with 100% accuracy.
Limited diarization accuracy: The current model (Deepgram’s diarization) misidentifies or fails to consistently assign speaker labels.
Lack of persistent speaker tags: Users must repeatedly tag speakers, as the system does not remember identities across different conversations.
Log in to comment and vote
Comments13
infdaze
Mar 3, 2025
Omi has a hard time If i’m listening to an external source like Youtube. It will often think I am talking and create erroneous entries in “About Me” and “Chat”
MooseSoup
Oct 11, 2025
•Merged request
•1 vote
Serious diarization issues
Multiple instances of other people being labeled as speaker zero even when I have no other speaker tags and a voice profile added. Reviewed audio files and something is chopping up the files and cutting speech parts off and that directly aligns with the changes back to speaker zero when it's not me.
At this point the way this is working it is no better than a personal note taking device where only I speak to it. My old compass AI device did that just fine. I want something I can build off of that will actually improve my life but I can't build off of broken transcripts and a device that can't recognize when I'm speaking vs someone else when my voice profile is the only one it has to go off of.
What is being said and who said it is the closest thing to truth the whole system has to go off of. This should be the foundation, it should be rock solid before anything else.
MooseSoup
Aug 11, 2025
Do I have a defective unit or is this all software?? I just had a 6 minute phone call on speaker. Raw transcript only says 1 minute 30 seconds. Transcript in app says 1 minute 9 seconds. Cannot tell the delifference between me and the caller merged some of what I said and what he said together and finally identified my voice on the last snippet it took of me. Then I go and look in my Google drive and listen to the audio and there's 2 30 second audio clips. Both are exactly the same clip…. I just don't know what to do anymore.
Ray Y
Oct 11, 2025
•Merged request
•2 votes
Omi cannot recognise other speakers
Started with not recognition own voice, then i let device keep running until battery runs out, turned on, “Introduced myself” again.
Now it can recognise my voice but not others. Even tho i have tagged the speaker before.
Please help
Ray Y
Aug 19, 2025
iPhone 13 Pro , IOS version 18.5
HollyAnn
Feb 25
•Merged request
•2 votes
Speaker Tagging
Request:
Ability to reset or remove all speaker assignments for an entire transcript.
Ability to select multiple messages and batch assign speakers in a single action.
Problem:
If a speaker is incorrectly assigned, each message must be manually selected and reassigned one by one.
If a speaker is incorrectly assigned, the “tag other segments” feature does not apply correctly, which makes correcting mistakes especially time consuming.
Related:
Speaker selection currently defaults to alphabetical order. It would be helpful if it remembered the last used speaker or defaulted to “You.”
nathan
Feb 25
•Merged request
•3 votes
Default Mic Speaker Attribution to User (Omi Desktop)
Problem / Context
When using Omi Desktop as my primary setup, the microphone input almost always represents me (the user). However, the system currently attempts speaker detection on microphone audio, which can result in unclear or incorrect speaker attribution in meeting notes.
In most real-world use cases:
The microphone = the user
All other voices come from system / speaker audio (calls, videos, meetings)
There is typically only one person speaking on the mic
Because of this, meeting notes would be significantly clearer if microphone audio were consistently attributed to the user by default.
Proposed Solution
Add an option to label microphone audio as the user by default, bypassing speaker detection for mic input unless explicitly enabled.
Microphone audio → Always attributed to the user’s name
System / speaker audio → Continue using speaker detection as usual
This assumes a single speaker on the mic (the user), which matches most desktop use cases
Optional Controls / Flexibility
To handle edge cases, add a toggle such as:
“Assume mic audio is me (default)”
ON: Mic audio is always attributed to the user
OFF: Enable speaker detection on mic audio (for shared mic scenarios)
This allows advanced users to override the default behavior when:
Multiple people share a microphone
A different speaker is using the mic
Group recordings are happening in the same physical space
Why This Matters
Much clearer and more accurate meeting notes
Reduces incorrect speaker labeling
Matches how users actually use Omi Desktop in practice
Simple default behavior with an escape hatch for edge cases
Summary
Assume one speaker on mic = the user, by default.
Let system audio handle everyone else.
Provide a toggle only when needed.
This small change would significantly improve transcript clarity and usability for desktop-first users.