Skip to main content

Improved Speaker Diarization & Persistent Speaker Tagging

Speaker diarization currently struggles with accurately identifying speakers, especially in single-speaker conversations where the system should confidently recognize one speaker at all times. Additionally, users should have the ability to tag speakers manually, and those tags should persist across conversations, reducing the need for repeated identification.

Challenges

  • Inconsistent speaker recognition: A user speaking solo should be identified with 100% accuracy.

  • Limited diarization accuracy: The current model (Deepgram’s diarization) misidentifies or fails to consistently assign speaker labels.

  • Lack of persistent speaker tags: Users must repeatedly tag speakers, as the system does not remember identities across different conversations.

Status: Completed13 comments

Log in to comment and vote

Comments13

  • infdaze

    •

    Mar 3, 2025

    Omi has a hard time If i’m listening to an external source like Youtube. It will often think I am talking and create erroneous entries in “About Me” and “Chat”

  • MooseSoup

    •

    Oct 11, 2025

    •

    Merged request

    •

    1 vote

    Serious diarization issues

    Multiple instances of other people being labeled as speaker zero even when I have no other speaker tags and a voice profile added. Reviewed audio files and something is chopping up the files and cutting speech parts off and that directly aligns with the changes back to speaker zero when it's not me.

    At this point the way this is working it is no better than a personal note taking device where only I speak to it. My old compass AI device did that just fine. I want something I can build off of that will actually improve my life but I can't build off of broken transcripts and a device that can't recognize when I'm speaking vs someone else when my voice profile is the only one it has to go off of.

    What is being said and who said it is the closest thing to truth the whole system has to go off of. This should be the foundation, it should be rock solid before anything else.

    • MooseSoup

      •

      Aug 11, 2025

      Do I have a defective unit or is this all software?? I just had a 6 minute phone call on speaker. Raw transcript only says 1 minute 30 seconds. Transcript in app says 1 minute 9 seconds. Cannot tell the delifference between me and the caller merged some of what I said and what he said together and finally identified my voice on the last snippet it took of me. Then I go and look in my Google drive and listen to the audio and there's 2 30 second audio clips. Both are exactly the same clip…. I just don't know what to do anymore.

  • Ray Y

    •

    Oct 11, 2025

    •

    Merged request

    •

    2 votes

    Omi cannot recognise other speakers

    Started with not recognition own voice, then i let device keep running until battery runs out, turned on, “Introduced myself” again.

    Now it can recognise my voice but not others. Even tho i have tagged the speaker before.

    Please help

    • Ray Y

      •

      Aug 19, 2025

      iPhone 13 Pro , IOS version 18.5

  • HollyAnn

    •

    Feb 25

    •

    Merged request

    •

    2 votes

    Speaker Tagging

    Request:

    • Ability to reset or remove all speaker assignments for an entire transcript.

    • Ability to select multiple messages and batch assign speakers in a single action.

    Problem:

    • If a speaker is incorrectly assigned, each message must be manually selected and reassigned one by one.

    • If a speaker is incorrectly assigned, the “tag other segments” feature does not apply correctly, which makes correcting mistakes especially time consuming.

    Related:

    • Speaker selection currently defaults to alphabetical order. It would be helpful if it remembered the last used speaker or defaulted to “You.”

  • nathan

    •

    Feb 25

    •

    Merged request

    •

    3 votes

    Default Mic Speaker Attribution to User (Omi Desktop)

    Problem / Context

    When using Omi Desktop as my primary setup, the microphone input almost always represents me (the user). However, the system currently attempts speaker detection on microphone audio, which can result in unclear or incorrect speaker attribution in meeting notes.

    In most real-world use cases:

    • The microphone = the user

    • All other voices come from system / speaker audio (calls, videos, meetings)

    • There is typically only one person speaking on the mic

    Because of this, meeting notes would be significantly clearer if microphone audio were consistently attributed to the user by default.


    Proposed Solution

    Add an option to label microphone audio as the user by default, bypassing speaker detection for mic input unless explicitly enabled.

    • Microphone audio → Always attributed to the user’s name

    • System / speaker audio → Continue using speaker detection as usual

    • This assumes a single speaker on the mic (the user), which matches most desktop use cases


    Optional Controls / Flexibility

    To handle edge cases, add a toggle such as:

    • “Assume mic audio is me (default)”

      • ON: Mic audio is always attributed to the user

      • OFF: Enable speaker detection on mic audio (for shared mic scenarios)

    This allows advanced users to override the default behavior when:

    • Multiple people share a microphone

    • A different speaker is using the mic

    • Group recordings are happening in the same physical space


    Why This Matters

    • Much clearer and more accurate meeting notes

    • Reduces incorrect speaker labeling

    • Matches how users actually use Omi Desktop in practice

    • Simple default behavior with an escape hatch for edge cases


    Summary

    Assume one speaker on mic = the user, by default.

    Let system audio handle everyone else.

    Provide a toggle only when needed.

    This small change would significantly improve transcript clarity and usability for desktop-first users.