Voices become people
MinuteMark locates speech, transcribes it, matches each passage to a voice and groups those passages into speakers. Every speaker gets a colour and a label you can rename.
Private speaker-aware transcription
Meetings, interviews and lectures become searchable transcripts that identify not only what was said, but who said it. Every model is already inside the app.
macOS 14 or later · Apple silicon + Intel · No account
Let’s start with the decision we need to make today.
The customer interviews point in one clear direction.
I’ll capture the next steps and send them this afternoon.
Raw transcription gives you words.
MinuteMark locates speech, transcribes it, matches each passage to a voice and groups those passages into speakers. Every speaker gets a colour and a label you can rename.
Change “Speaker 2” to a real name and it flows through the transcript and every export. Speaker identity stays consistent from review to handoff.
Tell MinuteMark the head count, or let it work the number out. A sensitivity control decides how readily similar voices are treated as one person.
Talk time is shown for each speaker in minutes and as a share of the conversation. Speech coverage shows whether a recording was mostly talk or mostly silence.
Consecutive fragments from the same person are merged into natural turns, so one paragraph is not split into a dozen rows. Search finds any word across the transcript.
Export Markdown with a talk-time table, plain text, WebVTT with speaker voice spans, CSV or structured JSON.
Offline by architecture
Every speech and speaker model ships inside MinuteMark. There is no account to create and no network connection involved in processing.
A recording with no speech is reported as such. A stretch that cannot be attributed to a voice is dropped instead of being assigned to the wrong speaker.
Five useful formats
MinuteMark for macOS
One purchase. No subscription. No account.