Skip to content
edict
Guides

How to transcribe audio to text on a Mac.

You have a file; you want words. On a modern Mac that takes seconds, costs somewhere between nothing and one payment, and doesn’t require the audio to leave the machine.

To transcribe audio to text on a Mac: drop the file into an on-device transcription app. Local models handle MP3, WAV, M4A, and video files, and on Apple silicon turn an hour of audio into text in well under a minute; Edict does it in about 24 seconds with speaker labels and SRT, VTT, or text export. Free routes exist too: MacWhisper’s free tier, Spokenly’s free local mode, and Apple’s Voice Memos transcription for its own recordings.

Last verified 2026-07-24

What you need

A Mac and the audio file. That’s the list; Apple silicon is what makes it fast, and the speed figures below assume it. The transcription models that used to justify cloud subscriptions run locally now; the practical questions left are which app, whether the transcript needs speakers separated, and where the audio is allowed to go.

Format almost never matters: MP3, WAV, M4A, and video containers like MP4 and MOV all transcribe directly.

The short version

  1. 01

    Pick the route. Free: MacWhisper’s free tier or Spokenly’s free local mode, both processing on the Mac and both covered in detail on the free-routes page; if the recording lives in Voice Memos, macOS 15 on Apple silicon transcribes it in place. One-time paid: Edict, when speaker labels or agent access matter.

  2. 02

    Drop the file onto the app. On Edict an hour of audio comes back in about 24 seconds, roughly 150x real time, with each speaker labeled.

  3. 03

    Export what the next step needs: plain text for notes and AI assistants, SRT or VTT when the timestamps should survive (subtitles, quotes you’ll need to find again in the recording).

  4. 04

    Check the hard names. Every model mishears product names and jargon; a one-minute skim against the original beats trusting any tool blindly, local or cloud.

Multiple voices need labels

A transcript of two people without speaker labels is a riddle. For interviews, meetings, and panels, use a tool with diarization: Edict does it on any tier, MacWhisper on Pro, both on the Mac. The source-specific walkthroughs on this site (Zoom recordings, lectures, voice memos) cover the finding-the-file part of each job.

When a cloud service is the right call

Collaboration, not transcription. Team libraries, shared highlights, meeting bots that join calls for you, and paid human review for transcripts that must be right; that’s the Otter and Rev end of the market, and if a team workspace is the point, a local app isn’t the tool. For the transcription itself, local models give nothing away, and the file never leaving your Mac is a property no service can offer. The full comparison of tools goes tool by tool.

The agent route

If an AI agent is part of your workflow, transcription stops being a manual step at all. Edict’s MCP server lets Claude Code or any MCP-capable agent submit an audio or video file, get the transcript back with speaker labels, and carry on working, with the processing done on your Mac. "Transcribe this interview and pull every quote about pricing" becomes one instruction instead of an afternoon.

Common questions

How do I transcribe an audio file to text for free on a Mac?

Two apps do it free on the Mac: MacWhisper’s free tier and Spokenly’s local mode. Voice Memos recordings can also be transcribed in place on macOS 15 with Apple silicon (not every country). Free cloud tools exist too, but they involve uploading the audio.

How long does transcription take?

On-device on Apple silicon: well under a minute for an hour of audio. Edict runs at roughly 150x real time, so a one-hour file takes about 24 seconds and a ten-minute file takes about four.

What audio formats can be transcribed?

The common ones all work directly: MP3, WAV, M4A, and video files like MP4 and MOV. If you have something exotic, convert it once with any audio tool and proceed.

How accurate is on-device transcription?

Whisper- and Parakeet-family models do the work on both sides of the network; the cloud has no better model waiting. Clear speech comes back clean; jargon and names are every model’s weak spot, which is why the last step of any transcription job is a skim for the hard words.

Can I automate transcription with an AI agent?

Yes. Over Edict’s MCP server an agent submits the file, the Mac does the transcription, and the agent works from the text: summaries, quote extraction, action items, whatever you asked for. No audio is uploaded at any point.

Apple Silicon (M1 or newer) · macOS 15 Sequoia or later · 7-day full trial · one-time purchase