Transcribe audio to text
Free, no sign-up, and your files never leave your device.
Drop your files here
Choose filesNothing is uploaded: everything happens in your browser
The first transcription downloads the speech recognition model once, then your browser keeps it: about 75 MB in standard accuracy, 240 MB in high accuracy. The video itself never leaves your device.
A voice memo, a WhatsApp message, a podcast, a meeting recorded on your phone: drop the file and speech recognition turns it into text, with a new paragraph at each silence. The recording stays on your device, which matters for a private conversation or a professional interview.
How to transcribe
- 1Drop your video or recording in the box, or click to pick it.
- 2Choose the format: plain text (TXT), SRT or VTT subtitles. The spoken language is detected automatically; you can also pick it in the settings.
- 3The first time, the speech recognition model downloads once and for all. After that, everything happens on your device: download the text or the subtitles.
The formats
MP3
The universal audio format. It compresses with loss, almost inaudibly at 192 kbps, and plays on every device, player and app. It is the format to aim for when sharing audio.
M4A
Apple's audio format, used by iPhone voice memos and iTunes. It is AAC in an MP4 container: good quality for a small size, but some players and websites do not accept it.
WAV
The uncompressed audio format of studios and editing software. It keeps full quality, but one minute of audio weighs about ten megabytes, ten times more than an MP3.
OGG
An open audio format, in Vorbis or Opus. It is the format of WhatsApp and Telegram voice messages. It is light and sounds good, but many players, editing tools and devices cannot play it.
TXT
Plain text, readable everywhere: the content of the recording, in paragraphs, with no formatting. Paste it into a document, an email, an article or an AI tool.
Your files stay with you
The conversion runs in your browser, with your device's decoders. No file is uploaded, no copy is kept, and nobody but you sees your photos or videos. You can check it in the Network tab of your browser's developer tools: none of your files goes out.
Frequently asked questions
Which files are accepted?
MP3, M4A (iPhone voice memos), WAV, OGG and Opus (WhatsApp and Telegram voice messages), FLAC, and also MP4, MOV and WebM videos.
Are the different speakers separated?
No, the text follows the conversation without naming who speaks. Silences mark the paragraphs, which helps spot changes of voice.
Is the converter really free?
Yes, with no sign-up, no watermark and no limit on the number of files. The conversion runs on your device and costs us nothing, so there is nothing to charge you for.
Are my files sent anywhere?
No. They are read and converted in your browser, then saved straight to your device. No server receives them.
Is there a maximum size?
The limit is your device's memory. Photos and audio files go through without trouble; for a video of several gigabytes, a computer does better than a phone.
Does it work on a phone?
Yes, on iPhone and Android, in Safari, Chrome or Firefox. For video, an up-to-date browser is needed: it provides the decoders being used.
What next, with AI
AI voice-over: the same text read by three engines
The same script read by three speech synthesis engines, ElevenLabs v3, MiniMax Speech 2.8 HD and OpenAI TTS HD, to choose your voice-over by ear. The full workflow, no account needed.
Open the tool