Live Transcription

Press record and your words stream straight to our AI — text appears seconds behind your voice, and a polished transcript lands when you stop.

No sign-up No watermark TXT · SRT · VTT exports Files auto-delete in 24h
Drag & drop your file here
or browse your files
Free: {0} files a day · up to {1} min & {2} MB each
0:00
Live draft
Words appear as you speak, transcribed live by our own AI — press Transcribe when you finish for the polished, editable version.
Microphone access was blocked — allow it in your browser's site settings, or upload a file instead.

Bigger files or more uploads? Free account: 5 files/day, 1-hour files · Pro: 10-hour files + speaker labels

0%
Uploading…
Keep this tab open — you'll be redirected to your transcript.
Talks & speechesReal-time notesIn-room speechAccessibility

Actually live: your speech pipes to our own AI while you talk

While you record, the page slices the microphone stream into short segments — a few seconds each, cut adaptively so the next one ships the moment the previous one comes back — and pipes each straight to our GPU servers, where the same Whisper-class model that powers the whole site transcribes it immediately. The text lands in the live feed seconds behind your voice and keeps flowing for as long as you keep talking. No browser speech service, no third-party dictation API — the live text is real Whisper output from our own hardware.

Because it only needs microphone recording, live mode works in every modern browser — Chrome, Edge, Safari, and Firefox, desktop and phone. Each audio segment is deleted from our servers the moment its text comes back.

The live feed vs the final transcript

Segment-by-segment transcription has one built-in limit: the model sees a few seconds at a time, so a word that straddles two segments can come out clipped, and punctuation across segment boundaries is guesswork. That's why the flow ends with one more step: the full recording is kept in your browser while you speak, and pressing Transcribe runs it through the model as a single take — full context, proper punctuation, timestamps, the editor, and TXT/SRT/VTT exports.

Use the live feed to confirm it's keeping up and to grab text the moment it's spoken (there's a copy button on the feed); use the final pass for anything you'll keep. For audio you already have as a file, upload mode on any tool page is the shorter path.

Frequently asked questions

Is this actually live, or recorded and sent afterwards?
Actually live: recording and transcription overlap. The page ships each few-second slice of audio to our GPUs while you're still speaking the next one, so text accumulates during the recording, not after it. Expect the feed to run a handful of seconds behind your voice — the time it takes a slice to finish, upload, and transcribe.
Why is the live text a few seconds behind me?
Three small delays stack up: a slice has to finish being recorded (three seconds or so), travel to the server, and pass through the model. That usually puts the feed about 4-8 seconds behind your voice — and if our GPUs are briefly busy, the page adapts by sending slightly longer slices instead of falling behind. Natural speech with pauses feels smooth.
Which browsers does live mode work in?
All modern ones — Chrome, Edge, Safari, and Firefox, on desktop and phones. It relies on standard microphone recording rather than a browser speech service, so it works in Firefox just as well as in Chrome.
Where does my voice go while live transcription runs?
To our servers only, and honestly described: live mode streams your audio to us in short segments as you speak — that's what makes it live — and each segment is deleted immediately after its text comes back. Nothing goes to Google, Apple, or any third-party speech service. The full recording stays in your browser until you press Transcribe, and anonymous transcripts auto-delete within 24 hours.
Can I use it to follow a meeting or lecture in the room?
As a personal aid, yes — put the device near the speaker and the feed tracks clear speech well, about ten seconds behind. Know the limits: one microphone, so distant voices and crosstalk degrade it, and it isn't a certified captioning service for accessibility compliance. Press Transcribe at the end and the full-context pass usually recovers what the live feed fumbled.
Which languages work live, and should I pick one?
The same 90+ languages as every transcription on this site — it's the same model. One tip specific to live mode: pick your language in the selector instead of leaving auto-detect on. Detection runs per segment, and seven seconds isn't much evidence, so pinning the language keeps the feed from wavering on short or accented phrases.