Go ad free
Go ad free

Free online transcription

Free speech to text converter.

Choosing the correct language can improve speed and spelling. Auto-detect works for multilingual audio.

Use a free Groq API key

Groq Free Plan

480 audio min/day

120 min/hour30 min/file here

Groq's current Whisper allowance. Limits apply to your Groq organization.

  1. 1Get key
  2. 2Paste it
  3. 3Validate

Shared limit: 5 minutes per 24 hours without an account, or 10 minutes when signed in. Your file is sent securely to Groq and is not stored by FreeTextoSpeech.

Free Speech to Text turns a short recording into editable text and timed captions. Upload common audio formats, choose the spoken language, and download the result without installing transcription software.

This page uses the same shared converter as Audio to Text, so transcription limits, privacy controls, and advertising stay consistent.

5–30 min

4 files

No

Go ad free
Speech to text guide turning audio into a transcript and captions for free speech to text

Visual guide

Transcription workflow for free speech to text.

Less replaying, more using

Why turn speech into text?

A recording preserves tone and detail, but finding one sentence can mean scrubbing through minutes of audio. A transcript makes the same material searchable, editable, and easier to share. It gives you a working document for notes, quotations, captions, summaries, or articles.

This free speech to text tool creates that first draft automatically. You still check proper names, numbers, and specialist vocabulary, but you avoid typing every sentence by hand. The editable result stays useful even when you need to correct a word or add speaker labels.

The converter also returns timed segments. That means one upload can provide readable copy for a document and caption files for a video editor or web player.

A four-step workflow

How free speech to text works.

The converter keeps the upload, language choice, editable transcript, and downloads in one place.

  1. Choose your recording

    Upload up to five minutes and 20 MB with the shared service, or 30 minutes and 25 MB with your own Groq key.

  2. Set the spoken language

    Leave automatic detection on, or select the language when you know it. A specific language can improve spelling and speed.

  3. Create the transcript

    The same speech-to-text engine used by our Audio to Text tool processes the recording and returns editable words plus timed segments.

  4. Review and export

    Correct names or specialist terms, add speaker labels, then copy the text or download TXT, SRT, VTT, or JSON.

Four useful exports

Choose the right transcript format.

You can download more than plain text from the same transcription.

FormatBest forWhat it includes
.txtEditing and repurposingClean transcript text
.srtVideo subtitlesNumbered, timed caption blocks
.vttWeb video captionsBrowser-friendly timed cues
.jsonApps and workflowsText, language, duration, and segments

Useful starting points

What can you transcribe?

Free speech to text is most useful when the recording is clear, short, and valuable enough to search or reuse later.

01

Interviews

Turn a recorded conversation into searchable notes, quotations, and a first editing pass.

02

Lectures and lessons

Create study notes or captions from clear classroom recordings and short teaching clips.

03

Meetings and voice notes

Convert decisions, reminders, and spoken ideas into text you can scan and organize.

04

Podcasts and video

Prepare show notes, subtitles, rough scripts, and searchable excerpts from published audio.

Accuracy starts before upload

Help the model hear the words.

Automatic transcription works best with direct speech, steady volume, and limited background sound. The model can handle many accents and languages, but it cannot recover words that the recording itself does not capture clearly.

  • Reduce competing sound.

    Move closer to the microphone and reduce music, traffic, fans, and room echo where possible.

  • Select the language.

    Automatic detection is convenient; choosing the known language can improve recognition and shorten processing.

  • Review high-stakes details.

    Check names, dates, numbers, brands, acronyms, and technical vocabulary before publishing the transcript.

Clear data handling

What happens to your audio?

Your browser uploads the selected recording to the FreeTextoSpeech API over HTTPS. The server checks the file type, size, and duration before forwarding the audio to Groq for transcription. FreeTextoSpeech's shared keys never enter your browser. If you opt to use your own key, it stays in your browser tab and is sent securely with the transcription request without being saved to our database.

FreeTextoSpeech does not save the recording or resulting transcript. For shared-key requests, we retain limited usage metadata—such as duration, file size, status, and an account identifier when signed in—to enforce the shared daily allowance and prevent abuse. Read the complete details in our privacy policy.

Tool interface mockup for free speech to text showing input controls and the finished result

In context

How the free speech to text tool presents input and output.

Common questions

Free speech to text FAQ.

Is free speech to text really free to use?

Yes. You can transcribe up to five minutes per 24 hours without an account. Signed-in users receive ten minutes per 24 hours. There is no paid plan required for these limits.

What audio files can I convert to text?

The converter accepts FLAC, MP3, MP4, MPEG, MPGA, M4A, OGG, WAV, and WEBM. Shared uploads allow five minutes and 20 MB; using your own Groq key raises the per-file limit to 30 minutes and 25 MB.

Can free speech to text create subtitles?

Yes. Along with editable plain text, you can download SRT and VTT caption files using the timestamped segments returned by the transcription model.

Does FreeTextoSpeech store my recording?

No. The selected file is transferred securely through the FreeTextoSpeech API to Groq for transcription. FreeTextoSpeech does not save the audio or transcript.

Will the transcript identify each speaker?

Automatic speaker diarization is not currently included. After transcription, use the Insert speaker control to add speaker labels while reviewing the editable transcript.

How accurate is automatic speech recognition?

Accuracy depends on recording quality, background noise, speaker clarity, accent, language selection, and specialist vocabulary. Always review names, numbers, and important quotations.