
Visual guide
The online studio layout used for text to audio.
Prepare a script as named sections, choose a voice, and download one joined WAV or MP3 plus the individual clips. No signup.
Text-to-audio covers everything from a quick voice memo to a full audiobook master. The output format you actually need depends on where the audio is going. FreeTextoSpeech generates lossless 24 kHz WAV by default because it is the cleanest source for any downstream workflow, keep the WAV as your master, derive MP3 or streaming formats when distribution requires them.
Related use cases
Paste a script, prepare its named sections, pick a voice, and generate the selected queue. Download one joined WAV or MP3 plus individual clips when you need them. Cloud credits are shown before generation, and signed-in users can use the private in-browser engine.
Up to 5,000 characters per request. Split longer scripts into chapters or sections, each one becomes a separate audio file.
54 Kokoro voices across 9 languages. Sample a few before committing, voice choice matters more than any post-processing.
Synthesis runs in your browser session. No queue, no waiting room, no signup. Most reads finish in a few seconds.
You get a lossless 24 kHz WAV file by default. Need MP3 instead? Use the dedicated /text-to-mp3 page for the conversion workflow.

Visual guide
The online studio layout used for text to audio.
Drop the WAV straight into Premiere, DaVinci Resolve, or CapCut. Lossless input means cleaner ducking, EQ, and noise gating later.
Use TTS for show intros, segment bumpers, ad reads, or full episodes. WAV master goes into your DAW, MP3 ships to the host.
Catch awkward sentences and pacing problems by listening to your manuscript before recording. Cheaper than reshooting a chapter.
Convert articles, PDFs, and course modules into audio for dyslexic readers, commuters, or learners who absorb better by ear.
WAV is lossless, generate once, derive MP3 or AAC for distribution. Going the other way (MP3 to WAV) does not recover the lost frequencies.
For voice-only content, 128 kbps MP3 is sonically indistinguishable from the WAV master in blind tests. No reason to pay the file-size cost of 320 kbps for podcasts or audiobooks.
Human speech tops out around 8 kHz, so 24 kHz captures everything with headroom. 44.1 kHz or 48 kHz is overkill for narration and just inflates file size.
A 5-minute read is about 7 MB as WAV, 3.7 MB as 128 kbps MP3, or 1.2 MB as 64 kbps Opus. Pick the format your delivery channel actually accepts.
For web playback, host an MP3 or AAC and stream it. For DAW work, download the WAV. Mixing the two, streaming a WAV, wastes bandwidth without sounding any better through laptop speakers.
Adjust pacing later in your editor. Compounding speed changes at both generation and editing introduces pitch artefacts that are hard to undo.

In context
A step-by-step visual guide for text to audio.
Still wondering? Get in touch →
When you need a lossy format for podcast hosts, email, and web playback.
Lossless 24 kHz master files for DAW and video editing workflows.
Long-form narration with full commercial rights and no watermark.
Studio-quality reads for intros, ad spots, and full episodes.
Indie audiobook narration with the WAV masters production needs.
Drop in a PDF and get clean narration of every page.
One paste, one click, lossless WAV in seconds.
A tiny favor
Allow ads for this site, then check again. Prefer no ads? Support us to unlock ad-free access and 2 million cloud characters.
Allow ads for freetexttospeech.net, then check again.
Already supporting us? Sign in to restore your perks.
Feedback
Tell us what you think
Bugs, ideas, or anything that would make FreeTextoSpeech better.