
Visual guide
The online studio layout used for AI dialogue generator.
Host and Guest already have voices. Write the conversation, generate, and download one joined audio file.
The dialogue studio starts with two speakers already loaded. Write the conversation, generate, and download one finished file. Each person keeps a distinct voice.
Related use cases
Write one line per turn. Host and Guest already have voices. Click Generate and download one joined WAV or MP3.
One line per turn. Host says the first line, Guest the second, and they keep alternating. NAME: line tags still work if you already write scripts that way.
Host and Guest already have contrasting voices. Rename them, pick a different voice, or add someone. You never detect speakers from the script.
One Generate click builds the conversation. Play it, then download a joined WAV or MP3.

Visual guide
The online studio layout used for AI dialogue generator.
Two-host podcast where one AI voice asks questions and the other explains. Build the script, generate Host A and Host B with contrasting voices, splice. No subscription, no waitlist.
Voice every character in your storyboard or Twine prototype with a different Kokoro voice. 54 voices means an entire small cast without hiring VAs for the rough cut.
Narrate the prose with one voice (River or Daniel), then swap in distinct character voices for direct quotes. Listeners track who is speaking without a "she said" tag every line.
Generate two-speaker dialogues in Spanish, French, Hindi, Japanese, Mandarin, and 4 more languages. Pair a male and female native voice for realistic A/B exchange drills.
Dialogue lives or dies on whether listeners can tell the two voices apart without thinking. These six pairings contrast on gender, accent, or character tone, pick one and start generating.
Host + expert guest
Best for
NotebookLM-style explainer podcasts. Sarah asks the questions, Adam delivers the authoritative answer. Most natural pairing for two-host US English shows.
Casual conversation
Best for
Lifestyle pods, friend-chat formats, café-scene dialogue in audiobooks. Both voices read warm, so the back-and-forth feels relaxed instead of formal.
Period drama or BBC-style
Best for
Audiobook scenes set in the UK, prestige documentary two-handers, history podcasts with a host-and-narrator structure. The matched accent keeps the world consistent.
Animation / game characters
Best for
Animated shorts, Twine prototypes, indie game NPC dialogue. Puck is mischievous, Sky is bright, clear character voices, not narrator voices.
Documentary narrator + interviewee
Best for
True-crime, history, science docs where a smooth narrator threads between recreated quotes. River carries the throughline, Nova plays the voice in the archive.
Co-host energy
Best for
Morning-show banter, news-and-chat formats, branded podcast intros where two hosts trade a cold open before the show kicks in.
Want to hear them? Browse all 54 voices →
The model handles the voice. The realism comes from how you script the beats, the silences, and the small reactions between characters. Six rules that cover the full pipeline.
Host takes the first line, Guest the second, and they keep alternating. Optional NAME: line tags still work if you already write scripts that way.
A comma is a quarter-beat, a period is a half-beat, an em dash is a real pause, an ellipsis is a held beat. If a character is hesitating, write "Well... I don't know if that's true." instead of "Well I don't know if that's true." The pause is what sells the hesitation, and Kokoro respects the punctuation.
Real conversational gap-time averages around 250 ms. Slap two TTS clips back-to-back and the dialogue sounds robotic; add 300 ms of true silence between turns in your editor and it sounds like two people talking. For tense scenes, drop to 100 ms. For thoughtful exchanges, push to 600 ms.
Truncate Character A's clip mid-word, then drag Character B's clip to start 50–100 ms before A ends. The brief overlap is the universal audio cue for "they cut in." A 100 ms crossfade on the overlap zone smooths the splice and the interruption reads as natural.
Pan Character A 15% left and Character B 15% right. Not enough to feel theatrical, enough that headphone listeners stop conflating who's speaking. Hard panning (50%+) only works for radio drama. For podcasts and audiobooks, 15% is the sweet spot.
No SSML, no <laugh> tag, write what you want spoken. "Hah hah hah" delivers a clean three-syllable laugh, "mmhm" reads as agreement, "ugh" lands as exasperation. Test the reaction as a 10-character solo generation, swap voices if one delivers it cleaner, then drop the WAV onto a separate track in your edit.
PlayHT and Murf both ship native multi-speaker modes, paste once, render one mix. The trade is monthly character caps, signup, and paid commercial license. Here is the honest read.
Multi-voice dialogue support
FreeTextoSpeech
Write one line per turn, Host and Guest already have voices, and export one joined conversation.
PlayHT / Murf free tiers
Native multi-speaker dialogue feature, but locked behind paid tiers and capped voice roster on free.
Voice variety per scene
FreeTextoSpeech
54 voices across 9 languages, enough for an entire ensemble cast.
PlayHT / Murf free tiers
Smaller free voice pool with most distinct character voices paywalled.
Cost for a 5-minute dialogue
FreeTextoSpeech
Free. No character cap on number of generations.
PlayHT / Murf free tiers
Counts against monthly character cap on free tier; longer scenes push you to a paid plan.
Commercial license on free tier
FreeTextoSpeech
Full commercial use allowed, no attribution.
PlayHT / Murf free tiers
Commercial use restricted on free tier, paid plan required for monetized podcasts and indie games.
Output format
FreeTextoSpeech
24 kHz WAV per voice, drop straight onto separate tracks in any DAW.
PlayHT / Murf free tiers
Compressed MP3 on free tier, often single-mix only.
Signup
FreeTextoSpeech
None. Open the page, paste, generate.
PlayHT / Murf free tiers
Email signup required, credit card on file for commercial features.
One-click multi-voice generation
FreeTextoSpeech
Native multi-speaker generation with loaded speakers, joined export, and optional Advanced controls.
PlayHT / Murf free tiers
Native multi-speaker dialogue mode on paid tiers (one paste, one render).
Generate runs each turn in order so the shared voice engine is not overloaded.

In context
A step-by-step visual guide for AI dialogue generator.
Still wondering? Get in touch →
Studio-quality reads for two-host podcast intros and segments.
Indie audiobook narration with character-dialogue voice swaps.
Why Kokoro voices read as human across long dialogue scenes.
Drop dialogue scenes straight into Premiere or DaVinci Resolve.
Generate it free, in under 10 minutes, with full commercial rights.
A tiny favor
Allow ads for this site, then check again. Prefer no ads? Support us to unlock ad-free access and 2 million cloud characters.
Allow ads for freetexttospeech.net, then check again.
Already supporting us? Sign in to restore your perks.
Feedback
Tell us what you think
Bugs, ideas, or anything that would make FreeTextoSpeech better.