Multi-voice TTS

AI Dialogue Generator

Host and Guest already have voices. Write the conversation, generate, and download one joined audio file.

Go ad free
Go ad free

Dialogue studio

Host and Guest already have voices. Write the conversation, generate, and download one file.

Working tool

Speakers

Already loaded. Change a voice only if you want a different person.

2 speakers

H
G

Line 1 is Host. Line 2 is Guest. They keep taking turns.

166 characters
Advanced options

Source files are parsed locally. In Cloud mode, only the selected text sections are sent for speech generation. In-browser mode keeps both the source and generated speech on this device.

Go ad free
Two-voice TTS

Multi-character dialogue from a flat script

The dialogue studio starts with two speakers already loaded. Write the conversation, generate, and download one finished file. Each person keeps a distinct voice.

The quick answer

Write one line per turn. Host and Guest already have voices. Click Generate and download one joined WAV or MP3.

Dialogue workflow

Write, then generate

  1. 01

    Write the conversation

    One line per turn. Host says the first line, Guest the second, and they keep alternating. NAME: line tags still work if you already write scripts that way.

  2. 02

    Change a voice only if you want

    Host and Guest already have contrasting voices. Rename them, pick a different voice, or add someone. You never detect speakers from the script.

  3. 03

    Generate and download

    One Generate click builds the conversation. Play it, then download a joined WAV or MP3.

FreeTextoSpeech studio mockup for AI dialogue generator with script, voice controls, and waveform...

Visual guide

The online studio layout used for AI dialogue generator.

When to use it

What people build with multi-voice TTS

04 scenarios
01 / 04

NotebookLM-style podcast clones

Two-host podcast where one AI voice asks questions and the other explains. Build the script, generate Host A and Host B with contrasting voices, splice. No subscription, no waitlist.

02 / 04

Animated shorts and game prototypes

Voice every character in your storyboard or Twine prototype with a different Kokoro voice. 54 voices means an entire small cast without hiring VAs for the rough cut.

03 / 04

Audiobook character dialogue

Narrate the prose with one voice (River or Daniel), then swap in distinct character voices for direct quotes. Listeners track who is speaking without a "she said" tag every line.

04 / 04

Language-learning conversations

Generate two-speaker dialogues in Spanish, French, Hindi, Japanese, Mandarin, and 4 more languages. Pair a male and female native voice for realistic A/B exchange drills.

Voice guide

Six voice pairings that contrast cleanly

Dialogue lives or dies on whether listeners can tell the two voices apart without thinking. These six pairings contrast on gender, accent, or character tone, pick one and start generating.

01 US English

Sarah + Adam

Host + expert guest

Best for

NotebookLM-style explainer podcasts. Sarah asks the questions, Adam delivers the authoritative answer. Most natural pairing for two-host US English shows.

02 US English

Bella + Liam

Casual conversation

Best for

Lifestyle pods, friend-chat formats, café-scene dialogue in audiobooks. Both voices read warm, so the back-and-forth feels relaxed instead of formal.

03 British English

Daniel + Emma

Period drama or BBC-style

Best for

Audiobook scenes set in the UK, prestige documentary two-handers, history podcasts with a host-and-narrator structure. The matched accent keeps the world consistent.

04 US English (character)

Puck + Sky

Animation / game characters

Best for

Animated shorts, Twine prototypes, indie game NPC dialogue. Puck is mischievous, Sky is bright, clear character voices, not narrator voices.

05 US English

River + Nova

Documentary narrator + interviewee

Best for

True-crime, history, science docs where a smooth narrator threads between recreated quotes. River carries the throughline, Nova plays the voice in the archive.

06 US English

Michael + Jessica

Co-host energy

Best for

Morning-show banter, news-and-chat formats, branded podcast intros where two hosts trade a cold open before the show kicks in.

Want to hear them? Browse all 54 voices →

Best practices

How to write dialogue that reads as natural TTS

The model handles the voice. The realism comes from how you script the beats, the silences, and the small reactions between characters. Six rules that cover the full pipeline.

  • 01

    Write one line per turn

    Host takes the first line, Guest the second, and they keep alternating. Optional NAME: line tags still work if you already write scripts that way.

  • 02

    Use punctuation to script the beats between lines

    A comma is a quarter-beat, a period is a half-beat, an em dash is a real pause, an ellipsis is a held beat. If a character is hesitating, write "Well... I don't know if that's true." instead of "Well I don't know if that's true." The pause is what sells the hesitation, and Kokoro respects the punctuation.

  • 03

    Leave 200–400 ms of silence between turns

    Real conversational gap-time averages around 250 ms. Slap two TTS clips back-to-back and the dialogue sounds robotic; add 300 ms of true silence between turns in your editor and it sounds like two people talking. For tense scenes, drop to 100 ms. For thoughtful exchanges, push to 600 ms.

  • 04

    Fake interruptions with a 50–100 ms overlap

    Truncate Character A's clip mid-word, then drag Character B's clip to start 50–100 ms before A ends. The brief overlap is the universal audio cue for "they cut in." A 100 ms crossfade on the overlap zone smooths the splice and the interruption reads as natural.

  • 05

    Pan voices slightly left and right for headphone clarity

    Pan Character A 15% left and Character B 15% right. Not enough to feel theatrical, enough that headphone listeners stop conflating who's speaking. Hard panning (50%+) only works for radio drama. For podcasts and audiobooks, 15% is the sweet spot.

  • 06

    Spell phonetic reactions for laughter and "uh-huh"

    No SSML, no <laugh> tag, write what you want spoken. "Hah hah hah" delivers a clean three-syllable laugh, "mmhm" reads as agreement, "ugh" lands as exasperation. Test the reaction as a 10-character solo generation, swap voices if one delivers it cleaner, then drop the WAV onto a separate track in your edit.

Honest comparison

FreeTextoSpeech vs PlayHT and Murf dialogue features

PlayHT and Murf both ship native multi-speaker modes, paste once, render one mix. The trade is monthly character caps, signup, and paid commercial license. Here is the honest read.

Multi-voice dialogue support

FreeTextoSpeech

Write one line per turn, Host and Guest already have voices, and export one joined conversation.

PlayHT / Murf free tiers

Native multi-speaker dialogue feature, but locked behind paid tiers and capped voice roster on free.

Voice variety per scene

FreeTextoSpeech

54 voices across 9 languages, enough for an entire ensemble cast.

PlayHT / Murf free tiers

Smaller free voice pool with most distinct character voices paywalled.

Cost for a 5-minute dialogue

FreeTextoSpeech

Free. No character cap on number of generations.

PlayHT / Murf free tiers

Counts against monthly character cap on free tier; longer scenes push you to a paid plan.

Commercial license on free tier

FreeTextoSpeech

Full commercial use allowed, no attribution.

PlayHT / Murf free tiers

Commercial use restricted on free tier, paid plan required for monetized podcasts and indie games.

Output format

FreeTextoSpeech

24 kHz WAV per voice, drop straight onto separate tracks in any DAW.

PlayHT / Murf free tiers

Compressed MP3 on free tier, often single-mix only.

Signup

FreeTextoSpeech

None. Open the page, paste, generate.

PlayHT / Murf free tiers

Email signup required, credit card on file for commercial features.

One-click multi-voice generation

FreeTextoSpeech

Native multi-speaker generation with loaded speakers, joined export, and optional Advanced controls.

PlayHT / Murf free tiers

Native multi-speaker dialogue mode on paid tiers (one paste, one render).

Generate runs each turn in order so the shared voice engine is not overloaded.

AI dialogue generator workflow showing paste text, pick a voice, generate, and download

In context

A step-by-step visual guide for AI dialogue generator.

FAQ

Ai Dialogue Generator FAQs

01

Can I generate two voices in a single request?

Yes. Write the conversation, give each speaker a different voice, and click Generate. The studio runs the turns in order and creates one joined audio file.
02

How do I keep voices from sounding similar?

Pick voices that contrast on at least two axes: gender, accent, or pitch. Sarah + Adam (US female warm + US male authoritative) reads cleanly. Bella + Daniel (US conversational + UK formal) reads even cleaner because the accent shift is unmistakable. Avoid pairing two same-accent same-gender voices, listeners lose the thread.
03

How do I fake an interruption between characters?

Use a short pause between turns in Advanced options. For a true overlap, download the joined file and place the incoming line a little early in your editor.
04

How do I handle laughter, sighs, or "uh-huh" reactions?

The model has no non-verbal sound tags. Write phonetic reactions such as “hah hah,” “mmhm,” or “ugh” as their own short speaker turns, then retry that turn if another spelling works better.
05

Can I publish AI-generated dialogue commercially?

Yes. The 24 kHz WAVs you download from FreeTextoSpeech are licensed for full commercial use, podcasts, monetized YouTube videos, paid audiobooks, indie games, ad reads. No attribution required and no royalty share. The license covers two-voice dialogue scenes the same way it covers single-voice narration.
06

How long can a dialogue scene be?

Each generation is capped at 5,000 characters, but you can run as many generations as you need. A typical two-host 30-minute podcast scene works out to roughly 4,500 words of dialogue split across both characters, about 12 generations, six per voice. Splice them in one project file and you have the full episode.
07

Does the dialogue sound natural across the cut between turns?

Yes, because the Kokoro model is deterministic per voice. Character A always sounds like Character A across every clip. The seam is in the silence between turns, not in the voice itself, leave 200–400 ms of room tone or true silence between speakers and the conversation breathes the way human dialogue does.
08

Can I do three or more characters?

Yes. Add a speaker in the studio, give them a distinct voice, and keep writing one line per turn. Three voices is fine; four starts to strain listener tracking unless one is clearly the narrator.

Still wondering? Get in touch →

Try it now

Build your first two-voice scene.

Generate it free, in under 10 minutes, with full commercial rights.