Mato Voice is Mato's own voice model, built for podcast conversation. You can hear it on the Mato Voice page, where one short script is rendered through Gemini, ElevenLabs and Mato Voice at matched loudness. From there, a show can clone a voice with the speaker's permission or design an original one in a guided Voice Studio session.
Today we are introducing Mato Voice.
Every show we have produced so far went through a third-party text-to-speech provider. The voices were clean and the words were right. In review, the same note kept coming back from our own team: the two hosts sounded like they were reading in turns.
A podcast has a shape that a transcript never shows. When an answer starts. When one host agrees quietly under the other. When a thought restarts halfway through a sentence. Read a good interview on paper and it looks ordinary. Hear it and you understand why people stayed for forty minutes.
We wanted a model we could direct toward those moments, and we wanted the choice of voice to be part of the same careful process. That is Mato Voice.
What is Mato Voice?
Mato Voice is our voice model for podcast conversation. It renders two-host dialogue from a script, and it is the model behind the Mato Voice clip in the listening comparison.
Three things ship with it:
- A listening comparison on the Mato Voice page, so you can judge the sound before you read another paragraph about it.
- Permissioned cloning, for hosts and recurring guests who have agreed to have their voice cloned.
- Original voice design in Mato Voice Studio, where a short text brief becomes a voice that never belonged to anyone.
Cloning and design both run as guided sessions with our team. The reason is further down.
Listen to the same script three ways
Speech is best judged with your ears, so the comparison comes first.
We wrote one short conversational script, a two-host exchange about a thermostat, and rendered it three times: once through Gemini, once through ElevenLabs and once through Mato Voice. The script and the speaker turns are identical in every clip. The published files are loudness-matched to the same target with no timeline edits, so you can play them back to back without reaching for the volume control.

One script, three renders. The clips on the Mato Voice page share the script, the speaker turns and the loudness target.
Each clip lists the model and the rendering recipe used. The voice identities differ by design: Gemini and ElevenLabs render their own voices, and Mato Voice renders the pair Mato uses for Amara and Derek from its talent lineup. The disclosure under each clip explains where every voice came from.
Two limits. This is a listening example with no score and no winner, and we are not claiming a measured margin over either provider. The script also carries no stage directions. Every system received the same fifteen lines with speaker labels only, so any laugh, hesitation or interruption you hear is that system's own choice. Listen for what you would listen for in your own show: whether the handoffs land, and whether the two voices sound like they share a room. Check the words against the script while you are at it.
What conversation asks of a voice
Our reviewers listen for four things, and they are the things we directed Mato Voice toward.
- Timing. Do the pauses and pickups support the meaning of the line, or do they only separate one read from the next?
- Backchannel. Small acknowledgements under the other speaker. The sounds that tell you someone is listening.
- Overlap. A quick agreement or an interruption that starts before the other host finishes.
- Repair. A restart or a self-correction, used sparingly, so written dialogue sounds spoken.
None of these is good on its own. Too little interaction and the episode feels assembled. Too much and it gets tiring on a commute. A measured interview wants clean handoffs; a lively co-host segment wants more energy. That is a creative decision about your show, and it belongs in the direction we take from you rather than in a default setting.
Two ways to get a voice
Clone a voice you have permission for
A host, a recurring guest or a narrator provides clean reference recordings and documented consent for that voice to be cloned. Permission comes first. We do not clone a voice without it, and we ask for the consent documentation before any work starts.
We then check the clone against the scripts your show publishes, because a single line makes almost any voice sound plausible. A full passage exposes names, sentence lengths, tonal turns and the phrases that carry a host's point of view.
Design a voice from a brief
If there is no recording, or you want a voice that belongs to the show rather than to a person, you start from a description. Warmth, pace, register, accent, role, and the kind of energy the voice should bring into the room. Our team turns that brief into voice options and auditions them with you on your own script lines.
What Voice Studio shows today
Mato Voice Studio is the public home for voice design. The page shows a set of designed voices, each made from a short text brief, so you can read the brief and hear the pre-rendered result side by side. It also has a brief builder: choose descriptors, add the context that never fits in a chip, and copy the result into a demo request.
Two things the page does not do, on purpose. It does not synthesize audio in your browser; every example there was rendered ahead of time. And it does not send anything to us until you attach the brief to a booking.
How a guided session runs

A Voice Studio session is run with our team, on the script lines your show publishes.
- You send the brief. Descriptors, reference material, the scripts you publish, and consent documentation if we are cloning.
- We run the studio with you. We review the direction, evaluate voice options on the script lines you publish and narrow the brief until it can carry a full episode.
- The voice goes into production. After your approval, our team connects it to your Mato workflow, and episodes continue through the review and publishing steps you already use.
We are keeping this guided for now. A voice becomes part of a show's identity, and the question that matters is whether your team is happy to hear it represent the show for the next hundred episodes. That takes a conversation and a passage from your own script.
Start with your own words
Listen to the comparison. Read the Studio briefs and hear what came out of them. Then bring us the passage from your show that is hardest to get right, with the proper names in it and the sentence that always sounds wrong. We will use it to decide together whether permissioned cloning or an original voice is the better starting point, and plan a session around it.
Bring the line that gives you trouble. Skip the one written for a demo reel.
Mato Voice
Hear it before you decide.
Play the same script through three systems, then bring us a passage from your own show.
Frequently asked questions
What is Mato Voice?
Mato Voice is Mato's own voice model for podcast conversation. It renders two-host dialogue from a script and powers the Mato Voice clip in the listening comparison. Voices for it come from permissioned cloning or from original voice design in a guided Voice Studio session.
Is the comparison on the Mato Voice page a benchmark?
No. It is one script rendered through Gemini, ElevenLabs and Mato Voice at matched loudness so you can listen back to back. There is no score, no winner and no measured claim of superiority. The script has no stage directions, so any laugh or interruption you hear is each system's own choice.
Can Mato clone my host's voice?
Yes, with the speaker's documented consent. Send clean reference recordings and the consent documentation with your brief, and we check the clone against the scripts your show publishes before it goes into production.
Can I design a voice without a recording?
Yes. Describe warmth, pace, register, accent and role in a brief. Our team turns the brief into voice options and auditions them with you on your own script lines in a guided session.
Does the Voice Studio page generate audio in my browser?
No. The examples on the page were rendered ahead of time from their text briefs. The brief builder writes text you can copy into a demo request; it does not synthesize audio or send anything to Mato on its own.
How do I get started?
Listen to the comparison, build a brief in Voice Studio and book a demo with a passage from your show. We plan the guided session around that passage.




