How it works
- 01
Paste your text
Any length. Articles, scripts, study notes, the whole thing rather than the first 500 characters.
- 02
Pick a voice and a pace
Local voices download once and then work offline. Your system voices need no download at all.
- 03
Listen, then download
MP3 for sharing and podcasts. WAV when the audio is going into an editor.
Local voices and system voices
Two engines sit behind this tool. Local voices are compact models that download once, between 20 and 73 MB each, and then run on your own processor. That is why they keep working with no connection, and why there is no per-character bill to pass on.
System voices come from your operating system through the Web Speech API. Nothing to download, and they start straight away. The catch is that they read through your speakers and hand back no file, so there is nothing to save. The quality also depends on what your machine happens to have installed.
Neither route sends your text anywhere. That matters more than it sounds. Most free text to speech sites post your text to an API. That makes them a poor choice for anything confidential, and it is why they cap you at a few hundred characters.
- Length
- No character limit. A full article or chapter goes through in one pass.
- Voices
- 42 local voices across 18 languages and regions, plus every voice your system provides.
- Output
- WAV or MP3. No watermark, and no attribution required.
- Offline
- Once a voice is cached, reading works with no connection at all.
Getting natural-sounding output
Punctuation does most of the work. Commas and full stops become pauses. Text written for the eye often runs long clauses together, and it reads far better once you break those into separate sentences.
Numbers, dates and money are turned into words by default, which is usually right and occasionally not. A version like 4.55 should be read as digits, and the first toggle controls that.
Acronyms are the other common snag. Familiar ones are said as words, unfamiliar ones are spelled out letter by letter, and the tool cannot always tell you which it did. Turn on Spell out acronyms and full stops go between the letters, which forces the spelled-out reading.