Private Text to Speech That Runs in Your Browser (What Leaves Your Device)
Want text to speech without sending your text to a server? Here's how in-browser AI voices work, exactly what leaves your device and what doesn't, and how to set it up for free.
Offline text to speech means the voice model runs on your own computer, so your text is turned into audio without being sent to a server. On Free Voice Reader, the local voices work this way: the model downloads to your browser once, and from then on speech is generated on your device, for free and with no account.
This page is about the privacy side: what gets downloaded, what stays put, and where the limits are. If you mainly care about getting past character caps, our guide to unlimited free text to speech covers that angle.
How can a website do text to speech without a server?
The website sends your browser a voice model, and your browser runs it. Modern browsers can run AI models locally using WebGPU (your graphics chip) or WebAssembly (your processor), so the heavy lifting happens on your machine.
Our local voices use two open models. Kokoro-82M is an 82-million-parameter model under the Apache 2.0 license, and Supertonic is a second option. Each runs in a background worker inside the page so the tool stays responsive while it generates.
What gets downloaded, and how big is it?
Only the model files come down, from Hugging Face, the first time you choose a local voice. Your browser stores them in its cache, so later visits load the model from your own disk.
The Kokoro download size depends on how your browser runs it. On most computers it uses a compressed version of the model, about 92 MB; if your browser runs it on the graphics chip (WebGPU), it uses the full-precision file, about 326 MB. Either way it's a one-time download per browser.
What leaves my device and what doesn't?
With a local voice, your text and the generated audio stay on your computer. Here's how that compares with the cloud voices on the same tool:
| Local voices (Kokoro, Supertonic) | Cloud voices | |
|---|---|---|
| Your text sent to a server | No | Yes, to generate the audio |
| Generated audio | Made on your device | Made in the cloud, sent back to you |
| Uploaded PDF, Word, EPUB or image | Text extracted in your browser | Text extracted in your browser, then sent for speech |
| One-time download | Model files (about 92 to 326 MB) | None |
| Account needed | No | No for free voices |
| Character limit | None | 5,000 per free conversion |
Two honest footnotes. First, like most websites, ours loads analytics when you open the page; that's about the visit, and the local voice path doesn't send your text anywhere. Second, you still need an internet connection to open the page itself, because the site isn't installable as an offline app yet.
Which documents can I read privately?
PDFs, Word files, EPUBs, plain text, Markdown and images all work. File parsing runs in your browser, and scanned pages are converted with Tesseract.js, an open-source OCR library, also in the browser.
Combine that with a local voice and the whole chain (open file, extract text, generate speech) happens on your computer. That's a reasonable fit for drafts, contracts, medical paperwork or anything else you'd rather not paste into a random website.
How do I set it up?
Pick the local voice option, download the model once, then paste or upload your text and press play. It takes a minute or two on a normal connection.
- Open freevoicereader.com.
- Choose the local (on-device) voice option.
- Click Download. The progress bar shows the model loading.
- Paste text or click Upload, pick a voice, and press play.
A recent desktop browser works best. Older laptops and phones can run it, just more slowly.
Which languages do the private voices support?
Kokoro covers US and UK English, Spanish, French, Hindi, Italian and Brazilian Portuguese. Supertonic covers English, Korean, Spanish, Portuguese and French.
For other languages, the cloud voices cover far more, with the privacy trade-off described above. Spanish speakers can see our Spanish text to speech guide for the full voice list.
Is local quality as good as cloud voices?
For everyday listening, Kokoro sounds natural and holds up well over long passages. The premium cloud voices (Google Neural2, WaveNet and Gemini on our subscriptions) are more expressive, and the gap shows most in dramatic reading and unusual names.
If you'd like a deeper comparison, we put Kokoro up against ElevenLabs in an earlier post.
Frequently asked questions
Is offline text to speech really free?
Yes. The local voices are free with no account and no character limit. You pay only in download time and your computer's effort.
Does it work with no internet at all?
Speech generation runs on your device once the model is cached, but you need a connection to load the web page. A fully offline app is a different product; our Mac app runs locally on macOS.
Where is the model stored, and can I delete it?
In your browser's site storage for freevoicereader.com. Clearing site data for the site in your browser settings removes it.
Can I save the audio?
MP3 download comes with Lifetime Boost or a subscription. See the text to MP3 guide.
Want to try it with something you'd rather keep private? Download a local voice and press play. Your text stays on your machine.
Transparency Notice: This article was written by AI, reviewed by humans. We fact-check all content for accuracy and ensure it provides genuine value to our readers.