Introduction
Text-to-Speech (TTS) technology has become a core part of modern applications — from accessibility tools and language learning platforms to podcast generators and voice assistants. Thanks to Python's gTTS (Google Text-to-Speech) library, you can convert any text into natural-sounding audio in over 30 languages with just a few lines of code.
In this tutorial you will build a polished TTS web app using gTTS and Streamlit. The finished app will let users type or paste any text, choose a language and speech speed, preview the audio directly in the browser, and download the generated MP3 file.
This project pairs naturally with our Translator App — you can translate text first, then feed the translation into this TTS app to hear it spoken aloud. It is also a core module of the larger SANGAM AI toolkit.
> 🎬 Watch the Full Video Tutorial:
> Watch the walkthrough on YouTube: Building a Simple Text-to-Speech Converter with Python and Streamlit
What is gTTS?
gTTS (Google Text-to-Speech) is a Python library that interfaces with Google Translate's TTS API. It sends your text to Google's servers and returns an MP3 audio stream, which you can then save or play back.
gTTS vs Other TTS Libraries
| Library | Free | Offline | Languages | Quality | Ease of Use |
|---|---|---|---|---|---|
| gTTS | ✅ Yes | ❌ No | 30+ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
pyttsx3 | ✅ Yes | ✅ Yes | System voices | ⭐⭐⭐ | ⭐⭐⭐⭐ |
| AWS Polly | ❌ Paid | ❌ No | 60+ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ |
| ElevenLabs | ❌ Paid | ❌ No | 28 | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ |
| Azure TTS | ❌ Paid | ❌ No | 100+ | ⭐⭐⭐⭐⭐ | ⭐⭐ |
| Coqui TTS | ✅ Yes | ✅ Yes | 20+ | ⭐⭐⭐⭐ | ⭐⭐ |
gTTS is the perfect choice for beginners and small projects: it is completely free, requires zero API key setup, and produces Google's polished TTS voices.
Prerequisites
Make sure you have the following installed before starting:
pip install streamlit gTTS| Package | Purpose |
|---|---|
streamlit | Web app framework |
gTTS | Google Text-to-Speech engine |
io | Built-in Python module for in-memory byte streams |
Supported Languages
gTTS supports over 30 languages and regional variants. Here is a selection of the most commonly used ones:
| Language | Code | Regional Variants Available |
|---|---|---|
| English | en | en-us, en-gb, en-au, en-in |
| Hindi | hi | — |
| Spanish | es | es-es, es-us, es-mx |
| French | fr | fr-fr, fr-ca |
| German | de | — |
| Portuguese | pt | pt-br, pt-pt |
| Japanese | ja | — |
| Korean | ko | — |
| Arabic | ar | — |
| Chinese | zh-CN | zh-TW |
| Italian | it | — |
| Russian | ru | — |
| Dutch | nl | — |
| Polish | pl | — |
For the full list, visit the gTTS documentation.
Project Structure
tts_app/
├── tts_app.py ← Main Streamlit application
└── requirements.txt ← pip dependenciesCreate requirements.txt:
streamlit
gTTSStep 1: Create the Basic App
Create tts_app.py and set up the page:
import streamlit as st
from gtts import gTTS, lang
import io
st.set_page_config(
page_title="Text to Speech",
page_icon="🔊",
layout="centered"
)
st.title("🔊 Text-to-Speech Converter")
st.write(
"Convert any text into natural-sounding speech. "
"Choose your language, adjust the speed, then preview or download the audio."
)Step 2: Build the Language Selector
gTTS provides a helper function lang.tts_langs() that returns a dictionary of all supported language codes and their names. We can use this to build a dynamic dropdown:
# Get all supported languages from gTTS
all_langs = lang.tts_langs() # Returns {"en": "English", "hi": "Hindi", ...}
# Flip to {name: code} for display
lang_name_to_code = {v: k for k, v in all_langs.items()}
lang_names = sorted(lang_name_to_code.keys())
col1, col2 = st.columns(2)
with col1:
selected_lang_name = st.selectbox(
"🌍 Language",
lang_names,
index=lang_names.index("English"),
)
lang_code = lang_name_to_code[selected_lang_name]
with col2:
slow_mode = st.radio(
"🐢 Speech Speed",
options=["Normal", "Slow"],
horizontal=True,
)
is_slow = slow_mode == "Slow"Step 3: Text Input Area
Add a large text area where users can type or paste any content:
st.divider()
text_input = st.text_area(
"📝 Enter your text:",
height=220,
placeholder=(
"Type or paste your text here...\n\n"
"Examples:\n"
"• A paragraph you want to hear read aloud\n"
"• A script for a video or podcast\n"
"• A translated phrase you want to pronounce correctly"
),
max_chars=5000,
)
char_count = len(text_input)
st.caption(f"{char_count} / 5,000 characters used")The 5,000 character limit prevents accidental very large requests to Google's TTS API.
Step 4: Generate and Play the Audio
This is the core of the app — converting text to speech and playing it in the browser:
if st.button("🎙️ Convert to Speech", type="primary", use_container_width=True):
if not text_input.strip():
st.warning("Please enter some text before converting.")
else:
with st.spinner("Generating audio..."):
try:
# Generate TTS audio
tts = gTTS(text=text_input, lang=lang_code, slow=is_slow)
# Write to an in-memory byte buffer (no temp files needed)
audio_buffer = io.BytesIO()
tts.write_to_fp(audio_buffer)
audio_buffer.seek(0)
st.success(f"✅ Audio generated in {selected_lang_name}!")
# Display the inline player
st.audio(audio_buffer, format="audio/mp3")
# Download button
audio_buffer.seek(0)
st.download_button(
label="⬇️ Download MP3",
data=audio_buffer,
file_name="speech.mp3",
mime="audio/mpeg",
use_container_width=True,
)
except Exception as e:
st.error(f"Something went wrong: {e}")
st.info(
"Tip: Check your internet connection — "
"gTTS requires access to Google's servers."
)Using io.BytesIO() instead of saving to a file on disk is the correct pattern for Streamlit: it keeps everything in memory and works perfectly in cloud environments where writing to disk is restricted.
Step 5: Add Multi-Language Demo
Let us add a helpful demo section that converts a greeting phrase into multiple languages at once — great for showcasing the app's capabilities:
st.divider()
with st.expander("🌐 Multi-Language Demo — Hear 'Hello' in 6 languages"):
demo_phrases = {
"English": ("en", "Hello! Welcome to the Text-to-Speech app."),
"Hindi": ("hi", "नमस्ते! टेक्स्ट-टू-स्पीच ऐप में आपका स्वागत है।"),
"Spanish": ("es", "¡Hola! Bienvenido a la aplicación de texto a voz."),
"French": ("fr", "Bonjour! Bienvenue dans l'application de synthèse vocale."),
"German": ("de", "Hallo! Willkommen bei der Text-to-Speech-App."),
"Japanese":("ja", "こんにちは!テキスト読み上げアプリへようこそ。"),
}
for language, (code, phrase) in demo_phrases.items():
col_a, col_b = st.columns([1, 3])
with col_a:
st.write(f"**{language}**")
with col_b:
if st.button(f"▶️ Play", key=f"demo_{code}"):
tts = gTTS(text=phrase, lang=code, slow=False)
buf = io.BytesIO()
tts.write_to_fp(buf)
buf.seek(0)
st.audio(buf, format="audio/mp3")Complete tts_app.py
Here is the complete, final version of the app:
import streamlit as st
from gtts import gTTS, lang
import io
st.set_page_config(page_title="Text to Speech", page_icon="🔊", layout="centered")
st.title("🔊 Text-to-Speech Converter")
st.write("Convert text to natural speech in 30+ languages. Preview and download the audio.")
all_langs = lang.tts_langs()
lang_name_to_code = {v: k for k, v in all_langs.items()}
lang_names = sorted(lang_name_to_code.keys())
col1, col2 = st.columns(2)
with col1:
selected_lang_name = st.selectbox("🌍 Language", lang_names, index=lang_names.index("English"))
lang_code = lang_name_to_code[selected_lang_name]
with col2:
is_slow = st.radio("🐢 Speed", ["Normal", "Slow"], horizontal=True) == "Slow"
st.divider()
text_input = st.text_area("📝 Enter your text:", height=220, max_chars=5000,
placeholder="Type or paste your text here...")
st.caption(f"{len(text_input)} / 5,000 characters")
if st.button("🎙️ Convert to Speech", type="primary", use_container_width=True):
if not text_input.strip():
st.warning("Please enter some text.")
else:
with st.spinner("Generating audio..."):
try:
tts = gTTS(text=text_input, lang=lang_code, slow=is_slow)
buf = io.BytesIO()
tts.write_to_fp(buf)
buf.seek(0)
st.success(f"✅ Ready in {selected_lang_name}!")
st.audio(buf, format="audio/mp3")
buf.seek(0)
st.download_button("⬇️ Download MP3", buf, "speech.mp3", "audio/mpeg",
use_container_width=True)
except Exception as e:
st.error(f"Error: {e}")
st.divider()
with st.expander("🌐 Multi-Language Demo"):
demos = {"English": ("en", "Hello! Welcome."), "Hindi": ("hi", "नमस्ते!"),
"Spanish": ("es", "¡Hola!"), "French": ("fr", "Bonjour!")}
for name, (code, phrase) in demos.items():
c1, c2 = st.columns([1, 3])
c1.write(f"**{name}**")
with c2:
if st.button("▶️ Play", key=f"d_{code}"):
t = gTTS(phrase, lang=code)
b = io.BytesIO()
t.write_to_fp(b)
b.seek(0)
st.audio(b, format="audio/mp3")Run the App
streamlit run tts_app.pyHow gTTS Works Internally
Under the hood, gTTS:
https://translate.google.com/translate_tts.BytesIO buffer.This is why a working internet connection is required — gTTS is not an offline library.
Common Use Cases
| Use Case | How to Adapt the App |
|---|---|
| Language Learning | Add phonetic transcription alongside audio |
| Accessibility Tool | Increase font size, add keyboard shortcuts |
| Podcast Script Reader | Upload a .txt script file and convert in bulk |
| Pronunciation Checker | Pair with the Translator App |
| E-learning Platform | Embed audio players next to quiz questions |
| Audio Digest | Summarise articles and convert to MP3 |
Troubleshooting
| Problem | Likely Cause | Fix |
|---|---|---|
No module named 'gtts' | Library not installed | Run pip install gTTS |
gTTSError: 429 | Rate limited by Google | Wait 30 seconds and retry |
| Empty or silent audio | Whitespace-only input | Add text_input.strip() check |
AssertionError: No text | Text is empty string | Validate before calling gTTS() |
| Truncated audio | Text too long | Split into paragraphs and convert separately |
| Language not found | Wrong language code | Use lang.tts_langs() to get valid codes |
Extending the App
Batch File Conversion
You can extend the app to convert an uploaded .txt file into speech:
uploaded = st.file_uploader("Upload a .txt file", type=["txt"])
if uploaded:
file_text = uploaded.read().decode("utf-8")
st.text_area("File Preview", file_text[:500] + "...", disabled=True)
if st.button("Convert File to Speech"):
tts = gTTS(text=file_text[:5000], lang=lang_code, slow=is_slow)
buf = io.BytesIO()
tts.write_to_fp(buf)
buf.seek(0)
st.audio(buf, format="audio/mp3")Speaking Numbers and Abbreviations
gTTS reads numbers as digits by default. To have it read 2025 as "twenty twenty-five" instead of "two thousand and twenty-five", convert it first:
import inflect
p = inflect.engine()
number_as_words = p.number_to_words(2025) # → "two thousand and twenty-five"Install inflect with pip install inflect.
Related Projects on This Blog
Conclusion
You have built a production-ready Text-to-Speech web app using Python, gTTS, and Streamlit. The app handles text input, language selection, speed control, in-browser audio playback, and MP3 downloads — all in under 50 lines of core logic.
TTS is a remarkably versatile tool. Whether you are building a language-learning assistant, an accessibility aid, or a voice-over generator for your content, the techniques in this tutorial give you a solid, extensible foundation.
Resources: