Instructions to use prasadvittaldev/pocket-tts-telugu-female-syspin with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use prasadvittaldev/pocket-tts-telugu-female-syspin with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("prasadvittaldev/pocket-tts-telugu-female-syspin") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Pocket TTS — Telugu (తెలుగు) · single voice syspin_female
86 MB. A complete Telugu text-to-speech model, 4-bit quantised, running
real-time on a CPU with no GPU anywhere in the loop.
Hear it first — sample_int4_g32.wav is generated by the
4-bit weights in this repo, not by the float model.
~110M parameters, 24 kHz, one voice, MIT.
Built by Prasad Vittaldev. Follow me on LinkedIn — I publish every new voice and language there as it launches.
Why this exists
Kyutai's pocket-tts is a genuinely remarkable piece of work: a text-to-speech model small enough and fast enough to run in real time on a CPU, released under MIT, with a training recipe that anyone can reproduce. That combination is rare — most speech models of this quality assume a GPU at inference and a restrictive licence.
What it shipped with, though, were European languages. India has none.
This repository is that gap closed for Telugu. Kyutai's architecture, training recipe and codec are used as published; what is added here is the part that did not exist — Telugu data preparation, forced alignment in native Telugu script, a Telugu tokenizer, a full teacher finetune, and a depth distillation down to a model this size. The result keeps everything that makes pocket-tts a great deal — CPU real-time, MIT, on-device, no API key, no per-character billing, nothing leaving the machine — and makes it speak Telugu.
Model
| Architecture | pocket-tts, 6-layer flow-matching backbone + Mimi codec |
| Parameters | ~110M (including the Mimi codec) |
| Sample rate | 24 kHz |
| Language | Telugu (తెలుగు) |
| Voice | syspin_female — fixed, see below |
| Speed | Measured 5.5x real-time on CPU alone (no GPU) at int8, on desktop x86. Generating a minute of speech takes about 11 seconds. |
The voice is fixed
This model was distilled on a single speaker, and voice cloning did not
survive that. It ships a voice_prompt.wav because the API requires one, but
supplying a different prompt will not change the voice — measured across
four prompts spanning 137-231 Hz, this model's output stayed inside a 9 Hz band:
| prompt clip | 24-layer teacher output | this model's output |
|---|---|---|
| 137 Hz | 136 Hz | 199 Hz |
| 205 Hz | 204 Hz | 207 Hz |
| 231 Hz | 185 Hz | 209 Hz |
| 210 Hz | 213 Hz | 208 Hz |
The teacher tracks its prompt; this student does not. That is the expected consequence of single-speaker distillation and it is why each voice is a separate repository rather than one model with selectable voices. If you want a different voice, use a different repo.
Files: start with the 4-bit one
| file | precision | download | notes |
|---|---|---|---|
model_int4_g32.safetensors |
int4 + fp16 scales | 86 MB | smallest. Unpack with the script below. |
model_bf16.safetensors |
bfloat16 | 209 MB | loads directly |
model.safetensors |
float32 | 418 MB | loads directly, no reason to prefer it |
The 4-bit file stores every backbone weight as two 4-bit codes per byte plus one
fp16 scale per group of 32, and keeps the Mimi audio codec in fp16 — quantising
a neural codec causes artefacts that have nothing to do with the language model.
Reconstruction is (code - 8) * scale.
What 4-bit buys and what it does not. It is a download and storage win: 86 MB instead of 418 MB, a 4.8x reduction. It is not a memory win — unpacking restores fp16 tensors, so RAM at inference matches the bf16 build. A native int4 kernel would change that, and none exists for pocket-tts today.
For speed, add --quantize to any of the three: dynamic int8 on attention
and feed-forward at load time, measured at 5.5x realtime against 3.5x realtime for float.
How far down can quantization go?
Every level below was generated from this model and listened to. Weight error is the mean relative error introduced in the 85.3 M backbone weights; duration is on one fixed 885-character Telugu passage where the reference is 74.3 s — a short clip means the model dropped text, which is an objective failure that needs no listening.
| precision | weight error | backbone size | duration | verdict |
|---|---|---|---|---|
| int8 | 0.4% | 81 MB | 74.3 s | shipped default, indistinguishable |
| int4 group 32 | 10.8% | 46 MB | 74.0 s | usable — no text lost |
| int4 group 128 | 14.0% | 42 MB | 72.0 s | usable, slightly lossy |
| int3 group 32 | 23.9% | 36 MB | 70.0 s | degrading — 5% of text lost |
| int3 group 128 | 30.6% | 32 MB | 64.4 s | degrading — 13% lost |
| int2 group 32 | 61.9% | 25 MB | 42.0 s | broken — 43% of the passage gone |
A sample_*.wav for each row is in this repo, all the same passage, so the
degradation is audible rather than merely tabulated.
int4 is the lowest usable precision and the audio holds up. It is not the
shipped default for a practical reason: pocket-tts implements int8 only, and
torchao's int4 CPU kernels need mslk, so there is no int4 runtime to point you
at. quantize_int4.py in this repo reproduces the int4 weights exactly if you
have a runtime that can use them. int3 and below are documented so nobody has to
repeat the experiment — they lose words.
Note the backbone is not the whole model: the Mimi codec adds ~20 M parameters and is deliberately left in float, because quantizing a neural audio codec introduces artefacts unrelated to the language model.
Usage
pip install pocket-tts
The 86 MB route
Download the 4-bit file, unpack it once, then use it like any other checkpoint:
huggingface-cli download prasadvittaldev/pocket-tts-telugu-female-syspin model_int4_g32.safetensors quantize_int4.py --local-dir .
python quantize_int4.py --unpack --src model_int4_g32.safetensors --dst model.safetensors
That writes a normal safetensors file the stock loader reads. Point a config at it, or just use the hosted config and the full-precision weights below.
Straight from the Hub
python -m pocket_tts generate --config hf://prasadvittaldev/pocket-tts-telugu-female-syspin/config.yaml --voice hf://prasadvittaldev/pocket-tts-telugu-female-syspin/voice_prompt.wav --text "గోదావరి నది తీరాన ఉన్న ఆ చిన్న గ్రామంలో సూర్యాస్తమయం చాలా అందంగా ఉంటుంది." --output-path out.wav --device cpu --quantize
--quantize is the int8 path: fastest at runtime, 5.5x realtime.
from pocket_tts import TTSModel
import numpy as np, soundfile as sf
model = TTSModel.load_model(config="hf://prasadvittaldev/pocket-tts-telugu-female-syspin/config.yaml", quantize=True)
model.to("cpu")
state = model.get_state_for_audio_prompt("voice_prompt.wav")
chunks = list(model.generate_audio_stream(
model_state=state, text_to_generate="గోదావరి నది తీరాన ఉన్న ఆ చిన్న గ్రామంలో సూర్యాస్తమయం చాలా అందంగా ఉంటుంది."))
sf.write("out.wav", np.concatenate([np.asarray(c).reshape(-1) for c in chunks]), 24000)
Run it in a browser or on a phone (ONNX)
An ONNX bundle is in onnx/. It was exported with
KevinAHM/pocket-tts-onnx-export
and verified against the PyTorch model tensor-for-tensor, then run end to end
through ONNX Runtime and listened to before publishing.
| graph | precision | size |
|---|---|---|
flow_lm_main_int4.onnx (+ .data) |
int4 (MatMulNBits, block 128) |
38.9 MB |
mimi_decoder_int8.onnx |
int8 | 21.6 MB |
flow_lm_flow_int8.onnx |
int8 | 9.5 MB |
mimi_encoder.onnx |
float32 | 37.9 MB |
text_conditioner.onnx |
float32 | 15.6 MB |
Measured on CPU with 4 threads, same sentence for each:
| precision | bundle | speed |
|---|---|---|
| int4 | 86 MB* | 3.98x real-time |
| int8 | 139 MB | 4.08x real-time |
| float32 | 419 MB | 3.06x real-time |
* 86 MB is the practical browser bundle: this model has a fixed voice, so
the 37.9 MB mimi_encoder — which exists only to encode a new voice prompt —
can be dropped once the prompt state is precomputed. Keep it if you want to
experiment with prompts, though note the voice will not change (see above).
int4 is 4.9x smaller than float32 and ~30% faster. It is not faster than int8,
because MatMulNBits dequantises to compute: 4-bit buys download size, and
in a browser that is usually the constraint that matters.
Two graphs stay in float32 on purpose. flow_lm_flow is 26 Gemm nodes and
zero MatMul, and MatMulNBits only rewrites MatMul; text_conditioner is
an embedding table, which weight-only quantisation skips. Neither is worth
graph surgery for the few MB involved.
Runtimes that can load this: ONNX Runtime Web for browsers, and sherpa-onnx for Android/iOS/Raspberry Pi with Kotlin, Swift, Java, Dart and C# bindings.
One important note on long text
Generate one sentence at a time and join the results. pocket-tts splits long
input internally, but it locates sentence ends using tokens from its own
tokenizer, and an Indic tokenizer does not produce the pieces it looks for — for
Telugu the boundary search finds nothing and the whole passage is pushed
through as a single chunk, which silently drops words. Splitting on . ? !
before you call the model
avoids this completely.
Audio samples
sample_bf16.wavsample_int2_g32.wavsample_int3_g128.wavsample_int3_g32.wavsample_int4_g128.wavsample_int4_g32.wavsample_int8.wav
How it was made
- Teacher. Kyutai's released 24-layer English model was finetuned on 124 hours of Telugu speech (15,000 steps). The text embedding is reset and relearned against a Telugu sentencepiece tokenizer, so nothing English-language survives in the text side; the English weights serve only as an acoustic starting point.
- Alignment. Word-level forced alignment in native Telugu script. The obvious choice, Meta's MMS multilingual aligner, is unusable here — it romanises, and its vocabulary contains no Telugu characters at all.
- Distillation. The 24-layer teacher was distilled down to this 6-layer student (40,000 steps), trained only on this single speaker's 53.0 hours / 20,599 clips. Depth is where the size is: the backbone is ~96% of the parameters.
Credits
This model would not exist without:
- kyutai-labs/pocket-tts — the architecture, the Mimi codec, the training and distillation recipes, released under MIT. All the hard modelling work is theirs.
- Speech corpora:
SPRINGLab/IndicTTS_Telugu— IndicTTS (SPRING Lab, IIT Madras)arpit-tiwari/syspin-telugu-tts— SYSPIN (IISc Bangalore)
Licence
MIT, matching upstream pocket-tts. The training corpora carry their own licences — check the dataset pages above before commercial use.
More voices and languages
More Telugu voices and more Indian languages are on the way. Follow Prasad Vittaldev on LinkedIn for each new model launch.
- Downloads last month
- -