Pocket TTS — Telugu (తెలుగు) · single voice syspin_female

86 MB. A complete Telugu text-to-speech model, 4-bit quantised, running

real-time on a CPU with no GPU anywhere in the loop.

Hear it first — sample_int4_g32.wav is generated by the 4-bit weights in this repo, not by the float model.

~110M parameters, 24 kHz, one voice, MIT.

Built by Prasad Vittaldev. Follow me on LinkedIn — I publish every new voice and language there as it launches.

Why this exists

Kyutai's pocket-tts is a genuinely remarkable piece of work: a text-to-speech model small enough and fast enough to run in real time on a CPU, released under MIT, with a training recipe that anyone can reproduce. That combination is rare — most speech models of this quality assume a GPU at inference and a restrictive licence.

What it shipped with, though, were European languages. India has none.

This repository is that gap closed for Telugu. Kyutai's architecture, training recipe and codec are used as published; what is added here is the part that did not exist — Telugu data preparation, forced alignment in native Telugu script, a Telugu tokenizer, a full teacher finetune, and a depth distillation down to a model this size. The result keeps everything that makes pocket-tts a great deal — CPU real-time, MIT, on-device, no API key, no per-character billing, nothing leaving the machine — and makes it speak Telugu.

Model

Architecture pocket-tts, 6-layer flow-matching backbone + Mimi codec
Parameters ~110M (including the Mimi codec)
Sample rate 24 kHz
Language Telugu (తెలుగు)
Voice syspin_female — fixed, see below
Speed Measured 5.5x real-time on CPU alone (no GPU) at int8, on desktop x86. Generating a minute of speech takes about 11 seconds.

The voice is fixed

This model was distilled on a single speaker, and voice cloning did not survive that. It ships a voice_prompt.wav because the API requires one, but supplying a different prompt will not change the voice — measured across four prompts spanning 137-231 Hz, this model's output stayed inside a 9 Hz band:

prompt clip 24-layer teacher output this model's output
137 Hz 136 Hz 199 Hz
205 Hz 204 Hz 207 Hz
231 Hz 185 Hz 209 Hz
210 Hz 213 Hz 208 Hz

The teacher tracks its prompt; this student does not. That is the expected consequence of single-speaker distillation and it is why each voice is a separate repository rather than one model with selectable voices. If you want a different voice, use a different repo.

Files: start with the 4-bit one

file precision download notes
model_int4_g32.safetensors int4 + fp16 scales 86 MB smallest. Unpack with the script below.
model_bf16.safetensors bfloat16 209 MB loads directly
model.safetensors float32 418 MB loads directly, no reason to prefer it

The 4-bit file stores every backbone weight as two 4-bit codes per byte plus one fp16 scale per group of 32, and keeps the Mimi audio codec in fp16 — quantising a neural codec causes artefacts that have nothing to do with the language model. Reconstruction is (code - 8) * scale.

What 4-bit buys and what it does not. It is a download and storage win: 86 MB instead of 418 MB, a 4.8x reduction. It is not a memory win — unpacking restores fp16 tensors, so RAM at inference matches the bf16 build. A native int4 kernel would change that, and none exists for pocket-tts today.

For speed, add --quantize to any of the three: dynamic int8 on attention and feed-forward at load time, measured at 5.5x realtime against 3.5x realtime for float.

How far down can quantization go?

Every level below was generated from this model and listened to. Weight error is the mean relative error introduced in the 85.3 M backbone weights; duration is on one fixed 885-character Telugu passage where the reference is 74.3 s — a short clip means the model dropped text, which is an objective failure that needs no listening.

precision weight error backbone size duration verdict
int8 0.4% 81 MB 74.3 s shipped default, indistinguishable
int4 group 32 10.8% 46 MB 74.0 s usable — no text lost
int4 group 128 14.0% 42 MB 72.0 s usable, slightly lossy
int3 group 32 23.9% 36 MB 70.0 s degrading — 5% of text lost
int3 group 128 30.6% 32 MB 64.4 s degrading — 13% lost
int2 group 32 61.9% 25 MB 42.0 s broken — 43% of the passage gone

A sample_*.wav for each row is in this repo, all the same passage, so the degradation is audible rather than merely tabulated.

int4 is the lowest usable precision and the audio holds up. It is not the shipped default for a practical reason: pocket-tts implements int8 only, and torchao's int4 CPU kernels need mslk, so there is no int4 runtime to point you at. quantize_int4.py in this repo reproduces the int4 weights exactly if you have a runtime that can use them. int3 and below are documented so nobody has to repeat the experiment — they lose words.

Note the backbone is not the whole model: the Mimi codec adds ~20 M parameters and is deliberately left in float, because quantizing a neural audio codec introduces artefacts unrelated to the language model.

Usage

pip install pocket-tts

The 86 MB route

Download the 4-bit file, unpack it once, then use it like any other checkpoint:

huggingface-cli download prasadvittaldev/pocket-tts-telugu-female-syspin model_int4_g32.safetensors quantize_int4.py --local-dir .
python quantize_int4.py --unpack --src model_int4_g32.safetensors --dst model.safetensors

That writes a normal safetensors file the stock loader reads. Point a config at it, or just use the hosted config and the full-precision weights below.

Straight from the Hub

python -m pocket_tts generate   --config hf://prasadvittaldev/pocket-tts-telugu-female-syspin/config.yaml   --voice  hf://prasadvittaldev/pocket-tts-telugu-female-syspin/voice_prompt.wav   --text   "గోదావరి నది తీరాన ఉన్న ఆ చిన్న గ్రామంలో సూర్యాస్తమయం చాలా అందంగా ఉంటుంది."   --output-path out.wav   --device cpu --quantize

--quantize is the int8 path: fastest at runtime, 5.5x realtime.

from pocket_tts import TTSModel
import numpy as np, soundfile as sf

model = TTSModel.load_model(config="hf://prasadvittaldev/pocket-tts-telugu-female-syspin/config.yaml", quantize=True)
model.to("cpu")
state = model.get_state_for_audio_prompt("voice_prompt.wav")
chunks = list(model.generate_audio_stream(
    model_state=state, text_to_generate="గోదావరి నది తీరాన ఉన్న ఆ చిన్న గ్రామంలో సూర్యాస్తమయం చాలా అందంగా ఉంటుంది."))
sf.write("out.wav", np.concatenate([np.asarray(c).reshape(-1) for c in chunks]), 24000)

Run it in a browser or on a phone (ONNX)

An ONNX bundle is in onnx/. It was exported with KevinAHM/pocket-tts-onnx-export and verified against the PyTorch model tensor-for-tensor, then run end to end through ONNX Runtime and listened to before publishing.

graph precision size
flow_lm_main_int4.onnx (+ .data) int4 (MatMulNBits, block 128) 38.9 MB
mimi_decoder_int8.onnx int8 21.6 MB
flow_lm_flow_int8.onnx int8 9.5 MB
mimi_encoder.onnx float32 37.9 MB
text_conditioner.onnx float32 15.6 MB

Measured on CPU with 4 threads, same sentence for each:

precision bundle speed
int4 86 MB* 3.98x real-time
int8 139 MB 4.08x real-time
float32 419 MB 3.06x real-time

* 86 MB is the practical browser bundle: this model has a fixed voice, so the 37.9 MB mimi_encoder — which exists only to encode a new voice prompt — can be dropped once the prompt state is precomputed. Keep it if you want to experiment with prompts, though note the voice will not change (see above).

int4 is 4.9x smaller than float32 and ~30% faster. It is not faster than int8, because MatMulNBits dequantises to compute: 4-bit buys download size, and in a browser that is usually the constraint that matters.

Two graphs stay in float32 on purpose. flow_lm_flow is 26 Gemm nodes and zero MatMul, and MatMulNBits only rewrites MatMul; text_conditioner is an embedding table, which weight-only quantisation skips. Neither is worth graph surgery for the few MB involved.

Runtimes that can load this: ONNX Runtime Web for browsers, and sherpa-onnx for Android/iOS/Raspberry Pi with Kotlin, Swift, Java, Dart and C# bindings.

One important note on long text

Generate one sentence at a time and join the results. pocket-tts splits long input internally, but it locates sentence ends using tokens from its own tokenizer, and an Indic tokenizer does not produce the pieces it looks for — for Telugu the boundary search finds nothing and the whole passage is pushed through as a single chunk, which silently drops words. Splitting on . ? ! before you call the model avoids this completely.

Audio samples

How it was made

  1. Teacher. Kyutai's released 24-layer English model was finetuned on 124 hours of Telugu speech (15,000 steps). The text embedding is reset and relearned against a Telugu sentencepiece tokenizer, so nothing English-language survives in the text side; the English weights serve only as an acoustic starting point.
  2. Alignment. Word-level forced alignment in native Telugu script. The obvious choice, Meta's MMS multilingual aligner, is unusable here — it romanises, and its vocabulary contains no Telugu characters at all.
  3. Distillation. The 24-layer teacher was distilled down to this 6-layer student (40,000 steps), trained only on this single speaker's 53.0 hours / 20,599 clips. Depth is where the size is: the backbone is ~96% of the parameters.

Credits

This model would not exist without:

Licence

MIT, matching upstream pocket-tts. The training corpora carry their own licences — check the dataset pages above before commercial use.

More voices and languages

More Telugu voices and more Indian languages are on the way. Follow Prasad Vittaldev on LinkedIn for each new model launch.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including prasadvittaldev/pocket-tts-telugu-female-syspin