Skip to content
Selected work
004/Case study

One Who Reads Aloud

Multi-voice audiobooks generated entirely on one 12GB GPU, at zero cost.

PL.004 · 2026
Shipped2026
Year
2026
Status
Shipped
Stack
Python · FastAPI · IndexTTS2 · Qwen3

The constraint

Producing a multi-voice audiobook means either a studio or cloud TTS billed per character, and both are out of reach for most of the people who would want one. The whole pipeline had to run on hardware someone already owns.

The hard part

Fitting LLM speaker diarization and emotional voice cloning inside 12GB of VRAM meant scheduling models in and out rather than holding them resident. Nothing leaves the machine, so there is no cloud TTS bill and no transcript sitting on someone else's server.