One Who Reads Aloud
Multi-voice audiobooks generated entirely on one 12GB GPU, at zero cost.
The constraint
Producing a multi-voice audiobook means either a studio or cloud TTS billed per character, and both are out of reach for most of the people who would want one. The whole pipeline had to run on hardware someone already owns.
The hard part
Fitting LLM speaker diarization and emotional voice cloning inside 12GB of VRAM meant scheduling models in and out rather than holding them resident. Nothing leaves the machine, so there is no cloud TTS bill and no transcript sitting on someone else's server.
