A voice-inference lab

Frontier speech, cheap enough to ship.

CelestLabs distills frontier text-to-speech until it streams on ordinary CPUs, then serves it in every region on earth. Not the most decorated instrument in the sky: the one everyone can afford to point.

01 · The log

How the lab came to be built.

Late 2025. We were prototyping a voice agent for a customer-support workflow. The model worked. The conversation worked. Then we priced a month of audio at the incumbent APIs, multiplied by a hundred concurrent users, and the number stopped us cold. Lifelike speech, the kind a real customer actually wants to hear, cost more than the entire rest of the stack put together.

That math does not only fail a scrappy startup. It fails an established product team rolling voice out across a million users, a public-service app trying to reach citizens in their own language, and a research group trying to ship anything at all. So the lab was founded on one question: what would it take to run a competitive speech model on a normal CPU, in a normal cloud region, at a price nobody flinches at?

Three months of benchmarking, distillation, and infrastructure work later, the first instrument was serving. feather: a 117-million-parameter flow-matching model, eight synthesis steps, no GPU anywhere in the path, and the first word out in under 100 ms. 0.19¢ per 1,000 characters, which is $1.90 for a million.

Then came starling, the expressive flagship: 107 languages, zero-shot voice cloning, $9.00 per million characters, and the same rule as feather. It has to run where your data already lives.

We do not chase leaderboards. Squeezing out the last decimal of "sounds human" is astronomy for its own sake; frontier models crossed that threshold already. We chart the sky so products can navigate it, and every claim on this site is measured on the hardware that serves you, not on a slide.

02 · The bet

Four observations we are willing to be wrong about.

Obs. 01

Sounds-human is solved.

Frontier speech crossed the believability threshold. The remaining gap to the most decorated systems is audible only to voiceover professionals, and it narrows every quarter. Competing on that last decimal is not where products are won.

Obs. 02

Cost and reach decide.

The variables that determine whether a voice product ships are what a character costs and where the model is allowed to run. At 0.19¢ per 1,000 characters, up to 50× cheaper than frontier voice models, whole categories of product stop being uneconomic.

Obs. 03

CPU is abundant.

GPUs are scarce, rented, and concentrated in a few regions. Ordinary 8-core CPUs exist in every cloud region, on-premise, and air-gapped. A model that streams faster than realtime on one, with the first word in under 100 ms, can go wherever your data must live.

Obs. 04

Ship checkpoints.

A lab earns trust by serving, not by teasing. feather and starling are both in production today; starling adds voice cloning across 107 languages. Each checkpoint is published with its measurements: RTF 0.18, 117M parameters, and the rest of the log.

03 · The founder

"I have spent my career making big models cheap to run, distilling them, and squeezing inference onto the humblest hardware that will hold it. CelestLabs is that instinct, pointed at voice."

Staff software engineer, most recently at xAI (eval and inference for Grok-4), and before that Coupang (air-gapped RAG and distilled 7B student models), Atlassian, Disney+ Hotstar, and Zomato, where scale meant billions of events a day. The same lesson kept repeating: the model is rarely the hard part. Serving it cheaply, everywhere, is. The next decade of software will not be typed, tapped, or stared at. It will be spoken, and the only thing between that world and this one is what it costs, and where it is allowed to run.

Aayush Gupta · Founder, CelestLabs
ex-xAI · Coupang · Atlassian · Disney+ Hotstar · Zomato
CelestLabs

Come see what the instrument measures.

ESTABLISHED · 2026
INSTRUMENTS · FEATHER + STARLING · SERVING