Emad Mostaque shows Taalas generating 14,000 tokens per second on podcast

Emad Mostaque demonstrated Taalas producing text at roughly 14,000 tokens per second, according to a tweet by Rohan Paul that captured a segment of “The Peter McCormack Show.” By contrast, ChatGPT currently operates between 50 and 150 tokens per second under typical conditions, a gap of at least two orders of magnitude.
If the figure holds under comparable conditions, it would imply a dramatic reduction in inference cost and latency for real-time applications. A throughput of 14,000 tokens per second could stream long responses with almost no perceptible delay and serve many more concurrent users on the same hardware. The source, however, does not disclose model size, context length, batch size, or hardware specifications, parameters without which the number alone conveys limited meaning.
The demonstration took place in a YouTube broadcast of a podcast, not in an independent test or a formal technical release. Rohan Paul, who posted the clip, is an observer rather than a reviewer and expressed enthusiasm about the speed. The accompanying article, technical report, or press release from Taalas provides no additional data, no quality metrics such as Perplexity, no stability measurements over time, and no detail on what exactly was measured (single stream, parallel streams, cache usage, etc.).
Full performance metrics have not been published, the running model’s parameter count and architecture are undisclosed, and it is unclear whether the speed persists with long prompts or only with short generation. Without these details, the claim cannot be directly compared to existing inference engines such as vLLM, TensorRT-LLM or solutions from Groq and Cerebras, which also report high numbers under specific conditions. It is also unknown whether Taalas relies on dedicated hardware, a compiler for existing models, or a combination of both.
Mostaque, formerly a founder of Stability AI, presents Taalas as his new venture. The history of speed announcements in inference suggests caution: podcast demos tend to select optimal settings and omit constraints that emerge in production. Until open data for reproducible testing are released, the 14,000-token-per-second figure remains a claim rather than a verified result.