How Cartesia is shipping real-time voice AI with Crusoe



About Cartesia: Inventing State Space Models
Cartesia is building the next generation of AI: ubiquitous, interactive intelligence that runs wherever you are. The company's founding team met as PhDs at the Stanford AI Lab, where they invented state space models (SSMs)—a fundamental new primitive for training large-scale foundation models. Over four years, they've scaled SSMs to state-of-the-art results across text, audio, video, images, and time-series data.
The Challenge: Training a New Model Architecture Demands Reliable Infrastructure
Building a model architecture from the ground up means relying on infrastructure that can keep pace. The Cartesia team needed three things from their cloud partner: strong price-to-performance for sustained training, a support team responsive enough to iterate alongside them, and operational reliability at the scale their training runs demanded.
The Solution: NVIDIA Hopper Clusters and Custom SLURM on Crusoe Cloud
Cartesia trains on NVIDIA’s Hopper GPU clusters on Crusoe Cloud. Cartesia built and optimized its own SSM inference stack to serve with low latency and high throughput at scale without impacting quality. Cartesia chose Crusoe for their 1) optimal price/performance offering; 2) responsive support team to ensure success; and 3) commitment to reliability at scale. When Cartesia needed SLURM and Crusoe didn't yet have the in-house expertise, Crusoe brought on a subject matter expert to build out a SLURM cluster—infrastructure that has since supported other Crusoe customers too.
The Results: Text-to-Speech Shipped at Sub-120ms Latency
On this foundation, Cartesia shipped Sonic, a text-to-speech model that generates high-quality, lifelike speech at sub-120ms latency—the fastest in its class at launch.Sonic offers multilingual support and features a unique voice ecosystem, where users can generate audio using voices from a diverse library, including applications such as customer support, gaming, entertainment, and content creation. Users can also customize voices with the native voice design studio with support for instant cloning and voice design. Sonic has been adopted across a wide array of verticals and will continue to power use cases for fast, reliable text-to-speech.
Bringing Real-Time Multimodal AI Models to the Edge
Cartesia’s platform has enabled developers to build real-time multimodal AI systems, most recently with their work in bringing their models to the edge. This effort allows users to run Cartesia’s models, including Sonic and other pretrained language models, directly on users’ devices. Cartesia’s work on efficient ML models will open the doors for creating interactive AI experiences that run locally on any device in a fast, secure, and personalized way.
As the world of AI offerings, use cases, hardware and platforms continue to evolve at lightning speed, relationships like this one; rooted in nimbleness, high-performance and customer success are crucial to an ever-changing landscape. So much of the competitive advantage of AI models depends on time to market and Crusoe was thrilled to support Cartesia with this effort. As the partnership has grown, Cartesia has tripled its footprint in the same Infiniband fabric and expanded storage instances. Looking ahead, Cartesia will continue to push boundaries in real time multimodal intelligence and Crusoe will continue to be a great partner in enabling the future of innovation.






.png)
.png)
.png)
















.png)
.png)
.png)










Related stories

Ready to build on Crusoe?
Talk to our team about your workload.
.png)

