Customers

Fireworks went from deal to thousands of nodes running in days on Crusoe Cloud

September 24, 2026
Inference
0:00/0:00
0:00
Observed uptime for Fireworks customers
99.99%
Average incident resolution
<4 hours
From signing to thousands of nodes online
Hours

We continue to work with Crusoe to maintain four nines availability for our customers, which helps our business tremendously.

Yun Jin
Head of AI Infrastructure

"Intelligence is cheap these days," says Chenyu Zhao, Co-founder of Fireworks. "So how do you differentiate? You bring your own data to the equation. The companies that win in this environment are going to be the ones that end up owning their own AI."

Fireworks makes this possible with a specialized platform that helps organizations fine-tune and serve state-of-the-art open models using their own data and domain expertise. Over the past year, the platform has seen explosive growth:

  • 40 trillion tokens served per day
  • $1B annualized revenue run rate
  • 10x growth in their GPU fleet size

But serving that much traffic at a global scale is a significant compute challenge. Every token has to run on a GPU somewhere. "Our capacity demands scale linearly with our demand for tokens," Zhao says. To keep up, Fireworks needed an infrastructure partner capable of scaling at the exact same speed.

The challenge: Reliability, flexibility, and partnership at the frontier

For Fireworks, a cloud provider must meet three non-negotiable requirements:

  • Reliability: Deliver the constant availability Fireworks needs to uphold uptime SLAs across a multi-region footprint.
  • Flexibility: Handle diverse, evolving workloads, from instant chat to heavy agentic tasks, and procure the NVIDIA GPUs needed to run increasingly massive models.
  • Partnership: Co-create novel solutions to the hard networking and data transfer problems that come with managing AI clusters at this scale.

The solution: Fireworks, Crusoe, and NVIDIA build at the frontier together

Today, Fireworks runs large-scale GPU capacity on Crusoe Cloud across multiple regions, powered by the NVIDIA accelerated computing platform. Crusoe supplies the full stack underneath Fireworks' inference engine: NVIDIA GPU clusters, Crusoe Managed Kubernetes, Crusoe Command Center for observability and automation, and an engineering team that works alongside Fireworks' own. Together, they are defining what it takes to run AI at this scale.

Reliability, engineered into the platform

For Fireworks, maximizing goodput, the share of the fleet doing useful work, is critical. Instead of building and managing infrastructure from scratch, Fireworks relies on Crusoe AutoClusters, part of Crusoe Managed Kubernetes, to protect goodput automatically. It continuously verifies node health and rotates in fresh nodes as needed, so their engineers can focus entirely on the inference stack.

"Using managed services like load balancers and Crusoe Managed Kubernetes reduces the burden on our infrastructure and engineering team by an incredible amount," Zhao says. "It allows us to bring these clusters up so much more quickly, so we can provide capacity to our customers from day one."

Crusoe Command Center gives the team full visibility into cluster operations with the ability to perform critical actions automatically. "By utilizing Crusoe Command Center, we have been able to cut significant hours in operating our clusters," says Yun Jin, Head of AI Infrastructure at Fireworks.

Flexibility, integrated across the full stack

Fireworks' workloads span the full spectrum of modern AI. "AI workloads today have become very diversified," Jin says. "Traditionally it started with the chat experience, where people expect much shorter response times. But today there are more and more agent workloads. Some require very large amounts of data but have higher latency tolerance, while others require extreme speed. The infrastructure needs to be versatile enough to handle all the different workloads while remaining highly efficient."

That versatility starts at the silicon level. Each new NVIDIA GPU generation follows a proven sequence: Crusoe delivers a test node or rack, Fireworks benchmarks the latest models on its inference engine, and the two teams scale to production. Fireworks' fastest turnaround from cluster delivery to serving customers was a matter of hours. NVIDIA's architecture makes that speed possible across every workload.

"We've found NVIDIA GPUs to be very versatile, not only in performance, but in supporting all the different workloads, whether training or inference, and different model architectures," Jin says. "We can always extract the highest performance from NVIDIA hardware."

Yet access to top-tier chips is only part of the equation. Bringing them into production requires aligning power, data centers, and networking simultaneously, a massive logistical hurdle.

"Today's AI infrastructure buildout is often bottlenecked on one or two components of the full stack, whether at the chip level or the power level," Jin says. "Crusoe's capability to vertically integrate the full stack end to end is a great advantage in helping us bring capacity online fast, at high quality. 

Partnership, minus the playbook

At this scale, there are no off-the-shelf answers. Operating ahead of the industry means standard playbooks don't exist for the complex networking and data transfer challenges of large-scale AI. Crusoe and Fireworks engineers often work on the same challenges side by side. When Fireworks opens a support ticket, the first response from a Crusoe engineer comes within minutes, and the median issue is resolved in under four hours.

"We partner with Crusoe to build the monitoring system together," Jin says. "We have to define the operation procedures together. We have to build the automated failure handling system together. At every step, we work with Crusoe to keep tuning the system and keep building the solution together, working almost as the same team to reach high availability.”

The impact: Deal to deployment in days

Elastic capacity and reliability are easy to promise and hard to prove. For Fireworks, the proof came when a strategic customer hit hypergrowth and needed thousands of nodes of new capacity on short notice.

"This was not tens of nodes or hundreds of nodes, but thousands of nodes on extremely short notice," Zhao says. "I remember being on the call with Crusoe late evening Friday night, Saturday morning, trying to figure out the best way to get this capacity. We were able to strike a deal within a couple of days, and we got those nodes up and running within a few hours of signing."

The experience shaped the advice Zhao gives other founders.

"For a new founder, I would optimize for speed of execution," he says. "Working with a partner like Crusoe that's willing to get on a call with you on a late evening over the weekend to deliver the capacity or level of support you need is absolutely invaluable. This industry is operating on double time."

The same speed carries through to Fireworks' customers.

"Our unique advantage is time to market," Zhao says. "Being able to access a huge amount of capacity from day one in the regions customers need allows us to go from POC to full-scale production on the order of days, as opposed to weeks."

For Jin, the impact adds up to infrastructure Fireworks can plan its growth around.

"To enable faster growth of our business, we require infrastructure to be flexible, to be always available at high reliability," Jin says. "When we need new generation hardware, we can work with Crusoe. We expand into new regions, we get the support from Crusoe. And we continue to work with Crusoe to maintain four nines availability for our customers, which helps our business tremendously."

What’s next: Planning for the NVIDIA Vera Rubin era 

The models keep growing, and the hardware underneath them keeps changing. Fireworks' roadmap is built around staying on the newest NVIDIA silicon the moment it arrives.

"The latest models are extremely beefy models that require running on the latest hardware, like NVIDIA GB300 GPUs," Zhao says. "As we go into 2027, we expect the newer models to run on NVIDIA Vera Rubin. What we're thinking about right now with Crusoe and our other providers is how we secure enough capacity to serve these models going forward."

Securing that capacity is no longer a transaction. It's a shared planning exercise, with Fireworks forecasting demand against Crusoe's roadmap of regions, data centers, and power.

"We're no longer a small startup, and neither is Crusoe," Zhao says. "Being able to work closely together as a partnership to figure out the capacity plan for the future is extremely valuable."

All of it serves the thesis Fireworks was founded on: that AI belongs to the companies who own it.

"When I think about who's going to own AI and intelligence in the future, I don't see that as just one company or a handful of companies," Zhao says. "Every business around the world is going to own their own intelligence, built on top of their own data, their own user base, specialized for the particular problems they're trying to solve. That gives me so much hope for how democratized this ecosystem is going to be."


Fireworks brings new capacity online in hours on Crusoe Cloud.

See what you can build on the same infrastructure.

‍Explore Crusoe Cloud.
‍Learn more about Fireworks. 

Trusted by:

Related stories

Software and app creation

Base44 cuts cost per token with Crusoe Managed Inference

Base44 runs its own fine-tuned model on Crusoe Managed Inference, serving the majority of its production traffic at frontier quality and a fraction of the cost.

Software and app creation
Read the story
Developer tools

How Windsurf reduced inference costs with Crusoe Cloud - VIDEO

Uncover how Windsurf cut GPU costs 50% and held 99.98% uptime on Crusoe Cloud, scaling to support 800,000+ developers on its AI code assistant.

Developer tools
Watch the story
Energy

How Neuralwatt increased AI inference throughput by 33% on Crusoe Cloud

See how Neuralwatt boosted AI inference throughput 33% on Crusoe Cloud while cutting idle GPU power draw by over 40%.

Energy
Read the story

Ready to build on Crusoe?

Talk to our team about your workload.