Customers

Base44 cuts cost per token with Crusoe Managed Inference

September 17, 2026
Software and app creation
Written
traffic running on Crusoe Managed Inference
50%+
users
10M+
ARR in the last 5 months
2X

"Our biggest fear was that cutting costs would cut the product quality. With Crusoe Managed Inference, performance and latency improved, and we're paying an order of magnitude less."

Maor Shlomo
Founder

Building software at the speed of a sentence

Most software is still built by teams of specialists over months. Base44 was founded to change that: a vibe coding platform that lets anyone describe an application in plain language and get working software back. 

Base44 grew $100 million to $200 million in annual recurring revenue in five months, a year after Wix acquired it, and has now passed 10 million users. Maor Shlomo, founder of Base44, runs it on a principle he calls compounding ambition: when AI hands you throughput, spend it on the biggest thing you can do, not on doing the same things faster.

For Base44, that meant owning the model behind the platform, and with it the unit economics of every app users build.

From renting intelligence to owning it

Every app built on Base44 runs through an LLM, and for its first year that meant frontier models. The generalist approach came with generalist economics. At Base44’s volume, frontier inference was among the largest costs in the business, and it scaled with every app users built.

To Base44 leadership, the problem was scope: Frontier models are trained to be good at everything, from GPU kernels to Chinese poetry. Base44 needs a model that is exceptional at exactly one thing: turning natural language into working web applications, with the platform's design and product decisions baked into the weights. "For many use cases, it might not be the best choice," Shlomo has said of frontier models, "especially if you can control and adapt and align a model to your needs."

Two things convinced Base44 the time to own the model was now. The company's growth generated tens of millions of real user interactions, enough data for fine-tuning to work. And open models got good enough at code generation that a fine-tuned iteration could match frontier output on Base44's core task. 

Base44 built Base 1, a custom model trained with reinforcement learning against real platform tasks: building and editing applications, scoring the outputs, and feeding the signal back into the weights. Owning the model created a new challenge: finding an inference provider that could serve it at Base44's scale and remain price performant.

From Small Scale PoC to Majority of Production Traffic

The engagement began with a proof of concept. Crusoe served a small share of Base44's production traffic, while running side-by-side evaluations against other managed inference providers.

Within two weeks, they had their answer. Base44 chose Crusoe for its reliability and price-performance.

As an agentic coding platform, its traffic carries the traces typical of that category, such as long system prompts and high cache hit rates, all originating from a single application.

If that traffic ran on a standard shared endpoint, those caching advantages would go to waste. So Crusoe set up dedicated capacity on Crusoe Managed Inference, tuned around Base44's workload instead of an average of everyone's. Crusoe built a tailored speculative decoding algorithm fitted to the patterns in Base44's own traffic, and added routing that keeps requests on the replica already holding their context without overloading any single GPU. The result is the same latency the product needs on fewer GPUs, and those savings came back as a lower cost per token.

Since then, Crusoe Managed Inference has run the majority of Base44's production traffic. "We change the model and the workload constantly, and most infrastructure partners can't move at that speed. Crusoe kept up. They didn't hand us endpoints and walk away, they got into the weeds to tune the workload with us," says Shlomo. 

What started as a trial is now the backbone of the platform. Today more than 50% of the traffic runs on Crusoe.

Frontier quality at a fraction of the cost

Moving away from frontier models can be risky. If the output gets noticeably worse, users will leave. But Base44 managed to migrate the bulk of their traffic while improving quality.

The product improved and the compute bills dropped. Output now beats the old frontier models at a fraction of the cost per token.

“Our biggest fear was that cutting costs would cut the product quality," says Shlomo. "But with Crusoe Managed Inference the performance and latency improved, and we're paying an order of magnitude less than we did with the big generalist models. Our users got a better and faster product, which was the whole point." 

Compounding ambition, compounding models

Base 1 is the first in a series. Shlomo has said Base44 has plans for larger models and deeper product integrations. And the platform itself keeps widening: mobile apps, games, slides, agents, and automations, each generating more of the interaction data the next model fine-tunes on.

More users generate more data, more data tunes better models, and better models need fine-tuning and inference capacity that scale on demand."Owning the model is our strategy. Running inference infrastructure is not. Crusoe manages the serving and the tuning, and the economics got better as we went from a small slice of our traffic to the majority of it. The only question left is how ambitious we can be."

Learn more about Crusoe Managed Inference and Base44.

Acknowledgments 

We'd like to thank our engineering team at Crusoe, specifically Yonatan Rimon and Amit Gruner, for working day and night to optimize this setup, as well as Ittai Zeidman and the rest of the Base44 team for the great collaboration.

Trusted by:

Related stories

World models

SpAItial hits 97% efficiency with Crusoe Cloud automation

Watch how SpAItial scales compute at over 97% efficiency, with Crusoe running the networking and control plane.

World models
Watch the story
Physical AI

Roboflow serves 300,000+ vision models on Crusoe Cloud

Roboflow uses Crusoe Cloud's high-performance infrastructure to power low-latency computer vision models for real-time physical AI applications.

Physical AI
Read the story
Computer Use Agents

Yutori serves browser agents on Crusoe Managed Inference

Yutori trains and serves its Navigator computer-use models on Crusoe. Learn how a cluster-wide KV cache changed cost per task for always-on background agents.

Computer Use Agents
Watch the story

Ready to build on Crusoe?

Talk to our team about your workload.