Base44 cuts cost per token with Crusoe Managed Inference




"Our biggest fear was that cutting costs would cut the product quality. With Crusoe Managed Inference, performance and latency improved, and we're paying an order of magnitude less."
Building software at the speed of a sentence
Most software is still built by teams of specialists over months. Base44 was founded to change that: a vibe coding platform that lets anyone describe an application in plain language and get working software back.
Base44 grew $100 million to $200 million in annual recurring revenue in five months, a year after Wix acquired it, and has now passed 10 million users. Maor Shlomo, founder of Base44, runs it on a principle he calls compounding ambition: when AI hands you throughput, spend it on the biggest thing you can do, not on doing the same things faster.
For Base44, that meant owning the model behind the platform, and with it the unit economics of every app users build.
From renting intelligence to owning it
Every app built on Base44 runs through an LLM, and for its first year that meant frontier models. The generalist approach came with generalist economics. At Base44’s volume, frontier inference was among the largest costs in the business, and it scaled with every app users built.
To Base44 leadership, the problem was scope: Frontier models are trained to be good at everything, from GPU kernels to Chinese poetry. Base44 needs a model that is exceptional at exactly one thing: turning natural language into working web applications, with the platform's design and product decisions baked into the weights. "For many use cases, it might not be the best choice," Shlomo has said of frontier models, "especially if you can control and adapt and align a model to your needs."
Two things convinced Base44 the time to own the model was now. The company's growth generated tens of millions of real user interactions, enough data for fine-tuning to work. And open models got good enough at code generation that a fine-tuned iteration could match frontier output on Base44's core task.
Base44 built Base 1, a custom model trained with reinforcement learning against real platform tasks: building and editing applications, scoring the outputs, and feeding the signal back into the weights. Owning the model created a new challenge: finding an inference provider that could serve it at Base44's scale and remain price performant.
From Small Scale PoC to Majority of Production Traffic
The engagement began with a proof of concept. Crusoe served a small share of Base44's production traffic, while running side-by-side evaluations against other managed inference providers.
Within two weeks, they had their answer. Base44 chose Crusoe for its reliability and price-performance.
As an agentic coding platform, its traffic carries the traces typical of that category, such as long system prompts and high cache hit rates, all originating from a single application.
If that traffic ran on a standard shared endpoint, those caching advantages would go to waste. So Crusoe set up dedicated capacity on Crusoe Managed Inference, tuned around Base44's workload instead of an average of everyone's. Crusoe built a tailored speculative decoding algorithm fitted to the patterns in Base44's own traffic, and added routing that keeps requests on the replica already holding their context without overloading any single GPU. The result is the same latency the product needs on fewer GPUs, and those savings came back as a lower cost per token.
Since then, Crusoe Managed Inference has run the majority of Base44's production traffic. "We change the model and the workload constantly, and most infrastructure partners can't move at that speed. Crusoe kept up. They didn't hand us endpoints and walk away, they got into the weeds to tune the workload with us," says Shlomo.
What started as a trial is now the backbone of the platform. Today more than 50% of the traffic runs on Crusoe.
Frontier quality at a fraction of the cost
Moving away from frontier models can be risky. If the output gets noticeably worse, users will leave. But Base44 managed to migrate the bulk of their traffic while improving quality.
The product improved and the compute bills dropped. Output now beats the old frontier models at a fraction of the cost per token.
“Our biggest fear was that cutting costs would cut the product quality," says Shlomo. "But with Crusoe Managed Inference the performance and latency improved, and we're paying an order of magnitude less than we did with the big generalist models. Our users got a better and faster product, which was the whole point."
Compounding ambition, compounding models
Base 1 is the first in a series. Shlomo has said Base44 has plans for larger models and deeper product integrations. And the platform itself keeps widening: mobile apps, games, slides, agents, and automations, each generating more of the interaction data the next model fine-tunes on.
More users generate more data, more data tunes better models, and better models need fine-tuning and inference capacity that scale on demand."Owning the model is our strategy. Running inference infrastructure is not. Crusoe manages the serving and the tuning, and the economics got better as we went from a small slice of our traffic to the majority of it. The only question left is how ambitious we can be."
Learn more about Crusoe Managed Inference and Base44.
Acknowledgments
We'd like to thank our engineering team at Crusoe, specifically Yonatan Rimon and Amit Gruner, for working day and night to optimize this setup, as well as Ittai Zeidman and the rest of the Base44 team for the great collaboration.


























Related stories

Ready to build on Crusoe?
Talk to our team about your workload.


.png)