How SpAItial scaled to 97% efficiency on Crusoe Cloud




“We benchmarked the cluster, and we realized that Crusoe matches the state-of-the-art speed in terms of both the GPUs and the storage. We decided to go with Crusoe because it was the best possible choice.”
When an injury ended David Novotný's professional volleyball career at 21, he sought another frontier to test himself against. He found it in AI research, earning a PhD in computer vision from Oxford with a focus on 3D deep learning. Driven by that same competitive spirit, he and his co-founders set out to pioneer physically grounded world models: AI capable of generating and reasoning about the visual appearance and physics of real as well as imagined environments.
They are achieving this through an automation-first engineering philosophy and building on managed infrastructure which frees their engineering team to ship rapidly.
What SpAItial builds
Unlike standard generative AI that relies on tokens or pixels, SpAItial's models operate natively in physical space-time to generate and reason about environments. While video models simply generate each scene view from scratch, SpAItial’s world models enable real-time interaction by generating a physically grounded representation of the scene itself. Echo-2, SpAItial’s latest model, represents environments with 3D Gaussian splats, a technique representing scenes as clouds of points which can be rendered immediately on edge, rather than waiting for a video model to generate each view from scratch.
This physical grounding is critical for robotics and physical AI, offering a direct path to embodied intelligence. Lack of training data is one of the largest constraints on physical AI, and real-world data collection can be slow and costly. SpAItial builds world models that generate and simulate synthetic environments where robots learn to act, move, and interact before real-world deployment. This addresses a core failure mode in robotics: skills learned in simulation often break down when real-world physics, lighting, and object variation don't match the training environment. This spatial reasoning capability also extends across other industries, from gaming and entertainment to CAD, engineering, construction, VR, and AR.
The challenge: frontier-scale infrastructure
Through 2025, the model SpAItial was training outgrew on-demand compute. A frontier world model has to train across a distributed cluster of machines that talk to each other over high-speed networking connecting top-tier GPUs, and on-demand services couldn't provide that.
SpAItial's training is compute-bound, with the model sharded across GPUs. One of the most significant determinants of scaling efficiency is the high speed interconnect between chips. As compute scales, the GPUs exchange data on every step, and near-linear scaling depends on fast RDMA, NVIDIA Quantum InfiniBand, and storage fast enough to keep the GPUs fed.
SpAItial's high-leverage infrastructure strategy made the constraint sharper. Every layer the team ran itself was time not spent improving the model training pipeline. For an engineering organization focused on time-to-market, the deciding factor is how much of the infrastructure a provider can absorb.
The solution: price-performant compute with Crusoe
SpAItial ran evaluations across a series of providers and chose Crusoe based on the results, benchmarking each option against its own flagship models before committing.
The cluster runs SpAItial's stack on NVIDIA Hopper GPUs. Crusoe manages the private VPC over NVIDIA Quantum InfiniBand, so the infrastructure team doesn't have to own the networking layer. Crusoe Managed Kubernetes handles the control plane and security patching the team would otherwise run itself.
"We prioritize a high-leverage engineering philosophy, so the question is always what we don't have to do ourselves. The control plane is the clearest opportunity. Crusoe Managed Kubernetes takes this load off us and ensures we can focus on differentiated work that helps the researchers and improves research efficiency," observed Saif Haq, Infrastructure Lead at SpAItial.
The impact: high leverage and 97% scaling efficiency
Operating with an automation-first mindset, SpAItial needed to move fast.
"From the first email that we sent Crusoe, within 11 days, we had already started the PoC," recalled Saif.
David observed, "It was the speed of back and forth between us and the Crusoe engineering team, which happened on Slack. The engineering team was available 24/7. And the cluster was pretty much ready. We didn't have to do any performance optimizations. The GPUs were running at top speed without any additional fixes. It was fairly seamless."
SpAItial benchmarked how the cluster scaled, with Saif adding, "During our benchmarks, we saw over 97% scaling efficiency with Crusoe.” At that efficiency, almost no compute is lost to interconnect overhead, and GPUs stay productive rather than idle as the cluster grows.
Rather than dedicating extensive resources to manage infrastructure at scale, SpAItial relies on Crusoe to handle the networking and control plane layers. This offloads the burden of infrastructure management and constant maintenance, empowering the team to prioritize the high-impact, differentiated engineering work essential to their model training.
"Not having to manage the networking layer takes a large burden off the team. Running bare metal without InfiniBand, using RDMA directly, adds a lot of work. Having this managed by Crusoe makes our lives a lot easier," said Saif.
David points to iteration speed as the biggest gain, "The biggest improvement we saw from working with Crusoe was the speed of experimentation" The team estimates it now iterates on training experiments roughly 30% faster.
Looking forward
SpAItial is building toward world models that give robots a place to learn before they act in the real world, and that let anyone generate and move through a space that doesn't physically exist. Every step in that direction is a larger model simulating more physics at higher fidelity, which turns progress into a scaling problem: more GPUs, more interconnect, and more storage moving in lockstep. Crusoe Managed Kubernetes manages that underlying complexity, ensuring the infrastructure stays performant and never bottlenecks the team’s shipping velocity.
That is the part SpAItial hands to Crusoe, so the engineers and researchers stay focused on the models and the AI infrastructure never becomes the ceiling.


















Related stories

Ready to build on Crusoe?
Talk to our team about your workload.
.png)

.jpg)