Better networking, 25% lower training costs: How Linum built a frontier video model on Crusoe





From the performing arts to frontier AI
Linum co-founders Sahil Chopra and Manu Chopra grew up at a school focused on the performing arts, but when diffusion models emerged in 2022, they pivoted toward a different mission: make animation accessible so that anyone can make their own shows and movies.
To tear down the financial walls of traditional production, they built a text-to-video model trained entirely from scratch, from the ground up, by just two people. In January 2026, Linum released Linum v2; a 2 billion parameter, Apache 2.0-licensed text-to-video model capable of generating 360p and 720p video.
.gif)
Crusoe powered the final, full-scale training run. The reason? Crusoe’s internet uplink speed.
The challenge: solving for data density
The first hurdle was the sheer density of video data. Generating just five seconds of 720p video requires processing roughly 110K tokens. To handle these massive context windows while maintaining the batch sizes needed for frontier training, Linum needed the larger and faster 141GB of HBM3e memory provided by NVIDIA H200 GPUs.
For a two-person team training on petabyte-scale video data, speed and cost were existential constraints. Video data is orders of magnitude larger than text, and every unnecessary data transfer adds up fast. Egress fees alone made certain providers unworkable. Crusoe doesn’t charge them.
Moving that volume to a new provider wasn't feasible on their timelines. Instead, Linum built an architecture that streams data live from Cloudflare R2 during training, caching it directly to NVMe drives on each node. When the network is fast enough, this approach eliminates migration overhead entirely and outperforms NFS storage.
The catch: poor uplinks don't degrade performance linearly. As Linum scaled to more nodes at other providers, throughput dropped noticeably. The streaming strategy simply wouldn't work without a provider whose public internet backbone was fast and reliable at scale.
"That awareness of the entire problem space is what I think leads to good networking results. If you've built datacenters, you're thinking about the uplinks versus other folks who are just renting out rack space from a third party provider. They don't have control on what those uplinks are. If you're dealing with multimodal data, networking is the first thing you should look at.”
The solution: an AI-first backbone network
Linum benchmarked multiple providers during their POC. Crusoe's internet uplink was the best performing that they tested. The backbone combines diverse transit providers with extensive peering, providing path diversity and reducing reliance on any single congested network.
The results: 25% lower training costs from faster networking and data throughput
During their POC, Linum found that data-loading bottlenecks at other providers would have made training 25% more expensive for the same work. Crusoe Cloud's backbone networking eliminated that overhead.
Crusoe engineers in the loop from day one
The POC told Linum everything they needed to know. Crusoe extended the original two-day trial to five days so the team could complete their benchmark after getting the system running. Linum ran their training on Slurm-managed clusters, which handled the job scheduling and multi-node orchestration that would otherwise have consumed engineering time they didn't have. From NVIDIA CUDA® configuration to cluster setup, Crusoe engineers were hands-on throughout the onboarding process.

That hands-on engagement continued into the Linum v2 training run, which ran over the holidays. When outages occurred, Crusoe communicated proactively about what was happening and what was being done. Critically, the team provided immediate clarity on fault attribution, so Linum knew the issue was on the infrastructure side and could stand down rather than debugging their own code.
Looking ahead: What's next for Linum and Crusoe
Linum’s next step is scaling to larger models. Linum’s goal hasn't changed: make animated filmmaking accessible to anyone with a story to tell. The infrastructure to get there just has to be good enough to keep up. Explore Crusoe Cloud.


































Related stories

How Cartesia built real-time multimodal AI on Crusoe
Discover how Cartesia shipped Sonic, the fastest text-to-speech model in its class, at sub-120ms latency on Crusoe Cloud.
.png)
How Odyssey trained world models on Crusoe Cloud
Learn how Odyssey cut infrastructure costs up to 4.4x on Crusoe Cloud and scaled its Explorer launch to hundreds of thousands of users.

How Neuralwatt increased AI inference throughput by 33% on Crusoe Cloud
See how Neuralwatt boosted AI inference throughput 33% on Crusoe Cloud while cutting idle GPU power draw by over 40%.
