The visual cortex of physical AI: how Roboflow serves 300,000+ vision models on Crusoe Cloud




"The support has been first class. We could actually talk to human beings at Crusoe. Crusoe gives us dedicated GPU capacity with the economics of a reservation and the responsiveness of a startup."
Making computers see
We take in most of the world through our eyes, but that’s not how computers consume information. Roboflow was founded in 2020 to change that, with a platform that lets any team build, train, and deploy computer vision models, whether they're a major AI lab or a railroad company.
Today, most of that data capture still runs through people, who observe the world and type what they see into a machine. Brad Dwyer, Roboflow's co-founder and CTO, believes that relationship is broken. "We think of Roboflow as the eyes of physical compute. If you can let the computer interact directly with the world, then it can start serving the humans, instead of the human serving the computer.”
A technology that once required a big-tech budget is now in the hands of teams across manufacturing, logistics, warehousing, robotics, and autonomous driving, plus thousands of researchers and hobbyists. And Roboflow's open-source tools, including its Apache-licensed inference server and Supervision library, extend that reach to vision developers worldwide.
Why vision doesn't scale like language
Dwyer frames the difference between vision and language AI in biological terms. "Perception and cognition are really two different problems to solve. If you look at biology, you can see the same thing: our brain has a visual cortex that is separate from our frontal lobe, and they're specialized in different things."
Language is thoroughly documented, so LLMs can learn much of it from the internet. Vision contends with the state of the natural world, where critical decisions are often edge cases: whether a part is assembled correctly inside an engine compartment is something only that manufacturer has ever seen, and the training data can't be scraped. "With vision, almost every valuable use case is a long-tail use case," Dwyer says.
So instead of one giant model for everyone, Roboflow serves thousands of customers, each running their own fine-tuned model: over 69,000 distinct computer vision models every month. Every one is expected to return predictions in under 100 milliseconds. "The key challenge is model residency: how many models can you bin pack into these GPUs?" says Sachin Agarwal, ML Infrastructure & Security Engineering Lead at Roboflow. "It's a very dynamic situation where models are being brought in and out of the fleet as users come and go."
Those compute needs became especially acute in fall 2025, when Roboflow partnered with Meta to launch the SAM 3 foundation model, far larger than its typical fine-tuned models, with demand impossible to predict. "We didn't know whether that was going to send us a hundred users, a thousand users, or a million users," Dwyer says.
From a two-week POC to a production fleet
After an initial introduction by Valor Equity Partners, the relationship moved quickly. The engagement began with a two-week proof of concept running research workloads on NVIDIA H100 GPUs and NVIDIA L40S GPUs on Crusoe Cloud. Within weeks, Roboflow signed for reserved capacity, and when Meta released SAM 3 in November, the model was operationalized on Crusoe on day one.
"The flexibility Crusoe gave us, supporting the initial launch of SAM 3 and then adapting as we went, was really helpful," Dwyer says. "Crusoe was agile and able to give us a more personal touch than some of the hyperscalers, and actually work with us on the launch."
Running SAM 3 in production convinced the team to bring Crusoe into bigger plans, starting with the serverless inference system its own team had built. "We were going to run this serverless system in another hyperscaler, but then we thought, wait a minute, we could run this on Crusoe," Agarwal says. "So we pivoted, and we've been increasing our footprint ever since."
Today, Roboflow workloads span across production inference and research, with high availability across two Crusoe data centers. The stack is broader than GPUs. CPU-heavy queueing and scheduling systems, key-value stores, and custom routing servers all run on Crusoe compute. Crusoe's network fabric with NVIDIA Quantum Infiniband networking sustains the constant, heavy traffic Roboflow's multicloud architecture demands. And Crusoe charges no egress fees, so data moves between clouds without a transfer penalty.
"There are a lot of providers selling GPUs out there, but ultimately the GPU is most useful when it's backed by reliable infrastructure," Agarwal says. "We've had good success with Crusoe: the network fabric is very performant, and we're able to pump a lot of gigabits on a constant basis."
From day one, Roboflow has consumed Crusoe entirely through Crusoe Managed Kubernetes (CMK). "Crusoe Managed Kubernetes has really simplified our infrastructure operational work. We don't need to stand up GPU Kubernetes clusters and maintain them. And it comes with all the bells and whistles: driver management, storage, DNS, the networking stack," Agarwal says. "We would not be on Crusoe if there was no CMK, to be honest."
Faster inference, leaner operations, better economics
Today, all of the 300,000+ computer vision models run on Crusoe infrastructure. Roboflow's serverless system processes 100s of millions of requests each month, with peaks of thousands of requests per second. Median latency is around 50 milliseconds, with the 95th percentile under one second, fast enough for real-time vision from defect detection to drone footage.
For Agarwal, the reserved capacity itself has become a strategic asset. "The reservations and the capacity that we've got in Crusoe is the product," he says. Latency-critical inference workloads always have instant headroom, while training jobs run against available capacity and can be preempted and resumed from checkpoints, keeping the reservation fully utilized.
Crusoe Managed Kubernetes has translated directly into engineering capacity. "Maintaining Kubernetes clusters is at least a two-to-three engineer job. It needs constant handholding," Agarwal says. "Crusoe Managed Kubernetes keeps getting upgraded for us. Just the QA on something that gets updated every three months is a full-time job. So that's a huge unlock in time for the team." Instead, the infrastructure team spent that time building the serverless inference system itself.
The economics point the same way. "You pay a huge premium at the hyperscalers," Dwyer says. "If you're a more sophisticated user, you want to capture part of that as margin, and it's hard to do that if your underlying infrastructure is too expensive. Crusoe offers more price-performant infrastructure."
When the team needs help, support often reaches out directly. "Crusoe is the only GPU cloud where I've had spare capacity show up in a Slack thread before we opened a support ticket," Agarwal says.
This responsiveness extends to everyday troubleshooting. "The support has been first class. We could actually talk to human beings at Crusoe," Agarwal adds. "When hardware fails, Crusoe can help us replace nodes in an hour. Even the highest level of hyperscaler support would take much longer. Crusoe gives us dedicated GPU capacity with the economics of a reservation and the responsiveness of a startup."
Seeing what comes next
Roboflow is betting on where vision AI goes from here: vision language models, video inference, and ever-larger foundation models. "LLMs are only about text, but 90 percent of our perception is through vision," Agarwal says. "It's really the next frontier."
That frontier is physical AI. According to Roboflow, most of the major robotics companies are Roboflow customers, many building their dataset and training stacks on the platform. And as robots, vehicles, and warehouses learn to see, the models behind them keep growing.
Dwyer puts a number on that future: the latency between your eye and your brain is roughly the same as between the edge and the cloud, just 40 to 50 milliseconds. "There's certainly a place where the frontal lobe of the AI organism runs in the cloud, and the visual cortex runs at the edge," he says. “As Roboflow's ambitions grow, the frontal lobe runs on Crusoe.”
"We like working with folks who are going to be long-term partners versus just commodity providers," Dwyer says. "It's been awesome how Crusoe has worked with us closely over the last year to meet our needs, and as an infrastructure partner as we launch new products."
Learn more about Crusoe Cloud and Roboflow.


























Related stories

SpAItial hits 97% efficiency with Crusoe Cloud automation
Watch how SpAItial scales compute at over 97% efficiency, with Crusoe running the networking and control plane.
.png)
Yutori serves browser agents on Crusoe Managed Inference
Yutori trains and serves its Navigator computer-use models on Crusoe. Learn how a cluster-wide KV cache changed cost per task for always-on background agents.

Linum built a text-to-video model on Crusoe with very real savings
Find out how Linum, a two-person team, trained a frontier text-to-video model on Crusoe Cloud and cut training costs 25%.

Ready to build on Crusoe?
Talk to our team about your workload.