How Yutori serves always-on agents on Crusoe Managed Inference
.png)



Elbow room in the mind
Watching a website for a price change. Tracking vendor quotes across portals. Moving data from an email into a CRM. These digital chores need to get done, and it doesn’t have to be by you.
Yutori builds AI agents that handle those chores, and its newest model, Navigator n2, trained on Crusoe Cloud, now does them across a full computer, not just the browser. The name is the mission: a Japanese word for the well-being that comes from mental spaciousness, or as co-founder and chief scientist Dhruv Batra puts it, “elbow room in the mind.”
For decades, people have adapted their workflow to how computers work. Yutori is reversing that relationship: its Navigator models see the screen as a person does and take the action a person would take, clicking a button, scrolling, typing into a form. Serving that model at scale and at the right price is an inference problem.
Why web agents need to see
Today, most software agents’ web interaction depends on APIs, structured interfaces that hand data directly to machines. If a website doesn’t offer one, the agent gets stuck. Millions of websites built, in every way imaginable, share one design constraint. They were built for people to look at. A small-town restaurant owner is busy running a restaurant, not modernizing an API for agents.
Batra believes that “if you and I can sit on a browser and solve a task, then machines should be able to operate it the same way.” Navigator uses computer vision to read the same pages people do. Yutori claims it near-matches frontier-model accuracy on browser control while running 2-4x faster at 3-5x lower cost. The more complex the task, the more that advantage grows.
Fourteen people building frontier models
Training the Navigator models in-house takes serious compute. Yutori is 14 people, and every hour spent managing infrastructure is an hour not spent on the model.
Serving Navigator is another challenge. Yutori’s Scouts run in the background: one monitoring hundreds of websites for a price change is a batch job, with no human waiting on the response. A single Scout task can mean hundreds of Navigator calls. Unlike chat agents, Yutori does not need to optimize for time-to-first-token. For background agents there are two indicators of success: that each task is cost effective, and that the fleet stays up around the clock. Cost comes down to throughput, because the more calls a set of GPUs can handle, the less each call costs.
The multistep workload is unusual too. Each step adds a small increment to a long, repeated input, and the output is a short action: click here, type this. The repetitive input means the system should read it once and cache the result. Servers that don’t share a cache redo that work on every request, leading to slower responses at higher cost.
One partner, from training to serving
Yutori’s relationship with Crusoe spans the full model lifecycle, from infrastructure to inference. It started in November 2025 with training on reserved capacity on NVIDIA GPUs. The team trains and post-trains its models in-house. As the models became ready to serve, the relationship expanded to Crusoe Managed Inference.
For serving, Yutori benchmarked managed inference providers on end-to-end latency across a range of request volumes. Crusoe came out ahead at production request rates and differentiated on guaranteed per-node SLAs, which was important for a product that can be run asynchronously. Yutori selected Crusoe Managed Inference in February 2026, and Navigator was in production by March.
Yutori signed on as a design partner, drawn by what Batra calls, “Crusoe’s unique expertise in the inference engine for deploying models.” The ask was to stay under a hard latency ceiling while maximizing throughput. “I gave Crusoe throughput and latency targets that would make it beneficial for us to work together, and Crusoe met those targets,” he says.
Crusoe’s engineers built a Tailored Deployment around Yutori’s cache-heavy workload. They tuned endpoint configuration, API features, and model lifecycle management across frequent checkpoint iterations. Crusoe mutualized the memory to give Yutori a cluster-size KV cache. Because every server shares one cache, no request starts from scratch.
Yutori started on a small endpoint and scaled up roughly 10x as Scouts moved into production and demand grew. Growth meant adding to the same deployment rather than re-benchmarking a new provider.
Frontier accuracy at less than a third of the cost
"Because Crusoe maximizes throughput, every token costs less: for Crusoe, for me, and for my customers.”
“Because Crusoe maximizes throughput, every token costs less: for Crusoe, for me, and for my customers,” Batra says.
The proof is in cost per task, measured on standard browser benchmarks such as filtering apartment listings on a real estate site. “On public benchmarks, a frontier model costs about $2.80 to complete a standard task. On Crusoe’s infrastructure, our n1.5 model costs $0.80 per task.” That difference decides whether it is cost-effective to use an agent at all. “If I came to you and said ‘I can get you a restaurant reservation, but it’s going to cost you $50,’ that product is dead on arrival,” notes Batra.
Crusoe has provided a reliable foundation for Yutori. “We are deeply appreciative of Crusoe’s high degree of reliability,” says Batra, “I don’t need to have on-call staff at 2AM trying to fix problems.”
From browser control to computer control
Yutori’s vision extends beyond the browser to agents that can do anything a person can do on a computer. Their latest model, Navigator n2 operates in a Linux or macOS sandbox. It clicks and types when a UI is the right tool, writing code when it isn’t, and sustaining hundreds of steps per task. In the future, Batra says, “We will be able to unlock minutes and then hours of knowledge work for enterprises.”
Learn more about Crusoe Managed Inference.
Learn more about Yutori.


































Related stories

How Neuralwatt increased AI inference throughput by 33% on Crusoe Cloud
See how Neuralwatt boosted AI inference throughput 33% on Crusoe Cloud while cutting idle GPU power draw by over 40%.

How Cartesia built real-time multimodal AI on Crusoe
Discover how Cartesia shipped Sonic, the fastest text-to-speech model in its class, at sub-120ms latency on Crusoe Cloud.

Ready to build on Crusoe?
Talk to our team about your workload.