Inference
that scales
Serve models with ultra-low latency and high throughput. Now with flexible deployment options that fit your workload.

Crusoe's inference engine is powered by MemoryAlloy™ technology, a unique cluster-native memory fabric that persists across nodes and enables intelligent routing to eliminate redundant prefill computation.

Flexible deployment options
Model hub




%204.png)
DeepSeek V4 Flash
DeepSeek V4 Pro
Gemma 4 31B it
GPT-OSS 120B
GPT-OSS 20B
Kimi K2.6
Llama 3.1 8B Instruct
Nemotron 3 Nano 30B A3B FP8
Nemotron 3 Nano Omni 30B A3B Reasoning
Nemotron 3 Super 120B A12B FP8
Nemotron 3 VoiceChat
Qwen3 8B
Qwen3.5 9B
Qwen3.5 2B
Qwen3.6 35B A3B
Yutori n2
Nemotron 3.5 Lightning
GLM 5.3 Flash
GLM 5.3
DeepSeek V4 Flash 0731
Qwen 3.5 4B
Qwen 3.8 27B
Unmatched performance, flexibility, and scalability
Crusoe Intelligence Foundry, designed for AI developers
Crusoe inference engine vs vLLM

Frequently
asked questions
Yes. Select the Tailored Deployments option to connect with us about using your own model.
You can also now fine-tune open-source models in Crusoe Intelligence Foundry using Serverless Fine-Tuning and go live in one click with Self-Serve Deployments.
Please select the Tailored Deployments option to connect with us about exploring a dedicated instance.
Self-Serve Deployments offers two profiles: throughput, optimized for cost and resource efficiency, and responsiveness, optimized for low latency.
View all currently supported models in Crusoe Intelligence Foundry. If the model you need isn't available yet, contact us about Tailored Deployments.

%203.png)