How green is Crusoe's inference engine?
When measuring on identical hardware, we found that tuning Crusoe's inference stack cut energy per output token by 26.9%, while MemoryAlloy cut energy by 23% in cache tests.

As models get bigger and inference traffic grows, AI's energy footprint is becoming one of the defining sustainability topics of our time. Every token served draws power, and at scale, that compounds into a significant environmental cost.
Currently, the industry’s dominant answer to “how efficient is the infrastructure serving AI inference?” is a single metric: Power Usage Effectiveness (PUE). PUE is a data center level metric that shows how efficiently a building delivers power to its IT equipment. It is useful for a bird’s eye view, but a layer removed from the inference stack itself. A PUE of 1.0 is the theoretical ideal because it would mean that 100% of the energy goes directly to the IT equipment without overhead like cooling. While that number would be impossibly unrealistic, the industry has spent years grinding real-world facilities closer to it. But it says nothing about how efficiently that power gets turned into tokens. Looking higher in the stack, application-layer techniques such as caching and prompt compression can reduce the raw number of tokens being served, but do not change the underlying efficiency of the inference stack.
This led to the question: can the inference stack itself be a lever for energy efficiency? This post walks through how we measured it on Crusoe’s production inference stack, and what we found.
Tokens per watt: a new metric for energy efficiency beyond PUE
Since PUE measures the efficiency of the building, not the workload, the implicit assumption underneath it – that a watt of compute is a watt of compute – doesn't hold for inference. Two GPUs drawing identical power can produce wildly different amounts of useful work depending on the model, the request pattern, and the software stack orchestrating them.
As a result, a new framing is filling the gap: tokens per watt. It may seem like a straightforward metric, but it’s also a surprisingly hard one to measure. Tokens per watt only means something when you specify which workload, which hardware, and which power boundary: GPU only, full server, or rack with cooling included. There's no agreed-upon methodology yet, and a tokens-per-watt number without that context is meaningless. These tokens must also meet a quality bar to be considered useful. In our case, every deployment we ship meets standard quality and intelligence benchmarks, validated against industry standards like GPQA and SWE-bench.
Connecting tokens per watt to energy
To measure the exact energy footprint of a specific inference job, we can also look at the inverse work metric: joules per output token. While tokens per watt represents your engine's overall efficiency rate, joules per token represents the energy cost of every token delivered.
With that framing in mind, we set out to answer the question: can the inference stack itself be used as a lever for sustainability, not just the building it runs in?
Inference stack optimizations
Alongside in-house optimizations natively baked into our high performance inference stack, Crusoe's inference stack exposes tunable knobs at distinct layers. In our experiments, we ran tests that adjusted these layers:
- Scheduling: how requests are dispatched across multiple replicas of a model
- Transport: how the KV cache and intermediate state move between GPUs. Execution, which is how the engine actually runs the model once a request lands on a GPU, lives here.
- MemoryAlloy™, our cluster-wide distributed KV cache that uses peer-to-peer GPU interconnects to skip the CPU round-trip competing caches rely on, is a feature that is offered in the transport layer. (Read more about MemoryAlloy and its benefits here). As explained below in the setup, MemoryAlloy was not adjusted as a part of our first experiment.
The setup
We ran two experiments:
- The head-to-head. Two side-by-side deployments of the same workload — one tuned across scheduling and transport, one running the stock defaults. MemoryAlloy was off in both deployments so we could test it in isolation against
vLLM’s default cache. - The MemoryAlloy isolation. Two passes over an identical prompt set with MemoryAlloy on, and the other using
vLLM.
For each experiment, we ran the same workload on the same hardware twice: once with the stack at sensible defaults — what you'd get out of the box with vLLM — and once with each layer deliberately set to our current production deployment. The delta is what we're measuring: how much further can you go by tuning the inference stack itself, rather than the hardware under it?
For all runs, we ran an identical text prompt stream (seed 42, 1024 in / 256 out, 256 concurrent requests per arm) using GLM-5.2-NVFP4 on two 8xB200 nodes. In other words, both arms saw byte-identical inputs at identical load, so any difference between them comes from the knob configuration and not the traffic hitting it. For each set, both were run at stock clocks, and shared the same instance type, nodepool, cluster, IB partition, and location (our Norway data center, which is powered 100% by hydroelectric energy). These workloads were the only ones being run on these nodes at these times.
How we measured power usage
To measure power usage from each workload, we sampled GPU power draw from inside the DCGM exporter pod running on each model's node, at 5-second intervals. We defined energy as mean power × window duration.
- Included: The GPU board power only for the 8 GPUs serving that model
- Excluded: Host CPU, NICs, chassis fans, PSU losses, facility cooling
There are tradeoffs with this approach; some studies will look at a broader picture versus just the GPU power itself. But for this experiment, we went with GPU-only. We did cross-check against the full server draw read from the node's BMC/IPMI sensors at 1Hz (GPUs, host CPU, NICs, local storage, fans), and saw similar percentage improvements — so the GPU-only number holds up as a proxy for this. We deliberately exclude facility cooling overhead, as that bleeds into PUE territory.
Last, the tokens served are counted from the inference server’s own metrics, so numerator and denominator come from the same time window.
Methodology
The vanilla workload utilized the following defaults, which come from an official vLLM recipe. The only difference is the KV cache dtype, which does not impact the result since all requests had zero cache hits in this experiment.
Before the head-to-head, we sweep each layer for the chosen workload to the deployment we use in production, which are settings that clear a fixed quality bar. That tuned configuration is the "after." The "before" is the same stack with each knob left at its conventional default.
Across both runs, only the dials and knobs configuration changes. Everything else — model weights, request distribution, sampling parameters — is identical. For the MemoryAlloy test, we ran a vLLM cache job against one with MemoryAlloy caching on, with all other dials equal.
Once the jobs run, we collect tokens per watt for each configuration over the steady-state window, excluding the first 60 seconds of each run to let caches warm.
If the tuned configuration is meaningfully more efficient and quality holds, we've shown that the inference stack is a legitimate energy-design surface, not just a software-performance one.
Results
Experiment 1: optimized vs. vanilla deployment
From our benchmarking, we found that the optimized deployment served the same number of output tokens per request, resulting in 17% higher overall throughput on 14.2% less power. This translates to 26.9% less energy per output token (1.686 → 1.232 J/tok) or a 36.9% increase in tokens per watt (0.593 → 0.812 TPW). It was also 28% faster per request end-to-end (22.7s → 16.3s).
Experiment 2: MemoryAlloy cache vs vLLM Cache
In our MemoryAlloy isolation test running two passes over an identical prompt set at 4 QPS, MemoryAlloy achieved a 30.7% cache hit rate (via its offload tier) compared to vLLM’s 19.8% (HBM prefix cache). MemoryAlloy makes better use of the GPU memory we already have: more previously computed context stays cached and reachable, so fewer cycles go to recomputing it. As a result, this increased tokens per watt by 30.0% (0.111 → 0.144 TPW) while reducing energy per output token by 23% (9.020 → 6.941 J/output token). As we increased the time, we also found that with higher cache hits, we can bring the energy required per generated token much smaller.
Takeaway
The question we opened with was whether the inference stack itself could be a lever for sustainability. On the same hardware, running the same workload, tuning across layers cut energy per output token by 26.9%. In a separate test, MemoryAlloy cut energy per token by 23% against its own vLLM control. By tailoring stack optimizations directly to specific customer production needs, the Crusoe Inference team maximizes token throughput and minimizes latency, proving that the inference stack itself is a primary lever for both peak performance and reduced energy consumption per token. We see additional opportunities to enhance sustainability through software optimization and will continue making improvements so the greener path is also the easier one for Crusoe Cloud customers.
Ready to run your inference workloads on a stack tuned for efficiency? Get started with Crusoe Managed Inference.
Frequently
asked questions
What is tokens per watt?
Tokens per watt is a measure of how much useful inference output a system produces for the power it draws, calculated as tokens generated per second divided by watts consumed. Crusoe's benchmark of GLM-5.2-NVFP4 on two NVIDIA B200 GPU nodes measured 0.812 tokens per watt on its tuned inference stack versus 0.593 on vLLM defaults.
How does joules per output token relate to tokens per watt?
It is the inverse metric. While tokens per watt represents your engine's overall efficiency rate, joules per token represents the energy cost of every token delivered.
Can software tuning make LLM inference more energy efficient?
Yes. Crusoe's benchmark demonstrated that tuning scheduling, transport layers, and KV-caching on identical GPUs reduced energy per output token without changing the underlying hardware or sacrificing model quality.
%20(2).jpg)




