White glove tokenization for frontier models
fastokens, Crusoe's open-source tokenizer, now has tuned paths for GLM-5.3, Kimi-K3 and DeepSeek-V4.1-Flash. No conversion needed: point it at the same files and get up to 8.4x faster throughput than HF tokenizers 1.0, already in vLLM, SGLang and NVIDIA Dynamo.

Crusoe’s open-source fastokens already loads any standard tokenizer and returns exactly the tokens the model expects. This post covers an upgrade tuned for three open frontier models: GLM-5.3, Kimi-K3 and DeepSeek-V4.1-Flash. We profiled each model's tokenizer on real workloads (web text, chat, long documents and a dozen scripts) and tuned the code paths each one actually uses. There is nothing to convert and no API change: point fastokens at the same model files and get the same tokens, faster. The numbers below are what that tuning buys against HF tokenizers 1.0 and gigatoken.
LLM tokenizer performance: fastokens on three open frontier models
- Faster on real text. On 3 GB of English and Chinese web text, chat and long documents, fastokens runs at up to 1.8 GB/s. From Rust that is 1.1–8.4× faster than HF tokenizers 1.0 and 1.2–7.9× faster than gigatoken. From Python, including building the id lists, it is 1.6–5.8× faster than both.
- Faster per request. A long-context prompt is tokenized in under 0.5 ms: 5–6× faster than HF tokenizers 1.0 and 2.2–3.5× faster than gigatoken. A chat conversation takes about 25 µs.
- Ahead on Hugging Face's benchmark. In tokbench (one thread, 26 corpora), fastokens beats the fastest other engine by 1.22× on GLM-5.3, 1.36× on DeepSeek-V4.1-Flash and 1.57× on Kimi-K3 in throughput (geometric means). It is at parity or ahead in 150 of 156 cells, and at most 5% behind in the other six. Against the current stable HF tokenizers (0.23) it is 7–199× faster.
- Available today on GitHub, and integrated into vLLM, SGLang and NVIDIA Dynamo.


Why it's fast: tokenizing for the model, not the file format
A general-purpose tokenizer interprets a model through a generic pipeline and pays for that generality on every byte. fastokens loads the same files, but when it recognizes the model's tokenizer family it switches to code written for that family:
- No regex engine. Each family's split rules are compiled into SIMD bit operations over 64-byte blocks (AVX-512 or AVX2), run in a single pass.
- BPE shaped by the vocabulary. Every short vocabulary word is pre-loaded into a per-thread cache. Merges are single 32-bit codes looked up in tables that stay in CPU cache. Chinese, Japanese and Korean merge whole characters wherever that model's vocabulary makes that faster.
- Special tokens. All special tokens are found in one SIMD multi-string search.
- Parallel inside one document. Long inputs are split where the model's regex can never cross, so even a single call uses every core. That is where the 6.9× on large one-call-per-document encodes comes from.

How we measured
- Corpus: 3 GB, 40% English web (C4), 20% Chinese web (mC4), 30% chat (ShareGPT) and 10% long documents (LongBench-v2). It is cut into documents of about 100, 2k and 245k tokens.
- Rules: each library runs in a fresh process with its default settings, after a warm-up on separate text. Only the encode calls are timed, and each figure is the median of 3 runs. The host is a 2-socket Xeon 8468V (Sapphire Rapids) with 176 vCPUs.
- tokbench: huggingface/tokbench at commit
2063586with its own adapters, plus our three models, with fastokens limited to one thread.
Get started with fastokens
- Use fastokens for frontier models on Crusoe Managed Inference without any setup, or from vLLM, SGLang and NVIDIA Dynamo.
- Contribute to fastokens on GitHub.
Frequently
asked questions
Crusoe Serverless Fine-Tuning bills per token processed during training, not per GPU-hour. GPU-hour pricing charges for the entire time a machine is reserved, including setup, idle time, queueing, and failures, so your cost depends on infrastructure efficiency you don't control. Token-based pricing ties spend directly to actual training work, making costs predictable from your dataset size and epoch count, with no GPUs to provision or right-size. You also only pay for what works: early stopping ends the job, and the billing, the moment your model stops improving, so you're never charged for epochs that add no accuracy.
What is fastokens?
fastokens is Crusoe's open-source tokenizer that loads any standard tokenizer and returns the same tokens the model expects, only faster. This update adds tuned code paths for three open frontier models, GLM-5.3, Kimi-K3 and DeepSeek-V4.1-Flash, based on profiling real workloads including web text, chat, long documents and code. There's no conversion step and no API change: fastokens reads the same model files and returns identical tokens.
How much faster is fastokens than other tokenizers?
On a 3 GB corpus of English and Chinese web text, chat and long documents, fastokens reaches up to 1.8 GB/s, which from Rust is 1.1–8.4x faster than HF tokenizers 1.0 and 1.2–7.9x faster than gigatoken, and from Python (including building the id lists) is 1.6–5.8x faster than both. Per request, a long-context prompt tokenizes in under 0.5 ms, 5–6x faster than HF tokenizers 1.0 and 2.2–3.5x faster than gigatoken, while a chat conversation takes about 25 microseconds. Against the current stable HF tokenizers 0.23, fastokens is 7–199x faster.
How does fastokens compare on Hugging Face's tokbench?
In tokbench, Hugging Face's own benchmark run on one thread across 26 corpora, fastokens beats the fastest other engine by 1.22x on GLM-5.3, 1.36x on DeepSeek-V4.1-Flash and 1.57x on Kimi-K3, measured as geometric means. It's at parity or ahead in 150 of the 156 test cells, and at most 5% behind in the remaining six.
Why is fastokens faster than a general-purpose tokenizer?
A general-purpose tokenizer runs the same generic pipeline for every model, paying for that generality on every byte. Once fastokens recognizes a model's tokenizer family, it switches to code written for that family: split rules compiled into SIMD bit operations instead of a regex engine, short vocabulary words cached per thread with merges resolved as single 32-bit codes, special tokens found in one SIMD multi-string search, and long documents split so even a single call uses every core.
Where can I use fastokens?
fastokens is available now on GitHub, and it's already integrated into vLLM, SGLang and NVIDIA Dynamo. Since it uses the same model files with no API change, adopting it just means pointing fastokens at those files.





