Cloud
Engineering
September 29, 2026

White glove tokenization for frontier models

fastokens, Crusoe's open-source tokenizer, now has tuned paths for GLM-5.3, Kimi-K3 and DeepSeek-V4.1-Flash. No conversion needed: point it at the same files and get up to 8.4x faster throughput than HF tokenizers 1.0, already in vLLM, SGLang and NVIDIA Dynamo.

Alon Kejzman Photo
Alon Kejzman
Staff Researcher
Omri Berkovitch Photo
Omri Berkovitch
Senior Manager, Research
Omer Landau Photo
Omer Landau
VP, Engineering
Michael YenChi Ho headshot
Michael YenChi Ho
Director, Technical Product Marketing
September 30, 2026
Isometric illustration of a CPU chip surrounded by performance charts, representing fastokens tokenizer benchmarks.

Crusoe’s open-source fastokens already loads any standard tokenizer and returns exactly the tokens the model expects. This post covers an upgrade tuned for three open frontier models: GLM-5.3, Kimi-K3 and DeepSeek-V4.1-Flash. We profiled each model's tokenizer on real workloads (web text, chat, long documents and a dozen scripts) and tuned the code paths each one actually uses. There is nothing to convert and no API change: point fastokens at the same model files and get the same tokens, faster. The numbers below are what that tuning buys against HF tokenizers 1.0 and gigatoken.

LLM tokenizer performance: fastokens on three open frontier models

  • Faster on real text. On 3 GB of English and Chinese web text, chat and long documents, fastokens runs at up to 1.8 GB/s. From Rust that is 1.1–8.4× faster than HF tokenizers 1.0 and 1.2–7.9× faster than gigatoken. From Python, including building the id lists, it is 1.6–5.8× faster than both.
  • Faster per request. A long-context prompt is tokenized in under 0.5 ms: 5–6× faster than HF tokenizers 1.0 and 2.2–3.5× faster than gigatoken. A chat conversation takes about 25 µs.
  • Ahead on Hugging Face's benchmark. In tokbench (one thread, 26 corpora), fastokens beats the fastest other engine by 1.22× on GLM-5.3, 1.36× on DeepSeek-V4.1-Flash and 1.57× on Kimi-K3 in throughput (geometric means). It is at parity or ahead in 150 of 156 cells, and at most 5% behind in the other six. Against the current stable HF tokenizers (0.23) it is 7–199× faster.
  • Available today on GitHub, and integrated into vLLM, SGLang and NVIDIA Dynamo.

Why it's fast: tokenizing for the model, not the file format

A general-purpose tokenizer interprets a model through a generic pipeline and pays for that generality on every byte. fastokens loads the same files, but when it recognizes the model's tokenizer family it switches to code written for that family:

  • No regex engine. Each family's split rules are compiled into SIMD bit operations over 64-byte blocks (AVX-512 or AVX2), run in a single pass.
  • BPE shaped by the vocabulary. Every short vocabulary word is pre-loaded into a per-thread cache. Merges are single 32-bit codes looked up in tables that stay in CPU cache. Chinese, Japanese and Korean merge whole characters wherever that model's vocabulary makes that faster.
  • Special tokens. All special tokens are found in one SIMD multi-string search.
  • Parallel inside one document. Long inputs are split where the model's regex can never cross, so even a single call uses every core. That is where the 6.9× on large one-call-per-document encodes comes from.

How we measured

  • Corpus: 3 GB, 40% English web (C4), 20% Chinese web (mC4), 30% chat (ShareGPT) and 10% long documents (LongBench-v2). It is cut into documents of about 100, 2k and 245k tokens.
  • Rules: each library runs in a fresh process with its default settings, after a warm-up on separate text. Only the encode calls are timed, and each figure is the median of 3 runs. The host is a 2-socket Xeon 8468V (Sapphire Rapids) with 176 vCPUs.
  • tokbench: huggingface/tokbench at commit 2063586 with its own adapters, plus our three models, with fastokens limited to one thread.

Get started with fastokens

  • Use fastokens for frontier models on Crusoe Managed Inference without any setup, or from vLLM, SGLang and NVIDIA Dynamo.
  • Contribute to fastokens on GitHub.

Latest articles

Chase Lochmiller - Co-founder, CEO
September 29, 2026
White glove tokenization for frontier models
Chase Lochmiller - Co-founder, CEO
September 28, 2026
Optimizing and Expanding American Energy Infrastructure through Co-located Power for AI Data Centers
Chase Lochmiller - Co-founder, CEO
September 28, 2026
10,195 AI clips in one weekend, powered by clean energy

Frequently
asked questions

What is fastokens?

fastokens is Crusoe's open-source tokenizer that loads any standard tokenizer and returns the same tokens the model expects, only faster. This update adds tuned code paths for three open frontier models, GLM-5.3, Kimi-K3 and DeepSeek-V4.1-Flash, based on profiling real workloads including web text, chat, long documents and code. There's no conversion step and no API change: fastokens reads the same model files and returns identical tokens.

How much faster is fastokens than other tokenizers?

On a 3 GB corpus of English and Chinese web text, chat and long documents, fastokens reaches up to 1.8 GB/s, which from Rust is 1.1–8.4x faster than HF tokenizers 1.0 and 1.2–7.9x faster than gigatoken, and from Python (including building the id lists) is 1.6–5.8x faster than both. Per request, a long-context prompt tokenizes in under 0.5 ms, 5–6x faster than HF tokenizers 1.0 and 2.2–3.5x faster than gigatoken, while a chat conversation takes about 25 microseconds. Against the current stable HF tokenizers 0.23, fastokens is 7–199x faster.

How does fastokens compare on Hugging Face's tokbench?

In tokbench, Hugging Face's own benchmark run on one thread across 26 corpora, fastokens beats the fastest other engine by 1.22x on GLM-5.3, 1.36x on DeepSeek-V4.1-Flash and 1.57x on Kimi-K3, measured as geometric means. It's at parity or ahead in 150 of the 156 test cells, and at most 5% behind in the remaining six.

Why is fastokens faster than a general-purpose tokenizer?

A general-purpose tokenizer runs the same generic pipeline for every model, paying for that generality on every byte. Once fastokens recognizes a model's tokenizer family, it switches to code written for that family: split rules compiled into SIMD bit operations instead of a regex engine, short vocabulary words cached per thread with merges resolved as single 32-bit codes, special tokens found in one SIMD multi-string search, and long documents split so even a single call uses every core.

Where can I use fastokens?

fastokens is available now on GitHub, and it's already integrated into vLLM, SGLang and NVIDIA Dynamo. Since it uses the same model files with no API change, adopting it just means pointing fastokens at those files.

Are you ready to build something amazing?