GigaToken: A Drop-in Replacement for 1000x Faster Language Model Tokenization

GigaToken: ~1000x faster Language model tokenization

GigaToken: A Drop-in Replacement for 1000x Faster Language Model Tokenization

I built GigaToken to solve the bottleneck of slow text processing in language modeling. This Rust-based tool offers a drop-in replacement for HuggingFace Tokenizers and Tiktoken, delivering speeds up to 1000 times faster on modern CPUs. By optimizing pretokenization with SIMD and minimizing Python overhead, you can tokenize massive datasets like Common Crawl in hours instead of days while maintaining exact compatibility with existing workflows.

At the rates we see on the EPYC CPU, you could tokenize the entirety of Common Crawl in just under 6.5 hours!

More from this day

2026-07-22