Retrofitting language models to operate over bytes

Subword tokenization, used by all leading LLMs, limits character-level understanding, biases generation, and favors English. Byte-level models have been proposed but rarely adopted because they train from scratch. This work introduces byteification: retrofitting an existing subword LLM to bytes with under 1% of a typical pretraining budget. The byteified models, including Bolmo 7B, outperform earlier byte-level LLMs and even surpass the original subword models on character understanding and some coding tasks.

Byteification establishes a connection between existing subword-level LLMs and byte-level LLMs.
  1. armcat

    There has been extensive research into token-free LLMs, but for some reason we are still operating in a token domain, so there is something to that.

    Byte Latent Transformers (BLT): https://arxiv.org/abs/2412.09871

    Charformer: https://arxiv.org/abs/2106.12672

  2. mrkn1

    subword tokenizers were never causal in the first place, so BPE was peeking at future bytes all along! TIL

  3. serioussecurity

    Wow nature got rolled. Should have stayed closer to their expertise. They were already being hustled by a lot of the applied AI work they were accepting.

More from this day

2026-10-11