Retrofitting language models to operate over bytes
Subword tokenization, used by all leading LLMs, limits character-level understanding, biases generation, and favors English. Byte-level models have been proposed but rarely adopted because they train from scratch. This work introduces byteification: retrofitting an existing subword LLM to bytes with under 1% of a typical pretraining budget. The byteified models, including Bolmo 7B, outperform earlier byte-level LLMs and even surpass the original subword models on character understanding and some coding tasks.
Byteification establishes a connection between existing subword-level LLMs and byte-level LLMs.
- armcat
There has been extensive research into token-free LLMs, but for some reason we are still operating in a token domain, so there is something to that.
Byte Latent Transformers (BLT): https://arxiv.org/abs/2412.09871
Charformer: https://arxiv.org/abs/2106.12672
- mrkn1
subword tokenizers were never causal in the first place, so BPE was peeking at future bytes all along! TIL
- serioussecurity
Wow nature got rolled. Should have stayed closer to their expertise. They were already being hustled by a lot of the applied AI work they were accepting.