
Nature study shows existing language models can be cheaply converted to read raw bytes
Researchers led by the Allen Institute for AI describe 'byteification', a two-stage method that retrofits standard language models to work on raw bytes instead of word-piece tokens. The converted models came close to their originals and beat earlier byte-level models of similar size.