This research introduces two methods to convert token logits to byte logits and conducts a large-scale study comparing byte-based and token-based language models across scaling dimensions. The study finds that while token models perform better initially, byte models eventually surpass them with increased compute, achieving superior data efficiency and performance on downstream tasks.