This paper establishes empirical scaling laws showing that language model loss follows power-law relationships with model size, dataset size, and compute, spanning over seven orders of magnitude. The research demonstrates that larger models are more sample-efficient and that optimal training involves large models on modest data, stopping before convergence.