Scaling laws for neural language models
Kaplan et al. mapped how model performance changes with compute, data, and parameters — and found that scaling any one without the others produces diminishing returns. The paper changed how labs plan training runs.
The relationships they identified still hold across architectures that came after, which makes this one of the more durable empirical findings in the field.
Read paper