Giant AI models don't always overfit despite having capacity to memorize training data because data structure matters more than model size, and training algorithms have implicit biases that favor simpler solutions. The double descent phenomenon shows test error can improve again after initial overfitting, while the distribution of noise across weak feature directions can preserve predictive performance on unseen examples.