Researchers introduce LIFT (Latent Information Feedback Transformer), a novel architecture that enables language models to feed deep-layer representations back to shallower layers during pretraining using teacher supervision from pretrained models. Experiments show LIFT consistently outperforms standard Transformers on language modeling and reasoning tasks while maintaining computational efficiency, demonstrating that models can effectively learn from deep-to-shallow feedback.