Researchers introduce LIFT (Latent Information Feedback Transformer), a novel architecture that enables language models to feed deep-layer representations back to shallower layers during pretraining using teacher supervision from pretrained models. Experiments show LIFT consistently outperforms standard Transformers on language modeling and reasoning tasks while maintaining computational efficiency, demonstrating that models can effectively learn from deep-to-shallow feedback.
Rho is a foundation model for vision-language-action robots that separates adaptation into two stages: first learning a specific robot's embodiment, then adapting to particular tasks. This two-stage approach reduces the finetuning data needed by half compared to baseline models, with midtrained variants matching or outperforming competing systems on three physical robots.