Researchers propose Pyramid-JiT (P-JiT), a decoder-only pixel-space architecture for training text-to-image models without VAEs, achieving better results than prior models while requiring 11.3× fewer training samples and 4.3× fewer GPU-hours. The work aims to reduce computational costs for video generation by exploring efficient token compression and scaling strategies inspired by language model architectures.