PSSA, a 2.7-billion-parameter language model built entirely in Rust, replaces transformer self-attention with a recurrent-convolutional architecture that processes tokens in linear time, avoiding quadratic memory scaling. Released under Apache 2.0, it achieved comparable performance to GPT-3.5 on standard benchmarks while excelling at long-context tasks up to 32k tokens, and runs 1.8× faster with 30% less RAM than PyTorch transformers.
PSSA is a non-transformer language model written in Rust from scratch that uses recurrent state-space layers, episodic memory, and plastic weights instead of attention. On matched parameters and corpus, it learns faster and generates text 12 times quicker than transformers, achieving better generalization with lower perplexity on held-out data.