Absolute-Address Partial Residency
Description
Caches, memoization, and lookup tables reuse previously computed information when that information
remains valid and the corresponding object can be identified again. This principle itself is not new and is not
claimed as a contribution of this record.
This record refers to the proposed structure as Absolute-Address Partial Residency (AAPR). AAPR asks
whether the same principle can be applied to execution relations in a trained dense Transformer. After
training, the model weights, architecture, and the learned function mapping an internal state to the next state
remain fixed during inference, while actual activations remain input-dependent. The invariance of the trained
model therefore does not by itself imply that execution paths can be reused. An additional condition must be
tested: can some input-dependent execution relations be observed once, recorded, and reused later?
To test this, post-hoc identifiers were assigned to all target channels of a trained dense Transformer.
Activation histories were recorded while passing multiple inputs through the unmodified model, and a static
transition mapping keyed by previous activation addresses was constructed. In an NF4 4-bit Gemma 3 4B
environment, closed-task validation using 80 of 2,560 MLP down_proj output channels per layer reproduced
the same top-1 result as an oracle that referenced the actual residual at every hop, across two seeds and
twelve hops for several task categories. A fact_geo failure caused by insufficient calibration coverage
recovered after adding fact-type examples and rebuilding only the transition map.
A path that physically computed only the selected rows matched a path that performed full computation and
then retained the same rows by masking. Validation answers were also preserved after portions of the original
L1-12 down_proj weights were physically released from VRAM.