This paper addresses vulnerabilities in mechanistic interpretability of large language models by showing that interpretable replacement networks (IRNs) produce unstable interpretations under minor input perturbations. The authors propose the first formal verification framework using reachability analysis to certify faithfulness bounds and demonstrate that verification-aware training improves interpretation reliability for safety auditing.