A randomized experiment found that generative AI access improved undergraduate test scores by 0.27 standard deviations immediately and one week later, with larger gains among students who used AI to explain concepts rather than generate text. Essay quality showed delayed improvements after AI was removed, driven by students shifting time toward reading and information-seeking rather than drafting.
A position paper argues that current AI agent development approaches undermine effective human oversight by both impeding oversight capabilities and degrading human cognitive skills through extended automation use. The authors propose design-level affordances and organizational protocols to support human overseers' critical judgment and counteract skill atrophy.