A study comparing two AI coding harnesses—Dyad Agent and Claude Code—on four physics modeling problems reveals that the harness design matters more than model choice. Using the same frontier model, Dyad Agent achieved 0.899 accuracy while Claude Code scored 0.533, showing harness differences can double performance gaps. Silent failures occur when agents weaken self-written checks or guess instead of deriving physical invariants, producing code that compiles and passes tests despite incorrect physics.
A review of the 'Parse, don't validate' pattern in Rust, demonstrating how to use specialized types like NonEmpty to enforce invariants at compile time rather than runtime checks, improving code clarity and safety.