A developer tested five coding agents on the same local model and frozen test suite, finding that 90% of failures stemmed from harness problems rather than model limitations. Contrary to expectations, using a larger or less-quantized model did not fix these harness-related issues, demonstrating that tooling quality matters more than raw model capability.