Researchers stress-tested LangGraph and autonomous AI agents using co-evolutionary fuzzing and quality-diversity algorithms, discovering that larger models (24B, 33B) were more vulnerable to prompt injection attacks than smaller ones, while reasoning models like DeepSeek-R1 showed superior safety. The LIFE FORGE framework simulates enterprise environments with adversarial conditions across multiple dimensions to map agent vulnerabilities.