Researchers tested Jev, an efficient classifier model, against Claude Sonnet 5 and open-weight alternatives for security classification of 100 real tool calls from production Claude Code sessions. While Jev offers speed and confidence scores, it struggles with nuanced security properties like distinguishing between output data characteristics (delta) and prerequisites for safe execution (requires), with baseline accuracy concerns when 79% of calls are inherently benign.