A head-to-head benchmark compares Laya and Jev language models on 751 identical test cases across 9 suites, finding Jev outperforms on multi-class and non-English tasks (intent 0.975 vs 0.725, toxic 1.000 vs 0.767) while Laya wins on agnews and mnli with zero cost and lower latency (180–660 ms vs 925–1068 ms). Emotion classification is weak on both models near 0.55 accuracy; gating at 0.85 confidence keeps 58% of Laya traffic at 0.878 accuracy and 78% of Jev at 0.917.
A new approach uses personal computer use data to train local LLMs that predict user judgment and writing patterns, reducing the effort required to prompt AI agents. In a two-week study, a specialized model achieved 17.1% semantic accuracy on next-write predictions at $0.3 per call, with a continually trained version reaching 3.0% accuracy at $0.01 per call, suggesting potential for scaling.