Semloop is a semantic loop detection tool for AI agent traces that uses an LLM judge to identify whether agents are making progress, improving upon deepeval's lexical similarity approach. It supports multiple backends including TypeSafe's Jev model and OpenAI-compatible APIs, achieving 90% accuracy on a benchmark dataset compared to 40% for the lexical baseline.
A benchmark test evaluated 16 AI models and web extraction services on their ability to correctly return null for missing fields rather than inventing data. Adding the instruction 'Do not guess' reduced hallucinated fields from 70.7% to 20.2%, with Firecrawl performing worst and plain fetch plus GPT-6 Luna performing best at low cost. A cheap verification method using GPT-6 Luna caught 20 of 24 hallucinated values for under a cent.