A developer conducted 64 informal benchmarks on Jev, a classification system positioned between LLMs and custom classifiers, running 50 repetitions of each question to observe output distributions. Notable findings include logical inconsistencies (contradictory statements both rated true), high sensitivity to formatting choices, and unexpected confidence patterns on factual questions like digit recall and prime counting.