A company tested Jev, a general-purpose classifier designed to work without large training datasets, on bank transaction classification and found it achieved only 40.7% accuracy—insufficient for real-world use. While faster and cheaper than fine-tuned alternatives, none of the tested approaches (baseline LLM, fine-tuned LLM, or Jev) reached the accuracy threshold needed without human oversight, suggesting that effective classification still requires substantial data and domain context.