Jev reportedly performs near Claude Sonnet on classification at far lower cost
A user testing TypeSafe AI’s small classifier reports promising results, while flagging possible training-data overlap in public benchmarks and a lack of real B2B sales data.
TLDR
A user evaluating TypeSafe AI’s Jev reports classification performance roughly on Claude Sonnet’s level while costing a couple of orders of magnitude less to run. The user says Jev performed on par with Opus on three public benchmarks, which may be in its training data, and alongside Sonnet and Opus on a fresh, hand-written test based on one Wikipedia article. On older customer-service and negotiation datasets, the user puts Jev a bit above Claude Haiku and slightly below Sonnet 5. Those datasets weren’t from real B2B sales, so the user calls the findings directional rather than representative of work at Lightfield.