Uber describes its feature logging framework for ensuring ML model consistency between training and inference. The system addresses data formatting mismatches, ETL fragility, and freshness gaps that cause model performance regressions, while tackling scale challenges from 8 million QPS prediction traffic.
A system-one model called Jev significantly improves entity resolution in high-throughput data pipelines, reducing costs by 99.56% and increasing throughput 7.35× while maintaining near-baseline accuracy when integrated into multi-stage review workflows. The authors demonstrate this on Ohio campaign finance data, where Jev resolves entities across donors, committees, and companies by adjudicating bundles of equivalent evidence.