SM
Sheby M
Loading portfolio experience
Loading portfolio experience
Business Operations · AI Agent, Dashboard, MVP
A benchmarking system comparing GPT and open-source models for resume parsing and candidate scoring across quality, speed, and cost.
The recruitment platform needed to choose between AI models without relying on small subjective tests or vendor claims.
I created a human-labeled benchmark dataset and pipeline that runs each model against the same schemas and scoring criteria, then compares accuracy, consistency, latency, failure rate, and cost.