Loading portfolio experience
Loading portfolio experience
Business Operations, Education · AI Agent, Dashboard, MVP
An AI benchmarking platform that evaluates OpenAI GPT and LLaMA models for resume parsing and candidate scoring across accuracy, consistency, speed, and operating cost.
The recruitment platform needed to select an AI model for extracting education, experience, skills, employment history, certifications, and contact information from resumes. Different models produced different output quality, candidate scores, response times, and costs, making it difficult to choose a model based only on small manual tests.
I developed a controlled benchmarking pipeline for testing OpenAI GPT and LLaMA models against a human-labeled resume dataset. Every model receives the same resumes, extraction schema, candidate criteria, and scoring instructions. The system compares extracted fields against verified reference data and measures accuracy, precision, recall, schema compliance, latency, token usage, cost, and scoring consistency. A dashboard presents model-level and field-level results and generates a recommendation based on configurable business priorities.