AIInterviewTraining logoAIInterview/Training
📊 Evaluation & ML Foundations
Core

Benchmarks and Their Limits

Public benchmarks like MMLU offer a shared yardstick, but they saturate, leak into training corpora, and stop tracking real ability once labs optimize for them. Contamination (test items in the training data) and Goodhart's law (a measure that becomes a target stops measuring) are why a high leaderboard score can mean nothing on your workload. AI, ML, and GenAI engineer interviews probe this to see whether you trust a number or build a private eval set on your own distribution.

a free account unlocks the core curriculum tier · no card
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
COMPANIES THAT ASSUME THIS
NEXT IN EVALUATION & ML FOUNDATIONSContrastive and Metric Learning