26What are MMLU, HumanEval, and GSM8K, and how do you interpret LLM benchmark scores?▼mediumOpenAIGoogleAnthropic1 replies◆ premiumPlenty of people cite benchmark numbers; almost nobody reads them carefully. What shows depth is explaining what each one measures and why leaderboard figures outrun real-world skill. The question the interviewer is saving is how you would spot contamination.Open full answer →
72How do you evaluate generative output quality (text and images) when there's no single correct answer?▼hardOpenAIBlack Forest LabsGoogle DeepMind1 replies◆ premiumFor open-ended generation there's no ground-truth string to match, so accuracy is meaningless. The field relies on a layered mix of automatic, model-based, and human metrics. Here is how to assemble a credible eval.Open full answer →
122How do you run a human evaluation you can actually trust?▼mediumOpenAIAnthropicScale AI◆ premiumMost teams run human eval as a vibe check with ten examples and a 1-5 slider. The signal is treating it as a designed experiment: pairwise, blinded, powered, and agreement-measured. Here is the protocol that survives scrutiny.Open full answer →