67Explain the basics of reinforcement learning (and how it differs from supervised learning).▼mediumGoogleOpenAIMeta1 replies◆ premiumRL sits under RLHF, robotics, and recommendation, and interviewers want the core framing. What matters is the agent-environment-reward loop and the three things that make it harder than supervised learning. Here is the answer.Open full answer →
117Your moderation model flags normal speech in other markets. How do you moderate across cultures?▼hardMetaGoogleTikTok◆ premiumA classifier trained on one culture's annotations does not generalize to another culture's speech, and the aggregate metric hides it. The signal is recognizing that the ground truth itself is the bug, then designing the policy and eval stack around that.Open full answer →