AIInterviewTraining logoAIInterview/Training
🧠 Foundations of LLMs & GenAI
Core

Alignment: Outer, Inner, and Scalable Oversight

Alignment is getting a system to pursue what we actually want rather than what we literally specified. Outer alignment asks whether the objective is right, and it fails as reward hacking, overoptimization, and sycophancy. Inner alignment asks whether the model internalized the goal or something merely correlated with it. Scalable oversight asks how humans supervise models they can no longer evaluate. AI, ML, and GenAI interviews probe this because RLHF, DPO, and Constitutional AI are mechanisms, and the frame beneath them is what tells you when they will fail.

a free account unlocks the core curriculum tier · no card
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
COMPANIES THAT ASSUME THIS