← 🧠 Foundations of LLMs & GenAI
Core
Alignment: Outer, Inner, and Scalable Oversight
Alignment is getting a system to pursue what we actually want rather than what we literally specified. Outer alignment asks whether the objective is right, and it fails as reward hacking, overoptimization, and sycophancy. Inner alignment asks whether the model internalized the goal or something merely correlated with it. Scalable oversight asks how humans supervise models they can no longer evaluate. AI, ML, and GenAI interviews probe this because RLHF, DPO, and Constitutional AI are mechanisms, and the frame beneath them is what tells you when they will fail.
a free account unlocks the core curriculum tier · no card
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
LLM & GenAI FundamentalsWhat is AI alignment, and how is it different from bolting on a safety filter?→LLM & GenAI FundamentalsWhat is Constitutional AI / RLAIF, and how does it differ from RLHF?→LLM & GenAI FundamentalsWalk through RLHF, then explain DPO and why it has largely displaced PPO-based RLHF.→LLM & GenAI FundamentalsExplain PPO and GRPO for LLM alignment. Why did GRPO drop the value model?→LLM & GenAI FundamentalsWhat is instruction tuning, and how does it differ from pretraining and alignment?→LLM & GenAI FundamentalsBeyond DPO: what are SimPO, KTO, and ORPO, and why do these alignment variants exist?→
COMPANIES THAT ASSUME THIS
