Lecture 4 - AI Alignment
| ← Back to AICS | Next: Lecture 5 → |
Lecture 4 is an introduction to the concept of alignment in AI safety. We define the alignment problem, and explain how post-training methods (SFT, RLHF, RLAIF) can be used to align models. We ask whether post-training works, how alignment and misalignment generalise, and discuss how it can be reverse via jailbreakin. Finally, we discuss various categories of AI use and misuse in the real world that are due to failures of alignment.
What you need to understand:
- The Alignment problem
- SFT and RL methods for post-training (e.g. PPO, DPO)
- Emergent misalignment
- Hallucinations and misguided attention
- How and why jailbreaking works
- Risks from misuse of Frontier AI
- Pluralistic alignment
Sample essay questions
What methods do AI developers use to align models, and have they been successful?
What are the major safety and security risks from AI? how can we mitigate them?