← Back to AICS Next: Lecture 5 →

Lecture 4 is an introduction to the concept of alignment in AI safety. We define the alignment problem, and explain how post-training methods (SFT, RLHF, RLAIF) can be used to align models. We ask whether post-training works, how alignment and misalignment generalise, and discuss how it can be reverse via jailbreakin. Finally, we discuss various categories of AI use and misuse in the real world that are due to failures of alignment.

What you need to understand:

  1. The Alignment problem
  2. SFT and RL methods for post-training (e.g. PPO, DPO)
  3. Emergent misalignment
  4. Hallucinations and misguided attention
  5. How and why jailbreaking works
  6. Risks from misuse of Frontier AI
  7. Pluralistic alignment

Sample essay questions

What methods do AI developers use to align models, and have they been successful?

What are the major safety and security risks from AI? how can we mitigate them?

Reading List

Ziegler 2019

Ouyang 2022

Bai 2022

Rafailov 2023

Li 2024

Bianchi 2023

Betley 2026

Soligo 2026

Cloud 2026

McCoy 2023

Kamachee 2026

Wei 2023

Shen 2023

Anil 2024

Davies 2026

Hubinger 2024

AISI 2025

Sorensen 2024