← Back to AICS Next: Lecture 3 →

This lecture unpacks how key-value memory works, how it is implemented in the transformer, and how it relates to fast associative memory in the brain. It explains how large neural networks allow abstractions to be formed, and discusses the distinction between in-weight and in-context learning, and how the latter may explain sample efficient learning in humans.

What you need to understand:

  1. Key-value memory and the transformer network
  2. What it means for language to be “grounded”
  3. Abstraction in neural representations (in LLMs or brains)
  4. The distinction between in-weight and in-context learning; sample-efficient learning
  5. Structure learning and neural scaffolds in brains and machines

Sample essay questions:

What are the computational principles that underpin modern AI’s success, and how do they mirror those of biological brains?

Discuss the difference between in-context and in-weight learning in transformer networks. Is this distinction useful for understanding natural intelligence?

Reading List

Summerfield & Stachenfeld 2026

Lake 2016

Hassabis 2017

Silver 2016

Bubeck 2023

Teyler DiScenna 1986

Brown 2020

Abramson 2020

Bender & Koller 2020

Piantadosi & Hill 2022

Goh 2021

Templeton 2024

Chan 2022

Pesnot-Lerousseau 2025

Gershman 2024

Whittington 2020

Huh 2024

Bosch 2026

Gurnee 2026