Lecture 7 - AI and human agency
In lecture 7 we consider the balance of power between humans and AI systems. We begin by considering the nature of human agency, and evidence that humans seek to maintain their freedom and control over the environment. We consider how AI may gradually erode human agency, as technology is embedded in critical systems and infrastructure. We explore how power can be exercised through technical systems (e.g. a bureaucracy) and by exercising influence. AI is now superhuman at persuasion, and exceptionally cyber capable. We consider the concept of instrumental convergence, evidence for AI “scheming” and assess the possibility of loss of control to AI.
What you need to understand:
- Empowerment
- Instrumental convergence
- Reward hacking / Specification gaming
- Rhetorical strategies for attitudinal vs. action persuasion
- AI scheming, e.g. sandbagging
- Model intentionality
Sample essay questions
What are the different ways that AI could disempower humans, either gradually or acutely?
What is the evidence for strong AI persuasion, and how could it impact society?