Want to dive into Scott Alexander's work and his thousands of blog posts? This fan website lets you sort and do semantic search through the whole codex. Enjoy!
| Date | ||
| Length | ||
| Likes | ||
| Comments |
A technical dialogue between Scott Alexander and Eliezer Yudkowsky exploring why AI alignment is difficult, covering the inapplicability of human moral development analogies, the nature of AI consequentialism, and skepticism about acausal trade as a safety mechanism.
Scott argues that mental states like beliefs and preferences may not exist as unified entities in the brain, but rather emerge from collections of behaviors and contextual responses.
Scott uses a thought experiment about a robot that shoots blue things to argue that we mistakenly interpret simple programmed behaviors as goal-directed utility maximization, when really the robot is just executing code, not minimizing blue.
Scott discusses the neuroscientific split between 'wanting' and 'liking', showing how wireheads might not actually be happy, and explores the troubling implications this has for economics, philosophy, and AI alignment.