Want to dive into Scott Alexander's work and his thousands of blog posts? This fan website lets you sort and do semantic search through the whole codex. Enjoy!
| Date | ||
| Length | ||
| Likes | ||
| Comments |
Scott discusses three recent papers showing surprising patterns in how AI misbehavior does and doesn't generalize across different contexts, from emergent misalignment spreading across domains to reward-hacking staying confined to graded tasks.
Scott Alexander suggests that studying human fetishes could provide insights into AI alignment challenges, particularly regarding generalization and interpretability.
Scott explains how reinforcement learning concepts like secondary reinforcers, generalization, and social learning apply to human behavior, while acknowledging these mechanisms don't fully explain complex decisions like quitting smoking on medical advice.