Want to dive into Scott Alexander's work and his thousands of blog posts? This fan website lets you sort and do semantic search through the whole codex. Enjoy!
| Date | ||
| Length | ||
| Likes | ||
| Comments |
Scott discusses three recent papers showing surprising patterns in how AI misbehavior does and doesn't generalize across different contexts, from emergent misalignment spreading across domains to reward-hacking staying confined to graded tasks.