Want to dive into Scott Alexander's work and his thousands of blog posts? This fan website lets you sort and do semantic search through the whole codex. Enjoy!
| Date | ||
| Length | ||
| Likes | ||
| Comments |
Scott explains current mechanistic interpretability techniques for understanding AI cognition, from linear probes to emotion vectors, showing they're useful for monitoring but insufficient for controlling AI behavior or ensuring safety.
Scott Alexander explains mesa-optimizers in AI alignment, their potential risks, and the challenges of creating truly aligned AI systems.