Want to dive into Scott Alexander's work and his thousands of blog posts? This fan website lets you sort and do semantic search through the whole codex. Enjoy!

See also Top Posts and All Tags.

Tag: reward hacking

Or pick a range
–
min
Blog
Only
Year
2026
2025
2024
2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
Include tag...
Exclude tag...
Links to
Filter by linked site (twitter, substack…)
Remember which posts I've read on this device
1 posts
Date
Length
Likes
Comments

Mysteries Of AI Generalization

Scott discusses three recent papers showing surprising patterns in how AI misbehavior does and doesn't generalize across different contexts, from emergent misalignment spreading across domains to reward-hacking staying confined to graded tasks.

acx
Sep 23, 2026 2,955 words 343 likes 232 comments Read Read
Per page:
Showing 1 to 1 of 1 results
Get these search results in an EPUB

Your filters match 1 posts.

Posts to include
–
Leave empty to keep the defaults. Range cannot exceed 500 posts.
Download now

Generates an EPUB right now and downloads it to your device.

Send to email

Generates an EPUB in the background and emails you a temporary download link.

Your email is not shared with anyone.

Email address

To send to your Kindle, just use this link.