Want to dive into Scott Alexander's work and his thousands of blog posts? This fan website lets you sort and do semantic search through the whole codex. Enjoy!

See also Top Posts and All Tags.

Tag: transformers

Minutes:
Pick a custom range (minutes). Leave a field empty for no limit.
–
Blog:
Year:
2026
2025
2024
2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
Tags:
Filter by tag...
Exclude tag...
5366 tags
Links:
Filter by linked site (twitter, substack…)
2 posts found
Compact Mode
Save Reads
Sep 24, 2026
acx
Read on
27 min • 4,121 words • Comments pending
Scott analyzes OpenAI's 'neuralese recurrence' technology that lets AI think in internal representations between processing steps, explaining the safety implications and arguing for clear taboos on recurrent architectures. Longer summary
Scott examines OpenAI's development of 'neuralese recurrence' in their Astra model, a technology that allows AI to think using internal representations rather than human-readable text between processing steps. He explains how transformers work with layers and chain-of-thought reasoning, why looping layers can create unmonitored thinking, and debates whether OpenAI's implementation crosses a dangerous threshold. The post concludes by discussing the need for clear taboos around recurrent AI architectures, drawing on Linchuan Zhang's argument that categorical boundaries work better than fuzzy thresholds for maintaining safety norms. Shorter summary
May 22, 2026
acx
Read on
6 min • 875 words • 415 comments • 314 likes • podcast (6 min)
Scott argues that even if AGI requires a new paradigm beyond LLMs, we shouldn't expect significant delays, since Lindy's Law suggests major paradigm shifts could occur within 3-5 years, and new paradigms typically emerge precisely when scaling hits limits. Longer summary
Scott addresses the objection that AGI is far off because LLMs need a 'new paradigm' to reach AGI. He traces the evolutionary tree of AI development from neural networks through transformers to modern LLMs, then applies Lindy's Law to show that even paradigm shifts as major as deep learning or transformers should be expected within 3-5 years at the 25th percentile. He argues this timeline is comparable to LLM-only predictions anyway. Scott also makes a subtler point: new paradigms historically emerge when old ones hit scaling limits, meaning they won't cause delays but rather continue progress from where scaling left off. He concludes that extrapolating from current LLM scaling remains the best forecasting method whether or not LLMs themselves reach AGI. Shorter summary
Per page:
Showing 1 to 2 of 2 results
Get these search results in an EPUB

Your filters match 2 posts.

Posts to include
–
Leave empty to keep the defaults. Range cannot exceed 500 posts.
Download now

Generates an EPUB right now and downloads it to your device.

Send to email

Generates an EPUB in the background and emails you a temporary download link.

Your email is not shared with anyone.

Email address

To send to your Kindle, just use this link.