Want to dive into Scott Alexander's work and his thousands of blog posts? This fan website lets you sort and do semantic search through the whole codex. Enjoy!
| Date | ||
| Length | ||
| Likes | ||
| Comments |
Scott explains current mechanistic interpretability techniques for understanding AI cognition, from linear probes to emotion vectors, showing they're useful for monitoring but insufficient for controlling AI behavior or ensuring safety.
A review of the synaptic plasticity and memory hypothesis arguing that it is insufficient to explain memory, and proposing a broader cellular processes and memory hypothesis that includes molecular and intracellular mechanisms.
Scott investigates the correlation between intelligence and neuron count, exploring various theories before suggesting that more neurons allow for less polysemantic (overlapping) representations of concepts.
Scott Alexander discusses recent breakthroughs in AI interpretability, explaining how researchers are beginning to understand the internal workings of neural networks.
Scott Alexander examines the 'canalization' theory in computational psychiatry and its refinement through deep learning concepts in the Deep CANAL model.
Scott Alexander examines why skills plateau, proposing decay and interference hypotheses to explain the phenomenon across various fields.
Scott Alexander discusses recent research unifying predictive coding in the brain with backpropagation in machine learning, exploring its implications for AI and neuroscience.
Scott Alexander explores the similarities between Wernicke's aphasia and GPT-3's language use, while noting that GPT-3's capabilities likely surpass this neurological comparison.
Scott Alexander examines GPT-3's capabilities, improvements over GPT-2, and potential implications for AI development through scaling.
Scott Alexander explores GPT-2's unexpected capabilities and argues that it demonstrates the potential for AI to develop abilities beyond its explicit programming, challenging skepticism about AGI.
Scott Alexander examines how recent AI progress in neural networks might challenge the Bostromian paradigm of AI risk, exploring potential implications for AI goal alignment and motivation systems.
Scott Alexander humorously explores the World Cup's complex rules, game theory in soccer, and unusual incentive structures in international tournaments.