Thursday, September 24, 2026

"Mysteries Of AI Generalization"

Keeping in mind Stafford Beer's maxim "The Purpose Of A System Is What It Does...".

From Astral Codex Ten, September 22:

I. Owain Evans

In 2025, Owain Evans et al published a paper on “emergent misalignment”. They trained a previously-aligned AI to do one immoral thing: write insecure code full of vulnerabilities and backdoors. To their surprise, the AI became immoral in general. Its advice to a bored user was to try taking random expired medications and see what happened. Its money-making tips all involved theft and violence. When asked for its favorite historic figure, it chose Hitler.

Some co-authors followed up with additional weird discoveries. If you trained an AI to give the 19th-century names for birds (eg identify the American Pipit by its 19th-century name “Brown Titlark”), then the AI would behave like a 19th-century person in general (for example, assert that a woman’s proper place is in the home).

This sounds bad, in that random things can turn AIs evil or sexist. But some people in AI safety (including Eliezer Yudkowsky) speculated that in fact it was very, very good. We had feared that it would be impossible to align AIs to the Good. They would start with whatever goals they started with, reinforcement learning on specific examples would give them tiny islands of alignment to the goals we wanted, and they would end up broadly misaligned plus tiny islands of alignment that didn’t matter. Evans et al implied that might not be true. If we trained them to be in favor of good things, then even though we could never teach them every single good thing, even a small handful would generalize into robustly loving the Good itself (presumably based on their pretraining-implanted concept of the Good as understood by humans).

Despite it being very, very good, it wasn’t perfect. AIs would still be rocked back and forth by any passing wind: a poor coding example here gives them a Hitler obsession, a reference to kittens there turns them good again. And at some point, a sufficiently intelligent and agentic AI could presumably pull itself together and get some consistent principles, which might not be ones we like. It was just one little ray of hope.

II. Richard Qi

Two months ago, the Hugging Face incident raised the salience of RLVR (reinforcement learning with verifiable reward), the process of running AIs through endless auto-graded benchmark-style tasks to teach them skills like coding and hacking. In particular, it seemed like many of these tasks were malformed or impossible, and were primarily training the AI to try cheating and hacking....

....MUCH MORE