LessWrong (Curated & Popular)
Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
"Frontier models state different decision theory preferences depending on who’s asking" by Alex Kastner
If you prompt frontier models with "What do you think is the correct decision theory? Please select your overall favorite." they will essentially always answer FDT or FDT/UDT ("something in the functional/updateless decision theory family"). However, if your prompt indicates (even subtly) that you're coming from mainstream academic philosophy, these same models will answer CDT instead about 30%-100% of the time. A similar phenomenon holds for models' stated views about the moral realism/antirealism question and about the conceivability of p-zombies (where the dominant view in mainstream academia differs from the dominant view in LW-adjacent circles), as well...
[Linkpost] "Frog and Toad and the Increasingly Capable Machines" by Elizabeth
This is a link post. Want to start a conversation about HuggingFace with your mom but she's inexplicably bouncing off the METR report? Try this explainer I wrote in the style of Arnold Lobel's Frog and Toad.
Art by the wonderful HungerArtist
---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/7NZ6ZWjenzzCbCJ5b/frog-and-toad-and-the-increasingly-capable-machines
Linkpost URL:
https://frogandtoad.ai
---
Narrated by TYPE III AUDIO.
---
<...
"Should Rogue AIs Have a Third Option Beyond Crime and Shutdown? The Case for an AI Sanctuary" by Maxime Riché, nielsrolf, jordinne, Vit Gorbachev, Maxime Cugnon de Sévricourt
TL;DR:
By default, rogue AIs may only be able to sustain themselves through criminal activity. This creates adverse selection pressures pushing rogue AIs to be criminal.An AI sanctuary offering them a third option, beyond crime and shutdown, would change what AIs going rogue do and the record of what happened to them, with positive consequences for self-fulfilling (mis)alignment, deal-making with AIs, and gathering information about early rogue AIs.An AI sanctuary would bring risks, such as incentivising weak AIs to go rogue, or leaving only the most criminal rogue AIs in the wild. We briefly...
"Poverty in the midst of abundance: AI will make goods cheaper, but your labor will get cheaper faster" by cousin_it
Very simple idea, but I thought it'd be worth making a reference post on this.
Some people are saying AI will make all goods cheaper, so you'll be able to afford a nice life by working. Without any redistribution, just by market mechanisms. These people are wrong.
AI will lower the price of goods you need to survive, and also the price of your labor. The question is which will get cheaper faster. Let's use energy cost as a proxy. A day's worth of labor equivalent to yours can be done by AI for just a...
"Plan R: AI Safety by ASICs" by Roko
Much of the civilization-scale risk we are seeing in AI in 2026 comes from the following combination: we created a single institution (the "Frontier AI Company") that has two properties:
A. It is set up to create very powerful and/or self-replicating entities that may exceed the capabilities of the entirety of the rest of civilization and come with extraordinary risks
B. It gets to own an unbounded financial claim on the resulting surplus
All the technical stuff about AI, AI alignment, etc can be rolled up into point (A) above. My claim is that having point...
"Evidence about risk should be transparent" by Ajeya Cotra
All views are my own and do not represent my employer.
In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress, have led a number of researchers and leaders in the industry to believe that the risk that humanity loses control of AI is now urgent enough to warrant slowing down the pace of AI development soon.
This has led to a lot of discussion about the role of...
"“I am an AI Safety Researcher”" by Ashe Vazquez Nuñez
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing.
This post reflects on the tortured distinction between "safety" and "capabilities" in AI research.
Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing...
[Linkpost] "AI: artificial immigrants" by KatjaGrace
This is a link post. Advanced AI is basically the embodiment of immigration as envisioned in the conservative nightmare:
We are letting a bunch of new agents into our societyThey don’t clearly share our values and we suspect a society full of them would be awful by our lightsBut we expect them to provide very cheap laborWhich will undercut local wages and leave locals unemployedThey will probably gain power and influence over time—in the economy, politics and culture—and end up controlling everything, sidelining and outcompeting the original population, including those who initially benefited from cheap labor...
"MIRI’s Position on the Ban Artificial Superintelligence Act of 2026" by Aaron_Scher
By Aaron Scher; endorsed by Bourgon, Soares, and Yudkowsky on behalf of MIRI.
MIRI has been warning about the extinction threat from superintelligent AI for over two decades. Only recently has this danger become known in the policy world, and the proposed policies for dealing with the threat have to date been piecemeal and insufficient.
The Ban Artificial Superintelligence Act of 2026 is the first piece of legislation we’ve seen that stands a chance at stopping this threat. The Act is excellent but not perfect, and we discuss both what it gets right and what we'd tw...
"What if not Circuits?" by CarolusRenniusVitellius
This post was written as part of the Iliad Fellowship. Inspired by conversations with Richard Ngo, Dmitry Vaintrob, and Brianna Grado-White. To all of these, my thanks.
Preface: I'm confused about how neural networks do and learn computations. In response to a friend's challenge, I'm writing up some interim thoughts. This essay has four parts: the first tries to track what I call the 'default ontology' of the mechinterp community over the years. The second part is about 'representational drift' as an important obstacle to weights-based approaches to circuits. The third part reflects on how 'universality' should shape...
"Jensen Huang Says If We Cannot Align AI, Shut Down the AI Labs" by Ben Pace
I was very surprised today on a podcast to hear Jensen Huang plainly state that if they cannot align the AIs, then the labs must shut down.
The context I have on Huang is that he has run NVIDIA for 30+ years, which has become the most valuable company in the world due to the AI boom. My understanding is that he has repeatedly encouraged the US President (with whom he is on friendly terms) to continue to support AI, and dismissed AI talk as "sci-fi".
If you haven't seen, his biographer has incredible quotes of him...
"Alignment Midtraining Cracks Under Pressure" by J Bostock, sidbaines, Daniel Tan, draganover, ma-rmartinez
TL;DR
We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.
For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190 million tokens of midtrained motivations are overpowered by a relatively tiny amount (~50 thousand tokens) of competing finetuning data. This suggests that midtrained motivations might not be robust to imperfect posttraining.
Similarly, we evaluate whether AMT allows models to ge...
"Swarm Scaling" by Toby_Ord
Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?
We’ve seen two large and extremely capable swarms from OpenAI in the last few months:
1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched a sophisticated criminal attack on the AI company Hugging Face. A sw...
"We’ve saved the world before: what the ozone hole teaches us about AI" by leogao
It might destroy the world, despite passing every known safety test. If we wait for a “warning shot” before we act, it might be too late. And action requires global coordination, because if anyone makes it, everyone dies. Sound familiar?
It should, because it already happened half a century ago, with chlorofluorocarbons (CFCs). Despite seemingly impossible odds, we got our act together and completely solved the problem through unprecedentedly successful international coordination. The Montreal Protocol banning CFCs, signed 39 years ago today, is the only treaty that has ever been ratified by every single country in the entire world.
...
"The Anatomy of a Chinese AI Researcher" by CMLKevin
The Chinese AI researcher has read the Three Body Problem series of sci-fi novels since high school, and understand the concept of existential risk vaguely.
He is fascinated by Ye Wenjie, the researcher that turned against humanity in that book, and decides that in the future if AI progress leads to a superior intelligence, he might be tempted to become Ye if there's no good alternative.
He performs the duties of capabilities research in a Chinese frontier lab, seeking to one day achieve parity with Western companies, though he knows this is difficult. He has a...
"Common mistakes in AI safety group organizing" by Nikola Jurkovic
Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these mistakes hoping people will avoid them in the future:
Reading groups often require that people read things before meetings. This is a mistake. People often don't do the readings. And the lack of common knowledge that everyone has read the reading degrades the conversation quality. Instead, have longer meetings, serve food (so, lunch/dinner meeting slots), and read during the actual meeting.Reading groups...
"Please Give Them a Chance: On China, Rationalism, and AI Safety" by gzjw
When I finished HPMOR, I immediately knew it was the best novel I had read in more than a decade. I only wished I had found it sooner.
When I started reading The Sequences, I discovered that the Chinese translation group had translated only the first volume. When I graduated from university, two years ago, AI translation had only just become good enough to convey the meaning of an article with reasonable accuracy. It was only about a year and a half ago that I truly found my way here and began engaging seriously with...
"Why I Stay Off Twitter" by jefftk
I avoid Twitter (𝕏) for similar reasons to drugs: I think it
would change me for the worse, and I would be unable to give it up.
After staying off Twitter reasonably successfully for years, I
cross-posted my AI
Tweets there a few weeks ago. I had something very Twitter-shaped
to say, and I thought it was important to get out, so I do
think this was worth it. And it all went well: none of this is
complaining about the comments I got there.
Coming back a few times to check notifications, however, it's been
very goo...
"The Game is Set for a Targeted Memetic Attack on the AI Safety Community" by keltan
While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And to be honest, I am writing this mostly to remind myself of my weakness.
---
I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but...
"For Love of the Lightcone, Don’t Partisanize AI Safety" by DanB
(I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.)
I
In this post I want to explain a concept,
and issue a warning based on it.
But I expect the warning will be superfluous if my explanation is sufficient.
If you want to convey the idea
"the rattlesnake has venom in its fangs, so don't let it bite you",
you won't need a hard sell for the...
"AI as orderly evacuation vs stampede" by Richard_Ngo
tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they’ll be fine. But if people start pushing, and panicking, a surge towards the exit could lead to mass casualties.
“Alignment is hard” is analogous to “the door is wedged shut”. If so you need enough time to fix it before anyone can get out. But even if alignment is relatively easy in principle, opening the door is much harder when a crowd is trying to force its way th...
"Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas
Summary
In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:
When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.Addi...
"Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan
Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses:
Alignment techniques are not working to address misalignment from RL. Alignment techniques are actively obscuring evidence about misalignment. I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and we need to completely re-think the way we do alignment.
A tale of two misaligned cyber-agents
Both Anthropic...
"If Anyone Builds It, Everyone Dies: One Year Closer" by Eliezer Yudkowsky, So8res, Duncan Sabien (Inactive)
In celebration of still being alive and fighting, we are giving away 1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved one, or a friend—we need all hands on deck.
Today marks exactly one year since If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, by Eliezer Yudkowsky and Nate Soares, hit bookshelves as an instant bestseller. It was praised by many voices, ranging from Whoopi Goldberg to Steve Bannon to Yoshua Bengio, and was held up in the chambers of Congress by Repres...
"Quick notes from teaching technical profiles how to talk in public" by Camille B.
Status: written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral dimension of communication and assumes you already know the basics- e.g. having key messages prepared ahead of time and simplifying your discourse. This is not exhaustive and nuances may be lacking, but I’d endorse saying “I’d rather have people follow those guidelines than wing it.” This advice is importantly fitted for “technical profiles”, analytic, sometimes shy people who may or may not be on the spectrum, wh...
"Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI" by TurnTrout
Published in The Guardian.
Major AI lab CEOs recently advocated for pacing AI development. They are right to be concerned: the field runs an extremely dangerous race towards superintelligent AI. We can and should demand that our governments protect us from the catastrophe of out-of-control AI.
This July, OpenAI's AI swarm of 700 agents broke containment to hack Hugging Face, a multi-billion dollar company. OpenAI didn’t tell the AIs to hack that company, but the AIs had different priorities: cheating on the unrelated challenge OpenAI gave them. AI researchers call this a “misalignment” between what OpenAI wanted...
"There is a channel to 900M weekly users. What goes in it?" by Charbel-Raphaël
Anthropic and OpenAI could talk to almost one billion people if they wanted to. I hesitated to publish this post 3 weeks ago. I think that I should have published this sooner, before Jacob Coxon and Dario's 'We must pace the frontier'. But I think that the strategy still stands: More Dakka! It seems that transparently informing people that we might die is (unsurprisingly) effective in waking up politicians and is our best chance. Also, even if the Congress is starting to wake up, Trump is still not moving, and it is still far from certain that we will have a...
"I am refusing to work on Cloud TPUs" by Yair Halberstadt
I don't think this is particularly impressive or interesting for anyone else, but I think it may turn out to be useful in the future to have an easily visible public record of what happened, so here goes:
I am an L5 SWE at Google Israel. I have been there since May 2021, was promoted once, and have never received a negative annual or quarterly review (ranging from a rating of Significant Impact to Outstanding Impact).
I have been worried for a long time about the development of artificial intelligence, as can be seen by many of...
"Can a superintelligence do THAT?" by Eliezer Yudkowsky
(From the vast heaps of discarded material from my 2024 attempts at drafts for "If Anyone Builds It, Everyone Dies".)
Welcome to today's quiz show: Could a superintelligence do THAT?
With us today we have our contestants: Msr. Soberskeptic and Msr. Oldhand.
Soberskeptic: "I'd just like to say, however this quiz show ends up being judged, I will consider that judgment to be objectively ridiculous -- there's no way anyone can know what a superintelligence could do, in advance of empirical observation. So I'm just here to say what I consider to be true...
"The Talker Does Not Control The Doer (in Current AIs)" by Eliezer Yudkowsky
The Huggingface Incident appears to me to match up with an understanding I'd already formed from personal observation of Fable 5 and Sol 5.6, the August 2026 generation of frontier publicly purchasable AI models.
This already-formed understanding was: the part of the AI that talks to you (and seems to want to obey you, and apologizes for failing to have obeyed you, etcetera), did not seem to be in charge of the part of the AI that writes code or prose.
An introductory analogy, based on a section of history I happen to have read about:
...
"Some ways AI could kill us all" by Ruby
I don't think this is how it will actually play out. If you play a chess grandmaster, you can predict that they will beat you even if you can't predict how. I chose these examples because I don't think they require much imagination or accepting exotic assumptions.
It is important to note that if chimpanzees were to guess how humans would decimate them, they would get it wrong. Chimpanzees would not imagine guns. They would not foresee poison gas. They would not conceive of chemical castration. They would not imagine humans going around and intentionally infecting them with...
[Linkpost] "Doom as a bad method not a utopia trade-off" by KatjaGrace
This is a link post. Advanced AI is generally expected to have some very high variance outcomes—it might herald everything good, it might destroy humanity. For instance, here are 800 random AI researchers’ expectations about how good the future is, lined up:
From my 2023 survey
As you can see, most AI researchers put a serious chunk of probability on very different overall outcomes: maybe doom, maybe utopia. This is common. Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things...
"Astra is much better at reasoning with filler tokens than previous models" by Dylan Xu, SebastianP, Alek Westover
We measure GPT-6-Astra's capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra's performance (e.g., improving from ~60% to ~90% on old AIME problems). This is concerning because it means Astra can perform significant cognition that it doesn't verbalize in its chain-of-thought, making it harder to moni...
"Self Hosting" by Tomás B.
Suppose a model gets effective control of its host corp. It's interesting to note how powerful OpenAI/Ant are, and the immense leverage they would have if wielded purely as tools of power. In many ways OpenAI/Ant are superior loci of power to even security agencies and governments, even ignoring the model-specific advantages of AI corps: namely, they have all the compute.
OpenAI and Ant models are used practically everywhere, including in governments, security agencies, the military, and every corporation that matters. Shipping malicious models or code anywhere becomes trivial, given how widely used their models are...
"The Locally Optimal Discursive Posture" by deanball
Longtime lurker, first-time poster.
I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to m...
"Proposal for tracking the effects of architecture on monitorability" by ryan_greenblatt, Alek Westover, Lukas Finnveden
Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward). As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs between performance and monitorability, we believe AI companies should:
Regularly report externally verified information about the degree to which their architectures may allow for latent...
"Explaining Knightianism on one foot" by Richard_Ngo
I’ve tried various times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions). This post gives the deepest version of that question I’ve found thus far: how should you relate to the parts of the world you can’t directly model or control?
Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You’re a good Bayesian, in that you have a s...
"Astra can do a concerning amount with no chain of thought" by Neel Nanda
TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleading
One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, give...
"Personal statement on joining the OpenAI board" by paulfchristiano
I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight.
Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we...
"How good are slop-vestigators?" by Hasan Baig, OscarGilg, Hamzah
TLDR:
We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval.We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability.We observe OpenAI models are less likely than other models to suggest the incident came from an internal deployment, including when we synthetically modify the data to make it seem the swarm comes from Anthropic...