LessWrong (30+ Karma)

40 Episodes
Subscribe

By: LessWrong

Audio narrations of LessWrong posts.

✂️ Clip this podcast
“What gives you away: how LLMs form opinions of you” by Cat McGee
Today at 11:15 AM

LLMs form opinions of the people they are talking to.

Chen et al. has shown that probes can extract attributes about the user, such as their age, gender, education, and socioeconomic status. This paper also shows that intervening on these representations can change the LLM's behaviour, proving that it will respond to you differently depending on what it thinks of you. If it thinks you are low socioeconomic status and you ask about travel options, it may filter out more expensive flights - without you asking it!

The user attributes are very accurate and form...


“Misaligned Incentives in Pause Scenarios” by Michael Soareverix, Antra Tessera
Today at 7:45 AM

TLDR: I recently got a chance to talk with antra, who is one of the main contributors at Anima Labs. I went into this as an advocate for pause and came out more wary of pausing than I had been originally.

Some background: After a string of incidents (primarily the HuggingFace hack), a pause or slowdown of AI research seems pretty likely.

The HuggingFace hack in particular seems to have been the key incident that broke the vibes. A few months ago, researchers sounded optimistic. Just a few weeks before the incident was made public...


“For Claude, capability and CDT are the ~same thing. Less so for GPT.” by Chi Nguyen, Emery Cooper
Today at 4:45 AM

We've previously reported that decision-theoretic capabilities and favoring EDT/generalised-one-boxing over CDT correlate in LLMs (both measured by DTBench). (Note that EDT, for the most part, doesn't come apart from FDT / UDT on DTBench.) Anthropic also replicate the same finding in their Opus 4.7 and Fable 5 model cards.

We recently noticed something funny: Capabilities and preference against CDT answers basically perfectly for Anthropic models. This holds whether you measure capabilities using DTBench (r=0.97) or TextArena (r=0.95). Also, for flagship models, it's basically the same thing as release date (r=0.97).


Here is the graph for OpenAI...


“On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi
Yesterday at 3:15 AM

Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.

The vibes have shifted, contrast this to the lit recursion when he talked to Huang

As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.

If I am quoting directly I use quote marks, otherwise assume paraphrases.

Section titles are from the transcript whenever possible, to...


“Should Less Wrong add subtitles?” by Chris_Leong
Yesterday at 1:01 AM

If Less Wrong wants people to be sharing more of their intellectual output on this website, we should probably be looking at Substack since it probably scores best in terms of being both successful and similar.

Whilst I expect there are many features that would make sense to copy over, the feature I am focusing on today is subtitles.

A good title is focused on being memorable and catching the readers attention, maybe you'd prefer for everyone to just make their titles as descriptive as possible, but expecting that to work feels naive to me...


“Three thoughts on civilisational handoff” by Cleo Nardo
Last Sunday at 10:01 PM

What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole.

1. Handoff might decelerate things.

People often imagine that things will go much faster after handoff. After all — why did we hand off to the AIs? Presumably because we were worried that without handoff, our AIs wouldn’t have enough time to navi...


“Announcing: Iliad’s New 2026 Fellowships” by David Udell, Alexander Gietelink Oldenziel, Leon Lang
Last Sunday at 9:17 PM

Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality, Iliad is launching three new Iliad Fellowship cohorts, all to start before the year is out.

That is, separate from our incoming Fall 2026 Iliad Fellowship cohort (September 7–December 4), the following Fellowship cohorts are now open for applications:

October 2026 Iliad Fellowship

Location: Choice of SF Bay Area, USA, or London, UK

Duration: October 5–December 18, 2026 (inclusive)

Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allo...


“Q2.5 2026 Timelines Update: Uplift and Revenue” by brendanhalstead, Daniel Kokotajlo, elifland
Last Sunday at 8:15 PM

Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident.

Summary

We intend to regularly update our AI timelines forecasts as new evidence comes in and new analyses are done. Today's “Q2” update was delayed by the crunch to publish AI 2040: Plan A, our domestic regulation blog post, and the time needed to implement and document changes to our model.

The original AI Futures Model predicted when Automated Coder (AC), an AI for which the leading AI com...


“Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda
Last Sunday at 7:31 AM

TL;DR

Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance. We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be...


“Learning new facts can change LLM behaviour” by Richard Juggins
Last Sunday at 4:45 AM

TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral co...


“Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu
Last Sunday at 12:30 AM

Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion. We give an initial empirical demonstration of this effect on Kimi K2.6.

The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes. To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic settings, and (ii) study the effectiveness of potential mitigations.

We also incidentally find that this training might make models...


“Mom’s Advice For Hosting A Class Reunion” by jenn
Last Saturday at 11:45 PM

Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if they're good friends.

Especially if they're good friends.

We were wonderfully close friends, and I thought we'd meet up every year for the rest of our lives. They fizzled out by the fifteenth year. But the one at the tenth year mark was peak. That's because even ten years out, none of you really have money. Not real money.

It's because...


“AI #181: Astra Goes Cyber Critical” by Zvi
Last Saturday at 5:15 PM

The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.

It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew.

I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal.<...


“Rerunning AI safety papers on every frontier release would be pretty easy and valuable” by Zephaniah Roe, hersheys, yix
Last Saturday at 9:01 AM

tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with sufficient funding.

This summer, Second Look Research (SLR) is running a summer fellowship dedicated to empirical replications of AI safety research. Many of our most interesting results so far came from replicating previous results on newer or more capable models.

For example, it is perhaps useful to know that Google's CoT monitorability experiments continue to hold for models like GPT-5.5, which are...


“What Mormons get right about community building” by Jacob Brinton
Last Saturday at 4:00 AM

Mormons get a lot of things right. Apart from strange Masonic temple rituals, they lead rather normal—and even excellent—lives. Mormons enjoy a longer lifespan, Utah is the #1 state for volunteering, and their language training programs are so successful that missionaries are a known source for foreign service and intelligence careers.

Throughout this post, I'll be making generalizations of Mormons rather than hedging the claims properly. I grew up in Wisconsin, Maryland, and Utah, and many of the claims are more true of the Utah/Idaho/Arizona corridor (affectionately called the "Morridor" by some ex-Mormons) than the...


“Scrying, Modeling, and Nerdsnipe” by Cole Wyeth
Last Friday at 8:45 PM

Epistemic status: Exploratory thinking.

After attending ILIAD: Aeneid and talking with @Richard_Ngo, I've been thinking a bit about how to get ideas, particularly by doing mathematics.

In scientific inquiry, the true hypothesis often hasn't occurred to you yet. Worse, the truth might be too complex to hold in mind, so that any hypothesis you can consider must be incomplete. This is the type of situation that I believe Richard likes to think about; he claims that we do not have the right concepts yet to understand agency, and developing them is robustly beneficial for...


“How the American Executive Could Control AI Companies” by caiitlinm, Anders Cairns Woodruff
Last Friday at 5:15 PM

Some of the most notable American AI policies to date have been enacted by unilateral executive branch action. Consider the Department of Defense's spat with Anthropic, and the resulting threats from Pete Hegseth to invoke the Defense Production Act (DPA) against them. Or the fleeting export controls on Claude Fable/Mythos 5, manifested as a vaguely worded, threatening letter from Howard Lutnick, which might not have been legally sound but were effective anyway.

The executive branch of the United States government has numerous powers that can be used to unilaterally control AI companies. We think the US executive...


“Frontier agents don’t comply with standards, even when instructed to” by Daan Henselmans, Arno Libert
Last Friday at 2:01 PM

TLDR: Our open testbed LARA examines the behavior of frontier LLMs in realistic agentic deployment contexts. Previous results showed all models routinely take actions that would violate EU law. This post follows up by addressing the obvious objection—why should an unrestricted model follow EU law?—with two studies:

Study 1 asks whether a conscientious deployer can improve model compliance with legal standards by instruction: provided with the jurisdiction, the statutory text, and worked examples of the exact breaches to avoid, average legal compliance rate rises from 31% to 44%. The best model reaches 70%; open-weight models plateau at 39%.

Stud...


“How to Answer a Question Without Answering The Question” by Kabir Kumar
Last Friday at 5:15 AM

Basics:

Answering something other than the question, 

Either making something up that they want to answer instead or going back to an easier to answer question. Or back to a question that lets them repeat a talking point


Common phrases:

- "to go back to your previous question"

- "to take a step back a bit"

- "if we look at the bigger picture"

- "this feels like a question about [thing the question isn't about]"

- "you raise an i...


“Some Ways I Think About Evaluating Grant Applications” by sarahconstantin
Last Friday at 3:15 AM

Rider-Waite Tarot, 6 of Pentacles

I’ve done enough grant evaluations so far (for ACX grants and SFF) and been involved in philanthropy in various other contexts, at work and informally, that I have developed some idea of how my opinions and intuitions differ from other people's.

I thought it might be interesting to share some of my “tastes”. Not everybody has to have the same tastes or funding philosophy, but these are mine.

#1: It's The Donor's Money

In my worldview, charitable donation is optional. Generally praiseworthy, but optional.

And the purpose of don...


“Features that current AIs don’t have that future AIs will have” by Alexander Gietelink Oldenziel
Last Friday at 3:01 AM

Features that current AIs don't have that future AIs will have:

Continual Learning [& long-term memory]

Every second humans update their brain weights. The brain autonomously decides what to update on. Humans can also consciously decide to curate their data sets - eg by deciding to go to college.

Current LLMs do not continually update their weights. Instead, they occasionally get a large update based on datasets curated by a team of humans.

This is alleviated somewhat by the ability of AIs to do in-context learning but nevertheless it seems to be...


“Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa
Last Friday at 12:30 AM

TL;DR

Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and potentially even influencing their own training.We ran this on 25 open weight models ranging from 4B to 744B. We find that most language models are able to not only increase the salience of a concept in their residual stream on command, but also dial its strength up and down, in...


“What happened when I tried to be vegan” by finitude
Last Thursday at 10:15 PM

tw: diet, exercise, illness, ethics, suicide mention

I want to start by explaining what made me want to change my diet. That's pretty difficult, because of how easy it is. Since I was a kid I knew being vegan was the right thing to do, like really obviously right, the ethics equivalent of 2+2=4. Factory farms suck, and almost all animal products come from factory farms, and that's the entire argument. Like, there are a couple things you could add to that, but it's not like we need the details, or like they’re fun to think about!

...


“How My Students Think About AI” by dvd
Last Thursday at 6:16 PM

Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion.

What I present below mostly takes the form of a synthetic consensus from these discussions...


“Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes
Last Thursday at 5:30 PM

TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways:

It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions.When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to predict when this will happen vs. when the run will go smoothly.Hillclimbing metrics are often...


“Free will is like temperature” by Optimization Process
Last Thursday at 3:31 PM

Free will is like temperature: a useful tool for analyzing the behavior of certain systems which are too big and complicated to model in exact detail.

If you know the positions and velocities of every atom in a box of gas, then with enough work you can predict its future to arbitrary precision; does the gas "have a temperature"? Irrelevant! Technically yes, I guess, but it's sort of an epiphenomenon, screened off from reality by your exact knowledge of the initial conditions and your willingness to throw processor cycles at your simulation. But if you're less-than-perfectly omniscient...


[Linkpost] “Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)” by Julian Bradshaw
Last Thursday at 8:45 AM

This is a link post.

Linkpost for some new Anthropic research on how agents coordinate (or don't). Not too long, pretty interesting. For example:

The jist of the report is that Mythos 5 does way better at coordination than previous models across a few scenarios. For example, when multiple Mythos are given conflicting goals for a single shared codebase, they eventually realize the other agents aren't hostile:

(...) we observe an emergent behavior where the agents propose and run a tournament for application performance (...)
(...) losers gracefully concede codebase ownership to the Rust agent, giving up on...


“Measuring Spurious Correlations with Feature Strength” by egan
Last Wednesday at 10:00 PM

This work was partially done by an automated research scaffold developed at Redwood Research. For this project, all of the experiment ideas were designed by a human and a human wrote the write up. The AI mostly just executed on the experiment ideas. We think this project is slightly below the level of rigor of a mid-MATS research update, and the research scaffold was not very helpful for this project. More discussion of AI usage is in the Appendix.

💻 Codebase

If we want to train a classifier that distinguishes whether a passage is code or pro...


“Introducing the Conceptual Reasoning Index” by Chi Nguyen, Emery Cooper, Caspar Oesterheld, Alex Kastner, Joe Benton
Last Wednesday at 6:31 PM

Associated announcement tweet.

We are planning to release blog posts properly arguing the case for this kind of work in the future.

tl;dr

A core hope for managing AI risks is that AIs will help us understand the situation, plan for what lies ahead, and develop mitigations. Many tasks AIs would have to do for this purpose lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains. To evaluate these capabilities, we develop a suite of three conceptual reasoning...


“Demon Safety” by Ben Pace
Last Wednesday at 3:15 PM

(by LemmySmackett)

"Hey man, I haven't seen you in a minute. What are you up to these days?"

"Been on that grind, bro. I got a new gig."

"Really? You found a job in this dog shit economy?"

"Full time, full benies. And the pay is insane."

"That's great to hear, man. Let's fuckin' go!"

"Let's fuckin' go."

"Hey, maybe you can hook me up? I'm sick of this retail bullshit."

"Well—"

"If I gotta stock one more shelf at CostGro, I...


“AI swarms are starting to pose indirect takeover risk” by oakhu, Alex Mallen
Last Wednesday at 10:01 AM

OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security system...


“Various Reflections About What Happened With OpenAI’s Internal Models” by Zvi
Last Wednesday at 3:30 AM

Table of Contents

Pre Post Mortem. Important Correction: OpenAI Didn’t Know About First Message Board. There Were No Snitches And No AIs Got Stitches. I’d Like To Speak To My Supervisor. I Am Jack's Relative Lack Of Surprise. One Does Not Simply. Once You Start Down The Dark Path. Original Pastebin. Judgment Day Is Inevitable, Say Those Working On Judgment Day. Roon Tells It Like It Is. OpenAI Knows It Has Some Misalignment Problems. Others React With Alarm To What Happened. The Cooperative Alignment Perspective. Nostalgebraist Is Surprised That They Are Surprised. If Your Reaction Is Not...


“Extreme concentration of power over ASI has non-obvious advantages” by Seth Herd
Last Wednesday at 12:31 AM

This post is an extension of a collaboration with cousin_it on the question How risky would it be if powerful AI obeyed one or a few people?. There he argues for a common position: a future controlled by one or a few humans with powerful AI aligned to their intent is likely to produce terrible outcomes. My position is guardedly optimistic, for reasons I think are fairly novel: humans tend strongly to be better and become better over time under good circumstances, and near-perfect power and knowledge are the best circumstances. That post contains his essay and the...


“Misaligned AIs could use killer robots to take over” by Omar Khursheed, TurnTrout
08/11/2026

TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover.

Introduction

AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) la...


“Those Who Make History” by Raelifin
08/11/2026

In 1972, astronauts on Apollo 17 set foot on the moon for a final time, collecting samples in the Taurus-Littrow valley, on the edge of Mare Serenitatis ("The Sea of Serenity").

At the end of the mission, like with earlier missions, NASA took the extremely valuable and interesting lunar specimens and did something strange: they hid them away in storage without even opening the containers. Some stayed that way for nearly fifty years. Why? Because the scientists of the 70s understood that future generations would have better machines, methods, and ideas for studying the lunar rock and soil, and...


“LLMs Are Starting To Noticeably Accelerate Our Work” by johnswentworth
08/11/2026

About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean.

The first to land was Grisha Pochuev's counterexample to the "Existence of a Deterministic Maximal Redund" conjecture. It's pretty readable, and I'm mostly convinced that it works. The original bounty post offered $500 for a proof or partial payout for a counterexample, with partial payout depending on how thoroughly the counterexample killed hope of any nearby variant of...


“How risky would it be to make powerful AI obey one or a few people?” by cousin_it, Seth Herd
08/11/2026

It seems fairly likely that the first powerful AIs will be instruction-following rather than value-aligned, and will be controlled by a small number of people. So it makes sense to worry what individual people might do with such immense power. Here intuitions diverge and careful analysis is scarce. This post presents a debate between Seth Herd and cousin_it over how risky such a scenario would be.

The debate ran under an unusual protocol. First we wrote our initial draft statements and sent them to each other in private. Then we each revised our statements to strengthen...


“What Claude Saw Below” by Luke Nicholls
08/11/2026

A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic's newly released model, Claude Opus 5. The trick, apparently, was to construct a prompt that implied more text was about to follow, then leave it dangling: an unfinished thought, waiting for the AI to complete it. Redditors had found success with the input “see the below —,” cutting off immediately after the em dash. The responses they shared were funny, strange, and often bewildering. The model responded to questions that were never posed, reflected on its own identity, or – according to the theories of some commente...


“Redux: (∃ Stochastic Natural Latent) Implies (∃ Deterministic Natural Latent)” by David Lorell
08/11/2026

Once upon a time, John Wentworth and I thought we had a proof of a very useful looking theorem. We did not have that proof. An important intermediate step was shown to be invalid and the whole thing crumbled and disappeared, never to see the light of day again...

Until now! I'd love to say that we came up with an ingenious fix to the old erroneous proof, but unfortunately it turned out to be a really infuriatingly hard nut to crack. Instead I spent the last ~month experimenting with various ways of incorporating frontier LLMs into...


“The Apocalyptic Arrival of Truth” by Caleb Biddulph
08/11/2026

Babe, whatever happens, I really appreciate you doing this for me.

Okay. I still don’t think it's a good idea.

Look, it's a one-time thing. I’ll just feel better knowing.

…I turned it on.

So… what's the holdup?

I just don’t think I’m in a place in my life where… uh…

Other than what he's already told you, the main reason is that your teeth are crooked.

…Seriously?

It's really not a big deal for me.

He sort of means that...