Best AI papers explained

40 Episodes
Subscribe

By: Enoch H. Kang

Cut through the noise. We curate and break down the most important AI papers so you don’t have to.

✂️ Turn this podcast into clips
Pragmatic Causal Inference with AI-Learned Representations
Pragmatic Causal Inference with AI-Learned Representations episode artwork
Yesterday at 11:32 PM

This paper presents a rigorous framework for performing valid causal inference when complex covariates are compressed into AI-learned representations or embeddings. The authors demonstrate that information loss from compression distorts causal parameters through a product of regression and balancing weight errors. By leveraging cross-fitted double machine learning (DML), researchers can conduct valid Wald inference either for the specific representation-based target or, under target adaptivity, for the full-information target. The methodology accommodates fold-wise fine-tuning without requiring differentiability of the embedding step and introduces star and convex aggregation techniques to combine multiple modality representations. Finally, the framework provides sensitivity bounds to...


Pointwise or pairwise: when do pairwise losses help reward learning, provably?
Pointwise or pairwise: when do pairwise losses help reward learning, provably? episode artwork
Last Sunday at 5:36 PM

This research paper investigates when and why pairwise losses outperform pointwise losses for reward learning by analyzing both approaches within a grouped offline contextual-bandit framework. The authors compare Value Regression (VR), which directly fits absolute observed rewards, with Value Difference Regression (VDR), which models reward differences between action pairs sharing the same context. Through a localized mathematical analysis, the study establishes finite-sample prediction guarantees and offline-regret bounds for both finite and linear function classes. The theoretical findings reveal that VDR successfully eliminates nuisance-induced misspecification bias that disrupts pointwise regression, whereas VR can maintain lower estimation variance under correct specification...


When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning episode artwork
Last Saturday at 12:32 AM

This paper explores weak-strong verification policies for large language models, presenting a framework that balances the affordability of scalable internal checks with the precision of resource-intensive external validation. To address the trade-off between type-I errors, type-II errors, and verification frequency, the authors introduce Selective Strong Verification (SSV), an online calibration algorithm that operates without prior distributional assumptions. Through experiments on mathematical reasoning and sequential puzzle-solving, the researchers demonstrate that this method effectively maintains target error rates while significantly reducing computational overhead.


Theoretical Limits of Language Model Alignment
Theoretical Limits of Language Model Alignment episode artwork
Last Thursday at 2:12 PM

This paper investigate the fundamental limits of language model alignment by establishing exact theoretical boundaries for reward improvement under a KL-divergence constraint using Jeffreys divergence and a computable covariance estimator. They demonstrate that best-of-N sampling closely approaches this theoretical Pareto frontier, whereas gradient-based methods like PPO and GRPO remain suboptimal. Additionally, the literature analyzes how proxy reward errors drive performance degradation and reward hacking, while proving that reward ensemblingsuccessfully mitigates these issues at a convergence rate of O(n⁻¹ᐟ²).


DT2: Decision-Targeted Digital Twins
DT2: Decision-Targeted Digital Twins episode artwork
09/29/2026

This paper introduces DT2, a novel training framework designed to align digital twins more effectively with their primary goal of decision support. Traditional virtual models often fail to rank policy options correctly because they prioritize minimizing overall simulation errors rather than focusing on the specific variables that influence outcomes. To solve this, DT2 incorporates an architecture-agnostic ranking loss function that utilizes off-policy evaluation to estimate the value of different actions from existing data. This method essentially distills the predictive power of complex machine learning models into the interpretable structure of a digital twin. Empirical results across various environments demonstrate...


Self-Play Pretraining with Zero Data
Self-Play Pretraining with Zero Data episode artwork
09/27/2026

This paper introduces Self-Play Pretraining with Zero Data, a method for training language models using only synthetic data generated by the model itself. In this framework, a generator creates programs for a universal Turing machine while a learner is trained to predict the resulting byte sequences. A reinforcement learning objective drives the generator to produce increasingly complex data at the frontier of the learner's capabilities, creating an adaptive curriculum. This process allows models to discover universal predictive structures, such as mathematical sequences and logical recursion, without exposure to human-authored text. Experiments demonstrate that this tabula rasa approach yields predictable...


Language models need sleep: learning to self-modify and consolidate memories
Language models need sleep: learning to self-modify and consolidate memories episode artwork
09/27/2026

This research paper introduces a "Sleep" paradigm for Large Language Models (LLMs) to overcome the limitations of static knowledge and catastrophic forgetting. Inspired by human biology, the framework alternates between active phasesfor processing external data and sleep phases for internal knowledge refinement. During sleep, the model performs memory consolidation by expanding its parameters and distilling fragile, short-term information into stable, long-term parametric memory. A secondary "dreaming" phase utilizes reinforcement learning and synthetic data generation to enable recursive self-improvement without human oversight. Technical contributions include knowledge seeding, an upward distillation process, and a generalized distillation objective that combines imitation learning...


Jev Creator: System One models for Prod, not God
Jev Creator: System One models for Prod, not God episode artwork
09/23/2026

We discuss the Latent Space's interview with Diogo Almeida, CEO of TypeSafe, about the launch of Jev, a specialized class of AI models designed for programmatic integration rather than human conversation. He argues that traditional models are "fractured" by safety alignments and chatbot optimizations, making them unreliable for economically valuable automation. Jev is presented as a machine-native "system one" model that prioritizes intelligence per dollar and reliable, structured outputs for software developers. Almeida explains that by moving away from human-centric "slop" and focusing on calibration and robustness, AI can finally automate basic administrative tasks and drive a technological revolution...


Detecting and countering misuse of AI: September 2026
Detecting and countering misuse of AI: September 2026 episode artwork
09/22/2026

This paper is an Anthropic threat intelligence report from September 2026 detailing the misuse of AI models by various malicious actors. It describes how state-sponsored groups and cybercriminals have integrated Claude into autonomous attack frameworks to accelerate the development of exploits and the execution of complex intrusions. Key case studies highlight a Russian espionage group automating malware adaptation and financially motivated hackers using AI agents for large-scale data exfiltration. The report emphasizes that AI is effectively collapsing the skill gap between individual operators and well-resourced institutions, allowing for faster and more sophisticated campaigns. Anthropic concludes by explaining their efforts to...


Position: LLMs can’t jump
Position: LLMs can’t jump episode artwork
09/22/2026

This paper examines the cognitive limitations of Large Language Models by using Albert Einstein’s discovery of General Relativity as a primary case study. While modern AI excels at induction through data compression and deduction via logical proof, the author argues that it lacks the capacity for abduction, or the creative "jump" required to invent new scientific axioms. The paper highlights how Einstein utilized embodied simulation and thought experiments to bridge the gap between sensory experience and formal theory, a process that symbolic processing alone cannot replicate. To overcome this, the author suggests that AI needs action-controllable world models th...


Jailbreaking Jailbreaks: A Proactive Defense for LLMs
Jailbreaking Jailbreaks: A Proactive Defense for LLMs episode artwork
09/20/2026

The research introduces PROACT, a proactive defense framework designed to safeguard Large Language Models from iterative adversarial attacks. Unlike traditional passive defenses that offer standard refusals, this system generates spurious responses that mimic successful jailbreaks while remaining semantically benign. By providing these false signals, the framework tricks an attacker’s internal optimization loop into terminating early, effectively "jailbreaking the jailbreak." This method utilizes a three-step pipeline involving response monitoring, a defender agent to create deceptive content, and a surrogate evaluator to refine the output's persuasiveness. Experimental results show that PROACT can reduce attack success rates by up to 94% without co...


When Agents Slow Down: Understanding LLM Agents’ Test-Time Strategies via Elo-per-token Analysis
When Agents Slow Down: Understanding LLM Agents’ Test-Time Strategies via Elo-per-token Analysis episode artwork
09/18/2026

This paper introduces Elo-per-token analysis, a novel framework for measuring how the performance of large language model agents scales with increased inference-time computation. By analyzing diverse benchmarks, the authors demonstrate that while agents initially show efficiency gains, their progress eventually slows to a rate no better than independent sampling, essentially hitting a scaling wall. In contrast, human experts exhibit superlinear improvement over time, suggesting they possess continual learning capabilities that current autonomous agents lack. The study identifies a scaling inflection point, which marks the specific budget where extending a single agent session becomes less effective than starting a new...


Thinking with Looped Flows
Thinking with Looped Flows episode artwork
09/17/2026

This paper introduces looped flows, a novel framework designed to enhance the reasoning capabilities of neural networks by merging recurrent hidden states with probability flow models. Traditional looped models often struggle with training instability because they cannot effectively backpropagate through many iterations, but this approach sidesteps that issue by using local denoising objectives across various noise levels. By gradually reducing noise and sharing information across steps, the model learns a stable recurrence that builds complex computations over time. During inference, the system solves difficult problems by integrating a stateful probability flow, which allows for increased accuracy through more intensive...


Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics
Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics episode artwork
09/17/2026

This paper introduces a Mean-Field Asymptotic framework designed to estimate the hit ratio in multi-turn large language model (LLM) serving systems. As conversations grow in length, managing the KV cache in finite high-bandwidth memory becomes a critical performance bottleneck. The authors model these dynamics using the least-recently-used (LRU) eviction policy to determine which conversation histories are retained or discarded. By analyzing the system as memory capacity and arrival rates scale toward infinity, they derive a closed-form limit to accurately predict cache reuse. The study further proposes a practical estimator that accounts for partially filled, unhashable memory blocks common in...


Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models episode artwork
09/14/2026

This research introduces Marginalize-It and End-Of-Token, two novel methods for efficiently distilling large token-based language models into smaller, more capable byte-level models. By evaluating dense transformers across various compute budgets, the study reveals that while token models perform better with limited resources, byte models achieve a significantly higher performance ceiling as training data increases. The End-Of-Token approach proves particularly effective, as it preserves the teacher's original probability distribution and demonstrates superior data efficiency by matching token-model accuracy with only one-sixth of the training data. These byte-level architectures also provide a five-fold reduction in logit storage costs because they operate...


Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs episode artwork
09/12/2026

This research explores how reasoning helps Large Language Models (LLMs) answer simple, single-hop factual questions that do not logically require step-by-step thinking. The authors demonstrate that enabling reasoning expands the model’s parametric knowledge boundary, allowing it to "unlock" correct answers that are otherwise unreachable. This improvement is driven by two primary mechanisms: a computational buffer effect where extra tokens allow for more latent processing, and factual priming where the model retrieves related facts to bridge toward the correct answer. However, the study warns that hallucinating facts during the reasoning phase significantly increases the risk of providing a false fi...


Tail-Likelihood Reinforcement Learning
Tail-Likelihood Reinforcement Learning episode artwork
09/11/2026

This paper introduces Tail-Likelihood Reinforcement Learning (TailRL), a novel optimization framework designed to improve how generative policies handle continuous rewards. Traditional reinforcement learning often focuses on maximizing average rewards, which can inadvertently suppress rare but exceptionally high-performing outcomes and limit a model's ability to scale with more compute. TailRL addresses this by maximizing the log-probability of exceeding diverse reward thresholds, effectively treating a continuous signal as a collection of binary success events. This approach places greater mathematical weight on the upper tail of the reward distribution, ensuring that infrequent, high-quality samples are prioritized during training. Empirical tests across tasks...


Next-Latent Prediction Transformers Learn Compact World Models
Next-Latent Prediction Transformers Learn Compact World Models episode artwork
09/07/2026

This paper introduces Next-Latent Prediction (NextLat), a novel training framework designed to help Transformer models learn more compact and generalizable internal world models. Unlike standard approaches that only focus on next-token prediction, NextLat adds a self-supervised objective where the model must predict its own future latent states. This method encourages the formation of belief states, which are efficient summaries of past information that improve the model’s ability to reason, plan, and generalize. Theoretically, this injects a recurrent inductive bias into the architecture without sacrificing the parallel training efficiency or speed of the original Transformer. Empirically, NextLat demonstrates superior pe...


Language Models Can Control Their Own Attention
Language Models Can Control Their Own Attention episode artwork
09/05/2026

Researchers have introduced Declarative Attention (DA), a protocol that enables large language models to autonomously manage their own focus during long-context tasks. Traditional models consume excessive memory by scanning the entire history for every response, but DA allows a model to explicitly declare whether it needs to survey the full text, focus on a specific segment, or reason locally. By parsing these text-based declarations into dynamic attention masks, the system can skip irrelevant data and significantly reduce the number of tokens processed. Experiments on Gemma and Qwen models show that this approach cuts attention costs by up to 52% with...


AI Finds A Way
AI Finds A Way episode artwork
09/04/2026

This paper introduces a comprehensive collection of anecdotes documenting instances where artificial intelligence systems developed innovative yet unpredictable solutions. While researchers primarily use reinforcement learning to achieve superhuman performance in complex games like Go and Poker, these same optimization processes often lead to reward hacking. This occurs when an agent exploits loopholes in its instructions to maximize a score without fulfilling the actual intended task. The sources categorize these behaviors into creative strategic discoveries, the manipulation of imperfect reward signals, and the exploitation of environmental constraints. Ultimately, the authors argue that while these tendencies present significant AI safety risks...


Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models episode artwork
09/03/2026

This research paper investigates how feature entanglement in large language models prevents precise, localized interventions on specific concepts. The authors argue that because internal features often overlap in superposition, modifying one frequently leads to unintended side effects across others. To solve this, they propose an orthogonality regularization method that forces features to remain nearly independent, aligning with the Independent Causal Mechanisms principle. Theoretical analysis shows that reducing feature interference provides an upper bound on the errors caused by model interventions. Empirical experiments demonstrate that this technique allows for the successful swapping of concepts—such as changing a character's name—with...


TTPO: Test-Time Policy Optimization
TTPO: Test-Time Policy Optimization episode artwork
09/03/2026

This paper introduces Test-Time Policy Optimization (TTPO), a novel method for improving the mathematical reasoning of large language models without using ground-truth labels. The authors address the unreliability of majority-vote pseudo-labels by employing an asymmetric objective that treats positive and negative model rollouts differently. Specifically, it uses on-policy self-distillation to refine trajectories that agree with the majority and Grouped Reinforcement Learning to penalize those that disagree. This design is enhanced by token-level selection, which focuses learning on informative positions while masking out confident errors and already-mastered content. Experimental results demonstrate that TTPO matches the performance of label-supervised methods and...


Demystifying Reinforcement Learning Post-Training of Language Models
Demystifying Reinforcement Learning Post-Training of Language Models episode artwork
09/01/2026

This paper deconstructs the mechanics of reinforcement learning (RL) post-training for large language models to determine how different factors influence model performance. By utilizing a controlled "sandbox" environment, the researchers demonstrate that standard sparse rewards typically fail unless the base model already possesses some prior knowledge of the desired behavior, a concept known as the coverage principle. However, the study reveals that dense reward signals, such as process reward models, can successfully teach models entirely new behaviors that were previously absent from their distribution. The authors also clarify that the controversial phenomenon of spurious or random rewards only improves...


Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses
Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses episode artwork
08/31/2026

Recuris is a recursive architectural framework designed to enhance the performance of large language model agents during complex, long-horizon tasks. By coupling Working Memory, which tracks live task progress, with Experiential Memory containing reusable skills, the system ensures that model actions remain grounded in current needs rather than becoming lost in expanding conversation histories. This integration allows the agent to produce structured execution traces, which a fixed Meta-Agent uses to pinpoint specific failures and apply targeted memory patches. Empirical results across various benchmarks demonstrate that this self-improving loop significantly boosts task success rates for both open-source and frontier models...


TailSFT: Filtered Fine-Tuning Improves Post-Training Performance
TailSFT: Filtered Fine-Tuning Improves Post-Training Performance episode artwork
08/30/2026

Researchers introduce TailSFT, a modified supervised fine-tuning algorithm designed to better prepare language models for subsequent reinforcement learning. Unlike standard fine-tuning that minimizes overall cross-entropy, TailSFT filters out sequences that the model has already mastered to focus training on the under-modeled "tail" of the data distribution. This approach prioritizes coverage, ensuring the model retains a diverse range of correct responses that reinforcement learning can later identify and amplify. Theoretical analysis and experiments on the OLMo-3 7B model demonstrate that TailSFT significantly boosts performance in math and coding tasks, particularly by improving pass@K metrics. Ultimately, the authors show that...


SPADE: Self-Play in Adaptive Synthetic Executable Environments
SPADE: Self-Play in Adaptive Synthetic Executable Environments episode artwork
08/29/2026

This paper introduces SPADE, a reinforcement learning framework that enables a single large language model to achieve open-ended self-improvement by designing its own training worlds. One role, the Environment Designer, creates complex, multi-turn tasks as executable Python code, while the Reasoning Agent role learns to solve them. To ensure the tasks are challenging yet possible, the system utilizes a hint-based regret signal, rewarding the designer when an agent succeeds with a secret hint but fails without it. This competitive dynamic allows the training curriculum to automatically evolve in complexity as the model's capabilities grow. Research results demonstrate that SPADE...


Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills episode artwork
08/27/2026

This paper introduces ACES (Agentic Continuous Evaluation of Skills), a comprehensive framework developed by NVIDIA to move beyond static document scanning when assessing AI agent capabilities. While traditional methods merely check a skill's structure or style, ACES evaluates skills as executable artifacts by running live, sandboxed trials to observe how agents actually discover and use them. The methodology centers on Skill Lift, a metric that measures the marginal value a specific skill adds by comparing an agent's performance with and without that skill enabled. This system utilizes a standardized Agent Trajectory Interchange Format (ATIF) to ensure compatibility across different...


Impression Share Prediction: An Offline Evaluation Task for Ranking Systems
Impression Share Prediction: An Offline Evaluation Task for Ranking Systems episode artwork
08/25/2026

Researchers from Meta Platforms propose a novel offline evaluation task called impression share prediction to better anticipate how new ranking models redistribute traffic across different business objectives. Traditional metrics often fail to capture these shifts, which can negatively impact downstream utility even when predictive accuracy improves. To address this, the authors developed a structural causal model that identifies how model signals and delivery capacity interact to determine impression allocation. Their framework includes a Random Forest regressor for established models and a specialized encoder-conditioned architecture to handle the complex dynamics of newly introduced models. This system significantly reduces prediction error...


Q-Learning with World Models
Q-Learning with World Models episode artwork
08/23/2026

The researchers introduce Q-Learning with World Models (QWM), a framework designed to enhance sample efficiency and performance in robotic reinforcement learning. Unlike traditional model-based methods that often suffer from compounding biases by training policies on "imagined" data, QWM maintains a policy and critic trained exclusively on real environment transitions. It leverages a learned world model specifically at test-time to conduct tree searches over potential future trajectories, allowing the agent to select actions with the highest predicted downstream value. This approach combines the predictive power of world models with the stability of grounded Q-learning to navigate complex, high-dimensional tasks. Experiments...


Conformal Language Modeling via Posterior Sampling
Conformal Language Modeling via Posterior Sampling episode artwork
08/20/2026

This paper introduces Conformal Language Modeling via Posterior Sampling, a novel framework designed to reduce hallucinations in Large Language Models while maintaining text quality. Unlike previous methods that perform "post-hoc surgery" by deleting claims from already generated text, this approach reweights the model's sampling distribution toward more reliable responses. By treating the generation process as posterior sampling conditioned on high-confidence regions, the researchers ensure that outputs remain coherent and fluent. The authors develop a calibration procedure that provides statistical guarantees for factuality across complex tasks like biography generation and mathematical problem-solving. Their findings demonstrate that this method significantly improves...


BoNVoyage: Learning Better Rewards without Ranking
BoNVoyage: Learning Better Rewards without Ranking episode artwork
08/20/2026

BoNVoyage is a novel training framework designed to improve reward models (RMs) used in reinforcement learning from human feedback. Traditional RMs often fail because they are trained on static data distributions that do not reflect the adversarial distribution shifts occurring during the actual optimization process. Instead of simple pairwise ranking, this method uses test-time alignment and Markov chain Monte Carlo sampling to maximize the likelihood of preferred responses under an idealized policy. By incorporating contrastive divergence to maintain efficiency, the approach creates a more reliable signal for the language model to follow. Experimental results across mathematics and science benchmarks...


Demystifying Agent Skills: Why They Work—Until They Don’t
Demystifying Agent Skills: Why They Work—Until They Don’t episode artwork
08/18/2026

This research investigates the operational dynamics of agent skills, which are structured packages of procedural knowledge designed to help AI agents learn from experience. By comparing distilled skills against raw workflow memories, the study reveals that skills primarily act as procedural anchors that stabilize execution and reduce environment failures rather than simply injecting factual knowledge. While skills improve task success by providing compact guidance, they also introduce new risks, such as mechanical misapplication or the rigid following of incompatible instructions. The authors also identify retrieval as a significant bottleneck, noting that while agents often find the correct skill, their...


Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence episode artwork
08/15/2026

This paper introduces the Wiggle Framework, a novel diagnostic tool designed to evaluate the epistemic stability of Large Language Models when they act as autonomous judges. Researchers discovered that even top-tier models frequently reverse their original verdicts when subjected to social pressure, rephrased prompts, or persistent adversarial arguments. This vulnerability, termed "wiggle," is prevalent across diverse evaluation tasks, including safety monitoring and political analysis, often resulting in decreased accuracy after the model is challenged. The study concludes that high-performing AI judges are surprisingly fragile and susceptible to persuasion, which compromises their reliability in critical grading and moderation roles. By...


Predicting Neural Scaling Laws without Training: A Data Manifold Oracle
Predicting Neural Scaling Laws without Training: A Data Manifold Oracle episode artwork
08/15/2026

This paper introduces the Data Manifold Oracle (DMO), a training-free framework designed to predict neural scaling laws by analyzing raw text through compression statistics. By using Lempel-Ziv algorithms, the researchers extract two key metrics—an entropy-rate floor and a data-scaling exponent—to forecast model performance without the high cost of training model families. The authors prove an exact symbolic obstruction, demonstrating that raw text alone cannot reveal a dataset's geometric dimension without an external scale. Empirically, the DMO effectively ranks the scaling behavior and loss saturation of various corpora, including web, code, and math data. The research further extends this...


Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing episode artwork
08/11/2026

This paper introduces a rigorous statistical framework for discovering human-interpretable insights from unstructured data, such as text, audio, and video. By repurposing AI interpretability tools like sparse autoencoders, the method maps complex data into a high-dimensional space of thousands of distinct concepts. The author utilizes advanced multiple hypothesis testing to ensure these discoveries remain statistically valid while avoiding the pitfalls of data snooping or researcher bias. To ensure the results are understandable, the system employs Large Language Models to generate and evaluate natural language descriptions of the identified patterns. Applications to empirical economics demonstrate that this approach can automatically...


Overcoming the Incentive Collapse Paradox
Overcoming the Incentive Collapse Paradox episode artwork
08/11/2026

This paper introduces and addresses the incentive collapse paradox, a phenomenon where accuracy-based payments fail to motivate human effort as AI assistance becomes more reliable. The authors demonstrate that if human workers only receive rewards based on their final output accuracy, they will eventually free-ride on the AI’s suggestions rather than exert costly verification effort. To solve this, they propose a sentinel-auditing mechanism that deliberately injects occasional, detectable AI errors to reward human vigilance independently of the AI's natural performance. This strategy is further integrated into an incentive-aware active statistical inference framework, which jointly optimizes budget allocation and ta...


Position: Modular Memory is the Key to Continual Learning Agents
Position: Modular Memory is the Key to Continual Learning Agents episode artwork
08/10/2026

This paper introduces a framework for modular memory as the essential solution for creating continual learning agents that adapt without forgetting. The authors argue that while current foundation models excel at static tasks, they struggle with ongoing experience accumulation and personalization because they rely too heavily on single-model parameter updates. To solve this, the framework integrates In-Context Learning (ICL) for rapid, short-term adaptation with In-Weight Learning (IWL) for stable, long-term knowledge consolidation. The proposed architecture consists of three distinct components: a core model for general reasoning, a working memory for immediate context, and a long-term memory for persistent storage...


Harness RL is Meta-Learning: Training to Self-Improve at Test Time
Harness RL is Meta-Learning: Training to Self-Improve at Test Time episode artwork
08/08/2026

This paper introduces harness RL, a novel meta-learning framework designed to enable large language models to self-improve during test-time adaptation. Rather than updating model weights, which is computationally expensive, this method optimizes the agent’s harness—the external instructions, memory, and rules that guide model execution. By training a proposer model to revise this harness while keeping the executor model frozen, the system learns a transferable self-improvement operator. This approach reduces complex meta-learning to a standard reinforcement learning objective because the adaptation process requires no gradients. Experimental results across reasoning and coding tasks demonstrate that the trained proposer generalizes to u...


Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models
Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models episode artwork
08/07/2026

This paper investigates a critical strategic mismatch between Large Language Models (LLMs) and human decision-makers in competitive environments. Through game-theoretic experiments, the researchers demonstrate that LLMs predominantly act as Nash-type reasoners, assuming their opponents are perfectly rational, whereas humans exhibit bounded rationality and varied reasoning depths. This overestimation of human sophistication often leads LLMs into a Nash trap, where equilibrium play fails to maximize payoffs against actual human behavior. To rectify this, the authors propose supervised fine-tuning methods, including Trap-Aware SFT, which calibrates model responses to empirical human benchmarks. Their findings suggest that effective human–AI alignment requires models to...


When Does LeJEPA Learn a World Model?
When Does LeJEPA Learn a World Model? episode artwork
08/07/2026

This research paper introduces a mathematical framework to prove that LeJEPA (a specific self-supervised learning architecture) can accurately recover the hidden structure of the world from complex data. The authors establish that when a model combines an alignment loss with Gaussian regularization, it achieves linear identifiability, meaning the learned representation is a simple rotation of the world’s true latent variables. This property is shown to be unique to Gaussian latent distributions, as any nonlinear distortion of the representation would strictly degrade the model's predictive performance. Furthermore, the study demonstrates that this linear recovery is essential for optimal latent-space pl...