每日AI论文
每个工作日更新,解读 Hugging Face Daily Papers(https://huggingface.co/papers)中得票最高的论文。播客文稿和音频均由 AI 生成。欢迎反馈与建议!邮箱:dailypapercast.ai@gmail.com 创作者: Jingwen Liang,3D 机器学习,https://www.linkedin.com/in/jingwen-liang/ Gengyu Wang,LLM ML,http://wanggengyu.com 英文版收听: Spotify:https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL Apple Podcasts:https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236 封面图片:Kawen Kuang https://kawen.art
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
🤗 Upvotes: 90 | cs.CL
作者:
Tao Feng, Fangxu Yu, Haozhen Zhang, Zhongjie Dai, Liangqi Yuan, Zijie Lei, Weizhi Zhang, Kunlun Zhu, Haodong Yue, Keyang Xuan, Ge Liu, Jiaxuan You
标题:
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Arxiv:
http://arxiv.org/abs/2608.06867v1
摘要:
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential deci...
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
🤗 Upvotes: 82 | cs.CV
作者:
Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao
标题:
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Arxiv:
http://arxiv.org/abs/2608.13546v1
摘要:
Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by...
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
🤗 Upvotes: 79 | cs.CV, cs.RO
作者:
DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang
标题:
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Arxiv:
http://arxiv.org/abs/2608.13489v1
摘要:
We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong...
DarwinX: Evolving Agent Harnesses Through Natural Selection
🤗 Upvotes: 59 | cs.NE, cs.AI, cs.LG, cs.SE
作者:
Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, Zhiyuan Hu, James Zhu, Phil Mui, Silvio Savarese, Ran Xu, Zeyuan Chen
标题:
DarwinX: Evolving Agent Harnesses Through Natural Selection
Arxiv:
http://arxiv.org/abs/2608.07545v1
摘要:
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evo...
Intern-S2-Preview: Scientific Agentic Foundation Model
🤗 Upvotes: 43 | cs.LG, cs.CL, cs.CV
作者:
Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Zi...
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
🤗 Upvotes: 39 | cs.CL, cs.AI
作者:
Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
标题:
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Arxiv:
http://arxiv.org/abs/2608.08975v1
摘要:
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR...
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
🤗 Upvotes: 33 | cs.CV
作者:
Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao
标题:
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Arxiv:
http://arxiv.org/abs/2608.13552v1
摘要:
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing lon...
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
🤗 Upvotes: 32 | cs.CV, cs.AI, cs.CL
作者:
Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li
标题:
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Arxiv:
http://arxiv.org/abs/2608.13560v1
摘要:
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical explo...
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
🤗 Upvotes: 29 | cs.AI
作者:
Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen
标题:
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Arxiv:
http://arxiv.org/abs/2608.12743v1
摘要:
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatia...
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
🤗 Upvotes: 216 | cs.CL
作者:
Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu, Jie Li, Yan Teng, Xingjun Ma, Xia Hu, Yu-Gang Jiang
标题:
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Arxiv:
http://arxiv.org/abs/2608.00677v1
摘要:
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety benchmarks often fail to capture these cumulative risks because they focus on short, stati...
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
🤗 Upvotes: 211 | cs.CL
作者:
Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng, Jin Jiang, Junsheng Zhang, Wenhao Wang
标题:
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Arxiv:
http://arxiv.org/abs/2608.11924v1
摘要:
Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skil...
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
🤗 Upvotes: 101 | cs.LG, cs.AI, cs.CL
作者:
Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke
标题:
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Arxiv:
http://arxiv.org/abs/2608.12307v1
摘要:
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study strong-to-weak scaffolding: whether a stronger buil...
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
🤗 Upvotes: 76 | cs.AI, cs.CL, cs.HC, cs.LG, cs.MA
作者:
Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
标题:
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Arxiv:
http://arxiv.org/abs/2608.12036v1
摘要:
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain...
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
🤗 Upvotes: 73 | cs.CL, cs.AI
作者:
Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu, Xin Yuan, Liming Zhu, Wenjie Zhang
标题:
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
Arxiv:
http://arxiv.org/abs/2608.05604v1
摘要:
Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time. As skill libraries grow, a central challenge is to expose the smallest sufficient executable context under a limited context budget. Existing systems struggle to reuse routines below the whole-skill level, preserve procedural cont...
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
🤗 Upvotes: 26 | cs.CL, cs.AI
作者:
Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong
标题:
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
Arxiv:
http://arxiv.org/abs/2608.08160v1
摘要:
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against unconstrained user interventions. To address this...
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
🤗 Upvotes: 23 | cs.CV
作者:
Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao, Junnan Liu, Weirong Huang, Mengyu Wang, Tianxiao Fu, Yikai Wang, Peng-Shuai Wang, Xiaojie Jin, Yao Zhao, Yunchao Wei
标题:
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Arxiv:
http://arxiv.org/abs/2608.12314v1
摘要:
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors t...
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
🤗 Upvotes: 175 | cs.AI
作者:
Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
标题:
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Arxiv:
http://arxiv.org/abs/2608.10915v2
摘要:
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whethe...
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
🤗 Upvotes: 120 | cs.CL
作者:
Qing Zong, Jiayu Liu, Junhao Shen, Zecong Tang, Linsi Wu, Yuxuan Liu, Rui Wang, Zhaowei Wang, Weiqi Wang, Cheng Qian, Xiusi Chen, Yangqiu Song
标题:
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Arxiv:
http://arxiv.org/abs/2608.10299v1
摘要:
Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic systems, a multi-component form of self-evolution in which multiple agents and their environme...
Beyond Pixels: From Video Priors to 4D Worlds
🤗 Upvotes: 108 | cs.CV
作者:
Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang
标题:
Beyond Pixels: From Video Priors to 4D Worlds
Arxiv:
http://arxiv.org/abs/2608.10744v1
摘要:
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or co...
Articulated Object Reconstruction from Rest-State Observation
🤗 Upvotes: 40 | cs.CV, cs.RO
作者:
Daeun Lee, Jaeah Lee, Woosung Kim, Haebeom Jung, Jaesik Park
标题:
Articulated Object Reconstruction from Rest-State Observation
Arxiv:
http://arxiv.org/abs/2607.27749v1
摘要:
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of moti...
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
🤗 Upvotes: 23 | cs.CV
作者:
Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, Hao Tang
标题:
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Arxiv:
http://arxiv.org/abs/2608.11205v1
摘要:
Fr\'echet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fr\'echet objectives can cause Fr\'echet hacking. The target metrics keep improving, but visual quality and Fr\'echet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrain...
Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
🤗 Upvotes: 22 | cs.AI, cs.LG
作者:
Changzhi Liu, Yilun Liu, Sikuan Yan, Volker Tresp, Yunpu Ma
标题:
Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Arxiv:
http://arxiv.org/abs/2608.07645v1
摘要:
Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self-modification from a single failure trajectory at a time, overlooking rich comparative signals available in the agent's expanding archive of past attempts. According to Mendelian principles of controlled inheritance, we introduce Mendel G\"odel Machine (M...
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
🤗 Upvotes: 275 | cs.LG, cs.CL
作者:
Mind Lab, :, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Kuss Koo, Jaron Lee, Andrew Lei, Alexy Li, Dawn Li, Lucian Li, Ray Li, Ricardo Li, Smith Li, Theo Li, Allen Lin, Elliot Lin, Fan Lin, Chen Ling, Kairus Liu, Kieran Liu, Logan Liu, Neo Liu, Xiang Liu, Yu...
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
🤗 Upvotes: 253 | cs.NE, cs.AI, cs.LG, stat.ML
作者:
Björn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemysław Uznański, Junlin Jiang, Rohan Phadke, Remigiusz Kinas, Richard Zhong
标题:
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Arxiv:
http://arxiv.org/abs/2608.09888v1
摘要:
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning. We evaluate the mo...
On-Policy Self-Distillation without Any Supervision
🤗 Upvotes: 90 | cs.LG
作者:
Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos
标题:
On-Policy Self-Distillation without Any Supervision
Arxiv:
http://arxiv.org/abs/2608.06296v2
摘要:
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-distillation. In this study, we show that on-policy self-distillation can be achieved using only a model's own generations via internal consistency. We...
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
🤗 Upvotes: 68 | cs.SE, cs.AI
作者:
Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev, Nikita Dragunov, Roman Yampolskiy, Andrei Kuznetsov
标题:
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
Arxiv:
http://arxiv.org/abs/2608.08311v2
摘要:
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement is itself a task, and completing one evolution cycle can schedule the next. In experience-driven core evolution, ordinary work a...
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
🤗 Upvotes: 37 | cs.AI, cs.LG
作者:
Taeil Kim, Kangsan Kim, Sung Ju Hwang
标题:
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Arxiv:
http://arxiv.org/abs/2608.07169v1
摘要:
Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory. AMD constructs three complementary m...
Motif 3: Technical Report
🤗 Upvotes: 36 | cs.AI
作者:
Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki Ryu
标题:
Motif 3: Technical Report
Arxiv:
http://arxiv.org/abs/2608.09119v1
摘要:
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer conta...
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
🤗 Upvotes: 26 | cs.CV, cs.AI
作者:
Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao
标题:
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Arxiv:
http://arxiv.org/abs/2608.09873v1
摘要:
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establi...
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
🤗 Upvotes: 25 | cs.CV, cs.AI
作者:
Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
标题:
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
Arxiv:
http://arxiv.org/abs/2608.07565v1
摘要:
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn imag...
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
🤗 Upvotes: 33 | cs.CL, cs.LG
作者:
Kejian Zhu, Zhuoran Jin, Shangqing Tu, Hongbang Yuan, Yushi Bai, Kang Liu, Juanzi Li, Jun Zhao
标题:
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
Arxiv:
http://arxiv.org/abs/2608.03573v2
摘要:
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the par...
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
🤗 Upvotes: 29 | cs.CV
作者:
Kejian Zhu, Zhuoran Jin, Dongqi Huang, Hongbang Yuan, Yupu Hao, Kang Liu, Jun Zhao
标题:
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
Arxiv:
http://arxiv.org/abs/2608.03571v2
摘要:
Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments does not always benefit. We further analyze the limitations in current multimodal environment distributions through a series of experiments. Based on these findings, we study how to build more effective training environment dis...
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
🤗 Upvotes: 22 | cs.CV
作者:
Zongchuang Zhao, Xin Zhou, Tianyang Xu, Zhengyang Sun, Kaixuan Zhou, Honglin Li, Dingkang Liang, Xiang Bai
标题:
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
Arxiv:
http://arxiv.org/abs/2608.07468v1
摘要:
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow...
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
🤗 Upvotes: 88 | cs.AI, cs.LG
作者:
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang
标题:
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Arxiv:
http://arxiv.org/abs/2608.05987v1
摘要:
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local...
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
🤗 Upvotes: 68 | cs.AI, cs.CL, cs.CV
作者:
Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
标题:
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Arxiv:
http://arxiv.org/abs/2607.28609v2
摘要:
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Veri...
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
🤗 Upvotes: 68 | cs.LG, cs.SD, q-bio.NC
作者:
Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova, Alex Ossadtchi
标题:
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
Arxiv:
http://arxiv.org/abs/2608.01481v1
摘要:
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. We build on a high-performing MEG-to-audio re...
WorldClaw: Agentic 3D Open-World Generation at Scale
🤗 Upvotes: 60 | cs.AI, cs.CV
作者:
Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang
标题:
WorldClaw: Agentic 3D Open-World Generation at Scale
Arxiv:
http://arxiv.org/abs/2608.05248v1
摘要:
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relatio...
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
🤗 Upvotes: 42 | cs.CV
作者:
Qifeng Zhang, Kaixiang Huang, Heng Dong, Huang Fang, Junting Chen, Junjie Zhu, Yonghang Chen, Zhiyu Zhang, Wei Li
标题:
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
Arxiv:
http://arxiv.org/abs/2608.05747v1
摘要:
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), a VQA benchmark for global spatial intelligence in video understanding, comprising human-verified questions deriv...
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
🤗 Upvotes: 39 | cs.AI
作者:
Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan Zhang, Xingshan Zeng, Weiwen Liu
标题:
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
Arxiv:
http://arxiv.org/abs/2608.06197v1
摘要:
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces extern...
ChronoVision: Temporal Reasoning via Latent State Reconstruction
🤗 Upvotes: 38 | cs.CV
作者:
Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao
标题:
ChronoVision: Temporal Reasoning via Latent State Reconstruction
Arxiv:
http://arxiv.org/abs/2608.05631v1
摘要:
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent ima...