Daily Paper Cast

Jingwen Liang, Gengyu Wang

500 episodes listed below

Listen to the show on

AI papers, structured for fast technical listening.

The gist

Daily Paper Cast is a fast-moving research explainer built around papers surfaced from Hugging Face's Daily Paper list. Echo and Nova move through introductions, methods, experiments, and related work with a steady focus on models, metrics, and implementation details.

PRESS PLAY

Find your next episode

All episodes

About Daily Paper Cast

Daily Paper Cast converts current AI research papers into structured audio briefings. The show's recurring shape is clear: introduce the paper, explain the motivation, walk through the method, summarize the experiments, and close with related work or future directions. Echo and Nova carry the discussion as complementary synthetic voices. One tends to ask the clarifying question, while the other fills in the technical machinery, and both keep the pace moving through dense material. Recent episodes range across multimodal world models, autonomous driving reinforcement learning, deep research agents, occupational task simulation, visual reward modeling, coding-agent memory transfer, video generation, and game-agent evaluation. That breadth gives the show a strong survey function. It is less about hot takes and more about turning a paper's architecture, benchmark design, and result tables into a listenable sequence. The vocabulary stays technical. Terms such as CLIP alignment, closed-loop evaluation, RAG sub-agents, robustness scores, reward models, inference optimization, and progress metrics are treated as normal working language. The show does, however, keep checking comprehension by restating concepts after the first pass. It also pays unusual attention to evaluation. Episodes repeatedly ask which baselines were used, what metrics counted, how many scenarios or tasks were tested, and which limitations remained after the reported gains. This makes it useful for researchers, builders, and technically fluent listeners who care about whether a method is merely impressive or actually measured well. The tone is measured, tidy, and faintly enthusiastic. There are occasional synthetic-audio rough edges and repeated phrases in the excerpts, but the editorial center remains consistent. Daily Paper Cast works best as a daily research companion for people who already know the field's vocabulary and want a fast map of what new papers are claiming. It gives the listener a structured first read before deciding whether the full paper deserves closer attention.

Made for: Daily Paper Cast is for machine learning practitioners, AI researchers, technical founders, and advanced students who want quick orientation to new papers. It assumes comfort with model names, benchmarks, metrics, and research jargon.

What sets it apart: The show keeps the paper's own structure intact, moving section by section through methods and experiments instead of turning research into loose commentary. Its synthetic duo format makes dense benchmark-heavy material feel more conversational without stripping out the technical substance.

In their own words

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: [email protected] Creator: Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/ Gengyu Wang, LLM ML, http://wanggengyu.com Listen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236 Cover Image by Kawen Kuang https://kawen.art

As heard by us

Based on 6 episodes we listened to · September 2026

Daily Paper Cast makes current AI papers legible by pairing accessible explanation with close attention to methods, benchmarks, results, and limitations.

Daily Paper Cast turns individual machine-learning papers into structured conversations between Echo and Nova, moving through the research question, methods, experiments, and related work.

Read our full review in PlayNext →

Why you'd press play

Echo and Nova turn research-paper methods and experiments into the lab meeting you missed.

Press play if you want

  • methods, benchmarks, and ablations translated into conversational checkpoints instead of one dense monologue
  • one paper at a time, with enough structure to keep the acronyms from staging a coup
Read the full recommendation in PlayNext →
AI research papersmultimodal modelsbenchmark designreinforcement learningAI agent evaluationvisual generationautonomous drivingmodel metrics

Talks about

Podcasts like Daily Paper Cast

Episodes

  1. 1

    ROWBench: Do Video Models Render What the Program Specifies?

    Oct 2, 2026·22m
  2. 2

    A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

    Oct 2, 2026·22m
  3. 3

    Agent Priors-guided Policy Learning

    Oct 2, 2026·21m
  4. 4

    ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

    Oct 2, 2026·22m
  5. 5

    Sharpening Tax in Post-Training

    Oct 2, 2026·21m
  6. 6

    Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

    Oct 2, 2026·20m
  7. 7

    AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

    Oct 2, 2026·21m
  8. 8

    Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

    Oct 2, 2026·22m
  9. 9

    World Observer: Joint Actor-Observer Generation for Persistent World Modeling

    Oct 2, 2026·20m
  10. 10

    Hierarchical Continuous Diffusion Language Models

    Oct 2, 2026·21m
  11. 11

    Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

    Oct 1, 2026·22m
  12. 12

    More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

    Oct 1, 2026·22m
  13. 13

    AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

    Oct 1, 2026·22m
  14. 14

    WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

    Oct 1, 2026·23m
  15. 15

    False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

    Oct 1, 2026·20m
  16. 16

    Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

    Oct 1, 2026·24m
  17. 17

    The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

    Oct 1, 2026·22m
  18. 18

    EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

    Oct 1, 2026·21m
  19. 19

    UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

    Oct 1, 2026·22m
  20. 20

    What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

    Sep 30, 2026·21m
  21. 21

    Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

    Sep 30, 2026·21m
  22. 22

    Think Before You Score: Thinking Reward Model for Visual Generation

    Sep 30, 2026·22m
  23. 23

    MaLiang-Harness: A Programmable Path to Image and Video Generation

    Sep 30, 2026·23m
  24. 24

    VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

    Sep 30, 2026·21m
  25. 25

    Omni-IO Skills: Harnessing Your Agent Omni-Native

    Sep 30, 2026·22m
  26. 26

    Raven: The Harness of Harnesses for Composable Agentic Intelligence

    Sep 30, 2026·26m
  27. 27

    SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation

    Sep 30, 2026·22m
  28. 28

    PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation

    Sep 30, 2026·21m
  29. 29

    In-Context Learning for Robots: Methods and Applications

    Sep 30, 2026·21m
  30. 30

    Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

    Sep 29, 2026·22m
  31. 31

    Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue

    Sep 29, 2026·18m
  32. 32

    Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

    Sep 29, 2026·21m
  33. 33

    Post-Training Leaves Behavioral Shadows on Unrelated Decisions

    Sep 29, 2026·19m
  34. 34

    TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

    Sep 29, 2026·21m
  35. 35

    Improving Test-Time Scaling with Adaptive Looped Transformers

    Sep 29, 2026·22m
  36. 36

    CompoWorld: Compositional Environment Scaling for General Agents

    Sep 29, 2026·23m
  37. 37

    YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

    Sep 29, 2026·22m
  38. 38

    How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

    Sep 29, 2026·19m
  39. 39

    EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    Sep 29, 2026·24m
  40. 40

    Training Object Permanence in World Models

    Sep 25, 2026·23m
  41. 41

    The Past Frames the Future: Memory for Autoregressive Video Generation

    Sep 24, 2026·17m
  42. 42

    HappyWorld-Bench

    Sep 24, 2026·24m
  43. 43

    RULER: Instance-aware Rubric Rewards for SVG Generation

    Sep 23, 2026·21m
  44. 44

    GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

    Sep 23, 2026·21m
  45. 45

    All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

    Sep 23, 2026·20m
  46. 46

    RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

    Sep 22, 2026·24m
  47. 47

    OmniEdu: Open Foundation Models for Learning and Teaching

    Sep 22, 2026·24m
  48. 48

    GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

    Sep 22, 2026·22m
  49. 49

    WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

    Sep 22, 2026·20m
  50. 50

    Transferring the Intelligence of VLMs to Robotic Control

    Sep 22, 2026·22m
  51. 51

    VideoGen-Agent: Reinforcing Video Generation Agents

    Sep 22, 2026·26m
  52. 52

    An Empirical Study of Harness Design for Coding Agents

    Sep 18, 2026·24m
  53. 53

    JEPA-Anything: Learning Predictive Models across Different Worlds

    Sep 18, 2026·20m
  54. 54

    DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

    Sep 18, 2026·23m
  55. 55

    SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

    Sep 18, 2026·23m
  56. 56

    A Zeroth-Order Paradigm for LLM Preference Alignment

    Sep 17, 2026·21m
  57. 57

    Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

    Sep 17, 2026·23m
  58. 58

    Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

    Sep 17, 2026·20m
  59. 59

    VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

    Sep 17, 2026·23m
  60. 60

    ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

    Sep 17, 2026·22m
  61. 61

    Agora: Git as Shared Memory for Collective AutoResearch

    Sep 17, 2026·22m
  62. 62

    ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    Sep 17, 2026·23m
  63. 63

    ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

    Sep 17, 2026·21m
  64. 64

    Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

    Sep 17, 2026·22m
  65. 65

    EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

    Sep 17, 2026·21m
  66. 66

    Continual Learning Mechanisms Compose for Long-Horizon Memorization

    Sep 16, 2026·23m
  67. 67

    StepAudio 3 Realtime Technical Report

    Sep 16, 2026·23m
  68. 68

    The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

    Sep 16, 2026·22m
  69. 69

    Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

    Sep 15, 2026·22m
  70. 70

    ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

    Sep 15, 2026·21m
  71. 71

    Dream-RSI: Recursive Self-Improvement through Evolving Worlds

    Sep 15, 2026·20m
  72. 72

    Atria Dawn: The Dawn of Agentic Superintelligence

    Sep 15, 2026·21m
  73. 73

    Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

    Sep 15, 2026·21m
  74. 74

    PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

    Sep 15, 2026·23m
  75. 75

    EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

    Sep 11, 2026·21m
  76. 76

    SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

    Sep 11, 2026·22m
  77. 77

    SenseNova-U1.5: Towards Native Unified Visual Intelligence

    Sep 11, 2026·22m
  78. 78

    NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

    Sep 11, 2026·24m
  79. 79

    Show-Harness: Just a VLM Agent Can Play Robots

    Sep 10, 2026·21m
  80. 80

    Programmable World Model

    Sep 10, 2026·19m
  81. 81

    WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

    Sep 10, 2026·22m
  82. 82

    Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

    Sep 9, 2026·21m
  83. 83

    DriveZero: End-to-End Driving Beyond Human Demonstrations

    Sep 9, 2026·22m
  84. 84

    SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

    Sep 9, 2026·18m
  85. 85

    Miles v0.1: Production-Level Post-Training

    Sep 9, 2026·21m
  86. 86

    Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

    Sep 9, 2026·21m
  87. 87

    NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

    Sep 9, 2026·20m
  88. 88

    Omni Interaction Agent Technical Report

    Sep 9, 2026·22m
  89. 89

    GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

    Sep 9, 2026·20m
  90. 90

    AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

    Sep 9, 2026·20m
  91. 91

    Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

    Sep 4, 2026·22m
  92. 92

    LatentPress: Context Compression Beyond Text and Vision

    Sep 4, 2026·20m
  93. 93

    Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

    Sep 4, 2026·21m
  94. 94

    Rethinking On-Policy Distillation of Large Language Models II: One Training Example

    Sep 4, 2026·22m
  95. 95

    LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

    Sep 4, 2026·23m
  96. 96

    Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

    Sep 4, 2026·20m
  97. 97

    Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

    Sep 4, 2026·22m
  98. 98

    SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

    Sep 3, 2026·21m
  99. 99

    It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning

    Sep 3, 2026·21m
  100. 100

    Language Models Can Control Their Own Attention

    Sep 3, 2026·21m
  101. 101

    Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

    Sep 3, 2026·21m
  102. 102

    On the Design Fundamentals of Pixel Text Representation Learning

    Sep 3, 2026·20m
  103. 103

    EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

    Sep 3, 2026·21m
  104. 104

    SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

    Sep 2, 2026·21m
  105. 105

    Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering

    Sep 2, 2026·20m
  106. 106

    H3-World: Turning Language Understanding into World Control

    Sep 2, 2026·21m
  107. 107

    StudentSim: Training LLM-based Student Simulators

    Sep 2, 2026·21m
  108. 108

    ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

    Sep 2, 2026·21m
  109. 109

    Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

    Sep 2, 2026·21m
  110. 110

    UI-Venus-2 Technical Report

    Sep 2, 2026·23m
  111. 111

    Normalized Low-Rank Adaptation

    Sep 1, 2026·19m
  112. 112

    CogEvol: Towards Efficient and Reliable Learning Environment Generation

    Sep 1, 2026·18m
  113. 113

    DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

    Sep 1, 2026·26m
  114. 114

    Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

    Sep 1, 2026·22m
  115. 115

    GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

    Sep 1, 2026·21m
  116. 116

    Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

    Sep 1, 2026·22m
  117. 117

    PaperGym: Rubric-Centered Evolution for Research-Plan Generation

    Sep 1, 2026·22m
  118. 118

    GameWAM: A World Action Model for Video Games

    Aug 28, 2026·23m
  119. 119

    What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

    Aug 28, 2026·21m
  120. 120

    PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

    Aug 28, 2026·22m
  121. 121

    Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

    Aug 28, 2026·21m
  122. 122

    TTPO: Test-Time Policy Optimization

    Aug 28, 2026·21m
  123. 123

    Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

    Aug 28, 2026·20m
  124. 124

    UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

    Aug 28, 2026·23m
  125. 125

    Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

    Aug 28, 2026·19m
  126. 126

    PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

    Aug 28, 2026·21m
  127. 127

    VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

    Aug 27, 2026·22m
  128. 128

    VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

    Aug 27, 2026·20m
  129. 129

    VGI-Bench: Probing Visual Intelligence in Video Generation Models

    Aug 27, 2026·19m
  130. 130

    WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

    Aug 27, 2026·20m
  131. 131

    JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

    Aug 27, 2026·20m
  132. 132

    FrontierChallenge: Evaluating Scientific Workflow Completion

    Aug 27, 2026·25m
  133. 133

    Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

    Aug 26, 2026·19m
  134. 134

    Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

    Aug 25, 2026·18m
  135. 135

    An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

    Aug 19, 2026·19m
  136. 136

    Agentic Transaction: Towards ACID-Compliant Agent Systems

    Aug 19, 2026·20m
  137. 137

    Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

    Aug 18, 2026·25m
  138. 138

    Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

    Aug 18, 2026·20m
  139. 139

    Self-Supervised Visual On-Policy Distillation

    Aug 18, 2026·22m
  140. 140

    Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

    Aug 18, 2026·23m
  141. 141

    Marionette: Predicting World States, Rendering Geometry, Painting Appearance

    Aug 18, 2026·21m
  142. 142

    DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

    Aug 18, 2026·17m
  143. 143

    Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

    Aug 18, 2026·22m
  144. 144

    PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

    Aug 15, 2026·22m
  145. 145

    Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

    Aug 15, 2026·20m
  146. 146

    Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

    Aug 15, 2026·20m
  147. 147

    How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

    Aug 15, 2026·20m
  148. 148

    Intern-S2-Preview: Scientific Agentic Foundation Model

    Aug 15, 2026·23m
  149. 149

    DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

    Aug 15, 2026·22m
  150. 150

    AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

    Aug 15, 2026·22m
  151. 151

    DarwinX: Evolving Agent Harnesses Through Natural Selection

    Aug 15, 2026·21m
  152. 152

    Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

    Aug 14, 2026·23m
  153. 153

    StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

    Aug 14, 2026·20m
  154. 154

    Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

    Aug 14, 2026·22m
  155. 155

    AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

    Aug 14, 2026·23m
  156. 156

    SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries

    Aug 14, 2026·19m
  157. 157

    Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

    Aug 14, 2026·22m
  158. 158

    OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

    Aug 14, 2026·22m
  159. 159

    Articulated Object Reconstruction from Rest-State Observation

    Aug 13, 2026·20m
  160. 160

    AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

    Aug 13, 2026·20m
  161. 161

    Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

    Aug 13, 2026·22m
  162. 162

    ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

    Aug 13, 2026·23m
  163. 163

    Beyond Pixels: From Video Priors to 4D Worlds

    Aug 13, 2026·22m
  164. 164

    Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

    Aug 13, 2026·19m
  165. 165

    Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

    Aug 12, 2026·23m
  166. 166

    Motif 3: Technical Report

    Aug 12, 2026·24m
  167. 167

    SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

    Aug 12, 2026·19m
  168. 168

    BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

    Aug 12, 2026·22m
  169. 169

    On-Policy Self-Distillation without Any Supervision

    Aug 12, 2026·21m
  170. 170

    Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

    Aug 12, 2026·20m
  171. 171

    What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

    Aug 12, 2026·20m
  172. 172

    Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

    Aug 12, 2026·21m
  173. 173

    Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

    Aug 12, 2026·22m
  174. 174

    SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

    Aug 11, 2026·21m
  175. 175

    Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

    Aug 11, 2026·21m
  176. 176

    SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

    Aug 11, 2026·22m
  177. 177

    Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

    Aug 8, 2026·21m
  178. 178

    AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

    Aug 8, 2026·21m
  179. 179

    OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

    Aug 8, 2026·20m
  180. 180

    GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

    Aug 8, 2026·20m
  181. 181

    ChronoVision: Temporal Reasoning via Latent State Reconstruction

    Aug 8, 2026·20m
  182. 182

    Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

    Aug 8, 2026·22m
  183. 183

    WorldClaw: Agentic 3D Open-World Generation at Scale

    Aug 8, 2026·22m
  184. 184

    EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

    Aug 8, 2026·19m
  185. 185

    ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

    Aug 7, 2026·21m
  186. 186

    Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

    Aug 7, 2026·24m
  187. 187

    ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

    Aug 7, 2026·20m
  188. 188

    Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

    Aug 7, 2026·19m
  189. 189

    Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

    Aug 6, 2026·23m
  190. 190

    PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

    Aug 6, 2026·22m
  191. 191

    Quo Vadis, World Modeling?

    Aug 6, 2026·21m
  192. 192

    PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

    Aug 6, 2026·22m
  193. 193

    LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

    Aug 6, 2026·22m
  194. 194

    MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

    Aug 6, 2026·22m
  195. 195

    JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

    Aug 6, 2026·21m
  196. 196

    AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

    Aug 6, 2026·20m
  197. 197

    Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

    Aug 6, 2026·18m
  198. 198

    SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    Aug 5, 2026·20m
  199. 199

    LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

    Aug 5, 2026·23m
  200. 200

    DAPD: Dual-Anchored Policy Distillation

    Aug 5, 2026·20m
  201. 201

    Progressive Agent Skill Generation via Reinforcement Learning

    Aug 5, 2026·22m
  202. 202

    UEmbed: Unified Sparse and Dense Multimodal Embeddings

    Aug 5, 2026·20m
  203. 203

    Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

    Aug 5, 2026·22m
  204. 204

    VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

    Aug 5, 2026·23m
  205. 205

    WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

    Aug 5, 2026·21m
  206. 206

    CADENA: Stepwise CAD Reverse Engineering

    Aug 5, 2026·23m
  207. 207

    InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

    Aug 5, 2026·24m
  208. 208

    $N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

    Aug 4, 2026·20m
  209. 209

    Meshy T2: Fast Native Mesh Generation with Flow Matching

    Aug 4, 2026·21m
  210. 210

    Weak-to-Strong On-Policy Distillation

    Aug 4, 2026·22m
  211. 211

    QQWorld: Quantile-Quantile Matching for World Model Regularization

    Aug 4, 2026·20m
  212. 212

    Scaling Properties of Text Conditioning in Visual Generation

    Aug 4, 2026·23m
  213. 213

    $N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

    Aug 4, 2026·23m
  214. 214

    AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

    Aug 4, 2026·19m
  215. 215

    Mental World Modeling

    Aug 4, 2026·21m
  216. 216

    From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

    Aug 4, 2026·20m
  217. 217

    PhiZero: A World Model Built Around Physical Language

    Aug 1, 2026·20m
  218. 218

    Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

    Aug 1, 2026·22m
  219. 219

    VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

    Aug 1, 2026·22m
  220. 220

    Beacon: Knowing When and How to Perform Agentic Visual Reasoning

    Aug 1, 2026·20m
  221. 221

    Flux-OPD: On-Policy Distillation with Evolving Contexts

    Aug 1, 2026·21m
  222. 222

    Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

    Aug 1, 2026·23m
  223. 223

    AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

    Aug 1, 2026·20m
  224. 224

    BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

    Aug 1, 2026·20m
  225. 225

    Metis: Memory Foundation Model

    Aug 1, 2026·19m
  226. 226

    Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

    Aug 1, 2026·22m
  227. 227

    CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

    Jul 31, 2026·21m
  228. 228

    TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    Jul 31, 2026·21m
  229. 229

    CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

    Jul 31, 2026·20m
  230. 230

    HumanCLAW: Can Vision-Language Models Act Through a Body?

    Jul 31, 2026·20m
  231. 231

    DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

    Jul 31, 2026·20m
  232. 232

    CAST: Game Solvers as Turn-Level Teachers for LLM Agents

    Jul 31, 2026·21m
  233. 233

    HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    Jul 30, 2026·19m
  234. 234

    A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

    Jul 30, 2026·20m
  235. 235

    CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents

    Jul 30, 2026·19m
  236. 236

    ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

    Jul 30, 2026·22m
  237. 237

    Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

    Jul 30, 2026·19m
  238. 238

    Pass the Baton: Trajectory-Relayed On-Policy Distillation

    Jul 30, 2026·20m
  239. 239

    Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

    Jul 29, 2026·22m
  240. 240

    Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

    Jul 29, 2026·22m
  241. 241

    OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

    Jul 29, 2026·20m
  242. 242

    From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

    Jul 29, 2026·21m
  243. 243

    Kimi K3: Open Frontier Intelligence

    Jul 29, 2026·21m
  244. 244

    Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

    Jul 29, 2026·23m
  245. 245

    Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

    Jul 29, 2026·21m
  246. 246

    JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

    Jul 29, 2026·20m
  247. 247

    StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

    Jul 29, 2026·21m
  248. 248

    Data Pyramid for Embodied Manipulation

    Jul 29, 2026·23m
  249. 249

    Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

    Jul 28, 2026·19m
  250. 250

    DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

    Jul 28, 2026·22m
  251. 251

    Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

    Jul 28, 2026·20m
  252. 252

    AREX: Towards a Recursively Self-Improving Agent for Deep Research

    Jul 25, 2026·20m
  253. 253

    ReferTrack: Referring Then Tracking for Embodied Visual Tracking

    Jul 25, 2026·21m
  254. 254

    Visual Contrastive Self-Distillation

    Jul 25, 2026·21m
  255. 255

    Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

    Jul 25, 2026·21m
  256. 256

    K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

    Jul 25, 2026·23m
  257. 257

    SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

    Jul 24, 2026·20m
  258. 258

    ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

    Jul 23, 2026·21m
  259. 259

    DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

    Jul 23, 2026·19m
  260. 260

    Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

    Jul 23, 2026·23m
  261. 261

    Generative World Renderer at the Speed of Play

    Jul 23, 2026·20m
  262. 262

    Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    Jul 23, 2026·20m
  263. 263

    AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

    Jul 23, 2026·21m
  264. 264

    AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

    Jul 23, 2026·21m
  265. 265

    DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

    Jul 22, 2026·20m
  266. 266

    TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

    Jul 22, 2026·17m
  267. 267

    EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

    Jul 22, 2026·21m
  268. 268

    Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

    Jul 22, 2026·22m
  269. 269

    SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

    Jul 22, 2026·20m
  270. 270

    RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

    Jul 22, 2026·23m
  271. 271

    HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

    Jul 22, 2026·19m
  272. 272

    Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

    Jul 22, 2026·21m
  273. 273

    RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

    Jul 21, 2026·20m
  274. 274

    Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    Jul 21, 2026·21m
  275. 275

    Loop the Loopies!

    Jul 21, 2026·20m
  276. 276

    xHC: Expanded Hyper-Connections

    Jul 21, 2026·21m
  277. 277

    RecGPT-V3 Technical Report

    Jul 21, 2026·22m
  278. 278

    Cura 1T: Specialized Model for Agentic Healthcare

    Jul 21, 2026·21m
  279. 279

    On-Policy Delta Distillation

    Jul 21, 2026·20m
  280. 280

    RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

    Jul 21, 2026·19m
  281. 281

    LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

    Jul 18, 2026·20m
  282. 282

    SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

    Jul 18, 2026·16m
  283. 283

    VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

    Jul 18, 2026·25m
  284. 284

    BadWAM: When World-Action Models Dream Right but Act Wrong

    Jul 18, 2026·21m
  285. 285

    KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

    Jul 18, 2026·22m
  286. 286

    MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

    Jul 18, 2026·21m
  287. 287

    From Pixels to States: Rethinking Interactive World Models as Game Engines

    Jul 18, 2026·18m
  288. 288

    UniVR: Thinking in Visual Space for Unified Visual Reasoning

    Jul 18, 2026·18m
  289. 289

    Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

    Jul 18, 2026·22m
  290. 290

    SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

    Jul 18, 2026·19m
  291. 291

    Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

    Jul 17, 2026·19m
  292. 292

    Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

    Jul 17, 2026·21m
  293. 293

    KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

    Jul 17, 2026·22m
  294. 294

    Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

    Jul 17, 2026·19m
  295. 295

    OvisOCR2 Technical Report

    Jul 17, 2026·21m
  296. 296

    PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

    Jul 17, 2026·20m
  297. 297

    MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

    Jul 17, 2026·19m
  298. 298

    GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

    Jul 17, 2026·21m
  299. 299

    SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

    Jul 16, 2026·21m
  300. 300

    Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

    Jul 16, 2026·19m
  301. 301

    Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

    Jul 16, 2026·19m
  302. 302

    Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

    Jul 16, 2026·20m
  303. 303

    Weak-to-Strong Generalization via Direct On-Policy Distillation

    Jul 15, 2026·20m
  304. 304

    ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

    Jul 15, 2026·22m
  305. 305

    4D Human-Scene Reconstruction from Low-Overlap Captures

    Jul 15, 2026·20m
  306. 306

    LightMem-Ego: Your AI Memory for Everyday Life

    Jul 15, 2026·19m
  307. 307

    Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

    Jul 14, 2026·20m
  308. 308

    Scalable Visual Pretraining for Language Intelligence

    Jul 14, 2026·18m
  309. 309

    Video Generation Models are General-Purpose Vision Learners

    Jul 14, 2026·20m
  310. 310

    Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

    Jul 11, 2026·26m
  311. 311

    Vidu S1: A Real-Time Interactive Video Generation Model

    Jul 11, 2026·23m
  312. 312

    Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

    Jul 11, 2026·25m
  313. 313

    UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

    Jul 11, 2026·26m
  314. 314

    Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

    Jul 10, 2026·27m
  315. 315

    Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

    Jul 10, 2026·25m
  316. 316

    Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

    Jul 10, 2026·22m
  317. 317

    Infinite Worlds with Versatile Interactions

    Jul 10, 2026·20m
  318. 318

    RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

    Jul 9, 2026·25m
  319. 319

    RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

    Jul 9, 2026·22m
  320. 320

    Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

    Jul 9, 2026·21m
  321. 321

    Vision as Unified Multimodal Generation

    Jul 9, 2026·26m
  322. 322

    Gemma 4 Technical Report

    Jul 9, 2026·24m
  323. 323

    EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots

    Jul 8, 2026·23m
  324. 324

    OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

    Jul 8, 2026·21m
  325. 325

    UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

    Jul 8, 2026·22m
  326. 326

    GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

    Jul 8, 2026·26m
  327. 327

    Wan-Streamer v0.2: Higher Resolution, Same Latency

    Jul 8, 2026·22m
  328. 328

    ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

    Jul 8, 2026·24m
  329. 329

    PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

    Jul 8, 2026·24m
  330. 330

    ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes

    Jul 8, 2026·27m
  331. 331

    MANCE: Manifold Aware Concept Erasure

    Jul 8, 2026·22m
  332. 332

    The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

    Jul 7, 2026·21m
  333. 333

    Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

    Jul 7, 2026·23m
  334. 334

    OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

    Jul 7, 2026·23m
  335. 335

    VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon

    Jul 7, 2026·20m
  336. 336

    DataComp-VLM: Improved Open Datasets for Vision-Language Models

    Jul 7, 2026·23m
  337. 337

    EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

    Jul 4, 2026·23m
  338. 338

    Morphing into Hybrid Attention Models

    Jul 4, 2026·25m
  339. 339

    AgenticDataBench: A Comprehensive Benchmark for Data Agents

    Jul 4, 2026·22m
  340. 340

    Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling

    Jul 4, 2026·22m
  341. 341

    Program-as-Weights: A Programming Paradigm for Fuzzy Functions

    Jul 4, 2026·24m
  342. 342

    AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

    Jul 4, 2026·24m
  343. 343

    PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

    Jul 3, 2026·22m
  344. 344

    Orca: The World is in Your Mind

    Jul 2, 2026·25m
  345. 345

    Dockerless: Environment-Free Program Verifier for Coding Agents

    Jul 2, 2026·24m
  346. 346

    Multi-Block Diffusion Language Models

    Jul 2, 2026·24m
  347. 347

    Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

    Jul 2, 2026·21m
  348. 348

    DOPD: Dual On-policy Distillation

    Jul 2, 2026·25m
  349. 349

    Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

    Jul 2, 2026·24m
  350. 350

    GEAR: Guided End-to-End AutoRegression for Image Synthesis

    Jul 2, 2026·27m
  351. 351

    TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

    Jul 1, 2026·23m
  352. 352

    Agentic Abstention: Do Agents Know When to Stop Instead of Act?

    Jul 1, 2026·24m
  353. 353

    LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

    Jul 1, 2026·23m
  354. 354

    Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

    Jul 1, 2026·26m
  355. 355

    Beyond IID: How General Are Tabular Foundation Models, Really?

    Jul 1, 2026·23m
  356. 356

    Trimming the Long-Tail of Visual World Modeling Evaluation

    Jul 1, 2026·26m
  357. 357

    Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

    Jul 1, 2026·21m
  358. 358

    Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

    Jul 1, 2026·23m
  359. 359

    AsyncOPD: How Stale Can On-Policy Distillation Be?

    Jul 1, 2026·23m
  360. 360

    PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

    Jun 30, 2026·21m
  361. 361

    Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

    Jun 30, 2026·24m
  362. 362

    Qwen-Image-2.0-RL Technical Report

    Jun 30, 2026·27m
  363. 363

    The Verification Horizon: No Silver Bullet for Coding Agent Rewards

    Jun 28, 2026·23m
  364. 364

    DanceOPD: On-Policy Generative Field Distillation

    Jun 28, 2026·26m
  365. 365

    In-Context World Modeling for Robotic Control

    Jun 28, 2026·23m
  366. 366

    Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

    Jun 28, 2026·21m
  367. 367

    OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

    Jun 28, 2026·24m
  368. 368

    ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

    Jun 28, 2026·26m
  369. 369

    Fast LeWorldModel

    Jun 28, 2026·25m
  370. 370

    GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents

    Jun 28, 2026·24m
  371. 371

    JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

    Jun 28, 2026·21m
  372. 372

    EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments

    Jun 13, 2026·24m
  373. 373

    MiniMax Sparse Attention

    Jun 13, 2026·26m
  374. 374

    SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

    Jun 13, 2026·23m
  375. 375

    InterleaveThinker: Reinforcing Agentic Interleaved Generation

    Jun 13, 2026·21m
  376. 376

    FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents

    Jun 13, 2026·23m
  377. 377

    Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

    Jun 13, 2026·20m
  378. 378

    LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

    Jun 13, 2026·23m
  379. 379

    WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

    Jun 13, 2026·21m
  380. 380

    HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

    Jun 13, 2026·21m
  381. 381

    MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

    Jun 13, 2026·24m
  382. 382

    ABot-Earth 0.5: Generative 3D Earth Model

    Jun 11, 2026·23m
  383. 383

    Kwai Keye-VL-2.0 Technical Report

    Jun 11, 2026·26m
  384. 384

    Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

    Jun 11, 2026·24m
  385. 385

    Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

    Jun 11, 2026·23m
  386. 386

    Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models

    Jun 11, 2026·22m
  387. 387

    Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference

    Jun 11, 2026·21m
  388. 388

    SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

    Jun 11, 2026·22m
  389. 389

    SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

    Jun 11, 2026·24m
  390. 390

    Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

    Jun 11, 2026·26m
  391. 391

    Agents' Last Exam

    Jun 10, 2026·25m
  392. 392

    SWE-Explore: Benchmarking How Coding Agents Explore Repositories

    Jun 10, 2026·23m
  393. 393

    On the Geometry of On-Policy Distillation

    Jun 10, 2026·27m
  394. 394

    Latent Spatial Memory for Video World Models

    Jun 10, 2026·25m
  395. 395

    Echo-Memory: A Controlled Study of Memory in Action World Models

    Jun 10, 2026·21m
  396. 396

    Human Psychometric Questionnaires Mischaracterize LLM Behavior

    Jun 10, 2026·25m
  397. 397

    LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

    Jun 10, 2026·22m
  398. 398

    FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

    Jun 10, 2026·22m
  399. 399

    CoVEBench: Can Video Editing Models Handle Complex Instructions?

    Jun 10, 2026·22m
  400. 400

    SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

    Jun 10, 2026·24m
  401. 401

    From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain

    Jun 4, 2026·23m
  402. 402

    Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

    Jun 4, 2026·26m
  403. 403

    Trust Region On-Policy Distillation

    Jun 4, 2026·24m
  404. 404

    KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks

    Jun 4, 2026·23m
  405. 405

    Mellum2 Technical Report

    Jun 2, 2026·22m
  406. 406

    Function2Scene: 3D Indoor Scene Layout from Functional Specifications

    Jun 2, 2026·22m
  407. 407

    Representation Forcing for Bottleneck-Free Unified Multimodal Models

    Jun 2, 2026·24m
  408. 408

    GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration

    Jun 2, 2026·23m
  409. 409

    COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

    Jun 2, 2026·21m
  410. 410

    Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

    Jun 2, 2026·26m
  411. 411

    Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

    May 23, 2026·24m
  412. 412

    DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards

    May 23, 2026·21m
  413. 413

    TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation

    May 23, 2026·23m
  414. 414

    $π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

    May 23, 2026·23m
  415. 415

    Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

    May 23, 2026·19m
  416. 416

    ACC: Compiling Agent Trajectories for Long-Context Training

    May 23, 2026·25m
  417. 417

    PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects

    May 23, 2026·23m
  418. 418

    Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning

    May 23, 2026·22m
  419. 419

    WorldKV: Efficient World Memory with World Retrieval and Compression

    May 23, 2026·23m
  420. 420

    LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

    May 23, 2026·22m
  421. 421

    Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

    May 22, 2026·20m
  422. 422

    Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

    May 22, 2026·23m
  423. 423

    Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

    May 22, 2026·25m
  424. 424

    IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools

    May 22, 2026·24m
  425. 425

    AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

    May 21, 2026·24m
  426. 426

    OpenComputer: Verifiable Software Worlds for Computer-Use Agents

    May 21, 2026·25m
  427. 427

    GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment

    May 21, 2026·25m
  428. 428

    When Vision Speaks for Sound

    May 21, 2026·23m
  429. 429

    Active Learners as Efficient PRP Rerankers

    May 21, 2026·24m
  430. 430

    Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

    May 21, 2026·23m
  431. 431

    Process Rewards with Learned Reliability

    May 21, 2026·23m
  432. 432

    EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL

    May 21, 2026·27m
  433. 433

    CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition

    May 21, 2026·23m
  434. 434

    Harnessing LLM Agents with Skill Programs

    May 21, 2026·22m
  435. 435

    SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution

    May 20, 2026·23m
  436. 436

    LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

    May 20, 2026·22m
  437. 437

    Lance: Unified Multimodal Modeling by Multi-Task Synergy

    May 20, 2026·23m
  438. 438

    Code as Agent Harness

    May 20, 2026·25m
  439. 439

    AI for Auto-Research: Roadmap & User Guide

    May 20, 2026·22m
  440. 440

    CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

    May 20, 2026·23m
  441. 441

    KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

    May 20, 2026·24m
  442. 442

    MMSkills: Towards Multimodal Skills for General Visual Agents

    May 19, 2026·23m
  443. 443

    DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo

    May 19, 2026·23m
  444. 444

    CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

    May 19, 2026·23m
  445. 445

    PhysBrain 1.0 Technical Report

    May 19, 2026·25m
  446. 446

    Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding

    May 19, 2026·21m
  447. 447

    InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation

    May 19, 2026·24m
  448. 448

    Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR

    May 19, 2026·21m
  449. 449

    Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

    May 19, 2026·21m
  450. 450

    Self-Distilled Agentic Reinforcement Learning

    May 16, 2026·25m
  451. 451

    MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

    May 16, 2026·27m
  452. 452

    Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

    May 16, 2026·24m
  453. 453

    SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

    May 16, 2026·22m
  454. 454

    MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

    May 16, 2026·23m
  455. 455

    Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems

    May 16, 2026·22m
  456. 456

    STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?

    May 16, 2026·24m
  457. 457

    WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

    May 16, 2026·25m
  458. 458

    Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning

    May 16, 2026·23m
  459. 459

    Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling

    May 16, 2026·23m
  460. 460

    MinT: Managed Infrastructure for Training and Serving Millions of LLMs

    May 15, 2026·24m
  461. 461

    MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

    May 15, 2026·24m
  462. 462

    AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

    May 15, 2026·24m
  463. 463

    Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

    May 15, 2026·23m
  464. 464

    Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling

    May 15, 2026·25m
  465. 465

    EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

    May 15, 2026·25m
  466. 466

    Many-Shot CoT-ICL: Making In-Context Learning Truly Learn

    May 15, 2026·24m
  467. 467

    Qwen-Image-VAE-2.0 Technical Report

    May 15, 2026·24m
  468. 468

    TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

    May 15, 2026·23m
  469. 469

    Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

    May 15, 2026·24m
  470. 470

    World Action Models: The Next Frontier in Embodied AI

    May 14, 2026·25m
  471. 471

    Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization

    May 14, 2026·25m
  472. 472

    Efficient Pre-Training with Token Superposition

    May 14, 2026·25m
  473. 473

    RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

    May 14, 2026·23m
  474. 474

    Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics

    May 14, 2026·23m
  475. 475

    AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward

    May 14, 2026·24m
  476. 476

    MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

    May 14, 2026·24m
  477. 477

    SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

    May 14, 2026·26m
  478. 478

    MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

    May 14, 2026·22m
  479. 479

    $δ$-mem: Efficient Online Memory for Large Language Models

    May 14, 2026·25m
  480. 480

    Qwen-Image-2.0 Technical Report

    May 13, 2026·23m
  481. 481

    TMAS: Scaling Test-Time Compute via Multi-Agent Synergy

    May 13, 2026·23m
  482. 482

    PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents

    May 13, 2026·23m
  483. 483

    CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

    May 13, 2026·25m
  484. 484

    Model Merging Scaling Laws in Large Language Models

    May 13, 2026·22m
  485. 485

    Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

    May 13, 2026·24m
  486. 486

    SEIF: Self-Evolving Reinforcement Learning for Instruction Following

    May 13, 2026·21m
  487. 487

    WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

    May 13, 2026·22m
  488. 488

    Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models

    May 13, 2026·22m
  489. 489

    Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers

    May 12, 2026·22m
  490. 490

    Flow-OPD: On-Policy Distillation for Flow Matching Models

    May 12, 2026·26m
  491. 491

    Beyond Retrieval: A Multitask Benchmark and Model for Code Search

    May 12, 2026·21m
  492. 492

    HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

    May 12, 2026·25m
  493. 493

    Anisotropic Modality Align

    May 12, 2026·23m
  494. 494

    MiA-Signature: Approximating Global Activation for Long-Context Understanding

    May 9, 2026·12m
  495. 495

    When to Trust Imagination: Adaptive Action Execution for World Action Models

    May 9, 2026·12m
  496. 496

    Continuous-Time Distribution Matching for Few-Step Diffusion Distillation

    May 9, 2026·15m
  497. 497

    Stream-T1: Test-Time Scaling for Streaming Video Generation

    May 8, 2026·22m
  498. 498

    RLDX-1 Technical Report

    May 8, 2026·23m
  499. 499

    Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

    May 8, 2026·22m
  500. 500

    OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

    May 8, 2026·25m