
Daily Paper Cast
Jingwen Liang, Gengyu Wang
500 episodes listed below
AI papers, structured for fast technical listening.
The gist
Daily Paper Cast is a fast-moving research explainer built around papers surfaced from Hugging Face's Daily Paper list. Echo and Nova move through introductions, methods, experiments, and related work with a steady focus on models, metrics, and implementation details.
PRESS PLAY
Find your next episode
About Daily Paper Cast
Daily Paper Cast converts current AI research papers into structured audio briefings. The show's recurring shape is clear: introduce the paper, explain the motivation, walk through the method, summarize the experiments, and close with related work or future directions. Echo and Nova carry the discussion as complementary synthetic voices. One tends to ask the clarifying question, while the other fills in the technical machinery, and both keep the pace moving through dense material. Recent episodes range across multimodal world models, autonomous driving reinforcement learning, deep research agents, occupational task simulation, visual reward modeling, coding-agent memory transfer, video generation, and game-agent evaluation. That breadth gives the show a strong survey function. It is less about hot takes and more about turning a paper's architecture, benchmark design, and result tables into a listenable sequence. The vocabulary stays technical. Terms such as CLIP alignment, closed-loop evaluation, RAG sub-agents, robustness scores, reward models, inference optimization, and progress metrics are treated as normal working language. The show does, however, keep checking comprehension by restating concepts after the first pass. It also pays unusual attention to evaluation. Episodes repeatedly ask which baselines were used, what metrics counted, how many scenarios or tasks were tested, and which limitations remained after the reported gains. This makes it useful for researchers, builders, and technically fluent listeners who care about whether a method is merely impressive or actually measured well. The tone is measured, tidy, and faintly enthusiastic. There are occasional synthetic-audio rough edges and repeated phrases in the excerpts, but the editorial center remains consistent. Daily Paper Cast works best as a daily research companion for people who already know the field's vocabulary and want a fast map of what new papers are claiming. It gives the listener a structured first read before deciding whether the full paper deserves closer attention.
Made for: Daily Paper Cast is for machine learning practitioners, AI researchers, technical founders, and advanced students who want quick orientation to new papers. It assumes comfort with model names, benchmarks, metrics, and research jargon.
What sets it apart: The show keeps the paper's own structure intact, moving section by section through methods and experiments instead of turning research into loose commentary. Its synthetic duo format makes dense benchmark-heavy material feel more conversational without stripping out the technical substance.
In their own words
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: [email protected] Creator: Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/ Gengyu Wang, LLM ML, http://wanggengyu.com Listen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXL Apple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-cast/id1777620236 Cover Image by Kawen Kuang https://kawen.art
As heard by us
Based on 6 episodes we listened to · September 2026
Daily Paper Cast makes current AI papers legible by pairing accessible explanation with close attention to methods, benchmarks, results, and limitations.
Daily Paper Cast turns individual machine-learning papers into structured conversations between Echo and Nova, moving through the research question, methods, experiments, and related work.
Why you'd press play
Echo and Nova turn research-paper methods and experiments into the lab meeting you missed.
Press play if you want
- methods, benchmarks, and ablations translated into conversational checkpoints instead of one dense monologue
- one paper at a time, with enough structure to keep the acronyms from staging a coup
Talks about
Podcasts like Daily Paper Cast


Real Vision: Finance & InvestingReal Vision Podcast NetworkMarkets, macro, AI, and crypto in motion.


AI DailyAmy IversonDaily AI news through the lens of trust, infrastructure, and real-world deployment.
Episodes
- 1
ROWBench: Do Video Models Render What the Program Specifies?
Oct 2, 2026·22m - 2
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Oct 2, 2026·22m - 3
Agent Priors-guided Policy Learning
Oct 2, 2026·21m - 4
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Oct 2, 2026·22m - 5
Sharpening Tax in Post-Training
Oct 2, 2026·21m - 6
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL
Oct 2, 2026·20m - 7
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
Oct 2, 2026·21m - 8
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States
Oct 2, 2026·22m - 9
World Observer: Joint Actor-Observer Generation for Persistent World Modeling
Oct 2, 2026·20m - 10
Hierarchical Continuous Diffusion Language Models
Oct 2, 2026·21m - 11
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Oct 1, 2026·22m - 12
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models
Oct 1, 2026·22m - 13
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
Oct 1, 2026·22m - 14
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
Oct 1, 2026·23m - 15
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Oct 1, 2026·20m - 16
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Oct 1, 2026·24m - 17
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
Oct 1, 2026·22m - 18
EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
Oct 1, 2026·21m - 19
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
Oct 1, 2026·22m - 20
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
Sep 30, 2026·21m - 21
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Sep 30, 2026·21m - 22
Think Before You Score: Thinking Reward Model for Visual Generation
Sep 30, 2026·22m - 23
MaLiang-Harness: A Programmable Path to Image and Video Generation
Sep 30, 2026·23m - 24
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
Sep 30, 2026·21m - 25
Omni-IO Skills: Harnessing Your Agent Omni-Native
Sep 30, 2026·22m - 26
Raven: The Harness of Harnesses for Composable Agentic Intelligence
Sep 30, 2026·26m - 27
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
Sep 30, 2026·22m - 28
PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
Sep 30, 2026·21m - 29
In-Context Learning for Robots: Methods and Applications
Sep 30, 2026·21m - 30
Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence
Sep 29, 2026·22m - 31
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
Sep 29, 2026·18m - 32
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Sep 29, 2026·21m - 33
Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Sep 29, 2026·19m - 34
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
Sep 29, 2026·21m - 35
Improving Test-Time Scaling with Adaptive Looped Transformers
Sep 29, 2026·22m - 36
CompoWorld: Compositional Environment Scaling for General Agents
Sep 29, 2026·23m - 37
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Sep 29, 2026·22m - 38
How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Sep 29, 2026·19m - 39
EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
Sep 29, 2026·24m - 40
Training Object Permanence in World Models
Sep 25, 2026·23m - 41
The Past Frames the Future: Memory for Autoregressive Video Generation
Sep 24, 2026·17m - 42
HappyWorld-Bench
Sep 24, 2026·24m - 43
RULER: Instance-aware Rubric Rewards for SVG Generation
Sep 23, 2026·21m - 44
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
Sep 23, 2026·21m - 45
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Sep 23, 2026·20m - 46
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Sep 22, 2026·24m - 47
OmniEdu: Open Foundation Models for Learning and Teaching
Sep 22, 2026·24m - 48
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Sep 22, 2026·22m - 49
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Sep 22, 2026·20m - 50
Transferring the Intelligence of VLMs to Robotic Control
Sep 22, 2026·22m - 51
VideoGen-Agent: Reinforcing Video Generation Agents
Sep 22, 2026·26m - 52
An Empirical Study of Harness Design for Coding Agents
Sep 18, 2026·24m - 53
JEPA-Anything: Learning Predictive Models across Different Worlds
Sep 18, 2026·20m - 54
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
Sep 18, 2026·23m - 55
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
Sep 18, 2026·23m - 56
A Zeroth-Order Paradigm for LLM Preference Alignment
Sep 17, 2026·21m - 57
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
Sep 17, 2026·23m - 58
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
Sep 17, 2026·20m - 59
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention
Sep 17, 2026·23m - 60
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
Sep 17, 2026·22m - 61
Agora: Git as Shared Memory for Collective AutoResearch
Sep 17, 2026·22m - 62
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Sep 17, 2026·23m - 63
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
Sep 17, 2026·21m - 64
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
Sep 17, 2026·22m - 65
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
Sep 17, 2026·21m - 66
Continual Learning Mechanisms Compose for Long-Horizon Memorization
Sep 16, 2026·23m - 67
StepAudio 3 Realtime Technical Report
Sep 16, 2026·23m - 68
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
Sep 16, 2026·22m - 69
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
Sep 15, 2026·22m - 70
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Sep 15, 2026·21m - 71
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Sep 15, 2026·20m - 72
Atria Dawn: The Dawn of Agentic Superintelligence
Sep 15, 2026·21m - 73
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
Sep 15, 2026·21m - 74
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
Sep 15, 2026·23m - 75
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Sep 11, 2026·21m - 76
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
Sep 11, 2026·22m - 77
SenseNova-U1.5: Towards Native Unified Visual Intelligence
Sep 11, 2026·22m - 78
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
Sep 11, 2026·24m - 79
Show-Harness: Just a VLM Agent Can Play Robots
Sep 10, 2026·21m - 80
Programmable World Model
Sep 10, 2026·19m - 81
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Sep 10, 2026·22m - 82
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Sep 9, 2026·21m - 83
DriveZero: End-to-End Driving Beyond Human Demonstrations
Sep 9, 2026·22m - 84
SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution
Sep 9, 2026·18m - 85
Miles v0.1: Production-Level Post-Training
Sep 9, 2026·21m - 86
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
Sep 9, 2026·21m - 87
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Sep 9, 2026·20m - 88
Omni Interaction Agent Technical Report
Sep 9, 2026·22m - 89
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation
Sep 9, 2026·20m - 90
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Sep 9, 2026·20m - 91
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Sep 4, 2026·22m - 92
LatentPress: Context Compression Beyond Text and Vision
Sep 4, 2026·20m - 93
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
Sep 4, 2026·21m - 94
Rethinking On-Policy Distillation of Large Language Models II: One Training Example
Sep 4, 2026·22m - 95
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
Sep 4, 2026·23m - 96
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
Sep 4, 2026·20m - 97
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
Sep 4, 2026·22m - 98
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
Sep 3, 2026·21m - 99
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning
Sep 3, 2026·21m - 100
Language Models Can Control Their Own Attention
Sep 3, 2026·21m - 101
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
Sep 3, 2026·21m - 102
On the Design Fundamentals of Pixel Text Representation Learning
Sep 3, 2026·20m - 103
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Sep 3, 2026·21m - 104
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
Sep 2, 2026·21m - 105
Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering
Sep 2, 2026·20m - 106
H3-World: Turning Language Understanding into World Control
Sep 2, 2026·21m - 107
StudentSim: Training LLM-based Student Simulators
Sep 2, 2026·21m - 108
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
Sep 2, 2026·21m - 109
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Sep 2, 2026·21m - 110
UI-Venus-2 Technical Report
Sep 2, 2026·23m - 111
Normalized Low-Rank Adaptation
Sep 1, 2026·19m - 112
CogEvol: Towards Efficient and Reliable Learning Environment Generation
Sep 1, 2026·18m - 113
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
Sep 1, 2026·26m - 114
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
Sep 1, 2026·22m - 115
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
Sep 1, 2026·21m - 116
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
Sep 1, 2026·22m - 117
PaperGym: Rubric-Centered Evolution for Research-Plan Generation
Sep 1, 2026·22m - 118
GameWAM: A World Action Model for Video Games
Aug 28, 2026·23m - 119
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
Aug 28, 2026·21m - 120
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Aug 28, 2026·22m - 121
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
Aug 28, 2026·21m - 122
TTPO: Test-Time Policy Optimization
Aug 28, 2026·21m - 123
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
Aug 28, 2026·20m - 124
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Aug 28, 2026·23m - 125
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
Aug 28, 2026·19m - 126
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Aug 28, 2026·21m - 127
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Aug 27, 2026·22m - 128
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Aug 27, 2026·20m - 129
VGI-Bench: Probing Visual Intelligence in Video Generation Models
Aug 27, 2026·19m - 130
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation
Aug 27, 2026·20m - 131
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
Aug 27, 2026·20m - 132
FrontierChallenge: Evaluating Scientific Workflow Completion
Aug 27, 2026·25m - 133
Towards a Densing Law for User Representation Learning at Billion-Scale Capacity
Aug 26, 2026·19m - 134
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Aug 25, 2026·18m - 135
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
Aug 19, 2026·19m - 136
Agentic Transaction: Towards ACID-Compliant Agent Systems
Aug 19, 2026·20m - 137
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Aug 18, 2026·25m - 138
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Aug 18, 2026·20m - 139
Self-Supervised Visual On-Policy Distillation
Aug 18, 2026·22m - 140
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Aug 18, 2026·23m - 141
Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Aug 18, 2026·21m - 142
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Aug 18, 2026·17m - 143
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Aug 18, 2026·22m - 144
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
Aug 15, 2026·22m - 145
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Aug 15, 2026·20m - 146
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Aug 15, 2026·20m - 147
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Aug 15, 2026·20m - 148
Intern-S2-Preview: Scientific Agentic Foundation Model
Aug 15, 2026·23m - 149
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Aug 15, 2026·22m - 150
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Aug 15, 2026·22m - 151
DarwinX: Evolving Agent Harnesses Through Natural Selection
Aug 15, 2026·21m - 152
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
Aug 14, 2026·23m - 153
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Aug 14, 2026·20m - 154
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Aug 14, 2026·22m - 155
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Aug 14, 2026·23m - 156
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
Aug 14, 2026·19m - 157
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Aug 14, 2026·22m - 158
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Aug 14, 2026·22m - 159
Articulated Object Reconstruction from Rest-State Observation
Aug 13, 2026·20m - 160
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Aug 13, 2026·20m - 161
Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Aug 13, 2026·22m - 162
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Aug 13, 2026·23m - 163
Beyond Pixels: From Video Priors to 4D Worlds
Aug 13, 2026·22m - 164
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
Aug 13, 2026·19m - 165
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Aug 12, 2026·23m - 166
Motif 3: Technical Report
Aug 12, 2026·24m - 167
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
Aug 12, 2026·19m - 168
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Aug 12, 2026·22m - 169
On-Policy Self-Distillation without Any Supervision
Aug 12, 2026·21m - 170
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
Aug 12, 2026·20m - 171
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
Aug 12, 2026·20m - 172
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Aug 12, 2026·21m - 173
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
Aug 12, 2026·22m - 174
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
Aug 11, 2026·21m - 175
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning
Aug 11, 2026·21m - 176
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
Aug 11, 2026·22m - 177
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
Aug 8, 2026·21m - 178
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Aug 8, 2026·21m - 179
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Aug 8, 2026·20m - 180
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
Aug 8, 2026·20m - 181
ChronoVision: Temporal Reasoning via Latent State Reconstruction
Aug 8, 2026·20m - 182
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
Aug 8, 2026·22m - 183
WorldClaw: Agentic 3D Open-World Generation at Scale
Aug 8, 2026·22m - 184
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
Aug 8, 2026·19m - 185
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
Aug 7, 2026·21m - 186
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Aug 7, 2026·24m - 187
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Aug 7, 2026·20m - 188
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Aug 7, 2026·19m - 189
Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
Aug 6, 2026·23m - 190
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
Aug 6, 2026·22m - 191
Quo Vadis, World Modeling?
Aug 6, 2026·21m - 192
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Aug 6, 2026·22m - 193
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
Aug 6, 2026·22m - 194
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Aug 6, 2026·22m - 195
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Aug 6, 2026·21m - 196
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
Aug 6, 2026·20m - 197
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Aug 6, 2026·18m - 198
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Aug 5, 2026·20m - 199
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Aug 5, 2026·23m - 200
DAPD: Dual-Anchored Policy Distillation
Aug 5, 2026·20m - 201
Progressive Agent Skill Generation via Reinforcement Learning
Aug 5, 2026·22m - 202
UEmbed: Unified Sparse and Dense Multimodal Embeddings
Aug 5, 2026·20m - 203
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Aug 5, 2026·22m - 204
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Aug 5, 2026·23m - 205
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Aug 5, 2026·21m - 206
CADENA: Stepwise CAD Reverse Engineering
Aug 5, 2026·23m - 207
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis
Aug 5, 2026·24m - 208
$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
Aug 4, 2026·20m - 209
Meshy T2: Fast Native Mesh Generation with Flow Matching
Aug 4, 2026·21m - 210
Weak-to-Strong On-Policy Distillation
Aug 4, 2026·22m - 211
QQWorld: Quantile-Quantile Matching for World Model Regularization
Aug 4, 2026·20m - 212
Scaling Properties of Text Conditioning in Visual Generation
Aug 4, 2026·23m - 213
$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
Aug 4, 2026·23m - 214
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Aug 4, 2026·19m - 215
Mental World Modeling
Aug 4, 2026·21m - 216
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Aug 4, 2026·20m - 217
PhiZero: A World Model Built Around Physical Language
Aug 1, 2026·20m - 218
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Aug 1, 2026·22m - 219
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Aug 1, 2026·22m - 220
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
Aug 1, 2026·20m - 221
Flux-OPD: On-Policy Distillation with Evolving Contexts
Aug 1, 2026·21m - 222
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Aug 1, 2026·23m - 223
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Aug 1, 2026·20m - 224
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
Aug 1, 2026·20m - 225
Metis: Memory Foundation Model
Aug 1, 2026·19m - 226
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
Aug 1, 2026·22m - 227
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Jul 31, 2026·21m - 228
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Jul 31, 2026·21m - 229
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Jul 31, 2026·20m - 230
HumanCLAW: Can Vision-Language Models Act Through a Body?
Jul 31, 2026·20m - 231
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Jul 31, 2026·20m - 232
CAST: Game Solvers as Turn-Level Teachers for LLM Agents
Jul 31, 2026·21m - 233
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Jul 30, 2026·19m - 234
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
Jul 30, 2026·20m - 235
CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents
Jul 30, 2026·19m - 236
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
Jul 30, 2026·22m - 237
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
Jul 30, 2026·19m - 238
Pass the Baton: Trajectory-Relayed On-Policy Distillation
Jul 30, 2026·20m - 239
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
Jul 29, 2026·22m - 240
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
Jul 29, 2026·22m - 241
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jul 29, 2026·20m - 242
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
Jul 29, 2026·21m - 243
Kimi K3: Open Frontier Intelligence
Jul 29, 2026·21m - 244
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
Jul 29, 2026·23m - 245
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
Jul 29, 2026·21m - 246
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Jul 29, 2026·20m - 247
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
Jul 29, 2026·21m - 248
Data Pyramid for Embodied Manipulation
Jul 29, 2026·23m - 249
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
Jul 28, 2026·19m - 250
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
Jul 28, 2026·22m - 251
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Jul 28, 2026·20m - 252
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Jul 25, 2026·20m - 253
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
Jul 25, 2026·21m - 254
Visual Contrastive Self-Distillation
Jul 25, 2026·21m - 255
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
Jul 25, 2026·21m - 256
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Jul 25, 2026·23m - 257
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Jul 24, 2026·20m - 258
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Jul 23, 2026·21m - 259
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Jul 23, 2026·19m - 260
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
Jul 23, 2026·23m - 261
Generative World Renderer at the Speed of Play
Jul 23, 2026·20m - 262
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Jul 23, 2026·20m - 263
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
Jul 23, 2026·21m - 264
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
Jul 23, 2026·21m - 265
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
Jul 22, 2026·20m - 266
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Jul 22, 2026·17m - 267
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
Jul 22, 2026·21m - 268
Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Jul 22, 2026·22m - 269
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
Jul 22, 2026·20m - 270
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
Jul 22, 2026·23m - 271
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
Jul 22, 2026·19m - 272
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Jul 22, 2026·21m - 273
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
Jul 21, 2026·20m - 274
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Jul 21, 2026·21m - 275
Loop the Loopies!
Jul 21, 2026·20m - 276
xHC: Expanded Hyper-Connections
Jul 21, 2026·21m - 277
RecGPT-V3 Technical Report
Jul 21, 2026·22m - 278
Cura 1T: Specialized Model for Agentic Healthcare
Jul 21, 2026·21m - 279
On-Policy Delta Distillation
Jul 21, 2026·20m - 280
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Jul 21, 2026·19m - 281
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Jul 18, 2026·20m - 282
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Jul 18, 2026·16m - 283
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Jul 18, 2026·25m - 284
BadWAM: When World-Action Models Dream Right but Act Wrong
Jul 18, 2026·21m - 285
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Jul 18, 2026·22m - 286
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Jul 18, 2026·21m - 287
From Pixels to States: Rethinking Interactive World Models as Game Engines
Jul 18, 2026·18m - 288
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Jul 18, 2026·18m - 289
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
Jul 18, 2026·22m - 290
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Jul 18, 2026·19m - 291
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Jul 17, 2026·19m - 292
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Jul 17, 2026·21m - 293
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
Jul 17, 2026·22m - 294
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
Jul 17, 2026·19m - 295
OvisOCR2 Technical Report
Jul 17, 2026·21m - 296
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Jul 17, 2026·20m - 297
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
Jul 17, 2026·19m - 298
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
Jul 17, 2026·21m - 299
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
Jul 16, 2026·21m - 300
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
Jul 16, 2026·19m - 301
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Jul 16, 2026·19m - 302
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Jul 16, 2026·20m - 303
Weak-to-Strong Generalization via Direct On-Policy Distillation
Jul 15, 2026·20m - 304
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Jul 15, 2026·22m - 305
4D Human-Scene Reconstruction from Low-Overlap Captures
Jul 15, 2026·20m - 306
LightMem-Ego: Your AI Memory for Everyday Life
Jul 15, 2026·19m - 307
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Jul 14, 2026·20m - 308
Scalable Visual Pretraining for Language Intelligence
Jul 14, 2026·18m - 309
Video Generation Models are General-Purpose Vision Learners
Jul 14, 2026·20m - 310
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
Jul 11, 2026·26m - 311
Vidu S1: A Real-Time Interactive Video Generation Model
Jul 11, 2026·23m - 312
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Jul 11, 2026·25m - 313
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
Jul 11, 2026·26m - 314
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Jul 10, 2026·27m - 315
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
Jul 10, 2026·25m - 316
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Jul 10, 2026·22m - 317
Infinite Worlds with Versatile Interactions
Jul 10, 2026·20m - 318
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
Jul 9, 2026·25m - 319
RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
Jul 9, 2026·22m - 320
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Jul 9, 2026·21m - 321
Vision as Unified Multimodal Generation
Jul 9, 2026·26m - 322
Gemma 4 Technical Report
Jul 9, 2026·24m - 323
EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots
Jul 8, 2026·23m - 324
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Jul 8, 2026·21m - 325
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
Jul 8, 2026·22m - 326
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
Jul 8, 2026·26m - 327
Wan-Streamer v0.2: Higher Resolution, Same Latency
Jul 8, 2026·22m - 328
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Jul 8, 2026·24m - 329
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
Jul 8, 2026·24m - 330
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Jul 8, 2026·27m - 331
MANCE: Manifold Aware Concept Erasure
Jul 8, 2026·22m - 332
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
Jul 7, 2026·21m - 333
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
Jul 7, 2026·23m - 334
OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
Jul 7, 2026·23m - 335
VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
Jul 7, 2026·20m - 336
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Jul 7, 2026·23m - 337
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
Jul 4, 2026·23m - 338
Morphing into Hybrid Attention Models
Jul 4, 2026·25m - 339
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Jul 4, 2026·22m - 340
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
Jul 4, 2026·22m - 341
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Jul 4, 2026·24m - 342
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
Jul 4, 2026·24m - 343
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
Jul 3, 2026·22m - 344
Orca: The World is in Your Mind
Jul 2, 2026·25m - 345
Dockerless: Environment-Free Program Verifier for Coding Agents
Jul 2, 2026·24m - 346
Multi-Block Diffusion Language Models
Jul 2, 2026·24m - 347
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
Jul 2, 2026·21m - 348
DOPD: Dual On-policy Distillation
Jul 2, 2026·25m - 349
Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views
Jul 2, 2026·24m - 350
GEAR: Guided End-to-End AutoRegression for Image Synthesis
Jul 2, 2026·27m - 351
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
Jul 1, 2026·23m - 352
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
Jul 1, 2026·24m - 353
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
Jul 1, 2026·23m - 354
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
Jul 1, 2026·26m - 355
Beyond IID: How General Are Tabular Foundation Models, Really?
Jul 1, 2026·23m - 356
Trimming the Long-Tail of Visual World Modeling Evaluation
Jul 1, 2026·26m - 357
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Jul 1, 2026·21m - 358
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Jul 1, 2026·23m - 359
AsyncOPD: How Stale Can On-Policy Distillation Be?
Jul 1, 2026·23m - 360
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
Jun 30, 2026·21m - 361
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
Jun 30, 2026·24m - 362
Qwen-Image-2.0-RL Technical Report
Jun 30, 2026·27m - 363
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
Jun 28, 2026·23m - 364
DanceOPD: On-Policy Generative Field Distillation
Jun 28, 2026·26m - 365
In-Context World Modeling for Robotic Control
Jun 28, 2026·23m - 366
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Jun 28, 2026·21m - 367
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
Jun 28, 2026·24m - 368
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Jun 28, 2026·26m - 369
Fast LeWorldModel
Jun 28, 2026·25m - 370
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
Jun 28, 2026·24m - 371
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
Jun 28, 2026·21m - 372
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
Jun 13, 2026·24m - 373
MiniMax Sparse Attention
Jun 13, 2026·26m - 374
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Jun 13, 2026·23m - 375
InterleaveThinker: Reinforcing Agentic Interleaved Generation
Jun 13, 2026·21m - 376
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
Jun 13, 2026·23m - 377
Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?
Jun 13, 2026·20m - 378
LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
Jun 13, 2026·23m - 379
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
Jun 13, 2026·21m - 380
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
Jun 13, 2026·21m - 381
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
Jun 13, 2026·24m - 382
ABot-Earth 0.5: Generative 3D Earth Model
Jun 11, 2026·23m - 383
Kwai Keye-VL-2.0 Technical Report
Jun 11, 2026·26m - 384
Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization
Jun 11, 2026·24m - 385
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Jun 11, 2026·23m - 386
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Jun 11, 2026·22m - 387
Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference
Jun 11, 2026·21m - 388
SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning
Jun 11, 2026·22m - 389
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
Jun 11, 2026·24m - 390
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Jun 11, 2026·26m - 391
Agents' Last Exam
Jun 10, 2026·25m - 392
SWE-Explore: Benchmarking How Coding Agents Explore Repositories
Jun 10, 2026·23m - 393
On the Geometry of On-Policy Distillation
Jun 10, 2026·27m - 394
Latent Spatial Memory for Video World Models
Jun 10, 2026·25m - 395
Echo-Memory: A Controlled Study of Memory in Action World Models
Jun 10, 2026·21m - 396
Human Psychometric Questionnaires Mischaracterize LLM Behavior
Jun 10, 2026·25m - 397
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
Jun 10, 2026·22m - 398
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Jun 10, 2026·22m - 399
CoVEBench: Can Video Editing Models Handle Complex Instructions?
Jun 10, 2026·22m - 400
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Jun 10, 2026·24m - 401
From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain
Jun 4, 2026·23m - 402
Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
Jun 4, 2026·26m - 403
Trust Region On-Policy Distillation
Jun 4, 2026·24m - 404
KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks
Jun 4, 2026·23m - 405
Mellum2 Technical Report
Jun 2, 2026·22m - 406
Function2Scene: 3D Indoor Scene Layout from Functional Specifications
Jun 2, 2026·22m - 407
Representation Forcing for Bottleneck-Free Unified Multimodal Models
Jun 2, 2026·24m - 408
GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration
Jun 2, 2026·23m - 409
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
Jun 2, 2026·21m - 410
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
Jun 2, 2026·26m - 411
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
May 23, 2026·24m - 412
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
May 23, 2026·21m - 413
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
May 23, 2026·23m - 414
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
May 23, 2026·23m - 415
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
May 23, 2026·19m - 416
ACC: Compiling Agent Trajectories for Long-Context Training
May 23, 2026·25m - 417
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
May 23, 2026·23m - 418
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
May 23, 2026·22m - 419
WorldKV: Efficient World Memory with World Retrieval and Compression
May 23, 2026·23m - 420
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
May 23, 2026·22m - 421
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
May 22, 2026·20m - 422
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
May 22, 2026·23m - 423
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
May 22, 2026·25m - 424
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
May 22, 2026·24m - 425
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
May 21, 2026·24m - 426
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
May 21, 2026·25m - 427
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
May 21, 2026·25m - 428
When Vision Speaks for Sound
May 21, 2026·23m - 429
Active Learners as Efficient PRP Rerankers
May 21, 2026·24m - 430
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
May 21, 2026·23m - 431
Process Rewards with Learned Reliability
May 21, 2026·23m - 432
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
May 21, 2026·27m - 433
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
May 21, 2026·23m - 434
Harnessing LLM Agents with Skill Programs
May 21, 2026·22m - 435
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
May 20, 2026·23m - 436
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
May 20, 2026·22m - 437
Lance: Unified Multimodal Modeling by Multi-Task Synergy
May 20, 2026·23m - 438
Code as Agent Harness
May 20, 2026·25m - 439
AI for Auto-Research: Roadmap & User Guide
May 20, 2026·22m - 440
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
May 20, 2026·23m - 441
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
May 20, 2026·24m - 442
MMSkills: Towards Multimodal Skills for General Visual Agents
May 19, 2026·23m - 443
DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo
May 19, 2026·23m - 444
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
May 19, 2026·23m - 445
PhysBrain 1.0 Technical Report
May 19, 2026·25m - 446
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
May 19, 2026·21m - 447
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
May 19, 2026·24m - 448
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
May 19, 2026·21m - 449
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
May 19, 2026·21m - 450
Self-Distilled Agentic Reinforcement Learning
May 16, 2026·25m - 451
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
May 16, 2026·27m - 452
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
May 16, 2026·24m - 453
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
May 16, 2026·22m - 454
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
May 16, 2026·23m - 455
Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems
May 16, 2026·22m - 456
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
May 16, 2026·24m - 457
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
May 16, 2026·25m - 458
Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning
May 16, 2026·23m - 459
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
May 16, 2026·23m - 460
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
May 15, 2026·24m - 461
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
May 15, 2026·24m - 462
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
May 15, 2026·24m - 463
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
May 15, 2026·23m - 464
Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling
May 15, 2026·25m - 465
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
May 15, 2026·25m - 466
Many-Shot CoT-ICL: Making In-Context Learning Truly Learn
May 15, 2026·24m - 467
Qwen-Image-VAE-2.0 Technical Report
May 15, 2026·24m - 468
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
May 15, 2026·23m - 469
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
May 15, 2026·24m - 470
World Action Models: The Next Frontier in Embodied AI
May 14, 2026·25m - 471
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
May 14, 2026·25m - 472
Efficient Pre-Training with Token Superposition
May 14, 2026·25m - 473
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
May 14, 2026·23m - 474
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics
May 14, 2026·23m - 475
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
May 14, 2026·24m - 476
MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
May 14, 2026·24m - 477
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
May 14, 2026·26m - 478
MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments
May 14, 2026·22m - 479
$δ$-mem: Efficient Online Memory for Large Language Models
May 14, 2026·25m - 480
Qwen-Image-2.0 Technical Report
May 13, 2026·23m - 481
TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
May 13, 2026·23m - 482
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents
May 13, 2026·23m - 483
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
May 13, 2026·25m - 484
Model Merging Scaling Laws in Large Language Models
May 13, 2026·22m - 485
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
May 13, 2026·24m - 486
SEIF: Self-Evolving Reinforcement Learning for Instruction Following
May 13, 2026·21m - 487
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
May 13, 2026·22m - 488
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
May 13, 2026·22m - 489
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers
May 12, 2026·22m - 490
Flow-OPD: On-Policy Distillation for Flow Matching Models
May 12, 2026·26m - 491
Beyond Retrieval: A Multitask Benchmark and Model for Code Search
May 12, 2026·21m - 492
HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents
May 12, 2026·25m - 493
Anisotropic Modality Align
May 12, 2026·23m - 494
MiA-Signature: Approximating Global Activation for Long-Context Understanding
May 9, 2026·12m - 495
When to Trust Imagination: Adaptive Action Execution for World Action Models
May 9, 2026·12m - 496
Continuous-Time Distribution Matching for Few-Step Diffusion Distillation
May 9, 2026·15m - 497
Stream-T1: Test-Time Scaling for Streaming Video Generation
May 8, 2026·22m - 498
RLDX-1 Technical Report
May 8, 2026·23m - 499
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
May 8, 2026·22m - 500
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
May 8, 2026·25m