Daily Paper Cast · Jingwen Liang, Gengyu Wang

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents

April 9, 2026·22 min·2 clips
How does COMBI achieve 17x speed-up in prompt learning while preventing context overload in parallel agents?
Four sections structure the ride. Echo and Nova open with COMBI as a paper on scaling prompt learning for self-improving language model agents, without touching model parameters. The intro starts with agents learning from inference-time context, then hits the obvious snag: context piles up fast. The methods section does most of the work. COMBI comes through as a parallel prompt learning framework, and the dynamic batch size controller gets treated as the piece that keeps it efficient and sturdy. The benchmarks are the meat. AppWorld tests multi-step API work through Task Goal Completion and Scenario Goal Completion, with 90 training tasks and a held-out test-normal split. TerminalBench 2.0 makes the test sharper: 89 command-line tasks, 60 DeepSeek 3.2 trajectories from HuggingFace, and average accuracy over three runs on 29 held-out tasks. Finance gets a separate pass. Finer checks fine-grained entity typing in XBRL documents, while Formula looks at numerical reasoning over structured filings. The spread matters. Software engineering, agentic workflows, and finance NLP all test whether scaled prompt learning travels beyond one tidy demo. Echo asks the setup questions. Nova answers with methods, metrics, and dataset structure, and the conversation keeps circling back to what the experiments show. The ending is upbeat, but not vague: COMBI is framed as a practical move toward scalable prompt learning, with better context management, parallel learning, and batch-size control doing the real work.

As heard by us

A clear look at COMBI's parallel prompt learning approach and its benchmark results.

Daily Paper Cast treats COMBI as a clean look at scalable prompt learning for self improving language model agents. It keeps the focus on the paper's core claim: agents can pick up task relevant knowledge from context at inference time without changing model parameters.

Read the full review in PlayNext →

Why you'd press play

You want a fast read on scaling prompt learning for self-improving language model agents.

Read the full recommendation in PlayNext →
Listen to the show on