ICLR 2026 in Rio: lost in multi-turn conversation
Notes from Rio built around the two outstanding papers, one proving transformers are exponentially more succinct than automata and one measuring a 39 percent drop when a task arrives across several turns. The second names a failure every agent builder already knew and finally gives it a number.
The week
ICLR ran from April 23 to 27 at Riocentro in Rio de Janeiro, and the outstanding paper awards were announced on the first day. The committee of twelve started from a longlist of 36 papers picked by the program chairs and by review score, consulted outside experts on each, cut to a shortlist of five and voted. Two papers won and one got an honourable mention, The Polar Express by Amsel, Persson, Musco and Gower, which derives optimal polynomial approximations for the matrix sign function used inside the Muon optimiser.
The two winners could hardly be more different in kind, and that is what made them a good pair to organise a week around. One is a theory result about what transformers can express compactly. The other is an empirical result about what they fail to do in an ordinary conversation.
Succinctness as the right measure
Pascal Bergsträßer, Ryan Cotterell and Anthony Lin ask a question that expressivity work usually skips. Two models may recognise the same languages, but one may need exponentially more description to do it. They prove that fixed-precision transformers are exponentially more succinct than linear temporal logic and than recurrent networks, and doubly exponentially more succinct than finite automata. They also give matching upper bounds. A transformer converts to an LTL formula with at most exponential blowup, which improves on the previous doubly exponential translation.
The consequence we found most useful is the negative one. Because transformers pack so much into so little, verification problems for them, including emptiness and equivalence checking, are EXPSPACE-complete. Anyone hoping that formal verification of small transformers would scale into a practical tool now has a clean reason why it will not, at least not by the automata route. The committee praised the paper for its conceptual strength, and we would add that it is rare to see a theory result that tells applied people what to stop trying.
The number everyone already felt
Philippe Laban, Hiroaki Hayashi, Yingbo Zhou and Jennifer Neville take instructions from six generation tasks, Python functions, text-to-SQL, API calls, elementary maths, table captioning and cited multi-document summaries, and shard each instruction into pieces that arrive one turn at a time. A simulated user reveals one shard per turn, and the model is scored at the end. The same instruction is also given in one turn, fully specified, and as a single concatenated bullet list, so the comparison isolates the multi-turn delivery from the information content.
Fifteen models from eight families, including GPT-4o, GPT-4.1, Gemini 2.5 Pro, Claude 3.7 Sonnet and Llama 3.3 70B, all degrade. The average drop is 39 percent across the six tasks over more than 200,000 simulated conversations, which cost about 5,000 dollars to run. The decomposition is the finding. Aptitude, the best the model can do, falls by about 16 percent. Unreliability, the spread between its best and worst attempts on the same task, rises by 112 percent. The model can still do the task. It just does it inconsistently, because it commits to an early guess, generates a solution before it has the whole specification, and then builds on that solution instead of revising it.
Why the measurement matters
Everyone who has built an agent has watched this happen. A user adds a constraint on turn four and the model patches its turn-two answer rather than starting over. What the paper adds is that the effect is large, that it survives every mitigation tried, and that it does not go away with scale. Lowering temperature to zero helps single-turn reliability and gives only minor gains in the multi-turn setting. Recap and snowball strategies that repeat prior context recover only 15 to 20 percent. Bigger models degrade about as much as small ones.
That last point is the one we would push on people who think the fix is the next model. If size did not help across a range from 8B to the frontier, the cause is somewhere in how these models are trained, and the authors argue for native multi-turn support in training rather than agent-side patches. Their practical advice for users is blunt. If a conversation has gone wrong, start a new one and put everything you know into the first message.
What we are taking home
The two papers meet in an odd place. One says transformers can represent a great deal in very little. The other says they cannot hold a partially specified task open for a few turns without prematurely collapsing it. Both are about commitment, in a sense, and we do not have a theory that connects them. We would like one.
The concrete thing we plan to do is add a sharded condition to our own agent evaluations. We report single-turn numbers because they are cheap and the benchmarks come that way. Laban and colleagues have shown that the single-turn number overstates what a user will see by something like 39 percent, and that the gap is mostly variance. An eval that only reports the mean of a distribution whose spread has doubled is hiding the result that matters.
Sources
From the foundation