By Gautam Kamath, on behalf of the 2025 TMLR Outstanding Paper Committee: Pablo Samuel Castro, Pin-Yu Chen, Vincent Dumoulin, Amir-massoud Farahmand, Andreas Kirsch, Jasper Lee, Jeffrey Pennington, Colin Raffel, Chang Xu.

The 2025 TMLR Outstanding Paper Committee is pleased to award the Outstanding Certification to the following paper:

The Mantis paper has multiple contributions. First, the authors provide a large-scale multi-image instruction tuning dataset. The authors further train a model which achieves state-of-the-art results, despite being trained on academic-level resources. Finally, they provide a human-curated benchmark that requires multi-image understanding. Conceptually, the paper demonstrates that interleaved multi-image instruction tuning can outperform massive pre-training. Furthermore, the committee recognized its substantial and immediate impact on the community, with the Mantis-Eval benchmark adopted by subsequent state-of-the-art models such as Qwen3-VL, Intern3.5-VL, and LLaVA-OneVision.

The Committee also recognizes three papers as Outstanding Certification Finalists:

Causal Reasoning and Large Language Models was lauded for its comprehensive empirical look at causal reasoning abilities of LLMs, opening this frontier for broader investigation. Cognitive Architectures for Language Agents (CoALA) provided a conceptual framework for connecting classic ideas in computational cognitive science and LLM agent research, synthesizing a fragmented body of work into an actionable set of insights and roadmap for the future. Gradient Scarcity in Graph Learning with Bilevel Optimization gave a mathematical characterization of why gradient scarcity occurs when learning graphs, providing deep insight into the nature of the problem. All papers listed above will receive a Featured Certification, if they did not already have one.

Selection Process

The selection process follows the same general approach used in previous years. All papers published in TMLR up until April 30, 2025, were considered, excluding those considered in previous years. This pool of papers was filtered down based on two criteria: if they were awarded a Featured Certification by their Action Editor, or if they received a large number of citations (determined using Google Scholar). Since TMLR has been rapidly expanding, this still resulted in too many papers to be considered by the committee, despite having a larger number of committee members than in previous years. For each paper that fit these criteria, we asked its Action Editor for their opinion of whether the paper should be considered for the Outstanding Certification, instructing them to err on the side of being permissive, as we didn’t want their sole opinion to carry too much weight unless they were very confident. For the papers which proceeded further, the Action Editors’ opinions were used as an expert review in deliberations.

At this point, we partitioned the resulting set of papers amongst the committee, with care for conflicts of interests. Each committee member was responsible for independently handling roughly six papers, and their instructions were to gather opinions from two additional experts in the area of the paper. These external opinions are essential, as experts in the area have the most informed evaluation of how significant a result is. Each committee member was instructed to advance one paper (or, exceptionally, zero or two papers) to the shortlist.

Eleven papers advanced to the shortlist.** Over the course of roughly two weeks, the committee was instructed to form opinions on the papers in the shortlist, using their own reading, expert comments, and discussion among the committee members. After voting and some additional discussion, committee members agreed upon the outcome above.

After the process, committee members noted that many of the shortlisted papers seemed to be focused on Large Language Models. On one hand, this reflects a popular direction that the community has coalesced around, and thus it is natural that the best work may be concentrated in this area. On the other hand, it may be a bias within the process itself. It was noted that LLM papers are more likely to have a high number of citations, and based on a high number of citations, even experts may be more inclined to comment on them more favourably than they would otherwise. We appreciate the committee’s feedback and will reflect upon whether there are ways to represent more styles of papers in the selection process.

Acknowledgments

I would like to extend my sincere thanks to the entire committee. TMLR takes the Outstanding Certification selection process very seriously, with a long and elaborate process as described above. The committee was able to meet these rigorous demands and participate actively and thoughtfully across the process, including on short notice or with rapid turnarounds at times. They have my genuine gratitude.

I would also like to thank Hugo Larochelle, who served to witness the selection process.

Finally, in previous years, we have listed all expert reviewers who were kind enough to share their opinions and expertise. Given that the recent OpenReview security incident is fresh in peoples’ minds, and the fact that we did not previously notify experts that their names might be publicly acknowledged, we felt that it might be a bit tone-deaf to reveal their names right now, even without explicit association to which papers they opined on. Nonetheless, they are free to list this information publicly if they so desire. In future years, we will explicitly warn experts that their name will appear in an acknowledgment.

*Note that I (Gautam Kamath) have an institutional conflict with several of the authors, being colleagues at the University of Waterloo. However, I was only involved in defining and overseeing the process, and did not participate in the discussions or voting process. To the best of my knowledge, my CoI did not substantially influence their decisions. Indeed, support from the committee for the winning paper was overwhelming.

**At this point, Colin Raffel was excused from the committee due to a strong CoI with one of the papers advanced (by another committee member) to the shortlist.