Eight years from side event to track

Last week NeurIPS announced that the Machine Learning Reproducibility Challenge will run as an official track at the December meeting in Sydney. The challenge started at ICLR 2018 under Joelle Pineau, and for eight years it has lived in workshops, journal special issues and volunteer effort. The announcement calls this the first time reproducibility science has had a dedicated home inside a major machine learning conference, and we think that framing is fair.

We have reviewed for MLRC and we have submitted to it. The work was always real. What it lacked was a line on a CV that a hiring committee would recognise. A student who spent four months failing to reproduce a headline result had, at best, a workshop paper to show for it. The incentive problem was never that people did not value replication. It was that the venues did not.

How the track actually works

The mechanics are unusual and worth reading closely. Papers do not go through the normal NeurIPS review cycle. They must be accepted at TMLR within an eligibility window that opened on June 20, 2025 and closes on September 30, 2026, and then pass what the announcement calls a light compatibility review by the MLRC committee. There are three ways in. Authors can file an expression of interest before their TMLR decision, with a soft deadline of June 4. They can self-nominate after acceptance. Or a TMLR area chair can nominate them. Notifications go out on October 7 and the conference runs December 6 to 13.

Routing through TMLR solves two problems at once. TMLR already has a rolling review process and a culture of accepting careful negative results, so the track inherits reviewers who know how to assess a replication rather than reviewers primed to ask what is novel. And it avoids adding a fifth or sixth thousand submissions to a conference review pool that is already strained. The cost is that the timeline is long. If you have not started a TMLR submission by now, you are unlikely to make the September cutoff.

The scope is broader than we expected. Eligible papers include direct reproductions and replications that test specific claims, generalisability studies that push a finding into new settings, meta-reproducibility studies, methods and tools for reproducibility research, and AI-assisted reproducibility studies. The announcement states plainly that negative results and partial failures to reproduce are as valuable as confirmations. Written into a call for papers at a top venue, that sentence does more than a decade of blog posts arguing the same thing.

What counts as a contribution now

The category we keep coming back to is AI-assisted reproducibility. A year ago we would have read that as a gimmick. Today a good fraction of the replication attempts we see in our own group start with an agent reading the paper, pulling the repository, and trying to run the pipeline before a human touches it. The agent is frequently wrong about why something failed, and it is still the fastest way to find out whether a repository runs at all. Making that a legitimate object of study, with its own failure modes documented, is the right call.

The generalisability category matters for a different reason. Most results in this field are not false so much as narrow. A method that works on the three datasets in the paper and falls over on a fourth is a common outcome and a useful one to publish. Under the old regime the fourth dataset was a footnote in someone else's paper or a tweet. Under this one it is a contribution.

The metadata mandate on the Datasets track

The same day, the Evaluations and Datasets track announced a second change. Croissant, the machine-readable metadata format that was already required for dataset submissions, must now include a Responsible AI section. The announcement describes a minimal set of fields covering dataset limitations, potential biases, intended use, and related considerations. Authors can put the information directly in the Croissant file or link clearly to the sections of the paper that cover it. Submissions missing the RAI fields will be flagged during review. There is an online RAI editor and a Croissant validator, both hosted on Hugging Face, to help authors get the format right.

The rationale given is standard: standardised information on how a dataset was built, what it should be used for, and where it breaks reduces misuse and misleading conclusions, and improves comparability and reuse. We agree with all of that and we still want to be careful about what a mandate like this can achieve.

Whether mandates change practice

Datasheets for datasets have existed as a proposal since 2018. Model cards are older than most of the models people use. The pattern with these documents is that they get filled in at submission time by whoever is nearest the deadline, and the limitations field says something like the dataset may not generalise to all domains. A required field in a JSON file is not harder to fill in badly than a required section in a PDF. The validator checks that the field exists, not that it is honest.

What might be different this time is the machine-readability. If the limitations and intended-use fields live in Croissant, tooling can read them. A training pipeline could refuse to ingest a dataset whose intended use excludes the task. A leaderboard could display the stated limitations next to the score. None of that exists yet, and the announcement does not promise it. But a text field in a PDF can never be acted on by software, and a structured field can. The mandate creates the possibility, and the community decides whether anything is built on top of it.

The test we would apply to both announcements is the same. In a year, count how many Datasets track papers have RAI fields that say something a reviewer could falsify, and count how many MLRC track papers report a failure to reproduce a result from a well-cited paper. If both numbers are healthy, the structures worked. If the RAI fields are boilerplate and the replications are all confirmations of easy targets, we will have learned that a venue can change what is publishable without changing what people are willing to write.

Sources

  1. NeurIPS blog: MLRC 2026, reproducibility as an official track at NeurIPS
  2. NeurIPS blog: Responsible AI metadata requirements for the Evaluations and Datasets track