The claim

Microsoft AI released MAI-Thinking-1 on Tuesday, a mixture-of-experts model with 35 billion active and one trillion total parameters, alongside MAI-Code-1-Flash. The announcement said the models were trained from the ground up on enterprise grade, clean and commercially licensed data, without distillation from third-party models, and built end to end by Microsoft using clean and appropriately licensed data. Simon Willison's first reaction, which he later walked back, was to ask whether these might be the first generally useful code models that did not train on an unlicensed dump of the web.

That is the natural reading of the words. Clean and appropriately licensed, to most people, means the training data was either licensed from its owners or otherwise cleared for use. So we opened the technical report to see what it says about where the data came from.

What the report says

Section 2.4 of the report says the base model was pretrained on 30 trillion tokens from a mixture of publicly available and licensed human-generated data covering web pages, public GitHub code, books, academic papers, news, and multilingual text. Appendix A gives the pipeline. The majority of the web corpus comes from a proprietary crawl of approximately 1.2 trillion pages, which policy, adult content, and blocklist filtering reduces to 794 billion pages, and which exact deduplication reduces to 423 billion documents. After fuzzy deduplication the crawl holds 73.4 billion English and 116.5 billion non-English documents. Common Crawl goes through the same pipeline, starting from roughly 300 billion pages and ending at 24.2 billion after deduplication against the proprietary corpus.

The licensing section, 2.4.1, is precise about what publicly available means. The crawler respects robots.txt and related meta tag controls. Sources that violate Microsoft's Responsible AI policies or appear on the USTR Notorious Markets list are excluded. Third-party datasets acquired through commercial agreements are subject to diligence on ownership and usage rights. The report also states that for privacy, legal, safety, and competitive reasons the full list of datasets and providers is not disclosed.

So the picture is a large web crawl that honours opt-out signals, plus Common Crawl, plus some licensed material of undisclosed size. That is a conventional and defensible data recipe. It is also a very different thing from what the launch copy implied.

What clean turns out to mean

Reading the report, clean is a statement about content quality and provenance hygiene rather than about rights. The authors chose not to use any language-model-generated synthetic data in pretraining, and they score pages with an AI-content detector and drop domains with extensive machine-generated text. They do not use open source training datasets and they exclude machine learning repositories such as huggingface.co from the web data. They filter for PII risk and safety. Those are real choices and some of them are unusual, and a model card that said this plainly would be informative.

Appropriately licensed is the phrase doing the heavier lifting. In the report it resolves to a crawl collected in accordance with applicable terms of use and industry standards for web controls. Whether robots.txt compliance makes a crawl appropriately licensed is a live legal argument in several jurisdictions, and reasonable people land on both sides. What is not reasonable is presenting that argument as settled in an announcement while the report itself uses more careful language.

Why the wording matters

There are two reasons a reader should care. The first is that provenance claims are now a competitive feature. Buyers in regulated industries ask whether a model was trained on licensed data, and a vendor that can say yes has an advantage. If the phrase can be satisfied by a filtered crawl of the open web, then it stops discriminating between vendors and starts being decoration. Simon's correction, and his admission that he did not cover this one well despite being at the Build conference when he wrote it, is a useful reminder that even careful readers take launch copy at face value on a busy day.

The second reason is that the report is good, and the gap between the report and the copy undersells it. The pipeline is described at a level of detail most labs no longer publish, down to the deduplication parameters and the fraction of English documents dropped by the quality model. The knowledge cutoff for each source family is tabulated. The authors say they are sharing these details to support a transparent and science-driven approach. That is exactly the disclosure the field should reward, and it deserves an announcement that matches it.

Open the PDF

The practical rule we take from this is simple and slightly tedious. When a model announcement makes a claim about data, find the section of the technical report that describes data collection and read it before repeating the claim. The report will usually be more honest than the headline, because it is written by the people who built the pipeline for readers who will build one. In this case it took us about ten minutes to find the two paragraphs that mattered.

What we would like to see from Microsoft and everyone else is a fixed vocabulary. Licensed should mean a contract exists. Public should mean crawled under opt-out controls. Synthetic-free should mean what MAI-Thinking-1 actually did. Three plain words in the model card, each with a number attached, would end most of these arguments before they start.

Sources

  1. Microsoft AI, MAI-Thinking-1: Building a Hill-Climbing Machine (technical report, June 2026)
  2. Simon Willison, weblog entries for June 2, 2026