The Foundation Model Transparency Index: scoring the labs on what they will not say
Stanford's index scores ten developers on 100 binary indicators and the best of them, Meta, gets 54. Reading notes on what the index measures, where the zeros cluster, and whether a scoreboard changes behaviour or just documentation.
A hundred yes or no questions
Rishi Bommasani and seven colleagues at Stanford's Center for Research on Foundation Models posted the Foundation Model Transparency Index on October 19. It is a list of 100 indicators, each answered yes or no from public material available as of September 15, applied to the flagship model of ten developers. The developers are OpenAI (GPT-4), Anthropic (Claude 2), Google (PaLM 2), Meta (Llama 2), Inflection (Inflection-1), Amazon (Titan Text), Cohere (Command), AI21 Labs (Jurassic-2), Hugging Face (BLOOMZ, as host of BigScience) and Stability AI (Stable Diffusion 2).
The indicators split into three domains. Upstream covers what went into the model, meaning data, data labour and compute, with 32 indicators. Model covers the artefact itself, meaning size, architecture, capabilities, limitations, risks and mitigations, with 33. Downstream covers distribution and use, meaning release, usage policies, feedback, impact and documentation, with 35. Within those there are 13 subdomains with at least three indicators each. Two researchers scored every indicator independently and resolved disagreements by discussion, and then every company got two weeks to contest scores in private.
The numbers
Meta scores 54, Hugging Face 53, OpenAI 48, Stability AI 47, Google 40, Anthropic 36, Cohere 34, AI21 Labs 25, Inflection 21 and Amazon 12. The mean is 37 with a standard deviation of 14.2. The authors see three clusters, four developers well above the mean, three around it, and three well below. Nobody clears 55 on a test where the questions are things like whether the model size is stated or whether there is a way for users to seek redress.
The distribution across domains is the more useful result. The mean upstream score is 7.2 out of 32, against 14.1 of 33 for the model domain and 15.7 of 35 for downstream. Three developers, AI21 Labs, Inflection and Amazon, receive zero across all 32 upstream indicators and Cohere gets three. No company at all scores a point for indicators about who created the data, its copyright and licence status, or copyright mitigations. The Impact subdomain averages 11 percent, the worst in the index, because no developer reports affected sectors, affected geographies or any usage figures.
Openness helps, and the paper is precise about where. Developers that release weights average 53 percent on upstream indicators against 9 percent for closed developers. On downstream indicators the gap almost vanishes, 49 percent versus 43 percent, even though API providers have far more control over how their models are used and could report on it. Two of the three open developers beat every closed one, and the third, Stability AI, sits one point below OpenAI.
What the rebuttal round tells you
All ten companies replied to the draft scores and eight contested specific ones, on average 8.75 indicators each. The contests raised scores by an average of 1.25 points per contesting developer. That is a small number, and the authors say most rebuttals surfaced misunderstandings about indicator definitions rather than overlooked disclosures. We read that as evidence the scoring is mostly stable under adversarial pressure, which is the main thing you want to know about an index before you trust the ranking.
It also tells you something about the ceiling. The authors note that 82 of the 100 indicators are satisfied by at least one developer and 71 by more than one. So the gap between 54 and 82 is made of practices someone already does. Nobody has to invent a new disclosure to get into the seventies, they have to copy a competitor.
Documentation theatre or moved behaviour
The obvious worry about a binary index is Goodhart's law, which the paper cites by name. An indicator like centralised model documentation can be satisfied with a page that exists and says very little. A developer can climb ten points by writing things down without changing what they do. The authors know this and devote a limitations section to gaming and to the general critique that composite indexes flatten concepts that are complicated.
We are less bothered by this than we expected to be, for two reasons. First, the indicators that nobody satisfies are the ones that resist theatre. You cannot write a page about data licences without stating what the licences are, and you cannot report affected sectors without counting them. Second, Bommasani has said in the accompanying HAI piece that the index deliberately does not rate corporate responsibility. A company gets a point for disclosing a bad practice, not for avoiding it. That keeps the instrument narrow and makes it hard to argue that a high score is a moral claim.
The failure mode we would watch for is subtler. If the next edition shows large gains on the cheap indicators and no movement on data, labour and impact, the index will have measured the cost of documentation rather than the willingness to disclose. The first edition already gives the baseline needed to check that, subdomain by subdomain.
What we would want next
The scoring justifications are public and the code is on GitHub, which is the part most rankings skip. We would like someone outside Stanford to rescore a subset of indicators blind and report agreement, since two internal scorers resolving disagreements by discussion is a thin reliability check for a result that will get quoted in policy hearings.
And we would want the second edition to add a column for whether each disclosure was verifiable by a third party, since a stated model size and a size anyone can check are different kinds of transparency. The index as it stands measures whether the labs say things. Whether what they say can be tested is the harder question, and it is the one this field will have to answer.
Sources
From the foundation