Two weeks from we/O to a post-mortem

AI Overviews launched to everyone in the United States at Google we/O in mid May. By the end of the month Elizabeth Reid, who runs Search, had published a post explaining what went wrong. That is a fast turnaround for a company that usually lets search changes settle for months, and it tells you how loud the reaction was. The two examples she chose to address directly were an answer about how many rocks a person should eat and an answer suggesting glue as a way to keep cheese on pizza.

Her explanation for each is specific and worth reading on its own terms. The rocks answer came from a query that almost nobody had ever typed, so the only page that matched was a piece of satire that a geological software site had republished. The glue answer came from a forum. Both are cases where the retrieval step did its job, found the most relevant text on the web, and handed it to a model that had no way to know the text was a joke or bad advice.

Data voids, satire and forums

Reid's post names three failure classes and we think the taxonomy is the most useful thing in it. The first is the data void, a query for which little or no good content exists, so whatever exists wins. The second is satire and humour, which is indistinguishable from fact at the level of a retrieved passage. The third is user-generated content, where a forum thread can be the most relevant document and the worst possible advice at the same time.

None of these are model failures in the usual sense. The model faithfully summarised a source that should never have been a source, and there was no hallucination involved. That distinction matters for anyone building a retrieval system, because the fixes live in retrieval and triggering rather than in the generator. Google says it shipped over a dozen changes, including better detection of queries where an overview is unhelpful, limits on satirical content, limits on user-generated content in advice-giving answers, and more triggering restrictions for queries with low utility. Health queries got extra guardrails.

The post also pushes back on part of the story. Reid says a large number of the screenshots that went around were faked, including some of the most dangerous-looking ones. We have no way to check that claim from outside, and we would treat it the way we treat any vendor's account of its own incident. But it is a fair reminder that the viral sample of an outage is not a random sample of it.

The base rate argument

Google's defence rests on rates. The post says the accuracy of AI Overviews is on par with featured snippets, that content policy violations appeared in fewer than one in every seven million unique queries, and that users are asking longer questions and clicking through to better pages. If you believe the numbers, the feature is working about as well as the thing it replaced.

The trouble is that a rate is the wrong unit for this surface. Search handles billions of queries, so a one in seven million rate is still hundreds of bad answers a day, each one screenshotted by someone who was actively hunting for it. A featured snippet that quotes a forum reads as a quote from a forum. An overview that says the same thing in Google's own voice reads as Google telling you to eat rocks. The error rate did not change much. The attribution of the error did, and that is what people reacted to.

There is a second problem with the featured snippet comparison. Snippets fail on the same data voids, but a snippet failure has been ambient for years and nobody was looking. Putting a model on the surface made everyone look at once. Some of what got found was probably already there.

What this means for anyone deploying on a wide surface

The lesson we take is about the shape of the deployment, and it generalises well beyond Google. When you ship a model behind a narrow product, your users are people who chose the product and roughly know what it does. When you ship it on a surface everybody already uses, your users include people who type nonsense to see what happens, and their queries land precisely on the data voids where your retrieval is weakest. At that scale the adversarial tail is a full-time audience.

The other lesson is that triggering is a first-class model decision. Most of Google's fixes are about when not to answer rather than how to answer better. That is a harder problem than it sounds, because a system that declines too often loses the benefit it was shipped for, and a system that declines too rarely ends up in the news. We would like to see someone publish a proper study of the triggering classifier's precision and recall on a public query sample, since that number, more than any accuracy figure, is what decides whether the surface is safe.

Next we want to know whether the fixes hold. Reid says the company will keep watching and adjusting. The data void problem does not go away, because new nonsense queries are free to generate and the web will always have exactly one page for some of them. A month from now, the test will be whether the next rocks answer is gone before anyone screenshots it.

Sources

  1. Elizabeth Reid, AI Overviews: About last week (Google, 30 May 2024)
  2. Wikipedia, AI Overviews