What the video showed

Google released Gemini on Wednesday with a video titled Hands-on with Gemini: Interacting with multimodal AI. In it a person draws a duck and the model narrates the sketch as it takes shape. The person plays rock paper scissors with hand gestures and the model names the game. Cups are shuffled and the model tracks the ball. Pictures of planets are laid out and the model says which order they should go in. The model speaks in a natural voice, responds within a beat of the action, and never needs to be told what to look at.

By Thursday, Parmy Olson at Bloomberg had noticed that the video's description did not match what the model could do, and TechCrunch wrote it up under the headline that the demo was faked. Google's own explanation of how the video was made describes the method. They captured footage to test Gemini's capabilities, then prompted the model with still image frames pulled from that footage and typed text. There was no voice input. There was no video input. The latency in the video is an editing decision.

The gap between the cut and the transcript

The details matter because the video's impressiveness lives in exactly the parts that were staged. In the rock paper scissors segment, the person makes gestures one at a time in silence and the model recognises the game. In the documented interaction, the model was shown all three gestures at once alongside the prompt: what do you think we are doing, hint, it is a game. Recognising a game from three labelled stills with a hint is a different task from recognising it from a live hand.

The planets segment has the same shape. In the video the model looks at the pictures and gives a quick ordering. In the documentation, it was told it was an expert on planets and asked to consider distance from the sun before ordering them. And for the duck drawing, which is the segment most people remember, TechCrunch could not find documentation of the interaction at all. That does not mean it was invented. It means the most striking part of the launch is the one with the least evidence behind it.

Google's answer

Google's position, delivered through a spokesperson who asked TechCrunch to change its headline, is that the video shows real outputs from Gemini and that the company was transparent about the edits. Oriol Vinyals of Google DeepMind said the video illustrates what multimodal user experiences built with Gemini could look like. Both statements are true in a narrow sense. The outputs were real outputs. Google did describe the method, for anyone who went looking for the description instead of watching the video. And the video does illustrate something that could exist.

The problem is that the video was the launch. A disclosure that lives somewhere else while the main artefact implies a capability the model does not have is a disclosure designed not to be read. Could look like is a phrase for a concept video, and the video was not labelled as one.

Why this matters beyond the video

We care about this less as a story about marketing and more as a story about evidence. The same launch shipped a set of benchmark claims, and those claims will be read by people who have just learned that the flagship demo was assembled to look better than the model performs. That is a reasonable update. If the demo was staged in ways the demo did not say, a reader is entitled to ask what else in the launch was arranged and not said. Benchmark reporting has plenty of quiet choices, prompt format, number of shots, whether the comparison model was run under the same conditions, and a lab that has just spent its credibility on a video has less to spend on those.

Demo culture has drifted toward this for years, and Google is not alone in it. But there is an asymmetry worth naming. A staged demo costs the lab a news cycle. It costs the field something more durable, which is the default assumption that a research organisation's public claims are made in good faith. That assumption is what lets outside researchers build on reported results without reproducing every one, and it erodes one launch at a time.

What disclosure should look like

The fix is not complicated and it does not require giving up the video. Put the method on the screen. If the input was still frames, say still frames in the first ten seconds, on the video itself, and show a frame being submitted. If the prompt included a hint, show the hint. If the latency was cut, show a timer. A viewer who sees a real interaction with real timing and real prompts, and is still impressed, has been persuaded honestly. A viewer who is impressed by an edit has been persuaded of something false, and will remember it when the product ships.

We would like the next multimodal launch, from any lab, to include one unedited session with the raw inputs visible, alongside whatever polished cut the marketing team wants. If the unedited session is boring, that is information too.

Sources

  1. TechCrunch, Google's best Gemini demo was faked