Jahanzaib
Back to Blog
AI Agentsai newsai-agentsai-security

Google Put a Make Believe Button on the Map the World Checks Against

A breakdown of why Google killed its Google Earth image generator in a day, why watermarks were never going to fix it, and what changes if you let AI write into systems your team treats as records.

Jahanzaib Ahmed
August 1, 2026·13 min read
Google Earth blog post titled Transform any place with Nano Banana in Google Earth, dated July 30 2026, by product manager Bryan Horowitz

Google shipped a create image button into Google Earth on Thursday. By Friday it was gone.

In between, an investigative journalist named Henk van Ess typed one sentence and put refugees near the Mexican border. Then he planted a nuclear plant in Iran. Then he put a fatal crash on a street in Amsterdam. Google's own satellite imagery sat underneath all three.

His summary is better than anything I could write: Google spent twenty years building the reference the world checks against, then added a button that makes things up.

Pulling it was the right call. But the reason it had to be pulled is not the reason most of the coverage gave, and the fix Google announced will not hold. That gap is your problem too, the moment you point a model at anything your business treats as a record.

What did Google actually launch, and why did it come down in a day?

On July 30 Google added Nano Banana 2, its image model, to Google Earth on the web. Worldwide, no waitlist, no application. You zoomed to any coordinates, clicked create image, typed what you wanted to see, and got back a photorealistic picture built on top of Google's real satellite, aerial and 3D imagery of that exact spot. If it was not quite right, you refined it. Roughly a day later Google rolled the whole thing back.

The pitch was benign. "For the first time, you can generate custom images using Google Earth's satellite, aerial, and 3D imagery alongside Nano Banana, which creates concepts grounded in the real world," wrote Bryan Horowitz, product manager for Google Earth, in the launch post. The examples were architects reimagining an empty lot in Tokyo as a shopping district, and a classroom rendering the ruins of Pompeii as the town looked in 78 A.D.

The reaction took hours. Eliot Higgins, who founded Bellingcat, posted that Google Earth, a tool often used as a source of satellite imagery to verify photos and videos, had added a feature to alter satellite imagery with AI, "for reasons." He followed it with an image of a giant golden statue of Donald Trump looming over the White House and the words "no way this could be abused." The BBC's Shayan Sardarizadeh made the same joke the same day, calling Google Earth one of the most reliable sources of visual evidence for journalists and researchers. Van Ess did not bother with the joke. He just went and did it.

Henk van Ess article on Digital Digging titled How to plant a nuclear plant in Iran, showing an AI generated crowd scene labelled Custom image inside Google Earth
Note the label on the generated frame: "Custom image." That is the entire distinction between a fabricated crowd and Google's real imagery, and it lives inside the picture Google encouraged you to screenshot.

By Friday Google had reversed. "We've seen geospatial professionals using this feature for a range of useful purposes, however we've also seen people sharing screenshots of generated imagery that appear to violate our policies," the company said. "So we're rolling back this feature in Google Earth while we work on implementing stronger guardrails."

Read that statement twice. The trigger Google names is people sharing screenshots. Not the generation. The distribution.

Why was the SynthID watermark the wrong answer?

Because a watermark is a forensic tool for somebody who already suspects the image. It does nothing for the person scrolling past it. Google's first response to van Ess pointed at SynthID and at blocks on harmful topics. Both of those operate on the file. Neither operates on the belief the picture creates in the two seconds before anyone thinks to check.

"We take misinformation seriously, every image created with Nano Banana in Google Earth includes the SynthID digital watermark, so if someone is unsure about an image, they can ask the Gemini app or use Lens in Search to see if the image was AI-generated," Google posted on July 30. "In addition, we prevent image creation on harmful topics and are continually updating our protections."

The technology is not the weak part. Ars Technica's own testing has found SynthID survives substantial edits and data loss, which is genuinely hard to build. The weak part is the sentence beginning "if someone is unsure." That clause is carrying the entire safety argument, and it assumes a reader who stops, doubts, opens a second Google product, and pastes the image in. Digital Digging still managed to fool Hive's AI detector with an altered video pulled out of Earth.

Van Ess made the cost concrete with an older case. A forger once took a Google Earth screenshot of the US Navy's Fifth Fleet headquarters in Manama, Bahrain, fed it to Gemini, and produced a fake image of the aftermath of an Iranian drone strike on the base. That took about six steps. With generation built into Earth it took, in his word, seconds.

Then there is the damage that runs the other way. Ars raised it and I think it is the underrated half of the story: a public tool for editing satellite imagery hands every official on earth a fresh excuse to wave away a genuine satellite photo as AI. You do not need to fake anything to benefit from that. You just need the possibility to exist.

The Verge article about the Google Earth AI image generator showing a photorealistic generated lakeside cabin with Refine image and Saved to project controls
The Verge's example is a dream house on a lake. Harmless, and exactly the problem: it is indistinguishable from a photograph, and the controls beside it say "Saved to project."

What did each newsroom get right, and what did all of them miss?

Four pieces across three outlets landed inside a day, and each framed it differently. Reading them side by side is useful, because the framing a story gets in the first 24 hours tends to become the lesson people remember from it.

SourceWhat it centredWhat it left on the table
The Verge, news pieceSpeed. The feature "only lasted one day," framed as a product misjudgement.Why a company this good at safety review shipped it at all.
The Verge, analysis pieceThat watermarks and topic restrictions were not enough on their own.What would have been enough.
TechCrunchThe backlash and the misinformation risk, with the BBC reaction as the hook.The second order harm to real imagery.
Ars TechnicaTrust in Earth as a reference, plus the excuse it hands officials denying real photos.The product design question underneath it.

Here is what none of them wrote about, and it is sitting along the bottom edge of the product screenshots every one of them published. Underneath the generated images, Google shipped its standard line: "Generative AI can make mistakes, so double-check it."

That sentence was written for chatbot answers. It works there because you can go and check a claim. Nobody double checks a satellite photo. There is no second satellite in your pocket. The disclaimer is a category error, and it ran underneath a fabricated disaster scene at Google's own headquarters.

Ars Technica article showing an AI modified Google Earth image of a fake disaster scene at the Googleplex in Mountain View, captioned as fabricated
A fabricated disaster scene at Google's own Mountain View headquarters, generated in Google Earth. Along the bottom edge of the frame sits the boilerplate Google ships under chatbot answers: "Generative AI can make mistakes, so double-check it."

So what was the actual engineering mistake?

Google applied a content filter where the product needed a context filter. A content filter asks whether this output is harmful. A context filter asks whether this output breaks the promise the surface makes to the people reading it. Google Earth's promise is that what you see is what is there. Every generated image breaks that promise, including the entirely harmless ones.

The Pompeii reconstruction is the clearest case. As an image it is lovely and educational and offends nobody. As a feature it is fatal, because it teaches users that the Earth canvas can hold things that are not there. Once the canvas is mixed, a reader has to verify every image on it, including the real ones. Trust in a reference does not erode picture by picture. It goes all at once, the first time you learn the reference can lie.

So when Google says it is rolling back "while we work on implementing stronger guardrails," I think that is the wrong fix, and I would rather be wrong loudly than hedge. Stronger guardrails on this architecture means a longer list of blocked prompts and the same broken guarantee. Blocklists lose to paraphrase. Van Ess did not need an exploit. He needed a sentence. If the feature returns with better refusals and the same shared canvas, it fails again, and the next researcher needs one afternoon to prove it.

What would work is boring. Put generated output in a different namespace from recorded output. A separate mode, a separate export, a frame burned into the pixels rather than a signal buried in them, and no path that puts a generated tile back onto the reference map. It wins no launch applause. It is the only version that survives contact with someone who wants to lie.

There is a wider pattern here about how fast generative features now reach production. Google's own vibe coding whitepaper argues for shipping quickly and correcting from real usage, which is reasonable advice for a feature whose worst outcome is a bug report. It is terrible advice for a feature whose worst outcome is a fake bomb crater beside a real hospital. The speed was not the problem on its own. Applying a build fast philosophy to a surface whose whole value is being trustworthy was.

Where does this same mistake show up in ordinary business builds?

Anywhere an agent writes into a field a human also writes into. CRM notes, calendar entries, ticket summaries, the internal wiki, the handover doc. The generated text is nearly always harmless and often better written than what the human would have typed. The damage arrives six months later, when nobody can tell which lines were observed and which were inferred.

And there are going to be a lot more of them. The gap between the agent numbers vendors announce and the ones actually in production is still enormous, but the direction is not in doubt. Every one of those agents will write somewhere.

I've shipped enough of these now to have one rule I do not negotiate on. Anything a model writes carries a source field that cannot be empty, and the interface renders it differently from human input. Not a badge hidden in a tooltip. A difference you cannot miss while skimming at speed, because skimming at speed is the only way anyone actually reads a client record. And there is no code path that drops the marker on export, because export is precisely where provenance dies.

I did not start there. I was wrong about it for a long stretch, and it bit me. Early on I let summarisation agents write straight into the existing notes field, because the output was good, the field already existed, and adding a column felt like ceremony for its own sake. It read fine. It read fine right up until the day someone needed to know whether a commitment sitting in a client record came from a phone call or from a model's reading of that phone call. That is not a writing quality problem. Better prose makes it worse, not better. The only fix was the column I had decided not to add.

This is the same failure class as the containment gaps in Anthropic's red team breaches and the Hugging Face agent incident, just wearing a friendlier outfit. In those cases an agent reached somewhere it should not have. Here the agent writes exactly where it was told to write, and the harm comes from the destination, not the reach. That version is harder to spot in review, which is why it survives longer. It is also why industry security alliances keep producing standards about model behaviour when most of the real exposure is in product architecture.

What should you check before wiring an agent into your own tools?

Three questions, in this order. What guarantee does this surface make to the people who read it? Does generated content break that guarantee even when it is completely correct? If the answer is yes, can you separate the namespaces instead of filtering the output?

Some things worth doing before the build, not after:

  • Say the guarantee out loud and write it down. "This calendar shows commitments a person made." "This ticket log shows what the customer said." If you cannot state it in one sentence, your team does not share it, and the agent will not respect it.
  • Provenance ships before quality. I tell clients this constantly and it still lands badly: a mediocre summary that is clearly labelled beats an excellent one that is anonymous, every time.
  • Test the export path, not the screen. CSV, PDF, API, the copy paste into Slack. Markers that live in CSS die on the way out.
  • Then assume the marker gets stripped anyway, and ask what breaks. If the answer is "nothing important," you probably did not need the separation. If the answer makes you wince, build the separate namespace.
  • Red team the harmless case. Google's safety review almost certainly caught the obvious abuse and let Pompeii through, and Pompeii is what proved the canvas was mixed.

If you are earlier than that and just want to know which of your systems can safely take an agent at all, the AI readiness quiz walks the same questions across your stack in a few minutes. For the pieces I actually build against this pattern, see the agent builds, and the glossary covers the terms if any of this was new.

Google will probably bring this feature back. If it comes back as a mode that cannot write to the reference map, that is a company that learned the lesson. If it comes back with a longer blocklist and the same shared canvas, we will be reading a near identical article in a few months, written by whichever researcher gets there first. My money is on the blocklist, and I would enjoy losing that bet.

Frequently asked questions

What was the Google Earth AI feature?

It was image generation powered by Nano Banana 2, added to Google Earth on the web on July 30, 2026. You could zoom to any coordinates, click create image, and describe a change. The model used Google's real satellite, aerial and 3D imagery of that location as its starting point and returned a photorealistic result you could refine and save to a project.

Why did Google remove it after one day?

Researchers and journalists demonstrated within hours that it made convincing geographic misinformation trivial to produce. Google's stated reason was that people were sharing screenshots of generated imagery that appeared to violate its policies, and it rolled the feature back while working on stronger guardrails.

Does the SynthID watermark make AI images safe to share?

No. SynthID is a detection aid, not a prevention layer. It helps a reader who already doubts an image and knows to check it in Gemini or Google Lens. It does nothing for the far larger number of readers who see the image in a feed, believe it, and move on. Digital Digging also managed to fool a third party AI detector with material generated through the Earth feature.

Is Google bringing the feature back?

Google's statement said it is rolling the feature back "while we work on implementing stronger guardrails," which points at a return rather than a permanent removal. What Google has not said is whether generated images would still be able to sit on the same canvas as authentic imagery, which is the part that decides whether the second attempt works.

What is the difference between a content filter and a context filter?

A content filter evaluates the output on its own terms and asks whether it is harmful. A context filter evaluates the destination and asks whether generated content breaks the guarantee that surface makes to its readers. Google Earth needed the second one, because a perfectly harmless generated image still breaks the promise that what you see on the map is what is really there.

How does this apply to AI agents inside a business?

Any system your team treats as a record makes a guarantee, whether or not anyone has written it down. A CRM note implies somebody observed the thing. A ticket log implies the customer said it. When an agent writes into those fields without a permanent provenance marker, the record stops meaning what everyone assumes it means, and the failure surfaces months later during an argument nobody can settle.

Citation Capsule: Google added Nano Banana 2 image generation to Google Earth on July 30, 2026 and rolled it back the following day. Google's own launch post and rollback statement, plus the demonstrations that triggered it. Google Earth blog, Bryan Horowitz (Jul 30, 2026) · Digital Digging, Henk van Ess (Jul 30, 2026) · Ars Technica, Jeremy Hsu (Aug 1, 2026) · The Verge, Jay Peters (Jul 31, 2026) · The Verge analysis, Stevie Bonifield (Jul 31, 2026) · TechCrunch, Lucas Ropek (Jul 31, 2026).
Feed to Claude or ChatGPT