+Field test

September 22, 2026

I Ran an AI Detector on Ten Pages That Rank. It Flagged Six. It Flagged Me Too.

Detectors are coin flips, the watermarks are moving inside the models, and Google now rejects review replies for sounding scripted. I pointed a structural tell detector at ten top-ranking pages and four of my own. It flagged the humans and it flagged the human. The only durable defense is effort you can show, and the rejection data says exactly what that looks like.

01The claim

The detectors lost. The platforms did not need them.

OpenAI shut down its own AI text classifier in July 2023 for a low rate of accuracy. It caught 26 percent of AI text and called 9 percent of human writing a machine. The company that makes the machine could not detect the machine. Every detector sold since then is asking you to believe it did better with less.

Meanwhile detection moved somewhere it can work: inside the models. Anthropic said in August that text from new Claude models carries an invisible watermark that travels when copied and may survive some editing. Google has had one on Gemini output for two years. The people who generate the text can grade the text. Third parties never could.

And the platforms stopped waiting for either. Since April, Google Business Profile holds every review reply for moderation and returns it as pending, approved, or rejected. One vendor analyzed 12,752 rejected replies from one month of that filter: 67 percent contained detectable AI boilerplate, 92.6 percent were replies to five-star reviews, identical replies sent to a hundred-plus reviews got caught as spam, and all 38 replies containing a hashtag were rejected. Every one. (Somebody put a hashtag in a review reply thirty-eight times. That person is not reading this. That person is fine.)

02The receipts

I pointed my own detector at the pages that rank. Then at me.

My agency runs a structural tell detector in its content pipeline. It does not guess who wrote a thing. It looks for shape: the list-of-three cadence, the not-this-but-that antithesis, the rhetorical question opener, bullet lists made of noun phrases, the generic wrap-up, and a short vocabulary list of phrases that read as machine. It is a linter. We use it to catch tics before a draft ships, not to accuse anyone of anything.

This morning I ran it over the ten top-ranking pages from a sameness test I did on two customer questions: HGTV, Sweeten, a Chicago tint shop, 3M, a regional design-build firm, and five more. Six of ten got flagged. Five for marching in threes. One for noun-phrase bullets. Two for the phrase when it comes to, which is apparently load-bearing at HGTV and a vinyl company. None crossed the rewrite threshold. All of them rank.

Then I ran it over my own four articles from this week, written by a human who bans em dashes for taste. Two of four flagged, same tic, the list of three. The detector could not tell HGTV from me, and to be fair to the detector, neither could you. That is the finding. A structural detector finds structure, and structure is what competent writing has. It is not evidence of origin. It is a horoscope that happens to mention rhythm.

03What it means

Nobody is penalized for using AI. They are penalized for being interchangeable.

Read the rejection data again. Boilerplate phrases. Identical replies at scale. Hashtags in a thank-you note. Those are not signs of a machine. They are signs of nobody being home. A human who pastes the same thanks so much for the kind words into 140 five-star reviews gets the same rejection as a script, because to the reader it is a script. Google's own guidance has said for years that appropriate use of AI or automation is not against its guidelines, and the filter data agrees. It is not hunting AI. It is hunting sameness.

The search side is the slow version of the same rule. John Mueller said this month that recovering from scaled content abuse tends to take time and significant effort to show the value. Show the value. That is the sentence. The penalty lands months after the pages went up, so the short-term traffic tells you nothing, and the recovery is not a humanizer pass, it is proving there was a reason for the page to exist.

Which puts the whole detector industry in a strange spot. The evasion loop, run the draft through a detector, hit the humanizer, delete the em dashes, run it again, produces text that is optimized for a coin flip and reads worse. The platforms are not looking at the coin. They are looking at whether the second reply is the same as the first.

04The playbook

Effort you can show, in four moves

One. Kill the evasion loop. Delete the humanizer. Keep the em-dash ban if you like the look, I do, but stop pretending it is a defense. Nothing in the rejection data cares about punctuation.

Two. Validate demand before you produce. The pages the spam update hits are the ones nobody searched for, written at volume because a quota said so. If a topic has no demand, the most human draft in the world is still a page with no reason to exist. Check the number first.

Three. Put the effort where a reader can see it. First-party numbers from your own jobs. A photo of the actual dented door. A named author with a license number. A specific answer in the first sentence. It is the same list as the originality test because it is the same problem: the filters, the models, and the readers all reward the page that has the thing in it.

Four. Review replies get the same rule. Draft with whatever you want, then make each reply reference something in the review it answers, never send the same sentence twice, and check the moderation state instead of assuming it published. A reply sitting in rejected for fifty days is a customer who thinks you ignored them.

05The tooling

Demand first, quality in the editor, oversight on the replies

The two checks that actually predict survival are the two I cannot do by hand at scale. In the Semrush Content Toolkit, Topic Finder validates that a topic has real demand before anyone writes a word, which is the whole difference between a content program and a quota. SEO Writing Assistant then scores the draft for originality, tone, and readability inside the editor, so the interchangeable draft gets caught before it ships instead of after a spam update finds it. That replaces the detector loop with a question that matters: is this worth publishing.

For the review side, Review Management in the Semrush Local Toolkit keeps the replies in one queue with a human reading each one before it goes out, which is exactly the duplicate-at-scale pattern the rejection filter is documented to catch.

06The verdict

The detector flagged HGTV. Write like someone was there.

What the engines did with pages that had a real number in them is in the citation test.

Fourteen pages is not a study. It is one linter and a morning. But it settled the question I had. A structural detector cannot see origin, it can see shape, and shape is what good writing has. The platforms know this, which is why they stopped detecting and started rejecting sameness. Detectors cannot see sameness. Readers can. Write for the reader and the filter takes care of itself.

You own a number that isn't moving. Let's move it.