Pew: 10% of the Web Is Written With AI — Authorship Detection Is Not GEO Detection
Pew Research analyzed 490,000 webpages and found AI authorship signs on 10% of the July 2026 crawl, and over a third among post-ChatGPT dated pages. Authorship detection and GEO detection answer different questions — neither decides whether you get cited.
This blog content may use AI tools for drafting and structuring, and is published after editorial review by the RanketAI Editorial Team.
Three-line summary
- AI writing is close to becoming the default on the web. Pew Research analyzed 490,000 webpages and found signs of AI authorship on 10% of pages in the July 2026 crawl, rising to over one-third among pages whose publication date falls after ChatGPT's release.
- "Who wrote this" and "what was this built to win" are separate questions. Pew's detector identifies the writer. GEO detection research identifies the intent behind the optimization. Being flagged by one says nothing about the other.
- Scrubbing surface signals is not the answer. Em dashes and Oxford commas appear in human writing too, and neither detector determines whether an AI answer will cite your page.
What Pew actually measured
On August 20, 2026, Pew Research Center Data Labs published "How Much of the Internet Is Written With AI?". This is not a survey — it is a large-scale analysis of a web archive.
| Component | Detail |
|---|---|
| Sample | 49 Common Crawl snapshots (2021-01 to 2026-07) x 10,000 pages each = 490,000 English pages |
| Detector | Open Pangram (editlens_Llama-3.2-3B), scoring 0 (human) to 1 (AI) |
| Threshold | 0.2 and above counted as meaningful signs of AI authorship |
| Cross-validation | 62,370 pages checked against Pangram commercial model 3.3 → 96% agreement, Cohen's kappa 0.61 |
| Coverage | Publicly accessible pages only — paywalled and login-gated content is underrepresented |
A threshold of 0.2 is deliberately low, and Pew spells out what that produces: a page a person wrote and then cleaned up with an AI tool lands in the same bucket as one the model produced from scratch.
AI detection models are probabilistic tools, and individual page classifications should not be treated as definitive. — Pew Research Center, methodology (2026-08-20)
10% of the web, and over a third of recent pages
All figures below come from the July 2026 crawl.
| Segment | Share with AI authorship signs |
|---|---|
| All pages | 10% |
| Pages with a post-ChatGPT publication date | Over one-third |
| .com domains | ~10% |
| .org domains | 4.6% |
| .edu domains | ~1% |
| .gov domains | ~1% |
The gap between the two headline numbers is a timing effect. The sample contains a large volume of pages from before 2021, which could not have been written with these tools at all. Read together, they say that the share is far higher among content being produced now.
One constraint matters here. Pew could parse a publication date for only 10-15% of pages, so "over one-third" describes dated content, not the web at large.
The domain split is easier to read. A .com page is about twice as likely to carry these signs as a .org page, and roughly ten times as likely as .edu or .gov. AI writing is concentrated where content is produced commercially and at volume — which also means most of the pages you compete against are being made the same way.
The style markers have already shifted
Separately from the classification scores, Pew counted how often specific language features appear per 10,000 words (2023 → 2026).
| Signal | 2023 | 2026 | Change |
|---|---|---|---|
| Em dashes (—) | 5.79 | 11.19 | ~2x |
| Oxford commas | 34.04 | 55.51 | +63% |
| AI-typical vocabulary ("delve", "interplay") | 11.94 | 26.02 | ~2x |
| Negative parallelism ("it's not X, it's Y") | 0.87 | 2.36 | ~2.7x |
The tempting conclusion is to strip those constructions out. Pew closes that door itself: taken one at a time, none of these features can identify a specific document, and human writers use all of them. The shift only means something across hundreds of thousands of pages.
Authorship detection and GEO detection answer different questions
This is the distinction that gets lost most often in practice. Two studies landed within days of each other, but they measure different things.
| Pew (Open Pangram) | GEO-Flag | |
|---|---|---|
| Question | Who wrote this page | What this page was built to win |
| Target | Traces of AI generation or AI editing | Strategic intervention aimed at AI citation |
| Prevalence | 10% of all pages / over a third of recent ones | 8.9% of pages / 16.4% of pages modified in 2026 |
| Text only polished with AI | Counted (threshold 0.2) | Trained to be excluded |
The same page can read as "AI-written" under Pew's method and "not optimized" under GEO-Flag, and the reverse holds too — a page a person wrote by hand, with mechanical optimization patterns layered on top, falls into the second case.
The figures and the practical line on the optimization side are covered in GEO Content Is Now Detectable.
So what does this mean for citations?
There is a question neither study answers: does a page flagged as AI-written become less likely to be cited in an AI answer? No direct measurement of that exists in either paper. What is established stops here.
- A substantial share of the web is AI-written (Pew)
- That is technically detectable (both studies)
- No search engine or AI platform has announced a penalty based on it
On citation probability itself, other research has produced answers. A survey reviewing 45 GEO studies reported cases where rewriting that manipulated form without adding substance actually reduced exposure (-9% / -16% / -6%). That result points at whether the content genuinely answers the question, not at what wrote it.
In other words, the writing tool is not the deciding variable. Whether the draft came from a model or from a person, what remains is whether the page carries the information an answer needs, in a form that can be extracted.
What to check now
- Rewording to dodge detection is not a priority. Removing em dashes neither evades a classifier nor earns a citation. With a 0.2 threshold and kappa of 0.61, any single page verdict is probabilistic to begin with.
- Review depth matters more than whether AI wrote the draft. The fact that AI-edited text is classified alongside AI-generated text also means the reverse: whatever produced the draft, quality diverges once a person verifies the facts, attributes the sources, and organizes the structure.
- Understand the competitive baseline numerically. At roughly 10% of
.compages and over a third of recent content, an average AI draft does not differentiate. What does: first-party data, your own measurements, and specific cases. - Check how your page actually reads first. Independent of tooling, the precondition for citation is whether an AI system can use your page while composing an answer.
Frequently asked questions
Will using AI to write get me penalized in search?
No penalty policy has been announced. The Pew study measures prevalence, and no search engine or AI platform has stated that it demotes AI-written content across the board. That said, now that detection has been demonstrated at scale, it is reasonable to treat "platforms may use this as one quality signal" as a long-term risk.
If a person writes it and AI polishes it, does that count as AI-written?
Under Pew's method, yes. The 0.2 threshold is set to capture mixed human-and-AI writing broadly, so a polished draft and a fully generated one land in the same category. GEO-Flag's detector was trained the other way — AI polishing alone does not raise its score. The answer depends on which detector you mean.
Can I avoid detection by removing em dashes and words like "delve"?
Not recommended. Pew states that these features cannot identify a specific document and that human writers use them. Removing them offers no guarantee the classification changes, and even if it did, citation probability is a separate matter. Time spent adjusting style is better spent strengthening the content — that helps on both fronts.
On a small budget, if AI drafting has to continue, what should come first?
Shifting resources from drafting to review is the more efficient move: verifying quantitative claims against primary sources, placing the answer sentence near the top, and adding your own data or cases. All three act directly on citation likelihood regardless of what produced the draft.
Is the 10% figure reliable?
The sample and validation are solid — 490,000 pages across 49 crawls, with 96% agreement against a separate commercial model. But Cohen's kappa of 0.61 is substantial agreement, not perfect agreement, and about 1% of pre-ChatGPT documents were false positives. Treat it as a reliable trend at the aggregate level and as unsuitable evidence for judging any individual page.
Do the same rates apply to non-English content?
Unknown. Pew's sample is limited to English pages. No comparable large-scale public study of AI authorship in other languages is available, and applying English-web figures directly to another language is not supported by evidence.
Wrapping up
Pew's analysis of 490,000 pages turned an impression — "AI is writing the web" — into a number: 10% overall, over a third of recent content. That number does not translate into a citation penalty. Authorship detection and optimization detection answer different questions, and neither one determines whether an AI answer will cite you.
One standard survives. Whatever produced it, does the page hold information that genuinely answers the question, in a form that can be extracted? RanketAI's free page structure diagnostic shows how your page reads right now.
Related reading
- GEO Content Is Now Detectable — 8.9% of Web Pages Already Optimized — the detection side that targets optimization intent, and where the practical line sits.
- Review of the 45-Study Critical GEO Survey — the backfire figures for form-only rewriting.
- What Is GEO (Generative Engine Optimization)? — verified optimization methods and the conditions under which they work.
- 9 AI Search Myths vs 15M Data Points — the record of shortcut strategies failing under measurement.
Execution Summary
| Item | Practical guideline |
|---|---|
| Core topic | Pew: 10% of the Web Is Written With AI — Authorship Detection Is Not GEO Detection |
| Best fit | Prioritize for geo workflows |
| Primary action | Standardize an input contract (objective, audience, sources, output format) |
| Risk check | Validate unsupported claims, policy violations, and format compliance |
| Next step | Store failures as reusable patterns to reduce repeat issues |
Frequently Asked Questions
How does the approach described in "Pew: 10% of the Web Is Written With AI —…" apply to real-world workflows?▾
Start with an input contract that requires objective, audience, source material, and output format for every request.
Is AI search suitable for individual practitioners, or does it require a full team effort?▾
Teams with repetitive workflows and high quality variance, such as geo, usually see faster gains.
What are the most common mistakes when first adopting AI search?▾
Before rewriting prompts again, verify that context layering and post-generation validation loops are actually enforced.
Data Basis
- Primary source: Pew Research Center Data Labs, "How Much of the Internet Is Written With AI?" (2026-08-20). The sample (49 Common Crawl snapshots from 2021-01 to 2026-07, 10,000 English pages each, 490,000 total), the detector (Open Pangram editlens_Llama-3.2-3B, scoring 0 to 1, with 0.2 and above classified as meaningful signs), the cross-validation (62,370 pages against Pangram commercial model 3.3, 96% agreement, Cohen's kappa 0.61), the prevalence figures (10% of the July 2026 crawl, over one-third among pages with post-ChatGPT publication dates), the domain breakdown and the linguistic marker trends were all read directly from the Pew report and its methodology page on 2026-08-25.
- Stated limits: publication dates could be parsed for only 10-15% of pages, so the post-ChatGPT figure describes dated content rather than the web as a whole. Open Pangram showed roughly 1% false positives on pre-ChatGPT documents. Text written by a person and then polished with AI falls in the same category as text generated by AI from scratch. Paywalled and login-gated pages are underrepresented. Individual page classifications are probabilistic estimates and are not treated as conclusive here.
- The GEO-Flag prevalence figures (8.9% / 16.4%) cited for contrast were verified against the arXiv abstract and body in our earlier article, and are used here only to establish that the two studies measure different things.
Key Claims and Sources
This section maps key claims to their supporting sources one by one for fast verification. Review each claim together with its original reference link below.
Claim:In the July 2026 crawl, 10% of all sampled webpages showed meaningful signs of AI authorship
Source:Pew Research Center Data Labs (2026-08-20)Claim:Among pages with a detectable post-ChatGPT publication date, the share rises to over one-third
Source:Pew Research Center Data Labs (2026-08-20)Claim:By domain, roughly 10% of .com pages, 4.6% of .org pages and about 1% of .edu and .gov pages carry those signs
Source:Pew Research Center Data Labs (2026-08-20)Claim:Per 10,000 words, em dashes rose from 5.79 to 11.19 and negative parallelism from 0.87 to 2.36 between 2023 and 2026
Source:Pew Research Center Data Labs (2026-08-20)Claim:The sample was 490,000 pages drawn from 49 Common Crawl snapshots, and cross-validation on 62,370 pages reached 96% agreement with Cohen kappa 0.61
Source:Pew Research Center methodology (2026-08-20)Claim:The GEO-Flag study flagged 8.9% of real-world pages, and 16.4% of pages modified in 2026, as GEO-optimized
Source:arXiv 2608.16824 (Chu et al., 2026-08-17, preprint)
External References
The links below are original sources directly used for the claims and numbers in this post. Checking source context reduces interpretation gaps and speeds up re-validation.
Is your site visible in AI search?
See for free how ChatGPT, Perplexity, and Gemini describe your brand.
Start Free Diagnosis →Related Posts
These related posts are selected to help validate the same decision criteria in different contexts. Read them in order below to broaden comparison perspectives.
Why Brand-Name Prompts Can't Measure AI Visibility — The "AI Knows Us" Illusion
With web search on, AI mentions a brand almost every time its name is in the prompt — near-zero signal as a KPI. Baselines by query type, the three real failure conditions, and query-design criteria for choosing an AI visibility tool.
89% of AI Search Demand Has No Clear Brand Owner — What 1,094 Categories Reveal (2026)
Kevin Indig tracked 1,094 US categories in ChatGPT for six months: only 15.2% had a clear brand owner and 89.3% of estimated demand sat in unowned categories, yet ownership held at a 90.4% monthly rate once claimed. How to claim a category before it locks.
AI Search Cannot Verify Your Business: What a 71-Business Audit Found (2026)
An audit of 71 verified local businesses found the average one meets only 15.6% of what AI needs to trust and retrieve it, and 17% had no AI-accessible presence at all. The five verification gaps, plus a six-point checklist to run this week.
AI Search Answers Without Referring: 9.4% Search Displacement, 5.2% Outbound Clicks
A new arXiv study finds wider ChatGPT Search access cut traditional search use by 9.4%, while only 5.2% of ChatGPT sessions produced an outbound click versus 31.1% for Google. Why brands should measure presence inside AI answers, not just site visits.
GEO Content Is Now Detectable — 8.9% of Web Pages Already Optimized (2026)
A detector now flags GEO-optimized content at F1 0.944. The GEO-Flag study flagged 8.9% of 10,095 real-world pages — and 16.4% of pages modified in 2026 — as GEO-optimized. We unpack the numbers and the line between mechanical and sustainable optimization.