AI Never Reads Your Whole Page. It Reads One Chunk. Here Is How To Own It
By David Sessford · 31 Aug 2026 · 12 min read
- Only 15% of the pages ChatGPT retrieves get cited. The other 85% were fetched, read and binned.
- 44.2% of ChatGPT citations come from the first 30% of a page. Your conclusion is the worst place to put your best line.
- Heavily cited text averages 20.6% proper nouns versus 5% to 8% for normal writing. Name names.
- 78.4% of question-linked citations came from headings. Make the H2 the question, answer it immediately underneath.
- 89.6% of prompts trigger follow-up searches, and 95% of those queries have zero search volume. You cover them with blocks, not keywords.
- The test for every block: cut it out, paste it in a blank doc, and see if it still makes sense to a stranger.
What content chunking actually is
Here is the thing nobody selling you an "AI SEO package" wants to explain. ChatGPT does not read your article. Neither does Perplexity. Neither does Google's AI Mode.
They chop it up.
Your page gets split into segments. Each segment gets turned into a vector, which is just a long list of numbers describing what that segment means. When somebody asks a question, the question becomes numbers too. The machine finds the segments whose numbers sit closest to the question's numbers, pulls those, and writes an answer out of them.
That segment is a chunk. And the chunk is the unit that wins or loses. Not the page. Not the domain. Not your 3,000 word masterpiece with the beautiful narrative arc.
Content chunking is the practice of writing your page as a series of self-contained blocks, each one able to survive being ripped out of context and dropped into an answer.
Most people are still optimising a page. The machine is grading paragraphs. That gap is where your competitor is beating you right now.
Why this matters more than your word count
Let me give you real numbers instead of vibes.
AirOps analysed 548,534 pages that ChatGPT retrieved across 15,000 prompts. Of everything the model went out and fetched, only 15% ended up cited in a final answer. 85% got pulled, read, and binned.
Read that again. Your page can rank. Your page can get crawled. Your page can get retrieved by the model in real time. And it still loses, 85 times out of 100, because when the machine looked at what it grabbed, nothing in there was quotable.
That is not a visibility problem. That is a writing problem.
The citation rate splits by query type too: 18.3% for product discovery, 16.9% for how-to questions, and only 11.3% for validation searches, the "is X any good" queries where someone is checking up on you. The queries where you most want to be named are the hardest ones to win.
Everyone in your niche is arguing about llms.txt files and schema plugins. Meanwhile the actual gate is whether one paragraph on your page can stand on its own two feet. We drill this exact rewrite inside the Dojo, and it is usually the fastest win a member gets in their first month.
The ski ramp: where citations actually come from
Kevin Indig ran the study that should have changed how everybody writes. His team analysed 1.2 million AI answers, isolating 18,012 verified citations out of roughly 3 million responses and 30 million citations, using sentence embeddings to trace each cited claim back to the exact sentence it came from.
The shape of the result is brutal and consistent. He calls it a ski ramp.
| Where it sits on your page | Share of citations | What that means for you |
|---|---|---|
| First 30% of content | 44.2% | Nearly half the prize is decided above the fold |
| Middle (30% to 70%) | 31.1% | Still live, if the blocks are clean |
| Final third | 24.7% | Sharp drop near the footer. Your conclusion is dead weight |
So the big conclusion you saved for the end, the payoff you built up to like a proper writer? The machine gave up on it. All that craft, spent on the section with the worst odds.
But there is a second layer, and this is the bit that catches people out. Inside a paragraph, the pattern flips. 53% of citations came from the middle of paragraphs, 24.5% from first sentences, and 22.5% from last sentences.
Translation: front-load at the article level, but do not force every paragraph to open with a punchline. Inside a paragraph, density beats theatre. Put the substance where the substance goes, and make sure there is substance at all.
How the machine cuts your page up, step by step
You cannot control the chunker. You can control what it finds when it cuts. Here is the sequence.
1. Fetch
The model or its retrieval layer grabs your HTML. Anything hidden behind a click, lazy-loaded on scroll, or rendered by JavaScript after the fact is a coin flip. Plain, server-rendered text is not glamorous. It is just what gets read.
2. Strip
Nav, footer, cookie banner and sidebar get thrown out. Your main content survives. If your key answer is sitting in a sidebar widget or an accordion, congratulations, you have written it for nobody.
3. Split
The text gets cut into chunks. Most production systems use a fixed window with a bit of overlap, and they respect structural boundaries where they can find them. Your headings, paragraph breaks and list items are the seams the knife follows. Give it clean seams or it cuts through the middle of your best sentence.
4. Embed
Each chunk becomes a vector. A chunk about three different things produces a muddy vector that matches nothing precisely. A chunk about one thing produces a sharp vector that matches that thing hard. This is the entire reason "one idea per block" works. It is not a style preference. It is maths.
5. Retrieve and rank
Query comes in, closest chunks come back. Then the model reads the shortlist and picks what to actually quote. This is the 85% graveyard. You made the shortlist and still lost.
6. Fan out
Here is the move most people have never heard of. ChatGPT does not just search your query. In the AirOps data, 89.6% of prompts triggered two or more follow-up searches, expanding 15,000 prompts into 43,233 queries. 32.9% of cited pages showed up only in those follow-ups, never in the original search. And 95% of those fan-out queries had zero traditional search volume.
Zero volume. Which means your keyword tool has never heard of them and never will. You cannot target fan-out queries with a spreadsheet. You cover them by writing chunks that answer the adjacent questions a curious person would ask next. That is the whole game, and it is why a page built as ten sharp answers outperforms a page built as one long essay.
The seven rules of a chunk that gets quoted
These come straight out of the citation data, not out of my head.
Rule 1. Define before you decorate
Cited passages were nearly twice as likely to contain a clear definition. "X is Y." "X refers to Z." Direct subject, verb, object. The hedge, the throat-clear, the "in today's fast-moving landscape" opener, all of it gets skipped. Say the thing.
Rule 2. Make your H2 the question
Cited content was twice as likely to contain a question mark, and 78.4% of question-linked citations came from headings. The model treats your H2 as the prompt and the paragraph underneath as the answer. So write the heading the way a human would type it, then answer it immediately underneath. No preamble between the two.
Rule 3. Name names
Typical English text runs 5% to 8% proper nouns. Heavily cited text averaged 20.6%. That is roughly triple. Brands, tools, people, places, versions, dates. Every named entity is a hook the machine can grab. "Several popular tools" is invisible. "Ahrefs, Semrush and Google Search Console" is a citation waiting to happen.
Rule 4. Neither robot nor ranter
Cited text clustered around a subjectivity score of 0.47, right between dry fact and pure opinion. The tone that wins reads like an analyst: here is the fact, here is what it means. Not a press release. Not a rant. Fact plus interpretation.
Rule 5. Write down, not up
Winning content averaged a Flesch-Kincaid grade level of 16. The stuff that lost averaged 19.1. Both are higher than a blog usually needs, because this is business writing, but the direction is unambiguous. Shorter sentences beat dense academic prose. Indig calls it a clarity tax. Pay it.
Rule 6. Every block survives amputation
This is the test. Take any block on your page. Cut it out. Paste it into a blank document. Does it still make sense to a stranger? If it opens with "this means that" or "as we saw above" or "the second reason is", it is dead on arrival. Chunks do not get to reference their neighbours.
Rule 7. Tables beat prose for anything comparative
Structured data extracts cleanly and unambiguously. If you are comparing options, prices, or timelines, a table gives the machine rows it can lift whole. A paragraph gives it a puzzle. This is the bit agencies quietly charge $2k a month for, and it is a formatting decision you can make this afternoon.
The rewrite: what a losing chunk becomes
Theory is cheap. Here is the actual move.
Before (invisible):
In today's competitive digital landscape, many businesses are wondering about the costs associated with their online efforts. There are a number of factors that can influence this, and it really depends on your specific situation. Below, we will explore some of the considerations that might affect what you end up paying.
Forty seconds of nothing. No definition, no entity, no number, no answer. The chunker will happily cut this out, embed it, and it will match precisely zero queries, because it is about nothing.
After (quotable):
How much does local SEO cost in 2026?
Local SEO costs $500 to $2,500 a month for most small businesses working with an agency, or $67 a month if you run it yourself with a system. The gap comes down to three things: how many locations you have, how competitive your city is, and whether you are fixing an existing Google Business Profile or building one from scratch.
Question as the heading. Answer in the first sentence. Two real numbers. Named entity. Then the qualification, placed after the answer instead of in place of it. That block can be lifted straight into an AI answer without a single edit, and it works just as well for a human skimming on a phone.
You do not need to rewrite your whole site. You need to do this to the twenty blocks that matter. We hand members the exact rewrite template and a checklist for finding those twenty blocks inside the Dojo.
How long should a chunk be?
Roughly 150 to 300 words, or one idea, whichever comes first.
People overthink this. There is a whole industry of RAG consultants selling clever "semantic chunking" that uses AI to find natural breakpoints. Researchers at Vectara tested it properly and published the result at NAACL 2025, in a paper with the wonderfully blunt title "Is Semantic Chunking Worth the Computational Cost?"
The answer was no. In realistic conditions, plain fixed-size chunking matched or beat the clever version across document retrieval, evidence retrieval and answer generation.
Which tells you something useful. If the sophisticated splitting method cannot reliably beat cutting every couple of hundred words, then the win is not in the splitting. The win is in whether the text you wrote holds together when it gets split. Stop looking for the trick. Write blocks that make sense on their own.
Chunk shapes that win and lose
| Element | Gets cited | Gets binned |
|---|---|---|
| Heading | The question a human would type | "Our Approach", "Key Considerations" |
| First sentence under it | Direct answer with a number or definition | Context-setting, throat-clearing |
| Block length | 150 to 300 words, one idea | 900 word wall, five ideas |
| Proper nouns | Around 20% of the text | "Various providers", "some experts" |
| Openers | Stands alone completely | "As mentioned above", "This means" |
| Comparisons | Table with labelled rows | Prose listing six things in a row |
| Best answer sits | In the first 30% of the page | In the conclusion |
| Tone | Fact plus interpretation | Hype, or lifeless spec sheet |
Do this in order
- Pick five pages that already get impressions. Not your favourites. The ones Search Console says are already being seen. You are upgrading proven pages, not gambling on new ones.
- Find the real question. For each page, write down the exact sentence a customer would type or say. Not the keyword. The sentence.
- Move the answer to the top. Put that question in as an H2 near the start, and answer it in 40 to 60 words directly underneath. If the answer already exists 1,400 words down, cut and paste it up. That is a ten minute job with a measurable payoff.
- Break the walls. Any paragraph over about 120 words gets split. Any section covering more than one idea gets its own H3.
- Run the amputation test on every block. Cut, paste into a blank doc, read it cold. Fix anything that only makes sense in place.
- Add the entities. Replace every vague noun with a name. Every "recently" with a date. Every "affordable" with a dollar figure you can stand behind.
- Turn comparisons into tables. Prices, options, pros and cons, timelines. All of it.
- Answer the next three questions. What would a curious person ask immediately after reading this? Add a short block for each. That is your fan-out coverage, and 95% of those queries will never show up in a keyword tool.
- Re-test after four weeks. Ask ChatGPT, Perplexity and Google AI Mode the original question. Note whether you appear. Repeat monthly.
Timeline and cost, honestly
I am not going to promise you a citation in thirty days, because nobody can and anybody who does is selling you something.
What I can tell you is the shape of the work. Restructuring one existing page properly takes 45 to 90 minutes once you know the moves. Five pages is an afternoon. AI systems that crawl live tend to pick up changes within days to a few weeks; systems relying on cached or trained data can take considerably longer, and some may never revisit.
| Route | Typical cost | What you get |
|---|---|---|
| Do it yourself, no system | $0 plus your weekends | Slow, guesswork, easy to do the wrong thing thoroughly |
| The Dojo | $67 a month | The templates, the checklists, the order to do it in |
| Typical agency retainer | $1,500 to $5,000 a month | Someone else does it, on their schedule |
| Done For You | Quoted per project | We do the restructure for you |
The honest version: the actual work here is not hard. It is fiddly, repetitive, and it needs doing in the right order. That is precisely why so few of your competitors have done it, and precisely why the window is open right now.
Who this is for, and who should skip it
Do this if you sell something people research before buying, you already have pages getting impressions, and your customers ask questions with real answers. Service businesses, local trades, B2B, software, consultants. This is your move.
Skip it if you have no content at all, in which case go write five genuinely useful pages first and come back. Skip it if your entire business runs on paid ads and you have no organic ambition. And skip it if you are hoping this replaces being any good, because a well-chunked page that says nothing worth quoting is still a page that says nothing worth quoting.
Five mistakes that will cost you
- Chunking a page nobody trusts. Structure amplifies authority, it does not create it. 55.8% of cited pages ranked in Google's top 20, and pages at position 1 were cited 3.5 times more often than pages outside the top 20. Retrieval and ranking are still entangled. Your SEO fundamentals and your links still count.
- Stuffing the top with keywords. Front-loading the answer is not front-loading the keyword. The data rewards clarity and entity density, not repetition.
- Turning your page into a FAQ list. Twenty one-line Q and As is not chunking, it is thin content wearing a costume. Each block needs enough substance to be worth quoting.
- Forgetting the human. If your page reads like a machine-readable spec sheet, the visitor who does click through will bounce. Good chunking makes pages easier to skim, not colder to read.
- Chunking once and walking away. Prices change, tools change, the models change. A cited chunk with a 2024 figure in it stops getting cited.
Tools that actually help
You need less than you think. Google Search Console tells you which pages already get impressions, which is where you start. Your own eyes and a blank document run the amputation test better than any software. For checking whether you actually appear, just open ChatGPT, Perplexity, Claude and Google AI Mode and ask the question yourself, monthly, and log it in a spreadsheet. That is a real tracking system and it costs nothing, which we covered in full in our guide to free AI visibility tracking.
Paid AI visibility platforms exist and some are decent, but buying one before you have fixed your chunks is buying a thermometer instead of a coat. Our current picks live in the Armory.
Risks and the honest counterargument
Two things worth saying out loud.
First, chunking is optimisation for a moving target. Retrieval architectures change. The specific numbers in this article describe systems as they were measured in 2026, and they will drift. What will not drift is the underlying logic: a self-contained, clearly-written block that answers a real question is easier to retrieve, easier to quote, and easier for a human to read. That has been true since long before anyone said the word "embedding" and it will outlive the current models.
Second, there is a fair objection that all this produces flat, samey, machine-pleasing content. It can. That is a real risk. The fix is not to write worse, it is to keep the opinions, the specifics and the voice, and simply put the answer first instead of last. Structure is the skeleton. Personality is still yours.
The 2026 view
Everyone is still optimising pages. The retrieval layer stopped caring about pages a while ago.
The operators who win the next two years will be the ones who think in blocks: dozens of tight, self-contained, quotable answers spread across a site, each one covering a question a real person actually asks, including the ones no keyword tool will ever show them. Your competitor is currently writing their 47th 2,000 word "ultimate guide" with the answer buried on line 900.
Let them. Go and rewrite twenty paragraphs instead.
Related reading: how to get AI to recommend your business, entity SEO if the machine has never heard of your brand, and ranking in Google AI Overviews.
Your move
Open your best page right now. Find the paragraph that answers the question your customers actually ask. Count how far down the page it is.
If it is not in the top third, you already know your first job.
Everything above is the mechanism, free, no gate. If you want the templates, the rewrite checklists, the tracker and the order to do it in, plus the rest of the system: Join The Dojo. $67/mo.
Frequently asked questions
What is content chunking in SEO?
Content chunking is writing your page as a series of self-contained blocks, each covering one idea in roughly 150 to 300 words. AI search systems split pages into segments, convert each into a vector, and retrieve the segments that best match a question. Chunking makes sure the segment they pull actually stands on its own.
How long should a chunk be?
Around 150 to 300 words, or one complete idea, whichever comes first. Vectara researchers testing chunking strategies at NAACL 2025 found that plain fixed-size chunking matched or beat more sophisticated semantic chunking in realistic conditions. The advantage comes from how you write, not how the text gets split.
Does chunking help with normal Google rankings too?
Yes, and the two reinforce each other. AirOps found 55.8% of cited pages ranked in Google's top 20, and pages at position 1 were cited 3.5 times more often than pages outside the top 20. Clear headings, direct answers and scannable blocks also help featured snippets and reduce bounce.
Where should I put my main answer on the page?
In the first 30% of the content. Kevin Indig's analysis of 1.2 million AI answers found 44.2% of citations came from the first 30% of a page, 31.1% from the middle, and only 24.7% from the final third. Put the answer under a question-shaped H2 near the top.
Will chunking make my writing sound robotic?
It can if you let it. The data actually rewards balanced tone, with cited text clustering around a subjectivity score of 0.47, which is fact plus interpretation rather than dry spec sheet or hype. Keep your opinions and specifics. Just move the answer to the front instead of the end.
How do I check if AI is quoting my content?
Ask it. Open ChatGPT, Perplexity, Claude and Google AI Mode, type the exact question a customer would ask, and note whether your business appears and which page it cites. Log it monthly in a spreadsheet. That is a genuine tracking system and it costs nothing.
How many pages should I chunk first?
Start with five pages that already get impressions in Google Search Console. You are upgrading proven pages rather than gambling on new ones. Restructuring one page properly takes 45 to 90 minutes once you know the moves, so five pages is roughly an afternoon.
David Sessford. Founder of The KillerEdge, 14+ years in SEO & online marketing More about the founders →
Sources
- Search Engine Land: 44% of ChatGPT citations come from the first third of content
- Growth Memo (Kevin Indig): The science of how AI pays attention
- Search Engine Land: Only 15% of pages retrieved by ChatGPT appear in final answers
- AirOps: The Influence of Retrieval, Fan-out, and Google SERPs on ChatGPT Citations
- NAACL 2025 Findings: Is Semantic Chunking Worth the Computational Cost?
Grab the free AI-Search Cheat Sheet
The exact moves to make ChatGPT, Gemini and Google AI name-drop YOU. Plus the weekly intel. Free.
Instant download. No spam. Quit anytime.
Want it done for you?
Learn it in The Dojo, or have us run it. Either way, you win.
Join The Dojo. $67/mo → ← Back to the blog