Summarize this documentation using AI
Key Takeaways
- A 90-day content review exists to answer one question: which formats earned citations, and which only earned effort.
- Score every published asset on four axes: retrievability, citation evidence, cluster contribution and commercial proximity.
- Most programmes discover that a small minority of formats carry nearly all the citation value. The review is how you find which minority.
- Judge the programme on citation share and branded demand, not on sessions. Pew found only 1 percent of visits to a page with an AI summary produced a click on a source link.
- End the review with three decisions: what to double, what to consolidate, and what to stop. A review without a stop list has not concluded anything.
Ninety days into a content programme you have enough published work to see a pattern and not yet enough sunk cost to be defensive about it. That is the right moment to stop and audit, and it is the moment most teams skip, because the calendar says publish and the calendar is easier to obey than a spreadsheet.
This is the framework we use to run that review. It is written for programmes optimising for answer engine visibility, where the traditional scoreboard of sessions and rankings tells you very little about whether the work is landing.
Why the standard content review fails for AEO programmes
The conventional quarterly content review ranks posts by pageviews, keeps the top decile, and quietly abandons the rest. That method breaks when the goal is citation rather than clicks.
Pew Research Center tracked 68,879 Google searches across 900 U.S. adults and found that when an AI summary was present, users clicked a traditional result in 8 percent of visits, against 15 percent when no summary appeared. Clicks on a source link inside the summary itself occurred in 1 percent of visits.
A page cited in a hundred AI answers and clicked once will sit near the bottom of a pageview-ranked report. Kill it on that evidence and you have deleted your most-quoted asset. The review has to measure the thing you were actually trying to do.
The four-axis scoring model
Score every asset published in the period on four axes, one to five each. The composite is less important than the pattern that emerges when you sort by each axis on its own.

Axis 1: Retrievability
Can a machine find, parse and lift a clean answer from this page?
- Does a self-contained answer appear in the first 40 to 60 words under a matching heading?
- Are headings phrased as the questions people actually ask?
- Is the page reachable by AI crawlers, confirmed in server logs rather than assumed?
- Does the page carry appropriate structured data, per Google's guidance on AI features?
Score 5 where a model could quote one paragraph and be entirely correct. Score 1 where the answer is distributed across the whole page and cannot be extracted without reading all of it.
Axis 2: Citation evidence
Run your fixed prompt set. Twenty to thirty buying-intent prompts, the same ones every month, across the engines that matter to your market. For each asset, record whether it was named, linked, or paraphrased without attribution.
Paraphrased-without-attribution is the most instructive category. It usually means the claim was good and the wording was ambiguous enough that the model felt no need to credit anyone. Those pages are candidates for a specificity rewrite rather than deletion.
Treat the evidence with appropriate scepticism. The Tow Center at Columbia tested eight AI search engines over 1,600 queries and found more than 60 percent of answers were incorrect, with over half of Gemini and Grok 3 responses citing fabricated or broken URLs. A single spot check proves nothing. A monthly trend across a fixed prompt set proves something.
Axis 3: Cluster contribution
Does this asset make its neighbours stronger, or does it sit alone?
- How many internal links point at it from other assets in the same cluster?
- How many does it send onward?
- Does it occupy a distinct query, or does it compete with something you already published?
Cannibalisation is the most common finding at ninety days, because the calendar rewards volume and volume produces near-duplicates. Two mediocre pages on the same question should become one strong page and a redirect. Our guide to auditing a lifecycle marketing programme applies the same consolidation logic to campaign inventory.
Axis 4: Commercial proximity
How many steps sit between reading this asset and buying something?
A definition page such as what is subscription retention is far from the transaction but wide at the top. A comparison page is close to it and narrow. Both are legitimate. What is not legitimate is a portfolio made entirely of one or the other, which is what an unaudited calendar reliably produces.
Reading the results
Sort the scored inventory four times, once per axis, and look for the three patterns that show up in nearly every review.
Pattern 1: the format concentration
Citation evidence almost never distributes evenly across formats. One or two formats carry most of it. Frequently it is definition anchors and data-backed comparisons, because both produce short, checkable, self-contained claims. Long narrative essays tend to underperform on retrievability for structural reasons rather than quality reasons.
The action is not to abandon the underperforming format. It is to restructure it so the extractable claim exists.
Pattern 2: the orphan cluster
Somewhere in the inventory there is a group of three to six assets on a real topic that nothing links to and that link to nothing. They were published in a burst, and the internal linking pass never happened. This is the cheapest fix available: an hour of linking usually moves cluster contribution more than a month of new publishing.
Pattern 3: the high-effort, low-return long tail
Every programme has assets that took three times the average effort and returned nothing on any axis. Look for what they share. Usually it is that they answered a question nobody was asking, and the review is the first time anyone checked.

Closing the review: three decisions
A review that ends in observations has not ended. Force it to three lists.
Double
The formats and clusters with demonstrated citation evidence. Commit specific capacity to them for the next ninety days, expressed as a number of assets, not an intention.
Consolidate
Overlapping assets that split authority. Merge into the stronger URL, redirect the weaker, and rebuild internal links to point at the survivor. Consolidation is the highest-return work in most reviews and the least satisfying to report, because the output is a smaller site.
Stop
The formats that cost the most and returned the least across all four axes. Write them down explicitly. A stop list is the only part of a content review that creates capacity, and it is the part teams reliably omit.
What to carry into the next quarter
Three commitments make the next review easier than this one.
- Freeze the prompt set. Changing prompts between reviews destroys comparability. Add new prompts as a separate tracked set.
- Record the crawler check. Log-level evidence that AI crawlers reached each new asset, captured at publish time rather than reconstructed later.
- Write the extractable claim first. If a brief cannot state, in one sentence, the claim a model should be able to lift, the asset is not ready to write. That single discipline moves retrievability more than any post-hoc optimisation.
For the underlying concepts this review depends on, see cohort analysis for the measurement mindset, lifecycle revenue for the commercial frame, and lifecycle marketing strategies for the programme context these content assets are meant to serve.
The bottom line
Ninety days is long enough to have evidence and short enough to act on it cheaply. The review works when it changes the calendar. Score honestly on four axes, accept that citation share and not sessions is the scoreboard, and finish with a stop list you actually honour.
The programmes that compound are not the ones that published the most in the first quarter. They are the ones that found their two working formats in the second and stopped paying for the other five.
Sources
- Google users are less likely to click on links when an AI summary appears in the results, Pew Research Center
- AI Search Has a Citation Problem, Tow Center for Digital Journalism, Columbia Journalism Review
- AI Features and Your Website, Google Search Central
- Top ways to ensure your content performs well in Google's AI experiences on Search, Google Search Central Blog
Frequently Asked Questions
How do you review an AEO content programme?
Score every asset published in the period on four axes from one to five: retrievability, citation evidence, cluster contribution and commercial proximity. Sort the inventory by each axis separately rather than by a composite, because the patterns appear within single axes. Close the review with three explicit lists: what to double, what to consolidate, and what to stop.
Why do pageviews fail as a content metric for AI search?
Because citation and clicks have decoupled. Pew Research Center found users clicked a source link inside an AI summary in 1 percent of visits, and clicked traditional results in 8 percent of visits when a summary appeared versus 15 percent when none did. A page cited in a hundred AI answers can rank near the bottom of a pageview report, so pruning on pageviews can delete your most-quoted asset.
How often should you audit content for answer engine visibility?
Every ninety days is the practical cadence. That is long enough to accumulate evidence across a fixed prompt set and short enough that changing course is still cheap. Run the same prompt set every month for measurement, but do the full four-axis scoring and the doubling, consolidation and stop decisions quarterly.
What does cannibalisation look like in an AEO content programme?
Two or more assets answering substantially the same question, each with moderate citation evidence and neither dominant. It is the most common finding at ninety days because publishing calendars reward volume and volume produces near-duplicates. The fix is to merge into the stronger URL, redirect the weaker one, and rebuild internal links to point at the survivor.
Can you trust AI citation data when auditing content?
Only as a trend, not as a single reading. The Tow Center tested eight AI search engines across 1,600 queries and found more than 60 percent of answers incorrect, with over half of Gemini and Grok 3 responses citing fabricated or broken URLs. Use a frozen prompt set run monthly so you are comparing like with like, and treat any single spot check as noise.

![Featured image: 25 Best Lifecycle Marketing Strategies - Based on Customer Lifecycle Stages [#2025 Update]](https://cdn.prod.website-files.com/69bbf9f7f2a3f75341357ebc/69f35b7250ddb533f7fcbd8b_69c7ce9f68e818fd4aa3dad6_698ae432a7625eb340e995a0_681eed717eea07585b2e5d5c_Lifecycle%25252520Marketing%25252520Strategies-1-min.png)
![Featured image: Are You Doing Lifecycle Marketing Right? [2025 Guide to Lifecycle Marketing]](https://cdn.prod.website-files.com/69bbf9f7f2a3f75341357ebc/69f35ba0096407058cd41531_69c7cea4126bca4d08f8985b_698ae439bbd09eb85ea274be_681f03dca921853a9f20dc7c_Lifecycle%25252520Marketing%25252520Challenges-1-min.png)