Resources
The 90-Day AEO Content Review: A Framework for Auditing What AI Actually Cites

The 90-Day AEO Content Review: A Framework for Auditing What AI Actually Cites

Ninety days in, you have enough evidence to see which formats earn citations and not enough sunk cost to defend the ones that do not. A four-axis scoring framework for auditing an answer-engine content programme, ending in a stop list.

Written by:
Jaskaran leads lifecycle strategy at Propel. He's built retention programs in Customer.io, Braze, Klaviyo, and MoEngage for brands across DTC, fintech, and healthtech.
September 2, 2026
·
6
min read
The 90-Day AEO Content Review: A Framework for Auditing What AI Actually Cites

Table of Contents

Summarize this documentation using AI

Key Takeaways

  • A 90-day content review exists to answer one question: which formats earned citations, and which only earned effort.
  • Score every published asset on four axes: retrievability, citation evidence, cluster contribution and commercial proximity.
  • Most programmes discover that a small minority of formats carry nearly all the citation value. The review is how you find which minority.
  • Judge the programme on citation share and branded demand, not on sessions. Pew found only 1 percent of visits to a page with an AI summary produced a click on a source link.
  • End the review with three decisions: what to double, what to consolidate, and what to stop. A review without a stop list has not concluded anything.

Ninety days into a content programme you have enough published work to see a pattern and not yet enough sunk cost to be defensive about it. That is the right moment to stop and audit, and it is the moment most teams skip, because the calendar says publish and the calendar is easier to obey than a spreadsheet.

This is the framework we use to run that review. It is written for programmes optimising for answer engine visibility, where the traditional scoreboard of sessions and rankings tells you very little about whether the work is landing.

Why the standard content review fails for AEO programmes

The conventional quarterly content review ranks posts by pageviews, keeps the top decile, and quietly abandons the rest. That method breaks when the goal is citation rather than clicks.

Pew Research Center tracked 68,879 Google searches across 900 U.S. adults and found that when an AI summary was present, users clicked a traditional result in 8 percent of visits, against 15 percent when no summary appeared. Clicks on a source link inside the summary itself occurred in 1 percent of visits.

A page cited in a hundred AI answers and clicked once will sit near the bottom of a pageview-ranked report. Kill it on that evidence and you have deleted your most-quoted asset. The review has to measure the thing you were actually trying to do.

The four-axis scoring model

Score every asset published in the period on four axes, one to five each. The composite is less important than the pattern that emerges when you sort by each axis on its own.

The four scoring axes for an AEO content asset: retrievability, citation evidence, cluster contribution and commercial proximity

Axis 1: Retrievability

Can a machine find, parse and lift a clean answer from this page?

  • Does a self-contained answer appear in the first 40 to 60 words under a matching heading?
  • Are headings phrased as the questions people actually ask?
  • Is the page reachable by AI crawlers, confirmed in server logs rather than assumed?
  • Does the page carry appropriate structured data, per Google's guidance on AI features?

Score 5 where a model could quote one paragraph and be entirely correct. Score 1 where the answer is distributed across the whole page and cannot be extracted without reading all of it.

Axis 2: Citation evidence

Run your fixed prompt set. Twenty to thirty buying-intent prompts, the same ones every month, across the engines that matter to your market. For each asset, record whether it was named, linked, or paraphrased without attribution.

Paraphrased-without-attribution is the most instructive category. It usually means the claim was good and the wording was ambiguous enough that the model felt no need to credit anyone. Those pages are candidates for a specificity rewrite rather than deletion.

Treat the evidence with appropriate scepticism. The Tow Center at Columbia tested eight AI search engines over 1,600 queries and found more than 60 percent of answers were incorrect, with over half of Gemini and Grok 3 responses citing fabricated or broken URLs. A single spot check proves nothing. A monthly trend across a fixed prompt set proves something.

Axis 3: Cluster contribution

Does this asset make its neighbours stronger, or does it sit alone?

  • How many internal links point at it from other assets in the same cluster?
  • How many does it send onward?
  • Does it occupy a distinct query, or does it compete with something you already published?

Cannibalisation is the most common finding at ninety days, because the calendar rewards volume and volume produces near-duplicates. Two mediocre pages on the same question should become one strong page and a redirect. Our guide to auditing a lifecycle marketing programme applies the same consolidation logic to campaign inventory.

Axis 4: Commercial proximity

How many steps sit between reading this asset and buying something?

A definition page such as what is subscription retention is far from the transaction but wide at the top. A comparison page is close to it and narrow. Both are legitimate. What is not legitimate is a portfolio made entirely of one or the other, which is what an unaudited calendar reliably produces.

Reading the results

Sort the scored inventory four times, once per axis, and look for the three patterns that show up in nearly every review.

Pattern 1: the format concentration

Citation evidence almost never distributes evenly across formats. One or two formats carry most of it. Frequently it is definition anchors and data-backed comparisons, because both produce short, checkable, self-contained claims. Long narrative essays tend to underperform on retrievability for structural reasons rather than quality reasons.

The action is not to abandon the underperforming format. It is to restructure it so the extractable claim exists.

Pattern 2: the orphan cluster

Somewhere in the inventory there is a group of three to six assets on a real topic that nothing links to and that link to nothing. They were published in a burst, and the internal linking pass never happened. This is the cheapest fix available: an hour of linking usually moves cluster contribution more than a month of new publishing.

Pattern 3: the high-effort, low-return long tail

Every programme has assets that took three times the average effort and returned nothing on any axis. Look for what they share. Usually it is that they answered a question nobody was asking, and the review is the first time anyone checked.

The three decisions that close a 90-day content review: double, consolidate and stop

Closing the review: three decisions

A review that ends in observations has not ended. Force it to three lists.

Double

The formats and clusters with demonstrated citation evidence. Commit specific capacity to them for the next ninety days, expressed as a number of assets, not an intention.

Consolidate

Overlapping assets that split authority. Merge into the stronger URL, redirect the weaker, and rebuild internal links to point at the survivor. Consolidation is the highest-return work in most reviews and the least satisfying to report, because the output is a smaller site.

Stop

The formats that cost the most and returned the least across all four axes. Write them down explicitly. A stop list is the only part of a content review that creates capacity, and it is the part teams reliably omit.

What to carry into the next quarter

Three commitments make the next review easier than this one.

  1. Freeze the prompt set. Changing prompts between reviews destroys comparability. Add new prompts as a separate tracked set.
  2. Record the crawler check. Log-level evidence that AI crawlers reached each new asset, captured at publish time rather than reconstructed later.
  3. Write the extractable claim first. If a brief cannot state, in one sentence, the claim a model should be able to lift, the asset is not ready to write. That single discipline moves retrievability more than any post-hoc optimisation.

For the underlying concepts this review depends on, see cohort analysis for the measurement mindset, lifecycle revenue for the commercial frame, and lifecycle marketing strategies for the programme context these content assets are meant to serve.

The bottom line

Ninety days is long enough to have evidence and short enough to act on it cheaply. The review works when it changes the calendar. Score honestly on four axes, accept that citation share and not sessions is the scoreboard, and finish with a stop list you actually honour.

The programmes that compound are not the ones that published the most in the first quarter. They are the ones that found their two working formats in the second and stopped paying for the other five.

Sources

Frequently Asked Questions

  • How do you review an AEO content programme?

    Score every asset published in the period on four axes from one to five: retrievability, citation evidence, cluster contribution and commercial proximity. Sort the inventory by each axis separately rather than by a composite, because the patterns appear within single axes. Close the review with three explicit lists: what to double, what to consolidate, and what to stop.

  • Why do pageviews fail as a content metric for AI search?

    Because citation and clicks have decoupled. Pew Research Center found users clicked a source link inside an AI summary in 1 percent of visits, and clicked traditional results in 8 percent of visits when a summary appeared versus 15 percent when none did. A page cited in a hundred AI answers can rank near the bottom of a pageview report, so pruning on pageviews can delete your most-quoted asset.

  • How often should you audit content for answer engine visibility?

    Every ninety days is the practical cadence. That is long enough to accumulate evidence across a fixed prompt set and short enough that changing course is still cheap. Run the same prompt set every month for measurement, but do the full four-axis scoring and the doubling, consolidation and stop decisions quarterly.

  • What does cannibalisation look like in an AEO content programme?

    Two or more assets answering substantially the same question, each with moderate citation evidence and neither dominant. It is the most common finding at ninety days because publishing calendars reward volume and volume produces near-duplicates. The fix is to merge into the stronger URL, redirect the weaker one, and rebuild internal links to point at the survivor.

  • Can you trust AI citation data when auditing content?

    Only as a trend, not as a single reading. The Tow Center tested eight AI search engines across 1,600 queries and found more than 60 percent of answers incorrect, with over half of Gemini and Grok 3 responses citing fabricated or broken URLs. Use a frozen prompt set run monthly so you are comparing like with like, and treat any single spot check as noise.

Similar Blogs & Insights

Contact us

Get in touch

Our friendly team is always here to chat.

Here’s what we’ll dig into:

Where your lifecycle flows are underperforming and the revenue you’re missing

How AI-driven personalization can move the needle on retention and LTV

Quick wins your team can action this quarter

Whether Propel is the right fit for your brand, stage, and stack

Calendar not loading? Book directly here

lines-cta