ChatGPT cited URLs that were 458 days newer than Google’s organic results—the strongest freshness preference of any platform we tested.
This study doesn’t contradict that narrative, but it does add an extra layer of nuance.
For instance, when we look at the search index, cited pages span a wide range of ages—the median is around 500 days (~1.3 years old), with some cited pages over 2,700 days old (~7.4 years old).
The median age is actually far lower than our initial freshness study linked above (958 days back in July vs 500 days in this dataset), suggesting that ChatGPT is skewing even younger in its citation preferences.
That said, we also found that non-cited pages are overwhelmingly very young.


So within a single prompt’s retrieval set, it’s the older, more established pages that tend to get cited, and the freshest content that tends to get discarded.
In other words, ChatGPT prefers fresh content, but tends to cite comparatively “older” content more often. That sounds counterintuitive, but both things can be true at the same time.
Across the broader population of AI citations, ChatGPT does skew fresher when compared against Google results, and even against it’s own citation preferences from only last year.
But within a given retrieval set, freshness alone isn’t enough. Relevance still does the heavy lifting.
A new page that matches fanout queries well will get cited. A new page that doesn’t will be retrieved, yet ignored.
It’s also worth pointing out that the pool of non-cited pages (~3M) across the search ref_type is far smaller than the cited group (~23M), which limits how confidently we can interpret the age gap.
Where freshness matters most is in “news”.
In this category, title relevance scores for cited and non-cited pages are nearly identical:


The AI can’t decide based on relevance alone, so it defaults to a temporal tie-breaker: page age. Cited news pages skew younger:


For news queries, younger pages have a clear advantage, even when relevance scores between cited and non-cited pages are similar.
Create the freshest news content using Firehose
If you publish news or time-sensitive content, freshness is non-negotiable.
Be the first to break news on certain stories using Ahrefs Firehose—our real-time web monitoring API that gives you a streaming feed of data from our huge crawler infrastructure.
For example, if you work in SaaS journalism, you can track content changes on pages like Google’s official blog, so you can be the first one to cover a new Google update as soon as it goes live.


Then, use Brand Radar’s Mentions history in the AI Responses report to track whether your ChatGPT visibility spikes after publication.


What this all means for being “citable”
The 1.4 million prompts paint a pretty clear picture. ChatGPT is an aggressive editor. It favors its general search index, uses semantic similarity to select and cite sources, and treats Reddit as a textbook it’s embarrassed to admit it read.
But the data also taught us a lesson in analytical caution.
Aggregate comparisons between “cited” and “non-cited” URLs can be misleading if the non-cited pool is dominated by a single source type with its own retrieval mechanics.
What initially looked like a paradox—less-optimized pages getting cited more—turned out to be a matter of dataset composition.
We would have got that one very wrong if we hadn’t isolated by ref_type.
Ultimately, the pages that get cited are the ones whose titles and content match the questions ChatGPT is asking behind the scenes, and that surface through the right retrieval channel.