Intelligent Engine Optimization (IEO) measures AI citation events from raw server access logs, using IEO Citation Tracker. Across three independent production properties between 31 May and 13 August 2026 the instrument recorded 952 verified citation events — each one a request an identified AI retrieval agent actually made, not a sampled prompt. This page states what those events can establish, and what they cannot. Version 2.0, 16 August 2026.
Three production properties · 31 May – 13 August 2026 · Version 2.0, 16 August 2026 · three claims retracted since v1.0
Read this first. Version 2.0 is shorter than version 1.0 because three published claims were withdrawn under adversarial review, not refined. A median latency figure, a correction to published waiting-period guidance, and a mechanism claim about same-day citation were each defeated and removed. What remains is what raw server access logs can support without a survival estimator, and it is less than this page claimed a week ago.
Every correction is dated in §9. The direction has been consistently unfavourable: every revision lowered a figure, narrowed a claim, or added a limit. That is not offered as a virtue. It is offered so the revisions can be reproduced from the raw observations and the stated rules, which is the only thing that makes any of it checkable.
Every widely quoted figure in AI-search measurement comes from one of two instruments: prompts run against an engine and scored for brand mentions, or vendor-side telemetry. Both are inference. Prompt polling cannot separate genuine channel volatility from its own sampling noise, and it measures queries an analyst invented rather than queries a person asked.
A server access log is the other side of the same event. When a retrieval agent fetches a page to answer a live question it leaves a line. That line is not a proxy for a citation — it is the fetch a citation is made from. It is also the only measurement surface that does not require an AI platform to share anything.
The trade, stated exactly. Log measurement sees every retrieval of your own pages and nothing about anyone else’s. Prompt polling sees the whole answer for a query nobody asked. Neither is complete. This work takes the log side because it is verifiable, and it accepts the consequence: an observed fetch establishes that an identified agent requested a page. It does not establish that an answer displayed it. Nothing on this page depends on the second claim.
Most published ratios in this field omit their denominator, which is why they cannot be compared with each other. Ours are stated here and used consistently. Where a convention differs from prior work, the difference is named rather than left to be discovered.
ChatGPT-User,
Claude-User, Perplexity-User, Shap-User, Google-Agent,
DuckAssistBot, YouBot, MistralAI-User. Verified against the engine’s published IP
ranges; a matching user-agent string alone is not sufficient and is counted separately as unverifiable.GPTBot, ClaudeBot,
CCBot, Bytespider, Amazonbot, Google-Extended. Nobody is waiting..html path that is not root, not protocol infrastructure, and not a structural or
taxonomy endpoint. See §2.1.Protocol infrastructure — excluded entirely. robots.txt, sitemaps, favicons and
/.well-known/ probes. A retrieval agent asking permission is the opposite of citing you, and none of these files
can appear in an answer because none is content. Applied retroactively to stored data, which lowered previously published
counts.
Root paths — excluded from the citation class. A request for / is a real retrieval event, and it
is administrative and navigational far more often than it is evidence of a specific citation: an agent fetches the root to
resolve a brand, verify a domain, find navigation, or fall back when a deep link fails. An earlier version of this page
counted root fetches as weak citation evidence. That was too generous. Root fetches are now reported separately, alongside
protocol infrastructure, and are not counted toward any content-page figure.
Structural and taxonomy endpoints — segregated, and not yet classified. Category archives, tag pages, pagination and feeds are HTML and are not articles. An agent hitting a category page to reach an article is performing navigation, not citing the category. These are being separated into their own denominator. Until that classification is complete, no content-page ratio on this page should be treated as final, and the figures in §3 carry that caveat.
Latency is measured from the first crawl visible in this log window, which is not necessarily the engine’s first contact with the page. A page crawled before 31 May, then fetched again on 2 June and cited on 3 June, records as a one-day interval when the real interval may be months.
The guard that existed was inadequate. Pages whose earliest observed fetch fell within two hours of the log’s first line were excluded. That catches only the boundary. A page crawled five days before the window opened passes straight through it, and every such page pulls measured latency down.
The correct control is a burn-in period. The first 14 to 30 days of the window serve only to register baseline crawl history, and only pages whose first observed crawl falls after that period enter the latency cohort. This has not yet been applied to the figures below. Every latency-related number on this page predates it and will change when it is computed. The cohort will be substantially smaller.
| IEO Engine | MoldMunchers | BubblesInTime | |
|---|---|---|---|
| Vertical | methodology | local service | consumer app |
| Window | 31 May – 13 Aug | 30 Jun – 13 Aug | 30 Jun – 13 Aug |
| Days | 75 | 45 | 45 |
| Total requests | 16,898 | 44,639 | 13,086 |
| AI-agent requests | 8,165 | 8,127 | 4,845 |
| Citation events | 339 | 571 | 42 |
| Days with ≥1 citation event | 69 / 75 | 45 / 45 | 21 / 45 |
| Pages crawled by AI | 289 | 454 | 357 |
| Pages with ≥1 citation | 47 | 78 | 10 |
| Crawl-to-citation yield | 16.3% | 17.2% | 2.8% |
Portfolio total: 74,623 requests, 21,137 AI-agent requests, 952 citation events. Yield figures are page-level and predate both the burn-in control (§2.2) and the structural-endpoint classification (§2.1).
Note the AI-agent row. The three properties received 8,165, 8,127 and 4,845 AI-agent requests despite MoldMunchers carrying nearly three times the total traffic of the other two combined. AI crawl volume tracks corpus size, not site popularity. The extra requests on MoldMunchers are commercial SEO crawlers and scrapers, not AI.
Comparing the training crawler to the live retrieval agent from the same vendor over the July window:
| Property | GPTBot | ChatGPT-User | Ratio |
|---|---|---|---|
| MoldMunchers | 271 | 287 | 1 : 1 |
| IEO Engine | 527 | 145 | 4 : 1 |
| BubblesInTime | 1,127 | 8 | 141 : 1 |
BubblesInTime was the most heavily ingested property in the set and the least retrieved. Heavy crawling is not a leading indicator of citation. The ratio is computable from one month of logs by anyone, and it is the single most useful thing on this page for a practitioner.
Stated as a limit: three properties in three verticals is not a sample from which the magnitude of that spread generalises. What it shows is that the two channels move independently, which a single “AI traffic” figure conceals.
Of 135 completed crawl-to-citation pairs — pages with both a first observed crawl and a subsequent citation-class retrieval inside the window — 31, or 23%, occurred within one day. That is a description of the completed set, not a rate over the crawled cohort.
Retracted: the median. Version 1.0 published a median of 14.0 days over those completed pairs, first as a point estimate and then as an upper bound. Both were wrong. Pages crawled and not yet cited are right-censored — still at risk, eventual interval unknown and possibly never. Computing a median from completed observations while excluding the active risk set is not a biased estimate of the cohort median; it is a different quantity that has been mislabelled. Disclosing the censoring does not repair it.
Computed 16 August 2026. Kaplan-Meier over a risk set of every distinct content page with a qualifying observed crawl — 1,105 pages, 136 events, 969 right-censored. The median is not reached. The survival function never falls to 0.50 because only 15.9% of crawled pages are ever cited inside the window. A median time-to-first-citation is undefined on this data. Full result, per-property breakdown and burn-in sensitivity: FN-009.
Cumulative share cited: 2.9% by one day, 4.2% by seven, 6.2% by fourteen, 10.6% by thirty, 15.0% by sixty. The derived observation set is published as a CSV so the curve can be recomputed. Under a 14-day burn-in the curve holds and same-day citation falls to 1.2%, which is what a left-censoring correction should do.
And the survival curve estimates a narrower thing than it appears to. It describes intervals among pages that were crawled. A page never crawled never enters the risk set, and that selection happens before the measured process begins. No figure derivable from these logs estimates the probability that an arbitrary published page is eventually cited.
What the same-day count does and does not establish. It is inconsistent with a model in which every citation requires a multi-day batch cycle after first contact as a necessary condition. It does not identify the path. Two remain live — a fetch triggered by a user at answer time, and insertion into a live retrieval index followed by selection — and server logs cannot separate them. A third possibility is not excluded either: a batch cycle begun before the window opened, completing inside it. The burn-in control in §2.2 exists to address that and has not yet been applied.
MoldMunchers produced at least one qualifying citation event on each of 45 consecutive days — 571 events across 78 distinct cited pages, roughly twelve a day, no blank day. IEO Engine ran 23 consecutive days and recorded events on 69 of 75.
This is a statement about observed daily activity and nothing more. It confirms the origin server was contacted by a citation-class agent every day in that interval. It does not establish continuous engine attention, stable ranking, uniform corpus coverage, or per-URL persistence — a rotating handful of frequently queried pages could sustain the streak while the rest of the corpus went untouched. An earlier version described this as the property remaining “inside the active retrieval loop”. That was a model of the mechanism, not an observation, and has been withdrawn.
Published figures put answer-surface citation turnover at 40–60% month over month with a half-life near ten days. Those measure a different layer — whether a URL renders as a visible citation for a given prompt. Both can be true at once, and if they are, the instability sits in selection and rendering rather than in whether engines still retrieve the corpus. That is a hypothesis this data is consistent with, not a finding it establishes.
The top-cited page on each property is a direct-answer page. On MoldMunchers, nine of the top twelve cited pages are cost or pricing guides, led by a 646-word comparison table at 69 citation events. On IEO Engine the leaders are a platform-specific method guide and a definition page. Pages that argue, narrate or describe accumulate crawl and earn nothing.
This is consistent with published research finding that statistics, citations and quotations produce the largest visibility gains. It is offered as an independently measured agreement, not an original finding, and it is descriptive: no control distinguishes content structure from topic demand.
Published crawl-to-citation ratios — 887:1, ~89% training versus ~8% retrieval, 20,000 pages per referral, 81.7% training-oriented — are request-weighted, and none states whether unidentified non-AI traffic was excluded.
Worked counter-example. A single MoldMunchers page took 1,445 requests in one month, 4.6% of the entire site’s traffic, from a residential-proxy pool:
1,445 requests status 200 on every one
1,385 (96%) identical Mac/Chrome user-agent
1,410 (98%) no referrer
~7,485b byte-identical response every time
hundreds of IPs dispersed across unrelated /16 blocks
one request every 2-5 minutes, around the clock
ramp 13 Jul -> peak 397 on 18 Jul -> gone by 25 Jul
Request-weighted, that page distorts the site ratio. Page-weighted, it moves the denominator by one.
The precise claim. Page weighting prevents repeated requests to one URL from mechanically dominating a page-level statistic. It does not eliminate selection bias in which pages enter the observed sample, and it is not immune to breadth inflation: a scraper fetching four hundred distinct pages once each enters a page-weighted denominator four hundred times. Page-weighted coverage is coverage of observed crawled pages, never of all published pages.
Both conventions are kept deliberately. The yield is page-weighted, because depth attacks are the common case. Contamination detection is request-weighted, because requests-per-page exposes a depth attack and requests-per-IP-per-page exposes a breadth one. A methodology that discards request-weighting loses its own contamination detector. Naming the weighting is necessary and not sufficient: crawl-budget behaviour and cache-invalidation horizons differ between crawlers, and two correctly-labelled ratios can still be non-comparable.
Published guidance: AI Overviews generally do not trigger a real-time crawl from a unique bot, and you will not see specific Google AI traffic in your logs.
Google-Agent appears in these logs as a citation-class fetcher, on all three properties, timestamped and
attributable to specific pages. That alone falsifies the general claim.
IEO additionally detects and counts Google AI Overview participation as its own evidence class. The detection method is not published, and that is a deliberate exception to how the rest of this page works — stated, with its own falsification protocol, at Google AI Overview is measurable. No figure on this page depends on it.
Crawl-to-citation yield across the three properties: 16.3%, 17.2% and 2.8%. Inverted, 83.7%, 82.8% and 97.2% of crawled pages were never retrieved by a citation-class agent.
The level is close to uninformative. Most pages on any site answer no question anyone is asking, and long-tail human curiosity would produce a large never-retrieved figure on a perfectly healthy corpus. That number on its own is not a diagnostic and is not offered as one.
The 14-point spread is a descriptive difference, not an anomaly. An earlier version called it an anomaly, which implies a cause. It has none established. The three properties operate in unrelated verticals with different query demand, competitive density, content depth, internal linking and indexation state, and any of those could produce a gap of that size. Comparing them without controls is comparing unlike things.
The control that would make this a finding is available and not applied: cross the never-retrieved set against Search Console position and impressions. A page ranking in the top ten, receiving impressions, crawled by AI agents and never retrieved is not explained by absent demand. That subset would be a content-structure diagnostic. Until it is computed, this section reports variance and claims nothing about its cause.
Three claims published in version 1.0 were defeated under review and removed rather than narrowed. They are recorded here because a page that publishes only what survived is not auditable.
Withdrawn: “median crawl-to-citation latency is 14.0 days.” And its successor, “at most 14.0 days.” A completed-pair median on right-censored data is the wrong quantity, and the upper-bound framing additionally had the direction of bias backwards: excluding pages not yet cited removes long intervals, so the completed-pair figure understates rather than overstates. See §4.2.
Withdrawn: “the 60-to-90-day waiting guidance is wrong by four-fold.” And its successor, “wrong as a necessary minimum latency.” That guidance is a decision threshold about when to conclude an intervention has failed — it concerns site-wide indexation and query-space stabilisation, not the minimum possible interval for a single page. Fast individual citations refute nothing anyone claimed. The narrowed version was a strawman and went with it.
Withdrawn: “no update cycle can produce a same-day citation.” This conflated foundation-model training updates with insertion into a live retrieval index. Real-time ingestion can crawl, embed and serve within seconds. See §4.2 for what replaced it.
Four hypotheses this project held and its own data destroyed, published because reporting only what survived would make every number above unauditable.
Predicted: pages asserting IEO Engine’s own methodology would be crawled and ignored while pages defining the field’s general concepts earned the citations. Measured: brand-specific pages 10 / 84 = 11.9%; general-concept pages 20 / 214 = 9.3%. Brand pages are 28% of the crawled corpus and take 35% of citation fetches. Result: over-represented, not under. The content strategy built on the prediction was abandoned.
Predicted: a page taking 1,445 crawls with zero retrievals was the flagship case of heavy AI ingestion producing nothing. Measured: not AI at all — the residential-proxy scraper in §5.1. Result: the headline example was wrong, and the investigation that corrected it produced the weighting argument, which is more useful than the claim it replaced.
Predicted: a citation-class agent requesting a URL that had never existed suggested the model held a structural map of the site. Measured: the page had existed and was removed in a content cull weeks earlier; a companion cluster of 27 similar 404s resolved the same way. Result: ordinary crawler memory of deleted URLs. Post-deletion crawl persistence ran at 6.6% and 2.0% of subsequent traffic on two properties a month after removal, which is worth measuring — but the original interpretation was unsupported.
Predicted: pages crawled but uncited within the observation window were dead weight and safe to remove. Measured: a cull removing 21% of one property’s corpus deleted three of the nine pages that had earned citations; a second property lost one cited page of 31. Result: any cull whose observation window is shorter than the citation interval will delete working pages. One deleted page was later restored on log evidence and has since earned citations again.
One measurement hazard worth publishing because it corrupts data silently: shared hosting can interpose a bot-verification
interstitial that returns HTTP 200 with a cacheable HTML body rather than a 429. Retrieval agents received these on
/robots.txt and /sitemap.xml repeatedly across all three properties. They are detectable by response
size on paths whose true size is known, and any log-based instrument that does not screen for them counts them as served pages.
The instrument is a text file on a shared host. The ingestion-versus-retrieval split in §4.1 — the most useful figure here — derives from two commands:
zgrep -c 'GPTBot' sslaccesslog*.gz # ingest class zgrep -c 'ChatGPT-User' sslaccesslog*.gz # citation class
Three checks before trusting the output. Screen for interstitials — compare response sizes on
/robots.txt; a cluster of ~7 KB responses on a file of a few hundred bytes means your host is serving
challenge pages as 200s. Screen for proxy pools — sort by request count per URL; a single page taking a
disproportionate share with one repeated user-agent, no referrer and a constant response size is a scraper. State your
window — any figure computed over fewer days than your citation interval will misclassify working pages as dead,
and any latency figure computed without a burn-in period is biased downward by an unknown amount.
Continuous measurement, including the classes these commands cannot separate, is what IEO Citation Tracker does.
| Version | Date | Change |
|---|---|---|
| 1.0 | 14 Aug 2026 | First publication. Window 31 May – 13 Aug 2026. |
| 1.1 | 16 Aug 2026 | Six corrections: distribution renamed from bimodal to zero-inflated; median restated as a censored-window floor; “first crawl” defined; streak reframed as a layer finding; page-weighting narrowed to depth inflation; never-retrieved reframed from level to spread; two scope limits added. |
| 1.2 | 16 Aug 2026 | Three further corrections: the same-day mechanism claim withdrawn as false; “first crawl” renamed first observed crawl; the four-fold bound restated as cohort-level. |
| 2.0 | 16 Aug 2026 | Rebuilt after five adversarial rounds across three model families. Three claims withdrawn outright rather than narrowed — the 14.0-day median in any form, the correction to 60-to-90-day guidance in any form, and the same-day mechanism claim. Left-censoring control upgraded from a two-hour boundary guard to a 14–30 day burn-in period, not yet applied. Root paths reclassified from weak citation evidence to administrative infrastructure. Structural and taxonomy endpoints segregated, classification incomplete. Crawl-to-citation yield defined explicitly at page level to prevent an event numerator over a page denominator. “Anomaly” replaced by “variance” throughout. Change-dates relabelled unverified operator-supplied metadata. Audit-trail position corrected: reproducibility establishes transparency, not validity. |
| 2.1 | 16 Aug 2026 | Kaplan-Meier computed and published as FN-009; the median is not reached. Fetch definition corrected from “HTTP 200 or 304” to HTTP 200 with a non-zero body — a 304 sends no content and cannot anchor an interval. Ever-cited rate falls 17.9% to 15.9% as a result. A host-interstitial byte band (6,754–7,279, derived from 292 challenge pages served on /robots.txt) is now screened; it cannot discriminate on content pages, so flagged rows are marked and retained rather than removed. The derived observation set is published with its README so the result can be recomputed independently. |
Figures are versioned and dated. Any figure quoted without its window and its version is quoted incorrectly.
Method note. All figures derive from raw SSL access logs across three independent production properties, 31 May – 13 August 2026, deduplicated across overlapping exports. Denominators are in §2, exclusions in §2.1, windows in §3. No figure comes from prompt polling, vendor telemetry or estimation. Corrections. Withdrawn claims are in §7, falsified predictions in §8, and the full revision history in §11. Reproducibility of those revisions is what makes them checkable; it is not itself evidence that the surviving figures are correct.