What the Surferstack experiment means for client GEO programs
A practitioner take on the Surferstack experiment, and what we actually change in client programs because of it.
We pay attention to experiments like Surferstack because they test the assumptions we run client programs on. Surferstack is a listicle website built entirely with Claude Code, and its CDN was connected to Promptwatch so the team could watch every AI crawler and every citation. In a July 2026 post, the founder noted that ChatGPT and Claude citations arrived at volume while Google Search Console showed essentially nothing, and that Bing correlated better with the AI results than Google did. The LinkedIn post on the Surferstack experiment is the source.
This is one site, and it is a listicle directory rather than a client site, so we do not treat it as a universal result. But it confirms something we already see in client work: a site can earn AI citations without earning traditional Google rankings first, and a measurement setup that only watches Google will miss the entire category. Here is what we actually change in client programs because of it.
We lead with crawler logs, not rankings
The first change is that we start every engagement with AI crawler access, not with a rank tracker. The how we audit AI crawler access before any GEO work guide is the process we run, because if the crawlers cannot read the site, nothing downstream works. The Surferstack result is the extreme version of why this matters: a site that gets cited needs to be crawled first, and the crawl is the signal you can actually measure from day one.
We run this on Promptwatch because its Agent Analytics crawler logs cover ChatGPTBot, ClaudeBot, PerplexityBot, GoogleOther, and the Meta AI crawler, with a crawl to citation path that connects a specific crawl to a specific citation. That path is what turns a log into a diagnosis rather than a list of hits.
We report on the three layers, not on one score
The second change is in how we report to clients. A single AI visibility score is a vanity number, and we do not use it. The what we promise clients on AI visibility guide is the honest version of what an engagement can and cannot guarantee, and the reporting follows the three layers: crawler logs, citation analytics, and visitor analytics. The attributing AI traffic measurement stack guide is how we close the loop on the traffic layer.
The Surferstack observation is the cleanest argument for this. A site with zero Google Search Console presence was earning AI citations at volume. If we reported to a client only on the Google layer, we would conclude nothing was happening, and we would be wrong. The citations were happening at the citation layer, fed by crawls at the crawler layer, and the traffic was arriving through a channel Google does not report.
We treat Bing as a second signal, not a measurement
The third change is that we watch Bing alongside Google, but we do not pretend it is a measurement. The Surferstack observation that Bing correlated better with AI citations than Google did is useful, because it means Bing Webmaster Tools is a better free signal than Google Search Console for this category. But a proxy is not a measurement, and we tell clients that.
The actual measurement is prompt level citation data. That means tracking which prompts return client pages as citations, how often those prompts are searched, and how difficult they are to win. The content gap analysis for AI answers guide is our process for turning that data into a content brief, which is the step that makes the measurement actionable.
We do not copy the listicle format blindly
The Surferstack site is a listicle directory, and listicles are a citation friendly format. The broader citation type data, summarised in the ChatGPT citation types for July 2026 report on our sister site, found product pages led ChatGPT citations that month at roughly a third of all citations, with listicles the fastest growing format. We use this to decide content format per prompt type, not to default every client to listicles.
A B2B SaaS client usually needs product pages and comparison pages more than listicles. A publisher client often benefits from listicles. The point is that the format decision follows the citation data for that client's prompts, not a blanket rule. The Surferstack guide to SEO for generative AI summaries is a useful read on the strategies that hold up, but the application is always client specific.
What we do not change
We do not change the part of the program that is already working. The 90 day GEO program guide is still the structure we use, because the structure is sound. The Surferstack experiment changes the measurement and the format emphasis, not the engagement model. We still start with an audit, move to gap analysis, produce content, publish, and measure. The difference is that the measurement now covers the three layers from the start, and the format decision follows the citation data.
We also do not promise clients the Surferstack result. A listicle directory built for AI extractability is not the same as a client site with a real brand and a real product. The experiment is a provocation about measurement, not a template for client work. The agency stack for AI search in 2026 guide is the stack we run, and it is built for client sites, not for experiments.
The takeaway for clients
If you are evaluating an agency for GEO work, the question to ask is how they measure. If the answer is Google Search Console and a rank tracker, you are running the Surferstack experiment blind. You will see the traffic that arrives and have no idea where it came from or why. The agencies that do this well are the ones that measure the layer where citations actually happen, and that close the loop between the gaps they find and the pages they publish. That is the standard we hold ourselves to, and the Surferstack experiment is a useful reminder of why.