How GEO Wins Actually Happen: Anatomy of the Published Case Studies
A close reading of the Crisp, Monks, and OpenUp GEO case studies, focused on mechanisms, evidence, and limits rather than testimonials.
GEO case studies are easy to read as promises. A conversion figure, a publishing cadence, or a top-cited page becomes the headline, while the operating work that produced the result disappears. That is the wrong way to use them.
We read case studies for mechanisms. What observation started the program? Which data changed a decision? What did the team publish or fix? Which outcome was measured? We also look for what the page does not establish, because a vendor-published customer story is not a controlled experiment.
The three Promptwatch case studies below describe different companies and different uses of the platform. Crisp, Monks, and OpenUp are not clients of 1001 SEO Media. Their reported results are not forecasts for our clients. Taken together, though, the stories give a useful picture of how a GEO program can move from an AI answer to a piece of work someone owns.
Crisp started with traffic economics
The published Crisp case study says the company noticed referrals from AI models at the end of 2024 and began using Promptwatch in early 2025. According to the case, AI traffic converted at twice the rate of Crisp's normal traffic. It also reports that Crisp scaled to publishing five to ten articles per day and monitored articles for crawler visits and citations.
Those facts belong to one company's report. The page does not provide the denominator, attribution settings, observation window for the conversion comparison, or a controlled counterfactual. We cannot turn Crisp's reported 2x into an expected conversion uplift for another business. We also cannot assume that publishing five to ten articles per day caused the conversion rate. The referral signal existed before the reported scale-up.
The useful mechanism is narrower. Crisp saw a valuable traffic source, measured visibility, inspected which sources appeared in answers, and adjusted distribution and content work. The case says citation data led the team to put more attention on Reddit and YouTube. That is a source-led decision. It is different from deciding in advance that every brand needs a daily article quota.
For an agency, the lesson is to begin with channel evidence. If AI referrals arrive, isolate them, inspect what brought them, and compare their behavior with a clearly defined baseline. Then trace the cited source. Volume comes after that diagnosis, if the organization can sustain accurate publishing.
Monks turned gaps into an operating system
The Monks case study is less about one page and more about managing client programs. It reports that Monks uses citation gaps, visibility and share metrics, and integrated dashboards that combine LLM visibility with traditional organic and paid data. The Answer Gap Report is described as an input to content roadmaps when competitors are cited and a client is absent.
The page opens with a reported market-share change that traditional metrics did not explain. It does not establish that missing AI citations caused that change. We would not repeat the case as proof of a universal causal link between AI visibility and market share. The defensible claim is that the unexplained change prompted Monks to add another measurement layer.
Its mechanism is organizational. A gap becomes visible, the agency identifies the source and query, a content or offsite response enters the roadmap, and leadership sees the work through an integrated report. The dashboard does not create the win. It gives different teams a shared object and prevents AI visibility from sitting in a separate spreadsheet that nobody uses.
There is another useful limit here. A share or visibility metric is an indicator inside a defined set of models and prompts. It is not total market share, and it should not be labeled that way in a client deck. The case is strongest when read as a reporting workflow: combine channels without pretending their metrics are interchangeable.
OpenUp found a topic outside its keyword list
Promptwatch's OpenUp case study describes a discovery problem. The company used prompt and search-volume data to prioritize GEO content. The story says a prompt-discovered topic became the top-cited result for its target prompt. It also reports that crawler logs and visit analytics reduced spreadsheet-heavy work.
This is a specific win: one topic, one target prompt, and a citation position reported in the case. The page does not claim that every OpenUp article became top cited, nor does it provide a broader conversion result for the program. It says leads from AI referrals are a goal. A goal is not the same as a measured outcome.
The mechanism begins before writing. Prompt data surfaced a question that conventional keyword planning had not suggested. The team decided it was relevant, wrote for it, and then checked citations. A generic instruction to write for AI is too vague. A verified question, target URL, and tracked prompt create a test.
Reduced spreadsheet work is an operational outcome. Faster access to crawler and visitor data can shorten diagnosis, but it should be reported as a process improvement unless the company has measured a business effect.
The common chain beneath three different stories
The cases start in different places. Crisp begins with conversion behavior from an existing referral source. Monks begins with an agency measurement gap. OpenUp begins with topic discovery. Their public outcomes should not be blended into one benchmark.
They do share a sequence:
- Observe a specific signal in prompts, citations, traffic, or reporting.
- Locate the gap at the level of a source, topic, model, or workflow.
- Change content, distribution, or measurement based on that gap.
- Watch the relevant prompt, page, crawl, or referral after the change.
That sequence is more transferable than any headline figure. It also explains why generic GEO checklists disappoint. Structured headings, schema, and crawler access may all help, but none tells a team which missing answer matters to its audience.
What a stronger GEO case study would include
When evaluating a case study, we want the fixed prompt set, models, locations, baseline dates, release dates, and exact outcome definition. For traffic claims, we want attribution rules and a comparison period. For citation claims, we want to know whether top cited refers to one prompt, a topic group, or the full monitor. We also want simultaneous changes recorded, including site migrations, product launches, and model updates.
Published customer stories rarely include every field. That does not make them useless. It limits the claim. A story can demonstrate that a workflow was possible without proving it will reproduce elsewhere.
This is how we use Promptwatch in client programs: as the measurement and diagnosis layer around work we can name. We freeze a baseline, log changes, follow pages from crawl to citation, and keep referral outcomes separate from visibility scores. A defensible win ties a relevant change to an owned action, with enough context to say what the evidence supports and what it does not.