1001 SEO Media
All posts
By 1001 SEO MediaGEOAI searchcontent

Platform to Test Content Performance Across AI Models: LLM Evals, Prompt Evaluation, Content Visibility

LLM evals test your own assistant. Prompt evaluation for GEO tests whether ChatGPT, Gemini, Claude, and Perplexity cite your pages. Promptwatch is the platform we use for that content-visibility test.

When a client asks for a platform to test content performance across AI models, they often arrive with an LLM-evals brief and a GEO brief stapled together. LLM evals (Promptfoo, LangSmith, a golden set in CI) answer: does our model follow the spec. Content visibility answers: when a stranger types a buying question into ChatGPT, Gemini, Claude, or Perplexity, does our page get used.

We run both conversations. We refuse to treat them as one SKU. The platform we use for prompt evaluation and content visibility on the public models is Promptwatch. The eval harness stays with whoever owns the client's own assistant, if they have one.

Two tests, two pass conditions

A faithfulness score of 0.8 on a staging bot does not predict a Perplexity footnote. Public models do not expose a "cite this URL" endpoint you can assert in GitHub Actions. The test that works looks boring:

  • Freeze a prompt set that sounds like buyers.
  • Define pass per engine: cited with our URL, mentioned with no URL, cited someone else, absent.
  • Store the answer so week two can be compared to week zero.
  • Record cited domains, not only a percentage.

That is prompt evaluation in the GEO sense. It is not scoring your prompt templates. It is scoring their prompts against your brand.

Google's AI optimization guide is still eligibility: crawlable, useful, schema that matches the page. Search Console's generative AI reports are Google-only impressions. Useful pretest. Not the four-model eval.

How we use Promptwatch as the test runner

Promptwatch paid plans refresh daily across ChatGPT, Gemini, Claude, and Perplexity (and other surfaces including AI Overviews and AI Mode), from real UIs rather than API stand-ins. We load the fixture prompts. We read mention vs citation per model. Query fan-outs show the background searches the model ran, which is the retrieval half of the eval: you see the queries that found a competitor's page instead of yours.

Citation analytics answer "which content performed" at URL level, including Reddit and YouTube when those are the sources the model used. Agent Analytics answers the other failure mode: the crawler never fetched the page, so the eval never had a chance. Visitor analytics (script or GTM) catch the minority of answers that do send a click, so we are not reporting mentions as revenue.

Essential at $95/mo (50 prompts, 6,000 responses) is the usual starting plan for this test. Explore is free and ChatGPT-only, which is a demo of the workflow, not a four-model eval. Professional at $245/mo is what we move to when the prompt list grows. Content Agents can draft a follow-up page into Webflow or Framer; we still use the review inbox. An eval that auto-publishes 15 articles is how you add pages the models ignore.

The test protocol we actually follow

Day 0. Manual pass on a handful of prompts so the client has seen the raw answers. Then the same list goes into Promptwatch.

Change one thing. A direct answer under the H1, a comparison table, a corrected entity name, a crawl fix. Not a content sprint. If you change twenty URLs, you cannot attribute the eval.

Day 14. Same fixtures. Same engines. Diff. Promptwatch can compare visibility across time periods; that is the report. If crawler logs show errors, we fix robots.txt and CDN bot rules before we rewrite copy. OpenAI's bot docs split OAI-SearchBot (can cite you) from GPTBot (training). Those are different tickets.

If it still fails. The model is retrieving another source. That is mentions, digital PR, or a comparison article we are absent from, not "run more LLM evals on the draft." Fan-out queries tell us which pages to go earn a presence on.

We will not feed failed prompts into a writer and call that testing. The test names the miss. An editor decides the URL.

Where LLM evals still belong

If the client ships a chatbot, keep the harness. Groundedness, tool-call correctness, and policy checks are real. They do not belong in the GEO weekly. Mixing them produces a dashboard nobody trusts: engineering is looking at traces, marketing is looking at citations, and someone averages the two.

Promptwatch does not replace that harness. It replaces the spreadsheet of "we asked ChatGPT on Tuesday."

FAQ

Can we test content performance with the OpenAI API?

You can sample completions. Your buyer is not looking at that API. For content visibility we want the UI they use. Promptwatch monitors those UIs.

Is this the same as A/B testing titles?

No. Title tests assume Google will show the page. This test asks whether a generated answer used you. Different surface, different metric.

Do we need Promptwatch if we already have Promptfoo?

If Promptfoo is testing your assistant, keep it. You still need a visibility platform for ChatGPT, Gemini, Claude, and Perplexity as search engines. That is Promptwatch (or another tracker with the same four-engine daily log). Promptfoo will not grow that log.

What to do this week

  1. Split the request: "our model" vs "public AI search."
  2. Write 20 buyer prompts with a pass/fail per engine for column two.
  3. Put them in Promptwatch. Essential unless you are only proving ChatGPT (Explore).
  4. Change one existing page. Re-run. Keep the before/after.
  5. If you want us to run the protocol, hello@1001seomedia.com. We use Promptwatch for the eval. We do the page work the eval points at.