Back to GEO LibraryAUDIT / SEARCH AND ANSWER SYSTEMS

AI search audit: what to check first

An AI search audit should locate the first broken link in the evidence chain. Publishing more content before that diagnosis can add cost without improving eligibility, accuracy or demand.

Updated
30/08/2026
Publisher
NobleJackal
Version
1.0
Source review
Primary sources reviewed
01

Freeze the question before collecting evidence

An audit begins with scope, not a crawler. Record the domain, canonical host, markets, languages, priority services and the commercial questions the business needs answered. Decide whether the review concerns technical eligibility, non-branded search discovery, brand accuracy, source citations, referrals or qualified enquiries. If those outcomes are mixed together, the final report will make activity look like progress.

Record the date and the available access as well. A public-site review can inspect rendered pages and HTTP behaviour. Search Console adds indexing and query evidence. Analytics can connect visits to actions. Server logs can show verified crawler requests. A model response is useful only as a dated sample. Each source answers a different question.

02

Test the eligibility layer first

Confirm that important URLs return the intended status, are not blocked, contain indexable text and declare a coherent canonical. Follow redirects to their final destination. Compare internal links, sitemap entries and canonical URLs rather than reviewing each file in isolation. A page included in a sitemap but canonicalised elsewhere is not a clean signal.

For multilingual sites, check every equivalent set as a unit. Each page needs the correct language, a same-language self-canonical and reciprocal alternates. Forced IP redirection can hide language versions from people and crawlers; stable URLs and visible language controls are easier to audit. Arabic also needs a genuine right-to-left reading experience, not only a translated string inside a left-to-right layout.

03

Reconcile the entity before optimising the wording

Collect the organisation name, legal operator, founder, services, markets, contact routes and other material facts from controlled sources. Then compare the homepage, company profile, service pages, author pages, structured data, public profiles and machine-readable files. Record contradictions instead of selecting the version that sounds most impressive.

Answer systems often have to resolve relationships between a brand, its legal entity, an individual, a sister company and a delivery partner. If the website blurs those relationships, more copy can amplify the error. The audit should identify the canonical fact, its owner, its evidence and the surfaces that require correction.

04

Evaluate whether the content deserves retrieval

A technically available page may still add little. Review whether it answers a recognisable question, defines the terms it uses, gives the reader enough context to act and shows where material claims come from. Pages that repeat category language without experience, method or evidence are easy to replace with another result.

Map existing pages before proposing new ones. Two articles competing for the same purpose can weaken internal clarity, while a missing comparison or buyer guide can leave a valuable decision unanswered. The recommendation may be to expand, merge, redirect or retire a page. An audit that always concludes ‘publish more’ has not examined the cost of duplication.

05

Separate crawler policy by purpose

Search retrieval, user-requested fetching and model training are different choices. For OpenAI, OAI-SearchBot concerns ChatGPT Search discovery, GPTBot concerns potential training use, and ChatGPT-User can be triggered by a person. Similar distinctions exist across platforms and may change, so the audit should use current publisher documentation rather than a copied robots template.

Check the CDN and firewall as well as robots.txt. A permissive file does not help if the hosting layer returns a challenge, rate limit or error to the crawler. Conversely, a successful request proves access on that occasion, not indexing or citation. The report should retain that boundary.

06

Sample AI answers without manufacturing a ranking

Build the query set before running it. Include branded, category, problem, comparison and method questions that a real buyer might ask. Fix the language, market, product surface, account state and collection date. Store the full answer and every cited URL, including unfavourable and absent results.

Score observations by type: mention, linked citation, accurate description, inclusion in a comparison, recommendation and referral. Do not compress them into one unexplained GEO score. Repeatability matters more than a dramatic screenshot, and a small sample should remain labelled as a small sample.

07

End with an action register, not a scorecard

Every material finding needs an affected URL, evidence, owner, priority, proposed change and verification method. Technical defects can be tested after release. Editorial claims require source and approval. Platform outcomes require time for recrawl and a dated observation window. Unknowns should stay unknown until the necessary access or evidence exists.

The final sequence should follow dependency, not excitement: remove hard eligibility barriers, correct entity facts, repair internal architecture, improve priority decision pages, then expand content where a distinct need remains. Measurement continues after publication. A launch is an intervention date, not proof of an outcome.

Practical checklist

  • Record scope, markets, languages and available evidence
  • Verify status, crawl, index, canonical and hreflang conditions
  • Reconcile organisation, person, service and partner facts
  • Map content purpose before adding URLs
  • Test search, retrieval and training policies separately
  • Preserve exact AI-answer samples and cited URLs
  • Assign every action an owner and verification state
FROM RESEARCH TO DELIVERY

Apply the GEO evidence to your own public system.

NobleJackal examines the path from crawl and indexing to entity clarity, source retrieval and measured AI-surface observations—without promising citations or positions.

Review the GEO agency service for Europe

Official sources

  1. Google Search CentralOptimizing your website for generative AI features on Google Search
  2. Google Search CentralAI features and your website
  3. Google Search CentralTell Google about localized versions of your page
  4. OpenAIOverview of OpenAI crawlers
  5. IndexNowIndexNow protocol documentation