Direct answer
To test a GEO claim, first record the exact claim and the entity being measured. Then preserve the scope, baseline, change made, fixed question set, system and version. The language, country, dates, raw answers, sources, assessment method and retest result must remain in the same evidence chain.
In plain language
Saying, 'We improved,' is not enough. You need the earlier mark, the work performed, the later result from the same test and the method used to calculate the mark. Without them, no one can check the success again.
Why this matters
Without a chain of records, the reason, size and scope of the change cannot be verified. The result may come from a selected example, a changed model or a different question.
Do not confuse
- A screenshot alone is not a complete evidence package.
- A measured change and a causal claim are different.
- A change on one platform cannot be generalised to every system.
What should you do?
- Give every claim a unique identifier.
- Lock the baseline before making the change.
- Record the change at file and version level.
- Preserve test conditions and raw outputs.
- Write the assessment rules in advance.
- Explain uncertainty between the result and its possible cause.
How do you audit it?
- Can the claim be reproduced?
- Are the baseline and later measurement comparable?
- Were the question set and model conditions the same?
- Does the raw output produce the score in the report?
- Does the claim exceed the boundary supported by the evidence?
Limit
A complete record does not by itself prove causation, but it provides the foundation needed for independent examination of the claim.
Remember in one sentence
Testable success preserves the measurement path as well as the result.
Sources for this record
- S01Aggarwal et al., *GEO: Generative Engine Optimization*Research
- S35W3C, *PROV-O: The PROV Ontology*Standard

