Direct answer
GEO success should not be reduced to one mention or one score. Access, retrieval, mentions, citations, recommendations, representation accuracy, evidential support, stability over time and human impact should be measured separately. Results should identify the system, version, language, country, date and fixed question set.
In plain language
You cannot call a student successful merely because they came to class. You examine attendance, correct answers, mistakes and improvement separately. Visibility is only one indicator in GEO; accurate, evidence-backed representation is the outcome that matters.
Why this matters
A single score hides the layer that failed. Mentions can rise while citation support or representation accuracy falls. Reporting only favourable numbers turns an audit into a marketing report.
Do not confuse
- Frequency and accuracy are different.
- Citation rate and citation support are different.
- Recommendation rate and suitable recommendation are different.
- One test result and stability over time are different.
What should you do?
- Fix the definition of success and the question set before measurement.
- Take a baseline measurement.
- Record the model, version, language, country, date and prompt.
- Keep the raw answers and sources.
- Score accuracy, completeness, currency and evidential support separately.
- Test again under the same conditions after a change.
How do you audit it?
- Were the metric formulae written in advance?
- Are the denominator and sample visible?
- Does the report include failed and adverse results?
- Was one model or one query generalised?
- Can the reported result be recalculated from the raw data?
Limit
NOMOS measures are proposed audit metrics; they must not be presented as an independently adopted official industry standard. Measurement does not reveal the full internal operation of a third-party system.
Remember in one sentence
GEO success is less about how often you appear than how accurately and evidentially you are represented.

