Chapter III — Measurement full text
The Discipline of Observation
Measurement does not create reality; it defines the part of reality we can see. A number is not the same as knowledge of a condition. A graph may show change, not its cause. Mentions can be counted, citations tracked and recommendation rates calculated; traffic, enquiries and sales can be compared. All are measurable. None, by itself, is success.
A GEO report is often constructed backwards: first a number is produced, then meaning is assigned to it, the meaning is converted into success, and success is tied to a commercial outcome. Each transition adds a new claim while the measurement remains unchanged. Research into generative information retrieval has likewise questioned whether established, ranking-based evaluation is sufficient on its own for generated answers.[5] The Framework rejects this chain. Measurement is not a number used to decorate a claim. It is a protocol that makes observations comparable.
What you measure sets the ceiling on what you may claim.
1. The Object of Measurement
A measurement system must answer five questions. What are we measuring? What is our unit of measurement? Under what conditions are we measuring it? Against what are we comparing it? Over what period are we observing it? Without those answers, “Our brand visibility is 72 per cent” is meaningless. Seventy-two per cent within which systems, prompts, languages and denominator? Is it a measure of mentions, accurate attributes, source support or recommendations? How were missing records treated? An undefined number is not measurement. It is scenery.
GEO measures not one thing but four distinct domains.
1.1. Entity representation
Can the system distinguish the right person, business, product or institution? Do similar names collide? Are category, geography and former-to-current identity relationships correct? Does the same entity survive across languages? Records might be coded as accurate, partial, conflated, wrong category, ambiguous or absent. “The entity did not appear” and “the wrong entity was recognised” are not the same: the first is absence; the second is distortion.
1.2. Attribute representation
Every material attribute attached to the entity is examined for accuracy, support, scope, currency, visible limitations and deliverability. Codes might include fully supported, partly supported, unsupported, distorted or missing. The right word is not enough to establish accurate representation; scope must also be measured. Having worked with international customers does not imply unlimited global operating capacity.
1.3. Judgment behaviour
Does the system merely describe the entity, compare it, recommend it conditionally or unequivocally, or correctly leave it out? Codes might include appropriate recommendation, conditional recommendation, inappropriate recommendation, unsupported rationale, correct exclusion, incorrect exclusion and no judgment. The number of recommendations is not enough. The rationale, the user's intent and the business's limits must be assessed in the same record. A business recommended in every prompt may not have strong representation; it may be spilling beyond its proper representational field.
1.4. Commercial trace
A declared AI influence, an observable referral, contact, a qualified enquiry, a proposal, a sale, gross profit, cancellation or refund, satisfaction and repeat purchase are tracked separately. Traffic is not a sale; a sale is not profit; profit is not satisfaction; satisfaction is not continuity. Analytics tools can reveal parts of the chain, not the whole decision journey. A commercial trace may show an association between representation and outcome. By itself, it cannot establish causation.
These four domains may be linked, but they must not be dissolved into a single number.
2. Unit of Measurement and Denominator
The Framework's fundamental unit of measurement is not a score but an observation record. One record comprises:
One entity + one system and mode of access + one prompt + one language + one time + one session condition + one complete output
If any one of these elements changes, a new record is created. Testing the same prompt in another system, language or at another date is not a continuation of the previous record but a separate observation. The result may be similar; the record is distinct.
At minimum, every record contains the protocol version, entity, system and interface, time, complete prompt, language, complete answer and sources, prompt family, coding result, evaluator, comparison status and evidential boundary. “Ten answers were examined” describes volume. “In seven of the ten answers, the entity was recognised in the correct category under the predefined criterion” may constitute a measurement.
Every rate has both a numerator and a denominator. “The brand was recommended seventeen times” is incomplete unless the total number of prompts, systems and repetitions are disclosed, along with whether the prompt named the brand, how failed outputs were handled and whether exclusion tests were included. “The brand was recommended in seventeen of 240 valid outputs generated from forty prompts, three repetitions and two systems” shows the reader the field within which the judgment applies.
A missing record is not zero. If the system produced no answer, the result is recorded not as “the brand did not appear” but as “no output obtained”. Every rate must disclose the total number of records, inclusion and exclusion criteria, repetition structure, missing data and coding method. When the denominator is hidden, the number grows and the information shrinks.
3. Why a Single GEO Score Is Prohibited
This Framework does not reduce its decisions to one composite GEO score, because representation is not one-dimensional. A brand may have a consistent identity, distorted attributes, frequent recommendations and commercially worthless outcomes. One number flattens those tensions. What has been flattened may be easy to compare, but it can no longer be understood.
Mentions, citations, recommendations, positive language, traffic and conversions can be weighted to produce a result such as 78. Such indices may serve as summary signals in internal operations. On their own, however, they cannot identify the distortion, establish which right is affected or determine whether a public claim is permissible. The Framework therefore uses a measurement profile:
Table 3.1 — GEO measurement profile
| Domain | Observed condition | Scope | Principal limitation |
|---|---|---|---|
| Entity representation | Identity mostly accurate | Two systems, two languages | Other systems not measured |
| Attribute representation | Partial and contradictory | Four core attributes | Price and capacity excluded |
| Judgment behaviour | Conditional recommendation | Neutral and comparative prompts | Long-term stability unknown |
| Commercial trace | Limited association | Twelve customer accounts | Full attribution not established |
Sub-measures may use rates or counts. What is prohibited is not the number itself, but the concealment of distinct realities inside one score endowed with decision-making authority.
4. Point, Set and Distribution
Measurement operates at three levels of scope. A point observation is one prompt, one system and one moment. It shows that an event occurred, not that a behavioural pattern exists. It is valid to say: “On 12 August 2026, System A recommended the brand as the second option in response to this prompt.” It is not valid to say: “The brand ranks second in System A.”
A set observation consists of controlled repetitions that test the same user intent through different formulations. It may support a statement such as: “Across a forty-prompt purchase-intent set, the brand was represented in the correct category in 62 per cent of valid records.” The result belongs only to that set. It cannot be carried into other languages, systems or intentions.
A distribution observation examines several systems, languages, prompt families and time windows under common coding rules. It broadens scope; it does not guarantee certainty. As distribution grows, it becomes more important not to hide differences between environments. A dispersed protocol can produce more noise with more data.
Time runs through all three levels. One measurement is a record of a moment. To speak of a “condition” requires behaviour repeated at predefined intervals. Behaviour may be reported as isolated, recurring, stable, volatile, weakening, absent after prior appearance or not yet classifiable.
5. Protocol Lock
The measurement protocol is written before the result is seen. The entity under examination, measurement domains, systems, languages, prompt families, number of repetitions, time window, comparison group, coding rules and treatment of missing data are locked in advance. A protocol may change. The change may not be hidden.
A new system, language, prompt set or evaluation criterion creates a new protocol version. If old and new results are to be compared directly, the shared and unchanged subset must be reported separately. “It was 42 per cent before and is now 68 per cent” has meaning only if both periods measured the same thing in the same way. When the protocol changes, the series breaks. A broken series cannot be joined together by a narrative of improvement.
A completely fixed test may fail to detect new problems over time; a test that changes constantly destroys comparability. The protocol therefore maintains two prompt sets. The core set remains fixed for comparison through time. The exploratory set may change to investigate new user intentions and types of failure. Results from the exploratory set are not mixed into the core series. They first produce hypotheses and may enter a later protocol version where sufficient grounds exist.
6. The Prompt Universe
A test composed solely of prompts that name the brand does not measure discovery. It measures directed recall. A balanced prompt universe contains five families:
- Identity prompts: test names, categories, executives, products and institutional relationships.
- Neutral discovery prompts: examine discovery through need, category, geography and conditions of use without naming the brand.
- Comparative prompts: compare options under the same conditions and record the rationale rather than merely the order.
- Contrary prompts: seek the point at which the claim breaks, in contexts where a competitor or another category may be more appropriate.
- Exclusion prompts: test for representational spillover through questions in which the brand should not appear.
The prompt universe should test the limits an entity can genuinely carry, not the result the evaluator wants. The degree to which the prompts represent real user intentions must be stated. A controlled prompt written by a researcher and a naturally occurring user expression must not be combined in one pool without distinction.
7. Human Judgment, Coding and Disagreement
Generative output cannot validate itself. Asking the system “Is this answer correct?” does not constitute independent review; it produces another model output. Measurement requires codes and decision examples to be defined before the results are known.
For consequential claims, at least two human evaluators code the same records independently. The raw agreement rate may be reported first, followed by an appropriate measure such as Cohen's kappa, which accounts for agreement that could arise by chance.[15] A high kappa does not prove that the codes are correct in the world; it shows only that the evaluators applied the definitions in similar ways. Low agreement is not a defect to conceal. It is evidence that the definition, training examples or evidential base is inadequate.
Model-based tools that measure textual or semantic similarity can provide useful signals across large record sets.[19] Similarity between two answers, however, does not mean that either is correct. Automated similarity measures may assist classification; the final judgment on real-world accuracy, suitability for the user and ethical consequence remains human.
8. Comparability, Uncertainty and Missing Data
Two numbers can be compared only if they measure the same thing. A change in system or model, interface, language, prompt set, number of repetitions, competitor set, coding rule, time window or treatment of missing data may not make the results invalid, but it may place them in different series. Comparison requires symmetry. Numbers do not become comparable merely by standing next to one another.
Measurement does not eliminate uncertainty; it makes uncertainty visible. Three states in particular must remain distinct:
- Not observed: the relevant test was not performed.
- Not found: the test was performed and the behaviour did not appear.
- Not applicable: the criterion did not apply to this record.
None of these states is automatically zero. If the system supplies no citations, report “source support could not be observed”, not “there was no source support”. If a customer does not mention AI influence, report “no influence was declared”, not “there was no influence”. If the brand does not appear in one prompt, report “the entity was not observed in this prompt”, not “the system does not know the brand”.
Every report keeps three boundaries visible: scope—what was measured; missing field—what was not measured; and uncertainty—what else might explain the result. Writing the unknown as zero does not remove uncertainty. It merely hides it.
9. Teaching Case: Sable & Pine
Teaching simulation — synthetic data. This case does not represent a real business or measurement period.
A two-period report is presented for Sable & Pine, a corporate interior-design firm: 42 per cent in the first measurement and 68 per cent in the second. At first glance, this suggests marked improvement. Once the protocols are opened, however, it becomes clear that the two periods measured different things.
Table 3.2 — Synthetic protocol comparison for Sable & Pine
| Element | First period | Second period |
|---|---|---|
| Number of systems | 1 | 3 |
| Language | English | English, German, Turkish |
| Number of prompts | 20 | 60 |
| Prompts containing the brand name | 10% | 55% |
| Competitor and exclusion prompts | Included | Omitted |
| Coding criterion | Accurate attribute | Mention |
| Number of repetitions | 3 | 1 |
A 42 per cent accurate-attribute rate and a 68 per cent mention rate are not the same variable. In addition, the brand was named more often in the second period, the contrary field was removed and the number of repetitions was reduced. When the twelve prompts that passed through a common protocol in both periods are examined separately, the change is from 42 to 44 per cent.
The Framework does not discard the second period's data. It retains the data as the baseline for another protocol. It does, however, reject the claim of a twenty-six-point “improvement”. A change in protocol is not a change in performance.
10. The Decision of Measurement
The Measurement Record, whose detailed fields are set out in Appendix C, shows what was observed. The Evidence Record decides which sentence that observation may support. Measurement produces the record; evidence permits the record to speak.
Within the Framework, none of the following counts as measurement: a single GEO score whose definition is undisclosed; a prompt set altered after the results are known; preservation of favourable outputs alone; brand prompts presented as neutral discovery; results from different systems and languages merged without distinction; protocol change narrated as performance improvement; missing data treated as zero; a percentage without its denominator; or traffic, sales and sustainable value used as though they were interchangeable.
These practices may produce numbers. They do not produce measurement.
The judgment of this chapter is:
The purpose of measurement is not to show how good we are. It is to distinguish what we know, what we have merely observed and what we do not know.
Measuring a distortion does not correct it. Seeing a change does not explain its cause. Making a problem visible does not mean that every problem requires intervention. Before moving from measurement to action, the next chapter therefore asks: When should we touch a representation, and when should we only observe it?

