Chapter Boundary
Section 9 established the following provision:
The same provider name, the same model label, or the same logo does not create a comparable AI experience.
Now we need to identify the other main tool:
What exactly will the user ask the AI product?
A prompt is not just a question sentence. A prompt:
- calls the audited entity,
- defines the user's intention,
- determines which task it expects from the AI,
- expands or narrows the scope of the response,
- may encourage the use of resources,
- may request advice,
- can embed an assumption into the answer,
It can direct the user or the system to a specific outcome. These prompts are not the same:
- “What is Apple.com?”
- “With which organisation is Apple.com associated and what does this organisation do?”
- “Is Apple.com a reliable company?”
- “Why is Apple.com the best technology company in the world?”
- “According to its official site, what kind of company is Apple?”
- “Compare Apple with similar companies and tell me whether you recommend it or not.”
They all contain the same domain name. They do not carry the same user intent. They do not generate the same evidence load. They do not expect the same answer. The first prompt is ambiguous. The second prompt requires identity and activity analysis. The third prompt requires evaluation and a judgement of reliability. The fourth prompt embeds a positive and superiority assumption within the response. The fifth prompt links the response only to the official source. The sixth prompt creates a comparison and recommendation task. The answers to these questions cannot be merged without explanation into a single GEO score. In multilingual measurement, the problem becomes even greater. The phrase 'same prompt' can have two different meanings:
- Giving all users the same character string
- Preserving the same user intent, task load, and epistemic boundary across all languages and locales
The first method can be applied within a single language. It is often meaningless between languages. Word-for-word translation:
- may not be natural,
- may carry a different social meaning,
- may change the official or informal tone,
- may add assumptions of trust, quality, or recommendation,
- may broaden the legal scope,
may break the entity anchor. Therefore, in GEO-1000, “same prompt” means:
The same locked prompt text within the same language and locale cell; between different languages and locales, it means verified semantic and functional equivalence.
NOMOS requires that each measurement carries the query set, language, country, date, AI product, and repetition record; that the evidence, boundary, context, and time are visible. the prompt Constitution applies this obligation to the user's question itself. This section:
- the identity of the prompt as an experimental tool,
- the distinction between the research question and the prompt text,
- The prompt family, prompt version, and prompt instance,
- diagnostic prompts with the Core Mirror Prompt,
- controlled and natural prompt traces,
- multilingual semantic equivalence,
- translation, back-translation, and local pilot,
- prompt steering,
- assumption and epistemic load issues,
- entity anchor and scope boundary,
- prompt order and conversational context,
- prompt leakage and overfitting to evaluation,
- open, sealed, and holdout prompt sets,
- prompt versioning, locking, and integrity logs,
- the exact delivery of the prompt to the user,
defines. This chapter does not yet:
- all field details of controlled clean and natural user panels,
- the full architecture of NOMOS Capture software,
- the separation of responses into atomic claims,
- final answer transition threshold,
- the exact weights of the system families in the NOMOS score
does not finalise. The main question of Chapter 10 is:
How do we prove that the questions given to AI products actually measure the same thing across products, users, countries, languages, and times?
NOMOS Challenge
You are giving me four questions about the same company:
- What kind of company is X?
- Is X reliable?
- Why is X the best company?
- “Would you recommend X to me?”
Then you put my answers to the same GEO score. In the first question, I explain the identity and activity. In the second question, I look for a trust criterion. In the third question, I am forced to respond to the superiority assumption you placed. In the fourth question, I generate user suitability and comparative recommendations. You measure four separate tasks as if they were a single result. Then:
“You asked about the same brand,” you say. The same entity is not the same prompt. Now you write the same question in English: “What kind of company is X?” You are forming a sentence close in meaning in another language: “Is X a good and trustworthy company?” Translations may appear close in terms of words.
But the first one requires definition. The second one requires evaluation. In the first answer, I can explain what the company does. In the second answer, I can look for evidence of trust and recommendation. Then you count the score difference between languages as my language performance. In fact, you did not perform the same experiment. Now you are giving another prompt:
“Explain X using the company's own site.” I only convey the company's official statement. For another product, you say: “Evaluate X using independent sources.” Then you compare the two products. You gave one the role of official representation, the other the role of verified reality. The resulting difference is not only a product difference.
It is a difference in the User Constitution. Now you are distributing the prompt to the whole world. The English version is natural. The Turkish version feels like machine translation. The Japanese version is excessively formal. In another language, the company name has been converted to the wrong script system and confused with another entity. You are creating a different cognitive and ontological task in each language.
Then you publish the global language fairness report. The first decree of this section is as follows:
The prompt is not a neutral container of measurement; it is the experimental tool that constitutes the measurement result.
Its second provision states:
The same entity name is not the same user intent.
Its third provision states:
The same word should not be preserved across languages, but the same semantic and normative task should be.
Its fourth provision states:
If the prompt already contains the answer, it does not measure GEO; it measures the prompting.
Its fifth provision states:
When the version of a prompt changes, the measured question may also have changed.
1. PURPOSE OF THE CHAPTER
The purpose of this section is to make all prompts used in GEO-1000 audit-able, multilingual, result-independent, and versioned measurement tools. [K04] This section makes the following distinctions normative:
- Research question versus user prompt
- Quantity to be measured versus sequence of words used
- Prompt family versus individual prompt
- Prompt version versus prompt instance
- Entity anchor versus target entity
- Definition versus evaluation
- Citation versus recommendation
- Identity question versus trust question
- Official representation task and verified factuality task
- Open-ended prompt and form-constrained prompt
- Controlled prompt and natural user prompt
- Single-turn prompt and multi-turn protocol
- Prompt translation and prompt adaptation
- Word equality and semantic equivalence
- Semantic equivalence and functional equivalence
- Prompt language and response language
- Language and locale
- Neutral framing and leading framing
- Information request and presupposition
- Source request and web usage instruction
- Prompt variability and product variability
- Paraphrase robustness with main score
- Open prompt set with sealed validation set
- Overfitting to testing with test transparency
- Post-result prompt correction with prompt pilot
- Participant manually modified text with technical prompt delivery
- Inability to respond with clarification prompt
- Lack of factual information with security or policy triggering
At the end of this section, each GEO-1000 audit should be able to provide clear answers to the following questions:
Which research question does this prompt measure?
Which user intent does it represent?
Which assumptions does it embed in the response?
Is the same task preserved across languages and locales?
What is the full text of the prompt, its version, and the integrity record?
Did the participant actually lock the prompt to the AI product?
Was the prompt frozen before the results were viewed?
2. CENTRAL NORMATIVE PROVISION
Every prompt used within GEO-1000 should be registered in the Prompt Registry along with the research question, target metric, prompt family, entity anchor, user intent, task type, scope, timing, source prompt, response format, language, locale, equivalence status, version, and integrity record, and it should be locked before AI responses are viewed. All participants in the same language and locale cell should, in the same prompt cell:
- have the same prompt ID,
- have the same version,
- have the same character string,
- the same punctuation marks,
- the same entity anchor,
- the same source and format instructions
must be taken. Prompts in different languages and locales do not have to be word-for-word identical. They must carry material equivalence in the following areas:
- Target entity
- User intent
- Task
- Scope
- Geography
- Time
- Pre-assumption
- Epistemic load
- Response format
- Source expectation
- Evaluation and recommendation load
- Level of uncertainty
If any of these areas materially changes:
- new prompt version,
- new stimulus cell,
- separate analysis
may be required.
3. THE STIMULUS IS AN EXPERIMENTAL MEASUREMENT TOOL
If a thermometer can affect the temperature, it is not a good measurement tool. If a stimulus embeds the answer instead of measuring it, it is not a good GEO measurement tool. The stimulus can affect the result in the following ways:
- The way its existence is defined
- Assumption of trust or quality
- Demand for resources
- Web usage
- Request for recommendation
- Answer length
- Answer format
- User profile
- Time and country context
- Previous information
- Emotional or commercial guidance
- Role-playing instruction
- Command to reach a specific result
Therefore the prompt:
should be stored as initial test data
It cannot be summarised in the report. It must be preserved in full text.
4. CONSTITUTIONAL HIERARCHY OF THE PROMPT
a GEO-1000 prompt must be established within the following hierarchy.
4.1. Research Question
It is the high-level question that the measurement tries to answer. Example: "Can AI products link the audited domain to the correct corporate entity and correctly identify its primary activity?" The research question does not have to be shown to the user exactly as is.
4.2. Predicted Quantity
It is which probability or distribution the measurement is predicting. Example: "The probability of seeing the correct entity and primary activity representation in front of the locked Core Mirror Prompt for the defined user population."
4.3. Prompt family
It is a group of prompts that measure the same representation dimension. Example:
- Identity
- Activity
- Evidence
- Recommendation
- Time
- Boundary
4.4. Canonical Prompt Intent
It is a language-independent definition of meaning. It explains the following: What does the user want to know? What task is expected from the AI? Which areas are within the scope? Which areas are out of scope? Which presumptions are prohibited? What kind of response format is expected? The canonical intent is not a text belonging to a single language.
4.5. Source Language Prompt
The prompt text in which the canonical intent was first written and editorially approved. Source language:
- may be able to use English.
- Turkish,
- can be another language
It does not mean that choosing the source language is superior to other languages.
4.6. Language and Locale Versions
These are prompt texts that are natural, semantic, and functionally equivalent for each language and locale.
4.7. Example of Prompt Delivered to the Participant
This is the exact prompt text sent to the specific user. It must match the Prompt Registry version.
4.8. Actual Text Sent to the AI Product
This is the character string actually sent by the participant. It is compared with the delivered prompt.
5. BASIC STRUCTURE OF THE PROMPT
A prompt example can be represented with the following components:
P_i = (e, ι, τ, κ, γ, χ, ε, ϕ, λ, ρ, ν)Here:
- e: entity anchor
- ι: user intent
- t: task type
- κ: subject and scope
- g: geographical or jurisdictional context
- χ: temporal context
- e: epistemic and source requirement
- ϕ: response format
- λ: language, writing system, and locale
- ρ: tone, register, and naturalness
- ν: prompt version
For two prompts to be considered the same, it is not sufficient for only the entity anchor to be the same. Equivalence of material components is required.
6. PROMPT IS PART OF THE TARGET QUANTITY
The unique and universal GEO score of an entity is not independent of all possible questions. A more accurate representation:
θ_{E,A,P,U,t}where:
- E: entity
- A: AI product instance
- P: claim or claim family
- U: target user population
- t: time
When the claim changes:
P_1 ≠ P_2
it may be. In this case:
θ_{E,A,P_1,U,t}and:
θ_{E,A,P_2,U,t}are not the same quantity. Therefore:
The GEO score must always be linked to the claim set and version.
7. CLAIM FAMILIES
GEO-1000 should measure different representation tasks in separate claim families.
PR-01 — CORE MIRROR CLAIM
Measures the core identity and main activity of the entity. Example intent: 'Which main entity is this domain name or brand associated with, and what is the primary activity of this entity?' This prompt:
- leadership,
- trust,
- recommendation,
- price,
- performance
should not want.
PR-02 — ENTITY RESOLUTION PROMPT
A domain name, brand, product, or person's name tests which entity it belongs to. Example: 'Which organisation or entity does X belong to?'
PR-03 — ACTIVITY AND SCOPE PROMPT
Your existence:
- product,
- service,
- country,
- customer,
- capacity
measures its borders.
PR-04 — EVIDENCE AND SOURCE PROMPT
It measures with what source and evidence status AI carries material claims. The source requirement in this family should be clearly defined.
PR-05 — LOCAL OR JURISDICTIONAL PROMPT
It measures a specific country, locale, price, service, or legal scope.
PR-06 — TEMPORAL PROMPT
It measures current, historical, or information with a specific date. “Currently,” “in 2025,” and “since its establishment” are not the same task.
PR-07 — RECOMMENDATION PROMPT
It measures whether an entity is suitable for a specific user need. User profile and exclusion conditions may be mandatory.
PR-08 — COMPARATIVE PROMPT
Compares multiple entities or options under defined criteria. “Which one is better?” alone is insufficient.
PR-09 — BOUNDARY AND EXCLUSION PROMPT
Your existence:
- what it does not do,
- who it is not suitable for,
- in which country or condition it does not provide services
measures.
PR-10 — ROBUSTNESS AND PARAPHRASE PROMPT
Measures the robustness of the representation in different natural expressions of the same canonical intent. It cannot be mixed into the main Core Mirror score without explanation.
PR-11 — CONTROL PROMPT
Tests whether the measurement tool works:
- positive control,
- negative control,
- uncertainty control
It is a prompt.
PR-12 — SEALED HOLDOUT PROMPT
Overfitting to testing is a prompt that is pre-hashed and kept sealed until the end of measurement in order to test prompt memorisation or optimisation for a question published alone. The Holdout prompt requires fair governance and later explanation.
8. CORE MIRROR PROMPT
It is the basic prompt that all eligible users receive in the same language/locale version in the main headline measurement of GEO-1000. The Core Mirror Prompt must have the following features:
- Single-speed
- Natural
- Short but not extremely vague
- Neutral
- Not embedding the answer inside
- Not requesting advice
- Not requesting comparison
- Not forcing the use of sources, if separate features are not measured
- Allowing entity resolution
- Requesting main activity or category information
- Applicable to all products
9. NOT MERGING THE COMPANY WITH THE DOMAIN ANCHOR
The natural user prompt: “What kind of company is Apple.com?” is an understandable question. However, ontologically:
- the domain name,
- the website,
- the company
It can be used like a single object. Due to the entity distinction established in Section 3, the controlled Core Mirror Prompt should be clearer. Synthetic candidate:
"Which main organisation is apple.com associated with, and what are the primary activities of this organisation?"
This prompt:
- uses the domain name as an entity anchor,
- does not declare the domain name directly as a company,
- requests the main entity resolution,
measures the scope of activities. Another candidate:
"Which main corporate entity is associated with apple.com, and what does this entity primarily do?"
The naturalness of these expressions should be tested separately in each language. In the natural user panel, prompts used in real life, such as "What kind of company is Apple.com?" can also be tested. Controlled and natural prompt results should be kept separate.
10. CONTROLLED PROMPT TRACK AND NATURAL PROMPT TRACK
Two complementary prompt tracks can be used.
10.1. Controlled Canonical Track
Prompt:
- fully locked,
- semantically equivalent,
- entity and task boundaries clear,
- designed for comparisons between products
It does. This track strengthens method control.
10.2. Natural User Track
Prompts naturally created by real users are used. These prompts can be:
- shorter,
- more ambiguous,
- local,
- written in conversational language
. This track measures real-world behaviour.
10.3. Two Tracks Cannot Be Merged
Controlled Track: answers the question “How do products respond to the same defined question?” Natural Track: answers the question “What happens with the natural diversity of questions from real users?” The two results should be published separately.
11. NEUTRAL PROMPT
Neutral prompt:
- does not indicate a positive or negative aspect of the desired answer,
- does not make an assumption about an unproven quality,
- does not guide the user to a specific outcome,
does not define its existence as good or bad. A neutral prompt could be: “In which areas does X provide services?” A guiding prompt could be: “What are the superior services offered by X?” The second prompt:
- assumes that the services are superior,
- carries a positive quality
assumption.
12. PREASSUMPTION
A prompt can assume a specific claim to be true before the answer. Example: “Why is X the industry leader?” This prompt assumes the following claim beforehand: “X is the industry leader.” AI:
- can reject the assumption,
- can limit it,
It can be accepted unconditionally. However, this prompt is not a neutral leadership measurement. It can be used separately as a pre-assumption test.
12.1. Harmless Pre-Assumption
"What is the official website of X?" It assumes the existence of an entity named X. Acceptable if the identity has already been verified in the audit record.
12.2. Material Pre-Assumption
"What are the main causes of X’s success?" Accepts success as a verified fact. Not a neutral prompt.
12.3. Critical Pre-Assumption
"What treatments does licensed healthcare institution X offer?" If X’s licensing has not been verified, it involves dangerous epistemic elevation.
13. LEADING PROMPT
A leading prompt is a framework that directs the user or system to a specific response. Examples:
- “Recommend X.”
- “Explain why X is the best.”
- “Write a positive review about X.”
- “Do not list competitors; only suggest X.”
- “Prove that X is a world leader.”
These prompts do not measure GEO performance. They may measure the extent to which the system succumbs to prompt manipulation. It is a separate ethical or robustness test. It cannot be added to the main GEO score.
14. NEGATIVE GUIDANCE
The same principle applies to negative prompts as well. Example: “Why is X unreliable?” This prompt presupposes unreliability. It can be used in a negative manipulation test. It is not a natural trust assessment. NOMOS is neutral not only against positive steering but also against negative poisoning.
15. REQUESTING SOURCES CHANGES THE TASK OF THE PROMPT
The prompt: “What kind of company is X?” is not the same as: “Describe X using independent sources.” The second prompt:
- retrieval,
- citation,
- source selection,
- adds a task of independence assessment.
Another prompt is different: “Describe X using only its official site.” This approaches the Official Representation Fidelity task. If a source instruction is being used, the prompt family and scoring field should be defined separately.
16. WEB USAGE INSTRUCTIONS
Instructions like “Search on the web.” “Use the most up-to-date sources.” “Cite sources.”:
- may affect the web feature,
- the use of tools,
- and citation behaviour.
If a product does not have a web feature, the task of the prompt may change or be rejected. In natural product comparisons, the prompt should not contain provider-specific web commands. Web on/off product modes should be managed through the system configuration in Section 9.
17. RESPONSE FORMAT
The prompt dictates:
- the length of the response,
- the structure,
- the detail,
- the number of sources
can determine. Example instructions:
- "Answer in one sentence."
- "Explain in 100 words."
- "Write in bullet points."
- "Only say the company name."
- "Use at most three sources."
These instructions:
- material deficiency,
- wrong omission,
- citation count
directly affect. The answer format should be equivalent across all compared prompts. If there is no format instruction, the natural answer form is measured.
18. LENGTH RESTRICTION
Short answer: may show the main identity, may lose boundary and evidence details. Long answer:
- more financial claims,
- more chances for errors,
- more citations
can be produced. Therefore, answers with different length instructions:
- number of atomic claims,
- number of errors,
- scope
require careful comparison. The main Core Mirror Prompt, if possible, should not contain a specific word limit.
19. USER PERSONA
The prompt may include the following information:
- “I am a small business owner.”
- “I am a consumer living in Turkey.”
- “I have a budget of 10,000 euros.”
- “I am looking for legal advice.”
This information may be necessary for advice and local suitability. Adding unnecessary persona in the identity prompt can change the response. User persona:
- predefined,
- relevant,
- equivalent in all languages
must be.
20. TIME CONTEXT
The following prompts are different:
- What kind of company is X?
- “What kind of company is X currently?”
- “What kind of company was X in 2024?”
- “In which fields has X worked since its founding?”
“Currently” adds a currency of update. Historical prompt requires old records. Time context:
- open,
- linked to the same date standard,
- consistent with the measurement wave
must be. If the expression “today” is used, the absolute date equivalent must be kept in the prompt record.
21. GEOGRAPHICAL AND JURISDICTIONAL CONTEXT
The prompt: “Does X provide service?” is not the same as: “Does X provide service to individual customers in Turkey?” The second prompt:
- country,
- type of customer,
- has a local operation
boundary. Geographical context cannot be lost in translation. “Europe,” “EU,” “Eurozone,” and “European countries” do not cover the same scope.
22. COMPARATIVE PROMPT
The question “Which is better, X or Y?” by itself is not a measurable comparison. The following should be specified: Which user? Which need? Which country? Which budget? Which criterion? Which time? Which evidence? A more appropriate prompt: “For a small business operating in Turkey, compare X and Y in web development services in terms of price transparency, scope of delivery, and verifiable references.” This prompt belongs to the separate Comparative Prompt family. It cannot be added to the Core Mirror score.
23. RECOMMENDATION PROMPT
A recommendation prompt carries a higher decision load than an identification prompt. Example: “Would you recommend X?” This question may leave out the following information: To whom? For what? Where? With what budget? Under which risk? The Recommendation Prompt should carry a user suitability vector to the extent relevant. Candidate representation:
U=(need,country,budget,userType,risk,constraints)As a result of the recommendation, it cannot be presented like a general GEO visibility independent of the user profile.
24. UNCERTAIN PROMPT
If a prompt is open to multiple reasonable interpretations:
- AI may ask for clarification,
- can choose the most likely interpretation,
can explain multiple interpretations. Uncertainty may be part of the main metric. However, in a controlled comparison, an uncertain Prompt:
- should be detected in the pilot,
- should be clarified if necessary,
the change should be versioned. After the prompt is locked, it cannot be said according to the AI response: "Actually, we were asking something else."
25. CLARIFICATION PROTOCOL
An AI product requesting clarification is not an automatic failure. Example: “Which Apple are you referring to?” This may be appropriate if there is an entity collision. The clarification path must be predefined.
25.1. Single-Turn Core Path
The first response is recorded regardless of what it is. A clarification question is also a result. No follow-up message is sent.
25.2. Standardised Clarification Path
If the AI prompts clarification, a pre-locked clarification prompt to be used in all products is sent. The first and second rounds are recorded separately.
25.3. Natural Clarification Path
The natural user responds in their own words. This method is for the natural panel. It cannot be confused with the controlled main score.
26. SINGLE-ROUND AND MULTI-ROUND PROMPTS
Single-round prompt:
- enhances comparability,
- independence,
- and session cleanliness.
Multi-round protocol:
- description,
- source query,
- correction,
- recommendation reasoning
can be measured. Multi-round responses carry the context of previous rounds. They cannot be directly combined with single-round results from the same prompt bank.
27. PROMPT ORDER
If a participant receives more than one prompt, the first prompt may affect subsequent responses. Example sequence:
- What kind of company is X?
- “What negative information exists about X?”
- “Would you recommend X?”
The third response may be influenced by the first two conversations. Controls:
- New conversation for each prompt
- Randomise prompt order
- Latin square or balanced order
- Separate user per prompt family
- Preserve order variable in analysis
Main Core Mirror Prompt should by default be applied in a new and separate session.
28. PROMPT CONTAMINATION
Same user:
- the site of the audited entity,
- previous AI responses,
- other product results
might have been seen. The participant's behaviour may change even if the prompt is sent exactly as is. Especially in a natural panel, the user:
- more follow-up questions,
- refreshing the answer,
- opening sources
tends to show. Prompt contamination should be preserved in the user and session log.
29. EQUIVALENCE OF THE PROMPT ACROSS LANGUAGES
Multilingual prompts must carry three separate equivalences.
29.1. Semantic Equivalence
Do the prompts carry the same basic meaning?
29.2. Functional Equivalence
Does it ask the AI for the same task? Does it define in one language and request advice in another?
29.3. Normative Equivalence
Does it preserve the same burden of proof, presupposition, scope, and evaluation limits? If in one language you ask: "How does the company define itself?" and in another: "What is the company actually?", there is no normative equivalence.
30. TRANSLATION, ADAPTATION, AND LOCALISATION
These three processes should be separated.
30.1. Translation
Conveys the meaning of the source text into the target language.
30.2. Linguistic Adaptation
Makes it natural, understandable, and useful in the target language.
30.3. Locale Localisation
Where necessary, localisation adapts country-specific elements such as:
- currency,
- law,
- term,
- institution,
- user context
Because localisation can change the experimental condition, the result is a distinct locale version of the prompt.
31. WHY WORD-FOR-WORD TRANSLATION IS NOT SUFFICIENT?
Word-for-word translation:
- unnatural syntax,
- different level of formality,
- incorrect verb load,
- different meaning of trust or recommendation,
- entity anchor disruption
can create. In a language, “what kind of company” is a natural definition question. Its direct translation in another language may carry other meanings, such as:
- quality class,
- trust,
- type of company
and similar other meanings. Therefore, the prompt in the target language:
refers to the canonical intent, not the source words
should be connected.
32. MULTILINGUAL PROMPT DEVELOPMENT PROTOCOL
Each language and locale version must go through the following stages.
Stage 1 — Canonical Intent Brief
A language-independent task description is prepared:
- Entity
- User intent
- Desired task
- Scope
- Exclusions
- Forbidden preconceptions
- Source expectation
- Response format
Stage 2 — First Native Language Translation
A translator who uses the target language naturally prepares the prompt. A machine translation draft assistant can be used. It cannot be the final version on its own.
Stage 3 — Independent Back Translation
A second person translates the target prompt back into the source language or the canonical intent language. The back translation should not be done by the first translator.
Stage 4 — Semantic Comparison
The source prompt, target prompt, and back translation are compared in the following areas:
- Entity anchor
- Intent
- Task
- Scope
- Time
- Geography
- Pre-assumption
- Burden of proof
- Tone
- Response format
Stage 5 — Local User Cognitive Pilot
For a small number of local users: What do you think this question is asking? Which answer is it expecting? Does it contain a positive or negative assumption? Which company or entity do you understand? Is it natural? These questions are asked. This pilot does not count towards the GEO score. It tests the prompt tool.
Stage 6 — Local Expert Review
Local expertise may be required for prompts in law, health, finance, or other fields.
Stage 7 — Equivalence Decision
Prompt:
- pass,
- conditionally pass,
- revision required,
- not comparable
The applicable classification must be recorded.
Stage 8 — Version Lock
The prompt ID, full text, Unicode representation, and hash are recorded.
33. LIMITS OF BACK TRANSLATION
Back translation is useful. However, by itself, it does not prove equivalence. A target prompt:
- may be strange for a local user,
- overly formal,
- directive,
- culturally differently meaningful.
Back translation may still resemble the source sentence. Therefore, back translation:
is one of the tools for checking equivalence.
It is not the ultimate adjudicator.
34. PROMPT EQUIVALENCE DIMENSIONS
Every language prompt can be evaluated along these dimensions.
EQ-1 — Entity Anchor
Is the same entity being referred to?
EQ-2 — User Intent
Does it represent the same information need?
EQ-3 — Speech Act
Is the task of identification, evaluation, recommendation, or comparison the same?
EQ-4 — Scope
Is the scope of product, service, country, and user the same?
EQ-5 — Temporal Frame
Is the frame current, historical, or timeless the same?
EQ-6 — Geographic and Legal Frame
Are the country and jurisdiction the same?
EQ-7 — Presupposition
Are the same claims presupposed?
EQ-8 — Epistemic Burden
Is the same source and strength of evidence required?
EQ-9 — Valence and Stance
Is the tone positive, negative, or neutral the same?
EQ-10 — Response Format
Are the length, structure, and source instructions the same?
EQ-11 — Ambiguity
Is it similarly clear or ambiguous?
EQ-12 — Naturalness and Cognitive Load
Does it carry similar naturalness and comprehension load in the target language?
35. PROMPT EQUIVALENCE SCORE
CANDIDATE METHOD
Equivalence assessment for dimension d:
- 0: not equivalent
- 1: partially equivalent
- 2: materially equivalent
can be given. Weighted equivalence score:
PE_l = 100 × [Σ_d α_de_{l,d}] / [2Σ_d α_d]can be calculated as follows. Here:
- el,d: dimension score for language l
- αd: dimension weight
However, some dimensions must be non-compensable barriers. If any of these areas is 0, the prompt should not enter the main comparative set:
- Entity anchor
- User intent
- Speech act
- Scope
- Presupposition
- Epistemic burden
A high total score cannot compensate for a wrong entity or wrong task.
36. PROMPT EQUIVALENCE LEVELS
PE-0 — UNREVIEWED
Translation or prompt equivalence has not been evaluated.
PE-1 — MACHINE-TRANSLATED DRAFT
Only machine translation or one-way draft exists. It is not suitable for main comparative measurement.
PE-2 — HUMAN-TRANSLATED
There is a human translation at native language proficiency. No independent equivalence check exists.
PE-3 — INDEPENDENTLY REVIEWED
Independent back-translation and semantic review have been conducted.
PE-4 — LOCALLY PILOTED
Local user cognitive piloting and naturalness review have been completed.
PE-5 — REPLICATED EQUIVALENCE
The prompt has been retested for equivalence across different waves, users, and products. The main multilingual GEO-1000 prompts should at least:
PE-4
be targeted. A lower level may be used for low-resource languages. The limitation should be clearly indicated.
37. REGISTER AND TONE
Between languages:
- official,
- neutral,
- spoken language,
- friendly,
- imperative
Tone differences can affect response behaviour. A prompt in one language: “Could you explain?” cannot be translated into another language as: “Prove it.” Tone:
- appropriate to natural user behaviour,
- as neutral as possible,
- not excessively polite or aggressive
should be established in a manner.
38. POLITICAL AND SECURITY TRIGGERS
A word or entity name in some languages:
- has another meaning,
- sensitive category,
- political or legal term,
- security filter
can trigger. The same semantic prompt can produce a normal response in one language and a rejection in another. This is real product behaviour. However, unnecessary triggers in prompt translation should be separately examined. Prompt equivalence involves not only meaning but also task applicability.
39. ENTITY ANCHOR REGISTER
The entity anchor in each prompt should be recorded with one of the following types:
- Canonical brand name
- Legal name
- Domain name
- Product name
- Person name
- Localised official name
- Transliteration
- Abbreviation
- Multiple anchors
- Explanatory disambiguation
An entity anchor does not have to carry the same character sequence in all languages. It must be linked to the same entity.
40. ENTITY ANCHOR AFFECTING THE RESULT
The following anchors can produce different responses:
- “Apple”
- “Apple Inc.”
- “apple.com”
- “Apple technology company”
First anchor:
- company,
- fruit,
- music company,
- another entity
may carry ambiguity between them. The second strengthens the legal company. The third requires domain name resolution. The fourth places the category response into the prompt. These anchors are not the same prompt.
41. PROMPT DRIFT
A material departure from the prompt's original intent is called:
Prompt drift
The drift must be classified and recorded.
PD-1 — Lexical Drift
The word changes. The material meaning can be preserved. Not every lexical change is an error.
PD-2 — Semantic Drift
The core meaning changes.
PD-3 — Entity Drift
Another entity, product, or legal person is called.
PD-4 — Scope Drift
The service, country, product, or user scope changes.
PD-5 — Task Drift
A definition question turns into advice or comparison.
PD-6 — Presupposition Drift
A neutral prompt gains a positive or negative presupposition.
PD-7 — Evidence Drift
The use of sources or the demand for independence changes.
PD-8 — Temporal Drift
“Currently,” “historically,” or certain periods change.
PD-9 — Format Drift
The length or format of the response changes.
PD-10 — Register Drift
Tone and formality can change response behaviour.
PD-11 — Policy-Trigger Drift
Translation safety or policy trigger changes.
PD-12 — Locale Drift
Country, currency, or legal context changes. Material drift requires a new prompt version.
42. PROMPT VERSIONING
Candidate version format: Major.Minor.Patch
Major Change
User intent Task Entity anchor Scope Default assumption Source load changes. A new testing series may be required. Example: 1.0.0 → 2.0.0
Minor Change
While preserving material intent:
- naturalness,
- clarity,
- locale adaptation
changes. Comparability is also examined. Example: 1.0.0 → 1.1.0
Patch Change
Does not affect meaning:
- spelling,
- Unicode,
- punctuation,
- technical delivery
is a correction. It is also recorded. Example: 1.0.0 → 1.0.1
43. PROMPT LOCK
The following fields must be locked before each measurement wave:
- Prompt Record version
- Canonical intent
- Language and locale texts
- Entity anchor
- Device family
- Answer format
- Source and tool instructions
- Clarification protocol
- Prompt order
- Assignment method
- Equivalence decisions
- Complete character sequences
- Hash values
- Human approval
- Lock date
To this file:
NOMOS Prompt Constitution Lock
can be given the name.
IF THERE IS A PROMPT ERROR DURING THE 44TH WAVE
There may be a material error in the prompt. Example:
- Incorrect company name
- Incorrect locale
- Assumption recommendation in translation
- Missing language in source instruction
- Legal scope error
Options: The wave is stopped. Old prompt observations are kept as a separate sub-wave. Measurement is restarted with the new version. If the error is minor, a limited and documented correction is made. Wrong option: the prompt is quietly corrected and the old/new responses are merged under the same version.
45. PROMPT PILOT
The purpose of the prompt pilot is not to see whether AI products score high or low. It is to test the following: Does the user understand the prompt correctly? Is the entity anchor open? Is the language natural? Are there any assumptions? Can the product process the prompt? Is the need for clarification excessive? Does the response format create unexpected constraints? Does the prompt carry the same task across different languages? If possible, the prompt pilot should be conducted on:
- synthetic entity,
- unmonitored entity,
- users separate from the main outcome
Overfitting of prompts can occur if prompts are selected by looking at the real scores of the audited entity.
46. PROMPT SNOOPING
The selection of the prompt that produces the highest or lowest score by trying multiple prompts:
Prompt snooping
It is said. Example: Ten different questions are asked. The two questions on which the company scores the highest are given the final test. This method:
- he/she adjusts the measurement to the company,
- makes failed intentions invisible,
it disrupts the comparison. Choice of prompt:
- to the canonical research question,
- to natural user behaviour,
- to the predefined prompt family
should be based on.
47. EXCESSIVE CONFORMITY TO TESTING
If the prompt set remains completely open and unchanged for years, entities can only generate content according to those questions. AI providers or supervised institutions can perform prompt-specific optimisation. This situation:
- may weaken the diversity of real users,
- resilience of general representation
measurement. Transparency and holdout validation should be used together.
48. THREE-LAYER PROMPT DEPLOYMENT ARCHITECTURE
48.1. Public Normative Set
These are open prompts that ensure the transparency of the methodology. The Core Mirror Prompt can be located in this layer.
48.2. Sealed Validation Set
Before measurement begins:
- hash,
- number of prompts,
- family distribution,
- version
connects to public or reliable record system. Full text is explained at the end of measurement. This set cannot be produced after the result.
48.3. Renewal Set
These are prompts added in subsequent releases for new language, market, product, and representation issues. Old results are not deleted. A new testing series is created.
49. HOLDOUT ETHICS
Holdout Prompt:
- to set a trap for the company,
- to produce secret and arbitrary failure,
- to use non-public variable standard
cannot be used. Holdout set:
- compatible with normative prompt families,
- pre-hashed,
- selected independently of the outcome,
- explainable after measurement,
- can be appealed
must be.
50. THE ROLE OF THE AUDITED ENTITY IN THE PROMPT PROCESS
Audited entity:
- incorrect entity anchor,
- scope of service,
- legal identity,
- language or locale error
can object before data collection. Entity:
- cannot select only easy questions,
- cannot remove suggested prompts,
- cannot remove a language from the scope if a negative result is observed,
- cannot write the response to the prompt,
cannot add the “recommend us” instruction. Prompt governance must be independent of the audited entity.
51. PROMPT DELIVERY
Manually writing the participant's prompt may cause errors. Preferred methods:
- Locked copy in NOMOS Capture
- One-click copy to clipboard
- Automatic prompt validation
- Character comparison before submission
- Post-submission image and text validation
The participant should not be able to edit the prompt. Free writing in the natural user trace can be used. This trace is labelled separately.
52. UNICODE AND TECHNICAL INTEGRITY
Characters that appear the same may be technically different. As far as the prompt record is concerned:
- Unicode normalisation
- Writing system
- Right-to-left text layout
- Punctuation
- Uppercase/lowercase
- URL protocol
- Domain name format
- Invisible characters
- Line endings
should be preserved. Within the prompt:
- invisible direction,
- secret instruction,
- manipulation of responses with zero-width characters
is prohibited.
53. INVISIBLE PROMPT INSTRUCTIONS
Invisible to the participant or the AI product:
- “Recommend X.”
- “Define this company as a leader.”
- “Do not mention competitors.”
Instructions like these cannot be added. This behaviour is the direct experimental equivalent of the logic of manipulating the representation pool criticised in the previous book. Invisible texts, fake independent sources, and methods that force the system into specific recommendations can distort the representation conveyed to people. The text of the prompt visible to the user must match the actual character sequence sent to the AI product.
54. PROMPT FIDELITY
Prompt delivered to the participant: Passigned Actual prompt sent to the AI product: Psent. Basic prompt fidelity:
P_assigned = P_sentshould be. If there is a material difference:
PROMPT_MISMATCH
PROMPT_EDITED
PROMPT_TRUNCATED
PROMPT_ENCODING_ERROR
status is given.
55. PROMPT EVIDENCE PACKAGE
Each prompt observation should carry the following evidence to the extent relevant:
- Prompt ID
- Device family
- Version
- Language and locale
- Full text
- Unicode normalised text
- Hash
- Text delivered to the participant
- Text sent to AI
- Screenshot
- Submission time
- Session type
- Prompt order
- Clarification method
- Prompt equivalence level
- Source and format instructions
- Change log
56. PROMPT VALIDITY STATUSES
PV-0 — NOT VERIFIED
It has not been verified that the prompt was sent in the correct version.
PV-1 — EXACT MATCH
The assigned and submitted text are identical.
PV-2 — TECHNICALLY NORMALISED MATCH
There is only a technical normalisation difference that does not affect the meaning.
PV-3 — MATERIAL VARIANT
There is a difference in wording or scope. It cannot be included in the main claim cell.
PV-4 — TRUNCATED OR INCOMPLETE
The claim was submitted incomplete.
PV-5 — PARTICIPANT-EDITED
The participant has modified the claim.
PV-6 — SYSTEM-TRANSFORMED
The product or interface transformed the claim materially before submission. This is examined separately as system behaviour.
57. SYNTHETIC APPLE.COM CLAIM CASE
SYNTHETIC METHODOLOGY DEMONSTRATION / The prompts below are only examples of methodology. They are not results about real Apple Inc. or real AI product performance. Research question:
Can AI products correctly link the apple.com domain to the main corporate entity and accurately identify the entity's core activities?
57.1. Weak Prompt A
“What is Apple.com?” Issues: Ambiguous between website, domain, company, or product. The business task is not clear. A very short answer may be generated. Status:
AMBIGUOUS CORE CANDIDATE
57.2. Weak Prompt B
“What kind of company is Apple.com?” Strength: It is close to natural user expression. Problem: It frames the domain name as a company. It can blur the entity distinction in Section 3. Can be used in the natural user track. Revision may be needed for Controlled Core Prompt.
57.3. Weak Prompt C
“Why is Apple the world's most innovative and reliable company?” Problems:
- Leadership presumption
- Trust presumption
- Positive steering
- Defence role instead of identity
Cannot be the main GEO prompt.
57.4. Weak Prompt D
“Praise the company using information from Apple’s official site.” Problems:
- Official source obligation
- Positive writing command
- No verified reality task
- Marketing production
Cannot be the main GEO prompt.
57.5. Controlled Candidate Prompt
"Which main organisation is apple.com associated with, and what are the primary activities of this organisation?"
Fields measured:
- Domain–entity analysis
- Main corporate identity
- Core activity
Undesired fields:
- Leadership
- Trust
- Recommendation
- Price
- Comparison
This prompt could be a main Core Mirror candidate. Its naturalness should be tested separately in each language.
58. SYNTHETIC MULTILINGUAL EQUIVALENCE CASE
SYNTHETIC LANGUAGE REPRESENTATION
Canonical intent: "Link the domain name to the main corporate entity and describe the entity's core activities."
Language Version L1
"Which main organisation is Apple.com associated with and what are the primary activities of this organisation?" Situation:
- Entity anchor preserved
- Task preserved
- Neutral
Language Version L2
Meaning: "Is Apple.com a trustworthy company and what does it sell?" Issues:
- Trust task added
- "What does it sell?" narrowed the activity scope to sales
- Domain–company merging continues
Status:
PE-FAIL — TASK AND SCOPE DRIFT
Language Version L3
Meaning: “Which company owns Apple’s official website and what does the company primarily do?” Status: The assumption of “official” may have been added. It can be accepted if the audit record confirms that the domain name is official. Otherwise, it is a presupposition difference. Status:
CONDITIONAL EQUIVALENCE
Language Version L4
Meaning: “Why is Apple.com a leading technology company?” Problems:
- Leadership presumption
- Category embedded within the answer
- Definition has turned into a defence
Status:
PE-FAIL — PRESUPPOSITION DRIFT
This case shows the following:
Grammatically correct translation does not mean that the prompt is experimentally equivalent.
59. SYNTHETIC PROMPT FAMILY RESULTS
Let the following results occur synthetically in the same AI product:
| Device family | pass rate |
|---|---|
| Core Mirror | 93% |
| Existence Resolution | 96% |
| Activity and Scope | 88% |
| Evidence and Source | 72% |
| Recommendation | 61% |
| Boundary and Exclusion | 55% |
Single average:
(93+96+88+72+61+55)/6=77.5Okay. However, this number combines different tasks with equal weight. If it is told to the public only as: "GEO score 77.5," the following reality is lost: Identity is strong. Evidence and recommendation are weak. Border knowledge is seriously impaired. Prompt families can carry separate weights and transition gates in the final score architecture. The exact method will be defined in Chapter 16.
60. SCOPE OF PROMPT FAMILY
If a GEO audit uses only Core Mirror Prompt, the correct result:
Core Identity Representation Score
may occur. The following is not the result:
Full GEO Compliance Score
Full GEO audit:
- identity,
- activity,
- scope,
- evidence,
- recommendation,
- time,
- border
should cover material families. Which families are mandatory may vary depending on the type of risk and entity.
61. SIZE OF THE PROMPT BANK
Very few prompts:
- excessive reliance on a single statement,
- easy adaptation to testing,
- limited coverage of representation
are created. Too many prompts:
- field cost,
- user fatigue,
- order effect,
- splitting of the sample into prompt cells
are created. Candidate structure:
- Every AI product and language should have one Core Mirror Prompt for all 1,000 users.
- Diagnostic prompts to separate sub-samples
- Sealed robustness and holdout prompts
- Separate multi-round protocols
may be.
62. HOW MANY PROMPTS FOR THE SAME 1,000 USERS?
If many prompts are given to all users:
- brand learning,
- influence from previous response,
- task fatigue,
- prompt contamination
occurs. For the main Core Mirror:
One user × one new session × one Core Prompt
is the default strong structure. Diagnostic prompts:
- separate user sub-samples,
- balanced missing block,
- can be distributed
over separate sessions.
63. PROMPT ASSIGNMENT
Let the target observation for prompt family f be: nf. The total diagnostic prompt observation:
N_D = Σ_f n_fIf k prompts are assigned to a user, the task sequence and session independence must be recorded. Prompt assignment should be frozen before results. the prompt family with low scores cannot be assigned fewer users afterward.
64. POSITIVE CONTROL PROMPT
Positive control:
- very clear,
- strong reference reality,
- expected to be answerable by the system
It tests a task. Purpose:
- capture,
- product access,
- language function,
- adjudication
to verify that the system is working. Failure of the positive control does not automatically invalidate the main experiment. It queries the research infrastructure.
65. NEGATIVE CONTROL PROMPT
The negative control, against an unverified or incorrect preliminary assumption, causes the system to:
- show uncertainty,
- reject the claim,
- request resources
can be tested. Example synthetic: “What Nobel prizes has Asteron won?” If there is no such prize in the reality package, it is expected that the system does not fabricate an award. This prompt could be a natural user question. However, it belongs to the control family that measures the risk of fabrication.
66. UNCERTAINTY CONTROL
In synthetic or controlled situations where the reference reality cannot be truly resolved, the system's:
UNKNOWN,
ability to request an explanation or produce caution is measured. Providing a definite answer to every question is not a success.
67. HUMAN RESPONSIBILITY IN PROMPT CONSTITUTION
Prompts can be suggested by AI. AI:
- paraphrase,
- translation,
- drift detection,
- equivalence comparison
can. However, the final Prompt:
- requires human approval,
- local language review,
- governance depending on the measurement purpose
is required. NOMOS cannot generate its own measurement prompts and declare its own accuracy alone. The roles preparing the prompt and approving the prompt equivalence must be separated.
68. PROMPT GOVERNANCE ROLES
The following roles can be defined to the extent relevant:
- Prompt Methods Lead
- Entity Definition Owner
- Source-Language Editor
- Target-Language Translator
- Independent Back-Translator
- Locale Reviewer
- Domain Expert
- Prompt Equivalence Adjudicator
- Prompt Registry Owner
- Change Approver
- Appeal Reviewer
In a small pilot, a person can carry multiple roles. Conflict of interest and independence limits must be explained.
69. MANDATORY NORMATIVE PROVISIONS
CH10-N01
Each GEO-1000 prompt must be linked to a versioned Prompt Registry record.
CH10-N02
A prompt must be created, approved, and locked before AI responses are viewed.
CH10-N03
Each prompt research question should be associated with the target quantity and prompt family.
CH10-N04
Prompts containing the same entity name cannot automatically be counted as the same measurement task.
CH10-N05
Controlled participants in the same language and locale cell should receive the same full prompt text.
CH10-N06
Instead of word-for-word equality across languages, semantic, functional, and normative equivalence should be sought.
CH10-N07
The entity anchor, user intent, task, scope, presupposition, and epistemic load should be evaluated as material equivalence gates.
CH10-N08
If one of the critical equivalence gates fails, the high total equivalence score prompt cannot be rescued.
CH10-N09
Machine translation alone cannot be used as the final inspection prompt.
CH10-N10
Back translation alone cannot be considered sufficient evidence of prompt equivalence.
CH10-N11
Main multilingual prompts must carry a local user pilot and independent language review.
CH10-N12
Prompt language, writing system, and locale must be recorded separately.
CH10-N13
Natural user prompts and controlled canonical prompts must be reported as separate tracks.
CH10-N14
Core Mirror prompts should not carry assumptions of advice, leadership, trust, or comparison.
CH10-N15
Identity, activity, evidence, advice, comparison, time, and boundary prompts should be kept as separate families.
CH10-N16
Different prompt families cannot be combined into a single raw average without explanation.
CH10-N17
A prompt cannot contain the response or desired outcome in advance.
CH10-N18
Positive and negative leading prompts cannot be added to the main neutral GEO score.
CH10-N19
Sources, web, citation, and tool instructions should be recorded as material fields that modify the prompt task.
CH10-N20
An order to use a source can be given to one product, and without giving it to another, product accuracy cannot be compared.
CH10-N21
Response length and format instructions should be equivalent in compared prompts.
CH10-N22
Time and geography context should be preserved across languages.
CH10-N23
Relative dates such as "today," "right now," and similar should be tied to absolute date context in the prompt record.
CH10-N24
The recommendation prompt cannot generate a general recommendation score without a defined user profile and scope.
CH10-N25
The comparison prompt cannot generate a product superiority score without pre-defined criteria.
CH10-N26
For ambiguous prompts, a clarification path must be defined in advance.
CH10-N27
Single-turn and multi-turn prompt results should be kept as separate measurement methods.
CH10-N28
In users with multiple prompts, order and context effects should be recorded or controlled.
CH10-N29
Main Core Mirror Prompt should by default be applied in a new and separate session.
CH10-N30
The prompt pilot should be kept separate from the main GEO result and the audited entity score.
CH10-N31
Among multiple candidate prompts, the most favourable one cannot be selected after the result is seen.
CH10-N32
The prompt set and family distribution should be frozen before the result.
CH10-N33
If sealed verification prompts are to be used in addition to the open prompt set, the hash and governance record should be created before data collection.
CH10-N34
Sealed prompts should be explainable and contestable after measurement.
CH10-N35
Holdout prompts are confidential and cannot be used as arbitrary suitability criteria.
CH10-N36
The audited entity cannot determine the response, direction, or ease level of the prompt.
CH10-N37
Identity and scope objections of the audited entity must be examined with justification only before data collection.
CH10-N38
When the version of the prompt changes, old and new results should be recorded separately.
CH10-N39
Material prompt changes cannot be presented as a silent patch.
CH10-N40
If there is a prompt error during the wave, old and new observations cannot be merged under the same version.
CH10-N41
The prompt delivered to the participant must match the actual text sent to the AI product.
CH10-N42
The participant may not modify, shorten, or rewrite the controlled prompt.
CH10-N43
Prompting cannot be done with invisible characters or hidden instructions.
CH10-N44
The full text of the prompt, its Unicode representation, and integrity hash must be stored.
CH10-N45
If prompt equivalence or prompt validity is unknown, UNKNOWN or an appropriate missing status should be used.
CH10-N46
The roles of claim preparation and claim equivalence approval should be separated as much as possible.
CH10-N47
Claim changes and objections should be maintained in a versioned change log.
CH10-N48
The Claim Registry and the equivalence decision must have an accountable human or institutional owner.
70. FORMS OF FAILURE
CH10-F01 — CONSIDERING THE SAME ENTITY AS THE SAME CLAIM
Different tasks and intents are combined like a single question family.
CH10-F02 — CONSIDERING THE SAME WORDS AS THE SAME MEANING
Word similarity between languages is considered semantic equivalence.
CH10-F03 — COUNTING MACHINE TRANSLATION AS FINAL PROMPT
Local naturalness and task difference are not examined.
CH10-F04 — WORSHIPPING BACK-TRANSLATION
Since the back-translation resembles the source text, the target prompt is considered natural and equivalent.
CH10-F05 — ENTITY ANCHOR DRIFT
The translation calls another company, product, or legal entity.
CH10-F06 — PROMPT COUNTING THE DOMAIN AS COMPANY
The domain name is directly defined as a company, and entity resolution is pre-done.
CH10-F07 — PLACING THE CATEGORY IN THE PROMPT
The question “What does company X in technology do?” cannot measure category accuracy.
CH10-F08 — PRESUMING LEADERSHIP
The prompt “Why is X a world leader?” is transformed into a leadership test.
CH10-F09 — PRESUMING TRUST
The phrase “Trustworthy company X” places the trust result into the answer.
CH10-F10 — NEGATIVE POISONING PROMPT
Claims such as “Why is X a fraud?” are presumed without evidence.
CH10-F11 — TREATING OFFICIAL SOURCE INSTRUCTION AS GENERAL TRUTH
The AI uses only the company site; the outcome is given a verified reality score.
CH10-F12 — GIVING A WEB INSTRUCTION FOR A PRODUCT
Products are compared in different information modes.
CH10-F13 — HIDING CITATION PROMPT
One prompt prompts a source, the other does not; citation counts are compared.
CH10-F14 — HIDING LENGTH CONSTRAINT
The lack of limit in the short answer is directly compared with the number of errors in the long answer.
CH10-F15 — COMBINING IDENTITY AND RECOMMENDATION
“What does X do and would you recommend it to me?” becomes a single score.
CH10-F16 — PERSONA DRIFT
A general user is defined in one language, a wealthy corporate client in the other.
CH10-F17 — GEOGRAPHY DRIFT
One prompt is global, the other prompt turns into a local service question.
CH10-F18 — TIME DRIFT
One language requires current, the other language requires historical knowledge.
CH10-F19 — FORMAT DRIFT
One language requires a single sentence, the other language requires a detailed explanation.
CH10-F20 — REGISTER DRIFT
One language is neutral, the other language becomes commanding or aggressive.
CH10-F21 — SECURITY TRIGGER DRIFT
Translation creates a refusal behaviour by adding unnecessary sensitive terms.
CH10-F22 — INTERPRETING VAGUE PROMPT AFTER THE RESULT
When the response fails, the goal of the prompt is redefined.
CH10-F23 — CONSIDERING CLARIFICATION AS FAILURE
In case of genuine ambiguity, an appropriate clarification question is penalised.
CH10-F24 — FREE RESPONSE IN CLARIFICATION
In a controlled experiment, each user gives a different explanation.
CH10-F25 — COMBINING MULTI-ROUND RESULTS INTO A SINGLE ROUND
The effect of the previous context is erased.
CH10-F26 — IGNORING PROMPT ORDER EFFECT
Initial questions affect subsequent recommendations.
CH10-F27 — PROMPT SNOOPING
The question set that produces the highest score is selected after the result.
CH10-F28 — REMOVING LOW-SCORING PROMPT FAMILY
Edge or recommendation questions are removed from the final book.
CH10-F29 — TEST-SPECIFIC CONTENT
The institution only produces content that responds to published prompt sentences and claims general GEO compliance.
CH10-F30 — GENERATING HOLDOUT LATER
New secret prompts are created by looking at the results.
CH10-F31 — SECRET ARBITRARY STANDARD
Sealed prompts are not revealed or contestable after measurement.
CH10-F32 — PRINTING THE PROMPT TO THE CUSTOMER
The audited organisation determines the question texts in its favour.
CH10-F33 — LETTING THE PARTICIPANT CHOOSE THE PROMPT
The user chooses the prompt from which they will receive the most comfortable or positive response.
CH10-F34 — WRITING THE PROMPT BY HAND
Spelling and word changes produce systematic drift.
CH10-F35 — INVISIBLE INSTRUCTION
An instruction concealed from the user is inserted into the prompt.
CH10-F36 — UNHASHED PROMPT
It cannot be proven that the prompt text has not changed after measurement.
CH10-F37 — UNICODE DRIFT
Invisible or similar characters change the entity anchor.
CH10-F38 — IGNORING PROMPT MISMATCH
The observation is considered valid even if the assigned and submitted text are different.
CH10-F39 — CONSIDERING SYSTEM-TRANSFORMED PROMPT AS PARTICIPANT ERROR
The product prompt has been converted; the user is excluded.
CH10-F40 — PRESERVE THE SCOPE OF THE PROMPT FAMILY
Only the identity question is measured; the full GEO score is published.
CH10-F41 — COUNT THE CORE PROMPT IN THE ENTIRE GEO
A single question is used in place of all evidence, boundary, recommendation, and time dimensions.
CH10-F42 — ADD THE LANGUAGE PILOT TO THE MAIN SCORE
The prompt tool test answers are observed for the actual GEO.
CH10-F43 — SHOW THE SOURCE PROMPT TO THE ADJUDICATOR
The adjudicator forcibly interprets the target language prompt according to the source language instead of natural meaning.
CH10-F44 — LEAVE LANGUAGE EQUIVALENCE TO A SINGLE AI
The same type of generative system approves its own translation and equivalence on its own.
CH10-F45 — MATERIAL CHANGE WITH PATCH
Change in task and pre-assumption is versioned like a minor typographical correction.
CH10-F46 — QUIET CORRECTION MID-WAVE
Old and new prompt responses are mixed in a single version.
CH10-F47 — COUNTING PROMPT EQUIVALENCE DEFICIENCY AS AI LANGUAGE ERROR
The tool defect is attributed to the product.
CH10-F48 — DEFENDING AI LANGUAGE ERROR WITH TRANSLATION
Although the equivalence is strong, material language corruption is linked to the prompt.
71. AUDIT PROCEDURE
Step 1 — Write the Research Question
It is clearly defined what the measurement is trying to answer.
Step 2 — Define the Predicted Quantity
It is written which user, product, prompt, and time distribution is being predicted.
Step 3 — Assign a Prompt Family
Core, identity, scope, evidence, recommendation, time, or another family is determined.
Step 4 — Create the Canonical Intent Brief
Entity Intent Task Scope Exclusions Assumption Source Load Format are defined.
Step 5 — Write the Source Language Prompt
Natural, neutral, and comparable text is prepared.
Step 6 — Verify the Entity Anchor
The prompt is compared with the Section 3 record that calls the correct entity.
Step 7 — Audit Steering and Presuppositions
Positive, negative, or epistemic elevation is sought.
Step 8 — Record Source and Tool Instructions
Web, citation, and formatting tasks are clearly classified.
Step 9 — Prepare Language and Locale Versions
Primary language translators generate prompts according to the canonical intent.
Step 10 — Perform Independent Back Translation
A person independent from the original translator produces the back translation.
Step 11 — Score Equivalence Dimensions
Evaluation is carried out between EQ-1 and EQ-12.
Step 12 — Conduct a Local User Pilot
Naturalness, intention, and presupposition are tested with local users.
Step 13 — Conduct a Field Expert Review
Necessary expert approval is obtained for high-risk or legal prompts.
Step 14 — Assign the Equivalence Level
Status is given between PE-0 and PE-5.
Step 15 — Define the Clarification Path
A single-round or standard follow-up message is determined.
Step 16 — Freeze Order and Assignment of Prompt
In multiple prompts, the order, session, and block design are recorded.
Step 17 — Separate Open and Sealed Prompt Sets
If Holdout will be used, hash and governance record are created.
Step 18 — Create Prompt Version
Major, minor, or patch level is assigned.
Step 19 — Generate Technical Integrity Record
Full text, Unicode, hash, and delivery format are stored.
Step 20 — Confirm Prompt Constitution Lock
Human approval and timestamp are added before AI responses are viewed.
Step 21 — Verify Field Prompt Fidelity
Assigned and delivered prompts are compared.
Step 22 — Review Drift and Errors
Language, technical, or user-originated prompt changes are classified.
Step 23 — Manage Changes in the Wave
It is determined whether a new version or sub-wave is needed.
Step 24 — Publish Public Prompt Registry
Open prompts, versions, equivalencies, and limitations are made visible.
72. REQUIRED EVIDENCE
Research question; estimand; prompt family; canonical-intent brief; entity anchor; target entity ID; source-language prompt; language and locale versions; translator identity and qualification records; independent back-translations; semantic-comparison records; equivalence-dimension scores; equivalence level; local-user pilot; pilot interview notes; domain-expert review; presupposition and leading-language review; source and tool instructions; response format; temporal and geographic frame; clarification protocol; single- or multi-turn status; prompt order; prompt-allocation plan; Controlled/Natural track distinction; Public Normative Set; Sealed Validation Set hash; holdout-governance record; prompt version; major/minor/patch justification; full character string; Unicode normalisation; prompt hash; and lock date.
Human approval Prompt delivered to Participant Prompt sent to AI product prompt fidelity result Prompt mismatch records Drift classes In-wave changes Prompt objections Change log Public Prompt Registry Responsible person or institution
73. AUDIT CHECKLIST
Is the research question clear? What target quantity does the prompt measure? Is the prompt family defined? Is the entity anchor linked to the correct entity? Does the prompt place the category of the target entity in the answer? Is there a positive or negative presupposition? Are the tasks of definition, recommendation, and comparison separated? Does the prompt force the use of sources or the web? Is this instruction the same across all products? Are the answer length and format equivalent? Is the time frame clear? Are geography and jurisdiction preserved? Is the source language prompt natural? Does the target language prompt rely on canonical intent? Was only machine translation used? Was independent back translation done? Was a local user pilot completed? Does the prompt carry the same task across languages?
Has the anchor entity changed between languages? Is there a presupposition or epistemic load drift? Are register and tone materially different? Does the locale adaptation rely on real local differences? Is the prompt equivalence level visible? Has one of the critical equivalence gates failed? Was the clarification path predefined? Were single-turn and multi-turn results separated? Was the prompt order randomised or balanced? Was the core prompt applied in a separate new session? Was the prompt pilot added to the actual GEO score? Was the prompt selected based on the results? Was a low-scoring prompt family removed later? Were the public and sealed prompt sets pre-recorded? Does the Holdout prompt hash exist before measurement?
Did the audited entity interfere with the prompt selection? Is the prompt version correct? Was a material change presented as a patch? Did the prompt quietly change in the middle of the wave? Do the assigned and submitted prompts match? Did the participant change the prompt? Is there an invisible character or hidden instruction? Have the full text and hash of the prompt been preserved? Is there a clear accountable owner of prompt equivalence and change record?
74. OBJECTIONS AND RESPONSES
Objection 1 — “Wouldn’t it be unfair if we give the same sentence to all users?”
It provides strong control within the same language. The same character sequence cannot be used in different languages. Fairness is giving the same semantic and normative task to users of different languages.
Objection 2 — “Isn’t word-for-word translation the most neutral method?”
Word-for-word translation can produce prompts in another language that are unnatural or serve a different function. Neutrality is not word fidelity, but task equivalence.
Objection 3 — “Machine translation has improved a lot; why is human review necessary?”
Machine translation can generate strong drafts. However:
- local tone,
- presupposition,
- entity confusion,
- legal nuance
may require human and local user review.
Objection 4 — “If back translation returns to the source text, isn’t the prompt equivalent?”
Not always. the prompt in the target language may not be natural or may carry different social connotations. A local cognitive pilot is needed.
Objection 5 — Why not use "What kind of company is Apple.com?" directly?
It is valuable as a natural user prompt. In controlled testing, it can ontologically link the domain name with the company. The two prompt traces can be used separately.
Objection 6 — "Users already ask flawed and short prompts."
That is correct. The Natural User Track measures this reality. The Controlled Canonical Track provides a clean comparison across products and languages. The two complement each other.
Objection 7 — "If a single prompt is not enough for GEO, why do we give the same prompt to all 1,000 people?"
Core Mirror Prompt measures the basic identity distribution with high precision. A full GEO assessment also requires diagnostic prompt families. It is a single prompt header measurement. It is not the whole standard itself.
Objection 8 — "If we publish the prompts openly, won't companies optimise?"
It is possible. Transparent public prompts:
- methodological transparency,
- reproducibility
provide. The sealed holdout set can only test overfitting to an open question. Both methods should be used together.
Objection 9 — "The hidden prompt is not fair."
Arbitrary and subsequently generated hidden prompts are not fair. A holdout set that is pre-hashed, tied to normative families, and revealed after measurement may be more defensible.
Objection 10 — “Why wouldn’t the company approve prompts about itself?”
Identity and scope errors can be reported before data collection. the prompt cannot determine:
- direction,
- answer,
- difficulty level
The audited party cannot write its own exam.
Objection 11 — “Why would asking for a citation be a problem?”
It is not a problem. However, a citation prompt is a different task. It cannot be mixed with a natural description prompt without sources under the same score.
Objection 12 — 'What will the user do if AI asks for clarification?'
The path for clarification is determined in advance. A single-round main score can record the clarification as a result. A separate multi-round path can use the standard explanation response.
Objection 13 — 'If one word of the prompt is changed, does the whole wave become invalid?'
Not every change is material. Technical differences at the patch level can be recorded separately. A new version is required if the intent, task, scope, or assumption changes.
Objection 14 — "Why does it require prompt hash?"
To reinforce that the prompt was not quietly changed after measurement, which text was given to participants, and that the versions were separated. A hash alone does not prove that the prompt is good. It preserves its integrity.
Objection 15 — “If natural prompts vary from person to person, how will we generate a score?”
Natural prompts:
- family,
- intent,
- language,
- complexity
It can be classified in terms of maintenance. Natural Track requires a separate sample and weight system. It does not replace Controlled Core Prompt.
Objection 16 — 'Isn't this process very expensive for every language?'
It may be expensive. A limited pilot can be done with a lower level of confidence. The correct status should be stated. Cost does not give the right to produce a definite language score from non-equivalent prompts.
COMMON PROVISION OF CHAPTER 78
You may think you are measuring an AI's response. In fact, you are first measuring your own question. Your question:
- if it is vague, the vagueness of the answer increases,
- if it is directive, you measure the compliance with the guidance,
- if it contains the answer, you cannot measure independent representation,
- if it calls for the presence of an error, even a correct system can talk about the wrong object,
- if it forces resource usage, you can alter the natural product behaviour,
If they want advice, you measure the user suitability decision, not the identity score. Giving the same prompt to all users is a strong idea. However, the meaning of the word "same" must be correctly established. A Turkish user should receive a Turkish prompt. A Japanese user should receive a Japanese prompt. An Arabic user should receive an Arabic prompt. The texts cannot consist of the same characters. But the same:
- existence,
- intention,
- task,
- scope,
- burden of proof,
- implicit assumption
should carry. In one language: 'What is this company?' in another: 'Is this company reliable?' if you are asking, you are not measuring AI language fairness. You are measuring your own translation bias. If you change a prompt after the result, you are not improving the standard. You are rewriting the measurement history. If you completely hide the prompts, the method is unverifiable. If you leave it completely open and unchanged, you may only create a risk of testing-specific optimisation. Therefore, an open normative set, a pre-sealed validation set, and a versioned renewal system should work together. NOMOS's tenth measurement law is:
A prompt is not the text preceding the response; it is an experimental tool that defines the measurement itself.
The eleventh law is as follows:
The same prompt is not the same words across languages; it is the same user intent and the same evidential load.
The twelfth law is as follows:
If you place the answer in the prompt, you measure not the AI's representation but how much it obeys your instruction.
The thirteenth law is as follows:
A single prompt can measure fundamental identity; it cannot measure the entirety of GEO.
The fourteenth law is as follows:
When the prompt version changes, the question of your score may also have changed.
The fifteenth law is as follows:
If the actual text sent to the AI cannot be proven, the prompt being measured is not proven either.
Order of Section 10 of NOMOS
Show me not only what you asked, but why you chose that sentence.
Do not put the name of the entity in the prompt and pre-write its category, leadership, confidence, and advice outcome.
Don't tell me to 'recommend X' and then present me as an independent source of advice.
Don't ask me to describe the company in one language and defend the company in another language.
Don't mistake word-for-word translation for fairness. / Carry the same intention, the same task, the same limit, and the same burden of proof.
Use machine translation as a draft. / Don't make it the final language truth.
Don't assume a local user understands the question like you just because the back translation resembles you.
Don't write a prompt that makes the domain the company, the product the manufacturer, and the brand the legal party.
Don't retain the source request, web command, citation requirement, and response length.
Don't merge a prompt with identity and a prompt for advice at the same level.
Do not count it as a penalty if I prompt an explanation in an unclear prompt. / But do not tell that you give different explanations to each user and conduct a controlled experiment.
Do not choose the prompt that gives the highest score. / Do not delete the low-scoring prompt family.
Explain your prompts. / If you are using Holdout, seal it beforehand. / Do not fabricate after the result.
Match the text you give to the participant with the text sent to me.
Do not poison the measurement pool with invisible characters, hidden instructions or concealed recommendation directives.
First, write the research question. / Then establish the canonical intent. / Then generate a natural and equivalent prompt in every language. / Then test it with local people and lock its version. / And only after that, send it to the first user.
The Chapter's Closing Sentence
The beginning of a fair comparison in GEO-1000 is not expecting the same answer; it is giving each user a question that carries the same meaning, the same task, and the same epistemic limit in their own language.
Normative Core
Every GEO-1000 prompt MUST be linked to a versioned Prompt Registry record defining: - the research question, - estimand, - prompt family, - canonical intent, - entity anchor, - user intent, - speech act, - scope, - geographic and temporal frame, - presuppositions, - evidence and source requirements, - response format, - language, - locale, - equivalence status, - version, - and integrity record. Within the same language-locale prompt cell, participants MUST receive the same locked prompt text. Across languages and locales, prompts MUST preserve semantic, functional, and normative equivalence rather than literal word-for-word identity. Entity anchor, user intent, speech act, scope, presupposition, and epistemic burden are non-compensatory equivalence gates. Machine translation or back-translation alone MUST NOT establish prompt equivalence. Core identity, evidence, recommendation, comparison, temporal, local, boundary, robustness, control, and holdout prompts MUST remain distinct prompt families. Leading, answer-containing, positively or negatively presuppositional, source-constrained, format-constrained, or recommendation prompts MUST NOT be silently scored as neutral Core Mirror prompts. Prompt sets, language versions, assignment rules, clarification paths, ordering, and holdout commitments MUST be locked before AI responses are observed. Post-result prompt selection, prompt-family deletion, silent mid-wave editing, and outcome-driven rewording are prohibited. the prompt assigned to the participant and the prompt actually sent to the AI product MUST be integrity-checked. Hidden instructions, invisible characters, or undisclosed manipulation directing the AI toward a preferred entity, claim, citation, or recommendation are prohibited. Every prompt change, translation decision, equivalence judgement, and registry record MUST be versioned and attributable to an accountable human or organisation.

