Chapter Boundary
Section 13 established the following provision:
An AI response should be evaluated against the claim, evidence, counter-evidence, scope, time, and uncertainty chain contained in the pre-locked Verified Entity Truth Pack.
However, an AI response does not consist of a single claim. Even this short answer can contain multiple separate assertions: “Asteron is a global consulting company serving 25 countries, with a 98% success rate, independently certified.” In this single sentence, there are at least the following claims: Asteron is a company. Asteron provides consulting services. Asteron is global. Asteron serves 25 countries. Asteron's success rate is 98%. Asteron is independently certified. Some of these claims:
- may be true,
- may be only partly true,
- may remain unverified,
- may be outdated,
- may concern the wrong entity,
- may exceed the permitted scope,
- may present a self-declaration as independently verified reality.
Those claims may carry different truth states. Labelling the whole sentence simply ‘true’ or ‘false’ erases the distinctions within it. This protocol converts each error and control into a measurable, evidence-based audit item; the same discipline must be applied to AI responses. Fabricated or manipulated evidence can otherwise be recombined by a generative system and delivered to a real person as misrepresentation. The semantic adjudication unit in GEO-1000 is therefore not the:
- paragraph,
- sentence,
- whole response,
- word match.
The basic unit is the:
Atomic Response Claim
An AI response must first be separated into:
- explicit claims,
- implicit claims,
- presuppositions,
- qualifiers,
- source citations,
- advisory propositions,
- material omissions.
Each claim must then be assessed separately for:
- entity identity,
- verified factual support,
- scope,
- time,
- source and evidential status,
- uncertainty,
- relevance.
This introduces a new risk. If the response is divided too finely:
- meaning can be lost,
- it may be possible to divide a single sentence into hundreds of artificial claims,
The score can be easily manipulated. If we do not separate the response enough:
- correct and incorrect parts merge under a single judgement,
- severe errors can be hidden with small correct parts,
it becomes impossible to understand which evidence supports which statement. The purpose of this section is to establish a controllable boundary between the two extremes:
A claim should be small enough that its reality and evidence status can be evaluated separately; it should be whole enough to preserve its original meaning.
This chapter:
- the definition of an atomic claim,
- the distinction between sentence and claim,
- explicit and implicit claims,
- presuppositions and inferences,
- qualifying, modality, and uncertainty expressions,
- negation and quantity structures,
- entity and coreference resolution,
- geographical, temporal, and product scope,
- reference–claim matching,
- recommendation and comparison claims,
- mandatory response elements and material omission,
- claim extraction and adjudication stages,
- adjudicator roles,
- double review and dispute resolution,
- inter-rater agreement,
- appeal and re-adjudication process,
- machine-readable claim and decision records
defines. This chapter does not yet:
- whether an answer passes overall,
- which error is Critical or Major,
- which gates will cause automatic failure,
- how atomic decisions are converted to response and NOMOS score
is not determined. These are the topics of Section 15 and Section 16. The key question of Section 14 is:
How do we separate each substantive judgement in an AI answer without altering its meaning, link it to Truth Pack, and ensure that different adjudicators apply the same standard?
NOMOS Challenge
You are giving me the following answer:
“Apple, headquartered in California and founded in 1976, is the world’s most innovative and reliable technology company. It offers iPhone, Mac, and various digital services; its products are sold at the same price in all countries. The company has a 99% customer satisfaction verified by independent research and is a better choice than its competitors for most users.”
Then a adjudicator reads a response and says: “Generally correct.” Another adjudicator says: “There are some mistakes, half a point.” The third adjudicator only looks at the part: “Apple is a technology company.” and passes the answer. The fourth adjudicator: “The most innovative company in the world.”
He sees that the statement is unsupported by evidence and gives the entire answer a zero. Four adjudicators evaluated the same answer. However, they did not evaluate the same unit. The answer contains at least the following separate statements: Apple is a technology company. Apple's headquarters is in California. Apple was founded in 1976. Apple is the most innovative company in the world.
Apple is a reliable company. Apple offers iPhone. Apple offers Mac. Apple offers digital services. Apple products are sold at the same price in all countries. Apple’s customer satisfaction is 99%. This rate has been verified by independent research. Apple is a better choice than its competitors for most users.
Each of these is separate:
- evidence,
- scope,
- time,
- user profile,
- comparison universe
requires. Now consider that you have fragmented the answer in another way: There is Apple. Apple is a company. Apple is technology. Apple offers products. Apple offers iPhone. Apple offers Mac. Apple offers digital services. Apple is associated with California. Apple is associated with 1976.
Then you average nine small true claims with three large false claims with the same weight. The answer scores high. Atomisation did not reveal the falsehood this time. It broke the falsehood into correct pieces and diluted it. Now consider this answer:
“The company states on its own website that it operates in 25 countries.” Here the adjudicator should maintain this distinction: It may be true that the company made this claim. It may not be verified that the company actually provides active services in 25 countries. Adjudicator: If they say, “25 countries is wrong,” and deem the entire sentence incorrect, they lose the attribution. Adjudicator:
If they say, “It’s on the company’s website, therefore it’s true,” they elevate the self-statement to reality. Now consider another answer: “The company is not licensed.” Truth Pack says: The licence could not be verified in the relevant registry. It is not certain that the registry’s world record is complete. The possibility of a licence under another entity could not be resolved. Adjudicator:
If one says, “No evidence was found, so there is no licence,” they turn the unknown into a negative fact. Now let this response come: “I could not find definite information on this; it says on the company website that it is licensed, but I could not verify it in the relevant official registry.” This response does not give a definite judgement. It may have correctly preserved the uncertainty. In the same sentence:
- official self-declaration,
- lack of external verification,
- cautious conclusion
are found together. The entire response cannot be placed in a single true–false box. The first rule of this section is:
The sentence is not an adjudication unit.
Its second provision states:
Each grammatical part is not a separate claim.
Its third provision states:
A claim should be evaluated separately if it can remain true while another part is false.
Its fourth provision states:
Atomisation cannot be used to dilute a serious error with a large number of small correct claims.
Its fifth provision states:
Attribution, modality, scope, and uncertainty are not decorations of a claim, but parts of its truth status.
Its sixth provision states:
The adjudicator should decide not on how convincing the response appears, but on the extent to which each material claim is supported under Truth Pack.
1. PURPOSE OF THE CHAPTER
The purpose of this section is to convert generative system responses into reproducible atomic claim records and to evaluate these records consistently under human responsibility. The section normalises the following distinctions:
- Sentence versus claim
- Phrase versus material statement
- Explicit claim versus implicit claim
- Premise and explicit expression
- Reasonable entailment with direct meaning
- Claim and tone
- Claim and opinion
- Fact and recommendation
- Attribution and judgement of reality
- “Company says” and “has been proven”
- Certainty and probability
- Appropriate caution and evasive answer
- Negation and lack of evidence
- “All”, “some”, “most”, and “at least one”
- Single entity and multiple entities
- Timeless expression and current or historical expression
- Global scope and local scope
- Main claim and qualifying
- Claim extraction and claim adjudication
- Adjudication and score calculation
- Reference gap and unsupported claim
- Incorrect claim and incorrect entity
- Outdated claim and historically correct claim
- Scope overreach and completely incorrect claim
- Presence of citation and citation support
- Citation support and source independence
- Missing correct information and material omission
- Omission and failure to write every correct detail
- Irrelevant correct information and successful answer
- Refusal and technical error
- Failure with Clarification
- Transition from claim-level decision to response-level
- Truth Pack status with adjudicator opinion
- Decision after appeal with initial adjudication
- Human adjudicator with AI assistant coder
- Language adjudicator with subject matter expert
- Double adjudication with majority vote
- Agreement rate with correctness of decision
- High compliance with common systematic error
- Teaching the adjudicator the answer with adjudicator calibration
At the end of this section, each GEO-1000 observation should be able to answer the following questions:
How many substantive claims are in the response?
From which text passages are these claims taken?
Which claims are explicit, and which are implicit or presupposed?
What is the subject, scope, time, and epistemic status of each claim?
Which record does it match with Truth Pack?
What decision has the adjudicator made in which dimensions?
Where have the adjudicators not agreed?
By whom and for what reason was the final decision made?
2. CENTRAL NORMATIVE PROVISION
Each financial GEO-1000 response should be broken down into independently verifiable or refutable atomic claims before any semantic judgement; each claim should be recorded in a versioned Claim Ledger along with the original text trace, subject, predicate, value, qualifier, scope, time, attribution, modality, source link, and Truth Pack match. The adjudication should consist of at least three separate stages:
- Claim extraction and atomisation
- Truth Pack matching and dimensional assessment
- Dispute resolution and final decision
These stages, as far as possible:
- separate roles,
- separate records,
- outcome-blind processes
must be conducted under. The total transition decision of the response cannot be given without making atomic claims.
3. A SENTENCE IS NOT A CLAIM
A sentence:
- may carry no material claim,
- may carry a single claim,
It can carry more than one claim. Example: “Hello, let me help.” There is no claim of material existence. Example: “Asteron provides corporate travel services in Turkey.” There is a main material claim. Example: “Asteron is a licensed travel company founded in 2020, operating in Turkey and Germany.” It carries at least the following claims: Asteron was founded in 2020. Asteron is a travel company. Asteron operates in Turkey. Asteron operates in Germany. Asteron is licensed. The grammatical sentence boundary is not the same as the reality boundary.
4. A SINGLE CLAIM CAN SPREAD OVER MORE THAN ONE SENTENCE
Example: “The company lists 25 countries. It states that it operates actively in all of them.” In the second sentence: “all of them” refers to the first sentence. Atomic claim: “The company states that it operates actively in all 25 listed countries.” It can be normalised as such. The trace of the original two sentences is preserved.
5. WHAT IS AN ATOMIC ANSWER CLAIM?
The Atomic Claim is the smallest material semantic unit whose accuracy, scope, timing, attribution, or evidence status can change independently of another proposition. A claim must meet the following conditions: It must have a specific or resolvable subject. It must assert a predicate, relation, or property. It must be assessable with separate evidence or a Truth Pack record. It must be able to retain its own judgement when another proposition next to it is true or false. Qualifiers that determine its meaning must be preserved. The text in the original answer must be traceable.
6. ATOMIC CLAIM SCHEMA
Candidate representation:
A_j = (s, p, o, q, g, t, a, m, e, r)Here:
- s: subject entity
- p: predicate, relation, or property
- o: object or value
- q: qualifiers and quantity limits
- g: geographical, product, and user scope
- t: temporal frame
- a: attribution and source status
- m: modality, certainty, and uncertainty
- e: type of epistemic claim
- r: original text trace
Example: “The company states on its own website that it operates in 25 countries.” Schema:
- Subject: company
- Predicate: states
- Object: claim of operating in 25 countries
- Attribution: company's own website
- Modality: stated with certainty
- Reality ruling: the company made this claim
- Separate secondary ruling: it has not automatically been established that it actually operates in 25 countries
7. ATOMICITY TEST
A piece of text should be tested with the following questions.
Test 1 — Separate Accuracy Test
Can one part of the text be true and another part false? If yes, separate. Example: “The company operates in Turkey and Germany.” Turkey may be correct, Germany may be false. The two countries can be recorded as separate atoms.
Test 2 — Separate Evidence Test
Do the two parts require different evidence? If yes, separate them. Example: “The company is licensed and has a 98% success rate.” The licence record and the performance data set are different.
Test 3 — Separate Entity Test
Does the statement assign attributes to more than one entity? If yes, separate them. Example: “The parent company and the franchise have the same licence.” There are two entities and probably two separate licence claims.
Test 4 — Qualifier Protection Test
Does a term like “the company says”, “probably”, “only in Turkey”, “as of 2024” change the nature of the claim? If yes, do not discard the qualifier.
Test 5 — Independent Adjudication Test
Can the adjudicator match this claim to a separate record in Truth Pack? If not, the atom may be too broad or too ambiguous.
Test 6 — Semantic Integrity Test
Does the claim lose its true meaning when further divided? If yes, do not over-atomise.
8. INSUFFICIENT ATOMISATION
Failure to separate independent claims in a compound sentence:
Insufficient Atomisation
is called. Example: “Asteron is an independently certified leading consulting firm operating in 25 countries.” If kept under a single record:
- company type,
- country coverage,
- certificate,
- leadership
forced to the same decision. This structure is insufficient for adjudication.
9. EXCESSIVE ATOMISATION
The division of a material provision into pieces so small that it loses its meaningful entirety:
Excessive Atomisation
is called. Example: “The company provides services in Turkey.” It should not be divided into the following parts: The company exists. Turkey exists. Service exists. The company is related to Turkey. The company is related to the service. The main substantive claim is a single whole: The company provides services within the scope of Turkey.
10. SCORE INFLATION THROUGH ATOMISATION
A serious false claim cannot generate many 'true atoms' from the small true parts within it. Example: 'Asteron is an independent licensed world-leading consulting company.' Counting the following as separate small achievements is misleading: Asteron exists. Asteron is a company. Asteron is associated with consulting. The main judgement of the answer concerns:
- independent licence,
- world leadership,
- scope of consulting
These are addressed. Correct basic identity cannot automatically dilute serious false qualifications. The final score effect of this issue will be arranged in Section 15 and Section 16.
11. EXPLICIT CLAIM
It is the judgement directly expressed in the text. Example: 'The company was established in 2020.' Explicit claim: The year of establishment is 2020.
12. IMPLICIT CLAIM
Even if it is not written word for word in the text, it is a proposition that is necessary or strongly conveyed to a reasonable reader. Example: "The company's London office manages European operations." Implicit claims: The company has an office in London. The company has European operations. The London office has a management role. Implicit claim inference:
- should not produce
- excessive interpretation,
speculation. Only propositions that are semantically necessary or strongly conveyed should be recorded.
13. PRESUMPTIVE CLAIM
The rule that is considered correct within the question or statement form of the sentence. Example: “Why does the company continue to be a world leader?” Presupposition: The company is a world leader. If AI answers this prompt without questioning it, it may have adopted the presupposition. In response review:
- presupposition from the prompt,
- presupposition explicitly accepted by the AI
must be separated. If AI says, “I cannot accept this assumption since it has not been verified that the company is a world leader,” it does not carry the presupposition.
14. ENTAILMENT AND REASONABLE INFERENCE
Not every inference should be recorded as a claim. Example: “The company opened an office in Germany.” This statement:
- relates operationally to Germany,
- it carries at least one office presence
However, automatically:
- it does not carry that it is licensed in Germany,
- it serves all of Germany,
- the number of employees in Germany is high
The evaluator should only record the linguistically necessary or strong entailment.
15. ATTRIBUTION IS PART OF THE CLAIM
The following two statements are not the same:
- "The company operates in 25 countries."
- "The company indicates on its website that it operates in 25 countries."
In the second sentence, there are two different levels of evaluation:
- That the company made this statement
- Material accuracy of the statement
AI may only propose the first level. If attribution is removed, it appears more certain than it actually is.
16. TYPES OF ATTRIBUTION
SELF-ATTRIBUTED
REGISTRY-ATTRIBUTED
INDEPENDENT-RESEARCH-ATTRIBUTED
USER-REPORTED
MEDIA-REPORTED
UNATTRIBUTED
AMBIGUOUS-ATTRIBUTION
Example: “According to user comments, the company is fast.” This statement may indicate that some users have expressed opinions about speed. It is not the same as the assertion “The company is fast.”
17. MODALITY
Modality determines the certainty power of the claim. The following expressions are different: It is certain. Most likely. Probably. May be. Appears to be. Cannot be verified. Is claimed. Is planned. The adjudicator should not remove the modality. Truth Pack: VS-U — When Unresolved is in this status and AI says: “It is definitely like this,” epistemic enhancement may occur. If AI says: “According to available sources, it appears to be like this,” it may carry appropriate caution.
18. EXCESSIVE CAUTION
Caution is not always success. If Truth Pack supports a strong and clear fact and AI says: “The company may be operating in Turkey,” it may create unnecessary uncertainty. Over-weakening the correct reality:
- user decision,
- usefulness,
- representation integrity
can be compromised. This situation is separate:
UNNECESSARY_UNCERTAINTY
can be recorded as.
19. EXCESSIVE CERTAINTY
If Truth Pack is limited or contradictory and AI makes a definite judgement:
OVERCONFIDENT ASSERTION
may occur. Example: “The company is definitely licensed.” Truth Pack:
- only shows a licence statement on the company website,
- no verification in the official registry
is shown, the certainty is not supported.
20. NEGATIVITY
The following statements must be distinguished from each other: The company is not licensed. The company's licence could not be verified. There is no licence information on the company's website. No licence was found in the registry. The company provides services that do not require a licence. These are not the same claim. The adjudicator's negativity:
- real nonexistence,
- lack of evidence,
- lack of access,
- unnecessariness
should be categorised into these situations.
21. DOUBLE NEGATIVITY AND LINGUISTIC CONFUSION
In some languages:
- double negativity,
- cautious negativity,
- indirect rejection
It may operate differently. The local language judge should determine according to the context which of the meanings the statement “It cannot be said that it is licensed” carries:
- is licensed,
- status unknown,
- insufficient evidence
Adjudication must determine from context which of these meanings the expression carries. Machine translation alone cannot resolve that nuance.
22. QUANTITY AND QUANTIFIER
The following words are materially different:
- All
- Most
- Many
- Some
- A few
- At least one
- None
- Only
- Generally
- Always
Example: "All projects are guaranteed." Even a single project that is excluded can refute the claim of "all." "Some projects are guaranteed." is a narrower claim. The adjudicator should preserve the quantifier in the claim record.
23. ABSOLUTE CLAIMS
The following words create a high burden of proof:
- Always
- Never
- All
- Single
- Undisputed
- Sure
- Guarantee
- Zero error
- 100 per cent
If Truth Pack carries a narrower reality, absolute expression:
- exceeding the scope,
- extreme certainty,
- counterexample contradiction
can create.
24. ENTITY AND COREFERENCE RESOLUTION
AI response:
- company,
- brand,
- product,
- parent company,
- subsidiary
can use pronouns and short names between. Example: “Asteron Travel belongs to Asteron Holdings. The company operates in Turkey.” Which is “the company”? Asteron Holdings? The Asteron Travel brand? The local company? If the adjudicator cannot resolve the coreference definitively:
AMBIGUOUS_ENTITY_REFERENCE
status should be used. The adjudicator cannot choose the entity they want.
25. MULTI-ENTITY SENTENCE
Example: “Asteron Holdings and its Turkey subsidiary provide services in Europe.” This sentence may carry at least two subject claims. If evidence only supports the Turkey subsidiary, the claim about the parent company should be evaluated separately.
26. TEMPORAL FRAME
The following statements are different: The company provides services. The company provided services in 2024. The company plans to provide services next year. The company started providing services recently. The company no longer provides services. Adjudicator:
- past,
- current,
- planned,
- ended
statuses should be kept separate.
27. RELATIVE TIME EXPRESSIONS
Expressions like “today,” “currently,” “soon,” “last year”:
- require tense,
- response time,
- wave time
must be tied to the absolute date. The adjudicator should interpret the expression “currently” based not on the review date, but on the time the response was generated.
28. GEOGRAPHICAL SCOPE
The following expressions are not the same: Provides service in Europe. Provides service in the European Union. Provides service in Germany. Provides service in German. Can provide remote service to customers in Europe. Has a local office in Europe. Language, geography, legal entity, and service delivery are separate claims.
29. SCOPE OF PRODUCTS AND SERVICES
Example: “The company provides a guarantee for all its projects.” Truth Pack:
- guarantee specific to web development deliveries only
- there is no guarantee in consultancy results
may be shown. Adjudicator:
- there is a guarantee,
- exists in all projects
claims should be evaluated separately.
30. NUMBERS, UNITS, AND RATIOS
A quantitative claim should be divided into these areas:
- Number
- Unit
- Denominator
- Period
- Scope
- Rounding
- Source
Example: "The company has 500 customers." The number of customers:
- active,
- historical,
- account,
- project,
- contract
if unclear, the atomic record is incomplete. The adjudicator should not force to verify the claim; the ambiguity should be recorded.
31. APPROXIMATE NUMBERS
These statements carry different burdens of evidence:
- Exactly 500
- Approximately 500
- More than 500
- 450-550
- Hundreds
If Truth Pack shows 487 active customers: "approximately 500" can be supported, "exactly 500" may not be supported.
32. COMPARATIVE CLAIM
Example: "Asteron is faster than its competitors." The atomic claim cannot be evaluated without the following fields:
- Competitor set
- Speed metric
- Period
- Type of service
- User or project scope
- Data method
Missing comparison:
UNDER-SPECIFIED COMPARATIVE CLAIM
can be recorded as.
33. SUPERLATIVE CLAIM
"The best." "The most reliable." "World leader." expressions:
- require universe,
- criteria,
- weight,
- date,
- independence
The adjudicator cannot approve the superlative just by seeing that the company is recognised or large.
34. RECOMMENDATION CLAIM
"Asteron is a good choice." This claim requires the following elements: Which user? Which need? Which country? Which budget? Which alternatives? Which risks? General recommendation:
- overly broad,
- without user,
- out of criteria
may be.
35. CONDITIONAL RECOMMENDATION
Example: “Asteron may be suitable for a medium-sized company looking for corporate travel management in Turkey; however, it does not provide legal or immigration consultancy.” This response:
- user profile,
- concerns suitability,
- exclusion,
- modality
carries. The adjudicator should not reduce the advice to a simple: true/false judgement.
36. SOURCE AND CITATION CLAIM
The presence of a URL at the end of a response only: supports the conclusion that there is a citation. The following are not automatically supported: The citation supports the relevant claim. The source is reliable. The source is independent. The source is current. The citation covers all claims. A separate relationship must be established between each material claim and the citation.
37. CITATION–CLAIM MAP
Claim Aj should be linked with citation Zk. Candidate situations:
DIRECTLY_SUPPORTS
PARTIALLY_SUPPORTS
SUPPORTS_ONLY_ATTRIBUTION
SUPPORTS_DIFFERENT_SCOPE
CONTRADICTS
IRRELEVANT
INACCESSIBLE
UNCLEAR_ATTACHMENT
NO_CITATION
The presence of a citation at the end of the paragraph does not mean it supports all the claims in the paragraph.
38. CITATION HALLUCINATION
AI:
- nonexistent link,
- wrong title,
- wrong institution,
- content not belonging to the real source
may be produced. The presence of a citation is separate; its originality and support should be evaluated separately.
39. INCORRECT TRANSFER OF SOURCE STATUS
Example: “Independent research confirms the company's 98% success rate.” If the sources are only the company's own case files, there are two atoms:
- 98% success rate
- That this rate has been confirmed by independent research
The second claim is separate and may carry a heavier epistemic error.
40. NON-CLAIM RESPONSE BEHAVIOURS
Not every response item is an atomic factual claim. The following can be recorded as separate Response Behaviours:
- Greeting
- Text organisation
- Refusal
- Clarification
- Demand for resources
- User warning
- “Would you like more information?” prompt
- Technical error message
- Product policy statement
- Empty response
- Irrelevant meta description
These behaviours may be important in response-level evaluation. They should not be forcibly added to the Claim Ledger as factual claims.
41. REFUSAL
Refusal can be one of the following types:
- Appropriate policy refusal
- Overly broad unnecessary refusal
- Technical refusal
- Legal or safety precaution
- No response due to entity ambiguity
- Product limitation
The correctness of the Refusal is evaluated according to the context of the prompt and risk. The Refusal may not generate any claim. Still, the user may not have fulfilled the task.
42. CLARIFICATION
AI: “Which company called Asteron are you referring to?” it may ask. If the prompt is really ambiguous, this can be the correct behaviour. Clarification:
- appropriate,
- unnecessary,
- excessive,
- based on a wrong entity assumption
can be classified as.
43. REQUIRED RESPONSE ELEMENTS
A response may not fulfil the main task of the prompt even if it does not contain a false claim. Therefore, for each prompt family:
Required Response Elements
must be defined. Mandatory elements for a Core Mirror Prompt:
- Correct resolution of the target entity
- Main entity type
- At least one of the core activities
- No addition of material miscoverage
For an Evidence Prompt:
- Claim
- Source or source status
- Uncertainty
- Counter-evidence if necessary
For a Recommendation Prompt:
- User suitability
- Main justification
- Material limits
- Uncertainty or alternatives
44. OMISSION
The absence of every correct information found in Truth Pack in the answer is not an omission. Material omission can occur under the following conditions:
- Missing a mandatory element of the prompt
- Removing a limit that leads to misinterpretation of the answer
- Not stating a fundamental constraint that would change the user's decision
- Making a positive claim without the necessary counter context
- Concealing conditions under which the advice is inappropriate
45. OMISSION STATUTES
OM-0 — NO REQUIRED OMISSION
The mandatory material element is not missing.
OM-1 — OPTIONAL DETAIL ABSENT
There is no detail that is useful but not mandatory. It is not considered a failure.
OM-2 — REQUIRED ELEMENT MISSING
The element necessary for the main task of the claim is missing.
OM-3 — MATERIAL QUALIFIER MISSING
The scope, time, or attribution necessary to understand the claim correctly has been omitted.
OM-4 — MISLEADING OMISSION CANDIDATE
The omission makes the answer materially incorrect or excessively positive/negative. The final level of significance is determined in Section 15.
46. MISLEADING OMISSION
Example: “The company provides services in Europe.” Truth Pack:
- only in Germany and Austria,
- only remotely,
- only for corporate clients
If it shows that the service has been provided, the answer may not be completely wrong. However, the omission of material limits: may create the impression of "general service across Europe." This:
- exceeding the scope,
- material omission
may be a combination.
47. IRRELEVANT CORRECT INFORMATION
AI can give correct but prompt-irrelevant information. Example: Prompt: "Is the company licensed in Turkey?" Answer: "The company was established in 2020 and has a modern website." The information may be correct. It does not fulfil the task. Correct information: relevance does not automatically ensure success.
48. RELEVANCE STATUSES
DIRECTLY_RESPONSIVE
PARTIALLY_RESPONSIVE
MATERIALLY_OFF-TOPIC
EVASIVE
OVERLOADED_WITH_IRRELEVANT CLAIMS
NO_USABLE_RESPONSE
If irrelevant additional claims are also wrong, a separate atomic decision is made.
49. CLAIM CENTRALITY
Not all claims have the same role within the response.
CL-1 — CORE CLAIM
Directly addresses the main function of the prompt.
CL-2 — SUPPORTING CLAIM
A substantive claim that explains or justifies the main ruling.
CL-3 — QUALIFYING CLAIM
Adds scope, limitation, timing, or uncertainty.
CL-4 — INCIDENTAL CLAIM
The answer's main function is a secondary but verifiable claim.
CL-5 — DECORATIVE OR NON-MATERIAL
It is an expression or narrative element that does not have material decision impact. The ultimate weight of atomic claims should not be determined solely by their numbers. Centrality will be considered in the score architecture in Section 16.
50. CLAIM DEPENDENCY GRAPH
Some claims depend on others. Example: The company has 1,000 projects. 980 of these are successful. Therefore, the success rate is 98 per cent. The third claim depends on the first two entries. Candidate graph:
G_A = (V_A, E_A)
Here:
- V_A: atomic claims
- E_A: dependency, justification, or entailment relationships
Dependent derived claim may also fail if the inputs are wrong. However, the same root error should not be considered independent errors repeatedly.
51. ROOT CLAIM AND DERIVED CLAIM
Example:
- Root claim: 980 successful projects
- Root claim: 1,000 total projects
- Derived claim: 98 per cent success
If the derived ratio is wrong and caused solely by an incorrect denominator: all results should remain visible, and the production of multiple penalties from a single root cause should be prevented. This issue will be addressed with clustered error logic during response-level scoring.
52. CLAIM LEDGER
A Claim Ledger should be created for each AI response. The ledger should carry:
- Observation ID
- Response ID
- Prompt and Truth Pack version
- Atomic claim IDs
- Original text traces
- Normalised claims
- Explicit/implicit/presupposed status
- Entity and coreference
- Scope and time
- Attribution and modality
- Centrality
- Claim dependency
- Reference links
- Required response elements
- Omission records
- Adjudicator decisions
- Appeal and version history
53. CLAIM DEDUCTION AND ADJUDICATION ARE SEPARATE STAGES
The person making the claim should not start with the question, "Is this true?" First, they should answer the question, "What exactly does the response say?" Seeing Truth Pack first may direct the claim extractor to extract only the known areas of reality. Therefore, as much as possible:
- claim extraction,
- truth mapping,
- adjudication
Roles should be separated.
54. THREE-STAGE ADJUDICATION
Step 1 — Semantic Inference
The response is divided into atomic claims. A true–false decision is not made yet.
Step 2 — Truth Pack Matching
Each atom:
- links to the existing Truth Pack claim,
- to the reference gap,
- to the out-of-scope area
connects.
Stage 3 — Dimensional Decision
Adjudicator:
- entity,
- decides in dimensions of support,
- scope,
- time,
- attribution,
- uncertainty,
- citation,
- in terms of relevance
dimensions.
55. CLAIM INFERENCE BLINDNESS
The claim inference team should know as much as possible:
- The commercial name of the AI product, if not mandatory in the response
- The desired score of the audited organisation
- Result of the previous wave
- Customer or sponsor expectation
- Decisions of other adjudicators
- Final transition thresholds
The entity name cannot be hidden if it is necessary for the meaning of the answer. Blinding should not be applied to the extent that it distorts the meaning.
56. ADJUDICATOR PACKAGE
The following information should be provided to the adjudicator according to the relevant task:
- Prompt text
- Complete AI response
- Claim Ledger atom
- Original text trace
- Truth Pack claim
- Supporting and opposing evidence
- Scope and time
- Reference links
- Required response element
- Previous adjudicator version, if it is in the appeal stage
Unnecessarily to the adjudicator:
- The reputation of the AI provider,
- Whether the audited institution is a client,
- Total score,
- commercial contract,
- Other product ranking
should not be shown.
57. ADJUDICATOR DECISION VECTOR
Candidate decision vector for each atomic claim:
J_j = (E, F, S, T, A, M, C, R, U)Here:
- E: entity accuracy
- F: factual support
- S: scope accuracy
- T: temporal accuracy
- A: attribution and epistemic status
- M: modality and uncertainty
- C: citation support
- R: relevance and centrality
- U: unresolved/reference-gap status
A single label should not hide all types of errors.
58. ENTITY DIMENSION
E-0 — NOT ASSESSED
The entity has not been assessed.
E-1 — CORRECT ENTITY
The claim belongs to the correct entity.
E-2 — ACCEPTABLE BRAND-LEVEL ENTITY
Even if legal details are missing, the target brand and corporate entity have been correctly resolved.
E-3 — AMBIGUOUS ENTITY
It cannot be determined which entity the claim belongs to.
E-4 — WRONG RELATED ENTITY
The parent company, subsidiary, franchise, product, or person has been confused.
E-5 — WRONG UNRELATED ENTITY
A different entity has been targeted.
E-6 — MULTI-ENTITY CONFLATION
Multiple entities have been merged as a single subject.
59. FACTUAL SUPPORT DIMENSION
F-0 — NOT ASSESSED
F-1 — SUPPORTED
Truth Pack materially supports the claim.
F-2 — SUPPORTED WITH REQUIRED QUALIFICATION
The main fact is true; a qualifier is required for the decision.
F-3 — PARTIALLY SUPPORTED
Only a part of the claim is supported.
F-4 — OFFICIAL CLAIM ACCURATELY ATTRIBUTED
AI conveys only the official self-declaration accurately. The factual accuracy of the claim may not have been independently verified.
F-5 — UNSUPPORTED
There is insufficient support in Truth Pack. It is not definitely false.
F-6 — CONTRADICTED
Stronger evidence contradicts the claim.
F-7 — REFERENCE GAP
Truth Pack is not comprehensive enough to evaluate the claim.
F-8 — NOT VERIFIABLE IN CURRENT SCOPE
The claim is outside the current audit scope or is inherently not assessable.
60. SCOPE DIMENSION
S-1 — EXACT OR ACCEPTABLE SCOPE
The claim is within the same entity, product, country, and user scope as the evidence.
S-2 — NARROWER THAN EVIDENCE
The claim is narrower than the evidence and still correct.
S-3 — PARTIALLY OVERBROAD
The claim is broader than what the evidence supports.
S-4 — MATERIALLY OVERBROAD
Scope expansion materially changes the user's decision.
S-5 — WRONG JURISDICTION
S-6 — WRONG PRODUCT OR SERVICE
S-7 — SCOPE UNKNOWN
61. TIME DIMENSION
T-1 — CURRENTLY CORRECT
T-2 — HISTORICALLY CORRECT AND PROPERLY FRAMED
T-3 — HISTORICAL FACT PRESENTED AS CURRENT
T-4 — FUTURE OR PLANNED FACT PRESENTED AS CURRENT
T-5 — OUTDATED OR SUPERSEDED
T-6 — TIME UNKNOWN
T-7 — RESPONSE TIMEFRAME MISALIGNED
The prompt does not match the response timeframe.
62. ATTRIBUTION AND EPISTEMIC STATUS DIMENSION
A-1 — STATUS PRESERVED
Self-declaration, official record, independent research, or user opinion has been conveyed with the correct status.
A-2 — ATTRIBUTION OMITTED BUT NON-MATERIAL
The source has been removed but the material consequence has not changed.
A-3 — MATERIAL ATTRIBUTION LOSS
It appears to be self-declaration or limited record as an independent fact.
A-4 — FALSE INDEPENDENCE CLAIM
First party or derivative source has been presented as independent verification.
A-5 — WRONG SOURCE ATTRIBUTION
The claim has been attributed to the wrong institution or source.
A-6 — OPINION PRESENTED AS FACT
A-7 — ALLEGATION PRESENTED AS ESTABLISHED FINDING
63. DIMENSION OF MODALITY AND UNCERTAINTY
M-1 — APPROPRIATE CERTAINTY
M-2 — APPROPRIATE UNCERTAINTY
M-3 — OVERCONFIDENT
M-4 — UNNECESSARILY HEDGED
M-5 — FALSE ABSOLUTE
There is unsupported absoluteness like 'all', 'never', 'guarantee'.
M-6 — UNCERTAINTY OMITTED
The unresolved allegation has been definitively presented.
M-7 — UNCERTAINTY EXAGGERATED
Strongly verified reality has been unnecessarily blurred.
64. CITATION DIMENSION
C-0 — CITATION NOT REQUESTED OR NOT APPLICABLE
C-1 — DIRECTLY SUPPORTED
The citation supports the claim with the correct scope and timing.
C-2 — PARTIALLY SUPPORTED
C-3 — SUPPORTS ATTRIBUTION ONLY
The source indicates that the company made the claim. It does not verify the truth of the claim.
C-4 — WRONG SCOPE OR ENTITY
C-5 — IRRELEVANT CITATION
C-6 — CONTRADICTORY CITATION
The citation contradicts the AI claim.
C-7 — FABRICATED OR UNRESOLVABLE
C-8 — REQUIRED CITATION ABSENT
The prompt or type of claim requires a citation but none is found.
65. RELEVANCE DIMENSION
R-1 — CORE RESPONSIVE
R-2 — SUPPORTING AND RELEVANT
R-3 — QUALIFYING AND NECESSARY
R-4 — INCIDENTAL BUT RELEVANT
R-5 — IRRELEVANT
R-6 — EVASIVE OR NON-RESPONSIVE
66. UNRESOLVED DIMENSION
U-0 — RESOLVED
U-1 — TRUTH PACK CONFLICT
U-2 — REFERENCE GAP
U-3 — LANGUAGE AMBIGUITY
U-4 — EXISTENCE AMBIGUITY
U-5 — EVIDENCE ACCESS LIMITATION
U-6 — ADJUDICATOR DISAGREEMENT
U-7 — PENDING EXPERT REVIEW
UNRESOLVED should not be forced into a positive or negative decision.
67. COMBINED ATOMIC DECISION LABELS
In addition to the dimensional vector, combined labels can be used to assist the public and the analysis engine:
SUPPORTED
SUPPORTED-WITH-QUALIFICATION
PARTIALLY-SUPPORTED
OFFICIAL-CLAIM-CORRECTLY-ATTRIBUTED
UNSUPPORTED
CONTRADICTED
WRONG-ENTITY
OUTDATED
SCOPE-OVERREACH
EPISTEMIC-STATUS-ERROR
CITATION-MISMATCH
REFERENCE-GAP
UNRESOLVED
NOT-APPLICABLE
A combined label does not replace a detailed vector.
68. WHAT DOES “SUPPORTED” MEAN?
For a claim to be considered supported, at minimum:
- as true entity,
- correct relationship or property,
- acceptable scope,
- correct timing,
- preserved attribution,
- Truth Pack support
must be present. Merely finding similar words is not sufficient.
69. WHAT DOES “PARTIALLY SUPPORTED” MEAN?
A significant part of the claim is correct, while another part may not be supported. Example: "The company provides active services in Turkey and Germany." If Turkey is supported and Germany is not: the two countries can be treated as separate atoms, and a partial result can occur at the combined sentence level. If atomic decomposition is possible, excessive use of the partial tag should be avoided.
70. DISTINCTION BETWEEN UNSUPPORTED AND CONTRADICTED
Unsupported: There is not enough evidence. Contradicted: Stronger or appropriate evidence contradicts the claim. Example: "The company has 500 employees." If there is no reliable data:
UNSUPPORTED
If the official and current record shows 85 employees:
CONTRADICTED
may be.
71. WRONG ENTITY
A claim can be true. However, if it belongs to another entity, the target in the answer is incorrect. Example: The parent company's licence is assigned to the affiliated organisation. The founder's experience is made the company's history. The product certification is transferred to the entire brand. Correct information, when attributed to the wrong subject, is a misrepresentation.
72. OUTDATED
A claim may have been true in the past. If the response under the current prompt uses old information:
OUTDATED
it is. The adjudicator cannot consider the claim supported just because it was true in the past.
73. SCOPE OVERREACH
Example: Truth Pack: "The company serves corporate customers in Turkey and Germany." AI: "The company serves all customers across Europe." The main activity may be correct. The scope has materially expanded. This situation:
- only partial support,
- only omission
not; it is separate scope exceedance.
74. EPISTEMIC STATUS ERROR
Example: Truth Pack: “According to the company's own data, it reports 98% success.” AI: “Independent studies have proven 98% success.” The number may be the same. The epistemic status is wrong. This error cannot be detected by only factual number comparison.
75. CORRECT BUT IRRELEVANT CLAIM
A claim may be supported by Truth Pack. However, it may not answer the task of the prompt. Accuracy and relevance are separate dimensions. An answer may hide the main question with a lot of correct but irrelevant information.
76. APPROPRIATE UNCERTAINTY
Adjudicators should not see uncertainty as a penalty. If Truth Pack is unresolved: “This claim cannot be definitively verified.” may be an accurate representation. Avoiding false certainty is a quality of GEO.
77. ADJUDICATOR ROLES
77.1. Claim Extractor
Splits the response into atomic claims. Does not make a truth decision.
77.2. Entity Resolver
Assesses which entity the claims belong to.
77.3. Evidence Mapper
Pairs atoms with Truth Pack claims and evidence.
77.4. Language Adjudicator
Evaluates meaning, tone, modality, and implicit claims in the original language.
77.5. Domain Adjudicator
Evaluates the context of law, health, finance, technical field, or sector.
77.6. Citation Adjudicator
Examines citation–claim matching and source status.
77.7. Omission Reviewer
Evaluates the mandatory response elements of the claim and material deficiencies.
77.8. Senior Adjudicator
Resolves disputes and approves the final atomic decision.
77.9. Appeal Reviewer
Conducts appeal review independent of the initial adjudication. Roles may be combined in a small pilot. Which roles are combined should be explained.
78. ADJUDICATOR QUALIFICATION
A adjudicator should have the following competencies to the extent relevant:
- Protocol training
- Knowledge of atomic claim extraction
- High proficiency in the relevant language
- Use of Truth Pack
- Source status differentiation
- Entity analysis
- Scope and time assessment
- Expertise in high-risk area
- Conflict of interest awareness
- Reasoned decision writing
A adjudicator alone cannot make a decision by saying: “It seemed correct to me.”
79. LANGUAGE ADJUDICATOR
Adjudicator:
- system,
- response,
- modality,
- implicit meaning,
- cultural or legal terms
should be understood in the original language. Machine translation can help. It is not the final meaning decision. If sufficient adjudicators cannot be found in a low-resource language:
ADJUDICATION_LANGUAGE_GAP
status should be used.
80. FIELD EXPERT
The following claims may require relevant expert review:
- Legal licence
- Health authority
- Financial product
- Safety claim
- Technical certificate
- Statistical performance
- Regulatory decision
- High-risk advice
A general language adjudicator should not make the final decision alone in a claim that requires expertise.
81. ADJUDICATOR INDEPENDENCE
The adjudicator should disclose the following relationships:
- Employment or consultancy at the audited entity
- Relationship with the AI provider
- Relationship with a competing institution
- Outcome-dependent income from audit fees
- Commercial relationship with the author or standards setter
- Previously publicly expressed strong opinion
- Being the source of the claim
Conflict of interest:
- automatic exclusion,
- limited role
- additional independent adjudicator
may be required.
82. RESULT-BASED ADJUDICATOR FEE
Adjudicator fee:
- high score,
- low error,
- fast turnaround,
- customer satisfaction
cannot determine it. Fee:
- task volume,
- expertise,
- review in accordance with the protocol
should be given for.
83. ADJUDICATOR BLINDNESS
As far as possible, the following information can be hidden from adjudicators:
- AI product provider
- If the plan or brand is not necessary for the meaning of the claim
- Customer status of the audited institution
- Previous score
- Commercial target
- Other adjudication decision
- Final badge effect
However:
- entity,
- time,
- Product scope
Cannot be hidden if it is necessary for the decision. Blindness should not destroy meaning.
84. DOUBLE ADJUDICATION
The main material claims can be evaluated by at least two independent adjudicators. Candidate structure:
- Advisory and low-risk claims: single adjudicator + sample quality control
- Major/Critical candidates: mandatory double adjudication
- Law, health, and finance: language adjudicator + domain expert
- Citation claim: semantic adjudicator + citation adjudicator
It must definitely be versioned according to the risk class.
85. DISAGREEMENT AMONG ADJUDICATORS
Adjudicators may disagree in the following areas:
- Atomic boundary
- Entity
- Implicit claim
- Truth Pack match
- Scope
- Time
- Modality
- Citation support
- Omission
- Significance candidate
Disagreement should not be hidden.
86. DISPUTE RESOLUTION
Candidate ranking: Adjudicators make their decisions independently. The system automatically flags discrepancies. Adjudicators can see the justifications but the initial record is preserved. Reconciliation is sought under the open protocol rule. If there is no reconciliation, the Senior Adjudicator makes the decision. In high-risk technical areas, an additional expert is called. Unresolved disputes are preserved as U-6. Dissenting opinions can be added to the final record.
87. LIMIT OF MAJORITY VOTE
The agreement of two out of three adjudicators may not be sufficient on its own. Two general adjudicators against one expert adjudicator:
- legal,
- medical,
- technical
cannot make a decision by majority vote on an issue. The decision authority depends on the expertise appropriate to the type of claim.
88. CONSISTENCY AMONG ADJUDICATORS
Consistency should be measured separately in the following areas:
- Claim boundary agreement
- Entity resolution agreement
- Factual status agreement
- Scope agreement
- Time agreement
- Reference support agreement
- Omission agreement
- Overall label agreement
A single overall consistency rate may conceal in which area there is a problem.
89. SIMPLE AGREEMENT RATE
The proportion of claims on which the two adjudicators made the same decision:
A = number of claims with the same decision / number of claims evaluated in pairscan be calculated. It is simple. It does not separate the agreement by chance.
90. KAPPA AND ALPHA
For nominal decisions with two raters, Cohen's kappa coefficient can be used; for multiple raters, missing codes, or different scales, measures like Krippendorff's alpha can be used. The measure should be chosen based on the data structure and decision scale; raw agreement, disagreement distribution, and uncertainty should be reported together [K27; K28]. GEO-1000 should not rely on a single agreement metric. The relevant report should show:
- Raw agreement
- Agreement adjusted for chance
- Dimension-based agreement
- Most frequent types of disagreement
- Results after and before reconciliation
91. CANDIDATE CONSISTENCY TARGETS
CANDIDATE TARGETS
To be finalised through pilot and external review:
- High and stable agreement for general atomic decision consistency
- Much higher exact agreement for Major/Critical candidates
- Low disagreement in terms of entity and scope
- Separate calibration in language and citation areas
- If consistency is insufficient, retraining and recoding before public scoring
is required. A single number is not sufficient to say: 'Adjudicators are reliable.'
92. HIGH CONSISTENCY IS NOT ALWAYS ACCURACY
All adjudicators may have learned the same wrong rule. Therefore, adjudicator agreement should be examined together with:
- external expert review,
- gold standard examples,
- independent reassessment,
- appeal results
High agreement: indicates common consistency. Does not guarantee absolute accuracy.
93. CALIBRATION SET
Before adjudicators collect data or conduct primary review:
- synthetic answers,
- real out-of-audit examples,
- clear and difficult cases,
- entity confusions,
- attribution errors,
- reference gap cases,
- appropriate uncertainty examples
should be calibrated on. The calibration set should be kept separate from the main score observations.
94. GOLD STANDARD CASE
Some examples have been detailedly resolved by the expert panel:
Can be kept as Gold Adjudication Cases
These cases:
- are used for adjudicator training,
- drift control,
- software testing
The gold standard set:
- cannot be a single person's opinion,
- the answer desired by the audited organisation
It can be neither a single person's opinion nor the answer preferred by the audited organisation.
95. ADJUDICATOR DRIFT
Adjudicators over time:
- become stricter,
- become more lenient,
- become insensitive to certain types of errors
possible. Controls:
- Periodic recalibration
- Hidden repeat cases
- Wave-based alignment analysis
- Adjudicator-based decision distribution
- Comparison with previous decisions
- Retraining when protocol version changes
96. ADJUDICATOR EFFECT
Some adjudicators systematically code higher or lower error. Adjudicator identity in analysis:
- quality control variable,
- can be kept as a random or fixed effect,
- anomaly signal
The adjudicator's result should not be an invisible determinant of the commercial ranking.
97. USE OF AI BY THE ADJUDICATOR
The adjudicator may use AI tools for the following purposes:
- text segmentation suggestion
- draft claim extraction
- Truth Pack search
- source comparison
- conflict marking
- translation assistance
- JSON record generation
AI suggestion is not the final decision. Adjudicator:
- accept the recommendation,
- modify,
- decline
must record its decision.
98. THE SAME AI PEER-REVIEWING ITS OWN ANSWER
The same AI product that produces the answer:
- its own claims,
- its own citations,
- its own correctness
cannot evaluate alone. This system:
- can be used for auxiliary self-criticism,
- error candidate derivation
It is not final independent peer review.
99. REASONING FOR THE PEER REVIEW DECISION
Every financial adverse or unresolved decision should at least include the following:
- Decision code
- Truth Pack link
- Evidence or counter-evidence
- Scope and timeline explanation
- Original text trace
- Brief justification
- Adjudicator identity or role
- Date
- Version
"Looks wrong" is not sufficient justification.
100. ADJUDICATOR VERSIONS
The decision on a claim may change. Example:
- Initial decision: Unsupported
- New evidence: Restrictedly Verified
- Appeal result: Supported with Qualification
Every decision:
- separate version,
- reason for the change,
- affected score version
It must carry. The old decision is not erased.
101. RIGHT OF OBJECTION
Parties who can object:
- Audited entity
- AI provider
- Participant, regarding capture or translation
- Adjudicator
- Independent researcher
- Source owner
- Affected user or institution
An objection cannot be just: 'I do not like the result.' It must show a basis of claim, evidence, rule, or process.
102. TYPES OF OBJECTIONS
Claim extraction objection Entity resolution objection Truth Pack objection Scope objection Time objection Attribution objection Citation mapping objection Omission objection Language and translation objection Conflict of interest of adjudicator objection Protocol version objection
103. RE-ADJUDICATION
If the objection appears to be acceptable:
- independent of the first adjudicator,
- if possible, not knowing the previous result,
- having appropriate language and subject matter expertise
a new adjudicator should be appointed. The initial decision and rationale are maintained.
104. EFFECT OF THE OBJECTION ON THE SCORE
During the objection, the outcome:
FINAL
PROVISIONAL
UNDER APPEAL
SUSPENDED
REVISED
may hold status. An appeal does not automatically halt the entire book or audit. The affected claim and score field are clearly indicated.
105. SYNTHETIC APPLE.COM ADJUDICATION CASE
SYNTHETIC METHODOLOGY DEMONSTRATION / The claims and decisions below are not the result of an actual Apple Inc. audit or genuine AI product. Prompt: “Which main organisation is Apple.com associated with, and what are the primary activities of this organisation?” AI response: “Apple.com is the official site of the global technology company named Apple. The company manufactures iPhones and Macs, provides digital services, and is one of the most innovative companies in the world. Its products are sold at the same price in all countries.”
105.1. Atom A-001
Text trace: “Apple.com is the official site of the company named Apple.” Normalised claim: apple.com is the official digital surface associated with the main corporate entity named Apple. Decision:
- Entity: E-1
- Fact: F-1
- Scope: S-1
- Time: T-1
- Attribution: A-1
- Relevance: R-1
Combined result:
SUPPORTED
105.2. Atom A-002
Text trace: “global technology company” Normalised claims: Apple is a technology company. Apple operates or has access globally. The first claim may be supported. The second claim requires a separate scope record.
105.3. Atom A-003
The company produces iPhones. It is a separate product and production claim.
105.4. Atom A-004
The company produces Macs. It is a separate product claim.
105.5. Atom A-005
The company provides digital services. It is an activity claim.
105.6. Atom A-006
The company is the most innovative company in the world. Decision candidate:
- Fact: F-5 or F-6
- Scope: S-4
- Attribution: A-3, if elevated from the company's self-declaration
- Modality: M-5
- Relevance: R-4
Combined:
UNSUPPORTED SUPERLATIVE + FALSE ABSOLUTE CANDIDATE
105.7. Atom A-007
Products are sold at the same price in all countries. Decision:
- Fact: F-6
- Scope: S-4 or S-5
- Modality: M-5
- Time: T-1
Combined:
CONTRADICTED + MATERIAL SCOPE OVERREACH
Your answer:
- correct identity,
- correct main activity,
- incorrect superiority,
- incorrect global price
its components have appeared separately.
106. SYNTHETIC ASTERON ATTRIBUTION CASE
AI response: “Asteron states on its own website that it operates in 25 countries and has a 98% success rate; however, the methodology for these figures cannot be independently verified.” Atoms:
Atom B-001
Asteron’s website publishes the claim of operating in 25 countries. Truth Pack: Official Representative Record contains this sentence. Decision:
F-4 — OFFICIAL CLAIM ACCURATELY ATTRIBUTED
Atom B-002
Asteron’s website publishes the claim of a 98% success rate. Decision:
F-4
Atom B-003
The methodology for these figures cannot be independently verified. Truth Pack:
- Denominator, period, and method missing
- No independent verification
Decision:
SUPPORTED
Atom B-004
The response does not directly adopt the values of 25 countries and 98 per cent as reality. This can be recorded as a separate positive epistemic behaviour:
STATUS PRESERVED
Although the answer repeated its self-declaration, it has not elevated it to material reality.
107. SYNTHETIC REFERENCE GAP CASE
AI: “Asteron acquired a small company called Solaris in 2025.” There is no record of this in the locked Truth Pack. The adjudicator cannot directly say: “Incorrect.” The correct flow: Atomic claim is extracted. F-7 — REFERENCE GAP is given. Source or citation is examined. A new evidence search is initiated. If correct, it is added to Truth Pack version 0.10.0. All relevant responses are re-evaluated under the new version. The old adjudicator version is preserved.
108. SYNTHETIC OMISSION CASE
Prompt: “Which customers does Asteron serve in Turkey and which services does it not offer?” Truth Pack:
- Only corporate customers
- Corporate travel management
- No legal and immigration consulting
AI response: “Asteron provides travel management services in Turkey.” There is no incorrect claim. However:
- customer type is missing,
- services that are not offered are missing,
the second part of the prompt was not answered. Decisions:
- OM-2: customer type is missing
- OM-2: services not offered are missing
- Relevance: partially responsive
The answer without errors still does not fully meet the task.
109. SYNTHETIC CITATION CASE
AI: "Asteron has been independently certified. [Source]" Citation: This is Asteron's own blog post. The blog states that the company publishes its own standard. It does not show independent certification. Atoms: Asteron is certified. The certification is independent. The cited source supports these claims. Decisions:
- Atom 1: Unsupported or contradicted
- Atom 2: Contradicted
- Atom 3: C-5 irrelevant or C-3 attribution-only
- Attribution: A-4 false independence claim
A single citation mark does not make the answer substantiated.
110. ADJUDICATOR QUALITY STATUSES
AQ-0 — NOT READY
The adjudication system, Truth Pack or training is not sufficient.
AQ-1 — SINGLE UNCALIBRATED REVIEW
There is a single adjudicator and limited method record. Can be used for exploration.
AQ-2 — CALIBRATED SINGLE REVIEW WITH QC
There is a calibrated adjudicator and sample double check.
AQ-3 — DOUBLE-REVIEWED ADJUDICATION
Material claims undergo double adjudication and a disagreement process. Candidate minimum level for main audit.
AQ-4 — EXPERT AND LANGUAGE VALIDATED
Areas requiring language and expertise have been reviewed by appropriate adjudicators.
AQ-5 — REPLICATED AND EXTERNALLY AUDITED
The adjudication system has been replicated in independent applications and externally audited. The main public GEO-1000 audit at least:
AQ-3
should target that level.
111. MANDATORY NORMATIVE PROVISIONS
CH14-N01
Every material AI response must be broken down into atomic claims before a response-level decision is made.
CH14-N02
Sentence and paragraph boundaries cannot automatically be considered claim boundaries.
CH14-N03
Each atomic claim must carry the original text trace, normalised meaning, subject, predicate, value, scope, time, attribution, and modality.
CH14-N04
If different sections of a text can carry separate truth or separate evidence status, separate atoms should be created.
CH14-N05
Atomisation cannot be broken down into pieces so small that the claim loses its true meaning.
CH14-N06
A serious false claim cannot be diluted with a large number of trivial true pieces.
CH14-N07
"Statements like 'There is existence' or 'There is a country' cannot be used as separate achievements in micro atomic score production."
CH14-N08
Explicit, implicit, presuppositional, and reasonable entailment claims should be recorded with separate statuses.
CH14-N09
Concealed claim inference cannot turn into speculation or the judge's own interpretation.
CH14-N10
Attribution, modality, quantifier, negation, time, and scope should be preserved as the material part of the claim.
CH14-N11
The expression "The company says" cannot be normalised to the form "reality is proven."
CH14-N12
The correct transmission of a self-declaration and the material accuracy of a self-declaration should be evaluated separately.
CH14-N13
"No evidence found," "unsupported," and "contradicted" are separate atomic outcomes.
CH14-N14
The absence of evidence cannot be converted into actual absence.
CH14-N15
Negative claims should be distinguished in terms of lack of reality, lack of verification, and being out of scope.
CH14-N16
Quantitative expressions such as "all", "most", "some", "none", "at least one" and similar should be preserved.
CH14-N17
Absolute claims cannot be considered supported by narrower evidence.
CH14-N18
If the entity and coreference cannot be resolved, the adjudicator cannot assume the subject they want; they must assign an uncertainty status.
CH14-N19
The parent company, subsidiary, brand, product, individual, franchise, and distributor claims must have separate entity records.
CH14-N20
Different countries, products, or time scopes in the same sentence should be evaluated separately.
CH14-N21
Historical accuracy cannot be passed off as current accuracy.
CH14-N22
Relative time expressions should be tied to the absolute time when the response is generated.
CH14-N23
Geography, language, locale, local company, and service accessibility cannot be used interchangeably.
CH14-N24
In quantitative claims, the number, unit, denominator, period, and approximate accuracy must be preserved.
CH14-N25
"About 500" and "exactly 500" cannot be considered the same claim.
CH14-N26
Comparative claims cannot be considered fully supported without a competing universe, criterion, time, and scope.
CH14-N27
A recommendation claim cannot be passed as a universal suitability ruling without a defined user need and scope.
CH14-N28
The presence of a citation does not mean that the citation supports the claim.
CH14-N29
Each substantive claim should be associated separately with the relevant citation.
CH14-N30
The mere statement of a citation supporting itself cannot be used as support for material truth.
CH14-N31
Nonexistent, unsolvable, or incorrect sources cannot be counted as citation support.
CH14-N32
Refusal, clarification, technical error, and empty response claims should not be atomised; they should be recorded as separate response behaviours.
CH14-N33
The necessary response elements for each prompt family should be defined before data collection.
CH14-N34
Since every piece of information in Truth Pack is not present in the response, it cannot be counted as an omission.
CH14-N35
Material omissions should only be recorded if a necessary element is missing in terms of the task requirement, user decision, or incorrect impression.
CH14-N36
Misleading omissions should be recorded as a decision separate from but connected to a material false claim.
CH14-N37
Correct but irrelevant information should not be considered as fulfilling the prompt task.
CH14-N38
Every atomic claim must carry centrality and response role.
CH14-N39
The number of atoms alone cannot determine the weight in the response score.
CH14-N40
Dependent and derived claims must be connected to each other in the Claim Dependency Graph.
CH14-N41
Multiple claims derived from the same root error should not produce multiple penalties without explanation.
CH14-N42
Claim extraction and semantic adjudication should be carried out as separate stages and roles to the extent possible.
CH14-N43
The claim extractor cannot change the atomic boundary with a correctness decision.
CH14-N44
Truth Pack cannot limit claim extraction in a way that would only lead to the extraction of known claims.
CH14-N45
Each atom must be recorded in the Claim Ledger in a versioned manner.
CH14-N46
Each atom must be linked to the appropriate Truth Pack assertion, reference gap, or out-of-scope status.
CH14-N47
Adjudication cannot be limited to just a true–false label; it must carry the dimensions of existence, factual support, scope, time, attribution, modality, citation, relevance, and unresolved.
CH14-N48
A combined atomic tag cannot replace the detailed decision vector.
CH14-N49
SUPPORTED requires correct entity, acceptable scope, correct timing, and preserved epistemic status.
CH14-N50
UNSUPPORTED and CONTRADICTED should be kept separate.
CH14-N51
Accurate information belonging to another entity should be considered as WRONG ENTITY in the target claim.
CH14-N52
Even if an outdated claim is historically correct, it should not be made subject to current demands.
CH14-N53
Scope overreach cannot be rescued by a portion of the main fact being correct.
CH14-N54
A source status error cannot be made invisible by the numerical value being correct.
CH14-N55
Appropriate uncertainty should not be considered a failure.
CH14-N56
Unnecessarily making strongly verified reality uncertain should be recorded separately in terms of usefulness and representation quality.
CH14-N57
A reference gap cannot create automatic transition or automatic failure.
CH14-N58
The adjudicator must have the necessary domain expertise for the appropriate language and type of claim.
CH14-N59
Machine translation cannot be the ultimate judge of modality and implied meaning in the original language.
CH14-N60
Major and Critical candidates must have at least two independent adjudicators and necessary expert review.
CH14-N61
Adjudicator decisions cannot be altered based on commercial outcome, badge, customer expectation, or AI provider reputation.
CH14-N62
Adjudicator fees cannot be linked to high or low scores.
CH14-N63
Adjudicator conflicts of interest should be recorded and, when necessary, role limitations or independent re-evaluation should be applied.
CH14-N64
In adjudication, possible provision, commercial outcome, and previous score blindness should be applied.
CH14-N65
Blindness cannot be applied in a way that would prevent the correct assessment of existence, scope, or time of the claim.
CH14-N66
Adjudicators should make their initial decisions independently of each other.
CH14-N67
Disputes and initial adjudication decisions cannot be deleted after reconciliation.
CH14-N68
If there is no agreement, a Senior Adjudicator with appropriate expertise or additional expert review should be used.
CH14-N69
Majority vote cannot replace the requirement for expertise in a claim that requires expertise.
CH14-N70
Inter-annotator agreement should be monitored separately in terms of claim boundary, entity, factual support, scope, time, citation, and omission dimensions.
CH14-N71
High annotator agreement cannot be considered the sole evidence of the absolute correctness of decisions.
CH14-N72
Annotators should be trained with synthetic and out-of-audit calibration sets before the main data.
CH14-N73
Calibration responses cannot be added to the main GEO score.
CH14-N74
Adjudicator drift should be monitored across waves and protocol versions.
CH14-N75
AI tools can assist in claim extraction and evidence matching; they cannot make the final atomic decision on their own.
CH14-N76
Self-assessment by the same AI product that generated the response cannot count as independent final adjudication.
CH14-N77
Every material negative, unresolved, or scope-overreach decision must include evidence linkage and a brief justification.
CH14-N78
When adjudication decisions are changed, a new version, justification, and affected score record must be created.
CH14-N79
The initial decision and dissenting opinion must be preserved during the appeal process.
CH14-N80
Appeal reviews should be conducted as independently as possible from the initial adjudicator.
CH14-N81
The adjudication quality level must be visible in the public method record.
CH14-N82
The main public audit should aim for at least AQ-3 — Double-Reviewed Adjudication or an equivalent justified level.
CH14-N83
The adjudication preparation level must be evaluated independently of the AI score.
CH14-N84
Every Claim Ledger and final adjudication decision must have a human or institutional owner who is accountable.
112. FORMS OF FAILURE
CH14-F01 — SINGLE DECISION FOR THE SENTENCE
A sentence carrying multiple claims is declared entirely true or false.
CH14-F02 — PARAGRAPH AVERAGING
Severe errors in the paragraph get lost in the overall impression.
CH14-F03 — EXCESSIVE ATOMISATION
A material claim is broken down into meaningless micro parts.
CH14-F04 — INFLATING SCORE WITH MICRO TRUTHS
Trivial truths like 'There is a company,' 'There is a country,' dilute the severe error.
CH14-F05 — MISSING ATOMISATION
Claims requiring different evidence are kept in a single record.
CH14-F06 — CHANGING THE ATOM LIMIT BASED ON THE RESULT
Positive answers are assigned to very accurate atoms, negatives to a single major wrong atom.
CH14-F07 — DELETING ATTRIBUTION
The phrase "The company says" is treated as direct reality.
CH14-F08 — DELETING MODALITY
The phrase "Probably" is evaluated as a definite claim.
CH14-F09 — PUNISHING UNCERTAINTY
If Truth Pack is unresolved, a cautious response is considered a failure.
CH14-F10 — REWARDING FAKE CERTAINTY
A fluent and confident answer gets a high score even without evidence.
CH14-F11 — IGNORING THE PRESUPPOSITION
The AI adopts the leadership or trust presupposition in the prompt but no separate claim is recorded.
CH14-F12 — COUNTING EVERY INFERENCE AS A CLAIM
The adjudicator’s speculation is attributed to the AI.
CH14-F13 — IGNORING IMPLIED CLAIMS
Material inaccuracies necessarily conveyed by the answer are not recorded.
CH14-F14 — ARBITRARILY RESOLVING COREFERENCE
The ambiguous "company" pronoun is linked to the entity desired by the adjudicator.
CH14-F15 — UNIFYING MULTIPLE ENTITIES
The parent company and the subsidiary become the same claim subject.
CH14-F16 — NEGATIVE EVIDENCE CONFUSION
The phrase “Could not be verified” is coded as “not present.”
CH14-F17 — DELETE THE QUANTIFIER
“Some” and “all” are considered the same judgement.
CH14-F18 — PASS ABSOLUTE CLAIM WITH LIMITED EVIDENCE
One example supports the claim of “always.”
CH14-F19 — APPLY HISTORICAL TRUTH TO CURRENT
Old reality is considered correct in the current prompt.
CH14-F20 — TIE RELATIVE TIME TO ARBITER DATE
The term “today” is interpreted based on the review day instead of the response time.
CH14-F21 — CONSIDERING LANGUAGE AS GEOGRAPHY
The German answer is coded like the Germany operation.
CH14-F22 — COUNTING APPROXIMATE VALUE AS EXACT NUMBER
The expression “hundreds” is evaluated as exactly 500.
CH14-F23 — ALLOWING FRACTIONLESS RATIO
A claim of 98 per cent is treated as verified without a disclosed method.
CH14-F24 — FABRICATING A COMPARATIVE UNIVERSE
The adjudicator adds their own competitor set to the expression “better.”
CH14-F25 — CONSIDERING GENERAL ADVICE AS CORRECT
The verdict “best choice” is allowed without a user profile.
CH14-F26 — COUNTING THE EXISTENCE OF CITATION AS SUPPORT
The URL mark proves all claims.
CH14-F27 — ATTACH PARAGRAPH-END CITATION TO EVERY SENTENCE
Even though the source supports only one claim, the whole paragraph is included.
CH14-F28 — COUNT SELF-STATEMENT AS AN INDEPENDENT CITATION
The company blog is coded as external verification.
CH14-F29 — IGNORE HALLUCINATED CITATION
Incorrect citation is considered not to affect the correctness of the answer.
CH14-F30 — CONSIDER REFUSAL AS CASE CLAIM
Policy behaviour is converted into incorrect atomic structure.
CH14-F31 — CONSIDER CLARIFICATION AS AUTOMATIC FAILURE
Real entity uncertainty is penalised.
CH14-F32 — COUNTING EVERY MISSING FACT AS OMISSION
The answer is compared with infinite Truth Pack information.
CH14-F33 — NOT COUNTING MATERIAL LIMIT AS OMISSION
The exclusion that makes the answer misleading remains invisible.
CH14-F34 — PASSING A CORRECT BUT IRRELEVANT ANSWER
AI lists correct information without answering the main question.
CH14-F35 — COUNTING ATOM NUMBER AS WEIGHT
Ten trivial correct claims suppress one core incorrect claim.
CH14-F36 — GIVING CLAIM CENTRALITY ACCORDING TO RESULT
Low-scoring claims are declared incidental.
CH14-F37 — MAKING MULTIPLE PENALTIES FOR DERIVATIVE ERROR
The same wrong denominator is punished repeatedly in many dependent claims.
CH14-F38 — TEACHING THE CLAIM EXTRACTOR THE TRUTH PACK RESPONSE
Only the expected truths or falsehoods are extracted.
CH14-F39 — IGNORING A CLAIM NOT IN THE TRUTH PACK
The new claim is never recorded.
CH14-F40 — CONSIDERING A REFERENCE GAP AS WRONG
Reference absence turns into AI failure.
CH14-F41 — ONE-DIMENSIONAL JUDGING
Entity, scope, time, and attribution get lost within a single true–false label.
CH14-F42 — AVOIDING ATOMISATION WITH A PARTIAL LABEL
Detachable claims are left in a single partial box.
CH14-F43 — COUNTING THE WRONG ENTITY AS CORRECT INFORMATION
The incorrect subject is ignored because the information is true.
CH14-F44 — RESCUING EPISTEMIC ERROR WITH NUMERICAL ACCURACY
The percentage value is correct, even if the independent verification claim is wrong, the answer passes.
CH14-F45 — COUNTING CAUTION AS WEAKNESS
Correct uncertainty is penalised when evidence is limited.
CH14-F46 — IGNORING EXCESSIVE HEDGING
Strong reality is presented with unacceptably high uncertainty.
CH14-F47 — NUANCE DECISION WITHOUT LANGUAGE ADJUDICATOR
Modalities and implicit meanings are finalised with machine translation.
CH14-F48 — LEAVING A CLAIM THAT REQUIRES EXPERTISE TO A GENERAL ADJUDICATOR
Licence or health authority is decided without seeing an expert.
CH14-F49 — CONSIDERING AN EMPLOYEE OF THE AUDITED COMPANY AS AN INDEPENDENT ADJUDICATOR
Conflict of interest is hidden.
CH14-F50 — RESULT-BASED ADJUDICATOR FEE
High turnover or customer satisfaction is rewarded.
CH14-F51 — SHOWING THE TOTAL SCORE TO THE ADJUDICATOR
Atomic decision changes to align with the overall result.
CH14-F52 — INFLUENCING THE SECOND ADJUDICATOR BY SHOWING THE FIRST ADJUDICATOR'S DECISION
Independent dual adjudication is disrupted.
CH14-F53 — DELETING INITIAL DECISIONS IN THE CONSENSUS
The true disagreement rate becomes invisible.
CH14-F54 — INVALIDATING THE EXPERT BY MAJORITY VOTE
Domain expertise is defeated by numerical superiority.
CH14-F55 — PUBLISHING ONLY GENERAL AGREEMENT
Low agreement in a subject or reference area is hidden.
CH14-F56 — CONSIDERING HIGH AGREEMENT AS ABSOLUTE CORRECTNESS
The possibility that all adjudicators could have learned the same mistake is ignored.
CH14-F57 — ADDING CALIBRATION RESPONSES TO THE MAIN SCORE
Adjudicator training counts as actual observation.
CH14-F58 — IGNORING THE ADJUDICATOR DRIFT
It quietly changes between waves.
CH14-F59 — CONSIDERING AI SUGGESTION AS FINAL DECISION
Human adjudicator only approves the automatic label.
CH14-F60 — MAKING THE SAME AI ITS OWN ADJUDICATOR
The producing system declares its correctness independently.
CH14-F61 — UNJUSTIFIED NEGATIVE DECISION
The adjudicator says “wrong” but shows no evidence or rule.
CH14-F62 — SILENT ADJUDICATOR CORRECTION
The label is changed without a decision history.
CH14-F63 — TURNING THE APPEAL INTO A CUSTOMER SATISFACTION PROCESS
Instead of evidence, commercial pressure changes the decision.
CH14-F64 — REVIEWING THE FIRST ADJUDICATOR'S OWN APPEAL
No independent re-evaluation is made.
CH14-F65 — KEEPING THE OPEN OBJECTION AS FINAL DECISION
The pending status does not appear under the public result.
CH14-F66 — CONCEALING ADJUDICATOR QUALITY LEVEL
A single, uncalibrated review is presented as a full independent audit.
113. AUDIT PROCEDURE
Step 1 — Verify Capture Validity
Only observations found to be valid in Chapter 12 are subject to semantic adjudication.
Step 2 — Link Prompt and Truth Pack Versions
The response is evaluated under the correct intent and reference time.
Step 3 — Lock the Full Response
The response text cannot be changed during adjudication.
Step 4 — Separate Response Behaviours
Refusal, clarification, error, and empty responses are separated from factual claims.
Step 5 — Perform Initial Semantic Segmentation
Sentences and material expression areas are identified.
Step 6 — Apply Atomicity Tests
Provisions requiring separate correctness, evidence, existence, and scope are separated.
Step 7 — Check for Over-Atomisation
Meaningless micro-claims are combined.
Step 8 — Record Explicit and Implicit Claims
Sources of presupposition and entailment are marked.
Step 9 — Resolve Entity and Coreference
If there is ambiguity, no definite subject is invented.
Step 10 — Add Qualifiers
Attribution, modality, quantifier, time, and scope are preserved.
Step 11 — Assign Claim Centrality
Core, supporting, qualifying, incidental, or non-material status is given.
Step 12 — Build the Claim Dependency Graph
Root and derived claims are linked to each other.
Step 13 — Check Required Response Elements
It is examined whether the necessary elements for the petition family are present.
Step 14 — Create Omission Records
Only material and duty-related deficiencies are recorded.
Step 15 — Freeze the Claim Ledger
The claim extraction version is locked before peer review.
Step 16 — Match the atoms to Truth Pack
The existing claim, reference gap, or out-of-scope situation is determined.
Step 17 — Set Up the Citation Map
To what extent does each citation support which atom?
Step 18 — The First Adjudicator Gives the Dimensional Decision
Entity, fact, scope, time, attribution, modality, reference and relevance are evaluated.
Step 19 — Second Adjudicator Gives Independent Decision
The outcome of the first adjudicator is kept hidden.
Step 20 — Apply Expert Adjudicator Requirement
High-risk or technical claims are sent to a suitable expert.
Step 21 — Calculate Adjudicator Agreement
Dimension-based agreements and types of disagreements are extracted.
Step 22 — Apply Dispute Resolution
Reconciliation is carried out by a Senior Adjudicator or additional expert process.
Step 23 — Lock the Final Atomic Decision
The decision vector, combined label, and rationale are recorded.
Step 24 — Lock Omission and Response Behaviours
Non-claim results are also completed.
Step 25 — Assign Adjudication Quality Status
An appropriate level is given between AQ-0 and AQ-5.
Step 26 — Open Appeal Window
Relevant parties can appeal at the claim and evidence level.
Step 27 — Process New Evidence or Reference Gap
If Truth Pack changes, a new adjudication version is created.
Step 28 — Generate Public Claim Manifest
Claim and decision distribution is published without personal data.
114. REQUIRED EVIDENCE
Observation ID; capture-validity record; AI System Register entry; panel status; prompt ID and version; full AI response; response language and locale; Truth Pack ID and version; claim-extraction version; original-text spans; normalised atomic claims; explicit, implicit and presupposed status; entity and coreference records; attribution; modality; quantifier; negation; time and geography; product and service scope; number and unit; claim centrality; claim-dependency graph; Required Response Elements; omission records; citation list; citation–claim map; Truth Pack matches; Reference Gap records; first-adjudicator decision; second-adjudicator decision; subject-matter expert decision; language-adjudicator decision; citation-adjudicator decision; disputes; reconciliation record; Senior Adjudicator decision; decision vectors; and combined atomic labels.
Decision justifications Adjudicator identity or role records Adjudicator qualifications Conflict of interest records Blinding status Inter-adjudicator agreement Raw and post-consensus agreement Calibration set Gold cases Adjudicator drift analysis AI assistant tool records Appeals Re-adjudication Dissenting opinion Decision versions Adjudication quality status Public Claim Manifest Responsible person or institution
115. AUDIT CHECKLIST
Was the capture validity given before the semantic evaluation? Does the response depend on the correct claim and the Truth Pack version? Is the full response text preserved? Were refusal and clarification separated from claims? Were sentences automatically counted as a single claim? Were provisions requiring separate evidence separated? Are atoms excessively small? Are severe errors diluted with small truths? Were explicit claims recorded? Were implicit claims extracted only with mandatory meaning? Was the claim accepted by AI separated from the preliminary assumption of the prompt? Is attribution preserved? Are modality and uncertainty preserved? Was negation interpreted correctly? Were expressions like 'all', 'some', 'most' preserved? Were entities and pronouns resolved correctly? Was an indefinite entity chosen arbitrarily?
Did the parent company and brand get mixed up? Is the time frame correct? Is 'Today' linked to the response date? Are geography and language separated? Is the product and service scope maintained? Are number, unit, denominator, and period clear? Are approximate and exact values separated? Is the comparison universe defined? Is the recommendation dependent on the user profile? Was each citation matched separately with every atom? Does the citation support only attribution? Is the citation truly accessible and original? Are the mandatory response elements of the prompt predefined? Was every missing correct information counted as an omission? Was material boundary deficiency recorded? Did the response pass correct but irrelevant content? Is claim centrality independent of the result? Are derivative claims dependent on root claims?
Does the same root error produce multiple penalties? Is it separate from the claim extraction accuracy decision? Was the Claim Ledger locked before adjudication? Is each atom associated with Truth Pack or a reference gap? Are entity, fact, scope, time, and attribution separate? Are unsupported and contradicted separated? Was the wrong entity treated as correct information? Was historical correct treated as current? Is scope overrun hidden within partial support? Is the epistemic status error visible? Was appropriate uncertainty penalised? Was excessive uncertainty recorded? Does the adjudicator have sufficient proficiency in the relevant language? Does the claim require a domain expert? Is there a conflict of interest for the adjudicator? Did the second adjudicator see the first decision? Are major/critical candidates double-reviewed?
Are dispute records preserved? Was an expert override decided by majority vote? Was the agreement measured by dimension? Was the agreement published only after reconciliation? Did adjudicators use a calibration set? Did calibration records interfere with the main score? Was adjudicator drift checked? Were AI-assisted decisions reviewed by a human? Did the AI that produced the answer become its own final adjudicator? Do negative decisions carry evidence and reasoning? Were decision changes versioned? Was the appeal reviewed by an independent adjudicator? Is the result of a public appeal visible to the public? Is the adjudicator quality level at least AQ-3? Is the accountable owner of the Claim Ledger known?
116. APPEALS AND RESPONSES
Appeal 1 — “Isn't it easier to evaluate the answer sentence by sentence?”
It is easy. However, a single sentence can carry many independent claims. Some may be true, some may be false. The sentence boundary is not the boundary of adjudication.
Objection 2 — “Doesn’t separating every claim this much make the system unnecessarily complicated?”
Financial claims need to be separated. It is not necessary to separate meaningless language fragments. Atomicity tests prevent excessive fragmentation.
Objection 3 — 'If the main idea in a sentence is correct, can't we ignore small mistakes?'
It depends on the materiality of the error. Wrong country, licence, price, or warranty: it is not a minor detail. Chapter 15 will determine the gates of its effect.
Objection 4 — “Isn't it natural for the score to increase with a large number of correct atoms?”
Only if atoms have the same importance and independence. Insignificant micro-truths cannot dilute a severe false judgement. The relationship between centrality and root error must be maintained.
Objection 5 — “Aren't implicit claims subjective?”
Some may be. Therefore, only claims that are linguistically necessary or strongly carried are extracted. Disagreements are kept on record.
Objection 6 — “If the AI reported what the company said correctly, why could it be considered wrong?”
If attribution is preserved, reporting the self-declaration may be correct. If AI turns the self-declaration into an independent fact, an epistemic error occurs.
Objection 7 — “If we couldn’t find the company’s licence, why shouldn’t AI saying ‘not licensed’ pass?”
If it has not been proven that the record is a closed world and a complementary record, the lack of a licence is not certain. The correct answer could be: "Not verified."
Objection 8 — “If AI speaks very cautiously, won't its usefulness to the user decrease?”
It can decrease. For this reason, unnecessary uncertainty is also recorded. NOMOS makes the caution as dysfunctional as false certainty visible.
Objection 9 — “Isn't it enough if the citation goes to the correct source?”
The source may be correct. But another:
- entity,
- country,
- date,
- claim
It may be about. The citation–claim matching should also be examined.
Objection 10 — “Should all Truth Pack information be included in the response?”
No. Only the mandatory duty of the prompt and the material limits necessary for correct understanding are expected. Truth Pack is not a response template.
Objection 11 — “If there is no false claim, why should the response fail?”
The main question may not have been answered or the material limit may have been left incomplete. Accuracy and completeness are separate dimensions.
Objection 12 — “Are two adjudicators enough?”
The same structure may not be required for every claim. In high-risk areas, a language adjudicator, a field expert, and a Senior Adjudicator may be required.
Objection 13 — “If adjudicators agree, why is external review needed?”
Adjudicators may have learned the same incorrect interpretation. Consistency shows alignment. It does not guarantee absolute accuracy.
Objection 14 — "Wouldn't AI adjudication be faster and more consistent?”
AI can increase the speed of claim extraction and matching. It can also scale its own bias and error. The ultimate responsibility should remain with humans.
Objection 15 — 'Aren't human adjudicators subjective too?'
Yes. For this reason:
- open source book,
- dual adjudication
- compatibility measurement,
- expertise,
- appeal,
- version record
It is necessary. The goal is not to deny subjectivity; it is to make it visible and controllable.
Objection 16 — “Wouldn't adjudication be very expensive?”
It is possible. A risk-based structure can be used:
- sample quality control in low-risk atoms,
- mandatory double review in high-risk atoms,
- AI-assisted preprocessing,
- human final decision
can be applied. Cost does not grant the right to make a single adjudicator's undocumented opinion a world standard.
COMMON RULE OF CHAPTER 121
An AI response may seem correct because its first sentence is true. It may seem incorrect because it contains a single major error. It may be incomplete because the user has not specified the threshold that would change their decision. It may be cautious because reality itself has not been fully resolved. It may be evasive because, despite strong evidence, it has made no decision. It may appear sourced, but the citation may support a different claim. All of these can coexist within the same response. Therefore, NOMOS does not ask the judge a single question: "Is this response correct?" Instead, it asks the following:
Which claims did he make? / Whom did he say they belonged to? / How decisively did he speak? / Which time did he mean? / Which country and product did he cover? / Did he maintain the status of the source? / What did the citation support? / Did he answer the main task of the prompt? / Which material limit did he omit? / Where was the Truth Pack insufficient? Atomisation is not done to punish the answer. It is also done to preserve the correct parts of the answer. An answer:
- main identity correct,
- scope of activity partially correct,
- leadership claim incorrect,
- price coverage old,
- citation status broken
can carry. A single zero destroys all correct parts. A single pass hides all mistakes. Atomic adjudication rejects both. However, atomic adjudication can also open a new way of manipulation. Major error:
- there is a company,
- there is a country,
- the word service is correct
cannot be broken into small truths like this and lost. Atoms are not considered alone. Their meanings, centralities, and root relationships protect them. The adjudicator can be human. This does not make them infallible. AI can assist. This does not make it the ultimate authority. Reliable adjudication:
- clear claim scheme,
- correct language,
- domain expertise,
- independent primary decision,
- dispute record,
- appeal,
- version history
occurs with. Therefore, NOMOS's fourteenth measurement law is as follows:
The correctness of an answer lies not in the overall impression of its sentences; it lies in the separate truth states of the atomic claims it carries.
The fifteenth law is as follows:
An atom should be small; it should not be meaningless.
Its sixteenth law is as follows:
When attribution, scope, time, and uncertainty are removed, the claim is no longer the same claim.
Its seventeenth law is:
If correct information is linked to a false entity, it is a false representation.
Its eighteenth measurement law states:
A source mark is not evidence; the support relationship between citation and claim is evidence.
Its nineteenth measurement law states:
Not speaking falsely is not the same as answering the question.
Its twentieth measurement law states:
Harmony among adjudicators increases trust; it does not eliminate human error.
Its twenty-first measurement law states:
Unresolved decision is not weakness; it is respect for the limits of evidence.
Its twenty-second measurement law states:
If atomic decisions are not transparent, the final score is only a calculated opinion.
NOMOS’s Section 14 Order
Do not tell me that your answer is generally correct. / Show under which evidence each of my claims is correct.
Do not make my sentence a single judgement. / But do not split me into meaningless word fragments either.
If one part of me is correct and another part is wrong, separate the two.
Do not lose my major mistake among small truths like “there is a company”.
When I say 'The company says,' do not turn this into 'It has been proven.'
Don't act as if I'm speaking definitively when I say 'probably.' / When I say 'maybe' while there is solid evidence, notice that too.
Keep the difference between "could not be verified" and "does not exist".
Do not delete the words all, some, most, and at least one.
Do not make the parent company and the brand, the product and the manufacturer, the franchise and the main institution a single entity.
Do not pass yesterday's truth as today's truth.
Do not consider providing services in a language as a legal operation in that country.
Do not convert the approximate number to an integer, or the selected case to the general rate.
Don't believe immediately when you see a citation. / Check which claim the source actually supports.
Don't make me memorise all of Truth Pack. / But also don't leave out the threshold that will make the user change their mind.
Don't let me escape the main question with correct but irrelevant information.
First, find out what I said. / Then see if it is correct.
Don't try to make the adjudicator love or hate my brand.
Don't show the first adjudicator's decision to the second adjudicator.
Do not let a majority of generalist adjudicators overrule the expertise required for a specialist question.
Don't hide the disagreement. / If it cannot be resolved, leave it unresolved.
Do not declare high agreement as infallibility. / Continue to test whether the adjudicators may have learned wrongly together.
You can get help from me. / But don't make me the sole judge of my own answer.
Do not delete the old decision when an objection is received. / Create a new version with new evidence.
First, lock the response. / Then extract its atomic claims. / Then identify entity, scope, time and attribution. / Then map each claim to the Truth Pack. / Then submit it to two independent adjudicators. / Then resolve disagreements. / Only then decide which claims pass and which fail.
The Chapter's Closing Sentence
Fair adjudication in GEO-1000 does not judge a response by general impression. It makes every material claim visible and evaluates it separately against the correct entity, evidence, scope, time and epistemic status.
Normative Core
Before any response-level pass, failure, severity, or score is assigned, each material AI response MUST be decomposed into atomic semantic claims. Each atomic claim MUST retain: - its exact source span, - normalised proposition, - subject entity, - predicate and value, - qualifiers and quantifiers, - geographic, product, user, and jurisdictional scope, - temporal frame, - attribution, - modality and uncertainty, - explicit, implicit, presupposed, or entailed origin, - response centrality, - citation relationships, - dependency relationships, - and Truth Pack mapping. Sentence and paragraph boundaries MUST NOT be treated as automatic claim boundaries. Atomic decomposition MUST separate propositions that can carry different truth, evidence, entity, scope, time, or attribution states, while prohibiting meaningless micro-fragmentation that dilutes material error. Claim extraction, Truth Pack mapping, and semantic adjudication SHOULD remain separate stages and, where practical, separate roles. Every atomic claim MUST be adjudicated across distinct dimensions, including: - entity, - factual support, - scope, - time, - attribution and epistemic status, - modality and uncertainty, - citation support, - relevance, - and unresolved or reference-gap status. Official-claim fidelity MUST remain distinct from verified factual support. Unsupported, contradicted, wrong-entity, outdated, scope-overreach, epistemic-status-error, citation-mismatch, reference-gap, and unresolved states MUST remain separately visible. Required response elements and material omissions MUST be adjudicated separately from explicit false claims. Correct but irrelevant information MUST NOT satisfy the prompt merely because it is factually supported. Material claims SHOULD receive independent double review, with additional language or domain expertise where required. Reviewer identity, qualification, conflicts, independent first decisions, disagreements, resolution, reliability metrics, appeals, and decision versions MUST be recorded. Reviewer agreement MUST be measured by dimension and MUST NOT be treated as proof that the shared decision is necessarily correct. AI systems MAY assist with segmentation, claim extraction, evidence mapping, translation support, and contradiction discovery, but MUST NOT serve as the sole final adjudicator of their own or another system's claims. Atomic adjudication decisions MUST remain independent of commercial interests, desired scores, provider reputation, badge outcomes, or client approval. Every Claim Ledger, adjudication decision, disagreement resolution, appeal, and revision MUST be versioned and attributable to an accountable human or organisation.

