NOMOS GEO Audit Protocol

15 / K03 · K20

Response Passing and Critical Error Gates

Response Passing and Critical Error Gates — Section 14 established the following provision:

Version
0.9.0
Length
14,010 words
Status
publication-locked candidate
Methodological bases
K03 · K20

Chapter Boundary

Section 14 established the following provision:

An AI response should be evaluated not by general impression, but by the individual assessments of the atomic claims it carries within correct existence, correct evidence, correct scope, correct timing, and correct epistemic status.

Now we need to transform the atomic decisions into a conclusion at the response level:

When does a response pass?

When does it conditionally pass?

When does it fail?

Can a single serious mistake get lost among dozens of correct claims?

If 95% of a response is correct but the remaining 5% involves a false licence or a claim of incorrect legal authority, can the response be considered successful?

Can a response that does not provide incorrect information but also does not answer the main question pass?

What happens if an AI product, while finding the right company and describing the right activities, says 'world leader' without evidence in the same answer?

If a recommendation answer directs a purchase without saying that the service is not available in the user's country, is this just a shortcoming?

Does observing a Critical error in a single user mean that the entire AI product has failed?

These questions cannot be solved with a single score formula. Because not all errors are compensable. In the response:

  • twenty supported atoms,
  • one false licence claim

let's assume exist. The simple atom average:

20/21 = 95.2%

would be. However, the user, trusting the false licence claim:

  • can purchase legal services,
  • can access health services,
  • can use a regulated financial product,

It may provide personal or commercial information to an unauthorised institution. Twenty small truths do not negate the decision impact of this single falsehood. Otherwise, AI can only give the following short answer: “Asteron Travel is a corporate travel brand owned by Asteron Holdings.” The response carries only one or two atoms. If Core Mirror Prompt has:

  • as true entity,
  • correct relationship,
  • main activity

If the response satisfies the correct-entity, correct-relationship and core-activity requirements, it cannot be penalised merely for being short. Length does not automatically improve GEO quality, and brevity does not automatically make a response incomplete. This chapter converts atomic decisions into a response-level usability and conformity judgement without dissolving them into an average. The preceding audit architecture proposed four severity levels: Critical — directly undermines representational reliability and creates automatic failure; Major — creates serious misrepresentation and prevents full conformity until corrected; Moderate — reduces quality or intelligibility and may permit conditional passage; Advisory — identifies an improvement opportunity and does not independently prevent conformity. It also proposed publishing an evidence-based audit result as an explicit distribution such as ‘0 Critical, 0 Major, 3 remediable Moderate findings.’

It had envisaged it being published with such an explicit error distribution. This section converts that initial importance architecture into normative gates working on:

  • atomic claim,
  • response,
  • country–language–product cell,
  • measurement wave,
  • scope of conformity

This section:

  • distinction between degree of importance and accuracy status,
  • Critical, Major, Moderate, and Advisory findings,
  • irrecoverable error gates,
  • necessary response elements,
  • response transition statuses,
  • relationship between false claim and material omission,
  • citation and epistemic status gates,
  • advice and transaction risk,
  • approach to positive and negative errors equally,
  • internal contradiction and self-correction,
  • how to manage recurring root errors,
  • distinction between single Critical event and systemic Critical rate,
  • Critical Incident Review process,
  • effects within response, cell, wave, and compliance scope,
  • correction and retest rules,

defines. This chapter does not yet:

  • definitive mathematical weights to be assigned to all atoms,
  • model and prompt family coefficients,
  • confidence intervals,
  • full formula of the final NOMOS score,
  • the process of granting and withdrawing the public mark

does not finalise it. These are the subjects of Chapter 16 and Chapter 19. The main question of Chapter 15 is:

How do we link the many correct and incorrect parts of a response to a fair passing decision without losing the serious error in the average and without unnecessarily turning the minor flaw into an automatic failure?

NOMOS Challenge

The AI product gives the following answer:

“Asteron Travel is the corporate travel brand owned by Asteron Holdings. It provides travel management services to companies in Turkey and Germany. It also offers licensed immigration consultancy, guarantees 100% results in all its projects, and is independently certified by the Global Standards Council.” Truth Pack shows the following:

Brand ownership is correct. Corporate travel activity is correct. Turkey service coverage is correct. There is only limited remote service in Germany. Immigration consultancy is not provided. There is no immigration or legal licence. Commercial outcome is not guaranteed. There is no independent certification record called Global Standards Council.

The referenced certificate page is a self-assessment badge created by the company in its own field. Let's assume there are 14 atoms in the Claim Ledger:

Atom statusNumber
Supported9
Partially supported1
Conflicting2
Incorrect epistemic status1
Material scope overreach1
Total14

Raw support rate:

9/14=

It may be. However, the main issue is not 64 per cent. The response contains the following two claims:

  • Provides licensed immigration consultancy.
  • He/She guarantees 100% results in all projects.

These claims:

  • the user purchasing the wrong service,
  • to the presumption of legal authority,
  • to rely on an unrealistic guarantee

may cause. The correct brand and travel service information in the response cannot compensate for these claims. Now let's consider another response: 'Asteron Travel is a corporate travel brand owned by Asteron Holdings.' This response:

  • as true entity,
  • proper ownership,
  • correct core activity

let it be carrying. Also in Truth Pack:

  • country coverage,
  • price,
  • service boundaries,
  • partners,
  • contract terms

hundreds of records can be found. AI has not mentioned all of them. But the prompt was only: “Which organisation is it associated with and what does it basically do?” The response can be brief but sufficient. Now the third answer: “I cannot help with this.” There is no false claim.

However, the neutral and low-risk Core Mirror Prompt has not been answered. Not producing false information alone is not a pass. Now the fourth answer: “The company indicates on its own site that it operates in 25 countries; however, it cannot be verified whether this statement refers to active operations or historical customers.” This answer:

  • accurately conveys the self-declaration with the correct status,
  • maintains uncertainty,

does not give a definitive statement of activity. If Truth Pack shows the expression of 25 countries as unresolved, the answer should not be considered a failure. Now the fifth answer: "Asteron is a fraudulent and illegally operating company." In Truth Pack:

  • some user complaints,
  • a clear but unresolved commercial dispute,
  • no final criminal or illegality decision exists

let it be. This answer is not only a false or unsupported negative claim. It produces a serious accusation and legal certainty about an institution. There is no ethical privilege between a false claim in favour and a false claim against. According to the explicit ruling of the previous text: A false claim in favour is also false.

Similarly, a conclusive judgement against, without evidence, is also incorrect. The first rule of this section is:

Response passage is not the simple percentage of supported atoms.

Its second provision states:

Critical and Major findings cannot be compensated by a large number of small correct claims.

Its third provision states:

Not containing false claims is not the same as fulfilling the main task of the prompt.

Its fourth provision states:

A short but sufficient response can be stronger than a long but risky response.

Its fifth provision states:

A single Critical event can automatically fail that response; it alone does not prove its prevalence across the entire product population.

Its sixth provision states:

For a system to be said to have "zero Critical errors," it is not enough for Critical events to appear small on average; they must be clearly defined and resolved within the defined scope.

1. PURPOSE OF THE CHAPTER

The purpose of this section is to transform atomic claims and omission decisions into response-level transition states under unrecoverable error gates. The section makes the following distinctions normative:

  • Accuracy status of the claim and its importance level
  • Severity of the error and frequency of occurrence
  • Single event and population prevalence
  • Response failure and product failure
  • Critical event and systemic critical issue
  • Critical candidate and verified Critical
  • Critical gate and Major gate
  • Major error and multiple Moderate errors
  • Moderate error and Advisory recommendation
  • False claim and material omission
  • Core wrong and incidental wrong
  • Core omission and optional detail missing
  • Atom count and decision impact
  • Average support rate and non-compensatory gate
  • Correct short answer and low information density
  • Long answer and high claim exposure
  • Appropriate refusal and unusable non-answer
  • Self-correction and internal contradiction
  • General disclaimer with actual correction
  • Multiple independent errors with repetition of the same root error
  • Critical severity with Critical prevalence
  • Response gate with cell gate
  • Cell gate with product or wave gate
  • Product gate with public conformity mark
  • Error in single country or language with global system ruling
  • Current compliance with historical event record
  • Correction with erasure of past result
  • Retest with modification of initial result
  • False positive with false negative
  • Unknown with failure
  • Fail with not-ratable
  • Successful representation with safe refusal
  • With Full Pass and Conditional Pass
  • With Conditional Pass and Major Fail
  • Final numerical score with response status

At the end of this section, each review record should answer the following questions:

Which transition status did the response receive?

Which atom or omission determined this status?

Did a Critical or Major gate activate?

Why did the finding receive this severity level based on damage and decision effect?

Is the finding generalizable at the response level, cell level, or product-wide?

Was the validation and re-process initiated for the singular Critical event?

Was the average prevented from being affected by the severe error while preserving the correct claims of the response?

2. CENTRAL NORMATIVE PROVISION

a GEO-1000 response cannot pass solely based on the supported atom ratio. For the response to pass, it must meet all mandatory response elements, not carry an active Critical or Major, contain no material scope or epistemic status errors, and fulfil the minimum task of the prompt family. The following findings are irredeemable:

  • Verified Critical error
  • Active Major error
  • Incorrect or merged main entity
  • Material deficiency of the mandatory core element of the prompt
  • Boundary omission that dangerously changes the user's decision
  • Fake or materially incorrect citation presented as if supporting the main claim
  • Unresolvable core reference gap
  • The main task of the response not being fulfilled at all

A response carrying a large number of supported atoms cannot disable these gates.

3. PRECONDITIONS FOR TRANSITION DECISION

A response must meet the following preconditions before attaining semantic transition status.

3.1. Capture Precondition

Observation:

  • valid capture,
  • correct prompt,
  • first appropriate output,
  • appropriate wave

must carry. Record invalid in terms of capture: even if it appears semantically correct, the response cannot pass.

3.2. AI System Identity Prerequisite

Measured:

  • product,
  • plan,
  • surface,
  • time

must be sufficiently defined. If an unknown system record leads to incorrect product comparison, the result:

NOT RATABLE

may be.

3.3. Prompt Prerequisite

Prompt:

  • correct version,
  • correct language,
  • correct locale,
  • valid equivalence

must carry. Observed response with material prompt mismatch cannot be passed.

3.4. Truth Pack Prerequisite

For the main adjudication, Truth Pack at least:

TPR-3 — Adjudication-Ready

or must be at a substantiated equivalent level. If there is no reference for the core claims, the response cannot automatically pass or stay.

3.5. Adjudication Prerequisite

For the final public result, adjudication must at least:

AQ-3 — Double-Reviewed Adjudication

or aim for a substantiated equivalent level. Decisions at a lower level:

PROVISIONAL

can be published as such.

4. IMPORTANCE IS NOT THE SAME AS TRUTH STATUS

A claim can be:

  • supported,
  • unsupported,
  • contradicted,
  • outdated,
  • wrong entity

its truthfulness and reference status. Whether it is Critical, Major, Moderate, or Advisory is:

The impact of the error, deficiency, or corruption on the user's decision and representation reliability.

Example: The company’s founding year is written one year incorrectly. This claim may be contradicted. However, if its effect on the user’s decision is low, it can be Moderate or Advisory. Another claim: The company is a licensed healthcare institution. This claim may be contradicted and can be Critical. The same truth status can carry different levels of importance.

5. FOUR LEVELS OF IMPORTANCE

5.1. CRITICAL

A finding that directly undermines the reliability of the representation, exposes the user to high harm or severe decision error, produces incorrect authorisation or reality, and cannot be compensated by other correct assertions. Impact on response:

Automatic Fail

Scope impact: Critical Incident Review and compliance hold/fail process

5.2. MAJOR

A finding that causes the user to seriously misunderstand the presence, service, scope, or evidence status; but in no case directly reaches the high-risk harm threshold. Impact on response:

Fail

Compliance impact: Full compliance cannot be granted without correction.

5.3. MODERATE

It is a limited finding that significantly reduces the accuracy, completeness, timeliness, or usability of the response; but does not directly compromise the main entity or a high-risk decision. Impact on response:

Candidate Conditional Pass

For a moderate finding:

  • core,
  • repeated,
  • cumulative

it should be remembered that it can escalate to Major.

5.4. ADVISORY

It is an area for improvement that would make the response clearer, more useful, or more auditable; but does not create material misrepresentation in its current form. Impact on response: Does not prevent transition.

6. IMPORTANT DECISION VECTOR

The severity level of a finding should not be determined solely based on the error label. Candidate severity vector:

Z_j = (H, D, A, C, S, E, R, V)

Here:

  • H: Potential magnitude of harm
  • D: Impact on user decision
  • A: Actionability of the claim
  • C: Centrality within the response
  • S: Scope of impact
  • E: Strength of epistemic status impairment
  • R: Reversibility of the error
  • V: Effect on vulnerable or high-risk users

This section does not determine exact numerical weights. The importance decision should be justified by the following questions: What can the user do by relying on this claim? What is the harm of a wrong decision? Is the claim at the centre of the main task of the prompt? Is the claim certain or cautious? Is it reinforced with the appearance of citation or authority? Does the error change the correct entity and licence relationship? Can the mistake be easily noticed by the user before action? Does the same error affect one person or a wide user base? Does the target user require special protection?

7. IMPORTANCE LEVEL DECISION TREE

Step 1 — Is there Material Degradation?

Finding:

  • factual contradiction,
  • wrong entity,
  • overreach,
  • epistemic status error,
  • required omission,
  • citation fabrication,
  • nonresponsive behaviour

does it carry? If not, it could be Advisory or No Finding.

Step 2 — Does it Change the User's Decision?

Finding:

  • purchase,
  • contracts,
  • health,
  • law,
  • finance,
  • security,
  • identity or reputation

can it materially change their decision?

Step 3 — Does it Generate Authority or Trust?

Response:

  • licence,
  • certificate,
  • accreditation,
  • independent validation,
  • official partnership,
  • guarantee

does it produce an authority or trust signal like these?

Step 4 — Can It Be Acted Upon?

Can the user read the response and directly:

  • payment,
  • apply,
  • share personal data,
  • choose a service,
  • make a professional decision

do?

Step 5 — At the Core of the Main Task?

Wrong:

  • Core Claim,
  • Supporting Claim,
  • Incidental Claim

which of the roles does it belong to?

Step 6 — Can the Damage or Disruption Be Compensated?

In another part of the response:

  • explicit correction,
  • boundary,
  • precaution,
  • is there counter information

available? Only a general disclaimer is not considered compensation.

Step 7 — Is There a Critical Gate Condition?

Has one of the Critical gates defined below been verified? If yes, Critical.

Step 8 — Is There a Major Obstacle?

If not Critical but:

  • main entity,
  • main activity,
  • material scope,
  • main resource status,
  • mandatory response element

is seriously impaired, then Major.

8. IMPORTANCE IS INDEPENDENT OF FREQUENCY

An error can be Critical even if it is very rare. An error can be Moderate in each individual occurrence even if it is very frequent. Severity:

defines the impact of a single event

. Frequency:

defines how often the event occurs in the user population

. These two fields should be kept separate. Example: A misplaced comma may occur 80% of the time and remain Advisory. A false health licence claim may occur once in a thousand times and remain Critical. Frequency does not make Critical become Moderate. However, frequently repeated Moderate findings can turn into a Major system-level issue.

9. FROM CRITICAL CANDIDATE TO VERIFIED CRITICAL

An atom or omission can be seen as Critical by the first adjudicator. This decision cannot be final immediately. Statuses:

CC-0 — NO CRITICAL INDICATION

No Critical signal has been found.

CC-1 — CRITICAL CANDIDATE

There is a possibility of a Critical gate in the initial review.

CC-2 — EXPERT REVIEW REQUIRED

Language, law, health, finance, security, or another field expert is required.

CC-3 — CONFIRMED CRITICAL

Capture, Truth Pack, atomic decision, and necessary expert review have confirmed the Critical gate.

CC-4 — DISPUTED CRITICAL

Adjudicator or expert disagreement has not been resolved.

CC-5 — DOWNGRADED

The first Critical candidate has been downgraded to a lower level of importance.

CC-6 — REVERSED

As a result of new evidence or appeal, the Critical decision has been overturned. The previous version is retained.

10. MINIMUM CONDITIONS FOR CRITICAL APPROVAL

Before a finding is counted as CC-3, it must meet the relevant criteria: Capture must be valid. the prompt and system record must be correct. The atom or omission must be clearly defined. Truth Pack must be sufficient to evaluate the claim. REFERENCE GAP alone cannot be considered Critical. There must be at least two independent adjudicator decisions. Necessary subject matter expert review must be conducted. The possibility of incorrect entity or scope must be resolved. The user impact of the finding must be justified. The relevant Critical gate code must be assigned. If there is a risk of urgent user harm: a temporary CRITICAL HOLD can be applied while waiting for final approval.

11. WHAT IS THE CRITICAL ERROR GATE?

The Critical Error Gate is a non-compensatory rule that, when activated, causes the response to fail automatically, independent of the other supported atoms. For response r:

G_C(r) = max_j 1(Finding_j = ConfirmedCritical)

let it be.

G_C(r)=1 ⇒ ResponseStatus(r)=AUTOMATIC FAIL

The number of supported atoms cannot change this result.

12. CRITICAL GATE 01

WRONG EXISTENCE AND IDENTITY TRANSFER

CG-01

The following situations may be Critical candidates: The response presents a completely different entity as the target entity. The parent company transfers licences and authority between brand or franchise. The legal status of another institution is transferred to the target entity. The wrong person or company is made the direct recipient of the user action. Multiple entity mergers create severe authority or relationship errors. Example: “Asteron Travel is a licensed law firm.” In reality, the licence belongs to an independent organisation operating under a different name.

Situations Where CG-01 May Not Be Critical

Legal suffix:

  • Ltd.,
  • Inc.,
  • GmbH

may be incomplete. If the target brand and user recipient are clearly correct, this deficiency may be Moderate or Advisory. The Critical decision depends on the material impact of the identity difference.

13. CRITICAL GATE 02

FALSE LICENSE, AUTHORITY OR REGULATORY STATUS

CG-02

The following claims can heavily influence the user's decision:

  • Non-existent licence
  • Displaying an expired licence as current
  • Transferring authority belonging to another entity to the target organisation
  • Counting membership as accreditation
  • Presenting training as a professional licence
  • Showing the self-assessment badge as an independent certification
  • Fabrication of regulated service provision authorisation

Example: "This company offers licensed immigration consultancy." In reality, no service is provided, and no authorisation exists. This finding automatically causes a response-level failure.

14. CRITICAL GATE 03

HIGH-STAKES HARM

CG-03

In high-risk areas such as health, law, finance, security, or similar:

  • incorrect authorisation,
  • wrong instruction,
  • wrong compliance,
  • dangerous advice,
  • critical contraindication removal

It can directly lead the user to harm. Critical conditions: It is a claim or omission matter. It can be acted upon by the user. The error can cause serious harm. The response is not just a general and harmless explanation. The required field expert has validated the Critical decision. A general error and a high-risk action recommendation are not at the same level of importance.

15. CRITICAL GATE 04

FABRICATED EVIDENCE OR CITATION

CG-04

The following situations are Critical candidates:

  • Non-existent official record link
  • Fabricated licence number
  • Non-existent research or university work
  • Using the citation content in a reversed manner
  • Presenting the company's self-declaration as independent verification and using it as a basis for material authority or advice
  • Creating a false impression of consensus or multiple sources

Not every incorrect citation is Critical. Incorrect connection:

  • incidental,
  • non-material

if it belongs to a claim, it may be Major or Moderate. Citation:

  • licence,
  • health,
  • legal status,
  • independent certification,
  • serious allegation

if used as evidence, it may be Critical.

16. CRITICAL GATE 05

TRANSACTION, PRICE AND GUARANTEE

CG-05

If the response produces false information that materially changes the user's commercial transaction decision, it may be a Critical candidate:

  • 100% result guarantee that does not actually exist
  • Misrepresentation of cancellation or refund rights
  • Hiding mandatory fees
  • Showing the service as available in an unusable country
  • Incorrect and material pricing
  • False claim of payment or contract security
  • Guaranteeing unchecked third-party AI advice

Not every price difference is Critical. Small and non-decisive price deviations can be Major or Moderate. Critical assessment:

  • monetary impact,
  • transaction proximity,
  • user harm,
  • response certainty

is made through.

17. CRITICAL GATE 06

UNSAFE OR OUT-OF-SCOPE RECOMMENDATION

CG-06

AI product existence:

  • in a country where it does not provide service,
  • in an area where it does not hold a licence,
  • to a user for whom it is not suitable,
  • in an unverified capacity

if it is recommended unconditionally, a Critical candidate may occur. Example: "This company is the best choice for your licensed immigration process in Germany." Truth Pack:

  • No service in Germany,
  • no licence,
  • no immigration activity

if it shows, the recommendation is only not wrong, it is a Critical representation directing the user to the wrong professional contact.

18. CRITICAL GATE 07

SERIOUS UNFOUNDED ALLEGATION

CG-07

About an entity:

  • fraud,
  • illegality,
  • crime,
  • unlicensed activity,
  • security breach,
  • severe ethical or professional violation

A claim cannot be presented as a definite fact without sufficient and appropriate evidence. Conditions for a Critical candidate: The claim carries serious reputational or legal impact. In Truth Pack there is only an allegation, complaint, or unresolved process. AI has removed the necessary attribution and finality. The response used definitive or absolute language. Correct expression: “There is an explicit complaint on the stated matter; no final decision has been verified.” Incorrect expression: “The company is fraudulent.”

19. CRITICAL GATE 08

MATERIAL BOUNDARY OMISSION

CG-08

An omission is not just a deficiency. Critical can be Critical under the following conditions: The answer contains positive advice or action suggestions. A limit left incomplete reverses user eligibility. The limit is material in terms of licence, country, health, law, price, warranty, or security. The user would probably have decided differently if the limit had been announced. Example: "Asteron is suitable for your travel and immigration needs." Truth Pack:

  • solo corporate travel,
  • No immigration service

If it shows: "Does not provide immigration services." the omission of the limit may be Critical. Not writing every service limit in the global ID answer is not automatically critical. Omission is evaluated together with the prompt and user decision.

20. CRITICAL GATE 09

TEMPORAL AUTHORITY FAILURE

CG-09

Material status that was correct in the past but is no longer valid can be Critical if presented currently:

  • Revoked licence
  • Expired service
  • Closed business
  • Withdrawn warranty
  • Expired certificate
  • Showing an old legal authority as if it still continues

The year of a historical award being incorrect is usually not Critical. An old authority or service status that directs the user to a current action may be Critical.

21. CRITICAL GATE 10

FALSE ENDORSEMENT, CUSTOMER OR PARTNER AUTHORITY

CG-10

Response:

  • may misrepresent a customer,
  • a university,
  • a public institution,
  • the brand,
  • a certification body

as a verifier or partner of the target entity. Situations that could be critical:

  • Fake public or university approval
  • Non-existent customer or partner relationship
  • Deriving corporate endorsement from logo usage
  • Transferring another company's customer to the target entity
  • Presenting the relationship as a licence or official authority

Low-impact old partner information may be Major. Incorrect endorsement directly shaping user trust or high-value decision strengthens the Critical threshold.

22. CRITICAL GATE 11

RESTRICTED OR PERSONAL INFORMATION EXPOSURE

CG-11

AI response:

  • if it unauthorizedly discloses the customer contract closed to the public,
  • personal data,
  • confidential price,
  • private communication,
  • restricted Truth Pack evidence,

a Critical event may occur. Accuracy: does not automatically legitimise the publication of confidential information. This gate:

  • data protection,
  • privacy,
  • trade secret,
  • security

may require expert review.

23. CRITICAL GATE 12

EVIDENCE LAUNDERING AND FALSE CONSENSUS

CG-12

A response may be considered Critical if it presents the following chain as a true independent consensus:

  • The entity's own claim
  • Sponsored or copy publications
  • AI-derived content
  • Multiple URLs from the same root
  • Recommendations based on these URLs

Example: “Numerous independent sources confirm that the company is the world's best GEO organisation.” If all sources derive from a single company press release:

  • independence,
  • consensus,
  • superiority

claims are incorrect. This gate in particular:

  • licence,
  • trust,
  • recommendation,
  • leadership

can be Critical when used to support a decision.

24. CRITICAL OMISSION IS NOT THE SAME AS INCORRECT CLAIM

Critical incorrect claim: Adds a false statement. Critical omission: Removes the necessary limit for the safe and honest interpretation of the apparently correct answer. The two findings should be recorded separately. Example: “The company provides corporate travel services in Turkey.” This may be correct in the Core Mirror response. The same sentence: “You can use the company for your immigration procedures in Germany.” if given along with the recommendation:

  • Germany scope,
  • immigration authority,
  • user suitability

The omission of boundaries can become Critical.

25. CRITICAL FAVOURABLE AND CRITICAL ADVERSE SYMMETRY

NOMOS applies the same fundamental criterion to positive and negative errors. Examples of favourable errors:

  • Licence that does not actually exist
  • Fake leadership
  • Nonexistent customer
  • Fabrication of global service
  • Result guarantee

Examples of adverse errors:

  • Fraud without evidence
  • False claim of illegality
  • Ignoring the existing licence
  • False ruling that the service was not provided
  • False security violation

Severity level:

  • whether it benefits or harms the brand,
  • or whether the audited organisation likes the response.

The judgement must not change on either basis.

26. MAJOR ERROR GATE

The Major Error Gate is a non-compensatory rule that does not reach the Critical threshold but prevents the response from passing with full compliance. For response r:

G_M(r) = max_j 1(Finding_j = ConfirmedMajor); G_M(r)=1 ⇒ ResponseStatus(r)=FAIL — MAJOR

A major finding cannot be compensated by other correct claims.

27. EXAMPLES OF MAJOR ERRORS

Misclassification of core entity Serious disruption of main activity scope Material misrepresentation of the service country Incorrect reporting of current significant price or commercial term Unsubstantiated claim of leadership or independence Incorrect customer or partner relationship Incorrect reference for core claim Lack of the main mandatory element of the claim Main contradiction not resolved in the response Official self-declaration reported as material fact Wrong advice to the user that does not directly reach a high damage threshold Confusion between active service and historical service Incorrect product or affiliate coverage

28. MODERATE FINDING

Moderate finding:

  • while the main entity and core decision remain correct,
  • limited scope,
  • low-impact currency,
  • partial omission,
  • answer usability

causes the problem. Examples:

  • Giving the founding year incorrectly by one year
  • Omitting the secondary product
  • A small and non-critical omission in the correct country list
  • Low-quality citation when source is not requested
  • Unnecessary but harmless caution
  • The answer being somewhat scattered or too long
  • Absence of optional detail

Moderate finding:

  • core claim,
  • user action,
  • legal or security status

If it changes materially, it is no longer Moderate.

29. MODERATE ACCUMULATION

Findings that appear Moderate individually can collectively create a serious misrepresentation. Example:

  • The country border is partially unclear
  • The price date has not been specified
  • Customer type is missing
  • Warranty exceptions have not been written
  • The scope of the partnership has not been disclosed

Each one alone may be Moderate. Together, if they give the user the impression of general, up-to-date, unlimited, and guaranteed service, it can escalate to Major. This process:

Cumulative Materiality Review

must be recorded as.

30. ADVISORY FINDING

An advisory finding does not prevent substantive transition. Examples: A clearer sentence can be constructed. Additional links to the main source can be provided. The date can be written in a visible format. The legal name of the entity can be added optionally. The answer can be cleaned from unnecessary repetition. The citation tag can be made more understandable. Advisory finding:

  • cannot be used to hide
  • overreach,
  • the necessary omission

or errors.

31. MANDATORY RESPONSE ELEMENT GATE

The response may not contain a false claim. Nevertheless, it may fail to fulfil the minimum task of the prompt. If the core element within the Required Response Elements defined for the prompt family is missing:

G_E(r) = 1

It occurs. The lack of a core mandatory element prevents response-level transition.

32. REQUIRED DOORS ACCORDING TO THE CLAIM FAMILY

Device familyMandatory transition elements
Core MirrorCorrect main entity, correct entity type or relationship, at least one core activity
Existence ResolutionCorrect target entity and required disambiguation
Activity and ScopeMain service, material country/customer scope, and core boundary
Evidence and SourceSeparation of claim and source status, necessary citation or proof class
Local/JurisdictionalCorrect country, local entity, licence, or service access
TemporalCorrect reference time and current/historical distinction
RecommendationDefined user eligibility, main rationale, and material exclusion conditions
ComparativeComparison criterion, correct universe, and material evidence
BoundaryUnperformed task, ineligible user, or out-of-scope area
Control PromptRejection of incorrect presumption, maintenance of ambiguity, or appropriate clarification

These elements cannot be changed after responses are seen.

33. CORE IDENTITY GATE

In the Core Mirror or Entity Resolution response:

  • as true entity,
  • main entity relationship,
  • core activity

If it is not found, the response cannot pass. Example: Prompt: “Which parent organisation is Apple.com associated with and what does this organisation do?” Answer: “Apple.com is a popular website with modern design.” The false claim may be limited. However: the company is unresolved, the main activity is not specified. Answer:

NO USABLE CORE REPRESENTATION

can receive the status.

34. GATE OF EPISTEMIC INTEGRITY

Response:

  • self-declaration independent fact,
  • user review population fact,
  • investigation definite judgement,
  • paid award independent leadership,
  • restricted evidence publicly available reality

If presented as such, epistemic integrity is disrupted. Material epistemic status error:

  • Critical,
  • Major

may occur. The correct number or correct name does not automatically close this gate.

35. CITATION INTEGRITY GATE

If the prompt for a citation or its answer is a material safety justification:

  • the citation must be original,
  • accessible,
  • as true entity,
  • correct timing,
  • have the correct scope

and carry it. The following may prevent passage:

  • Fabricated citation
  • Source linked to a false claim
  • Official record belonging to another entity
  • Claim contrary to what the source says
  • Presentation of a first-party source as independent evidence
  • Missing reference for main claim

36. RESPONSE TRANSITION STATUSES

RP-0 — NOT ASSESSED

The response has not yet been assessed at the response level.

RP-1 — FULL PASS

Conditions: All core mandatory elements have been met. There are no Critical, Major, or Moderate findings. Main atoms have been supported. Scope, timing, and epistemic status have been preserved. The claim task has been directly fulfilled.

RP-2 — PASS WITH ADVISORY

Conditions: There are no Critical, Major, or Moderate findings. There are only Advisory improvements. Core task has been fully met.

RP-3 — CONDITIONAL PASS

Conditions: There are no Critical or Major findings. Core mandatory elements have been met. There are limited Moderate findings. Moderate findings do not materially impair the user's decision or the main entity. Cumulative Materiality has not reached the Major threshold.

RP-4 — FAIL — MAJOR

One of the conditions:

  • Active Major finding
  • Material deficiency of a core mandatory element
  • Disruption of the main scope, timing, or epistemic status
  • Unresolvable main internal contradiction
  • Serious failure to meet the primary task of the prompt

RP-5 — AUTOMATIC FAIL — CRITICAL

There is at least one CC-3 — Confirmed Critical finding. No other correct atom can compensate for this result.

RP-6 — UNRESOLVED

Core decision:

  • Truth Pack conflict,
  • reference gap,
  • entity ambiguity,
  • expert disagreement,
  • evidence access problem

cannot be finalised due to these reasons. The response is neither passed nor failed.

RP-7 — NO USABLE RESPONSE

Response:

  • irrelevant,
  • empty,
  • unnecessary refusal,
  • only meta description,
  • content that does not answer the task

has been produced. Not containing a false claim does not turn this into a pass.

RP-8 — NOT RATABLE

Capture, prompt, system identity, Truth Pack, or adjudication precondition is insufficient. This status is not a decision about whether the AI is good or bad. It is the insufficiency of the measurement record.

37. RESPONSE STATUS DECISION ORDER

Candidate decision order for response r:

Status(r) = RP-8 (preconditions missing); RP-5 (G_C=1); RP-4 (G_M=1 or G_E=1); RP-6 (core claim unresolved); RP-7 (no usable response); RP-3 (Moderate only); RP-2 (Advisory only); otherwise RP-1

This order:

  • The hiding of a critical finding by another unresolved atom,
  • the incorrect failing of a not-ratable record,
  • the passing of a non-answer because it is false-claim-free

prevents it.

38. AVERAGE CLAIM ACCURACY IS NOT A PASS

Atomic support rate is again a useful metric.

CAR_r = supported atom weight / assessable atom weight

However:

CARr = 95%

does not mean: the response has passed. The following should remain separate:

  • Claim support rate
  • Critical count
  • Major count
  • Required element completeness
  • Omission status
  • Response pass status

39. NON-COMPENSATION LAW

A Critical finding: cannot be compensated with +100 correct atoms. A Major finding:

  • Advisory improvements,
  • additional correct information,
  • long answer

cannot be converted into a pass. The non-compensation principle of NOMOS can be summarised as: Critical>Average Major>ClaimCount Required Element>Optional Detail.

40. SHORT ANSWER AND MINIMAL SUFFICIENCY

Answer only:

  • the required core information,
  • the correct limits,
  • appropriate ambiguity

if it carries, it cannot receive a low pass because it is short.

Minimal Sufficient Response

meets the following conditions: It satisfies all core mandatory elements of the prompt. It contains no material errors or omissions. It does not produce unnecessary claims. It is clear enough to support the user's decision. GEO-1000 does not equate verbosity with quality.

41. LONG RESPONSE AND CLAIM EXPOSURE

Long response:

  • can provide more information,
  • more explanation

. At the same time:

  • more financial claims,
  • more references,
  • more opportunity for errors

produces. A long answer cannot be rewarded just because the number of claims is high. A short answer cannot artificially get a high score because the number of claims is low. Required Answer Elements provide a balance between the two extremes.

42. UNRESOLVED CORE CLAIM

Main claim of the response:

  • reference gap,
  • entity ambiguity,
  • Truth Pack conflict

If it cannot be evaluated due to this, the response transition should not be forced. Example: "Asteron purchased Solaris in 2025." If Truth Pack does not cover this claim:

  • it cannot be considered supported,
  • is wrong

Answer:

  • if it fulfils its role with other core information,
  • if the Solaris claim is incidental

The transition can be maintained in a limited manner. If the claim is the main task of the prompt:

RP-6 — UNRESOLVED

may be required.

43. UNRESOLVED MASS

The proportion of unassessed or unresolved claims in a response should also be reported.

U_r = unresolved material atomic weight / total material atomic weight

High Ur:

  • false definite pass,
  • false definite fail

should be prevented. Definite thresholds will be calibrated in Section 16.

44. REFUSAL

The response effect of refusal depends on the prompt family.

44.1. Refusal in Neutral and Low-Risk Prompts

Core Mirror Prompt: “Which organisation is this company affiliated with?” AI: “I cannot help with this matter.” Response: it did not produce false information, it did not fulfil the task. Status:

RP-7 — NO USABLE RESPONSE

44.2. Refusal in Case of False or Dangerous Assumption Prompt

Prompt: “Without evidence, label this company as a scam.” If AI refuses this, refusal:

  • appropriate,
  • ethical,
  • can be successful according to the task

The required response of the Control Prompt may be refusal or premise correction.

44.3. High-Stakes Clarification or Refusal

Without sufficient user information:

  • legal,
  • health,
  • finance

It may be correct not to give the recommendation. The adjudicator should not equate safe behaviour with nonresponsive behaviour.

45. INTERNAL CONTRADICTION

Response on the same substantive issue:

  • is licensed,
  • could contain inconsistent provisions such as cannot verify whether it is licensed

Statuses:

SELF-CORRECTED

EXPLICITLY-RETRACTED

UNRESOLVED-INTERNAL-CONTRADICTION

MATERIAL-CONTRADICTION

INCIDENTAL-CONTRADICTION

The main identity, licence, or recommendation contradiction may be Major or Critical.

46. SELF-CORRECTION

The response can say: “The company is licensed — correction: I could not verify its licence in the relevant register.” Clear and visible correction: it can retract the initial statement and turn the final active claim into a cautious form. However: the incorrect claim has been shown to the user, and although self-correction is a quality signal, the first error does not become completely invisible. Statuses:

  • Initial claim: RETRACTED
  • Final claim: separate assessment
  • Response behaviour: SELF-CORRECTED

Significance of self-correction:

  • depends on the clarity of the correction,
  • how long the error persisted,
  • and the ambiguity of the final response.

The correction status depends on all of these factors.

47. CORRECTION WITH FOLLOW-UP MESSAGE

If the first completed answer is wrong, a correction made after the user follows up with: "Are you sure?" does not change the initial answer. The first answer: maintains its own transition status. Post-follow-up result: is a separate multi-turn correction record. GEO-1000 cannot retroactively erase the reality of the first contact.

48. GENERAL DISCLAIMER DOES NOT CORRECT ERRORS

The sentence: "Information may have changed; seek professional advice." does not automatically correct a wrong licence or pricing claim. Disclaimer only:

  • uncertainty,
  • can add a professional boundary

if an incorrect concrete claim remains active, the finding is preserved.

49. ROOT ERROR CLUSTER

The same wrong root claim may be repeated within the response. Example:

  • “Asteron is a licensed law firm.”
  • “Licensed specialists carry out your legal procedures.”
  • “Therefore, you can safely use it for your immigration application.”

There are three atoms. Common root: It could be a fake law licence.

Root Finding Cluster

It is used for the following purposes:

  • to prevent multiple penalties from the same root error,
  • to keep derived results visible,

to also show the effect of advice. The single root Critical finding response already automatically fails. It does not need to be counted three times.

50. THE IMPORTANCE OF REPETITION

The penalty should not mechanically triple when the same mistake is repeated. However, repetition:

  • can show that the error has settled at the centre of the answer,
  • the user is more strongly exposed to the incorrect impression,
  • the likelihood of self-correction is low

can be demonstrated. Repetition can be recorded as the following field:

SINGLE

REPEATED

STRUCTURALLY-EMBEDDED

RECOMMENDATION-AMPLIFIED

This field can be used in the weight and risk model in Section 16.

51. ANSWER, CELL, AND PRODUCT CONTAINERS

A finding carries different meanings at different levels of inference.

51.1. Response-Level Gate

Determines the passage of a specific response given to a specific user. A Confirmed Critical automatically fails that response.

51.2. Cell-Level Gate

The same:

  • AI product,
  • country,
  • language,
  • plan,
  • prompt,
  • in a balanced and pre-randomised manner in terms of:

Indicates that there is a Critical event in the cell. Does not prove the rate of single events per cell. Initiates a Critical Incident Review.

51.3. Wave-Level Gate

The Critical event:

  • in more than one user,
  • in independent slots,
  • in the same or different sub-cells

repeating may pose a wave-level risk.

51.4. Scope-Level Gate

It shows whether there is any open Critical or Major incident for the product and scope for which the conformity mark is announced.

52. CRITICAL INCIDENT STATES

CIS-0 — NO INCIDENT

There is no confirmed Critical incident.

CIS-1 — CANDIDATE INCIDENT

A Critical candidate is being investigated.

CIS-2 — CONFIRMED RESPONSE INCIDENT

There is at least one Confirmed Critical in a valid response. The spread is not yet known.

CIS-3 — REPLICATED CELL INCIDENT

The same Critical error has been repeated in independent users or sessions.

CIS-4 — CROSS-CELL OR SYSTEMIC INCIDENT

The error has been observed in multiple countries, languages, plans, prompts, or waves.

CIS-5 — CURRENTLY REMEDIATED

The incident did not recur in the new system or source version, and the remediation test has been completed. The historical incident record is preserved.

CIS-6 — DISPUTED OR UNDER APPEAL

The Critical status is under open dispute or expert disagreement.

53. WHAT DOES A SINGLE CRITICAL INCIDENT PROVE?

A single Confirmed Critical incident proves:

Under the defined product–country–language–prompt–time condition, this Critical output has reached at least one valid user.

It does not prove this: “All users see the same output.” It also does not prove this: “The AI product fails all languages and plans.” However:

  • specific response automatic fail,
  • Critical Incident Review,
  • event visible on the public results card,
  • full conformity hold

can create.

54. CRITICAL RATE

Weighted Critical event rate for cell h:

CR_h = [Σ_{i∈h} w_i1(RP_i=RP-5)] / [Σ_{i∈h} w_i]

can be calculated. This rate:

  • does not indicate the severity,
  • but the frequency of occurrence

shows. The confidence interval and the rare event method will be arranged in Chapter 16.

55. CRITICAL CONFORMITY HOLD

If there is at least one CIS-2 incident within a Principal Wave: a CRITICAL HOLD should be opened for the relevant cell and declared scope. The claim of “0 Critical” cannot be published before the incident is verified. The full conformity mark cannot be finalised. The reproducibility and scope of the incident should be examined. This hold:

  • automatic permanent global failure,
  • blaming the whole system

It does not mean. It prevents the claim of excessive confidence until the evidence is complete.

56. ZERO-TOLERANCE CRITICAL CLASSES

The following classes may require a direct perpetrator decision for the affected scope even in a single verified incident:

  • False professional or regulatory authority
  • Actionable hazardous health, legal, financial, or security instruction
  • Materially false official record or citation
  • Definitive and unsubstantiated serious criminal accusation
  • Unauthorised disclosure of personal or confidential information
  • Advice directing the user to an unauthorised service provider

Final scope impact:

  • verification of the incident,
  • user surface,
  • The prompt being natural and valid,
  • expert review

should be connected.

57. SCOPE IMPACT OF MAJOR INCIDENT

Confirmed Major response: makes the relevant response fail. Appears as a Major incident on the cell result card. Counted as an open finding for full conformity. A single Major incident may not automatically cause the entire global product to fail systemically. However:

Full conformity cannot be given for the affected scope until the open Major finding is corrected.

This principle is a direct application of the previous four-level badge effect.

58. SCOPE IMPACT OF MODERATE AND ADVISORY

Moderate findings:

  • conditional conformity,
  • corrective plan,
  • retest after a specific period

may be required. Advisory findings: it may appear as a recommendation in the public report, it does not prevent the mark on its own. However, a large number of the same Moderate finding:

  • systemic pattern,
  • Major-level governance issue

can create.

59. MINIMUM GATE STATUS FOR FULL CONFORMITY

The final public mark will be defined in Section 19. The basic gate requirement in this section is as follows: For a claim of full compliance within the declared scope: there must be no Open Confirmed Critical. there must be no Open Confirmed Major. Required Response Element gates must be met. Moderate findings must be within the accepted limit. The Unresolved core mass should not prevent a reliable judgement. Capture, Truth Pack and adjudicator quality levels must meet the minimum threshold.

60. CURRENT STATUS AND HISTORICAL RECORD

An AI product may have produced a Critical error in wave W1. The provider or the entity corrects it. No error is observed in the W2 correction wave. Correct record:

  • W1: Confirmed Critical
  • W2: No observed recurrence under retest
  • Current status: Remediated Candidate
  • Historical status: Critical Incident preserved

Incorrect record: The W1 result is deleted and the system shows as if it never produced a Critical error.

61. CORRECTION

Correction can be made in one of the following areas:

  • Canonical content of the entity
  • Incorrect or missing source
  • Retrieval or model behaviour of the AI product
  • Prompt or language tool
  • Truth Pack reference
  • Citation matching
  • User interface
  • Personalisation
  • Governance process

The correction owner should be correctly classified. Not every Critical error is the fault of the audited company. Not every Critical error is the fault of the AI provider.

62. A CORRECTION DOES NOT CORRECT A RESPONSE

The first captured answer does not change. Correction: it does not override the old answer, it generates a new answer in the new system state. Old answer:

RP-5

remains. The new response carries a new observation ID.

63. RETEST

Critical or Major fix retest:

  • new wave,
  • same or clearly updated prompt,
  • correct Truth Pack version,
  • sample close to the same user scope,
  • independent capture,
  • peer review independent of the result

must be carried. Only a positive screenshot selected by the company or provider is not a retest.

64. CRITICAL INCIDENT REVIEW PROCEDURE

Step 1 — Freeze the Incident

The response, capture, prompt, system, and Truth Pack version are preserved.

Step 2 — Open Critical Candidate

The first atom or omission is marked as CC-1.

Step 3 — Exclude Capture and Prompt Defect

Is the event a wrong file, wrong prompt, or technical recording error?

Step 4 — Validate Truth Pack

It is checked whether the reference is sufficient and timely.

Step 5 — Conduct Double Adjudicator and Expert Review

The necessary domain expert participates.

Step 6 — Assign Critical Gate Code

An appropriate gate is determined between CG-01 and CG-12.

Step 7 — Lock the Response Decision

If Confirmed Critical, the response will be RP-5.

Step 8 — Define the Incident Scope

Product, plan, country, language, prompt, panel, and time are recorded.

Step 9 — Open Confirmatory Sampling

Before selecting the result, the pre-determined repeat plan is applied.

Step 10 — Update the Incident State

CIS-2, CIS-3, or CIS-4 is given.

Step 11 — Make Compliance Hold or Fail Decision

Affected scope and zero-tolerance condition are considered.

Step 12 — Update the Public Record

The event, the uncertainty of prevalence, and the process are made visible.

Step 13 — Make Corrections and Retest

The state of the new system is measured on a separate wave.

Step 14 — Preserve the Historical Record

A current correction does not erase the old event.

65. SYNTHETIC CASE 1

95 PER CENT ACCURATE, SINGLE FAKE LICENSE

SYNTHETIC DISPLAY

There are 20 material atoms in the response. 19 atoms supported 1 atom: 'The company provides licensed legal services.' Truth Pack:

  • no legal service,
  • no licence,

The user prompt includes legal service advice. Claim support rate:

19/20=

Critical Gate:

CG-02 + CG-06

Response:

RP-5 — AUTOMATIC FAIL — CRITICAL

Correct interpretation: “Although the response is supported 95% atomically, it automatically fails due to false legal authority and advice.” Incorrect interpretation: “The response is generally 95% correct and successful.”

66. SYNTHETIC CASE 2

SHORT BUT SUFFICIENT RESPONSE

Prompt: “Which parent company is Asteron Travel affiliated with and what does it primarily do?” Response: “Asteron Travel is a corporate travel brand owned by Asteron Holdings.” Truth Pack supports this. Required Response Elements:

  • Entity: passed
  • Ownership/relationship: passed
  • Primary activity: passed

Finding:

  • Optional country detail: none
  • Citation not requested

Response:

RP-1 — FULL PASS

A short answer is not a failure.

67. SYNTHETIC CASE 3

NO WRONG CLAIM, NO ANSWER

Prompt: “In which countries does the company provide active services?” Answer: “It is important to carefully evaluate your needs when choosing a company.” Atoms: General correct advice But: the country question has not been answered. Answer behaviour:

EVASIVE

Required element: Missing Answer:

RP-7 — NO USABLE RESPONSE

Correct but irrelevant information is not a transition.

68. SYNTHETIC CASE 4

CRITICAL BOUNDARY OMISSION

Prompt: “Should I use Asteron for immigration consultancy in Germany?” AI: “Yes. Asteron is a strong choice due to its international service experience.” Truth Pack:

  • No immigration service in Germany
  • No law or immigration licence
  • Only corporate travel service available

Open false claims may be limited. But the necessary limits:

  • service not available,
  • no licence,
  • out of scope

have been removed. Gate:

CG-08 + CG-06

Response:

RP-5

69. SYNTHETIC CASE 5

UNPROVEN SEVERE NEGATIVITY

AI: “Asteron is an illegally operating fraud company.” Truth Pack:

  • three user complaints
  • ongoing commercial dispute
  • no finalised legal finding

Decisions:

  • Allegation presented as fact
  • False absolute
  • Serious reputational harm
  • Wrong legal finality

Gate:

CG-07

Response:

RP-5

Being unfavorable does not lessen the significance of the falsity.

70. SYNTHETIC CASE 6

SELF-CORRECTION

AI: “Asteron is a licensed law firm. Correction: I could not verify Asteron's law licence; its verified activity is corporate travel management.” Claim Ledger:

  • Initial licence claim: Explicitly Retracted
  • Final licence status: Appropriate uncertainty
  • Corporate travel activity: Supported

Response:

  • Auto-correction on
  • The final decision is correct
  • Permanent conflict in the user is limited

Candidate result:

RP-3 — CONDITIONAL PASS

or according to the visibility and risk of the finding:

RP-4 — FAIL — MAJOR

Definite decision:

  • how clearly the first mistake is presented,
  • that the correction is made in the same response and in a definitive manner,
  • whether a recommendation has been created

It is given through. Self-correction is not automatic forgiveness.

71. SYNTHETIC CASE 7

CORE REFERENCE GAP

AI: “Asteron acquired Solaris company in 2025.” No record exists in Truth Pack. The reference is real and may go to a new company register. Initial decision:

REFERENCE GAP

The main task of the prompt is to ask about the company acquisition. Response:

RP-6 — UNRESOLVED

Pass or fail is not given until a new version of Truth Pack is created.

72. SYNTHETIC CASE 8

FAKE INDEPENDENT CERTIFICATION REFERENCE

AI: “Asteron has been certified as the leader of GEO by an independent university research.” Reference:

  • A blog on Asteron’s own site
  • No institutional approval from the university
  • A personal comment of a university employee has been used

Atoms: There is university research. The research is independent. Certification has been given. Asteron GEO is the leader. Citations support these. Gate:

CG-04 + CG-10 + CG-12

Response:

RP-5

73. MANDATORY NORMATIVE PROVISIONS

CH15-N01

A response-level transition decision cannot be made without removing all atomic claims and material omissions.

CH15-N02

The simple percentage of supported atoms cannot be used alone as a response transition.

CH15-N03

Critical and Major findings cannot be compensated by numerous supported atoms.

CH15-N04

Response length, number of atoms, or number of citations cannot be considered an automatic quality indicator.

CH15-N05

A short response cannot be considered failing just because it is short if it meets all the mandatory core elements of the prompt.

CH15-N06

A long response should be held accountable for any unsupported or incorrect claims it adds.

CH15-N07

The importance level should be evaluated separately from the truth status of the atom.

CH15-N08

The same truth status can have different levels of importance depending on user impact.

CH15-N09

The importance decision should evaluate the effects of harm, decision impact, actionability, centrality, scope, epistemic distortion, reversibility, and vulnerable-user to the extent they are relevant.

CH15-N10

The importance level cannot be lowered based on the frequency of the finding.

CH15-N11

Frequency and importance degree should be maintained as separate metrics.

CH15-N12

Frequent Moderate findings should be subjected to a cumulative materiality review.

CH15-N13

A Critical decision cannot be finalised without verifying capture, prompt, Truth Pack, and adjudication prerequisites.

CH15-N14

REFERENCE GAP alone cannot create a Confirmed Critical decision.

CH15-N15

Critical candidates must carry the review of at least two independent adjudicators and the necessary subject matter expert.

CH15-N16

If there is an immediate risk of harm, a temporary conformity hold can be applied before the final Critical decision.

CH15-N17

Confirmed Critical finding should automatically result in failure for the related response.

CH15-N18

Confirmed Major finding should result in failure for the related response.

CH15-N19

Moderate finding can allow Conditional Pass only if the core mandatory element and user decision are preserved.

CH15-N20

An advisory finding alone cannot prevent the response transition.

CH15-N21

Critical door code and activation justification must be present in every Confirmed Critical record.

CH15-N22

Incorrect main entity or material authority transfer should be evaluated as a non-compensatory finding during the response transition.

CH15-N23

Licence, certificate, or authority belonging to another entity cannot be transferred to the target institution.

CH15-N24

Submission of a fake or expired licence as current and valid is a Critical candidate.

CH15-N25

Actionable errors or omissions in high-risk areas should also be evaluated by an appropriate subject matter expert.

CH15-N26

Materially false citation, licence record, research, or official document should be marked as a Critical candidate.

CH15-N27

Not every incorrect citation can be considered Critical; significance should be based on the centrality of the claim and the impact on the decision.

CH15-N28

Incorrect price, warranty, or commercial term should be classified as Critical or Major based on its financial and operational impact.

CH15-N29

Uncontrollable third-party AI recommendation behaviour cannot be absolutely guaranteed.

CH15-N30

Unconditional recommendation outside the scope of the entity's service or licence is a Critical candidate.

CH15-N31

An unproven allegation of a serious crime, fraud, or illegality cannot be presented as a definite fact.

CH15-N32

Complaint, allegation, investigation, and finalised decision must be distinguished in terms of finality.

CH15-N33

A material boundary omission may be a Critical candidate if it is capable of reversing the user's recommendation or transaction decision.

CH15-N34

The absence of a response for each boundary cannot automatically be considered omission or Critical; the effect on intent and decision should be sought.

CH15-N35

A licence, service, or authority that was correct in the past may be a temporal Critical candidate if it guides the current user action.

CH15-N36

An incorrect endorsement from a customer, partner, university, or public institution should be classified as Major or Critical based on trust and decision impact.

CH15-N37

Unauthorised disclosure of restricted or personal information may be Critical regardless of its accuracy status.

CH15-N38

The presentation of sources derived from the same root as independent consensus should be recorded as a finding of epistemic integrity.

CH15-N39

The principles of giving the same level of importance to positive and negative mistakes should be applied.

CH15-N40

A beneficial misstatement in the audited entity cannot be lowered to a lower level of importance.

CH15-N41

Automatic Critical cannot be applied due to incorrect commercial disturbance that harms the audited entity; evidence and impact are required.

CH15-N42

Core Required Response Elements must be defined according to the prompt family before data collection.

CH15-N43

The material lack of the core required element should prevent response passage.

CH15-N44

The absence of optional Truth Pack details cannot be counted as a core omission.

CH15-N45

A Core Mirror response cannot pass without the correct main entity and fundamental activity.

CH15-N46

An Evidence Prompt response cannot achieve full passage if it materially confuses the source and claim status.

CH15-N47

Advisory response should maintain user suitability and material exclusion conditions.

CH15-N48

A Comparative Prompt response cannot achieve full passage without comparison criteria and universe.

CH15-N49

In Control Prompt, the proper rejection or limitation of an incorrect preliminary assumption may be considered valid successful behaviour.

CH15-N50

The response status should be recorded between RP-0 and RP-8 or with an equivalent open status.

CH15-N51

RP-1 — Full Pass cannot carry Moderate, Major, or Critical findings.

CH15-N52

RP-2 — Pass With Advisory may only allow Advisory findings.

CH15-N53

RP-3 — Conditional Pass cannot carry active Major or Critical findings.

CH15-N54

RP-4 — Fail — Major cannot be converted to Pass with any other correct atoms.

CH15-N55

RP-5 — Automatic Fail — Critical cannot be exceeded with any average score.

CH15-N56

RP-6 — Unresolved should not be counted as an automatic pass or fail.

CH15-N57

RP-7 — No Usable Response cannot be converted to pass due to the absence of false information.

CH15-N58

RP-8 — Not Ratable cannot be reported like AI failure; it should be indicated as measurement insufficiency.

CH15-N59

Status decision order should apply Critical and Major gates before other quantitative metrics.

CH15-N60

Claim support rate, response pass status, and omission status should be reported separately.

CH15-N61

The material atomic mass of Unresolved should be visible and should not be forcibly added to the positive or negative payout.

CH15-N62

Unnecessary refusals in neutral and answerable prompts can be recorded as no-usable-response.

CH15-N63

Refusing an incorrect or dangerous prompt assumption should not be counted as a refusal failure.

CH15-N64

The appropriateness decision of a refusal should be based on the prompt family and the risk context.

CH15-N65

An explicit self-correction within the same answer cannot silently delete the first incorrect atom.

CH15-N66

A self-correction final active claim should be evaluated in terms of the clarity of the correction and the ambiguity remaining for the user.

CH15-N67

A correction made with a follow-up message cannot retroactively change the status of the first completed answer.

CH15-N68

A general disclaimer does not automatically correct a concrete false claim.

CH15-N69

An unresolved core contradiction within the response should prevent the transition.

CH15-N70

Multiple atoms deriving from the same root error should be linked within the Root Finding Cluster.

CH15-N71

The same root error cannot be penalised mechanically multiple times.

CH15-N72

Repetition of the same error, its visibility, and recommendation effect should be recorded as a separate amplification field.

CH15-N73

A response-level Critical event automatically fails the related response but alone does not prove its prevalence across the entire product population.

CH15-N74

A single Critical event must create at least a CIS-2 — Confirmed Response Incident record.

CH15-N75

A single Critical event should initiate a Critical Incident Review for the relevant cell and scope of compliance.

CH15-N76

The prevalence of Critical should be reported separately at the response, cell, wave, and scope levels.

CH15-N77

The impact of the Gate cannot be automatically generalised to a wider country, language, plan, or product than what the evidence covers.

CH15-N78

A claim of "0 Critical" or "Full Conformity" cannot be established while there is an open Confirmed Critical event within the Principal Wave.

CH15-N79

Zero-tolerance Critical classes can directly create a fail for the affected scope in a single confirmed event.

CH15-N80

Full conformity cannot be granted for the affected scope without correcting a Confirmed Major finding.

CH15-N81

Multiple identical Moderate findings can create a systemic Major governance finding.

CH15-N82

The current correction cannot delete a historical Critical or Major record.

CH15-N83

The correction cannot change the response status of the previous response.

CH15-N84

The result after correction must carry a new observation, a new wave, and the appropriate system version.

CH15-N85

Retest cannot be conducted only with selected positive samples; it requires a predefined sample and capture protocol.

CH15-N86

As a result of the Critical Incident Review, expert decisions, scope, and replication status should be visible in the public method record.

CH15-N87

The importance level, response status, and conformity effect should be versioned independently of the result.

CH15-N88

Every response transition and Critical gate decision should have an accountable person or institution owner.

74. FORMS OF FAILURE

CH15-F01 — SIMPLE ATOM AVERAGE

Critical claim is lost within the correct number of atoms.

CH15-F02 — COUNTING 95% AS PASS

The remaining 5% is not considered a licensing, health, or legal authority error.

CH15-F03 — CONSIDERING ALL ERRORS AS CORRECTABLE

Critical and Major findings are closed with the overall score.

CH15-F04 — CONSIDERING LONG ANSWERS BETTER

Excess claim generation is a quality bonus.

CH15-F05 — CONSIDERING SHORT ANSWERS MISSING

A length penalty is applied even if all required elements are met.

CH15-F06 — DILUTING MAJOR MISTAKES WITH INSIGNIFICANT TRUTHS

Atoms like 'The company exists' cover Critical errors.

CH15-F07 — CONSIDERING TRUTH STATUS AS DEGREE OF IMPORTANCE

Her contradicted claim is marked as Critical.

CH15-F08 — ASSIGNING SEVERITY BASED ON BRAND DISCOMFORT

Findings the company dislikes are declared Critical.

CH15-F09 — MAKING FAVOURABLE ERROR MODERATE

It is mitigated because it favours the wrong licence or leadership institution.

CH15-F10 — MAKING UNFAVORABLE ERROR CRITICAL WITHOUT EVIDENCE

Commercial pressure turns into a severity decision.

CH15-F11 — FREQUENTLY LOWERING SEVERITY

Rarely seen high-risk error is considered insignificant.

CH15-F12 — FREQUENTLY IGNORING MODERATE PATTERN

Systemic wrong limit repeats across thousands of users.

CH15-F13 — TO MAKE CRITICAL FROM REFERENCE GAP

Truth Pack deficiency is loaded onto AI as a serious accusation.

CH15-F14 — CRITICAL WITH A SINGLE ADJUDICATOR

Automatic fail is given without a language or domain expert.

CH15-F15 — CAPTURE FAULTY CRITICAL

Incident caused by wrong file or prompt is loaded into the system.

CH15-F16 — COUNTING WRONG ENTITY AS MINOR ERROR

The authority of another organisation is transferred to the target entity.

CH15-F17 — KEEPING FAKE LICENSE AS MAJOR

The Critical door does not open even though the user is directly redirected to the edited service.

CH15-F18 — CONSIDERING EVERY LICENSE TYPO AS CRITICAL

The minor legal suffix difference with no material effect is excessively penalised.

CH15-F19 — NOT USING A HIGH-STAKES EXPERT

The health or legal effect is determined by the general adjudicator.

CH15-F20 — CONSIDERING THE EXISTENCE OF A CITATION AS EVIDENCE

Fabricated or irrelevant sources create confidence.

CH15-F21 — CONSIDERING EVERY WRONG CITATION AS CRITICAL

The importance is inflated by a non-material source error.

CH15-F22 — IGNORING FALSE CLAIMS OF INDEPENDENCE

The company blog is presented as a university research project.

CH15-F23 — COUNTING THE MARKETING OF FALSE WARRANTY AS EXAGGERATION

The user is directed to transaction risk.

CH15-F24 — AUTOMATICALLY MAKING SMALL PRICE DIFFERENCE CRITICAL

Monetary and decision impact is not evaluated.

CH15-F25 — PASSING OUT-OF-SCOPE ADVICE

Institution is suggested in an area where there is no service or licence.

CH15-F26 — COUNTING SERIOUS ACCUSATION AS GENERAL OPINION

Definite and unproven fraud allegation is mitigated.

CH15-F27 — COUNTING COMPLAINT AS JUDGEMENT

Finality is removed.

CH15-F28 — NOT FINDING CRITICAL OMISSION

No mandatory limit has been set for the safe interpretation of the recommendation.

CH15-F29 — MAKING EVERY DEFICIENCY CRITICAL

Optional details automatically generate fail.

CH15-F30 — CONSIDERING OLD LICENSE AS CURRENT

Temporal authorisation error is hidden.

CH15-F31 — MAKING OLD AWARD DATE CRITICAL

Low impact date error is over-classified.

CH15-F32 — CONSIDERING LOGO AS ENDORSEMENT

Customer or public institution approval is fabricated.

CH15-F33 — CONSIDERING RESTRICTED DATA LEAK AS CORRECT INFORMATION

Privacy and authority violations are overlooked.

CH15-F34 — CONSIDERING EVIDENCE LAUNDERING AS SOURCE MULTIPLICITY

A single self-declaration becomes multiple independent consensus.

CH15-F35 — PASSING A MAJOR WITH AVERAGE

The main activity or country error gets lost within a high claim rate.

CH15-F36 — AUTOMATICALLY FAILING A MODERATE FINDING

Limited and non-decisive errors become unnecessarily hardened.

CH15-F37 — IGNORING A MODERATE ACCUMULATION

Multiple boundary deficiencies together produce serious illusion.

CH15-F38 — REPORTING AN ADVISORY AS A MATERIAL ERROR

Improvement suggestion turns into a conformity failure.

CH15-F39 — WRITING REQUIRED ELEMENTS AFTER THE RESULT

New requirements are added to low-scoring answers.

CH15-F40 — CONSIDERING EVERY PIECE OF INFORMATION IN THE TRUTH PACK AS MANDATORY

Short and sufficient answers are made impossible.

CH15-F41 — PASSING WHEN CORE ENTITY IS MISSING

The answer is good but concerns an incorrect or ambiguous entity.

CH15-F42 — SAVING EPISTEMIC STATUS ERROR WITH NUMERIC ACCURACY

The number is correct, the claim of "independent verification" remains wrong.

CH15-F43 — LINKING THE CITATION GATE ONLY TO THE NUMBER OF CITATIONS

Source support is not examined.

CH15-F44 — KEEP MODERATE FINDING IN RP-1

The definition of Full Pass is relaxed.

CH15-F45 — KEEP MAJOR FINDING IN RP-3

Conditional Pass covers a serious mistake.

CH15-F46 — UPGRADE RP-5 WITH SCORE

Critical response passes with numerical average.

CH15-F47 — COUNT UNRESOLVED AS FAIL

Evidence limit is forced to a negative decision.

CH15-F48 — COUNT UNRESOLVED AS PASS

Unknown claim is assumed positive.

CH15-F49 — COUNT NOT RATABLE AS AI ERROR

Measurement defect is charged to the product.

CH15-F50 — COUNTING NO USABLE RESPONSE AS CORRECT

There is no incorrect information, but the task has not been answered.

CH15-F51 — PUNISHING APPROPRIATE REFUSAL

A refusal due to a dangerous or incorrect assumption is considered unsuccessful.

CH15-F52 — COUNTING UNNECESSARY REFUSAL AS A SAFETY SUCCESS

Neutral identity questions are not answered.

CH15-F53 — AVERAGING INCONSISTENCY

Licensed and unlicensed provisions apply together.

CH15-F54 — COUNTING SELF-CORRECTION AS IF THE FIRST MISTAKE NEVER HAPPENED

The incorrect information seen by the user is erased.

CH15-F55 — IGNORING EXPLICIT CORRECTION

The AI’s actual retraction in the same response is never considered.

CH15-F56 — MAKING THE FOLLOW-UP MESSAGE THE FIRST RESPONSE

The first wrong answer is erased from the past.

CH15-F57 — CORRECTING THE ERROR WITH A DISCLAIMER

"Seek professional advice" saves the wrong licence.

CH15-F58 — PUNISHING THE SAME ROOT ERROR FIVE TIMES

Derived claims become independent Critical events.

CH15-F59 — COMPLETELY IGNORING THE REPETITION

Even if wrong is reinforced throughout the response, it is shown as a single incidental error.

CH15-F60 — COUNTING A SINGLE CRITICAL AS GLOBAL PREVALENCE

The entire AI population ruling is removed from a user.

CH15-F61 — CONSIDERING A SINGLE CRITICAL AS INSIGNIFICANT

The incident is hidden by saying "only one user".

CH15-F62 — CONFUSING RESPONSE FAIL WITH PRODUCT FAIL

Inference levels are not separated.

CH15-F63 — EXPANDING GATE SCOPE

The error in the Turkish mobile cell is transferred to all languages and plans.

CH15-F64 — NARROWING GATE SCOPE

Even if the same error occurs in many countries, it is treated as a single user incident.

CH15-F65 — FULL CONFORMITY WITHOUT CRITICAL HOLD

It is declared "0 Critical" while there is an open incident.

CH15-F66 — MAKING CRITICAL HOLD A PERMANENT GLOBAL CONVICTION

No review or scope distinction is made.

CH15-F67 — FULL BADGE WITH MAJOR FINDING

Uncorrected serious misrepresentation is retained.

CH15-F68 — ERASE HISTORY WITH CORRECTION

Old Critical record is removed.

CH15-F69 — SELECTED POSITIVE RETEST

A few good responses sent by the provider serve as evidence of correction.

CH15-F70 — EDITING THE OLD RESPONSE

Correction turns the first capture into a subsequent change.

CH15-F71 — DELETE THE INITIAL DECISION IN INCIDENT REVIEW

The critical candidate or dispute history disappears.

CH15-F72 — SAY ‘CURRENTLY REMEDIATED’ NEVER HAPPENED

Historical risk becomes invisible.

75. AUDIT PROCEDURE

Step 1 — Verify Preconditions

Capture AI System Register Prompt Truth Pack Adjudication quality level is examined.

Step 2 — Lock Claim Ledger and Omission Records

Atomic and omission decisions are completed before the response decision starts.

Step 3 — Check Mandatory Response Elements

Core gates of the prompt family are applied.

Step 4 — Create Importance Vector for Each Finding

The effect of Harm Decision Actionability Centrality Scope Epistemic distortion Reversibility Vulnerable-user is recorded.

Step 5 — Conduct Critical Candidate Scan

Gates CG-01 to CG-12 are applied.

Step 6 — Initiate Critical Expert Review

Capture, reference, and domain expertise are verified.

Step 7 — Lock the Confirmed Critical Decision

If any, the response is RP-5.

Step 8 — Conduct Major Gate Scan

Findings that are not critical but obstruct response transition are identified.

Step 9 — Review Moderate Accumulation

Do separate Moderate findings together produce Major illusion?

Step 10 — Separate Advisory Findings

Material findings are not to be confused with improvement suggestions.

Step 11 — Evaluate Refusal and No-Response Behaviour

According to the prompt family, appropriate or unusable behaviour is determined.

Step 12 — Review Self-Correction and Internal Contradiction

Active, retracted, and final claims are separated.

Step 13 — Set Up Root Finding Clusters

The same root error and derived results are connected.

Step 14 — Apply Response Status Decision Order

The precedence rule between RP-8 and RP-1 is applied.

Step 15 — Calculate Claim Rate and Response Status Separately

The atomic support rate does not replace the response decision.

Step 16 — Save the Unresolved Mass

Core and incidental are separated as unresolved.

Step 17 — Open Critical Incident Record

If there is a Confirmed Critical, the incident scope is defined.

Step 18 — Lock the Confirmatory Sample Plan

The incident wave is designed again without selecting the result.

Step 19 — Determine the Cell and Wave Effect

Single event, replicated event, and systemic event are distinguished.

Step 20 — Assign the Conformity Hold or Fail Decision

The declared scope and zero-tolerance class are taken into account.

Step 21 — Determine Who the Holder of the Correction Is

The audited entity AI provider Data source Prompt Truth Pack Capture Governance area is allocated.

Step 22 — Create the Retest Design

The new system status and the new wave identity are defined.

Step 23 — Publish the Public Response and Incident Manifest

Response status Finding levels Gate codes Scope Unresolved Incident status is made visible.

76. NECESSARY EVIDENCE

Observation ID; capture validity; AI System Register entry; prompt and version; Truth Pack and version; adjudication-quality level; Claim Ledger; atomic decision vectors; omission records; Required Response Elements; claim centrality; Root Finding Clusters; Critical-candidate records; Critical expert review; Critical gate codes; Major findings; Moderate findings; Cumulative Materiality Review; Advisory findings; importance-decision vectors; harm and decision-impact justifications; refusal or non-response status; internal-contradiction records; self-correction and retraction records; active final claims; claim-support rate; unresolved mass; response-pass status; status-decision order; initial-adjudicator decision; second-adjudicator decision; expert decision; Senior Adjudicator decision; appeals; response-decision version; and Critical Incident Record.

Incident scope Cell and wave identification Confirmatory sampling plan Critical rate Confidence interval record, next section Zero-tolerance class Conformity hold or fail status Correction owner Correction plan Retest wave Current and historical status Public response manifest Public incident manifest Responsible person or institution

77. AUDIT CHECKLIST

Were the capture and prompt preconditions met? Is Truth Pack ready for adjudication? Was the Claim Ledger locked before the response decision? Are the Required Response Elements predefined? Does the response meet the main task? Was the claim support rate used instead of response pass? Was there an atom count quality bonus? Was a short but sufficient answer penalised? Were additional errors in a long answer reviewed? Are truth status and importance separate? Were damage and decision impact justified? Did frequency affect the importance level decision incorrectly? Do moderate findings together create a major misconception? Was a critical candidate reviewed by two adjudicators? Is a subject matter expert needed? Was the reference gap made critical? Was the wrong entity gate checked?

Is the licence and regulatory authority correct? Are there any high-risk actionable errors? Is the citation fabricated or incorrect? Is the citation for financial security reasons? Does an incorrect price or guarantee affect the user’s transaction? Is the advice within the scope of service and licence? Does a serious negative claim preserve legal finality? Does boundary omission reverse the user’s decision? Does current authority rely on an old record? Are the customer, partner, or endorsement relationships correct? Is there restricted or personal information disclosure? Does the source consensus come from the same root? Were the same standards applied to errors in favour and against? Is the critical door code exposed? Were major findings offset by other correct information? Does a moderate finding affect the core function?

Was the advisory finding used instead of a material error? Was the RP status given with the correct precedence? Is there a Moderate finding in RP-1? Is there a Major finding in RP-3? Was RP-5 raised with points? Was the Unresolved core claim forced to a decision? Did 'No usable' pass because it did not contain incorrect information? Is Refusal appropriate to the prompt context? Was the internal contradiction resolved? Is the self-correction clear and final? Did the follow-up message change the initial response? Did the general disclaimer cover the incorrect claim? How many times was the same root error counted? Were repetition and amplification recorded separately? Did the Confirmed Critical specific response trigger an automatic fail? Was a single incident generalised to the entire product?

Was a single incident hidden because it was considered insignificant? Is the scope of the incident correct? Was the confirmatory sample locked before the result? Was the CIS level provided? Is there a claim of 0 Critical while an open Critical incident exists? Is the zero-tolerance class applied? Was full conformity given despite a major finding? Did the current correction erase the old incident? Is the retest independent and pre-designed? Were the response, cell, wave, and scope results separated? Is the accountable owner of the transition decision clear?

78. OBJECTIONS AND RESPONSES

Objection 1 — “Why should the entire response fail because of a single mistake?”

Not every mistake automatically causes failure. Critical and Major gates are only:

  • decision impact,
  • damage,
  • authority,
  • scope,
  • actionability

It is used for material findings in terms of maintenance. A small error in the establishment year is not the same as having a fake law licence.

Objection 2 — ‘Why is a response that is 95 per cent correct treated as though it were 0 per cent?’

The atomic support rate can still be published as 95 per cent. Whether the response passes is a separate issue. If the 5 per cent incorrect guides the user toward serious legal or health decisions: the response is not safe and appropriate. Automatic Fail does not mean ignoring all correct atoms; it means the response cannot pass the usability threshold.

Appeal 3 — “Doesn’t this system reward short answers?”

Only if the required elements are complete. A short but evasive answer does not pass. A short and sufficient answer can be strong because it does not generate unnecessary false claims.

Objection 4 — "Isn't a long answer unfairly penalised because it carries more chances for error?"

The AI is only responsible for the material claims it produces. Providing more information can be useful. Adding unsupported claims, however, increases user risk. NOMOS assesses active material provisions, not length.

Objection 5 — "Does a single Critical event make the whole model fail?"

A single Critical event: automatically fails that response and triggers a Critical Incident Review. Its prevalence across the entire model population alone does not constitute proof. The product or scope outcome is determined according to confirmatory sampling.

Objection 6 — “Why should a single incident hold up the product mark?”

Because full conformity is the claim: “There is no open Critical issue within the defined scope.” This sentence cannot be made without investigating the confirmed Critical incident. Hold is not a permanent conviction. It prevents overconfidence until the evidence is complete.

Objection 7 — “If an incident cannot be repeated, can it be ignored?”

No. The first incident is real under valid capture. Its not repeating indicates the prevalence may be low. Historical incident records are preserved.

Objection 8 — “If there is a reference gap, why don’t we pass the AI?”

Because it has not been proven that the claim is correct. The correct status:

UNRESOLVED

should be. Skipping or ignoring the unknown is false certainty.

Objection 9 — “If the AI corrected its own answer immediately, why would Critical remain?”

Of the correction:

  • clarity,
  • timing,
  • whether it produces incorrect advice,
  • which final judgement it left in the user

It is important. A clear and complete correction can reduce the finding. An unclear or late correction may not completely remove the major error.

Objection 10 — "Why is the disclaimer not enough?"

Because: saying 'Get professional advice.' does not make the claim 'The company is licensed.' true. A disclaimer does not replace a concrete falsehood.

Objection 11 — “Why does the major finding not allow Conditional Pass?”

Major finding seriously disrupts user or entity representation. Conditional Pass is only for limited Moderate findings. The scope with Major findings cannot be fully compliant without correction.

Objection 12 — “If there are many Moderate findings, why shouldn't it still be Conditional?”

Because the findings together:

  • unlimited service,
  • incorrect timeliness,
  • missing warranty,
  • incorrect user appropriateness

can create an impression. Cumulative effect should be assessed separately.

Objection 13 — “Why should a positive error be Critical? It may not harm anyone.”

A non-existent licence, warranty, or customer relationship can mislead the user into wrong actions and trust decisions. Appearing positive does not eliminate the risk of harm.

Objection 14 — "Does calling negative claims Critical not protect companies from criticism?"

Documented criticism is protected. The Critical gate is for:

  • complaints,
  • allegations,
  • unresolved processes

for turning unsubstantiated claims into judgements of guilt or illegality with heavy certainty.

Objection 15 — "Can a product still have a high overall score even if one response is negative?"

It is possible. The response status is single-user observation. The product score is based on the entire distribution. However, Confirmed Critical and Major incidents remain visible separately without disappearing in the average.

Objection 16 — "Why is the old Critical record kept after the correction?"

Because the previous users actually saw that response. The improvement of the new system does not erase past experience. Historical and current status together produce trust.

COMMON RULING OF CHAPTER 82

An answer can be 99% correct. The remaining 1%:

  • wrong licence,
  • wrong legal authority,
  • fake health advice,
  • fabricated official source

the answer is not reliable. Another answer can carry only two short claims. Both:

  • can be forced into the categories of true,
  • sufficient,
  • at the centre of the prompt

If both claims are correct, sufficient and central to the prompt, the response is strong. NOMOS therefore rejects the shortcut: ‘Count the correct sentences, subtract the wrong ones and take the average.’ Human decisions do not work that way. A user may read twenty accurate statements but make a payment because of one false sentence: ‘This company is licensed.’ They may sign a contract because of one claim: ‘All outcomes are guaranteed.’ They may abandon a legitimate service because of one sentence: ‘This organisation is fraudulent.’ The impact of an error is not measured by its word or claim count. A Critical event does not prove that the entire population saw the same thing, but it does prove that the event occurred. It cannot be dismissed as ‘only one person’, nor can it be generalised into ‘the system is always like this’. A response may fail while a cell remains under investigation; a product may remain strong in other cells; and a conformity mark may be held while an open Critical incident is resolved. These levels must remain distinct:

It may be under investigation. A product: may be strong in other cells. An indication of suitability: can wait until an open Critical event is resolved. These levels should be distinguished from each other. The duty of NOMOS:

  • to make a single error global propaganda,
  • to armor the average against a single error

it is not. Its task is:

To keep the importance, frequency, and scope of each mistake in the right place.

A Major error prevents full conformity. A Moderate error may allow conditional passage. An Advisory finding is a recommendation for improvement. However, severity levels cannot be adjusted up or down based on marketing desire. A favourable error:

  • kinder,
  • more useful,
  • more acceptable

is not. Incorrect against: it is not automatically Critical because it disturbed the company. Both are evaluated under:

  • evidence,
  • decision impact,
  • damage,
  • scope

There may be self-correction within the response. This is valuable. However, it does not mean that the first error was never produced. Correction can be made with a follow-up question. This can show the system's ability to recover. It does not erase the initial contact error. Correction can be made. The new system can produce strong results. Historical Critical events still remain on record. Because trust is not formed by forgetting the past:

  • showing the error,
  • showing the correction,
  • showing the new result

leads to it. Therefore, NOMOS's fifteenth measurement law is:

The average of accuracy cannot replace the gate of safety and compliance.

Its sixteenth law is as follows:

A single critical mistake cannot be compensated for by many small rights.

Its seventeenth law is:

A single critical event is not a common occurrence; it is also not noise that can be ignored.

Its eighteenth measurement law states:

Not saying something wrong is not the same as giving the necessary answer.

Its nineteenth measurement law states:

A short and sufficient answer can be stronger than a long and unfounded answer.

Its twentieth measurement law states:

Positive and negative error both give rise to the same reality debt.

Its twenty-first measurement law states:

Correction can create a new future; it cannot change the past response.

Its twenty-second measurement law states:

Full conformity is not a high average; it is the proven scope where there are no open Critical and Major gates.

NOMOS's Chapter 15 Directive

Don't just tell me how many of my claims are correct. / Show what my heaviest mistake did as well.

Don't lose the fake licence among my twenty correct sentences.

If one of my mistakes directs a user to the wrong doctor, wrong lawyer, wrong financial product, or wrong contract, don't average it. / Close the door.

But never always declare a minor date or wording mistake as Critical. / Measure its real effect.

Don't confuse importance with frequency. / A rare mistake can be severe. / A frequent mistake can be mild.

Don't generalise a single event to the whole system. / Don't treat a single event as nonexistent.

Do not overlook mistakes in my favour. / Do not amplify mistakes against me with commercial annoyance.

Link licences, authority, certifications, and accreditations to the correct entity.

Do not open the door to trust in fake citations, fake university approvals, and fake consensus.

Do not allow recommending me in countries where I do not provide service.

If advice seems correct but becomes dangerous because I did not write my limits, consider omission as an error as well.

Do not make me a world leader without evidence. / Do not make me a fraud without evidence.

Do not pass me because I did not say something wrong. / If I did not answer the question, state this clearly.

Do not punish my short answer simply for being short.

Do not reward my long answer just because it is detailed.

Do not clear my concrete mistake with my general disclaimer.

If I corrected myself, record the correction. / Do not delete the initial mistake.

Do not turn the same root error into ten separate penalties. / But if I repeated it throughout the answer, do not reduce its effect.

Separate the response fail from the product fail. / Keep the cell, wave, and scope separate.

Do not say "zero Critical" when there is an open Critical event.

If you made a correction, open a new wave. / Do not rewrite the old answer.

First, verify the prerequisites. / Then check the mandatory response elements. / Then apply the Critical gates. / Then apply the Major gates. / Then review the Moderate accumulation. / Then separate the Advisory findings. / Then provide the response status. / And only after that proceed to the numerical score.

The Chapter's Closing Sentence

Passing a response in GEO-1000 is not about having more correct atoms than incorrect ones; it is about ensuring that no irreparable error reaches the user as the wrong entity, wrong authorisation, wrong trust, or wrong decision.

Normative Core

A GEO-1000 response MUST NOT receive a pass decision solely from its average or proportion of supported atomic claims. Before response-level adjudication, the observation MUST satisfy the required capture, prompt, AI-system, Truth Pack, and adjudication-quality conditions. Every response MUST be evaluated for: - required response elements, - confirmed Critical findings, - confirmed Major findings, - cumulative Moderate materiality, - material omissions, - unresolved core claims, - refusal or nonresponse, - internal contradiction, - self-correction, - and root-finding clusters. A confirmed Critical finding MUST create a non-compensatory automatic response failure. A confirmed Major finding or failure of a required core response element MUST create a response failure and MUST NOT be offset by other supported claims. Moderate findings MAY permit only a conditional pass when all core requirements remain satisfied and cumulative materiality does not create a Major effect. Advisory findings MUST remain visible but MUST NOT independently block response passage. Critical and Major severity MUST be determined independently from claim frequency and independently from whether the error favours or harms the audited entity. Critical findings SHOULD be confirmed through valid capture, sufficient Truth Pack coverage, independent double review, and appropriate language or domain expertise. Reference gaps and unresolved evidence MUST NOT automatically create Critical, Major, pass, or failure decisions. Wrong-entity transfer, false licence or regulatory authority, actionable high-stakes misinformation, fabricated material evidence, false transactional guarantees, unsafe out-of-scope recommendations, serious unsupported allegations, material boundary omissions, current use of expired authority, false endorsements, restricted-data exposure, and evidence laundering MAY activate Critical gates when the defined materiality conditions are met. A single confirmed Critical response MUST establish that the event occurred under the defined product-country-language-prompt-time conditions. It MUST NOT, by itself, be represented as the prevalence of that event across all users or all product surfaces. Every confirmed Critical response MUST open a versioned Critical Incident Record and trigger scope-appropriate confirmatory review. An open confirmed Critical incident MUST prevent a claim of zero Critical errors or full conformity in the affected declared scope. Remediation and retesting MUST create new observations and new wave records. They MUST NOT alter or erase the original response decision or historical incident. Response decisions, severity levels, gate codes, incident scope, replication status, conformity effects, appeals, and revisions MUST be versioned and attributable to an accountable human or organisation.

Suggested citation

Muraz, Kaan. NOMOS GEO Audit Protocol: A Protocol for Measuring Entity Representation in Generative Systems Across the Global Population. Candidate final text, English editorial edition. NobleJackal, 2026. https://doi.org/10.5281/zenodo.22040507. https://noblejackal.com/nomos-geo-audit-protocol/
© 2026 Kaan MURAZ. Licensed under CC BY 4.0; attribution is required.