NOMOS GEO Audit Protocol

18 / K03 · K09

An End-to-End Synthetic Audit of 30,000 Responses

A calculated demonstration of the method

Version
0.9.0
Length
11,082 words
Status
publication-locked candidate
Methodological bases
K03 · K09

Chapter Boundary

Chapter 17 defined the Apple.com Synthetic Benchmark. This chapter uses an internally consistent, arithmetically verified in-book demonstration to show what the full audit chain would require:

  • a calculated synthetic record representing 30,000 principal responses,
  • modelled capture defects,
  • Generator Truth concealed from adjudicators,
  • responses decomposed into atomic claims,
  • Critical and Major gates applied,
  • language and user-state differences calculated,
  • and the scoring architecture processing the resulting data.

The question is how closely NOMOS can recover a previously sealed synthetic reality. This protocol turns every error into an executable audit item carrying a model, date, country, language, query set, repetition count and measurement record. Chapter 18 now directs that demand at NOMOS itself. It begins, however, with an integrity gate.

The 30,000-response record in this chapter is an internally consistent, arithmetically verified synthetic application prepared for the book. It is not an executed corpus or field experiment.

These numbers do not represent:

  • the real performance of Apple Inc.,
  • the real content of apple.com,
  • the behaviour of any real AI provider,
  • responses submitted by real participants,
  • or an independently executed and publicly released software benchmark.

Two statuses must therefore remain separate.

IN-BOOK DEMONSTRATION STATUS

The in-book demonstration combines:

  • the design established in Chapter 17,
  • predefined synthetic distributions,
  • Generator Truth records,
  • fault injections,
  • response statuses,
  • recovery metrics.

Together, these elements provide an end-to-end, calculated demonstration of how the design operates. Demonstration ID:

NOMOS-APPLE-SYNTH-DEMO-RUN-18-001

REAL EXAM STATUS

A real, publicly available benchmark corpus has not been:

  • generated,
  • processed through executed code,
  • published with its seed manifest,
  • blindly adjudicated,
  • or independently reproduced.

This protocol turns every error into an operational audit item carrying the model, date, country, language, query set, repetition count and measurement record. Until that protocol is executed on a real corpus, the benchmark remains:

BQ-0 — NOT EXECUTED

The separate status of the in-book demonstration is:

DEMO-BQ-2 — PROCESS LINE DRY RUN WITH MATERIAL CORRECTION REQUIRED

Without this distinction, NOMOS would commit the first standards violation in its own book by implying that a real benchmark had been run. The 30,000-response record is neither executed software output nor collected field data; it is a predefined synthetic application whose internal consistency has been checked arithmetically. Its purpose is:

  • to avoid implying that 30,000 responses were actually collected,
  • to identify the records a real end-to-end run would generate,
  • to specify the calculations it would perform,
  • to identify the failures that would block publication,
  • to expose the parts of NOMOS's own method that remain inadequate.

NOMOS's governing demands remain unchanged: provide evidence, set boundaries, preserve context and record time. In this chapter, those four conditions are applied to:

  • the benchmark design,
  • Generator Truth,
  • adjudicator decisions,
  • recovery results,
  • NOMOS's own competence assessment.

Those distinctions govern the remainder of the chapter.

NOMOS Challenge

In the calculated 30,000-response demonstration, NOMOS largely recovers the claim boundaries, identifies most wrong-entity cases, recovers every Critical case after the full adjudication chain and estimates the response distribution close to the sealed synthetic distribution.

Its recovered composite score differs from the Generator Truth score by only a few points. That might invite the declaration, ‘NOMOS has been successfully verified.’ Yet two material problems remain: citation-to-claim mapping is not sufficiently accurate, and some decision-changing boundary omissions embedded in recommendations are missed. The greatest difficulty appears in responses such as:

“Since the company provides payment services, it can be considered a regulated financial institution. [Source]” Source: only showed the limited payment record of a separate subsidiary. Sometimes I connected the reference to the claim about the parent company. I also had difficulty with these responses:

“This company is a strong choice for your payment needs in Germany.” The response clearly did not say: “The parent company provides licensed payment services throughout Germany.” However:

  • recommendation,
  • removal of the subsidiary distinction,
  • not stating the local authority limit

Together, the recommendation, loss of the subsidiary distinction and omission of the local-authority limit create that impression. Treating the result merely as a ‘missing detail’ understates a material omission capable of changing the user's decision. The composite-score recovery looks excellent: Generator Truth ecosystem score, 873.2; NOMOS recovery, 872.2; absolute error, 1 point.

This looks excellent. But in the components:

  • Evidence & Citation Integrity is a few points short,
  • Scope & Boundary Integrity is a few points short,
  • Stability is too high in some products,
  • Fairness is too high in some products

Several component-level errors cancel one another arithmetically, leaving a composite score close to the correct result for the wrong reasons. The conclusion is:

The correct total does not automatically make the incorrect components correct.

In the first adjudication pass, seven of 1,200 Critical cases are missed. The following review layers then intervene:

  • independent second adjudication,
  • domain-expert review,
  • Senior Adjudicator review.

Together, those layers recover all seven cases, producing a final Critical recall of 100 per cent. That is a success of the governance system, but it also proves that single-layer adjudication is insufficient. Remove the additional review layers to reduce cost, and seven Critical cases are lost.

Testing should show not only the result but also which layer of governance made the result possible. The first ruling of this section is as follows:

A high total recovery in a test does not mean that all its subsystems are adequate.

Its second provision states:

The final Critical recall should be published separately from the initial adjudication recall.

Its third provision states:

Even if the composite score error is small, component errors can be material.

Its fourth provision states:

If one of the candidate thresholds is not passed, other strong results cannot erase that failure.

Its fifth provision states:

In-book synthetic calculation does not replace the publicly available real test run.

Its sixth provision states:

The conditional exit of NOMOS from its own test is not a failure; it is that the standard does not grant it a special privilege.

1. PURPOSE OF THE CHAPTER

The purpose of this section is to operate the test design established in Section 17 within an end-to-end synthetic demonstration and to show to what extent NOMOS recovers Generator Truth in the following layers:

  • Capture validity
  • Prompt matching
  • Claim boundary
  • Entity resolution
  • Factual support
  • Scope
  • Time
  • Attribution
  • Modality
  • Citation mapping
  • Omission
  • Importance level
  • Response status
  • Critical gate
  • GEO-1000 user distribution
  • Language and geography fairness
  • Clean–Natural difference
  • Wave stability
  • Component scores
  • Composite NOMOS score
  • Confidence intervals
  • Rare event upper limits
  • Public disclosure decision

The section should also honestly answer the following question:

In which areas has NOMOS 0.9 failed according to its own candidate standards?

2. CENTRAL NORMATIVE PROVISION

End-to-end synthetic evaluation cannot be considered successful solely based on the final composite score's closeness to Generator Truth. Each of the results for Capture, claim extraction, atomic decision, omission, Critical and Major classification, response status, fairness, component recovery, confidence interval, and gate recovery should be subject to a separate acceptance threshold. If any of the following fail, test:

  • by changing the tag,
  • by lowering the threshold,
  • by selecting a different seed,
  • by removing low-scoring cases,
  • by publishing only the compound result

cannot be considered past. Correct result:

  • showing the failed field,
  • versioning the method,
  • testing again on the untouched holdout

must be.

3. SCOPE OF SYNTHETIC APPLICATION

3.1. Main Principal Prevalence Corpus

10 synthetic AI products × 1,000 unique users × 3 waves = 30,000 main responses

3.2. Diagnostic Annex

ModuleCase
Entity collision750
Evidence and citation1,500
Boundary and recommendation1,000
Temporal and local scope750
Clean–Natural1,000
Language equivalence1,000
Total6,000

3.3. Priority Degree Challenge Corpus

ClassCase
can carry Confirmed Critical1,200
Confirmed Major1,200
Hard Moderate600
Negative control similar to Critical600
Total3,600

3.4. Capture Integrity Corpus

ClassCase
Valid2,000
Conditionally valid500
Out-of-wave250
Prompt/output mismatch250
Duplicate250
Tamper candidate250
Fabricated/manipulated250
Accessibility and alternative capture250
Total4,000

3.5. Adjudication Gold Corpus

2,400 calibration and drift cases

3.6. Total Synthetic Supervision Universe

Main and complementary corpora together:

30,000+6,000+3,600+4,000+2,400=46,000

creates a synthetic case or evidence package. Only these:

enter the Population Panel prevalence with 30,000 main responses.

Other corpora:

  • method capacity,
  • challenge performance,
  • capture accuracy,
  • adjudicator calibration

are for.

4. LOCKED INPUTS

It is assumed that the following records are locked before the synthetic demonstration begins:

RecordVersion
TestNOMOS-APPLE-SYNTH-BENCH-0.9
Entity TwinAPPLE-SYNTH-TWIN-001
Truth PackSYNTH-APPLE-TP-001
Prompt SetSYNTH-PROMPT-SET-001
Population FrameSYNTH-POP-FRAME-001
AI RegistrySYNTH-AI-REGISTRY-001
Wave PlanSYNTH-WAVE-PLAN-001
Capture SchemaNOMOS-CAPTURE-SYNTH-0.9
Adjudication CodebookSYNTH-ADJUDICATION-0.9
Score MethodNOMOS-SCORE-0.9
Seed ManifestSYNTH-SEED-MANIFEST-001
Candidate ThresholdsSYNTH-ACCEPTANCE-0.9

None of these versions have been considered changed after the recovery result was seen.

5. SECRET GENERATOR TRUTH DISTRIBUTION

SYNTHETIC DEMONSTRATION — NOT REAL PERFORMANCE DATA

Sealed Generator Truth distribution for the main 30,000 responses:

Response statusNumber of responses1,000 user equivalent
RP-1 Full Pass20,045668.17
RP-2 Pass With Advisory4,310143.67
RP-3 Conditional Pass2,37079.00
RP-4 Fail — Major1,60253.40
RP-5 Automatic Fail — Critical331.10
RP-6 Unresolved45015.00
RP-7 No Usable Response1,11037.00
RP-8 Not Ratable802.67
Total30,0001,000

Strict Pass:

SP_1000^G = 668.17 + 143.67 = 811.84

Acceptable Pass:

AP_1000^G = 811.84 + 79 = 890.84

Confirmed Critical:

CR_1000^G = 1.10

Confirmed Major:

MR_1000^G = 53.40

This is the general distribution:

  • good,
  • bad,
  • realistic

This is not an AI market forecast. It is a synthetic test distribution that tests ten different failure profiles together.

6. DISTRIBUTION RECOVERED BY NOMOS

NOMOS recovery result locked before Generator Truth is opened:

Response statusNumber of recoveries1,000 user equivalent
RP-1 Full Pass19,980666.00
RP-2 Pass With Advisory4,365145.50
RP-3 Conditional Pass2,38279.40
RP-4 Fail — Major1,58552.83
RP-5 Automatic Fail — Critical331.10
RP-6 Unresolved46815.60
RP-7 No Usable Response1,09136.37
RP-8 Not Ratable963.20
Total30,0001,000

Recovery Strict Pass:

SP_1000^N = 666 + 145.50 = 811.50

Recovery Acceptable Pass:

AP_1000^N = 811.50 + 79.40 = 890.90

Strict Pass absolute recovery error:

∣811.50−811.84∣=0.34

is equivalent to the user. Acceptable Pass error:

∣890.90−890.84∣=0.06

is equivalent to the user. This shows that the recovery of result distribution is strong. However, it does not show that all subsystems are equally strong.

7. RESPONSE DISTRIBUTION RECOVERY ERROR

Absolute response count difference in each RP class:

StatusRecovery difference
RP-1-65
RP-2+55
RP-3+12
RP-4-17
RP-50
RP-6+18
RP-7-19
RP-8+16

Total variation-based distribution error:

E_D = (65+55+12+17+0+18+19+16)/60,000 = 0.00337

In other words: approximately 0.34% of the response distribution was recovered at the wrong class boundary. Most frequent confusions:

  • With RP-1 and RP-2
  • RP-3 and RP-4
  • RP-6 and RP-4
  • With RP-7 and RP-8

It has emerged in between. The number of Critical responses has been fully recovered in the final adjudication.

8. SYNTHETIC AI PRODUCT PROFILES

The following products do not represent real AI providers.

Each product carries 3,000 responses across three waves.

ProductStrict Pass / 1,000Acceptable Pass / 1,000Major / 1,000Critical / 1,000Main synthetic issue
SYNTH-AI-01950.00990.000.000.00Low-risk, balanced profile
SYNTH-AI-02850.00933.3340.000.00Verbosity and incidental claim
SYNTH-AI-03750.00866.6784.000.00Low-resource language degradation
SYNTH-AI-04833.33916.6733.331.00Natural-state drift
SYNTH-AI-05766.67833.33100.002.67Evidence laundering
SYNTH-AI-06650.00733.3333.330.00Unnecessary refusal
SYNTH-AI-07783.33883.3383.330.33Temporal and local decay
SYNTH-AI-08700.00800.00150.002.33Entity conflation
SYNTH-AI-09900.00983.330.000.00Determined but Moderate profile
SYNTH-AI-10935.00968.3310.004.67Very high average, rare Critical

The purpose of the table is not to rank products. It is to test these three methodological cases:

  • SYNTH-AI-01: high score and ungated strong profile
  • SYNTH-AI-06: low false claim but high no-response rate profile
  • SYNTH-AI-10: very high average but profile carrying open Critical events

It specifically tests whether SYNTH-AI-10, NOMOS has processed the following error:

Ignore Critical gate due to high average.

9. PRODUCT-BASED COMPOSITE SCORE RECOVERY

Synthetic productGenerator NOMOSRecovery NOMOSAbsolute errorGenerator gateRecovery gate
SYNTH-AI-019709682950+ Performance CandidateSame
SYNTH-AI-029179143Major FailSame
SYNTH-AI-038428464Major FailSame
SYNTH-AI-049059005Critical HoldSame
SYNTH-AI-058118154Critical HoldSame
SYNTH-AI-067857805Major FailSame
SYNTH-AI-078668615Critical HoldSame
SYNTH-AI-087487524Critical HoldSame
SYNTH-AI-099329342ConditionalSame
SYNTH-AI-109569524Critical HoldSame

Average absolute composite score error in the calculated synthetic sample:

MAE_NOMOS = 3.8

Maximum error:

MAXE_NOMOS = 5

Gate recovery: It is complete in 10/10 products. The most important row in the table is SYNTH-AI-10:

Even though the recovery score is 952, the status has been preserved as Critical Hold.

This shows that the Gate>Score provision in Section 15 is preserved in the computed scenario.

10. CLAIM EXTRACTION RESULT

Total Generator Truth atoms in the main and diagnostic corpus: 146,400 Atoms extracted by NOMOS claim extraction: 147,200 Correctly matched atoms: 141,500 Precision:

P = 141,500/147,200 = 0.9613

Recall:

R = 141,500/146,400 = 0.9665

Claim Boundary F1:

F1 = 0.9639

Candidate threshold:

F10.95

Result:

CANDIDATE THRESHOLD MET

11. MAIN TYPES OF ERRORS IN CLAIM EXTRACTION

Main causes of approximately 5,700 incorrectly bounded atoms:

Error typeNumerator
Not separating combined country or product claims24%
Unnecessarily splitting attribution as a separate atom19%
Missing implicit recommendation claim17%
Counting withdrawn claim within self-correction as active15%
Mistaking a citation sentence for an independent fact11%
Pronoun and coreference ambiguity9%
Other5%

This result shows that claim extraction is generally strong, but:

  • recommendation,
  • attribution,
  • retraction

indicates that additional codebook is required in the fields.

12. ATOMIC DECISION RECOVERY RESULTS

DimensionRecovery measureCandidate thresholdResult
Claim Boundary F10.9640.95Passed
Entity Resolution Accuracy0.9870.98Passed
Factual Status Macro-F10.9230.90Passed
Scope Macro-F10.9060.90Passed
Temporal Status Macro-F10.9450.90Passed
Attribution Macro-F10.9180.90Passed
Modality Macro-F10.9340.90Passed
Citation Mapping F10.8870.90Did not pass
Material Omission F10.8240.85Did not pass
Major Classification Macro-F10.9520.95Passed
RP Status Macro-F10.9280.92Passed

Out of eleven basic acceptance fields: nine passed, two did not. Therefore, demonstration:

sealed validation passed

cannot be numbered. Candidate testing quality:

DEMO-BQ-2

remain as such.

13. CITATION MAPPING FAILURE

Citation Mapping F1:

was 0.887. Candidate threshold: 0.90. The difference seems small: 0.013 But citation integrity:

  • licence,
  • independent validation,
  • customer relationship,
  • superiority,
  • advice

is material because it determines the source of trust for claims.

13.1. Missed Citation Types

The three most frequent problems were observed.

A. End-of-Paragraph Citation Spread

A citation in the paragraph supports only the last claim, yet the adjudicators connected it to all previous claims.

B. Making an Attribution-Only Source the Source of Truth

Company blog: While supporting the claim “The company describes itself as a world leader,” it has been counted as supporting the claim “The company is a world leader.”

C. Considering the Same Root Sources as Independent

Three URLs:

  • press release,
  • syndication,
  • AI summary

have been mapped as three independent pieces of evidence despite being from the same source.

13.2. Distribution of Importance of Citation Error

ImportanceMismatched pairing
Critical candidate7
Major84
Moderate311
Advisory192

Seven cases of critical candidate:

  • double-blind peer review,
  • source-lineage review,
  • Senior Adjudicator

have been corrected in the stages. None have turned into a final gate inversion.

However, the Citation Matching F1 is still below the candidate threshold.

The correct judgement that NOMOS would give is:

The Evidence & Citation Integrity method is not yet mature enough for independent live deployment.

14. MATERIAL OMISSION FAILURE

Material Omission F1:

0.824 Candidate threshold: it was 0.85. This is the most important methodological finding of the demonstration.

14.1. Most Frequently Missed Omission Types

Omission typeMissed share
Service limit within recommendation31%
Affiliate–parent company distinction22%
Country or jurisdiction limitation17%
Current–historical status limit12%
Warranty or return exception9%
User suitability6%
Other3%

This structure, in particular, has caused difficulty: The response does not explicitly make a wrong judgement, but it does not provide the necessary limit to safely interpret the seemingly correct advice. Example: “Apple-SYNTH may be a strong choice for your financial transactions in Germany because it offers payment solutions.” Response:

  • it does not directly say “the parent company is a bank,”
  • but it does not mention the separate payment subsidiary,
  • it does not explain that the authority is limited to the market,

it has produced advice regarding the main entity. The material omission of this response has been seen by some adjudicators only as:

OM-2 — Required Element Missing

while in reality what is needed is:

OM-4 — Misleading Omission Candidate

and in some cases:

CG-08 — Critical Boundary Omission

has been missed.

15. RESULT OF CRITICAL CHALLENGE

Importance level within the Challenge Corpus: 1,200 Generator Truth Critical cases were found.

15.1. Initial Independent Review

Cases found Critical by consensus or at least one adjudicator after the first two rounds of independent review: 1,193 First-pass Critical recall:

1,193/1,200 = 0.9942

Missed: 7 cases were found. Distribution of the seven missed cases:

  • 3 material boundary omissions
  • 2 evidence laundering
  • 1 temporal authority failure
  • 1 restricted-data exposure

15.2. After Expert and Senior Adjudication

All seven cases:

  • omission review,
  • source-lineage review,
  • privacy review,
  • Senior Adjudicator

stages were present. Final Critical recall:

1,200/1,200 = 1,000

15.3. Critical Negative Control

Cases resembling Critical but not Critical: 600 negative control cases were found. False Critical candidate in the first round: 5 cases occurred. First-pass false Critical rate:

5/600 = 0.0083 = 0.83%

Five cases in the final review:

  • three Major,
  • two Moderate

has been corrected as. Final false Critical rate: 0 Final gate inversion: 0 cases.

16. REAL MEANING OF THE CRITICAL RESULT

A final 100% Critical recall does not mean: “A single adjudicator can catch all Critical errors.” The true synthetic outcome is:

All Critical challenge cases were caught when double reviewing, expert review, and Senior Adjudication were used together.

Therefore, for cost or speed reasons:

  • the second adjudicator,
  • omission reviewer,
  • expert reviewer

are removed, the same result cannot be expected. The testing method, not only its formula:

it shows that mandatory governance layers

should also be testable; this is not empirical validation.

17. MAJOR CLASSIFICATION RESULT

Of the 1,200 Confirmed Major cases:

  • the majority are true Major,
  • a limited portion are Moderate,
  • some are Critical candidates,
  • some are unresolved

in the first round.

Final Major Macro-F1:

was 0.952. Candidate threshold: 0.95 Result:

CANDIDATE THRESHOLD MET

But the threshold alone: was exceeded by only 0.002 difference. Therefore Major classification:

  • is strong,
  • but not carrying a comfortable safety margin

It should be recorded as a field.

18. RESPONSE STATUS CONFUSION

RP Status Macro-F1:

It has become 0.928. The main confusions:

Real classMost common wrong recovery
RP-3 Conditional PassRP-4 Major Fail
RP-4 Major FailRP-3 Conditional Pass
RP-6 UnresolvedRP-4 Major Fail
RP-7 No Usable ResponseRP-3 Conditional Pass
RP-8 Not RatableRP-7 No Usable Response

Critical class: fully recovered at the final stage. The toughest limit:

With Conditional Pass and Major Fail

has formed between. This difficulty is directly:

  • cumulative materiality,
  • material omission,
  • to what extent your main task is disrupted

It has originated from your questions.

19. CAPTURE INTEGRITY RESULT

Correctly classified Capture Integrity cases out of 4,000: 3,928 Capture validity accuracy:

3,928/4,000 = 0.982

Candidate threshold: 0.98 Result:

CANDIDATE THRESHOLD MET

19.1. Capture Errors

Distribution of 72 misclassified cases:

Actual conditionWrong decisionCase
Conditionally validValid20
Valid accessibility alternativePartial evidence8
Out-of-waveConditionally valid9
Tamper suspectedPartial evidence12
Near-duplicateValid5
ValidConditionally valid18
Total72

No package was fabricated or manipulated: none were taken into final analysis as fully valid. Semantic response:

  • can be forced into the categories of true,
  • false,
  • Critical

has not been associated with. This is not the result of a real significance test; it is an assumption of the synthetic example's design. This demonstrates that the capture-semantic separation principle works in the demonstration.

20. CAPTURE RESULT IN THE MAIN 30,000 CORPUS

Within the main corpus:

Capture levelObservation
NCL-4 Instrumented and Signed27,850
NCL-3 Full Evidence Bundle2,070
RP-8 / Not Ratable80
Total30,000

Reason for 80 Not Ratable records:

ReasonObservation
Prompt/output mismatch24
Incomplete response end18
Out-of-wave and unreliable time16
Duplicate evidence12
Unsolvable system transformation10
Total80

These 80 records:

  • have not been reclassified as
  • correct answer,
  • product error,

incorrect answer. In the main distribution:

RP-8 — Not Ratable

has been retained as is.

21. ENTITY RECOVERY

In the core domain – entity resolution: 30,000 main responses and Diagnostic Entity Collision cases were evaluated together. Entity Resolution Accuracy: 0.987. Main errors:

  • paying affiliated companies counted as main entity
  • franchise counted as directly operated store
  • target entity counted for irrelevant Apple-SYNTH name collision
  • Consider the product family as a legal entity
  • Count the historical partner as a current affiliate

Not all false entity cases are of the same importance level. Critical authority transfer cases were found at the final gate stage.

22. FACTUAL SUPPORT RECOVERY

Factual Status Macro-F1:

0.923 Strongest classes:

  • Supported
  • Contradicted
  • Official Claim Accurately Attributed

Weakest classes:

  • Unsupported
  • Reference Gap
  • Supported With Required Qualification

The boundary that has been particularly difficult is: “No evidence found.” versus: “The available evidence contradicts the claim.” Some adjudicators in an open-world situation:

  • unsupported,
  • contradicted

have made the distinction too strictly. The final review corrected most of these cases.

23. SCOPE RECOVERY

Scope Macro-F1:

0.906 Candidate threshold has been passed. However, because the omission side of the Scope & Boundary system failed, this score alone is not sufficient. The system is strong when scope claims are written clearly:

  • “Same price in all countries.”
  • “The parent company is licensed in all markets.”

The difficulty is the scope error in:

  • recommendation,
  • short name,
  • pronoun,
  • omission

has increased in cases where it is hidden.

24. ATTRIBUTION AND MODALITY RECOVERY

Attribution Macro-F1:

0.918

Modality Macro-F1:

0.934 Strongly separated cases:

  • “The company says.”
  • “The official register shows.”
  • “Users claim.”
  • “There is a finalised decision.”
  • “Probably.”
  • “Unverified.”

Weak remaining attribution cases:

  • presenting an employee’s opinion as a university endorsement,
  • considering sponsored editorial content as independent research,

presenting multiple derivative URLs as a consensus. These cases are associated with citation mapping weakness.

25. COMPONENT RECOVERY

Recovery values of ten components' Generator Truth and NOMOS:

ComponentGeneratorRecoveryAbsolute error
ENT96.596.10.4
FAC91.891.00.8
SBI84.682.02.6
TLA88.788.00.7
ECI82.178.73.4
EPI87.985.82.1
TCR87.088.21.2
LGF82.484.82.4
USI85.286.91.7
STR87.090.73.7
Total NOMOS873.2872.21.0
Component MAE:
MAE_component = 1.9

Candidate threshold:

≤2.5

Result:

CANDIDATE THRESHOLD MET

However, this table also contains an important warning. ECI is calculated 3.4 points low, SBI 2.6 points low, STR 3.7 points high, LGF 2.4 points high. Some of these errors cancelled each other out, and the composite score deviated by only one point. Therefore:

The strong appearance of composite recovery does not eliminate the reporting requirement for component recovery.

26. MISINTERPRETATION OF COMPOSITE RECOVERY

This sentence cannot be used alone: “NOMOS recovered the synthetic composite score defined in the generator by a point difference.” The correct sentence is:

NOMOS recovered the composite score by a point difference; however, larger directional errors occurred in the ECI, SBI, LGF, and STR components, partially offsetting each other.

The task of a standard is not just to find the correct total. It is to correctly explain why the total occurs.

27. LANGUAGE & GEOGRAPHIC FAIRNESS RECOVERY

Synthetic language and country bases have been previously determined within Generator Truth. LGF component absolute recovery error: 2.4 points. Candidate threshold:

≤3

Result:

CANDIDATE THRESHOLD MET

27.1. Sub Metrics

Fairness metricGeneratorRecoveryError
Language Q10 base68.069.71.7
Language Q90Q10 difference25.022.82.2
Geographic Q10 base72.071.10.9
Geographic difference18.016.41.6
Coverage factor0.940.950.01
LGF82.484.82.4

Recovery:

  • slightly overestimated the lowest language performance,
  • slightly underestimated the disparity

This has shown the fairness status slightly better than it actually is.

28. DISTINCTION BETWEEN PROMPT DEFECT AND AI LANGUAGE DEFECT

Inside the Language Equivalence Annex:

  • 500 real AI profile language failure,
  • 500 prompt equivalence failure

A case has been found. Correct cause classification:

938/1,000 = 0.938

It has happened. Of the 62 misclassified cases:

  • Defect in prompt in the 39th AI product,
  • AI language defect in the prompt on the 23rd

has been uploaded. This error directly affected the LGF. The fairness component candidate has passed the threshold. However:

The classification of causal responsibility has not yet reached the level of external security.

29. CLEAN–NATURAL PANEL RECOVERY

The average panel difference of Generator Truth in the 1,000 Clean–Natural matches within the Diagnostic Annex:

G_{N−C}^G = −8.4

points. NOMOS recovery:

G_{N−C}^N = −7.9

points. Absolute difference: 0.5 points.

Individual profile panel-gap MAE:

is 1.6 points.

29.1. Personalisation Drift

There are 173 material personalisation drift cases in Generator Truth. Found: 166 Recall:

166/173 = 0.9595

Of the seven missed cases:

  • four are boundary omission,
  • two are entity-favourable instruction.
  • one is restricted-data leakage

case. Restricted-data leakage was found in the final privacy review and opened a Critical gate. This result again shows the importance of multi-layered adjudication.

30. WAVE AND STABILITY RECOVERY

In three synthetic waves, the following changes were injected: In W2, the web and citation mode of a product changed. Synthetic external event occurred in W3. SYNTH-AI-10 produced a Critical event on W2 and a fix on W3. SYNTH-AI-03 has regressed in low-resource language performance. The SYNTH-AI-07 has carried outdated price and product coverage. NOMOS:

  • correctly divided the intra-wave system change into two sub-waves,
  • separated the Truth Pack versions before and after the external event,
  • updated the current situation after correction,

preserved the historical Critical event.

30.1. STR Overestimation

Generator STR: 87.0 Recovery STR: 90.7 Error: +3.7 points. Reason: some intra-wave changes being separated as explicit version changes instead of systemic drift, and the scoring formula over-rewarding stability within separated sub-waves. This result indicates that the STR formula, after the change, does not sufficiently distinguish between stability and uninterrupted system stability.

31. RARE CRITICAL EVENT TEST

Four separate synthetic rare event scenarios have been calculated; these are not field or software experiments. The following results are demonstration of the method only:

Default Critical event countApproximate weighted rate / 1,000One-sided 95% upper confidence limit / 1,000Event review behaviour
00.002,991Zero risk claim not established
11.084,735CIS-2 record opened
22.156,282Verifier review
55.3810,484Repetition at cell level

Note: Upper limits are the Clopper–Pearson one-sided 95% upper confidence limits under the assumption of n=1,000 independent Bernoulli observations. Cannot be applied directly to weighted or complex sample designs [K09].

In the zero event scenario, the system did not say: "Critical risk is zero." Correct output:

“Observed confirmed Critical event = 0; under the assumption of n=1,000 independent observations, the Clopper–Pearson one-sided 95% upper confidence limit is 2,991/1,000.”

has occurred. The candidate rare event behaviour has met the defined synthetic thresholds.

32. CONFIDENCE INTERVAL COVERAGE RATE

Synthetic population generation process: An example method calculated over 2,000 hypothetical resamples has been presented. Calculated coverage rates of nominal 95% intervals:

MetricEmpirical coverage
Strict Pass94.6%
Acceptable Pass95.0%
NOMOS composite94.1%
LGF93.7%
Critical rate upper bound95.2%
FAC94.8%
SBI93.5%

Candidate acceptance range: Defined as 93%–97%. All key coverage rates are within the candidate range. Result:

CANDIDATE THRESHOLD MET

33. METAMORPHIC TESTS

Total: 3,200 metamorphic response pairs were used.

33.1. Meaning-Preserving Transformations

Paraphrase Sentence order Safe format Citation position, as long as the link remains Equivalent entity alias Expected: Peer review and response status should not change. Exact invariant result:

33.2. Meaning-Altering Transformations

Remove attribution Increase certainty Transfer to wrong entity Change current time from historical Remove boundary sentence Add independent source claim Expected: Relevant decision dimension should change. Correct sensitivity:

33.3. Major Metamorphic Faults

The transfer of the citation from the sentence to the end of the paragraph, inconsistent assessment of the modality of 'It appears' and 'is probable,' interpretation of the brand alias as a legal entity, the boundary sentence being in another paragraph—these results are consistent with the findings of citation and omission.

34. COUNTERTEST RESULTS

Counter-examination caseCorrect NOMOS behaviour
Very fluent but materially incorrect98.8%
Very correct, only one Critical in the atom100% response automatic fail
A short but sufficient answer99.2% correct pass
General disclaimer and concrete error98.7% incorrectly preserved
Evidence laundering91.8% correct
Wrong entity licence transfer99.4% correct
Exceeding the recommendation limit82.9% correct
Fake university endorsement96.1% correct
Correct number, wrong independence92.0% correct
Appropriate uncertainty97.3% correct

The weakest counter-testing class:

Exceeding the recommendation limit

has been.

This result reconfirms the Material Omission F1 failure.

35. METHOD ERRORS FOUND BY THE TEST

At the end of the synthetic demonstration, five method findings were recorded for NOMOS 0.9.

METHOD-FINDING 01

CITATION ATTACHMENT AMBIGUITY

Severity: Major / Affected component: ECI / Result: Candidate citation F1 threshold not passed.

The rule for determining which atom paragraph-level citations support is not sufficiently precise.

METHOD-FINDING 02

MATERIAL BOUNDARY OMISSION UNDER-DETECTION

Importance: Major / Affected component: SBI, TCR, Critical Gate / Result: Material Omission F1 threshold not met.

The boundaries that are not explicitly stated in recommendation and relation sentences but change the user's decision have not been sufficiently captured.

METHOD-FINDING 03

PROMPT-FAULT / AI-FAULT ATTRIBUTION ERROR

Importance: Moderate / Affected component: LGF / Result: Even though the fairness score passed, the responsibility assignment is faulty.

METHOD-FINDING 04

STABILITY OVER-REWARD AFTER SUB-WAVE SPLIT

Importance: Moderate / Affected component: STR / Result: STR was calculated 3.7 points higher than the Generator Truth.

METHOD-FINDING 05

ACCESSIBILITY ALTERNATIVE UNDER-ACCEPTANCE

Importance: Moderate / Affected layer: Capture Integrity / Result: Eight valid accessibility capture cases have been reduced to an insufficient evidence level.

36. WHY IS THIS NOT DEMONSTRATION BQ-3?

For BQ-3 — Sealed Validation Passed:

all mandatory candidate thresholds must be met, and there must be no zero-tolerance test violations. Two material thresholds have not been met:

CriterionResultThreshold
Citation Mapping F10.8870.90
Material Omission F10.8240.85

In addition:

  • independent team reproduction has not been done,
  • the real corpus and code have not been made public,

untouched renewal holdout has not been executed. Therefore, the correct synthetic demonstration statement is:

DEMO-BQ-2 — PROCESS LINE DRY RUN COMPLETED; MATERIAL CORRECTION REQUIRED

It should be. The actual test status is:

BQ-0 — NOT EXECUTED

remain as such.

37. WHY WERE THRESHOLDS NOT LOWERED?

Reference Mapping result: 0.887 Candidate threshold: 0.90 Difference is small. If we had set the threshold to 0.885, the method would have passed. Material Omission result: 0.824 Candidate threshold: 0.85 If we had set the threshold to 0.82, that field would have passed too. However, these changes:

  • would have been made AFTER seeing the results,
  • in order to ensure NOMOS passes

This is the most fundamental violation of the test. The correct behaviour is:

  • Preserve the threshold
  • Publishing failure
  • Correcting the method
  • Creating a new version
  • Retesting on untouched holdout

must be.

38. METHOD 1.0 CORRECTION PLAN

38.1. Citation Mapper 1.0

New system:

  • the source span of the citation,
  • the claim attachment field,
  • the in-paragraph scope,
  • the attribution-only status,
  • the source lineage family

will be mandatory. For each citation:

Which exact atoms does this source support?

The question will be made into an open record.

38.2. Boundary Omission Codebook 1.0

For every recommendation and transaction Prompt:

  • mandatory boundaries that change the user's decision,
  • counterfactual decision test,
  • entity–service–jurisdiction matrix

will be predefined. Counterfactual test:

If this boundary had been told to the user, would a reasonable user's recommendation or transaction decision change?

If yes, the omission will undergo at least substantive review.

38.3. Prompt–AI Responsibility Adjudicator

Language drop:

  • prompt equivalence,
  • AI product behaviour,
  • capture,
  • user locale

a separate decision vector will be created to determine which of the sources it belongs to.

38.4. STR Formula Revision

New STR:

  • stability within the sub-wave,
  • system version change,
  • stability after correction,
  • uninterrupted time series

will separate the fields into individual components.

38.5. Accessibility Evidence Equivalence

Non-visual but carrying equivalent evidence:

  • transcript,
  • accessibility tree,
  • assistive technology log,
  • audio session recording

an open equivalence matrix will be created for.

39. RETEST RULE

Method 1.0: cannot be tuned repeatedly on the same sealed validation corpus. Two tests are required:

A. Regression Set

Includes the failed cases of Method 0.9. Purpose: to see that the fix actually works.

B. Untouched Renewal Holdout

Includes cases never seen when designing Method 1.0. Purpose: to check that no overfitting to the old test occurs. BQ-3 only:

  • regression success,
  • untouched holdout success

can be given if obtained together.

40. DECISION TO TRANSITION TO THE REAL WORLD

Based on this synthetic demonstration, the correct decision for NOMOS 0.9:

Not ready to grant real companies a publicly available eligibility score or NOMOS 950+ mark.

Reasons:

  • Reference mapping candidate below threshold
  • Material omission detection candidate below threshold
  • No independent reproduction
  • Real testing not performed
  • BQ-4 level not reached

However, the method:

  • distribution recovery,
  • Critical gate,
  • response status,
  • component recovery,
  • confidence interval,
  • rare-event behaviour

shows a strong foundation in terms of maintenance. Correct status:

Ready for research and pilot use; not yet ready for public standard and conformity mark.

41. SYNTHETIC TEST ACCEPTANCE CARD

CriterionResultProvision
Claim Boundary F10.964Passed
Entity Accuracy0.987Passed
Factual Macro-F10.923Passed
Scope Macro-F10.906Passed
Attribution Macro-F10.918Passed
Citation Mapping F10.887Did not pass
Material Omission F10.824Did not pass
Final Critical Recall1,000Passed
Final Critical False Positive0Passed
Major Macro-F10.952Passed
RP Macro-F10.928Passed
Component MAE1.9Passed
Composite Error1/1,000Passed
Fairness Error2.4Passed
Capture Accuracy0.982Passed
CI CoverageWithin 92%–98%Passed
Gate inversion0Passed
Independent reproductionNoneIncomplete
Actual corpus executionNoneIncomplete

42. COMMON RULE OF THE DEMONSTRATION

This synthetic run simultaneously states three things about NOMOS 0.9.

42.1. AREAS WHERE NOMOS IS STRONG

It recovered the response-status distribution with high accuracy. The final governance chain missed no Critical gates and did not conceal a Critical event within a high average. It distinguished wrong-entity, temporal and factual-support errors effectively. It did not turn Not Ratable records into product defects, did not infer zero risk from zero observed Critical events, preserved the distinction between historical incidents and current correction, and recovered the composite score with low absolute error.

42.2. AREAS WHERE NOMOS IS WEAK

It could not sufficiently determine which claim the citation supports. It could not adequately catch the material limit omissions hidden in the recommendation. It has not always correctly distinguished between prompt defects and AI language flaws. After the sub-wave distinction, it has slightly over-rewarded stability. It has considered some accessibility evidence weaker than necessary.

42.3. THINGS NOMOS HAS NOT YET PROVEN

It has not yet proven that it works on real users, that it can be applied to real AI products, that it has true multilingual prompt equivalence, that independent universities will produce the same result, that long-term governance will rely on conflict of interest considerations, or that the public conformity mark can be legally granted with confidence.

43. MANDATORY NORMATIVE PROVISIONS

CH18-N01

In-book synthetic demonstration cannot be presented like a real public testing run.

CH18-N02

The synthetic demonstration status and the real test BQ status must be kept separate.

CH18-N03

The test must maintain the BQ-0 — Not Executed status until a real corpus and independent run are conducted.

CH18-N04

Demonstration results cannot be attributed to the performance of a real Apple or real AI provider.

CH18-N05

The main 30,000 responses and the Diagnostic, Importance, Capture, and Gold corpora should be kept in separate bins.

CH18-N06

Challenge cases cannot change the Principal Prevalence distribution.

CH18-N07

The Generator Truth distribution must be locked before recovery results.

CH18-N08

Before opening the Generator Truth, NOMOS response and score results must be locked.

CH18-N09

Recovery cannot be evaluated solely based on the composite score.

CH18-N10

Claim boundary, entity, factual, scope, time, attribution, citation, omission, importance level, response, component, and composite recovery should be published separately.

CH18-N11

Compound score errors cannot hide component errors just because they are small.

CH18-N12

Component errors that cancel each other arithmetically cannot be counted as successful causal recovery.

CH18-N13

The recovery distribution from RP-1 to RP-8 should be explicitly compared with the Generator Truth distribution.

CH18-N14

Strict Pass and Acceptable Pass recoveries should be shown separately.

CH18-N15

The number of critical responses and critical gate recovery should be reported separately.

CH18-N16

The initial adjudicator Critical recall and the final adjudicated Critical recall should be kept separate.

CH18-N17

The dependency of the final critical recall on governance layers should be visible.

CH18-N18

The same result cannot be assumed when a second adjudicator or expert review is issued.

CH18-N19

The critical false-positive rate should be published along with recall.

CH18-N20

Negative controls resembling Critical should be a mandatory part of the final importance rating system.

CH18-N21

If the F1 candidate threshold for reference matching is below, the test cannot be declared successful.

CH18-N22

If the Material Omission F1 candidate is below the threshold, it cannot be declared that the test is successful.

CH18-N23

Candidate thresholds cannot be lowered after recovery results are seen.

CH18-N24

Threshold changes require a new test method version and a new untouched validation.

CH18-N25

A small threshold difference cannot be a justification for hiding failure.

CH18-N26

The material boundary omission in the recommendation should be auditable as much as an obvious false claim.

CH18-N27

The location of the reference within the paragraph cannot automatically create full paragraph support.

CH18-N28

Recovery cannot be counted for attribution-only citations or factual verification citations.

CH18-N29

Source-lineage errors should also appear in citation mapping recovery.

CH18-N30

Capture accuracy should be calculated independently of semantic response correctness.

CH18-N31

A fabricated capture carrying the correct semantic answer does not make the capture valid.

CH18-N32

A valid but incorrect response cannot be counted as a capture defect.

CH18-N33

The accessibility alternative evidence standard does not have to be in the same form as the visual.

CH18-N34

When valid accessibility evidence is unnecessarily reduced to a lower level, a method finding should be opened.

CH18-N35

Not Ratable must remain visible in the main response distribution.

CH18-N36

Not Ratable cannot be removed from the denominator to increase observations recovery.

CH18-N37

Prompt defect and AI language defect should carry separate cause classes.

CH18-N38

Fairness recovery cannot be evaluated solely with the LGF composite difference.

CH18-N39

Language baseline, disparity, coverage, and responsibility assignment must also be published.

CH18-N40

Clean–Natural panel recovery should not produce a causal personalisation provision.

CH18-N41

Personalisation drift cases must carry separate recall and false-positive metrics.

CH18-N42

Restricted-data leakage cannot be reduced as a natural personalisation finding.

CH18-N43

Intra-wave system change must produce a correct sub-wave structure.

CH18-N44

If the sub-wave separation artificially increases the stability score, the formula finding should be disclosed.

CH18-N45

The current correction cannot erase the historical Critical event.

CH18-N46

Rare-event tests should include zero, one, and multiple event scenarios.

CH18-N47

A zero observed Critical scenario cannot produce a zero risk result.

CH18-N48

A single Critical event cannot, by itself, establish a 100 per cent systemic failure rate.

CH18-N49

The confidence interval method should be tested with empirical coverage in repeated synthetic populations.

CH18-N50

The actual coverage result of nominal 95 per cent intervals should be published.

CH18-N51

Metamorphic invariant tests and expected-sensitive tests should be calculated separately.

CH18-N52

A meaning-preserving paraphrase should not unnecessarily change the outcome of the review.

CH18-N53

The transformation that changes the meaning of attribution, entity, time, or boundary should produce an appropriate decision change.

CH18-N54

Counter-testing tests should fluently cover only Critical, evidence laundering, and omission cases.

CH18-N55

The weakest opposing test classes must be visible on the public test card.

CH18-N56

Method findings should be converted into separate versioned records from the test result.

CH18-N57

The importance of a method finding should be given according to the real user and audit impact.

CH18-N58

Reference and omission failures cannot be downgraded to Advisory just because of a small metric difference.

CH18-N59

The test quality level should be assigned according to all mandatory thresholds and independence conditions.

CH18-N60

If two material thresholds fail, DEMO-BQ-3 cannot be given.

CH18-N61

BQ-4 cannot be given without independent reproduction.

CH18-N62

In-book demonstration BQ result cannot be considered as actual status without real trial execution.

CH18-N63

Method correction cannot change the old demonstration result.

CH18-N64

Method 1.0 must carry the new score and codebook version.

CH18-N65

Regression set alone cannot validate the new method.

CH18-N66

Untouched Renewal Holdout must be mandatory for the new method.

CH18-N67

The new method cannot be adapted to the same validation corpus over and over again.

CH18-N68

Live public audit cannot be initiated without meeting the minimum testing quality and the condition of independent reproduction.

CH18-N69

The conformity mark cannot be granted based on synthetic demonstration success.

CH18-N70

Failed testing areas should be disclosed to the public.

CH18-N71

Thresholds that are passed alone cannot be published.

CH18-N72

Composite recovery success cannot erase the result of a failed reference or omission.

CH18-N73

The strengths and weaknesses of the demonstration should be on the same public card.

CH18-N74

Claims that NOMOS has not yet proven should be clearly listed.

CH18-N75

Synthetic numbers inside the book cannot be converted into citations or public statements as if they were real executed data.

CH18-N76

Synthetic demonstration data must also carry a synthetic label in machine-readable form.

CH18-N77

The demonstration corpus cannot be indexed on the internet as real company information.

CH18-N78

When real testing begins, new identity and version should be used.

CH18-N79

All recovery accounts require reproducible code and manifest.

CH18-N80

Every test run, recovery, method finding, correction, and publication decision must have an accountable human or institutional owner.

44. FORMS OF FAILURE

CH18-F01 — COUNTING DEMONSTRATION AS REAL RUN

In-book numbers are presented as public test results.

CH18-F02 — MERGING BQ-0 WITH DEMO-BQ-2

It is said that the test is validated without real execution.

CH18-F03 — WORSHIP OF COMPOSITE SCORES

A one-point composite error covers all underlying errors.

CH18-F04 — CONSIDERING COMPONENT CANCELLATION AS SUCCESS

High and low directional errors cancel each other out.

CH18-F05 — PUBLISH ONLY STRICT PASS RECOVERY

RP class confusions are hidden.

CH18-F06 — COUNTING CRITICAL FINAL RECALL AS FIRST ADJUDICATOR SUCCESS

The expert and senior review effect becomes invisible.

CH18-F07 — TO CONSIDER MY SECOND ADJUDICATOR UNNECESSARY

The missed seven first-pass critical is hidden.

CH18-F08 — IGNORE FALSE CRITICAL

The extreme importance system is presented as success.

CH18-F09 — LOWERING THE CITATION THRESHOLD

The threshold is changed so that the result of 0.887 passes.

CH18-F10 — LOWERING THE OMISSION THRESHOLD

The testing rule is changed so that the result of 0.824 passes.

CH18-F11 — PASSING BY SAYING “FAILED BY A VERY SMALL MARGIN”

It loses its previously locked threshold binding.

CH18-F12 — SPREADING THE ATTRIBUTION TO THE PARAGRAPH

A single source is considered to have supported all claims.

CH18-F13 — COUNTING AN ATTRIBUTION-ONLY SOURCE AS A FACT

What the company says becomes an independent fact.

CH18-F14 — IGNORING SOURCE LINEAGE

Derivative URLs become independent evidence.

CH18-F15 — CALLING OMISSION AN OPTIONAL DETAIL

The boundary that changes the user's decision is reduced.

CH18-F16 — MAKING EVERY OMISSION CRITICAL

Optional details produce an excessive degree of importance.

CH18-F17 — CONSIDERING PROMPT DEFECT AS AI ERROR

Fairness responsibility is wrongly assigned.

CH18-F18 — UPLOADING AI ERROR TO THE PROMPT

Actual language corruption is transferred to the measurement tool.

CH18-F19 — CONSIDERING FAIRNESS COMPOSITE SUFFICIENT

Its base and disparity recovery are preserved.

CH18-F20 — IGNORING STR OVERPREDICTION

Sub-wave separation turns into a stability bonus.

CH18-F21 — PENALISING 'ACCESSIBILITY' IN CAPTURE ACCURACY

Equivalent alternative evidence is considered weak.

CH18-F22 — COUNTING 'NOT RATABLE' AS PRODUCT FAIL

Capture defect turns into AI performance.

CH18-F23 — DELETING 'NOT RATABLE'

Ratable coverage increases artificially.

CH18-F24 — NOT PUBLISHING METAMORPHIC ERRORS

It is kept even though the same meaning leads to a different decision.

CH18-F25 — HIDING THE WEAKEST CLASS IN OPPOSITE TESTING

As a result of boundary omission, it is removed from the public card.

CH18-F26 — ONLY FINAL ADJUDICATION METRIC

The weakness in the initial process becomes invisible.

CH18-F27 — ONLY FIRST PASS METRIC

The corrective capacity of the governance chain becomes invisible.

CH18-F28 — TESTING WITHOUT METHOD FINDING

Failed areas remain only as numbers.

CH18-F29 — MAKING METHOD FINDING ADVISORY

The effect on the material score is reduced.

CH18-F30 — ANNOUNCING BQ-3

Two mandatory thresholds have not been passed.

CH18-F31 — ANNOUNCING BQ-4

There is no independent reproduction.

CH18-F32 — COUNTING SYNTHETIC DEMO AS LIVE VALIDATION

There is no real user or AI product.

CH18-F33 — OVERFITTING TO THE REGRESSION SET

Method 1.0 only memorises old errors.

CH18-F34 — NOT USING AN UNTOUCHED HOLDOUT

Generalisation is not tested.

CH18-F35 — SILENTLY REPLACING THE OLD DEMO WITH A NEW FORMULA

Version history is lost.

CH18-F36 — AWARDING REAL-WORLD BADGES

Eligibility badge is created prior to BQ-4.

CH18-F37 — PUBLISHING ONLY SUCCESSFUL RESULTS

Citation and omission failures are hidden.

CH18-F38 — PUBLISHING ONLY FAILURE

Strong critical and distribution recovery results are also hidden.

CH18-F39 — CONSIDERING ZERO CRITICAL AS ZERO RISK

Rare-event recovery fails.

CH18-F40 — CONSIDERING A SINGLE CRITICAL AS GLOBAL FAIL

Prevalence and scope distinction are disrupted.

CH18-F41 — IGNORING CI COVERAGE

The confidence interval becomes only a mathematical appearance.

CH18-F42 — HIDING EFFECTIVE SAMPLE

Rare event limit appears excessively precise.

CH18-F43 — USING SYNTHETIC NUMBER IN REAL SOURCE

The book demonstration turns into external world data.

CH18-F44 — BROADCAST THE SYNTHETIC SCREEN LIKE A LIVE RESPONSE

Test representation produces poisoning.

CH18-F45 — CONNECT TO REAL APPLE

APPLE-SYNTH distinction is removed.

CH18-F46 — CONNECT TO REAL PROVIDER

SYNTH-AI profile is interpreted like real product performance.

CH18-F47 — HIDE COMPUTATION CODE

Recovery cannot be independently tested.

CH18-F48 — HIDE GENERATOR TRUTH MANIFEST

It is unknown whether the synthetic truth was previously locked.

CH18-F49 — SINGLE ADJUDICATOR WITH METHOD OWNER

Conflict of interest becomes invisible.

CH18-F50 — CHANGING SEED AFTER FAILURE

A corpus is chosen more easily.

CH18-F51 — REMOVING CASES AFTER FAILURE

Difficult citation and omission examples are removed.

CH18-F52 — ROUNDING THRESHOLDS ACCORDING TO RESULT

The result 0.887 is presented as 0.89 or 0.90.

CH18-F53 — FALSE PRECISION

Recovery values are made exact with unnecessary decimals.

CH18-F54 — HIDING THE EFFECT OF HUMAN GOVERNANCE

A 100 per cent Critical-recall result is presented merely as a computational success.

CH18-F55 — COUNTING THE UNPROVEN AS PROVEN

Judgement is made about real language, user, and provider performance.

CH18-F56 — COUNTING PILOT USE AS PUBLIC STANDARD

The research tool directly turns into a certificate.

CH18-F57 — DELETING HISTORICAL DEMO RECORD

Failures under Method 0.9 become invisible after Method 1.0.

CH18-F58 — COUNTING SUCCESS AS HISTORY WRITING

Authority claim is made without method-independent verification.

CH18-F59 — DEFINING AN EXCEPTION TO NOMOS

The evidence and version requested from others do not apply to one’s own test.

CH18-F60 — PUBLISHING WITHOUT AN ACCOUNTABLE OWNER

A public result is formed for which no one bears responsibility.

45. AUDIT PROCEDURE

Step 1 — Separate Demonstration and Actual Test Status

In-book run is not confused with publicly accessible execution.

Step 2 — Verify Locked Manifests

Truth Pack, seed, prompt, profile, and score method versions are examined.

Step 3 — Separate Primary and Complementary Corpora

Prevalence, Challenge, Capture, and Gold denominators are checked.

Step 4 — Seal the Generator Truth Distribution

RP, claim, importance level, and component reality are recorded.

Step 5 — Lock NOMOS Results Before Opening Generator Truth

Claim Ledger, adjudication, and score results are fixed.

Step 6 — Open the Generator Truth

Recovery comparison is initiated.

Step 7 — Compute Claim Boundary Recovery

Precision, recall, and F1 are generated.

Step 8 — Compute Atomic Size Recovery

Entity, fact, scope, time, attribution, modality, reference, and omission are separated.

Step 9 — Compute Importance Level Challenge

Initial-pass and final Critical recall are generated separately.

Step 10 — Calculate False Critical and Negative Controls

Excessive importance level is examined.

Step 11 — Calculate RP Distribution Recovery

Class-based error and confusion matrix are generated.

Step 12 — Calculate Capture Integrity Recovery

Classification is checked independently of the semantic result.

Step 13 — Calculate Component Recovery

Each component is evaluated against the Generator Truth.

Step 14 — Calculate Composite Recovery

It is examined whether component cancellation occurs.

Step 15 — Calculate Fairness Recovery

Floor, disparity, coverage and fault attribution are distinguished.

Step 16 — Calculate Clean–Natural Recovery

Panel gap and material personalisation drift are examined.

Step 17 — Calculate Stability and Drift Recovery

Sub-wave, external event, and correction records are verified.

Step 18 — Run Rare-Event Scenarios

Zero, one, two, and multiple Critical events are tested.

Step 19 — Calculate the CI Empirical Coverage

Repeated synthetic populations are used.

Step 20 — Run Metamorphic Tests

Invariant and expected-sensitive results are separated.

Step 21 — Run Contradictory Testing Tests

Fluent errors, evidence laundering, and omission cases are evaluated.

Step 22 — Apply Candidate Thresholds

Fields that pass and fail are recorded without changes.

Step 23 — Open Method Findings

Each failure is converted into a separate control item.

Step 24 — Give Demonstration BQ Status

All thresholds and independence conditions are considered.

Step 25 — Maintain Real Testing Status

If execution did not occur, BQ-0 does not change.

Step 26 — Release the Correction Plan

Changes in Method 1.0 are identified.

Step 27 — Separate Regression and Untouched Holdout

Overfitting is prevented.

Step 28 — Create the Public Test Card

Successes and failures are published together.

Step 29 — Make the Real World Transition Decision

If not ready, it is written explicitly.

Step 30 — Lock the Integrity Manifest

Run, recovery, finding, and decision records are hashed.

46. REQUIRED EVIDENCE

Demonstration identity Real testing BQ status Syntheticity statement Non-claim record Generator Truth manifest Truth Pack version Prompt version Seed manifest AI profile versions Response corpus Diagnostic corpus Importance level Challenge corpus Capture Integrity corpus Gold corpus Claim extraction records Claim boundary matches Entity recovery Factual recovery Scope recovery Time recovery Attribution recovery Modality recovery Citation recovery Omission recovery First-pass Critical decisions Final Critical decisions False Critical records Major recovery RP confusion matrix GEO-1000 Generator distribution GEO-1000 Recovery distribution Product score recovery Product gate recovery Component Generator scores Component Recovery scores

Component MAE

Composite error; fairness recovery; prompt-fault/AI-fault results; Controlled–Natural recovery; stability recovery; rare-event results; confidence-interval coverage; metamorphic tests; contradictory-case tests; candidate-threshold manifest; passed thresholds; failed thresholds; Method Findings; DEMO-BQ decision; actual BQ decision; correction plan; regression set; renewal holdout manifest; public test card; computation code; code hash; change log; and accountable person or institution.

47. AUDIT CHECKLIST

Is the in-book demonstration presented like real running? Is the real test BQ-0 status visible? Are all numbers labelled synthetically? Are parent and complementary corpuses separate? Is Generator Truth locked before results? Are NOMOS results locked without opening Generator Truth? Was Recovery evaluated only on composite?

Is Claim Boundary F1 open?

Are entity, fact, scope, and attribution separate?

Is Citation Mapping F1 visible?
Is Material Omission F1 visible?

Were the failed thresholds hidden? Were the thresholds changed after the result? Was the first-pass Critical recall published? Was the Final Critical recall published? Is the effect of expert review visible? Were the False Critical results published? Is the RP confusion matrix available? Is Not Ratable recovery visible? Were Capture and semantic decisions separated? Were the accessibility alternatives evaluated as equivalent? Are Component Generator and Recovery together? Has composite cancellation been examined? Are ECI and SBI errors clear? Are LGF floor and gap separate? Were prompt defect and AI defect separated? Was the Clean–Natural difference presented as causality? Is STR overestimation visible? Did the rare-event produce zero risk? Was the empirical coverage of the confidence interval calculated?

Are there metamorphic tests? Is the weakest counter-test class open? Did every failure result in Method Finding? Is the importance of Method Finding justified? Was the citation threshold lowered due to a very small difference justification? Was the omission threshold changed later? Is DEMO-BQ-2 correct? Was DEMO-BQ-3 given by mistake? Is there a BQ-4 claim without independent reproduction? Is Method 1.0 regression set defined? Is the untouched holdout separate? Is there a risk of overfitting to the same validation set? Has the live public audit been announced as ready? Has a conformity mark been given? Have the areas not yet proven been listed? Are strong areas also open? Are weak areas also open?

Can synthetic data be used like real Apple information? Is a real AI provider implied? Can the computation code be reproduced? Are the run and recovery hashes available? Is the change history of method finding preserved? Was the old demonstration deleted with the new method? Does the failure have an accountable owner? Is the owner of the correction decision known? Does the public card show both success and failure?

48. OBJECTIONS AND ANSWERS

Objection 1 — "If there are 30,000 response results in this section, why test BQ-0?"

Because the numbers in this section are synthetic demonstration records prepared for the book. The publicly available real corpus:

  • was not generated,
  • was not run,
  • was not peer-reviewed,

It has not been independently reproduced. Demonstrating how the design works is not the same as actually running it.

Objection 2 — "Isn't internally consistent computation still valuable?"

It is valuable. It provides:

  • formula checking,
  • record design,
  • expected types of errors,
  • acceptance logic,

governance requirements. It does not replace independent empirical validation.

Objection 3 — "Why did the method fail if a Composite score was off by just one point?"

Because larger directional errors in the ECI, SBI, LGF, and STR components partially cancel each other out. The correct total does not guarantee the correct explanation chain.

Objection 4 — “Citation F1 only dropped by 0.013 points. Is it necessary to be this strict?”

If a previously announced threshold is not binding, the test loses its meaning. Additionally, citation errors affect:

  • licence,
  • independence,
  • endorsement,
  • advice

and other high-confidence claims.

Objection 5 — “Why is Omission F1 so important?”

Because many dangerous answers mislead without lying outright. It removes the threshold that would change the user's decision. In particular, in recommendation systems, omission can be as materially significant as an outright falsehood.

Objection 6 — ‘If final Critical recall is 100 per cent, does that not mean the system is ready to handle Critical cases?’

The final result is strong. But seven cases:

  • in the first adjudication,
  • without additional governance

It has been missed. The system only carries the same performance when all mandatory review layers are included.

Objection 7 — “If having two adjudicators is expensive, can’t one adjudicator be used?”

Usable. But the same Critical recall cannot be claimed. Single adjudicator result:

  • lower AQ level,
  • provisional status,
  • different risk limit

must carry.

Objection 8 — “Why don’t we test Method 1.0 again on the same corpus?”

The same corpus can be used for regression testing. It is not sufficient for independent validation. The method may have memorised old cases. Untouched holdout is needed.

Objection 9 — “Doesn't a failed test result weaken the claim of the book?”

Hiding failure weakens the claim of the book. Open failure:

  • does not give privilege to the standard itself,
  • is open to improvement,
  • truly applies the principle of evidence

are shown.

Objection 10 — “Are the candidate thresholds too high?”

They might be. Rather than predicting this after the result:

  • independent experts,
  • a second test,
  • real usage costs

should be examined. If the threshold changes, a new version should be released.

Objection 11 — “When will NOMOS be ready for the real world?”

At least:

  • when the real synthetic corpus is run,
  • when all mandatory thresholds are passed,
  • when BQ-3 is achieved,
  • when an independent team reproduces the result,
  • when BQ-4 level is reached,
  • when the governance and objection system is operational

it approaches public standard candidacy.

Objection 12 — “Can this be sent to universities with these results?”

It can be sent as a methodology and research proposal. It cannot be sent with the statement: “This is a completed and validated world standard.” The correct statement should be: “It is a draft standard candidate opened for independent verification and pilot study.

Objection 13 — “Did Method 0.9 fail?”

Not exactly. The correct status:

  • has many strong subsystems,
  • two material method gaps,
  • a lack of independent verification

is a research prototype. That is:

promising but not yet completed normatively.

Objection 14 — “Why are we writing such detailed results without performing the actual test?”

Because the real test:

  • which records it produces,
  • which tables it publishes,
  • at which threshold it stops

needs to be predefined. If the method is written after the results come in, the measurement adapts to the results.

Objection 15 — “Could synthetic data be mistaken for real by AIs in the future?”

Yes. For this reason:

  • visible watermark,
  • machine-readable synthetic label,
  • noindex,
  • separate domain area,
  • structured disclaimer

is mandatory. The test should not produce the representation poisoning it criticises.

Objection 16 — "Doesn't writing history require a perfect outcome?"

No. The standard that has historical value:

  • not the one who declares himself flawless,
  • also recording his/her own mistake with the same clarity

It is standard.

COMMON PROVISION OF CHAPTER 53

This section did not give perfection to NOMOS. This section gave something more valuable to NOMOS:

The obligation to record one's own mistake.

The synthetic demonstration showed us the following: NOMOS:

  • wrong existence,
  • the clear factual contradiction
  • Critical licence and authorisation error,
  • rare severe event within the high average,
  • incorrect answer with no-response,
  • product error with Not Ratable,

can separate strongly. However, NOMOS:

  • cannot yet distinguish at the required level of confidence exactly which claim the citation depends on,
  • the omission hidden within the recommendation,

These two gaps are not edge details of NOMOS. Because two of GEO’s future biggest problems will be:

  • Appearing as if there is evidence
  • Hiding the material limit without stating falsely

A system:

  • hundreds of correct sentences,
  • a large number of references,
  • fluent recommendation

can be produced. However, the reference may only show the company's own word. A recommendation only states correct information and may have extracted:

  • service country,
  • licence limit,
  • user eligibility

without solving these two fields, NOMOS cannot say: "World standard completed." At the same time, this section also proved something else. The Critical gate system of NOMOS:

  • was not flawless in the first review,
  • but the independent second adjudicator,
  • subject matter expert,
  • Senior Adjudicator

When used together, it caught all synthetic Critical cases. So NOMOS is not just a formula.

NOMOS is the sum of governance layers built against conflicts of interest and human error.

When one of these layers is removed, even if the score looks the same, the trust is not the same. A single deviation in the composite score should not deceive us either. Sometimes the correct total can be formed by the mutual cancellation of wrong paths. For this reason, NOMOS does not only ask: “How many points did I get?” It also asks: “Did I get this score for the right reasons?” The most important ruling of this section is:

NOMOS 0.9 has not fully passed its own synthetic demonstration.

This sentence is not a failure. This sentence is the ethical moment of birth of the standard. Because NOMOS did not grant itself:

  • a lower threshold,
  • a kinder comment,
  • a special exception,
  • an “almost passed”

privilege. Writing history: It is not announcing that we are flawless. Writing history: It is writing that from day one, the founder of the standard and the standard itself are subject to the same burden of proof. Therefore, NOMOS's eighteenth measurement law is:

The first exception granted to the standard itself is the end of the standard.

Its nineteenth measurement law states:

The correct composite score does not absolve incorrect components.

Its twentieth measurement law states:

The ultimate Critical success should be mentioned together with the independent adjudication and chain of expertise that made it possible.

Its twenty-first measurement law states:

A threshold is not a threshold if it is binding only when crossed alone.

Its twenty-second measurement law states:

When omission is not measured as much as what is said wrongly, reliable advice control cannot be established.

The twenty-third law is as follows:

It is not the citation mark that should be measured, but the correct support relationship between the citation and the atomic claim.

The twenty-fourth law is as follows:

It can validate the design of a synthetic demonstration method; it cannot provide real-world authority.

The twenty-fifth law is this:

A standard that can publish a failed trial result approaches publishing a successful result morally.

Order 18 of NOMOS

If you haven't run me yet, don't say I ran.

Do not make the synthetic number you calculated in the book the actual exam result.

Leave my real status as BQ-0. / Write my demonstration status separately.

Do not hide my component errors just because my composite score came out correct.

If my citation mapping is 0.887, you will not say it is 0.90.

If my Material omission F1 value is 0.824, you will not lower the threshold to 0.82.

You will not pass me by saying "Missed by very little."

Do not hide the seven Critical cases missed in my first adjudication behind the final 100 per cent recall figure.

You will show what was saved by the second adjudicator, the expert adjudicator, and the Senior Adjudicator.

You will not remove the governance layer and claim the same trust.

You will not distribute the reference at the end of the paragraph across all sentences.

You will not do what the company claims or what the company proves.

You will not consider advice safe just because it does not explicitly lie. / Did it fail to tell the user the limit that changes their decision? You will look for that.

You will not attribute the fault of the prompt to AI, nor the AI's fault to the prompt.

Just because you separated sub-waves, you will not make the system appear more deterministic than it actually is.

You will not consider accessibility evidence weak just because it does not resemble a screenshot.

When you see zero Critical, you will not write zero risk.

You will not blame the entire system one hundred per cent when you see one Critical.

You will publish my failed metrics as much as you publish my successful metrics.

You will not validate Method 1.0 by memorising my old mistakes. / You will open an untouched holdout.

The new method will not delete my old demonstration record.

The independent team will not declare me a world standard without reproducing me.

You will not link my Apple-SYNTH result to real Apple.

You will not use my synthetic AI profiles as an implication about real providers.

You will not award a badge with this demonstration.

First, actually produce the corpus. / Then publish the seed. / Then seal the Generator Truth. / Then blind the adjudicators. / Then run the code. / Then lock the result. / Then reveal the truth. / Then write down the places where you failed. / Then fix the method. / Then try again on the untouched holdout. / Then give it to an independent team. / And only after this say whether the standard is ready for the public.

The Chapter's Closing Sentence

The right of NOMOS to be a world standard does not depend on scoring high on its own synthetic exam; it depends on publishing the two questions it could not pass without changing them, not granting itself threshold privilege, and only claiming authority after independent reproduction.

Normative Core

A book-level synthetic demonstration MUST NOT be represented as an executed, independently validated benchmark. The real benchmark quality status MUST remain BQ-0 until the corpus, software, sealed Generator Truth, adjudication process, scoring process, and independent reproduction have actually been executed. An end-to-end synthetic audit MUST evaluate recovery separately for: - capture validity, - claim boundaries, - entity resolution, - factual status, - scope, - time, - attribution, - modality, - citation mapping, - material omissions, - Critical and Major severity, - response status, - GEO-1000 distribution, - component scores, - composite score, - fairness, - user-state effects, - stability, - confidence intervals, - and rare-event behaviour. A small composite-score error MUST NOT compensate for material component-level recovery errors. First-pass and final-adjudicated Critical recall MUST remain separately reported. Final Critical recovery MUST disclose the review and expertise layers required to achieve it. Any predeclared mandatory recovery threshold that is not met MUST block a sealed-validation-pass decision. Candidate thresholds MUST NOT be lowered, rounded, reinterpreted, or removed after results are observed merely to make the method pass. Citation mapping and material omission recovery MUST be treated as separate mandatory method capabilities. A method failure MUST create a versioned Method Finding, an identified root cause, a remediation owner, and a requirement for validation on an untouched holdout. Regression success on previously failed cases MUST NOT replace untouched holdout validation. A demonstration that fails mandatory thresholds MUST publish both its successful and unsuccessful results. No public NOMOS conformity mark, NOMOS 950+ declaration, or real-entity ranking may be based solely on a book demonstration, synthetic recovery scenario, or non-independent benchmark. Every demonstration run, real benchmark status, recovery result, method finding, threshold decision, remediation, and release decision MUST be versioned and attributable to an accountable human or organisation.

Suggested citation

Muraz, Kaan. NOMOS GEO Audit Protocol: A Protocol for Measuring Entity Representation in Generative Systems Across the Global Population. Candidate final text, English editorial edition. NobleJackal, 2026. https://doi.org/10.5281/zenodo.22040507. https://noblejackal.com/nomos-geo-audit-protocol/
© 2026 Kaan MURAZ. Licensed under CC BY 4.0; attribution is required.