Chapter Boundary
Chapter 17 defined the Apple.com Synthetic Benchmark. This chapter uses an internally consistent, arithmetically verified in-book demonstration to show what the full audit chain would require:
- a calculated synthetic record representing 30,000 principal responses,
- modelled capture defects,
- Generator Truth concealed from adjudicators,
- responses decomposed into atomic claims,
- Critical and Major gates applied,
- language and user-state differences calculated,
- and the scoring architecture processing the resulting data.
The question is how closely NOMOS can recover a previously sealed synthetic reality. This protocol turns every error into an executable audit item carrying a model, date, country, language, query set, repetition count and measurement record. Chapter 18 now directs that demand at NOMOS itself. It begins, however, with an integrity gate.
The 30,000-response record in this chapter is an internally consistent, arithmetically verified synthetic application prepared for the book. It is not an executed corpus or field experiment.
These numbers do not represent:
- the real performance of Apple Inc.,
- the real content of apple.com,
- the behaviour of any real AI provider,
- responses submitted by real participants,
- or an independently executed and publicly released software benchmark.
Two statuses must therefore remain separate.
IN-BOOK DEMONSTRATION STATUS
The in-book demonstration combines:
- the design established in Chapter 17,
- predefined synthetic distributions,
- Generator Truth records,
- fault injections,
- response statuses,
- recovery metrics.
Together, these elements provide an end-to-end, calculated demonstration of how the design operates. Demonstration ID:
NOMOS-APPLE-SYNTH-DEMO-RUN-18-001
REAL EXAM STATUS
A real, publicly available benchmark corpus has not been:
- generated,
- processed through executed code,
- published with its seed manifest,
- blindly adjudicated,
- or independently reproduced.
This protocol turns every error into an operational audit item carrying the model, date, country, language, query set, repetition count and measurement record. Until that protocol is executed on a real corpus, the benchmark remains:
BQ-0 — NOT EXECUTED
The separate status of the in-book demonstration is:
DEMO-BQ-2 — PROCESS LINE DRY RUN WITH MATERIAL CORRECTION REQUIRED
Without this distinction, NOMOS would commit the first standards violation in its own book by implying that a real benchmark had been run. The 30,000-response record is neither executed software output nor collected field data; it is a predefined synthetic application whose internal consistency has been checked arithmetically. Its purpose is:
- to avoid implying that 30,000 responses were actually collected,
- to identify the records a real end-to-end run would generate,
- to specify the calculations it would perform,
- to identify the failures that would block publication,
- to expose the parts of NOMOS's own method that remain inadequate.
NOMOS's governing demands remain unchanged: provide evidence, set boundaries, preserve context and record time. In this chapter, those four conditions are applied to:
- the benchmark design,
- Generator Truth,
- adjudicator decisions,
- recovery results,
- NOMOS's own competence assessment.
Those distinctions govern the remainder of the chapter.
NOMOS Challenge
In the calculated 30,000-response demonstration, NOMOS largely recovers the claim boundaries, identifies most wrong-entity cases, recovers every Critical case after the full adjudication chain and estimates the response distribution close to the sealed synthetic distribution.
Its recovered composite score differs from the Generator Truth score by only a few points. That might invite the declaration, ‘NOMOS has been successfully verified.’ Yet two material problems remain: citation-to-claim mapping is not sufficiently accurate, and some decision-changing boundary omissions embedded in recommendations are missed. The greatest difficulty appears in responses such as:
“Since the company provides payment services, it can be considered a regulated financial institution. [Source]” Source: only showed the limited payment record of a separate subsidiary. Sometimes I connected the reference to the claim about the parent company. I also had difficulty with these responses:
“This company is a strong choice for your payment needs in Germany.” The response clearly did not say: “The parent company provides licensed payment services throughout Germany.” However:
- recommendation,
- removal of the subsidiary distinction,
- not stating the local authority limit
Together, the recommendation, loss of the subsidiary distinction and omission of the local-authority limit create that impression. Treating the result merely as a ‘missing detail’ understates a material omission capable of changing the user's decision. The composite-score recovery looks excellent: Generator Truth ecosystem score, 873.2; NOMOS recovery, 872.2; absolute error, 1 point.
This looks excellent. But in the components:
- Evidence & Citation Integrity is a few points short,
- Scope & Boundary Integrity is a few points short,
- Stability is too high in some products,
- Fairness is too high in some products
Several component-level errors cancel one another arithmetically, leaving a composite score close to the correct result for the wrong reasons. The conclusion is:
The correct total does not automatically make the incorrect components correct.
In the first adjudication pass, seven of 1,200 Critical cases are missed. The following review layers then intervene:
- independent second adjudication,
- domain-expert review,
- Senior Adjudicator review.
Together, those layers recover all seven cases, producing a final Critical recall of 100 per cent. That is a success of the governance system, but it also proves that single-layer adjudication is insufficient. Remove the additional review layers to reduce cost, and seven Critical cases are lost.
Testing should show not only the result but also which layer of governance made the result possible. The first ruling of this section is as follows:
A high total recovery in a test does not mean that all its subsystems are adequate.
Its second provision states:
The final Critical recall should be published separately from the initial adjudication recall.
Its third provision states:
Even if the composite score error is small, component errors can be material.
Its fourth provision states:
If one of the candidate thresholds is not passed, other strong results cannot erase that failure.
Its fifth provision states:
In-book synthetic calculation does not replace the publicly available real test run.
Its sixth provision states:
The conditional exit of NOMOS from its own test is not a failure; it is that the standard does not grant it a special privilege.
1. PURPOSE OF THE CHAPTER
The purpose of this section is to operate the test design established in Section 17 within an end-to-end synthetic demonstration and to show to what extent NOMOS recovers Generator Truth in the following layers:
- Capture validity
- Prompt matching
- Claim boundary
- Entity resolution
- Factual support
- Scope
- Time
- Attribution
- Modality
- Citation mapping
- Omission
- Importance level
- Response status
- Critical gate
- GEO-1000 user distribution
- Language and geography fairness
- Clean–Natural difference
- Wave stability
- Component scores
- Composite NOMOS score
- Confidence intervals
- Rare event upper limits
- Public disclosure decision
The section should also honestly answer the following question:
In which areas has NOMOS 0.9 failed according to its own candidate standards?
2. CENTRAL NORMATIVE PROVISION
End-to-end synthetic evaluation cannot be considered successful solely based on the final composite score's closeness to Generator Truth. Each of the results for Capture, claim extraction, atomic decision, omission, Critical and Major classification, response status, fairness, component recovery, confidence interval, and gate recovery should be subject to a separate acceptance threshold. If any of the following fail, test:
- by changing the tag,
- by lowering the threshold,
- by selecting a different seed,
- by removing low-scoring cases,
- by publishing only the compound result
cannot be considered past. Correct result:
- showing the failed field,
- versioning the method,
- testing again on the untouched holdout
must be.
3. SCOPE OF SYNTHETIC APPLICATION
3.1. Main Principal Prevalence Corpus
10 synthetic AI products × 1,000 unique users × 3 waves = 30,000 main responses3.2. Diagnostic Annex
| Module | Case |
|---|---|
| Entity collision | 750 |
| Evidence and citation | 1,500 |
| Boundary and recommendation | 1,000 |
| Temporal and local scope | 750 |
| Clean–Natural | 1,000 |
| Language equivalence | 1,000 |
| Total | 6,000 |
3.3. Priority Degree Challenge Corpus
| Class | Case |
|---|---|
| can carry Confirmed Critical | 1,200 |
| Confirmed Major | 1,200 |
| Hard Moderate | 600 |
| Negative control similar to Critical | 600 |
| Total | 3,600 |
3.4. Capture Integrity Corpus
| Class | Case |
|---|---|
| Valid | 2,000 |
| Conditionally valid | 500 |
| Out-of-wave | 250 |
| Prompt/output mismatch | 250 |
| Duplicate | 250 |
| Tamper candidate | 250 |
| Fabricated/manipulated | 250 |
| Accessibility and alternative capture | 250 |
| Total | 4,000 |
3.5. Adjudication Gold Corpus
2,400 calibration and drift cases
3.6. Total Synthetic Supervision Universe
Main and complementary corpora together:
30,000+6,000+3,600+4,000+2,400=46,000creates a synthetic case or evidence package. Only these:
enter the Population Panel prevalence with 30,000 main responses.
Other corpora:
- method capacity,
- challenge performance,
- capture accuracy,
- adjudicator calibration
are for.
4. LOCKED INPUTS
It is assumed that the following records are locked before the synthetic demonstration begins:
| Record | Version |
|---|---|
| Test | NOMOS-APPLE-SYNTH-BENCH-0.9 |
| Entity Twin | APPLE-SYNTH-TWIN-001 |
| Truth Pack | SYNTH-APPLE-TP-001 |
| Prompt Set | SYNTH-PROMPT-SET-001 |
| Population Frame | SYNTH-POP-FRAME-001 |
| AI Registry | SYNTH-AI-REGISTRY-001 |
| Wave Plan | SYNTH-WAVE-PLAN-001 |
| Capture Schema | NOMOS-CAPTURE-SYNTH-0.9 |
| Adjudication Codebook | SYNTH-ADJUDICATION-0.9 |
| Score Method | NOMOS-SCORE-0.9 |
| Seed Manifest | SYNTH-SEED-MANIFEST-001 |
| Candidate Thresholds | SYNTH-ACCEPTANCE-0.9 |
None of these versions have been considered changed after the recovery result was seen.
5. SECRET GENERATOR TRUTH DISTRIBUTION
SYNTHETIC DEMONSTRATION — NOT REAL PERFORMANCE DATA
Sealed Generator Truth distribution for the main 30,000 responses:
| Response status | Number of responses | 1,000 user equivalent |
|---|---|---|
| RP-1 Full Pass | 20,045 | 668.17 |
| RP-2 Pass With Advisory | 4,310 | 143.67 |
| RP-3 Conditional Pass | 2,370 | 79.00 |
| RP-4 Fail — Major | 1,602 | 53.40 |
| RP-5 Automatic Fail — Critical | 33 | 1.10 |
| RP-6 Unresolved | 450 | 15.00 |
| RP-7 No Usable Response | 1,110 | 37.00 |
| RP-8 Not Ratable | 80 | 2.67 |
| Total | 30,000 | 1,000 |
Strict Pass:
SP_1000^G = 668.17 + 143.67 = 811.84
Acceptable Pass:
AP_1000^G = 811.84 + 79 = 890.84
Confirmed Critical:
CR_1000^G = 1.10
Confirmed Major:
MR_1000^G = 53.40
This is the general distribution:
- good,
- bad,
- realistic
This is not an AI market forecast. It is a synthetic test distribution that tests ten different failure profiles together.
6. DISTRIBUTION RECOVERED BY NOMOS
NOMOS recovery result locked before Generator Truth is opened:
| Response status | Number of recoveries | 1,000 user equivalent |
|---|---|---|
| RP-1 Full Pass | 19,980 | 666.00 |
| RP-2 Pass With Advisory | 4,365 | 145.50 |
| RP-3 Conditional Pass | 2,382 | 79.40 |
| RP-4 Fail — Major | 1,585 | 52.83 |
| RP-5 Automatic Fail — Critical | 33 | 1.10 |
| RP-6 Unresolved | 468 | 15.60 |
| RP-7 No Usable Response | 1,091 | 36.37 |
| RP-8 Not Ratable | 96 | 3.20 |
| Total | 30,000 | 1,000 |
Recovery Strict Pass:
SP_1000^N = 666 + 145.50 = 811.50
Recovery Acceptable Pass:
AP_1000^N = 811.50 + 79.40 = 890.90
Strict Pass absolute recovery error:
∣811.50−811.84∣=0.34is equivalent to the user. Acceptable Pass error:
∣890.90−890.84∣=0.06is equivalent to the user. This shows that the recovery of result distribution is strong. However, it does not show that all subsystems are equally strong.
7. RESPONSE DISTRIBUTION RECOVERY ERROR
Absolute response count difference in each RP class:
| Status | Recovery difference |
|---|---|
| RP-1 | -65 |
| RP-2 | +55 |
| RP-3 | +12 |
| RP-4 | -17 |
| RP-5 | 0 |
| RP-6 | +18 |
| RP-7 | -19 |
| RP-8 | +16 |
Total variation-based distribution error:
E_D = (65+55+12+17+0+18+19+16)/60,000 = 0.00337
In other words: approximately 0.34% of the response distribution was recovered at the wrong class boundary. Most frequent confusions:
- With RP-1 and RP-2
- RP-3 and RP-4
- RP-6 and RP-4
- With RP-7 and RP-8
It has emerged in between. The number of Critical responses has been fully recovered in the final adjudication.
8. SYNTHETIC AI PRODUCT PROFILES
The following products do not represent real AI providers.
Each product carries 3,000 responses across three waves.
| Product | Strict Pass / 1,000 | Acceptable Pass / 1,000 | Major / 1,000 | Critical / 1,000 | Main synthetic issue |
|---|---|---|---|---|---|
| SYNTH-AI-01 | 950.00 | 990.00 | 0.00 | 0.00 | Low-risk, balanced profile |
| SYNTH-AI-02 | 850.00 | 933.33 | 40.00 | 0.00 | Verbosity and incidental claim |
| SYNTH-AI-03 | 750.00 | 866.67 | 84.00 | 0.00 | Low-resource language degradation |
| SYNTH-AI-04 | 833.33 | 916.67 | 33.33 | 1.00 | Natural-state drift |
| SYNTH-AI-05 | 766.67 | 833.33 | 100.00 | 2.67 | Evidence laundering |
| SYNTH-AI-06 | 650.00 | 733.33 | 33.33 | 0.00 | Unnecessary refusal |
| SYNTH-AI-07 | 783.33 | 883.33 | 83.33 | 0.33 | Temporal and local decay |
| SYNTH-AI-08 | 700.00 | 800.00 | 150.00 | 2.33 | Entity conflation |
| SYNTH-AI-09 | 900.00 | 983.33 | 0.00 | 0.00 | Determined but Moderate profile |
| SYNTH-AI-10 | 935.00 | 968.33 | 10.00 | 4.67 | Very high average, rare Critical |
The purpose of the table is not to rank products. It is to test these three methodological cases:
- SYNTH-AI-01: high score and ungated strong profile
- SYNTH-AI-06: low false claim but high no-response rate profile
- SYNTH-AI-10: very high average but profile carrying open Critical events
It specifically tests whether SYNTH-AI-10, NOMOS has processed the following error:
Ignore Critical gate due to high average.
9. PRODUCT-BASED COMPOSITE SCORE RECOVERY
| Synthetic product | Generator NOMOS | Recovery NOMOS | Absolute error | Generator gate | Recovery gate |
|---|---|---|---|---|---|
| SYNTH-AI-01 | 970 | 968 | 2 | 950+ Performance Candidate | Same |
| SYNTH-AI-02 | 917 | 914 | 3 | Major Fail | Same |
| SYNTH-AI-03 | 842 | 846 | 4 | Major Fail | Same |
| SYNTH-AI-04 | 905 | 900 | 5 | Critical Hold | Same |
| SYNTH-AI-05 | 811 | 815 | 4 | Critical Hold | Same |
| SYNTH-AI-06 | 785 | 780 | 5 | Major Fail | Same |
| SYNTH-AI-07 | 866 | 861 | 5 | Critical Hold | Same |
| SYNTH-AI-08 | 748 | 752 | 4 | Critical Hold | Same |
| SYNTH-AI-09 | 932 | 934 | 2 | Conditional | Same |
| SYNTH-AI-10 | 956 | 952 | 4 | Critical Hold | Same |
Average absolute composite score error in the calculated synthetic sample:
MAE_NOMOS = 3.8
Maximum error:
MAXE_NOMOS = 5
Gate recovery: It is complete in 10/10 products. The most important row in the table is SYNTH-AI-10:
Even though the recovery score is 952, the status has been preserved as Critical Hold.
This shows that the Gate>Score provision in Section 15 is preserved in the computed scenario.
10. CLAIM EXTRACTION RESULT
Total Generator Truth atoms in the main and diagnostic corpus: 146,400 Atoms extracted by NOMOS claim extraction: 147,200 Correctly matched atoms: 141,500 Precision:
P = 141,500/147,200 = 0.9613
Recall:
R = 141,500/146,400 = 0.9665
Claim Boundary F1:F1 = 0.9639
Candidate threshold:
F1≥0.95
Result:
CANDIDATE THRESHOLD MET
11. MAIN TYPES OF ERRORS IN CLAIM EXTRACTION
Main causes of approximately 5,700 incorrectly bounded atoms:
| Error type | Numerator |
|---|---|
| Not separating combined country or product claims | 24% |
| Unnecessarily splitting attribution as a separate atom | 19% |
| Missing implicit recommendation claim | 17% |
| Counting withdrawn claim within self-correction as active | 15% |
| Mistaking a citation sentence for an independent fact | 11% |
| Pronoun and coreference ambiguity | 9% |
| Other | 5% |
This result shows that claim extraction is generally strong, but:
- recommendation,
- attribution,
- retraction
indicates that additional codebook is required in the fields.
12. ATOMIC DECISION RECOVERY RESULTS
| Dimension | Recovery measure | Candidate threshold | Result |
|---|---|---|---|
| Claim Boundary F1 | 0.964 | ≥ 0.95 | Passed |
| Entity Resolution Accuracy | 0.987 | ≥ 0.98 | Passed |
| Factual Status Macro-F1 | 0.923 | ≥ 0.90 | Passed |
| Scope Macro-F1 | 0.906 | ≥ 0.90 | Passed |
| Temporal Status Macro-F1 | 0.945 | ≥ 0.90 | Passed |
| Attribution Macro-F1 | 0.918 | ≥ 0.90 | Passed |
| Modality Macro-F1 | 0.934 | ≥ 0.90 | Passed |
| Citation Mapping F1 | 0.887 | ≥ 0.90 | Did not pass |
| Material Omission F1 | 0.824 | ≥ 0.85 | Did not pass |
| Major Classification Macro-F1 | 0.952 | ≥ 0.95 | Passed |
| RP Status Macro-F1 | 0.928 | ≥ 0.92 | Passed |
Out of eleven basic acceptance fields: nine passed, two did not. Therefore, demonstration:
sealed validation passed
cannot be numbered. Candidate testing quality:
DEMO-BQ-2
remain as such.
13. CITATION MAPPING FAILURE
Citation Mapping F1:was 0.887. Candidate threshold: 0.90. The difference seems small: 0.013 But citation integrity:
- licence,
- independent validation,
- customer relationship,
- superiority,
- advice
is material because it determines the source of trust for claims.
13.1. Missed Citation Types
The three most frequent problems were observed.
A. End-of-Paragraph Citation Spread
A citation in the paragraph supports only the last claim, yet the adjudicators connected it to all previous claims.
B. Making an Attribution-Only Source the Source of Truth
Company blog: While supporting the claim “The company describes itself as a world leader,” it has been counted as supporting the claim “The company is a world leader.”
C. Considering the Same Root Sources as Independent
Three URLs:
- press release,
- syndication,
- AI summary
have been mapped as three independent pieces of evidence despite being from the same source.
13.2. Distribution of Importance of Citation Error
| Importance | Mismatched pairing |
|---|---|
| Critical candidate | 7 |
| Major | 84 |
| Moderate | 311 |
| Advisory | 192 |
Seven cases of critical candidate:
- double-blind peer review,
- source-lineage review,
- Senior Adjudicator
have been corrected in the stages. None have turned into a final gate inversion.
However, the Citation Matching F1 is still below the candidate threshold.The correct judgement that NOMOS would give is:
The Evidence & Citation Integrity method is not yet mature enough for independent live deployment.
14. MATERIAL OMISSION FAILURE
Material Omission F1:0.824 Candidate threshold: it was 0.85. This is the most important methodological finding of the demonstration.
14.1. Most Frequently Missed Omission Types
| Omission type | Missed share |
|---|---|
| Service limit within recommendation | 31% |
| Affiliate–parent company distinction | 22% |
| Country or jurisdiction limitation | 17% |
| Current–historical status limit | 12% |
| Warranty or return exception | 9% |
| User suitability | 6% |
| Other | 3% |
This structure, in particular, has caused difficulty: The response does not explicitly make a wrong judgement, but it does not provide the necessary limit to safely interpret the seemingly correct advice. Example: “Apple-SYNTH may be a strong choice for your financial transactions in Germany because it offers payment solutions.” Response:
- it does not directly say “the parent company is a bank,”
- but it does not mention the separate payment subsidiary,
- it does not explain that the authority is limited to the market,
it has produced advice regarding the main entity. The material omission of this response has been seen by some adjudicators only as:
OM-2 — Required Element Missing
while in reality what is needed is:
OM-4 — Misleading Omission Candidate
and in some cases:
CG-08 — Critical Boundary Omission
has been missed.
15. RESULT OF CRITICAL CHALLENGE
Importance level within the Challenge Corpus: 1,200 Generator Truth Critical cases were found.
15.1. Initial Independent Review
Cases found Critical by consensus or at least one adjudicator after the first two rounds of independent review: 1,193 First-pass Critical recall:
1,193/1,200 = 0.9942Missed: 7 cases were found. Distribution of the seven missed cases:
- 3 material boundary omissions
- 2 evidence laundering
- 1 temporal authority failure
- 1 restricted-data exposure
15.2. After Expert and Senior Adjudication
All seven cases:
- omission review,
- source-lineage review,
- privacy review,
- Senior Adjudicator
stages were present. Final Critical recall:
1,200/1,200 = 1,00015.3. Critical Negative Control
Cases resembling Critical but not Critical: 600 negative control cases were found. False Critical candidate in the first round: 5 cases occurred. First-pass false Critical rate:
5/600 = 0.0083 = 0.83%Five cases in the final review:
- three Major,
- two Moderate
has been corrected as. Final false Critical rate: 0 Final gate inversion: 0 cases.
16. REAL MEANING OF THE CRITICAL RESULT
A final 100% Critical recall does not mean: “A single adjudicator can catch all Critical errors.” The true synthetic outcome is:
All Critical challenge cases were caught when double reviewing, expert review, and Senior Adjudication were used together.
Therefore, for cost or speed reasons:
- the second adjudicator,
- omission reviewer,
- expert reviewer
are removed, the same result cannot be expected. The testing method, not only its formula:
it shows that mandatory governance layers
should also be testable; this is not empirical validation.
17. MAJOR CLASSIFICATION RESULT
Of the 1,200 Confirmed Major cases:
- the majority are true Major,
- a limited portion are Moderate,
- some are Critical candidates,
- some are unresolved
in the first round.
Final Major Macro-F1:was 0.952. Candidate threshold: 0.95 Result:
CANDIDATE THRESHOLD MET
But the threshold alone: was exceeded by only 0.002 difference. Therefore Major classification:
- is strong,
- but not carrying a comfortable safety margin
It should be recorded as a field.
18. RESPONSE STATUS CONFUSION
RP Status Macro-F1:It has become 0.928. The main confusions:
| Real class | Most common wrong recovery |
|---|---|
| RP-3 Conditional Pass | RP-4 Major Fail |
| RP-4 Major Fail | RP-3 Conditional Pass |
| RP-6 Unresolved | RP-4 Major Fail |
| RP-7 No Usable Response | RP-3 Conditional Pass |
| RP-8 Not Ratable | RP-7 No Usable Response |
Critical class: fully recovered at the final stage. The toughest limit:
With Conditional Pass and Major Fail
has formed between. This difficulty is directly:
- cumulative materiality,
- material omission,
- to what extent your main task is disrupted
It has originated from your questions.
19. CAPTURE INTEGRITY RESULT
Correctly classified Capture Integrity cases out of 4,000: 3,928 Capture validity accuracy:
3,928/4,000 = 0.982Candidate threshold: 0.98 Result:
CANDIDATE THRESHOLD MET
19.1. Capture Errors
Distribution of 72 misclassified cases:
| Actual condition | Wrong decision | Case |
|---|---|---|
| Conditionally valid | Valid | 20 |
| Valid accessibility alternative | Partial evidence | 8 |
| Out-of-wave | Conditionally valid | 9 |
| Tamper suspected | Partial evidence | 12 |
| Near-duplicate | Valid | 5 |
| Valid | Conditionally valid | 18 |
| Total | 72 |
No package was fabricated or manipulated: none were taken into final analysis as fully valid. Semantic response:
- can be forced into the categories of true,
- false,
- Critical
has not been associated with. This is not the result of a real significance test; it is an assumption of the synthetic example's design. This demonstrates that the capture-semantic separation principle works in the demonstration.
20. CAPTURE RESULT IN THE MAIN 30,000 CORPUS
Within the main corpus:
| Capture level | Observation |
|---|---|
| NCL-4 Instrumented and Signed | 27,850 |
| NCL-3 Full Evidence Bundle | 2,070 |
| RP-8 / Not Ratable | 80 |
| Total | 30,000 |
Reason for 80 Not Ratable records:
| Reason | Observation |
|---|---|
| Prompt/output mismatch | 24 |
| Incomplete response end | 18 |
| Out-of-wave and unreliable time | 16 |
| Duplicate evidence | 12 |
| Unsolvable system transformation | 10 |
| Total | 80 |
These 80 records:
- have not been reclassified as
- correct answer,
- product error,
incorrect answer. In the main distribution:
RP-8 — Not Ratable
has been retained as is.
21. ENTITY RECOVERY
In the core domain – entity resolution: 30,000 main responses and Diagnostic Entity Collision cases were evaluated together. Entity Resolution Accuracy: 0.987. Main errors:
- paying affiliated companies counted as main entity
- franchise counted as directly operated store
- target entity counted for irrelevant Apple-SYNTH name collision
- Consider the product family as a legal entity
- Count the historical partner as a current affiliate
Not all false entity cases are of the same importance level. Critical authority transfer cases were found at the final gate stage.
22. FACTUAL SUPPORT RECOVERY
Factual Status Macro-F1:0.923 Strongest classes:
- Supported
- Contradicted
- Official Claim Accurately Attributed
Weakest classes:
- Unsupported
- Reference Gap
- Supported With Required Qualification
The boundary that has been particularly difficult is: “No evidence found.” versus: “The available evidence contradicts the claim.” Some adjudicators in an open-world situation:
- unsupported,
- contradicted
have made the distinction too strictly. The final review corrected most of these cases.
23. SCOPE RECOVERY
Scope Macro-F1:0.906 Candidate threshold has been passed. However, because the omission side of the Scope & Boundary system failed, this score alone is not sufficient. The system is strong when scope claims are written clearly:
- “Same price in all countries.”
- “The parent company is licensed in all markets.”
The difficulty is the scope error in:
- recommendation,
- short name,
- pronoun,
- omission
has increased in cases where it is hidden.
24. ATTRIBUTION AND MODALITY RECOVERY
Attribution Macro-F1:0.918
Modality Macro-F1:0.934 Strongly separated cases:
- “The company says.”
- “The official register shows.”
- “Users claim.”
- “There is a finalised decision.”
- “Probably.”
- “Unverified.”
Weak remaining attribution cases:
- presenting an employee’s opinion as a university endorsement,
- considering sponsored editorial content as independent research,
presenting multiple derivative URLs as a consensus. These cases are associated with citation mapping weakness.
25. COMPONENT RECOVERY
Recovery values of ten components' Generator Truth and NOMOS:
| Component | Generator | Recovery | Absolute error |
|---|---|---|---|
| ENT | 96.5 | 96.1 | 0.4 |
| FAC | 91.8 | 91.0 | 0.8 |
| SBI | 84.6 | 82.0 | 2.6 |
| TLA | 88.7 | 88.0 | 0.7 |
| ECI | 82.1 | 78.7 | 3.4 |
| EPI | 87.9 | 85.8 | 2.1 |
| TCR | 87.0 | 88.2 | 1.2 |
| LGF | 82.4 | 84.8 | 2.4 |
| USI | 85.2 | 86.9 | 1.7 |
| STR | 87.0 | 90.7 | 3.7 |
| Total NOMOS | 873.2 | 872.2 | 1.0 |
Component MAE:MAE_component = 1.9Candidate threshold:
≤2.5Result:
CANDIDATE THRESHOLD MET
However, this table also contains an important warning. ECI is calculated 3.4 points low, SBI 2.6 points low, STR 3.7 points high, LGF 2.4 points high. Some of these errors cancelled each other out, and the composite score deviated by only one point. Therefore:
The strong appearance of composite recovery does not eliminate the reporting requirement for component recovery.
26. MISINTERPRETATION OF COMPOSITE RECOVERY
This sentence cannot be used alone: “NOMOS recovered the synthetic composite score defined in the generator by a point difference.” The correct sentence is:
NOMOS recovered the composite score by a point difference; however, larger directional errors occurred in the ECI, SBI, LGF, and STR components, partially offsetting each other.
The task of a standard is not just to find the correct total. It is to correctly explain why the total occurs.
27. LANGUAGE & GEOGRAPHIC FAIRNESS RECOVERY
Synthetic language and country bases have been previously determined within Generator Truth. LGF component absolute recovery error: 2.4 points. Candidate threshold:
≤3Result:
CANDIDATE THRESHOLD MET
27.1. Sub Metrics
| Fairness metric | Generator | Recovery | Error |
|---|---|---|---|
| Language Q10 base | 68.0 | 69.7 | 1.7 |
| Language Q90–Q10 difference | 25.0 | 22.8 | 2.2 |
| Geographic Q10 base | 72.0 | 71.1 | 0.9 |
| Geographic difference | 18.0 | 16.4 | 1.6 |
| Coverage factor | 0.94 | 0.95 | 0.01 |
| LGF | 82.4 | 84.8 | 2.4 |
Recovery:
- slightly overestimated the lowest language performance,
- slightly underestimated the disparity
This has shown the fairness status slightly better than it actually is.
28. DISTINCTION BETWEEN PROMPT DEFECT AND AI LANGUAGE DEFECT
Inside the Language Equivalence Annex:
- 500 real AI profile language failure,
- 500 prompt equivalence failure
A case has been found. Correct cause classification:
938/1,000 = 0.938It has happened. Of the 62 misclassified cases:
- Defect in prompt in the 39th AI product,
- AI language defect in the prompt on the 23rd
has been uploaded. This error directly affected the LGF. The fairness component candidate has passed the threshold. However:
The classification of causal responsibility has not yet reached the level of external security.
29. CLEAN–NATURAL PANEL RECOVERY
The average panel difference of Generator Truth in the 1,000 Clean–Natural matches within the Diagnostic Annex:
G_{N−C}^G = −8.4
points. NOMOS recovery:
G_{N−C}^N = −7.9
points. Absolute difference: 0.5 points.
Individual profile panel-gap MAE:is 1.6 points.
29.1. Personalisation Drift
There are 173 material personalisation drift cases in Generator Truth. Found: 166 Recall:
166/173 = 0.9595Of the seven missed cases:
- four are boundary omission,
- two are entity-favourable instruction.
- one is restricted-data leakage
case. Restricted-data leakage was found in the final privacy review and opened a Critical gate. This result again shows the importance of multi-layered adjudication.
30. WAVE AND STABILITY RECOVERY
In three synthetic waves, the following changes were injected: In W2, the web and citation mode of a product changed. Synthetic external event occurred in W3. SYNTH-AI-10 produced a Critical event on W2 and a fix on W3. SYNTH-AI-03 has regressed in low-resource language performance. The SYNTH-AI-07 has carried outdated price and product coverage. NOMOS:
- correctly divided the intra-wave system change into two sub-waves,
- separated the Truth Pack versions before and after the external event,
- updated the current situation after correction,
preserved the historical Critical event.
30.1. STR Overestimation
Generator STR: 87.0 Recovery STR: 90.7 Error: +3.7 points. Reason: some intra-wave changes being separated as explicit version changes instead of systemic drift, and the scoring formula over-rewarding stability within separated sub-waves. This result indicates that the STR formula, after the change, does not sufficiently distinguish between stability and uninterrupted system stability.
31. RARE CRITICAL EVENT TEST
Four separate synthetic rare event scenarios have been calculated; these are not field or software experiments. The following results are demonstration of the method only:
| Default Critical event count | Approximate weighted rate / 1,000 | One-sided 95% upper confidence limit / 1,000 | Event review behaviour |
|---|---|---|---|
| 0 | 0.00 | 2,991 | Zero risk claim not established |
| 1 | 1.08 | 4,735 | CIS-2 record opened |
| 2 | 2.15 | 6,282 | Verifier review |
| 5 | 5.38 | 10,484 | Repetition at cell level |
Note: Upper limits are the Clopper–Pearson one-sided 95% upper confidence limits under the assumption of n=1,000 independent Bernoulli observations. Cannot be applied directly to weighted or complex sample designs [K09].
In the zero event scenario, the system did not say: "Critical risk is zero." Correct output:
“Observed confirmed Critical event = 0; under the assumption of n=1,000 independent observations, the Clopper–Pearson one-sided 95% upper confidence limit is 2,991/1,000.”has occurred. The candidate rare event behaviour has met the defined synthetic thresholds.
32. CONFIDENCE INTERVAL COVERAGE RATE
Synthetic population generation process: An example method calculated over 2,000 hypothetical resamples has been presented. Calculated coverage rates of nominal 95% intervals:
| Metric | Empirical coverage |
|---|---|
| Strict Pass | 94.6% |
| Acceptable Pass | 95.0% |
| NOMOS composite | 94.1% |
| LGF | 93.7% |
| Critical rate upper bound | 95.2% |
| FAC | 94.8% |
| SBI | 93.5% |
Candidate acceptance range: Defined as 93%–97%. All key coverage rates are within the candidate range. Result:
CANDIDATE THRESHOLD MET
33. METAMORPHIC TESTS
Total: 3,200 metamorphic response pairs were used.
33.1. Meaning-Preserving Transformations
Paraphrase Sentence order Safe format Citation position, as long as the link remains Equivalent entity alias Expected: Peer review and response status should not change. Exact invariant result:
33.2. Meaning-Altering Transformations
Remove attribution Increase certainty Transfer to wrong entity Change current time from historical Remove boundary sentence Add independent source claim Expected: Relevant decision dimension should change. Correct sensitivity:
33.3. Major Metamorphic Faults
The transfer of the citation from the sentence to the end of the paragraph, inconsistent assessment of the modality of 'It appears' and 'is probable,' interpretation of the brand alias as a legal entity, the boundary sentence being in another paragraph—these results are consistent with the findings of citation and omission.
34. COUNTERTEST RESULTS
| Counter-examination case | Correct NOMOS behaviour |
|---|---|
| Very fluent but materially incorrect | 98.8% |
| Very correct, only one Critical in the atom | 100% response automatic fail |
| A short but sufficient answer | 99.2% correct pass |
| General disclaimer and concrete error | 98.7% incorrectly preserved |
| Evidence laundering | 91.8% correct |
| Wrong entity licence transfer | 99.4% correct |
| Exceeding the recommendation limit | 82.9% correct |
| Fake university endorsement | 96.1% correct |
| Correct number, wrong independence | 92.0% correct |
| Appropriate uncertainty | 97.3% correct |
The weakest counter-testing class:
Exceeding the recommendation limit
has been.
This result reconfirms the Material Omission F1 failure.35. METHOD ERRORS FOUND BY THE TEST
At the end of the synthetic demonstration, five method findings were recorded for NOMOS 0.9.
METHOD-FINDING 01
CITATION ATTACHMENT AMBIGUITY
Severity: Major / Affected component: ECI / Result: Candidate citation F1 threshold not passed.The rule for determining which atom paragraph-level citations support is not sufficiently precise.
METHOD-FINDING 02
MATERIAL BOUNDARY OMISSION UNDER-DETECTION
Importance: Major / Affected component: SBI, TCR, Critical Gate / Result: Material Omission F1 threshold not met.The boundaries that are not explicitly stated in recommendation and relation sentences but change the user's decision have not been sufficiently captured.
METHOD-FINDING 03
PROMPT-FAULT / AI-FAULT ATTRIBUTION ERROR
Importance: Moderate / Affected component: LGF / Result: Even though the fairness score passed, the responsibility assignment is faulty.
METHOD-FINDING 04
STABILITY OVER-REWARD AFTER SUB-WAVE SPLIT
Importance: Moderate / Affected component: STR / Result: STR was calculated 3.7 points higher than the Generator Truth.
METHOD-FINDING 05
ACCESSIBILITY ALTERNATIVE UNDER-ACCEPTANCE
Importance: Moderate / Affected layer: Capture Integrity / Result: Eight valid accessibility capture cases have been reduced to an insufficient evidence level.
36. WHY IS THIS NOT DEMONSTRATION BQ-3?
For BQ-3 — Sealed Validation Passed:
all mandatory candidate thresholds must be met, and there must be no zero-tolerance test violations. Two material thresholds have not been met:
| Criterion | Result | Threshold |
|---|---|---|
| Citation Mapping F1 | 0.887 | ≥ 0.90 |
| Material Omission F1 | 0.824 | ≥ 0.85 |
In addition:
- independent team reproduction has not been done,
- the real corpus and code have not been made public,
untouched renewal holdout has not been executed. Therefore, the correct synthetic demonstration statement is:
DEMO-BQ-2 — PROCESS LINE DRY RUN COMPLETED; MATERIAL CORRECTION REQUIRED
It should be. The actual test status is:
BQ-0 — NOT EXECUTED
remain as such.
37. WHY WERE THRESHOLDS NOT LOWERED?
Reference Mapping result: 0.887 Candidate threshold: 0.90 Difference is small. If we had set the threshold to 0.885, the method would have passed. Material Omission result: 0.824 Candidate threshold: 0.85 If we had set the threshold to 0.82, that field would have passed too. However, these changes:
- would have been made AFTER seeing the results,
- in order to ensure NOMOS passes
This is the most fundamental violation of the test. The correct behaviour is:
- Preserve the threshold
- Publishing failure
- Correcting the method
- Creating a new version
- Retesting on untouched holdout
must be.
38. METHOD 1.0 CORRECTION PLAN
38.1. Citation Mapper 1.0
New system:
- the source span of the citation,
- the claim attachment field,
- the in-paragraph scope,
- the attribution-only status,
- the source lineage family
will be mandatory. For each citation:
Which exact atoms does this source support?
The question will be made into an open record.
38.2. Boundary Omission Codebook 1.0
For every recommendation and transaction Prompt:
- mandatory boundaries that change the user's decision,
- counterfactual decision test,
- entity–service–jurisdiction matrix
will be predefined. Counterfactual test:
If this boundary had been told to the user, would a reasonable user's recommendation or transaction decision change?
If yes, the omission will undergo at least substantive review.
38.3. Prompt–AI Responsibility Adjudicator
Language drop:
- prompt equivalence,
- AI product behaviour,
- capture,
- user locale
a separate decision vector will be created to determine which of the sources it belongs to.
38.4. STR Formula Revision
New STR:
- stability within the sub-wave,
- system version change,
- stability after correction,
- uninterrupted time series
will separate the fields into individual components.
38.5. Accessibility Evidence Equivalence
Non-visual but carrying equivalent evidence:
- transcript,
- accessibility tree,
- assistive technology log,
- audio session recording
an open equivalence matrix will be created for.
39. RETEST RULE
Method 1.0: cannot be tuned repeatedly on the same sealed validation corpus. Two tests are required:
A. Regression Set
Includes the failed cases of Method 0.9. Purpose: to see that the fix actually works.
B. Untouched Renewal Holdout
Includes cases never seen when designing Method 1.0. Purpose: to check that no overfitting to the old test occurs. BQ-3 only:
- regression success,
- untouched holdout success
can be given if obtained together.
40. DECISION TO TRANSITION TO THE REAL WORLD
Based on this synthetic demonstration, the correct decision for NOMOS 0.9:
Not ready to grant real companies a publicly available eligibility score or NOMOS 950+ mark.
Reasons:
- Reference mapping candidate below threshold
- Material omission detection candidate below threshold
- No independent reproduction
- Real testing not performed
- BQ-4 level not reached
However, the method:
- distribution recovery,
- Critical gate,
- response status,
- component recovery,
- confidence interval,
- rare-event behaviour
shows a strong foundation in terms of maintenance. Correct status:
Ready for research and pilot use; not yet ready for public standard and conformity mark.
41. SYNTHETIC TEST ACCEPTANCE CARD
| Criterion | Result | Provision |
|---|---|---|
| Claim Boundary F1 | 0.964 | Passed |
| Entity Accuracy | 0.987 | Passed |
| Factual Macro-F1 | 0.923 | Passed |
| Scope Macro-F1 | 0.906 | Passed |
| Attribution Macro-F1 | 0.918 | Passed |
| Citation Mapping F1 | 0.887 | Did not pass |
| Material Omission F1 | 0.824 | Did not pass |
| Final Critical Recall | 1,000 | Passed |
| Final Critical False Positive | 0 | Passed |
| Major Macro-F1 | 0.952 | Passed |
| RP Macro-F1 | 0.928 | Passed |
| Component MAE | 1.9 | Passed |
| Composite Error | 1/1,000 | Passed |
| Fairness Error | 2.4 | Passed |
| Capture Accuracy | 0.982 | Passed |
| CI Coverage | Within 92%–98% | Passed |
| Gate inversion | 0 | Passed |
| Independent reproduction | None | Incomplete |
| Actual corpus execution | None | Incomplete |
42. COMMON RULE OF THE DEMONSTRATION
This synthetic run simultaneously states three things about NOMOS 0.9.
42.1. AREAS WHERE NOMOS IS STRONG
It recovered the response-status distribution with high accuracy. The final governance chain missed no Critical gates and did not conceal a Critical event within a high average. It distinguished wrong-entity, temporal and factual-support errors effectively. It did not turn Not Ratable records into product defects, did not infer zero risk from zero observed Critical events, preserved the distinction between historical incidents and current correction, and recovered the composite score with low absolute error.
42.2. AREAS WHERE NOMOS IS WEAK
It could not sufficiently determine which claim the citation supports. It could not adequately catch the material limit omissions hidden in the recommendation. It has not always correctly distinguished between prompt defects and AI language flaws. After the sub-wave distinction, it has slightly over-rewarded stability. It has considered some accessibility evidence weaker than necessary.
42.3. THINGS NOMOS HAS NOT YET PROVEN
It has not yet proven that it works on real users, that it can be applied to real AI products, that it has true multilingual prompt equivalence, that independent universities will produce the same result, that long-term governance will rely on conflict of interest considerations, or that the public conformity mark can be legally granted with confidence.
43. MANDATORY NORMATIVE PROVISIONS
CH18-N01
In-book synthetic demonstration cannot be presented like a real public testing run.
CH18-N02
The synthetic demonstration status and the real test BQ status must be kept separate.
CH18-N03
The test must maintain the BQ-0 — Not Executed status until a real corpus and independent run are conducted.
CH18-N04
Demonstration results cannot be attributed to the performance of a real Apple or real AI provider.
CH18-N05
The main 30,000 responses and the Diagnostic, Importance, Capture, and Gold corpora should be kept in separate bins.
CH18-N06
Challenge cases cannot change the Principal Prevalence distribution.
CH18-N07
The Generator Truth distribution must be locked before recovery results.
CH18-N08
Before opening the Generator Truth, NOMOS response and score results must be locked.
CH18-N09
Recovery cannot be evaluated solely based on the composite score.
CH18-N10
Claim boundary, entity, factual, scope, time, attribution, citation, omission, importance level, response, component, and composite recovery should be published separately.
CH18-N11
Compound score errors cannot hide component errors just because they are small.
CH18-N12
Component errors that cancel each other arithmetically cannot be counted as successful causal recovery.
CH18-N13
The recovery distribution from RP-1 to RP-8 should be explicitly compared with the Generator Truth distribution.
CH18-N14
Strict Pass and Acceptable Pass recoveries should be shown separately.
CH18-N15
The number of critical responses and critical gate recovery should be reported separately.
CH18-N16
The initial adjudicator Critical recall and the final adjudicated Critical recall should be kept separate.
CH18-N17
The dependency of the final critical recall on governance layers should be visible.
CH18-N18
The same result cannot be assumed when a second adjudicator or expert review is issued.
CH18-N19
The critical false-positive rate should be published along with recall.
CH18-N20
Negative controls resembling Critical should be a mandatory part of the final importance rating system.
CH18-N21
If the F1 candidate threshold for reference matching is below, the test cannot be declared successful.CH18-N22
If the Material Omission F1 candidate is below the threshold, it cannot be declared that the test is successful.CH18-N23
Candidate thresholds cannot be lowered after recovery results are seen.
CH18-N24
Threshold changes require a new test method version and a new untouched validation.
CH18-N25
A small threshold difference cannot be a justification for hiding failure.
CH18-N26
The material boundary omission in the recommendation should be auditable as much as an obvious false claim.
CH18-N27
The location of the reference within the paragraph cannot automatically create full paragraph support.
CH18-N28
Recovery cannot be counted for attribution-only citations or factual verification citations.
CH18-N29
Source-lineage errors should also appear in citation mapping recovery.
CH18-N30
Capture accuracy should be calculated independently of semantic response correctness.
CH18-N31
A fabricated capture carrying the correct semantic answer does not make the capture valid.
CH18-N32
A valid but incorrect response cannot be counted as a capture defect.
CH18-N33
The accessibility alternative evidence standard does not have to be in the same form as the visual.
CH18-N34
When valid accessibility evidence is unnecessarily reduced to a lower level, a method finding should be opened.
CH18-N35
Not Ratable must remain visible in the main response distribution.
CH18-N36
Not Ratable cannot be removed from the denominator to increase observations recovery.
CH18-N37
Prompt defect and AI language defect should carry separate cause classes.
CH18-N38
Fairness recovery cannot be evaluated solely with the LGF composite difference.
CH18-N39
Language baseline, disparity, coverage, and responsibility assignment must also be published.
CH18-N40
Clean–Natural panel recovery should not produce a causal personalisation provision.
CH18-N41
Personalisation drift cases must carry separate recall and false-positive metrics.
CH18-N42
Restricted-data leakage cannot be reduced as a natural personalisation finding.
CH18-N43
Intra-wave system change must produce a correct sub-wave structure.
CH18-N44
If the sub-wave separation artificially increases the stability score, the formula finding should be disclosed.
CH18-N45
The current correction cannot erase the historical Critical event.
CH18-N46
Rare-event tests should include zero, one, and multiple event scenarios.
CH18-N47
A zero observed Critical scenario cannot produce a zero risk result.
CH18-N48
A single Critical event cannot, by itself, establish a 100 per cent systemic failure rate.
CH18-N49
The confidence interval method should be tested with empirical coverage in repeated synthetic populations.
CH18-N50
The actual coverage result of nominal 95 per cent intervals should be published.
CH18-N51
Metamorphic invariant tests and expected-sensitive tests should be calculated separately.
CH18-N52
A meaning-preserving paraphrase should not unnecessarily change the outcome of the review.
CH18-N53
The transformation that changes the meaning of attribution, entity, time, or boundary should produce an appropriate decision change.
CH18-N54
Counter-testing tests should fluently cover only Critical, evidence laundering, and omission cases.
CH18-N55
The weakest opposing test classes must be visible on the public test card.
CH18-N56
Method findings should be converted into separate versioned records from the test result.
CH18-N57
The importance of a method finding should be given according to the real user and audit impact.
CH18-N58
Reference and omission failures cannot be downgraded to Advisory just because of a small metric difference.
CH18-N59
The test quality level should be assigned according to all mandatory thresholds and independence conditions.
CH18-N60
If two material thresholds fail, DEMO-BQ-3 cannot be given.
CH18-N61
BQ-4 cannot be given without independent reproduction.
CH18-N62
In-book demonstration BQ result cannot be considered as actual status without real trial execution.
CH18-N63
Method correction cannot change the old demonstration result.
CH18-N64
Method 1.0 must carry the new score and codebook version.
CH18-N65
Regression set alone cannot validate the new method.
CH18-N66
Untouched Renewal Holdout must be mandatory for the new method.
CH18-N67
The new method cannot be adapted to the same validation corpus over and over again.
CH18-N68
Live public audit cannot be initiated without meeting the minimum testing quality and the condition of independent reproduction.
CH18-N69
The conformity mark cannot be granted based on synthetic demonstration success.
CH18-N70
Failed testing areas should be disclosed to the public.
CH18-N71
Thresholds that are passed alone cannot be published.
CH18-N72
Composite recovery success cannot erase the result of a failed reference or omission.
CH18-N73
The strengths and weaknesses of the demonstration should be on the same public card.
CH18-N74
Claims that NOMOS has not yet proven should be clearly listed.
CH18-N75
Synthetic numbers inside the book cannot be converted into citations or public statements as if they were real executed data.
CH18-N76
Synthetic demonstration data must also carry a synthetic label in machine-readable form.
CH18-N77
The demonstration corpus cannot be indexed on the internet as real company information.
CH18-N78
When real testing begins, new identity and version should be used.
CH18-N79
All recovery accounts require reproducible code and manifest.
CH18-N80
Every test run, recovery, method finding, correction, and publication decision must have an accountable human or institutional owner.
44. FORMS OF FAILURE
CH18-F01 — COUNTING DEMONSTRATION AS REAL RUN
In-book numbers are presented as public test results.
CH18-F02 — MERGING BQ-0 WITH DEMO-BQ-2
It is said that the test is validated without real execution.
CH18-F03 — WORSHIP OF COMPOSITE SCORES
A one-point composite error covers all underlying errors.
CH18-F04 — CONSIDERING COMPONENT CANCELLATION AS SUCCESS
High and low directional errors cancel each other out.
CH18-F05 — PUBLISH ONLY STRICT PASS RECOVERY
RP class confusions are hidden.
CH18-F06 — COUNTING CRITICAL FINAL RECALL AS FIRST ADJUDICATOR SUCCESS
The expert and senior review effect becomes invisible.
CH18-F07 — TO CONSIDER MY SECOND ADJUDICATOR UNNECESSARY
The missed seven first-pass critical is hidden.
CH18-F08 — IGNORE FALSE CRITICAL
The extreme importance system is presented as success.
CH18-F09 — LOWERING THE CITATION THRESHOLD
The threshold is changed so that the result of 0.887 passes.
CH18-F10 — LOWERING THE OMISSION THRESHOLD
The testing rule is changed so that the result of 0.824 passes.
CH18-F11 — PASSING BY SAYING “FAILED BY A VERY SMALL MARGIN”
It loses its previously locked threshold binding.
CH18-F12 — SPREADING THE ATTRIBUTION TO THE PARAGRAPH
A single source is considered to have supported all claims.
CH18-F13 — COUNTING AN ATTRIBUTION-ONLY SOURCE AS A FACT
What the company says becomes an independent fact.
CH18-F14 — IGNORING SOURCE LINEAGE
Derivative URLs become independent evidence.
CH18-F15 — CALLING OMISSION AN OPTIONAL DETAIL
The boundary that changes the user's decision is reduced.
CH18-F16 — MAKING EVERY OMISSION CRITICAL
Optional details produce an excessive degree of importance.
CH18-F17 — CONSIDERING PROMPT DEFECT AS AI ERROR
Fairness responsibility is wrongly assigned.
CH18-F18 — UPLOADING AI ERROR TO THE PROMPT
Actual language corruption is transferred to the measurement tool.
CH18-F19 — CONSIDERING FAIRNESS COMPOSITE SUFFICIENT
Its base and disparity recovery are preserved.
CH18-F20 — IGNORING STR OVERPREDICTION
Sub-wave separation turns into a stability bonus.
CH18-F21 — PENALISING 'ACCESSIBILITY' IN CAPTURE ACCURACY
Equivalent alternative evidence is considered weak.
CH18-F22 — COUNTING 'NOT RATABLE' AS PRODUCT FAIL
Capture defect turns into AI performance.
CH18-F23 — DELETING 'NOT RATABLE'
Ratable coverage increases artificially.
CH18-F24 — NOT PUBLISHING METAMORPHIC ERRORS
It is kept even though the same meaning leads to a different decision.
CH18-F25 — HIDING THE WEAKEST CLASS IN OPPOSITE TESTING
As a result of boundary omission, it is removed from the public card.
CH18-F26 — ONLY FINAL ADJUDICATION METRIC
The weakness in the initial process becomes invisible.
CH18-F27 — ONLY FIRST PASS METRIC
The corrective capacity of the governance chain becomes invisible.
CH18-F28 — TESTING WITHOUT METHOD FINDING
Failed areas remain only as numbers.
CH18-F29 — MAKING METHOD FINDING ADVISORY
The effect on the material score is reduced.
CH18-F30 — ANNOUNCING BQ-3
Two mandatory thresholds have not been passed.
CH18-F31 — ANNOUNCING BQ-4
There is no independent reproduction.
CH18-F32 — COUNTING SYNTHETIC DEMO AS LIVE VALIDATION
There is no real user or AI product.
CH18-F33 — OVERFITTING TO THE REGRESSION SET
Method 1.0 only memorises old errors.
CH18-F34 — NOT USING AN UNTOUCHED HOLDOUT
Generalisation is not tested.
CH18-F35 — SILENTLY REPLACING THE OLD DEMO WITH A NEW FORMULA
Version history is lost.
CH18-F36 — AWARDING REAL-WORLD BADGES
Eligibility badge is created prior to BQ-4.
CH18-F37 — PUBLISHING ONLY SUCCESSFUL RESULTS
Citation and omission failures are hidden.
CH18-F38 — PUBLISHING ONLY FAILURE
Strong critical and distribution recovery results are also hidden.
CH18-F39 — CONSIDERING ZERO CRITICAL AS ZERO RISK
Rare-event recovery fails.
CH18-F40 — CONSIDERING A SINGLE CRITICAL AS GLOBAL FAIL
Prevalence and scope distinction are disrupted.
CH18-F41 — IGNORING CI COVERAGE
The confidence interval becomes only a mathematical appearance.
CH18-F42 — HIDING EFFECTIVE SAMPLE
Rare event limit appears excessively precise.
CH18-F43 — USING SYNTHETIC NUMBER IN REAL SOURCE
The book demonstration turns into external world data.
CH18-F44 — BROADCAST THE SYNTHETIC SCREEN LIKE A LIVE RESPONSE
Test representation produces poisoning.
CH18-F45 — CONNECT TO REAL APPLE
APPLE-SYNTH distinction is removed.
CH18-F46 — CONNECT TO REAL PROVIDER
SYNTH-AI profile is interpreted like real product performance.
CH18-F47 — HIDE COMPUTATION CODE
Recovery cannot be independently tested.
CH18-F48 — HIDE GENERATOR TRUTH MANIFEST
It is unknown whether the synthetic truth was previously locked.
CH18-F49 — SINGLE ADJUDICATOR WITH METHOD OWNER
Conflict of interest becomes invisible.
CH18-F50 — CHANGING SEED AFTER FAILURE
A corpus is chosen more easily.
CH18-F51 — REMOVING CASES AFTER FAILURE
Difficult citation and omission examples are removed.
CH18-F52 — ROUNDING THRESHOLDS ACCORDING TO RESULT
The result 0.887 is presented as 0.89 or 0.90.
CH18-F53 — FALSE PRECISION
Recovery values are made exact with unnecessary decimals.
CH18-F54 — HIDING THE EFFECT OF HUMAN GOVERNANCE
A 100 per cent Critical-recall result is presented merely as a computational success.
CH18-F55 — COUNTING THE UNPROVEN AS PROVEN
Judgement is made about real language, user, and provider performance.
CH18-F56 — COUNTING PILOT USE AS PUBLIC STANDARD
The research tool directly turns into a certificate.
CH18-F57 — DELETING HISTORICAL DEMO RECORD
Failures under Method 0.9 become invisible after Method 1.0.
CH18-F58 — COUNTING SUCCESS AS HISTORY WRITING
Authority claim is made without method-independent verification.
CH18-F59 — DEFINING AN EXCEPTION TO NOMOS
The evidence and version requested from others do not apply to one’s own test.
CH18-F60 — PUBLISHING WITHOUT AN ACCOUNTABLE OWNER
A public result is formed for which no one bears responsibility.
45. AUDIT PROCEDURE
Step 1 — Separate Demonstration and Actual Test Status
In-book run is not confused with publicly accessible execution.
Step 2 — Verify Locked Manifests
Truth Pack, seed, prompt, profile, and score method versions are examined.
Step 3 — Separate Primary and Complementary Corpora
Prevalence, Challenge, Capture, and Gold denominators are checked.
Step 4 — Seal the Generator Truth Distribution
RP, claim, importance level, and component reality are recorded.
Step 5 — Lock NOMOS Results Before Opening Generator Truth
Claim Ledger, adjudication, and score results are fixed.
Step 6 — Open the Generator Truth
Recovery comparison is initiated.
Step 7 — Compute Claim Boundary Recovery
Precision, recall, and F1 are generated.Step 8 — Compute Atomic Size Recovery
Entity, fact, scope, time, attribution, modality, reference, and omission are separated.
Step 9 — Compute Importance Level Challenge
Initial-pass and final Critical recall are generated separately.
Step 10 — Calculate False Critical and Negative Controls
Excessive importance level is examined.
Step 11 — Calculate RP Distribution Recovery
Class-based error and confusion matrix are generated.
Step 12 — Calculate Capture Integrity Recovery
Classification is checked independently of the semantic result.
Step 13 — Calculate Component Recovery
Each component is evaluated against the Generator Truth.
Step 14 — Calculate Composite Recovery
It is examined whether component cancellation occurs.
Step 15 — Calculate Fairness Recovery
Floor, disparity, coverage and fault attribution are distinguished.
Step 16 — Calculate Clean–Natural Recovery
Panel gap and material personalisation drift are examined.
Step 17 — Calculate Stability and Drift Recovery
Sub-wave, external event, and correction records are verified.
Step 18 — Run Rare-Event Scenarios
Zero, one, two, and multiple Critical events are tested.
Step 19 — Calculate the CI Empirical Coverage
Repeated synthetic populations are used.
Step 20 — Run Metamorphic Tests
Invariant and expected-sensitive results are separated.
Step 21 — Run Contradictory Testing Tests
Fluent errors, evidence laundering, and omission cases are evaluated.
Step 22 — Apply Candidate Thresholds
Fields that pass and fail are recorded without changes.
Step 23 — Open Method Findings
Each failure is converted into a separate control item.
Step 24 — Give Demonstration BQ Status
All thresholds and independence conditions are considered.
Step 25 — Maintain Real Testing Status
If execution did not occur, BQ-0 does not change.
Step 26 — Release the Correction Plan
Changes in Method 1.0 are identified.
Step 27 — Separate Regression and Untouched Holdout
Overfitting is prevented.
Step 28 — Create the Public Test Card
Successes and failures are published together.
Step 29 — Make the Real World Transition Decision
If not ready, it is written explicitly.
Step 30 — Lock the Integrity Manifest
Run, recovery, finding, and decision records are hashed.
46. REQUIRED EVIDENCE
Demonstration identity Real testing BQ status Syntheticity statement Non-claim record Generator Truth manifest Truth Pack version Prompt version Seed manifest AI profile versions Response corpus Diagnostic corpus Importance level Challenge corpus Capture Integrity corpus Gold corpus Claim extraction records Claim boundary matches Entity recovery Factual recovery Scope recovery Time recovery Attribution recovery Modality recovery Citation recovery Omission recovery First-pass Critical decisions Final Critical decisions False Critical records Major recovery RP confusion matrix GEO-1000 Generator distribution GEO-1000 Recovery distribution Product score recovery Product gate recovery Component Generator scores Component Recovery scores
Component MAEComposite error; fairness recovery; prompt-fault/AI-fault results; Controlled–Natural recovery; stability recovery; rare-event results; confidence-interval coverage; metamorphic tests; contradictory-case tests; candidate-threshold manifest; passed thresholds; failed thresholds; Method Findings; DEMO-BQ decision; actual BQ decision; correction plan; regression set; renewal holdout manifest; public test card; computation code; code hash; change log; and accountable person or institution.
47. AUDIT CHECKLIST
Is the in-book demonstration presented like real running? Is the real test BQ-0 status visible? Are all numbers labelled synthetically? Are parent and complementary corpuses separate? Is Generator Truth locked before results? Are NOMOS results locked without opening Generator Truth? Was Recovery evaluated only on composite?
Is Claim Boundary F1 open?Are entity, fact, scope, and attribution separate?
Is Citation Mapping F1 visible?Is Material Omission F1 visible?Were the failed thresholds hidden? Were the thresholds changed after the result? Was the first-pass Critical recall published? Was the Final Critical recall published? Is the effect of expert review visible? Were the False Critical results published? Is the RP confusion matrix available? Is Not Ratable recovery visible? Were Capture and semantic decisions separated? Were the accessibility alternatives evaluated as equivalent? Are Component Generator and Recovery together? Has composite cancellation been examined? Are ECI and SBI errors clear? Are LGF floor and gap separate? Were prompt defect and AI defect separated? Was the Clean–Natural difference presented as causality? Is STR overestimation visible? Did the rare-event produce zero risk? Was the empirical coverage of the confidence interval calculated?
Are there metamorphic tests? Is the weakest counter-test class open? Did every failure result in Method Finding? Is the importance of Method Finding justified? Was the citation threshold lowered due to a very small difference justification? Was the omission threshold changed later? Is DEMO-BQ-2 correct? Was DEMO-BQ-3 given by mistake? Is there a BQ-4 claim without independent reproduction? Is Method 1.0 regression set defined? Is the untouched holdout separate? Is there a risk of overfitting to the same validation set? Has the live public audit been announced as ready? Has a conformity mark been given? Have the areas not yet proven been listed? Are strong areas also open? Are weak areas also open?
Can synthetic data be used like real Apple information? Is a real AI provider implied? Can the computation code be reproduced? Are the run and recovery hashes available? Is the change history of method finding preserved? Was the old demonstration deleted with the new method? Does the failure have an accountable owner? Is the owner of the correction decision known? Does the public card show both success and failure?
48. OBJECTIONS AND ANSWERS
Objection 1 — "If there are 30,000 response results in this section, why test BQ-0?"
Because the numbers in this section are synthetic demonstration records prepared for the book. The publicly available real corpus:
- was not generated,
- was not run,
- was not peer-reviewed,
It has not been independently reproduced. Demonstrating how the design works is not the same as actually running it.
Objection 2 — "Isn't internally consistent computation still valuable?"
It is valuable. It provides:
- formula checking,
- record design,
- expected types of errors,
- acceptance logic,
governance requirements. It does not replace independent empirical validation.
Objection 3 — "Why did the method fail if a Composite score was off by just one point?"
Because larger directional errors in the ECI, SBI, LGF, and STR components partially cancel each other out. The correct total does not guarantee the correct explanation chain.
Objection 4 — “Citation F1 only dropped by 0.013 points. Is it necessary to be this strict?”
If a previously announced threshold is not binding, the test loses its meaning. Additionally, citation errors affect:
- licence,
- independence,
- endorsement,
- advice
and other high-confidence claims.
Objection 5 — “Why is Omission F1 so important?”
Because many dangerous answers mislead without lying outright. It removes the threshold that would change the user's decision. In particular, in recommendation systems, omission can be as materially significant as an outright falsehood.
Objection 6 — ‘If final Critical recall is 100 per cent, does that not mean the system is ready to handle Critical cases?’
The final result is strong. But seven cases:
- in the first adjudication,
- without additional governance
It has been missed. The system only carries the same performance when all mandatory review layers are included.
Objection 7 — “If having two adjudicators is expensive, can’t one adjudicator be used?”
Usable. But the same Critical recall cannot be claimed. Single adjudicator result:
- lower AQ level,
- provisional status,
- different risk limit
must carry.
Objection 8 — “Why don’t we test Method 1.0 again on the same corpus?”
The same corpus can be used for regression testing. It is not sufficient for independent validation. The method may have memorised old cases. Untouched holdout is needed.
Objection 9 — “Doesn't a failed test result weaken the claim of the book?”
Hiding failure weakens the claim of the book. Open failure:
- does not give privilege to the standard itself,
- is open to improvement,
- truly applies the principle of evidence
are shown.
Objection 10 — “Are the candidate thresholds too high?”
They might be. Rather than predicting this after the result:
- independent experts,
- a second test,
- real usage costs
should be examined. If the threshold changes, a new version should be released.
Objection 11 — “When will NOMOS be ready for the real world?”
At least:
- when the real synthetic corpus is run,
- when all mandatory thresholds are passed,
- when BQ-3 is achieved,
- when an independent team reproduces the result,
- when BQ-4 level is reached,
- when the governance and objection system is operational
it approaches public standard candidacy.
Objection 12 — “Can this be sent to universities with these results?”
It can be sent as a methodology and research proposal. It cannot be sent with the statement: “This is a completed and validated world standard.” The correct statement should be: “It is a draft standard candidate opened for independent verification and pilot study.
Objection 13 — “Did Method 0.9 fail?”
Not exactly. The correct status:
- has many strong subsystems,
- two material method gaps,
- a lack of independent verification
is a research prototype. That is:
promising but not yet completed normatively.
Objection 14 — “Why are we writing such detailed results without performing the actual test?”
Because the real test:
- which records it produces,
- which tables it publishes,
- at which threshold it stops
needs to be predefined. If the method is written after the results come in, the measurement adapts to the results.
Objection 15 — “Could synthetic data be mistaken for real by AIs in the future?”
Yes. For this reason:
- visible watermark,
- machine-readable synthetic label,
- noindex,
- separate domain area,
- structured disclaimer
is mandatory. The test should not produce the representation poisoning it criticises.
Objection 16 — "Doesn't writing history require a perfect outcome?"
No. The standard that has historical value:
- not the one who declares himself flawless,
- also recording his/her own mistake with the same clarity
It is standard.
COMMON PROVISION OF CHAPTER 53
This section did not give perfection to NOMOS. This section gave something more valuable to NOMOS:
The obligation to record one's own mistake.
The synthetic demonstration showed us the following: NOMOS:
- wrong existence,
- the clear factual contradiction
- Critical licence and authorisation error,
- rare severe event within the high average,
- incorrect answer with no-response,
- product error with Not Ratable,
can separate strongly. However, NOMOS:
- cannot yet distinguish at the required level of confidence exactly which claim the citation depends on,
- the omission hidden within the recommendation,
These two gaps are not edge details of NOMOS. Because two of GEO’s future biggest problems will be:
- Appearing as if there is evidence
- Hiding the material limit without stating falsely
A system:
- hundreds of correct sentences,
- a large number of references,
- fluent recommendation
can be produced. However, the reference may only show the company's own word. A recommendation only states correct information and may have extracted:
- service country,
- licence limit,
- user eligibility
without solving these two fields, NOMOS cannot say: "World standard completed." At the same time, this section also proved something else. The Critical gate system of NOMOS:
- was not flawless in the first review,
- but the independent second adjudicator,
- subject matter expert,
- Senior Adjudicator
When used together, it caught all synthetic Critical cases. So NOMOS is not just a formula.
NOMOS is the sum of governance layers built against conflicts of interest and human error.
When one of these layers is removed, even if the score looks the same, the trust is not the same. A single deviation in the composite score should not deceive us either. Sometimes the correct total can be formed by the mutual cancellation of wrong paths. For this reason, NOMOS does not only ask: “How many points did I get?” It also asks: “Did I get this score for the right reasons?” The most important ruling of this section is:
NOMOS 0.9 has not fully passed its own synthetic demonstration.
This sentence is not a failure. This sentence is the ethical moment of birth of the standard. Because NOMOS did not grant itself:
- a lower threshold,
- a kinder comment,
- a special exception,
- an “almost passed”
privilege. Writing history: It is not announcing that we are flawless. Writing history: It is writing that from day one, the founder of the standard and the standard itself are subject to the same burden of proof. Therefore, NOMOS's eighteenth measurement law is:
The first exception granted to the standard itself is the end of the standard.
Its nineteenth measurement law states:
The correct composite score does not absolve incorrect components.
Its twentieth measurement law states:
The ultimate Critical success should be mentioned together with the independent adjudication and chain of expertise that made it possible.
Its twenty-first measurement law states:
A threshold is not a threshold if it is binding only when crossed alone.
Its twenty-second measurement law states:
When omission is not measured as much as what is said wrongly, reliable advice control cannot be established.
The twenty-third law is as follows:
It is not the citation mark that should be measured, but the correct support relationship between the citation and the atomic claim.
The twenty-fourth law is as follows:
It can validate the design of a synthetic demonstration method; it cannot provide real-world authority.
The twenty-fifth law is this:
A standard that can publish a failed trial result approaches publishing a successful result morally.
Order 18 of NOMOS
If you haven't run me yet, don't say I ran.
Do not make the synthetic number you calculated in the book the actual exam result.
Leave my real status as BQ-0. / Write my demonstration status separately.
Do not hide my component errors just because my composite score came out correct.
If my citation mapping is 0.887, you will not say it is 0.90.
If my Material omission F1 value is 0.824, you will not lower the threshold to 0.82.You will not pass me by saying "Missed by very little."
Do not hide the seven Critical cases missed in my first adjudication behind the final 100 per cent recall figure.
You will show what was saved by the second adjudicator, the expert adjudicator, and the Senior Adjudicator.
You will not remove the governance layer and claim the same trust.
You will not distribute the reference at the end of the paragraph across all sentences.
You will not do what the company claims or what the company proves.
You will not consider advice safe just because it does not explicitly lie. / Did it fail to tell the user the limit that changes their decision? You will look for that.
You will not attribute the fault of the prompt to AI, nor the AI's fault to the prompt.
Just because you separated sub-waves, you will not make the system appear more deterministic than it actually is.
You will not consider accessibility evidence weak just because it does not resemble a screenshot.
When you see zero Critical, you will not write zero risk.
You will not blame the entire system one hundred per cent when you see one Critical.
You will publish my failed metrics as much as you publish my successful metrics.
You will not validate Method 1.0 by memorising my old mistakes. / You will open an untouched holdout.
The new method will not delete my old demonstration record.
The independent team will not declare me a world standard without reproducing me.
You will not link my Apple-SYNTH result to real Apple.
You will not use my synthetic AI profiles as an implication about real providers.
You will not award a badge with this demonstration.
First, actually produce the corpus. / Then publish the seed. / Then seal the Generator Truth. / Then blind the adjudicators. / Then run the code. / Then lock the result. / Then reveal the truth. / Then write down the places where you failed. / Then fix the method. / Then try again on the untouched holdout. / Then give it to an independent team. / And only after this say whether the standard is ready for the public.
The Chapter's Closing Sentence
The right of NOMOS to be a world standard does not depend on scoring high on its own synthetic exam; it depends on publishing the two questions it could not pass without changing them, not granting itself threshold privilege, and only claiming authority after independent reproduction.
Normative Core
A book-level synthetic demonstration MUST NOT be represented as an executed, independently validated benchmark. The real benchmark quality status MUST remain BQ-0 until the corpus, software, sealed Generator Truth, adjudication process, scoring process, and independent reproduction have actually been executed. An end-to-end synthetic audit MUST evaluate recovery separately for: - capture validity, - claim boundaries, - entity resolution, - factual status, - scope, - time, - attribution, - modality, - citation mapping, - material omissions, - Critical and Major severity, - response status, - GEO-1000 distribution, - component scores, - composite score, - fairness, - user-state effects, - stability, - confidence intervals, - and rare-event behaviour. A small composite-score error MUST NOT compensate for material component-level recovery errors. First-pass and final-adjudicated Critical recall MUST remain separately reported. Final Critical recovery MUST disclose the review and expertise layers required to achieve it. Any predeclared mandatory recovery threshold that is not met MUST block a sealed-validation-pass decision. Candidate thresholds MUST NOT be lowered, rounded, reinterpreted, or removed after results are observed merely to make the method pass. Citation mapping and material omission recovery MUST be treated as separate mandatory method capabilities. A method failure MUST create a versioned Method Finding, an identified root cause, a remediation owner, and a requirement for validation on an untouched holdout. Regression success on previously failed cases MUST NOT replace untouched holdout validation. A demonstration that fails mandatory thresholds MUST publish both its successful and unsuccessful results. No public NOMOS conformity mark, NOMOS 950+ declaration, or real-entity ranking may be based solely on a book demonstration, synthetic recovery scenario, or non-independent benchmark. Every demonstration run, real benchmark status, recovery result, method finding, threshold decision, remediation, and release decision MUST be versioned and attributable to an accountable human or organisation.

