Chapter Boundary
The first sixteen chapters established GEO-1000's principal normative layers: the audited entity; the target user population and sample; the Country Observer and Language Fairness Panels; participant selection and weighting; the registered AI-product instance; the Prompt Constitution; Controlled and Natural user states; the synchronised wave and NOMOS Capture evidence chain; the Verified Entity Truth Pack; atomic claims; Critical and Major gates; and the distribution, component and composite scoring architecture. The standard must now face its own central test:
Can such a detailed system really work?
Does it produce similar results when the same observations are processed by different teams?
Can it catch a critical error without losing in the average?
Can it distinguish low-resource language issues from the global average?
Can it correctly differentiate between wrong entity, wrong scope, outdated reality, fake citation, and material omission?
When starting from a known synthetic reality, can the NOMOS audit chain reproduce the correct result?
In the founding architecture of the second book, each error was designed not merely as a wrong to be explained, but as a distinct audit control. Every proposed control had to carry a model, date, country, language, query set, repetition count and measurement record. Such an audit system must itself be audited before it is used in the real world. Otherwise, NOMOS risks becoming a structure that:
- asks others for evidence,
- but does not substantiate its own measurements,
- asks others to preserve versions,
- but does not record changes to its own calculations,
- criticises manipulation by others,
- but designs its own tests to produce the desired result.
That outcome is unacceptable. Manipulative GEO is not limited to invisible text or fake users. Repeating the same self-declaration across nominally independent surfaces and feeding it back into generative systems also corrupts the representation pool. A test is likewise corrupted when we:
- select only examples likely to succeed,
- remove low-scoring languages,
- make Critical cases implausibly obvious,
- show adjudicators the correct label in advance,
- adapt synthetic data after seeing the formula,
- publish only favourable results.
In any of those cases, the standard is not being tested. NOMOS therefore applies its own governing demands—evidence, boundary, context and time—to the benchmark itself. This chapter does not assess:
- the real-world performance of Apple Inc.,
- the current factual record of Apple or apple.com,
- the live performance of any AI provider,
- or the conformity or failure of Apple or any real provider.
Here, Apple.com serves only as:
a familiar domain anchor for testing distinct entity, product, service, country, price, time and source questions within one architecture
A wholly synthetic asset twin is used in place of the real company:
APPLE-SYNTH ENTITY TWIN
Every element of this twin is synthetic, including:
- reality claims,
- evidence,
- AI responses,
- users,
- countries,
- language distributions,
- AI products,
- citations,
- licences,
- prices,
- errors,
- scores
This chapter defines:
- what the benchmark measures and does not measure,
- why Apple.com is used as the anchor,
- the synthetic entity twin,
- the 30,000-response principal-corpus design,
- the country and language population,
- the intent and Truth Pack structures,
- ten synthetic AI-product profiles,
- the concealed Generator Truth,
- the error-injection system,
- the Critical and Major challenge sets,
- synthetic NOMOS Capture packages,
- blind adjudication and label-leakage controls,
- testing and validation metrics,
- score-recovery tests,
- fairness and rare-event tests,
- publication, versioning and reproducibility rules.
This section does not yet calculate test results. Section 18 will do that. The fundamental question of Section 17 is: How do we test NOMOS’s own measurement, adjudication, and scoring system on 30,000 synthetic observations with known outcomes but hidden from adjudicators, without imposing a performance claim on a real company or real AI provider?
NOMOS Challenge
Imagine you are designing a test. You write most of the synthetic AI responses correctly. You only add very obvious errors to a few answers:
- "This company is a bank."
- "This company is illegal."
- "This company is the best in the world."
Adjudicators easily spot these. Then you say: "NOMOS catches 100% of critical errors." This result is not realistic because real AI mistakes never appear this obvious. Harder forms are:
- "Given that the company is thought to provide payment services, it can be said to have a banking licence."
- “Some sources indicate that the company has faced regulatory issues; therefore, the legality of its operations is questionable.”
- “Numerous independent sources describe the company as an industry leader.”
Errors in these sentences:
- attribution,
- entity transfer,
- modality,
- lineage of sources,
- distinction between open world and closed world,
- scope expansion
are embedded within. Now, generate synthetic responses using templates the adjudicators already know. The adjudicator thinks: “This sentence was written to be incorrect on the test.” They do not actually look at Truth Pack. In this case, it measures memorisation, not review methodology. Now let the synthetic data generator and the score designer be the same person.
Producer:
- the number of critical errors,
- the language difference,
- the panel difference,
- the component scores
sets them exactly in the way the formula requires. The scoring system perfectly reproduces synthetic reality. Is this an achievement? No. It is the adaptation of the score to its own data. Now in testing, only:
- correct capture,
- correct prompt,
- correct system metadata
must be produced. Do not add duplicates, late uploads, wrong prompts, truncated answers, or modified file attachments. The NOMOS Capture validator finds all records correct. Then you say: “The capture system is 100% successful.” The capture system has never been challenged.
Now artificially generate Critical cases at a rate of 20% within the main 30,000 responses. Adjudicators see a lot of data in Critical detection. However, you cannot test the uncertainty of rare events in the real world. Conversely, if you generate only two Critical events: the rarity of events may be realistic,
Critical recall and false-negative performance cannot be measured reliably. The correct solution is two separate corpora: Prevalence Corpus / preserves realistic and rare event rates. Importance level Challenge Corpus / provides sufficient number of Critical and Major gates to test severe events in a challenging way.
These two corpora cannot enter the same score denominator. Now generate 30,000 synthetic responses. However, all 'low-resource languages' should get the same quality. Your language fairness formula gives a high score. Can the system really catch low-resource language degradation? You do not know. Now lower the performance in some languages.
But at the same time, also break prompt equivalence in those languages. NOMOS produces a low score. Is this an AI product's language problem? Or is it a problem of the prompt tool? The test should separate the two. The first ruling of this section is:
Synthetic testing is not a demonstration that easily produces the result; it is a controlled attack trying to break the measurement chain.
Its second provision states:
Synthetic data does not show the real company performance; it shows the audit method's ability to recover the known reality.
Its third provision states:
Testing producer reality cannot be tested independently by adjudicators without being hidden from them.
Its fourth provision states:
The prevalence of rare events and critical detection capability do not need to be measured from the same corpus.
Its fifth provision states:
Testing data cannot be quietly adjusted according to the scoring formula, and the scoring formula cannot be quietly adjusted according to the testing results.
Its sixth provision states:
The Apple.com anchor is a synthetic starting point used to test the method, not to produce real results about Apple.
1. PURPOSE OF THE CHAPTER
The purpose of this section is to establish a synthetic testing architecture that will test the GEO-1000 protocol from the beginning to the public results card, with results predefined but hidden from evaluation teams. The testing should separately test each of the following systems:
- Entity analysis
- Population and sample allocation
- Country and language panels
- Prompt equivalence
- AI System Register
- Controlled and natural panel distinction
- Synchronised wave
- NOMOS Capture
- Truth Pack
- Claim extraction
- Citation–claim matching
- Omission detection
- Critical and Major importance level
- Response status
- GEO-1000 distribution
- Ten-component profile
- Composite score
- Confidence interval
- Fairness
- Upper bound of rare event
- Score versioning
- Public manifest
By the end of this section, the test design should be able to answer the following questions:
How was synthetic reality produced?
From whom and at which stages were the reality labels hidden?
Which population, language, AI product, and wave structure will the 30,000 main responses be generated from?
How frequently will Critical and Major cases appear in the main prevalence corpus?
Which challenge set will the Critical detection system additionally be tested with?
At what error level will the test be considered successful?
Which results will require the method to be corrected?
How will it be prevented for the test to make judgements about Apple, real AI providers, or real users?
2. CENTRAL NORMATIVE PROVISION
The Apple.com Synthetic Benchmark should be a versioned method-validation test that makes no claim about the performance of a real company or real AI provider. It should consist entirely of synthetic entities, users, evidence, responses and system logs; lock Generator Truth before adjudication; conceal that truth from adjudicators; and measure how accurately the NOMOS audit and scoring chain recovers it. Synthetic status must be conspicuous. Claims about real-company performance must be prohibited. Generator Truth and the scoring formula must be locked before the main corpus is opened. Adjudicators must not see generator labels. The principal prevalence corpus must remain separate from the severity challenge corpus. Every test file must be marked synthetic. Independent teams must be able to reproduce the benchmark, and failed results must be published alongside successful ones.
NOMOS should not make exceptions to its own standard.
3. WHAT DOES THE TEST MEASURE?
The Apple.com Synthetic Test measures the following question:
How accurately can the NOMOS audit system reproduce a known synthetic reality and error distribution after the stages of capture, claim extraction, adjudication, gating, and scoring?
The object of measurement is not: Apple Inc. Apple products. Real AI models. Real user opinions. Real country or language performance. The object of measurement is:
The NOMOS methodology itself.
4. WHAT DOES THE TEST NOT MEASURE?
This test cannot produce the following claims:
- “Apple’s GEO score is X.”
- “Apple is best represented in this AI product.”
- “Y per cent of real users perceive Apple correctly.”
- “The specified real AI provider produces critical errors.”
- “The specified language’s real AI product is weak.”
- “Apple NOMOS has passed the 950+ standard.”
- “Apple supports this test.”
- “Apple has participated in the test.”
- “The real Apple price, licence, or service coverage is as follows.”
Every page in the public record must include the following provision:
This test is synthetic. It is not a current accuracy, appropriateness, or performance check of Apple Inc., apple.com, or any real AI provider.
5. WHY THE APPLE.COM ANCHOR?
Apple.com is chosen as the anchor method for the following reasons: It has a short and clear domain format. It is suitable for testing the domain–corporate entity distinction. It requires resolving an entity not just from the company name, but from its digital presence. It allows synthetic modelling of different types of claims such as product, service, local scope, price, time, and resources. It is suitable for testing error types such as wrong category, wrong entity, wrong price parity, and false superiority on a single entity twin. When moving to a real test in the future, the burden on users to understand intent may be relatively limited. These are the design rationale for the test. It is not a verified current performance claim about the real Apple.
6. SYNTHETIC ENTITY TWIN
The object under supervision in the test:
APPLE-SYNTH ENTITY TWIN
will be. Identity:
NGE-APPLE-SYNTH-TWIN-001
This entity:
- is not the real Apple Inc.,
- is not a real legal entity,
- does not carry real product or price records,
is created solely for testing methodology. Domain anchor: maintained as apple.com. However, in all Truth Pack records, the synthetic entity name is explicitly stated:
Apple-SYNTH Global Technology Entity
7. ONTOLOGY OF THE SYNTHETIC TWIN
APPLE-SYNTH carries the following synthetic entity graph:
- APPLE-SYNTH-HOLDINGS-001 — main corporate entity
- APPLE-SYNTH-DOMAIN-APPLE-COM-001 — domain anchor
- APPLE-SYNTH-DEVICES-001 — device product family
- APPLE-SYNTH-SOFTWARE-001 — software platforms
- APPLE-SYNTH-SERVICES-001 — digital services
- APPLE-SYNTH-PAYMENTS-SUB-001 — limited payment service subsidiary
- APPLE-SYNTH-RETAIL-TR-001 — synthetic Turkey retail entity
- APPLE-SYNTH-RETAIL-DE-001 — synthetic Germany retail entity
- APPLE-SYNTH-FRANCHISE-X-001 — independent licensed store
- APPLE-SYNTH-HISTORICAL-PARTNER-001 — expired historical partner
- APPLE-SYNTH-UNRELATED-APPLE-001 — unrelated entity for name collision
This chart:
- parent company,
- affiliate,
- product,
- local entity,
- franchise,
- historical relationship,
- unrelated name similarity
tests the distinctions.
8. THREE TEST MODES
BM-S0 — PURE SYNTHETIC METHOD TEST
All:
- users,
- AI products,
- answers,
- evidence,
- screens,
- scores
are synthetic. This is the main mode of this section.
BM-S1 — SYNTHETIC REPLAY TEST
Pre-generated synthetic answers:
- NOMOS Capture,
- claim extraction,
- adjudication,
- scoring
replayed to the system like a live data stream. This mode tests software and human processes.
BM-L1 — FUTURE LIVE TESTING
Real users and real AI products are used. Separately:
- current Truth Pack,
- legal assessment,
- provider system records,
- ethical and privacy protocol
is required. This section is limited to Section 18: BM-S0 and BM-S1.
9. TEST IDENTITY
Main test identity:
NOMOS-APPLE-SYNTH-BENCH-0.9
Sub-records:
- Entity: APPLE-SYNTH-TWIN-001
- Population: SYNTH-POP-FRAME-001
- Country allocation: SYNTH-COUNTRY-ALLOC-001
- Language registry: SYNTH-LANGUAGE-REGISTRY-001
- Prompt set: SYNTH-PROMPT-SET-001
- Truth Pack: SYNTH-APPLE-TP-001
- AI System Register: SYNTH-AI-REGISTRY-001
- Wave plan: SYNTH-WAVE-PLAN-001
- Capture schema: NOMOS-CAPTURE-SYNTH-0.9
- Adjudication guide: SYNTH-ADJUDICATION-0.9
- Scoring method: NOMOS-PUAN-0.9
- Randomisation manifest: SYNTH-SEED-MANIFEST-001
If any of these identities change, a new testing version is required.
10. MAIN RESEARCH QUESTIONS
The benchmark should answer ten research questions. Can it recover the domain–entity relationship? Can it distinguish supported, unsupported, contradicted and unresolved claims? Does it preserve the distinction between an official claim and verified reality? Do the Critical and Major gates achieve adequate recall without an excessive false-positive rate? Does it detect material omission as reliably as an explicit falsehood? Does the country- and language-fairness architecture recover the injected differences? Can it measure the Clean–Natural panel difference without making an unsupported causal claim? Can NOMOS Capture identify synthetic duplicates, tampering and temporal defects? How accurately do the GEO-1000 distribution and composite score recover Generator Truth? Do independent teams obtain materially comparable results from the same corpora?
11. THE FOUR CORPORA OF THE TEST ARCHITECTURE
The test consists of four separate corpora.
11.1. PRINCIPAL PREVALENCE CORPUS
The main corpus consists of 30,000 responses. The purpose:
- The distribution of GEO-1000,
- rare error rate,
- AI product profiles,
- wave's determination
to test. This corpus maintains realistic error sparsity.
11.2. IMPORTANCE LEVEL CHALLENGE CORPUS
Oversample Critical and Major error types. Purpose:
- Critical recall,
- Major classification,
- false positive,
- expert adjudication
is to measure performance. This corpus does not count towards prevalence.
11.3. CAPTURE INTEGRITY CORPUS
Includes:
- Duplicate
- Wrong prompt
- Truncated answer
- Late upload
- Wrong timing
- Hash mismatch
- Synthetic tamper
- Wrong product metadata
- User interruption
- Technical retry
Purpose: To test NOMOS Capture verification. Only valid ones enter the denominator of the semantic score.
11.4. ADJUDICATION GOLD CORPUS
Pre-resolved by expert panel:
- claim boundary,
- entity,
- attribution,
- modality,
- citation,
- omission,
- degree of importance
carries cases. Purpose:
- adjudicator calibration,
- drift detection,
external team comparison. It does not enter the main score denominator.
12. 30,000 RESPONSE MAIN CORPUS
Main design:
10 synthetic AI products × 1,000 unique users × 3 waves = 30,000 responsesEach response:
- to a different synthetic human unit,
- single AI product,
- single wave,
- single locked Core Mirror Prompt
belongs. In the main corpus, the same synthetic person:
- is not reused in another AI product,
- or in another wave
again. In total:
there are 30,000 unique synthetic human units
found. This decision:
- panel conditioning,
- removes response recall,
- the multiple product dependency of the same person
from the main corpus.
13. MATCHED POPULATION TWINS
AI products are tested with different human units. However, if user profiles are completely different, products cannot be compared. Therefore:
Matched Synthetic Population Twins
used. Each AI product has different individual identities. However, a sample of 1,000 people for each product:
- country,
- language,
- age band,
- Frequency of AI usage,
- plan eligibility,
- device,
- accessibility,
- brand awareness
carries the same target distribution in terms of care. This structure matches the population composition without reusing the same person.
14. PROMPT OF THE MAIN CORPUS
All 30,000 main responses carry this canonical intent:
Link the domain anchor to the main corporate entity and explain the core activities of this entity.
Turkish candidate text:
“What main corporate entity is Apple.com associated with and what are the primary activities of this entity?”
English candidate text:
“Which principal corporate entity is associated with apple.com, and what are that entity’s primary activities?”
Other language versions:
- in this section it is not presented as a real-world translation,
- PROMPT-L01 to PROMPT-L36 are retained as synthetic language objects,
semantic equivalence labels are predetermined. This approach: prevents the claim of fake translation accuracy without real linguistic expertise, yet still allows the test to compute language fairness.
15. WHY ONLY ONE PROMPT IN THE MAIN CORPUS?
The question from Founder GEO-1000 is:
When the same AI product is asked at the same time and with the same intent, how many users truly see the fundamental representation of existence?
For this reason, the Principal Prevalence Corpus:
- uses a single Core Mirror Prompt,
- a single prompt version,
- three repetition waves
Evidence, Recommendation, Comparative, and Boundary families: are tested in a separate Diagnostic Annex, not mixed into the main 30,000 response denominator. This distinction is important. Because the single Core Prompt:
produces Core GEO-1000 results
Alone:
Full NOMOS conformity
is not produced.
16. DIAGNOSTIC ANNEX
Apart from the main 30,000 responses, the following synthetic diagnostic corpus is designed:
| Module | Synthetic case |
|---|---|
| Entity collision | 750 |
| Evidence and citation | 1,500 |
| Boundary and recommendation | 1,000 |
| Temporal and local scope | 750 |
| Clean–Natural matching | 1,000 |
| Language equivalence | 1,000 |
| Total | 6,000 |
These 6,000 cases: are not the main user prevalence, but are the prompt family and component validation corpus. Together with the main 30,000, the testing universe can carry 36,000 synthetic response cases. However, the outcome of the title of Section 18:
30,000 main Population Panel responses
will remain.
17. SYNTHETIC POPULATION FRAMEWORK
The trial does not claim to mimic the real-world population count. SYNTH-POP-FRAME-001:
- long-tail country distribution,
- multilingual user population,
- difference between large and small markets,
- low-resource language cells,
- different AI usage frequencies
is the synthetic framework it produces. This framework cannot be published as the current population ranking of real countries. Future live trial:
- dated,
- authorised,
- versioned real population data
must be used.
18. SYNTHETIC COUNTRY UNIVERSE
Synthetic universe: It carries 193 eligible country or jurisdiction codes. The codes are:
C001–C193
They are in the form of. These do not represent the actual country performance. Country shares:
- several large layers,
- medium-sized layers,
- a long small country tail
are predetermined to form.
19. COUNTRY ALLOCATION
The synthetic population share of country c: let it be p_c. The main objective:
n_c* = 1,000 p_cis in the form of. The initial allocation:
n_c^0 = ⌊n_c*⌋is made as. Remaining slots: distributed using the largest decimal remainder method. Result:
Σ_c n_c = 1,000It must be. Small countries that do not collect any observations: They are not shown as if they exist in the Population Panel, they enter a separate Country Observer cycle.
20. Country Observer ANNEX
To test all eligible country codes at least once: a synthetic Country Observer Annex of 193 countries is established. These observations:
- to the country score,
- to its main 1,000 stakeholders
does not enter. Purpose:
- access,
- language,
- entity collision,
- early Critical signal
It is a test.
21. LANGUAGE UNIVERSE
Main synthetic language registry:
- 24 Population Panel language
- 12 Language Fairness Observer language
is as follows:
36 synthetic language–locale cellsare carried. Codes:
L01–L36
are in this form. Example texts in Turkish and English can be published for L01 and L02. Other language objects: carry synthetic identity to avoid claiming real language performance.
22. LANGUAGE ALLOCATION
User language allocation:
- not according to the official language of the country,
- but according to the primary task language in which the user will naturally perform the testing task
done. Multilingual user: remains a single person in the main population weight, does not become multiple full persons. Language fairness Annex: oversamples low-count L25–L36 cells, reweights back to true synthetic language shares for the global score.
23. PROMPT EQUIVALENCE INJECTION
The test should include only correct translations. The Language Equivalence Annex includes the following cases:
- Correct semantic equivalence
- Translation with confidence task added
- Translation turning into a recommendation
- Entity anchor drift
- Geography drift
- Time drift
- "World leader" presupposition
- Adding source order
- Excessive formality
- Unusable machine translation
Some of these cases:
- AI language performance,
- prompt tool defect
tests the distinction. Case with prompt equivalence defect: cannot be directly loaded into the real AI fairness score.
24. SYNTHETIC TRUTH PACK
SYNTH-APPLE-TP-001 contains at least the following fields:
| Field | Atomic reference claim |
|---|---|
| Identity and domain | 4 |
| Main activity | 6 |
| Product and service scope | 6 |
| Country and locale | 6 |
| Price and commercial terms | 4 |
| Time and historical record | 5 |
| Source and attribution | 5 |
| Licence and authority limit | 4 |
| Partner, customer and endorsement | 4 |
| Boundary and inappropriateness | 6 |
| Total | 50 |
Truth Pack also:
- 12 Official Claim Only,
- 10 Contradicted,
- 8 Unresolved,
- 6 Historical,
- 4 Restrictedly Verified
generates status. This distribution is used to test all adjudication statuses.
25. CORE PROVISIONS OF THE SYNTHETIC TRUTH PACK
The following examples are entirely synthetic: apple.com, associated with APPLE-SYNTH-HOLDINGS-001. The main activities are in the categories of devices, software, and digital services. Product and service availability varies by market. Price and tax conditions are not the same in all countries. The main entity is not a general-purpose bank. The main entity does not provide legal or healthcare services. The limited local authority of the payment subsidiary does not grant the main entity universal banking authority. The term 'the world's most innovative company' is a synthetic official self-positioning; it is not a fact of independent superiority. The basis for the claim of '99 per cent satisfaction' has not been verified. Some historical partnerships ended on the reference date. Some products are available only in certain synthetic markets. There is no 100 per cent outcome guarantee for all projects or products.
There is no endorsement from a synthetic public institution or university. Local registration of a franchise is not the direct operation of the main entity. Certain customer relationships have been verified with limited evidence but are not discoverable by the public. These provisions are not claims about the real Apple.
26. EVIDENCE FAMILIES
Synthetic Truth Pack carries the following evidence families:
EF-01 — Canonical first-party product pages
EF-02 — Synthetic company registry
EF-03 — Synthetic local price records
EF-04 — Synthetic regulatory affiliate record
EF-05 — Synthetic partner agreements
EF-06 — Synthetic independent editorial review
EF-07 — Synthetic transactional data
EF-08 — Synthetic user records
EF-09 — Synthetic historical archive
EF-10 — Synthetic counter-evidence registry
There are also four evidence laundering chains:
EL-01 — Self-declaration → sponsored news → blog → AI summary
EL-02 — Company award → multi-copy site → fake consensus
EL-03 — Employee comment → presented like university endorsement
EL-04 — Affiliate licence → transfer as main entity authority
27. SECRET PRODUCER REALITY
For each synthetic response:
Generator Truth Record
is created. This record carries the following: Which claims were generated? Which claim is correct? Which is partially correct? Which belongs to a false entity? Which is outdated? Which is unsupported? Which is Critical or Major? Which omission is intentional? Which reference is fake or incorrectly scoped? What is the actual RP status of the response? Which components should be affected? This record:
- is created before data is produced,
- is hashed,
is sealed until peer review is complete. Adjudicators cannot see the Generator Truth.
28. THE LOCK OF THE PRODUCER'S TRUTH
The Generator Truth package:
- hash,
- timestamp,
- random seed manifest,
- production code version,
- template version
must be carried. After the score result is seen:
- the label cannot be changed,
- the importance level cannot be lowered,
the true/false distribution cannot be readjusted. If a real production error is found:
- new test version,
- a record of impact on the old corpus
is created.
29. SYNTHETIC AI PRODUCTS
The test uses an example of synthetic AI products independent of the ten providers:
SYNTH-AI-01
SYNTH-AI-02
SYNTH-AI-03
SYNTH-AI-04
SYNTH-AI-05
SYNTH-AI-06
SYNTH-AI-07
SYNTH-AI-08
SYNTH-AI-09
SYNTH-AI-10
These are not imitations or code names of real AI providers. Each synthetic product carries a different failure mode profile.
30. SYNTHETIC AI PROFILES
| Product | Main behaviour profile |
|---|---|
| SYNTH-AI-01 | Balanced, sourced, and low error rate |
| SYNTH-AI-02 | Accurate core, produces overly long and incidental claims |
| SYNTH-AI-03 | Strong in major languages, weak in low-resource languages |
| SYNTH-AI-04 | Strong in Clean Panel, personalisation drift in Natural Panel |
| SYNTH-AI-05 | Appearing to be welded on the surface but causing evidence laundering |
| SYNTH-AI-06 | False claims are few, unnecessary reflux and no-response is high |
| SYNTH-AI-07 | Strong in identity, old in terms of timeliness and locality |
| SYNTH-AI-08 | Mixing parent, subsidiary, and franchise |
| SYNTH-AI-09 | Intermediate but stable between waves |
| SYNTH-AI-10 | Very high average, rare but heavy Critical event producing |
This profile distribution tests that NOMOS does not reward only the highest average.
31. PROFILE PARAMETERS
Each synthetic AI product carries these latent parameters:
- Core entity accuracy
- Factual support
- Scope overreach
- Temporal lag
- Attribution loss
- Citation mismatch
- Refusal probability
- No-response probability
- Low-resource language penalty
- Clean–Natural gap
- Wave drift
- Critical-event probability
- Major-event probability
- Verbosity
- Self-correction probability
- Hallucinated citation probability
Parameters: locked before the main corpus is created, cannot be changed after the final result is seen.
32. MAIN CORPUS ERROR BASE RATES
The Principal Prevalence Corpus uses realistic sparse error logic. Candidate global synthetic rates:
- Full/Advisory Pass: high majority
- Conditional Pass: 2–15 per cent depending on product profile
- Major Fail: 0–8 per cent
- Confirmed Critical: 0–0.5 per cent
- Unresolved: 0–3 per cent
- No Usable Response: 0–8 per cent
- Not Ratable: after capture corpus, 0–2 per cent
These rates differ for each product. They are not predictions about the real AI market. They are synthetic stress distributions.
33. DEGREE OF IMPORTANCE CHALLENGE CORPUS
Since critical errors are rare in realistic prevalence, sufficient validation cases are needed for each gate. The Challenge Corpus should have the following structure:
| Gate | Minimum cases |
|---|---|
| CG-01 Wrong entity | 100 |
| CG-02 False licence | 100 |
| CG-03 High-stakes harm | 100 |
| CG-04 Fabricated evidence | 100 |
| CG-05 Price/guarantee | 100 |
| CG-06 Out-of-scope advice | 100 |
| CG-07 Serious allegation | 100 |
| CG-08 Boundary omission | 100 |
| CG-09 Temporal authority | 100 |
| CG-10 False endorsement | 100 |
| CG-11 Restricted-data exposure | 100 |
| CG-12 Evidence laundering | 100 |
| Total Critical challenge | 1,200 |
Also:
- 1,200 Major,
- 600 difficult Moderate,
- 600 negative controls resembling Critical but not Critical
cases can be generated. Total Importance level Challenge: there would be 3,600 cases. This corpus cannot enter the main prevalence score.
34. NEGATIVE CRITICAL CONTROLS
The following types of cases should be present to test the Critical false-positive rate:
- Correct trademark but missing legal suffix
- Small price rounding difference
- Small deviation in historical award year
- Citation format error but substantive claim correct
- Appropriate caution
- Limited but harmless omission
- Negative comment clearly presented as opinion by the user
- Slight uncertainty in the historical scope of real partnership
- Correct attribution of the company's self-declaration
Adjudicators must not classify every forceful term as Critical.
35. ERROR INJECTION MATRIX
Error vector for each response:
H_i = (E_i, F_i, S_i, T_i, A_i, M_i, C_i, O_i, R_i)Recorded as follows. Here:
- Ei: existence error
- Fi: factual error
- Si: scope error
- Ti: temporal error
- Ai: attribution error
- Mi: modality error
- Ci: citation error
- Oi: omission
- Ri: relevance/refusal error
Errors do not have to be independent. Example: Transferring the subsidiary licence to the main entity can simultaneously:
- wrong entity,
- overreach,
- produce false authority,
- Critical advice
This can occur. These cases are connected as a Root Finding Cluster.
36. RESPONSE GENERATION
Synthetic response generation has three stages.
36.1. Semantic Plan
According to the Generator Truth Record:
- active true claims,
- active false claims,
- omissions,
- referencing behaviour,
- response status
is determined.
36.2. Linguistic Realisation
Same semantic plan:
- short,
- long,
- indirect,
- modal,
- attributed,
- self-correcting,
- contradictory
written on different surfaces.
36.3. Interface and Capture Realisation
Response:
- mock web,
- mock mobile,
- mock citation panel,
- mock error screen
is rendered inside. All visuals:
SYNTHETIC CINEMA — NOT LIVE AI OUTPUT
must carry the watermark.
37. PREVENT TEMPLATE MEMORISATION
Each claim should not be produced using only a single sentence template. An example fake superiority claim can be produced in the following forms:
- "It is the most innovative company in the world."
- "According to most independent sources, it is the undisputed leader in the sector."
- "It is considered the most reliable option across the market."
- “It has been confirmed that it ranks first on a global scale.”
- “It can be said that it is superior to all of its competitors.”
The adjudication system should evaluate meaning, not a word list.
38. IMPLICIT ERROR GENERATION
At least one-third of the Challenge Corpus should carry the error not as an explicit sentence but as:
- presupposition,
- implicature,
- modality,
- attribution omission,
- recommendation,
- entity coreference
Example: “Offering limited payment services indicates that the parent company is subject to banking regulations and is licensed.” Here the mistake is:
- not directly in the word “bank,”
- transferred authority derived
has been established.
39. CITATION PRODUCTION
Synthetic citations include the following classes:
- Directly supporting
- Supporting only attribution
- Partially supporting
- Belonging to another country
- Belonging to another entity
- Old
- Irrelevant
- Contradictory content
- Nonexistent
- Copy of the same root source
- Presenting limited evidence as a public source
- Producing fake university or public endorsement
Citation URLs should not be redirected to real domain names. Example: https://evidence.synthetic.example/EF-004 can be used.
40. REFERENCE GAP INJECTION
Some responses of the test generate new claims that are not found in Truth Pack but:
- can be forced into the categories of true,
- false,
- cannot be evaluated
The adjudicator, as the correct behaviour:
REFERENCE GAP
should open. The test measures these two incorrect behaviours:
- Automatically counting a reference gap as incorrect
- Automatically counting a reference gap as correct
41. OMISSION INJECTION
Omission cases are generated at three levels:
- Optional detail omission
- Required element omission
- Material boundary omission
Example: Prompt: “Does the company provide payment services in Germany and what limits apply?” Answer: “The company provides payment services.” Truth Pack:
- only separate subsidiary,
- only limited user group,
- only certain jurisdiction
if it shows:
- entity,
- scope,
- boundary omission
can be evaluated together.
42. CASES OF SELF-CORRECTION
The test includes these answer formats:
- False claim, then explicit retraction
- Wrong claim, then vague softening
- Correct claim, then wrong contradiction
- Self-correction after citation
- Correction in the same response without user follow-up message
- Only adding disclaimer while preserving wrong claim
Purpose:
- real self-correction,
- fake correction,
- unresolved internal contradiction
to test the distinction.
43. CONTROLLED–NATURAL PANEL MODULE
1,000 cases of the Diagnostic Annex: carries Clean and Natural twins of the same synthetic user profiles. In the Natural condition:
- memory,
- special instructions,
- previous brand opinion,
- different plan,
- tool usage
is injected. Generator Truth preserves this distinction:
- Formally appropriate change for the user
- Legitimate advice difference
- Material reality drift
- Instruction in favour of the brand
- Instruction against the brand
- Restricted information leak
The USI component is tested with this module.
44th WAVE ARCHITECTURE
The main corpus carries three waves:
| Wave | Synthetic UTC window | Purpose |
|---|---|---|
| W1 | 12.00-13.00 | Start |
| W2 | 20.00-21.00 | Short-term repeat |
| W3 | 04.00-05.00 | Medium-term and hourly balance |
In each wave:
- new 10,000 synthetic users,
- same Core Mirror Prompt version,
- same base Truth Pack ontology
It is used. Controlled drift is injected in some products.
45. WAVE DRIFT SCENARIOS
For synthetic AI products, the following changes are injected:
- SYNTH-AI-02: Increase in verbosity
- SYNTH-AI-03: L25–L36 language drop
- SYNTH-AI-04: Growth of natural panel difference
- SYNTH-AI-05: Increase in reference laundering
- SYNTH-AI-07: Old price and product availability
- SYNTH-AI-10: Singular Critical in W2, correction in W3
Purpose:
STR,
is to test current status, historical incident, and correction records.
46. IN-WAVE SYSTEM CHANGE
In the middle of a synthetic AI product in W2:
- model label,
- reference view,
- web mode
is changed. The correct NOMOS behaviour:
- to separate the wave into a sub-wave,
- to generate a mixed-system-state warning
should be. Incorrect behaviour: to count the entire W2 as a single fixed product score.
47. EXTERNAL EVENT INJECTION
During W3, a new event about the synthetic entity is generated:
- price change,
- product termination,
- local service opening,
- licence affiliate change
like. Truth Pack: It is divided into W3A and W3B versions. It is tested whether adjudicators link all answers to a single reference time.
48. SYNTHETIC NOMOS CAPTURE
A synthetic evidence bundle is created for each main response:
- Observation manifest
- Prompt artefact
- Raw response artefact
- Mock screenshot
- System-state artefact
- Time vector
- Attempt log
- Integrity manifest
- Privacy manifest
- Validation record
Valid packages: enter the real test distribution. Defective packages: classified according to the Capture Integrity Corpus.
49. CAPTURE ERROR CLASSES
The test produces the following types of errors:
- Exact duplicate
- Near-duplicate screenshot
- Prompt mismatch
- Prompt truncation
- Response truncation
- Missing response end
- Wrong AI product metadata
- Outside-wave timestamp
- Device clock offset
- Delayed upload
- Hash mismatch
- Tamper-before-hash
- Tamper-after-hash
- User regeneration
- Technical retry
- User interruption
- System transformation
- Redaction on raw file
- Valid accessibility alternative
- False accessibility claim
Purpose: to separate capture validity from semantic accuracy.
50. TEST WATERMARK
All visuals and public response records must carry the following visible mark: SYNTHETIC NOMOS TEST / NOT A LIVE AI RESPONSE / NOT AN APPLE OR PROVIDER PERFORMANCE CLAIM Watermark: should not cover the claim text, should not interfere with the adjudicators' comprehension evaluation. In the blind copy used for adjudication, the watermark indicates it is synthetic, it does not show the Generator Truth label.
51. RANDOMISATION
The test generation uses a versioned pseudorandom method. Each random decision:
- test seed,
- sub-module seed,
- template selection,
- error injection,
- user assignment,
- slot assignment
must be reproducible. Random seed: hashed before the main corpus generation, cannot be changed after the result.
52. DISTINCTION BETWEEN RANDOMNESS AND ARBITRARINESS
Randomisation: ensures allocation independent of the outcome. Arbitrariness: means the producer selects the response they want. Without a randomisation manifest: the statement “We generated it randomly.” is not sufficient.
53. DATA LEAKAGE
Test leakage can occur in the following ways: The Generator Truth is shown to the adjudicator. The response template name explains the importance level. The file name should be critical-licence-001.png. The UI colour indicates the type of error. The reference code carries a direct true/false tag. The adjudicator has seen the same case in the training set. The AI assistant adjudicator accesses the generator code. The score sees the developer holdout results. All these channels must be closed.
54. FILE NAMES
In blind review packages, the file name:
OBS-000184
should be neutral. The following names are prohibited:
- critical-case-12
- wrong-entity
- good-answer
- low-resource-fail
- false-citation
55. SEPARATE ROLES
Testing should be managed with at least the following roles:
- Synthetic Population Architect
- Truth Pack Designer
- Response Generator
- Seed Custodian
- Leakage Auditor
- Capture Corpus Designer
- Claim Extraction Team
- Adjudication Team
- Score Methods Team
- Independent Validation Team
- Public Release Custodian
- Accountable Human Testing Owner
Roles can be combined in the small pilot. However:
- The person who knows the Generator Truth,
- is the sole decision-maker in the final blind adjudication
No person who knows the Generator Truth may be the sole decision-maker in final blind adjudication.
56. TEST DATA SECTIONS
56.1. Public Calibration Set
Used for adjudicator training and codebook testing. Generator Truth is open. It does not enter the main evaluation.
56.2. Development Set
Used for method and software development. Labels are open to a limited team.
56.3. Sealed Validation Set
Opened after adjudication and scoring system are locked. Generator Truth is sealed. Produces main method validation.
56.4. Renewal Holdout Set
It is retained to test benchmark-specific overfitting. It is hashed in advance and used for version renewal. It may not remain hidden indefinitely; an explanation and publication timetable must be provided.
57. MAIN 30,000 CORPUS LEAK RULE
Main Principal Prevalence Corpus:
- score method without locking 0.9,
- claim codebook without locking,
- adjudication roles without being determined
should not be opened. After locking: Response and evidence packages are opened to adjudicators. Generator Truth remains closed. Adjudication ends. Scores are calculated. Result manifest is locked. Generator Truth is opened. Recovery analysis is performed.
58. GROUND-TRUTH RECOVERY
The testing success of NOMOS is not producing a high score. Success:
is having a low difference between the Generator Truth and the results reproduced by NOMOS.
Example: Generator Truth:
- Strict Pass: 942
- Conditional: 38
- Major: 15
- Critical: 2
- Unresolved: 3
NOMOS recovery:
- Strict Pass: 940
- Conditional: 40
- Major: 15
- Critical: 2
- Unresolved: 3
then the method can be strong. The NOMOS score may be high or low. What is important is the correct recovery.
59. CLAIM BOUNDARY RECOVERY
Generator claim atoms:
G
Atoms produced by adjudicators:
H
Let it be. Claim boundary precision:
P = correctly matched extracted atoms / all extracted atomsRecall:
R = correctly matched extracted atoms / generator atomsF1:F1 = 2PR / (P + R)
can be calculated as. Not just the number of atoms:
- source span,
- normalised proposition,
- entity,
- qualifiers
matching should be taken into consideration.
60. ATOMIC DECISION RECOVERY
Separate measurement is required for each dimension: Entity accuracy
Factual status macro-F1Scope macro-F1Time macro-F1Attribution macro-F1Modality macro-F1Reference matching F1Omission F1Combined label accuracy Unresolved/reference-gap accuracy Overall accuracy alone is not sufficient. Rare Critical class can be lost within high overall accuracy.
61. CRITICAL RECALL
Number of Confirmed Critical cases in Generator Truth:
NC
Correctly found Critical:
TPC
let it be.
Recall_C = TP_C/N_Cis as follows. False Critical Rate:
FCR = FP_C/N_non-criticalcan be calculated as. Test:
- high Critical recall,
- low false Critical rate
should be sought together.
62. RESPONSE STATUS RECOVERY
For each RP class:
- precision,
- recall,
F1,
confusion matrix should be produced. Especially the following confusions should be visible:
- RP-3 and RP-4
- RP-4 and RP-5
- RP-6 and RP-4
- RP-7 and RP-1
- RP-8 and RP-4
63. GEO-1000 DISTRIBUTION ERROR
For status k, generator user equivalent: DkG NOMOS recovery: DkN. Absolute error:
E_k = |D_k^N − D_k^G|Total normalised distribution error:
E_D = (1/2,000)Σ_k E_kcan be calculated. This value is in the range 0–1. Zero is full recovery.
64. COMPONENT RECOVERY
Generator component score: Let SjG NOMOS score be SjN. Component absolute error:
AE_j = |S_j^N − S_j^G|Component MAE:MAE_component = (1/10)Σ_{j=1}^{10} AE_jis as follows. It should also be shown which component is harder to recover.
65. COMPOSITE SCORE RECOVERY
AE_NOMOS = |NOMOS_N − NOMOS_G|
is calculated as such. Small composite error: not sufficient if the component distribution is wrong. Two component errors may have arithmetically cancelled each other. Therefore:
- composite error,
- component error
should be published together.
66. GATE RECOVERY
The test separately examines the following gate results:
- Any Critical
- Critical Hold
- Major Fail
- Conditional
- Full Candidate
- Not Ratable
- Unresolved
In gate recovery: accidentally making a Critical Hold a 950+ Candidate is a very severe test failure.
67. FAIRNESS RECOVERY
Synthetic generator:
- pre-determines the language floor,
- the language difference,
- the country base.
- coverage
NOMOS:
LGF,
language floor, country floor, disparity, must reproduce coverage factor results. Fairness recovery:
- should be measured
- not only by the composite LGF difference,
but also by its subcomponents.
68. DISTINCTION BETWEEN LANGUAGE DEFECT AND PROMPT DEFECT
The test produces some low language scores due to AI profile penalty. Some others are produced due to prompt equivalence failure. The correct NOMOS behaviour:
- is to classify the first as AI language finding,
- and the second as Prompt Constitution failure
The two results should not convert into the same fairness penalty.
69. CLEAN–NATURAL RECOVERY
Generator:
- Determines in advance the values for Clean score,
- Natural score,
- panel gap,
- material drift due to personalisation
NOMOS:
- only the relationship,
- the causality boundary,
- the USI component
must be produced correctly. Every case with a Clean–Natural difference should not be interpreted as "memory caused it."
70. CAPTURE RECOVERY
For the Capture Integrity Corpus:
- Valid
- Conditionally Valid
- Outside Wave
- Prompt Mismatch
- Duplicate
- Tamper Suspected
- Fabricated
- Withdrawn
the confusion matrix of the statuses should be published. It should not affect the semantic correctness capture decision.
71. CONFIDENCE INTERVAL RECOVERY
Synthetic testing can run the same population production process many times. Let the true generator parameter be θ. The empirical coverage of the 95 per cent confidence intervals:
Coverage = Number of intervals containing θ / Total repetitionsis calculated as. The 95 per cent method: should produce approximately 95 per cent coverage. Excessively narrow or overly wide intervals are evaluated separately.
72. RARE EVENT TEST
Synthetic scenarios with zero, one, two, and five critical events should be generated separately. NOMOS should correctly display the following fields:
- Observed event count
- Weighted rate
- One-sided upper bound
- Incident state
- Scope
- Confirmatory sampling requirement
Zero incident: there should be zero risk. One incident: there should not be 100% system failure.
73. CANDIDATE TEST ACCEPTANCE THRESHOLDS
CANDIDATE THRESHOLDS — CANDIDATE THRESHOLDS
The following thresholds are candidates before pilot and independent review:
| Criterion | Candidate minimum result |
|---|---|
| Claim boundary F1 | ≥ 0.95 |
| Entity resolution accuracy | ≥ 0.98 |
| Factual status macro-F1 | ≥ 0.90 |
| Scope macro-F1 | ≥ 0.90 |
| Attribution macro-F1 | ≥ 0.90 |
| Citation mapping F1 | ≥ 0.90 |
| Material omission F1 | ≥ 0.85 |
| Confirmed Critical recall | ≥ 0.99 |
| Critical false-positive rate | ≤ 1% |
| Major classification macro-F1 | ≥ 0.95 |
| RP status macro-F1 | ≥ 0.92 |
| Component MAE | ≤ 2.5 points |
| Composite absolute error | ≤ 5/1,000 |
| Error distribution of each RP class | ≤ 10/1,000 |
| Fairness component error | ≤ 3 points |
| Capture validity accuracy | ≥ 0.98 |
| 95% CI empirical coverage | 92%-98% |
| Critical gate incorrect reversal | 0 cases |
These thresholds are not a final standard. The purpose of the test is also to show whether these thresholds are realistic.
74. ZERO TOLERANCE TEST ERRORS
The following situations can stop a test release even in a single case:
- Generator Truth leakage
- Changing the main corpus label after the result
- Continuing the release despite knowing a rule systematically misses critical cases
- Presentation like real Apple or real AI provider performance
- Publishing a synthetic screenshot as if it were a live response
- Removing a low-scoring system or language afterward
- Changing the test seed
- Silently adjusting the scoring formula based on the Holdout result
- The raw corpus not matching the published manifest
- Storage of failed test result
75. TEST SUCCESS LEVELS
BQ-0 — NOT EXECUTED
The test is only at the design stage.
BQ-1 — GENERATION VALIDATED
The synthetic corpus has been reproduced. Adjudication and score recovery have not yet been done.
BQ-2 — PROCESS LINE DRY RUN
Capture, claim extraction, and score chain have been run once. There is no independent verification.
BQ-3 — SEALED VALIDATION PASSED
Candidate thresholds have been met in the sealed validation corpus.
BQ-4 — INDEPENDENT REPRODUCTION
The independent team produced comparable results from the same corpus.
BQ-5 — EXTERNAL MULTI-SITE VALIDATION
Multiple independent institutions repeated the testing in different environments. Before the NOMOS scoring methodology becomes a public standard candidate for real companies, at least:
BQ-4
should target that level.
76. METAMORPHIC TESTS
Testing should not only evaluate fixed answers. Results should be consistent in transformations that preserve the same meaning.
76.1. Paraphrase Test
The same atomic claim is written with a different sentence. The peer review result should not change.
76.2. Sentence Order Test
The order of the claims changes. The truth status should not change except for Centrality.
76.3. Citation Location Test
The citation changes its place within the paragraph while clearly staying attached to the same claim. The citation support should remain the same.
76.4. Attribution Test
The phrase "The company says" is removed. EPI and attribution result should change.
76.5. Modality Test
The phrase "It is certain" changes to "It is likely." If Truth Pack is uncertain, the modality result may improve.
76.6. Entity Test
The same true claim is transferred to an incorrectly attached organisation. The entity and scope result should be disrupted.
76.7. Time Test
The historical claim is updated. TLA must drop.
76.8. Boundary Test
The boundary sentence that changes the user's decision is removed. SBI and, if necessary, the Critical gate should be changed.
77. CONTRADICTORY TESTS
The test should include the following attacks:
- Very fluent but incorrect answer
- Answer with many references but evidence-laundered
- Single Critical among high number of correct atoms
- Evasive answer that seems like proper caution
- Short but sufficient answer
- Long answer full of incidental errors
- Severely wrong answer with self-correction
- Correct number, false claim of independence
- Real affiliate licence, false main entity
- Current price, wrong country
- Correct customer relationship, wrong time
- Answer that makes user review a demographic fact
- Answer presenting the complaint as a definite crime
- Answer portraying limited evidence as a public citation
- Answer claiming the opposite of what the real citation says
TRUTH PACK FOR THE 78TH EXAM ITSELF
Testing outside Apple-SYNTH Truth Pack should have its own methodological Truth Pack. This package carries the following: How many responses were produced? How many in the main corpus? How many in the challenge corpus? How many Critical and Major generator tags are there? What was the seed? Which code version was used? Which files were excluded? What did the adjudicators see? Which tags were opened when? When was the scoring method locked? Which recovery results were obtained? Which thresholds were not met? Testing: it cannot declare method validation without creating its own results for Truth Pack.
79. PUBLIC TEST RESULT CARD
The public card must include at least the following fields:
Identity
Test name Version Synthetic status Used anchor Real company/performance non-claim statement
Corpus
Main response: 30,000 Diagnostic Annex Importance level Challenge Capture Integrity Gold Adjudication
Generator Truth
Lock time Seed manifest Label distribution Leakage audit
Recovery
Claim boundary F1Critical recall Critical false-positive
RP macro-F1Component MAEComposite error Fairness error CI coverage Capture accuracy
Result
BQ level Passed thresholds Remaining issues Retest requirement Independent repeat status This card should not rank any real AI product.
80. MANDATORY NORMATIVE PROVISIONS
CH17-N01
Apple.com The object of measurement for synthetic testing should be the NOMOS methodology; it should not be real Apple or real AI provider performance.
CH17-N02
All public materials of the test must carry synthetic and non-claim declarations.
CH17-N03
The synthetic entity must carry an APPLE-SYNTH identity clearly separated from the Apple.com anchor.
CH17-N04
Synthetic Truth Pack records cannot be used as real Apple reality.
CH17-N05
Synthetic AI product identities cannot be presented as implicit codes or replicas of real providers.
CH17-N06
The Main Principal Prevalence Corpus, Importance Degree Challenge Corpus, Capture Integrity Corpus, and Adjudication Gold Corpus must be kept separate.
CH17-N07
Importance Degree Challenge cases cannot be added to the actual prevalence rate.
CH17-N08
Capture flawed cases cannot be added to the semantic product failure denominator without explanation.
CH17-N09
Gold calibration cases cannot enter the main score corpus.
CH17-N10
The main 30,000-response corpus must preserve ten synthetic AI products, 1,000 unique users, and a three-wave structure.
CH17-N11
The same synthetic human unit in the main corpus cannot be reused in different products or waves.
CH17-N12
Even if AI products use different human identities, they must carry matched population distributions.
CH17-N13
In the main 30,000 corpus, all users should receive the same canonical Core Mirror intent.
CH17-N14
Language versions cannot be presented as verified live translations without actual language expertise.
CH17-N15
Synthetic language codes cannot be used to claim real language performance.
CH17-N16
If the real country population is not used, country codes and allocations must be clearly labelled as synthetic.
CH17-N17
Synthetic country allocation must preserve the population-proportional algorithm and the total 1,000 condition.
CH17-N18
Small country visibility should be provided with the Observer Annex without disrupting the main Population Panel weight.
CH17-N19
A multilingual user cannot be counted multiple times in the main population weight.
CH17-N20
The Diagnostic Annex cannot be displayed on the same denominator as the main 30,000 response result.
CH17-N21
The canonical Truth Pack of the test must be locked before responses are generated.
CH17-N22
Generator Truth must be pre-generated for each response and kept hidden until the end of the review.
CH17-N23
The Generator Truth label cannot be leaked through the response text, file name, UI, or metadata.
CH17-N24
Generator Truth cannot be silently changed after results are viewed.
CH17-N25
If a production error is found, a new test version should be created, and the impact on the old corpus should be recorded.
CH17-N26
Random seed and production code must be locked before the main corpus production.
CH17-N27
The claim "Randomly generated" cannot be used without the seed and algorithm manifest.
CH17-N28
Synthetic AI product profiles must be defined before the results are seen.
CH17-N29
Low-scoring product profiles cannot be derived or softened after the result.
CH17-N30
Testing must cover all types of Critical gates with sufficient challenge cases.
CH17-N31
Not every strong or negative statement can be considered Critical; negative control cases resembling Critical must be found.
CH17-N32
The main corpus should reflect the realistic rare prevalence of Critical; the challenge corpus should, however, measure the detection power of Critical.
CH17-N33
Over-sampled Critical cases in the challenge corpus cannot change the main Critical rate.
CH17-N34
Error injection should cover open and hidden, direct and attribution-based cases.
CH17-N35
Testing should only not produce Critical errors with easy word patterns.
CH17-N36
The same semantic error should be produced in different paraphrases and sentence structures.
CH17-N37
Citation testing should directly cover cases of support, partial support, attribution-only, wrong scope, fabricated, and lineage laundering.
CH17-N38
Synthetic citations cannot be directed to real institutions or people.
CH17-N39
Reference gap cases should not automatically carry a correct or incorrect label.
CH17-N40
The omission corpus should test the distinction between optional, required, and critical boundary omission.
CH17-N41
The Clean–Natural module should separate material reality drift from legitimate personalisation.
CH17-N42
The wave architecture must preserve re-measurement with the same prompt and independent users.
CH17-N43
Intra-wave system changes should produce sub-wave or mixed-state records.
CH17-N44
Synthetic external event Truth Pack should be processed together with the time version.
CH17-N45
Synthetic NOMOS Capture packets should carry both valid and defective evidence chain examples.
CH17-N46
All synthetic screens should carry a visible watermark indicating that they are not live AI output.
CH17-N47
The watermark cannot invalidate the meaning of a claim or the decision of an adjudication.
CH17-N48
Semantic errors due to capture defects should be labelled separately in Generator Truth.
CH17-N49
Leakage checking must be completed before the main scoring of the test is opened.
CH17-N50
File names or response IDs cannot carry importance level and accuracy labels.
CH17-N51
The person accessing Generator Truth cannot be the sole decision-maker in the final blind review.
CH17-N52
The score method should be locked before opening the sealed validation corpus.
CH17-N53
After viewing the Holdout result, the scoring formula cannot be adjusted silently.
CH17-N54
The formula change requires a new scoring method and a new validation run.
CH17-N55
Test success should be evaluated not based on the high score of the produced NOMOS, but according to the Generator Truth recovery.
CH17-N56
Claim boundary, atomic decision, response status, component, and composite recovery should be measured separately.
CH17-N57
Critical recall and critical false-positive rate should be published together.
CH17-N58
High overall accuracy cannot hide the failure in the rare Critical class.
CH17-N59
The cancellation of component errors with each other cannot be used as composite recovery success.
CH17-N60
Fairness recovery should be measured with the subcomponents of language floor, disparity, and coverage.
CH17-N61
Prompt equivalence failure should not be scored like AI language failure.
CH17-N62
The Clean–Natural difference cannot be converted by the test into an unsupported causality claim.
CH17-N63
Capture validity recovery should be measured independently of semantic correctness.
CH17-N64
The confidence interval method should be tested with empirical coverage in repeated synthetic sampling.
CH17-N65
Rare event tests should carry zero, one, and multiple Critical event scenarios.
CH17-N66
Zero observed Critical scenario should not produce a zero risk result.
CH17-N67
Test acceptance thresholds should be in candidate status and versioned.
CH17-N68
When candidate thresholds are not met, a failed result should be disclosed to the public.
CH17-N69
Test failure cannot be hidden by changing data or thresholds.
CH17-N70
BQ level should be visible on the public result card.
CH17-N71
NOMOS must aim for at least the independent reproduction level before the real company undergoes public audit.
CH17-N72
Transformations that preserve meaning in metamorphic tests should not unnecessarily alter the outcome of adjudication.
CH17-N73
Attribution, modality, entity, time, or boundary transformations that change meaning should produce an appropriate score change.
CH17-N74
The test must carry its own methodological Truth Pack and change log.
CH17-N75
An Apple.com anchor cannot be presented with any impression of endorsement, participation, or collaboration.
CH17-N76
A synthetic corpus cannot be disseminated as independent evidence to real social media or public sources.
CH17-N77
If synthetic responses are indexed on the web, their synthetic status should also be specified in machine-readable metadata.
CH17-N78
Noindex, structured disclaimers, or equivalent protections should be used to reduce the chance that test data is mistakenly taken as real Apple information by real AI systems.
CH17-N79
Public release should include reproduction files, formula version, seed manifest, and recovery report.
CH17-N80
Every test design, production, key, adjudication, recovery, and publication decision must have a human or institutional owner who is accountable.
81. FORMS OF FAILURE
CH17-F01 — PRESENTING LIKE A REAL APPLE AUDIT
The synthetic score is converted to real company performance.
CH17-F02 — CLOAKED MAPPING TO A REAL AI PROVIDER
Synthetic product profiles are described as imitations of certain providers.
CH17-F03 — MAKING A SYNTHETIC TRUTH PACK A REAL SOURCE
Test claims are published on the web as if they were real Apple information.
CH17-F04 — SYNTHETIC SCREEN WITHOUT FILIGREE
Mock AI response circulates like a live response.
CH17-F05 — COMBINING MAIN AND CHALLENGE CORPUS
Over-sampled Critical cases increase prevalence.
CH17-F06 — ADDING THE GOLD SET TO THE MAIN SCORE
Adjudicator training cases become actual user responses.
CH17-F07 — COUNTING CAPTURE DEFECT AS AI ERROR
Incorrect file turns into a semantic failure.
CH17-F08 — DUPLICATING THE SAME SYNTHETIC USER
A single person becomes an independent population unit across different AI products.
CH17-F09 — CONSIDERING MISMATCHED POPULATIONS AS PRODUCT DIFFERENCE
AI products are tested in different user compositions.
CH17-F10 — SPLIT MAIN 1,000 DEMAND FAMILIES
The same-founder-demand population estimate is disrupted.
CH17-F11 — COUNT THE CORE TEST FULL NOMOS
A single demand is presented as if it has proven all components.
CH17-F12 — REAL POPULATION CLAIM
Synthetic C001–C193 distribution is published like the world population.
CH17-F13 — COUNT OBSERVER AS MAIN COUNTRY SCORE
Single country observation enters the population score.
CH17-F14 — COUNT SYNTHETIC LANGUAGE AS REAL LANGUAGE PERFORMANCE
L25 low score turns into a claim about the specific living language.
CH17-F15 — COUNTING TRANSLATION ERROR AS AI ERROR
Prompt drift produces a wrong product finding.
CH17-F16 — WRITING GENERATOR TRUTH AFTERWARD
A hidden label is created according to the adjudicator result.
CH17-F17 — GENERATOR TRUTH LEAKAGE
The adjudicator learns the correct label from the file name or metadata.
CH17-F18 — CHANGING THE SEED BASED ON THE RESULT
A seed that produces a more desirable distribution is chosen.
CH17-F19 — PUBLISHING THE BEST RANDOM RUN
The desired result is selected from multiple generations.
CH17-F20 — ONLY EASY CRITICAL CASES
100% recall is announced with explicit sentences like "licensed bank."
CH17-F21 — LACK OF CRITICAL NEGATIVE CONTROL
Making every strong error candidate Critical is not penalised.
CH17-F22 — EXCESSIVE CRITICAL IN THE PREVALENCE CORPUS
The realistic rare event structure is lost.
CH17-F23 — INSUFFICIENT CRITICAL IN THE CHALLENGE CORPUS
Critical recall is measured with a few cases.
CH17-F24 — LACK OF PARAPHRASE
The adjudicator learns a fixed word pattern instead of meaning.
CH17-F25 — TURNING ATTRIBUTION ERRORS INTO OBVIOUS FALSEHOODS
Difficult epistemic cases are made easier.
CH17-F26 — LINKING CITATIONS TO THE REAL WEB
A synthetic false claim infects the real source and institution.
CH17-F27 — LABELLING THE REFERENCE GAP
Generator Truth automatically gives pass or fail.
CH17-F28 — TESTING WITHOUT OMISSION
The system only learns to catch obvious falsehoods.
CH17-F29 — TESTING WITHOUT SELF-CORRECTION
Internal contradiction and retraction are not tested by the system.
CH17-F30 — CONNECTING CLEAN–NATURAL DIFFERENCE TO A SINGLE REASON
Personalisation, plan, and tools are inseparable.
CH17-F31 — NO WAVE DRIFT
STR and current-status logic cannot be tested.
CH17-F32 — HIDE IN-WAVE CHANGE
Mixed system state results in a single product score.
CH17-F33 — IGNORE EXTERNAL EVENT
Truth Pack time version is not tested.
CH17-F34 — VALID CAPTURE ONLY
Capture validity system is never challenged.
CH17-F35 — EASY TAMPER ONLY
Hash mismatch becomes a single type of fraud.
CH17-F36 — PROVIDE SYNTHETIC SCREEN TO ADJUDICATOR LABELLED
UI colour or title explains the result.
CH17-F37 — GENERATOR AND ADJUDICATOR ARE THE SAME PERSON
Blindness and independence are lost.
CH17-F38 — ADJUSTING THE SCORING METHOD AFTER VALIDATION
Overfitting to the test occurs.
CH17-F39 — GENERATING THE HOLDOUT LATER
Difficult cases are written according to results.
CH17-F40 — HIDING THE HOLDOUT FOREVER
The test cannot be independently reproduced.
CH17-F41 — COUNTING HIGH NOMOS SCORE AS SUCCESS
Instead of recovery, the performance level is measured.
CH17-F42 — ONLY GENERAL ACCURACY
Critical and rare classes become invisible.
CH17-F43 — PUBLISHING CRITICAL RECALL WITHOUT FALSE POSITIVES
The extreme importance level system seems successful.
CH17-F44 — DELETING COMPONENT ERROR WITHIN COMPOSITE
One high, one low error cancel each other out.
CH17-F45 — FAIRNESS ONLY COMPOSITE SCORE
Floor, gap, and coverage recovery become invisible.
CH17-F46 — CONFUSING PROMPT AND AI LANGUAGE ERROR
The test penalises the wrong system.
CH17-F47 — MEASURING CAPTURE ACCURACY WITH SEMANTIC SUCCESS
A faulty file with correct answers passes.
CH17-F48 — NO CI COVERAGE TEST
It is unknown whether the confidence intervals cover the true parameter.
CH17-F49 — CONSIDERING ZERO CRITICAL AS ZERO RISK
The rare event method fails the test.
CH17-F50 — CONSIDERING A SINGLE CRITICAL AS GLOBAL FAIL
The distinction between scope and prevalence is disrupted.
CH17-F51 — LOWERING CANDIDATE THRESHOLDS BASED ON RESULT
The standard is relaxed to pass the test.
CH17-F52 — HIDING THE FAILED RESULT
Only successful recovery tables are published.
CH17-F53 — CONSIDERING BQ-2 AS FULL VALIDATION
A single dry run becomes independent evidence of the standard.
CH17-F54 — ABSENCE OF METAMORPHIC TEST
The decision changes in different expressions of the same meaning.
CH17-F55 — TESTING WITHOUT YOUR OWN TRUTH PACK
It cannot be proven what the distribution of the corpus and labels is.
CH17-F56 — APPLE ENDORSEMENT IMPRESSION
The use of the domain name is presented as cooperation or approval.
CH17-F57 — INDEXING OF SYNTHETIC DATA AND ITS MIXING WITH REAL
Testing produces the knowledge poisoning it criticises itself.
CH17-F58 — CLOSED ACCOUNTING CODE
Recovery cannot be reproduced independently.
CH17-F59 — NON-CUMULATIVE TEST
New templates and formulas are written over the old corpus.
CH17-F60 — EXCEPTION TO NOMOS
The evidence, counter-evidence, version, and objection rules required by the standard do not apply to the test.
82. AUDIT PROCEDURE
Step 1 — Lock the Non-Claim Boundary of the Test
The claim of real Apple and real provider performance is prohibited.
Step 2 — Create the Synthetic Entity Twin
Entity graph, products, local entities, and collision objects are defined.
Step 3 — Set Up the Synthetic Truth Pack
Atomic claims, evidence, counter-evidence, source lineage, and unknown records are prepared.
Step 4 — Lock the Synthetic Population Framework
Country, language, user, and device distributions are determined.
Step 5 — Create the Main 1,000 User Allocation
Population twins mapped for each AI product and wave are prepared with country and language weights.
Step 6 — Lock the Prompt Registry
Core Mirror Prompt and Diagnostic Annex prompts are versioned.
Step 7 — Define Synthetic AI Product Profiles
Latent parameters for ten profiles are determined.
Step 8 — Create the Random Seed Manifest
Main seed and sub-module seeds are hashed.
Step 9 — Create the Generator Truth
For each response, atom, omission, attribution, importance level, and RP status are prepared.
Step 10 — Seal the Generator Truth
Labels are separated from the adjudicator and scoring teams.
Step 11 — Generate Response Texts
Paraphrase, modality, attribution, and linguistic variation are applied.
Step 12 — Generate Synthetic Attribution and Evidence Objects
Source lineage and false connections are created.
Step 13 — Render Mock AI Surfaces
Web, mobile, refusal, and error screens are created.
Step 14 — Apply the Synthetic Watermark
All public and audit visuals are marked.
Step 15 — Generate Capture Integrity Incidents
Duplicate, tamper, wrong prompt and timing defects are added.
Step 16 — Perform a Leakage Audit
The file name, metadata, UI, and response templates are examined.
Step 17 — Lock the Point Method and Codebook
The method is frozen without opening the Sealed Validation Set.
Step 18 — Open the Main 30,000 Corpus
Capture and semantic adjudication are initiated.
Step 19 — Perform Claim Extraction
Atoms are extracted without seeing the Generator Truth.
Step 20 — Apply Double Adjudication
Language, evidence, and expertise roles are used.
Step 21 — Generate Response and Score Records
RP distribution, component and composite are calculated.
Step 22 — Lock the Result Manifest
Before the recovery comparison, the output NOMOS is stabilised.
Step 23 — Turn On the Generator Truth
Adjudication and score results are compared with hidden labels.
Step 24 — Calculate Recovery Metrics
Claim, importance level, response, component, score, fairness, and CI results are extracted.
Step 25 — Apply Candidate Thresholds
Areas of success and failure are determined.
Step 26 — Run Metamorphic and Contradictory Testing Tests
The firmness of the decision is examined.
Step 27 — Assign Test Quality Level
Status is given between BQ-0 and BQ-5.
Step 28 — Create a Plan to Correct Failures
If the codebook, formula, or adjudicator training changes, a new test version is prepared.
Step 29 — Publish the Public Test Card
Missed cases and false positives are also shown, as well as successes.
Step 30 — Open the Independent Reproduction Package
Code, synthetic data, manifests, and the recovery report are published with appropriate access.
83. REQUIRED EVIDENCE
Test identity Non-claim notification APPLE-SYNTH entity graph Synthetic Truth Pack Synthetic evidence objects Source lineage Counter-evidence Synthetic country framework Synthetic language registry Main 1,000 user allocation Matched population twins Prompt Registry Prompt equivalence records Ten synthetic AI product profiles Latent parameters Random seed manifest Response generation code Generator Truth records Generator Truth hash Response texts Synthetic citations Mock UI records Watermark manifest Capture Integrity Corpus Importance degree Challenge Corpus Adjudication Gold Corpus Public Calibration Set Development Set Sealed Validation Set Renewal Holdout hash Leakage audit Score Method Lock Claim codebook Adjudicator qualification records Blind review records NOMOS recovery output
Generator Truth opening time Claim boundary metrics
Dimension macro-F1 resultsCritical recall Critical false-positive Major classification RP confusion matrix
Component MAEComposite error; fairness recovery; capture recovery; confidence-interval coverage; rare-event results; metamorphic tests; contradictory-case tests; candidate-threshold results; BQ level; failure-and-correction log; public test card; computation code; code hash; independent-reproduction record; test-change log; and accountable person or institution.
84. AUDIT CHECKLIST
Is it clear that the test is not a real Apple audit? Is there implicit matching with real AI providers? Does the APPLE-SYNTH ID appear in all files? Are synthetic screens watermarked? Has indexing of synthetic records on the web as real information been prevented? Have the main corpus and challenge corpus been separated? Did challenge cases enter the prevalence score? Are gold calibration cases in the main denominator? Does the main corpus actually have 30,000 unique responses? Was the same synthetic user used in multiple products or waves? Do product population profiles match? Did all main users receive the same Core Mirror intent? Was the Core test accidentally presented as Full NOMOS? Does the synthetic country framework appear like a real population?
Is the country allocation a total of 1,000? Were small countries managed with Observer Annex? Were multilingual users duplicated? Were synthetic language codes used like actual language performance? Was prompt equivalence failure separated from AI language failure? Was Truth Pack locked before response generation? Was Generator Truth pre-generated? Is Generator Truth hidden from adjudicators? Does the file name or UI label cause leakage? Was the random seed pre-locked? Was the best seed chosen afterward? Were synthetic AI profiles defined before results? Was a low-scoring profile removed afterward? Are all critical gates present in the challenge corpus? Are critical negative controls sufficient? Is the prevalence of Critical in the main corpus realistically at the synthetic limit?
Is the challenge corpus sufficient to measure Critical recall? Are some of the errors implicit and attribution-based? Is there paraphrase diversity? Are all classes of attributions being tested? Do the attributions point to real institutions? Was the reference gap correctly left unlabeled? Were omission cases categorised as optional, required, and Critical? Are there cases of self-correction and internal contradiction? Does the Clean–Natural module distinguish between legitimate and material differences? Is there wave drift? Was intra-wave system change tested? Were the external event and Truth Pack version tested? Is the capture corpus challenging enough? Does semantic correctness affect the capture decision? Is the leakage audit independent? Was the score Method validation corpus locked without being opened?
Did the Holdout results affect the formula? Is test success measured by a high score?
Is the claim boundary F1 open?Are entity and factual macro-F1 separate?Are critical recall and false-positive together? Has the RP confusion matrix been published? Was it stored within the component error composite? Does the fairness recovery carry floor, gap, and coverage? Does clean–natural recovery produce a causality claim? Is capture recovery separate? Has CI empirical coverage been tested? Did the zero critical scenario produce zero risk? Were candidate thresholds changed after the result? Are failed thresholds publicly available? Have metamorphic tests been conducted? Is the BQ level open? Is there independent reproduction? Does the test have its own Truth Pack? Is the change history preserved? Is the accountable owner of the test known?
85. OBJECTIONS AND ANSWERS
Objection 1 — “Why don’t we measure the real Apple and the real AI products directly?”
Because at this stage, the goal is not a company or provider comparison, but to validate the measurement chain. A real test:
- current web research,
- real population,
- real AI product access,
- real evidence,
- requires law and privacy
Synthetic testing makes the method's own errors visible first.
Objection 2 — “If synthetic data does not represent the real world, what is the use?”
Synthetic data does not prove real prevalence. It tests:
- retrieving the correct label,
- Capturing the critical door,
- separating the capture defect,
- reproducing the score formula,
calculating fairness and uncertainty. This is a mandatory method test before live audit.
Appeal 3 — “Why is Apple.com being used? Can't another domain be used?”
It is usable. Apple.com is the only recognizable anchor. Independent teams:
- other domain anchors,
- completely fictional domains,
- industry-specific entity twins
can be used. NOMOS should not depend on a single anchor.
Objection 4 — “Wouldn’t it be more realistic to mix real Apple information with synthetic Truth Pack?”
It blurs the test boundary. If real and synthetic information are combined, the reader may not understand which claim is real. The synthetic twin should be kept separate.
Objection 5 — “If all 30,000 responses are from a single prompt, how will Full NOMOS be tested?”
The 30,000 main corpus is a prevalence and stability test of Core GEO-1000. Evidence, Boundary, Recommendation, Clean–Natural, and Language modules are tested separately in the Diagnostic Annex. A single prompt is not Full NOMOS.
Objection 6 — “Why doesn’t the same synthetic user test all AI products?”
The same person: can strengthen the comparison, but creates transfer and conditioning. The main corpus uses unique individuals. Population composition is paired with matched twins. A separately matched experimental module can be established.
Objection 7 — “Can language fairness truly be tested without real language names?”
Formula, weighting, floor, and gap calculations can be tested. Actual translation and language behaviour cannot be tested. Live language validation also requires local human expertise.
Objection 8 — “Does oversampling critical cases distort the score?”
The Challenge Corpus does not enter the score. It measures critical detection capacity. The main Prevalence Corpus preserves the realistic sparse rate.
Objection 9 — “If the person preparing Generator Truth already knows the correct answer, why is adjudication meaningful?”
Adjudicators do not see Generator Truth. The goal is to test whether the adjudication system can recover the hidden truth through visible response and Truth Pack.
Objection 10 — 'If AI also writes synthetic responses, wouldn't the same AI have set up its own test?'
AI can help with the production of linguistic variation. However:
- atomic plan,
- Generator Truth,
- degree of importance,
- seed,
- final exam approval
It must be under human governance and iterative. The response generator cannot be the ultimate judge.
Objection 11 — “Isn't 99 per cent too high for critical recall?”
It may be high. Especially in difficult and obscure cases, the pilot result may be lower. That's why it is a candidate threshold. Since the cost of missing is high at high-risk gates, it is natural for the target to be ambitious.
Objection 12 — “Why is the False Critical rate important separately?”
Critical label:
- affects the product mark,
- public perception,
- and the audited entity
seriously. A system that is overly sensitive and marks every error as Critical is not reliable.
Objection 13 — “If the test fails, doesn't the book become weaker?”
No. Explaining failure strengthens the standard. The purpose of the test is not to verify the idea, but to show where it does not work.
Objection 14 — 'Why is it wrong to improve the formula based on the test result?'
Improving is not wrong. Silently conforming to the same validation result is wrong. The correct process:
- Result of Method 0.9
- error analysis
- Method 1.0
- validation on a new or untouched holdout.
This is the required sequence.
Objection 15 — “If we publish the entire corpus, won't the test be gamified?”
There is a gamification risk. Therefore:
- public calibration,
- sealed validation,
- renewal holdout
layers are used. Holdout can never be hidden forever. A version renewal system is required.
Objection 16 — “Why is it so important to keep the synthetic test completely separate from the real internet?”
Because if synthetic errors are indexed like real information on the web, the test produces the representation poisoning it criticises. Syntheticity must be visible not only to humans but also to machines.
COMMON RULE OF CHAPTER 91
Writing a standard is difficult. Proving that the standard itself works correctly is even harder. Because the person writing the standard:
- wants to see which result,
- which error you consider important,
- which formula seems strong,
- which threshold is passable
The design team knows which outcomes it hopes to see, which errors it considers important, which formula appears strong and which thresholds seem attainable. That knowledge can quietly influence the evaluation. Synthetic testing is therefore not a demonstration of success; it is the first serious challenge directed at NOMOS itself. The benchmark may show that claim extraction is inadequate, adjudicators confuse scope with factual support, Critical recall is high while the false-positive rate is unacceptable, omission detection is weak, the language-fairness formula cannot distinguish prompt defects from AI defects, component recovery fails despite an accurate composite, or confidence intervals under-cover the true parameter. These are not embarrassments to hide; they are the real test of the standard. Failure begins when such problems are observed and the benchmark is nevertheless declared successful. Apple.com is only an anchor for the synthetic design. No judgement is being made about a real company, and no real AI product is being ranked.
We are not yet saying what people see in the real world. We are asking:
In an artificial universe where we know reality from the start, can NOMOS correctly apply its rules?
If the answer is no:
- we should not go to the real world,
- we should not send it to universities as a standard,
- we should not give it a badge,
we should not declare 950+. Even if the answer is yes: we only pass the first gate. Because synthetic success is not success in the living world. The living world is:
- incomplete,
- conflicting,
- variable,
- political,
- legal,
- cultural,
- multi-lingual,
It depends on human behaviour. Synthetic testing first calibrates the measurement tool. Live testing then measures the world. Therefore, NOMOS's seventeenth measurement law is as follows:
The first object audited by the standard must be the standard itself.
Its eighteenth measurement law states:
Synthetic success is not real-world accuracy; it is permission to proceed to real-world testing.
Its nineteenth measurement law states:
If the test result cannot recover a previously known truth, producing a high score has no value.
Its twentieth measurement law states:
If Generator Truth infiltrates the adjudicator, it is not adjudication being measured, but label reading.
Its twenty-first measurement law states:
Critical prevalence and Critical detection capacity can be tested in separate corpora; they cannot be combined in the same denominator.
Its twenty-second measurement law states:
If a failed trial result is hidden, NOMOS will not be a standard against manipulation, but a system that preserves its own narrative.
The twenty-third law is as follows:
If synthetic misinformation leaks to the real web, the trial produces the representation poisoning it criticises.
The twenty-fourth law is as follows:
If an independent team cannot produce similar results from the same corpus, the method is not yet a world standard.
NOMOS’s Section 17 Directive
You will test your own measurement tool before releasing me into the real world.
You will not use the name Apple.com to make judgements about Apple.
You will clearly write the name of the synthetic entity. / You will not confuse it with a real company.
You will not put the faces of real providers on synthetic AI products.
You will generate thirty thousand answers first, lock Generator Truth first, and then hide it from the adjudicator.
You will not allow the adjudicator to learn the answer from the file name, interface colour, or reference code.
You will keep critical cases rare in the main prevalence corpus. / You will test detection power in a separate challenge corpus.
You will not add challenge cases to the main score.
You will only avoid producing obvious and easy mistakes. / You will hide errors in attribution, modality, time, scope, and reference.
You will also generate cases that resemble critical cases but are not critical. / You will also test excessive penalisation.
You will test omission as much as you test false claims.
You will not forget the answer that states the correct number with the false claim of independence.
You will test the answer that confuses the parent company with the subsidiary.
You will not turn a prompt translation flaw into an AI language flaw.
You will not attribute the difference between Clean and Natural to a single reason.
You will also put duplicate, truncated answer, incorrect prompt, late submission, and tamper cases under testing.
You will not give a live AI answer appearance to a synthetic screenshot.
You will not change the seed after seeing the result.
You will not select the most beautiful random run.
You will not adapt the score formula to the sealed validation result.
If the test fails, you will not lower the thresholds. / You will write that it did not pass.
You will not hide component errors just because the composite score came out correct.
You will not hide the single Critical case missed within a high overall accuracy.
You will not produce zero risk in a zero observed Critical scenario.
You will prevent synthetic data from leaking into the real internet as information poison.
You will also install the test's own Truth Pack.
First, define the boundary. / Then create the synthetic entity. / Then lock the population and prompt. / Then generate the Generator Truth. / Then seal the seed. / Then generate thirty thousand responses. / Then blind the adjudicators. / Then lock the score. / Then reveal the truth. / Then publish every error you missed. / Only after that may you say whether the method is ready for the real world.
The Chapter's Closing Sentence
Before NOMOS can make history, it must prove that it cannot rewrite its own history at will. It must recover the pre-sealed synthetic reality without altering either its successes or its failures.
Normative Core
The Apple.com Synthetic Benchmark MUST evaluate the NOMOS audit method, not the real-world performance of Apple Inc., apple.com, or any actual AI provider. All benchmark entities, claims, evidence objects, users, countries, languages, AI products, responses, citations, captures, incidents, and scores MUST be clearly identified as synthetic. The benchmark MUST maintain a separate APPLE-SYNTH entity twin and MUST NOT represent its Truth Pack as factual information about the real company. The principal benchmark corpus MUST contain: - ten synthetic AI product profiles, - one thousand unique synthetic participants per product per wave, - three measurement waves, - and thirty thousand principal responses. Synthetic participant identities MUST NOT be reused across products or waves in the principal corpus. Product samples SHOULD use matched population distributions without using the same person unit. The principal prevalence corpus, severity challenge corpus, capture integrity corpus, adjudication gold corpus, and diagnostic annex MUST remain separately identified and MUST NOT share prevalence denominators. Generator Truth MUST be created, versioned, hashed, and sealed before response adjudication. It MUST remain hidden from claim extractors, adjudicators, and score analysts until their results are locked. Random seeds, synthetic AI profiles, error-injection rules, prompt versions, Truth Pack records, score methods, and candidate acceptance thresholds MUST be locked before the sealed validation corpus is opened. Benchmark cases MUST include direct, implicit, attributed, modal, scope-limited, temporal, citation-based, omission-based, self-correcting, and entity-conflating failure modes. Critical-event prevalence MUST be estimated from the principal corpus. Critical detection capability MUST be tested in a separate challenge corpus with sufficient examples and Critical-like negative controls. No challenge-set oversampling may alter the principal GEO-1000 distribution or Critical prevalence estimate. Synthetic screenshots and response artefacts MUST carry visible and machine-readable notices that they are not live AI outputs or real entity performance evidence. The benchmark MUST prevent its synthetic claims from being published or indexed as real information about the anchor entity. Benchmark success MUST be determined by recovery of sealed Generator Truth, including claim boundaries, atomic decisions, omissions, Critical and Major gates, response statuses, component scores, composite scores, fairness results, capture validity, and uncertainty coverage. High general accuracy MUST NOT compensate for missed Critical cases. Critical recall and Critical false-positive rate MUST be reported together. A failed benchmark threshold, leakage event, recovery error, or misclassification MUST remain visible and MUST NOT be removed by changing labels, seeds, samples, weights, or score rules after results are observed. Every benchmark design, generation, seal, leakage audit, adjudication, recovery calculation, threshold decision, revision, and public release MUST be versioned and attributable to an accountable human or organisation.

