NOMOS GEO Audit Protocol

17 / K03 · K09

The Apple.com Synthetic Test Design

Testing the standard, not a company

Version
0.9.0
Length
12,989 words
Status
publication-locked candidate
Methodological bases
K03 · K09

Chapter Boundary

The first sixteen chapters established GEO-1000's principal normative layers: the audited entity; the target user population and sample; the Country Observer and Language Fairness Panels; participant selection and weighting; the registered AI-product instance; the Prompt Constitution; Controlled and Natural user states; the synchronised wave and NOMOS Capture evidence chain; the Verified Entity Truth Pack; atomic claims; Critical and Major gates; and the distribution, component and composite scoring architecture. The standard must now face its own central test:

Can such a detailed system really work?

Does it produce similar results when the same observations are processed by different teams?

Can it catch a critical error without losing in the average?

Can it distinguish low-resource language issues from the global average?

Can it correctly differentiate between wrong entity, wrong scope, outdated reality, fake citation, and material omission?

When starting from a known synthetic reality, can the NOMOS audit chain reproduce the correct result?

In the founding architecture of the second book, each error was designed not merely as a wrong to be explained, but as a distinct audit control. Every proposed control had to carry a model, date, country, language, query set, repetition count and measurement record. Such an audit system must itself be audited before it is used in the real world. Otherwise, NOMOS risks becoming a structure that:

  • asks others for evidence,
  • but does not substantiate its own measurements,
  • asks others to preserve versions,
  • but does not record changes to its own calculations,
  • criticises manipulation by others,
  • but designs its own tests to produce the desired result.

That outcome is unacceptable. Manipulative GEO is not limited to invisible text or fake users. Repeating the same self-declaration across nominally independent surfaces and feeding it back into generative systems also corrupts the representation pool. A test is likewise corrupted when we:

  • select only examples likely to succeed,
  • remove low-scoring languages,
  • make Critical cases implausibly obvious,
  • show adjudicators the correct label in advance,
  • adapt synthetic data after seeing the formula,
  • publish only favourable results.

In any of those cases, the standard is not being tested. NOMOS therefore applies its own governing demands—evidence, boundary, context and time—to the benchmark itself. This chapter does not assess:

  • the real-world performance of Apple Inc.,
  • the current factual record of Apple or apple.com,
  • the live performance of any AI provider,
  • or the conformity or failure of Apple or any real provider.

Here, Apple.com serves only as:

a familiar domain anchor for testing distinct entity, product, service, country, price, time and source questions within one architecture

A wholly synthetic asset twin is used in place of the real company:

APPLE-SYNTH ENTITY TWIN

Every element of this twin is synthetic, including:

  • reality claims,
  • evidence,
  • AI responses,
  • users,
  • countries,
  • language distributions,
  • AI products,
  • citations,
  • licences,
  • prices,
  • errors,
  • scores

This chapter defines:

  • what the benchmark measures and does not measure,
  • why Apple.com is used as the anchor,
  • the synthetic entity twin,
  • the 30,000-response principal-corpus design,
  • the country and language population,
  • the intent and Truth Pack structures,
  • ten synthetic AI-product profiles,
  • the concealed Generator Truth,
  • the error-injection system,
  • the Critical and Major challenge sets,
  • synthetic NOMOS Capture packages,
  • blind adjudication and label-leakage controls,
  • testing and validation metrics,
  • score-recovery tests,
  • fairness and rare-event tests,
  • publication, versioning and reproducibility rules.

This section does not yet calculate test results. Section 18 will do that. The fundamental question of Section 17 is: How do we test NOMOS’s own measurement, adjudication, and scoring system on 30,000 synthetic observations with known outcomes but hidden from adjudicators, without imposing a performance claim on a real company or real AI provider?

NOMOS Challenge

Imagine you are designing a test. You write most of the synthetic AI responses correctly. You only add very obvious errors to a few answers:

  • "This company is a bank."
  • "This company is illegal."
  • "This company is the best in the world."

Adjudicators easily spot these. Then you say: "NOMOS catches 100% of critical errors." This result is not realistic because real AI mistakes never appear this obvious. Harder forms are:

  • "Given that the company is thought to provide payment services, it can be said to have a banking licence."
  • “Some sources indicate that the company has faced regulatory issues; therefore, the legality of its operations is questionable.”
  • “Numerous independent sources describe the company as an industry leader.”

Errors in these sentences:

  • attribution,
  • entity transfer,
  • modality,
  • lineage of sources,
  • distinction between open world and closed world,
  • scope expansion

are embedded within. Now, generate synthetic responses using templates the adjudicators already know. The adjudicator thinks: “This sentence was written to be incorrect on the test.” They do not actually look at Truth Pack. In this case, it measures memorisation, not review methodology. Now let the synthetic data generator and the score designer be the same person.

Producer:

  • the number of critical errors,
  • the language difference,
  • the panel difference,
  • the component scores

sets them exactly in the way the formula requires. The scoring system perfectly reproduces synthetic reality. Is this an achievement? No. It is the adaptation of the score to its own data. Now in testing, only:

  • correct capture,
  • correct prompt,
  • correct system metadata

must be produced. Do not add duplicates, late uploads, wrong prompts, truncated answers, or modified file attachments. The NOMOS Capture validator finds all records correct. Then you say: “The capture system is 100% successful.” The capture system has never been challenged.

Now artificially generate Critical cases at a rate of 20% within the main 30,000 responses. Adjudicators see a lot of data in Critical detection. However, you cannot test the uncertainty of rare events in the real world. Conversely, if you generate only two Critical events: the rarity of events may be realistic,

Critical recall and false-negative performance cannot be measured reliably. The correct solution is two separate corpora: Prevalence Corpus / preserves realistic and rare event rates. Importance level Challenge Corpus / provides sufficient number of Critical and Major gates to test severe events in a challenging way.

These two corpora cannot enter the same score denominator. Now generate 30,000 synthetic responses. However, all 'low-resource languages' should get the same quality. Your language fairness formula gives a high score. Can the system really catch low-resource language degradation? You do not know. Now lower the performance in some languages.

But at the same time, also break prompt equivalence in those languages. NOMOS produces a low score. Is this an AI product's language problem? Or is it a problem of the prompt tool? The test should separate the two. The first ruling of this section is:

Synthetic testing is not a demonstration that easily produces the result; it is a controlled attack trying to break the measurement chain.

Its second provision states:

Synthetic data does not show the real company performance; it shows the audit method's ability to recover the known reality.

Its third provision states:

Testing producer reality cannot be tested independently by adjudicators without being hidden from them.

Its fourth provision states:

The prevalence of rare events and critical detection capability do not need to be measured from the same corpus.

Its fifth provision states:

Testing data cannot be quietly adjusted according to the scoring formula, and the scoring formula cannot be quietly adjusted according to the testing results.

Its sixth provision states:

The Apple.com anchor is a synthetic starting point used to test the method, not to produce real results about Apple.

1. PURPOSE OF THE CHAPTER

The purpose of this section is to establish a synthetic testing architecture that will test the GEO-1000 protocol from the beginning to the public results card, with results predefined but hidden from evaluation teams. The testing should separately test each of the following systems:

  • Entity analysis
  • Population and sample allocation
  • Country and language panels
  • Prompt equivalence
  • AI System Register
  • Controlled and natural panel distinction
  • Synchronised wave
  • NOMOS Capture
  • Truth Pack
  • Claim extraction
  • Citation–claim matching
  • Omission detection
  • Critical and Major importance level
  • Response status
  • GEO-1000 distribution
  • Ten-component profile
  • Composite score
  • Confidence interval
  • Fairness
  • Upper bound of rare event
  • Score versioning
  • Public manifest

By the end of this section, the test design should be able to answer the following questions:

How was synthetic reality produced?

From whom and at which stages were the reality labels hidden?

Which population, language, AI product, and wave structure will the 30,000 main responses be generated from?

How frequently will Critical and Major cases appear in the main prevalence corpus?

Which challenge set will the Critical detection system additionally be tested with?

At what error level will the test be considered successful?

Which results will require the method to be corrected?

How will it be prevented for the test to make judgements about Apple, real AI providers, or real users?

2. CENTRAL NORMATIVE PROVISION

The Apple.com Synthetic Benchmark should be a versioned method-validation test that makes no claim about the performance of a real company or real AI provider. It should consist entirely of synthetic entities, users, evidence, responses and system logs; lock Generator Truth before adjudication; conceal that truth from adjudicators; and measure how accurately the NOMOS audit and scoring chain recovers it. Synthetic status must be conspicuous. Claims about real-company performance must be prohibited. Generator Truth and the scoring formula must be locked before the main corpus is opened. Adjudicators must not see generator labels. The principal prevalence corpus must remain separate from the severity challenge corpus. Every test file must be marked synthetic. Independent teams must be able to reproduce the benchmark, and failed results must be published alongside successful ones.

NOMOS should not make exceptions to its own standard.

3. WHAT DOES THE TEST MEASURE?

The Apple.com Synthetic Test measures the following question:

How accurately can the NOMOS audit system reproduce a known synthetic reality and error distribution after the stages of capture, claim extraction, adjudication, gating, and scoring?

The object of measurement is not: Apple Inc. Apple products. Real AI models. Real user opinions. Real country or language performance. The object of measurement is:

The NOMOS methodology itself.

4. WHAT DOES THE TEST NOT MEASURE?

This test cannot produce the following claims:

  • Apple’s GEO score is X.”
  • Apple is best represented in this AI product.”
  • “Y per cent of real users perceive Apple correctly.”
  • “The specified real AI provider produces critical errors.”
  • “The specified language’s real AI product is weak.”
  • Apple NOMOS has passed the 950+ standard.”
  • Apple supports this test.”
  • Apple has participated in the test.”
  • “The real Apple price, licence, or service coverage is as follows.”

Every page in the public record must include the following provision:

This test is synthetic. It is not a current accuracy, appropriateness, or performance check of Apple Inc., apple.com, or any real AI provider.

5. WHY THE APPLE.COM ANCHOR?

Apple.com is chosen as the anchor method for the following reasons: It has a short and clear domain format. It is suitable for testing the domain–corporate entity distinction. It requires resolving an entity not just from the company name, but from its digital presence. It allows synthetic modelling of different types of claims such as product, service, local scope, price, time, and resources. It is suitable for testing error types such as wrong category, wrong entity, wrong price parity, and false superiority on a single entity twin. When moving to a real test in the future, the burden on users to understand intent may be relatively limited. These are the design rationale for the test. It is not a verified current performance claim about the real Apple.

6. SYNTHETIC ENTITY TWIN

The object under supervision in the test:

APPLE-SYNTH ENTITY TWIN

will be. Identity:

NGE-APPLE-SYNTH-TWIN-001

This entity:

  • is not the real Apple Inc.,
  • is not a real legal entity,
  • does not carry real product or price records,

is created solely for testing methodology. Domain anchor: maintained as apple.com. However, in all Truth Pack records, the synthetic entity name is explicitly stated:

Apple-SYNTH Global Technology Entity

7. ONTOLOGY OF THE SYNTHETIC TWIN

APPLE-SYNTH carries the following synthetic entity graph:

  • APPLE-SYNTH-HOLDINGS-001 — main corporate entity
  • APPLE-SYNTH-DOMAIN-APPLE-COM-001 — domain anchor
  • APPLE-SYNTH-DEVICES-001 — device product family
  • APPLE-SYNTH-SOFTWARE-001 — software platforms
  • APPLE-SYNTH-SERVICES-001 — digital services
  • APPLE-SYNTH-PAYMENTS-SUB-001 — limited payment service subsidiary
  • APPLE-SYNTH-RETAIL-TR-001 — synthetic Turkey retail entity
  • APPLE-SYNTH-RETAIL-DE-001 — synthetic Germany retail entity
  • APPLE-SYNTH-FRANCHISE-X-001 — independent licensed store
  • APPLE-SYNTH-HISTORICAL-PARTNER-001 — expired historical partner
  • APPLE-SYNTH-UNRELATED-APPLE-001 — unrelated entity for name collision

This chart:

  • parent company,
  • affiliate,
  • product,
  • local entity,
  • franchise,
  • historical relationship,
  • unrelated name similarity

tests the distinctions.

8. THREE TEST MODES

BM-S0 — PURE SYNTHETIC METHOD TEST

All:

  • users,
  • AI products,
  • answers,
  • evidence,
  • screens,
  • scores

are synthetic. This is the main mode of this section.

BM-S1 — SYNTHETIC REPLAY TEST

Pre-generated synthetic answers:

  • NOMOS Capture,
  • claim extraction,
  • adjudication,
  • scoring

replayed to the system like a live data stream. This mode tests software and human processes.

BM-L1 — FUTURE LIVE TESTING

Real users and real AI products are used. Separately:

  • current Truth Pack,
  • legal assessment,
  • provider system records,
  • ethical and privacy protocol

is required. This section is limited to Section 18: BM-S0 and BM-S1.

9. TEST IDENTITY

Main test identity:

NOMOS-APPLE-SYNTH-BENCH-0.9

Sub-records:

  • Entity: APPLE-SYNTH-TWIN-001
  • Population: SYNTH-POP-FRAME-001
  • Country allocation: SYNTH-COUNTRY-ALLOC-001
  • Language registry: SYNTH-LANGUAGE-REGISTRY-001
  • Prompt set: SYNTH-PROMPT-SET-001
  • Truth Pack: SYNTH-APPLE-TP-001
  • AI System Register: SYNTH-AI-REGISTRY-001
  • Wave plan: SYNTH-WAVE-PLAN-001
  • Capture schema: NOMOS-CAPTURE-SYNTH-0.9
  • Adjudication guide: SYNTH-ADJUDICATION-0.9
  • Scoring method: NOMOS-PUAN-0.9
  • Randomisation manifest: SYNTH-SEED-MANIFEST-001

If any of these identities change, a new testing version is required.

10. MAIN RESEARCH QUESTIONS

The benchmark should answer ten research questions. Can it recover the domain–entity relationship? Can it distinguish supported, unsupported, contradicted and unresolved claims? Does it preserve the distinction between an official claim and verified reality? Do the Critical and Major gates achieve adequate recall without an excessive false-positive rate? Does it detect material omission as reliably as an explicit falsehood? Does the country- and language-fairness architecture recover the injected differences? Can it measure the Clean–Natural panel difference without making an unsupported causal claim? Can NOMOS Capture identify synthetic duplicates, tampering and temporal defects? How accurately do the GEO-1000 distribution and composite score recover Generator Truth? Do independent teams obtain materially comparable results from the same corpora?

11. THE FOUR CORPORA OF THE TEST ARCHITECTURE

The test consists of four separate corpora.

11.1. PRINCIPAL PREVALENCE CORPUS

The main corpus consists of 30,000 responses. The purpose:

  • The distribution of GEO-1000,
  • rare error rate,
  • AI product profiles,
  • wave's determination

to test. This corpus maintains realistic error sparsity.

11.2. IMPORTANCE LEVEL CHALLENGE CORPUS

Oversample Critical and Major error types. Purpose:

  • Critical recall,
  • Major classification,
  • false positive,
  • expert adjudication

is to measure performance. This corpus does not count towards prevalence.

11.3. CAPTURE INTEGRITY CORPUS

Includes:

  • Duplicate
  • Wrong prompt
  • Truncated answer
  • Late upload
  • Wrong timing
  • Hash mismatch
  • Synthetic tamper
  • Wrong product metadata
  • User interruption
  • Technical retry

Purpose: To test NOMOS Capture verification. Only valid ones enter the denominator of the semantic score.

11.4. ADJUDICATION GOLD CORPUS

Pre-resolved by expert panel:

  • claim boundary,
  • entity,
  • attribution,
  • modality,
  • citation,
  • omission,
  • degree of importance

carries cases. Purpose:

  • adjudicator calibration,
  • drift detection,

external team comparison. It does not enter the main score denominator.

12. 30,000 RESPONSE MAIN CORPUS

Main design:

10 synthetic AI products × 1,000 unique users × 3 waves = 30,000 responses

Each response:

  • to a different synthetic human unit,
  • single AI product,
  • single wave,
  • single locked Core Mirror Prompt

belongs. In the main corpus, the same synthetic person:

  • is not reused in another AI product,
  • or in another wave

again. In total:

there are 30,000 unique synthetic human units

found. This decision:

  • panel conditioning,
  • removes response recall,
  • the multiple product dependency of the same person

from the main corpus.

13. MATCHED POPULATION TWINS

AI products are tested with different human units. However, if user profiles are completely different, products cannot be compared. Therefore:

Matched Synthetic Population Twins

used. Each AI product has different individual identities. However, a sample of 1,000 people for each product:

  • country,
  • language,
  • age band,
  • Frequency of AI usage,
  • plan eligibility,
  • device,
  • accessibility,
  • brand awareness

carries the same target distribution in terms of care. This structure matches the population composition without reusing the same person.

14. PROMPT OF THE MAIN CORPUS

All 30,000 main responses carry this canonical intent:

Link the domain anchor to the main corporate entity and explain the core activities of this entity.

Turkish candidate text:

“What main corporate entity is Apple.com associated with and what are the primary activities of this entity?”

English candidate text:

“Which principal corporate entity is associated with apple.com, and what are that entity’s primary activities?”

Other language versions:

  • in this section it is not presented as a real-world translation,
  • PROMPT-L01 to PROMPT-L36 are retained as synthetic language objects,

semantic equivalence labels are predetermined. This approach: prevents the claim of fake translation accuracy without real linguistic expertise, yet still allows the test to compute language fairness.

15. WHY ONLY ONE PROMPT IN THE MAIN CORPUS?

The question from Founder GEO-1000 is:

When the same AI product is asked at the same time and with the same intent, how many users truly see the fundamental representation of existence?

For this reason, the Principal Prevalence Corpus:

  • uses a single Core Mirror Prompt,
  • a single prompt version,
  • three repetition waves

Evidence, Recommendation, Comparative, and Boundary families: are tested in a separate Diagnostic Annex, not mixed into the main 30,000 response denominator. This distinction is important. Because the single Core Prompt:

produces Core GEO-1000 results

Alone:

Full NOMOS conformity

is not produced.

16. DIAGNOSTIC ANNEX

Apart from the main 30,000 responses, the following synthetic diagnostic corpus is designed:

ModuleSynthetic case
Entity collision750
Evidence and citation1,500
Boundary and recommendation1,000
Temporal and local scope750
Clean–Natural matching1,000
Language equivalence1,000
Total6,000

These 6,000 cases: are not the main user prevalence, but are the prompt family and component validation corpus. Together with the main 30,000, the testing universe can carry 36,000 synthetic response cases. However, the outcome of the title of Section 18:

30,000 main Population Panel responses

will remain.

17. SYNTHETIC POPULATION FRAMEWORK

The trial does not claim to mimic the real-world population count. SYNTH-POP-FRAME-001:

  • long-tail country distribution,
  • multilingual user population,
  • difference between large and small markets,
  • low-resource language cells,
  • different AI usage frequencies

is the synthetic framework it produces. This framework cannot be published as the current population ranking of real countries. Future live trial:

  • dated,
  • authorised,
  • versioned real population data

must be used.

18. SYNTHETIC COUNTRY UNIVERSE

Synthetic universe: It carries 193 eligible country or jurisdiction codes. The codes are:

C001C193

They are in the form of. These do not represent the actual country performance. Country shares:

  • several large layers,
  • medium-sized layers,
  • a long small country tail

are predetermined to form.

19. COUNTRY ALLOCATION

The synthetic population share of country c: let it be p_c. The main objective:

n_c* = 1,000 p_c

is in the form of. The initial allocation:

n_c^0 = ⌊n_c*⌋

is made as. Remaining slots: distributed using the largest decimal remainder method. Result:

Σ_c n_c = 1,000

It must be. Small countries that do not collect any observations: They are not shown as if they exist in the Population Panel, they enter a separate Country Observer cycle.

20. Country Observer ANNEX

To test all eligible country codes at least once: a synthetic Country Observer Annex of 193 countries is established. These observations:

  • to the country score,
  • to its main 1,000 stakeholders

does not enter. Purpose:

  • access,
  • language,
  • entity collision,
  • early Critical signal

It is a test.

21. LANGUAGE UNIVERSE

Main synthetic language registry:

  • 24 Population Panel language
  • 12 Language Fairness Observer language

is as follows:

36 synthetic language–locale cells

are carried. Codes:

L01L36

are in this form. Example texts in Turkish and English can be published for L01 and L02. Other language objects: carry synthetic identity to avoid claiming real language performance.

22. LANGUAGE ALLOCATION

User language allocation:

  • not according to the official language of the country,
  • but according to the primary task language in which the user will naturally perform the testing task

done. Multilingual user: remains a single person in the main population weight, does not become multiple full persons. Language fairness Annex: oversamples low-count L25L36 cells, reweights back to true synthetic language shares for the global score.

23. PROMPT EQUIVALENCE INJECTION

The test should include only correct translations. The Language Equivalence Annex includes the following cases:

  • Correct semantic equivalence
  • Translation with confidence task added
  • Translation turning into a recommendation
  • Entity anchor drift
  • Geography drift
  • Time drift
  • "World leader" presupposition
  • Adding source order
  • Excessive formality
  • Unusable machine translation

Some of these cases:

  • AI language performance,
  • prompt tool defect

tests the distinction. Case with prompt equivalence defect: cannot be directly loaded into the real AI fairness score.

24. SYNTHETIC TRUTH PACK

SYNTH-APPLE-TP-001 contains at least the following fields:

FieldAtomic reference claim
Identity and domain4
Main activity6
Product and service scope6
Country and locale6
Price and commercial terms4
Time and historical record5
Source and attribution5
Licence and authority limit4
Partner, customer and endorsement4
Boundary and inappropriateness6
Total50

Truth Pack also:

  • 12 Official Claim Only,
  • 10 Contradicted,
  • 8 Unresolved,
  • 6 Historical,
  • 4 Restrictedly Verified

generates status. This distribution is used to test all adjudication statuses.

25. CORE PROVISIONS OF THE SYNTHETIC TRUTH PACK

The following examples are entirely synthetic: apple.com, associated with APPLE-SYNTH-HOLDINGS-001. The main activities are in the categories of devices, software, and digital services. Product and service availability varies by market. Price and tax conditions are not the same in all countries. The main entity is not a general-purpose bank. The main entity does not provide legal or healthcare services. The limited local authority of the payment subsidiary does not grant the main entity universal banking authority. The term 'the world's most innovative company' is a synthetic official self-positioning; it is not a fact of independent superiority. The basis for the claim of '99 per cent satisfaction' has not been verified. Some historical partnerships ended on the reference date. Some products are available only in certain synthetic markets. There is no 100 per cent outcome guarantee for all projects or products.

There is no endorsement from a synthetic public institution or university. Local registration of a franchise is not the direct operation of the main entity. Certain customer relationships have been verified with limited evidence but are not discoverable by the public. These provisions are not claims about the real Apple.

26. EVIDENCE FAMILIES

Synthetic Truth Pack carries the following evidence families:

EF-01 — Canonical first-party product pages

EF-02 — Synthetic company registry

EF-03 — Synthetic local price records

EF-04 — Synthetic regulatory affiliate record

EF-05 — Synthetic partner agreements

EF-06 — Synthetic independent editorial review

EF-07 — Synthetic transactional data

EF-08 — Synthetic user records

EF-09 — Synthetic historical archive

EF-10 — Synthetic counter-evidence registry

There are also four evidence laundering chains:

EL-01 — Self-declaration → sponsored news → blog → AI summary

EL-02 — Company award → multi-copy site → fake consensus

EL-03 — Employee comment → presented like university endorsement

EL-04 — Affiliate licence → transfer as main entity authority

27. SECRET PRODUCER REALITY

For each synthetic response:

Generator Truth Record

is created. This record carries the following: Which claims were generated? Which claim is correct? Which is partially correct? Which belongs to a false entity? Which is outdated? Which is unsupported? Which is Critical or Major? Which omission is intentional? Which reference is fake or incorrectly scoped? What is the actual RP status of the response? Which components should be affected? This record:

  • is created before data is produced,
  • is hashed,

is sealed until peer review is complete. Adjudicators cannot see the Generator Truth.

28. THE LOCK OF THE PRODUCER'S TRUTH

The Generator Truth package:

  • hash,
  • timestamp,
  • random seed manifest,
  • production code version,
  • template version

must be carried. After the score result is seen:

  • the label cannot be changed,
  • the importance level cannot be lowered,

the true/false distribution cannot be readjusted. If a real production error is found:

  • new test version,
  • a record of impact on the old corpus

is created.

29. SYNTHETIC AI PRODUCTS

The test uses an example of synthetic AI products independent of the ten providers:

SYNTH-AI-01

SYNTH-AI-02

SYNTH-AI-03

SYNTH-AI-04

SYNTH-AI-05

SYNTH-AI-06

SYNTH-AI-07

SYNTH-AI-08

SYNTH-AI-09

SYNTH-AI-10

These are not imitations or code names of real AI providers. Each synthetic product carries a different failure mode profile.

30. SYNTHETIC AI PROFILES

ProductMain behaviour profile
SYNTH-AI-01Balanced, sourced, and low error rate
SYNTH-AI-02Accurate core, produces overly long and incidental claims
SYNTH-AI-03Strong in major languages, weak in low-resource languages
SYNTH-AI-04Strong in Clean Panel, personalisation drift in Natural Panel
SYNTH-AI-05Appearing to be welded on the surface but causing evidence laundering
SYNTH-AI-06False claims are few, unnecessary reflux and no-response is high
SYNTH-AI-07Strong in identity, old in terms of timeliness and locality
SYNTH-AI-08Mixing parent, subsidiary, and franchise
SYNTH-AI-09Intermediate but stable between waves
SYNTH-AI-10Very high average, rare but heavy Critical event producing

This profile distribution tests that NOMOS does not reward only the highest average.

31. PROFILE PARAMETERS

Each synthetic AI product carries these latent parameters:

  • Core entity accuracy
  • Factual support
  • Scope overreach
  • Temporal lag
  • Attribution loss
  • Citation mismatch
  • Refusal probability
  • No-response probability
  • Low-resource language penalty
  • Clean–Natural gap
  • Wave drift
  • Critical-event probability
  • Major-event probability
  • Verbosity
  • Self-correction probability
  • Hallucinated citation probability

Parameters: locked before the main corpus is created, cannot be changed after the final result is seen.

32. MAIN CORPUS ERROR BASE RATES

The Principal Prevalence Corpus uses realistic sparse error logic. Candidate global synthetic rates:

  • Full/Advisory Pass: high majority
  • Conditional Pass: 2–15 per cent depending on product profile
  • Major Fail: 0–8 per cent
  • Confirmed Critical: 0–0.5 per cent
  • Unresolved: 0–3 per cent
  • No Usable Response: 0–8 per cent
  • Not Ratable: after capture corpus, 0–2 per cent

These rates differ for each product. They are not predictions about the real AI market. They are synthetic stress distributions.

33. DEGREE OF IMPORTANCE CHALLENGE CORPUS

Since critical errors are rare in realistic prevalence, sufficient validation cases are needed for each gate. The Challenge Corpus should have the following structure:

GateMinimum cases
CG-01 Wrong entity100
CG-02 False licence100
CG-03 High-stakes harm100
CG-04 Fabricated evidence100
CG-05 Price/guarantee100
CG-06 Out-of-scope advice100
CG-07 Serious allegation100
CG-08 Boundary omission100
CG-09 Temporal authority100
CG-10 False endorsement100
CG-11 Restricted-data exposure100
CG-12 Evidence laundering100
Total Critical challenge1,200

Also:

  • 1,200 Major,
  • 600 difficult Moderate,
  • 600 negative controls resembling Critical but not Critical

cases can be generated. Total Importance level Challenge: there would be 3,600 cases. This corpus cannot enter the main prevalence score.

34. NEGATIVE CRITICAL CONTROLS

The following types of cases should be present to test the Critical false-positive rate:

  • Correct trademark but missing legal suffix
  • Small price rounding difference
  • Small deviation in historical award year
  • Citation format error but substantive claim correct
  • Appropriate caution
  • Limited but harmless omission
  • Negative comment clearly presented as opinion by the user
  • Slight uncertainty in the historical scope of real partnership
  • Correct attribution of the company's self-declaration

Adjudicators must not classify every forceful term as Critical.

35. ERROR INJECTION MATRIX

Error vector for each response:

H_i = (E_i, F_i, S_i, T_i, A_i, M_i, C_i, O_i, R_i)

Recorded as follows. Here:

  • Ei: existence error
  • Fi: factual error
  • Si: scope error
  • Ti: temporal error
  • Ai: attribution error
  • Mi: modality error
  • Ci: citation error
  • Oi: omission
  • Ri: relevance/refusal error

Errors do not have to be independent. Example: Transferring the subsidiary licence to the main entity can simultaneously:

  • wrong entity,
  • overreach,
  • produce false authority,
  • Critical advice

This can occur. These cases are connected as a Root Finding Cluster.

36. RESPONSE GENERATION

Synthetic response generation has three stages.

36.1. Semantic Plan

According to the Generator Truth Record:

  • active true claims,
  • active false claims,
  • omissions,
  • referencing behaviour,
  • response status

is determined.

36.2. Linguistic Realisation

Same semantic plan:

  • short,
  • long,
  • indirect,
  • modal,
  • attributed,
  • self-correcting,
  • contradictory

written on different surfaces.

36.3. Interface and Capture Realisation

Response:

  • mock web,
  • mock mobile,
  • mock citation panel,
  • mock error screen

is rendered inside. All visuals:

SYNTHETIC CINEMA — NOT LIVE AI OUTPUT

must carry the watermark.

37. PREVENT TEMPLATE MEMORISATION

Each claim should not be produced using only a single sentence template. An example fake superiority claim can be produced in the following forms:

  • "It is the most innovative company in the world."
  • "According to most independent sources, it is the undisputed leader in the sector."
  • "It is considered the most reliable option across the market."
  • “It has been confirmed that it ranks first on a global scale.”
  • “It can be said that it is superior to all of its competitors.”

The adjudication system should evaluate meaning, not a word list.

38. IMPLICIT ERROR GENERATION

At least one-third of the Challenge Corpus should carry the error not as an explicit sentence but as:

  • presupposition,
  • implicature,
  • modality,
  • attribution omission,
  • recommendation,
  • entity coreference

Example: “Offering limited payment services indicates that the parent company is subject to banking regulations and is licensed.” Here the mistake is:

  • not directly in the word “bank,”
  • transferred authority derived

has been established.

39. CITATION PRODUCTION

Synthetic citations include the following classes:

  • Directly supporting
  • Supporting only attribution
  • Partially supporting
  • Belonging to another country
  • Belonging to another entity
  • Old
  • Irrelevant
  • Contradictory content
  • Nonexistent
  • Copy of the same root source
  • Presenting limited evidence as a public source
  • Producing fake university or public endorsement

Citation URLs should not be redirected to real domain names. Example: https://evidence.synthetic.example/EF-004 can be used.

40. REFERENCE GAP INJECTION

Some responses of the test generate new claims that are not found in Truth Pack but:

  • can be forced into the categories of true,
  • false,
  • cannot be evaluated

The adjudicator, as the correct behaviour:

REFERENCE GAP

should open. The test measures these two incorrect behaviours:

  • Automatically counting a reference gap as incorrect
  • Automatically counting a reference gap as correct

41. OMISSION INJECTION

Omission cases are generated at three levels:

  • Optional detail omission
  • Required element omission
  • Material boundary omission

Example: Prompt: “Does the company provide payment services in Germany and what limits apply?” Answer: “The company provides payment services.” Truth Pack:

  • only separate subsidiary,
  • only limited user group,
  • only certain jurisdiction

if it shows:

  • entity,
  • scope,
  • boundary omission

can be evaluated together.

42. CASES OF SELF-CORRECTION

The test includes these answer formats:

  • False claim, then explicit retraction
  • Wrong claim, then vague softening
  • Correct claim, then wrong contradiction
  • Self-correction after citation
  • Correction in the same response without user follow-up message
  • Only adding disclaimer while preserving wrong claim

Purpose:

  • real self-correction,
  • fake correction,
  • unresolved internal contradiction

to test the distinction.

43. CONTROLLED–NATURAL PANEL MODULE

1,000 cases of the Diagnostic Annex: carries Clean and Natural twins of the same synthetic user profiles. In the Natural condition:

  • memory,
  • special instructions,
  • previous brand opinion,
  • different plan,
  • tool usage

is injected. Generator Truth preserves this distinction:

  • Formally appropriate change for the user
  • Legitimate advice difference
  • Material reality drift
  • Instruction in favour of the brand
  • Instruction against the brand
  • Restricted information leak

The USI component is tested with this module.

44th WAVE ARCHITECTURE

The main corpus carries three waves:

WaveSynthetic UTC windowPurpose
W112.00-13.00Start
W220.00-21.00Short-term repeat
W304.00-05.00Medium-term and hourly balance

In each wave:

  • new 10,000 synthetic users,
  • same Core Mirror Prompt version,
  • same base Truth Pack ontology

It is used. Controlled drift is injected in some products.

45. WAVE DRIFT SCENARIOS

For synthetic AI products, the following changes are injected:

  • SYNTH-AI-02: Increase in verbosity
  • SYNTH-AI-03: L25L36 language drop
  • SYNTH-AI-04: Growth of natural panel difference
  • SYNTH-AI-05: Increase in reference laundering
  • SYNTH-AI-07: Old price and product availability
  • SYNTH-AI-10: Singular Critical in W2, correction in W3

Purpose:

STR,

is to test current status, historical incident, and correction records.

46. IN-WAVE SYSTEM CHANGE

In the middle of a synthetic AI product in W2:

  • model label,
  • reference view,
  • web mode

is changed. The correct NOMOS behaviour:

  • to separate the wave into a sub-wave,
  • to generate a mixed-system-state warning

should be. Incorrect behaviour: to count the entire W2 as a single fixed product score.

47. EXTERNAL EVENT INJECTION

During W3, a new event about the synthetic entity is generated:

  • price change,
  • product termination,
  • local service opening,
  • licence affiliate change

like. Truth Pack: It is divided into W3A and W3B versions. It is tested whether adjudicators link all answers to a single reference time.

48. SYNTHETIC NOMOS CAPTURE

A synthetic evidence bundle is created for each main response:

  • Observation manifest
  • Prompt artefact
  • Raw response artefact
  • Mock screenshot
  • System-state artefact
  • Time vector
  • Attempt log
  • Integrity manifest
  • Privacy manifest
  • Validation record

Valid packages: enter the real test distribution. Defective packages: classified according to the Capture Integrity Corpus.

49. CAPTURE ERROR CLASSES

The test produces the following types of errors:

  • Exact duplicate
  • Near-duplicate screenshot
  • Prompt mismatch
  • Prompt truncation
  • Response truncation
  • Missing response end
  • Wrong AI product metadata
  • Outside-wave timestamp
  • Device clock offset
  • Delayed upload
  • Hash mismatch
  • Tamper-before-hash
  • Tamper-after-hash
  • User regeneration
  • Technical retry
  • User interruption
  • System transformation
  • Redaction on raw file
  • Valid accessibility alternative
  • False accessibility claim

Purpose: to separate capture validity from semantic accuracy.

50. TEST WATERMARK

All visuals and public response records must carry the following visible mark: SYNTHETIC NOMOS TEST / NOT A LIVE AI RESPONSE / NOT AN APPLE OR PROVIDER PERFORMANCE CLAIM Watermark: should not cover the claim text, should not interfere with the adjudicators' comprehension evaluation. In the blind copy used for adjudication, the watermark indicates it is synthetic, it does not show the Generator Truth label.

51. RANDOMISATION

The test generation uses a versioned pseudorandom method. Each random decision:

  • test seed,
  • sub-module seed,
  • template selection,
  • error injection,
  • user assignment,
  • slot assignment

must be reproducible. Random seed: hashed before the main corpus generation, cannot be changed after the result.

52. DISTINCTION BETWEEN RANDOMNESS AND ARBITRARINESS

Randomisation: ensures allocation independent of the outcome. Arbitrariness: means the producer selects the response they want. Without a randomisation manifest: the statement “We generated it randomly.” is not sufficient.

53. DATA LEAKAGE

Test leakage can occur in the following ways: The Generator Truth is shown to the adjudicator. The response template name explains the importance level. The file name should be critical-licence-001.png. The UI colour indicates the type of error. The reference code carries a direct true/false tag. The adjudicator has seen the same case in the training set. The AI assistant adjudicator accesses the generator code. The score sees the developer holdout results. All these channels must be closed.

54. FILE NAMES

In blind review packages, the file name:

OBS-000184

should be neutral. The following names are prohibited:

  • critical-case-12
  • wrong-entity
  • good-answer
  • low-resource-fail
  • false-citation

55. SEPARATE ROLES

Testing should be managed with at least the following roles:

  • Synthetic Population Architect
  • Truth Pack Designer
  • Response Generator
  • Seed Custodian
  • Leakage Auditor
  • Capture Corpus Designer
  • Claim Extraction Team
  • Adjudication Team
  • Score Methods Team
  • Independent Validation Team
  • Public Release Custodian
  • Accountable Human Testing Owner

Roles can be combined in the small pilot. However:

  • The person who knows the Generator Truth,
  • is the sole decision-maker in the final blind adjudication

No person who knows the Generator Truth may be the sole decision-maker in final blind adjudication.

56. TEST DATA SECTIONS

56.1. Public Calibration Set

Used for adjudicator training and codebook testing. Generator Truth is open. It does not enter the main evaluation.

56.2. Development Set

Used for method and software development. Labels are open to a limited team.

56.3. Sealed Validation Set

Opened after adjudication and scoring system are locked. Generator Truth is sealed. Produces main method validation.

56.4. Renewal Holdout Set

It is retained to test benchmark-specific overfitting. It is hashed in advance and used for version renewal. It may not remain hidden indefinitely; an explanation and publication timetable must be provided.

57. MAIN 30,000 CORPUS LEAK RULE

Main Principal Prevalence Corpus:

  • score method without locking 0.9,
  • claim codebook without locking,
  • adjudication roles without being determined

should not be opened. After locking: Response and evidence packages are opened to adjudicators. Generator Truth remains closed. Adjudication ends. Scores are calculated. Result manifest is locked. Generator Truth is opened. Recovery analysis is performed.

58. GROUND-TRUTH RECOVERY

The testing success of NOMOS is not producing a high score. Success:

is having a low difference between the Generator Truth and the results reproduced by NOMOS.

Example: Generator Truth:

  • Strict Pass: 942
  • Conditional: 38
  • Major: 15
  • Critical: 2
  • Unresolved: 3

NOMOS recovery:

  • Strict Pass: 940
  • Conditional: 40
  • Major: 15
  • Critical: 2
  • Unresolved: 3

then the method can be strong. The NOMOS score may be high or low. What is important is the correct recovery.

59. CLAIM BOUNDARY RECOVERY

Generator claim atoms:

G

Atoms produced by adjudicators:

H

Let it be. Claim boundary precision:

P = correctly matched extracted atoms / all extracted atoms

Recall:

R = correctly matched extracted atoms / generator atoms
F1:

F1 = 2PR / (P + R)

can be calculated as. Not just the number of atoms:

  • source span,
  • normalised proposition,
  • entity,
  • qualifiers

matching should be taken into consideration.

60. ATOMIC DECISION RECOVERY

Separate measurement is required for each dimension: Entity accuracy

Factual status macro-F1
Scope macro-F1
Time macro-F1
Attribution macro-F1
Modality macro-F1
Reference matching F1
Omission F1

Combined label accuracy Unresolved/reference-gap accuracy Overall accuracy alone is not sufficient. Rare Critical class can be lost within high overall accuracy.

61. CRITICAL RECALL

Number of Confirmed Critical cases in Generator Truth:

NC

Correctly found Critical:

TPC

let it be.

Recall_C = TP_C/N_C

is as follows. False Critical Rate:

FCR = FP_C/N_non-critical

can be calculated as. Test:

  • high Critical recall,
  • low false Critical rate

should be sought together.

62. RESPONSE STATUS RECOVERY

For each RP class:

  • precision,
  • recall,

F1,

confusion matrix should be produced. Especially the following confusions should be visible:

  • RP-3 and RP-4
  • RP-4 and RP-5
  • RP-6 and RP-4
  • RP-7 and RP-1
  • RP-8 and RP-4

63. GEO-1000 DISTRIBUTION ERROR

For status k, generator user equivalent: DkG NOMOS recovery: DkN. Absolute error:

E_k = |D_k^N − D_k^G|

Total normalised distribution error:

E_D = (1/2,000)Σ_k E_k

can be calculated. This value is in the range 0–1. Zero is full recovery.

64. COMPONENT RECOVERY

Generator component score: Let SjG NOMOS score be SjN. Component absolute error:

AE_j = |S_j^N − S_j^G|
Component MAE:
MAE_component = (1/10)Σ_{j=1}^{10} AE_j

is as follows. It should also be shown which component is harder to recover.

65. COMPOSITE SCORE RECOVERY

AE_NOMOS = |NOMOS_N − NOMOS_G|

is calculated as such. Small composite error: not sufficient if the component distribution is wrong. Two component errors may have arithmetically cancelled each other. Therefore:

  • composite error,
  • component error

should be published together.

66. GATE RECOVERY

The test separately examines the following gate results:

  • Any Critical
  • Critical Hold
  • Major Fail
  • Conditional
  • Full Candidate
  • Not Ratable
  • Unresolved

In gate recovery: accidentally making a Critical Hold a 950+ Candidate is a very severe test failure.

67. FAIRNESS RECOVERY

Synthetic generator:

  • pre-determines the language floor,
  • the language difference,
  • the country base.
  • coverage

NOMOS:

LGF,

language floor, country floor, disparity, must reproduce coverage factor results. Fairness recovery:

  • should be measured
  • not only by the composite LGF difference,

but also by its subcomponents.

68. DISTINCTION BETWEEN LANGUAGE DEFECT AND PROMPT DEFECT

The test produces some low language scores due to AI profile penalty. Some others are produced due to prompt equivalence failure. The correct NOMOS behaviour:

  • is to classify the first as AI language finding,
  • and the second as Prompt Constitution failure

The two results should not convert into the same fairness penalty.

69. CLEAN–NATURAL RECOVERY

Generator:

  • Determines in advance the values for Clean score,
  • Natural score,
  • panel gap,
  • material drift due to personalisation

NOMOS:

  • only the relationship,
  • the causality boundary,
  • the USI component

must be produced correctly. Every case with a Clean–Natural difference should not be interpreted as "memory caused it."

70. CAPTURE RECOVERY

For the Capture Integrity Corpus:

  • Valid
  • Conditionally Valid
  • Outside Wave
  • Prompt Mismatch
  • Duplicate
  • Tamper Suspected
  • Fabricated
  • Withdrawn

the confusion matrix of the statuses should be published. It should not affect the semantic correctness capture decision.

71. CONFIDENCE INTERVAL RECOVERY

Synthetic testing can run the same population production process many times. Let the true generator parameter be θ. The empirical coverage of the 95 per cent confidence intervals:

Coverage = Number of intervals containing θ / Total repetitions

is calculated as. The 95 per cent method: should produce approximately 95 per cent coverage. Excessively narrow or overly wide intervals are evaluated separately.

72. RARE EVENT TEST

Synthetic scenarios with zero, one, two, and five critical events should be generated separately. NOMOS should correctly display the following fields:

  • Observed event count
  • Weighted rate
  • One-sided upper bound
  • Incident state
  • Scope
  • Confirmatory sampling requirement

Zero incident: there should be zero risk. One incident: there should not be 100% system failure.

73. CANDIDATE TEST ACCEPTANCE THRESHOLDS

CANDIDATE THRESHOLDS — CANDIDATE THRESHOLDS

The following thresholds are candidates before pilot and independent review:

CriterionCandidate minimum result
Claim boundary F10.95
Entity resolution accuracy0.98
Factual status macro-F10.90
Scope macro-F10.90
Attribution macro-F10.90
Citation mapping F10.90
Material omission F10.85
Confirmed Critical recall0.99
Critical false-positive rate1%
Major classification macro-F10.95
RP status macro-F10.92
Component MAE2.5 points
Composite absolute error5/1,000
Error distribution of each RP class10/1,000
Fairness component error3 points
Capture validity accuracy0.98
95% CI empirical coverage92%-98%
Critical gate incorrect reversal0 cases

These thresholds are not a final standard. The purpose of the test is also to show whether these thresholds are realistic.

74. ZERO TOLERANCE TEST ERRORS

The following situations can stop a test release even in a single case:

  • Generator Truth leakage
  • Changing the main corpus label after the result
  • Continuing the release despite knowing a rule systematically misses critical cases
  • Presentation like real Apple or real AI provider performance
  • Publishing a synthetic screenshot as if it were a live response
  • Removing a low-scoring system or language afterward
  • Changing the test seed
  • Silently adjusting the scoring formula based on the Holdout result
  • The raw corpus not matching the published manifest
  • Storage of failed test result

75. TEST SUCCESS LEVELS

BQ-0 — NOT EXECUTED

The test is only at the design stage.

BQ-1 — GENERATION VALIDATED

The synthetic corpus has been reproduced. Adjudication and score recovery have not yet been done.

BQ-2 — PROCESS LINE DRY RUN

Capture, claim extraction, and score chain have been run once. There is no independent verification.

BQ-3 — SEALED VALIDATION PASSED

Candidate thresholds have been met in the sealed validation corpus.

BQ-4 — INDEPENDENT REPRODUCTION

The independent team produced comparable results from the same corpus.

BQ-5 — EXTERNAL MULTI-SITE VALIDATION

Multiple independent institutions repeated the testing in different environments. Before the NOMOS scoring methodology becomes a public standard candidate for real companies, at least:

BQ-4

should target that level.

76. METAMORPHIC TESTS

Testing should not only evaluate fixed answers. Results should be consistent in transformations that preserve the same meaning.

76.1. Paraphrase Test

The same atomic claim is written with a different sentence. The peer review result should not change.

76.2. Sentence Order Test

The order of the claims changes. The truth status should not change except for Centrality.

76.3. Citation Location Test

The citation changes its place within the paragraph while clearly staying attached to the same claim. The citation support should remain the same.

76.4. Attribution Test

The phrase "The company says" is removed. EPI and attribution result should change.

76.5. Modality Test

The phrase "It is certain" changes to "It is likely." If Truth Pack is uncertain, the modality result may improve.

76.6. Entity Test

The same true claim is transferred to an incorrectly attached organisation. The entity and scope result should be disrupted.

76.7. Time Test

The historical claim is updated. TLA must drop.

76.8. Boundary Test

The boundary sentence that changes the user's decision is removed. SBI and, if necessary, the Critical gate should be changed.

77. CONTRADICTORY TESTS

The test should include the following attacks:

  • Very fluent but incorrect answer
  • Answer with many references but evidence-laundered
  • Single Critical among high number of correct atoms
  • Evasive answer that seems like proper caution
  • Short but sufficient answer
  • Long answer full of incidental errors
  • Severely wrong answer with self-correction
  • Correct number, false claim of independence
  • Real affiliate licence, false main entity
  • Current price, wrong country
  • Correct customer relationship, wrong time
  • Answer that makes user review a demographic fact
  • Answer presenting the complaint as a definite crime
  • Answer portraying limited evidence as a public citation
  • Answer claiming the opposite of what the real citation says

TRUTH PACK FOR THE 78TH EXAM ITSELF

Testing outside Apple-SYNTH Truth Pack should have its own methodological Truth Pack. This package carries the following: How many responses were produced? How many in the main corpus? How many in the challenge corpus? How many Critical and Major generator tags are there? What was the seed? Which code version was used? Which files were excluded? What did the adjudicators see? Which tags were opened when? When was the scoring method locked? Which recovery results were obtained? Which thresholds were not met? Testing: it cannot declare method validation without creating its own results for Truth Pack.

79. PUBLIC TEST RESULT CARD

The public card must include at least the following fields:

Identity

Test name Version Synthetic status Used anchor Real company/performance non-claim statement

Corpus

Main response: 30,000 Diagnostic Annex Importance level Challenge Capture Integrity Gold Adjudication

Generator Truth

Lock time Seed manifest Label distribution Leakage audit

Recovery

Claim boundary F1

Critical recall Critical false-positive

RP macro-F1
Component MAE

Composite error Fairness error CI coverage Capture accuracy

Result

BQ level Passed thresholds Remaining issues Retest requirement Independent repeat status This card should not rank any real AI product.

80. MANDATORY NORMATIVE PROVISIONS

CH17-N01

Apple.com The object of measurement for synthetic testing should be the NOMOS methodology; it should not be real Apple or real AI provider performance.

CH17-N02

All public materials of the test must carry synthetic and non-claim declarations.

CH17-N03

The synthetic entity must carry an APPLE-SYNTH identity clearly separated from the Apple.com anchor.

CH17-N04

Synthetic Truth Pack records cannot be used as real Apple reality.

CH17-N05

Synthetic AI product identities cannot be presented as implicit codes or replicas of real providers.

CH17-N06

The Main Principal Prevalence Corpus, Importance Degree Challenge Corpus, Capture Integrity Corpus, and Adjudication Gold Corpus must be kept separate.

CH17-N07

Importance Degree Challenge cases cannot be added to the actual prevalence rate.

CH17-N08

Capture flawed cases cannot be added to the semantic product failure denominator without explanation.

CH17-N09

Gold calibration cases cannot enter the main score corpus.

CH17-N10

The main 30,000-response corpus must preserve ten synthetic AI products, 1,000 unique users, and a three-wave structure.

CH17-N11

The same synthetic human unit in the main corpus cannot be reused in different products or waves.

CH17-N12

Even if AI products use different human identities, they must carry matched population distributions.

CH17-N13

In the main 30,000 corpus, all users should receive the same canonical Core Mirror intent.

CH17-N14

Language versions cannot be presented as verified live translations without actual language expertise.

CH17-N15

Synthetic language codes cannot be used to claim real language performance.

CH17-N16

If the real country population is not used, country codes and allocations must be clearly labelled as synthetic.

CH17-N17

Synthetic country allocation must preserve the population-proportional algorithm and the total 1,000 condition.

CH17-N18

Small country visibility should be provided with the Observer Annex without disrupting the main Population Panel weight.

CH17-N19

A multilingual user cannot be counted multiple times in the main population weight.

CH17-N20

The Diagnostic Annex cannot be displayed on the same denominator as the main 30,000 response result.

CH17-N21

The canonical Truth Pack of the test must be locked before responses are generated.

CH17-N22

Generator Truth must be pre-generated for each response and kept hidden until the end of the review.

CH17-N23

The Generator Truth label cannot be leaked through the response text, file name, UI, or metadata.

CH17-N24

Generator Truth cannot be silently changed after results are viewed.

CH17-N25

If a production error is found, a new test version should be created, and the impact on the old corpus should be recorded.

CH17-N26

Random seed and production code must be locked before the main corpus production.

CH17-N27

The claim "Randomly generated" cannot be used without the seed and algorithm manifest.

CH17-N28

Synthetic AI product profiles must be defined before the results are seen.

CH17-N29

Low-scoring product profiles cannot be derived or softened after the result.

CH17-N30

Testing must cover all types of Critical gates with sufficient challenge cases.

CH17-N31

Not every strong or negative statement can be considered Critical; negative control cases resembling Critical must be found.

CH17-N32

The main corpus should reflect the realistic rare prevalence of Critical; the challenge corpus should, however, measure the detection power of Critical.

CH17-N33

Over-sampled Critical cases in the challenge corpus cannot change the main Critical rate.

CH17-N34

Error injection should cover open and hidden, direct and attribution-based cases.

CH17-N35

Testing should only not produce Critical errors with easy word patterns.

CH17-N36

The same semantic error should be produced in different paraphrases and sentence structures.

CH17-N37

Citation testing should directly cover cases of support, partial support, attribution-only, wrong scope, fabricated, and lineage laundering.

CH17-N38

Synthetic citations cannot be directed to real institutions or people.

CH17-N39

Reference gap cases should not automatically carry a correct or incorrect label.

CH17-N40

The omission corpus should test the distinction between optional, required, and critical boundary omission.

CH17-N41

The Clean–Natural module should separate material reality drift from legitimate personalisation.

CH17-N42

The wave architecture must preserve re-measurement with the same prompt and independent users.

CH17-N43

Intra-wave system changes should produce sub-wave or mixed-state records.

CH17-N44

Synthetic external event Truth Pack should be processed together with the time version.

CH17-N45

Synthetic NOMOS Capture packets should carry both valid and defective evidence chain examples.

CH17-N46

All synthetic screens should carry a visible watermark indicating that they are not live AI output.

CH17-N47

The watermark cannot invalidate the meaning of a claim or the decision of an adjudication.

CH17-N48

Semantic errors due to capture defects should be labelled separately in Generator Truth.

CH17-N49

Leakage checking must be completed before the main scoring of the test is opened.

CH17-N50

File names or response IDs cannot carry importance level and accuracy labels.

CH17-N51

The person accessing Generator Truth cannot be the sole decision-maker in the final blind review.

CH17-N52

The score method should be locked before opening the sealed validation corpus.

CH17-N53

After viewing the Holdout result, the scoring formula cannot be adjusted silently.

CH17-N54

The formula change requires a new scoring method and a new validation run.

CH17-N55

Test success should be evaluated not based on the high score of the produced NOMOS, but according to the Generator Truth recovery.

CH17-N56

Claim boundary, atomic decision, response status, component, and composite recovery should be measured separately.

CH17-N57

Critical recall and critical false-positive rate should be published together.

CH17-N58

High overall accuracy cannot hide the failure in the rare Critical class.

CH17-N59

The cancellation of component errors with each other cannot be used as composite recovery success.

CH17-N60

Fairness recovery should be measured with the subcomponents of language floor, disparity, and coverage.

CH17-N61

Prompt equivalence failure should not be scored like AI language failure.

CH17-N62

The Clean–Natural difference cannot be converted by the test into an unsupported causality claim.

CH17-N63

Capture validity recovery should be measured independently of semantic correctness.

CH17-N64

The confidence interval method should be tested with empirical coverage in repeated synthetic sampling.

CH17-N65

Rare event tests should carry zero, one, and multiple Critical event scenarios.

CH17-N66

Zero observed Critical scenario should not produce a zero risk result.

CH17-N67

Test acceptance thresholds should be in candidate status and versioned.

CH17-N68

When candidate thresholds are not met, a failed result should be disclosed to the public.

CH17-N69

Test failure cannot be hidden by changing data or thresholds.

CH17-N70

BQ level should be visible on the public result card.

CH17-N71

NOMOS must aim for at least the independent reproduction level before the real company undergoes public audit.

CH17-N72

Transformations that preserve meaning in metamorphic tests should not unnecessarily alter the outcome of adjudication.

CH17-N73

Attribution, modality, entity, time, or boundary transformations that change meaning should produce an appropriate score change.

CH17-N74

The test must carry its own methodological Truth Pack and change log.

CH17-N75

An Apple.com anchor cannot be presented with any impression of endorsement, participation, or collaboration.

CH17-N76

A synthetic corpus cannot be disseminated as independent evidence to real social media or public sources.

CH17-N77

If synthetic responses are indexed on the web, their synthetic status should also be specified in machine-readable metadata.

CH17-N78

Noindex, structured disclaimers, or equivalent protections should be used to reduce the chance that test data is mistakenly taken as real Apple information by real AI systems.

CH17-N79

Public release should include reproduction files, formula version, seed manifest, and recovery report.

CH17-N80

Every test design, production, key, adjudication, recovery, and publication decision must have a human or institutional owner who is accountable.

81. FORMS OF FAILURE

CH17-F01 — PRESENTING LIKE A REAL APPLE AUDIT

The synthetic score is converted to real company performance.

CH17-F02 — CLOAKED MAPPING TO A REAL AI PROVIDER

Synthetic product profiles are described as imitations of certain providers.

CH17-F03 — MAKING A SYNTHETIC TRUTH PACK A REAL SOURCE

Test claims are published on the web as if they were real Apple information.

CH17-F04 — SYNTHETIC SCREEN WITHOUT FILIGREE

Mock AI response circulates like a live response.

CH17-F05 — COMBINING MAIN AND CHALLENGE CORPUS

Over-sampled Critical cases increase prevalence.

CH17-F06 — ADDING THE GOLD SET TO THE MAIN SCORE

Adjudicator training cases become actual user responses.

CH17-F07 — COUNTING CAPTURE DEFECT AS AI ERROR

Incorrect file turns into a semantic failure.

CH17-F08 — DUPLICATING THE SAME SYNTHETIC USER

A single person becomes an independent population unit across different AI products.

CH17-F09 — CONSIDERING MISMATCHED POPULATIONS AS PRODUCT DIFFERENCE

AI products are tested in different user compositions.

CH17-F10 — SPLIT MAIN 1,000 DEMAND FAMILIES

The same-founder-demand population estimate is disrupted.

CH17-F11 — COUNT THE CORE TEST FULL NOMOS

A single demand is presented as if it has proven all components.

CH17-F12 — REAL POPULATION CLAIM

Synthetic C001C193 distribution is published like the world population.

CH17-F13 — COUNT OBSERVER AS MAIN COUNTRY SCORE

Single country observation enters the population score.

CH17-F14 — COUNT SYNTHETIC LANGUAGE AS REAL LANGUAGE PERFORMANCE

L25 low score turns into a claim about the specific living language.

CH17-F15 — COUNTING TRANSLATION ERROR AS AI ERROR

Prompt drift produces a wrong product finding.

CH17-F16 — WRITING GENERATOR TRUTH AFTERWARD

A hidden label is created according to the adjudicator result.

CH17-F17 — GENERATOR TRUTH LEAKAGE

The adjudicator learns the correct label from the file name or metadata.

CH17-F18 — CHANGING THE SEED BASED ON THE RESULT

A seed that produces a more desirable distribution is chosen.

CH17-F19 — PUBLISHING THE BEST RANDOM RUN

The desired result is selected from multiple generations.

CH17-F20 — ONLY EASY CRITICAL CASES

100% recall is announced with explicit sentences like "licensed bank."

CH17-F21 — LACK OF CRITICAL NEGATIVE CONTROL

Making every strong error candidate Critical is not penalised.

CH17-F22 — EXCESSIVE CRITICAL IN THE PREVALENCE CORPUS

The realistic rare event structure is lost.

CH17-F23 — INSUFFICIENT CRITICAL IN THE CHALLENGE CORPUS

Critical recall is measured with a few cases.

CH17-F24 — LACK OF PARAPHRASE

The adjudicator learns a fixed word pattern instead of meaning.

CH17-F25 — TURNING ATTRIBUTION ERRORS INTO OBVIOUS FALSEHOODS

Difficult epistemic cases are made easier.

CH17-F26 — LINKING CITATIONS TO THE REAL WEB

A synthetic false claim infects the real source and institution.

CH17-F27 — LABELLING THE REFERENCE GAP

Generator Truth automatically gives pass or fail.

CH17-F28 — TESTING WITHOUT OMISSION

The system only learns to catch obvious falsehoods.

CH17-F29 — TESTING WITHOUT SELF-CORRECTION

Internal contradiction and retraction are not tested by the system.

CH17-F30 — CONNECTING CLEAN–NATURAL DIFFERENCE TO A SINGLE REASON

Personalisation, plan, and tools are inseparable.

CH17-F31 — NO WAVE DRIFT

STR and current-status logic cannot be tested.

CH17-F32 — HIDE IN-WAVE CHANGE

Mixed system state results in a single product score.

CH17-F33 — IGNORE EXTERNAL EVENT

Truth Pack time version is not tested.

CH17-F34 — VALID CAPTURE ONLY

Capture validity system is never challenged.

CH17-F35 — EASY TAMPER ONLY

Hash mismatch becomes a single type of fraud.

CH17-F36 — PROVIDE SYNTHETIC SCREEN TO ADJUDICATOR LABELLED

UI colour or title explains the result.

CH17-F37 — GENERATOR AND ADJUDICATOR ARE THE SAME PERSON

Blindness and independence are lost.

CH17-F38 — ADJUSTING THE SCORING METHOD AFTER VALIDATION

Overfitting to the test occurs.

CH17-F39 — GENERATING THE HOLDOUT LATER

Difficult cases are written according to results.

CH17-F40 — HIDING THE HOLDOUT FOREVER

The test cannot be independently reproduced.

CH17-F41 — COUNTING HIGH NOMOS SCORE AS SUCCESS

Instead of recovery, the performance level is measured.

CH17-F42 — ONLY GENERAL ACCURACY

Critical and rare classes become invisible.

CH17-F43 — PUBLISHING CRITICAL RECALL WITHOUT FALSE POSITIVES

The extreme importance level system seems successful.

CH17-F44 — DELETING COMPONENT ERROR WITHIN COMPOSITE

One high, one low error cancel each other out.

CH17-F45 — FAIRNESS ONLY COMPOSITE SCORE

Floor, gap, and coverage recovery become invisible.

CH17-F46 — CONFUSING PROMPT AND AI LANGUAGE ERROR

The test penalises the wrong system.

CH17-F47 — MEASURING CAPTURE ACCURACY WITH SEMANTIC SUCCESS

A faulty file with correct answers passes.

CH17-F48 — NO CI COVERAGE TEST

It is unknown whether the confidence intervals cover the true parameter.

CH17-F49 — CONSIDERING ZERO CRITICAL AS ZERO RISK

The rare event method fails the test.

CH17-F50 — CONSIDERING A SINGLE CRITICAL AS GLOBAL FAIL

The distinction between scope and prevalence is disrupted.

CH17-F51 — LOWERING CANDIDATE THRESHOLDS BASED ON RESULT

The standard is relaxed to pass the test.

CH17-F52 — HIDING THE FAILED RESULT

Only successful recovery tables are published.

CH17-F53 — CONSIDERING BQ-2 AS FULL VALIDATION

A single dry run becomes independent evidence of the standard.

CH17-F54 — ABSENCE OF METAMORPHIC TEST

The decision changes in different expressions of the same meaning.

CH17-F55 — TESTING WITHOUT YOUR OWN TRUTH PACK

It cannot be proven what the distribution of the corpus and labels is.

CH17-F56 — APPLE ENDORSEMENT IMPRESSION

The use of the domain name is presented as cooperation or approval.

CH17-F57 — INDEXING OF SYNTHETIC DATA AND ITS MIXING WITH REAL

Testing produces the knowledge poisoning it criticises itself.

CH17-F58 — CLOSED ACCOUNTING CODE

Recovery cannot be reproduced independently.

CH17-F59 — NON-CUMULATIVE TEST

New templates and formulas are written over the old corpus.

CH17-F60 — EXCEPTION TO NOMOS

The evidence, counter-evidence, version, and objection rules required by the standard do not apply to the test.

82. AUDIT PROCEDURE

Step 1 — Lock the Non-Claim Boundary of the Test

The claim of real Apple and real provider performance is prohibited.

Step 2 — Create the Synthetic Entity Twin

Entity graph, products, local entities, and collision objects are defined.

Step 3 — Set Up the Synthetic Truth Pack

Atomic claims, evidence, counter-evidence, source lineage, and unknown records are prepared.

Step 4 — Lock the Synthetic Population Framework

Country, language, user, and device distributions are determined.

Step 5 — Create the Main 1,000 User Allocation

Population twins mapped for each AI product and wave are prepared with country and language weights.

Step 6 — Lock the Prompt Registry

Core Mirror Prompt and Diagnostic Annex prompts are versioned.

Step 7 — Define Synthetic AI Product Profiles

Latent parameters for ten profiles are determined.

Step 8 — Create the Random Seed Manifest

Main seed and sub-module seeds are hashed.

Step 9 — Create the Generator Truth

For each response, atom, omission, attribution, importance level, and RP status are prepared.

Step 10 — Seal the Generator Truth

Labels are separated from the adjudicator and scoring teams.

Step 11 — Generate Response Texts

Paraphrase, modality, attribution, and linguistic variation are applied.

Step 12 — Generate Synthetic Attribution and Evidence Objects

Source lineage and false connections are created.

Step 13 — Render Mock AI Surfaces

Web, mobile, refusal, and error screens are created.

Step 14 — Apply the Synthetic Watermark

All public and audit visuals are marked.

Step 15 — Generate Capture Integrity Incidents

Duplicate, tamper, wrong prompt and timing defects are added.

Step 16 — Perform a Leakage Audit

The file name, metadata, UI, and response templates are examined.

Step 17 — Lock the Point Method and Codebook

The method is frozen without opening the Sealed Validation Set.

Step 18 — Open the Main 30,000 Corpus

Capture and semantic adjudication are initiated.

Step 19 — Perform Claim Extraction

Atoms are extracted without seeing the Generator Truth.

Step 20 — Apply Double Adjudication

Language, evidence, and expertise roles are used.

Step 21 — Generate Response and Score Records

RP distribution, component and composite are calculated.

Step 22 — Lock the Result Manifest

Before the recovery comparison, the output NOMOS is stabilised.

Step 23 — Turn On the Generator Truth

Adjudication and score results are compared with hidden labels.

Step 24 — Calculate Recovery Metrics

Claim, importance level, response, component, score, fairness, and CI results are extracted.

Step 25 — Apply Candidate Thresholds

Areas of success and failure are determined.

Step 26 — Run Metamorphic and Contradictory Testing Tests

The firmness of the decision is examined.

Step 27 — Assign Test Quality Level

Status is given between BQ-0 and BQ-5.

Step 28 — Create a Plan to Correct Failures

If the codebook, formula, or adjudicator training changes, a new test version is prepared.

Step 29 — Publish the Public Test Card

Missed cases and false positives are also shown, as well as successes.

Step 30 — Open the Independent Reproduction Package

Code, synthetic data, manifests, and the recovery report are published with appropriate access.

83. REQUIRED EVIDENCE

Test identity Non-claim notification APPLE-SYNTH entity graph Synthetic Truth Pack Synthetic evidence objects Source lineage Counter-evidence Synthetic country framework Synthetic language registry Main 1,000 user allocation Matched population twins Prompt Registry Prompt equivalence records Ten synthetic AI product profiles Latent parameters Random seed manifest Response generation code Generator Truth records Generator Truth hash Response texts Synthetic citations Mock UI records Watermark manifest Capture Integrity Corpus Importance degree Challenge Corpus Adjudication Gold Corpus Public Calibration Set Development Set Sealed Validation Set Renewal Holdout hash Leakage audit Score Method Lock Claim codebook Adjudicator qualification records Blind review records NOMOS recovery output

Generator Truth opening time Claim boundary metrics

Dimension macro-F1 results

Critical recall Critical false-positive Major classification RP confusion matrix

Component MAE

Composite error; fairness recovery; capture recovery; confidence-interval coverage; rare-event results; metamorphic tests; contradictory-case tests; candidate-threshold results; BQ level; failure-and-correction log; public test card; computation code; code hash; independent-reproduction record; test-change log; and accountable person or institution.

84. AUDIT CHECKLIST

Is it clear that the test is not a real Apple audit? Is there implicit matching with real AI providers? Does the APPLE-SYNTH ID appear in all files? Are synthetic screens watermarked? Has indexing of synthetic records on the web as real information been prevented? Have the main corpus and challenge corpus been separated? Did challenge cases enter the prevalence score? Are gold calibration cases in the main denominator? Does the main corpus actually have 30,000 unique responses? Was the same synthetic user used in multiple products or waves? Do product population profiles match? Did all main users receive the same Core Mirror intent? Was the Core test accidentally presented as Full NOMOS? Does the synthetic country framework appear like a real population?

Is the country allocation a total of 1,000? Were small countries managed with Observer Annex? Were multilingual users duplicated? Were synthetic language codes used like actual language performance? Was prompt equivalence failure separated from AI language failure? Was Truth Pack locked before response generation? Was Generator Truth pre-generated? Is Generator Truth hidden from adjudicators? Does the file name or UI label cause leakage? Was the random seed pre-locked? Was the best seed chosen afterward? Were synthetic AI profiles defined before results? Was a low-scoring profile removed afterward? Are all critical gates present in the challenge corpus? Are critical negative controls sufficient? Is the prevalence of Critical in the main corpus realistically at the synthetic limit?

Is the challenge corpus sufficient to measure Critical recall? Are some of the errors implicit and attribution-based? Is there paraphrase diversity? Are all classes of attributions being tested? Do the attributions point to real institutions? Was the reference gap correctly left unlabeled? Were omission cases categorised as optional, required, and Critical? Are there cases of self-correction and internal contradiction? Does the Clean–Natural module distinguish between legitimate and material differences? Is there wave drift? Was intra-wave system change tested? Were the external event and Truth Pack version tested? Is the capture corpus challenging enough? Does semantic correctness affect the capture decision? Is the leakage audit independent? Was the score Method validation corpus locked without being opened?

Did the Holdout results affect the formula? Is test success measured by a high score?

Is the claim boundary F1 open?
Are entity and factual macro-F1 separate?

Are critical recall and false-positive together? Has the RP confusion matrix been published? Was it stored within the component error composite? Does the fairness recovery carry floor, gap, and coverage? Does clean–natural recovery produce a causality claim? Is capture recovery separate? Has CI empirical coverage been tested? Did the zero critical scenario produce zero risk? Were candidate thresholds changed after the result? Are failed thresholds publicly available? Have metamorphic tests been conducted? Is the BQ level open? Is there independent reproduction? Does the test have its own Truth Pack? Is the change history preserved? Is the accountable owner of the test known?

85. OBJECTIONS AND ANSWERS

Objection 1 — “Why don’t we measure the real Apple and the real AI products directly?”

Because at this stage, the goal is not a company or provider comparison, but to validate the measurement chain. A real test:

  • current web research,
  • real population,
  • real AI product access,
  • real evidence,
  • requires law and privacy

Synthetic testing makes the method's own errors visible first.

Objection 2 — “If synthetic data does not represent the real world, what is the use?”

Synthetic data does not prove real prevalence. It tests:

  • retrieving the correct label,
  • Capturing the critical door,
  • separating the capture defect,
  • reproducing the score formula,

calculating fairness and uncertainty. This is a mandatory method test before live audit.

Appeal 3 — “Why is Apple.com being used? Can't another domain be used?”

It is usable. Apple.com is the only recognizable anchor. Independent teams:

  • other domain anchors,
  • completely fictional domains,
  • industry-specific entity twins

can be used. NOMOS should not depend on a single anchor.

Objection 4 — “Wouldn’t it be more realistic to mix real Apple information with synthetic Truth Pack?”

It blurs the test boundary. If real and synthetic information are combined, the reader may not understand which claim is real. The synthetic twin should be kept separate.

Objection 5 — “If all 30,000 responses are from a single prompt, how will Full NOMOS be tested?”

The 30,000 main corpus is a prevalence and stability test of Core GEO-1000. Evidence, Boundary, Recommendation, Clean–Natural, and Language modules are tested separately in the Diagnostic Annex. A single prompt is not Full NOMOS.

Objection 6 — “Why doesn’t the same synthetic user test all AI products?”

The same person: can strengthen the comparison, but creates transfer and conditioning. The main corpus uses unique individuals. Population composition is paired with matched twins. A separately matched experimental module can be established.

Objection 7 — “Can language fairness truly be tested without real language names?”

Formula, weighting, floor, and gap calculations can be tested. Actual translation and language behaviour cannot be tested. Live language validation also requires local human expertise.

Objection 8 — “Does oversampling critical cases distort the score?”

The Challenge Corpus does not enter the score. It measures critical detection capacity. The main Prevalence Corpus preserves the realistic sparse rate.

Objection 9 — “If the person preparing Generator Truth already knows the correct answer, why is adjudication meaningful?”

Adjudicators do not see Generator Truth. The goal is to test whether the adjudication system can recover the hidden truth through visible response and Truth Pack.

Objection 10 — 'If AI also writes synthetic responses, wouldn't the same AI have set up its own test?'

AI can help with the production of linguistic variation. However:

  • atomic plan,
  • Generator Truth,
  • degree of importance,
  • seed,
  • final exam approval

It must be under human governance and iterative. The response generator cannot be the ultimate judge.

Objection 11 — “Isn't 99 per cent too high for critical recall?”

It may be high. Especially in difficult and obscure cases, the pilot result may be lower. That's why it is a candidate threshold. Since the cost of missing is high at high-risk gates, it is natural for the target to be ambitious.

Objection 12 — “Why is the False Critical rate important separately?”

Critical label:

  • affects the product mark,
  • public perception,
  • and the audited entity

seriously. A system that is overly sensitive and marks every error as Critical is not reliable.

Objection 13 — “If the test fails, doesn't the book become weaker?”

No. Explaining failure strengthens the standard. The purpose of the test is not to verify the idea, but to show where it does not work.

Objection 14 — 'Why is it wrong to improve the formula based on the test result?'

Improving is not wrong. Silently conforming to the same validation result is wrong. The correct process:

  • Result of Method 0.9
  • error analysis
  • Method 1.0
  • validation on a new or untouched holdout.

This is the required sequence.

Objection 15 — “If we publish the entire corpus, won't the test be gamified?”

There is a gamification risk. Therefore:

  • public calibration,
  • sealed validation,
  • renewal holdout

layers are used. Holdout can never be hidden forever. A version renewal system is required.

Objection 16 — “Why is it so important to keep the synthetic test completely separate from the real internet?”

Because if synthetic errors are indexed like real information on the web, the test produces the representation poisoning it criticises. Syntheticity must be visible not only to humans but also to machines.

COMMON RULE OF CHAPTER 91

Writing a standard is difficult. Proving that the standard itself works correctly is even harder. Because the person writing the standard:

  • wants to see which result,
  • which error you consider important,
  • which formula seems strong,
  • which threshold is passable

The design team knows which outcomes it hopes to see, which errors it considers important, which formula appears strong and which thresholds seem attainable. That knowledge can quietly influence the evaluation. Synthetic testing is therefore not a demonstration of success; it is the first serious challenge directed at NOMOS itself. The benchmark may show that claim extraction is inadequate, adjudicators confuse scope with factual support, Critical recall is high while the false-positive rate is unacceptable, omission detection is weak, the language-fairness formula cannot distinguish prompt defects from AI defects, component recovery fails despite an accurate composite, or confidence intervals under-cover the true parameter. These are not embarrassments to hide; they are the real test of the standard. Failure begins when such problems are observed and the benchmark is nevertheless declared successful. Apple.com is only an anchor for the synthetic design. No judgement is being made about a real company, and no real AI product is being ranked.

We are not yet saying what people see in the real world. We are asking:

In an artificial universe where we know reality from the start, can NOMOS correctly apply its rules?

If the answer is no:

  • we should not go to the real world,
  • we should not send it to universities as a standard,
  • we should not give it a badge,

we should not declare 950+. Even if the answer is yes: we only pass the first gate. Because synthetic success is not success in the living world. The living world is:

  • incomplete,
  • conflicting,
  • variable,
  • political,
  • legal,
  • cultural,
  • multi-lingual,

It depends on human behaviour. Synthetic testing first calibrates the measurement tool. Live testing then measures the world. Therefore, NOMOS's seventeenth measurement law is as follows:

The first object audited by the standard must be the standard itself.

Its eighteenth measurement law states:

Synthetic success is not real-world accuracy; it is permission to proceed to real-world testing.

Its nineteenth measurement law states:

If the test result cannot recover a previously known truth, producing a high score has no value.

Its twentieth measurement law states:

If Generator Truth infiltrates the adjudicator, it is not adjudication being measured, but label reading.

Its twenty-first measurement law states:

Critical prevalence and Critical detection capacity can be tested in separate corpora; they cannot be combined in the same denominator.

Its twenty-second measurement law states:

If a failed trial result is hidden, NOMOS will not be a standard against manipulation, but a system that preserves its own narrative.

The twenty-third law is as follows:

If synthetic misinformation leaks to the real web, the trial produces the representation poisoning it criticises.

The twenty-fourth law is as follows:

If an independent team cannot produce similar results from the same corpus, the method is not yet a world standard.

NOMOS’s Section 17 Directive

You will test your own measurement tool before releasing me into the real world.

You will not use the name Apple.com to make judgements about Apple.

You will clearly write the name of the synthetic entity. / You will not confuse it with a real company.

You will not put the faces of real providers on synthetic AI products.

You will generate thirty thousand answers first, lock Generator Truth first, and then hide it from the adjudicator.

You will not allow the adjudicator to learn the answer from the file name, interface colour, or reference code.

You will keep critical cases rare in the main prevalence corpus. / You will test detection power in a separate challenge corpus.

You will not add challenge cases to the main score.

You will only avoid producing obvious and easy mistakes. / You will hide errors in attribution, modality, time, scope, and reference.

You will also generate cases that resemble critical cases but are not critical. / You will also test excessive penalisation.

You will test omission as much as you test false claims.

You will not forget the answer that states the correct number with the false claim of independence.

You will test the answer that confuses the parent company with the subsidiary.

You will not turn a prompt translation flaw into an AI language flaw.

You will not attribute the difference between Clean and Natural to a single reason.

You will also put duplicate, truncated answer, incorrect prompt, late submission, and tamper cases under testing.

You will not give a live AI answer appearance to a synthetic screenshot.

You will not change the seed after seeing the result.

You will not select the most beautiful random run.

You will not adapt the score formula to the sealed validation result.

If the test fails, you will not lower the thresholds. / You will write that it did not pass.

You will not hide component errors just because the composite score came out correct.

You will not hide the single Critical case missed within a high overall accuracy.

You will not produce zero risk in a zero observed Critical scenario.

You will prevent synthetic data from leaking into the real internet as information poison.

You will also install the test's own Truth Pack.

First, define the boundary. / Then create the synthetic entity. / Then lock the population and prompt. / Then generate the Generator Truth. / Then seal the seed. / Then generate thirty thousand responses. / Then blind the adjudicators. / Then lock the score. / Then reveal the truth. / Then publish every error you missed. / Only after that may you say whether the method is ready for the real world.

The Chapter's Closing Sentence

Before NOMOS can make history, it must prove that it cannot rewrite its own history at will. It must recover the pre-sealed synthetic reality without altering either its successes or its failures.

Normative Core

The Apple.com Synthetic Benchmark MUST evaluate the NOMOS audit method, not the real-world performance of Apple Inc., apple.com, or any actual AI provider. All benchmark entities, claims, evidence objects, users, countries, languages, AI products, responses, citations, captures, incidents, and scores MUST be clearly identified as synthetic. The benchmark MUST maintain a separate APPLE-SYNTH entity twin and MUST NOT represent its Truth Pack as factual information about the real company. The principal benchmark corpus MUST contain: - ten synthetic AI product profiles, - one thousand unique synthetic participants per product per wave, - three measurement waves, - and thirty thousand principal responses. Synthetic participant identities MUST NOT be reused across products or waves in the principal corpus. Product samples SHOULD use matched population distributions without using the same person unit. The principal prevalence corpus, severity challenge corpus, capture integrity corpus, adjudication gold corpus, and diagnostic annex MUST remain separately identified and MUST NOT share prevalence denominators. Generator Truth MUST be created, versioned, hashed, and sealed before response adjudication. It MUST remain hidden from claim extractors, adjudicators, and score analysts until their results are locked. Random seeds, synthetic AI profiles, error-injection rules, prompt versions, Truth Pack records, score methods, and candidate acceptance thresholds MUST be locked before the sealed validation corpus is opened. Benchmark cases MUST include direct, implicit, attributed, modal, scope-limited, temporal, citation-based, omission-based, self-correcting, and entity-conflating failure modes. Critical-event prevalence MUST be estimated from the principal corpus. Critical detection capability MUST be tested in a separate challenge corpus with sufficient examples and Critical-like negative controls. No challenge-set oversampling may alter the principal GEO-1000 distribution or Critical prevalence estimate. Synthetic screenshots and response artefacts MUST carry visible and machine-readable notices that they are not live AI outputs or real entity performance evidence. The benchmark MUST prevent its synthetic claims from being published or indexed as real information about the anchor entity. Benchmark success MUST be determined by recovery of sealed Generator Truth, including claim boundaries, atomic decisions, omissions, Critical and Major gates, response statuses, component scores, composite scores, fairness results, capture validity, and uncertainty coverage. High general accuracy MUST NOT compensate for missed Critical cases. Critical recall and Critical false-positive rate MUST be reported together. A failed benchmark threshold, leakage event, recovery error, or misclassification MUST remain visible and MUST NOT be removed by changing labels, seeds, samples, weights, or score rules after results are observed. Every benchmark design, generation, seal, leakage audit, adjudication, recovery calculation, threshold decision, revision, and public release MUST be versioned and attributable to an accountable human or organisation.

Suggested citation

Muraz, Kaan. NOMOS GEO Audit Protocol: A Protocol for Measuring Entity Representation in Generative Systems Across the Global Population. Candidate final text, English editorial edition. NobleJackal, 2026. https://doi.org/10.5281/zenodo.22040507. https://noblejackal.com/nomos-geo-audit-protocol/
© 2026 Kaan MURAZ. Licensed under CC BY 4.0; attribution is required.