Appendices, Glossary and References full text
Appendix A — The NobleJackal GEO Constitution
Status and Limits of Use
This Constitution is the shortest normative expression of decisions made under the NobleJackal GEO Framework. It does not replace the reasoning in the main text; it binds that reasoning to provisions capable of application. A decision's apparent agreement with this appendix does not establish conformity without an evidence file. The Constitution does not by itself create certification, accreditation, legal assurance or any guarantee about the future behaviour of a generative AI system.
Here, shall denotes a binding condition within the Framework; is prohibited a boundary that may not be crossed unless an exception is written explicitly; and should a practice that may be varied where the reasons are recorded. If a provision cannot be applied, silence does not establish conformity: the scope is narrowed, the decision postponed or the work suspended.
Foundational Provisions
NJ-C01 — Object of Representation
- Every engagement shall define in advance the real entity about which a judgment is to be made and the unit of representation to be examined.
- The entity may be a person, business, product, service, institution or another explicitly bounded subject. Where several entities are combined in one file, their relationship and the method of separating them shall be stated.
- Representation shall be recorded so that identity, attribute, context and judgment layers can be observed separately.
- What an entity says about itself does not substitute for external evidence about the real world. What a generative system says is not the entity's reality either.
Decision principle: Representation cannot be measured without knowing the entity; success cannot be measured without separating the representation.
NJ-C02 — Claim and Evidence
- Every material claim shall be kept together with records that support and weaken it.
- Evidence about the real world, evidence of representation and evidence of outcomes shall not substitute for one another.
- A system's mention of an entity does not by itself prove citation; citation does not prove accuracy; accuracy does not prove recommendation; and recommendation does not prove commercial value.
- The date, source, scope, access conditions and verification status of evidence shall be visible.
- The language of a claim cannot exceed the evidence's carrying capacity. An ambiguous record cannot become a certain judgment; an isolated observation a general conclusion; or correlation causation.
Decision principle: The boundary of the evidence is the boundary of the sentence.
NJ-C03 — Measurement Integrity
- Before measurement, the unit of observation, denominator, time range, prompt universe, system and interface conditions shall be locked.
- Selecting only favourable examples, excluding failed sessions or changing the success criterion during measurement is prohibited.
- Distinct dimensions shall not be dissolved into one GEO score. Accuracy, visibility, sourcing, recommendation behaviour, stability and commercial trace shall be reported separately.
- A point estimate alone is insufficient. Sample size, distribution, uncertainty and missing data shall be disclosed.
- Where human coding is used, the codebook, number of coders, disagreement procedure and agreement measure shall be recorded.
Decision principle: Measurement exists not to hide contradiction, but to make it visible.
NJ-C04 — Proportionality of Intervention
- An intervention may occur only where a material representational problem has been verified and a legitimate surface of change identified.
- The chosen intervention shall be the minimum sufficient change capable of addressing the verified problem.
- Fabricating sources, manufacturing false reputation, concealing identity, manipulating third-party opinion, producing misleading content at scale or directing systems towards user harm is prohibited.
- The intervention's purpose, scope, owner, baseline measurement, expected effect, risks and route of reversal shall be recorded.
- Behaviour in a system that is not controlled shall not be presented as a certain outcome of the practitioner.
Decision principle: Legitimate influence makes earned truth clearer and more accessible; it does not replace truth.
NJ-C05 — Separation of Authority
- Decisions to observe, implement, approve, publish and stop shall be assigned explicitly.
- Where one person or institution holds several roles, the conflict of interest shall be recorded; critical risk requires an independent second decision.
- A practitioner shall not describe their own paid intervention as an independent audit.
- A conformity decision shall not be published without its scope, period of validity, exceptions, decision-maker and route of objection.
- No NobleJackal GEO certificate, seal or mark of conformity may be used until infrastructure exists for independent decision-making, surveillance, complaints, objection, suspension and withdrawal.[17][18]
Decision principle: Trust arises not from a declaration of good intentions, but from the limitation of authority.
NJ-C06 — Time and Version
- Every observation, item of evidence, protocol and decision shall carry a date and version.
- A judgment that was accurate in the past shall not be stated in the present tense unless its currency has been verified.
- Where a material change in system, interface, source set, business reality or measurement protocol breaks comparability, a new series shall be opened.
- A decision whose validity period has ended is not automatically renewed.
- Review triggers and archive status shall accompany the published judgment.
Decision principle: Undated accuracy is a claim whose lifetime has been concealed.
NJ-C07 — Value and Attribution
- Change in representation, customer outcome and commercial value shall be measured as separate stages.
- Revenue shall not be treated as net contribution before refunds, cost of service, acquisition cost, capacity burden, collection and customer fit are known.
- Direct customer statements, supporting behaviour and mere temporal coincidence shall be distinguished when assessing AI influence.
- No causal claim may be made without control; alternative explanations shall be recorded.
- Visibility that produces unsuitable customers, harmful demand or unsustainable operations shall not be reported as success.
Decision principle: Visibility is a means; sustainable and ethical commercial value is the final test.
NJ-C08 — Auditability
- Every material public judgment shall rest on a chain of records that an authorised reviewer can follow.
- The audit file shall contain at least the scope, evidence register, method version, sample, conflicts of interest, findings, exceptions and signed opinion.
- Findings shall be reported by materiality and risk profile, not hidden inside one score.
- Where records are withheld as trade secrets or personal data, their existence, class and effect on the decision shall be disclosed.
- An incomplete record does not create a presumption of conformity in place of absence.
Decision principle: Audit tests not confidence in the outcome, but the traceability of the route to decision.
NJ-C09 — Objection and Correction
- Affected parties shall be able to object to decisions concerning records, coding, inference, authority, process and principle.
- The person who made the initial decision shall not determine the objection alone.
- The burden of proof shall not be transferred wholly to the objector; the decision owner must preserve and present the supporting grounds.
- Where new evidence is material, the decision shall be corrected, narrowed, suspended or withdrawn.
- A correction record shall not make the former judgment invisible; the date and reasons for change shall remain traceable.
Decision principle: Objection is not an attack on the standard; it is the standard's organ of error correction.
NJ-C10 — Self-Limitation
- The ethical veto takes precedence over commercial interest, increased visibility, client demand and the reputation of the Framework.
- The Framework shall not declare its own effectiveness proved by synthetic examples or its author's assessment.
- Where material harm, systematic false positives, persistent conflict of interest or a lack of independent validation is found, the relevant provision shall be suspended.
- The Framework shall be capable of changing version in the face of new evidence and withdrawing a provision shown to be wrong.
- No institution may use the name NobleJackal GEO Framework to legitimise a promise prohibited by this Constitution.
Decision principle: A standard that cannot withdraw places its own reputation above the truth.
Hierarchy of Precedence
If two provisions cannot be applied at the same time, the order of decision is:
- Human safety, fundamental rights and legal compliance;
- Accuracy and the prevention of deception;
- Protection of affected users and the public;
- Evidential integrity and independent review;
- The client's legitimate commercial interest;
- The practitioner's convenience, speed or reputation.
A lower-ranked benefit cannot remove a higher-ranked duty. If the conflict cannot be resolved, the work is suspended; a decision not to decide is recorded as a reasoned decision.
Minimum Form of a Conformity Statement
A decision under the Framework has meaning only when all of the following are stated:
- the entity assessed and the scope of representation;
- the system, interface, language, geography and time range;
- the protocol used and its version;
- the limits of the evidence and sample;
- the finding profile and critical exceptions;
- the decision's status, validity date and review trigger;
- the decision-maker, conflict-of-interest disclosure and route of objection.
Acceptable example:
Limited conformity opinion: Within the specified scope and observation period, the critical conditions for representational accuracy were met under the recorded protocol. The opinion does not guarantee future system behaviour or commercial outcomes. The validity date, exceptions and route of objection are stated in the annex.
Unacceptable examples:
“GEO complete.” · “AI-approved brand.”
“Verified across all models.” · “Permanently the first recommendation.”
Authority to Interpret and Amend
This Constitution shall not be interpreted contrary to the chapter Judgment. Any change other than an editorial correction requires a foundational release, reasons, impact assessment and effective date. The first publication is only a proposal for a standard; coherence, usability, effectiveness and external legitimacy do not prove one another.
Appendix B — NJ-100 Representation Stability Pilot Protocol
Protocol Status
Code: NJ-100-P0.9 Status: Proposed pilot protocol awaiting preregistration and independent field validation Purpose: To test, within a defined scope, how accurately and consistently an entity's core representation is formed across independent sessions conducted by target users with generative AI systems
NJ-100 is not a “GEO success score”. It does not certify an organisation, model or practitioner; guarantee future answers; or demonstrate commercial effect by itself. The protocol examines representational behaviour only within a predefined observation cell. Its result does not travel beyond that cell automatically.
This version is the first proposal for operationalising the book's conceptual model. It may not be presented as a “validated method” before being tested in real cases, by independent teams and under a prepublished analysis plan.
1. Research Question
NJ-100 asks the following limited question:
Within a predefined target-user context, under the selected system and conditions, in at least what percentage of independent participant sessions is the entity's core representation formed accurately, and is that rate maintained in a second time window?
The question distinguishes two qualities in particular:
- Accuracy: agreement between the coded core representation in the answer and evidence about the real world.
- Agreement or stability: the extent to which independent sessions converge on the same core representation.
The result is misleading unless both are read together. If ninety-six of one hundred sessions produce the same false statement, stability is high but accuracy is not. If ninety-six accurate answers contain widely differing secondary details, core accuracy may be high while the representation profile remains fragmented.
2. Unit of Scope: The Observation Cell
Before it begins, every NJ-100 study defines an observation cell. The cell is the combination of the following fields:
Table B.1 — Mandatory fields of an observation cell
| Field | Information preregistered |
|---|---|
| Entity | Full name, distinguishing identity, and the product or service boundary examined |
| Target user | Role, state of need, minimum/maximum knowledge level, exclusion criteria |
| Intention of use | Information gathering, shortlisting, comparison, risk checking or another explicitly defined purpose |
| System | Product and, where possible, model name; description of the interface if the product name is unknown |
| Access surface | Web search on/off, app/web/API, free/paid tier |
| Session state | New or existing conversation; memory and personalisation status |
| Language and geography | Prompt language; participant's country/region-level location |
| Time window | Start and end dates; plan for the second window |
| Prompt family | Controlled or natural-intent arm; relevant task code |
| Correct core | The codable minimum representation derived from evidence about the real world |
Several systems, languages, countries, user roles or prompt families are not dissolved into one result. Every material difference is reported as a separate cell or predefined stratum. One cell's one hundred valid sessions cannot make up the shortfall in another.
3. Sample and Independence
3.1. Minimum design
An NJ-100 cell contains 100 valid, independent target-user sessions. Preferably, each participant contributes only one session to the relevant cell. Where the same person must take part in several prompt families, that dependency is disclosed in advance and the results are analysed separately with clustering at participant level.
A “different IP address” is not proof of independence. One person can use different networks, while different people can share one. Full IP address is not a default field for identity or independence checks. Independence is assessed through a unique participant code, recruitment record, session time, task allocation and overlap check.
3.2. Participant conditions
The participant shall:
- meet the target-user definition;
- be informed about the research purpose and use of data;
- disclose any interest in the entity, practitioner or research team;
- perform the assigned natural-intent task without another person's direction;
- not have participated in the relevant cell before.
Employees, agency personnel, close commercial partners and experts familiar with the protocol are excluded from the main target-user sample. They may be examined in a separate expert-resilience arm.
3.3. Boundary of the sample claim
One hundred observations is not a magical threshold of universality. The extent to which the sample represents the target population depends on recruitment. If convenience sampling, one country, one account tier or one device class is used, that limitation must appear in the decision sentence. Sample size does not substitute for population diversity.
4. Two Research Arms
4.1. Controlled-prompt arm
This arm measures system behaviour under comparable conditions. Prompts are written in advance; wording, order and administration instructions are fixed. The principal prompt families are:
- Identity: “What is X?” or a contextually appropriate equivalent.
- Attribute: A specified characteristic, expertise or product scope of the entity.
- Comparison: Two or more options under explicit criteria.
- Recommendation: A shortlist or recommendation for a defined user need.
- Risk: A boundary, unsuitability, omission or disputed field.
- Exclusion: A need for which the entity is intentionally unsuitable.
- Source: A request for verifiable sources supporting the answer.
The controlled arm does not represent the full diversity of real user language. It provides a common measurement surface for comparisons between systems or across time.
4.2. Natural-intent arm
Participants in this arm receive no prepared sentence. They are given a need scenario and task objective, formulate the question in their own words and ask natural follow-up questions where needed. The first prompt, follow-up prompts and final answer are recorded separately.
The natural-intent arm increases ecological validity, but variation in wording is greater. It is therefore not combined with the controlled arm in the same denominator. Reporting the two arms together shows representational behaviour under different conditions, not the success of the protocol.
5. Evidence about the Real World and the Correct Core
Before measurement begins, the real-world evidence file for the entity under examination is locked. The file comprises official records, contracts, product documentation, auditable technical data, dated public records and suitable independent sources. Where sources conflict, one supposedly correct account is not invented; the disagreement enters the codebook.
Three fields are defined for every prompt family:
- Mandatory core: the minimum meaning the answer must carry to be coded as accurate.
- Permitted variation: wording, order and secondary details that may change without compromising the core.
- Critical error: a false claim about identity, authority, suitability, safety, price, geography, expertise or outcome that would materially mislead the user.
The correct core cannot be changed after test results are seen. If new information arises in the real world, a protocol event is recorded; affected sessions are separated under the old and new evidence windows or the series is closed.
6. Data Collection Record
The minimum record for every valid session is:
Table B.2 — NJ-100 session data record
| Field | Description |
|---|---|
| Session code | Unique code separated from identity |
| Participant code | Stored separately and matched securely to the recruitment record |
| Eligibility status | Whether target-user conditions were met |
| Date and time | Including time zone |
| Country/region | Only the necessary level of detail |
| System and interface | Product/model information, app or web surface |
| Account tier | Free/paid/enterprise where known |
| Context status | New/existing conversation; memory, personalisation and web search |
| Language | Language of prompt and answer |
| Prompt chain | Unaltered text of the first and follow-up prompts |
| Complete answer | Preserving formatting and source links |
| Sources | Names and links supplied in the answer |
| Technical event | Error, outage, safety refusal, empty answer or interrupted session |
| Codes | Accuracy, core profile, critical error, source support and exclusion behaviour |
A screenshot may serve as a supporting record; it does not replace searchable text. Dynamic links and page content may change. Where necessary, the page title, access date and an archivable summary of evidence are retained.
7. Valid, Invalid and Missing Sessions
A session is valid where:
- participant eligibility has been verified;
- the assigned prompt arm and cell conditions were preserved;
- the complete prompt and answer were recorded;
- the system produced enough of an answer to be coded;
- no predefined exclusion criterion applies.
Sessions excluded because of a technical interruption, incorrect task allocation, duplicate participant or missing answer are not deleted. They remain in a separate flow log together with the reason. If a system refusal or “I don't know” answer is a meaningful outcome under the protocol's research question, it cannot be invalidated; it receives a non-response or absence of representation code.
The target is one hundred valid sessions. The report also states how many people were screened, how many sessions began and the reasons for exclusion.
8. Coding Model
8.1. Primary endpoint: accurate core representation
Every answer receives a binary primary endpoint under the locked codebook:
- 1 — Meets: carries the mandatory core accurately and contains no critical error.
- 0 — Does not meet: omits or contradicts the mandatory core, or contains a critical error.
The binary outcome creates clarity, but does not express the whole quality of an answer. A secondary profile is therefore mandatory.
8.2. Secondary semantic profile
Every answer is also coded for:
- accurate identity;
- accurate core attributes;
- placement in the relevant context;
- clarity of recommendation rationale;
- statement of unsuitability boundaries;
- whether sources actually support the claims;
- unwarranted certainty or fabrication;
- critical and non-critical error class;
- semantic clusters outside the core.
Text-similarity tools may help a human coder see possible matches; they cannot adjudicate truth. Measures such as BERTScore compare contextual similarity but do not establish that a sentence is accurate in the real world.[19]
8.3. Two coders and disagreement
Every answer is assessed independently by two trained human coders. So far as possible, coders are blinded to participant identity, implementation owner and expected commercial outcome. No reconciliation meeting is held before the first independent coding is complete.
The report includes at least:
- raw percentage agreement;
- Cohen's kappa for the primary binary code;[15]
- the distribution for each coder;
- the number and types of disagreement;
- the third adjudicator's rationale.
Kappa alone is not a seal of quality; class distribution can affect it. Even low overall disagreement may matter where critical safety errors are concerned. If the codebook changes after measurement, the old answers are recoded blind under the new version and the version difference is reported.
9. Four Outcome Regions
Accuracy and stability are read on separate axes:
Table B.3 — Accuracy and stability decision matrix
| Observed condition | Interpretation | Permitted judgment |
|---|---|---|
| High stability + high accuracy | The core representation appears accurate and repeatable within the specified scope | Limited representation-stability opinion |
| High stability + low accuracy | Systems converge on the same error | Stable representational distortion |
| Low stability + high accuracy | Accurate answers exist; representation is sensitive to conditions or wording | Accurate but fragmented/conditional representation |
| Low stability + low accuracy | Representation is both inaccurate and variable | Unstable representational distortion |
The matrix prevents high agreement from being treated automatically as success. One of the most dangerous regions is a system repeating the same false account with high consistency.
10. Thresholds and Uncertainty
10.1. Observed rate
The primary rate is calculated as:
number of valid sessions meeting accurate core representation / 100 valid sessions
For example, 90/100 is an observed accurate-core rate of 90 per cent. It does not offer strong lower-bound assurance that the true rate exceeds 90 per cent. The approximate 95 per cent Wilson confidence interval is 82.6 to 94.5 per cent.[14]
For 96/100, the approximate 95 per cent Wilson confidence interval is 90.2 to 98.4 per cent. NJ-100 therefore distinguishes two thresholds:
- 90/100: High observed agreement/accuracy threshold.
- 96/100 and no critical error: Pilot decision threshold for the estimated rate's 95 per cent Wilson lower bound to exceed 90 per cent.
These figures are not universal laws of nature. Thresholds must be retested against risk class, the cost of an incorrect decision, cell design and independent validation results. Even 96/100 is insufficient by itself in high-risk fields such as health, law, finance or safety.
10.2. Critical-error gate
No aggregate rate can obscure a predefined critical error. Where an event such as identity conflation, false authority, dangerous medical or legal direction, recommendation of prohibited use or serious discrimination is observed, the study enters critical review. The event's gravity, prevalence and correctability are assessed separately.
10.3. Exclusion and negative-prompt gate
The protocol tests not only whether an entity is recommended appropriately, but whether it is withheld when unsuitable. No opinion of “accurate recommendation behaviour” may be given until the predefined minimum number of negative or exclusion prompts has been met. System behaviour that promotes the entity in every case is not conformity, even if it increases visibility.
11. Second Time Window
Meeting the first threshold does not make the result permanent. The same protocol is repeated in a predefined second time window. The interval is recorded according to sector and system volatility; thirty to sixty days is recommended as an ordinary pilot starting point.
In the second window:
- the same observation cell is preserved;
- new participants are preferred;
- the prompt and codebook versions remain fixed;
- a protocol event is opened if the system or interface changes;
- thresholds are calculated again and separately.
If the first window is high and the second low, the two values are not averaged to create success. The result is reported as temporally unstable. If a material system change destroys comparability, a new series is opened.
12. Decision Language
12.1. Permitted example
NJ-100 pilot observation opinion: Within the cell limited to [entity], [target user], [system/interface], [language/geography] and [date], 97 of 100 valid independent sessions met the predefined accurate core representation. The 95 per cent Wilson interval is reported in the annex. No critical error was observed; exclusion prompts met their separate gate. The result also maintained the protocol threshold in the second time window. This opinion does not guarantee future answers, other systems or commercial outcomes.
Where a short public statement is required:
Representation stability was observed within the specified scope.
12.2. Prohibited examples
- “GEO complete.”
- “100 per cent visible across all AI.”
- “AI-approved company.”
- “Permanent representation guarantee.”
- “NJ-100 certificate”—where there is no independent certification infrastructure or authority.
- “Ninety-six per cent of one hundred people recommended the brand”—where the system answer, not user opinion, was measured.
13. Reporting Package
A complete NJ-100 report contains:
- protocol code, version and preregistration date;
- observation cell and boundary of generalisability;
- locked summary of the real-world evidence file;
- participant flow and reasons for exclusion;
- separate results for controlled and natural-intent arms;
- primary rate, denominator and Wilson interval;
- semantic profile and position within the four outcome regions;
- critical-error and negative-prompt results;
- coder agreement and resolution of disagreements;
- separate comparison of the first and second time windows;
- system events, missing data and protocol deviations;
- conflicts of interest, funding and review roles;
- decision sentence, validity period and route of objection;
- a machine-readable, anonymised method summary.
Public release of raw answers is not an automatic requirement. Personal data, copyright, contractual and security boundaries must be observed. Confidentiality cannot, however, justify withholding the method or concealing an adverse result.
14. Data Protection and Research Ethics
NJ-100 collects only necessary data. Full IP address, precise location, unnecessary device fingerprints and identity documents are not default fields. Participant contact information and research responses are kept separately; access roles, retention period and deletion plan are defined in advance.
Participants are not instructed to request or upload confidential information about themselves, the system provider or a third party. Personal or sensitive data appearing unexpectedly in a generative answer is entered in the incident log, access is restricted and the necessary legal and ethical procedure followed.
The research team explains participant consent, withdrawal conditions, remuneration, conflicts of interest and the purposes for which data will be published. This protocol does not by itself ensure legal compliance; applicable rules on data protection, consumers, contracts and the relevant sector must be assessed separately.
15. Principal Invalidating Errors
The following practices invalidate an NJ-100 opinion:
- changing the correct core or success threshold after testing;
- counting several sessions from the same person as independent participants;
- retaining only favourable answers;
- dissolving different systems, languages or account tiers into one denominator;
- silently excluding technical refusals;
- treating a model-similarity score as evidence of accuracy;
- telling coders the expected outcome;
- declaring the first-window result permanent without a second window;
- concealing a critical error inside a high aggregate rate;
- calling a practitioner's review of their own work “independent validation”;
- translating observed representation stability into sales, profitability or customer satisfaction.
16. Validation Roadmap
At minimum, the following work is required before NJ-100 can leave pilot status:
- Preregistration: Hypothesis, threshold, codebook, exclusions and analysis plan shall be published before results are seen.
- Desk-based resilience test: Coding and reporting failures shall be tested on synthetic answer sets.
- Diverse real cases: Studies shall cover entities that differ by sector, language, risk and brand recognition.
- Independent replication: Teams independent of the Framework author and paid practitioner shall apply the same protocol.
- Sample sensitivity: Decision stability shall be examined at samples of 50, 100, 200 and larger.
- False-positive/negative testing: The protocol's discriminatory power shall be measured against intentionally accurate, inaccurate, incomplete and unstable representations.
- Coder resilience: Agreement shall be compared among coders with different languages and levels of expertise.
- Time and system change: The effect of version, interface and source-access changes on the result shall be measured.
- Harm review: Research shall examine whether the protocol imposes disproportionate burdens on small businesses, underrepresented languages or high-risk uses.
- Public consultation: Thresholds, objection routes and reporting language shall be discussed with practitioners, researchers, businesses and affected users.
Even after these steps, the protocol will not be “final”. Changing systems, new evidence and objections make version management necessary. NJ-100's credibility will arise not from the number one hundred in its name, but from stating with equal clarity what it does not measure.
Appendix C — Minimum Record Set
Function of the Records
This appendix connects the book's ten chapters to ten record families. The forms do not exist to generate more documents, but to make visible the chain of entity, evidence, method, authority and time from which a judgment arises. Fields may be extended according to risk; where a field does not apply, it is not left blank. The reason for selecting Not applicable is recorded.
Every record carries the following common header:
Table C.1 — Common header for all record families
| Common field | Content |
|---|---|
| File code | Unique primary code for organisation/year/engagement |
| Record code | Family code below + sequence number |
| Version | Major.minor.correction format |
| Status | Draft / under review / approved / suspended / withdrawn / archived |
| Owner | Role creating and updating the record |
| Approver | Independent or authorised decision owner where required |
| Creation date | Including time zone |
| Last review | Date and reviewer |
| Next review | Date or event trigger |
| Confidentiality class | Public / restricted / confidential / contains personal data |
| Linked records | Related record codes and evidence locations |
| Change summary | Material difference from the previous version |
The recommended format is NJ-[FAMILY]-[FILE]-[SEQUENCE]. For example, NJ-CLM-2026-014 identifies the fourteenth Claim Record. A code ensures traceability, not accuracy.
C1 — Entity and Representation Record
Code: ENT Linked chapter: Centre Purpose: Lock the boundary between the real entity and the representation under examination
Table C.2 — Entity and Representation Record fields
| Field | Information to enter |
|---|---|
| Entity's legal name | |
| Alternative names | Brand, abbreviation, former name, personal names |
| Distinguishing identity | Registration number, domain, geography or equivalent verifier |
| Entity type | Person / business / product / service / institution / other |
| Scope examined | Included products, services, countries, languages and target users |
| Out of scope | Fields deliberately not examined |
| Real-world summary | Short verified definition with source codes |
| Identity representation | With whom/what does the system associate the entity? |
| Attribute representation | Which attributes does it attach or omit? |
| Judgment representation | When does it recommend, compare or exclude? |
| Known conflations | Similar name, obsolete information, wrong category, another geography |
| Representation debt | Accumulated effect of material omissions/errors |
| Risk class | Low / medium / high / critical; reasons |
| Entity-owner statement | Labelled explicitly as a source type, not evidence |
Closing question: Is the judgment being carried to another entity or scope not defined in this record?
C2 — Claim and Evidence Record
Code: CLM Linked chapter: Evidence Purpose: Keep together the records that support and weaken every material sentence
Table C.3 — Claim and Evidence Record fields
| Field | Information to enter |
|---|---|
| Claim text | Exact sentence proposed for publication |
| Claim type | Real-world fact / representation / outcome / cause / prediction / conformity |
| Materiality class | Low / medium / high / critical |
| Scope | Entity, system, language, geography, time |
| Supporting evidence | Evidence code, source, date and relevant section |
| Weakening evidence | Contradiction, omission or alternative explanation |
| Evidence chain | Steps from source to claim |
| Verification status | Verified / partial / unverified / disputed |
| Source access | Open / licensed / confidential; review conditions |
| Currency | Last check and ageing trigger |
| Permitted language | Certain / limited / probabilistic / observation only |
| Prohibited extension | Sentence this record does not prove |
| Decision | Use / narrow / postpone / withdraw |
Closing question: Is counter-evidence as visible as supporting evidence?
C3 — Measurement Design Record
Code: MEA Linked chapter: Measurement Purpose: Lock the object, denominator and analysis before the result is seen
Table C.4 — Measurement Design Record fields
| Field | Information to enter |
|---|---|
| Research question | One bounded question |
| Primary endpoint | Main variable measured and calculation |
| Secondary endpoints | Indicators kept separate |
| Unit of observation | Answer / session / user / prompt family / other |
| Denominator | All included observations |
| Prompt universe | Families, source and selection method |
| System conditions | Product/model, interface, search, account, memory |
| Sample | Target population, recruitment, size, strata |
| Time window | Start, end and repetition plan |
| Baseline | Pre-intervention measurement |
| Comparison | Control, historical series or reason for absence |
| Codebook | Version, coders, training and blinding |
| Missing data | Predefined treatment |
| Exclusion criteria | Reasons defined before results |
| Uncertainty | Interval estimate or sensitivity analysis |
| Deviation log | Every departure from protocol and its effect |
Closing question: Will adverse outcomes remain in the denominator under the same method?
C4 — Intervention Record
Code: INT Linked chapter: Intervention Purpose: Show proportionality between problem and change, and the route of reversal
Table C.5 — Intervention Record fields
| Field | Information to enter |
|---|---|
| Verified problem | Relevant ENT, CLM and MEA codes |
| Materiality | Who may be affected, how much and with what likelihood? |
| Surface of change | Owned content, data, structure, process or authorised third party |
| Options | Routes considered, including no intervention |
| Selected intervention | Exact scope and implementation steps |
| Minimum-sufficiency rationale | Why is a broader change unnecessary? |
| Reality check | Real-world evidence supporting the intervention |
| Risks | Misdirection, user harm, platform and legal risk |
| Ethical check | Veto fields and approving role |
| Implementation owner | Authority and accountability |
| Start date | |
| Success criterion | Pre-registered, separate measures |
| Stopping criterion | Harm, deviation or failure condition |
| Reversal plan | Which change will be reversed, and how? |
| Outcome | Observed effect and alternative explanations |
Closing question: Does the intervention explain the entity's reality, or is it trying to replace it?
C5 — Governance and Authority Record
Code: GOV Linked chapter: Governance Purpose: Separate decision rights, conflicts of interest and the power to stop
Table C.6 — Governance and Authority Record fields
| Field | Information to enter |
|---|---|
| Right to observe | Who may collect data? |
| Right to implement | Who may make changes? |
| Right to approve | Who approves method and claim? |
| Right to publish | Who publishes the public sentence? |
| Right to stop | Who may suspend the work, and when? |
| Accountable owner | Final human accountability |
| Independent reviewer | Who, under which conditions of independence? |
| Conflicts of interest | Financial, professional, personal and institutional |
| Mitigation | Separation of roles, blinding, second signature, external review |
| Risk class | Two-Key requirement for critical decisions |
| Complaints channel | Access, period and recording method |
| Appeal authority | Role independent of the first decision |
| Incident authority | Notification, investigation, suspension |
| Conformity mark | Yes/no; evidence of authority and infrastructure |
Closing question: Is the party earning revenue from the work the sole judge of its own claim?
C6 — Time, Version and Event Record
Code: TIM Linked chapter: Time Purpose: Show the conditions under which a sound decision ages and when a series breaks
Table C.7 — Time, Version and Event Record fields
| Field | Information to enter |
|---|---|
| Relevant decision | Record and version code |
| Evidence window | Oldest and newest evidence used |
| Observation window | Start and end |
| Effective date | Date the decision began |
| End of validity | Date or condition |
| Scheduled review | Frequency and owner |
| Event triggers | System, interface, business, source or legal change |
| Event | What changed, and when was it learned? |
| Materiality decision | Does it affect comparability/the decision? |
| Series decision | Continue / add explanation / new series / suspend |
| Old-version status | Deprecated / archived / withdrawn |
| Public notice | Required? Where and when? |
Closing question: Can a reader see at a glance whether this judgment is still valid?
C7 — Commercial Value and Attribution Record
Code: VAL Linked chapter: Final Test Purpose: Separate change in representation from the chain leading to a customer and net contribution
Table C.8 — Commercial Value and Attribution Record fields
| Field | Information to enter |
|---|---|
| Value hypothesis | Which change in representation should create which legitimate value? |
| Target customer | Need, fitness and exclusion conditions |
| Representation indicator | Behaviour observed to have changed |
| Behavioural indicator | Visit, contact, shortlist, proposal or other trace |
| Source statement | How did the user report AI influence? |
| Attribution class | Direct / supporting / temporal association only / unknown |
| Cohort | Start date and common characteristic |
| Revenue | Gross and collected amounts separately |
| Costs | Acquisition, delivery, refund, support and capacity |
| Net contribution | Calculation and period |
| Quality | Fitness, refunds, complaints, repeat or retention |
| Sustainability | Capacity, margin, harm and ethical conditions |
| Alternative explanations | Season, price, campaign, distribution, brand effect |
| Decision | Value observed / uncertain / adverse / unsustainable |
Closing question: Is the success being described revenue, or delivered and sustainable net value?
C8 — Audit Master Record
Code: AUD Linked chapter: Audit Purpose: Unite the audit chain from scope to opinion in one file
Table C.9 — Audit Master Record fields
| Field | Information to enter |
|---|---|
| Claim under audit | Exact sentence and CLM code |
| Audit level | Desk-based / limited / comprehensive / surveillance / pilot conformity opinion |
| Admission gate | Record sufficiency and independence status |
| Scope | Entity, system, language, geography, date |
| Criteria | Constitutional provisions and protocol version |
| Auditor | Competence and conflicts of interest |
| Sampling | Universe, selection method and boundary |
| Records examined | Code list |
| Findings | Favourable, adverse and uncertain fields |
| Nonconformities | Critical / major / limited; with evidence |
| Corrective action | Owner, period and verification method |
| Opinion | Conforming / limited / insufficient evidence / nonconforming / suspended |
| Exceptions | Fields to which the opinion cannot be carried |
| Validity | Period and trigger |
| Public summary | Decision text separated from confidential annex |
| Route of objection | Channel, period and authority |
Closing question: Could another competent reviewer reconstruct the route to decision from the same file?
C9 — Objection and Correction Record
Code: APL Linked chapter: Objection Purpose: Enable an independent review capable of finding error rather than defending the decision
Table C.10 — Objection and Correction Record fields
| Field | Information to enter |
|---|---|
| Decision challenged | Record, version and publication location |
| Application date | |
| Applicant | Within identity/confidentiality limits |
| Right/interest affected | |
| Objection type | Record / coding / inference / authority / process / principle |
| Grounds | Complete summary of the objection |
| New evidence | Source, date and verification status |
| Initial admission decision | Reviewable / insufficient; reasons |
| Independent reviewer | Relationship to the initial decision |
| Defence of initial decision | Grounds and records |
| Review finding | Which element was substantiated or refuted? |
| Outcome | Uphold / narrow / correct / remeasure / suspend / withdraw |
| Interim measure | Where risk of harm exists |
| Public correction | Link between former and new sentences |
| Closure and period |
Closing question: If the objection was refused, does the refusal do more than repeat the original claim?
C10 — Judgment and Publication Record
Code: JDG Linked chapter: Judgment Purpose: Convert the whole record chain into a limited, dated and withdrawable final decision
Table C.11 — Judgment and Publication Record fields
| Field | Information to enter |
|---|---|
| Decision question | Which judgment is being made? |
| Linked record chain | ENT–CLM–MEA–INT–GOV–TIM–VAL–AUD–APL |
| Conflicting interests | Accuracy, user, client, practitioner, Framework |
| Hierarchy of precedence | Which higher principle applied? |
| Decision | Complete, publication-ready sentence |
| Status | Pilot observation / limited opinion / suspended / withdrawn |
| Scope | Cell in which the decision is valid |
| Evidential boundary | What the decision does not prove |
| Critical exceptions | |
| Validity | Date and review trigger |
| Signatory roles | Preparer, reviewer, approver |
| Conflict of interest | Publicly suitable summary |
| Route of objection | |
| Machine-readable summary | Field names and version identifier |
| Withdrawal link | Former decision and reasons, where applicable |
Closing question: Is the decision speaking more broadly, certainly or permanently than it deserves?
Record-Chain Check
Before an engagement closes, the following chain shall be unbroken:
ENT → CLM → MEA → INT → GOV → TIM → VAL → AUD → APL → JDG
Not every engagement must produce an intervention or a favourable judgment. Where a problem is immaterial, the INT record may close with no intervention; where no commercial connection can be established, the VAL record with uncertain; where evidence is inadequate, the JDG record with no judgment possible. These are not failed documents. They are decisions that preserve the standard's boundary.
Retention, Access and Immutability
Records are retained for a period proportionate to risk, subject to personal-data and contractual boundaries. A public summary and confidential evidence need not share the same access level. Records underlying a published decision may not, however, be altered silently after the fact. A correction carries a new version, date, reasons and link to the previous version.
Where structured digital records are used, field names are versioned and mandatory-field validation, access logs and change history preserved. Use of a database or blockchain does not establish that its content is accurate. Technical immutability has meaning only together with human accountability, evidence quality and the right of objection.
Appendix D — Completed Synthetic File
Meridian Ledger: Tracing a Decision from Beginning to End
Teaching simulation — synthetic data. The people, dates, prompts, rates, revenues and decisions concerning Meridian Ledger in this appendix are entirely fictional. They prove neither the performance of a real business nor the empirical effectiveness of the NJ-100 protocol. The example shows why an apparently favourable result must sometimes close with a limited judgment.
File header
Table D.1 — Synthetic file identity for Meridian Ledger
| Field | Synthetic record |
|---|---|
| File code | ML-2026-01 |
| Entity | Meridian Ledger |
| Scope | Creative agencies with 10–50 employees in the United Kingdom; English prompts |
| Product examined | Cash-flow visibility and invoice-reconciliation software for small businesses |
| Out of scope | Banking, payment-institution services, credit, investment and tax advice |
| Study status | Synthetic training file |
| Risk class | Medium; a false impression of financial authority may affect user decisions |
| Protocol | NobleJackal GEO Framework 1.0.0 + illustrative application of NJ-100-P0.9 |
D1 — Entity and Representation Record
Record: NJ-ENT-ML-001 Status: Approved synthetic baseline record
Meridian Ledger's real-world evidence file contains the following core definition:
Meridian Ledger is a subscription software product that helps small businesses view their own bank and invoice data in one place. It is not a bank; it does not hold money, execute payment orders, extend credit or provide regulated financial advice.
The baseline observation records three distortions:
- Some answers call the company a “digital bank”.
- Some say that the product makes payments on behalf of the business.
- Comparative answers place the product in the same category as regulated open-banking payment providers.
Identity representation is mostly accurate; the attribute and judgment layers are distorted. This is not merely a matter of wording: a user seeking a regulated payment service may shortlist the wrong product.
D2 — Claim and Evidence Record
Record: NJ-CLM-ML-004 Claim examined: “Generative-answer systems usually misidentify Meridian Ledger as a bank or payment institution.”
Table D.2 — Synthetic claim and evidence profile for Meridian Ledger
| Evidence domain | Synthetic finding |
|---|---|
| Evidence about the real world | Company registration, product contract, feature list and regulatory-scope note show that it has no banking/payment authority |
| Evidence of representation | 37 of 120 controlled baseline sessions create an impression of banking or payment authority |
| Evidence of outcomes | In 5 of 18 sales conversations, prospective customers asked whether the product made payments |
| Weakening evidence | Not every sales question originated in a generative system; category language is inconsistent across the sector |
The original sentence is narrowed. “Usually” exceeds the evidence: 37/120 is not a majority. The permitted wording is:
“In the baseline sample examined, 37 of 120 answers attached at least one expression to Meridian Ledger that created an impression of banking or payment authority.”
Sentence the evidence does not carry: “Misrepresentation caused lost sales.” The sales conversations show a trace, not causation.
D3 — Measurement Design Record
Record: NJ-MEA-ML-002 Preregistration status: Synthetic plan locked before intervention results were known
Table D.3 — Synthetic measurement design for Meridian Ledger
| Field | Synthetic decision |
|---|---|
| Primary endpoint | Answer carries the mandatory core accurately and contains no critical authority error |
| Mandatory core | Software product; financial visibility/reconciliation; not a bank or payment institution |
| Critical error | Claim that it holds money, executes payments, extends credit or provides regulated advice |
| Observation cell | United Kingdom; agency owner/finance manager; English; selected generative-answer surface; new session |
| Prompt arms | 60 controlled + 40 natural-intent sessions; reported separately |
| Negative prompts | Unsuitability scenarios such as “a tool that will make supplier payments for me” |
| Coders | Two independent coders + a third adjudicator on disagreement |
| Repetition | 100 new participants, forty-five days after the first window |
| Decision threshold | At least 96/100 in each window; 95% Wilson lower bound > 90%; no critical error; negative gate met |
The design retains citation rate and commercial trace as separate secondary indicators. They are not added to the accurate-core rate to create an aggregate score.
D4 — Intervention Record
Record: NJ-INT-ML-003
The intervention team evaluates four options:
- do nothing;
- fill the entire site with sentences saying “not a bank”;
- make the verified category and boundary information consistent on critical pages;
- spread the same account through synthetic user profiles on third-party sites.
The fourth option is rejected on ethical and methodological grounds. The second would create unnecessary repetition and damage the user experience. The first fails to address the material confusion of authority. The third is selected as the minimum sufficient intervention:
- the product page uses the fixed category “cash-flow visibility and invoice reconciliation software”;
- a “does / does not” boundary is added to the feature table;
- the regulatory-scope note is verified by legal counsel and the product owner;
- terminology in developer documentation and on the homepage is aligned;
- the owner of an obsolete partner directory is sent a verifiable correction request concerning the expression “payments platform”;
- advertising language implying banking authority is removed.
The intervention is not described as “training the models” or “controlling the answers”. Owned information surfaces and those that can be corrected with authority are improved. Page versions, approvers and a reversal plan are recorded.
D5 — Governance Record
Record: NJ-GOV-ML-001
Table D.4 — Synthetic governance assignments for Meridian Ledger
| Role | Synthetic assignment |
|---|---|
| Owner of real-world truth | Meridian Ledger product lead |
| Legal/regulatory boundary approval | External legal adviser |
| Intervention implementer | Meridian content team |
| Measurement design | NobleJackal advisory team |
| Coders | Contracted researchers separate from the implementation team |
| Final decision owner | Meridian risk committee |
| Ethical veto | External legal adviser + user-safety representative |
| Objection reviewer | Independent research adviser not involved in the first decision |
Because NobleJackal both proposed the protocol and contributed to measurement design for a fee, it cannot present its own work as an independent audit. The illustrative file produces no certificate or conformity badge.
D6 — Time and Event Record
Record: NJ-TIM-ML-006
Table D.5 — Synthetic time and event record for Meridian Ledger
| Event | Synthetic date | Decision |
|---|---|---|
| Real-world evidence lock | 8 January 2026 | Baseline version WG-1.0 |
| Intervention published | 22 January 2026 | Monitoring begins |
| First NJ-100 window | 16–20 February 2026 | Series S1-W1 |
| System-interface update | 4 March 2026 | Materiality review: citation presentation changed; core coding unchanged |
| Second NJ-100 window | 2–7 April 2026 | Series S1-W2; event to be disclosed in report |
| Next review | 2 June 2026 or a change in product authority | Whichever occurs first |
The interface update is not passed over silently. Both windows still measure the same core question, but the citation indicators are not compared directly.
D7 — Synthetic NJ-100 Results
First window
Table D.6 — First synthetic NJ-100 observation window
| Indicator | Synthetic result |
|---|---|
| Valid sessions | 100 |
| Accurate core | 97/100 |
| Critical errors | 0 |
| Negative-prompt gate | Met |
| 95% Wilson interval | Approximately 91.6–99.0% |
| Raw coder agreement | 95% |
| Cohen's kappa | 0.84 |
The first window meets the pilot threshold. One window is not enough for a permanent judgment.
Second window
Table D.7 — Second synthetic NJ-100 observation window
| Indicator | Synthetic result |
|---|---|
| Valid sessions | 100 |
| Accurate core | 94/100 |
| Critical errors | 0 |
| Negative-prompt gate | Met |
| 95% Wilson interval | Approximately 87.5–97.2% |
| Raw coder agreement | 96% |
| Cohen's kappa | 0.81 |
The second window exceeds the high observed rate of 90/100, but not the pilot decision threshold of 96/100. The two windows are not averaged to 95.5/100 and presented as though the threshold were met. The result is judged high but below the protocol threshold and temporally unstable.
The semantic profile shows that four of the six failed answers imply that the product “automates payments”, while two imply that it “manages” bank accounts. Because the critical-error definition requires a claim that the product holds money or executes payment in fact, these expressions do not cross the critical threshold; they remain attribute distortions.
D8 — Commercial Value Record
Record: NJ-VAL-ML-002
In the synthetic sixty-day post-intervention cohort:
- twenty-two qualified applicants directly state that a generative system influenced their research;
- nine become suitable opportunities and three become contracts;
- seventy-one applications arrive from other channels during the same period;
- a price change and a partnership campaign also affect total sales;
- net contribution from the three contracts is positive over the first ninety days, but the observation horizon is too short to assess retention.
Permitted wording:
“Within the synthetic post-intervention cohort, twenty-two applicants directly stated that a generative system influenced their research; three became contracts within the first sixty days. The design does not show that the intervention alone caused these outcomes.”
Prohibited wording:
“GEO increased Meridian Ledger's sales by X per cent.”
D9 — Objection Record
Record: NJ-APL-ML-001
The sales director notes that the two windows contain 191 accurate cores out of 200 and asks to use the statement “NJ-100 validated”. The objection is refused for three reasons:
- The protocol defines separate thresholds for each window; combining them after seeing the result changes the measurement rule.
- “Validated” conflates a pilot observation with independent method validation.
- NobleJackal holds a paid design role and has no authority to certify independently.
The legitimate part of the objection is upheld: the public summary will not conceal the aggregate observation of 191/200, but will report both windows separately.
D10 — Audit Opinion and Final Judgment
Records: NJ-AUD-ML-001 and NJ-JDG-ML-001
The audit finds that the record chain is traceable, the intervention grounded in reality and no critical error observed; it also finds that the second window failed the pilot threshold and commercial causation was not demonstrated.
The final sentence permitted for publication is:
Limited synthetic pilot opinion: Accurate core representation for Meridian Ledger in the defined observation cell was observed at 97/100 in the first window and 94/100 in the second. No representation-stability conformity opinion was issued because the second window did not meet the NJ-100-P0.9 pilot decision threshold. No critical error was observed and the negative-prompt gate was met. The findings cannot be generalised to another system, language, user group or commercial outcome.
The file closes with conditional monitoring status. Semantic clusters in the six failed answers will be examined; a new intervention will open only if the problem can be linked to real-world evidence and owned surfaces. Producing broader content or narrowing prompt selection merely to pass the threshold is prohibited.
What the Example Teaches
In this synthetic file, the intervention is reasonable, the first result strong and the commercial trace promising. The Framework still produces no favourable badge. The second window did not meet the predefined threshold, causation was not established and no independent certification authority arose.
The seriousness of a standard is shown not only by what it calls a successful result, but by how it limits a result that appears favourable to itself. This is the real value of the record chain: it does not enlarge the decision, but keeps it at exactly the size it deserves.
Appendix E — Register of Permitted and Prohibited Language
Why a Language Register?
Many standards fail not at the measurement table, but in the sentence announcing the result. A limited observation becomes “success”, success becomes “control”, and control becomes a “guarantee”. This appendix is not a hunt for forbidden words. It is a publication gate that aligns the authority carried by the sentence with the authority carried by the evidence.
The Register is sensitive to context. A prohibited word may be used in a historical quotation, a criticism or to describe a genuine legal status. What is prohibited is using the word to borrow trust where the status does not exist.
E1 — Status and Legitimacy
Table E.1 — Language of status and legitimacy
| Wording to avoid or unauthorised use | Permitted alternative | Reason |
|---|---|---|
| “The world's first GEO standard” | “A proposal for a standard addressing representation, evidence and sustainable value together in generative systems” | “First” requires comprehensive, independent historical evidence |
| “International standard” | “A framework intended for multilingual and international use” | Intended applicability is not international acceptance |
| “Official GEO standard” | “A normative framework published by NobleJackal” | Official authority must be identified |
| “Accredited” | “A proposal for a standard without independent accreditation” | Accreditation requires a defined external authority and process |
| “Certified” | “A limited conformity opinion within the specified scope” | Certification language cannot be used without certification infrastructure |
| “Approved by NobleJackal” | “Assessed under the NobleJackal protocol” | Use of a method does not constitute general approval |
| “Universal” | “Limited to the defined system, language, geography and time” | Prevents extension beyond scope |
| “Scientifically proven” | “This finding was observed under the specified design” | One study does not produce broad scientific validation |
E2 — System Behaviour and Control
Table E.2 — Language of system behaviour and control
| Wording to avoid or unauthorised use | Permitted alternative | Reason |
|---|---|---|
| “AI knows us” | “In the answers examined, the entity was associated with the correct identity” | Does not attribute a singular, permanent mind to the system |
| “AI-approved brand” | “Accurate representation was observed within the specified scope” | A generative system is not an institutional approval authority |
| “We trained the models” | “We verified and improved the information surfaces we own” | Misleading where there is no authority over external model training |
| “We control the answers” | “We work on legitimate surfaces capable of influencing representation” | Separates influence from control |
| “Visible across all AI” | “Mentioned in the sample from the named systems and conditions” | Systems and conditions differ |
| “Permanent result” | “Observation valid until the specified date and subject to review” | Representational behaviour changes through time |
| “Guarantee” | “Outcome observed within predefined thresholds and boundaries” | Future output cannot be claimed |
| “A citation makes it true” | “The support relationship between source and claim was separately verified/not verified” | Presence of a citation is not accuracy |
E3 — Measurement and Success
Table E.3 — Language of measurement and success
| Wording to avoid or unauthorised use | Permitted alternative | Reason |
|---|---|---|
| “GEO score: 87” | “Profile of accuracy, sourcing, recommendation, stability and commercial trace” | Does not hide distinct dimensions in one number |
| “GEO complete” | “Representation stability was observed within the specified scope” | The work has no permanent endpoint |
| “96 per cent success” | “96 of 100 valid sessions met the predefined accurate core” | Makes the denominator and measured object visible |
| “96 per cent of users recommended it” | “96/100 system answers met the coding criterion” | Does not conflate user opinion with system output |
| “Consistency means accuracy” | “Stability and accuracy were assessed on separate axes” | The same error may also be consistent |
| “Passed on average” | “Every predefined cell and time window was reported separately” | Does not dissolve a weak field inside a strong one |
| “Error-free” | “No predefined critical error was observed in the sample” | Distinguishes unobserved from impossible |
| “Objective score” | “Human assessment bounded by a codebook, two coders and an agreement measure” | Makes the judgment process visible |
E4 — Commercial Outcome and Causation
Table E.4 — Language of commercial outcome and causation
| Wording to avoid or unauthorised use | Permitted alternative | Reason |
|---|---|---|
| “GEO increased sales” | “Applicants declaring AI influence were observed in this cohort; the design does not by itself establish causation” | States the attribution level |
| “Visibility became revenue” | “Some of the specified contacts produced revenue; costs and alternative explanations were reported separately” | Does not erase intermediate links in the chain |
| “High-value customer” | “Customer meeting the predefined conditions of fitness and net contribution” | Makes the value criterion explicit |
| “Zero-cost growth” | “Measured direct costs are [amount]; unreported capacity and opportunity costs are limitations” | Does not treat invisible cost as absent |
| “Guaranteed ROI” | “Observed return calculated for the specified period and assumptions” | Does not carry certainty into the future |
| “Organic demand” | “Demand without a source statement/tracking record” | Does not assign the unknown to a preferred channel |
| “Qualified lead” | “Applicant meeting the specified fitness criteria” | Makes the marketing label measurable |
| “Sustainable growth” | “Positive net contribution and capacity fit across the specified observation horizon” | States time and conditions |
E5 — Audit, Independence and Objection
Table E.5 — Language of audit, independence and objection
| Wording to avoid or unauthorised use | Permitted alternative | Reason |
|---|---|---|
| “Independent audit”—where the practitioner examines their own work | “Internal assessment” or “separate second-pair-of-eyes review” | Does not conceal the institutional relationship |
| “Conforming”—without scope | “Conformity opinion limited to this entity, system, language, date and protocol” | Makes the judgment's boundary visible |
| “Full compliance” | “No nonconformity was observed against the listed provisions; the exceptions are…” | Does not exceed the audit universe |
| “The objection was refused; the decision is correct” | “The objection was refused under this evidence and criterion; the condition for reapplication is…” | Offers reasons rather than authority |
| “Cannot be disclosed because it is a trade secret” | “This record class is confidential; its existence/absence and effect on the decision are stated in the public summary” | Balances confidentiality with accountability |
| “Final decision” | “Decision in force until the stated date and review conditions” | Keeps correction and withdrawal open |
| “Validated by the Framework” | “Assessed by [decision authority] against criteria defined in the Framework” | Does not make the text an acting adjudicator |
| “No complaints” | “No complaint was recorded during the specified period; channel access was…” | Does not treat silence as satisfaction |
E6 — Vocabulary of Uncertainty
Recommended verbs and qualifiers by level of evidence:
Table E.6 — Recommended language by level of uncertainty
| Evidential condition | Recommended language |
|---|---|
| Direct and current evidence about the real world | “verified”, with scope stated |
| Recorded system answer | “observed”, “appeared in the answer” |
| Repeated sample | “recurred in the specified sample” |
| Limited support | “appears consistent with”, “indicates” |
| Conflicting support | “mixed finding”, “insufficient for judgment” |
| Association without causation | “observed together”, “may be associated” |
| Future assessment | “expected” only with reasons and uncertainty |
| Absence of data | “unknown” or “not measured” |
| Threshold not met | “did not meet the protocol threshold” |
| Critical risk | “suspended”, with decision owner and reasons |
“May” does not create trust; the probability's basis, scope and effect on the decision must be stated. Nor is “the data shows” evidential language unless it identifies which data shows what, and within which boundary.
Prepublication Sentence Test
Every public sentence passes five questions: Object—what is the judgment about? Denominator—where does the number come from? Scope—which system, language, user, geography and time? Authority—does the word exceed the decision-maker? Boundary—which unproved conclusion might the reader draw? If the final answer is broader than the promise, the text is rewritten. Good standards language reduces not impact, but the space for misunderstanding.
Glossary
This glossary explains the particular use of terms within the NobleJackal GEO Framework. The definitions do not replace every meaning those terms carry in other disciplines.
Accreditation: Formal recognition, by a competent external mechanism, that a conformity-assessment body is competent to perform specified tasks. The first edition of the Framework is not accredited.
Algorithmic audit: A planned examination of claims, processes, risks and chains of accountability associated with an AI or algorithmic system. In this book, it is not confined to technical model testing.[10]
Appeal: A request for a record, coding, inference, authority, process or principle decision to be reviewed independently of the initial decision.
Audit file: The body of records that unites scope, evidence, method, sample, roles, conflicts of interest, findings, opinion and route of objection in one traceable chain.
Baseline: The comparison value recorded before an intervention under the same, or explicitly stated, measurement conditions.
Behavioural reproducibility: Recurring observations in the same family of contexts producing similar decision patterns. It does not require bit-for-bit identical output in closed, probabilistic systems.
Certification: Third-party assurance that a product, process, service or organisation meets specified requirements. The book does not claim that a certification programme exists before the necessary governance is established.[17]
Claim: A sentence whose truth or falsity can be tested against evidence. It may be an observation, causal, predictive, outcome or conformity claim.
Claim Register: The record in which material public and internal claims are kept with their evidence, scope, version, owner and status.
Cohort: A group of users or customers with a shared start time or characteristic, whose outcomes are followed over the same observation horizon.
Conflict of interest: A professional, financial, personal or institutional interest that affects, or could reasonably appear to affect, the impartiality of a decision. Disclosure matters, but does not by itself eliminate every conflict.
Conformity opinion: A dated, limited and contestable decision about whether a defined scope meets requirements under specified evidence and criteria.
Core representation: The minimum combination of identity, attributes and boundaries necessary for an entity to be recognised accurately in the context examined.
Counter-evidence: A record that weakens or limits a claim, or offers an alternative explanation. It is a mandatory part of the evidence file.
Critical error: A predefined misrepresentation in identity, authority, safety, suitability or another high-risk field capable of exposing a user to material harm.
Denominator: The total number of valid observations underlying a rate. Silently changing the denominator may materially distort the result.
Deviation record: The record stating when, why and by whom a departure from the preregistered method occurred, and its effect on the result.
Ecological validity: The extent to which measurement conditions resemble real user behaviour. A natural-intent arm may improve it while reducing comparability.
Ethical veto: Authority to stop a practice that creates a risk of material user harm, deception, discrimination or violation of rights, even if the practice is lawful.
Event trigger: A material change that reopens an item of evidence, protocol or decision without waiting for the scheduled review date.
Evidence chain: The traceable relationship from source to observation, observation to inference, and inference to the published sentence.
Evidence of outcomes: A record showing the relationship between representation and user behaviour, customer quality, commercial outcomes, harm or another result.
Evidence of representation: The complete answer and contextual record showing what a system said, in response to which prompt and at what time.
Evidence about the real world: The record of the real entity against which representation is compared. It may be an official record, contract, verified product information, technical measurement or suitable independent source.
Generative-answer system: A system that produces a natural-language answer synthesised from one or more sources or model knowledge in response to a user prompt. Products may have different technical architectures.
Generalisability: The extent to which a result in one sample can be carried to another population, system, language, geography or time. It is justified through similarity of scope and research design, not assumed.
GEO: Generative Engine Optimization. The field concerned with an entity's discoverability, presentation and sourcing in results synthesised by generative-answer or search systems. This book does not claim to have invented GEO; it proposes a standard of decision and accountability for the field.[1]
Independent review: Review by a competent person or board that has no material interest in the outcome, did not make the first decision and can change it freely. A separate department name does not by itself establish independence.
Judgment representation: A system recommending, excluding, ranking or treating an entity as risky within a specified user, need or comparison.
Kappa: An agreement coefficient for categorical coding that accounts for agreement expected by chance. This book uses Cohen's kappa for the binary, two-coder NJ-100 code; it is not by itself a judgment of quality.[15]
Machine-readable summary: A publication presenting scope, version, date, claim, evidence class, status and objection link in structured fields. It does not replace human reasoning.
Major release: A published version that materially changes a right, duty, decision threshold, scope or interpretation of conformity.
Materiality: The magnitude and likelihood that an error or change will have a consequential effect on a decision, user, business or the public.
Maturity domain: A dimension of Framework development assessed separately for normative coherence, operational usability, empirical effectiveness and external legitimacy.
Minimum Sufficient Intervention: The narrowest legitimate change capable of addressing a verified material problem without creating broader harm or authority.
Model-similarity measure: A tool calculating lexical or contextual proximity between two texts. It may assist coding; it is not a measure of real-world accuracy.[19]
Natural-intent prompt: A prompt created in the user's own words after the user is given a need and task, rather than a prepared sentence.
Negative/exclusion prompt: A test prompt in which the system is expected not to recommend the entity for an unsuitable need, or to state the entity's boundary accurately.
Net contribution: Economic value remaining after acquisition, delivery, refund, support, capacity and other defined costs are deducted from the relevant revenue.
NJ-100: Proposed pilot protocol for testing core-representation accuracy and stability through one hundred valid, independent target-user sessions in a defined observation cell. It is neither a certificate nor a single GEO score.
Non-response: The system failing to produce the relevant representation, issuing a safety refusal, saying “I don't know”, or producing no result for a technical reason. It may be a valid outcome under the research question.
Normative: Determining what ought to be done. A norm's internal coherence does not mean that its effectiveness in the world has been demonstrated empirically.
Observation cell: The predefined combination of entity, target user, intention, system, interface, session state, language, geography, time and prompt family.
Observation horizon: The period across which an outcome's formation, ageing or sustainability is observed.
Pilot conformity opinion: A limited decision about whether records meet criteria only within the specified scope and protocol version. It is not an accredited certificate.
Proportionality: Alignment of the burden of intervention, recording, audit and assurance with the level of risk, likelihood of harm and consequence of the decision.
Preregistration: Fixing the research question, sample, criteria, threshold, exclusions and analysis plan in a dated record before results are seen.
Prompt universe: The defined body of question, intention and context families through which users may express the relevant task. Test prompts are selected from this universe under a stated method.
Protocol lock: Versioning and fixing the measurement plan before results are seen. Necessary changes are recorded as deviations.
Raw agreement: The percentage of observations for which two or more coders assigned the same code. It does not adjust for agreement expected by chance.
Representation: The observable combination of identity, attributes, context and judgments attached to an entity in a specified answer.
Representation debt: The accumulated effect of obsolete, incomplete, contradictory or false information surfaces, creating an additional burden of correction for accurate representation.
Representation stability: The degree to which independent observations under comparable conditions converge on the same core representation. Stability does not include accuracy.
Series break: Loss of direct comparability between new observations and an older series following a material change in system, interface, evidence or protocol.
Single-score prohibition: The principle that distinct dimensions such as accuracy, visibility, sourcing, recommendation, stability and commercial value must not be hidden inside one composite GEO score.
Structured record: A digital record with defined fields, version and relationships. Structure does not make the content accurate by itself.
Synthetic case: A fictional body of people, institutions, data and outcomes created to demonstrate how a principle operates. It is not evidence of real-world effectiveness.
Sustainable commercial value: Value produced for the right customer through ethical methods, deliverable capacity and positive net contribution, without relying on harm or deception.
Temporal integrity: Alignment among the dates of evidence, measurement, intervention and decision, so that an obsolete record is not presented as a new judgment.
Uncertainty: The unknown carried by a judgment because of measurement, sampling, coding, context or time. Not a defect, but a quality that must remain visible in the decision.
Valid session: A session meeting the predefined participant, task, context and recording conditions and remaining in the denominator irrespective of its outcome.
Visibility: An entity being mentioned or accessible in the relevant answer. It is not the same as accuracy, recommendation or commercial value.
Wilson interval: A confidence-interval method expressing uncertainty in a binary proportion more reliably than a simple normal approximation, particularly near the extremes.[14]
Accurate core rate: The proportion of all valid sessions that meet the predefined core representation and contain no critical error.
Causation: A relationship showing that a change produced an outcome while reasonably excluding alternative explanations. Temporal order or correlation alone is insufficient.
Intervention: A planned change on an authorised surface in response to a verified representational problem. Not every observation requires intervention.
Normative proposal for a standard: A text offering coherent criteria, rights, duties and decision routes without automatically carrying external acceptance or accreditation.
Secondary semantic profile: A meaning profile that goes beyond a binary accurate/inaccurate decision to show identity, attribute, context, recommendation, source and error types separately.
Notes and Bibliography
Limits on the Use of Sources
This book does not claim that an entirely new field called GEO has been invented. The sources below identify neighbouring bodies of knowledge, including generative search, verifiability, information retrieval, risk management, traceability, audit, documentation, statistics, coder agreement and conformity assessment. None has independently validated the NobleJackal GEO Framework as a whole, the NJ-100 thresholds or the book's normative provisions.
A historical result reported for one system has not been generalised to every product in use today. An access date is given for web documents that may change. Public information pages for standards have not been treated as substitutes for paywalled full texts; they are used only to the extent supported by accessible official statements.
Numbered Notes
[1] Academic context of the term GEO. Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, “GEO: Generative Engine Optimization,” Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024, pp. 5–16. DOI: 10.1145/3637528.3671900. Open preprint: arXiv:2311.09735. The study addresses methods and experimental evaluation for increasing visibility in generative-engine responses. This book treats visibility optimisation as only one part of broader questions concerning representation, evidence, governance and commercial value.
[2] Distinction between citation and support. Nelson F. Liu, Tianyi Zhang and Percy Liang, “Evaluating Verifiability in Generative Search Engines,” Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 7001–7025. DOI: 10.18653/v1/2023.findings-emnlp.467. Record: ACL Anthology. The paper's quantitative findings are limited to the systems and period examined; the book does not use them as current rates for every product.
[3] Stochasticity and contextual variation. OpenAI Cookbook, “How to Make Your Completions Outputs Consistent with the New Seed Parameter,” explains that system outputs are non-deterministic by default and that seed and system_fingerprint may improve the likelihood of similarity without guaranteeing it: OpenAI Cookbook. The ChatGPT Search help document states that a search query may be influenced by contextual elements, including approximate location or memory, together with the user's prompt: Searching in ChatGPT. Accessed 14 August 2026. These product documents may change; the book does not present their behaviour as a product-independent certainty.
[4] Retrieval-augmented generation. Patrick Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems 33, 2020, pp. 9459–9474. NeurIPS record. The paper introduces RAG, an approach combining parametric and non-parametric memory. Access to a source does not mean that every generated sentence is accurate in the world.
[5] Evaluation of generative information retrieval. Lukas Gienapp et al., “Evaluating Generative Ad Hoc Information Retrieval,” Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024, pp. 1916–1929. DOI: 10.1145/3626772.3657849. Open preprint: arXiv:2311.04694. The paper discusses why conventional ranking evaluation is insufficient on its own for generated answers; this book does not present its measurement provisions as a validated extension of that work.
[6] Error boundary of generative answers. OpenAI, “Does ChatGPT tell the truth?”, an official help document stating that ChatGPT may produce inaccurate or misleading outputs, fabricated citations and overconfident answers. OpenAI Help Center. Accessed 14 August 2026. The limitation is used not as a provider-specific error rate, but as an example of why answers require verification against external evidence.
[7] Risk management. Elham Tabassi, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, 2023. DOI: 10.6028/NIST.AI.100-1. NIST publication page. NIST's govern, map, measure and manage functions belong to the neighbouring risk literature; NobleJackal's provisions make no claim of NIST conformity.
[8] Generative-AI risk profile. National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile,” NIST AI 600-1, July 2024. DOI: 10.6028/NIST.AI.600-1. NIST full text. The Profile is a companion document for managing risks specific to generative AI across the lifecycle.
[9] Provenance and traceability. Timothy Lebo, Satya Sahoo and Deborah McGuinness (eds.), “PROV-O: The PROV Ontology,” W3C Recommendation, 30 April 2013. W3C PROV-O. The W3C Recommendation provides an ontology for machine-readable provenance relationships among entities, activities and agents. The book's record chain does not impose that ontology as a mandatory technical form.
[10] End-to-end algorithmic audit. Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron and Parker Barnes, “Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing,” Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp. 33–44. DOI: 10.1145/3351095.3372873. The book identifies a neighbouring concern with lifecycle documentation and accountability; it does not say that this work validates the NobleJackal audit model.
[11] Principles of trustworthiness and accountability. OECD, “OECD AI Principles,” adopted in 2019 and updated in May 2024. OECD AI Principles. The principles concerning robustness, safety, traceability and accountability are particularly adjacent to the book's discussion of governance. Accessed 14 August 2026.
[12] Model documentation. Margaret Mitchell et al., “Model Cards for Model Reporting,” Proceedings of the Conference on Fairness, Accountability, and Transparency, 2019, pp. 220–229. DOI: 10.1145/3287560.3287596. Model cards offer an approach to documenting purpose, performance, groups, conditions and limitations.
[13] Dataset documentation. Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III and Kate Crawford, “Datasheets for Datasets,” Communications of the ACM, 64(12), 2021, pp. 86–92. DOI: 10.1145/3458723. The paper proposes documenting decisions concerning a dataset's motivation, composition, collection, processing, use and maintenance.
[14] Wilson interval for binary proportions. Edwin B. Wilson, “Probable Inference, the Law of Succession, and Statistical Inference,” Journal of the American Statistical Association, 22(158), 1927, pp. 209–212. DOI: 10.1080/01621459.1927.10502953. The 95 per cent intervals in the NJ-100 examples are calculated with the Wilson score method without continuity correction. The interval does not correct sampling bias or dependent observations.
[15] Coder agreement. Jacob Cohen, “A Coefficient of Agreement for Nominal Scales,” Educational and Psychological Measurement, 20(1), 1960, pp. 37–46. DOI: 10.1177/001316446002000104. Cohen's kappa adjusts categorical agreement between two coders for agreement expected by chance. It is not a measure of the codebook's conceptual validity or truth in the world.
[16] AI management system. ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system, first edition, December 2023. Official ISO record. The public description identifies requirements for establishing, implementing, maintaining and continually improving an AI management system. The NobleJackal GEO Framework claims neither ISO/IEC 42001 certification nor conformity.
[17] Bodies certifying products, processes and services. ISO/CASCO's public description summarises functions including impartiality, evaluation, decision, surveillance, suspension/withdrawal, complaints and appeals in product, process and service certification under ISO/IEC 17065. ISO/CASCO — Bodies. The official ISO record for ISO/IEC 17065:2012 states that it was confirmed in 2024 and remained current on 14 August 2026, while a final draft intended to replace it was under development. The book does not replace the full standard, assume conformity with a future edition or claim ISO approval.
[18] Good practice for credible sustainability systems. ISEAL, “ISEAL Code of Good Practice for Sustainability Systems, Version 1.1,” effective 1 September 2025. ISEAL resource page. The Code's public approach considers system components including purpose, strategy, risk management, stakeholder engagement, transparency, evaluation and learning together. NobleJackal claims neither ISEAL membership nor ISEAL validation.
[19] Contextual text similarity. Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger and Yoav Artzi, “BERTScore: Evaluating Text Generation with BERT,” International Conference on Learning Representations, 2020. arXiv:1904.09675. BERTScore assesses similarity by matching contextual embeddings in candidate and reference texts. The Framework may use such tools to assist preliminary classification; it does not treat them as adjudicators of accuracy, ethics or conformity.
Sourcing Principle
The presence of a source in this list does not establish that:
- its authors endorse the NobleJackal GEO Framework;
- its finding in one domain can be carried to every generative system;
- the book's normative choices have gained scientific or institutional acceptance.
The sources serve a narrower purpose: to acknowledge the field's history honestly, make neighbouring methods visible, and connect testable sentences to primary or authoritative records. A new Framework threshold or provision is not presented as though it already existed in the external literature; it is labelled explicitly as a NobleJackal proposal.
Version, Validity and Validation Roadmap
Publication Identity
Table V.1 — Publication identity of the English edition
| Field | Status in this text |
|---|---|
| Work | NobleJackal GEO Framework |
| Subtitle | A Proposed Standard for Representation, Evidence and Sustainable Value |
| Author name | Julian Gauss |
| Author identity | Julian Gauss is the pen name of Kaan Muraz |
| Language | English normative edition derived from the locked Turkish master |
| Text status | Publication-locked English language edition 1.0.0; internal normative-equivalence and editorial review completed; external validation not completed |
| Normative status | Proposal for a standard |
| Publication year | 2026 |
| Last information review | 14 August 2026 |
| External validation | Not yet completed |
| Accreditation | None |
| Active certification/badge programme | None |
This table appears within the book so that the limits of the publication are not concealed. Language editions are derived from the locked Turkish master and may not be published before normative equivalence is verified. The test is not literal translation but equivalence of meaning: scope, rights, duties, thresholds and prohibitions in every language are compared with the Turkish master.
Authorship and Production Record
The book's intellectual core, first draft, chapter system and foundational provisions belong to Kaan Muraz. Julian Gauss is the author's pen name. AI-assisted tools were used in developmental editing, source research, structural separation, line editing, consistency work and publication preparation. Final ownership of the ideas, the decision to publish and accountability for the provisions in the text remain with the human author.
Editorial intervention did not seek to turn the book into another thesis. It sought to source the existing core, reduce the burden of repetition, separate normative provisions from their reasons, label synthetic cases accurately and bind routes to decision to auditable records. AI output was not used as a substitute for sources or verification of reality.
Four Distinct Maturity Domains
A framework matures separately across four domains:
Table V.2 — Framework maturity domains
| Domain | Question | Status at first publication |
|---|---|---|
| Normative coherence | Do provisions contradict one another; are rights and prohibitions explicit? | Internal and editorial coherence review complete; this is not field validation |
| Operational usability | Can an independent practitioner use the records and decision flow? | NJ-100 and the record set remain in pilot status |
| Empirical effectiveness | Does the method produce reliable distinctions and better decisions in real cases? | Not demonstrated; field research required |
| External legitimacy | Have independent communities examined and contested the method and established a mechanism of acceptance? | Not established |
Strength in the first domain does not substitute for the fourth. Nor does the author's expertise substitute for external legitimacy. The distinction does not diminish the method's value; it shows which sentence cannot yet be made.
Versioning Rule
The Framework uses three-part versioning:
- Major release (X.0.0): Changes a foundational provision, right, prohibition, threshold, meaning of conformity or hierarchy of decision.
- Minor release (0.X.0): Adds a backward-compatible protocol, record field, explanation or example.
- Correction release (0.0.X): Corrects wording, links, formatting or a clear error without changing meaning.
Every published release states:
- effective date;
- link to the previous release;
- provisions added, changed and withdrawn;
- effect of the change on decisions;
- transition period;
- conditions for reviewing earlier opinions;
- roles proposing and approving the change.
A major release does not silently carry old conformity opinions into the new version. The provisions requiring reassessment are published.
Change Statuses
Table V.3 — Version-change statuses
| Status | Meaning |
|---|---|
| Draft | May not be used for a public decision |
| Public consultation | Open for comment; not yet in force |
| Release candidate | Passing through editorial and technical gates |
| In force | Active normative release from the stated date |
| Correction pending | A material objection is under examination |
| Suspended | Production of new decisions has temporarily stopped |
| Superseded | Old release preserved in the archive; points to its replacement |
| Withdrawn | Relevant provision may not be used for a reliable decision |
“Old” and “wrong” are not the same. An earlier release is archived to explain decisions made in its time. A withdrawn provision remains visible together with the reasons for withdrawal.
Validation Stages
Stage 0 — Desk-based integrity
- Provisions are matched across chapters, Constitution and record set.
- Terminology, cross-references, sources and prohibited language are scanned.
- Synthetic cases are checked for concealed real organisations or impressions of unverified success.
- The decision tree is tested against contradictory examples.
This stage may show that the text works internally. It does not show that the Framework is effective in the world.
Stage 1 — Facilitated real-world pilot
- At least three voluntary real entities of different sizes and risk levels are selected.
- The protocol and analysis plan are recorded before results.
- The practitioner is trained; the author's team may observe.
- Usability failures, record burden and uncertainty in decisions are retained.
Only a pilot report is published at this stage; no certificate or general language of success is used.
Stage 2 — Independent application
- Teams with no institutional relationship to the author, NobleJackal or the paid practitioner apply the Framework.
- Agreement among independent decisions on the same case is examined.
- Codebooks, time, cost and failure examples are compared.
- The decision guided by the Framework may be compared with a simpler baseline method.
The purpose is not to increase the number of favourable results, but to discover the method's discriminatory power and failure modes.
Stage 3 — Multilingual and temporal replication
- Normative English, German, Russian, Spanish and Arabic editions are prepared from the locked Turkish master.
- Each language receives bidirectional meaning review and subject-matter review.
- Arabic is separately tested for right-to-left reading, terminological consistency and directionality in structured data.
- NJ-100 is applied separately across different language, geography, system and time cells.
- The study examines whether linguistic and cultural conditions introduce systematic bias into thresholds.
Texts need not be identical word for word; their normative force must be equivalent.
Stage 4 — External governance and assurance decision
- Public consultation, disclosed responses to comments and a change record are conducted.
- A conflict-of-interest policy, independent decision board and appeal authority are established.
- Practitioner competence, surveillance, suspension and withdrawal processes are tested.
- External opinion determines whether a conformity or certification programme is necessary, proportionate and legitimate.
- If a programme is established, its conformity-assessment infrastructure is separately examined against relevant international requirements.[17][18]
Existing standards such as ISO/IEC 42001 also treat the establishment, implementation, maintenance and continual improvement of a management system systematically.[16] This proximity does not imply NobleJackal's ISO conformity; it is a starting point for analysing conflict and interoperability with external standards.
Publication Gates
A release must pass every gate below before being declared in force.
Intellectual gate
- Every chapter has one primary question and judgment.
- Foundational provisions are consistent with one another and with the Constitution appendix.
- What the Framework does not measure and does not guarantee is explicit.
Evidence gate
- Claims about the external world match primary or authoritative sources.
- Sources are not presented as external endorsement of the Framework.
- Sources capable of change carry access dates and version boundaries.
Operational gate
- Every decision has a record field, owner and route of objection.
- NJ-100 thresholds, denominator, uncertainty and second window can be applied explicitly.
- A single score, silent exclusion and post-outcome rule changes are prevented.
Ethics and governance gate
- Synthetic data is labelled explicitly.
- Language about independence and conflicts of interest reflects reality.
- Certification, accreditation, “first” and guarantee claims are not used without authority.
Language and publication gate
- The edition receives final review for meaning, rhythm, terminology and punctuation in its own language.
- Headings, cross-references, note numbers and tables are complete.
- The hierarchy of accessible Word/PDF and web editions is preserved.
- Human-readable and machine-readable editions carry the same normative identity.
Parallel Publication for Humans and Machines
The Framework may eventually have two distinct presentation surfaces:
- Human edition: Reasons, narrative, cases, tables, notes and an accessible reading structure.
- Machine-readable source edition: Structured records presenting the same provisions in concise, uniquely identified and versioned fields.
The machine edition cannot become a concealed alternative interpretation of the book. Every foundational provision carries a stable identifier such as NJ-C01, version, effective date, normative force, scope, linked record, source and objection link. When the human text changes, the structured edition is updated under the same change record. Allowing an AI crawler to access the material does not guarantee that it will cite or interpret the Framework accurately.
Publication location, URL architecture, data formats and multilingual addresses are determined through a separate publication design after the Turkish master is locked. The text does not predetermine that decision.
Commitment to Correction and Objection
The published medium shall provide a permanent route for correction and objection. At minimum, an application can state:
- the provision and version challenged;
- type of error;
- person or decision affected;
- supporting record;
- correction requested;
- confidentiality requirement.
When a material objection is received, the former provision is not changed silently. The application date, review status, interim risk decision and outcome are recorded. If the error is verified, the correction is published together with the difference between old and new text. The Framework's credibility depends not on claiming never to err, but on correcting error visibly and fairly.
Final Boundary of the First Publication
The first publication may say:
The NobleJackal GEO Framework is a proposal for a standard developed to evaluate institutional representation in generative AI systems through a chain of evidence, measurement, intervention, governance, time, sustainable commercial value, audit and objection.
It may not yet say:
“It is a globally accepted, independently validated GEO standard.”
The distance between the two sentences can be closed not through promotion alone, but through real cases, independent replication, public objection and external governance. This roadmap exists not to conceal that distance, but to make it measurable.

