NOMOS GEO-QA / English edition

NGQ-062 / Measurement, audit and conformance

How should variability, uncertainty, sample size and confidence intervals be reported for probabilistic AI outputs?

Short answerOne answer is one observation; a conclusion about the system requires repetitions, distribution, scope and uncertainty together.
VERSION
0.3.1
STATUS
founder edition · released
PRIMARY SOURCES
4

Direct answer

A probabilistic system may give different answers to the same or similar question at different times. Report the number of repetitions, execution conditions, result distribution, error types and uncovered cases. Sample size is not a fixed number chosen for marketing; it depends on the decision, error difference to be detected, expected variability and accepted uncertainty. Use a confidence interval or another statistical summary only when the method's assumptions are met.

In plain language

You cannot tell whether a die is fair from one throw. Rolling a six does not mean it will always roll six. Throw it many times, examine the distribution and state how uncertain you still are.

Why this matters

One good screenshot does not prove success, and one bad answer does not prove constant failure. If variability and uncertainty are hidden, large conclusions are drawn from small samples and comparisons become misleading.

Do not confuse

  • A model confidence signal is the model's own output; it is not a confidence interval for the measurement result.
  • Measurement uncertainty concerns how well the observed result estimates the underlying rate.
  • Sampling error may arise because the examined sample does not represent the target population fully.
  • System variation is change in outputs under the same conditions or over time.
  • Evidence uncertainty means that the Truth Pack itself is incomplete or conflicting.

What should you do?

  1. Define the target population, decision threshold and error difference you need to detect.
  2. Run pilot repetitions under the same protocol to collect initial error-rate and variability data.
  3. Justify sample size through desired precision, risk, variability and available resources.
  4. Keep question, language, country, model, version, date and repeat conditions fixed or explicitly stratified.
  5. Report counts, rates, distributions, outliers and critical-error counts rather than an average alone.
  6. When assumptions are suitable, report a confidence interval with its method and level; otherwise limit the result to a descriptive account.
  7. State unrepresented groups, failed runs and limits on generalisation.

How do you audit it?

  • Was one output presented as the system's continuing behaviour?
  • Can the number of repetitions and execution conditions be seen?
  • Was sample size justified by the decision purpose and expected variability?
  • Did the average conceal critical errors or subgroup differences?
  • Was the confidence interval calculated with a method appropriate to the data and assumptions?
  • Were model confidence and measurement confidence confused?
  • Was the result overgeneralised to unobserved languages, countries, models or times?

Limit

NOMOS GEO does not prescribe one minimum number of repetitions or one sample size for every study. Independence, distribution and system-stability assumptions may not be met fully in third-party generative systems; reduce the strength of the statistical claim and seek specialist methodological review where appropriate.

Remember in one sentence

One answer is a sample; state the result together with the sample's limits.

Sources for this record

CITATION RECORD

Muraz, K. (2026). NOMOS GEO-QA: Canonical Question Registry (English Edition, v0.3.1). NobleJackal. https://noblejackal.com/nomos-geo-qa/
© 2026 Kaan MURAZ. All rights reserved.