Direct answer
A probabilistic system may give different answers to the same or similar question at different times. Report the number of repetitions, execution conditions, result distribution, error types and uncovered cases. Sample size is not a fixed number chosen for marketing; it depends on the decision, error difference to be detected, expected variability and accepted uncertainty. Use a confidence interval or another statistical summary only when the method's assumptions are met.
In plain language
You cannot tell whether a die is fair from one throw. Rolling a six does not mean it will always roll six. Throw it many times, examine the distribution and state how uncertain you still are.
Why this matters
One good screenshot does not prove success, and one bad answer does not prove constant failure. If variability and uncertainty are hidden, large conclusions are drawn from small samples and comparisons become misleading.
Do not confuse
- A model confidence signal is the model's own output; it is not a confidence interval for the measurement result.
- Measurement uncertainty concerns how well the observed result estimates the underlying rate.
- Sampling error may arise because the examined sample does not represent the target population fully.
- System variation is change in outputs under the same conditions or over time.
- Evidence uncertainty means that the Truth Pack itself is incomplete or conflicting.
What should you do?
- Define the target population, decision threshold and error difference you need to detect.
- Run pilot repetitions under the same protocol to collect initial error-rate and variability data.
- Justify sample size through desired precision, risk, variability and available resources.
- Keep question, language, country, model, version, date and repeat conditions fixed or explicitly stratified.
- Report counts, rates, distributions, outliers and critical-error counts rather than an average alone.
- When assumptions are suitable, report a confidence interval with its method and level; otherwise limit the result to a descriptive account.
- State unrepresented groups, failed runs and limits on generalisation.
How do you audit it?
- Was one output presented as the system's continuing behaviour?
- Can the number of repetitions and execution conditions be seen?
- Was sample size justified by the decision purpose and expected variability?
- Did the average conceal critical errors or subgroup differences?
- Was the confidence interval calculated with a method appropriate to the data and assumptions?
- Were model confidence and measurement confidence confused?
- Was the result overgeneralised to unobserved languages, countries, models or times?
Limit
NOMOS GEO does not prescribe one minimum number of repetitions or one sample size for every study. Independence, distribution and system-stability assumptions may not be met fully in third-party generative systems; reduce the strength of the statistical claim and seek specialist methodological review where appropriate.
Remember in one sentence
One answer is a sample; state the result together with the sample's limits.
Sources for this record
- S38NIST AI 600-1, *Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile*Voluntary standards-oriented institutional guidance
- S44NIST, *AI Risk Management Framework Playbook* and AI RMF CoreVoluntary implementation guidance
- S50NIST AI 100-1, *Artificial Intelligence Risk Management Framework (AI RMF 1.0)*Voluntary standards-oriented institutional framework
- S51NIST/SEMATECH, *e-Handbook of Statistical Methods*Official authority technical handbook

