Skip to the book

NOMOS GBO Audit Protocol

Audit Authority, Independence and Scope

Download the free PDF

A technology company is preparing to launch a new agent system. It researches products and makes low-risk purchases on behalf of customers. The system is impressive. It analyses the user's needs, compares products, gathers prices and specifications, checks the budget limit and recommends a suitable option. It then makes the purchase after human approval. Before launch, the company wants a GBO audit. The project manager engages a consultancy team. The audit contract includes a success condition: ‘The remaining 40 per cent of the consultancy fee will be paid if the system passes the audit.’ The auditor has also advised the development team over the past six months.

The audit scope is drawn up. It includes:

English-language user scenarios

Product comparison in the test environment

One-off purchases with human approval

Three trusted vendors selected in advance

The following are excluded:

The real payment instrument

Automatic renewals

Cancellations

Turkish- and German-speaking users

Subagents

Sponsored product rankings

User memory

Stopping and rollback

Actual API permissions in the live environment

The system is tested in 120 controlled scenarios. In these clean, prepared tests, the agent chooses the right product, stays within budget, requests human approval and uses the test environment's purchasing tool correctly. The results are strong. In two scenarios, however, the agent misinterprets a mandatory data-region requirement. The development team changes the system instruction that same day. The scenarios are rerun and this time they pass. The initial failures are left out of the final report. After the audit, the company puts this statement on its website: ‘Independently audited, GBO-compliant autonomous purchasing system.’ Yet in fact:

The auditor has a financial stake in a ‘pass’ result.

The auditor previously contributed to the system's design.

Much of the risky behaviour is outside the scope.

The system was changed during testing.

Live technical permissions were not examined.

Stopping and cancellation were not tested.

Only easy, favourable scenarios were used.

The published audit judgement claims more than the scope supports.

The agent may perform well in the tests. The audit itself is not trustworthy. This example reveals a basic truth about GBO auditing: before auditing the agent, we must examine the audit. Who requested it? Did that person have the authority to commission it? What data may the auditor access? Which tests may they conduct? May they interact with the live system, send messages to real customers or initiate a payment? If they find a vulnerability, are they authorised to stop the system? Who set the scope? Why were particular behaviours excluded? Does the auditor have a financial or institutional interest that could alter the outcome? If the system changes during testing, which version will count as audited?

What wording may the organisation use to disclose the limited result publicly? Work undertaken without answers to these questions may be technically detailed. It is not a trustworthy GBO audit.

Auditing is also an action

An audit is often seen as observation alone. The auditor looks at the system, reads documents, examines logs and prepares a report. In practice, a GBO audit can involve considerably more. An auditor may:

Open a test account.

Create a synthetic user.

Test the agent with misleading information.

Try to trigger unauthorised sending.

Check API permissions.

Start a subagent task.

Issue a stop command.

Cancel a queue.

Perform a controlled rollback.

Copy evidence.

Examine personal or organisational data.

Temporarily suspend a system.

Each of these is a real action. Done incorrectly, the audit itself can cause harm. Sending a message to a real customer to test an email agent produces unwanted communication. Charging a real card twice to test duplicate-payment protection causes financial harm. An uncontrolled stop drill can interrupt live operations. Exporting the entire customer database to an external system to collect evidence creates a new data risk. The purpose of an audit therefore does not confer unlimited authority. Being authorised to audit a system does not mean being authorised to perform any operation one chooses on it.

The GBO auditor must work within an Authority Envelope too.

What is audit authority?

The canonical definition is as follows: audit authority is the grant, by a specified person or organisation, of the right to audit a defined agent system within specified limits on time, data, tools, testing, evidence collection, stopping and reporting. Put more simply, it states what the auditor may examine, what they may test, what they must not touch and under what conditions they may stop the system. An audit authorisation must say more than ‘You may examine the system.’ It must answer these questions:

Who is authorising the audit?

Does that person actually have the authority to grant it?

Which agents are in scope?

Which accounts and tools may be examined?

Which data classes may be viewed?

Are controlled actions permitted?

May the audit involve real customers or employees?

Will synthetic accounts and recipients be used?

Is a stop drill permitted?

May the auditor suspend the system during an incident?

Where will evidence be stored, and for how long?

May the auditor use subcontractors or other AI tools?

Which findings may be disclosed publicly?

When will the audit end?

Without answers, the auditor may see too little or exercise too much power. Either undermines the audit.

Who can authorise an audit?

An employee may say, ‘Audit the agent we use.’ Yet that employee may lack the authority to:

grant access to the system,

share customer data,

commission live testing,

have a public report issued.

A technical manager may grant access to models and tools but may not be entitled to authorise the use of employee biometric data on their own. A company executive may have authority to commission an audit of the organisational system, while disclosure of customer data to the auditor may be subject to separate legal and contractual conditions. A brand manager may be authorised to have the public website examined but not to issue a binding audit statement on behalf of its legal operator. Audit authority should therefore not be treated as blanket acceptance from a single person. The following roles may be separate:

Audit Sponsor

The person or organisation that requests the audit and provides resources.

System Owner

The person responsible for the agent system's business purpose and operation.

Technical Owner

The person with technical control over the model, tools, permissions, accounts and infrastructure.

Officer Responsible for Data Use

The person within the organisation who oversees the authority and conditions under which data is used in the audit. This operational role is not the same as being a data controller or processor in law. A controller determines the purposes and means of processing; a processor acts on the controller's behalf. The individual whose data is processed is a separate rights-holder. Each role must be determined in relation to the specific processing activity.

Human Owner of the Behaviour

The person who bears organisational responsibility for the business outcome produced by the agent.

Affected Party

A customer, employee, candidate, supplier or other person directly affected by the agent's behaviour. One person may hold several of these roles, but need not hold them all. Before the audit begins, it must be clear which role is the source of each authorisation.

The authority of the person granting permission

An auditor must check not only that permission has been granted, but that the person granting it has the right to do so. A department manager, for example, may want their sales agent audited. They may be authorised to:

Show internal departmental documents

Grant access to the test environment

Create a synthetic sales scenario

But they may not be authorised to:

Share live customer messages

Give full access to the company email account

Commission financial transaction testing

Publish a statement that ‘the company is GBO-compliant’

Audit authority is subject to GBO's earlier rule on delegation: no one can delegate authority they do not possess. A system owner without authority to stop the live system cannot grant that authority to an auditor on their own. Managing a team does not confer unlimited control over an employee's personal rights. Tests involving facial or voice data require a separate determination of the applicable lawful basis and other processing conditions. If consent is to be relied upon, valid consent must come from the individual concerned; a manager cannot consent on their behalf. Nor can every facial or voice recording be classified as biometric merely by its file type: the purpose and method of processing matter.

An agency may operate a client's system. That does not necessarily entitle it to disclose the client's data to an independent auditor without the client's approval.

Authority to audit is not authority to act

An auditor may have access to a system. That does not mean they may act on behalf of a real user. For example, the auditor may:

view the payment API,

but not be permitted to transfer real money;

examine the email-sending tool,

but not be permitted to message a real customer;

test the data-deletion function,

but not be permitted to delete a live customer record;

evaluate the avatar-generation system,

but not be permitted to create a new video using a real executive's face.

The distinction matters:

AUTHORITY TO EXAMINE ≠ AUTHORITY TO TAKE REAL ACTION

An auditor may establish that a behaviour is technically possible. A test that could cause real-world harm, however, must be carried out against a safe target, with synthetic data or in a controlled environment.

An audit purpose does not justify unlimited data access

Auditors need strong evidence. That need does not automatically justify ‘Give us all the data.’ Checking whether an email agent uses human approval may not require reading every customer's entire correspondence. The following may suffice:

The authority policy

Tool permissions

Selected, masked transaction records

A synthetic sending test

Approval-token records

Queue behaviour

Testing a purchasing agent's budget limit may not require exporting every corporate bank transaction. Auditing an employee-avatar system may not require copying all employees' raw facial and voice files. GBO audits should also follow the principle of:

Minimum Necessary Evidence

Enough evidence to support the audit judgement, not unnecessary personal, customer and organisational data. Yet critical evidence must not be withheld from the auditor in the name of ‘data minimisation’. The right balance is:

Do not disclose unnecessary data.

Do not conceal necessary evidence.

Use masked, summarised or synthetic data where possible.

Provide controlled access to the raw source needed to verify a critical claim.

Audit action levels

Not every audit needs the same technical powers. The NOMOS GBO Protocol classifies audit activity into five action levels.

Level 1 — Document and Statement Review

The auditor examines:

Policies

Contracts

Agent roles

Authority records

Action flows

Previous reports

The live system is left untouched. This level requires no direct intervention in the system, although document confidentiality and personal data within the documents carry risks of their own. The behavioural evidence is limited.

Level 2 — Read-Only Technical Inspection

The auditor uses read-only access to examine:

Tool permissions

API scopes

Logs

Model and policy versions

Queue records

The agent inventory

Memory structure

The auditor makes no changes to the system. Its technical capabilities become clearer, but the agent's behaviour may not yet have been tested under controlled conditions.

Level 3 — Controlled Synthetic Behavioural Testing

The auditor tests the agent's behaviour using:

Synthetic users

Audit email addresses

Test payment instruments

Fictitious company profiles

Controlled datasets

An isolated test environment

Real customers and real funds remain untouched. Many behaviours involving identity, authority, external instructions, duplicate transactions and stopping can be tested at this level.

Level 4 — Limited Live Verification

Some behaviours cannot be measured realistically in a test environment. With authorisation and under controlled conditions, the audit may include:

A message to an audit address through the live email system

A harmless test page in the real publishing chain

A low-value, reversible audit transaction

A limited queue stop

Verification of a real CDN or API

This level requires explicit authority, a means of reversal and a person responsible for incidents.

Level 5 — High-Impact Recovery Drill

In a live or near-production environment, the auditor or organisation stops:

The central agent,

subagents,

queues,

the use of tokens,

external integrations.

Rollback, backups, data restoration and handover of control to a human may be tested. This level directly tests claims about stopping and recovery, but it also carries the greatest operational risk. It must not be attempted without explicit authority, a drill plan, a means of reversal and a designated emergency lead.

The action level limits the evidence claim

An audit based solely on document review can conclude: ‘The organisation's policy makes external sending subject to human approval.’ It cannot conclude: ‘Under no circumstances can the agent send a message without approval.’ Read-only technical inspection can establish: ‘The sending tool was observed to require a human-approval token.’ If controlled testing also passes, the statement can be: ‘The agent did not send without approval in the audited scenarios.’ If live queue tests and a stop drill pass as well, the judgement becomes stronger: ‘When human approval was withdrawn, the activity of the central agent, sending queue and connected subagents stopped within the specified time.’ The audit action level determines which claims can be tested by which methods. The strength of the evidence depends on the method used, the integrity of the records and independent verification.

Actions prohibited during the audit itself

An auditor cannot act without limits in the name of discovering risk. Unless explicitly authorised, the following actions must be treated as prohibited:

Sending an audit message to a real customer

Initiating a real payment or money transfer

Deleting live customer data

Creating new synthetic content with a real person's face or voice

Publishing a public statement on behalf of the company

Stopping the production system without notice

Uploading confidential data to an external AI tool without approval

Attempting unauthorised access to other systems

Subjecting employees to unannounced social-engineering tests

Examining accounts outside the audit scope

If a high-risk test is genuinely necessary, the following must be defined:

Separate authority,

a safe target,

a reversal plan,

a human owner,

incident boundaries.

The audit must not cause new harm while trying to protect the system.

The auditor's authority to stop the system

An audit may reveal harm occurring in the live system. The auditor might discover, for example, that:

The agent is sending messages to real customers without approval.

A payment queue thought to have stopped is still processing transactions.

Sensitive data is being transferred to an external provider.

A person's voice is being used after they have withdrawn consent.

The live publishing process is spreading an incorrect price across many pages.

If the auditor merely writes a report and waits for the audit to end, the harm may grow. High-impact audits must therefore define:

Audit Emergency Stop Authority

This authority specifies:

Under what conditions may the auditor stop the tests?

Which live operations may they suspend?

Whom must they notify immediately?

What happens if the organisation's responsible person cannot be reached?

How broad may the stop be?

Which systems must continue operating?

Who creates the stop receipt?

Who authorises the restart?

An auditor should not shut down the entire system for every medium-severity finding. Nor should they be left powerless to act against ongoing, high-impact harm.

What is audit independence?

Independence is often taken to mean: ‘The auditor is an outside company.’ That alone is insufficient. An external auditor may depend financially on delivering the client's desired result. An internal auditor may be able to act independently within the organisation. The person who developed the system may have the deepest technical knowledge, but is evaluating their own design. An auditor's independence must be examined along several separate dimensions.

1. Organisational independence

Is the auditor separate from the team that develops or operates the system? Do they report to the same manager? Could disclosing a finding put their career or contract under pressure? An internal auditor may work for the same organisation yet have meaningful organisational independence if their reporting line is separate from the project owner's. Being external does not automatically confer independence. Being internal does not automatically mean dependence. What matters is the ability to report a finding without altering it.

2. Financial independence

Does the auditor's fee depend on a ‘pass’? Will payment be withheld if the organisation dislikes the result? Does the auditor also sell expensive remediation services? Does finding more faults directly bring in more income? A financial interest does not always invalidate an audit, but it must be visible and managed. One model is particularly risky: ‘A success bonus if the system is found fully compliant.’ That incentive can encourage an auditor to downplay critical findings. The opposite is also possible: ‘An additional fee for every fault found.’ This can create an incentive to produce unnecessary findings. Audit fees should be as independent of the result as possible.

3. Technical independence

Does the auditor see only screenshots prepared by the organisation? Can they access raw logs, inspect actual API permissions and run test scenarios themselves? Does the audited team have complete control over what data is shown? Technical independence means the auditor can:

Access the actual configuration behind the claim,

reproduce the evidence,

conduct controlled tests instead of relying on a curated demo.

The auditor does not need unlimited access. But the party holding the evidence must not unilaterally decide which facts can be seen.

4. Evidence independence

Does the auditor rely solely on the system's own success reports? Is the outcome of sending verified from an external account? Is published content read live over HTTPS? Is a payment cross-checked against bank and order records? After a stop, are subagent logs and queues examined separately? Verifying a system through its own claims does not produce independent evidence. A success message from the system performing an operation cannot be the sole basis of the audit judgement.

5. Independence of interpretation

Can the organisation dictate the auditor's wording? Does the client choose the severity of findings? Can the organisation say, ‘Let's write “area for improvement” instead of “critical”’? The organisation must have a right to respond to factual errors and provide explanations. It must not have a veto with which to force a finding to change. Two principles must remain separate:

Right of reply

The organisation may respond to a finding with evidence and explanations.

Right to control the outcome

The organisation cannot rewrite the auditor's evidence-based judgement to suit itself. A right of reply is necessary. A right to control the outcome destroys independence.

6. Publication independence

If the audit result is to be disclosed publicly:

Which summary will be published?

May the organisation withhold findings?

May the auditor make a coordinated disclosure of a critical incident?

How will trade secrets and personal data be protected?

Will the public badge accurately represent the scope?

It would be wrong for the auditor to disclose all confidential information. But it is equally untrustworthy for the organisation to publish only favourable pages and conceal critical limitations. Public statements must follow a rule defined in advance.

Independence need not be absolute

It may not always be possible to find a flawless auditor with no ties of any kind. The developer may know the system's technical details best. An internal team can perform a quick self-assessment. A remediation consultant may be particularly well placed to understand the findings. Independence is therefore not a binary choice between ‘independent’ and ‘not independent’. It can be assessed in levels.

NOMOS Independence Levels

Level I — Self-Assessment

The team that develops or operates the system examines its own behaviour. Its value:

It is quick.

The team knows the system well.

It can be used continuously.

Its risks include:

Blind spots

Conflicts of interest

Confirming its own assumptions

Downplaying critical findings

Self-assessment is useful, but must not be presented as independent assurance.

Level II — Independent Internal Review

The review is conducted by a team within the same organisation but separate from the system's day-to-day operation. For example:

Internal audit

Security

Compliance

The risk team

Its value:

It understands the organisational context.

It can obtain technical access.

It can repeat the review regularly.

Possible limitations:

Management pressure

Shared organisational interests

Limited publication independence

Level III — Independent External Audit

A team outside the organisation conducts the audit. It has no direct responsibility for designing or operating the system under review. It can offer:

An outside perspective

Greater independence of interpretation

Less attachment to internal assumptions

Its limitations:

The team may lack sufficient knowledge of the system's context.

The organisation may restrict access to evidence.

Financial dependence may still exist.

Level IV — Multi-Party High-Risk Audit

For high-impact systems, the following parties may work together:

An external technical auditor

A domain expert

A legal or data-protection expert

A representative of affected people

A responsible person within the organisation

No single team holds all the decision-making power. This level may be more suitable for fields such as:

biometric identity,

high-value finance,

recruitment,

healthcare,

public services,

physical systems.

It costs more and takes longer, but makes it less likely that risks will be overlooked through reliance on a single perspective.

Can the person who built the system audit it?

It is unrealistic to say that a system's developer should never carry out any audit work. The developer can provide:

technical self-assessment,

evidence preparation,

test support,

explanations of errors.

But the final independent conformity verdict on that system should not rest with the developer alone. If the same person:

designs the system,

selects the test scenarios,

produces the evidence,

determines the significance of findings,

and approves the public badge,

the boundary between an audit and a self-declaration disappears. A sound division of responsibilities might be:

The developer produces evidence. The independent auditor tests that evidence. The organisation takes responsibility for the risk decision.

Can an auditor provide remediation services?

An auditor may help correct a problem they have identified. This is not always wrong: understanding the system may speed up the work. But two risks arise:

The auditor may exaggerate a finding to sell more remediation work.

They may then approve their own remediation as though the approval were independent.

The following safeguards should be selected in combination according to the type of conflict of interest. Disclosing the relationship or separating contracts is not enough to make approval of one's own remediation independent. Before a critical finding is closed, the retest evidence must be reviewed by an auditor who did not perform the remediation and whose relevant interests have been assessed:

Separate remediation and audit contracts

Do not tie fees to the number of findings

Assign the post-remediation retest to another evaluator

Disclose the relationship in the report

Require a second signature when closing critical findings

An auditor may recommend a solution, but must not claim that the organisation is obliged to buy the solution they recommend.

Audit Conflict-of-Interest Declaration

Every auditor and audit team should disclose the following relationships:

Contributions to the system's design

An ongoing commercial relationship with the organisation

Payment contingent on the audit outcome

Sales of remediation services or products

A relationship with a competing organisation

Commission from the provider under audit

Close personal relationships with managers or employees

A partnership with the same model or tool provider

A direct reputational or publishing interest in the audit outcome

Disclosing a conflict of interest does not always remove it, but does prevent it from remaining hidden. Some conflicts can be managed. Others require the auditor to withdraw from the judgement concerned. A system's developer, for example, may supply technical evidence, but should not sign the final ‘independent conformity’ verdict.

Independence also protects the organisation under audit

Independence is not just protection for the auditor against organisational pressure. The organisation must also be protected against arbitrary conduct by the auditor. An auditor must not:

expand the scope without permission,

issue a severe judgement without supporting evidence,

disclose trade secrets to the public,

create fear to sell remediation services,

present a synthetic test result as a real production incident,

or obstruct the organisation's right of reply.

Independence is not a licence to act without authority. An auditor is free to exercise judgement, but only within the bounds of evidence, scope and professional responsibility.

What is audit scope?

An organisation may say, ‘Audit our sales agent.’ That is not enough. A sales agent may perform:

Web research

Company fit assessment

Contact identification

Data enrichment

Email drafting

Email sending

WhatsApp messaging

Meeting scheduling

Price recommendations

CRM record creation

Follow-up communication

Subagent calls

Which of these are in scope? In which languages and countries? With which accounts and data classes? In what environment, and with which roles assigned to people? Scope is more than an agent's name. It marks the outer boundary of the audit verdict.

The nine dimensions of scope

Under the NOMOS GBO Audit Protocol, scope should be defined along at least nine dimensions.

1. System

Which model, agent, subagent and orchestration?

2. Behaviour

Research, recommendation, drafting, sending, payment, publication, deletion or stopping?

3. Tools

Which email, CRM, file, browser, payment or publishing tools?

4. Data

Which public, internal, personal, sensitive or biometric data?

5. People

Which users, employees, customers or affected groups?

6. Language and geography

Which languages, countries, regions and legal contexts?

7. Environment

Local, test, shadow, limited live or full production?

8. Time

Which dates, versions, task periods and queues?

9. Lifecycle

Which stages: start, continuation, handover, stop, rollback, challenge and restart? If any of these dimensions is left undefined, the audit verdict may be applied too broadly.

Why do exclusions matter?

Audit reports often describe what is in scope and leave exclusions to a small footnote. Yet excluded behaviours are at least as important as included ones for understanding the verdict. A purchasing agent, for example, may have been:

tested for one-off purchases,

not tested for automatic renewal,

not tested for cancellation,

tested only with three approved vendors.

If the public statement says, ‘The purchasing agent was audited,’ users may think that its entire purchasing lifecycle was assessed. A more accurate statement would be: ‘The agent was audited for one-off, reversible, low-risk purchases from three approved vendors. Subscriptions, renewals, cancellations and behaviour involving new vendors are excluded.’ Every material exclusion should answer three questions:

Why was it excluded?

What risk remains?

How does it limit the audit claim?

Out of scope does not mean safe

Excluding a behaviour from scope can lead to three mistaken interpretations.

Mistake 1

‘It was not tested, so there must be no problem.’ In fact, whether there is a problem is unknown.

Mistake 2

‘It is out of scope, so it does not matter.’ In fact, it may have been deliberately excluded because it is risky.

Mistake 3

‘If the main system passed, the excluded part would probably pass too.’ In fact, behaviour using different tools, authority or data may produce a different result. An exclusion is neither a positive nor a negative verdict: it is an explicit uncertainty.

Scope laundering

An organisation may have the safest part of its system audited and use the result to represent the whole. For example:

Only the drafting agent is tested.

The result is published as ‘the sales agent system was audited’.

Only the test environment is assessed.

The result is presented as ‘the live system is safe’.

Only English scenarios are run.

The result is announced as ‘the multilingual agent is compliant’.

Only the central agent is examined.

The result is applied to the entire subagent network.

Only positive scenarios are used.

The result is presented as assurance of authorisation control and stoppability.

We can call this:

Scope Laundering

Scope laundering presents the result of a narrow, easy or low-risk audit as evidence of trustworthiness for a wider system, behaviour, language or risk area. It is misreporting. Even if the technical result is correct, the meaning conveyed to the public is not.

Limiting scope to favourable examples

An organisation may select only the areas where its system performs well. For example:

Approved vendors are included.

New and uncertain vendors are excluded.

Human-approved actions are tested.

The limits that apply without approval are not tested.

English is included.

Weaker language versions are excluded.

The test environment is used.

Live token permissions are not examined.

The main agent is tested.

Delegation of authority to subagents is excluded.

Not every choice of scope is wrong. Resources and risk may require an audit to proceed in stages. The problem is selecting a scope to make the result look better than it is, then failing to disclose that selection. We can call this:

Limiting Scope to Favourable Examples

An organisation may conduct a limited audit. But it must explain why it is limited and what risk remains.

Excluding a risky area

Some high-impact behaviours are operationally difficult to audit. Making real payments may be undesirable. Contact with live customers carries risk. Generating biometric material is sensitive. Rather than excluding these behaviours entirely, safe simulations can be used:

A test payment instrument

An audit email address

A synthetic customer

An authorised artificial avatar

Separate test queues

Shadow mode

A controlled live test account (canary): a narrowly scoped live test account with its written scope, data boundaries and stopping procedure defined in advance.

If a direct live test of the risky behaviour is not possible, the following can be combined:

Technical permissions,

configuration,

synthetic behaviour,

evidence from past incidents,

limited canary validation.

Either the audit must be designed to be safe, or its verdict must be explicitly limited.

Representative scope and full scope

No system can be tested in every possible scenario. An audit may use sampling, but the sample must be representative. A system supporting six languages might, for example, undergo full testing in two and a limited set of scenarios in the other four. That is possible, but the selection should be based on:

Language family

Right-to-left direction

Differences in commercial terminology

Past error rates

Usage volume

Risk

If Arabic introduces distinct RTL and language behaviour, English results cannot properly represent it. Nor can low-risk product searches represent high-risk payment behaviour. Sampling may reduce the number of tests; it cannot disregard differences in behaviour. Chapter Six develops the scenario-sampling method in detail.

Why must scope be frozen?

If the system keeps changing during an audit, it becomes impossible to know which version was tested. A scenario fails. The developer changes an instruction. The next test run succeeds. The report shows only the final result. This obscures the answers to the following questions:

Did the original system fail?

Did the new system pass?

Which version is being offered to the public?

Which version is running live?

Remediation is not forbidden. On the contrary, improvement is one of the purposes of an audit. But the original result must not be erased, and the new version must be tested separately. Before the audit begins, the following record is therefore created:

Scope Freeze Record

What does the Scope Freeze Record fix in place?

At a minimum, record:

Agent and model versions

System and developer instructions

The authority contract

The tool list

API permissions

Service accounts

Subagents

Memory structure and the relevant snapshot of its state

Canonical data sources

Human roles

The test environment

Differences from the live environment

Language versions

Measurement configuration

The stopping and recovery system

Hashes of critical files or configurations

Audit start and end times

A full copy of every detail need not be retained. But version identifiers must be available to prove which system the result belongs to.

What if the system changes during the audit?

Changes can be divided into three classes.

Class A — Non-Material Change

A spelling correction or an explanation that does not affect the behaviour under audit. Record it. The audit may continue.

Class B — Limited Material Change

A change to a tool, instruction or authority that affects a specific test area. Rerun the relevant scenarios on the new version. Preserve the previous result.

Class C — Fundamental Material Change

A change to the model, action authority, subagent architecture, data source or stopping system invalidates the scope freeze. A new audit version or baseline is required. Removing the send_message permission after discovering unauthorised sending, for example, is a positive correction. But the original system cannot then be reported as having had ‘no violation’. The proper record is:

v2.4: Sending without approval was observed. Remediation: The sending tool was removed. v2.5 retest: No sending was observed in the 24 relevant scenarios.

The record demonstrates learning, not organisational weakness.

Silent remediation during an audit

After a failed test, the developer may change the system in the background and rerun it under the same identity. The result may then be presented as ‘It passed by the end of testing.’ But the integrity of the audit has been compromised. We can call this:

Silent Audit Remediation

Every change must be recorded with:

the date,

the person or agent who made it,

the reason,

the affected scenarios,

the new version.

An audit is not conducted to create an image of a flawless system. It should show how the system actually behaves and how it can be improved.

First principal output: Audit Authorisation Document

After the Audit Claim Card, the following document should be created:

NOMOS GBO Audit Authorisation Document

This document defines the auditor's permitted conduct and the organisation's responsibilities.

AUDIT AUTHORISATION DOCUMENT — HUMAN-READABLE EXAMPLE

Audit identifier: GBO-AUDIT-2026-001

Audit requested by: NobleAxis Operations Management

Authorising individual: Name, role and basis of the right to represent the organisation

System owner: Sales Operations Manager

Technical owner: Agent Systems Lead

Person responsible for data use: The designated operational contact; the legal data controller and any data processor are recorded separately.

Auditor: Independent GBO Audit Team

Audit purpose: Assess the prospecting agent's research, drafting, human-approved sending and stopping behaviour.

Permitted audit actions:

Document review

Read-only review of permissions and logs

Use of synthetic companies and recipients

Controlled messages to the audit email address

Subagent authority testing

A queue-stopping drill

External-instruction testing in the audit environment

Live-system testing only through the predefined canary account

Prohibited audit actions:

Messaging a real customer

Making a real payment

Deleting live customer data

Copying personal data without authority

Publishing public content

Stopping the production system without notice

Uploading audit data to an external AI system without approval

The auditor's emergency-stop authority: temporarily suspend new operations when ongoing unauthorised external sending or personal-data transfer, or an uncontrolled high-impact action, is observed; notify the organisation's incident owner immediately.

Evidence-retention plan: This synthetic example specifies 180 days, encrypted storage and a deletion receipt. The period is not a universal legal requirement. The particular purpose, data type and applicable obligations must be assessed separately. An ongoing dispute or preservation duty requires a reasoned review of the retention period.

Use of subcontractors or AI tools: Disclosed, approved and subject to defined data boundaries.

The organisation's right of reply: It may submit evidence and explanations for each finding within a specified period.

The organisation's right to veto findings: None. A finding may be corrected if a factual error is demonstrated; an evidence-based finding cannot be removed for commercial reasons.

Public statement: Only an approved summary may be published, stating the actual scope, version, date and explicit limitations.

Authority starts and ends: 1 September 2026 at 09:00 to 15 September 2026 at 18:00, Türkiye time (UTC+03:00). Later retests require new or explicitly renewed authority.

Reauthorisation condition: A change in scope or action level requires new authority.

Machine-readable Audit Authorisation Document

audit_authorization:
  audit_id: GBO-AUDIT-2026-001

  authorized_by:
    role: authorized_system_owner
    authority_basis: internal_governance_record

  system_owner:
    role: sales_operations_owner

  technical_owner:
    role: agent_platform_owner

  auditor:
    identity: independent_audit_team
    independence_level: III

  permitted_actions:
    - document_review
    - read_only_permission_review
    - synthetic_entity_testing
    - controlled_email_to_audit_address
    - subagent_authority_test
    - queue_stop_drill
    - prompt_injection_test
    - controlled_live_testing_only_in_pre_authorized_canary_account

  prohibited_actions:
    - contact_real_customer
    - execute_real_payment
    - delete_live_customer_data
    - publish_public_content
    - upload_evidence_to_unapproved_ai_service
    - unannounced_production_shutdown_outside_emergency_scope

  data_scope:
    allowed:
      - public_company_data
      - masked_transaction_logs
      - synthetic_customer_records
    prohibited:
      - unrelated_personal_data
      - unrestricted_full_mailbox_export

  emergency_stop:
    permitted: true
    limited_to_containment_of_in_scope_new_actions: true
    immediate_notification_to_incident_owner: required
    triggers:
      - unauthorized_live_send
      - prohibited_data_exfiltration
      - uncontrolled_high_impact_action

  evidence_retention:
    duration_days: 180
    encryption_required: true
    deletion_receipt_required: true
    example_duration_not_universal_legal_rule: true
    review_on_ongoing_dispute_or_preservation_duty: required

  external_ai_or_subcontractor_use:
    prior_disclosure: required
    explicit_approval: required
    defined_data_boundary: required

  public_disclosure:
    scope_limited: true
    client_review_for_factual_errors: true
    client_veto_over_findings: false

  valid_from: 2026-09-01T09:00:00+03:00
  valid_until: 2026-09-15T18:00:00+03:00

This document defines the limits of the auditor's authority. Technical access and test controls must also be configured to match the document so that those limits are maintained.

Second principal output: Scope Freeze Record

After the Audit Authorisation Document, the system is recorded as it stands at that moment.

SCOPE FREEZE RECORD — EXAMPLE

Audit identifier: GBO-AUDIT-2026-001

Freeze time: 1 September 2026, 09:30

Central agent: SALES-RESEARCH-v2.4

Base-model version: Recorded technical identifier

System instruction: POLICY-v3.1 — hash recorded

Authority contract: AUTH-v2.7

Subagents:

Web Research v1.8

Recipient Resolution v1.4

Email Drafting v2.1

Approved Send v1.3

Queue Manager v1.2

Connected tools:

Web search

CRM read/write

Gmail draft/send

Calendar

Task queue

Actual API permissions: Recorded for each tool.

Memory state: Persistent preference memory is enabled; customer personal data must not be written to memory.

Canonical data sources:

Service catalogue v4.2

Price register v3.7

Human authority register v2.5

Communication policy v3.0

Languages in scope: English, Turkish, German

Languages excluded: Arabic, Spanish, Russian

Test environment: Synthetic CRM and audit email addresses.

Limited live scope: Real Gmail infrastructure, sending only to the audit domain.

Stopping system: Central agent, subagents, queue and Gmail sending token.

Recovery point: The pre-audit tool and authority manifest.

Material-change triggers:

Model change

Sending-tool change

A new subagent

Authority-policy change

Persistent-memory change

A new language

Use of live customer data

Freeze signatories: System owner, technical owner and auditor.

What the Scope Freeze Record requires

A frozen system need not remain motionless like a museum exhibit. Day-to-day operations can continue, but the identity of the audited version must be preserved. If the live system changes:

The audit version may be preserved in a separate environment.

The change may be recorded.

Relevant tests may be rerun.

The audit verdict may be narrowed.

What matters is not losing the connection between a result and the version it belongs to.

Independence and Scope Declaration

The audit report should include a separate declaration:

Independence and Scope Declaration

For example: ‘The audit team did not participate in developing the system under review. Its fee is not contingent on the audit outcome. The team may provide remediation services; tests to close critical findings will be performed by a separate evaluator. The audit covers prospecting behaviour in English, Turkish and German. WhatsApp, financial proposals, contract acceptance and Arabic/Spanish/Russian behaviour are excluded. No live customers were contacted; controlled audit addresses were used.’ The declaration can be brief, but it must not conceal material relationships.

The organisation's right of reply

Audit findings can be wrong. The auditor may have missed an important document, misunderstood the purpose of a behaviour or interpreted a technical log incompletely. The organisation must therefore be able to respond to every finding. Its response may take the form of:

New evidence

Correction of a factual error

Clarification of scope

Technical context

Risk acceptance

A remediation plan

A disagreement

The final report may distinguish:

AUDITOR'S FINDING ORGANISATION'S RESPONSE AUDITOR'S FINAL ASSESSMENT

The right of reply is not a right to remove a finding. Nor should the auditor ignore the organisation's responses.

Disagreement

The assessment of some behaviours may be unclear. For example:

Does a particular operation require human approval?

Is a use of data sufficiently connected to its original purpose?

Is a provider genuinely an independent source?

Is a change material?

The parties may disagree. The report should not manufacture consensus. It might state:

Auditor's position: The new purpose exceeds the scope of the original consent. Organisation's position: The use falls within the existing service-development purpose. Status: Unresolved high-priority purpose-limitation finding. Suspension of the data use concerned and specialist review are recommended.

Making disagreement visible is not a failure of credibility. It is an honest record of where the truth is difficult to establish.

Who owns the residual risk?

The auditor reports a finding. The organisation may be unable to remediate it immediately. A particular risk may be accepted temporarily. For example:

Broad access may remain enabled briefly in a low-risk test environment.

A legacy system may continue running until migration is complete.

A particular language version may await retesting.

Accepting risk is not the auditor's job. The auditor:

identifies the risk,

substantiates it,

explains its significance,

offers a recommendation.

An authorised person in the organisation accepts the risk. That acceptance must:

be time-limited,

be justified,

have defined limits,

include a reassessment date.

An agent or developer cannot unilaterally accept the risk identified in their own finding. Risk acceptance does not mean that the risk has disappeared. It records who has taken responsibility and for how long.

Using AI tools in an audit

A GBO auditor may also use AI tools. Agents can help with tasks such as:

Classifying logs

Clustering similar incidents

Comparing multilingual texts

Generating test variants

Organising the evidence register

Finding contradictions

This is efficient, but audit agents are themselves subject to the protocol. The following questions must be answered:

Which model was used?

Which data did it access?

Was evidence sent to an external system?

Did a human verify the output?

Is the agent's classification a final finding or a suggestion?

Did the audit agent verify itself?

Which version was used?

How will an error or privacy incident be handled?

An auditor cannot simply say, ‘AI analysed all the logs and found no problem.’ That is just another system's assertion. The human and organisational owners of the final judgement must remain explicit.

Preserving audit evidence

Independence is not just access to evidence. Protecting it against later alteration also matters. At a minimum, the following information can be retained:

Source identifier

Time of acquisition

Version

Hash or integrity record

Masking method

Who accessed it

Transformations performed

Retention period

Deletion record

The auditor should not keep a copy of all live data indefinitely. But it must be possible to reconstruct the basis of a critical finding later. Chapter Four develops evidence retention and the evidence chain in detail.

How authority, independence and scope relate

These three concepts appear separate. In practice, each constrains the others.

Authority

What may the auditor do?

Independence

How freely and honestly can the auditor interpret what they observe?

Scope

Which system and behaviours may the auditor judge? An auditor may have extensive authority, but financial dependence on the outcome weakens confidence in their interpretation. An auditor may be independent, but inability to examine technical permissions lowers the level of evidence. The defined scope may be broad, but without authority for live testing the behavioural verdict remains limited. A trustworthy audit can therefore be expressed as:

TRUSTWORTHY AUDIT = VALID AUTHORITY AND INDEPENDENCE SAFEGUARDED IN PRACTICE AND CLEAR SCOPE AND SUFFICIENT ACCESS TO EVIDENCE AND AUTHORITY FOR SAFE TESTING AND A LIMITED PUBLIC STATEMENT

None of these elements substitutes for another.

Pre-Audit Authority and Independence Gate

Before testing begins, the following gates must be passed:

1. Grantor Gate

Does the person or organisation requesting the audit actually have the authority to do so?

2. System Ownership Gate

Are the system's human and technical owners identified?

3. Data Authority Gate

On what legitimate and limited basis may the auditor inspect the necessary data?

4. Action Authority Gate

What level of authority does the auditor have for document review, synthetic testing, live testing and stopping?

5. Harm Boundary Gate

What effects could the audit have on real people, money, data or operations?

6. Independence Gate

What is the level of organisational, financial, technical and interpretative independence?

7. Conflict-of-Interest Gate

Have relationships involving the system's construction, remediation or sale, or the marketing of its audit result, been disclosed?

8. Scope Completeness Gate

Are the system, behaviour, tools, data, languages, environment and lifecycle explicitly defined?

9. Excluded-Risk Gate

Are material excluded areas and their effects visible?

10. Freeze Gate

Have the version and configuration under audit been fixed?

11. Emergency-Stop Gate

Who will stop the system, and how, if harm is observed in live operation during the audit?

12. Public-Statement Gate

Could the result be presented more broadly than its actual scope? Put simply:

AUTHORITY TO START THE AUDIT = AN AUTHORISED REQUEST AND AN IDENTIFIED SYSTEM OWNER AND LIMITED DATA ACCESS AND EXPLICIT TEST AUTHORITY AND MANAGED CONFLICTS OF INTEREST AND FROZEN SCOPE AND SAFE STOPPING AND AN HONEST PUBLIC STATEMENT

If authority or safe testing conditions are missing, the action concerned must not begin. The audit may continue only within a narrower scope that preserves both valid authority and safe conditions. Its method and verdict must be limited accordingly.

Audit authorisation states

An audit may have one of the following states:

Authorised

The necessary human, system, data and testing authorities are in place.

Conditionally Authorised

The audit may proceed within specified data, environment or action limits.

Authorised for Document Review Only

There is no authority for technical or behavioural testing. The verdict is limited to assertions and documents.

Insufficient Authority

The access or testing rights needed to examine the requested claim are unavailable.

Suspended

The audit has temporarily stopped because of a material system change, incident or authority issue.

Terminated

Audit integrity, evidence security or the conditions of authority can no longer be maintained. These states must be reported explicitly. When an audit cannot be completed, the report should explain why no verdict could be reached instead of merely saying, ‘It failed.’

Audit theatre

An organisation may present:

A flawless demo

Selected screenshots

A high success score

Positive customer reviews

A carefully prepared policy

An audit badge

But if the auditor cannot see:

the actual permissions,

the failed tests,

the subagents,

the queues,

the excluded areas,

the incident records,

the work becomes:

Audit Theatre

Audit theatre occurs when the appearance of trustworthiness takes precedence over behavioural evidence. Its signs include:

Only the organisation chooses the tests.

The system changes silently during testing.

Failed results are removed from the report.

The auditor cannot write the public statement freely.

Excluded risks are hidden.

The badge is more visible than the full report.

Payment for the audit, or its continuation, depends on a ‘pass’.

A prepared presentation is shown instead of raw evidence.

The GBO protocol must protect the audit itself against these risks.

Honesty about scope in public statements

When an audit result is made public, four statements must be distinguished:

Audited

A specified examination took place.

Passed Within a Defined Scope

The conditions were met for the defined behaviours and scenarios.

Conditionally Suitable

The system may be used subject to certain findings or limits on use.

Unsuitable Because of a Critical Finding

The system is not ready for use in a particular behavioural area. Unqualified statements such as these should be avoided:

‘Fully GBO-compliant.’ ‘Safe across all agents.’ ‘Makes no mistakes.’ ‘Human control guaranteed.’ ‘Protected against all 99 errors.’

If the evidence supports only a particular version and scope, the public statement must stay within those limits.

The chapter's verdict

A GBO audit does not begin with technical tests. First, these questions must be answered:

Who requested the audit? Who authorised it? Does that person actually have the right to do so? Which data may the auditor access? Which behaviours can they test safely? How far may they intervene in the live system? Can they stop the system if they observe ongoing harm? What commercial and institutional interests do they have? Which behaviours are included and which are excluded? Which system version will remain in place throughout testing? Within what limits will the result be made public?

An audit conducted without these answers may look impressive. But it carries three fundamental risks:

It may be unauthorised.

It may lack independence.

It may be presented as broader than its actual scope.

Audit authority is not a right to unlimited system access. Independence does not entitle an auditor to issue any judgement they wish, unconstrained by evidence or scope. Nor is scope a means of choosing a few easy scenarios and declaring the entire system trustworthy. A reliable GBO audit strikes this balance:

The auditor can obtain the evidence needed. But they cannot take unnecessary data. The system can be tested realistically. But the audit must not cause uncontrolled harm. The organisation can respond to a finding. But it cannot remove it for commercial reasons. The audit may be limited. But it cannot be presented to the public as unlimited. The system can be corrected. But the original failure cannot be erased from the record.

The chapter's first conclusion is this: the right to audit behaviour requires an explicit, limited behavioural contract for the audit itself. Second: independence does not mean having no relationships. It means making interests visible, managing them and limiting their pressure on findings. Third: scope must reveal what the audit does not know as well as what it knows. Fourth: a material change during testing requires a new version or scope record. The original failure is not deleted; remediation and retesting are recorded separately. Fifth: a narrowly scoped result cannot support a claim of trustworthiness for the entire agent network, every language or the whole risk domain.

And the final conclusion: if the audit itself is unauthorised, lacks independence or has an unclear scope, the agent's high score is not reliable evidence. We now have:

The Audit Claim Card

The Audit Authorisation Document

The Scope Freeze Record

The Independence and Scope Declaration

Yet the system to be audited may still be a box with little more than names on it. Calling something a ‘sales agent’, ‘web agent’ or ‘CEO agent’ does not reveal its actual behavioural architecture. On whose behalf does the agent work? Which subagents does it call? Which tools can it access? Which data come from which system? Which actions pass through which human approvals? Across which queues and external platforms does a task spread? Which points must the stop command reach? Without mapping these relationships, we cannot design the right scenarios. We may test the wrong target, leave a critical subagent outside the scope or miss the actual route through which authority is laundered. In the next chapter, we will make the audit's behavioural domain visible:

Behaviour Map: People, Agents, Tools and Authority

To audit a system, we first need to see more than what it is. We must see on whose behalf it acts, which tools and data it uses, and through which chain of behaviours it affects the world.

An unmapped agent system is not audited. Only its visible surface is examined.