# AI-Native Operating Architecture for SMEs — Document Set Overview

| | |
|---|---|
| **Version** | v1.0 (Final) |
| **Date** | 2026-08-10 |
| **Scope** | Model operating and technology architecture for organisations with 1–500 employees |
| **Horizon** | Up to 24 months, calibrated to readiness, risk and size |
| **Status** | Reference model — adapt and approve locally |

---

## 1. Purpose

This document set defines a practical reference architecture and adoption program for an organisation that wants AI to become part of normal operations rather than remain a collection of employee chat tools and pilots.

The set covers:

- the target operating and technology architecture;
- a phased implementation plan;
- data, security and AI governance;
- benefits measurement and financial conversion; and
- the minimum evidence required before an AI-enabled workflow is allowed to act in production.

It is a **reference model, not a product recipe or a promise that every organisation will become fully AI-native in 24 months**. The normal Month-24 target is a governed, measured and orchestrated operating model (maturity L4 in Doc 01), with selective AI-native characteristics in the functions where the economics and controls justify them.

The title uses “SME” commercially. Legal and statistical definitions vary by jurisdiction; the S3 band includes organisations that some markets classify as lower mid-market rather than SME.

## 2. Target organisation profile

The model is designed for an organisation that:

- has **1–500 employees**, using the internal size bands in §6;
- operates a predominantly SaaS and public-cloud estate;
- uses Google Workspace or Microsoft 365 as its primary identity and collaboration estate;
- is a **buyer and integrator of models**, not a frontier-model developer;
- has conventional functions such as finance, sales, marketing, support, people, operations and possibly product/engineering; and
- is not subject to a control regime that requires a materially different reference architecture.

“Not heavily regulated” does **not** mean unregulated. Privacy, employment, discrimination, consumer, records, intellectual-property, cybersecurity and contractual obligations still apply. Health, financial services, government, defence, critical infrastructure, safety-critical operations and similarly regulated deployments require a separate overlay and may require different technology choices, validation and segregation.

## 3. What “AI-native” means in this model

AI-native is an operating condition, not a licence tier or a count of agents. An organisation is moving towards AI-native operation when the following are true in the relevant functions:

1. **Work is managed as outcomes and workflows.** Each production workflow has an owner, trigger, inputs, controls, outputs, service levels, fallback path and unit economics.
2. **Models advise; deterministic controls authorise side effects.** A model may interpret, classify, draft, plan or propose an action. Business rules, policy checks, approvals and the transaction executor determine whether an action is permitted and commit it safely.
3. **Every material actor and action is attributable.** Human users, logical agents, runtime workloads and approvers are distinguishable; authority is least-privilege, time-bounded and auditable.
4. **Authoritative data is accessible through governed interfaces.** Operational data remains in systems of record. Organisational knowledge is curated, permission-aware, attributable to sources and subject to freshness and deletion controls.
5. **Production AI is evaluated and operated like a service.** Each system has acceptance tests, production monitoring, incident handling, rollback or manual fallback, change control and retirement criteria.
6. **Economics are measured per outcome.** Capacity created, cash converted, quality effects and total cost of ownership are reported separately. Adoption and token use are inputs, not benefits.
7. **The organisation continuously redesigns work.** It removes, simplifies and standardises work before automating it, then uses production evidence to improve processes, controls and role design.

AI-native therefore does **not** mean maximum autonomy. A mature organisation often chooses deterministic automation or human judgment over an agent because that is safer, cheaper and more reliable.

## 4. Three distinct AI operating patterns

The documents deliberately distinguish three patterns that are often conflated:

| Pattern | Primary user | Typical purpose | Default control posture |
|---|---|---|---|
| **Employee assistant** | A person | Search, analysis, drafting, coding and individual productivity | User remains accountable; no unattended side effects |
| **Bounded AI workflow** | A function or process | Repeatable task with explicit inputs, outputs and controls | Production owner, evaluations, deterministic action boundary and operational support |
| **Autonomous action class** | A background workflow | Narrow, repetitive, low-impact action within hard limits | Exceptional; only after sustained evidence, continuous monitoring and immediate pause capability |

An employee assistant is not automatically a production agent. A production agent is not granted one blanket autonomy level: **execution mode is assigned per action class**.

## 5. The document set

| # | Document | Primary question |
|---|---|---|
| 01 | **Target Architecture** | What is the target technical and operating architecture, and what controls sit between a model and a business transaction? |
| 02 | **Phased Implementation Plan** | How should the capability and workflow portfolio be introduced, gated and scaled over up to 24 months? |
| 03 | **Data, Security & AI Governance** | How are AI systems, data, identities, tools, risks, incidents, affected people and suppliers governed? |
| 04 | **Benefits Realisation** | How are capacity, service, quality, cash benefits and total cost measured without double counting? |

## 6. Size bands used throughout

The bands set a default level of architectural depth, not a mandatory shopping list.

| Band | Headcount | Default posture |
|---|---:|---|
| **S1 — Micro / very small** | 1–20 | One primary sanctioned assistant; native SaaS AI and managed automation; fractional ownership; no custom platform unless a specific workflow pays for it |
| **S2 — Small** | 21–100 | Managed workflow capability; first bounded production agents; lightweight central logging; partner or fractional engineering; data platform only where use cases require it |
| **S3 — Medium / lower mid-market** | 101–500 | Small internal platform capability; standard action-control pattern; central model/API controls where justified; durable workflows; formal portfolio, assurance and benefits routines |

A highly technical 15-person software company may need selected S2 controls. A 200-person organisation with few knowledge workflows may not need the full S3 platform. Use the readiness and complexity triggers in Doc 02 rather than headcount alone.

## 7. Non-negotiable design positions

1. **Simplify before automating.** Remove unnecessary work, fix the source process and use deterministic automation before introducing probabilistic reasoning.
2. **Use the minimum sufficient autonomy.** T3 is reserved for narrow, reversible, low-impact and normally non-sensitive action classes. Open-ended autonomous operation is outside the target state.
3. **Separate reasoning from execution.** Models do not hold broad credentials or directly improvise transactions. They emit typed proposals; a policy and transaction layer validates and executes them.
4. **Identity is necessary but not sufficient.** A logical agent ID, a workload principal and a delegating user are separate concepts. User-facing access preserves the user’s downstream permissions; background workflows use narrowly scoped workload identity.
5. **MCP is an interface, not a security boundary or enterprise service bus.** Use it where an AI host needs a standard tool interface, with pinned protocol versions, explicit authorisation, input/output validation and policy enforcement. Native APIs, events and workflow integrations remain the default for system-to-system processing.
6. **Systems of record remain authoritative.** RAG, vector indexes, model memory and chat history are not systems of record and are never used as transaction truth.
7. **Permission-aware retrieval is necessary but incomplete.** Security trimming, source provenance, freshness, permission-change propagation and deletion from indexes are all required.
8. **Do not create a prompt-and-response data lake.** Audit logs record sufficient metadata and governed references to reconstruct material actions. Raw content is retained only for a defined purpose, for the shortest appropriate period, with redaction and access controls.
9. **Model changes are production changes.** Model, prompt, tool, data-source and policy changes require proportionate regression testing, controlled rollout and rollback.
10. **One primary assistant pattern per user cohort.** Do not buy overlapping enterprise assistants for everyone by default. Add specialist products only where incremental value exceeds licence, change and governance cost.
11. **Local evidence outranks generic productivity claims.** AI performance is task- and context-dependent. Every material workflow must earn its business case against a local baseline and quality guardrails.
12. **Capacity and cash are separate ledgers.** Time released is an operational benefit. It becomes a financial benefit only when a documented conversion changes an approved cost, hiring or revenue plan.

## 8. How to adapt the model

1. Confirm the applicable organisation profile and document the exclusions.
2. Complete the readiness assessment in Doc 02 Appendix B.
3. Establish the legal and contractual obligations register for the jurisdictions, employees, customers and data involved.
4. Select the size-band defaults, then override them using workflow volume, integration complexity, data sensitivity and risk.
5. Record the architecture decisions in Doc 01, including the action-control, identity, logging and autonomy-ceiling decisions.
6. Inventory sanctioned, embedded and shadow AI; select the primary employee-assistant pattern by cohort.
7. Shortlist workflows using Doc 02’s “eliminate → simplify → deterministic → AI-assisted → agentic” decision sequence.
8. Localise the governance thresholds, prohibited uses, records rules, approval authorities and incident obligations in Doc 03.
9. Build the local cost and benefits baseline in Doc 04 before committing to portfolio savings.
10. Re-plan the phases against business capacity. S1 may reach its appropriate steady state in 6–12 months and can use **Appendix A** as its working summary; S3 commonly uses the full horizon.

## 9. Interpretation of requirement words

Within the document set:

- **MUST** identifies a baseline control required for claimed conformance to this reference model.
- **SHOULD** identifies the normal position; deviations require a recorded rationale.
- **MAY** identifies an optional capability selected by need.

Product names are examples current at the review date, not endorsements. Procurement must revalidate product features, contractual terms, pricing, residency and integration maturity.

## 10. Reference baseline at 10 August 2026

The model is designed to be compatible with, and proportionate to, the following current sources of good practice:

- Australian Government, **Guidance for AI Adoption** (six essential practices);
- ASD’s Australian Cyber Security Centre and international partners, **Careful Adoption of Agentic AI Services** (May 2026);
- NIST **AI Risk Management Framework 1.0** and **Generative AI Profile (NIST AI 600-1)**;
- NIST **Cybersecurity Framework 2.0** and **Secure Software Development Framework 1.1**;
- ISO/IEC **42001:2023** (AI management systems), **23894:2023** (AI risk management) and **42005:2025** (AI system impact assessment);
- OWASP **Top 10 for LLM Applications 2026** and **Top 10 for Agentic Applications 2026**;
- Model Context Protocol specification dated **2026-07-28**, together with OAuth 2.0 Security Best Current Practice (RFC 9700);
- OAIC guidance on privacy and commercially available AI products; and
- jurisdiction-specific requirements, including the EU AI Act where the organisation is in scope.

Conformance to this model is not certification against any of those frameworks and is not legal advice.

## 11. Glossary

| Term | Meaning in this document set |
|---|---|
| **Action broker / transaction boundary** | Deterministic service that receives a typed proposed action, checks policy and authority, executes the permitted transaction idempotently, and records the result. |
| **Action class** | A specific type of operation with a defined scope and impact, such as “add an internal CRM note” or “post a three-way-matched invoice below a threshold”. Execution mode is assigned at this level. |
| **Agent** | A software component that uses a model to interpret context, choose or propose steps, and interact with approved tools within a bounded workflow. |
| **AI system** | The complete deployed system: models, prompts, data, tools, workflow code, policies, interfaces, people and operating procedures. |
| **AI system register** | Organisation-wide inventory of procured, embedded and built AI systems, their owners, purpose, risk, data, suppliers, tests and lifecycle status. |
| **Autonomy / execution modes (T0–T3)** | Retrieve → Draft → Act with approval → Autonomous within a bounded policy. They describe execution, not overall risk. |
| **Authoritative source map** | Record of which system, table, field or document class is authoritative for a business fact. |
| **Banked financial benefit** | A finance-validated change to an approved cost, hiring, revenue or cash plan, supported by evidence and net of attributable cost. |
| **Context** | Information supplied to a model for one execution, including instructions, retrieved data and tool results. It is minimised and labelled by trust level. |
| **Control plane** | Registries, policy, identity, deployment, evaluation, budgets, approvals and administrative controls governing AI execution. |
| **DLP** | Data loss prevention controls that detect or block sensitive data movement. |
| **DPIA / impact assessment** | Proportionate assessment of privacy and broader effects on people, groups and rights before deployment and after material change. |
| **Evaluation set** | Versioned representative, edge and adversarial test cases with expected outcomes or calibrated rubrics. “Golden set” is retained as an informal synonym. |
| **Gateway** | Optional control point for model API credentials, routing, quotas, version policy, traces and cost allocation. It does not make models behaviourally interchangeable. |
| **HITL** | Human-in-the-loop control in which an appropriately authorised person reviews a specific proposed action before execution. |
| **Idempotency** | Property that permits a transaction to be retried without creating a duplicate side effect. |
| **IdP / SSO / SCIM** | Identity provider / single sign-on / automated account lifecycle provisioning. |
| **iPaaS** | Integration platform used primarily for deterministic application workflows. |
| **Logical agent ID** | Stable registry identity for an AI system or workflow, distinct from the runtime credential and any human on whose behalf it acts. |
| **MCP** | Model Context Protocol, a standard interface through which an AI host can access declared resources and tools. It does not by itself supply least privilege, business policy or safe execution. |
| **Model memory** | Persisted information used across executions. It is treated as a governed data store, not an informal model feature. |
| **Policy decision point** | Deterministic component that evaluates whether a proposed action is permitted given identity, data, action, value, context and risk. |
| **RAG** | Retrieval-augmented generation: retrieval of governed sources to ground a model response. It is not a database or authorisation system. |
| **Security trimming** | Applying the effective permissions of the requesting principal at retrieval time so prohibited content is not returned. |
| **SoR** | System of record — the authoritative application or data field for a defined business fact. |
| **TCO** | Total cost of ownership, including licences, usage, implementation, integration, data, evaluation, security, change, support and decommissioning. |
| **Unit cost** | Net cost per completed and quality-accepted outcome, such as per invoice, ticket or proposal. |
| **Workflow owner** | Accountable business owner for the workflow’s outcome, controls, evidence, cost and lifecycle. |
| **Workload identity** | Short-lived, non-human runtime identity used by a background service or agent, scoped to its approved resources and actions. |

## 12. Change log

| Version | Date | Change |
|---|---|---|
| v0.1 | 2026-08-10 | Initial four-document model |
| v0.2 | 2026-08-10 | Added operating model, runtime sequence, ADR template, workflow catalogue, readiness assessment and evaluation guidance |
| **v0.3** | **2026-08-10** | Resolved the Month-24 maturity contradiction; separated employee assistants, bounded workflows and autonomous action classes; introduced a deterministic action-control boundary; separated logical, delegated and workload identity; narrowed MCP’s role; strengthened RAG, logging, evaluation, continuity and supplier controls; aligned governance to August 2026 guidance; and separated operational capacity from cash benefits |
| **v1.0** | **2026-08-10** | Final release. Adopted the v0.3 third-party review in full after independent verification of its August-2026 citations (MCP 2026-07-28, OWASP LLM and Agentic 2026 editions, ASD/ACSC agentic-AI guidance, EU AI Omnibus dates, Australian ADM transparency commencement). Corrected the review register’s readiness-dimension count (eight, not nine); added the S1 minimum viable pattern (00 Appendix A); added value-story presentation guidance (04 §2.2) and study-anchored productivity reference points (04 §12) |

---

## Appendix A — The S1 minimum viable pattern (1–20 people)

This appendix is the whole model at S1 scale. An S1 organisation should not attempt to run the program as written for S3; the four documents remain the reference when a question of detail arises.

### A.1 Do these seven things

1. **One assistant, business terms.** Choose one sanctioned assistant — suite AI or one enterprise chat product — on business/team terms with training-on-your-data disabled and central administration. Personal accounts never touch company data.
2. **Identity basics.** Universal MFA, a shared password manager, central billing, and a written joiner/leaver checklist that is actually executed on the day someone leaves.
3. **Sharing hygiene.** Lock external-sharing defaults, review "anyone with link" access on anything sensitive, and know who owns the Drive or SharePoint estate.
4. **A one-page policy.** Approved tools; data that must never be pasted into AI (customer records, financials, HR matters, credentials); verify before you use or send; how to report a mistake — no blame.
5. **Three sheets.** A systems/tools register, a workflow register and a benefits/cost sheet. Spreadsheets are fine; keeping them current is the control.
6. **Two or three workflows.** Screen candidates with eliminate → simplify → deterministic → AI. Prefer high-frequency, low-risk tasks with objectively checkable outputs. Baseline volume, touch time and cost before switching anything on.
7. **A monthly 30-minute review.** Adoption, quality problems, spend against value, and kill/keep decisions.

### A.2 Non-negotiables at any size

- Models draft; a person commits (T1 is the default). Any automated write action — even inside a vendor tool — needs the approval discipline and a way to switch it off.
- No agent or automation ever runs on a person's credentials.
- Capacity is not cash: time saved becomes money only when a bill, contract, licence or planned hire actually changes.
- If it touches customer money, bank details, employment decisions or legal commitments, a person decides — every time.

### A.3 Skip until a specific workflow pays for it

Warehouse, model gateway, custom RAG/vector stack, code-first agents, multi-agent frameworks, and self-hosted anything.

### A.4 A 6–12 month path

- **Months 0–1:** items 1–5 above; select workflows; capture baselines.
- **Months 2–4:** assistant adopted in earnest; first two workflows live as drafts and checks (T0/T1); measure against baseline.
- **Months 5–8:** keep what works, kill what does not; consider one vendor-managed T2 automation (for example invoice capture with owner approval) if the evidence supports it.
- **Months 9–12:** consolidate overlapping tools; bank any real cash effects (subscriptions cancelled, outsourced work reduced); decide what year two looks like.

### A.5 When to graduate to S2 controls

Triggers, not headcount: the first custom write-enabled automation; grounding C3 data beyond suite-native capability; a workflow whose failure would materially harm a customer; or more than roughly five production workflows.


---

# 01 — Target Operating and Technology Architecture

| | |
|---|---|
| **Version** | v1.0 (Final) |
| **Date** | 2026-08-10 |
| **Target** | Governed, measured and orchestrated human+AI operations |
| **Companion docs** | 00 Overview · 02 Implementation Plan · 03 Data, Security & AI Governance · 04 Benefits Realisation |

---

## 1. Purpose and scope

This document defines the target-state architecture for a predominantly SaaS-based organisation with 1–500 employees that buys and integrates AI capabilities rather than training frontier models.

It defines:

- the operating and technical principles;
- a six-plane reference model;
- execution modes for AI-enabled action classes;
- capability depth by size band;
- reference deployment patterns for Google- and Microsoft-anchored estates;
- integration, identity, action-control and observability patterns;
- non-functional requirements;
- the target operating model and maturity model; and
- the architecture decisions each adopting organisation must record.

The normal Month-24 ambition is maturity **L4 — Orchestrated**, with selective L5 characteristics. “Full AI-native transformation in 24 months” is not a required claim and should not be used as a program success measure.

Out of scope: frontier-model training, safety-critical control, physical/OT systems, unrestricted autonomous agents, and regulated-sector overlays that materially alter the reference design.

## 2. Current-state archetype

A typical starting estate has:

- 20–120 SaaS applications, with a smaller set carrying most operational work;
- central identity on major products but password-based or shared access on the long tail;
- years of accumulated Drive, SharePoint, group and external-sharing permission debt;
- personal or consumer AI use without consistent contractual protection, inventory or monitoring;
- duplicated employee assistants and embedded AI features purchased without a portfolio decision;
- brittle point-to-point automations owned by individuals;
- operational data trapped in systems of record and knowledge dispersed across documents, tickets, email and people; and
- benefits claims based on anecdotes, adoption or “minutes saved”, without local baselines or financial conversion.

The target state does not replace every core SaaS product. It creates a governed execution fabric around authoritative applications, improves permission and data discipline, introduces a deterministic boundary between model reasoning and business transactions, and operates AI systems as measured services.

## 3. Architecture principles

| # | Principle | Architectural implication |
|---|---|---|
| P1 | **Outcome before technology** | Every production capability starts with a workflow, owner, baseline, quality requirement and economic hypothesis. |
| P2 | **Eliminate → simplify → standardise → automate → add AI** | Agentic technology is not used to preserve a broken or unnecessary process. |
| P3 | **Deterministic where possible; probabilistic where valuable** | Rules perform calculations, validations, routing and authorisation. Models interpret ambiguity, language and unstructured content. |
| P4 | **Models propose; controlled services transact** | A typed action proposal passes through deterministic policy, approval and transaction controls before any material side effect. |
| P5 | **Minimum sufficient autonomy** | Each action class receives the lowest execution mode that produces the required business value. T3 is exceptional. |
| P6 | **Identity and authority are explicit** | Logical agent, workload principal, delegating user and approver are separately attributable. No ambient or shared authority. |
| P7 | **Authoritative sources stay authoritative** | RAG, caches and model memory do not become systems of record. Writes occur only through governed application interfaces. |
| P8 | **Permission-correct, purpose-limited context** | Retrieval applies effective user/workload permissions, source curation, trust labels, minimisation, provenance, freshness and deletion. |
| P9 | **Open interfaces, controlled boundaries** | Native APIs/events first; iPaaS for deterministic workflows; MCP for approved agent-tool interfaces; browser/computer use only by exception. |
| P10 | **Evidence before promotion** | Offline evaluation, shadow operation, production SLOs and incident history gate T2/T3 promotion and model changes. |
| P11 | **Observability without uncontrolled data replication** | Audit metadata, governed references and business records are separated from short-lived diagnostic traces. |
| P12 | **Graceful degradation** | Every material workflow has timeouts, retries, idempotency, manual fallback and a tested pause or rollback path. |
| P13 | **Buy capability; build thin control and differentiation** | Custom development concentrates on proprietary workflow logic, policy and integration—not generic chat, vector databases or model hosting. |
| P14 | **Cost per accepted outcome** | Architecture captures model, platform, human-review, rework and support cost at workflow level. |
| P15 | **Exit is designed, not assumed** | Data, prompts, evaluations, action schemas, policies and runbooks are portable; model/provider switches are regression-tested. |

## 4. Six-plane reference model

```text
┌─────────────────────────────────────────────────────────────────────┐
│ P1  EXPERIENCE & COLLABORATION                                      │
│     Employee assistants · Embedded AI · Workflow inboxes ·          │
│     Authenticated approval and exception surfaces                   │
├─────────────────────────────────────────────────────────────────────┤
│ P2  WORKFLOWS, AGENTS & ACTION CONTROL                              │
│     Durable workflow engine · Agent runtime · Context assembler ·   │
│     Typed action broker · Policy decision point · Tool interfaces · │
│     Approvals · Queues · Schedulers · Registry · Pause controls     │
├─────────────────────────────────────────────────────────────────────┤
│ P3  MODELS, CONTEXT & EVALUATION                                    │
│     Approved model catalogue · API/gateway · Routing · Prompt and   │
│     context assets · Evaluation · Red-team tests · Caching          │
├─────────────────────────────────────────────────────────────────────┤
│ P4  DATA, KNOWLEDGE & INTEGRATION                                   │
│     Systems of record · APIs/events/iPaaS · Warehouse when needed · │
│     ACL-correct knowledge · Metadata · Metrics · Governed memory    │
├─────────────────────────────────────────────────────────────────────┤
│ P5  IDENTITY, SECURITY & TRUST                         cross-cutting │
│     IdP · User delegation · Workload identity · Authorisation ·     │
│     Secrets · DLP · Egress · Endpoint and supply-chain security     │
├─────────────────────────────────────────────────────────────────────┤
│ P6  GOVERNANCE, OBSERVABILITY & OPERATIONS              cross-cutting│
│     AI system register · Policies · Audit · Traces · SLOs · FinOps ·│
│     Incidents · Assurance · Continuity · Benefits telemetry         │
└─────────────────────────────────────────────────────────────────────┘
```

P5 and P6 apply to every capability in P1–P4. P2 is the critical safety boundary: model output is not a transaction until deterministic controls authorise and execute it.

### 4.1 P1 — Experience and collaboration

**Purpose:** give people a coherent, secure way to use AI and supervise workflows without creating a new interface for every use case.

| Capability | Requirement | S1 | S2 | S3 |
|---|---|---|---|---|
| Primary employee assistant | One primary sanctioned pattern per user cohort, under business terms and admin control | One product for eligible staff | One primary product; specialists by evidence | Cohort-based portfolio; duplication reviewed quarterly |
| Suite-native AI | AI embedded in Workspace/M365 for mail, documents, meetings and search | Use where licence value is clear | Target eligible roles | Broad where measured; not automatic for every role |
| Embedded SaaS AI | CRM, support, finance, design, engineering and HR features | Review before enablement | Inventory and approve | Portfolio-managed; embedded agents in AI register |
| Workflow inbox | Queue for proposals, exceptions and failed cases | Existing task/chat surface | Authenticated central queue | SLA-tracked queue with routing and cover |
| Approval surface | Shows source evidence, proposed change, policy result, materiality, reversibility and alternatives | Authenticated link/card | Integrated with workflow engine | Step-up authentication and separation-of-duties where required |
| Custom UI | Thin interface over a proprietary workflow | Avoid | Rare | Selective; do not rebuild general chat |

**Positions**

- Do not automatically buy both a suite assistant and a general enterprise assistant for every employee. Run a cohort-level value and TCO decision.
- Personal or consumer accounts MUST NOT process non-public company data.
- Customer-facing AI MUST disclose that it is AI where required by law, policy or reasonable user expectation, and MUST provide a route to a person for material issues.
- Approval by email reply is not sufficient for material transactions unless the approval mechanism cryptographically binds the authenticated approver, proposal and decision.

### 4.2 P2 — Workflows, agents and action control

**Purpose:** execute repeatable business work safely and reliably.

| Capability | Description | S1 | S2 | S3 |
|---|---|---|---|---|
| Deterministic workflow engine | Triggers, states, rules, retries, timeouts and exception routing | Native SaaS/iPaaS | Managed iPaaS or durable workflow service | Standard durable workflow runtime for material paths |
| Agent runtime | Runs bounded model-assisted steps | Vendor-managed | Vendor-managed plus first code-first agents | Standard managed runtime; isolated environments by trust domain |
| Context assembler | Selects minimum authorised instructions and data, with trust and classification metadata | Product-native | Reusable component for custom workflows | Standard service/pattern |
| **Typed action broker** | Accepts declared action schema; validates, authorises, executes, verifies and records side effects | Vendor controls or no custom writes | Required for custom T2/T3 writes | Standard control boundary |
| Policy decision point | Deterministic rules on principal, action, data, value, customer, time, risk and approval authority | Simple allowlists/thresholds | Central rules per workflow | Policy-as-code or equivalent, versioned and tested |
| Tool interfaces | Native API, event, iPaaS action or MCP tool; no direct credential exposure to model | Curated vendor connectors | Registered tools and actions | Versioned tool catalogue with trust-domain boundaries |
| Durable execution | Checkpoints, idempotency, retry policy, dead-letter handling and compensation | For material flows | Required for T2 | Required for all material T2/T3 |
| Human oversight | Approval, exception review, sampling and intervention | Existing authorised role | Tooled queue and cover | SLA, workload, quality and separation-of-duties controls |
| Registry and deployment | Owner, purpose, versions, identities, tools, data, mode, tests, limits, runbook | Register in sheet/tool | Authoritative register | Register integrated with CI/change workflow |
| Pause controls | Per-workflow/action pause, credential revocation and global emergency controls | Documented/tested | Automated where possible | Automated anomaly pause plus tested recovery |

#### 4.2.1 Execution modes

Execution modes describe **what an action class may do**, not its overall legal, privacy or human-impact risk. A single workflow can contain different modes.

| Mode | Name | Permitted behaviour | Minimum controls |
|---|---|---|---|
| **T0** | Retrieve / analyse | Read permitted sources; search, classify, summarise or calculate; no external or SoR write | Effective permissions; source provenance; output labelled; logging |
| **T1** | Draft / recommend | Produce a draft, recommendation or typed proposal; an authorised person independently commits or sends it | T0 controls; representative evaluation; clear human accountability |
| **T2** | Act after approval | Execute the exact typed action approved by an authorised person | Dedicated/delegated identity; policy checks; authenticated approval; idempotent executor; verification; exception path; full action audit |
| **T3** | Autonomous within bounded policy | Execute a declared low-impact action class without per-action approval | Sustained T2/shadow evidence; hard scope/value/rate/time limits; continuous monitoring and sampling; auto-pause; reversible/compensatable action; named owner and fallback |

**T4 open-ended autonomy**—self-defined goals, unrestricted tool discovery, dynamic privilege expansion, unsupervised sub-agent creation or broad administrative access—is outside this model and prohibited by default.

#### 4.2.2 T3 ceiling

T3 SHOULD be limited to action classes that are:

- low-impact and normally non-sensitive;
- narrow and objectively testable;
- reversible or automatically compensatable;
- bounded by deterministic rules, rate and value limits;
- observable before harmful scale can accumulate; and
- operationally unnecessary to route through a person each time.

Examples: applying an internal support tag, updating a non-material CRM note field, or sending an internal reminder. Releasing payments, changing bank details, final employment decisions, legal commitments and privileged infrastructure changes are not T3 targets in this reference model.

### 4.3 P3 — Models, context and evaluation

**Purpose:** provide controlled access to models and the assets needed to assess their behaviour.

| Capability | Requirement | S1 | S2 | S3 |
|---|---|---|---|---|
| Approved model catalogue | Approved products/models, permitted data, regions, retention, use cases and owner | Tool register | Central catalogue | Version policy and deployment aliases |
| Model/API access | Enterprise terms; central ownership of API accounts and keys | Direct vendor products/API | Managed project and quotas | Gateway when trigger criteria are met |
| Gateway | Credentials, quotas, version routing, traces and cost allocation—not a false promise of model interchangeability | Usually no | Optional | Normally yes for multi-workflow/multi-provider API use |
| Task routing | Model selected by evaluated task class, latency, data and cost | Manual/product default | Workflow configuration | Policy-driven with fallback and regression evidence |
| Prompt and context assets | Versioned system instructions, schemas, examples, policy references and context rules | Controlled templates | Git or managed registry | CI/change-controlled deployment artefacts |
| Evaluation harness | Offline test sets, adversarial cases, shadow comparisons and production sampling | Manual/managed | Repeatable harness | CI-integrated plus production evaluation pipeline |
| Model-change management | Pin where possible; detect vendor updates; regression, canary and rollback | Vendor monitoring | Change calendar | Controlled aliases, canary and fallback |
| Caching | Only for semantically safe, permission-compatible and non-volatile requests | Product-native | Selective | Policy-aware; no cross-principal leakage |
| Fine-tuning | Exception only after retrieval, prompt, tool and process limits are demonstrated | No | Rare | Business case, data rights and lifecycle controls required |

**Gateway trigger criteria:** consider a gateway when at least two of the following apply: multiple production API workflows; more than one model provider; material usage spend; required central traces/quotas; data-residency routing; or frequent controlled model switching. A self-hosted gateway adds operational and security burden and MUST have an owner, SLO, patching and recovery plan.

#### 4.3.1 Evaluation dimensions

A production evaluation is multi-dimensional. Depending on the workflow it covers:

- task correctness and completeness;
- groundedness, source attribution and citation correctness;
- extraction or classification precision/recall;
- policy compliance and refusal behaviour;
- tool selection, argument validity and action success;
- prompt-injection and data-exfiltration resistance;
- fairness or differential performance where people may be affected;
- human-review acceptance, override and rework;
- latency, availability and cost per accepted outcome; and
- recovery from tool, model and data-source failure.

Twenty to fifty cases are a useful seed set, not universal production evidence. Sample size and observed runtime volume MUST match risk, variability and the tolerated error rate. Doc 02 Appendix C defines the evidence pattern.

### 4.4 P4 — Data, knowledge and integration

**Purpose:** provide authoritative, permission-correct and operationally reliable information to humans and AI systems.

| Capability | Requirement | S1 | S2 | S3 |
|---|---|---|---|---|
| Authoritative source map | System/field/document class, owner, quality, retention and permitted use | Lightweight register | Domain map | Stewarded metadata and data contracts for critical domains |
| Operational integration | Native APIs/events and deterministic workflows; writes only through supported interfaces | Native/iPaaS | Managed connectors/API | Durable services/events for material paths |
| Analytical data | Central store only where cross-system analytics, history or measurement require it | Usually native reports | Use-case-led starter warehouse | Modelled warehouse/lakehouse where justified |
| Metrics layer | Agreed definitions for business and workflow KPIs | Finance-owned sheet/model | Reusable definitions | Semantic/metrics layer for critical KPIs |
| Knowledge retrieval | Curated sources, security trimming, provenance, freshness and citations | Suite-native search/grounding | Curated assistant sources | Enterprise search/RAG only where incremental value is proven |
| Index lifecycle | Permission changes and deletions propagate; stale sources expire; index access is tested | Vendor assurance | Tested controls | Automated reconciliation and leak tests |
| Retrieval trust | Retrieved content labelled by source/trust; embedded instructions treated as untrusted data | Product controls | Workflow rules | Standard context-isolation pattern |
| Model memory | Off by default; if used, purpose, owner, ACL, retention, correction and deletion are explicit | Avoid shared memory | User/workflow-scoped | Governed store; no unbounded self-written memory |
| Data quality | Freshness, validity, reconciliation and exception checks on data used for decisions | Source controls | Checks on used datasets | Automated tests, ownership and alerting |

**Positions**

- “One system of record per entity” is a useful simplification but not always accurate. Record authority at the domain or field level where lifecycle systems differ.
- A warehouse is not an AI prerequisite. Build it when analytics, cross-system truth or measurement requires it.
- A vector index is a derived data store. It inherits privacy, residency, retention, access, deletion and incident obligations from its source content.
- Security trimming at query time is preferred. Cached ACL snapshots MUST have bounded staleness and reconciliation.
- Retrieved documents, web pages, tickets, emails and tool descriptions are data, not instructions. They cannot override system policy.
- AI-generated summaries MUST link back to authoritative sources when they may inform a material decision.

### 4.5 P5 — Identity, security and trust

| Capability | Target requirement |
|---|---|
| Human identity | Google Workspace or Entra ID as primary IdP; SSO for material systems; universal MFA; phishing-resistant authentication for privileged and financial roles; tested joiner/mover/leaver process |
| Logical AI identity | Stable `system_id`, `workflow_id` and `agent_id` in the register and every event, independent of runtime implementation |
| Delegated user access | Interactive assistants and tools preserve the requesting user’s effective downstream permissions; no service-account elevation hidden behind a user interface |
| Workload identity | Background workflows use narrowly scoped, short-lived workload credentials; static keys are avoided; production and non-production identities are separated |
| Approval identity | The approver’s authenticated identity and business authority are verified; approval does not transfer the approver’s reusable credential to the agent |
| Authorisation | Action- and resource-specific scopes; deny by default; no token passthrough; audience/issuer validation; step-up authentication for material actions |
| Secrets | Managed secrets service; no credentials in prompts, traces, source code or tool descriptions; rotation and emergency revocation |
| Network/egress | Code-first runtimes restrict outbound destinations; tools are allowlisted; web access, code execution and file handling are isolated by risk |
| DLP and minimisation | Context and output scanning where proportionate; C4 data redacted or field-minimised; no general assistant access by default |
| Supply chain | Models, frameworks, MCP servers, plugins, packages and images inventoried, reviewed, pinned where possible and monitored for change |
| Endpoint/SaaS posture | Managed devices/browsers for C3+ work; third-party OAuth grants and admin roles reviewed; risky sharing defaults disabled |

#### 4.5.1 MCP security profile

Where MCP is used for remote enterprise tools, the implementation SHOULD conform to the 2026-07-28 specification and MUST:

- require authorisation for non-public HTTP resources and tools;
- follow current OAuth security best practice, including issuer, audience and redirect validation;
- avoid passing upstream tokens through to downstream services;
- bind scopes to declared resources and actions rather than a broad “all tools” grant;
- validate and constrain every tool input, rate limit calls and sanitise outputs;
- display and approve material tool inputs before execution at T2;
- pin server/tool versions or otherwise detect material schema/description changes;
- treat tool metadata and results as untrusted content; and
- emit action events to the organisation’s audit path.

MCP does not replace application authorisation, the policy decision point or the action broker.

### 4.6 P6 — Governance, observability and operations

| Capability | Target requirement |
|---|---|
| AI system register | Covers procured, embedded and built systems; owner, supplier, purpose, data, users, action classes, risk, impact assessment, tests, versions, controls and lifecycle status |
| Policy implementation | Acceptable use, tool tiers, data rules, action limits, approvals, disclosure and retention mapped to configuration or a logged process |
| Audit events | Immutable or tamper-evident metadata sufficient to attribute and reconstruct material T2/T3 decisions and transactions |
| Diagnostic traces | Separate, access-controlled and short-lived; raw prompts/outputs retained only where necessary and redacted where possible |
| Business records | Approved outputs, sent communications and executed transactions retained under the applicable business record schedule |
| Operational SLOs | Accuracy/quality, action success, policy violations, latency, queue age, availability, cost and human-review load per workflow |
| FinOps | Seat utilisation, model/platform usage, human-review and rework cost allocated to workflows; budgets and anomaly alerts |
| Incident response | AI-specific detection and playbooks integrated with normal cyber, privacy and business incident management |
| Continuity | Manual fallback, queued-work handling, provider outage procedures, restore tests and decommissioning path |
| Assurance | Risk-based self-assessment, action reconstruction, access review, red-team testing and independent review for higher-impact systems |
| Benefits telemetry | Accepted outcomes, unit cost, capacity, quality and financial conversion evidence feed Doc 04 |

#### 4.6.1 Three different records

Do not collapse the following into one observability store:

1. **Audit event:** who/what acted, authority, policy decision, action, result, references and cost.
2. **Diagnostic trace:** detailed model/tool interaction used to debug or evaluate, with short retention and restricted access.
3. **Business record:** the communication, document, approval or transaction that the organisation must retain under normal records rules.

This separation reduces privacy and breach impact while preserving accountability.

### 4.7 Canonical T2 runtime sequence

The AP invoice example demonstrates the target pattern. Every production workflow follows the same control shape, even when the business tools differ.

```text
 1. Trigger          The AP mailbox/API creates a workflow instance and idempotency key.
 2. Authenticate     Runtime obtains the approved workload identity; the logical workflow
                     and agent IDs are attached to the trace.
 3. Classify         Input is malware-scanned and classified; sender/source trust and data
                     classes are recorded. Document content is treated as untrusted data.
 4. Retrieve         Context assembler fetches only the invoice, approved vendor/PO fields
                     and applicable policy/version through read-scoped interfaces.
 5. Interpret        Model extracts fields and proposes a typed action. It does not post.
 6. Validate         Deterministic code validates schema, arithmetic, tax rules, duplicate
                     status, vendor bank-detail status and three-way match.
 7. Decide policy    Policy service returns allow-with-approval, reject or exception based on
                     amount, match status, vendor state, authority and risk.
 8. Approve          An authorised AP approver sees source, proposed record diff, policy result,
                     exceptions and reversibility in an authenticated surface.
 9. Execute          Action broker uses a write-scoped credential to post the exact approved
                     transaction, with idempotency and timeout/retry controls.
10. Verify           Result is read back from the finance SoR; failure routes to exception and
                     any partial effect is compensated or clearly flagged.
11. Record           Audit event captures identities, versions, input/output references,
                     policy and approval decision, action/result and cost. The invoice and
                     finance record remain the business records.
12. Learn            Accepted/rejected cases, overrides, failures, review effort and unit cost
                     feed monitoring, evaluations and the benefits register.
```

At T1, steps 7–9 end with a human manually committing the draft. At T3, step 8 is omitted only for a bounded action class whose policy, evidence and limits have been approved. All other controls remain.

### 4.8 Control plane and execution plane

**Control plane:** system and agent registers; approved model/tool catalogue; identities and scopes; prompt/policy/evaluation versions; deployments; budgets; approvals; kill/pause controls; and audit configuration.

**Execution plane:** workflow instances; model calls; retrieval; tool invocations; transactions; queues; exceptions and results.

Administrative access to the control plane MUST be more restricted than ordinary workflow operation. A compromised model or workflow MUST NOT be able to modify its own policy, tools, scopes, tests, budgets or audit configuration.

## 5. Reference deployment patterns by size

### 5.1 S1 — Micro / very small

Minimum viable architecture:

- one primary sanctioned employee assistant or suite-native AI pattern;
- business accounts, MFA, central billing and an acceptable-use policy;
- native SaaS automation or managed iPaaS for deterministic flows;
- no custom agent with write access unless a specific, measured workflow justifies support and control overhead;
- lightweight AI system/source/benefit registers;
- password manager/secrets service and documented joiner/leaver process;
- vendor-managed retrieval over carefully selected sources; and
- manual fallback for every automation.

Avoid a warehouse, self-hosted gateway, self-hosted vector stack or multi-agent framework unless a clear workload pays for it.

### 5.2 S2 — Small

Add when justified:

- a managed workflow/agent runtime and reusable action-control pattern;
- first read-only and T1 agents, then T2 action classes after evidence;
- central log destination and repeatable evaluation harness;
- dedicated workload identities and managed secrets;
- curated knowledge sources with tested security trimming;
- starter analytical store only for cross-system reporting/measurement use cases;
- a partner or fractional engineer who owns runbooks and support; and
- formal monthly portfolio, risk and benefit review.

### 5.3 S3 — Medium / lower mid-market

Add:

- standard durable workflow runtime, action broker and policy pattern;
- versioned tool catalogue and bounded MCP estate where useful;
- model gateway when the trigger criteria in §4.3 are met;
- central model, prompt, policy and evaluation release process;
- broader analytical/metrics layer where business use requires it;
- structured SLOs, on-call/escalation and incident exercises;
- automated access, permission, registry and spend checks; and
- independent review for higher-impact workflows.

## 6. Vendor-aligned examples

These are coherent examples, not mandatory stacks. Features and commercial terms MUST be revalidated at procurement.

### 6.1 Google-anchored estate

| Capability | S1 | S2 | S3 |
|---|---|---|---|
| Identity/collaboration | Google Workspace | Workspace + automated lifecycle to major SaaS | Workspace + formal access reviews and central log export |
| Employee AI | Workspace with Gemini **or** selected business assistant as primary | Primary cohort pattern; ChatGPT Business/Enterprise or Claude Team/Enterprise only where incremental | Cohort portfolio; suite and specialist assistants rationalised quarterly |
| Automation/workflows | Native SaaS, Apps Script where supportable, managed iPaaS | Managed iPaaS / Cloud Run or equivalent for bounded code | Durable workflow service + Cloud Run/container runtime or equivalent |
| Agent building | Vendor-managed builders only | Managed builders + first code-first workflow | Code-first standard for material workflows; builders for low-risk citizen use |
| Model APIs | Direct managed project | Central project, quotas and logging | Gateway/API management where trigger criteria apply |
| Data/analytics | Native reports | BigQuery or equivalent only when use cases require | BigQuery + tested transformations/metrics where justified |
| Knowledge | Workspace-native grounding on curated sources | Curated enterprise assistant sources | Enterprise search/RAG only if native capability is insufficient and ACL controls are proven |
| Secrets/logging | Business password manager; admin logs | Secret Manager + Cloud Logging or equivalent | Central security/observability projects, retention tiers and dashboards |

### 6.2 Microsoft-anchored estate

| Capability | S1 | S2 | S3 |
|---|---|---|---|
| Identity/collaboration | Microsoft 365 + Entra ID | Entra lifecycle to major SaaS | Entra access reviews/entitlement controls and central log export |
| Employee AI | Microsoft 365 Copilot **or** selected business assistant as primary | Primary cohort pattern; alternate assistant only where incremental | Cohort portfolio; duplication and agent proliferation reviewed quarterly |
| Automation/workflows | Power Automate/native SaaS | Power Automate/Copilot Studio plus bounded Azure runtime as needed | Durable workflow/runtime on managed Azure services; Copilot Studio for appropriate low-code cases |
| Agent building | Vendor-managed builders | Managed builders + first code-first workflow | Code-first standard for material workflows; low-code under environment and DLP controls |
| Model APIs | Direct managed project | Central subscription/project, quotas and logging | Gateway/API management where trigger criteria apply |
| Data/analytics | Native reports/Power BI | Fabric or equivalent only when use cases require | Modelled Fabric/lakehouse and governed metrics where justified |
| Knowledge | M365/Graph grounding on curated sources | Curated SharePoint/enterprise sources | Enterprise search/RAG only if incremental value and security trimming are proven |
| Secrets/logging | Business password manager; audit logs | Key Vault + Log Analytics or equivalent | Central security/observability workspaces, retention tiers and dashboards |

### 6.3 Neutral assets to own

Regardless of vendor, the organisation SHOULD own or be able to export:

- workflow definitions and typed action schemas;
- prompt, policy and context templates;
- evaluation cases, rubrics and results;
- source and authority maps;
- AI system/agent registry records;
- benefits baselines and outcome telemetry;
- architecture decisions and runbooks; and
- business data in supported, documented formats.

## 7. Integration decision hierarchy

For each integration or workflow need, use this sequence:

1. **Eliminate or simplify the process.**
2. **Use a native capability** in the authoritative application where it meets the control and economic need.
3. **Use a supported API and event/webhook** for deterministic application behaviour.
4. **Use iPaaS or a durable workflow service** for cross-application rules and state.
5. **Expose a bounded tool interface, including MCP where appropriate,** when a model-assisted host needs to select or use a tool.
6. **Use controlled computer/browser automation** only when no supported interface exists, for a time-limited exception with strong isolation and monitoring.
7. **Write custom point-to-point code** only with a named owner, tests, SLO, runbook and exit plan.

**Prohibited anti-patterns:** human credential reuse; agent-admin accounts; raw model output into SQL/shell/HTML or transactions; runtime discovery of unapproved tools; unsupervised browser automation for sensitive work; mega-tools exposing an entire SaaS API; indexing without effective ACLs; and approvals that do not bind the exact proposed action.

## 8. Build versus buy

Build a custom production component only when all of the following are true:

1. the workflow is strategically differentiating or commercially material;
2. the business volume or risk justifies engineering ownership;
3. configurable products cannot meet the control and outcome requirements at acceptable TCO;
4. the process and authoritative data are sufficiently stable;
5. an accountable business owner and technical owner exist;
6. the organisation can test, secure, monitor, support and retire the component; and
7. the risk-adjusted business case beats buying, simplifying or not doing it.

Even then, build the thinnest differentiating layer. Prefer managed models, identity, queues, secrets, databases and observability.

Record significant decisions as ADRs. A “proof of concept” is not exempt from data, security or contractual controls if it uses real C3/C4 data or can cause a real side effect.

## 9. Non-functional requirements

Targets are set **per workflow**, based on impact and business criticality.

| Area | Baseline requirement |
|---|---|
| Correctness/quality | Acceptance criteria and error budget; objective checks where possible; model output alone is not evidence |
| Safety/policy | Zero tolerance for unauthorised action classes; policy violations auto-pause material workflows |
| Availability | SLO, business hours/24×7 need, provider dependency and degraded/manual path documented |
| Recovery | RTO/RPO where relevant; retry, dead-letter, compensation and restoration tested |
| Auditability | 100% of T2/T3 actions attributable to logical system, runtime principal, policy version and approver/authority |
| Privacy | Purpose limitation, minimum context, data rights, retention, deletion and residency appropriate to the use |
| Security | Threat model; least privilege; isolated untrusted content; egress/tool restrictions; secure software and supply-chain practices |
| Performance | Latency and queue-age targets; slow model paths do not block critical operations without fallback |
| Cost | Maximum cost per accepted outcome; seat and usage budgets; anomaly alerting; human-review/rework included |
| Accessibility | Human-facing AI and approval surfaces meet applicable accessibility requirements |
| Explainability/evidence | Sources, material factors and policy results available at the level appropriate to impact; no uncalibrated confidence score presented as assurance |
| Maintainability | Owner, versions, test suite, deployment/rollback, runbook and decommissioning path |
| Portability | Data and control assets exportable; provider changes subject to regression rather than assumed compatibility |

## 10. Maturity model

Maturity measures controlled business operation, not the number or autonomy of agents.

| Level | Name | Characteristics |
|---|---|---|
| **L1** | Ad hoc | Shadow or individual AI use; no reliable inventory, policy or outcome evidence |
| **L2** | Sanctioned | Approved employee tools, basic policy, ownership, AI literacy and initial data controls |
| **L3** | Measured | Curated knowledge; several T0/T1 workflows with baselines, evaluations, unit costs and monitoring |
| **L4** | Orchestrated | Production T2 workflows across functions; standard action boundary; reliable operations, assurance, portfolio and benefit conversion; selective bounded T3 |
| **L5** | Adaptive operating model | Functions are deliberately designed around human+AI workflows; closed-loop process improvement, role/capacity planning and investment decisions use production evidence; autonomy remains risk-bounded |

**Month-24 target:** L4 for S2/S3, with L5 characteristics in selected functions. S1 targets the economically appropriate subset of L3/L4 controls, not a miniature enterprise platform.

## 11. Target operating model

### 11.1 Accountabilities

| Role/duty | Accountability |
|---|---|
| Executive Sponsor | Strategy, risk appetite, funding, workforce principles and material exceptions |
| AI/Automation Lead | Portfolio, reference architecture, operating cadence and cross-functional capability |
| AI System Owner | End-to-end accountability for a deployed AI system, including supplier, controls, tests, monitoring and retirement |
| Workflow Owner | Business outcome, process design, backlog, acceptance criteria, unit cost and benefit evidence |
| Technical Owner | Runtime, integrations, action controls, deployment, SLOs, incidents and runbook |
| Data/Knowledge Owner | Authority, classification, access, quality, source curation, retention and correction |
| Risk/Privacy/Security Owner | Proportionate risk/impact assessment, control review, legal/privacy coordination and incident oversight |
| Authorised Approver | Makes a specific business decision at T2 within delegated authority; not merely an “AI supervisor” |
| Review/Sampling Duty | Reviews production samples and exceptions; separate from approval where independence is required |
| Finance Partner | Baselines, TCO, financial conversion and sign-off of banked benefits |

S1/S2 may combine duties, but accountability remains explicit. Suppliers can perform work; they do not absorb the organisation’s accountability.

### 11.2 Human+AI team design

Humans concentrate on judgment, relationships, exceptions, accountability and process improvement. AI systems perform retrieval, interpretation, drafting and bounded execution where evidence supports it.

Do not set supervision capacity by “number of agents”. Size it using:

- approval and exception volume;
- minutes and cognitive effort per decision;
- queue-age and service-level requirements;
- sampling burden and error rate;
- materiality and separation-of-duties needs; and
- cover for leave and incidents.

A queue that forces cursory approvals is a failed control. Sustained very-fast approvals, high approval rates without variation, backlog growth, low override rates inconsistent with known model error, or missed evidence review trigger redesign or demotion.

### 11.3 Managerial changes

Managers:

- own workflow output and exception quality rather than monitoring “AI usage” in isolation;
- decide where released capacity is explicitly redeployed;
- maintain human expertise and fallback capability for material work;
- ensure affected employees receive training and role-change clarity; and
- treat AI systems as governed production capacity with owners, budgets, SLOs and performance history.

## 12. Month-24 outcome statement

A successful S3 implementation should normally demonstrate:

- one primary sanctioned employee-assistant pattern per relevant cohort, with licence value and use measured;
- an organisation-wide AI system register including embedded and supplier-provided AI;
- curated, permission-correct knowledge sources with provenance, freshness and deletion controls;
- 5–15 production AI-enabled workflows across at least three functions, scaled down by band and business need;
- a majority of material write actions at T2, with T3 confined to approved bounded action classes;
- a standard deterministic action-control pattern for custom T2/T3 workflows;
- every production workflow owned, versioned, evaluated, monitored, budgeted, supportable and pausable;
- outcome quality, unit cost, human-review effort and total cost reported monthly;
- capacity and financial benefits reported separately and finance-validated; and
- at least two functions redesigned around human+AI workflows rather than merely automating legacy steps.

The target is not “the majority of all work is autonomous”. It is that AI-enabled work is commercially useful, technically supportable and governed as normal operations.

## 13. Sector and jurisdiction overlays

The six-plane model remains useful in regulated contexts, but additional obligations can materially change implementation, including:

- deployment location and supplier eligibility;
- segregation and key management;
- validation depth and independence;
- audit and records retention;
- human decision requirements and contestability;
- model explainability and transparency;
- incident notification; and
- prohibited or high-risk uses.

Do not describe these as governance-only “bolt-ons”. For health, finance, public sector, defence, critical infrastructure, children’s data or safety-relevant use, commission a dedicated architecture and legal/control overlay before selecting products or autonomy ceilings.

## 14. Architecture decisions to record

| ADR | Decision |
|---|---|
| ADR-1 | Identity and collaboration estate; SSO and lifecycle exceptions |
| ADR-2 | Primary employee-assistant pattern by cohort; specialist-product criteria |
| ADR-3 | AI system register and ownership model |
| ADR-4 | Action-control pattern: proposal schema, policy decision point, broker/executor and approval surface |
| ADR-5 | Delegated-user versus workload-identity patterns; credential lifetime and scope |
| ADR-6 | Model providers, approved versions, contractual posture and change-notification approach |
| ADR-7 | Gateway trigger and product/operating model |
| ADR-8 | Workflow/agent runtime and durable execution standard |
| ADR-9 | Integration hierarchy and MCP security/profile/version policy |
| ADR-10 | Authoritative source map, analytical platform trigger and metrics layer |
| ADR-11 | Knowledge/RAG sources, security trimming, freshness, deletion and provenance approach |
| ADR-12 | Model memory policy |
| ADR-13 | Evaluation framework, sample/evidence rules and production SLOs |
| ADR-14 | Execution-mode ceilings and prohibited action classes |
| ADR-15 | Audit, trace and business-record destinations and retention |
| ADR-16 | Privacy, residency, disclosure, affected-person and legal obligations |
| ADR-17 | Business continuity, provider fallback and manual operating paths |
| ADR-18 | Build/buy and exit strategy for custom components |

### 14.1 ADR template

```text
ADR-<n>: <title>
Status:       Proposed | Accepted | Superseded-by-<n>
Date:         <yyyy-mm-dd>
Owner:        <accountable role>
Review:       <date and event-based triggers>

Context:
  What outcome, constraints, risks, volumes and obligations drive the decision?

Options:
  2–4 credible options, including “do nothing/simplify”, with TCO and control trade-offs.

Decision:
  The selected option and its scope, stated unambiguously.

Controls and evidence:
  Required tests, limits, approvals, monitoring, fallback and acceptance criteria.

Consequences:
  Commitments, residual risks, lock-in and operating ownership.

Revisit triggers:
  Headcount/volume threshold, supplier or law change, incident, failed SLO, economics,
  material model/tool change or new capability.
```

ADRs, prompts, policies, action schemas, evaluations and runbooks SHOULD be versioned together. They are the operational memory of the architecture.


---

# 02 — Phased Implementation Plan

| | |
|---|---|
| **Version** | v1.0 (Final) |
| **Date** | 2026-08-10 |
| **Planning horizon** | Up to 24 months; outcome-gated rather than calendar-driven |
| **Companion docs** | 00 Overview · 01 Target Architecture · 03 Data, Security & AI Governance · 04 Benefits Realisation |

---

## 1. Delivery approach

This is a capability and operating-model program, not a one-off technology implementation.

1. **Run a portfolio, not a parade of pilots.** Every candidate is selected, funded, gated, operated and retired through one process.
2. **Redesign before automation.** For each workflow ask: can the work be eliminated, simplified, standardised or handled deterministically before adding a model?
3. **Separate program maturity from workflow maturity.** The organisation can be in Phase 2 while an individual workflow remains T0/T1 or is retired.
4. **Use thin, end-to-end slices.** Deliver one measurable outcome with identity, controls, evaluation, operations and benefits—not a broad demonstration with no production path.
5. **Promote action classes, not whole agents.** T2/T3 evidence applies to a declared action with fixed scope and limits.
6. **Adoption is necessary, not sufficient.** Usage is monitored, but only accepted outcomes and converted benefits justify scale.
7. **Fund in stages.** Release larger platform and licence commitments only when the portfolio demonstrates demand and value.
8. **Plan for failure and retirement.** Manual fallback, pause, rollback, supplier exit and decommissioning are designed before production.
9. **Treat workforce trust as an operating dependency.** Communicate role effects and measurement principles before broad deployment.

Phases below are indicative. Readiness, integration complexity, legal obligations and organisational change capacity determine duration. S1 organisations may reach an appropriate steady state in 6–12 months; S3 commonly uses the full horizon.

## 2. Program governance

| Role/duty | Accountability | Typical incumbent |
|---|---|---|
| Executive Sponsor | Strategy, funding, risk appetite, workforce principles and material exceptions | CEO, COO or relevant executive |
| AI/Automation Lead | Portfolio, architecture adherence, delivery cadence, register and cross-functional capability | Fractional in S1/S2; dedicated in S3 |
| AI Council / Steering Group | Approves higher-impact use cases, T3 action classes, material exceptions and portfolio investment/retirement | Sponsor, AI Lead, IT/security, privacy/legal as needed, finance and rotating function owner |
| Function Product Owner | Prioritises outcomes, process redesign, adoption and benefit evidence | Function head or delegate |
| Workflow Owner | Owns one production workflow’s acceptance criteria, controls, unit cost and lifecycle | Existing operational manager or product owner |
| Technical Owner | Build/configure, integrations, action boundary, SLOs, deployment, support and incident response | Internal engineer, IT lead or contracted partner |
| Data/Knowledge Owner | Authority, access, quality, source curation, retention and correction | SoR or corpus owner |
| Security/Privacy/Legal Owner | Proportionate control and impact review; obligations and incidents | Existing leads or external adviser/MSP |
| Finance Partner | Baselines, TCO, benefit validation and financial conversion | CFO/finance lead |
| Employee/Change Lead | AI literacy, role guidance, consultation, champions and feedback | People lead or change duty-holder |

**Decision cadence**

- weekly or fortnightly delivery review for active workflows;
- monthly portfolio, risk, cost and incident review;
- quarterly investment and benefits-conversion review; and
- annual strategy, policy, risk appetite and supplier review.

S1 may combine these into one monthly operating review. Governance effort should scale with risk and portfolio size, not reproduce enterprise bureaucracy.

## 3. Workstreams

| WS | Workstream | Scope |
|---|---|---|
| WS1 | Strategy & Portfolio | AI intent, inventory, workflow selection, process redesign, investment and retirement |
| WS2 | Identity & Security | SSO/MFA/lifecycle, delegated and workload identity, secrets, endpoint, egress, action security and logging |
| WS3 | Data & Knowledge | Authority map, permission hygiene, source registry, retrieval controls, data quality, metrics and analytical data where needed |
| WS4 | Platform & Integration | Employee tools, deterministic workflow engine, agent runtime, action broker/policy, tool interfaces, gateway and observability |
| WS5 | Workflow Delivery | Design, build/configure, evaluate, shadow, deploy, operate and improve individual workflows |
| WS6 | People & Adoption | AI literacy, role-based practice, manager guidance, champions, support, workforce dialogue and accessibility |
| WS7 | Governance & Assurance | Policy, AI system register, risk/impact assessments, supplier review, change control, incidents, contestability and assurance |
| WS8 | Benefits & FinOps | Baselines, TCO, capacity/outcome ledger, financial conversion, unit economics and investment review |

## 4. Two-dimensional delivery model

### 4.1 Program phases

The organisation progresses through four maturity phases:

- **Phase 0 — Establish control and evidence**
- **Phase 1 — Governed augmentation**
- **Phase 2 — Controlled execution**
- **Phase 3 — Institutionalise and selectively automate**

### 4.2 Workflow lifecycle

Every workflow follows this lifecycle regardless of program phase:

```text
Discover → Eliminate/simplify → Define contract → Risk/impact assess →
Build/configure → Offline evaluate → Shadow → T0/T1 production →
T2 controlled action (if justified) → T3 bounded action (exceptional) →
Operate/improve → Retire
```

A workflow can be killed at any gate. Failed experiments are useful only if they are closed, documented and stop consuming attention or licences.

## 5. Phase 0 — Establish control and evidence (nominal Months 0–3)

**Objective:** know what AI is already present, make the estate safe enough to amplify, establish accountable ownership, capture baselines and select the first evidence-producing portfolio.

### 5.1 Activities

**WS1 — Strategy & Portfolio**

- State the business intent and explicit non-goals; confirm that unrestricted autonomous operation is out of scope.
- Inventory employee tools, embedded SaaS AI, custom automations, supplier AI and known shadow use.
- Complete the readiness assessment (Appendix B).
- Map 10–20 candidate task families; run the eliminate/simplify/deterministic screen before scoring AI candidates.
- Select 2–4 Phase 1 workflows, including at least one low-risk, high-frequency use case with objective acceptance criteria.

**WS2 — Identity & Security**

- Audit SSO, MFA, privileged access, joiner/mover/leaver and external OAuth grants.
- Require universal MFA; move privileged and financial roles towards phishing-resistant authentication.
- Close shared-account and ex-employee access; document the unsupported SaaS long tail.
- Establish managed secrets and the delegated-user versus workload-identity patterns.
- Choose the audit destination and minimum event schema.

**WS3 — Data & Knowledge**

- Lock down risky external-sharing defaults and remediate the highest-risk links/groups/sites.
- Define the local data classification scheme and owners.
- Create the authoritative source and AI-grounding source registers.
- Identify personal information, C4/regulated data and retention constraints in candidate workflows.
- Do not index broad corpora merely to demonstrate search.

**WS4 — Platform & Integration**

- Select one primary employee-assistant pattern per initial cohort under appropriate business terms.
- Select a deterministic automation/workflow platform suited to existing skills and estate.
- Define the standard typed-action, policy and approval pattern before custom write-enabled agents.
- Record the initial ADRs in Doc 01 §14.
- Do not procure a warehouse, model gateway or custom RAG stack without a use-case trigger.

**WS5 — Workflow Delivery**

- Define workflow contracts: outcome, trigger, source truth, action classes, quality, mode ceiling, exception path, SLO and unit cost.
- Build representative evaluation seed sets from real historical cases.
- Establish kill criteria before build.

**WS6 — People & Adoption**

- Publish the acceptable-use policy and a no-blame shadow-AI disclosure period.
- Deliver foundation AI literacy: capability limits, privacy, security, verification, disclosure and incident reporting.
- Explain what activity/outcome data will and will not be used for; workflow measurement is not covert individual surveillance.
- Nominate champions by team/function where useful, not by a rigid ratio.

**WS7 — Governance & Assurance**

- Establish the AI system register and use-case risk/impact process.
- Identify legal and contractual obligations, including privacy, employment, records, IP and customer requirements.
- Approve tool tiers, prohibited/default-excluded uses, supplier checklist and incident playbooks.
- Document any imminent legal change that affects the program and assign an owner.

**WS8 — Benefits & FinOps**

- Capture workflow baselines, quality/rework, volumes and current unit costs.
- Inventory AI, SaaS, external-services and delivery costs.
- Establish separate capacity/outcome and financial ledgers (Doc 04).
- Define how model, platform, human-review and rework cost will be allocated to workflows.

### 5.2 Deliverables

- approved program charter and risk appetite;
- current AI/system inventory and initial register;
- readiness assessment and remediation plan;
- policy set v1 and legal/obligations register;
- identity, permission and external-sharing remediation report;
- sanctioned employee-tool decision and supplier terms;
- architecture decision log and action-control design;
- baseline/TCO pack; and
- Phase 1 portfolio briefs with workflow contracts, evaluation and kill criteria.

### 5.3 Gate 0 → 1

All applicable criteria are required:

- [ ] Executive Sponsor, AI Lead and workflow owners are named with time and authority.
- [ ] Material AI systems and embedded AI are inventoried; shadow disclosure period completed.
- [ ] MFA is universal; privileged access and leaver controls are tested; critical shared accounts are removed.
- [ ] Highest-risk permission/sharing findings are remediated or have time-bound accepted treatment.
- [ ] Data classification, source ownership and initial legal/privacy obligations are documented.
- [ ] Sanctioned tools operate under reviewed business terms and central administration.
- [ ] Policy, risk/impact intake, incident route and AI system register are live.
- [ ] Local baselines and TCO assumptions exist for all Phase 1 workflows.
- [ ] Each selected workflow has a kill criterion, owner, source truth, manual fallback and mode ceiling.

A target such as “95% SSO coverage” can guide remediation, but gate approval is risk-based: all material systems MUST have acceptable identity and lifecycle controls, even if a harmless long-tail application remains outside SSO.

## 6. Phase 1 — Governed augmentation (nominal Months 3–9)

**Objective:** establish useful employee adoption and 3–5 measured T0/T1 workflows, prove source and evaluation controls, and learn where AI genuinely improves work.

### 6.1 Activities

- Deploy the primary assistant to defined eligible cohorts; use staged licence allocation and reclaim unused seats.
- Curate knowledge sources and test effective permissions, source provenance, freshness and deletion using low-privilege identities.
- Deliver 3–5 T0/T1 workflows across no more than 2–3 functions initially.
- Use models for ambiguous interpretation/drafting and deterministic code for calculations, validation and policy.
- Operate material workflows in shadow mode before relying on outputs.
- Implement repeatable evaluations and production telemetry: quality, acceptance, override, rework, latency and cost.
- Establish a workflow support queue and named fallback.
- Run role-based training using the actual tools and workflows, followed by office hours/guilds and manager coaching.
- Start monthly portfolio and quarterly benefits reviews; retire licences and experiments that fail value or adoption thresholds.
- For S2/S3, introduce the first read-scoped tool/MCP interfaces only where native integrations do not suffice.
- Build an analytical store only where a selected workflow or benefit baseline requires cross-system history.

### 6.2 Gate 1 → 2

- [ ] Eligible-cohort adoption is sufficient to assess value; unused licences are being reclaimed rather than hidden by blanket targets.
- [ ] At least three workflows have representative evaluation results and measured outcome/quality change versus baseline.
- [ ] Source-permission leak tests pass; citations/provenance are available for material knowledge outputs.
- [ ] All live systems are registered, owned, risk/impact-assessed and supported.
- [ ] Unit cost includes model/platform, human-review and rework effort.
- [ ] No unresolved critical/high incident or overdue control treatment blocks scale.
- [ ] At least one workflow has been killed, narrowed or redesigned where evidence did not justify continuation—or the council has documented why all survived.
- [ ] Candidate T2 action classes pass the production-readiness checklist in Appendix D and have completed shadow operation.

No organisation should interpret “high assistant weekly active use” as a universal objective. The target is sustained, appropriate use among eligible roles that produces accepted outcomes at acceptable TCO.

## 7. Phase 2 — Controlled execution (nominal Months 9–18)

**Objective:** move selected action classes from drafts to T2 execution, harden runtime and operational controls, and scale only workflows with demonstrated value.

### 7.1 Activities

**Architecture and operations**

- Implement the standard action broker/policy/approval pattern for custom write-enabled workflows.
- Use dedicated short-lived workload identity or preserved delegated-user identity as appropriate.
- Add durable execution: idempotency, retries, timeouts, dead-letter handling, verification and compensation.
- Implement authenticated approvals with evidence, record diff, authority checks and separation of duties where required.
- Establish SLOs, error budgets, alerting, on-call/escalation, pause/recovery and provider-outage procedures.
- Expand central model/API controls or introduce a gateway only when Doc 01 trigger criteria are met.
- Restrict runtime egress, code execution and tool access; pin and review MCP/tool versions.

**Workflow portfolio**

- Promote proven action classes to T2; do not promote an entire agent wholesale.
- Add workflows in finance, support, sales, operations and engineering only where local data/process maturity permits.
- Prefer action classes with clear ground truth, high frequency, bounded scope and reversible writes.
- Run canary deployment, intensified sampling and rollback for model, prompt, tool or policy changes.
- Track approval queue load and quality; redesign if approvers become a bottleneck or rubber stamp.

**People and benefits**

- Train authorised approvers and reviewers on specific failure modes, evidence and intervention.
- Make role and capacity changes explicit; allocate released capacity to named priorities.
- Convert benefits only through Doc 04 evidence rules; finance signs off cash impacts.
- Consolidate software or external spend only after the replacement capability has met SLOs for a sustained period.

### 7.2 Gate 2 → 3

- [ ] At least 2 (S1), 4 (S2) or 5 (S3) production workflows include a T2 action class with accepted quality and economics; volume is adjusted to actual business need.
- [ ] Every T2 action is attributable, policy-checked, bound to an authenticated approval and verified in the SoR.
- [ ] Production identities, scopes, tools, budgets, versions, runbooks and pause procedures match the register.
- [ ] SLOs, error budgets, human-review load and cost are reported monthly.
- [ ] Incident and restore/pause exercises have passed; manual fallback remains usable.
- [ ] Benefits reported as financial are finance-validated and not double counted with capacity.
- [ ] Each proposed T3 action class meets the additional criteria in Appendix D, including low impact, bounded scope, sustained evidence and auto-pause.

## 8. Phase 3 — Institutionalise and selectively automate (nominal Months 18–24)

**Objective:** embed the operating model, redesign selected end-to-end processes, permit only justified T3 action classes, and establish a sustainable Year-3 portfolio.

### 8.1 Activities

- Promote a small number of eligible action classes to T3; keep sensitive, material, rights-affecting and irreversible actions at T1/T2 or outside AI execution.
- Redesign 2–3 end-to-end processes around human judgment, deterministic controls and AI-enabled work rather than copying legacy steps.
- Consolidate redundant AI, automation and point products based on proven replacement and exit evidence.
- Run role-evolution, reskilling and internal-mobility plans where task compression is durable.
- Conduct structured assurance against the organisation’s selected baseline (for example NIST AI RMF and the Australian six essential practices); obtain independent review for higher-impact systems.
- Complete model/tool/provider exit tests for critical workflows.
- Refresh risk appetite, policies, legal obligations, supplier register, evaluation sets and architecture decisions.
- Approve Year-3 operating budget and portfolio based on accepted outcomes, TCO and risk—not forecast enthusiasm.

### 8.2 Month-24 exit criteria

- [ ] Maturity L4 is achieved for the relevant scale; selected functions show L5 operating characteristics.
- [ ] The AI system register, source register, action controls and change records are complete and current.
- [ ] T2/T3 action spot-checks reconstruct identity, policy, evidence, approval, execution and result end-to-end.
- [ ] All T3 action classes remain within the Doc 01 ceiling and meet current SLO/error-budget evidence.
- [ ] Manual fallback, supplier outage and decommissioning paths have been tested for material workflows.
- [ ] Capacity, service/quality and financial benefits are reported separately; TCO and net cash effect reconcile to finance.
- [ ] Steady-state roles, support, assurance, training and budgets are approved.

## 9. Workflow selection framework

### 9.1 Mandatory decision sequence

For every task family:

1. **Eliminate:** Is the output or control still needed?
2. **Simplify:** Can policy, form design, data quality or process ownership remove the work?
3. **Standardise:** Can templates, master data or a single SoR reduce variation?
4. **Deterministic automation:** Can rules, APIs, OCR or conventional software solve it more reliably?
5. **AI assistance:** Does language, judgment under ambiguity or unstructured content justify a model?
6. **Agentic execution:** Does allowing the system to choose/use tools create additional net value after control and support cost?

### 9.2 Weighted score

Score 1–5 and multiply by the weight.

| Criterion | Weight | What a high score means |
|---|---:|---|
| Outcome value | 20% | High volume/cost, service or risk impact; clear unit of outcome |
| Process suitability | 15% | Stable enough to define, but contains ambiguity AI can help resolve |
| Data/source readiness | 15% | Authoritative, accessible, lawful, representative and sufficiently clean inputs |
| Control and reversibility | 15% | Actions are bounded, testable, reversible/compensatable and can be policy-checked |
| Risk/impact fit | 15% | Low/moderate impact with proportionate controls; no default-excluded use |
| Owner/adoption pull | 10% | Named owner and users actively want the outcome and can change the process |
| Repeatability/scale | 10% | Frequent enough to support evidence and repay implementation/operation |

Apply a **complexity penalty** for open-ended goals, many tools, internet/email exposure, long plans, dynamic sub-agents, browser automation, C4 data, weak ground truth, low frequency or material external commitments.

High score does not override a prohibited use or legal obligation. The initial portfolio SHOULD contain mostly T0/T1 workflows and deterministic automation, not a quota of agents.

### 9.3 Portfolio balance

Maintain a mix of:

- **employee capability:** assistants and training;
- **knowledge access:** curated T0 search/analysis;
- **operational workflows:** T1/T2 with measurable units;
- **foundation work:** permissions, source quality, identity and integration; and
- **experiments:** small, time-boxed and explicitly disposable.

Do not let platform work consume the program without workflow demand, or let workflow pilots bypass platform controls.

## 10. Indicative timeline

```text
Month:      0──3             3──9                    9──18                    18──24
Program:    CONTROL          AUGMENT                  CONTROLLED EXECUTION     INSTITUTIONALISE
            inventory        employee capability     T2 action classes       selective T3
            identity/data    T0/T1 evidence           durable operations      process redesign
            policy/baseline  curated knowledge        action boundary         assurance/steady state

Per workflow:
            discover/simplify → contract/risk → evaluate/shadow → T0/T1 → T2 → T3 if justified → retire
```

Calendar progress never substitutes for gate evidence.

## 11. Resourcing by size band

Figures are indicative delivery capacity, not a requirement to create separate jobs. Existing employees, an MSP and specialist partners can supply the capacity, but internal accountability cannot be outsourced.

| Capability | S1 (1–20) | S2 (21–100) | S3 (101–500) |
|---|---|---|---|
| Sponsor / business ownership | Founder/COO duty | Executive + function owners | Executive sponsor + portfolio owners |
| AI/Automation Lead | 0.1–0.3 FTE, combined role | 0.3–0.7 FTE | ~1.0 FTE |
| Automation/agent engineering | Partner bursts; avoid standing platform | 0.5–1.0 FTE equivalent internal/partner | 1–3 FTE depending on portfolio/complexity |
| Data/analytics | Existing analyst/finance; use-case only | 0.2–0.5 FTE | 0.5–1.5 FTE |
| Security/privacy/legal | MSP/adviser as needed | 0.1–0.3 FTE equivalent | 0.3–1.0 FTE across existing specialists |
| Change/enablement | Owner + team champions | 0.1–0.3 FTE | 0.3–0.7 FTE |
| Finance/benefits | Existing finance duty | 0.1–0.2 FTE | 0.2–0.4 FTE |

Do not add the top of every range to create an assumed headcount requirement. Workload volume, supplier mix, code ownership and risk determine the actual shape.

## 12. Budget and commercial controls

Build the budget by category rather than using a universal “AI cost per employee”:

- employee-assistant licences and specialist seats;
- API/model usage and reserved/committed spend;
- automation, agent, gateway, search/RAG and observability platforms;
- data integration/storage/analytics required by selected workflows;
- implementation and integration;
- internal product/engineering/data/security/change capacity;
- evaluation, red-team and independent assurance;
- support, incident response and business continuity;
- training and workforce transition; and
- exit/decommissioning.

Commercial guardrails:

- stage seat purchases and reclaim unused licences;
- avoid multi-year volume commitments before measured use unless the discount and exit terms justify them;
- require model/version-change, incident, subprocessor, deletion, export and audit-log terms appropriate to the use;
- distinguish subscription entitlements from metered agent/API usage;
- cap cost per workflow and cost per accepted outcome;
- include human approval, rework and exception handling in unit economics; and
- fund platform components only against a multi-workflow demand case or mandatory control need.

## 13. Principal program risks

| Risk | Response |
|---|---|
| Permission debt becomes an internal data breach | Phase 0 remediation; curated sources; security trimming and leak/deletion tests before grounding |
| Employee tools proliferate faster than governance | One primary cohort pattern; inventory embedded AI; reclaim seats; DLP and procurement controls |
| Agentic solution chosen for a rules problem | Mandatory eliminate/simplify/deterministic screen and architecture review |
| Model output bypasses business controls | Typed proposal, deterministic policy and action broker; no raw output to transactions |
| Prompt injection or poisoned content drives action | Trust-labelled context, tool limits, isolation, output validation, T2 approval and T3 ceiling |
| Pilot purgatory | Workflow contract, production owner, gate dates and kill criteria at intake |
| “Golden set” overstates production reliability | Representative/adversarial tests, holdout, shadow mode, production SLOs and sampled review |
| Human approval becomes a rubber stamp | Evidence-rich UI, authority checks, queue/load metrics, review quality triggers and demotion |
| Model/provider change silently degrades workflows | Version monitoring/pinning, regression, canary, fallback and rollback |
| Token or licence cost expands without outcome | Per-workflow allocation, budgets, seat reclamation and cost-per-accepted-outcome limits |
| Benefits double counted | Separate capacity and financial ledgers; finance validates conversion |
| Workforce trust collapses | Early disclosure, no covert individual surveillance, clear redeployment principles, training and consultation |
| Key-person/partner dependency | Owned artefacts, runbooks, source access, deployment rights, exit test and secondary cover |
| Custom platform overwhelms SME capacity | Size-band defaults, managed foundations and build/buy gate |
| Legal obligations emerge after design | Obligations register, material-change review and privacy/legal participation at intake |

## 14. Sequencing rules

1. Inventory and risk appetite before broad procurement.
2. Identity, permissions and supplier terms before non-public data is grounded or shared.
3. Baseline and workflow contract before build.
4. Deterministic process and action boundary before write-enabled model steps.
5. Offline evaluation before shadow; shadow before reliance; sustained T2 evidence before T3.
6. Read scope before write scope; narrow action before broader action.
7. Operations, incident and fallback readiness before production.
8. Financial conversion only after sustained evidence and an approved change to a plan or cost.
9. Tool consolidation only after replacement SLOs hold and exit/restore are tested.
10. Decommission credentials, indexes, data and licences when a workflow retires.

## 15. First 90 days

1. Appoint sponsor, AI Lead and finance/risk counterparts; approve scope and risk appetite.
2. Inventory employee, embedded, supplier and shadow AI; open a time-limited no-blame disclosure channel.
3. Complete readiness, SSO/MFA/lifecycle and high-risk sharing reviews; begin remediation.
4. Publish acceptable-use, data/tool tier and incident rules; create the AI system register.
5. Select the primary employee-assistant pattern for one or two eligible cohorts under reviewed terms.
6. Map and baseline 10–20 task families; select 2–4 workflows after simplify/deterministic screening.
7. Define the standard workflow contract, evaluation pattern and typed action-control architecture.
8. Deliver foundation literacy and role-based practice; establish support and feedback.
9. Ship one or two low-risk T0/T1 workflows through offline evaluation and shadow mode.
10. Hold the first evidence gate: scale, narrow, redesign or kill each workflow.

## 16. Monthly portfolio scorecard

Report by function and workflow where applicable:

- eligible users, active use and licence utilisation;
- workflows/action classes by lifecycle stage and execution mode;
- accepted outcome volume and quality/SLO attainment;
- evaluation pass rates, production overrides, rework and sampling failures;
- incidents, policy violations, pause events and overdue treatments;
- approval/exception queue age, decision effort and escalation;
- model/platform/human-review cost and cost per accepted outcome;
- capacity created, service/quality changes and finance-validated cash benefits;
- training/role coverage and user/affected-person feedback; and
- permission, source, identity, supplier and register control status.

---

## Appendix A — Reference workflow catalogue

The catalogue suggests candidates; it does not pre-approve them. Local process, data, law, risk and economics determine selection and mode ceiling.

| ID | Workflow | Initial mode | Possible ceiling | Critical controls and evidence | Primary unit |
|---|---|---|---|---|---|
| WF-01 | Support reply drafting | T1 | T2 send for pre-approved known categories | Grounded KB, customer context minimisation, tone/correctness, escalation, authenticated approval | Accepted cost/ticket |
| WF-02 | Ticket triage/tag/routing | T0/T2 | T3 for reversible internal routing | Routing accuracy/recall by category, priority false-negative limit, easy undo and queue monitoring | Time to correct queue |
| WF-03 | KB draft from resolved tickets | T1 | T1 | Source citations, redaction, editor acceptance and stale-content controls | Accepted article cost |
| WF-04 | AP invoice extraction/coding | T1 | T2 posting; T3 only for deterministic matched class if locally approved | Duplicate, arithmetic/tax, vendor/PO/master-data checks, bank-detail controls, amount authority and verification | Cost/accepted invoice |
| WF-05 | AR reminder drafting | T1 | T2 send | Ledger truth, dispute/ hardship/strategic-account exclusions, tone, contact consent and approval | Cost/collection cycle |
| WF-06 | Management reporting pack | T1 | T1 | Figures reconciled to governed metrics; commentary linked to data; finance sign-off | Close/preparation hours |
| WF-07 | Meeting-to-CRM notes/tasks | T1/T2 | T3 for low-impact internal fields | Consent/transcript rules, field-level validation, no forecast/commitment overwrite, undo | Admin cost/meeting |
| WF-08 | Account research brief | T0/T1 | T1 | Source quality/date, factuality, conflict and privacy rules | Prep cost/meeting |
| WF-09 | Proposal/tender first draft | T1 | T1 | Approved facts/pricing, source/template control, legal/commercial review | Cost/accepted draft |
| WF-10 | Marketing content variants | T1 | T2 publish for low-risk pre-approved channels only | Brand, claim, rights, disclosure and channel approval | Cost/accepted asset |
| WF-11 | Candidate application summary | T1 | **T1 hard ceiling** | No ranking/recommendation; job-related rubric; privacy, bias/adverse-impact tests; human decision and contestability | Recruiter time/application |
| WF-12 | Employee onboarding pack/tasks | T1/T2 | T2 | HRIS truth, least data, role/template completeness, approval and failed-task queue | Cost/hire |
| WF-13 | Contract first-pass issue spotting | T1 | **T1 hard ceiling** | Clause library, recall-weighted evaluation, jurisdiction limits and qualified legal review | Cost/reviewed contract |
| WF-14 | Coding assistance and PR summary | T1 | T2 for opening a PR; human merge | Repository scope, secrets/code controls, tests/security scan, maintainability and human review | Cost/accepted change |
| WF-15 | Organisation knowledge Q&A | T0 | T0 | Curated sources, security trimming, citations, freshness/deletion and “not found” behaviour | Cost/accepted answer |

**Default-excluded examples:** final hiring/termination/performance decisions; ranking people without a lawful and approved design; payment release or bank-detail change; signing contracts; unreviewed legal/medical/financial advice; unrestricted shell/database/admin access; biometric emotion or personality inference; and dynamic creation of unapproved tools, privileges or sub-agents.

## Appendix B — Readiness assessment

Score 1–5 before program commitment and at each phase gate.

| # | Dimension | A score of 5 | A score of 1 | Score |
|---|---|---|---|---|
| R1 | Strategy and leadership | Named sponsor, funded outcomes, clear risk appetite and willingness to redesign work | Tool purchase delegated downward; no outcome or change commitment | /5 |
| R2 | Identity and cyber baseline | Material apps under strong identity/lifecycle controls; privileged access and incidents managed | Password/shared-account sprawl; leavers or admin access uncontrolled | /5 |
| R3 | Data and permission readiness | Authority/owners known; sharing controlled; candidate data lawful, accessible and sufficiently clean | Unknown authority; broad sharing; no owner; poor input quality | /5 |
| R4 | Process and integration readiness | Stable outcome/process, supported APIs/events and clear exception path | Highly variable undocumented process; browser-only critical system | /5 |
| R5 | Measurement and finance | Baselines, outcome units, quality and TCO can be measured; finance engaged | Anecdotal value only; no volume/cost/quality evidence | /5 |
| R6 | Governance, privacy and legal | Inventory, risk/impact route, obligations, supplier controls and incident process exist | No accountable owner or awareness of obligations/affected people | /5 |
| R7 | Skills and workforce trust | Role-based capability, curious owners, clear measurement/redeployment principles | Fear/prohibition or uncontrolled enthusiasm; no time to learn | /5 |
| R8 | Delivery and operations | Technical owner, test/deploy/support/fallback capability and partner exit rights | Demo capability only; no production owner or support route | /5 |

**Interpretation (/40)**

| Total | Reading | Response |
|---:|---|---|
| 34–40 | Strong starting position | Phase 0 may be compressed, but gates remain |
| 26–33 | Standard | Run Phase 0 as designed and treat low dimensions explicitly |
| 18–25 | Remediation-weighted | Extend Phase 0; limit real-data pilots; reduce portfolio breadth |
| <18 | Not program-ready | Run focused identity/data/governance/process remediation and re-score |

Any dimension ≤2 becomes a named prerequisite with an owner. A high total does not compensate for a critical identity, privacy or operational weakness in the selected workflow.

## Appendix C — Evaluation and evidence pattern

### C.1 Define the workflow contract

Document:

- exact outcome and acceptance unit;
- in-scope and out-of-scope cases;
- authoritative sources and expected outputs;
- action schema and deterministic validations;
- tolerated errors by type and their cost/impact;
- escalation/refusal behaviour;
- latency, cost and availability targets; and
- mode ceiling and promotion/demotion rules.

### C.2 Build representative tests

1. Start with roughly 30–50 real cases for a low-risk seed set; use more where the task has many categories or edge conditions.
2. Stratify routine, uncommon, difficult and known-failure cases rather than sampling only “clean” examples.
3. Add adversarial cases for internet, email, document, tool or memory exposure: indirect prompt injection, data exfiltration requests, malformed input and poisoned sources.
4. Define exact expected values where possible; otherwise use a 3–5 criterion rubric calibrated with domain experts.
5. Keep an untouched holdout and version cases, labels, rubric and reviewer decisions.
6. Check differential performance where people or protected cohorts may be affected.

### C.3 Evaluate the whole system

Test not just the model response but:

- retrieval permissions and source quality;
- tool selection and arguments;
- deterministic policy outcomes;
- transaction success/idempotency/compensation;
- human approval information and authority;
- incident/fallback behaviour; and
- total cost and review/rework.

### C.4 Shadow and production evidence

- Run against live inputs without relying on or executing output.
- Compare to actual human/process outcomes and capture false positives, false negatives and abstentions.
- Size evidence to risk and tolerated error. For common T3 candidates, expect hundreds of representative shadow/T2 observations rather than a small static set; use confidence intervals or conservative upper bounds for material error rates.
- Deploy T2/T3 changes by canary where possible and intensify review after changes.
- Feed real failures, overrides, incidents and new input classes into the test set.

### C.5 Promotion and demotion

Promotion requires all applicable Appendix D controls, accepted offline and production evidence, an owner and council approval at T3. Demote or pause when:

- a critical policy/security violation occurs;
- an error-budget or quality threshold is breached;
- input/data/model/tool change invalidates evidence;
- approval/sampling reveals systematic failure;
- cost per accepted outcome exceeds the stop threshold; or
- the manual fallback is unavailable.

## Appendix D — Production-readiness checklist

### D.1 Required before T0/T1 production

**Business and people**

- [ ] Workflow owner, users, outcome, scope, exclusions and manual fallback are documented.
- [ ] Process has been simplified and deterministic alternatives considered.
- [ ] Affected employees/users are trained; disclosure/contestability requirements are met.

**Data and legal**

- [ ] Sources, authority, classifications, lawful purpose, retention, rights and residency are known.
- [ ] Retrieval permissions, provenance, freshness and deletion have been tested.
- [ ] Privacy/impact assessment is complete where triggered.

**System and security**

- [ ] System is registered; supplier/model/tool versions and responsibilities are documented.
- [ ] Identity, least privilege, secrets, egress and environment separation are implemented.
- [ ] Threat model covers injection, disclosure, poisoning, tool misuse, supply chain and failure.

**Evaluation and operations**

- [ ] Workflow-level evaluation and holdout meet acceptance criteria.
- [ ] Logs/traces/records and retention are configured proportionately.
- [ ] SLOs, cost allocation, support, incident, pause and decommissioning routes exist.

### D.2 Additional before T2

- [ ] Typed action schema, deterministic validations and policy decision are versioned and tested.
- [ ] Approver is authenticated, authorised and shown the exact proposal, evidence, materiality and exceptions.
- [ ] Executor uses narrow write authority, idempotency, verification and compensation/exception handling.
- [ ] Shadow evidence covers representative live cases and known attacks/failures.
- [ ] Approval and exception workload is viable with cover and service levels.
- [ ] Every action can be reconstructed without retaining unnecessary raw sensitive content.

### D.3 Additional before T3

- [ ] Action class is narrow, low-impact, normally non-sensitive and reversible/compensatable.
- [ ] Open-ended planning, dynamic tools/sub-agents and privilege expansion are impossible by design.
- [ ] Sustained T2/shadow evidence is sufficient for the tolerated error and impact, not merely a calendar period.
- [ ] Hard value/rate/time/customer/data limits and anomaly auto-pause are enforced outside the model.
- [ ] Continuous sampling, error budget, owner review and escalation are operating.
- [ ] Pause, rollback/compensation and manual continuity have been exercised.
- [ ] Council approval records residual risk and expiry/review date.


---

# 03 — Data, Security & AI Governance Plan

| | |
|---|---|
| **Version** | v1.0 (Final) |
| **Date** | 2026-08-10 |
| **Posture** | Proportionate governance for a non-heavily-regulated SME |
| **Companion docs** | 00 Overview · 01 Target Architecture · 02 Implementation Plan · 04 Benefits Realisation |

> This plan provides a governance operating model, not legal advice. The adopting organisation remains responsible for identifying the laws, contracts, industrial instruments, professional duties and sector requirements that apply to each use case.

## 1. Purpose and principles

Governance exists to make useful adoption repeatable and safe. It should reduce ambiguity, concentrate decision-making and keep control effort proportional to consequence.

1. **Accountability is named.** Every sanctioned AI system, source and production workflow has a business owner and an operational/control owner.
2. **Proportionality governs effort.** Controls scale with data sensitivity, affected-person impact, autonomy, transaction value, reversibility, scale, criticality and legal exposure.
3. **Identity and authority are explicit.** The organisation can identify the human initiator, automated workload and authority used for a consequential act.
4. **The model does not authorise itself.** Deterministic policy and transaction controls sit between model reasoning and side effects.
5. **Data use is purpose-bound and minimised.** Access to a large corpus or broad API is not justified merely because an agent could use it.
6. **Human control must be meaningful.** Oversight requires competence, context, time, authority, independence and a practical intervention route.
7. **Testing continues in production.** Offline evaluation is necessary but insufficient; overrides, complaints, failures, drift and impacts are monitored.
8. **Telemetry is metadata-first.** Sensitive content is not copied into a new observability repository by default.
9. **Policies are implemented.** Each policy maps to a setting, register, workflow, review, control test or evidence record.
10. **Incidents improve the system.** Staff can report mistakes and near misses without blame; serious matters are contained and escalated quickly.
11. **Users and affected people receive material information.** AI use, limitations and human-review routes are disclosed where relevant or required.
12. **Governance is simplified when it stops adding control value.** The model avoids a parallel bureaucracy detached from existing risk, IT, privacy, finance and people processes.

## 2. Governance operating model

| Role | Responsibilities |
|---|---|
| Executive Sponsor | Sets risk appetite; approves policy and material exceptions; owns major customer, workforce and risk decisions |
| AI Lead | Runs the governance cadence; maintains portfolio/register integrity; coordinates first-line assessments and evidence |
| Function Product Owner | Owns function outcomes, workflow portfolio, adoption and benefit evidence |
| Workflow Owner | Owns a production workflow’s purpose, controls, quality, cost, incidents, review and retirement |
| Security/IT Owner | Identity, secrets, SaaS posture, runtime security, logging, vendor security and incident response |
| Privacy/Data Owner | Purpose, lawful handling, classification, source authority, access, retention/deletion and privacy impact |
| People/Employment Owner | Workforce consultation, role impacts, recruitment/people use cases, discrimination and employee communications |
| Finance Partner | TCO, benefit evidence, transaction authority and segregation-of-duties alignment |
| AI Council | Approves moderate/high-impact systems, T2/T3 promotion, exceptions, major vendor/model changes and remediation priorities |
| All staff/contractors | Use approved tools, protect data, verify work as trained and report incidents/near misses |

### 2.1 Decision rights

| Decision | Default approver |
|---|---|
| Low-impact T0/T1 use case using an already approved tool and data boundary | AI Lead + workflow/data owner |
| Moderate-impact use case, new C3 data source or customer-visible workflow | AI Council |
| High-impact use case, C4 processing, people-affecting system or material legal/financial action | AI Council + Executive Sponsor and specialist advice as required |
| T2 production approval | AI Council or delegated production-readiness authority within documented risk appetite |
| Any T3 action class | AI Council; sponsor where impact is high or policy requires |
| New model/provider/tool tier | Security/Privacy owners + AI Lead; council for material dependency or higher-impact use |
| Policy exception | AI Council, named owner, compensating controls and expiry date |
| Emergency pause/kill | Workflow owner, Security/IT, AI Lead or incident commander according to playbook—no council meeting required |

Low-risk decisions may be delegated, but accountability and evidence may not be delegated away.

## 3. Minimum policy and register set

### 3.1 Policies

| Policy/standard | Purpose | Owner | Default review |
|---|---|---|---|
| Acceptable Use of AI | Approved tools, prohibited data/actions, verification, disclosure and incident reporting | AI Lead | 6-monthly or material change |
| Data Classification & AI Handling | C1–C4 labels, source, prompt, retrieval, output, storage and sharing rules | Security/Privacy | Annual or law/change |
| AI Use-Case & Impact Assessment | Intake, risk rating, impact triggers, approval and evidence | AI Lead / Privacy | Annual |
| Agent & Workflow Operating Standard | Tiers, identities, tools, policy enforcement, limits, operations, promotion/demotion and retirement | AI Lead / IT | 6-monthly |
| Model, Tool & Vendor Standard | Due diligence, approved list, data terms, controls, cost, change and exit | Security/Commercial | Annual and renewal |
| Evaluation & Release Standard | Test design, evidence, change classification, production monitoring and rollback | AI Lead / Engineering | 6-monthly |
| Logging, Monitoring & Records | Metadata schema, sensitive-content policy, retention, access and business-record treatment | Security/Privacy | Annual |
| AI Incident Response | Detection, severity, containment, evidence, notification, communication, restoration and learning | Security/IT | Annual + after incident |
| Human Oversight & Transparency | Approver competence/load, disclosure, contestability, accessibility and escalation | AI Lead / People | Annual |
| Exceptions & Retirement | Exception expiry, compensating controls, decommissioning and evidence disposal | AI Lead | Annual |

### 3.2 Registers

The organisation maintains, at minimum:

1. **AI-system/tool register** — sanctioned employee AI, embedded SaaS AI, models, APIs and material features.
2. **Workflow/agent register** — every production workflow, action classes, tier, owner, identity, data, tools, model/version, controls, evals, limits, runbook and review.
3. **Source register** — grounded corpora/data sources, owner, authority, classification, allowed purposes, ACL model, freshness, retention and deletion.
4. **Vendor/subprocessor register** — contract, security, data-use, geography, model-change, support, cost and exit facts.
5. **Risk/impact assessment register** — current assessments, residual risks, decisions and reassessment triggers.
6. **Exception register** — exception, owner, rationale, compensating controls, expiry and closure.
7. **Incident/near-miss register** — class, severity, affected data/people, containment, notification decision, remediation and lessons.
8. **Benefits/TCO register** — maintained under Document 04 and linked to the workflow record.

A spreadsheet is acceptable for S1/S2 if access, versioning, ownership and review are reliable. Tool sophistication is not a governance outcome.

## 4. Data governance

### 4.1 Classification and AI handling

| Level | Examples | Default AI rule |
|---|---|---|
| **C1 — Public** | Published marketing, public product information, released policies | Any sanctioned tool; unsanctioned tools only where policy permits and no non-public context is added |
| **C2 — Internal** | General working documents, non-sensitive SOPs, internal communications | Tier A or B tools; approved grounding sources |
| **C3 — Confidential** | Customer data, non-public financials, contracts, ordinary employee data, support tickets, source code | Tier A only; purpose-bound access; registered source/workflow; ACL and retention controls |
| **C4 — Restricted** | Payroll detail, health/sensitive data, disciplinary/M&A material, highly sensitive legal data, credentials/secrets | Purpose-specific Tier A deployment only with high-impact approval; general-purpose chat prohibited; stronger minimisation/isolation/monitoring |

**Absolute rule:** credentials, API keys, access tokens, private keys, recovery codes and equivalent secrets MUST NOT be placed in model prompts, retrieval indexes, memories, eval datasets or content-bearing telemetry. They belong in a secrets system and are referenced indirectly.

Other C4 data may be processed only where the workflow is named and approved, the minimum fields are used, provider/deployment terms and location are acceptable, access is isolated, content logging is disabled or specifically controlled, and an impact/privacy assessment supports the use.

Labels SHOULD be applied through Drive labels, Purview sensitivity labels or equivalent controls. Location defaults and ownership are more reliable than expecting every user to label every item perfectly.

### 4.2 AI tool tiers

| Tier | Definition | Mandatory requirements | Maximum default data |
|---|---|---|---|
| **A — Approved enterprise** | Contracted/approved service or deployment with sufficient administrative, data and security control | Due-diligence checklist passed; approved terms/configuration; identity/admin; incident and exit route; register entry | C3; C4 only by named-workflow exception |
| **B — Approved limited** | Useful tool with known limitations that are acceptable for low-sensitivity use | Owner, limited review, explicit permitted use, SSO/admin where available, no hidden expansion to C3 | C2 |
| **C — Unapproved/personal** | Personal accounts, public consumer tools or unreviewed extensions/plugins | No company non-public data; blocked/restricted where practical | C1 only |

A product does not become Tier A solely because it states that customer data is not used for training. Retention, human review, subprocessors, location, access, logs, model changes, IP, service levels, security and exit also matter.

### 4.3 Permission hygiene

1. **Discover** — at least quarterly for material estates: external/public links, broad groups, orphaned drives/sites, stale access, risky OAuth grants and shared accounts.
2. **Prioritise** — risk-rank by classification, audience, volume, owner and whether the source is AI-grounded.
3. **Remediate** — C3/C4 and AI-grounded findings within a short defined SLA; lower-risk findings within a proportionate SLA.
4. **Prevent** — safe sharing defaults, allowlisted external collaboration where appropriate, owner at creation and controlled high-sensitivity locations.
5. **Verify** — before grounding, run adversarial access tests using representative identities; re-test after material permission or retrieval changes.
6. **Propagate** — define and test how source access revocation, deletion and retention expiry reach indexes, caches and derived stores.

A knowledge system that copies content under a broad service identity and filters only in the user interface is not permission-correct.

### 4.4 Data minimisation and purpose limitation

For each model/retrieval/tool step:

- identify the minimum fields and time range required;
- use purpose-built read models or field-level projections where practical;
- exclude unrelated free text, attachments and historical records;
- redact or tokenise direct identifiers where the task can still work;
- prevent output from being reused for a new purpose without assessment; and
- prohibit persistent “memory” for C3/C4 content unless explicitly required, owned, accessible, correctable and deletable.

### 4.5 Records, retention and telemetry

Different artefacts have different purposes and schedules.

| Artefact | Default treatment |
|---|---|
| Sent communication, filed document, approved decision or executed transaction | Business record; same schedule as human-produced equivalent |
| Proposed action and T2 approval | Retain with or link to the resulting business record for the applicable audit period |
| Consequential action event metadata | Starting default: at least 24 months, unless law/business schedule requires longer or risk supports shorter |
| Operational trace metadata | Starting default: ~90 days online, with aggregated trends retained longer as needed |
| Raw prompts, retrieved passages, tool arguments/results and model outputs in telemetry | **Off by default**; if approved for debugging/evaluation, minimise/redact, tightly restrict and retain for the shortest practical period (often 7–30 days) |
| Evaluation datasets | Versioned and retained while needed for assurance; personal/confidential data minimised and access-controlled |
| Hidden model reasoning / chain-of-thought | Do not request or retain as an audit requirement; retain final rationale, evidence and policy decisions instead |

These are model defaults, not universal legal periods. The organisation documents the purpose, authority, access and deletion method for each log class.

### 4.6 Residency, transfer and deletion

- Record provider processing/storage locations and material subprocessors.
- Match residency and transfer controls to applicable law, customer contracts and data classification.
- Test tenant/export deletion and contract termination deletion for material vendors.
- Define cache, embedding, index, backup and evaluation-copy deletion behaviour—not only source deletion.
- Maintain a data-flow record for C3/C4 production workflows.

## 5. AI-specific governance

### 5.1 Use-case intake and risk rating

Every material workflow completes the intake template in §8.1. Rate each dimension Low, Moderate or High; the overall rating is normally the highest dimension unless an approved method justifies otherwise.

| Dimension | Low | Moderate | High |
|---|---|---|---|
| Data | C1–C2 | C3 | C4 or sensitive data at scale |
| Autonomy | T0–T1 | T2 | T3 or open-ended delegation |
| Effect on people | No material individual effect | Advice/interaction may influence a person | Decision or material influence on rights, access, employment, price, service or reputation |
| Reversibility | Immediate/complete | Reversible with effort or delay | Irreversible, difficult to detect or difficult to remediate |
| External/transaction exposure | Internal, informational | Customer-visible or low-value transaction | Financial/legal commitment, public statement or high-value action |
| Scale/criticality | Small, non-critical | Repeated or operationally important | Large-scale, systemic, safety/continuity critical |
| Novelty/dependency | Mature approved pattern | New model/source/tool or material dependency | Unproven architecture, opaque third party or cascading dependencies |
| Legal/contract status | Clearly permitted | Conditions/notice required | High-risk/prohibited/uncertain use or specialist advice required |

**Approval:** Low → AI Lead/owners; Moderate → council; High → council + sponsor and specialist input as appropriate. T3 is not automatically prohibited for every high-rated system, but the burden of evidence and control increases materially; many high-impact decisions will remain human.

### 5.2 AI impact assessment

Complete an AI impact assessment, proportionate to ISO/IEC 42005-style questions, where a system:

- makes or materially informs decisions about individuals;
- may create differential outcomes across groups;
- operates at material scale or in a power-imbalanced context;
- is customer/public-facing and may mislead or materially influence behaviour;
- combines data in a novel way or infers sensitive attributes;
- creates material safety, financial, legal, environmental or societal effects; or
- is designated Moderate/High by the intake.

The assessment covers:

`purpose · affected people/groups · intended benefit · foreseeable use/misuse · data and proxy risks · error/harm pathways · accessibility · transparency · human control · contestability · monitoring · residual risk · decision`

A DPIA/privacy assessment may be required as well. One does not automatically replace the other.

### 5.3 Autonomy controls

| Control | T0 Retrieve | T1 Draft | T2 Act with approval | T3 Autonomous within policy |
|---|---|---|---|---|
| Principal | Delegated user or read-only workload | Delegated user or workload | Dedicated workload plus initiator attribution | Dedicated workload; no shared human session |
| Authority | Read/search/classify only | No consequential side effect | Exact approved typed action | Only approved action class within policy |
| Human role | Verify where task requires | Human independently sends/commits/decides | Competent person approves before execution | Human supervises portfolio, exceptions and sampled outcomes |
| Policy enforcement | Source ACL and read scope | Data/tool/output constraints | Runtime identity, value, SoD, precondition and action checks | Same as T2 plus hard caps and auto-pause |
| Evaluation | Retrieval/task fitness | Representative offline + pilot | End-to-end, security and controlled-live evidence | Sustained T2 evidence through complete business cycle and representative volume |
| Monitoring | Usage/quality spot review | Quality and user correction | 100% action event monitoring plus approval/exception trends | 100% automated monitoring plus risk-based human sample with minimum counts |
| Limits | Query/token/data limits | + output/channel limits | + transaction/action/rate/spend limits | + strict daily/period caps, anomaly and dependency circuit breakers |
| Change | Owner review | Regression on material change | Logged approval and regression before release | Council-approved material changes; canary/rollback |
| Review | At least annual | 6-monthly or change | Quarterly owner/control review | Frequent owner review; monthly council visibility initially |

#### 5.3.1 Promotion, demotion and pause

Promotion is recorded against the specific workflow version and action class. It requires evidence, residual-risk acceptance and a rollback owner.

A workflow is automatically paused or demoted on any defined trigger, including:

- evaluation or production error budget breach;
- Sev-1/2 incident or credible data-exposure concern;
- policy/permission change invalidating the approved scope;
- unexplained anomaly in action, cost, target, destination or tool use;
- monitoring or audit failure;
- material provider/model change without completed regression; or
- sampling, complaint or override evidence showing the control assumptions no longer hold.

Restoration requires containment, cause analysis, control/eval update and authorised release.

### 5.4 Workflow/agent lifecycle and change control

```text
Propose → assess value/risk/impact → design authority and controls
→ build/configure → evaluate → controlled pilot → production-readiness review
→ register/approve tier → operate/monitor → change/promote/demote → retire
```

The authoritative register record includes:

- system/workflow ID, name, business purpose and prohibited uses;
- owners and affected user/customer populations;
- environment and service dependencies;
- human initiator and workload identity model;
- data sources, classifications, retention and legal/impact assessment;
- tools/actions, scopes and policy limits;
- model/provider/version/configuration and prompt/schema versions;
- tier by action class and promotion history;
- evaluation suite, latest results, production error budget and monitoring;
- SLO, budgets, exception queue, fallback, kill/restore and runbook;
- vendor/contracts and exit dependencies;
- benefit/TCO link; and
- last/next review and retirement status.

**Material changes** include prompts/system instructions, models/configuration, retrieval method or sources, schemas, tool allowlists, scopes, policy rules, approval design, data class, affected population, transaction limits and deployment region. They trigger proportionate regression and approval.

### 5.5 Model, tool and vendor onboarding

The due-diligence checklist covers:

#### Data and privacy

- [ ] DPA/privacy terms and purpose are acceptable
- [ ] Customer data and metadata use, model training/improvement and human review are understood and configured
- [ ] Retention, deletion, backup and export behaviour is documented
- [ ] Processing locations, subprocessors and cross-border transfers are acceptable
- [ ] C3/C4 use boundaries and content-logging controls are clear

#### Security and administration

- [ ] Appropriate security assurance (e.g. SOC 2 Type II, ISO/IEC 27001 or compensating assessment)
- [ ] SSO, MFA, SCIM/lifecycle, role separation and audit export appropriate to use
- [ ] API/workload authentication and least-privilege scope supported
- [ ] Vulnerability, incident contact and breach-notification commitments
- [ ] Plugin/connector/tool governance and tenant controls

#### AI/product behaviour

- [ ] Models/features and material limitations identified
- [ ] Provider change-notification and deprecation practices understood
- [ ] Grounding, permissions, citation and admin behaviour tested—not assumed from marketing
- [ ] Evaluation, version pinning or rollback options adequate for criticality
- [ ] Safety/abuse and customer-support escalation route known

#### Commercial and legal

- [ ] Seat, usage, agent/tool and overage pricing modelled under expected and peak load
- [ ] Minimum term/commitment, renewal, suspension and price-change terms understood
- [ ] IP, confidentiality, output and indemnity terms reviewed where material
- [ ] Availability/support/SLA adequate for dependency
- [ ] Accessibility and regional availability considered
- [ ] Export, migration, termination and deletion path tested or credibly demonstrated

The vendor is assigned a tool tier and permitted data/action boundary. Approval is not permanent; renewal and material feature changes trigger review.

### 5.6 Human oversight standards

Meaningful oversight has six elements:

1. **Competence** — the person understands the domain, task, likely AI failure modes and their accountability.
2. **Context** — they see source evidence, proposed action, material differences, policy results, uncertainty/flags and consequences.
3. **Authority** — they are authorised under existing delegation and segregation-of-duties rules.
4. **Time and workload** — volume and interface permit real review; median review time, rejection and escalation are monitored.
5. **Independence** — high-consequence review is not distorted by targets that reward automatic approval.
6. **Intervention** — reject, edit, escalate, pause and correct are practical; affected people have a human route where appropriate.

A T2 approval card SHOULD show in one view:

`trigger/input · source evidence · proposed action/diff · validations/policy checks · material uncertainty/flags · consequence · approve/reject/edit/escalate`

For financial, legal, people and high-value customer actions, existing approval matrices and segregation of duties remain in force. AI cannot approve its own proposed action or satisfy a two-person control twice.

#### Fatigue controls

- set workflow-specific approval capacity based on volume and review time;
- batch only where each item remains reviewable and risk permits;
- rotate and provide backup approvers;
- monitor implausibly fast review, near-100% approvals, low evidence opening and overdue queues;
- move back to T1 or redesign if T2 creates an unsustainable approval factory; and
- never use poor approval quality as an argument to remove approval and call the workflow T3.

### 5.7 Transparency, disclosure and contestability

Provide information proportionate to context and law:

- employees know which tools are sanctioned, what is monitored and how AI affects their work;
- customers/users know when they are interacting with AI where that fact is material or required;
- AI-generated public or synthetic content is labelled where required by law, platform or policy;
- people affected by a material AI-assisted decision can obtain appropriate information and human review/complaint handling;
- disclosures explain purpose, role of AI and limits without implying certainty or hiding accountability; and
- accessibility and non-AI alternatives are maintained where necessary.

## 6. Security controls

### 6.1 Human and non-human identity

**Humans:** central IdP, universal MFA, phishing-resistant methods for privileged/high-impact roles, rapid joiner/mover/leaver controls, periodic access review and monitored risky OAuth grants.

**Interactive AI:** preserve user identity and source ACLs wherever the use is personal/delegated. Do not replace user permissions with a broad backend identity merely for convenience.

**Automated workflows:** dedicated workload identity per independently operated production trust boundary; minimal scopes; short-lived credentials or workload federation where supported; separate identities by environment; no shared human accounts or browser sessions.

**Attribution:** record both the human initiator and executing workload when a human triggers an automated action.

### 6.2 Policy enforcement and side-effect boundary

A model may produce a **typed proposed action**. Before execution, a deterministic control validates:

- authenticated principal and initiator;
- approved workflow/environment/version;
- action allowlist and tool version;
- data classification and permitted purpose;
- role/delegation and segregation of duties;
- transaction value/rate/time/recipient limits;
- current business-record preconditions;
- schema, sanitisation and destination; and
- whether human approval is required and still valid.

The policy decision and rule version are logged. Prompt instructions, model confidence or an MCP tool description are not an authorisation decision.

### 6.3 LLM and agentic threat controls

This control map should be maintained against both the current **OWASP Top 10 for LLM Applications** and **OWASP Top 10 for Agentic Applications**.

| Threat family | Primary controls |
|---|---|
| Prompt injection / agent-goal hijack | Treat retrieved, user and tool content as untrusted; separate instruction/data channels; context trust labels; least-privilege tools; deterministic policy; approval for consequential actions; adversarial tests. Pattern filters are supplementary only. |
| Tool misuse / excessive agency | Typed action schemas; explicit action allowlist; minimal scopes; transaction/rate/spend caps; sandbox; preconditions; no free-form shell/SQL/payment/HR execution |
| Identity and privilege abuse | Dedicated workload identity; delegated-user preservation; secure OAuth; no token passthrough; audience/scope validation; short-lived credentials; privileged access review |
| Insecure output handling | Parse and validate outputs; parameterised APIs/queries; HTML/content sanitisation; no direct model output into command, code, SQL or browser execution |
| Sensitive information disclosure | Purpose/field minimisation; ACL-aware retrieval; redaction; DLP/egress; content logging off; no secrets; response filtering where justified |
| Data, memory or RAG poisoning | Curated/owned sources; restricted write access; provenance; freshness/authority ranking; source-integrity monitoring; memory review/deletion; conflicting-source handling |
| Agentic supply chain / rogue tools | Approved tool/MCP registry; owner; version pinning; code review; SBOM/dependency scanning; signing/provenance where available; sandbox and egress limits |
| Insecure inter-agent communication | Avoid unnecessary multi-agent design; authenticate peers; typed messages; provenance; bounded delegation; no inherited broad authority; loop/cascade limits |
| Cascading failure / unexpected behaviour | Durable orchestration; state/step limits; recursion/loop caps; circuit breakers; exception queues; compensation; dependency isolation; auto-pause |
| Model denial-of-service / runaway cost | Input/output size limits; quotas; timeouts; concurrency controls; per-workflow budgets; anomaly alerts and hard stop |
| Overreliance and human-factor failure | Task-specific training; uncertainty/evidence; outcome evals; meaningful oversight; complaint/appeal; quality guardrails and periodic manual benchmark |
| Model/version drift | Approved version/config; regression tests; canary; provider-change monitoring; rollback/demotion and reapproval |

### 6.4 MCP and tool-adapter security

MCP is not a security boundary. Production MCP servers/clients MUST use the current specification and security guidance appropriate to their transport, including robust OAuth-based authorisation for HTTP deployments where supported.

Required controls:

- approved server registry, owner and business purpose;
- explicit client/server trust and endpoint allowlist;
- audience-bound, scoped and short-lived tokens; no credential or token passthrough;
- defence against confused-deputy, redirect and authorisation-server mix-up risks;
- per-action scopes and deterministic policy enforcement outside tool descriptions;
- schema validation and safe handling of tool outputs;
- version pinning, review, dependency/SBOM and change notification;
- tenant/data boundary validation and no implicit cross-user data sharing;
- egress, rate, action and spend limits;
- audit correlation across client, server, policy and downstream system; and
- rapid server/action disable capability.

The same principles apply to proprietary tool/function interfaces.

### 6.5 Logging and audit

Minimum consequential-action metadata:

```text
trace_id · timestamp · environment · workflow/version · workload identity
human initiator (if any) · model/provider/version/config · prompt/schema version
retrieval source references · tool/action/version · proposed-action reference/hash
policy decision/rule version · approval/approver · result/reconciliation
exception/override · latency · usage/cost · business-record/evidence references
```

Controls:

- raw content capture disabled by default;
- explicit purpose and approval for content-bearing traces;
- field redaction/tokenisation before export where practical;
- separate access roles for operators, security, evaluators and business users;
- encryption, integrity and retention controls on the log destination;
- no secrets, full credentials or hidden reasoning;
- quarterly reconstruction of representative T2/T3 actions; and
- alerting on missing logs for workflows that require complete action evidence.

Optional higher-assurance controls include append-only/WORM storage, signed events or hash chaining, but only where the threat/risk justifies the operational cost.

### 6.6 Runtime, endpoint and egress controls

- managed devices/browsers for users handling C3/C4 where proportionate;
- restrict unapproved browser extensions, plugins and OAuth apps;
- code-first runtimes use separate networks/projects/subscriptions and minimal outbound destinations;
- sandbox model-generated code and untrusted files;
- do not expose internal administrative endpoints or broad network access to an agent;
- dependency and container scanning for custom components;
- secure configuration, patching and backups aligned to existing cyber program;
- apply the Essential Eight or another recognised cyber baseline according to risk—do not assume one maturity level fits every SME; and
- test manual fallback and recovery as a cyber-resilience control.

### 6.7 Secure development and release

For custom components and workflow code:

- version control, peer review and protected production branches;
- secrets scanning, dependency/SBOM and vulnerability scanning;
- unit/integration/end-to-end and adversarial tests;
- reproducible configuration and environment separation;
- signed/provenanced builds where practical;
- release approval proportional to tier;
- canary/feature flag/rollback for material changes; and
- supported runtime, dependency and model deprecation monitoring.

## 7. Incident response

### 7.1 Incident classes

| Class | Examples | Indicative severity |
|---|---|---|
| Data/privacy exposure | Wrong-user retrieval, C3/C4 disclosure, unapproved processing, telemetry leak | Sev 1–2 |
| Unauthorised/incorrect action | Outside-scope write, wrong recipient/amount/status, bypassed approval | Sev 1–2 |
| Injection/tool compromise | Untrusted content hijacks goal or tool; malicious/rogue MCP server | Sev 1–2 |
| People/impact harm | Unfair or misleading decision/influence, inaccessible process, denied review | Sev 1–2 depending on impact |
| Quality regression | Material hallucination, extraction/routing collapse, stale source causing error | Sev 2–3 |
| Operational/cascading failure | Loop, duplicate/partial transactions, dependency outage, backlog | Sev 2–3 |
| Runaway cost/abuse | Budget breach, credential misuse, uncontrolled usage | Sev 2–3 |
| Shadow-AI exposure | Company data in unapproved tool/account | Sev 2–3 based on data and exposure |

The organisation should align severity definitions with its existing incident framework rather than create incompatible parallel terminology.

### 7.2 Response flow

```text
Detect/report → contain (pause/kill/revoke/block) → preserve evidence
→ assess data, people, transaction and dependency blast radius
→ correct/reverse/notify users as operationally required
→ determine legal, privacy, contractual, employment and customer notification
→ remediate controls/data/evals → controlled restore → post-incident learning
→ update registers, assessments, runbooks and benefit/disbenefit records
```

### 7.3 Immediate containment authority

Named incident roles may pause a workflow, revoke its identity, disable a tool/server, block a source or revert to manual operation without prior council approval. Restoration follows authorised change control.

### 7.4 Drills

- workflow-level pause/kill and manual fallback: at least quarterly for consequential T2/T3;
- data-exposure or injection tabletop: at least annually, and after material architecture change;
- restore/reconciliation test: according to workflow criticality; and
- evidence reconstruction: quarterly sample.

## 8. Templates

### 8.1 Use-case intake

```text
Workflow name / owner / function / version
Business problem, current process and baseline
Target users and people/customers affected
Outcome, unit cost and kill criteria
Data sources, fields, classifications, owner and allowed purpose
Model/provider, tools/actions, identity and scopes
Autonomy tier by action class
Reversibility, transaction/external exposure, scale and criticality
Privacy/PII and AI-impact triggers
Applicable jurisdictions, customer terms and professional duties
Human oversight, disclosure, review/appeal and accessibility
Evaluation, production monitoring and error budget
SLO, exception/fallback, kill/restore and incident owner
TCO/benefit link
Risk rating → approval path → decision/review date
```

### 8.2 AI impact assessment

```text
System/workflow and decision role
Affected individuals/groups and context/power relationship
Intended benefits and evidence
Foreseeable normal use, misuse and out-of-scope use
Data, proxies, inference and representativeness
Error types, distribution, severity and detectability
Fairness/non-discrimination and accessibility considerations
Transparency, explanation and contestability
Human authority, competence, workload and intervention
Security, privacy, safety and dependency pathways
Monitoring, complaint and remediation plan
Residual impacts, mitigations, decision and reassessment trigger
```

### 8.3 Workflow/agent register record

```text
ID / name / owner / purpose / prohibited use / environment
User/affected populations
Human initiator and workload identity
Data/source/classification/retention
Tools/actions/scopes/policy limits
Model/provider/version/config; prompt/schema versions
Tier by action class and promotion history
Eval suite/results/error budget/production monitoring
SLO/budgets/limits/exception queue/fallback/kill/restore
Impact/privacy/vendor/contract links
Runbook/support/incident owner
Benefit/TCO link
Last review / next review / status / retirement record
```

### 8.4 Exception record

```text
Exception ID / policy/control / workflow/vendor
Reason standard position cannot be met
Risk and affected data/actions/people
Compensating controls
Owner and approver
Effective date / expiry / review trigger
Remediation or exit plan
Closure evidence
```

## 9. Assurance and framework alignment

### 9.1 Reference posture as at 10 August 2026

- **Australia:** use the National AI Centre’s six Essential AI Practices as the local baseline: accountability, impact understanding, risk management, essential information, testing/monitoring and human control.
- **NIST:** map governance to the current AI RMF 1.0 functions—Govern, Map, Measure, Manage—and the NIST AI 600-1 Generative AI Profile. AI RMF 1.0 is under revision, so maintain an update trigger rather than claiming timeless conformance.
- **ISO:** keep the management system **alignment-ready** with ISO/IEC 42001. Use ISO/IEC 42005-style AI system impact assessment for material people/societal impacts. Seek certification only where customer, market or risk value justifies it.
- **Security:** maintain control mapping to the current OWASP Top 10 for LLM Applications and Agentic Applications, plus the organisation’s existing cyber framework.
- **Australia cyber baseline:** use the ASD Essential Eight or another recognised baseline proportionately. A universal maturity-level prescription for all SMEs is inappropriate.

### 9.2 Key risk indicators

Report trend, target and overdue action for:

- percentage of sanctioned AI systems/workflows correctly registered;
- production workflows with current owner, eval and review;
- unresolved permission and risky OAuth findings by classification;
- DLP/privacy events and content-logging exceptions;
- evaluation/error-budget breaches and drift;
- approval workload, median review time, rejection/override and rubber-stamp indicators;
- T3 sample failures, auto-pauses and action-limit breaches;
- incidents/near misses by class and time to contain/restore;
- unapproved AI use and remediation;
- vendor/model material changes awaiting assessment;
- overdue exceptions and impact/privacy reassessments; and
- staff/approver competency coverage.

### 9.3 Assurance cadence

| Cadence | Assurance activity |
|---|---|
| Monthly | Council KRI review, incidents, exceptions, vendor/model changes and promote/demote decisions |
| Quarterly | Workflow register recertification, access/OAuth review, action reconstruction, permission test, T2/T3 control health and benefits/TCO review |
| Six-monthly | Policy, architecture, training and threat-model refresh |
| Annual | Structured framework assessment, impact/privacy portfolio review, incident exercise, material exit/export and resilience tests |
| Triggered | Material model/tool/source/scope/change, law/contract change, incident, new affected population or transaction class |

## 10. Jurisdiction and obligation watchpoints — August 2026

### 10.1 Australia

- The Privacy Act applies to AI handling personal information where the organisation is an APP entity. Use purpose limitation, minimisation, appropriate notice, security and vendor due diligence.
- From **10 December 2026**, Australian APP entities using personal information in automated decision-making that could significantly affect rights or interests will have additional privacy-policy transparency obligations. The 24-month program must identify affected workflows and prepare policy disclosures before commencement.
- Employment, discrimination, surveillance, consumer, professional and workplace-consultation obligations may apply even where there is no AI-specific law.
- Recruitment and people workflows require particular scrutiny; the base architecture prohibits automated ranking/rejection in the starter catalogue.

### 10.2 European Union

- The EU AI Act became generally applicable on **2 August 2026**, with earlier provisions already in force and specific later dates for some high-risk obligations.
- Following the 2026 AI Omnibus changes, high-risk rules for relevant stand-alone systems are scheduled from **2 December 2027**, and certain product-embedded systems from **2 August 2028**. Applicability must be checked against the final rules, role (provider/deployer), location, affected people and use case.
- AI literacy and relevant transparency duties may already apply. Employment and other Annex III-type systems should be designed now for inventory, impact, data, logging, human oversight and documentation even where the later date applies.

### 10.3 Other jurisdictions and contracts

Maintain a simple applicability matrix covering:

`where organisation operates · where affected people are · provider/deployer role · decision domain · data location · customer/sector terms · required notice/review/records · responsible owner`

Do not infer that a vendor’s “compliant” marketing transfers legal accountability to the adopting organisation.

## 11. Proportional operating burden

The objective is effective control, not a fixed governance headcount. Indicative steady-state coordination effort may be:

- S1: embedded duties, often 0.05–0.2 FTE plus MSP/legal specialist bursts;
- S2: roughly 0.2–0.6 FTE distributed across AI Lead, IT/privacy, finance and function owners; and
- S3: roughly 0.5–1.5 FTE equivalent across the portfolio, excluding engineering and ordinary operational supervision.

These are planning ranges, not benchmarks. High-impact or regulated portfolios require more; a low-risk managed-tool portfolio may require less. If review effort grows without reducing incidents, uncertainty or decision time, simplify the control design.

## 12. Primary reference points

- [Australian Government — Guidance for AI Adoption](https://www.ai.gov.au/staying-safe-and-responsible/essential-ai-practices/guidance-ai-adoption-implementation-guidance)
- [ASD/ACSC — Careful adoption of agentic AI services](https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/careful-adoption-of-agentic-ai-services)
- [OAIC — Privacy and commercially available AI](https://www.oaic.gov.au/privacy/privacy-guidance-for-organisations-and-government-agencies/guidance-on-privacy-and-the-use-of-commercially-available-ai-products)
- [OAIC — automated decision-making transparency consultation](https://www.oaic.gov.au/engage-with-us/consultations/consultation-on-guidance-for-transparency-in-automated-decision-making)
- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework)
- [ISO/IEC 42001 — AI management systems](https://www.iso.org/standard/42001)
- [ISO/IEC 42005 — AI system impact assessment](https://www.iso.org/standard/42005)
- [EU AI Act implementation](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai)
- [MCP specification 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28)
- [OpenTelemetry — handling sensitive data](https://opentelemetry.io/docs/security/handling-sensitive-data/)
- [OWASP Top 10 for LLM Applications 2026](https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/)
- [OWASP Top 10 for Agentic Applications 2026](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/)


---

# 04 — Benefits Realisation Plan

| | |
|---|---|
| **Version** | v1.0 (Final) |
| **Date** | 2026-08-10 |
| **Goal** | Prove workflow value, convert eligible value through explicit routes, report full cost and prevent double counting |
| **Companion docs** | 00 Overview · 01 Target Architecture · 02 Implementation Plan · 03 Data, Security & AI Governance |

## 1. Principles

1. **Baseline before business-case approval.** Discovery may begin earlier; no quantified benefit is approved without a credible current-state measure.
2. **Capacity is not cash.** Reduced task effort is reported as capacity created. It becomes a cash effect only through an evidenced spend, hiring or workforce route.
3. **Capacity and the cash route are not additive.** When capacity enables hiring avoidance, attrition reshaping or in-housing, the underlying hours remain one resource and are never summed twice.
4. **Count net of full TCO.** Licences and tokens are only part of cost. Include implementation, integration, controls, training, support, overlap, failure/rework and retirement.
5. **Separate outcome classes.** Cash, capacity, revenue, quality/risk and disbenefits are reported in separate columns. A board total combines only compatible categories.
6. **Use actual adoption and sustained performance.** A laboratory time saving is reduced by utilisation, exceptions, rework, quality guardrails and operating downtime.
7. **Bank only evidenced cash effects.** Forecast, committed and banked are distinct states; finance co-signs banked values.
8. **No double counting across workflows.** Shared tasks, headcount plans, external spend and software licences have one card/owner and a clear allocation rule.
9. **Protect quality, risk and resilience.** A saving that increases material error, complaint, security, safety or employee/customer harm is not a net benefit.
10. **Use conservative confidence and sensitivity.** AI productivity effects vary materially by task, user, model, process and implementation; local evidence overrides generic claims.
11. **Redeploy transparently.** In a growing SME, the default use of capacity is usually growth, service, quality and backlog—not covert redundancy.
12. **Re-baseline after sustained change.** Yesterday’s benefit becomes tomorrow’s normal operating baseline.

## 2. Benefit and cost taxonomy

| ID | Class | Definition | Cash treatment | Evidence standard |
|---|---|---|---|---|
| **B1** | Capacity created | Net internal hours removed from in-scope work after adoption, exception, rework and quality adjustment | **Non-cash** until allocated through a route | Repeated measurement versus baseline; owner accepts available capacity |
| **B2** | External spend reduction | Agency, contractor, outsourced service or professional-fee spend reduced | Cash benefit | Contract/PO/statement-of-work reduced or invoice run-rate falls, net of replacement cost |
| **B3** | Hiring avoidance or attrition-based reshaping | A pre-existing funded role is deferred/cancelled, or a vacancy is not backfilled after durable task compression | Cash benefit | Pre-existing workforce plan or approved vacancy; budget/headcount decision recorded; capacity sufficiency demonstrated |
| **B4** | Software/licence consolidation | Subscription, seat, platform or point tool retired/right-sized | Cash benefit | Renewal/order/invoice reduced; replacement and migration cost included |
| **B5** | Direct unit-cost reduction | Non-labour transaction cost, error/rework cost, postage/processing or other per-unit cost falls; labour may be included only if not also claimed through B1/B3 | Cash benefit | Component-level unit-cost evidence sustained for an approved period |
| **B6** | Revenue enablement | Greater throughput, conversion, retention, price realisation or faster time-to-revenue | Separate; cash only with credible attribution | Leading indicators by default; bank only where finance accepts causal evidence and avoids overlap |
| **B7** | Quality, risk and experience | Fewer errors, faster service, better compliance, resilience, employee/customer experience or reduced risk exposure | Usually non-cash; cash where avoided/rework cost is measured | Defined quality/risk metric, baseline and sustained change; monetisation method documented |
| **B8** | AI total cost of ownership | All one-time and recurring cost attributable to the capability and its operation | Negative cash line | Actuals/forecast from finance, contracts, cloud and allocated internal/partner effort |
| **D1** | Disbenefits | New rework, review burden, delay, incidents, complaints, lock-in, lost flexibility, quality decline or change cost | Negative cash and/or non-cash | Measured or provisioned; linked to workflow and corrective action |

### 2.1 Board reporting rule

The board-level cash number is:

```text
Banked B2 + B3 + B4 + cash-qualified B5 + cash-qualified B6/B7
− actual B8 − cash-valued D1
```

B1 capacity created is shown separately. It is never added to B3 or labour-based B5. Non-cash B6/B7 and non-cash D1 are reported as operational outcomes, not hidden.

### 2.2 Presenting the value story

The cash number in §2.1 is deliberately conservative and must never be presented alone. Every board or owner report shows three lines together, in this order:

1. **Net cash** — the §2.1 formula, finance-validated;
2. **Capacity created and where it went** — the ledger destinations in §4.4, the operating dividend of the program; and
3. **Quality, risk and service movements, including disbenefits** — the guardrail evidence.

In a growth-mode SME the cash line is commonly negative in Year 1 and modest at Month 24 (§8). The commercial case is normally carried by capacity redeployed to funded priorities, service and quality effects, and the option value of a governed capability. Leading with cash alone understates a sound program; leading with capacity alone overstates it. Present all three lines every quarter, and let §8.4 supply the interpretation.

## 3. Where value may arise

The table is a discovery map, not a benchmark. It deliberately avoids asserting universal percentage gains.

| Function | Candidate workflow families | Primary baseline | Likely value routes | Critical counter-metrics |
|---|---|---|---|---|
| Finance | Invoice capture/match/post, AR reminders, expense review, reporting preparation, close support | Volume, touch/cycle time, exception/rework, outsourced cost, cost per transaction | B1, B2, B3, B5 | Duplicate/error, tax/coding accuracy, close quality, fraud/SoD |
| Customer support | Triage, draft replies, known-issue actions, KB drafting, QA summaries | Handle time, first response, resolution, recontact, backlog, outsourced L1 | B1, B2, B3, B5, B7 | Resolution quality, complaints, escalation, customer effort |
| Sales | Research, call prep, notes-to-CRM, proposal drafts, hygiene | Admin hours, proposal cycle, CRM completeness, pipeline coverage | B1, B3, B6 | Accuracy, customer trust, selling time actually redeployed |
| Marketing | Brief/asset drafts, variants, reporting, routine publishing | Internal/agency hours, cost per asset, cycle, output volume | B1, B2, B4, B6 | Brand/legal quality, rework, channel performance |
| People/HR | Administrative summaries, policy/JD drafts, onboarding packs, employee FAQ | Admin time, recruiter/outsourcing spend, time to onboard | B1, B2, B7 | Fairness, privacy, decision quality, employee experience |
| Operations/admin | Scheduling, document processing, vendor communications, SOP drafting | Volume, queue time, touch time, manual errors | B1, B3, B4, B5 | Missed exceptions, service continuity, supplier/customer impact |
| Engineering/product | Coding support, tests, PR drafts, documentation, issue analysis | Cycle time, review time, defects, throughput, contractor spend | B1, B2, B3, B6, B7 | Escaped defects, review burden, security, maintainability |
| Legal/commercial | First-pass issue spotting, template population, obligation extraction | Internal/counsel time, turnaround, issue recall | B1, B2, B7 | Missed material issues, false reassurance, confidentiality |
| Cross-functional | Knowledge search, meeting actions, internal drafting | Search/admin time, answer success, meeting follow-through | B1, B4, B7 | Leakage, unsupported answers, meeting/action quality |

### 3.1 Software-consolidation discovery

Review:

- duplicate general AI assistants;
- transcription/meeting tools now covered by a sanctioned suite;
- single-purpose drafting, research, OCR or summarisation tools;
- overlapping automation/iPaaS subscriptions;
- unused or low-value premium seats;
- duplicate enterprise-search or chatbot products;
- bespoke pilots that became vendor-supported features; and
- development/test subscriptions left running after experiments.

Do not assume a fixed percentage of SaaS spend is recoverable. Use actual contracts, renewal dates, utilisation, migration cost and feature gaps.

## 4. Measurement system

### 4.1 Baselines

For each selected workflow, record:

- annual/monthly volume and seasonality;
- current touch time and elapsed cycle time;
- role mix and fully loaded cost methodology;
- exception, correction, rework and escalation rates;
- service/quality/customer metrics;
- external spend and software cost attributable to the process;
- planned hires or vacancies relevant to future capacity;
- current controls and manual review cost; and
- confidence, sample period and known limitations of the baseline.

Use a practical method: system event data, ticket/CRM/finance analytics, a representative work sample, structured time study or a combination. Avoid covert employee surveillance; measure workflows and outcomes.

### 4.2 Ongoing measures

| Layer | Measures |
|---|---|
| Adoption/utilisation | Eligible users/transactions, successful completion, depth, abandonment and fallback |
| Workflow effort | Human touch time, approval time, exception/rework and support effort |
| Service/outcome | Cycle time, throughput, backlog, resolution, accuracy and customer/user result |
| Quality/risk | Error severity, complaints, incidents, override, bias/impact, reconciliation and control failure |
| Reliability | Availability, failed/partial runs, retries, exception age and manual fallback |
| Unit economics | Fully defined cost per outcome and marginal cost at volume |
| B8 TCO | Licences, model usage, platform, people, partner, security/privacy, change, support, overlap and retirement |
| Benefits | B1 capacity, B2–B7 by evidence state, D1 and cash/net result |

### 4.3 Capacity-created formula

A practical calculation is:

```text
Gross hours released
= eligible transaction volume × (baseline touch time − new touch time)

Net capacity created
= gross hours released
× successful adoption/utilisation rate
− additional approval, exception, rework, support and recovery hours
```

Convert to FTE-equivalent or loaded-cost equivalent for planning, but keep it labelled **capacity**, not cash.

Where quality changes, adjust the calculation or fail the benefit gate. Faster work with more material errors is not capacity creation.

### 4.4 Capacity ledger

Every measured pool of released time is allocated once.

`Period · function/role · source workflow(s) · net hours · confidence · owner · destination · cash route (if any) · evidence · status`

Allowed destinations:

- **R1 — growth/service redeployment**: more pipeline, product, service, customer or market work;
- **R2 — backlog/quality/resilience**: documentation, controls, technical debt, training, leave coverage or service improvement;
- **R3 — external-spend in-housing**: supports B2, but the internal hours are not also added as a cash saving;
- **R4 — hiring avoidance**: supports B3 against a documented plan;
- **R5 — attrition-based reshaping**: supports B3 when a vacancy is not backfilled;
- **R6 — retained operating slack**: resilience or workload normalisation, reported honestly as such; and
- **R7 — not realised**: fragmented, absorbed, lost to demand growth or not accepted by the manager.

### 4.5 Benefit card

```text
Card ID / benefit class / workflow(s) / owner
Baseline value, date, method and confidence
Mechanism and counter-metrics
Forecast: low/base/high and evidence grade
Capacity dependency and ledger allocation, if any
Cash conversion route and date
One-time and recurring B8 cost allocation
Disbenefits/risks and quality guardrails
Status: Hypothesis → Baseline validated → Measured → Committed → Banked → Reversed/closed
Evidence links / finance sign-off / banked amount and date
Overlap check with other cards
```

### 4.6 Evidence grades

| Grade | Meaning | Use |
|---|---|---|
| **E0 — Hypothesis** | Directional assumption with no local baseline | Discovery only; not in committed plan |
| **E1 — Baseline validated** | Current state measured with known confidence | Forecast input |
| **E2 — Pilot measured** | Controlled local result, not yet sustained at production scale | Risk-adjusted forecast |
| **E3 — Production sustained** | Result sustained over representative volume/business cycle | May move to committed route |
| **E4 — Committed** | Budget/spend/headcount owner has approved conversion action and date | Committed benefit, not yet banked |
| **E5 — Banked** | P&L/budget/contract effect occurred and finance verified it | Board cash benefit |

A forecast MAY be probability-weighted for planning, but a banked actual is never probability-discounted.

## 5. Full AI total cost of ownership (B8)

### 5.1 Cost categories

| Category | Examples |
|---|---|
| Employee licences | Suite copilot, enterprise assistant, specialist AI seats |
| Model/agent usage | Tokens, searches, agent actions, tool calls, embeddings, storage and overages |
| Platform/integration | iPaaS, gateway, observability, retrieval/search, warehouse, queues and managed compute |
| Build/configuration | Internal engineering/analysis, implementation partner and test-data preparation |
| Security/privacy/assurance | Vendor review, impact/DPIA, testing, logging, DLP, penetration/adversarial testing and audit |
| Change/people | Training, communications, workflow redesign, consultation and manager/approver time |
| Run/support | Monitoring, incident response, exception handling, evaluations, prompt/source maintenance and on-call/owner time |
| Transition/overlap | Dual licences, parallel process, migration, data clean-up and temporary productivity dip |
| Failure/disbenefit | Rework, customer remediation, incident cost, wrong transactions and risk allowance |
| Exit/retirement | Export, migration, contract termination, deletion verification and decommissioning |

Allocate shared cost to workflows using a documented driver such as active users, actions, compute, support effort or attributable use. Do not hide platform-team cost outside the AI business case.

### 5.2 Economic gate

Before production, each workflow records:

- one-time investment and annual run cost;
- low/base/high capacity and cash outcomes;
- adoption, quality and volume sensitivities;
- risk/disbenefit allowance;
- incremental and fully allocated unit cost;
- expected payback and break-even volume;
- contract/minimum-commitment and exit exposure; and
- kill thresholds and review date.

A sensible default for discretionary SME automation is a **base-case payback within approximately 24 months**, or a documented strategic, customer, quality or risk reason for accepting longer. The organisation may set a different hurdle, but should do so explicitly.

## 6. Conversion routes

### 6.1 H1 — Capacity redeployment

**Mechanism:** released hours are formally assigned to funded priorities, service improvement, backlog or resilience.

**Treatment:** valuable operating outcome, usually non-cash. It becomes a cash/revenue benefit only through a separately evidenced route.

**Evidence:** manager and receiving priority accept the capacity; work allocation or output changes; no simultaneous B3 claim unless the allocation is adjusted.

### 6.2 H2 — External-spend reduction

**Mechanism:** reduce outsourced work, agency scope, contractor volume or professional fees because an internal AI-supported process replaces it.

**Treatment:** B2 cash benefit.

**Evidence:** contract/PO/invoice reduction, net of internal capacity, tooling, quality-control and transition cost. Capacity used to in-house the work is allocated R3 and not added separately.

### 6.3 H3 — Hiring avoidance

**Mechanism:** a pre-existing funded role is deferred or cancelled because demonstrated capacity absorbs the expected workload.

**Treatment:** B3 cash benefit against the approved workforce plan.

**Evidence:** documented prior role plan, capacity/volume analysis, budget decision and service/quality guardrails. The same released capacity cannot also support another B3 card.

### 6.4 H4 — Attrition-based reshaping

**Mechanism:** after sustained task compression and role redesign, a natural vacancy is not backfilled or is replaced with a different/lower-cost capability.

**Treatment:** B3 cash benefit after vacancy and budget decision.

**Evidence:** durable workload evidence, role redesign, service-risk assessment, org/budget change and people-process compliance.

### 6.5 H5 — Software/licence retirement

**Mechanism:** retire/right-size point tools, overlapping assistants, unused premium seats or duplicate platforms.

**Treatment:** B4 cash benefit.

**Evidence:** renewal/order/invoice reduced; migration, replacement usage and lost functionality included.

### 6.6 H6 — Direct unit-cost reduction

**Mechanism:** reduce non-labour transaction cost, error/rework, manual processing fee or other variable cost.

**Treatment:** B5 cash benefit.

**Evidence:** component-level baseline and sustained unit-cost change. Labour components must reconcile to the capacity ledger and B3 cards.

## 7. Redeployment and people plan

### 7.1 Default capacity allocation policy

The council may set a default planning split, reviewed annually. An illustrative growth-oriented policy is:

| Destination | Illustrative share | Examples |
|---|---:|---|
| Growth/customer capacity | 40–55% | More account coverage, faster roadmap, new offers, service improvement |
| Backlog/quality/resilience | 20–30% | Documentation, controls, technical debt, leave coverage, training |
| AI/process capability | 10–20% | Workflow ownership, evaluation, data quality and automation improvement |
| Cashable workforce route | 0–20% | Hiring avoidance or attrition-based reshaping where durable and appropriate |

These are policy ranges, not benefits. The actual allocation follows strategy, demand and workforce obligations.

### 7.2 People principles

- no-surprise role and capacity conversations;
- measure workflows, not individual keystrokes;
- do not use pilot time studies as covert performance management;
- reskill and redesign before assuming role removal;
- internal mobility before external hiring where skills can transfer;
- consult employees/representatives where required;
- protect realistic operating slack and recovery capacity;
- disclose where AI monitoring or decision support affects staff; and
- record where capacity went so “efficiency” does not become an unspoken workload increase.

## 8. Worked example — 100-FTE organisation

This example demonstrates the accounting logic. It is **not** a benchmark or forecast for a real organisation.

### 8.1 Assumptions

- 100 FTE; 70 roles have material eligible knowledge/process work.
- Average fully loaded cost: **$140,000**.
- Eligible payroll: **$9.8 million**; total people cost: approximately **$14.0 million**.
- Current external service candidates: **$600,000/year**.
- Current SaaS spend: **$700,000/year**.
- A pre-existing two-year growth plan includes roles that may become B3 candidates.
- All values are AUD, rounded, before tax and financing effects.

### 8.2 Year 1 — investment and proof

| Line | Conservative | Base | Stretch |
|---|---:|---:|---:|
| Average gross time release across eligible roles | 3% | 5% | 7% |
| B1 gross capacity equivalent | $294k | $490k | $686k |
| Capacity realisation ratio after adoption/rework/support | 35% | 45% | 55% |
| **B1 net capacity created — reported separately** | **$103k** | **$221k** | **$377k** |
| B2 external spend reduction | $50k | $100k | $160k |
| B3 hiring/attrition benefit | $0 | $0 | $140k |
| B4 software consolidation | $15k | $30k | $50k |
| B5 direct unit-cost reduction | $10k | $25k | $50k |
| **Gross cash benefit — excludes B1** | **$75k** | **$155k** | **$400k** |
| B8 full AI TCO | ($300k) | ($360k) | ($450k) |
| **Net cash Year 1** | **($225k)** | **($205k)** | **($50k)** |

**Reading:** a negative Year-1 cash result is commercially plausible because setup, change, overlap and capability costs arrive before most cash routes. Capacity created is useful but is not added to the cash result.

### 8.3 Month-24 annual run-rate

| Line | Conservative | Base | Stretch |
|---|---:|---:|---:|
| Gross time release across eligible roles | 8% | 12% | 16% |
| B1 gross capacity equivalent | $784k | $1,176k | $1,568k |
| Capacity realisation ratio | 45% | 60% | 70% |
| **B1 net capacity created — reported separately** | **$353k** | **$706k** | **$1,098k** |
| B2 external spend reduction | $120k | $220k | $320k |
| B3 hiring/attrition benefit | $140k | $280k | $420k |
| B4 software consolidation | $35k | $60k | $90k |
| B5 direct unit-cost reduction | $40k | $100k | $180k |
| **Gross annual cash benefit — excludes B1** | **$335k** | **$660k** | **$1,010k** |
| B8 annual run-rate TCO | ($380k) | ($480k) | ($620k) |
| **Net annual cash run-rate at Month 24** | **($45k)** | **+$180k** | **+$390k** |
| Net cash as % of total people cost | (0.3%) | 1.3% | 2.8% |

### 8.4 Reconciliation and interpretation

- B1 is **not additive** to the cash lines. In the base case, some of the $706k capacity equivalent may support the $280k hiring-avoidance route, the $220k external-spend route and non-cash growth/backlog work. The capacity ledger assigns those hours once.
- The base case reaches a positive **annual run-rate** by Month 24; it does not imply cumulative cash payback has already occurred. Cumulative payback depends on when contracts, vacancies and licences actually change.
- The result is highly sensitive to external spend, planned hiring, workflow concentration, adoption, exception burden and internal delivery cost. A firm with little outsourced spend or growth hiring may create substantial capacity but little near-term cash.
- Do not scale the example linearly by headcount. Small firms have higher fixed-cost and fragmented-capacity effects; larger firms may have stronger scale but more integration, governance and change cost.

## 9. Reporting rhythm

### Monthly

- workflow unit costs and B8 actuals;
- adoption, quality, exception and disbenefit trends;
- B1 capacity measured and allocated;
- benefit-card movements and forecast variance;
- spend/seat/usage anomalies; and
- workflow kill/remediation candidates.

### Quarterly portfolio and harvest review

1. Reconcile scorecard, TCO and material D1 disbenefits.
2. Review each E2–E4 benefit card and overlap checks.
3. Approve or reject capacity allocations and cash conversion actions.
4. Finance co-signs E5 banked cash benefits.
5. Reforecast low/base/high cases using current evidence.
6. Promote, redesign, pause or retire workflows based on net value and risk.
7. Communicate to staff where meaningful capacity has been redeployed.

### Annual

- re-baseline sustained processes;
- reset unit-cost targets, capacity assumptions and cost allocations;
- reconcile cumulative cash investment/benefit and Month-24/run-rate view;
- review workforce and contract plans;
- test whether shared platform costs remain justified; and
- roll forward the three-year portfolio case.

## 10. Anti-patterns and rejected claims

1. Multiplying “minutes saved per email” by every employee and calling the total cash savings.
2. Adding B1 capacity value to B3 hiring avoidance or labour-based B5 savings.
3. Counting a role as avoided when it was never funded or planned before the AI case.
4. Counting a software list price instead of the actual renewal/invoice reduction.
5. Ignoring implementation, control, support, overlap, approval and exception cost.
6. Calling adoption, prompts submitted or agents created a benefit.
7. Claiming revenue uplift without a counterfactual or credible attribution.
8. Treating a short pilot result as a sustained production result.
9. Banking a benefit while quality, complaints, risk or workload has materially worsened.
10. Claiming 100% of measured time as available capacity.
11. Scaling a 100-FTE example linearly to 10 or 500 FTE.
12. Leaving shared platform/team cost unallocated while reporting workflow savings.
13. Counting the same external contract reduction against multiple workflows.
14. Treating avoided hypothetical risk as cash without an accepted valuation method.
15. Keeping a negative-value workflow because it is strategically fashionable.

## 11. Quarterly review agenda

| Item | Time guide |
|---|---:|
| Portfolio outcome, quality, incidents and B8 actuals | 10 min |
| Capacity ledger reconciliation and destination decisions | 10 min |
| Benefit-card review: Measured → Committed → Banked | 20 min |
| Workflow commercial gates: scale, redesign, pause or retire | 10 min |
| People, consultation, reskilling and workload guardrail check | 10 min |
| Contract, licence and reinvestment decisions | 10 min |

## 12. Reference note on productivity evidence

External studies are useful for forming hypotheses but show materially different effects by task and context. For example, field evidence has found meaningful gains in some customer-support settings, while controlled research with experienced open-source developers has found negative productivity effects under particular tools and tasks. The commercial conclusion is not that AI “works” or “does not work” universally; it is that local workflow design, user capability, process fit and measurement determine the result.

As calibration rather than planning inputs: the NBER customer-support field study found roughly a 14% average productivity gain, concentrated among less-experienced agents (around 34%), while the METR randomised study found experienced open-source developers were about 19% *slower* with AI assistance on their own repositories. Published task-level effects therefore span roughly −20% to +35% or more depending on task, user experience, tooling and process fit — a range wide enough that no generic percentage belongs in a business case.

- [NBER — Generative AI at Work](https://www.nber.org/papers/w31161)
- [METR — Early-2025 AI and experienced open-source developer productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)
