# 02 — Phased Implementation Plan

| | |
|---|---|
| **Version** | v1.0 (Final) |
| **Date** | 2026-08-10 |
| **Planning horizon** | Up to 24 months; outcome-gated rather than calendar-driven |
| **Companion docs** | 00 Overview · 01 Target Architecture · 03 Data, Security & AI Governance · 04 Benefits Realisation |

---

## 1. Delivery approach

This is a capability and operating-model program, not a one-off technology implementation.

1. **Run a portfolio, not a parade of pilots.** Every candidate is selected, funded, gated, operated and retired through one process.
2. **Redesign before automation.** For each workflow ask: can the work be eliminated, simplified, standardised or handled deterministically before adding a model?
3. **Separate program maturity from workflow maturity.** The organisation can be in Phase 2 while an individual workflow remains T0/T1 or is retired.
4. **Use thin, end-to-end slices.** Deliver one measurable outcome with identity, controls, evaluation, operations and benefits—not a broad demonstration with no production path.
5. **Promote action classes, not whole agents.** T2/T3 evidence applies to a declared action with fixed scope and limits.
6. **Adoption is necessary, not sufficient.** Usage is monitored, but only accepted outcomes and converted benefits justify scale.
7. **Fund in stages.** Release larger platform and licence commitments only when the portfolio demonstrates demand and value.
8. **Plan for failure and retirement.** Manual fallback, pause, rollback, supplier exit and decommissioning are designed before production.
9. **Treat workforce trust as an operating dependency.** Communicate role effects and measurement principles before broad deployment.

Phases below are indicative. Readiness, integration complexity, legal obligations and organisational change capacity determine duration. S1 organisations may reach an appropriate steady state in 6–12 months; S3 commonly uses the full horizon.

## 2. Program governance

| Role/duty | Accountability | Typical incumbent |
|---|---|---|
| Executive Sponsor | Strategy, funding, risk appetite, workforce principles and material exceptions | CEO, COO or relevant executive |
| AI/Automation Lead | Portfolio, architecture adherence, delivery cadence, register and cross-functional capability | Fractional in S1/S2; dedicated in S3 |
| AI Council / Steering Group | Approves higher-impact use cases, T3 action classes, material exceptions and portfolio investment/retirement | Sponsor, AI Lead, IT/security, privacy/legal as needed, finance and rotating function owner |
| Function Product Owner | Prioritises outcomes, process redesign, adoption and benefit evidence | Function head or delegate |
| Workflow Owner | Owns one production workflow’s acceptance criteria, controls, unit cost and lifecycle | Existing operational manager or product owner |
| Technical Owner | Build/configure, integrations, action boundary, SLOs, deployment, support and incident response | Internal engineer, IT lead or contracted partner |
| Data/Knowledge Owner | Authority, access, quality, source curation, retention and correction | SoR or corpus owner |
| Security/Privacy/Legal Owner | Proportionate control and impact review; obligations and incidents | Existing leads or external adviser/MSP |
| Finance Partner | Baselines, TCO, benefit validation and financial conversion | CFO/finance lead |
| Employee/Change Lead | AI literacy, role guidance, consultation, champions and feedback | People lead or change duty-holder |

**Decision cadence**

- weekly or fortnightly delivery review for active workflows;
- monthly portfolio, risk, cost and incident review;
- quarterly investment and benefits-conversion review; and
- annual strategy, policy, risk appetite and supplier review.

S1 may combine these into one monthly operating review. Governance effort should scale with risk and portfolio size, not reproduce enterprise bureaucracy.

## 3. Workstreams

| WS | Workstream | Scope |
|---|---|---|
| WS1 | Strategy & Portfolio | AI intent, inventory, workflow selection, process redesign, investment and retirement |
| WS2 | Identity & Security | SSO/MFA/lifecycle, delegated and workload identity, secrets, endpoint, egress, action security and logging |
| WS3 | Data & Knowledge | Authority map, permission hygiene, source registry, retrieval controls, data quality, metrics and analytical data where needed |
| WS4 | Platform & Integration | Employee tools, deterministic workflow engine, agent runtime, action broker/policy, tool interfaces, gateway and observability |
| WS5 | Workflow Delivery | Design, build/configure, evaluate, shadow, deploy, operate and improve individual workflows |
| WS6 | People & Adoption | AI literacy, role-based practice, manager guidance, champions, support, workforce dialogue and accessibility |
| WS7 | Governance & Assurance | Policy, AI system register, risk/impact assessments, supplier review, change control, incidents, contestability and assurance |
| WS8 | Benefits & FinOps | Baselines, TCO, capacity/outcome ledger, financial conversion, unit economics and investment review |

## 4. Two-dimensional delivery model

### 4.1 Program phases

The organisation progresses through four maturity phases:

- **Phase 0 — Establish control and evidence**
- **Phase 1 — Governed augmentation**
- **Phase 2 — Controlled execution**
- **Phase 3 — Institutionalise and selectively automate**

### 4.2 Workflow lifecycle

Every workflow follows this lifecycle regardless of program phase:

```text
Discover → Eliminate/simplify → Define contract → Risk/impact assess →
Build/configure → Offline evaluate → Shadow → T0/T1 production →
T2 controlled action (if justified) → T3 bounded action (exceptional) →
Operate/improve → Retire
```

A workflow can be killed at any gate. Failed experiments are useful only if they are closed, documented and stop consuming attention or licences.

## 5. Phase 0 — Establish control and evidence (nominal Months 0–3)

**Objective:** know what AI is already present, make the estate safe enough to amplify, establish accountable ownership, capture baselines and select the first evidence-producing portfolio.

### 5.1 Activities

**WS1 — Strategy & Portfolio**

- State the business intent and explicit non-goals; confirm that unrestricted autonomous operation is out of scope.
- Inventory employee tools, embedded SaaS AI, custom automations, supplier AI and known shadow use.
- Complete the readiness assessment (Appendix B).
- Map 10–20 candidate task families; run the eliminate/simplify/deterministic screen before scoring AI candidates.
- Select 2–4 Phase 1 workflows, including at least one low-risk, high-frequency use case with objective acceptance criteria.

**WS2 — Identity & Security**

- Audit SSO, MFA, privileged access, joiner/mover/leaver and external OAuth grants.
- Require universal MFA; move privileged and financial roles towards phishing-resistant authentication.
- Close shared-account and ex-employee access; document the unsupported SaaS long tail.
- Establish managed secrets and the delegated-user versus workload-identity patterns.
- Choose the audit destination and minimum event schema.

**WS3 — Data & Knowledge**

- Lock down risky external-sharing defaults and remediate the highest-risk links/groups/sites.
- Define the local data classification scheme and owners.
- Create the authoritative source and AI-grounding source registers.
- Identify personal information, C4/regulated data and retention constraints in candidate workflows.
- Do not index broad corpora merely to demonstrate search.

**WS4 — Platform & Integration**

- Select one primary employee-assistant pattern per initial cohort under appropriate business terms.
- Select a deterministic automation/workflow platform suited to existing skills and estate.
- Define the standard typed-action, policy and approval pattern before custom write-enabled agents.
- Record the initial ADRs in Doc 01 §14.
- Do not procure a warehouse, model gateway or custom RAG stack without a use-case trigger.

**WS5 — Workflow Delivery**

- Define workflow contracts: outcome, trigger, source truth, action classes, quality, mode ceiling, exception path, SLO and unit cost.
- Build representative evaluation seed sets from real historical cases.
- Establish kill criteria before build.

**WS6 — People & Adoption**

- Publish the acceptable-use policy and a no-blame shadow-AI disclosure period.
- Deliver foundation AI literacy: capability limits, privacy, security, verification, disclosure and incident reporting.
- Explain what activity/outcome data will and will not be used for; workflow measurement is not covert individual surveillance.
- Nominate champions by team/function where useful, not by a rigid ratio.

**WS7 — Governance & Assurance**

- Establish the AI system register and use-case risk/impact process.
- Identify legal and contractual obligations, including privacy, employment, records, IP and customer requirements.
- Approve tool tiers, prohibited/default-excluded uses, supplier checklist and incident playbooks.
- Document any imminent legal change that affects the program and assign an owner.

**WS8 — Benefits & FinOps**

- Capture workflow baselines, quality/rework, volumes and current unit costs.
- Inventory AI, SaaS, external-services and delivery costs.
- Establish separate capacity/outcome and financial ledgers (Doc 04).
- Define how model, platform, human-review and rework cost will be allocated to workflows.

### 5.2 Deliverables

- approved program charter and risk appetite;
- current AI/system inventory and initial register;
- readiness assessment and remediation plan;
- policy set v1 and legal/obligations register;
- identity, permission and external-sharing remediation report;
- sanctioned employee-tool decision and supplier terms;
- architecture decision log and action-control design;
- baseline/TCO pack; and
- Phase 1 portfolio briefs with workflow contracts, evaluation and kill criteria.

### 5.3 Gate 0 → 1

All applicable criteria are required:

- [ ] Executive Sponsor, AI Lead and workflow owners are named with time and authority.
- [ ] Material AI systems and embedded AI are inventoried; shadow disclosure period completed.
- [ ] MFA is universal; privileged access and leaver controls are tested; critical shared accounts are removed.
- [ ] Highest-risk permission/sharing findings are remediated or have time-bound accepted treatment.
- [ ] Data classification, source ownership and initial legal/privacy obligations are documented.
- [ ] Sanctioned tools operate under reviewed business terms and central administration.
- [ ] Policy, risk/impact intake, incident route and AI system register are live.
- [ ] Local baselines and TCO assumptions exist for all Phase 1 workflows.
- [ ] Each selected workflow has a kill criterion, owner, source truth, manual fallback and mode ceiling.

A target such as “95% SSO coverage” can guide remediation, but gate approval is risk-based: all material systems MUST have acceptable identity and lifecycle controls, even if a harmless long-tail application remains outside SSO.

## 6. Phase 1 — Governed augmentation (nominal Months 3–9)

**Objective:** establish useful employee adoption and 3–5 measured T0/T1 workflows, prove source and evaluation controls, and learn where AI genuinely improves work.

### 6.1 Activities

- Deploy the primary assistant to defined eligible cohorts; use staged licence allocation and reclaim unused seats.
- Curate knowledge sources and test effective permissions, source provenance, freshness and deletion using low-privilege identities.
- Deliver 3–5 T0/T1 workflows across no more than 2–3 functions initially.
- Use models for ambiguous interpretation/drafting and deterministic code for calculations, validation and policy.
- Operate material workflows in shadow mode before relying on outputs.
- Implement repeatable evaluations and production telemetry: quality, acceptance, override, rework, latency and cost.
- Establish a workflow support queue and named fallback.
- Run role-based training using the actual tools and workflows, followed by office hours/guilds and manager coaching.
- Start monthly portfolio and quarterly benefits reviews; retire licences and experiments that fail value or adoption thresholds.
- For S2/S3, introduce the first read-scoped tool/MCP interfaces only where native integrations do not suffice.
- Build an analytical store only where a selected workflow or benefit baseline requires cross-system history.

### 6.2 Gate 1 → 2

- [ ] Eligible-cohort adoption is sufficient to assess value; unused licences are being reclaimed rather than hidden by blanket targets.
- [ ] At least three workflows have representative evaluation results and measured outcome/quality change versus baseline.
- [ ] Source-permission leak tests pass; citations/provenance are available for material knowledge outputs.
- [ ] All live systems are registered, owned, risk/impact-assessed and supported.
- [ ] Unit cost includes model/platform, human-review and rework effort.
- [ ] No unresolved critical/high incident or overdue control treatment blocks scale.
- [ ] At least one workflow has been killed, narrowed or redesigned where evidence did not justify continuation—or the council has documented why all survived.
- [ ] Candidate T2 action classes pass the production-readiness checklist in Appendix D and have completed shadow operation.

No organisation should interpret “high assistant weekly active use” as a universal objective. The target is sustained, appropriate use among eligible roles that produces accepted outcomes at acceptable TCO.

## 7. Phase 2 — Controlled execution (nominal Months 9–18)

**Objective:** move selected action classes from drafts to T2 execution, harden runtime and operational controls, and scale only workflows with demonstrated value.

### 7.1 Activities

**Architecture and operations**

- Implement the standard action broker/policy/approval pattern for custom write-enabled workflows.
- Use dedicated short-lived workload identity or preserved delegated-user identity as appropriate.
- Add durable execution: idempotency, retries, timeouts, dead-letter handling, verification and compensation.
- Implement authenticated approvals with evidence, record diff, authority checks and separation of duties where required.
- Establish SLOs, error budgets, alerting, on-call/escalation, pause/recovery and provider-outage procedures.
- Expand central model/API controls or introduce a gateway only when Doc 01 trigger criteria are met.
- Restrict runtime egress, code execution and tool access; pin and review MCP/tool versions.

**Workflow portfolio**

- Promote proven action classes to T2; do not promote an entire agent wholesale.
- Add workflows in finance, support, sales, operations and engineering only where local data/process maturity permits.
- Prefer action classes with clear ground truth, high frequency, bounded scope and reversible writes.
- Run canary deployment, intensified sampling and rollback for model, prompt, tool or policy changes.
- Track approval queue load and quality; redesign if approvers become a bottleneck or rubber stamp.

**People and benefits**

- Train authorised approvers and reviewers on specific failure modes, evidence and intervention.
- Make role and capacity changes explicit; allocate released capacity to named priorities.
- Convert benefits only through Doc 04 evidence rules; finance signs off cash impacts.
- Consolidate software or external spend only after the replacement capability has met SLOs for a sustained period.

### 7.2 Gate 2 → 3

- [ ] At least 2 (S1), 4 (S2) or 5 (S3) production workflows include a T2 action class with accepted quality and economics; volume is adjusted to actual business need.
- [ ] Every T2 action is attributable, policy-checked, bound to an authenticated approval and verified in the SoR.
- [ ] Production identities, scopes, tools, budgets, versions, runbooks and pause procedures match the register.
- [ ] SLOs, error budgets, human-review load and cost are reported monthly.
- [ ] Incident and restore/pause exercises have passed; manual fallback remains usable.
- [ ] Benefits reported as financial are finance-validated and not double counted with capacity.
- [ ] Each proposed T3 action class meets the additional criteria in Appendix D, including low impact, bounded scope, sustained evidence and auto-pause.

## 8. Phase 3 — Institutionalise and selectively automate (nominal Months 18–24)

**Objective:** embed the operating model, redesign selected end-to-end processes, permit only justified T3 action classes, and establish a sustainable Year-3 portfolio.

### 8.1 Activities

- Promote a small number of eligible action classes to T3; keep sensitive, material, rights-affecting and irreversible actions at T1/T2 or outside AI execution.
- Redesign 2–3 end-to-end processes around human judgment, deterministic controls and AI-enabled work rather than copying legacy steps.
- Consolidate redundant AI, automation and point products based on proven replacement and exit evidence.
- Run role-evolution, reskilling and internal-mobility plans where task compression is durable.
- Conduct structured assurance against the organisation’s selected baseline (for example NIST AI RMF and the Australian six essential practices); obtain independent review for higher-impact systems.
- Complete model/tool/provider exit tests for critical workflows.
- Refresh risk appetite, policies, legal obligations, supplier register, evaluation sets and architecture decisions.
- Approve Year-3 operating budget and portfolio based on accepted outcomes, TCO and risk—not forecast enthusiasm.

### 8.2 Month-24 exit criteria

- [ ] Maturity L4 is achieved for the relevant scale; selected functions show L5 operating characteristics.
- [ ] The AI system register, source register, action controls and change records are complete and current.
- [ ] T2/T3 action spot-checks reconstruct identity, policy, evidence, approval, execution and result end-to-end.
- [ ] All T3 action classes remain within the Doc 01 ceiling and meet current SLO/error-budget evidence.
- [ ] Manual fallback, supplier outage and decommissioning paths have been tested for material workflows.
- [ ] Capacity, service/quality and financial benefits are reported separately; TCO and net cash effect reconcile to finance.
- [ ] Steady-state roles, support, assurance, training and budgets are approved.

## 9. Workflow selection framework

### 9.1 Mandatory decision sequence

For every task family:

1. **Eliminate:** Is the output or control still needed?
2. **Simplify:** Can policy, form design, data quality or process ownership remove the work?
3. **Standardise:** Can templates, master data or a single SoR reduce variation?
4. **Deterministic automation:** Can rules, APIs, OCR or conventional software solve it more reliably?
5. **AI assistance:** Does language, judgment under ambiguity or unstructured content justify a model?
6. **Agentic execution:** Does allowing the system to choose/use tools create additional net value after control and support cost?

### 9.2 Weighted score

Score 1–5 and multiply by the weight.

| Criterion | Weight | What a high score means |
|---|---:|---|
| Outcome value | 20% | High volume/cost, service or risk impact; clear unit of outcome |
| Process suitability | 15% | Stable enough to define, but contains ambiguity AI can help resolve |
| Data/source readiness | 15% | Authoritative, accessible, lawful, representative and sufficiently clean inputs |
| Control and reversibility | 15% | Actions are bounded, testable, reversible/compensatable and can be policy-checked |
| Risk/impact fit | 15% | Low/moderate impact with proportionate controls; no default-excluded use |
| Owner/adoption pull | 10% | Named owner and users actively want the outcome and can change the process |
| Repeatability/scale | 10% | Frequent enough to support evidence and repay implementation/operation |

Apply a **complexity penalty** for open-ended goals, many tools, internet/email exposure, long plans, dynamic sub-agents, browser automation, C4 data, weak ground truth, low frequency or material external commitments.

High score does not override a prohibited use or legal obligation. The initial portfolio SHOULD contain mostly T0/T1 workflows and deterministic automation, not a quota of agents.

### 9.3 Portfolio balance

Maintain a mix of:

- **employee capability:** assistants and training;
- **knowledge access:** curated T0 search/analysis;
- **operational workflows:** T1/T2 with measurable units;
- **foundation work:** permissions, source quality, identity and integration; and
- **experiments:** small, time-boxed and explicitly disposable.

Do not let platform work consume the program without workflow demand, or let workflow pilots bypass platform controls.

## 10. Indicative timeline

```text
Month:      0──3             3──9                    9──18                    18──24
Program:    CONTROL          AUGMENT                  CONTROLLED EXECUTION     INSTITUTIONALISE
            inventory        employee capability     T2 action classes       selective T3
            identity/data    T0/T1 evidence           durable operations      process redesign
            policy/baseline  curated knowledge        action boundary         assurance/steady state

Per workflow:
            discover/simplify → contract/risk → evaluate/shadow → T0/T1 → T2 → T3 if justified → retire
```

Calendar progress never substitutes for gate evidence.

## 11. Resourcing by size band

Figures are indicative delivery capacity, not a requirement to create separate jobs. Existing employees, an MSP and specialist partners can supply the capacity, but internal accountability cannot be outsourced.

| Capability | S1 (1–20) | S2 (21–100) | S3 (101–500) |
|---|---|---|---|
| Sponsor / business ownership | Founder/COO duty | Executive + function owners | Executive sponsor + portfolio owners |
| AI/Automation Lead | 0.1–0.3 FTE, combined role | 0.3–0.7 FTE | ~1.0 FTE |
| Automation/agent engineering | Partner bursts; avoid standing platform | 0.5–1.0 FTE equivalent internal/partner | 1–3 FTE depending on portfolio/complexity |
| Data/analytics | Existing analyst/finance; use-case only | 0.2–0.5 FTE | 0.5–1.5 FTE |
| Security/privacy/legal | MSP/adviser as needed | 0.1–0.3 FTE equivalent | 0.3–1.0 FTE across existing specialists |
| Change/enablement | Owner + team champions | 0.1–0.3 FTE | 0.3–0.7 FTE |
| Finance/benefits | Existing finance duty | 0.1–0.2 FTE | 0.2–0.4 FTE |

Do not add the top of every range to create an assumed headcount requirement. Workload volume, supplier mix, code ownership and risk determine the actual shape.

## 12. Budget and commercial controls

Build the budget by category rather than using a universal “AI cost per employee”:

- employee-assistant licences and specialist seats;
- API/model usage and reserved/committed spend;
- automation, agent, gateway, search/RAG and observability platforms;
- data integration/storage/analytics required by selected workflows;
- implementation and integration;
- internal product/engineering/data/security/change capacity;
- evaluation, red-team and independent assurance;
- support, incident response and business continuity;
- training and workforce transition; and
- exit/decommissioning.

Commercial guardrails:

- stage seat purchases and reclaim unused licences;
- avoid multi-year volume commitments before measured use unless the discount and exit terms justify them;
- require model/version-change, incident, subprocessor, deletion, export and audit-log terms appropriate to the use;
- distinguish subscription entitlements from metered agent/API usage;
- cap cost per workflow and cost per accepted outcome;
- include human approval, rework and exception handling in unit economics; and
- fund platform components only against a multi-workflow demand case or mandatory control need.

## 13. Principal program risks

| Risk | Response |
|---|---|
| Permission debt becomes an internal data breach | Phase 0 remediation; curated sources; security trimming and leak/deletion tests before grounding |
| Employee tools proliferate faster than governance | One primary cohort pattern; inventory embedded AI; reclaim seats; DLP and procurement controls |
| Agentic solution chosen for a rules problem | Mandatory eliminate/simplify/deterministic screen and architecture review |
| Model output bypasses business controls | Typed proposal, deterministic policy and action broker; no raw output to transactions |
| Prompt injection or poisoned content drives action | Trust-labelled context, tool limits, isolation, output validation, T2 approval and T3 ceiling |
| Pilot purgatory | Workflow contract, production owner, gate dates and kill criteria at intake |
| “Golden set” overstates production reliability | Representative/adversarial tests, holdout, shadow mode, production SLOs and sampled review |
| Human approval becomes a rubber stamp | Evidence-rich UI, authority checks, queue/load metrics, review quality triggers and demotion |
| Model/provider change silently degrades workflows | Version monitoring/pinning, regression, canary, fallback and rollback |
| Token or licence cost expands without outcome | Per-workflow allocation, budgets, seat reclamation and cost-per-accepted-outcome limits |
| Benefits double counted | Separate capacity and financial ledgers; finance validates conversion |
| Workforce trust collapses | Early disclosure, no covert individual surveillance, clear redeployment principles, training and consultation |
| Key-person/partner dependency | Owned artefacts, runbooks, source access, deployment rights, exit test and secondary cover |
| Custom platform overwhelms SME capacity | Size-band defaults, managed foundations and build/buy gate |
| Legal obligations emerge after design | Obligations register, material-change review and privacy/legal participation at intake |

## 14. Sequencing rules

1. Inventory and risk appetite before broad procurement.
2. Identity, permissions and supplier terms before non-public data is grounded or shared.
3. Baseline and workflow contract before build.
4. Deterministic process and action boundary before write-enabled model steps.
5. Offline evaluation before shadow; shadow before reliance; sustained T2 evidence before T3.
6. Read scope before write scope; narrow action before broader action.
7. Operations, incident and fallback readiness before production.
8. Financial conversion only after sustained evidence and an approved change to a plan or cost.
9. Tool consolidation only after replacement SLOs hold and exit/restore are tested.
10. Decommission credentials, indexes, data and licences when a workflow retires.

## 15. First 90 days

1. Appoint sponsor, AI Lead and finance/risk counterparts; approve scope and risk appetite.
2. Inventory employee, embedded, supplier and shadow AI; open a time-limited no-blame disclosure channel.
3. Complete readiness, SSO/MFA/lifecycle and high-risk sharing reviews; begin remediation.
4. Publish acceptable-use, data/tool tier and incident rules; create the AI system register.
5. Select the primary employee-assistant pattern for one or two eligible cohorts under reviewed terms.
6. Map and baseline 10–20 task families; select 2–4 workflows after simplify/deterministic screening.
7. Define the standard workflow contract, evaluation pattern and typed action-control architecture.
8. Deliver foundation literacy and role-based practice; establish support and feedback.
9. Ship one or two low-risk T0/T1 workflows through offline evaluation and shadow mode.
10. Hold the first evidence gate: scale, narrow, redesign or kill each workflow.

## 16. Monthly portfolio scorecard

Report by function and workflow where applicable:

- eligible users, active use and licence utilisation;
- workflows/action classes by lifecycle stage and execution mode;
- accepted outcome volume and quality/SLO attainment;
- evaluation pass rates, production overrides, rework and sampling failures;
- incidents, policy violations, pause events and overdue treatments;
- approval/exception queue age, decision effort and escalation;
- model/platform/human-review cost and cost per accepted outcome;
- capacity created, service/quality changes and finance-validated cash benefits;
- training/role coverage and user/affected-person feedback; and
- permission, source, identity, supplier and register control status.

---

## Appendix A — Reference workflow catalogue

The catalogue suggests candidates; it does not pre-approve them. Local process, data, law, risk and economics determine selection and mode ceiling.

| ID | Workflow | Initial mode | Possible ceiling | Critical controls and evidence | Primary unit |
|---|---|---|---|---|---|
| WF-01 | Support reply drafting | T1 | T2 send for pre-approved known categories | Grounded KB, customer context minimisation, tone/correctness, escalation, authenticated approval | Accepted cost/ticket |
| WF-02 | Ticket triage/tag/routing | T0/T2 | T3 for reversible internal routing | Routing accuracy/recall by category, priority false-negative limit, easy undo and queue monitoring | Time to correct queue |
| WF-03 | KB draft from resolved tickets | T1 | T1 | Source citations, redaction, editor acceptance and stale-content controls | Accepted article cost |
| WF-04 | AP invoice extraction/coding | T1 | T2 posting; T3 only for deterministic matched class if locally approved | Duplicate, arithmetic/tax, vendor/PO/master-data checks, bank-detail controls, amount authority and verification | Cost/accepted invoice |
| WF-05 | AR reminder drafting | T1 | T2 send | Ledger truth, dispute/ hardship/strategic-account exclusions, tone, contact consent and approval | Cost/collection cycle |
| WF-06 | Management reporting pack | T1 | T1 | Figures reconciled to governed metrics; commentary linked to data; finance sign-off | Close/preparation hours |
| WF-07 | Meeting-to-CRM notes/tasks | T1/T2 | T3 for low-impact internal fields | Consent/transcript rules, field-level validation, no forecast/commitment overwrite, undo | Admin cost/meeting |
| WF-08 | Account research brief | T0/T1 | T1 | Source quality/date, factuality, conflict and privacy rules | Prep cost/meeting |
| WF-09 | Proposal/tender first draft | T1 | T1 | Approved facts/pricing, source/template control, legal/commercial review | Cost/accepted draft |
| WF-10 | Marketing content variants | T1 | T2 publish for low-risk pre-approved channels only | Brand, claim, rights, disclosure and channel approval | Cost/accepted asset |
| WF-11 | Candidate application summary | T1 | **T1 hard ceiling** | No ranking/recommendation; job-related rubric; privacy, bias/adverse-impact tests; human decision and contestability | Recruiter time/application |
| WF-12 | Employee onboarding pack/tasks | T1/T2 | T2 | HRIS truth, least data, role/template completeness, approval and failed-task queue | Cost/hire |
| WF-13 | Contract first-pass issue spotting | T1 | **T1 hard ceiling** | Clause library, recall-weighted evaluation, jurisdiction limits and qualified legal review | Cost/reviewed contract |
| WF-14 | Coding assistance and PR summary | T1 | T2 for opening a PR; human merge | Repository scope, secrets/code controls, tests/security scan, maintainability and human review | Cost/accepted change |
| WF-15 | Organisation knowledge Q&A | T0 | T0 | Curated sources, security trimming, citations, freshness/deletion and “not found” behaviour | Cost/accepted answer |

**Default-excluded examples:** final hiring/termination/performance decisions; ranking people without a lawful and approved design; payment release or bank-detail change; signing contracts; unreviewed legal/medical/financial advice; unrestricted shell/database/admin access; biometric emotion or personality inference; and dynamic creation of unapproved tools, privileges or sub-agents.

## Appendix B — Readiness assessment

Score 1–5 before program commitment and at each phase gate.

| # | Dimension | A score of 5 | A score of 1 | Score |
|---|---|---|---|---|
| R1 | Strategy and leadership | Named sponsor, funded outcomes, clear risk appetite and willingness to redesign work | Tool purchase delegated downward; no outcome or change commitment | /5 |
| R2 | Identity and cyber baseline | Material apps under strong identity/lifecycle controls; privileged access and incidents managed | Password/shared-account sprawl; leavers or admin access uncontrolled | /5 |
| R3 | Data and permission readiness | Authority/owners known; sharing controlled; candidate data lawful, accessible and sufficiently clean | Unknown authority; broad sharing; no owner; poor input quality | /5 |
| R4 | Process and integration readiness | Stable outcome/process, supported APIs/events and clear exception path | Highly variable undocumented process; browser-only critical system | /5 |
| R5 | Measurement and finance | Baselines, outcome units, quality and TCO can be measured; finance engaged | Anecdotal value only; no volume/cost/quality evidence | /5 |
| R6 | Governance, privacy and legal | Inventory, risk/impact route, obligations, supplier controls and incident process exist | No accountable owner or awareness of obligations/affected people | /5 |
| R7 | Skills and workforce trust | Role-based capability, curious owners, clear measurement/redeployment principles | Fear/prohibition or uncontrolled enthusiasm; no time to learn | /5 |
| R8 | Delivery and operations | Technical owner, test/deploy/support/fallback capability and partner exit rights | Demo capability only; no production owner or support route | /5 |

**Interpretation (/40)**

| Total | Reading | Response |
|---:|---|---|
| 34–40 | Strong starting position | Phase 0 may be compressed, but gates remain |
| 26–33 | Standard | Run Phase 0 as designed and treat low dimensions explicitly |
| 18–25 | Remediation-weighted | Extend Phase 0; limit real-data pilots; reduce portfolio breadth |
| <18 | Not program-ready | Run focused identity/data/governance/process remediation and re-score |

Any dimension ≤2 becomes a named prerequisite with an owner. A high total does not compensate for a critical identity, privacy or operational weakness in the selected workflow.

## Appendix C — Evaluation and evidence pattern

### C.1 Define the workflow contract

Document:

- exact outcome and acceptance unit;
- in-scope and out-of-scope cases;
- authoritative sources and expected outputs;
- action schema and deterministic validations;
- tolerated errors by type and their cost/impact;
- escalation/refusal behaviour;
- latency, cost and availability targets; and
- mode ceiling and promotion/demotion rules.

### C.2 Build representative tests

1. Start with roughly 30–50 real cases for a low-risk seed set; use more where the task has many categories or edge conditions.
2. Stratify routine, uncommon, difficult and known-failure cases rather than sampling only “clean” examples.
3. Add adversarial cases for internet, email, document, tool or memory exposure: indirect prompt injection, data exfiltration requests, malformed input and poisoned sources.
4. Define exact expected values where possible; otherwise use a 3–5 criterion rubric calibrated with domain experts.
5. Keep an untouched holdout and version cases, labels, rubric and reviewer decisions.
6. Check differential performance where people or protected cohorts may be affected.

### C.3 Evaluate the whole system

Test not just the model response but:

- retrieval permissions and source quality;
- tool selection and arguments;
- deterministic policy outcomes;
- transaction success/idempotency/compensation;
- human approval information and authority;
- incident/fallback behaviour; and
- total cost and review/rework.

### C.4 Shadow and production evidence

- Run against live inputs without relying on or executing output.
- Compare to actual human/process outcomes and capture false positives, false negatives and abstentions.
- Size evidence to risk and tolerated error. For common T3 candidates, expect hundreds of representative shadow/T2 observations rather than a small static set; use confidence intervals or conservative upper bounds for material error rates.
- Deploy T2/T3 changes by canary where possible and intensify review after changes.
- Feed real failures, overrides, incidents and new input classes into the test set.

### C.5 Promotion and demotion

Promotion requires all applicable Appendix D controls, accepted offline and production evidence, an owner and council approval at T3. Demote or pause when:

- a critical policy/security violation occurs;
- an error-budget or quality threshold is breached;
- input/data/model/tool change invalidates evidence;
- approval/sampling reveals systematic failure;
- cost per accepted outcome exceeds the stop threshold; or
- the manual fallback is unavailable.

## Appendix D — Production-readiness checklist

### D.1 Required before T0/T1 production

**Business and people**

- [ ] Workflow owner, users, outcome, scope, exclusions and manual fallback are documented.
- [ ] Process has been simplified and deterministic alternatives considered.
- [ ] Affected employees/users are trained; disclosure/contestability requirements are met.

**Data and legal**

- [ ] Sources, authority, classifications, lawful purpose, retention, rights and residency are known.
- [ ] Retrieval permissions, provenance, freshness and deletion have been tested.
- [ ] Privacy/impact assessment is complete where triggered.

**System and security**

- [ ] System is registered; supplier/model/tool versions and responsibilities are documented.
- [ ] Identity, least privilege, secrets, egress and environment separation are implemented.
- [ ] Threat model covers injection, disclosure, poisoning, tool misuse, supply chain and failure.

**Evaluation and operations**

- [ ] Workflow-level evaluation and holdout meet acceptance criteria.
- [ ] Logs/traces/records and retention are configured proportionately.
- [ ] SLOs, cost allocation, support, incident, pause and decommissioning routes exist.

### D.2 Additional before T2

- [ ] Typed action schema, deterministic validations and policy decision are versioned and tested.
- [ ] Approver is authenticated, authorised and shown the exact proposal, evidence, materiality and exceptions.
- [ ] Executor uses narrow write authority, idempotency, verification and compensation/exception handling.
- [ ] Shadow evidence covers representative live cases and known attacks/failures.
- [ ] Approval and exception workload is viable with cover and service levels.
- [ ] Every action can be reconstructed without retaining unnecessary raw sensitive content.

### D.3 Additional before T3

- [ ] Action class is narrow, low-impact, normally non-sensitive and reversible/compensatable.
- [ ] Open-ended planning, dynamic tools/sub-agents and privilege expansion are impossible by design.
- [ ] Sustained T2/shadow evidence is sufficient for the tolerated error and impact, not merely a calendar period.
- [ ] Hard value/rate/time/customer/data limits and anomaly auto-pause are enforced outside the model.
- [ ] Continuous sampling, error budget, owner review and escalation are operating.
- [ ] Pause, rollback/compensation and manual continuity have been exercised.
- [ ] Council approval records residual risk and expiry/review date.
