# 04 — Benefits Realisation Plan

| | |
|---|---|
| **Version** | v1.0 (Final) |
| **Date** | 2026-08-10 |
| **Goal** | Prove workflow value, convert eligible value through explicit routes, report full cost and prevent double counting |
| **Companion docs** | 00 Overview · 01 Target Architecture · 02 Implementation Plan · 03 Data, Security & AI Governance |

## 1. Principles

1. **Baseline before business-case approval.** Discovery may begin earlier; no quantified benefit is approved without a credible current-state measure.
2. **Capacity is not cash.** Reduced task effort is reported as capacity created. It becomes a cash effect only through an evidenced spend, hiring or workforce route.
3. **Capacity and the cash route are not additive.** When capacity enables hiring avoidance, attrition reshaping or in-housing, the underlying hours remain one resource and are never summed twice.
4. **Count net of full TCO.** Licences and tokens are only part of cost. Include implementation, integration, controls, training, support, overlap, failure/rework and retirement.
5. **Separate outcome classes.** Cash, capacity, revenue, quality/risk and disbenefits are reported in separate columns. A board total combines only compatible categories.
6. **Use actual adoption and sustained performance.** A laboratory time saving is reduced by utilisation, exceptions, rework, quality guardrails and operating downtime.
7. **Bank only evidenced cash effects.** Forecast, committed and banked are distinct states; finance co-signs banked values.
8. **No double counting across workflows.** Shared tasks, headcount plans, external spend and software licences have one card/owner and a clear allocation rule.
9. **Protect quality, risk and resilience.** A saving that increases material error, complaint, security, safety or employee/customer harm is not a net benefit.
10. **Use conservative confidence and sensitivity.** AI productivity effects vary materially by task, user, model, process and implementation; local evidence overrides generic claims.
11. **Redeploy transparently.** In a growing SME, the default use of capacity is usually growth, service, quality and backlog—not covert redundancy.
12. **Re-baseline after sustained change.** Yesterday’s benefit becomes tomorrow’s normal operating baseline.

## 2. Benefit and cost taxonomy

| ID | Class | Definition | Cash treatment | Evidence standard |
|---|---|---|---|---|
| **B1** | Capacity created | Net internal hours removed from in-scope work after adoption, exception, rework and quality adjustment | **Non-cash** until allocated through a route | Repeated measurement versus baseline; owner accepts available capacity |
| **B2** | External spend reduction | Agency, contractor, outsourced service or professional-fee spend reduced | Cash benefit | Contract/PO/statement-of-work reduced or invoice run-rate falls, net of replacement cost |
| **B3** | Hiring avoidance or attrition-based reshaping | A pre-existing funded role is deferred/cancelled, or a vacancy is not backfilled after durable task compression | Cash benefit | Pre-existing workforce plan or approved vacancy; budget/headcount decision recorded; capacity sufficiency demonstrated |
| **B4** | Software/licence consolidation | Subscription, seat, platform or point tool retired/right-sized | Cash benefit | Renewal/order/invoice reduced; replacement and migration cost included |
| **B5** | Direct unit-cost reduction | Non-labour transaction cost, error/rework cost, postage/processing or other per-unit cost falls; labour may be included only if not also claimed through B1/B3 | Cash benefit | Component-level unit-cost evidence sustained for an approved period |
| **B6** | Revenue enablement | Greater throughput, conversion, retention, price realisation or faster time-to-revenue | Separate; cash only with credible attribution | Leading indicators by default; bank only where finance accepts causal evidence and avoids overlap |
| **B7** | Quality, risk and experience | Fewer errors, faster service, better compliance, resilience, employee/customer experience or reduced risk exposure | Usually non-cash; cash where avoided/rework cost is measured | Defined quality/risk metric, baseline and sustained change; monetisation method documented |
| **B8** | AI total cost of ownership | All one-time and recurring cost attributable to the capability and its operation | Negative cash line | Actuals/forecast from finance, contracts, cloud and allocated internal/partner effort |
| **D1** | Disbenefits | New rework, review burden, delay, incidents, complaints, lock-in, lost flexibility, quality decline or change cost | Negative cash and/or non-cash | Measured or provisioned; linked to workflow and corrective action |

### 2.1 Board reporting rule

The board-level cash number is:

```text
Banked B2 + B3 + B4 + cash-qualified B5 + cash-qualified B6/B7
− actual B8 − cash-valued D1
```

B1 capacity created is shown separately. It is never added to B3 or labour-based B5. Non-cash B6/B7 and non-cash D1 are reported as operational outcomes, not hidden.

### 2.2 Presenting the value story

The cash number in §2.1 is deliberately conservative and must never be presented alone. Every board or owner report shows three lines together, in this order:

1. **Net cash** — the §2.1 formula, finance-validated;
2. **Capacity created and where it went** — the ledger destinations in §4.4, the operating dividend of the program; and
3. **Quality, risk and service movements, including disbenefits** — the guardrail evidence.

In a growth-mode SME the cash line is commonly negative in Year 1 and modest at Month 24 (§8). The commercial case is normally carried by capacity redeployed to funded priorities, service and quality effects, and the option value of a governed capability. Leading with cash alone understates a sound program; leading with capacity alone overstates it. Present all three lines every quarter, and let §8.4 supply the interpretation.

## 3. Where value may arise

The table is a discovery map, not a benchmark. It deliberately avoids asserting universal percentage gains.

| Function | Candidate workflow families | Primary baseline | Likely value routes | Critical counter-metrics |
|---|---|---|---|---|
| Finance | Invoice capture/match/post, AR reminders, expense review, reporting preparation, close support | Volume, touch/cycle time, exception/rework, outsourced cost, cost per transaction | B1, B2, B3, B5 | Duplicate/error, tax/coding accuracy, close quality, fraud/SoD |
| Customer support | Triage, draft replies, known-issue actions, KB drafting, QA summaries | Handle time, first response, resolution, recontact, backlog, outsourced L1 | B1, B2, B3, B5, B7 | Resolution quality, complaints, escalation, customer effort |
| Sales | Research, call prep, notes-to-CRM, proposal drafts, hygiene | Admin hours, proposal cycle, CRM completeness, pipeline coverage | B1, B3, B6 | Accuracy, customer trust, selling time actually redeployed |
| Marketing | Brief/asset drafts, variants, reporting, routine publishing | Internal/agency hours, cost per asset, cycle, output volume | B1, B2, B4, B6 | Brand/legal quality, rework, channel performance |
| People/HR | Administrative summaries, policy/JD drafts, onboarding packs, employee FAQ | Admin time, recruiter/outsourcing spend, time to onboard | B1, B2, B7 | Fairness, privacy, decision quality, employee experience |
| Operations/admin | Scheduling, document processing, vendor communications, SOP drafting | Volume, queue time, touch time, manual errors | B1, B3, B4, B5 | Missed exceptions, service continuity, supplier/customer impact |
| Engineering/product | Coding support, tests, PR drafts, documentation, issue analysis | Cycle time, review time, defects, throughput, contractor spend | B1, B2, B3, B6, B7 | Escaped defects, review burden, security, maintainability |
| Legal/commercial | First-pass issue spotting, template population, obligation extraction | Internal/counsel time, turnaround, issue recall | B1, B2, B7 | Missed material issues, false reassurance, confidentiality |
| Cross-functional | Knowledge search, meeting actions, internal drafting | Search/admin time, answer success, meeting follow-through | B1, B4, B7 | Leakage, unsupported answers, meeting/action quality |

### 3.1 Software-consolidation discovery

Review:

- duplicate general AI assistants;
- transcription/meeting tools now covered by a sanctioned suite;
- single-purpose drafting, research, OCR or summarisation tools;
- overlapping automation/iPaaS subscriptions;
- unused or low-value premium seats;
- duplicate enterprise-search or chatbot products;
- bespoke pilots that became vendor-supported features; and
- development/test subscriptions left running after experiments.

Do not assume a fixed percentage of SaaS spend is recoverable. Use actual contracts, renewal dates, utilisation, migration cost and feature gaps.

## 4. Measurement system

### 4.1 Baselines

For each selected workflow, record:

- annual/monthly volume and seasonality;
- current touch time and elapsed cycle time;
- role mix and fully loaded cost methodology;
- exception, correction, rework and escalation rates;
- service/quality/customer metrics;
- external spend and software cost attributable to the process;
- planned hires or vacancies relevant to future capacity;
- current controls and manual review cost; and
- confidence, sample period and known limitations of the baseline.

Use a practical method: system event data, ticket/CRM/finance analytics, a representative work sample, structured time study or a combination. Avoid covert employee surveillance; measure workflows and outcomes.

### 4.2 Ongoing measures

| Layer | Measures |
|---|---|
| Adoption/utilisation | Eligible users/transactions, successful completion, depth, abandonment and fallback |
| Workflow effort | Human touch time, approval time, exception/rework and support effort |
| Service/outcome | Cycle time, throughput, backlog, resolution, accuracy and customer/user result |
| Quality/risk | Error severity, complaints, incidents, override, bias/impact, reconciliation and control failure |
| Reliability | Availability, failed/partial runs, retries, exception age and manual fallback |
| Unit economics | Fully defined cost per outcome and marginal cost at volume |
| B8 TCO | Licences, model usage, platform, people, partner, security/privacy, change, support, overlap and retirement |
| Benefits | B1 capacity, B2–B7 by evidence state, D1 and cash/net result |

### 4.3 Capacity-created formula

A practical calculation is:

```text
Gross hours released
= eligible transaction volume × (baseline touch time − new touch time)

Net capacity created
= gross hours released
× successful adoption/utilisation rate
− additional approval, exception, rework, support and recovery hours
```

Convert to FTE-equivalent or loaded-cost equivalent for planning, but keep it labelled **capacity**, not cash.

Where quality changes, adjust the calculation or fail the benefit gate. Faster work with more material errors is not capacity creation.

### 4.4 Capacity ledger

Every measured pool of released time is allocated once.

`Period · function/role · source workflow(s) · net hours · confidence · owner · destination · cash route (if any) · evidence · status`

Allowed destinations:

- **R1 — growth/service redeployment**: more pipeline, product, service, customer or market work;
- **R2 — backlog/quality/resilience**: documentation, controls, technical debt, training, leave coverage or service improvement;
- **R3 — external-spend in-housing**: supports B2, but the internal hours are not also added as a cash saving;
- **R4 — hiring avoidance**: supports B3 against a documented plan;
- **R5 — attrition-based reshaping**: supports B3 when a vacancy is not backfilled;
- **R6 — retained operating slack**: resilience or workload normalisation, reported honestly as such; and
- **R7 — not realised**: fragmented, absorbed, lost to demand growth or not accepted by the manager.

### 4.5 Benefit card

```text
Card ID / benefit class / workflow(s) / owner
Baseline value, date, method and confidence
Mechanism and counter-metrics
Forecast: low/base/high and evidence grade
Capacity dependency and ledger allocation, if any
Cash conversion route and date
One-time and recurring B8 cost allocation
Disbenefits/risks and quality guardrails
Status: Hypothesis → Baseline validated → Measured → Committed → Banked → Reversed/closed
Evidence links / finance sign-off / banked amount and date
Overlap check with other cards
```

### 4.6 Evidence grades

| Grade | Meaning | Use |
|---|---|---|
| **E0 — Hypothesis** | Directional assumption with no local baseline | Discovery only; not in committed plan |
| **E1 — Baseline validated** | Current state measured with known confidence | Forecast input |
| **E2 — Pilot measured** | Controlled local result, not yet sustained at production scale | Risk-adjusted forecast |
| **E3 — Production sustained** | Result sustained over representative volume/business cycle | May move to committed route |
| **E4 — Committed** | Budget/spend/headcount owner has approved conversion action and date | Committed benefit, not yet banked |
| **E5 — Banked** | P&L/budget/contract effect occurred and finance verified it | Board cash benefit |

A forecast MAY be probability-weighted for planning, but a banked actual is never probability-discounted.

## 5. Full AI total cost of ownership (B8)

### 5.1 Cost categories

| Category | Examples |
|---|---|
| Employee licences | Suite copilot, enterprise assistant, specialist AI seats |
| Model/agent usage | Tokens, searches, agent actions, tool calls, embeddings, storage and overages |
| Platform/integration | iPaaS, gateway, observability, retrieval/search, warehouse, queues and managed compute |
| Build/configuration | Internal engineering/analysis, implementation partner and test-data preparation |
| Security/privacy/assurance | Vendor review, impact/DPIA, testing, logging, DLP, penetration/adversarial testing and audit |
| Change/people | Training, communications, workflow redesign, consultation and manager/approver time |
| Run/support | Monitoring, incident response, exception handling, evaluations, prompt/source maintenance and on-call/owner time |
| Transition/overlap | Dual licences, parallel process, migration, data clean-up and temporary productivity dip |
| Failure/disbenefit | Rework, customer remediation, incident cost, wrong transactions and risk allowance |
| Exit/retirement | Export, migration, contract termination, deletion verification and decommissioning |

Allocate shared cost to workflows using a documented driver such as active users, actions, compute, support effort or attributable use. Do not hide platform-team cost outside the AI business case.

### 5.2 Economic gate

Before production, each workflow records:

- one-time investment and annual run cost;
- low/base/high capacity and cash outcomes;
- adoption, quality and volume sensitivities;
- risk/disbenefit allowance;
- incremental and fully allocated unit cost;
- expected payback and break-even volume;
- contract/minimum-commitment and exit exposure; and
- kill thresholds and review date.

A sensible default for discretionary SME automation is a **base-case payback within approximately 24 months**, or a documented strategic, customer, quality or risk reason for accepting longer. The organisation may set a different hurdle, but should do so explicitly.

## 6. Conversion routes

### 6.1 H1 — Capacity redeployment

**Mechanism:** released hours are formally assigned to funded priorities, service improvement, backlog or resilience.

**Treatment:** valuable operating outcome, usually non-cash. It becomes a cash/revenue benefit only through a separately evidenced route.

**Evidence:** manager and receiving priority accept the capacity; work allocation or output changes; no simultaneous B3 claim unless the allocation is adjusted.

### 6.2 H2 — External-spend reduction

**Mechanism:** reduce outsourced work, agency scope, contractor volume or professional fees because an internal AI-supported process replaces it.

**Treatment:** B2 cash benefit.

**Evidence:** contract/PO/invoice reduction, net of internal capacity, tooling, quality-control and transition cost. Capacity used to in-house the work is allocated R3 and not added separately.

### 6.3 H3 — Hiring avoidance

**Mechanism:** a pre-existing funded role is deferred or cancelled because demonstrated capacity absorbs the expected workload.

**Treatment:** B3 cash benefit against the approved workforce plan.

**Evidence:** documented prior role plan, capacity/volume analysis, budget decision and service/quality guardrails. The same released capacity cannot also support another B3 card.

### 6.4 H4 — Attrition-based reshaping

**Mechanism:** after sustained task compression and role redesign, a natural vacancy is not backfilled or is replaced with a different/lower-cost capability.

**Treatment:** B3 cash benefit after vacancy and budget decision.

**Evidence:** durable workload evidence, role redesign, service-risk assessment, org/budget change and people-process compliance.

### 6.5 H5 — Software/licence retirement

**Mechanism:** retire/right-size point tools, overlapping assistants, unused premium seats or duplicate platforms.

**Treatment:** B4 cash benefit.

**Evidence:** renewal/order/invoice reduced; migration, replacement usage and lost functionality included.

### 6.6 H6 — Direct unit-cost reduction

**Mechanism:** reduce non-labour transaction cost, error/rework, manual processing fee or other variable cost.

**Treatment:** B5 cash benefit.

**Evidence:** component-level baseline and sustained unit-cost change. Labour components must reconcile to the capacity ledger and B3 cards.

## 7. Redeployment and people plan

### 7.1 Default capacity allocation policy

The council may set a default planning split, reviewed annually. An illustrative growth-oriented policy is:

| Destination | Illustrative share | Examples |
|---|---:|---|
| Growth/customer capacity | 40–55% | More account coverage, faster roadmap, new offers, service improvement |
| Backlog/quality/resilience | 20–30% | Documentation, controls, technical debt, leave coverage, training |
| AI/process capability | 10–20% | Workflow ownership, evaluation, data quality and automation improvement |
| Cashable workforce route | 0–20% | Hiring avoidance or attrition-based reshaping where durable and appropriate |

These are policy ranges, not benefits. The actual allocation follows strategy, demand and workforce obligations.

### 7.2 People principles

- no-surprise role and capacity conversations;
- measure workflows, not individual keystrokes;
- do not use pilot time studies as covert performance management;
- reskill and redesign before assuming role removal;
- internal mobility before external hiring where skills can transfer;
- consult employees/representatives where required;
- protect realistic operating slack and recovery capacity;
- disclose where AI monitoring or decision support affects staff; and
- record where capacity went so “efficiency” does not become an unspoken workload increase.

## 8. Worked example — 100-FTE organisation

This example demonstrates the accounting logic. It is **not** a benchmark or forecast for a real organisation.

### 8.1 Assumptions

- 100 FTE; 70 roles have material eligible knowledge/process work.
- Average fully loaded cost: **$140,000**.
- Eligible payroll: **$9.8 million**; total people cost: approximately **$14.0 million**.
- Current external service candidates: **$600,000/year**.
- Current SaaS spend: **$700,000/year**.
- A pre-existing two-year growth plan includes roles that may become B3 candidates.
- All values are AUD, rounded, before tax and financing effects.

### 8.2 Year 1 — investment and proof

| Line | Conservative | Base | Stretch |
|---|---:|---:|---:|
| Average gross time release across eligible roles | 3% | 5% | 7% |
| B1 gross capacity equivalent | $294k | $490k | $686k |
| Capacity realisation ratio after adoption/rework/support | 35% | 45% | 55% |
| **B1 net capacity created — reported separately** | **$103k** | **$221k** | **$377k** |
| B2 external spend reduction | $50k | $100k | $160k |
| B3 hiring/attrition benefit | $0 | $0 | $140k |
| B4 software consolidation | $15k | $30k | $50k |
| B5 direct unit-cost reduction | $10k | $25k | $50k |
| **Gross cash benefit — excludes B1** | **$75k** | **$155k** | **$400k** |
| B8 full AI TCO | ($300k) | ($360k) | ($450k) |
| **Net cash Year 1** | **($225k)** | **($205k)** | **($50k)** |

**Reading:** a negative Year-1 cash result is commercially plausible because setup, change, overlap and capability costs arrive before most cash routes. Capacity created is useful but is not added to the cash result.

### 8.3 Month-24 annual run-rate

| Line | Conservative | Base | Stretch |
|---|---:|---:|---:|
| Gross time release across eligible roles | 8% | 12% | 16% |
| B1 gross capacity equivalent | $784k | $1,176k | $1,568k |
| Capacity realisation ratio | 45% | 60% | 70% |
| **B1 net capacity created — reported separately** | **$353k** | **$706k** | **$1,098k** |
| B2 external spend reduction | $120k | $220k | $320k |
| B3 hiring/attrition benefit | $140k | $280k | $420k |
| B4 software consolidation | $35k | $60k | $90k |
| B5 direct unit-cost reduction | $40k | $100k | $180k |
| **Gross annual cash benefit — excludes B1** | **$335k** | **$660k** | **$1,010k** |
| B8 annual run-rate TCO | ($380k) | ($480k) | ($620k) |
| **Net annual cash run-rate at Month 24** | **($45k)** | **+$180k** | **+$390k** |
| Net cash as % of total people cost | (0.3%) | 1.3% | 2.8% |

### 8.4 Reconciliation and interpretation

- B1 is **not additive** to the cash lines. In the base case, some of the $706k capacity equivalent may support the $280k hiring-avoidance route, the $220k external-spend route and non-cash growth/backlog work. The capacity ledger assigns those hours once.
- The base case reaches a positive **annual run-rate** by Month 24; it does not imply cumulative cash payback has already occurred. Cumulative payback depends on when contracts, vacancies and licences actually change.
- The result is highly sensitive to external spend, planned hiring, workflow concentration, adoption, exception burden and internal delivery cost. A firm with little outsourced spend or growth hiring may create substantial capacity but little near-term cash.
- Do not scale the example linearly by headcount. Small firms have higher fixed-cost and fragmented-capacity effects; larger firms may have stronger scale but more integration, governance and change cost.

## 9. Reporting rhythm

### Monthly

- workflow unit costs and B8 actuals;
- adoption, quality, exception and disbenefit trends;
- B1 capacity measured and allocated;
- benefit-card movements and forecast variance;
- spend/seat/usage anomalies; and
- workflow kill/remediation candidates.

### Quarterly portfolio and harvest review

1. Reconcile scorecard, TCO and material D1 disbenefits.
2. Review each E2–E4 benefit card and overlap checks.
3. Approve or reject capacity allocations and cash conversion actions.
4. Finance co-signs E5 banked cash benefits.
5. Reforecast low/base/high cases using current evidence.
6. Promote, redesign, pause or retire workflows based on net value and risk.
7. Communicate to staff where meaningful capacity has been redeployed.

### Annual

- re-baseline sustained processes;
- reset unit-cost targets, capacity assumptions and cost allocations;
- reconcile cumulative cash investment/benefit and Month-24/run-rate view;
- review workforce and contract plans;
- test whether shared platform costs remain justified; and
- roll forward the three-year portfolio case.

## 10. Anti-patterns and rejected claims

1. Multiplying “minutes saved per email” by every employee and calling the total cash savings.
2. Adding B1 capacity value to B3 hiring avoidance or labour-based B5 savings.
3. Counting a role as avoided when it was never funded or planned before the AI case.
4. Counting a software list price instead of the actual renewal/invoice reduction.
5. Ignoring implementation, control, support, overlap, approval and exception cost.
6. Calling adoption, prompts submitted or agents created a benefit.
7. Claiming revenue uplift without a counterfactual or credible attribution.
8. Treating a short pilot result as a sustained production result.
9. Banking a benefit while quality, complaints, risk or workload has materially worsened.
10. Claiming 100% of measured time as available capacity.
11. Scaling a 100-FTE example linearly to 10 or 500 FTE.
12. Leaving shared platform/team cost unallocated while reporting workflow savings.
13. Counting the same external contract reduction against multiple workflows.
14. Treating avoided hypothetical risk as cash without an accepted valuation method.
15. Keeping a negative-value workflow because it is strategically fashionable.

## 11. Quarterly review agenda

| Item | Time guide |
|---|---:|
| Portfolio outcome, quality, incidents and B8 actuals | 10 min |
| Capacity ledger reconciliation and destination decisions | 10 min |
| Benefit-card review: Measured → Committed → Banked | 20 min |
| Workflow commercial gates: scale, redesign, pause or retire | 10 min |
| People, consultation, reskilling and workload guardrail check | 10 min |
| Contract, licence and reinvestment decisions | 10 min |

## 12. Reference note on productivity evidence

External studies are useful for forming hypotheses but show materially different effects by task and context. For example, field evidence has found meaningful gains in some customer-support settings, while controlled research with experienced open-source developers has found negative productivity effects under particular tools and tasks. The commercial conclusion is not that AI “works” or “does not work” universally; it is that local workflow design, user capability, process fit and measurement determine the result.

As calibration rather than planning inputs: the NBER customer-support field study found roughly a 14% average productivity gain, concentrated among less-experienced agents (around 34%), while the METR randomised study found experienced open-source developers were about 19% *slower* with AI assistance on their own repositories. Published task-level effects therefore span roughly −20% to +35% or more depending on task, user experience, tooling and process fit — a range wide enough that no generic percentage belongs in a business case.

- [NBER — Generative AI at Work](https://www.nber.org/papers/w31161)
- [METR — Early-2025 AI and experienced open-source developer productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)
