I keep seeing the same gap in management's use of AI. Reports become shorter. Emails become more diplomatic. Someone builds an impressive dashboard that cannot trace its numbers back to the company's actual records.

Meanwhile, the team still checks the same documents, chases the same approvals and copies information between the same systems.

For a COO, CFO or head of shared services, useful AI adoption means changing that work. Train people on their own cases, observe the process where it happens, connect approved company data and measure whether the operation improves. Personal convenience is a good beginning. It is a poor place to declare the implementation finished.

The productivity gain may stop at the manager's desk

Writing an email faster has value. So does understanding a difficult report without spending an afternoon on it. These are sensible uses of ChatGPT, and employees should not need an enterprise transformation programme to benefit from them.

The mistake is treating those gains as evidence that the business process has changed.

In McKinsey's August 2026 State of AI survey, 80% of respondents said AI improved their individual productivity. Only 37% reported a positive contribution to their organisation's EBIT. Those are different measures, not a conversion rate, and a survey cannot tell us what caused the gap. It does show why management should ask for evidence beyond usage statistics.

A faster report can still describe a slow operation. If the approval queue takes five days, saving ten minutes on the commentary does little for the customer waiting at the other end.

This is where I would narrow the conversation about enterprise AI. Focus on businesses with established systems, recurring transactions and expensive handoffs. Finance operations, procurement, service delivery and document-heavy industrial businesses have concrete work to examine. They also have consequences when the work is wrong.

Four places where a useful experiment stops too early

The following are illustrative operating scenarios, not claims about named Amalgama clients. Each starts with a reasonable use of AI. The question is what happens next.

Article data table
Starting pointWhat remains unresolvedA useful operating changeEvidence to collect
A manager summarises the monthly reportPeople still reconcile inconsistent exports and explain unexplained differencesBuild the report from governed data, calculate variances in code and let AI draft commentary linked to those figuresPreparation time, reconciliation errors, correction rate and reporting delay
A team uses AI to write supplier emailsMissing documents and unanswered requests still require manual chasingClassify incoming requests, check required documents, route exceptions and prepare replies from the case recordComplete submissions, repeat contacts, overdue cases and total handling time
An attractive dashboard is generated from a sample spreadsheetThe definitions, refresh process and underlying records are unclearAgree metric definitions, connect trusted sources and make each exception traceable to the relevant transactionData freshness, reconciliation with source systems and whether flagged issues get resolved
Someone pastes an invoice into a chatbotThe result still needs checking, retyping and approvalExtract proposed fields, validate totals, match supplier and purchase-order records, then route discrepancies for reviewCost per completed invoice, duplicate detection, rework and exception backlog

Notice how much of the useful version is ordinary systems work. Identity, data access, validation and routing do not become optional because a language model is involved.

For invoice processing automation, a model can help interpret inconsistent descriptions or document layouts. Arithmetic belongs in deterministic checks. Payment authority belongs in the company's approval controls. A fluent explanation should not be able to overrule either.

Give the dashboard a job before giving it a redesign

A management dashboard needs a decision attached to it.

If a chart says supplier performance is deteriorating, a manager should be able to see which orders are late, how lateness is defined, when the data was refreshed and who owns the next action. If none of that is available, the dashboard may be a useful prototype. It is not yet a management control.

Keep the calculation layer separate from the explanation layer. Calculate operational metrics from agreed definitions and controlled queries. Let AI explain a verified change, retrieve relevant records or suggest questions for investigation. Show when evidence is missing.

This matters because confidence can survive a wrong answer. In a 2023 experiment involving 758 BCG consultants, GPT-4 improved performance on creative product innovation tasks. On a deliberately difficult business problem-solving task, participants using it performed worse than the control group. BCG reported a 23% decline in performance for that task.

That is a result for a particular model and experiment, not a score for today's AI products. The management lesson is narrower. Success at drafting a convincing answer does not establish competence at making the underlying decision. Test the task you intend to delegate.

Start with a gemba walk through the work

Before asking which AI platform to buy, sit with the people who process the cases.

The Lean Enterprise Institute describes a gemba walk as direct observation and inquiry before action. In an office, the relevant place may be an ERP screen, a shared inbox or the spreadsheet someone maintains because the official system does not quite work.

An operations specialist demonstrates a delivery discrepancy while an analyst observes and takes notes.

Follow a case from arrival to completion across departments. Include ordinary cases and awkward exceptions. Ask the employee to show the work rather than describe the procedure from memory.

  • Where do you look for the information you need?
  • What do you copy, compare or re-enter?
  • Which cases wait, and what are they waiting for?
  • When do you ask an experienced colleague for help?
  • Which errors come back from the next team?
  • What happens when the system is unavailable?

Separate hands-on work from waiting time. Ten minutes of processing and three days in an approval queue are different problems. A better extraction model will not fix an absent approver.

Do this with the team, not to the team. Employees need to know that the exercise is about understanding the process, not catching them using the wrong spreadsheet. Otherwise, management sees the documented workflow and misses the one that actually keeps the business running.

The output should be a specific record of the process. What arrives, what information is needed, which decisions are made, where exceptions go and how completion is recorded. Only then can you judge which parts need AI.

Teach people to handle the difficult case

Generic prompt training can help employees get started. Operational AI training needs a different test of competence.

Can someone recognise that a document is incomplete? Can they find the source of an answer, correct an extracted field and escalate a case without losing its history? Do they understand which data the approved tool may process?

Build the training around representative work. Use authorised, suitably protected examples, including ambiguous requests and failures. Teach a finance reviewer to spot an incorrect supplier match. Teach a service agent to distinguish a policy exception from a routine answer. Teach managers to interpret process results without turning every saved minute into a cash-saving claim.

A practice loop uses a protected real case, checks the AI answer against its source, corrects or escalates the exception and records the lesson.

Make employees practise the fallback, too. A successful implementation still has days when a model, connector or source system is unavailable.

There is evidence that assistance within actual work can help people learn. The 2023 NBER working paper on generative AI in customer support studied 5,179 agents and reported a 14% average increase in issues resolved per hour. The gain was larger for less experienced and lower-skilled workers, with little effect for the most experienced workers. These figures refer to that working-paper version and deployment, not a universal productivity promise.

For management, the useful question is where experienced employees hold knowledge that newer colleagues repeatedly need. An assistant drawing on approved guidance can make that knowledge easier to use. It still needs a way to recognise when the guidance does not answer the case.

Prove the capacity before cutting the team

A convincing coding demo or a batch of generated layouts can make an IT or design team look suddenly oversized. Before removing roles, check who will maintain the code, investigate failures, test accessibility and handle the cases the demonstration avoided.

The Reuters investigation into Meta's AI workforce overhaul, published in August 2026, offers a caution. Facebook's parent envisaged smaller AI-assisted teams and changes to engineering and design roles. It cut 10% of employees in May but cancelled planning for a second wave in November. Reuters reported internal doubts about the expected productivity gains, while noting that it could not establish exactly why Zuckerberg changed course. Meta described the programme as a mix of cost reduction, team redesign and redeployment.

Workforce optimisation can produce lasting gains. It is one of my core areas of experience. I have led headcount reductions involving thousands of people across large organisations without deterioration in long-term operating efficiency. The test is whether the operation continues to perform after the savings appear in the budget.

Before deciding who the redesigned operation needs, assess AI adaptability through actual work. Give people comparable training, approved tools and representative tasks. Look at whether they can:

  • Produce an acceptable result with less total effort, including review and corrections.
  • Catch unsupported answers, faulty code or design decisions that fail the brief.
  • Learn the revised process and handle exceptions without constant supervision.
  • Protect company data, document their work and recognise when human help is necessary.

An employee in a lower salary band may prove more effective with AI than someone whose existing title suggests greater value. At the current cost, that person may contribute far more than the old role allowed. Treat this as something to test, not an assumption about cheaper employees. Salary, seniority and enthusiasm for AI are all poor substitutes for observed results.

Use those observations to plan training, redeployment and, where justified, phased reductions. Managers should make the decisions, with service quality, workload and operational resilience checked over time. Cutting the people who know how to repair the process can make the next efficiency programme rather expensive.

Pick one high-volume operation worth changing

Start with a recurring operation whose result someone can check. Supplier onboarding, invoice exceptions, service-request triage and document completeness checks are plausible candidates. A process owner still needs to test their suitability against local data and risks.

Choose a bounded slice. For supplier onboarding, that might mean identifying missing documents and preparing a request for them. It does not have to include approving the supplier, changing bank details or deciding contractual terms.

Estimate the opportunity using observed volumes and handling time, then include the work the new system creates. Review, corrections and ongoing maintenance count.

Consider a hypothetical operation with 10,000 cases a month. If half the cases qualify for assistance and each saves four minutes before review, the gross opportunity is about 333 hours. If those assisted cases still need one minute of new review each, about 250 hours remain before accounting for rework, maintenance and other overheads.

Those hours are capacity, not automatically cash. The business must decide whether it will reduce overtime, clear a backlog, absorb growth or move people to other work. If nothing changes in staffing, expenditure or output, a spreadsheet should not quietly promote the capacity gain into profit.

Use an AI automation ROI calculator to challenge the assumptions, then replace estimates with pilot observations. Model rollout time and gradual adoption. A tool installed on Monday rarely delivers a full month's benefit by Tuesday.

Make the pilot finish a piece of work

A useful pilot begins with a real trigger and ends in the system where the business records completion.

For the supplier-document example, the workflow could look like this.

  1. An incoming request creates or updates a case with a stable identifier.
  2. The system retrieves only the records and document requirements that the user is authorised to access.
  3. AI proposes document classifications and extracted fields, linked to their sources.
  4. Explicit rules check completeness and contradictions. Unclear cases go to a named reviewer.
  5. An authorised person approves any consequential action. The approved result is written back to the case, with a record of what changed.

Supplier requests pass through permitted records, AI proposals and rule checks. Exceptions go to a named reviewer before authorised approval and an audited case update.

Test against historical cases with known outcomes before allowing live actions. Include duplicate submissions, missing attachments, inconsistent names and documents containing misleading instructions. Start in a mode where staff can compare suggestions with their normal decisions. Expand only when the evidence supports it.

Track the whole result. Faster classification is irrelevant if corrections increase downstream. Measure completion time, review effort and error severity alongside model accuracy. Compare similar case types and keep a manual fallback available.

Our guide to moving enterprise AI from pilot to scale covers the production controls in more detail. Management's job at this stage is to agree who can stop the rollout and what evidence would justify doing so.

This is where operating experience earns its place

My background spans twenty years across enterprise technology, consulting and business ownership, including work across banking, energy and global IT delivery. That experience shapes the questions I ask before an AI build starts.

Who owns the process when it crosses departments? Which figures does finance accept? What must remain available during a system change? Where does an exception go when the person who normally handles it is away?

These are familiar implementation questions. AI adds new failure modes, but it does not remove the organisational ones.

For Amalgama, the useful engagement is with a business that already has work worth improving. A finance or operations team with recurring volume, corporate data, multiple handoffs and an accountable owner. The work combines management training, process observation and technical delivery around an operation that can be measured.

You can read more about the experience behind that approach. The practical starting point is smaller than a company-wide AI strategy. Bring one expensive queue, its process owner and a representative set of cases to a workflow assessment.

We can then work out what should be simplified, what can be automated and where AI genuinely helps. If the result is another dashboard, at least it should explain why the queue is finally getting shorter.