From Pilot to Performance: Adopting AI as one of the team
- November 8, 2025
- Posted by: Justin Prince
- Categories: Applied Technology, Human Resources and Industrial Relations
A simple, human‑centric roadmap to make AI actually work
Meet the AI Geek (close cousin on the CPU side to the IR Geek)
The AI Geek shares the IR Geek’s allergy to theatre and love of results – same dry humour, less patience for window dressing. They’re the person who quietly says, “Don’t expect a 10× return from a 10‑second setup,” draws a neat box on the whiteboard, and asks, “What job are we actually hiring AI to do?” From there it’s about turning demos into outcomes: clear roles, real context, and a weekly coaching rhythm.
Measure twice, cut once.
A 30‑minute brief beats a 30‑second one – like giving an artist time to create something worth framing.
The Museum of Interesting Experiments
Why AI pilots succeed and then nothing happens
Somewhere in an office right now there is a slide deck from an AI pilot that finished eighteen months ago. The results were good. The technology worked. Three or four people still talk about it fondly. Nothing came of it.
That deck is not evidence of failure. It is evidence of something more awkward, which is a pilot that answered the wrong question and was congratulated for it.
The question the pilot answered was whether the model could do the task. The answer was yes, because the answer is almost always yes, and it has been yes for a while now. The question nobody asked was what would have to change about the way the work is done before that capability produced anything a customer or a finance director could see.
The six weeks everyone recognises
The pattern is consistent enough to set your watch by.
Monday of week one, the leadership team launches the discovery phase. There is a town hall, a budget, a partner logo and a handful of demonstrations that make the room go quiet in a good way. Enthusiasm is genuine. Nobody is being cynical.
By week six the pilot has wrapped and the report is honest. Some steps got faster. Some edge cases surfaced that nobody had thought about. The team now has a much clearer picture of what good looks like, which is exactly what a pilot is for. The next steps are obvious and they are all unglamorous: write down what the AI is actually responsible for, assemble the reference material it needs, decide which outputs a human has to sign off before they leave the building, and wire the whole thing into the place where people already do the work.
Then comes the part nobody schedules. Momentum quietly evaporates. A budget line gets trimmed. The project sponsor picks up something urgent. Nobody makes a decision to stop, because stopping is not what happens. What happens is that pausing starts to feel like the sensible, adult, risk-managed option, and the pilot slides into the museum of interesting experiments, where it sits next to the intranet redesign and the last customer portal.
The diagnosis is not complicated. The organisation did the exciting part and skipped the operational part. It installed a new capability into an unchanged process and then wondered why the process behaved the way it always has.
Purpose, or write the role before you turn it on
Every organisation in the country knows how to bring on a new team member. There is a position description, a reporting line, a probation period, a set of things the person is allowed to decide alone and a set of things they are not. None of this is exotic. It is the ordinary machinery of getting work done through someone other than yourself.
AI gets none of it. It gets a login and a hopeful email.
A one page charter fixes most of this, and it does not need to be elegant. It needs to say what outcome changes for a customer or a cost line, what the tool is explicitly not for, and, most importantly, who decides what. Split that last one three ways. There is work the system can complete on its own, work where it acts and a human reviews a sample, and work where a human approves before anything moves. Write examples against each tier rather than definitions, because definitions get argued about and examples get used.
The useful discipline here is that prompts are not incantations. They are specifications. A vague specification produces vague output, and then everyone stands around the demonstration blaming the model.
Structure, or the workflow is the thing you are actually changing
This is where most programs quietly go wrong, because it is the part that requires someone to give something up.
Map the handovers honestly. Where does the AI produce something, where does a human take it, and where does it go back. Then look for the steps that only exist because the old process needed them. There will be several. There is usually at least one review stage that was invented to catch a category of error the new workflow no longer produces, and it will survive the redesign anyway because removing it feels like removing a control.
Removing it is the point. If the flow ends up with the same number of steps it started with, nothing has been redesigned. Something has been decorated.
Three things need attention at once, and in this order. The people first, because the humans should end up with the better job rather than the leftover one, which means judgment, exceptions and relationships instead of assembly. The process second, with gates set by risk rather than by seniority, and some agreement on what speed and accuracy actually need to be. The technology last, and mostly this means connecting the thing to the data that is authoritative rather than the data that was easiest to reach.
Context, or the part that nobody wants to own
The performance difference between an AI implementation that works and one that produces confident nonsense is rarely the model. It is the reference material.
Ten to twenty worked examples, including the bad ones and why they are bad. A glossary that settles the terms your organisation uses differently from everyone else. Links to the documents that are actually authoritative, as distinct from the documents that are merely available on the shared drive. The known edge cases with the answers you would accept.
Version it the way you version code. If a change alters performance, it gets a number and a note explaining what changed and why. This sounds fussy. It is fussy. It is also the difference between an implementation you can debug and one you can only apologise for.
Feedback, or confidence is not competence
An AI system will tell you it is right in exactly the same tone whether it is right or not. So will a manager, which is why IR practitioners spend so much of their working life on the gap between the two, and it is the same gap here.
The fix is a weekly rhythm and it takes under an hour. Pull a random sample of outputs. Argue briefly about which ones are good. Update the charter, the examples and the tests to reflect the argument. Publish a short scoreboard showing accuracy against your reference examples, the rework rate with its top three causes, and cycle time against last week.
Skip this and quality decays, reliably and without anyone noticing until a customer notices first. Nothing degrades faster than an automated process that nobody has looked at since the launch.
The first month, realistically
Week one is for choosing and writing. Pick two tasks that are narrow and high in volume, because that is where variance and repetition meet and where a result will actually be visible. Baseline what they cost today in time and in errors, since a program without a baseline can only ever be defended with anecdotes. Draft the charter.
Week two is the reference material and the data connections, plus a small evaluation set you can rerun every time something changes.
Week three is a live pilot with a small group, daily tracking of quality and rework, the first weekly review, and the deletion of at least one step that the data says is no longer earning its place.
Week four is publishing. Share the charter, the controls and the numbers. Tie it to a recognised standard, lightly, because governance should make the correct path the convenient one rather than serve as its own separate ceremony. Then choose the next two workflows.
The point of the month is not to scale a pilot. It is to establish a way of working that the second, fifth and twentieth use case can inherit.
The five decisions that cannot be delegated
There are only five, and consultants cannot make any of them for you.
Which two workflows go first, and on what reasoning. Who owns the reference material, by name, because material owned by everyone is maintained by nobody. What the risk posture is for each tier, with examples rather than principles. Which single metric you will celebrate publicly when it moves. And what the organisation is going to stop doing, because a program that only ever adds things is a program that is not really changing anything, and everybody in the building can tell.
The closing observation
Measure twice, pour once. Concrete is unforgiving about preparation and completely indifferent to how enthusiastic you were on the day.
The organisations getting real value out of this are not the ones with the best models. Everyone has access to broadly the same models, and the gap between the good ones narrows every quarter. They are the ones that treated the technology as a colleague who needed a role, a set of reference materials, some rules about what they could decide alone, and a conversation once a week about how it was going.
None of that is technically difficult. All of it is boring. Boring wins, which is not a new observation in this business, merely one that has found a new place to be true.
Somewhere in Australia a leadership team is about to approve a second pilot to answer a question the first pilot already answered. It will look impressive. It will finish in six weeks. It will go straight into the museum, where the lighting is good and nothing ever changes.
Appendix: evidence and further reading
Productivity and workflow redesign
Generative AI at Work, Quarterly Journal of Economics 2025, open access. https://academic.oup.com/qje/article-abstract/140/2/889/7990658. A large field study finding roughly a 14 to 15 per cent lift in support agent productivity, with the largest gains going to the least experienced staff.
Generative AI at Work, NBER working paper. https://www.nber.org/papers/w31161. The working paper version, with full methods and the heterogeneity results.
The Impact of AI on Developer Productivity, controlled experiment, arXiv. https://arxiv.org/abs/2302.06590. Developers completed a coding task around 55.8 per cent faster with Copilot.
GitHub, Measuring the impact of GitHub Copilot. https://resources.github.com/learn/pathways/copilot/essentials/measuring-the-impact-of-github-copilot/. A readable summary of multiple studies and the metrics used.
Dell’Acqua and others, Navigating the Jagged Technological Frontier, Harvard Business School working paper. https://www.hbs.edu/ris/Publication%20Files/24-013_d9b45b68-9e74-42d6-a1c6-c72fb70c7282.pdf. A consultant field experiment showing strong gains inside the technology’s strength zone and measurable risk at its edges.
Adoption, culture and leadership
MIT Sloan Management Review and BCG, Learning to Manage Uncertainty, With AI. https://sloanreview.mit.edu/projects/learning-to-manage-uncertainty-with-ai/. Organisations that pair organisational learning with AI learning perform better under uncertainty. Full report at https://web-assets.bcg.com/c1/a7/af0e57dc4b47a31eb7409d981d3e/mitsmr-bcg-ai-report-november-2024.pdf.
Harvard Business Review, Your Organization Isn’t Designed to Work with GenAI. https://hbr.org/2024/02/your-organization-isnt-designed-to-work-with-genai. Treat the technology as an assistive agent and redesign the work around it.
Harvard Business Review, Stop Tinkering with AI. https://hbr.org/2023/01/stop-tinkering-with-ai. Pilots create no value until they scale into redesigned work.
Harvard Business Review, The Gen AI Playbook for Organizations. https://hbr.org/2025/11/the-gen-ai-playbook-for-organizations. Current guidance on what leaders should be asking and doing.
Governance and standards
NIST AI Risk Management Framework. https://www.nist.gov/itl/ai-risk-management-framework. Four functions: govern, map, measure, manage. Playbook at https://www.nist.gov/itl/ai-risk-management-framework/nist-ai-rmf-playbook.
ISO/IEC 42001:2023, AI management systems. https://www.iso.org/standard/42001. The management system standard for AI, structured much as ISO 27001 is for information security. Australian adoption note at https://www.standards.org.au/blog/spotlight-on-as-iso-iec-42001-2023.
European Parliament, EU AI Act implementation timeline. https://www.europarl.europa.eu/RegData/etudes/ATAG/2025/772906/EPRS_ATA%282025%29772906_EN.pdf. Useful for anyone planning review cycles against the phased obligations.
Macro context
Stanford HAI, AI Index 2025. https://hai.stanford.edu/assets/files/hai_ai_index_report_2025.pdf. A neutral synthesis covering adoption, economics and policy, suitable for board reading.
LevelUp
https://lvlup.au/, https://lvlup.au/how-we-work/, https://lvlup.au/how-we-work/applied-technology/