Read it once end to end, then use the flash cards. Every card links back to its chapter.
≈ 111.4 min of reading48 cards · 34 interview questions · 130 chapters
0 of 82 reviewed
The map: one idea per chapter
M00 · Course Introduction
Welcome and How This Program Works — Level 1 builds AI decision-makers, not engineers. Ten questions in order, one route for every role, and a method that turns reading into action.
Why Executives Need AI Literacy — Specialists can say what AI can do; only leaders decide what the organization does with it. Literacy is judgment, questions and accountability.
From AI Awareness to AI Leadership — Awareness changes what you know; leadership changes what you do. Four behaviors mark the shift: sponsor, ask for evidence, redesign work, own decisions.
Your AI Leadership Starting Point — Know where you start. Rate what you can explain and show, not what you feel, then choose one gap and one real problem to carry.
M01 · The AI Revolution
AI Is Changing Everything — Almost everyone has AI; few have changed the business. Advantage comes from redesigning the work around it, as it did with electricity.
The Four Industrial Revolutions — AI is new; the pattern is not. General-purpose technologies pay off only after organizations build the complements - processes, skills and new ways of working.
Why AI, Why Now? — AI is old research; readiness is new. Web-scale data, specialized chips, foundation models and easy access converged, for everyone at once.
The AI Adoption Curve — AI adoption is a curve of depth, not only spread. Your position depends on where the AI lives, not on how many people have access.
AI Leaders vs AI Followers — The best models sit close together, so behavior decides who leads: choose a real problem, own it, change the work and prove the result.
The AI Competitive Advantage — A rival can buy your AI model in weeks. Advantage lives in the layers around it, which are slow to copy and slower together.
The Speed of AI Adoption — Generative AI spreads at employee speed because a first try needs nothing new. Advantage goes to organizations whose own learning keeps up.
AI Is Already Inside Your Organization — AI is already inside the organization - embedded, official and informal. See it first, give it owners, then buy.
Generative AI Changes the Game — Generative AI did not invent AI. It made language the interface, so almost anyone can use it - widening opportunity and responsibility.
From Copilots to AI Agents — AI is moving from answers to action. Copilots draft while people act; agents act themselves, so permissions must match the action.
AI and the Future of Work — The unit of change is the task, not the job. Leaders choose automate or augment task by task, redesign roles and tell people the truth.
The AI Talent Gap — The AI talent gap is a missing capability across four layers, not a shortage of experts. You cannot hire your way out of it.
The AI Transformation Challenge — A working model is the easy part. Value needs six components holding together, and the weakest one sets the ceiling.
Where Should We Start With AI? — Where you start decides what your organization learns about AI. Start where the work matters, feedback is fast and mistakes can be undone.
Measuring AI Business Value — Activity is not value. AI is worth what the business measurably changes because of it, and only measurement shows which.
The Economics of AI — The model price is one line of the bill. Judge what an outcome costs, and give one owner the whole bill.
M02 · Understanding AI & Generative AI
What Exactly Is Artificial Intelligence? — AI is a label for a family of capabilities. Judge any system by what it does, how reliably, and what happens when it is wrong.
AI vs Automation — Automation follows rules people wrote. AI follows patterns learned from data. They fail differently, so design the simplest combination, with people owning exceptions.
AI vs Machine Learning — AI is the field; machine learning is one way to build it. Ask which parts learn, how they learn, and where their answer key comes from.
The Major Types of AI — AI labels answer four questions: what it does, how it is built, how general it is and how much it does alone.
How Machine Learning Learns — A model learns patterns by reducing its error on known examples. They pay only if they hold on unseen cases and keep holding.
Deep Learning — The Executive Mental Model — Deep learning is machine learning with many-layered neural networks that learn their own features from raw data. Depth is capacity, not a guarantee.
Training vs Inference — Training creates capability; inference delivers it on every request, without changing the model, and carries much of the running cost.
Data, Models and Compute — AI capability is data, model and compute working like factors in a product. The weakest corner sets the ceiling.
Foundation Models — A foundation model is trained broadly once, then adapted to many tasks. The model is the bottom layer; value and accountability sit above it.
Large Language Models — An LLM predicts the next token from context. It supplies the language; your systems of record supply the facts.
Tokens, Context and Embeddings — Tokens are what a model reads and bills; context is what it sees now; embeddings find the right content by meaning.
Multimodal AI — Multimodal AI works across text, images, audio and video. Add a modality when it holds information that changes the decision.
How Generative AI Works — Generative AI composes new output from learned patterns, drawing on chance at every step. It gives the likely answer, not the checked one.
From AI Assistants to AI Agents — An assistant answers a request. An agent pursues a goal through a loop of steps, tools and checks, so judge it on the whole job.
M03 · Business Value of AI
What Does AI Value Actually Mean? — AI value is the measurable change in business results that AI causes, compared with what would have happened without it, net of cost.
Where AI Creates Revenue — AI creates revenue only when it changes behavior. Count the incremental part, measured against what would have happened anyway.
Where AI Reduces Cost — AI reduces cost only when a named budget line falls, net of what AI costs, with quality held. Hours saved are capacity, not cash.
AI and Customer Experience — Better customer experience becomes business value only when customer behavior changes. Aim AI at the effort customers waste, wherever in the journey it costs most.
AI and Workforce Productivity — A faster task creates capacity, not productivity. It pays only when the whole job improves and management decides what the time is for.
From AI Capability to Business Outcome — A capability becomes valuable only when it causes a measurable change in a business outcome. Every arrow between the two is an assumption.
Productivity vs Realized Capacity — Time saved is potential value. It becomes realized capacity only when management decides where it goes and measures the result.
Leading vs Lagging AI Metrics — Leading metrics steer; lagging metrics judge. Link them in one chain, check that leaders really lead, and give guardrails early warnings.
Building the AI Business Case — A business case is an argument that an option beats business as usual on expected value, with its range, its breaking points and its evidence shown.
AI ROI and Value Realization — ROI is calculated twice, as a forecast and as a result. Benefit leaks while cost stays, so realization must be owned and judged forward.
Total Business Impact — Total impact maps every effect across six dimensions, then counts each economic benefit once, where the money lands and against the plan.
AI Value Scorecard — An AI value scorecard is one page that ends in a decision - realized against expected value, confidence on every figure, rules agreed in advance.
Module 03 Synthesis — Choosing Value Over Hype — Hype skips links between capability and cash. Walk every claim through definition, scope, adoption, capture and cost, then let evidence choose the verb.
M04 · AI Strategy
What Is an Enterprise AI Strategy? — An AI strategy is a few hard choices - a diagnosis, a guiding policy, refusals and shared capability - that trace back to the business strategy.
AI Strategy vs AI Adoption — Strategy is choices; adoption is behavior. Usage only counts when it changes an important workflow that serves a strategic priority.
Start With Business Strategy — Start AI strategy with how the business wins and what constrains it, then ask whether AI materially changes that constraint.
Defining AI Ambition — AI ambition is the choice of how much AI should change the business. The right level fits strategy, readiness, budget and risk appetite.
Finding Strategic AI Opportunities — Walk down from the strategy to where value leaks, size the pool roughly, then ask whether AI beats the best alternative.
Build vs Buy vs Partner — Buying is the default. Name the few capabilities to own, decide layer by layer, and keep a way out.
Centralized vs Federated AI — Ask where each AI capability belongs, not who owns AI. Centralize for scale and reuse; federate for context and ownership.
AI Platform Strategy — An AI platform exists to create leverage. Build the smallest shared platform that repeated demand justifies, and judge it by the effort it removes.
Data Strategy for AI — Having data is not enough. Start from the AI outcome you chose, find the data gap that blocks it, and fix that first.
AI Talent and Capability Strategy — Talent can be hired; capability has to be built. Blueprint the capabilities, set role targets, measure the gap, then decide the sourcing.
AI Operating Model — Strategy sets direction. The operating model - owners, rights, funding and routines - makes execution repeatable for every use case.
AI Decision Rights and Accountability — Deploying AI moves decision rights. Choose where they land - many responsible, one accountable owner, authority scaled to stakes, the right to stop assigned.
AI and Competitive Advantage — When rivals can buy the same AI, customers keep the gains. Advantage needs a gap in cost or value and a guard that keeps it open.
The Enterprise AI Use-Case Landscape — Enterprise AI is five patterns repeated in every function. Read it by workflow and outcome, not by tool or department.
AI Use-Case Discovery and Design — AI projects fail on the problem, not the model. Map the work, give each task to a rule, AI or a person, then screen it.
AI in Finance — Finance AI should draft, forecast and detect at speed, while the authority to approve and pay is designed deliberately, never granted by default.
AI in Human Resources — Use AI freely for HR service work; when it shapes someone's job, pay or opportunity, a named person decides and outcomes are tested.
AI in Sales and Marketing — AI gives sellers time back and makes messages almost free. The value is better commercial decisions, approved claims and fair use of signals.
AI in Customer Service — Answering is not resolving. Service AI succeeds when the problem is fixed with less effort, within limits, and it knows when to hand over.
AI in Operations and Supply Chain — Operational AI ends in a physical move. A signal creates value only when someone may act on it in time.
AI in Software Engineering — AI makes writing code much faster. Value arrives only when verification, review and release keep pace, so measure delivery and fund the checks.
AI in Product Development — Ideas and prototypes are now nearly free. The scarce resource is evidence of what customers do, and people still choose what to build.
AI in Cybersecurity — Attackers use AI to go faster. Use it to decide faster, within written limits, and judge it by the risk it removes.
AI in Knowledge Management — AI makes knowledge cheaper to capture, combine and deliver. It cannot decide which knowledge is right - that is leadership work.
AI in Document and Data Processing — Document AI pays when trapped information becomes checked data a process uses, routed by error cost and traceable to its page.
AI Agents and Intelligent Workflows — Score an agent on three outcomes - verified, handed off, silently wrong. Price the silent errors, write its mandate, design its handoffs.
M06 · AI Risks
The AI Risk Landscape — AI risk is the business impact when an AI-enabled system is wrong. Accuracy belongs to the model; risk belongs to the use.
Accuracy, Hallucination and Reliability — AI will be fluent and sometimes wrong. Reliability is designed around the model - measured on real work, grounded, checked and able to decline.
Bias and Fairness — Accurate is not the same as fair. Check what the system predicts, test results by group, choose the fairness objective and keep monitoring.
Privacy and Confidential Data — AI opens new paths for sensitive information. Design privacy into the whole data path instead of assuming it from an approved tool.
Security and AI Attacks — Language is now an attack surface. We cannot make AI impossible to fool, so we limit what a fooled system can reach and do.
Intellectual Property and Copyright — AI-generated does not mean we own it - or that anyone can. Rights depend on inputs, provider terms, law and human contribution.
Model and Third-Party Risk — AI dependency is business dependency. Critical AI providers need an owner, proportionate controls and a tested way out.
Operational and Workforce Risk — AI can make an organization faster and more fragile at once. Keep the human capability to check the AI and to run without it.
Agentic AI and Autonomous Actions — When AI can act, risk moves from what it says to what it does. Scope its authority and grant more only on evidence.
Human Oversight and AI Incidents — A human in the loop is not control. Oversight needs authority, information, time and tools, and incidents need a rehearsed response.
What Is AI Governance? — AI governance is the system that decides who may make which AI decisions, under what rules and controls, with what accountability.
Why Responsible AI Matters — Responsible AI earns informed trust, and trust drives adoption. Valuable AI that people do not trust, or trust blindly, will not scale.
AI Governance Principles — Seven stable principles guide AI decisions whatever the model. Control depth follows the risk, and every principle needs an owner and a control.
AI Policy vs AI Governance — Policy states what must be true about AI. Governance is the system that makes it true and keeps it true.
AI Roles, Ownership and Accountability — Accountability for AI is designed or it defaults. Each line gets a distinct job; each material system gets one named owner who can stop it.
The AI Governance Operating Model — Centralize what must be consistent, delegate what needs business context, and let risk decide how far each decision travels.
AI Inventory and Risk Classification — The inventory shows what AI exists. Classifying each use, legal tier first, decides how closely it is governed.
AI Lifecycle Governance — An approval is a snapshot of a moving system. Govern the AI you run, from discovery to retirement, with one named owner.
AI Evaluation and Approval Gates — Evaluation is evidence about the conditions tested. Approval is an owned decision to proceed, under conditions, with triggers that reopen it.
AI Policies, Standards, Monitoring and Audit — A policy says what must be true. Standards make it testable, controls make it happen, monitoring shows it, and independent audit proves it.
The Economics of AI — AI is metered. Judge it by what each unit of value costs, across all cost layers, and whether that holds at ten times the use.
Understanding AI Total Cost of Ownership — Total cost of ownership is build plus run plus operate plus change over the system's whole life. The vendor quote is one line.
Data, Integration and Operational Costs — The model is priced per request. Data, integration and operations are priced by sources, connections, change and exceptions, and usually cost more.
Experimentation vs Production Economics — An experiment buys information, priced by the decision it informs. Production buys reliable delivery, priced by full cost of ownership.
AI Unit Economics and Economics at Scale — Scale multiplies what each unit leaves behind. Know contribution per unit, break-even volume and the assumption that moves both.
Cost Optimization and AI FinOps — AI FinOps is a loop that makes AI spend visible, owned and improvable. Aim for value per unit of spend, not the smallest bill.
Assessing Enterprise AI Readiness — Do not ask whether you are AI-ready. Ask whether you are ready for a specific ambition, what it needs, and which gap holds everything back.
The AI Maturity Model — Maturity is not pilots or spend. It is how reliably you run and repeat governed, measured AI value, at the level your strategy requires.
Prioritizing the AI Portfolio — Prioritizing AI is capital allocation: choose the combination your people can finish, stage the money behind evidence, and stop what fails its test.
From AI Pilot to Production to Scale — Four proofs: a pilot shows it is feasible, validation that it is valuable, production that it is operable, scale that it is repeatable.
Building the 90-Day AI Plan — A 90-day AI plan is a contract: few owned outcomes, long poles started in week one, a not-yet list, and a production decision on Day 90.
Building the Multi-Year AI Roadmap — A multi-year AI roadmap sequences decisions under falling certainty: a quarterly contract, a one-year commitment, a direction beyond, held up by named assumptions.
Where AI Is Going Next — AI already reasons, perceives and acts. Watch four dials - reliability, cost, autonomy, physical reach - and act when one crosses your written threshold.
The Evolution of AI Models — Top AI capability is crowded and changes hands in months. Match each workload to a model tier, test releases on your own cases, keep models swappable.
Multimodal and Agentic AI — Seeing and doing are becoming one AI system. The leadership choices are the door it uses, what it may watch, and where a person signs.
AI + Robotics — A wrong movement is not a wrong answer. Start from the physical workflow, keep safety outside the model, and judge cost per successful task.
AI-Native Products and Business Models — Sell the hole, not the drill. Price the unit that tracks your cost and the customer's result; defend with what rivals cannot rent.
Module 10 Synthesis — The Future AI Leader — Prepare the organization, not the forecast. Watch thresholds, test on your own work, commit at the speed of undo, and review on a date.
Level 1 builds AI decision-makers, not engineers. Ten questions in order, one route for every role, and a method that turns reading into action.
Four levels share one vocabulary. Level 1 has 130 chapters in eleven modules; each module after the introduction answers one executive question, in a deliberate order.
Every chapter is written once and published for different moments. Routes: the full sequence, a role track or the cheat sheet.
Use each chapter in five steps - understand, question, apply, discuss, act. Testing yourself beats rereading; a card a week later is spaced practice.
Say it: “The goal is not to finish the program. It is to decide better because of it.”
Red flag: Treating the program as reading to finish, or as a tool tutorial, instead of practice on real decisions.
Specialists can say what AI can do; only leaders decide what the organization does with it. Literacy is judgment, questions and accountability.
AI decisions about opportunity, investment, risk, people and strategy reach the executive table. Literacy means judgment, questions and accountability, not coding.
Low literacy fails three ways: abdicate (leave it to IT), overreach (fund everything) or over-trust (believe fluent output).
Delegate how AI is built and run; keep whether, why and how much risk. Boards need enough fluency to ask for AI risk reporting.
Say it: “You do not need to build AI. You need to be able to own the decision.”
Red flag: Saying AI is the technology team's job, or approving an AI proposal because the demo was impressive.
Awareness changes what you know; leadership changes what you do. Four behaviors mark the shift: sponsor, ask for evidence, redesign work, own decisions.
Personal AI fluency is awareness and literacy, not leadership. Sweden's prime minister used AI often and still faced the question of who decides.
Four behaviors: sponsor with money and time; ask for evidence against a baseline; redesign the work, not just the task; give every AI-touched decision a named owner.
Alcoa already knew safety mattered. Its lost-workday rate fell from 1.86 per 100 employees to 0.2 (0.5 in another account) once its leader changed what he did.
Say it: “Awareness changes what you know. Leadership changes what you do.”
Red flag: Pointing to your own daily AI use, or to the number of tools deployed, as proof that you are leading on AI.
Know where you start. Rate what you can explain and show, not what you feel, then choose one gap and one real problem to carry.
Confidence is a poor guide: people's ratings of their own understanding fall once they try to explain, and self-ratings track performance only modestly (r = 0.29).
Rate seven specific questions - understanding, value, strategy, risk, governance, economics, execution - at three levels: heard of it, can explain it, have used it.
Start where your role depends on a question you cannot yet answer with evidence; choose a track by gap, and carry one real business problem.
Say it: “A low score is not a problem. An unknown starting point is.”
Red flag: Rating yourself by how confident you feel, or starting with your strongest area because progress there feels quick.
AI is new; the pattern is not. General-purpose technologies pay off only after organizations build the complements - processes, skills and new ways of working.
Schwab's Fourth Industrial Revolution is a fusion of physical, digital and biological technologies, with AI as one driver. Perez counts five revolutions since 1771.
Each revolution moves from installation (a speculative frenzy) through a turning point to deployment, where the broad gains arrive.
The productivity paradox recurs - electricity, computers, now AI. Gains go to firms that pair the technology with new work practices.
Say it: “Technology changes capability. Leadership changes outcomes.”
Red flag: Treating market excitement or spending as evidence of value, or assuming the gains will follow adoption automatically.
AI is old research; readiness is new. Web-scale data, specialized chips, foundation models and easy access converged, for everyone at once.
Four distinct forces converged: web-scale public data, specialized chips enabling training at scale, general foundation models, and access through the cloud and plain language.
Building at the frontier keeps getting costlier while using models gets far cheaper, so most organizations rent the capability instead of building it.
Access is shared by every competitor, so advantage moves from obtaining AI to applying it to your own data, processes and people.
Say it: “AI did not arrive overnight. The world around it became ready, for everyone at once.”
Red flag: Saying AI arrived with one product, that we must build our own model, or that we should wait for the next model before learning.
The best models sit close together, so behavior decides who leads: choose a real problem, own it, change the work and prove the result.
Access no longer separates competitors: in March 2026 the top closed model led the best open model by only 3.3 percent.
Eight behaviors in four moves: choose a problem and few priorities; own it with a sponsor and built-in controls; change the work and equip people; prove, then scale or stop.
Leader and follower are patterns, not identities: 40 percent of management-practice variation lies within the same firm. Diagnose function by function.
Say it: “Access to AI is spreading. Behavior still decides who leads.”
Red flag: Claiming leadership because the company has licenses, a platform or many pilots - access and activity instead of behavior and outcomes.
AI is already inside the organization - embedded, official and informal. See it first, give it owners, then buy.
Much AI at work is quiet. It ranks, scores, forecasts and routes inside software you bought; assistants are only the visible part.
AI arrives in three buckets - embedded, official and informal - and increasingly by software update. Leadership usually sees only the official bucket, and first counts run low.
Bought is not owned. If a system shapes customers, money or people, someone inside answers for it, and the law addresses the user too.
Say it: “Our AI transformation started before we named it. First we see what is running and who owns it; then we choose.”
Red flag: Saying the AI journey starts when we pick a platform, as if nothing were already running or the vendor owned the outcome.
A working model is the easy part. Value needs six components holding together, and the weakest one sets the ceiling.
BCG's AI leaders spend about 10% on algorithms, 20% on technology and data and 70% on people and processes. The model is the smallest share.
Six components must hold together: technology, data, people, process, governance and leadership. They multiply, so one missing component caps the value.
Business leaders own outcomes; AI teams enable them. Governance designed early is what lets AI leave the lab.
Say it: “The model is rarely what fails. Value appears only when all six components hold.”
Red flag: Saying the pilot failed because the model was wrong, or that the AI team owns the transformation.
Where you start decides what your organization learns about AI. Start where the work matters, feedback is fast and mistakes can be undone.
"Where can we use AI?" gets the same answer everywhere. "Where do we start?" is the leadership decision, because sponsors, specialists and patience are scarcer than ideas.
Avoid four wrong starts - easiest first, everything at once, biggest first, copy the headline - and projects whose results arrive only in years.
Judge a first win by value plus learning - it matters, it can be finished, its errors are caught before they count, and it earns trust.
Say it: “Start where you will know within months whether it works, and where a person catches the mistakes.”
Red flag: Starting wherever is easiest or most exciting, announcing it as a transformation, or giving every department its own pilot.
AI is a label for a family of capabilities. Judge any system by what it does, how reliably, and what happens when it is wrong.
AI is the capability of a computer system to perform tasks that normally require human thinking. The OECD and EU definitions center on inference from input to output.
Six capabilities form the family - perceive, understand, learn, reason, generate, act - and every real system has an uneven profile.
Translate every 'AI-powered' claim with six questions: task, input, output, evidence, failure, outcome. Regulators now punish false AI claims.
Say it: “Don't ask whether it's AI. Ask what it can do, how reliably, and what happens when it's wrong.”
Red flag: Treating 'AI-powered' as an answer, or arguing about whether a system is 'really AI' instead of asking what it reliably does.
An LLM predicts the next token from context. It supplies the language; your systems of record supply the facts.
One model does many tasks because pretraining gives broad patterns, post-training teaches it to follow instructions and the prompt sets the task.
What a model learned is patterns, not records. Exact, current or verifiable facts, including references, come from systems of record at run time.
Reasoning models work through steps before answering - stronger on multi-step problems, slower and costlier, and no better at recall. Visible steps are not proof.
Say it: “The model supplies the language. Your systems supply the facts.”
Red flag: Asking the model for facts, figures or references and trusting the fluent answer, or making it the system of record.
Scale multiplies what each unit leaves behind. Know contribution per unit, break-even volume and the assumption that moves both.
Contribution is value per unit minus complete variable cost, review included. Break-even volume is fixed cost divided by contribution.
Scale spreads fixed cost only over units that arrive, and can raise unit cost through harder cases, review and step costs.
Averages hide heavy users; fit the price structure to the cost structure, and measure the most sensitive assumption, usually value, first.
Say it: “Scale multiplies what each unit leaves behind. Know that number before you grow.”
Red flag: Approving a scale-up because cost per request fell and usage grew, with no contribution per unit, no break-even volume and no scenario range.
AI already reasons, perceives and acts. Watch four dials - reliability, cost, autonomy, physical reach - and act when one crosses your written threshold.
Reasoning, multimodal and agentic AI are in production in 2026; in McKinsey's 2026 survey, 40% of respondents at firms above 1 billion dollars in revenue reported scaling agents.
Reliability is the binding dial: computer-use agents succeed on about 66% of everyday tasks, and METR's 80% horizons are about five times shorter than 50% ones.
Prices for a fixed level of capability fell 9x to 900x a year, so parked ideas need written thresholds, owners and quarterly re-tests.
Say it: “Prepare for capabilities, not predictions. Watch the threshold, not the announcement.”
Red flag: Presenting reasoning, multimodal AI or agents as future steps, or setting strategy by predicted dates instead of written thresholds.
Prepare the organization, not the forecast. Watch thresholds, test on your own work, commit at the speed of undo, and review on a date.
In January 2026, 56% of CEOs reported no revenue or cost benefit from AI; the 12% reporting both had stronger foundations, not better models.
Every module reduces to three threads: redesign the work, demand evidence before scale, and give every outcome an owner.
Nike's 2001 planning failure shows the risk: an unproved system wired into a decision that could not be undone.
Say it: “Prepare the organization, not the forecast.”
Red flag: Answering "what is our AI future?" with a prediction about which model or vendor will win.
Frameworks to draw from memory
Three levels of changeSix questions for any AI claimCapability-to-value chainThe strategy kernelFive use-case patternsFour risk multipliersSeven decision rightsSix cost layersReady for what - score with evidenceAdopt / Trial / Watch / Park radarUnderstand, question, apply, discuss, act
Terms
10-20-70 rule
BCG's description of how AI leaders spend resources: about 10% on algorithms, 20% on technology and data, 70% on people and processes.
90-day AI plan
A contract for one quarter: a few measurable outcomes, one owner each, the dependencies they need and three decision dates.
Abstention
The system declines, asks a clarifying question or escalates when evidence is missing, unreadable or unclear.
Agent loop
Plan, act, observe, decide - repeated until the goal is met or a stopping condition ends it.
Agent mandate
The written delegation for an agent - its outcome, systems, actions it may take alone, handoff rules and business owner.
Agentic AI
An AI system that pursues a goal in a loop - planning, calling tools and acting - without a new human instruction at each step.
AI adoption curve
The path from owning AI tools to redesigned work, as AI moves from licenses to people, into processes and into how the business competes.
AI advantage stack
Six layers from foundation models to customer experience; the higher the layer, the harder it is to copy.
AI agent
A system that pursues a goal by choosing and taking a sequence of actions through tools, adjusting to each result.
AI awareness
Knowing that AI matters to your industry and organization, without yet knowing what it can reliably do or what to change.
AI follower
An organization with the same AI that starts with tools, spreads effort across disconnected pilots and counts activity instead of outcomes.
AI governance
The system of decision rights, accountability, policies, controls and oversight that keeps AI use responsible and aligned with business objectives.
AI inventory
A living record of every AI use, with its owner, purpose, data, provider, users, actions, legal and internal tier, and controls.
AI leader
An organization whose habits - problem choice, ownership, redesign and proof - turn widely available AI into measured changes in how work is done.
AI leadership
Changed behavior: sponsoring AI work, asking for evidence, redesigning processes and owning the decisions AI touches.
AI literacy
Understanding what AI can and cannot reliably do well enough to reason about its use, risks and requirements.
AI portfolio
All AI experiments, products, automations and shared capabilities the enterprise funds, managed together against one budget and one pool of people.
AI risk
The potential business impact when an AI-enabled system produces, amplifies or acts on an incorrect, unsafe, unauthorized or inappropriate outcome.
AI system (OECD and EU)
A machine-based system that infers from its input how to generate outputs such as predictions, content, recommendations or decisions.
AI transformation
Changing how an organization works so that an AI capability produces repeatable business results, not just a successful demonstration.
AI translator
A person with enough domain knowledge and AI literacy to turn a business problem into sound AI work, or to say no.
AI value
The measurable business benefit AI causes by changing work, compared with what would have happened without it, net of cost.
AI washing
Claiming that a product or service uses AI, or uses it more capably, than it actually does.
AI-enabled professional
Someone who uses AI well, safely and critically in their own role without being a technical specialist.
Allocation rule
One written method for charging shared platform costs to AI systems, applied to every case so nothing is hidden or counted twice.
Anchor problem
One real business problem, not an AI project, that a leader carries through every module as a learning reference.
Artificial intelligence
The capability of a computer system to perform tasks that normally require human thinking, such as recognizing, predicting, generating or deciding within limits.
Attention
The Transformer mechanism that lets each token weigh the earlier tokens in its context, so context steers the prediction.
Augmentation
AI raises the speed or quality of a person's work while that person stays in control and owns the outcome.
Automate versus augment
Whether AI replaces a task or strengthens the person doing it. The same exposure can shrink hiring or raise performance.
Automation
A machine performs a repetitive task end to end, usually to cut cost, time or errors.
Benefit owner
The named business leader accountable for a benefit line; finance validates the money.
Benefit variance
The gap between planned and realized benefit, split by cause: adoption, benefit per use or conversion.
Blast radius
Everything a manipulated AI system could access, change, send or trigger before someone stops it.
Bottom-up adoption pressure
Employees adopt a technology before the organization approves it, so leaders must see and bound use, not only introduce it.
Bounded experiment
A visible trial with a few written rules, a measure and a date to keep, change or stop it.
Break-even volume
Fixed cost divided by contribution per unit: the volume at which total value covers total cost.
Brussels effect
Firms adopting EU rules worldwide because one global standard is cheaper than several (Anu Bradford, 2020).
Build
Developing a significant AI capability inside the organization, owning its behavior, data and roadmap.
Business owner
The leader with authority over the process who is accountable for the AI outcome, as distinct from the team that builds the AI.
Business productivity
Valuable output produced relative to the resources used: people's time, technology, capital and outside services.
Buy
Acquiring a mature capability from the market; it still needs decisions on data, security, integration and governance.
Calibration
How well a system's stated confidence matches how often it is actually right.
Capability
What the technology can do - generate, classify, predict, retrieve or optimize - independent of any business workflow.
Capability dials
The four things still changing fast: reliability, cost per task, autonomy and reach into the physical world.
Capability profile
What a particular system does reliably, task by task - strong at some tasks, weak or blind at others.
Capability radar
A list sorting each tracked capability into Adopt, Trial, Watch or Park, moved only by evidence.
Capability versus reliability
What a system can do on its best day versus how dependably it does it; fluent output can still be wrong.
Capability-to-value chain
Capability, adoption, behavior change, outcome, value. Each link is necessary; none is sufficient on its own.
Capacity creation
Time and attention AI frees. It becomes value only when management assigns it to output, quality, growth or lower cost.
Catastrophic forgetting
A neural network losing skills it had when it is trained on new material, which is why every retrained version needs testing.
Cheat sheet
One page per level that gathers every chapter's card: one idea, key points, terms, frameworks, red flags and interview questions.
Competitive advantage
Producing at lower cost than rivals, or delivering more perceived value, or a mix of the two (Rumelt).
Complements
The processes, skills, data and organizational changes that a general-purpose technology needs before it pays off.
Constraint layer
The capability layer whose weakness currently limits what the organization can achieve with AI.
Contribution per unit
Value of one unit minus the variable cost of delivering it, including expected review and failures; what remains to cover fixed cost.
Controlled retirement
Switching an AI system off deliberately - access removed, records archived, data handled, contracts ended, users told.
Convergence
Several technologies maturing at roughly the same time, so that each one makes the others useful.
Copilot
An AI assistant that drafts, suggests or summarizes while a person stays in control and takes the action.
Copy test
Asking what would still be hard to copy if a competitor got exactly your AI model tomorrow.
Cost of a first try
What a person must buy, install, learn or ask permission for before trying a technology; for generative AI it is close to zero.
Counterfactual
What would have happened without the AI initiative; estimated with a baseline, comparison team or staggered rollout.
Data poisoning
Planting tainted material in the data a model learns from so it misbehaves later, often long after the plant.
Decision rights
Clear answers to who may propose, build, approve, deploy, change and stop an AI system, and who is accountable for its outcome.
Deployer
Under the EU AI Act, an organization that uses an AI system under its own authority. It has duties separate from the provider's.
Differentiation test
Ask whether we would still have an advantage if rivals had this capability tomorrow. If yes, it is not our edge.
Digital Omnibus on AI
The 2026 EU regulation amending the AI Act; it moved high-risk dates to December 2027 and August 2028.
Diseconomies of scale
Cost per unit rising with volume, through harder cases, rising review shares, step costs or lower value per unit.
Embedded AI
AI built into software bought for another purpose - forecasting, ranking, routing or fraud scoring - often without the AI label.
Enabler
A shared capability, such as clean master data, that several initiatives need and that has little standalone return.
Enterprise AI readiness
The organization's ability to turn a specific AI ambition into repeatable, governed, economically sustainable work.
Enterprise AI strategy
A coordinated set of choices about where AI matters, what the organization builds and funds, and what it will not do.
Escalation of commitment
The tendency to invest more in a failing course of action one is personally responsible for.
Evidence question
A request that names a claim, a baseline, a threshold, a guardrail and the decision the result will settle.
Excessive agency
OWASP's name for a system with more functionality, permissions or autonomy than its job needs.
Executive AI literacy
Enough understanding of AI to judge its strengths and limits, ask the questions that expose value and risk, and own the decision.
Executive sponsor
A senior leader who owns an AI initiative's business outcome, can fund it into production and has the authority to stop it.
Exit strategy
A plan for leaving a vendor or partner - data portability, migration effort, alternative suppliers and contract terms.
Explain test
Rate your confidence, explain the topic in three plain steps, point to a decision where you used it, then re-rate.
Feedback loop
The time between doing the work and knowing whether it worked. A first win needs one measured in weeks or months, not years.
Fine-tuning
Further training a model on examples to change its behavior or specialization; not a way to keep facts current.
First win
A first AI project chosen to deliver real value and to teach the organization: hard enough to matter, realistic enough to finish.
Fixed workflow
Steps set in advance by code, with AI used inside single steps; cheaper and easier to test than an agent.
Fixed-cost absorption
Fixed cost spread over more units as volume grows; low adoption leaves the same cost on fewer units.
Flywheel
A loop where use creates feedback, feedback improves the offer and a better offer brings more use.
Foundation model
A general model trained once on broad data at scale and adapted to many tasks, such as drafting, summarizing and translating.
General-purpose technology
A technology that is pervasive, keeps improving and spawns complementary innovation, such as steam, electricity or computers.
Generative AI
General-purpose AI that drafts, summarizes and transforms content in response to requests in ordinary language.
Grounded answer
An answer built from retrieved evidence, with sources a person can check against each claim.
Groundedness
Whether an answer is supported by the source the system was supposed to use, such as current company policy.
Hallucination
Fluent model output that is false or unsupported, such as an invented fact, figure or reference.
Handoff
A case the agent stops on and passes to a named person, with the request, what it checked and why it stopped.
If-then plan
A commitment of the form 'when X happens, I will do Y', which makes follow-through more likely than a general intention.
Illusion of explanatory depth
The tendency to feel you understand how something works far better than you can actually explain it.
Imperfect imitability
Barney's term for resources rivals cannot easily copy, because of history, causal ambiguity or social complexity.
Indirect prompt injection
Hostile instructions planted in a document, email or web page that the AI later reads; the attacker never talks to it.
Inference
Using a trained model to turn an input into an output. The parameters stay unchanged.
Informal AI
AI tools employees bring themselves - personal accounts, extensions, departmental subscriptions - outside approval. Often called shadow AI.
Installation and deployment
Perez's two periods of a revolution: speculative build-out first, broad productive use later, often after a crash.
Isolating mechanism
Whatever stops rivals from closing an advantage, such as contracts, relationships, reputation, scale or tacit know-how.
Jagged frontier
The uneven boundary of AI capability: similar-looking tasks can fall inside it, where AI helps, or outside it, where AI hurts.
Language as the interface
Reaching AI by stating the outcome you want in ordinary words, instead of learning code, query languages or specialist screens.
Large language model
A large neural network trained on vast text and code to predict the next token from context; the base of most AI assistants.
Leader's loop
Watch thresholds, test on your own work, commit at the speed a decision can be undone, and review on a fixed date.
Least privilege
Running each program or agent with only the privileges its task requires, and nothing because it might be useful later.
Legal tier
The category the EU AI Act assigns to a use - prohibited, high-risk, transparency or minimal - regardless of internal scoring.
Lethal trifecta
Private data, untrusted content and external communication in one system; together they let an attacker steal data.
Lifecycle governance
The policies, decisions, controls, reviews and accountability applied to an AI system across its whole life, from discovery to retirement.
Long pole
A dependency on the critical path, often in another function, such as a security review or data agreement, that sets the earliest finish.
Material change
A change that alters an AI system's risk, impact, data, users, autonomy or decision consequences, and so requires reassessment.
Metered cost
Cost that moves with each request, page of context, answer and human check, unlike a license that sits still.
Minimum required readiness
Building the capabilities the current stage and the next stage need, rather than every foundation before starting.
Mixed diagnosis
An honest assessment that names where an organization leads and where it follows, function by function, with one strength and one gap.
Model serving
Running a trained model in production: accepting requests, scaling with demand and returning answers fast enough and cheaply enough.
Net AI value
Total benefit minus total cost of an AI capability, with both sides tested at expected scale.
Net present value
Future net cash flows discounted at the required rate of return, minus the investment.
Not-yet list
The written list of deferred work, each item with its reason and the quarter in which it will be reconsidered.
Official AI
AI tools the organization licensed as AI, with a contract, a named owner and a usage policy.
Over-trust
Approving AI output because it sounds fluent and confident, without checks matched to the cost of an error.
Oversight duty
A director's duty to make a good-faith effort to have reporting on mission-critical risks and to act on red flags.
Parity
An investment every competitor can make. It protects the business but does not set it apart.
Partner
Combining our domain knowledge and data with a specialist's skills to create a capability neither could easily make alone.
Pattern mix
The combination of create, understand, predict, decide and act that a problem needs, as opposed to the one a team already owns.
Payback
The time it takes to recover the investment; it ignores the time value of money and flows after payback.
Permission-aware retrieval
Search that applies the asking user's access rights before any content reaches the model.
Planning fallacy
The tendency to underestimate how long one's own work will take, even when asked for a worst-case estimate.
Post-training
Training after pretraining, using examples and human ratings, that teaches a model to follow instructions helpfully and safely.
Posture
A steady way of operating that keeps an organization able to benefit whichever way AI moves, instead of betting on one forecast.
Practice testing
Recalling material instead of rereading it. It improves retention after a week, even though rereading feels more reassuring.
Pre-mortem
Before committing, imagine the initiative has already failed and list why. Prospective hindsight surfaces more reasons than asking what might go wrong.
Preemption
Federal law overriding state law. In US AI policy it is being sought by executive action, which does not by itself repeal state laws.
Process carries the AI
The workflow itself depends on AI, so results no longer rely on individuals remembering to use a tool.
Productivity paradox
Wide adoption of a technology with little measured productivity gain, as Solow observed for computers in 1987.
Prompt injection
Untrusted text that steers an AI system to act against its owner's intent, typed directly or hidden in content it reads.
Proportionate governance
Controls that rise with the impact of an AI use, so low-risk experiments move fast and high-impact uses get stronger review.
Quality-adjusted productivity
Output change times quality change. Thirty percent more output at 20 percent lower quality is only about a 4 percent gain.
Readiness record
For each dimension, the current score, the evidence, the level the ambition requires, and the action that closes the gap.
Realized ROI
Measured benefit minus actual full cost, divided by actual full cost; known only after the money is spent.
Reasoning model
A language model trained, largely with reinforcement learning, to work through intermediate steps before answering; stronger on multi-step problems, slower and costlier.
Rebound effect
When a resource gets cheaper to use, total use can grow so much that total spending rises. Also called the Jevons paradox.
Recoverable error
A mistake that is seen and fixed before it counts, such as a draft a specialist checks before it is sent.
Relevant range
The band of activity within which a fixed or step cost stays flat. Forecasts are valid only inside it.
Repeat-run reliability
How often an agent succeeds every time the same task is run again, not just once.
Retrieval-augmented generation (RAG)
Retrieving relevant, permitted information and giving it to a model as context before it generates an answer.
Review date
The date on which a commitment is decided again: keep, change or stop. A planned stop is a result, not a failure.
Risk classification
Assigning each AI use a tier that decides who reviews it, who approves it and how closely it is monitored.
Risk multipliers
Reversibility, reach, detectability and speed - the four questions that make the same error trivial or serious.
Role track
A curated list of 10 to 14 Level 1 chapters for one role, such as CEO, CFO, COO, risk and legal, or CTO and CIO.
Role-specific AI literacy
Enough understanding to judge AI output and redesign the work in your own role. It does not mean learning to code.
Rules versus access
Rules govern what you may do in a market; access governs whether you can obtain chips, compute, models or markets at all.
Silent error
A task the agent completed wrongly without anyone noticing; found only later, or by sampling completed work.
Six components
Technology, data, people, process, governance and leadership - the parts that must hold together before AI creates business value.
Socio-technical
NIST's term for AI risk arising from technology together with how, where and by whom it is used.
Spaced practice
Reviewing material in short sessions spread over time, such as a two-minute card a week later, rather than in one sitting.
Spread and depth
Spread counts who has access or uses AI. Depth asks whether the work itself now depends on it.
State
The agent's running record of where the job stands, so it neither repeats nor skips steps.
Step cost
A cost that stays flat within a range of activity, then jumps to a new level, such as one more reviewer or capacity block.
Stop rule
The result, written before a test starts, that would make the team halt or redesign the use case.
Strategic refusal
An explicit decision about what the organization will not do with AI, so scarce talent, data and money go to the priorities.
Strategy kernel
Rumelt's three parts of a good strategy: a diagnosis of the critical obstacle, a guiding policy and coherent actions.
System of record
Where the business keeps the official version of something, such as the CRM, billing, ledger or ticketing system.
Task decomposition
Breaking a role into the tasks it actually contains, so AI's impact is judged task by task rather than by job title.
Task productivity
How much faster or better one task is done with AI. An input to business results, not proof of them.
Task-level suitability
Deciding for each task in a workflow whether a fixed rule, AI, a person or removal is the right answer.
Test-time compute
Extra computation a model spends while answering, such as reasoning step by step; it raises cost and latency per answer.
Three clocks
Technology (a capability appears), employee (people use it) and enterprise (the organization sees, bounds and scales it). Leaders close the last gap.
Three task states
Human-led (the person does it), AI-assisted (the person uses AI and stays accountable), AI-automated (AI does it; the person handles exceptions).
Threshold
The quality, cost, speed and risk line a capability must cross for one named workflow before you act.
Time horizon
METR's measure: the length of task, in human expert time, that a model completes at a given success rate.
Total cost of ownership
Everything it costs to build, run, operate and change an AI system over its whole life, including migration and retirement.
Trace-back test
Checking that an AI initiative links up through a priority and the AI strategy to the business strategy before it is funded as a bet.
Traditional AI
Task-specific systems that predict, classify or optimize, usually built into a business application rather than used directly.
Training
Repeatedly adjusting a model's parameters on example data until its outputs improve; this creates the model's capability.
Training at scale
Teaching a model on vast data with large clusters of specialized chips; at the frontier, its cost keeps rising.
Transformation
Redesigning the process, roles, controls and measures around what AI makes possible, not just adding a tool.
Unit of economics
The business unit a case is measured in - per case, document, customer or shipment - for both cost and value.
Use case
A specific application of an AI capability to improve a business activity, decision or workflow, with a user and a measurable outcome.
Use-by-market map
A table of each AI use against each market served, showing the duties and dates that apply in each.
Use-case chain
Role, current problem, AI capability, new workflow, measurable outcome. A proposal missing a link is not yet a use case.
Use-case design brief
One page naming the problem and measure, four roles, workflow map, task split, data and systems, and the smallest test with a stop rule.
Value hypothesis
A testable claim: if this capability, for this workflow, then this outcome improves by this much, while a constraint holds.
Weakest link
The dimension whose gap limits the whole ambition, however strong the others are; after Kremer's O-ring theory.
Web-scale data
The public internet's text, code and images used as training material; it does not include your company's private information.
Workflow integration
AI wired into the systems, hand-offs and sign-offs where work happens, rather than used in a separate chat tab.
Interview questions
Answer out loud first, then flip the card.
I would start with the gap between adoption and value. McKinsey's 2025 survey found that 88% of organizations use AI, but only about 6% attribute 5% or more of EBIT to it, and those high performers were nearly three times as likely to have redesigned their workflows. Electricity followed the same pattern: factories had motors for decades before productivity rose, and it rose when they rebuilt the floor around them. So my ask is not a bigger platform budget. It is one important process redesigned end to end, with a business owner, a baseline and a 90-day measure, and clear points where people stay accountable.
Avoid
Leading with vendor names or model features
Promising headcount cuts
Saying we will wait until the technology settles
Follow-ups
Which process would you start with, and why?
How would you know in twelve months that it worked?
I would stop treating access as an advantage. The best closed and open models are within a few points of each other, so any competitor can match our model within months. I would run four moves as a loop. Choose a handful of business problems, not a long list of ideas. Own each with a senior sponsor who can fund it into production and stop it, with privacy, security and oversight built in from the first experiment. Change the workflow and train people in the real job. Prove the result against a baseline, then scale what beats it and stop what does not. Then I would say honestly where we follow today, function by function, and close the gap that matters most.
Avoid
Leading with which model or vendor to buy
Counting pilots or licenses as progress
Treating follower as an insult rather than a diagnosis
Follow-ups
Which of the eight behaviors is weakest in your organization today?
When did you last stop an AI initiative, and what did you learn?
AI itself is not new for us: we have run forecasting, fraud and routing models for years, mostly inside applications. What changed is the interface. The model behind ChatGPT's launch was largely in place already; what drew a million users in five days was being able to ask in plain language. So almost anyone can now use AI, which widens opportunity, because a first try no longer needs a new system. It also widens responsibility: more people can make fluent mistakes, put data in the wrong place, or face our customers through a chatbot. I would ask for three things: keep our existing models in the jobs they do well, say which kind of AI owns which job, and set one rule for generated work - a named person checks it before it leaves the team.
Avoid
Calling generative AI the whole of AI
Treating easy to use as safe to deploy
Leading with vendor or model names
Follow-ups
Which existing system would you protect from being bypassed?
What would you tell employees on their first day of access?
I would start from the assumption that the model is probably not the main problem, and test the other five components. Data: is the information the pilot used available, clean and agreed every day, or did someone curate it for the demo? People: who uses it on a normal Tuesday, and are they still measured on the old process? Process: did the workflow change, including exceptions and handoffs, or did we bolt a tool onto the old path? Governance: do we know what it may see and decide, and who is accountable when it is wrong? Leadership: is there one business owner with authority over the process, or does every function own a piece? The component with no clear answer is usually the reason. I would fix that before spending more on technology.
Avoid
Blaming the model or switching vendors first
Assuming more training alone will fix adoption
Leaving ownership of the result with the AI team
Follow-ups
Which component would you fix first, and who owns it?
What would you expect to see change in the first month?
I would be neither impressed nor dismissive; I would translate the label. 'AI-powered' says nothing about the task, the input, the reliability or the failure behavior. So I would ask six questions: what exactly it does, what input it needs and what it cannot read, what it produces, what evidence shows its reliability on work like ours, what happens when it is wrong, and which business outcome it changes. I would press hardest on failure, because silent, confident errors are the expensive ones; a lawyer was sanctioned in 2023 for filing cases a chatbot invented. Then I would test it on a sample of our own material before any commitment.
Avoid
Debating whether it is 'real AI'
Accepting a headline accuracy figure without asking how it was measured
Dismissing the product because the label sounds like hype
Follow-ups
Which of the six questions matters most for a legal or financial use?
What would a good answer to 'what happens when it is wrong?' sound like?
I would separate the two phases. Training built the capability once, usually at the provider. Every question after that is inference, and inference uses compute every time, so cost grows with volume, with how much text goes in and out, and with how much reasoning the model does per answer. At Google, inference took about three-fifths of machine learning energy. Before approving, I would ask the pilot team for requests per user, cost and response time per request, and the share of requests that genuinely need a reasoning model. Then I would model full volume, set the reasoning effort by type of request and track cost per resolved case, not per question.
Avoid
Assuming cost is fixed once the model exists
Using the most powerful setting for every request
Measuring cost per question instead of per outcome
Follow-ups
What would you measure in the pilot to forecast full cost?
When is a reasoning model worth its extra cost?
I separate the language from the facts. Drafting, summarizing and adjusting tone are what these models do best, and I want that value. But the model holds patterns, not records, so I ask where every fact in the output comes from - prices, stock, policy, figures, references - and insist they are fetched from our systems at the moment of the request, not recalled. Then I ask what happens when it is confidently wrong, who checks high-stakes output before it leaves the team, whether those people are trained, and what baseline we measure against. Tromso shows the cost of skipping this: 11 of 18 references in a public report did not exist. With live facts, checks and an owner, I approve a limited pilot.
Avoid
Treating a fluent demo as proof of accuracy
Letting the model supply prices, policy or references from memory
No named owner for wrong output
Follow-ups
Which outputs would a person check first?
What would you measure in the first month?
I'd start with the gap we are closing. If it is access to policies, prices or procedures that change and differ by role, training is the wrong tool. Research comparing the two found retrieval beat fine-tuning for both familiar and new facts. Every change would mean retraining, old and new versions would blur, and anything trained into the model can no longer be restricted by role. I'd keep documents in their authoritative systems, retrieve what each user may see at question time, and require a source and an effective date on every answer. Live records come through APIs. Fine-tuning stays an option if evaluation shows a behavior gap - format, tone or a specialized task - rather than a knowledge gap.
Avoid
Accepting 'the model will know our business' as the goal
Planning to handle permissions after launch
Treating a vector database as the strategy
Follow-ups
What evidence would make you reconsider fine-tuning?
How would you test that permissions hold before launch?
First, is it really an agent? It should be given a goal, choose its own next step, act through tools and adjust to the results. If the steps are the same every time, a fixed workflow is cheaper and easier to test. Then the design: which tools it can reach, which exact rules run in software rather than in the model, and the five ways the loop stops - goal met, limit reached, person needed, tool failed, policy blocks. Finally, evidence: end-to-end success on real tasks, wrong actions, escalations and cost per completed task. I would approve a narrow start on a low rung of autonomy and widen it only on that evidence.
Avoid
Approving because the demo looked impressive
Accepting step accuracy as job accuracy
No answer on when the agent stops
Follow-ups
Which step would you move into exact software first?
What evidence would justify more autonomy?
I would say adoption is necessary but it is not the result. Daily users tell us the capability is in use; they do not tell us whether work changed or whether any outcome moved. I would walk the chain: which workflows changed, which outcome metric moved against the baseline or comparison team we set, and what that is worth net of the full cost. Where we have no baseline, I would say so and set one now, for example by staggering the next rollout. Then I would report two or three outcomes, not a usage chart.
Avoid
Treating usage or satisfaction as value
Quoting time saved as money saved
Claiming a gain with no comparison
Follow-ups
Which outcome metric would you report first?
What if the baseline was never measured?
I would say the pilot has created capacity, not savings yet. First I'd check that the two hours cover the whole job, including review and rework, and that quality held; self-reported time is a weak measure on its own. Then I'd show what we decided to do with the time: more output from the same team, a backlog cleared, a planned hire avoided or an outside contract ended. Only the last two reduce cost directly. I'd report an outcome measure, such as files completed per week or customer waiting time, not hours saved or usage. If nobody has decided where the hours go, the honest answer is that there are no savings yet, and that is the next decision to make.
Avoid
Multiplying hours saved by salary and calling it savings
Promising headcount cuts equal to the percentage gain
Citing usage or prompts as evidence
Follow-ups
Which outcome measure would you put on the dashboard?
What would make you stop the rollout?
I would split the decision in two. Backward: explain the gap by cause. How much is adoption, how much is lower benefit per use, such as extra review, how much is freed capacity that was never converted into lower spend, and how much is cost overrun? Each has a different owner and fix. Forward: the money already spent is sunk, so I compare the remaining benefit with the remaining cost, under the current capture rate and with a costed fix. If continuing still adds value, I continue and fix, name a business owner for each benefit line, and review monthly. I would hold any expansion until adoption reaches plan, and feed the realized numbers into our next business case.
Avoid
Deciding on money already spent
Declaring failure without splitting the variance
Leaving the benefit with the AI team
Follow-ups
Which variance would you check first?
What evidence would finance use to confirm the benefit?
I would say we have a lot of activity and the start of a strategy. Forty projects show we are learning, not that we are aligned. I would bring the board a one-sentence diagnosis of the business obstacle AI must help overcome, three to five priorities, and a written list of what we will not do. Then I would trace every project to those priorities: the ones that trace become the portfolio, the useful general tools move to the ordinary software budget, and the rest stop. BCG's 2024 survey found that AI leaders pursued about half as many opportunities as others, so I would expect the count to fall and the returns to become visible.
Avoid
Citing the number of projects as proof of strategy
Leading with a model or vendor choice
Promising a strategy document with no refusals in it
Follow-ups
Which project would you stop first, and why?
What stays with the business units?
Probably, but not by default. I would split the capability into layers and run two tests on each. Differentiation: if competitors had exactly this tomorrow, would we still have an advantage? If yes, it is not our edge and buying is sensible. If no, we should own it - build if we can run it for years, partner if we cannot, and keep the data and workflow on our side. Control: how badly would a change in the vendor's price, behavior or roadmap hurt us? Usually we buy the model and infrastructure, own the data and workflow, and make sure the contract gives us an exit.
Avoid
Choosing on benchmark scores or price alone
Saying strategic means we must build everything
No view on exit or data ownership
Follow-ups
Which layer would you insist on owning?
What would make you revisit the decision in two years?
I would sort the portfolio with two questions: does each initiative clearly move our cost or the value customers see, and how quickly could a well-funded rival match it? Most will land in parity. That is fine; we fund those fast and cheaply and judge them on speed and unit cost, but we do not call them strategy, because once rivals buy the same tools the savings pass to customers through price. For the few that remain, I look for the guard: a contract, a customer relationship, reputation, scale or know-how that stops the gap closing. The best candidates feed a loop where each round of use makes the offer better, and someone is accountable for keeping that loop turning.
Avoid
Naming a model or vendor as the advantage
Treating early launch as a lasting lead
No distinction between parity and advantage
Follow-ups
Which of our current initiatives would you relabel as parity?
What would our equivalent of a paid-per-hour contract be?
I would not fund eighty. BCG found that AI leaders pursue about half as many opportunities as their peers and scale twice as many. I would ask each sponsor to fill the same chain: whose work changes, what the problem costs today, which capability helps, how the workflow changes and what number will move. Proposals that cannot fill it go back. I would then check the mix: if nearly all are drafting tools, our costliest problems are missing, so I would add the two or three cross-functional workflows that cost us most, each with an owner. From that list I would fund a small, balanced set with baselines, and report workflows improved, not pilots launched.
Avoid
Launching all of them because competitors are moving
Funding only the easiest ones
Letting each department pick in isolation
Follow-ups
What evidence would make you stop a funded use case?
How would you find a valuable workflow nobody submitted?
I would neither approve nor reject the tool yet. I would ask what problem it solves: whose work, which outcome, by how much. Then I would fund a short mapping of the workflow as it really runs, because the time is usually lost in waiting and handoffs. Each task gets its own answer - some are fixed rules ordinary software can do, some suit AI, some stay with people. The result goes on one page with a baseline and a stop rule, and must pass value, feasibility, risk and strategic fit. Then I fund the smallest test. Competitors having one is a reason to look, not to buy.
Avoid
Approving because competitors have one
Starting with vendor or model selection
No baseline or stop rule
Follow-ups
What would make you stop the test?
Which task in that workflow would you give to a rule?
I'd ask what happens to the misses. If the 90 percent agent acts anyway when it is wrong, its tenth case is a silent error, found weeks later at a high price. The 70 percent agent's misses are visible handoffs that a person resolves today. So I'd price both: the cost of a silent error against the cost of a handoff, using a weekly sample of completed work to measure silent errors, because logs will not show them. In many workflows the agent that stops when unsure wins. Then I'd check it succeeds when the same case is run again, that its handoffs carry context, and that a business owner signs the mandate.
Avoid
Picking the higher completion rate by reflex
Measuring silent errors from the agent's own logs
No named business owner
Follow-ups
When would the 90 percent agent be the better choice?
How would you sample completed work?
I would thank them for the number and then ask what the system does with its output and who acts on it. What happens in the two percent, and who is affected? Can we undo the result, how far could one error spread, and would we notice an error nobody complains about? How many cases a day does it touch? Who can stop it? If the errors are reversible, contained and visible, I can approve. If not, I would ask for a design that drafts first and automates in steps. The 98 percent does not decide; the use does.
Avoid
Approving because 98 percent sounds high
Rejecting because it is not 100 percent
Leaving the decision to the technical team
Follow-ups
Which multiplier worries you most here, and why?
What evidence would let you widen automation?
Not on that number alone. First, accurate on what: who wrote the test questions, and does their mix match how people will really use it? I want results by slice, because a test that under-samples one process averages its failures away. Second, how often is it wrong and how often does it decline? A system that admits uncertainty is easier to make reliable than one that guesses. Third, what checks sit around the model: does a rule confirm that quoted values appear in the cited source, does it abstain when evidence is missing, and who gets the escalation? If those answers fit the workflow, I would roll out with monitoring by slice. The errors belong to us, not to the vendor.
Avoid
Accepting the average as proof
Assuming the vendor owns the errors
Requiring a human on every answer regardless of risk
Follow-ups
How would you rebuild the test set?
Which errors would you want to catch automatically?
I would assume someone will eventually manipulate it, because any supplier email can carry hidden instructions. So my questions are about harm, not cleverness. This design combines untrusted input, a sensitive system and the power to move money in one session, which is the dangerous combination. Which leg can we remove? Ideally it proposes payments and changes but cannot execute them. Does it run under its own identity with only the permissions it needs? Do bank-detail changes need a human call-back to a known number? Is every tool call logged, and has someone tried to break it? Who can stop it, and how fast? With those answers I would approve a narrow first scope and widen it only on evidence.
Avoid
Accepting 'the vendor secures the model' as the answer
Relying on a prompt filter
No plan to stop it quickly
Follow-ups
Which of the three capabilities would you remove first?
Who owns the emergency stop?
I would ask which rung it needs - generate, recommend, plan, execute or operate autonomously - and why. Then where its actions sit on blast radius and reversibility, and what the worst action is with the access requested. I want least privilege under its own identity, not a shared login; hard caps on value, volume and spend; and a person approving anything external, financial or hard to undo. External content is data, never instructions. Those controls must live in the systems it touches, not only in its prompt, and a second AI is not my check. It starts one rung lower than requested and climbs on evidence. If those answers exist, I approve a scoped start; if not, it stays at recommend.
Avoid
Approving because pilot accuracy was high
Requiring human approval on every single action
Relying on a second agent as the only check
Follow-ups
What evidence would move it up one rung?
How would you stop it in the middle of a bad run?
I would not point to the policy. I would show whether we pass the hundred-uses test for the AI uses that matter: which ones exist, who owns each one, what each touches, which controls apply, and who can stop it. Where we can, we are governed. Where we cannot, we have adoption without governance, and I would say so. Then I would set out the plan: find what is in use, name owners and decision rights inside our existing delegation of authority, and set controls in proportion to impact, so low-risk experiments keep moving while customer-facing and high-impact uses get stronger review. The test I would offer the board: if we launched a hundred new uses tomorrow, would we know who decides and who can switch each one off?
Avoid
Pointing to a policy document as proof
Saying Legal or IT owns it
Proposing a ban until everything is reviewed
Follow-ups
What would you do in the first two weeks?
How do you stop governance from slowing adoption?
Not yet. First, how was the register built? If it came from the project portfolio, it misses AI inside software we already buy, tools staff built and agents. I would run discovery across procurement, cloud bills, security tools and business units. Second, who graded it? When the same questionnaire decides the legal category and the effort required, the answer drifts low; the Dutch Court of Audit said as much. I would have legal decide the legal tier first - prohibited stops, anything in Annex III stays high - then classify each use by impact, data, consequence, autonomy and reach, before controls. Finally I would ask to see one high-tier system and prove it got a different review from a low one.
Avoid
Accepting a nearly all-low register at face value
Classifying by vendor or model name
Letting controls lower the tier
Follow-ups
Which systems would you classify first after discovery?
How would you prove the tiers change the review?
I would not assume the original approval still holds. A new model can change behavior, and moving from staff to customers changes who is affected, what must be disclosed and what data flows. Where I can, I pull the system back to the approved scope while the owner runs our change route: impact assessment, a fresh look at the risk level, evaluation of the changed system and a decision by whoever holds authority at that level. I would not switch off the part that was approved and works. Then I fix the gap: a written definition of material change so the next one is caught.
Avoid
It was approved, so it is fine
Switching everything off by reflex
Not naming who decides
Follow-ups
Which changes would you treat as not material?
What if the old model version has been retired?
I would thank them for real evidence, then ask for the production case. First, which cost layers are in the estimate - model and compute only, or also data, integration, operations and governance? Second, what share of outputs will people review, at what cost, and what rate would make the case stop paying? Third, what is the unit - full cost and realized value per resolved inquiry - and does the gap hold at the expected volume? Fourth, how much released time has someone committed to turn into avoided hiring or more output? If the rebuilt case still pays, I would approve a staged rollout with a cost-per-unit target, not an open budget.
Avoid
Multiplying the pilot bill by the volume increase
Booking all saved minutes as cash
Arguing only about the model price
Follow-ups
What review rate would make you stop the rollout?
Model prices keep falling. Why might the bill still rise?
I would ask the team to restate it as total cost of ownership over a stated life, say three years: build, run, operate and change. Build is engineering, integration, data preparation and a test set. Run is the quote plus cloud, storage and our share of the platform. Operate is the people who check exceptions, monitoring, security and support. Change is model updates: hosted models are retired on the provider's schedule, often within 18 months. Then two questions: where are the step costs, such as the next reviewer or capacity block, and at what volume do they arrive? And which allocation rule puts shared platform costs in, so nothing is hidden or counted twice? Only then would I set cost against value.
Avoid
Negotiating the quote down and approving
Treating staff time as free because salaries are already paid
Ignoring model change because it is not yet scheduled
Follow-ups
Which family would you expect to be largest, and why?
How would you allocate a platform team that serves five AI systems?
I would ask for the unit economics in the business's own unit, a completed case or a served customer, not a request. What is one unit worth in realized value, and what does it cost to deliver, including review and failures? That gives contribution per unit. Then fixed cost and break-even volume, with today's volume marked against it. Then the case at a fifth, two, five and ten times today, with step costs, mix and review rate recomputed rather than scaled. Finally, which assumption moves the result most, usually value per unit, and which segment loses money today. If the case holds, I would approve in stages with contribution per unit reported each quarter.
Avoid
Treating usage growth as proof of value
Multiplying today's unit cost by ten
Follow-ups
What would make cost per unit rise as volume grows?
How would you value one completed case?
Ready for what is the first question. I would segment the twenty by ambition and stage: some are experiments with a low bar, some need production readiness, and any high-risk use carries legal duties. For each group I would assess the eight dimensions with evidence, not opinion, and look for the gaps several initiatives share, such as data ownership, an approval route or production engineering. Those shared gaps become the first roadmap items. Lower-risk work proceeds under control while they close. I would report the weakest link, not an average score, so the board sees what actually limits us.
Avoid
Answering yes because the pilots went well
Freezing everything until the organization is perfect
Quoting an average readiness score
Follow-ups
Which gap would you close first, and why?
Who would you put on the assessment team?
I would not start from the budget, because people, not money, are usually the limit. First I would drop anything that serves none of our stated outcomes. Then I would place the rest on value and feasibility, with ranges and a confidence note, and use a score only to structure the debate. Next I would check capacity: if the shared data team can carry four builds at once, funding fifteen just makes each one slower. I would find the enabler several proposals depend on and fund it first. Each funded item gets an owner, a next gate and a stop rule written now. The rest are deferred with a trigger or stopped.
Avoid
Ranking all 38 by ROI
Spreading the budget across all 15 at once
No stop rules or capacity limit
Follow-ups
What would you do with the sponsor of a stopped proposal?
How would you value the data clean-up?
I would keep the ambition and shrink the quarter. First I would check what the shared specialists can finish in ninety days, and leave slack, because plans run late. Then I would pick three to five outcomes across three lanes: one use case to deliver toward a production decision, one bet with a single question to answer, and only the capability those two need. Each gets a business owner and a result Day 90 can settle. In week one we would request the security review, data access, legal assessment and any works-council consultation. Everything else goes on a written not-yet list with a reason and a quarter. Day 90 decides go or no-go for production, not scale.
Avoid
Accepting the full list to show ambition
Making the AI team the owner of every outcome
Promising scale on Day 90
Follow-ups
What would you cut first if Day 30 shows the plan slipping?
How do you explain the not-yet list to the people whose ideas wait?
No. Reasoning models shipped in 2024, and in McKinsey's 2026 survey 40 percent of respondents at the largest firms reported scaling agents, so those are today's tools. What still moves is reliability, cost per completed task, autonomy and physical reach. I would rewrite the roadmap around named workflows: for each, the error rate it tolerates, the cost per task it can bear and the controls it needs. Then I would keep a radar - adopt, trial, watch, park - with an owner and a quarterly re-test, so we act when a threshold is crossed rather than when a product is announced.
Avoid
Treating agents or reasoning as future capabilities
Committing to dated predictions
Tracking announcements instead of thresholds
Follow-ups
Which workflow would you write the first threshold for?
Who should own the radar?
I would not wait, because several duties are already in force. Korea's AI Basic Act applies since January 2026, and the EU's transparency duties since August 2026; both require telling customers they are dealing with AI. I would ask for a one-page map of the use against each market, with dated duties and an owner. Disclosure, labeling and logging of which model answered I would build once for every market. Where rules conflict or a model is unavailable in a market, I would design a switch. And I would check that a tested second model exists, since access to models varies by country: Meta's license for its multimodal Llama models excluded EU-based companies. Dates do move, as the EU's high-risk date did in 2026, so the map gets reviewed whenever a duty arrives or changes.
Avoid
Waiting for the rules to settle
One global design with no market check
Assuming US federal policy has removed state duties
Follow-ups
Which duty would you design for first, and why?
What would make you stay out of a market?
I would not give them a single forecast, because even AI researchers moved theirs by 13 years in one year. I would show three things. First, the thresholds our plans depend on, such as an error rate on our own cases, a cost per completed task and the EU AI Act dates for our markets, each with an owner. Second, our commitments sorted by how hard they are to undo: no-regret moves now, options, and the few big bets with their evidence. Third, the review date and stop rule on each. Then I would ask them to hold us to that loop every quarter. That is how we stay among the minority of firms getting both revenue and cost gains.
Avoid
Naming a winning model or vendor as the strategy
Promising certainty
No owners or review dates
Follow-ups
Which threshold would change your plan first?
What have you stopped in the last year?
Red flags to avoid
Treating AI as a tool purchase and calling the rollout a transformation while roles, measures and controls stay the same.
Claiming leadership because the company has licenses, a platform or many pilots - access and activity instead of behavior and outcomes.
Calling generative AI the whole of AI, or handing it to everyone because it is easy, with no rule for checking what it produces.
Saying the pilot failed because the model was wrong, or that the AI team owns the transformation.
Treating 'AI-powered' as an answer, or arguing about whether a system is 'really AI' instead of asking what it reliably does.
Saying the model is already trained so using it is nearly free, or that it learns from every conversation.
Asking the model for facts, figures or references and trusting the fluent answer, or making it the system of record.
Proposing to train the model on every company document so it knows the business, with permissions to be sorted out later.
Calling a chatbot an agent, or judging an agent on its accuracy at one step instead of the whole job.
Reporting logins, prompts or pilot accuracy as proof that an AI initiative created value.
Treating hours saved as cash saved, or announcing that a 20 percent productivity gain means 20 percent fewer staff.
Treating the approved ROI as achieved, or deciding whether to continue by looking at the money already spent.
Presenting a count of AI projects, a vendor roadmap or a slogan such as "AI-first" as the AI strategy.
One word for the whole AI effort - "AI is strategic, so we build it all" or "a vendor sells it, so we buy it all".
Presenting an AI tool every rival can license as a strategic advantage, or treating being first as if it were being ahead.
Starting from a tool or a department list, then reporting the number of AI pilots as progress.
Funding an AI proposal from its demo before anyone has mapped the workflow or written a baseline and a stop rule.
Scaling the agent with the highest completion rate without asking what its misses are or what they cost.
Approving or rejecting an AI system on its accuracy score alone, without asking what happens when it is wrong.
Approving a system on one overall accuracy figure without asking who wrote the test and how it scores on each slice of real work.
Saying the model provider handles security, or that a better prompt filter will solve prompt injection.
Giving an agent its user's full access, or a high rung on day one, because the demo worked.
Answering "we have an AI policy" or "Legal handles it" when asked who decides what AI is acceptable.
Classifying the model or vendor instead of the use, or saying a prohibited practice can be approved with stronger controls.
Calling a model swap, new data, new users or new permissions "implementation details" and leaving the original approval in place.
Approving an AI case on the model price and its falling trend, with no unit, no review rate and no ten-times test.
Presenting the vendor quote as the annual cost of the AI system, with no model-change budget and no named step costs.
Approving a scale-up because cost per request fell and usage grew, with no contribution per unit, no break-even volume and no scenario range.
Announcing a readiness percentage without being able to say ready for what, or what evidence supports it.
Funding every sponsored idea at once, ranked by projected ROI, with no capacity limit and no written stop rules.
"Our first 90 days will launch two use cases, test three bets, stand up governance and build the platform."
Presenting reasoning, multimodal AI or agents as future steps, or setting strategy by predicted dates instead of written thresholds.
"We'll wait until the AI rules settle" or "the federal order means state laws no longer apply".
Answering "what is our AI future?" with a prediction about which model or vendor will win.
Treating the program as reading to finish, or as a tool tutorial, instead of practice on real decisions.
Saying AI is the technology team's job, or approving an AI proposal because the demo was impressive.
Pointing to your own daily AI use, or to the number of tools deployed, as proof that you are leading on AI.
Rating yourself by how confident you feel, or starting with your strongest area because progress there feels quick.
Treating market excitement or spending as evidence of value, or assuming the gains will follow adoption automatically.
Saying AI arrived with one product, that we must build our own model, or that we should wait for the next model before learning.
Reporting licenses, pilots and active users as proof of transformation, or giving the whole enterprise one score.
Naming the vendor, the model or "our data" as the advantage without saying what a competitor could not copy.
Equating a fast license rollout with transformation, or proposing to ban AI until the policy is ready.
Saying the AI journey starts when we pick a platform, as if nothing were already running or the vendor owned the outcome.
Calling a renamed chatbot an agent, or limiting an agent by instruction instead of by enforced permission.
Answering "How many jobs will AI replace?" with a global number, or promising that no job will change.
Answering the talent question with a hiring number, as if one central AI team were the whole answer.
Starting wherever is easiest or most exciting, announcing it as a transformation, or giving every department its own pilot.