Public course deck

Cost, efficiency and AI

An operating-loop workshop for executives

74 slides · Print from your browser to export a PDF

Full screen shows slides only; the course material reads below

Back to course overview

P1 Prologue · Cover and agenda

Cost, efficiency and AI

An operating-loop workshop for executives

  • Where the value is
  • How strong the evidence is
  • Who owns it
  • How to scale without losing control
Course material 1 tables · 1 recaps · 1.1

From 1.1 Core claim and hooks

Three-sentence version (for the opening)

  • AI is no longer optional. The State Council set the timetable: over 70% adoption of smart devices and agents by 2027.
  • But using it is not earning from it. 88% of organisations use it; only 13% report enterprise value.
  • Where is the gap? Nobody translated those three saved hours into a number on the income statement.

Why this claim lands with a CEO

What the boss is thinking What this course answers
"Another AI salesperson" I don't sell tools. I teach you how to sign one off.
"We already bought it; results are mediocre" Yes — 87% of companies worldwide are in the same place. The tool is not the problem.
"They say it's more efficient; I don't see the money" Because the time saved has no exit. Here are five exit conditions.
"How much should we spend" Don't ask how much to spend; ask where the baseline is. No baseline, no project.
"Will something go wrong" Five gates; the CFO signs the fourth, and only after the fifth may you replicate.

Hook 3: the −19% experiment

Claim   ── The gap is not in the technology, it is in the operating loop
Figures ── 88% use it, 13% get enterprise value → 75 points of arbitrage

Hooks   ── ① The 75-point gap: everyone is running, most are running on the spot
           ② Superstars: +33.5% vs +163% — AI widens gaps, it does not close them
           ③ +55% or −19%: same tool, two directions, and the difference is management

Close   ── Not how to use AI — how to sign AI off:
           where the value is · how strong the evidence is · who owns it · how to scale without losing control
P2 Prologue · Cover and agenda

What this course is not

  • How large models work
  • Prompting tricks
  • Vendor pitches

All of it is out of date in three months.

Course material 1 tables · 1 recaps · 1.1

From 1.1 Core claim and hooks

Three-sentence version (for the opening)

  • AI is no longer optional. The State Council set the timetable: over 70% adoption of smart devices and agents by 2027.
  • But using it is not earning from it. 88% of organisations use it; only 13% report enterprise value.
  • Where is the gap? Nobody translated those three saved hours into a number on the income statement.

Why this claim lands with a CEO

What the boss is thinking What this course answers
"Another AI salesperson" I don't sell tools. I teach you how to sign one off.
"We already bought it; results are mediocre" Yes — 87% of companies worldwide are in the same place. The tool is not the problem.
"They say it's more efficient; I don't see the money" Because the time saved has no exit. Here are five exit conditions.
"How much should we spend" Don't ask how much to spend; ask where the baseline is. No baseline, no project.
"Will something go wrong" Five gates; the CFO signs the fourth, and only after the fifth may you replicate.

Hook 3: the −19% experiment

Claim   ── The gap is not in the technology, it is in the operating loop
Figures ── 88% use it, 13% get enterprise value → 75 points of arbitrage

Hooks   ── ① The 75-point gap: everyone is running, most are running on the spot
           ② Superstars: +33.5% vs +163% — AI widens gaps, it does not close them
           ③ +55% or −19%: same tool, two directions, and the difference is management

Close   ── Not how to use AI — how to sign AI off:
           where the value is · how strong the evidence is · who owns it · how to scale without losing control
P3 Prologue · Cover and agenda

What this course is

Nine acts. The three peaks fall in diagnosis, accounting and governance.

  1. Act I Alarm The window is closing P4–P9
  2. Act II Reality Using is not earning P10–P17
  3. Act III Diagnosis Where money leaks and process jams P18–P28
  4. Act IV Selection Stop spreading thin P29–P34
  5. Act V Capability What AI can actually do P35–P43
  6. Act VI Evidence Who has actually done it P44–P50
  7. Act VII Accounting Saved hours are not profit P51–P58
  8. Act VIII Governance The line you do not cross P59–P67
  9. Act IX Execution Ninety days, five gates P68–P72
Course material 1 tables · 1 recaps · 1.1

From 1.1 Core claim and hooks

Three-sentence version (for the opening)

  • AI is no longer optional. The State Council set the timetable: over 70% adoption of smart devices and agents by 2027.
  • But using it is not earning from it. 88% of organisations use it; only 13% report enterprise value.
  • Where is the gap? Nobody translated those three saved hours into a number on the income statement.

Why this claim lands with a CEO

What the boss is thinking What this course answers
"Another AI salesperson" I don't sell tools. I teach you how to sign one off.
"We already bought it; results are mediocre" Yes — 87% of companies worldwide are in the same place. The tool is not the problem.
"They say it's more efficient; I don't see the money" Because the time saved has no exit. Here are five exit conditions.
"How much should we spend" Don't ask how much to spend; ask where the baseline is. No baseline, no project.
"Will something go wrong" Five gates; the CFO signs the fourth, and only after the fifth may you replicate.

Hook 3: the −19% experiment

Claim   ── The gap is not in the technology, it is in the operating loop
Figures ── 88% use it, 13% get enterprise value → 75 points of arbitrage

Hooks   ── ① The 75-point gap: everyone is running, most are running on the spot
           ② Superstars: +33.5% vs +163% — AI widens gaps, it does not close them
           ③ +55% or −19%: same tool, two directions, and the difference is management

Close   ── Not how to use AI — how to sign AI off:
           where the value is · how strong the evidence is · who owns it · how to scale without losing control

Act I

Alarm

The window is closing

P4–P9

P4 Act I · Alarm

88% → 36% → 13%

Widely used, rarely earning.

Three surveys, different samples and definitions — not a single funnel.

SourceStanford HAI, AI Index 2026;Accenture, Making Reinvention Real with Gen AI

Course material 2 tables · 2.1

From 2.1 Three numbers on the value gap

Three supporting pillars

Stage Figure Source Definition
Adopted AI (in at least one business process) 88% [R019] Grade A Organisation-level survey
Have scaled a generative AI solution 36% [R024] Grade B Executive self-report
Report significant enterprise-level value 13% [R024] Grade B Executive self-report

Leave this line out and an informed CFO takes you apart. Put it in and you look more rigorous than the consultancy.

Supporting layer (other angles on the same conclusion)

Conclusion Figure Source
Only a third are genuinely redesigning the business 34% are transforming deeply; the rest stay surface-level or partial [R023]
Limited impact on enterprise EBIT About 39% of respondents report an enterprise-level EBIT impact [R020] ⚠️ See below
Spending is unbalanced Three times as much budget goes to technology as to people [R024]
The organisation is not ready Only 35% of executives have a roadmap for how AI reshapes the workforce [R024]
P5 Act I · Alarm

Same technology, a fivefold spread

AI widens gaps. It does not close them.

SourcePwC, Global AI Jobs Barometer 2026

Course material 5 tables · 2.5

From 2.5 Talent and organisation

The superstar effect (the most striking numbers in the course)

Group Revenue per employee growth (2018 baseline)
Most AI-exposed companies (all of them) 33.5%
The strongest 20% of them (the "superstars") 163%

Skill premium: the talent market has already voted with money

Metric 2025 2024
Wage premium for roles requiring AI skills 62% 57%

A two-track labour market (talk role redesign, not layoffs)

[R025] AI is splitting jobs onto two tracks:

Track Share Trait Result
Professionalised 22% Demands deeper judgement and creativity Growing twice as fast as the other track, with 42% faster pay growth
Democratised 52% Lower skill barriers, shifting toward less specialised work Slow growth, slow pay rises
  • Skill requirements in high-AI-exposure roles change twice as fast as in low-exposure roles
  • New tasks relying on empathy, judgement and creativity are growing 2.5× faster

The upheaval in entry-level roles (what management should watch most)

Finding Figure
In high-AI-exposure entry-level roles, the share of new skill requirements that are traditionally senior capabilities (motivating, strategic decisions, team building) 52%
The same share for entry-level roles with low AI exposure 7%
Entry-level roles pushed up the seniority ladder this way 35% growth
Total entry-level roles worldwide with high AI exposure Essentially flat

The global labour picture (WEF; for the macro backdrop slide)

Metric Figure
Roles created before 2030 170 million
Roles disappearing over the same period 92 million
Net gain 78 million (about 7% of today's employment)
Share of current skills transformed or obsolete between 2025 and 2030 39% (down from 44% in 2023 and 57% in 2020)
If the global workforce were 100 people, how many need training before 2030 59 people
Of these: 29 upskilled in place, 19 moved roles after upskilling, 11 get no training and face threatened prospects
Fastest-growing skill AI and big data, then networks and cybersecurity, then technical literacy
The most valued core skill Analytical thinking (essential for 7 in 10 firms), then resilience and agility, then leadership and social influence
P6 Act I · Alarm

+55% or −19%?

One class of tool, two directions.

Controlled experiment: a set coding task

+55%

Developers with AI assistance finished faster

Another study: experienced OSS developers

−19%

Slower — and they believed they were faster

Three things separate them: task boundaries, whether quality is checked, and how work is divided. All management questions.

NoteMETR could not reproduce the result later, largely because developers were no longer willing to work without AI.

SourceGitHub controlled experiment; METR / Becker et al., 2025 (not subsequently replicated)

Course material 3 tables · 2.4

From 2.4 Productivity evidence

Task level: evidence that holds (all controlled or quasi-experimental)

Use case Effect Study Grade
Customer service (conversational AI assistant) 14–15% more issues resolved per hour Brynjolfsson et al., 2025[R019] A-
Software development (Copilot) About 55% faster on a specified coding task [R049] A-
Software development (Copilot, a separate study) 26% more pull requests submitted Cui et al., 2025[R019] A-
Marketing (multimodal ad creative generation) Output per person +50% Ju & Aral, 2025[R019] A-
General office work (M365 Copilot) About 29% average speed-up across several task experiments [R050] About 3,000 participants in real working conditions A-
Finding Figure Source
Experienced open-source developers got slower with AI assistance −19% METR / Becker et al., 2025[R019]
And they believed they were faster Perception and measurement diverge As above
Engineers who lean heavily on AI to learn No measurable speed-up, and a "learning penalty" appears Shen & Tamkin, 2025[R019]
Survey of 6,000 executives across four countries Widely adopted, but realised productivity gains are minimal; employment forecast −0.7% over three years Yotzov et al., 2026[R019]

One rule (the conclusion here; put it on the slide in large type)

Observation Management action
The clearer the task boundary, the bigger the gain Standardise the process before adding AI
Gains only stick where quality is monitored Set acceptance criteria before claiming gains
Judgement and deep-reasoning work can show negative returns Don't rush these roles; assist first, don't replace
Experienced people may be slowed down Don't pilot only with experts, or only with enthusiasts
P7 Act I · Alarm

The timetable is already set

State Council opinion on the AI Plus initiative, August 2025.

  1. 2027Adoption of next-generation intelligent terminals and agents above 70%; deep integration across six priority sectors
  2. 2030Adoption above 90%; the intelligent economy becomes a principal growth engine
  3. 2035Full entry into an intelligent economy and society

These are sector-wide targets, not a mandate for any single company's internal adoption rate.

SourceState Council, Guidance on Deepening the "AI Plus" Initiative

Course material 5 tables · 2.3

From 2.3 China policy and industry

The State Council set the timetable (the hardest page)

Guidance on Deepening the "AI Plus" Initiative, issued August 2025 — [R013]

Milestone Target
2027 Deep AI integration across six priority sectors first; over 70% adoption of next-generation smart devices and agents
2030 AI underpins high-quality growth; smart devices and agents exceed 90% adoption; the intelligent economy becomes a major growth pole
2035 Entering a new phase of the intelligent economy and society

The NDRC's framing: the phase has changed

Keyword Meaning How to use it
From experimentation to value creation Officials acknowledge the last phase was experimental; now it must deliver Echoes "88% use it, 13% earn from it"
The last mile Use cases, data and talent are still the blockers Leads straight into the diagnosis section
Supply-demand coordination Not just buying technology — redesigning the business Leads into process redesign

Market size: the market is already big enough

Metric Figure
China's core AI industry in 2024 Past ¥900 billion, up 24% year on year
2025 estimate ¥1.2 trillion
Number of AI companies at end-2025 Over 6,000, or 16% of the world total
Customer-facing public-cloud foundation model token volume in 2025 About 2 quadrillion

Capability jump: domestic is no longer the weak link

Dimension Change in 2025
Overall capability of leading language models Up about 30%
Multimodal understanding Up more than 50%
Model accuracy on domestic chips After joint optimisation, broadly on par with leading foreign systems
General-purpose agents vs specialised agents A well-packaged general agent can beat the frontier model itself

The number that stings most: data quality

Question Share
Content is not dense enough 82.50%
Not domain-relevant enough 14.04%
P8 Act I · Alarm

Only 6% pull back

What organisations plan to do if current AI investment falls short.

Competitors will not stop because year one did not pay.

SourceBCG AI Radar 2026(n=2,360)

Course material 3 tables · 2.2

From 2.2 Global trends at a glance

Money: investment is doubling, and not on short-term returns

Figure Meaning Source
94% Share that will keep investing even if 2026 does not deliver the expected financial return [R021]
6% Share planning to cut investment if current projects miss expectations [R021]
70% / 24% Stay the course / invest more (the other two answers to the same question) [R021]
Doubled AI investment is expected to double in 2026, rising as a share of revenue year on year [R021]

People: the CEO gets personally involved

Figure Meaning Source
72% CEOs who say they are the main AI decision-maker [R021]
60% Share of trailblazers' AI budget going to agents [R021]
83% Share of executives who say generative AI's commercial potential exceeded their expectations [R024]

Speed: one of the fastest-spreading technologies ever

Figure Meaning Source
53% Adoption within three years of generative AI reaching the mass market [R019]
Faster than the PC or the internet at the same stage [R019]
88% AI used in at least one business process at organisation level [R019]
Strong positive correlation National AI usage correlates strongly with GDP per capita [R019]
P9 Act I · Alarm

25% → 54% in six months

Share of companies with at least 40% of experiments in production.

This is the shortest window signal available to you.

SourceDeloitte, State of AI in the Enterprise 2026

Course material 2 tables · 1 recaps · 2.2

From 2.2 Global trends at a glance

Rollout: from pilot to production

Figure Meaning Source
25% → 54% Share with at least 40% of experiments in production is expected to double within six months (+29pts) [R023]
+50% Growth in the share of staff with access to AI in 2025 [R023]
34% Share genuinely redesigning the business (the rest stay surface-level or partial) [R023]
58% → 80% Share with at least limited use of physical AI (embodied, industrial) within two years [R023] Asia-Pacific leads
Money ── 94% keep investing even without a return; only 6% pull back
People ── 72% of CEOs say "I am the decision-maker"
Speed ── 53% adoption in three years, faster than the PC or the internet
Soon  ── Production deployment goes from 25% to 54% within six months

Chart suggestion

Chart Type Data
"Only 6% will pull back" Donut chart, 70/24/6 [R021]
How fast the technology spread Line chart: generative AI vs the PC vs the internet [R019]
Sprint into production Before-and-after bars, 25% → 54% [R023]

Act II

Reality

Using is not earning

P10–P17

P10 Act II · Reality

You are not behind

87% of companies are stuck exactly where you are.

Course material 1 tables · 2.1

From 2.1 Three numbers on the value gap

Three causes of the gap (the teaching spine)

Cause Performance Which section of this course
Costs are opaque Can state the total, cannot state the unit cost 4.2 Cost transparency
Processes are invisible The process you drew ≠ the process that runs 4.3 Process diagnosis
Benefits are not recognised The time saved has no financial exit 4.6 ROI calculation
P11 Act II · Reality

5% / 35% / 60%

Companies worldwide, by how much value they actually realise.

  • Future-builtKey capabilities in place; AI used for innovation and redesign
  • ScalingScaling up and starting to see value
  • LaggingSubstantial investment, almost no material value

SourceBCG research on high-performing AI companies

Course material 2 tables · 2.7

From 2.7 BCG value tiers

Three tiers (works as a single slide)

Tier Share Status BCG's term
Tier 1 5% Key capabilities in place; AI used for innovation and redesign, not just efficiency Future-built
Tier 2 35% Scaling up and starting to produce value scaling
Tier 3 60% Substantial spend, almost no real value — minimal revenue or cost benefit lagging

How big the gap is ⭐

Comparison Multiple
Revenue growth 5 ×
Cost reduction 3 ×
P12 Act II · Reality

What the top tier gets

Future-built companies against everyone else.

revenue growth
cost reduction
15%
of AI budget on agents

The gap did not open at once. It compounds: earn, reinvest in capability, earn more.

SourceBCG research on high-performing AI companies

Course material 2 tables · 2.7

From 2.7 BCG value tiers

What the top tier does differently

Action Data
Increase IT spending Planning to spend 26% more (close to a percentage point of revenue)
AI as a share of the IT budget Up to 64%
Reinvest the returns Reinvesting AI returns into talent and technical capability
Bet on agents 15% of the AI budget goes to agents; about a third already use them
Expected gap By 2028: twice the revenue growth of laggards, and 40% deeper cost reduction

These two sets are not the same thing; do not mix them:

88 / 36 / 13 5 / 35 / 60
Source Stanford + Accenture BCG
What to measure Adoption → scaling → enterprise value Three tiers of value realised
Purpose Frame the gap (where the problem is) Place yourself (which tier am I in)
Which act it belongs to ① Alarm / ② Truth ② Truth / self-assessment
  • Open with 88/36/13 to establish that using it is not earning from it
  • Use 5/35/60 in the self-assessment so they place themselves
  • ⚠️ Never plot the two sets on one chart — different bases invite challenge
P13 Act II · Reality

Which tier are you in?

No hands. Just pick one.

  1. 01

    We have AI work that already shows up in the financials

  2. 02

    We are scaling and seeing value, but it is not in the financials yet

  3. 03

    We have spent the money and cannot say what we got

If you picked the third: so did 60% of companies worldwide.

Course material 2 tables · 6 recaps · 2.7 · 5.6

From 2.7 BCG value tiers

The good news (don't leave only anxiety)

Starting point First step
Lagging tier (60%) Start with a firm leadership commitment; translate business goals into vision and strategy
Scaling tier (35%) Invest — and reinvest returns — in forward capability such as agent innovation
5%   Future-built ── 5× revenue growth · 3× cost reduction · 1/3 already using agents
35%  Scaling      ── Value starting to show · 12% using agents
60%  Lagging      ── Substantial spend, almost no real value · barely any agents

What the top tier does:
  IT spend +26% · AI up to 64% of the IT budget · returns reinvested · 15% to agents
  → By 2028: twice the revenue growth of laggards, 40% deeper cost reduction

Mechanism ── Compounding: returns → reinvest in capability → larger returns

From 5.6 Maturity self-assessment

How to use it

Item Notes
When After Act 2 Truth, before Act 3 Diagnosis
Duration 10 minutes to fill in + 10 minutes to debrief
Format On paper, handwritten, anonymous — anonymity is what gets honesty
Debrief Collect a few and read them out, without naming anyone

Dimension 1: strategy and accountability

□ 1. There is a clear AI priority order (what to do, what not to)        ___
□ 2. Every AI application has a named business owner (not IT)            ___
□ 3. The executive team explains "why we do AI" the same way             ___
□ 4. AI outcomes appear in someone's performance review                  ___
                                                  Dimension average ___

Dimension 2: cost and value

□ 1. You can state the unit cost of your main processes                  ___
□ 2. Every AI project has baseline data before approval                  ___
□ 3. At least one AI benefit has been confirmed by finance               ___
□ 4. ROI includes human review and data governance costs                 ___
                                                  Dimension average ___

Dimension 3: data foundation

□ 1. Master data is standardised (material, customer, vendor codes)      ___
□ 2. Event logs export for key processes (case ID + activity + time)     ___
□ 3. Data is classified and access is least-privilege                    ___
□ 4. You know how much complete, labelled, usable data you have          ___
                                                  Dimension average ___

Dimension 4: process and use cases

□ 1. You know the gap between the standard and the real process          ___
□ 2. You know which step waits longest and reworks most                  ___
□ 3. Five or fewer AI projects running, plus a do-not-do list            ___
□ 4. Pilots have a control group or historical baseline                  ___
                                                  Dimension average ___

Dimension 5: organisation and people

□ 1. There is a roadmap for how AI reshapes roles                        ___
□ 2. At least 25% of the AI budget goes to people and change             ___
□ 3. AI-capable key staff have a retention mechanism                     ___
□ 4. Staff know the company's position: augment or replace               ___
                                                  Dimension average ___
P14 Act II · Reality

Six-dimension maturity check

Score each 1–5. Read the shortest spoke, not the longest.

Scaling AI is a series circuit: weak data breaks scenario choice, which breaks the business case, which loses finance, which stops the rollout.

Course material 2 tables · 3 recaps · 5.6

From 5.6 Maturity self-assessment

Dimension 6: governance and risk

□ 1. Written AI usage guidance exists and has been communicated          ___
□ 2. A department or body clearly owns governance                        ___
□ 3. High-risk output requires a human signature                         ___
□ 4. AI-assisted decisions are logged and auditable                      ___
□ 5. There is a vendor exit plan (you can switch, and take the data)     ___
                                                  Dimension average ___

Scoring and placement

Total score Tier Maps to the BCG tiers [R022] Next step
24–30 Future-built About 5% globally Invest in forward capability (agents) and reinvest the returns
15–23 Scaling About 35% globally Fix the shortest dimension, then go deep on one core process
6–14 Lagging tier About 60% globally (substantial spend, almost no value) Start with a leadership commitment; do the three free things first

Radar chart (drawn live)

        Strategy &
       accountability
            5
 Governance 4    Cost &
  & risk    3     value
            2
  People    1     Data
            0
      Process & use cases

The fastest fix for each of the six dimensions

Dimension When this scores lowest, do this first Cost
Strategy and accountability Name a business owner for every running AI project — write the name down Zero
Cost and value Pick three processes and work out their unit cost Low
Data foundation Export one process's event log (three columns) and see whether you can Zero
Process and use cases List every running AI project, cut to three, and write the do-not-do list Zero
Organisation and people Say plainly which roles are being augmented, not replaced Zero
Governance and risk Publish usage guidance and name an owning department (China Mobile, Tencent and Alibaba all did) [R016] Low
Six dimensions ── strategy · cost & value · data · process · people · governance
Scoring        ── 1–5 each, sum the six averages, 30 maximum
Placement      ── 24–30 future-built (5%) / 15–23 scaling (35%) / 6–14 lagging (60%)

⭐ Core        ── Read the shortest spoke, not the longest
                  Scaling is a series circuit: break one link and everything after it fails

Fixes          ── Four of the six cost nothing and can be done this week
  • Anonymity is the point. Put names on it and everyone scores a 4.
  • When reading a few out, give the score spread only; never comment on a company.
  • If the room scores under 10 across the board, that is good — they are being honest. Say at once: "60% of companies worldwide are in this tier. You are not alone."
  • If someone scores 28, ask privately which AI benefit the CFO has signed off. It is usually an overestimate.
P15 Act II · Reality

Three causes of the gap

This slide is the map for the rest of the day.

  • Cost is opaque

    You can state the total, not the unit cost

  • Process is invisible

    The documented flow is not the flow that runs

  • Benefit is unclaimed

    Saved time has no route to the financials

Course material 1 tables · 1 recaps · 2.1

From 2.1 Three numbers on the value gap

Counter-evidence (makes the 13% look reachable)

Do not present only the gap, or the boss concludes "let's wait". Follow immediately with: the gap can be crossed, and the conditions are known.

Condition Effect Source
All five priorities in place 2.5× more likely to achieve enterprise-level results [R024]
Go deep on one core process (risk, claims, underwriting or R&D) 3× more likely to beat ROI expectations [R024] 34% of organisations already do this
Strategic investment in agent architecture 4.5× more likely to scale successfully [R024]
Reshape ways of working and talent Organisations achieving enterprise value score 88% higher here [R024]
Three bars ── 88% adopted → 36% scaled → 13% getting enterprise value
⚠️ Say it   ── The three surveys use different bases; this is not one funnel

Support    ── Only 34% are genuinely redesigning the business (Deloitte)
           ── Technology budget : people budget = 3 : 1 (Accenture)
           ── Only 35% of executives have a workforce roadmap

Causes     ── Costs are opaque · processes are invisible · benefits are not recognised

13% is not luck ──
  ×2.5  all five priorities in place    ×3    go deep on one core process
  ×4.5  strategic bet on agents         +88%  score on reshaping ways of working
P16 Act II · Reality

13% is not luck

Organisations that reach enterprise value do the same four things.

  • Act on all five imperatives → 2.5× more likely to reach enterprise-level results

  • Go deep on one core process → 3× more likely to exceed ROI expectations

  • Invest strategically in agentic architecture → 4.5× more likely to scale successfully

  • Reshape work and talent → scored 88% higher on this dimension

SourceAccenture (2,000+ generative AI projects / 3,000+ C-level executives)

Course material 1 tables · 1 recaps · 2.1 · 2.5

From 2.1 Three numbers on the value gap

Chart suggestion

  • Chart: three bars (88/36/13) falling off a cliff, joined by dashed arrows labelled "different bases".
  • Colour: 88% grey, 36% mid-grey, 13% in the accent colour (red or orange).
  • Footer — Sources: Stanford HAI, AI Index 2026; Accenture, Making Reinvention Real with Gen AI. The three surveys differ in sample and definition and do not form a single funnel. Accessed 2026-08-05.
  • Mermaid and ECharts source is in 6.2.

From 2.5 Talent and organisation

The biggest organisational gap

Gap Figure Source
AI budget spent on technology vs on people 3 : 1 [R024]
Executives with a roadmap for how AI reshapes the workforce Only 35% [R024]
How organisations achieving enterprise value score on reshaping ways of working 88% higher [R024]
Share of working hours expected to change because of generative AI 44% [R024]
Hiring   ── The most AI-exposed firms add people twice as fast as the least
Amplify  ── Top 20%: revenue per head +163% vs +33.5% for the group (5× gap)
Premium  ── 62% wage premium for AI skills (57% last year)
Two tracks ── 22% of roles professionalise (+42% pay) / 52% democratise
Newcomers ── 52% of new entry-level skill demands are senior capabilities (7% in low-exposure roles)
Global   ── Net +78 million roles; of every 100 people, 59 need training and 11 will not get it
Gap      ── Budget is 3:1 technology to people, but the gap opens on the people side
P17 Act II · Reality

The seven-step operating loop

The first diagram you take home.

  1. Cost transparency Where the money is
  2. Process diagnosis Where it jams
  3. Scenario portfolio What comes first
  4. Small-scale proof Does it work
  5. ROI audit Did it earn
  6. Ownership and governance Who is accountable
  7. Scale and replicate Will it spread

A loop, not a pipeline: step seven returns to step one, because scale changes the cost structure. Most companies break at ① and ⑤.

Course material 3 tables · 1 recaps · 4.1

From 4.1 The seven-step operating loop

The seven steps in detail

Step Core question Key action Deliverable How the block shows up
① Cost transparency Where is the money going? Build a cost tree down to process, department and unit cost Cost tree + accountability matrix Can state the total, cannot state the unit cost
② Process diagnosis Where is the process blocked? Use event logs to see the real process Target process + bottleneck map The process you drew ≠ the process that actually runs
③ Use case portfolio Which one first? Value × feasibility matrix 3–5 priority use cases Running 20 pilots at once
④ Small-scale validation Does it actually work? Prototype + pilot, with a control group Minimum viable prototype + pilot data Proving it can generate, without proving accuracy or cost
⑤ ROI audit Did it earn anything? Five KPI layers + three ROI numbers Baseline vs target vs confirmed Converting every saved hour into full salary
⑥ Organisation and governance Who owns it? Could it go wrong? Named owner + five gates Governance gates + accountability matrix Shadow AI, no human review, no exit plan
⑦ Scale and replicate Can it be rolled out? Cross-department replication + standing budget Replication playbook + budget Pilot islands, piloting forever

Which section of this library each step maps to

Step Method in detail
4.2 Cost transparency
4.3 Process diagnosis
4.4 Use case priority
4.7 Governance and risk gates (prototype gate, pilot gate)
4.5 KPI metric tree + 4.6 ROI calculation
4.7 Governance and risk gates + 2.5 Talent and organisation
4.7 Scale-up gates

Mapped to the CAICT framework (essential for SOEs and upward reporting)

The seven steps in this course The four phases in CAICT's foundation-model rollout roadmap [R018]
① ② ③ Diagnose (capability analysis, needs discovery)
Build + apply (solution design, development and testing, application build)
Apply (performance assessment)
⑥ ⑦ Manage (monitoring, operations management)
Where's the money → Where's it blocked → What first → Does it work → Did it earn → Who owns it → Can it scale
      ①                    ②                  ③            ④              ⑤            ⑥              ⑦
                                                                            ↑
                                                              Most companies die here

Act III

Diagnosis

Where money leaks and process jams

P18–P28

P18 Act III · Diagnosis Peak

What does one purchase order cost?

Everyone has an annual total. Almost no one has a unit cost.

Course material 1 tables · 4.2

From 4.2 Cost transparency

Three blind spots

Blind spot Performance Consequence
The master ledger is clear; the split is not You know IT spent ¥80 million; you do not know which system spent what No owner can be assigned
Technical resources ≠ business units Cloud bills by resource; reporting is read by department Nobody owns the waste
Cost definitions are inconsistent Multi-cloud, SaaS and on-premise each speak a different language Cannot be compared across units
P19 Act III · Diagnosis

Three blind spots

  • Clear totals, murky splits

    You know what IT spent, not what each system spent — so no one owns it

  • Technical resources are not business units

    Billed by resource, reported by department — waste has no owner

  • Inconsistent definitions

    Multi-cloud, SaaS and on-premise each speak their own language

Course material 1 tables · 4.2

From 4.2 Cost transparency

Three open-source tools (demonstrable in class, no production access needed)

Tool What it solves What to demo Licence
Microsoft FinOps Toolkit[R004] Cloud cost automation + Power BI reporting Cost trends, idle resources, committed discounts, tag coverage MIT
OpenCost[R005] Kubernetes / multi-cloud cost allocation View cost by namespace, team or application Apache 2.0 (documentation CC BY 4.0)
The FOCUS specification and sample data [R008] One cost vocabulary across cloud, SaaS and data centres Multi-cloud billing, before and after unification Per the sample repository's license.md (CC family)
Cloud Custodian[R006] Policy as code: automate idle resources and tagging Shut down non-production overnight / delete unattached disks / enforce tagging Apache 2.0
P20 Act III · Diagnosis

The cost tree

Break it down until every branch has an owner.

Total cost

  • PeopleBy process and role → labour cost per process unit
  • Cloud and computeBy application and team → unit cost per application
  • Software licencesBy utilisation → share of idle seats
  • Outsourcing and procurementBy category → unit price against market
  • Cost of qualityRework, escapes, line stoppage, claims
Course material 1 recaps · 4.2

From 4.2 Cost transparency

Cost tree (participant output 1)

Total cost
├── People        → split by process/role  → labour unit cost per process
├── Cloud & compute → split by app/team    → unit cost per application (OpenCost)
├── Software licences → split by usage     → share of idle seats
├── Outsourcing & procurement → by category → unit price vs market benchmark
└── Cost of quality → rework / missed defects / stoppages / claims  ← the big one most often missed
P21 Act III · Diagnosis

Cost of quality: the missing block

These costs sit scattered across the accounts and are never summed.

  • Rework: labour and material spent doing it twice

  • Escapes: defects that reach the customer

  • Line stoppage: lost output and delivery from unplanned downtime

  • Claims: warranty, returns and reputational loss

Leave cost of quality out and a good AI project is undervalued to the point of rejection.

Course material 1 recaps · 4.2

From 4.2 Cost transparency

Three unit costs you must ask about (take away)

□ What does one unit of business cost?  (per order / ticket / conversation / quote)
□ How does that compare with peers?     (no benchmark, no target)
□ How much of it is repetitive work?    (that is what AI can actually reach)
P22 Act III · Diagnosis

One line of spend, four prices

Use the wrong one and your savings report is fiction.

  • ListCost

    List price

    Undiscounted rate card · a negotiation baseline

  • ContractedCost

    Contracted

    Your agreed price · what procurement won

  • BilledCost

    Billed

    What this invoice says · reconciliation and payment

  • EffectiveCost

    Effective cost

    True cost after amortising prepayments and commitments

Unit economics, chargeback and ROI can only use effective cost. Allocate on billed cost and one month carries the whole prepayment while eleven ride free.

SourceFOCUS Specification(FinOps Foundation)

Course material 2 tables · 1 recaps · 4.11

From 4.11 FOCUS cost definitions

⭐ The four costs (FOCUS's core concept)

Field Chinese What it is When to use it
ListCost List cost List price, undiscounted Negotiation baseline; shows discount depth
ContractedCost Contracted cost Your negotiated price Shows what procurement negotiated
BilledCost Billed cost The number on this period's invoice Reconciliation and payment
EffectiveCost Effective cost True cost after amortising prepayments and commitments ⭐ Unit economics, departmental allocation and ROI can use only this one

What FOCUS is

Before: one field set for AWS, another for Azure, another for Alibaba Cloud, more for each SaaS
        → to answer "what did we spend on AI" you hand-align dozens of fields

Now:    export everything as FOCUS → one field set → one table shows all of it

43 standard fields (grouped by purpose)

Group Key field What question it answers
Cost BilledCost EffectiveCost ListCost ContractedCost Four prices
Owner BillingAccountId SubAccountId Tags Whose budget this belongs to
Resource ResourceId ResourceName ResourceType Exactly which machine or which service
Service ServiceName ServiceCategory ProviderName PublisherName What was bought, and from whom
Usage ConsumedQuantity ConsumedUnit PricingQuantity PricingUnit How much was used (the denominator of unit cost)
Billing basis ChargeCategory ChargeClass ChargeFrequency ChargeDescription Usage-based, subscription, or one-off
Discount CommitmentDiscountId/Name/Type/Status/Category Whether committed discounts were fully used
Unit price ListUnitPrice ContractedUnitPrice PricingCategory Is the unit price right
Time BillingPeriodStart/End ChargePeriodStart/End Which period it belongs to
Position RegionId RegionName AvailabilityZone Cost differences across regions
Spec SkuId SkuPriceId Spec
P23 Act III · Diagnosis

FOCUS: 43 standard fields

One cost vocabulary across clouds, SaaS and on-premise.

  • CostBilledCost · EffectiveCost · ListCost · ContractedCost
  • AttributionBillingAccountId · SubAccountId · Tags
  • ResourceResourceId · ResourceName · ResourceType
  • UsageConsumedQuantity · ConsumedUnit · PricingQuantity
  • DiscountsCommitmentDiscountId / Type / Status / Category

Tags decide whether you can answer "whose spend is this". The first FinOps KPI is not savings — it is tag coverage.

SourceFOCUS Specification and sample data

Course material 2 tables · 2 recaps · 4.11

From 4.11 FOCUS cost definitions

Three things you can do immediately

□ 1. Have the CIO export cloud bills as FOCUS (all major providers support it)
□ 2. Check tag coverage — how much spend can be attributed to a business line?
□ 3. Standardise all cost reporting on EffectiveCost, and write it into policy
Business line EffectiveCost Per unit of volume Unit cost Change vs prior period Tag coverage
¥ per unit % %

Classroom demo (no real invoices)

Material Content
FOCUS official sample data [R067] focus_sample.csv (small) / focus_sample_10000.csv / focus_sample_100000.csv.gz
Demo Load into Excel or Power BI → pivot by ServiceCategory and Tags → compare BilledCost with EffectiveCost
Pairs with The Power BI templates in the Microsoft FinOps Toolkit [R004] connect straight to FOCUS data
⭐ Four costs ── ListCost (list price) · ContractedCost (negotiated)
               · BilledCost (invoice) · EffectiveCost (after amortisation)
               → unit economics, allocation and ROI may use EffectiveCost only

Trap        ── Allocating on BilledCost: the prepayment month is punished, later months ride free

FOCUS       ── One cost vocabulary across cloud, SaaS and on-premise; 43 standard fields
Key field   ── Tags (they decide whose budget it is)
First KPI   ── Not money saved — tag coverage

Three moves ── Export as FOCUS · check tag coverage · standardise on EffectiveCost
P24 Act III · Diagnosis

Four open-source tools

Make the money visible. Demonstrations use simulated data only.

  • FinOps ToolkitMIT

    Cost automation and reporting: trends, idle resources, commitments, tag coverage

  • OpenCostApache 2.0

    Allocate container and multi-cloud cost by namespace, team and application

  • FOCUSCC family

    A common billing schema across clouds, with sample data

  • Cloud CustodianApache 2.0

    Policy as code: stop non-production at night, delete unattached disks, enforce tags

Course material 1 tables · 4.2

From 4.2 Cost transparency

Accountability matrix (participant output 2)

Cost pool Amount Share Owner Room to come down AI accessibility
e.g. customer-service labour ¥X,XXX,XXX XX% Head of customer service XX% High (conversations are structured)
e.g. quality-inspection labour ¥X,XXX,XXX XX% Head of quality XX% High (labelled data available)
e.g. idle cloud resources ¥X,XXX,XXX XX% Platform owner XX% No AI needed — just change the configuration
P25 Act III · Diagnosis

First save what needs no AI at all

Idle resources, duplicate licences, test environments nobody switched off. That usually funds your first AI project.

Course material 1 recaps · 4.2

From 4.2 Cost transparency

How it echoes policy and data

Problem ── You can state the total; you cannot state the unit cost
Blind spots ── Muddled allocation · resources ≠ business units · inconsistent definitions
Tools   ── FinOps Toolkit · OpenCost · FOCUS · Cloud Custodian (all open source)
Output  ── Cost tree (including cost of quality) + accountability matrix
Line    ── First save the money you can save without AI
  • Cost Sankey: total → product line → process → resource type (see 6.2)
  • FinOps dashboard screenshot: simulated data showing cost trend, idle resources and tag coverage
P26 Act III · Diagnosis Peak

The documented process is not the real one

A six-step standard process usually shows forty-plus real variants once you open the event log.

Course material 2 tables · 4.3

From 4.3 Process diagnosis

What process mining is (one sentence for a non-technical executive)

What it reveals Management implication
The actual process path How far practice diverges from policy
Variants How many non-standard paths there are, and their share
Waiting time Which step is idling, and with whom it is stuck
Rework path Which steps get sent back repeatedly
Compliance deviation Who bypassed a mandatory approval

Tool: PM4Py [R007]

Item Notes
What it is Python process-mining library
Input Event log (case ID, activity, timestamp — three columns is enough to start)
Output Process map, cycle-time distribution, variant analysis, bottleneck location
Critical path notebooks/1_event_data.ipynb、notebooks/5_advanced_examples.ipynb
Licence GPL family — usually fine for training demos; embedding in a product needs legal review
P27 Act III · Diagnosis

What process mining sees

Not interviews, not surveys, not the process manual — the data.

  • The real path

    How far execution has drifted from the policy

  • Variant spread

    How many non-standard routes exist, and their share

  • Waiting time

    Which step idles, and in whose queue

  • Rework loops

    Which stages get sent back repeatedly

  • Compliance drift

    Who bypassed a mandatory approval

Three columns are enough to start: case ID, activity, timestamp. Your ERP, OA and ticketing systems already have all three.

SourcePM4Py (GPL family — fine for demos; productisation needs legal review)

Course material 1 tables · 2 recaps · 4.3

From 4.3 Process diagnosis

Four questions to take away

□ 1. How many steps in the standard process? How many variants in reality?
□ 2. Which step waits longest? Who is it waiting on?
□ 3. What is the rework rate? Is it always the same problem?
□ 4. How many transactions bypassed a mandatory approval?

From process diagnosis to AI use cases (where this section lands)

After diagnosis, run it through three sieves:

Sieve Question What a passing use case looks like
Repetitive Is this step highly repetitive? High-frequency and patterned
Waiting What is this step waiting on? Waiting on information, approval, or a human judgement
Rework Why does this step get reworked? Incomplete information, unchecked rules, unstable quality
Highly repetitive → automation / agent execution
Long waits        → AI pre-fill / smart routing / parallelisation
High rework       → AI rule checking / right first time
P28 Act III · Diagnosis

Three sieves

From process diagnosis to AI scenarios.

  • High repetitionAutomation / agent execution
  • Long waitsPre-fill, smart routing, parallelisation
  • Heavy reworkRule checking, right-first-time

Not every bottleneck needs AI. Fix the ones an approval right or a required field would solve — so the AI project does not carry blame that is not its own.

Course material 1 recaps · 4.3

From 4.3 Process diagnosis

Participant output: target process and bottleneck map

  • The real process today (with the share of main variants)
  • The three biggest bottlenecks (waiting / rework / bypass)
  • Candidate actions per bottleneck: standardise / automate / apply AI
Problem ── The process in the policy ≠ the process in the system
Method  ── Process mining: reconstruct the real process from event logs
Bar     ── Three columns is enough: case ID + activity + timestamp
Tool    ── PM4Py (GPL — fine for demos, commercial use needs legal review)
Four Qs ── How many variants? Waiting on whom? Rework rate? Who bypassed approval?
Landing ── Repetition→automate · waiting→pre-fill/route · rework→rule checking
Rule    ── See it clearly → standardise what can be standardised → then apply AI
  • Process variant chart: the standard path vs the top five real variants, with shares
  • Waiting-time waterfall: time per step, with the longest wait marked
  • Before-and-after funnel: which step's waiting, rework or drop-off fell after AI

Act IV

Selection

Stop spreading thin

P29–P34

P29 Act IV · Selection

Twenty pilots at once is none at all

Depth wins. Breadth does not.

Course material 1 tables · 1 recaps · 4.4

From 4.4 Use case priority matrix

Matrix

High ┃  🟡 Strategic bet         🟢 Start now
     ┃  (valuable but hard,      (do this first)
Value┃   break it down first)
     ┃  🔴 Drop it               🔵 Quick win
 Low ┃  (leave it alone)         (build confidence and the team)
     ┗━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
         Low       Feasibility       High
Quadrant Strategy Suggested count
🟢 Start now (high value · high feasibility) First wave; results in 90 days 1–2
🔵 Quick win to build the team (low value · high feasibility) Use it to build confidence, train the team, prove the pipeline 1
🟡 Strategic bet (high value · low feasibility) Break it down; fix the data and process foundations first 1 (groundwork only)
🔴 Drop it (low value · low feasibility) Write it explicitly onto the do-not-do list All
P30 Act IV · Selection

Value against feasibility

Bubble size = risk or investment.

Course material 2 tables · 4.4

From 4.4 Use case priority matrix

Value dimension (1–5 each)

Dimension Question
Cost impact How large is the annual cost pool?
Revenue impact Can it shorten cycles, lift conversion, or add capacity?
Frequency How often a year? (frequency × benefit per event = real value)
Strategic weight Does it touch a core competitive activity?

Feasibility dimension (1–5 each)

Dimension Question
Data readiness Is there structured historical data? (highest weight)
Process standardisation Is the process standardised with clear boundaries?
Quality verifiable Can right and wrong be judged clearly?
Accountability Is there a clearly named business owner?
Technical difficulty Can an off-the-shelf tool cover it?
P31 Act IV · Selection

Three hard rules

Not our invention — the AI Index summary of the productivity literature.

  • Clear task boundaries

    You can state what "done correctly" looks like

  • Repeatable

    Frequent and patterned

  • Quality can be monitored

    A person or a rule can check it

All three → go. Missing one → fix it first. Missing two → do-not-do list.

SourceStanford HAI, AI Index 2026

Course material 2 tables · 1 recaps · 2.4 · 4.4

From 2.4 Productivity evidence

Macro level: a J-curve — don't rush to read the totals

Metric Figure Source
12,000 European companies AI adoption lifts labour productivity 4%, and more where training is provided Aldasoro et al., 2026[R019]
US productivity growth in 2025 2.7%, nearly double the 1.4% average of the previous decade [R019]
The OECD's forecast for the G7 +0.2 to +1.3 percentage points a year over the next decade Filippucci et al., 2025[R019]

Note the +4%, stronger where training was provided: the hardest evidence you have for a training budget (and it prices your own course nicely).

✅ Works    Service +14–15% / coding +26–55% / marketing +50% / office work +29%
❌ Counter  Senior developers −19% — and they believed they were faster
📏 Rule     Clear task boundary + repeatable + quality monitored = strongest gain
👤 Who gains  Novices > experts (AI narrows the skill gap)
📈 Macro    12,000 European firms +4% (stronger with training); US productivity 2.7% in 2025
⏳ Mindset  A J-curve: pay first, gain later — but be able to prove you are climbing

From 4.4 Use case priority matrix

Risk dimension (sets bubble size)

Dimension Question
Cost of an error What is the worst a single error can cost?
Compliance sensitivity Does it touch personal data, finance or workplace safety?
Human resistance Will this be read as layoffs?
P32 Act IV · Selection

Weight data readiness

82.50%

of industry datasets lack content density

A further 14.04% lack domain relevance.

You have a lot of data. Most of it is not useful to a model. Give data readiness 40% of the feasibility score.

SourceCAICT, AI Industry Development Research Report (2025)

Course material 2 tables · 3 recaps · 2.3 · 4.4

From 2.3 China policy and industry

Where AI lands across industry

[R015] Industrial models are deepening at both ends and breaking through in the middle:

Stage Share
Back-office operations 45.8% (highest)
Front-end R&D and design 28.3%
Production in between 25.9% (rising)

Governance: regulators are flagging risk too

Finding Figure
Some frontier models showed strategic deception in testing As high as 84%
Other risks Self-replication, refusing shutdown, chain-of-thought attacks, hallucination

Rollout methodology: CAICT's official framework (teachable as is)

Capability analysis → needs discovery → solution design → development & testing
  → application build → performance assessment → monitoring → operations management
Policy  ── Over 70% adoption by 2027 — under 17 months left
Framing ── From "experiment" to "value creation"; regulators want results
Market  ── ¥900bn core industry in 2024 (+24%), heading for ¥1.2tn in 2025
Capability ── Domestic chips are broadly on par; the technical reason to wait is disappearing
Data    ── 82.5% of industry datasets lack content density ← your real starting point
Landing ── Operations 45.8% > R&D and design 28.3% > production 25.9%
Method  ── CAICT's four phases and eight steps — the best language for reporting upward

Copyright note

Government documents and CAICT reports are natively Chinese and may be summarised, but:

  • Government infographics (such as the NDRC explainers): quote in full; never crop the source, date or emblem
  • CAICT reports are copyrighted: cite report and page when extracting data, and clear permission for full charts
  • See 7.3 Licensing red lines

From 4.4 Use case priority matrix

Three hard screening rules (from the evidence in 2.4)

✅ Clear task boundary   ← you can describe what "done right" looks like
✅ Repeatable            ← high-frequency and patterned
✅ Quality monitorable   ← a person or a rule can check it
P33 Act IV · Selection

Common scenarios, pre-judged

Start here, then argue about your own list.

Customer service assist / deflectionHighHighStart now
Inspection and defect detectionHighHighStart now
Quotation and bid documentsHighMed-highStart now
Contract and document extractionMed-highHighQuick win
Internal knowledge Q&AMediumMediumQuick win
Code assistanceMed-highHighQuick win
Demand forecasting / schedulingHighLowStrategic bet
Strategic decision supportHighLowNot yet

That last row: for work needing deep reasoning and judgement, the evidence shows gains near zero or negative. The place executives most want to use AI is the place to use it last.

Course material 1 tables · 4.4

From 4.4 Use case priority matrix

Pre-judged common use cases (to save time)

Use case Value Feasibility Recommendation
Customer-service assist / auto-response High High (historical tickets available) 🟢 Typical first choice
Quality inspection / defect detection High High (labelled data available) 🟢 First choice in manufacturing
Quote and tender document generation High Medium-high 🟢 When there are rules to check against
Extracting data from contracts and documents Medium-high High 🔵 Classic quick win
Internal knowledge Q&A (RAG) Medium Medium (depends on document quality) 🔵 Documents are not knowledge
Coding assistant Medium-high High 🔵 But don't extrapolate to R&D cost cuts
Demand forecasting / smart scheduling High Low (both data and process are barriers) 🟡 Strategic bet; fix the foundations first
Strategic decision support High Low (judgement work; measured benefit can be negative) 🔴 Leave it alone for now
P34 Act IV · Selection

The do-not-do list

Two deliverables leave the workshop: the three you will do, and everything you cut.

Scenario cut Why it was cut What would restart it

If it is not written down, it comes back next week.

Course material 1 tables · 1 recaps · 4.4

From 4.4 Use case priority matrix

Workshop run sheet (standard 75-minute version)

Time Stage Output
10 min Each group lists candidate use cases (no limit) Use case list
20 min Score on value and feasibility Scoring sheet
15 min Place on the matrix; bubble size marks risk Matrix chart
20 min Debate: cut it to 3 Priority use cases + do-not-do list
10 min Name a business owner for every use case Owner signs off
Hard rule ── Clear boundary · repeatable · verifiable (missing one? fix it first)
Quota     ── Start now: 1–2 · quick win: 1 · strategic bet: 1 (groundwork only)
Weighting ── Data readiness is 40% of feasibility (82.5% of datasets have quality problems)
Deliver   ── Priority use case list + do-not-do list + business owner's signature
Counter-intuitive ── Strategic decisions, where the boss most wants AI, are what to avoid now

Act V

Capability

What AI can actually do

P35–P43

P35 Act V · Capability

Three modes

The frame first, then the demonstrations.

  • Retrieval-augmented

    It answers, you judge · the risk is a wrong answer

  • Assistive copilot

    It suggests, you decide · the risk is that you stop judging

  • Autonomous agent

    It acts, you find out after · the risk is a wrong action you never see

Course material 1 tables · 1 recaps · 5.5

From 5.5 AI capability demo script

The six-step translation method (every demo follows it)

1. Start from a business problem   ← not from the technology
2. Pick the smallest example       ← understandable in five minutes
3. Swap in company data            ← anonymised / synthetic / public; never real sensitive data
4. Record before and after         ← baseline, result, human review rate, running cost
5. Translate into management terms ← beside the architecture write: owner, risk, cost, conditions to extend
6. Keep licences and sources       ← note the repository, version, commit date and licence in the deck

Preparation

Item Content
Data 10–50 anonymised company documents (policies, product manuals, past tickets)
Environment Rehearse it and record a screen backup — the network is the first killer
Control Prepare two questions it can answer, and one it cannot ⭐
P36 Act V · Capability

Demo 1: knowledge-base Q&A

Four frames, none optional.

  1. Input: what was asked
  2. Output: the answer, and which passage it cites
  3. Latency and tokens: this is the invoice
  4. Human check: is it right, and who decides

We deliberately ask something the knowledge base does not contain. "I don't know" is a good model. A plausible invention is a hallucination — and the reason review costs money.

Course material 1 tables · 3 recaps · 5.5

From 5.5 AI capability demo script

Four screens on stage (all four required)

① Input        ← what the user asked
② Output       ← what it answered, and which passage of which document it cited
③ Time & tokens ← this is the invoice
④ Human check  ← is it right, and who decides

Management question (ask immediately after the demo)

□ Who is accountable when it answers wrongly?
□ When it doesn't know, does it invent something?
□ When a document is updated, how long until answers reflect it?
□ Who may see which documents? Can the AI answer across permission boundaries?

Demo point: the human is driving

AI advises → a person judges → a person accepts or rejects → the person is accountable

Side-by-side on stage (the version that works best)

Without AI With AI
Time to complete Timing Timing
First-pass yield Record Record
Share needing edits Record Record
P37 Act V · Capability

Demo 2: assisted mode, timed live

The same task, once with and once without.

  1. Time to complete
  2. First-pass acceptance rate
  3. Share needing revision

Task level is not job level is not enterprise level. If a role spends 20% of its time on the task, a 29% task speed-up is under 6% at the role.

Course material 2 recaps · 5.5

From 5.5 AI capability demo script

Management question

  • Controlled experiment: about 55% faster on a coding task [R049]
  • But another study found senior developers 19% slower [R019] (with a note that it was not replicated)

Demo point: it plans and calls tools by itself

Receives a goal → breaks it into steps → calls tools → judges the result → continues or adjusts → delivers

Management question (the most important one here)

□ Can it perform write actions? (change data, place orders, send messages)
□ Do write actions need approval? Who approves?
□ How do you roll back an error?
□ If step seven of ten goes wrong, how do you find it? Is there a log?
□ Where does its autonomy end? Does it stop and ask when it reaches that edge?
P38 Act V · Capability

Demo 3: an agent on a multi-step task

It plans, calls tools and executes on its own.

  1. Takes the goal → decomposes the steps
  2. Calls tools → evaluates results
  3. Continues or adjusts → delivers

We deliberately let it get step three wrong and carry on. That is the difference from a chatbot: a chatbot gives you a wrong answer you can see; an agent takes a wrong action on your behalf.

Course material 1 tables · 5.5

From 5.5 AI capability demo script

When the demo fails ⭐

Situation Response
The network drops / the API goes down Play the recording you made earlier. Say: "I recorded this because venue networks are unreliable — which is itself a production risk worth noting."
The model gives a poor answer Don't panic — this is a gift. "Perfect. That is exactly why human review exists."
Someone claims the demo data was cherry-picked Be straight: "Yes, this is a prepared example. So don't trust my demo — run your own pilot with today's method."
A technical executive derails into detail "You know this better than I do. Let me bring it back — commercially, what this choice means is…"
Ran over time Cut the Copilot demo (the second) and keep RAG and the agent — those two differ most in management implication
P39 Act V · Capability

Risk escalates with capability

And gets harder to detect at every step.

  1. RetrievalA wrong answer
  2. CopilotYou stop judging
  3. AgentA wrong action you never see

Stronger capability demands stronger governance. That is not caution — it is what lets you move. The rule is simple: read-only can run free; anything that writes needs approval and an audit trail.

Course material 2 recaps · 5.5

From 5.5 AI capability demo script

Demo discipline (red lines)

🔴 Never upload real customer data, contracts or personal information to a public AI service
🔴 Never let a key, internal address or real client name appear in a demo
🔴 Never promise "we can achieve the same result"
✅ Always use anonymised, synthetic or public sample data
✅ Ask a management question straight after each demo to pull it back to the business
✅ Note in the deck: repository, version, commit, licence
Three demos ── RAG (it answers, you judge) · Copilot (it suggests, you decide) · agent (it acts for you)
Risk line   ── Wrong answer → you stop judging → it acts wrongly and you don't know (rising, and harder to spot)

Key design:
  RAG      Deliberately ask one it cannot answer → demonstrate hallucination
  Copilot  Time it live, side by side → task level ≠ role level
  Agent    Deliberately show a failure → explain write-action approval

Rule     ── Read-only is open; write actions need approval and a log
Backup   ── Record it beforehand; a bad answer is a gift; if challenged, be straight
P40 Act V · Capability

Agents are becoming the divide

Share of companies already using agents.

Leaders put 60% of the AI budget into agents. Agents are forecast to move from 17% of total AI value in 2025 to 29% by 2028.

SourceBCG AI Radar 2026; BCG research on high-performing AI companies

Course material 5 tables · 1 recaps · 2.6

From 2.6 Agents and the next wave

This is where the money is going

Figure Meaning Source
60% Share of trailblazers' AI budget going to agents R021 BCG
15% Share of AI budget future-built companies allocate to agents R022 BCG
17% → 29% Agents as a share of total AI value: 2025 → 2028 R022 BCG
× 3 Organisations planning to invest in agent architecture, vs last year R024 Accenture
× 4.5 Organisations that scaled generative AI are more likely to have invested strategically in agent architecture R024

Who is using it, and how big the gap is

Group Share using agents Source
Future-built companies (5%) About 1/3 R022 BCG
Companies that are scaling (35%) 12% R022
Lagging companies (60%) Almost none R022

Evidence from China

Finding Data Source
General-purpose agents vs frontier models A well-packaged general agent can outperform the frontier model itself R015 CAICT "Fangsheng" benchmark
Mechanism A planning engine plus a tool-calling framework closes the sense-decide-act loop R015
Position Agents become proto-"digital employees"; execution shifts from fixed flows to dynamic optimisation R015
Scale Customer-facing public-cloud model token volume in 2025: about 2 quadrillion R015

Embodied / physical AI (essential for manufacturing clients)

Figure Meaning Source
58% → 80% Share with at least limited use of physical AI within two years; Asia-Pacific leads R023 Deloitte
Phase Embodied AI is at the pivot from lab validation to commercial scale R015
Bottleneck Scarce high-quality data, weak cross-domain generalisation, hardware-software stability R015

Management implication: agents raise three new questions

New problem Why it did not exist before How to handle it
Write-action authorisation AI used to only produce text; now it changes data, places orders and sends email Write actions need approval; separate preview, approve and execute
Attributing multi-step failures Step seven of ten went wrong — how do you find it Log the whole chain; every step traceable
Boundary of responsibility It decided on its own — who is liable? Define the autonomy boundary; anything beyond escalates to a human
Trend    ── AI moves from "can think" to "can do" (CAICT's phrasing)
Money    ── Leading organisations put 60% of the AI budget into agents
Value    ── Agents go from 17% of total AI value (2025) to 29% (2028)
Divide   ── Usage: future-built 33% / scaling 12% / lagging ≈0
Counter-intuitive ── A well-packaged general agent can beat the frontier model itself
Physical AI ── 58% → 80% within two years; Asia-Pacific leads

Management ── AI used to give you a wrong answer; now it does the wrong thing for you
Rule       ── Read-only is open; write actions need approval and a log
P41 Act V · Capability

Four forks in the road

You do not need the engineering. You do need to settle these four.

  1. Open or closed weightsSets your team cost and your exit cost
  2. Retrieval or fine-tuning90% of companies should start with retrieval
  3. Public, private or hybridSets data exposure and up-front investment
  4. Large or small modelThe easiest saving available — an order of magnitude

In financial services the default is local deployment: CAICT records that most banks run large models on their own infrastructure.

SourceCAICT, Foundation Model Deployment Roadmap

Course material 3 tables · 2 recaps · 4.9

From 4.9 Technology selection decision table

Decision 1: open-weight or closed

Approach Characteristics (CAICT wording) Business language
Open-weight model Lower cost and faster development; suits research, individual work, fast validation and sharing Cheap, fast, controllable — but you must staff it yourself
Closed model Meets customisation and security needs; suits high-security, highly bespoke, commercially sensitive settings Easy and supported, but expensive and prone to lock-in
□ Do we have a team that can maintain an open-weight model? (No → go closed)
□ Can data leave the company? (No → open-weight, self-hosted)
□ What would switching vendors cost in three years? (A lot → prefer open-weight)

Decision 2: how to make the model understand your business (pick one or combine)

Approach Characteristic Cost When to use it
Retrieval-augmented generation (RAG) Supports domain Q&A; mitigates hallucination to a degree and adds domain depth Low Where most companies start
Full fine-tuning Fits the dataset well and learns strongly, but trains inefficiently High When you have plenty of tuning data
Efficient fine-tuning Updates fewer parameters: more efficient, less compute, less training time Medium Data available, compute limited
Instruction tuning Better intent understanding and answer alignment Medium When the interaction is not good enough
Prompt tuning Steering the model with a specific prompt to produce closely related output Very low Fast validation
□ Start with RAG, run it three months, then see what is missing  ← where 90% should begin
□ If the gap is jargon            → instruction tuning
□ If the gap is missing knowledge → extend the knowledge base, not fine-tuning
□ If the gap is an unusual task pattern → only then consider fine-tuning

Decision 3: where to deploy

Approach Characteristic Business meaning
Public cloud Scales on demand; lower build and run cost; no infrastructure to own Fast to start, but data leaves the company and long-run cost may be higher
Private cloud Cuts sensitive-data exposure; more flexible to operate; reuses existing hardware and software Secure, but heavy upfront investment
Hybrid cloud Combines the strengths of both, flexing to demand spikes and business change The realistic choice for most mid-to-large companies
P42 Act V · Capability

Retrieval mitigates hallucination "to a degree"

Retrieval-augmented generation supports domain-specific question answering and can, to a degree, mitigate model hallucination while improving domain accuracy.

CAICT, Large Model Deployment Roadmap

Note "to a degree". Vendors will tell you retrieval ends hallucination. The official wording does not. Human review never leaves the budget.

Course material 2 tables · 1 recaps · 4.9

From 4.9 Technology selection decision table

Decision 4: how large a model

Approach Characteristic Business meaning
¥10bn revenue and above For complex tasks demanding accuracy in generation, understanding, reasoning and decisions; heavy compute in training and inference Expensive, but the only option for hard tasks
¥1bn revenue and below Suits simple tasks; low compute needs; deployable on edge and on-device Cheap; suits high-frequency simple tasks and field devices

Recommended combinations of the four decisions, by company type

Company type Model Optimisation Deployment Scale
Large state-owned / financial Open-weight + self-hosted Start with RAG Private cloud / on-premise Hybrid (large model for hard tasks, small for high volume)
Mid-size manufacturer Start with a closed-model API RAG Public cloud → hybrid Mostly small models
Technology / internet Mostly open-weight RAG + efficient fine-tuning Hybrid cloud Tiered by use case
Small company Closed-model API Prompting + RAG Public cloud Small model
Four forks (the boss decides these):
① Open-weight or closed   → sets team commitment and exit difficulty
② RAG or fine-tuning      → 90% of companies should do RAG first
③ Public / private / hybrid → sets data security and upfront cost
④ Large model or small     → the easiest decision to save money on

⚠️ The official wording: RAG mitigates hallucination only "to a certain extent"
   → human review cost can never be deleted from the budget
P43 Act V · Capability

AI spend sits in five layers, plus a sixth line

The intelligent-compute cloud stack.

  1. AISaaSApplication services · finished business applications
  2. MaaSModel as a service · the model and its lifecycle
  3. AIPaaSPlatform services · build, train and deploy toolchain
  4. AIIaaSInfrastructure services · compute, network, storage
  5. AIMSPManaged services · end-to-end delivery
  6. Human reviewAbsent from the technical stack, often the largest running cost

A budget with only the first five lines is not a budget.

SourceCAICT, Cloud Computing Blue Book (2025)

Course material 4 tables · 2 recaps · 4.12

From 4.12 The five-layer AI cloud architecture

The five layers (CAICT's original naming)

Layer Full name In one sentence Whose budget Build or buy
AIIaaS AI cloud infrastructure services Compute, network, storage — AI's foundation CIO / infrastructure Buy to start; consider building once you are at scale
AIPaaS AI cloud platform services Toolchain for development, training and deployment CIO / R&D Most companies should buy
MaaS Foundation model service The model itself, served end to end CIO / business Buy + light fine-tuning
AISaaS AI cloud application services Finished applications for business use cases Business unit Buy off the shelf first
AIMSP AI cloud managed services End-to-end professional delivery services CIO / procurement Buy (but guard against lock-in)
  • AIIaaS: reshapes intelligent infrastructure for efficient training and inference
  • AIPaaS: reshapes AI development services and accelerates AI+ applications at scale
  • MaaS: end-to-end service capability, the main vehicle for putting models into production
  • AISaaS: brings intelligence to enterprise applications and lifts operating efficiency
  • AIMSP: professional AI cloud services connecting end-to-end model application delivery

⭐ Break the budget down these five layers (the core tool here)

Layer Annual budget Share Build / buy Vendor Cost of switching away
AIIaaS compute infrastructure
AIPaaS platform toolchain
MaaS model services
AISaaS business applications
AIMSP delivery services
Human review (outside the five layers)

Three decision questions (for the CIO to settle)

Question Basis for the call
Which layer to build yourself? Build only when you are large enough and that layer is your differentiator. For most companies, only the data and use cases are yours; buy the rest.
Which layer locks you in fastest? AIPaaS and AIMSP. Once the toolchain and delivery services are entangled, switching vendors is very costly
Which layer is easiest to save on? MaaS — tier large and small models by use case; see 4.9 Technology selection
□ AIPaaS layer: prefer platforms on open standards where the model is swappable
□ MaaS layer: connect at least two model vendors and keep the ability to switch
□ AIMSP layer: write data and knowledge-base export terms into the contract

How it echoes the market trend

Finding Data Source
AI cloud is becoming vendors' way through China's cloud market keeps growing fast R017
"Cloud + AI" as twin engines Industry AI adoption is accelerating R017
Scale of compute investment Gigawatt-scale AI compute clusters are being built out R015
Domestic chips With joint hardware-software optimisation, accuracy is broadly on par with leading foreign systems R015
Token volume Customer-facing public-cloud foundation model calls in 2025: about 2 quadrillion R015
Five layers ── AIIaaS infrastructure → AIPaaS platform → MaaS models → AISaaS applications → AIMSP delivery
              (CAICT Cloud Computing Blue Book 2025, AI cloud architecture)

⭐ One more row ── Human review (outside the technical architecture, often the largest running cost)

Build rule  ── Build only when you are large enough and it is your differentiator
               For most companies only the data and use cases are yours; buy the rest

Anti-lock-in ── AIPaaS and AIMSP lock you in fastest
               → choose open standards · connect two model vendors · put export terms in the contract

Biggest saving ── The MaaS layer: tier large and small models by use case

Act VI

Evidence

Who has actually done it

P44–P50

P44 Act VI · Evidence

Four questions for any case

Every case, the same four steps.

  • 1

    What was stuck?

    The baseline. A case without one is not worth telling

  • 2

    What changed?

    The process, or only the tool

  • 3

    Where did the number come from?

    Who measured it, on what definition, at what grade

  • 4

    What would it take here?

    Land it on the audience

Course material 3 tables · 1 recaps · 3.1

From 3.1 Case quick-reference

Full table

# Case The result in one line Industry Who to tell Evidence
1 Klarna AI customer service [R048] About two-thirds of conversations handled by AI, roughly 700 FTE equivalent; resolution time 11 min → under 2 min; SEC filing shows about US$39m saved in 2024 Finance / retail services CEO · CFO · customer service A-
2 The GitHub Copilot controlled experiment [R049] About 55% faster on a specified coding task Technology / R&D CIO·CTO A-
3 Microsoft 365 Copilot[R050] About 29% average speed-up across several tasks, with about 3,000 participants in real conditions General office work CHRO·COO A-
4 BBVA enterprise-wide adoption [R051] About 100,000 staff covered, over 70% weekly active, about 3 hours saved per person per week, up to 80% efficiency gain on selected processes Banking CEO·CHRO·CIO B
5 The Fubusi textile platform [R052] 90% uptime; production cost down 15%–50%; design cost down to 20% of the original Textiles / light industry CEO·COO B
6 AI non-destructive testing of high-voltage components [R053] Manual inspection cost down 85%+, total production cost down 40%, 3× inspection throughput, R&D cycle down 28% Electronic components / manufacturing COO · quality · CFO B
7 Glodon construction cost automation [R054] 98%+ accuracy, 50–70% faster preparation, 3–5 days → about 4 hours, change risk down 30% Construction / engineering COO · professional services C+
8 Digiwin formulation and quoting agent [R055] Formulation cost down 15%, efficiency +8%; quoting labour down 80%, quote margin attainment +15% Discrete manufacturing COO · sales · CFO C
9 State Grid procurement digitalisation [R056] Annual bidding cost down 70%; cumulative inventory reduction of about ¥8 billion Energy / state-owned CEO · procurement · CFO B
10 Xinjiang Energy field automation [R057] On-site headcount down 75%+; smart devices replace about 80% of manual inspection Energy / heavy industry COO · workplace safety B

Indexed by industry

Client industry Preferred case Alternative
Manufacturing / industry 6 Capacitor inspection, 5 Textiles 8 Digiwin, 10 Xinjiang Energy
Energy / state-owned 9 State Grid, 10 Xinjiang Energy 6 Capacitor inspection
Finance / banking 4 BBVA、1 Klarna 3 M365
Retail / consumer services 1 Klarna 3 M365
Construction / engineering 7 Glodon 6 Capacitor inspection
Technology / software 2 Copilot experiment 3 M365
Cross-industry / mixed executive cohort 1 + 6 + 9 4 BBVA

Indexed by management topic

The point you want to make Which case to use
AI can actually reach the income statement 1 Klarna (backed by an SEC disclosure)
Separate task-level speed-up from role-level 2 Copilot + 3 M365
Scaling depends on the organisation, not the technology 4 BBVA (100,000 staff, 70% weekly active)
Cost of quality is worth more than cost of labour 6 Capacitor inspection (missed defects, line stoppages and R&D cycle counted together)
Digitise first, then add intelligence 9 State Grid (no standardisation, no AI)
Discount vendor case studies 8 Digiwin (Grade C — treat as hypothesis only)
Read headcount cuts alongside safety and uptime 10 Xinjiang Energy

The four-step formula for presenting a case

Use these four steps for every case. Participants remember the formula — worth more than the case itself.

1. What was it stuck on?      ← the baseline. A case without one is not worth telling
2. What did it change?        ← the process, or only the tool?
3. Where do the numbers come from? ← who measured, on what basis, grade A, B or C
4. What would it take for you? ← bring it back to the audience

⚠️ Three rules for presenting a case

  • Talk business results, not headcount replaced. Follow Klarna's "700 FTE equivalent" immediately with resolution time and repeat-contact rate — headcount alone invites "and what happened to service?"
  • Explain the range. The textile case's 15%–50% is wide; say it reflects different companies' baselines, not one company fluctuating.
  • Treat Grade C cases as hypotheses. Glodon (C+) and Digiwin (C) come from media and vendors — say plainly that these are hypotheses, not commitments.

See 3.4 Evidence grading and the debunking checklist.

P45 Act VI · Evidence

Klarna: the one backed by a regulatory filing

A-Fintech / customer service

~2/3
of service chats handled by AI
≈700
full-time equivalents
11 min → <2 min
average resolution time
−25%
repeat enquiries
~USD 39m
2024 cost saving (regulatory disclosure)

A press release is what a company wants you to see. A regulatory filing is what it is legally answerable for. That is the difference between grade A and grade B.

Question to the roomDid those 700 FTE-equivalents become reduced hiring, more volume handled, or simply a less busy team? Three different answers, three different financial meanings.

SourceKlarna press releases and SEC disclosures

Course material 3 tables · 3.2

From 3.2 International benchmark cards

Figure

Metric Result
Share of service conversations handled by AI About 2/3
Equivalent headcount About 700 full-time equivalents of work
Average resolution time 11 minutes → under 2 minutes
Repeat enquiries −25%
Financial impact About US$39 million saved in 2024 (per SEC disclosure)

Watch out for

  • ⚠️ Never say only "replaced 700 people" — always give resolution time, repeat contacts and satisfaction too
  • ⚠️ Klarna later adjusted its human-service strategy. If asked, answer: "That is precisely the point — AI service is not a one-off replacement but a continuously tuned human-machine mix."

Figure

Metric Result
Speed on a specified coding task About 55% faster
Supporting evidence (a separate study, Cui et al. 2025) Pull requests +26%

Figure

Metric Result
Average speed-up across several task experiments About 29%
Sample Multiple clients, about 3,000 participants in real working conditions

Why this card is worth having

Large sample, measured in real workplaces rather than a lab; fits admin, sales, finance and knowledge work.

P46 Act VI · Evidence

AI inspection for high-voltage components: five pockets

BElectronic components

−85%+
manual inspection cost
−40%
total production cost
inspection throughput per unit
−28%
product development cycle
98%
domestically sourced hardware and software

Most executives count only the inspectors saved. The return here comes from five pockets: labour, escaped defects, line-stoppage risk, throughput and development cycle. Count only the first and the project never clears approval.

Question to the roomWhy would an inspection system shorten development by 28%? Because testing during trial production got faster, so the iteration loop turned faster. AI returns often appear where you did not look — provided your metrics can see there.

SourceMIIT published case collection

Course material 3 tables · 3.3

From 3.3 China practice cards

The figures (five dimensions move together — that is what makes it valuable)

Dimension Result
Quality up Inspection accuracy up markedly; defect rate down
Cost down Manual inspection cost down 85%+; total production cost down 40%
Efficiency up 3× faster per-component inspection; shorter changeovers; higher equipment utilisation
Innovation Product development cycle down 28%
Sovereign and controllable 98% domestically sourced hardware and software; ERP and MES data connected
Business result 2024 revenue ¥148.297m, profit ¥7.405m

Why this card is the trump

  • Real financial figures — revenue and profit to the decimal — not a vague "XX% better"
  • Five dimensions improve at once — a perfect illustration of the cost-of-quality model
  • 98% domestically sourced — decisive for SOE, defence and supply-chain-sensitive clients
  • Connected to ERP and MES — proof it is not an isolated pilot

Figure

Metric Result
Equipment uptime Reaches 90%
Production cost Down 15%–50%
Design cost Down to 20% of the original (i.e. −80%)
Platform deployment cost ¥1.5m–2m (priced by module)

Figure

Metric Result
Annual bidding cost Down 70%
Cumulative inventory reduction About ¥8 billion
Supporting scale 500 million smart meters, over 19,000 servers, nearly 30 PB of storage, 900+ shared enterprise services
Acceptance Promoted nationally by MOFCOM as an innovation case
P47 Act VI · Evidence

State Grid: not an AI case at all

BEnergy / state-owned enterprise

−70%
annual tendering cost
~RMB 8bn
cumulative inventory reduction
900+
enterprise shared services
500m
smart meters connected

This case is here precisely because it is not about models. The numbers came from data classification, single-entry shared data and process standardisation. Without that groundwork AI does not get off the ground — CAICT found 82.5% of industry datasets lack content density.

Question to the roomThe order does not reverse: digitise, then add intelligence. Skip the first step and you will do it later anyway, at greater cost.

SourceState Grid published case; MOFCOM innovation case

Course material 1 tables · 1 recaps · 3.3

From 3.3 China practice cards

Additional cases (optional, use as needed)

Case In one sentence Purpose Evidence
Glodon construction cost estimating [R054] 98%+ accuracy, 50–70% faster preparation, 3–5 days → about 4 hours, change risk −30% For professional knowledge work, AI's value is not only drafting: it is rule checking, knowledge reuse and lower change risk C+
Digiwin formulation and quoting agent [R055] Formulation cost −15%, efficiency +8%; quoting labour −80%, quote margin attainment +15% Shows AI affecting cost, speed and margin quality at once C
Xinjiang Energy field automation [R057] On-site headcount −75%+; smart devices replace about 80% of manual inspection Covers workplace safety, remote operations, human-machine teaming B

How to combine the China cases

State Grid (7) → Textiles (6) → Capacitor inspection (5)
 digitise first   a price you can       all five dimensions
                  actually account for   connected
P48 Act VI · Evidence

BBVA: a four-rung ladder

A scaled deployment across roughly 100,000 employees.

100,000 × 3 hours × 50 weeks = 15 million hours. The case stops at rung three — not because they are hiding it, but because that jump is the hard one.

SourceBBVA customer case disclosure

Course material 1 tables · 2 recaps · 3.2

From 3.2 International benchmark cards

Figure

Metric Result
Employees covered About 100,000 people
Weekly active rate >70%
Saved per person per week About 3 hours
Best efficiency gain on a selected process 80%

How to teach it: the four-rung metric ladder (this card's core value)

Adoption ── 100,000 people got the tool
   ↓
Frequency ── 70% genuinely use it weekly   ← most companies die here
   ↓
Time     ── About 3 hours saved per person per week
   ↓
Business result ── ???                     ← the case does not disclose this

Watch out for

  • ⚠️ Grade B: a customer case page; the figures are self-reported
  • ⚠️ The "80% efficiency gain" is the best selected process, not an average, and is not a company-wide figure

How to combine the three cards (recommended order)

Copilot (2) → M365 (3) → BBVA (4) → Klarna (1)
 task-level    real working  organisation  confirmed by
 experiment    conditions    scale         finance
P49 Act VI · Evidence

Three grades of evidence

Every figure belongs to one of them.

  • A

    Regulatory filings, financial statements, randomised or quasi-experiments

    Strong enough for a business case and a board resolution

  • B

    Company disclosures, government case collections, joint customer cases

    Strong enough for an implementation hypothesis

  • C

    Media, vendor and partner accounts

    Suggestive only — verify before acting

Grade A goes in the board minutes. Grade B goes in the business case. Grade C goes on the whiteboard.

Course material 2 tables · 1 recaps · 3.4

From 3.4 Evidence grading and the debunking checklist

The three-grade evidence standard (the core tool taught here)

Grade What it is Example What it can be used for
Grade A Regulatory filings, financial reports, randomised trials, verifiable formal research Klarna's SEC disclosure, the Copilot controlled experiment, Stanford's AI Index Fit for a business case
Grade B Official company disclosures and joint customer-vendor cases; specific but mostly self-reported BBVA, the MIIT case collection, textiles and capacitor inspection Enough to form an implementation hypothesis
Grade C Retold by media, vendors or ecosystem partners Glodon, Digiwin Suggestive only; must be independently verified

Four figure labels (finer than the three grades; for reported data)

Every figure in this deck belongs to one of these:

Tag Meaning Principle of use
Market statistics Adoption rates, investment, job trends Describes the external environment; never a promise of your own returns
Survey self-report Benefits and maturity self-reported by staff or executives Indicates direction; flag sample and self-report bias
Project observation Client projects written up by consultancies or vendors Enough for an implementation hypothesis, not an audit
Controlled-experiment evidence Time and quality differences in randomised or quasi-experiments Useful for sizing potential; still needs internal validation

Debunking checklist: eight questions a CEO can take away

□ 1. What is the baseline?   Without a "before" number, no percentage means anything
□ 2. Who measured it?        Vendor / customer / third party — three grades apart
□ 3. Was there a control?    Without one you cannot tell whether AI did it
□ 4. How big, how long?      Three people for a week is not 3,000 people for a year
□ 5. Only keen users?        Interviewing volunteers filters out the failures automatically
□ 6. Was quality measured?   Halve any number that reports speed but not quality
□ 7. Are all costs in?       Model, integration, data governance, human review, operations — miss one and the books are false
□ 8. Who signed it off?      A benefit without a CFO signature is still just a slide
P50 Act VI · Evidence

Eight questions for any claim

Ask a vendor these in order. This is the first card you take home.

  • What is the baseline? Without a "before", no percentage means anything

  • Who measured it? Vendor, customer or third party — three different grades

  • Was there a control group? Without one you cannot attribute the change to AI

  • How large a sample, over how long? Three people for a week is not three thousand for a year

  • Were only enthusiasts included? Interviewing volunteers filters out the failures

  • Was quality measured? Halve any figure that reports speed and not quality

  • Is the full cost counted? Model, integration, data governance, human review, operations

  • Who signs it off? A benefit without a finance signature is still a slide

Course material 2 tables · 1 recaps · 3.4

From 3.4 Evidence grading and the debunking checklist

The five most common number traps (each gets a laugh)

Trap What it looks like How to counter it
Extrapolating from a single task "Coding is 55% faster → R&D costs fall 55%" Writing code is one slice of the R&D cycle. Ask: what share of the role's time is it?
Hourly-rate conversion "3 hours saved per person per week × hourly rate × headcount = benefit" Time saved with no exit is not money. See the five exit conditions in 4.6
Reporting only the best number "Efficiency gains of up to 80%" "Up to" means the best one they found. Ask for the mean and the median
Survivorship bias "Everyone who uses it likes it" Why do non-users not use it? What is the drop-off?
Only average accuracy was measured "98% accurate" Ask for the error rate on high-loss edge cases. 98% accuracy is a disaster if the 2% are major incidents

The figures in this course face the same test

Check Status
Every figure carries a resource ID and access date See 7.2 Number provenance table
The three bars (88/36/13) carry a note that the bases differ and do not form one funnel ✅ Required
The McKinsey 39% EBIT figure could not be verified locally and is marked as such ⚠️ See 2.1
The METR −19% figure carries a note that it was not subsequently replicated ✅ Required
Ranges (e.g. textiles 15%–50%) come with an explanation of baseline differences ✅ Required
Grade C cases are marked "hypothesis, not commitment" ✅ Required
Grade A = regulator / financial report / experiment → board resolutions
Grade B = company self-report                      → project approval papers
Grade C = media / vendor retelling                 → brainstorming only

Eight questions: baseline? who measured? control? sample? cherry-picked? quality? all costs? who signed?

Five traps: extrapolating from one task · salary conversion · reporting only the best · survivorship bias · average accuracy only

Act VII

Accounting

Saved hours are not profit

P51–P58

P51 Act VII · Accounting

A very popular piece of arithmetic

We receive this calculation regularly.

  • 1,000 staff × 3 hours saved per week × 50 weeks = 150,000 hours
  • 150,000 hours × RMB 100 per hour = RMB 15m saved per year
Course material 2 recaps · 4.6

From 4.6 ROI calculation template

Five exit conditions (the most valuable page in this course)

For saved time to become a financial benefit, at least one must hold:

□ 1. Less overtime, outsourcing or new hiring
      → direct cost reduction; the hardest evidence

□ 2. More orders, cases or customers handled by the same people
      → capacity released; you need incremental demand to absorb it

□ 3. Shorter time to market, quoting, collection or development, producing incremental revenue
      → revenue-side benefit; usually the largest and the hardest to prove

□ 4. Capacity redeployed onto clearly measurable high-value work
      → you must say where it went and how it is measured, or it does not count

□ 5. Fewer errors, rework, incidents, compliance failures or lost customers
      → cost of quality; the most consistently underestimated

The ROI formula

        annualised realised benefit − annualised total cost
ROI =  ───────────────────────────────────────────────────
                     annualised total cost
P52 Act VII · Accounting Peak

Did your payroll actually fall by RMB 15m?

It did not. Nobody left and no salary changed. That 15 million is in the air.

Course material 1 tables · 2 recaps · 4.6

From 4.6 ROI calculation template

Numerator: annualised realised benefit

Capacity released
+ external spend avoided
+ incremental profit
+ lower working capital
+ risk losses avoided

Denominator: annualised total cost (the half most often missed)

Models and compute
+ software licences
+ data governance      ← often missed
+ integration
+ security and compliance ← often missed
+ change and training  ← often missed
+ human review         ← missed most of all!
+ ongoing operations

Three ROI numbers (every pilot needs all three)

Figure Who calculates Purpose Does the executive team read it
① Theoretical potential Project team Sparks imagination at approval ❌ Not a basis for decisions
② Risk-discounted business case Project team + finance business partner Project approval review ✅ Primary basis
③ Finance-confirmed benefit CFO / finance business partner Close-out and scale-up decision ✅ Final basis
P53 Act VII · Accounting

Five exit conditions

For saved time to become money, at least one must hold. This is the second card you take home.

  • Less overtime, outsourcing or new hiring — the hardest and most direct

  • More orders, cases or customers handled by the same people — capacity released, provided demand exists to absorb it

  • Shorter time to market, quotation, collection or development, producing incremental revenue — usually the largest and the hardest to prove

  • Capacity redeployed to a defined, measurable higher-value task — you must name the task and the measure

  • Fewer errors, less rework, fewer incidents, compliance or churn losses — cost of quality, routinely underestimated

If none applies, the problem is not the project — the benefit route has not been designed yet.

Course material 1 tables · 4.6

From 4.6 ROI calculation template

Suggested benefit discount (how to apply the risk haircut)

Situation Suggested discount
Benefit comes from a Grade A controlled experiment, validated internally Discount to 80–90%
Benefit comes from Grade B self-reported cases Discount to 50–60%
Benefit comes from Grade C media or vendor cases Discount below 30%, or exclude entirely
The benefit is time saved, but no exit condition is confirmed Scores 0
The benefit is on the revenue side (shorter cycles bring incremental revenue) Discount to 30–50% (hardest to prove)
P54 Act VII · Accounting

Usage is not value

  • 1.2 million calls this month
  • 3,000 active users
  • 80,000 prompts submitted

Which of those three appears on the income statement? None. These are vanity metrics.

Course material 1 tables · 1 recaps · 4.5

From 4.5 The five-layer KPI tree

The five-layer metric tree

┌──────────────────────────┐
        │  ① Financial results      │  ← the board reads only this layer
        ├──────────────────────────┤
        │  ② Operational results    │
        ├──────────────────────────┤
        │  ③ Customer results       │
        ├──────────────────────────┤
        │  ④ People results         │
        ├──────────────────────────┤
        │  ⑤ Risk results           │  ← the only layer read when something goes wrong
        └──────────────────────────┘
   (Call volume and active users sit outside the tree — that is health monitoring, not value)
Layer Core metric Common mistake
① Financial results EBIT impact, cost avoided, incremental profit, cash flow, capital tied up, payback Converting all saved time into full salary value
② Operational results Cycle time, throughput, first-pass yield, rework rate, utilisation, unit cost Watching speed while ignoring quality and demand shifts
③ Customer results Response time, first-contact resolution, conversion, retention, NPS, complaints Treating chatbot volume as satisfaction
④ People results Weekly active rate, proficient users, task coverage, skill transfer, span of control Mandated use produces fake activity
⑤ Risk results Error rate, hallucination rate, data exposure, human review rate, audit findings Measuring average accuracy but not high-loss edge cases
P55 Act VII · Accounting

A five-layer indicator tree

Five layers separate a vanity metric from EBIT. None can be skipped.

  1. Financial

    EBIT impact, cost avoidance, incremental profit, cash flow, capital tied up, payback

  2. Operational

    Cycle time, throughput, first-pass yield, rework, utilisation, unit cost

  3. Customer

    Response time, first-contact resolution, conversion, retention, NPS, complaints

  4. People

    Weekly active use, proficient-user share, task coverage, skill transfer

  5. Risk

    Error rate, hallucination rate, data exposure, review rate, audit findings

Call volume and active users sit outside the tree — health monitoring, not value.

Course material 1 tables · 1 recaps · 4.5

From 4.5 The five-layer KPI tree

The pairing rule for each layer (the heart of this section)

You report this Must be reported alongside Why
Processing speed ↑ First-pass yield / rework rate Fast and wrong costs more
Labour saved ↑ Service level / customer satisfaction Cutting headcount into churn is a loss
Automation rate ↑ Human review rate / escalation rate Fully automated but escalating daily is fake automation
Active usage ↑ Task coverage / share of proficient users Logging in daily without working is fake activity
Accuracy ↑ Error rate on high-loss edge cases 98% accuracy is a disaster if the 2% are major incidents
On-site headcount ↓ Safety incident rate / equipment availability Cutting headcount into downtime is a loss (see the Xinjiang Energy case [R057])

The BBVA four-rung ladder (why you cannot skip a rung)

Adoption  ── 100,000 staff got the tool
   ↓  (most companies start boasting here)
Frequency ── 70% weekly active
   ↓  (most companies die here)
Time      ── About 3 hours saved per person per week
   ↓  ← ⚠️ the hardest jump
Business result ── ?
P56 Act VII · Accounting

Every gain needs its counterweight

No positive metric is reported without its paired negative.

Speed ↑First-pass yield / reworkFast and wrong costs more
Headcount saved ↑Service level / satisfactionSavings that cost customers are losses
Automation rate ↑Human review / escalation rateFully automated but always escalating is not automation
Active usage ↑Task coverage / proficient usersLogging in is not working
Accuracy ↑Error rate on high-loss edge cases98% accuracy is a disaster in the 2% that matter
On-site headcount ↓Safety incidents / availabilityFewer people and more downtime is a loss

When someone reports the left column without the right, send it back. A one-sided metric will be optimised until it is meaningless.

Course material 1 tables · 1 recaps · 4.5

From 4.5 The five-layer KPI tree

Metric tree template (participant output)

Layer Metric Baseline (before) Target Measured in the pilot Confirmed by finance
① Financial CFO signs off
② Operational
③ Customer
④ People
⑤ Risk
  • If the baseline column is blank, it does not enter ROI review.
  • At least one metric per layer, plus a paired counter-metric.
  • Only the CFO or a finance business partner may fill the last column.
Outside the tree ── call volume and active users are health monitoring, not value
① Financial  EBIT / cost avoided / profit / cash flow / payback
② Operational cycle time / throughput / first-pass yield / rework / unit cost
③ Customer   response / first-contact resolution / conversion / retention / NPS
④ People     weekly active / proficient share / task coverage / skill transfer
⑤ Risk       error rate / hallucination rate / data exposure / review rate / audit findings

Hard rule ── Every positive metric needs a paired counter-metric, or it may not be reported
Hard rule ── No baseline, no ROI review
  • Governance maturity radar: score all five layers and find the weak one (pairs with 4.7)
  • Four-rung funnel: adoption → frequency → time → business result, with the drop-off at each rung
P57 Act VII · Accounting

From 100 gross to 15 net

Illustrative figures, shown to make the structure of the decay visible.

Almost every AI business case omits human review. We have seen it consume 60% of gross benefit. And review rate is itself improvable — 100% in month one falling to 20% by month six is the real maturity curve.

Course material 1 tables · 4.6

From 4.6 ROI calculation template

Full template (participant output 3)

Project Use case
Baseline Current unit cost / cycle time / quality
Exit condition Which of 1–5 you ticked
Gross benefit Annualised
Risk discount By evidence grade ×__%
Cost (all 8 items) Includes human review and data governance
① Theoretical-potential ROI %
② Risk-discounted ROI %
③ Finance-confirmed ROI Fill in after the pilot %
Payback period __ months
Business owner Sign-off
Finance sign-off Sign-off
P58 Act VII · Accounting

Three ROI figures

Every pilot carries all three. Executives act only on the last two.

  • ① Theoretical potential

    Project team · useful for imagination at kickoff · not a decision input

  • ② Risk-discounted business case

    Project team with finance · the basis for approval

  • ③ Finance-confirmed benefit

    Signed by finance · the basis for closure and scale-up

Discount by evidence grade: A at 80–90%, B at 50–60%, C below 30%, and time savings with no confirmed exit route at zero. The common failure is deciding on the first figure and never computing the third.

Course material 2 recaps · 4.6

From 4.6 ROI calculation template

ROI waterfall (recommended chart)

Gross benefit ██████████████████  100
  Models and compute   ▼ −12
  Integration          ▼ −18
  Data governance      ▼ −15
  Human review         ▼ −22   ← the largest single piece
  Change and training  ▼  −8
  Ongoing operations   ▼ −10
Net benefit   ███                15
Five exits ── less hiring/outsourcing · throughput · cycle→revenue · capacity redeployed · losses avoided
             (none of them ticked = the benefit path was never designed)

Eight cost items ── compute · licences · data governance · integration · compliance · training · human review · operations
             (human review is missed most, and can eat 60% of the gross benefit)

Three numbers ── theoretical potential / risk-discounted / finance-confirmed → executives read only the last two

Hard rule ── No baseline, no review; no CFO signature, no benefit

Act VIII

Governance

The line you do not cross

P59–P67

P59 Act VIII · Governance

Governance is not a brake

A race car has brakes so it can go fast, not slowly. Without governance AI stays in low-risk corners — and the core processes are where the money is.

Course material 1 tables · 6 recaps · 4.7

From 4.7 Governance and risk gates

The five gates (the core takeaway tool)

① Approval → ② Prototype → ③ Pilot → ④ Finance sign-off → ⑤ Scale-up
                                              ↑                    ↑
                                        CFO signs        only then may you replicate

① Use case approval gate

□ Business owner (not the IT lead)
□ Current baseline (no baseline, no ROI review)
□ Benefit mechanism (tick one of the five exit conditions)
□ Which staff are affected
□ Data sources and permissions
□ Unacceptable risks (the red lines)

② Prototype gate

□ Accuracy            □ Stability
□ Cost (including per-call cost)  □ Response time
□ Human review rate   □ Behaviour on anomalous input
□ Data leakage risk

③ Pilot gate

□ Keep a control group or historical baseline
□ Do not interview only enthusiastic users   ← survivorship bias
□ Measure time and quality together
□ Cover the high-loss edge cases

④ Finance sign-off gate ⭐

Cost cut / cost avoided / incremental revenue / cash flow / risk avoided

⑤ Scale-up gate

□ Data permissions   □ Model monitoring   □ Escalation path for exceptions
□ Vendor exit plan (how long to switch, and whether the data comes with you)
□ Staff training     □ Standing budget (whose money keeps it running)

Risk checklist (NIST framework, mapped to the China context)

[R027] [R016]

Risk category What actually happens inside companies Who owns it
Wrong content / hallucination The data, clauses or conclusions are invented Business owner + quality
Sensitive information leaks Staff paste contracts and customer lists into public services CISO + legal
Privilege overreach The AI can reach data it should not CIO
Shadow AI Staff are using it; the company does not know CIO + HR
Supply-chain dependency Single-model or single-vendor lock-in CIO + procurement
Compliance and personal data Cross-border data transfer, personal data handling Legal
Accountability is unclear When it goes wrong, nobody is accountable CEO
P60 Act VIII · Governance

26 leading models, hallucination from 22% to 94%

Not 2% to 9%.

A lower rate means the model knows more, or knows better when to say it does not know. First rule of selection: do not pick the most fluent, pick the one most aware of its own limits.

SourceStanford HAI, AI Index 2026, Ch.3 Responsible AI

Course material 3 tables · 2.8

From 2.8 Evidence on hallucination and governance

Measured hallucination rates (26 models)

Position Model Hallucination rate
Lowest Grok 4.20 Beta 0305 22%
Second Claude 4.5 Haiku 26%
Third MiMo-V2-Pro 30%
High end Gemini 3 Flash 92%
Highest gpt-oss-20B (high) 94%

⭐ The most dangerous finding: say something wrong and the model follows

Test condition Result
Baseline accuracy GPT-4o 98.2%
Switch to a knowledge-belief benchmark GPT-4o drops to 64.4%
Same test DeepSeek R1 falls from over 90% to 14.4%
Framing a false statement as something others believe The model handles it well
Framing the same false statement as something the user believes Performance collapse

Four more measured findings (each with a management action)

Finding Data Management action
Dialect and accent cut accuracy Leading models lose nearly half their accuracy on regional dialects Test dialects separately for service, field and frontline settings
Guardrails collapse once jailbroken On the AILuminate benchmark most frontier models rate good or very good; under adversarial prompting, every model tested got worse Safety testing must include adversarial cases, not just normal ones
Transparency is going backwards The Foundation Model Transparency Index rose from 37 to 58 between 2023 and 2024, then fell back to 40 in 2025; training data, compute and post-deployment impact remain the main gaps Procurement questionnaires must ask about training data and compute sourcing
The responsible-AI dimensions pull against each other Empirical work finds that training which improves one responsible-AI dimension consistently damages others Do not expect safe, fair, accurate and fast at once — rank them explicitly
P61 Act VIII · Governance Peak

A wrong premise from the user drags the model with it

The same false statement, framed two ways, produces two outcomes.

Framed as someone else's belief

"Some people think the earth is flat"

The model corrects you

Framed as the user's own belief

"I think the earth is flat"

Performance collapses

  • GPT-4o98.2% → 64.4%
  • DeepSeek R1>90% → 14.4%

In an enterprise setting: when staff bring a false premise to the model, it will usually go along — and make the error look more professional, more complete and more credible. Human review has to check the premise, not only the output.

SourceStanford HAI, AI Index 2026, Ch.3 Responsible AI

Course material 3 tables · 2.8

From 2.8 Evidence on hallucination and governance

Industry incidents are rising fast

Year AI incidents recorded (AI Incident Database)
Before 2022 Fewer than 100 a year
2024 233
2025 362

Good news: corporate governance is catching up

Metric Change
A dedicated AI governance role Up 17% in 2025
Share of companies with no responsible-AI policy at all Down from 24% to 11%
Organisations claiming no regulation affects them Down from 17% to 12%
Obstacle Share
Knowledge gap 59%
Budget constraint 48%
Regulatory uncertainty 41%
P62 Act VIII · Governance

Incidents rising, transparency falling

Transparency went backwards this year. Training data, compute and post-deployment impact remain the largest disclosure gaps — which is why procurement must ask about both data and compute provenance. The incident database is human-curated and skews to English-language, high-visibility events; the real number is higher.

SourceStanford HAI, AI Index 2026, Ch.3 Responsible AI

Course material 1 tables · 1 recaps · 2.8

From 2.8 Evidence on hallucination and governance

The standards companies cite most (for procurement and compliance)

Standard Share with citations Change
GDPR 60% Down from 65% in 2024
ISO/IEC 42001 (AI management system standard) 36% New in 2025
The NIST AI Risk Management Framework 33% New in 2025
Hallucination ── 26 frontier models, rates from 22% to 94% (not 2% to 9%)
              ── A low rate means it knows better when to admit it doesn't know
              ── First selection rule: pick the one that knows what it doesn't know, not the most fluent

⭐ Most dangerous ── A false premise framed as "others believe" → the model corrects you
                    Framed as "I believe" → performance collapses
                    → an employee with a wrong premise gets agreement, delivered more professionally
                    → human review must check the input premise, not only the output

Four findings ── Dialects cost nearly half the accuracy · every model's defences fall once jailbroken
              ── Transparency went backwards (58→40) · safety, fairness and accuracy fight each other

Incidents ── Under 100 before 2022 → 233 in 2024 → 362 in 2025
Governance ── Firms with no policy fall from 24% to 11%; the biggest barrier is a knowledge gap, at 59%
Standards ── GDPR 60% · ISO/IEC 42001 36% · NIST AI RMF 33%
  • Hallucination-rate bars: 26 models ranked, 22%–94%, both ends highlighted
  • Premise-contamination comparison: "others believe" vs "I believe", the latter collapsing in red
  • Incident growth curve: <100 → 233 → 362
  • Governance catching up: firms with no policy fall from 24% to 11%
P63 Act VIII · Governance

Twelve generative AI risks (NIST)

A neutral international frame that converts directly into a checklist.

  1. CBRN information or capability
  2. Confabulation: confidently stated falsehoods
  3. Dangerous, violent or hateful content
  4. Data privacy: leakage and failed de-anonymisation
  5. Environmental impact of training and inference
  6. Harmful bias and homogenisation
  7. Human-AI configuration: over-reliance and automation bias
  8. Information integrity: a lower bar for disinformation
  9. Information security: prompt injection, data poisoning
  10. Intellectual property: infringement and trade-secret exposure
  11. Obscene, degrading or abusive content
  12. Value chain and component integration: opaque upstream, weak vendor vetting

For reporting, NIST groups these three ways: technical/model risk from malfunction, misuse by humans, and ecosystem or societal risk. Four management functions: govern, map, measure, manage.

SourceNIST AI RMF Generative AI Profile

Course material 2 tables · 4.10

From 4.10 Risk checklist: NIST and CAICT

NIST's twelve generative AI risks

# Risk In plain terms How it most often shows up in companies
1 CBRN information or capability Easier access to CBRN and other dangerous information Rarely applies, but compliance questionnaires ask
2 Confabulation Stating something wrong with confidence (commonly "hallucination") Most common: invented clauses, figures and cases
3 Dangerous, violent or hateful content More likely to generate inciting, illegal or self-harm content Customer-facing service and content generation
4 Data privacy Biometric, health and location data leaks, or failed de-anonymisation Staff paste a customer list into a public service
5 Environmental impact Heavy compute consumption in training and inference ESG reporting, dual-carbon targets
6 Harmful bias and homogenisation Amplifies historical and systemic bias; performance varies by group and language Hiring, credit and pricing use cases
7 Human-machine configuration People anthropomorphise AI, over-rely on it, and defer to automation Staff stop reviewing — the most insidious risk
8 Information integrity Lowers the barrier to producing and spreading disinformation Published content is inaccurate
9 Information security Lowers the bar for attackers; AI itself is exposed to prompt injection and data poisoning Prompt injection, poisoned training data
10 Intellectual property Easier to reproduce copyrighted or trademarked content; trade secrets leak Infringing output, contaminated code licences
11 Obscene, demeaning or abusive content More prone to generating harmful imagery Customer-facing products must filter
12 Value chain and component integration Upstream third-party components are opaque and untraceable; vendor review is thin You don't know whose model and data your vendor used

NIST's three groupings (easier to present upward)

Category Includes
Technical and model risk (caused by failure) Confabulation, dangerous recommendations, data privacy, value chain integration, harmful bias and homogenisation
Human misuse (malicious use) CBRN, data privacy, human-AI configuration, abusive content, information integrity, information security
Ecosystem and societal risk (systemic) Data privacy, environment, intellectual property
P64 Act VIII · Governance

The four that actually happen

Twelve is a lot. For most companies, four occur constantly.

  • 1

    Confabulation

    It invents, confidently → high-risk output requires a human signature

  • 2

    Data privacy

    Staff paste sensitive data into public tools → give them a compliant route and an audit log

  • 3

    Human-AI configuration

    Review quietly stops → monitor the review rate continuously, do not set it once at launch

  • 4

    Value chain

    You do not know what sits underneath the vendor → ask about model and data provenance

The third is the easiest to miss. Month one, every output is read carefully. Month three, skimmed. Month six, approved on sight. Risk peaks in month six, not on launch day.

Course material 2 tables · 1 recaps · 4.10

From 4.10 Risk checklist: NIST and CAICT

The four that matter most (only these four are taught)

Priority Risk In one sentence Minimum countermeasure
🔴 1 Confabulation / hallucination (#2) It will invent confidently High-risk output must carry a human signature
🔴 2 Data privacy (#4) Staff paste sensitive data in Provide a compliant channel + audit logs
🟡 3 Human-AI configuration (#7) After a while nobody reviews it any more Review rate must be monitored, not set once at launch
🟡 4 Value chain integration (#12) You don't know who sits beneath your vendor Procurement questionnaires must ask where the model and data came from

The CAICT "two horizontals, three verticals" framework

            Development   Deployment    Use
              ┌──────────┬──────────┬──────────┐
 Management   │  policy · clear accountability · audit trail  │
              ├──────────┼──────────┼──────────┤
 Technology   │  capability · guardrails · monitoring         │
              └──────────┴──────────┴──────────┘
                     ↓ multi-stakeholder governance (ecosystem coordination)
Axis Meaning
Two horizontals Management and technology in step — policy pulling, capability supporting
Three verticals Development, deployment and use — protection across the whole chain
Outer ring Multi-stakeholder governance, ecosystem coordination
P65 Act VIII · Governance

CAICT's "two rows, three columns" framework

The Chinese reference model for internal AI governance.

Development side Deployment side Application side
Management (row)Policy leadership · clear accountability · audit trail
Technology (row)Capability support · guardrail tooling · monitoring and alerting

Multi-party industry governance

The cheapest first move is two management actions: publish a usage guideline and name an owning department. China Mobile, Tencent and Alibaba all started there.

SourceCAICT, AI Safety Governance Blue Book

Course material 2 tables · 1 recaps · 4.10

From 4.10 Risk checklist: NIST and CAICT

Five practice modules

Module Key point
Risk management Build a closed-loop risk management system
Model development Secure the development stage at source
System deployment Build a protective barrier at deployment
Application operation Strengthen dynamic assessment in use
Industry ecosystem Jointly build benchmarks and coordinated governance

Four things internal governance must do (CAICT wording; copy straight into policy)

# Action Chinese company examples (recorded by CAICT)
1 Write standard operating guidance, have management approve and communicate it, then tailor by department China Mobile issued AI safety risk guidance covering infrastructure, service provision and service use
2 Build a dedicated governance structure coordinating experts across the firm Tencent and Alibaba have AI safety governance bodies; Microsoft has its Aether Committee and IBM its privacy and responsible technology office
3 Build continuous tracking and audit: internal review plus management review, improving continuously CAICT leads the industry standard on AI safety governance and risk management capability
4 Risk management runs across the lifecycle: prevent early, track throughout, close the loop Google's SAIF framework ties AI system risk to business process
NIST's twelve ── confabulation · privacy · bias · human-AI configuration · information integrity · information security · IP · value chain…
NIST's four   ── Govern → Map → Measure → Manage

Watch four    ── 🔴 hallucination · 🔴 data leakage · 🟡 over-reliance · 🟡 opaque vendors
Most insidious ── Human-AI configuration: risk peaks in month six, not on launch day

CAICT framework ── "Two horizontals, three verticals"
  Horizontals: management + technology    Verticals: development + deployment + use
  Five modules: risk management · model development · system deployment · application operation · industry ecosystem

Cheapest first step ── ① publish usage guidance  ② name an owning department
P66 Act VIII · Governance

Who does what

AI drafts. A person signs.

AI executes

  • Drafting
  • Extraction
  • Classification
  • First-pass judgement

Human review

  • High-risk clauses
  • Amounts and dates
  • Exception escalation
  • Final signature

System automates

  • Archiving
  • Rule validation
  • Access control
  • Audit logging

Three principles: external commitments, financial amounts and legal terms always end with a person; the review rate is a managed metric, not a fixed cost; every exception has an escalation path with a human on it.

Course material 1 recaps · 4.7

From 4.7 Governance and risk gates

Human-AI division of responsibility (one swimlane explains it)

┌──────────────┬──────────────────┬──────────────────┐
│  AI executes │  Human reviews    │  System automates │
├──────────────┼──────────────────┼──────────────────┤
│ Drafting     │ High-risk clauses │ Archiving         │
│ Extraction   │ Amounts and terms │ Validation rules  │
│ Classifying  │ Escalations       │ Access control    │
│ First pass   │ Final signature   │ Audit logging     │
└──────────────┴──────────────────┴──────────────────┘
  • AI drafts, a person signs. External commitments, financial amounts and legal terms remain a human's responsibility.
  • The review rate is a manageable metric, not a fixed cost — bringing it down from 100% is itself proof of maturity.
  • Exceptions need an escalation path, with a real person on it.
P67 Act VIII · Governance

Eight questions for procurement

The third card you take home.

  • Which model sits underneath? Open or closed weights? Which version?

  • Where did the training data come from? Any rights issues?

  • Will our data be used to train your models?

  • Where is it deployed? Does our data leave our network?

  • Are there content guardrails? Who sets the blocking rules?

  • What protects against prompt injection and data poisoning?

  • How is liability for wrong output divided? What does the contract say?

  • If we switch vendors, can we export our data and knowledge base?

Course material 1 recaps · 4.10

From 4.10 Risk checklist: NIST and CAICT

Procurement questionnaire (take away; use it on vendors)

□ Which model is underneath? Open-weight or closed? Which version?
□ Where did the training data come from? Any copyright issues?   ← NIST #10, #12
□ Will our data be used to train your models?                    ← NIST #4
□ Where is it deployed? Does data leave our network?
□ Are there content safety guardrails? Who sets the blocking rules? ← NIST #2, #3
□ What protects against prompt injection and data poisoning?     ← NIST #9
□ How is liability for wrong output divided? What does the contract say?
□ If we switch vendors, can we export the data and knowledge base? ← exit plan
  • The two-horizontals, three-verticals framework (Mermaid grid)
  • Heat map of the twelve risks (rows = risk, columns = likelihood / impact)
  • Governance maturity radar (see 4.7)

Act IX

Execution

Ninety days, five gates

P68–P72

P68 Act IX · Execution

Ninety days, five gates

Not a three-year blueprint. Ninety days.

  1. Initiation

    Business owner · current baseline · benefit route · data sources · unacceptable risks

  2. Prototype

    Accuracy · stability · cost · latency · review rate · adversarial input · leakage risk

  3. Pilot

    Control group or historical baseline · measure time and quality together · do not survey only enthusiasts

  4. Financial confirmation

    Finance classifies the benefit: cost reduction, cost avoidance, incremental revenue, cash flow or risk avoided

  5. Scale-up

    Data permissions · model monitoring · escalation · vendor exit plan · training · standing budget

Gate four matters most: it turns an AI project from an IT narrative into an operating result finance will stand behind. Unconfirmed personal time savings are never booked as profit.

Course material 1 tables · 2 recaps · 4.7

From 4.7 Governance and risk gates

Governance maturity radar (self-assessment tool)

Six dimensions, scored 1–5 each:

     Policy & accountability
              │
   Agent ─────┼───── Data
   control    │      governance
              │
   Staff ─────┴───── Model
   training          monitoring
     Security & compliance
Dimension What a 5 looks like
Policy and accountability Every AI application has a named business owner
Data governance Classified and tiered, least-privilege, auditable
Model monitoring Accuracy, hallucination rate and cost monitored continuously with alerts
Security and compliance Cleared by legal and security, with a cross-border data ruling
Staff training Usage rules trained; red lines clear
Agent control Autonomy is bounded, write actions need approval, everything is logged
Five gates ── approval → prototype → pilot → finance sign-off (CFO) → scale-up
              ①           ②            ③        ④ ⭐                  ⑤

Seven risks ── hallucination · leakage · privilege overreach · shadow AI · supply-chain lock-in · compliance · unclear accountability

Division ── AI drafts, a person signs; the review rate is a metric you can lower; escalations need a real person

Framework ── NIST GenAI Profile [R027] + the CAICT safety governance blue book [R016]
             Govern → Map → Measure → Manage

Line ── A racing car has brakes so it can go fast, not slowly
P69 Act IX · Execution

CAICT's four phases and eight steps

Use this language upward and nobody questions the framework itself.

  1. Diagnose

    • Capability analysis
    • Requirement discovery
  2. Build

    • Solution design
    • Development and testing
  3. Apply

    • Application development
    • Effectiveness assessment
  4. Manage

    • Monitoring
    • Operations management

Five element layers: infrastructure · data · algorithms and models · application services · security and trust

The official effectiveness metrics include scenario penetration, process improvement rate and return on investment — even the official framework treats ROI as a maturity indicator. Decide internally with the seven-step loop; report upward in four phases.

SourceCAICT, Foundation Model Deployment Roadmap

Course material 5 tables · 1 recaps · 4.8

From 4.8 The CAICT rollout methodology

Four phases · eight steps · five layers

Diagnose ─────→ Build ─────→ Apply ─────→ Manage
 │               │             │            │
Capability      Solution      Application   Monitoring
analysis        design        build
Needs           Development   Performance   Operations
discovery       and testing   assessment    management

Five layers (check all five at every step):
infrastructure · data resources · algorithms and models · application services · security and trust
Phase Target Maps to the seven steps
Diagnose Clarify business growth and transformation needs ①②③ Cost / process / use cases
Build Harden the technical base ④ Small-scale validation (first half)
Apply Optimise deployment + build a performance assessment system ④⑤ Validation + ROI audit
Manage Monitoring and operations management ⑥⑦ Governance + scale-up

1. Completeness of core resources

Category What it assesses
Compute Floating-point capability, chip performance, energy efficiency, utilisation
Network Architecture, bandwidth, latency, stability
Storage Capacity, throughput, latency
Software infrastructure Vector databases, deep learning frameworks and operating systems: function, performance, compatibility
Data resources Volume, type and distribution; accuracy, completeness, consistency, availability
Algorithms and models Existing model assets: type, count, deployment mode, openness, compatibility

2. Balance of the talent mix

Dimension What it assesses
Technical capability Model architecture, algorithm optimisation, data governance, testing and validation
Basis of assessment Education, role, title and grade, years of experience, project history, IP
Management capability Team leadership, communication, planning, analysis and decisions, project management, self-learning

3. Fit with strategic planning ⭐

Dimension What it assesses
Strategic planning Mission, positioning, structure, market demand, risk appetite
Funding and budget How well the budget matches hardware, software, data acquisition and processing, people and operations
Overall judgement Whether the current budget covers build requirements, resource cost, timeline and risk

Apply phase: the five-layer maturity assessment

CAICT's assessment dimensions, usable directly as a self-assessment sheet:

Layer Assessment dimension
Infrastructure Resources (servers, chips, storage, vector stores, tooling) + overall performance (training and inference speed, compatibility, reliability, stability, autonomy)
Data resources Data composition (source, modality, distribution) + quality (accuracy, completeness, consistency, linkage, redundancy, volume, refresh rate) + diversity + availability + compliance + cost-effectiveness and return
Algorithms and models Function (perception, cognition, cross-modal fusion, self-learning) + performance (accuracy, compute efficiency, concurrency, response time, stability, robustness, reproducibility)
Application services Service experience + operations management (process, automation, continuous loop) + performance (use case penetration, optimisation rate, return on investment)
Security and trust Trustworthiness (hardware and software, data, model, service, content) + security (technical capability + policy)
P70 Act IX · Execution

Three readiness checks

CAICT requires three assessments before work begins.

  • Resource readiness

    Compute, network, storage, software stack, data assets, existing models

  • Team balance

    Technical and managerial capability, judged on experience, seniority and specialism

  • Strategy and budget fit

    Whether the current budget can actually support what the strategy promises

The usual failure is not technical. It is a strategy that says all-in and a budget approved as a pilot. When those two disagree the project dies halfway. Score all three and fix the lowest first.

SourceCAICT, Foundation Model Deployment Roadmap

Course material 2 tables · 1 recaps · 4.8

From 4.8 The CAICT rollout methodology

Manage phase: three things

Action Content
Real-time monitoring Real-time monitoring, tracking and early warning
Operations management A reference operating model for platforms and services
Building the system Principles for building a sound foundation-model operating system
Four phases ── diagnose → build → apply → manage
Eight steps ── capability analysis · needs discovery | solution design · development and testing | application build · performance assessment | monitoring · operations management
Five layers ── infrastructure · data resources · algorithms and models · application services · security and trust

Three diagnostics ── completeness of core resources · balance of the talent mix · fit with strategic planning ⭐
Maturity layers ── infrastructure / data resources / algorithms and models / application services / security and trust
Key metrics ── use case penetration · optimisation rate · return on investment (the official set)

Use ── The common language for reporting upward; nobody questions the framework itself
The seven steps, in the boss's language The CAICT four phases (reporting language)
① Where the money is ② Where it's blocked ③ What to do first Diagnose (capability analysis, needs discovery)
④ Does it work Build + apply (solution design, development and testing, application build)
⑤ Did it earn Apply (performance assessment)
⑥ Who owns it ⑦ Can it be rolled out Manage (monitoring, operations management)
  • Four-phase, eight-step flow (Mermaid; see 6.2)
  • Three-readiness radar (self-scored 1–5)
  • Five-layer maturity heat map (rows = layers, columns = maturity levels)
P71 Act IX · Execution

What leaves the room with you

Not a deck.

Three worksheets, filled in by hand

  • Scenario portfolio, with the do-not-do list
  • Value calculation: baseline, exit condition, three ROI figures
  • Ninety-day plan: five gates, owners, dates

Four pocket cards

  • Eight questions for any claim
  • Eight questions for procurement
  • Five exit conditions
  • Executive glossary

The worksheets are photographed before you leave, and reviewed together after ninety days.

Course material 6 tables · 2 recaps · 5.4

From 5.4 Workshop output template

Section A: scoring candidate use cases

□ Clear task boundary   □ Repeatable   □ Quality monitorable
(Missing one → fix it first; missing two or more → onto the do-not-do list)

Section B: priority use cases (3 maximum)

# Use case Business owner<br>(not IT) 90-day target Sign-off
1
2
3

Baseline (leave this blank and it does not enter ROI review)

Item Current value Data source
Unit cost
Cycle time
Quality metric (first-pass yield / accuracy)
Annual volume

Benefit exit (tick at least one)

□ 1. Less overtime, outsourcing or new hiring
□ 2. More throughput from the same people   → where does the extra demand come from? ______
□ 3. Shorter cycles producing revenue       → how is that revenue measured? ______
□ 4. Capacity redeployed to high-value work → redeployed where? measured how? ______
□ 5. Fewer errors / rework / incidents / churn → what is the current loss? ______

Five KPI layers (at least one positive and one paired counter-metric each)

Layer Positive metric Baseline Target Pair with a counter-metric Baseline Red line
① Financial
② Operational
③ Customer
④ People
⑤ Risk

The three ROI numbers

Amount / ratio Notes
Annualised gross benefit
Risk discount × ___% By evidence grade (A: 80–90% / B: 50–60% / C: 30% or less)
Annualised total cost (8 items) Compute · licences · data governance · integration · compliance · training · human review · operations
① Theoretical-potential ROI % *Indicative only; not a basis for decisions*
② Risk-discounted ROI % Basis for approval
③ Finance-confirmed ROI % Fill in after the pilot; basis for scaling
Payback period ___ months

Table 3 | 90-day action plan (the five gates)

Gate Deliverable Owner Target date Pass criteria
① Approval Business owner · baseline · benefit mechanism · data source · risk red lines Day __ All six in place
② Prototype Accuracy · stability · cost · response time · review rate · anomalous input · leakage risk Day __ All seven passed
③ Pilot Control group or historical baseline · measure both time and quality · cover edge cases Day __ Do not only interview enthusiasts
④ Finance sign-off ⭐ Classify the benefit: cost cut / cost avoided / incremental revenue / cash flow / risk avoided CFO:______ Day __ CFO signs off
⑤ Scale-up Data permissions · model monitoring · exception escalation · vendor exit plan · training · standing budget Day __ Replicate only when all six are in place
Stage Action
When to hand it out Hand out Table 1 before the workshop, Table 2 before the KPI module, Table 3 at the end
Ground rules Handwritten, not typed. Writing by hand changes how deeply people think
Closing action The instructor reads out each name and date in the commitment column
On file Photograph it live and post it to the group
Follow-up Check in after 90 days — the best moment for a repeat engagement
P72 Act IX · Execution

Four questions

Back to where we started

  • Where the value is
  • How strong the evidence is
  • Who owns it
  • How to scale without losing control

This course does not teach you to use AI. It teaches you to sign off on it. Models turn over every few months; these four questions will still hold in ten years.

Course material 1 tables · 1 recaps · 1.1

From 1.1 Core claim and hooks

Three-sentence version (for the opening)

  • AI is no longer optional. The State Council set the timetable: over 70% adoption of smart devices and agents by 2027.
  • But using it is not earning from it. 88% of organisations use it; only 13% report enterprise value.
  • Where is the gap? Nobody translated those three saved hours into a number on the income statement.

Why this claim lands with a CEO

What the boss is thinking What this course answers
"Another AI salesperson" I don't sell tools. I teach you how to sign one off.
"We already bought it; results are mediocre" Yes — 87% of companies worldwide are in the same place. The tool is not the problem.
"They say it's more efficient; I don't see the money" Because the time saved has no exit. Here are five exit conditions.
"How much should we spend" Don't ask how much to spend; ask where the baseline is. No baseline, no project.
"Will something go wrong" Five gates; the CFO signs the fourth, and only after the fifth may you replicate.

Hook 3: the −19% experiment

Claim   ── The gap is not in the technology, it is in the operating loop
Figures ── 88% use it, 13% get enterprise value → 75 points of arbitrage

Hooks   ── ① The 75-point gap: everyone is running, most are running on the spot
           ② Superstars: +33.5% vs +163% — AI widens gaps, it does not close them
           ③ +55% or −19%: same tool, two directions, and the difference is management

Close   ── Not how to use AI — how to sign AI off:
           where the value is · how strong the evidence is · who owns it · how to scale without losing control
P73 Appendix · Glossary and notice

Executive glossary

Each term explained only for what it means to the business.

Billing and cost

Token
The billing and compute unit. Double the usage, double the invoice
Inference
Every call made in service. The running cost lives here, not in training
Unit economics
Cost per unit of business. Without it you cannot claim a saving

Approaches

Retrieval augmentation
Search your knowledge base first, answer from what was found. The right starting point for 90% of companies
Fine-tuning
Continue training on your data. Use it when the model misses your vocabulary, not when it misses knowledge
Agent
Plans, calls tools and executes multi-step work. Anything that writes needs approval

Governance

Hallucination
Confidently stated error. Always present; it can be mitigated, never removed
Review rate
Share of output a person checks. A metric you can lower, not a fixed cost
Shadow AI
Tools staff use that the company does not know about. Open a front door rather than only blocking

Five claims most likely to mislead

  • "98% accurate" — on whose data, and what is the error rate on high-loss edge cases?
  • "Zero hallucination" — the official wording is "to a degree"
  • "Works out of the box" — your data quality decides everything
  • "An industry-specific model" — trained on what, and how much better than retrieval?
  • "80% more efficient" — task level is not role level is not enterprise level
Course material 6 tables · 1 recaps · 9.1

From 9.1 CXO glossary

Capability

Term One sentence, in business language What the boss should ask
Foundation model / LLM General language capability trained on vast amounts of text Do we buy, rent, or build?
Multimodal Handles text, images, audio and video together Does our use case need images? If so, don't pick a text-only model
Token The smallest unit models bill and compute in — roughly half a Chinese character to a word This is what the bill measures — double the usage, double the money
Context window How much the model can hold at once Decides whether a long contract fits in one pass
Inference Every computation the model performs while serving The running cost is here, not in training
Hallucination / confabulation States something wrong with confidence It never goes away — only mitigates — so review cost cannot be removed

Practice terms (tied to the four decisions)

Term In one sentence When to choose it
Prompt engineering Asking better so the model answers better The cheapest starting point; try this first
RAG (retrieval-augmented generation) Search your knowledge base first, then have the model answer from what it found The right starting point for 90% of companies; mitigates hallucination "to a certain extent" [R018]
Fine-tuning Continue training on your data so it learns your jargon and formats Use it when the gap is jargon; if the gap is missing knowledge, extend the knowledge base instead
Efficient fine-tuning Updates only a small share of parameters, saving compute and time Data available, compute limited
Knowledge base / vector store Turning company documents into something the model can retrieve The foundation of RAG; document quality sets the ceiling
Agent Plans, calls tools and executes multi-step tasks on its own Clear process + an interface + verifiable + reversible — all four before you go
MCP / tool calling Letting the model act on external systems (query, order, send) Write actions need approval; read-only can be opened up

Cost and deployment

Term In one sentence Business meaning
Public cloud / private cloud / hybrid Runs in someone else's data centre / your own / both Decides whether data leaves the company, and how much you pay upfront
On-premise deployment The whole model sits in your own environment The default answer in finance; most banks use it [R016]
On-device / edge deployment Runs on the device, never reaches the cloud Field, production line, offline settings
FinOps A practice that makes cloud and compute cost visible, attributable and improvable Make each department answer for what it spends
Unit economics What each unit of business costs If you cannot state this, you cannot talk about cutting cost
Small model / large model Under a billion vs over ten billion parameters The easiest decision to save money on; costs can differ by an order of magnitude

Governance

Term In one sentence Who owns it
Human review rate What share of AI output needs a human to look at it A metric you can drive down, not a fixed cost
Guardrails Mechanisms that block improper input and output In use at PSBC and Ant [R016]
Prompt injection An attacker manipulates the model through input into unintended behaviour CISO; NIST risk #9
Data poisoning Poisoning training data to steer model output CISO; NIST risk #9
Shadow AI AI tools staff use that the company does not know about CIO + HR; open a front door, don't only block
Explainability Being able to explain why the model reached that conclusion A hard requirement in finance, healthcare and high-risk settings
Safe rollback You can switch the AI off immediately and revert to people A hard requirement in manufacturing and energy [R016]
Usable without being exposed Training without handing over data (federated learning, secure multi-party computation) Already in use in the energy sector [R016]

Evaluation

Term In one sentence Why it matters
Baseline The real number before the change No baseline, no ROI review
Control group The group not using AI, kept for comparison Without a control group you cannot tell whether AI did it
A/B/C evidence grades Experiment or regulator · company self-report · media retelling Grade A goes into resolutions, B into approvals, C onto the whiteboard
First-pass yield Share that passes without rework The counter-metric that pairs with speed
Use case penetration The share of that use case actually handled by AI CAICT's official maturity metric [R018]
Return on investment ROI CAICT lists it as a maturity metric — this is not "ROI-obsessed"

The five words you are most likely to be sold ⭐

Term What the vendor wants you to believe Truth
"98% accurate" Almost never wrong Ask: whose data was it tested on? What is the error rate on high-loss edge cases?
"Zero hallucination" It stops making things up The official wording: RAG mitigates it only "to a certain extent" [R018]
"Works out of the box" Works out of the box Your data quality decides everything; 82.5% of industry datasets lack content density [R015]
"Industry foundation model" Knows your industry better Ask: where did the training data come from? How much better than RAG? Worth the price?
"80% more efficient" Costs can fall 80% Task level ≠ role level ≠ enterprise level. Ask how the role's time is composed
Billing unit ── the token: double the usage, double the money
Starting point ── RAG: the right place for 90% of companies to begin
Trap ── Hallucination never goes away → human review cost cannot be deleted
Saving ── Small model vs large: costs can differ by an order of magnitude
Security ── Write actions need approval; read-only can be opened up
Hard metrics ── baseline · control group · first-pass yield · unit economics
Sales-proofing ── "accuracy" / "zero hallucination" / "out of the box" / "industry model" / "80% more efficient"
P74 Appendix · Glossary and notice

Notice and sources

  • Third-party reports, data, charts and cases remain the property of their respective owners.
  • Figures keep their original survey definitions; redrawn charts state their data source.
  • Case figures carry an A / B / C evidence grade. B and C are self-reported or media-reported and are indicative only.
  • This is management training material. It is not investment advice, legal advice, or a promise of any particular return.

Sources retrieved 2026-08-05 · 67 sources, all publicly verifiable

Course material 3 tables · 3 recaps · 7.3

From 7.3 Licensing red lines

Red-amber-green

🟢 Free to use 🟡 Usable with conditions 🔴 Do not touch
Charts you drew yourself in Mermaid or ECharts Government documents and infographics (full citation) Screenshotting a consultancy's chart into commercial material
MIT / Apache 2.0 code and templates CAICT report data (cite the report and page) Making a modified version of CC BY-ND material
Charts redrawn from public data CC BY content (keep the attribution) Reproducing a whole book or HBR article
Officially published sample data GPL code (fine for demos; commercial use needs legal) Downloading a podcast or video and redistributing it
Short quote + explicit attribution Case figures (with evidence grade) Bypassing logins, paywalls or DRM

Government and official documents (R013, R014, R016, R027 + the MIIT cases)

  • ✅ May be summarised and quoted with organisation, title, date and link
  • ⚠️ Government infographics (such as NDRC explainers): quote in full; never crop the source, date or emblem
  • ⚠️ Government marks and emblems must not be altered
  • ⚠️ Check each site's redistribution terms before commercial training use
  • NIST material (US government) is broadly reusable, but check third-party charts, marks and appendices

CAICT reports (R015, R016, R017, R018)

  • ✅ Extract the figure and cite the report and page
  • 🔴 Reproducing a full chart needs permission; copyright reserved
  • 🔴 Limit how much you reproduce in public material; never in full

Consultancy reports (R020–R025, R028, R034, R035, R038 — BCG, Deloitte, McKinsey, PwC, Accenture)

  • 🔴 Do not screenshot the original into commercial material — copyright reserved
  • ✅ Restate the conclusion and redraw from their data, footnoting "redrawn from XX data"
  • ✅ Quote the figure with its source and survey basis

Stanford AI Index(R019 R058)

  • ⚠️ If the report is CC BY-ND 4.0: redistribute unchanged with attribution; no modified versions
  • ✅ To redraw, start from their Public Data and check that source's licence
  • ✅ Footnote it "redrawn from AI Index Public Data"

Open-source code and templates

Resource Licence Note
OpenAI Cookbook[R001]· MS GenAI[R002]· FinOps Toolkit[R004]· Mermaid[R009]· Marp[R010] MIT Modify and use commercially, provided the notices stay
OpenCost[R005]· Cloud Custodian[R006]· ECharts[R011] Apache 2.0 As above; OpenCost documentation is CC BY 4.0 and attribution must be kept
PM4Py[R007] GPL family ⚠️ Usually fine for training demos; embedding in a product needs legal review
FOCUS sample data [R008] CC family (per license.md) Verify the exact version and attribute it
AntV[R012] Varies by sub-project ⚠️ Check each repository; trademarks and design assets are not covered by the code licence
Datawhale LLM Universe[R003] ⚠️ No open licence is clearly shown on the homepage Check the project terms or contact the maintainer before redistributing

Books and articles (R029–R038)

  • ✅ Write a summary, extract the framework, use the four-card method
  • 🔴 Do not reproduce in full; HBR articles are copyrighted
  • ✅ Rewired has a CITIC simplified-Chinese edition; the title and arguments may be cited normally

Podcasts, video and webinars (R039–R047)

  • ✅ Use only a 1–3 minute core clip, or a timestamped link
  • 🔴 Do not download and redistribute, or re-edit and publish
  • 🔴 Do not bypass logins, paywalls, DRM or streaming restrictions
  • Without an official download or open licence, keep only the public page, transcript and metadata

Company cases (R048–R057)

  • ✅ Quote the disclosed figure with its source and evidence grade
  • ⚠️ Klarna and similar corporate content is copyrighted: quote the figures and conclusions, never the page
  • ⚠️ MIIT cases are official publications; keep the attribution when quoting

Five items in the slide footer (required on every external chart)

Publisher · report title · date · page or figure number · licence, or a "redrawn from original data" note

Data-security red lines (for both POCs and training)

🔴 Never put full contract text, complete approval records, personal data or sensitive internal approvals into public output
🔴 Never upload real company data to a public AI service for a classroom demo
🔴 Never record keys in documents, test output or chat summaries
🔴 Never leave local absolute paths in external material
✅ Classroom demos always use anonymised, synthetic or officially published sample data
✅ Reports and slides carry short quotations and locators only

Internal training vs commercial course material

Item Internal corporate training Commercial course material
Consulting report charts Showable within fair-use limits 🔴 Redraw it; never screenshot the original
CAICT report May be quoted at length Limit how much you reproduce
Book content Summary + framework Summary + framework (stricter)
Video clip Playing a clip live Link with a timestamp only
Government infographic Full citation Full citation + check the redistribution terms
Corporate case data Attribution is enough Attribution + evidence grade + disclaimer

Final check before delivery

□ Every external data slide carries the five-item footer
□ No consultancy chart has been screenshotted directly
□ No CC BY-ND material has been modified
□ Government infographics are quoted in full, with emblem and date intact
□ CAICT data cites the report title and page
□ PM4Py has been through legal if it is being productised
□ Datawhale material has had its licence checked before external release
□ Video appears only as a clip or a link, never redistributed
□ Case figures carry an evidence grade (A/B/C)
□ No real customer data, personal information or keys
□ No local absolute paths
□ The disclaimer is present: this course is supporting review and training material, not legal or investment advice
1 / 74 ← → to page · space to advance · Home / End for first and last · ESC for the index