The meeting that ends with "it feels about 30% faster"
Late July, in the evening meeting room of a mid-sized manufacturer. On the whiteboard, half-erased in red marker, were the words "AI Rollout — 6-Month Review." The team lead running the project stood in front of five executives and said, "It feels about 30% faster." One of the executives clicked his pen a couple of times, then quietly asked, "Where does that 30% come from?" The team lead paused. Everyone in the room already knew the answer. Nowhere. It was a gut feel.
If the previous piece on hidden operating costs was about "the reality of spending," this one is the opposite side. Spending lands neatly on the approval sheet; benefits almost never do. That asymmetry is the real bottleneck for AI budget approvals.

Why benefits only get talked about as feelings
Talk to enough practitioners and the reason is almost always the same: no one measured the pre-rollout state. How many minutes did it actually take to draft one document? How many hands touched a single customer inquiry? What was the rework rate? Without a baseline, there is no delta. The "30% faster" instinct may be real, but there is no way to convert that 30% into hours.
When you run a rollout review without a baseline, what you are left with are qualitative signals: "the mood is better," "the team seems happier." That is not bad news for an executive, but it is thin evidence when you are asking for the next tranche of budget. A recent U.S. survey found that 56% of CEOs said they saw "no revenue or cost benefit at all" from AI investments. The number is less damning than it looks — in many cases the benefit existed, but the metrics to capture it did not.
Gartner's line that "only 28% of AI projects fully meet ROI expectations" gets recycled a lot, but reading it as "AI does not work" is only half the sentence. The real sentence is: there was no yardstick to judge whether AI worked, from day one.
A scorecard that gets approved — four lines is enough
The metrics do not need to be complicated. If anything, the opposite. A four-line scorecard per workflow is enough. It is close to a frame Atlassian's Teamwork Lab has been publishing, and it drops cleanly onto a Korean approval sheet.
- Speed — processing time (minutes per unit). Before rollout vs. after.
- Quality — rework rate, or the rejection rate on submitted work. "First-draft acceptance rate" works too.
- Cost — cost per unit (KRW). Fully-loaded hourly rate × time saved.
- Adoption — actual usage rate (weekly active users / assigned users). Satisfaction surveys are secondary.
Four lines and the approval sheet changes completely. "Feels about 30%" becomes "average draft time 42 min → 19 min, rework rate 18% → 11%, cost per unit ₩8,400 → ₩3,900, WAU 62%." There is no room left for a pen to click.
ROI in one line — hourly cost × time saved × reuse rate
The math is simpler than it looks. The more complicated you make it, the less the approver trusts it. Take one document-drafting workflow.
Assume a fully-loaded hourly cost of ₩30,000 per team member (salary, benefits, and overhead all in). Each unit saves 23 minutes. The team processes 400 units per month. And the actual "reuse rate" — how often the tool was really pulled off the shelf again — is 70%. The calculation: ₩30,000 × (23/60) × 400 × 0.7 ≈ ₩3.22M per month. Subtract ₩800K in license and API costs, and net benefit lands around ₩2.42M per month, or roughly ₩29M per year.
Three things about this number are honest. First, the conversion from time to money did not skip depreciation or overhead — it uses the fully-loaded rate. Second, the "time saved" is multiplied against actual used units (via reuse rate), not eligible units. Gut-feel ROI almost always assumes 100% reuse. Third, license costs are subtracted. The moment you present gross benefit instead of net, credibility on the approval sheet collapses.
Two traps
Teams that start tracking metrics fall into two familiar traps.
The first is measuring only right after rollout. A three-month review that does not show a revenue lift often gets written off as failure, but recent international studies suggest 55% of AI's long-term value comes from indirect benefits — employee satisfaction, competitive positioning, innovation capacity. That 55% will never show up in a three-month review. The scorecard needs "6-month remeasurement" and "12-month remeasurement" slots built in from day one.
The second trap is chasing only revenue growth. Domestic surveys show the top AI benefits are efficiency and productivity gains (66%), stronger data-driven decision-making (53%), and cost reduction (40%) — with revenue growth way down at 20%. Most benefits, in other words, live on the cost side of the ledger. Yet approval sheets still ask for revenue-growth scenarios. That is how you miss the actual signal. If you do not also track where the saved time went — reallocated to other work? less overtime? new projects opened up? — the benefit gets logged as "disappeared somewhere."
Stage gates — approving in phases, not in full
Once metrics are in place, the grammar of approval itself changes. Instead of the binary "full approval vs. rejection," you move to stage gates: each next tranche of budget opens based on realized metrics.
One version looks like this. Phase 1: ₩60M over three months — target metrics are "≥20% speed gain and ≥50% adoption on three target workflows." Pass the gate and Phase 2 opens: ₩180M over six months — target is "≥25% reduction in cost per unit, quality metrics held steady." Pass that and Phase 3 goes org-wide.
The power of this frame is in the approver's psychology. From an executive's chair, signing off ₩300M in one shot is far harder than signing off ₩60M now with "we decide the next step from the numbers we see." From the practitioner's side, a project that has cleared a gate carries much stronger political defense inside the organization than one that got the full ₩300M upfront. A recent international study finding that "organizations using structured frameworks see 3.5× returns within 24 months and 3× odds of positive return" points in the same direction.
The rise of the "value realization" role
To absorb this shift organizationally, the term Value Realization Office has started to appear more often abroad. Literally: an office whose job is to track whether deployed technology actually produces value. In Korea, few departments carry that name yet, but over the past few months we have seen large-corp CDO organizations spinning up similar roles under names like "AI Business Planning Team" or "AI Performance Management Cell." Usually one or two people. But the difference between organizations that have those one or two and organizations that do not is obvious the moment you sit in the approval room. One side speaks in numbers; the other still speaks in "it feels."
Mid-sized companies do not need to formalize this as a department. But even designating a single "metrics owner" per rollout project changes the approval sheet completely. That person updates the four-line scorecard every month and stands in front of the executives at each gate. Projects without this role — no matter how well the rollout goes — do not get their second round of budget.
Closing — coming next
Setting up metrics is, in the end, a declaration that "we will not speak in feelings anymore." And oddly, that declaration affects not just the approval room but morale inside the project. When "what we did" is visible every month, the team moves differently. Without metrics, the good work does not get counted, and neither does the bad. Eventually everyone gets tired.
The next piece closes this series as a retrospective — the actual total cost after year one, and three lessons. Setting the original budget sheet next to the year-end reconciliation, what diverged, and what the divergence teaches — a full accounting of the whole series.
If you want help figuring out where to start on an ROI metric set, how to measure your organization's baseline, and how to distill it into a four-line scorecard — reach out to 5years+ for a metrics-set diagnosis. We can walk through Korean and Japanese rollout cases, with real examples of scorecards that have actually gone up for board approval.