"The PoC went great, but then nothing happened after that." That's what one executive told me last month at a dinner following a board meeting at a Seoul manufacturer. The whiteboard in the meeting room still had the pilot metrics written in green, and the team lead was smiling — but his eyes looked a little tired. Eight weeks to the demo, and for the five months since, nobody has touched the tool.
Which brings us back to the very problem we opened this series with. Why do the five failure patterns we covered last time keep repeating themselves? In most cases, it isn't a lack of technology — it's the absence of an agreement on what deliverables need to be produced at which stage. Now, at the end of an eight-part run, it's time to tie the thread. If we've walked through concepts, categories, ROI, stack, cost, security, and failure patterns, there is only one question left. "So starting Monday, what exactly do we do, and how?"

Why 87% Never Reach Production
In a 2024 survey, Gartner reported that 87% of AI projects fail to reach production. S&P Global's 2025 figures put the average prototype-to-production time at eight months, with only around 48% actually crossing the finish line. The number 87% sounds terrifying, but flipped around it says this: the truly hard stretch isn't the eight weeks of plugging in a model — it's the five or six months that come after.
Around the same time, Gartner published another figure from a different angle. Thirty percent of generative AI projects are scrapped after PoC, with the main causes being a lack of governance and unclear ROI. On the flip side, organizations that follow a structured pilot process are 2.5x more likely to report ROI within twelve months. In the end, success is decided not by "how well you tuned the model," but by "how well you designed the stages."
To be honest, when I first entered this market I thought it was the opposite. I assumed that with a good model and good data, everything else would naturally fall into place. In reality, that "naturally falling into place" part eats up half the project budget.
Stage 0 — Three Things to Pin Down Before You Start
Recently the industry has taken to calling this stretch "Phase 0." Teams that lock down three things in writing before starting a PoC are the ones that escape pilot purgatory: evaluation criteria, deployment pipeline, and governance approval chain. By contrast, teams that opt for "let's build the coolest demo first and harden it later" mostly stall out around month five.
Operationally, a single half-day workshop is enough. What are the quantitative metrics we'll use to judge that this tool succeeded? What pipeline does it need to pass through to go to production? When personal data or confidential documents are involved, who signs off last? If you don't have answers to those three questions, you're not ready to start a PoC. This is where the ROI and KPI design from Episode 3 and the security and governance from Episode 6 come together.
Stage 1 — Diagnosis (2–3 weeks)
Diagnosis starts in the meeting room, but the answers are always on the front line. When we run diagnosis with a client, we work three tracks in parallel. On the work track, we observe how each team's hours are spent — where the day drains away and how many repetitive tasks there are. On the data track, we map where internal documents, tickets, and manuals live and in what format. On the systems track, we check how far the access-control system and audit logging have been built out.
Only two deliverables are needed at this stage. The first is a table of three to five candidate use cases, each annotated with estimated time savings and risks. The second is a single paragraph defining success for the first pilot target. A 30-item pile of documents is actually a warning sign — it means the team hasn't decided what matters.
Stage 2 — PoC (6–8 weeks)
The industry consensus is that a well-scoped internal tool can be production-ready in eight to twelve weeks. The first six to eight of those are the PoC window. Which combination of RAG, fine-tuning, and prompting we compared in Episode 2 to run with should already be settled in the diagnosis phase. Teams that overhaul their architecture at this stage generally don't make it to production.
The exit criteria for a PoC aren't "the demo worked" — they're three answers. First, when 20 to 30 real users ran the tool for two weeks, how much did the KPIs defined in diagnosis actually move? Second, how were the wrong answers, hallucinations, and permission incidents that came up logged and categorized? Third, are the budget, staffing, and governance approvals needed for the next stage secured? If those three questions don't have answers, it's better to end the PoC and redesign than to extend it.
Stage 3 — Full Build (2–3 months)
The part that most often collapses during the full build isn't, surprisingly, the model — it's the integrations. Corporate SSO, document repositories, ticketing systems, audit-log servers, notification channels. Miss any one of those five and user churn starts within three weeks of deployment. In the domestic cases Hello T observed in 2026, the answer to "why isn't anyone using it, even though the PoC succeeded?" almost always traced back to one of those five integration points.
The deliverables for this stretch are three: a production pipeline (deploy, rollback, A/B), an operations dashboard (response quality, cost, usage, refusal rate), and user onboarding materials (a guide of two pages or less plus a 30-minute session). I've seen more than a few teams burned by underestimating that third one. Even a great tool won't be reopened once someone hits a wall in the first 30 minutes.
Stage 4 — Operations and Ongoing Development (continuous)
Presenc AI's June 2026 report shows that 85–90% of large enterprises are already running at least one LLM in production, a jump from 65% a year earlier. The message is unmistakable. "Whether to adopt" is a debate that's already over, and the competition of the next three years will be decided by "how well you operate it."
Operations run on three rhythms. In the weekly rhythm, you watch usage, refusal rate, and quality scores. In the monthly rhythm, you reconcile the cost structure from Episode 5 against the actual invoice, and inspect prompts, search indexes, and document freshness. In the quarterly rhythm, you reassess use cases and decide which new departments and tasks to bring on. If your team can't internalize this rhythm and instead leaves it entirely to the partner, the tool will become a fossil within six months.
Five Questions to Ask When Choosing a Partner
When picking a partner to walk with you from diagnosis to operations, there are five questions every practitioner should ask. If they can't answer these concretely, you should think twice, no matter how good the demo looks.
- How are the SLA and response procedures at the six-month operations mark written into the contract?
- How is responsibility split for model updates, prompt changes, and document-freshness management?
- What's the roadmap and documentation for handing operations back over to your internal team?
- What are the rollback, data-deletion, and audit-log retention procedures if things go wrong?
- What reusable assets (pipelines, evaluation sets, governance documents) remain when you expand to the next use case?
The fifth question matters most. A good partner leaves organizational muscle behind when the project ends. A bad one makes sure nothing runs without them.
Tying the Thread — Eight Episodes on One Page
That's eight episodes. We started with the concepts and categories of a proprietary-data LLM in Episode 1, moved through methods, ROI, stack, cost, security, and failure patterns, and arrived at today's execution plan. If you've followed all eight, one thing should be clear by now: adopting your own LLM is not a "technology deployment project" — it's "building a new muscle inside the organization."
In a market where 30% is scrapped after PoC and 87% never reaches production, the requirement for landing on the winning side isn't a flashy model. It's boringly, faithfully walking through the four stages of diagnosis, PoC, full build, and operations. Instead of raising a toast because the demo wrapped in eight weeks, the question to ask is whether that demo is still being opened every day six months later.
At 5years+, we work with mid-sized and SMB companies across Korea and Japan across the full stretch — from diagnosis through operations. If you'd like a roadmap tailored to your own organization, please share your current situation in our free adoption-roadmap design consultation. The series ends here, but the real execution starts now.