TLDR: The companies profiting from AI scaled the work they had redesigned first, kept people on the decisions that carry risk, and grew only the parts that stayed reliable.
The speed trap
Speed was the first thing every leadership team noticed about generative artificial intelligence (AI), and speed is where the trap was hiding. A pilot stood up over a weekend. A support queue cleared overnight. A campaign that used to take a fortnight shipped in an afternoon. The dashboards looked spectacular, and the profit-and-loss statement (P&L) stayed flat.
The Massachusetts Institute of Technology’s Project NANDA put a number on the distance between motion and money. Its 2025 study, The GenAI Divide: State of AI in Business, found that only 5 per cent of enterprise AI deployments produced measurable P&L impact, while the other 95 per cent stalled before the books changed.1 The decisive factor sat in integration: how well a company folded AI into its workflows, its governance, and the way its teams actually decide. The same report found that pilots built by pairing internal specialists with outside expertise reached production 67 per cent of the time, against 22 per cent for builds left to the technology team alone.
That gap is the whole story. Scaling fast produces a demo; scaling smart produces a system that holds at ten times the volume.
What “smarter” means: redesign over acceleration
McKinsey’s 2025 survey of nearly 2,000 leaders across 105 countries draws the same line from the opposite direction. Eighty-eight per cent of organisations now use AI in at least one function, about a third have begun to scale it enterprise-wide, and 39 per cent report any effect on company-level earnings before interest and taxes (EBIT).2 Agents sharpen the picture further: 62 per cent of companies run experiments with them, and 23 per cent have reached real scale.
The teams that cross from pilot to profit share one move. They treat AI as a reason to redesign the work itself. The losing move bolts a faster engine onto the work they already had. They rewrite the process, decide where judgment lives, and instrument the output so quality stays visible. Acceleration drops a model onto a leaky funnel and ships the leak at speed. Redesign rebuilds the funnel so the model has somewhere dependable to plug in.
For a small or medium-sized enterprise (SME), that is the liberating part. Enterprise-wide scale is optional. A focused team that redesigns one revenue-critical process and grows only the parts that stay dependable will out-earn a giant running ninety stalled pilots.
Klarna scaled fast, then rebuilt the human layer
The cautionary case already exists, and it belongs to a company brave enough to go first. When Klarna launched its OpenAI-powered assistant in 2024, the early numbers were screenshot material: the bot handled 2.3 million conversations in its first month and cut average resolution time from eleven minutes to under two, absorbing two-thirds of the support volume.3 By 2025 the company counted roughly 60 million dollars in savings and a cost-per-transaction cut by about 40 per cent.
Then the curve bent. Pushed onto harder cases, the system produced hallucinations on edge cases and softer satisfaction scores on complex, emotional tickets even when the answers were technically right, and Klarna began rehiring human agents.4 Chief executive Sebastian Siemiatkowski said plainly that cost-first automation produces lower quality, and committed that a customer will always reach a person on request.
Read it correctly and Klarna is a calibration story with a happy second act. The volume work, the password resets and the where-is-my-order queries, scaled cleanly and stays automated. The judgment work, the dispute and the distressed customer, went back to people. The takeaway holds for any operator: automate the volume layer, and keep people on the layer where consistency and trust are the actual product.
Where AI scales cleanly: high-volume, rule-clear work
The operational headaches that ease under AI share a shape. They run on high volume, repeatable logic, and a clear definition of a good answer. That is where scale becomes an asset.
Top-of-funnel screening is the cleanest win. A fund drowning in inbound decks, or a sales team buried in leads, can point a memory-equipped agent at the flow, hand it the thesis or the ideal-customer profile once, and let it rank, summarise, and surface the few that deserve human time, with the reasoning attached. Tier-one customer support is the second, the layer Klarna kept. Compliance and onboarding checks (identity, sanctions, ownership tracing) are the third, where the agent gathers and drafts while a person signs off. Content and campaign production is the fourth, where variants and versions parallelise with ease.
The quality evidence on this kind of work comes, unusually, from a controlled experiment. In a randomised controlled trial (RCT) of 758 Boston Consulting Group consultants published in Organization Science, those working with AI completed 12.2 per cent more tasks, moved 25.1 per cent faster, and produced work rated about 40 per cent higher in quality.5 The figure that matters most for anyone scaling a team: the weakest performers improved by 43 per cent. AI lifted the floor faster than the ceiling, which is the working definition of consistency.
Quality and consistency are engineered on purpose
The firms that grow while protecting the customer experience (CX) treat quality the way a factory floor treats it: something to design, measure, and improve on purpose. Manufacturing already named the disciplines, Kaizen for continuous improvement and Six Sigma for defect reduction. AI makes both practical at software speed, once the organisation wires them in.
Two mechanisms carry the load. The first is memory. An agent that remembers the account, the brand voice, the last three interactions, and the correction it received last week delivers the same standard on the ten-thousandth touch as on the first. Consistency at scale is a memory problem before it is an intelligence problem. The second is relevance, and it pays in cash. McKinsey finds that companies which personalise well generate roughly 40 per cent more revenue from those activities than the companies that leave it on the table.6 Persistent-memory agents make that practical for a small team, tailoring the next touch across an entire pipeline. That is choice architecture in the sense Richard Thaler and Cass Sunstein gave the term, the deliberate design of the next decision, applied across a whole funnel where it once shaped a single landing page.7
Here is where growing clients report the deepest impact. The headline cost cut is real and finite. The compounding effect is the prize: every interaction lands more relevant than the last, the team levels up as the load grows, and the brand sounds like itself whether it answers ten enquiries a day or ten thousand.
The pattern behind the 5%
Strip away the logos and the winners rhyme. They redesigned a process before they scaled it. They drew a hard line between volume work, which they automated, and judgment work, which they augmented and kept under human sign-off. They invested in memory and proprietary context, so the agent grew more reliable over time as well as faster. And they bought outcomes, instrumenting quality so they could see the moment growth started to cost it, the exact gauge Klarna’s customers had to supply by hand.
That is sustainable scaling: output that climbs while quality, consistency, and trust hold their line. It takes longer to stand up than a weekend pilot, and it is the only version that survives contact with a real customer base.
Where Kainjoo stands
At Kainjoo, we build brand and technology for regulated industries, where a single inconsistent answer carries real cost. Our view is that the durable advantage in this cycle belongs to the operator who owns the memory, the data, and the trust that agents plug into. We treat AI as a Kaizen and Six Sigma instrument: automate the volume, instrument the quality, keep sensitive work on infrastructure you control, and keep a human signature on the decisions that carry risk. Build memory worth keeping, and you get harder to replace as you grow.
What leaders should do now
The move that separates the profitable 5 per cent is specific, and it depends on your seat.
- Founders and chief executives: choose one revenue-critical process and rebuild it around AI from the ground up. Scale only the layer that stays reliable, and hold the line when the volume win tempts you to automate the judgment layer too.
- Operations and Revenue Operations (RevOps) leaders: point agents at high-volume, rule-clear work (screening, tier-one support, first-pass compliance, production), and instrument quality so any slip shows up the day it happens, well before it surfaces in churn.
- CX and brand leaders: make memory the priority. An agent that holds account history and brand voice keeps the experience steady as volume climbs. Route complex and emotional tickets to people.
- Regulated and quality-led teams: run AI as a Kaizen and Six Sigma instrument. Keep sensitive data on infrastructure you control, keep a human signature on the final decision, and let the agent gather and draft the first pass.
FAQ
What separates scaling smart from scaling fast? It means growing output while quality, consistency, and trust hold steady, by redesigning the process around AI and scaling only the parts that stay reliable.
Which operational work becomes easier with AI first? High-volume, rule-clear work: deal and lead screening, tier-one support, first-pass compliance checks, and content production.
How do companies protect quality and customer experience while scaling? They engineer it, with persistent memory for consistency, instrumented quality metrics, and people kept on complex, high-risk decisions.
Where do growing companies see the greatest impact? In compounding relevance: personalised interactions that lift revenue, a team that levels up, and a brand that stays consistent from the first interaction to the ten-thousandth.
Which work stays with people? Taste, the difficult conversation, and the final approval stay with people. Agents generate the options and run the first pass.
References
- finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html — MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (coverage of the 5%/95% finding)
- mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai — McKinsey, The State of AI in 2025: Agents, Innovation, and Transformation
- langchain.com/blog/customers-klarna — Klarna AI assistant: volume and resolution-time results
- blog.promptlayer.com/klarna-customer-service-from-ai-first-to-human-hybrid-balance — Klarna’s quality recalibration and human-hybrid model
- pubsonline.informs.org/doi/10.1287/orsc.2025.21838 — Dell’Acqua et al., “Navigating the Jagged Technological Frontier,” Organization Science (Harvard/BCG randomised controlled trial)
- mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying — McKinsey, The Value of Getting Personalization Right (or Wrong) Is Multiplying
- Thaler, R. H. & Sunstein, C. R. (2008), Nudge: Improving Decisions About Health, Wealth, and Happiness, Yale University Press — choice architecture