The bad AI strategy meeting follows a familiar pattern. Someone presents several ideas that each seem reasonable on their own. Every one looks practical in isolation. The presentation rarely ranks them, and almost never names the existing process that would actually change.
Tool adoption often passes for capital allocation. Founders feel prudent after approving experiments with AI. Consultants stay current by listing agents on the roadmap. Operators stay covered once each department runs a pilot. Model access was never the scarce resource. Attention, process change, review capacity, trust, and the patience to embed a new workflow always were.
Market data makes the gap more pressing. Stanford's Human-Centered Artificial Intelligence (HAI) 2026 AI Index reports organizational AI adoption at 88 percent in surveyed organizations. McKinsey's late-2025 State of AI survey found the same split: nearly nine in ten respondents reported using AI, and roughly two-thirds still had not scaled it across the enterprise. International Business Machines (IBM)'s 2025 CEO study added the boardroom view: only a minority of chief executives said recent AI initiatives had delivered expected return on investment (ROI) or reached enterprise scale.
So AI strategy is a capital-allocation discipline. Decide which workflows deserve automation, which deserve assistance, and which should wait until operating conditions improve. The unit of analysis is the business decision that could move faster, cost less, carry less risk, or gain better information once the surrounding workflow changes.
If the workflow is not worth redesigning, it is not worth automating.
Automate-assist-ignore labels mark only the final call. The harder test is whether the proposed workflow carries enough strategic value, operational fit, proprietary context, and governance capacity to justify any change at all.
What Current AI Changes, And What It Does Not
The strongest frontier models have expanded the range of feasible work. They draft, classify, retrieve, compare, translate, summarize, follow narrow procedures, inspect documents, call tools, and coordinate multi-step tasks at a level that fits real operations. Anthropic's Economic Index shows business API usage leaning toward automation patterns. OpenAI's enterprise case library shows banks, pharmaceutical companies, and customer-service teams embedding models into repeated workflows. The old objection that the technology still needs years of improvement is weaker than it was two years ago.
Managerial judgment remains the bottleneck. Plausible sales briefs still leave open which accounts define a good market. Policy summaries still leave open when a loyal customer deserves an exception. Pricing evidence still leaves open whether the company should trade margin for adoption. In workflows that sit near identity, risk, trust, or scarce commercial judgment, "can AI perform this task alone?" is the wrong first question.
The hard limits are operational. Exception routing stays brittle when policies conflict with customer value. Audit trails often lag the confidence of a generated answer. Tacit company context lives in half-remembered deal reviews, Slack threads, and political agreements outside any retrieval corpus. Tool execution carries different risk from drafting. Feedback loops can improve a system or train it on the wrong signal. Outright failure is one risk; success at the wrong layer of responsibility is another.
Place the system in the chain of responsibility first. Automate when the output is bounded and the harm of error is low. Assist when judgment must remain visible. Wait when data is weak, the process is political, the owner is absent, or the decision itself remains confused — those conditions make automation premature.
Named deployments point the same way. Morgan Stanley's wealth-management assistant stands out because it placed retrieval, summarization, evaluations, compliance controls, and advisor review around a high-frequency knowledge workflow — not because it made finance autonomous. Klarna's customer-service assistant shows the upside of narrow automation when the routine repeats, spans languages, and produces measurable results; it also shows why vendor and company numbers need a claim ceiling. Banco Bilbao Vizcaya Argentaria (BBVA)'s 2026 OpenAI case reads more like an operating change than a chatbot rollout: leadership training, security and legal alignment, and selected workflow efficiency gains. Amazon's operations updates supply the clearest caution. Project Eluna is framed as decision support for fulfillment teams. Blue Jay left operations after an update even though the underlying technology continued. Serious firms stage autonomy, measure fit, and keep the option to stop. That is the pattern worth copying, success or failure labels aside.
The useful screen stays plain. Before approving an AI pilot, ask four non-overlapping questions.
1. Strategic Value
Name the business outcome that improves if the workflow changes. The answer must point to a revenue, cost, risk, speed, or learning decision. "Productivity" is too loose. "Account executives prepare for enterprise calls in half the time and ask better second questions" comes closer to an investment thesis.
High-value workflows sit near decisions that compound — account pursuit, feature shipping, claim language, risk escalation, customer promises kept. Low-value workflows often produce more internal text without altering a commercial decision. The most deceptive cases are administratively painful and strategically sterile: they keep employees busy enough to want relief without justifying redesign.
2. Workflow Suitability
Some work fits AI intervention more readily than others. Strong candidates repeat, stay bounded, carry context, and allow review. Weak candidates appear rarely, stay ambiguous, carry political sensitivity, or remain so poorly understood that automation would only encode confusion.
The repeatability test matters because custom one-off work hides process debt. The reviewability test matters because AI systems fail fluently. A workflow that no responsible human can review is a liability with a nicer interface.
3. Advantage Source
Ask whether the output improves because it belongs to this company. If a generic model and public data can produce the same result, the project may still save time — it is unlikely to become a strategic asset.
Advantage can come from proprietary customer context, operating cadence, accumulated judgments, distribution, or a workflow that teaches the company faster than competitors can copy. Without one of those sources, the company is probably renting a commodity productivity gain. Price and govern that gain as a convenience. Do not mistake it for strategy.
4. Governance And Adoption
AI pilots fail for reasons that go well beyond model weakness. They fail when nobody owns the exception path, the review rule, the data scope, the user habit, or the kill rule. A pilot without governance becomes a demo that lingers in the stack.
Keep the adoption test concrete. Name the Tuesday-morning user, the allowed inputs, the actions they may take from the output, the reviewer, and the result that would force a stop.
Where The C-Suite Is Still Stuck
In boardrooms and founder meetings, AI can already produce useful work. The open question is where that work should enter the company's operating system. Three questions keep coming back.
- What counts as return on investment? A faster draft is visible. A better decision, avoided rework, or shorter sales cycle is harder to attribute. Many pilots never define the pre-AI baseline they claim to improve, so success becomes a mood rather than a measure.
- Who owns the changed workflow? Chief information officers can buy tools. Sales, finance, operations, legal, and product leaders own the habits that make those tools matter. Vague ownership turns adoption into optional enthusiasm.
- Where does autonomy stop? Leaders want capability and accountability together. Recommendation, execution, exception handling, and audit trail still lack a practical line.
Institutional friction is real. Procurement may sit with the chief information officer while the business unit absorbs process change. Legal and compliance own much of the downside. Finance wants a baseline that rarely exists. Strategic sacrifice lands with the chief executive when a workflow changes how customers, employees, or partners experience the firm. Slogans are plentiful. What is scarce is agreement on who owns measurement, exceptions, and failure before autonomy is approved.
The screen starts before tool selection. It forces the executive argument into the open. A serious AI pilot should sound like "we will reduce enterprise account qualification time from four hours to forty minutes, without letting the system invent relationship context or change pursuit priority without a human owner" — not like "we will use agents for sales."
Four Named Patterns
The public record is uneven. The cleanest numbers often come from vendor or company case studies; failures are less carefully documented. Four patterns remain useful enough for an executive screen.
Morgan Stanley: assist where trust and proprietary context matter
Morgan Stanley is the anchor case for assistance. Its AI @ Morgan Stanley Assistant and Debrief products sit in wealth management, a domain where wrong answers can damage trust and compliance. Public OpenAI materials emphasize evaluations, retrieval over a large internal corpus, regression testing, expert feedback, advisor review, and high adoption among advisor teams.
Screen guidance here is to assist knowledge work inside a controlled advisory workflow. Automating financial advice is the wrong frame. Advisors spend less time searching and summarizing, and more time preparing for client conversations. Advantage comes from proprietary intellectual capital and relationship context. Governance runs through evaluation-led deployment with humans finalizing outputs. Founders should copy that model before chasing autonomous agents.
Klarna: automate only where the routine is measurable
Klarna is the cleanest public support-automation case, and the one that most needs a source warning. Headline figures are company-reported through Klarna and OpenAI: millions of first-month conversations, a large full-time-agent workload equivalence, faster resolution, and a projected profit improvement. Treat those numbers as directional evidence without treating them as audited proof.
The executive lesson is still sharp. Customer support can be a candidate for narrow automation when intents are repeated, policies are retrievable, outcomes are measurable, and escalation rules are explicit. Leaders who treat containment as customer judgment create risk. The screen should pin down which promises the assistant may never make, which exceptions must escalate, and which quality signal would force a rollback.
BBVA and Moderna: scale starts as workflow redesign
BBVA and Moderna show the middle ground between personal productivity and autonomous execution. BBVA's 2026 OpenAI case describes a broad ChatGPT Enterprise rollout, senior-leader training, security, legal, and compliance alignment, and selected workflow efficiency gains. Moderna's enterprise case in OpenAI's 2025 report points to a narrower strategic workflow: Target Product Profile drafting and review, where AI helps extract facts and structure drafts so teams can pressure-test tradeoffs earlier.
Neither case treats software licenses as strategy on their own. Adoption becomes interesting when it attaches to named work, accountable teams, and decision quality. "Weekly active usage" is a weak executive metric by itself. Ask which meeting, handoff, risk review, or product decision changed.
Amazon: a stop rule is part of the strategy
Amazon's operations examples sharpen the "ignore" and "stop" side of the screen. The company has announced AI forecasting, mapping, robotics, and agentic tools for fulfillment operations. Project Eluna is framed as an assistant that helps teams make better decisions from facility data. Blue Jay, a robotics system announced in 2025, later carried an update saying Amazon was no longer using it in operations, while continuing to use underlying technology.
That public stop is disciplined automation. Physical operations expose demo limits faster than slideware does. Safety, throughput, space, reliability, and human workflow all have to clear the bar at once. Founders should read the update as permission to include a kill rule. Costly pilots are the ones that linger because nobody defined failure. Stopped pilots are cheaper.
Three Composite Calls
The following composites translate those named patterns into founder-scale decisions. Their value is in the contrast.
Case 1: The B2B founder who wants an account-research agent
The workflow looks attractive because it is repeated and painful. Before each sales call, the team gathers company news, funding history, executive changes, product signals, and likely pain points. A generic model can speed the work. Strategic value comes from joining public signals to the company's own win and loss notes and founder judgment.
Failure is subtle. Agents produce handsome briefs. Teams feel prepared. Nobody notices that the brief optimizes for visible facts over buying intent. Funding history, headcount growth, and executive quotes are easy to retrieve. Proprietary signal is harder — trigger events that led to closed revenue, objections that killed deals, founder-level hunches later confirmed.
Assist is the right call. Let AI prepare a briefing, suggest hypotheses, and surface contradictions. Humans still choose the opening question and decide whether the account deserves pursuit. Measure whether discovery improves; research that only looks nicer is the wrong target. Advantage lives in the firm's accumulating map of what actually predicts a good customer.
Case 2: The support leader who wants policy automation
This workflow is more bounded. Customers ask repeated questions. Policies exist. Exceptions can be escalated. Improvising promises the company cannot keep is the obvious danger.
Exception handling is the hard part; the first answer is easy. Systems must separate "the policy says no" from "this customer is about to churn, the policy is silent, and a human should trade short-term concession for lifetime value." Support automation disappoints when it treats policy retrieval as the whole job.
Narrow automation is the right call. Let AI retrieve and draft from approved policy, and constrain the action scope. When the answer touches refund discretion, legal exposure, security, or an angry high-value account, hand off rather than perform confidence. Track first-contact resolution, escalation quality, and promise violations. Ticket deflection alone is a weak scorecard.
Case 3: The CEO who wants pricing recommendations
Pricing feels like a tempting executive AI use case because it is analytical, commercial, and high gain. The model can summarize competitors and customer segments. Pricing is still a promise about positioning, margin, sales motion, and willingness to lose certain customers.
State-of-the-art models help mostly upstream. Clustering objections from call notes, comparing packaging language, finding inconsistencies in the sales deck, and drafting sensitivity scenarios are all fair uses. Strategic sacrifice embedded in a price stays with executives. Low prices may buy adoption and weaken enterprise credibility. High prices may improve margin and slow the learning loop. Usage-based models may align value and complicate procurement.
The right call is ignore or assist only. Use AI to assemble inputs and scenario ranges. Keep executive judgment on what kind of company the price is meant to create. The pilot is useful only if it makes the tradeoffs more legible.
Control Points
A useful executive screen separates four control points. The model may produce an output, recommend an action, execute an action, or learn from the result and change future behavior. Those are separate decisions.
Many companies jump from output to execution because the demo makes the path look continuous. Each control point changes the risk profile. Drafting a customer reply is one risk. Sending it is another. Applying a refund is another. Updating the policy because the refund worked is another again.
| Control point | Question | Common control |
|---|---|---|
| Output | Can AI draft, retrieve, classify, or compare safely? | Source citation, formatting rules, human review. |
| Recommendation | Can AI rank options without hiding the assumptions? | Confidence bands, counter-arguments, named owner. |
| Execution | Can AI act without creating unrecoverable harm? | Approval thresholds, escalation rules, audit trail. |
| Learning | Can the system improve without drifting from doctrine? | Feedback cadence, dataset review, kill rule. |
The Three Calls
The screen should force one of three decisions.
- Automate when the work is repeated, bounded, low-harm, easy to review, and owned by someone with authority to change the process.
- Assist when the workflow has high value and judgment, taste, risk, or relationship context must remain explicit.
- Ignore when context advantage is weak, the adoption path is poor, error is expensive, or no named business decision exists.
| Workflow | Likely call | Reason |
|---|---|---|
| Sales account research | Assist | It improves preparation speed, but a human still owns judgment and relationship context. |
| Customer support policy lookup | Automate narrowly | The work is repeated, bounded, reviewable, and can escalate exceptions. |
| Strategic pricing recommendation | Ignore or assist only | The consequence of a confident error is high and the context is commercially sensitive. |
A Public Worksheet Preview
The lightweight version is enough to reject bad pilots. Score each factor from one to five: decision weight, workflow repeatability, context advantage, error tolerance, adoption path, and learning-loop quality. A high score is not automatic permission to automate. It means the workflow deserves a sharper pilot design.
A Worked Example
Consider enterprise account research. On first pass it looks automatable: the inputs are public, the task repeats, and the output is easy to format. A lazy score might call it 22 out of 25 and approve an autonomous agent that researches, prioritizes, drafts outreach, and updates the customer relationship management system.
The honest score is lower. Decision weight is high, because better qualification can change sales focus. Repeatability is high. Context advantage depends on whether the system can use proprietary win and loss evidence rather than generic firmographics. Error tolerance is middling because a bad brief wastes an executive conversation. Ownership is often weak because sales ops, account executives, and founders all touch the workflow. Learning-loop quality is poor unless the company records which hypotheses actually predicted revenue.
| Factor | Naive score | Honest score | Why it changes |
|---|---|---|---|
| Decision leverage | 5 | 5 | Better qualification can change which accounts receive scarce founder or enterprise-sales attention. |
| Workflow repeatability | 5 | 4 | The research pattern repeats, but enterprise context varies by segment and buyer politics. |
| Context advantage | 4 | 3 | Public data is easy; proprietary buying-intent evidence may be sparse or unstructured. |
| Error tolerance | 4 | 2 | A confident but wrong brief can damage a high-value meeting. |
| Ownership and learning | 4 | 2 | Unless someone records outcomes and updates the rubric, the system learns little from reality. |
The call changes from automate to assist. Let AI draft the brief, cite sources, surface contradictions, and propose opening questions. Keep account priority, personalized claims, and customer-relationship-management writes under human review until the team has proved that the generated signals predict better discovery. Workflows can still earn a pilot. Autonomous versions should wait until those proofs exist.
Operator rule: name the business decision, accountable human, review control point, learning loop, and kill rule — or withhold pilot status.
That rule can feel slower than approving a tool. It prevents a costlier delay: three months of scattered pilots that leave the company with more software, more demos, and no better theory of where AI should change the business.
Source Notes
- Stanford HAI 2026 AI Index: broad AI adoption raises the premium on deciding where AI should enter operations.
- McKinsey State of AI Global Survey 2025: many organizations use AI, but scaling remains thinner than adoption.
- IBM CEO Study 2025: CEO-reported AI ambition sits beside return on investment, scaling, data quality, and leadership constraints.
- Anthropic Economic Index, September 2025: enterprise API usage skews toward automation patterns, especially in specialized business tasks.
- OpenAI / Morgan Stanley case: advisor-facing AI scaled through evals, retrieval, compliance controls, and human review.
- OpenAI / Klarna case: customer-service automation figures are company-reported directional evidence without audited proof.
- OpenAI / BBVA case and OpenAI Enterprise AI 2025 report: broad adoption becomes strategically relevant when tied to governed workflow redesign.
- Amazon operations AI and robotics update: agentic and robotics pilots show why stop rules belong in the investment screen.