Buying AI used to mean picking a supplier. You signed for a model, handed out access, and that was the strategy. That world is gone. Buyers now face frontier proprietary systems, open-weight models, managed cloud services, local installs, specialized agents, and suite copilots — each with different economics, oversight load, update cycles, and failure modes. Treating them as interchangeable is how you get expensive default everywhere and real control nowhere.

The open-model debate doesn't help much at the purchasing desk. One camp frames open weights as the path to competition and national autonomy; the other frames them as uncontrolled capability with security baggage. Both have substance. Neither gives you a rule for what to put on which workflow next Tuesday.

What you actually need is a model mix table: a plain record that assigns each task the capability, oversight, and accountability it deserves. Stanford Human-Centered AI Institute (HAI)'s 2026 AI Index shows the performance gap between leading U.S. and Chinese systems shrinking, agentic tools moving fast, and transparency for the strongest models still patchy. The Organisation for Economic Co-operation and Development (OECD) has pointed out that "open source" doesn't transfer cleanly from ordinary software to foundation models — openness comes in degrees. The National Telecommunications and Information Administration (NTIA) report on open weights lands in a similar place: open models have real advantages, but governments and buyers still need monitoring and the ability to respond when risks shift. Organizations have to make the same call internally.

The strategic mistake is letting one model category quietly become the default for every kind of work.

The New Buying Problem

For a founder or enterprise buyer, model choice is a portfolio problem. A frontier API often fits work where reasoning depth, tool connections, vendor help, and fast upgrades matter more than unit cost. An open-weight model often fits work where data location, customization, offline operation, cost stability, or legal jurisdiction matter more than peak score. A bundled copilot fits work that already lives in email, documents, code repos, CRM, or support software. A smaller specialist model fits narrow, repetitive jobs where "good enough" is actually enough.

Those assignments should move as prices fall, benchmarks flatten, internal data improves, rules clarify, and failure evidence piles up. Independent evaluators like Artificial Analysis show why annual procurement cycles are a bad fit: quality, price, latency, context length, and open-weight availability keep shifting. A model that looked inadequate last quarter can be fine now. One that looked cheap at the API layer can get expensive once review effort, retries, context volume, and agent loops show up on the real cost table.

The executive job is to stop model use from becoming unmanaged infrastructure. Every AI workflow should have an explicit class, owner, rationale, fallback, review load, and exit condition — written down, not assumed.

What to track

A useful model mix table groups work by decision rights, not by vendor logo.

Work Class Likely Model Choice Control Test False Positive
Frontier judgment support Frontier API or managed enterprise model Does quality matter enough to justify premium cost and vendor dependency? Using the best model because the task feels important but has no measured quality threshold.
Repeatable internal production Open-weight, hosted open model, or smaller tuned model Can the company govern data, cost, latency, and failure recovery better with more control? Assuming self-hosting is cheaper while ignoring engineering, security, and evaluation work.
Office and workflow assistance Suite copilot or application-native AI Does the AI live where the work record, permissions, and audit trail already live? Buying convenience that creates invisible shadow work and unclear responsibility.
Regulated or sensitive work Constrained deployment with explicit assurance and human checkpoints Can the buyer prove provenance, access control, monitoring, and rollback? Treating a vendor security page as a substitute for workflow-level assurance.
Low-value commodity drafting Cheapest adequate model or no model Does the output save enough time after review to justify any AI path? Automating marginal work because the marginal token price looks low.

Open Is Not One Thing

"Open-source AI" is a bag of different claims. A model might disclose code, weights, architecture, training data, evaluations, prompts, or deployment steps — or only some of those. Research use might be allowed while commercial use is barred. Download can be easy while running at scale is hard. Experimentation can be easy while data sources, safety tests, and training provenance stay murky.

That's why the terminology warning actually matters. Software conventions don't map onto foundation models. Source code can be inspected, changed, and redistributed under known rules. Model "openness" often means you get weights without a full story of how they were produced, what data shaped them, or how behavior will shift after later fine-tunes.

Score openness by purpose, not as a binary badge. The useful question is whether the access you get is enough for the job you intend to run.

The Sovereignty Temptation

Sovereignty talk has real weight now. Governments and large organizations want more say over data flows, domestic capability, regulatory compliance, and supply dependence. Linux Foundation research on sovereign AI notes that open source often gets cast as a foundation for autonomy because it offers flexibility and visibility. Cross-border organizations should take that signal seriously rather than wave it off as politics.

Even so, sovereignty can turn into theater. An organization can claim the goal while still depending on foreign chips, clouds, advisers, and model updates, with a thin internal team that can't really assess or sustain the stack. Local hosting alone isn't sovereignty. Open weights alone aren't either. You need the ability to run, examine, update, secure, and exit the system without unacceptable reliance on someone else.

Smaller companies usually don't need a national posture. They need to move data when they must, set control points, manage suppliers, and keep costs from drifting.

The Cost Trap

Open-weight models can cut the sticker price of capability. Cost doesn't disappear — it moves. Per-token spend becomes hosting, GPUs, engineering time, monitoring, evaluation, compliance, and incident handling. Sometimes that trade is a win. Sometimes it's just duplicated infrastructure you didn't need.

A cost table should track six lines: model access, compute, integration, evaluation, review effort, and switching cost. When those stay hidden, buyers compare brochure prices instead of operating reality.

Frontier providers still have a clean argument in many cases. A closed model can be rational when the vendor carries the infrastructure load, ships updates quickly, offers enterprise controls, and gives you a real support path. Closed doesn't mean worse, and open doesn't mean cheap. The test is whether the control you gain is worth the operational load you take on.

The Governance Test

The National Institute of Standards and Technology (NIST) AI Risk Management Framework and Generative AI Profile give a simple sequence: govern, map, measure, manage. Applied to model selection, every production workflow needs a short record of its own.

  1. Name the work this model performs.
  2. State why this model class beats a cheaper, more controlled, or more capable alternative.
  3. List data allowed into the model and data that is forbidden.
  4. Set the quality threshold that justifies production use.
  5. Measure the review burden the workflow creates.
  6. Define the failure that triggers rollback, replacement, or human-only operation.
  7. Name the owner of the next model-mix review.

That last item is the one people skip. Model assignments go stale. A sound choice in January can be waste by July. Without a scheduled check, the mix hardens into a fossil of earlier assumptions.

The Executive Move

Don't accept a single-model plan from the AI team. Ask for a model mix table that covers the ten workflows where AI use actually matters. For each one, require model class, rationale, cost line, control line, quality threshold, data scope, review load, fallback, and next review date.

Then look at the shape of the portfolio. Frontier on every workflow means capability without cost discipline. Open-weight everywhere can mean control standing in for quality or support. Suite copilots on everything can mean the application vendor quietly owns the record of future work. And if no work category is marked "no model," governance is incomplete.

Strong buyers treat model selection like portfolio management. Frontier capability goes where it changes judgment. Open and smaller models cover work where control and economics dominate. Vendor bundles stay where the workflow record already lives. AI gets declined where review effort exceeds the return.

The whole doctrine is three observations held at once: capability is rented, control has to be maintained, and default assignments should get reviewed on purpose — not only when something breaks.

Source Notes