Insights
Inside a frontier partnership: How enterprises should think about the seven levers of model progress
Adnan Masood, PhD | Chief AI Architect, UST
AI success isn't determined by the next model release. It comes from mastering seven interconnected levers that turn AI progress into measurable business outcomes, competitive advantage, and sustained enterprise value.
Adnan Masood, PhD | Chief AI Architect, UST
Enterprise AI strategy fails most often at the level of mental model. Leadership teams track model announcements the way they once tracked competitor earnings, treating each release as a discrete event to be evaluated, piloted, and priced. But model progress is not a sequence of events. It is a system with seven distinct levers: compute, data, architecture, distillation, reinforcement learning, harness, and product. Leaders who understand how these levers interact make better decisions about budgets, talent, vendor relationships, and timing. Leaders who see only the headline release make expensive mistakes on all four. After a decade of leading AI programs across banking, insurance, and healthcare, I have come to view this framework as the difference between companies that compound AI gains and companies that perpetually pilot.
DIVIDER
The framework at a glance
DIVIDER
01 Compute
Raw computational scale remains the primary engine of model improvement, and it is now spent in three places: pretraining, post-training, and inference inside the product itself. The third category deserves a CFO’s direct attention. Modern models spend variable amounts of reasoning on each task, which means an identical request can cost five cents or five dollars depending on how much thinking is authorized. Most enterprises still budget AI as if it were per-seat software with a predictable slope. It behaves more like cloud egress a decade ago: metered, invisible in the moment, and eventually the subject of a difficult finance review.
The discipline that works is measuring cost per business outcome rather than cost per API call. Cost per resolved claim. Cost per modernized COBOL module. Cost per closed ticket. Token economics become manageable only at that grain. UST’s token commitment within its Anthropic partnership reflects this logic: committing real volume forces an organization to measure consumption properly, and measurement is where inference economics stop being a surprise.
DIVIDER
02 Data
Synthetic data has become the growth dimension of model training, with human data serving as its seed. The implication that matters for enterprises is quieter than the headline. Models increasingly learn inside realistic environments, simulated versions of real work where an agent can attempt tasks, fail safely, and improve. An environment for a claims adjudication agent is a sandboxed replica of the policy administration system, seeded with genuine edge cases and scored by criteria a domain expert will sign.
No vendor sells that environment, because it encodes the buyer’s business rules and decades of judgment that live partly in a mainframe and partly in the heads of a few people nearing retirement. Proprietary data is a seed. The environment built around it is the moat, and building it is a leadership decision about where institutional knowledge gets formalized before it walks out the door.
DIVIDER
03 Architecture
Model architecture is now designed for quality, inference cost, and context length simultaneously, and the pattern with the largest commercial consequence is this: capability keeps getting cheaper after it stops getting better. Whatever premium model a company’s hardest workload requires today, a lower-cost tier will absorb that workload on a fairly predictable schedule.
Model routing decisions therefore carry a shelf life measured in months, which argues for a standing evaluation capability rather than an annual bake-off. UST-Eval was built on this premise: when a new Claude tier ships, a client’s benchmark suite gets rerun and produces a defensible routing answer quickly, because the savings compound and the window to capture them is short. Organizations that hardcode a model name into two hundred call sites pay a tax they have not yet noticed.
DIVIDER
04 Distillation
New models are increasingly built from previous models rather than from scratch, with specialized teacher models combined into a single student that inherits their capabilities at lower cost. Two findings from this work deserve an executive audience. A stronger model does not automatically teach better. And the question of how to combine multiple teachers remains unsolved.
Enterprises face the identical problem with human expertise. Which underwriter, which engineer, which compliance officer should an AI system learn from, in what sequence, without one perspective overwriting another? If teacher combination is an open problem for model weights, it is entirely unaddressed in most corporate knowledge architectures. The firms that solve it for their own institutional knowledge will hold an advantage that no model release can erase.
DIVIDER
05 Reinforcement learning
Reinforcement learning remains the most powerful tool for shaping model behavior and the slowest part of the improvement loop. Weight updates tend to be incremental, and lightweight adaptation frequently covers whatever change was actually required. This finding functions as a warning label for the custom fine-tuning program in many FY27 plans. Most of what enterprises seek from fine-tuning is available through better context, better tools, and better evaluation, at a fraction of the cost and none of the maintenance burden. In my experience, roadmaps redirected on this point do not ask to go back.
DIVIDER
06 Harness
The harness is everything wrapped around the model: search, tools, context, memory, orchestration, verification. Improving the harness frequently outperforms upgrading the model, and yet organizations routinely misattribute the gain. When a team reports that a new model lifted accuracy from 71 percent to 88 percent, one question separates real signal from noise: did anything else change? Almost always, someone rewrote the retrieval strategy or the system prompt in the same release. Seven-figure procurement decisions have rested on comparisons where the harness changed underneath the model, with the vendor accepting the credit gladly.
The harness is also where durability lives. Each new model absorbs some of what was built the year before, and each new class of work demands harness that does not yet exist. That cycle rewards companies treating agentic harness engineering as a permanent discipline. It is the reasoning behind UST’s decision to make its AlphaAI platforms, led by CodeCrafter, Claude-native, and to certify engineers through Anthropic’s forward-deployed engineering program. Model weights are rented. The harness is owned.
DIVIDER
07 Product
Usability regularly matters more than benchmark performance, and real interaction data, evidence of how people actually work, is the strongest single input to improvement. The cycle is unforgiving: deploy at scale, capture real usage, evaluate, improve, deploy again. A system that never reaches real users generates no usage data, therefore no evaluation signal, therefore no improvement, therefore no case for expansion, therefore another pilot.
This is the capability overhang in its natural habitat: the model is ready, the organization is not, and the gap gets billed to the technology. The eleven-month pilot is a leadership failure wearing a technical costume.
DIVIDER
The lever that runs through all seven
Human involvement is not a transition cost on the way to automation. It is a permanent input to every lever above. People define the evaluations that encode what a business values. People supply the seed data worth learning from. People build the applications that widen what a model can improve. And people ensure the whole system remains aligned with what a regulator, a board, and a customer would recognize as legitimate. Alignment at the frontier is the lab’s responsibility, and it is central to why UST partnered with Anthropic. Alignment with an enterprise’s own obligations belongs to the enterprise, which is why governance must sit inside the delivery model rather than beside it.
Leaders who manage all seven as a portfolio, rather than waiting on any single one, are the ones for whom each model release compounds instead of merely arriving.
Stop reacting to AI. Start orchestrating it.
The enterprises creating lasting AI advantage are managing the systems behind model progress, not just adopting the latest models. Discover how UST AlphaAI helps organizations build scalable AI foundations, accelerate deployment, and capture measurable business value.
DIVIDER
References and further reading
Anthropic. "Building Effective AI Agents." — distinguishes workflows, where models and tools follow predefined code paths, from agents, where models direct their own tool use. The practical starting point for harness design.
Anthropic. "Effective Context Engineering for AI Agents." — frames the discipline as curating the optimal set of tokens during inference rather than optimizing prompt wording.
Anthropic. "Lessons on Building Effective Human-Agent Teams." — covers verification patterns, task checklists, and how autonomy gets extended per task type after repeated success. Useful for the human involvement section.
MIT Project NANDA. The GenAI Divide: State of AI in Business 2025. Coverage: finds that roughly 95 percent of enterprise pilots stall with little measurable P&L impact, and attributes the gap to a learning and integration problem rather than model quality. The empirical backing for the capability overhang argument.
Anthropic Economic Index. — ongoing measurement of real deployment patterns. The September 2025 chapter is the most relevant to enterprise buyers: it examines enterprise API use and finds that capability, ease of deployment, and economic value drive adoption more than cost per task.
Anthropic Economic Index report, June 2026. — the most recent chapter, including survey work on how people experience AI at work.