Organizations don't actually start AI-native IT operations transformation by picking a model. In practice, the starting point is almost always one of three things: extend the script-based or runbook automation already in place into something agentic, adopt a SaaS vendor's agent library or a platform like ServiceNow, or — for a smaller subset — buy enterprise licenses of a frontier model and start pointing it at operations problems.

Each of those is a reasonable instinct. None of them, on their own, is an architecture. And without clear objectives and guardrails behind that instinct, the result is a familiar pattern: a shadow-IT problem, the same way shadow cloud adoption played out a decade ago — capability spreading faster than governance, spend accumulating without a shared measure of what it's actually buying. Token spend without an outcome architecture behind it is exactly that pattern, one layer up the stack.

This is a reference architecture built to close that gap — from having actually run it, not designed it on a whiteboard.

Start by measuring what you have, not what you're buying

The fix isn't a better vendor choice. It's sequencing: build intelligence about your current state before deciding what to automate. That first pass is usually uncomfortable. It shows, in hard numbers, how labor-intensive current operations really are, and what that labor intensity is actually costing in business outcomes: MTTR, end-user experience, the ratio of human effort to resolved value.

That current-state intelligence is what should define your goals and objectives for the future state — not the other way around. Skip it, and you're back to the shadow-IT pattern: automation tooling adopted because it's available, not because it's aimed at a measured gap. With it, every subsequent claim — every dollar saved, every ticket eliminated, every hour of SME time freed up — is measured against a number you established before you changed anything.

The shift the architecture is built to produce

Once current-state intelligence defines the target, the goal is to generate actionable future-state intelligence — the kind that drives a specific structural shift in how work gets done:

A labor-driven execution layer → an SME-driven human-assurance layer.

That's not a rebrand of the same team. It's a change in what the humans are actually for. In the labor-driven state, people execute — they work tickets, they run playbooks, they are the throughput. In the target state, the execution layer is elimination-led agentic AI: it removes root causes first, automates what's left, and executes autonomously. The humans left in the loop are SMEs providing assurance — judgment on the cases the system flags as genuinely uncertain, not throughput on the cases it already knows how to close.

That shift only holds together if there's a closed loop back to intelligence generation. Every SME-assurance event isn't just a resolved ticket — it's a labeled signal about a gap: missing evidence, missing correlated data, missing instrumentation. That signal retrains the intelligence layer, which improves what the execution layer can handle autonomously next time, which further shrinks what SMEs need to touch. The loop compounds. It doesn't plateau after the first automation win — and it doesn't sprawl the way ungoverned agent adoption does either, because every gap it finds routes back to the same intelligence layer instead of spinning up another disconnected tool.

Current-state intelligence Measure labor intensity, MTTR, and effort-to-value before automating anything
Future-state objectives The measured gap defines the target — not a vendor template
Elimination-led execution Agentic AI removes root causes, automates what's left, executes autonomously
SME assurance Humans judge only the genuinely uncertain cases

↻ every SME-assurance event feeds back into current-state intelligence

Why the outcome targets have to be specific

Vague AI outcomes ("efficiency," "faster resolution") don't survive contact with a Board asking for token-level ROI. The outcome targets in this architecture are deliberately narrow and measurable:

  • Do not escalate to a human until absolutely necessary — escalation is the exception, not the workflow.
  • When escalation does happen, it's not a hand-off — feedback modifies the agentic layer and gets executed, closing the loop rather than parking it.
  • MTTR compression for proactive, automatable tickets.
  • End-user delight for reactive incidents and self-service-portal requests, where experience matters as much as resolution.
  • Business outcomes as the outer measure everything else has to roll up into.

Because a human is removed from the resolution path for automatable work, MTTR stops being a distribution you report after the fact and becomes math you can state in advance: agent execution time plus any workflow approval wait, both known quantities. That's the difference between "we're fast" and "we can tell you, before we act, how many minutes this will take." One is a marketing claim. The other is a guarantee the token spend is earning its keep.

Why this only runs on open-weight models — and never off the shelf

Two things have to both be true for this loop to work, and neither is available from a frontier model behind a shared API, and neither is available from a generic SaaS agent library either.

First, the models have to be trained and tuned to your operating model and your business-outcome requirements — not a generic industry template. The intelligence-generation model and the execution model are each purpose-built: one toward the outcome being pursued, the other toward execution and the contextual feedback that trains the first. That level of specificity requires actually training and modifying the model, not prompting a shared one. It also means this is never a one-size-fits-all deployment — the same architecture produces a different tuned system for every organization it's built for.

Second, every retraining cycle makes the resulting intelligence more valuable and more specific to the organization that built it. That accumulated intelligence is IP. Run the same loop through a frontier-model API or a licensed vendor platform and you're training a model you don't own, on infrastructure you don't control, and handing the accumulated value to someone else's platform.

Open-weight, locally hosted models aren't a cost-saving layer bolted onto this architecture. They're the only way the loop — and the IP it generates — can belong to the organization running it.

The economics, stated plainly

None of this pencils out as a pricing model if the system is static. A one-time-trained automation layer saves money once and then drifts as the environment changes around it. A system built around continuous current-state-to-future-state intelligence generation doesn't drift — it compounds, because every gap it finds becomes a training signal instead of a standing exception.

That's what makes outcome-based, token-economics pricing viable instead of aspirational: you're not billing for hours because the entire design goal is needing fewer of them over time, and you're not billing for raw compute because the token spend is tied directly to a declared, measurable outcome. The reference architecture — current-state intelligence defining future-state objectives, elimination-led execution, SME assurance, and the feedback loop connecting all of it — is what gives token investment a real chance of returning the business outcome the Board is actually looking for.

Key Takeaways

  • Extending existing automation, adopting a vendor's agent library, or licensing a frontier model are reasonable instincts — but none of them is an architecture, and without one, token spend sprawls the same way shadow cloud adoption once did.
  • Sequencing matters: current-state intelligence — labor intensity, MTTR, effort-to-value — has to define future-state objectives, not the other way around.
  • The target structural shift is from a labor-driven execution layer to an SME-driven human-assurance layer, where elimination-led agentic AI executes and humans provide judgment only on genuinely uncertain cases.
  • Every SME-assurance event is a labeled signal that retrains the intelligence layer, compounding capability instead of plateauing after the first automation win.
  • The loop only produces owned IP on open-weight, locally hosted models trained on your own operating model — a frontier API or licensed platform hands the accumulated value to someone else's infrastructure.
  • Outcome-based, token-economics pricing only works when the system compounds rather than drifts — which is what makes MTTR a number you can state in advance, not just report after the fact.

This is part of an ongoing series on AI-native IT operations. Read the full framework at murthymalapaka.com/insights.