Every enterprise IT function is chartered around the same objective: enable peak business performance through the application and infrastructure layers, bound together by integrations, process, and the right skills. That's the job description, whether or not it's written down that way.

So when I run statistical clustering — agglomerative clustering, specifically — against a portfolio's ticket history, I'm not looking for automation candidates. I'm looking for what's standing between the operation and that peak state. And what comes out the other side is almost embarrassingly legible.

The clusters aren't subtle. No-action patterns, closed by a monitoring tool or event manager but still assigned to a human anyway. Identity and access management — password resets, account lockouts, access requests — sitting on top of an MFA and conditional-access stack the organization already paid for. End-user productivity tickets — software installs, VPN, VDI — that are really a story about self-service maturity, not incident volume. None of this requires exotic analysis to find. It requires someone willing to look at the pattern and ask why it still exists.

Here's the part that should bother every enterprise more than it seems to. I've watched this same intelligence surface for the same customer across multiple years. Not a one-time finding buried in a report nobody read — a recurring signal, labeled for elimination, present on the transformation backlog, and never prioritized. That's not a data problem. That's an organization that has made peace with paying for work it has already proven doesn't need to exist.

Elimination is not automation

Before going further, the two words need to stop being used interchangeably, because the difference is the entire argument.

Elimination is a one-time program that permanently removes a workload and gets absorbed into the platform itself. It has no steady state. Automation is a recurring activity — it still runs the workload, every time, just faster and with less human effort. One makes the ticket stop existing. The other makes the ticket cheaper to keep having.

Elimination vs automation — what actually differs
Dimension Elimination Automation
Nature of the work One-time program Recurring activity
What happens to the workload Removed permanently, absorbed into the platform Still runs, every time, just faster
Ongoing cost None — no steady-state execution Continues indefinitely — compute, licensing, tokens
What it fixes The root cause the workload exists for The friction of executing the workload
End state The ticket stops existing The ticket exists, faster

Every cluster in this article is being evaluated against that distinction, not against how much faster GenAI can execute it.

Three clusters, three different jobs

Not every eliminable pattern gets eliminated the same way. The clustering sorts cleanly into three categories, and conflating them is where most transformation programs stall out.

Pure elimination — the insurance-policy ticket. Event-driven and monitoring tickets where a human is only there to double-check that the monitoring tool got the threshold right. That's not triage. That's an insurance policy against the tool being wrong, executed manually, forever. The fix isn't a smarter human workflow — it's handing the verification itself back to the monitoring and event-management layer: correct the thresholds, correct the wait-time logic, remove the human checkpoint. These are close to cookie-cutter across organizations, which means a GenAI layer can drive the transformation with minimal human involvement — not by generating a ticket response, but by generating and validating the configuration change that makes the ticket stop existing.

Assisted elimination — identity and access. The root cause here usually isn't a missing capability. It's an org with MFA, Windows Hello, and conditional access already deployed, still generating lockout and reset volume because end-user contact data is stale, or automation can't handle the exception path — a password reset that trips a policy restriction and lands the user in a lockout instead. GenAI can build and execute the fix here too, but this tier needs more human involvement: verification, exception testing, and end-user willingness to actually adopt the corrected flow. Still a one-time program. Just not a zero-touch one.

Empowerment, not automation — the productivity layer. Software installs, VPN, VDI. The clustering here doesn't point to an automation opportunity at all — it points to a self-service and configuration gap. Config-assisted self-service portals, control-panel-level fixes, device-level configuration rollouts. Human-assisted to build, but one-time. Once it's built, there's no steady-state workload left to automate, which is precisely the point elimination is making that automation can't: the goal was never to do the work faster. It was to make the work stop recurring.

The three clusters and how each one gets removed
Cluster Pattern Root cause Who does the work Outcome
Pure elimination Event-driven / monitoring tickets, human as insurance policy Verification checkpoint, not a real judgment call GenAI generates and validates the config change, minimal human involvement Human checkpoint removed entirely
Assisted elimination IAM — password resets, lockouts, access requests Stale contact data, unhandled exception paths, not a missing capability GenAI builds the fix; humans verify, test exceptions, drive adoption One-time program, not zero-touch
Empowerment Software installs, VPN, VDI Self-service and configuration gap, not incident volume Human-assisted rollout of self-service portals / device config No steady-state workload left to automate

What elimination actually clears the way for

This is the piece that ties back to peak business performance directly, and it's the reason I don't treat elimination as a cost exercise.

Once you remove these three clusters from the operational picture, what's left is a much smaller, much more meaningful set of events — and that smaller set is exactly what real observability was supposed to be built for in the first place. Convert what remains into metric streams. Layer anomaly detection on top. Correlate those anomalies to upstream business process performance, not just infrastructure health. Run real-time root cause identification and triage on that smaller signal — and you're no longer reacting to incidents, you're avoiding them before they touch the business process they were always going to hit.

Eliminate the three clusters No-action / monitoring, IAM, end-user productivity
A smaller, meaningful event set Converted into metric streams with anomaly detection layered on top
Correlation to business process performance Real-time root cause identification and triage on the remaining signal
Incident avoidance The business process is never touched in the first place

That's not a new idea for me — it's the same incident-avoidance thesis I was building a decade ago, before "AIOps" was a category. What's changed is that GenAI now makes the surrounding work — the elimination programs that have to happen first — tractable at a fraction of the cost and time. The subject-matter-expertise gap that used to stall these projects for years is largely gone. What's needed now is human assistance to verify and test, and end-user willingness to adopt — both solvable, both far cheaper than the alternative.

And this is where I'd draw the line most enterprises are currently missing: that root-cause, business-process-correlated observability program, and the end-user empowerment program, are usually buried underneath the automation and agentic backlog — not because they're less valuable, but because they're less visible, and because "add an agent" has become the reflexive answer to any ticket volume problem, eliminable or not.

Why the finding sits on the backlog for years

If elimination is this visible, why does it stay unprioritized across multiple contract cycles for the same customer? Two things have to line up before an organization actually acts on it, and most of the time, neither does.

The first is a triggering event. Elimination doesn't happen inside the steady state of a contract — it happens at the moments when the operating model itself is up for renegotiation: contract renewals, RFPs, and M&A. Renewals and RFPs open a window because the commercial construct is being rewritten anyway, and elimination can be built into the new terms instead of fought for inside the old ones. M&A is the stronger case of the two, but specifically when the outcome is insourcing — bringing operations back in-house removes the party whose commercial interest runs counter to elimination in the first place. Absent one of these events, elimination is asking an incumbent operating model to argue for its own shrinkage, mid-contract, with no mechanism forcing the conversation.

The second is what the contract actually measures. Look at how managed services contracts are structured, especially outsourced ones: ARC (Annual Reduction Commitment), RRC, and CSI clauses, all of it ultimately anchored to money — a cost reduction percentage, a rate card adjustment, a fee-at-risk number. None of it is anchored to the service being consumed, and none of it is anchored to business outcome. That's a strong, structural link, and it's precisely why elimination doesn't get incentivized under these constructs. A provider committing to an annual cost reduction has no commercial reason to eliminate a ticket cluster that's contributing to volume-based revenue, and a CSI program built around arbitrary annual improvement targets rewards incremental tinkering, not structural removal of the work. The contract is measuring the wrong thing, so it's optimizing for the wrong thing.

This is where the budget conversation and the elimination conversation meet. IT budgets have been shrinking for several cycles now, and that trend isn't reversing. Under an ARC/RRC/CSI construct anchored to cost, a shrinking budget just means doing the same lights-on work for less — there's no mechanism in that model for the budget to move toward better business performance, because the model was never measuring business performance to begin with.

My proposition is straightforward: replace the cost-anchored target with an incremental business-outcome target, and the entire budget equation changes. Instead of negotiating how much cheaper the lights-on operation gets next year, the negotiation becomes how much measurable business performance improvement the operating model delivers — and elimination is the only credible mechanism to fund that shift, because it's the one lever that converts operating spend into freed capacity without asking for new investment. A budget conversation anchored to business outcome, not cost reduction, is what finally gives elimination a commercial reason to happen inside the contract instead of waiting for the contract to end.

The agent rental market is chasing the wrong backlog

I recently sat through a SaaS vendor pitch offering 600 pre-built agents to handle enterprise ticket volume. I didn't see a case for most of them — not because agents aren't capable, but because a meaningful share of what they're being pitched to handle is exactly the clustered work above: work that shouldn't exist, sitting on platforms the customer already bought, waiting on a one-time program that's been sitting on the backlog for years.

Renting an agent to triage a ticket that a corrected monitoring threshold would have prevented isn't automation. It's a subscription on top of a gap you already knew how to close. I expect the ROI conversation on a lot of these deployments to look very different after a couple of rounds of POCs, once someone on the buying side asks the obvious question: why are we paying — in license fees and in tokens — for a workload we already had the intelligence to eliminate?

The path to peak business performance isn't agent-led. It's elimination-led, with agents and automation earning their place only on what's left once the eliminable work is gone.

Key Takeaways

  • Elimination is a one-time program with no steady state — the workload is removed and absorbed into the platform. Automation is a recurring activity that still runs the workload, just faster. One makes the ticket stop existing; the other makes it cheaper to keep having.
  • Agglomerative clustering against ticket history surfaces the same three patterns repeatedly: no-action monitoring tickets, IAM volume sitting on top of an already-paid-for MFA and conditional-access stack, and end-user productivity tickets that are really a self-service maturity gap.
  • Each cluster needs a different job — pure elimination driven largely by GenAI-generated config change, assisted elimination that needs human verification and end-user adoption, and empowerment through self-service rather than automation at all.
  • Removing those clusters is what makes real observability possible: a smaller event set converted into metric streams, anomalies correlated to business process performance, and incidents avoided instead of resolved faster.
  • The finding sits on the backlog for years because elimination needs a triggering event — renewal, RFP, or M&A with insourcing — and because ARC/RRC/CSI constructs anchor to cost, not to business outcome. Replace the cost-anchored target with an incremental business-outcome target and elimination finally has a commercial reason to happen.
  • Renting agents to handle eliminable work is a subscription on top of a gap you already knew how to close. The path to peak business performance is elimination-led, with agents earning their place on what's left.

Part of an ongoing series on AI-native IT operations. Read the full framework at murthymalapaka.com/insights.