VentureBeat ran a piece this week on TypeSafe's Jev that deserves attention from every CIO managing an AI budget. The core argument is that most AI inside enterprise software isn't writing anything. It's deciding: which queue a ticket belongs in, whether a document is relevant, whether a transaction looks wrong. And those decisions are often made by having a language model generate text that then gets parsed back into a label.

I agree completely. I'd take it further, because we learned this the hard way while building an AI-native IT operations platform.

Operations was always a decision problem

When a ticket arrives, the questions are simple. What class is it? Which configuration item? What's the cause? Should this work be eliminated, automated or assisted? None of these needs an essay.

Our journey started six years ago with clustering and classification. We grouped tickets into recurring patterns and trained classifiers to assign dispositions, working through models like BERT and XGBoost along the way. There was no LLM anywhere in the stack.

So when the article argues that engineers who built classifiers, cascades and evaluation pipelines before 2022 are relevant again, that describes our path exactly.

The decision signal isn't only text

One caveat in the piece matters more than it seems: Jev is in early access and accepts text only.

In operations, the signal for a decision is split. Telemetry tells you what the system is doing. The ticket tells you what a human noticed. A label drawn from text alone is a guess. The same label, informed by the metric stream behind it, is a diagnosis.

That's why our platform fuses numeric and textual data. Numeric features from telemetry feed the classifiers alongside what's extracted from the ticket text, and both anchor to a common ontology of machine, method and model.

Where we did use LLMs, and why that isn't a contradiction

We did use language models to reach dispositions, through prompt engineering and retrieval. What made it work was how tightly the model was constrained.

We moved from ticket-level clustering to record-by-record extraction constrained by a taxonomy, using a small open-weight model. The vocabularies were closed: a three-class top-level taxonomy, about 19 service-request categories, and roughly 22 action verbs. The model wasn't writing anything. It was choosing from a fixed set, and the harness around it enforced determinism. In practice it was a classifier wearing generative clothing.

Just as important, this ran in batch, in an offline intelligence-generation layer. We paid for tokens once, at pattern-discovery time, not on every runtime decision.

The real goal: take the model out of the runtime path

This is where I'd push the Jev conversation forward. A cheaper decision model is progress. A decision that no longer needs a model at all is better.

Our design lets patterns graduate:

  • Discover. Probabilistic extraction surfaces candidate patterns.
  • Confirm. A human validates resolution time and hop count across enough instances.
  • Graduate. The pattern becomes a deterministic, vetted playbook, executed through a secure execution gateway with no language model in the runtime path.
  • Demote. A demotion path catches drift when patches, re-architected services or reorganized assignment groups break the rule.

Language models never generate commands. The cheapest decision is one that no longer needs a model. The cheapest ticket is one that never gets created.

Calibration is where humans belong

The article rightly stresses calibrated confidence. If a system can't tell when it's likely to be wrong, it can't reliably decide when to hand a case to a human.

We treated human assistance as a first-class tool the agent can call, not an out-of-band interrupt. Every human-assurance event also feeds back into the system. It reveals the gaps in what the system can see: missing log files for a failed batch job, change-request data behind an end-user issue, or absent monitoring for a performance problem. The escalation itself becomes training signal.

Route by size, structure and time

The article describes cascades in which a cheap classifier decides whether a request goes to a small local model, a specialist, or a frontier API. That matches what we built.

Playbook lookup runs in three tiers: deterministic first, then structured, then retrieval. A small model handles extraction, and a larger one is reserved for judging novel incidents. Heavy discovery runs in batch, while only a thin real-time slice supports execution. Work is routed by model size, data structure and timing, much like a distributed computing system, instead of sending everything down one expensive path.

Before you automate a decision, ask whether it should exist

This week HCLSoftware announced it will acquire Robotiq.ai, an RPA platform, to give its agentic orchestration layer the ability to act inside applications that lack APIs. Technically, that's a sensible gap to fill. Agents need hands, and plenty of enterprise systems still offer only a screen.

But the combination should give leaders pause. RPA automates a process exactly as it exists today. Putting an agent on top makes the automation smarter without asking whether the work is legitimate in the first place. An intelligent decision layer driving scripted execution against a task that shouldn't exist only produces waste faster.

In our platform, every pattern was first tested against three dispositions: eliminate, automate, or assist. Elimination came first because it's the only option that permanently removes cost. A cheaper label, a smarter bot and a faster agent all still assume the work should happen. Intelligence should start by questioning that assumption.

What I'd tell a CIO this quarter

  • Question legitimacy first. Before automating any process or decision, ask whether it should exist. Only the business holds that decision right.
  • Audit your LLM traffic. Separate decisions from generation.
  • Fuse the data where the signal lives. For operations, that means telemetry plus text, not text alone.
  • Design a graduation path. Move proven decisions out of models and into deterministic execution.
  • Own the models. Open-weight models tuned to your operating model keep the intelligence you generate as your IP.

The result for us was that we onboarded more customers year over year while keeping people and technology costs flat. Jev confirms the direction the industry is taking. The next step is a decision layer that learns its way out of the loop, and starts by asking whether the work should exist at all.

Part of an ongoing series on AI-native operations and building intelligence that drives peak business performance.