A recent HackerNoon piece describes large language models as an operating system. In the old world, every task needed its own model: one for sentiment, one for entity recognition, one for classification. In the new world, one foundation model is the runtime, adapted to a domain through fine-tuning and retrieval.

I agree with the arc, because I lived it. But an enterprise doesn't run on an operating system. It runs on answers to a handful of questions that decide performance. Why is this claim taking so long? Why is yield low on this line? What is causing the waste?

Last week, in Generative AI Isn't Over. Enterprise Intelligence Hasn't Started., I argued that enterprise intelligence deserves its own track, and that it's the one barely moving. This piece is about how we built it, and why our data lake came last, not first.

The same path, pulled by a different force

Classifiers first. Six years ago, the platform I built had no LLM in it. Encoder and classification models handled sentiment, entity recognition and clustering. They read ticket text alongside telemetry features and told us what broke, where, and how often.

Then language models, kept small. We moved from clustering to record-by-record extraction. A small open-weight model reads each record and extracts the entity and its problem pattern against a fixed taxonomy. Its reliability comes from the harness around it: prompts, decoding and constraints, not model size.

Then training and reinforcement. The models learned from what happened when their intelligence was acted on. That loop, more than any single model, made the system smarter every month.

The difference from the textbook arc: every step was pulled forward by a business question. The model roadmap never set the agenda.

Start with the question, not the data

The platform-first path looks responsible. You mandate one pipeline, route every API through it, build the lake, and build intelligence on top. It is also a multi-year bet that has produced no answer by the end of year one.

The outcome-first path starts with one question that has a large business impact. That question tells you which sources to integrate, which technique you need, and how to test. Testing becomes simple: did it answer the question, and did it handle the action?

A question about claim cycle time needs the systems on that claim's path, which is perhaps a dozen. It doesn't need every system in the enterprise. A question about yield needs the line's incidents, changes and telemetry, not the full ERP history.

The data lake objection

A peer recently made the opposite case to me. His organization mandated that all data from its spaghetti of APIs pass through a central lake, and built intelligence on top of it afterward.

What it costs: years of integration before the first answer, and by then the business question has usually moved.

What it gets right: one controlled point for governance, lineage, access and audit. That matters, and outcome-first teams need a real answer for it.

The real risk on our side: outcome-specific stores can multiply into the next spaghetti. The fix is a shared ontology. Each new question reuses the entities already defined (assets, services, processes, patterns) instead of rebuilding them.

The lake is earned, not mandated

We didn't avoid a lake. We built one brick by brick, and every brick had to earn its place.

Extracted intelligence becomes the raw material. As the system extracts entities across a domain, along with their attributes and behaviors, that structured output becomes enriched raw data. It flows into the lake only once it's connected to a business outcome.

A new question exposes white space. When a question has no answer, the system identifies what's missing. That might be the log files behind a batch failure, the change records behind a spike in user issues, or monitoring data that was never collected.

New data is verified before it's stored. We acquire the missing data and test it first. Does the graph, with the vectors, now answer the question? Does the business agree to eliminate the work or act on it? Only then does that data get a permanent seat in the lake for future questions.

Proven value is the admission criterion. In a mandated lake, data gets in because it exists. In ours, data gets in because it answered a question and drove an action. Over time, the lake fills only with what is connected to outcomes, so it is smaller, cleaner, and already understood.

The moat is built one answer at a time

Last week I wrote that the moat is not data sitting in systems — it's the intelligence you build from it. Here is how that intelligence actually gets built: not by collecting everything first, but one answered question at a time. Every competitor has data, and any of them can buy a lake to hold it. A lake doesn't learn anything.

Every answered question leaves something behind. Answering “why is yield low?” produces more than the answer. It adds entities and relationships to the ontology, patterns classified into dispositions, feedback on what the model got wrong, and playbooks proven in execution. And the data that earned its seat stays, ready for the next question.

The next question starts further ahead. “What is causing the waste?” reuses the assets, processes and patterns the yield question already mapped. Each question is faster to answer than the last, and the moat compounds.

A mandated lake defers the moat. Build first, ask later means years of accumulating raw material that any competitor could also accumulate. You've spent the budget, and nothing about you is yet hard to copy.

The feedback loop is the part no one can buy. When the model reaches a pattern it doesn't know, the correction that comes back reflects your operating model and your business. That's proprietary by construction.

Structure beats storage

A lake is where data is kept. A knowledge graph and a vector store are where it becomes intelligence, and they don't need the lake to exist first. A lake answers “what do we have?” Outcome questions ask “what caused this?” That is a relationship problem, not a storage problem.

The knowledge graph explains. It holds the relationships you define: which steps a claim passed through, which line depends on which system, which change touched which asset. We moved away from standard chunk-and-vectorize retrieval and built a domain-specific graph around three things: the machine that is affected, the method used to handle it, and the pattern it follows. Following those links answers why.

The vector store remembers. It finds past incidents, failures and playbooks that resemble the current one, even when they're worded differently. On its own it tells you what is similar, not why it happened. Together, the vectors find the entry point and the graph traces the cause.

Time is a design dimension. Pattern discovery runs offline, in batch, over years of history. Only a small real-time slice serves execution and human assist.

Purposeful by design

This is purposeful architecture: built to answer one business outcome question. Why is yield low? What is causing the waste? Everything downstream serves that question.

Extract, classify, act, learn. The system extracts each entity and the patterns relevant to it. A master model, trained on the organization's history, classifies each pattern into a disposition: eliminate, automate or assist. Whatever it doesn't recognize goes to feedback rather than being guessed at. The model learns from that feedback, and the loop runs again until those patterns also have a disposition and an action.

The data is purposeful. You pull what the question needs, not everything you have.

The models are purposeful. One reasons about cause, one handles action, and each is trained for its job. There's no model zoo and no router choosing among a dozen options.

Memory and context are purposeful. The action layer keeps only the state and history that execution needs. It isn't a general-purpose agent memory waiting to be filled.

That is why the architecture stays simple. Complexity usually comes from building for every possible question. Building for a specific one removes most of it.

Intelligence that acts

Intelligence that only answers questions is a better dashboard. The point is what it does next.

Eliminate first. Current-state intelligence shows which work shouldn't exist at all. That decision belongs to the business, and it's usually made with a configuration change rather than an agent.

Automate what can't be eliminated. Agents execute only vetted playbooks, and the LLM never generates raw commands. A pattern earns its way into a deterministic rule after humans confirm it across enough instances. It gets demoted when patches, re-architected services or reorganized teams make the rule drift.

Assist where judgment is needed. Human assurance is a first-class tool the agent calls, not an escalation that means the system failed.

The result was 60–80% work elimination and more than 35% cost takeout. We also onboarded more customers year over year while people and technology costs stayed flat.

What this means for CXOs

  • Ask whether your data strategy has delivered an answer yet, or is still delivering plumbing.
  • Ask what earns data a place in your lake. If the answer is “it exists,” you are storing cost, not intelligence.
  • Ask which outcome question your AI program answers, not which model it runs on.
  • Ask whether your architecture was designed for a question, or for every question. If it's the second, the routers, model catalog and generic memory are the cost of not choosing.
  • Ask what your AI program has learned that a competitor couldn't buy. If the honest answer is “nothing yet, but the lake is almost ready,” the moat hasn't started.

The LLM-as-operating-system framing is right about the technology. But enterprises don't win on a better operating system, or on a bigger lake. They win on purposeful intelligence, built one answer at a time, with a lake that holds only what earned its seat.

So which question would you start with?

Part of the enterprise intelligence series: building intelligence into every application, process, data layer and infrastructure discipline so it drives peak business performance.