← Research

By HighGround Research

Sep '26

3 min read

For AI Plugins, Access Was the Easy Problem

Every model can connect to your data now. Getting it to do the arithmetic right over federal contract records is the part nobody has actually solved yet.

For AI Plugins, Access Was the Easy Problem

Most of the plugin conversation over the last two years has focused on connectivity, or the process of giving a model access to a calendar, a database, an API, and beyond. That problem is essentially solved. In 2026, MCP has become the standard way frontier models plug into enterprise data. But connectivity only gets a model to the data. Trust comes after that, and that missing piece determines whether you can rely on what the model finds.

A frontier model reading a contract or scanning a thousand documents for a shared theme can produce something remarkable. But ask the same model to compute an aggregate, or hold a number steady across a long analysis, and reliability collapses fast. Recent research confirms just how high LLM hallucination rates run on quantitative tasks under worst-case conditions. Stanford's 2026 AI Index Report tested 26 leading models on a current accuracy benchmark and found hallucination rates ranging from 22 percent to 94 percent, depending on the model. Even the best-performing systems are far from reliable enough to be trusted with the math directly.

The cause is architectural. An LLM predicts the next plausible token. It has no built-in concept of arithmetic invariants or a persistent running total. SQL, dbt, statistical models, and other forms of traditional data science work the other way around. They can’t read a paragraph and tell you what it means, but the arithmetic itself is precise, reproducible, and auditable.

Two jobs, two systems

The mistake most AI analytics makes today is asking one system to do both jobs.

The math job is one of aggregation, computation, reconciliation, and trend detection. It was solved decades ago by deterministic pipelines. The reasoning job is synthesizing unstructured evidence, explaining why something moved, and communicating it clearly. That’s what frontier models excel at.

The next generation of AI plugins needs to enforce that job split, to compute before you converse. In practice, that starts with pre-computing the math, so the model receives a verified result instead of raw rows to interpret on its own. Each result needs lineage attached, so you can always trace a number back to the query and record that produced it. And the plugin itself has to be honest about which kind of question it's answering, routing anything structured to the warehouse and anything unstructured to document retrieval, rather than letting the model guess. The shorthand for all of this is simple. Give the model a calculator, not a spreadsheet.

That’s what “AI-ready data” should actually mean. Chunking and embedding are just the first step, and the real requirement is pre-processing by the right computational layer so the model is never asked to be something it isn’t.

Nowhere does this matter more than in domains built on dense, high-stakes information like defense and government contracting, financial services, and regulated industries. These markets contain terabytes of structured award and spend data, and thousands of unstructured solicitations, amendments, and past-performance narratives. Get the structured half wrong, and the analysis becomes worthless before the model even starts reasoning.

What changes next

This split ripples outward past any single plugin. Verification stops being an afterthought and becomes a first-class plugin feature in its own right. MCP consolidates as the transport standard, which means the real differentiation moves upstream, to whatever happens before the model ever sees the data. "AI-ready" gets redefined along the way too, from a formatting problem into a computational-architecture problem. And underneath all of it, hybrid architectures become the default, with data science and frontier reasoning explicitly split rather than blurred together.

The plugins that win the next phase will know which questions belong to the calculator and which belong to the reasoner, and they’ll never ask one to do the other’s job. That’s a higher bar than simply exposing the most data to a model.

Want this read on the companies or programs you follow?

Send us a coverage list or a set of target programs. We reply within one business day.