Five Thousand Sources, One API, and the Problem That Starts After Integration
Every platform can unify government data sources now. The open problem is getting a general model to reason correctly over the result, and defense finance breaks them.

The FY2027 justification books carry a line labeled "FY 2027 mandatory request" across 35 program elements, holding $124.9 billion in reconciliation money. The label uses the word "request." The money has not been enacted. Point a general-purpose language model at that line and ask whether the funding is in law, and the model will report that it is. The word "mandatory" triggers a reasoning pattern trained on federal budget conventions, where mandatory spending refers to entitlement programs that do not require annual appropriation. In defense, the term means something different. It means money the department is asking for through reconciliation, dependent on a legislative vehicle that has not moved. The model reads the label correctly. But it draws the wrong conclusion from it, and the conclusion sounds authoritative because the underlying number is real.
That failure is specific to this year's budget. It did not exist in FY2026, when last year's justification books did not print reconciliation money alongside the discretionary request. It will also shift again next year as the labels and legislative status change. A general model trained on budget conventions will not catch the distinction unless someone builds the correction into the reasoning layer. That correction requires understanding what reconciliation money means in the defense context, how the defense usage diverges from the domestic budget, and what legislative status the FY2027 mandatory lines actually carry. The model can see every number on the page. What it lacks is the domain knowledge to interpret what the numbers mean, and that gap does not close with better data access.
What integration does not fix
This is the problem that appears after integration, and it is the problem that the current market conversation largely skips. Every serious data company selling into the defense analytics space makes the same pitch: thousands of structured and unstructured government sources, reconciled and exposed through a single API. The pitch is true. It has also become ordinary enough that it no longer separates one platform from another. The companies that built these integrations did hard, legitimate engineering work. Reconciling entity names across registration systems is a difficult task. So is normalizing budget line identifiers across fiscal years and matching contract actions to the awards they modify. All of it is necessary. The question is what happens after the sources are joined and a model starts reasoning over the unified result. That is where the failures appear, in patterns specific to the domain.
The mandatory-request confusion is one of those patterns. Contract data produces another that is equally consequential and equally invisible to a general model. A federal contract record carries multiple financial fields, and two of them look similar enough to cause consistent misreads. The contract ceiling records the maximum obligated value over the life of the contract. The cumulative total-value field records the running position as of each modification. A model reading a modification notice without understanding the distinction will misread what happened on every large contract action it encounters. The PAC-3 MSE contract from earlier in this series shows the problem in full. An initial undefinitized action posted at $4.7 billion in April. Then a seven-year modification posted at $53.86 billion in July, taking the total ceiling to $58.62 billion. The cumulative field on the modification record captures the running contract value. A model that reads it as the modification amount will report a contract that grew by $53.86 billion in a single action, when what actually happened is that a ceiling was set for the full seven-year term. Every number in the record is correct. The error is entirely in the interpretation. It requires knowing which field carries the increment and which carries the running total. That knowledge lives in the domain, not in the data.
Where entity resolution breaks
Entity resolution is a third area where general models fail consistently, and for anyone trying to size a company's federal book of business it is the most consequential. A large defense contractor holds its federal work across dozens or hundreds of subsidiary registrations, joint ventures, and legacy entity records that were never consolidated after an acquisition. Pulling a company's federal revenue from a single registration will undercount its actual book by a wide margin. The consolidation index described elsewhere in this series found 722 awards that changed hands across 232 entities in a single trailing twelve-month period. The novation record that tracks those ownership changes is the only reliable signal that two registrations belong to the same parent company. A general model querying a company name against the contract database will return whatever the name match produces. Resolving the full entity graph requires reading modification reason codes at the action level across tens of millions of records and linking them to a parent structure that the database itself does not maintain. The model has access to every contract record in the system. It still cannot tell you the size of a company's federal book without the entity resolution layer that sits between the raw data and the answer.
These failures are not unusual. They're the first things that go wrong whenever a capable model meets defense finance data. The reason is consistent. A term like "mandatory" means one thing in the domestic budget and something else entirely in defense. A contract field logs a running total under a name that looks like an increment. A registry tracks entities by name rather than by the ownership chains that actually connect them. In every case, the data is accessible and machine-readable. The reasoning needed to interpret it is domain knowledge nobody built in.
Why a bigger model does not help
The instinct in the market right now is to point a larger general model at more sources and treat the reasoning as solved by scale. These failure types cut against that instinct in three ways. The errors are specific to the domain, so a better general model does not fix them. They recur predictably, enough to be catalogued and tested against each new model version. And they are consequential at the scale of real decisions. A model that reports Golden Dome as funded in law when $17.1 billion of its $17.5 billion depends on an unenacted reconciliation request is producing an answer that could move a position. The error is not a hallucination. The model read the data correctly and reasoned about it in a way that the domain does not support.
The practical case for a narrower model built on this category of data rests on these failure types rather than on an abstract preference for specialization. A model trained on the domain will flag "mandatory request" in the FY2027 defense justification book as unenacted reconciliation money rather than reporting it as funded law. The same model, trained on contract data conventions, will read the total-value field on a modification as cumulative, not incremental. Trained on the entity graph, it will resolve a contractor's federal book against the full novation history instead of a single-registration snapshot. Each correction is a piece of domain knowledge applied at a specific point in the reasoning chain. Each is testable by handing the model a record and checking whether the output matches the domain-correct reading. The failures are checkable because they are narrow, well-defined, and specific, and that is what makes them fixable in a model built to carry the domain rather than a general one scaled to cover everything.
What stays with the analyst
What remains genuinely unsolved sits further from the current conversation and closer to the boundary where any model, narrow or general, needs a human alongside it. A domain-tuned model will still encounter records where the correct interpretation depends on context that lives outside any structured field. A Senate authorizer's zero on a procurement line could mean the committee eliminated the program or expects it funded through reconciliation. The only way to resolve that ambiguity is to read the committee report language, which is unstructured legislative prose written by staff who are not optimizing for machine readability. A contract description referencing "missile defense engineering support" could be Golden Dome work or legacy MDA activity that predates the program by a decade. Distinguishing the two requires reading the description against the funding history and the vendor's prior work, which is a judgment call rather than a pattern match.
The data unification problem is solved across the industry. Narrowing a model to the domain closes the reasoning failures that this series has documented across budget labels, contract fields, and entity records. The judgment problem is the part that remains open. It’s the ambiguity that requires a human reader to hold two possible interpretations in mind and weigh them against context that no model fully carries. Getting the first two layers right is what determines whether the analyst spends their time on that judgment work or on catching errors that a properly built system should never surface in the first place.
Send us a coverage list or a set of target programs. We reply within one business day.
