The Answer Isn't a Bigger LLM. It's the Layer Underneath.
By Realila
When an AI gets something wrong, the standing instinct is to wait for a bigger model. A week ago we argued that the wait misses the point: the scarce resource is not intelligence but grounded intelligence, reasoning that rests on information someone has refined, kept current, and made traceable. This week the clearest evidence yet arrived, and it came from an earnings call. Not from the headline numbers, remarkable as they were: from something Palantir's chief technology officer said about an open model.
An open model beat the frontier
On the call, Shyam Sankar described standing up NVIDIA's Nemotron Ultra (an open-weight model, unmodified, with no post-training) inside Palantir's platform. Within twenty-four hours it was, in his words, "better than frontier": it outperformed the leading closed models on five production tasks. On public benchmark tables the same model sits below the frontier. By those numbers, he said, the result should not have been possible.
His explanation was that the benchmarks measure the wrong thing. They are right about what they measure (general exam questions) but those are not the tasks his customers are trying to solve. Graded on the customer's actual work, with the customer's actual data and rules within reach, the open model won.
The model had not changed. The ground underneath it had. That is the whole story of this company, and this week's results put a number on it.
The layer underneath
Palantir does not train frontier models. What it builds is the layer a model stands on: an ontology (a live, machine-readable model of how an organisation actually works). Its data. The business logic that governs that data. The actions the organisation can take. The rules about who may see and do what. All of it held in one current, governed layer that any AI model can reason over.
Readers of our last piece will recognise this. It is the refining we described (collect completely, reconcile against the source of record, know what each number can and cannot support, make every figure traceable) done at industrial scale and handed to the model as a foundation. The conversational layer sits on top. The substrate is the product.
Two consequences follow, and the Nemotron result is both of them at work.
The model becomes an interchangeable part. When the substrate does the grounding, model choice turns into an engineering decision (cost, speed, where the data must live) instead of a capability ceiling. Customers plug in open or closed models, switch between them, and keep what the work produces: the fine-tuned weights, the metadata, the reasoning traces.
Evaluation moves from leaderboards to your reality. The platform lets a customer build the benchmark that actually matters (their tasks, their data, their definition of correct) and grade every model against it. That is the only way a result like Nemotron's can even be discovered. It is also, we would argue, the only evaluation that means anything once real money is on the line.
How it reaches the customer
The delivery model is as distinctive as the product. Palantir embeds forward-deployed engineers inside the customer's operation, to learn the domain before software gets written. Its signature onboarding is a five-day workshop in which a prospective customer builds working applications on its own data, supervised by those engineers. Sales cycles that once ran a year or more now close in days, because the product is demonstrated on the buyer's own reality, not a demo dataset.
The business model, and this week's proof
Palantir's stated position is that it wants to be paid as a share of the value its software creates (not per token, click, or unit of consumption) on the argument that usage alone says little about results. The quarter suggests customers accept the bargain. Revenue of US$1.94 billion, up 93% on a year earlier: the fastest growth in the company's history. GAAP net income of US$1.06 billion, a 55% margin. The US commercial segment grew 149%. Full-year guidance was raised to about US$8.15 billion, an 82% growth rate.
The number that speaks most directly to whether the layer works in production is net dollar retention: 157%. The customers Palantir already had a year ago now spend 57% more, before counting anyone added since. That is expansion earned after deployment, inside accounts, where the software either produced results or did not. And it came, the shareholder letter notes, with a small and shrinking sales headcount. When a deployed product performs, the performance does the selling.
The same layer, at person scale
Palantir builds this for ministries and multinationals. We build it for a person facing a property decision in Singapore, and the mechanism is the same, because the failure it prevents is the same.
What people call hallucination is, in our domain, a fluent number with nothing underneath it: a price the model half-remembers, a duty rate from an old forum post, arithmetic improvised mid-sentence. Our platform is built so those moments cannot reach the person. The model never recalls prices from its training. It reads our refined record: more than a million sale records and 1.6 million rental records, reconciled against the official statistics, every figure traceable to the transaction that produced it. Calculations (duties, loan limits, affordability) run in verified code the model calls; it presents the result, it does not improvise the arithmetic. Regulated rates come from a maintained policy reference with effective dates, never from the model's memory. And our evaluation is theirs, scaled down: the first user is a practising, licensed agent running the platform inside live transactions, where a wrong number surfaces at a negotiating table rather than in a demo. Palantir calls that pattern the forward-deployed engineer. At our scale it has a simpler name: the founder.
The result is the one Sankar described. The model gets better (noticeably, measurably better) without getting bigger, because it is standing on better ground.
Models will keep improving, for everyone at once, and that is good news that decides nothing: the same tide lifts every boat. What compounds is the layer underneath: refined, governed, current, and graded against the tasks that actually matter. A week ago we ended by asking what an answer is resting on, and whether you can see it. This week, one of the most valuable software companies in the world reported what happens when that question is the entire product.
Next in this series: AI vs Human Is the Wrong Question
Realila Labs is building intelligence for decisions that matter. Be the first to know when Realila launches.