Thinking

We Have LLMs Now. Why Is It Still Hard to Get a Real Answer?

By Realila

Ask a language model almost anything and you will get an answer. It will be fluent, well-structured, and delivered without hesitation. For a large class of questions, it will also be right. This has quietly reset what people expect: if the machine can explain quantum tunnelling and debug my code and draft my contract, surely it can tell me whether to buy this flat.

It cannot. And the reason is worth understanding, because it is not a temporary limitation waiting on a bigger model.

The test that breaks the illusion

Pick a question that actually costs you something to get wrong. Not "explain compound interest" but "is the asking price on this specific unit reasonable." Not "how do bonds work" but "given what I hold, what should I do this month."

Ask a model. You will get an answer. Read it carefully and notice what it is made of: general principles, plausible-sounding heuristics, a hedge, and a recommendation to consult a professional. What it is not made of is the thing that would actually settle the question, which is what comparable units in that building sold for in the last six months, on what lease, at what floor, and how that compares to the quarter before.

The model does not have that. It was not trained on it, it cannot look it up in any form it can trust, and in most markets that record is not sitting anywhere in a usable state. So it does what it is built to do: it produces the most plausible text.

Fluency is not grounding

Here is the part that makes this dangerous rather than merely disappointing.

A language model's confidence is a property of its language modelling. It is not a signal about evidence. The same measured, authoritative tone comes out whether the answer is anchored in six thousand verified transactions or in the vague residue of a property blog from three years ago. The prose does not get shakier when the ground disappears.

Humans read confidence as a proxy for reliability, because among humans it usually is. Someone who answers a specific factual question briskly and precisely has usually earned that with knowledge. That heuristic is deeply trained into us and it fails completely here. The most common way people get hurt by these tools is not wild hallucination, which is easy to spot. It is a reasonable-sounding answer, delivered with exactly the same fluency as a well-grounded one, that happens to rest on nothing.

"Just give it search" is not the fix

The obvious response is to hand the model sources. Let it search, let it retrieve, let it cite.

This helps, and for many questions it is enough. For consequential decisions in a real market, it inherits a problem that has nothing to do with AI: the sources themselves are a mess.

Information about any market where money changes hands is scattered across dozens of places, thick with noise, and shaped by whoever is publishing it. The same number gets presented two ways depending on the interest behind it. A headline index can fall while the median price rises, and both statements are true, and which one you lead with tells the reader a different story about the same quarter. Records get revised. Some get withdrawn after publication and quietly disappear from the source, leaving anyone who copied them earlier holding figures that no longer exist.

Point a retrieval system at that and you get a model that confidently repeats the mess, spin included, with citations. The citations make it feel more grounded. They do not make it more true.

Refining has to happen before the reasoning

The thing that is actually missing is unglamorous. Someone has to do the work of turning raw information into something a decision can rest on:

Collect it completely, which means understanding that the record fills in over weeks and that the most recent period is always provisional. A number that will move is not the same as a number that is wrong, but reporting it as final is.

Reconcile it against the source of record. When your own computed figure contradicts the official one, the useful assumption is that you are wrong. We had exactly this happen recently: our regional price cut said one thing, the official index said the opposite. The temptation was to trust our own, because it was more granular and it was ours. The contradiction turned out to be a tagging fault in our pipeline that had silently left thousands of records unclassified. Had we published, we would have been confidently, fluently wrong.

Know what each number can and cannot support. A raw median moves with the mix of what sold, not only with prices. Reading it as a price change is one of the most common errors in market commentary, and it is the kind of mistake a language model will happily reproduce because it appears everywhere in the text it learned from.

Make every figure traceable. If an answer cannot be followed back to the record that produced it, nobody can check it, and an answer nobody can check is a guess wearing better clothes.

None of that is model work. It is data work, done upstream, before a language model ever sees the question.

Where the models are genuinely extraordinary

This is not an argument that LLMs are overrated. Give a model a clean, current, well-structured substrate and ask it to reason over that, and it is remarkable: it will explain what the numbers mean for one particular situation, hold several constraints at once, and answer follow-ups in the person's own terms. That is real, and it is new, and it is the part that used to require an expensive hour with a specialist.

The reasoning layer arrived first and it arrived cheap. The layer underneath it did not.

So the scarce resource in AI-assisted decisions is not intelligence. It is grounded intelligence: a reasoning engine sitting on top of information that has actually been refined, kept current, and made traceable. Every credible AI product in a domain that matters is, underneath, mostly a data problem wearing a conversational interface.

What we are building toward

Information is a raw resource. On its own it is scattered, noisy, and shaped by competing interests, which makes it hard to turn into a good decision, especially without the experience to tell signal from spin.

Realila Labs exists to do the processing. Our flagship, Realila, does it for property: the full transaction record, refined, reconciled against the official statistics, and traced back to source, so that when someone asks what a home is worth or whether now is the time, the answer rests on something. The conversational part is the easy half. The half that took the work is making sure there is something true underneath it.

If a model gives you a confident answer to a question that matters, the useful next question is not whether it sounds right. It is: what is this resting on, and can I see it?


Next in this series: The Answer Isn't a Bigger LLM. It's the Layer Underneath.


Realila Labs is building intelligence for decisions that matter. Be the first to know when Realila launches.