The Signal & The Noise · Vol. 06
I asked my own system the same question five times, over the same folder of documents, with the same model. Four different answers came back, and a fifth run that produced nothing at all.
Somewhere in your organisation this week, somebody will ask an AI system a question about your own business. What did we agree with that supplier. What have we committed to. Where is this project stuck.
They will get an answer, it will look finished, and nobody, including the person who asked, will be able to tell you whether it was complete.
That question sits underneath everything my company does. Samture builds the knowledge layer that makes an organisation's own AI trustworthy, and the honest way to describe the problem is as a question rather than a product: can your organisation prove what it knows? I was not willing to put that to anybody else before putting it to myself, so I asked my own system a question I have to answer properly about once a month. What is outstanding with each partner, since when, and who is the contact. I asked it five times, over the same three hundred and thirty eight documents, with the same model.
Four of the five gave me an answer, and all four were different. One named two partners the others left out. The fifth gave me nothing at all, having walked into a folder it could not open and stopped there, after twenty minutes in which nothing said anything was missing.
Every one of those answers looked thorough. Had I asked once, which is what all of us do, I would have taken whichever one I happened to get and gone into my week with it.
The problem is not the model, and a better one will not fix it.
When a system searches your documents it reads them again from scratch for every question, works out again what they mean, and does that under time pressure using whichever files it happened to open. Reading is a judgement. Ask twice, get two judgements.
What changes it is unglamorous and it is not a purchase. You read the material once, deliberately, and write down what you found. This agreement runs until this date. This commitment was made by this person. This project is blocked by this question. Every fact carries the document it came from and whether a human being has confirmed it.
And before any of that, you agree what your own words mean. What counts as an agreement rather than a conversation. What makes something a commitment rather than an intention. Written down once, so that the same document filed twice lands in the same place.
That written agreement is the whole thing, and it is not a technical document. It is a description of how your organisation actually works, and nobody outside it can write it for you. For Samture the machinery cost about four dollars and half an hour, and runs on a laptop. The expensive part was the afternoon I spent deciding what my own words mean.
There are names for all of this once it exists. The layer that fixes what may exist and how the pieces relate is an ontology, the layer that records what each term means and excludes is its semantics, and the result is a knowledge graph built on linked data standards that have been sitting in this industry for twenty years. I am putting those words at the end of this section rather than the beginning on purpose. They are the consequence of taking the question seriously, and they are not the answer to it.
The basis, so you can weigh it yourself: one question, five runs on each route, the same model on both sides, on 4 September 2026, over three hundred and thirty eight documents.
The answer stopped moving. Nought out of five runs agreed on the searching route, and five out of five agreed on the other. That is the finding. Everything else below is a consequence of it, because what changed is the nature of the operation: finding an answer stopped being a search, which is a fresh judgement every time, and became a retrieval, which returns the same rows or tells you there are none.
Looking things up went from between eighty and two hundred and fifteen seconds to under a second. Writing the answer out still takes about a minute either way, so the end to end time barely moved. And a caveat with a date on it, because this is the point of the whole article: that store has since doubled and the same question now takes four and a half seconds. It is a reading taken on a Thursday, not a property of the method.
95% less material went through the machine. Answering that one question by searching pushed roughly 1.1 million words past the model, because each turn carries everything gathered so far back past it again. The written question used 55,500. That is the figure I would put in front of anybody planning to run this at scale, because it is paid on every question, by every person, every day.
57% came off the bill, from $1.17 to $0.50 a question. Less than the 95%, because going back over material a model has already seen is cheap. Both numbers are true and they measure different things, so I am giving you both rather than the flattering one. Loading the whole archive in the first place cost around four dollars, and that is paid once rather than on every question, forever.
One thing this does not do is remove the model's capacity to invent. What it removes is the opportunity. On the second route the model never sees the documents. It is handed a structured set of rows and writes them out, so it cannot introduce a party, a date or an amount that is not in that set. That is controlled retrieval rather than hallucination-free AI, and the distinction matters, because the first is something you can demonstrate to an auditor and the second is a marketing claim.
Almost the entire conversation about this is about which model, which platform, which benchmark, and how often it makes things up.
Most AI evaluation answers whether a model was right. Very little of it answers whether your organisation can reproduce, explain and defend the answer it was given, and in a regulated business that second question is the one that eventually gets asked, by a board, an auditor or a regulator who wants to know where a number came from.
Since I did this, the same question returns the same rows every time and I can trace any line in any answer back to its document in a couple of seconds. That is the saving that matters most, and it has little to do with money. It is the difference between believing something about your own business and being able to show it.
What follows is not a list of things that happened to go wrong with me. It is the first week of output from a system built to catch exactly this kind of thing, pointed at a business that had never been examined this way before. Every one of these was found because something was counting, and every one is now closed or has a check on it.
I have a rule I take seriously. What a client pays us must never arrive later than what we owe a supplier, because otherwise we are financing the gap ourselves. I turned that rule into a check the system runs against itself.
It reported no violations at all, across a hundred and eighty eight agreements.
Then I looked underneath. Forty seven of those agreements record what we owe the supplier. Two record what the client owes us. None record both. The check had never once had anything to compare, and a report of nothing found reads exactly like a report of nothing wrong.
There is now a second check that reports that gap, and my clean sheet became forty seven open items without one thing in the business having changed.
I would put that in front of anybody running a contract portfolio, because the same is almost certainly true of yours. And because it is the general lesson. A control can function perfectly and still be controlling nothing, so the question to ask of every one you own is not whether it works, but what it would fail to see.
Two more, and both transfer.
A third of the tools I thought were running were not there. Asking the system what is connected returned a list, and roughly a third of it had been unplugged months earlier and never taken off. Nobody had lied about it. The list had simply never been asked to prove itself against reality.
And this one: ask what something can do, not what it is allowed to do today. I have a rule that anything able to send, delete or pay has to ask me before it acts. When I put that question to the system, the answer it gave me was wrong, because the code had worked it out from the wrong list. Two questions that look identical and give opposite answers: what is this capable of, and what is it cleared for right now. For a security review only the first one counts, and I had been answering the second without noticing. The controls themselves were set correctly, so nothing had gone wrong. What was broken was my ability to demonstrate that nothing had.
None of this is a Gulf problem, and it is sharper here than anywhere I have worked.
Two weeks ago I wrote about a report in which sixteen per cent of organisations in the Emirates had testing, auditing and risk management in place for the AI they are already running. Dubai has committed to delivering half of all government services through AI agents inside two years. Adoption is moving faster than the layer underneath it, which is a good problem to have and a short window in which to solve it.
And it does scale, though not in the order people expect. It scales conceptually before it scales technically. First you decide what your organisation means by its own words, then you make those meanings machine readable, then you connect them to the evidence. An organisation with a hundred thousand documents, twenty systems and fifteen definitions of the same term does not need a bigger machine than I used. It needs that first decision made properly, by people with the authority to make it, before anything is built.
Which leaves the question of who does that work. When Samture does this for a client, we contract for the result and hold the standard the work is measured against. The people who build to that standard are our Excellence Centre, and it is not a plan or a building. It is around two hundred and fifty certified practitioners who can do this work in the Gulf today, delivering through our partner network. There are nowhere near enough of them for what this region is about to ask, which is why we are training more, here, rather than importing them in five years from somewhere that got there first.
Can your organisation prove what it knows? Not describe it, not believe it, prove it: the same answer twice, with its source attached, and an honest list of what nobody has checked yet. That is the whole of what Samture builds, and this month I built it for myself first, because I would rather find these things on my own material than on somebody else's.
So do the experiment rather than take my word for it, because it costs you an afternoon and it will tell you where you stand. Pick the question your organisation asks its own systems most often, the one somebody senior relies on, ask it five times this week in five separate sessions, and lay the five answers side by side.
If they come back the same, I would genuinely like to know how you did it. If they come back different, you have just learned something about your organisation that no dashboard was ever going to tell you, and that is the point at which this stops being an article and becomes a conversation worth having. Write to me at team@samture.com or send me a message on LinkedIn.
Where does your own foundation actually stand? An honest look at the eight domains where a foundation either holds or quietly gives way, and at the distance between what you have bought and what has landed.
Take the Capability IndexThis is one system observed closely rather than a study of many, and the figures are from Samture's own business. The count of certified practitioners is an informed estimate being counted this month against a written definition and a fixed date; the result will be published whatever it says.
The Signal & The Noise
A fortnightly read on what is actually changing in AI across the UAE and the Gulf, written personally. No content calendar, no generated filler.