
A Model Has No Null
A model cannot return an empty set. Ask it something it has no basis for and it fills the hole anyway, because filling holes is the whole job. Two vacuums come out of that, and only one of them gets talked about. Here is how I engineer the abstention the architecture will not give me.
Ask a model what a roof costs and it will tell you fifteen thousand dollars. I buy that roof for seven.
I've watched people treat that as a hallucination problem. It isn't. The model has never seen my invoices, my crew, or my supplier. It has no basis for an answer. What it does not have, and this is the part worth building around, is any way to say so.
There is no null in the output. Every response is a completion. The architecture cannot return an empty set, cannot hold a blank, cannot answer "insufficient data" unless somebody builds it a door marked insufficient data and points at it. Absent that door, it does the only thing it can do: it produces the most probable continuation. Confidence is not a setting it chose. Confidence is the only mode it has.
Once you see it that way, a whole category of failure collapses into one shape.
Two vacuums, and only one of them gets talked about
The data vacuum is the famous one. No cost history, so it fills with retail. That's the roof. Everybody in this space has written about it, myself included.
The resistance vacuum is the one nobody architects for, and it's the more dangerous of the two.
I watched it happen in real time last week. An operator in a session I run had loaded photos of a house and said he was leaning toward a wholetale. The model agreed. He asked what to do with the kitchen. Don't spend money there, it said. He asked what it thought about putting new doors on that same kitchen. It liked that too, and explained why.
His summary: it never tells me no.
That's not a different bug. That's the identical bug with a different input missing. He gave it a house with no cost basis and it invented retail. He gave it a decision with no opposition and it invented agreement. Both times it filled a hole, because filling holes is the whole job.
The reason the second one is worse is that a wrong price announces itself eventually. A sub quotes you, the invoice lands, reality arrives with a number attached. A wrong agreement never announces itself at all. You just proceed, feeling validated, and the cost of it shows up months later as a decision you can no longer trace back to the moment you should have been argued out of.
So build the null yourself
If the system cannot abstain, abstention becomes your job to engineer. Three constraints, in the order I'd add them.
Constrain the source. The model gets my cost export and an instruction with no wiggle in it: use these rates, do not deviate, and if a number is not in here, flag the hole rather than fill it. That last clause is doing all the work. It manufactures the null the architecture won't give you. A gap I can see costs me a phone call. A plausible number I can't see costs me the job.
Constrain the shape. Give it room and it will line-item a scope into oblivion, and forty items on a job you've never priced yourself all look equally credible. That isn't rigor, it's camouflage. Fake precision needs space to hide in, so cap the output at the handful of categories you cut checks for. Constraining the shape to match your real decisions is what makes an estimate something you can argue with instead of something you can only accept or reject.
Constrain the arbiter. The thing that produces should not be the thing that approves. In forVEX this is literal: rehab estimates run on a deterministic engine, and the model handles capture and narration but never touches the arithmetic. Same shape on the content side, where a writer produces and a separate grader scores against a fixed rubric and sends it back. The grader has no stake in the draft having been good. That indifference is the entire value.
You can approximate all three in a chat window without building anything. Load your own numbers first. Cap the detail. And decline the first answer once, every time, as a standing habit rather than a thing you remember when you're already suspicious.
The strongest version of that last one I've seen wasn't clever. An operator asked for sourcing help, got a list back where every option carried a mediocre rating, and simply refused it and asked for better. He bought off the second list and came in well under the quote he'd been working from. He didn't write a prompt. He declined an answer. The skill is entirely in noticing that the first response is a draft, not a verdict.
Who checks the checker
Here's the part I don't have a clean answer to, and I'd rather say so than run a victory lap.
The grader is also a model. It also has no null. Ask it to score something and it will produce a confident number whether or not it has grounds for one, for exactly the same reason the pricer did. I've built a checker that inherits the disease it was hired to treat.
What keeps it useful is that its job is narrow and its rubric is fixed, so there's far less vacuum available to fill. What keeps it honest is nothing internal at all. It's me, reading outcomes, noticing when a piece the grader liked landed flat. The loop only closes on contact with the real world, and the real world is not in the system.
Which is the thing I'd want any operator building this stuff to sit with. You can constrain the source, the shape, and the arbiter, and you will still be the last check in the chain. Not because the tooling is immature and will improve. Because there is no point at which a system that cannot say "I don't know" gets to be the final word on anything that costs you money.
The model doesn't sign the check, carry the holding cost, or eat the overage. You do, every time, and that asymmetry is the entire reason you stay in the loop. Any tool that appears to be offering to take it off your hands is the one to watch.
Build the null. Then keep standing behind it.
The working version of this, including the three pushback prompts I use and the exact instruction that forces a flag instead of a fill, ran in the No-Hype AI newsletter: It told me fifteen. I buy it for seven.