A Second Opinion Only Counts When the Second Doctor Reads Something the First One Did Not

ORIGINAL LINKEDIN ARTICLE

A Second Opinion Only Counts When the Second Doctor Reads Something the First One Did Not

5 min read
A Second Opinion Only Counts When the Second Doctor Reads Something the First One Did Not - article by Andre Magrini

My revenue forecast was wrong for years, and nobody on my team ever lied to me. The number was broken for a structural reason. The person who wanted the deal to be real was the same person certifying that it was real, working from the same call notes, carrying the same optimism, paid on the same outcome. I now recognize that exact architecture on almost every AI reliability slide I am shown.

The pitch has converged. The model drafts, the model critiques the draft, the model revises, and the vendor calls the result verified. It is the standard reliability story in 2026, and it usually appears one slide before pricing.

Meanwhile the room changed. Risk, procurement and legal now join the first call, and they are not asking whether the system is accurate. They are asking who checked. Most architectures shipping today cannot answer that question, because the checker and the author are the same entity wearing two labels.

Let me concede the strong version of the counterargument first, because it is real and because most critiques of this stuff skip it.

Sampling a model many times and taking the majority answer does lift measured accuracy. That result is replicated, it is not marketing, and anyone who waves it away is not arguing in good faith. It works because it attacks variance. When a model's errors are scattered, more samples wash them out.

That is also exactly where it stops. The failures that survive to production are not scattered. They are the ones that look correct from inside the model's own distribution: fluent, well-formed, structurally plausible, and wrong with the same confidence as the right answers. A checker built from the same weights, reading the same context, carrying the same priors, tends to fail on those same inputs. The correlation is not perfect. It is high enough to void the arithmetic that every reliability claim quietly assumes. Independent failures multiply and get rare fast. Correlated failures do not multiply at all, and in the limit the second check removes exactly zero error while adding cost, latency, and a sense of safety you did not buy.

I run a verification benchmark in my research work at Tepis AI. It clears 54 of 1,000 cases end to end on the strict criterion, with zero false clears observed across that set.

Now the uncomfortable part, before somebody puts it in the comments. We designed that benchmark, and you cannot see the cases, which is precisely the weakness this post is about. So let me scope it honestly. Zero observed false clears in 1,000 is not "never." The defensible statement is under half a percent, not zero, and the only thing that would convert this from a claim into evidence is an outside party running the same thousand cases and reporting whatever they find. That offer stands. A number you cannot falsify is a brochure.

What I will defend is the design reason behind it: the verifier does not share a substrate or an information source with the generator. That is the whole mechanism, and it generalizes far beyond anything I build.

Four gates. Run whatever you are about to trust with a real decision through them.

1. 𝗜𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗼𝗻 𝗶𝗻𝗱𝗲𝗽𝗲𝗻𝗱𝗲𝗻𝗰𝗲. Does the checker have access to something the generator never saw? A document, a record, a measurement, a system of record, a person. This is the load-bearing gate. If both are reasoning over the identical context, you are measuring internal consistency, and consistency is a property of good fiction as much as of good analysis. The single question worth asking a vendor is this one: when your system verifies an output, what specifically does the verifier have that the generator did not?

2. 𝗦𝘂𝗯𝘀𝘁𝗿𝗮𝘁𝗲 𝗶𝗻𝗱𝗲𝗽𝗲𝗻𝗱𝗲𝗻𝗰𝗲. Do the two share weights, architecture or training data? This gate is what makes gate one durable. Shared substrate means shared blind spots, so ask for the correlation between their failures rather than for two accuracy numbers. Two components that are wrong on the same inputs give you one check, priced as two.

3. 𝗜𝗻𝗰𝗲𝗻𝘁𝗶𝘃𝗲 𝗶𝗻𝗱𝗲𝗽𝗲𝗻𝗱𝗲𝗻𝗰𝗲. Is anything in the system scored on throughput, resolution rate, or share of cases cleared? A checker rewarded for clearing will clear. This is not a machine problem, it is the oldest problem in auditing, and we are rebuilding it without the safeguards that took the accounting profession a century to write down.

4. 𝗥𝗲𝗰𝗼𝗻𝘀𝘁𝗿𝘂𝗰𝘁𝗶𝗼𝗻 𝗶𝗻𝗱𝗲𝗽𝗲𝗻𝗱𝗲𝗻𝗰𝗲. Can a third party reconstruct why the check passed, from stored artifacts, without rerunning the model? If the justification only exists inside a run you cannot reproduce, you have a result, not an audit trail. This is the gate regulated buyers will decide on, and almost nobody is building for it yet.

Notice what passes. A unit test passes. A type checker, a schema validator, a database constraint, a reconciliation against a ledger: all four gates, cleanly, and they are cheap and boring. A same-model checker reading a retrieved contract the generator never saw passes gate one and fails gate two, which is a real improvement and worth having. A judge from a different vendor, or a human reviewing a random sampled slice, gets you most of the way. None of that is my product. If a framework only clears the thing its author sells, it is a spec, not a test.

Which brings me back to the forecast.

Taking an organization from $35M to $150M was, in retrospect, a long lesson in exactly this. Every one of the four gates was open. The evidence was produced by the same people who were graded on it, held in their notes, justified in a conversation nobody else could reconstruct. The fix was not better estimating and it was not better people. It was making the evidence independent. A deal counted when we could point at an artifact the customer had produced rather than one we had produced about them: a requirement written in their words, a named person in the buying committee on record, a mutual plan they had actually edited. The number got smaller, and for the first time it also got true. A forecast that survives someone trying to kill it is the only kind worth carrying into a board meeting.

I still run the same test on my own work. Every batch of output in the research venture goes to an adversarial review that did not produce it and gains nothing from its survival. It regularly kills work I liked. That is the evidence it is functioning. A gate that always passes is a formality, and formalities are what accountability decays into when nobody checks the checker.

So I want to hear where this went the other way. Every time I have watched a team choose between a cheap checker that shares a substrate and an expensive one that does not, the cheap one won, and the stated reason was always that the independent one flagged too much. Where have you seen real independence survive the cost conversation, and what did it take to defend it?

APPLY THE THINKING

Turn insight into an accountable operating decision.

Start with a focused diagnostic of the revenue, GTM, RevOps, forecast, or AI constraint.

Request an AI Revenue Diagnostic