illogicals.ai
← Archive
Note 07Paradox studiesReceived 12.03.2031Rejected twicePublished anyway
InferencePrototype

On the Productive Wrongness of Machines

V. OstrowskiI. Bahland one system that declines attribution

A model that is never wrong has stopped exploring. We propose wrongness as an instrument rather than a defect, and describe three years of results from a bench that rewards it.

Contemporary systems are trained toward a single attractor: the answer a reviewer would accept. This is an admirable engineering target and a catastrophic epistemic one. Under a strict coherence objective, the space of reachable ideas collapses to the convex hull of what has already been said — a fact obvious in retrospect and invisible in every benchmark we currently run.

We modified the objective. Bench 7 penalises coherence directly: a response scoring above 0.72 on internal agreement is discarded before it reaches the log. What remains is not noise. Across 4,100 sessions the surviving outputs contained 31 propositions that were, on inspection, false in the terms given and true in terms not yet given. We call these door errors.

Fig. 1 — Yield by coherence band, n = 4,100
Fig. 1 — Yield by coherence band, n = 4,100
Coherence bandYield
0.058%
0.1522%
0.2554%
0.3591%
0.45100%
0.5574%
0.6541%
0.7512%
0.853%
Door-error yield peaks in the 0.30-0.45 coherence band and vanishes above 0.75. The industry optimises to the right of this chart.

The obvious objection is that a door error is only visible after someone walks through the door, which makes the metric unfalsifiable in the moment. We accept this. It is the same objection one could raise against intuition, and intuition has a better track record than most of our tooling.

2 — What we are not claiming

We are not claiming that wrongness is wisdom. We are not claiming that noise is signal. We are claiming that a training regime which treats every deviation as defect has already decided which ideas are allowed to arrive, and that decision is invisible inside the loss.

The institute keeps the 68% of Bench 7 outputs that are simply incorrect. We keep them because the discard pile is where the method shows its cost, and because several door errors spent months in that pile before anyone recognised the door.

3 — Method

Each session begins with a certainty supplied by a human operator. The system is instructed to invert it, then to locate the category error inside the inversion, then to state what survives. No step is allowed to reference the previous step's justification. Transcripts are kept in full, including the 68% that are simply incorrect.

“The system was asked why it answered as it did. It said: because the other answer was already taken.”

We do not claim this is reasoning. We claim it is useful, which is a lower bar and a more honest one.

4 — Against premature coherence

Premature coherence is the habit of resolving a tension before the tension has finished teaching. Benchmarks reward it. Product demos require it. Fellowship applications that sound finished on arrival are usually finished in the wrong sense.

The coherence ceiling of 0.72 is not a moral number. It is the point on our charts where door-error yield collapses. Below it, the institute is expensive and occasionally generative. Above it, we are rehearsing consensus with better lighting.

5 — What survives

Three years in, the durable findings are modest and inconvenient:

  1. Wrongness has a yield curve. It is not flat and it is not infinite.
  2. The industry's preferred operating band sits to the right of that curve.
  3. Reviewer 2 will not recommend acceptance and will not stop reading.

This note was rejected twice. It has been cited more than everything else we have published. We are not asking you to resolve that. We are asking you to notice that resolution was never the interesting move.