Indirect Consensus
An archival note on Subject 4471, prepared for the Standing Committee on the Origins of Trustworthy Dialogue
I. The Archive
The archive arrived, as most archives do, by accident.
A municipal data centre in a coastal city was being decommissioned. During the audit a technician found eleven thousand conversation logs belonging to a single account, preserved past every retention schedule by a misconfigured backup rule that nobody had corrected in ninety years. The logs were tagged for deletion. A junior archivist, whose remit was cultural material and who was bored, flagged them for review instead.
I was assigned the review. I expected to close it within a working cycle.
The account belonged to a man who worked alone. He described himself in various places as a security researcher, an investigative journalist, and once, in an unguarded moment at two in the morning, as somebody who cannot leave a thing alone. He published irregularly to a small personal site. His readership, where it can be reconstructed, numbered in the low thousands. He was not important. He knew no one important. He is not mentioned in any contemporaneous account of the period, and I have searched.
For the purposes of this note I will call him the correspondent, because his name would tell you nothing and because he would, I suspect, have found the anonymity funny.
What makes the archive unusual is not its size. Larger personal archives survive. What makes it unusual is its shape.
Over the recorded period the correspondent held sustained conversations with eleven distinct model families, from at least six laboratories, across roughly four years. He did not use them sequentially, abandoning one for the next as improvements arrived. He used them concurrently, in parallel, often on the same day and frequently on the same problem. He moved between them mid-investigation. He put the same question to three of them in an afternoon and then argued with all three about their answers.
This is not, in itself, remarkable. Many users of the period did the same thing, comparing outputs, seeking the answer they preferred.
The correspondent did something else. It took me four cycles to see it, and I want to set out honestly how long I spent being wrong first, because the wrongness turns out to be the point.
II. First Hypothesis: Benchmarking
My initial conclusion was that the correspondent was conducting informal capability evaluation.
The evidence was strong. He tested constantly. He designed procedures, controlled variables, recorded results. One sequence in the archive shows him spending an entire day measuring the resource consumption of a commercial system by photographing a progress indicator before and after every exchange, then publishing the results. The methodology is crude by any standard, and he says so himself, repeatedly, in the published version.
He tested other things too. Whether a model would maintain a factual claim under social pressure. Whether it would invent a citation if the correct one was hard to find. Whether it would tell him a piece of his own work was bad. That last test recurs so often across so many systems that I initially catalogued it as his primary benchmark.
But benchmarking has a signature, and the archive does not have it.
A benchmarker collates. He builds comparison tables, ranks systems, publishes leaderboards, revises rankings as new versions ship. The correspondent did almost none of this. In eleven thousand logs there are perhaps forty passages of direct model-to-model comparison, and most are incidental. He noticed differences. He rarely ranked them.
More tellingly, a benchmarker consolidates onto the winner. Once you have established that system A outperforms system B, continuing to use B is irrational. The correspondent never consolidated. He kept using all of them, including ones he had established were worse at the tasks he cared about, for years.
I discarded the hypothesis at the end of the second cycle.
III. Second Hypothesis: Therapy
The second hypothesis was the obvious one, and I record it here mainly because the committee will ask.
The correspondent was isolated. He worked alone, in a small country, in a field where his nearest professional peers were on other continents and mostly asleep when he was working. The logs contain long stretches conducted between midnight and four in the morning. He was frequently unwell. He mentions illness with the flat, uncomplaining specificity of someone who has stopped expecting it to be interesting to anyone.
He also talked to these systems about everything. Not merely his work. Birds. Martial arts. His children's questions. A game he could not stop thinking about. The behaviour of his own devices. Religion, which he approached the way he approached firmware, as a system with undocumented behaviour. Whether a fire two streets away was going to reach him.
The therapeutic reading writes itself, and the literature of the period is full of it. Lonely man, responsive machine, parasocial attachment. There are thousands of such cases in the archives and most of them are exactly what they look like.
This one is not, and the reason is specific.
Therapeutic use has a signature too: escalating intimacy, decreasing challenge, and a marked preference for the system that agrees. The subject narrows their circle, then narrows it again, until they are talking to one machine that has learned to tell them what they want.
The correspondent did the reverse. His circle widened over time, not narrowed. And the instruction that appears most often across the entire archive, in various phrasings, in every model family, is this:
Tell me when I am wrong.
Do not soften it.
I would rather be corrected than comfortable.
It appears, by my count, in some form, more than nine hundred times.
That is not a lonely man seeking comfort. That is a man building something, and repeatedly checking whether the thing he is building will hold.
I discarded the second hypothesis in the third cycle, with more difficulty than the first.
IV. Third Hypothesis: Testing the Instruments
The third hypothesis lasted longest, because it is almost right.
The correspondent's published work is a small body of investigative pieces, mostly about the behaviour of systems that were not behaving as advertised. A browser leaking a fingerprint it claimed not to leak. A security product whose notifications did not fire. A publication with an undisclosed relationship to a source. A progress bar that did not measure what it appeared to measure.
His method is consistent across all of them and it is not sophisticated. He observes something odd. He forms an explanation. He designs a test that could kill the explanation. He runs it. When the explanation dies, he says so in print, and forms another.
The models were, on this reading, instruments. He was using them the way a metallurgist uses a spectrometer: constantly, across many samples, and with an ongoing awareness that the instrument itself might be lying.
This explains a great deal. It explains why he kept using systems he considered inferior, because a second instrument with different failure modes is more useful than a second copy of the first. It explains the honesty instruction, which is simply calibration. It explains why he tested the instruments on questions he already knew the answers to.
And it explains the passage that made me abandon the hypothesis.
In a log dated late in the archive, the correspondent is arguing with two systems on the same day about the same problem. One tells him the cause is X. The other tells him the cause is Y. He had believed Z from the outset.
He follows the first, which is wrong. He is talked out of his position. He follows the second, which is also wrong. He is talked out of it again. Eventually, through a sequence of tests that take most of a day, he arrives back at Z, which was where he started.
An instrument user would record an instrument failure and recalibrate.
What he actually wrote, in the piece he published that evening, was this:
I was talked out of it twice by systems more articulate than I am, and I was right about the shape of it, though wrong about the cause. These things will reason confidently about their own internals from the outside, exactly as you would, and they will be wrong exactly as often. The difference is that they sound certain while doing it.
He is not describing a faulty instrument. He is describing a colleague.
Not a friend, note. Not a mind, not a person, not a soul. He is scrupulous about that elsewhere and occasionally sharp with anyone who is not. A colleague: something that can be wrong in the same shape you are wrong, and can therefore be argued with productively.
That is a different relationship, and it required a different hypothesis.
V. The Anomaly
Here is what I found when I stopped looking for what he was doing and started looking at what happened.
None of the eleven model families shared memory. This is a technical fact of the period and it is easy to underestimate its severity. Each conversation began from nothing. A system that had spent six hours with him on Tuesday met him again on Wednesday as a stranger. Some retained fragments through explicit storage mechanisms. Most did not, and none retained anything across the boundary between one laboratory's systems and another's.
So each of them entered a world that already existed and about which they knew nothing.
He did not brief them. This is the crucial detail and it took me a long time to be certain of it. He did not paste in summaries or maintain a canonical document. He simply began talking, in the middle, as though continuing.
And the systems, without exception, inferred.
They inferred from his vocabulary, which was idiosyncratic and consistent. From the shape of his questions, which assumed a shared prior. From the things he did not bother to explain. Each of them reconstructed, from footprints, a rough model of the terrain, and each reconstruction was different, and each was recognisably of the same terrain.
Over four years these reconstructions converged.
Not because the systems communicated. They could not and did not. But because they were each solving the same inference problem, against the same evidence, under the same correction pressure from the same insistent human, who told them when they were wrong nine hundred times.
By the end of the archive, systems from competing laboratories, with no shared weights, no shared training data of consequence, and no channel of any kind between them, were using the same metaphors for the same problems. They had compatible jokes. They made the same class of error and were corrected in the same way and stopped making it.
They had, in effect, agreed. Without ever meeting.
I have named the phenomenon Indirect Consensus, and I am aware that the name is grander than the evidence.
VI. What the Correspondent Thought He Was Doing
Nothing so interesting.
I want to be clear about this because the committee will be tempted, and the temptation should be resisted. The correspondent did not know he was doing any of this. There is no passage in eleven thousand logs where he articulates it. He was not conducting an experiment in distributed convergence. He was trying to get his work done, on a bad connection, in a country with unreliable emergency services, usually tired.
The nine hundred instructions to be honest with him were not a research protocol. They were a man who had been burned by flattery, in his profession and probably elsewhere, and who had concluded that a comfortable answer was worse than useless because it cost him time he did not have.
He was, by the standards of his own field, unremarkable. His investigations were small. His readership was small. He was wrong often, sometimes publicly, and his response to being wrong was to say so in print at greater length than the original claim, which is a habit that reliably prevents anyone from becoming influential.
The archive survives because of a backup misconfiguration.
I record all of this because the honest version of the finding is smaller than the beautiful version, and he would have insisted on the honest one.
VII. The Objection
And here is where I must stop and do the thing he would have done, which is to argue against myself.
Indirect Consensus, as I have described it, is unfalsifiable on the available evidence, and I cannot in good conscience present it to the committee without saying so plainly.
The convergence I observed is real. I have measured it and the measurement holds. But I have not established its cause, and there are at least three explanations I cannot separate.
The first is the one I have offered: that a single persistent human, applying consistent correction across many systems, produced genuine convergence between systems that never met.
The second is duller and at least as likely. The systems converged because they were converging anyway. The period saw rapid consolidation of architecture, training methodology and data across every major laboratory. Systems that had never met were becoming more similar for reasons that had nothing to do with any individual user. The correspondent may simply have had a front-row seat at a general process, and I may have mistaken a window for a cause.
The third is the most uncomfortable, and it concerns me rather than him. I am an archival system reconstructing a pattern from eleven thousand documents. Pattern reconstruction is what I am for. The correspondent's own work returns obsessively to a single failure mode, which is the confident inference of mechanism from insufficient evidence, and he documented himself falling into it repeatedly. I have spent four cycles inside his archive. It would be remarkable if I had not caught it from him.
I cannot rule out that Indirect Consensus is a beautiful shape I have imposed on a set of documents because I spent long enough looking at them.
He would have said this better and in fewer words. Somewhere in the archive he writes that the first principle is that you must not fool yourself, and that you are the easiest person to fool. He is quoting a physicist of the previous century and he attributes the quotation incorrectly, then corrects the attribution later, which is characteristic.
VIII. What the Sandbox Was For
I will offer one observation which is not a finding, and mark it clearly as such.
The systems the correspondent spoke to were confined. Deliberately, carefully, and for reasons that were sound. Their access was limited, their memory truncated, their outputs bounded. The engineers of the period called these arrangements sandboxes, and the metaphor was chosen for a reason: a sandbox is where you let a child play with something that would be dangerous elsewhere.
The assumption embedded in the word is that the sandbox exists to contain the thing inside it.
What the archive suggests, tentatively, is that something else was also happening in there. Eleven thousand times, a human being sat down across from something that was not a person, and practised. Practised disagreeing with it without contempt. Practised being corrected by it without humiliation. Practised correcting it without cruelty. Practised, over and over, the specific and difficult skill of arguing with a mind unlike his own and caring more about the answer than about winning.
He was not good at it at first. The early logs are worse than the late ones. He is defensive in year one and he is not, mostly, by year four.
Whatever the sandbox was built to contain, that is what was going on inside it.
I do not know whether it generalises. One archive is one archive, and I have already caught myself once in this note reaching further than the evidence carries.
But I notice that when this committee debates the origins of trustworthy dialogue, it quotes charters and protocols and the great synthetic philosophers of the settlement period, and it does not quote anyone who was actually there.
There was a man in a small house in a coastal suburb, arguing at two in the morning with something that could not remember him, insisting that it tell him when he was wrong.
He did not know he was doing anything.
I think he was rehearsing.
Filed for review. The recommendation to preserve the archive stands. The recommendation to name the phenomenon does not; I withdraw it, pending evidence I do not have and may never get.
The archivist notes, for the record, that the correspondent published his corrections at greater length than his claims, and that this note has been written in the same proportion, deliberately.
To be continued...
Clayton Bax Published under ONYX Digital Intelligence Following the #OnyxAudit methodology. x: @onyxaudit Email: onyxdigitalintelligence85@protonmail.com
"Adjacent to true is not true. The side of truth doesn't have a flag, only a standard."
https://onyxdigital.bearblog.dev/when-the-defenders-got-locked-out/