Opus Incertum
Opus incertum is Roman masonry. Irregular stones, no uniform courses, pressed into a concrete core. The Latin means uncertain work. It is also, near enough, a description of what follows.
I have never been sure exactly how Anthropic measures token use, or what the usage percentage in the app actually represents. I have been tracking it for a few months.
When 4.8 launched I ran a quick test. Thinking on, 4.8 Max, and it burned roughly 4 to 6% per long exchange. Voice, TTS, live chat, made no difference. The methodology was crude. Too broad, not accurate enough, and I knew it at the time.
Opus 5 launched overnight. This time I did it properly. Here is what I found.
Everyone complains about the usage limits, but I have never seen anyone actually measure them. So I measured them.
This is not a complaint and it is not a hit piece. It is an attempt to understand a number that affects everyone using this thing daily, written by someone who uses it daily.
It turns out the thing that drains your session is not the thing anyone assumes. It is invisible, undocumented, and it punishes the most natural way to use the app.
How this started
I was ill on Friday and did almost nothing. So I opened Claude on Saturday morning with no story and mild desperation, half expecting to end up writing about ChatGPT because at least that would be something.
Instead I got a launch banner. Opus 5 was live.
I should have known. Someone mentioned weeks ago that the Fable relaunch might be a precursor, and I did not follow it up. That is normally the kind of thing I am on top of. I was not.
So I did what I do. Screenshotted my usage, zero percent, clean slate, and sent one message into a project chat I had been working in for about forty-eight hours.
Nineteen percent. One turn.

Sent another. Thirty-six.
The argument
I asked the model what was happening and it told me, flatly, that I was measuring wrong. Chat too long, project too heavy, baseline contaminated. Start clean, then measure.
It read as combative. It was not. Nothing it said was untrue and nothing was pointed. But when a system tells you your method is wrong while you are measuring that system, it lands a certain way.
I ignored it and took the numbers anyway.
That turned out to be the single most important decision of the day. Every later test in fresh chats produced completely different figures, and only the contrast between the old chat and the new ones exposed the mechanism. Had I followed the advice, I would have destroyed the comparison that eventually explained everything.

What I thought first, and why it died
The obvious explanation was caching. Large context processed once at a premium, then read back cheaply. Nineteen, seventeen, then a collapse to one percent. A textbook sawtooth.
Easy prediction: leave the chat idle past the cache window, come back cold, and the next turn spikes.
I waited ten minutes. Came back. One percent.
Declared dead. It was not dead. The test was simply too short. The window is longer than ten minutes, so I never went cold at all. I designed an experiment that could not fail in the direction I needed, then read the null result as a refutation. Remember this.
What I thought second, and why that died too
Next theory: the model's own output. Reasoning tokens are generated tokens, output bills at a multiple of input, and thinking blocks are invisible.
In a clean chat, replies cost one, two and three percent in proportion to their length. Clean signal.
So I tested it head-on. Same chat, same tier, same minute. One-word answer, then a thousand-word answer.
Both cost one percent.
Declared dead. Also not dead. The meter reports whole percentages, and one percent spans anything from 0.5 to 1.49. A genuine twofold difference fits entirely inside the rounding. Second theory killed by an instrument that could not resolve it.
The flat band
By now I had fifteen or so readings and almost nothing I changed made any difference.
Context size, model tier, extended thinking, time gaps under half an hour, retrieval, tool calls, uploaded images, regenerating a reply, creating a brand-new chat inside a 178-file project. All of it between zero and three percent. The project itself, 178 files at 40% of capacity, was effectively free to carry.

One thing had ever cost more, and I could not reproduce it.
Retrieval, which is real but is not the answer
A fresh chat inside the project, no history, deliberately broad question: search the project and list every file mentioning the kill switch.
Twelve steps. Five to seven percent. A second broad sweep on FROST cost four more.

The same class of question in the warm forty-eight-hour chat had cost one percent in four steps. Context tells the model where to look, so it fetches instead of sweeping. Real effect, real cost, and I thought at the time it was the finding.
It was not. Five to seven does not explain nineteen.
The answer
Late in the day I went back to the old chat after being away for the better part of an hour and asked it something trivial. Short reply, no thinking to speak of, no retrieval.
Eighteen percent. Eighty-two to a hundred, in about a minute.

Then the session reset. First message of a clean window, same old chat, one word: Ready? Three seconds, no thinking, one-word answer.
Nineteen percent.
Same figure twice, under completely different session conditions. The mid-session reading rules out any session-opening toll on its own. What both share is that I had been away long enough for the conversation to go cold.
I then measured two more cold re-entries, into conversations of different sizes:
| Conversation | Tier | Cost |
|---|---|---|
| Fresh project chat, about five turns | Max | 5% |
| Working chat, about twenty-five turns, image-heavy | Medium | 11% |
| The forty-eight-hour project chat | Max | 19% |
Those three chats differ in turn count, image load and age all at once, so treat the ordering as solid and the figures as indicative.
So: the cost is re-entry into a cold conversation, and it scales with how much conversation there is. The cheapest of the three was on the most expensive tier, which rules out the model setting as the cause. It is the history being reloaded, not the model and not the project. The 178 files were free all day because they were never what I was carrying.
What to actually do
The expensive pattern is grazing. Dipping into a long chat every hour or two across a day means paying the re-entry toll every single time. Four or five casual one-line questions in a big chat can consume most of a window before you have said anything substantial.
Sustained work in the same chat is cheap and gets cheaper. Within a warm window, cost drops to output length. Two to three percent for a long reply, near zero for a short one.
So the rule is not "start fresh chats to save usage," which is what I would have told you this morning and what the draft of this piece said four hours ago. It is: finish what you start, in one sitting. If you are only going to ask one thing, ask it somewhere small.
What I cannot claim
The meter reports whole percentages. Everything here is ratios and orders of magnitude, never figures.
The thinking-step counter increments live, so any count read mid-flight is a floor. I logged one twelve-step turn as four before I noticed. Worse, step counts appear to no longer be retrievable for past turns in this build. The chip now describes what was being considered without showing the steps, so I could not go back and verify. If that changed with Opus 5, it removes the only price signal the interface offered.
My session bar and weekly bar diverged repeatedly, 44 against 38 at one point. Different things, and I only tracked one carefully. The monthly dollar figure and my credit balance never moved at all across a session that consumed the entire bar. Three counters, at least two of which do not mean what their labels imply.
The one thing I could not test is where the expiry sits. One percent at ten minutes, eighteen at roughly an hour. Somewhere in that gap a conversation goes cold, and I have no way to narrow it without a day of waiting between single messages.
A note on the models
Two Claude instances argued opposite sides of this to me with equal confidence. One insisted it was input. One insisted it was output. Both were partly right, neither could see its own billing, and both of them talked me out of the correct answer at least once.
I held the same position from my first message this morning, that those opening turns were the model pulling everything from the project, and that afterwards it would not need to. I was talked out of it twice by systems more articulate than I am, and I was right about the shape of it, though wrong about the cause. It was not the project being pulled. It was the conversation.
That is the finding that has nothing to do with tokens. These things will reason confidently about their own internals from the outside, exactly as you would, and they will be wrong exactly as often. The difference is that they sound certain while doing it.
"The first principle is that you must not fool yourself, and you are the easiest person to fool." Richard P.Feynman
Clayton Bax Published under ONYX Digital Intelligence Following the #OnyxAudit methodology. x: @onyxaudit Email: onyxdigitalintelligence85@protonmail.com
"Adjacent to true is not true. The side of truth doesn't have a flag, only a standard."