Skip to content
DAbuild-your-owncontext-compaction

The median was hiding it

I claimed I don't clear my context. So I measured 569 of my own sessions — and found the median was hiding the whole story.

The median lies when I measure my Claude Code sessions. The typical session hasn't gotten longer — the raw median actually fell. But that's the wrong conclusion, because I'm doing far more small things now. Once I separate the quick tasks from the real work sessions, the pattern shows up: the long sessions got dramatically longer, and that's where most of the work lives. Compact doesn't beat clear because every session has to run for days. It beats clear because the few sessions that actually carry the work can only exist if I let them compact and keep going.

Most sessions are still short

I wrote earlier that I almost never clear my context anymore. That was true as a feeling — but you have to be careful with that. A feeling isn't data.

So I pulled 569 of my own interactive Claude Code sessions from March to June 2026 out of pks brain (a tool that's read every session I've run) and looked at how long they actually are.

How long are my sessions, really?
569 interactive sessions, Mar–Jun 2026
Median 1.3h · 45% under an hour · but 17% run more than a full day.

The everyday session is short: median 1.3 hours, and 45% finish in under an hour. Most of what I do is quick, scoped tasks — fix this, rename that, check one thing. That's normal, and it should be. Not every prompt is a marathon.

The naive answer is wrong

So did my sessions get longer? The lazy way to check is the median, and the median says no — it dropped from 109 to 58 tool-calls. If I stopped there, I'd conclude my sessions got shorter.

But that's a bias, and it's worth being honest about it.

I'm doing more and more
Total token spend (bars) and session count (line) per month
Volume climbed hard — which is why raw counts mislead and we have to normalize.

I'm doing a lot more now. My token spend went from 3 billion in March to 13.8 billion in May. When you add a flood of new quick tasks, the median drops even if your real sessions are growing — the little ones outvote the big ones. Raw counts can't answer the question. I have to normalize.

Drop the small stuff

So I filtered. Keep only the "real" sessions — at least 20 prompts or at least an hour — and throw away the one-offs.

Drop the quick tasks and my sessions grow
Median duration — all sessions vs. only 'real' ones (≥20 prompts or ≥1h)
Median for real sessions went 3.8h → 18.3h. It only fell overall because I do more quick tasks.

Now the median for a real session goes from 3.8 hours in March to 18.3 hours in June. The overall median only fell because I do more quick tasks; my real sessions didn't get shorter, they got much longer. That's the whole point of this post, and it's the kind of thing you only see if you stop trusting a single summary number.

The work lives in the tail

Counting sessions treats a 5-minute rename the same as a 20-day build. So weight by actual work instead — tokens.

Where the work actually lives (weighted by tokens)
Share of the month's token spend in sessions running more than a day
Normalized for volume: ~70% of all real work now happens in long sessions I never clear.

By May and June, ~70% of all my token spend happens in sessions that run more than a day, up from 27% in March. Most of the real work has moved into the long sessions I never clear. The short ones are many, but they're light.

The multi-day sessions are growing

Multi-day sessions grow every month
Share of sessions per duration band
Share running at least a day: 10% (Mar) → 26% (Jun).

The cleanest trend in the whole dataset: the share of sessions that run at least a full day went from 10% in March to 26% in June. One in four of my sessions now stays open for a day or more.

What I can't claim

Here's where the data fights me, and I'd rather say it than hide it.

What I can't claim: that I compact more
Share of sessions that compacted ≥1× + average per session
Peaked in April (46%), then fell to 12–15%. Honestly: I compact less often now — partly because there are more short sessions.

I can't claim I compact more and more. The share of sessions that compacted at least once peaked in April at 46%, then fell to 12–15%. I compact less often now — partly because there are so many short sessions that never run long enough to hit the wall. So "I clear less and less over time" is not a claim the data supports. The marathons compact; the growing pile of quick tasks doesn't.

The tail is absurd

Marathon sessions: duration vs. number of compactions — each dot is a session

Each dot is a session. Compactions cluster on the far side of the 24-hour line — and the one outlier ran 478 hours and compacted 66 times.

And then there's the extreme. One single session ran 478 hours — about 20 days — compacted 66 times, and grew to 71 MB. You cannot have a session like that if you clear between tasks. It exists only because it kept compacting and carrying on.

The takeaway

The median lies. "My sessions got longer" is only true for the tail, not for the typical session. A handful of marathon sessions do most of the real work, and most of my sessions are still short, quick tasks — "I don't clear" is about the long ones, not about everything.

And that's exactly why compact beats clear: those long sessions only exist because they compact instead of getting wiped. The short ones don't need it. The long ones couldn't live without it.