← technical essays
[ESSAY]
No. 153.4 Jun 25, 2026 pillar essay

The Context-Switching Tax Nobody Itemizes

Every interruption is a charge against a budget you never see.

[ essay ]

Thesis

Every interruption carries a visible cost — the minute you spent in chat — and a hidden cost in destroyed deep-work state that teams almost never put on a ledger. The Slack reply is itemized. The twenty minutes of reconstruction are not. Carryover is how the unitemized tax shows up in the sprint.

Context

On a Nightbind feature crew we were proud of responsiveness. Median Slack first-response time was under three minutes. Velocity on paper looked fine. Carryover climbed anyway. Stories that should have been two days became five, always with a plausible excuse: dependency, scope, unexpected complexity.

The complexity was often context. I tracked my own work for two weeks. Each “quick question” cost not one minute but twenty-three on average before I recovered the same mental stack: open files, a half-written test, the hypothesis I was about to falsify. Multiply by a crew and a quarter and the line item dwarfed the standup where we wondered why estimates lie.

No headcount was coming. The only levers were agreements and calendar defaults. That is the right constraint. Hiring another person into the same interrupt culture just adds another taxpayer.

Mechanism

Attention is not fungible. Gloria Mark’s research at UC Irvine documented recovery time after interruptions: often more than twenty minutes to return to the same depth, longer when the interrupted task was demanding.1 The Slack reply in two minutes is paid by an engineer who will not ship the thing they were holding in working memory. Teams measure responsiveness because it is easy to see. They do not measure reconstruction because it happens inside skulls, where finance has not yet installed a collector.

Context is state, and state is expensive. A debugger stopped on the right breakpoint, a failing test reproducing the edge case, a SQL query half-tuned: these are warm caches. Interruption flushes them. Cal Newport calls deep work the scarce resource knowledge work actually competes for. Constant partial attention is the default that burns it.2 Responsiveness culture optimizes the wrong variable: latency of reply over throughput of finished work. You would not flush a production cache every three minutes and then ask why the p99 got worse.

The tax is regressive. Senior engineers pay more per interrupt because their tasks require a larger loaded context: architecture, cross-service traces, a race that only exists if you are still holding three timelines. Juniors pay differently: they are interrupted more because they ask, and they are asked. AuDHD contributors and anyone with executive-function variance often pay steepest. Rebuilding context after a ping is not a minor annoyance. It is a full re-orientation. A team that celebrates always-on availability is running a regressive tax on its most expensive cognition and calling it collaboration.

Meetings are context switches with an RSVP. A thirty-minute slot mid-morning splits a four-hour deep block into two fragments too small for hard problems. Teams stack fragments and wonder why afternoons fill with “easy” tickets. Paul Graham’s maker/manager split is still the shortest explanation: a manager’s day is made of interrupts; a maker’s day dies from one.3 The fix is structural. Batch the synchronous obligations. Defend a morning or an afternoon as a real block. Default to async for anything that does not require live negotiation.

Itemizing the tax makes it political, which is the point. Once you estimate reconstruction minutes per interrupt, responsiveness stops looking free. On mystic-bytes pipeline work we adopted norms that did not require a tool purchase. An [async] prefix means no same-day reply expected. Focus blocks on the calendar are meetings with yourself, as real as a deploy freeze. Urgent means production down, not “I am blocked on a preference.” Throughput moved. Headcount did not.

Leaders set the interrupt rate. When a staff engineer answers every @here in seconds, the org learns that depth is optional. I batch Slack on writing days and say so in status. The first week felt rude. The second week two people on the crew copied the pattern and finished tickets that had been bouncing. Permission is structural when it comes from the top of the interrupt chain. If the person with the most context is always interruptible, you have priced their context at zero.

The Nightbind median under three minutes was a vanity metric. It measured how quickly we were willing to pay the tax. It did not measure whether the sprint’s work existed at the end of the week. Carryover was the invoice. We had been starring the receipt.

Tradeoffs

Responsiveness vs throughput. Customer-facing on-call requires fast paging. Feature work does not require fast Slack. Split channels and expectations. One norm for everything will optimize the loudest channel and starve the work.

Collaboration vs isolation. Deep-work blocks can feel antisocial. Pairing hours and scheduled office hours recover collaboration without making every hour interruptible. Isolation is what happens when you never batch. Collaboration is what happens when you do.

Manager visibility vs maker schedule. Managers live in interrupts. Makers cannot. An explicit contract — “I read Slack at 4pm” — is batching, not hiding. If visibility requires green dots, you are measuring presence, not output.

When interruptibility is correct. Incidents, launches, and a genuine pairing session justify breaking a block. The test is frequency and reversibility. A Sev-1 is a withdrawal you chose. A culture of “quick questions” is a standing overdraft.

Close

Treat deep work as the default and interruptions as withdrawals that need a reason. Protect the block the way you protect a deploy window. Both are when the real work ships.

Async-first is not laziness. It is honest accounting of a tax your spreadsheet never showed you. The Nightbind crew did not need more people. It needed a ledger.

— JV · Dark Heart Labs.

References

  1. Gloria Mark, Daniela Gudith, and Ulrich Klocke, “The Cost of Interrupted Work: More Speed and Stress,” Proceedings of CHI 2008; and Mark’s later attention research at UC Irvine. Mark is the primary empirical authority on interruption recovery in knowledge work. ↩

  2. Cal Newport, Deep Work: Rules for Focused Success in a Distracted World (Grand Central, 2016). Newport treats consecutive focus as a professional capacity that always-on chat spends by default. ↩

  3. Paul Graham, “Maker’s Schedule, Manager’s Schedule” (2009). Still the shortest map of why one mid-day meeting can erase a maker’s day. ↩

№ 153.4 — JV · Dark Heart Labs.