← technical essays
[ESSAY]
No. 6.4 Feb 11, 2026 pillar essay

Testing Is a Conversation

The suite tells you what the system is actually willing to promise.

[ essay ]

Thesis

A test suite is a conversation between past-you and future-you about what the system is actually willing to promise. It is not a quality-assurance artifact you hang on the wall after the work is done. Every assertion is a sentence the codebase has accepted as binding. Green CI is that sentence, repeated, until someone changes the subject.

Context

The merge queue on a Nightbind-adjacent project had a fast path: green main, small diff, skip the heavy integration lane. Then the fast path stopped being fast in the only way that matters. One operator, one terminal, no rollback window. A bad merge shipped because the lane’s assumptions were never written as promises.

The suite was green. It tested mock shapes and private helpers. It did not test what callers may rely on: changelog present, semver valid, consumer fixtures matching the published schema. Future-me, reading a checkmark, believed a promise the system never made. The conversation had been happening. It was about the wrong thing.

When I rewrote the lane guards as contract tests, the tone changed. Past-me stopped whispering “probably fine.” Future-me got sentences a stranger could read at a terminal without opening anyone’s head. mystic-bytes taught the same lesson later on cover hashes and slug uniqueness. The tests that saved us were the ones that sounded like a release manager talking, not like a mocking library talking.

Mechanism

Tests are executable specifications. Kent Beck framed test-driven development as design feedback: you write the promise first, then negotiate with reality until the code keeps it.1 The suite is the negotiation record. Commits are proposals. Tests are the thread where past-you replies, in a language CI can grade.

Glenford Myers argued decades earlier that testing can show the presence of defects, not their absence.2 That sounds pessimistic until you notice the job. The suite is not omniscience. It is an explicit list of which risks you are willing to detect and which you are accepting. A conversation with omissions is still a conversation. Silent omissions are lies that look like coverage percentages.

Wrong tests document wrong promises. Suites that assert HTTP 200 on the happy path and ignore error shapes teach future-you that failures are decorative. Suites that snapshot entire JSON blobs teach that any change is breakage, even when the contract got better. Suites that mock everything teach that integration is someone else’s problem. That someone is usually the on-call engineer, alone, without a rollback.

I have maintained suites that were a second codebase. Every refactor of a private helper required rewriting six tests that had never protected a caller. Those tests were not careful. They were clingy. They froze an accident of implementation and called it safety. The conversation they recorded was “do not move this furniture,” not “do not break this door.”

The right tests read like an API for trust. Name the behavior a stakeholder would pay you to keep. “Merge queue rejects empty changelog” is a promise to release managers. “Consumer fixture matches published entry schema” is a promise to downstream callers. “Fast path skips the heavy lane only when contract tests pass” is a promise to the operator at the terminal. If you cannot say the sentence without jargon, the test is probably about your feelings toward the code, not about the contract.

Jerry Weinberg treated quality as a relationship among people, not a property of disks.3 Testers, developers, and operators negotiate what “good enough” means under constraint. Automated tests automate part of that relationship. They do not remove the need to choose which promises matter. Coverage tools will happily grade a suite that promises nothing a human cares about.

One operator, no rollback is the extreme case where the conversation must be finished before the click. A manual checklist plus a missing automated contract is folklore. Folklore does not survive a tired Friday. The merge button should invoke promises you can read. If the only person who understands the fast path is the person who wrote it, you do not have a suite. You have a diary.

On mystic-bytes I now ask one question before adding a test: who is the other party in this conversation? If the answer is “the mock,” I delete the draft. If the answer is “the next person who ships a cover,” I keep it.

Tradeoffs

Coverage vs meaningful promises. High line coverage with shallow assertions is noise you pay to maintain. Prefer fewer tests that encode invariants callers depend on over many tests that encode how a function happened to be written on a Tuesday.

Speed vs fidelity. The fast path exists because full integration is expensive. Contract tests at the boundary buy speed without lying, if they track real caller expectations and not wishful stubs. Stubs that always return 200 are a conversation with yourself.

Brittleness vs precision. Overspecified tests raise the maintenance tax. Underspecified tests raise the incident tax. Tune to the stability of the boundary. Public contracts deserve precision. Internal refactor surfaces deserve behavior-level assertions that survive a rename.

When not to test. Exploratory spikes and throwaway prototypes may not have earned promises yet. Write that absence down. Do not let green CI imply guarantees you declined to write. An empty suite that admits it is empty is more honest than a green suite that tests the furniture.

Close

Write tests as if you are dictating the system’s contract, because you are. Future-you is a stranger with fewer context switches and less patience. Past-you owes them sentences, not vibes.

Suites that test the right things become the most useful documentation a project ever has. Suites that test the wrong things become a second codebase you maintain to justify a false sense of safety. The merge queue does not care which one you chose. The operator at the terminal does.

— JV · Dark Heart Labs.

References

  1. Kent Beck, Test-Driven Development: By Example (Addison-Wesley, 2002). Beck’s framing of TDD as design feedback is the canonical reference for tests as specification rather than after-the-fact inspection. ↩

  2. Glenford J. Myers, The Art of Software Testing (Wiley, 1979; later editions). Myers is the standard authority on what testing can and cannot prove, which is the honest frame for a suite that claims to be a conversation about risk. ↩

  3. Gerald M. Weinberg, Perfect Software and Other Illusions about Testing (Dorset House, 2008). Weinberg treats testing as human communication about quality and shared expectations, not as a property of the disk. ↩

№ 6.4 — JV · Dark Heart Labs.