Matthew Boston

Behavior Tests Tell You Why the Code Exists

September 3, 2024

You change a method, the suite passes, and you ship. A week later someone reports that a feature you’d never heard of has stopped working. The code behind it was sitting right there in your diff, and nothing told you why it existed. Behavior tests close that gap. Each one names a requirement the system has to meet and checks that it still does.

Tests that check the wiring

Here’s a test you’ve probably seen some version of:

```ruby it “cancels the subscription” do expect(repository).to receive(:find).with(42).and_return(subscription) expect(subscription).to receive(:update!).with(status: “cancelled”) expect(mailer).to receive(:cancellation_email).with(subscription)

CancelSubscription.new(repository, mailer).call(42) end ```

It passes, and it describes exactly how CancelSubscription works today. That’s the trouble with it. Rename update!, move the email to a background job, or load the record with a different query, and this test fails even though the customer sees no difference.

It also can’t answer the questions a future reader will have. Is the email a requirement or a courtesy? What does cancelling actually guarantee? The test restates the implementation line by line, so it knows nothing the code doesn’t already say.

Tests that state the requirement

Here’s the same feature tested by its behavior:

```ruby describe CancelSubscription do it “stops future renewal charges” do subscription = create_subscription(renews_on: Date.new(2024, 10, 1))

CancelSubscription.call(subscription.id)

expect(Billing.charges_due_on(Date.new(2024, 10, 1))).to be_empty   end

it “keeps access until the end of the paid period” do subscription = create_subscription(paid_through: Date.new(2024, 9, 30))

CancelSubscription.call(subscription.id)

expect(subscription.reload.active_on?(Date.new(2024, 9, 15))).to be(true)   end

it “sends the customer a cancellation confirmation” do # … end end ```

Each test goes in through the front door and checks an outcome someone cares about. None of them mention the repository or the order of method calls.

The second test is where the “why” shows up. Someone who simplifies cancellation to revoke access immediately gets a failure named “keeps access until the end of the paid period.” They learn the delay is deliberate before they ship, instead of from a support ticket afterward.

Run the suite with rspec --format documentation and those names print as plain sentences under CancelSubscription. That output is a spec. A new engineer can read it on their first day and learn what the billing system promises without reverse-engineering it from the code.

Testing should enable refactoring

Jimmy Bogard put it in one line: “Testing should enable refactoring, not prevent it.”

A refactor changes structure and leaves behavior alone. That’s the definition. So if a refactor breaks forty tests, those forty tests were checking structure. Teams in that position learn the obvious lesson and stop refactoring. The code hardens in place, and the suite that was supposed to make change safe has made it expensive.

Behavior tests don’t care how the code is arranged. Split the class, inline the helper, swap the query, and the suite stays green as long as the outcomes hold. When it goes red, something a user would notice has changed, and that’s exactly when you want to hear about it.

Mocks still have a job at the edges of the system: the payment gateway, the email provider, anything across a network where a real call would be slow or unreliable. Inside your own code, mock the world, never the subject. The moment a test mocks the thing it’s testing, it’s back to checking wiring.

When the requirement changes

Requirements move, and this is where behavior tests pay off a second time. Say the business decides cancellations should take effect immediately. You open the suite, find “keeps access until the end of the paid period,” and rewrite it to describe the new rule. It fails. You change the code until it passes. Finding the test that encoded the old requirement took seconds, because its name was the requirement.

With wiring tests, the same change means reading through mock expectations, working out which ones encode the old rule and which are incidental, and hoping you found them all. A suite only works as an index of what the system promises if each test is named after a promise.

Name every test after the reason the code exists. Whoever changes that code next will know what they’re about to break.