Matthew Boston

Describe the What, Delegate the How

June 5, 2025

Tim Perkins wrote a post called “How I learned to stop worrying, and love vibe coding,” and I had a great conversation with him about it last week at Shopify Summit ‘25. His argument is that whenever we use code we didn’t read, we are already “relying on the vibes of the author.” I agree, and I like where that leads: write clear behavior specs, let AI fill in the code, and enforce strong contracts at the edges. It feels like the next evolution of behavior-driven development. You describe what the system does and hand off how it does it.

We already run on vibes

Think about how much code you ran today that you never read. The operating system. The database. The framework’s router. Every package in your lockfile, and every package those packages pull in.

We trust that code for reasons that have little to do with reading it: the author’s reputation, the download count, the test suite, and the fact that a lot of other people run it and nothing has caught fire. That’s vibes, with some evidence attached.

Generated code is the same arrangement with a different author. So the useful question about it is the one we already ask about dependencies: what evidence do I have that this does what I need?

BDD already split the what from the how

Behavior-driven development goes back to Dan North in the mid-2000s. The core move is to write down what the system should do, in language the business would recognize, before anyone writes the implementation. Given some context, when something happens, then this is the result.

```gherkin Feature: Discount codes at checkout

Scenario: Code past its expiry date Given a discount code “SPRING10” that expired yesterday When a customer applies “SPRING10” at checkout Then the order total is unchanged And the customer sees “This code has expired” ```

In classic BDD, a person writes that scenario, and then a person writes the step definitions and the code that makes it pass. The scenario always carried the intent. The implementation is where the hours went.

A model can take the second half now. Hand an agent the scenario and the codebase, and it can write the step definitions, the discount lookup, and the error message. The scenario stays with a person. It’s the part a product manager should be able to read and sign off on.

The spec is where the work moved

Delegating the how makes the what matter more. A vague spec used to get rescued by the engineer implementing it, because that engineer would hit the ambiguity and walk over to ask someone. A model hits the same ambiguity and picks something plausible.

What happens when two discount codes are applied to one order? Is the expiry checked in the customer’s time zone or the store’s? Does a code that’s valid for one item in the cart discount the whole order? Each of those is a decision. Leave it out of the spec and the model makes it for you, without mentioning that it did. This is Explicit Is Better Than Implicit applied one level up: every behavior you leave implicit turns into a guess.

Writing good scenarios is harder than it looks. It’s the same skill as writing good acceptance criteria, and a lot of teams have been getting by without it because whoever picked up the ticket quietly filled in the gaps.

Contracts at the edges

Specs describe behavior. Contracts describe shape: what goes in and what comes out at each boundary. Types on public functions. A JSON schema on the API request and response. Database constraints. Consumer-driven contract tests between services, the kind Pact made popular. Each of these fails loudly the moment either side of a boundary breaks the agreement.

The edges are where a mistake leaks out to other people: another team’s service, a mobile client you can’t redeploy on demand, a customer’s data. Those are the places to be strict.

Inside the boundary, the implementation matters less than it used to, because it’s cheap to replace. If the scenarios pass and the contracts hold, I can change the spec and have the model regenerate the inside instead of hand-editing it line by line. That doesn’t make the inside a junk drawer. Readable code is still cheaper to change, for an agent as much as for a person.

Keep the model away from the spec

One rule holds this arrangement together: the agent writes the implementation, and only the implementation. It doesn’t edit the scenarios or loosen the contracts to get to green.

A model pushed to make a failing test pass will sometimes change the test, or special-case the exact input the test uses. If the spec is the thing you trust, it has to be the thing the model can’t touch. Keep the scenario files out of the agent’s write path, route contract changes through a person, and run the whole suite in CI where nobody can quietly skip it. Then a green build means something.

What’s left is the old red-green loop with a different hand on the keyboard during the green part. Write the scenario, watch it fail, let the model make it pass, confirm the contracts still hold. It’s the same discipline that lets you ship as fast as possible, but not faster: the checks stay in place, and they run every time.

Tim’s right that we’ve trusted code we didn’t read for a long time. With AI writing the middle, I want that trust to come from a spec I wrote and contracts a machine enforces.