Whether You Think AI Can or Can't, You're Right
There’s a line often attributed to Henry Ford: “Whether you think you can, or you think you can’t, you’re right.” Swap in AI and it still holds. Whether you think AI can write working software or you think it can’t, you’ll find the evidence to prove it, because how you use the tool decides which result you get.
Two runs of the same experiment
Picture two engineers with the same model and the same codebase.
The first one opens the agent and types a sentence describing the feature. No pointers to the relevant code, no conventions, no definition of done. The agent produces something. It invents a helper that already exists under a different name, ignores the pattern every other module follows, and writes tests that assert the mock returned what the mock was told to return. Nobody reviews it closely, or someone does and finds garbage. Conclusion: AI can’t code.
The second engineer writes a short spec, points the agent at the code it needs to read, makes sure it can run the tests after every change, and reviews the result before merging. The agent ships a working feature. Conclusion: AI can code.
Both engineers are reporting what they saw. The difference is in the setup. The first engineer designed an experiment that could only fail, then reported the failure as a finding about the tool.
Treat it like a new hire
Think about how you’d bring a capable engineer onto your team. You wouldn’t hand them a one-line ticket on day one, skip the codebase tour, and merge their first PR without reading it. And if you did and the PR was bad, you’d blame the onboarding before you blamed the engineer.
An agent is a new hire who starts over every session. It knows a lot about programming and nothing about your codebase: your conventions, the decision the team made last quarter, the utility module everyone uses for dates. Agents are stateless, so every session is day one, and the onboarding has to be written down where the agent can read it.
A new hire needs clear specs, tight feedback loops, and an actual review of their work. So does the agent.
Clear specs
A spec for an agent can be short. It should say what the problem is, what done looks like, where the change belongs, and what to leave alone. Something like this:
```markdown ## Goal Let users export invoices as CSV from the billing page.
Done when
- The export includes every invoice visible under the current filters
- Amounts use the account’s currency format
- Tests cover an empty export and a filtered export
Constraints
- Follow the existing CSV export in the reports module
- No new dependencies ```
Most of that is context the agent can’t infer on its own. The parts that apply to every feature, the conventions and house rules, belong in a SKILL.md so you write them once. The parts specific to this feature come from doing the research before the agent writes a line.
Tight feedback loops
New hires get better fast when they find out quickly that something’s wrong. A failing test five seconds after a bad change teaches more than a review comment two days later.
Agents work the same way, only more so. An agent that can run lint and the relevant tests in seconds catches its own mistakes and fixes them before you see them. That’s why coding agents need a faster feedback loop than most teams have. The skeptic’s setup usually has no loop at all. The agent writes the code, declares itself done, and the first signal it ever gets is the skeptic’s verdict.
An actual review
You read a new hire’s pull requests carefully, at least until you know how they work. Do the same for the agent. Review is where you catch the plausible but wrong change: code that looks reasonable, passes a shallow test, and misses what the feature was for. Point your attention at whether the change solved the problem it was meant to solve, and hold it to the same bar as any other code, because code quality still matters.
If nobody reviewed it, the bad code that shipped is a process failure, and the tool took the blame for it.
The variable is the engineer
None of this is new. Good teams wrote specs, kept their test suites fast, and reviewed each other’s code long before agents showed up. What changed is the cost of skipping it. A human who gets a vague ticket walks over and asks a question. An agent fills the gap with a confident guess and keeps going.
So when someone holds up a pile of AI-generated garbage as proof the tools don’t work, the useful question is what the agent was given to work with.
The model is the same for everyone. How much of your own engineering discipline you bring to the conversation decides what comes out.