LeetCode Interviews Are on Borrowed Time
If LeetCode-style puzzles are the main way you decide whether to hire an engineer, you’re probably passing on some of the best people available. You’re also leaning on a test that GitHub Copilot and ChatGPT can already pass. I think it’s only a matter of time before these interviews stop telling you anything useful. The engineers worth hiring are problem solvers, leaders, and strategists, and the interview should go looking for that.
What the puzzle measures
A typical round: 45 minutes, a shared editor, and a problem with a known optimal solution. Find the longest substring without repeating characters. Merge k sorted lists. The candidate who has seen the pattern recently, a sliding window or a heap, gets there. The candidate who hasn’t spends half the hour rediscovering it while someone watches.
That does tell you something. The candidate can write code that runs and can talk about Big O. Past that, it mostly tells you how many hours they spent on practice problems last month.
Compare it to an ordinary week of engineering work. Reading code someone else wrote. Figuring out what a vague ticket is asking for. Tracing a production bug through several services. Deciding whether a new dependency is worth carrying. Convincing a teammate that the schema has to change. Very little of that looks like merging k sorted lists from memory against a timer.
The tools already pass
Paste one of those classic problems into ChatGPT or Copilot Chat and you get a working solution in seconds, usually with a complexity analysis attached. These problems are some of the most written-about code on the internet, so of course the models are good at them.
That has two consequences. First, the skill being tested is one the job is steadily handing to tools. Second, the interview usually forbids the tools the candidate will use on their first day. You end up measuring how well someone works without the setup they’ll have, which is a strange way to predict how they’ll work with it.
There’s also nothing stopping a remote candidate from keeping an assistant open in another window. You can try to police that, or you can stop relying on a test that a chatbot passes.
Hire for the parts the tools don’t do
When code gets cheap to produce, the judgment around it is what’s left to hire for.
Problem solvers take an ambiguous request, ask the questions that narrow it down, and find the version worth building. Leaders, at any level, move a group toward a decision; they explain tradeoffs clearly and make the people around them better. Strategists know which problem is worth solving this quarter and what a choice will cost a year from now.
You won’t see any of that while someone inverts a binary tree. You have to design the interview to surface it.
Interviews that look like the job
A few formats that do a better job, and all of them can be standardized:
- A work-sample task. Give the candidate a small, working repository and a realistic change: add an endpoint, fix a failing test, handle an edge case. Keep it to an hour or two. Let them use their normal tools, Copilot included, and spend the follow-up conversation on why they made the choices they did.
- Pairing on a real-ish problem. Hand over a small app with a bug in it and work through it together. Watch how they read unfamiliar code and how they test a hunch. The first ten minutes of someone debugging tells you more than an hour of puzzle-solving.
- A code review exercise. Give them a pull request with a few planted problems: a missing test, an N+1 query, a confusing name, user input concatenated into a SQL string. See what they catch and how they word the feedback. Review is a big part of the job, and it gets bigger as more of the code comes from tools.
- A system design discussion. Skip the URL shortener everyone has rehearsed and pick something close to your own domain. Look for questions about requirements and load before the boxes and arrows, and for honest answers about what breaks first.
Letting candidates use AI tools in these exercises is a feature. You get to watch whether they read what came back and whether they notice when it’s wrong. That’s a skill worth hiring for now.
Puzzles are cheap and consistent
The best argument for LeetCode rounds is operational. Any engineer can run one, and every candidate gets a comparable question.
Work samples can be just as consistent. Use the same exercise for every candidate in a role, and write the rubric before the first interview. It takes more setup up front. That setup is cheap compared to a bad hire, or to the great engineer you never met because they didn’t grind practice problems that month.
Puzzle rounds were always a proxy for engineering ability, and the proxy keeps getting easier for a tool to pass. Test for the work you’re hiring someone to do.