AIMay 28, 2026·10 min read·By ProgramAlpha Team

LLMs, agents and the new engineering interview

Interviews changed when LLMs became part of the stack. How to prepare: tool-calling loops, evals, context engineering, and the judgement that separates an operator from a passenger.

The engineering interview changed the day LLMs stopped being a novelty and became part of the stack. It is not that the classic rounds vanished — a software engineer still needs algorithms, systems thinking, and clean debugging — it is that the target moved. Companies now want engineers who can build software with a model in it: systems that call an LLM, feed it context, check its output, and survive a production pager rotation. The candidate who cannot speak that language is now interviewing against the calendar instead of against the market.

The interview no longer asks whether you can type code. It asks whether you can ship software that includes a probabilistic component and still keep its guarantees.

Why the interview changed

Five years ago, "AI" in an interview meant a token ML question. Today the bar is not research; it is integration. Hiring managers have discovered the hard way that a model API is easy to hit and brutal to operate. Context windows fill with junk, tools get called with the wrong arguments, latency targets get wrecked by streaming, and cost balloons because nobody built an eval. The new interview tests the skills that make those integrations survive: context engineering, tool calling, evaluation, and resource judgement.

The agent loop is the new breadth-first search

The single most-repeated construct in LLM interviews is the agent loop, and failing to know its shape is like failing to know binary search was in 2019. The loop is embarrassingly short:

while not finished:
    response = model.complete(messages, tools=registry)
    if response.tool_call:
        result = run_tool(response.tool_call)
        messages.append(tool_result(result))
    else:
        return extract_answer(response)

That is the whole core. The interview depth comes from the details around it: how you bound the loop so a model that keeps requesting tools cannot burn your token budget; how you validate tool arguments before they hit your real API; how you return errors so the model can recover instead of retry in a straight line; and how you decide when finished is actually true. A candidate who can draw that loop, name its two failure modes, and sketch the safety rail has already cleared most of the round.

Context engineering is the new system design

System design still matters, but the AI variant has two extra dimensions. The first is what goes into the context window — retrieval, not memorisation. Nobody pastes the whole corpus; you retrieve the few chunks that matter, you rank them, you budget token space for instructions versus data, and you design for the day the window fills. The second dimension is tooling: which capabilities your agent exposes, how a tool gets described so the model chooses it reliably, and how function-calling schemas stay in sync with the code that runs them. These two dimensions are genuinely new interview territory, and generic system-design prep alone no longer covers the pattern. The good news is they are learnable: rather than memorising ten architecture diagrams, you learn to make one call, check one boundary, and read one cost line.

Evals are the new unit tests

Ask a strong candidate "how do you know your LLM feature works?" and the interview hangs on the answer. "We tested it once and it looked fine" loses the round. The senior answer is an evaluation harness: a golden set of inputs with expected outputs, a rubric, and a threshold your CI blocks merges on.

def test_tool_call_when_ambiguous():
    reply = agent("Book a flight tomorrow if possible")
    assert reply.tool_call.name == "flights.search"
    assert not reply.hallucinated_policy

The framing that separates engineers from prompt-writers: an LLM feature is not a function you can unit test for determinism, so you test the distribution. Regression of refusal rate, drift of latency percentiles, accuracy on a golden set that grows every time a customer finds an embarrassing edge case. Candidates who describe their eval set as a living asset, with owner-verified labels and tracked thresholds, demonstrate exactly what a build wants. If you have never written one, build a ten-case golden set this week; it changes how you see every model call afterwards.

What stays stubbornly human

The new interview is not an invitation to ignore the old one. Three human skills keep their value, and interviewers probe for them harder now, because the model makes them scarcer. Debugging across a boundary: when the output is wrong, is it the prompt, the retrieval, the tool, or the model? Tracing a probabilistic failure to a deterministic cause is the hardest thing an AI engineer does. System degradation design: you cannot ship a checkout flow that depends on a model returning 200 — you design the fallback, the timeout, the human-in-the-loop path, and the observability that tells you which leg you are on. Judgement about when not to use AI: the candidate who can say "this feature is best handled by a hash table and a cron job" in an LLM-heavy interview is the candidate who looks senior; the candidate who wraps a neural net around every requirement looks like a cargo cult.

How to prepare, concretely

First, build the loop: one weekend, one small agent over a public API, with tool calling, a bound loop, and a three-line eval. That single artifact teaches more than ten YouTube walkthroughs. Second, rehearse the three follow-ups interviewers love: how do you prevent infinite tool loops; how do you keep the context window from filling; how do you measure regression when there is no deterministic answer. Third, rediscover the classic skills with an AI lens — strong algorithms and systems fundamentals still gate the interview, but now they are tested next to your ability to place them correctly, including when a model is the wrong tool for a job.

The question bank has added the AI-era prompts worth drilling, the interview strategy guide sequences how to present a build-it-from-scratch agent project you have actually shipped, and the deeper platform patterns live inside the courses so the architecture vocabulary arrives before the interviewer uses it.

The verdict

The engineering interview did not get harder because it got softer; it got newer. The core judgement — read the problem, name the shape, build the smallest thing that survives production — is untouched. What is new is the subject matter: how you thread a probabilistic component through a deterministic system without letting it break the guarantees. Candidates who treat the LLM round as one more engineering problem, solvable with loops, evals, and boundaries, will find it the most honest round on the list. Candidates who treat it as a prompt contest will find the whole interview as kind as a pager at 3 a.m. — which is to say, as kind as the systems you are asked to design.

Keep reading