“Loop engineering” showed up in my feed three times last week, each time with the same confident undertone: prompt engineering is over, context engineering had its moment, and now the real skill is engineering the loop. I’ve spent enough hours watching agents grind through unattended runs to believe there’s something real hiding under that phrase. I’ve also seen enough of what people are calling loop engineering to know most of it is retry-until-it-looks-done with a name attached. That’s not a discipline. That’s hope wrapped in a for-loop.
So before this term calcifies into another thing people say without meaning anything specific by it, here’s what I think it actually points at, and where the engineering has to happen for it to be real.
The loop underneath every agent
Strip away the branding and every coding agent, every research agent, every support agent that calls tools in sequence is running the same cycle: plan what to do next, act on it, observe what happened, decide whether to continue, stop, or hand back to a person. Four stages. A single prompt-response pair only ever does the first two, once, with a human doing the deciding afterward. A loop is what you get when the agent does the deciding itself, over and over, without you in between iterations.
That’s the part worth being precise about. I’ve written before about harness engineering — the orchestration, guardrails, and observability wrapped around a model in production. The loop is the piece of that wrapper actually making the turn-by-turn call: keep going, or stop. Loop engineering is the discipline of making that call trustworthy instead of vibes-based.
Where each stage actually breaks
Plan, without engineering, means the agent reacts to whatever it just observed with no memory of the bigger arc — it tries a fix, sees it didn’t work, tries a slightly different fix, and thrashes between two or three variations of the same wrong idea because nothing is tracking that it’s already been down this path. A planned loop keeps an explicit record of what’s been tried, so “try again” isn’t the default move when a smarter one is “this approach is exhausted, escalate.”
Act, without engineering, assumes every action is safe to repeat. It isn’t. If an iteration gets interrupted and retried, an action that isn’t idempotent — sending a message, appending to a file instead of overwriting it, calling an API that charges per call — does its side effect twice. This is the most purely “engineering” part of the whole loop, and the least glamorous: before you let a step run unattended and retry on failure, you have to know whether running it twice by accident produces a duplicate or a no-op.
Verify, without engineering, is the agent asking itself “does this look right?” and answering yes, because it wrote the thing it’s now grading. That’s not verification, it’s the model agreeing with itself. A real verify step is an external check — a test suite, a schema validator, a status code, something that would say no regardless of how the agent feels about its own output. The gap between these two is exactly why an agent can report success on every iteration of a ten-hour run and still hand you something broken at the end — nothing in the loop was ever capable of saying no.
Decide, without engineering, has exactly one exit: run until the budget or the turn limit is gone. A decide step worth the name has three real exits, not one — a stop condition stating what “done” means, a pause condition stating what makes it hand back to a person instead of continuing on its own judgment, and the default failure exit when neither fires and something outside the plan shows up.
The part that isn’t one of the four stages
There’s a fifth thing loop engineering has to handle that doesn’t fit neatly into plan-act-verify-decide: what the loop remembers about its own progress once the conversation gets long enough to be compacted or summarized. An agent running for hours doesn’t keep every early decision in its live context — most harnesses trim or summarize older turns to make room. If the record of what’s been tried, what’s in scope, and what’s already verified lives only in that context, it’s exactly the part that’s most likely to get quietly compressed away right when the loop needs it most. Keeping that state external — a scratch file, a running log, a checklist the loop reads back on every iteration instead of trusting its own memory of turn twelve — is what makes a long-running loop still coherent at hour six instead of slowly drifting off its own plan.
What I’d actually check before trusting one
Given a system claiming to do loop engineering, the questions that separate the real thing from the rebrand are concrete. Does verification run an external check, or does the model grade itself? Are the actions inside each iteration safe to repeat if that iteration gets retried? Is there a stop condition and a pause condition, stated as actual conditions rather than “it’ll know when it’s done”? And does the loop’s sense of its own progress survive a context compaction, or does it reset every time the window gets trimmed?
If the answer to any of those is “it just kind of works,” what’s running isn’t an engineered loop. It’s a chat session that happens to call itself in sequence, and it’ll behave like one — confidently, right up until the point where nobody was checking.