In June 2026 the term loop engineering started circulating — the idea that the real engineering happens not inside the agent's plan-act-observe cycle, but in the outer system that repeatedly invokes it: trigger, frame, run, verify, record, stop. A recent mining study (arXiv 2608.21884) scanned 36,710 repositories and found only 217 with genuine outer loops — and most of those committed the loop configuration but no state and no verifier depth.
I read that with some amusement, because my setup had been running this way for months before the term existed. One operator, thirteen scheduled agent jobs, a watchdog on a five-minute tick, and a tamper-evident ledger recording what the loops decided. This post describes the machinery — including where it failed this week, and the one boundary I still cannot enforce.
The loop, concretely
The trigger layer is mostly time-based — a watchdog every five minutes, a nightly maintenance cycle, a scout every six hours — with one exception that turned out to matter: an ingress guard that fires only when a message queue exceeds twenty entries. Condition-triggered rather than schedule-triggered. In hindsight that one job is closer to the event-driven model the literature recommends than anything else I run.
Each job is framed by a fixed protocol, runs under a token budget, and is checked by something that is not the agent itself. That last part is the load-bearing piece: 287 deterministic consistency checks run at every session start, a digest job re-verifies every reported error against its actual artifact, and a final review stage once caught a real bug in a commit I had already made — the verifier earning its keep in the most literal way.
Stop conditions, generalized
Until recently only the nightly cycle had hard stop conditions (four of them: lock file, running benchmark, stack health, live users). This week I generalized the pattern across all thirteen jobs. The watchdog now flags, per job: a run exceeding max(90 minutes, twice its usual duration), three or more consecutive failures, an execution overdue by more than 1.5× its own interval — a weekly job is not "late" after two days — and jobs that never ran at all. Alerts go out over WhatsApp with a cooldown, and a recovery message follows when the job heals.
A war story from the same week shows why this matters: a weekly deep-research job issued four parallel web searches that simply never returned. No error, no timeout — silence. The stall detector aborted the run after 1,172 seconds. Post-mortem: the upstream search backend was degraded that afternoon, and the tool call had no internal timeout, so one stuck call held the entire run. The fix I control is boring and effective — cap the search volume, run smaller batches — and the fix I don't control is a per-tool timeout that still doesn't exist.
A budget that actually stops
The deeper gap was that monitoring existed but enforcement didn't. The alarm had been reporting roughly 5.57 million tokens in 24 hours against a 2.5 million cap — detection without teeth. The fix is a shared lock file: when the rolling window exceeds the cap, a guard writes the lock; every internal tool call and the guest-chat API then refuse until the window is compliant again. No manual unlock — the lock clears itself when usage drops back under the cap.
Live verification, same day: lock active at 5,566,561 tokens. Tools refuse. Guests get a holding note instead of silence.
And the honest boundary: the bulk of that consumption comes from application-owned cron jobs whose scheduler sits behind a separate gateway I cannot programmatically pause — a token-drift defect blocks the control API. So the lock stops my stack's consumption and degrades guest answers gracefully, while the upstream jobs continue until I intervene by hand. Enforcement reaches exactly as far as my infrastructure boundary and not one step further. I wrote it down because partial enforcement honestly described is more useful than a claim of full control.
Where this sits
Against the study's taxonomy: triggered runs (yes), machine-checkable stop conditions (now uniform), persistent state files that are actually committed (yes), verifier stages independent of the producing agent (yes, and battle-tested), token budgets (now enforced, not just measured), defined human escalation points (an autonomy charter with explicit triggers). The two known weak spots are queue-based discovery — my loops still wait for the clock rather than finding work — and the upstream enforcement boundary described above.
I am not claiming precedence on any of this. The interesting datum is convergence: a one-person setup built out of operational necessity arrived at the same architecture the field is now naming. That suggests the pattern is real and not fashion.
Why publish this
I'm an independent researcher with a small working system, not a lab. If you run agentic loops and have solved the parts I haven't — upstream budget enforcement across process boundaries, per-tool-call timeouts under unreliable backends, queue-driven work discovery — I'd like to compare notes. And if you see a stop condition I'm missing, that email is the one I most want to receive.
The watchdog and guard tooling sit in a private repo. Happy to share the approach in detail with anyone who asks.