Agent in Production
IntegrationsLong read

Scheduling Recurring Agents for Maintenance and Migration Tasks

Agents need dedicated infrastructure designed for stateful, multi-step work, not just cron triggers.

Contributing Editor · · 9 min read
Cover illustration for “Scheduling Recurring Agents for Maintenance and Migration Tasks”
Integrations · October 2, 2026 · 9 min read · 2,118 words

An agent that fails partway through a task exposes this gap, and that is the whole argument this piece makes: that recurring agents need production infrastructure, not a trigger.

Why recurring agent tasks outgrow cron infrastructure

Cron has handled scheduled work reliably for decades, and the instinct to reach for it when an agent needs to run nightly or weekly is not a mistake born of ignorance. It's the correct generalization from everything that came before agents: batch jobs, backup scripts, report generators, all of them stateless, all of them either succeeding or failing in a way that a log line could capture. Pointing a cron trigger at an agent looks like the obvious next step from that history. The trouble starts because agent sessions aren't one-shot stateless requests the way a batch script is. Production traces from June 2026 show that agent execution forms a sequential chain, where each tool call's output feeds the next LLM call, and 87% of LLM invocations in these systems are agent-initiated rather than triggered directly by a user. That structure shifts the unit of work from an individual request to an entire session, and most serving infrastructure, built for independent, short-lived, stateless calls, is not built to carry a session through hours or days of dependent steps.

Cron handles time. It does not handle the four things a recurring agent actually needs: state that survives between runs, coordination across agents that depend on one another, structured capture of failures with retry logic, and an audit trail that can reconstruct what happened and why. The gap between what cron offers and what agents need is not cosmetic. A cron job that fails leaves a line in a log file. An agent that fails partway through a workflow can leave a compliance issue, a missed deadline, or a customer refund sitting downstream of that failure.

Consider a nightly data-quality agent auditing a warehouse. If it runs fresh every night with no memory of yesterday's baseline, it has no way to detect whether today's numbers represent an anomaly, because it has nothing to compare against and no mechanism to escalate when variance crosses a threshold. Running it on a bare schedule doesn't give it a cadence; it erases the information the task depends on. The same gap appears in coordination: a report agent finishes at 6 a.m., and an email agent is supposed to send that report out afterward, not alongside it. By mid-2026, conversations inside agent engineering communities had already moved past debating how autonomous agents should be and settled into debating scheduling patterns, state recovery, and cost governance instead, which is a sign that the industry recognizes this as an infrastructure problem rather than a scheduling inconvenience.

What a production loop looks like structurally

Production recurring agent work settles into a supervisor-worker loop rather than a single long-running process grinding through a task start to finish. A central orchestrator tracks goals, deadlines, and dependencies, then hands off bounded pieces of work, like code search, patch generation, testing, or security review, to specialized workers, each operating in its own context window. This delegation solves a real technical constraint: subagents keep the controller's own context window from filling up with the details of every subtask, since the controller offloads focused work to workers and only takes back a report when they finish.

Google Labs research offers a three-level taxonomy that helps place where most production systems actually sit today. By definition, Level 1 agents run only when someone prompts them directly. Level 2, labeled "Scheduled," run off a schedule or a predefined trigger, and may filter or batch their outputs according to a developer's interruption preferences, but they don't carry learning across contexts. Level 3, "Situation Aware," agents watch a continuous stream of events and update a model of the developer's preferences from ongoing feedback. Most production recurring tasks running today are Level 2, which matters because it sets honest expectations: these systems run on triggers and schedules, they don't yet adapt their own behavior from accumulated feedback the way a Level 3 system would.

Four loop types cover the space of recurring-task design. Heartbeat loops run continuously on a short interval and suit monitoring work: watching logs, checking service health, scanning for configuration drift. Cron loops run at specific times and suit batch work like daily code review, weekly dependency audits, or morning standup summaries; Claude Code ships native /loop and cron scheduling, and Codex ships an Automations tab that supports recurring schedules along with subagent spawning. Hook loops trigger off external events, a pull request, a failed CI run, a message in Slack, and run once per trigger rather than on a timer. Goal loops iterate until a defined success condition is met, then stop, which suits migrations or refactors where the scope isn't fully known at the outset, provided the loop has both a stop condition and a ceiling on how many iterations it's allowed to run.

Five components recur across effective loop implementations. Each iteration runs inside its own isolated git worktree, so a broken run damages a copy of the codebase rather than the branch everyone else depends on. Instructions live in reusable, version-controlled skill files that the loop references, rather than being inlined as prompt text baked into the schedule itself. Connectors built on the Model Context Protocol give the loop a single, consistent way to reach external tools, databases, issue trackers, deployment systems, instead of a different integration for each. Subagents let the controller decompose work and assign it appropriately: a security-reviewer subagent might run a strong model at high reasoning effort, while a file scanner runs a fast, cheap model, and each subagent carries its own context window and its own tool permissions. State tracking, whether through file-based checkpoints, git history, or an external store, keeps the loop from repeating work it already finished and lets it resume cleanly after a failure.

A migration task shows why the stop condition and the iteration ceiling aren't optional extras. An agent migrating files off an old API pattern needs to keep running until no files match that pattern anymore, and it needs a hard ceiling on how many iterations it's allowed regardless of whether the condition is met. Without both pieces in place, the loop has no exit, and a system with no exit is a system that eventually runs until something external stops it.

Session isolation as the foundation

Recurring agents run without a person watching each step, often with write access to codebases, databases, and outside services, and every one of those sessions needs to run in its own isolated environment. Isolation is what makes it safe to let the loop run unattended on a schedule.

Agents generate and run code as a normal part of doing their work, and that makes the untrusted-code problem inherent to the loop itself rather than an edge case to guard against. A recurring agent running every night is, in effect, running untrusted code at scale on a fixed schedule, repeatedly, without a human reviewing each execution before it happens.

Isolation at the infrastructure level means three concrete things. Each session gets its own filesystem and network namespace, with no access to the host system. Credentials get minted fresh for that session and revoked the moment it ends, instead of persisting as long-lived keys shared across every run. Tool access is scoped to what the task needs, not granted broadly to cover every possible future use the team can imagine.

The sandbox tooling that has emerged in 2026 reflects these requirements directly. Modal runs a large volume of concurrent sessions on a container stack tuned for fast startup, and Ramp uses Modal Sandboxes to run background coding agents that generate code changes and commit them or open pull requests with the result. Blaxel offers a persistent "agent computer" that stays on standby and resumes quickly, without charging for compute while it sits idle. NVIDIA introduced OpenShell at GTC 2026, wrapping agents like Claude Code and Codex in kernel-level isolation governed by declarative YAML security policies. None of these approaches treats isolation as an add-on. Each one builds it into the layer where the agent actually executes, which is the only layer where it can meaningfully apply.

Agent configuration in the repo, not in a UI

A recurring agent that runs on a schedule is a piece of infrastructure, no different in kind from a deployment pipeline or a database migration script, and infrastructure that lives only inside a dashboard is infrastructure nobody has reviewed, versioned, or can roll back when it misbehaves.

The pattern that addresses this treats agent definitions, the schedule, the prompt, the model choice, tool permissions, stop conditions, and iteration caps, as YAML files committed to the repository, subject to the same review process as any other code change. kagent's API illustrates what this looks like concretely: Kubernetes-native manifests using apiVersion: kagent.dev/v1alpha2 and kind: Agent combine GitHub and AWS Terraform tool servers into a single declarative agent definition that understands its own infrastructure requirements and can generate the code and pull requests it needs. Oracle introduced its Open Agent Specification in October 2025 as a framework-agnostic declarative language, defining building blocks for standalone agents, structured workflows, and multi-agent compositions in portable JSON/YAML targeting any compatible framework or runtime.

Putting agent configuration under version control changes what a team can actually do when something goes wrong. A pull request that changes the scope or model tier of a nightly dependency-audit agent goes through the same review a code change would, and reviewers can see precisely what the agent will do differently before it ships. Rolling back a bad configuration becomes a git revert rather than a support ticket routed to whoever happens to own the scheduling dashboard. The configuration file doubles as documentation: the prompt, the stop condition, the tool list, and the budget cap all sit in the repo next to the code the agent operates on, visible to anyone who needs to understand what the agent is authorized to do.

For teams moving a recurring agent out of experimentation and into something that touches shared infrastructure, the requirement is specific: its definition needs to live in version control, as a YAML file any engineer can read, propose changes to, and audit after an incident. A dashboard toggle cannot offer any of that.

Cost controls that enforce rather than alert

Diagram: Five Enforced Cost Controls for Recurring Agent Loops. Visualizes: Visualize a layered architecture of five cost-control mechanisms that must stack together to govern a recurring agent loop.

A recurring agent running without enforced budget limits is a bill waiting to be scheduled, because agent loops burn through tokens at a rate far beyond standard chat interactions, and that consumption compounds across iterations, subagents, and retries before anyone has a chance to notice.

Agent workflows consume substantially more tokens per completed task than a standard chat exchange does, and teams that price their agent infrastructure using chat-based assumptions tend to find this out the expensive way. The clearest illustration: four agents entered an infinite loop in November 2025, ran for over a week without anyone catching it, and generated a bill of $47,000. Tool failures make this worse in a specific, measurable way: production traces show that a failed tool call tends to trigger a retry loop that multiplies compute cost by roughly four times for the affected turns.

Alerting tells a team what already happened. Enforcement stops the agent before its next LLM call goes out. That distinction is the entire argument for building cost controls into the runtime rather than bolting dashboards on top of it. One mechanism for this is per-key budget enforcement: when a key hits its cap, an AI Gateway returns an HTTP 402 and rejects any further requests on that key, which turns a runaway loop into a rejected request instead of a surprise invoice arriving at the end of the month. A fuller architecture layers five separate controls: a ceiling on each individual request, a rolling budget tracked per session, a monthly cap applied per key, routing that sends work to cheaper models where a cheaper model suffices, and circuit breakers that halt a loop outright when something looks wrong.

Model-tier routing deserves particular attention because it cuts cost without touching output quality in any way that matters: a loop step that scans files for a pattern match doesn't need the same model as a subagent reviewing security-sensitive code, and assigning the cheaper model to the cheaper task reduces loop-wide spend without asking the system to do anything differently.

The common thread across isolation, version control, and cost enforcement is the same: recurring agents inherit the authority of the humans who deployed them, and the infrastructure built to run those agents has to make that authority accountable at every layer, not just convenient to configure on a calendar.

Sources

  1. Agentic Coding Needs Proactivity, Not Just Autonomy
  2. Why Agent Scheduling Is Your Next Infrastructure Problem - DEV Community
  3. Loop Engineering: How to Build AI Agent Loops That Run Themselves
  4. AI Agent Architecture 2026: Building Production-Grade Systems — Patterns, Benchmarks, and Lessons from 10,000-Agent Swarms - DEV Community
  5. Agentic Coding in the Wild: Characterizing GitHub Copilot at Production Scale
Filed underIntegrations

More in Integrations