Routing Different GitHub Events to Different Agent Configurations

Routing different GitHub events, pull requests, issues, push hooks, failed CI runs, to purpose-built agent configurations is a necessity for teams running agents in production. It's a correctness requirement, the same way type-checking or input validation is a correctness requirement, not a style preference. Most teams start with the opposite assumption: one webhook handler, one agent config, every event shoved through the same pipe as though a pull request and a push were interchangeable inputs. That assumption breaks fast, because a PR review event and a CI failure event carry different payloads, need different context windows, justify different compute budgets, and call for entirely different outputs.
The architecture makes the stakes concrete. In the standard webhook routing model, every POST triggers an agent run by default: the payload becomes a prompt, the agent chews on it, tokens get spent. That cost lands whether the event actually needed reasoning or just needed someone (something) to send a notification. The inefficiency costs money, but the bigger problem is behavioral. A code review agent set loose on a push event doesn't just waste compute, it produces the wrong output entirely, and an issue-triage agent pointed at a CI failure has no relevant context to work with in the first place. That's fast enough to matter. That's broken.
"Purpose-built configuration" Purpose-built configuration means something specific here, a different label slapped on the same runner would not capture it. It means different system prompts, different tool permissions, different budgets, and different places the output actually lands. Mapping event type to configuration is its own engineering discipline, with patterns that can be learned, reused, and audited, and the rest of this piece exists to prove that out.
The full spectrum of GitHub events that can trigger agent runs
GitHub exposes a wide set of event types that can trigger an agent run, and each one hands the agent a different shape of information. Pull request events, opened, updated, closed, assigned for review, carry a diff, a list of changed files, and PR metadata. Push and commit events carry commit SHAs, the ref, and the file paths touched, but no review context. Check and workflow events, the completed-or-failed CI signal, carry job logs, check status, and the commit tied to that run. Issue events carry only the issue body and its metadata, and no code. Dependency events flag manifest changes, obsolete packages, vulnerable versions. And scheduled or background triggers need no human action whatsoever, they fire on a cron or a manual dispatch and the agent works entirely on its own clock.
Whether an agent handles an event well or badly depends on what actually varies across this list. Some events carry code, some carry only text. Some have a human sitting there waiting on a synchronous response; a scheduled job has nobody watching. And the cost of an agent getting it wrong swings wildly, from "mildly annoying label" to "merged the wrong fix into a branch."
Given that spread, the sane move is to configure only the events a given workflow actually needs, rather than subscribing an agent to the entire firehose of repository activity. That's the first enforcement point for routing, before any YAML file or model choice even enters the picture. It also reshapes the permission surface, because a review-only agent needs read access to source and the ability to comment on a PR, an agent acting on push events needs write access to branches, and an agent triaging CI failures needs read access to Actions. Understanding what each event actually carries in its payload is the precondition for everything that follows: you can't design the agent's behavior until you know what it's working with.
Requirements each event type imposes on an agent, and the divergence behind them
Pull request review is the heaviest lift of any event type. The agent needs the full diff, file history, the project's coding standards, and a security checklist, more context than any other trigger demands. The output has to land as inline PR comments, check run results, and a summary review, structured enough for a human to scan quickly and fast enough that latency actually matters, because someone is often sitting there waiting on the result. That's also the one place where spending more tokens per event is defensible on its face: a missed security defect costs vastly more than the extra reasoning tokens it would have taken to catch it. GitHub's own Lifecycle team leaned into this by building an agentic pipeline where agents classify issues, dig through the codebase, draft implementation plans, and open draft pull requests across more than 50 services, with the review stage carrying the most reasoning load in that whole chain. Faros AI's analysis of more than 10,000 developers found that review time still jumps 91% even in high-AI-adoption teams.
Issue events sit at the opposite end. There's no diff to load, so pulling the whole repository into context for a labeling task is pure waste. What the agent needs instead is the issue body, the label taxonomy, and whatever's already in the backlog, and what it produces is mostly structured metadata, labels applied, an assignee suggested, a related issue linked. None of that calls for a heavyweight reasoning model, a lighter and faster one does the job and costs a fraction as much. But the same event type can flip into a completely different job. The moment an issue gets assigned to the agent rather than simply filed, the agent now has to branch, write code, and open a draft PR, which is a different configuration entirely, triggered by a different subtype of the same event: "issue assigned" behaves nothing like "issue opened," even though both arrive through the same webhook.
CI failure events land in the middle. The agent needs the failing job's logs, the commit that broke things, and the diff between the last green run and this one, narrower than a full PR review but it still demands real log parsing and cause-and-effect reasoning. Success here is measured by whether the output is actionable: a proposed fix, a draft PR, or a triage comment that actually explains what broke and why. GitHub has been pushing in exactly this direction, expanding "Fix with Copilot" across failing Actions runs so it connects an error back to the change that caused it and drafts a correction.
Push events are where the discipline pays off the most, because they fire constantly and mostly don't need deep reasoning at all, just a lightweight check, a triggered downstream workflow, maybe a changelog update. Route every push through a reasoning-heavy config and token spend spirals fast, since this is by far the highest-frequency event class in most repos. The escape hatch here is a setting like deliver_only: true, which skips the agent loop entirely and just sends a plain notification, exactly the right fit for a push event that needs nothing more.
Scheduled and background events are the widest-ranging category of all, a dependency audit looks nothing like a migration sweep, so these need the most explicit, most deliberate scoping. They're also the safest place to fan work out across many parallel subagents, since Anthropic's introduction of dynamic parallel subagent orchestration in Claude Code (May 2026) lets a job split across tens or even hundreds of subagents at once, and doing that on a schedule, with nobody waiting and no deploy hanging in the balance, is a much safer bet than doing it mid-PR. Rakuten put roughly this shape of task to work at real scale: Claude Code autonomously implemented a complex activation vector extraction method within the vLLM codebase, 12.5 million lines of code, finishing in 7 hours at 99.9% numerical accuracy DEV Community. That's a scheduled, high-budget, isolated run, and it would be reckless to trigger something with that footprint off a single incoming pull request.
How declarative configuration files encode the routing decisions
The routing logic described above has to live somewhere other than a team's collective memory, and GitHub's native tooling gives it a home. The copilot-setup-steps.yml file inside.github/workflows/ defines the environment an agent runs in, the runs-on key sets which runner it lands on, and that job executes before the agent's main work even starts. For PR review specifically, a separate copilot-code-review.yml file can override that default: if it's absent, review falls back to copilot-setup-steps.yml, and if it's present, it takes priority. That's the first native mechanism for actually routing different event types to different configurations, and it cuts both ways: not writing a copilot-code-review.yml file is itself a routing decision, one that quietly folds review behavior into general-agent behavior. Fine for a small repo, a real liability at scale.
A second layer, AGENTS.md, carries hierarchical per-repo instructions that guide codebase navigation, testing standards, review criteria, and engineering practices. This is where event-specific behavior gets written down in plain language. Teams managing this across many repositories have started reaching for tools like Caliber, a CLI that fingerprints a project, generates and syncs configs including AGENTS.md, and scores how good those configs actually are.
Governance sits above both of those layers. Enterprise accounts can centrally define plugin rules, permission modes, and model selection through a managed settings file, version-controlled the same way any other infrastructure config would be. Org admins can also set a default runner across every repository without touching each one individually, and lock that setting so no individual repo can quietly override it, which is what keeps a routing decision from unraveling the first time a contributor changes teams.
The frontier of this pattern is already visible. AgentSPEX, out of UIUC, is a declarative YAML spec and execution language for LLM-agent workflows, with typed steps, branching, loops, and explicit state management, aimed at teams whose event orchestration has outgrown a handful of workflow files. And on the cost side, an open-source project called LLM-Cost-Autopilot keeps model IDs and prices in a config/models.yaml file, complete with a pricing_updated_at field per model, and runs a separate weekly workflow to train, evaluate, and publish reports on its own routing decisions. Different corner of the problem, same underlying idea: routing logic checked into the repo, reviewed like any other code change, and legible in git history rather than living in someone's head.
Sandbox and runner design: matching isolation level to event risk
Isolation is the physical expression of the same argument made above about permissions and budgets: risk isn't uniform across event types, so the sandbox shouldn't be uniform either DEV Community. A scheduled migration touching a 12.5-million-line codebase carries a blast radius that an issue-labeling run simply doesn't, and treating both the same way causes a governance failure DEV Community.
GitHub's cloud sandbox, an ephemeral, GitHub-hosted environment built on Azure Container Apps, gives full machine isolation, sessions that can be snapshotted, and effectively zero local attack surface, billed by usage. That's the right tool for high-risk, high-context work: large-scale migrations, security audits, anything where a mistake could ripple wide. That's the everyday baseline, the right fit for lower-risk work like issue triage or routine review, not the exception case.
Self-hosted runners, configured through copilot-setup-steps.yml, earn their place when an agent needs to reach internal resources a cloud runner simply can't touch: private package registries, internal APIs, that sort of thing. Each task still runs in an ephemeral workspace, so nothing leaks from one run into the next. For teams running at real scale, Actions Runner Controller scale sets can sit behind the same runs-on: target in that same config file, deploy ARC, stand up a scale set, point the config at it, and network ACLs, scanning, and logging all apply automatically at the runner level, with no per-event tuning required.
The failure mode that ephemeral infrastructure is specifically built to close off is state leakage. If sandboxing isn't enforced and different event types happen to land on the same runner back to back, a PR review session could, in principle, inherit environment state left behind by an earlier CI triage run. Ephemeral runners don't mitigate that risk, they remove the entire class of bug outright. Local sandboxing is included with standard GitHub Copilot seats, runs on macOS, Linux, and Windows with a consistent experience built on Microsoft MXC (Microsoft eXecution Container), and is appropriate as the everyday baseline for lower-risk events like issue triage or routine code review.
Budget and token spend: why event type should determine cost ceiling
Push events make the cost argument almost too easy: they're the highest-frequency event in most repositories, and pointing all of them at a reasoning-heavy model with no budget cap turns ordinary day-to-day development into runaway LLM spend. The deliver_only: true setting mentioned earlier isn't only a behavior control, it's a cost control in its own right, and for a push event that only needs a notification, it cuts token spend to zero. Auto model selection carries a built-in 10% discount with configurable reasoning levels, so encoding the model choice into each event's own config is a direct lever on the bill, rather than using a universal default.
Budget caps belong at the event level as well as the org level. A per-session cap keeps one runaway scheduled task from eating the month's entire budget, and a per-event-type cap keeps a high-frequency class like push events from crowding out the runs that actually matter. Faros AI's analysis of over 10,000 developers found that high-AI-adoption teams complete 21% more tasks and merge 98% more pull requests, yet review time still jumps 91%, with no significant company-wide improvement in throughput or quality. Read together, that's not an argument for spending less, it's an argument for spending on purpose. The point of per-event cost control isn't frugality, it's steering the budget toward the 21% gain and away from the 91% overhead it would otherwise get buried under.
Observability: logging tool calls and diffs per event type so routing decisions stay auditable
None of the routing logic above means anything if nobody can see it working. A team that routes push events to deliver-only and PR events to a full reasoning model has made a real engineering decision, and like any real engineering decision, it needs a paper trail: which event triggered which config, which tools the agent called, what diff or output it produced, and what that run cost. Logged per event type, that record turns a routing scheme from a guess into something a team can actually defend, tune, and hand off. Without it, the routing table is just a belief about how the system behaves, not a fact anyone can check.


