Weekly DevOps Roundup: Agent guardrails, runners, and tokens
This week's DevOps roundup centers on turning agent-driven automation into something you can safely run in production, with reference architectures that emphasize shared inventory and identity, policy-enforced tool access, sandboxed execution, and measurable telemetry for safety and cost. On the CI/CD side, GitHub Actions changes will force practical maintenance work, including macOS 14 runner retirement brownouts, stricter minimum runner versions, and broader retention rules that now apply to checks and run metadata. Security and supply-chain updates focus on identity and trust boundaries, from stateless GitHub App tokens and TLS compatibility deadlines to more automatable vulnerability reporting and advisory workflows. We close with platform tooling updates that aim for predictable automation, including pay-as-you-go GitHub-hosted agents in Azure Pipelines, improved azd configuration layering, and an open source Azure Policy Linter built for CI.
This Week's Overview
- Governing and containing AI agents in production
- GitHub Actions and CI runner operational changes
- macOS 14 runner image retirement (deadline and brownouts)
- Self-hosted runner minimum version enforcement (now in effect)
- Retention now applies to checks, runs, and statuses (not just logs/artifacts)
- Actions Runner Controller 0.15.0 for Kubernetes runner scale sets
- Code coverage uploads skip (instead of fail) on new branches without PRs
- Dependabot can be pinned to specific runner settings per repository
- GitHub security reporting and advisory workflows get more structured (and more automatable)
- Identity, tokens, and supply-chain hardening in GitHub and Azure DevOps
- Copilot and agent-driven developer workflow updates (GitHub and IDEs)
- DevOps platform tooling: Azure Pipelines, azd, and policy linting
- Other DevOps News
Governing and containing AI agents in production
Teams are converging on a similar shape for “agents in the enterprise”: central inventory and identity, policy-driven runtime controls, and end-to-end telemetry for both safety and cost. This week had multiple concrete reference architectures that treat agents like any other production workload, with clear boundaries around what tools they can call, how they authenticate, and how you prove what happened after the fact, extending last week's focus on agent-first delivery into the platform guardrails you need once agents start touching real environments.
On the Microsoft side, a federated model targets multitenant realities by using Agent 365 as a centralized inventory across Microsoft Entra tenants, paired with Entra Agent ID for agent identity. Enforcement happens at runtime through an Azure API Management AI gateway (with Open Policy Agent (OPA) for policy decisions), while Microsoft Purview provides governance hooks and Foundry/Copilot Studio cover build and operations. If you have agents spanning business units, the takeaway is that “tenant sprawl” does not need to become “policy sprawl” if inventory, identity, and enforcement are treated as shared platform services.
For action-taking agents, the security story focused on execution boundaries rather than prompt rules. An AKS-based reference design separates the Model Context Protocol (MCP) tool interface from sandbox lifecycle management (OpenSandbox) and runtime isolation (Kata Containers via AKS pod sandboxing), and it compares that approach with Azure Container Apps Sandboxes and Dynamic Sessions. Practically, this pushes you to design tool servers and sandboxes as independent components so you can rotate credentials (via a credential vault), change isolation technology, and still keep the agent-tool contract stable.
The operational side is catching up too: one comparison walks through three proof-of-concepts for agent evaluation and observability, from self-hosted Langfuse on an Azure VM to a Foundry-hosted agent with Azure-native tracing/evaluation, to an Azure Container Apps agent instrumented with OpenTelemetry into Application Insights and Log Analytics. The common thread is that agent “correctness” and “safety” are increasingly treated as measurable signals (traces, evaluations, costs), not human gut checks, which matters when agent automations run on schedules or react to production incidents.
- Build and govern AI agents across a multitenant organization
- When AI Starts Taking Action: Building Execution Boundaries with OpenSandbox and AKS
- Comparing Three Approaches to AI Agent Evaluation and Observability
GitHub Actions and CI runner operational changes
Runner lifecycle and retention policies got several updates that will force near-term maintenance, especially for macOS pipelines and self-hosted fleets. At the same time, GitHub is tightening how long workflow metadata sticks around and improving the ergonomics of operating runners at scale (including Kubernetes-based runner scale sets).
macOS 14 runner image retirement (deadline and brownouts)
GitHub Actions will retire the macOS 14 runner image on November 2, 2026, and the announcement includes brownout windows where macOS 14 jobs will fail. The immediate action item is to inventory workflows using the affected labels and migrate to supported macOS arm64 runner labels called out in the post.
If you have conditional matrices or reusable workflows that pin runs-on labels, check those templates first since they often fan out across many repos. Treat the brownout windows as a test harness: move a subset of jobs early, validate toolchains (Xcode, SDKs, signing), then complete the rollout before the cutoff.
Self-hosted runner minimum version enforcement (now in effect)
GitHub adjusted the enforcement timeline for minimum self-hosted runner versions on GitHub Enterprise Cloud, with full enforcement starting September 29, 2026. If your runners fall behind, expect jobs to fail once they cross the enforced minimum, so you need an upgrade process that keeps pace with deprecations rather than reacting to outages.
The post points to documentation and a REST API that exposes runner deprecation dates, which is useful for building a simple “runner fleet compliance” check in your internal monitoring. For large estates, this is the kind of change that benefits from treating runners like pets-and-cattle: immutable images and automated rollouts beat manual upgrades.
Retention now applies to checks, runs, and statuses (not just logs/artifacts)
Actions retention settings now also govern checks, workflow runs, and statuses, aligning their cleanup behavior with artifacts and logs. GitHub is keeping existing enterprise and org caps, and public repositories still have a 90-day maximum, but the key point is that more of your CI/CD metadata lifecycle is now controlled by retention rather than lingering indefinitely.
If you rely on old check runs for audit trails, debugging regressions, or compliance evidence, revisit retention settings and your downstream archiving strategy. This is a good trigger to export the signals you truly need (for example, SARIF findings, test summaries, deployment provenance) into a system designed for long-lived retention, especially as agents increasingly generate and act on CI signals rather than humans scanning logs.
Actions Runner Controller 0.15.0 for Kubernetes runner scale sets
Actions Runner Controller 0.15.0 adds several knobs aimed at operating runner scale sets more predictably on Kubernetes: in-place patch upgrades, configurable shutdown and rate limits, and lower reconciliation/API payload overhead. These are the kinds of changes that reduce “controller thrash” and help keep GitHub API usage and cluster churn under control as scale increases.
If you run bursty workloads or auto-scale runners based on queue depth, the shutdown and rate limit controls are worth reviewing because they can reduce job disruption during scaling events. Observability improvements also matter here because they help you tie together GitHub queue behavior, controller actions, and cluster resource pressure.
Code coverage uploads skip (instead of fail) on new branches without PRs
GitHub Code Quality updated the upload-code-coverage action so pushes to new branches without an open pull request skip the upload and report the reason instead of failing CI. Default-branch uploads and supported PR event uploads stay the same, so this mainly removes friction for experimentation branches and short-lived feature spikes.
If your workflow currently treats missing coverage as a hard failure, this change shifts the failure mode from “pipeline broken” to “signal absent” on those specific triggers. You may want to pair it with a PR-only coverage gate (or a merge queue requirement) if coverage is part of your merge criteria.
Dependabot can be pinned to specific runner settings per repository
GitHub added repository-level settings controlling which Actions runners Dependabot uses for version and security updates, including runner type plus optional custom labels and runner groups. This gives platform teams a cleaner way to route Dependabot workloads away from constrained runners or into hardened runner pools.
If you have private network dependencies, special tooling, or compliance requirements for dependency update jobs, this setting can remove the need for per-repo workflow hacks. It also pairs well with runner group policies so you can keep dependency automation inside the same trust boundaries as your main CI.
GitHub security reporting and advisory workflows get more structured (and more automatable)
GitHub is tightening the private vulnerability reporting and repository security advisory loop with changes that make reports more consistent, reduce abuse, and add richer API surfaces for integrations. For maintainers and AppSec teams, the theme is “less manual triage, more predictable data, and better in-timeline collaboration without leaking sensitive investigation notes.”
Private vulnerability reports: structured forms and rate limits
Private vulnerability reporting now supports structured forms with required fields, including a minimum-length proof of concept, optional CWE enforcement, and SECURITY.md guidance. Custom forms can be enforced for REST API submissions via a new endpoint that returns the enforced form, which helps tooling (and reporters) avoid submitting the wrong shape of data.
GitHub also introduced daily rate limits on new private vulnerability reports to reduce bulk and automated submissions, with per-repo customization and allow-listing for trusted reporters in GitHub Advanced Security settings. If you have a high-profile repo that attracts spammy reports, this gives you a throttle and a trust mechanism without turning off private reporting entirely.
Repository security advisories: confidential comments and a comments REST API preview
Repository security advisories now support confidential comments that are only visible to users with write access, which makes it easier to keep sensitive investigation notes inside the advisory timeline. The feature is audited, and you cannot toggle confidentiality after posting, so teams should decide up front when a note is meant to be private.
In parallel, GitHub shipped a public preview REST API for reading, adding, and editing advisory comments (including advisories created from private vulnerability reports) and added comment count fields in advisory responses. Together, these changes make it easier to build internal triage bots, synchronize notes into ticketing systems, or automate reminders while still keeping the most sensitive discussion in GitHub.
- Confidential comments on repository security advisories
- Repository security advisory comments API in public preview
SecurityAdvisory GraphQL API gets richer fields and filters
GitHub expanded the SecurityAdvisory GraphQL object with new fields (including CVE and NVD/GitHub review timestamps) and added severities and isWithdrawn filters on the securityAdvisories query. The practical win is fewer REST fallbacks when building dashboards or syncing advisory data into internal vulnerability management systems.
If you track remediation SLAs, the new timestamps can help you compute lead times and publication latency more accurately. Filtering withdrawn advisories and targeting specific severity bands also reduces post-processing and makes it cheaper to run frequent sync jobs.
Identity, tokens, and supply-chain hardening in GitHub and Azure DevOps
Several changes and incident learnings this week landed on the same message: DevOps compromises often start with identity and end with pipeline trust. Updates span GitHub App auth, TLS compatibility, and npm publishing hygiene, alongside a real-world incident write-up showing how attackers chain identity access into build systems and then into production credentials.
GitHub App installation tokens go stateless by default
GitHub completed the rollout of stateless GitHub App installation tokens, now issued by default in the ghs_APPID_JWT format. These tokens are significantly longer than the legacy format, so integrations need to validate storage limits (databases, secrets managers), proxy/header limits, and log redaction patterns that assume a shorter token.
GitHub also noted that the X-GitHub-Stateless-S2S-Token header will be deprecated on November 30, 2026. If you built custom middleware around that header, plan to migrate now so the deprecation does not become an outage later in the year.
X25519-only TLS support ends for GHE.com (data residency) on Oct 7
GitHub Enterprise Cloud with data residency will stop accepting TLS clients configured for X25519-only key agreement on October 7, 2026. GitHub will continue supporting FIPS-approved P-256 and P-384, so the fix is usually a client configuration change rather than a platform migration, but it needs to happen before the cutoff to avoid HTTPS connection failures.
This can affect older or tightly locked-down TLS stacks in enterprise proxies, Java runtimes, or custom clients that pinned curves. The safest approach is to test connectivity from any non-browser automation (CI runners, scanners, integration services) that talks to your data-residency GHE.com endpoint.
npm trusted publishing can (optionally) manage dist-tags via OIDC
npm trusted publishing now supports opt-in dist-tag management via short-lived OIDC credentials. This removes a common reason teams kept long-lived npm access tokens around after adopting OIDC-based publishing: updating dist-tags like latest, next, or release channels post-publish.
If your release pipeline tags are handled by a separate job (or a manual “promote” step), this change lets you keep that flow while still eliminating persistent secrets. Review your npm workflow permissions and only enable dist-tag permissions where you actually need them.
Incident analysis: identity compromise to pipeline trust to Kubernetes credential theft
Microsoft Defender Experts (DART) detailed how Storm-3068 used a self-service password reset compromise to gain persistent identity access, pivoted into Azure DevOps, and abused trusted pipelines to harvest Kubernetes credentials (kubeconfig). The value for DevOps teams is the concrete chain: identity weakness → DevOps access → pipeline trust abuse → production credential exposure, which is exactly the path many organizations under-model.
The hardening steps called out in the post are a good checklist: tighten identity recovery flows, reduce standing privileges, scrutinize which pipelines are “trusted” and what they can access, and treat cluster credentials as high-risk secrets that should be short-lived and scoped. If you already have pipeline OIDC federation and workload identity available, this is another nudge to replace static kubeconfigs with federated access wherever possible.
Copilot and agent-driven developer workflow updates (GitHub and IDEs)
This week connected product updates with the day-to-day reality of agent workflows: long-running agent sessions, multi-folder work, cleaning up agent-created worktrees, and moving changes from implementation all the way through merge. The practical shift is that “agent mode” is being treated as a full lifecycle feature, not a chat box, and that creates new places to standardize how teams work, picking up directly from last week's emphasis on Copilot spanning IDE, terminal, and PR workflows.
Copilot in VS Code: September 2026 releases and VS Code 1.140/1.141 agent workflow touches
GitHub summarized Copilot updates for VS Code v1.136-v1.140, including Agents window improvements, scheduled automations, and “agent merge” to help carry changes from implementation through pull request merge. The same update called out better session navigation and cleanup, plus improved workspace, Dev Container, and GitHub-context chat flows, which matters when your agent sessions need repository context and repeatability.
A separate VS Code 1.140 video highlighted Copilot-specific features like HydraFusion and improvements to multi-folder sessions and session composer controls. On the maintenance side, VS Code 1.141 (Insiders) adds a Copilot Chat command to review and remove inactive agent worktrees, which should help keep repos from accumulating stale worktrees after iterative agent runs.
- GitHub Copilot in VS Code, September 2026 releases
- Visual Studio Code and GitHub Copilot - What's new in 1.140
- Visual Studio Code 1.141 (Insiders)
Visual Studio: BYOM for agent mode, NuGet audit fixes, and a Git agent with MCP support
The September Visual Studio update introduced Bring Your Own Model (BYOM) for Copilot agent mode, which is a meaningful knob for teams balancing capability, compliance, and cost. It also added Copilot-assisted fixes for NuGet Audit warnings directly from the Error List, tightening the loop between security signals and developer action.
On the workflow side, Visual Studio added a Git agent for pull request exploration with MCP server support, and it improved C# debugging for compound if conditions plus Podman attach-to-process support. For teams using MCP to connect tools and context providers, the Git agent note is a reminder to treat MCP endpoints like production dependencies (versioning, auth, monitoring), not ad-hoc local scripts.
“Work after the agent finishes”: making delegated changes reviewable and auditable
A GitHub AES (Agentic Engineering System) write-up argued that even when an agent produces a green PR, teams can still accumulate “handoff debt” if reviewers cannot quickly see evidence, intent, and release impact. It proposes three checks to keep agent work safe to accept: evidence packets (what the agent observed and verified), enforced stopping points (where humans must review), and release receipts (what shipped and why).
This is less about tooling and more about process, but it fits the reality of agent merge and scheduled automations. If you are turning on “agent continues after PR” features, pairing them with explicit evidence and stopping points is how you keep velocity without losing accountability, echoing last week's reminder that repo readiness and enforcement matter as much as model quality.
DevOps platform tooling: Azure Pipelines, azd, and policy linting
Cost controls, repeatability, and guardrails showed up as practical tooling improvements rather than new concepts. The common thread is that teams want more predictable automation: predictable runner costs, predictable project configuration layering, and predictable policy quality before deployment.
Azure Pipelines: GitHub-hosted agents with pay-as-you-go pricing are GA
Azure Pipelines announced GA for GitHub-hosted agents with pay-as-you-go pricing, including GA macOS SKUs and preview Linux/Windows SKUs. The post goes deep on what VM images are available, how to enable and reference the agent pool in YAML, and how to track minutes and costs using Azure Cost Management.
If you have intermittent workloads (release-only pipelines, mobile builds, or “burst” demand) and do not want to maintain self-hosted capacity, this gives you a first-class costed alternative with clearer accounting. It also gives teams a way to separate “critical pipelines” from “best-effort pipelines” by selecting different SKUs and controlling spend rather than fighting for shared runner capacity.
Azure Developer CLI (azd): config layering, extension contracts, and concurrency controls
The September 2026 azd release recap included dependency-aware extension uninstall, versioned gRPC extension contracts, new azure.yaml project layering, per-phase concurrency limits, and local-socket support for external authentication hosts (via AZD_AUTH_ENDPOINT). For platform teams, azure.yaml layering and concurrency controls are the highlights because they make it easier to standardize templates while still letting app teams override safely, reinforcing last week's guided Azure app build flow that leaned on az and azd to keep agent-driven provisioning repeatable.
If you maintain internal azd templates, versioned extension contracts reduce the risk that a CLI update silently breaks your tooling. Local-socket auth host support is also useful for locked-down dev environments where network-based auth endpoints are undesirable.
Azure Policy Linter is open source (and CI-friendly)
Azure Policy Linter is now open source, installable as a .NET global tool, and it can run against one or many Azure Policy definition files. It outputs JSON and supports rule sets to catch common authoring issues early, which makes it straightforward to wire into pull request validation.
If your organization uses Azure Policy at scale, linting shifts policy quality left in the same way unit tests shift code quality left. The real benefit is consistency: fewer broken deployments due to policy syntax mistakes and fewer “works in one repo but not another” rule variations.
Other DevOps News
GitHub shipped several workflow and governance updates that do not fit neatly into the themes above but are still worth tracking, especially if you build internal developer platforms on top of GitHub APIs and organization-wide standards. This also included a few “quality of life” security hygiene options that reduce accidental leaks and make repositories easier to navigate for contributors.
- How to protect your repo in 2 minutes with gh-secure
- GitHub async merge API generally available
- New dashboard experience now the default
- Bring business context with external custom properties
- Accessibility statements highlighted on repository overview
- GitHub Advanced Security trials for GitHub Team
- Meet the Hosted Skills Canvas: Build, Run, and Debug in GitHub Copilot
- Enabling Consistent AI-Assisted Engineering with GitHub Copilot Plugins
- Episode 2: Reduce restarts with Hotpatch for Windows Server 2025 | The Azure Arc Check-In
- Bring Your Azure DevOps Repos to Azure SRE Agent Plugins
- Updates to Copilot Code Reviews for Azure Repos
- Highlights from Git 2.56
- How we found 24 Android vulnerabilities using our open source AI security agent
- How GitHub's Tiny Wins team tackles the AI slop problem | S02E04 | The GitHub Podcast
- Get started with the GitHub Copilot app: a free, hands-on course
- How GitHub Copilot app fixes CI failures automatically
- GitHub Universe Day 1 Keynote
- GitHub Universe Day 2 Keynote
- 10 technical talks I'm excited about at GitHub Universe 2026
- Lakeflow in Azure Databricks
- One Command Opens Claude Code, Another Opens Codex: How Azure Databricks Governs Coding Agents
- A New Model Costs 60 Percent More by Default: Isolating the Financial Risk of Testing Frontier AI
- Microsoft and IBM bring enterprise-scale workload orchestration to Azure with IBM Spectrum Symphony
- VS Code Live: Release Recap