Weekly DevOps Roundup: CI runtime shifts and governed AI agents
This week in DevOps, GitHub Actions made breaking-but-necessary runtime and API changes (Node.js 24 for JavaScript actions, artifact visibility updates, and new query count limits) that will affect pipelines, dashboards, and automation. GitHub also pushed collaboration and security forward with richer PR triage, review-stage metrics, SSH hardening, and proof-of-presence controls for sensitive enterprise actions. On the platform side, Azure sharpened the agent operations story with Foundry governance, Container Apps Sandboxes (microVM isolation with egress control and OTLP export), and AKS updates for isolated and confidential AI workloads. Across it all, the practical theme is treating agents, CI, and security controls like production systems: pin versions, instrument everything, and design for policy and scale.
This Week's Overview
- GitHub Actions and CI runtime changes you need to handle now
- GitHub Actions API and artifact visibility changes (less confusion, different assumptions)
- GitHub Copilot app: agents, observability, and scaling the PR surface
- GitHub collaboration and automation: issues, PRs, and throughput metrics
- GitHub security posture updates: SSH hardening and proof-of-presence controls
- Code scanning tooling: CodeQL 2.27.1 updates and a packaging deprecation
- Building and governing AI agents on Azure: Foundry, Container Apps Sandboxes, and real-world patterns
- Microsoft Foundry updates: model choice, voice agents, resilience, and optimization loops
- Azure Container Apps Sandboxes (GA): microVM isolation with egress controls and telemetry export
- Agent-first platform architecture: governance plane vs execution plane
- Foundry Hosted Agents + MCP in practice: deployment boundaries and debugging failure modes
- AKS for AI workloads: isolation, confidential GPUs, and burst capacity
- Azure resilience thinking: safer deployments, chaos testing, and AI dependency risk
- Infrastructure-as-code workflow improvements: auto-generated Bicep module docs
- Database and streaming operations: GoldenGate on Azure and Spark state repartitioning
- Other DevOps News
GitHub Actions and CI runtime changes you need to handle now
Following last week's push to tighten CI guardrails (for example, cache-mode GA to reduce cache poisoning risk), GitHub Actions runners no longer provide Node.js 20, and JavaScript actions now run on Node.js 24 by default. For action maintainers, that means updating your action.yml to runs.using: node24 (and validating dependencies, bundling, and runtime assumptions against Node 24) because the previous opt-out environment variable has been removed.
Workflow authors should expect breakage if they depend on older JavaScript actions that have not been rebuilt for Node 24, or if they run on self-hosted environments that are now out of scope (the changelog calls out dropped support for older macOS and ARM32 self-hosted runners). If you run regulated pipelines, this is a good moment to pin action versions, audit your marketplace dependencies, and add a canary workflow that exercises critical actions on every runner image you use.
GitHub Actions API and artifact visibility changes (less confusion, different assumptions)
Two related GitHub Actions changes landed this week: the UI/REST API now hides expired artifacts, and workflow run queries now cap visible counts at “2,500+” when matches exceed 2,500. Together, these updates reduce misleading results in the UI, but they can change automation that previously inferred state from artifact listings or relied on exact run counts.
Expired artifacts no longer appear in UI or REST API
Expired artifacts are no longer shown in the workflow run summary UI and are no longer returned by the artifacts REST API endpoints. If you have scripts that list artifacts to validate retention, calculate storage usage, or decide whether a job “produced outputs”, you will need to switch to recording artifact metadata at creation time (or emitting explicit job outputs) instead of expecting expired artifacts to remain visible.
This should also reduce confusion for teams trying to reconcile retention policies with billing and storage expectations. Practically, you can treat the artifacts API as “current artifacts only”, and document retention behavior in your pipeline runbooks so incident responders do not chase artifacts that have already expired.
Workflow run queries now report “2,500+” for large result sets
GitHub Actions workflow run queries in the REST API and UI will now report counts as “2,500+” when matches exceed 2,500. Paginated results still work (up to 1,000 items returned across pages), but you should stop treating the count as an exact number once you cross that threshold.
If you use counts for dashboards or alerting (for example, “number of failed runs in the last N days”), update your queries to compute counts from returned pages, or aggregate elsewhere (SIEM/observability, data warehouse exports, etc.). GitHub says the change reduces misleading counts caused by query timeouts and improves performance, so expect it to be more stable at scale but less precise at a glance.
GitHub Copilot app: agents, observability, and scaling the PR surface
Building on last week's Copilot thread around stricter sandbox governance and enterprise policy controls in IDEs, Copilot's DevOps story kept shifting from “IDE helper” toward “workflow participant” this week, with updates that make agents easier to run where your code lives and easier to monitor like any other service. The throughline is operational maturity: sandboxing, tracing, and ergonomics for large code review surfaces.
Running Copilot agents inside WSL (and using Git worktrees for parallel work)
GitHub published a practical walkthrough for connecting the GitHub Copilot app to WSL on Windows and running Copilot coding agents directly inside Ubuntu. The guide leans on Git worktrees to let you run parallel feature work safely (multiple working directories backed by the same repo) while keeping the agent's changes isolated and reviewable.
The workflow emphasizes in-app preview and diff verification, which matters if you are delegating multi-file refactors or repetitive changes to an agent. For teams standardizing Windows developer machines but shipping Linux workloads, this is a straightforward way to keep agent execution closer to your production-like toolchain without abandoning Windows as the primary desktop OS.
OpenTelemetry support for Copilot app sessions (enterprise-managed)
The GitHub Copilot app now supports OpenTelemetry configuration through enterprise-managed settings, letting organizations export agent activity traces to compatible observability backends. That is a meaningful step if you need to troubleshoot agent sessions the same way you debug CI jobs or production services (what tools were invoked, where time was spent, what failed, and under which policy).
If your org already standardizes on OTLP pipelines, this opens the door to correlating agent work with repo events (PRs/issues), CI runs, and even incident timelines. The practical next step is to decide what “good” looks like (latency, error rate, tool failure rate) and create lightweight dashboards so agent automation does not become a blind spot.
Weekly Copilot releases: model choice, local sandboxing, and agent/session polish
Following last week's weekly release emphasis on model routing and IDE policy surfaces, GitHub's Copilot weekly releases (September 21) added new model choices (Claude Opus 5.5, GPT-6 Sol/Luna, and Grok 4.7), plus local sandboxing and broader agent/session improvements across Slack/Teams, JetBrains, and VS Code. Model choice is becoming an operational control as much as a UX feature, since different models can have different latency/cost/quality profiles for the same workflow.
Local sandboxing complements the shift toward running agents against real repos and tooling, since it creates clearer execution boundaries when agents run tasks that touch code, credentials, or build tools. If you are piloting agentic workflows, treat model selection and sandboxing as part of your “environment configuration” (similar to choosing runner images or container bases), not just a per-user preference.
Rendering million-line pull requests in the Copilot app
GitHub shared a deep dive on how the Copilot app rebuilt its PR diff surface to handle extreme cases (million-line PRs with hundreds of inline comments). The work included split geometry for code vs dynamic blocks, viewport-scoped measurement, scroll anchoring, and CI-backed performance instrumentation to prevent regressions.
For DevOps teams that regularly review generated changes (dependency bumps, codegen output, mass refactors), this is not just UX trivia. It reduces the “tool falls over” risk when agents or automation produce large diffs, and it reinforces a pattern: agent tooling needs the same performance budgets and observability discipline as production UI.
GitHub collaboration and automation: issues, PRs, and throughput metrics
GitHub shipped several small-but-practical changes around the work tracking loop (issues → PRs → reviews → merge) and the reporting you need to manage it. The key theme is making collaboration state more structured and more automatable through APIs, while improving the UI for high-volume teams.
Issue automation: intents for metadata triage, plus new saved views and relationships
GitHub demonstrated “issue intents” for automating issue metadata triage, including confidence thresholds, reviewing the agent's reasoning, and applying routine updates with agentic workflows surfaced through GitHub Actions logs. If you already run automation for labels/projects/assignees, the important detail is the reviewable reasoning and the notion of thresholds, which gives you a control point for “auto-apply vs queue for human review”.
Separately, GitHub made private saved views for repository issues available, and the “Relates to” issue relationship is now generally available across REST/GraphQL APIs, webhooks, timeline events, and search (issues and projects). Together, these features improve how teams build personal triage workflows while still emitting structured relationship data that automation and reporting can consume.
- How to automate issue metadata with GitHub issue intents
- Private saved views for repository issues and “Relates to” issue relationship is generally available
PR UI and review analytics: faster triage, better stage breakdowns
Building on last week's PR list refresh preview, GitHub's refreshed repository pull requests page is now generally available, adding improved filtering, advanced boolean search (including nested queries), a collapsible sidebar, bulk PR actions, and richer PR context like status check counts and unread update indicators. If you operate a repo with many concurrent PRs, the combination of nested search and bulk actions is the practical win because it lets maintainers run “queue management” directly in the UI rather than bouncing between filters and per-PR pages.
On the reporting side, the usage metrics API expanded its Copilot usage metrics reporting by adding a pull_request_review_times array to repos-1-day rows. It breaks review time into three stages (ready-to-first-review, first-to-final-review, final-review-to-merge) and reports median and p90, which makes it easier to separate “reviewer availability” from “iteration churn” and “merge bureaucracy” when you are trying to improve throughput.
- Refreshed repository pull requests page generally available
- Usage metrics API adds pull request review stages
GitHub security posture updates: SSH hardening and proof-of-presence controls
GitHub's security updates this week focus on reducing risk from legacy cryptography and compromised sessions, extending last week's theme of “block unsafe actions by default” from repos and pipelines into authentication and admin workflows. The practical impact is that teams should expect some “this worked yesterday” failures if they have older SSH clients/keys or rely on long-lived authenticated browser sessions for sensitive actions.
SSH: deprecations, stronger defaults, and post-quantum hybrid key exchange
GitHub announced upcoming SSH security changes that deprecate SHA-1-based RSA signatures (ssh-rsa), remove diffie-hellman-group-exchange-sha256, and enforce a 3072-bit minimum RSA key size for new uploads. If you still use RSA keys, plan to move to RSA with SHA-2 signatures (or modern key types like Ed25519) and validate your fleet's OpenSSH versions and crypto policies ahead of enforcement.
GitHub is also enabling the post-quantum hybrid key exchange mlkem768x25519-sha256 on github.com and some GitHub Enterprise Cloud regions. This will mostly be transparent for updated clients, but it is another reason to inventory SSH client versions on build agents and developer workstations so you can catch negotiation failures before they block deploys.
Proof of presence (preview) for high-impact actions in Enterprise Cloud EMU
GitHub introduced a public preview “proof of presence” control for GitHub Enterprise Cloud EMU enterprises using Microsoft Entra ID SSO. It requires fresh re-authentication or MFA before high-impact actions, building on GitHub sudo mode to reduce risk from stolen sessions and long-lived tokens.
For enterprises with strict change control, this is a policy lever worth testing alongside privileged workflow protections and environment rules. Plan for user experience impacts (more prompts) and document which actions trigger proof-of-presence so on-call responders are not surprised during time-sensitive incidents.
Code scanning tooling: CodeQL 2.27.1 updates and a packaging deprecation
CodeQL 2.27.1 shipped query improvements across multiple ecosystems (including C/C++, C#, and GitHub Actions) and added Kotlin 2.4.20 support, plus extractor and data-flow model accuracy improvements. If you maintain custom queries or rely on precise taint/data-flow behavior, expect some findings to shift as modeling improves, so treat this as a “baseline may change” update and review diffs in alert volume.
Building on last week's CodeQL 2.27.0 note (including the new Linux ARM64 support), GitHub also deprecated the all-platform CodeQL bundle starting with CodeQL CLI 2.27.0, with removal planned for mid-March 2027. If you have automation that downloads a single bundle for heterogeneous runners (including Linux ARM64), you will need to switch to platform-specific downloads and update caching and artifact strategies accordingly.
- CodeQL 2.27.1 adds C and C++ query and Kotlin 2.4.20 support
- Deprecation notice: All-platform CodeQL bundle
Building and governing AI agents on Azure: Foundry, Container Apps Sandboxes, and real-world patterns
Building on last week's DevOps focus on agent governance (Foundry updates, cost attribution, and MCP tooling), Azure's agent story got more concrete this week, with product updates and architecture guidance that separate “governance” (identity, evaluation, tracing, policy) from “execution” (isolated runtimes, egress control, deployment). If you are responsible for platform guardrails, these posts offer a clear direction: treat agents like production workloads with dedicated governance planes, not like IDE features.
Microsoft Foundry updates: model choice, voice agents, resilience, and optimization loops
Microsoft Foundry announced updates spanning expanded frontier model availability, voice agents, and long-running hosted agent resilience. It also added observability, evaluation, and optimization capabilities aimed at continuously improving agent quality, latency, and cost, with stronger governance controls (including items like toolboxes and network egress controls, plus integration points such as Azure API Management's AI Gateway).
For DevOps teams, the operational hook is the move toward repeatable measurement and policy for agents: you can evaluate versions, run controlled rollouts, and enforce constraints the same way you manage services. That is especially relevant when agents can call tools that create tickets, modify infra, or access production diagnostics.
Azure Container Apps Sandboxes (GA): microVM isolation with egress controls and telemetry export
Azure Container Apps Sandboxes are now generally available, providing hardware-isolated microVM sandboxes designed for running untrusted code and agent workloads. The GA announcement highlights per-sandbox egress controls, private networking, snapshots, volumes, and opt-in telemetry export to Azure monitoring destinations via OpenTelemetry (OTLP), with IaC support called out for Bicep and Terraform (ACA provider).
This is a practical runtime option when you need stronger isolation than standard containers, but still want a managed deployment surface for short-lived tasks (agent runs, tool execution, untrusted transforms). If your security model depends on network control, the per-sandbox egress policy support is the detail to evaluate, since agent tools often need outbound access that you may want to tightly constrain.
Agent-first platform architecture: governance plane vs execution plane
Microsoft published an “agent-first” platform pattern that explicitly splits agent governance (Microsoft Foundry with Entra Agent ID, tracing, evaluation) from execution (Azure Container Apps Sandboxes running per-task isolated microVM environments). The architecture is motivated by scaling autonomous agents without collapsing security and compliance controls, especially when tasks vary widely in trust level and required permissions.
If you are designing internal agent platforms, the actionable takeaway is to standardize identity and policy at the governance layer, then allow multiple execution backends as long as they meet isolation and telemetry requirements. This makes it easier to evolve runtime choices over time (containers, microVM sandboxes, Kubernetes jobs) without rewriting governance.
Foundry Hosted Agents + MCP in practice: deployment boundaries and debugging failure modes
Building on last week's MCP “how to build a server” tutorial, a detailed guide walked through building “Fibey Field Ops” using Microsoft Foundry Hosted Agents, Model Context Protocol (MCP), Azure Container Apps, and Azure AI Search, with identity handled through Microsoft Entra ID and Managed Identity. The value here is the engineering boundary-setting: tool schemas vs skills, hosted integration contracts, staged deployment with azd, and practical debugging of identity, streaming, and retrieval failures (including SSE and OpenTelemetry).
If you are struggling to operationalize agent prototypes, this kind of “where it breaks” guidance is what turns a demo into a service. Use it as a checklist for runbooks: verify identity paths, validate retrieval quality, and instrument streaming paths so you can spot timeouts and partial failures.
AKS for AI workloads: isolation, confidential GPUs, and burst capacity
Azure Kubernetes Service updates this month focused on running AI workloads with stronger isolation, better identity and storage security, and more flexible scaling. This matters for DevOps teams because AI workloads often combine sensitive data paths, expensive accelerators, and spiky batch-style execution.
AKS product innovations: Kata sandboxing, confidential GPU, GitOps, and scaling options
Microsoft highlighted new AKS capabilities aimed at AI workloads, including pod sandboxing with Kata Containers, confidential GPU support (AMD SEV-SNP + NVIDIA H100), and security improvements for identity and storage (including workload identity and Azure Files CSI updates). The announcement also calls out GitOps via an Argo CD extension for AKS, burst scaling with Virtual Nodes v2, and new efficiency/hyperscale options like AKS Hyperscale Configuration and Flex Nodes.
If you operate mixed-trust clusters, Kata-based pod sandboxing is a key control point for isolating workloads without fully splitting clusters. For accelerator-heavy deployments, confidential GPU support and tighter identity defaults are worth evaluating early, because they can influence node pool design, scheduling constraints, and admission policy.
Virtual Nodes on ACI: serverless burst compute plus confidential containers
A separate post introduced next-generation AKS virtual nodes backed by Azure Container Instances (ACI), positioning ACI as a serverless compute layer for bursty Kubernetes workloads while keeping standard kubectl/Helm workflows. It also shows how to enable confidential containers using confcom-generated Confidential Container Environment (CCE) policies (Rego) and hardware-backed attestation with AMD SEV-SNP.
The practical DevOps angle is capacity management: you can keep steady-state workloads on regular node pools while bursting certain jobs onto ACI without changing your deployment tooling. If you run regulated workloads, the confidential container policy and attestation flow is the part to test, since it impacts how you package images, define policies, and validate runtime integrity.
Topology-aware GPU infrastructure via Dapple + Azure (AKS as control plane)
Microsoft also described collaborating with Dapple to offer topology-aware, purpose-built GPU infrastructure through an Azure-native operating model. AKS acts as the operational control plane, and a Dapple Operator bridges Kubernetes intent and the underlying infrastructure topology (including high-performance networking like InfiniBand).
If you have ever hit performance cliffs due to GPU placement or interconnect topology, this is a signal that “GPU scheduling” is moving beyond simple node labels and into topology-aware provisioning. Expect GitOps patterns to matter more here, because infrastructure intent (topology constraints) needs to be expressed and reconciled consistently.
Azure resilience thinking: safer deployments, chaos testing, and AI dependency risk
Azure's resilience messaging this week doubled down on a point SRE teams already live with: an architecture diagram is a plan, not evidence. The posts tie resilience to policy, continuous validation, and explicit handling of AI/agent non-determinism as a new reliability variable.
Mark Russinovich revisited a 2014 bug that nearly caused a worldwide Azure outage and connected it to the creation of a safe deployment policy as a core resilience practice. In parallel, Azure published guidance arguing that diagrams capture intent but do not prove resilience, pointing to health models, infrastructure as code, and chaos testing as the continuous validation loop as workloads and AI dependencies evolve. A longer discussion with Russinovich reinforced that AI agents introduce non-determinism that changes reliability assumptions and requires guardrails and clearer accountability models.
- The 2014 Bug That Almost Broke Azure
- Your architecture diagram is not your resilience
- Resilience at Cloud Scale: Azure CTO on Outages, Hardware, and AI
Infrastructure-as-code workflow improvements: auto-generated Bicep module docs
Bicep gained an experimental documentation generator: bicep docs generate. It generates module documentation directly from compiled Bicep modules and supports customization via Scriban templates, which helps teams standardize what “module docs” should include (inputs/outputs, examples, notes) without hand-editing markdown.
The post also shows how to run it across a repo with patterns and how to enforce up-to-date docs in CI, turning documentation drift into a failing check instead of a recurring review comment. If you maintain internal module registries or Azure Verified Modules-like catalogs, this is a useful building block for making IaC changes more reviewable and less tribal.
Database and streaming operations: GoldenGate on Azure and Spark state repartitioning
Oracle and Databricks both published operator-focused updates that reduce friction in ongoing operations rather than day-one setup. The theme is “make change safer after you are in production”: better observability for replication and more flexibility for streaming state.
Oracle AI Database@Azure expanded with an Azure-native management experience for Oracle GoldenGate, including Azure portal/API lifecycle operations and generally available observability through Azure Monitor metrics and Azure Log Analytics events. The post also calls out GoldenGate connectivity options, including Azure Event Hubs, Kafka options, and Microsoft Fabric destinations, which helps when you are stitching replication into event-driven architectures.
On the streaming side, Databricks Runtime 18 for Spark Structured Streaming supports on-demand state repartitioning, letting you change spark.sql.streaming.stateStore.partitions without abandoning checkpoints. The guide explains how to validate impact using StreamingQueryProgress metrics, which is a practical way to tune state store performance (including RocksDB-backed stores) as load and key distribution change over time.
- Oracle AI Database@Azure Expands with Azure Native Oracle GoldenGate User Experience, Observability
- The Partition Count You Picked on Day 1 Doesn't Have to Follow You Forever
Other DevOps News
Building on last week's emphasis on inventorying and governing developer platform risk (from secret scanning rules to threat mapping), GitHub added enterprise-grade inventory support by enabling credential inventory exports (keys and tokens) from enterprise settings or via a paginated REST API, with metadata and filtering designed for incident response workflows. The key operational move is correlating inventory data with audit log activity so you can prioritize rotation and investigation based on actual usage, not just presence.
Several posts focused on agentic workflow adoption in day-to-day engineering: guidance on moving from code completion to multi-step agent workflows, a survey-backed argument for better measurement to reduce wasted compute (including an agentic workflow that opens draft PRs with benchmark evidence), and VS Code agent extensibility patterns (tools, permissions, sandboxing, MCP, and packaging via Agent Plugins). If you are planning to expand agent use, these are useful inputs for policy (permissions and approvals), implementation (MCP/tooling), and change management (measurement and review).
Finally, Microsoft published practical SRE guidance for configuring Azure SRE Agent (data connections, response plans, approvals, permissions), and a separate architecture post described Azure KARS as a reference stack for running coding-agent CLIs as governed Kubernetes workloads with identity brokering, L7 egress policy, budgets, content safety, and tamper-evident audit logs. These complement the broader Foundry and sandbox announcements by focusing on the on-call realities: routing, approvals, tool access, and auditability.
- GitHub Enterprise adds credential inventory exports
- How to move AI from code completion to agentic workflows
- Developers want more efficient software. Here’s what over 1000 GitHub users told us they need.
- VS Code Learn: Extending Agents
- Getting the best out of Azure SRE Agent
- When the Coding Agent Leaves the Codebase: Building a Multi-Runtime AI Agent Infra with Azure KARS