Weekly GitHub Copilot Roundup - Multi-model, agents, and governance
This week's Weekly GitHub Copilot Roundup is about Copilot growing up operationally: more models (GPT-6.1 Sol and Claude Sonnet 5.5), clearer routing with HydraFusion patterns, and more reasons to treat model choice as a governed dependency. Agents kept expanding beyond the IDE with dynamic workflows, scheduled automations, and a new “computer use” preview that brings desktop apps into scope (and into security reviews). Code review automation moved forward with API triggers and a new default effort level, while MCP-based extensibility matured with hosted skills, shared canvases, and managed automation patterns that emphasize evidence, auditability, and controlled execution.
This Week's Overview
- Multi-model Copilot gets real: GPT-6.1 Sol, Claude Sonnet 5.5, and HydraFusion routing
- Agents and automation across Copilot, VS Code, and the Copilot app
- Code review automation and governance: APIs, defaults, and Azure Repos parity
- Copilot extensibility and MCP: hosted skills, shared canvases, and managed automation
- Model changes and cost controls: deprecations, routing, and billing visibility
- Execution boundaries and safety for action-taking agents
- Repo security and AI-assisted security research
- Other GitHub Copilot News
Multi-model Copilot gets real: GPT-6.1 Sol, Claude Sonnet 5.5, and HydraFusion routing
GitHub Copilot expanded its model lineup again, with Claude Sonnet 5.5 now generally available and OpenAI's GPT-6.1 Sol rolling out as a generally available option across Copilot surfaces. Both announcements emphasize admin control through Copilot model policies and usage-based billing details, which means teams should validate what is enabled for their org before telling developers to “just pick it” in a model picker.
At the same time, Copilot is leaning harder into multi-model orchestration with HydraFusion, which moved from last week's “here's how it works” storyline into a research preview you can actually enable in VS Code and the GitHub Copilot app. HydraFusion introduces explicit execution patterns (Single, Cascade, Critique) so the system can route prompts and optionally run “critique” passes, making model choice and workflow structure more transparent for developers who care about repeatability.
If you are standardizing Copilot across a company, this week's practical work is governance: decide which models you allow, how usage-based billing affects budgets, and when to enable HydraFusion for specific roles (for example, critique flows for code review and cascade flows for multi-step refactors). Expect some short-term drift in team behavior as people experiment with model selection and HydraFusion patterns, so it helps to set defaults and publish guidance early.
- Claude Sonnet 5.5 in GitHub Copilot
- GPT-6.1 Sol in GitHub Copilot
- HydraFusion in VS Code and the GitHub Copilot app
- Inside Project HydraFusion: multi-model orchestration in the GitHub Copilot CLI
- Visual Studio Code and GitHub Copilot - What's new in 1.140
Agents and automation across Copilot, VS Code, and the Copilot app
This week reinforced a clear theme: Copilot is not just “chat in the IDE” anymore, its surfaces are converging around agents, orchestration, and end-to-end workflow automation. The biggest changes land in three buckets: repeatable workflows, agents that can take actions on your machine, and IDE features that help manage the mess agents can create.
Dynamic workflows in Copilot CLI, the Copilot app, and the Copilot SDK
Dynamic workflows landed as a new way to define repeatable, multi-agent orchestration in code, spanning Copilot CLI, the GitHub Copilot app, and the GitHub Copilot SDK. The key idea is structure: parallel steps, checkpoints, and explicit handoffs, which helps teams encode “how we do X here” (like dependency upgrades or release prep) instead of relying on ad-hoc prompting.
For developers building internal automation, the interesting part is that workflows are meant to be packaged and reused, not just run once. That shifts the work from prompt craft to maintainable workflow definitions, so you can review changes, version them, and roll them out across a team the same way you manage other developer tooling.
“Computer use” preview: Copilot interacting with desktop apps
Following last week's push to make agent sessions portable across the Copilot app, VS Code, and the CLI, GitHub shipped a public preview of “computer use” for Copilot CLI and the GitHub Copilot app on macOS and Windows, allowing Copilot to click, type, and navigate desktop workflows with user approval and permission controls. On macOS, that includes system-level permissions like Accessibility and Screen Recording, which should immediately trigger a security and IT review for managed devices.
For teams, the practical implication is that agent capability is moving outside the repo and IDE boundary into “whatever is on screen.” If you plan to enable this, define guardrails up front (what apps are in scope, how credentials are handled, and what evidence is captured), then treat it like any other automation that can modify state on a developer machine.
VS Code agent workflow updates: scheduled automations, agent merge, and workspace hygiene
Building on last week's focus on connected sessions and Agent Merge as the bridge from agent output to merged PRs, the September Copilot updates in VS Code (v1.136-v1.140) focused on making agent-driven work easier to run and easier to finish. Highlights include Agents window improvements, scheduled automations, and “agent merge” to help push changes through implementation into pull request merge, plus better session navigation and cleanup and improved flows for workspace, Dev Container, and GitHub-context chat.
VS Code Insiders 1.141 adds a small but real quality-of-life feature: a Copilot Chat command to review and remove inactive agent worktrees. If your team is experimenting with agentic development, worktrees can pile up quickly, so having first-party cleanup helps keep local repos manageable without inventing your own scripts.
Copilot app learning path and “agent keeps working after PR” pattern
Microsoft published a free, open source “GitHub Copilot app for Beginners” course that teaches how to run coding agents in the desktop app and validate their work via diffs, tests, previews, and PR checks, echoing last week's theme that agent work needs repeatable validation loops, not just chat transcripts. It is a sign that the Copilot app workflow (agents + worktrees + validation loops) is becoming standardized enough to teach as a repeatable process.
GitHub also highlighted a concrete “keep working after PR” loop: the Copilot app can fix CI failures and address reviewer comments automatically using agent merge. If you adopt this pattern, make sure your repo policies still force humans to own final approval, and ensure the agent is gated by tests and checks rather than trusting the narrative in chat.
- Get started with the GitHub Copilot app: a free, hands-on course
- How GitHub Copilot app fixes CI failures automatically
Code review automation and governance: APIs, defaults, and Azure Repos parity
Copilot code review is becoming more automatable and more configurable, which matters if you want consistent review coverage without forcing everyone into the same UI flow. The GitHub side added API triggers and changed defaults, while Azure Repos gained better effort controls and auditability.
GitHub Copilot code review: API triggers and default “Balanced” effort
This follows last week's GA improvements to Copilot code review UX (tracking findings across pushes and clearer resolution behavior) by adding API triggers and a more opinionated default: GitHub Copilot code review can now be triggered via GitHub's REST and GraphQL APIs, which opens the door to wiring reviews into internal bots, scripts, or CI pipelines. GitHub also changed the default review effort level to Balanced, while still allowing configuration at enterprise, org, repo, and personal scopes (so teams can standardize defaults but allow exceptions).
If you already have custom PR automation (for example, labeling, CODEOWNERS routing, or policy checks), this API support lets you treat Copilot review as another step in your PR workflow rather than a manual button click. The “Balanced” default also matters operationally because it changes cost and latency expectations compared to a lighter review mode, so you may want to explicitly set effort in tooling to avoid surprises.
Azure Repos: effort levels, smoother suggestion application, and audit log events
Azure Repos Copilot Code Review improvements include configurable effort levels and a smoother path to resolving review comments when applying suggestions. Azure DevOps also added audit log events for tracking enablement and configuration changes, which is important for regulated teams that need to prove who changed what and when.
If you run mixed GitHub and Azure DevOps estates, this narrows the gap in how you govern AI review behavior across platforms. The new audit events are especially useful if you centralize monitoring or compliance reporting, since “Copilot was enabled” becomes a trackable, queryable event rather than tribal knowledge.
Copilot extensibility and MCP: hosted skills, shared canvases, and managed automation
Tooling around Model Context Protocol (MCP) and Copilot extensibility kept expanding, with new ways to build skills, distribute them across teams, and connect managed automation backends. The common thread is moving from one-off local tools to governed, inspectable integrations that teams can share.
Hosted Skills Canvas in the Copilot app (Azure Functions + MCP)
Building on last week's MCP governance thread (secure hosting patterns, spec updates, and usage metrics visibility), the Azure Functions Hosted Skills canvas in the GitHub Copilot app adds a local, event-driven workflow for building, running, and debugging hosted skills, with triggers, logs, and evidence-backed outputs. It also covers deploying with Azure Developer CLI (azd) and exposing the hosted skill as an MCP tool endpoint, including identity options like Managed Identity.
For developers, this makes “turn a useful script into a reusable agent tool” feel closer to normal app development: you can run locally, inspect logs, then deploy and wire it into Copilot via MCP. It also nudges teams toward evidence-driven agent outputs (logs, artifacts, traces) instead of trusting the agent's summary.
Azure Canvases: plugin-based Copilot workspace for Azure tasks
Azure Canvases for GitHub Copilot was announced as a plugin-based workspace pairing Copilot chat with interactive dashboards and guided workflows for Azure tasks. Examples include deploying Functions, querying resources across subscriptions, and reviewing cost and governance signals (such as an “Azure Cost Health Check”).
This is a step toward agent UX that is not only chat: teams get repeatable UI surfaces where they can see context, review signals, and take guided actions. If you are responsible for platform engineering, Azure Canvases hints at how you might package internal “golden path” workflows as something developers can run with guardrails rather than copying wiki steps.
Managed browser automation via MCP (Playwright Workspaces)
Microsoft shared a guide for connecting Playwright Workspaces (cloud-hosted browsers) to agent applications using MCP, with an example using the GitHub Copilot app. The workflow emphasizes human approval before submission, evidence capture, and diagnostics (logs, traces, screenshots, recordings) to make agent-driven browser actions debuggable.
If your Copilot integrations need to touch flaky web UIs (admin portals, legacy apps, vendor consoles), running automation in a managed browser environment plus MCP gives you a cleaner separation between agent intent and execution. The built-in evidence trail is useful when you need to explain what happened after an agent “clicked the wrong thing.”
Distributing consistent agent behavior with Copilot plugins
This extends last week's focus on reusable agent skills and measuring real customization usage by describing a distribution mechanism: Microsoft published practical guidance for using GitHub Copilot plugins to distribute versioned agent instructions, skills, and MCP tool connections across teams. It covers what to package, how to publish via an internal marketplace, and how to govern updates and precedence, which addresses a common pain point: every team ends up with slightly different prompts, tools, and expectations.
For orgs trying to standardize AI-assisted engineering, plugins are a governance mechanism as much as a technical feature. Versioning and rollout control matter when a plugin change can alter tool access or agent behavior across hundreds of developers.
Model changes and cost controls: deprecations, routing, and billing visibility
Between model deprecations inside Copilot and broader Microsoft updates around model routing and FinOps, this week had a strong “ops” angle. Developers will feel it as “why did my model disappear” and platform teams will feel it as “how do we keep spend and governance under control.”
Deprecations in Copilot model selection
This is the follow-through on last week's mid-October deprecation planning advice: GitHub deprecated several models across all GitHub Copilot experiences on October 2, 2026, and provided suggested replacement models. Copilot Enterprise admins may need to enable replacement models via model policies so they show up in Copilot Chat model selectors in VS Code and on github.com.
The practical takeaway is to treat model availability as a managed dependency, not a personal preference. If your onboarding docs mention a specific model, update them quickly, and if you rely on model-specific behavior (context length, tool calling style, cost), plan small evaluation runs when replacements roll out.
FinOps and platform updates: billing metrics, routing, and governance knobs
Sonia Cuff's September 2026 FinOps roundup called out pricing and cost-control updates across Azure, Azure AI Foundry/Azure OpenAI, and GitHub Copilot, including expanded usage and billing metrics. In parallel, John Savill's Microsoft AI update highlighted Azure AI Foundry model additions, Model Router changes, agent features, governance controls, and Copilot platform updates that connect directly to how teams standardize and govern agent behavior.
For teams adopting usage-based billing models in Copilot, better metrics change what you can manage: you can start correlating cost with org/repo usage patterns, then decide where to enforce model policies or effort levels. If you already route requests through model routers or proxies, these updates are a reminder to keep routing logic and budgets aligned with real developer workflows, not just a lab benchmark.
Execution boundaries and safety for action-taking agents
As Copilot gains more ways to take action (desktop control, browser automation, hosted skills), the question shifts to “where does code execute and what is it allowed to touch.” Two deep dives from kinfey focus on designing boundaries that make agent execution auditable and constrainable without killing productivity.
Sandboxing agent actions with OpenSandbox + AKS (and alternatives)
A reference architecture shows how to separate MCP tool interfaces from sandbox lifecycle management (OpenSandbox) and runtime isolation (Kata Containers), implemented on AKS with pod sandboxing. The post compares this approach with Azure Container Apps Sandboxes and Azure Container Apps Dynamic Sessions, and discusses related pieces like credential vaulting.
If you are moving from “agent suggests code” to “agent executes tasks,” this type of separation is the difference between a demo and a production system. You want tight control over credentials, network egress, and filesystem access, and you want a clear story for incident response when an agent does something unexpected.
Decision harnesses and typed gates with Jev + model routers
A second prototype compares a traditional Copilot tool-calling loop with a Jev-based “decision harness” that emits typed choices and confidence signals before execution, then uses Python gates to restrict tools and enforce workflow boundaries. The write-up includes concrete tool chains, token/latency observations, and cautions about benchmarking and safety claims.
The core idea is to make “should we do the dangerous thing” a first-class step with structured outputs, not a hidden part of a chat transcript. If you are building custom Copilot SDK tools or MCP servers, typed decisions plus explicit thresholds can be a practical pattern for reducing accidental destructive actions.
Repo security and AI-assisted security research
GitHub Security Lab continued pushing lightweight tooling for everyday repo hygiene alongside heavier agent-driven research workflows. The common thread is reducing the chance of human mistakes (like leaking secrets) while making security work more repeatable.
gh-secure: one-command repo hardening from GitHub Security Lab
GitHub highlighted gh-secure as a one-command way to improve repository security hygiene and reduce accidental secret leaks in public code. It can be used via the Copilot app, Copilot CLI, or GitHub CLI, which makes it easy to bake into “new repo” checklists or quick triage when a repo is about to be open-sourced.
If you maintain templates or internal scaffolding, consider adding gh-secure to your bootstrap scripts so baseline protections are not optional. Teams will get the most value when this is paired with education on secret handling (rotating leaked keys, using OIDC, and avoiding long-lived tokens).
Taskflow Agent case study: 24 Android vulnerabilities found with targeted LLM taskflows
GitHub Security Lab shared how its open-source Taskflow Agent used targeted LLM taskflows to audit Android apps, resulting in 24 disclosed vulnerabilities. The post includes a Codespaces-based runbook, two detailed vulnerability walkthroughs, and practical notes on false positives and severity estimation, plus details on Copilot licensing/tokens in the workflow.
For security teams, the key takeaway is that “agentic” security work still benefits from narrow, well-scoped taskflows rather than open-ended chatting. For developers, it is a reminder that AI-based findings need the same triage discipline as any static or dynamic analysis: reproduce, validate impact, and document evidence.
Other GitHub Copilot News
GitHub Universe content and adjacent IDE updates reinforced the direction of travel: Copilot features are increasingly about agents, memory, evaluation, and controlled context distribution, not just faster autocomplete, building on last week's arc of making agent behavior governable through policies, metrics, and repeatable workflows. Visual Studio's September update adds more knobs for enterprise teams (like Bring Your Own Model for Copilot agent mode) and tightens the workflow between IDE diagnostics and Copilot assistance.
- GitHub Universe Day 1 Keynote
- 10 technical talks I'm excited about at GitHub Universe 2026
- Visual Studio September Update - Power Your Workflow with Your Model
- Keep Your AI Accurate with Microsoft Learn MCP Server
- VS Code Live: Release Recap
- Voices from the Skies: Build Your Own Personal Assistant with GitHub Copilot App