Weekly AI Roundup: Copilot workflows, governance, and data foundations
This week's AI roundup tracks a clear shift from chat-based assistance to action-taking systems, with GitHub Copilot adding dynamic workflows, desktop “computer use”, multi-model orchestration (HydraFusion), and more automation across code review and PR merge. In parallel, Microsoft and partners focused on the controls teams need when agents can execute real changes, including identity and policy across Entra tenants, runtime enforcement via gateways, and isolation with sandboxes and microVMs. On the data side, Fabric IQ and OneLake updates reinforced that agent-ready context depends on governed semantics and permissions, while Azure SQL Hyperscale pushed vector search deeper into mainstream databases. We close with practical guidance on measuring cost and quality (routing eval toolkits, FinOps controls) and the security work needed to keep agent-driven development reviewable and accountable.
This Week's Overview
- GitHub Copilot expands from “chat” to orchestrated, action-taking workflows
- Dynamic workflows for Copilot CLI, the Copilot app, and the Copilot SDK
- “Computer use” brings desktop UI automation (with permissions and approval)
- HydraFusion brings multi-model orchestration to Copilot Chat surfaces
- Copilot in VS Code: more agent operations, automations, and PR merge support
- Copilot code review becomes automatable via API (and defaults to Balanced)
- Copilot models: new options arrive while older picks are removed
- Governing AI agents: identity, policy, observability, and execution boundaries
- Data foundations for agentic AI: Fabric IQ, OneLake governance, and vector search at scale
- Cost, routing, and evaluation: making model choice measurable
- Building agents and agentic apps: MCP tools, hosted skills, and agent-centric UI
- AI security and trustworthy operations: threats, accountability, and repo hygiene
- AI infrastructure and deployment patterns: private GPU extension and serverless GPU benchmarks
- Other Artificial Intelligence News
GitHub Copilot expands from “chat” to orchestrated, action-taking workflows
GitHub spent this week making it clearer that Copilot is becoming a workflow engine, not just an autocomplete tool (building on last week's push toward connected sessions across the Copilot app, IDE, terminal, and PRs). New previews focus on multi-agent coordination (dynamic workflows), multi-model routing (HydraFusion), and the ability for Copilot to take actions in your local desktop apps (computer use) with explicit user approval and OS-level permissions.
For developers, the practical shift is that Copilot is moving closer to “delegate a task end-to-end” patterns: create a plan, run steps in parallel, checkpoint, hand off between agents, and then carry changes through PR merge and follow-up fixes. The flip side is that teams now need stronger governance around model selection, tool permissions, evidence collection, and “what happens after the PR is green” so agent work stays reviewable.
Dynamic workflows for Copilot CLI, the Copilot app, and the Copilot SDK
Dynamic workflows (preview) let you define repeatable orchestration in code, including parallel steps, checkpoints, and structured handoffs between agents (a direct extension of last week's theme that agent work needs explicit handoffs, repo readiness, and repeatable patterns rather than better prompts). This is aimed at teams that want something more deterministic than “prompt, then hope the agent follows the same plan next time,” and it aligns with Copilot Extensions and the Copilot SDK so workflows can live alongside your internal tooling.
If you already script around Copilot CLI, the key idea is turning those scripts into first-class workflow definitions that can be reused and versioned. Expect this to pair naturally with internal policy controls (which models are allowed, which tools can run, and where human approval is required).
“Computer use” brings desktop UI automation (with permissions and approval)
GitHub Copilot “computer use” is now in public preview for Copilot CLI and the GitHub Copilot app on macOS and Windows. It lets Copilot click, type, and navigate desktop application workflows, but it is gated by user approval and platform permission controls (including macOS Accessibility permissions and screen recording permissions).
This moves Copilot closer to “operator” style automation for dev workflows that still live outside your terminal and IDE (for example, updating settings in GUI tools, copying data between apps, or stepping through onboarding flows), and it sharpens last week's least-privilege conversation around MCP/tool access into a broader “user can do anything” risk when the tool is the desktop. It also raises the bar for least-privilege design, since UI automation can easily become “do anything my user can do” if you do not scope it carefully.
HydraFusion brings multi-model orchestration to Copilot Chat surfaces
HydraFusion (research preview) is now available beyond Copilot CLI, rolling into VS Code and the GitHub Copilot app. It supports multiple execution patterns such as Single (one model), Cascade (try a smaller/cheaper model then escalate), and Critique (a second model reviews or critiques the first output), with an emphasis on transparency in why a model was selected.
The developer implication is that “which model should I use?” starts to become a runtime decision based on intent, cost, and quality trade-offs, instead of a static dropdown choice, which follows last week's rollout of Auto model selection tiers and budget escalation as the admin side of the same multi-model story. If your org uses model policies, you will want to confirm the allowed model set still supports your HydraFusion flows and that usage-based billing expectations are clear.
- HydraFusion in VS Code and the GitHub Copilot app
- Inside Project HydraFusion: multi-model orchestration in the GitHub Copilot CLI
- Visual Studio Code and GitHub Copilot - What's new in 1.140
Copilot in VS Code: more agent operations, automations, and PR merge support
The September 2026 Copilot updates for VS Code (v1.136-v1.140) leaned into day-to-day agent management: Agents window improvements, better session navigation and cleanup, and new flows for working in Dev Containers and with GitHub context in chat, continuing last week's VS Code agent-first thread around making sessions portable and reproducible (including Dev Containers). GitHub also called out “scheduled automations” and “agent merge” to help move changes from implementation through pull request merge, and even continue after opening a PR (for example fixing CI failures or addressing review feedback).
If you are piloting agent-driven development, these are the kinds of features that reduce friction when an agent creates multiple worktrees, leaves behind sessions, or needs to keep working after CI feedback arrives. Teams should treat these as workflow changes, not just UI polish, and update conventions around branching, evidence (tests, logs), and when human review is mandatory.
- GitHub Copilot in VS Code, September 2026 releases
- How GitHub Copilot app fixes CI failures automatically
- Visual Studio Code 1.141 (Insiders)
- VS Code Live: Release Recap
Copilot code review becomes automatable via API (and defaults to Balanced)
Copilot code review can now be triggered via the GitHub REST API and GraphQL API, which makes it easier to wire into internal tools, scripts, or custom PR automation (a natural follow-on to last week's code review iteration tracking improvements that made review output easier to manage at scale). GitHub also changed the default review effort level to Balanced, while keeping configuration options at enterprise, org, repo, and personal scopes.
The API addition matters most when you want consistent enforcement (for example, “run AI review on every PR opened against main”) or when you want to collect review metadata for tracking. If you previously tuned effort levels to control noise and latency, double-check your defaults after this change so teams do not get surprised by a different review depth.
Copilot models: new options arrive while older picks are removed
GitHub continues to evolve the Copilot model lineup, with GPT-6.1 Sol rolling out as generally available and Claude Sonnet 5.5 now generally available as well. In the same week, GitHub deprecated several “selected models” across Copilot experiences (effective October 2, 2026) and pointed to replacement models, which may require Copilot Enterprise admins to update model policies so the alternatives show up in model selectors in VS Code and on github.com (tightening the loop from last week's warning about mid-October deprecations into an immediate “update policies now” change).
For developers, this is now a recurring operational task: treat model availability like a dependency that can change, and ensure your org has a process for updating allowed models, communicating changes, and rerunning key evals when defaults shift. If you use usage-based billing, align model choices with cost expectations and make sure teams can see usage at the right scope.
- GPT-6.1 Sol in GitHub Copilot
- Claude Sonnet 5.5 in GitHub Copilot
- Selected models in GitHub Copilot deprecated
Governing AI agents: identity, policy, observability, and execution boundaries
A consistent theme across Microsoft and GitHub content this week was that agents need more than prompts and tools - they need identity, inventory, runtime controls, and audit trails. The most concrete guidance focused on multitenant governance patterns (Entra tenants), runtime enforcement via gateways and policies, and isolating action-taking agents inside sandboxes and microVMs.
For teams deploying agents in production, the practical takeaway is to design “what can run, where, and under what identity” as a first-class architecture concern. MCP (Model Context Protocol) keeps showing up as the interoperability layer for tools and context, but the governance story depends on policy engines (for example OPA), telemetry (OpenTelemetry), and hardened execution environments (Kata, microVM sandboxes).
Federated governance across Entra tenants with Agent 365, APIM, Purview, OPA, and OpenTelemetry
Microsoft shared a federated reference architecture for governing AI agents across multiple Microsoft Entra tenants. The design uses Agent 365 for centralized inventory, ties agents to identity via Microsoft Entra Agent ID, and combines Foundry and Copilot Studio with Microsoft Purview for governance, plus an Azure API Management (APIM) AI gateway for runtime enforcement, evaluation, and operations.
This architecture picks up where last week's MCP authorization patterns left off by showing how identity, policy, and telemetry fit together as a repeatable enterprise control plane across many tenants. The architecture explicitly calls out policy (including Open Policy Agent) and observability (OpenTelemetry) so teams can standardize controls across tenants without centralizing every workload into a single directory. If you operate in regulated or multi-subsidiary environments, this is a useful blueprint for “distributed build, centralized guardrails.”
Execution boundaries for action-taking agents with OpenSandbox + AKS pod sandboxing (Kata)
When agents take actions (not just generate text), isolation becomes the safety boundary. A reference implementation using AKS separates MCP tool interfaces from sandbox lifecycle management (OpenSandbox) and runtime isolation (Kata Containers), and it compares the approach with Azure Container Apps Sandboxes and Dynamic Sessions.
This extends last week's “agents need verified execution” message into concrete infrastructure primitives, showing how to separate tool authorization from the environment that actually runs untrusted code. The core idea is to make tool-calling capability explicit and constrain the blast radius: isolate runtime, control credentials (including vaulting), and keep lifecycle operations separate from the agent's tool surface. If you are building agent tools that run untrusted code or touch sensitive systems, this is a strong reminder that “tool permissions” and “runtime isolation” are different layers you need together.
Azure Container Apps Sandboxes as a managed isolation option
Azure Container Apps Sandboxes were highlighted as an isolation primitive for hosting agent workloads, using microVM-based isolation plus egress proxy/rules and logging support. The walkthrough content focused on lifecycle concepts, pricing, and managing sandbox groups, positioning this as a security-focused hosting option when agents need network controls and stronger tenant isolation.
If you are evaluating AKS vs Container Apps for agent runtimes, this builds on last week's discussion of sandboxing and evaluation boundaries by pointing to a more managed way to get isolation without assembling every component yourself. If you are evaluating AKS vs Container Apps for agent runtimes, ACA Sandboxes are shaping up as the “managed secure-by-default” path, while AKS gives you more composability when you need custom sandbox lifecycle logic or tighter Kubernetes integration.
Data foundations for agentic AI: Fabric IQ, OneLake governance, and vector search at scale
Microsoft's Fabric and SQL announcements this week reinforced that “agent-ready” systems depend on governed semantics, provenance, and permissioned context, not just access to raw tables, echoing last week's Fabric IQ over MCP examples that showed ontology-driven retrieval and normalization in practice. Fabric IQ is positioned as the semantic and governance layer for Copilot and agent experiences, while OneLake expands governed access patterns (shortcuts, mirroring, APIs for Delta/Iceberg) and adds policies and auditing.
On the retrieval side, Azure SQL Database Hyperscale continues to push vector search into core data platforms via DiskANN-based indexes, with claims up to 1B rows and guidance on filtering behavior and DML support. For developers building RAG systems, the message is that vector search and semantic context are becoming first-class capabilities in the database and analytics stack, reducing the need for bespoke pipelines.
Fabric IQ and Real-Time Intelligence for “agent-ready” governed context
Multiple FabCon/SQLCon updates emphasized Fabric IQ as a governed semantic/context layer for Copilot and agents, including ontology-based semantic modeling with governance and CI/CD support. Fabric Real-Time Intelligence updates added operational features such as Observability Insights and agentic investigation patterns (for example Operations Agent/Investigator), tying real-time signals to AI operations.
This is the platform-level continuation of last week's hands-on posts about using Fabric IQ via MCP for agent retrieval, but with more emphasis on governed rollout and operational ownership (CI/CD, observability, and investigation loops). If you are building enterprise agents, this pushes architecture toward “publish trusted data products + semantics, then let agents consume them with permissioning and auditability.” It is also a signal that governance controls are being threaded into Copilot experiences, not kept separate in admin-only tooling.
- FabCon and SQLCon 2026 in Barcelona: Building the data foundation for Microsoft Copilot and agents
- Trusted AI starts with Microsoft Fabric, Real-Time Intelligence, and IQ
- Bringing governed analytics into the flow of work: Fabric Analytics at FabCon Europe 2026
OneLake governance and interoperability expand (including governed table access APIs)
OneLake updates focused on expanding shortcuts and mirroring, improving partner interoperability, and adding on-demand billing options. There were also new APIs for governed access to Delta Lake and Apache Iceberg tables, plus broader catalog governance features (policies, auditing, and security improvements) aimed at controlling access across platforms.
This connects cleanly to last week's focus on permissioned context for agents: if OneLake becomes the governed substrate, the same access and provenance rules can flow into the data agents you build on top. For teams trying to avoid “copy data into every tool,” these changes push toward a single governed lake with multiple compute engines and agent experiences sitting on top. If you are introducing agents, the key benefit is consistent permissions and provenance across the contexts you feed into prompts and tools.
- FabCon and SQLCon Barcelona 2026: What’s new in Microsoft OneLake and its rapidly growing ecosystem
- Build, deploy, and govern Microsoft Fabric at scale
Vector indexing in Azure SQL Database Hyperscale (DiskANN) reaches GA for PaaS and Fabric SQL
Azure SQL Database Hyperscale is spotlighting DiskANN-based vector indexing with performance claims up to 1B rows and sub-second search in best-case scenarios, along with details on filtering behavior and DML support. The engineering team also stated vector index support is generally available in Azure SQL PaaS and Microsoft Fabric SQL, reinforcing that vector search is moving into mainstream relational platforms.
For developers building RAG, this makes it more viable to keep embeddings and retrieval close to transactional data without standing up a separate vector database, complementing last week's theme that operational simplicity matters (evaluation, observability, and fewer moving parts). It also suggests you should pay attention to operational constraints (index maintenance, filter selectivity, and write patterns) since “vector in the database” changes the performance profile of mixed workloads.
Cost, routing, and evaluation: making model choice measurable
This week had a noticeable emphasis on tooling that turns model selection into an engineering decision instead of taste, building on last week's Copilot cost controls (Auto tiers, budget escalation) by adding more explicit evaluation harnesses for routing. Azure AI Foundry Model Router and its Auto Evaluation toolkit were highlighted as a way to compare routing decisions against a baseline model (for example GPT-5), using metrics like quality, cost, latency, and model distribution, producing an HTML dashboard to guide trade-offs.
Alongside this, FinOps guidance and Databricks governance patterns pointed to budgets, rate limits, and telemetry as the controls you need when experimenting with newer frontier models that may be more expensive by default. The common thread is that “try a new model” now needs guardrails and measurement from day one.
Model Router Auto Evaluation: test if routing beats “always use the big model”
The Model Router Auto Evaluation toolkit is designed to quantify whether Azure AI Foundry Model Router is actually improving value relative to a fixed baseline model. It frames evaluation across quality, cost, latency, and value, and produces an HTML dashboard with charts to make routing behavior and trade-offs visible.
This mirrors last week's Copilot-side move toward runtime model choice (Auto tiers) by giving teams a repeatable way to prove that routing is helping rather than quietly concentrating spend on a larger model. If your team is debating “do we really need the biggest model for this workflow,” this kind of harness is useful because it turns the answer into a repeatable benchmark you can rerun when prompts, tools, or model availability changes. It also helps catch cases where a router silently concentrates traffic on a more expensive model than expected.
FinOps and spend controls for AI workloads (Azure OpenAI/Foundry and Copilot)
Microsoft's September 2026 FinOps roundup pointed to pricing and cost-control updates across Azure AI Foundry/Azure OpenAI and GitHub Copilot, including model availability and deployment pricing changes plus expanded usage and billing metrics. Separate guidance content reiterated the basics that still matter most in practice: budgets, alerts, and monitoring, backed by tools like the Azure Pricing Calculator.
This extends last week's “cost and policy levers” theme from Copilot into the wider Azure AI estate, which matters if you are mixing Copilot usage-based billing with first-party model endpoints in the same budget conversations. If you run AI features in production, the lesson is to treat cost controls as part of the deployment checklist, not a follow-up task. Usage-based billing (both in cloud model endpoints and Copilot) makes visibility and guardrails a prerequisite for safe rollout.
Databricks governance patterns: routing, budgets, and telemetry for multiple coding agents
Azure Databricks described how Unity Gateway CLI (ug) can standardize authentication, routing, and governance for multiple coding agents by reusing Unity Catalog permissions, adding rate limits (QPM/TPM), and enforcing account budgets. The posts also covered Smart Routing (Beta) and cost-risk isolation patterns for newly launched models using layered budgets and OpenTelemetry-based tracking.
This is a data-platform analogue to last week's agent governance and observability focus: route all agent/model traffic through enforceable gateways, then use telemetry to decide what to scale. Even if you are not a Databricks customer, the governance pattern generalizes well: unify authZ around existing data permissions, route model traffic through a gateway that can enforce budgets, and export telemetry so cost and quality decisions are based on evidence. This fits neatly with the broader theme that multi-model environments need centralized controls.
- One Command Opens Claude Code, Another Opens Codex: How Azure Databricks Governs Coding Agents
- A New Model Costs 60 Percent More by Default: Isolating the Financial Risk of Testing Frontier AI
- Not Every AI Decision Needs a Chat: Running a Pure Decision Model Straight from SQL
Building agents and agentic apps: MCP tools, hosted skills, and agent-centric UI
This week had several “how-to” updates for building agents that do real work: MCP-based tool connections, hosted skills you can run and debug locally before deployment, and UI components that treat agents as first-class participants in the product experience, continuing last week's thread about standardizing skills and MCP tool access while security and packaging catch up. The net effect is that the ecosystem is filling in the missing middle between “prompt in chat” and “ship a governed, observable agent.”
For developers, this is where architecture choices start to show up in code: how you package tools, how you debug tool runs, where human approvals live, and how you represent agent steps in the UI (streamed blocks, evidence, and state). It also ties back to governance because every new tool integration is a new security boundary you need to manage.
Hosted Skills Canvas in the GitHub Copilot app (Azure Functions + MCP)
The Azure Functions Hosted Skills Canvas (preview) in the GitHub Copilot app provides a local, event-driven workflow for building, running, and debugging hosted skills with triggers, logs, and evidence-backed outputs. It also covers deployment to Azure via Azure Developer CLI (azd) and exposing the hosted skill as an MCP tool endpoint.
This builds on last week's “skills as reusable units” guidance by giving teams a more standardized, developer-friendly place to implement and validate those skills before they become shared tooling. This is a concrete step toward making agent tools feel like normal developer artifacts: versioned, testable, debuggable, and deployable. If you are building internal automations, the canvas approach can help you standardize how skills emit logs and evidence so reviews focus on behavior, not trust.
Managed browser automation via Playwright Workspaces over MCP
A step-by-step guide showed how to connect Microsoft Playwright Workspaces (Playwright Cloud Browsers) to agent applications using MCP, with an example using the GitHub Copilot app. It emphasized human approval before submission plus evidence capture and diagnostics (logs, traces, screenshots, recordings) for observability and debugging.
This is a concrete continuation of last week's push for verified execution and agent observability: browser automation is where “show me what happened” artifacts often matter more than the agent's summary. For agent builders, this is a practical blueprint for “safe browser automation”: run the browser remotely, collect artifacts automatically, and build approval gates into workflows where a mistaken submission would be costly. It also reinforces MCP's role as the standard connector between agents and external tools.
Copilot plugins for distributing versioned skills, instructions, and MCP tool connections
Guidance on GitHub Copilot plugins focused on how to distribute versioned agent instructions, skills, and MCP tool connections across teams, including publishing via an internal marketplace and governing updates and precedence. The goal is consistency: teams should not be hand-configuring the same prompt/tool bundle in every repo or IDE.
This follows directly from last week's “documented skills + repo readiness” theme by giving orgs a packaging and rollout mechanism that can be reviewed and updated like any other dependency. If you have multiple teams building agent workflows, plugins are a governance tool as much as a productivity tool. They give you a packaging unit for “what the agent is allowed to do and how it should behave,” which can be reviewed like any other dependency.
Agentic UI with Blazor AI Components (experimental in .NET 11 RC1)
Microsoft introduced experimental Blazor AI Components in .NET 11 RC1 for building “agentic UI” where agents participate in the UI beyond a chat box. The components cover patterns like streamed content blocks, tool rendering, approvals, and shared/predictive state, with an end-to-end sample that uses Microsoft.Extensions.AI, AG-UI, Microsoft Agent Framework, ASP.NET Core, Foundry, and .NET Aspire.
This complements last week's emphasis on evidence, evaluations, and approvals by showing how those concepts can be first-class UI elements, not just hidden logs behind a chat transcript. For app developers, the big change is that you can design the UI around steps, artifacts, and approvals rather than free-form chat transcripts. That makes it easier to build software that is auditable (what did the agent do?), safer (where did the user approve?), and more usable (outputs rendered as UI, not text).
- Build Agentic UI with the new Blazor AI components
- Blazor Community Standup: Build agentic experiences with the Blazor AI Components
AI security and trustworthy operations: threats, accountability, and repo hygiene
Security coverage this week split into two tracks: real-world threat reporting (phishing and malware campaigns) and a growing focus on securing AI agents themselves through identity, access control, monitoring, and accountability, extending last week's “secure tooling and repo readiness” thread into both threat intel and practical hardening steps. Microsoft’s Digital Defense Report themes reinforced that AI is part of attacker workflows now, which raises the urgency for defenders to treat AI systems (and especially action-taking agents) as security-relevant infrastructure.
On the developer side, GitHub continued to push “security hygiene by default” tooling, including one-command repo hardening and AI-assisted security research workflows. The message is consistent: the faster you can apply guardrails and gather evidence, the safer it is to adopt agentic tooling.
Digital Defense Report themes: interconnected risk and securing AI agents
Microsoft highlighted 2026 threat themes where identities, infrastructure, apps, and supply chains create interconnected risk, and it explicitly called out AI’s growing role in attacker workflows. The guidance emphasizes securing AI agents through identity, access control, and monitoring, including threats like prompt injection and AI-assisted vulnerability discovery.
This reinforces last week's theme that agent security is not abstract: identity, authorization, and telemetry are the control points that make tool use and outcomes defensible in real incidents. A companion piece focused on government resilience priorities, including faster response to AI-accelerated threats and planning for incident spread via identity compromise. Even if you are not in the public sector, the operational takeaway is that identity security and telemetry are foundational, because they are the pivot points when incidents cascade.
- Insights from the 2026 Microsoft Digital Defense Report
- Preparing governments for an era of interconnected cyber risk
Threat intel: Star Blizzard’s RedFlick technique and CosmicPulse backdoor
Microsoft Threat Intelligence reported that Star Blizzard shifted to higher-volume phishing and introduced the RedFlick delivery chain, using scheduled tasks for persistence to deploy the CosmicPulse backdoor. The post included mitigations, Microsoft Defender detections, hunting queries, and indicators of compromise (IOCs), with references to Microsoft Sentinel ASIM queries for defenders.
For security teams, this is immediately actionable content: update detections, run hunts, and validate scheduled-task monitoring coverage. For developers and platform engineers, it is another reminder that the systems your agents and CI/CD pipelines rely on (identity, endpoints, email) are exactly where attackers are investing.
GitHub security tooling: gh-secure and AI-assisted vuln research
GitHub highlighted gh-secure from the GitHub Security Lab as a one-command way to improve repo security hygiene and reduce accidental secret leaks in public code, available via the Copilot app, Copilot CLI, or GitHub CLI. Separately, GitHub shared how its open-source Taskflow Agent was used to audit Android apps and disclose 24 vulnerabilities, with a Codespaces-based runbook and notes on false positives and severity estimation.
This aligns with last week's emphasis on treating agent work as operational engineering: packaged runbooks, reproducible environments, and reviewable evidence, not one-off “AI said so” findings. The combination is useful: gh-secure is about preventing basic mistakes at scale, while the Taskflow Agent story shows how to operationalize LLM taskflows for targeted audits without treating the model as an oracle. If you adopt either, build in review loops and evidence collection so humans can validate findings before acting.
- How to protect your repo in 2 minutes with gh-secure
- How we found 24 Android vulnerabilities using our open source AI security agent
AI infrastructure and deployment patterns: private GPU extension and serverless GPU benchmarks
AI infrastructure updates this week focused on how to scale GPU capacity and where different model types fit operationally, continuing last week's interest in production-grade agent systems where performance, isolation, and operating model choices show up as engineering constraints. Microsoft outlined AI Points of Presence (AI PoPs) paired with Azure ExpressRoute to privately connect distributed, third-party GPU capacity into Azure while keeping a consistent operating model.
At the application layer, a case study benchmarked a self-hosted bounded decision model (served by Ollaya) on Azure Container Apps serverless GPU against GPT-5.4 Nano via Azure OpenAI. The point was not “local is always better,” but that closed-set decision models can outperform general LLM calls on latency, cost, and predictability when the problem is bounded.
- Scaling distributed AI infrastructure with Azure’s ExpressRoute and AI Points of Presence (AI PoPs)
- Run an Ollaya decision model on Azure Container Apps
Other Artificial Intelligence News
Microsoft’s monthly AI and Azure video roundups continued to be a useful “what changed” sweep, including Azure AI Foundry model additions, Model Router changes, agent features (A2A, routines, voice agents), and governance controls plus Copilot platform items across GitHub Copilot and VS Code agent workflows. John Savill’s Azure Weekly Update also included an AI model availability note alongside broader platform changes, and it called out Azure SQL Database vectors in the wider Azure update stream.
On the platform side, SQL Server on Azure Local reached general availability for connected and disconnected operations, targeting sovereign and edge environments, with Foundry Local on Azure Local in preview to run AI inference close to SQL Server data while keeping processing on-premises. If you are building hybrid or disconnected AI solutions, this is a key “where can inference run” milestone.
Microsoft also continued pushing MCP as the connective tissue for grounding and tooling: the Microsoft Learn MCP Server aims to help assistants validate answers against current Learn documentation, while Azure AI Foundry Toolbox demos showed MCP-governed toolsets with approval-gated actions to separate read access from execution, picking up on last week's theme that MCP adoption is accelerating while authorization and evidence trails become the real differentiators.