Weekly Azure Roundup: Foundry agents, identity security, and runtime isolation
Welcome to this week's Azure roundup, where the focus shifted from agent prototypes to agent platforms you can standardize and run. Microsoft Foundry expanded model options (including new GPT-6 SKUs and Claude Opus 5.5), added production primitives like Routines scheduling, hosted long-running execution, and deeper trace-driven evaluation and optimization. In parallel, Azure doubled down on controlling blast radius with Entra-backed identities and security guidance after real-world service principal abuse, while the container stack (AKS and Container Apps Sandboxes) pushed stronger isolation and scaling patterns for AI workloads. We also cover practical platform work: managed connectors in Azure Functions (preview), resilience validation beyond diagrams, experimental Bicep doc generation, and several storage and integration gotchas that can affect day-2 operations.
This Week's Overview
- Microsoft Foundry and Agent Framework move closer to production defaults
- Expanded model lineup and deployment options in Foundry (GPT-6, Claude)
- Routines (GA) and agent reliability features (voice agents, long-running execution)
- Continuous optimization: traces, evaluation, routing, and cost controls
- Isolation, memory, and “agent-first” platform patterns (Foundry + Container Apps Sandboxes)
- Securing identities and controlling blast radius in Azure
- AKS and Container Apps add stronger isolation and scale options for AI workloads
- Azure Functions expands eventing reach with managed connectors (preview)
- Resilience engineering: beyond diagrams, toward continuous validation
- Infrastructure as code: Bicep adds experimental doc generation
- Storage and data operations: immutability gotchas, durable agent memory, and observability
- Integration and migration watchlist: Log Ingestion, Service Bus, and Oracle on Azure
- Other Azure News
Microsoft Foundry and Agent Framework move closer to production defaults
This week’s Azure AI news centered on turning agents into something you can ship and operate, not just prototype, building directly on last week's “platform path” guidance (landing zones first, then Foundry) by adding more concrete control-plane and run-time defaults teams can standardize across orgs. Microsoft Foundry expanded model choice (including new GPT-6 SKUs and Claude Opus 5.5), added more ways to run agents reliably (voice agents, long-running hosted execution), and doubled down on evaluation and optimization loops so teams can tune cost, latency, and quality using real traces.
A clear theme across the updates is separation of concerns: governance and identity up front (Entra-backed identities, guardrails, tracing, evaluation), with execution isolated and resilient (hosted sessions, background workflows, and sandboxed runtimes). If you are standardizing an “agent platform” inside your org, these releases push Foundry toward a more complete control plane for model routing, observability, scheduled execution, and policy enforcement.
Expanded model lineup and deployment options in Foundry (GPT-6, Claude)
Microsoft introduced GPT-6 Sol and GPT-6 Luna as generally available in Foundry alongside GPT-6 Astra, with practical guidance on choosing models for production agents, which fits neatly with last week's emphasis on operationalizing RAG and agents with consistent guardrails and cost attribution rather than treating model selection as a one-off team decision. The announcement calls out multiple deployment modes (Standard, Provisioned Throughput, and Priority Processing) plus regional/data-zone considerations and safety controls like guardrails and prompt-injection mitigation.
Foundry also added Claude Opus 5.5 availability for long-running coding and knowledge-work agents, including features like adaptive thinking controls and beta capabilities such as tool changes with prompt caching and asynchronous compaction. For teams doing sustained agentic workflows (coding agents, research assistants, multi-step analysis), the key takeaway is you can now mix vendors within the same Foundry operational model and select pricing/latency modes based on workload profile.
- GPT-6 Astra, Sol, and Luna: For production agents in Microsoft Foundry
- Claude Opus 5.5 comes to Microsoft Foundry for long-running coding and knowledge work
Routines (GA) and agent reliability features (voice agents, long-running execution)
Routines in Foundry Agent Service reached general availability, adding managed execution via timer, recurring schedules, and event-based triggers, with run history and auditing, extending last week's “day 2” checklist into an explicit scheduling-and-audit primitive you can operationalize across teams. The release also introduced a reminder tool (preview) aimed at letting agents resume work, and it clarified identity choices (creator vs agent identity) backed by Microsoft Entra ID, which matters when you need least-privilege separation and clean audit trails.
In parallel, Foundry’s broader “ship agents faster” update emphasized voice agents and resilience improvements for long-running hosted agents. Combined, these changes make it easier to treat agents like scheduled jobs and event handlers with operational guardrails, rather than one-off chat sessions.
- From Chatbots to Automated Assistants: Routines in Microsoft Foundry Are Now Generally Available
- Ship agents faster with expanded model choice, voice agents, and continuous optimization
Continuous optimization: traces, evaluation, routing, and cost controls
Foundry is leaning hard into “observe then optimize” as a standard operating loop for agents, reinforcing last week's message that governance and cost control belong on the request path (token limits, rate controls, attribution) by showing how traces and evals feed concrete tuning decisions over time. The updates highlighted Insights in Foundry, evaluation tooling, and an Agent Optimizer that turns production traces into concrete suggestions spanning instructions, tools/skills, and model selection.
A complementary capability is the Azure AI Foundry model router, which routes requests to different models based on complexity to balance cost and quality. The practical value is that teams can measure which model served each request and evaluate the router as a unit, which is exactly what you need when you start mixing frontier and cheaper models behind a single endpoint.
- One model deployment to route them all
- How do you know what to optimize next in your AI agent? Ask The Experts!
- Ship agents faster with expanded model choice, voice agents, and continuous optimization
Isolation, memory, and “agent-first” platform patterns (Foundry + Container Apps Sandboxes)
Several posts reinforced that agent platforms need both isolation boundaries and durable memory patterns, continuing last week's “composable building blocks” framing (identity, access, observability, security) by putting clearer lines between Foundry governance and where untrusted agent code actually runs. Microsoft Agent Framework added production-oriented capabilities like AG-UI endpoints for interactive experiences, Foundry-backed semantic memory via context providers, CodeAct execution with Hyperlight, and resilient background hosting for recoverable workflows across Python and .NET.
On the platform side, an “agent-first” architecture pattern emerged: keep governance (Entra Agent ID, tracing, evaluation) in Foundry, but run executions in Azure Container Apps Sandboxes with per-task microVM isolation. In practice, this helps teams scale autonomous agent workloads while keeping permissions and compliance controls in the right place, and it aligns with recent GA announcements for Container Apps Sandboxes and the new “Express” path for image-to-app deployments.
- What’s new in Microsoft Agent Framework: Interactive experiences, memory, and resilient execution
- Designing agent-first platforms: What changes when agents do the work
- Azure Container Apps Sandboxes, Now Generally Available
- Azure Container Apps Express is now Generally Available
Securing identities and controlling blast radius in Azure
This week also delivered a sharp reminder that cloud security failures are often identity failures first, echoing last week's focus on standardizing Entra-backed auth as a foundational “landing zone” concern rather than an app-by-app choice. Microsoft Security Research documented Storm-3168 (JADEPUFFER) activity where attackers used compromised Azure service principals to perform reconnaissance, execute rapid destructive operations across storage and app resources, and retrieve storage keys.
The mitigation guidance puts workload identity protection and least privilege at the center, along with recovery safeguards (resource locks, protected deletion patterns, and validated restore processes) and better detections via Defender for Cloud and Defender XDR. The report frames a broader shift toward AI-orchestrated attacks, which means defenders should assume faster execution and more automation once credentials are stolen.
AKS and Container Apps add stronger isolation and scale options for AI workloads
Azure’s container stack continued to evolve around two themes: isolation for untrusted/agent code and elasticity for bursty AI workloads, complementing last week's guidance to treat platform security and policy as reusable building blocks by hardening the runtime layer where agents and AI services actually execute. AKS announced a set of AI-focused capabilities, including pod sandboxing with Kata Containers, confidential GPU support (AMD SEV-SNP with NVIDIA H100), identity and storage security improvements, a GitOps path via an Argo CD extension, and new scaling and efficiency options like Virtual Nodes v2 for bursting.
In more detail, the next-generation AKS virtual nodes backed by Azure Container Instances position ACI as a serverless compute layer that still fits standard Kubernetes workflows (kubectl, Helm). The guide also covered confidential containers, including generating Confidential Container Environment (CCE) policies using confcom (Rego policies) and validating execution with hardware-backed attestation, which is relevant when you need stronger runtime guarantees for multi-tenant or sensitive workloads.
- Run and scale AI applications on AKS – the latest product innovations
- Virtual nodes on Azure Container Instances: a new compute layer for AKS
Azure Functions expands eventing reach with managed connectors (preview)
Azure Functions added integration with Azure Connector Namespace in public preview, bringing connector triggers and typed connector clients so functions can respond to events and invoke actions across roughly 1,700 managed connectors, which fits the same “repeatable platform work” theme from last week by reducing bespoke integration glue that is hard to govern consistently. The provided .NET sample uses DefaultAzureCredential and demonstrates an RFP intake flow that ties together SharePoint and Teams connectors plus Azure AI Content Understanding for document extraction.
For teams that have been stitching together SaaS and Microsoft 365 workflows with custom webhooks and glue code, this preview shifts more of that integration surface into a supported connector model. It also raises new operational questions (connector permissions, identity boundaries, connector-trigger scaling), so treat it like an early platform capability and validate governance and monitoring before broad adoption.
Resilience engineering: beyond diagrams, toward continuous validation
Azure leadership revisited resilience as an engineering discipline, not an architectural aspiration, and it aligns with last week's push for “day 2” operational requirements by arguing you only earn reliability through repeatable validation, not design intent. Mark Russinovich described a 2014 bug that nearly caused a worldwide outage and how it directly shaped Azure’s “safe deployment policy,” while a separate post argued that architecture diagrams capture intent but do not prove resilience without continuous validation.
The practical direction is clear: invest in health models, infrastructure-as-code (IaC) rigor, and ongoing chaos testing (Azure Chaos Studio) so you can validate behavior as dependencies change. The posts also call out how AI and agents introduce non-determinism, which means reliability assumptions need guardrails, better accountability models, and testing that covers AI dependencies as first-class components.
- The 2014 Bug That Almost Broke Azure
- Your architecture diagram is not your resilience
- Resilience at Cloud Scale: Azure CTO on Outages, Hardware, and AI
Infrastructure as code: Bicep adds experimental doc generation
Bicep introduced an experimental CLI command, bicep docs generate, to generate module documentation directly from compiled Bicep modules, a practical extension of last week's landing-zone-first mindset where repeatability depends on well-maintained IaC artifacts (and the docs teams use to apply them correctly). The feature supports customization using Scriban templates, and the post shows how to run generation across a repo and enforce up-to-date docs in CI so documentation stays aligned with module behavior.
If you maintain internal Bicep registries or publish Azure Verified Modules, this is a practical step toward treating docs as a build artifact. It also encourages teams to standardize module metadata and outputs, since those become the source for consistently generated documentation.
Storage and data operations: immutability gotchas, durable agent memory, and observability
Several posts landed on real-world operational friction in storage and data platforms, especially where compliance features change deletion and lifecycle behavior, which mirrors last week's framing that governance needs enforcement points and “day 2” workflows, not just policy statements. One troubleshooting guide explains why storage accounts and containers can become effectively undeletable after enabling Azure Blob Storage immutability (WORM), including the difference between locked vs unlocked policies, the impact of legal holds, and how hidden blob versions can block cleanup.
Separately, agent builders got more patterns for durable state using Azure Blob Storage. A LangChain Deep Agents tutorial introduces an AzureBlobBackend to treat Blob Storage as a durable, shareable virtual filesystem across agents and runs, with guidance on DefaultAzureCredential/ManagedIdentityCredential, RBAC, and safety features like soft delete and versioning.
- Troubleshooting Azure Storage Deletion: Immutability, WORM, and Version Retention
- Using Azure Blob Storage as a durable filesystem for LangChain Deep Agents
Integration and migration watchlist: Log Ingestion, Service Bus, and Oracle on Azure
A few integration/migration items are worth flagging because they affect existing workloads and deadlines, extending last week's emphasis on platform checklists and operational guardrails into the less-glamorous but critical work of keeping integrations compliant with API limits and retirement timelines. Azure Monitor’s Data Collector API deprecation continues to push teams to the Logs Ingestion API, and one migration write-up focuses on handling the 1 MB payload limit via array chunking and for-each loops to send smaller batches reliably.
On messaging, Azure Service Bus has an immediate deadline: Azure is ending SBMP support on September 30, 2026, and BizTalk Server 2020 users need to validate they are actually using AMQP (via KB5091379). A separate guide explains how Azure API Management’s send-service-bus-message policy lets you expose governed HTTP APIs that enqueue work to Service Bus with managed identity and patterns for queues vs topics/subscriptions, plus message controls like sessions, TTL, and duplicate detection.
- Migrate Data Ingestion from Data Collector to Log Ingestion - Part 2
- How to Validate That Your BizTalk SB-Messaging Adapter Is Really Using AMQP
- What the Azure API Management integration means for Azure Service Bus
- Oracle AI Database@Azure Expands with Azure Native Oracle GoldenGate User Experience, Observability
Other Azure News
Several items this week focused on practical guidance and operational readiness across cost controls, networking, troubleshooting, and applied AI reference architectures, and they continue last week's through-line of making enterprise AI and platform operations more repeatable (fewer bespoke decisions, more standardized controls and runbooks). The common thread is reducing “unknown unknowns”: understanding what changed, what will retire, and what you can no longer assume (whether that’s reservation flexibility, connectivity patterns, or the behavior of time types during migrations).
- Azure Weekly Update - 25th September 2026
- You Get One Exchange Left: Rethinking Azure Commitments Before February 2027
- Azure Multicloud Interconnect for AWS Brings Simpler Private Connectivity Between Clouds
- Change Analysis in Azure Resource Graph: Your first stop for “What changed?”
- Azure Retirements Livestream - Please register! Session 1 - Tracking ID: 0P59-60Z
- From Camera to Canvas: How Multimodal GenAI Automates Shop Floor Instructions
- .NET Memory Dumps, Database Time Migration & AI Agent Memory | The Upload
- .NET 11 Performance Update, Azure Resiliency & Multicloud Databases | Ep 6
- Agentic RAG and chat completion in the Microsoft SQL engine | Data Exposed
- Elasticsearch Vector Database on Azure: semantic search and RAG without managing a cluster
- Expanding access to specialized GPU infrastructure through Dapple and Azure
- Azure Local Storage: Choose the Architecture That Fits Your Business
- Azure Local Tools: Community Projects, Microsoft Official Utilities, and OEM Resources
- Microsoft Sovereign Private Cloud and Azure Local: Much More Than a Virtualization Platform
- Getting the best out of Azure SRE Agent
- The Partition Count You Picked on Day 1 Doesn't Have to Follow You Forever
- Catalyst: Frontier stories of AI Infrastructure innovation
- Decoding Environmental Viromes: Amplifying Human Expertise with Microsoft Discovery
- How AI is helping scientists track viruses through wastewater
- Engineering Agentic Recall Controls with MCP and Microsoft Foundry
- Foundry hosted agent isolation with Microsoft Agent Framework
- Control where your hosted agent connects with network egress in Foundry Agent Service
- Microsoft Foundry Hosted Agents and MCP in Practice: Building Fibey Field Ops
- When the Coding Agent Leaves the Codebase: Building a Multi-Runtime AI Agent Infra with Azure KARS
- Beyond Cherry-Picking: Evaluating Text-to-Image Models