When my agent is slow or wrong, where do I even look?
Microsoft Developer walks through a practical workflow for figuring out why an AI agent is slow or producing incorrect answers, using tracing in Azure AI Foundry and Azure Application Insights.
Overview
The video focuses on diagnosing two common agent problems:
- The agent takes too long to answer.
- The agent answers with confidence but is wrong.
The core idea is that without traces you’re guessing whether the issue is:
- The model
- A tool call
- The agent’s instructions
With tracing enabled, you can inspect an agent run end-to-end and pinpoint the exact step that was slow or incorrect.
Enabling tracing with Azure AI Foundry + Application Insights
- Connect Application Insights to your Microsoft (Azure) AI Foundry project.
- This enables server-side tracing with no code changes.
Reading an agent run span-by-span
Once tracing is enabled, you can inspect a single agent run as a sequence of spans, including:
- The overall agent invocation
- Each model call
- Every tool call, including:
- Tool arguments
- Tool result
- Token usage
- Latency
This lets you identify:
- The slow step (where latency is coming from)
- The wrong step (where the agent’s behavior diverged)
Querying traces to find patterns with KQL
After identifying an issue in a single run, the video shows using KQL queries in Application Insights to determine whether:
- It was a one-off bad run, or
- It’s a recurring pattern across runs (systemic issue)