Stop guessing which model your agent needs
Microsoft Developer shows how to use the Microsoft Foundry model catalog to choose an AI model based on measured results instead of intuition, comparing deployments with the same evaluation while tracking quality, latency, token usage, and cost.
Overview
The video walks through selecting the right model for an agent using Microsoft Foundry’s model catalog and a repeatable evaluation approach.
Key points covered:
- Using the model catalog to pick candidates based on required capabilities
- Comparing different deployment options by swapping only the deployed model
- Re-running the same evaluation across models to compare outcomes consistently
- Reviewing model availability and fine-tuning considerations
- Using benchmarks as an input to model selection
- Configuring a model deployment
- Measuring real-run signals:
- Latency
- Token usage
- Output quality
- Comparing cost across models
- Managing quotas and optimizing model choice for production
Links:
- Foundry portal: https://aka.ms/foundry-portal
- Playlist: https://aka.ms/InsideMicrosoftFoundryPlaylist