Am I paying for a giant model I don't need?
Microsoft Developer explains how to evaluate Azure AI Foundry Model Router against a baseline model (for example GPT-5) to understand cost, latency, and quality trade-offs, using the open-source Model Router Auto Evaluation toolkit and its HTML dashboard output.
Overview
The video focuses on deciding whether a single large “giant” model is worth paying for, or whether routing requests across a selection of models via Model Router can reduce cost without unacceptable quality loss.
It demonstrates using the open-source Model Router Auto Evaluation toolkit to compare Model Router with an existing baseline model across multiple dimensions:
- Quality
- Cost
- Latency
- Value
- Model distribution (how traffic is routed across models)
The output is an HTML dashboard containing 8 charts intended to help answer whether model routing is a better fit than sticking with a single large model.
What’s covered (video chapters)
Save costs with Model Router
- Frames the core question: are you paying for a larger model than your workload needs?
- Introduces Model Router as a way to route to “the best model for the task” with cost/quality trade-offs.
Model Router Auto Evaluation
- Introduces the evaluation toolkit used to compare Model Router vs a baseline model.
- Highlights the comparison dimensions: quality, cost, latency, value, and model distribution.
Configure models and pricing
- Shows configuring the models involved in the comparison.
- Includes setting up pricing inputs so cost comparisons reflect the workload.
Prepare the evaluation dataset
- Covers preparing a dataset used to run the evaluation.
Run and view the evaluation
- Runs the evaluation and views results.
- Produces an HTML dashboard with multiple charts.
Compare cost, latency, and quality
- Reviews the dashboard outputs to compare Model Router against the baseline model.
Analyze the quality breakdown
- Digs into how quality varies across the evaluation, supporting a more granular decision than a single aggregate score.
Choose the right trade-offs
- Uses the evaluation results to decide whether to keep a single large model or adopt routing across multiple models based on acceptable trade-offs.