Am I paying for a giant model I don't need?

Microsoft Developer explains how to evaluate Azure AI Foundry Model Router against a baseline model (for example GPT-5) to understand cost, latency, and quality trade-offs, using the open-source Model Router Auto Evaluation toolkit and its HTML dashboard output.

Overview

The video focuses on deciding whether a single large “giant” model is worth paying for, or whether routing requests across a selection of models via Model Router can reduce cost without unacceptable quality loss.

It demonstrates using the open-source Model Router Auto Evaluation toolkit to compare Model Router with an existing baseline model across multiple dimensions:

The output is an HTML dashboard containing 8 charts intended to help answer whether model routing is a better fit than sticking with a single large model.

What’s covered (video chapters)

Save costs with Model Router

Model Router Auto Evaluation

Configure models and pricing

Prepare the evaluation dataset

Run and view the evaluation

Compare cost, latency, and quality

Analyze the quality breakdown

Choose the right trade-offs