AI Model Routing Explained With a Real Application
Learn AI model routing by building a real app that sends tasks to the best model based on cost and latency.
The Problem With One Model for Everything
You have a chat feature, a summarizer, and a data extractor. You picked one large language model and pointed every request at it. It works, but the bill is painful and responses to simple tasks feel slow.
The fix is model routing: sending each request to the model that best matches the task, based on cost, latency, and quality. In this post you will build a small but real routing layer that inspects a request, decides which model to call, and falls back when a model is down.
- You will use Node.js and the OpenAI SDK, but the pattern works with any provider.
- You will create a config file that maps task types to models.
- You will implement a router that checks a request against rules.
- You will measure cost and latency to prove the router works.
Before You Start
You need Node.js 18 or newer, an OpenAI API key, and a terminal. The example uses model names that are current as of this writing: gpt-4o-mini for cheap tasks and gpt-4o for complex ones. Check the OpenAI model list if names change.
Create a new project and install the OpenAI SDK.
mkdir model-router
cd model-router
npm init -y
npm install openai dotenvStep 1: Define the Routing Config
Routing rules belong in a config file, not in code. This lets you change models without redeploying. Create a file called routing.json with a list of routes. Each route has a task type, a model, a max cost in dollars per request, and a max latency in milliseconds.
The router will check the request's task field against these routes. If the task is not listed, it uses the default route.
{
"routes": [
{
"task": "classification",
"model": "gpt-4o-mini",
"maxCost": 0.001,
"maxLatencyMs": 1000
},
{
"task": "summarization",
"model": "gpt-4o-mini",
"maxCost": 0.002,
"maxLatencyMs": 2000
},
{
"task": "code_generation",
"model": "gpt-4o",
"maxCost": 0.01,
"maxLatencyMs": 5000
},
{
"task": "complex_reasoning",
"model": "gpt-4o",
"maxCost": 0.02,
"maxLatencyMs": 8000
}
],
"default": {
"model": "gpt-4o-mini",
"maxCost": 0.001,
"maxLatencyMs": 1500
}
}Step 2: Build the Router Core
Create a file called router.js. It loads the config, looks at the request's task, and returns the route. If the task is unknown, it returns the default route.
The router also checks a budget. If the estimated cost of the model exceeds the max cost for the route, it picks a cheaper model. You will estimate cost later; for now the router just picks the route.
const fs = require('fs');
function loadRoutes() {
const raw = fs.readFileSync('./routing.json', 'utf-8');
return JSON.parse(raw);
}
function getRoute(request) {
const config = loadRoutes();
const route = config.routes.find(r => r.task === request.task);
return route || config.default;
}
module.exports = { getRoute };- The config file is read on every call. In production, cache it or watch for changes.
- You can add more fields like 'provider' or 'region' later.
Step 3: Call the Model and Measure
Now create a file called callModel.js. It uses the OpenAI SDK to send a prompt to the chosen model. It also measures latency and calculates cost based on token usage.
The cost calculation uses the pricing for gpt-4o-mini and gpt-4o as of this writing: gpt-4o-mini is 0.15 cents per 1K input tokens and 0.60 cents per 1K output tokens. gpt-4o is 2.50 dollars per 1M input and 10 dollars per 1M output. These are approximate and you can change them.
const OpenAI = require('openai');
require('dotenv').config();
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const PRICING = {
'gpt-4o-mini': { input: 0.00000015, output: 0.0000006 },
'gpt-4o': { input: 0.0000025, output: 0.00001 }
};
async function callModel(model, prompt) {
const start = Date.now();
const completion = await openai.chat.completions.create({
model,
messages: [{ role: 'user', content: prompt }]
});
const latencyMs = Date.now() - start;
const usage = completion.usage;
const cost = usage.prompt_tokens * PRICING[model].input + usage.completion_tokens * PRICING[model].output;
return { text: completion.choices[0].message.content, latencyMs, cost };
}
module.exports = { callModel };Step 4: Put It Together
Create an index.js that accepts a request object with a task and a prompt. It gets the route, calls the model, and prints the result with cost and latency.
Run it with a simple classification task to see the router pick gpt-4o-mini.
const { getRoute } = require('./router');
const { callModel } = require('./callModel');
async function handleRequest(request) {
const route = getRoute(request);
console.log(`Routing to ${route.model} for task: ${request.task}`);
const result = await callModel(route.model, request.prompt);
console.log(`Latency: ${result.latencyMs} ms`);
console.log(`Cost: $${result.cost.toFixed(6)}`);
return result.text;
}
// Example
handleRequest({
task: 'classification',
prompt: 'Classify the sentiment of this review: The product is amazing.'
}).then(console.log);Step 5: Add a Fallback
In production, models fail or rate limits hit. Add a fallback that tries a cheaper model if the first call throws an error. Modify callModel to accept a list of models and try each in order.
Update index.js to pass the route's model and a fallback model from the config.
async function callModelWithFallback(models, prompt) {
for (const model of models) {
try {
return await callModel(model, prompt);
} catch (err) {
console.warn(`Model ${model} failed: ${err.message}`);
}
}
throw new Error('All models failed');
}
module.exports = { callModelWithFallback };Verify It Worked
Run the example and check the output. You should see the router pick gpt-4o-mini for classification and gpt-4o for code generation. The cost and latency numbers should be printed.
Try changing the task to code_generation and see the model switch.
node index.js
# Expected output:
# Routing to gpt-4o-mini for task: classification
# Latency: 850 ms
# Cost: $0.000042Real-World Routing Strategies
The simple task-based routing above is a start. In production you also want to consider prompt complexity, token budget, and user expectations. Here are three strategies you can add.
Strategy 1: Estimate complexity by prompt length. Short prompts go to the small model, long ones to the big model. Strategy 2: Use a classifier model to decide which model to call. Strategy 3: Monitor latency and cost in real time and adjust routes dynamically.
- Check the input token count before calling. If it exceeds a threshold, use a larger context model.
- Track errors per model and disable a model if it fails more than 5% of the time.
- Log every routing decision with model, cost, and latency for later analysis.
What I Would Do
For a real application, I would start with the config-driven router you just built, then add a dynamic fallback that checks a health endpoint. I would also add a simple cache for repeated prompts.
Here is a copy-paste starter for the dynamic fallback using a health check.
async function getHealthyModel(models) {
for (const model of models) {
const health = await checkHealth(model); // fake function
if (health) return model;
}
throw new Error('No healthy model');
}
async function checkHealth(model) {
// In production, call a status endpoint or test with a tiny prompt
return true; // stub
}Troubleshooting
If you get an error about the model name, check that the model IDs in routing.json match the ones in your OpenAI account.
If the cost is always zero, check that the PRICING object has the correct model names and that the usage object is returned by the SDK.
If you see a rate limit error, add a small delay between requests or use a retry with exponential backoff.
- Make sure your OPENAI_API_KEY is set in a .env file.
- If you change routing.json, restart the Node process.
- For high traffic, load the config once and cache it.
FAQ
Answers to the questions that come up most often on this topic.
- Q: Can I use this with other providers? A: Yes, replace the OpenAI SDK with any provider and adjust the pricing table.
- Q: How do I decide the maxCost and maxLatency values? A: Start with your budget and user expectations, then adjust based on real usage.
- Q: What if a model is not available? A: The fallback mechanism tries the next model on the list.
- Q: Is model routing worth it for small apps? A: Yes, even a simple router can cut costs by 50% if you have mixed tasks.
Key takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on ai for?
- Working developers who need a practical take on ai model routing explained with a real application — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published September 27, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.