How to Build a Cheap AI Agent Without Making It Dumb
Learn to build a cost-effective AI agent with a minimal loop, tool schema, and guardrails, using open-source models and practical code examples.
The problem: your agent is either expensive or useless
You want to build an AI agent that can actually do things: look up data, call APIs, maybe write files. But when you wire it to a powerful model like GPT-4 or Claude, the token bill climbs fast. Every tool call eats context, and soon you are paying for a conversation that goes nowhere.
So you try a cheap model. It follows instructions poorly, invents tool arguments, and loops on the same mistake. The agent is dumb.
The trick is not to pick one model. It is to design the agent so that a small model can be reliable. You do this by constraining the output, validating every step, and keeping the context small. This article shows a concrete pattern you can copy: a minimal agent loop in TypeScript that uses a local or cheap model, with a tool schema, a validator, and a retry mechanism.
Before you start
You need Node.js 18 or later, and an OpenAI-compatible API endpoint. The examples use the OpenAI SDK, but the pattern works with any provider that exposes a chat completions endpoint, including Ollama, LM Studio, or a local vLLM server.
Set an environment variable for your API key. For a local server, you can set it to a dummy value.
export OPENAI_API_KEY=your-key-here
# For local Ollama, you can set a dummy key:
export OPENAI_BASE_URL=http://localhost:11434/v1- Use a model that supports JSON mode or function calling. For local models, install Ollama and pull a small model like phi3:mini.
- Keep the model context window in mind. Phi-3 mini has 4K context, so your prompts must be tight.
- Install the OpenAI SDK: npm install openai
Step 1: Define a minimal tool schema
An agent is only as good as its tools. For a cheap model, you want tools with simple, flat arguments. Avoid nested objects and enums unless necessary. Here is a schema for two tools: one to get the current time, and one to add two numbers.
const tools = [
{
type: "function",
function: {
name: "get_current_time",
description: "Get the current time in ISO 8601 format",
parameters: {
type: "object",
properties: {},
additionalProperties: false
}
}
},
{
type: "function",
function: {
name: "add_numbers",
description: "Add two numbers together",
parameters: {
type: "object",
properties: {
a: { type: "number" },
b: { type: "number" }
},
required: ["a", "b"],
additionalProperties: false
}
}
}
];- Keep descriptions short. The model has to read them every turn.
- Use additionalProperties: false to force the model to stick to the schema.
- Avoid enums unless you have a strong reason. They confuse small models.
Step 2: Write the agent loop
The core loop is simple: send the conversation plus the tool schema to the model, check if it wants to call a tool, execute the tool, append the result, and repeat. The trick is to limit the number of iterations and to validate the tool call arguments before executing.
import OpenAI from "openai";
const client = new OpenAI();
async function runAgent(userMessage: string) {
const messages: OpenAI.Chat.ChatCompletionMessageParam[] = [
{ role: "system", content: "You are a helpful assistant. Use the provided tools when needed." },
{ role: "user", content: userMessage }
];
const maxIterations = 5;
for (let i = 0; i < maxIterations; i++) {
const response = await client.chat.completions.create({
model: "gpt-4o-mini", // or "phi3:mini" for local
messages,
tools,
tool_choice: "auto",
temperature: 0
});
const message = response.choices[0].message;
messages.push(message);
if (message.tool_calls) {
for (const toolCall of message.tool_calls) {
const result = executeTool(toolCall);
messages.push({
role: "tool",
tool_call_id: toolCall.id,
content: result
});
}
} else {
return message.content;
}
}
throw new Error("Agent exceeded max iterations");
}
function executeTool(toolCall: OpenAI.Chat.Completions.ChatCompletionMessageToolCall): string {
const args = JSON.parse(toolCall.function.arguments);
switch (toolCall.function.name) {
case "get_current_time":
return new Date().toISOString();
case "add_numbers":
return String(args.a + args.b);
default:
throw new Error(`Unknown tool: ${toolCall.function.name}`);
}
}- Set temperature to 0 for deterministic tool calls.
- Always push the assistant message to the conversation before the tool results.
- Limit iterations to avoid infinite loops.
Step 3: Validate tool arguments before execution
When the tool returns an error message, the model sees it and can correct itself. This is called self-correction, and it works surprisingly well with small models if the error message is specific.
You can also use a JSON schema validator like zod to validate arguments before execution. That is a bit more code, but it gives you structured errors.
function executeTool(toolCall: OpenAI.Chat.Completions.ChatCompletionMessageToolCall): string {
let args: any;
try {
args = JSON.parse(toolCall.function.arguments);
} catch (e) {
return "Error: invalid JSON. Please re-call the tool with valid JSON arguments.";
}
switch (toolCall.function.name) {
case "get_current_time":
return new Date().toISOString();
case "add_numbers": {
if (typeof args.a !== "number" || typeof args.b !== "number") {
return "Error: a and b must be numbers. Please re-call with numeric values.";
}
return String(args.a + args.b);
}
default:
return `Error: unknown tool ${toolCall.function.name}`;
}
}Step 4: Keep the context small
The biggest cost driver is the number of tokens you send on every request. If your agent has a long conversation history, the prompt grows and every call becomes more expensive and slower.
For a cheap agent, you should trim the conversation to the last N messages. A good starting point is to keep the system prompt, the last user message, and the most recent tool results. You can also summarize older parts of the conversation, but that adds complexity.
function trimMessages(messages: OpenAI.Chat.ChatCompletionMessageParam[], maxMessages: number) {
if (messages.length <= maxMessages) return messages;
// Always keep the system message
const system = messages.find(m => m.role === "system");
const rest = messages.filter(m => m.role !== "system");
const trimmed = rest.slice(-(maxMessages - 1));
return system ? [system, ...trimmed] : trimmed;
}- Start with a maxMessages of 10. Adjust based on your model's context window.
- Keep the system prompt short and to the point.
- If you need long-term memory, use a vector store, but do not stuff the whole history into the prompt.
Step 5: Add a simple guardrail
A cheap agent can be tricked into doing something it should not. You can add a lightweight check that rejects tool calls to dangerous tools or unexpected arguments. For example, if your agent has a tool that can delete files, you might want to require a confirmation flag.
Here is a simple permission check that only allows a tool called delete_file if the user explicitly asked for it.
function isAllowed(toolName: string, userMessage: string): boolean {
if (toolName === "delete_file") {
return /delete|remove/i.test(userMessage);
}
return true;
}
// In the loop, before executing:
if (!isAllowed(toolCall.function.name, userMessage)) {
messages.push({
role: "tool",
tool_call_id: toolCall.id,
content: "Error: this action is not allowed unless you explicitly ask for it."
});
continue;
}- Do not rely on the model to decide permissions. Enforce them in code.
- Log every tool call for audit.
- You can extend this to a role-based system later.
Verify it worked
Run the agent with a few test prompts. Here is a quick smoke test you can paste into a script.
node -e "
const { runAgent } = require('./agent');
runAgent('What time is it?')
.then(console.log)
.catch(console.error);
"- Expect the model to call get_current_time and return a time string.
- Try a prompt that requires two tool calls, like 'Add 5 and 3, then tell me the time.'
- Check that the agent does not loop more than 5 times.
What I would do: recommended setup
For a production-grade but still cheap agent, I would use a local model for development and a small hosted model for production. Here is a copy-paste starter configuration for Ollama with phi3:mini, which you can run on a laptop.
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Pull a small model
ollama pull phi3:mini
# Run the agent with the local model
export OPENAI_BASE_URL=http://localhost:11434/v1
export OPENAI_API_KEY=ollama
node agent.js- Use phi3:mini for development and gpt-4o-mini for production if you need more reliability.
- Set a timeout on every API call to avoid hanging.
- Monitor token usage with a simple counter.
Troubleshooting
If the agent is not calling tools, try the following.
- Check that the model supports function calling. Some small models do not.
- Simplify your tool descriptions. The model may be ignoring tools it does not understand.
- Lower the temperature to 0.
- If the model returns malformed JSON, add a retry loop that sends the error back to the model.
FAQ
Answers to the questions that come up most often on this topic.
- Q: What is the cheapest model that can do function calling? A: For hosted, gpt-4o-mini is a good balance. For local, phi3:mini or llama3.2:3b work if you keep the schema simple.
- Q: Why does my agent loop forever? A: Usually because it does not understand the tool result. Make sure the tool result is short and clear, and always include an error message on failure.
- Q: Can I use this pattern with Python? A: Yes, the same logic applies. Use the openai python package and the same tool schema.
- Q: How do I reduce costs further? A: Cache repeated tool results, trim context aggressively, and use a smaller model for simple tasks.
Next step
Create a file called agent.js with the code from this article, run the smoke test, and then extend it with your own tools. The pattern is solid, and you can now build a cheap agent that is not dumb.
Key takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on ai-agents for?
- Working developers who need a practical take on how to build a cheap ai agent without making it dumb — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published September 28, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.