My 2026 AI Developer Stack
A hands-on guide to building a practical AI developer stack for 2026: models, agents, orchestration, and observability with copy-paste code.
The problem: AI coding tools are multiplying faster than you can evaluate them
Every week there is a new model, a new agent framework, or a new MCP server. You have a production codebase, a CI pipeline, and a budget. You need a stack that actually ships code, not one that demos well.
This article is my 2026 setup: the models, the orchestration layer, the agent loop, and the observability. Everything here is reproducible. You can copy the configs, run the commands, and adapt them to your own projects.
I focus on practical choices: what I run locally, what I call via API, and how I keep costs under control. The stack is built around open standards and avoids lock-in where possible.
Before you start: what you need
You need a machine with Node.js 20+ and Python 3.11+. You also need API keys for at least one LLM provider. I use OpenAI and Anthropic, but the patterns work with any provider that exposes a chat completions endpoint.
Create a project directory and set up a virtual environment. All the code in this article is tested against the versions listed in the requirements file.
mkdir ai-stack-2026 && cd ai-stack-2026
python3 -m venv .venv && source .venv/bin/activate
pip install openai anthropic pydantic rich requests
npm init -y
npm install @modelcontextprotocol/sdk typescript tsxStep 1: Pick your models for different jobs
Not every task needs the biggest model. I use a three-tier strategy: a small model for linting and formatting, a medium model for refactoring and boilerplate, and a large model for architecture and complex debugging.
As of early 2026, my choices are: GPT-5-mini for fast tasks, Claude Sonnet 4.5 for everyday coding, and Claude Opus 4.5 for the hard stuff. Check the latest pricing and capabilities before you commit, because this changes quickly.
Here is a routing function that sends prompts to the right model based on complexity. Save it as router.py.
import os
from openai import OpenAI
from anthropic import Anthropic
openai_client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
anthropic_client = Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
def route(prompt: str, task: str = "general"):
if task == "quick":
return openai_client.chat.completions.create(
model="gpt-5-mini",
messages=[{"role": "user", "content": prompt}]
).choices[0].message.content
elif task == "complex":
return anthropic_client.messages.create(
model="claude-opus-4-5",
max_tokens=4096,
messages=[{"role": "user", "content": prompt}]
).content[0].text
else:
return anthropic_client.messages.create(
model="claude-sonnet-4-5",
max_tokens=2048,
messages=[{"role": "user", "content": prompt}]
).content[0].text- Use the small model for tasks that are deterministic and fast, like generating a regex or formatting a JSON snippet.
- Use the medium model for code review, refactoring, and writing tests.
- Use the large model for debugging elusive bugs, designing system architecture, or generating complex migrations.
Step 2: Set up a local agent with tool calling
A single model call is not enough. You want an agent that can read files, run tests, and iterate. I use a minimal agent loop with the Anthropic API because it supports tool use natively.
Here is a minimal agent that can run shell commands and read files. Save it as agent.py.
import os
import subprocess
from anthropic import Anthropic
client = Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
tools = [
{
"name": "run_command",
"description": "Run a shell command and return the output",
"input_schema": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]
}
},
{
"name": "read_file",
"description": "Read a file from the current directory",
"input_schema": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"]
}
}
]
def run_agent(prompt: str):
messages = [{"role": "user", "content": prompt}]
while True:
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=4096,
tools=tools,
messages=messages
)
if response.stop_reason != "tool_use":
print(response.content[0].text)
break
tool_use = next(block for block in response.content if block.type == "tool_use")
if tool_use.name == "run_command":
result = subprocess.run(tool_use.input["command"], shell=True, capture_output=True, text=True)
output = result.stdout + result.stderr
elif tool_use.name == "read_file":
try:
with open(tool_use.input["path"]) as f:
output = f.read()
except Exception as e:
output = str(e)
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": [{"type": "tool_result", "tool_use_id": tool_use.id, "content": output}]})
if __name__ == "__main__":
run_agent("In this repo, find the failing test and fix it. Run the tests to verify.")- The loop continues until the model stops requesting tools.
- You can extend the tools list with any function you want, like a database query or a linter.
- Always cap the number of iterations to avoid infinite loops.
Step 3: Give the agent context with MCP servers
You can then tell your agent to use these tools by name. For example, you can ask it to look up a table schema from the database and then generate a migration.
The MCP SDK allows you to build your own servers too. If you have a internal API, wrap it as an MCP server so your agent can call it directly.
{
"mcpServers": {
"docs": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-filesystem", "./docs"]
},
"db": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-postgres", "postgresql://user:pass@localhost:5432/mydb"]
}
}
}Step 4: Add code generation to your editor
You can also run a local proxy that routes requests based on the prompt length. Here is a simple FastAPI server that does this. Save it as proxy.py.
from fastapi import FastAPI, Request
import httpx
app = FastAPI()
@app.post("/v1/chat/completions")
async def chat(request: Request):
body = await request.json()
prompt = body["messages"][-1]["content"]
if len(prompt) > 2000:
model = "claude-sonnet-4-5"
base = "https://api.anthropic.com/v1/messages"
else:
model = "gpt-5-mini"
base = "https://api.openai.com/v1/chat/completions"
# forward the request with the selected model
# ... (implementation depends on provider API)
return {"model": model, "message": "forwarded"}Step 5: Set up CI with AI code review
The review.py script sends the diff to the model and posts comments. Here is a minimal version.
import os
import sys
from openai import OpenAI
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
def review_diff(diff: str):
prompt = f"You are a senior code reviewer. Review this diff and suggest improvements.\n\n{diff}"
response = client.chat.completions.create(
model="gpt-5-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.2
)
print(response.choices[0].message.content)
if __name__ == "__main__":
diff = sys.argv[1]
review_diff(diff)- Set a max diff size to avoid sending huge files to the API.
- Use the small model for style checks and the medium model for logic issues.
- Only post comments on lines that changed to reduce noise.
Step 6: Observe and evaluate your AI usage
For cost tracking, I use a simple script that reads the log and aggregates by model. You can adapt it to your provider's pricing.
import json
from collections import defaultdict
costs = {
"gpt-5-mini": (0.15, 0.60), # per 1M input, output tokens
"claude-sonnet-4-5": (3.00, 15.00),
"claude-opus-4-5": (15.00, 75.00)
}
def estimate_cost(entry):
model = entry.get("model", "")
tokens = entry.get("tokens", {})
input_tokens = tokens.get("input", 0)
output_tokens = tokens.get("output", 0)
if model in costs:
return input_tokens/1e6 * costs[model][0] + output_tokens/1e6 * costs[model][1]
return 0
# Load log and aggregate
with open("ai_log.jsonl") as f:
total = 0.0
for line in f:
entry = json.loads(line)
total += estimate_cost(entry)
print(f"Total estimated cost: ${total:.2f}")Recommended setup: my full stack in one command
Run docker-compose up and you have a full stack: a routing proxy, an MCP server for docs, and a logger. You can then point your editor and CI to the proxy.
This is the setup I use daily. It is not perfect, but it is reproducible and easy to tweak.
version: '3.8'
services:
proxy:
build: .
ports:
- "8000:8000"
environment:
- OPENAI_API_KEY=${OPENAI_API_KEY}
- ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY}
mcp:
image: node:20
command: npx -y @modelcontextprotocol/server-filesystem /app/docs
volumes:
- ./docs:/app/docs
logger:
image: python:3.11
volumes:
- .:/app
working_dir: /app
command: python logging.pyFAQ
Answers to the questions that come up most often on this topic.
- Q: Do I need to use both OpenAI and Anthropic? A: No, you can stick to one provider. The routing idea still works if you use different model sizes from the same provider.
- Q: How do I keep costs down? A: Use the small model for most tasks, cache responses where possible, and set a monthly budget alert in your provider dashboard.
- Q: Is MCP ready for production? A: It is stable enough for internal tools. For external-facing agents, you need to add authentication and rate limiting.
- Q: Can I use local models instead? A: Yes, if you have the hardware. The same routing code works with Ollama or vLLM endpoints.
- Q: How do I keep my prompts and configs in sync? A: Store them in a shared repo and use environment variables for secrets.
Try it on code.live
If you want to test a model routing or a prompt quickly, use the API Mock Data Generator on code.live to create sample inputs for your agent, or use the JSON Formatter to clean up your config files.
Key takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on ai for?
- Working developers who need a practical take on my 2026 ai developer stack — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published September 29, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.