Why Your AI Agent Forgets Everything Between Tasks
Learn why AI agents lose context between tasks and how to fix it with memory, state, and observability patterns.
The problem: your agent starts from zero every time
You built an AI agent that can answer questions, run tools, and even write code. But every time you send it a new task, it acts like it has never met you. It asks for the same context, re-reads the same files, and repeats the same mistakes. This is not a bug in your prompt. It is a fundamental design issue: most agents are stateless by default. Each call to the model is independent, and the model has no memory of previous interactions unless you explicitly carry that context forward.
In this article, I will show you why agents lose context, how to implement memory at three levels (conversation, task, and long-term), and how to add observability so you can see what your agent actually remembers. You will walk away with a minimal agent loop that persists state and a checklist to audit your own agent.
Before you start
You will need a working Node.js environment (v18 or later) and an OpenAI API key (or any compatible API). We will use the official OpenAI SDK, but the patterns apply to any LLM provider.
Create a new project directory and install the dependencies:
mkdir agent-memory-demo
cd agent-memory-demo
npm init -y
npm install openai dotenv- Make sure you have Node 18+ installed (check with node -v).
- Create a .env file and add your OPENAI_API_KEY.
- We will use the gpt-4o-mini model for cost efficiency, but you can swap it out.
Step 1: The naive agent loop
Run it with node naive-agent.js. The second question will likely return something like 'I do not know your name.' because the model has no memory of the first exchange. This is the core problem.
import OpenAI from 'openai';
import 'dotenv/config';
const openai = new OpenAI();
async function runAgent(userInput) {
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
messages: [{ role: 'user', content: userInput }],
});
return response.choices[0].message.content;
}
// Test: ask two related questions
console.log(await runAgent('My name is Alice.'));
console.log(await runAgent('What is my name?'));Step 2: Add conversation memory
Now the second call will correctly return 'Alice' because the full history is included. However, this only works within a single process. If the agent restarts, or if you have multiple workers, the history is lost. Also, the history grows unboundedly, which will eventually hit token limits and increase cost.
import OpenAI from 'openai';
import 'dotenv/config';
const openai = new OpenAI();
const messages = [];
async function runAgent(userInput) {
messages.push({ role: 'user', content: userInput });
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
messages,
});
const reply = response.choices[0].message.content;
messages.push({ role: 'assistant', content: reply });
return reply;
}
console.log(await runAgent('My name is Alice.'));
console.log(await runAgent('What is my name?'));- Store the messages array in a database or a file to survive restarts.
- Consider using a sliding window: keep the last N messages, or summarize older ones.
- Be aware of token limits: a long conversation will exceed the model context window.
Step 3: Persist state to a file
To make memory survive across process restarts, you can persist the conversation to disk. Here is a simple implementation using a JSON file. In a real system you would use a database like SQLite or Redis, but the pattern is the same.
Create a file called persistent-memory.js:
import OpenAI from 'openai';
import 'dotenv/config';
import fs from 'fs';
const openai = new OpenAI();
const MEMORY_FILE = './memory.json';
function loadMemory() {
if (fs.existsSync(MEMORY_FILE)) {
return JSON.parse(fs.readFileSync(MEMORY_FILE, 'utf8'));
}
return [];
}
function saveMemory(messages) {
fs.writeFileSync(MEMORY_FILE, JSON.stringify(messages));
}
async function runAgent(userInput) {
const messages = loadMemory();
messages.push({ role: 'user', content: userInput });
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
messages,
});
const reply = response.choices[0].message.content;
messages.push({ role: 'assistant', content: reply });
saveMemory(messages);
return reply;
}
console.log(await runAgent('My name is Alice.'));
// Restart the process, then:
console.log(await runAgent('What is my name?'));- Run the script twice: the second run should still remember the name.
- This is a minimal example; in production use a real database with atomic writes.
- Consider adding a user ID to separate conversations for different users.
Step 4: Add task-level memory with a summarizer
Conversation history is not enough for long-running agents that perform multiple tasks. You need to maintain a task state: what has been done, what remains, and any decisions made. A common pattern is to keep a summary that is updated after each step.
Here is an agent loop that maintains a running summary of the task state. It uses a second LLM call to summarize the conversation after each turn, and then uses that summary in the next request.
import OpenAI from 'openai';
import 'dotenv/config';
const openai = new OpenAI();
let taskSummary = '';
async function summarize(messages) {
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
messages: [
{ role: 'system', content: 'Summarize the conversation into a concise task state. Include key facts, decisions, and what remains to be done.' },
...messages,
],
});
return response.choices[0].message.content;
}
async function runAgent(userInput) {
const messages = [
{ role: 'system', content: `Current task state: ${taskSummary || 'No state yet.'}` },
{ role: 'user', content: userInput },
];
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
messages,
});
const reply = response.choices[0].message.content;
taskSummary = await summarize([...messages, { role: 'assistant', content: reply }]);
return reply;
}
// Example usage
console.log(await runAgent('We need to write a Python script that reads a CSV and prints the sum of a column.'));
console.log(await runAgent('What is the current plan?'));- The summary is updated after each turn, so the agent always has a compact representation of the task state.
- This reduces token usage compared to sending the full history every time.
- In a real system, persist the summary to a database along with a task ID.
Step 5: Long-term memory with a vector store
In a real agent, you would store these memories with metadata like user ID and timestamp, and inject the retrieved memories into the system prompt before each call.
import OpenAI from 'openai';
import 'dotenv/config';
const openai = new OpenAI();
const memoryStore = []; // [{ text, embedding }]
async function embed(text) {
const response = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: text,
});
return response.data[0].embedding;
}
function cosineSimilarity(a, b) {
const dot = a.reduce((sum, val, i) => sum + val * b[i], 0);
const normA = Math.sqrt(a.reduce((sum, val) => sum + val * val, 0));
const normB = Math.sqrt(b.reduce((sum, val) => sum + val * val, 0));
return dot / (normA * normB);
}
async function remember(text) {
const embedding = await embed(text);
memoryStore.push({ text, embedding });
}
async function recall(query, topK = 3) {
const queryEmbedding = await embed(query);
return memoryStore
.map((item) => ({ text: item.text, score: cosineSimilarity(queryEmbedding, item.embedding) }))
.sort((a, b) => b.score - a.score)
.slice(0, topK)
.map((item) => item.text);
}
// Example
await remember('User prefers Python for data tasks.');
const relevant = await recall('What language does the user prefer?');
console.log(relevant);Step 6: Observability - see what the agent remembers
You cannot fix what you cannot see. Add logging to your agent to track the exact messages sent to the model, the context used, and the state after each step. This will help you debug why the agent forgets something.
Here is a simple logging wrapper that prints the conversation history and the task summary before each API call.
import OpenAI from 'openai';
import 'dotenv/config';
const openai = new OpenAI();
function logContext(messages, summary) {
console.log('--- Context sent to model ---');
console.log('Summary:', summary);
messages.forEach((m, i) => console.log(`[${i}] ${m.role}: ${m.content.slice(0, 100)}`));
}
async function runAgent(userInput, summary) {
const messages = [
{ role: 'system', content: `Current task state: ${summary}` },
{ role: 'user', content: userInput },
];
logContext(messages, summary);
const response = await openai.chat.completions.create({
model: 'gpt-4o-mini',
messages,
});
return response.choices[0].message.content;
}- Log the full messages array, not just the latest user input.
- Log token usage (available in the API response) to monitor cost.
- In production, use structured logging (JSON) and a tracing tool like LangSmith or OpenTelemetry.
Recommended setup for production agents
This setup gives you a stateless API server that can handle multiple users and tasks without losing context.
-- Schema example for SQLite
CREATE TABLE conversations (
id INTEGER PRIMARY KEY,
user_id TEXT NOT NULL,
role TEXT NOT NULL,
content TEXT NOT NULL,
created_at DATETIME DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE tasks (
task_id TEXT PRIMARY KEY,
user_id TEXT NOT NULL,
summary TEXT NOT NULL,
status TEXT DEFAULT 'in_progress',
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP
);- Use a conversation table with columns: id, user_id, role, content, created_at.
- Use a task table with columns: task_id, user_id, summary, status, updated_at.
- Use a vector store (like pgvector) for long-term memories, keyed by user_id.
- Before each agent step, load the last N messages, the task summary, and top-K relevant memories.
- After each step, save the new messages, update the summary, and store any new facts as embeddings.
Troubleshooting: common causes of forgotten context
Run the observability logging from Step 6 to see exactly what the model receives. If the context is missing, trace where it got dropped.
- You are not including the system prompt with the task state in every request.
- You are trimming the conversation history too aggressively, cutting off important early details.
- Your summarizer is losing key facts because the prompt is too vague.
- You are using a vector store but not injecting the retrieved memories into the prompt.
- You are not persisting state across process restarts or load-balanced instances.
- You are hitting token limits and the model truncates the context silently.
FAQ
Answers to the questions that come up most often on this topic.
- Q: Why does my agent forget if I use a chat completion API? A: The API is stateless by design. You must pass the full history each time, or use a memory system.
- Q: How much history should I keep? A: It depends on your model context window. Use a sliding window of the last 20 messages, or summarize older ones.
- Q: Is a vector store necessary? A: Only if you need long-term recall across sessions. For single-session tasks, conversation memory and task summaries are enough.
- Q: Can I use a database like Redis for memory? A: Yes, Redis is great for shared state and fast access, but for durable storage use a SQL database.
- Q: How do I handle multiple users? A: Always include a user_id in your memory storage and filter by it when loading context.
Next step: run the demo
Now that you understand the patterns, implement them in your agent. Start by adding conversation memory to your existing loop, then persist it, then add a summarizer. The final test: run your agent, restart the process, and ask it something from the previous session. If it remembers, you have solved the problem.
Run the persistent-memory.js example and verify it works across restarts:
node persistent-memory.js
# Run it again after a restart
node persistent-memory.jsKey takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on ai-agents for?
- Working developers who need a practical take on why your ai agent forgets everything between tasks — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published September 25, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.