How to Make AI Explain a Codebase Without Uploading the Whole Repository
Learn practical techniques to give AI coding assistants context about your codebase without exposing or uploading the entire repository, using local tools and selective file sharing.
Before you start
You are working on a large repository and want AI to explain how a feature is implemented. You have tried pasting a few files, but the AI lacks context and gives generic answers. Uploading the whole repo is not an option because of size limits or confidentiality.
This article shows you how to build a local context bundle that you can share with any AI chat or coding assistant. You will use command-line tools to extract structure, summaries, and key files without exposing secrets or exceeding token limits.
You will need a Unix-like shell (Linux, macOS, or WSL on Windows), tree, git, and optionally jq. All commands are run from the root of your repository.
tree --version
git --version
jq --version- Check that tree is installed: run tree --version
- Check that git is installed: run git --version
- Optional: install jq with your package manager (sudo apt install jq on Debian/Ubuntu)
Step 1: Generate a repository map
The first thing an AI needs is a high-level view of the project structure. Instead of uploading every file, create a tree of the repository that excludes generated folders and dependency directories.
The tree command can output a plain-text representation of the directory structure. Use the -I flag to ignore common noise like node_modules, .git, and build outputs.
tree -I 'node_modules|.git|dist|build|__pycache__|.venv|venv' -L 3 > tree.txt- Adjust the -L flag to control depth. A depth of 3 is usually enough for a first pass.
- Review tree.txt to make sure sensitive files are not listed. If they are, add them to the ignore pattern.
- Keep this file small. It should be a few hundred lines at most.
Step 2: Extract key file contents with git
Rather than copying files manually, use git to list the files that are actually tracked. This avoids including untracked files that might contain secrets or temporary changes.
You can then select the most relevant files based on your question. For example, if you want to understand the authentication flow, you would look for files with auth in their names or in an auth directory.
git ls-files | grep -E '(auth|login|session)' | head -20- Use git ls-files to get a list of all tracked files.
- Pipe to grep to filter by keywords related to your question.
- Limit the output to a manageable number of files, for example with head -20.
Step 3: Bundle selected files into a single context file
Now you have a list of relevant files. You need to combine their contents into a single text file that you can paste into an AI chat. This file should include the file paths as headers so the AI knows which file each block belongs to.
The following script loops over a list of files, adds a header, and appends the content to context.txt. It skips binary files and files larger than a few hundred kilobytes to keep the context under token limits.
#!/bin/bash
# files.txt contains one file path per line
while IFS= read -r file; do
if [ -f "$file" ] && [ $(wc -c < "$file") -lt 200000 ]; then
echo "===== $file =====" >> context.txt
cat "$file" >> context.txt
echo "" >> context.txt
fi
done < files.txtgit ls-files | grep -E '(auth|login|session)' > files.txt
bash bundle.sh
wc -l context.txt- The script skips files larger than 200 KB to avoid exceeding token limits.
- You can adjust the size limit based on your AI tool's context window.
- Always review context.txt before sharing it to ensure no sensitive data is included.
Step 4: Ask the AI with a focused prompt
With the context file ready, you can now ask the AI to explain the relevant code. The quality of the answer depends heavily on how you frame the question. Provide the context and ask for a specific explanation.
Here is an example prompt that works well. It asks the AI to trace the login flow and explain the key components, without requiring the entire repository.
I have a repository with the following structure and relevant files. Please explain how the login flow works, referencing specific files and functions. Focus on the authentication logic and how it integrates with the rest of the system.
[Paste the contents of tree.txt and context.txt here]- Be specific about what you want the AI to explain: a flow, a function, or an architecture decision.
- Include the tree and context in the same message so the AI has both structure and code.
- If the AI asks for more files, use the same git grep approach to extract them and add to the context.
Step 5: Automate the process with a script
If you do this often, create a reusable script that generates the context for any directory. This script takes a keyword and a depth as arguments, and produces a context file ready for pasting.
The script combines the previous steps: it generates the tree, finds relevant files, and bundles them. You can save it as make-context.sh and run it whenever needed.
#!/bin/bash
# usage: ./make-context.sh <keyword> [depth]
KEYWORD=${1:?Usage: $0 keyword [depth]}
DEPTH=${2:-3}
TREE_FILE="tree.txt"
CONTEXT_FILE="context.txt"
# Generate tree
tree -I 'node_modules|.git|dist|build|__pycache__|.venv|venv' -L $DEPTH > $TREE_FILE
# Find relevant files
git ls-files | grep -E "$KEYWORD" > files.txt
# Bundle
echo "===== TREE =====" > $CONTEXT_FILE
cat $TREE_FILE >> $CONTEXT_FILE
echo "" >> $CONTEXT_FILE
while IFS= read -r file; do
if [ -f "$file" ] && [ $(wc -c < "$file") -lt 200000 ]; then
echo "===== $file =====" >> $CONTEXT_FILE
cat "$file" >> $CONTEXT_FILE
echo "" >> $CONTEXT_FILE
fi
done < files.txt
echo "Context written to $CONTEXT_FILE"
wc -l $CONTEXT_FILE- Run chmod +x make-context.sh to make it executable.
- Use ./make-context.sh auth to generate context for authentication-related files.
- The script prints the line count so you can estimate token usage.
Verify it worked
After running the script, you should have a context.txt file that contains a tree and the contents of relevant files. Check that the file does not include any secrets or files you did not intend to share.
Paste the contents into your AI assistant and ask a test question. For example, ask: What does the login function do? The answer should reference the specific files and functions you included.
- Open context.txt and scan for any sensitive information like API keys or passwords.
- Check that the file list in files.txt matches your expectations.
- If the AI gives a vague answer, you may need to include more files or refine your prompt.
Troubleshooting
If tree is not installed, you can use find as a fallback. The find command can produce a similar listing, though it is less readable.
If git ls-files returns nothing, make sure you are in a git repository and that files are tracked. Use git status to check.
If the context file is too large for your AI tool, reduce the depth of the tree or filter files more aggressively.
find . -type f -not -path './node_modules/*' -not -path './.git/*' | head -50- Use find with -not -path to exclude directories.
- Consider using ripgrep (rg) for faster search: rg --files | grep keyword
- If you have many files, sort by size and include only the largest or most recently modified.
Recommended setup
For a consistent workflow, I recommend keeping the make-context.sh script in a central location, such as ~/bin, and adding it to your PATH. This way you can generate context from any repository without copying the script.
Here is a copy-paste starter for the script. Save it as ~/bin/make-context, make it executable, and use it as shown.
cat > ~/bin/make-context << 'EOF'
#!/bin/bash
KEYWORD=${1:?Usage: make-context keyword [depth]}
DEPTH=${2:-3}
TREE_FILE="tree.txt"
CONTEXT_FILE="context.txt"
tree -I 'node_modules|.git|dist|build|__pycache__|.venv|venv' -L $DEPTH > $TREE_FILE
git ls-files | grep -E "$KEYWORD" > files.txt
echo "===== TREE =====" > $CONTEXT_FILE
cat $TREE_FILE >> $CONTEXT_FILE
echo "" >> $CONTEXT_FILE
while IFS= read -r file; do
if [ -f "$file" ] && [ $(wc -c < "$file") -lt 200000 ]; then
echo "===== $file =====" >> $CONTEXT_FILE
cat "$file" >> $CONTEXT_FILE
echo "" >> $CONTEXT_FILE
fi
done < files.txt
echo "Context written to $CONTEXT_FILE"
wc -l $CONTEXT_FILE
EOF
chmod +x ~/bin/make-context- Add export PATH="$HOME/bin:$PATH" to your .bashrc or .zshrc.
- Run make-context auth 4 to get a deeper tree.
- Always review the generated context before sharing it with an external AI service.
FAQ
Here are answers to common questions about this approach.
- Q: Can I use this with any AI tool? A: Yes, the context file is plain text, so you can paste it into any chat interface or use it with local models.
- Q: What if my repository has many files? A: Use a more specific keyword or increase the depth of the tree. You can also filter by file type, for example grep -E '\.(ts|js)$'.
- Q: How do I avoid leaking secrets? A: Review the context file manually, and use git ls-files to ensure only tracked files are included. You can also add a pre-scan for common secret patterns.
- Q: What is the token limit? A: It depends on the AI model. As a rule of thumb, keep the context under 50,000 characters. The script skips files larger than 200 KB to help.
- Q: Can I automate this in CI? A: Yes, you can run the script in a CI pipeline and upload the context as an artifact, but be cautious about storing sensitive code in logs.
Next steps
You now have a repeatable way to give AI the context it needs without exposing your whole repository. The next time you need an explanation, run make-context with a keyword, review the output, and paste it into your assistant.
Start with a real question about your codebase. Run the script, inspect the generated context, and ask the AI to explain a specific function. You will likely get a much more accurate answer than pasting a single file.
make-context authKey takeaways
- Apply one concrete change from this post before collecting more reading.
- Prefer browser-side tools when the work involves secrets, tokens, or PII.
- Document the why next to the how so the next reviewer inherits context.
FAQ
- Who is this guide on ai-coding for?
- Working developers who need a practical take on how to make ai explain a codebase without uploading the whole repository — not a marketing overview. Skim the sections, apply one tip, then come back when you hit an edge case.
- Do I need an account to use the related tools?
- No. code.live tools run in your browser with no signup. Nothing you paste is uploaded to a server for the client-side utilities linked from this post.
- How often is this article updated?
- This post was published October 3, 2026. Fundamentals stay stable; check linked tool pages and official docs when version-specific behavior matters.