← Writing / March 8, 2026 / 6 min read
CodeContexter: Packing a Whole Codebase Into One LLM-Ready File
How I built a fast, safety-first Rust CLI that walks an entire repository, respects .gitignore, redacts secrets before they leave your machine, and estimates the token budget, all in a single streamed pass.
Every time I wanted an LLM to reason about a whole project, I hit the same wall. The model needs context, but a project is not one file. It is hundreds of files scattered across folders, plus lock files, build artifacts, and a .env I would very much like to never paste into a chat window. Copying files by hand is slow, lossy, and dangerous. So I wrote CodeContexter: a Rust CLI that walks a repository and packs it into a single structured file that a model can actually read.
The whole thing is one binary and one job: take a directory, produce a clean, deduplicated, secret-free document with a file tree and every file’s contents. Here is what turned out to be interesting under the hood.
The walk has to be ignore-aware, not just fast
The naive version of this tool recursively reads every file. That version is useless, because it will happily dump node_modules, target, and a 40MB lock file into your context and blow the budget on noise.
So the walk is built on the ignore crate, the same directory traversal library that powers ripgrep. It gives me .gitignore semantics for free:
let walker = WalkBuilder::new(&root_path) .hidden(false) .git_ignore(true) .follow_links(false) // prevent symlink loops and duplication .overrides(overrides) .build();Two of those flags are deliberate. hidden(false) means I do not skip dotfiles, because config like .eslintrc or .github/workflows is often exactly the context you want. And follow_links(false) closes a nasty failure mode: a symlink that points back up the tree can send a recursive walker into an infinite loop or silently duplicate half the repo. Turning link-following off makes the walk finite and honest.
The one thing .gitignore gets wrong for this use case is secrets. Plenty of repos commit or fail to ignore a stray .pem or id_rsa, and I do not want the walk to depend on the user having a perfect ignore file. So before the walk starts, I layer in hard-coded overrides that force-exclude the dangerous stuff regardless of what .gitignore says:
let hard_coded_excludes = vec![ "!*.env", "!*.env.*", "!*.pem", "!*.key", "!id_rsa", "!id_ed25519", "!*.p12", "!*.pfx",];The ! prefix in an override inverts it to an exclude, so these patterns are the tool’s own opinion about what should never be aggregated, applied on top of the user’s .gitignore and any --exclude globs they pass. Belt and suspenders. The file-level redaction below is the belt.
Redaction is defense in depth, not the only defense
Excluding secret files handles the obvious case. It does not handle the API key someone hardcoded in the middle of a Python module. For that, every file’s contents pass through a sanitizer before they are written out.
The sanitizer is a small set of compiled regexes, initialized once and reused across every file via a OnceLock so I am not recompiling patterns in a hot loop:
static SECRET_PATTERNS: OnceLock<Vec<Regex>> = OnceLock::new();let patterns = SECRET_PATTERNS.get_or_init(|| vec![ Regex::new(r"-----BEGIN [A-Z ]+ PRIVATE KEY-----").unwrap(), Regex::new(r"AKIA[0-9A-Z]{16}").unwrap(), Regex::new(r"(?i)sk-[a-zA-Z0-9]{20,}").unwrap(), Regex::new(r"gh[pousr]-[a-zA-Z0-9]{36}").unwrap(), Regex::new(r#"(?i)(api_key|secret|token|password)\s*[:=]\s*["'][a-zA-Z0-9]{32,}["']"#).unwrap(),]);These target the shapes that leak most: RSA and other PEM private key headers, AWS access keys (the AKIA prefix), OpenAI and Stripe style sk- keys, GitHub tokens across their prefix family (ghp, gho, ghu, ghs, ghr), and the generic api_key = "..." assignment pattern. Anything matching gets replaced with [REDACTED SECRET]. This is pattern matching, not proof, so the tool tells you to review output before sharing it. But it means the common ways a key escapes are covered by default, without the user thinking about it.
Deciding what is even worth including
Before a file’s bytes matter, the tool has to decide whether the file belongs at all. A few cheap filters do most of the work.
Empty files are dropped on a metadata check before any read. Binary files are caught by sampling: I read the file, look at up to the first 8192 bytes, and if any of them is a null byte, I treat it as binary and skip it.
fn is_binary(content: &[u8]) -> bool { let len = std::cmp::min(content.len(), 8192); content[..len].contains(&0)}That heuristic is crude and it is exactly right for this job. Real source code effectively never contains a null byte in its first 8KB, and images, compiled objects, and archives almost always do. Whitespace-only files are dropped after read, since a file that trims to nothing adds zero signal and non-zero tokens.
Large files get a different treatment. Anything over 1MB would dominate the budget, so instead of including or dropping it wholesale, the tool keeps the first 50 and last 50 lines and drops the middle with a marker noting how many lines were omitted. You usually want the imports and the shape of a big generated file, not its ten thousand middle lines.
Token accounting so you know before you paste
Every artifact carries a token estimate, and the tool sums them into the header. The estimate is deliberately simple: characters divided by four.
const CHARS_PER_TOKEN: usize = 4;let token_estimate = content_str.len() / CHARS_PER_TOKEN;This is not a real tokenizer, and it does not try to be. Pulling in a model-specific BPE tokenizer would add a heavy dependency, tie the output to one model’s vocabulary, and slow the whole thing down for a number that is a budgeting hint, not a billing figure. The four-characters-per-token approximation is close enough to tell you at a glance whether the output fits in a context window, which is the only decision the number needs to support.
Parallelism, and streaming instead of buffering
File discovery is sequential because the directory walk is inherently ordered, but reading and processing every file is embarrassingly parallel. That phase runs across all cores with rayon: collected_paths.par_iter() fans the per-file work out, and a progress bar increments as each file lands. On a large repo this is where the wall-clock time is won.
Output is streamed, not assembled in memory. Rather than building one giant string and writing it at the end, the tool writes each artifact directly through a BufWriter as it goes. That keeps memory flat even on huge repositories, since the peak footprint is the artifacts themselves plus a small buffer, not a second full copy of the concatenated output.
The output format is tuned for how models read code
The default output is Markdown, with JSON and XML available. Markdown is the default for a reason: it is what these models were trained on the most, and fenced code blocks with a language hint (rust, python) are the strongest signal you can give a model about where one file ends and the next begins.
The document leads with a header line that states the file count and total token estimate, then a text fenced project tree so the model sees the structure before the contents, then each file as its own section with a metadata line (language, line count, token estimate, and a truncation flag when relevant) above its fenced body. JSON and XML exist for programmatic consumers, and the XML path carefully escapes &, <, >, quotes, and apostrophes so file contents cannot break the document.
Why Rust was the right call
This tool touches the filesystem hard, needs to be safe by default, and wants to fan work across cores. Rust gives me all three without compromise: the ignore crate for correct traversal, rayon for parallelism that is a one-line change, and a compiler that will not let me leak a buffer or race a shared regex table. The result starts instantly, holds flat memory on repos of any size, and finishes fast enough that the token count is printed before you have finished reading the command you typed.
It is MIT licensed and it is the open-source project I reach for most, because the alternative is pasting files into a chat one at a time and hoping I did not include the wrong one.
Say you want to ask an AI a question about a piece of software. Not a tiny script, but a real project: a website, an app, a tool. To answer well, the AI needs to actually see the project. And here is the catch that trips everyone up: a software project is not one file. It is thousands of files, scattered across folders inside folders, the way a house has things in every drawer, closet, and cupboard.
Handing all of that to an AI by hand is miserable. You would open files one at a time, copying, pasting, and hoping you did not miss the one that explained everything. So I built a small tool called CodeContexter that does the gathering for you.
Packing a house into one labeled box
The best way to picture it is moving day. Imagine you have to pack an entire house into a single box for a mover, and you want that box so well organized that the mover understands the whole house just by looking inside.
That is the job. CodeContexter walks through every room of your project, picks up everything worth keeping, and lays it out in one neat document. At the top it even draws a little map of the house, a simple outline of which folders hold what, so whoever opens the box sees the layout before digging into the contents.
Leaving out the junk
Not everything in a house is worth packing. Real projects are full of clutter that machines generate automatically: giant stockpiles of downloaded parts, leftover build scraps, empty files, and files that are not really text at all, like images that would look like nonsense.
CodeContexter knows to skip all of that. It reads the same “do not pack this” list that programmers already keep for their own tools, and adds its own good judgment on top. Blank files get left behind. Files that are not readable text get left behind. The result is a box with the useful things in it and none of the packing peanuts.
There is one more clever bit. Some files are simply enormous, like a piece of furniture that will not fit through the door. Instead of dropping those entirely, the tool keeps the beginning and the end and leaves a note about how much of the middle it set aside. You still get the shape of the thing without letting it crowd out everything else.
Never packing your passport
This is the part I care about most. Hidden inside almost every software project are secrets: passwords and digital keys, the equivalent of the spare key to your front door. You absolutely do not want those handed to an outside AI service by accident.
So CodeContexter is careful in two ways, like a friend helping you pack who keeps an eye out for your passport. First, it knows the usual hiding spots for secret files and refuses to pack them at all, even if you forgot to tell it not to. Second, and this is the important one, it reads through the contents of your files looking for things shaped like a password or a key, and blacks them out before anything leaves your computer. Where a secret used to be, the document just says it was removed.
It is not magic, and I say so plainly: always glance at the box before you ship it. But the common ways a secret slips out are covered without you having to think about it.
Telling you if it fits
AI systems can only read so much at once. There is a limit, like a mover who can carry exactly one box and not an ounce more. Pack too much and the top of the pile simply falls off and gets ignored.
So when CodeContexter finishes, it tells you roughly how big the box is in terms the AI cares about. That single number lets you know at a glance whether your whole project will fit, or whether you need to trim it down first. No surprises after you have already handed it over.
Fast, and built to stay out of your way
I wrote this in a programming language chosen for speed and safety, which is a fancy way of saying two things. It finishes almost instantly, even on a big project, so using it never feels like a chore. And it stays light on your computer no matter how large the project is, because it writes the box out steadily as it goes rather than holding the whole thing at once.
That is really the whole idea. Asking an AI about your software should be as simple as asking about anything else. You should not have to become a professional packer first, and you certainly should not worry about accidentally mailing off your keys. CodeContexter packs the box, labels it, leaves out the junk, hides the valuables, and tells you if it fits. Then you can just ask your question.