Resume as code, part two: one file or many?
After I published Resume as code and shared it on LinkedIn, one of the comments made a fair point: why a whole folder of files? A single CV file would be enough as the source of truth.
I’d wondered the same thing myself, and was already thinking about collapsing everything into one structured markdown file, or even YAML. So instead of answering from instinct, I went and measured. This post is the result: what one file versus many really costs when AI tools (Claude, Codex, ChatGPT and friends) are the ones reading and editing it, and what changes when the data behind this approach is much bigger than one person’s work history.
The short version is that for my resume, the commenter is mostly right. The reasons to split are real, but they mostly aren’t about tokens, and they only start to matter once the data grows.
The actual numbers
My entire docs/career/ folder (profile, eight role files, the project bank and
four variants, in two languages) is about 46 KB, or roughly 12,000 tokens.
The biggest file is the profile at 10 KB.
Any current model holds that in its context with room to spare. And splitting only saves tokens if a task reads part of the data. Tailoring a resume to a job description reads everything: the profile, every role, the full project bank. So when an agent does that job, it consumes about the same amount of context either way. One file is even slightly cheaper, since it’s one read instead of ten.
So on raw performance and context usage, splitting doesn’t win at this size.
Where the difference shows up: how the tool reads
The comparison gets more interesting when you look at which AI tool is reading.
Chat tools (ChatGPT, claude.ai, Gemini) favour the single file. You upload one attachment and the model sees all of it. With several files, especially when they’re stored as “project knowledge”, the tool may chunk and retrieve them, so the model sees whichever pieces the retrieval step thought were relevant. For a task that needs everything, like matching a job posting against your whole history, partial retrieval is the worst outcome: a project that would have been the best evidence just never shows up.
Agentic tools (Claude Code, Codex) read either shape fine. The difference appears when they edit:
- A single file is full of repeated text. Every role has a
Tech stack:line and a## Summary [en]heading. Edits made by find-and-replace become ambiguous, and the tool has to work harder to target the right occurrence. - Weaker tools sometimes respond by rewriting the whole file, and in the process quietly reword a role you never asked them to touch.
- With one file per role, a bad edit is contained to that role. The diff is
small,
git blamestays meaningful, and reviewing what the agent changed takes seconds.
Accuracy is a wash at this size. Models do lose some recall in the middle of very long contexts, but 12,000 tokens is nowhere near where that bites.
Prompt caching doesn’t care much about file count either. It works on prefixes, so what matters is that stable content comes first and the parts you edit come later.
So at resume scale, the folder structure costs nothing and doesn’t gain
anything on performance. What it gains is maintenance: small diffs, a private
## Notes scratchpad per role, English and Spanish side by side for each job,
and a small blast radius when an agent makes a mistake.
When the source is much bigger
This is where the two approaches really part ways. Imagine the same idea applied to a consultancy with two hundred people’s career histories, a product catalogue, or an internal knowledge base. The data is bigger, and most tasks only need a slice of it.
Context ceiling and cost. With one file, every call pays for the whole corpus. At a few hundred thousand tokens it either doesn’t fit at all, or every request is expensive and slow.
Selective loading. This is the real reason to split. The pattern that scales is a small index that is always loaded, with detail loaded only when it’s needed. My setup already has a tiny version of it: the variant files and the project tags act as the index, and the role and project text is the detail. Claude Code’s own skills work the same way: the model sees a one-line description of every skill, and reads the full instructions only for the one it’s about to use.
Retrieval quality. Files are natural chunk boundaries. An agent using grep, or a retrieval pipeline, gets one complete, self-contained unit, like a whole role or a whole project, instead of an arbitrary slice that cuts a project in half.
Concurrency. Several people or agents editing one big file means merge conflicts. Separate files rarely collide.
Validation. At scale you want each record checked against a schema, so that one malformed entry fails on its own instead of breaking the parse of everything.
A rough rule of thumb I’d use:
- Under about 20,000–30,000 tokens: pick whatever is easier to maintain. One file is fine.
- Above about 50,000–100,000 tokens, or when most tasks need only a subset: split the data and keep an index the model can always see.
Markdown or YAML?
The other half of my original idea was switching to YAML for real structure. Here’s how the options compare:
- Plain YAML: you get structure and parsing for free. The problem is prose. Role summaries are paragraphs, and long text in YAML means block scalars, quoting colons, and indentation that a model can break while editing. Most of a resume is prose, so that’s a weak spot exactly where the content lives.
- Structured markdown: the format models read and write most naturally. The
price is a custom parser, which
build-resume.pyis. - Markdown with YAML front matter: structured fields like dates, tech and tags go in the front matter, where they can be validated. Prose goes in the body. Astro’s content collections, which this site is built on, use the same pattern.
Writing this post, I realised I’m already mostly on the third option. Each role
file has front matter (heading, start, end, tech) and a markdown body.
The exceptions are projects.md, which keeps every project in one file with
a tags: line under each heading, and the profile. The project bank is the
part of the system most likely to grow, and the part where the tags really
are an index, so it’s the first candidate if I ever split further: one file
per project, front matter for tags.
What “single source of truth” means
I think the comment and my setup are closer than they look. “Single source of
truth” means one canonical place for each fact, not one file. Every
sentence on my resume exists in exactly one place in docs/career/; the
markdown, docx, PDF and this website’s work history are all generated from
it. In that sense, the folder is a single source of truth.
For one person’s resume, one well-structured file would work just as well, and for a chat-based workflow it would even work better. The folder pays off in edit safety today, and in the index-plus-details shape that lets the same idea grow to data that no longer fits in a single context window.