Versioned prompts in TypeScript: one typed registry instead of scattered strings
Prompts scattered across routes make quality drops a mystery. A typed, versioned registry makes every prompt edit reviewable, testable and cache-safe.
The first time a user told me "the assistant got worse this week," I had no way to answer them. The system prompt was a string literal inside a route handler, it had been edited three times that month by two people, and nothing in the logs said which wording produced which answer. Versioned prompts in TypeScript fix that for the price of one file: every prompt lives in a typed registry, carries an id and a version number, and goes through code review like any other change. When quality drops, you can at least name the suspect.
Versioned prompts as one typed registry
Put every system prompt in one module, keyed by name, with the version next to the text it describes:
export const PROMPTS = {
chatAssistant: {
id: "chat-assistant",
version: 2,
description: "General assistant behind the demo chat endpoint.",
system: () =>
[
"You are the AI assistant inside a SaaS product built on the Shipwright starter kit.",
"Answer the user's questions directly and concisely.",
"If you don't know something, say so plainly instead of guessing.",
].join("\n"),
},
supportBot: {
id: "support-bot",
version: 1,
description: "Customer-support prompt exercised by the eval suite.",
system: (vars: { product: string }) =>
[
`You are the customer support assistant for ${vars.product}.`,
"- Refunds are available within 30 days of purchase.",
// ...policies
].join("\n"),
},
} as const;
export type PromptKey = keyof typeof PROMPTS;
Three details carry most of the value:
systemis a function, not a string. A prompt that needs variables declares them in its signature, soPROMPTS.supportBot.system()without aproductis a compile error instead of a prompt that says "support assistant for undefined" in production.as constkeeps the literals.PromptKeybecomes the union"chatAssistant" | "supportBot", so anything that takes a prompt name rejects typos at build time.- Lines joined from an array. Diffs on a prompt edit show exactly which sentence changed, one line per instruction, which is what makes review of a prompt change practical.
Routes import the prompt instead of owning it. The chat route in this kit is one line of prompt plumbing:
system: [
{
type: "text",
text: PROMPTS.chatAssistant.system(),
cache_control: { type: "ephemeral" },
},
],
What a version bump is for
The version number is not decoration. It is the answer to "which prompt produced this output," and it only works if you are strict about when it moves. My rule: bump on any edit that could change behavior (a new instruction, a reworded policy, a removed line), and do not bump for a fixed typo in a comment. When in doubt, bump; numbers are free.
Then stamp the id and version anywhere you will look later. In this kit the version lives on the registry entry and nowhere else yet, so the wiring is yours to add, and it is small. Print it in the eval header so a red run tells you what it ran against:
const p = PROMPTS.supportBot;
console.log(`Eval suite: ${p.id}@v${p.version}`);
const ok = await runSuite(suite);
process.exit(ok ? 0 : 1);
The same pair belongs on your usage rows or request logs. The kit already records the model on every usage event, which matters because a quality change has two common causes: you edited the prompt, or the model under it changed. With both stamped, a regression report narrows to one of them in a single query instead of an afternoon of guessing. If you have not wired evals yet, running them as a CI gate is the natural next step: the registry gives the suite one real prompt to test instead of a copy that drifts.
Deterministic prompts are a caching prerequisite
The registry has a second job that has nothing to do with quality: money. Prompt caching on Claude is a byte-exact prefix match. If the system prompt is identical across requests, repeat reads of that prefix bill at a fraction of the normal input rate; if one byte differs, the cache misses and you pay full price for everything after the difference.
The most common way to break that is innocent-looking: a "Current date: ..." line at the top of the system prompt, or a request ID for debugging. Every request now has a unique prefix and the cache never hits. Keeping prompts as pure functions in one file makes this rule enforceable in review: a new Date() inside a registry entry is easy to spot, while the same line buried in a route handler is not. Per-user or per-request data belongs in the messages, after the cached prefix. The prompt caching cost math post walks through what a hit is actually worth.
Things that bite
- Variables in the system prompt are cache-breakers too.
supportBottakesproduct, which is fine because it is the same for every request to that deployment. A variable that changes per user (their name, their plan) splits the cache into one entry per value. Put it in the conversation instead. - A version number without history is just a number. You do not need a changelog file: keep one prompt per registry entry and
git log -pon the file is the changelog, with authors and dates. - Do not move prompts to a CMS to "let product edit them." You lose the typecheck, the code review, and the eval gate in one move, which are the three reasons the registry exists. If non-engineers need to propose wording, they can open a pull request against one file.
- Comments drift like any other docs. A header comment saying the version gets logged somewhere is a claim, not wiring. Grep for the call before you trust it in a postmortem.
Shipwright, the Next.js and Claude starter kit this blog documents, ships this registry alongside the eval runner, prompt caching on the chat route, and per-request cost metering, so the prompt layer starts out reviewable instead of scattered. You can see the chatAssistant prompt answering in the live demo.