A Living Glossary for Agent-Driven Development

Get Site as Markdown

The more I build with agentic coding environments, the more precision matters — even though, or rather exactly because, I work in natural language. Every sentence I write can turn into a source-code change, so vague wording turns into vague diffs.

Describing things exactly is harder than it looks. And a glossary I maintain by hand rots within weeks: I rename something in code and forget the docs, or the docs describe behavior from three versions ago.

How nice could it be if to just type jl:domain.output.jxlinfo-sidecar and a feature/view/visuell element could be mentiond without any contradiction.

So I built a small system that is easy to maintain because it mostly maintains itself: keywords are defined once, directly at the implementation in source comments, referenced everywhere else with a fixed syntax, and a generator extracts them into one always up-to-date glossary file. The result feels like a living glossary — I can name, describe, and reference every concept exactly, together with the AI. Dot notation gives me levels, a project-owned prefix keeps keys stable and even addressable across projects. It goes as far as invisible references inside Markdown docs and clickable key links in GitHub issues.

I’m sure that a lot of people have already built different approaches to living glossaries for agent-driven development. This is just my approach.

I am still experimenting with this; if you built something similar — or better — I want to hear about it.

Living example

The system runs in production on its own first subject — the glossary below is generated, not hand-written:

What to look at: the <!-- generated - do not edit --> header, the facet chapters (jl:domain, jl:view, jl:tech, jl:build), and the ref-count column — count 0 marks orphans, high counts mark hot concepts like jl:view.queue. Every description there traces back to one source comment at the implementing code.

Prerequisites

  • Go toolchain (the extractor and generator are Go; my repo also uses go generate).
  • The example repo: dhcgn/jxleet, generated file docs/GLOSSARY.md (44 keys on 13.09.2026).
  • The reusable skills: dhcgn/personal-dev-skills (glossary-init, glossary-create, glossary-search, plus glossary-create-issue).
  • Optional: rg (with git grep fallback) for the search scripts.

How it works

1. Define once, in code, at the implementation

Every user-visible element, process step, and feature gets one Key + Description pair. The key starts with a project prefix (jl: for jxleet), then dot-namespaced lowercase segments:

  • jl:domain.route.transcode
  • jl:domain.output.replace
  • jl:view.queue
  • jl:tech.update.notify-only

The first segment is a mandatory facet (domain, view, tech, build) that becomes a chapter in the generated file; facets and their order live in a small hand-edited docs/GLOSSARY.topics.yaml. Keys are stable once merged — a rename updates all references in the same change.

Short terms use a single line in the file’s native comment syntax:

// jl:domain.route.transcode=JPEG repacked losslessly with --lossless_jpeg=1; djxl restores the original byte for byte.
# jl:domain.preset.rule-fallback=Trailing "*" rule; without it unmatched files are skipped and reported.

Longer descriptions use a ----fenced block:

// Finalize commits the encoder's temp output ...
//
// ---
// jl:Key: jl:domain.output.replace
// jl:Description: >
//
//	Write a temp file in the target directory, decode-verify it ...
//
// ---
func Finalize(...) ... {

Supported comment styles: Go/TypeScript //, Svelte // + <!-- --> + /* */, CSS /* */, YAML/PowerShell #, Markdown <!-- -->. Only comment payloads match — never string literals or plain prose. Test fixtures, vendored dirs, and the generated file itself are excluded.

2. Reference it with ref:, everywhere else

Definitions answer “what is it called”; references answer “we mean that”:

// Each value names a preset file (ref:jl:domain.preset).
Every file takes one `ref:jl:domain.route`, never a bare format name.

Code is scanned in comments only; Markdown is scanned as full prose. A reference with no matching definition lands in a Broken references section of the generated file — and generation exits non-zero, so dangling pointers can never merge silently. Refs to bare facets (jl:domain) never resolve; only full defined keys do.

Docs get the same mechanism as invisible tags — one HTML comment per section, zero visual clutter:

<!-- ref:jl:domain.route.transcode -->

3. Generate, and let the gates keep it fresh

A small generator walks the repo, merges definitions, groups by facet chapter, sorts alphabetically, and writes docs/GLOSSARY.md with Key + ref-count + Description tables. The ref-count shows how many ref: pointers aim at each key, so orphans (count 0) and hot concepts stand out.

Staleness fails the gate: a test regenerates into memory and diffs against the committed file, so go test ./... fails when someone edits a comment without regenerating:

go generate ./...
go test ./...

Nothing here is Go-specific in principle: the generator is a small text walk (collect comment payloads, match key patterns, render tables), so a script in whatever language you prefer — Python, PowerShell, TypeScript — does the same job. For me it was a close call; I picked Go’s built-in go generate because the directive lives next to the code it regenerates and needs no extra runner wired into CI.

4. Use it with agents and in issues

Three skills chain the workflow (glossary-init → glossary-create → glossary-search):

  • glossary-init bootstraps a repo from zero (copy extractor, pick prefix, topics file, seed 3–5 definitions, wire gates, extend AGENTS.md).
  • glossary-create is the daily convention: define once at the anchor file, ref: everywhere else.
  • glossary-search ships read-only finder scripts with identical behavior in both shells — scripts/glossary-search.sh and scripts/glossary-search.ps1 — so humans and agents validate and navigate every key definition and ref: the same way:
scripts/glossary-search.sh defs                                  # all definitions
scripts/glossary-search.sh refs                                  # all ref: mentions
scripts/glossary-search.sh usages jl:domain.output.replace       # everything about one key
scripts/glossary-search.sh keys                                  # unique tokens, diff against the .md table for gaps
scripts/glossary-search.ps1 defs
scripts/glossary-search.ps1 refs
scripts/glossary-search.ps1 usages jl:domain.output.replace
scripts/glossary-search.ps1 keys

Output is plain file:line:text; the scripts need only rg (preferred) or git on PATH, and already exclude test fixtures, node_modules, and the generated file itself.

The prefix is a parameter (glossary.New("acme")), so other repos reuse the same tooling with their own namespace — that is what makes keys referenceable across projects.

In GitHub issues, every key becomes clickable through a repo-scoped code-search link, so the reader sees exactly what I mean instead of my paraphrase:

https://github.com/search?q=repo%3Adhcgn%2Fjxleet+%22jl%3Aview.queue%22&type=code

I know this is something every LSP (Language Server Protocol) is able to handle, this is just my lightweight, language-agnostic approach that works across different shells and CI environments.

Verification

go generate ./...
git diff --exit-code -- docs/GLOSSARY.md
go test ./...
  • No diff: comments and glossary are in sync.
  • Tests green: no stale output, no duplicate definitions with differing text, no broken ref: pointers.
  • Spot check without the generator: pick a key in docs/GLOSSARY.md and confirm every pointer resolves — scripts/glossary-search.sh usages <key> (bash) or scripts/glossary-search.ps1 usages <key> (PowerShell).

Limitations

  • Tested in a Go-led repo (Go, TypeScript/Svelte frontend, YAML, PowerShell, Markdown). Other stacks reuse the parsing ideas but need their own walk.
  • Facet chapters are hand-curated in GLOSSARY.topics.yaml; keys with an unknown facet render under Uncategorized with a warning instead of failing — strictness is a deliberate follow-up.
  • Wording of prose is convention, not CI-gated: AGENTS.md declares the glossary normative for naming (SHOULD-use keys), agents ask and fix the description on ambiguity.
  • Early-stage experiment: the skill set and the extractor will keep evolving; check the repos for the current state.