Data Dictionary
A data dictionary, in the classic sense, is the catalog of a system's data elements: names, meanings, and where each is used. This page is the site-wide version. It collects the industry terms and house coinages the published corpus reuses, gives each a one-line general definition, and states what this site uses it for. Market entries describe measurement only; nothing here is investment advice.
AI and agents
Vocabulary for building with language models: how work is specified, routed, checked, and paid for.
- Context architecture#
The design of what information a model or reader receives for a task, and what it never sees.
Used here: The site's central working bet: each task loads a curated spec slice instead of the whole corpus, so the model spends its budget on judgment instead of fetching. Implemented as short router files, per-workspace contexts, and named seams between them. See: the context architecture case study, the essay that named it.
- Context window#
The finite amount of text a language model can consider in one call.
Used here: Treated as the scarce resource every other concern competes for. Routing files stay short, workspaces load only their own slice, and long-term memory lives in files outside the window. See: SDD is context management.
- Engineered refusal#
A designed "I do not know" or "I will not answer" output, tested like any other feature.
Used here: Ballast declines out-of-corpus questions, personal-finance advice, and injection attempts with stated reasons. Over-refusal and under-refusal are scored separately, so the gate cannot be passed by refusing everything. See: the Ballast case study.
- Eval gate#
A fixed battery of graded questions that a change must pass before it ships.
Used here: Answer quality is enforced as a build gate: faithfulness, refusal correctness, and hallucination rate over a labeled question set, with results kept in a dated, committed metrics ledger. A regression or a stale ledger fails the merge. See: the Ballast case study.
- Frontier model#
The top, most capable and most expensive tier of a model provider's lineup.
Used here: Spent only where its judgment changes the outcome: planning, briefing, review. After a parallel-agent wave burned through a monthly quota in days, frontier models were banned from fanning out by written policy. See: the model-routing series, the cost ledger.
- Grounding#
Restricting a model's answers to supplied, citable source material.
Used here: The marathon coach receives ninety days of real Garmin history and a prompt that forbids inventing any mileage or pace not in it. Ballast answers only from retrieved passages and shows the citations. See: the coach case study, the Ballast case study.
- Guardrail#
A check a system cannot skip, enforced in code rather than left to discipline.
Used here: Used in two senses the site deliberately parallels: pre-execution safety checks that block unsafe automated operations, and input and output screens that make a model decline. Both must be countable, explainable, and owned. See: the guardrail design essay, the Ballast case study.
- Human gate#
An action class an autonomous agent may never start unattended: spending money, deploying, publishing outward.
Used here: Autonomy expands by adding gates, never by trusting the model more. The backlog loop carries explicit human-gate items, and the gates are re-evaluated each model generation, because a gate the model no longer needs is just latency. See: Autonomy is mostly knowing when to stop.
- Spec-driven development (SDD)#
A workflow where a versioned written spec precedes the code and serves as the acceptance target.
Used here: The site's governing method. Every feature starts as a numbered spec with acceptance criteria and an out-of-scope list, agents implement against it, and review asks whether the diff matches the spec, a tractable question, instead of whether it is good, an unbounded one. See: Lesson 1: the spec directory, the SDD platform case study.
- Standing file#
The markdown file an agent reads at the start of every session, such as CLAUDE.md or AGENTS.md.
Used here: Kept to pointers and rules of conduct. Feature intent lives in specs, because a bloated standing file gets ignored and every token in it is spent on every turn. See: Lesson 1: the spec directory.
- Subagent#
A spawned agent handed one self-contained unit of work with its own fresh context.
Used here: The fan-out unit for parallel work. A subagent's report is treated as a claim, not a result: the supervising session reruns the checks and reads the diff before anything merges. See: the orchestration essay, the parallel-agents essay.
- Supervisor-worker split (model routing)#
Assigning planning and review to a stronger model and execution volume to cheaper ones.
Used here: Written policy, not habit: the frontier session triages the backlog, writes briefs, and reads every diff, while cheaper models type. The cost ledger showed supervision running many times the typing bill, because supervision is reading. See: the model-routing series, the cost ledger.
Engineering practice
The working discipline the site runs on: specs, gates, and the checks that keep them from rotting.
- Acceptance criteria#
The observable states, written before the build starts, that decide whether work is done.
Used here: Written early so the goalposts cannot move once results arrive, and graded line by line at review. A spec without them is a wish. See: Lesson 1: the spec directory.
- Architecture decision record (ADR)#
A dated document recording one decision, the alternatives considered, and the consequences.
Used here: The reasoning layer of the spec corpus. ADRs amend each other with explicit supersedes lines, so a fresh session stops re-proposing what was already rejected. See: Lesson 2: decision records.
- Audit trail (append-only)#
A record that is only ever added to, never rewritten.
Used here: Scores, theses, guardrail outcomes, and deploy events are frozen with their input hashes, so any past decision can be reconstructed exactly and hindsight cannot rewrite history after the fact. See: Composite what you trust.
- Blast radius#
How much breaks when one component fails or one action is wrong.
Used here: An architecture organizer: infrastructure stacks are split by it, credentials are scoped to bound it, and rewrites proceed one axis at a time so a wrong change has somewhere to stop. See: the solo-fleet security baseline.
- CI (continuous integration)#
The automated validation that runs on every proposed change.
Used here: Treated as a metered budget, not free infrastructure. Hosted CI minutes are priced for human typing speed, so the fleet mirrors the pipeline locally and reserves the cloud run for the final record. See: the CI cost essay, the local-first CI essay.
- Config pair#
Two settings in different places that must change together or production breaks.
Used here: The canonical incident: a framework trailing-slash setting and the edge rewrite function drifted apart, and every sub-route served the home page. The fleet keeps a written list of pairs and fails checks when only one side lands. See: the context architecture essay.
- Drift#
The silent divergence of two artifacts that are supposed to agree: spec and code, enum copies across repos, local and hosted CI.
Used here: The standard fix is a machine check at the seam, not vigilance. Where two artifacts must agree, a test or a gate compares them on every run. See: the method-repo essay.
- False green#
A check reporting success while the work it guards is dead.
Used here: A scheduled job once reported success for weeks while its program failed every run, because the wrapper returned the redirection's exit code. Health verdicts here require three independent signals: result, config, and staleness. See: A summer of sharpening.
- Gate (merge gate)#
An enforced checkpoint that refuses a bad state without relying on anyone's memory.
Used here: PR validation failing is the gate doing its job, not friction. Merges require a recorded full local run, and new work ships dark behind flags until deliberately flipped. See: the local-first CI essay.
- Planted failure#
A deliberately broken artifact, planted on a schedule, that the watchdog must catch.
Used here: The answer to who checks the checker: a canary job fails on purpose every week, and silence during the plant is itself an incident. A check never seen failing is treated as already dead. See: A summer of sharpening.
- Worktree#
An isolated git working tree that shares one repository's history.
Used here: Each parallel agent works in its own worktree with its own real dependency install, so concurrent sessions never share a checkout. Junctioned dependency trees are banned after a recursive delete followed one. See: the parallel-agents essay.
The publishing pipeline
How the site and its data ship: static files, scheduled jobs, and baked bundles instead of live APIs.
- As-of stamp#
The date a data artifact describes, rendered next to the numbers it supports.
Used here: Every market readout carries one. A value without an as-of date is treated as a claim the site refuses to make. See: the daily readout.
- Baked file (public bundle)#
Data written to a static JSON artifact by a scheduled job, instead of fetched by the browser from a live API.
Used here: The pattern behind every market page: a nightly job validates and publishes one public bundle, and the pages read that. The browser never calls the data API. See: the markets hub, the daily readout.
- CDN invalidation#
The cache purge at the edge network that ends a deploy so visitors receive the new files.
Used here: Every frontend deploy finishes with a full-site CloudFront invalidation, and the invalidation logs double as the deploy audit trail when the CI API undercounts. See: the phone-first deploy essay.
- FRED#
The Federal Reserve's public economic data service, including the daily Treasury par-yield series.
Used here: The source behind the yield-curve surfaces: the nightly job joins eleven constant-maturity series and publishes only days where all eleven report. See: the yield-curve surface.
- Lambda#
AWS's serverless function runtime: code that runs on demand with no server to operate.
Used here: Runs the site's Python backends and exporters. The fifteen-minute ceiling is a hard design constraint that shapes how collectors chunk and rotate their work. See: the site architecture essay.
- Nightly exporter#
The scheduled backend job that re-scores public data each night and republishes the static bundles.
Used here: The reason the market pages can say they refresh every night, and the reason deploys must exclude the data prefix: the exporter writes to the same bucket the site syncs. See: the markets hub, the daily readout.
- Static export#
Building the site to plain pre-rendered files served from object storage, with no application server.
Used here: The deploy shape of the whole site. It is why route lists are build-time constants, why data ships as baked files, and why a page can be audited byte for byte before it ships. See: the site architecture essay.
Markets and measurement
The market suite's scoring and statistics vocabulary. All of it describes measurement; none of it is investment advice.
- Altman Z-Score#
A published bankruptcy-distance score computed from five balance-sheet ratios.
Used here: Sits inside the Health factor with teeth: below the distress zone the composite is attenuated toward a floor, so a weak balance sheet cannot hide behind strong scores elsewhere. See: the methodology page, Market's Best.
- Backtest#
Replaying an explicit rule against historical data to measure what it would have done.
Used here: An instrument for killing false beliefs before money is involved. Every engine change must clear the offline panel, and published claims carry their costs, their window, and their hash. See: Backtesting without fooling yourself.
- Black-Scholes#
The textbook model for pricing a European option from spot, strike, time, rate, and volatility.
Used here: Powers the 3D Greeks surface, computed live in the browser, and the C++ learning pricer whose output is cross-checked against published reference prices. See: the Greeks surface, the pricer essay.
- Composite score#
The single 0 to 100 weighted blend that ranks each graded stock, mapped to a letter grade.
Used here: The ranking key for the whole market suite. Thresholds render from the engine's contract file rather than hand-typed copy, so the public pages cannot drift from the math. See: the methodology page, Market's Best.
- Correlation regime#
How tightly assets move together, dialed from calm to crisis.
Used here: The correlation-regime surface shows diversification evaporating as pairwise correlations converge to one in a crisis, first on synthetic data, then on real sector funds. See: the correlation surface.
- Drawdown#
Peak-to-trough decline: how far a value has fallen from its high.
Used here: Two jobs here: a market-state gate that suppresses sell-side recommendations during panics, and a 3D topology of simulated crash paths for calibrating how deep a bad run digs. See: the finance case study, the drawdown surface.
- Four factors (Quality, Valuation, Momentum, Health)#
Quality, Valuation, Momentum, and Health: the independent dimensions behind every composite score.
Used here: Scored separately so a blend never hides the story, and missing data redistributes weight across the surviving factors instead of becoming a confident zero. See: the methodology page.
- Implied vs realized volatility (variance risk premium)#
Implied volatility is what option prices expect; realized volatility is what happened. Their persistent gap is the variance risk premium.
Used here: The price-of-fear chart shades that gap back to 1990, and the volatility cone shows where today's realized move sits against its own history. See: the VRP chart, the volatility cone.
- Information coefficient (IC)#
The per-date rank correlation between scores and the returns that followed.
Used here: The site's basic unit of evidence for whether the ranker knows anything, computed nightly and per factor. An IC is not a net-of-cost portfolio return, and the site says so wherever it reports one. See: the methodology page, the factor monitor.
- Insider buying (SEC Form 4)#
Open-market purchases by a company's own officers and directors, disclosed on SEC Form 4.
Used here: Plotted against crowd chatter to separate quiet accumulation from all-talk names in the insiders-versus-crowd quadrant. See: the insiders quadrant.
- Newey-West standard errors and effective n#
A significance correction for autocorrelated series, and the count of how many independent observations a correlated series really contains.
Used here: Applied to the engine's monthly readings, it turned an apparently overwhelming t-statistic into about two independent observations. No factor weight is promoted unless its holdout cell clears the effective-n floor. See: the Newey-West essay.
- Option Greeks#
The sensitivities of an option's price: delta to the underlying, gamma to delta, theta to time, vega to volatility, rho to rates.
Used here: Taught as geometry rather than formulas: gamma is a ridge at the strike sharpening toward expiry, theta is the downhill slope in the time direction, on a surface the reader can rotate. See: the Greeks surface, the reader's guide.
- Out-of-sample test#
Measuring a rule on data that played no part in forming it.
Used here: The gate that killed the site's best-looking signal. For model judgment that cannot be replayed, a live resolver logs dated, immutable predictions and grades them as the future arrives. See: Backtesting without fooling yourself.
- Overfitting and PBO#
Tuning to historical noise until the past looks predictable; PBO is the probability that a selected backtest is an overfit artifact.
Used here: The default assumption about any good-looking backtest until it survives the checklist. The autopsy series computes PBO across claim variants before believing any of them. See: the autopsy series.
- Pre-registration#
Writing the hypothesis, metric, and threshold down before running the test.
Used here: Applied to backtests and to the author's own theses, which are frozen in an append-only journal so yesterday's reasoning cannot be edited to fit today's outcome. See: the pocket quant essay, Backtesting without fooling yourself.
Return per unit of risk; the deflated version discounts the result for how many strategies were tried.
Used here: The standing lesson that a real signal can still be a non-strategy: a 0.42 Sharpe over a short window is statistically indistinguishable from luck once the search is priced in. See: Backtesting without fooling yourself.
- Survivorship bias and point-in-time data#
Testing only on today's winners because the dead companies vanished from the dataset; point-in-time data reconstructs what was knowable on each past date.
Used here: Load-bearing infrastructure, not a footnote: the site rebuilt index membership from SEC filings after catching a free dataset doing survivor-only backfill, and reports affected magnitudes as upper bounds. See: Backtesting without fooling yourself, the methodology page.
- Yield curve inversion (2s10s)#
Short-term Treasury yields sitting above long-term ones, the bond market's classic recession warning; 2s10s is the ten-year minus two-year spread.
Used here: Reduced to one explicit rule on the daily readout, with the spread shown day over day and back to 1976, and a synthetic textbook mode that animates an inversion forming when the real curve is flat. See: the daily readout, the yield-curve surface.
Running and fitness
The sports-science terms the marathon coach computes from Garmin data.
- ACWR (acute:chronic workload ratio)#
Recent training load divided by the longer-term baseline load.
Used here: The coach's core safety metric. Past the hard line the rules engine forces a rest or cross-training call regardless of anything else, and a metric that fails to load stays empty rather than collapsing to a fake zero. See: the coach case study, the coach demo.
- Recovery signals#
The daily wellness inputs: resting heart rate, heart-rate variability, sleep, soreness, readiness.
Used here: The narrow slice that gates hard workouts. Curation is defensive: the rules engine reads only what loaded, and absence is a stated state, never a zero. See: the Garmin wiring essay.
- TRIMP (training impulse)#
A training-load formula combining session duration with heart-rate intensity.
Used here: Computed one way at read time so CSV imports and Garmin-polled activities share a single number, with no source-dependent drift and no silently zeroed load. See: the coach case study.
House vocabulary
Coinages this site reuses as named concepts. Each is defined by how it is used here, not by industry convention.
- The backlog as program#
The BACKLOG.md file that holds both the work queue and the protocol for the agent processing it.
Used here: Rules live in the header, items carry dependencies and acceptance criteria below, and state lives in the file itself, so an unattended loop is safe to leave alone. See: the orchestration essay.
- The board#
The nightly-graded public universe of large-cap stocks.
Used here: Every market page is framed as a view of the board: the sectors read it by group, the stock pages read it by name, and the APEX constellation draws it in three dimensions. See: the markets hub, the APEX constellation.
- "Educational, not advice"#
The standing disclaimer closing every market surface.
Used here: The site publishes measurement and engineering, never recommendations. It claims no demonstrated edge net of costs and labels its own magnitudes as survivorship-affected upper bounds. See: the markets hub, the methodology page.
- The engine#
The site's own deterministic factor-scoring system for stocks.
Used here: Backs the public market pages and the private finance workspace. Its methodology page is the public fact sheet, and its model layer narrates conclusions it is never allowed to reach by itself. See: the methodology page.
- Epistemic status#
The manifest's declaration of how proven an artifact is: production, on the record, in progress, or experiment.
Used here: Rendered as badges across the Explore views, so a reader can tell a running system from a sketch at a glance. See: the site-organization essay.
- Evidence labels (Public, Representative, Private)#
The per-claim provenance tags at the end of every case study.
Used here: Public claims link their artifact, representative claims link a demo or write-up, and private employer work is deliberately unlinked. A missing link is meant to read as integrity, not vagueness. See: the case studies.
- The fleet#
The owner's collective noun for the sites, engines, collectors, and agents he operates solo.
Used here: The unit that monitoring, CI policy, and the security baseline are standardized across, because a team of one needs the machine to be the colleague who catches mistakes. See: Leverage by subtraction, the solo-fleet security baseline.
- Honest absence (tri-state)#
The rendering rule that every data panel distinguishes loading, absent, and ready.
Used here: Absence is a stated state with an as-of date, never a fabricated zero and never a crash. A name that falls off the board keeps a dated notice page instead of a 404 between deploys. See: the daily readout, the sectors overview.
- The investment committee#
Six fixed investor personas that narrate the engine's factor scores in each investor's style.
Used here: They explain the math without recomputing it. A decision record pinned each persona to its own factor after the committee started diverging from the scores. See: the committee tool, the finance case study.
- Market's Best#
The public leaderboard of the engine's highest current grades.
Used here: The standing exhibit for the nightly scan, with a factor-breakdown drawer per name and a market map of the whole scanned universe. See: Market's Best.
- Narrate-not-decide#
The boundary where deterministic math owns the conclusion and the model only explains it.
Used here: The same boundary in the finance engine and the marathon coach: the model is told the numbers, never asked to derive them, and can never overrule a threshold. See: the coach case study, the methodology page.
- Receipts#
Published, hash-pinned evidence artifacts behind a public claim.
Used here: The answer to a viral claim is a definition, a window, a cost assumption, and a hash, so a reader can verify without trusting the author. See: the autopsy series.
- Synthetic mode#
A labeled textbook dataset, always one toggle away, on every real-data visualization.
Used here: The page never silently fakes data: when live data is absent or a concept needs a clean case, the synthetic mode says so on its face. See: the yield-curve surface, the correlation surface.
- The wall#
The structural separation between data allowed to move the grade and data only allowed to be displayed.
Used here: Trusted fundamentals move the composite; prediction markets and crowd sentiment are display-only. A test asserts the grade stays byte-identical when sentiment swings, and a signal crosses only through pre-registration and out-of-sample proof. See: Composite what you trust.