A nine-lesson series on running an autonomous software team of one, managed from chat. Each lesson ships one small, copyable artifact, and the nine compose into a working system by the end.
An eight-minute vibe-coding challenge at a Mouse AI League mentor mixer, where I volunteer as a mentor and judge: an AI agent built a French flashcard app in eight minutes, and the speed of the machine is what makes product judgment the…
About ten on one subscription, and the model is never what stops you. Four shared resources under the agents set the ceiling: the link to the vendor and its quota window, the git checkout, the port your tests bind, and the file every bra…
Ballast declines 24 of its 89 golden questions on purpose and answers the other 65 with citations. As of the August 16 ledger the answer-or-decline call was right on all 89, and one of the 65 answers was unfaithful to its sources. Where…
Thirty days of running an agent fleet from chat: 8.7 billion tokens, 70,227 calls, 2,861 sessions, a 98 percent cache hit rate. At each vendor's list price that traffic bills $7,566. Three subscriptions covered it for $599. Where the tok…
I split a day of development across two AI subscriptions: judgment on the Anthropic meter, typing on Moonshot's. Five merged PRs later, here is where the tokens went, how K3's pricing actually compares, and how I'm tuning the mix.
I built a self-healing RAG pipeline, a guardrails gateway, and an eval gate as one system, then threw a 44-question battery at it. Zero hallucinations, because the behavior that matters most is refusal. Here is how trust got built into t…
A year of building and operating a small fleet of finance and content products almost entirely through an AI coding agent. What worked, what was hard, the honest failures (including a flagship signal that measured nothing and an edge tha…
Prompt caching looks like a flag you flip for a cheaper bill. It is really the reuse of a stored prompt prefix, governed by three rules, and applying it across four parts of my own system showed where it pays, where it quietly does nothi…
Four days after I said goodbye to Opus, an export-control directive pulled Fable 5 offline and the fallback became the workhorse again. What I shipped in the window, what it cost, and the model-tiering plan for when Fable comes back.
I handed a backlog to Claude Fable, told it once it could merge, and let it run. It shipped seventeen items across five repos. The line that mattered was not in the work it finished. It was in the work it refused to touch.
Anthropic shipped Claude Fable 5 and Mythos 5: same model, two names, one safeguard layer apart. What the new frontier model means for running agents in production.
Dumping the whole corpus into an AI agent makes it worse, not better. The fix is architectural: each task loads a curated slice, not everything you have. Here is the method, and the same move at three different layers: specs, sensor data…
A small ARM box that started as a local LLM experiment and ended up a self-governing node: private retrieval, a resident agent under a written constitution, a code-enforced safety fence, and a nightly job where it audits itself and files…
Anthropic published a guide on building a session-level orchestration mode. I built it two ways, on the CLI and on the API, and then hit the part the guide does not cover: an orchestrator that fans out is useless without a backlog of rea…
How I replaced manual CSV exports with a live Garmin data feed for my AI marathon coach: a scheduled unofficial-API poller, resilient session handling, and the design calls that keep training and recovery data fresh and trustworthy.
Spec-driven development reads like a methodology for controlling AI agents. It isn't. It's a methodology for managing context across stateless sessions. The spec is the persistent memory.
Two production sites, a blog, and two personal AI projects, shipped this week from a phone. The chain is voice dictation into Perplexity Computer, a spec, then Claude Code on the web. The interaction model is the story.
How I built a personal AI coaching system for marathon training, layering deterministic guardrails over an LLM narrative engine, ingesting Garmin FIT files, and designing for my own injury history.
Fourth lesson in a series on running an AI-powered software team of one. A session pointed at one repository commits there and nowhere else, so a project that lives in one folder cannot be delegated safely. Split it into one home per sta…
Third lesson in a series on running an AI-powered software team of one. A rule you write in a README gets broken by the first session that never reads it. A folder of one-script checks, run by git before every push, turns your most-repea…
Second lesson in a series on running an AI-powered software team of one. A fresh agent session cannot see what you rejected, so it proposes it again. One page per settled choice, an index, and one standing line make every session read yo…
The third post in the SDD numbers series is not about velocity. Since June the platform's effort moved into four quiet systems: an externalized memory, a statistical screen that refuses to flatter me, a point-in-time data floor, and a pr…
The full path to an agent platform that edits its own instructions is already published, scattered across seven articles that never mention each other. I run that architecture as one person: supervisor and worker tiers, verification guar…
First lesson in a series on running an AI-powered software team of one. Before you ask an agent to build, give your project a home for intent: a small directory of numbered specs grown from one template your agent fills in and you approv…
SpecSelf looks like a set of features: coherence checks, persona rotation, review cadences, an audit trail. Every one of them was implied by ten frontmatter fields decided on day one. A life-OS is not a feature list. It is a schema decis…
I built a method as a public repo and the product that runs it as two private ones, and none of them treated the others as a source of truth. The domain enum lived in four places. A persona drifted between its lens file and its API contr…
Three months ago I laid out spec-driven development and the folder architecture that makes it work. Most of the playbook held up in daily production use. Three ideas that essay never mentioned turned out to matter more than anything in i…
In one working session I designed a read-only advisor and spent the rest of it being the customer of another. Both ran on the same short discipline: surface the decision and never seize it, verify before you assert, pilot before you fan…
A team holds its hard-won knowledge across many heads. A solo operator holds it in one, and that one forgets. The fix is to externalize memory into structured records the tools read by default, so the system remembers what the person can…
The most dangerous result is the one you want to be true. Your own review is compromised by the same motivation that produced the finding, so the fix is a standing skeptic whose job is to refute, not confirm, before you act on anything.
The instinct with agentic tooling is to add: more agents, more skills, more clever prompts. The leverage runs the other way. Here is the test I use to decide whether a piece of work should be a script, a hook, a skill, or an agent, and w…
Boris Cherny mapped five execution archetypes on the Claude Code team, and noted they cut across job titles. His framework describes a team dividing labor across people. Run a fleet alone and the same five split a different way: across y…
Why good validation reports every problem at once instead of failing on the first one, and how to build the accumulator, phasing, and structured errors that make it work.
Why spec-driven development and structured folder architecture are the missing infrastructure for AI-assisted engineering: methodology, common mistakes, and where to start.
A practitioner's review of Doug Kerwin's Enterprise Vibe Coding Playbook, why AI as a thinking partner, not a replacement, is the framework enterprise engineering teams need.
After GitHub Actions burned a month of minutes in two days, the gate moved to the workstation. A pre-push hook runs a quick tier on every push, every green run writes a record tagged with its tier and commit, and a merge guard refuses an…
The navigation on caskeycoding.com is a build artifact: one public content graph, several switchable views projected from it, and CI checks that keep the graph and the routes from drifting apart.
A Windows Task Scheduler default fed my Claude Code memory pipeline its own transcripts for nine days: 21,264 junk sessions, 38 GB of exhaust, and an 18,428-session backlog that was really 1,924. The durable fix: postconditions that outr…
My kids collect Pokemon cards, and one of them asked what a Charizard is actually worth. Answering it turned into a weekend build: a chart that puts graded cards and the S&P 500 on one indexed axis, and admits out loud where its data is…
You cannot out-staff a security team when you are the whole team. But the failures that actually end a solo operation are a short, known list, and each has a cheap defense you set up once. Here is the catastrophic floor I stood up in an…
GitHub Actions' default minute allowance is priced for a team that types at human speed. At agent velocity the bill breaks before the engineering does. Here is how a forced workaround, a local CI mirror plus local deploys, became the bet…
A high-level tour of the technologies running this site: Next.js on CloudFront, Python Lambdas behind API Gateway, DynamoDB plus S3, Anthropic's API with a Bedrock fallback, and AWS CDK wiring it together.
Where this blog started: owning enterprise monitoring at Prudential and Amazon, an automation mishap that paged a whole support queue for ten minutes, and the throughline that still runs through everything I build, make the safe path the…
My build log said the loop was vectorized, and it had: a trivial loop at the end of the benchmark. The pricing loop, the one with exp and log in it, stayed scalar until I wrote the four-wide math by hand and the engine went from about 11…
The Greeks manifold prices a long call on a 60 by 60 Black-Scholes grid over spot and time to expiry, draws profit and loss as height against a premium fixed at the strike and the longest tenor, and colors each cell by one Greek. Gamma i…
The yield curve surface reads 11 FRED constant-maturity Treasury series from one month to thirty years, keeps the 60 most recent complete daily curves, and labels the curve inverted when the 30-year yield is more than 0.1 percentage poin…
A Newey-West standard error corrects a t-statistic for autocorrelation. On my engine's 12-month panel it turned 18 monthly readings with a t-statistic of negative 18.8 into about two independent observations. The arithmetic, the lag choi…
A viral post claimed a simple 5/10 moving-average strategy that has not lost in eight years. I codified the most charitable readings of the claim, ran each one over SPY with costs, and checked the claim against its own success metric. It…
The order block is trading social media's favorite glossary card. I wrote the definition down as code, ran it across 30 large-cap names and five years of daily bars, and measured what the retest entry actually earns net of costs. The ver…
I built a quant research platform, then built an agent to operate it: a scheduled Claude session that reads the boards, keeps a pre-registered track record, and texts me three times a day without ever saying buy.
I wrote a C++ options pricer to learn low-latency numerics. The first clean version priced fifteen million options a second; getting to 215 million was less about clever code and more about being wrong, in public with myself, about where…
A 3D render crossed my feed once and stuck with me, so I tried to see an option the same way: as a surface I could grab and turn, not a number. That turned into five market visualizations on one shared trick, a compliance rule the archit…
A screensaver joke, every stock a fish, grew into a six-lens market board I leave running on a wall. It lives in one HTML file on purpose, its data is baked because the browser is not allowed to fetch it, and the feature that finally mad…
I stopped staring at market dashboards. A set of alarms now watches a dozen signal dimensions across the market and taps me on the shoulder only when something actually needs a decision.
Every system that fuses signals into one consequential number has a fault line: the data you trust enough to composite into a grade versus the data you only trust enough to watch. How I drew that boundary in my personal finance engine, a…
A backtest's job is not to find an edge. It is to stop you from believing in one that is not there. The toolkit I used to test my own trading engine, and the part where it killed my single best signal.
Two posts ago I bet that keeping my portfolio reviewer's engine deterministic and auditable was worth it. This is where that bet paid off: because the engine is replayable, I could run a simulated market crash through the real production…
A personal portfolio reviewer where the scoring is deterministic and the AI only narrates. The architecture that held up after I had to rewrite the model it was built on, and why that boundary is the whole point.
Two weeks after I shipped a post about a scoring engine I'd built, I rewrote the spec it was based on. Here's what I learned, and why I had an AI agent do the literature review.
Spec-driven development, pointed at a life. Why my principles and goals live as markdown an AI agent reasons over, and the one rule that makes handing an agent your life safe: it can challenge the record, but it never writes it.
The Knicks won their first title since 1973, decades before I was born. A lifelong fan on the grandfather who handed down the wait, the lean years that taught patience, and the roster I half-built in my head a decade before it came true.