About Eric Caskey
Senior Software Engineer at Amazon. Safety-critical platforms, and the engineering notebook that goes with them.
The short version
I am a Senior Software Engineer at Amazon, on a platform team whose job is to be a force multiplier for service teams. I joined in June 2022 and built a monitoring lifecycle platform that brought 3 million monitors under standardized management across 2,750+ application stages, while hiring and mentoring eight new-to-industry engineers in New York. Since August 2024 I have owned that platform and architected the workflow orchestration control plane whose guardrails validate every change before it runs, for development teams in multiple global regions. Those guardrails have run 500,000 safety checks so far.
The part I care most about is the safety layer. A pre-execution engine runs concurrent checks before every automated operation: change control, regional isolation, monitoring readiness, pipeline health. A just-in-time composite monitor aggregates the alarm sources an execution depends on into one ephemeral, execution-scoped check, and an environment-isolation guardrail cannot be opted out of. An incorrect yes from a system like this is a production incident, so the system is built to earn every yes.
I also drove the platform's spec-driven AI strategy: a specification-as-code system as the single source of truth, and a hub-and-spoke setup where specialist coding agents inherit a curated slice of context instead of the whole corpus. The bet is that context architecture beats documentation dumps. So far it has held.
Before Amazon
Nine years at Prudential Financial, from Tier 3 help desk to Senior Remote Access SRE, keeping enterprise VPN and MFA services quiet: a VPN serving 60,000 corporate users through the early months of COVID, an MFA self-service portal used more than 18,000 times in six languages, over 300,000 activation codes generated by a QR enrollment tool I built as a side project in 2013 and ran for eight years, and over 200,000 administrative actions automated. Most of that work was removing failure modes rather than papering over them.
Nights and weekends
This site is where I build in the open and write honestly about how it goes: a finance reviewer where the model narrates numbers it never derives, an open-source RAG system engineered to refuse when it should, a C++ options engine, a marathon coach grounded in real training data, and the market visualizations in the Lab. The case studies carry the architecture and the decisions; the blog carries the field notes, including the failures.
What I believe
- Guardrails, not booleans. Every guardrail I build has to say not just what failed, but why, what would make it pass, and who owns remediation.
- Structure before code. The specs, folder structure, and decision logs that precede implementation decide whether the system holds up; the AI follows the structure, not the other way around.
- Scale is the evidence. A title tells you where someone sits, not what they own; the systems, and the numbers behind them, carry the argument.
- Earned lessons only. I write what I have learned by getting it wrong first, in production.
Mission
Build the systems other engineers depend on. Write honestly about how it goes. Outside of engineering: a dad, a marathoner, and a constant reader on AI and distributed systems.
The long version, in my own words, is on ericcaskey.com/about.