The Harness, Not the Chaos
Introduce the failure-lab harness
How a repeatable Azure failure-lab harness turns one-off chaos experiments into self-serve drills with inject contracts, alerts, and recovery paths.
Engineering leader with 20+ years architecting secure enterprise Azure platforms. I designed, secured, and led the platform behind a $150M global managed-services business. I build at the frontier, write about what holds up in production, and still ship the code.
I architect and lead the platforms enterprises depend on. At Rackspace I built, secured, and ran the governance control plane behind a $150M global managed-services business, managing 29 million resources, 315,000 policy baselines, and 600,000 alert baselines across 36,000+ tenants, 5,000 subscriptions, and 13 global sites.
The outcomes are what I'm measured on. I refactored that platform from $56K to $9K per month, an 84% cut while increasing capacity, drove mean-time-to-resolution from five hours to fifteen minutes with AIOps, cut deployments from 48 minutes to 6, and passed every annual Azure Expert MSP audit with zero findings. I grew the team from myself to eight engineers and championed AI adoption across the org.
I lead by building. I architect the systems my teams run on and stay close to the code, so my technical decisions hold up and my team trusts them. That work runs from the enterprise identity practice I stood up on Entra ID and federation, through the Zero Trust security model, to the agentic systems I build now: MCP services, custom agent coding skills, and Borg, my open-source and enterprise agent-memory platform.
Security is part of how I build platforms, not a review gate bolted on afterward. I set the operating model for Zero Trust, IAM governance, SIEM/SOAR, DevSecOps, and compliance readiness across a multi-tenant Azure estate where automation had to be trusted at fleet scale.
Designed a multi-layered Zero Trust model across WAF, API Management, network segmentation, managed identities, Entra ID Conditional Access, PIM, RBAC, and federated identity patterns.
Established multi-workspace Microsoft Sentinel operations across 500+ customer tenants, with telemetry onboarding standards, content deployment pipelines, AIOps correlation, and governed remediation.
Standardized secure delivery with Bicep IaC, GitHub Actions, self-hosted runners, SAST, DAST, container and dependency scanning, policy-as-code, detection-as-code, and pre-commit enforcement.
Built MCP security patterns mapped to the OWASP LLM Top 10, hardening tenant scoping, OAuth/SAML authentication, token handling, and audit trace correlation for AI-facing platform surfaces.
My AI work is real engineering, not a list of tools I've tried. I built the agentic surface of a production platform, led AI adoption across my team, and created open-source and enterprise AI infrastructure with published benchmarks. The work below is grouped by where it happened: production, leadership, and product engineering.
Built the platform's programmatic and agentic layer: REST APIs, Model Context Protocol services, agent frameworks, and custom agent coding skills, with the documentation to make it usable inside and out.
Integrated an AIOps framework using machine-learning event correlation that cut mean-time-to-resolution from five hours to fifteen minutes across a multi-tenant estate.
Served as my team's AI champion and trainer, running demos and proofs of concept, upskilling engineers, and integrating AI-assisted workflows into production engineering practice.
Built a Postgres-native memory stronghold with an Apache-2.0 open-source edition and a commercial enterprise deployment. Benchmarked at 10/10 task success and 91.3% retrieval precision. See details ↓
I build with agents daily, writing custom agent coding skills and MCP integrations that extend what they can do and speed up delivery across the stack.
I design how agents fit into real engineering systems, using namespace isolation, token-budgeted context, drift detection integration, and bitemporal fact supersession in production Sentinel pipelines. AIOps detects and recommends; a separately governed automation plane validates and remediates.
Architected and led a versioned governance control plane managing 5,000 subscriptions, 29 million resources, 315,000 policy baselines, and 600,000 alert baselines across 36,000+ tenants and 13 global sites, underpinning a $150M managed-services business.
Built out the Identity and Access Management practice on Azure AD, ADFS, and Microsoft Identity Manager, delivering single sign-on and federated identity at scale across cloud and on-premises.
Local practice lead for the Austin business unit, owning delivery across Azure, Office 365, and federation, managing the full project lifecycle from sales to closure.
Progressed from enterprise support advisor to systems engineer, designing private-cloud, identity, virtualization, and systems-management solutions across Microsoft and VMware platforms.
Built the networking, systems, and university IT consulting foundation that later shaped my enterprise architecture work.
Began my career designing and supporting business networks and telecommunications systems for Comercial e Inversiones SuperMart and Agencias Panamericanas de Sula.
The team I built is the part of this work I'm proudest of. I care about it more than any platform I've shipped.
I grew my engineering function from one person to eight. Hiring was only the start. The harder, more important work was building a place where strong engineers could do their best work and want to stay.
I mentor directly rather than manage from an org chart. I have been the escalation point and the person who teaches since my early engineering roles. As my team's AI champion and trainer, I ran the demos, proofs of concept, and hands-on sessions that got everyone comfortable with tools that were changing fast.
I protect my team's focus, give credit generously, and take the heat when something breaks. People do their best work when they feel valued, trusted, and appreciated, so that is the environment I work to create.
A memory stronghold for AI coding agents, available as an Apache-2.0 open-source project and a commercial enterprise deployment. Every session across Claude Code, Codex, Copilot, and Kiro flows into one Postgres knowledge graph, so the next agent arrives already briefed. No Qdrant, no Neo4j, no sync daemons.
Its five-stage borg_think pipeline classifies intent, retrieves across facts, episodes, and graph relationships, ranks on relevance, recency, stability, and provenance, then compiles token-budgeted context for the target model.
Built for enterprise deployment with Entra JWT, RBAC, namespace isolation, and per-compilation audit trails. Seven MCP tools cover context compilation, learning, recall, fetch, soft deletion, action guardrails, and deterministic project summaries.
Field notes on the platforms, operating models, and engineering decisions that hold up at enterprise scale.
I write for practitioners making the architecture work and leaders accountable for what happens after it ships.
Explore the complete writing library →Introduce the failure-lab harness
How a repeatable Azure failure-lab harness turns one-off chaos experiments into self-serve drills with inject contracts, alerts, and recovery paths.
Establish the full platform story
How a versioned control plane keeps thousands of Azure subscriptions and millions of resources aligned to an approved governance baseline.
Frame the agent-memory failure mode
Why stateless coding agents repeat old mistakes, why vector RAG can trade amnesia for stale misinformation, and why memory needs to be compiled instead of searched.
Explain the main architecture decision
A practical architecture decision record comparing serverless orchestration with Kubernetes for bursty, fleet-scale governance workloads.
Show the repeatable infrastructure foundation
Why repeatable infrastructure, modular Bicep, and controlled ARM deployments turn approved architecture into enforceable fleet state.