Lessons from the Edge
Decades in IT have taught me: Designing for failure wins.
By Bill Vass, Chief Technology Officer at Booz Allen
Velocity Magazine | V5. Summer 2026
Download the article for extended discussion of deployment methodology, national security and critical infrastructure resilience, and real-world platform examples, or download the full edition of Velocity Magazine for more insights for innovators.
Modern resilience assumes systems will fail—IT leaders need to focus on containing impact and recovering quickly, not chasing perfect uptime. Teams that measure, rehearse, and learn recover faster—leveraging visibility, disciplined postmortems, and automation to turn failure into continuous improvement.
Download the article for extended discussion of deployment methodology, national security and critical infrastructure resilience, and real-world platform examples, or download the full edition of Velocity Magazine for more insights for innovators.
For decades, technology leaders have chased uptime like it was the ultimate prize. But here's the truth: Failure is inevitable. After 48 years building and operating distributed systems—from the Pentagon to Sun Microsystems, startups, AWS, and now Booz Allen—I've learned something you can count on:
Systems will fail. Networks fail. Hardware breaks. Air-conditioning units overheat. Power drops. Software behaves in ways no one predicted. The specifics don't matter—but your response does. Resilience doesn't come from preventing failure. It comes from surviving it well.
lesson one
Once you accept that failure is normal, your focus shifts. Instead of wondering if something will break, you manage how much breaks when it does. Today's cloud, edge, AI pipelines, and distributed data stores multiply both scale and complexity—every dependency is another opportunity for things to go wrong.
Resilient systems don't rely on strong hardware or perfect networks. They assume the opposite. Stateless, load-balanced services are forgiving by design—if one node fails, traffic shifts; if one system degrades, the rest keep running. And deployment becomes methodical: start with one instance, let it bake, roll back if it fails, then expand zone by zone, region by region.
lesson two
The most resilient systems I've seen are intentionally segmented. Multi-availability zone and multi-region deployment aren't about redundancy for its own sake—they're about containment. Bulkheads keep one failure from sinking the entire ship. Circuit breakers stop bad dependencies from dragging healthy services down.
Zero-trust principles reinforce the same idea from a security angle: never assume a component is healthy just because it exists inside the perimeter. Real resilience stacks multiple, independent, heterogeneous layers—network paths, firewalls, encryption—so a flaw in one doesn't compromise them all. This matters especially in national security and critical infrastructure, where systems must degrade gracefully and keep working long enough for humans to intervene.
about the author
Chief Technology Officer, Booz Allen
An award-winning technology leader who has built and operated distributed systems at massive scale for robotics startups, Fortune 100 companies, and the federal government—including roles at the Pentagon, Sun Microsystems, and AWS.
lesson three
Within minutes of an outage, I can tell whether a team will recover quickly or spiral. The difference is visibility, experience, and discipline. The best teams have deeply instrumented systems—they know what normal looks like, so anomalies stand out immediately. They investigate what failed, why it failed, and how to prevent recurrence, then automate those root causes away.
AI and automation are now supercharging this process. AI-assisted operations can accelerate log analysis, find patterns humans might miss, and speed diagnosis when seconds matter—but only if systems are observable to begin with. Over time, with the right inputs, agentic systems could move toward self-healing, detecting root issues and automatically initiating corrective action.
lesson four
Not every system deserves the same level of protection—pretending otherwise wastes time and money. Before an outage happens, classify your applications: which are mission critical, which are business critical, and which can tolerate some downtime. Every level of resilience carries tradeoffs. Live-live replication across regions is expensive. Hot-cold deployments cost less but accept some downtime. Backup-and-restore models are cheaper still, but recovery may take hours. Decide ahead of time what risks you're willing to take. Resilience should be intentional, not reactive
takeaway
Today's systems span hyperscale environments, on-prem networks, edge nodes, and devices operating far from reliable connectivity. In national security and defense, systems must keep operating even when the network is degraded, jammed, or gone entirely. Resilience can't depend on constant connectivity, perfect data, or centralized control. It must be designed in from the start.
The article PDF includes extended discussion of deployment methodology, national security and critical infrastructure resilience, and real-world platform examples.
New edition | v5. summer 2026
cover story
Securing enterprises against AI threats requires disrupting operating models, enriching detection, and strengthening resilience—because attacks now unfold in minutes, not days.
tech spotlight
Why trust must be designed, governed, and validated—not assumed.
mission spotlight
Cybersecurity must go beyond compliance to defeat new threats.
in conversation
An interview with Raghu Raghuram, managing partner at a16z.
emerging trends
Formal methods and automated reasoning are reshaping software and AI security.
lessons from the edge
Resilience doesn't come from preventing failure, it comes from surviving it well.
tech watch
Trusting more (but revealing less) with zero-knowledge proofs for government.
in conversation
An interview with Raghu Raghuram, managing partner at a16z.
emerging trends
Formal methods and automated reasoning are reshaping software and AI security.
lessons from the edge
Resilience doesn't come from preventing failure, it comes from surviving it well.
tech watch
Trusting more (but revealing less) with zero-knowledge proofs for government.
New edition | v5. summer 2026
cover story
Securing enterprises against AI threats requires disrupting operating models, enriching detection, and strengthening resilience—because attacks now unfold in minutes, not days.
tech spotlight
Why trust must be designed, governed, and validated—not assumed.
mission spotlight
Cybersecurity must go beyond compliance to defeat new threats.
in conversation
An interview with Raghu Raghuram, managing partner at a16z.
emerging trends
Formal methods and automated reasoning are reshaping software and AI security.
lessons from the edge
Resilience doesn't come from preventing failure, it comes from surviving it well.
tech watch
Trusting more (but revealing less) with zero-knowledge proofs for government.
in conversation
An interview with Raghu Raghuram, managing partner at a16z.
emerging trends
Formal methods and automated reasoning are reshaping software and AI security.
lessons from the edge
Resilience doesn't come from preventing failure, it comes from surviving it well.
tech watch
Trusting more (but revealing less) with zero-knowledge proofs for government.
New edition | v5. summer 2026
cover story
Securing enterprises against AI threats requires disrupting operating models, enriching detection, and strengthening resilience—because attacks now unfold in minutes, not days.
tech spotlight
Why trust must be designed, governed, and validated—not assumed.
mission spotlight
Cybersecurity must go beyond compliance to defeat new threats.
in conversation
An interview with Raghu Raghuram, managing partner at a16z.
emerging trends
Formal methods and automated reasoning are reshaping software and AI security.
lessons from the edge
Resilience doesn't come from preventing failure, it comes from surviving it well.
tech watch
Trusting more (but revealing less) with zero-knowledge proofs for government.
in conversation
An interview with Raghu Raghuram, managing partner at a16z.
emerging trends
Formal methods and automated reasoning are reshaping software and AI security.
lessons from the edge
Resilience doesn't come from preventing failure, it comes from surviving it well.
tech watch
Trusting more (but revealing less) with zero-knowledge proofs for government.
New edition | v5. summer 2026
cover story
Learn how CISOs are rebuilding to keep pace with AI-powered attacks.
tech spotlight
Why trust must be designed, governed, and validated—not assumed.
mission spotlight
Cybersecurity must go beyond compliance to defeat new threats.
in conversation
An interview with Raghu Raghuram, managing partner at a16z.
emerging trends
Formal methods and automated reasoning are reshaping software and AI security.
lessons from the edge
Resilience doesn't come from preventing failure, it comes from surviving it well.
tech watch
Trusting more (but revealing less) with zero-knowledge proofs for government.