Booz Allen CTO Bill Vass on why resilience beats perfection—and how designing for failure strengthens mission systems.

Lessons from the Edge

Resilience Tops Perfection

Decades in IT have taught me: Designing for failure wins.

By Bill Vass, Chief Technology Officer at Booz Allen
Velocity Magazine | V5. Summer 2026

abstract image of resilient technology

Speed Read ↗︎

  • Modern resilience assumes systems will fail—IT leaders need to focus on containing impact and recovering quickly, not chasing perfect uptime. 
  • Good architecture limits the blast radius: segmentation, zero trust, and multi-region design strengthen both operational resilience and security.
  • Teams that measure, rehearse, and learn recover faster—leveraging visibility, disciplined postmortems, and automation to turn failure into continuous improvement.
This is a web summary.

Download the article for extended discussion of deployment methodology, national security and critical infrastructure resilience, and real-world platform examples, or download the full edition of Velocity Magazine for more insights for innovators.

Speed Read ↗︎

Modern resilience assumes systems will fail—IT leaders need to focus on containing impact and recovering quickly, not chasing perfect uptime. Teams that measure, rehearse, and learn recover faster—leveraging visibility, disciplined postmortems, and automation to turn failure into continuous improvement.

This is a web summary.

Download the article for extended discussion of deployment methodology, national security and critical infrastructure resilience, and real-world platform examples, or download the full edition of Velocity Magazine for more insights for innovators.

For decades, technology leaders have chased uptime like it was the ultimate prize. But here's the truth: Failure is inevitable. After 48 years building and operating distributed systems—from the Pentagon to Sun Microsystems, startups, AWS, and now Booz Allen—I've learned something you can count on:

Systems will fail. Networks fail. Hardware breaks. Air-conditioning units overheat. Power drops. Software behaves in ways no one predicted. The specifics don't matter—but your response does. Resilience doesn't come from preventing failure. It comes from surviving it well.

lesson one

Assume Everything Fails

Once you accept that failure is normal, your focus shifts. Instead of wondering if something will break, you manage how much breaks when it does. Today's cloud, edge, AI pipelines, and distributed data stores multiply both scale and complexity—every dependency is another opportunity for things to go wrong.

Resilient systems don't rely on strong hardware or perfect networks. They assume the opposite. Stateless, load-balanced services are forgiving by design—if one node fails, traffic shifts; if one system degrades, the rest keep running. And deployment becomes methodical: start with one instance, let it bake, roll back if it fails, then expand zone by zone, region by region.

lesson two

Limit the Blast Radius

The most resilient systems I've seen are intentionally segmented. Multi-availability zone and multi-region deployment aren't about redundancy for its own sake—they're about containment. Bulkheads keep one failure from sinking the entire ship. Circuit breakers stop bad dependencies from dragging healthy services down.

Zero-trust principles reinforce the same idea from a security angle: never assume a component is healthy just because it exists inside the perimeter. Real resilience stacks multiple, independent, heterogeneous layers—network paths, firewalls, encryption—so a flaw in one doesn't compromise them all. This matters especially in national security and critical infrastructure, where systems must degrade gracefully and keep working long enough for humans to intervene. 

about the author

Bill Vass

Chief Technology Officer, Booz Allen

An award-winning technology leader who has built and operated distributed systems at massive scale for robotics startups, Fortune 100 companies, and the federal government—including roles at the Pentagon, Sun Microsystems, and AWS.

lesson three

You Can't Fix What You Can't See

Within minutes of an outage, I can tell whether a team will recover quickly or spiral. The difference is visibility, experience, and discipline. The best teams have deeply instrumented systems—they know what normal looks like, so anomalies stand out immediately. They investigate what failed, why it failed, and how to prevent recurrence, then automate those root causes away.

AI and automation are now supercharging this process. AI-assisted operations can accelerate log analysis, find patterns humans might miss, and speed diagnosis when seconds matter—but only if systems are observable to begin with. Over time, with the right inputs, agentic systems could move toward self-healing, detecting root issues and automatically initiating corrective action.

High-reliability organizations don't depend on luck. They're methodical. They take small, deliberate steps—every single time.

lesson four

Define What Matters Before it Breaks

Not every system deserves the same level of protection—pretending otherwise wastes time and money. Before an outage happens, classify your applications: which are mission critical, which are business critical, and which can tolerate some downtime. Every level of resilience carries tradeoffs. Live-live replication across regions is expensive. Hot-cold deployments cost less but accept some downtime. Backup-and-restore models are cheaper still, but recovery may take hours. Decide ahead of time what risks you're willing to take. Resilience should be intentional, not reactive

takeaway

Build for Reality, Not Perfection

Today's systems span hyperscale environments, on-prem networks, edge nodes, and devices operating far from reliable connectivity. In national security and defense, systems must keep operating even when the network is degraded, jammed, or gone entirely. Resilience can't depend on constant connectivity, perfect data, or centralized control. It must be designed in from the start.

Where to Begin:
  1. Start with one system that matters.
  2. Map its dependencies honestly, not optimistically.
  3. Instrument it so you can see when it's under stress.
  4. Practice failure—on purpose—before it happens.
  5. Learn from every incident and bake those lessons back into the design.
  6. Repeat.

Read the Full Article

The article PDF includes extended discussion of deployment methodology, national security and critical infrastructure resilience, and real-world platform examples.

New edition | v5. summer 2026

Explore the New Velocity

cover story

Reimagining Cyber for a Faster Fight

Securing enterprises against AI threats requires disrupting operating models, enriching detection, and strengthening resilience—because attacks now unfold in minutes, not days. 

graphic representing AI agent

tech spotlight

How Can You Trust Agentic AI? Start with Engineering

Why trust must be designed, governed, and validated—not assumed.

image of city infrastructure

mission spotlight

Infrastructure Under Attack: The Zero Trust Imperative

Cybersecurity must go beyond compliance to defeat new threats.

graphic of technology intersecting with Washington DC

in conversation

Infrastructure to Impact with Raghu Raghuram

An interview with Raghu Raghuram, managing partner at a16z.

abstract image of math

emerging trends

The Math that Makes Technology Trustworthy

Formal methods and automated reasoning are reshaping software and AI security.

graphic representing resilient technology

lessons from the edge

Resilience Tops Perfection: Desiging for Failure Wins

Resilience doesn't come from preventing failure, it comes from surviving it well.

image of digital fingerprint

tech watch

Don't Take My Word For It: Zero-Knowledge Proofs

Trusting more (but revealing less) with zero-knowledge proofs for government.

graphic of technology intersecting with Washington DC

in conversation

Infrastructure to Impact with Raghu Raghuram

An interview with Raghu Raghuram, managing partner at a16z.

abstract image of math

emerging trends

The Math that Makes Technology Trustworthy

Formal methods and automated reasoning are reshaping software and AI security.

graphic representing resilient technology

lessons from the edge

Resilience Tops Perfection: Desiging for Failure Wins

Resilience doesn't come from preventing failure, it comes from surviving it well.

image of digital fingerprint

tech watch

Don't Take My Word For It

Trusting more (but revealing less) with zero-knowledge proofs for government.

New edition | v5. summer 2026

Explore the New Velocity

cover story

Reimagining Cyber for a Faster Fight

Securing enterprises against AI threats requires disrupting operating models, enriching detection, and strengthening resilience—because attacks now unfold in minutes, not days. 

graphic representing AI agent

tech spotlight

How Can You Trust Agentic AI? Start with Engineering

Why trust must be designed, governed, and validated—not assumed.

image of city infrastructure

mission spotlight

Infrastructure Under Attack: The Zero Trust Imperative

Cybersecurity must go beyond compliance to defeat new threats.

graphic of technology intersecting with Washington DC

in conversation

Infrastructure to Impact with Raghu Raghuram

An interview with Raghu Raghuram, managing partner at a16z.

abstract image of math

emerging trends

The Math that Makes Technology Trustworthy

Formal methods and automated reasoning are reshaping software and AI security.

graphic representing resilient technology

lessons from the edge

Resilience Tops Perfection: Desiging for Failure Wins

Resilience doesn't come from preventing failure, it comes from surviving it well.

image of digital fingerprint

tech watch

Don't Take My Word For It: Zero-Knowledge Proofs

Trusting more (but revealing less) with zero-knowledge proofs for government.

graphic of technology intersecting with Washington DC

in conversation

Infrastructure to Impact with Raghu Raghuram

An interview with Raghu Raghuram, managing partner at a16z.

abstract image of math

emerging trends

The Math that Makes Technology Trustworthy

Formal methods and automated reasoning are reshaping software and AI security.

graphic representing resilient technology

lessons from the edge

Resilience Tops Perfection: Desiging for Failure Wins

Resilience doesn't come from preventing failure, it comes from surviving it well.

image of digital fingerprint

tech watch

Don't Take My Word For It

Trusting more (but revealing less) with zero-knowledge proofs for government.

New edition | v5. summer 2026

Explore the New Velocity

cover story

Reimagining Cyber for a Faster Fight

Securing enterprises against AI threats requires disrupting operating models, enriching detection, and strengthening resilience—because attacks now unfold in minutes, not days. 

graphic representing AI agent

tech spotlight

How Can You Trust Agentic AI? Start with Engineering

Why trust must be designed, governed, and validated—not assumed.

image of city infrastructure

mission spotlight

Infrastructure Under Attack: The Zero Trust Imperative

Cybersecurity must go beyond compliance to defeat new threats.

graphic of technology intersecting with Washington DC

in conversation

Infrastructure to Impact with Raghu Raghuram

An interview with Raghu Raghuram, managing partner at a16z.

abstract image of math

emerging trends

The Math that Makes Technology Trustworthy

Formal methods and automated reasoning are reshaping software and AI security.

graphic representing resilient technology

lessons from the edge

Resilience Tops Perfection: Desiging for Failure Wins

Resilience doesn't come from preventing failure, it comes from surviving it well.

image of digital fingerprint

tech watch

Don't Take My Word For It: Zero-Knowledge Proofs

Trusting more (but revealing less) with zero-knowledge proofs for government.

graphic of technology intersecting with Washington DC

in conversation

Infrastructure to Impact with Raghu Raghuram

An interview with Raghu Raghuram, managing partner at a16z.

abstract image of math

emerging trends

The Math that Makes Technology Trustworthy

Formal methods and automated reasoning are reshaping software and AI security.

graphic representing resilient technology

lessons from the edge

Resilience Tops Perfection: Desiging for Failure Wins

Resilience doesn't come from preventing failure, it comes from surviving it well.

image of digital fingerprint

tech watch

Don't Take My Word For It

Trusting more (but revealing less) with zero-knowledge proofs for government.

New edition | v5. summer 2026

Explore the New Velocity

graphic of technology intersecting with Washington DC

cover story

Reimagining Cyber for a Faster Fight

Learn how CISOs are rebuilding to keep pace with AI-powered attacks.

graphic representing AI agent

tech spotlight

How Can You Trust Agentic AI? Start with Engineering

Why trust must be designed, governed, and validated—not assumed.

image of city infrastructure

mission spotlight

Infrastructure Under Attack: The Zero Trust Imperative

Cybersecurity must go beyond compliance to defeat new threats.

graphic of technology intersecting with Washington DC

in conversation

Infrastructure to Impact with Raghu Raghuram

An interview with Raghu Raghuram, managing partner at a16z.

abstract image of math

emerging trends

The Math that Makes Technology Trustworthy

Formal methods and automated reasoning are reshaping software and AI security.

graphic representing resilient technology

lessons from the edge

Resilience Tops Perfection: Desiging for Failure Wins

Resilience doesn't come from preventing failure, it comes from surviving it well.

image of digital fingerprint

tech watch

Don't Take My Word For It: Zero-Knowledge Proofs

Trusting more (but revealing less) with zero-knowledge proofs for government.