Apollo 13
…and the Difference Between Working and Continuing to Work
We all know the famous line, “Okay Houston, we’ve had a problem here,” but perhaps fewer people are familiar with another quote: “Our mission was called ‘a successful failure,’ in that we returned safely but never made it to the moon.” — Jim Lovell, Mission Commander
Even today, NASA describes the Apollo 13 mission as a successful failure. The word successful still sounds unusual because we are accustomed to considering a project successful only when it achieves the objective for which it was designed. Apollo 13 never reached the Moon, yet it continues to be studied as one of the most extraordinary missions in the history of space exploration.
Why? And what does this have to do with cybersecurity?
Let’s start from the beginning.
On April 13, 1970, Apollo 13 was approximately 300,000 kilometers from Earth—quite a long way from home. Everything was proceeding according to plan until the explosion of an oxygen tank in one of the spacecraft’s five propulsion systems changed everything, triggering a complex chain of failures and forcing Mission Control in Houston to abort the lunar landing.
The explosion did not merely damage the propulsion system. It also compromised the Command Module—the only spacecraft designed to bring the crew safely back to Earth—which gradually lost its ability to generate electrical power and produce water. The mission’s original objective ceased to exist almost instantly. Only one priority remained, one that until then had been secondary: bringing the three astronauts home alive.
NASA’s solution has since become part of history.
The Aquarius Lunar Module, originally designed to transport two astronauts to the lunar surface and sustain them for approximately forty-five hours, was transformed into a lifeboat capable of supporting three astronauts for nearly ninety hours during their journey back to Earth. No one had ever designed that spacecraft to perform such a role.
Every procedure written for the mission suddenly became obsolete. None of the documented scenarios or predefined plans matched what had actually happened. Mission Control found itself facing an unprecedented situation and had to develop an entirely new operational plan.
NASA describes this effort with a sentence that perfectly captures the scale of the challenge:
“It was necessary to write entirely new procedures and test them in the simulator before transmitting them to the crew.”
Apollo 13 is often remembered as a triumph of improvisation. Yet, reading NASA’s technical documentation reveals a different story. The term improvisation does not accurately describe what actually happened. Engineers focused on real-time data and the operational context. Information transmitted by the spacecraft was interpreted in Houston, reproduced as faithfully as possible, and transformed into a sufficiently accurate representation of the situation. Each proposed solution was tested in the simulator—allowing engineers to experiment without increasing the risks faced by the crew in space—and only then communicated to the astronauts.
The goal, therefore, was not to make the fastest decision, but the one most consistent with the actual conditions of the mission. The distinction is fundamental because Apollo 13 was not saved simply by the amount of information available or by the speed at which decisions were made—as we are often inclined to believe is the best way to deal with similar situations—but by the ability to transform fragmented information into a shared understanding that was continuously updated and aligned with what was actually happening.
This brings us to the key point: the very same principle applies to cybersecurity.
Today’s digital infrastructures generate an enormous volume of data. Logs, alerts, telemetry, network traffic, and endpoint events provide visibility into virtually every component of an organization. However, the availability of this information does not necessarily translate into an understanding of the overall state of the infrastructure.
A system can be highly visible while, at the same time, remain poorly understood.
This gap becomes particularly evident in complex environments, where IT systems, OT networks, IoT devices, cloud services, and supply chain components continuously interact, and where an event rarely remains confined to the asset in which it originated. It can propagate, alter the behavior of other systems, disrupt a physical process, or compromise a critical function without any of these consequences being immediately apparent from a single alert.
Detecting an anomaly is therefore only the first step. It is essential to understand its context, reconstruct the relationships between the affected assets, assess its impact, and determine the response that will contain the incident without compromising what must continue to operate.
Speed alone does not solve the problem. A rapid response built on an incomplete understanding of the situation is still the wrong response—only delivered faster.
For Houston, every new piece of information changed what Mission Control knew about the mission, every simulation reduced uncertainty, and every procedure validated before becoming operational reduced risk. The true value of the simulator did not lie simply in replicating the spacecraft, but in enabling NASA to observe the consequences of a decision before that decision could produce irreversible effects in space.
Cyber resilience is built on the same principle. According to NIST, it is the ability to anticipate, withstand, recover from, and adapt to adverse conditions, attacks, or compromises affecting systems that rely on digital resources. It is therefore not defined by the absence of incidents, but by the ability to continue achieving the mission’s objectives, even in a compromised environment. If the objective is to protect enterprise infrastructure, this does not eliminate the possibility that it may be attacked; rather, it requires ensuring that, even in such circumstances, it is not compromised.
This changes the way we define success in cybersecurity.
For years, we have measured success primarily by the events we managed to prevent: blocked attacks, remediated vulnerabilities, denied access attempts, and intercepted malware. These are essential prerequisites, but they describe only what happens when security controls perform the function for which they were designed.
Resilience is measured at a different moment—when a defense is bypassed, a component becomes unavailable, or the real-world scenario no longer matches the one that was anticipated. It is at that point that the ability to observe, understand, and respond becomes more important than the ability to execute a predefined procedure.
Apollo 13 returned safely to Earth because NASA continuously transformed knowledge while the incident was still unfolding, converting it into decisions that were both validated and effective.
The Lunar Module, which would never reach the surface of the Moon, kept the crew alive. The simulators, originally designed to prepare astronauts for a predefined mission, were used to develop procedures for scenarios that had never been anticipated.
Deprived of its original mission plan, Mission Control never stopped managing the mission. How? By preserving its ability to understand what was happening. That is the difference between working and continuing to work.
Cybersecurity faces the same distinction today. Protecting a system means ensuring that it performs the function for which it was designed. Building resilience means ensuring that, when the original plan no longer applies, the organization can continue to operate by understanding what is happening and responding in time to find an alternative course of action.