Now Cleared for Failure: Tunnel Vision in Software Verification

Software testing efforts are mainly focused on the application code itself. However, recent major failures were due to configuration changes or system data. In this article, Sergei Cherednichenko discusses how quality assurance efforts should widen their scope from an application-centric perspective to a system-wide reliability verification.

Author: Sergei Cherednichenko, https://linkedin.com/in/sergei-cherednichenko/

Several recent high-profile outages, including Cloudflare’s oversized bot-management feature file and GitHub’s policy access change, have resulted from configuration changes or system data rather than application code. Yet much of the industry’s testing investment remains focused on source code rather than system-level failures. Verification techniques such as property-based testing, executable reference models, and deterministic simulation have traditionally required significant time and resources, limiting adoption outside organizations such as Amazon Web Services (AWS) and Google. As artificial intelligence (AI) reduces some of those barriers, engineering teams have an opportunity to expand testing beyond code and place greater emphasis on system verification.

Complexity requires system verification over code verification

Complex systems have non-linear feedback relations between system elements, creating uncertain dynamics that static risk assessments do not readily capture. The more complex the system, the more difficult risk detection becomes. GitHub is a prime example of a complex distributed system. GitHub repositories and services host most modern software development tools, with an average of over 43 million pull requests each month globally in 2025. Most issues in these systems are not related to code but to deployment and dependency issues. Figure 1 demonstrates that software bugs, labeled as internal, are secondary to external dependencies and deployment issues, which scale with complexity.

Now Cleared for Failure: Tunnel Vision in Software Verification

Figure 1: Incidents by primary cause. Source: Data, GitHub outage analysis May 2025-April 2026; Graphic courtesy of Sergei Cherednichenko.

Such systems are susceptible to metastable failures that are particularly difficult to guard against and recover from. They occur when a distributed system becomes trapped in a degraded condition by a self-sustaining feedback loop, such as repeated retries that increase load and cause further failures, even after the original trigger is removed.

The February 2, 2026, GitHub outage caused virtual machine (VM) operations to fail across all regions, as a storage policy access change triggered a cascading metastable failure. GitHub identified additional failures and remedies in its May 2026 availability report. Remediation included infrastructure steps to increase monitoring and failover testing, break monolith code into isolated services, improve elastic capacity, and shift production testing toward system behavior verification to complement code verification.

Integrated production is system production: Metastable failures in Cloudflare

In November 2025, a Cloudflare feature file doubled in size, automatically shipped, and then propagated across servers. Upon failure, the configuration file deployed and pushed servers into a buffer fail-state, triggering a metastable failure. At its peak, it caused a three-hour service disruption during which 26 million 5xx error status requests per second were served. After Cloudflare provided a known-good file, it still took two more hours to fully restore because the services had to be manually restarted.

Remediation steps included hardening configuration files through system testing, adding feature-level global kill-switches, setting system resource consumption limits on error reports, and testing error conditions across all core proxy modules. All system-level remediations aligned with metastable failure research suggestions to solve root-cause structural issues, including feedback loops, not just triggers.

Automated traffic from crawlers and AI agents has overwhelmed systems developed, tested, and adapted to human behavior. As system behavior changes, it’s critical for testing to also adapt to verify new behavior constraints. Distributed systems and site reliability engineering (SRE) offer ready tools and best practices for these emergent risks.

Engineering constraints and testing: Site reliability at Google

Google’s approach to SRE is software engineering with systems engineering skills: “SRE is a specific implementation of DevOps with some idiosyncrasies.” Google caps SRE operational task time at 50%. If more operational task time is required, developers are looped in on operations bugs to cross-train until the task drops below 50% operational load. This simple task-allocation limit enables a durable focus on engineering and ensures that availability, latency, monitoring, and response issues are embedded in the development team’s work.

Target reliability and error budgeting are metrics that guide SRE teams to meet service level objectives (SLOs). This allows teams to understand the service’s or product’s tolerance for unreliability, emphasizing that 100% reliability is rarely desired, as most local hardware and internet service providers (ISPs) operate at 99% or 99.9% reliability. Improving beyond this level yields diminishing returns and is not usually perceptible to end users. Accepting a certain number of errors from known failures helps set operational flags and provide early signals for review.

This risk-management practice translates well to management by demonstrating downtime and risk exposure for feature failure. Google maintains high reliability while pursuing incremental innovation by using the SLO to control the release rate.

Approaching testing as site reliability enables teams to define development objectives and SLOs, which they then use to design and monitor the operational aspects of the products and services they offer. As seen in the Cloudflare and GitHub outages, features did not cause the metastable failures. Operational changes degraded service and led to the outages.

Simulation and property-based testing

Property-based testing (PBT) is a functional programming technique that specifies a function’s output and then stress-tests it using randomly generated domains. The Python libraries Hypothesis and its extension HypoFuzz are PBT tools that enable standard library and customized PBT with fuzzing on specified functions. While unit tests state how to test, properties specify what to test. For example, this short Hypothesis property-based test snippet ensures that configuration files don’t exceed supported limits (Figure 2):

Now Cleared for Failure: Tunnel Vision in Software Verification

Figure 2: Python-Hypothesis PBT example for a config file feature size. Graphic courtesy of Sergei Cherednichenko.

Rapid simulation of production environments is another tool that large language model (LLM) code generation has made more accessible. These system simulations enable developers to build and maintain discrete system reference models to simulate deployment and use PBT to test system behavior. Together, simulation and PBT can test for metastable failures similar to the GitHub and Cloudflare outages.

LLM tools are democratizing access to these resource-intensive system verification tools. Full system simulation is possible, but the cost-benefit of engineering, criticality, and maintainability still needs to be weighed against risk. Feature or product releases with high-risk failure scenarios that can damage reputation, such as metastable failures, lockouts, and outages, are better candidates for simulation. Alternatively, certain critical subsystems can be maintained in a sandbox as a stable-state alternative-infrastructure simulation. This service-isolation approach to system simulation may enable testing that targets the highest-risk interfaces while minimizing overhead.

From specification to reliability

A research team evaluating LLM agents in software development suggested viewing software correctness as a distribution of outcomes rather than a binary property of code. Binary testing is still an essential step in quality engineering but is no longer sufficient for deployment. Testing to distributions of specified outcomes in a system-of-systems is the new quality paradigm.

Software testing can only verify what it knows to measure; unit testing is limited to discrete input/output (I/O) coverage, while PBT can broadly test domains. Formal specification enables stress testing products in simulation beyond the standard production battery. Adapted development pipelines sharpen human value-add and improve documentation while taming AI development capacities toward properly bounded development targets. Specification pipelining can enable fuller coverage in quality testing and provide metastable resilience across distributed systems.

Adopting an SRE mentality in production can free traditional software testing habits from developer-generated unit-test thinking and secure the resiliency that large service providers build their names on. Incremental use of these techniques creates opportunities for development teams to achieve greater code coverage and fault testing in their continuous integration and continuous delivery (CI/CD) processes, while providing core documentation to support enriched experimentation and capability permutations.

About the Author

Sergei Cherednichenko is a senior software engineer with 11 years of experience in developing SaaS products, distributed systems, and testing solutions for startups and internet-scale companies. He holds a Master of Science in computer engineering from State University ITMO of Saint-Petersburg, Russia. Connect with Sergei on LinkedIn.

Be the first to comment

Leave a Reply

Your email address will not be published.


*


This site uses Akismet to reduce spam. Learn how your comment data is processed.