After more than 300,000 production penetration tests (pentests), our company has learned something that may surprise people watching the recent wave of autonomous security announcements.
The hardest problem in autonomous security isn’t teaching a machine how to attack. It’s teaching an AI-based system how to operate safely, predictably, and repeatedly inside production environments where mistakes have consequences.
Finding an attack path is an engineering problem. Building a platform that organizations trust to operate against healthcare systems, financial institutions, manufacturers, and critical infrastructure is an operational one. The difference only becomes apparent after years of running at scale.
As the industry embraces AI agents, autonomous red teaming, and machine-speed operations, much of the conversation remains focused on capability. Can a machine identify a path to compromise? Can it chain weaknesses together? Can it achieve the same outcome as a human operator?
Those are reasonable questions. They are not the questions security leaders ultimately care about.
Security leaders need confidence that a platform can operate safely in production, consistently produce meaningful results, and help teams make better decisions about risk. In our experience, that’s where the real challenge begins.
Since 2019, NodeZero® has executed more than 300,000 production pentests across thousands of environments. Those engagements have reinforced a lesson that continues to surface.
The biggest security challenges rarely come from what organizations cannot see. They come from separating signal from noise.
Most organizations are not struggling to find vulnerabilities
The security industry has spent decades improving visibility.
Organizations have vulnerability scanners, attack surface management platforms, cloud security tools, exposure management programs, and countless dashboards filled with findings. Most security teams are not suffering from a lack of information. This issue is: They’re struggling to determine which information matters.
Attackers do not think in terms of individual findings. They think in terms of outcomes. They identify a weakness, combine it with another weakness, move through the environment, and pursue an objective. The path matters more than any individual step along the way.
Security teams often inherit the opposite problem. Thousands of findings arrive in a dashboard, each evaluated independently, with little context around how those weaknesses might connect. As a result, teams spend significant time debating severity while attackers focus on exploitability.
The difference sounds subtle, but it changes everything. Severity describes a vulnerability. Exploitability describes risk.
Experience changes how you evaluate risk
Trust isn’t built on promises, it’s built on the deep experience gained from executing hundreds of thousands of pentests. Over time, recurring patterns begin to emerge regardless of industry, technology stack, or organizational maturity.
We’ve seen organizations trust legacy tools that require enormous effort to remediate vulnerabilities that had little practical impact, while overlooking seemingly minor weaknesses that ultimately enabled significant compromise. That happens because risk rarely exists as a single vulnerability. It exists in the way weaknesses interact with one another.
In a financial services environment, a single compromised credential led to 586 critical impacts across 115 hosts, including three separate domain compromises. Viewed independently, the credential did not appear particularly significant. Viewed as part of an attack path, it became something entirely different.
In another cloud environment, the path to full Entra ID tenant compromise did not require a common vulnerabilities and exposures (CVE) or zero-day exploit. The weaknesses involved were already known. Existing tools had identified them. What was missing was an understanding of how those weaknesses could be chained together and what that chain of events meant for the organization.
We have also seen organizations discover that the initial compromise was not the most important part of the assessment. In one education environment, the larger question was how far an attacker could move after gaining access. Measuring blast radius exposed paths to systems and data that were never expected to be reachable from the original point of compromise.
These examples reinforce the same lesson. The challenge is rarely finding weaknesses. The challenge is knowing which weaknesses matter before an attacker does. That kind of judgment isn’t built from demonstrations or benchmarks. It’s earned through years of operating in production environments and seeing how real attack paths emerge across thousands of organizations.
Click here to learn why reliability is the key to building trust and how to get there.
What 7 years of autonomous pentesting taught us
Most organizations do not need another source of findings. They already have more findings than they can realistically address. What they need is confidence in what is actually exploitable, how attackers would use it, and whether their fixes reduced risk.
That’s the lesson we’ve learned from years of operating in production environments.
And it’s the problem we have solved.










