Finding a dangerous software flaw is getting easier. Proving that it is real, preparing a repair and getting it to users remain difficult. In his Black Hat USA 2026 keynote, Arizona State University researcher Yan Shoshitaishvili described a laboratory that, by his estimate, discovers vulnerabilities roughly ten times faster than it can prepare disclosures.

This is a gap between discovery and protection. For a developer, an AI finding becomes useful when it explains the conditions that trigger a flaw, what an attacker could do and how a repair closes that path. A warning count does not provide those answers.

Why “find bugs” is an incomplete task

A language model can read code and suggest dozens of suspicious locations. Some will be invented; others may be real defects. Even an accurate list covers only the kinds of flaw the model recognises under that task definition. Other kinds can remain unseen.

Shoshitaishvili distinguishes three ways to increase findings: improve an analysis method, inspect more programs, or use a different method on the same code. The last often uncovers what the first method never looked for. A new tool with more findings cannot immediately be declared better: the order of inspection also affects the result.

For example, fuzzing tests a program with many automatically selected inputs and can expose cases that crash it. Code analysis can uncover different problems. A model may also know an old bug from training material; finding a familiar defect is then weak evidence of performance on unknown vulnerabilities.

Define the vulnerability before searching

The laboratory’s approach centres on vulnerability properties. A researcher describes the relationship between data, permissions and program actions that produces a flaw. For some defects, the source of data matters; for others, whether an attacker can influence a command; for others, what happens when two actions run at the same time.

This gives an agent a specific task. An agent here is a model that uses tools, examines a program and tests hypotheses. A human defines which conditions make a finding dangerous and evaluates the evidence.

What Android flaws revealed about OpenHarmony

The laboratory applied this approach to OpenHarmony, an open-source operating system. In a USENIX Security 2026 paper, the team reviewed 116 Android publications, extracted 56 unique design-level vulnerabilities and showed that OpenHarmony was vulnerable to 24 of them. Consequences included privacy breaches, stealthy privilege escalation and reliability problems.

In the keynote account, agents helped review the research history, while vulnerability properties and OpenHarmony code were inspected manually at this stage. That distinction matters: the published result cannot be attributed entirely to an autonomous model. It shows that a new implementation can retain an old design flaw without copying its source code.

How the workflow changed Linux results

In Shoshitaishvili’s account of Linux kernel testing, several dozen models initially found about 300 potential local privilege-escalation vulnerabilities. These are flaws that let a user without administrative permissions gain higher privileges on the same computer. That is a specific threat model, rather than any kernel defect.

The team then connected three models through a workflow with planning and cross-checking. The count rose to about 600; the speaker says those findings were reviewed manually. Adding properties from historical and fresh vulnerabilities pushed the reported count above a thousand.

These figures describe the laboratory’s experience. The keynote does not provide conditions that make the sequence an independent comparative model benchmark. Shoshitaishvili himself warns that his counting criteria differ from public figures for another system. The supported conclusion is narrower: how the search was organised substantially affected its results.

Why discovery has not yet delivered protection

A disclosure should include a substantiated analysis and a proposed fix. With discovery running at what Shoshitaishvili estimates to be ten times the disclosure-preparation rate, the backlog becomes a separate problem. The laboratory refuses to send an unprocessed stream of findings.

In his account, details are published for flaws with assigned CVEs, the commonly used vulnerability identifiers. For the rest, the team leaves hashes, short digital values that record findings. The team is discussing temporary patches with manufacturers; these patches might help before the main update. The researcher has no finished replacement for the customary disclosure process.

What a Rust rewrite can retain

Rust helps prevent many memory-access errors. A program can handle memory correctly while still checking permissions incorrectly, mishandling state or implementing a cryptographic algorithm wrongly. These are logic flaws.

Shoshitaishvili asked agents to rewrite widely used C libraries, including libssl, libpng and libxml. The resulting code built and ran. Yet known logic flaws reappeared even when agents were explicitly forbidden to carry over old defects. Subsequent checks found what the instruction had failed to prevent.

The subsequently published SafeLibs project offers compatible Rust replacements for Ubuntu libraries. Its site states plainly that the code has not been read by humans and translation can introduce new non-memory bugs. This is a research experiment with substantial limits, rather than a validated secure replacement for production libraries.

The work that remains for researchers

Shoshitaishvili favours access to powerful tools for defenders and considers cybersecurity education especially valuable now. This is his position: agents expand how much a specialist can do, while task definition and evaluation require knowledge of the system.

He closes by returning to research prototypes. An agent can run an experiment, observe failure under new conditions and help change a program. A human sets the test conditions and decides what counts as evidence. In vulnerability research, that boundary is decisive: a security assurance requires testing the path an attacker can actually use.