AI can give security teams more ways to test an environment. The useful question is what we do with what it finds.

The capability shift deserves our attention. In its April 2026 research, Anthropic reported that Claude Mythos Preview could discover and exploit software vulnerabilities after an initial prompt, including weaknesses in widely used operating systems. These are the developer’s reported results from specific testing conditions, not a guarantee of performance in every enterprise environment.

For Canadian financial services, the concern extends beyond a more capable testing tool. OSFI’s frontier AI bulletin describes a threat environment where vulnerability discovery and exploitation can move faster, putting pressure on established patching and response practices.

Defenders should use the capability, too

Our response should include applying these capabilities to our own authorized security testing. AI can help investigate weaknesses and explore possible attack paths. We should evaluate where it adds coverage and repeatability, rather than assume either that it replaces existing testing or that existing testing is enough.

I see this as a way to extend the reach of security teams. Give experienced people more capacity to test, investigate, and validate, while retaining their responsibility for scope, interpretation, and the decisions that follow.

The advantage is knowing our environment

A model may identify a technical weakness. Our teams know which service supports a critical business process, where sensitive data sits, which privileges matter, and which dependencies make a change difficult. Some of that knowledge can be supplied through inventories and system documentation. Some requires a conversation with the people who operate the service.

Consider two apparently similar findings. One affects a disposable development workload. The other opens a route into a service that clients depend on. The technical description may look alike; the urgency, consequences, and remediation approach can be very different.

More testing capability becomes more valuable when it is connected to a better understanding of what we are protecting.

Build a loop from testing to action

The path forward is a continuous loop: identify exposure, test meaningful scenarios, validate the evidence, connect it to business impact, and act. After remediation, test again. Feed what we learn back into engineering patterns, access controls, monitoring, and future testing.

This also means evaluating the testing capability itself. Did it discover something useful? Was the finding reproducible? Did it stay within scope? Did the result change a decision or improve a control? A convincing report is not the same thing as demonstrated risk reduction.

Protect the service while testing it

Testing needs explicit authorization, constrained access, agreed targets, and operating limits. Where an activity could disrupt a service, use an appropriate test environment or a tightly controlled exercise. A tool’s ability to act autonomously should not determine how much authority we give it.

We also need a credible route from discovery to remediation. Fixing a serious exposure may require an urgent change; it may also require staged deployment, temporary containment, or coordination with a provider. The decision should account for both the risk of compromise and the risk of interrupting the business.

My position is that we should move deliberately, with urgency. Combine frontier AI’s technical capabilities with our knowledge of the environment, and measure success by the exposures we remove and the services we keep dependable.

References & further reading

Primary guidance for the concepts and controls discussed here.