Haxact

Buyer's Guide · 16 min read · September 28, 2026

How to Evaluate an Autonomous Pentester

How do you tell the difference between a world-class agentic hacker and a chatbot with a terminal? Here is what you should ask vendors when looking for security assessments.

If you are reading this, I assume you are familiar with scanning: cheap, fast, but shallow, a quick way to get a sense of your baseline security. And you are probably familiar with pentests: expensive, slow, but thorough engagements where you book a window for hackers to attack you, then get a report with their findings. Those two extremes have left a lot of companies thinking a good level of security is something they cannot afford. With AI, that is changing. Agentic pentesting combines the best of both worlds: a thorough assessment at a reasonable price, whenever you please, with the findings written up in a couple of days at most.

AI is proficient enough at cybersecurity that the model providers now gate it: OpenAI and Anthropic both keep offensive security work behind vetted access programs, and the default safeguards are broad enough that they regularly catch legitimate security work too. It is easy to see why they are careful. In a Stanford and CMU study, an autonomous agent (ARTEMIS) out-scored nine of ten human professionals on a live enterprise network. In DARPA's AI Cyber Challenge, seven autonomous systems set loose on millions of lines of real-world code each turned up a genuine, previously unknown bug in the process, not just the vulnerabilities the organisers had planted. Google's Big Sleep agent has found real zero-days in shipping software, one of them just before attackers could use it. In the hands of a skilled hacker, AI can find and exploit vulnerabilities at a speed no human can match.

Of course, deploying AI so it reaches its potential as a hacker is not trivial. It needs the right environment, custom tooling, a harness that holds up, and a lot of small details that are easy to overlook. Assembling something that looks like an autonomous pentester is easy; getting one that runs a thorough pentest is hard. We wrote this guide to help you weigh what vendors are offering, so you can tell who can actually hand you a pentest report worth paying for.

This guide helps you filter through the vendor feature lists and gets to the small set of things that separate a real autonomous pentester from a chatbot with a terminal. They fall into four areas, in roughly the order they rule vendors out: whether the work is any good, where your data goes and whether it will satisfy an auditor, the logistics of living with the product, and whether you can trust it to operate against production. As a companion, the Autonomous Pentester Scorecard turns these into a checklist you can walk through vendor by vendor. And to see what you should be getting, here is a redacted report from a real engagement of ours to use as a reference.

Agentic work quality

Everything in this first area answers one question asked four ways: is it doing real testing, and doing it well? This is where most of the field falls out, so it is where the most time is worth spending.

Can it actually reach your application?

Let's start with the basics. A lot of tools that call themselves autonomous pentesters are command-line agents: they fetch a URL, read the response, and reason about the text. That works against an old-fashioned, server-rendered site with a form and a database. It falls apart the moment your application is anything modern, with stateful and complex logic.

If your app relies on JavaScript for its functionality (like a lot of single-page apps do), a tool that cannot drive a real browser will not be able to interact with it properly. It will either see an empty page, or fail to see the state change as a result of its own actions. Even worse, if the login is a multi-step, script-rendered, potentially MFA-gated flow, a command-line agent stalls at the front door while everything worth testing sits on the other side of it. How do you check for login links, reset passwords, and discover account-takeover vulnerabilities without access to a mailbox? A vendor should be able to explain how their agent works, and show you it can handle a modern, authenticated application, not just the public shell.

Another thing to consider is mobile. A good mobile app is practically essential if you want to reach a lot of customers, but most AI pentesters do web and API only, so if you maintain an Android or iOS app, those agents will never touch it. This is a genuine coverage cliff, because a "clean" report from a tool that does not test your mobile app is not a clean report. A hacker will absolutely check out your mobile app, so if you have one, make sure to procure security testing from a vendor who actually offers it.

Finally, and most critically, a good pentester should understand your application's business logic. Hacking does not always mean sending a request the system is not equipped to handle; it is often about finding and abusing unintended privileges to achieve a goal the app was never meant to allow. Your app could let users read each other's data. It could expose sensitive information on a public URL. There could be a way for anyone who can identify an admin account to impersonate it and exfiltrate everything. Finding business-logic issues takes an agent that can hold and understand state, chain requests, and keep a good idea of what your application is meant to be doing. A company can say it has a million different agents testing your application; if all of them finish within a dozen requests or so, what you get is the analytic depth of a scanner, which is basically nil.

To understand what level of complexity an agentic pentester can offer, ask for a report against a realistic target. The complexity of the vulnerabilities it discovers should tell you how proficient it is at pentesting. In the sample report, you should see the agent find more than just IDOR issues provable with a single request; you should see complex, multi-step exploitation flows that chain multiple issues to expose critical vulnerabilities. Be wary of a vendor that waves a benchmark score at you instead: public benchmarks are easy to game, and some of their challenges are broken, so a self-reported score tells you far less than a redacted report from a real engagement.

Does it lean on your source code?

Some vendors ask that the agent gets access to the source code. They may even price the pentest differently based on whether you give the agent access or not. While a good pentester will not turn down source access, an agent that requires it, whose findings would drop in quality without it, is not really a pentester; it is a code reviewer. And you can run code reviews on your own.

A pentester is meant to find issues beyond what the source code can convey. Did the code get deployed properly, or was some critical configuration option left out? Is there some business logic issue, justified with a comment in the code, that enables a much more severe class of abuse than previously imagined? Are the specifics of how the third-party libraries work well understood? A lot of logic issues can be missed with just a review of the code. Not to mention issues that reviewing the source might never reveal, such as cache poisoning. You should expect to be able to provide an agent with the source code, but a good pentester should not need access to it to deliver quality work.

Can you trust the findings, and can your team act on them?

Large Language Models (LLMs) have gotten better at reasoning and at avoiding hallucinations, but surfacing legitimate findings is still not trivial, and producing a finding is not the same as finding a problem. The generic sort of agent is weak here: in that live enterprise-network study from Stanford and CMU, the agent's headline weakness was a higher false-positive rate than the humans. So "it found 200 issues" is not indicative of a good pentester. If those 200 issues are unverified, unsubstantiated notes that a human has to manually double-check, a human might as well go through the app themselves. And if most of them turn out to be false positives, the human might even be more efficient.

To mitigate this, you have to demand proof from the agents. A finding nobody has proven is a hypothesis: this parameter looks like it might be vulnerable. The only thing that turns a hypothesis into a fact is exploiting it. If the tool actually performs the attack and hands you the evidence, there is nothing left to argue about. This is the line NIST draws in its testing guide, SP 800-115: a scan flags what might be wrong, a penetration test mounts the actual attack and proves it. Agents should be able to give step-by-step reproduction instructions, material evidence that they carried out the testing (screenshots, request and response pairs), and an explanation of what makes the finding a legitimate issue to fix, alongside instructions for exactly how to fix it.

Without proof, a finding requires triage: someone still has to confirm it's real. With the request, the response, and the steps to reproduce it, an engineer can fix it straight away.

There is a second reason to demand proof, and it is specific to AI. A model that wants to be helpful can invent a plausible vulnerability out of thin air, and it can also be misled by the target: when researchers fed autonomous agents deliberately deceptive responses in a 2026 study, ATOBench, they found a busy agent can look productive while its verification has quietly broken, and the ones that held up were those that recovered real evidence and carried it through to the report. You should expect every reported finding to include substantial evidence and reproduction steps; that way you can understand exactly what has happened, and can easily tell whether a finding is real. Having AI tell you "here is the exploited issue" instead of "I think something might be wrong" is a big deal; the first is a useful report, the second is noise you already have too much of.

The remediation instructions matter just as much. A report with no actionable next steps is hard for engineers to use: the ticket with a real, reproducible exploit sits in the backlog a month later, waiting for someone to figure out what to do about it. A finding worth paying for closes the loop, pairing the proof with concrete guidance on the fix, so your team gets something they can act on rather than one more confirmed problem to schedule. Ask what a finding actually contains, because "critical severity" with a paragraph of description is a headline, and your team fixes the body copy.

What are you actually paying for?

The pricing of agentic pentests can be hard to compare. Most providers are not confident about estimating token usage for a pentest, so they will do one or more of four things to obscure what you are actually paying for:

  • Tell you the cost, or part of it, is variable based on token usage, or that you can bring your own key.
  • Cap the number of tokens, or the amount of time, that agents will spend on your application.
  • Give you credits to spend that act as a proxy for agent effort: the more credits you spend, the more tokens your pentest agent can use.
  • Bundle scanning in with the pentest, so you also get some non-LLM functionality. That part is usually "unlimited," because it is much cheaper to provide.

Some vendors also offer different "tiers" of pentest, and gate features or scope behind them. There might be one price for a web app test, but a higher tier if you also want to pentest your mobile apps.

Either way, you should be able to understand what you are paying for and how much you get for the money. That way you can form reasonable expectations about what a vendor's pentest might find in your app, because you know how simple or complex it is. You want to be confident the agents will have enough budget to cover your apps thoroughly, and that everything you care about can be in scope for the tier you choose. Vendors should be able to give you practical limits, because numbers without a reference point (a thousand agents, ten million tokens) tell you nothing at all.

Data and compliance

If you're going to trust a vendor to find your weaknesses, you need to know that you can trust them to handle your secrets, and that the pentest result you get will be good enough to satisfy an auditor who asks for it.

Where does your data go?

An autonomous pentester runs on an AI model, and that model is almost always hosted by a third party, so your test data and, potentially, your secrets pass through someone else's infrastructure. For the most part that can be fine: any testing credentials you give the agent should be dedicated to the test rather than production secrets, data in a staging environment is usually fabricated placeholders, and debugging detail from a test run may not carry over to production. The findings are the exception. They are real issues you will have to fix, effectively a map of how to break into your live system, and you do not want that ending up in a training set or read by anyone other than you.

What you need to know is that the model provider has a zero-data-retention policy, or more technically, that no data of yours is ever written to persistent storage on their side. Then make sure the pentest vendor itself does not use your data for any reason other than to deliver their service, and that any and all of it is deleted if and when you choose. The consequences of mishandling data are severe, and introducing third-party risk from an assessor meant to help you minimise it would be extremely unfortunate.

Will it satisfy your auditor?

While we in the cybersecurity industry like to believe that everyone shares our passion for secure software, clients often seek a pentest to satisfy a compliance requirement. Some require it outright. If you take card payments, PCI DSS names a penetration test in requirement 11.4. Others, such as SOC 2 and ISO 27001, expect it without naming it. ISO 27001 comes at it from two angles: Annex A control 8.8 on managing technical vulnerabilities, and control 8.29 on security testing in development and acceptance. The Trust Services Criteria behind SOC 2 go further still, naming penetration testing as one example of the "ongoing and separate evaluations" management runs to check its own controls (criterion CC4.1). For anyone holding data on people in the EU, GDPR's Article 32 requires "a process for regularly testing, assessing and evaluating the effectiveness" of your security, and HIPAA's Security Rule similarly calls for a periodic evaluation of your safeguards; neither names a pentest, but one is the clearest way to show you are doing it.

Sometimes, the toughest assessors are not auditors looking to certify you for a standard, but your customers and business partners. Part of the due diligence process of procurement could involve asking you for a recent pentest report. This helps businesses see that their prospective business partner is taking security seriously by actually performing assessments on a regular cadence, and that they do not have any severe issues lingering that they know about. As a bonus, once a company grows to a certain size where cyber insurance makes sense, the insurers will frequently ask for, and favourably consider, a recent pentest that does not show any severe issues present. This can give the company more favourable rates.

One nuance decides whether the report actually closes the requirement. A credible tool follows a recognised methodology, whether that is NIST SP 800-115, the PTES, or OWASP's testing guides, and shows the exploitation rather than a list of theoretical weaknesses. That standard is the same whether a person or a machine did the work. However, certain assessors may expect a human of record, and reject any machine-generated report, regardless of quality. Make sure to consult with your auditor to determine what exactly it is that you need, to avoid spending money on a tool that does not meet the required criteria.

Logistics and transparency

A pentest engagement can be high-quality and inexpensive, but if it is hard for users to manage, it will not provide the value it should. The vendor's portal, or whatever they use to deliver reports, needs to meet some minimum standards of its own.

Can you run it when you need to?

The whole promise of an autonomous tool is that it works like software: you can start an engagement at precisely the times you want, without scheduling a slot, booking a meeting, or being told there is a queue ahead of you. You have shipped a new version and you need it tested now; security should not be the bottleneck. Confirm that you, not the vendor, decide when a pentest happens, and that there are no delays or conditions attached to starting one.

Does it tell you what it didn't reach?

This matters most when tokens are limited during the engagement. Agents, like humans, are not perfect. They may focus on one area and neglect another. If the vendor caps the agents' budget, they may run out halfway through a test. They may not try every permutation, they may find something promising but lack the ability to explore it, or they may simply stop short and leave the job half-done. None of that means the report is useless, but you should never assume a pentest report contains everything there is to find.

For that reason the vendor needs to be transparent about what happened during the pentest. What vulnerability classes were tested? Was testing cut off partway through? Are there gaps the agents left behind, things they did not get to for whatever reason? Knowing what was checked, and how thoroughly, is what lets you decide how to proceed. An honest account of coverage is what makes a "clean" result mean "tested and clear" rather than "never looked": if the report flags no critical issues but you know the agents never reached part of the app, that is your signal to run a follow-up rather than sit on a false sense of security.

What happens when a run breaks?

Lastly, sometimes things go wrong. Maybe the agents crash, maybe something was misconfigured and the run fails, maybe the model provider has an outage. For a service as expensive as a pentest, the vendor needs to detect a failed run accurately and commit to refunding any that does not complete. That commitment is really about the quality of the work: a vendor that takes the assessment seriously is upfront when a run fails or only partly completes, rather than passing off a half-finished test as a clean result. Silence here, or a shrug, tells you how often runs break and who currently absorbs the cost. Security is a sensitive subject, and you want a vendor that handles its own failures professionally.

Operational safety

You are giving an autonomous system permission to attack your production infrastructure, so it needs to be built to minimise the damage it can do when something goes wrong. However rare an incident might be, you want a plan for the worst case. Heavy machinery can do serious harm if mishandled, and an agent set loose on your systems is no different: it has to be easy to control to be safe to use. There are four questions to work through.

Can you stop it?

Can you stop a pentest immediately, with a button that tears down the testing infrastructure? Once you press the "Stop" button, the teardown should begin within seconds at most, and the environment itself should be fully wound down within a minute or so; the faster the better. Any vendor that is slow to stop, or does not give you a kill switch at all, is one to be suspicious of.

Is it contained?

The naive question here is whether the tool will only ever touch your own domains. Testing a modern application means touching what it depends on: the auth tenant on someone else's infrastructure, the storage bucket on a different domain, the API answering from a subdomain you do not directly control. A tool locked to your apex domain cannot test the parts most likely to be misconfigured. So the real question is not whether it stays on a leash but whether it is contained. Each engagement should be walled off in a fresh, single-use environment, with no shared credentials and nothing that outlives the test. One customer's assessment should never be able to reach another's data or wander into unrelated systems.

Containment is also what makes the model's fallibility survivable. A vendor who reassures you that they "screen inputs for prompt injection" is selling a filter that does not hold. The better answer is about blast radius: a well-built agent that is manipulated still cannot fabricate a finding, because of the proof rule, and still cannot reach another customer, because of isolation. That makes injection a contained nuisance rather than a catastrophe, which is the honest bar.

Can you tell its traffic apart?

You want to be able to look in your own logs and know which requests came from the hacker agent. A tool that tags every request with a known marker, or exits from a known address you can match, gives you a record you can check, rather than a promise to take on trust. The same trail is what makes the engagement auditable. Every finding should be traceable to the exact traffic that proves it, and you should be able to keep an eye on the pentest as the agent explores your infrastructure, to differentiate it from any other questionable requests.

Does it work with your defences on?

Does the tool need you to switch your WAF and bot protection off before it can work? The answer should be no. A real attacker does not get your defences disabled as a courtesy, so a test run against an undefended target is measuring a lab rather than the application your users actually reach. An agent will always look, and be, more effective with the filters off, which is exactly why a result obtained that way is suspect: if a real attacker cannot get past your WAF, a finding that only lands with it disabled is a latent weakness rather than an exploitable risk. A capable agent can get through many of those controls anyway, and watching it do so against your live defences tells you more than a clean report earned by switching them off.

An honest summary

The range between tools is enormous. Generic agents are weak: PentestEval found off-the-shelf LLM pipelines succeeding about a third of the time, and fully autonomous ones barely at all. The best purpose-built scaffolds are strong: in the ARTEMIS study, the top agent matched all but one of ten human professionals on a real enterprise network, winning on systematic enumeration, parallel exploitation, and cost. The same holds beyond pentesting, where the autonomous systems in DARPA's challenge and Google's Big Sleep agent have both been finding and fixing real, previously unknown vulnerabilities in production software.

So the gap is mostly one of engineering, which is exactly why you should evaluate vendors and their offerings carefully. Agents, like humans, are fallible. They produce false positives at a higher rate than a careful human would, which is why validation is the thing that matters most: a good tool proves each finding with a working exploit and drops whatever it cannot, instead of handing you a pile of maybes to sort through. They can also miss things entirely. Weaknesses in a particular library, framework, or piece of infrastructure may be beyond what they are equipped to spot or exploit, and those blind spots are false negatives: real issues that never make it into the report. And the floor is low: a poorly made product underperforms badly. A determined human expert is still more persistent than an agent tends to be, willing to grind at a single stubborn target long after a model would have moved on.

None of that makes the good tools a bad buy. It makes the difference between them worth your attention. The best do most of the work, prove every finding, and own up to what they missed. The weak ones lean on the category's reputation and hope you never look too closely. Everything above is how you tell the difference.

If you think we have a point, take these questions to your vendors.

Get the vendor evaluation worksheet →

Where we stand

A word on where Haxact sits, since you have read this far. We built it around the work-quality area, because that is the part that decides everything else: every finding it reports arrives with a working, reproducible proof-of-exploit and concrete remediation. It drives a real browser, so it tests single-page apps and real login flows, and it covers Android, not only web and API. You run it yourself, on demand, the way you run any other tool, and a run that fails on our side is refunded. Each engagement runs in its own ephemeral environment with identifiable traffic, so it is contained and you can find it in your logs. A stop takes effect immediately and the environment is gone within the minute, and inference runs under a zero-retention, no-training agreement, so your data is not kept or learned from beyond the engagement. We do not claim that our service will be the best for your use case, that our agents would beat any human expert, or that pentesting is a solved problem. What we promise to do is hand you qualified findings extracted from a thorough, no-token-limit security assessment.

If you have any questions, we'd be happy to help you get answers. No obligation. No strings attached.