Agents Don’t Have Favorite Attacks: Two Ethical Hackers on Agentic Pentesting

Agents Don't Have Favorite Attacks: Reflectiz founders webinar on agentic pentesting
Share article
twitter linkedin medium facebook

By their own admission, Reflectiz CEO Idan Cohen and CTO Ysrael Gurt could talk about hacking all day long. Yet after years running the company together, this webinar is the first time they’ve done it on camera, and the timing is no accident: AI pentesting agents have arrived, for attackers and defenders alike.

Both men built their careers breaking into systems for ethical reasons. Idan led the BugSec hacking group and ran product and R&D at Cynet before Reflectiz. Ysrael’s record includes hacking Google, Facebook, and Microsoft, plus three XSS bugs in Gmail. Google’s Hall of Fame ranked him in its top 25, and Forbes Israel named him to its 30 Under 30. As he puts it, he’s “an expert in doing it the wrong way.”

Then they built a company to defend others against less scrupulous attackers. After a decade of agentless web monitoring, the founders have returned to their roots with Offensive Hub, Reflectiz’s agentic pentesting solution. In just under an hour, they cover what it has already found, why it works differently from a human tester, and what you should ask any vendor in this suddenly crowded market. Here are a few highlights.

TL;DR
  • An AI pentesting agent found a critical authentication bypass on a major insurer’s live website, then stopped at its guardrails.
  • Agents don’t have favorite attacks. They run the full test plan, while human testers lean on the methods they know best.
  • The harness matters more than the model, and the cheapest model per token isn’t always the cheapest to run.
  • Test critical flows such as login in production, with guardrails. Attackers’ agents already do.

A critical login flaw

Offensive Hub is already finding serious flaws on customer sites. One stood out: a critical authentication bypass on the production website of a major international insurance company. Left unchecked, it would have let anyone log in as any customer and read their data.

“The sheer existence of an authentication bypass in a financial website in production is something that was supposed to be almost impossible,” says Ysrael. “But today everything is possible.”

The agent’s analyzer read the site’s JavaScript and spotted that the request triggering a one-time passcode (OTP) let the user control where the code was sent. Its planner built a dedicated attack around that. Drawing on tooling Reflectiz has built over a decade to receive OTPs by SMS and email, the agent tried it, and the login succeeded. Crucially, it stopped there because its guardrails kicked in.

Humans have favorites. Agents don’t.

In one way, it’s no big deal that the agent found the flaw. A good human pentester might have done the same. What is a big deal is that it was running four other attack plans too. Humans have favorite methods of attack. Ysrael’s was always XSS, but an agent doesn’t play favorites. It builds the full test plan upfront and runs every item. Nothing gets skipped.

Attackers don’t book a test window

The founders also discuss the time you actually pay for. Idan explains how a typical ten-day engagement often shrinks to six or seven days of real testing. And because tests are booked weeks ahead, a few times a year, new pages can sit waiting to go live.

But attackers’ AI agents have no schedule. As a defender, you need continuous agentic pentesting that covers everything. An attacker only needs one flaw to cause damage, so Offensive Hub maps every page and link, then runs every check in its playbook.

Years of groundwork. A few months to build.

Idan and Ysrael had toyed with automated pentesting for years but never believed it would deliver usable value for customers. The old tools fired off thousands of results, and a human sifted through the noise to find two real ones. Then Ysrael tried the latest models. He saw that with the right guidance, an LLM could actually understand what it was looking at and decide what to test. A few intense months later, Reflectiz had its agent.

The hardest part, they argue, isn’t sending a payload; any LLM can do that in seconds. It’s getting around a real website, logging in, and handling one-time passcodes and 2FA. Offensive Hub does that in a real browser, reaching the business logic that surface-level tools never see. Reflectiz has been scanning websites at scale for a decade, so that groundwork was already done. “Nobody has agentic pentesting with years of experience,” says Idan.

Offensive Hub sits alongside Security Hub, Privacy Hub, and the PCI Module, so a CISO can see web risk in one place rather than across seven tools.

Which LLM? The wrong question

The most common question in Reflectiz sales meetings is which LLM Offensive Hub uses, but Ysrael thinks this is the wrong question. An agent is a brain (the model) inside a body (the harness). Reflectiz isn’t trying to beat OpenAI at building brains; it’s building the best harness that any model can plug into.

Models change almost weekly now, and the latest isn’t always the greatest for every task. Idan often finds Sonnet does a particular job better than Opus, so Reflectiz benchmarks model by model, task by task, and picks the best fit for each. Ysrael adds that the cheapest model per token isn’t always the cheapest to run. A model that’s cheap per token but burns through a million tokens to scan a single page could end up costing ten times more than a pricier model that does the same job in 10,000 tokens.

Why not point a frontier model at your own site? The founders say go ahead and try. It might find five real vulnerabilities, but what about the other eleven still waiting for someone else’s agent? Without the right harness, you get findings, not coverage.

Production or pre-production?

Pentesting traditionally happens in pre-production, at night or on weekends, and many CISOs balk at the idea of letting an agent loose on a live site. Few relish having to tell the board why the website went down for two hours in the middle of shopping season, but as Idan points out: “You’re already running dozens of AI agents in production. You’re just not getting the results.” His point: attackers already run agents against your site, but they won’t stop at two hours, and they won’t write you a report afterward. Better to have a trusted expert do it, with guardrails.

Reflectiz can start in pre-production and validate a finding against production in one click. But the founders recommend testing in production for critical areas such as the login page. They predict that within a year, everyone will be there, with armies of agents attacking and defending every site. (If you think anti-bot protection will keep attackers’ agents out, listen in to hear how long it took an agent to get past one.)

What to ask before you buy

With new agentic pentesting vendors raising money every week, the founders close with a buyer’s checklist. Here are their five questions, with what a strong answer looks like.

  1. Was it built for the web? It drives a real browser through logins, one-time passcodes, and multi-step flows, not raw requests.
  2. What enforces coverage? A full test plan built upfront, with every item run and logged, so you can see what was tested.
  3. Can you dial the cost? You set the depth and frequency per site, and cost tracks what you actually test.
  4. Where are the guardrails? Configurable limits on what the agent can do, and a hard stop once it proves a flaw.
  5. Can you act on the results? Validated findings with fix guidance your developers can use, not a PDF of raw output.

They work through each in the webinar, with real examples of the tools that fail them.

What AI pentesting agents still can’t do

Ysrael and Idan don’t claim humans are finished. When it comes to attack, agents still lack the spark of intuition: the hunch a very experienced tester gets from seeing unrelated things and somehow knowing they offer a way in. Agents also can’t sit down with a developer and ask: “If I could transfer money this way, would you allow it?” Without that conversation, an agent doesn’t get the full story of the business logic it’s testing. That, they say, is the next level, and they hint it may not be far off.

As for human pentesters, the founders expect them to move up a level: guiding agents, answering their questions, and overseeing twenty websites instead of one. Rather than producing another PDF report, they’ll have time to work with developers on actually fixing things.

Watch the full conversation

This is only a taste. In the full webinar, Idan and Ysrael go deeper into how the recon, analyzer, and planner sub-agents work together, how the guardrails are designed, the anti-bot story, and plenty more from two people who’ve spent their careers doing things the wrong way, for the right reasons.

FAQs

Can AI pentesting agents replace human pentesters?

No. Agents cover more ground than one tester, but they still lack a senior tester’s intuition, and they can’t ask developers how the business logic should behave. Reflectiz’s founders expect human pentesters to move up a level: guiding agents, answering their questions, and overseeing twenty sites instead of one.

Which LLM is best for agentic pentesting?

There’s no single best model. Models change almost weekly, and the newest isn’t always best for every task, so Reflectiz benchmarks models task by task and picks the best fit for each. The harness around the model, which handles navigation, logins, and coverage, matters more than the model itself.

Why can a cheaper LLM cost more to run?

Price per token is only half the math. A model that’s cheap per token but uses a million tokens to scan one page can cost ten times more than a pricier model that does the same job in 10,000 tokens. Compare models on cost per completed task, not cost per token.

How did an AI agent find an authentication bypass?

Offensive Hub’s analyzer read the site’s JavaScript and saw that the one-time passcode request let the user choose where the code was sent. Its planner built an attack around that, the agent received the code through Reflectiz’s OTP tooling, and the login succeeded. Its guardrails then stopped the test.

How do guardrails keep a pentest agent safe in production?

Guardrails set hard limits on what the agent can do. In the authentication bypass case, the agent proved the login worked, then stopped. Reflectiz can also start in pre-production and validate a finding against production in one click, so teams can test critical flows like login under control.

What should you ask an agentic pentesting vendor?

Ask five questions: Was it built for the web? What enforces coverage? Can you dial the cost? Where are the guardrails? Can you act on the results? A strong vendor answers each with proof, such as a real-browser demo, a coverage record, and validated findings.

Subscribe to our newsletter

Stay updated with the latest news, articles, and insights from Reflectiz.

AI Has Changed The Web.

Are You Ready for What’s Next?

Third-party code shifts by the hour. Supply-chain compromises strike without warning. AI-driven web attacks now evolve faster than traditional security can ever keep up.

Reflectiz delivers the continuous, real-time visibility needed to expose the risks traditional tools miss entirely.

Zero code changes. Zero access to your data. Ultimate peace of mind.

Try for free