Skip to content
October 30, 2023

AI-Enhanced Reconnaissance: Fueling Sophisticated Security Breaches

Attackers use models to turn public DNS, GitHub, and job posts into a working map of you. The incident looks sophisticated when the first step was free.

Brackish Security5 min read

The write-up that called the last incident “sophisticated” usually skipped the first week. Nobody needed a new exploit. They needed a list: which of your names still resolve, which of those names still answer, and which of those answers belong to a system nobody has logged into since the last reorg.

That list used to take a patient operator a few days. A model does the same correlation in an afternoon, across more sources than most security teams review in a quarter. The breach looks advanced because the targeting was. The targeting was cheap.

Recon was always public. AI made it exhaustive.

Classic external reconnaissance is not a vulnerability. It is inventory. Certificate transparency logs, DNS, reverse WHOIS, GitHub, job ads, vendor status pages, Shodan and Censys banners — none of that requires a foothold. It requires your name.

What changed is not the data. It is the cost of reading all of it at once.

A human skims the interesting subdomains and moves on. A model will ingest the full set, pair staging. and dev. names with the same TLS issuer as production, notice that a careers page is still hiring for a specific VPN appliance, and draft the pretext that cites the project name from last month’s engineering blog. That is reconnaissance. It is also why the later email does not look like phishing.

We start the same way on an external test. DNSDumpster and certificate logs are still the first pass. The difference on the attacker side is they do not stop at a screenshot of the map. They keep going until every name has a role, a probable owner, and a likely login.

What the model is actually doing

Skip the version where a chatbot “hacks the company.” Operators use models for three boring jobs at volume.

Normalize the noise. Paste a dump of subdomains, HTML titles, and error pages and ask it to group them: mail gateway, forgotten WordPress, API with a Swagger UI on a non-standard port. That used to be a junior analyst with a spreadsheet.

Name the people and the stack. Job posts, LinkedIn, conference talks, and GitHub orgs are a living architecture diagram. “We are migrating from Duo to …” is a gift. So is a public Terraform example that still has the account alias in it. The model’s job is a short brief: identity provider, VPN, ticketing, the SaaS that will accept a password reset for the helpdesk.

Write the next packet. Once the map exists, the same model drafts the lures, the OpenAPI probes, and the wordlists that match your actual app names instead of admin and test. That is why AI phishing no longer has the old tells. The recon is what killed the tells.

None of this requires a leaked 0-day. It requires your public surface to be larger than the surface you think you have.

Why the breach then looks sophisticated

Incident reports like a story with a clever exploit. Boards like that story too. It implies the defenders were up against something exotic.

Most of the time the path is ordinary once you have the map:

The operator who found those things with a model did not invent a new class of attack. They compressed the time between “I have a domain” and “I have a working theory of this company.” That is the sophistication. It is industrial, not magical.

If your external pentest still starts from a two-page list of production URLs, you are testing the environment you wish you had. Attackers are testing the one certificate transparency already published.

What to assume is already enumerated

Treat these as already in someone else’s notes:

Then ask a meaner question: when did we last look at that list the way an outsider does? Not the CMDB. The CT log. If the answer is “the last pentest,” check whether that test was allowed to wander off the in-scope URLs. If it was not, the recon phase of the next real incident is untested.

Attack surface management is the continuous version of that check. A point-in-time external test still matters, but only if the tester is allowed to start where the model starts: with everything that already has your name on it.

Do not wait for the write-up to say sophisticated

If you want a useful control this quarter:

  1. Pull your own CT and DNS picture. Kill names that should not resolve. The DNSDumpster walkthrough is the afternoon version of that.
  2. Put production, staging, and vendor portals behind the same identity bar you claim in the policy. A leftover basic-auth box is a recon finding, not an edge case.
  3. When you scope the next external test, say the tester may follow any hostname that identifies as you. If that makes legal nervous, you already know the gap.

The model did not make your perimeter interesting. It made ignoring the interesting parts indefensible.

Want this tested against your environment?

Reading about an attack path is not the same as knowing whether yours holds. We can tell you which it is.

Scope an engagement