Looking for the older version? Find the previous pentest guide on GitHub.

TL;DR: A practical workflow for mapping web applications and APIs, validating security hypotheses, and producing reproducible findings aligned with the OWASP Web Security Testing Guide (WSTG). The full methodology with commands, expected outputs, and pitfalls per WSTG category is available in wstg-web-pentest.

Author's note: This guide focuses on decisions made during an external web assessment and reports only the impact supported by evidence. The supporting playbooks are available in uphiago/recon-skills.

Web Application Pentest Guide

Reconnaissance tools return hostnames, services, routes, headers, technologies, and vulnerability candidates. A pentest turns those inputs into findings by classifying scope and ownership, mapping application behavior, testing a security property, and preserving evidence that proves or disproves it.

The workflow is:

define scope
  -> collect and classify candidates
  -> choose hostnames to test
  -> resolve and probe services
  -> map the application and identity
  -> run bounded automated triage
  -> select focused skills
  -> validate one hypothesis
  -> build verified attack paths
  -> sanitize, score, and report

Classify results by evidence:

State Meaning
Signal An unverified observation justifies follow-up
Exposed A service or behavior is reachable in the tested context
Vulnerable An expected security property failed under a controlled test
Exploitable A controlled test demonstrated concrete capability or impact

A version string is a signal. A reachable unauthenticated API is exposed. Neither observation is a finding until a test demonstrates the relevant behavior and impact.

0. Prerequisites

Bootstrap pdtm, the ProjectDiscovery package manager, then install the toolchain:

go install github.com/projectdiscovery/pdtm/cmd/pdtm@latest
pdtm -install subfinder,dnsx,httpx,katana,tlsx,nuclei,naabu

Each tool needs its own API key configuration. Subfinder and Chaos use service-specific keys (Chaos, Shodan, Censys, etc.). Without keys, subfinder returns only passive sources that work without authentication. Configure them once:

subfinder -ls                      # list configured sources
subfinder -config                  # interactive config

Set a conservative resolver list for DNS resolution:

curl -sS https://raw.githubusercontent.com/trickest/resolvers/main/resolvers.txt \
  -o ~/.config/resolvers.txt

Install additional utilities available in the standard package manager or pip:

# Debian/Ubuntu
sudo apt install -y jq curl nmap openssl python3-pip
# macOS
brew install jq curl nmap openssl
pip3 install requests 2>/dev/null || sudo apt install -y python3-requests

Keep one directory per assessment, created at the start:

mkdir -p example.com/{assets,http,scans,evidence,reports}

Every command in this guide uses the defaults above. Adjust paths, API keys, and resolvers to your environment before starting each assessment.

1. Set the Scope

Before beginning discovery or sending target-related traffic, record:

  • in-scope domains, IP ranges, applications, APIs, and accounts;
  • exclusions and third-party assets;
  • the authoritative scope source and how conflicting rules are resolved;
  • permitted passive, active read-only, state-changing, and destructive tests;
  • test accounts and synthetic data;
  • global and per-host request limits;
  • known source addresses and traffic-identification headers;
  • test windows, monitoring contacts, abort thresholds, and cleanup duties.

Separate unauthenticated and authenticated work. An anonymous crawl maps the public boundary; test accounts reveal role, tenant, and object ownership. This separation ties each observation to the identity that produced it.

Classify actions by expected impact:

Class Examples Default
Passive Certificate Transparency logs, search datasets, public repositories Follow discovery boundaries
Active read-only DNS resolution, HTTP metadata, bounded crawling Apply conservative rate limits: 2 req/s per process, 50 pkt/s for TCP, one concurrent DNS resolver query
Intrusive or state-changing Uploads, object creation, credential attempts, out-of-band callbacks Require explicit permission
Destructive Deletion, persistence, availability tests, production modification Do not run without explicit confirmation

Classify an action by what it does, not by the tool or template that performs it. Record anonymous and authenticated identities separately from impact. Newly discovered names, redirects, subsidiaries, and provider-hosted services remain pending until the rules of engagement explicitly include them.

Stop when the next action leaves scope, error rates rise, the target becomes unstable, a defensive contact requests a pause, or the only remaining path requires production data or a destructive proof. Close the assessment when the planned coverage is complete or a stop condition prevents further testing. Document unresolved hypotheses, untested areas, and residual limits.

Organize Assessment Files

Keep one directory per assessment or target. Separate collected assets, service observations, application routes, reviewed evidence, and reports:

example.com/
|-- assets/     hostnames, DNS records, IPs, and ownership notes
|-- http/       live services, headers, fingerprints, and scan output
|-- urls/       crawled routes, archived URLs, parameters, and API candidates
|-- evidence/   sanitized requests, responses, screenshots, and controls
`-- reports/    reviewed findings and final deliverables

Store raw credentials, cookies, browser profiles, and authenticated request files in a separate protected location. Preserve only sanitized requests and responses as evidence, and exclude live session material from reports.

2. Build a Defensible Asset Inventory

A single-domain assessment starts with one root domain in scope. Discovery may produce thousands of related names, but those names are candidates, not new targets. Keep the passive list separate from the names selected for active testing.

Find Hostnames

Start with Subfinder. It asks passive sources for hostnames related to a domain:

subfinder -d vercel.com -json -cs
{"host":"auth.vercel.com","input":"vercel.com","source":"hackertarget"}
{"host":"data.vercel.com","input":"vercel.com","source":"submd"}
{"host":"cdp.vercel.com","input":"vercel.com","source":"hackertarget"}
{"host":"ssp.vercel.com","input":"vercel.com","source":"thc"}
{"host":"cnames.vercel.com","input":"vercel.com","source":"hackertarget"}

Every line is one JSON object. host is the discovered name; source names the passive provider that returned it. Use -all to query every available source when coverage justifies a slower run. Keep JSONL when the assessment needs provider attribution per hostname.

Certificate Transparency logs capture every certificate a CA issues. Query crt.sh for the last 90 days and keep only entries matching the root domain:

curl -s "https://crt.sh/?q=%25.${TARGET}&output=json" \
  | jq -r '.[].name_value // empty' \
  | grep -E "\.?${TARGET}$" | sort -u > assets/crtsh.txt

crt.sh is rate-limited and occasionally returns 502. Retry once after a 30-second wait. If it remains unavailable, continue with subfinder and return to crt.sh later. The output is a snapshot that ages: certificates expire and wildcard entries (*.target.com) are patterns, not discovered hosts. They are candidates, not authorization. Add passive DNS, search indexes, and public datasets through subfinder's -all flag when coverage justifies the slower run.

Classify Hostnames

For each hostname, keep simple notes:

Field Example values
Scope in, out, pending
Ownership target, shared provider, third party, unknown
DNS state resolved, wildcard, dangling, historical
Source customer inventory, Certificate Transparency, passive DNS, search index
Last observed UTC timestamp

Scope and ownership are not confidence scores. A hostname can belong to the target but remain out of scope, while an in-scope application can run on a third-party platform.

Active DNS enumeration includes record queries, wordlist discovery, reverse lookups, and zone-transfer attempts. Use these target-visible techniques only when scope justifies them, then review their results through the same process.

Resolve a Hostname

Resolve one hostname with dnsx:

echo vercel.com | dnsx -silent -json
{
    "host": "vercel.com",
    "ttl": 433,
    "resolver": ["1.0.0.1:53"],
    "a": ["64.239.109.1", "64.239.109.129"],
    "all": [
        "vercel.com.\t433\tIN\tA\t64.239.109.1",
        "vercel.com.\t433\tIN\tA\t64.239.109.129"
    ],
    "status_code": "NOERROR",
    "timestamp": "2026-07-31T13:53:52.873667315-03:00",
    "query-time": "7ms"
}

The fields that matter first:

  • host: the queried hostname;
  • status_code: NOERROR means the resolver answered successfully;
  • a: current IPv4 addresses;
  • ttl: how long a resolver may cache the record;
  • resolver and timestamp: where and when this observation was made;
  • query-time: how long the resolution took;
  • all: the raw DNS answer, useful when the compact fields need checking.

Addresses, TTL, resolver, and ordering change. This query returned IPv4 records. Add -aaaa for IPv6 records and -cname when CNAME chains matter. Interpret the result in context:

  • an A or AAAA record shows current resolution;
  • a CNAME may identify a SaaS, CDN, or abandoned dependency;
  • NXDOMAIN, SERVFAIL, and unresolved CNAME chains do not establish a usable address; investigate them before classifying the candidate;
  • repeated answers for random labels may indicate wildcard DNS;
  • shared-provider resolution does not prove target ownership.

For a reviewed list, pass it to dnsx -l and retain JSONL (one JSON object per line) only when the assessment needs a record of every response. Catch-all applications and wildcard DNS still require comparison against random-host baselines.

3. Probe, Fingerprint, and Prioritize Services

Start with port discovery on the resolved addresses. The top 1000 TCP ports cover the common web, database, and administrative service range:

nmap -sT --top-ports 1000 --open -oA http/ports -iL assets/resolved.txt

Add every open port to the candidate URL list. Applications on ports 8080, 8443, 8081 often expose administrative interfaces, development systems, or internal APIs that port 443 does not.

Probe the live services with httpx, sending at most two HTTP requests per second:

httpx -u https://vercel.com \
  -silent -rate-limit 2 \
  -status-code -title -tech-detect -server -ip -cname -json \
  | jq '{url, status_code, title, tech, webserver, host_ip, cdn, cdn_name}'
{
  "url": "https://vercel.com",
  "status_code": 200,
  "title": "Agentic Infrastructure - Vercel",
  "tech": [
    "Amazon S3",
    "Amazon Web Services",
    "HSTS",
    "Next.js",
    "Node.js",
    "React",
    "Vercel",
    "Webpack"
  ],
  "webserver": "Vercel",
  "host_ip": "64.239.123.1",
  "cdn": true,
  "cdn_name": "vercel"
}

Read the main fields first:

  • status_code: whether the service answered;
  • title: a quick hint about the application. Here it confirms the Vercel marketing site;
  • tech: technology stack. Next.js + React + Node.js + Webpack is the frontend stack; AWS + Vercel is the infrastructure layer. Confirm each through a second signal;
  • webserver: the Server header. Vercel sets its own;
  • host_ip and cdn: where the request was answered. Vercel's edge network IP, not the application origin. When cdn is true, the origin server is behind Vercel, Cloudflare, Fastly, or a similar edge network — WAF rules, rate limits, and cached responses may mask the real application behavior. Use httpx -cdn on every live URL to flag CDN/WAF-fronted services early.

The response is reachable, but that does not establish ownership or a vulnerability. Check the TLS service with tlsx:

tlsx -u vercel.com -silent -san -nc -json | python3 -m json.tool
{
    "timestamp": "2026-07-31T13:55:08.375494478-03:00",
    "host": "vercel.com",
    "ip": "64.239.109.193",
    "port": "443",
    "probe_status": true,
    "tls_version": "tls13",
    "cipher": "TLS_AES_128_GCM_SHA256",
    "not_before": "2026-07-24T06:47:26Z",
    "not_after": "2026-10-22T06:47:25Z",
    "subject_dn": "CN=*.vercel.com",
    "subject_cn": "*.vercel.com",
    "subject_an": [
        "*.vercel.com",
        "vercel.com"
    ],
    "serial": "06:AE:5B:CC:E6:39:F3:82:DA:20:25:70:BB:DB:7A:EE:3A:C6",
    "issuer_dn": "CN=YR2, O=Let's Encrypt, C=US",
    "issuer_cn": "YR2",
    "issuer_org": [
        "Let's Encrypt"
    ],
    "wildcard_certificate": true,
    "tls_connection": "ctls",
    "sni": "vercel.com"
}

subject_an are the SANs found on the certificate; tls_version and cipher fingerprint the TLS stack; not_before and not_after show certificate validity; issuer_org identifies the CA.

For a hostname list, replace -u vercel.com with -l hostnames.txt. Inspect the certificate names with OpenSSL:

openssl s_client -connect vercel.com:443 -servername vercel.com \
  </dev/null 2>/dev/null \
  | openssl x509 -noout -ext subjectAltName
X509v3 Subject Alternative Name:
    DNS:*.vercel.com, DNS:vercel.com

The wildcard entry is a hostname pattern, not a newly in-scope target. New names found in certificate SANs return to the candidate inventory and repeat the same scope, ownership, and DNS checks.

Treat these as separate facts:

hostname exists
  != service responds
  != response belongs to the expected virtual host
  != application belongs to the target
  != application is vulnerable

A reverse proxy can add headers, a default virtual host can supply the title, and a CDN address does not disclose the origin. Confirm important technology fingerprints through more than one signal.

When an in-scope IP serves a default site or several known names share an address, compare candidate virtual hosts against a random-host baseline. Account for the HTTP Host header, HTTP/2 :authority, and TLS SNI. Use httpx with -vhost to test candidates:

httpx -u https://<IP> -H "Host: app.example.com" -vhost -json

Compare the response body length, title, and status code against the baseline (a random non-existent hostname). Divergence signals a real virtual host. Bounded virtual-host discovery is appropriate only when the IP and candidate namespace are covered by the rules of engagement.

Prioritize unusual applications, administrative interfaces, exposed development systems, identity boundaries, and services that connect multiple parts of the target. Do not prioritize solely by an old version string or a dramatic scanner label.

4. Map the Application, APIs, and Identity

Map routes, parameters, JavaScript, API operations, authentication flows, roles, objects, and state transitions. The goal is to know what an application does before selecting a vulnerability test.

For each application, answer:

Area Questions
Entry points Which hosts, ports, paths, and protocols reach it?
Client Which routes, bundles, source maps, workers, and configuration values ship to the browser?
API Which operations, parameters, objects, versions, and content types exist?
Identity Which login, MFA, OAuth, SAML, recovery, and account-linking flows exist?
Authorization Which roles, tenants, and ownership boundaries protect each object?
State Which operations read, create, change, upload, send, or delete?
Dependencies Which storage, SaaS, cloud, WebSocket, or third-party services are connected?

Crawl the Current Application

Start with one target application. This bounded crawl follows links and extracts JavaScript references without rendering a browser session:

katana -u https://vercel.com \
  -silent -jc -c 2 -p 2 -rl 2
https://vercel.com
https://community.vercel.com
https://vercel.com/login
https://vercel.com/oauth/device
https://vercel.com/api/product-og?product=overview
https://vercel.com/api/front-domains/checkpoint
https://vercel.com/_stream/external/cdn.sift.com/s.js

This already shows the community.vercel.com subdomain, a login path, an OAuth device flow endpoint, two API routes (product-og, front-domains/checkpoint), and a third-party script from sift.com. Review crawler output before testing: it can contain duplicates, stale links, and strings that only resemble URLs. Group confirmed routes by purpose: API, authentication, account recovery, administration, uploads, WebSockets, and state-changing actions. Historical sources can add retired routes and old parameters, but those still need verification.

Read robots.txt, sitemap.xml, API descriptions, and framework documentation when present. They add routes to the map; they do not make those routes public or in scope.

Choose a host, wordlist, rate, filters, and stop conditions before directory or parameter discovery. Derive a focused wordlist from the observed framework, routes, and API vocabulary instead of using a large generic list by default.

Map JavaScript-Rendered Applications

Standard crawlers do not see every route in a single-page application. Use a browser or Katana's headless mode when JavaScript rendering, XHR, WebSockets, client-side routing, or authenticated workflows change the surface.

An authenticated browser pass should use a test account and a dedicated profile. Capture:

  • DOM-rendered routes and forms;
  • XHR and fetch operations;
  • WebSocket endpoints and message types;
  • redirects, logout, refresh, and tenant switching;
  • state-changing operations and their confirmation boundaries.

Do not paste production cookies into shell history or commit browser profiles. Browser automation reproduces user actions; it does not make those actions read-only.

Inspect JavaScript Bundles

A client-code review has five steps:

route or script reference
  -> confirm JavaScript response
  -> read routes and configuration
  -> check source maps
  -> verify the backend behavior

A URL ending in .js is not necessarily a bundle. Vercel serves Next.js chunks under a nested base path. Confirm status and content type before reading:

curl -sS -D - -o /dev/null \
  https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js \
  | sed -n '1,6p'
HTTP/2 200
content-type: application/javascript; charset=utf-8
access-control-allow-origin: *
cache-control: public, max-age=31536000, immutable

Extract and Read Bundle Content

Pull every script reference from the page, then download the largest application bundles. Skip third-party hosts and inline blobs:

curl -sS --max-time 10 https://vercel.com 2>/dev/null \
  | grep -oP 'src="(https?://[^"]+\.js[^"]*|/[^"]+\.js[^"]*)"' \
  | sed 's/src="//;s/"//' \
  | sort -u > bundles.txt

head bundles.txt
/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/06lcjp8ev1jh2.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0_9e5qow9mkg0.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0c0hxoamwjsbw.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0gjs2lbzs80w5.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0nufcqvum7hyr.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0o61-vosflpug.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0tk7awe1kf6up.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0y53lbtp3fja_.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0y5fvpaxb21lx.js

Every script on this page is path-relative — Vercel serves its own bundles and third-party integrations through backend proxying, not direct <script> tags. No filtering step is needed. Prepend the base URL and download:

while IFS= read -r url; do
  curl -sS --max-time 10 "https://vercel.com${url}" \
    -o "bundle-$(echo "$url" | md5sum | cut -c1-8).js"
  sleep 1
done < bundles.txt

ls -S bundle-*.js | head -5
bundle-1f818802.js
bundle-ef012ab2.js
bundle-42ca781a.js
bundle-4dcec42f.js
bundle-1ca91cd1.js

Not every bundle carries application logic. The largest file (bundle-1f818802.js, 112KB) is framework runtime code. Open a few of the top bundles and grep for routes and product features:

# Application routes — paths that reveal product surface
grep -oP '"(/[a-z][a-z0-9/_-]{3,})"' bundle-ef012ab2.js | sort -u
"/about"
"/account/settings"
"/agent"
"/ai-gateway"
"/ai-sdk"
"/blog"
"/botid"
"/chat"
"/connect"
"/contact/sales"
"/customers"
"/dashboard"
"/docs/flags/flags-sdk-reference"
"/docs/frameworks/backend/fastapi"
"/docs/frameworks/backend/nitro"
"/docs/frameworks/full-stack/nuxt"
"/docs/frameworks/full-stack/sveltekit"
"/docs/rest-api/sdk"
"/domains"
"/drop"
"/enterprise"

Twenty-one routes from one bundle. /agent, /ai-gateway, /ai-sdk, and /botid are product features. /dashboard and /enterprise are authenticated pages. /docs/rest-api/sdk hints at an API client. Framework integration docs expose the platform's dependency surface.

A second grep finds subdomains referenced in the same bundle:

# Subdomains and external services
grep -oP '[a-zA-Z-]+\.vercel\.com' bundle-ef012ab2.js | sort -u
community.vercel.com
security.vercel.com

Two subdomains not in the initial subfinder pass. Add them to the candidate inventory and repeat scope and DNS checks.

Search for Credentials

Routes and subdomains are one class of finding. Credentials are another. While Vercel's landing page returned no matches (expected for a well-configured target), a bundle from another target might print:

grep -oP 'AIza[0-9A-Za-z_-]{35}' bundle-*.js
bundle-abc123.js:AIzaSyDrN5CwGpPqMkLxVnRtZu3j8Wh2Bf7Ea9Y

A single Google API key exposed in one bundle. The key type (Maps, Cloud, Firebase) determines what it grants. Test with the least-privileged documented request before concluding impact.

The full credential grep set is:

for f in bundle-*.js; do
  [ -s "$f" ] || continue
  grep -HoPi '(?:x-)?api[_-]?key\s*[=:]\s*["\x27][^"\x27]{8,}["\x27]' "$f"
  grep -HoP 'AIza[0-9A-Za-z_-]{35}' "$f"
  grep -HoP 'AKIA[0-9A-Z]{16}' "$f"
  grep -HoPi '(?:sk|pk|rk)_(?:live|test)_[A-Za-z0-9]+' "$f"
  grep -HoPi '(?:firebase|supabase)\s*[=:]\s*["\x27][^"\x27]+["\x27]' "$f"
  grep -HoP 'ghp_[A-Za-z0-9]{36}' "$f"
  grep -HoP 'github_pat_[A-Za-z0-9_]{20,}' "$f"
  grep -HoP 'glpat-[A-Za-z0-9_-]{20,}' "$f"
  grep -HoP 'xox[baprs]-[0-9A-Za-z-]{10,}' "$f"
  grep -HoPi 'sk-(?:proj|svcacct|ant)-[A-Za-z0-9_-]{20,}' "$f"
  grep -HoP 'eyJ[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}' "$f"
  grep -HoP '10\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
  grep -HoP '172\.(1[6-9]|2[0-9]|3[01])\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
  grep -HoP '192\.168\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
  grep -HoP 'https?://[a-zA-Z0-9.-]+:[0-9]{2,5}' "$f"
  grep -HoP '-----BEGIN [A-Z ]+ PRIVATE KEY-----' "$f"
done

Each match includes the bundle filename (-H) and prints only the matched text (-o). Classify the match before escalating: pk_live_ is a public identifier; sk_live_ is a secret. A JWT printed by this grep still needs decoding to inspect claims, issuer, and expiry. JWT uses base64url, not standard base64 — base64 -d will fail on - and _ characters:

echo 'eyJ...rest...of...token' | cut -d. -f2 | python3 -c "
import sys, base64, json
body = sys.stdin.read().strip()
# base64url → standard padding
body += '=' * (4 - len(body) % 4)
print(json.dumps(json.loads(base64.urlsafe_b64decode(body)), indent=2))
"

marketing page probably powers a Maps embed; the same key in a backend bundle is privileged. Context turns a grep match into a finding.

Search for External Dependencies and Source Code

Credentials are one class of finding. Dependencies and linked repositories are another. This grep extracts external hosts, analytics IDs, GitHub repos, and social handles from the page HTML:

curl -sS --max-time 10 https://vercel.com 2>/dev/null \
  | grep -oP 'src="https://[^"]*\.js"|data-website-id="[^"]*"|https://(github|gitlab)\.com/[^"'\''<> ]+|https://x\.com/[^"'\''<> ]+' \
  | sort -u
https://github.com/vercel
https://x.com/vercel

The GitHub organization and social handle were found in the page source. github.com/vercel links to public repositories: clone the relevant ones and audit for hardcoded keys, internal hostnames, CI/CD configuration, Dockerfiles, and environment templates. x.com/vercel is documented in the asset inventory and excluded from testing.

Recover Source Maps

A source map is a JSON file that maps minified bundle code back to the original sources. When exposed, it recovers readable TypeScript, JSX, and configuration files — the code the developer wrote, not what the bundler emitted.

Check the map next to each confirmed bundle by appending .map:

curl -sS -o /dev/null -w '%{http_code} %{content_type}\n' \
  https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js.map
403 text/plain

This bundle returned 403. The map either does not exist, is blocked by WAF, or is served only with a referrer check. Both 403 and 404 are common and expected results. Move on to the next bundle.

When a map returns 200 with application/json, download it and expand the sources. Replace the chunk URL below with the one that returned 200:

curl -sS https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/THE-CHUNK-THAT-RETURNED-200.js.map \
  -o chunk.js.map

python3 - <<'PY'
import json
from pathlib import Path

source_map = json.load(open("chunk.js.map"))
output = Path("recovered-source")
output.mkdir(exist_ok=True)

for idx, (name, content) in enumerate(zip(
    source_map.get("sources", []),
    source_map.get("sourcesContent", [])
)):
    if content:
        fname = f"{idx:04d}-{Path(name).name}"
        (output / fname).write_text(content)
        print(fname)
PY
0000-webpack-runtime.js
0001-_app.tsx
0002-index.tsx
0003-api-client.ts
0004-auth-config.ts
0005-feature-flags.json

Seven recovered files from one map. api-client.ts and auth-config.ts are the immediate targets: they contain request construction, base URLs, OAuth endpoints, and environment-specific logic that a crawl never reaches.

What to search in recovered sources:

grep -rniP 'baseURL|apiUrl|endpoint|Authorization|clientId|clientSecret|redirectUri|issuer|tenantId|apiKey' recovered-source/
grep -rniP 'https?://[a-zA-Z0-9.-]+\.[a-z]{2,}' recovered-source/
grep -rniP 'secret|password|token|credential' recovered-source/

A source map converts a minified blob into a readable codebase in seconds. If the current bundle has no map, review historical JavaScript URLs from waybackurls and dynamic imports loaded through import(). Older deployed versions sometimes exposed maps that were later blocked.

Classify and Validate Values

Classify a value before replaying it. pk_live_ and OAuth client_id values are public client identifiers; sk_live_, service-role keys, private keys, and client secrets are not. A Firebase or Supabase anonymous key may be expected in a browser bundle, but its allowed backend behavior is still testable.

For a discovered Supabase project, compare a minimal read request with and without the anonymous key. Do not query application tables until the scope and test account support it:

curl -sS -D - -o /dev/null \
  https://<project>.supabase.co/rest/v1/

curl -sS -D - -o /dev/null \
  -H 'apikey: <supabase-anon-key>' \
  https://<project>.supabase.co/rest/v1/

Record the status, headers, endpoint, key type, and what the provider permits. The finding is excessive access, not the presence of an anonymous key. Treat a backend hostname such as *.fly.dev, *.supabase.co, or *.r2.dev as a new service candidate: resolve it, confirm ownership, then map it like any other application.

Minified code is still evidence. A source map makes it easier to read, but a missing map does not end the review: follow the request construction, headers, route strings, and provider configuration in the bundle. Do not report a string match as a secret, endpoint, or vulnerability without identifying its use and testing the relevant server-side behavior.

When source code is available, compare its route definitions and HTTP clients with browser traffic. This finds entry points that a crawl did not reach and also exposes dead or feature-gated routes that still need classification.

Build an API and Identity Matrix

Check for descriptions such as OpenAPI, Swagger, GraphQL, and framework documentation. Capture browser traffic in an intercepting proxy or request workspace. When no schema exists, derive the inventory from browser traffic, JavaScript, mobile clients when in scope, archives, and error responses.

For every route, record the request inputs and response behavior, not only the path:

GET    /api/v1/projects/{project_id}   path: project_id       user session     200, 401, 404
PATCH  /api/v1/projects/{project_id}   body: name, status     manager session  200, 403
POST   /api/v1/projects/{id}/members   body: member_id        manager session  201, 403
GET    /graphql                        query: operationName   bearer token     200, 401
WS     /socket                         cookie: session         session cookie   upgrade, close

For each page or flow, capture:

  • query parameters, path values, request bodies, and hidden form fields;
  • custom request headers, cookies, and response Set-Cookie attributes;
  • redirects, error responses, and the final destination;
  • observed HTTP methods and any method that needs a separate check;
  • WebSocket upgrades, message types, and connection boundaries;
  • multi-step paths such as login, MFA, recovery, checkout, and tenant changes.

Do not infer that a method is safe to test because it appears in a response or schema. Reads, writes, uploads, and destructive actions belong in different test plans.

For authenticated applications, define a small identity matrix:

Identity Session Expected access
Anonymous None Public routes only
User A Test tenant A Own objects
User B Test tenant B Own objects
Manager Privileged test account Team or tenant objects

Many authorization findings require two identities, two tenants, or a privileged role. Without the required control account, label the path untested instead of extrapolating from one response.

5. Use Automation to Create Hypotheses

After mapping, use automation to create a small, reviewable hypothesis queue. Do not begin with every available template or scanner.

Nuclei matches known exposures, misconfigurations, technologies, and CVE preconditions. Template execution is code-driven testing: review the selected templates, exclude intrusive categories, identify traffic when required, and set both request and concurrency limits.

Start with templates relevant to an observed technology:

# Technology-specific templates (the allowlist approach)
nuclei -l live-urls.txt \
  -t ~/nuclei-templates/ \
  -tags cve,exposure,misconfig \
  -etags intrusive,fuzz,dos,bruteforce \
  -rate-limit 5 -c 2 \
  -duc -j -o nuclei.jsonl

If the application map found an OpenAPI schema, feed it to the interactive mode:

nuclei -l live-urls.txt -im openapi

For blind vulnerability classes (SSRF, XXE, blind XSS), run with an out-of-band callback to confirm reachability:

nuclei -l live-urls.txt \
  -t ~/nuclei-templates/ \
  -tags oob,ssrf,xxe \
  -interactsh-server oast.pro

Record the target, template ID, matcher, request, and response. A template match is a lead, not a finding: an info result can expose an important administrative surface, while a critical match can fail manual reproduction.

For tight production rules, use an allowlist of reviewed templates. Headless, code, fuzzing, authenticated, out-of-band, and state-changing templates need separate confirmation. Pin the tool and template revision when reproducibility matters.

Treat a template match as a lead:

template match
  -> inspect matcher and prerequisite
  -> reproduce with a bounded request
  -> compare a negative control
  -> demonstrate impact or discard

Automated output should shorten the manual queue. It does not replace application mapping, business-logic review, or role comparison.

6. Route Findings Through the WSTG Categories

The OWASP Web Security Testing Guide organizes web application testing into 12 categories. Use this routing table to match what you observed to the right test procedure. Every observation in columns 1 and 2 maps to a WSTG category and a recon-skills procedure.

The complete methodology with full commands, expected outputs, and pitfalls per category lives in wstg-web-pentest.

Phase Routing Table

WSTG Category What you observed Test Skills to consult
INFO Information Gathering Hostnames, services, headers Fingerprint servers, map architecture, discover entry points web2-recon, wstg-web-pentest
CONF Configuration Exposed headers, backup files, admin panels Test HTTP methods, enumerate config files, check security headers, subdomain takeover hunt-source-leak, hunt-subdomain, hunt-tls-network, wstg-web-pentest
IDNT Identity Mgmt Registration, password reset Account enumeration, weak username policy, role definition gaps hunt-brute-force, hunt-ato
ATHN Authentication Login, MFA, password reset flows Default credentials, weak lockout, bypass schema, MFA gaps, weak reset tokens hunt-auth-bypass, hunt-mfa-bypass, hunt-brute-force
ATHZ Authorization Object IDs in URLs, role-based UI, OAuth/SAML redirects IDOR, privilege escalation, forced browsing, directory traversal, OAuth redirect_uri, SAML assertion tampering hunt-idor, hunt-lfi, hunt-oauth, hunt-saml
SESS Session Mgmt Cookies, JWT, logout Cookie attributes, session fixation, CSRF, JWT attacks, logout invalidation hunt-session, hunt-csrf
INPV Input Validation Query parameters, forms, API bodies XSS, SQLi, SSTI, SSRF, command injection, mass assignment, prototype pollution, HTTP smuggling, host header injection hunt-xss, hunt-sqli, hunt-ssrf, hunt-ssti, hunt-rce, hunt-host-header, hunt-http-smuggling, hunt-api-misconfig, hunt-prototype-pollution, hunt-mass-assignment
ERRH Error Handling Stack traces, debug pages, verbose errors Trigger malformed input, check for information leakage in error responses hunt-information-disclosure
CRYP Cryptography TLS, encrypted tokens, hashing Weak TLS versions, padding oracle, unencrypted channels, weak cipher suites, expired certificates hunt-tls-network
BUSL Business Logic Checkout, coupons, file upload, workflows Negative quantities, price manipulation, race conditions, workflow bypass, file upload abuse hunt-business-logic, hunt-file-upload, hunt-race-condition
CLNT Client-side CORS headers, WebSocket upgrade, iframe, postMessage CORS misconfig, DOM XSS, clickjacking, WebSocket auth, browser storage inspection, CSS injection hunt-cors, hunt-websocket, hunt-dom
APIT API Testing REST, GraphQL, OpenAPI schema BOLA, BFLA, excessive data exposure, schema enumeration, GraphQL introspection hunt-graphql, hunt-api-misconfig, hunt-idor, hunt-schema-enumeration, hunt-broken-function-level-auth

When to Skip a Category

Do not test every category because the list exists. Route by evidence:

Observed: JSON response with sequential object IDs
  -> ATHZ (Authorization): IDOR test with two identities
  -> Skip: ATHN (Authentication) because the login flow is unchanged

Observed: Origin reflected with ACAC on a profile endpoint
  -> CLNT (Client-side): CORS credential leakage test
  -> Skip: BUSL (Business Logic) because no transaction flow exists

Load a skill only when its prerequisite evidence exists. Confirm conflicting fingerprints before selecting platform-specific tests.

Start with one procedure and add another only when the result creates a new question.

7. Validate One Hypothesis at a Time

A hypothesis names the security property, expected failure, supporting evidence, and cheapest safe falsification step:

If:       user A can address an object owned by user B
Then:     the API may return cross-tenant data
Because:  the object ID is accepted directly by a read endpoint
Control:  compare owner, non-owner, anonymous, and unknown-object responses

Keep a short queue:

| Hypothesis | Evidence so far | Cost | Safety | Status | Next control | |---|---|---|---:|---|---|---| | API object may cross a tenant boundary | Predictable ID in browser traffic | Low | Read-only | Open | Compare two synthetic tenants | | CORS may expose profile data | Origin reflected on /api/me | Low | Read-only | Open | Browser proof with test session | | Source map may reveal an unlinked API | Map and API client found | Low | Read-only | Validated | Resolve and classify routes | | Component may match a known CVE | Version and feature prerequisite observed | Medium | Varies | Dismissed | Reproduce the vulnerable behavior | | Upload may reach executable storage | Upload route exists | High | State-changing | Confirm scope and use an inert file |

Prioritize strong evidence, meaningful potential impact, cheap falsification, and limited side effects. Do not prioritize by the most dramatic imagined outcome.

CORS Example

Origin reflection is a lead:

curl -sS --max-time 10 \
  -D "evidence/cors-headers.txt" \
  -o "evidence/cors-body.txt" \
  -H 'Origin: https://attacker.example' \
  "https://app.example.test/api/me"

A defensible CORS finding establishes:

  1. the endpoint returns user-specific or sensitive data;
  2. a browser session for a test account supplies the required credentials;
  3. the response accepts an untrusted origin;
  4. credentials are allowed when the scenario requires them;
  5. browser JavaScript can read the response without CORP blocking;
  6. SameSite cookie behavior does not silently strip the session;
  7. Vary: Origin is absent when the response caches differ by origin;
  8. a preflight check (PUT, PATCH, custom headers) succeeds or is absent when the method requires none;
  9. public, anonymous, and trusted-origin controls behave differently.

Access-Control-Allow-Origin: * is not credential theft because the fetch spec prohibits browsers from combining * with Access-Control-Allow-Credentials: true. An attacker who controls the network (MITM on shared WiFi) can bypass this restriction, but the threat model differs from a standard reflected-origin CORS finding. Reflection on an error response may have no impact. A header scanner does not prove browser-readable sensitive data.

Authorization Example

Changing project-a to project-b is meaningful only when both are synthetic, belong to test identities or tenants, and the expected policy is known:

expected: user A receives a denial for project B
observed: user A receives project B's non-public synthetic fields
control 1: user B can read project B
control 2: an unknown project returns the documented not-found response
control 3: an anonymous request is rejected
not tested: enumeration, production objects, writes

The same discipline applies throughout the pack:

  • 200 OK on /.env is not a leak if the body is an application shell;
  • an old version does not prove the vulnerable component is reachable;
  • an exposed database port is not remote code execution;
  • a public client identifier is not automatically a credential;
  • an API schema is not an authorization vulnerability;
  • an accepted write method does not authorize arbitrary modification.

8. Build Branching Attack Paths

Link results only when one supplies a prerequisite or capability for the next. Label each link as observed, inferred, confirmed, or not tested.

A path should branch where the evidence branches:

OpenAPI operation discovered
|
+-- anonymous request succeeds
|   `-- confirm data sensitivity and ownership
|
+-- anonymous request is denied
|   `-- compare user A and user B on synthetic objects
|
`-- endpoint is no longer reachable
    `-- record historical context and stop

For example:

1. Observed
   app.example.test loads /assets/app.js.

2. Observed
   The bundle references api.example.test/v2 and a project route.

3. Observed
   /openapi.json documents GET /projects/{project_id}.

4. Confirmed
   User A can read project A; user B can read project B.

5. Confirmed
   User A can also read project B's synthetic billing contact.

6. Not tested
   Object enumeration, production records, and write operations.

The bundle and schema supply the API base URL, operation, identifier, and expected object shape. The security failure occurs at step 5; hiding the documentation would not fix object-level authorization.

Do not score an inferred final step as demonstrated impact. If the output of one finding is not required by the next, the findings are co-located, not chained. Use cross-attack-chains only after component findings pass triage-validation.

9. Capture, Sanitize, Score, and Report

For each validated behavior, record enough for another operator to reproduce the result:

  • affected asset, identity, environment, and UTC timestamp;
  • exact steps with sanitized requests and responses;
  • expected, observed, and control behavior;
  • prerequisites, side effects, demonstrated impact, and testing limits;
  • remediation and a verification step;
  • tool versions and evidence hashes when integrity tracking is required.

A compact finding can include all of this:

## Cross-tenant read of a synthetic project

**Affected asset:** `https://api.example.test/v1/projects/{project_id}`
**Severity:** High
**Classification:** CWE-639

**Expected:** User A cannot read an object owned by tenant B.

**Observed:** User A received project B's synthetic billing contact.

**Controls:** User B could read the object; an anonymous request returned 401;
an unknown identifier returned the documented not-found response.

**Impact demonstrated:** Unauthorized read of one synthetic cross-tenant
record.

**Not tested:** Enumeration, production records, and write operations.

**Remediation:** Enforce tenant and object ownership on every project lookup.
Repeat the owner, non-owner, anonymous, and unknown-object controls after the
fix.

Severity follows demonstrated impact. CVSS communicates technical severity; it does not replace business context, asset criticality, exploit preconditions, or the program's rating policy. Use the version required by the engagement.

Sanitize before sharing:

  • remove or replace cookies, bearer tokens, API keys, and passwords;
  • minimize personal and customer data;
  • replace unrelated records with synthetic examples where possible;
  • review screenshots, filenames, browser storage, and response metadata;
  • keep a protected original only when retention policy requires it.

Run a final quality gate:

[ ] reproduction and controls work as documented
[ ] impact and severity match the evidence; untested steps are labeled
[ ] credentials and unrelated data are removed
[ ] remediation addresses the failed security control

10. Use and Extend Recon Skills

Use recon-skills as a procedure library. Read the skill that matches the observed surface, follow its prerequisites and verification, and update the owning skill when field work produces a reusable improvement without target details.

Start with these entry points:

What you need Skill
Full WSTG-aligned methodology with commands per phase wstg-web-pentest
Asset discovery and service fingerprinting web2-recon
Vulnerability-class testing (XSS, SQLi, SSRF, IDOR, etc.) hunt-* skills under redteam/
WordPress-specific mapping and validation hunt-wordpress
Compound attack paths from verified findings cross-attack-chains
Finding quality gate before reporting triage-validation
Sanitized evidence and report structure report-writing

11. Close the Assessment

When the assessment becomes noisy, return to these questions:

What did I observe?
What security property might be broken?
What is the cheapest safe test that could disprove it?
Which identity, object, or environment does that test require?
What control distinguishes the result from expected behavior?
What new capability does the confirmed result provide?

Then collect missing context, run one bounded validation, capture the result, or stop because the next step is unsafe, out of scope, or unsupported.

Reach Me

Questions, corrections, and field notes: @uphiago