Looking for the older version? Find the previous pentest guide on GitHub.
TL;DR: A practical workflow for mapping web applications and APIs, validating security hypotheses, and producing reproducible findings aligned with the OWASP Web Security Testing Guide (WSTG). The full methodology with commands, expected outputs, and pitfalls per WSTG category is available in
wstg-web-pentest.
Author's note: This guide focuses on decisions made during an external web assessment and reports only the impact supported by evidence. The supporting playbooks are available in uphiago/recon-skills.
Web Application Pentest Guide
Reconnaissance tools return hostnames, services, routes, headers, technologies, and vulnerability candidates. A pentest turns those inputs into findings by classifying scope and ownership, mapping application behavior, testing a security property, and preserving evidence that proves or disproves it.
The workflow is:
define scope
-> collect and classify candidates
-> choose hostnames to test
-> resolve and probe services
-> map the application and identity
-> run bounded automated triage
-> select focused skills
-> validate one hypothesis
-> build verified attack paths
-> sanitize, score, and report
Classify results by evidence:
| State | Meaning |
|---|---|
| Signal | An unverified observation justifies follow-up |
| Exposed | A service or behavior is reachable in the tested context |
| Vulnerable | An expected security property failed under a controlled test |
| Exploitable | A controlled test demonstrated concrete capability or impact |
A version string is a signal. A reachable unauthenticated API is exposed. Neither observation is a finding until a test demonstrates the relevant behavior and impact.
0. Prerequisites
Bootstrap pdtm, the ProjectDiscovery package manager, then install the toolchain:
go install github.com/projectdiscovery/pdtm/cmd/pdtm@latest
pdtm -install subfinder,dnsx,httpx,katana,tlsx,nuclei,naabu
Each tool needs its own API key configuration. Subfinder and Chaos use service-specific keys (Chaos, Shodan, Censys, etc.). Without keys, subfinder returns only passive sources that work without authentication. Configure them once:
subfinder -ls # list configured sources
subfinder -config # interactive config
Set a conservative resolver list for DNS resolution:
curl -sS https://raw.githubusercontent.com/trickest/resolvers/main/resolvers.txt \
-o ~/.config/resolvers.txt
Install additional utilities available in the standard package manager or pip:
# Debian/Ubuntu
sudo apt install -y jq curl nmap openssl python3-pip
# macOS
brew install jq curl nmap openssl
pip3 install requests 2>/dev/null || sudo apt install -y python3-requests
Keep one directory per assessment, created at the start:
mkdir -p example.com/{assets,http,scans,evidence,reports}
Every command in this guide uses the defaults above. Adjust paths, API keys, and resolvers to your environment before starting each assessment.
1. Set the Scope
Before beginning discovery or sending target-related traffic, record:
- in-scope domains, IP ranges, applications, APIs, and accounts;
- exclusions and third-party assets;
- the authoritative scope source and how conflicting rules are resolved;
- permitted passive, active read-only, state-changing, and destructive tests;
- test accounts and synthetic data;
- global and per-host request limits;
- known source addresses and traffic-identification headers;
- test windows, monitoring contacts, abort thresholds, and cleanup duties.
Separate unauthenticated and authenticated work. An anonymous crawl maps the public boundary; test accounts reveal role, tenant, and object ownership. This separation ties each observation to the identity that produced it.
Classify actions by expected impact:
| Class | Examples | Default |
|---|---|---|
| Passive | Certificate Transparency logs, search datasets, public repositories | Follow discovery boundaries |
| Active read-only | DNS resolution, HTTP metadata, bounded crawling | Apply conservative rate limits: 2 req/s per process, 50 pkt/s for TCP, one concurrent DNS resolver query |
| Intrusive or state-changing | Uploads, object creation, credential attempts, out-of-band callbacks | Require explicit permission |
| Destructive | Deletion, persistence, availability tests, production modification | Do not run without explicit confirmation |
Classify an action by what it does, not by the tool or template that performs it. Record anonymous and authenticated identities separately from impact. Newly discovered names, redirects, subsidiaries, and provider-hosted services remain pending until the rules of engagement explicitly include them.
Stop when the next action leaves scope, error rates rise, the target becomes unstable, a defensive contact requests a pause, or the only remaining path requires production data or a destructive proof. Close the assessment when the planned coverage is complete or a stop condition prevents further testing. Document unresolved hypotheses, untested areas, and residual limits.
Organize Assessment Files
Keep one directory per assessment or target. Separate collected assets, service observations, application routes, reviewed evidence, and reports:
example.com/
|-- assets/ hostnames, DNS records, IPs, and ownership notes
|-- http/ live services, headers, fingerprints, and scan output
|-- urls/ crawled routes, archived URLs, parameters, and API candidates
|-- evidence/ sanitized requests, responses, screenshots, and controls
`-- reports/ reviewed findings and final deliverables
Store raw credentials, cookies, browser profiles, and authenticated request files in a separate protected location. Preserve only sanitized requests and responses as evidence, and exclude live session material from reports.
2. Build a Defensible Asset Inventory
A single-domain assessment starts with one root domain in scope. Discovery may produce thousands of related names, but those names are candidates, not new targets. Keep the passive list separate from the names selected for active testing.
Find Hostnames
Start with Subfinder. It asks passive sources for hostnames related to a domain:
subfinder -d vercel.com -json -cs
{"host":"auth.vercel.com","input":"vercel.com","source":"hackertarget"}
{"host":"data.vercel.com","input":"vercel.com","source":"submd"}
{"host":"cdp.vercel.com","input":"vercel.com","source":"hackertarget"}
{"host":"ssp.vercel.com","input":"vercel.com","source":"thc"}
{"host":"cnames.vercel.com","input":"vercel.com","source":"hackertarget"}
Every line is one JSON object. host is the discovered name; source
names the passive provider that returned it. Use -all to query every
available source when coverage justifies a slower run. Keep JSONL when
the assessment needs provider attribution per hostname.
Certificate Transparency logs capture every certificate a CA issues. Query crt.sh for the last 90 days and keep only entries matching the root domain:
curl -s "https://crt.sh/?q=%25.${TARGET}&output=json" \
| jq -r '.[].name_value // empty' \
| grep -E "\.?${TARGET}$" | sort -u > assets/crtsh.txt
crt.sh is rate-limited and occasionally returns 502. Retry once after
a 30-second wait. If it remains unavailable, continue with subfinder
and return to crt.sh later. The output is a snapshot that ages: certificates expire and wildcard
entries (*.target.com) are patterns, not discovered hosts. They are
candidates, not authorization. Add passive DNS, search indexes, and public
datasets through subfinder's -all flag when coverage justifies the slower
run.
Classify Hostnames
For each hostname, keep simple notes:
| Field | Example values |
|---|---|
| Scope | in, out, pending |
| Ownership | target, shared provider, third party, unknown |
| DNS state | resolved, wildcard, dangling, historical |
| Source | customer inventory, Certificate Transparency, passive DNS, search index |
| Last observed | UTC timestamp |
Scope and ownership are not confidence scores. A hostname can belong to the target but remain out of scope, while an in-scope application can run on a third-party platform.
Active DNS enumeration includes record queries, wordlist discovery, reverse lookups, and zone-transfer attempts. Use these target-visible techniques only when scope justifies them, then review their results through the same process.
Resolve a Hostname
Resolve one hostname with dnsx:
echo vercel.com | dnsx -silent -json
{
"host": "vercel.com",
"ttl": 433,
"resolver": ["1.0.0.1:53"],
"a": ["64.239.109.1", "64.239.109.129"],
"all": [
"vercel.com.\t433\tIN\tA\t64.239.109.1",
"vercel.com.\t433\tIN\tA\t64.239.109.129"
],
"status_code": "NOERROR",
"timestamp": "2026-07-31T13:53:52.873667315-03:00",
"query-time": "7ms"
}
The fields that matter first:
host: the queried hostname;status_code:NOERRORmeans the resolver answered successfully;a: current IPv4 addresses;ttl: how long a resolver may cache the record;resolverandtimestamp: where and when this observation was made;query-time: how long the resolution took;all: the raw DNS answer, useful when the compact fields need checking.
Addresses, TTL, resolver, and ordering change. This query returned IPv4
records. Add -aaaa for IPv6 records and -cname when CNAME chains matter.
Interpret the result in context:
- an A or AAAA record shows current resolution;
- a CNAME may identify a SaaS, CDN, or abandoned dependency;
- NXDOMAIN, SERVFAIL, and unresolved CNAME chains do not establish a usable address; investigate them before classifying the candidate;
- repeated answers for random labels may indicate wildcard DNS;
- shared-provider resolution does not prove target ownership.
For a reviewed list, pass it to dnsx -l and retain JSONL (one JSON object per
line) only when the assessment needs a record of every response. Catch-all
applications and wildcard DNS still require comparison against random-host
baselines.
3. Probe, Fingerprint, and Prioritize Services
Start with port discovery on the resolved addresses. The top 1000 TCP ports cover the common web, database, and administrative service range:
nmap -sT --top-ports 1000 --open -oA http/ports -iL assets/resolved.txt
Add every open port to the candidate URL list. Applications on ports 8080, 8443, 8081 often expose administrative interfaces, development systems, or internal APIs that port 443 does not.
Probe the live services with httpx, sending at most two HTTP requests per second:
httpx -u https://vercel.com \
-silent -rate-limit 2 \
-status-code -title -tech-detect -server -ip -cname -json \
| jq '{url, status_code, title, tech, webserver, host_ip, cdn, cdn_name}'
{
"url": "https://vercel.com",
"status_code": 200,
"title": "Agentic Infrastructure - Vercel",
"tech": [
"Amazon S3",
"Amazon Web Services",
"HSTS",
"Next.js",
"Node.js",
"React",
"Vercel",
"Webpack"
],
"webserver": "Vercel",
"host_ip": "64.239.123.1",
"cdn": true,
"cdn_name": "vercel"
}
Read the main fields first:
status_code: whether the service answered;title: a quick hint about the application. Here it confirms the Vercel marketing site;tech: technology stack. Next.js + React + Node.js + Webpack is the frontend stack; AWS + Vercel is the infrastructure layer. Confirm each through a second signal;webserver: theServerheader. Vercel sets its own;host_ipandcdn: where the request was answered. Vercel's edge network IP, not the application origin. Whencdnistrue, the origin server is behind Vercel, Cloudflare, Fastly, or a similar edge network — WAF rules, rate limits, and cached responses may mask the real application behavior. Usehttpx -cdnon every live URL to flag CDN/WAF-fronted services early.
The response is reachable, but that does not establish ownership or a vulnerability. Check the TLS service with tlsx:
tlsx -u vercel.com -silent -san -nc -json | python3 -m json.tool
{
"timestamp": "2026-07-31T13:55:08.375494478-03:00",
"host": "vercel.com",
"ip": "64.239.109.193",
"port": "443",
"probe_status": true,
"tls_version": "tls13",
"cipher": "TLS_AES_128_GCM_SHA256",
"not_before": "2026-07-24T06:47:26Z",
"not_after": "2026-10-22T06:47:25Z",
"subject_dn": "CN=*.vercel.com",
"subject_cn": "*.vercel.com",
"subject_an": [
"*.vercel.com",
"vercel.com"
],
"serial": "06:AE:5B:CC:E6:39:F3:82:DA:20:25:70:BB:DB:7A:EE:3A:C6",
"issuer_dn": "CN=YR2, O=Let's Encrypt, C=US",
"issuer_cn": "YR2",
"issuer_org": [
"Let's Encrypt"
],
"wildcard_certificate": true,
"tls_connection": "ctls",
"sni": "vercel.com"
}
subject_an are the SANs found on the certificate; tls_version and
cipher fingerprint the TLS stack; not_before and not_after show
certificate validity; issuer_org identifies the CA.
For a hostname list, replace -u vercel.com with -l hostnames.txt. Inspect
the certificate names with OpenSSL:
openssl s_client -connect vercel.com:443 -servername vercel.com \
</dev/null 2>/dev/null \
| openssl x509 -noout -ext subjectAltName
X509v3 Subject Alternative Name:
DNS:*.vercel.com, DNS:vercel.com
The wildcard entry is a hostname pattern, not a newly in-scope target. New names found in certificate SANs return to the candidate inventory and repeat the same scope, ownership, and DNS checks.
Treat these as separate facts:
hostname exists
!= service responds
!= response belongs to the expected virtual host
!= application belongs to the target
!= application is vulnerable
A reverse proxy can add headers, a default virtual host can supply the title, and a CDN address does not disclose the origin. Confirm important technology fingerprints through more than one signal.
When an in-scope IP serves a default site or several known names share an
address, compare candidate virtual hosts against a random-host baseline.
Account for the HTTP Host header, HTTP/2 :authority, and TLS SNI. Use
httpx with -vhost to test candidates:
httpx -u https://<IP> -H "Host: app.example.com" -vhost -json
Compare the response body length, title, and status code against the baseline (a random non-existent hostname). Divergence signals a real virtual host. Bounded virtual-host discovery is appropriate only when the IP and candidate namespace are covered by the rules of engagement.
Prioritize unusual applications, administrative interfaces, exposed development systems, identity boundaries, and services that connect multiple parts of the target. Do not prioritize solely by an old version string or a dramatic scanner label.
4. Map the Application, APIs, and Identity
Map routes, parameters, JavaScript, API operations, authentication flows, roles, objects, and state transitions. The goal is to know what an application does before selecting a vulnerability test.
For each application, answer:
| Area | Questions |
|---|---|
| Entry points | Which hosts, ports, paths, and protocols reach it? |
| Client | Which routes, bundles, source maps, workers, and configuration values ship to the browser? |
| API | Which operations, parameters, objects, versions, and content types exist? |
| Identity | Which login, MFA, OAuth, SAML, recovery, and account-linking flows exist? |
| Authorization | Which roles, tenants, and ownership boundaries protect each object? |
| State | Which operations read, create, change, upload, send, or delete? |
| Dependencies | Which storage, SaaS, cloud, WebSocket, or third-party services are connected? |
Crawl the Current Application
Start with one target application. This bounded crawl follows links and extracts JavaScript references without rendering a browser session:
katana -u https://vercel.com \
-silent -jc -c 2 -p 2 -rl 2
https://vercel.com
https://community.vercel.com
https://vercel.com/login
https://vercel.com/oauth/device
https://vercel.com/api/product-og?product=overview
https://vercel.com/api/front-domains/checkpoint
https://vercel.com/_stream/external/cdn.sift.com/s.js
This already shows the community.vercel.com subdomain, a login path,
an OAuth device flow endpoint, two API routes (product-og,
front-domains/checkpoint), and a third-party script from sift.com.
Review crawler output before testing: it can contain duplicates, stale
links, and strings that only resemble URLs. Group confirmed routes by
purpose: API, authentication, account recovery, administration, uploads,
WebSockets, and state-changing actions. Historical sources can add
retired routes and old parameters, but those still need verification.
Read robots.txt, sitemap.xml, API descriptions, and framework documentation
when present. They add routes to the map; they do not make those routes public
or in scope.
Choose a host, wordlist, rate, filters, and stop conditions before directory or parameter discovery. Derive a focused wordlist from the observed framework, routes, and API vocabulary instead of using a large generic list by default.
Map JavaScript-Rendered Applications
Standard crawlers do not see every route in a single-page application. Use a browser or Katana's headless mode when JavaScript rendering, XHR, WebSockets, client-side routing, or authenticated workflows change the surface.
An authenticated browser pass should use a test account and a dedicated profile. Capture:
- DOM-rendered routes and forms;
- XHR and fetch operations;
- WebSocket endpoints and message types;
- redirects, logout, refresh, and tenant switching;
- state-changing operations and their confirmation boundaries.
Do not paste production cookies into shell history or commit browser profiles. Browser automation reproduces user actions; it does not make those actions read-only.
Inspect JavaScript Bundles
A client-code review has five steps:
route or script reference
-> confirm JavaScript response
-> read routes and configuration
-> check source maps
-> verify the backend behavior
A URL ending in .js is not necessarily a bundle. Vercel serves
Next.js chunks under a nested base path. Confirm status and content type
before reading:
curl -sS -D - -o /dev/null \
https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js \
| sed -n '1,6p'
HTTP/2 200
content-type: application/javascript; charset=utf-8
access-control-allow-origin: *
cache-control: public, max-age=31536000, immutable
Extract and Read Bundle Content
Pull every script reference from the page, then download the largest application bundles. Skip third-party hosts and inline blobs:
curl -sS --max-time 10 https://vercel.com 2>/dev/null \
| grep -oP 'src="(https?://[^"]+\.js[^"]*|/[^"]+\.js[^"]*)"' \
| sed 's/src="//;s/"//' \
| sort -u > bundles.txt
head bundles.txt
/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/06lcjp8ev1jh2.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0_9e5qow9mkg0.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0c0hxoamwjsbw.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0gjs2lbzs80w5.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0nufcqvum7hyr.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0o61-vosflpug.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0tk7awe1kf6up.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0y53lbtp3fja_.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0y5fvpaxb21lx.js
Every script on this page is path-relative — Vercel serves its own bundles
and third-party integrations through backend proxying, not direct <script>
tags. No filtering step is needed. Prepend the base URL and download:
while IFS= read -r url; do
curl -sS --max-time 10 "https://vercel.com${url}" \
-o "bundle-$(echo "$url" | md5sum | cut -c1-8).js"
sleep 1
done < bundles.txt
ls -S bundle-*.js | head -5
bundle-1f818802.js
bundle-ef012ab2.js
bundle-42ca781a.js
bundle-4dcec42f.js
bundle-1ca91cd1.js
Not every bundle carries application logic. The largest file
(bundle-1f818802.js, 112KB) is framework runtime code. Open a few
of the top bundles and grep for routes and product features:
# Application routes — paths that reveal product surface
grep -oP '"(/[a-z][a-z0-9/_-]{3,})"' bundle-ef012ab2.js | sort -u
"/about"
"/account/settings"
"/agent"
"/ai-gateway"
"/ai-sdk"
"/blog"
"/botid"
"/chat"
"/connect"
"/contact/sales"
"/customers"
"/dashboard"
"/docs/flags/flags-sdk-reference"
"/docs/frameworks/backend/fastapi"
"/docs/frameworks/backend/nitro"
"/docs/frameworks/full-stack/nuxt"
"/docs/frameworks/full-stack/sveltekit"
"/docs/rest-api/sdk"
"/domains"
"/drop"
"/enterprise"
Twenty-one routes from one bundle. /agent, /ai-gateway, /ai-sdk,
and /botid are product features. /dashboard and /enterprise are
authenticated pages. /docs/rest-api/sdk hints at an API client.
Framework integration docs expose the platform's dependency surface.
A second grep finds subdomains referenced in the same bundle:
# Subdomains and external services
grep -oP '[a-zA-Z-]+\.vercel\.com' bundle-ef012ab2.js | sort -u
community.vercel.com
security.vercel.com
Two subdomains not in the initial subfinder pass. Add them to the candidate inventory and repeat scope and DNS checks.
Search for Credentials
Routes and subdomains are one class of finding. Credentials are another. While Vercel's landing page returned no matches (expected for a well-configured target), a bundle from another target might print:
grep -oP 'AIza[0-9A-Za-z_-]{35}' bundle-*.js
bundle-abc123.js:AIzaSyDrN5CwGpPqMkLxVnRtZu3j8Wh2Bf7Ea9Y
A single Google API key exposed in one bundle. The key type (Maps, Cloud, Firebase) determines what it grants. Test with the least-privileged documented request before concluding impact.
The full credential grep set is:
for f in bundle-*.js; do
[ -s "$f" ] || continue
grep -HoPi '(?:x-)?api[_-]?key\s*[=:]\s*["\x27][^"\x27]{8,}["\x27]' "$f"
grep -HoP 'AIza[0-9A-Za-z_-]{35}' "$f"
grep -HoP 'AKIA[0-9A-Z]{16}' "$f"
grep -HoPi '(?:sk|pk|rk)_(?:live|test)_[A-Za-z0-9]+' "$f"
grep -HoPi '(?:firebase|supabase)\s*[=:]\s*["\x27][^"\x27]+["\x27]' "$f"
grep -HoP 'ghp_[A-Za-z0-9]{36}' "$f"
grep -HoP 'github_pat_[A-Za-z0-9_]{20,}' "$f"
grep -HoP 'glpat-[A-Za-z0-9_-]{20,}' "$f"
grep -HoP 'xox[baprs]-[0-9A-Za-z-]{10,}' "$f"
grep -HoPi 'sk-(?:proj|svcacct|ant)-[A-Za-z0-9_-]{20,}' "$f"
grep -HoP 'eyJ[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}' "$f"
grep -HoP '10\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
grep -HoP '172\.(1[6-9]|2[0-9]|3[01])\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
grep -HoP '192\.168\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
grep -HoP 'https?://[a-zA-Z0-9.-]+:[0-9]{2,5}' "$f"
grep -HoP '-----BEGIN [A-Z ]+ PRIVATE KEY-----' "$f"
done
Each match includes the bundle filename (-H) and prints only the
matched text (-o). Classify the match before escalating:
pk_live_ is a public identifier; sk_live_ is a secret.
A JWT printed by this grep still needs decoding to
inspect claims, issuer, and expiry. JWT uses base64url, not standard
base64 — base64 -d will fail on - and _ characters:
echo 'eyJ...rest...of...token' | cut -d. -f2 | python3 -c "
import sys, base64, json
body = sys.stdin.read().strip()
# base64url → standard padding
body += '=' * (4 - len(body) % 4)
print(json.dumps(json.loads(base64.urlsafe_b64decode(body)), indent=2))
"
marketing page probably powers a Maps embed; the same key in a backend bundle is privileged. Context turns a grep match into a finding.
Search for External Dependencies and Source Code
Credentials are one class of finding. Dependencies and linked repositories are another. This grep extracts external hosts, analytics IDs, GitHub repos, and social handles from the page HTML:
curl -sS --max-time 10 https://vercel.com 2>/dev/null \
| grep -oP 'src="https://[^"]*\.js"|data-website-id="[^"]*"|https://(github|gitlab)\.com/[^"'\''<> ]+|https://x\.com/[^"'\''<> ]+' \
| sort -u
https://github.com/vercel
https://x.com/vercel
The GitHub organization and social handle were found in the page source.
github.com/vercel links to public repositories: clone the relevant ones
and audit for hardcoded keys, internal hostnames, CI/CD configuration,
Dockerfiles, and environment templates. x.com/vercel is documented in the
asset inventory and excluded from testing.
Recover Source Maps
A source map is a JSON file that maps minified bundle code back to the original sources. When exposed, it recovers readable TypeScript, JSX, and configuration files — the code the developer wrote, not what the bundler emitted.
Check the map next to each confirmed bundle by appending .map:
curl -sS -o /dev/null -w '%{http_code} %{content_type}\n' \
https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js.map
403 text/plain
This bundle returned 403. The map either does not exist, is blocked by WAF, or is served only with a referrer check. Both 403 and 404 are common and expected results. Move on to the next bundle.
When a map returns 200 with application/json, download it and expand
the sources. Replace the chunk URL below with the one that returned 200:
curl -sS https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/THE-CHUNK-THAT-RETURNED-200.js.map \
-o chunk.js.map
python3 - <<'PY'
import json
from pathlib import Path
source_map = json.load(open("chunk.js.map"))
output = Path("recovered-source")
output.mkdir(exist_ok=True)
for idx, (name, content) in enumerate(zip(
source_map.get("sources", []),
source_map.get("sourcesContent", [])
)):
if content:
fname = f"{idx:04d}-{Path(name).name}"
(output / fname).write_text(content)
print(fname)
PY
0000-webpack-runtime.js
0001-_app.tsx
0002-index.tsx
0003-api-client.ts
0004-auth-config.ts
0005-feature-flags.json
Seven recovered files from one map. api-client.ts and auth-config.ts
are the immediate targets: they contain request construction, base URLs,
OAuth endpoints, and environment-specific logic that a crawl never reaches.
What to search in recovered sources:
grep -rniP 'baseURL|apiUrl|endpoint|Authorization|clientId|clientSecret|redirectUri|issuer|tenantId|apiKey' recovered-source/
grep -rniP 'https?://[a-zA-Z0-9.-]+\.[a-z]{2,}' recovered-source/
grep -rniP 'secret|password|token|credential' recovered-source/
A source map converts a minified blob into a readable codebase in seconds.
If the current bundle has no map, review historical JavaScript URLs from
waybackurls and dynamic imports loaded through import(). Older deployed
versions sometimes exposed maps that were later blocked.
Classify and Validate Values
Classify a value before replaying it. pk_live_ and OAuth client_id values
are public client identifiers; sk_live_, service-role keys, private keys, and
client secrets are not. A Firebase or Supabase anonymous key may be expected
in a browser bundle, but its allowed backend behavior is still testable.
For a discovered Supabase project, compare a minimal read request with and without the anonymous key. Do not query application tables until the scope and test account support it:
curl -sS -D - -o /dev/null \
https://<project>.supabase.co/rest/v1/
curl -sS -D - -o /dev/null \
-H 'apikey: <supabase-anon-key>' \
https://<project>.supabase.co/rest/v1/
Record the status, headers, endpoint, key type, and what the provider permits.
The finding is excessive access, not the presence of an anonymous key. Treat a
backend hostname such as *.fly.dev, *.supabase.co, or *.r2.dev as a new
service candidate: resolve it, confirm ownership, then map it like any other
application.
Minified code is still evidence. A source map makes it easier to read, but a missing map does not end the review: follow the request construction, headers, route strings, and provider configuration in the bundle. Do not report a string match as a secret, endpoint, or vulnerability without identifying its use and testing the relevant server-side behavior.
When source code is available, compare its route definitions and HTTP clients with browser traffic. This finds entry points that a crawl did not reach and also exposes dead or feature-gated routes that still need classification.
Build an API and Identity Matrix
Check for descriptions such as OpenAPI, Swagger, GraphQL, and framework documentation. Capture browser traffic in an intercepting proxy or request workspace. When no schema exists, derive the inventory from browser traffic, JavaScript, mobile clients when in scope, archives, and error responses.
For every route, record the request inputs and response behavior, not only the path:
GET /api/v1/projects/{project_id} path: project_id user session 200, 401, 404
PATCH /api/v1/projects/{project_id} body: name, status manager session 200, 403
POST /api/v1/projects/{id}/members body: member_id manager session 201, 403
GET /graphql query: operationName bearer token 200, 401
WS /socket cookie: session session cookie upgrade, close
For each page or flow, capture:
- query parameters, path values, request bodies, and hidden form fields;
- custom request headers, cookies, and response
Set-Cookieattributes; - redirects, error responses, and the final destination;
- observed HTTP methods and any method that needs a separate check;
- WebSocket upgrades, message types, and connection boundaries;
- multi-step paths such as login, MFA, recovery, checkout, and tenant changes.
Do not infer that a method is safe to test because it appears in a response or schema. Reads, writes, uploads, and destructive actions belong in different test plans.
For authenticated applications, define a small identity matrix:
| Identity | Session | Expected access |
|---|---|---|
| Anonymous | None | Public routes only |
| User A | Test tenant A | Own objects |
| User B | Test tenant B | Own objects |
| Manager | Privileged test account | Team or tenant objects |
Many authorization findings require two identities, two tenants, or a privileged role. Without the required control account, label the path untested instead of extrapolating from one response.
5. Use Automation to Create Hypotheses
After mapping, use automation to create a small, reviewable hypothesis queue. Do not begin with every available template or scanner.
Nuclei matches known exposures, misconfigurations, technologies, and CVE preconditions. Template execution is code-driven testing: review the selected templates, exclude intrusive categories, identify traffic when required, and set both request and concurrency limits.
Start with templates relevant to an observed technology:
# Technology-specific templates (the allowlist approach)
nuclei -l live-urls.txt \
-t ~/nuclei-templates/ \
-tags cve,exposure,misconfig \
-etags intrusive,fuzz,dos,bruteforce \
-rate-limit 5 -c 2 \
-duc -j -o nuclei.jsonl
If the application map found an OpenAPI schema, feed it to the interactive mode:
nuclei -l live-urls.txt -im openapi
For blind vulnerability classes (SSRF, XXE, blind XSS), run with an out-of-band callback to confirm reachability:
nuclei -l live-urls.txt \
-t ~/nuclei-templates/ \
-tags oob,ssrf,xxe \
-interactsh-server oast.pro
Record the target, template ID, matcher, request, and response. A
template match is a lead, not a finding: an info result can expose an
important administrative surface, while a critical match can fail
manual reproduction.
For tight production rules, use an allowlist of reviewed templates. Headless, code, fuzzing, authenticated, out-of-band, and state-changing templates need separate confirmation. Pin the tool and template revision when reproducibility matters.
Treat a template match as a lead:
template match
-> inspect matcher and prerequisite
-> reproduce with a bounded request
-> compare a negative control
-> demonstrate impact or discard
Automated output should shorten the manual queue. It does not replace application mapping, business-logic review, or role comparison.
6. Route Findings Through the WSTG Categories
The OWASP Web Security Testing Guide
organizes web application testing into 12 categories. Use this routing table to
match what you observed to the right test procedure. Every observation in
columns 1 and 2 maps to a WSTG category and a recon-skills procedure.
The complete methodology with full commands, expected outputs, and pitfalls
per category lives in wstg-web-pentest.
Phase Routing Table
| WSTG Category | What you observed | Test | Skills to consult |
|---|---|---|---|
| INFO Information Gathering | Hostnames, services, headers | Fingerprint servers, map architecture, discover entry points | web2-recon, wstg-web-pentest |
| CONF Configuration | Exposed headers, backup files, admin panels | Test HTTP methods, enumerate config files, check security headers, subdomain takeover | hunt-source-leak, hunt-subdomain, hunt-tls-network, wstg-web-pentest |
| IDNT Identity Mgmt | Registration, password reset | Account enumeration, weak username policy, role definition gaps | hunt-brute-force, hunt-ato |
| ATHN Authentication | Login, MFA, password reset flows | Default credentials, weak lockout, bypass schema, MFA gaps, weak reset tokens | hunt-auth-bypass, hunt-mfa-bypass, hunt-brute-force |
| ATHZ Authorization | Object IDs in URLs, role-based UI, OAuth/SAML redirects | IDOR, privilege escalation, forced browsing, directory traversal, OAuth redirect_uri, SAML assertion tampering | hunt-idor, hunt-lfi, hunt-oauth, hunt-saml |
| SESS Session Mgmt | Cookies, JWT, logout | Cookie attributes, session fixation, CSRF, JWT attacks, logout invalidation | hunt-session, hunt-csrf |
| INPV Input Validation | Query parameters, forms, API bodies | XSS, SQLi, SSTI, SSRF, command injection, mass assignment, prototype pollution, HTTP smuggling, host header injection | hunt-xss, hunt-sqli, hunt-ssrf, hunt-ssti, hunt-rce, hunt-host-header, hunt-http-smuggling, hunt-api-misconfig, hunt-prototype-pollution, hunt-mass-assignment |
| ERRH Error Handling | Stack traces, debug pages, verbose errors | Trigger malformed input, check for information leakage in error responses | hunt-information-disclosure |
| CRYP Cryptography | TLS, encrypted tokens, hashing | Weak TLS versions, padding oracle, unencrypted channels, weak cipher suites, expired certificates | hunt-tls-network |
| BUSL Business Logic | Checkout, coupons, file upload, workflows | Negative quantities, price manipulation, race conditions, workflow bypass, file upload abuse | hunt-business-logic, hunt-file-upload, hunt-race-condition |
| CLNT Client-side | CORS headers, WebSocket upgrade, iframe, postMessage | CORS misconfig, DOM XSS, clickjacking, WebSocket auth, browser storage inspection, CSS injection | hunt-cors, hunt-websocket, hunt-dom |
| APIT API Testing | REST, GraphQL, OpenAPI schema | BOLA, BFLA, excessive data exposure, schema enumeration, GraphQL introspection | hunt-graphql, hunt-api-misconfig, hunt-idor, hunt-schema-enumeration, hunt-broken-function-level-auth |
When to Skip a Category
Do not test every category because the list exists. Route by evidence:
Observed: JSON response with sequential object IDs
-> ATHZ (Authorization): IDOR test with two identities
-> Skip: ATHN (Authentication) because the login flow is unchanged
Observed: Origin reflected with ACAC on a profile endpoint
-> CLNT (Client-side): CORS credential leakage test
-> Skip: BUSL (Business Logic) because no transaction flow exists
Load a skill only when its prerequisite evidence exists. Confirm conflicting fingerprints before selecting platform-specific tests.
Start with one procedure and add another only when the result creates a new question.
7. Validate One Hypothesis at a Time
A hypothesis names the security property, expected failure, supporting evidence, and cheapest safe falsification step:
If: user A can address an object owned by user B
Then: the API may return cross-tenant data
Because: the object ID is accepted directly by a read endpoint
Control: compare owner, non-owner, anonymous, and unknown-object responses
Keep a short queue:
| Hypothesis | Evidence so far | Cost | Safety | Status | Next control |
|---|---|---|---:|---|---|---|
| API object may cross a tenant boundary | Predictable ID in browser traffic | Low | Read-only | Open | Compare two synthetic tenants |
| CORS may expose profile data | Origin reflected on /api/me | Low | Read-only | Open | Browser proof with test session |
| Source map may reveal an unlinked API | Map and API client found | Low | Read-only | Validated | Resolve and classify routes |
| Component may match a known CVE | Version and feature prerequisite observed | Medium | Varies | Dismissed | Reproduce the vulnerable behavior |
| Upload may reach executable storage | Upload route exists | High | State-changing | Confirm scope and use an inert file |
Prioritize strong evidence, meaningful potential impact, cheap falsification, and limited side effects. Do not prioritize by the most dramatic imagined outcome.
CORS Example
Origin reflection is a lead:
curl -sS --max-time 10 \
-D "evidence/cors-headers.txt" \
-o "evidence/cors-body.txt" \
-H 'Origin: https://attacker.example' \
"https://app.example.test/api/me"
A defensible CORS finding establishes:
- the endpoint returns user-specific or sensitive data;
- a browser session for a test account supplies the required credentials;
- the response accepts an untrusted origin;
- credentials are allowed when the scenario requires them;
- browser JavaScript can read the response without CORP blocking;
- SameSite cookie behavior does not silently strip the session;
Vary: Originis absent when the response caches differ by origin;- a preflight check (PUT, PATCH, custom headers) succeeds or is absent when the method requires none;
- public, anonymous, and trusted-origin controls behave differently.
Access-Control-Allow-Origin: * is not credential theft because the
fetch spec prohibits browsers from combining * with
Access-Control-Allow-Credentials: true. An attacker who controls the
network (MITM on shared WiFi) can bypass this restriction, but the
threat model differs from a standard reflected-origin CORS finding.
Reflection on an error response may have no impact. A header scanner
does not prove browser-readable sensitive data.
Authorization Example
Changing project-a to project-b is meaningful only when both are synthetic,
belong to test identities or tenants, and the expected policy is known:
expected: user A receives a denial for project B
observed: user A receives project B's non-public synthetic fields
control 1: user B can read project B
control 2: an unknown project returns the documented not-found response
control 3: an anonymous request is rejected
not tested: enumeration, production objects, writes
The same discipline applies throughout the pack:
200 OKon/.envis not a leak if the body is an application shell;- an old version does not prove the vulnerable component is reachable;
- an exposed database port is not remote code execution;
- a public client identifier is not automatically a credential;
- an API schema is not an authorization vulnerability;
- an accepted write method does not authorize arbitrary modification.
8. Build Branching Attack Paths
Link results only when one supplies a prerequisite or capability for the next. Label each link as observed, inferred, confirmed, or not tested.
A path should branch where the evidence branches:
OpenAPI operation discovered
|
+-- anonymous request succeeds
| `-- confirm data sensitivity and ownership
|
+-- anonymous request is denied
| `-- compare user A and user B on synthetic objects
|
`-- endpoint is no longer reachable
`-- record historical context and stop
For example:
1. Observed
app.example.test loads /assets/app.js.
2. Observed
The bundle references api.example.test/v2 and a project route.
3. Observed
/openapi.json documents GET /projects/{project_id}.
4. Confirmed
User A can read project A; user B can read project B.
5. Confirmed
User A can also read project B's synthetic billing contact.
6. Not tested
Object enumeration, production records, and write operations.
The bundle and schema supply the API base URL, operation, identifier, and expected object shape. The security failure occurs at step 5; hiding the documentation would not fix object-level authorization.
Do not score an inferred final step as demonstrated impact. If the output of one
finding is not required by the next, the findings are co-located, not chained.
Use cross-attack-chains only after component findings pass
triage-validation.
9. Capture, Sanitize, Score, and Report
For each validated behavior, record enough for another operator to reproduce the result:
- affected asset, identity, environment, and UTC timestamp;
- exact steps with sanitized requests and responses;
- expected, observed, and control behavior;
- prerequisites, side effects, demonstrated impact, and testing limits;
- remediation and a verification step;
- tool versions and evidence hashes when integrity tracking is required.
A compact finding can include all of this:
## Cross-tenant read of a synthetic project
**Affected asset:** `https://api.example.test/v1/projects/{project_id}`
**Severity:** High
**Classification:** CWE-639
**Expected:** User A cannot read an object owned by tenant B.
**Observed:** User A received project B's synthetic billing contact.
**Controls:** User B could read the object; an anonymous request returned 401;
an unknown identifier returned the documented not-found response.
**Impact demonstrated:** Unauthorized read of one synthetic cross-tenant
record.
**Not tested:** Enumeration, production records, and write operations.
**Remediation:** Enforce tenant and object ownership on every project lookup.
Repeat the owner, non-owner, anonymous, and unknown-object controls after the
fix.
Severity follows demonstrated impact. CVSS communicates technical severity; it does not replace business context, asset criticality, exploit preconditions, or the program's rating policy. Use the version required by the engagement.
Sanitize before sharing:
- remove or replace cookies, bearer tokens, API keys, and passwords;
- minimize personal and customer data;
- replace unrelated records with synthetic examples where possible;
- review screenshots, filenames, browser storage, and response metadata;
- keep a protected original only when retention policy requires it.
Run a final quality gate:
[ ] reproduction and controls work as documented
[ ] impact and severity match the evidence; untested steps are labeled
[ ] credentials and unrelated data are removed
[ ] remediation addresses the failed security control
10. Use and Extend Recon Skills
Use recon-skills as a procedure library. Read the skill that matches the observed surface, follow its prerequisites and verification, and update the owning skill when field work produces a reusable improvement without target details.
Start with these entry points:
| What you need | Skill |
|---|---|
| Full WSTG-aligned methodology with commands per phase | wstg-web-pentest |
| Asset discovery and service fingerprinting | web2-recon |
| Vulnerability-class testing (XSS, SQLi, SSRF, IDOR, etc.) | hunt-* skills under redteam/ |
| WordPress-specific mapping and validation | hunt-wordpress |
| Compound attack paths from verified findings | cross-attack-chains |
| Finding quality gate before reporting | triage-validation |
| Sanitized evidence and report structure | report-writing |
11. Close the Assessment
When the assessment becomes noisy, return to these questions:
What did I observe?
What security property might be broken?
What is the cheapest safe test that could disprove it?
Which identity, object, or environment does that test require?
What control distinguishes the result from expected behavior?
What new capability does the confirmed result provide?
Then collect missing context, run one bounded validation, capture the result, or stop because the next step is unsafe, out of scope, or unsupported.
Useful Links
- OWASP WSTG (Web Security Testing Guide)
- OWASP WSTG Testing Categories
- WSTG-aligned Pentest Skill
- Recon Skills
- ProjectDiscovery documentation
- Nuclei Templates
- PortSwigger Web Security Academy
- SecLists
- CVSS v4.0
- OWASP Top 10 (2021)
- OWASP API Security Top 10
- HackTricks
Reach Me
Questions, corrections, and field notes: @uphiago