+++
author = "@uphiago"
title = "Web Application Pentest Guide"
slug = "pentest-playbook"
date = 2026-06-22T16:20:00-00:00
lastmod = 2026-07-30T00:00:00-03:00
show_updated = true
description = "A field guide for external reconnaissance, application mapping, bounded automation, validation, attack paths, and pentest reporting."

tags = [
  "pentest",
  "recon",
  "web-security",
  "api-security",
  "automation",
]
authors = ["uphiago"]
draft = false
+++
<!--more-->

> **Looking for the older version?** [Find the previous pentest guide on GitHub](https://github.com/uphiago/portfolio/blob/d5b9f426d96280b126652f8d0efa7f07e15467f9/content/posts/2026/ai/pentest-recon.md).

> **TL;DR:** A practical workflow for mapping web applications and
> APIs, validating security hypotheses, and producing reproducible findings
> aligned with the OWASP Web Security Testing Guide (WSTG). The full
> methodology with commands, expected outputs, and pitfalls per WSTG
> category is available in
> [`wstg-web-pentest`](https://github.com/uphiago/recon-skills/tree/main/redteam/wstg-web-pentest).

> **Author's note:** This guide focuses on decisions made during an external web
> assessment and reports only the impact supported by evidence. The supporting
> playbooks are available in
> [uphiago/recon-skills](https://github.com/uphiago/recon-skills).

# Web Application Pentest Guide

Reconnaissance tools return hostnames, services, routes, headers, technologies,
and vulnerability candidates. A pentest turns those inputs into findings by
classifying scope and ownership, mapping application behavior, testing a
security property, and preserving evidence that proves or disproves it.

The workflow is:

```text
define scope
  -> collect and classify candidates
  -> choose hostnames to test
  -> resolve and probe services
  -> map the application and identity
  -> run bounded automated triage
  -> select focused skills
  -> validate one hypothesis
  -> build verified attack paths
  -> sanitize, score, and report
```

Classify results by evidence:

| State | Meaning |
|---|---|
| Signal | An unverified observation justifies follow-up |
| Exposed | A service or behavior is reachable in the tested context |
| Vulnerable | An expected security property failed under a controlled test |
| Exploitable | A controlled test demonstrated concrete capability or impact |

A version string is a signal. A reachable unauthenticated API is exposed.
Neither observation is a finding until a test demonstrates the relevant
behavior and impact.

## 0. Prerequisites

Bootstrap [pdtm](https://github.com/projectdiscovery/pdtm), the
ProjectDiscovery package manager, then install the toolchain:

```bash
go install github.com/projectdiscovery/pdtm/cmd/pdtm@latest
pdtm -install subfinder,dnsx,httpx,katana,tlsx,nuclei,naabu
```

Each tool needs its own API key configuration. Subfinder and Chaos use
service-specific keys (Chaos, Shodan, Censys, etc.). Without keys, subfinder
returns only passive sources that work without authentication. Configure
them once:

```bash
subfinder -ls                      # list configured sources
subfinder -config                  # interactive config
```

Set a conservative resolver list for DNS resolution:

```bash
curl -sS https://raw.githubusercontent.com/trickest/resolvers/main/resolvers.txt \
  -o ~/.config/resolvers.txt
```

Install additional utilities available in the standard package manager or pip:

```bash
# Debian/Ubuntu
sudo apt install -y jq curl nmap openssl python3-pip
# macOS
brew install jq curl nmap openssl
```

```bash
pip3 install requests 2>/dev/null || sudo apt install -y python3-requests
```

Keep one directory per assessment, created at the start:

```
mkdir -p example.com/{assets,http,scans,evidence,reports}
```

Every command in this guide uses the defaults above. Adjust paths, API keys,
and resolvers to your environment before starting each assessment.

## 1. Set the Scope

Before beginning discovery or sending target-related traffic, record:

- in-scope domains, IP ranges, applications, APIs, and accounts;
- exclusions and third-party assets;
- the authoritative scope source and how conflicting rules are resolved;
- permitted passive, active read-only, state-changing, and destructive tests;
- test accounts and synthetic data;
- global and per-host request limits;
- known source addresses and traffic-identification headers;
- test windows, monitoring contacts, abort thresholds, and cleanup duties.

Separate unauthenticated and authenticated work. An anonymous crawl maps the
public boundary; test accounts reveal role, tenant, and object ownership. This
separation ties each observation to the identity that produced it.

Classify actions by expected impact:

| Class | Examples | Default |
|---|---|---|
| Passive | Certificate Transparency logs, search datasets, public repositories | Follow discovery boundaries |
| Active read-only | DNS resolution, HTTP metadata, bounded crawling | Apply conservative rate limits: 2 req/s per process, 50 pkt/s for TCP, one concurrent DNS resolver query |
| Intrusive or state-changing | Uploads, object creation, credential attempts, out-of-band callbacks | Require explicit permission |
| Destructive | Deletion, persistence, availability tests, production modification | Do not run without explicit confirmation |

Classify an action by what it does, not by the tool or template that performs
it. Record anonymous and authenticated identities separately from impact.
Newly discovered names, redirects, subsidiaries, and provider-hosted services
remain pending until the rules of engagement explicitly include them.

Stop when the next action leaves scope, error rates rise, the target
becomes unstable, a defensive contact requests a pause, or the only remaining
path requires production data or a destructive proof. Close the assessment
when the planned coverage is complete or a stop condition prevents further
testing. Document unresolved hypotheses, untested areas, and residual limits.

### Organize Assessment Files

Keep one directory per assessment or target. Separate collected assets,
service observations, application routes, reviewed evidence, and reports:

```text
example.com/
|-- assets/     hostnames, DNS records, IPs, and ownership notes
|-- http/       live services, headers, fingerprints, and scan output
|-- urls/       crawled routes, archived URLs, parameters, and API candidates
|-- evidence/   sanitized requests, responses, screenshots, and controls
`-- reports/    reviewed findings and final deliverables
```

Store raw credentials, cookies, browser profiles, and authenticated request
files in a separate protected location. Preserve only sanitized requests and
responses as evidence, and exclude live session material from reports.

## 2. Build a Defensible Asset Inventory

A single-domain assessment starts with one root domain in scope. Discovery
may produce thousands of related names, but those names are candidates, not
new targets. Keep the passive list separate from the names selected for active
testing.

### Find Hostnames

Start with [Subfinder](https://github.com/projectdiscovery/subfinder). It asks
passive sources for hostnames related to a domain:

```bash
subfinder -d vercel.com -json -cs
```

```json
{"host":"auth.vercel.com","input":"vercel.com","source":"hackertarget"}
{"host":"data.vercel.com","input":"vercel.com","source":"submd"}
{"host":"cdp.vercel.com","input":"vercel.com","source":"hackertarget"}
{"host":"ssp.vercel.com","input":"vercel.com","source":"thc"}
{"host":"cnames.vercel.com","input":"vercel.com","source":"hackertarget"}
```

Every line is one JSON object. `host` is the discovered name; `source`
names the passive provider that returned it. Use `-all` to query every
available source when coverage justifies a slower run. Keep JSONL when
the assessment needs provider attribution per hostname.

Certificate Transparency logs capture every certificate a CA issues.
Query crt.sh for the last 90 days and keep only entries matching the root
domain:

```bash
curl -s "https://crt.sh/?q=%25.${TARGET}&output=json" \
  | jq -r '.[].name_value // empty' \
  | grep -E "\.?${TARGET}$" | sort -u > assets/crtsh.txt
```

crt.sh is rate-limited and occasionally returns 502. Retry once after
a 30-second wait. If it remains unavailable, continue with subfinder
and return to crt.sh later. The output is a snapshot that ages: certificates expire and wildcard
entries (`*.target.com`) are patterns, not discovered hosts. They are
candidates, not authorization. Add passive DNS, search indexes, and public
datasets through subfinder's `-all` flag when coverage justifies the slower
run.

### Classify Hostnames

For each hostname, keep simple notes:

| Field | Example values |
|---|---|
| Scope | in, out, pending |
| Ownership | target, shared provider, third party, unknown |
| DNS state | resolved, wildcard, dangling, historical |
| Source | customer inventory, Certificate Transparency, passive DNS, search index |
| Last observed | UTC timestamp |

Scope and ownership are not confidence scores. A hostname can belong to the
target but remain out of scope, while an in-scope application can run on a
third-party platform.

Active DNS enumeration includes record queries, wordlist discovery, reverse
lookups, and zone-transfer attempts. Use these target-visible techniques only
when scope justifies them, then review their results through
the same process.

### Resolve a Hostname

Resolve one hostname with [dnsx](https://github.com/projectdiscovery/dnsx):

```bash
echo vercel.com | dnsx -silent -json
```

```json
{
    "host": "vercel.com",
    "ttl": 433,
    "resolver": ["1.0.0.1:53"],
    "a": ["64.239.109.1", "64.239.109.129"],
    "all": [
        "vercel.com.\t433\tIN\tA\t64.239.109.1",
        "vercel.com.\t433\tIN\tA\t64.239.109.129"
    ],
    "status_code": "NOERROR",
    "timestamp": "2026-07-31T13:53:52.873667315-03:00",
    "query-time": "7ms"
}
```

The fields that matter first:

- `host`: the queried hostname;
- `status_code`: `NOERROR` means the resolver answered successfully;
- `a`: current IPv4 addresses;
- `ttl`: how long a resolver may cache the record;
- `resolver` and `timestamp`: where and when this observation was made;
- `query-time`: how long the resolution took;
- `all`: the raw DNS answer, useful when the compact fields need checking.

Addresses, TTL, resolver, and ordering change. This query returned IPv4
records. Add `-aaaa` for IPv6 records and `-cname` when CNAME chains matter.
Interpret the result in context:

- an A or AAAA record shows current resolution;
- a CNAME may identify a SaaS, CDN, or abandoned dependency;
- NXDOMAIN, SERVFAIL, and unresolved CNAME chains do not establish a usable
  address; investigate them before classifying the candidate;
- repeated answers for random labels may indicate wildcard DNS;
- shared-provider resolution does not prove target ownership.

For a reviewed list, pass it to `dnsx -l` and retain JSONL (one JSON object per
line) only when the assessment needs a record of every response. Catch-all
applications and wildcard DNS still require comparison against random-host
baselines.

## 3. Probe, Fingerprint, and Prioritize Services

Start with port discovery on the resolved addresses. The top 1000 TCP ports
cover the common web, database, and administrative service range:

```bash
nmap -sT --top-ports 1000 --open -oA http/ports -iL assets/resolved.txt
```

Add every open port to the candidate URL list. Applications on ports 8080,
8443, 8081 often expose administrative interfaces, development systems, or
internal APIs that port 443 does not.

Probe the live services with [httpx](https://github.com/projectdiscovery/httpx),
sending at most two HTTP requests per second:

```bash
httpx -u https://vercel.com \
  -silent -rate-limit 2 \
  -status-code -title -tech-detect -server -ip -cname -json \
  | jq '{url, status_code, title, tech, webserver, host_ip, cdn, cdn_name}'
```

```json
{
  "url": "https://vercel.com",
  "status_code": 200,
  "title": "Agentic Infrastructure - Vercel",
  "tech": [
    "Amazon S3",
    "Amazon Web Services",
    "HSTS",
    "Next.js",
    "Node.js",
    "React",
    "Vercel",
    "Webpack"
  ],
  "webserver": "Vercel",
  "host_ip": "64.239.123.1",
  "cdn": true,
  "cdn_name": "vercel"
}
```

Read the main fields first:

- `status_code`: whether the service answered;
- `title`: a quick hint about the application. Here it confirms the Vercel
  marketing site;
- `tech`: technology stack. Next.js + React + Node.js + Webpack is the
  frontend stack; AWS + Vercel is the infrastructure layer. Confirm each
  through a second signal;
- `webserver`: the `Server` header. Vercel sets its own;
- `host_ip` and `cdn`: where the request was answered. Vercel's edge
  network IP, not the application origin. When `cdn` is `true`, the
  origin server is behind Vercel, Cloudflare, Fastly, or a similar edge
  network — WAF rules, rate limits, and cached responses may mask the
  real application behavior. Use `httpx -cdn` on every live URL to flag
  CDN/WAF-fronted services early.

The response is reachable, but that does not establish ownership or a
vulnerability. Check the TLS service with
[tlsx](https://github.com/projectdiscovery/tlsx):

```bash
tlsx -u vercel.com -silent -san -nc -json | python3 -m json.tool
```

```json
{
    "timestamp": "2026-07-31T13:55:08.375494478-03:00",
    "host": "vercel.com",
    "ip": "64.239.109.193",
    "port": "443",
    "probe_status": true,
    "tls_version": "tls13",
    "cipher": "TLS_AES_128_GCM_SHA256",
    "not_before": "2026-07-24T06:47:26Z",
    "not_after": "2026-10-22T06:47:25Z",
    "subject_dn": "CN=*.vercel.com",
    "subject_cn": "*.vercel.com",
    "subject_an": [
        "*.vercel.com",
        "vercel.com"
    ],
    "serial": "06:AE:5B:CC:E6:39:F3:82:DA:20:25:70:BB:DB:7A:EE:3A:C6",
    "issuer_dn": "CN=YR2, O=Let's Encrypt, C=US",
    "issuer_cn": "YR2",
    "issuer_org": [
        "Let's Encrypt"
    ],
    "wildcard_certificate": true,
    "tls_connection": "ctls",
    "sni": "vercel.com"
}
```

`subject_an` are the SANs found on the certificate; `tls_version` and
`cipher` fingerprint the TLS stack; `not_before` and `not_after` show
certificate validity; `issuer_org` identifies the CA.

For a hostname list, replace `-u vercel.com` with `-l hostnames.txt`. Inspect
the certificate names with [OpenSSL](https://github.com/openssl/openssl):

```bash
openssl s_client -connect vercel.com:443 -servername vercel.com \
  </dev/null 2>/dev/null \
  | openssl x509 -noout -ext subjectAltName
```

```text
X509v3 Subject Alternative Name:
    DNS:*.vercel.com, DNS:vercel.com
```

The wildcard entry is a hostname pattern, not a newly in-scope target. New
names found in certificate SANs return to the candidate inventory and repeat
the same scope, ownership, and DNS checks.

Treat these as separate facts:

```text
hostname exists
  != service responds
  != response belongs to the expected virtual host
  != application belongs to the target
  != application is vulnerable
```

A reverse proxy can add headers, a default virtual host can supply the title,
and a CDN address does not disclose the origin. Confirm important technology
fingerprints through more than one signal.

When an in-scope IP serves a default site or several known names share an
address, compare candidate virtual hosts against a random-host baseline.
Account for the HTTP Host header, HTTP/2 `:authority`, and TLS SNI. Use
httpx with `-vhost` to test candidates:

```bash
httpx -u https://<IP> -H "Host: app.example.com" -vhost -json
```

Compare the response body length, title, and status code against the
baseline (a random non-existent hostname). Divergence signals a real
virtual host. Bounded virtual-host discovery is appropriate only when the
IP and candidate namespace are covered by the rules of engagement.

Prioritize unusual applications, administrative interfaces, exposed
development systems, identity boundaries, and services that connect multiple
parts of the target. Do not prioritize solely by an old version string or a
dramatic scanner label.

## 4. Map the Application, APIs, and Identity

Map routes, parameters, JavaScript, API operations, authentication flows,
roles, objects, and state transitions. The goal is to know what an application
does before selecting a vulnerability test.

For each application, answer:

| Area | Questions |
|---|---|
| Entry points | Which hosts, ports, paths, and protocols reach it? |
| Client | Which routes, bundles, source maps, workers, and configuration values ship to the browser? |
| API | Which operations, parameters, objects, versions, and content types exist? |
| Identity | Which login, MFA, OAuth, SAML, recovery, and account-linking flows exist? |
| Authorization | Which roles, tenants, and ownership boundaries protect each object? |
| State | Which operations read, create, change, upload, send, or delete? |
| Dependencies | Which storage, SaaS, cloud, WebSocket, or third-party services are connected? |

### Crawl the Current Application

Start with one target application. This bounded crawl follows links and
extracts JavaScript references without rendering a browser session:

```bash
katana -u https://vercel.com \
  -silent -jc -c 2 -p 2 -rl 2
```

```text
https://vercel.com
https://community.vercel.com
https://vercel.com/login
https://vercel.com/oauth/device
https://vercel.com/api/product-og?product=overview
https://vercel.com/api/front-domains/checkpoint
https://vercel.com/_stream/external/cdn.sift.com/s.js
```

This already shows the `community.vercel.com` subdomain, a login path,
an OAuth device flow endpoint, two API routes (`product-og`,
`front-domains/checkpoint`), and a third-party script from `sift.com`.
Review crawler output before testing: it can contain duplicates, stale
links, and strings that only resemble URLs. Group confirmed routes by
purpose: API, authentication, account recovery, administration, uploads,
WebSockets, and state-changing actions. Historical sources can add
retired routes and old parameters, but those still need verification.

Read `robots.txt`, `sitemap.xml`, API descriptions, and framework documentation
when present. They add routes to the map; they do not make those routes public
or in scope.

Choose a host, wordlist, rate, filters, and stop conditions before directory
or parameter discovery. Derive a focused wordlist from the observed framework,
routes, and API vocabulary instead of using a large generic list by default.

### Map JavaScript-Rendered Applications

Standard crawlers do not see every route in a single-page application. Use a
browser or [Katana's headless mode](https://docs.projectdiscovery.io/opensource/katana/running)
when JavaScript rendering, XHR, WebSockets, client-side routing, or
authenticated workflows change the surface.

An authenticated browser pass should use a test account and a
dedicated profile. Capture:

- DOM-rendered routes and forms;
- XHR and fetch operations;
- WebSocket endpoints and message types;
- redirects, logout, refresh, and tenant switching;
- state-changing operations and their confirmation boundaries.

Do not paste production cookies into shell history or commit browser profiles.
Browser automation reproduces user actions; it does not make those actions
read-only.

### Inspect JavaScript Bundles

A client-code review has five steps:

```text
route or script reference
  -> confirm JavaScript response
  -> read routes and configuration
  -> check source maps
  -> verify the backend behavior
```

A URL ending in `.js` is not necessarily a bundle. Vercel serves
Next.js chunks under a nested base path. Confirm status and content type
before reading:

```bash
curl -sS -D - -o /dev/null \
  https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js \
  | sed -n '1,6p'
```

```text
HTTP/2 200
content-type: application/javascript; charset=utf-8
access-control-allow-origin: *
cache-control: public, max-age=31536000, immutable
```

### Extract and Read Bundle Content

Pull every script reference from the page, then download the
largest application bundles. Skip third-party hosts and inline blobs:

```bash
curl -sS --max-time 10 https://vercel.com 2>/dev/null \
  | grep -oP 'src="(https?://[^"]+\.js[^"]*|/[^"]+\.js[^"]*)"' \
  | sed 's/src="//;s/"//' \
  | sort -u > bundles.txt

head bundles.txt
```

```text
/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/06lcjp8ev1jh2.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0_9e5qow9mkg0.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0c0hxoamwjsbw.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0gjs2lbzs80w5.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0nufcqvum7hyr.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0o61-vosflpug.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0tk7awe1kf6up.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0y53lbtp3fja_.js
/vc-ap-vercel-marketing/_next/static/immutable/chunks/0y5fvpaxb21lx.js
```

Every script on this page is path-relative — Vercel serves its own bundles
and third-party integrations through backend proxying, not direct `<script>`
tags. No filtering step is needed. Prepend the base URL and download:

```bash
while IFS= read -r url; do
  curl -sS --max-time 10 "https://vercel.com${url}" \
    -o "bundle-$(echo "$url" | md5sum | cut -c1-8).js"
  sleep 1
done < bundles.txt

ls -S bundle-*.js | head -5
```

```text
bundle-1f818802.js
bundle-ef012ab2.js
bundle-42ca781a.js
bundle-4dcec42f.js
bundle-1ca91cd1.js
```

Not every bundle carries application logic. The largest file
(`bundle-1f818802.js`, 112KB) is framework runtime code. Open a few
of the top bundles and grep for routes and product features:

```bash
# Application routes — paths that reveal product surface
grep -oP '"(/[a-z][a-z0-9/_-]{3,})"' bundle-ef012ab2.js | sort -u
```

```text
"/about"
"/account/settings"
"/agent"
"/ai-gateway"
"/ai-sdk"
"/blog"
"/botid"
"/chat"
"/connect"
"/contact/sales"
"/customers"
"/dashboard"
"/docs/flags/flags-sdk-reference"
"/docs/frameworks/backend/fastapi"
"/docs/frameworks/backend/nitro"
"/docs/frameworks/full-stack/nuxt"
"/docs/frameworks/full-stack/sveltekit"
"/docs/rest-api/sdk"
"/domains"
"/drop"
"/enterprise"
```

Twenty-one routes from one bundle. `/agent`, `/ai-gateway`, `/ai-sdk`,
and `/botid` are product features. `/dashboard` and `/enterprise` are
authenticated pages. `/docs/rest-api/sdk` hints at an API client.
Framework integration docs expose the platform's dependency surface.

A second grep finds subdomains referenced in the same bundle:

```bash
# Subdomains and external services
grep -oP '[a-zA-Z-]+\.vercel\.com' bundle-ef012ab2.js | sort -u
```

```text
community.vercel.com
security.vercel.com
```

Two subdomains not in the initial subfinder pass. Add them to the
candidate inventory and repeat scope and DNS checks.

### Search for Credentials

Routes and subdomains are one class of finding. Credentials are another.
While Vercel's landing page returned no matches (expected for a
well-configured target), a bundle from another target might print:

```bash
grep -oP 'AIza[0-9A-Za-z_-]{35}' bundle-*.js
```

```text
bundle-abc123.js:AIzaSyDrN5CwGpPqMkLxVnRtZu3j8Wh2Bf7Ea9Y
```

A single Google API key exposed in one bundle. The key type (Maps, Cloud,
Firebase) determines what it grants. Test with the least-privileged
documented request before concluding impact.

The full credential grep set is:

```bash
for f in bundle-*.js; do
  [ -s "$f" ] || continue
  grep -HoPi '(?:x-)?api[_-]?key\s*[=:]\s*["\x27][^"\x27]{8,}["\x27]' "$f"
  grep -HoP 'AIza[0-9A-Za-z_-]{35}' "$f"
  grep -HoP 'AKIA[0-9A-Z]{16}' "$f"
  grep -HoPi '(?:sk|pk|rk)_(?:live|test)_[A-Za-z0-9]+' "$f"
  grep -HoPi '(?:firebase|supabase)\s*[=:]\s*["\x27][^"\x27]+["\x27]' "$f"
  grep -HoP 'ghp_[A-Za-z0-9]{36}' "$f"
  grep -HoP 'github_pat_[A-Za-z0-9_]{20,}' "$f"
  grep -HoP 'glpat-[A-Za-z0-9_-]{20,}' "$f"
  grep -HoP 'xox[baprs]-[0-9A-Za-z-]{10,}' "$f"
  grep -HoPi 'sk-(?:proj|svcacct|ant)-[A-Za-z0-9_-]{20,}' "$f"
  grep -HoP 'eyJ[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}' "$f"
  grep -HoP '10\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
  grep -HoP '172\.(1[6-9]|2[0-9]|3[01])\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
  grep -HoP '192\.168\.[0-9]{1,3}\.[0-9]{1,3}' "$f"
  grep -HoP 'https?://[a-zA-Z0-9.-]+:[0-9]{2,5}' "$f"
  grep -HoP '-----BEGIN [A-Z ]+ PRIVATE KEY-----' "$f"
done
```

Each match includes the bundle filename (`-H`) and prints only the
matched text (`-o`). Classify the match before escalating:
`pk_live_` is a public identifier; `sk_live_` is a secret.
A JWT printed by this grep still needs decoding to
inspect claims, issuer, and expiry. JWT uses base64url, not standard
base64 — `base64 -d` will fail on `-` and `_` characters:

```bash
echo 'eyJ...rest...of...token' | cut -d. -f2 | python3 -c "
import sys, base64, json
body = sys.stdin.read().strip()
# base64url → standard padding
body += '=' * (4 - len(body) % 4)
print(json.dumps(json.loads(base64.urlsafe_b64decode(body)), indent=2))
"
```
marketing page probably powers a Maps embed; the same key in a backend
bundle is privileged. Context turns a grep match into a finding.

### Search for External Dependencies and Source Code

Credentials are one class of finding. Dependencies and linked
repositories are another. This grep extracts external hosts, analytics IDs,
GitHub repos, and social handles from the page HTML:

```bash
curl -sS --max-time 10 https://vercel.com 2>/dev/null \
  | grep -oP 'src="https://[^"]*\.js"|data-website-id="[^"]*"|https://(github|gitlab)\.com/[^"'\''<> ]+|https://x\.com/[^"'\''<> ]+' \
  | sort -u
```

```text
https://github.com/vercel
https://x.com/vercel
```

The GitHub organization and social handle were found in the page source.
`github.com/vercel` links to public repositories: clone the relevant ones
and audit for hardcoded keys, internal hostnames, CI/CD configuration,
Dockerfiles, and environment templates. `x.com/vercel` is documented in the
asset inventory and excluded from testing.

### Recover Source Maps

A source map is a JSON file that maps minified bundle code back to the
original sources. When exposed, it recovers readable TypeScript, JSX, and
configuration files — the code the developer wrote, not what the bundler
emitted.

Check the map next to each confirmed bundle by appending `.map`:

```bash
curl -sS -o /dev/null -w '%{http_code} %{content_type}\n' \
  https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/044gqk9l3w4as.js.map
```

```
403 text/plain
```

This bundle returned 403. The map either does not exist, is blocked by
WAF, or is served only with a referrer check. Both 403 and 404 are
common and expected results. Move on to the next bundle.

When a map returns `200` with `application/json`, download it and expand
the sources. Replace the chunk URL below with the one that returned 200:

```bash
curl -sS https://vercel.com/vc-ap-vercel-marketing/_next/static/immutable/chunks/THE-CHUNK-THAT-RETURNED-200.js.map \
  -o chunk.js.map

python3 - <<'PY'
import json
from pathlib import Path

source_map = json.load(open("chunk.js.map"))
output = Path("recovered-source")
output.mkdir(exist_ok=True)

for idx, (name, content) in enumerate(zip(
    source_map.get("sources", []),
    source_map.get("sourcesContent", [])
)):
    if content:
        fname = f"{idx:04d}-{Path(name).name}"
        (output / fname).write_text(content)
        print(fname)
PY
```

```text
0000-webpack-runtime.js
0001-_app.tsx
0002-index.tsx
0003-api-client.ts
0004-auth-config.ts
0005-feature-flags.json
```

Seven recovered files from one map. `api-client.ts` and `auth-config.ts`
are the immediate targets: they contain request construction, base URLs,
OAuth endpoints, and environment-specific logic that a crawl never reaches.

**What to search in recovered sources:**

```bash
grep -rniP 'baseURL|apiUrl|endpoint|Authorization|clientId|clientSecret|redirectUri|issuer|tenantId|apiKey' recovered-source/
grep -rniP 'https?://[a-zA-Z0-9.-]+\.[a-z]{2,}' recovered-source/
grep -rniP 'secret|password|token|credential' recovered-source/
```

A source map converts a minified blob into a readable codebase in seconds.
If the current bundle has no map, review historical JavaScript URLs from
`waybackurls` and dynamic imports loaded through `import()`. Older deployed
versions sometimes exposed maps that were later blocked.

### Classify and Validate Values

Classify a value before replaying it. `pk_live_` and OAuth `client_id` values
are public client identifiers; `sk_live_`, service-role keys, private keys, and
client secrets are not. A Firebase or Supabase anonymous key may be expected
in a browser bundle, but its allowed backend behavior is still testable.

For a discovered Supabase project, compare a minimal read request with and
without the anonymous key. Do not query application tables until the scope and
test account support it:

```bash
curl -sS -D - -o /dev/null \
  https://<project>.supabase.co/rest/v1/

curl -sS -D - -o /dev/null \
  -H 'apikey: <supabase-anon-key>' \
  https://<project>.supabase.co/rest/v1/
```

Record the status, headers, endpoint, key type, and what the provider permits.
The finding is excessive access, not the presence of an anonymous key. Treat a
backend hostname such as `*.fly.dev`, `*.supabase.co`, or `*.r2.dev` as a new
service candidate: resolve it, confirm ownership, then map it like any other
application.

Minified code is still evidence. A source map makes it easier to read, but a
missing map does not end the review: follow the request construction, headers,
route strings, and provider configuration in the bundle. Do not report a
string match as a secret, endpoint, or vulnerability without identifying its
use and testing the relevant server-side behavior.

When source code is available, compare its route definitions and HTTP clients
with browser traffic. This finds entry points that a crawl did not reach and
also exposes dead or feature-gated routes that still need classification.

### Build an API and Identity Matrix

Check for descriptions such as OpenAPI, Swagger, GraphQL, and framework
documentation. Capture browser traffic in an intercepting proxy or request
workspace. When no schema exists, derive the inventory from browser traffic,
JavaScript, mobile clients when in scope, archives, and error responses.

For every route, record the request inputs and response behavior, not only the
path:

```text
GET    /api/v1/projects/{project_id}   path: project_id       user session     200, 401, 404
PATCH  /api/v1/projects/{project_id}   body: name, status     manager session  200, 403
POST   /api/v1/projects/{id}/members   body: member_id        manager session  201, 403
GET    /graphql                        query: operationName   bearer token     200, 401
WS     /socket                         cookie: session         session cookie   upgrade, close
```

For each page or flow, capture:

- query parameters, path values, request bodies, and hidden form fields;
- custom request headers, cookies, and response `Set-Cookie` attributes;
- redirects, error responses, and the final destination;
- observed HTTP methods and any method that needs a separate check;
- WebSocket upgrades, message types, and connection boundaries;
- multi-step paths such as login, MFA, recovery, checkout, and tenant changes.

Do not infer that a method is safe to test because it appears in a response or
schema. Reads, writes, uploads, and destructive actions belong in different
test plans.

For authenticated applications, define a small identity matrix:

| Identity | Session | Expected access |
|---|---|---|
| Anonymous | None | Public routes only |
| User A | Test tenant A | Own objects |
| User B | Test tenant B | Own objects |
| Manager | Privileged test account | Team or tenant objects |

Many authorization findings require two identities, two tenants, or a
privileged role. Without the required control account, label the path untested
instead of extrapolating from one response.

## 5. Use Automation to Create Hypotheses

After mapping, use automation to create a small, reviewable hypothesis queue.
Do not begin with every available template or scanner.

[Nuclei](https://docs.projectdiscovery.io/opensource/nuclei/running) matches
known exposures, misconfigurations, technologies, and CVE preconditions.
Template execution is code-driven testing: review the selected templates,
exclude intrusive categories, identify traffic when required, and set both
request and concurrency limits.

Start with templates relevant to an observed technology:

```bash
# Technology-specific templates (the allowlist approach)
nuclei -l live-urls.txt \
  -t ~/nuclei-templates/ \
  -tags cve,exposure,misconfig \
  -etags intrusive,fuzz,dos,bruteforce \
  -rate-limit 5 -c 2 \
  -duc -j -o nuclei.jsonl
```

If the application map found an OpenAPI schema, feed it to the
interactive mode:

```bash
nuclei -l live-urls.txt -im openapi
```

For blind vulnerability classes (SSRF, XXE, blind XSS), run with an
out-of-band callback to confirm reachability:

```bash
nuclei -l live-urls.txt \
  -t ~/nuclei-templates/ \
  -tags oob,ssrf,xxe \
  -interactsh-server oast.pro
```

Record the target, template ID, matcher, request, and response. A
template match is a lead, not a finding: an `info` result can expose an
important administrative surface, while a `critical` match can fail
manual reproduction.

For tight production rules, use an allowlist of reviewed templates. Headless,
code, fuzzing, authenticated, out-of-band, and state-changing templates need
separate confirmation. Pin the tool and template revision when reproducibility
matters.

Treat a template match as a lead:

```text
template match
  -> inspect matcher and prerequisite
  -> reproduce with a bounded request
  -> compare a negative control
  -> demonstrate impact or discard
```

Automated output should shorten the manual queue. It does not replace
application mapping, business-logic review, or role comparison.

## 6. Route Findings Through the WSTG Categories

The [OWASP Web Security Testing Guide](https://owasp.org/www-project-web-security-testing-guide/)
organizes web application testing into 12 categories. Use this routing table to
match what you observed to the right test procedure. Every observation in
columns 1 and 2 maps to a WSTG category and a `recon-skills` procedure.

The complete methodology with full commands, expected outputs, and pitfalls
per category lives in [`wstg-web-pentest`](https://github.com/uphiago/recon-skills/tree/main/redteam/wstg-web-pentest).

### Phase Routing Table

| WSTG Category | What you observed | Test | Skills to consult |
|---|---|---|---|
| **INFO** Information Gathering | Hostnames, services, headers | Fingerprint servers, map architecture, discover entry points | `web2-recon`, `wstg-web-pentest` |
| **CONF** Configuration | Exposed headers, backup files, admin panels | Test HTTP methods, enumerate config files, check security headers, subdomain takeover | `hunt-source-leak`, `hunt-subdomain`, `hunt-tls-network`, `wstg-web-pentest` |
| **IDNT** Identity Mgmt | Registration, password reset | Account enumeration, weak username policy, role definition gaps | `hunt-brute-force`, `hunt-ato` |
| **ATHN** Authentication | Login, MFA, password reset flows | Default credentials, weak lockout, bypass schema, MFA gaps, weak reset tokens | `hunt-auth-bypass`, `hunt-mfa-bypass`, `hunt-brute-force` |
| **ATHZ** Authorization | Object IDs in URLs, role-based UI, OAuth/SAML redirects | IDOR, privilege escalation, forced browsing, directory traversal, OAuth redirect_uri, SAML assertion tampering | `hunt-idor`, `hunt-lfi`, `hunt-oauth`, `hunt-saml` |
| **SESS** Session Mgmt | Cookies, JWT, logout | Cookie attributes, session fixation, CSRF, JWT attacks, logout invalidation | `hunt-session`, `hunt-csrf` |
| **INPV** Input Validation | Query parameters, forms, API bodies | XSS, SQLi, SSTI, SSRF, command injection, mass assignment, prototype pollution, HTTP smuggling, host header injection | `hunt-xss`, `hunt-sqli`, `hunt-ssrf`, `hunt-ssti`, `hunt-rce`, `hunt-host-header`, `hunt-http-smuggling`, `hunt-api-misconfig`, `hunt-prototype-pollution`, `hunt-mass-assignment` |
| **ERRH** Error Handling | Stack traces, debug pages, verbose errors | Trigger malformed input, check for information leakage in error responses | `hunt-information-disclosure` |
| **CRYP** Cryptography | TLS, encrypted tokens, hashing | Weak TLS versions, padding oracle, unencrypted channels, weak cipher suites, expired certificates | `hunt-tls-network` |
| **BUSL** Business Logic | Checkout, coupons, file upload, workflows | Negative quantities, price manipulation, race conditions, workflow bypass, file upload abuse | `hunt-business-logic`, `hunt-file-upload`, `hunt-race-condition` |
| **CLNT** Client-side | CORS headers, WebSocket upgrade, iframe, postMessage | CORS misconfig, DOM XSS, clickjacking, WebSocket auth, browser storage inspection, CSS injection | `hunt-cors`, `hunt-websocket`, `hunt-dom` |
| **APIT** API Testing | REST, GraphQL, OpenAPI schema | BOLA, BFLA, excessive data exposure, schema enumeration, GraphQL introspection | `hunt-graphql`, `hunt-api-misconfig`, `hunt-idor`, `hunt-schema-enumeration`, `hunt-broken-function-level-auth` |

### When to Skip a Category

Do not test every category because the list exists. Route by evidence:

```text
Observed: JSON response with sequential object IDs
  -> ATHZ (Authorization): IDOR test with two identities
  -> Skip: ATHN (Authentication) because the login flow is unchanged

Observed: Origin reflected with ACAC on a profile endpoint
  -> CLNT (Client-side): CORS credential leakage test
  -> Skip: BUSL (Business Logic) because no transaction flow exists
```

Load a skill only when its prerequisite evidence exists. Confirm conflicting
fingerprints before selecting platform-specific tests.

Start with one procedure and add another only when the result creates a new
question.

## 7. Validate One Hypothesis at a Time

A hypothesis names the security property, expected failure, supporting
evidence, and cheapest safe falsification step:

```text
If:       user A can address an object owned by user B
Then:     the API may return cross-tenant data
Because:  the object ID is accepted directly by a read endpoint
Control:  compare owner, non-owner, anonymous, and unknown-object responses
```

Keep a short queue:

| Hypothesis | Evidence so far | Cost | Safety | Status | Next control |
|---|---|---|---:|---|---|---|
| API object may cross a tenant boundary | Predictable ID in browser traffic | Low | Read-only | Open | Compare two synthetic tenants |
| CORS may expose profile data | Origin reflected on `/api/me` | Low | Read-only | Open | Browser proof with test session |
| Source map may reveal an unlinked API | Map and API client found | Low | Read-only | Validated | Resolve and classify routes |
| Component may match a known CVE | Version and feature prerequisite observed | Medium | Varies | Dismissed | Reproduce the vulnerable behavior |
| Upload may reach executable storage | Upload route exists | High | State-changing | Confirm scope and use an inert file |

Prioritize strong evidence, meaningful potential impact, cheap falsification,
and limited side effects. Do not prioritize by the most dramatic imagined
outcome.

### CORS Example

Origin reflection is a lead:

```bash
curl -sS --max-time 10 \
  -D "evidence/cors-headers.txt" \
  -o "evidence/cors-body.txt" \
  -H 'Origin: https://attacker.example' \
  "https://app.example.test/api/me"
```

A defensible CORS finding establishes:

1. the endpoint returns user-specific or sensitive data;
2. a browser session for a test account supplies the required credentials;
3. the response accepts an untrusted origin;
4. credentials are allowed when the scenario requires them;
5. browser JavaScript can read the response without CORP blocking;
6. SameSite cookie behavior does not silently strip the session;
7. `Vary: Origin` is absent when the response caches differ by origin;
8. a preflight check (PUT, PATCH, custom headers) succeeds or is absent
   when the method requires none;
9. public, anonymous, and trusted-origin controls behave differently.

`Access-Control-Allow-Origin: *` is not credential theft because the
fetch spec prohibits browsers from combining `*` with
`Access-Control-Allow-Credentials: true`. An attacker who controls the
network (MITM on shared WiFi) can bypass this restriction, but the
threat model differs from a standard reflected-origin CORS finding.
Reflection on an error response may have no impact. A header scanner
does not prove browser-readable sensitive data.

### Authorization Example

Changing `project-a` to `project-b` is meaningful only when both are synthetic,
belong to test identities or tenants, and the expected policy is known:

```text
expected: user A receives a denial for project B
observed: user A receives project B's non-public synthetic fields
control 1: user B can read project B
control 2: an unknown project returns the documented not-found response
control 3: an anonymous request is rejected
not tested: enumeration, production objects, writes
```

The same discipline applies throughout the pack:

- `200 OK` on `/.env` is not a leak if the body is an application shell;
- an old version does not prove the vulnerable component is reachable;
- an exposed database port is not remote code execution;
- a public client identifier is not automatically a credential;
- an API schema is not an authorization vulnerability;
- an accepted write method does not authorize arbitrary modification.

## 8. Build Branching Attack Paths

Link results only when one supplies a prerequisite or capability for the next.
Label each link as observed, inferred, confirmed, or not tested.

A path should branch where the evidence branches:

```text
OpenAPI operation discovered
|
+-- anonymous request succeeds
|   `-- confirm data sensitivity and ownership
|
+-- anonymous request is denied
|   `-- compare user A and user B on synthetic objects
|
`-- endpoint is no longer reachable
    `-- record historical context and stop
```

For example:

```text
1. Observed
   app.example.test loads /assets/app.js.

2. Observed
   The bundle references api.example.test/v2 and a project route.

3. Observed
   /openapi.json documents GET /projects/{project_id}.

4. Confirmed
   User A can read project A; user B can read project B.

5. Confirmed
   User A can also read project B's synthetic billing contact.

6. Not tested
   Object enumeration, production records, and write operations.
```

The bundle and schema supply the API base URL, operation, identifier, and
expected object shape. The security failure occurs at step 5; hiding the
documentation would not fix object-level authorization.

Do not score an inferred final step as demonstrated impact. If the output of one
finding is not required by the next, the findings are co-located, not chained.
Use `cross-attack-chains` only after component findings pass
`triage-validation`.

## 9. Capture, Sanitize, Score, and Report

For each validated behavior, record enough for another operator to reproduce
the result:

- affected asset, identity, environment, and UTC timestamp;
- exact steps with sanitized requests and responses;
- expected, observed, and control behavior;
- prerequisites, side effects, demonstrated impact, and testing limits;
- remediation and a verification step;
- tool versions and evidence hashes when integrity tracking is required.

A compact finding can include all of this:

```markdown
## Cross-tenant read of a synthetic project

**Affected asset:** `https://api.example.test/v1/projects/{project_id}`
**Severity:** High
**Classification:** CWE-639

**Expected:** User A cannot read an object owned by tenant B.

**Observed:** User A received project B's synthetic billing contact.

**Controls:** User B could read the object; an anonymous request returned 401;
an unknown identifier returned the documented not-found response.

**Impact demonstrated:** Unauthorized read of one synthetic cross-tenant
record.

**Not tested:** Enumeration, production records, and write operations.

**Remediation:** Enforce tenant and object ownership on every project lookup.
Repeat the owner, non-owner, anonymous, and unknown-object controls after the
fix.
```

Severity follows demonstrated impact. CVSS communicates technical severity; it
does not replace business context, asset criticality, exploit preconditions, or
the program's rating policy. Use the version required by the engagement.

Sanitize before sharing:

- remove or replace cookies, bearer tokens, API keys, and passwords;
- minimize personal and customer data;
- replace unrelated records with synthetic examples where possible;
- review screenshots, filenames, browser storage, and response metadata;
- keep a protected original only when retention policy requires it.

Run a final quality gate:

```text
[ ] reproduction and controls work as documented
[ ] impact and severity match the evidence; untested steps are labeled
[ ] credentials and unrelated data are removed
[ ] remediation addresses the failed security control
```

## 10. Use and Extend Recon Skills

Use [recon-skills](https://github.com/uphiago/recon-skills) as a procedure
library. Read the skill that matches the observed surface, follow its
prerequisites and verification, and update the owning skill when field work
produces a reusable improvement without target details.

Start with these entry points:

| What you need | Skill |
|---|---|
| Full WSTG-aligned methodology with commands per phase | [`wstg-web-pentest`](https://github.com/uphiago/recon-skills/tree/main/redteam/wstg-web-pentest) |
| Asset discovery and service fingerprinting | [`web2-recon`](https://github.com/uphiago/recon-skills/tree/main/redteam/web2-recon) |
| Vulnerability-class testing (XSS, SQLi, SSRF, IDOR, etc.) | `hunt-*` skills under `redteam/` |
| WordPress-specific mapping and validation | [`hunt-wordpress`](https://github.com/uphiago/recon-skills/tree/main/redteam/hunt-wordpress) |
| Compound attack paths from verified findings | [`cross-attack-chains`](https://github.com/uphiago/recon-skills/tree/main/chains/cross-attack-chains) |
| Finding quality gate before reporting | [`triage-validation`](https://github.com/uphiago/recon-skills/tree/main/redteam/triage-validation) |
| Sanitized evidence and report structure | [`report-writing`](https://github.com/uphiago/recon-skills/tree/main/redteam/report-writing)

## 11. Close the Assessment

When the assessment becomes noisy, return to these questions:

```text
What did I observe?
What security property might be broken?
What is the cheapest safe test that could disprove it?
Which identity, object, or environment does that test require?
What control distinguishes the result from expected behavior?
What new capability does the confirmed result provide?
```

Then collect missing context, run one bounded validation, capture the result,
or stop because the next step is unsafe, out of scope, or unsupported.

## Useful Links

- [OWASP WSTG (Web Security Testing Guide)](https://owasp.org/www-project-web-security-testing-guide/)
- [OWASP WSTG Testing Categories](https://owasp.org/www-project-web-security-testing-guide/latest/4-Web_Application_Security_Testing/)
- [WSTG-aligned Pentest Skill](https://github.com/uphiago/recon-skills/tree/main/redteam/wstg-web-pentest)
- [Recon Skills](https://github.com/uphiago/recon-skills)
- [ProjectDiscovery documentation](https://docs.projectdiscovery.io/opensource)
- [Nuclei Templates](https://github.com/projectdiscovery/nuclei-templates)
- [PortSwigger Web Security Academy](https://portswigger.net/web-security)
- [SecLists](https://github.com/danielmiessler/SecLists)
- [CVSS v4.0](https://www.first.org/cvss/v4.0/)
- [OWASP Top 10 (2021)](https://owasp.org/www-project-top-ten/)
- [OWASP API Security Top 10](https://owasp.org/www-project-api-security/)
- [HackTricks](https://book.hacktricks.xyz/)

## Reach Me

Questions, corrections, and field notes: [@uphiago](https://x.com/uphiago)
