The question a Web pentest answers
Every run is tied to one repository, one exact URL, and one point in time. The runbook states the question directly: How safely did the exact authorized website behave under this Web pentest?
That is a different question from the one the Security audit asks. The Security audit reads the repository and, at most, sends a passive request to a linked site to record headers and status codes. The Web pentest audit exercises the deployed application: it walks public workflows, sends crafted requests to the routes it discovers, and keeps a proof record for each behavior it can reproduce. Repository code is used to understand the product and to avoid false alarms, but a line of code never becomes a pentest finding on its own.
The answer belongs to the deployed revision that was reachable during the run. Deploying a fix and running the audit again is the only way to change it. An issue, a branch, or an unmerged pull request records that work started; it does not move the score.
One High access-control gap lets a visitor read another customer's booking. Three Medium and two Low findings keep this site below the healthy band.
What Enji Guard collects
The task carries a repository, its URL, and one Web pentest target URL stored exactly as the product supplied it. Enji Guard does not canonicalize, follow, or rewrite that string. A missing or ambiguous target is a task failure: nothing is probed, no substitute host is chosen, and no successful score is published because the identity of the target cannot be established.
With the target fixed, preflight clones the repository, records the default branch and full commit, and collects a bounded slice of provider history: open and closed issues, open, closed, and merged pull or merge requests, filtered by security terms such as auth, XSS, SSRF, SQL injection, IDOR, access control, and secrets. Comments and closing notes are read for the relevant matches so later stages can avoid re-reporting a known finding or repeating a fix pattern that maintainers rejected. Preflight records facts about that history; it does not decide whether a prior finding is valid.
Preflight also fingerprints the project from manifests, lockfiles, container files, and the target's observable nature, then prepares the security tooling the detected stack needs: specialized target probing, recognition of known vulnerability and misconfiguration patterns, and source or dependency analysis alongside manual browser and HTTP work. Tools are installed outside the repository from pinned, integrity-checked sources and only asked for a version or self-check at this stage. A tool that cannot be obtained through a trusted path is recorded as a coverage limitation rather than replaced with an unverified download.
During the run itself Enji Guard collects the surface it can observe on the exact target: public routes, API calls made by the client bundle, forms, cookies, headers, same-target redirects, visible errors, and access barriers such as login walls, one-time codes, geo-blocking, or maintenance responses. Each active check that reveals a behavior is kept as a proof record with the request shape, reproduction steps, observed response, time, access conditions, plausible impact, and confidence.
What the audit deliberately does not do
Active does not mean unrestricted. The rules of engagement forbid denial of service, load or resource-exhaustion tests, brute force and credential stuffing, spam, persistence or backdoors, lateral movement, credential theft, social engineering, destructive or material data changes, and exfiltration of real user data. Automated probing defaults to a cap of ten requests per second per endpoint. If probing starts to cause visible instability, the run stops active work, records the limitation, and moves on to reporting.
Scope is equally narrow. Only the exact target URL is tested. Subdomains, linked domains, third-party services, unrelated IP addresses, and external redirect destinations are recorded as out of scope, never followed as new targets. Test accounts are used only when the task or repository provides them, and the run does not self-register accounts on a production system.
The audit is assessment-only. It creates no repository files, issues, comments, branches, commits, pull requests, deployments, or accounts. It does not repeat the repository Security audit under another name: code paths are inspected only where they connect to the external surface, and repository hardening that has no externally observable effect stays with Security. It also does not claim to replace a contracted manual penetration test for a high-risk system; it keeps the deployed site under recurring, bounded observation.
Two failure modes are kept apart on purpose. A target that cannot be identified is a task failure with no result. A correctly identified target that turns out to have no meaningful surface to exercise, because it is down, fully behind authentication, or blocked, still produces an ordinary result: zero findings, a fail-closed score of zero, and plain wording that nothing meaningful was verified.
How the audit works
The runbook runs as a chain of five stages. Each stage repeats the authorization boundary it inherited, reads only the saved state of the stages before it, and validates that repository identity, revision, and target URL still agree before doing anything.
- 1
Target, repository, history, and tool preflight
Resolve the exact target URL and repository revision, collect bounded provider history as facts, fingerprint the project, and prepare trusted tooling without running scans or contacting the site.
- 2
Repository-assisted analysis
Inspect routes, handlers, auth and session code, middleware, uploads, payments, admin and debug surfaces, and configuration connected to the target. Classify each internal candidate with a stable identifier as needs external confirmation, not externally relevant, duplicate of existing work, already fixed, or manual review useful. Nothing is confirmed here, and the site is not contacted.
- 3
Target understanding and attack plan
Visit only the exact URL, map its public surface and access barriers, give every internal candidate one disposition, add target-derived vectors when observed behavior justifies them, and record how seven risk categories are covered, not applicable, or blocked.
- 4
Active probing
Exercise real workflows before firing payloads, test the highest-value vectors first, pivot when new endpoints or errors appear, re-check tool results by hand, and write one proof record per reproducible behavior observed on the exact target.
- 5
Adjudication, score, report, and handoff
Merge candidates, plan dispositions, proof records, and provider history into one canonical assessment; deduplicate by root behavior; assign severity; compute the score; write the report, summary, and findings; compare with the previous run for the same site; and save the bounded context that Web pentest Autofix will read.
Coverage is planned against seven categories
The attack plan is target-specific rather than a fixed checklist, but before it is finalized the run has to say, for each of seven web risk categories, whether at least one planned vector covers it, whether it does not apply to this target, or whether an access barrier blocks it. A category counts as covered only when a concrete vector addresses it; the plan is not allowed to invent an irrelevant vector to fill the row.
| Category | What it groups |
|---|---|
identity_and_access | Sign-in, session, token, reset, invitation, role, and object-level access behavior. |
input_and_output_handling | Injection, reflected or stored output, parsing, and encoding on visible inputs. |
server_side_interactions | Outbound requests, callbacks, imports, and other server-side fetches a visitor can influence. |
file_data_and_tenant_boundaries | Uploads, downloads, exports, storage paths, and separation between customers or tenants. |
business_workflows | Multi-step flows such as booking, checkout, or approval that can be skipped or replayed. |
operational_and_client_exposure | Debug endpoints, verbose errors, exposed configuration, client bundle secrets, and header hygiene. |
observed_components | Frameworks, libraries, and services identified on the target and their known weaknesses. |
Prepared tools support this plan; they do not drive it. Manual, browser, and HTTP-client verification remains the baseline, scanners run only for scoped, light, target-specific checks inside the rules of engagement, and a tool result that manual or code-context review does not support is recorded as a dead end, not a finding.
From proof record to confirmed finding
A proof record is the unit of evidence. It exists only for behavior observed in the current run against the exact target, and it names the in-scope URL or endpoint, the redacted request or browser action, reproduction steps another operator could follow safely, the minimal non-destructive proof, the observed response, the time and access conditions, a plausible impact, a confidence level, and a remediation direction. Destructive payloads and full sensitive response bodies never enter user-facing output.
The final stage owns classification. Every candidate and observation ends up as confirmed, likely but not proven, duplicate of existing work, already fixed, not externally relevant, or useful for manual review. A confirmed finding must cite at least one proof record from this run. Repository code, an isolated local reproduction, scanner output, a previous report, or an intermediate label can explain or strengthen a finding, but none of them can create one or lower the score by itself.
One root behavior is counted once
Repeated manifestations of the same behavior are merged before scoring. If the same missing ownership check shows up on three booking endpoints, the report may list the representative routes, but the score counts one finding at the highest supported severity. Each finding receives a stable key derived from the target, the normalized weakness class, the affected behavior, and its evidence identity, so a rerun can tell a stable finding from a new one without relying on wording or list position.
Existing provider work does not remove a deduction. If an issue or an unmerged pull request already describes the behavior and the deployed site still exhibits it, the finding stays in the score and the report records the relation so downstream work can reuse that issue instead of opening another.
Reports
62 / 100September 9, 2026Current
Executive Summary
- One High finding lets an unauthenticated visitor read another customer's booking by changing the reference in the URL.
- Score: 62/100 — Some issues. No Critical finding was confirmed under the stated coverage.
- The full fictional run confirms 1 High, 3 Medium, and 2 Low findings; two representative rows appear below.
Audit Target
- Website: https://staging.lumen-booking.example
- Repository: lumen/booking-web (fictional)
- Branch: main
- Checked commit: 3d9f71c2a4e8b06d5f1c7a92e4b3d8f60a1c5e77
- Authorization: Verified site with recorded consent; public, unauthenticated surface only
What Was Checked
- Public booking search and lookup, account sign-in and password reset, the JSON API used by the client bundle, cookies and response headers, error handling on malformed input, and the static upload path, at ten requests per second or less per endpoint.
Findings
| Priority | Finding | Impact | Affected surface |
|---|---|---|---|
| High | Booking lookup returns another customer's reservation when the reference in the URL is changed. | A visitor who knows or guesses a reference can read names, dates, and partial payment details without signing in. | GET /api/bookings/{reference} |
| Medium | Password reset accepts unlimited attempts without a rate limit. | Reset tokens can be guessed faster than the current token length supports. | POST /api/auth/reset |
How the score is calculated
Let C, H, M, and L be the numbers of deduplicated confirmed findings at each severity. Severity comes from demonstrated current impact, exploitability, access prerequisites, affected data and users, and confidence. Useful observations that did not reach confirmation stay in the report as observations or limitations and never enter these counts.
critical = C
high = H
medium = M
low = L
score = 100
score -= critical * 45
score -= high * 16
score -= medium * 6
score -= low * 2
if critical > 0:
score = min(score, 39)
score = clamp(score, 0, 100)Inputs and deductions are integers and the result is an integer; there is no rounding step. One Critical finding imposes a ceiling of 39 regardless of the other counts. A High finding has no ceiling of its own. The final clamp keeps every result between 0 and 100. Repository-only findings, duplicate manifestations, unconfirmed hypotheses, coverage barriers, and the number of checks attempted create no additional deduction.
Display bands
- 70–100: good; the product label is Healthy.
- 40–69: warn; the product label is Some issues.
- 0–39: bad; the product label is Needs attention.
A worked example: 62 out of 100
The fictional lumen/booking-web run against https://staging.lumen-booking.example confirms no Critical finding, one High, three Medium, and two Low deduplicated findings.
0 confirmed findings × 45
01 confirmed finding × 16
-163 confirmed findings × 6
-182 confirmed findings × 2
-4100 - 0 - 16 - 18 - 4
62No Critical ceiling applies and the clamp changes nothing. The same arithmetic explains the fictional history of this site. On August 12 one Critical and one Medium finding gave 100 − 45 − 6 = 49, which the Critical ceiling reduced to 39. On August 26, after the Critical behavior was fixed and redeployed, two High, two Medium, and one Low finding gave 54. On September 9 one High, three Medium, and two Low gave 62. Each number describes what the deployed site did on that day, not how much time had passed.
Previous runs explain movement; they do not set the score
The current findings, score, summary, and report are drafted and validated before any previous result is read. Only then does Enji Guard fetch the earlier context for the same repository and the same exact URL. That comparison can name stable, new, and resolved findings, surface an obvious same-method arithmetic or classification inconsistency, and explain a material change in coverage. It must not anchor, average, or smooth the current score toward the previous one, and a large delta is never by itself a reason to revise.
How the result appears in Enji Guard
The Web pentest card is site-scoped. A repository can have several linked websites, and the card shows one of them at a time: the selector at the top names the exact URL, the dropdown lists every linked site with its latest score or Not started, and switching sites swaps the score, chips, and narrative without mixing them. Two sites on one repository never overwrite or average each other.
The rest of the card follows the shared audit layout: the 0–100 ring, one of the labels Healthy, Some issues, or Needs attention, four severity chips named Critical, High, Medium, and Low with an em dash for zero, and the actions Run audit now, Recurring audit, and Recurring autofixes. Run audit now stays disabled until the selected site is verified, and the first run for a site opens the consent dialog shown above.
Opening the audit leads to the report for the selected site. The report view names the website, repository, and branch, shows the score, and lets a reader switch between linked sites. The report body keeps the exact target, the authorization boundary, the tested and discovered surface, confirmed findings with redacted evidence, the Findings table with its Score calculation, coverage and access limitations, and concrete improvement steps. Copy and Download act on that saved report, and the history view lists every run for that one site so a score can be read against the run that produced it.
A recurring Web pentest is configured per site in automation settings. Adding a site there triggers the same consent dialog, and the schedule cannot be saved until every selected site has its confirmation. Recurring autofixes are enabled per site in the same place, only for verified sites, and each site keeps its own issue or pull-request preference for that work.
Where the audit ends and Web pentest Autofix begins
The audit ends with read-only artifacts: a score, a report, a findings list, and a bounded, target-specific context that names the repository, the exact URL, and the artifacts of this run. For each confirmed finding with a plausible repository remediation the assessment also records a fix candidate: affected components and paths, a defensive problem statement, the expected safe behavior, allowed and excluded scope, repository checks, and non-destructive verification. Attack plans, raw payloads, and instructions to repeat probing are excluded from that candidate on purpose.
Web pentest Autofix is the separately published improvement that reads this context. It computes the expected context for the same repository and URL, validates it against the current task, fetches the exact findings artifact by identifier, and works on one finding. It never falls back to the Security audit's queue, never runs a new pentest, and never contacts the website to reproduce or extend an exploit; only repository-local, non-destructive checks are allowed.
Issue first, pull request only when bounded
The improvement revalidates the current default branch, the affected code, provider history, and duplicate work, then classifies the selection as current, stale, duplicate, already fixed, manual review, or not fixable. In issue-only mode it creates, reuses, or comments on at most one issue. In issue and PR mode it may add one dedicated branch and one pull or merge request, but only for a current, bounded, defensive change with a credible repository-local verification path. It never merges, never deploys, and never pushes to the default branch. A stale, duplicate, or unfixable finding ends as a reported domain outcome, not a forced patch.
Choose which autofixes to run right now.
Require the booking session or account to match the reference before GET /api/bookings/{reference} returns a record, and add a test for the cross-customer request.
Ready to fix.Apply the existing per-IP limiter to POST /api/auth/reset and cover the limit in the auth tests.
Ready to fix.Set the attributes in the shared cookie helper and verify the header in the session tests.
Ready to fix.Every improvement report states the same rule the audit does: an issue or an unmerged pull request leaves the Web pentest score unchanged. The site has to be redeployed and tested again, and only that new run can confirm that the externally observed behavior is gone.
Limits, reruns, and renewed consent
The result is bounded by what the run could reach. Login walls, one-time codes, WAF or geo-blocking, maintenance pages, and the absence of test accounts reduce declared coverage and are listed as limitations; they do not invalidate findings that were confirmed on the public surface. Public and unauthenticated testing is the current scope, so weaknesses behind a sign-in that Enji Guard has no credentials for remain unassessed and are described that way.
- A tool that could not be prepared through a trusted, pinned source is a recorded coverage gap, never a reason to accept an unverified binary.
- Probing that begins to destabilize the target stops, and the affected vectors are recorded as unverified beyond the safe evidence gathered so far.
- A missing previous run makes the current run the baseline; a previous run with a different scenario, route availability, or auth barrier is compared with those differences named as limitations.
Rerun when the deployment changes
Rerun after a fix is deployed, after the site's routes, authentication, or infrastructure change, and on the recurring cadence chosen for the site. Each run is substantially independent: a score can rise when a behavior disappears, stay flat when an issue is open but the site is unchanged, or fall when a new release exposes a new behavior. A recorded consent covers the specific site and scope it was given for; the Authorized Testing terms require a fresh confirmation when the target, ownership, or scope changes, when a repository or website is disconnected and reconnected, or after twelve months.
Keeping the live product in the green zone
Run on a schedule against a verified site, the Web pentest audit notices when the live application leaves the green zone, hands one confirmed finding at a time to a human-reviewed fix, and credits the improvement only after the next deployment has been tested. Coding agents that ship features on top of that repository inherit a site whose externally observable behavior is checked as often as its code.
Enji Guard