Files
compliance-scanner-agent/docs/features/control-mapping.md
T
Sharang ParnerkarandClaude Opus 4.8 9c8f864a1d
CI / Check (push) Skipped
CI / Check (pull_request) Successful in 5m47s
CI / Detect Changes (pull_request) Skipped
CI / Deploy Agent (pull_request) Skipped
CI / Deploy Dashboard (pull_request) Skipped
CI / Deploy Docs (pull_request) Skipped
CI / Deploy MCP (pull_request) Skipped
docs(control-mapping): MCP emission loop + flags now default-on
- Add 'Emitting over MCP' section: the oscal_assessment tool, breakpilot's
  /v1/cra/oscal-from-scanner pull, and the tenant-context operational note
  (task_local lost across rmcp's session spawn -> bind tenant to the session).
- Configuration: semantic + grounded passes are now default-on (validated live),
  still no-op without BREAKPILOT_BASE_URL.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-22 13:14:59 +02:00

137 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Compliance Control Mapping
Control mapping connects the scanner's raw output — deterministic tool findings and the code itself — to the **compliance controls** each piece of evidence supports. A hardcoded credential stops being just "CWE-798 from semgrep" and becomes evidence for *"cra-ai-8: no default passwords"* and, at scale, master control *`mc-31761` hardcoded_secrets_detection*. Findings carry those references (`control_refs`) into the dashboard and out over the MCP server as OSCAL, so the compliance report is built from real, grounded findings rather than a questionnaire.
## The core principle: tools detect, the LLM judges
The design has one rule, borrowed from the ZeroFalse / IRIS line of research: **deterministic tools are the detectors; the LLM is only ever a grounded false-positive filter, never the thing that finds the issue.**
- A tool (semgrep, gitleaks, syft/osv, ZAP, nuclei) detects deterministically.
- An **authored, human-reviewed lookup table** (`control-map`) maps that detection to the control(s) it's evidence for.
- The LLM enters last, to *confirm or refute* the mapping against the actual code — and every surviving verdict is anchored to a verbatim snippet by the grounding gate.
This keeps hallucination out of detection. The LLM supplies cross-language, cross-stack pattern *recognition*; the surrounding machinery supplies determinism.
## Coverage model
Every control lands in one of three buckets, recorded in the `control-map` LUT (`control-map/data/cra_control_map.json`) and never decided by an LLM:
| Bucket | Meaning |
| --- | --- |
| `covered` | An existing tool's scan surfaces findings for this control |
| `needs_tooling` | Code-checkable, but no off-the-shelf tool digs it out — we author a detector or use the grounded surface check |
| `not_code_checkable` | A design/process property — out of static-scan scope |
For the **CRA** framework (40 controls) the split is **13 covered · 8 needs_tooling · 19 not_code_checkable**. The 16 originally-uncovered controls were resolved as a hybrid:
- **4 custom semgrep detectors** (`cra-ai-1`, `7`, `10`, `14`) — secure-by-default, weak password hashing, insecure session cookies, weak data-at-rest ciphers. Shipped in the binary and matched back to controls **by rule id** so a broad CWE can't over-attribute.
- **8 grounded surface checks** (`cra-ai-6`, `11`, `12`, `24`, `27`, `28`, `29`, `30`) — the absence-based controls (no rate limiting, no security logging, no update-signature check…) that have no syntactic pattern.
- **4 marked not_code_checkable** (`cra-ai-2`, `3`, `4`, `5`) — minimal attack surface, secure architecture, least privilege, tamper protection.
At scale, the **master-controls** corpus (breakpilot's deduped clusters, exported as OSCAL) currently provides **~2,882 code-checkable controls** (2,143 `network` + 739 `source_code`), matched semantically.
## The three mapping paths
```mermaid
flowchart TD
T[Deterministic tools\nsemgrep · gitleaks · syft/osv · ZAP] --> F[Findings]
F --> B["Stage 5b — LUT triage\ncontrols_for(tool, cwe / rule_id)"]
F --> C["Stage 5c — Semantic\nembed region+intent → top-K master controls"]
R[Repo source] --> D["Stage 5d — Grounded surface\nretrieve surface for absence-based controls"]
B --> J{{Grounded LLM judge\ntemp 0 · verbatim snippet}}
C --> J
D --> J
J -->|snippet grounds in region| S[Stamp control_refs]
J -->|refuted / ungrounded| X[Dropped]
```
All three paths converge on the same **grounded judge** and the same **grounding gate**. They differ only in how candidate (finding/region, control) pairs are produced.
### Stage 5b — deterministic LUT triage
The default path. A tool finding is matched to controls via `control_map.controls_for_finding(tool, cwe, rule_id)`; the judge then confirms each mapped control against the code region. Outcomes: `Confirmed([ids])` (stamp them), `FalsePositive` (drop the finding), or `Unmapped` (keep it untagged). Runs whenever `BREAKPILOT_BASE_URL` is set.
### Stage 5c — semantic retrieval (master-controls scale)
Master controls carry no CWE, so they can't be LUT-mapped. Instead we map by *similarity*: embed every control's requirement text once (cached), then for each finding retrieve the top-K nearest controls and hand them to the judge. Gated behind `BREAKPILOT_SEMANTIC_MAPPING` (default off). See [Semantic retrieval](#semantic-retrieval-in-detail).
### Stage 5d — grounded surface checks (absence-based controls)
Some controls are violated by an *absence* — no rate limiting on login, no security logging, no signature check on an update. There's no pattern for semgrep to match, so we deterministically retrieve the code **surface** the control governs (a login route, a logging setup, update/download code) by identifier/route terms, and let the judge decide whether the control holds there. Produces net-new, already-grounded findings. Gated behind `BREAKPILOT_GROUNDED_CHECKS` (default off).
## The grounding gate
No matter the path, a verdict becomes a finding only if it survives `compliance_core::control_check::ground`:
1. The judge runs at **temperature 0** with a closed prompt and must quote the offending code **verbatim** into `snippet`.
2. That snippet must appear **literally** in the retrieved region — otherwise the verdict is dropped.
3. The finding's line is **recomputed from the match**; the model's own line number is never trusted.
4. Verdicts are cached by content hash, so re-scans reproduce.
The model is allowed to be smart; it is never trusted.
## Semantic retrieval in detail
1. **Embed the corpus once.** Each control's requirement text is embedded with `bge-multilingual-gemma2` (3584-dim — multilingual matters, the master controls are in German while code is English). The embedding backend caps input arrays at 25 per request, so `embed()` chunks at 16; the whole `ControlIndex` is persisted to `snapshot_dir` keyed by a **corpus hash**, so only the first scan after a catalog change pays the embedding cost.
2. **Build the query from the finding's intent, not just the code.** The retrieval query is `finding.title + finding.description + region`, not the raw region. This is the single most important tuning: two findings in one file share overlapping windows and, on the code alone, embed alike and collapse onto the same controls. The finding's own words ("brute-force protection" vs "weak hash") carry the discriminating signal. The raw region still goes to the judge for grounding.
3. **Retrieve → judge → ground.** Top-K nearest by cosine, each judged against the region, each grounded.
## Worked examples
Both examples are from the live end-to-end verification (`c5_semantic_live.rs`) against the real ~2,882-control corpus.
### Example 1 — a small auth file (the tuning story)
Two findings in one `auth.py`: a weak `hashlib.md5(password)` hash and a login endpoint with no brute-force protection.
| Finding | Region-only retrieval | Intent-enriched retrieval |
| --- | --- | --- |
| Weak md5 hash | 19874, 20683, 23149, 29985 | **`mc-23149`** (eliminate weak unsalted hashes) at rank 1, + `mc-21634` salted hashing |
| Login w/o brute-force protection | *identical 4, reordered* | newly surfaces **`mc-19984`** brute_force_protection + **`mc-23186`** account_lockout |
Region-only retrieval gave both findings the *same* four password-hashing controls — the brute-force finding never found its real controls because its window is saturated with `password` tokens. Enriching the query with the finding's intent fixed it: the brute-force finding now pulls the correct rate-limiting / lockout controls out of the 2,882.
### Example 2 — four topically distinct vulnerabilities
| Finding | Top matched controls | Family |
| --- | --- | --- |
| SQL injection (string-concat query) | `sql_injection_prevention`, `sql_injection`, `parameterized_queries`, input_sanitization | input-validation ✓ |
| Hardcoded API credential | `hardcoded_secrets_detection`, credential_scanning, secrets_detection | credentials ✓ |
| TLS verification disabled (`verify=False`) | `https_enforcement`, `configuration_verification`, transport config | transport-encryption ✓ |
| Insecure deserialization (`pickle.loads`) | `deserialization`, `deserialization_testing`, `deserialization_security` | deserialization ✓ |
Every finding maps to its exact control family, with the most specific control often at the top, and the four sets are distinct.
## Known limitations
- **Absence findings are weak for semantic retrieval.** Similarity matches what code *is about*, not what it *lacks*; a "missing rate limiting" finding embeds like login code. This is exactly why the grounded surface path (Stage 5d) exists — it decides presence/absence at a retrieved surface rather than by embedding distance.
- **Generic catch-all controls co-occur.** `mc-20890 secure_development_security_code_review` appears in the top-K for many code-security findings because it is semantically near almost all of them. It's harmless (the judge grounds it, and it never crowds out the specific controls — the SQLi example didn't get it) but is a candidate for future down-weighting.
- **Corpus classification noise.** The master-controls `verification_method` classification is imperfect — e.g. a documentation control (`eu_declaration_accuracy`) is currently tagged `source_code`. That's a corpus-side data-quality issue, separate from the mapping engine.
## Emitting over MCP — closing the loop
Findings don't just land in the dashboard; they flow to breakpilot-compliance as OSCAL over the scanner's MCP server, so the compliance report is assembled from real, control-tagged findings.
- The MCP server exposes an **`oscal_assessment`** tool: given a `repo_id`, it emits a standard OSCAL 1.1 assessment-results document for that repo's findings — mapped findings target their controls via the stamped `control_refs`, and unmapped findings are reported **as-is** (as observations), so nothing is lost.
- breakpilot pulls it: `POST /v1/cra/oscal-from-scanner` calls `oscal_assessment` over MCP (Streamable HTTP + bearer) and consumes the pre-computed OSCAL — rather than pulling raw findings and re-assessing.
**Operational note — tenant context over HTTP.** The MCP server is multi-tenant; the bearer token resolves a tenant whose per-tenant database the tools query. rmcp's Streamable HTTP transport runs each session's tool calls in a `tokio::spawn`ed task, and `task_local`s do **not** cross a spawn — so binding the tenant in a per-request middleware `task_local` leaves tool handlers with no context (every call fails `no tenant context`). The fix is to bind the tenant to the **per-session server instance** at creation (the factory runs in the request scope before the spawn), not to a per-request task_local. Until this was fixed, the loop silently failed over HTTP and consumers fell back to demo data.
## Configuration
| Variable | Effect |
| --- | --- |
| `BREAKPILOT_BASE_URL` | breakpilot-compliance root; enables control ingest + all mapping passes. **Unset disables all control mapping** — findings are produced without `control_refs`. |
| `BREAKPILOT_SEMANTIC_MAPPING` | Stage 5c (semantic master-controls mapping). **Default on** (validated live). |
| `BREAKPILOT_GROUNDED_CHECKS` | Stage 5d (grounded surface checks). **Default on** (validated live). |
| `BREAKPILOT_SNAPSHOT_DIR` | Where OSCAL catalog snapshots and the cached control-embedding index live. |
The semantic and grounded passes default **on** now that both are validated live; each is still a no-op if `BREAKPILOT_BASE_URL` is unset or the catalog is unreachable, so they only ever add coverage. The live verifications live in `compliance-agent/tests/c5_semantic_live.rs` and `grounded_surface_live.rs` (ignored; run with `--ignored`).
## Appendix — the master-controls data pipeline
The master-controls corpus is produced by breakpilot-compliance and pulled as an OSCAL catalog from `GET /api/compliance/v1/oscal/catalog?framework=master-controls`. Two operational lessons are worth recording, because they cost real time to diagnose:
- **The catalog is served from `breakpilot_db`, not `postgres`.** Diagnostics run against the wrong database will look clean while the app serves something else entirely. Confirm the app's datname (`pg_stat_activity`) before trusting any count or `EXPLAIN`.
- **A constraint-less dump triplicated the master-control tables.** Restored without their PK/unique constraints, `master_controls` / `mc_verification` / `master_control_members` accumulated identical rows 3× (the same artifact migration `158` fixed for `doc_check_controls`). That inflated the catalog to ~26k dup'd controls and, with the indexes also missing, drove the export query to a >120s / 502. The fix (breakpilot migration `160`) ctid-dedups each table by its natural key and restores the constraints + indexes so it can't recur; the export query was also rewritten set-based (a single windowed pass instead of a per-row correlated subquery). After dedup: 41,850 → 13,950 master controls, catalog **25,938 → 2,882** code-checkable, endpoint **502 → 200 in ~3s**.