# Compliance Control Mapping Control mapping connects the scanner's raw output — deterministic tool findings and the code itself — to the **compliance controls** each piece of evidence supports. A hardcoded credential stops being just "CWE-798 from semgrep" and becomes evidence for *"cra-ai-8: no default passwords"* and, at scale, master control *`mc-31761` hardcoded_secrets_detection*. Findings carry those references (`control_refs`) into the dashboard and out over the MCP server as OSCAL, so the compliance report is built from real, grounded findings rather than a questionnaire. ## The core principle: tools detect, the LLM judges The design has one rule, borrowed from the ZeroFalse / IRIS line of research: **deterministic tools are the detectors; the LLM is only ever a grounded false-positive filter, never the thing that finds the issue.** - A tool (semgrep, gitleaks, syft/osv, ZAP, nuclei) detects deterministically. - An **authored, human-reviewed lookup table** (`control-map`) maps that detection to the control(s) it's evidence for. - The LLM enters last, to *confirm or refute* the mapping against the actual code — and every surviving verdict is anchored to a verbatim snippet by the grounding gate. This keeps hallucination out of detection. The LLM supplies cross-language, cross-stack pattern *recognition*; the surrounding machinery supplies determinism. ## Coverage model Every control lands in one of three buckets, recorded in the `control-map` LUT (`control-map/data/cra_control_map.json`) and never decided by an LLM: | Bucket | Meaning | | --- | --- | | `covered` | An existing tool's scan surfaces findings for this control | | `needs_tooling` | Code-checkable, but no off-the-shelf tool digs it out — we author a detector or use the grounded surface check | | `not_code_checkable` | A design/process property — out of static-scan scope | For the **CRA** framework (40 controls) the split is **13 covered · 8 needs_tooling · 19 not_code_checkable**. The 16 originally-uncovered controls were resolved as a hybrid: - **4 custom semgrep detectors** (`cra-ai-1`, `7`, `10`, `14`) — secure-by-default, weak password hashing, insecure session cookies, weak data-at-rest ciphers. Shipped in the binary and matched back to controls **by rule id** so a broad CWE can't over-attribute. - **8 grounded surface checks** (`cra-ai-6`, `11`, `12`, `24`, `27`, `28`, `29`, `30`) — the absence-based controls (no rate limiting, no security logging, no update-signature check…) that have no syntactic pattern. - **4 marked not_code_checkable** (`cra-ai-2`, `3`, `4`, `5`) — minimal attack surface, secure architecture, least privilege, tamper protection. At scale, the **master-controls** corpus (breakpilot's deduped clusters, exported as OSCAL) currently provides **~2,882 code-checkable controls** (2,143 `network` + 739 `source_code`), matched semantically. ## The three mapping paths ```mermaid flowchart TD T[Deterministic tools\nsemgrep · gitleaks · syft/osv · ZAP] --> F[Findings] F --> B["Stage 5b — LUT triage\ncontrols_for(tool, cwe / rule_id)"] F --> C["Stage 5c — Semantic\nembed region+intent → top-K master controls"] R[Repo source] --> D["Stage 5d — Grounded surface\nretrieve surface for absence-based controls"] B --> J{{Grounded LLM judge\ntemp 0 · verbatim snippet}} C --> J D --> J J -->|snippet grounds in region| S[Stamp control_refs] J -->|refuted / ungrounded| X[Dropped] ``` All three paths converge on the same **grounded judge** and the same **grounding gate**. They differ only in how candidate (finding/region, control) pairs are produced. ### Stage 5b — deterministic LUT triage The default path. A tool finding is matched to controls via `control_map.controls_for_finding(tool, cwe, rule_id)`; the judge then confirms each mapped control against the code region. Outcomes: `Confirmed([ids])` (stamp them), `FalsePositive` (drop the finding), or `Unmapped` (keep it untagged). Runs whenever `BREAKPILOT_BASE_URL` is set. ### Stage 5c — semantic retrieval (master-controls scale) Master controls carry no CWE, so they can't be LUT-mapped. Instead we map by *similarity*: embed every control's requirement text once (cached), then for each finding retrieve the top-K nearest controls and hand them to the judge. Gated behind `BREAKPILOT_SEMANTIC_MAPPING` (default off). See [Semantic retrieval](#semantic-retrieval-in-detail). ### Stage 5d — grounded surface checks (absence-based controls) Some controls are violated by an *absence* — no rate limiting on login, no security logging, no signature check on an update. There's no pattern for semgrep to match, so we deterministically retrieve the code **surface** the control governs (a login route, a logging setup, update/download code) by identifier/route terms, and let the judge decide whether the control holds there. Produces net-new, already-grounded findings. Gated behind `BREAKPILOT_GROUNDED_CHECKS` (default off). ## The grounding gate No matter the path, a verdict becomes a finding only if it survives `compliance_core::control_check::ground`: 1. The judge runs at **temperature 0** with a closed prompt and must quote the offending code **verbatim** into `snippet`. 2. That snippet must appear **literally** in the retrieved region — otherwise the verdict is dropped. 3. The finding's line is **recomputed from the match**; the model's own line number is never trusted. 4. Verdicts are cached by content hash, so re-scans reproduce. The model is allowed to be smart; it is never trusted. ## Semantic retrieval in detail 1. **Embed the corpus once.** Each control's requirement text is embedded with `bge-multilingual-gemma2` (3584-dim — multilingual matters, the master controls are in German while code is English). The embedding backend caps input arrays at 25 per request, so `embed()` chunks at 16; the whole `ControlIndex` is persisted to `snapshot_dir` keyed by a **corpus hash**, so only the first scan after a catalog change pays the embedding cost. 2. **Build the query from the finding's intent, not just the code.** The retrieval query is `finding.title + finding.description + region`, not the raw region. This is the single most important tuning: two findings in one file share overlapping windows and, on the code alone, embed alike and collapse onto the same controls. The finding's own words ("brute-force protection" vs "weak hash") carry the discriminating signal. The raw region still goes to the judge for grounding. 3. **Retrieve → judge → ground.** Top-K nearest by cosine, each judged against the region, each grounded. ## Worked examples Both examples are from the live end-to-end verification (`c5_semantic_live.rs`) against the real ~2,882-control corpus. ### Example 1 — a small auth file (the tuning story) Two findings in one `auth.py`: a weak `hashlib.md5(password)` hash and a login endpoint with no brute-force protection. | Finding | Region-only retrieval | Intent-enriched retrieval | | --- | --- | --- | | Weak md5 hash | 19874, 20683, 23149, 29985 | **`mc-23149`** (eliminate weak unsalted hashes) at rank 1, + `mc-21634` salted hashing | | Login w/o brute-force protection | *identical 4, reordered* | newly surfaces **`mc-19984`** brute_force_protection + **`mc-23186`** account_lockout | Region-only retrieval gave both findings the *same* four password-hashing controls — the brute-force finding never found its real controls because its window is saturated with `password` tokens. Enriching the query with the finding's intent fixed it: the brute-force finding now pulls the correct rate-limiting / lockout controls out of the 2,882. ### Example 2 — four topically distinct vulnerabilities | Finding | Top matched controls | Family | | --- | --- | --- | | SQL injection (string-concat query) | `sql_injection_prevention`, `sql_injection`, `parameterized_queries`, input_sanitization | input-validation ✓ | | Hardcoded API credential | `hardcoded_secrets_detection`, credential_scanning, secrets_detection | credentials ✓ | | TLS verification disabled (`verify=False`) | `https_enforcement`, `configuration_verification`, transport config | transport-encryption ✓ | | Insecure deserialization (`pickle.loads`) | `deserialization`, `deserialization_testing`, `deserialization_security` | deserialization ✓ | Every finding maps to its exact control family, with the most specific control often at the top, and the four sets are distinct. ## Known limitations - **Absence findings are weak for semantic retrieval.** Similarity matches what code *is about*, not what it *lacks*; a "missing rate limiting" finding embeds like login code. This is exactly why the grounded surface path (Stage 5d) exists — it decides presence/absence at a retrieved surface rather than by embedding distance. - **Generic catch-all controls co-occur.** `mc-20890 secure_development_security_code_review` appears in the top-K for many code-security findings because it is semantically near almost all of them. It's harmless (the judge grounds it, and it never crowds out the specific controls — the SQLi example didn't get it) but is a candidate for future down-weighting. - **Corpus classification noise.** The master-controls `verification_method` classification is imperfect — e.g. a documentation control (`eu_declaration_accuracy`) is currently tagged `source_code`. That's a corpus-side data-quality issue, separate from the mapping engine. ## Emitting over MCP — closing the loop Findings don't just land in the dashboard; they flow to breakpilot-compliance as OSCAL over the scanner's MCP server, so the compliance report is assembled from real, control-tagged findings. - The MCP server exposes an **`oscal_assessment`** tool: given a `repo_id`, it emits a standard OSCAL 1.1 assessment-results document for that repo's findings — mapped findings target their controls via the stamped `control_refs`, and unmapped findings are reported **as-is** (as observations), so nothing is lost. - breakpilot pulls it: `POST /v1/cra/oscal-from-scanner` calls `oscal_assessment` over MCP (Streamable HTTP + bearer) and consumes the pre-computed OSCAL — rather than pulling raw findings and re-assessing. **Operational note — tenant context over HTTP.** The MCP server is multi-tenant; the bearer token resolves a tenant whose per-tenant database the tools query. rmcp's Streamable HTTP transport runs each session's tool calls in a `tokio::spawn`ed task, and `task_local`s do **not** cross a spawn — so binding the tenant in a per-request middleware `task_local` leaves tool handlers with no context (every call fails `no tenant context`). The fix is to bind the tenant to the **per-session server instance** at creation (the factory runs in the request scope before the spawn), not to a per-request task_local. Until this was fixed, the loop silently failed over HTTP and consumers fell back to demo data. ## Configuration | Variable | Effect | | --- | --- | | `BREAKPILOT_BASE_URL` | breakpilot-compliance root; enables control ingest + all mapping passes. **Unset disables all control mapping** — findings are produced without `control_refs`. | | `BREAKPILOT_SEMANTIC_MAPPING` | Stage 5c (semantic master-controls mapping). **Default on** (validated live). | | `BREAKPILOT_GROUNDED_CHECKS` | Stage 5d (grounded surface checks). **Default on** (validated live). | | `BREAKPILOT_SNAPSHOT_DIR` | Where OSCAL catalog snapshots and the cached control-embedding index live. | The semantic and grounded passes default **on** now that both are validated live; each is still a no-op if `BREAKPILOT_BASE_URL` is unset or the catalog is unreachable, so they only ever add coverage. The live verifications live in `compliance-agent/tests/c5_semantic_live.rs` and `grounded_surface_live.rs` (ignored; run with `--ignored`). ## Appendix — the master-controls data pipeline The master-controls corpus is produced by breakpilot-compliance and pulled as an OSCAL catalog from `GET /api/compliance/v1/oscal/catalog?framework=master-controls`. Two operational lessons are worth recording, because they cost real time to diagnose: - **The catalog is served from `breakpilot_db`, not `postgres`.** Diagnostics run against the wrong database will look clean while the app serves something else entirely. Confirm the app's datname (`pg_stat_activity`) before trusting any count or `EXPLAIN`. - **A constraint-less dump triplicated the master-control tables.** Restored without their PK/unique constraints, `master_controls` / `mc_verification` / `master_control_members` accumulated identical rows 3× (the same artifact migration `158` fixed for `doc_check_controls`). That inflated the catalog to ~26k dup'd controls and, with the indexes also missing, drove the export query to a >120s / 502. The fix (breakpilot migration `160`) ctid-dedups each table by its natural key and restores the constraints + indexes so it can't recur; the export query was also rewritten set-based (a single windowed pass instead of a per-row correlated subquery). After dedup: 41,850 → 13,950 master controls, catalog **25,938 → 2,882** code-checkable, endpoint **502 → 200 in ~3s**.