Files
sharang ea516cc054
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Agent (push) Skipped
CI / Deploy Dashboard (push) Skipped
CI / Deploy MCP (push) Skipped
CI / Deploy Docs (push) Successful in 57s
docs(control-mapping): MCP emission loop + default-on flags (#227)
2026-07-22 11:23:51 +00:00

13 KiB
Raw Permalink Blame History

Compliance Control Mapping

Control mapping connects the scanner's raw output — deterministic tool findings and the code itself — to the compliance controls each piece of evidence supports. A hardcoded credential stops being just "CWE-798 from semgrep" and becomes evidence for "cra-ai-8: no default passwords" and, at scale, master control mc-31761 hardcoded_secrets_detection. Findings carry those references (control_refs) into the dashboard and out over the MCP server as OSCAL, so the compliance report is built from real, grounded findings rather than a questionnaire.

The core principle: tools detect, the LLM judges

The design has one rule, borrowed from the ZeroFalse / IRIS line of research: deterministic tools are the detectors; the LLM is only ever a grounded false-positive filter, never the thing that finds the issue.

  • A tool (semgrep, gitleaks, syft/osv, ZAP, nuclei) detects deterministically.
  • An authored, human-reviewed lookup table (control-map) maps that detection to the control(s) it's evidence for.
  • The LLM enters last, to confirm or refute the mapping against the actual code — and every surviving verdict is anchored to a verbatim snippet by the grounding gate.

This keeps hallucination out of detection. The LLM supplies cross-language, cross-stack pattern recognition; the surrounding machinery supplies determinism.

Coverage model

Every control lands in one of three buckets, recorded in the control-map LUT (control-map/data/cra_control_map.json) and never decided by an LLM:

Bucket Meaning
covered An existing tool's scan surfaces findings for this control
needs_tooling Code-checkable, but no off-the-shelf tool digs it out — we author a detector or use the grounded surface check
not_code_checkable A design/process property — out of static-scan scope

For the CRA framework (40 controls) the split is 13 covered · 8 needs_tooling · 19 not_code_checkable. The 16 originally-uncovered controls were resolved as a hybrid:

  • 4 custom semgrep detectors (cra-ai-1, 7, 10, 14) — secure-by-default, weak password hashing, insecure session cookies, weak data-at-rest ciphers. Shipped in the binary and matched back to controls by rule id so a broad CWE can't over-attribute.
  • 8 grounded surface checks (cra-ai-6, 11, 12, 24, 27, 28, 29, 30) — the absence-based controls (no rate limiting, no security logging, no update-signature check…) that have no syntactic pattern.
  • 4 marked not_code_checkable (cra-ai-2, 3, 4, 5) — minimal attack surface, secure architecture, least privilege, tamper protection.

At scale, the master-controls corpus (breakpilot's deduped clusters, exported as OSCAL) currently provides ~2,882 code-checkable controls (2,143 network + 739 source_code), matched semantically.

The three mapping paths

flowchart TD
    T[Deterministic tools\nsemgrep · gitleaks · syft/osv · ZAP] --> F[Findings]
    F --> B["Stage 5b — LUT triage\ncontrols_for(tool, cwe / rule_id)"]
    F --> C["Stage 5c — Semantic\nembed region+intent → top-K master controls"]
    R[Repo source] --> D["Stage 5d — Grounded surface\nretrieve surface for absence-based controls"]
    B --> J{{Grounded LLM judge\ntemp 0 · verbatim snippet}}
    C --> J
    D --> J
    J -->|snippet grounds in region| S[Stamp control_refs]
    J -->|refuted / ungrounded| X[Dropped]

All three paths converge on the same grounded judge and the same grounding gate. They differ only in how candidate (finding/region, control) pairs are produced.

Stage 5b — deterministic LUT triage

The default path. A tool finding is matched to controls via control_map.controls_for_finding(tool, cwe, rule_id); the judge then confirms each mapped control against the code region. Outcomes: Confirmed([ids]) (stamp them), FalsePositive (drop the finding), or Unmapped (keep it untagged). Runs whenever BREAKPILOT_BASE_URL is set.

Stage 5c — semantic retrieval (master-controls scale)

Master controls carry no CWE, so they can't be LUT-mapped. Instead we map by similarity: embed every control's requirement text once (cached), then for each finding retrieve the top-K nearest controls and hand them to the judge. Gated behind BREAKPILOT_SEMANTIC_MAPPING (default off). See Semantic retrieval.

Stage 5d — grounded surface checks (absence-based controls)

Some controls are violated by an absence — no rate limiting on login, no security logging, no signature check on an update. There's no pattern for semgrep to match, so we deterministically retrieve the code surface the control governs (a login route, a logging setup, update/download code) by identifier/route terms, and let the judge decide whether the control holds there. Produces net-new, already-grounded findings. Gated behind BREAKPILOT_GROUNDED_CHECKS (default off).

The grounding gate

No matter the path, a verdict becomes a finding only if it survives compliance_core::control_check::ground:

  1. The judge runs at temperature 0 with a closed prompt and must quote the offending code verbatim into snippet.
  2. That snippet must appear literally in the retrieved region — otherwise the verdict is dropped.
  3. The finding's line is recomputed from the match; the model's own line number is never trusted.
  4. Verdicts are cached by content hash, so re-scans reproduce.

The model is allowed to be smart; it is never trusted.

Semantic retrieval in detail

  1. Embed the corpus once. Each control's requirement text is embedded with bge-multilingual-gemma2 (3584-dim — multilingual matters, the master controls are in German while code is English). The embedding backend caps input arrays at 25 per request, so embed() chunks at 16; the whole ControlIndex is persisted to snapshot_dir keyed by a corpus hash, so only the first scan after a catalog change pays the embedding cost.
  2. Build the query from the finding's intent, not just the code. The retrieval query is finding.title + finding.description + region, not the raw region. This is the single most important tuning: two findings in one file share overlapping windows and, on the code alone, embed alike and collapse onto the same controls. The finding's own words ("brute-force protection" vs "weak hash") carry the discriminating signal. The raw region still goes to the judge for grounding.
  3. Retrieve → judge → ground. Top-K nearest by cosine, each judged against the region, each grounded.

Worked examples

Both examples are from the live end-to-end verification (c5_semantic_live.rs) against the real ~2,882-control corpus.

Example 1 — a small auth file (the tuning story)

Two findings in one auth.py: a weak hashlib.md5(password) hash and a login endpoint with no brute-force protection.

Finding Region-only retrieval Intent-enriched retrieval
Weak md5 hash 19874, 20683, 23149, 29985 mc-23149 (eliminate weak unsalted hashes) at rank 1, + mc-21634 salted hashing
Login w/o brute-force protection identical 4, reordered newly surfaces mc-19984 brute_force_protection + mc-23186 account_lockout

Region-only retrieval gave both findings the same four password-hashing controls — the brute-force finding never found its real controls because its window is saturated with password tokens. Enriching the query with the finding's intent fixed it: the brute-force finding now pulls the correct rate-limiting / lockout controls out of the 2,882.

Example 2 — four topically distinct vulnerabilities

Finding Top matched controls Family
SQL injection (string-concat query) sql_injection_prevention, sql_injection, parameterized_queries, input_sanitization input-validation ✓
Hardcoded API credential hardcoded_secrets_detection, credential_scanning, secrets_detection credentials ✓
TLS verification disabled (verify=False) https_enforcement, configuration_verification, transport config transport-encryption ✓
Insecure deserialization (pickle.loads) deserialization, deserialization_testing, deserialization_security deserialization ✓

Every finding maps to its exact control family, with the most specific control often at the top, and the four sets are distinct.

Known limitations

  • Absence findings are weak for semantic retrieval. Similarity matches what code is about, not what it lacks; a "missing rate limiting" finding embeds like login code. This is exactly why the grounded surface path (Stage 5d) exists — it decides presence/absence at a retrieved surface rather than by embedding distance.
  • Generic catch-all controls co-occur. mc-20890 secure_development_security_code_review appears in the top-K for many code-security findings because it is semantically near almost all of them. It's harmless (the judge grounds it, and it never crowds out the specific controls — the SQLi example didn't get it) but is a candidate for future down-weighting.
  • Corpus classification noise. The master-controls verification_method classification is imperfect — e.g. a documentation control (eu_declaration_accuracy) is currently tagged source_code. That's a corpus-side data-quality issue, separate from the mapping engine.

Emitting over MCP — closing the loop

Findings don't just land in the dashboard; they flow to breakpilot-compliance as OSCAL over the scanner's MCP server, so the compliance report is assembled from real, control-tagged findings.

  • The MCP server exposes an oscal_assessment tool: given a repo_id, it emits a standard OSCAL 1.1 assessment-results document for that repo's findings — mapped findings target their controls via the stamped control_refs, and unmapped findings are reported as-is (as observations), so nothing is lost.
  • breakpilot pulls it: POST /v1/cra/oscal-from-scanner calls oscal_assessment over MCP (Streamable HTTP + bearer) and consumes the pre-computed OSCAL — rather than pulling raw findings and re-assessing.

Operational note — tenant context over HTTP. The MCP server is multi-tenant; the bearer token resolves a tenant whose per-tenant database the tools query. rmcp's Streamable HTTP transport runs each session's tool calls in a tokio::spawned task, and task_locals do not cross a spawn — so binding the tenant in a per-request middleware task_local leaves tool handlers with no context (every call fails no tenant context). The fix is to bind the tenant to the per-session server instance at creation (the factory runs in the request scope before the spawn), not to a per-request task_local. Until this was fixed, the loop silently failed over HTTP and consumers fell back to demo data.

Configuration

Variable Effect
BREAKPILOT_BASE_URL breakpilot-compliance root; enables control ingest + all mapping passes. Unset disables all control mapping — findings are produced without control_refs.
BREAKPILOT_SEMANTIC_MAPPING Stage 5c (semantic master-controls mapping). Default on (validated live).
BREAKPILOT_GROUNDED_CHECKS Stage 5d (grounded surface checks). Default on (validated live).
BREAKPILOT_SNAPSHOT_DIR Where OSCAL catalog snapshots and the cached control-embedding index live.

The semantic and grounded passes default on now that both are validated live; each is still a no-op if BREAKPILOT_BASE_URL is unset or the catalog is unreachable, so they only ever add coverage. The live verifications live in compliance-agent/tests/c5_semantic_live.rs and grounded_surface_live.rs (ignored; run with --ignored).

Appendix — the master-controls data pipeline

The master-controls corpus is produced by breakpilot-compliance and pulled as an OSCAL catalog from GET /api/compliance/v1/oscal/catalog?framework=master-controls. Two operational lessons are worth recording, because they cost real time to diagnose:

  • The catalog is served from breakpilot_db, not postgres. Diagnostics run against the wrong database will look clean while the app serves something else entirely. Confirm the app's datname (pg_stat_activity) before trusting any count or EXPLAIN.
  • A constraint-less dump triplicated the master-control tables. Restored without their PK/unique constraints, master_controls / mc_verification / master_control_members accumulated identical rows 3× (the same artifact migration 158 fixed for doc_check_controls). That inflated the catalog to ~26k dup'd controls and, with the indexes also missing, drove the export query to a >120s / 502. The fix (breakpilot migration 160) ctid-dedups each table by its natural key and restores the constraints + indexes so it can't recur; the export query was also rewritten set-based (a single windowed pass instead of a per-row correlated subquery). After dedup: 41,850 → 13,950 master controls, catalog 25,938 → 2,882 code-checkable, endpoint 502 → 200 in ~3s.