Compare commits

...
Author SHA1 Message Date
Sharang ParnerkarandClaude Fable 5 d03f027fa1 docs(features): control mapping pipeline (grounded in the C5 live results)
CI / Check (push) Skipped
CI / Check (pull_request) Successful in 5m53s
CI / Detect Changes (pull_request) Skipped
CI / Deploy Agent (pull_request) Skipped
CI / Deploy Dashboard (pull_request) Skipped
CI / Deploy Docs (pull_request) Skipped
CI / Deploy MCP (pull_request) Skipped
Extensive feature doc for the compliance control-mapping engine:
- core principle (tools detect, LLM judges/grounds — never detects)
- coverage model + the CRA hybrid (4 semgrep / 8 grounded / 4 not-code-checkable)
- the three mapping paths (LUT 5b / semantic 5c / grounded-surface 5d) with a
  mermaid flow, and the shared grounding gate
- semantic retrieval detail incl. the query-enrichment tuning
- two worked examples from the live C5 run (auth file + varied vulns)
- known limitations (absence findings, catch-all controls, corpus noise)
- config flags + an appendix on the master-controls data pipeline war-story
  (dump-triplication -> migration 160 dedup; 502 -> 200)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 09:23:04 +02:00
sharang a25d41c3e5 feat(controls): enrich semantic retrieval query with finding intent + C5 live test (#222)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Deploy MCP (push) Skipped
CI / Deploy Agent (push) Failing after 4s
2026-07-22 07:14:49 +00:00
sharang 60601d8215 fix(llm): chunk embed() requests under the backend batch cap (#221)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Deploy MCP (push) Skipped
CI / Deploy Agent (push) Failing after 5s
2026-07-21 15:41:07 +00:00
6 changed files with 341 additions and 9 deletions
+12 -3
View File
@@ -234,13 +234,22 @@ pub async fn semantic_stamp_findings(
let Some(region) = fetch_region(repo_path, &file, line) else { let Some(region) = fetch_region(repo_path, &file, line) else {
continue; continue;
}; };
let region_emb = match llm.embed(vec![region.content.clone()]).await { // Retrieve on the finding's intent + the code, not the region alone: two
// findings in one file share overlapping windows and otherwise embed alike,
// collapsing onto the same controls. The finding's title/description carry
// the discriminating signal (e.g. "brute-force protection" vs "weak hash").
// The raw `region` still goes to the judge for snippet grounding.
let query = format!(
"{}\n{}\n\n{}",
finding.title, finding.description, region.content
);
let query_emb = match llm.embed(vec![query]).await {
Ok(mut embs) => match embs.pop() { Ok(mut embs) => match embs.pop() {
Some(v) => v, Some(v) => v,
None => continue, None => continue,
}, },
Err(e) => { Err(e) => {
tracing::warn!(error = %e, "region embed failed; skipping finding"); tracing::warn!(error = %e, "query embed failed; skipping finding");
continue; continue;
} }
}; };
@@ -248,7 +257,7 @@ pub async fn semantic_stamp_findings(
.check( .check(
&index, &index,
&region, &region,
&region_emb, &query_emb,
SEMANTIC_TOP_K, SEMANTIC_TOP_K,
&finding.repo_id, &finding.repo_id,
) )
+7 -5
View File
@@ -22,18 +22,20 @@ impl<J: ControlJudge> SemanticControlChecker<J> {
Self { judge } Self { judge }
} }
/// Map a code region to the controls it violates. `region_embedding` is the /// Map a code region to the controls it violates. `query_embedding` is the
/// region's embedding (the caller computes it via the LLM); the top-`k` /// caller-supplied retrieval embedding — typically the finding's intent
/// nearest controls in `index` are judged and grounded. /// (title/description) plus the region, so retrieval keys on what the finding
/// is *about*, not just the ambient code. The top-`k` nearest controls in
/// `index` are then judged against the raw `region` and grounded.
pub async fn check( pub async fn check(
&self, &self,
index: &ControlIndex, index: &ControlIndex,
region: &CandidateRegion, region: &CandidateRegion,
region_embedding: &[f64], query_embedding: &[f64],
k: usize, k: usize,
repo_id: &str, repo_id: &str,
) -> Vec<Finding> { ) -> Vec<Finding> {
let candidates = index.nearest(region_embedding, k); let candidates = index.nearest(query_embedding, k);
let mut findings = Vec::new(); let mut findings = Vec::new();
for spec in &candidates { for spec in &candidates {
let verdict = self.judge.judge(spec, region).await; let verdict = self.judge.judge(spec, region).await;
+49 -1
View File
@@ -22,6 +22,11 @@ struct EmbeddingData {
index: usize, index: usize,
} }
/// Max inputs per embedding request. The bge/OpenAI-like backends cap the input
/// array (bge-multilingual-gemma2 rejects >25 with "batch size overflow"), so we
/// chunk larger corpora — a whole control catalog (~1.8k) would otherwise 500.
const EMBED_BATCH_SIZE: usize = 16;
// ── Embedding implementation ─────────────────────────────────── // ── Embedding implementation ───────────────────────────────────
impl LlmClient { impl LlmClient {
@@ -29,8 +34,21 @@ impl LlmClient {
&self.embed_model &self.embed_model
} }
/// Generate embeddings for a batch of texts /// Generate embeddings for a batch of texts, chunking into backend-sized
/// requests and preserving input order across chunks.
pub async fn embed(&self, texts: Vec<String>) -> Result<Vec<Vec<f64>>, AgentError> { pub async fn embed(&self, texts: Vec<String>) -> Result<Vec<Vec<f64>>, AgentError> {
if texts.is_empty() {
return Ok(Vec::new());
}
let mut out = Vec::with_capacity(texts.len());
for chunk in texts.chunks(EMBED_BATCH_SIZE) {
out.extend(self.embed_batch(chunk.to_vec()).await?);
}
Ok(out)
}
/// Embed one backend-sized batch (≤ [`EMBED_BATCH_SIZE`]) in a single request.
async fn embed_batch(&self, texts: Vec<String>) -> Result<Vec<Vec<f64>>, AgentError> {
let url = format!("{}/v1/embeddings", self.base_url.trim_end_matches('/')); let url = format!("{}/v1/embeddings", self.base_url.trim_end_matches('/'));
let request_body = EmbeddingRequest { let request_body = EmbeddingRequest {
@@ -72,3 +90,33 @@ impl LlmClient {
Ok(data.into_iter().map(|d| d.embedding).collect()) Ok(data.into_iter().map(|d| d.embedding).collect())
} }
} }
#[cfg(test)]
mod tests {
use super::*;
use secrecy::SecretString;
fn client() -> LlmClient {
LlmClient::new(
"http://unused".into(),
SecretString::from(String::new()),
"m".into(),
"e".into(),
)
}
#[tokio::test]
async fn empty_input_makes_no_request() {
// Must short-circuit before any HTTP call (base_url is unroutable).
let out = client().embed(Vec::new()).await.unwrap();
assert!(out.is_empty());
}
#[test]
fn batch_size_is_within_backend_cap() {
assert!(
EMBED_BATCH_SIZE <= 25,
"must stay under the bge 25-input cap"
);
}
}
+145
View File
@@ -0,0 +1,145 @@
//! C5 live verification — the semantic master-controls path end to end against the
//! deployed api-dev catalog. Ignored (hits api-dev + LiteLLM). Run explicitly:
//!
//! set -a; . ./.env; set +a
//! BREAKPILOT_BASE_URL=https://api-dev.breakpilot.ai \
//! cargo test -p compliance-agent --test c5_semantic_live -- --ignored --nocapture
//!
//! Pulls the live master-controls catalog, embeds the corpus (chunked), then for a
//! couple of real vulnerable findings retrieves the nearest master controls and
//! grounded-judges them, stamping master-control refs.
mod common;
use std::sync::Arc;
use compliance_agent::llm::LlmClient;
use compliance_core::config::BreakpilotConfig;
use compliance_core::models::finding::{Finding, Severity};
use compliance_core::models::scan::ScanType;
use secrecy::SecretString;
fn env(k: &str) -> String {
std::env::var(k).unwrap_or_else(|_| panic!("env {k} must be set for the live C5 test"))
}
fn mk_finding(file: &str, line: u32, title: &str) -> Finding {
let mut f = Finding::new(
"repo-c5".into(),
format!("{file}:{line}"),
"semgrep".into(),
ScanType::Sast,
title.into(),
title.into(),
Severity::High,
);
f.file_path = Some(file.into());
f.line_number = Some(line);
f
}
#[tokio::test]
#[ignore = "live: requires deployed api-dev master-controls (fetch+parse only, no LLM)"]
async fn c5_ingest_master_controls_catalog() {
use compliance_agent::controls::OscalControlsProvider;
let provider = OscalControlsProvider::new(
reqwest::Client::new(),
env("BREAKPILOT_BASE_URL"),
None,
std::env::temp_dir().join("c5-ingest-snap"),
);
let doc = provider
.load_master_controls()
.await
.expect("pull + parse master-controls catalog");
let controls = doc.to_controls();
println!(
"\n=== C5 ingest: {} master controls parsed ===",
controls.len()
);
for c in controls.iter().take(4) {
let text: String = c.text.chars().take(90).collect();
println!(" {} | {} | {}", c.id, c.title, text);
}
assert!(
!controls.is_empty(),
"expected a non-empty master-control corpus"
);
}
#[tokio::test]
#[ignore = "live: requires deployed api-dev master-controls + LiteLLM"]
async fn c5_semantic_stamps_master_control_refs() {
let llm = Arc::new(LlmClient::new(
env("LITELLM_URL"),
SecretString::from(env("LITELLM_API_KEY")),
env("LITELLM_MODEL"),
env("LITELLM_EMBED_MODEL"),
));
let mut config = common::dev_config("mongodb://unused".into(), "c5".into());
let snapshot = std::env::temp_dir().join("c5-oscal-snap");
config.breakpilot = BreakpilotConfig {
base_url: Some(env("BREAKPILOT_BASE_URL")),
token: None,
snapshot_dir: snapshot.to_string_lossy().into_owned(),
semantic_mapping: true,
grounded_control_checks: false,
};
// Fixture repo with recognizable code-checkable surfaces.
let repo = std::env::temp_dir().join("c5-fixture-repo");
let _ = std::fs::remove_dir_all(&repo);
std::fs::create_dir_all(repo.join("app")).expect("mkdir");
std::fs::write(
repo.join("app/auth.py"),
concat!(
"import hashlib\n",
"\n",
"def store_password(user, password):\n",
" # weak, unsalted password hashing\n",
" digest = hashlib.md5(password.encode()).hexdigest()\n",
" db.save(user, digest)\n",
"\n",
"@app.route('/login', methods=['POST'])\n",
"def login():\n",
" u = request.form['username']\n",
" p = request.form['password']\n",
" return 'ok' if check(u, p) else ('bad', 401)\n",
),
)
.expect("write fixture");
let mut findings = vec![
mk_finding("app/auth.py", 5, "Weak password hash (md5, unsalted)"),
mk_finding(
"app/auth.py",
9,
"Login endpoint without brute-force protection",
),
];
let tagged =
compliance_agent::controls::semantic_stamp_findings(&config, llm, &repo, &mut findings)
.await;
println!("\n=== C5 semantic master-controls stamping ===");
for f in &findings {
println!(
" {:50} {}:{:?} -> {:?}",
f.title,
f.file_path.as_deref().unwrap_or(""),
f.line_number,
f.control_refs
);
}
println!("findings that gained >=1 master-control ref: {tagged}");
let _ = std::fs::remove_dir_all(&repo);
// Live corpus — assert only that the path runs and stamps at least one ref.
assert!(
tagged >= 1,
"expected at least one finding to gain a master-control ref"
);
}
+1
View File
@@ -36,6 +36,7 @@ export default withMermaid(defineConfig({
{ text: 'Pentest Architecture', link: '/features/pentest-architecture' }, { text: 'Pentest Architecture', link: '/features/pentest-architecture' },
{ text: 'AI Chat', link: '/features/ai-chat' }, { text: 'AI Chat', link: '/features/ai-chat' },
{ text: 'Code Knowledge Graph', link: '/features/graph' }, { text: 'Code Knowledge Graph', link: '/features/graph' },
{ text: 'Compliance Control Mapping', link: '/features/control-mapping' },
{ text: 'MCP Integration', link: '/features/mcp-server' }, { text: 'MCP Integration', link: '/features/mcp-server' },
], ],
}, },
+127
View File
@@ -0,0 +1,127 @@
# Compliance Control Mapping
Control mapping connects the scanner's raw output — deterministic tool findings and the code itself — to the **compliance controls** each piece of evidence supports. A hardcoded credential stops being just "CWE-798 from semgrep" and becomes evidence for *"cra-ai-8: no default passwords"* and, at scale, master control *`mc-31761` hardcoded_secrets_detection*. Findings carry those references (`control_refs`) into the dashboard and out over the MCP server as OSCAL, so the compliance report is built from real, grounded findings rather than a questionnaire.
## The core principle: tools detect, the LLM judges
The design has one rule, borrowed from the ZeroFalse / IRIS line of research: **deterministic tools are the detectors; the LLM is only ever a grounded false-positive filter, never the thing that finds the issue.**
- A tool (semgrep, gitleaks, syft/osv, ZAP, nuclei) detects deterministically.
- An **authored, human-reviewed lookup table** (`control-map`) maps that detection to the control(s) it's evidence for.
- The LLM enters last, to *confirm or refute* the mapping against the actual code — and every surviving verdict is anchored to a verbatim snippet by the grounding gate.
This keeps hallucination out of detection. The LLM supplies cross-language, cross-stack pattern *recognition*; the surrounding machinery supplies determinism.
## Coverage model
Every control lands in one of three buckets, recorded in the `control-map` LUT (`control-map/data/cra_control_map.json`) and never decided by an LLM:
| Bucket | Meaning |
| --- | --- |
| `covered` | An existing tool's scan surfaces findings for this control |
| `needs_tooling` | Code-checkable, but no off-the-shelf tool digs it out — we author a detector or use the grounded surface check |
| `not_code_checkable` | A design/process property — out of static-scan scope |
For the **CRA** framework (40 controls) the split is **13 covered · 8 needs_tooling · 19 not_code_checkable**. The 16 originally-uncovered controls were resolved as a hybrid:
- **4 custom semgrep detectors** (`cra-ai-1`, `7`, `10`, `14`) — secure-by-default, weak password hashing, insecure session cookies, weak data-at-rest ciphers. Shipped in the binary and matched back to controls **by rule id** so a broad CWE can't over-attribute.
- **8 grounded surface checks** (`cra-ai-6`, `11`, `12`, `24`, `27`, `28`, `29`, `30`) — the absence-based controls (no rate limiting, no security logging, no update-signature check…) that have no syntactic pattern.
- **4 marked not_code_checkable** (`cra-ai-2`, `3`, `4`, `5`) — minimal attack surface, secure architecture, least privilege, tamper protection.
At scale, the **master-controls** corpus (breakpilot's deduped clusters, exported as OSCAL) currently provides **~2,882 code-checkable controls** (2,143 `network` + 739 `source_code`), matched semantically.
## The three mapping paths
```mermaid
flowchart TD
T[Deterministic tools\nsemgrep · gitleaks · syft/osv · ZAP] --> F[Findings]
F --> B["Stage 5b — LUT triage\ncontrols_for(tool, cwe / rule_id)"]
F --> C["Stage 5c — Semantic\nembed region+intent → top-K master controls"]
R[Repo source] --> D["Stage 5d — Grounded surface\nretrieve surface for absence-based controls"]
B --> J{{Grounded LLM judge\ntemp 0 · verbatim snippet}}
C --> J
D --> J
J -->|snippet grounds in region| S[Stamp control_refs]
J -->|refuted / ungrounded| X[Dropped]
```
All three paths converge on the same **grounded judge** and the same **grounding gate**. They differ only in how candidate (finding/region, control) pairs are produced.
### Stage 5b — deterministic LUT triage
The default path. A tool finding is matched to controls via `control_map.controls_for_finding(tool, cwe, rule_id)`; the judge then confirms each mapped control against the code region. Outcomes: `Confirmed([ids])` (stamp them), `FalsePositive` (drop the finding), or `Unmapped` (keep it untagged). Runs whenever `BREAKPILOT_BASE_URL` is set.
### Stage 5c — semantic retrieval (master-controls scale)
Master controls carry no CWE, so they can't be LUT-mapped. Instead we map by *similarity*: embed every control's requirement text once (cached), then for each finding retrieve the top-K nearest controls and hand them to the judge. Gated behind `BREAKPILOT_SEMANTIC_MAPPING` (default off). See [Semantic retrieval](#semantic-retrieval-in-detail).
### Stage 5d — grounded surface checks (absence-based controls)
Some controls are violated by an *absence* — no rate limiting on login, no security logging, no signature check on an update. There's no pattern for semgrep to match, so we deterministically retrieve the code **surface** the control governs (a login route, a logging setup, update/download code) by identifier/route terms, and let the judge decide whether the control holds there. Produces net-new, already-grounded findings. Gated behind `BREAKPILOT_GROUNDED_CHECKS` (default off).
## The grounding gate
No matter the path, a verdict becomes a finding only if it survives `compliance_core::control_check::ground`:
1. The judge runs at **temperature 0** with a closed prompt and must quote the offending code **verbatim** into `snippet`.
2. That snippet must appear **literally** in the retrieved region — otherwise the verdict is dropped.
3. The finding's line is **recomputed from the match**; the model's own line number is never trusted.
4. Verdicts are cached by content hash, so re-scans reproduce.
The model is allowed to be smart; it is never trusted.
## Semantic retrieval in detail
1. **Embed the corpus once.** Each control's requirement text is embedded with `bge-multilingual-gemma2` (3584-dim — multilingual matters, the master controls are in German while code is English). The embedding backend caps input arrays at 25 per request, so `embed()` chunks at 16; the whole `ControlIndex` is persisted to `snapshot_dir` keyed by a **corpus hash**, so only the first scan after a catalog change pays the embedding cost.
2. **Build the query from the finding's intent, not just the code.** The retrieval query is `finding.title + finding.description + region`, not the raw region. This is the single most important tuning: two findings in one file share overlapping windows and, on the code alone, embed alike and collapse onto the same controls. The finding's own words ("brute-force protection" vs "weak hash") carry the discriminating signal. The raw region still goes to the judge for grounding.
3. **Retrieve → judge → ground.** Top-K nearest by cosine, each judged against the region, each grounded.
## Worked examples
Both examples are from the live end-to-end verification (`c5_semantic_live.rs`) against the real ~2,882-control corpus.
### Example 1 — a small auth file (the tuning story)
Two findings in one `auth.py`: a weak `hashlib.md5(password)` hash and a login endpoint with no brute-force protection.
| Finding | Region-only retrieval | Intent-enriched retrieval |
| --- | --- | --- |
| Weak md5 hash | 19874, 20683, 23149, 29985 | **`mc-23149`** (eliminate weak unsalted hashes) at rank 1, + `mc-21634` salted hashing |
| Login w/o brute-force protection | *identical 4, reordered* | newly surfaces **`mc-19984`** brute_force_protection + **`mc-23186`** account_lockout |
Region-only retrieval gave both findings the *same* four password-hashing controls — the brute-force finding never found its real controls because its window is saturated with `password` tokens. Enriching the query with the finding's intent fixed it: the brute-force finding now pulls the correct rate-limiting / lockout controls out of the 2,882.
### Example 2 — four topically distinct vulnerabilities
| Finding | Top matched controls | Family |
| --- | --- | --- |
| SQL injection (string-concat query) | `sql_injection_prevention`, `sql_injection`, `parameterized_queries`, input_sanitization | input-validation ✓ |
| Hardcoded API credential | `hardcoded_secrets_detection`, credential_scanning, secrets_detection | credentials ✓ |
| TLS verification disabled (`verify=False`) | `https_enforcement`, `configuration_verification`, transport config | transport-encryption ✓ |
| Insecure deserialization (`pickle.loads`) | `deserialization`, `deserialization_testing`, `deserialization_security` | deserialization ✓ |
Every finding maps to its exact control family, with the most specific control often at the top, and the four sets are distinct.
## Known limitations
- **Absence findings are weak for semantic retrieval.** Similarity matches what code *is about*, not what it *lacks*; a "missing rate limiting" finding embeds like login code. This is exactly why the grounded surface path (Stage 5d) exists — it decides presence/absence at a retrieved surface rather than by embedding distance.
- **Generic catch-all controls co-occur.** `mc-20890 secure_development_security_code_review` appears in the top-K for many code-security findings because it is semantically near almost all of them. It's harmless (the judge grounds it, and it never crowds out the specific controls — the SQLi example didn't get it) but is a candidate for future down-weighting.
- **Corpus classification noise.** The master-controls `verification_method` classification is imperfect — e.g. a documentation control (`eu_declaration_accuracy`) is currently tagged `source_code`. That's a corpus-side data-quality issue, separate from the mapping engine.
## Configuration
| Variable | Effect |
| --- | --- |
| `BREAKPILOT_BASE_URL` | breakpilot-compliance root; enables control ingest + Stage 5b. Unset disables all control mapping. |
| `BREAKPILOT_SEMANTIC_MAPPING` | Enables Stage 5c (semantic master-controls mapping). Default off. |
| `BREAKPILOT_GROUNDED_CHECKS` | Enables Stage 5d (grounded surface checks). Default off. |
| `BREAKPILOT_SNAPSHOT_DIR` | Where OSCAL catalog snapshots and the cached control-embedding index live. |
The semantic and grounded passes are gated because they are the heavier, less deterministic paths; they stay off until verified live against a deployed catalog. The live verification lives in `compliance-agent/tests/c5_semantic_live.rs` (ignored; run with `--ignored`).
## Appendix — the master-controls data pipeline
The master-controls corpus is produced by breakpilot-compliance and pulled as an OSCAL catalog from `GET /api/compliance/v1/oscal/catalog?framework=master-controls`. Two operational lessons are worth recording, because they cost real time to diagnose:
- **The catalog is served from `breakpilot_db`, not `postgres`.** Diagnostics run against the wrong database will look clean while the app serves something else entirely. Confirm the app's datname (`pg_stat_activity`) before trusting any count or `EXPLAIN`.
- **A constraint-less dump triplicated the master-control tables.** Restored without their PK/unique constraints, `master_controls` / `mc_verification` / `master_control_members` accumulated identical rows 3× (the same artifact migration `158` fixed for `doc_check_controls`). That inflated the catalog to ~26k dup'd controls and, with the indexes also missing, drove the export query to a >120s / 502. The fix (breakpilot migration `160`) ctid-dedups each table by its natural key and restores the constraints + indexes so it can't recur; the export query was also rewritten set-based (a single windowed pass instead of a per-row correlated subquery). After dedup: 41,850 → 13,950 master controls, catalog **25,938 → 2,882** code-checkable, endpoint **502 → 200 in ~3s**.