Compare commits

..
Author SHA1 Message Date
Sharang ParnerkarandClaude Opus 4.7 eb88656c43 feat(dashboard): UI for managing MCP tokens
CI / Check (pull_request) Failing after 4m33s
CI / Detect Changes (pull_request) Has been skipped
CI / Deploy Agent (pull_request) Has been skipped
CI / Deploy Dashboard (pull_request) Has been skipped
CI / Deploy Docs (pull_request) Has been skipped
CI / Deploy MCP (pull_request) Has been skipped
Adds /mcp-tokens page that lets a logged-in user mint, list, and
revoke bearer tokens for the MCP server. Stacks on #92 (which added
the agent endpoints + middleware) — once both land, the loop is
closed: a user can copy a token from the dashboard straight into
their Claude Desktop / Cursor / ChatGPT MCP config.

UX
- "Create Token" button → inline form with name input.
- On submit, server function calls `POST /api/v1/mcp-tokens`. The
  raw token is shown ONCE in a prominent yellow banner with a copy
  button and a "won't be shown again" warning, then the user
  dismisses it manually.
- List view: card per token with name, prefix `mcpt_xxxx…`, created
  date, last_used (or "never"). Revoked tokens render dimmed with a
  "revoked" pill. Active tokens have a trash button → confirm
  modal → soft delete.
- Toast feedback on create/revoke success/failure.

Files
- infrastructure/mcp_tokens.rs (new) — three #[server] functions:
  fetch_mcp_tokens, create_mcp_token, revoke_mcp_token. All go
  through agent_client so the Keycloak Bearer is auto-attached;
  the agent then enforces tenant scoping on every endpoint.
- pages/mcp_tokens.rs (new) — the page component itself.
- app.rs — adds Route::McpTokensPage at /mcp-tokens.
- pages/mod.rs, infrastructure/mod.rs — module + re-export wiring.

Timestamp format
- The agent serializes BSON DateTime as extended JSON
  `{"$date":{"$numberLong":"..."}}`. Page has a small helper that
  accepts that shape, plain ISO strings, or anything else
  (best-effort). Same approach used elsewhere in the dashboard so
  there's no new dependency.

Test plan
- cargo fmt --all clean
- cargo clippy -p compliance-dashboard --features server
  -- -D warnings clean
- cargo clippy -p compliance-dashboard --features web
  --no-default-features -- -D warnings clean
- cargo check on both feature sets clean

Followup
- No sidebar entry yet (matches mcp_servers — settings-style
  pages are reached via direct URL today). Worth adding a
  Settings sub-menu in a separate UX pass.
- Token expiry + per-tool scope when those land on the agent side
  will need a small UI for the create modal (extra fields).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-30 17:27:36 +02:00
55 changed files with 297 additions and 6018 deletions
-13
View File
@@ -7,17 +7,4 @@ ignore = [
# not a realistic attack surface here. Revisit when mongodb bumps hickory.
"RUSTSEC-2026-0118", # NSEC3 loop, no fix available upstream
"RUSTSEC-2026-0119", # O(n²) name compression, fixed in hickory-proto >=0.26.1
# rmcp 0.16.0 — DNS rebinding in Streamable HTTP server transport (missing
# Host header validation). Patched in rmcp >= 1.4.0, which is a major API
# version jump from our pin; rmcp shipped 0.x → 1.x → 2.x in three months
# and the migration touches every tool handler + the auth middleware we
# just landed in #92. Threat model in our deployment: the MCP server is
# exposed at a public hostname (comp-mcp-dev.meghsakha.com) behind orca's
# TLS-terminating ingress with per-tenant bearer auth — the attack model
# (browser DNS-rebinding into localhost MCP server) doesn't directly apply.
# Defense-in-depth Host-header check is still a worthwhile follow-up.
# FOLLOW-UP: bump rmcp to 2.x in a dedicated PR (M7.3 follow-up, sized
# multi-hour due to API surface change).
"RUSTSEC-2026-0189",
]
+10 -67
View File
@@ -9,33 +9,15 @@ on:
env:
CARGO_TERM_COLOR: always
RUSTFLAGS: "-D warnings"
# Compile cache: sccache -> Hetzner S3 (breakpilot-sccache), runner-independent
# and persistent across CI runs (own key prefix). Reuses the shared cluster S3
# creds (same bucket as werkpilot). Requires repo secrets HETZNER_S3_ACCESS_KEY
# and HETZNER_S3_SECRET_KEY.
# sccache caches compilation artifacts within a job so that compiling
# both --features server and --features web shares common crate work.
RUSTC_WRAPPER: /usr/local/bin/sccache
SCCACHE_BUCKET: breakpilot-sccache
SCCACHE_ENDPOINT: https://nbg1.your-objectstorage.com
SCCACHE_REGION: auto
SCCACHE_S3_USE_SSL: "true"
SCCACHE_S3_KEY_PREFIX: compliance-scanner
AWS_ACCESS_KEY_ID: ${{ secrets.HETZNER_S3_ACCESS_KEY }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.HETZNER_S3_SECRET_KEY }}
# compliance-agent depends on tramiton-core via git; use the system git so the
# credential rewrite below (see "Configure git auth ...") is honored on fetch.
CARGO_NET_GIT_FETCH_WITH_CLI: "true"
# Throttle cargo so a ~670-crate concurrent download burst doesn't 429 the
# Kellnr mirror: fewer concurrent connections (HTTP/1.1) + more retries.
CARGO_NET_RETRY: "10"
CARGO_HTTP_MULTIPLEXING: "false"
SCCACHE_DIR: /tmp/sccache
# Cancel superseded PR runs, but NEVER cancel main-branch runs — those build and
# deploy per-service images, and cancelling one merge's deploy when the next
# merge lands leaves a service un-deployed (as happened between two back-to-back
# merges). So cancel-in-progress only for pull_request events.
# Cancel in-progress runs for the same branch/PR
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
cancel-in-progress: true
jobs:
# ---------------------------------------------------------------------------
@@ -54,44 +36,16 @@ jobs:
git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}"
git checkout FETCH_HEAD
# Resolve crates.io deps through the self-hosted Kellnr mirror (cached,
# crates.io-independent). Git deps (tramiton-core) are unaffected — source
# replacement only applies to crates.io-sourced crates.
- name: Use Kellnr crates.io mirror
run: |
: "${CARGO_HOME:=/usr/local/cargo}"
mkdir -p "$CARGO_HOME"
{
echo '[source.crates-io]'
echo 'replace-with = "kellnr"'
echo '[registries.kellnr]'
echo 'index = "sparse+https://crates.meghsakha.com/api/v1/cratesio/"'
} >> "$CARGO_HOME/config.toml"
env:
RUSTC_WRAPPER: ""
- name: Install tools
run: |
rustup component add rustfmt clippy
curl -fsSL https://github.com/mozilla/sccache/releases/download/v0.10.0/sccache-v0.10.0-x86_64-unknown-linux-musl.tar.gz \
| tar xz --strip-components=1 -C /usr/local/bin/ sccache-v0.10.0-x86_64-unknown-linux-musl/sccache
curl -fsSL https://github.com/mozilla/sccache/releases/download/v0.9.1/sccache-v0.9.1-x86_64-unknown-linux-musl.tar.gz \
| tar xz --strip-components=1 -C /usr/local/bin/ sccache-v0.9.1-x86_64-unknown-linux-musl/sccache
chmod +x /usr/local/bin/sccache
cargo install cargo-audit --locked
env:
RUSTC_WRAPPER: ""
# compliance-agent has a git dependency on tramiton-core (a private repo on
# this Gitea instance). Rewrite its SSH URL to HTTPS + a PAT so the runner
# can fetch it. Requires the repo secret TRAMITON_FETCH_TOKEN (a Gitea PAT
# with read:repository, owned by a user with access to sharang/tramiton).
# (Honored on fetch because CARGO_NET_GIT_FETCH_WITH_CLI=true uses system git.)
- name: Configure git auth for private tramiton dependency
run: |
git config --global \
url."https://sharang:${{ secrets.TRAMITON_FETCH_TOKEN }}@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"
env:
RUSTC_WRAPPER: ""
# Format (no compilation needed)
- name: Format
run: cargo fmt --all --check
@@ -194,18 +148,13 @@ jobs:
image: docker:27-cli
steps:
- name: Build, push and trigger orca redeploy
env:
# PAT for fetching the private tramiton-core git dependency during the
# image build (injected as a BuildKit secret, never baked into a layer).
TRAMITON_FETCH_TOKEN: ${{ secrets.TRAMITON_FETCH_TOKEN }}
run: |
apk add --no-cache git curl openssl
git init && git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}" && git checkout FETCH_HEAD
IMAGE=registry.meghsakha.com/compliance-agent
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login registry.meghsakha.com -u "${{ secrets.REGISTRY_USERNAME }}" --password-stdin
DOCKER_BUILDKIT=1 docker build --secret id=tramiton_token,env=TRAMITON_FETCH_TOKEN \
-f Dockerfile.agent -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker build -f Dockerfile.agent -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker push "$IMAGE:latest" && docker push "$IMAGE:${GITHUB_SHA}"
PAYLOAD=$(printf '{"ref":"refs/heads/main","repository":{"full_name":"sharang/compliance-scanner-agent"},"head_commit":{"id":"%s","message":"deploy agent"}}' "${GITHUB_SHA}")
SIG=$(printf '%s' "$PAYLOAD" | openssl dgst -sha256 -hmac "${{ secrets.ORCA_WEBHOOK_SECRET }}" | awk '{print $2}')
@@ -220,16 +169,13 @@ jobs:
image: docker:27-cli
steps:
- name: Build, push and trigger orca redeploy
env:
TRAMITON_FETCH_TOKEN: ${{ secrets.TRAMITON_FETCH_TOKEN }}
run: |
apk add --no-cache git curl openssl
git init && git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}" && git checkout FETCH_HEAD
IMAGE=registry.meghsakha.com/compliance-dashboard
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login registry.meghsakha.com -u "${{ secrets.REGISTRY_USERNAME }}" --password-stdin
DOCKER_BUILDKIT=1 docker build --secret id=tramiton_token,env=TRAMITON_FETCH_TOKEN \
-f Dockerfile.dashboard -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker build -f Dockerfile.dashboard -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker push "$IMAGE:latest" && docker push "$IMAGE:${GITHUB_SHA}"
PAYLOAD=$(printf '{"ref":"refs/heads/main","repository":{"full_name":"sharang/compliance-scanner-agent"},"head_commit":{"id":"%s","message":"deploy dashboard"}}' "${GITHUB_SHA}")
SIG=$(printf '%s' "$PAYLOAD" | openssl dgst -sha256 -hmac "${{ secrets.ORCA_WEBHOOK_SECRET }}" | awk '{print $2}')
@@ -265,16 +211,13 @@ jobs:
image: docker:27-cli
steps:
- name: Build, push and trigger orca redeploy
env:
TRAMITON_FETCH_TOKEN: ${{ secrets.TRAMITON_FETCH_TOKEN }}
run: |
apk add --no-cache git curl openssl
git init && git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}" && git checkout FETCH_HEAD
IMAGE=registry.meghsakha.com/compliance-mcp
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login registry.meghsakha.com -u "${{ secrets.REGISTRY_USERNAME }}" --password-stdin
DOCKER_BUILDKIT=1 docker build --secret id=tramiton_token,env=TRAMITON_FETCH_TOKEN \
-f Dockerfile.mcp -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker build -f Dockerfile.mcp -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker push "$IMAGE:latest" && docker push "$IMAGE:${GITHUB_SHA}"
PAYLOAD=$(printf '{"ref":"refs/heads/main","repository":{"full_name":"sharang/compliance-scanner-agent"},"head_commit":{"id":"%s","message":"deploy mcp"}}' "${GITHUB_SHA}")
SIG=$(printf '%s' "$PAYLOAD" | openssl dgst -sha256 -hmac "${{ secrets.ORCA_WEBHOOK_SECRET }}" | awk '{print $2}')
Generated
+13 -115
View File
@@ -692,9 +692,6 @@ dependencies = [
"tower-http",
"tracing",
"tracing-subscriber",
"tramiton-core",
"tramiton-repro",
"tramiton-sbom",
"urlencoding",
"uuid",
"walkdir",
@@ -1121,9 +1118,9 @@ dependencies = [
[[package]]
name = "crossbeam-epoch"
version = "0.9.20"
version = "0.9.18"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2d6914041f254d6e9176c01941b21115dcfb7089e55135a35411081bd106ef3f"
checksum = "5b82ac4a3c2ca9c3460964f020e1402edd5753411d7737aa39c3714ad1b5420e"
dependencies = [
"crossbeam-utils",
]
@@ -2103,7 +2100,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "39cab71617ae0d63f51a36d69f866391735b51691dbda63cf6f96d042b63efeb"
dependencies = [
"libc",
"windows-sys 0.52.0",
"windows-sys 0.61.2",
]
[[package]]
@@ -3698,7 +3695,7 @@ version = "0.50.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7957b9740744892f114936ab4a57b3f487491bbeafaf8083688b16841a4240e5"
dependencies = [
"windows-sys 0.60.2",
"windows-sys 0.61.2",
]
[[package]]
@@ -3769,15 +3766,6 @@ dependencies = [
"syn",
]
[[package]]
name = "object"
version = "0.36.7"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "62948e14d923ea95ea2c7c86c71013138b66525b86bdc08d2dcc262bdb497b87"
dependencies = [
"memchr",
]
[[package]]
name = "octocrab"
version = "0.44.1"
@@ -4209,7 +4197,7 @@ version = "3.4.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "219cb19e96be00ab2e37d6e299658a0cfa83e52429179969b0f0121b4ac46983"
dependencies = [
"toml_edit 0.23.10+spec-1.0.0",
"toml_edit",
]
[[package]]
@@ -4294,9 +4282,9 @@ dependencies = [
[[package]]
name = "quinn-proto"
version = "0.11.15"
version = "0.11.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4fcb935c5bec503c2f0e306bdd3e58bb9029dcb14fa8d9ac76e3a5256ac0763e"
checksum = "434b42fec591c96ef50e21e886936e66d3cc3f737104fdb9b737c40ffb94c098"
dependencies = [
"bytes",
"getrandom 0.3.4",
@@ -4679,7 +4667,7 @@ dependencies = [
"errno",
"libc",
"linux-raw-sys 0.4.15",
"windows-sys 0.52.0",
"windows-sys 0.59.0",
]
[[package]]
@@ -4692,7 +4680,7 @@ dependencies = [
"errno",
"libc",
"linux-raw-sys 0.12.1",
"windows-sys 0.52.0",
"windows-sys 0.61.2",
]
[[package]]
@@ -5008,15 +4996,6 @@ dependencies = [
"syn",
]
[[package]]
name = "serde_spanned"
version = "0.6.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bf41e0cfaf7226dca15e8197172c295a782857fcb97fad1808a166870dee75a3"
dependencies = [
"serde",
]
[[package]]
name = "serde_urlencoded"
version = "0.7.1"
@@ -5570,10 +5549,10 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "82a72c767771b47409d2345987fda8628641887d5466101319899796367354a0"
dependencies = [
"fastrand",
"getrandom 0.3.4",
"getrandom 0.4.1",
"once_cell",
"rustix 1.1.4",
"windows-sys 0.52.0",
"windows-sys 0.61.2",
]
[[package]]
@@ -5831,27 +5810,6 @@ dependencies = [
"tokio",
]
[[package]]
name = "toml"
version = "0.8.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dc1beb996b9d83529a9e75c17a1686767d148d70663143c7854d8b4a09ced362"
dependencies = [
"serde",
"serde_spanned",
"toml_datetime 0.6.11",
"toml_edit 0.22.27",
]
[[package]]
name = "toml_datetime"
version = "0.6.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "22cddaf88f4fbc13c51aebbf5f8eceb5c7c5a9da2ac40a13519eb5b0a0e8f11c"
dependencies = [
"serde",
]
[[package]]
name = "toml_datetime"
version = "0.7.5+spec-1.1.0"
@@ -5861,20 +5819,6 @@ dependencies = [
"serde_core",
]
[[package]]
name = "toml_edit"
version = "0.22.27"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "41fe8c660ae4257887cf66394862d21dbca4a6ddd26f04a3560410406a2f819a"
dependencies = [
"indexmap 2.13.0",
"serde",
"serde_spanned",
"toml_datetime 0.6.11",
"toml_write",
"winnow",
]
[[package]]
name = "toml_edit"
version = "0.23.10+spec-1.0.0"
@@ -5882,7 +5826,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "84c8b9f757e028cee9fa244aea147aab2a9ec09d5325a9b01e0a49730c2b5269"
dependencies = [
"indexmap 2.13.0",
"toml_datetime 0.7.5+spec-1.1.0",
"toml_datetime",
"toml_parser",
"winnow",
]
@@ -5896,12 +5840,6 @@ dependencies = [
"winnow",
]
[[package]]
name = "toml_write"
version = "0.1.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5d99f8c9a7727884afe522e9bd5edbfc91a3312b36a77b5fb8926e4c31a41801"
[[package]]
name = "tonic"
version = "0.12.3"
@@ -6148,46 +6086,6 @@ dependencies = [
"wasm-bindgen",
]
[[package]]
name = "tramiton-core"
version = "0.4.1"
source = "git+ssh://git@gitea.meghsakha.com:22222/sharang/tramiton.git?tag=v0.4.1#ae4fc1376279f9edb9882605b20877335e7ba8ba"
dependencies = [
"serde",
"tempfile",
"thiserror 1.0.69",
"toml",
"walkdir",
]
[[package]]
name = "tramiton-repro"
version = "0.4.1"
source = "git+ssh://git@gitea.meghsakha.com:22222/sharang/tramiton.git?tag=v0.4.1#ae4fc1376279f9edb9882605b20877335e7ba8ba"
dependencies = [
"serde",
"serde_json",
"sha2",
"tempfile",
"thiserror 1.0.69",
"toml",
"tramiton-core",
"walkdir",
]
[[package]]
name = "tramiton-sbom"
version = "0.4.1"
source = "git+ssh://git@gitea.meghsakha.com:22222/sharang/tramiton.git?tag=v0.4.1#ae4fc1376279f9edb9882605b20877335e7ba8ba"
dependencies = [
"object",
"serde",
"serde_json",
"sha2",
"tramiton-core",
"tramiton-repro",
]
[[package]]
name = "tree-sitter"
version = "0.24.7"
@@ -6746,7 +6644,7 @@ version = "0.1.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c2a7b1c03c876122aa43f3020e6c3c3ee5c05081c9a00739faf7503aeba10d22"
dependencies = [
"windows-sys 0.48.0",
"windows-sys 0.61.2",
]
[[package]]
+2 -41
View File
@@ -2,22 +2,7 @@ FROM rust:1.94-bookworm AS builder
WORKDIR /app
COPY . .
# compliance-agent depends on the private tramiton-core git repo. Authenticate
# the fetch with a PAT passed as a BuildKit secret (never baked into a layer).
# Build with: DOCKER_BUILDKIT=1 docker build --secret id=tramiton_token,env=TRAMITON_FETCH_TOKEN ...
RUN --mount=type=secret,id=tramiton_token \
if [ -s /run/secrets/tramiton_token ]; then \
git config --global \
url."https://sharang:$(cat /run/secrets/tramiton_token)@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"; \
fi && \
CARGO_NET_GIT_FETCH_WITH_CLI=true cargo build --release -p compliance-agent
# A throwaway stage that packs a real nix store (store paths + the validity DB)
# into a compressed bootstrap tarball. Only the tarball is copied into the final
# image, so we don't carry a raw /nix copy layer.
FROM nixos/nix:latest AS nixseed
RUN tar -C / -czf /nix-bootstrap.tar.gz nix
RUN cargo build --release -p compliance-agent
FROM debian:bookworm-slim
RUN apt-get update && apt-get install -y ca-certificates libssl3 git curl python3 python3-pip npm golang-go php-cli && rm -rf /var/lib/apt/lists/*
@@ -46,30 +31,7 @@ RUN pip3 install --break-system-packages semgrep
# Install ruff for Python linting
RUN pip3 install --break-system-packages ruff
# Real nix for the tramiton reproducible-build firmware SBOM.
#
# nix-portable's proot fallback can't run here: user namespaces are blocked by
# the container's default seccomp/apparmor profile, and orca exposes no way to
# relax it. So ship a *real* nix and disable its build sandbox
# (`sandbox = false`) — a plain gcc/make firmware build needs no user namespace,
# so it runs fine under the locked-down profile with no proot involved.
#
# The store is shipped as a bootstrap tarball and seeded onto /nix at first
# start (see docker/agent-entrypoint.sh), so a persistent /nix volume survives
# redeploys. A missing/broken nix just falls back to the analysis-only SBOM.
COPY --from=nixseed /nix-bootstrap.tar.gz /opt/nix-bootstrap.tar.gz
ENV PATH="/nix/var/nix/profiles/default/bin:${PATH}"
RUN mkdir -p /etc/nix && printf '%s\n' \
'experimental-features = nix-command flakes' \
'sandbox = false' \
'build-users-group =' \
'substituters = https://cache.nixos.org' \
'trusted-public-keys = cache.nixos.org-1:6NCHdD59X431o0gWypbMrAURkbJ16ZPMQFGspcDShjY=' \
> /etc/nix/nix.conf
COPY --from=builder /app/target/release/compliance-agent /usr/local/bin/compliance-agent
COPY docker/agent-entrypoint.sh /usr/local/bin/agent-entrypoint.sh
RUN chmod +x /usr/local/bin/agent-entrypoint.sh
# Copy documentation for the help chat assistant
COPY --from=builder /app/README.md /app/README.md
@@ -81,6 +43,5 @@ RUN mkdir -p /data/compliance-scanner/ssh
EXPOSE 3001 3002
# Seeds /nix (fresh volume) from the bootstrap tarball, then runs the agent.
ENTRYPOINT ["/usr/local/bin/agent-entrypoint.sh"]
ENTRYPOINT ["compliance-agent"]
+1 -10
View File
@@ -7,16 +7,7 @@ ARG DOCS_URL=/docs
WORKDIR /app
COPY . .
ENV DOCS_URL=${DOCS_URL}
# compliance-agent (a workspace member) depends on the private tramiton-core git
# repo, so the workspace resolve needs it even to build the dashboard.
# Authenticate the fetch with a PAT passed as a BuildKit secret.
RUN --mount=type=secret,id=tramiton_token \
if [ -s /run/secrets/tramiton_token ]; then \
git config --global \
url."https://sharang:$(cat /run/secrets/tramiton_token)@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"; \
fi && \
CARGO_NET_GIT_FETCH_WITH_CLI=true dx build --release --package compliance-dashboard
RUN dx build --release --package compliance-dashboard
FROM debian:bookworm-slim
RUN apt-get update && apt-get install -y ca-certificates libssl3 && rm -rf /var/lib/apt/lists/*
+1 -10
View File
@@ -2,16 +2,7 @@ FROM rust:1.94-bookworm AS builder
WORKDIR /app
COPY . .
# compliance-agent (a workspace member) depends on the private tramiton-core git
# repo, so the workspace resolve needs it even to build the mcp binary.
# Authenticate the fetch with a PAT passed as a BuildKit secret.
RUN --mount=type=secret,id=tramiton_token \
if [ -s /run/secrets/tramiton_token ]; then \
git config --global \
url."https://sharang:$(cat /run/secrets/tramiton_token)@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"; \
fi && \
CARGO_NET_GIT_FETCH_WITH_CLI=true cargo build --release -p compliance-mcp
RUN cargo build --release -p compliance-mcp
FROM debian:bookworm-slim
RUN apt-get update && apt-get install -y ca-certificates libssl3 && rm -rf /var/lib/apt/lists/*
-10
View File
@@ -10,16 +10,6 @@ workspace = true
compliance-core = { workspace = true, features = ["mongodb", "telemetry", "axum"] }
compliance-graph = { path = "../compliance-graph" }
compliance-dast = { path = "../compliance-dast" }
# Native firmware build/target detection for bare-metal & RTOS artifacts.
# Same-company IP, used directly (not via CLI) so the whole tramiton suite is
# available to the onboarding classifier. NOTE: CI must be able to fetch this
# private repo (see the git-auth step in .gitea/workflows/ci.yml).
tramiton-core = { git = "ssh://git@gitea.meghsakha.com:22222/sharang/tramiton.git", tag = "v0.4.1" }
# tramiton-repro drives the reproducible build (NixBackend seal_and_build) that
# yields a sealed lock; `libraries_from_inputs` is the analysis-only fallback.
tramiton-repro = { git = "ssh://git@gitea.meghsakha.com:22222/sharang/tramiton.git", tag = "v0.4.1" }
# tramiton-sbom renders the bill of materials from a sealed lock (+ binary SCA).
tramiton-sbom = { git = "ssh://git@gitea.meghsakha.com:22222/sharang/tramiton.git", tag = "v0.4.1" }
serde = { workspace = true }
serde_json = { workspace = true }
tokio = { workspace = true }
+1 -25
View File
@@ -63,31 +63,7 @@ impl ComplianceAgent {
let db = self.db_pool.for_tenant_id(tenant_id).await?;
let orchestrator =
PipelineOrchestrator::new(self.config.clone(), db, self.llm.clone(), self.http.clone());
if self.config.unified_pipeline {
orchestrator.run_target(repo_id, trigger).await
} else {
orchestrator.run(repo_id, trigger).await
}
}
/// Run a scan for an onboarded target through the unified pipeline,
/// unconditionally.
///
/// Unlike [`Self::run_scan`], this does *not* consult the
/// `unified_pipeline` transition flag: the caller (the `/targets/{id}/scan`
/// endpoint) operates on `onboarded_targets` by construction, so it must
/// always dispatch to `run_target` regardless of how the legacy paths
/// (scheduler, webhooks, `/repositories/{id}/scan`) are configured.
pub async fn run_target_scan(
&self,
tenant_id: &str,
target_id: &str,
trigger: compliance_core::models::ScanTrigger,
) -> Result<(), crate::error::AgentError> {
let db = self.db_pool.for_tenant_id(tenant_id).await?;
let orchestrator =
PipelineOrchestrator::new(self.config.clone(), db, self.llm.clone(), self.http.clone());
orchestrator.run_target(target_id, trigger).await
orchestrator.run(repo_id, trigger).await
}
/// Run a PR review: scan the diff and post review comments.
-1
View File
@@ -9,7 +9,6 @@ pub mod help_chat;
pub mod issues;
pub mod mcp_tokens;
pub mod notifications;
pub mod onboarding;
pub mod pentest_handlers;
pub use pentest_handlers as pentest;
pub mod repos;
@@ -1,393 +0,0 @@
//! Onboarding API — CRUD for unified targets, artifact add, classification, and
//! the scan-applicability matrix. The wizard (and future integrations) drive
//! onboarding through these endpoints.
use std::collections::HashMap;
use std::sync::Arc;
use axum::extract::{Extension, Path, Query};
use axum::http::StatusCode;
use axum::Json;
use mongodb::bson::{doc, oid::ObjectId, to_bson};
use serde::{Deserialize, Serialize};
use compliance_core::models::{
Artifact, ArtifactKind, ComplianceProfile, OnboardedTarget, PlcFormat, TargetScanConfig,
TargetType,
};
use compliance_core::scan_matrix::{applicable_scans, supports_pentest};
use compliance_core::tenant_ctx::TenantCtx;
use crate::agent::ComplianceAgent;
use crate::classify::{classify_target, MockFirmwareDetector};
use super::dto::tenant_db;
use super::{collect_cursor_async, ApiResponse, PaginationParams};
type AgentExt = Extension<Arc<ComplianceAgent>>;
/// A client-supplied artifact spec. The server builds the [`Artifact`] (and its
/// id) from it, so clients never set internal fields.
#[derive(Deserialize)]
pub struct ArtifactInput {
pub kind: ArtifactKind,
pub source_ref: String,
#[serde(default)]
pub branch: Option<String>,
#[serde(default)]
pub plc_format: Option<PlcFormat>,
}
impl ArtifactInput {
fn build(&self) -> Artifact {
let s = self.source_ref.clone();
match self.kind {
ArtifactKind::GitRepo => {
Artifact::git_repo(s, self.branch.clone().unwrap_or_else(|| "main".to_string()))
}
ArtifactKind::LiveUrl => Artifact::live_url(s),
ArtifactKind::FirmwareImage => Artifact::firmware_image(s),
ArtifactKind::SourceArchive => Artifact::source_archive(s),
ArtifactKind::MobilePackage => Artifact::mobile_package(s),
ArtifactKind::ContainerImage => Artifact::container_image(s),
ArtifactKind::PlcProject => {
Artifact::plc_project(s, self.plc_format.unwrap_or(PlcFormat::PlcopenXml))
}
ArtifactKind::PlaintextDescription => Artifact::plaintext(s),
}
}
}
#[derive(Deserialize)]
pub struct CreateTargetRequest {
pub name: String,
pub target_type: TargetType,
#[serde(default)]
pub description: Option<String>,
#[serde(default)]
pub artifacts: Vec<ArtifactInput>,
}
#[derive(Deserialize)]
pub struct UpdateTargetRequest {
pub name: Option<String>,
pub target_type: Option<TargetType>,
pub scan_config: Option<TargetScanConfig>,
pub compliance_profile: Option<ComplianceProfile>,
pub scan_schedule: Option<String>,
}
/// One applicable-scan option, serialized for the wizard.
#[derive(Serialize)]
pub struct ScanOptionDto {
pub scan: String,
pub default_on: bool,
pub rationale: String,
pub required_artifact: Option<String>,
pub blocked_reason: Option<String>,
}
#[derive(Serialize)]
pub struct ApplicableScansResponse {
pub scans: Vec<ScanOptionDto>,
pub pentest_supported: bool,
}
fn parse_oid(id: &str) -> Result<ObjectId, StatusCode> {
ObjectId::parse_str(id).map_err(|_| StatusCode::BAD_REQUEST)
}
/// GET /api/v1/targets — list onboarded targets (paginated).
#[tracing::instrument(skip_all)]
pub async fn list_targets(
Extension(agent): AgentExt,
tenant: TenantCtx,
Query(params): Query<PaginationParams>,
) -> Result<Json<ApiResponse<Vec<OnboardedTarget>>>, StatusCode> {
let db = tenant_db(&agent, &tenant).await?;
let skip = (params.page.saturating_sub(1)) * params.limit as u64;
let total = db
.onboarded_targets()
.count_documents(doc! {})
.await
.unwrap_or(0);
let targets = match db
.onboarded_targets()
.find(doc! {})
.skip(skip)
.limit(params.limit)
.await
{
Ok(cursor) => collect_cursor_async(cursor).await,
Err(e) => {
tracing::warn!("Failed to fetch onboarded targets: {e}");
Vec::new()
}
};
Ok(Json(ApiResponse {
data: targets,
total: Some(total),
page: Some(params.page),
}))
}
/// POST /api/v1/targets — create an onboarded target.
#[tracing::instrument(skip_all)]
pub async fn create_target(
Extension(agent): AgentExt,
tenant: TenantCtx,
Json(req): Json<CreateTargetRequest>,
) -> Result<Json<ApiResponse<OnboardedTarget>>, StatusCode> {
let mut target = OnboardedTarget::new(req.name, req.target_type);
target.description = req.description;
target.artifacts = req.artifacts.iter().map(ArtifactInput::build).collect();
let db = tenant_db(&agent, &tenant).await?;
let res = db
.onboarded_targets()
.insert_one(&target)
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
target.id = res.inserted_id.as_object_id();
Ok(Json(ApiResponse {
data: target,
total: None,
page: None,
}))
}
/// GET /api/v1/targets/{id} — fetch one target.
#[tracing::instrument(skip_all, fields(target_id = %id))]
pub async fn get_target(
Extension(agent): AgentExt,
tenant: TenantCtx,
Path(id): Path<String>,
) -> Result<Json<ApiResponse<OnboardedTarget>>, StatusCode> {
let oid = parse_oid(&id)?;
let db = tenant_db(&agent, &tenant).await?;
let target = db
.onboarded_targets()
.find_one(doc! { "_id": oid })
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?
.ok_or(StatusCode::NOT_FOUND)?;
Ok(Json(ApiResponse {
data: target,
total: None,
page: None,
}))
}
/// PATCH /api/v1/targets/{id} — update mutable fields.
#[tracing::instrument(skip_all, fields(target_id = %id))]
pub async fn update_target(
Extension(agent): AgentExt,
tenant: TenantCtx,
Path(id): Path<String>,
Json(req): Json<UpdateTargetRequest>,
) -> Result<Json<ApiResponse<OnboardedTarget>>, StatusCode> {
let oid = parse_oid(&id)?;
let db = tenant_db(&agent, &tenant).await?;
let mut set = doc! { "updated_at": mongodb::bson::DateTime::now() };
if let Some(name) = req.name {
set.insert("name", name);
}
if let Some(tt) = req.target_type {
set.insert(
"target_type",
to_bson(&tt).map_err(|_| StatusCode::BAD_REQUEST)?,
);
}
if let Some(sc) = req.scan_config {
set.insert(
"scan_config",
to_bson(&sc).map_err(|_| StatusCode::BAD_REQUEST)?,
);
}
if let Some(cp) = req.compliance_profile {
set.insert(
"compliance_profile",
to_bson(&cp).map_err(|_| StatusCode::BAD_REQUEST)?,
);
}
if let Some(ss) = req.scan_schedule {
set.insert("scan_schedule", ss);
}
db.onboarded_targets()
.update_one(doc! { "_id": oid }, doc! { "$set": set })
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
get_target(Extension(agent), tenant, Path(id)).await
}
/// DELETE /api/v1/targets/{id} — remove the target and its findings/scans.
#[tracing::instrument(skip_all, fields(target_id = %id))]
pub async fn delete_target(
Extension(agent): AgentExt,
tenant: TenantCtx,
Path(id): Path<String>,
) -> Result<Json<serde_json::Value>, StatusCode> {
let oid = parse_oid(&id)?;
let db = tenant_db(&agent, &tenant).await?;
db.onboarded_targets()
.delete_one(doc! { "_id": oid })
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
// Cascade the collections keyed by repo_id == target id (best-effort).
let by_repo = doc! { "repo_id": &id };
let _ = db.findings().delete_many(by_repo.clone()).await;
let _ = db.scan_runs().delete_many(by_repo.clone()).await;
let _ = db.sbom_entries().delete_many(by_repo.clone()).await;
let _ = db.cve_alerts().delete_many(by_repo).await;
Ok(Json(serde_json::json!({ "status": "deleted" })))
}
/// POST /api/v1/targets/{id}/artifacts — attach an artifact (by reference).
#[tracing::instrument(skip_all, fields(target_id = %id))]
pub async fn add_artifact(
Extension(agent): AgentExt,
tenant: TenantCtx,
Path(id): Path<String>,
Json(input): Json<ArtifactInput>,
) -> Result<Json<ApiResponse<OnboardedTarget>>, StatusCode> {
let oid = parse_oid(&id)?;
let db = tenant_db(&agent, &tenant).await?;
let artifact = to_bson(&input.build()).map_err(|_| StatusCode::BAD_REQUEST)?;
db.onboarded_targets()
.update_one(
doc! { "_id": oid },
doc! { "$push": { "artifacts": artifact }, "$set": { "updated_at": mongodb::bson::DateTime::now() } },
)
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
get_target(Extension(agent), tenant, Path(id)).await
}
/// GET /api/v1/targets/{id}/applicable-scans — the scan-applicability matrix.
#[tracing::instrument(skip_all, fields(target_id = %id))]
pub async fn applicable_scans_for_target(
Extension(agent): AgentExt,
tenant: TenantCtx,
Path(id): Path<String>,
) -> Result<Json<ApiResponse<ApplicableScansResponse>>, StatusCode> {
let oid = parse_oid(&id)?;
let db = tenant_db(&agent, &tenant).await?;
let target = db
.onboarded_targets()
.find_one(doc! { "_id": oid })
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?
.ok_or(StatusCode::NOT_FOUND)?;
let scans = applicable_scans(&target)
.into_iter()
.map(|o| ScanOptionDto {
scan: o.scan.to_string(),
default_on: o.default_on,
rationale: o.rationale,
required_artifact: o.required_artifact.map(|k| k.to_string()),
blocked_reason: o.blocked_reason,
})
.collect();
Ok(Json(ApiResponse {
data: ApplicableScansResponse {
scans,
pentest_supported: supports_pentest(target.target_type),
},
total: None,
page: None,
}))
}
/// POST /api/v1/targets/{id}/detect — classify the target from its artifacts.
///
/// This is the lightweight pass: it classifies from artifact kinds without
/// ingesting (cloning) sources, so it returns immediately. Deep detection (after
/// ingest, with tramiton firmware analysis) is a follow-up background step.
#[tracing::instrument(skip_all, fields(target_id = %id))]
pub async fn detect_target(
Extension(agent): AgentExt,
tenant: TenantCtx,
Path(id): Path<String>,
) -> Result<Json<ApiResponse<OnboardedTarget>>, StatusCode> {
let oid = parse_oid(&id)?;
let db = tenant_db(&agent, &tenant).await?;
let mut target = db
.onboarded_targets()
.find_one(doc! { "_id": oid })
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?
.ok_or(StatusCode::NOT_FOUND)?;
// No ingested working paths here → kind-based classification only; the mock
// firmware detector is never invoked (no firmware working path present).
let empty = HashMap::new();
let detector = MockFirmwareDetector { detection: None };
let classification = classify_target(&target, &empty, &detector)
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
let classification_bson =
to_bson(&classification).map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
db.onboarded_targets()
.update_one(
doc! { "_id": oid },
doc! { "$set": { "classification": classification_bson, "updated_at": mongodb::bson::DateTime::now() } },
)
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?;
target.classification = Some(classification);
Ok(Json(ApiResponse {
data: target,
total: None,
page: None,
}))
}
/// POST /api/v1/targets/{id}/scan — trigger a scan for the target.
///
/// Dispatches to the unified pipeline when `UNIFIED_PIPELINE` is set (else the
/// legacy path). Runs in the background and returns immediately.
#[tracing::instrument(skip_all, fields(target_id = %id))]
pub async fn trigger_target_scan(
Extension(agent): AgentExt,
tenant: TenantCtx,
Path(id): Path<String>,
) -> Result<Json<serde_json::Value>, StatusCode> {
let oid = parse_oid(&id)?;
let db = tenant_db(&agent, &tenant).await?;
// 404 if the target doesn't exist for this tenant.
if db
.onboarded_targets()
.find_one(doc! { "_id": oid })
.await
.map_err(|_| StatusCode::INTERNAL_SERVER_ERROR)?
.is_none()
{
return Err(StatusCode::NOT_FOUND);
}
let agent_clone = (*agent).clone();
let tenant_id = tenant.0.tenant_id.clone();
tokio::spawn(async move {
// Always the unified target pipeline — this endpoint is about an
// onboarded target by construction, independent of the global
// `unified_pipeline` transition flag used by the legacy paths.
if let Err(e) = agent_clone
.run_target_scan(
&tenant_id,
&id,
compliance_core::models::ScanTrigger::Manual,
)
.await
{
tracing::error!("Manual target scan failed for {id}: {e}");
}
});
Ok(Json(serde_json::json!({ "status": "scan_triggered" })))
}
-27
View File
@@ -25,33 +25,6 @@ pub fn build_router() -> Router {
"/api/v1/repositories/{id}/webhook-config",
get(handlers::get_webhook_config),
)
// Unified onboarding targets (#131).
.route(
"/api/v1/targets",
get(handlers::onboarding::list_targets).post(handlers::onboarding::create_target),
)
.route(
"/api/v1/targets/{id}",
get(handlers::onboarding::get_target)
.patch(handlers::onboarding::update_target)
.delete(handlers::onboarding::delete_target),
)
.route(
"/api/v1/targets/{id}/artifacts",
post(handlers::onboarding::add_artifact),
)
.route(
"/api/v1/targets/{id}/applicable-scans",
get(handlers::onboarding::applicable_scans_for_target),
)
.route(
"/api/v1/targets/{id}/detect",
post(handlers::onboarding::detect_target),
)
.route(
"/api/v1/targets/{id}/scan",
post(handlers::onboarding::trigger_target_scan),
)
.route("/api/v1/findings", get(handlers::list_findings))
.route("/api/v1/findings/{id}", get(handlers::get_finding))
.route(
-217
View File
@@ -1,217 +0,0 @@
//! Firmware classification via tramiton.
//!
//! tramiton is the company's firmware build/repro engine; we do not re-implement
//! its detection. We depend on `tramiton-core` directly (same-company IP) and run
//! its provider analysis in-process behind a [`FirmwareDetector`] port, mapping
//! tramiton's `BuildPlan` onto a [`TargetType`]. A deterministic
//! [`MockFirmwareDetector`] backs the tests so CI unit tests need neither the
//! tramiton sources nor a real firmware tree.
use std::path::Path;
use compliance_core::error::CoreError;
use compliance_core::models::{DetectedFact, TargetType};
use compliance_core::traits::ClassifierVerdict;
/// A minimal firmware-detection summary, mapped from tramiton's `BuildPlan`.
/// Kept small and tramiton-independent so the classifier and the test mock don't
/// need to construct a full tramiton plan.
#[derive(Debug, Clone, Default)]
pub struct FirmwareDetection {
/// The detecting provider (e.g. `zephyr`, `cmake`, `source-archaeology`).
pub provider: String,
/// Detection confidence: `low` | `medium` | `high`.
pub confidence: String,
/// Build-system label (e.g. `Zephyr`, `ESP-IDF`, `CMake`).
pub build_system: String,
/// Framework, when known (`zephyr`, `esp-idf`, `bare-metal`, ...).
pub framework: Option<String>,
/// Target board / MCU / arch.
pub target: FirmwareTarget,
/// Unresolved gaps in the plan.
pub gaps: Vec<String>,
}
/// The detected firmware target (board / MCU / arch).
#[derive(Debug, Clone, Default)]
pub struct FirmwareTarget {
/// Board name.
pub board: Option<String>,
/// MCU part.
pub mcu: Option<String>,
/// Architecture.
pub arch: Option<String>,
}
/// A source of tramiton firmware detection.
#[allow(async_fn_in_trait)]
pub trait FirmwareDetector: Send + Sync {
/// Run detection over a path, returning a firmware detection if tramiton
/// could form a build plan.
async fn detect(&self, path: &Path) -> Result<Option<FirmwareDetection>, CoreError>;
}
/// Uses `tramiton-core` in-process. The analysis is blocking (filesystem walk),
/// so it runs on a blocking thread to avoid stalling the async runtime. A path
/// with no recognizable build system yields `Ok(None)`.
pub struct TramitonNative;
impl FirmwareDetector for TramitonNative {
async fn detect(&self, path: &Path) -> Result<Option<FirmwareDetection>, CoreError> {
let path = path.to_path_buf();
let plan = tokio::task::spawn_blocking(move || {
let repo = tramiton_core::Repo::new(&path);
tramiton_core::provider::analyze(&repo)
})
.await
.map_err(|e| CoreError::Other(format!("tramiton detect task join error: {e}")))?
.map_err(|e| CoreError::Other(format!("tramiton analyze error: {e}")))?;
Ok(plan.map(|bp| detection_from_build_plan(&bp)))
}
}
/// Map tramiton's `BuildPlan` onto our minimal detection summary.
fn detection_from_build_plan(bp: &tramiton_core::BuildPlan) -> FirmwareDetection {
FirmwareDetection {
provider: bp.provider.clone(),
confidence: bp.confidence.to_string(),
build_system: bp.build_system.label().to_string(),
framework: bp.framework.clone(),
target: FirmwareTarget {
board: bp.target.board.clone(),
mcu: bp.target.mcu.clone(),
arch: bp.target.arch.clone(),
},
gaps: bp.gaps.clone(),
}
}
/// Map a firmware detection to a target type. Framework/build-system signals
/// distinguish RTOS from bare-metal from Yocto.
pub fn detection_to_target_type(det: &FirmwareDetection) -> TargetType {
let framework = det.framework.as_deref().unwrap_or("").to_lowercase();
let build_system = det.build_system.to_lowercase();
let signal = format!("{framework} {build_system} {}", det.provider.to_lowercase());
const RTOS: [&str; 6] = ["zephyr", "esp-idf", "freertos", "nuttx", "riot", "chibios"];
if signal.contains("bitbake") || signal.contains("yocto") || signal.contains("openembedded") {
TargetType::EmbeddedLinuxYocto
} else if RTOS.iter().any(|k| signal.contains(k)) {
TargetType::FirmwareRtos
} else {
TargetType::FirmwareBareMetal
}
}
/// Map tramiton's confidence label to a `[0,1]` score.
fn confidence_score(label: &str) -> f32 {
match label.to_lowercase().as_str() {
"high" => 0.9,
"medium" => 0.6,
"low" => 0.3,
_ => 0.4,
}
}
/// Turn a firmware detection into a classifier verdict, carrying the MCU / board
/// / build-system as facts.
pub fn detection_to_verdict(det: &FirmwareDetection) -> ClassifierVerdict {
let target_type = detection_to_target_type(det);
let mut facts = vec![DetectedFact::new(
"build_system",
det.build_system.clone(),
"tramiton",
)];
if let Some(fw) = &det.framework {
facts.push(DetectedFact::new("framework", fw.clone(), "tramiton"));
}
if let Some(mcu) = &det.target.mcu {
facts.push(DetectedFact::new("mcu", mcu.clone(), "tramiton"));
}
if let Some(board) = &det.target.board {
facts.push(DetectedFact::new("board", board.clone(), "tramiton"));
}
if let Some(arch) = &det.target.arch {
facts.push(DetectedFact::new("arch", arch.clone(), "tramiton"));
}
ClassifierVerdict {
target_type,
confidence: confidence_score(&det.confidence),
facts,
rationale: format!(
"tramiton detected build system '{}'{}",
det.build_system,
det.framework
.as_ref()
.map(|f| format!(" (framework {f})"))
.unwrap_or_default()
),
}
}
/// A deterministic [`FirmwareDetector`] for tests — returns a preset detection.
pub struct MockFirmwareDetector {
/// The detection to return (or `None` for "no detection").
pub detection: Option<FirmwareDetection>,
}
impl FirmwareDetector for MockFirmwareDetector {
async fn detect(&self, _path: &Path) -> Result<Option<FirmwareDetection>, CoreError> {
Ok(self.detection.clone())
}
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
fn detection(build_system: &str, framework: Option<&str>) -> FirmwareDetection {
FirmwareDetection {
provider: build_system.to_string(),
confidence: "high".to_string(),
build_system: build_system.to_string(),
framework: framework.map(|s| s.to_string()),
target: FirmwareTarget {
mcu: Some("stm32f429".to_string()),
..Default::default()
},
gaps: Vec::new(),
}
}
#[test]
fn zephyr_maps_to_rtos() {
assert_eq!(
detection_to_target_type(&detection("zephyr", Some("zephyr"))),
TargetType::FirmwareRtos
);
}
#[test]
fn bare_cmake_maps_to_bare_metal() {
assert_eq!(
detection_to_target_type(&detection("cmake", Some("bare-metal"))),
TargetType::FirmwareBareMetal
);
}
#[test]
fn bitbake_maps_to_yocto() {
assert_eq!(
detection_to_target_type(&detection("bitbake", None)),
TargetType::EmbeddedLinuxYocto
);
}
#[test]
fn verdict_carries_mcu_fact_and_confidence() {
let v = detection_to_verdict(&detection("esp-idf", Some("esp-idf")));
assert_eq!(v.target_type, TargetType::FirmwareRtos);
assert!((v.confidence - 0.9).abs() < f32::EPSILON);
assert!(v
.facts
.iter()
.any(|f| f.key == "mcu" && f.value == "stm32f429"));
}
}
-357
View File
@@ -1,357 +0,0 @@
//! Heuristic target-type classification from artifact kinds and source markers.
//!
//! Complements the tramiton firmware detector: this handles web / backend /
//! mobile / desktop / PLC by sniffing manifest files and file extensions in the
//! ingested code trees, plus strong priors from the artifact kinds themselves
//! (a PLC-project artifact is a PLC target; an `.ipa` is an iOS app).
use std::collections::HashSet;
use std::fs;
use std::path::Path;
use compliance_core::error::CoreError;
use compliance_core::models::{ArtifactKind, DetectedFact, TargetType};
use compliance_core::traits::{ClassificationInput, ClassifierVerdict, TargetClassifier};
/// Max directory depth scanned for marker files.
const SCAN_DEPTH: usize = 2;
/// Markers collected from a code tree.
#[derive(Default)]
struct Markers {
files: HashSet<String>,
dirs: HashSet<String>,
exts: HashSet<String>,
}
impl Markers {
fn has_file(&self, name: &str) -> bool {
self.files.contains(name)
}
fn has_ext(&self, ext: &str) -> bool {
self.exts.contains(ext)
}
fn any_dir_ends_with(&self, suffix: &str) -> bool {
self.dirs.iter().any(|d| d.ends_with(suffix))
}
}
/// Recursively collect marker file/dir/extension names up to [`SCAN_DEPTH`].
fn collect_markers(root: &Path) -> Markers {
let mut m = Markers::default();
scan_dir(root, 0, &mut m);
m
}
fn scan_dir(dir: &Path, depth: usize, m: &mut Markers) {
let Ok(entries) = fs::read_dir(dir) else {
return;
};
for entry in entries.flatten() {
let path = entry.path();
let name = entry.file_name().to_string_lossy().to_lowercase();
if path.is_dir() {
m.dirs.insert(name);
if depth < SCAN_DEPTH {
scan_dir(&path, depth + 1, m);
}
} else {
if let Some(ext) = path.extension() {
m.exts.insert(ext.to_string_lossy().to_lowercase());
}
m.files.insert(name);
}
}
}
/// Whether a `package.json` at `root` looks like a front-end app.
fn package_json_is_frontend(root: &Path) -> bool {
let Ok(content) = fs::read_to_string(root.join("package.json")) else {
return false;
};
let c = content.to_lowercase();
["react", "next", "vue", "@angular", "svelte", "vite"]
.iter()
.any(|f| c.contains(f))
}
/// The heuristic classifier: artifact-kind priors + source-tree markers.
pub struct HeuristicClassifier;
impl HeuristicClassifier {
/// Verdicts from the artifact kinds alone (no filesystem needed).
fn kind_priors(&self, input: &ClassificationInput<'_>) -> Vec<ClassifierVerdict> {
let mut out = Vec::new();
for a in input.artifacts {
let lower = a.source_ref.to_lowercase();
match a.kind {
ArtifactKind::PlcProject => out.push(verdict(
TargetType::PlcSps,
0.85,
"PLC project artifact",
vec![],
)),
ArtifactKind::MobilePackage => {
let (tt, why) = if lower.ends_with(".ipa") {
(TargetType::IosApp, "iOS package (.ipa)")
} else {
(TargetType::AndroidApp, "Android package (.apk/.aab)")
};
out.push(verdict(tt, 0.85, why, vec![]));
}
ArtifactKind::ContainerImage => out.push(verdict(
TargetType::BackendService,
0.4,
"container image",
vec![],
)),
ArtifactKind::FirmwareImage => out.push(verdict(
TargetType::FirmwareBareMetal,
0.35,
"firmware image (pending tramiton detection)",
vec![],
)),
ArtifactKind::LiveUrl if input.artifacts.len() == 1 => {
out.push(verdict(TargetType::WebApp, 0.3, "live URL only", vec![]))
}
_ => {}
}
}
out
}
/// Verdicts from scanning the ingested code trees for manifest markers.
fn source_verdicts(&self, input: &ClassificationInput<'_>) -> Vec<ClassifierVerdict> {
let mut out = Vec::new();
for a in input.artifacts {
if !matches!(a.kind, ArtifactKind::GitRepo | ArtifactKind::SourceArchive) {
continue;
}
let Some(path) = input.working_paths.get(&a.id) else {
continue;
};
let m = collect_markers(path);
// Mobile (checked first — strongest signal).
if m.has_file("androidmanifest.xml") || m.has_ext("apk") || m.has_ext("aab") {
out.push(verdict(
TargetType::AndroidApp,
0.8,
"Android manifest / gradle",
facts_lang("kotlin/java"),
));
}
if m.any_dir_ends_with(".xcodeproj")
|| m.has_file("info.plist")
|| m.has_file("podfile")
|| m.has_ext("ipa")
{
out.push(verdict(
TargetType::IosApp,
0.8,
"Xcode project / Info.plist",
facts_lang("swift/objc"),
));
}
// Desktop.
if m.has_ext("sln")
|| m.has_ext("csproj")
|| m.has_ext("vcxproj")
|| m.has_ext("desktop")
{
out.push(verdict(
TargetType::DesktopApp,
0.7,
"desktop project files",
facts_lang("dotnet/native"),
));
}
// PLC.
if m.has_ext("st") {
out.push(verdict(
TargetType::PlcSps,
0.8,
"Structured Text sources",
facts_lang("iec-61131-3"),
));
}
// Web vs backend from package.json.
if m.has_file("package.json") {
if package_json_is_frontend(path) {
out.push(verdict(
TargetType::WebApp,
0.65,
"package.json with a front-end framework",
facts_lang("javascript"),
));
} else {
out.push(verdict(
TargetType::BackendService,
0.55,
"package.json (no front-end framework)",
facts_lang("javascript"),
));
}
}
// Backend languages.
for (file, lang) in [
("cargo.toml", "rust"),
("go.mod", "go"),
("pom.xml", "java"),
("requirements.txt", "python"),
("pyproject.toml", "python"),
] {
if m.has_file(file) {
out.push(verdict(
TargetType::BackendService,
0.6,
"backend build manifest",
facts_lang(lang),
));
}
}
// Container-only.
if m.has_file("dockerfile") && out.is_empty() {
out.push(verdict(
TargetType::BackendService,
0.4,
"Dockerfile",
facts_lang("container"),
));
}
}
out
}
}
impl TargetClassifier for HeuristicClassifier {
fn name(&self) -> &str {
"heuristic"
}
async fn classify(
&self,
input: &ClassificationInput<'_>,
) -> Result<Vec<ClassifierVerdict>, CoreError> {
let mut out = self.kind_priors(input);
out.extend(self.source_verdicts(input));
Ok(out)
}
}
fn verdict(
target_type: TargetType,
confidence: f32,
rationale: &str,
facts: Vec<DetectedFact>,
) -> ClassifierVerdict {
ClassifierVerdict {
target_type,
confidence,
facts,
rationale: rationale.to_string(),
}
}
fn facts_lang(lang: &str) -> Vec<DetectedFact> {
vec![DetectedFact::new("language", lang, "heuristic")]
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
use compliance_core::models::Artifact;
use std::collections::HashMap;
use std::path::PathBuf;
struct Scratch(PathBuf);
impl Scratch {
fn new() -> Self {
let p = std::env::temp_dir().join(format!("cs-classify-{}", uuid::Uuid::new_v4()));
fs::create_dir_all(&p).expect("mkdir");
Self(p)
}
}
impl Drop for Scratch {
fn drop(&mut self) {
let _ = fs::remove_dir_all(&self.0);
}
}
async fn classify_tree(setup: impl FnOnce(&Path)) -> Vec<ClassifierVerdict> {
let scratch = Scratch::new();
setup(&scratch.0);
let artifact = Artifact::git_repo("https://git/x", "main");
let mut wp = HashMap::new();
wp.insert(artifact.id.clone(), scratch.0.clone());
let artifacts = vec![artifact];
let input = ClassificationInput {
artifacts: &artifacts,
working_paths: &wp,
description: None,
};
HeuristicClassifier
.classify(&input)
.await
.expect("classify")
}
#[tokio::test]
async fn frontend_package_json_is_webapp() {
let v = classify_tree(|root| {
fs::write(
root.join("package.json"),
r#"{"dependencies":{"react":"18"}}"#,
)
.unwrap();
})
.await;
assert!(v.iter().any(|x| x.target_type == TargetType::WebApp));
}
#[tokio::test]
async fn cargo_toml_is_backend() {
let v = classify_tree(|root| {
fs::write(root.join("Cargo.toml"), "[package]\nname='x'").unwrap();
})
.await;
assert!(v
.iter()
.any(|x| x.target_type == TargetType::BackendService));
}
#[tokio::test]
async fn android_manifest_is_android() {
let v = classify_tree(|root| {
fs::write(root.join("AndroidManifest.xml"), "<manifest/>").unwrap();
})
.await;
assert!(v.iter().any(|x| x.target_type == TargetType::AndroidApp));
}
#[tokio::test]
async fn structured_text_is_plc() {
let v = classify_tree(|root| {
fs::write(root.join("main.st"), "PROGRAM main END_PROGRAM").unwrap();
})
.await;
assert!(v.iter().any(|x| x.target_type == TargetType::PlcSps));
}
#[tokio::test]
async fn ipa_artifact_prior_is_ios() {
let artifacts = vec![Artifact::mobile_package("app.ipa")];
let wp = HashMap::new();
let input = ClassificationInput {
artifacts: &artifacts,
working_paths: &wp,
description: None,
};
let v = HeuristicClassifier
.classify(&input)
.await
.expect("classify");
assert!(v.iter().any(|x| x.target_type == TargetType::IosApp));
}
}
-226
View File
@@ -1,226 +0,0 @@
//! Target classification.
//!
//! Runs the classifier registry over a target's artifacts and their ingested
//! working paths, then merges and ranks the verdicts into a [`Classification`].
//! The registry is the heuristic classifier (artifact kinds + source markers)
//! plus the tramiton firmware detector (behind a [`FirmwareDetector`] port).
mod firmware;
mod language;
pub use firmware::{
FirmwareDetection, FirmwareDetector, FirmwareTarget, MockFirmwareDetector, TramitonNative,
};
pub use language::HeuristicClassifier;
use std::collections::HashMap;
use std::path::PathBuf;
use compliance_core::error::CoreError;
use compliance_core::models::{
ArtifactKind, Classification, DetectedFact, OnboardedTarget, TargetType, TargetTypeCandidate,
};
use compliance_core::traits::{ClassificationInput, ClassifierVerdict, TargetClassifier};
use firmware::detection_to_verdict;
/// Classify a target from its artifacts and their ingested working paths, using
/// the heuristic classifier plus the tramiton firmware detector. Verdicts are
/// merged (max confidence per target type) and ranked into a [`Classification`].
pub async fn classify_target<D: FirmwareDetector>(
target: &OnboardedTarget,
working_paths: &HashMap<String, PathBuf>,
firmware_detector: &D,
) -> Result<Classification, CoreError> {
let input = ClassificationInput {
artifacts: &target.artifacts,
working_paths,
description: target.description.as_deref(),
};
let mut verdicts = Vec::new();
let mut detected_by = Vec::new();
let heuristic = HeuristicClassifier.classify(&input).await?;
if !heuristic.is_empty() {
detected_by.push("heuristic".to_string());
}
verdicts.extend(heuristic);
// Tramiton firmware detection over firmware / code working paths.
let mut tramiton_used = false;
for artifact in &target.artifacts {
if !matches!(
artifact.kind,
ArtifactKind::FirmwareImage | ArtifactKind::GitRepo | ArtifactKind::SourceArchive
) {
continue;
}
let Some(path) = working_paths.get(&artifact.id) else {
continue;
};
if let Some(detection) = firmware_detector.detect(path).await? {
verdicts.push(detection_to_verdict(&detection));
tramiton_used = true;
}
}
if tramiton_used {
detected_by.push("tramiton".to_string());
}
Ok(rank(verdicts, detected_by, target.target_type))
}
/// Merge verdicts by target type (keeping the max confidence and its rationale),
/// dedupe facts, rank by descending confidence, and assemble a [`Classification`].
/// Falls back to the declared type when no verdict is produced.
fn rank(
verdicts: Vec<ClassifierVerdict>,
detected_by: Vec<String>,
fallback: TargetType,
) -> Classification {
let mut best: HashMap<TargetType, (f32, String)> = HashMap::new();
let mut facts: Vec<DetectedFact> = Vec::new();
for verdict in verdicts {
for fact in verdict.facts {
if !facts
.iter()
.any(|e| e.key == fact.key && e.value == fact.value)
{
facts.push(fact);
}
}
let entry = best
.entry(verdict.target_type)
.or_insert((0.0, String::new()));
if verdict.confidence > entry.0 {
*entry = (verdict.confidence, verdict.rationale);
}
}
let mut candidates: Vec<TargetTypeCandidate> = best
.into_iter()
.map(
|(target_type, (confidence, rationale))| TargetTypeCandidate {
target_type,
confidence,
rationale,
},
)
.collect();
// Descending confidence; ties broken by type name for deterministic ordering.
candidates.sort_by(|a, b| {
b.confidence
.partial_cmp(&a.confidence)
.unwrap_or(std::cmp::Ordering::Equal)
.then_with(|| a.target_type.to_string().cmp(&b.target_type.to_string()))
});
let suggested = candidates
.first()
.map(|c| c.target_type)
.unwrap_or(fallback);
Classification {
suggested,
candidates,
facts,
detected_by,
detected_at: chrono::Utc::now(),
confirmed: false,
}
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
use compliance_core::models::Artifact;
use std::fs;
use std::path::Path;
struct Scratch(PathBuf);
impl Scratch {
fn new() -> Self {
let p = std::env::temp_dir().join(format!("cs-classify-mod-{}", uuid::Uuid::new_v4()));
fs::create_dir_all(&p).expect("mkdir");
Self(p)
}
}
impl Drop for Scratch {
fn drop(&mut self) {
let _ = fs::remove_dir_all(&self.0);
}
}
fn no_firmware() -> MockFirmwareDetector {
MockFirmwareDetector { detection: None }
}
#[tokio::test]
async fn backend_repo_classifies_as_backend() {
let scratch = Scratch::new();
fs::write(scratch.0.join("go.mod"), "module x").unwrap();
let artifact = Artifact::git_repo("https://git/x", "main");
let mut wp = HashMap::new();
wp.insert(artifact.id.clone(), scratch.0.clone());
let mut target = OnboardedTarget::new("x".to_string(), TargetType::WebApp);
target.artifacts.push(artifact);
let c = classify_target(&target, &wp, &no_firmware())
.await
.expect("classify");
assert_eq!(c.suggested, TargetType::BackendService);
assert!(c.detected_by.contains(&"heuristic".to_string()));
assert!(!c.confirmed);
}
#[tokio::test]
async fn firmware_detector_verdict_ranks_top() {
let scratch = Scratch::new();
fs::write(scratch.0.join("fw.bin"), b"x").unwrap();
let artifact =
Artifact::firmware_image(scratch.0.join("fw.bin").to_string_lossy().to_string());
let mut wp = HashMap::new();
wp.insert(artifact.id.clone(), scratch.0.clone());
let mut target = OnboardedTarget::new("fw".to_string(), TargetType::FirmwareBareMetal);
target.artifacts.push(artifact);
let detector = MockFirmwareDetector {
detection: Some(FirmwareDetection {
provider: "zephyr".to_string(),
confidence: "high".to_string(),
build_system: "zephyr".to_string(),
framework: Some("zephyr".to_string()),
target: FirmwareTarget {
mcu: Some("nrf52840".to_string()),
..Default::default()
},
gaps: vec![],
}),
};
let c = classify_target(&target, &wp, &detector)
.await
.expect("classify");
// tramiton's high-confidence RTOS verdict beats the weak firmware prior.
assert_eq!(c.suggested, TargetType::FirmwareRtos);
assert!(c.detected_by.contains(&"tramiton".to_string()));
assert!(c.facts.iter().any(|f| f.key == "mcu"));
}
#[tokio::test]
async fn no_signal_falls_back_to_declared_type() {
let scratch = Scratch::new();
let _ = Path::new(&scratch.0);
let target = OnboardedTarget::new("empty".to_string(), TargetType::DesktopApp);
let wp = HashMap::new();
let c = classify_target(&target, &wp, &no_firmware())
.await
.expect("classify");
assert_eq!(c.suggested, TargetType::DesktopApp);
assert!(c.candidates.is_empty());
}
}
-8
View File
@@ -45,14 +45,6 @@ pub fn load_config() -> Result<AgentConfig, AgentError> {
.unwrap_or_else(|| "0 0 * * * *".to_string()),
git_clone_base_path: env_var_opt("GIT_CLONE_BASE_PATH")
.unwrap_or_else(|| "/tmp/compliance-scanner/repos".to_string()),
artifact_store_base_path: env_var_opt("ARTIFACT_STORE_BASE_PATH")
.unwrap_or_else(|| "/data/compliance-scanner/artifacts".to_string()),
// Defaults ON: the unified onboarded-target pipeline is now the primary
// path (no legacy `repositories` data in production). Set
// `UNIFIED_PIPELINE=0` to fall back to the legacy repository pipeline.
unified_pipeline: env_var_opt("UNIFIED_PIPELINE")
.map(|v| v == "1" || v.eq_ignore_ascii_case("true"))
.unwrap_or(true),
ssh_key_path: env_var_opt("SSH_KEY_PATH")
.unwrap_or_else(|| "/data/compliance-scanner/ssh/id_ed25519".to_string()),
keycloak_url: env_var_opt("KEYCLOAK_URL"),
-61
View File
@@ -179,23 +179,6 @@ impl DatabasePool {
.collect())
}
/// Tenant ids for every provisioned tenant database, derived by stripping
/// the `<prefix>_` from the database names. Skips the admin database
/// (`<prefix>__admin`). Hash-fallback names (very long tenant_ids) are lost
/// at the cluster level and cannot be recovered here — in practice tenant
/// ids are UUIDs and never hit that path. Used by the migration CLI's
/// `--all` mode.
pub async fn list_tenant_ids(&self) -> Result<Vec<String>, AgentError> {
let prefix = format!("{}_", self.db_prefix);
Ok(self
.list_tenant_db_names()
.await?
.into_iter()
.filter_map(|n| n.strip_prefix(&prefix).map(str::to_string))
.filter(|id| !id.starts_with('_'))
.collect())
}
/// Drop the database for a specific tenant. Used by GDPR delete
/// and tenant offboarding. Idempotent — dropping a non-existent
/// database is a no-op at the driver level.
@@ -445,36 +428,6 @@ impl Database {
)
.await?;
// onboarded_targets: multikey on artifact source ref (webhook + dedupe
// lookup). Non-unique — "one git URL per tenant" is enforced in the
// create handler, since a unique multikey index on an array field has
// null-collision caveats.
self.onboarded_targets()
.create_index(
IndexModel::builder()
.keys(doc! { "artifacts.source_ref": 1 })
.build(),
)
.await?;
// onboarded_targets: multikey on artifact kind
self.onboarded_targets()
.create_index(
IndexModel::builder()
.keys(doc! { "artifacts.kind": 1 })
.build(),
)
.await?;
// onboarded_targets: target_type filter
self.onboarded_targets()
.create_index(
IndexModel::builder()
.keys(doc! { "target_type": 1 })
.build(),
)
.await?;
tracing::info!("Database indexes ensured");
Ok(())
}
@@ -531,20 +484,6 @@ impl Database {
self.inner.collection("dast_targets")
}
/// The unified onboarding targets that replace `repositories` and
/// `dast_targets`. Ids are preserved from the legacy collections during
/// migration so downstream `repo_id` / `target_id` references keep resolving.
pub fn onboarded_targets(&self) -> Collection<OnboardedTarget> {
self.inner.collection("onboarded_targets")
}
/// A typed handle to an arbitrary collection by name. For bookkeeping
/// collections without a dedicated model (e.g. `schema_migrations`,
/// `onboarding_migration_log`).
pub fn collection_named<T: Send + Sync>(&self, name: &str) -> Collection<T> {
self.inner.collection(name)
}
pub fn dast_scan_runs(&self) -> Collection<DastScanRun> {
self.inner.collection("dast_scan_runs")
}
-154
View File
@@ -1,154 +0,0 @@
//! Content-addressed blob storage and archive extraction for ingest.
//!
//! Blobs are stored at `<base>/blobs/<sha[0:2]>/<sha>` and deduplicated by
//! digest; per-run working directories live under `<base>/work/`.
use std::fs::{self, File};
use std::io::{self, Read};
use std::path::{Path, PathBuf};
use sha2::{Digest, Sha256};
use crate::error::AgentError;
/// Read buffer size for streaming hashes/copies (64 KiB).
const BUF_LEN: usize = 64 * 1024;
/// Stream-hash a file with SHA-256, returning the lowercase-hex digest and the
/// byte length. Streams so large firmware images never load fully into memory.
pub fn hash_file(path: &Path) -> Result<(String, u64), AgentError> {
let mut file = File::open(path)?;
let mut hasher = Sha256::new();
let mut buf = [0u8; BUF_LEN];
let mut total: u64 = 0;
loop {
let n = file.read(&mut buf)?;
if n == 0 {
break;
}
hasher.update(&buf[..n]);
total += n as u64;
}
Ok((hex::encode(hasher.finalize()), total))
}
/// Copy `src` into the content-addressed blob store under `base`, returning the
/// stored path. Idempotent: an already-present blob is not rewritten.
pub fn store_file(base: &Path, src: &Path, sha: &str) -> Result<PathBuf, AgentError> {
if sha.len() < 2 {
return Err(AgentError::Other(format!("invalid content hash '{sha}'")));
}
let dir = base.join("blobs").join(&sha[0..2]);
fs::create_dir_all(&dir)?;
let dest = dir.join(sha);
if !dest.exists() {
fs::copy(src, &dest)?;
}
Ok(dest)
}
/// Extract a zip archive into `dest` (created if needed). `enclosed_name`
/// sanitizes each entry path, so this is safe against zip-slip traversal.
pub fn extract_zip(archive: &Path, dest: &Path) -> Result<(), AgentError> {
let file = File::open(archive)?;
let mut zip =
zip::ZipArchive::new(file).map_err(|e| AgentError::Other(format!("open zip: {e}")))?;
fs::create_dir_all(dest)?;
for i in 0..zip.len() {
let mut entry = zip
.by_index(i)
.map_err(|e| AgentError::Other(format!("read zip entry: {e}")))?;
// `enclosed_name` returns `None` for traversal-unsafe paths — skip them.
let Some(rel) = entry.enclosed_name() else {
continue;
};
let out = dest.join(rel);
if entry.is_dir() {
fs::create_dir_all(&out)?;
} else {
if let Some(parent) = out.parent() {
fs::create_dir_all(parent)?;
}
let mut outfile = File::create(&out)?;
io::copy(&mut entry, &mut outfile)?;
}
}
Ok(())
}
/// The working directory for one artifact of a target: `<base>/work/<target>/<artifact>`.
pub fn work_dir(base: &Path, target_id: &str, artifact_id: &str) -> PathBuf {
base.join("work").join(target_id).join(artifact_id)
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
/// A unique scratch directory, removed on drop.
struct Scratch(PathBuf);
impl Scratch {
fn new() -> Self {
let p = std::env::temp_dir().join(format!("cs-ingest-{}", uuid::Uuid::new_v4()));
fs::create_dir_all(&p).expect("mkdir scratch");
Self(p)
}
fn path(&self) -> &Path {
&self.0
}
}
impl Drop for Scratch {
fn drop(&mut self) {
let _ = fs::remove_dir_all(&self.0);
}
}
#[test]
fn hash_is_stable_and_reports_size() {
let dir = Scratch::new();
let f = dir.path().join("a.bin");
fs::write(&f, b"hello world").expect("write");
let (sha, size) = hash_file(&f).expect("hash");
assert_eq!(size, 11);
// Known SHA-256 of "hello world".
assert_eq!(
sha,
"b94d27b9934d3e08a52e52d7da7dabfac484efe37a5380ee9088f7ace2efcde9"
);
}
#[test]
fn store_is_content_addressed_and_idempotent() {
let base = Scratch::new();
let src = base.path().join("src.bin");
fs::write(&src, b"payload").expect("write");
let (sha, _) = hash_file(&src).expect("hash");
let p1 = store_file(base.path(), &src, &sha).expect("store");
let p2 = store_file(base.path(), &src, &sha).expect("store again");
assert_eq!(p1, p2);
assert!(p1.ends_with(&sha));
assert!(p1.starts_with(base.path().join("blobs").join(&sha[0..2])));
assert_eq!(fs::read(&p1).expect("read"), b"payload");
}
#[test]
fn extract_zip_writes_entries() {
let base = Scratch::new();
let archive = base.path().join("a.zip");
{
let file = File::create(&archive).expect("create");
let mut w = zip::ZipWriter::new(file);
let opts: zip::write::SimpleFileOptions = Default::default();
w.start_file("dir/hello.txt", opts).expect("start");
io::Write::write_all(&mut w, b"hi").expect("write");
w.finish().expect("finish");
}
let dest = base.path().join("out");
extract_zip(&archive, &dest).expect("extract");
assert_eq!(
fs::read_to_string(dest.join("dir/hello.txt")).expect("read"),
"hi"
);
}
}
-334
View File
@@ -1,334 +0,0 @@
//! Artifact ingest.
//!
//! Normalizes each [`Artifact`] on an [`OnboardedTarget`] into a local working
//! path plus recorded metadata (content hash, size, discovered facts) that the
//! classifier and scanners consume. Every blob is SHA-256 hashed — that digest
//! is also the reconciliation key against sibling products (a firmware sha256
//! matches tramiton's `Artifact.sha256`).
mod blob;
use std::collections::HashMap;
use std::path::{Path, PathBuf};
use compliance_core::models::{Artifact, ArtifactKind, DetectedFact, OnboardedTarget};
use compliance_core::AgentConfig;
use crate::error::AgentError;
use crate::pipeline::git::{GitOps, RepoCredentials};
/// The paths and identifiers an ingest needs. Decoupled from the full
/// [`AgentConfig`] so ingest is testable without a complete config.
pub struct IngestContext<'a> {
/// Base directory for content-addressed blobs and working dirs.
pub artifact_store_base: &'a Path,
/// Base directory for git clones.
pub git_clone_base: &'a str,
/// Default SSH key path (used when an artifact provides none).
pub ssh_key_path: &'a str,
/// The id of the target these artifacts belong to (namespaces working dirs).
pub target_id: &'a str,
}
impl<'a> IngestContext<'a> {
/// Build an ingest context from the agent config for a given target.
pub fn from_config(config: &'a AgentConfig, target_id: &'a str) -> Self {
Self {
artifact_store_base: Path::new(&config.artifact_store_base_path),
git_clone_base: &config.git_clone_base_path,
ssh_key_path: &config.ssh_key_path,
target_id,
}
}
}
/// The result of ingesting one artifact.
pub struct IngestedArtifact {
/// The artifact this corresponds to ([`Artifact::id`]).
pub artifact_id: String,
/// The artifact kind.
pub kind: ArtifactKind,
/// Local working path (clone dir, extracted dir, or blob file). `None` for
/// artifacts with no on-disk form (live URL, plaintext, container ref).
pub working_path: Option<PathBuf>,
/// SHA-256 of the content (blobs) or git head SHA (git repos).
pub content_hash: Option<String>,
/// Stored blob size in bytes, when applicable.
pub size_bytes: Option<u64>,
/// Facts discovered during ingest.
pub facts: Vec<DetectedFact>,
}
/// All ingested artifacts for a target, keyed by artifact id.
pub struct IngestSet {
/// The ingested artifacts, keyed by [`Artifact::id`].
pub by_artifact: HashMap<String, IngestedArtifact>,
}
impl IngestSet {
/// The working paths of every ingested artifact that has one — the input the
/// classifier expects.
pub fn working_paths(&self) -> HashMap<String, PathBuf> {
self.by_artifact
.iter()
.filter_map(|(id, a)| a.working_path.clone().map(|p| (id.clone(), p)))
.collect()
}
/// The ingest result for a specific artifact.
pub fn get(&self, artifact_id: &str) -> Option<&IngestedArtifact> {
self.by_artifact.get(artifact_id)
}
}
/// Ingest every artifact on a target.
pub fn ingest_all(
target: &OnboardedTarget,
ctx: &IngestContext<'_>,
) -> Result<IngestSet, AgentError> {
let mut by_artifact = HashMap::new();
for artifact in &target.artifacts {
let ingested = ingest_artifact(artifact, ctx)?;
by_artifact.insert(artifact.id.clone(), ingested);
}
Ok(IngestSet { by_artifact })
}
/// Ingest a single artifact, dispatching on its kind.
pub fn ingest_artifact(
artifact: &Artifact,
ctx: &IngestContext<'_>,
) -> Result<IngestedArtifact, AgentError> {
match artifact.kind {
ArtifactKind::GitRepo => ingest_git(artifact, ctx),
ArtifactKind::SourceArchive | ArtifactKind::MobilePackage | ArtifactKind::PlcProject => {
ingest_blob(artifact, ctx, true)
}
ArtifactKind::FirmwareImage => ingest_blob(artifact, ctx, false),
ArtifactKind::ContainerImage => Ok(metadata_only(
artifact,
DetectedFact::new("container_ref", artifact.source_ref.as_str(), "ingest"),
)),
ArtifactKind::LiveUrl => Ok(metadata_only(
artifact,
DetectedFact::new("live_url", artifact.source_ref.as_str(), "ingest"),
)),
ArtifactKind::PlaintextDescription => Ok(metadata_only(
artifact,
DetectedFact::new(
"description_len",
artifact.source_ref.len().to_string(),
"ingest",
),
)),
}
}
/// Clone (or fetch) a git artifact, recording the head SHA as the content hash.
fn ingest_git(
artifact: &Artifact,
ctx: &IngestContext<'_>,
) -> Result<IngestedArtifact, AgentError> {
let creds = credentials_for(artifact, ctx.ssh_key_path);
let git_ops = GitOps::new(ctx.git_clone_base, creds);
let repo_path = git_ops.clone_or_fetch(&artifact.source_ref, &artifact.id)?;
let head = GitOps::get_head_sha(&repo_path).ok();
Ok(IngestedArtifact {
artifact_id: artifact.id.clone(),
kind: artifact.kind,
working_path: Some(repo_path),
content_hash: head,
size_bytes: None,
facts: Vec::new(),
})
}
/// Store a blob artifact content-addressed. When `extract` is set and the blob
/// is a zip container (source archive, APK/AAB/IPA), also unpack it into a
/// working directory; otherwise the working path is the stored blob.
fn ingest_blob(
artifact: &Artifact,
ctx: &IngestContext<'_>,
extract: bool,
) -> Result<IngestedArtifact, AgentError> {
let base = ctx.artifact_store_base;
let src = local_source(artifact)?;
let (sha, size) = blob::hash_file(&src)?;
let stored = blob::store_file(base, &src, &sha)?;
let mut facts = Vec::new();
let working_path = if extract {
let dest = blob::work_dir(base, ctx.target_id, &artifact.id);
match blob::extract_zip(&stored, &dest) {
Ok(()) => dest,
Err(e) => {
// Not a zip (e.g. a tar.gz source archive) — keep the blob and
// note it so later stages can decide what to do.
facts.push(DetectedFact::new(
"archive_unextracted",
e.to_string(),
"ingest",
));
stored.clone()
}
}
} else {
stored.clone()
};
Ok(IngestedArtifact {
artifact_id: artifact.id.clone(),
kind: artifact.kind,
working_path: Some(working_path),
content_hash: Some(sha),
size_bytes: Some(size),
facts,
})
}
/// An artifact with no on-disk form: record a single fact, no hash/path.
fn metadata_only(artifact: &Artifact, fact: DetectedFact) -> IngestedArtifact {
IngestedArtifact {
artifact_id: artifact.id.clone(),
kind: artifact.kind,
working_path: None,
content_hash: None,
size_bytes: None,
facts: vec![fact],
}
}
/// The local file backing a blob artifact: its `stored_path` if already
/// uploaded, else its `source_ref` interpreted as a filesystem path.
fn local_source(artifact: &Artifact) -> Result<PathBuf, AgentError> {
let path = artifact
.stored_path
.as_deref()
.unwrap_or(artifact.source_ref.as_str());
let path = PathBuf::from(path);
if !path.exists() {
return Err(AgentError::Other(format!(
"artifact {} source not found at {}",
artifact.id,
path.display()
)));
}
Ok(path)
}
/// Build git credentials from an artifact's auth plus a default SSH key path.
fn credentials_for(artifact: &Artifact, default_ssh_key_path: &str) -> RepoCredentials {
let auth = artifact.auth.as_ref();
RepoCredentials {
ssh_key_path: auth
.and_then(|a| a.ssh_key_path.clone())
.or_else(|| Some(default_ssh_key_path.to_string())),
auth_token: auth.and_then(|a| a.secret.clone()),
auth_username: auth.and_then(|a| a.username.clone()),
}
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
use compliance_core::models::{ArtifactAuth, TargetType};
/// A unique scratch directory, removed on drop.
struct Scratch(PathBuf);
impl Scratch {
fn new() -> Self {
let p = std::env::temp_dir().join(format!("cs-ingest-mod-{}", uuid::Uuid::new_v4()));
std::fs::create_dir_all(&p).expect("mkdir scratch");
Self(p)
}
}
impl Drop for Scratch {
fn drop(&mut self) {
let _ = std::fs::remove_dir_all(&self.0);
}
}
fn ctx_for<'a>(store: &'a Path, target_id: &'a str) -> IngestContext<'a> {
IngestContext {
artifact_store_base: store,
git_clone_base: "/tmp/cs-ingest-test-repos",
ssh_key_path: "/tmp/cs-ingest-test-ssh",
target_id,
}
}
#[test]
fn firmware_blob_is_hashed_and_stored() {
let scratch = Scratch::new();
let store = scratch.0.join("store");
let fw = scratch.0.join("fw.bin");
std::fs::write(&fw, b"firmware-bytes").expect("write");
let ctx = ctx_for(&store, "t1");
let artifact = Artifact::firmware_image(fw.to_string_lossy().to_string());
let out = ingest_artifact(&artifact, &ctx).expect("ingest");
assert_eq!(out.kind, ArtifactKind::FirmwareImage);
assert_eq!(out.size_bytes, Some(14));
let sha = out.content_hash.expect("hash");
assert_eq!(sha.len(), 64);
// working path is the content-addressed blob
let wp = out.working_path.expect("working path");
assert!(wp.starts_with(store.join("blobs")));
}
#[test]
fn live_url_has_no_blob() {
let scratch = Scratch::new();
let store = scratch.0.join("store");
let ctx = ctx_for(&store, "t1");
let artifact = Artifact::live_url("https://example.com");
let out = ingest_artifact(&artifact, &ctx).expect("ingest");
assert!(out.working_path.is_none());
assert!(out.content_hash.is_none());
assert!(out.facts.iter().any(|f| f.key == "live_url"));
}
#[test]
fn ingest_all_collects_working_paths() {
let scratch = Scratch::new();
let store = scratch.0.join("store");
let fw = scratch.0.join("fw.bin");
std::fs::write(&fw, b"abc").expect("write");
let ctx = ctx_for(&store, "t1");
let mut target = OnboardedTarget::new("t".to_string(), TargetType::FirmwareBareMetal);
target
.artifacts
.push(Artifact::firmware_image(fw.to_string_lossy().to_string()));
target.artifacts.push(Artifact::live_url("https://x"));
let set = ingest_all(&target, &ctx).expect("ingest all");
assert_eq!(set.by_artifact.len(), 2);
// Only the firmware artifact yields a working path.
assert_eq!(set.working_paths().len(), 1);
}
#[test]
fn credentials_prefer_artifact_auth() {
let mut artifact = Artifact::git_repo("https://git/x", "main");
artifact.auth = Some(ArtifactAuth {
method: "token".to_string(),
username: Some("bob".to_string()),
secret: Some("pat".to_string()),
..Default::default()
});
let creds = credentials_for(&artifact, "/default/ssh/key");
assert_eq!(creds.auth_token.as_deref(), Some("pat"));
assert_eq!(creds.auth_username.as_deref(), Some("bob"));
}
#[test]
fn credentials_fall_back_to_default_ssh_key() {
let artifact = Artifact::git_repo("git@host:x.git", "main");
let creds = credentials_for(&artifact, "/default/ssh/key");
assert_eq!(creds.ssh_key_path.as_deref(), Some("/default/ssh/key"));
assert!(creds.auth_token.is_none());
}
}
-3
View File
@@ -2,13 +2,10 @@
pub mod agent;
pub mod api;
pub mod classify;
pub mod config;
pub mod database;
pub mod error;
pub mod ingest;
pub mod llm;
pub mod migrate;
pub mod pentest;
pub mod pipeline;
pub mod rag;
+1 -54
View File
@@ -1,50 +1,4 @@
use compliance_agent::{agent, api, config, database, migrate, scheduler, ssh, webhooks};
/// Run the `migrate onboarding` subcommand and exit. Backfills (or reverts) the
/// unified `onboarded_targets` collection per tenant.
///
/// Usage: `compliance-agent migrate onboarding [--all | --tenant <id>] [--dry-run] [--revert]`
async fn run_migration(
args: &[String],
pool: &database::DatabasePool,
) -> Result<(), compliance_agent::error::AgentError> {
if args.get(2).map(String::as_str) != Some("onboarding") {
eprintln!(
"usage: compliance-agent migrate onboarding [--all | --tenant <id>] [--dry-run] [--revert]"
);
std::process::exit(2);
}
let has = |flag: &str| args.iter().any(|a| a == flag);
let dry_run = has("--dry-run");
let revert = has("--revert");
let tenant = args
.iter()
.position(|a| a == "--tenant")
.and_then(|i| args.get(i + 1))
.cloned();
let tenants: Vec<String> = if has("--all") {
pool.list_tenant_ids().await?
} else if let Some(t) = tenant {
vec![t]
} else {
eprintln!("specify --all or --tenant <id>");
std::process::exit(2);
};
for tenant_id in tenants {
let db = pool.for_tenant_id(&tenant_id).await?;
if revert {
migrate::onboarding::revert(&db).await?;
println!("[{tenant_id}] reverted onboarding backfill");
} else {
let report = migrate::onboarding::backfill_onboarded_targets(&db, dry_run).await?;
let prefix = if dry_run { "(dry-run) " } else { "" };
println!("[{tenant_id}] {prefix}{report:?}");
}
}
Ok(())
}
use compliance_agent::{agent, api, config, database, scheduler, ssh, webhooks};
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
@@ -77,13 +31,6 @@ async fn main() -> Result<(), Box<dyn std::error::Error>> {
let db_pool =
database::DatabasePool::connect(&config.mongodb_uri, &config.mongodb_database).await?;
// One-shot subcommands run and exit without starting the servers.
let args: Vec<String> = std::env::args().collect();
if args.get(1).map(String::as_str) == Some("migrate") {
run_migration(&args, &db_pool).await?;
return Ok(());
}
let agent = agent::ComplianceAgent::new(config.clone(), db_pool);
tracing::info!("Starting scheduler...");
-8
View File
@@ -1,8 +0,0 @@
//! One-time data migrations.
//!
//! Currently just the onboarding backfill ([`onboarding`]), which folds the
//! legacy `repositories` and `dast_targets` collections into the unified
//! `onboarded_targets` collection, preserving `_id` so every downstream record
//! keyed by `repo_id` / `target_id` keeps resolving.
pub mod onboarding;
-406
View File
@@ -1,406 +0,0 @@
//! Backfill: legacy `repositories` + `dast_targets` → `onboarded_targets`.
//!
//! The transforms here are **id-preserving**: an [`OnboardedTarget`] keeps the
//! same `_id` as the `TrackedRepository` / `DastTarget` it came from, so every
//! downstream collection keyed by that hex id (findings, sbom, scan_runs,
//! graph, dast_*, pentest_*) keeps resolving with zero row rewrites, and
//! existing webhook URLs keep working. The mapping functions are pure and unit
//! tested; the DB orchestration (idempotent per-tenant backfill + revert) is a
//! thin driver over them.
use compliance_core::models::{
Artifact, ArtifactKind, DastTarget, DastTargetType, GitArtifactConfig, IssueTrackerConfig,
OnboardedTarget, TargetType, TrackedRepository, WebArtifactConfig,
};
use futures_util::TryStreamExt;
use mongodb::bson::{doc, Document};
use crate::database::Database;
use crate::error::AgentError;
/// Marker id in `schema_migrations` recording that the backfill has run.
const MIGRATION_MARKER: &str = "onboarding_backfill_v1";
/// Summary of a backfill run.
#[derive(Debug, Clone, Default, PartialEq, Eq)]
pub struct MigrationReport {
/// Repositories turned into onboarded targets.
pub repos_migrated: u64,
/// DAST targets folded into an existing (repo-linked) target as a LiveUrl.
pub dast_targets_folded: u64,
/// DAST targets with no repo link, migrated as standalone targets.
pub dast_targets_standalone: u64,
/// Records skipped because a target with that `_id` already existed.
pub skipped_existing: u64,
}
/// Map a legacy `DastTargetType` to a unified [`TargetType`]. REST/GraphQL APIs
/// are backend services; a browser app is a web app.
fn target_type_for_dast(kind: &DastTargetType) -> TargetType {
match kind {
DastTargetType::WebApp => TargetType::WebApp,
DastTargetType::RestApi | DastTargetType::GraphQl => TargetType::BackendService,
}
}
/// Build the LiveUrl artifact for a DAST target (its base URL + crawl config +
/// auth). Shared by fold-in and standalone migration.
pub fn dast_to_artifact(dast: &DastTarget) -> Artifact {
let mut artifact = Artifact::live_url(dast.base_url.clone());
artifact.web = Some(WebArtifactConfig {
target_kind: dast.target_type.clone(),
excluded_paths: dast.excluded_paths.clone(),
max_crawl_depth: dast.max_crawl_depth,
rate_limit: dast.rate_limit,
allow_destructive: dast.allow_destructive,
});
artifact.auth = dast.auth_config.clone().map(Into::into);
artifact
}
/// Map a `TrackedRepository` to an onboarded target, preserving `_id`. The git
/// remote becomes a `GitRepo` artifact carrying the repo's branch, watermark,
/// and auth; tracker config folds into `scan_config`.
///
/// `target_type` is a safe default (`BackendService`) — the classifier can
/// refine it later; `classification` is left `None` (unconfirmed).
pub fn repo_to_target(repo: &TrackedRepository) -> OnboardedTarget {
let mut target = OnboardedTarget::new(repo.name.clone(), TargetType::BackendService);
target.id = repo.id;
let mut artifact = Artifact::git_repo(repo.git_url.clone(), repo.default_branch.clone());
artifact.git = Some(GitArtifactConfig {
default_branch: repo.default_branch.clone(),
last_scanned_commit: repo.last_scanned_commit.clone(),
local_path: repo.local_path.clone(),
});
if repo.auth_token.is_some() || repo.auth_username.is_some() {
artifact.auth = Some(compliance_core::models::ArtifactAuth {
method: "token".to_string(),
username: repo.auth_username.clone(),
secret: repo.auth_token.clone(),
..Default::default()
});
}
target.artifacts.push(artifact);
if repo.tracker_type.is_some() {
target.scan_config.issue_tracker = Some(IssueTrackerConfig {
tracker_type: repo.tracker_type.clone(),
owner: repo.tracker_owner.clone(),
repo: repo.tracker_repo.clone(),
token: repo.tracker_token.clone(),
});
}
target.scan_schedule = repo.scan_schedule.clone();
target.webhook_enabled = repo.webhook_enabled;
target.webhook_secret = repo.webhook_secret.clone();
target.findings_count = repo.findings_count;
target.created_at = repo.created_at;
target.updated_at = repo.updated_at;
target
}
/// Append a DAST target's LiveUrl artifact onto an existing (repo-derived)
/// target. If the repo default was `BackendService` but the DAST target is a
/// browser web app, promote the type to `WebApp`.
pub fn fold_dast_into_target(target: &mut OnboardedTarget, dast: &DastTarget) {
if matches!(dast.target_type, DastTargetType::WebApp)
&& target.target_type == TargetType::BackendService
{
target.target_type = TargetType::WebApp;
}
if !target.has(ArtifactKind::LiveUrl) {
target.artifacts.push(dast_to_artifact(dast));
}
}
/// Map a repo-less DAST target to a standalone onboarded target, preserving `_id`.
pub fn dast_to_standalone_target(dast: &DastTarget) -> OnboardedTarget {
let mut target =
OnboardedTarget::new(dast.name.clone(), target_type_for_dast(&dast.target_type));
target.id = dast.id;
target.artifacts.push(dast_to_artifact(dast));
target.created_at = dast.created_at;
target.updated_at = dast.updated_at;
target
}
/// Whether the onboarding backfill has already been applied to this database.
pub async fn already_applied(db: &Database) -> Result<bool, AgentError> {
let found = db
.collection_named::<Document>("schema_migrations")
.find_one(doc! { "_id": MIGRATION_MARKER })
.await?;
Ok(found.is_some())
}
/// Backfill `onboarded_targets` from `repositories` + `dast_targets` for one
/// tenant database.
///
/// Id-preserving and **idempotent**: targets that already exist (by `_id`) are
/// skipped, so re-running is safe. With `dry_run`, computes the report without
/// writing. The legacy collections are never deleted; the only mutation outside
/// `onboarded_targets` is the history relink of folded DAST targets, which is
/// logged so [`revert`] can undo it.
pub async fn backfill_onboarded_targets(
db: &Database,
dry_run: bool,
) -> Result<MigrationReport, AgentError> {
let mut report = MigrationReport::default();
// 1. repositories -> onboarded_targets (preserve _id, skip existing).
let mut repos = db.repositories().find(doc! {}).await?;
while let Some(repo) = repos.try_next().await? {
let Some(id) = repo.id else { continue };
if db
.onboarded_targets()
.find_one(doc! { "_id": id })
.await?
.is_some()
{
report.skipped_existing += 1;
continue;
}
if !dry_run {
db.onboarded_targets()
.insert_one(repo_to_target(&repo))
.await?;
}
report.repos_migrated += 1;
}
// 2. dast_targets -> fold into the linked repo target, or migrate standalone.
let mut dasts = db.dast_targets().find(doc! {}).await?;
while let Some(dast) = dasts.try_next().await? {
let Some(dast_id) = dast.id else { continue };
let repo_oid = dast
.repo_id
.as_deref()
.and_then(|r| mongodb::bson::oid::ObjectId::parse_str(r).ok());
let linked = match repo_oid {
Some(oid) => db.onboarded_targets().find_one(doc! { "_id": oid }).await?,
None => None,
};
match (linked, repo_oid) {
// Fold into an existing repo-derived target.
(Some(mut target), Some(oid)) => {
if target.has(ArtifactKind::LiveUrl) {
report.skipped_existing += 1; // already folded on a prior run
continue;
}
fold_dast_into_target(&mut target, &dast);
if !dry_run {
db.onboarded_targets()
.replace_one(doc! { "_id": oid }, &target)
.await?;
relink_history(db, &dast_id.to_hex(), &oid.to_hex()).await?;
}
report.dast_targets_folded += 1;
}
// No linked repo target: migrate as a standalone target (keeps _id).
_ => {
if db
.onboarded_targets()
.find_one(doc! { "_id": dast_id })
.await?
.is_some()
{
report.skipped_existing += 1;
continue;
}
if !dry_run {
db.onboarded_targets()
.insert_one(dast_to_standalone_target(&dast))
.await?;
}
report.dast_targets_standalone += 1;
}
}
}
if !dry_run {
db.collection_named::<Document>("schema_migrations")
.update_one(
doc! { "_id": MIGRATION_MARKER },
doc! { "$set": { "applied_at": mongodb::bson::DateTime::now() } },
)
.upsert(true)
.await?;
}
Ok(report)
}
/// Relink DAST scan runs and pentest sessions from the old DAST target id to the
/// unified target id, logging each move so [`revert`] can undo it.
///
/// Note: if multiple DAST targets fold into the same repo target, revert
/// restores only the last-logged mapping — a rare edge. The source collections
/// (`repositories`, `dast_targets`) are never deleted, so no data is lost.
async fn relink_history(db: &Database, old_id: &str, new_id: &str) -> Result<(), AgentError> {
db.dast_scan_runs()
.update_many(
doc! { "target_id": old_id },
doc! { "$set": { "target_id": new_id } },
)
.await?;
db.pentest_sessions()
.update_many(
doc! { "target_id": old_id },
doc! { "$set": { "target_id": new_id } },
)
.await?;
db.collection_named::<Document>("onboarding_migration_log")
.insert_one(doc! { "old_target_id": old_id, "new_target_id": new_id })
.await?;
Ok(())
}
/// Undo the backfill: replay the relink log in reverse, drop `onboarded_targets`
/// and the log, and clear the marker. The legacy collections are untouched, so
/// this restores the pre-migration state.
pub async fn revert(db: &Database) -> Result<(), AgentError> {
let log = db.collection_named::<Document>("onboarding_migration_log");
let mut cursor = log.find(doc! {}).await?;
while let Some(entry) = cursor.try_next().await? {
if let (Ok(old), Ok(new)) = (
entry.get_str("old_target_id"),
entry.get_str("new_target_id"),
) {
db.dast_scan_runs()
.update_many(
doc! { "target_id": new },
doc! { "$set": { "target_id": old } },
)
.await?;
db.pentest_sessions()
.update_many(
doc! { "target_id": new },
doc! { "$set": { "target_id": old } },
)
.await?;
}
}
db.onboarded_targets().drop().await?;
log.drop().await?;
db.collection_named::<Document>("schema_migrations")
.delete_one(doc! { "_id": MIGRATION_MARKER })
.await?;
Ok(())
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
use compliance_core::models::{DastAuthConfig, TrackerType};
fn repo() -> TrackedRepository {
let mut r = TrackedRepository::new("acme".to_string(), "https://git/acme.git".to_string());
r.id = Some(mongodb::bson::oid::ObjectId::new());
r.default_branch = "develop".to_string();
r.last_scanned_commit = Some("abc123".to_string());
r.auth_token = Some("pat".to_string());
r.auth_username = Some("bob".to_string());
r.tracker_type = Some(TrackerType::Gitea);
r.tracker_owner = Some("acme".to_string());
r.findings_count = 7;
r
}
fn dast(repo_id: Option<String>, kind: DastTargetType) -> DastTarget {
let mut d = DastTarget::new(
"acme-web".to_string(),
"https://acme.example.com".to_string(),
kind,
);
d.id = Some(mongodb::bson::oid::ObjectId::new());
d.repo_id = repo_id;
d.max_crawl_depth = 5;
d.auth_config = Some(DastAuthConfig {
method: "bearer".to_string(),
login_url: None,
username: None,
password: None,
token: Some("tok".to_string()),
headers: None,
});
d
}
#[test]
fn repo_maps_preserving_id_and_git_artifact() {
let r = repo();
let t = repo_to_target(&r);
assert_eq!(t.id, r.id); // id preserved
assert_eq!(t.findings_count, 7);
assert_eq!(t.scan_schedule, r.scan_schedule);
let git = t.code_artifact().expect("git artifact");
assert_eq!(git.kind, ArtifactKind::GitRepo);
assert_eq!(git.source_ref, "https://git/acme.git");
let gc = git.git.as_ref().expect("git config");
assert_eq!(gc.default_branch, "develop");
assert_eq!(gc.last_scanned_commit.as_deref(), Some("abc123"));
let auth = git.auth.as_ref().expect("auth");
assert_eq!(auth.secret.as_deref(), Some("pat"));
assert_eq!(auth.username.as_deref(), Some("bob"));
assert_eq!(
t.scan_config
.issue_tracker
.as_ref()
.and_then(|it| it.tracker_type.clone()),
Some(TrackerType::Gitea)
);
}
#[test]
fn standalone_dast_maps_preserving_id_and_live_url() {
let d = dast(None, DastTargetType::WebApp);
let t = dast_to_standalone_target(&d);
assert_eq!(t.id, d.id);
assert_eq!(t.target_type, TargetType::WebApp);
let url = t.live_url().expect("live url");
assert_eq!(url.source_ref, "https://acme.example.com");
let web = url.web.as_ref().expect("web config");
assert_eq!(web.max_crawl_depth, 5);
assert_eq!(
url.auth.as_ref().and_then(|a| a.secret.clone()),
Some("tok".to_string())
);
}
#[test]
fn rest_api_dast_maps_to_backend_service() {
let d = dast(None, DastTargetType::RestApi);
assert_eq!(
dast_to_standalone_target(&d).target_type,
TargetType::BackendService
);
}
#[test]
fn fold_adds_live_url_and_promotes_webapp() {
let mut t = repo_to_target(&repo());
assert_eq!(t.target_type, TargetType::BackendService);
fold_dast_into_target(&mut t, &dast(Some("x".to_string()), DastTargetType::WebApp));
assert_eq!(t.target_type, TargetType::WebApp); // promoted
assert!(t.has(ArtifactKind::LiveUrl));
assert!(t.has(ArtifactKind::GitRepo));
}
#[test]
fn fold_is_idempotent_on_live_url() {
let mut t = repo_to_target(&repo());
let d = dast(Some("x".to_string()), DastTargetType::WebApp);
fold_dast_into_target(&mut t, &d);
fold_dast_into_target(&mut t, &d);
let live_urls = t
.artifacts
.iter()
.filter(|a| a.kind == ArtifactKind::LiveUrl)
.count();
assert_eq!(live_urls, 1);
}
}
-2
View File
@@ -328,7 +328,6 @@ mod tests {
scan_schedule: String::new(),
cve_monitor_schedule: String::new(),
git_clone_base_path: String::new(),
artifact_store_base_path: String::new(),
ssh_key_path: String::new(),
keycloak_url: None,
keycloak_realm: None,
@@ -342,7 +341,6 @@ mod tests {
pentest_imap_password: None,
admin_api_token: None,
tenant_registry_url: None,
unified_pipeline: false,
}
}
@@ -1,152 +0,0 @@
//! Firmware SBOM via tramiton.
//!
//! Phase 2 (full, the default): drive a **reproducible build** with tramiton's
//! `NixBackend` — `analyze` → `seal_and_build` → a sealed lock whose libraries
//! are pinned and whose firmware artifact carries a content hash — then render
//! the SBOM from the lock plus deep binary SCA of pre-compiled inputs. This is
//! the complete bill of materials (toolchain + every fetched library + the
//! firmware image), the same one `tramiton sbom` produces.
//!
//! Phase 1 fallback (analysis-only): when no nix backend is available or the
//! build fails, fall back to the resolvable libraries + toolchain from the build
//! plan alone (no build). A scan therefore always yields *something*, and a nix
//! that can't run in the deployment never breaks a scan.
use std::path::Path;
use compliance_core::models::{SbomEntry, TargetType};
use tramiton_repro::ReproBackend;
use tramiton_sbom::ComponentKind;
/// Whether firmware SBOM applies to this target family.
pub fn is_firmware_target(target_type: TargetType) -> bool {
matches!(
target_type,
TargetType::FirmwareBareMetal | TargetType::FirmwareRtos | TargetType::EmbeddedLinuxYocto
)
}
/// Build SBOM entries for a firmware target from its source tree. Prefers a full
/// reproducible build (sealed lock); falls back to analysis-only. Returns an
/// empty vector when tramiton cannot even form a build plan.
pub async fn firmware_sbom_entries(path: &Path, repo_id: &str) -> Vec<SbomEntry> {
let p = path.to_path_buf();
let repo = repo_id.to_string();
// The whole analyze → seal → build → render sequence is blocking (it shells
// out to nix), so keep it off the async runtime. Bound it: a firmware build
// that hangs must not wedge the scan (the orphaned task is abandoned).
let handle = tokio::task::spawn_blocking(move || build_sbom_blocking(&p, &repo));
match tokio::time::timeout(std::time::Duration::from_secs(900), handle).await {
Ok(Ok(entries)) => entries,
Ok(Err(e)) => {
tracing::warn!(repo_id, error = %e, "Firmware SBOM: task join error");
Vec::new()
}
Err(_) => {
tracing::warn!(repo_id, "Firmware SBOM: build exceeded 15m; skipping");
Vec::new()
}
}
}
fn build_sbom_blocking(path: &Path, repo_id: &str) -> Vec<SbomEntry> {
let repo = tramiton_core::Repo::new(path);
let plan = match tramiton_core::provider::analyze(&repo) {
Ok(Some(bp)) => bp,
Ok(None) => return Vec::new(),
Err(e) => {
tracing::warn!(repo_id, error = %e, "Firmware SBOM: tramiton analyze failed");
return Vec::new();
}
};
// Phase 2: reproducible build → sealed lock → complete SBOM.
if let Some(backend) = tramiton_repro::NixBackend::detect() {
match tramiton_repro::seal_and_build(&backend, &plan, path) {
Ok(lock) => {
let mut sbom = tramiton_sbom::Sbom::from_lock(&lock, repo_id);
// Deep binary SCA of any pre-compiled inputs in the tree.
sbom.components.extend(tramiton_sbom::binary::scan(path));
let entries = sbom_to_entries(&sbom, repo_id);
tracing::info!(
repo_id,
backend = backend.name(),
count = entries.len(),
"Firmware SBOM: sealed reproducible build"
);
return entries;
}
Err(e) => {
tracing::warn!(repo_id, error = %e, "Firmware SBOM: reproducible build failed; falling back to analysis-only")
}
}
} else {
tracing::info!(
repo_id,
"Firmware SBOM: no nix backend available; analysis-only SBOM"
);
}
// Phase 1 fallback: analysis-only (toolchain + resolvable libraries).
analysis_entries(&plan, repo_id)
}
/// Map a rendered [`tramiton_sbom::Sbom`] (primary firmware + components) into
/// our [`SbomEntry`] rows. Source-file (`File`) components are dropped — they are
/// build inputs, not a dependency inventory.
fn sbom_to_entries(sbom: &tramiton_sbom::Sbom, repo_id: &str) -> Vec<SbomEntry> {
let mut entries = Vec::new();
if let Some(primary) = &sbom.primary {
entries.push(component_to_entry(primary, repo_id));
}
for c in &sbom.components {
if matches!(c.kind, ComponentKind::File) {
continue;
}
entries.push(component_to_entry(c, repo_id));
}
entries
}
fn component_to_entry(c: &tramiton_sbom::Component, repo_id: &str) -> SbomEntry {
let manager = match c.kind {
ComponentKind::Firmware => "firmware",
ComponentKind::Library => "library",
ComponentKind::Toolchain => "toolchain",
ComponentKind::File => "file",
};
let mut entry = SbomEntry::new(
repo_id.to_string(),
c.name.clone(),
c.version.clone().unwrap_or_default(),
manager.to_string(),
);
entry.purl = c.source.clone();
entry
}
/// Analysis-only components from the build plan: the cross-toolchain plus the
/// resolvable fetched libraries, without a build.
fn analysis_entries(bp: &tramiton_core::BuildPlan, repo_id: &str) -> Vec<SbomEntry> {
let mut entries = Vec::new();
if let Some(id) = bp.toolchain.id.clone() {
let version = bp.toolchain.version.clone().unwrap_or_default();
entries.push(SbomEntry::new(
repo_id.to_string(),
id,
version,
"toolchain".to_string(),
));
}
for lib in tramiton_repro::lock::libraries_from_inputs(&bp.inputs) {
let mut entry = SbomEntry::new(
repo_id.to_string(),
lib.name,
lib.revision,
"library".to_string(),
);
entry.purl = lib.source;
entries.push(entry);
}
entries
}
-2
View File
@@ -1,7 +1,6 @@
pub mod code_review;
pub mod cve;
pub mod dedup;
pub mod firmware_sbom;
pub mod git;
pub mod gitleaks;
mod graph_build;
@@ -9,7 +8,6 @@ mod issue_creation;
pub mod lint;
pub mod orchestrator;
pub mod patterns;
pub mod plan;
mod pr_review;
pub mod sbom;
pub mod semgrep;
@@ -15,7 +15,6 @@ use crate::pipeline::git::GitOps;
use crate::pipeline::gitleaks::GitleaksScanner;
use crate::pipeline::lint::LintScanner;
use crate::pipeline::patterns::{GdprPatternScanner, OAuthPatternScanner};
use crate::pipeline::plan::build_scan_plan;
use crate::pipeline::sbom::SbomScanner;
use crate::pipeline::semgrep::SemgrepScanner;
@@ -420,320 +419,6 @@ impl PipelineOrchestrator {
Ok(new_count)
}
/// Unified entry point (behind `UNIFIED_PIPELINE`): run a scan for an
/// `OnboardedTarget`. Mirrors [`Self::run`] but sources the target from
/// `onboarded_targets` and dispatches by the scan plan.
#[tracing::instrument(skip_all, fields(target_id = %target_id, trigger = ?trigger))]
pub async fn run_target(
&self,
target_id: &str,
trigger: ScanTrigger,
) -> Result<(), AgentError> {
let oid = mongodb::bson::oid::ObjectId::parse_str(target_id)
.map_err(|e| AgentError::Other(e.to_string()))?;
let target = self
.db
.onboarded_targets()
.find_one(doc! { "_id": oid })
.await?
.ok_or_else(|| AgentError::Other(format!("Onboarded target {target_id} not found")))?;
let scan_run = ScanRun::new(target_id.to_string(), trigger);
let insert = self.db.scan_runs().insert_one(&scan_run).await?;
let scan_run_id = insert
.inserted_id
.as_object_id()
.map(|id| id.to_hex())
.unwrap_or_default();
let result = self.run_target_pipeline(&target, &scan_run_id).await;
match &result {
Ok(count) => {
self.db
.scan_runs()
.update_one(
doc! { "_id": &insert.inserted_id },
doc! { "$set": {
"status": "completed",
"current_phase": "completed",
"new_findings_count": *count as i64,
"completed_at": mongodb::bson::DateTime::now(),
} },
)
.await?;
// Refresh the target's cached findings count. The shared pipeline
// (Stage 7) increments `repositories`, which the unified path does
// not use, so set the accurate total on the target itself.
let total = self
.db
.findings()
.count_documents(doc! { "repo_id": target_id })
.await
.unwrap_or(*count as u64);
self.db
.onboarded_targets()
.update_one(
doc! { "_id": oid },
doc! { "$set": {
"findings_count": total as i64,
"updated_at": mongodb::bson::DateTime::now(),
} },
)
.await?;
}
Err(e) => {
tracing::error!(target_id, error = %e, "Unified scan pipeline failed");
self.db
.scan_runs()
.update_one(
doc! { "_id": &insert.inserted_id },
doc! { "$set": {
"status": "failed",
"error_message": e.to_string(),
"completed_at": mongodb::bson::DateTime::now(),
} },
)
.await?;
}
}
result.map(|_| ())
}
/// Run the applicable scans for a target. For a code target this reuses the
/// full legacy pipeline over the code artifact (clone → SAST umbrella →
/// triage → persist → issues → DAST); firmware/PLC/mobile scanners are
/// follow-ups (#128/#129/#130). Returns the number of new findings.
async fn run_target_pipeline(
&self,
target: &OnboardedTarget,
scan_run_id: &str,
) -> Result<u32, AgentError> {
let target_id = target.id.map(|id| id.to_hex()).unwrap_or_default();
let plan = build_scan_plan(target);
tracing::info!(
target_id = %target_id,
target_type = %target.target_type,
planned_steps = plan.steps.len(),
"Unified pipeline: scan plan built"
);
// Ingest + classify (tramiton for firmware) and store the detected type.
self.classify_and_store(target, &target_id, scan_run_id)
.await;
// Provision a DAST target from a LiveUrl artifact so DAST fires for
// wizard-created targets, not just migrated ones.
self.ensure_dast_target(target, &plan).await;
match target.code_artifact() {
Some(code) if code.kind == ArtifactKind::GitRepo => {
let repo = repo_view_from_target(target, code);
let new_count = self.run_pipeline(&repo, scan_run_id).await?;
self.finalize_target(target, &repo, new_count).await?;
Ok(new_count)
}
Some(_) => {
tracing::warn!(
target_id = %target_id,
"Unified pipeline: source-archive scanning not yet wired; skipping"
);
Ok(0)
}
None => {
// No code to scan. Firmware/PLC/mobile static scanners land in
// #128/#129/#130; DAST for a running URL still works when a
// DastTarget row exists (migrated targets).
tracing::info!(
target_id = %target_id,
"Unified pipeline: no code artifact; attempting DAST only"
);
self.update_phase(scan_run_id, "dast_scanning").await;
self.maybe_trigger_dast(&target_id, scan_run_id).await;
Ok(0)
}
}
}
/// Ingest the target's artifacts, classify (tramiton for firmware/RTOS/Yocto,
/// heuristics otherwise), and store the detected classification on the target.
/// Best-effort — never fails the scan.
async fn classify_and_store(
&self,
target: &OnboardedTarget,
target_id: &str,
scan_run_id: &str,
) {
self.update_phase(scan_run_id, "classification").await;
let ctx = crate::ingest::IngestContext::from_config(&self.config, target_id);
let ingest_set = match crate::ingest::ingest_all(target, &ctx) {
Ok(set) => set,
Err(e) => {
tracing::warn!(target_id, error = %e, "Unified pipeline: ingest for classification failed");
return;
}
};
let working_paths = ingest_set.working_paths();
match crate::classify::classify_target(
target,
&working_paths,
&crate::classify::TramitonNative,
)
.await
{
Ok(classification) => {
tracing::info!(
target_id,
suggested = %classification.suggested,
"Unified pipeline: classified target"
);
if let (Some(oid), Ok(bson)) = (target.id, mongodb::bson::to_bson(&classification))
{
let _ = self
.db
.onboarded_targets()
.update_one(
doc! { "_id": oid },
doc! { "$set": { "classification": bson } },
)
.await;
}
}
Err(e) => {
tracing::warn!(target_id, error = %e, "Unified pipeline: classification failed")
}
}
// Analysis-based firmware SBOM: for embedded targets, derive components
// (resolved libraries + cross-toolchain) from tramiton's build-plan
// analysis over the already-ingested source — no build, no binary
// upload. Best-effort; empty when no build plan forms.
if crate::pipeline::firmware_sbom::is_firmware_target(target.target_type) {
if let Some(code) = target.code_artifact() {
if let Some(path) = working_paths.get(&code.id) {
let entries =
crate::pipeline::firmware_sbom::firmware_sbom_entries(path, target_id)
.await;
if !entries.is_empty() {
let _ = self
.db
.sbom_entries()
.delete_many(doc! { "repo_id": target_id })
.await;
for entry in &entries {
let filter = doc! {
"repo_id": &entry.repo_id,
"name": &entry.name,
"version": &entry.version,
};
if let Ok(d) = mongodb::bson::to_document(entry) {
let _ = self
.db
.sbom_entries()
.update_one(filter, doc! { "$set": d })
.upsert(true)
.await;
}
}
tracing::info!(
target_id,
count = entries.len(),
"Firmware SBOM: stored components from tramiton analysis"
);
}
}
}
}
}
/// If the target has a `LiveUrl` artifact and DAST is planned, provision a
/// `DastTarget` (keyed by `repo_id` = target id) so the existing DAST trigger
/// fires for wizard-created targets. Idempotent.
async fn ensure_dast_target(
&self,
target: &OnboardedTarget,
plan: &crate::pipeline::plan::ScanPlan,
) {
if !plan.has(ScanType::Dast) {
return;
}
let (Some(url), Some(oid)) = (target.live_url(), target.id) else {
return;
};
let target_id = oid.to_hex();
if self
.db
.dast_targets()
.find_one(doc! { "repo_id": &target_id })
.await
.ok()
.flatten()
.is_some()
{
return; // already provisioned
}
let kind = url
.web
.as_ref()
.map(|w| w.target_kind.clone())
.unwrap_or(DastTargetType::WebApp);
let mut dast = DastTarget::new(target.name.clone(), url.source_ref.clone(), kind);
dast.repo_id = Some(target_id);
if let Some(web) = &url.web {
dast.excluded_paths = web.excluded_paths.clone();
dast.max_crawl_depth = web.max_crawl_depth;
dast.rate_limit = web.rate_limit;
dast.allow_destructive = web.allow_destructive;
}
if let Some(auth) = &url.auth {
dast.auth_config = Some(DastAuthConfig {
method: auth.method.clone(),
login_url: auth.login_url.clone(),
username: auth.username.clone(),
password: None,
token: auth.secret.clone(),
headers: auth.headers.clone(),
});
}
if let Err(e) = self.db.dast_targets().insert_one(&dast).await {
tracing::warn!(error = %e, "Unified pipeline: failed to provision DAST target");
}
}
/// Sync the onboarded-target document after a scan: bump `findings_count`
/// and advance the git artifact's `last_scanned_commit` watermark.
async fn finalize_target(
&self,
target: &OnboardedTarget,
repo: &TrackedRepository,
new_count: u32,
) -> Result<(), AgentError> {
let oid = match target.id {
Some(id) => id,
None => return Ok(()),
};
self.db
.onboarded_targets()
.update_one(
doc! { "_id": oid },
doc! {
"$inc": { "findings_count": new_count as i64 },
"$set": { "updated_at": mongodb::bson::DateTime::now() },
},
)
.await?;
let repo_path = std::path::Path::new(&self.config.git_clone_base_path).join(&repo.name);
if let (Ok(sha), Some(code)) = (GitOps::get_head_sha(&repo_path), target.code_artifact()) {
self.db
.onboarded_targets()
.update_one(
doc! { "_id": oid, "artifacts.id": &code.id },
doc! { "$set": { "artifacts.$.git.last_scanned_commit": sha } },
)
.await?;
}
Ok(())
}
pub(super) async fn update_phase(&self, scan_run_id: &str, phase: &str) {
if let Ok(oid) = mongodb::bson::oid::ObjectId::parse_str(scan_run_id) {
let _ = self
@@ -751,37 +436,6 @@ impl PipelineOrchestrator {
}
}
/// Build a legacy `TrackedRepository` view from an onboarded target's code
/// artifact, so the unified pipeline can reuse the existing repo pipeline. The
/// inverse of the migration's `repo_to_target`. `_id` is preserved so findings
/// and DAST lookups resolve against the same key.
fn repo_view_from_target(target: &OnboardedTarget, code: &Artifact) -> TrackedRepository {
let mut repo = TrackedRepository::new(target.name.clone(), code.source_ref.clone());
repo.id = target.id;
if let Some(git) = &code.git {
repo.default_branch = git.default_branch.clone();
repo.last_scanned_commit = git.last_scanned_commit.clone();
repo.local_path = git.local_path.clone();
}
if let Some(auth) = &code.auth {
repo.auth_token = auth.secret.clone();
repo.auth_username = auth.username.clone();
}
if let Some(it) = &target.scan_config.issue_tracker {
repo.tracker_type = it.tracker_type.clone();
repo.tracker_owner = it.owner.clone();
repo.tracker_repo = it.repo.clone();
repo.tracker_token = it.token.clone();
}
repo.scan_schedule = target.scan_schedule.clone();
repo.webhook_enabled = target.webhook_enabled;
repo.webhook_secret = target.webhook_secret.clone();
repo.findings_count = target.findings_count;
repo.created_at = target.created_at;
repo.updated_at = target.updated_at;
repo
}
/// Extract the scheme + host from a git URL.
/// e.g. "https://gitea.example.com/owner/repo.git" -> "https://gitea.example.com"
/// e.g. "ssh://git@gitea.example.com:22/owner/repo.git" -> "https://gitea.example.com"
@@ -806,50 +460,3 @@ pub(super) fn extract_base_url(git_url: &str) -> Option<String> {
None
}
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
use compliance_core::models::{
Artifact, ArtifactAuth, IssueTrackerConfig, TargetType, TrackerType,
};
#[test]
fn repo_view_preserves_id_git_auth_and_tracker() {
let mut target = OnboardedTarget::new("acme".to_string(), TargetType::WebApp);
target.id = Some(mongodb::bson::oid::ObjectId::new());
target.findings_count = 3;
target.scan_config.issue_tracker = Some(IssueTrackerConfig {
tracker_type: Some(TrackerType::Gitea),
owner: Some("acme".to_string()),
repo: Some("web".to_string()),
token: Some("tt".to_string()),
});
let mut artifact = Artifact::git_repo("https://git/acme.git", "develop");
if let Some(git) = artifact.git.as_mut() {
git.last_scanned_commit = Some("abc123".to_string());
}
artifact.auth = Some(ArtifactAuth {
method: "token".to_string(),
username: Some("bob".to_string()),
secret: Some("pat".to_string()),
..Default::default()
});
target.artifacts.push(artifact);
let code = target.code_artifact().expect("code artifact");
let repo = repo_view_from_target(&target, code);
assert_eq!(repo.id, target.id); // preserved
assert_eq!(repo.git_url, "https://git/acme.git");
assert_eq!(repo.default_branch, "develop");
assert_eq!(repo.last_scanned_commit.as_deref(), Some("abc123"));
assert_eq!(repo.auth_token.as_deref(), Some("pat"));
assert_eq!(repo.auth_username.as_deref(), Some("bob"));
assert_eq!(repo.tracker_type, Some(TrackerType::Gitea));
assert_eq!(repo.tracker_owner.as_deref(), Some("acme"));
assert_eq!(repo.findings_count, 3);
}
}
+1 -1
View File
@@ -215,7 +215,7 @@ fn scan_with_patterns(
repo_id.to_string(),
fingerprint,
scanner_name.to_string(),
scan_type,
scan_type.clone(),
pattern.title.clone(),
pattern.description.clone(),
pattern.severity.clone(),
-195
View File
@@ -1,195 +0,0 @@
//! The scan plan.
//!
//! [`build_scan_plan`] turns an [`OnboardedTarget`] into the concrete ordered
//! list of scans to run, each bound to the artifact it consumes. It intersects
//! the scan-applicability matrix ([`applicable_scans`]) with the target's
//! `scan_config` overrides: a scan runs when its required artifact is present and
//! it is either on by default or explicitly enabled, and is not explicitly
//! disabled. This is the decision engine the unified pipeline (`run_target`)
//! executes.
use compliance_core::models::{Artifact, ArtifactKind, OnboardedTarget, ScanPhase, ScanType};
use compliance_core::scan_matrix::applicable_scans;
/// One scan to run, bound to the artifact it operates on.
#[derive(Debug, Clone, PartialEq)]
pub struct ScanStep {
/// The scan to run.
pub scan_type: ScanType,
/// The pipeline phase to report while it runs.
pub phase: ScanPhase,
/// The id of the artifact this scan consumes ([`Artifact::id`]).
pub artifact_id: String,
}
/// The ordered set of scans to run for a target.
#[derive(Debug, Clone, Default, PartialEq)]
pub struct ScanPlan {
/// The scans, in matrix order.
pub steps: Vec<ScanStep>,
}
impl ScanPlan {
/// Whether the plan contains a step for the given scan type.
pub fn has(&self, scan: ScanType) -> bool {
self.steps.iter().any(|s| s.scan_type == scan)
}
/// Whether the plan is empty (nothing to run).
pub fn is_empty(&self) -> bool {
self.steps.is_empty()
}
}
/// Build the scan plan for a target: matrix defaults ∩ `scan_config`, each scan
/// bound to the artifact it consumes. Scans whose required artifact is absent, or
/// that are disabled, or off-by-default and not explicitly enabled, are dropped.
pub fn build_scan_plan(target: &OnboardedTarget) -> ScanPlan {
let enabled = &target.scan_config.enabled_scans;
let disabled = &target.scan_config.disabled_scans;
let mut steps = Vec::new();
for option in applicable_scans(target) {
// Required artifact missing → not runnable.
if option.blocked_reason.is_some() {
continue;
}
// Explicit opt-out wins.
if disabled.contains(&option.scan) {
continue;
}
// Run if on by default, or explicitly enabled.
if !option.default_on && !enabled.contains(&option.scan) {
continue;
}
let Some(artifact) = resolve_artifact(target, option.required_artifact) else {
continue;
};
steps.push(ScanStep {
scan_type: option.scan,
phase: phase_for(option.scan),
artifact_id: artifact.id.clone(),
});
}
ScanPlan { steps }
}
/// Resolve the artifact a scan consumes. A "code" requirement (represented by
/// `GitRepo`) is satisfied by a git repo *or* a source archive.
fn resolve_artifact(target: &OnboardedTarget, required: Option<ArtifactKind>) -> Option<&Artifact> {
match required {
Some(ArtifactKind::GitRepo) => target.code_artifact(),
Some(kind) => target.first_of(kind),
None => target.code_artifact().or_else(|| target.artifacts.first()),
}
}
/// The pipeline phase reported while a given scan runs.
fn phase_for(scan: ScanType) -> ScanPhase {
match scan {
ScanType::Sast => ScanPhase::Sast,
ScanType::Sbom => ScanPhase::SbomGeneration,
ScanType::Cve => ScanPhase::CveScanning,
ScanType::Gdpr | ScanType::OAuth => ScanPhase::PatternScanning,
ScanType::SecretDetection => ScanPhase::SecretDetection,
ScanType::Lint => ScanPhase::LintScanning,
ScanType::CodeReview => ScanPhase::CodeReview,
ScanType::Graph => ScanPhase::GraphBuilding,
ScanType::Dast => ScanPhase::DastScanning,
ScanType::FirmwareStatic => ScanPhase::FirmwareStatic,
ScanType::PlcControlLogic => ScanPhase::PlcAnalysis,
ScanType::MobileStatic => ScanPhase::MobileStatic,
ScanType::ContainerScan => ScanPhase::ContainerScan,
}
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
use compliance_core::models::{PlcFormat, TargetType};
fn target(target_type: TargetType, artifacts: Vec<Artifact>) -> OnboardedTarget {
let mut t = OnboardedTarget::new("t".to_string(), target_type);
t.artifacts = artifacts;
t
}
fn step_for<'a>(plan: &'a ScanPlan, scan: ScanType) -> Option<&'a ScanStep> {
plan.steps.iter().find(|s| s.scan_type == scan)
}
#[test]
fn webapp_with_code_and_url_runs_sast_and_dast_bound_to_the_right_artifacts() {
let code = Artifact::git_repo("https://git/x", "main");
let url = Artifact::live_url("https://x");
let (code_id, url_id) = (code.id.clone(), url.id.clone());
let t = target(TargetType::WebApp, vec![code, url]);
let plan = build_scan_plan(&t);
let sast = step_for(&plan, ScanType::Sast).expect("sast planned");
assert_eq!(sast.artifact_id, code_id);
assert_eq!(sast.phase, ScanPhase::Sast);
let dast = step_for(&plan, ScanType::Dast).expect("dast planned");
assert_eq!(dast.artifact_id, url_id);
}
#[test]
fn webapp_without_url_omits_dast() {
let t = target(TargetType::WebApp, vec![Artifact::git_repo("u", "main")]);
let plan = build_scan_plan(&t);
assert!(plan.has(ScanType::Sast));
assert!(!plan.has(ScanType::Dast));
}
#[test]
fn code_scan_binds_to_source_archive_when_no_git_repo() {
let arc = Artifact::source_archive("src.zip");
let arc_id = arc.id.clone();
let t = target(TargetType::BackendService, vec![arc]);
let plan = build_scan_plan(&t);
let sast = step_for(&plan, ScanType::Sast).expect("sast planned");
assert_eq!(sast.artifact_id, arc_id);
}
#[test]
fn firmware_sbom_and_cve_bind_to_the_firmware_image() {
let fw = Artifact::firmware_image("fw.bin");
let fw_id = fw.id.clone();
let t = target(TargetType::FirmwareBareMetal, vec![fw]);
let plan = build_scan_plan(&t);
let sbom = step_for(&plan, ScanType::Sbom).expect("sbom planned");
assert_eq!(sbom.artifact_id, fw_id);
assert!(step_for(&plan, ScanType::FirmwareStatic).is_some());
assert!(!plan.has(ScanType::Dast));
}
#[test]
fn plc_plans_only_control_logic() {
let plc = Artifact::plc_project("p.xml", PlcFormat::PlcopenXml);
let t = target(TargetType::PlcSps, vec![plc]);
let plan = build_scan_plan(&t);
assert_eq!(plan.steps.len(), 1);
assert_eq!(plan.steps[0].scan_type, ScanType::PlcControlLogic);
assert_eq!(plan.steps[0].phase, ScanPhase::PlcAnalysis);
}
#[test]
fn disabled_scan_is_dropped_and_off_by_default_can_be_enabled() {
let mut t = target(TargetType::WebApp, vec![Artifact::git_repo("u", "main")]);
t.scan_config.disabled_scans = vec![ScanType::Lint];
// CodeReview is off by default for web; enable it explicitly.
t.scan_config.enabled_scans = vec![ScanType::CodeReview];
let plan = build_scan_plan(&t);
assert!(!plan.has(ScanType::Lint));
assert!(plan.has(ScanType::CodeReview));
assert!(plan.has(ScanType::Sast));
}
#[test]
fn no_code_artifact_yields_empty_plan_for_web() {
let t = target(TargetType::WebApp, vec![]);
let plan = build_scan_plan(&t);
assert!(plan.is_empty());
}
}
+18 -198
View File
@@ -7,18 +7,11 @@ use crate::agent::ComplianceAgent;
use crate::database::Database;
use crate::error::AgentError;
/// Default tenant the scheduler runs against when neither the tenant
/// registry nor `SCHEDULER_TENANT_IDS` are configured. Matches the
/// dev-injector default so a bare `cargo run` has the scheduler
/// scanning whatever lives in `<prefix>_dev`.
/// Default tenant the scheduler runs against when `SCHEDULER_TENANT_IDS`
/// isn't set. Matches the dev-injector default so a bare `cargo run` has
/// the scheduler scanning whatever lives in `<prefix>_dev`.
const DEFAULT_SCHEDULER_TENANT_ID: &str = "dev";
/// Request timeout when fetching the live tenant list from the
/// registry. Kept short — if the registry is slow we'd rather fall
/// back to env-configured ids and finish the tick than block the
/// scheduler loop.
const REGISTRY_FETCH_TIMEOUT_SECS: u64 = 5;
pub async fn start_scheduler(agent: &ComplianceAgent) -> Result<(), AgentError> {
let sched = JobScheduler::new()
.await
@@ -31,12 +24,7 @@ pub async fn start_scheduler(agent: &ComplianceAgent) -> Result<(), AgentError>
let agent = scan_agent.clone();
Box::pin(async move {
tracing::info!("Scheduled scan triggered");
let tenants = scheduler_tenants(&agent).await;
tracing::debug!(
tenant_count = tenants.len(),
"Scheduled scan: tenants resolved"
);
for tenant_id in tenants {
for tenant_id in scheduler_tenants() {
scan_all_repos(&agent, &tenant_id).await;
}
})
@@ -54,12 +42,7 @@ pub async fn start_scheduler(agent: &ComplianceAgent) -> Result<(), AgentError>
let agent = cve_agent.clone();
Box::pin(async move {
tracing::info!("CVE monitor triggered");
let tenants = scheduler_tenants(&agent).await;
tracing::debug!(
tenant_count = tenants.len(),
"CVE monitor: tenants resolved"
);
for tenant_id in tenants {
for tenant_id in scheduler_tenants() {
monitor_cves(&agent, &tenant_id).await;
}
})
@@ -75,14 +58,9 @@ pub async fn start_scheduler(agent: &ComplianceAgent) -> Result<(), AgentError>
.await
.map_err(|e| AgentError::Scheduler(format!("Failed to start scheduler: {e}")))?;
let tenants = scheduler_tenants(agent).await;
let source = if agent.config.tenant_registry_url.is_some() {
"tenant-registry (env fallback)"
} else {
"env (SCHEDULER_TENANT_IDS)"
};
let tenants = scheduler_tenants();
tracing::info!(
"Scheduler started: scans='{}', CVE monitor='{}', tenant source={source}, tenants={tenants:?}",
"Scheduler started: scans='{}', CVE monitor='{}', tenants={tenants:?}",
agent.config.scan_schedule,
agent.config.cve_monitor_schedule,
);
@@ -93,40 +71,10 @@ pub async fn start_scheduler(agent: &ComplianceAgent) -> Result<(), AgentError>
}
}
/// Tenants the scheduler iterates each tick.
///
/// Resolution order:
/// 1. **Tenant registry** at `agent.config.tenant_registry_url`
/// (`GET /v1/tenants`). Fresh on every tick — picks up newly
/// provisioned tenants without an agent restart.
/// 2. **`SCHEDULER_TENANT_IDS`** env (comma-separated) — fallback when
/// the registry is unreachable, the response is malformed, or no
/// registry URL is configured.
/// 3. **`DEFAULT_SCHEDULER_TENANT_ID`** (`"dev"`) — last-ditch fallback
/// so the scheduler keeps doing something useful in dev.
///
/// We never panic out of this function — the scheduler must keep
/// firing even if the registry is offline.
async fn scheduler_tenants(agent: &ComplianceAgent) -> Vec<String> {
if let Some(url) = agent.config.tenant_registry_url.as_deref() {
match fetch_tenants_from_registry(&agent.http, url).await {
Ok(v) if !v.is_empty() => return v,
Ok(_) => {
tracing::warn!("tenant-registry returned empty list; falling back to env");
}
Err(e) => {
tracing::warn!(
url = %url,
error = %e,
"tenant-registry fetch failed; falling back to env"
);
}
}
}
tenants_from_env()
}
fn tenants_from_env() -> Vec<String> {
/// Tenants the scheduler iterates each tick. From `SCHEDULER_TENANT_IDS`
/// (comma-separated), or `DEFAULT_SCHEDULER_TENANT_ID` if unset. M7.2-D
/// will replace this with a pull from the tenant-registry.
fn scheduler_tenants() -> Vec<String> {
std::env::var("SCHEDULER_TENANT_IDS")
.ok()
.map(|s| {
@@ -140,134 +88,6 @@ fn tenants_from_env() -> Vec<String> {
.unwrap_or_else(|| vec![DEFAULT_SCHEDULER_TENANT_ID.to_string()])
}
/// Shape we accept from the registry. Liberal in what we accept:
/// the registry can return any field shape as long as either `id` or
/// `tenant_id` is present. Other fields are ignored.
#[derive(serde::Deserialize)]
struct RegistryTenant {
#[serde(alias = "tenant_id")]
id: String,
/// Filter out non-running tenants if status is present. Missing
/// status defaults to "active" so older registry deployments keep
/// working.
#[serde(default = "default_status")]
status: String,
}
fn default_status() -> String {
"active".to_string()
}
#[derive(serde::Deserialize)]
struct RegistryListResponse {
data: Vec<RegistryTenant>,
}
async fn fetch_tenants_from_registry(
http: &reqwest::Client,
base_url: &str,
) -> Result<Vec<String>, String> {
let url = format!("{}/v1/tenants", base_url.trim_end_matches('/'));
let resp = http
.get(&url)
.timeout(std::time::Duration::from_secs(REGISTRY_FETCH_TIMEOUT_SECS))
.send()
.await
.map_err(|e| format!("request failed: {e}"))?;
if !resp.status().is_success() {
return Err(format!("registry returned {}", resp.status()));
}
let body: RegistryListResponse = resp
.json()
.await
.map_err(|e| format!("invalid JSON: {e}"))?;
Ok(filter_active(body.data))
}
/// Frozen/Archived tenants don't need scheduled scans; the M7.1
/// status gate would 402/410 anyway. Skip them so we don't waste
/// cycles. Active / trial / demo / anything-else-unknown all run.
fn filter_active(rows: Vec<RegistryTenant>) -> Vec<String> {
rows.into_iter()
.filter(|t| !matches!(t.status.as_str(), "frozen" | "archived"))
.map(|t| t.id)
.collect()
}
#[cfg(test)]
mod tests {
use super::*;
fn tenant(id: &str, status: &str) -> RegistryTenant {
RegistryTenant {
id: id.to_string(),
status: status.to_string(),
}
}
#[test]
fn filter_active_keeps_running_skips_frozen_archived() {
let rows = vec![
tenant("a", "active"),
tenant("b", "trial"),
tenant("c", "demo"),
tenant("d", "frozen"),
tenant("e", "archived"),
tenant("f", "weird-but-not-known-dead"),
];
let out = filter_active(rows);
assert_eq!(out, vec!["a", "b", "c", "f"]);
}
#[test]
fn deserialize_registry_response_accepts_id_or_tenant_id() {
let body = r#"{"data":[
{"id":"a","status":"active"},
{"tenant_id":"b","status":"trial"},
{"id":"c"}
]}"#;
let parsed: RegistryListResponse = serde_json::from_str(body).unwrap();
assert_eq!(parsed.data.len(), 3);
assert_eq!(parsed.data[0].id, "a");
assert_eq!(parsed.data[1].id, "b");
assert_eq!(parsed.data[2].id, "c");
// Default status for the third entry should be "active"
assert_eq!(parsed.data[2].status, "active");
}
/// Combined into a single test: cargo runs tests in parallel and
/// env vars are process-global, so two separate tests touching
/// `SCHEDULER_TENANT_IDS` race each other. Doing both checks in
/// one test keeps them in a deterministic order.
#[test]
fn tenants_from_env_resolution() {
std::env::remove_var("SCHEDULER_TENANT_IDS");
assert_eq!(
tenants_from_env(),
vec![DEFAULT_SCHEDULER_TENANT_ID.to_string()],
"unset → default"
);
std::env::set_var("SCHEDULER_TENANT_IDS", "acme, globex ,,hello");
let out = tenants_from_env();
std::env::remove_var("SCHEDULER_TENANT_IDS");
assert_eq!(
out,
vec!["acme", "globex", "hello"],
"splits + trims + drops empty"
);
std::env::set_var("SCHEDULER_TENANT_IDS", "");
let out = tenants_from_env();
std::env::remove_var("SCHEDULER_TENANT_IDS");
assert_eq!(
out,
vec![DEFAULT_SCHEDULER_TENANT_ID.to_string()],
"empty → default"
);
}
}
/// Resolve the per-tenant database. Logs and returns `None` on failure
/// so the loop in the caller can continue with other tenants.
async fn tenant_db(agent: &ComplianceAgent, tenant_id: &str) -> Option<Database> {
@@ -288,25 +108,25 @@ async fn scan_all_repos(agent: &ComplianceAgent, tenant_id: &str) {
None => return,
};
let cursor = match db.onboarded_targets().find(doc! {}).await {
let cursor = match db.repositories().find(doc! {}).await {
Ok(c) => c,
Err(e) => {
tracing::error!("Failed to list targets for tenant '{tenant_id}': {e}");
tracing::error!("Failed to list repos for tenant '{tenant_id}': {e}");
return;
}
};
let targets: Vec<_> = cursor.filter_map(|r| async { r.ok() }).collect().await;
let repos: Vec<_> = cursor.filter_map(|r| async { r.ok() }).collect().await;
for target in targets {
let target_id = target.id.map(|id| id.to_hex()).unwrap_or_default();
for repo in repos {
let repo_id = repo.id.map(|id| id.to_hex()).unwrap_or_default();
if let Err(e) = agent
.run_target_scan(tenant_id, &target_id, ScanTrigger::Scheduled)
.run_scan(tenant_id, &repo_id, ScanTrigger::Scheduled)
.await
{
tracing::error!(
"Scheduled scan failed for {} (tenant '{tenant_id}'): {e}",
target.name
repo.name
);
}
}
+2 -5
View File
@@ -25,9 +25,8 @@ impl TestServer {
let mongodb_uri = std::env::var("TEST_MONGODB_URI")
.unwrap_or_else(|_| "mongodb://root:example@localhost:27017/?authSource=admin".into());
// Unique db-name prefix per run. Must fit the pool's 30-char cap
// (`<prefix>_<32 hex>` <= 63), so use a 16-hex-char suffix.
let db_name = format!("t_{}", &uuid::Uuid::new_v4().simple().to_string()[..16]);
// Unique database name per test run to avoid collisions
let db_name = format!("test_{}", uuid::Uuid::new_v4().simple());
let db_pool = DatabasePool::connect(&mongodb_uri, &db_name)
.await
@@ -45,7 +44,6 @@ impl TestServer {
scan_schedule: String::new(),
cve_monitor_schedule: String::new(),
git_clone_base_path: "/tmp/compliance-scanner-tests/repos".into(),
artifact_store_base_path: "/tmp/compliance-scanner-tests/artifacts".into(),
ssh_key_path: "/tmp/compliance-scanner-tests/ssh/id_ed25519".into(),
github_token: None,
github_webhook_secret: None,
@@ -70,7 +68,6 @@ impl TestServer {
pentest_imap_password: None,
admin_api_token: None,
tenant_registry_url: None,
unified_pipeline: false,
};
let agent = ComplianceAgent::new(config, db_pool);
@@ -2,6 +2,5 @@ mod cascade_delete;
mod dast;
mod findings;
mod health;
mod onboarding;
mod repositories;
mod stats;
@@ -1,115 +0,0 @@
use crate::common::TestServer;
use serde_json::json;
#[tokio::test]
async fn create_list_and_applicable_scans() {
let server = TestServer::start().await;
// Initially empty.
let resp = server.get("/api/v1/targets").await;
assert_eq!(resp.status(), 200);
let body: serde_json::Value = resp.json().await.unwrap();
assert_eq!(body["data"].as_array().unwrap().len(), 0);
// Create a web-app target with a git repo + a live URL.
let resp = server
.post(
"/api/v1/targets",
&json!({
"name": "acme-web",
"target_type": "web_app",
"artifacts": [
{ "kind": "git_repo", "source_ref": "https://git/acme.git", "branch": "main" },
{ "kind": "live_url", "source_ref": "https://acme.example.com" }
]
}),
)
.await;
assert_eq!(resp.status(), 200);
let body: serde_json::Value = resp.json().await.unwrap();
let id = body["data"]["_id"]["$oid"].as_str().unwrap().to_string();
assert!(!id.is_empty());
assert_eq!(body["data"]["artifacts"].as_array().unwrap().len(), 2);
// List returns it.
let resp = server.get("/api/v1/targets").await;
let body: serde_json::Value = resp.json().await.unwrap();
assert_eq!(body["data"].as_array().unwrap().len(), 1);
// Applicable scans: SAST present + DAST offered (live URL present), pentest supported.
let resp = server
.get(&format!("/api/v1/targets/{id}/applicable-scans"))
.await;
assert_eq!(resp.status(), 200);
let body: serde_json::Value = resp.json().await.unwrap();
let scans = body["data"]["scans"].as_array().unwrap();
let names: Vec<&str> = scans.iter().filter_map(|s| s["scan"].as_str()).collect();
assert!(names.contains(&"sast"));
assert!(names.contains(&"dast"));
assert_eq!(body["data"]["pentest_supported"], true);
server.cleanup().await;
}
#[tokio::test]
async fn detect_classifies_a_plc_target() {
let server = TestServer::start().await;
// A PLC project artifact is a strong kind-based signal.
let resp = server
.post(
"/api/v1/targets",
&json!({
"name": "line-controller",
"target_type": "backend_service", // deliberately wrong; detect should suggest PLC
"artifacts": [
{ "kind": "plc_project", "source_ref": "line.xml", "plc_format": "plcopen_xml" }
]
}),
)
.await;
let body: serde_json::Value = resp.json().await.unwrap();
let id = body["data"]["_id"]["$oid"].as_str().unwrap().to_string();
let resp = server
.post(&format!("/api/v1/targets/{id}/detect"), &json!({}))
.await;
assert_eq!(resp.status(), 200);
let body: serde_json::Value = resp.json().await.unwrap();
assert_eq!(body["data"]["classification"]["suggested"], "plc_sps");
server.cleanup().await;
}
#[tokio::test]
async fn add_artifact_and_delete_target() {
let server = TestServer::start().await;
let resp = server
.post(
"/api/v1/targets",
&json!({ "name": "svc", "target_type": "backend_service" }),
)
.await;
let body: serde_json::Value = resp.json().await.unwrap();
let id = body["data"]["_id"]["$oid"].as_str().unwrap().to_string();
// Attach a git repo.
let resp = server
.post(
&format!("/api/v1/targets/{id}/artifacts"),
&json!({ "kind": "git_repo", "source_ref": "https://git/svc.git" }),
)
.await;
assert_eq!(resp.status(), 200);
let body: serde_json::Value = resp.json().await.unwrap();
assert_eq!(body["data"]["artifacts"].as_array().unwrap().len(), 1);
// Delete it.
let resp = server.delete(&format!("/api/v1/targets/{id}")).await;
assert_eq!(resp.status(), 200);
let resp = server.get(&format!("/api/v1/targets/{id}")).await;
assert_eq!(resp.status(), 404);
server.cleanup().await;
}
@@ -1,156 +0,0 @@
// Integration tests for the onboarding backfill migration.
//
// Requires MongoDB (set TEST_MONGODB_URI if not at the default).
// Not run in CI (which is `--lib` only) — run locally:
// cargo test -p compliance-agent --test e2e migration
use compliance_agent::database::{Database, DatabasePool};
use compliance_agent::migrate::onboarding;
use compliance_core::models::{
ArtifactKind, DastTarget, DastTargetType, TargetType, TrackedRepository,
};
use mongodb::bson::{doc, Document};
async fn fresh_db() -> (DatabasePool, String, Database) {
let uri = std::env::var("TEST_MONGODB_URI")
.unwrap_or_else(|_| "mongodb://root:example@localhost:27017/?authSource=admin".into());
// Prefix must fit the pool's 30-char cap (`<prefix>_<32 hex>` <= 63).
let prefix = format!("t_{}", &uuid::Uuid::new_v4().simple().to_string()[..16]);
let pool = DatabasePool::connect(&uri, &prefix)
.await
.expect("connect mongo");
let db = pool.for_tenant_id("t1").await.expect("tenant db");
(pool, prefix, db)
}
async fn cleanup(pool: &DatabasePool, prefix: &str) {
if let Ok(names) = pool.client().list_database_names().await {
for n in names {
if n.starts_with(prefix) {
pool.client().database(&n).drop().await.ok();
}
}
}
}
#[tokio::test]
async fn backfill_folds_relinks_is_idempotent_and_reversible() {
let (pool, prefix, db) = fresh_db().await;
// Seed a repo.
let repo = TrackedRepository::new("acme".into(), "https://git/acme.git".into());
let repo_id = db
.repositories()
.insert_one(repo)
.await
.expect("insert repo")
.inserted_id
.as_object_id()
.expect("repo oid");
// A DAST target linked to the repo (folds + promotes to WebApp + relinks).
let mut linked = DastTarget::new(
"acme-web".into(),
"https://acme.example.com".into(),
DastTargetType::WebApp,
);
linked.repo_id = Some(repo_id.to_hex());
let linked_id = db
.dast_targets()
.insert_one(linked)
.await
.expect("insert linked dast")
.inserted_id
.as_object_id()
.expect("linked oid");
// A repo-less DAST target (standalone).
let standalone = DastTarget::new(
"acme-api".into(),
"https://api.acme.com".into(),
DastTargetType::RestApi,
);
let standalone_id = db
.dast_targets()
.insert_one(standalone)
.await
.expect("insert standalone dast")
.inserted_id
.as_object_id()
.expect("standalone oid");
// A DAST scan run pointing at the linked target — should be relinked to the repo.
db.collection_named::<Document>("dast_scan_runs")
.insert_one(doc! { "target_id": linked_id.to_hex(), "status": "completed" })
.await
.expect("insert dast run");
// --- Backfill ---
assert!(!onboarding::already_applied(&db).await.unwrap());
let report = onboarding::backfill_onboarded_targets(&db, false)
.await
.expect("backfill");
assert_eq!(report.repos_migrated, 1);
assert_eq!(report.dast_targets_folded, 1);
assert_eq!(report.dast_targets_standalone, 1);
assert!(onboarding::already_applied(&db).await.unwrap());
// Repo target: preserved _id, has git + folded live-url, promoted to WebApp.
let repo_target = db
.onboarded_targets()
.find_one(doc! { "_id": repo_id })
.await
.unwrap()
.expect("repo target");
assert!(repo_target.has(ArtifactKind::GitRepo));
assert!(repo_target.has(ArtifactKind::LiveUrl));
assert_eq!(repo_target.target_type, TargetType::WebApp);
// Standalone target: preserved _id, live-url, backend service.
let standalone_target = db
.onboarded_targets()
.find_one(doc! { "_id": standalone_id })
.await
.unwrap()
.expect("standalone target");
assert!(standalone_target.has(ArtifactKind::LiveUrl));
assert_eq!(standalone_target.target_type, TargetType::BackendService);
// The DAST run was relinked from the old dast id to the repo (unified) id.
let run = db
.collection_named::<Document>("dast_scan_runs")
.find_one(doc! {})
.await
.unwrap()
.expect("run");
assert_eq!(run.get_str("target_id").unwrap(), repo_id.to_hex());
// --- Idempotent: re-run migrates nothing new ---
let again = onboarding::backfill_onboarded_targets(&db, false)
.await
.expect("backfill again");
assert_eq!(again.repos_migrated, 0);
assert_eq!(again.dast_targets_folded, 0);
assert_eq!(again.dast_targets_standalone, 0);
assert!(again.skipped_existing >= 2);
// --- Revert: onboarded targets gone, relink undone, marker cleared ---
onboarding::revert(&db).await.expect("revert");
assert_eq!(
db.onboarded_targets()
.count_documents(doc! {})
.await
.unwrap(),
0
);
let run_after = db
.collection_named::<Document>("dast_scan_runs")
.find_one(doc! {})
.await
.unwrap()
.expect("run");
assert_eq!(run_after.get_str("target_id").unwrap(), linked_id.to_hex());
assert!(!onboarding::already_applied(&db).await.unwrap());
cleanup(&pool, &prefix).await;
}
@@ -7,4 +7,3 @@
// Or nightly: (via CI with MongoDB service container)
mod api;
mod migration;
-8
View File
@@ -24,9 +24,6 @@ pub struct AgentConfig {
pub scan_schedule: String,
pub cve_monitor_schedule: String,
pub git_clone_base_path: String,
/// Base directory for content-addressed artifact blobs and per-run working
/// dirs (`<base>/blobs/<sha[0:2]>/<sha>`, `<base>/work/<target>/<artifact>/`).
pub artifact_store_base_path: String,
pub ssh_key_path: String,
pub keycloak_url: Option<String>,
pub keycloak_realm: Option<String>,
@@ -49,11 +46,6 @@ pub struct AgentConfig {
/// of tenants to iterate. When `None` or unreachable, scheduler
/// falls back to `SCHEDULER_TENANT_IDS` env (M7.2-C).
pub tenant_registry_url: Option<String>,
/// When true, `run_scan` dispatches to the unified `run_target` pipeline
/// (reads `onboarded_targets`) instead of the legacy repository pipeline.
/// Env `UNIFIED_PIPELINE`. Defaults on; set `UNIFIED_PIPELINE=0` to use the
/// legacy repository pipeline.
pub unified_pipeline: bool,
}
#[derive(Clone, Debug, Serialize, Deserialize)]
-1
View File
@@ -2,7 +2,6 @@ pub mod config;
pub mod db;
pub mod error;
pub mod models;
pub mod scan_matrix;
#[cfg(feature = "telemetry")]
pub mod telemetry;
pub mod tenant;
-7
View File
@@ -9,7 +9,6 @@ pub mod issue;
pub mod mcp;
pub mod mcp_token;
pub mod notification;
pub mod onboarding;
pub mod pentest;
pub mod repository;
pub mod sbom;
@@ -32,12 +31,6 @@ pub use issue::{IssueStatus, TrackerIssue, TrackerType};
pub use mcp::{McpServerConfig, McpServerStatus, McpTransport};
pub use mcp_token::{McpToken, McpTokenView};
pub use notification::{CveNotification, NotificationSeverity, NotificationStatus};
pub use onboarding::{
default_compliance_profile, Artifact, ArtifactAuth, ArtifactKind, Classification,
ComplianceFramework, ComplianceProfile, DetectedFact, ExternalRef, ExternalSystem,
GitArtifactConfig, IssueTrackerConfig, OnboardedTarget, PlcArtifactConfig, PlcFormat,
TargetScanConfig, TargetType, TargetTypeCandidate, WebArtifactConfig,
};
pub use pentest::{
AttackChainNode, AttackNodeStatus, AuthMode, CodeContextHint, Environment, IdentityProvider,
PentestAuthConfig, PentestConfig, PentestEvent, PentestMessage, PentestSession, PentestStats,
-753
View File
@@ -1,753 +0,0 @@
//! The unified onboarding model.
//!
//! An [`OnboardedTarget`] is the single source of truth for anything the scanner
//! can analyze. It records *what kind of software* the target is ([`TargetType`]),
//! the concrete [`Artifact`]s that were provided for it (a git repo, a firmware
//! image, a live URL, a PLC project, ...), the classifier's verdict, and the scan
//! configuration. It replaces the older git-only `TrackedRepository` and the
//! standalone `DastTarget`, both of which fold into this type as artifacts.
use std::collections::HashMap;
use chrono::{DateTime, Utc};
use serde::{Deserialize, Serialize};
use super::dast::{DastAuthConfig, DastTargetType};
use super::issue::TrackerType;
use super::pentest::{Environment, PentestConfig, PentestStrategy};
use super::scan::ScanType;
/// The family of software a target belongs to.
///
/// Targets look endlessly varied but fall into a small enumerable set classified
/// by where the analyzable signal lives. This drives the scan-applicability
/// matrix and the onboarding wizard's type selection.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum TargetType {
/// Browser-facing web application (front end + server).
WebApp,
/// Headless backend service / API (REST, GraphQL, gRPC).
BackendService,
/// Desktop application (Windows/macOS/Linux GUI or CLI binary).
DesktopApp,
/// Android application (APK / AAB).
AndroidApp,
/// iOS application (IPA).
IosApp,
/// Bare-metal embedded firmware (no operating system).
FirmwareBareMetal,
/// Embedded firmware running on an RTOS (Zephyr, FreeRTOS, ...).
FirmwareRtos,
/// Embedded Linux built with Yocto / OpenEmbedded (BSP + image).
EmbeddedLinuxYocto,
/// Programmable logic controller software (IEC 61131-3, PLCopen / SPS).
PlcSps,
}
impl std::fmt::Display for TargetType {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::WebApp => write!(f, "web_app"),
Self::BackendService => write!(f, "backend_service"),
Self::DesktopApp => write!(f, "desktop_app"),
Self::AndroidApp => write!(f, "android_app"),
Self::IosApp => write!(f, "ios_app"),
Self::FirmwareBareMetal => write!(f, "firmware_bare_metal"),
Self::FirmwareRtos => write!(f, "firmware_rtos"),
Self::EmbeddedLinuxYocto => write!(f, "embedded_linux_yocto"),
Self::PlcSps => write!(f, "plc_sps"),
}
}
}
/// The kind of artifact provided for a target.
///
/// Which scans are possible is a function of the target type *and* which of
/// these are present (SAST needs code, DAST needs a running URL, and so on).
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum ArtifactKind {
/// A git repository (cloned for static analysis).
GitRepo,
/// A source archive (zip / tarball) with no live git remote.
SourceArchive,
/// A firmware image or binary blob.
FirmwareImage,
/// A mobile package: Android APK/AAB or iOS IPA.
MobilePackage,
/// An OCI/Docker container image reference.
ContainerImage,
/// A reachable running instance (base URL / endpoint) for dynamic testing.
LiveUrl,
/// A PLC project: PLCopen XML or Structured Text source.
PlcProject,
/// Free-form plaintext describing the target (feeds classification only).
PlaintextDescription,
}
impl std::fmt::Display for ArtifactKind {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::GitRepo => write!(f, "git_repo"),
Self::SourceArchive => write!(f, "source_archive"),
Self::FirmwareImage => write!(f, "firmware_image"),
Self::MobilePackage => write!(f, "mobile_package"),
Self::ContainerImage => write!(f, "container_image"),
Self::LiveUrl => write!(f, "live_url"),
Self::PlcProject => write!(f, "plc_project"),
Self::PlaintextDescription => write!(f, "plaintext_description"),
}
}
}
/// Credentials attached to an artifact.
///
/// This folds both `TrackedRepository`'s git auth (`auth_token` / `auth_username`
/// / SSH key) and `DastAuthConfig`'s HTTP auth (form / bearer / cookie) into one
/// shape so a single artifact carries whatever it needs to be fetched or probed.
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct ArtifactAuth {
/// Auth method: `none` | `token` | `basic` | `bearer` | `cookie` | `form` | `ssh`.
#[serde(default)]
pub method: String,
/// Username (git user, basic-auth user, or `x-access-token` for PATs).
pub username: Option<String>,
/// The secret credential: PAT, password, or bearer token. Encrypted at rest.
pub secret: Option<String>,
/// Path to an SSH private key for git-over-SSH.
pub ssh_key_path: Option<String>,
/// Login URL for form-based authentication.
pub login_url: Option<String>,
/// Extra headers to send when authenticating / probing.
pub headers: Option<HashMap<String, String>>,
}
impl From<DastAuthConfig> for ArtifactAuth {
fn from(c: DastAuthConfig) -> Self {
Self {
method: c.method,
username: c.username,
// Prefer a bearer token; otherwise fall back to the password.
secret: c.token.or(c.password),
ssh_key_path: None,
login_url: c.login_url,
headers: c.headers,
}
}
}
/// Git-specific configuration for a [`ArtifactKind::GitRepo`] artifact.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct GitArtifactConfig {
/// Branch to scan.
pub default_branch: String,
/// Commit SHA of the last completed scan (change-detection watermark).
pub last_scanned_commit: Option<String>,
/// Local clone path once the repo has been fetched.
pub local_path: Option<String>,
}
impl GitArtifactConfig {
/// Config for a fresh git artifact on the given branch.
pub fn on_branch(branch: impl Into<String>) -> Self {
Self {
default_branch: branch.into(),
last_scanned_commit: None,
local_path: None,
}
}
}
impl Default for GitArtifactConfig {
fn default() -> Self {
Self::on_branch("main")
}
}
/// Dynamic-analysis configuration for a [`ArtifactKind::LiveUrl`] artifact.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct WebArtifactConfig {
/// Whether the endpoint is a web app, REST API, or GraphQL API.
pub target_kind: DastTargetType,
/// URL paths to exclude from crawling / scanning.
#[serde(default)]
pub excluded_paths: Vec<String>,
/// Maximum crawl depth.
pub max_crawl_depth: u32,
/// Rate limit in requests per second.
pub rate_limit: u32,
/// Whether destructive methods (DELETE / PUT) are permitted.
#[serde(default)]
pub allow_destructive: bool,
}
impl Default for WebArtifactConfig {
fn default() -> Self {
Self {
target_kind: DastTargetType::WebApp,
excluded_paths: Vec::new(),
max_crawl_depth: 3,
rate_limit: 10,
allow_destructive: false,
}
}
}
/// The source format of a PLC project artifact.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum PlcFormat {
/// PLCopen XML project export.
PlcopenXml,
/// IEC 61131-3 Structured Text source.
StructuredText,
}
/// PLC-specific configuration for a [`ArtifactKind::PlcProject`] artifact.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct PlcArtifactConfig {
/// The project source format.
pub format: PlcFormat,
}
/// A single fact discovered about a target by ingest or classification
/// (e.g. `language=rust`, `build_system=cmake`, `mcu=stm32f429`).
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct DetectedFact {
/// The fact name.
pub key: String,
/// The fact value.
pub value: String,
/// What produced the fact (e.g. `tramiton`, `language-fingerprint`).
pub source: String,
}
impl DetectedFact {
/// Build a fact from its parts.
pub fn new(
key: impl Into<String>,
value: impl Into<String>,
source: impl Into<String>,
) -> Self {
Self {
key: key.into(),
value: value.into(),
source: source.into(),
}
}
}
/// One concrete thing provided for a target: code, a binary, a URL, etc.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct Artifact {
/// Stable per-artifact id (UUID v4) — scan steps reference this.
pub id: String,
/// What kind of artifact this is.
pub kind: ArtifactKind,
/// The source reference: git URL, blob id, live URL, or image ref.
pub source_ref: String,
/// Optional human-friendly label.
pub display_name: Option<String>,
/// Content-addressed storage path once ingested (blobs only).
pub stored_path: Option<String>,
/// SHA-256 of the ingested content (git artifacts store the head SHA).
pub content_hash: Option<String>,
/// Size of the stored blob in bytes.
pub size_bytes: Option<u64>,
/// Credentials for fetching or probing this artifact.
pub auth: Option<ArtifactAuth>,
/// Git configuration (present for [`ArtifactKind::GitRepo`]).
pub git: Option<GitArtifactConfig>,
/// Dynamic-analysis configuration (present for [`ArtifactKind::LiveUrl`]).
pub web: Option<WebArtifactConfig>,
/// PLC configuration (present for [`ArtifactKind::PlcProject`]).
pub plc: Option<PlcArtifactConfig>,
/// Facts discovered about this artifact by ingest / classification.
#[serde(default)]
pub detected: Vec<DetectedFact>,
/// When this artifact was last ingested.
#[serde(default, with = "super::serde_helpers::opt_bson_datetime")]
pub ingested_at: Option<DateTime<Utc>>,
}
impl Artifact {
/// A bare artifact of the given kind and source reference.
fn bare(kind: ArtifactKind, source_ref: impl Into<String>) -> Self {
Self {
id: uuid::Uuid::new_v4().to_string(),
kind,
source_ref: source_ref.into(),
display_name: None,
stored_path: None,
content_hash: None,
size_bytes: None,
auth: None,
git: None,
web: None,
plc: None,
detected: Vec::new(),
ingested_at: None,
}
}
/// A git-repository artifact tracking the given branch.
pub fn git_repo(url: impl Into<String>, branch: impl Into<String>) -> Self {
let mut a = Self::bare(ArtifactKind::GitRepo, url);
a.git = Some(GitArtifactConfig::on_branch(branch));
a
}
/// A live-URL artifact with default crawl settings.
pub fn live_url(url: impl Into<String>) -> Self {
let mut a = Self::bare(ArtifactKind::LiveUrl, url);
a.web = Some(WebArtifactConfig::default());
a
}
/// A firmware-image artifact referenced by name (blob ingested later).
pub fn firmware_image(source_ref: impl Into<String>) -> Self {
Self::bare(ArtifactKind::FirmwareImage, source_ref)
}
/// A source-archive artifact referenced by name (blob ingested later).
pub fn source_archive(source_ref: impl Into<String>) -> Self {
Self::bare(ArtifactKind::SourceArchive, source_ref)
}
/// A mobile-package artifact (APK/AAB/IPA) referenced by name.
pub fn mobile_package(source_ref: impl Into<String>) -> Self {
Self::bare(ArtifactKind::MobilePackage, source_ref)
}
/// A container-image artifact referenced by OCI ref.
pub fn container_image(source_ref: impl Into<String>) -> Self {
Self::bare(ArtifactKind::ContainerImage, source_ref)
}
/// A PLC-project artifact in the given format.
pub fn plc_project(source_ref: impl Into<String>, format: PlcFormat) -> Self {
let mut a = Self::bare(ArtifactKind::PlcProject, source_ref);
a.plc = Some(PlcArtifactConfig { format });
a
}
/// A plaintext-description artifact (classification input only).
pub fn plaintext(text: impl Into<String>) -> Self {
Self::bare(ArtifactKind::PlaintextDescription, text)
}
}
/// One ranked candidate produced by the classifier.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TargetTypeCandidate {
/// The candidate target type.
pub target_type: TargetType,
/// Confidence in `[0.0, 1.0]`.
pub confidence: f32,
/// Why this candidate was proposed.
pub rationale: String,
}
/// The classifier's verdict for a target: a suggested type plus ranked
/// alternatives and the facts the decision rested on.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct Classification {
/// The top-ranked target type.
pub suggested: TargetType,
/// All candidates, sorted by descending confidence.
#[serde(default)]
pub candidates: Vec<TargetTypeCandidate>,
/// Facts gathered during classification.
#[serde(default)]
pub facts: Vec<DetectedFact>,
/// Which classifiers contributed (e.g. `["tramiton", "language-fingerprint"]`).
#[serde(default)]
pub detected_by: Vec<String>,
/// When classification ran.
#[serde(with = "super::serde_helpers::bson_datetime")]
pub detected_at: DateTime<Utc>,
/// Whether a human confirmed the suggestion.
#[serde(default)]
pub confirmed: bool,
}
/// Issue-tracker linkage, migrated from `TrackedRepository`'s `tracker_*` fields.
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct IssueTrackerConfig {
/// The tracker platform.
pub tracker_type: Option<TrackerType>,
/// Tracker owner / organization.
pub owner: Option<String>,
/// Tracker repository / project.
pub repo: Option<String>,
/// Per-target tracker access token.
pub token: Option<String>,
}
/// How a target should be scanned.
///
/// `enabled_scans` / `disabled_scans` override the scan-applicability matrix
/// defaults; the pentest and tracker blocks reuse the existing wizard config.
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct TargetScanConfig {
/// Scans explicitly turned on (empty means "use matrix defaults").
#[serde(default)]
pub enabled_scans: Vec<ScanType>,
/// Scans explicitly turned off.
#[serde(default)]
pub disabled_scans: Vec<ScanType>,
/// Target environment (gates destructive / active testing).
#[serde(default)]
pub environment: Environment,
/// Whether destructive tests are permitted for this target.
#[serde(default)]
pub allow_destructive: bool,
/// Pentest strategy selector.
pub strategy: Option<PentestStrategy>,
/// Full pentest wizard configuration.
pub pentest: Option<PentestConfig>,
/// Issue-tracker linkage.
pub issue_tracker: Option<IssueTrackerConfig>,
}
/// A sibling product in the company suite that may already hold authoritative
/// data for a target. compliance-scanner reconciles with these rather than
/// recomputing what they already know.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum ExternalSystem {
/// Reproducible-build & firmware compliance engine (build plan, SBOM, VEX,
/// attestation).
Tramiton,
/// Code assistant (downstream remediation consumer).
Werkpilot,
/// Compliance-controls RAG (atomic controls derived from laws).
BreakpilotCompliance,
}
impl std::fmt::Display for ExternalSystem {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::Tramiton => write!(f, "tramiton"),
Self::Werkpilot => write!(f, "werkpilot"),
Self::BreakpilotCompliance => write!(f, "breakpilot_compliance"),
}
}
}
/// A link from this target to a record in a sibling product, used to reconcile
/// existing evidence instead of recomputing it.
///
/// For tramiton, `project_id` is the shared cross-product key and
/// `subject_sha256` matches a firmware artifact's [`Artifact::content_hash`]
/// (which equals tramiton's `Artifact.sha256`).
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ExternalRef {
/// Which sibling product this reference points at.
pub system: ExternalSystem,
/// The sibling product's project identifier, if known.
pub project_id: Option<String>,
/// Content digest of the subject artifact (firmware sha256), if known.
pub subject_sha256: Option<String>,
/// Reconciliation status: `linked` | `reconciled` | `unavailable`.
#[serde(default)]
pub status: String,
/// Opaque, offline-verifiable entitlement grant (e.g. tramiton's signed
/// `LicenseGrant`), if the tenant provided one.
pub license_grant: Option<String>,
/// When evidence was last reconciled from this system.
#[serde(default, with = "super::serde_helpers::opt_bson_datetime")]
pub last_reconciled_at: Option<DateTime<Utc>>,
}
impl ExternalRef {
/// A freshly linked (not yet reconciled) reference to a sibling system.
pub fn linked(system: ExternalSystem) -> Self {
Self {
system,
project_id: None,
subject_sha256: None,
status: "linked".to_string(),
license_grant: None,
last_reconciled_at: None,
}
}
}
/// A regulatory / standards framework a target must comply with. Drives which
/// controls the mapping engine pulls from the [`crate::traits::ControlsProvider`].
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum ComplianceFramework {
/// EU Cyber Resilience Act.
Cra,
/// IEC 62443 (industrial automation & control systems security).
Iec62443,
/// EU General Data Protection Regulation.
Gdpr,
/// SOC 2.
Soc2,
/// ISO/IEC 27001.
Iso27001,
/// EU Radio Equipment Directive (RED) cybersecurity articles.
RedDirective,
}
impl std::fmt::Display for ComplianceFramework {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::Cra => write!(f, "cra"),
Self::Iec62443 => write!(f, "iec_62443"),
Self::Gdpr => write!(f, "gdpr"),
Self::Soc2 => write!(f, "soc2"),
Self::Iso27001 => write!(f, "iso_27001"),
Self::RedDirective => write!(f, "red_directive"),
}
}
}
/// The compliance scope of a target: which frameworks apply and, optionally, the
/// jurisdiction. Captured at onboarding (with per-target-type defaults from
/// [`default_compliance_profile`]).
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct ComplianceProfile {
/// Applicable frameworks.
#[serde(default)]
pub frameworks: Vec<ComplianceFramework>,
/// Free-form jurisdiction (e.g. `eu`, `us`, `de`).
pub jurisdiction: Option<String>,
}
/// The sensible default compliance scope for a target type. Firmware / PLC /
/// embedded default to CRA + IEC 62443; software defaults to GDPR + SOC 2.
pub fn default_compliance_profile(target_type: TargetType) -> ComplianceProfile {
use ComplianceFramework::{Cra, Gdpr, Iec62443, Soc2};
let frameworks = match target_type {
TargetType::PlcSps
| TargetType::FirmwareBareMetal
| TargetType::FirmwareRtos
| TargetType::EmbeddedLinuxYocto => vec![Cra, Iec62443],
TargetType::WebApp | TargetType::BackendService => vec![Gdpr, Soc2],
TargetType::DesktopApp | TargetType::AndroidApp | TargetType::IosApp => {
vec![Gdpr, Cra]
}
};
ComplianceProfile {
frameworks,
jurisdiction: None,
}
}
/// A target onboarded for scanning: the unified replacement for the legacy
/// `TrackedRepository` (SAST) and `DastTarget` (DAST) records.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct OnboardedTarget {
/// Mongo id. Preserved from the legacy record during migration so every
/// downstream collection keyed by `repo_id` / `target_id` keeps resolving.
#[serde(rename = "_id", skip_serializing_if = "Option::is_none")]
pub id: Option<bson::oid::ObjectId>,
/// Human-friendly name.
#[serde(default)]
pub name: String,
/// The software family this target belongs to.
pub target_type: TargetType,
/// Optional free-form description (also a classification input).
pub description: Option<String>,
/// The artifacts provided for this target.
#[serde(default)]
pub artifacts: Vec<Artifact>,
/// The classifier's verdict, once run.
pub classification: Option<Classification>,
/// How this target should be scanned.
#[serde(default)]
pub scan_config: TargetScanConfig,
/// The compliance scope (applicable frameworks / jurisdiction).
#[serde(default)]
pub compliance_profile: ComplianceProfile,
/// Links to sibling products (tramiton, ...) holding reconcilable evidence.
#[serde(default)]
pub external_refs: Vec<ExternalRef>,
/// Cron schedule for recurring scans, if any.
pub scan_schedule: Option<String>,
/// Whether inbound webhooks are enabled for this target.
#[serde(default)]
pub webhook_enabled: bool,
/// HMAC secret for verifying inbound webhooks.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub webhook_secret: Option<String>,
/// Cached count of findings across this target's scans.
#[serde(default)]
pub findings_count: u32,
/// Creation timestamp.
#[serde(
default = "chrono::Utc::now",
with = "super::serde_helpers::bson_datetime"
)]
pub created_at: DateTime<Utc>,
/// Last-update timestamp.
#[serde(
default = "chrono::Utc::now",
with = "super::serde_helpers::bson_datetime"
)]
pub updated_at: DateTime<Utc>,
}
impl OnboardedTarget {
/// A new target of the given type with a freshly generated webhook secret.
pub fn new(name: String, target_type: TargetType) -> Self {
let now = Utc::now();
let webhook_secret = uuid::Uuid::new_v4().to_string().replace('-', "");
Self {
id: None,
name,
target_type,
description: None,
artifacts: Vec::new(),
classification: None,
scan_config: TargetScanConfig::default(),
compliance_profile: default_compliance_profile(target_type),
external_refs: Vec::new(),
scan_schedule: None,
webhook_enabled: false,
webhook_secret: Some(webhook_secret),
findings_count: 0,
created_at: now,
updated_at: now,
}
}
/// The first artifact of the given kind, if present.
pub fn first_of(&self, kind: ArtifactKind) -> Option<&Artifact> {
self.artifacts.iter().find(|a| a.kind == kind)
}
/// Whether the target has at least one artifact of the given kind.
pub fn has(&self, kind: ArtifactKind) -> bool {
self.artifacts.iter().any(|a| a.kind == kind)
}
/// The primary code artifact (git repo or source archive), if any.
pub fn code_artifact(&self) -> Option<&Artifact> {
self.artifacts
.iter()
.find(|a| matches!(a.kind, ArtifactKind::GitRepo | ArtifactKind::SourceArchive))
}
/// The live-URL artifact, if any.
pub fn live_url(&self) -> Option<&Artifact> {
self.first_of(ArtifactKind::LiveUrl)
}
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
fn sample_target() -> OnboardedTarget {
let mut t = OnboardedTarget::new("acme-web".to_string(), TargetType::WebApp);
t.artifacts.push(Artifact::git_repo(
"https://git.example.com/acme.git",
"main",
));
t.artifacts
.push(Artifact::live_url("https://acme.example.com"));
t
}
#[test]
fn onboarded_target_bson_round_trip() {
let t = sample_target();
let b = bson::to_bson(&t).expect("serialize");
let back: OnboardedTarget = bson::from_bson(b.clone()).expect("deserialize");
let b2 = bson::to_bson(&back).expect("re-serialize");
assert_eq!(b, b2);
}
#[test]
fn enum_display_is_snake_case() {
assert_eq!(
TargetType::FirmwareBareMetal.to_string(),
"firmware_bare_metal"
);
assert_eq!(TargetType::PlcSps.to_string(), "plc_sps");
assert_eq!(ArtifactKind::PlcProject.to_string(), "plc_project");
assert_eq!(ArtifactKind::MobilePackage.to_string(), "mobile_package");
}
#[test]
fn helpers_locate_artifacts() {
let t = sample_target();
assert!(t.has(ArtifactKind::GitRepo));
assert!(t.live_url().is_some());
assert!(t.code_artifact().is_some());
assert!(!t.has(ArtifactKind::FirmwareImage));
assert_eq!(
t.first_of(ArtifactKind::GitRepo).map(|a| a.kind),
Some(ArtifactKind::GitRepo)
);
}
#[test]
fn new_target_generates_webhook_secret() {
let t = OnboardedTarget::new("t".to_string(), TargetType::BackendService);
let secret = t.webhook_secret.expect("secret present");
assert_eq!(secret.len(), 32);
assert!(!secret.contains('-'));
}
#[test]
fn dast_auth_folds_into_artifact_auth() {
let dast = DastAuthConfig {
method: "bearer".to_string(),
login_url: Some("https://x/login".to_string()),
username: Some("user".to_string()),
password: Some("pw".to_string()),
token: Some("tok".to_string()),
headers: None,
};
let auth = ArtifactAuth::from(dast);
assert_eq!(auth.method, "bearer");
// Bearer token wins over password.
assert_eq!(auth.secret.as_deref(), Some("tok"));
assert_eq!(auth.login_url.as_deref(), Some("https://x/login"));
}
#[test]
fn each_artifact_gets_a_unique_id() {
let a = Artifact::firmware_image("fw.bin");
let b = Artifact::firmware_image("fw.bin");
assert_ne!(a.id, b.id);
}
#[test]
fn firmware_default_profile_is_cra_and_62443() {
let p = default_compliance_profile(TargetType::FirmwareBareMetal);
assert!(p.frameworks.contains(&ComplianceFramework::Cra));
assert!(p.frameworks.contains(&ComplianceFramework::Iec62443));
}
#[test]
fn webapp_default_profile_is_gdpr_and_soc2() {
let p = default_compliance_profile(TargetType::WebApp);
assert!(p.frameworks.contains(&ComplianceFramework::Gdpr));
assert!(p.frameworks.contains(&ComplianceFramework::Soc2));
}
#[test]
fn new_target_gets_default_profile_and_no_external_refs() {
let t = OnboardedTarget::new("fw".to_string(), TargetType::FirmwareRtos);
assert!(!t.compliance_profile.frameworks.is_empty());
assert!(t.external_refs.is_empty());
}
#[test]
fn external_ref_linked_defaults() {
let r = ExternalRef::linked(ExternalSystem::Tramiton);
assert_eq!(r.system, ExternalSystem::Tramiton);
assert_eq!(r.status, "linked");
assert!(r.project_id.is_none());
assert!(r.last_reconciled_at.is_none());
}
}
+1 -19
View File
@@ -3,7 +3,7 @@ use serde::{Deserialize, Serialize};
use super::repository::ScanTrigger;
#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq)]
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq)]
#[serde(rename_all = "lowercase")]
pub enum ScanType {
Sast,
@@ -16,14 +16,6 @@ pub enum ScanType {
SecretDetection,
Lint,
CodeReview,
/// Static analysis of a firmware image (unpack + component CVE).
FirmwareStatic,
/// Control-logic security analysis of PLC / SPS programs.
PlcControlLogic,
/// Static analysis of a mobile package (APK / AAB / IPA).
MobileStatic,
/// Static analysis of a container image.
ContainerScan,
}
impl std::fmt::Display for ScanType {
@@ -39,10 +31,6 @@ impl std::fmt::Display for ScanType {
Self::SecretDetection => write!(f, "secret_detection"),
Self::Lint => write!(f, "lint"),
Self::CodeReview => write!(f, "code_review"),
Self::FirmwareStatic => write!(f, "firmware_static"),
Self::PlcControlLogic => write!(f, "plc_control_logic"),
Self::MobileStatic => write!(f, "mobile_static"),
Self::ContainerScan => write!(f, "container_scan"),
}
}
}
@@ -59,8 +47,6 @@ pub enum ScanRunStatus {
#[serde(rename_all = "snake_case")]
pub enum ScanPhase {
ChangeDetection,
ArtifactIngest,
Classification,
Sast,
SbomGeneration,
CveScanning,
@@ -69,10 +55,6 @@ pub enum ScanPhase {
LintScanning,
CodeReview,
GraphBuilding,
FirmwareStatic,
PlcAnalysis,
MobileStatic,
ContainerScan,
LlmTriage,
IssueCreation,
DastScanning,
-379
View File
@@ -1,379 +0,0 @@
//! The scan-applicability matrix.
//!
//! Which scans are possible for a target is a function of its [`TargetType`] and
//! which [`ArtifactKind`]s are actually present: SAST needs code, DAST needs a
//! running URL, firmware-static analysis needs a firmware image, and so on. This
//! module encodes that as a table — one rule set per target type — and resolves
//! it against a concrete [`OnboardedTarget`] into a list of [`ScanOption`]s the
//! onboarding wizard and the scan pipeline both consume.
use crate::models::{ArtifactKind, OnboardedTarget, ScanType, TargetType};
/// What an artifact a scan needs in order to run.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum ArtifactRequirement {
/// Source code — a git repo or a source archive.
Code,
/// A reachable running instance (live URL / endpoint).
RunningUrl,
/// A firmware image / binary blob.
Firmware,
/// A PLC project (PLCopen XML or Structured Text).
Plc,
/// A mobile package (APK / AAB / IPA).
Mobile,
/// A container image.
Container,
/// No specific artifact required.
Any,
}
/// A static rule: this scan applies to a target type, needs this artifact, and
/// defaults on/off. The rationale explains the entry to the user.
#[derive(Debug, Clone, Copy)]
pub struct ScanRule {
/// The scan this rule governs.
pub scan: ScanType,
/// Whether the scan is on by default (only when its artifact is present).
pub default_on: bool,
/// Human-readable explanation of what the scan does here.
pub rationale: &'static str,
/// The artifact the scan consumes.
pub requires: ArtifactRequirement,
}
impl ScanRule {
const fn new(
scan: ScanType,
default_on: bool,
rationale: &'static str,
requires: ArtifactRequirement,
) -> Self {
Self {
scan,
default_on,
rationale,
requires,
}
}
}
/// A resolved scan choice for a specific target: a rule intersected with the
/// artifacts actually present. `blocked_reason` is `Some` when the required
/// artifact is missing.
#[derive(Debug, Clone)]
pub struct ScanOption {
/// The scan.
pub scan: ScanType,
/// Whether to pre-select the scan (false when blocked).
pub default_on: bool,
/// Why the scan is offered.
pub rationale: String,
/// The artifact kind the scan needs, if any specific one.
pub required_artifact: Option<ArtifactKind>,
/// Set when the required artifact is absent, explaining the block.
pub blocked_reason: Option<String>,
}
/// The SAST umbrella: every static-analysis sub-scan that runs over source code.
fn sast_umbrella() -> Vec<ScanRule> {
use ArtifactRequirement::Code;
vec![
ScanRule::new(
ScanType::Sast,
true,
"Static analysis (Semgrep) over source",
Code,
),
ScanRule::new(
ScanType::Sbom,
true,
"Software bill of materials from source",
Code,
),
ScanRule::new(
ScanType::Cve,
true,
"Match dependencies against known CVEs",
Code,
),
ScanRule::new(
ScanType::SecretDetection,
true,
"Scan source for committed secrets",
Code,
),
ScanRule::new(ScanType::Lint, true, "Language linters over source", Code),
ScanRule::new(
ScanType::Gdpr,
true,
"GDPR data-handling pattern checks",
Code,
),
ScanRule::new(
ScanType::OAuth,
true,
"OAuth misconfiguration patterns",
Code,
),
ScanRule::new(
ScanType::Graph,
true,
"Build the code graph for impact analysis",
Code,
),
ScanRule::new(
ScanType::CodeReview,
false,
"LLM code review over changed source",
Code,
),
]
}
/// The rule set for a target type. Scans that are never applicable to a type are
/// simply absent (e.g. DAST is not listed for a PLC target).
pub fn rules_for(target_type: TargetType) -> Vec<ScanRule> {
use ArtifactRequirement::{Firmware, Mobile, Plc, RunningUrl};
match target_type {
TargetType::WebApp | TargetType::BackendService => {
let mut r = sast_umbrella();
r.push(ScanRule::new(
ScanType::Dast,
true,
"Dynamic scan of the running endpoint",
RunningUrl,
));
r
}
TargetType::DesktopApp => sast_umbrella(),
TargetType::AndroidApp | TargetType::IosApp => {
let mut r = sast_umbrella();
r.push(ScanRule::new(
ScanType::MobileStatic,
true,
"Static analysis of the mobile package (manifest, permissions, libs)",
Mobile,
));
r
}
TargetType::FirmwareBareMetal | TargetType::FirmwareRtos => {
let mut r = sast_umbrella();
r.push(ScanRule::new(
ScanType::FirmwareStatic,
true,
"Unpack and statically analyze the firmware image",
Firmware,
));
r.push(ScanRule::new(
ScanType::Sbom,
true,
"SBOM from the firmware image (binwalk / tramiton)",
Firmware,
));
r.push(ScanRule::new(
ScanType::Cve,
true,
"Match firmware components against known CVEs",
Firmware,
));
r
}
TargetType::EmbeddedLinuxYocto => {
let mut r = sast_umbrella();
r.push(ScanRule::new(
ScanType::FirmwareStatic,
true,
"EMBA / binwalk static analysis of the image",
Firmware,
));
r.push(ScanRule::new(
ScanType::Sbom,
true,
"SBOM from image layers / recipes",
Firmware,
));
r.push(ScanRule::new(
ScanType::Cve,
true,
"Match image components against known CVEs",
Firmware,
));
r.push(ScanRule::new(
ScanType::Dast,
false,
"Dynamic scan of exposed network services (if any)",
RunningUrl,
));
r
}
TargetType::PlcSps => vec![ScanRule::new(
ScanType::PlcControlLogic,
true,
"Control-logic security rules over the PLC program",
Plc,
)],
}
}
/// Whether an active penetration test is applicable to this target type.
///
/// Pentest runs as its own session (not a [`ScanType`] scan) and needs a
/// reachable running target, so it is offered only for the network-reachable
/// families.
pub fn supports_pentest(target_type: TargetType) -> bool {
matches!(
target_type,
TargetType::WebApp
| TargetType::BackendService
| TargetType::AndroidApp
| TargetType::IosApp
| TargetType::EmbeddedLinuxYocto
)
}
/// The representative artifact kind a requirement is satisfied by.
fn representative_kind(req: ArtifactRequirement) -> Option<ArtifactKind> {
match req {
ArtifactRequirement::Code => Some(ArtifactKind::GitRepo),
ArtifactRequirement::RunningUrl => Some(ArtifactKind::LiveUrl),
ArtifactRequirement::Firmware => Some(ArtifactKind::FirmwareImage),
ArtifactRequirement::Plc => Some(ArtifactKind::PlcProject),
ArtifactRequirement::Mobile => Some(ArtifactKind::MobilePackage),
ArtifactRequirement::Container => Some(ArtifactKind::ContainerImage),
ArtifactRequirement::Any => None,
}
}
/// Whether the target carries an artifact that satisfies the requirement.
fn requirement_satisfied(req: ArtifactRequirement, target: &OnboardedTarget) -> bool {
match req {
ArtifactRequirement::Code => target.code_artifact().is_some(),
ArtifactRequirement::RunningUrl => target.has(ArtifactKind::LiveUrl),
ArtifactRequirement::Firmware => target.has(ArtifactKind::FirmwareImage),
ArtifactRequirement::Plc => target.has(ArtifactKind::PlcProject),
ArtifactRequirement::Mobile => target.has(ArtifactKind::MobilePackage),
ArtifactRequirement::Container => target.has(ArtifactKind::ContainerImage),
ArtifactRequirement::Any => true,
}
}
/// Resolve the matrix for a concrete target into the scans it can run, marking
/// any whose required artifact is missing as blocked.
pub fn applicable_scans(target: &OnboardedTarget) -> Vec<ScanOption> {
rules_for(target.target_type)
.into_iter()
.map(|rule| {
let satisfied = requirement_satisfied(rule.requires, target);
let required_artifact = representative_kind(rule.requires);
let blocked_reason = if satisfied {
None
} else {
Some(match required_artifact {
Some(kind) => format!("no {kind} artifact provided"),
None => "required artifact missing".to_string(),
})
};
ScanOption {
scan: rule.scan,
default_on: rule.default_on && satisfied,
rationale: rule.rationale.to_string(),
required_artifact,
blocked_reason,
}
})
.collect()
}
#[cfg(test)]
#[allow(clippy::expect_used, clippy::unwrap_used)]
mod tests {
use super::*;
use crate::models::{Artifact, PlcFormat};
fn target_with(target_type: TargetType, artifacts: Vec<Artifact>) -> OnboardedTarget {
let mut t = OnboardedTarget::new("t".to_string(), target_type);
t.artifacts = artifacts;
t
}
fn option<'a>(opts: &'a [ScanOption], scan: ScanType) -> Option<&'a ScanOption> {
opts.iter().find(|o| o.scan == scan)
}
#[test]
fn webapp_with_code_and_url_offers_sast_and_dast() {
let t = target_with(
TargetType::WebApp,
vec![
Artifact::git_repo("u", "main"),
Artifact::live_url("http://x"),
],
);
let opts = applicable_scans(&t);
let sast = option(&opts, ScanType::Sast).expect("sast offered");
assert!(sast.default_on && sast.blocked_reason.is_none());
let dast = option(&opts, ScanType::Dast).expect("dast offered");
assert!(dast.default_on && dast.blocked_reason.is_none());
}
#[test]
fn webapp_without_url_blocks_dast() {
let t = target_with(TargetType::WebApp, vec![Artifact::git_repo("u", "main")]);
let opts = applicable_scans(&t);
let dast = option(&opts, ScanType::Dast).expect("dast listed");
assert!(!dast.default_on);
assert!(dast.blocked_reason.is_some());
assert_eq!(dast.required_artifact, Some(ArtifactKind::LiveUrl));
}
#[test]
fn firmware_offers_firmware_static_and_not_dast() {
let t = target_with(
TargetType::FirmwareBareMetal,
vec![Artifact::firmware_image("fw.bin")],
);
let opts = applicable_scans(&t);
let fw = option(&opts, ScanType::FirmwareStatic).expect("firmware static offered");
assert!(fw.default_on && fw.blocked_reason.is_none());
assert!(option(&opts, ScanType::Dast).is_none());
}
#[test]
fn plc_offers_only_control_logic() {
let t = target_with(
TargetType::PlcSps,
vec![Artifact::plc_project("p.xml", PlcFormat::PlcopenXml)],
);
let opts = applicable_scans(&t);
assert_eq!(opts.len(), 1);
assert_eq!(opts[0].scan, ScanType::PlcControlLogic);
assert!(opts[0].default_on);
}
#[test]
fn pentest_support_matches_reachable_families() {
assert!(supports_pentest(TargetType::WebApp));
assert!(supports_pentest(TargetType::BackendService));
assert!(!supports_pentest(TargetType::PlcSps));
assert!(!supports_pentest(TargetType::FirmwareBareMetal));
assert!(!supports_pentest(TargetType::DesktopApp));
}
#[test]
fn every_target_type_has_at_least_one_rule() {
for tt in [
TargetType::WebApp,
TargetType::BackendService,
TargetType::DesktopApp,
TargetType::AndroidApp,
TargetType::IosApp,
TargetType::FirmwareBareMetal,
TargetType::FirmwareRtos,
TargetType::EmbeddedLinuxYocto,
TargetType::PlcSps,
] {
assert!(!rules_for(tt).is_empty(), "{tt} has no rules");
}
}
}
-51
View File
@@ -1,51 +0,0 @@
//! The target-classification port.
//!
//! A [`TargetClassifier`] inspects a target's artifacts (and optionally their
//! ingested working directories) and proposes one or more [`ClassifierVerdict`]s
//! — a target type, a confidence, and the facts the decision rested on. Concrete
//! classifiers live in the agent (language/build-system fingerprinting, a
//! firmware detector backed by tramiton, etc.); a registry merges and ranks
//! their verdicts. This mirrors the [`crate::traits::Scanner`] port so the two
//! read the same way.
use std::collections::HashMap;
use std::path::PathBuf;
use crate::error::CoreError;
use crate::models::{Artifact, DetectedFact, TargetType};
/// Everything a classifier needs to reason about a target.
pub struct ClassificationInput<'a> {
/// The artifacts declared for the target.
pub artifacts: &'a [Artifact],
/// Ingested working paths, keyed by [`Artifact::id`]. Absent for artifacts
/// with no on-disk form (e.g. a live URL).
pub working_paths: &'a HashMap<String, PathBuf>,
/// Free-form description of the target, if provided.
pub description: Option<&'a str>,
}
/// A single classifier's proposal for a target.
pub struct ClassifierVerdict {
/// The proposed target type.
pub target_type: TargetType,
/// Confidence in `[0.0, 1.0]`.
pub confidence: f32,
/// Facts that informed the proposal.
pub facts: Vec<DetectedFact>,
/// Human-readable explanation.
pub rationale: String,
}
/// A source of target-type classification.
#[allow(async_fn_in_trait)]
pub trait TargetClassifier: Send + Sync {
/// Stable identifier for this classifier (recorded in `detected_by`).
fn name(&self) -> &str;
/// Propose zero or more ranked verdicts for the given input.
async fn classify(
&self,
input: &ClassificationInput<'_>,
) -> Result<Vec<ClassifierVerdict>, CoreError>;
}
-45
View File
@@ -1,45 +0,0 @@
//! The compliance-controls provider port.
//!
//! The mapping engine turns findings into compliance status against a corpus of
//! controls. That corpus is pluggable: the built-in OSCAL catalog by default, or
//! a tenant-owned RAG of atomic controls derived from laws
//! (`breakpilot-compliance`) when available. A [`ControlsProvider`] abstracts the
//! source so the mapping engine does not hardcode a catalog.
use crate::error::CoreError;
use crate::models::ComplianceFramework;
/// A control retrieved from a controls corpus.
#[derive(Debug, Clone)]
pub struct Control {
/// Stable control identifier (e.g. an OSCAL control id or a RAG chunk id).
pub id: String,
/// The framework this control belongs to.
pub framework: ComplianceFramework,
/// Short human-readable title.
pub title: String,
/// The control text / requirement.
pub text: String,
/// Free-form source reference (catalog name, law citation, ...).
pub source: Option<String>,
}
/// A query for relevant controls.
pub struct ControlQuery<'a> {
/// Frameworks in scope for the target.
pub frameworks: &'a [ComplianceFramework],
/// Free-text describing what to map (a finding summary, a component, ...).
pub context: &'a str,
/// Maximum number of controls to return.
pub limit: usize,
}
/// A source of compliance controls (built-in OSCAL catalog, breakpilot RAG, ...).
#[allow(async_fn_in_trait)]
pub trait ControlsProvider: Send + Sync {
/// Stable identifier for this provider.
fn name(&self) -> &str;
/// Retrieve the controls most relevant to the query.
async fn controls(&self, query: &ControlQuery<'_>) -> Result<Vec<Control>, CoreError>;
}
-58
View File
@@ -1,58 +0,0 @@
//! The external-evidence provider port.
//!
//! A sibling product (tramiton, for firmware) may already hold authoritative
//! analysis for an artifact. An [`EvidenceProvider`] lets compliance-scanner
//! *reconcile* that evidence — a build plan, an SBOM, a VEX document, a
//! reproducible-build lock, an attestation — instead of recomputing it. The key
//! used to match is the artifact content digest (a firmware sha256, which equals
//! [`crate::models::Artifact::content_hash`]).
//!
//! Concrete providers live in the agent (a tramiton CLI shell-out today, a cloud
//! client later) plus a deterministic mock for tests, so nothing here depends on
//! an external binary.
use std::path::Path;
use crate::error::CoreError;
use crate::models::{Artifact, ExternalSystem};
/// A single reconcilable evidence document fetched from a sibling product.
#[derive(Debug, Clone)]
pub struct EvidenceDocument {
/// What the document is: `build_plan` | `sbom` | `vex` | `lock` | `attestation`.
pub kind: String,
/// The document's format (e.g. `cyclonedx-1.5`, `openvex-0.2.0`, `toml`, `json`).
pub format: String,
/// The raw document payload.
pub content: String,
}
/// The evidence a provider could return for a target's artifact.
#[derive(Debug, Clone, Default)]
pub struct ReconciledEvidence {
/// The sibling's project identifier, if resolved.
pub project_id: Option<String>,
/// The subject content digest the evidence pertains to.
pub subject_sha256: Option<String>,
/// The documents fetched (any of build plan / SBOM / VEX / lock / attestation).
pub documents: Vec<EvidenceDocument>,
}
/// A source of externally-held, reconcilable evidence for an artifact.
#[allow(async_fn_in_trait)]
pub trait EvidenceProvider: Send + Sync {
/// Which sibling product this provider integrates.
fn system(&self) -> ExternalSystem;
/// Whether this provider can handle the given artifact + working path
/// (e.g. tramiton handles firmware images / embedded source trees).
fn handles(&self, artifact: &Artifact, working_path: Option<&Path>) -> bool;
/// Reconcile existing evidence for the artifact, keyed by its content digest.
/// Returns `Ok(None)` when the provider has nothing for this artifact.
async fn reconcile(
&self,
artifact: &Artifact,
working_path: Option<&Path>,
) -> Result<Option<ReconciledEvidence>, CoreError>;
}
-6
View File
@@ -1,16 +1,10 @@
pub mod classifier;
pub mod controls;
pub mod dast_agent;
pub mod evidence;
pub mod graph_builder;
pub mod issue_tracker;
pub mod pentest_tool;
pub mod scanner;
pub use classifier::{ClassificationInput, ClassifierVerdict, TargetClassifier};
pub use controls::{Control, ControlQuery, ControlsProvider};
pub use dast_agent::{DastAgent, DastContext, DiscoveredEndpoint, EndpointParameter};
pub use evidence::{EvidenceDocument, EvidenceProvider, ReconciledEvidence};
pub use graph_builder::{LanguageParser, ParseOutput};
pub use issue_tracker::IssueTracker;
pub use pentest_tool::{PentestTool, PentestToolContext, PentestToolResult};
-4
View File
@@ -12,10 +12,6 @@ pub enum Route {
OverviewPage {},
#[route("/repositories")]
RepositoriesPage {},
#[route("/targets")]
TargetsPage {},
#[route("/onboard")]
OnboardingPage {},
#[route("/findings")]
FindingsPage {},
#[route("/findings/:id")]
@@ -24,15 +24,10 @@ pub fn Sidebar() -> Element {
icon: rsx! { Icon { icon: BsSpeedometer2, width: 18, height: 18 } },
},
NavItem {
label: "Targets",
route: Route::TargetsPage {},
label: "Repositories",
route: Route::RepositoriesPage {},
icon: rsx! { Icon { icon: BsFolder2Open, width: 18, height: 18 } },
},
NavItem {
label: "Onboard",
route: Route::OnboardingPage {},
icon: rsx! { Icon { icon: BsPlusCircle, width: 18, height: 18 } },
},
NavItem {
label: "Findings",
route: Route::FindingsPage {},
@@ -10,7 +10,6 @@ pub mod issues;
pub mod mcp;
pub mod mcp_tokens;
pub mod notifications;
pub mod onboarding;
pub mod pentest;
#[allow(clippy::too_many_arguments)]
pub mod repositories;
@@ -1,139 +0,0 @@
//! Server functions for the onboarding wizard — proxy to the agent's
//! `/api/v1/targets` endpoints.
use dioxus::prelude::*;
use serde::{Deserialize, Serialize};
/// One artifact the wizard collects for a target.
#[derive(Debug, Clone, Serialize, Deserialize, Default, PartialEq)]
pub struct ArtifactInputDto {
pub kind: String,
pub source_ref: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub branch: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub plc_format: Option<String>,
}
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct TargetsResponse {
pub data: Vec<serde_json::Value>,
pub total: Option<u64>,
}
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct TargetResponse {
pub data: serde_json::Value,
}
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct ApplicableScansData {
#[serde(default)]
pub scans: Vec<serde_json::Value>,
#[serde(default)]
pub pentest_supported: bool,
}
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
pub struct ApplicableScansResponse {
pub data: ApplicableScansData,
}
/// List onboarded targets.
#[server]
pub async fn fetch_targets() -> Result<TargetsResponse, ServerFnError> {
let resp = super::agent_client::agent_get("/api/v1/targets")
.await?
.send()
.await
.map_err(|e| ServerFnError::new(e.to_string()))?;
resp.json()
.await
.map_err(|e| ServerFnError::new(e.to_string()))
}
/// Create a target with the collected artifacts.
#[server]
pub async fn create_target(
name: String,
target_type: String,
description: Option<String>,
artifacts: Vec<ArtifactInputDto>,
) -> Result<TargetResponse, ServerFnError> {
let body = serde_json::json!({
"name": name,
"target_type": target_type,
"description": description,
"artifacts": artifacts,
});
let resp = super::agent_client::agent_request(reqwest::Method::POST, "/api/v1/targets")
.await?
.json(&body)
.send()
.await
.map_err(|e| ServerFnError::new(e.to_string()))?;
resp.json()
.await
.map_err(|e| ServerFnError::new(e.to_string()))
}
/// Run kind-based classification on a target.
#[server]
pub async fn detect_target(id: String) -> Result<TargetResponse, ServerFnError> {
let resp = super::agent_client::agent_request(
reqwest::Method::POST,
&format!("/api/v1/targets/{id}/detect"),
)
.await?
.send()
.await
.map_err(|e| ServerFnError::new(e.to_string()))?;
resp.json()
.await
.map_err(|e| ServerFnError::new(e.to_string()))
}
/// Fetch the scan-applicability matrix for a target.
#[server]
pub async fn fetch_applicable_scans(id: String) -> Result<ApplicableScansResponse, ServerFnError> {
let resp = super::agent_client::agent_get(&format!("/api/v1/targets/{id}/applicable-scans"))
.await?
.send()
.await
.map_err(|e| ServerFnError::new(e.to_string()))?;
resp.json()
.await
.map_err(|e| ServerFnError::new(e.to_string()))
}
/// Delete a target (and cascade its findings / SBOM / scan runs / CVE alerts).
#[server]
pub async fn delete_target(id: String) -> Result<serde_json::Value, ServerFnError> {
let resp = super::agent_client::agent_request(
reqwest::Method::DELETE,
&format!("/api/v1/targets/{id}"),
)
.await?
.send()
.await
.map_err(|e| ServerFnError::new(e.to_string()))?;
resp.json()
.await
.map_err(|e| ServerFnError::new(e.to_string()))
}
/// Trigger a scan for a target.
#[server]
pub async fn trigger_target_scan(id: String) -> Result<serde_json::Value, ServerFnError> {
let resp = super::agent_client::agent_request(
reqwest::Method::POST,
&format!("/api/v1/targets/{id}/scan"),
)
.await?
.send()
.await
.map_err(|e| ServerFnError::new(e.to_string()))?;
resp.json()
.await
.map_err(|e| ServerFnError::new(e.to_string()))
}
+5 -5
View File
@@ -20,7 +20,7 @@ pub fn FindingsPage() -> Element {
let mut selected_ids = use_signal(Vec::<String>::new);
let repos = use_resource(|| async {
crate::infrastructure::onboarding::fetch_targets()
crate::infrastructure::repositories::fetch_repositories(1)
.await
.ok()
});
@@ -86,14 +86,14 @@ pub fn FindingsPage() -> Element {
}
select {
onchange: move |e| { repo_filter.set(e.value()); page.set(1); },
option { value: "", "All Targets" }
option { value: "", "All Repositories" }
{
match &*repos.read() {
Some(Some(resp)) => rsx! {
for t in &resp.data {
for repo in &resp.data {
{
let id = t.get("_id").and_then(|o| o.get("$oid")).and_then(|s| s.as_str()).unwrap_or_default().to_string();
let name = t.get("name").and_then(|n| n.as_str()).unwrap_or_default().to_string();
let id = repo.id.as_ref().map(|id| id.to_hex()).unwrap_or_default();
let name = repo.name.clone();
rsx! {
option { value: "{id}", "{name}" }
}
-4
View File
@@ -12,13 +12,11 @@ pub mod impact_analysis;
pub mod issues;
pub mod mcp_servers;
pub mod mcp_tokens;
pub mod onboarding;
pub mod overview;
pub mod pentest_dashboard;
pub mod pentest_session;
pub mod repositories;
pub mod sbom;
pub mod targets;
pub use chat::ChatPage;
pub use chat_index::ChatIndexPage;
@@ -34,10 +32,8 @@ pub use impact_analysis::ImpactAnalysisPage;
pub use issues::IssuesPage;
pub use mcp_servers::McpServersPage;
pub use mcp_tokens::McpTokensPage;
pub use onboarding::OnboardingPage;
pub use overview::OverviewPage;
pub use pentest_dashboard::PentestDashboardPage;
pub use pentest_session::PentestSessionPage;
pub use repositories::RepositoriesPage;
pub use sbom::SbomPage;
pub use targets::TargetsPage;
@@ -1,402 +0,0 @@
use dioxus::prelude::*;
use crate::components::page_header::PageHeader;
use crate::infrastructure::onboarding::{
create_target, detect_target, fetch_applicable_scans, trigger_target_scan, ArtifactInputDto,
};
/// (value, label, one-line description) for the 9 target families.
const TARGET_TYPES: &[(&str, &str, &str)] = &[
("web_app", "Web Application", "Front end + server"),
("backend_service", "Backend / API", "REST, GraphQL, gRPC"),
("desktop_app", "Desktop App", "Windows / macOS / Linux"),
("android_app", "Android App", "APK / AAB"),
("ios_app", "iOS App", "IPA"),
(
"firmware_bare_metal",
"Firmware — bare metal",
"No operating system",
),
("firmware_rtos", "Firmware — RTOS", "Zephyr, FreeRTOS, ..."),
(
"embedded_linux_yocto",
"Embedded Linux / Yocto",
"BSP + image",
),
("plc_sps", "PLC / SPS", "IEC 61131-3"),
];
/// (value, label) for the artifact kinds a user can attach.
const ARTIFACT_KINDS: &[(&str, &str)] = &[
("git_repo", "Git repository"),
("source_archive", "Source archive (zip)"),
("firmware_image", "Firmware image"),
("mobile_package", "Mobile package (APK/IPA)"),
("container_image", "Container image"),
("live_url", "Live URL"),
("plc_project", "PLC project"),
("plaintext_description", "Description (text)"),
];
const STEP_LABELS: &[&str] = &["Target type", "Artifacts", "Review", "Done"];
/// One row in the applicable-scans list on the success step.
#[component]
fn ScanRow(scan: serde_json::Value) -> Element {
let name = scan
.get("scan")
.and_then(|v| v.as_str())
.unwrap_or("?")
.to_string();
let rationale = scan
.get("rationale")
.and_then(|v| v.as_str())
.unwrap_or("")
.to_string();
let blocked = scan
.get("blocked_reason")
.and_then(|v| v.as_str())
.map(String::from);
let default_on = scan
.get("default_on")
.and_then(|v| v.as_bool())
.unwrap_or(false);
let badge_class = if blocked.is_some() {
"badge badge-info"
} else if default_on {
"badge badge-success"
} else {
"badge"
};
rsx! {
div { style: "display: flex; gap: 8px; align-items: center; padding: 6px 0;",
span { class: "{badge_class}", "{name}" }
span { style: "opacity: 0.8;", "{rationale}" }
if let Some(b) = blocked {
span { style: "opacity: 0.6; font-style: italic;", "— {b}" }
}
}
}
}
fn kind_label(kind: &str) -> &str {
ARTIFACT_KINDS
.iter()
.find(|(v, _)| *v == kind)
.map(|(_, l)| *l)
.unwrap_or(kind)
}
fn type_label(value: &str) -> &str {
TARGET_TYPES
.iter()
.find(|(v, _, _)| *v == value)
.map(|(_, l, _)| *l)
.unwrap_or(value)
}
#[component]
pub fn OnboardingPage() -> Element {
let mut step = use_signal(|| 0usize);
let mut name = use_signal(String::new);
let mut target_type = use_signal(String::new);
let mut description = use_signal(String::new);
let mut artifacts = use_signal(Vec::<ArtifactInputDto>::new);
// "Add artifact" mini-form.
let mut new_kind = use_signal(|| "git_repo".to_string());
let mut new_source = use_signal(String::new);
let mut new_branch = use_signal(|| "main".to_string());
// Create + result state.
let mut creating = use_signal(|| false);
let mut error = use_signal(|| Option::<String>::None);
let mut scans = use_signal(Vec::<serde_json::Value>::new);
let mut suggested = use_signal(|| Option::<String>::None);
let mut created_id = use_signal(|| Option::<String>::None);
let mut scan_msg = use_signal(|| Option::<String>::None);
let step_now = step();
let can_advance_type = !name().trim().is_empty() && !target_type().trim().is_empty();
let has_artifacts = !artifacts().is_empty();
rsx! {
PageHeader {
title: "Onboard a target",
description: "Add a target, attach its artifacts, and see which scans apply.",
}
// Stepper.
div { class: "wizard-steps",
for (i, label) in STEP_LABELS.iter().enumerate() {
div {
class: if i == step_now { "wizard-step wizard-step-active" } else { "wizard-step" },
span { class: "wizard-step-dot", "{i + 1}" }
span { class: "wizard-step-label", "{label}" }
}
}
}
if let Some(err) = error() {
div { class: "card", style: "border-color: var(--danger, #d33); margin-bottom: 12px;",
div { class: "card-header", "Error" }
div { style: "padding: 12px;", "{err}" }
}
}
div { class: "card",
// ---- Step 0: target type + name ----
if step_now == 0 {
div { class: "card-header", "What kind of software is this?" }
div { style: "padding: 16px;",
div { class: "form-group",
label { "Name" }
input {
r#type: "text",
placeholder: "acme-web",
value: "{name}",
oninput: move |e| name.set(e.value()),
}
}
div {
style: "display: grid; grid-template-columns: repeat(auto-fill, minmax(200px, 1fr)); gap: 12px; margin-top: 12px;",
for (value, tlabel, tdesc) in TARGET_TYPES.iter().copied() {
div {
class: "card",
style: if target_type() == value {
"padding: 12px; cursor: pointer; border: 2px solid var(--accent, #3b82f6);"
} else {
"padding: 12px; cursor: pointer;"
},
onclick: move |_| target_type.set(value.to_string()),
div { style: "font-weight: 600;", "{tlabel}" }
div { style: "font-size: 0.85em; opacity: 0.7;", "{tdesc}" }
}
}
}
}
}
// ---- Step 1: artifacts ----
if step_now == 1 {
div { class: "card-header", "Attach artifacts" }
div { style: "padding: 16px;",
div { style: "display: flex; gap: 8px; flex-wrap: wrap; align-items: flex-end;",
div { class: "form-group", style: "margin: 0;",
label { "Kind" }
select {
value: "{new_kind}",
oninput: move |e| new_kind.set(e.value()),
for (value, klabel) in ARTIFACT_KINDS.iter().copied() {
option { value: "{value}", "{klabel}" }
}
}
}
div { class: "form-group", style: "margin: 0; flex: 1; min-width: 240px;",
label { "Reference (URL / path / text)" }
input {
r#type: "text",
placeholder: "https://git.example.com/acme.git",
value: "{new_source}",
oninput: move |e| new_source.set(e.value()),
}
}
if new_kind() == "git_repo" {
div { class: "form-group", style: "margin: 0;",
label { "Branch" }
input {
r#type: "text",
value: "{new_branch}",
oninput: move |e| new_branch.set(e.value()),
}
}
}
button {
class: "btn btn-secondary",
onclick: move |_| {
let kind = new_kind();
if !new_source().trim().is_empty() {
let branch = if kind == "git_repo" { Some(new_branch()) } else { None };
artifacts.write().push(ArtifactInputDto {
kind,
source_ref: new_source(),
branch,
plc_format: None,
});
new_source.set(String::new());
}
},
"+ Add"
}
}
div { style: "margin-top: 16px;",
if has_artifacts {
for (i, a) in artifacts().iter().enumerate() {
div {
style: "display: flex; justify-content: space-between; align-items: center; padding: 8px 12px; border: 1px solid var(--border, #333); border-radius: 6px; margin-bottom: 6px;",
span {
span { style: "opacity: 0.7;", "{kind_label(&a.kind)}: " }
"{a.source_ref}"
}
button {
class: "btn btn-ghost-danger btn-sm",
onclick: move |_| { artifacts.write().remove(i); },
"Remove"
}
}
}
} else {
div { style: "opacity: 0.6;", "No artifacts yet. Add at least one." }
}
}
}
}
// ---- Step 2: review + create ----
if step_now == 2 {
div { class: "card-header", "Review" }
div { style: "padding: 16px;",
div { class: "wizard-summary",
div { strong { "Name: " } "{name()}" }
div { strong { "Type: " } "{type_label(&target_type())}" }
div { strong { "Artifacts:" } }
ul {
for a in artifacts() {
li { "{kind_label(&a.kind)}: {a.source_ref}" }
}
}
}
div { style: "margin-top: 12px; opacity: 0.7; font-size: 0.9em;",
"The applicable scans (SAST / DAST / firmware / PLC) are shown after the target is created."
}
}
}
// ---- Step 3: created ----
if step_now == 3 {
div { class: "card-header", "Target onboarded" }
div { style: "padding: 16px;",
p {
strong { "{name()}" }
" was created."
if let Some(s) = suggested() {
span { " Suggested type from detection: " strong { "{type_label(&s)}" } "." }
}
}
h4 { style: "margin-top: 16px;", "Applicable scans" }
if scans().is_empty() {
div { style: "opacity: 0.6;", "No scans available (no code / URL / firmware artifact present)." }
} else {
for s in scans() {
ScanRow { scan: s }
}
}
if let Some(msg) = scan_msg() {
div { style: "margin-top: 8px; color: var(--success, #2a2);", "{msg}" }
}
div { style: "margin-top: 16px; display: flex; gap: 8px;",
button {
class: "btn btn-primary",
onclick: move |_| {
if let Some(id) = created_id() {
scan_msg.set(Some("Scan triggered...".to_string()));
spawn(async move {
match trigger_target_scan(id).await {
Ok(_) => scan_msg.set(Some(
"Scan started — findings will appear as it runs.".to_string(),
)),
Err(e) => scan_msg.set(Some(format!("Failed to start scan: {e}"))),
}
});
}
},
"Run scan"
}
button {
class: "btn btn-secondary",
onclick: move |_| {
step.set(0);
name.set(String::new());
target_type.set(String::new());
description.set(String::new());
artifacts.write().clear();
scans.write().clear();
suggested.set(None);
created_id.set(None);
scan_msg.set(None);
error.set(None);
},
"Onboard another"
}
}
}
}
}
// ---- Footer navigation ----
if step_now < 3 {
div { style: "display: flex; justify-content: space-between; margin-top: 16px;",
button {
class: "btn btn-back",
disabled: step_now == 0,
onclick: move |_| { if step() > 0 { step.set(step() - 1); } },
"Back"
}
if step_now < 2 {
button {
class: "btn btn-primary",
disabled: (step_now == 0 && !can_advance_type) || (step_now == 1 && !has_artifacts),
onclick: move |_| step.set(step() + 1),
"Next"
}
} else {
button {
class: "btn btn-primary",
disabled: creating(),
onclick: move |_| {
let n = name();
let tt = target_type();
let desc = description();
let arts = artifacts();
let d = if desc.trim().is_empty() { None } else { Some(desc) };
creating.set(true);
error.set(None);
spawn(async move {
match create_target(n, tt, d, arts).await {
Ok(resp) => {
let id = resp
.data
.get("_id")
.and_then(|o| o.get("$oid"))
.and_then(|s| s.as_str())
.map(String::from);
if let Some(id) = id {
created_id.set(Some(id.clone()));
if let Ok(sc) = fetch_applicable_scans(id.clone()).await {
scans.set(sc.data.scans);
}
if let Ok(det) = detect_target(id).await {
suggested.set(
det.data
.get("classification")
.and_then(|c| c.get("suggested"))
.and_then(|s| s.as_str())
.map(String::from),
);
}
}
step.set(3);
}
Err(e) => error.set(Some(e.to_string())),
}
creating.set(false);
});
},
if creating() { "Creating..." } else { "Create target" }
}
}
}
}
}
}
+230 -1
View File
@@ -23,6 +23,20 @@ async fn async_sleep_5s() {
#[component]
pub fn RepositoriesPage() -> Element {
let mut page = use_signal(|| 1u64);
let mut show_add_form = use_signal(|| false);
let mut name = use_signal(String::new);
let mut git_url = use_signal(String::new);
let mut branch = use_signal(|| "main".to_string());
let mut auth_token = use_signal(String::new);
let mut auth_username = use_signal(String::new);
let mut show_auth = use_signal(|| false);
let mut ssh_public_key = use_signal(String::new);
let mut show_tracker = use_signal(|| false);
let mut tracker_type_val = use_signal(String::new);
let mut tracker_owner_val = use_signal(String::new);
let mut tracker_repo_val = use_signal(String::new);
let mut tracker_token_val = use_signal(String::new);
let mut adding = use_signal(|| false);
let mut toasts = use_context::<Toasts>();
let mut confirm_delete = use_signal(|| Option::<(String, String)>::None); // (id, name)
let mut edit_repo_id = use_signal(|| Option::<String>::None);
@@ -50,7 +64,222 @@ pub fn RepositoriesPage() -> Element {
rsx! {
PageHeader {
title: "Repositories",
description: "Legacy git repositories. Onboard new targets from Targets / Onboard.",
description: "Tracked git repositories",
}
div { style: "margin-bottom: 16px;",
button {
class: "btn btn-primary",
onclick: move |_| show_add_form.toggle(),
if show_add_form() { "Cancel" } else { "+ Add Repository" }
}
}
if show_add_form() {
div { class: "card",
div { class: "card-header", "Add Repository" }
div { class: "form-group",
label { "Name" }
input {
r#type: "text",
placeholder: "my-project",
value: "{name}",
oninput: move |e| name.set(e.value()),
}
}
div { class: "form-group",
label { "Git URL" }
input {
r#type: "text",
placeholder: "https://github.com/org/repo.git or git@github.com:org/repo.git",
value: "{git_url}",
oninput: move |e| git_url.set(e.value()),
}
}
div { class: "form-group",
label { "Default Branch" }
input {
r#type: "text",
placeholder: "main",
value: "{branch}",
oninput: move |e| branch.set(e.value()),
}
}
// Private repo auth section
div { style: "margin-top: 8px;",
button {
class: "btn btn-ghost",
style: "font-size: 12px; padding: 4px 8px;",
onclick: move |_| {
let opening = !show_auth();
show_auth.toggle();
if opening {
// Fetch SSH key every time the section opens
ssh_public_key.set(String::new());
spawn(async move {
match crate::infrastructure::repositories::fetch_ssh_public_key().await {
Ok(key) => ssh_public_key.set(key),
Err(_) => ssh_public_key.set("(not available)".to_string()),
}
});
}
},
if show_auth() { "Hide auth options" } else { "Private repository?" }
}
}
if show_auth() {
div { class: "auth-section", style: "margin-top: 12px; padding: 12px; border: 1px solid var(--border-subtle); border-radius: 8px;",
// SSH deploy key display
div { style: "margin-bottom: 12px;",
label { style: "font-size: 12px; color: var(--text-secondary);",
"For SSH URLs: add this deploy key (read-only) to your repository"
}
div {
class: "copyable",
style: "margin-top: 4px; padding: 8px; background: var(--bg-secondary); border-radius: 4px;",
code {
style: "font-size: 11px; word-break: break-all; user-select: all;",
if ssh_public_key().is_empty() {
"Loading..."
} else {
"{ssh_public_key}"
}
}
if !ssh_public_key().is_empty() {
crate::components::copy_button::CopyButton { value: ssh_public_key(), small: true }
}
}
}
// HTTPS auth fields
p { style: "font-size: 12px; color: var(--text-secondary); margin-bottom: 8px;",
"For HTTPS URLs: provide an access token (PAT) or username/password"
}
div { class: "form-group",
label { "Auth Token / Password" }
input {
r#type: "password",
placeholder: "ghp_xxxx or personal access token",
value: "{auth_token}",
oninput: move |e| auth_token.set(e.value()),
}
}
div { class: "form-group",
label { "Username (optional, defaults to x-access-token)" }
input {
r#type: "text",
placeholder: "x-access-token",
value: "{auth_username}",
oninput: move |e| auth_username.set(e.value()),
}
}
}
}
// Issue tracker config section
div { style: "margin-top: 8px;",
button {
class: "btn btn-ghost",
style: "font-size: 12px; padding: 4px 8px;",
onclick: move |_| show_tracker.toggle(),
if show_tracker() { "Hide tracker options" } else { "Issue tracker?" }
}
}
if show_tracker() {
div { class: "auth-section", style: "margin-top: 12px; padding: 12px; border: 1px solid var(--border-subtle); border-radius: 8px;",
p { style: "font-size: 12px; color: var(--text-secondary); margin-bottom: 8px;",
"Configure an issue tracker to auto-create issues from findings"
}
div { class: "form-group",
label { "Tracker Type" }
select {
value: "{tracker_type_val}",
onchange: move |e| tracker_type_val.set(e.value()),
option { value: "", "None" }
option { value: "github", "GitHub" }
option { value: "gitlab", "GitLab" }
option { value: "gitea", "Gitea" }
option { value: "jira", "Jira" }
}
}
div { class: "form-group",
label { "Owner / Namespace" }
input {
r#type: "text",
placeholder: "org-name",
value: "{tracker_owner_val}",
oninput: move |e| tracker_owner_val.set(e.value()),
}
}
div { class: "form-group",
label { "Repository / Project" }
input {
r#type: "text",
placeholder: "repo-name",
value: "{tracker_repo_val}",
oninput: move |e| tracker_repo_val.set(e.value()),
}
}
div { class: "form-group",
label { "Tracker Token (PAT)" }
input {
r#type: "password",
placeholder: "ghp_xxxx / glpat-xxxx",
value: "{tracker_token_val}",
oninput: move |e| tracker_token_val.set(e.value()),
}
}
}
}
button {
class: "btn btn-primary",
disabled: adding(),
onclick: move |_| {
let n = name();
let u = git_url();
let b = branch();
let tok = {
let v = auth_token();
if v.is_empty() { None } else { Some(v) }
};
let usr = {
let v = auth_username();
if v.is_empty() { None } else { Some(v) }
};
let tt = { let v = tracker_type_val(); if v.is_empty() { None } else { Some(v) } };
let t_owner = { let v = tracker_owner_val(); if v.is_empty() { None } else { Some(v) } };
let t_repo = { let v = tracker_repo_val(); if v.is_empty() { None } else { Some(v) } };
let t_tok = { let v = tracker_token_val(); if v.is_empty() { None } else { Some(v) } };
adding.set(true);
spawn(async move {
match crate::infrastructure::repositories::add_repository(n, u, b, tok, usr, tt, t_owner, t_repo, t_tok).await {
Ok(_) => {
toasts.push(ToastType::Success, "Repository added");
repos.restart();
}
Err(e) => toasts.push(ToastType::Error, e.to_string()),
}
adding.set(false);
});
show_add_form.set(false);
show_auth.set(false);
show_tracker.set(false);
name.set(String::new());
git_url.set(String::new());
auth_token.set(String::new());
auth_username.set(String::new());
tracker_type_val.set(String::new());
tracker_owner_val.set(String::new());
tracker_repo_val.set(String::new());
tracker_token_val.set(String::new());
},
if adding() { "Validating..." } else { "Add" }
}
}
}
// ── Delete confirmation dialog ──
+9 -9
View File
@@ -28,9 +28,9 @@ pub fn SbomPage() -> Element {
let mut diff_repo_a = use_signal(String::new);
let mut diff_repo_b = use_signal(String::new);
// ── Targets for dropdowns ──
// ── Repos for dropdowns ──
let repos = use_resource(|| async {
crate::infrastructure::onboarding::fetch_targets()
crate::infrastructure::repositories::fetch_repositories(1)
.await
.ok()
});
@@ -114,14 +114,14 @@ pub fn SbomPage() -> Element {
select {
class: "sbom-filter-select",
onchange: move |e| { repo_filter.set(e.value()); page.set(1); },
option { value: "", "All Targets" }
option { value: "", "All Repositories" }
{
match &*repos.read() {
Some(Some(resp)) => rsx! {
for repo in &resp.data {
{
let id = repo.get("_id").and_then(|o| o.get("$oid")).and_then(|s| s.as_str()).unwrap_or_default().to_string();
let name = repo.get("name").and_then(|n| n.as_str()).unwrap_or_default().to_string();
let id = repo.id.as_ref().map(|id| id.to_hex()).unwrap_or_default();
let name = repo.name.clone();
rsx! { option { value: "{id}", "{name}" } }
}
}
@@ -476,8 +476,8 @@ pub fn SbomPage() -> Element {
Some(Some(resp)) => rsx! {
for repo in &resp.data {
{
let id = repo.get("_id").and_then(|o| o.get("$oid")).and_then(|s| s.as_str()).unwrap_or_default().to_string();
let name = repo.get("name").and_then(|n| n.as_str()).unwrap_or_default().to_string();
let id = repo.id.as_ref().map(|id| id.to_hex()).unwrap_or_default();
let name = repo.name.clone();
rsx! { option { value: "{id}", "{name}" } }
}
}
@@ -498,8 +498,8 @@ pub fn SbomPage() -> Element {
Some(Some(resp)) => rsx! {
for repo in &resp.data {
{
let id = repo.get("_id").and_then(|o| o.get("$oid")).and_then(|s| s.as_str()).unwrap_or_default().to_string();
let name = repo.get("name").and_then(|n| n.as_str()).unwrap_or_default().to_string();
let id = repo.id.as_ref().map(|id| id.to_hex()).unwrap_or_default();
let name = repo.name.clone();
rsx! { option { value: "{id}", "{name}" } }
}
}
-339
View File
@@ -1,339 +0,0 @@
//! Targets page — lists onboarded targets (the unified `OnboardedTarget`
//! records the wizard creates) and lets the user run a scan, inspect the
//! classification / applicable scans, or delete a target.
//!
//! Creation lives in the Onboard wizard (`/onboard`); this page is the
//! "where is my target, and what did the scan find" surface.
use dioxus::prelude::*;
use dioxus_free_icons::icons::bs_icons::*;
use dioxus_free_icons::Icon;
use crate::components::page_header::PageHeader;
use crate::components::toast::{ToastType, Toasts};
use crate::infrastructure::onboarding::{
delete_target, fetch_applicable_scans, fetch_targets, trigger_target_scan,
};
/// Prettify a snake_case target-type value into a human label.
fn pretty_type(v: &str) -> String {
match v {
"web_app" => "Web Application".into(),
"backend_service" => "Backend / API".into(),
"desktop_app" => "Desktop App".into(),
"android_app" => "Android App".into(),
"ios_app" => "iOS App".into(),
"firmware_bare_metal" => "Firmware — bare metal".into(),
"firmware_rtos" => "Firmware — RTOS".into(),
"embedded_linux_yocto" => "Embedded Linux / Yocto".into(),
"plc_sps" => "PLC / SPS".into(),
other => other.replace('_', " "),
}
}
fn str_at<'a>(v: &'a serde_json::Value, key: &str) -> &'a str {
v.get(key).and_then(|x| x.as_str()).unwrap_or("")
}
fn target_id(t: &serde_json::Value) -> String {
t.get("_id")
.and_then(|o| o.get("$oid"))
.and_then(|s| s.as_str())
.unwrap_or_default()
.to_string()
}
/// The applicable-scans matrix for one target, fetched on expand.
#[component]
fn TargetScans(id: String) -> Element {
let scan_id = id.clone();
let scans = use_resource(move || {
let id = scan_id.clone();
async move { fetch_applicable_scans(id).await.ok() }
});
let snapshot = scans.read().clone();
match &snapshot {
Some(Some(resp)) => {
let rows = resp.data.scans.clone();
let pentest = resp.data.pentest_supported;
rsx! {
div { style: "margin-top: 8px;",
if rows.is_empty() {
div { style: "opacity: 0.6;", "No scans available (no code / URL / firmware artifact present)." }
}
for s in rows {
{
let name = str_at(&s, "scan").to_string();
let rationale = str_at(&s, "rationale").to_string();
let blocked = s.get("blocked_reason").and_then(|b| b.as_str()).map(String::from);
let default_on = s.get("default_on").and_then(|b| b.as_bool()).unwrap_or(false);
let badge = if blocked.is_some() {
"badge badge-info"
} else if default_on {
"badge badge-success"
} else {
"badge"
};
rsx! {
div { style: "display: flex; gap: 8px; align-items: center; padding: 4px 0;",
span { class: "{badge}", "{name}" }
span { style: "opacity: 0.8; font-size: 0.9em;", "{rationale}" }
if let Some(b) = blocked {
span { style: "opacity: 0.6; font-style: italic; font-size: 0.9em;", "— {b}" }
}
}
}
}
}
if pentest {
div { style: "margin-top: 6px; opacity: 0.75; font-size: 0.85em;",
"Active pentest is supported for this target type."
}
}
}
}
}
Some(None) => rsx! { div { style: "opacity: 0.6;", "Failed to load applicable scans." } },
None => rsx! { div { style: "opacity: 0.6;", "Loading scans..." } },
}
}
#[component]
pub fn TargetsPage() -> Element {
let mut toasts = use_context::<Toasts>();
let mut scanning_ids = use_signal(Vec::<String>::new);
let mut expanded_ids = use_signal(Vec::<String>::new);
let mut confirm_delete = use_signal(|| Option::<(String, String)>::None);
let mut targets = use_resource(move || async move { fetch_targets().await.ok() });
rsx! {
PageHeader {
title: "Targets",
description: "Onboarded targets and what their scans found. Add new targets from Onboard.",
}
div { style: "margin-bottom: 16px; display: flex; gap: 8px;",
Link { to: crate::app::Route::OnboardingPage {}, class: "btn btn-primary",
"+ Onboard a target"
}
button {
class: "btn btn-secondary",
onclick: move |_| targets.restart(),
"Refresh"
}
}
// ── Delete confirmation ──
if let Some((del_id, del_name)) = confirm_delete() {
div { class: "modal-overlay",
div { class: "modal-dialog",
h3 { "Delete Target" }
p { "Delete " strong { "{del_name}" } "?" }
p { class: "modal-warning",
"This permanently removes the target and its findings, SBOM entries, scan runs, and CVE alerts."
}
div { class: "modal-actions",
button {
class: "btn btn-secondary",
onclick: move |_| confirm_delete.set(None),
"Cancel"
}
button {
class: "btn btn-danger",
onclick: move |_| {
let id = del_id.clone();
let name = del_name.clone();
confirm_delete.set(None);
spawn(async move {
match delete_target(id).await {
Ok(_) => {
toasts.push(ToastType::Success, format!("{name} deleted"));
targets.restart();
}
Err(e) => toasts.push(ToastType::Error, e.to_string()),
}
});
},
"Delete"
}
}
}
}
}
{
let targets_snapshot = targets.read().clone();
match &targets_snapshot {
Some(Some(resp)) => {
let rows = resp.data.clone();
if rows.is_empty() {
rsx! {
div { class: "card", style: "padding: 24px; text-align: center;",
p { style: "opacity: 0.7;", "No targets yet." }
Link { to: crate::app::Route::OnboardingPage {}, class: "btn btn-primary",
"Onboard your first target"
}
}
}
} else {
rsx! {
div { class: "card",
div { class: "table-wrapper",
table {
thead {
tr {
th { "Name" }
th { "Type" }
th { "Detected" }
th { "Artifacts" }
th { "Findings" }
th { "Actions" }
}
}
tbody {
for t in rows {
{
let id = target_id(&t);
let name = str_at(&t, "name").to_string();
let ttype = pretty_type(str_at(&t, "target_type"));
let suggested = t
.get("classification")
.and_then(|c| c.get("suggested"))
.and_then(|s| s.as_str())
.map(pretty_type);
let artifacts = t
.get("artifacts")
.and_then(|a| a.as_array())
.cloned()
.unwrap_or_default();
let findings = t
.get("findings_count")
.and_then(|n| n.as_u64())
.unwrap_or(0);
let facts = t
.get("classification")
.and_then(|c| c.get("facts"))
.and_then(|f| f.as_array())
.cloned()
.unwrap_or_default();
let is_scanning = scanning_ids().contains(&id);
let is_expanded = expanded_ids().contains(&id);
let id_scan = id.clone();
let id_exp = id.clone();
let id_del = id.clone();
let name_del = name.clone();
let artifacts_detail = artifacts.clone();
rsx! {
tr {
td { strong { "{name}" } }
td { "{ttype}" }
td {
if let Some(sug) = suggested.clone() {
span { class: "badge badge-success", "{sug}" }
} else {
span { style: "opacity: 0.5;", "" }
}
}
td { "{artifacts.len()}" }
td { "{findings}" }
td { style: "display: flex; gap: 4px;",
button {
class: "btn btn-ghost",
title: "Details",
onclick: move |_| {
let mut ids = expanded_ids();
if ids.contains(&id_exp) {
ids.retain(|i| i != &id_exp);
} else {
ids.push(id_exp.clone());
}
expanded_ids.set(ids);
},
Icon { icon: BsInfoCircle, width: 16, height: 16 }
}
button {
class: if is_scanning { "btn btn-ghost btn-scanning" } else { "btn btn-ghost" },
title: "Run scan",
disabled: is_scanning,
onclick: move |_| {
let id = id_scan.clone();
let mut ids = scanning_ids();
ids.push(id.clone());
scanning_ids.set(ids);
spawn(async move {
match trigger_target_scan(id.clone()).await {
Ok(_) => toasts.push(ToastType::Success, "Scan triggered — findings appear as it runs. Use Refresh."),
Err(e) => toasts.push(ToastType::Error, e.to_string()),
}
let mut ids = scanning_ids();
ids.retain(|i| i != &id);
scanning_ids.set(ids);
});
},
if is_scanning {
span { class: "spinner" }
} else {
Icon { icon: BsPlayCircle, width: 16, height: 16 }
}
}
button {
class: "btn btn-ghost btn-ghost-danger",
title: "Delete target",
onclick: move |_| {
confirm_delete.set(Some((id_del.clone(), name_del.clone())));
},
Icon { icon: BsTrash, width: 16, height: 16 }
}
}
}
if is_expanded {
tr {
td { colspan: "6",
div { style: "padding: 12px 8px;",
h4 { style: "margin: 0 0 6px;", "Artifacts" }
if artifacts_detail.is_empty() {
div { style: "opacity: 0.6;", "No artifacts." }
}
for a in artifacts_detail {
div { style: "font-size: 0.9em; padding: 2px 0;",
span { style: "opacity: 0.7;", "{str_at(&a, \"kind\")}: " }
span { style: "font-family: monospace;", "{str_at(&a, \"source_ref\")}" }
}
}
if !facts.is_empty() {
h4 { style: "margin: 12px 0 6px;", "Detected facts" }
for f in facts {
div { style: "font-size: 0.9em; padding: 2px 0;",
span { style: "font-family: monospace;", "{str_at(&f, \"key\")}={str_at(&f, \"value\")}" }
span { style: "opacity: 0.5;", " ({str_at(&f, \"source\")})" }
}
}
}
h4 { style: "margin: 12px 0 6px;", "Applicable scans" }
TargetScans { id: id.clone() }
}
}
}
}
}
}
}
}
}
}
}
}
}
}
Some(None) => rsx! {
div { class: "card", p { "Failed to load targets." } }
},
None => rsx! {
div { class: "loading", "Loading targets..." }
},
}
}
}
}
-19
View File
@@ -1,19 +0,0 @@
#!/usr/bin/env bash
# Seed the nix store on first start, then run the agent.
#
# The firmware-SBOM pipeline drives a real `nix` build (tramiton NixBackend).
# The image ships the store as a bootstrap tarball rather than baking /nix, so a
# persistent /nix volume (mounted empty on first deploy) gets populated once and
# then survives redeploys. Seeding is best-effort: if it fails, the agent still
# starts and firmware SBOMs fall back to analysis-only.
if [ ! -e /nix/store ]; then
echo "agent-entrypoint: seeding /nix store from image bootstrap..."
mkdir -p /nix
if tar -C / -xzf /opt/nix-bootstrap.tar.gz; then
echo "agent-entrypoint: /nix store seeded."
else
echo "agent-entrypoint: WARN nix seed failed; firmware SBOM will use analysis-only fallback."
fi
fi
exec compliance-agent "$@"