Compare commits

..
Author SHA1 Message Date
Sharang ParnerkarandClaude Fable 5 61c6eab54e docs: document Nuclei/ZAP as decided deterministic detectors under the agentic DAST (not yet implemented)
CI / Check (push) Skipped
CI / Check (pull_request) Failing after 2m55s
CI / Detect Changes (pull_request) Skipped
CI / Deploy Agent (pull_request) Skipped
CI / Deploy Dashboard (pull_request) Skipped
CI / Deploy Docs (pull_request) Skipped
CI / Deploy MCP (pull_request) Skipped
Spec v0.2.1 decision (2026-08-31): keep the agentic DAST/pentest layer, integrate
Nuclei then ZAP baseline underneath as deterministic detectors feeding the
control-map LUT; offline vuln DB (Trivy/Grype) on the on-prem runner trigger.
Clearly marked as planned so the docs stay truthful until the code lands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgxGHn22YEfQz5fLHSHkLv
2026-08-31 15:12:41 +02:00
Sharang ParnerkarandClaude Fable 5 b7b9c812ab docs: align tool inventory with what actually runs (no Grype, no ZAP/nuclei)
CI / Check (push) Skipped
CI / Check (pull_request) Failing after 3m1s
CI / Detect Changes (pull_request) Skipped
CI / Deploy Agent (pull_request) Skipped
CI / Deploy Dashboard (pull_request) Skipped
CI / Deploy Docs (pull_request) Skipped
CI / Deploy MCP (pull_request) Skipped
CVE matching is done directly against OSV.dev (by purl) and NVD (CVSS, CODESYS
CPE) in pipeline/cve.rs; Grype is not installed or invoked. DAST/pentest use the
in-house compliance-dast agents, not ZAP or nuclei. Found while reconciling the
Certifai requirements spec (Collectives) with the code.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EgxGHn22YEfQz5fLHSHkLv
2026-08-31 15:05:06 +02:00
sharang 0834547a74 chore(ci): repoint registry/git/cargo meghsakha.com -> breakpilot.com (#231)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 5s
CI / Deploy Docs (push) Skipped
CI / Deploy Dashboard (push) Failing after 4m47s
CI / Deploy MCP (push) Failing after 15s
CI / Deploy Agent (push) Failing after 18s
2026-08-06 08:42:45 +00:00
sharang effc3080e0 fix(orchestrator): refresh control_refs on PLC re-scans (#230)
CI / Detect Changes (push) Successful in 3s
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Deploy MCP (push) Skipped
CI / Deploy Agent (push) Successful in 3m49s
CI / Check (push) Skipped
2026-07-22 17:06:08 +00:00
sharang 51bab61f77 fix(orchestrator): run semantic control mapping on PLC findings (#229)
CI / Deploy MCP (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Check (push) Skipped
CI / Deploy Agent (push) Successful in 4m3s
2026-07-22 13:16:53 +00:00
sharang 5bdc35ee92 fix(orchestrator): refresh control_refs on existing findings during re-scan (#228)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Agent (push) Successful in 4m13s
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Deploy MCP (push) Skipped
2026-07-22 12:24:36 +00:00
sharang ea516cc054 docs(control-mapping): MCP emission loop + default-on flags (#227)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Agent (push) Skipped
CI / Deploy Docs (push) Successful in 57s
CI / Deploy Dashboard (push) Skipped
CI / Deploy MCP (push) Skipped
2026-07-22 11:23:51 +00:00
sharang 7d5c95ddb8 fix(mcp): bind tenant to session — bearer context was lost over HTTP (#226)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 2s
CI / Deploy Agent (push) Skipped
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Deploy MCP (push) Successful in 1m57s
2026-07-22 09:30:07 +00:00
sharang 3e233da128 feat(controls): promote grounded controls to covered + enable LLM passes by default (#224)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Docs (push) Skipped
CI / Deploy Agent (push) Successful in 7m53s
CI / Deploy Dashboard (push) Successful in 7m10s
CI / Deploy MCP (push) Successful in 1m50s
2026-07-22 07:48:51 +00:00
sharang a72f79e557 ci: fix cosign signing (install fallback, env-style login, portal sign step) (#225)
CI / Deploy Docs (push) Canceled after 0s
CI / Check (push) Canceled after 0s
CI / Detect Changes (push) Canceled after 0s
CI / Deploy Agent (push) Canceled after 0s
CI / Deploy Dashboard (push) Canceled after 0s
CI / Deploy MCP (push) Canceled after 0s
2026-07-22 07:46:34 +00:00
sharang 182dec69b8 docs(features): compliance control-mapping pipeline (#223)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Agent (push) Skipped
CI / Deploy Dashboard (push) Skipped
CI / Deploy MCP (push) Skipped
CI / Deploy Docs (push) Failing after 4s
2026-07-22 07:30:10 +00:00
sharang a25d41c3e5 feat(controls): enrich semantic retrieval query with finding intent + C5 live test (#222)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Deploy MCP (push) Skipped
CI / Deploy Agent (push) Failing after 4s
2026-07-22 07:14:49 +00:00
sharang 60601d8215 fix(llm): chunk embed() requests under the backend batch cap (#221)
CI / Check (push) Skipped
CI / Detect Changes (push) Successful in 3s
CI / Deploy Dashboard (push) Skipped
CI / Deploy Docs (push) Skipped
CI / Deploy MCP (push) Skipped
CI / Deploy Agent (push) Failing after 5s
2026-07-21 15:41:07 +00:00
23 changed files with 770 additions and 119 deletions
+26 -15
View File
@@ -7,6 +7,13 @@ on:
pull_request:
env:
# registry + cosign creds via env, NOT inline ${{ }}: the Harbor robot
# username contains '$', which sh expands when interpolated into the
# script (robot$ci-push -> robot-push) => docker login unauthorized.
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
COSIGN_KEY: ${{ secrets.COSIGN_KEY }}
COSIGN_PASSWORD: ${{ secrets.COSIGN_PASSWORD }}
CARGO_TERM_COLOR: always
RUSTFLAGS: "-D warnings"
# Compile cache: sccache -> Hetzner S3 (breakpilot-sccache), runner-independent
@@ -65,7 +72,7 @@ jobs:
echo '[source.crates-io]'
echo 'replace-with = "kellnr"'
echo '[registries.kellnr]'
echo 'index = "sparse+https://crates.meghsakha.com/api/v1/cratesio/"'
echo 'index = "sparse+https://crates.breakpilot.com/api/v1/cratesio/"'
} >> "$CARGO_HOME/config.toml"
env:
RUSTC_WRAPPER: ""
@@ -87,8 +94,8 @@ jobs:
- name: Configure git auth for private tramiton dependency
run: |
git config --global \
url."https://sharang:${{ secrets.TRAMITON_FETCH_TOKEN }}@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"
url."https://sharang:${{ secrets.TRAMITON_FETCH_TOKEN }}@git.breakpilot.com/".insteadOf \
"ssh://git@git.breakpilot.com:22222/"
env:
RUSTC_WRAPPER: ""
@@ -206,12 +213,13 @@ jobs:
apk add --no-cache git curl openssl
git init && git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}" && git checkout FETCH_HEAD
IMAGE=repo.meghsakha.com/certifai/compliance-agent
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login repo.meghsakha.com -u "${{ secrets.REGISTRY_USERNAME }}" --password-stdin
IMAGE=repo.breakpilot.com/certifai/compliance-agent
echo "$REGISTRY_PASSWORD" | docker login repo.breakpilot.com -u "$REGISTRY_USERNAME" --password-stdin
DOCKER_BUILDKIT=1 docker build --secret id=tramiton_token,env=TRAMITON_FETCH_TOKEN \
-f Dockerfile.agent -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker push "$IMAGE:latest" && docker push "$IMAGE:${GITHUB_SHA}"
command -v cosign >/dev/null 2>&1 || { curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 && chmod +x /usr/local/bin/cosign; }
{ command -v cosign >/dev/null 2>&1 || curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 || wget -qO /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64; } || echo "::warning::cosign fetch failed"
chmod +x /usr/local/bin/cosign 2>/dev/null || true
cosign sign --yes --key env://COSIGN_KEY "$IMAGE:latest" || echo "::warning::cosign failed"
PAYLOAD=$(printf '{"ref":"refs/heads/main","repository":{"full_name":"sharang/compliance-scanner-agent"},"head_commit":{"id":"%s","message":"deploy agent"}}' "${GITHUB_SHA}")
SIG=$(printf '%s' "$PAYLOAD" | openssl dgst -sha256 -hmac "${{ secrets.ORCA_WEBHOOK_SECRET }}" | awk '{print $2}')
@@ -232,12 +240,13 @@ jobs:
apk add --no-cache git curl openssl
git init && git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}" && git checkout FETCH_HEAD
IMAGE=repo.meghsakha.com/certifai/compliance-dashboard
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login repo.meghsakha.com -u "${{ secrets.REGISTRY_USERNAME }}" --password-stdin
IMAGE=repo.breakpilot.com/certifai/compliance-dashboard
echo "$REGISTRY_PASSWORD" | docker login repo.breakpilot.com -u "$REGISTRY_USERNAME" --password-stdin
DOCKER_BUILDKIT=1 docker build --secret id=tramiton_token,env=TRAMITON_FETCH_TOKEN \
-f Dockerfile.dashboard -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker push "$IMAGE:latest" && docker push "$IMAGE:${GITHUB_SHA}"
command -v cosign >/dev/null 2>&1 || { curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 && chmod +x /usr/local/bin/cosign; }
{ command -v cosign >/dev/null 2>&1 || curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 || wget -qO /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64; } || echo "::warning::cosign fetch failed"
chmod +x /usr/local/bin/cosign 2>/dev/null || true
cosign sign --yes --key env://COSIGN_KEY "$IMAGE:latest" || echo "::warning::cosign failed"
PAYLOAD=$(printf '{"ref":"refs/heads/main","repository":{"full_name":"sharang/compliance-scanner-agent"},"head_commit":{"id":"%s","message":"deploy dashboard"}}' "${GITHUB_SHA}")
SIG=$(printf '%s' "$PAYLOAD" | openssl dgst -sha256 -hmac "${{ secrets.ORCA_WEBHOOK_SECRET }}" | awk '{print $2}')
@@ -256,11 +265,12 @@ jobs:
apk add --no-cache git curl openssl
git init && git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}" && git checkout FETCH_HEAD
IMAGE=repo.meghsakha.com/certifai/compliance-docs
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login repo.meghsakha.com -u "${{ secrets.REGISTRY_USERNAME }}" --password-stdin
IMAGE=repo.breakpilot.com/certifai/compliance-docs
echo "$REGISTRY_PASSWORD" | docker login repo.breakpilot.com -u "$REGISTRY_USERNAME" --password-stdin
docker build -f Dockerfile.docs -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker push "$IMAGE:latest" && docker push "$IMAGE:${GITHUB_SHA}"
command -v cosign >/dev/null 2>&1 || { curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 && chmod +x /usr/local/bin/cosign; }
{ command -v cosign >/dev/null 2>&1 || curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 || wget -qO /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64; } || echo "::warning::cosign fetch failed"
chmod +x /usr/local/bin/cosign 2>/dev/null || true
cosign sign --yes --key env://COSIGN_KEY "$IMAGE:latest" || echo "::warning::cosign failed"
PAYLOAD=$(printf '{"ref":"refs/heads/main","repository":{"full_name":"sharang/compliance-scanner-agent"},"head_commit":{"id":"%s","message":"deploy docs"}}' "${GITHUB_SHA}")
SIG=$(printf '%s' "$PAYLOAD" | openssl dgst -sha256 -hmac "${{ secrets.ORCA_WEBHOOK_SECRET }}" | awk '{print $2}')
@@ -281,12 +291,13 @@ jobs:
apk add --no-cache git curl openssl
git init && git remote add origin "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}.git"
git fetch --depth=1 origin "${GITHUB_SHA}" && git checkout FETCH_HEAD
IMAGE=repo.meghsakha.com/certifai/compliance-mcp
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login repo.meghsakha.com -u "${{ secrets.REGISTRY_USERNAME }}" --password-stdin
IMAGE=repo.breakpilot.com/certifai/compliance-mcp
echo "$REGISTRY_PASSWORD" | docker login repo.breakpilot.com -u "$REGISTRY_USERNAME" --password-stdin
DOCKER_BUILDKIT=1 docker build --secret id=tramiton_token,env=TRAMITON_FETCH_TOKEN \
-f Dockerfile.mcp -t "$IMAGE:latest" -t "$IMAGE:${GITHUB_SHA}" .
docker push "$IMAGE:latest" && docker push "$IMAGE:${GITHUB_SHA}"
command -v cosign >/dev/null 2>&1 || { curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 && chmod +x /usr/local/bin/cosign; }
{ command -v cosign >/dev/null 2>&1 || curl -sSfLo /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64 || wget -qO /usr/local/bin/cosign https://github.com/sigstore/cosign/releases/download/v2.4.3/cosign-linux-amd64; } || echo "::warning::cosign fetch failed"
chmod +x /usr/local/bin/cosign 2>/dev/null || true
cosign sign --yes --key env://COSIGN_KEY "$IMAGE:latest" || echo "::warning::cosign failed"
PAYLOAD=$(printf '{"ref":"refs/heads/main","repository":{"full_name":"sharang/compliance-scanner-agent"},"head_commit":{"id":"%s","message":"deploy mcp"}}' "${GITHUB_SHA}")
SIG=$(printf '%s' "$PAYLOAD" | openssl dgst -sha256 -hmac "${{ secrets.ORCA_WEBHOOK_SECRET }}" | awk '{print $2}')
+2 -2
View File
@@ -8,8 +8,8 @@ COPY . .
RUN --mount=type=secret,id=tramiton_token \
if [ -s /run/secrets/tramiton_token ]; then \
git config --global \
url."https://sharang:$(cat /run/secrets/tramiton_token)@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"; \
url."https://sharang:$(cat /run/secrets/tramiton_token)@git.breakpilot.com/".insteadOf \
"ssh://git@git.breakpilot.com:22222/"; \
fi && \
CARGO_NET_GIT_FETCH_WITH_CLI=true cargo build --release -p compliance-agent
+2 -2
View File
@@ -13,8 +13,8 @@ ENV DOCS_URL=${DOCS_URL}
RUN --mount=type=secret,id=tramiton_token \
if [ -s /run/secrets/tramiton_token ]; then \
git config --global \
url."https://sharang:$(cat /run/secrets/tramiton_token)@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"; \
url."https://sharang:$(cat /run/secrets/tramiton_token)@git.breakpilot.com/".insteadOf \
"ssh://git@git.breakpilot.com:22222/"; \
fi && \
CARGO_NET_GIT_FETCH_WITH_CLI=true dx build --release --package compliance-dashboard
+2 -2
View File
@@ -8,8 +8,8 @@ COPY . .
RUN --mount=type=secret,id=tramiton_token \
if [ -s /run/secrets/tramiton_token ]; then \
git config --global \
url."https://sharang:$(cat /run/secrets/tramiton_token)@gitea.meghsakha.com/".insteadOf \
"ssh://git@gitea.meghsakha.com:22222/"; \
url."https://sharang:$(cat /run/secrets/tramiton_token)@git.breakpilot.com/".insteadOf \
"ssh://git@git.breakpilot.com:22222/"; \
fi && \
CARGO_NET_GIT_FETCH_WITH_CLI=true cargo build --release -p compliance-mcp
+19 -11
View File
@@ -112,9 +112,9 @@ async fn build_specs(provider: &OscalControlsProvider) -> HashMap<String, Contro
/// the grounded judge decide whether the control holds there. Returns net-new
/// findings, each already tagged with its control and grounded to a real snippet.
///
/// Gated: the orchestrator runs this only when `breakpilot.grounded_control_checks`
/// is set. Absence detection is the least deterministic path (the judge decides
/// presence/absence, not a syntactic pattern), so it stays off until tuned live.
/// The orchestrator runs this when `breakpilot.grounded_control_checks` is set
/// (on by default). Validated live; it covers the 8 absence-based CRA controls
/// (the judge decides presence/absence, grounded to a real snippet).
pub async fn grounded_surface_findings(
config: &AgentConfig,
llm: Arc<LlmClient>,
@@ -173,11 +173,10 @@ fn fetch_region(repo_path: &Path, file: &str, line: u32) -> Option<CandidateRegi
/// ~13.6k master-control corpus (which has no CWE to LUT on). Returns the number
/// of findings that gained a master-control ref.
///
/// Gated: the orchestrator runs this only when `breakpilot.semantic_mapping` is
/// set (default off, flipped on once the master-controls catalog is live). The
/// control embedding index is built once and cached to `snapshot_dir` keyed by
/// corpus hash ([`ControlIndex::load_or_build`]), so only the first scan after a
/// catalog change pays the embedding cost.
/// The orchestrator runs this when `breakpilot.semantic_mapping` is set (on by
/// default). The control embedding index is built once and cached to
/// `snapshot_dir` keyed by corpus hash ([`ControlIndex::load_or_build`]), so only
/// the first scan after a catalog change pays the embedding cost.
pub async fn semantic_stamp_findings(
config: &AgentConfig,
llm: Arc<LlmClient>,
@@ -234,13 +233,22 @@ pub async fn semantic_stamp_findings(
let Some(region) = fetch_region(repo_path, &file, line) else {
continue;
};
let region_emb = match llm.embed(vec![region.content.clone()]).await {
// Retrieve on the finding's intent + the code, not the region alone: two
// findings in one file share overlapping windows and otherwise embed alike,
// collapsing onto the same controls. The finding's title/description carry
// the discriminating signal (e.g. "brute-force protection" vs "weak hash").
// The raw `region` still goes to the judge for snippet grounding.
let query = format!(
"{}\n{}\n\n{}",
finding.title, finding.description, region.content
);
let query_emb = match llm.embed(vec![query]).await {
Ok(mut embs) => match embs.pop() {
Some(v) => v,
None => continue,
},
Err(e) => {
tracing::warn!(error = %e, "region embed failed; skipping finding");
tracing::warn!(error = %e, "query embed failed; skipping finding");
continue;
}
};
@@ -248,7 +256,7 @@ pub async fn semantic_stamp_findings(
.check(
&index,
&region,
&region_emb,
&query_emb,
SEMANTIC_TOP_K,
&finding.repo_id,
)
+7 -5
View File
@@ -22,18 +22,20 @@ impl<J: ControlJudge> SemanticControlChecker<J> {
Self { judge }
}
/// Map a code region to the controls it violates. `region_embedding` is the
/// region's embedding (the caller computes it via the LLM); the top-`k`
/// nearest controls in `index` are judged and grounded.
/// Map a code region to the controls it violates. `query_embedding` is the
/// caller-supplied retrieval embedding — typically the finding's intent
/// (title/description) plus the region, so retrieval keys on what the finding
/// is *about*, not just the ambient code. The top-`k` nearest controls in
/// `index` are then judged against the raw `region` and grounded.
pub async fn check(
&self,
index: &ControlIndex,
region: &CandidateRegion,
region_embedding: &[f64],
query_embedding: &[f64],
k: usize,
repo_id: &str,
) -> Vec<Finding> {
let candidates = index.nearest(region_embedding, k);
let candidates = index.nearest(query_embedding, k);
let mut findings = Vec::new();
for spec in &candidates {
let verdict = self.judge.judge(spec, region).await;
+60 -7
View File
@@ -232,9 +232,9 @@ impl PipelineOrchestrator {
// Stage 5c: semantic control mapping — scale path for the master-controls
// corpus (no CWE to LUT on): embed each finding's region, retrieve the
// nearest master controls, grounded-judge, and stamp confirmed refs. Gated
// (default off) as the corpus embedding + per-finding judging is the heavy
// path; enabled once verified live against a deployed master-controls catalog.
// nearest master controls, grounded-judge, and stamp confirmed refs. On by
// default (validated live); the corpus embedding is cached so only the
// first scan after a catalog change pays it.
if self.config.breakpilot.semantic_mapping {
self.update_phase(scan_run_id, "semantic_control_mapping")
.await;
@@ -256,8 +256,8 @@ impl PipelineOrchestrator {
// rate limiting, no security logging, no update-signature check) have no
// syntactic pattern to match, so we retrieve the code surface each governs
// and let the grounded judge decide whether it holds, producing net-new
// findings already tagged + grounded. Gated (default off): absence
// detection is the least deterministic path, kept off until tuned live.
// findings already tagged + grounded. On by default (validated live); it
// covers the 8 absence-based CRA controls.
if self.config.breakpilot.grounded_control_checks {
self.update_phase(scan_run_id, "grounded_control_checks")
.await;
@@ -277,8 +277,10 @@ impl PipelineOrchestrator {
}
}
// Dedup against existing findings and insert new ones
// Dedup against existing findings: insert first-seen ones, and refresh the
// control mappings on ones we've seen before.
let mut new_count = 0u32;
let mut refreshed_count = 0u32;
let mut new_findings: Vec<Finding> = Vec::new();
for mut finding in all_findings {
finding.scan_run_id = Some(scan_run_id.to_string());
@@ -293,8 +295,25 @@ impl PipelineOrchestrator {
finding.id = result.inserted_id.as_object_id();
new_findings.push(finding);
new_count += 1;
} else if !finding.control_refs.is_empty() {
// Re-scan refresh: a mapping pass (newly enabled or tuned) computed
// control_refs for a finding first seen before mapping ran. Persist
// them onto the existing row — the insert path alone never would.
self.db
.findings()
.update_one(
doc! { "fingerprint": &finding.fingerprint },
doc! { "$set": { "control_refs": finding.control_refs.clone() } },
)
.await?;
refreshed_count += 1;
}
}
if refreshed_count > 0 {
tracing::info!(
"[{repo_id}] Refreshed control_refs on {refreshed_count} existing findings"
);
}
// Remove stale SBOM entries for this repo before reinserting
if !sbom_entries.is_empty() {
@@ -567,7 +586,21 @@ impl PipelineOrchestrator {
let Some(path) = ingest_set.get(&a.id).and_then(|ia| ia.working_path.clone()) else {
continue;
};
all_findings.extend(crate::pipeline::plc::analyze_tree(&path, target_id));
let mut source_findings = crate::pipeline::plc::analyze_tree(&path, target_id);
// Control mapping for the PLC path (run_plc_scan is separate from
// run_pipeline, which does its own mapping). PLC findings carry
// file_path/line/cwe, so the semantic pass reads each region under this
// source's `path` and stamps master-control refs. The LUT + grounded
// surface passes are code-pattern / CRA-specific and don't apply to
// IEC 61131-3 control logic, so only the semantic pass runs here.
crate::controls::semantic_stamp_findings(
&self.config,
self.llm.clone(),
&path,
&mut source_findings,
)
.await;
all_findings.extend(source_findings);
// Control-application SBOM: CODESYS libraries + runtime from a
// `.projectarchive` (uploaded, or committed in the working tree).
let archive = a
@@ -592,6 +625,7 @@ impl PipelineOrchestrator {
);
let mut new_count = 0u32;
let mut refreshed_count = 0u32;
for mut finding in all_findings {
finding.scan_run_id = Some(scan_run_id.to_string());
if self
@@ -603,8 +637,27 @@ impl PipelineOrchestrator {
{
self.db.findings().insert_one(&finding).await?;
new_count += 1;
} else if !finding.control_refs.is_empty() {
// Re-scan refresh: mirror run_pipeline — persist newly-computed
// control_refs onto a PLC finding first seen before the semantic
// pass ran. The insert path alone never would, so without this a
// PLC re-scan can only pick up mappings via a delete + re-add.
self.db
.findings()
.update_one(
doc! { "fingerprint": &finding.fingerprint },
doc! { "$set": { "control_refs": finding.control_refs.clone() } },
)
.await?;
refreshed_count += 1;
}
}
if refreshed_count > 0 {
tracing::info!(
target_id,
"Refreshed control_refs on {refreshed_count} existing PLC findings"
);
}
if !all_sbom.is_empty() {
if let Err(e) = self
+125
View File
@@ -0,0 +1,125 @@
//! C5 example 2 — exploratory (not a committed regression test). Four topically
//! distinct findings, to see whether tuned semantic retrieval maps each to the
//! right master-control family. Run:
//! export ... (LITELLM_* + BREAKPILOT_BASE_URL)
//! cargo test -p compliance-agent --test c5_example2 -- --ignored --nocapture
mod common;
use std::sync::Arc;
use compliance_agent::llm::LlmClient;
use compliance_core::config::BreakpilotConfig;
use compliance_core::models::finding::{Finding, Severity};
use compliance_core::models::scan::ScanType;
use secrecy::SecretString;
fn env(k: &str) -> String {
std::env::var(k).unwrap_or_else(|_| panic!("env {k} must be set"))
}
fn mk(file: &str, line: u32, title: &str, desc: &str) -> Finding {
let mut f = Finding::new(
"repo-c5b".into(),
format!("{file}:{line}"),
"semgrep".into(),
ScanType::Sast,
title.into(),
desc.into(),
Severity::High,
);
f.file_path = Some(file.into());
f.line_number = Some(line);
f
}
fn write(repo: &std::path::Path, rel: &str, body: &str) {
let p = repo.join(rel);
if let Some(parent) = p.parent() {
std::fs::create_dir_all(parent).unwrap();
}
std::fs::write(p, body).unwrap();
}
#[tokio::test]
#[ignore = "live: api-dev + LiteLLM"]
async fn c5b_varied_findings() {
let llm = Arc::new(LlmClient::new(
env("LITELLM_URL"),
SecretString::from(env("LITELLM_API_KEY")),
env("LITELLM_MODEL"),
env("LITELLM_EMBED_MODEL"),
));
let mut config = common::dev_config("mongodb://unused".into(), "c5b".into());
config.breakpilot = BreakpilotConfig {
base_url: Some(env("BREAKPILOT_BASE_URL")),
token: None,
snapshot_dir: std::env::temp_dir()
.join("c5-oscal-snap")
.to_string_lossy()
.into_owned(),
semantic_mapping: true,
grounded_control_checks: false,
};
let repo = std::env::temp_dir().join("c5b-fixture-repo");
let _ = std::fs::remove_dir_all(&repo);
write(
&repo,
"app/db.py",
"import sqlite3\n\ndef get_user(username):\n q = \"SELECT * FROM users WHERE name = '\" + username + \"'\"\n return conn.execute(q)\n",
);
write(
&repo,
"app/config.py",
"# service config\nAPI_KEY = \"sk_live_51H8xYz3kQ9v2bNmR7wT4uSpQ\"\nDB_HOST = \"db.internal\"\n",
);
write(
&repo,
"app/net.py",
"import requests\n\ndef fetch(url):\n return requests.get(url, verify=False, timeout=5)\n",
);
write(
&repo,
"app/ser.py",
"import pickle\n\ndef load_state(blob):\n return pickle.loads(blob)\n",
);
let mut findings = vec![
mk(
"app/db.py",
4,
"SQL injection via string-concatenated query",
"User input is concatenated directly into a SQL statement, allowing SQL injection.",
),
mk(
"app/config.py",
2,
"Hardcoded API credential in source",
"A live API key is hardcoded in source code instead of a secret store.",
),
mk(
"app/net.py",
4,
"TLS certificate verification disabled",
"requests is called with verify=False, disabling TLS certificate validation.",
),
mk(
"app/ser.py",
3,
"Insecure deserialization with pickle.loads",
"Untrusted data is deserialized with pickle.loads, allowing remote code execution.",
),
];
let tagged =
compliance_agent::controls::semantic_stamp_findings(&config, llm, &repo, &mut findings)
.await;
println!("\n=== C5 example 2: varied findings ===");
for f in &findings {
println!(" {:52} -> {:?}", f.title, f.control_refs);
}
println!("tagged: {tagged}/4");
let _ = std::fs::remove_dir_all(&repo);
assert!(tagged >= 1);
}
+145
View File
@@ -0,0 +1,145 @@
//! C5 live verification — the semantic master-controls path end to end against the
//! deployed api-dev catalog. Ignored (hits api-dev + LiteLLM). Run explicitly:
//!
//! set -a; . ./.env; set +a
//! BREAKPILOT_BASE_URL=https://api-dev.breakpilot.ai \
//! cargo test -p compliance-agent --test c5_semantic_live -- --ignored --nocapture
//!
//! Pulls the live master-controls catalog, embeds the corpus (chunked), then for a
//! couple of real vulnerable findings retrieves the nearest master controls and
//! grounded-judges them, stamping master-control refs.
mod common;
use std::sync::Arc;
use compliance_agent::llm::LlmClient;
use compliance_core::config::BreakpilotConfig;
use compliance_core::models::finding::{Finding, Severity};
use compliance_core::models::scan::ScanType;
use secrecy::SecretString;
fn env(k: &str) -> String {
std::env::var(k).unwrap_or_else(|_| panic!("env {k} must be set for the live C5 test"))
}
fn mk_finding(file: &str, line: u32, title: &str) -> Finding {
let mut f = Finding::new(
"repo-c5".into(),
format!("{file}:{line}"),
"semgrep".into(),
ScanType::Sast,
title.into(),
title.into(),
Severity::High,
);
f.file_path = Some(file.into());
f.line_number = Some(line);
f
}
#[tokio::test]
#[ignore = "live: requires deployed api-dev master-controls (fetch+parse only, no LLM)"]
async fn c5_ingest_master_controls_catalog() {
use compliance_agent::controls::OscalControlsProvider;
let provider = OscalControlsProvider::new(
reqwest::Client::new(),
env("BREAKPILOT_BASE_URL"),
None,
std::env::temp_dir().join("c5-ingest-snap"),
);
let doc = provider
.load_master_controls()
.await
.expect("pull + parse master-controls catalog");
let controls = doc.to_controls();
println!(
"\n=== C5 ingest: {} master controls parsed ===",
controls.len()
);
for c in controls.iter().take(4) {
let text: String = c.text.chars().take(90).collect();
println!(" {} | {} | {}", c.id, c.title, text);
}
assert!(
!controls.is_empty(),
"expected a non-empty master-control corpus"
);
}
#[tokio::test]
#[ignore = "live: requires deployed api-dev master-controls + LiteLLM"]
async fn c5_semantic_stamps_master_control_refs() {
let llm = Arc::new(LlmClient::new(
env("LITELLM_URL"),
SecretString::from(env("LITELLM_API_KEY")),
env("LITELLM_MODEL"),
env("LITELLM_EMBED_MODEL"),
));
let mut config = common::dev_config("mongodb://unused".into(), "c5".into());
let snapshot = std::env::temp_dir().join("c5-oscal-snap");
config.breakpilot = BreakpilotConfig {
base_url: Some(env("BREAKPILOT_BASE_URL")),
token: None,
snapshot_dir: snapshot.to_string_lossy().into_owned(),
semantic_mapping: true,
grounded_control_checks: false,
};
// Fixture repo with recognizable code-checkable surfaces.
let repo = std::env::temp_dir().join("c5-fixture-repo");
let _ = std::fs::remove_dir_all(&repo);
std::fs::create_dir_all(repo.join("app")).expect("mkdir");
std::fs::write(
repo.join("app/auth.py"),
concat!(
"import hashlib\n",
"\n",
"def store_password(user, password):\n",
" # weak, unsalted password hashing\n",
" digest = hashlib.md5(password.encode()).hexdigest()\n",
" db.save(user, digest)\n",
"\n",
"@app.route('/login', methods=['POST'])\n",
"def login():\n",
" u = request.form['username']\n",
" p = request.form['password']\n",
" return 'ok' if check(u, p) else ('bad', 401)\n",
),
)
.expect("write fixture");
let mut findings = vec![
mk_finding("app/auth.py", 5, "Weak password hash (md5, unsalted)"),
mk_finding(
"app/auth.py",
9,
"Login endpoint without brute-force protection",
),
];
let tagged =
compliance_agent::controls::semantic_stamp_findings(&config, llm, &repo, &mut findings)
.await;
println!("\n=== C5 semantic master-controls stamping ===");
for f in &findings {
println!(
" {:50} {}:{:?} -> {:?}",
f.title,
f.file_path.as_deref().unwrap_or(""),
f.line_number,
f.control_refs
);
}
println!("findings that gained >=1 master-control ref: {tagged}");
let _ = std::fs::remove_dir_all(&repo);
// Live corpus — assert only that the path runs and stamps at least one ref.
assert!(
tagged >= 1,
"expected at least one finding to gain a master-control ref"
);
}
@@ -0,0 +1,92 @@
//! Live validation of the grounded surface path (Stage 5d) for absence-based CRA
//! controls. Ignored (hits api-dev CRA catalog + LiteLLM). Run:
//! export ... (LITELLM_* + BREAKPILOT_BASE_URL)
//! cargo test -p compliance-agent --test grounded_surface_live -- --ignored --nocapture
//!
//! Builds a fixture whose code surfaces trigger several absence-based controls
//! (no rate limiting, no security logging, unverified update) and checks that the
//! grounded checker produces control-tagged findings.
mod common;
use std::sync::Arc;
use compliance_agent::llm::LlmClient;
use compliance_core::config::BreakpilotConfig;
use secrecy::SecretString;
fn env(k: &str) -> String {
std::env::var(k).unwrap_or_else(|_| panic!("env {k} must be set"))
}
fn write(repo: &std::path::Path, rel: &str, body: &str) {
let p = repo.join(rel);
if let Some(parent) = p.parent() {
std::fs::create_dir_all(parent).unwrap();
}
std::fs::write(p, body).unwrap();
}
#[tokio::test]
#[ignore = "live: api-dev CRA catalog + LiteLLM"]
async fn grounded_surface_flags_absence_controls() {
let llm = Arc::new(LlmClient::new(
env("LITELLM_URL"),
SecretString::from(env("LITELLM_API_KEY")),
env("LITELLM_MODEL"),
env("LITELLM_EMBED_MODEL"),
));
let mut config = common::dev_config("mongodb://unused".into(), "grounded".into());
config.breakpilot = BreakpilotConfig {
base_url: Some(env("BREAKPILOT_BASE_URL")),
token: None,
snapshot_dir: std::env::temp_dir()
.join("grounded-snap")
.to_string_lossy()
.into_owned(),
semantic_mapping: false,
grounded_control_checks: true,
};
let repo = std::env::temp_dir().join("grounded-fixture-repo");
let _ = std::fs::remove_dir_all(&repo);
// cra-ai-11: login endpoint with no rate limiting / lockout
write(
&repo,
"app/auth.py",
"@app.route('/login', methods=['POST'])\ndef login():\n u = request.form['username']\n p = request.form['password']\n if authenticate(u, p):\n return redirect('/')\n return 'bad credentials', 401\n",
);
// cra-ai-24: privileged admin action with no security/audit logging
write(
&repo,
"app/admin.py",
"@app.route('/admin/delete_user', methods=['POST'])\ndef admin_delete_user():\n uid = request.form['uid']\n db.users.delete_one({'_id': uid})\n return 'ok', 200\n",
);
// cra-ai-28/29/30: firmware update applied without signature / checksum verification
write(
&repo,
"app/updater.py",
"def apply_firmware_update(url):\n blob = download(url)\n install_firmware(blob)\n reboot_device()\n",
);
let findings =
compliance_agent::controls::grounded_surface_findings(&config, llm, &repo, "repo-grounded")
.await;
println!("\n=== Grounded surface findings ({}) ===", findings.len());
for f in &findings {
println!(
" {:24} {}:{:?} {}",
f.control_refs.join(","),
f.file_path.as_deref().unwrap_or(""),
f.line_number,
f.title
);
}
let _ = std::fs::remove_dir_all(&repo);
assert!(
!findings.is_empty(),
"expected the grounded pass to flag at least one absence-based control"
);
}
+9 -8
View File
@@ -76,14 +76,15 @@ pub struct BreakpilotConfig {
/// Directory for catalog snapshots.
pub snapshot_dir: String,
/// Enable the master-controls **semantic** mapping pass (embed regions,
/// retrieve nearest controls, grounded-judge). Off by default: it is the
/// scale path and stays gated until verified live against a deployed
/// master-controls catalog.
/// Enable the master-controls **semantic** mapping pass (embed regions,
/// retrieve nearest controls, grounded-judge). On by default — validated live
/// against the deployed master-controls catalog. Still a no-op unless
/// `base_url` is set and the catalog is reachable.
pub semantic_mapping: bool,
/// Enable the **grounded surface** pass for absence-based controls (retrieve
/// the code surface a control governs, judge whether it holds). Off by
/// default: absence detection is the least deterministic path and stays gated
/// until tuned against live scans.
/// the code surface a control governs, judge whether it holds). On by default
/// — validated live; it covers the 8 absence-based CRA controls that no
/// syntactic rule can.
pub grounded_control_checks: bool,
}
@@ -93,8 +94,8 @@ impl Default for BreakpilotConfig {
base_url: None,
token: None,
snapshot_dir: "/data/compliance-scanner/oscal".to_string(),
semantic_mapping: false,
grounded_control_checks: false,
semantic_mapping: true,
grounded_control_checks: true,
}
}
}
+16 -9
View File
@@ -42,7 +42,19 @@ async fn main() -> Result<(), Box<dyn std::error::Error>> {
let pool_for_factory = pool.clone();
let service = StreamableHttpService::new(
move || Ok(ComplianceMcpServer::new(pool_for_factory.clone())),
move || {
// The factory runs in the request task, still inside the bearer
// middleware's `TENANT_ID` scope, and BEFORE rmcp spawns the
// session task (which would lose the task_local). So bind the
// tenant into the session's server instance here, once.
let tenant_id = auth::current_tenant_id().ok_or_else(|| {
std::io::Error::other("no tenant context when creating MCP session")
})?;
Ok(ComplianceMcpServer::new(
pool_for_factory.clone(),
tenant_id,
))
},
Arc::new(LocalSessionManager::default()),
StreamableHttpServerConfig::default(),
);
@@ -69,16 +81,11 @@ async fn main() -> Result<(), Box<dyn std::error::Error>> {
tenant_id = %synth_tenant,
"stdio transport — using synthetic tenant id; DO NOT use in production"
);
let server = ComplianceMcpServer::new(pool);
let server = ComplianceMcpServer::new(pool, synth_tenant);
let transport = rmcp::transport::stdio();
use rmcp::ServiceExt;
auth::TENANT_ID
.scope(synth_tenant, async {
let handle = server.serve(transport).await?;
handle.waiting().await?;
Ok::<_, Box<dyn std::error::Error>>(())
})
.await?;
let handle = server.serve(transport).await?;
handle.waiting().await?;
}
Ok(())
+9 -13
View File
@@ -2,37 +2,33 @@ use rmcp::{
handler::server::wrapper::Parameters, model::*, tool, tool_handler, tool_router, ServerHandler,
};
use crate::auth::current_tenant_id;
use crate::database::{Database, DatabasePool};
use crate::tools::{dast, findings, oscal, pentest, sbom};
pub struct ComplianceMcpServer {
pool: DatabasePool,
/// Tenant this session serves. Bound once at session creation (the HTTP
/// factory reads the bearer-set tenant while still in the request scope;
/// stdio passes a synthetic id) — NOT a per-request `task_local`, which is
/// lost across the `tokio::spawn` that runs the Streamable-HTTP session.
tenant_id: String,
#[allow(dead_code)]
tool_router: rmcp::handler::server::router::tool::ToolRouter<Self>,
}
impl ComplianceMcpServer {
/// Resolve the per-tenant `Database` from the bearer-set
/// `task_local`. Every tool handler calls this; missing context
/// surfaces as `internal_error` because it means the auth
/// middleware was misconfigured (handler ran without scope).
/// The per-tenant `Database` for this session.
fn tenant_db(&self) -> Result<Database, rmcp::ErrorData> {
let tenant_id = current_tenant_id().ok_or_else(|| {
rmcp::ErrorData::internal_error(
"no tenant context — bearer middleware not in chain".to_string(),
None,
)
})?;
Ok(self.pool.for_tenant_id(&tenant_id))
Ok(self.pool.for_tenant_id(&self.tenant_id))
}
}
#[tool_router]
impl ComplianceMcpServer {
pub fn new(pool: DatabasePool) -> Self {
pub fn new(pool: DatabasePool, tenant_id: String) -> Self {
Self {
pool,
tenant_id,
tool_router: Self::tool_router(),
}
}
+80 -24
View File
@@ -52,9 +52,16 @@
{
"control": "cra-ai-6",
"title": "Integritaetspruefung",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-7",
@@ -138,16 +145,30 @@
{
"control": "cra-ai-11",
"title": "Brute-Force-Schutz",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-12",
"title": "Rollenbasierte Autorisierung",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-13",
@@ -307,9 +328,16 @@
{
"control": "cra-ai-24",
"title": "Security-Logging",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-25",
@@ -328,30 +356,58 @@
{
"control": "cra-ai-27",
"title": "Log-Integritaet und -Aufbewahrung",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-28",
"title": "Sichere Update-Mechanismen",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-29",
"title": "Update-Authentizitaet",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-30",
"title": "Update-Integritaet",
"scans": [],
"note": "absence-based — no syntactic pattern; covered by the grounded surface check (retrieve surface + LLM judge), gated (BREAKPILOT_GROUNDED_CHECKS) pending live tuning",
"status": "needs_tooling"
"scans": [
{
"tool": "grounded-control-check",
"scan_type": "code_review",
"cwe": [],
"rules": []
}
],
"note": "covered by the grounded surface check (retrieve code surface + grounded LLM judge decides presence/absence); validated live",
"status": "covered"
},
{
"control": "cra-ai-31",
+12 -8
View File
@@ -178,11 +178,13 @@ mod tests {
}
#[test]
fn every_bucket_is_represented() {
fn covered_and_not_checkable_are_populated() {
let s = ControlMap::cra().unwrap().summary();
assert!(s.covered > 0);
assert!(s.needs_tooling > 0);
assert!(s.not_code_checkable > 0);
// needs_tooling is now empty: every code-checkable control is either
// tool-covered or covered by the grounded surface pass.
assert_eq!(s.needs_tooling, 0);
}
#[test]
@@ -215,14 +217,16 @@ mod tests {
}
#[test]
fn coverage_reflects_the_b_track_split() {
fn coverage_after_grounded_promotion() {
let s = ControlMap::cra().unwrap().summary();
// 9 already tool-covered + B1's 4 custom-semgrep controls.
assert_eq!(s.covered, 13);
// The 8 grounded surface controls stay needs_tooling until live-tuned.
assert_eq!(s.needs_tooling, 8);
// B3 marked the 4 pure-architectural controls not code-checkable.
// 9 off-the-shelf + 4 custom-semgrep + 8 grounded surface controls (promoted
// after the grounded path was validated live).
assert_eq!(s.covered, 21);
// Nothing left as needs_tooling — every code-checkable control is covered.
assert_eq!(s.needs_tooling, 0);
// The 4 pure-architectural controls remain not code-checkable.
assert_eq!(s.not_code_checkable, 19);
assert_eq!(s.total(), 40);
}
#[test]
+1
View File
@@ -36,6 +36,7 @@ export default withMermaid(defineConfig({
{ text: 'Pentest Architecture', link: '/features/pentest-architecture' },
{ text: 'AI Chat', link: '/features/ai-chat' },
{ text: 'Code Knowledge Graph', link: '/features/graph' },
{ text: 'Compliance Control Mapping', link: '/features/control-mapping' },
{ text: 'MCP Integration', link: '/features/mcp-server' },
],
},
+136
View File
@@ -0,0 +1,136 @@
# Compliance Control Mapping
Control mapping connects the scanner's raw output — deterministic tool findings and the code itself — to the **compliance controls** each piece of evidence supports. A hardcoded credential stops being just "CWE-798 from semgrep" and becomes evidence for *"cra-ai-8: no default passwords"* and, at scale, master control *`mc-31761` hardcoded_secrets_detection*. Findings carry those references (`control_refs`) into the dashboard and out over the MCP server as OSCAL, so the compliance report is built from real, grounded findings rather than a questionnaire.
## The core principle: tools detect, the LLM judges
The design has one rule, borrowed from the ZeroFalse / IRIS line of research: **deterministic tools are the detectors; the LLM is only ever a grounded false-positive filter, never the thing that finds the issue.**
- A tool (semgrep, gitleaks, syft/osv, the PLC linter; the DAST agents today, with **Nuclei and ZAP planned** as deterministic web/OT detectors underneath them — see [Tools & Scanners](/reference/tools#planned-integrations-decided-2026-08-31-not-yet-in-the-code)) detects.
- An **authored, human-reviewed lookup table** (`control-map`) maps that detection to the control(s) it's evidence for.
- The LLM enters last, to *confirm or refute* the mapping against the actual code — and every surviving verdict is anchored to a verbatim snippet by the grounding gate.
This keeps hallucination out of detection. The LLM supplies cross-language, cross-stack pattern *recognition*; the surrounding machinery supplies determinism.
## Coverage model
Every control lands in one of three buckets, recorded in the `control-map` LUT (`control-map/data/cra_control_map.json`) and never decided by an LLM:
| Bucket | Meaning |
| --- | --- |
| `covered` | An existing tool's scan surfaces findings for this control |
| `needs_tooling` | Code-checkable, but no off-the-shelf tool digs it out — we author a detector or use the grounded surface check |
| `not_code_checkable` | A design/process property — out of static-scan scope |
For the **CRA** framework (40 controls) the split is **13 covered · 8 needs_tooling · 19 not_code_checkable**. The 16 originally-uncovered controls were resolved as a hybrid:
- **4 custom semgrep detectors** (`cra-ai-1`, `7`, `10`, `14`) — secure-by-default, weak password hashing, insecure session cookies, weak data-at-rest ciphers. Shipped in the binary and matched back to controls **by rule id** so a broad CWE can't over-attribute.
- **8 grounded surface checks** (`cra-ai-6`, `11`, `12`, `24`, `27`, `28`, `29`, `30`) — the absence-based controls (no rate limiting, no security logging, no update-signature check…) that have no syntactic pattern.
- **4 marked not_code_checkable** (`cra-ai-2`, `3`, `4`, `5`) — minimal attack surface, secure architecture, least privilege, tamper protection.
At scale, the **master-controls** corpus (breakpilot's deduped clusters, exported as OSCAL) currently provides **~2,882 code-checkable controls** (2,143 `network` + 739 `source_code`), matched semantically.
## The three mapping paths
```mermaid
flowchart TD
T[Deterministic tools\nsemgrep · gitleaks · syft/osv · DAST · PLC linter] --> F[Findings]
F --> B["Stage 5b — LUT triage\ncontrols_for(tool, cwe / rule_id)"]
F --> C["Stage 5c — Semantic\nembed region+intent → top-K master controls"]
R[Repo source] --> D["Stage 5d — Grounded surface\nretrieve surface for absence-based controls"]
B --> J{{Grounded LLM judge\ntemp 0 · verbatim snippet}}
C --> J
D --> J
J -->|snippet grounds in region| S[Stamp control_refs]
J -->|refuted / ungrounded| X[Dropped]
```
All three paths converge on the same **grounded judge** and the same **grounding gate**. They differ only in how candidate (finding/region, control) pairs are produced.
### Stage 5b — deterministic LUT triage
The default path. A tool finding is matched to controls via `control_map.controls_for_finding(tool, cwe, rule_id)`; the judge then confirms each mapped control against the code region. Outcomes: `Confirmed([ids])` (stamp them), `FalsePositive` (drop the finding), or `Unmapped` (keep it untagged). Runs whenever `BREAKPILOT_BASE_URL` is set.
### Stage 5c — semantic retrieval (master-controls scale)
Master controls carry no CWE, so they can't be LUT-mapped. Instead we map by *similarity*: embed every control's requirement text once (cached), then for each finding retrieve the top-K nearest controls and hand them to the judge. Gated behind `BREAKPILOT_SEMANTIC_MAPPING` (default off). See [Semantic retrieval](#semantic-retrieval-in-detail).
### Stage 5d — grounded surface checks (absence-based controls)
Some controls are violated by an *absence* — no rate limiting on login, no security logging, no signature check on an update. There's no pattern for semgrep to match, so we deterministically retrieve the code **surface** the control governs (a login route, a logging setup, update/download code) by identifier/route terms, and let the judge decide whether the control holds there. Produces net-new, already-grounded findings. Gated behind `BREAKPILOT_GROUNDED_CHECKS` (default off).
## The grounding gate
No matter the path, a verdict becomes a finding only if it survives `compliance_core::control_check::ground`:
1. The judge runs at **temperature 0** with a closed prompt and must quote the offending code **verbatim** into `snippet`.
2. That snippet must appear **literally** in the retrieved region — otherwise the verdict is dropped.
3. The finding's line is **recomputed from the match**; the model's own line number is never trusted.
4. Verdicts are cached by content hash, so re-scans reproduce.
The model is allowed to be smart; it is never trusted.
## Semantic retrieval in detail
1. **Embed the corpus once.** Each control's requirement text is embedded with `bge-multilingual-gemma2` (3584-dim — multilingual matters, the master controls are in German while code is English). The embedding backend caps input arrays at 25 per request, so `embed()` chunks at 16; the whole `ControlIndex` is persisted to `snapshot_dir` keyed by a **corpus hash**, so only the first scan after a catalog change pays the embedding cost.
2. **Build the query from the finding's intent, not just the code.** The retrieval query is `finding.title + finding.description + region`, not the raw region. This is the single most important tuning: two findings in one file share overlapping windows and, on the code alone, embed alike and collapse onto the same controls. The finding's own words ("brute-force protection" vs "weak hash") carry the discriminating signal. The raw region still goes to the judge for grounding.
3. **Retrieve → judge → ground.** Top-K nearest by cosine, each judged against the region, each grounded.
## Worked examples
Both examples are from the live end-to-end verification (`c5_semantic_live.rs`) against the real ~2,882-control corpus.
### Example 1 — a small auth file (the tuning story)
Two findings in one `auth.py`: a weak `hashlib.md5(password)` hash and a login endpoint with no brute-force protection.
| Finding | Region-only retrieval | Intent-enriched retrieval |
| --- | --- | --- |
| Weak md5 hash | 19874, 20683, 23149, 29985 | **`mc-23149`** (eliminate weak unsalted hashes) at rank 1, + `mc-21634` salted hashing |
| Login w/o brute-force protection | *identical 4, reordered* | newly surfaces **`mc-19984`** brute_force_protection + **`mc-23186`** account_lockout |
Region-only retrieval gave both findings the *same* four password-hashing controls — the brute-force finding never found its real controls because its window is saturated with `password` tokens. Enriching the query with the finding's intent fixed it: the brute-force finding now pulls the correct rate-limiting / lockout controls out of the 2,882.
### Example 2 — four topically distinct vulnerabilities
| Finding | Top matched controls | Family |
| --- | --- | --- |
| SQL injection (string-concat query) | `sql_injection_prevention`, `sql_injection`, `parameterized_queries`, input_sanitization | input-validation ✓ |
| Hardcoded API credential | `hardcoded_secrets_detection`, credential_scanning, secrets_detection | credentials ✓ |
| TLS verification disabled (`verify=False`) | `https_enforcement`, `configuration_verification`, transport config | transport-encryption ✓ |
| Insecure deserialization (`pickle.loads`) | `deserialization`, `deserialization_testing`, `deserialization_security` | deserialization ✓ |
Every finding maps to its exact control family, with the most specific control often at the top, and the four sets are distinct.
## Known limitations
- **Absence findings are weak for semantic retrieval.** Similarity matches what code *is about*, not what it *lacks*; a "missing rate limiting" finding embeds like login code. This is exactly why the grounded surface path (Stage 5d) exists — it decides presence/absence at a retrieved surface rather than by embedding distance.
- **Generic catch-all controls co-occur.** `mc-20890 secure_development_security_code_review` appears in the top-K for many code-security findings because it is semantically near almost all of them. It's harmless (the judge grounds it, and it never crowds out the specific controls — the SQLi example didn't get it) but is a candidate for future down-weighting.
- **Corpus classification noise.** The master-controls `verification_method` classification is imperfect — e.g. a documentation control (`eu_declaration_accuracy`) is currently tagged `source_code`. That's a corpus-side data-quality issue, separate from the mapping engine.
## Emitting over MCP — closing the loop
Findings don't just land in the dashboard; they flow to breakpilot-compliance as OSCAL over the scanner's MCP server, so the compliance report is assembled from real, control-tagged findings.
- The MCP server exposes an **`oscal_assessment`** tool: given a `repo_id`, it emits a standard OSCAL 1.1 assessment-results document for that repo's findings — mapped findings target their controls via the stamped `control_refs`, and unmapped findings are reported **as-is** (as observations), so nothing is lost.
- breakpilot pulls it: `POST /v1/cra/oscal-from-scanner` calls `oscal_assessment` over MCP (Streamable HTTP + bearer) and consumes the pre-computed OSCAL — rather than pulling raw findings and re-assessing.
**Operational note — tenant context over HTTP.** The MCP server is multi-tenant; the bearer token resolves a tenant whose per-tenant database the tools query. rmcp's Streamable HTTP transport runs each session's tool calls in a `tokio::spawn`ed task, and `task_local`s do **not** cross a spawn — so binding the tenant in a per-request middleware `task_local` leaves tool handlers with no context (every call fails `no tenant context`). The fix is to bind the tenant to the **per-session server instance** at creation (the factory runs in the request scope before the spawn), not to a per-request task_local. Until this was fixed, the loop silently failed over HTTP and consumers fell back to demo data.
## Configuration
| Variable | Effect |
| --- | --- |
| `BREAKPILOT_BASE_URL` | breakpilot-compliance root; enables control ingest + all mapping passes. **Unset disables all control mapping** — findings are produced without `control_refs`. |
| `BREAKPILOT_SEMANTIC_MAPPING` | Stage 5c (semantic master-controls mapping). **Default on** (validated live). |
| `BREAKPILOT_GROUNDED_CHECKS` | Stage 5d (grounded surface checks). **Default on** (validated live). |
| `BREAKPILOT_SNAPSHOT_DIR` | Where OSCAL catalog snapshots and the cached control-embedding index live. |
The semantic and grounded passes default **on** now that both are validated live; each is still a no-op if `BREAKPILOT_BASE_URL` is unset or the catalog is unreachable, so they only ever add coverage. The live verifications live in `compliance-agent/tests/c5_semantic_live.rs` and `grounded_surface_live.rs` (ignored; run with `--ignored`).
## Appendix — the master-controls data pipeline
The master-controls corpus is produced by breakpilot-compliance and pulled as an OSCAL catalog from `GET /api/compliance/v1/oscal/catalog?framework=master-controls`. Two operational lessons are worth recording, because they cost real time to diagnose:
- **The catalog is served from `breakpilot_db`, not `postgres`.** Diagnostics run against the wrong database will look clean while the app serves something else entirely. Confirm the app's datname (`pg_stat_activity`) before trusting any count or `EXPLAIN`.
- **A constraint-less dump triplicated the master-control tables.** Restored without their PK/unique constraints, `master_controls` / `mc_verification` / `master_control_members` accumulated identical rows 3× (the same artifact migration `158` fixed for `doc_check_controls`). That inflated the catalog to ~26k dup'd controls and, with the indexes also missing, drove the export query to a >120s / 502. The fix (breakpilot migration `160`) ctid-dedups each table by its natural key and restores the constraints + indexes so it can't recur; the export query was also rewritten set-based (a single windowed pass instead of a per-row correlated subquery). After dedup: 41,850 → 13,950 master controls, catalog **25,938 → 2,882** code-checkable, endpoint **502 → 200 in ~3s**.
+4
View File
@@ -92,3 +92,7 @@ Filters can be combined. A count indicator shows how many findings match the cur
::: tip
Findings marked as **Confirmed** exploitable were verified with a successful attack payload. **Unconfirmed** findings show suspicious behavior that may indicate a vulnerability but could not be fully exploited.
:::
## Deterministic detectors (planned)
The DAST engine above is agentic: an LLM drives crawler, browser and testing tools and decides what to try next. That gives depth and code-aware exploitation, but not run-to-run reproducibility. The next step (decided 2026-08-31, not yet implemented) adds two deterministic open-source detectors **under** the agents: **Nuclei** (template checks incl. ICS/OT and default-credential templates) first, then an **OWASP ZAP** baseline scan. Their findings will appear alongside agent findings, carry CWE + compliance `control_refs`, and seed the agent's context so it verifies and chains instead of rediscovering. See [Tools & Scanners](/reference/tools#planned-integrations-decided-2026-08-31-not-yet-in-the-code).
+2 -2
View File
@@ -26,7 +26,7 @@ Filters can be combined. Results are paginated with 20 findings per page.
| Severity | Color-coded badge: Critical (red), High (orange), Medium (yellow), Low (green), Info (blue) |
| Title | Short description of the vulnerability (clickable) |
| Type | SAST, SBOM, CVE, GDPR, OAuth, Secrets, or Code Review |
| Scanner | Tool that found the issue (e.g. Semgrep, Grype) |
| Scanner | Tool that found the issue (e.g. Semgrep, Syft/OSV) |
| File | Source file path where the issue was found |
| Status | Current triage status |
@@ -73,7 +73,7 @@ If the finding has been pushed to an issue tracker (GitHub, GitLab, Gitea, Jira)
| Type | Source | Description |
|------|--------|-------------|
| **SAST** | Semgrep | Code-level vulnerabilities found through static analysis |
| **SBOM** | Syft + Grype | Vulnerable dependencies identified in your software bill of materials |
| **SBOM** | Syft + OSV.dev/NVD | Vulnerable dependencies identified in your software bill of materials |
| **CVE** | NVD | Known CVEs matching your dependency versions |
| **GDPR** | Custom rules | Personal data handling and consent issues |
| **OAuth** | Custom rules | OAuth/OIDC misconfigurations and insecure token handling |
+1 -1
View File
@@ -6,7 +6,7 @@ The SBOM (Software Bill of Materials) feature provides a complete inventory of a
A Software Bill of Materials is a list of every component (library, package, framework) that your software depends on, along with version numbers, licenses, and known vulnerabilities. SBOMs are increasingly required for compliance audits, customer security questionnaires, and supply chain transparency.
Certifai generates SBOMs automatically during each scan using Syft for dependency extraction and Grype for vulnerability matching.
Certifai generates SBOMs automatically during each scan using Syft for dependency extraction and OSV.dev + NVD for vulnerability matching.
## Packages Tab
+2 -2
View File
@@ -8,7 +8,7 @@ When a scan is triggered, Certifai runs through these phases in order:
1. **Clone** -- pulls the latest code from the Git remote (or clones it for the first time)
2. **SAST** -- runs static analysis using Semgrep with rules covering OWASP, GDPR, OAuth, secrets, and general security patterns
3. **SBOM** -- extracts all dependencies using Syft, identifying packages, versions, licenses, and known vulnerabilities via Grype
3. **SBOM** -- extracts all dependencies using Syft, identifying packages, versions, licenses, and known vulnerabilities via OSV.dev + NVD
4. **CVE Check** -- cross-references dependencies against the NVD database for known CVEs
5. **Graph Build** -- parses the codebase to construct a code knowledge graph of functions, classes, and their relationships
6. **AI Triage** -- new findings are reviewed by an LLM that assesses severity, considers blast radius using the code graph, and generates remediation guidance
@@ -52,7 +52,7 @@ A full scan runs multiple analysis engines, each producing different types of fi
| Scan Type | What It Detects | Scanner |
|-----------|----------------|---------|
| **SAST** | Code-level vulnerabilities (injection, XSS, insecure crypto, etc.) | Semgrep |
| **SBOM** | Dependency inventory, outdated packages, known vulnerabilities | Syft + Grype |
| **SBOM** | Dependency inventory, outdated packages, known vulnerabilities | Syft + OSV.dev/NVD |
| **CVE** | Known CVEs in dependencies cross-referenced against NVD | NVD API |
| **GDPR** | Personal data handling issues, consent violations | Custom rules |
| **OAuth** | OAuth/OIDC misconfigurations, insecure token handling | Custom rules |
+2 -2
View File
@@ -58,8 +58,8 @@ An open-source static analysis tool that finds bugs and enforces code standards
**Syft**
An open-source tool for generating SBOMs from container images and filesystems. Used by Certifai to extract dependency information.
**Grype**
An open-source vulnerability scanner for container images and filesystems. Used by Certifai to match dependencies against known vulnerabilities.
**OSV.dev**
Google's open distributed vulnerability database, queried by package URL. Certifai uses it (together with NVD) to match SBOM components against known vulnerabilities.
## Protocols
+16 -6
View File
@@ -24,15 +24,14 @@ Semgrep produces SAST-type findings with file paths, line numbers, and rule desc
Syft output feeds into both the SBOM feature and the vulnerability scanning pipeline.
## Grype -- Vulnerability Scanning
## OSV.dev + NVD -- Vulnerability Matching
[Grype](https://github.com/anchore/grype) is an open-source vulnerability scanner that matches your dependencies against known vulnerability databases. It takes Syft's SBOM output and cross-references it against:
Certifai matches every SBOM component directly against two public vulnerability sources (no separate scanner binary):
- National Vulnerability Database (NVD)
- GitHub Advisory Database
- OS-specific advisory databases
- [OSV.dev](https://osv.dev/) -- batch queried by package URL (purl) for ecosystem advisories (npm, PyPI, crates.io, Go, Maven, ...)
- [NVD](https://nvd.nist.gov/) -- queried per CVE for the CVSS v3.1 base score, and by CPE for CODESYS runtime versions found in PLC projects
Grype produces SBOM-type findings with CVE identifiers, severity ratings, and links to advisories.
Matches are stored as CVE alerts with CVSS scores and re-checked hourly, so newly published CVEs against an unchanged dependency still raise a notification.
## Custom OAuth Scanner
@@ -97,3 +96,14 @@ When you mark findings as false positives or provide developer feedback, this in
::: tip
The AI triage is a starting point, not a final verdict. Always review the rationale and code evidence before acting on a finding. See [Understanding Findings](/guide/findings#human-in-the-loop) for more on the human-in-the-loop workflow.
:::
## Planned integrations (decided 2026-08-31, not yet in the code)
The product spec keeps an **OSS-only** tooling policy and a control-mapping rule of *tools detect, the LLM judges*. Two deterministic detectors are therefore being added **underneath** the agentic DAST/pentest layer — the agents stay on top for context-seeded exploitation, chaining and explanation:
| Tool | Role | Status |
|------|------|--------|
| [Nuclei](https://github.com/projectdiscovery/nuclei) | Template-driven checks (CVE probes, default credentials, exposed panels, misconfigurations) including ICS/OT templates for WebVisu / OpenPLC / HMI endpoints. Runs as a DAST phase and as a Werkbank job with vendored templates so it works on-prem. | Planned — tracked as an issue |
| [OWASP ZAP](https://www.zaproxy.org/) | Baseline (passive) and, behind the destructive-tests flag, active scan for reproducible spider + rule coverage; results seed the pentest agent. | Planned — follows Nuclei |
Both feed the same `control-map` lookup table as Semgrep, so their findings receive compliance `control_refs` through the grounded judge. An **offline vulnerability database** (Trivy preferred, Grype as alternative) is planned for the on-prem Werkbank runner, which cannot reach the OSV.dev / NVD APIs. Until these land, DAST findings come exclusively from the in-house agents described above.