Skip to content

[aw-failures] [aw-fix] P1: Threat-detection conclude step runs threat-detect when detection is skipped → exit 127 false-failur [Content truncated due to length] #47016

Description

@github-actions

Fix conclude_threat_detection.sh — guard line 41 so threat-detect conclude never runs when RUN_DETECTION=false; it exits 127 and marks otherwise-successful runs as failed.

This is a deterministic gh-aw product bug in the generated detection job, not an inference/model/firewall issue. It produces false CI failures whenever threat detection is skipped.

Problem statement

When a workflow's detection job short-circuits (RUN_DETECTION=false, DETECTION_AGENTIC_EXECUTION_OUTCOME=skipped), the conclude step still invokes the external threat-detect binary. That binary is not on PATH when detection is skipped, so the step dies with command not found / exit code 127 and the whole run is reported as failed — even though the agent completed all its work successfully.

Affected workflows and run IDs

Workflow Run Engine Symptom
Daily Sub-Agent Model Resolution Audit §29806493363 Copilot (gpt-5.4-mini) detection job exit 127; agent exit 1

The agent job itself ran to completion (11m42s, 11.3M input tokens) and produced full analysis output. Only the detection-conclude step and the downstream failure propagation broke the run.

Root cause

actions/setup/sh/conclude_threat_detection.sh:41 calls the binary unconditionally:

# line 13 guard only covers: RUN_DETECTION=true AND result file missing
if [ "${RUN_DETECTION:-false}" = "true" ] && [ ! -f "${RESULT_FILE}" ]; then
  ... exit early ...
fi

threat-detect conclude --result-file "${RESULT_FILE}"   # line 41 — runs even when RUN_DETECTION=false

With set -euo pipefail (line 4) and RUN_DETECTION=false, execution falls straight through the line-13 block to line 41. threat-detect is not installed on the skipped-detection path, so:

/home/runner/work/_temp/gh-aw/actions/conclude_threat_detection.sh: line 41: threat-detect: command not found
##[error]Process completed with exit code 127.

The comment at pkg/workflow/threat_detection_job.go:137 states that when run_detection=false the conclusion "short-circuits with conclusion=skipped" — the shell script does not implement that short-circuit. This surfaced when the external threat-detect binary path (feature-flagged, see pkg/constants/feature_constants.go:89) was enabled. Existing tests (pkg/workflow/threat_detection_conclude_script_test.go) only cover RUN_DETECTION=true, so the skipped path was never exercised.

Supporting evidence (audit + audit-diff)
  • audit run 29806493363: errors=[{file:"detection", exit code 127}, {file:"agent", exit code 1}]; engine copilot/gpt-5.4-mini; no model_not_supported/http_400 conditions actually triggered.
  • audit-diff (base 29229167122 → 29806493363): new_domain_count=0, status_change_count=0, has_anomalies=false. No firewall/MCP/model drift — confirms the failure is the detection script, not inference.
  • Detection log tail: RUN_DETECTION: false, DETECTION_AGENTIC_EXECUTION_OUTCOME: skipped, then the command not found line above.

Proposed remediation

Guard the conclude invocation so it only runs when detection actually executed. Minimal fix:

if [ "${RUN_DETECTION:-false}" != "true" ]; then
  echo "conclusion=skipped" >> "${GITHUB_OUTPUT}"
  echo "GH_AW_DETECTION_CONCLUSION=skipped" >> "${GITHUB_ENV}"
  exit 0
fi

threat-detect conclude --result-file "${RESULT_FILE}"

This aligns the script with the run_detection=false → conclusion=skipped contract already documented in threat_detection_job.go. Recompile all .lock.yml files after the change.

Success criteria / verification

  • New test in threat_detection_conclude_script_test.go with RUN_DETECTION=false and no threat-detect on PATH → script exits 0 and emits conclusion=skipped.
  • Re-run "Daily Sub-Agent Model Resolution Audit" (or any workflow with detection skipped) → detection job succeeds; run is not marked failed.
  • Existing RUN_DETECTION=true tests still pass; threat-detect is still invoked on the real detection path.

Existing-issue correlation

Not covered by any open agentic-workflows issue. Distinct from the Copilot model-resolution family (#46301, #46326, #46647, #46846) — those are model/pricing/SDK errors; this is a detection-script control-flow bug. No existing issues were closed because none showed fresh evidence of being fixed.

Also observed this window (P2 — monitor, no fix issue yet)
  • Daily VulnHunter Scan §29807147048 (Claude Opus): agent step failed while actively reading source files — log tail shows normal activity with no error marker, consistent with abrupt termination / engine timeout mid-scan. Single occurrence; watch for recurrence before opening a fix issue.
  • All other window "failures" were cancellations, not defects: 3× Tidy (concurrency) + ~90 PR-triggered runs cancelled together at 03:32. No action.

References:

Generated by 🔍 [aw] Failure Investigator (6h) · 253 AIC · ⌖ 33.9 AIC · ⊞ 5.1K ·

  • expires on Jul 27, 2026, 11:50 PM UTC-08:00

Metadata

Metadata

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions