* Run workflows that call Claude on the egress-firewall runner in auto permission mode
Move every workflow job that signs in to the Claude API from
ubuntu-latest to GitHub's egress-firewall runner (ubuntu-24.04-firewall),
and pass --permission-mode auto to the Claude Code action in each of
them except claude.yml.
In auto permission mode Claude Code reviews each tool call that needs
permission and that the allowed tools do not cover, and runs it only if
Claude Code's safety review passes it. Until now these runs, which have nobody to
ask, refused every such call. The issue triage step, which anyone can
start by opening an issue, also passes a --disallowedTools list for
tools it never needs.
claude.yml answers @claude mentions, and for those the action sets
--permission-mode acceptEdits itself. Its permission mode is unchanged.
Allowed tools, settings, models, triggers and permissions are unchanged.
The network allow list for the firewall follows in the next change.
* Add the network allow list, a security check and CLAUDE.md guidance
- .github/egress-firewall.yaml: the hosts that jobs on the
egress-firewall runner may reach, in enforce mode, each with what
uses it.
- .github/workflows/workflow-hardening.yml and
.github/scripts/check_workflow_hardening.py: a check that fails when a
job that calls Claude is not on the egress-firewall runner, does not
pass --permission-mode auto, or when the allow list is missing, empty, not in
enforce mode or names a host with '*'.
claude.yml is listed as exempt from the permission mode rule, with
the reason.
- CLAUDE.md: a "Security hardening for GitHub Actions" section so that
new and edited workflows keep these protections.
* Pin the model for issue triage
The step named no model, so it ran on Claude Code's default model. That
default changed with Claude Code 2.1.280 to a model that is not available
to this CI account. Pass --model claude-opus-5, the model this repository's test-*.yml
workflows (ANTHROPIC_MODEL) and the Python SDK's example runs are
pinned to for the same reason.
This is independent of the security hardening in the two changes before
it.
* Check settings files and --settings for a permission mode
The security check looked for defaultMode only in an inline 'settings'
input. Also read the file a 'settings' input or a --settings flag in
claude_args names, when it is inside the repository, and fail if it sets
a permission mode.
* Flag defaultMode anywhere in a job that calls Claude
A settings file that an earlier step writes at run time can't be read by
the step-level check, so also fail when the job's definition mentions
defaultMode at all.
* Fail the check when settings name a file outside the repository
The check can't read such a file, so it can't tell whether it sets a
permission mode.
* Run @claude mentions in auto permission mode
The action sets --permission-mode acceptEdits for @claude mentions and
appends the workflow's claude_args after it, so the --permission-mode auto
in claude.yml wins. Drop claude.yml from EXEMPT_FROM_AUTO_MODE.
The check now also fails when a step in auto mode names a model older than
claude-opus-4-6: Claude Code does not run auto mode on those and falls back
to its default permission mode.
---------
Co-authored-by: Claude <noreply@anthropic.com>
The test workflows name no model, so they run on the Claude Code CLI's
default model. Since the bump to Claude Code 2.1.280 that default is Opus 5.5,
which the CI account's API access refuses, so the tests fail and the release
job never runs. Setting ANTHROPIC_MODEL at the workflow level keeps the tests
independent of the CLI's default model, the same fix the Agent SDK for Python
used.
Co-authored-by: Claude Code <noreply@anthropic.com>
The download_job_log MCP tool called
client.actions.downloadJobLogsForWorkflowRun() with no timeout and no
AbortController. @octokit/rest@21 runs on Node's native fetch, which has
no default timeout, and Octokit only cancels a request when the caller
passes request.signal. If the log blob fetch stalls, that await never
resolves and never rejects.
This tool is always enabled in tag mode (src/modes/tag/index.ts), so a
"fix the failing CI" run that calls get_ci_status -> get_workflow_run_details
-> download_job_log can hang on this one await with nothing to recover it.
It's headless, so the run only ends when the Actions job-level
timeout-minutes kills it, burning the whole job budget with the tracking
comment stuck at "Claude Code is working...".
Sibling fetch in src/github/utils/image-downloader.ts (fetchImage) already
got this treatment in #1625 via a timeout-driven AbortController. Same
shape of call: fetch a GitHub-hosted resource by ID from untrusted PR/CI
content. This mirrors that fix for github-actions-server.ts.
Extracted the download+write logic into an exported downloadJobLog()
function (with an injectable timeoutMs) so the timeout path is directly
testable, and guarded the module's entrypoint side effects with
import.meta.main, matching the pattern already used by the other
entrypoints in src/entrypoints/.
* fix: use paths param in delete_files prompt example
The tag-mode prompt told the model to call delete_files with
"files", but the MCP tool schema and handler expect "paths".
Co-authored-by: Cursor <cursoragent@cursor.com>
* test: prove delete_files prompt against the live MCP schema
The old {files} payload is rejected by the same Zod shape the
tool registers; the generated prompt example now parses cleanly.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test: cover delete_files schema edges and the sibling commit_files tool
Confirm the old files-only payload still fails, types and required
fields are enforced, and commit_files was not inverted by the fix.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: RESILIENCE Agentic Solutions <286555414+WeAreResilience@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
On Windows, git's default core.autocrlf=true checks out CRLF, and the contributor commands in CONTRIBUTING.md then fail: bun run format:check reports 191 files because Prettier defaults to endOfLine "lf", and three tests that match LF-terminated content fail.
All tracked blobs are already LF, so this changes checkout behavior only and produces no renormalization diff. Every CI job runs on ubuntu-latest, where it is a no-op.