Rework tag_deploy/release_rpms: surface Forge errors, fix tag-triggered RPM builds (el8/el9/el10) - #44
Conversation
Mirrors #44 so the two branches merge cleanly in either order. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCnDsYaJDLiP8z8tafz9Tp
|
Added a second robustness improvement prompted by the gpasswd post-mortem: the built module archive is now uploaded as a workflow artifact before the Forge POST, so every tag run — pass or fail — leaves the exact tarball downloadable from the run page (previously a failed upload left nothing: no artifact, and the GitHub release carries no assets). Mirrored to #42 as before. |
|
Third improvement: the module archive is now also attached to the GitHub release ( |
hcaballero2
left a comment
There was a problem hiding this comment.
Skeptical review — I went looking for the usual failure modes in "attach assets + surface API errors" changes and verified the PR's claims against the actual branch contents rather than the description. What checks out:
- Ordering is sound.
deploy-to-puppet-forgehasneeds: [create-github-release], so the release exists beforegh release uploadruns — no race. And attaching before the Forge POST is the right order, not just for evidence-preservation: the inverse order plus a transient attach failure would force a job re-run whose Forge POST would then 409 on the already-published version. - The prerelease claim is accurate: the skip is the job-level
if: needs.create-github-release.outputs.prerelease != 'yes', so the new artifact/attach steps can't run for prereleases (where no non-draft release semantics would apply). - The token claim is accurate:
create-github-releasealready doesgh release createwith the defaultGITHUB_TOKEN, sogh release uploadneeds nothing the workflow doesn't already rely on. actions/upload-artifact@v7is the current major (v7.0.1), andif-no-files-found: errormeans a missing tarball now fails fast at the artifact step instead of producing a cryptic curl--form file=@error two steps later.- The "mirrored to #42" claim is true —
openvox9-ruby4-template-refreshcarries the identical three-step change (verified lines 188/196/202-215 of itstag_deploy.yml). - Error handling holds up at the transport level too: GitHub
run:steps default tobash -e, so a non-HTTP curl failure (DNS, TLS) fails the step at thehttp_code=assignment with--show-erroroutput; thecaseonly needs to cover HTTP-level failures, which is exactly what the old--failwas eating. Failing on 3xx (which--faildid not) is a behavior change, but a correct one for a POST that should never redirect.
Two small suggestions inline; neither blocks.
curl --fail discards the response body, so a failed publish reports only an HTTP status (the simp-gpasswd 2.0.0 release died with a bare 403). Capture the body and status, print both, and fail on any non-2xx result so the Forge's own error message lands in the job log. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCnDsYaJDLiP8z8tafz9Tp
A failed Forge upload previously left nothing to download - the tarball existed only on the runner and the GitHub release carries no assets. Upload it before attempting the Forge POST so every tag run, pass or fail, leaves the exact archive available from the run page. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCnDsYaJDLiP8z8tafz9Tp
Release assets are permanent and publicly downloadable, unlike workflow artifacts (authenticated, expiring). Uploaded before the Forge POST so a failed publish still leaves the exact archive on the release. --clobber keeps job re-runs idempotent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XCnDsYaJDLiP8z8tafz9Tp
The deployed workflows have moved on (PUPPET_VERSION '~> 8', Ruby 3.4.9, checkout@v7, github-script@v9, rake pupmod:build instead of pdk build) while the template still described the Puppet-7 era. Reconciling before the sync keeps the rollout diff down to the intended changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Every tag push has been dispatching two release_rpms.yml runs hardcoded to the retired centos7/centos8 containers, overriding #80's el8 default — all tag-triggered RPM builds fail. The caller templates now dispatch once with no OS override, and release_rpms.yml owns the build-OS list via a build_container_oses input (default '["el8","el9","el10"]', narrowable on manual runs). Per-release work (release lookup/creation, the clean-input asset wipe) is split into a resolve-release job so the parallel per-OS legs cannot race on release autocreation or wipe each other's uploads, and both build inputs are required: false so omitting them in dispatch calls is well-defined. Also fixes validate-inputs writing '{name}={value}' to GITHUB_OUTPUT — invalid syntax, so prebuild_suffix/build_semver never populated and prerelease tags were treated as full releases. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
EL7 is gone fleet-wide and the per-OS choice now lives in release_rpms.yml, so tag_deploy_github-rpms-el7-el8.yml had no remaining purpose. pkg-r10k and simp-adapter fall back to the standard simp_unknown presets, whose absent list already removes the file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
f713fb1 to
346cdc3
Compare
Closes #84.
This started as the Forge-deploy robustness work (first three commits) and grew into the full tag-and-release rework, per our convention of combining the template change and the workflow-run session config in one PR so they're tested together:
curl --faildiscarded the response body, so a failed Forge publish reported only a bare HTTP status (the simp-gpasswd 2.0.0 release died with an unexplained 403). The deploy step now prints the response body and status and fails on any non-2xx result.tag_deploy.ymlwith the deployed fleet — the template still described the Puppet-7/PDK era (PUPPET_VERSION '~> 7', Ruby 2.7, checkout@v5,pdk build); deployed copies have long since moved to'~> 8', Ruby 3.4.9, checkout@v7, andrake pupmod:build. Reconciling first keeps the rollout diff down to the intended changes (verified: the merged diff on a real clone touches nothing else).release_rpms.ymlruns hardcoded to the retiredcentos7/centos8containers, overriding Pull RPM build containers from ghcr.io/simp/simp-<os>-build (fixes fleet-wide RPM build failures) #80'sel8default; all 14 tag-triggered RPM builds since 2026-07-28 failed. All three caller templates now dispatch once with no OS override, andrelease_rpms.ymlowns the build-OS list via abuild_container_osesinput (default'["el8","el9","el10"]', narrowable on manual runs). Per-release work is split into aresolve-releasejob so the parallel per-OS legs can't race on release autocreation or thecleanasset wipe; both build inputs are nowrequired: false. Output changes from two RPMs per tag to three (el8/el9/el10) — intended.validate-inputswrote{name}={value}toGITHUB_OUTPUT(invalid syntax), soprebuild_suffix/build_semvernever populated: prerelease tags were treated as full releases and the RPM release-tag customization never ran. Nowname=value.release_rpms.yml, it had no remaining purpose; pkg-r10k and simp-adapter fall back to the standardsimp_unknownpresets, whose absent list removes the deployed file on their next apply sync.20260812-tag-deploy-rpm-matrix.yaml(nowlatest) — mergestag_deploy.yml+release_rpms.ymlfleet-wide (scoped, Renovate-managed scalars preserved as always).Verified: 210 rspec + 19 BoltSpec plan examples green; all plans parse and load; YAML validates; local e2e against a real pupmod-simp-aide clone produces exactly the intended diff with all deployed Renovate values (checkout@v7, github-script@v9, ubuntu-24.04) preserved.
Note: the
simp-*asset repos (simp-doc, simp-utils, …) carry the same caller templates but aren't in the dynamic inventory or permitted project types — they pick up the fixed templates on their nextapply_puppet_rolesync, and the session config notes this. Once merged and rolled out, the failed releases in #84's evidence table can be rebuilt by re-dispatchingrelease_rpms.yml— no re-tagging needed.🤖 Generated with Claude Code