Diagnose CI/CD pipeline failures, analyze Jenkins build logs, and troubleshoot deployment issues...
Diagnoses CI/CD pipeline issues for the Elohim project using Jenkins. Covers two access paths β Jenkins MCP (preferred, primary) and direct WebFetch against public URLs (fallback for contexts where MCP isn't loaded) β plus retrieval workflows for sprint-reports, cucumber reports, and build-to-commit correlation.
Jenkins is OIDC-protected, with two access paths and sharply different guardrails:
1. Anonymous (default β MCP and unauthenticated WebFetch both run this way). No Authorization header β explicit auth triggers an interactive OIDC login flow and breaks the connection. Consequences:
getBuild, getBuildLog, searchBuildLog, getJob, getJobs, getBuildChangeSets, getTestResults, getFlakyFailures, getStatus, getBuildScm, getJobScm β anonymous has Overall.Read + Job.Read on https://jenkins.ethosengine.com. Same for public WebFetch URLs (address book below) β no auth needed for reads.mcp__jenkins__triggerBuild and mcp__jenkins__updateBuild don't work β anonymous lacks Job.Build. For retriggers, push a commit (optionally [build:<pipeline>] tagged β see "Triggering builds" below).2. Authenticated (JENKINS_USERNAME + JENKINS_TOKEN, orchestrator-only). A direct curl -u "$JENKINS_USERNAME:$JENKINS_TOKEN" call can do what anonymous can't, including parameterized builds (e.g. RESET_STORAGE=true for schema-drift recovery, where [build:*] tags are insufficient because they carry pipeline membership but not parameter values). Reserved for the shift Opus orchestrator, usable autonomously (no per-use user confirmation) but only against verified Jenkins state β the orchestrator must KNOW from actual reads that triggering won't cause interruptions or storms. Guessing is disqualifying. Full guardrails and workflow: "Parameterized rebuild (authenticated)" below.
MCP vs WebFetch: prefer MCP (mcp__jenkins__*) when loaded β it's structured (log search, test-result enumeration, changeset queries) and always anonymous-mode. Fall back to WebFetch against the URL patterns below when MCP isn't in your tool list (subagent contexts that don't inherit the parent's MCP set, or a session before postStart finishes). Detect by checking your tool list for mcp__jenkins__*; if absent, go straight to WebFetch rather than spending a turn hoping it appears.
This is the workflow run 90% of the time. Don't grep the repo, don't guess URLs β hit these in order:
Confirm the commit is on origin
git log --oneline origin/dev..HEAD # empty? already pushed.
git log --oneline -1 origin/dev # note the SHA
Find the latest build of the affected pipeline β the job index page lists the last N builds with status:
WebFetch https://jenkins.ethosengine.com/job/elohim-genesis/job/dev/
prompt: "List last 3 builds with build number, status, timestamp.
Also list any archived artifacts under genesis/a2o/reports/."
Returns build numbers like #935, status, timestamp, and the artifact tree for the latest. The index page is your address book β always start here.
Confirm your SHA landed in that build:
WebFetch https://jenkins.ethosengine.com/job/<pipeline>/job/dev/<N>/api/json?tree=changeSets[items[commitId,msg]],actions[causes[shortDescription]]
prompt: "List commit SHAs in this build's changeset, and the build cause."
changeSets[].items[].commitId: build reflects your fix β if findings haven't changed, the fix itself is wrong.actions[].causes[] shows what triggered it (upstream pipeline, scheduled run, your push, or a manual triggerBuild).actions tells you which commit it actually ran against.To check whether several commits all landed, diff git log --oneline origin/dev~10..origin/dev against the build's changeSets list β anything above the last listed commitId lands in the next build.
Retrieve the sprint-report (A2O pipeline only) for the findings delta:
WebFetch https://jenkins.ethosengine.com/job/elohim-genesis/job/dev/<N>/artifact/genesis/a2o/reports/sprint-report.md
prompt: "Return verbatim. Preserve summary table, pillar headers, all fingerprints."
Compare fingerprints against the previous build's report β fingerprints are stable across runs (the same hash repeats for the same error signature across builds).
Narrate the delta, not the absolute numbers. "Fingerprint X dropped from 26 β 0" or "new fingerprint Y appeared" is the signal β not "99 failures is bad."
Never guess these. Copy from here. Base URL: https://jenkins.ethosengine.com. Branches: dev (main), feature branches become job/feat-xxxx, etc. β URL-encode slashes as /job/.
| Path | Template | Example |
|---|---|---|
| Job index (build list) | /job/<pipeline>/job/<branch>/ |
/job/elohim-genesis/job/dev/ |
| Specific build | /job/<pipeline>/job/<branch>/<N>/ |
/job/elohim-genesis/job/dev/935/ |
| Latest build (alias) | /job/<pipeline>/job/<branch>/lastBuild/ |
useful when you don't know N |
| Last successful | /job/<pipeline>/job/<branch>/lastSuccessfulBuild/ |
skips failed builds |
| Build artifact | /job/.../<N>/artifact/<repo-relative-path> |
β¦/935/artifact/genesis/a2o/reports/sprint-report.md |
| Console log (raw) | /job/.../<N>/consoleText |
plain text β huge, prefer searchBuildLog via MCP |
| API (JSON) | /job/.../<N>/api/json?tree=<fields> |
tree=changeSets[items[commitId]] for SHAs |
| Pipeline | Artifact | Full path |
|---|---|---|
elohim-genesis |
Sprint-report (markdown) | genesis/a2o/reports/sprint-report.md |
elohim-genesis |
Sprint-report (JSON) | genesis/a2o/reports/sprint-report.json |
elohim-genesis |
Cucumber HTML | genesis/a2o/reports/cucumber-report.html |
elohim-genesis |
Cucumber JSON | genesis/a2o/reports/cucumber-report.json |
elohim-genesis |
Coverage gap report | genesis/a2o/reports/coverage-gap-report.json |
elohim-genesis |
Per-scenario console errors | genesis/a2o/reports/console/<scenario>.json |
elohim-app |
Test results | check job's test trend; surface via mcp__jenkins__getTestResults |
Flow: elohim-orchestrator (webhook-triggered changeset analysis) dispatches elohim-holochain, elohim-edge, elohim-app in parallel, each feeding elohim-genesis (seed + a2o verification), then post-deploy health checks. elohim-steward is manual-only (Tauri desktop).
| Job name | Purpose | Produces | Typical duration |
|---|---|---|---|
elohim-orchestrator |
Changeset analysis, downstream trigger | dispatch decision | 3-5 min β supersedes an in-flight run if a new push lands first |
elohim-holochain |
DNA compilation, hApp packaging + k8s deploy | elohim.happ |
15-25 min β fast build, slower deploy (pod restart + readiness probes) |
elohim-edge |
Docker builds: doorway + elohim-storage (heavy cargo build --release) + deploy |
doorway:tag, storage:tag images |
25-35 min β the long pole; start here when estimating end-to-end time |
elohim-app |
Angular build, static assets | dist/elohim-app |
β |
elohim-genesis |
Content seeding + a2o cucumber run + sprint-report | seed verification, sprint-report | 10-15 min β dominated by E2E scenario timeouts; a broken alpha inflates this |
elohim-steward |
Tauri desktop app | β | manual trigger only |
| Full cascade (push β sprint-report) | ~45-60 min β budget a full hour from git push to reading a fresh sprint-report |
Despite the name, elohim-edge uses elohim/holochain/Jenkinsfile (not the DNA-only one) β it's what builds Rust images and deploys. The DNA-only job is elohim-holochain, using elohim/holochain/dna/Jenkinsfile. Authoritative nameβjenkinsPathβchangePatterns mapping: genesis/orchestrator/Jenkinsfile PIPELINES map.
Key Jenkinsfile locations: root orchestrator /projects/elohim/Jenkinsfile; pipeline controller /projects/elohim/genesis/orchestrator/Jenkinsfile; DNA/hApp builds /projects/elohim/holochain/Jenkinsfile; seeding pipeline /projects/elohim/genesis/Jenkinsfile; desktop app /projects/elohim/steward/Jenkinsfile.
| Branch pattern | Environment | Doorway URL |
|---|---|---|
dev, feat-*, claude-* |
Alpha | doorway-alpha.elohim.host |
staging-* |
Staging | doorway-staging.elohim.host |
main |
Production | doorway.elohim.host |
Use ScheduleWakeup.delaySeconds per what you're waiting for β don't poll inside a turn:
dev-latest points at last green). holochain-green + edge-not-yet-green = the deploy is running old code β needs either a fresh holochain trigger after edge, or a rollout restart to pick up the new tag.dev-latest tag race. StatefulSet pods only re-pull on pod restart. With imagePullPolicy: Always + dev-latest, k8s pulls the new image only when the pod restarts β a rolling restart happens on spec change (env var, mount, resource bump), NOT on image tag content change alone. A deploy relying on dev-latest moving must also change something in the pod spec, or include an explicit kubectl rollout restart statefulset/....changePatterns β it compares against each pipeline's own last-built baseline commit and skips if that baseline already contained the file in its current state. A manifest-only annotation change can look "new" to git but not to the pipeline's baseline. Escape hatch: the [build:*] tag (below).mcp__jenkins__triggerBuild is unavailable to the anonymous MCP user. The canonical trigger surface is the GitHub webhook landing on elohim-orchestrator, which analyzes the changeset and dispatches downstream pipelines.
Retry a failed build for the same commit β empty commit + push:
git commit --allow-empty -m "ci: retrigger [build:edge]"
git push
Force a pipeline whose changeset analysis missed β the orchestrator's webhook-trigger branch parses the HEAD commit message for [build:<pipeline>] tags and force-adds matching pipelines regardless of changeset analysis:
| Tag | Pipeline |
|---|---|
[build:edge] |
elohim-edge (Rust doorway + storage + deploy) |
[build:dna] |
elohim-holochain (DNA/hApp only) |
[build:app] |
elohim (Angular) |
[build:genesis] |
elohim-genesis (a2o + seed) |
[build:sophia] |
elohim-sophia |
[build:steward] |
elohim-steward |
[build:all] |
every non-manualOnly pipeline |
Comma-separated forms work: [build:edge,genesis]. Use this when the orchestrator's changeset analysis is verifiably wrong, you need a rolling restart the normal deploy wouldn't cause, or you're testing CI changes without meaningful source changes. Don't abuse [build:all] β each pipeline costs minutes; prefer the narrowest tag.
Trigger with parameters: the webhook path only uses pipeline-defined defaults β it can't vary SKIP_SEEDING, ENVIRONMENT, RESET_STORAGE, etc. For that, use the authenticated path below (editing Jenkinsfile defaults and pushing also works, but only for permanent changes, not one-off operational unblocks).
The only sanctioned way for Claude to trigger parameterized Jenkins builds. Credentials live in env (JENKINS_USERNAME, JENKINS_TOKEN, JENKINS_URL) and authenticate against Job.Build, which anonymous MCP lacks. Read both halves below before reaching for the curl.
Who may use this: only the shift Opus orchestrator β even within /shift, never subagents. ci-observer and ci-investigator are read-only instruments that feed Jenkins state to the orchestrator; they never invoke triggers.
When it's appropriate: the [build:*] tag mechanism forces pipeline membership but can't pass parameter values. Use the authenticated path when a parameterized stage gates the actual fix (canonical case: elohim-genesis Seed Database stage's RESET_STORAGE=true, which clears content.db to recover from schema drift (see feedback_seed_lock_means_schema_drift) β tag-based retrigger inherits the default false and reproduces the failure), when a diagnostic rebuild needs a non-default value (SKIP_SEEDING=true, narrowed STEPS=..., alternate ENVIRONMENT), or when the webhook path simply can't carry the operational intent. Don't reach for it when an empty [build:<pipeline>] commit would do.
Hard guardrails β structural preconditions, not procedural confirmations:
getJob/getStatus reads that triggering won't stomp a concurrent build or cause a storm. "Guessed," "should be fine," "MCP didn't respond but probably ok" β do not curl; re-check after a wait or bail with an explicit question.elohim-orchestrator, elohim-genesis, elohim-edge, elohim-holochain, elohim) via mcp__jenkins__getJob jobFullName="<pipeline>/dev" and confirm lastBuild.building: false. Any failed or ambiguous read means the precondition isn't met β don't proceed. Record the queue snapshot in the journal as evidence.elohim-orchestrator with parameters β it dispatches downstream pipelines, multiplying storm risk. Stick to leaf pipelines (elohim-genesis is canonical; elohim-edge is acceptable but rarely needs parameters).RESET_STORAGE=true runs kubectl exec rm content.db && kubectl delete pod for each alpha human (genesis/Jenkinsfile:285-334) β it stomps any concurrent reader/writer. Tighten guardrail #2 to its strictest reading: if anything alpha-touching is building: true, defer.$JENKINS_TOKEN (same for $JENKINS_USERNAME/$JENKINS_URL). If a curl needs quoting in the journal, use the env-var placeholders, never resolved values β capture only the resulting build number / queue id.curl -sS -u "$JENKINS_USERNAME:$JENKINS_TOKEN" "$JENKINS_URL/api/json?tree=mode" | head -c 200
Jenkins JSON back (not an HTML login page or 403) means credentials are good. Otherwise journal the failure mode and bail with a question β don't retry blindly. This is a sanity check, not a permission gate.The workflow: (1) Diagnose β Opus, with ci-observer/ci-investigator evidence, identifies a parameterized rebuild as the right move. (2) Verify preconditions β guardrails #2-#5: queue check on each alpha pipeline, recent-trigger check, leaf-not-root check, destructive-parameter strict reading; defer or bail on partial verification, never curl on it. (3) Verify credentials (first use of session only β guardrail #7). (4) Issue the curl (pattern below); record the invocation in the journal using env-var placeholders. (5) Capture the response β Jenkins returns a queue item Location header on success (Location: .../queue/item/<n>/); record the queue id and resulting build number, and update the recent-trigger timestamp for guardrail #3. (6) Re-enter the observation loop β normal observer/investigator dispatch applies from here.
Curl pattern β Jenkins API tokens generally bypass CSRF crumb requirements:
curl -sS -X POST \
-u "$JENKINS_USERNAME:$JENKINS_TOKEN" \
-D /tmp/jenkins-trigger-headers.txt \
"$JENKINS_URL/job/elohim-genesis/job/dev/buildWithParameters?RESET_STORAGE=true"
# Check the Location header for the queue item
grep -i '^Location:' /tmp/jenkins-trigger-headers.txt
If the install requires a crumb (403, "No valid crumb"):
CRUMB=$(curl -sS -u "$JENKINS_USERNAME:$JENKINS_TOKEN" \
"$JENKINS_URL/crumbIssuer/api/json" | jq -r '.crumb')
curl -sS -X POST \
-u "$JENKINS_USERNAME:$JENKINS_TOKEN" \
-H "Jenkins-Crumb: $CRUMB" \
"$JENKINS_URL/job/elohim-genesis/job/dev/buildWithParameters?RESET_STORAGE=true"
The exact crumb requirement is install-specific β first-use verification (guardrail #7) tells you which path applies.
Pass multiple parameters as separate query params: .../buildWithParameters?RESET_STORAGE=true&SKIP_SEEDING=false&STEPS=all (URL-encode any value with spaces/special characters). Branched job paths mirror the multibranch structure, e.g. /job/elohim-genesis/job/dev/, /job/elohim-edge/job/dev/, /job/elohim-orchestrator/job/dev/ β append buildWithParameters?... for parameters or build for defaults.
Failure modes:
| Symptom | Likely cause | Recovery |
|---|---|---|
| 401 Unauthorized | Wrong username/token | Check env vars; do not retry blindly |
| 403 Forbidden, "No valid crumb" | CSRF protection on, token alone insufficient | Use the crumb pattern above |
| 403 Forbidden, "missing Job.Build permission" | Token lacks build permission | Stop. User must regenerate token with correct scope |
| 404 Not Found | Wrong job path (branch encoding, pipeline name) | Verify with curl ... /api/json?tree=jobs[name] |
| 200 OK but no Location header | Build was queued but routing changed | Check getJob for the new build |
| Thought | Reality |
|---|---|
| "Jenkins MCP isn't connected, I can't check the build." | MCP is anonymous-mode and connects automatically on workspace start. If truly missing in a subagent, fall back to WebFetch β public Jenkins reads need no auth. |
"Let me call mcp__jenkins__triggerBuild to retry." |
Anonymous can't trigger builds. Push an empty commit with [build:<pipeline>]; the orchestrator webhook dispatches it. |
"I'll add Authorization: Basic <token> to make MCP authenticated." |
Don't β OIDC intercepts the attempt and 302s into a redirect loop. Anonymous reads cover the entire diagnostic workflow. |
| "Let me grep the repo for the error." | The repo doesn't know what happened in CI. Fetch the sprint-report or console log. |
| "Let me spawn a ci-investigator subagent to check." | It may or may not have MCP loaded; either way it falls back to WebFetch if missing. Prefer ci-observer (Haiku) for surface scans; reserve ci-investigator (Sonnet) for cross-build correlation or low-confidence escalation. |
| "I'll assume the build ran against my latest commit." | Assume nothing β correlate SHA via changeSets[].items[].commitId before reading findings as proof-of-fix. |
| "99 β 90 failures means my fix barely worked." | Check fingerprints, not raw counts. A cascade-root fix drops several fingerprints at once; net failures can stay flat if pendings shift. |
| "The build status is green, so we're done." | Sprint-report aggregator is non-blocking β build can be SUCCESS with scenario failures. Always pull the report. |
| "I'll just read cucumber-report.json, it has everything." | Hundreds of KB of raw scenarios. Start with sprint-report.md (ranked, deduplicated); drill into cucumber only for specific stack traces. |
The aggregator (genesis/a2o/scripts/build-sprint-report.ts) produces fingerprinted, deduplicated findings. Fetch baseline and new reports via the artifact address book above, then:
Summary table β scenarios / passed / failed / pending gives the gross signal.POST /auth/register 503 fingerprint means every scenario needing a registered human also fails (surfacing as POST /auth/login 401 Invalid credentials because the fixture human was never created). Fixing the root often collapses most of the report in one shot; expect occurrences: to drop across multiple fingerprints together.Reading a fingerprint: each finding has a 12-hex fingerprint (SHA-256 of the normalized error), a normalized error message, an occurrence count, and the affected-scenarios list. Same error across runs β same fingerprint, even with different runtime IDs/timestamps (the aggregator strips UUIDs, ports, hashes) β a structural change to the error message is what produces a different fingerprint.
Map the failed Jenkins stage name to the component at fault: "Build DNAs" β Rust/WASM; "Build App" β Angular/TypeScript; "Seed Content" β doorway/conductor connection; "Deploy" β k8s/Docker; "E2E VERIFICATION (API)" β a2o scenarios (pull sprint-report.md).
| Class | Grep the log for | Common causes | Fix |
|---|---|---|---|
| DNA / WASM build | error[E, cannot find, unresolved |
Missing RUSTFLAGS for getrandom backend; incompatible dep versions; zome syntax errors | Confirm RUSTFLAGS='--cfg getrandom_backend="custom"'; verify Cargo.lock committed; check zome source |
| Angular build | error TS, Cannot find module |
Type mismatches after model changes; missing imports; circular deps | Run npm run build locally; check type sync between elohim-service and elohim-app; verify imports resolve |
| Seeding β connection | ETIMEDOUT, WebSocket, connection refused |
Doorway not ready; wrong admin URL; network policy blocking | Check doorway health endpoint; verify HOLOCHAIN_ADMIN_URL; check k8s pod status |
| Seeding β schema | missing required, validation failed |
Content files missing id/title fields | Run npm run validate in genesis/seeder; fix content files |
| Docker build | COPY failed, RUN failed, denied |
Missing build artifacts from a prior stage; Harbor registry auth; Dockerfile syntax | Check prior-stage artifacts; verify Harbor credentials; lint the Dockerfile |
mcp__jenkins__getStatus # overall Jenkins health
mcp__jenkins__getJobs # list jobs
mcp__jenkins__getJob jobFullName="elohim-genesis/dev" # job info (omit buildNumber args for latest below)
mcp__jenkins__getBuild jobFullName="elohim-genesis/dev" buildNumber=935
mcp__jenkins__getBuildLog jobFullName="elohim-genesis/dev" limit=-200 # tail
mcp__jenkins__searchBuildLog jobFullName="elohim-genesis/dev" pattern="error|failed|Exception|panic" ignoreCase=true contextLines=3
mcp__jenkins__getTestResults jobFullName="elohim-app/dev" onlyFailingTests=true
mcp__jenkins__getBuildChangeSets jobFullName="elohim-genesis/dev" buildNumber=935
mcp__jenkins__getFlakyFailures jobFullName="elohim-genesis/dev"
After genesis completes, verify deployment health via stats:dev/stats:prod commands, doorway health endpoints, and application smoke tests.