Start and manage the Elohim P2P Framework local development environment. Orchestrates conductor (identity/provenance), storage (content), and doorway (unified API)...
The root justfile is the public developer interface. It coordinates the
Holochain conductor (identity/provenance), elohim-storage (content and blob
projection), and doorway (HTTP/WS projection) without exposing crate-specific
paths, RUSTFLAGS, or cargo-target placement.
Run from the repository root:
just --list
just dev start # isolated single-peer stack
just dev start isolated true # start and seed
just dev status
just dev stop
The positional dev parameters are action profile seed build. Profiles:
isolated β local island DHT; the safe default.alpha β joins alpha via its bootstrap/signal endpoints and deployed hApp.Use just dev start isolated false true to force component rebuilds. Native
storage and doorway builds go to their explicit cargo-pool release slots;
DNA/WASM builds remain in-tree because hc dna pack requires ./target.
The alpha-shaped local topology is two doorways (A :8888 alpha stand-in,
bootstrap+signal owner; B :8889 apex/elohim.host stand-in, jessica-primary)
plus N conductor/storage peers (default matthew, jessica, james), all on
loopback. The canonical Dowell household fixture executes persistently under the workspace's gitignored
genesis/local-dev/household-dowell; its conductors/ subtree keeps each peer's conductor database and Lair keystore beside storage and archive state. MESH_DIR preserves an explicit alternative. A
loopback mongod (:27017, dbpath $MESH_DIR/mongo) fronts the doorways
so both doorways boot archive-backed (Mongo-side projection archive:
app_file_cache / warm-shell ShellArchive / resolver store, one database per
doorway: doorway-a, doorway-b). Without the binary (MONGOD_BIN unset and
no mongod on $PATH/~/bin) the doorways run archive-less with an INERT
warm-shell store β the production shape 18a65fd0d found un-wired β and
mesh status says so. Both doorways launch with --dev-mode --dev-signal-subscriber (env twin DEV_SIGNAL_SUBSCRIBER), and hc-mesh.sh
still passes it, but the flag is a no-op since f64d8c5bf β it parses
and is otherwise ignored. Signal subscription is gated by
should_subscribe_to_signals(projection_writer) alone (one input, no mode
flag): a projection WRITER subscribes to the multi-peer signal subscriber
under every stage, dev or fleet, and status.json compute.peers[]
populates from that β the surface the peer-conductor-resilience a2o reads. A
read replica (PROJECTION_WRITER=false) never subscribes. The image ships
mongod (che-devworkspaces udi-plus);
a2o resolves alpha-A/elohim.host to E2E_DOORWAY_ALPHA/E2E_DOORWAY_B,
so the failover feature runs against the local pair unchanged. Ordinary start resumes a complete stopped household without changing conductor or storage keys. Partial conductor/storage/archive state refuses; dual/iroh peers require iroh.key, while captured libp2p-only peers do not. The archive posture is persisted so archive-backed state cannot restart with a missing DB/runtime and an intentionally archive-less household remains archive-less. MESH_RESET=1 just mesh start is the stopped-only full-recast arm. A stopped legacy /tmp/elohim-local-mesh moves to the persistent root only when the destination is absent, then remains as a compatibility symlink for captured absolute paths. A complete stopped Dowell conductor sandbox moves peer-by-peer from the historical shared elohim/holochain/local-dev root into conductors/, preserving unrelated single-peer state and leaving per-peer compatibility links; partial or competing state refuses:
just mesh start # mongod β doorway A β doorway B β conductors β storage peers
just mesh preflight # every start_all refusal, checked BEFORE anything launches β one ok/REFUSED line per check
just mesh wait [--timeout N] # block until ready, or exit 1 on the first REFUSED line in the start log (default timeout 900s)
just mesh status
just mesh probe
just mesh quiesce
just mesh recovery <warm|cold> <peer> [--label k=v] # single recovery run (hc-mesh-recovery.sh)
just mesh recovery-matrix # cycle the recovery scenario library (MESH_PEER_TRANSPORTS + hc-mesh-recovery.sh)
just mesh stop
MESH_RELAY_BIN=<dir>/bin/iroh-relay just mesh start # holochain 0.7: the conductors need a REAL iroh-relay (see below)
**Conductor selection (2026-09-07):** `hc-mesh.sh` auto-detects the pinned fork at `$MESH_TOOLS_DIR/hc-fork-<submodule-pin12>/bin` (pin read from the `elohim/holochain-conductor` gitlink) and REFUSES a conductor whose line differs from the DNA's `hdk` line (`assert_conductor_matches_dna`; `MESH_ALLOW_TOOLCHAIN_SKEW=1` overrides). The workspace image ships stock 0.7.0 in `/opt/holochain/bin` (rebuilt 2026-09-07); the stock binary is a fallback, the fork is what the mesh launches.
**Detached start, preflight, and wait (2026-09-08).** `just mesh start` runs `preflight` first, then re-execs itself detached (`setsid nohup`, its own session) and returns in about a second β a `start` that blocked inline in the calling shell/tool-task used to get reaped along with the whole mesh process group when the tool call's own timeout fired (the 2026-09-07 fixtures-clone hand-off). `MESH_FOREGROUND=1` keeps the old inline/blocking behavior exactly. `just mesh preflight` (also run automatically by `start`) checks every refusal `start_all` can hit β binaries incl. the fork-pair pin, per-peer transport capability, every mesh-owned port free-or-ours, toolchain/DNA-line parity β before anything is generated or launched; a port already served by a live, PID-recorded process of this same mesh reads `ok ... (reusing)`, not a refusal. `just mesh wait [--timeout N]` (default 900s) polls the same readiness ladders `start_all` always blocked on and tails `start.log`, exiting 1 immediately the moment a line matches a refusal pattern instead of waiting out the full timeout. `MESH_RELAY_BIN` resolution now has a fallback: when unset, `detect_relay_bin()` picks the newest `$MESH_TOOLS_DIR/iroh-relay-*/bin/iroh-relay` (`sort -V`) β the same path `start_all`'s relay launch already used, so `preflight`'s relay check and the actual launch agree; a relay this exact mesh already launched reads `ok iroh-relay: already up on :3340 (reusing)` regardless of whether `MESH_RELAY_BIN` resolves in the calling shell.
**Stale-binary refusal (2026-09-11).** Preflight also REFUSES a pool `elohim-storage`/`doorway` binary older than the newest TRACKED source file under its crate's `src` (`assert_binary_newer_than_source`, compared by mtime, not commit time β a build-then-commit is current). `just gate` runs `cargo test --lib --bins`, which never refreshes a `--bin` artifact, so a pool slot commonly goes stale after a source fix lands: rebuild it directly (`cargo build --bin elohim-storage` with `RUSTFLAGS="--cfg getrandom_backend=\"custom\""`, or `cargo build --bin doorway` with `RUSTFLAGS=""`) in the pool slot. `MESH_ALLOW_STALE_BINARY=1` overrides for a deliberate measurement of an older binary.
MESH_TRANSPORT_BACKEND=libp2p|dual|iroh selects the elohim-storage Track-2
backend for the whole household run (default dual since 2026-08-23, matching the
alpha fleet; the conductor's own kitsune2 transport is a separate layer β iroh only
since holochain 0.7, homed to a relay). For example:
MESH_TRANSPORT_BACKEND=dual just mesh start
MESH_TRANSPORT_BACKEND=dual just mesh prologue
MESH_TRANSPORT_BACKEND=dual just test mesh features/dataplane/content-sync.feature
just test mesh-browserjust test mesh runs the mesh cucumber profile, which is tagged
@e2e and not @wip and not @browser and not @browser-only β it EXCLUDES every
browser scenario. Until the mesh-browser profile existed, the only browser
profile pointed at env: 'alpha' (the deployed fleet), so the sign-in portal
was never exercised against a mesh this run owns, even though MESH_PORTAL=1
serves it and the doorway proxies it at <doorway>/threshold.
just mesh start && just mesh prologue # portal served, cast seeded
just test mesh-browser '@auth' # the same act, through a real browser
It sets E2E_DEVICE_MODE=playwright and points E2E_APP_URL at the doorway.
That matters: on the mesh the DOORWAY serves the app itself, so app and portal
share one origin and a portal returnUrl is an ordinary same-origin redirect.
The doorwayToAppUrl default (localhost:8888 -> localhost:4200) is the
split-origin local-dev shape β ng serve beside a doorway β and stays the
default for that workflow.
Two things this lane needs that are easy to get wrong. The cast must be seeded
(just mesh prologue): without it every fixture login fails
INVALID_CREDENTIALS, which reads like an auth regression and is an empty
substrate. And the portal runs ng serve --live-reload false, because the
doorway proxies /threshold/* and a hot-reload WebSocket cannot traverse that
proxy β it fails its handshake against the proxied 200 and emits console errors
on every page, which a lane asserting a clean console after login cannot tell
apart from product errors.
dual and iroh require the selected STORAGE_BIN to be built with
--features "p2p p2p-iroh"; start refuses a default-feature binary and prints
the exact cargo-pool build command. mesh status prints transport=<mode> per
peer from the live/captured process environment, storage-restart preserves that
captured mode (including with MESH_RESTART_APPLY_PROFILE=1), and each household
sprint report stamps the mode in JSON and its Markdown header so runs remain comparable.
mesh status also prints one footprint <role> <name> rss=<MB> cpu=<pct> line per running
conductor, storage, doorway, and mongod process, followed by footprint total rss=<MB>. The values
come from ps and describe what is running, not the next-launch configuration.
hc-mesh.sh join-peer <fresh-name> appends ONE additional conductor + storage
peer to an already-running warm mesh β the late-join staging regime (regime 3 of
mesh-fixture-fidelity). It never restarts or reconfigures an incumbent, and it
refuses duplicate names, partial/cold meshes, and occupied derived ports before
launch. The receipt probe is genesis/a2o/scripts/late-joiner-receipt.ts (run
from genesis/a2o): exact signed NodeId discovery by every warm incumbent, no
restarts, bounded by three announce cadences.
hc-mesh.sh blocks [peer...] reads each peer's BlockSpan rows and the
rejected ops behind them (elohim/holochain/tools/hc-dbtool, resolved from its own
cargo-pool slot via DBTOOL_BIN; MESH_BLOCKS_DNA=<hash,...> forces a
rejected read for a DNA no block row names). It is READ-ONLY. Holochain 0.7
blocks the AUTHOR'S CELL until Timestamp::max() when one op that author wrote
integrates as invalid, and exposes no unblock β no admin call, no HDK host fn β
so a permanent refusal looks exactly like an unreachable peer: null arcs, zero
completed gossip rounds, nothing in the log. Lifting is a deliberate three-step
operator sequence, because hc-dbtool refuses to write while any live process
holds conductor.db open: hc-mesh.sh blocks <peer> -> hc-mesh.sh stop ->
hc-dbtool --databases <local-dev>/<peer>/databases unblock --cell <dna>:<agent> --yes
(omit --yes for a dry run) -> hc-mesh.sh start.
mesh quiesce measures an already-running mesh and records its bounded result
(one line per run, including the wall-clock, verdict, knobs and an io_baseline
write-throughput probe) under ${MESH_DIR:-<repo>/genesis/local-dev/household-dowell}. It never
starts or stops peers. The underlying maintained interfaces are
hc-mesh.sh start|stop|status|probe and hc-mesh-quiesce.sh; do not copy
their pacing environment into new npm aliases.
just mesh recovery-matrix cycles the checked-in mesh-recovery-scenarios.tsv
library across MESH_RECOVERY_SHAPES and MESH_RECOVERY_RUNS, alternates the
two-peer recovering slot, and reshapes only when peers/doorways change. Narrow
a run with MESH_RECOVERY_SCENARIOS; opt into per-run local quiesce records
with MESH_RECOVERY_QUIESCE=1. It composes the per-peer transport interface
(MESH_PEER_TRANSPORTS in hc-mesh.sh) with the single-peer recovery
primitive (hc-mesh-recovery.sh). Drive one scenario/run directly with
just mesh recovery <warm|cold> <peer> [--label k=v] before reaching for the
full matrix β pass --label scenario=homo-iroh (or homo-libp2p/homo-dual) so
recovery-timeline.py --table groups the row instead of filing it <unlabeled>; a red
poll prints recovery-detail: P1-bad=N [id=absent|<hash>β¦] P2-bad=β¦ on stderr naming
the failing ids, so a plateau is never a bare bit. Transport self-awareness is on by default; for a
before/after pair restart the arms with MESH_RESTART_ENV_OVERLAY="ELOHIM_TRANSPORT_SELECTION=off" (static
prior only; sampling and /p2p/status.transportPaths stay on) and read elohim_transport_route_total{reason}
on the recovering peer. The matrix seeds its notion of the current shape from the LIVE
mesh (matrix: live shape <peers>/<doorways>) so it only reshapes
(stop/start/prologue) when a scenario genuinely needs a different shape β
override with MESH_RECOVERY_LIVE_SHAPE="<peers-csv>/<0|1>", and point the
probe base at MESH_RECOVERY_PROBE_BASE (default 8090). A reshape is judged
by whether the mesh SERVES (every peer /health + /db/stats.contentCount>0,
both doorways 200 on /db/content/elohim-host-landing when doorways=1),
bounded by MESH_RECOVERY_RESHAPE_VERIFY_SECS (default 180s;
MESH_RECOVERY_RESHAPE_VERIFY_STUB=ok|fail for tests) β the prologue's own
exit code is logged only as advisory, since its seeder post-flight is a known
false red. Reshape retries per shape are bounded by
MESH_RECOVERY_RESHAPE_RETRIES (default 1); once exhausted, every scenario
still needing that shape gets FAIL(reshape) rows instead of another
regenerate attempt, and DOORWAY_B_PORT (default 8889) is honored alongside
DOORWAY_PORT throughout.
Fleet identity fidelity (MESH_DOORWAY_GATEWAY_SCOPING, default 1). A doorway
GATEWAY-SCOPES identifiers when it has a DOORWAY_URL: register AND login re-qualify the
local part with its own domain (auth_routes.rs gateway_domain + normalize_identifier),
so GET /auth/me answers matthew.dowell@alpha.elohim.host, not the string that was typed.
Every deployed doorway runs that way. Both mesh doorways launched WITHOUT the variable until
2026-08-29, so they stored identifiers verbatim and the household mesh was structurally
incapable of reproducing the fleet's naming β which is how genesis #1519, not a local run,
discovered a portal scenario asserting a bare name. mesh start now passes
DOORWAY_URL=http://localhost:$DOORWAY_PORT (doorway A) and :$DOORWAY_B_PORT (doorway B),
so a mesh human is susan@localhost. Set MESH_DOORWAY_GATEWAY_SCOPING=0 to run the
verbatim shape; the variable is then OMITTED rather than set empty, because an empty
DOORWAY_URL still reads as present to clap and would leak "doorwayUrl": "" into auth
responses. Scenarios must never derive the scoped name β read it from the auth response
(AuthResponse.identifier) or use genesis/a2o/src/framework/doorway-identity.ts; a
TypeScript re-implementation of gateway_domain is a second home for one rule and is wrong
wherever a test reaches a doorway at an address other than its configured one.
The sign-in portal (MESH_PORTAL, THRESHOLD_PORT). The doorway forwards
/threshold/* to THRESHOLD_URL with the path INTACT, and its binary default is
http://localhost:8081; on the fleet that is the doorway-app nginx sidecar. mesh start
now serves doorway-app on THRESHOLD_PORT (default 8081) under /threshold, so
http://localhost:$DOORWAY_PORT/threshold/login answers exactly as the deployed sidecar
does. MESH_PORTAL=0 skips it. Without it that path is a 502, and because
GET /auth/authorize 302s every unauthenticated caller to /threshold/login, the whole
OAuth authorization-code flow dead-ends locally β which is why the chaperone portal was
never exercisable before a push. It is launched detached and NOT waited on (~40s to first
paint; nothing else in the mesh depends on it), its port joins mesh_owned_ports so
mesh stop reaps it, and mesh status probes it THROUGH the doorway β a 502 there means
the proxy has no portal behind it. Drive it through the doorway, never against
THRESHOLD_PORT directly: doorway-app's environment.doorwayUrl is '' (same-origin),
so its API calls follow whatever serves it and would 404 against the dev server.
portal-restart (2026-09-11). The portal is a bare ng serve with no supervisor of its own, so the workspace RAM guard sheds it like any other fat node process β the only symptom was every @browser scenario timing out inside threshold-register-display-name with nothing naming the cause. just mesh portal-restart reaps the recorded pid (or notes it was already shed), relaunches, and BLOCKS until first paint (MESH_PORTAL_WAIT, default 180s) β unlike start, its caller is waiting on this surface specifically. mesh status now distinguishes "no portal on :8081" from a doorway proxy fault instead of printing one flat down, and just test mesh-browser REFUSES before launch when <doorway>/threshold/login isn't 200, naming this arm. hc-mesh-recovery.sh's backpressure witness reads
the CONDUCTOR log ($LOCAL_DEV_DIR/.sandbox_run_log[.<peer>]) for
conductor_receipt_max_s (JSON null when no receipt-latency line falls in
the window) and records conductor_receipt_scope per peer (per-peer vs
mesh-wide); it captures each peer's environ/exe BEFORE inflicting loss and
refuses (exit 5) without one, and the resulting record carries zome_path
(alive/dead/inconclusive/unknown) from the restart's zome probe.
The public-name membership authority (MESH_MEMBERSHIP, MESH_MEMBERSHIP_NAME,
MESH_MEMBERSHIP_PROBE_SECS, BEACON_BIN). Shared membership is the set of ORIGINS
currently eligible to serve one public name. On the fleet relay-addr-beacon projects that
set into Cloudflare as multi-A records; nothing about the set is decided by DNS β DNS is its
projection. mesh start now stages the SAME apparatus with a projection target the household
owns: one beacon leg per doorway (--sink file), owner alpha β http://localhost:8888 and
owner apex β http://localhost:8889, both writing exactly their OWN entry into
$MESH_DIR/membership/elohim.local.json. Same reconcile_membership call, same serving probe
(HTTP 200 on the doorway's own /health), same join2/leave3 hysteresis β only the sink
differs. The probe cadence is short here (MESH_MEMBERSHIP_PROBE_SECS, default 3s) so a
withdraw resolves in ~10s and a rejoin in ~6s; the fleet's 15s cadence would not fit an
acceptance window. A leg needs NO address detection (membership is origin-keyed), so no egress
echo endpoint is contacted and it works offline. start_membership_beacons also writes
$MESH_DIR/membership/authority.json, which the Prologue copies into the household fixture as
membershipAuthority β that is how features/dataplane/doorway-apex-transition.feature
resolves the public name THROUGH the document (try the advertised origins in order, stick to
the first that serves) instead of through a hardcoded doorway port, which the feature's own
preamble refuses as certification. mesh stop reaps the legs by recorded pid, mesh status
prints the currently-eligible set plus each leg's liveness, and mesh preflight REFUSES
without the binary (cd doorway/relay-addr-beacon && just gate, or BEACON_BIN=<path>).
MESH_MEMBERSHIP=0 stages nothing β and the absence is honest: the apex-transition scenarios
then fail naming a household that owns no membership authority, rather than reading a set
nothing maintains. WAN ingress continuity remains a separate prerequisite: the document proves
the routing decision, never that the resolved address is reachable.
hc-mesh-prologue.sh)hc-mesh.sh brings the mesh's PROCESSES up; it does not cast the household.
Once just mesh start reports both doorways healthy, run the Prologue to
seed the Act I substrate a2o's @act:i scenarios need β named conductor
identities, the base corpus rows every later leg patches bytes onto
(elohim-host-landing, lamad-spa, evolution-of-trust β seed a row here or
its stage leg 404s and the scenarios that need it env-red on the
precondition), the landing + lamad-spa bundles (browser AND the landing's SSR
server bundle, whose serverBlobHash is then stamped on EVERY peer's row β
the field is a diesel-direct deploy-projection artifact no sync plane
carries, so without the per-peer stamp doorway B's declared read stays NULL
and resiliency-saga ch06's cross-doorway scenario pends forever), the full
CI-order seed chain (identities cast BEFORE seed-humans on the mesh β doorway A's hosted pool is matthew's conductor, so hosted registrations must not claim it first), and the household fixture
manifest genesis/a2o/src/framework/fixtures/household-mesh.ts resolves
against, and Act I's own cast β the drill fixtures two resilience features
name (heal-target, chaos-ladder) with their household custody promises,
and the co-steward agreement (seed-household-costeward.ts) the saga's last
chapters count. On alpha that agreement is authored at run time by chapter 5
with adam as co-steward; the household mesh has no such author, so the
Prologue casts jessica co-stewarding the landing EPR instead:
just mesh start # bring the mesh up first β the Prologue never starts/stops it
./app/elohim-app/scripts/hc-mesh.sh prologue # (or: bash hc-mesh-prologue.sh directly)
just mesh prologue is routed by the root justfile (mesh recipe whitelist); hc-mesh.sh prologue is the same entry point.
Hosted provisioning is real (2026-09-11). A --dev-mode doorway used to skip
provisioning entirely and every hosted registration rode a singleton-Human
recovery path (one shared key); should_provision removed that path by
design, so the doorway must install a real cell per hosted human from a
bundle that exists on this host. hc-mesh.sh wires HAPP_BUNDLE_PATH at
the packed bundle the sandboxes install (the binary default,
/app/elohim.happ, is a container path that is NotFound on the mesh), and
POOL_COMPUTE_URL / POOL_COMPUTE_TOKEN / POOL_COMPUTE_PERFORMER name
the storage peer that NOTARIZES each hosted cell as a promise
(POST /api/v1/compute/grants) β all three required together or the leg is
skipped silently. The grant surface refuses unless
ELOHIM_COMPUTE_LOCAL_API=1 on that peer, the bearer matches its
ELOHIM_COMPUTE_LOCAL_TOKEN (a fixed dev default,
MESH_COMPUTE_LOCAL_TOKEN overrides it), and X-Verified-Performer is the
peer's OWN cell actor (same_actor) β peer_agent_key supplies that,
cached to disk because the doorways boot BEFORE the conductors and cannot
know it on a cold start (reconcile_doorway_pool_performer corrects and
restarts the one doorway that booted wrong, once storage answers).
Cast scope (MESH_HOSTED_CAST=lane|all). seed-humans.ts's
HOUSEHOLD_HOSTED_CAST allow-list (14 names β the humans some household-lane
a2o scenario actually signs in as) is the default cast once DOORWAY_URL is
loopback; a deployed fleet (alpha) keeps the full standing cast unchanged.
MESH_HOSTED_CAST=lane forces the allow-list even against a remote doorway;
MESH_HOSTED_CAST=all restores the full standing cast for someone who has
the RAM (~23GB of conductor heap for 29). DOORWAY_MAX_AGENTS_PER_CONDUCTOR
(MESH_DOORWAY_MAX_AGENTS, default 25) is the per-conductor hosted-agent
ceiling this cast is sized against β the fleet default of 50 is an operator
ceiling on kitsune2's per-space gossip budget that stalls arc convergence
past ~30 agents on one conductor; 25 = the 14-name lane cast + 3
prologue-hosted-* registrants (seed-hosted-humans.ts) + 8 headroom for
story-created humans that register and close again within one scenario.
Each hosted human costs ~786MB of conductor heap that closing a session does
not free β recycle it between hosted lanes with
just mesh conductors-restart && just mesh storage-restart <peers>.
Build stamp (version.json). package-angular-check.py refuses a
browser or SSR archive whose version.json is absent or whose commit is
empty, and for kind=server additionally requires the server stamp to equal
the browser one. Only the CI Jenkinsfile and package-angular.mjs build
(a side effect of a full rebuild) normally write that file β a plain
pnpm build/ng build, what a household dist is built with, leaves none,
so every stage leg is refused before a byte is uploaded. The Prologue now
stamps version.json on any locally-built dist BEFORE staging
(stamp_build_version, mirroring package-angular.mjs's five fields
exactly), but never overwrites a stamp that is already there β a dist from
just dev package or unpacked from CI carries the authoritative bytes.
The cast fix β named CONDUCTOR_URLS. An unnamed loopback conductor URL
(ws://localhost:4445,ws://localhost:4455,...) resolves by first-reachable-
wins in seed-conductor-identities.ts, which can cast the wrong human onto
the wrong conductor (observed 2026-08-21: Adam cast onto james's conductor,
zeroing household participants downstream). hc-mesh.sh's conductor_csv()
produces the named form instead β matthew=ws://localhost:4445,jessica=ws:// localhost:4455,james=ws://localhost:4465 β and mesh_seed_env() (source
hc-mesh.sh, then call it) exports it as CONDUCTOR_URLS alongside
HOLOCHAIN_ADMIN_URL / STORAGE_URL / DOORWAY_URL / PEER_STORAGE_URLS /
SEEDER_TARGET_PEERS / API_KEY_ADMIN β one source of truth for both the
Prologue script and an operator's shell. just mesh status prints the same
named form on its probe env: line.
Backlog record: genesis/data/timeline/backlog/mesh-prologue-cast-and-env-gaps.md.
cucumber-js -p local <files> runs the WHOLE suite: a profile's paths in
cucumber.mjs MERGE with CLI positionals rather than being replaced by them,
so pointing the local profile at one feature directory still loads every
path the profile itself declares. To re-measure a narrow set of scenarios
after a fix, either:
# an EMPTY .mjs config file, path relative to the REPO ROOT, so no profile paths merge in
pnpm exec cucumber-js --config path/to/empty.mjs features/dataplane/one.feature
# or: keep the profile's env/worldParameters, narrow by scenario NAME instead of path
pnpm exec cucumber-js -p local --name '^exact scenario title$'
Never assume a directory argument alone narrows the run under a named profile β it is additive, not a filter.
The mesh runs a declared dev-tier pacing profile (a preproduction-stakes
declaration, never a prod default) exported to the storage peers by
hc-mesh.sh β each knob overridable via its MESH_* twin:
PROJECTION_RECONCILE_SECS=30, ACQUISITION_RECONCILE_SECS=10 (acquisition + provide pin
reconcile tick β prod default 60s; chapter 11's exhaustion wait is bounded by it, 290 s β 50 s),
CONTEST_BACKOFF_SECONDS=120,
HEAL_MISSING_BACKOFF_SECONDS=60, ELOHIM_EVIDENCE_ABSENT_BACKOFF_SECS=600,
ELOHIM_HEAD_CORPUS_DIGEST=1, ELOHIM_ADOPT_BEFORE_AUTHOR=1 (cross-peer
head divergence has no adopt discharge without it), ADOPT_CONTEST_FANOUT=1
(concurrent declares race the conductor chain head; serialized lands
first-try), ELOHIM_NETWORK_STAKES=simulacra (the explicit Simulacra
declaration). The conductor side gets kitsune2 k2Gossip intervals patched
to 1000ms test cadence post-generate. The profile block in hc-mesh.sh is
the authoritative knob list. Beside the profile, every storage peer gets the three seed-only levers
ALLOW_SEED_NETWORK_STAKES=1, ALLOW_SEED_DELEGATES_COMPUTE=1, ALLOW_SEED_SHARD_MANIFEST=1
(the last unlocks PUT /admin/seed/shard-manifest; grandma-photos's four scenarios pend on a 403
without it) β mesh-only preproduction levers, never a prod default.
HOLOCHAIN_BIN) β and why the hc CLI is half of itAlpha runs the conductor fork; a mesh on stock is not a proving ground for it, and
hc-mesh.sh status says so on both the conductor RUNNING: and conductor NEXT LAUNCH: lines
([FORK] / [STOCK β alpha runs the fork, so this mesh is NOT at parity]).
The CLI is not a detail: hc sandbox REWRITES conductor-config.yaml in its own version's
schema (that is how -f pins admin ports), so a stock hc in front of a fork conductor hands the
conductor a file it refuses to parse. Therefore:
HOLOCHAIN_BIN takes a binary OR a directory holding holochain + hc; the matching hc
goes on PATH for generate AND run (the old code did run only β that asymmetry is how a
fork conductor got 0.6.0-schema configs and three hc sandbox run panics).start and conductors-restart refuse a mismatched pair and print both versions.
The check is on the FULL version, not the major.minor line β 0.6.0 and 0.6.3 agree on 0.6 and are
still schema-incompatible. MESH_ALLOW_TOOLCHAIN_SKEW=1 overrides for a deliberate experiment.network β¦ quic <RELAY_URL> on holochain 0.7 (iroh is the
only transport; the 0.6 stock webrtc <SIGNAL_URL> grammar is gone). mesh_network_args() fills the
relay argument from mesh_relay_url() β MESH_FORK_RELAY_URL, default the LOCAL iroh-relay the
script launches at http://localhost:$MESH_RELAY_PORT/ (3340). The relay is load-bearing, not a
NAT nicety (kitsune2 0.5: a conductor homes to relay_url at boot and dials only peers whose
advertised relay matches its own exactly). Measured 2026-09-03: three loopback 0.7 conductors with
relay_url pointed at the doorway β the 0.6-era "parseable placeholder" β booted clean and sat at
0 connections for a whole prologue. Knobs: MESH_RELAY_BIN (the iroh-relay 1.0.3 binary,
built with RUSTFLAGS="" cargo install iroh-relay --version 1.0.3 --locked --features server --root <dir>
β without --features server the install "succeeds" and installs nothing; unset = first
iroh-relay on PATH), MESH_RELAY_PORT (3340), MESH_RELAY=0 (skip the launch; you then own
MESH_FORK_RELAY_URL). just mesh status prints a relay row; the proof the mesh is REALLY up is
dump_network_stats on each admin port showing connections > 0, not three "ready" conductors.$MESH_DIR/fork-bin, $REPO_ROOT/.fork-bin, /opt/elohim/fork-bin β opt-in
homes only. A local fork build in the cargo-pool slot is passed explicitly:
HOLOCHAIN_BIN=/projects/.cargo-target-pool/family/dev/crates/dev/release just mesh start.app_ports:[]. just mesh stop then generate fresh with the fork's own hc.
Background: backlog mesh-prologue-cast-and-env-gaps.md (parity attempts 1β3).conductors-restart is still half an operation: it leaves each storage peer's app-websocket
handles pointing at a conductor that no longer honors them. Follow it with storage-restart and
confirm with zome-probe. The bridge supervisor re-mints the three supervised roles, but the
PeerStatus heartbeat holds a fourth, unsupervised client β so /health conductor.zomePath flaps
liveβdead on a ~60 s cycle until storage restarts (backlog
storage-stale-app-interface-token-after-conductor-restart.md).
storage-restart [peerβ¦] re-execs each storage peer in place from its captured /proc environ
(conductors untouched; a chaos-re-keyed AGENT_PUBKEY survives). The mesh runs the
parked mesh copy (/projects/.cargo-target-pool/family/dev/elohim__elohim-storage/mesh-bin/elohim-storage,
built with --features "p2p p2p-iroh" in the dev slot and copied there, where just gate's
cargo test cannot overwrite it) β hc-mesh.sh's print_iroh_build_command prints the exact pair when a
Rust cure must reach the mesh.
doorway-restart <a|b|c> (just mesh doorway-restart, 2026-09-26) is the doorway twin: it re-execs one
doorway on the binary now at its path with its captured environment. The gate's cargo test does not
refresh that binary either, so rebuild it first (CARGO_TARGET_DIR=<pool>/doorway__doorway-service/dev RUSTFLAGS="" cargo build --bin doorway).
A live peer's binary is read from /proc/<pid>/exe and recorded beside the environ
($MESH_DIR/storage-restart/<name>.exe); a DEAD peer is restored from that record, then a running
sibling's exe, then STORAGE_BIN β so the default path no longer has to exist. Before those
fallbacks, each peer consumes its per-storage-dir release slot at
$MESH_DIR/<peer>/release-adoption/slot/elohim-storage.next; the adjacent .next.json is printed as
the adoption receipt. A healthy boot archives both with an .applied-<timestamp> suffix and records
the archived slot executable; a failed boot archives them as .failed-<timestamp> and restores the
previous exe record, so no candidate can become a restart loop. The restart FAILS
(non-zero, storage-restart FAILED for: β¦) when any requested peer has no usable capture or does
not answer /health by port afterwards; an empty .environ is named out loud (a capture taken with
fs.copyFile on procfs is 0 bytes β the 2026-08-22 cascade; read to EOF instead, which is what
genesis/a2o/src/framework/fixtures/process-control.ts writeRestartCapture does). The captured
environment is authoritative; MESH_RESTART_APPLY_PROFILE=1 overlays THIS script's dev-tier
pacing knobs on the re-exec (a knob added after boot reaches a running mesh without regenerating
it; never touches AGENT_PUBKEY), and MESH_RESTART_ENV_OVERLAY="K=V K=V" overlays ad-hoc keys.
Mesh starts default ALLOW_COORDINATOR_UPDATE=true (MESH_ALLOW_COORDINATOR_UPDATE overrides) so
the rung-1 coordinator hot-swap vehicle (POST /admin/coordinators/sync +
scripts/ci/fleet-coordswap.sh) works out of the box. Run a local swap through the guarded mesh path:
app/elohim-app/scripts/hc-mesh.sh coordswap --happ <bundle> --peers <roster> [--apply]. Before
status, probe, start, join-peer, conductors-restart, or that coordswap pass-through, the
harness checks every live conductor sandbox exists and that no open handle into it is deleted. An
orphaned-data-root result refuses mutation and names path/mode evidence plus the explicit
stop/restore remediation; only MESH_RESET=1 just mesh start deliberately recasts identities. A peer restarted via storage-restart
re-execs the CAPTURED environ, so a mesh booted before that flag existed needs the overlay
(MESH_RESTART_ENV_OVERLAY="ALLOW_COORDINATOR_UPDATE=true") or a full mesh restart to accept swaps.
The re-exec closes inherited fds β₯3 first β a caller holding a flock (the a2o
mesh lock $MESH_DIR/a2o.lock, which serializes concurrent agents' mesh-touching
commands) otherwise leaks the lock into the long-lived peer (2026-08-22: three peers owned the lock
for 40 min). The restart re-resolves every peer's pid AND agentPubKey into the household fixture
(fixture-refresh does only that): the a2o chaos drills kill and verify peers BY THAT PID, so a
stale fixture reads as kill ESRCH, and custody commitments name providers by agent key.
just mesh monitor (hc-mesh-monitor.py, port 4210 via the mesh-monitor
devfile endpoint; honors MESH_MONITOR_PORT) serves the one-page live
dashboard: component liveness, per-peer convergence gauges, a gate-legs
panel mirroring fleet-quiesce-gate.sh's exact PASS predicate, log tails,
and a phase/progress status bar.
hc-mesh-spin-detector.sh, hc-mesh-chaos-rekey.sh)The alpha conductor spin (sys-validation retrying unfetchable dependencies every 10 s,
read-pool saturation logged at ~1000 lines/s β backlog
alpha-conductor-sys-validation-spin-unfetchable-deps.md) is measured and staged at the desk:
bash app/elohim-app/scripts/hc-mesh-spin-detector.sh --window 20 --cycles 9 # verdict SPIN|QUIET + JSON
bash app/elohim-app/scripts/hc-mesh-chaos-rekey.sh --peer james --tag act1 --phase author # james-originated content, referenced by the others
bash app/elohim-app/scripts/hc-mesh-chaos-rekey.sh --peer james --tag act1 --phase rekey # destroy james's chain (conductor + key), keep his storage DB
bash app/elohim-app/scripts/hc-mesh-chaos-rekey.sh --peer james --tag act1 --phase measure # detector for 3+ min on the survivors
The detector reads per-conductor CPU from /proc, log-rate spectroscopy from .sandbox_run_log
(saturation / No peers to fetch / missing dependencies with the count extracted) and says SPIN
when the missing-dependency count is non-decreasing across β₯3 cycles with saturation above threshold.
It is validated against alpha-shaped synthetic logs both ways. The three conductors are children of
ONE hc sandbox run parent: a single-peer kill may drag the others down; the script restarts them
from their existing sandboxes (no regenerate, no re-key) and says so loudly, because it changes the
measure. Scenario: genesis/a2o/features/resilience/conductor-validation-spin.feature.
Shutdown ownership: mesh start/restart paths persist PID + process-start identity under
$MESH_DIR/pids/<role>-<name>. just mesh stop validates those records, merges them with listeners on
the configured mesh TCP/UDP ports, and terminates only that exact set. Service-name pgrep -f patterns
remain only as a warned, /proc/exe plus mesh argv/cwd-validated compatibility fallback for pre-PID
meshes; a shell whose argv merely mentions a service path is not a candidate. The recovery matrix calls
stop synchronously β it no longer needs a separate-session workaround to survive shutdown.
just test mesh # whole Act I lane (@e2e, not @wip/@browser) under cluster-state.act1-household.yaml
just test mesh features/dataplane/doorway-failover.feature
just test mesh '@act:i and @dataplane'
just test mesh sources hc-mesh.sh's mesh_seed_env and exports the Prologue's a2o env block; a
scope argument makes it write a paths-less config so the run is actually scoped (cucumber merges a
profile's paths with positionals otherwise). A roster-backed lane then gives every configured storage peer one shared 75-second window to return the Prologue anchor through its authenticated lamad/content_store read rail; malformed, error, or incomplete responses refuse before scenario setup. MESH_ALLOW_NO_PROLOGUE=1 explicitly bypasses this seeded-anchor check. This establishes lamad zome-call readiness only; scenario assertions still prove write admission. @act:<i|ii|iii|host> resolves to the act's baseline caps;
an undeclared @requires: cap warns loudly once per run. Spec: genesis/a2o/LAYERS.md.
Destructive steps (kill/restart/pin/delete) ride ONE gate β substrate-scope.ts destructiveAllowed():
the lane's declared owned-substrate cap (true only in cluster-state.act1-household.yaml), with
A2O_ALLOW_DESTRUCTIVE=1|0 as the operator override, never fail-open. So on this lane they RUN; two
consequences are already handled by hc-mesh.sh: doorways launch with generous, overridable
DOORWAY_MEMBRANE_{SHAPE,CHALLENGE,BAN}_THRESHOLD (every a2o request is one loopback client and the
churn scenarios exceed the binary's 1200/min ban β a tripped membrane answers 403 x-membrane:deny
to the rest of the lane), and storage-restart refreshes fixture pids. A scoped run overwrites the
full-lane cucumber JSON unless you pass your own CUCUMBER_JSON_REPORT=<path>; every run still mints
its own run-identified sprint report. Process control (find a pid by exact argv on
its port, capture environ+exe to EOF, SIGTERMβSIGKILL, re-exec the same argv/env/cwd) lives ONCE in
src/framework/fixtures/process-control.ts; a custody provider matches a peer by libp2p peerId OR
the fixture's agentPubKey (reconcile/custody.rs accepts either namespace).
just test mesh and just mesh start|stop claim the household lease through berth before
touching the mesh β the mesh is one per workspace, so this is the router, not an add-on.
BERTH_CLASS=verify|measure (default verify), BERTH_TTL (default 1800 s; verify's max is
3600 s β a longer hold is a measure and is refused as one). A live holder in another session
refuses in under a second (exit 3), naming the holder. Exit 4 (no session resolvable β a mooring
carries the Claude pid and its start time, and a claim resolves by process ancestry; from a
non-Claude shell pass --session or BERTH_SESSION) is the ONLY exit code a caller proceeds
past unleased; every other failure propagates.
A verify lane is fenced by timeout at its TTL: the lane re-enters under timeout --signal=INT --kill-after=30, so past the lease it is killed (INT, then KILL 30 s later) and the recipe ends
BUDGET-EXCEEDED: <class> lane ran past its <ttl>s lease, exit 3. Scope the lane down, or send it
through just measure instead of raising the TTL.
--class measure is refused on the dev berth (just test mesh BERTH_CLASS=measure and just mesh matrix|recovery|recovery-matrix all claim mesh --class measure first) and the refusal
names just measure <scope>. MEASURE_ON_DEV_BERTH=1 is the declared override: it claims (ttl
BERTH_TTL, default 3600 s for this class), writes a kind: override row to the berth ledger,
and emits dev-berth-held-by-measure@1 immediately β that measure renders on the
push-delivers-within-budget headline line. matrix, recovery, and recovery-matrix are
measure-class for exactly this reason: windows, restarts, and repeated shapes are what the dev
berth was never meant to hold.
just mesh start holds the daemon mesh class instead β no TTL, taken over only when the
holder's mooring is dead. A session holding it runs its own just test mesh lanes as already
covered (a renew, not a new claim). just mesh stop refuses under another live holder (owner-check,
exit 3, holder named); MESH_STOP_FORCE=1 overrides, on the record.
just measure <feature-path> [--on jessica|adam] [--gap <id>]
(genesis/agentic/compute/measure.sh) is the developer verb that leaves the dev berth entirely:
it authors ONE feature path (a tag expression is refused; v1 is one path) as a peer-executed
stage, submits it to the provider that holds the stores, and returns immediately. Sub-verbs:
grant --on jessica|adam, worker --on jessica|adam, status, poll, fixture --on jessica|adam <feature-path>. It refuses before launch, naming the fix, rather than failing
mid-run β missing compute-executor/ark binaries, an unresolved provider. MEASURE_DRY_RUN=1
prints the resolved env and the exact command(s) without running them. --on adam refuses today,
by name, listing the operator items still open (adam is a k8s pod on shem, provisioned by the
cluster operator) β only jessica is a live provider. The listener writes the brit validate ref,
sprint-report-peer-stage-*.json, the gap fulfil, and the habit delta when the stage's completion
arrives; a rung-H (household stand-in) result is labelled as such β it never claims offload that
wasn't proven.
berth status shows every lease with its class, held-for duration, and holder, and marks an
unclaimed-but-expired lease EXPIRED. berth say [--to SESSION] TEXT is the cross-session
channel for asking a holder for a window instead of guessing or forcing.
just gate # changed projects vs origin/dev + worktree
just gate elohim-storage # one manifest project
just gate doorway/doorway-service
just test app
build-manifest.json gate.projects owns both detection and typed local
execution. The shared runner resolves explicit cargo-pool workspaces and the
crate-specific RUSTFLAGS; direct native cargo build/test/check/clippy
without CARGO_TARGET_DIR is intentionally denied by the disk guard.
just seed validate
just seed apply local
just seed stats
just seed diagnose
just look page http://localhost:4200/epr/elohim-protocol
just look graphos list
There is currently no content-seed dry-run. Historical --dry-run and
--validate-only flags were ignored by seed.ts and could perform real
writes; use just seed validate for the non-writing schema check.
just status runtime
just status habits
just status saga
| Service | Default | Probe |
|---|---|---|
| Angular | 4200 | page request |
| Doorway | 8888 | /health, /status, /db/stats |
| Doorway health watchdog | 8079 (A) / 8089 (B) | /health, /ready, /health/serving on their own OS-thread runtime (DOORWAY_A_HEALTH_PORT/DOORWAY_B_HEALTH_PORT; alpha runs 8079) |
| Conductor app | 4445 | WebSocket |
| Conductor admin | dynamic | elohim/holochain/local-dev/.hc_ports |
| Storage | 8090 | /health, /db/stats |
| Doorway B (mesh) | 8889 | /health |
| mongod (mesh) | 27017 | tcp open; $MESH_DIR/logs/mongod.log |
app/elohim-app/scripts/hc-start.sh is the single-peer owner. Its native
builds are pool-aware; its DNA builds intentionally are not redirected.DOORWAY_AUTH=auto|secure|keyless, default auto): with a
mongod (MONGOD_BIN, MONGO_PORT default 27017; dbpath
.local-dev/mongo) it runs SECURE β per-workspace JWT_SECRET +
API_KEY_ADMIN generated once under .local-dev/doorway/, chaperone
provisions, no --dev-mode; without one it runs KEYLESS (native local-first;
--dev-mode passed only as the startup declaration the config validator
requires). DOORWAY_AUTH=secure fails fast when no mongod is found. Both
--conductor-url (app :4445) and --conductor-admin-url (hc sandbox's
random admin port) are passed explicitly β the app_port-1 derivation does
not hold for hc sandbox.hc sandbox --piped reads the lair passphrase from stdin. -f pins admin
ports, -r pins app ports, and -n/-d create named sandboxes.ws://signal.localhost:8888). This includes
join-alpha: CONDUCTOR_SIGNAL_URL defaults to wss://signal.alpha.elohim.host
(the fleet's own value); wss://doorway-alpha.elohim.host/signal panics the
conductor at boot (parsing tx5 sig url InvalidLastSymbol) β proven 2026-08-28.HTTP_PORT; pass --http-port.ENABLE_P2P=true, P2P_PORT, and the conductor agent key.--storage-url plus comma-separated --storage-urls./db/paths route.join-alpha workspace-stack idempotency (2026-09-06). hc-start.sh's join-alpha profile now reuses a conductor whose recorded admin AND app ports (.hc_ports) both answer a live TCP probe instead of starting a second sandbox on the pinned join-alpha app port (4485) β the prior code fell through to a fresh hc sandbox generate whenever hc sandbox call --running failed against a healthy but CLI-schema-mismatched conductor. .hc_ports is now deleted only when NOTHING answers its recorded admin port; unconditional deletion had erased the one structural signal a live-but-undetected conductor leaves for workspace-to-fleet-release.steps.ts. The stack also prefers the cargo-pool's DEBUG elohim-storage/doorway binary over a cold release build when no release binary exists yet β FORCE_BUILD=1 (equivalent to --build) still forces a compile. A new DOORWAY_PORT knob (default 8888) joins STORAGE_PORT (8090) so a workspace peer can run BESIDE the household mesh, which owns both defaults: STORAGE_PORT=8093 DOORWAY_PORT=8889 NETWORK_PROFILE=join-alpha ./hc-start.sh. And the join-alpha stock-conductor refusal now states the real 0.7 reason β alpha is a TWO-RELAY fleet and the ethosengine fork carries a cross-relay preflight fix stock kitsune2 lacks, so a stock conductor lands partitioned rather than merely unconnected β and names the harbor extraction path (elohim-edgenode:conductor-<hc12>, layer 25/26 ALONE β extracting every layer lets the base image's stock binaries overwrite the fork) as a fleet-parity pair without a 45-minute fork build, replacing the retired tx5 "never connects" measurement. (When the household mesh is up, this is superseded automatically by the T3 workspace-peer auto-offset below.)
Persistent join-alpha workspace peer (2026-09-30). just dev conductor alpha always uses the deterministic t3-join-alpha sandbox rooted under elohim/holochain/local-dev, whether or not the household mesh is currently running. With a live mesh, storage and doorway default to 8095 and 8898; its conductor app socket defaults to 4485, and the optional Agent SDK uses 8096 to avoid storage. Explicit port values win. Join-alpha storage and the conductor sandbox share the durable local-dev root, so a normal stop/start resumes the registered sandbox in place with the same conductor identity and chain. The script refuses to overwrite incomplete or inconsistently registered sandbox state. On a genuinely new sandbox, enrollment/key creation requires an explicit CONDUCTOR_ENROLL=1; omit that flag for every resume. Join-alpha resumes pin the admin socket to the recorded .hc_ports port with hcβs global --force-admin-ports option, preserving existing storage and doorway pairing. CONDUCTOR_ADMIN_PORT=39097 explicitly selects a port from 1 through 65535 for enrollment or resume; invalid values refuse before launch. A running conductor at a different admin port refuses the override until the owned conductor is stopped. A first enrollment with neither override nor recorded port keeps hcβs random-port default. Export sourced environment assignments (set -a; source /private/path/join-alpha.env; set +a) before invoking the launcher so its child receives the persistent profile settings. Example first enrollment: CONDUCTOR_ENROLL=1 just dev conductor alpha; subsequent runs use just dev conductor alpha. For workspace service parity, pass the deployed fork pair with HOLOCHAIN_BIN and HC_CLI_BIN, and set STORAGE_PORT=8095 DOORWAY_PORT=8898 CONDUCTOR_APP_PORT=4485 ELOHIM_AGENT_PORT=8096 when running beside the household. Isolated development keeps its existing random sandbox/default-port behavior. For join-alpha, doorway hosted provisioning defaults to the durable deployed hApp bundle selected by the profile; HAPP_BUNDLE_PATH=/absolute/path/to/app.happ explicitly overrides the doorway bundle only. Set STORAGE_HAPP_PATH=/absolute/path/to/app.happ to retain an explicit coordinator candidate across storage restarts: hc-start.sh passes it as storage --happ-path in both ordinary and release-channel launches. The path must be a readable regular file before launch; omitting it keeps the existing storage default. Coordinator synchronization still respects ALLOW_COORDINATOR_UPDATE. The conductor sandbox still installs or resumes the profile-selected conductor hApp, and a normal resume does not regenerate its identity.
Env-tunable conductor-ready wait + scoped stop (2026-09-08). CONDUCTOR_READY_TIMEOUT (default 240s, was a hardcoded 45Γ1s loop) bounds how long hc-start.sh waits for the "admin_port" line in .sandbox_log, printing the measured elapsed seconds on ready β this only raises the ceiling, the isolated case still returns the moment the line appears. just dev stop no longer runs a blind pkill -x holochain + fuser -k 8888/tcp 8090/tcp 8095/tcp (catastrophic beside a running mesh, whose conductors are also named holochain and whose doorway A/storage matthew sit on exactly those ports): hc-start.sh now records the pid it resolved for each role (conductor, storage, agent-sdk, doorway) beside the sandbox in .hc-start-pids/, validated against /proc/<pid>/stat's start-tick, so hc-start.sh --stop (wired from just dev stop) reaps only a process this exact script incarnation actually started β a household mesh running alongside is left untouched.
just status runtime
fuser 8888/tcp 8090/tcp 4445/tcp
cat elohim/holochain/local-dev/.hc_ports
curl -s http://localhost:8888/status | jq .
If a port is held by an old binary, stop the stack and restart it before
trusting wire shapes. If the shell's prestart needs unavailable wasm-pack
but the generated package already exists, the specialist escape hatch is:
cd app/elohim-app
pnpm exec ng serve --proxy-config proxy.conf.mjs --disable-host-check
| File | Purpose |
|---|---|
justfile |
public eight-verb interface |
app/elohim-app/scripts/hc-start.sh |
single-peer stack |
app/elohim-app/scripts/hc-mesh.sh |
local multi-peer mesh |
app/elohim-app/scripts/hc-mesh-quiesce.sh |
bounded quiesce measure |
app/elohim-app/scripts/hc-mesh-prologue.sh |
Act I Prologue cast (seeds an already-running mesh) |
app/elohim-app/scripts/hc-mesh-spin-detector.sh |
conductor spin detector (CPU + log-rate spectroscopy β SPIN/QUIET + JSON) |
app/elohim-app/scripts/hc-mesh-chaos-rekey.sh |
stage the unfetchable-dependency class: author on one peer, re-key it, measure the survivors |
app/elohim-app/scripts/hc-mesh-perf-watch.sh |
continuous 15 s timing watch: per-service CPU, direct-vs-doorway latency, breaker state, storage zome path; writes $MESH_DIR/perf/watch.jsonl + SPIKE lines to perf/watch.spikes. Run it after start β it is how the doorway first-SSR-render stall was found |
genesis/orchestrator/gate-runner.mjs |
manifest gate selection/execution |
genesis/agentic/bin/pool-lib.sh |
cargo-pool family and slot authority |
elohim/holochain/local-dev/.hc_ports |
local conductor ports |
MESH_CONDUCTOR_LAUNCH=ark, 2026-09-02)Each household conductor runs as the child of an ark (the tevah compute envelope; ARK_BIN
defaults to /projects/.cargo-target-pool/family/dev/elohim/dev/debug/ark, else command -v ark;
build with cd elohim && CARGO_TARGET_DIR=/projects/.cargo-target-pool/family/dev/elohim/dev RUSTFLAGS="" cargo build -p elohim-ark).
The script writes <peer>/ark/manifest.json + berth.json per peer, records pids/ark-<peer>, and
mesh_conductor_pid <peer> reads the child pid from <peer>/ark/passport.json; just mesh status
shows conductor(ark) <peer> pid= incarnation= ready= rows. Two refusals are new in this mode:
start_all refuses to wipe a data root while an ark or conductor pid survives, and a peer's
incarnation read-back from an existing passport is fail-closed (malformed passport β that peer's
launch aborts). Toolchain parity is skipped in ark mode (like direct); jq is required.
The conductor readiness ladder ends by hashing the kernel-observed executable (/proc/<pid>/exe)
and comparing it with the manifest pin; a staged replacement path is not running identity.
Rebuild ark before using a mesh script that declares executable_identity: older arks refuse
the unknown probe. A failed readiness rung is witnessed and judged by restart policy rather
than reported as an intentional successful stop. This does not implement binary adoption or rollback.
Death drill: kill -9 $(mesh_conductor_pid jessica) then
$ARK_BIN witness ls --berth elohim/holochain/local-dev/jessica/ark/berth.json (witness within ~1 s;
the ark restarts the conductor; the ark's incarnation is unchanged). Rolling the arks onto a new binary:
SIGTERM the ark-* pids (each stops its conductor with SIGINT + grace), then
MESH_CONDUCTOR_LAUNCH=ark just mesh conductors-restart; afterwards just mesh storage-restart <peer>
for any storage peer the restart flags with a stale app-interface token.
direct and ark modes, 2026-09-23)A fork conductor from the 61565f320 lineage serves its own OpenTelemetry instruments as
Prometheus text at GET /metrics when HOLOCHAIN_PROMETHEUS_LISTEN=host:port is set
(hc_conductor_workflow_{duration,run_total,in_flight}, hc_db_connections_use_time,
hc_holochain_p2p_request_duration, hc_ribosome_zome_call_duration, β¦). The mesh exports it
only from the per-process launch modes: direct and ark give each conductor
127.0.0.1:$(conductor_metrics_port <index>) = 9464 matthew, 9465 jessica, 9466 james, and the
preflight/join-peer port checks include those ports in exactly those modes. The default
hc sandbox run supervisor launches every conductor from one environment, so it exposes none β
curl -s localhost:9464/metrics | grep hc_ returning nothing under the default mode is expected,
not a broken exporter. An older fork binary ignores the variable.
Every storage peer starts with ELOHIM_RUNTIME_CONFIG_PATH=<mesh>/<peer>/runtime-config.toml
(created empty by start, kept across storage-restart). elohim-storage's runtime-config watcher
is OFF unless that variable names a file, and the a2o release ceremony writes
ELOHIM_RELEASE_CHANNELS = "β¦" into exactly that path then POSTs /admin/runtime-config/reload β
a peer started without it answers /admin/adoption with sweeps: 0 forever and rung-5 station 1
times out on "the household's runtime follows release channel β¦". A flag flip or a channel follow
lands on the RUNNING peer within one poll, no restart.
To build and package an Angular EPR app, run just dev package app/elohim-app (or the app directory). The command builds first and stamps both outputs after success. The SDK packager checks browser assets, version context and the SSR entry contract before producing content-addressed archives. It leaves the build directory intact and does not upload or author a head. Local packaging and CI staging use the same packing implementation. Package checks establish artifact structure; just test mesh remains the proof of peer propagation, doorway rendering/cache behavior and browser boot. See elohim/sdk/scripts/package-app.mjs --help for checking existing builds, explicit output and single-bundle options.
The built-in adapter is Angular, not a universal EPR-app requirement. EPR_APP_ADAPTER=./my-adapter.mjs just dev package ./my-client supplies a local adapter with build, layout, package validation and runtime-check hooks. The common layer archives and hashes; adapter checks own framework/host compatibility. Native desktop and non-browser Wasm do not inherit SSR or browser bootstrap requirements. Packaging does not execute the separate runtime check or certify delivery; the SDK guide at elohim/sdk/scripts/package-app.md defines the extension contract and its verification API.