MeshForge NOC (Network Operations Center) assistant for LoRa mesh network development...
Scope note (2026-06-09): CLAUDE.md,
.claude/rules/security.md, and.claude/foundations/persistent_issues.mdare auto-loaded into every session β do NOT restate them here. This skill carries only operational reference that lives nowhere else. Version: readsrc/__version__.py, never hardcode it here (a stale copy sat at 0.5.5-beta while main was 0.6.1-beta). Handler surface: readcapability_index.md(next to this file) β auto-generated from the liveget_all_handlers()registry and test-pinned, so the count never goes stale by hand again (it drifted 60β64β96 when hardcoded). Routerβprimer upgrade 2026-07-14 (cross-model arc): domain model, fleet facts, harness layer, and the MF lint index now live here so a smaller model reading ONE skill sees the whole operating picture.
Measured 2026-10-04 over 112 sessions (08-15β10-04): skills here fire when a
hook or a slash command NAMES them β description-matching fired this skill 0
times. So the short form of this table lives in CLAUDE.md (loaded every turn);
this is the long form. Re-count at each freeze review:
grep -ohE '"name":"Skill","input":\{"skill":"[^"]+"' ~/.claude/projects/-opt-meshforge/*.jsonl | sort | uniq -c.
| Question shape | Method | Why this, not the obvious move |
|---|---|---|
| Something broke / failing / "why does X" | git log -S'<error text>' + grep -n <symptom> .claude/foundations/persistent_issues*.md FIRST, then mattpocock-skills:diagnosing-bugs |
the root cause is often already found and its cure undeployed (becd34cb was re-found 09-26) |
| Build or change behaviour | mattpocock-skills:tdd |
the test must FAIL on the old code before it passes on the new β a test that only ever passed pins the author, not the code |
| Define / decide / plan ("what is 1.0", "should we") | mattpocock-skills:grilling β result into .claude/ROADMAP.md |
turns a wish into checks that can FAIL; an unmeasurable goal has no distance |
| Module / seam / interface design | mattpocock-skills:codebase-design; decision record β engineering:architecture (ADR in .claude/plans/adr_*.md) |
|
| Upstream / library / protocol fact | mattpocock-skills:research |
read the PINNED source, never the version string ("2.8 is newer" β "2.8 is fixed") |
| Diff ready / "review this" | code-review |
β οΈ SEQUENTIAL on VolcanoAI (see Fleet Facts); two contextless readers one after another keep the independence |
| RF / radio / traceroute / SNR / mesh behaviour | this skill β measure at BOTH ends (the originator's journal holds the back leg) | an app screen is a rendering, not a measurement; label physics as physics |
| Box down / unreachable / probe blind | persistent_issues "Quick diagnostic tells" table, then ssh + the box's own uptime |
you may have observed a PATH, not a box |
| Session start / mid-session refresh | /warmstart |
|
| Wrapping up | memory-health, then the end-of-session double tap |
Two incompatible mesh ecosystems, one NOC bridging them:
meshtasticd (system service) owns the LoRa radio.
Surfaces: PhoneAPI TCP :4403 (single-consumer β a second reader eats
packets, #17-class), Web UI :9443, MQTT JSON uplink (per-box broker).
TX goes through send_text_direct() (meshtastic_protobuf_client.py);
connections through MeshtasticConnection (connection_manager.py). Never
read /api/v1/fromradio outside the one owner.rnsd (system service) is the ONE RNS host per box (shared
instance, AF_UNIX @rns/<instance>); every app is a client of it. LXMF is
the messaging layer; NomadNet the human client. rns/lxmf are MeshForge-owned
forks pinned by # MF-FORK-PIN in requirements/rns.txt (+mf.N markers).
ALL construction goes through open_reticulum() (utils/rns_init.py) β the
guarded chokepoint (MF019).src/gateway/rns_bridge.py bridges the two: mesh text β
LXMF (long_name in subject, [Mesh:xxxx] prefix) and LXMF β mesh
(@id/@short_name directed downlink). CanonicalMessage
(canonical_message.py) is the shared contract with MeshAnchor β byte-locked
twin file. SQLite retry queue in message_queue.py; delivery honesty via
compute_confirmation_view (#74 β never cross-population rates).meshforge-map serves :5000; collectors feed it; boxes
federate via each other's /api/status. Serving must never block on
collection (response byte-caches, #70/#71)./var/lib/meshforge/watchdog.json
β mini-dudeai (user unit, 30s rule loop) β ntfy pages. Silence is a failure
mode: cron verdicts (cron_verdict.sh β #78 probe), synth soak, tracer RTT.Where truth lives: journals + /api/status + scripts β never synthesis.
Radio RX truth = grep 'Received text msg' in the meshtasticd journal (json
greps miss via_mqtt, #75 trap). Service state = check_service() only.
| Service | Port | Protocol | Notes |
|---|---|---|---|
| meshtasticd TCP API | 4403 | TCP | PhoneAPI β single consumer (#17); never probe casually (#75/#76) |
| meshtasticd Web UI | 9443 | HTTPS | guarded against HAT-overlay port theft (#58) |
| RNS shared instance | 37428 | TCP | legacy port; live IPC is the AF_UNIX @rns/<instance> socket β owner must be rnsd (#69): sudo ss -xnpl | grep "@rns/" |
| MeshForge map | 5000 | HTTP | federator on VolcanoAI; /api/status is the probe surface |
| HamClock Live / API | 8081 / 8082 | HTTP | |
| MQTT | 1883 | TCP | per-box broker islands β no fleet consensus |
rnsd, meshtasticd, meshforge-map are system services (sudo systemctl β¦).
meshforge-mini-dudeai and nomadnet are user units (systemctl --user β¦,
logs via journalctl _SYSTEMD_USER_UNIT=<unit>).check_service() from utils.service_check (MF008)./etc/reticulum/config: sudo systemctl restart rnsd (authkey
derives from identity, #37) β then restart RNS-using services, and never
rapid-cycle rnsd fleet-wide (#69 race window).~/.config/meshforge/fleet_hosts β never trust a copy (this line
carried a 07-14 list missing moc4 + lehua until 10-04). Manager VolcanoAI;
watchers pane (NOT posture β that is scripts/fleet_posture.py check): PYTHONPATH=src python3 -m mini_dudeai.rollup (or /warmstart).config.proto; this line said moc3 only, and the stale roster helped cost an
hour); moc runs the cross-preset bridge. ST is the throughput leg for
ST<>meshforge<>RNS. Preset "drift" checks must best-match over BOTH
templates β and any RF-only witness must take segment as INPUT, or it
reports a healthy cross-preset radio as silent (the claw watch list did,
2026-07-30). Cross-segment visibility needs MQTT/RNS bridging, never RF.ssh kiai (rtun ProxyCommand); no LAN-routable HTTP until the dstnat lands./opt/meshforge checkout for the lab/mini user units. No MF watchdog there β
mini runs MINI_DUDEAI_ENABLE_WATCHDOG=0 (declared absent β error).git log -S'fan-out' for the cause, then measure
headroom incl. swap (10-04: swap 2015/2047 MB used). Never test it BY fanning
out on the manager.^src/ diffs β never during a soak; the
no-restart deploy is targeted git pull --ff-only per box. Always pull every
box after git push (divergence failure mode).Each menu action is a self-contained handler in src/launcher_tui/handlers/,
dispatched by handler_registry.py. Context arrives via set_context() (stored
as self.ctx); execute() receives the selected action tag, not the context:
from handler_protocol import BaseHandler # bare import: src/launcher_tui on path
class MyHandler(BaseHandler):
handler_id = "my_handler"
menu_section = "system" # which menu it appears under
def menu_items(self): # (tag, label, feature_flag_or_None)
return [("mything", "My Thing β does the thing", None)]
def execute(self, action: str): # action == the selected tag
self.ctx.report_action(ok, "Done", "It worked", "Failed", "It did not")
New handlers must be appended in handlers/__init__.py:get_all_handlers() or
they are silently dead UI (TestHandlerReachability guards this). The full
command surface β every section, tag, label, and feature flag β is in
capability_index.md (next to this file); grep it to answer "can the TUI
do X?", then open the handler. Regenerate it after touching any handler:
python3 scripts/gen_capability_index.py.
scripts/meshforge-launcher.sh # Primary interface (TUI) β sets PYTHONPYCACHEPREFIX
python3 src/standalone.py # Zero-dependency RF tools
python3 scripts/lint.py --all # Blocking gate (MF rules)
python3 -m pytest tests/ -v # Full suite
python3 scripts/parity_check.py # MeshForge<->MeshAnchor drift
python3 scripts/db_audit.py # DBSpec inventory (MF013)
The full suite runs ~18 min on the manager β longer than a Bash timeout, and the
idle reaper kills long background shells. Run it as a transient unit and read
the exit code from a FILE it wrote, never a | tail stream:
systemd-run --user --unit=mf-suite-$$ --working-directory=/opt/meshforge bash -c "python3 -m pytest tests/ -q -p no:cacheprovider >$S/suite.out 2>&1; echo \$? >$S/suite.rc".
Local green is not CI green: CI has no operator ~/.gitconfig, no fleet, no
~/.config β re-check CI for the exact HEAD before deploying.
Any model, any session β these commands re-derive state; never trust a summary:
bash scripts/honest_status.sh # THE check of record (exit 0 = green; 2/UNKNOWN β pass)
bash scripts/harness_audit.sh # is the second-brain spine itself wired? (9 legs)
PYTHONPATH=src python3 -m mini_dudeai.rollup # all-boxes posture, freshness re-derived now
git config core.hooksPath # must say .githooks (found OFF once β #29 spine dormant)
The harness is the portability layer, not the model (.claude/rules/model_advisor.md):
~/calibration_ledger.jsonl) records
VERIFIED claims per model so drift is measurable. Tag every completion claim
VERIFIED / BELIEVED / UNKNOWN (.claude/rules/calibrated_claims.md, auto-loaded)..claude/audits/review_provenance.md, never fake the pass.evals/local_brain/ case
(.claude/rules/honest_failure_modes.md point 10; check the consumer's live state first)..claude/foundations/harness_map.md (hooks β claim-gate β ledger β
mini β probes β paging; truth-source oracle table; SPOF list).Authoritative one-liners live in the scripts/lint.py header docstring β re-read
it when touching lint-adjacent code; new rules land there first. Grouped digest:
Path.home() (sudo breaks it). MF014 no
operator-specific values in src/templates/scripts/docs. MF015 no LAN IPs in
published docs (security.md).shell=True Β· MF003 no bare except: Β·
MF004 subprocess needs timeout= Β· MF010 daemon loops use
_stop_event.wait() not time.sleep().TCPInterface() Β· MF008 service
state via check_service() only Β· MF009 RNS.Reticulum() needs configdir= Β·
MF019 RNS construction ONLY via open_reticulum() chokepoint Β· MF023 map
collector interface creation only via the bounded helper.safe_import for first-party Β· MF011 repair
logic placement Β· MF013 SQLite via connect_tuned() + DBSpec Β· MF016 test
patch seams (utils.paths, not src.utils.paths) Β· MF018 no TUI
shell-escapes (in-domain principle) Β· MF020 never discard
apply_config_and_restart() result Β· MF021 mini-dudeai is observation-only Β·
MF022 installers route pip/apt through install_common.sh..claude/foundations/persistent_issues.md (auto-loaded).claude/foundations/domain_architecture.md.claude/INDEX.md.claude/research/src/utils/knowledge_base.pyTo answer "has this fleet seen X before?" search the WHOLE corpus (persistent issues + archive, foundations, rules, research, docs, memory topic files) in one deterministic shot β no INDEX to consult, no grep guessing:
PYTHONPATH=src python3 -m mini_dudeai.offline_oracle --retrieve-only "<question>"
BM25-ranked excerpts with paths land in ~1s. Drop --retrieve-only for a
cited local-LLM answer when Ollama is reachable (tier L; works with the
frontier away). In-app: TUI β dashboard β "offline oracle". Answers are
citation-gated β an answer citing nothing it was shown degrades to the
retrieval list, honestly.