NetGrimoire Infrastructure Control — v2.2026-08-20
▶ Tap a tile or pick from a list — everything here works with taps only, no typing
ControlLive
Refresh Gatus
Refresh All Backups
Refresh Everything
Cleanup All
Cleanup Gatus
Cleanup Backup
Cleanup Cron
Regenerate the backup stanza for one specific stack.
DeployLive
All stacks under swarm/ and compose/<host> — pick one and start it. Runs docker stack deploy / docker compose up -d directly, no CI/CD check/fix gate.
New-service pilot (gremlin.deploy.stage): normalizes a pasted compose snippet to Gremlin's stack standard, asks inline for anything it can't infer (stack name, UID/GID method, storage host, Homepage group/icon/description), shows the final staged YAML for review, then commits to traveler/services for CI/CD to check and deploy on your confirm — all on this page. ntfy only pings gremlin-alerts once staging succeeds or fails.
BackupLive
Pick a stack and run its backup. Stacks without gremlin.backup.enable configured return an error.
Backup Status
DockerLive
Cluster-wide Swarm summary, or merged swarm-task + standalone-container status for one node.
Tail last 100 lines for a Swarm service or standalone container.
Full inspect output for a Swarm service or standalone container.
Stacks diun has flagged with a pending image update. Versioned-tag stacks with a gremlin.feed label get a computed next-version target — verified against the registry before ever being offered, never guessed. Moving tags (e.g. latest) just need a redeploy to pick up the new digest. Everything else is flagged for manual review. Clicking Update bumps the tag (or the version-date label for a redeploy) in swarm/<stack>.yaml, commits, and pushes — Gremlin CI/CD validates and deploys from there, same as any other push.
Fleet-wide proactive scan (not diun-dependent) of every container running a moving tag (:latest, :dev, etc, fleet-wide, Swarm services and standalone containers alike). Resolves each one's actual running image digest against the registry's current digest for that tag — reports up to date / behind / unknown, plus version/build-date labels on each side where the image reports them. Anything found behind gets a Gremlin AI risk read (qwen2.5:14b, low/medium/high + reasoning) before an Update button is offered. Swarm services already tracked in swarm/*.yaml can be updated the same way as the panel above (redeploy — bumps gremlin.version to force a re-pull); standalone/non-GitOps containers are flag-only, no automated path. SSHes every host plus a per-candidate registry check and an AI call per stale image — this can take several minutes on a full fleet scan.
Homelable DiagramLive
Full Canvas Regen
Generate devices/cabling from Netgrimoire/Codex/Network/Port-Assignments.md. Preview shows the diff first; Apply writes the files and triggers a canvas regen.
HomepageLive
Reconcile the live homepage dashboard against current gremlin/*.yaml directive files — removes entries whose source now has homepage.skip:true, plus orphans left behind by a renamed or deleted device file. Preview shows what would go first; Apply writes and commits the change.
Fleet BootstrapLive
Brings a host with a local break-glass account (e.g. wizard, key-trusted, password-required sudo, now also LLDAP-registered so LDAP wins when reachable) or a genuinely fresh machine (username+password login) into the fleet: SSSD/LLDAP auth config, gremlin's SSH key, per-user passwordless sudo for gremlin/graymutt. Password is used only for this run and is never stored — see traveler/services/ansible/README.md.
Fleet Patch StatusLive
Read-only Ansible audit (apt upgradable + reboot-required + running kernel) cross-referenced against each host's actual running workloads, then AI risk-tiered (low/medium/high) via local Ollama — never cloud. Also runs weekly (Sunday 03:00 UTC). "Scan All Hosts" and each host's own "Check Now" always run a genuinely fresh check, never a cached result. Nothing is ever auto-applied — every host requires its own Apply tap, regardless of risk tier. Apply runs apt upgrade only (never dist-upgrade) and never reboots, even if the report says reboot-required — Reboot is its own button, its own confirm, and its own moment, and the up/down badge next to each host shows you it coming back live.
Error/warn lines for a service over a lookback window (default 1h).
Most recent lines for a service (last 6h window).
All logs (containers + journald) for a node over a lookback window.
Mail LookupLive
Traces a message's full journey through Postfix (queued, spam-scanned, delivered or rejected) by sender/recipient address or domain. Writes a saved report too.
Per-sender classify & act for phil@pncharris.com's inbox — vital (flag), unsubscribe (RFC 8058 one-click / mailto, then trash), and delete (trash) are live; important/normal/ignore are placeholders. Triage uncategorized senders, review or edit existing rules, and pull the latest inbox mail on demand — nothing is fetched automatically, only on refresh or Morning Briefing generation.
Generated reports — health checks, security audits, and docs like the znas filesystem pilot. Filter by type, status, severity, or tags; review a report in place, mark it reviewed or actioned, and export it as markdown to paste into Wiki.js.
DocsDisabled
Generate or update the wiki doc for a service — fetches upstream docs + NetGrimoire config.
Ingest DocLive
Fetch a URL, convert to markdown, and save it into the wiki under Reference/. Pasting the same URL again updates the existing page.
Doc BuildLive
Run the two-stage doc build (/doc-mechanical + /doc-design) for a service from the doc-audit checklist. Runs in background on gremlin-worker — result posted to ntfy gremlin-audits.
IngestLive
Run knowledge ingest for one repo — starts in background on docker4.
Ingest All Repos
ComicsLive
Runs for real, no dry-run — resolves ComicVine matches and pushes high-confidence ones into ComiXed; ambiguous matches land in Gremlin Reports for review. Starts in background on docker4.
ROMsLive
Preprocessor for the RomM library. Drop ROMs in
/export/Data/media/games-staging/inbox/<platform>/, then validate:
archive integrity, one ClamAV pass, an executable-inside-archive audit, and DAT
hash matching. clean deletes only positively-broken files —
anything that merely matches no DAT is quarantined, never deleted.
On-demand Kopia backup for any service — live step progress (sql dump, snapshot, etc), pass/fail, and a full fleet status table (last attempt, result) for every service.
Media RestoreLive
Restore Video
Restore Photos
ZFSPlanned
Snapshot
Sync
Mount
Dismount
Pool Status
Pocket GrimoireLive
Pocket Green
TODO: put pocket-green on the same 12hr / 90min timer Green's own auto-lock already runs on. Not built yet.
Home Station
Bring pocket up natively on znas. None of these steps touch the external drives.
Step 1 — Mount / Unlock
Unlock vault/pocket/green — tap to enter passphrase
—
Step 2 — Deploy
Builds one stack (or every stack in the queue) directly on znas — Gremlin SSHs in and runs pocket-deploy.sh itself, pocket-n8n is never involved (it can't deploy itself).
Step 3 — DNS: point warden + satchel at znas — both become CNAMEs to znas.netgrimoire.com
DNS → znas
Prepare to Deploy
Take pocket down on znas, sync everything onto the drives, export, and hand off to the road.
Attached Drives
zpool status for the three pocket drives (pocket-service / pocket-strongbox / pocket-cache) — which are currently imported on znas.
Step 1 — Dismount: Stop
Calls pocket-lifecycle.sh shutdown on znas — full teardown including pocket-n8n/webhost, safe for physical drive removal.
—
Step 2 — Prep
Calc reports space/counts for a category (read-only); Run does the actual prep work — git pull for dev, direct Radarr pocket-tag symlinks at /pocket/media/movies for movies, direct Sonarr pocket-tag symlinks at /pocket/media/tv for tv, symlink verification for green-movies.
Step 3 — Attach Drives
Actually imports all three pocket pools (pocket-service / pocket-strongbox / pocket-cache) if not already imported — as opposed to the read-only status check above.
Step 4 — Sync
Runs one catalog target (or the full sweep if left on "Full sweep") between znas and whichever pocket drive holds it — syncoid or rsync, local to znas, pocket never involved. Shows space needed vs. available before transferring. Full sweep is required before export will run.
—
No sync running
Last sync per resource
Step 5 — Export / Dismount Drives
Spot-checks the synced services content, then zpool exports the selected drive(s) — refuses if Sync hasn't reported success within the last hour, unless the sync check is overridden below.
—
Step 6 — DNS: point warden + satchel at Pocket — warden.pocket.lan → warden.pncharris.com, satchel.pocket.lan → satchel.pncharris.com
What's Gremlin doing in the next 24 hours (n8n schedules only — v1).
Security AuditLive
Runs all four layers (firewall, port sweep, port–service mapping, Forgejo/MailCow/OPNsense pilot) in sequence. Takes a couple of minutes — results land in the Action Log and as ntfy notifications on gremlin-audits, not here.
Fleet ShutdownIrreversible
Powers off the entire fleet — docker3, docker4, docker5, DockerPi1,
OPNsense (router/firewall), then znas last. Runs on znas, which announces (ntfy + GALDS)
and gives a 60-second abortable grace window before anything irreversible happens, then
gracefully stops mailcow/immich/docker5/znas's own compose stacks before powering off
each host in turn. There is no remote way to bring the fleet back up once this
completes — recovery needs a physical power-on, then
gremlin-fleet-recover.sh
(runs automatically 6 minutes after znas boots).
Requires typing the exact confirmation phrase in the prompt that follows — the click alone does nothing.