Skip to content

Architecture

This page covers Docker Compose patterns, container security rules, networking, and directory conventions. For host-level setup (UID/GID allocation, storage, multi-server deployment), see Infrastructure. For development workflow (Renovate, commits, releases), see Contributing.

Compose File Standards

Every service in this repo follows these conventions:

services:
  example:
    image: registry/image:tag@sha256:...   # Always pin to digest
    container_name: example                # Explicit name for predictable references
    env_file:
      - .env                               # SOPS-decrypted secrets
      - ../shared/env/tz.env               # Shared timezone
    user: "3100:3100"                     # Hardcoded UID:GID (see Infrastructure § UID/GID Allocation)
    deploy:
      restart_policy:
        condition: on-failure             # Restart only on crash (non-zero exit)
        max_attempts: 3                   # At most 3 automatic retries; not a rolling window
        window: 120s                      # Swarm-only; ignored by standalone Compose
    networks:
      - <service>-frontend                 # Traefik-facing network
    mem_limit: ${MEM_LIMIT:-<default>}     # Prevent runaway memory
    pids_limit: 100                        # Prevent fork-bomb DoS
    security_opt:
      - no-new-privileges=true             # Block privilege escalation
    cap_drop:
      - ALL                                # Drop every capability …
    # cap_add:                            # … re-add only what is provably needed
    #   - NET_BIND_SERVICE
    read_only: true                        # Immutable root filesystem
    tmpfs:
      - /tmp                               # Writable scratch space
    healthcheck:                           # Required for --wait deploys
      test: [...]
      interval: 15s
      timeout: 5s
      retries: 3
      start_period: 10s
    labels:
      - "traefik.enable=true"              # Opt-in to Traefik discovery
      - "traefik.http.routers...middlewares=chain-auth@file"

Restart policy under standalone Docker Compose:

  • Compose maps condition and max_attempts to the Docker Engine restart policy, but drops window and delay. The example above therefore becomes on-failure:3, not three retries per 120 seconds.
  • Moby checks restartCount < MaximumRetryCount and increments the count for each automatic restart. Running for at least 10 seconds resets only the restart backoff delay, not the cumulative retry count; an explicit start resets the restart manager. Failures separated by days can therefore exhaust the cap and leave a container stopped until an explicit start or recreation.
  • Health checks report readiness/health; an unhealthy state alone does not trigger an Engine restart. scripts/dccd.sh is a deployment tool, not a watchdog: by default, it skips deployment when the repository is unchanged and some containers are running, even if another container has stopped.

Source references: Compose v2.39.4 getRestartPolicy, Moby v28.5.0 ShouldRestart, and explicit-start handling.

Key rules:

  • Images must always include an explicit registry prefix (e.g. docker.io/library/busybox, ghcr.io/gethomepage/homepage). Bare image names like busybox or user/image are not allowed — Docker's implicit docker.io default is not reliable across runtimes and Renovate cannot enforce the correct registry without it
  • Images are digest-pinned (@sha256:...) — Renovate manages updates via PRs
  • Every imported image must receive an upstream dependency maintenance review; pinning, a fresh tag, and runtime hardening do not establish dependency security.
  • When the same image is published on several registries, prefer the docker.io copy — see Image Selection: Registry Preference below.
  • Prefer the smallest, most hardened image variant available for a given version tag. When multiple variants are published, choose according to this priority order:
  • Hardened (e.g., -hardened, Chainguard distroless/static images, Docker Hub Hardened Images at dhi.io/<name>) — minimal attack surface, no shell, stripped of unnecessary OS components. Note that hardened variants are sometimes published under a different image name or registry rather than as a tag suffix on the official image (e.g., eclipse-mosquitto has a hardened build at dhi.io/eclipse-mosquitto). Always check the Docker Hub Hardened Images catalog and Chainguard for a hardened alternative before falling back to Alpine. Note: Images from dhi.io require Docker Hub authentication (docker login dhi.io with your Docker Hub username and a personal access token). Tag format is <name>:<version>-<os> (e.g., dhi.io/traefik:3.6.14-debian13). Runtime variants run as a nonroot user (UID 65532) and have no shell; health checks must use CMD format (not CMD-SHELL).
  • Alpine (e.g., 2.1.2-alpine) — musl-based, ~5 MB base layer, no unnecessary tools
  • Slim (e.g., 2.1.2-slim) — Debian-based with non-essential packages removed
  • Standard (e.g., 2.1.2) — only when no smaller or hardened variant exists, or when the application requires glibc/full Debian (e.g., for apt at runtime or native library dependencies)

When choosing a variant, verify it provides the required functionality (some Alpine builds omit optional compiled modules). Document any exception in a comment in the compose file.

If an image publishes additional non-standard variant suffixes (e.g., -openssl, -jdk, -bookworm, -ubi9) that are not covered by the priority order above, ask the user which variant is preferred before selecting one — the implications (library compatibility, licence, FIPS compliance, etc.) are context-dependent. - read_only: true with tmpfs mounts for writable paths - no-new-privileges on every container, no exceptions - cap_drop: ALL on every container — this is a hard security requirement. If a container needs a specific capability, declare cap_add with only the minimum required capability and add a comment on the container in the compose file explaining why the exception is necessary - Memory limits with env-var overrides for per-environment tuning - pids_limit on every container to prevent fork-bomb DoS - Health checks are mandatory — dccd.sh uses docker compose up --wait - Volumes mounted :ro wherever the container only reads - ./config volumes must always be mounted :ro — config files are git-tracked and must never be modified by a container at runtime. If a service needs to write config at runtime, copy the file from ./config to ./data in an init container and mount the ./data copy read-write (see the gatus pattern). Any exception requires explicit approval and a comment in the compose file explaining why

Image Selection: Upstream Dependency Maintenance

Security goal: every container we import has effective upstream maintenance of its language/framework/runtime dependencies and its base-image/OS packages. This covers application, database, parser, backup, init, and other sidecar images, plus imported build-stage/base images. Apply this review to adoption and image updates; the goal also applies to existing imports, but is not a claim that the fleet has already been audited.

Required Evidence

Record a compact, dated review in the consuming service's README or a linked review issue. Identify each image's upstream source, exact digest and target platform, review date, maintenance flag, confidence, evidence links, gaps, and adoption decision. Use dated links to concrete changes, tests, releases, and artifacts rather than only a project home page.

Evidence What to establish
Dependency tracking Renovate, Dependabot, or an equivalent documented effective process covers both language/runtime dependencies and base-image/OS dependencies. Inspect the actual coverage, not just whether a bot configuration exists.
Effective delivery Recent dependency updates were merged, tested, released, and incorporated into rebuilt images. Assess recency against the upstream support/release cadence; open bot PRs or a newly pushed tag alone are insufficient.
Exact artifact Inspect the selected digest/platform and its SBOM or package inventory with a current vulnerability scan. Verify affected and fixed versions actually present in the artifact; a fixed repository lockfile or source commit does not prove the published image contains the fix. Record scan date/tool/database freshness and distinguish raw entries from deduplicated findings.
Published policies Check maintenance/release commitments, supported versions, security reporting and disclosure channels, and end-of-life policies for the application, language runtime, and base distribution. Record missing information as a gap, not an invented policy.

Assign one maintenance flag with high, medium, or low confidence, explaining the confidence and linking the dated evidence:

Flag Meaning
well maintained Evidence demonstrates effective dependency tracking and delivery into supported artifacts across both dependency layers.
partial Maintenance exists, but coverage or delivery has evidenced gaps, such as unrebuilt images or neglected runtime/base updates.
unknown Evidence is insufficient or inaccessible; record what remains unverified.
unmaintained Positive evidence of abandonment, unsupported/EOL components without maintenance, or a persistently ineffective update process.

No bot does not mean unmaintained: a documented, effective manual or vendor process may qualify. Conversely, a configured bot is not proof of maintenance. The flag describes the evidence, not adoption approval; even a well-maintained project can publish a vulnerable artifact. Unresolved evidence requires review, not an automatic safe or unmaintained classification.

Adoption Gate and Remediation Ownership

Known, fixable HIGH/CRITICAL findings relevant to the selected artifact and deployment block adoption, image updates, and rollout, unless an explicit, reviewed, documented, narrowly justified exception is approved. Investigate package/version applicability, affected code paths, reachability, exploit prerequisites, and deployed configuration. Do not assume an internal network, authentication, or dropped capabilities makes a vulnerable dependency unreachable. Unknown reachability is not evidence of safety.

This is not a blanket global zero-CVE requirement. Document non-applicability, false positives, and residual risks with evidence; do not silently suppress findings. Runtime, restore, and hardening tests complement this assessment but cannot clear a dependency-security blocker.

An exception must identify the approving reviewer and date, exact image digest/platform and findings, applicability analysis, narrow deployment scope, mitigations, residual risk, upstream tracking, and an expiry/review date with exit criteria. It must not waive all future findings or silently carry across image changes.

Remediation belongs upstream. Seek upstream dependency updates and rebuilt releases, then verify the fixes in the exact replacement artifact and rerun relevant regressions. Do not take ownership of downstream patched dependency locks or forks as an adoption workaround. Local packaging/startup wrappers do not replace upstream dependency maintenance. Without an acceptable upstream artifact or approved exception, retain the integration as a blocked candidate, not a deployable service; gate any documented rollout commands.

Image Selection: Docker Hardened Images

DHI images at dhi.io/<image>:<tag> are preferred whenever the catalog lists the upstream — they provide hardened bases, signed SBOMs, and SLSA Build Level 3 provenance, and are free under the Community tier (Apache 2.0). This preference does not waive the exact-artifact maintenance and vulnerability review above. Pulls require docker login dhi.io; scripts/dccd.sh performs that login automatically when DHI credentials are present. Current DHI consumers: services/adguard/compose.yaml (redis), services/alloy/compose.yaml, services/traefik/compose.yaml.

Two caveats apply when adopting a DHI (or any new) image — the full procedure lives in the new-docker-app skill at .github/skills/new-docker-app/SKILL.md:

  1. Verify multi-arch before adoption. This homelab runs both svlnas (x86_64) and svlazext (arm64). Confirm the catalog page lists LINUX/AMD64 and LINUX/ARM64. DHI has previously republished tags as amd64-only manifest digests for short windows, breaking arm64 hosts (see commits b3cfb8c, dbd9f59, d3660cb).
  2. Initial commit must be tag-only — no @sha256 digest. Renovate runs on an amd64 GitHub runner and pins the digest on its next pass. Pinning manually from a stale snapshot can lock the deployment to an amd64-only digest if the registry has not yet republished as multi-arch. The tag itself stays pinned (e.g. 8.6.2-debian13) for reproducibility.

Image Selection: Registry Preference

Many images are published to several registries at once (e.g. docker.io/linuxserver/sonarr, lscr.io/linuxserver/sonarr, ghcr.io/linuxserver/sonarr). When the copies are functionally equivalent, pick the docker.io one. This is a tie-breaker, not an absolute rule — it never overrides the variant priority in Key rules above.

Why. Renovate only derives a release timestamp for a Docker dependency when the registry host is exactly https://index.docker.io. For that host it queries the Docker Hub API at hub.docker.com/v2/repositories/... and reads the tag_last_pushed field; for every other host it takes the generic code path against the OCI Registry V2 /tags/list endpoint, which by specification returns tag names and nothing else. Renovate documents this in its Docker datasource: "Only supported on Docker Hub: The release timestamp is determined from the tag_last_pushed field in the results." Note that tag_last_pushed comes from the separate hub.docker.com API — it is Hub-specific and is not part of the OCI registry specification.

This is a Renovate limitation keyed on the registry hostname, not an omission by the other registries. Docker almost certainly holds push timestamps for dhi.io too — the DHI web portal displays them — but Renovate never asks for them, because the host is not index.docker.io. Renovate also refuses to trust the OCI org.opencontainers.image.created label as a substitute, because the publisher controls that value and could forge it to bypass the age gate.

For updates subject to the 14-day policy, the soak is enforced either way: images from other registries fall back to the required pr-cooldown status check described in Contributing § Enforcing the soak without a trusted timestamp. Docker Hub mutable/rolling digest exceptions can flow sooner, as described in Contributing § Update timing policy. The argument for docker.io is therefore narrower than "gate versus no gate":

  • The native soak measures the actual publish time. pr-cooldown measures age from the PR branch HEAD committer date, which is a proxy and restarts when a new head commit is written, including a rebase that preserves the image target. The frozen-candidate cohort, expanded from Homepage to include Traefik, both Immich app images in their existing group, and Home Assistant's rolling :beta digest, stops ordinary automatic Renovate rewrites to existing candidate branches. Its exact five-package local rule leaves the 14-day gate and accrued HEAD age unchanged; manual rejection or refresh remains available for known-bad candidates or needed security fixes.
  • The fallback has more moving parts — a scheduled workflow, a required status check, and a branch-ruleset entry. ADR 0005 in DevSecNinja/.github calls pr-cooldown a load-bearing control.

Tradeoff. Docker Hub applies pull rate limits to anonymous and free-tier accounts; ghcr.io and lscr.io do not. scripts/dccd.sh re-pulls images on every deploy, so this is a real operational cost on a self-hosted box. Digest pinning keeps pulls cache-friendly, which is what makes the tradeoff acceptable — but it is a genuine tradeoff, and a service that is redeployed very frequently is a fair reason to stay on ghcr.io.

Exceptions that outrank this preference:

  • Hardened images. dhi.io is a separate registry from docker.io, and hardened variants rank first in the variant priority order, so DHI images cannot follow this preference.
  • Images published on only one registry.
  • A smaller or hardened variant that exists only on ghcr.io (or another non-Hub registry) — variant size and hardening win over registry choice.

Precedent. All LinuxServer images were moved from lscr.io to docker.io for exactly this reason (PR #641). The current manifest digest was verified to be identical across docker.io, lscr.io, and ghcr.io beforehand, confirming this was a registry change and not a switch to a different build lineage.

Healthchecks on Distroless Bases

Hardened bases vary in how minimal they are:

  • Debian-13 minimal variants (e.g. dhi.io/redis, dhi.io/traefik) include /bin/sh (dash) but no curl, wget, nc, or bash.
  • Binary-only variants (e.g. dhi.io/alloy) ship only the application binary — no shell, no probe utilities at all.

For Debian-13-minimal images, when the application itself does not expose an HTTP probe binary, verify the listener via /proc/net/tcp (with /proc/net/tcp6 as IPv6 fallback) using only sh:

healthcheck:
  test:
    - CMD
    - sh
    - -c
    - grep -q ':3039 ' /proc/net/tcp || grep -q ':3039 ' /proc/net/tcp6
  interval: 30s
  timeout: 5s
  retries: 3
  start_period: 10s

The port is the listening port in uppercase hex (e.g. 12345 decimal → 3039). Confirm /bin/sh is actually present (docker run --rm --entrypoint sh <image> -c true) before adopting this pattern.

For binary-only images where no in-container probe is possible, set healthcheck: { disable: true } with a comment explaining why. scripts/dccd.sh (via compose_up_wait_tolerant) treats "no healthcheck" as successful once the container reaches running, so the deploy still completes. Pair this with external monitoring (Gatus + Traefik logs + the application's self-metrics) for actual liveness signal. The canonical example is services/alloy/compose.yaml.

Volume Permissions: Init Container Pattern

Named Docker volumes and bind-mounted directories are created as root:root by Docker. A container with a hardcoded non-root user: and cap_drop: ALL has no CAP_CHOWN and cannot fix this at runtime — it will fail to write on first deploy.

Rule: Any service with both of the following requires a <app>-init container:

  1. user: "<UID>:<GID>" (explicit non-root)
  2. At least one writable volume (named volume or bind mount)

The init container runs as root, chowns the volume paths to the service's UID:GID, and exits before the main container starts. The main service declares depends_on: <app>-init: condition: service_completed_successfully.

Bind-mount directories (./data, ./backups) that are runtime-only (gitignored) must be included in the init container's chown command, even when the main container mounts a path inside them as :ro. A host-level chown (e.g. a TrueNAS dataset permission reset) can make those directories unreadable or untraversable. Init restores ownership on its declared paths; this does not guarantee recursive repair on every deploy. Open Archiver prepares directory roots only during routine startup and requires explicit stopped-writer repair for restored descendants.

Git-tracked ./config directories must NEVER be chowned or chmod'd by an init container. Doing so changes file ownership away from the deploy user and causes git pull to fail with error: unable to unlink old '...': Permission denied. Config files checked out by git are already world-readable (644 files, 755 directories), so any container user can read them without ownership changes. If a service needs to write config at runtime, copy the file from ./config to ./data in the init container and mount the ./data copy into the main container (see the gatus pattern).

# Pattern — copy this block, adjust container_name, UID:GID, command paths, and volumes
# IMPORTANT: only chown ./data (runtime) paths — NEVER chown ./config (git-tracked)
<app>-init:
  image: docker.io/library/busybox:1.37.0@sha256:1487d0af5f52b4ba31c7e465126ee2123fe3f2305d638e7827681e7cf6c83d5e
  container_name: <app>-init
  env_file:
    - path: .env          # Decrypted from secret.sops.env
      required: false
  restart: "no"
  network_mode: none
  mem_limit: 64m
  pids_limit: 50
  security_opt:
    - no-new-privileges=true
  cap_drop:
    - ALL
  cap_add:
    - CHOWN    # Required to chown volume paths
  read_only: true
  command:
    - "sh"
    - "-c"
    - |-
      chown -Rv <UID>:<GID> /data
  volumes:
    - ./data:/data

For services that only chown runtime-only paths (named Docker volumes, ./data/), the chmod 775/664 step and FOWNER/DAC_OVERRIDE capabilities can be omitted — only CHOWN is needed. Docker creates named volumes and ./data/ directories as root:root 755, so UID 0 is always the owner and can traverse them without DAC_OVERRIDE.

Exception — external bind-mount paths with non-root ownership: If the bind-mount source is a host directory owned by a non-root user (e.g., a TrueNAS dataset with truenas_admin:truenas_admin 770), UID 0 inside the container matches neither owner nor group and has no permissions. Busybox chown -Rv opens the directory before chowning it, which fails without DAC_OVERRIDE. Add DAC_OVERRIDE to any init container that chowns such a path.

Exceptions — images that manage their own permissions:

  • s6-overlay images (LinuxServer, tiredofit/db-backup) start as root and chown their own directories during their own init phase, so they generally do not need a dedicated external init container. Karakeep is an exception: its existing app init container pre-owns the ./backups/db-backup child because the backup image resets the read-write parent mount root to root ownership.
  • nfrastack/db-backup manages backup output ownership internally. Verify the selected image's supported identity settings; Open Archiver uses DBBACKUP_USER/DBBACKUP_GROUP with a pre-init account check, not the ignored legacy USER_DBBACKUP variables.
  • Database images (postgres, MongoDB) that start as root can initialise their own data-directory ownership. Direct non-root deployments, such as Open Archiver's PostgreSQL and Valkey containers, still require pre-owned runtime paths.

Services using this pattern:

Service Init container Volumes chown'd
_bootstrap content-init /mnt/archive-pool/content (full tree: mkdir + chown :3200 + setgid 2775)
adguard adguard-init ./data/work, ./data/conf
adguard adguard-unbound-init ./data/unbound (generated template output)
alloy alloy-init ./data (WAL + queue)
changedetection changedetection-init docker.io/library/busybox:1.38.0; chowns ./data mounted at /datastore → PUID/PGID 3131:3131, then applies u=rwX,g=,o=
dawarich dawarich-init Validates required decrypted values; chowns ./data/public, ./data/storage, ./data/watched, ./data/app-tmp, ./data/sidekiq-tmp → 3128:3128; ./data/redis → 999:999
dozzle dozzle-init ./data
frigate frigate-init Seeds ./config/config.yml → ./data/config/ on first deploy (cp -n)
gatus gatus-init Copies ./config/config.yaml → ./data/sidecar-config/ (config mounted :ro)
home-assistant home-assistant-init Seeds ./config/configuration.yaml → ./data/config/ on first deploy (cp -n)
homepage (removed) None — config is git-tracked and read-only; no init needed
immich immich-init /mnt/archive-pool/private/photos/immich (+ DAC_OVERRIDE), ./data/model-cache
karakeep karakeep-init docker.io/library/busybox:1.38.0; creates and chowns ./backups/db-backup, then chowns ./data/karakeep and ./data/meilisearch → 3130:3130
matter-server matter-server-init ./data
memos memos-init ./data → PUID/PGID 3129:3129
metube metube-init ./data/state
mosquitto mosquitto-init ./data/data, ./data/log
open-archiver open-archiver-init docker.io/library/busybox:1.38.0; chowns directory roots ./data/archive, ./data/scratch, ./data/meilisearch → UID/GID 3132:3132; ./data/postgres → 70:70; ./data/valkey → 999:1000; applies u=rwX,g=,o= non-recursively by default and mode 0700 to ./data
openclaw openclaw-init ./data (chown to 3127:3127)
outline outline-init ./data/data (chown to UID 1000 — image-internal node user)
spottarr spottarr-chown ./data
traefik traefik-init ./data/acme
traefik-forward-auth traefik-forward-auth-init ./data
wmbusmeters wmbusmeters-init ./data/logs, ./data/state

changedetection-init validates that DOMAINNAME is populated and is neither CHANGE_ME nor GENERATE before touching permissions. It runs without a network and retains only CHOWN, FOWNER, and DAC_OVERRIDE for ownership, permission tightening, and traversal of private paths on repeat deployments. It only changes ./data, never ./config. The app uses a direct user: identity rather than PUID/PGID environment variables.

dawarich-init uses docker.io/library/busybox:1.38.0. Its container mount targets are /public, /storage, /watched, /app-tmp, /sidekiq-tmp, and /redis; it never changes ownership under ./config. Before changing permissions, it rejects any required decrypted value that is empty or equals CHANGE_ME. It retains only CHOWN, FOWNER, and DAC_OVERRIDE; DAC_OVERRIDE is required to traverse mode 770 runtime paths on repeat deployments. PostGIS and Redis depend on successful init completion, preventing a placeholder database password from initializing PostGIS; the application and worker start only after their backends are healthy.

karakeep-init uses docker.io/library/busybox:1.38.0 to validate the required domain, NextAuth secret, Meilisearch master key, mobile proxy token, and backup passphrase before assigning ./data/karakeep, ./data/meilisearch, and the newly created ./backups/db-backup child to 3130:3130. It mounts ./backups at /backups so it can prepare the child without changing the parent mount root. It never changes ownership under ./config.

open-archiver-init validates DOMAINNAME and the eight generated secret variables before changing permissions, rejecting empty values, GENERATE, and CHANGE_ME. Both encryption keys must encode exactly 32 bytes as 64 hexadecimal characters. The init runs without networking and retains only CHOWN, FOWNER, and DAC_OVERRIDE for private runtime paths. It mounts only ./data; tracked ./config is never chowned or written. Runtime identities use direct user: settings, not PUID/PGID environment variables. Routine init explicitly sets OPEN_ARCHIVER_REPAIR_PERMISSIONS=false and uses umask 077: it creates and sets ownership/modes on the five directory roots, plus mode 0700 on /data, without walking the archive or database contents. Restored trees require an operator-only run with OPEN_ARCHIVER_REPAIR_PERMISSIONS=true, after stopping all writers and preserving the existing state. Only the five named runtime trees are repaired recursively; see the recovery procedure. The separate open-archiver-migrate one-shot waits for healthy PostgreSQL and must exit successfully before the app starts.


Exception — hardened direct-non-root override:

The upstream Dawarich image normally starts as root and uses its PUID/PGID entrypoint to prepare data before dropping privileges. This stack instead prepares every writable path with dawarich-init, then runs both the Rails application and Sidekiq worker directly as 3128:3128. The web container explicitly uses web-entrypoint.sh with bin/rails server -p 3000 -b ::; the worker explicitly uses sidekiq-entrypoint.sh. Both containers use a read-only root filesystem, drop all capabilities, and set no-new-privileges=true. This bypasses the upstream root-start permission model while preserving its role-specific startup scripts and required writable paths.

The Rails health check sends X-Forwarded-Proto: https with its internal HTTP request. Because APPLICATION_PROTOCOL=https enables Rails force_ssl, omitting the header redirects the probe instead of returning the expected health response. The Sidekiq health check invokes the image's Ruby interpreter to inspect /proc/1/cmdline and verify that PID 1 contains sidekiq. The image does not include pgrep/procps, so the check does not depend on those utilities.

Karakeep follows a similar direct-process pattern. The published all-in-one image normally starts its web and worker processes as root under s6-overlay. This deployment sets USING_LEGACY_SEPARATE_CONTAINERS=true and launches the split web and worker processes directly as 3130:3130; the web command applies database migrations before serving traffic. Meilisearch also runs as 3130:3130. The web, worker, and Meilisearch containers use read-only root filesystems, drop all capabilities, and enforce resource limits. The root-start s6 backup sidecar omits a read-only root filesystem and uses the restricted capability set documented in Karakeep Network and Access Model.

Karakeep enables browser crawling with CRAWLER_HEADLESS_BROWSER=true and BROWSER_WEB_URL=http://karakeep-chrome:9222. changedetection.io defaults to html_webdriver with PLAYWRIGHT_DRIVER_URL=http://172.30.100.22:9222. Both applications wait for healthy Chrome; Chrome waits for its healthy proxy.

Both browser definitions use dhi.io/playwright:1.63.0-debian13 and are active by default, without profiles, with a mandatory Chromium 153.0.8010.52 minimum and no active version exception. The browser runs as 65532:65532 with a read-only root, init: true, dropped capabilities, no-new-privileges, tmpfs scratch, a fixed 512-task cap (processes and threads), and ${CHROME_MEM_LIMIT:-2048m}. Its Squid proxy adds ${BROWSER_PROXY_MEM_LIMIT:-256m} and a separate 100-PID limit. The browser cap provides initial headroom, not a production capacity guarantee. Changes require a reviewed Compose edit; no environment override is supported. The shared Node launcher checks Chromium's version before launch and provides an IPv4 TCP CDP relay on port 9222 to loopback port 9223. See Browser Egress Policy.

Exceptions — s6-overlay and root-start containers:

Some images cannot use read_only: true or user: because their init system (s6-overlay) requires a writable root filesystem and starts as root before dropping privileges internally. cap_drop: ALL is still required for these images — only the specific capabilities that s6-overlay needs are re-added via cap_add. Each such container must include a comment block in the compose file explaining the deviation. This applies to:

  • LinuxServer images (e.g., unifi-network-application, plex) — use PUID/PGID environment variables for internal privilege dropping; omit user: and read_only. Add back CHOWN, SETUID, SETGID, and SETPCAP via cap_add.
  • LinuxServer socket-proxy — runs as root by design to proxy the Docker socket. Does not support custom users, mods, or scripts. Omit cap_drop: ALL; no-new-privileges and read_only are still applied.
  • tiredofit/db-backup — uses USER_DBBACKUP/GROUP_DBBACKUP for internal privilege dropping; omit user: and read_only.
  • nfrastack/db-backup — starts as root with a writable root filesystem, then selects an internal backup identity. Verify the image-supported variables rather than assuming legacy settings work; Open Archiver uses DBBACKUP_USER/DBBACKUP_GROUP and a pre-init account check.
  • mvance/unbound — starts as root and drops privileges to the _unbound user internally; its startup script generates unbound.conf and creates subdirectories at runtime, so omit user: and read_only.
  • meeb/tubesync — uses s6-overlay: tubesync-config-init sets the app user's PUID:PGID and prepares app-owned /config directories (mode 0755) and /run/app (mode 0700); service startup scripts finish root-level setup before dropping privileges to app. Omit user: and read_only:. Retain cap_drop: ALL and no-new-privileges; add back CHOWN, DAC_OVERRIDE, FOWNER, SETUID, SETGID, and SETPCAP via cap_add. FOWNER permits chmod after chown; DAC_OVERRIDE lets root create /config/state/hat and access /run/app despite their app ownership and restrictive modes.
  • ghcr.io/home-assistant/home-assistant — uses s6-overlay (confirmed by s6-rc log lines). Omit user: and read_only:. Add back CHOWN, SETUID, SETGID, SETPCAP via cap_add (standard s6-overlay set). Also add NET_RAW — required by HA's built-in DHCP watcher integration, which opens raw AF_PACKET sockets to track devices; without it HA logs [Errno 1] Operation not permitted at startup and the DHCP integration stops working. No TrueNAS service account or init container is required — s6-overlay manages /config ownership internally.
  • ghcr.io/esphome/esphome — compiles C++ firmware at runtime using platformio, downloading platform packages and managing build artifacts across /config/.esphome/. Requires extensive filesystem writes; omit user: and read_only:. cap_drop: ALL is applied; no additional capabilities are needed.
  • ghcr.io/blakeblackshear/frigate — runs as root; manages its own internal processes (nginx, go2rtc, detector workers) and requires access to hardware devices (GPU, optional Coral TPU). Omit user: and read_only:. cap_drop: ALL is applied; no additional capabilities are needed.
  • ghcr.io/bitwarden/lite — bundles nginx + multiple .NET services (admin, api, identity, icons, notifications, sso, scim, events) under supervisord. The entrypoint runs as root to generate the IdentityServer PFX certificate, create the in-container bitwarden user from PUID/PGID, chown /etc/bitwarden and the nginx/supervisor paths, then exec su-exec into supervisord as the runtime user. Omit user: and read_only:. Add CHOWN, DAC_OVERRIDE, FOWNER, SETGID, SETUID, and SETPCAP via cap_add (chown chain + privilege drop).

Each exception is documented with a comment block in the compose file explaining why the deviation is necessary.

Minimum cap_add for LinuxServer/s6-overlay images:

Capability Why it is needed
CHOWN s6-overlay chowns mounted volumes (e.g., /config) to PUID:PGID at startup
SETUID s6-overlay calls setuid() to drop from root to PUID
SETGID s6-overlay calls setgid() to drop from root to PGID
SETPCAP s6-overlay clears the bounding capability set before exec-ing the application daemon

All other default Docker capabilities (NET_RAW, NET_BIND_SERVICE, MKNOD, AUDIT_WRITE, SYS_CHROOT, FSETID, FOWNER, DAC_OVERRIDE, KILL) remain dropped unless an explicitly documented app-specific requirement justifies adding them back.

Pitfalls specific to s6-overlay images:

  • read_only: true silently breaks PUID/PGID. s6-overlay writes the UID/GID entries to /etc/passwd and /etc/group during startup before dropping privileges. With read_only: true those writes fail silently and the container continues running as the image default (UID 911 for LinuxServer Plex), ignoring PUID/PGID entirely. Always omit read_only for s6-overlay images when working with subprcesses in the container such as Plex that need access to volumes.

  • group_add does not grant supplementary groups to the application process. group_add adds GIDs to the credentials of PID 1 (s6-overlay, which runs as root). When s6-overlay drops privileges to run the application it re-initialises the process's supplementary groups from /etc/group inside the container — where the host-only GID does not exist. The result is that the application process has no membership in the added group. To grant an s6-overlay image membership in a host GID, set that GID as PGID (primary group) or ensure the image's own group-setup mechanism adds it. The correct approach for LinuxServer images is to set the desired GID via the PGID env var; s6-overlay will then create the /etc/group entry and the application process will run with that GID.

Config Template Substitution: Envsubst Init Containers

Some services need secrets or environment-specific values (domain names, API keys) injected into their configuration files at deploy time. Since these config files are committed to Git as templates with ${VAR} placeholders, a separate init container processes them before the main service starts.

Pattern: An <app>-init container mounts ./config as /templates:ro, runs envsubst.sh to replace ${VAR} placeholders with values from secret.sops.env, and writes the processed output to ./data/. The main container then mounts the processed file from data/ as :ro.

This keeps secrets out of Git (the template only contains placeholder names) while the processed config with real values lives in data/ which is gitignored.

Services using this pattern:

Service Init container Template → Output
adguard (unbound) adguard-unbound-init config/unbound/*.conf → data/unbound/*.conf
traefik-forward-auth traefik-forward-auth-init config/config.yaml → data/config.yaml

AdGuard Unbound Startup Wrapper

The adguard-unbound resolver overrides its entrypoint with /sbin/tini -- /bin/sh -ec. At each container start, the wrapper uses sed -i on the ephemeral /usr/local/unbound/unbound.conf to re-enable the conf.d and zones.d include-toplevel directives, validates that both exact directives are present, then runs exec /entrypoint to preserve the image's upstream initialization and privilege drop. DAC_OVERRIDE is added only because sed must rewrite this config within the image filesystem.

No generated base config is bind-mounted. This avoids the first-deploy bind-mount race and ensures each Renovate image update starts with that image's current defaults; the wrapper changes only the two required include directives instead of masking new defaults with an older persistent copy. The healthcheck requires unbound-checkconf -o control-enable to return yes, proving that the conf.d include made the mounted remote-control config effective.

Networking: DNS Resolution

The chosen architecture separates host bootstrap DNS from everyday LAN client DNS. Public resolution for TrueNAS boot, recovery, and image pulls must not depend on a DNS container hosted on that same NAS.

flowchart LR
    Host["TrueNAS host"] --> Gateway["UniFi gateway"]
    Gateway -->|"public queries"| Public["Public upstream resolver"]
    Gateway -->|"known infrastructure names"| Records["Small manual record set"]
    Client["LAN clients via VLAN DHCP"] --> AdGuard["NAS-local AdGuard"]
    AdGuard --> Unbound["Co-located Unbound<br/>local records + recursion"]
    Unbound -->|"public queries"| Hierarchy["Public DNS hierarchy"]

The gateway independently serves a small infrastructure record set alongside public forwarding. This covers known host/container dependencies, not the full client DNS namespace or filtering policy; it is not complete client DNS redundancy. There is no automatic record synchronisation with Unbound. See DNS configuration and record maintenance for the chosen settings and record owners.

Why the Paths Are Separate

The Azure DNS VM became unavailable after its credits were exhausted, leaving an unreachable remote secondary that contributed to long DNS delays. Separately, uncached NAS queries through the gateway received immediate REFUSED responses with EDE23 Network Error when the gateway used NAS-local AdGuard as its upstream, even though router-originated, cached, and direct AdGuard queries worked. Switching the gateway upstream to a public resolver worked; the exact forwarding fault was never proven, and this split avoids both that path and NAS bootstrap coupling.

DNS server lists are not guaranteed ordered primary/backup failover. The inspected gateway runtime used dnsmasq's all-servers setting to query all eligible upstreams; this is an observation of that configuration, not every UniFi firmware version. DHCP clients select resolvers according to their own OS behaviour, not that gateway flag. An unreachable resolver can delay or fail queries; mixing public and private resolvers in a client list can bypass filtering and return NXDOMAIN or NODATA for private names.

Docker DNS Is Not Independent Split DNS

The tracked Compose files have no dns, dns_search, or dns_opt overrides. On custom networks, Docker's embedded resolver handles Docker service names; other lookups default to the host's upstream resolvers unless an out-of-repository daemon/runtime override changes this. Verify each affected running container's effective resolver after host DNS changes rather than assuming immediate adoption.

Docker database/cache names and Traefik's Docker backend/forward-auth names do not require private-domain DNS. IP-addressed Gatus checks with a Host header also avoid hostname resolution for their target, but this does not prove that HTTP integrations or alerts work. Homepage server-side monitors, Gatus alert webhooks, and deployment status reporting still need the infrastructure records. Browser links resolve on the client, not in the Homepage container.

Desired State: Two Independent Local Resolvers

The second local resolver is not deployed. It should run AdGuard with its own resolving backend on separate always-on hardware or a suitable router-capable host, with independent storage and power where practical. Another VM/container on this TrueNAS, or a resolver forwarding back to its AdGuard/Unbound stack, would not provide the required failure-domain independence.

Either local resolver must provide public resolution and identical internal answers while the other is unavailable. Client VLANs should then advertise both local resolver addresses, with matching filtering/client policies and internal records. Host bootstrap must remain independent of NAS-local DNS. See future configuration and the validation gate before advertising a second resolver.

Networking: Per-Service Isolation

Each service gets its own frontend network (e.g., echo-server-frontend, homepage-frontend). Traefik joins approved, activated services' frontend networks individually. A blocked candidate such as Open Archiver does not add its network dependency to shared Traefik before a separately reviewed activation change.

Why not a single shared traefik-public network?

Network-level isolation. With per-service networks, containers cannot communicate with each other — only with Traefik. A shared network would let any compromised container reach every other service. The trade-off is that adding a new service requires adding its network to Traefik's compose file.

Services that need Docker API access get a dedicated internal backend network with a socket proxy (e.g., homepage-backend). The same pattern applies to databases and other backing services — they sit on an internal backend network with internal: true, preventing external routing and ensuring only the application container can reach them.

Exception: arr-stack-backend

The arr stack (Radarr, Sonarr, Bazarr, Lidarr, Prowlarr, qBittorrent, SABnzbd, Spottarr) shares a single arr-stack-backend internal bridge network so the apps can communicate directly for API calls (e.g., Prowlarr pushing indexer results to Sonarr). This network is created by the _bootstrap service and referenced as external: true by each arr app. All internet traffic still exits through each app's dedicated VLAN 70 macvlan network — the backend bridge is internal: true and carries no internet route.

Exception: dawarich-backend

Dawarich and Alloy share the external dawarich-backend internal bridge: Dawarich uses it for PostGIS, Redis, backup, and exporter traffic, while Alloy uses it only to scrape dawarich-db-exporter. The _bootstrap stack creates the network before Alloy and Dawarich because its directory sorts first in a full dccd.sh deployment. This ownership and ordering prevent Alloy from failing on a fresh deployment while keeping the network internal. Redis remains backend-only. The database backup container also stays backend-only and sets ENABLE_NOTIFICATIONS=FALSE; the remaining NOTIFICATIONS_EMAIL_* variables configure only Dawarich application email.

dawarich-db-backup uses the maintained docker.io/nfrastack/db-backup:4.9.2 compatibility release. It runs backup-now in one-shot MODE=MANUAL mode on every full dccd.sh deployment. The backup job depends on the healthy dawarich Rails service rather than only PostgreSQL, so startup migrations, data migrations, and seeding finish before the dump can run. DEFAULT_COMPRESSION=ZSTD, DEFAULT_CHECKSUM=SHA1, DEFAULT_ENCRYPT=TRUE, and DEFAULT_ENCRYPT_PASSPHRASE=${DB_ENC_PASSPHRASE} produce a ZSTD-compressed, GPG-encrypted PostgreSQL dump with a SHA1 sidecar. DEFAULT_CLEANUP_TIME=2880 retains the artifacts for 2880 minutes. Runtime testing successfully decrypted the dump and restored it into a fresh PostgreSQL database.

Version 5.0.0 was intentionally not selected because runtime restore validation failed with an invalid bigint conversion. Version 4.9.2 preserves the proven v4 workflow while using the maintained nfrastack image and repository.

Exception: iot-backend

The IoT stack (Home Assistant, Mosquitto, ESPHome, Frigate, wmbusmeters) shares a single iot-backend internal bridge network so the services can communicate directly. For example, wmbusmeters publishes MQTT messages to Mosquitto, Home Assistant subscribes to MQTT topics, and Frigate sends events via MQTT. This network is created by the _bootstrap service and referenced as external: true by each IoT app. The backend bridge is internal: true and carries no internet route. Matter Server is excluded — it uses network_mode: host for mDNS device discovery and Thread border router communication.

Browser Egress Policy (Public Websites Only)

Current decision: browser execution enabled with the version floor enforced. Both apps use dhi.io/playwright:1.63.0-debian13 and a dedicated Squid proxy, active by default without profiles. Normal Compose deployment recreates the changed browser definitions. Initial DHI adoption remains tag-only under the repository rule; Renovate will add a digest pin. No custom image publication is needed.

A fresh pull of the unchanged DHI tag to a disposable rootful Podman VM reported Chromium 153.0.8010.52 from the actual binary, meeting the configured fixed-version floor.

Both app-specific browser command sets then returned valid CDP discovery without a version override, with the fixed 512-task cap, UID 65532, read-only root and 2 GiB memory. Both stopped cleanly. These isolated checks used no network access; they do not revalidate complete app/proxy workflows. The old cached 153.0.8010.47 build was rejected with exit code 1 without an override.

Verified image reference Digest
Registry image index sha256:41e897d34bf35331d360bf9a63ecb2aa2d29e38139ab1640897947db674a66c4
Linux amd64 manifest sha256:b7f5663a4c94286c6181b031b9625d02f25066cd0004b1b3b56688e48da3b8b9

These are verification references, not Compose digest pins. This confirms the binary version only, not a production deployment or full application/proxy revalidation with this build.

The previously inspected DHI image contained Chromium 153.0.8010.47-2~deb13u1. Its confirmed Linux issue was the high-severity V8 CVE-2026-93377. At that review, the CNA description of CVE-2026-93372 was Android-specific; critical Linux applicability was not confirmed. The known-exploited CVE-2026-85046 finding concerned the former Chromium 151 image, not the September 17 issue. CVE-2026-93377 was not listed in CISA KEV when checked; absence from that catalog does not establish safety.

services/shared/config/browser/launch.mjs retains Google's September 17 fixed-version floor, 153.0.8010.52. Neither browser definition sets BROWSER_ALLOW_UNPATCHED_VERSION, so the floor is mandatory for both stacks: older cached 153.0.8010.47 images fail closed. The launcher's exact-version exception compatibility remains in code but is dormant; it is not an active deployment exception.

Meeting the version floor is not a complete security assessment

The verified binary meets the fixed-version floor for the reviewed V8 issue. This does not establish that all CVEs are fixed, validate production TrueNAS, or eliminate browser risk. Squid and container hardening do not replace browser security updates.

Existing-app deployment: with the aliases already sourced, run dccd-app karakeep, dccd-app changedetection, then dccd-all. Verify the actual Chromium version in both containers is at least 153.0.8010.52, with no exception warning, and verify service health and browser workflows. The unchanged tag alone does not prove which build is running. Do not restore the obsolete override to start an older cached image.

The launcher mounts read-only at /opt/browser/launch.mjs and runs as 65532:65532. After its version check passes, it launches Chromium with --no-sandbox on loopback port 9223 and relays internal port 9222 to it. A Node health check validates /json/version and its WebSocket URL.

changedetection.io and Karakeep each use a dedicated <app>-browser-proxy. Both use the operator-approved Canonical image docker.io/ubuntu/squid:7.2-26.04_edge@sha256:77585a8f42b57872b96c6c10de8d0199642256b9390beeb6e99afa5de8a71e1c and mount services/shared/config/browser/squid.conf read-only at /etc/squid/squid.conf.

Chrome joins only its app's internal browser network. Only the Squid proxy joins both that control network and the existing browser-egress bridge. Chrome has no direct internet route, published port, or frontend membership. The proxy also has no published port or frontend membership. App clients retain their existing networks; this policy does not filter non-browser Basic HTTP paths or optional AI calls.

Control Configuration
Browser proxy --proxy-server=http://<app>-browser-proxy:3128
Implicit loopback bypass Removed with --proxy-bypass-list=<-loopback>
Non-proxied UDP paths --disable-quic and --force-webrtc-ip-handling-policy=disable_non_proxied_udp
Allowed ports 80 and 443; CONNECT only to 443
Destinations Eligible public IPv4/native IPv6, excluding configured private, loopback, link-local, reserved/special-use, and transition ranges
IPv6 transitions IPv4-mapped addresses checked against IPv4 policy; NAT64 and 6to4 excluded by eligibility/deny rules
Request checks host_verify_strict on; cache-manager access denied
Proxy runtime Explicit 65534:65534, read-only root, cap_drop: ALL, no-new-privileges, 100 PIDs, ${BROWSER_PROXY_MEM_LIMIT:-256m}, /tmp tmpfs only
State and logs No disk cache or persistent state; access/store logs disabled; diagnostics to stderr; pinger disabled
Readiness Bundled Perl sends a loopback-destination HTTP request and requires 403, not just an open socket

Chrome depends on a healthy proxy, and both applications depend on healthy Chrome. The denial probe tests one policy path, not every destination rule or browser workflow.

Both browser and proxy declare config.watch=../shared/config/browser and config.sha256=${CONFIG_HASH:-}. The config-change mechanism recreates enabled consumers on their next dccd deployment when the shared launcher or policy changes. There are no new host firewall rules, host dependencies, secrets, databases, or persistent-volume init containers.

Public websites only is an intentional restriction

Internal-site crawling/monitoring is unsupported. Do not add DIRECT fallbacks, proxy bypass rules, or additional Chrome egress networks to work around denied requests. Unreviewed per-watch proxy overrides are outside this policy's assurance and must not be used to bypass it.

The policy targets browser HTTP(S) redirect/subresource SSRF and removes Chrome's direct internet route. It is not full native-code compromise containment: the launcher uses --no-sandbox, and the internal control network necessarily includes app clients and the proxy. A compromised native browser can reach those peers. Neither the flags nor Squid make browser zero-days safe or guarantee protection for arbitrary application fetch/proxy overrides. Complete DNS-rebinding protection has not been established.

Standalone tests of the pinned Squid image passed public HTTP, certificate-validated HTTPS, and controlled denial tests for private literals/hostnames, IPv4-mapped IPv6, NAT64, 6to4, forbidden ports, and disallowed CONNECT. These tests exercised the controlled proxy without contacting real LAN services.

Browser Runtime Validation

The following historical checks predate both the fixed 512-task browser cap and removal of the version exception. The DHI browser/proxy configuration passed synthetic runtime checks on a rootful Podman test VM, using the cached DHI image with actual Chromium 153.0.8010.47. Both browser logs emitted the explicit unpatched-version exception warning.

  • Startup and isolation: both apps, browsers, and proxies were healthy. Browser UID 65532, proxy UID 65534, read-only browser/proxy roots, 100-PID limits, no published host ports, and internal-only Chrome network membership were confirmed.
  • changedetection.io: an API-created public-site browser watch persisted history and a PNG. Python Playwright passed public HTTPS, a JavaScript button click, and screenshot capture through its DHI browser.
  • Karakeep client: the worker image's actual Playwright 1.59.1 client resolved the configured browser URL to IPv4 and passed public HTTPS, JavaScript, and screenshot checks through its DHI browser.
  • Private-destination probes: both actual browsers blocked private image/subresource and redirect tests against a controlled canary. Positive-control clients reached the canary; browser probes added zero connections.
  • Restart recovery: both browsers were restarted without restarting the apps. Public fetches, reconnection, and the changedetection watch succeeded again; the canary connection count remained unchanged.

These checks did not validate production TrueNAS, real Entra/SSO or TLS termination, the full Karakeep signup/bookmark job flow, or changedetection.io's browser-step UI/visual selector. Those checks remain pending. The historical file-state restore evidence remains separate. These checks used the older, unpatched image, not the newly verified 153.0.8010.52 build, and do not close the DNS-rebinding assurance gap or prove native-code compromise containment. See the test-runtime requirement.

changedetection.io Network and Access Model

Browser fetching and both browser/proxy services are enabled by default in Compose; no profile activation is required.

Network Members and access
changedetection-browser App, dedicated Chrome, and browser proxy; internal, IPv4-only CDP and HTTP proxy traffic
changedetection-browser-egress Browser proxy only; outbound connections to permitted public websites
changedetection-frontend App and Traefik; HTTPS UI/API routing and app outbound traffic

Traefik routes https://changedetection.${DOMAINNAME} to the app's fixed internal port 5000 through chain-auth@file. The entire UI and API use this SSO-protected router: there are no bypass routers or Gatus integration. Neither app nor Chrome publishes a host port. Chrome has no Traefik labels and does not join the frontend network.

flowchart LR
    User[UI or API client] --> Proxy[Traefik: chain-auth]
    Proxy -->|frontend: port 5000| App[changedetection]
    App -->|internal browser network: CDP| Chrome[changedetection-chrome]
    Chrome -->|HTTP proxy| Squid[changedetection-browser-proxy]
    Squid --> Egress[Dedicated egress bridge]
    Egress --> Web[Public websites only]

changedetection-chrome uses the same DHI tag as Karakeep, but uses a separate container/network stack. Neither browser nor proxy instance is shared between apps. The browser's configured identity is 65532:65532; the proxy remains 65534:65534.

Browser control reservation Compose value
Subnet 172.30.100.16/29
Dynamic allocation range 172.30.100.16/30
Chrome ipv4_address 172.30.100.22 (outside the dynamic range)
App PLAYWRIGHT_DRIVER_URL http://172.30.100.22:9222; HTTP CDP discovery

The reservation provides an IP-based HTTP CDP endpoint, avoiding Docker hostname Host headers and a hardcoded transient WebSocket ID. HTTP discovery obtains the current endpoint on each connection. IPv4-only control matches the shared launcher's relay. Keep the subnet/range, Chrome address, and app driver URL coordinated.

DEFAULT_FETCH_BACKEND=html_webdriver selects browser fetching with ${FETCH_WORKERS:-2} adjustable workers. Startup waits for successful init and healthy Chrome. Existing Basic HTTP watch settings are not automatically migrated; explicitly select the browser fetcher where wanted. DHI browser-step and visual-selector integration remain pending validation. The public-website-only policy and residual risks apply to this stack; arbitrary per-watch proxy overrides are not covered.

Operational setup and file-state backup limitations are documented in the changedetection.io service guide.

Karakeep Network and Access Model

Karakeep uses karakeep-frontend for Traefik access to the web container and outbound worker crawling and optional AI calls. The internal karakeep-backend network carries communication between the web, worker, and Meilisearch containers. Meilisearch is backend-only; only the web container is exposed through Traefik.

CRAWLER_HEADLESS_BROWSER=true and BROWSER_WEB_URL=http://karakeep-chrome:9222 enable browser crawling. The web service depends on healthy Chrome; Chrome and Squid have no profiles.

Network Members and access
karakeep-backend Web, workers, and Meilisearch; internal application and search traffic
karakeep-browser Web, workers, Chrome, and browser proxy; internal CDP and HTTP proxy traffic
karakeep-browser-egress Browser proxy only; outbound connections to permitted public websites
karakeep-frontend Web and workers; Traefik access and their outbound traffic

The karakeep-browser control network is IPv4-only, matching the shared launcher's CDP relay. IPv4-only control is not an SSRF mitigation. The browser definition uses only this internal network; the separate karakeep-browser-egress bridge belongs only to the filtering proxy.

Chrome does not join the frontend or backend networks and has no published ports or Traefik labels. Over the internal browser network it shares CDP and HTTP proxy reachability with Karakeep web, workers, and the proxy. It cannot route directly to the internet.

Residual browser and SSRF risk

Chrome processes attacker-controlled pages using the shared launcher's --no-sandbox option. This is an explicit residual browser-engine risk, not a safe mode. Non-root execution, container hardening, resource limits, and network separation reduce blast radius; they do not eliminate browser exploitation. A native-code compromise can still reach app/proxy peers on the shared internal control network.

Karakeep validates HTTP(S) URLs, resolved A/AAAA addresses, redirects, and browser subrequests against private and reserved ranges. This deployment does not configure internal hostname allowlists. The plain HTTP crawler pins its validated DNS result. The browser HTTP(S) path adds Squid checks rather than relying only on app-side validation before browser resolution. The public-website-only policy targets redirect/subresource SSRF without claiming complete DNS-rebinding or native-code compromise containment. Plain HTTP crawling and optional AI calls retain their existing app egress, outside this browser proxy. Review the upstream security considerations and reviewed source revision before changing these controls. Browser configuration and the mandatory version floor are documented in the Karakeep service documentation.

The one-shot karakeep-db-backup sidecar is network-isolated. It starts with the tiredofit image's s6 supervisor as root, then maps USER_DBBACKUP=3130 and GROUP_DBBACKUP=3130 to drop the backup process to Karakeep's app-owned identity. It reads db.db from a read-only mount. The sidecar mounts ./backups at /backup-data and writes to the pre-owned /backup-data/db-backup child because the image resets read-write mount roots to root ownership; host output remains ./backups/db-backup.

The sidecar drops all capabilities and adds only CHOWN, DAC_OVERRIDE, FOWNER, SETGID, SETUID, and SETPCAP. Reduced-capability runtime testing proved DAC_OVERRIDE necessary for s6 to create root-owned runtime paths before dropping privileges. The remaining capabilities are the documented s6 path-preparation and privilege-drop set.

The standard karakeep-rtr web router applies chain-auth@file, and Karakeep retains its own local user authentication behind that middleware. All unspecified web and API requests use this router.

The higher-priority karakeep-mobile-rtr permits the official mobile app to bypass interactive Forward Auth only when the request supplies X-Karakeep-Proxy-Token: ${KARAKEEP_MOBILE_PROXY_TOKEN} and its method and path are in this exact envelope:

Method Path
GET /api/health
GET /api/version
GET /api/assets/{single-path-segment}
POST /api/assets
GET /api/trpc/{allowlisted-query-or-query-batch}
POST /api/trpc/{allowlisted-mutation-or-mutation-batch}

The GET tRPC allowlist contains config.clientConfig; bookmarks.getBookmarks, bookmarks.getBookmark, bookmarks.searchBookmarks, bookmarks.getReadingProgress; lists.list, lists.stats, lists.get, lists.getListsOfBookmark; tags.list, tags.get; highlights.getAll, highlights.getForBookmark; and users.whoami, users.settings, users.stats.

The POST tRPC allowlist contains apiKeys.exchange, apiKeys.validate, apiKeys.revoke; bookmarks.createBookmark, bookmarks.updateBookmark, bookmarks.deleteBookmark, bookmarks.updateTags, bookmarks.summarizeBookmark, bookmarks.updateReadingProgress; lists.create, lists.edit, lists.delete, lists.addToList, lists.removeFromList, lists.leaveList; tags.create, tags.update, tags.delete; highlights.create, highlights.update, highlights.delete; and users.updateSettings, users.deleteAccount. Each tRPC request may name one allowlisted procedure or a comma-separated batch made entirely from the corresponding method's list.

After a match, the router applies chain-no-auth@file and karakeep-mobile-strip; the latter removes X-Karakeep-Proxy-Token before the request reaches Karakeep. A missing or incorrect token, a different method, or any unlisted path falls through to karakeep-rtr and SSO.

Karakeep's discovery/version/config and apiKeys.exchange and apiKeys.validate procedures are public upstream, but the mobile router token-gates them at Traefik. Assets and other user-data routes still require a Karakeep API key or session and the applicable scope checks. No dedicated mobile sync or push endpoints exist. /signup also remains on the standard router because the native app opens it in a browser without custom headers.

flowchart LR
    Request --> Match{Correct token, method, and allowlisted path?}
    Match -->|No| SSO[chain-auth]
    SSO --> Standard[Web UI or unspecified API]
    Match -->|Yes| Bypass[chain-no-auth]
    Bypass --> Strip[Strip proxy token]
    Strip --> App{Karakeep route auth}
    App -->|Public upstream| Discovery[Discovery or API-key exchange]
    App -->|Protected| Data[API key/session and scope checks]

The mobile app accepts arbitrary custom header name/value pairs in its Server Address form and applies them to tRPC, health/version, upload, and Karakeep-local asset requests. Settings are stored in Expo SecureStore, but their values are visible in the configuration UI. Configure the server as https://karakeep.${DOMAINNAME} with header name X-Karakeep-Proxy-Token and the decrypted KARAKEEP_MOBILE_PROXY_TOKEN value.

Rotate the token by updating the encrypted variable through the approved SOPS workflow, deploying, and then replacing the value on every device. The old token stops working immediately when the deployment takes effect. The Karakeep service documentation provides the operator procedure and upstream source references pinned to the audited mobile commit.

Open Archiver Network and Access Model

Blocked candidate; production adoption is not approved. The 2026-09-27 per-image review records all eight image inputs, exact-digest/platform scan inventories, maintenance evidence, confidence, and gaps. The inventory gap is closed, but every image remains NOT APPROVED. The published derivative has fixable HIGH/CRITICAL dependencies. The hardening and synthetic results below do not establish a secure image or authorize deployment. Remediation must remain upstream; no patched dependency fork will be maintained here. Require a verified upstream-fixed artifact or explicit operator acceptance through a narrowly documented exception under the adoption policy. The integration is implemented with no TrueNAS deployment. Image-bootstrap commit 831954bc4caac4edee212c2bf07cbabca6267c35 records the published derivative image's provenance. Findings are recorded in services/open-archiver/README.md.

Open Archiver uses https://open-archiver.${DOMAINNAME} with chain-auth@file (ItalyPaleAle's Traefik Forward Auth, not Authelia), then the app-local open-archiver-access Forward Auth middleware with condition Role("open-archiver-access"), followed by local authentication and MFA. The role gate covers every request, including GET /setup and POST /api/v1/auth/setup; a missing role denies access. Built-in SSO is not available in OSS. Before Custom App creation or route exposure, create the enabled Users/Groups role in the Entra registration used by ${AZURE_CLIENT_ID} and assign only the bootstrap administrator in its Enterprise application. The shared main portal remains unchanged. Follow the bootstrap role and fresh-session acceptance procedure; keep the gate after local setup locks and MFA is enabled. There are no published host ports or UI/API/monitoring bypasses. Open Archiver intentionally has no Gatus integration or unauthenticated monitoring router; in-container health checks remain enabled. Adding any route, including monitoring, requires review.

Shared Traefik currently has neither an open-archiver-frontend attachment nor an external declaration for that network. All eight Open Archiver services, including init, migration, and backup, have no Compose profile and are selected normally. Initial enrollment is manual TrueNAS Custom App creation: dccd.sh -t skips a missing app config directory, but generic/unscoped dccd outside TrueNAS mode and raw Compose do not have that enrollment guard and select this stack. Use the existing TrueNAS aliases and app-first rollout, not generic discovery.

Per-image adoption and runtime clearance and a reviewed Compose pin of the approved current-source, published and digest-verified derivative are still required. Assign only the Entra bootstrap administrator before Custom App creation. Keep shared Traefik's entries deferred until the Custom App creates open-archiver-frontend; then add the attachment and external declaration to tracked services/traefik/compose.yaml through normal review before the final dccd-all. The include-only Custom App YAML retains services: {}. Do not create untracked overrides or manually connect networks. The diagram below describes the intended flow after onboarding, not a currently exposed service.

Network Members and access
open-archiver-backend Internal: app, migrations, PostgreSQL, Valkey, Meilisearch, and database backup
open-archiver-frontend App only: outbound mailbox-provider connectivity; Traefik attachment awaits reviewed activation
open-archiver-parser Internal: app and Tika only; no database, queue, search, or Traefik membership
flowchart LR
    User --> Gate["Traefik + chain-auth"]
    Gate --> Role["App-local open-archiver-access role gate"]
    Role --> App["Open Archiver: local auth + MFA"]
    App --> Backend["Internal: PostgreSQL / Valkey / Meilisearch"]
    App --> Parser["Separate internal network: Tika OCR"]
    App --> Mail["Mailbox providers"]
Component Security boundary
App Five direct Node children under a non-root supervisor; build-time dependencies, separate migrations, read-only root, private disk-backed scratch
Backup Approved writable-root exception because /init writes /etc/bash/bashrc; drops all capabilities except CHOWN, DAC_OVERRIDE, FOWNER, SETUID, SETGID; mounts only backup output
PostgreSQL Native psql \getenv in read-only config/init-database.sql; nonsuperuser database owner and authenticated SELECT 1 health check; admin password never passed to app/migrations/backup
Tika Full OCR image; no archive mount, secrets, frontend membership, or direct internet route. Parser exploits can still reach its app peer
Valkey / Meilisearch Private backend state: queue/transient MFA data and rebuildable but sensitive searchable plaintext, respectively

All containers retain no-new-privileges and a 100-task PID limit. Identity details cover the backup pre-init account check and DBBACKUP_USER/DBBACKUP_GROUP selection. Detailed runtime settings and image tradeoffs remain in services/open-archiver/README.md.

Earlier local rootless Podman testing passed init/migrations, all five long-running service health checks, and synthetic database/archive/queue recovery. Compose pins the published GHCR derivative, anonymously pulled and recreated healthy. Repeated init, restart persistence, full reindex, and real HTTPS proxy-boundary checks passed. Proxy tests used official upstream Traefik 3.7.10 (not the DHI build), the earlier app labels/repository middleware, and a synthetic Forward Auth responder. These results do not validate the new role gate, non-recursive routine init/explicit repair mode, revised Dockerfile/publication workflow, or Valkey memory limits under load. The pinned derivative is unchanged; revised image source has not yet been published or runtime tested. Actual Entra role configuration, authorization checks, and TrueNAS deployment still require operator validation under issue #789. See Restore Open Archiver: database dumps alone are not a coordinated archive backup.

Dawarich Split Authentication Routers

Dawarich's main HTTPS router uses chain-auth-dawarich@file as defense in depth for the web UI, including protection against exposure of the upstream seeded demo administrator. The chain combines rate limiting, Forward Auth, and middlewares-dawarich-secure-headers. Its restrictive CSP permits connect-src 'self' https: wss:. HTTPS permits upstream's user-configurable vector, raster, and style basemap URLs; custom basemap endpoints must use HTTPS, and HTTP custom origins remain intentionally blocked. Explicit wss: permits ActionCable updates for the live map, family locations, tracks, and notifications because browsers do not consistently treat 'self' as authorizing WSS. worker-src stays limited to 'self' blob: for MapLibre, and the other directives remain restrictive. All requests use this router unless one of two higher-priority, chain-no-auth@file routers matches:

Router Match Application authentication
Mobile health Correct X-Dawarich-Proxy-Token and /api/v1/health None; returns no user data
Mobile data Correct X-Dawarich-Proxy-Token and any other allowlisted /api/v1 mobile resource User's Dawarich API key
Third-party ingestion POST to exactly /api/v1/owntracks/points, /api/v1/overland/batches, or /api/v1/traccar/points Dawarich API key

The mobile allowlist contains health, users/me, plan, settings/mobile, points, timeline, tracks, visits, stats, insights, digests, demo_data, families, photos, countries, maps, tiles, and places. The proxy token comes from DAWARICH_MOBILE_PROXY_TOKEN in the SOPS-encrypted service environment and limits the Forward Auth bypass to configured official clients. /api/v1/health intentionally skips API-key authentication upstream because the iOS app uses it for server and version discovery; it requires only the proxy token at Traefik and returns no user data. Every other allowlisted mobile data endpoint requires both the proxy token and the user's Dawarich API key.

The main router also covers /sidekiq. With SELF_HOSTED=true on Dawarich 1.14.2, that dashboard requires a signed-in Dawarich administrator account rather than separate dashboard credentials.

flowchart LR
    Request --> Mobile{Mobile header and allowlisted path?}
    Mobile -->|Yes| Health{Health endpoint?}
    Health -->|Yes| Discovery[Discovery; no user data]
    Health -->|No| MobileAPI[Mobile data with Dawarich API key]
    Mobile -->|No| Ingest{POST to exact ingestion path?}
    Ingest -->|Yes| IngestAPI[Ingestion API with Dawarich API key]
    Ingest -->|No| Entra[chain-auth-dawarich]
    Entra --> Default[Web UI or other route]

GPSLogger and PhoneTrack use the OwnTracks endpoint. API login and registration, /api-docs, non-POST requests to ingestion paths, and all other endpoints remain on the main router. Requests with a missing or incorrect mobile proxy token also fall through to Forward Auth. This prevents the known seeded demo@dawarich.app / safepassword credentials from being exchanged for an API key without first passing SSO; the credentials must still be changed immediately after first login.

Third-party clients may send API keys in query strings. Because query strings can be recorded in proxy access logs, access to those logs must be restricted and exposed keys must be rotated.

Cloudflare Tunnel through Traefik

Services that need internet exposure without opening inbound ports use Cloudflare Tunnel (cloudflared) combined with Traefik. The cloudflared agent establishes an outbound-only connection to Cloudflare's edge network, then forwards requests to Traefik over the shared Docker network. Traefik applies its standard label-based routing and middleware chain before reaching the backend service.

Traffic flow:

Internet → Cloudflare edge → cloudflared container → Traefik → backend service

All three containers (cloudflared, Traefik, and the backend) share the same frontend network (e.g., <app>-frontend). In the Cloudflare Zero Trust dashboard, the tunnel target is set to https://traefik with noTLSVerify enabled (Traefik presents a self-signed certificate on this hop; TLS is terminated at Cloudflare's edge for the external client). The backend service carries standard Traefik labels (e.g., chain-no-auth@file for a public API) so Traefik routes by Host header as usual.

No services currently use this pattern — the cloudflared compose stack is retained but paused. When a new public-facing service is added, cloudflared must be re-attached to that app's frontend network.

Why route through Traefik instead of directly to the backend?

  • Traefik middleware (rate limiting, headers, CORS) applies consistently whether traffic arrives from Cloudflare Tunnel or from the LAN.
  • Observability: all requests appear in Traefik's access logs and metrics.
  • One routing model: every service is configured via Traefik labels — no split between "Traefik services" and "tunnel-only services."

When to use Cloudflare Tunnel + Traefik vs Traefik alone:

Criteria Traefik only Cloudflare Tunnel + Traefik
Internal services (LAN only) ✓ —
Auth-protected public services ✓ ✓ (with chain-auth@file)
Public APIs (no auth) ✓ ✓ Preferred (zero inbound ports)
Requires inbound ports Yes (80, 443) No (outbound only)
TLS termination (external) Traefik (Let's Encrypt) Cloudflare edge

Gatus Internal Monitoring Entrypoint

Auth-protected services (those using chain-auth@file) redirect Gatus health checks to the OAuth login page, causing false-negative alerts. To monitor these services without bypassing network isolation, Traefik exposes a dedicated monitoring entrypoint on port 8444.

How it works:

  1. Each auth-protected service declares a secondary Traefik router (<app>-monitor) that listens on the monitoring entrypoint and uses chain-no-auth@file instead of chain-auth@file.
  2. Gatus sends HTTP requests to http://172.30.100.6:8444 with the appropriate Host header. The IP is Traefik's static address on gatus-frontend (set via ipv4_address in traefik/compose.yaml) to avoid Docker DNS falling through to the host's external resolver.
  3. Traefik routes the request to the target container over that service's dedicated frontend network.
  4. Gatus endpoints also set client.ignore-redirect: true as defense-in-depth — if the monitoring router were misconfigured, Gatus would still detect a redirect rather than silently passing.

Security model — three independent layers:

  1. Unpublished port. Port 8444 does not appear in Traefik's ports: mapping — the internet cannot reach it regardless of middleware misconfiguration.
  2. Docker network isolation. Only containers that share a Docker network with Traefik can open a TCP connection to it. However, Traefik joins every service's frontend network, so this alone is insufficient — any container could reach :8444.
  3. Entrypoint-level ipAllowList. The monitoring entrypoint in traefik.yml applies monitoring-ipallowlist@file (defined in middlewares.yml) which restricts source IPs to the gatus-frontend subnet (172.30.100.0/29). This runs before any router-level middleware, so it blocks all traffic from other frontend subnets. The gatus-frontend network uses a fixed IPAM subnet in gatus/compose.yaml to make this deterministic.

All three layers must be defeated for a container on another frontend network to bypass auth via the monitoring entrypoint.

Why not a shared secret header instead of an IP allowlist? Any container with a socket proxy (Dozzle, Homepage) could read the secret from docker inspect labels, so the secret is only as strong as the weakest socket-proxy consumer. The IP allowlist is infrastructure-derived, not stored anywhere a container can read it, and cannot be brute-forced.

Traffic flow comparison:

Browser request (authenticated):
  Browser → Host:443 → Traefik :443 → chain-auth → (sonarr-frontend) → Sonarr

Gatus health check (internal monitoring):
  Gatus → (gatus-frontend 172.30.100.0/29) → Traefik :8444 [172.30.100.6]
    → ipAllowList ✓ → chain-no-auth → (sonarr-frontend) → Sonarr

Other container attempting to use monitoring entrypoint:
  Sonarr → (sonarr-frontend 172.x.x.x) → Traefik :8444
    → ipAllowList ✗ (403 Forbidden)

Services with monitoring routers: All services monitored by Gatus have a -monitor router on the monitoring entrypoint. This includes both auth-protected services (which need it to bypass forward-auth) and no-auth services (which need it to avoid TLS/SNI issues when checking by IP address). Only the Gatus service itself is excluded (gatus.enabled=false).

Configuration locations (keep in sync when changing the subnet):

File What to update
services/gatus/compose.yaml gatus-frontend network ipam.config[0].subnet
services/traefik/compose.yaml Traefik gatus-frontend ipv4_address
services/traefik/config/rules/middlewares.yml monitoring-ipallowlist.ipAllowList.sourceRange
services/traefik/config/traefik.yml Comment documenting the subnet (for reference)

Alloy Metrics Scrape Entrypoint

Alloy scrapes Traefik's Prometheus metrics over a second internal entrypoint, metrics, listening on port 8082. The entrypoint is not host-published — only containers on a Docker network Traefik joins can reach it. An entrypoint-level metrics-ipallowlist@file middleware further restricts source IPs to the alloy-frontend subnet (172.30.100.8/29), pinned in services/alloy/compose.yaml. This mirrors the Gatus monitoring-entrypoint security model (unpublished port + Docker network isolation + entrypoint-level ipAllowList).

The metrics.prometheus block in traefik.yml enables addRoutersLabels, addServicesLabels, and addEntryPointsLabels so per-router, per-service, and per-entrypoint cardinality is exposed. Each Alloy instance scrapes the local Traefik instance running on the same host (works for both svlnas and svlazext, since Traefik's compose.svlazext.yaml already lists alloy-frontend in its network list).

Alloy additionally scrapes per-app postgres_exporter sidecars for Dawarich (postgres_dawarich job → dawarich-db-exporter:9187), Immich (postgres_immich job → immich-db-exporter:9187), and Outline (postgres_outline job → outline-db-exporter:9187) at 60s intervals. The exporters live in each app's own compose file and reuse that app's existing *_DB_PASSWORD secret — no DB credentials are added to Alloy's secret.sops.env. Reachability is provided by joining Alloy to the dawarich-backend, immich-backend, and outline-backend Docker networks (declared external: true in services/alloy/compose.yaml); _bootstrap creates the internal dawarich-backend before Alloy and Dawarich, and compose.svlazext.yaml drops all three app networks via networks: !override, since the apps are not deployed on svlazext. Dawarich's Rails exporter is disabled with PROMETHEUS_EXPORTER_ENABLED=false; Alloy scrapes only postgres_dawarich from dawarich-db-exporter, not the application container. Unifi is intentionally not included — it uses MongoDB, not Postgres.

Alloy also runs the built-in prometheus.exporter.github component (10m scrape interval, job=integrations/github_exporter) which polls the GitHub REST API for DevSecNinja/truenas-apps and DevSecNinja/dotfiles repo-level stats: rate-limit headroom, stars/forks/watchers, open PR/issue counts, repo size. No new compose service or network — it's an in-process exporter that calls api.github.com over the host's public egress. Auth is a fine-grained GitHub PAT (Metadata + Issues + Pull requests read-only on the listed repos), stored as GITHUB_API_TOKEN in services/alloy/secret.sops.env. Single-host scrape: a discovery.relabel keep rule on HOSTNAME_OVERRIDE filters the target list to empty on every host except svlnas, so only one Alloy instance polls the API (avoids duplicated metrics and doubled API quota). Pairs with the official Grafana Cloud "GitHub integration" dashboards (GitHub API Usage, GitHub Repository Stats).

Alloy additionally tails the host systemd journal via loki.source.journal, so dccd.sh deploy logs (already emitted to journald via logger -t dccd) and host-level signals not visible from container logs — sshd auth events, smartd disk warnings, kernel/OOM messages, ZFS events — land in Loki alongside container logs. The component reads /var/log/journal directly: the persistent journal directory, /run/log/journal, and /etc/machine-id are bind-mounted read-only into the container, and group_add adds the host's systemd-journal group GID to UID 3125 so it can read mode-0640 journal files owned by root:systemd-journal. The numeric GID differs per host (102 on svlnas, 999 on svlazext) and is hardcoded per host in compose.yaml and compose.svlazext.yaml respectively — same per-host override pattern as HOSTNAME_OVERRIDE, since secret.sops.env decrypts to the same plaintext on every host and so cannot carry per-host values. The pipeline promotes __journal_syslog_identifier, __journal__transport, and __journal_priority_keyword to indexed labels (syslog_identifier, transport, level) alongside the static host, instance=$HOSTNAME_OVERRIDE, and job=integrations/node_exporter. The unit label is populated transport-aware: when transport=syslog it takes __journal_syslog_identifier, otherwise it takes __journal__systemd_unit. This split exists because TrueNAS' cron spawns the user shell inside a transient session-<N>.scope unit, so __journal__systemd_unit is set to that scope rather than empty for dccd.sh lines; preferring the syslog identifier when transport is syslog makes dccd (and any other logger -t … producer) selectable from the Grafana Cloud "Linux Server / Logs" dashboard's unit dropdown, while real services (sshd.service, smartd.service, etc.) continue to use their proper unit names because they don't use the syslog transport — the same job/instance pair carried by the host metrics stream, so Grafana Cloud's "Linux Server" integration "Logs" dashboard works against this Alloy without modification. __journal__boot_id is intentionally not promoted (bounded but a fresh stream per reboot accumulates over months); every other journal field stays on the log line, keeping Loki stream cardinality flat. Two loki.process drop stages run on the stream: a match stage suppresses high-volume, zero-signal pam_unix .* session (opened|closed) messages from CRON and systemd-logind, and a 167h stale-entry drop stage matches the docker pipeline's Cloud Free 168h ingest-window protection. The source itself is capped with max_age = "12h0m0s" so a fresh tail (first start, or after losing its cursor) cannot read far enough back to produce rejected batches. The component's read cursor lives at /var/lib/alloy/data/<component>/positions.yml (already on the chowned ./data mount), so restarts resume cleanly. Both svlnas and svlazext have persistent journaling enabled and ship journal logs.

Subnets and entrypoints in use:

Entrypoint Port Source range allowed Purpose
monitoring 8444 gatus-frontend (172.30.100.0/29) Gatus auth-free health checks
metrics 8082 alloy-frontend (172.30.100.8/29) Alloy Prometheus scrape of Traefik

Configuration locations (keep in sync when changing the subnet):

File What to update
services/alloy/compose.yaml alloy-frontend network ipam.config[0].subnet
services/traefik/config/rules/middlewares.yml metrics-ipallowlist.ipAllowList.sourceRange
services/traefik/config/traefik.yml metrics entrypoint address + metrics.prometheus
services/alloy/config/config.alloy prometheus.scrape "traefik" target host:port

Docker Socket Proxy

Services never mount /var/run/docker.sock directly. Instead, each gets its own LinuxServer socket-proxy instance with minimal permissions:

  • CONTAINERS=1 — read container metadata only
  • POST=0 — read-only, no mutations (Homepage)
  • Separate proxy per service to prevent lateral movement

Why one proxy per service instead of sharing?

If Traefik and Homepage shared one proxy, compromising either would grant the attacker the union of both permission sets. Separate proxies enforce least privilege per consumer.

Docker Compose Profiles

Repository policy: reserve profiles for genuinely optional runtime workloads, such as surveillance. Do not automatically add a staging, adoption, or onboarding profile to normal TrueNAS apps. Initial enrollment uses manual Custom App creation through the existing app-first rollout.

TrueNAS mode (dccd.sh -t) skips an app when /mnt/.ix-apps/app_configs/<app>/versions is missing. This enrollment guard belongs to dccd's TrueNAS mode, not ordinary Compose. Open Archiver has no profile and normal Compose selection includes all eight services. Generic/unscoped dccd outside TrueNAS mode and raw Compose do not check that enrollment and select this stack. Use the sourced TrueNAS aliases and the existing onboarding sequence; lack of a Custom App is not an inactive-by-default guarantee outside TrueNAS mode. Manual enrollment does not waive image-adoption or vulnerability-risk review.

Services with a profiles: key are excluded from default service selection when no matching profile is active and no service is explicitly targeted. dccd.sh skips stacks with zero selected services in both TrueNAS and generic modes. Profiles are selection controls, not authorization or risk approval: explicit named-service targeting and --profile '*' bypass default selection.

Services using profiles:

Profile Services Purpose
surveillance frigate-init, frigate NVR — only needed when away from home

Activating a Profile

The examples below apply to the optional surveillance workload.

Environment variable (recommended): Docker Compose natively reads the COMPOSE_PROFILES variable. Set it before running dccd.sh:

export COMPOSE_PROFILES=surveillance

Multiple profiles can be comma-separated:

export COMPOSE_PROFILES=surveillance,other

On TrueNAS, for operational profiles such as surveillance: Add the export to the cron job that runs dccd.sh, or to the deployment user's shell startup file for interactive runs. dccd forwards it as Compose CLI profile flags. This does not establish that the TrueNAS UI receives the variable.

CLI flag (one-off): For a single manual run without persisting:

docker compose --profile surveillance up -d

Deactivating a Profile

Unset the variable or remove the export line:

unset COMPOSE_PROFILES

Removing a profile prevents its services from being selected for a normal deployment; it does not guarantee that already-running containers are stopped or removed. Explicitly stop the affected services and verify their state. For the surveillance stack, an explicit teardown is:

docker compose --profile surveillance down

Gatus Monitoring

Once profiled containers are stopped and removed, their Traefik labels are not discoverable by the Gatus sidecar. Merely deactivating a profile does not prove that old containers or their monitoring labels have disappeared.

Config-Change Container Recreation

scripts/dccd.sh computes a hash of watched configuration files and exports it as CONFIG_HASH. Services can use this hash in Docker Compose service labels to automatically recreate/restart containers when configuration changes.

How to opt in:

  1. Specify what to watch via the config.watch label — this is an app-level setting that controls which path to hash. Set it exactly once at the top of the service list:
  2. Omit or set config.watch=config (default) — hash the entire ./config directory
  3. Set config.watch=none — disable config-change detection (no hashing)
  4. Set config.watch=config/subdir — hash only a specific config path (relative to service directory)

  5. On all services that should recreate when config changes, add a label with the config.sha256 key:

    services:
      example:
        labels:
          - "config.watch=config/app.conf"    # App-level setting (set exactly once)
          - "config.sha256=${CONFIG_HASH:-}"  # Recreate when config changes
      example-init:
        labels:
          - "config.sha256=${CONFIG_HASH:-}"  # Also references the same hash
    

How it works:

  • On each deploy, dccd.sh reads config.watch labels across services; the first one found controls a single CONFIG_HASH value
  • The hash is computed as a deterministic SHA256 of the watched path and exported as ${CONFIG_HASH}
  • Multiple services can reference ${CONFIG_HASH} in their labels; all see the same hash value
  • When config changes, Docker Compose recreates/restarts all services that reference ${CONFIG_HASH}
  • If the watched path does not exist or is set to none, CONFIG_HASH is empty, and service labels referencing it keep an empty value — no container recreation is triggered

Generated Config and Git-Tracked Paths

CONFIG_HASH computation itself does not write files. Services that are recreated by this mechanism must write any generated or substituted config only under ignored runtime paths such as ./data/ or ./backups/, never under ./config/ or other git-tracked paths. Since services/**/data/ and services/**/backups/ are gitignored, generated outputs (e.g., Unbound's ./data/unbound/) will not block git pull when containers are restarted.

Directory Conventions

Each service follows a consistent layout:

services/<service>/
  compose.yaml       # Service definition — committed to Git
  secret.sops.env    # SOPS-encrypted secrets — committed to Git
  config/            # Static configuration — committed to Git
  data/              # Runtime data — NOT committed to Git
  backups/           # Backup output — NOT committed to Git

compose.yaml defines the service: images, networks, volumes, labels, and resource limits. It is the source of truth for how the service runs and is always committed to Git.

secret.sops.env stores secrets (API keys, passwords, tokens) encrypted with SOPS + Age. Because the values are encrypted, the file is safe to commit to Git. At deploy time, the CD script decrypts it to .env which is excluded from Git.

config/ holds files that you author and version-control: configuration files, rule sets, and any other inputs the container reads at startup. For example, Traefik's config/ contains traefik.yml and the dynamic rules under rules/. These are mounted :ro into the container because the container should only read them, never write to them.

data/ holds files that are produced or mutated by the running container: databases, certificates, caches, state files, and other dynamic output. This directory lives only on the host machine and is excluded from Git via .gitignore. It is mounted read-write so the container can persist its runtime state across restarts.

Named Docker volumes are not used in this repo. All persistent container data uses bind mounts to ./data/ (or a subdirectory of it). This ensures that TrueNAS ZFS snapshots — taken at the dataset level — capture all container state without needing to snapshot opaque Docker-managed volumes. It also makes data locations explicit and auditable from the host filesystem.

ZFS datasets do not need to be created manually for individual services. The full dataset hierarchy is established once during initial setup (see README.md § Setup). Each services/<app>/ directory lives on the vm-pool/apps dataset (or a child dataset created at setup time). TrueNAS handles snapshots and replication of these datasets automatically — no per-service backup containers are needed for file-level data (only for databases, which require consistent pg_dump / mongodump exports).

vm-pool/homes is a native-host sibling of vm-pool/apps, not app storage. The approved Personal Home Folders design uses one shared Multiprotocol dataset (NFSv4/Passthrough, case-sensitive, Atime Off, Exec On), with automatic SMB/NFS shares disabled at creation. Its administrative root grants ordinary users only non-inheriting traversal, not listing or creation. For each personal account, Create Home Directory checked with parent /mnt/vm-pool/homes makes TrueNAS create the ordinary /mnt/vm-pool/homes/<username> directory with the personal owner/private group and explicit 0700. Verify the saved path and folder ACLs; no per-user dataset or separate manual home creation is needed. Existing/restored homes instead use the full path with creation unchecked.

Separately create each private Files/ directory and individual SMB share; .ssh, .config, and shell dotfiles remain SSH/local-only. Folder/file ACLs provide privacy between ordinary users, not protection from administrators or a jailed SSH shell. Optional user quotas are ownership-based; dataset properties, encryption, snapshots, and replication are shared. A shared dataset rollback affects every user. No app/service accounts or the truenas_admin home and boot-time mirror are changed. Snapshot/replication coverage of homes and off-site copies of each user's files must be verified, not inferred from a pool-level task. Host setup and tests are pending.

backups/ holds database backup files produced by backup sidecars such as tiredofit/db-backup and nfrastack/db-backup. Like data/, this directory is excluded from Git and mounted read-write. Each backup type gets its own subdirectory (e.g., backups/db-backup/).

Secret Management

Secrets are encrypted with SOPS + Age and stored in git as secret.sops.env. The CD script decrypts them to .env at deploy time using an Age key stored on the TrueNAS host.

PUID and PGID are hardcoded directly in each service's compose.yaml so they are visible, auditable, and self-documenting. See Infrastructure § UID/GID Allocation.

Shared Environment Files

Reusable env files live in services/shared/env/ and are referenced via relative paths in env_file blocks. They are committed to Git because they contain no secrets.

File Purpose When to include
tz.env Sets TZ=Europe/Amsterdam Every container, including changedetection app, init, Chrome, and browser proxy; all eight Open Archiver containers

UID and GID values are not stored in shared env files or in secret.sops.env. They are hardcoded directly in each service's compose.yaml (in the user: directive and init container commands) so they are visible, auditable, and not treated as secrets. See Infrastructure § UID/GID Allocation for the full allocation table.

Shell Script Logging

Utility scripts under scripts/ source the vendored lib/log.sh library (pinned to a tagged release) for consistent, colorized, timestamped output with severity-aware routing (INFO/STATE/RESULT/HINT/STEP to stdout; WARN/ERROR/FATAL to stderr).

Every script sources lib/log.sh and sets a LOG_TAG matching the script's purpose:

_SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=lib/log.sh disable=SC1091
. "${_SCRIPT_DIR}/lib/log.sh"
# shellcheck disable=SC2034
LOG_TAG="dccd"

compose_up_wait_tolerant logs actual Compose failure output through log_data ERROR on stderr. On TrueNAS, dump_project_logs_tail adds selected container state and failed health probe output alongside log tails; the summary reports deployment failure rather than assuming a timeout. Cron may hide standard output, but must not hide standard error. Diagnostics omit the full container environment; review logs and probe output for secrets before sharing.

Helpers

Severity helpers (pick by intent, not by colour):

Helper Use for
log_error Failures that abort the operation
log_warn Recoverable issues, fallbacks, deprecations
log_info Neutral progress and informational messages
log_state An action that is in progress (e.g. "Deploying")
log_result A completed outcome or summary line
log_hint Suggested next steps for the user
log_step Numbered steps inside a wizard or multi-stage flow

Structural helpers for visual grouping:

Helper Renders
log_banner <title> [KIND] Boxed banner for major phase boundaries
log_rule [KIND] <title> Titled horizontal divider
log_sep [KIND] Plain horizontal rule

KIND is one of INFO, STATE, RESULT, HINT, STEP, WARN, ERROR and selects the colour.

Constraints

  • Functions that return data via stdout (printf '%s' "$value") must redirect any internal log calls to stderr — otherwise log lines contaminate the captured value:
get_thing() {
    log_info "looking up thing" >&2
    printf '%s' "${value}"
}
  • Heredoc / grouped redirects to a file ({ echo …; echo …; } > FILE) must keep using plain echo. Those blocks produce file content, not log output, and routing them through log_* would prepend timestamps and tags into the written file.

  • GitHub Actions annotation scripts (scripts/gha-image-age-check.sh, scripts/gha-trivy-image-scan.sh) intentionally use plain echo so that ::warning:: / ::error:: workflow commands reach the runner verbatim. Do not migrate these to log.sh.

Refreshing the vendored copy

scripts/lib/log.sh is vendored from a pinned upstream release. To bump it, run:

bash scripts/update-log-sh.sh

The refresher script downloads the configured release, verifies the SHA-256 against scripts/lib/log.sh.sha256, and updates both files in place. Commit the result as a chore change.