> ## Documentation Index
> Fetch the complete documentation index at: https://docs.swarmd.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Configuration

> Application settings, secrets, SMTP, and pointing the chart at infrastructure you already run.

# Configuration

Everything that isn't the database layout or ingress: how settings reach the
services, where secrets come from, and how to substitute your own
infrastructure.

***

## Your own values file

Whatever we send you — the plain chart, or a values file sized for your
cluster — is a starting point, not something to edit in place. Keep your
overrides in a file you own, in your own git repo, and layer it on top.

```bash theme={null}
helm install swarmd $SWARMD_CHART -n swarmd \
  -f swarmd-baseline.yaml \      # what we sent you
  -f my-overrides.yaml \         # yours — wins on any key it sets
  --set licence.existingSecretRef.name=swarmd-licence
```

**Later `-f` wins, and `--set` beats every `-f`.** You only need to write the
keys you are changing; everything else falls through to the file underneath
and then to the chart's defaults.

That separation is what makes the next upgrade cheap. When we send a new
baseline, you replace one file and your overrides are untouched — instead of
diffing your edits back out of a file that has moved on.

### Sizing one service

`services.<name>.resources` merges onto `services.defaults.resources`, so you
can raise one number without restating the block:

```yaml theme={null}
services:
  defaults:
    resources:
      requests: { cpu: 100m, memory: 512Mi }
      limits:   { cpu: "1",  memory: 1Gi }

  audit:
    resources:
      requests: { memory: 2Gi }     # cpu request still 100m
      limits:   { memory: 3Gi }     # cpu limit still "1"
```

### How the layering actually merges

<AccordionGroup>
  <Accordion title="Maps merge key by key">
    Setting one key leaves its siblings alone:

    ```yaml theme={null}
    # baseline
    appSettings:
      LOG_LEVEL: info
      TZ: UTC

    # yours
    appSettings:
      LOG_LEVEL: debug
    ```

    Result: `LOG_LEVEL: debug`, `TZ: UTC`. Nesting goes as deep as you like.
  </Accordion>

  <Accordion title="Lists replace wholesale — they do not append">
    This is the one that surprises people:

    ```yaml theme={null}
    # baseline
    global:
      image:
        pullSecrets: [{ name: a }, { name: b }]

    # yours
    global:
      image:
        pullSecrets: [{ name: c }]
    ```

    Result: `[{ name: c }]`. `a` and `b` are **gone** — not appended to.

    Any list is affected: `pullSecrets`, `tolerations`, `ingress.hosts`. If you
    mean to add to a list, restate the whole thing including the entries you
    want to keep.
  </Accordion>

  <Accordion title="Empty is not the same as absent">
    Omitting a key inherits it. Setting it to `{}`, `[]` or `null`
    *overrides* it with emptiness — which is how you deliberately clear
    something the baseline set, and also how people accidentally wipe a block
    they meant to leave alone.
  </Accordion>
</AccordionGroup>

<Warning>
  **Helm does not carry values forward between releases.** An upgrade that omits
  a `-f` you passed at install time reverts those settings to the chart defaults
  — silently, because from Helm's point of view you asked for the defaults.

  Pass the same `-f` files on every `helm upgrade`. `--reuse-values` looks like
  the shortcut and is not one: it also skips new defaults introduced by the
  chart you are upgrading *to*, so a fresh install and an upgraded install end
  up configured differently. Keep the files, pass them every time.
</Warning>

### Checking what is actually applied

```bash theme={null}
helm -n swarmd get values swarmd        # just your overrides
helm -n swarmd get values swarmd -a     # merged with every chart default
```

The second is the real answer to "what is this cluster running with", and the
first thing worth looking at when a setting seems not to have taken.

<Tip>
  Preview before applying: `helm diff upgrade` (from the `helm-diff` plugin) or
  `helm upgrade --dry-run=server`. Never `helm template` against a live install
  — see [Upgrades](/self-hosting/upgrades#the-one-that-bricks-installs).
</Tip>

***

## Application settings

Anything that would otherwise be a bespoke env var — encryption keys,
timeouts, log levels, model config — goes under `appSettings`, and lands as an
env var on every service.

Each entry takes one of three shapes:

```yaml theme={null}
appSettings:
  # 1. A plain value
  LOG_LEVEL: debug

  # 2. From a Secret you already have
  OPENAI_API_KEY:
    secretRef: my-openai-secret
    key: api-key

  # 3. From a ConfigMap you already have
  MODEL_CONFIG:
    configMapRef: llm-config
    key: models.json
```

These render to native `env:` entries — scalars as `value:`, maps as
`valueFrom.secretKeyRef` or `valueFrom.configMapKeyRef`. A map that is neither
shape fails the install rather than rendering something odd.

### Per-service overrides

`services.<name>.settings` takes the same three shapes and **wins** over
`appSettings` for the same key:

```yaml theme={null}
appSettings:
  LOG_LEVEL: info              # everything runs at info

services:
  relay:
    settings:
      LOG_LEVEL: debug         # ...except relay
      MODEL_CONFIG:
        configMapRef: llm-config
        key: models.json
```

<Note>
  The defaults ship `LOG_LEVEL=info` globally and wire the keys the services
  need from the chart-generated Secret, each only to the service that reads it:
  `SWARMD_ENCRYPTION_KEY` to registry, relay and teams, and
  `AUDIT_HASHCHAIN_SIGNING_KEYS` to audit (under `services.audit.settings`).
  That is enough for a first install to work with no input. Override them with
  your own key management when you go to production — see
  [Bring your own encryption keys](#bring-your-own-encryption-keys).
</Note>

***

## Secrets

The chart generates a Secret named `swarmd-generated-credentials` at install:

| Key | Used by |
| - | - |
| `postgres-username` / `postgres-password` | Every DB-backed service |
| `keycloak-admin-username` / `keycloak-admin-password` | Services calling the Keycloak Admin API |
| `clickhouse-username` / `clickhouse-password` | Populated even when ClickHouse is off, so enabling it later doesn't rotate anything |
| `encryption-key` | This install's 32-character symmetric key. Signs UI sessions, and is the `v2` entry of `encryption-keys` |
| `encryption-keys` | The key list registry, relay and teams encrypt stored credentials with (`SWARMD_ENCRYPTION_KEY`): `v1:` the built-in key used by installs before 0.4.1, kept so older values still decrypt, and `v2:` this install's key, used for every new write |
| `service-client-secret-<service>` | Each platform service's Keycloak client secret. The service writes it to Keycloak when it starts |
| `audit-hashchain-signing-key` | Signs audit hash-chain checkpoints (Immutable proof, sealed retirements). Rotating it breaks verification of every checkpoint signed before the rotation |

Passwords are 24 random alphanumeric characters. On `helm upgrade` the chart
reads the existing Secret via `lookup` and reuses the values — **upgrades do
not rotate credentials**.

The Secret carries `helm.sh/resource-policy: keep`, so `helm uninstall` leaves
it behind.

<Warning>
  **Back up `encryption-key` and `encryption-keys` separately from your
  database.** They encrypt data at rest in registry, relay and teams. A database
  restore without the matching keys leaves you with rows nobody can decrypt.
</Warning>

### Supplying your own

```yaml theme={null}
global:
  generatedSecrets:
    enabled: false
    name: swarmd-generated-credentials
```

Then pre-create a Secret of that name with the keys above. Useful when your
platform team owns secret material, or you sync from Vault or OpenBao with an
operator. The `encryption-keys` and `service-client-secret-*` keys are
optional in your own Secret: without them the services fall back to built-in
defaults, which work but are not secret, so add them before production.

### Bring your own encryption keys

The services read `SWARMD_ENCRYPTION_KEY` as a comma-separated list of
`v<N>:<base64 of 32 bytes>` entries. The highest version encrypts new values;
every listed version can decrypt. Generate a key with `openssl rand -base64 32`
and set it as a setting, which replaces the generated list:

```yaml theme={null}
appSettings:
  SWARMD_ENCRYPTION_KEY:
    secretRef: my-swarmd-keys
    key: keys     # e.g. "v1:<old key>,v2:<old key>,v3:<new key>"
```

To rotate, append a higher version and keep the old ones. Never remove a
version that still has rows encrypted with it.

***

## Agents and MCP servers on private addresses

The relay refuses to call private, in-cluster, loopback or link-local
addresses (SSRF protection). An agent or MCP server running in your cluster
(`http://invoice-agent.agents.svc.cluster.local:8080`) or on your private
network is refused with *Blocked request to private/internal address* until
its host is listed:

```yaml theme={null}
services:
  relay:
    ssrfAllowedHosts: "*.agents.svc.cluster.local,crm-mcp.corp.internal"
```

* Entries are exact host names, or `*.` patterns that match any host in that
  domain (`*.agents.svc.cluster.local` matches `invoice.agents.svc.cluster.local`).
* A host admitted by a pattern still may not resolve to loopback, link-local
  (cloud metadata) or `0.0.0.0`.
* Public hosts need no entry.
* Keep patterns to the namespaces your agents live in. `*.svc.cluster.local`
  would let anyone who can register an agent reach every service in the
  cluster through the relay, Keycloak included.

The chart's install notes remind you when this list is empty.

***

## SMTP

tenant-auth sends email for verification, password reset and team invites.
With `smtp.enabled: false` — the default — every one of those is **logged and
dropped**. Users can register, but nothing arrives.

<Warning>
  **Without SMTP you have a one-account deployment.** Sign-up and login work, so
  the first admin gets in fine — but that is as far as it goes:

  | | Without SMTP |
  | - | - |
  | Sign up, log in | ✅ works |
  | **Invite anyone else** | ❌ the invite link never arrives, so **you cannot add a second user** |
  | **Password reset** | ❌ the link never arrives, so a forgotten password can only be fixed by a Keycloak admin |

  Fine for a solo evaluation. Not fine for anything a team touches, and not
  something to discover after handing the platform over.
</Warning>

### Nobody is locked out, though

`smtp.enabled` also sets the realm's `verifyEmail` flag, so with SMTP off
Keycloak never raises the "verify your email" required action and sign-up
followed by login works normally. Verification is genuinely off, not skipped by
a special case: tenant-auth does refuse a login for an unverified user, but only
when Keycloak reports the account as not fully set up, which cannot happen while
`verifyEmail` is false.

Those users are still *recorded* as unverified, so they show as `PENDING` in
user listings. That is cosmetic — it gates nothing.

```yaml theme={null}
smtp:
  enabled: true
  host: smtp.postmarkapp.com
  port: 587
  from: "no-reply@swarmd.example.com"
  credentialsSecret: smtp-creds     # Secret with `username` and `password`
```

```bash theme={null}
kubectl -n swarmd create secret generic smtp-creds \
  --from-literal=username='<smtp-user>' \
  --from-literal=password='<smtp-pass>'
```

The install refuses to render with `smtp.enabled: true` and any of `host`,
`from` or `credentialsSecret` missing.

<Warning>
  **Turning SMTP on later does not start enforcing email verification.**
  `verifyEmail` is written into the realm when the realm is *created*, and the
  bootstrap Job is create-only — on every subsequent install and upgrade it finds
  the realm, logs `already exists — nothing to do`, and changes nothing. So a
  deployment that first came up without SMTP keeps `verifyEmail: false` even
  after you enable SMTP and upgrade, and existing unverified users carry on
  logging in.

  Verified on a live install: enable SMTP, upgrade, and an account created before
  the change still logs in without verifying.

  Enabling SMTP does make invites and password resets start working, which is
  usually the reason for turning it on. If you also want verification enforced,
  change it on the realm yourself through Keycloak's admin console or API — the
  chart cannot do it for you without overwriting realm customisations you may
  have made.
</Warning>

<Note>
  `swarmd.tenant-auth.auto-verify-email-patterns` marks matching addresses
  verified at creation — a comma-separated list of globs, e.g.
  `*@yourcompany.com`. Useful once verification is enforced but you want your own
  domain to skip it. Empty by default, so nobody is auto-verified.
</Note>

***

## Bring your own infrastructure

Every optional component follows the same three-state pattern:
`enabled: false` · `enabled: true, deploy: true` · `enabled: true,
deploy: false` plus an `external` block. The chart fails the render if you
choose the third and leave a required URL empty.

### Bring your own Keycloak

```yaml theme={null}
keycloak:
  deploy: false
  external:
    serverUrl: http://keycloak.internal:8080        # in-cluster reachable
    publicUrl: https://auth.example.com             # what browsers see
    adminSecret: my-keycloak-admin                  # username + password keys
```

`adminSecret` holds master-realm admin credentials. Registry, relay,
tenant-auth and teams call the Admin API to manage agent service accounts, so
this is not optional.

<Warning>
  **External Postgres forces external Keycloak.** `keycloak.deploy=true` with
  `postgres.deploy=false` is unsupported and fails the render — the chart can't
  derive host and port from an arbitrary JDBC URL for the Keycloak container.
</Warning>

### PII detection

Both analyzers are **off by default**. With neither enabled, prompts pass
through the LLM path unfiltered.

```yaml theme={null}
presidio:
  enabled: true
  deploy: true            # in-cluster, pulled from Microsoft's public registry

piiranha:
  enabled: true
  deploy: false           # point at one you run
  external:
    url: http://piiranha.internal:5002
```

<Warning>
  **Piiranha's analyzer image sits outside the release set.** It is a
  third-party model image, not built or versioned with the platform, and it is
  not covered by the testing that goes into a Swarmd release. It is also large —
  roughly 1.3 GiB — so first pull is slow and it wants a node with room for the
  model in memory.

  Run it yourself and point the chart at it:

  ```yaml theme={null}
  piiranha:
    enabled: true
    deploy: false
    external:
      url: http://piiranha.internal:5002
  ```

  If you would rather have it in-cluster, set `piiranha.image.repository` to
  your own copy so you control what version you are running.

  Presidio has no such constraint — its image comes from Microsoft's public
  registry, so `deploy: true` is fine.
</Warning>

### a2a-payments

Off by default, and the only service that genuinely cannot start without
configuration: its `X402_WALLET_PRIVATE_KEY` is `@NotBlank`-validated with no
degraded mode.

```yaml theme={null}
services:
  a2aPayments:
    enabled: true
    settings:
      X402_WALLET_PRIVATE_KEY:
        secretRef: my-wallet
        key: private-key
```

### Turning services off

```yaml theme={null}
services:
  teams:
    enabled: false
  a2aPayments:
    enabled: false
```

Gateway, registry, relay, audit and tenant-auth are the core — disabling any
of them gives you a platform that doesn't work.

<Warning>
  **Keep `notification` on.** Registry, audit and tenant-auth hand governance
  notices (suspension, retirement), serious-incident deadline alerts and monitor
  alerts to notification-service through an outbox. With it off they retry
  against a Service that does not exist and then drop the notice — nothing
  fails visibly. The install notes warn you when it is off. Channels deliver by
  e-mail (needs [SMTP](#smtp)), Slack or Microsoft Teams; *Webhook* destinations
  have no sender yet, so the API refuses to create one.
</Warning>

***

## Settings that bite later

Some settings don't block startup but fail the first time a feature is used.
The chart can't enforce them, because the service boots fine without them.

| Service | Setting | Symptom if missing |
| - | - | - |
| `teams` | `TEAMS_BOT_APP_ID`, `TEAMS_BOT_APP_SECRET` | First Teams webhook returns `500` |
| `teams` | `OPENAI_API_KEY` | LLM routing call fails |
| `registry` | `MCP_OAUTH_REDIRECT_URI` | MCP dynamic client registration errors when an admin triggers it |

Provide them under `services.<name>.settings` when you enable the feature.

***

## Scaling: core and worker

Every DB-backed service ships as one image that can run three ways, and the
chart **splits every one of them by default**:

| Workload | Runs | Sized by |
| - | - | - |
| `<service>` | The API. `SWARMD_RUNTIME_MODE=core`. | `services.<name>.replicaCount` / `.resources` |
| `<service>-worker` | Schedulers and outbox drainers. `SWARMD_RUNTIME_MODE=worker`. | `services.<name>.worker.replicaCount` / `.resources` |
| `<service>-migrate` | Flyway, once, then exits. Runs before both. | — |

The worker's web server stays up so health probes work, but the Service never
routes traffic to it. Worker sizing falls back to the service's own
`replicaCount` / `resources` when left empty.

```yaml theme={null}
services:
  audit:
    replicaCount: 3            # three API pods
    worker:
      enabled: true
      replicaCount: 2          # two worker pods, sized independently
      resources:
        requests: { cpu: 500m, memory: 1Gi }
```

<Note>
  **Why this is the default.** In the combined topology, scaling the API to *N*
  replicas also runs *N* copies of every `@Scheduled` job — outbox drainers and
  monitor evaluators racing each other, which is a data problem rather than a
  performance one. Splitting also stops a slow scheduler starving the API's
  thread pool, and takes Flyway out of pod startup so a long migration can't
  trip readiness on every replica at once.
</Note>

### Collapsing a service back

```yaml theme={null}
services:
  teams:
    worker:
      enabled: false      # single container: API + schedulers + boot migration
```

Worth doing on a laptop, where the extra pod per service costs more than the
isolation buys you. Setting it on every service takes the default install
from 18 Deployments to 11.

<Warning>
  Don't run a collapsed service at `replicaCount > 1`. That is exactly the
  scheduler-duplication case the split exists to prevent.
</Warning>

***

## Startup ordering

Every service talks to Keycloak while it boots — the Admin API for
service-account management, the JWK set for token validation — and exits
non-zero if it cannot reach it. On a cold install that is a race the services
lose, so each one is held at a `wait-for-keycloak` init container until
Keycloak's realm answers its OIDC discovery endpoint.

You will see pods sit in `Init:0/1` for roughly a minute on a first install.
That is the gate working; **expect zero restarts**.

```yaml theme={null}
services:
  defaults:
    waitForKeycloak: true          # default
    waitImage: busybox:1.36        # must be pullable before the service image
```

Set `waitForKeycloak: false` per service (or on `defaults`) if you sequence
startup another way. Migration Jobs are never gated — they run Flyway and
never touch Keycloak.

<Note>
  The gate checks the **realm's discovery document**, not the TCP port.
  Keycloak accepts connections well before the realm-bootstrap Job has imported
  the realm, and a service starting in that window fails exactly as it would
  have without the gate.
</Note>

***

## JVM sizing

The chart sizes the JVM against the container memory limit:

```yaml theme={null}
appSettings:
  JAVA_TOOL_OPTIONS: "-XX:MaxRAMPercentage=75.0 -XX:InitialRAMPercentage=50.0 -XX:+ExitOnOutOfMemoryError"
```

<Warning>
  **Do not remove this while lowering the memory limit.** Without
  `MaxRAMPercentage` the JVM falls back to 25% of the limit — a 512Mi limit
  gives a 128Mi heap, which Spring Boot with \~40 JPA repositories cannot start
  in. The symptom is a container that reaches "Starting service \[Tomcat]" and
  is then OOMKilled (exit 137) in a crash loop.
</Warning>

On a running install these services sit at **593–697 MiB** working set, which
is why the default limit is 1Gi and the default request 512Mi. If you raise
the limit, the heap grows with it; if you lower it below \~768Mi, lower
`MaxRAMPercentage` too or the JVM will size a heap the container cannot hold.

No garbage collector is pinned deliberately — JVM ergonomics picks one suited
to the container, which matters more here than in a large deployment because
this chart is routinely run with small limits.

***

## Observability

```yaml theme={null}
global:
  observability:
    prometheus:
      enabled: true
      path: /actuator/prometheus
    tracing:
      enabled: true
      endpoint: "http://otel-collector.observability:4317"
```

Prometheus scrape annotations land on every service pod; tracing points the
OTLP exporter at your collector.

***

## Verifying a change before you apply it

`$SWARMD_CHART` is the chart file we sent you — see
[Quickstart → Step 2](/self-hosting/quickstart).

```bash theme={null}
helm lint $SWARMD_CHART

helm template swarmd $SWARMD_CHART \
  -f my-values.yaml \
  --set licence.key=LIC-TEST | kubectl apply --dry-run=client -f -
```

The chart's validation runs during `template`, so a bad combination fails here
— before it touches your cluster.

<Warning>
  `helm template` renders **without a cluster connection**, which is why the
  `--set licence.key=LIC-TEST` above is needed and why the output is only good
  for checking your values. Never apply a `helm template` render to a live
  install: the chart cannot read back the credentials Secret in that mode and
  the output carries freshly generated passwords. Use
  `helm upgrade --dry-run=server` to preview a real change — see
  [Upgrades](/self-hosting/upgrades#the-one-that-bricks-installs).
</Warning>

Unpacked the tarball (`tar xzf`)? Both commands work against the `./swarmd`
directory the same way.

***

## Checks before anything is installed

`helm install` and `helm upgrade` validate your values before they apply
anything. A failed check stops the command with what is wrong, why it matters
and what to change:

```
✗ services.telemetry.enabled=true but clickhouse.enabled=false.

Why: Agent telemetry (OpenTelemetry logs and steps) is stored only in ClickHouse; …

Fix: Either keep ClickHouse on (clickhouse.enabled: true — the default), or turn telemetry off as well: …
```

| Refused when | Because |
| - | - |
| Agent telemetry is on and ClickHouse is off | ClickHouse is telemetry's only store |
| An upgrade would create ClickHouse while audit, registry or telemetry use the worker split | Their migration Jobs would run before ClickHouse exists and hang. The message prints the flags for running that one upgrade without the split |
| `services.relay.ssrfAllowedHosts` has a scheme, port, path or space in an entry | Entries are compared with the host name only, so it would never match |
| `services.relay.ssrfAllowedHosts` allows the whole cluster (`*`, `*.svc.cluster.local`, …) | The relay could then reach Keycloak, Postgres and Swarmd's internal APIs |
| A Secret named in your values does not exist, or lacks a key (licence, SMTP, external Postgres/ClickHouse/Keycloak, or your own credentials Secret when `generatedSecrets.enabled: false`) | Pods would otherwise fail with `CreateContainerConfigError` after the install was applied, or fall back to defaults that are not secret |
| An external Postgres, Keycloak, ClickHouse, Presidio or Piiranha is missing its URL or credentials, SMTP is on without host/from/credentials, or the licence is missing | The service could not start or would silently drop work |

Checks that need to see the cluster (Secrets, an existing ClickHouse) only run
when Helm has a live connection; under `helm template` they are skipped.

Valid choices that switch features off do not stop the install, but the
install notes say what you lose: ClickHouse off (no monitors or alerts),
notification-service off (no notices or alerts delivered), SMTP off (one
account), and an empty relay allow-list (in-cluster agents refused).

After installing or upgrading, `helm test swarmd -n swarmd --logs` checks the
running install the same way (see [Upgrades › Verifying afterwards](/self-hosting/upgrades#verifying-afterwards)).
For air-gapped clusters, mirror `curlimages/curl` and set `tests.image.repository`.

### What the platform says at runtime

The services answer the same way when something is not set up:

| You do | You get |
| - | - |
| Create or enable a monitor with ClickHouse off | `409` — monitors are not evaluated without ClickHouse, and the setting to change |
| Create a *Webhook* notification destination | `400` — nothing delivers to webhooks yet; use e-mail, Slack or Teams |
| Send an agent's telemetry with telemetry off | `404 TELEMETRY_NOT_DEPLOYED` from the gateway, naming `services.telemetry.enabled` |
| Send telemetry while telemetry-service is down | `503 TELEMETRY_UNAVAILABLE` |
| Message an agent on a private address that is not allowed | "The relay is not allowed to reach this agent's address"; the relay log names the host and `services.relay.ssrfAllowedHosts` |

***

## Next

<CardGroup cols={2}>
  <Card title="Presets" icon="layer-group" href="/self-hosting/presets">
    Tested values files that combine all of this.
  </Card>

  <Card title="Databases" icon="database" href="/self-hosting/databases">
    Layouts, external Postgres, ClickHouse and backups.
  </Card>

  <Card title="Upgrades" icon="arrow-up-right-dots" href="/self-hosting/upgrades">
    Moving between chart versions without taking the install down.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.