Skip to main content

Configuration

Everything that isn’t the database layout or ingress: how settings reach the services, where secrets come from, and how to substitute your own infrastructure.

Your own values file

Whatever we send you — the plain chart, or a values file sized for your cluster — is a starting point, not something to edit in place. Keep your overrides in a file you own, in your own git repo, and layer it on top.
Later -f wins, and --set beats every -f. You only need to write the keys you are changing; everything else falls through to the file underneath and then to the chart’s defaults. That separation is what makes the next upgrade cheap. When we send a new baseline, you replace one file and your overrides are untouched — instead of diffing your edits back out of a file that has moved on.

Sizing one service

services.<name>.resources merges onto services.defaults.resources, so you can raise one number without restating the block:

How the layering actually merges

Setting one key leaves its siblings alone:
Result: LOG_LEVEL: debug, TZ: UTC. Nesting goes as deep as you like.
This is the one that surprises people:
Result: [{ name: c }]. a and b are gone — not appended to.Any list is affected: pullSecrets, tolerations, ingress.hosts. If you mean to add to a list, restate the whole thing including the entries you want to keep.
Omitting a key inherits it. Setting it to {}, [] or null overrides it with emptiness — which is how you deliberately clear something the baseline set, and also how people accidentally wipe a block they meant to leave alone.
Helm does not carry values forward between releases. An upgrade that omits a -f you passed at install time reverts those settings to the chart defaults — silently, because from Helm’s point of view you asked for the defaults.Pass the same -f files on every helm upgrade. --reuse-values looks like the shortcut and is not one: it also skips new defaults introduced by the chart you are upgrading to, so a fresh install and an upgraded install end up configured differently. Keep the files, pass them every time.

Checking what is actually applied

The second is the real answer to “what is this cluster running with”, and the first thing worth looking at when a setting seems not to have taken.
Preview before applying: helm diff upgrade (from the helm-diff plugin) or helm upgrade --dry-run=server. Never helm template against a live install — see Upgrades.

Application settings

Anything that would otherwise be a bespoke env var — encryption keys, timeouts, log levels, model config — goes under appSettings, and lands as an env var on every service. Each entry takes one of three shapes:
These render to native env: entries — scalars as value:, maps as valueFrom.secretKeyRef or valueFrom.configMapKeyRef. A map that is neither shape fails the install rather than rendering something odd.

Per-service overrides

services.<name>.settings takes the same three shapes and wins over appSettings for the same key:
The defaults ship ENCRYPTION_KEY and AUDIT_HASHCHAIN_SIGNING_KEYS wired to the chart-generated Secret, plus LOG_LEVEL=info — enough for a first install to work with no input. Override them with your own key management when you go to production.

Secrets

The chart generates a Secret named swarmd-generated-credentials at install: Passwords are 24 random alphanumeric characters. On helm upgrade the chart reads the existing Secret via lookup and reuses the values — upgrades do not rotate credentials. The Secret carries helm.sh/resource-policy: keep, so helm uninstall leaves it behind.
Back up encryption-key separately from your database. It encrypts data at rest in registry, relay and teams. A database restore without the matching key leaves you with rows nobody can decrypt.

Supplying your own

Then pre-create a Secret of that name with the keys above. Useful when your platform team owns secret material, or you sync from Vault or OpenBao with an operator.

SMTP

tenant-auth sends email for verification, password reset and team invites. With smtp.enabled: false — the default — every one of those is logged and dropped. Users can register, but nothing arrives.
Without SMTP you have a one-account deployment. Sign-up and login work, so the first admin gets in fine — but that is as far as it goes:Fine for a solo evaluation. Not fine for anything a team touches, and not something to discover after handing the platform over.

Nobody is locked out, though

smtp.enabled also sets the realm’s verifyEmail flag, so with SMTP off Keycloak never raises the “verify your email” required action and sign-up followed by login works normally. Verification is genuinely off, not skipped by a special case: tenant-auth does refuse a login for an unverified user, but only when Keycloak reports the account as not fully set up, which cannot happen while verifyEmail is false. Those users are still recorded as unverified, so they show as PENDING in user listings. That is cosmetic — it gates nothing.
The install refuses to render with smtp.enabled: true and any of host, from or credentialsSecret missing.
Turning SMTP on later does not start enforcing email verification. verifyEmail is written into the realm when the realm is created, and the bootstrap Job is create-only — on every subsequent install and upgrade it finds the realm, logs already exists — nothing to do, and changes nothing. So a deployment that first came up without SMTP keeps verifyEmail: false even after you enable SMTP and upgrade, and existing unverified users carry on logging in.Verified on a live install: enable SMTP, upgrade, and an account created before the change still logs in without verifying.Enabling SMTP does make invites and password resets start working, which is usually the reason for turning it on. If you also want verification enforced, change it on the realm yourself through Keycloak’s admin console or API — the chart cannot do it for you without overwriting realm customisations you may have made.
swarmd.tenant-auth.auto-verify-email-patterns marks matching addresses verified at creation — a comma-separated list of globs, e.g. *@yourcompany.com. Useful once verification is enforced but you want your own domain to skip it. Empty by default, so nobody is auto-verified.

Bring your own infrastructure

Every optional component follows the same three-state pattern: enabled: false · enabled: true, deploy: true · enabled: true, deploy: false plus an external block. The chart fails the render if you choose the third and leave a required URL empty.

Bring your own Keycloak

adminSecret holds master-realm admin credentials. Registry, relay, tenant-auth and teams call the Admin API to manage agent service accounts, so this is not optional.
External Postgres forces external Keycloak. keycloak.deploy=true with postgres.deploy=false is unsupported and fails the render — the chart can’t derive host and port from an arbitrary JDBC URL for the Keycloak container.

PII detection

Both analyzers are off by default. With neither enabled, prompts pass through the LLM path unfiltered.
Piiranha’s analyzer image sits outside the release set. It is a third-party model image, not built or versioned with the platform, and it is not covered by the testing that goes into a Swarmd release. It is also large — roughly 1.3 GiB — so first pull is slow and it wants a node with room for the model in memory.Run it yourself and point the chart at it:
If you would rather have it in-cluster, set piiranha.image.repository to your own copy so you control what version you are running.Presidio has no such constraint — its image comes from Microsoft’s public registry, so deploy: true is fine.

a2a-payments

Off by default, and the only service that genuinely cannot start without configuration: its X402_WALLET_PRIVATE_KEY is @NotBlank-validated with no degraded mode.

Turning services off

Gateway, registry, relay, audit and tenant-auth are the core — disabling any of them gives you a platform that doesn’t work.

Settings that bite later

Some settings don’t block startup but fail the first time a feature is used. The chart can’t enforce them, because the service boots fine without them. Provide them under services.<name>.settings when you enable the feature.

Scaling: core and worker

Every DB-backed service ships as one image that can run three ways, and the chart splits every one of them by default: The worker’s web server stays up so health probes work, but the Service never routes traffic to it. Worker sizing falls back to the service’s own replicaCount / resources when left empty.
Why this is the default. In the combined topology, scaling the API to N replicas also runs N copies of every @Scheduled job — outbox drainers and monitor evaluators racing each other, which is a data problem rather than a performance one. Splitting also stops a slow scheduler starving the API’s thread pool, and takes Flyway out of pod startup so a long migration can’t trip readiness on every replica at once.

Collapsing a service back

Worth doing on a laptop, where the extra pod per service costs more than the isolation buys you. Setting it on every service takes the default install from 18 Deployments to 11.
Don’t run a collapsed service at replicaCount > 1. That is exactly the scheduler-duplication case the split exists to prevent.

Startup ordering

Every service talks to Keycloak while it boots — the Admin API for service-account management, the JWK set for token validation — and exits non-zero if it cannot reach it. On a cold install that is a race the services lose, so each one is held at a wait-for-keycloak init container until Keycloak’s realm answers its OIDC discovery endpoint. You will see pods sit in Init:0/1 for roughly a minute on a first install. That is the gate working; expect zero restarts.
Set waitForKeycloak: false per service (or on defaults) if you sequence startup another way. Migration Jobs are never gated — they run Flyway and never touch Keycloak.
The gate checks the realm’s discovery document, not the TCP port. Keycloak accepts connections well before the realm-bootstrap Job has imported the realm, and a service starting in that window fails exactly as it would have without the gate.

JVM sizing

The chart sizes the JVM against the container memory limit:
Do not remove this while lowering the memory limit. Without MaxRAMPercentage the JVM falls back to 25% of the limit — a 512Mi limit gives a 128Mi heap, which Spring Boot with ~40 JPA repositories cannot start in. The symptom is a container that reaches “Starting service [Tomcat]” and is then OOMKilled (exit 137) in a crash loop.
On a running install these services sit at 593–697 MiB working set, which is why the default limit is 1Gi and the default request 512Mi. If you raise the limit, the heap grows with it; if you lower it below ~768Mi, lower MaxRAMPercentage too or the JVM will size a heap the container cannot hold. No garbage collector is pinned deliberately — JVM ergonomics picks one suited to the container, which matters more here than in a large deployment because this chart is routinely run with small limits.

Observability

Prometheus scrape annotations land on every service pod; tracing points the OTLP exporter at your collector.

Verifying a change before you apply it

$SWARMD_CHART is the chart file we sent you — see Quickstart → Step 2.
The chart’s validation runs during template, so a bad combination fails here — before it touches your cluster.
helm template renders without a cluster connection, which is why the --set licence.key=LIC-TEST above is needed and why the output is only good for checking your values. Never apply a helm template render to a live install: the chart cannot read back the credentials Secret in that mode and the output carries freshly generated passwords. Use helm upgrade --dry-run=server to preview a real change — see Upgrades.
Unpacked the tarball (tar xzf)? Both commands work against the ./swarmd directory the same way.

Next

Presets

Tested values files that combine all of this.

Databases

Layouts, external Postgres, ClickHouse and backups.

Upgrades

Moving between chart versions without taking the install down.