> ## Documentation Index
> Fetch the complete documentation index at: https://docs.swarmd.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Upgrades

> Moving between chart versions without taking your install down — what is preserved, what to check first, and how to roll back.

# Upgrades

An upgrade is one command. This page is about the handful of things that make
that command safe to run against an install with real data in it.

```bash theme={null}
# The version you are moving TO — the newer tarball we sent you.
export SWARMD_CHART=~/Downloads/swarmd-<new-version>.tgz

helm upgrade swarmd $SWARMD_CHART -n swarmd \
  -f my-values.yaml \
  --wait --timeout 15m
```

***

## Getting the new version

We send you a new `swarmd-<version>.tgz` for each release, with notes covering
what changed and anything that needs care on the way to it. There is no
registry to poll and no `--version` flag: **the file is the version.**

That has a couple of practical consequences.

* **Keep old tarballs.** `helm rollback` restores the previous *revision* from
  what Helm already has, so it works without the file. Reinstalling an older
  release from scratch does not — that needs the tarball again, and asking us
  for it is a slower path than keeping a copy.
* **Name your releases after the file.** `helm -n swarmd list` shows the chart
  version, so a cluster can always tell you what it is running; keeping the
  tarballs in one directory makes it just as easy in the other direction.

<Note>
  Read the release notes for every version between where you are and where you
  are going — not just the target. Helm applies one jump; the notes are written
  per version.
</Note>

Anything a particular version needs from you on the way through is listed
under [Version notes](#version-notes) at the end of this page.

<Tip>
  Check the file before you run it against a live install:

  ```bash theme={null}
  helm show chart $SWARMD_CHART
  ```

  It prints the name, version and appVersion in two lines, which is the quickest
  way to catch a stale download or the wrong file in `~/Downloads`.
</Tip>

***

## What survives an upgrade

| | Preserved? | Notes |
| - | - | - |
| Generated credentials | ✅ | The chart reads the existing Secret and reuses it. See [the rotation guard](#the-one-that-bricks-installs). |
| Postgres data | ✅ | The PVC is untouched. Schema changes run as migration Jobs, or at pod start when a service runs without a worker. |
| Keycloak realm | ✅ | The bootstrap Job checks whether the realm exists and exits early. Customisations you made through the Admin UI are never overwritten. |
| ClickHouse data | ✅ | PVC untouched. |
| Your `values.yaml` | ⚠️ | Only what you pass. Helm does not merge in the values from the previous release unless you use `--reuse-values`, which has its own surprises — pass the same `-f` file every time instead. |

***

## The one that bricks installs

The chart generates `swarmd-generated-credentials` on first install and, on
every upgrade, reads the existing Secret back so nothing rotates. That read is
a Helm `lookup`, and **`lookup` only works when Helm has a live connection to
the cluster.**

When it doesn't, every credential falls through to a freshly generated random
value. Applying that render rotates the Postgres password, the Keycloak admin
password and `encryption-key` in a single step. The first two stop the
platform; the third is worse, because registry, relay and teams encrypt rows
at rest with that key and **a rotated key makes existing rows permanently
unreadable**. No backup of the database helps — you need the old key.

The chart now refuses rather than letting that happen:

```
Refusing to upgrade: could not read the existing Secret
"swarmd-generated-credentials" in namespace "swarmd".
```

Two situations produce it.

<AccordionGroup>
  <Accordion title="helm upgrade --dry-run">
    Plain `--dry-run` is client-side and disables `lookup`. Use
    `--dry-run=server` when you want to preview an upgrade — it keeps the
    cluster connection, so the Secret is read and the render matches what a
    real upgrade would apply.

    ```bash theme={null}
    helm upgrade swarmd $SWARMD_CHART -n swarmd \
      -f my-values.yaml --dry-run=server
    ```
  </Accordion>

  <Accordion title="GitOps that renders with `helm template`">
    Argo CD renders charts with `helm template` by default and applies the
    output. `lookup` returns nothing in that mode, and because the render
    reports itself as an *install* the guard cannot tell it apart from a
    genuine first install — so it does not fire, and the rotation happens
    silently.

    **If you deploy through a template-and-apply GitOps tool, do not let the
    chart own its credentials.** Turn generation off and manage the Secret
    yourself:

    ```yaml theme={null}
    global:
      generatedSecrets:
        enabled: false
        name: swarmd-generated-credentials
    ```

    Then pre-create a Secret of that name with the keys listed in
    [Configuration → Secrets](/self-hosting/configuration#secrets), sourced
    from your secret manager (Vault, OpenBao, External Secrets). This is the
    right shape for GitOps anyway — the credentials stop being a side effect
    of a render.
  </Accordion>
</AccordionGroup>

<Warning>
  **Back up `encryption-key` before your first upgrade**, separately from your
  database, and keep it somewhere a cluster failure cannot take with it.

  ```bash theme={null}
  kubectl -n swarmd get secret swarmd-generated-credentials \
    -o jsonpath='{.data.encryption-key}' | base64 -d
  ```
</Warning>

***

## Before you run it

<Steps>
  <Step title="Take a database backup">
    Migrations run forwards only. There is no `down` script, so a rollback of
    the chart does not roll back the schema — see [Rolling back](#rolling-back).
    For the single in-cluster Postgres (`schema-per-service` with
    `sharedDatabase: swarmd`):

    ```bash theme={null}
    PGUSER=$(kubectl -n swarmd get secret swarmd-generated-credentials -o jsonpath='{.data.postgres-username}' | base64 -d)
    PGPASS=$(kubectl -n swarmd get secret swarmd-generated-credentials -o jsonpath='{.data.postgres-password}' | base64 -d)
    for db in swarmd keycloak; do
      kubectl -n swarmd exec deploy/swarmd-postgres -- \
        env PGPASSWORD="$PGPASS" pg_dump -U "$PGUSER" -Fc "$db" > "$db-before-upgrade.dump"
    done
    ```

    Note the current revision too — it is the one to
    [roll back](#rolling-back) to:

    ```bash theme={null}
    helm -n swarmd history swarmd --max 1
    ```

    Other layouts: see [Databases › Backups](/self-hosting/databases).
  </Step>

  <Step title="Back up the keys">
    ```bash theme={null}
    kubectl -n swarmd get secret swarmd-generated-credentials -o yaml > swarmd-credentials-backup.yaml
    ```

    Keep it outside the cluster. A database restore is useless without the
    matching `encryption-key` / `encryption-keys`.
  </Step>

  <Step title="Confirm you can read the credentials Secret">
    ```bash theme={null}
    kubectl -n swarmd get secret swarmd-generated-credentials
    ```

    If this is missing, stop. Restore it before upgrading.
  </Step>

  <Step title="Preview the change">
    ```bash theme={null}
    helm diff upgrade swarmd $SWARMD_CHART -n swarmd \
      -f my-values.yaml
    ```

    `helm diff` is a plugin (`helm plugin install https://github.com/databus23/helm-diff`)
    and is the single most useful thing you can install for this. Failing
    that, `--dry-run=server` and read the output.
  </Step>

  <Step title="Check you have headroom for the ECR pull">
    The licence bootstrap Job refreshes the image-pull Secret as a
    `pre-upgrade` hook, so an expired registry token is not a problem. A
    licence that has *expired* is — the Job fails and the upgrade stops before
    touching any workload. Check the expiry on your licence first.
  </Step>
</Steps>

***

## What happens during the rollout

Services are replaced in place rather than surged: `maxSurge: 0`,
`maxUnavailable: 1`.

That means a single-replica service is **briefly unavailable** while its pod
is replaced — roughly 40–70 seconds for a JVM service. It is deliberate. The
Kubernetes default starts a second pod before retiring the first, which
doubles the memory every service wants at the same moment; on a node sized for
the steady state the kubelet starts OOMKilling, and it kills the *old* pods
too. A routine upgrade takes down a working install.

If you need a genuinely zero-downtime rollout, run more than one replica and
put the surge back on that service — and size the node for the surge:

```yaml theme={null}
services:
  gateway:
    replicaCount: 2
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
```

Expect to see, in order:

1. Migration Jobs run to completion (`pre-upgrade` hooks, Flyway, forwards only).
   A service with `worker.enabled: false` — every service in
   `values-prismforce-local.yaml` — has no Job and migrates as its pod starts
   instead.
2. Pods replaced one service at a time, each held at `Init:0/1` until
   Keycloak's realm answers.
3. The realm bootstrap Job for the new revision runs and exits early, because
   the realm already exists.

**Zero restarts is the expected outcome.** A pod that restarts during an
upgrade is a signal, not noise — check `kubectl -n swarmd describe pod` for
exit code 137 (OOMKilled) and see [JVM sizing](/self-hosting/configuration#jvm-sizing).

***

## Rolling back

First find the revision you are going back to. **Always name it.** Without a
revision, `helm rollback` goes to the *previous* revision, and that is only the
version you came from if nothing has been applied since. After a first
rollback attempt, or a second `helm upgrade` to change a value, "previous" is
the new version again.

```bash theme={null}
helm -n swarmd history swarmd
# REVISION  CHART         APP VERSION  DESCRIPTION
# 1         swarmd-0.3.3  0.3.3        Install complete
# 2         swarmd-0.4.2  0.4.2        Upgrade complete

helm rollback swarmd 1 -n swarmd --wait --timeout 15m
```

This restores that chart revision — image tags, resource limits, settings,
all of it.

<Warning>
  **It does not roll back the database.** Flyway migrations are forwards-only,
  so after a rollback the schema is still the new one while the images are the
  old ones. That is fine for additive migrations and is not fine for anything
  that dropped or renamed a column, or for data the older version cannot read.

  Treat `helm rollback` on its own as the fast path for a configuration mistake
  or a bad image *within the same version*, and your database backup as the
  recovery path when you are going back a version. [Version notes](#version-notes)
  call out every version where a rollback needs a restore.
</Warning>

### Rolling back with a restore

Run this from the directory holding the dumps you took
[before upgrading](#before-you-run-it), straight after the `helm rollback`
above. It is self-contained: it reads the Postgres credentials itself, and it
puts every workload back at the replica count it had.

```bash theme={null}
NS=swarmd
PGUSER=$(kubectl -n $NS get secret swarmd-generated-credentials -o jsonpath='{.data.postgres-username}' | base64 -d)
PGPASS=$(kubectl -n $NS get secret swarmd-generated-credentials -o jsonpath='{.data.postgres-password}' | base64 -d)
pg() { kubectl -n $NS exec -i deploy/swarmd-postgres -- env PGPASSWORD="$PGPASS" "$@"; }

# 1. Remember each workload's replica count, then stop everything but Postgres.
REPLICAS=$(kubectl -n $NS get deploy -l app.kubernetes.io/instance=swarmd \
  -o jsonpath='{range .items[*]}{.metadata.name}={.spec.replicas}{" "}{end}')
kubectl -n $NS scale deploy -l app.kubernetes.io/instance=swarmd --replicas=0
kubectl -n $NS scale deploy/swarmd-postgres --replicas=1
kubectl -n $NS rollout status deploy/swarmd-postgres --timeout=5m

# 2. Recreate each database empty and restore into it.
for db in swarmd keycloak; do
  pg dropdb -U "$PGUSER" --if-exists --force "$db"
  pg createdb -U "$PGUSER" -O "$PGUSER" "$db"
  pg pg_restore -U "$PGUSER" -d "$db" --exit-on-error --single-transaction < "$db-before-upgrade.dump"
done

# 3. Start everything again.
for r in $REPLICAS; do kubectl -n $NS scale "deploy/${r%%=*}" --replicas="${r##*=}"; done
kubectl -n $NS rollout status deploy -l app.kubernetes.io/instance=swarmd --timeout=15m
```

<Warning>
  **Restore into an empty database, never over the upgraded one.** An earlier
  version of this page used `pg_restore --clean`. That cannot drop tables the
  newer version added when they reference the older ones, so it skips them,
  then appends the backup's rows into tables that were never emptied: agents
  appear twice, primary keys are lost, and the registry answers 500 for every
  agent. `--exit-on-error --single-transaction` makes any such failure stop the
  restore with nothing applied, instead of finishing with errors you might not
  see.
</Warning>

Anything created after the upgrade is lost, which is why it is worth upgrading
in a quiet window and checking the result straight away. Then run the checks
in [Verifying afterwards](#verifying-afterwards) again.

What is *not* in the database dumps, and what happens to it:

* **`swarmd-generated-credentials`.** The rollback puts the Secret back as the
  older release wrote it, so keys the newer version added are removed. That
  is safe: `encryption-keys` is derived from `encryption-key`, and the
  per-service Keycloak client secrets are regenerated and re-applied to
  Keycloak by each service at start, so upgrading again rebuilds both. The
  keys the older version uses are never changed.
* **ClickHouse.** From 0.4.2 its volume is kept when a rollback removes
  ClickHouse, and its data is still there when you upgrade again. Rolling back
  *from 0.4.1* deletes it.

***

## Skipping versions

Upgrading across several versions in one jump works — migrations are
cumulative and run in order. What it costs you is the ability to tell which
version broke something.

On an install carrying data you care about, step through one minor version at
a time and let each one settle. On a rebuildable environment, jump.

***

## Verifying afterwards

```bash theme={null}
# Every pod Running, zero restarts
kubectl -n swarmd get pods

# Migration Jobs completed, not failed
kubectl -n swarmd get jobs

# The release is at the version you asked for
helm -n swarmd list
```

Then run the chart's own checks against the running install:

```bash theme={null}
helm test swarmd -n swarmd --logs
```

It checks every enabled service is healthy, Keycloak's realm answers, service
tokens carry the audience the platform requires, the gateway reports telemetry
correctly, and the UI serves its sign-in page. Each failure says what is wrong,
why it matters and what to do. It is read-only and safe to run at any time.

Then exercise one real path — log in to the UI, or call an agent through the
gateway. Pods being `Running` only tells you the JVMs started.

***

## Version notes

### From 0.3 to 0.4

Go straight from 0.3.x to the latest 0.4 release.

<Steps>
  <Step title="Start from the 0.4 values file">
    0.4 turns on `notification-service` (governance notices and alerts depend
    on it), ClickHouse (monitors, alerts, LLM traffic figures and agent
    telemetry depend on it) and agent telemetry. All three are on by default
    from 0.4.1. A values file that turns any of them off still upgrades, but
    those features quietly do nothing. Re-apply your own changes on top and compare:

    ```bash theme={null}
    diff my-0.3-values.yaml values-prismforce-local.yaml
    ```
  </Step>

  <Step title="Tell the relay where your agents are">
    If your agents or MCP servers run in your cluster or on a private network,
    list them, or the relay refuses them with *Blocked request to
    private/internal address*:

    ```yaml theme={null}
    services:
      relay:
        ssrfAllowedHosts: "*.my-agents.svc.cluster.local"
    ```

    See [Agents and MCP servers on private addresses](/self-hosting/configuration#agents-and-mcp-servers-on-private-addresses).
  </Step>

  <Step title="Plan for the new pods">
    `swarmd-notification`, `swarmd-clickhouse` and `swarmd-telemetry` take the
    PrismForce layout to about 6 GiB of memory requests (a 12 GiB node) and
    add a 20 GiB ClickHouse volume.
  </Step>

  <Step title="Split installs: turn the split off for this one upgrade">
    If your services run with the worker split (`services.*.worker.enabled:
            true`, the chart default) and ClickHouse was off on 0.3, the chart refuses
    the upgrade before it starts: the migration Jobs would run before
    ClickHouse exists and hang. The message prints the flags to add — run this
    upgrade with `worker.enabled: false` for audit, registry and telemetry,
    then put the split back in your next upgrade.
    `values-prismforce-local.yaml` has no worker split and is not affected.
  </Step>

  <Step title="After the upgrade">
    * **Re-save stored credentials.** Before 0.4.1, stored credentials (LLM
      provider keys, the API keys, bearer tokens and OAuth secrets used to call
      agents and MCP servers) were encrypted with a built-in default key
      instead of your install's own. 0.4.1 encrypts new values with your key and
      still reads the old ones; re-entering them re-encrypts them.
    * **Ask users to sign in again** — a web UI session-isolation fix ships in
      0.4.
    * **Review monitors on `ERROR` events.** Failed LLM calls now count as
      errors.
    * **Custom channel integrations:** the Art. 50 AI notice, once you switch
      it on for an agent, arrives as its own message. See
      [AI disclosure](/sdks/typescript/ai-disclosure).
    * Notification channels deliver by e-mail, Slack or Microsoft Teams.
      *Webhook* destinations have no sender yet and can no longer be created.
  </Step>
</Steps>

<Warning>
  **Rolling back to 0.3 needs a database restore** — use
  [Rolling back with a restore](#rolling-back-with-a-restore). A `helm rollback`
  on its own leaves 0.3 running on the 0.4 schema, and:

  * evaluation sets cannot be listed or run: relay migration V053 drops
    `verdicts_simulated`, which the 0.3 relay still reads (HTTP 500);
  * tenant admins cannot load their own profile (`GET /tenant-auth/v1/me` answers
    400, `No enum constant … EntityType.TELEMETRY`), because 0.4 grants a
    permission 0.3 does not know;
  * credentials saved after the upgrade are encrypted with a key 0.3 does not
    use;
  * rolling back from 0.4.1 deletes the ClickHouse volume (0.4.2 keeps it).
</Warning>

Also fixed in 0.4.1, applied automatically by the upgrade: agent and service
tokens carry the audience registry and relay require (agents calling the
platform API were refused on 0.3), and each service's Keycloak client gets its
own generated secret instead of a built-in default.

Fixed in 0.4.2:

* **No outage straight after the upgrade.** On 0.4.1 a service could fetch its
  token while its Keycloak client was still being updated, then use that token
  for up to four minutes; on an upgrade from 0.3 every message to an agent
  failed with `Validation failed: [401]` until it expired. Services now discard
  that token once their client is complete, and replace any service token a
  platform API refuses.
* **`helm test` passes without SMTP.** Notification's mail health check is
  switched off when `smtp.enabled` is false; it reported the service as
  unhealthy on every install that had not configured SMTP yet.
* **Audit trails no longer contain credentials.** `Authorization`, `Cookie` and
  API-key headers are recorded as `[REDACTED]`. Events recorded before 0.4.2
  keep what they captured — audit events are append-only — but the tokens in
  them expired within five minutes of being issued.
* **The ClickHouse volume survives a rollback** (see above).

***

## Next

<CardGroup cols={2}>
  <Card title="Configuration" icon="sliders" href="/self-hosting/configuration">
    Secrets, JVM sizing, and the core/worker split.
  </Card>

  <Card title="Databases" icon="database" href="/self-hosting/databases">
    Layouts, external Postgres and backups.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.