> ## Documentation Index
> Fetch the complete documentation index at: https://docs.swarmd.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent telemetry

> Turn on telemetry-service so agents can send their own OpenTelemetry logs and steps, stored in your ClickHouse beside the audit trail.

# Agent telemetry

`telemetry-service` takes agents' own OpenTelemetry log lines and spans and
stores them in ClickHouse, keyed by the same correlation id as the audit trail.
The trace view then shows what each agent logged, and the steps it took, on the
same timeline as the policy decisions. The developer side is described in
[Agent telemetry (Python SDK)](/sdks/python/telemetry).

It is a **preview** in 0.4.x. The service is on by default; nothing is
collected until an agent opts in.

<Note>
  **Nothing leaves your cluster.** Agents post to your gateway, the gateway
  forwards to `telemetry-service`, and it writes to your ClickHouse. It calls no
  other service and nothing at Swarmd.
</Note>

***

## Before you start

| Need | Why |
| - | - |
| ClickHouse (`clickhouse.enabled`, on by default) | It is the only store. The chart refuses telemetry without it. |
| About 2 GiB more memory | ClickHouse (512 Mi–2 Gi) plus one telemetry pod (\~320 Mi requested). |
| Disk for 30 days of logs | Retention is fixed at 30 days from ingest. Size the ClickHouse volume for your agents' log volume, not just audit history. |
| Agents on `swarmd-google-adk` / `swarmd-sdk` 0.4.0+ | Older SDKs do not send telemetry. |

**External ClickHouse** works the same way as for audit history: it must
define a cluster named `swarmd_cluster` and the `{shard}`/`{replica}` macros,
because the tables are `ReplicatedMergeTree ... ON CLUSTER swarmd_cluster`. The
bundled ClickHouse already does this.

***

## Turn it on

**It is on by default** from 0.4.1, together with ClickHouse, so on most
installs there is nothing to do on the cluster: go straight to
[turning it on in your agents](#turn-it-on-in-your-agents).

### If ClickHouse is off

ClickHouse is telemetry's only store, so the chart refuses
`services.telemetry.enabled: true` with `clickhouse.enabled: false` before
installing anything, and says why. Either turn ClickHouse on (the default),
or turn telemetry off as well:

```yaml theme={null}
services:
  telemetry:
    enabled: false
```

With telemetry off, `/telemetry/*` on the gateway answers `404` with code
`TELEMETRY_NOT_DEPLOYED` and a message naming these settings, and the trace
view says telemetry is not enabled, rather than either reporting an error. If
telemetry is on but its pods are down, the gateway answers `503
TELEMETRY_UNAVAILABLE` instead.

### Turn it on in your agents

Set `SWARMD_TELEMETRY_ENABLED=true` and, for steps,
`SWARMD_TELEMETRY_STEPS=true` on each agent and restart it. See the
[SDK page](/sdks/python/telemetry). Every agent, including ones registered
before the upgrade, already has `TELEMETRY:WRITE`.

<Note>
  **An upgrade that would create ClickHouse on a split install is refused.**
  With the worker split (`services.*.worker.enabled: true`), migrations run in
  `pre-upgrade` Jobs, which Helm runs before it creates ClickHouse, so they
  would wait for a ClickHouse that does not exist yet. The chart stops the
  upgrade before it starts and prints the exact flags to use: run that one
  upgrade with `worker.enabled: false` for audit, registry and telemetry (they
  then migrate when their pods start, after ClickHouse is ready), and put the
  split back in your next upgrade. Installs without the split, such as
  `values-prismforce-local.yaml`, are not affected.
</Note>

***

## The values

| Key | Default | Meaning |
| - | - | - |
| `services.telemetry.enabled` | `true` | Deploy the service (only when ClickHouse is on). |
| `services.telemetry.captureEnabled` | `true` | Register the ingest and read endpoints. Set it back to `false` to stop accepting telemetry without uninstalling anything; the pods roll and keep running. |
| `services.telemetry.worker.enabled` | `true` | Core + worker + migrate Job, like the other services. Set `false` on single-node installs. |
| `services.telemetry.resources` | chart default | Requests/limits for the pod. |
| `services.telemetry.settings.TELEMETRY_MAX_BODY_CHARS` | `65536` | Largest log body accepted. Larger records are **rejected, not truncated**. |
| `services.telemetry.settings.TELEMETRY_MAX_ATTRIBUTE_CHARS` | `16384` | Largest serialised attribute map per record or span. |
| `services.telemetry.settings.TELEMETRY_MAX_RECORDS_PER_REQUEST` | `10000` | Records past this in one request are rejected. |

When telemetry is enabled the chart points the gateway at it. When it is not,
platform-ui is told so and the trace view says "Agent telemetry is not enabled
on this deployment" instead of reporting an error.

***

## The endpoints

Agents reach it through the gateway, on the same host as the rest of the API:

| Method and path | Auth | What |
| - | - | - |
| `POST /telemetry/v1/logs` | Agent token, `TELEMETRY:WRITE` | OTLP/HTTP **JSON** log export |
| `POST /telemetry/v1/traces` | Agent token, `TELEMETRY:WRITE` | OTLP/HTTP **JSON** trace export |
| `GET /telemetry/v1/logs` | User token, `TELEMETRY:READ` | Read back log lines |
| `GET /telemetry/v1/spans` | User token, `TELEMETRY:READ` | Read back steps |

* **JSON only.** No protobuf and no gRPC. If you put your own OpenTelemetry
  Collector in front, use the `otlphttp` exporter with `encoding: json` and
  forward the agent's bearer token.
* **Tenant comes from the token,** never from the payload. A token without a
  tenant is refused.
* Ingest always answers `200`. Records it refuses (too large, or a span without
  a trace or span id) are counted in the OTLP `partialSuccess` reply. The rest
  of the batch is stored.

***

## Who can read it

`TELEMETRY:READ` is granted to the **Tenant Administrator** group only. Editors
and Viewers see "You do not have permission to read telemetry" on traces until
you add it to their group under **Manage › Groups**. Agents can write but can
never read.

***

## Limits of the preview

* **No redaction of reasoning content.** If an agent sets
  `SWARMD_TELEMETRY_REASONING=true`, full prompts and completions are stored.
  Keep it to test tenants.
* **No per-tenant quotas, sampling or rate limit.** A noisy agent can fill the
  ClickHouse volume. Watch its disk. A full ClickHouse volume also stops audit
  history and schema migrations.
* **Retention is fixed at 30 days** from ingest time. An agent's clock cannot
  shorten or lengthen it.
* **Logs and spans only.** No metrics.
* **Python only.** Automatic capture is for Google ADK agents. LangChain agents
  can send log lines by calling the SDK, and the TypeScript SDK does not send
  telemetry.

***

## Troubleshooting

<AccordionGroup>
  <Accordion title="Traces say telemetry is not enabled">
    `services.telemetry.enabled` is false in the values the release was
    rendered with (`helm -n swarmd get values swarmd -a`). The gateway says
    the same to agents: `/telemetry/*` answers `404 TELEMETRY_NOT_DEPLOYED`.
  </Accordion>

  <Accordion title="Telemetry is enabled but nothing arrives">
    Check the agent first. It logs `[telemetry] shipping logs for <name> to <url>`
    at startup when telemetry is on, and `[telemetry] off` or `runtime not configured`
    when it is not. The URL must be your gateway (`SWARMD_BASE_URL`). If the
    agent's own `logger.info()` lines are missing but warnings arrive, raise
    `LOG_LEVEL` to `info` (SDK 0.4.0+).
  </Accordion>

  <Accordion title="Some steps are missing">
    A step whose attributes are larger than `TELEMETRY_MAX_ATTRIBUTE_CHARS` is
    rejected. This mostly happens to `call_llm` steps with reasoning content on.
    Raise the limit in `services.telemetry.settings`, or leave reasoning off.
  </Accordion>

  <Accordion title="The telemetry migrate Job never completes">
    It is waiting for ClickHouse. Either ClickHouse is being created in the
    same upgrade on a split install and the chart did not hold it back (see
    [Turn it on](#turn-it-on)), or an
    external ClickHouse lacks the `swarmd_cluster` definition.
  </Accordion>
</AccordionGroup>

***

## Next

<CardGroup cols={2}>
  <Card title="SDK side" icon="python" href="/sdks/python/telemetry">
    The variables agents set, and what they capture.
  </Card>

  <Card title="Databases" icon="database" href="/self-hosting/databases">
    ClickHouse alongside Postgres, and backups.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.