Skip to main content

Agent telemetry

telemetry-service takes agents’ own OpenTelemetry log lines and spans and stores them in ClickHouse, keyed by the same correlation id as the audit trail. The trace view then shows what each agent logged, and the steps it took, on the same timeline as the policy decisions. The developer side is described in Agent telemetry (Python SDK). It is a preview in 0.4.x. The service is on by default; nothing is collected until an agent opts in.
Nothing leaves your cluster. Agents post to your gateway, the gateway forwards to telemetry-service, and it writes to your ClickHouse. It calls no other service and nothing at Swarmd.

Before you start

External ClickHouse works the same way as for audit history: it must define a cluster named swarmd_cluster and the {shard}/{replica} macros, because the tables are ReplicatedMergeTree ... ON CLUSTER swarmd_cluster. The bundled ClickHouse already does this.

Turn it on

It is on by default from 0.4.1, together with ClickHouse, so on most installs there is nothing to do on the cluster: go straight to turning it on in your agents.

If ClickHouse is off

ClickHouse is telemetry’s only store, so the chart refuses services.telemetry.enabled: true with clickhouse.enabled: false before installing anything, and says why. Either turn ClickHouse on (the default), or turn telemetry off as well:
With telemetry off, /telemetry/* on the gateway answers 404 with code TELEMETRY_NOT_DEPLOYED and a message naming these settings, and the trace view says telemetry is not enabled, rather than either reporting an error. If telemetry is on but its pods are down, the gateway answers 503 TELEMETRY_UNAVAILABLE instead.

Turn it on in your agents

Set SWARMD_TELEMETRY_ENABLED=true and, for steps, SWARMD_TELEMETRY_STEPS=true on each agent and restart it. See the SDK page. Every agent, including ones registered before the upgrade, already has TELEMETRY:WRITE.
An upgrade that would create ClickHouse on a split install is refused. With the worker split (services.*.worker.enabled: true), migrations run in pre-upgrade Jobs, which Helm runs before it creates ClickHouse, so they would wait for a ClickHouse that does not exist yet. The chart stops the upgrade before it starts and prints the exact flags to use: run that one upgrade with worker.enabled: false for audit, registry and telemetry (they then migrate when their pods start, after ClickHouse is ready), and put the split back in your next upgrade. Installs without the split, such as values-prismforce-local.yaml, are not affected.

The values

When telemetry is enabled the chart points the gateway at it. When it is not, platform-ui is told so and the trace view says “Agent telemetry is not enabled on this deployment” instead of reporting an error.

The endpoints

Agents reach it through the gateway, on the same host as the rest of the API:
  • JSON only. No protobuf and no gRPC. If you put your own OpenTelemetry Collector in front, use the otlphttp exporter with encoding: json and forward the agent’s bearer token.
  • Tenant comes from the token, never from the payload. A token without a tenant is refused.
  • Ingest always answers 200. Records it refuses (too large, or a span without a trace or span id) are counted in the OTLP partialSuccess reply. The rest of the batch is stored.

Who can read it

TELEMETRY:READ is granted to the Tenant Administrator group only. Editors and Viewers see “You do not have permission to read telemetry” on traces until you add it to their group under Manage › Groups. Agents can write but can never read.

Limits of the preview

  • No redaction of reasoning content. If an agent sets SWARMD_TELEMETRY_REASONING=true, full prompts and completions are stored. Keep it to test tenants.
  • No per-tenant quotas, sampling or rate limit. A noisy agent can fill the ClickHouse volume. Watch its disk. A full ClickHouse volume also stops audit history and schema migrations.
  • Retention is fixed at 30 days from ingest time. An agent’s clock cannot shorten or lengthen it.
  • Logs and spans only. No metrics.
  • Python only. Automatic capture is for Google ADK agents. LangChain agents can send log lines by calling the SDK, and the TypeScript SDK does not send telemetry.

Troubleshooting

services.telemetry.enabled is false in the values the release was rendered with (helm -n swarmd get values swarmd -a). The gateway says the same to agents: /telemetry/* answers 404 TELEMETRY_NOT_DEPLOYED.
Check the agent first. It logs [telemetry] shipping logs for <name> to <url> at startup when telemetry is on, and [telemetry] off or runtime not configured when it is not. The URL must be your gateway (SWARMD_BASE_URL). If the agent’s own logger.info() lines are missing but warnings arrive, raise LOG_LEVEL to info (SDK 0.4.0+).
A step whose attributes are larger than TELEMETRY_MAX_ATTRIBUTE_CHARS is rejected. This mostly happens to call_llm steps with reasoning content on. Raise the limit in services.telemetry.settings, or leave reasoning off.
It is waiting for ClickHouse. Either ClickHouse is being created in the same upgrade on a split install and the chart did not hold it back (see Turn it on), or an external ClickHouse lacks the swarmd_cluster definition.

Next

SDK side

The variables agents set, and what they capture.

Databases

ClickHouse alongside Postgres, and backups.