Agent telemetry
telemetry-service takes agents’ own OpenTelemetry log lines and spans and
stores them in ClickHouse, keyed by the same correlation id as the audit trail.
The trace view then shows what each agent logged, and the steps it took, on the
same timeline as the policy decisions. The developer side is described in
Agent telemetry (Python SDK).
It is a preview in 0.4.x. The service is on by default; nothing is
collected until an agent opts in.
Nothing leaves your cluster. Agents post to your gateway, the gateway
forwards to
telemetry-service, and it writes to your ClickHouse. It calls no
other service and nothing at Swarmd.Before you start
External ClickHouse works the same way as for audit history: it must
define a cluster named
swarmd_cluster and the {shard}/{replica} macros,
because the tables are ReplicatedMergeTree ... ON CLUSTER swarmd_cluster. The
bundled ClickHouse already does this.
Turn it on
It is on by default from 0.4.1, together with ClickHouse, so on most installs there is nothing to do on the cluster: go straight to turning it on in your agents.If ClickHouse is off
ClickHouse is telemetry’s only store, so the chart refusesservices.telemetry.enabled: true with clickhouse.enabled: false before
installing anything, and says why. Either turn ClickHouse on (the default),
or turn telemetry off as well:
/telemetry/* on the gateway answers 404 with code
TELEMETRY_NOT_DEPLOYED and a message naming these settings, and the trace
view says telemetry is not enabled, rather than either reporting an error. If
telemetry is on but its pods are down, the gateway answers 503 TELEMETRY_UNAVAILABLE instead.
Turn it on in your agents
SetSWARMD_TELEMETRY_ENABLED=true and, for steps,
SWARMD_TELEMETRY_STEPS=true on each agent and restart it. See the
SDK page. Every agent, including ones registered
before the upgrade, already has TELEMETRY:WRITE.
An upgrade that would create ClickHouse on a split install is refused.
With the worker split (
services.*.worker.enabled: true), migrations run in
pre-upgrade Jobs, which Helm runs before it creates ClickHouse, so they
would wait for a ClickHouse that does not exist yet. The chart stops the
upgrade before it starts and prints the exact flags to use: run that one
upgrade with worker.enabled: false for audit, registry and telemetry (they
then migrate when their pods start, after ClickHouse is ready), and put the
split back in your next upgrade. Installs without the split, such as
values-prismforce-local.yaml, are not affected.The values
When telemetry is enabled the chart points the gateway at it. When it is not,
platform-ui is told so and the trace view says “Agent telemetry is not enabled
on this deployment” instead of reporting an error.
The endpoints
Agents reach it through the gateway, on the same host as the rest of the API:- JSON only. No protobuf and no gRPC. If you put your own OpenTelemetry
Collector in front, use the
otlphttpexporter withencoding: jsonand forward the agent’s bearer token. - Tenant comes from the token, never from the payload. A token without a tenant is refused.
- Ingest always answers
200. Records it refuses (too large, or a span without a trace or span id) are counted in the OTLPpartialSuccessreply. The rest of the batch is stored.
Who can read it
TELEMETRY:READ is granted to the Tenant Administrator group only. Editors
and Viewers see “You do not have permission to read telemetry” on traces until
you add it to their group under Manage › Groups. Agents can write but can
never read.
Limits of the preview
- No redaction of reasoning content. If an agent sets
SWARMD_TELEMETRY_REASONING=true, full prompts and completions are stored. Keep it to test tenants. - No per-tenant quotas, sampling or rate limit. A noisy agent can fill the ClickHouse volume. Watch its disk. A full ClickHouse volume also stops audit history and schema migrations.
- Retention is fixed at 30 days from ingest time. An agent’s clock cannot shorten or lengthen it.
- Logs and spans only. No metrics.
- Python only. Automatic capture is for Google ADK agents. LangChain agents can send log lines by calling the SDK, and the TypeScript SDK does not send telemetry.
Troubleshooting
Traces say telemetry is not enabled
Traces say telemetry is not enabled
services.telemetry.enabled is false in the values the release was
rendered with (helm -n swarmd get values swarmd -a). The gateway says
the same to agents: /telemetry/* answers 404 TELEMETRY_NOT_DEPLOYED.Telemetry is enabled but nothing arrives
Telemetry is enabled but nothing arrives
Check the agent first. It logs
[telemetry] shipping logs for <name> to <url>
at startup when telemetry is on, and [telemetry] off or runtime not configured
when it is not. The URL must be your gateway (SWARMD_BASE_URL). If the
agent’s own logger.info() lines are missing but warnings arrive, raise
LOG_LEVEL to info (SDK 0.4.0+).Some steps are missing
Some steps are missing
A step whose attributes are larger than
TELEMETRY_MAX_ATTRIBUTE_CHARS is
rejected. This mostly happens to call_llm steps with reasoning content on.
Raise the limit in services.telemetry.settings, or leave reasoning off.The telemetry migrate Job never completes
The telemetry migrate Job never completes
It is waiting for ClickHouse. Either ClickHouse is being created in the
same upgrade on a split install and the chart did not hold it back (see
Turn it on), or an
external ClickHouse lacks the
swarmd_cluster definition.Next
SDK side
The variables agents set, and what they capture.
Databases
ClickHouse alongside Postgres, and backups.
