Troubleshooting
Start by establishing where the path broke. The two commands that answer that:
sudo constellation-agent status # is the agent running at all?
sudo constellation-agent verify # does a metric make it all the way to topology?
verify distinguishes the two halves for you:
| Output | Broke at |
|---|---|
FAIL: telemetry socket is missing | The agent — it is not running or never created the socket. |
FAIL: metric entered the local socket but proof … was not read back | Delivery — local handoff worked, ingest did not. |
READY: telemetry proof … was read back | Nothing. The agent path is healthy; the problem is in your publisher. |
Permission denied
Symptom. Connecting to the socket raises Permission denied (EACCES).
Cause. The publishing user is not in the constellation group, or is in it but running in a session that predates the change.
# 1. Is the user in the group at all?
id -nG constellation-publisher | tr ' ' '\n' | grep -x constellation
# 2. Add it if not
sudo usermod -aG constellation constellation-publisher
# 3. The step people miss — the change is not live until a new session
sudo systemctl restart constellation-publisher
# or log out and back in for an interactive user
Group membership is resolved at process start. A running process keeps the groups it started with, so usermod alone changes nothing for it. If id -nG shows the group but your application still gets EACCES, the application has not been restarted.
Verify the socket and its directory look right:
ls -ld /run/constellation-agent # expect drwxrws--- telegraf constellation (2770)
ls -l /run/constellation-agent/telemetry.sock
No such file or directory
Symptom. Connecting raises No such file or directory (ENOENT).
Causes, in the order worth checking:
- The agent is not running.
sudo constellation-agent status. If the unit is dead,sudo systemctl restart telegrafand checksudo constellation-agent logs. - You are using the wrong path. Read it from the env file rather than a hardcoded literal:
. /etc/constellation/socket.env && echo "$CONSTELLATION_SOCKET_PATH"
- The runtime directory is missing after a reboot. It is recreated by
systemd-tmpfilesfrom/usr/lib/tmpfiles.d/constellation-agent.conf. If it is absent:sudo systemd-tmpfiles --create /usr/lib/tmpfiles.d/constellation-agent.confsudo systemctl restart telegraf
Note that a directory your user cannot traverse can surface as ENOENT rather than EACCES, so check group membership too.
systemd reports an unknown ImportCredential key
Symptom. The journal contains:
/usr/lib/systemd/system/telegraf.service:12: Unknown key name 'ImportCredential' in section 'Service', ignoring.
Cause. The host's systemd is older than the telegraf.service vendor unit installed by the Telegraf package. ImportCredential belongs to that vendor unit, not to the Constellation Fleet Agent configuration, and systemd is explicitly ignoring the unsupported directive.
Confirm that the Fleet Agent still starts and completes an end-to-end proof:
sudo constellation-agent status
sudo constellation-agent verify
If both commands succeed, the warning does not affect Fleet Agent authentication because the token is loaded from /etc/constellation/agent.env. Record the OS, systemd, and Telegraf versions in the host baseline and upgrade the OS/systemd package before production rather than editing /usr/lib/systemd/system/telegraf.service by hand. If the service does not start, use sudo constellation-agent logs and treat the first error after this warning as the actionable failure.
The write succeeds but nothing arrives
This is the most common report, and it follows directly from the delivery contract: a successful write only means the local kernel accepted the bytes. It does not mean the line was parsed, queued, accepted, or made visible.
Work down this list:
# 1. Is the whole path healthy independent of your code?
sudo constellation-agent verify
# 2. Watch what the agent actually receives
sudo constellation-agent debug on
sudo constellation-agent logs -f
# ... run your publisher ...
sudo constellation-agent debug off
With debug on, the agent logs the raw line protocol it received. Compare it byte for byte against what you meant to send. The usual culprits are below.
Missing or empty entity_id
Symptom. Delivery fails with invalid_input:entity_id (gRPC) or 400 {"error":"invalid_input","field":"entity_id"} (HTTP).
The platform rejects the request before storing an unqueryable record. Validate it at the producer for a faster, local error:
if not tags.get("entity_id"):
raise ValueError("entity_id is required and must be non-empty")
Also confirm the value actually matches an asset:
- Check spelling and case —
GS-001andgs-001are different entities. - Confirm the asset exists in the console under Assets or Topology.
- Watch for an interpolation bug producing a literal
None,null,undefined, or empty string.
Parse errors
A malformed line can be discarded after your write returns success. Turn on debug mode to see what the agent received.
Common malformations:
# Missing the space that separates tags from fields
ground_station,entity_id=GS-001snr=12.7
# Unescaped comma inside a tag value — parses as a second tag
ground_station,entity_id=GS-001,site=Primary, Site snr=12.7
# Unescaped space inside a tag value — ends the tag section early
ground_station,entity_id=GS-001,site=Primary Site snr=12.7
# Boolean written as a number — stored as a float, not a bool
ground_station,entity_id=GS-001 locked=1
# Unquoted string — not a valid float, bool, or integer
ground_station,entity_id=GS-001 modcod=16APSK-3/4
# No fields at all
ground_station,entity_id=GS-001
The corrected forms:
ground_station,entity_id=GS-001 snr=12.7
ground_station,entity_id=GS-001,site=Primary\,\ Site snr=12.7
ground_station,entity_id=GS-001 locked=true
ground_station,entity_id=GS-001 modcod="16APSK-3/4"
The full rules are in Socket protocol. If you are hand-rolling an encoder, check it against the reference implementation there — particularly the escaping table, which differs by position.
To check an encoding without sending it:
constellation-agent emit ground_station --entity-id GS-001 --field lat=34.2 --field lon=-118.2 --field snr=12.7 --print
Clock and timestamp errors
Symptom. Data arrives but lands at the wrong time — often 1970, or far in the future.
Almost always a unit error. Nanoseconds are 19 digits at present:
1785528000 seconds -> reads as 1970
1785528000000 milliseconds -> reads as 1970
1785528000000000 microseconds -> reads as 1970
1785528000000000000 nanoseconds -> correct
Multiply up rather than guessing:
ns = ms * 1_000_000 # from milliseconds
ns = s * 1_000_000_000 # from seconds
If the device clock itself is wrong — common on embedded hardware after a reboot, before NTP resync — omit the timestamp entirely and let the agent stamp arrival time:
ground_station,entity_id=GS-001 snr=12.7
That is strictly better than sending a confidently wrong timestamp. Check timedatectl on the host if arrival times are also skewed.
Data appears late
Symptom. Metrics arrive in bursts, or minutes after they were written.
This is normally the agent buffering through an outage and then draining, which is working as designed. Confirm:
sudo constellation-agent logs 500 | grep -i -e buffer -e retry -e "connection refused" -e timeout
du -sh /var/lib/constellation/buffer
If the buffer is large and growing, egress is failing. If it is draining, wait.
Baseline latency is about flush_interval plus a fixed transit/ingest floor — measured ~0.6s at the default 1s. If steady-state latency is much worse than that with an empty buffer, look at CONSTELLATION_FLUSH_INTERVAL and at network round-trip to the API.
Throughput ceiling
Symptom. Sustained publishing above a few hundred metrics per second falls behind and the buffer grows without an outage.
The default disk buffer strategy fsyncs per metric, capping sustained ingest near 290 metrics/s on NVMe and lower on eMMC or SD storage. That is the ceiling, not a bug.
If you need more, reinstall with the memory strategy and accept its trade-off — up to one flush_interval of telemetry lost if the node dies:
curl -fsSL https://install.constellation.space/agent | \
sudo env CONSTELLATION_API_TOKEN="$TOKEN" CONSTELLATION_BUFFER_STRATEGY=memory sh
Also make sure you are not paying avoidable cost per sample: reuse one connection, batch several newline-terminated metrics per write, and never shell out to emit in a loop.
Agent not running after reboot
sudo systemctl status telegraf
sudo journalctl -u telegraf -b --no-pager | tail -50
Check in order:
- Is the unit enabled?
sudo systemctl is-enabled telegraf - Does the runtime directory exist?
ls -ld /run/constellation-agent - Do both env files exist?
ls -l /etc/constellation/
A missing /etc/constellation/agent.env means the credential is gone and the agent cannot authenticate — re-run the installer with a valid token.
Delivery failures in the logs
| Log signature | Action |
|---|---|
401 / UNAUTHENTICATED | Stop the failing publisher and replace the invalid or revoked key. |
403 / PERMISSION_DENIED | Check telemetry:write and account ingestion access before resuming. |
429 | Inspect auth_lockout versus rate_limited and follow retry guidance. |
INVALID_ARGUMENT | Correct the rejected field before resending the batch. |
| connection refused / timeout | Check DNS, outbound port 443, and TLS connectivity. |
Inspect the response or RPC status before retrying. The agent retries transient gRPC failures, drops permanent rejections, and does not replay a partially accepted batch. See Operations for transport behavior.
Record types and delivery
Keep a field's type consistent across writes. If a numeric field needs a boolean representation, use a separate field name. Include entity_id on each record and use tags for identifiers.
A successful local socket write does not confirm delivery. Check the relevant entity and fields in topology, in addition to running sudo constellation-agent verify. See Socket protocol.
Collecting evidence for support
constellation-agent --version 2>/dev/null || grep AGENT_VERSION /etc/constellation/agent.env
sudo constellation-agent status
sudo constellation-agent verify
sudo constellation-agent logs 300 > /tmp/agent-logs.txt
Include the exact line protocol you are sending (from emit --print or debug-mode logs), the entity_id involved, and the output of verify. The verify result is the single most useful line, because it separates "the agent is broken" from "my publisher is broken".
Related
- Install — install the service and grant publisher access
- Send one metric — validate the local write path
- Verify it — prove end-to-end delivery
- Equipment contracts — the ACU and modem field contracts, including the
rssi_dbmanddoppler_hzgotchas - Socket protocol — escaping, types, and limits
- Operations — status, logs, upgrades, and buffering