Skip to main content

Troubleshooting

Start by establishing where the path broke. The two commands that answer that:

sudo constellation-agent status # is the agent running at all?
sudo constellation-agent verify # does a metric make it all the way to topology?

verify distinguishes the two halves for you:

OutputBroke at
FAIL: telemetry socket is missingThe agent — it is not running or never created the socket.
FAIL: metric entered the local socket but proof … was not read backDelivery — local handoff worked, ingest did not.
READY: telemetry proof … was read backNothing. The agent path is healthy; the problem is in your publisher.

Permission denied​

Symptom. Connecting to the socket raises Permission denied (EACCES).

Cause. The publishing user is not in the constellation group, or is in it but running in a session that predates the change.

# 1. Is the user in the group at all?
id -nG constellation-publisher | tr ' ' '\n' | grep -x constellation

# 2. Add it if not
sudo usermod -aG constellation constellation-publisher

# 3. The step people miss — the change is not live until a new session
sudo systemctl restart constellation-publisher
# or log out and back in for an interactive user

Group membership is resolved at process start. A running process keeps the groups it started with, so usermod alone changes nothing for it. If id -nG shows the group but your application still gets EACCES, the application has not been restarted.

Verify the socket and its directory look right:

ls -ld /run/constellation-agent # expect drwxrws--- telegraf constellation (2770)
ls -l /run/constellation-agent/telemetry.sock

No such file or directory​

Symptom. Connecting raises No such file or directory (ENOENT).

Causes, in the order worth checking:

  1. The agent is not running. sudo constellation-agent status. If the unit is dead, sudo systemctl restart telegraf and check sudo constellation-agent logs.
  2. You are using the wrong path. Read it from the env file rather than a hardcoded literal:
    . /etc/constellation/socket.env && echo "$CONSTELLATION_SOCKET_PATH"
  3. The runtime directory is missing after a reboot. It is recreated by systemd-tmpfiles from /usr/lib/tmpfiles.d/constellation-agent.conf. If it is absent:
    sudo systemd-tmpfiles --create /usr/lib/tmpfiles.d/constellation-agent.conf
    sudo systemctl restart telegraf

Note that a directory your user cannot traverse can surface as ENOENT rather than EACCES, so check group membership too.

systemd reports an unknown ImportCredential key​

Symptom. The journal contains:

/usr/lib/systemd/system/telegraf.service:12: Unknown key name 'ImportCredential' in section 'Service', ignoring.

Cause. The host's systemd is older than the telegraf.service vendor unit installed by the Telegraf package. ImportCredential belongs to that vendor unit, not to the Constellation Fleet Agent configuration, and systemd is explicitly ignoring the unsupported directive.

Confirm that the Fleet Agent still starts and completes an end-to-end proof:

sudo constellation-agent status
sudo constellation-agent verify

If both commands succeed, the warning does not affect Fleet Agent authentication because the token is loaded from /etc/constellation/agent.env. Record the OS, systemd, and Telegraf versions in the host baseline and upgrade the OS/systemd package before production rather than editing /usr/lib/systemd/system/telegraf.service by hand. If the service does not start, use sudo constellation-agent logs and treat the first error after this warning as the actionable failure.

The write succeeds but nothing arrives​

This is the most common report, and it follows directly from the delivery contract: a successful write only means the local kernel accepted the bytes. It does not mean the line was parsed, queued, accepted, or made visible.

Work down this list:

# 1. Is the whole path healthy independent of your code?
sudo constellation-agent verify

# 2. Watch what the agent actually receives
sudo constellation-agent debug on
sudo constellation-agent logs -f
# ... run your publisher ...
sudo constellation-agent debug off

With debug on, the agent logs the raw line protocol it received. Compare it byte for byte against what you meant to send. The usual culprits are below.

Missing or empty entity_id​

Symptom. Delivery fails with invalid_input:entity_id (gRPC) or 400 {"error":"invalid_input","field":"entity_id"} (HTTP).

The platform rejects the request before storing an unqueryable record. Validate it at the producer for a faster, local error:

if not tags.get("entity_id"):
raise ValueError("entity_id is required and must be non-empty")

Also confirm the value actually matches an asset:

  • Check spelling and case — GS-001 and gs-001 are different entities.
  • Confirm the asset exists in the console under Assets or Topology.
  • Watch for an interpolation bug producing a literal None, null, undefined, or empty string.

Parse errors​

A malformed line can be discarded after your write returns success. Turn on debug mode to see what the agent received.

Common malformations:

# Missing the space that separates tags from fields
ground_station,entity_id=GS-001snr=12.7

# Unescaped comma inside a tag value — parses as a second tag
ground_station,entity_id=GS-001,site=Primary, Site snr=12.7

# Unescaped space inside a tag value — ends the tag section early
ground_station,entity_id=GS-001,site=Primary Site snr=12.7

# Boolean written as a number — stored as a float, not a bool
ground_station,entity_id=GS-001 locked=1

# Unquoted string — not a valid float, bool, or integer
ground_station,entity_id=GS-001 modcod=16APSK-3/4

# No fields at all
ground_station,entity_id=GS-001

The corrected forms:

ground_station,entity_id=GS-001 snr=12.7
ground_station,entity_id=GS-001,site=Primary\,\ Site snr=12.7
ground_station,entity_id=GS-001 locked=true
ground_station,entity_id=GS-001 modcod="16APSK-3/4"

The full rules are in Socket protocol. If you are hand-rolling an encoder, check it against the reference implementation there — particularly the escaping table, which differs by position.

To check an encoding without sending it:

constellation-agent emit ground_station --entity-id GS-001 --field lat=34.2 --field lon=-118.2 --field snr=12.7 --print

Clock and timestamp errors​

Symptom. Data arrives but lands at the wrong time — often 1970, or far in the future.

Almost always a unit error. Nanoseconds are 19 digits at present:

1785528000 seconds -> reads as 1970
1785528000000 milliseconds -> reads as 1970
1785528000000000 microseconds -> reads as 1970
1785528000000000000 nanoseconds -> correct

Multiply up rather than guessing:

ns = ms * 1_000_000 # from milliseconds
ns = s * 1_000_000_000 # from seconds

If the device clock itself is wrong — common on embedded hardware after a reboot, before NTP resync — omit the timestamp entirely and let the agent stamp arrival time:

ground_station,entity_id=GS-001 snr=12.7

That is strictly better than sending a confidently wrong timestamp. Check timedatectl on the host if arrival times are also skewed.

Data appears late​

Symptom. Metrics arrive in bursts, or minutes after they were written.

This is normally the agent buffering through an outage and then draining, which is working as designed. Confirm:

sudo constellation-agent logs 500 | grep -i -e buffer -e retry -e "connection refused" -e timeout
du -sh /var/lib/constellation/buffer

If the buffer is large and growing, egress is failing. If it is draining, wait.

Baseline latency is about flush_interval plus a fixed transit/ingest floor — measured ~0.6s at the default 1s. If steady-state latency is much worse than that with an empty buffer, look at CONSTELLATION_FLUSH_INTERVAL and at network round-trip to the API.

Throughput ceiling​

Symptom. Sustained publishing above a few hundred metrics per second falls behind and the buffer grows without an outage.

The default disk buffer strategy fsyncs per metric, capping sustained ingest near 290 metrics/s on NVMe and lower on eMMC or SD storage. That is the ceiling, not a bug.

If you need more, reinstall with the memory strategy and accept its trade-off — up to one flush_interval of telemetry lost if the node dies:

curl -fsSL https://install.constellation.space/agent | \
sudo env CONSTELLATION_API_TOKEN="$TOKEN" CONSTELLATION_BUFFER_STRATEGY=memory sh

Also make sure you are not paying avoidable cost per sample: reuse one connection, batch several newline-terminated metrics per write, and never shell out to emit in a loop.

Agent not running after reboot​

sudo systemctl status telegraf
sudo journalctl -u telegraf -b --no-pager | tail -50

Check in order:

  1. Is the unit enabled? sudo systemctl is-enabled telegraf
  2. Does the runtime directory exist? ls -ld /run/constellation-agent
  3. Do both env files exist? ls -l /etc/constellation/

A missing /etc/constellation/agent.env means the credential is gone and the agent cannot authenticate — re-run the installer with a valid token.

Delivery failures in the logs​

Log signatureAction
401 / UNAUTHENTICATEDStop the failing publisher and replace the invalid or revoked key.
403 / PERMISSION_DENIEDCheck telemetry:write and account ingestion access before resuming.
429Inspect auth_lockout versus rate_limited and follow retry guidance.
INVALID_ARGUMENTCorrect the rejected field before resending the batch.
connection refused / timeoutCheck DNS, outbound port 443, and TLS connectivity.

Inspect the response or RPC status before retrying. The agent retries transient gRPC failures, drops permanent rejections, and does not replay a partially accepted batch. See Operations for transport behavior.

Record types and delivery​

Keep a field's type consistent across writes. If a numeric field needs a boolean representation, use a separate field name. Include entity_id on each record and use tags for identifiers.

A successful local socket write does not confirm delivery. Check the relevant entity and fields in topology, in addition to running sudo constellation-agent verify. See Socket protocol.

Collecting evidence for support​

constellation-agent --version 2>/dev/null || grep AGENT_VERSION /etc/constellation/agent.env
sudo constellation-agent status
sudo constellation-agent verify
sudo constellation-agent logs 300 > /tmp/agent-logs.txt

Include the exact line protocol you are sending (from emit --print or debug-mode logs), the entity_id involved, and the output of verify. The verify result is the single most useful line, because it separates "the agent is broken" from "my publisher is broken".