Observability¶
Logstrm exposes a common Prometheus contract from the Data Plane and Manager. The contract is identical for Kubernetes and VM deployments; only target discovery differs.
Endpoints¶
| Component | Kubernetes Service port | VM address | Path |
|---|---|---|---|
| Data Plane | api |
127.0.0.1:8090 |
/metrics |
| Manager | http |
127.0.0.1:8091 |
/metrics |
The Manager metrics endpoint is public in the same way as the health endpoint. Administrative API routes under /api/v1/ remain protected by the configured authentication middleware.
Metric contract¶
Metric names retain the slimstream_ prefix for compatibility with existing dashboards, alerts and integrations. Treat these as stable technical identifiers; the product name is Logstrm. Environment variables documented with the SLIMSTREAM_ prefix follow the same compatibility convention. Do not rename these identifiers in monitoring configuration when referring to the Logstrm product.
Data Plane¶
slimstream_events_ingested_total— events that matched a pipeline; userate()for matched-event throughput, not raw source payload count.slimstream_events_emitted_total— events handed to configured emitters. TheEmitter.Emitinterface has no delivery result, so this counter does not confirm successful downstream delivery. Fan-out counts one handoff per emitter invocation.slimstream_events_dropped_total— events rejected before any emitter handoff, with a bounded reason and pipeline label.slimstream_bytes_ingested_total— input payload bytes measured at supported source boundaries. A payload can contain zero, one or many events; this byte counter and the matched-event counter have different units and scopes.slimstream_bytes_saved_total— estimated reduction between compact JSON encodings of an event before and after transformations. It is not measured destination wire traffic and must not be interpreted as vendor-billed bytes or realized cost savings.slimstream_bytes_emitted_total— reserved metric family; output bytes are not currently recorded consistently across emitters.slimstream_emitter_request_body_bytes_total{emitter_name,destination_type}— serialized HTTP request-body bytes submitted to the HTTP client per attempt for generic HTTP JSON, Sentinel DCR, Splunk HEC and Elasticsearch Bulk NDJSON. Retries count each submitted attempt; Sentinel DCR counts the compressed body when gzip is enabled. SQS request bytes are not measured.slimstream_emitter_payload_bytes_total{emitter_name,destination_type}— serialized JSON object-payload bytes submitted to S3, GCS and Azure Blob upload SDKs per attempt, plus each serialized Pub/Sub message payload passed toPublish. Excludes Pub/Sub message attributes, protocol framing, request metadata and transport overhead; it is not wire or billing bytes. Retries count each call's message payload again.slimstream_emitter_accepted_events_total{emitter_name,destination_type}— HTTP 2xx batch acceptance for generic HTTP JSON, Sentinel DCR and Splunk HEC; Elasticsearch Bulk items with a successful item status; SQS entries listed as successful in the service response; S3/GCS/Azure Blob batch events after successful upload SDK completion; and individual Pub/Sub messages whose asynchronousPublishResult.Getsucceeds. These are attempt-level outcomes: retry may count the same event again, and acceptance is not proof of downstream consumption, unique delivery, or end-to-end delivery.slimstream_emitter_failed_events_total{emitter_name,destination_type}— Elasticsearch Bulk items with a non-2xx item status, SQS entries explicitly listed as failed by the service response, and Pub/Sub messages whose asynchronousPublishResult.Getreturns an error. These are explicit per-item outcomes; a publish transport/cancellation error without an individual result is not independently classified, and retries may count an explicit failure repeatedly.slimstream_emitter_acknowledged_events_total{emitter_name,destination_type}— Splunk HEC events counted only after a positive HEC protocol ACK when ACK mode is enabled. HTTP 2xx without a positive ACK does not increment this counter.slimstream_processing_duration_seconds— processing latency histogram.slimstream_routing_duration_secondsandslimstream_transform_duration_seconds— pipeline stage latency histograms.slimstream_emitter_errors_totalandslimstream_emitter_retries_total— connector error and retry counters, labelled by emitter.slimstream_circuit_breaker_rejected_total— batches fast-failed while an emitter circuit is open.slimstream_circuit_breaker_transitions_total{emitter,state}— emitter breaker state transitions (open,half_open,closed).slimstream_rate_limit_wait_seconds— time spent waiting for per-emitter batch-rate tokens.slimstream_dlq_depth— current number of entries waiting in the DLQ.slimstream_dlq_size_bytes— current physical DLQ size.slimstream_dlq_events_totalandslimstream_dlq_retries_total— DLQ counters.slimstream_dlq_quarantined_total— malformed DLQ lines moved out of replay.slimstream_dlq_write_failures_total— delivery failures that could not be durably queued. A non-zero rate means the event was neither delivered nor persisted.slimstream_active_pipelines— active pipeline count.slimstream_config_reloads_total— configuration reload counter.
Manager¶
slimstream_manager_rollouts_total{node,status}— completed rollout attempts.slimstream_manager_rollout_duration_seconds{node}— rollout duration histogram.slimstream_manager_node_status{node,status}— one for the current node status and zero for other known statuses.slimstream_manager_rollout_nodes{status}— node counts from the latest completed rollout.
Manager status values are success, failed, cancelled, and unknown. A cancelled rollout did not start that node because the request context ended. Existing metric names are unchanged.
A rollout response with X-Logstrm-Persistence: failed means the push already ran, but the Manager could not persist the rollout row or node status. The response body still contains the node results. The Manager log records rollout result persistence failed with the version ID and error, without node credentials.
Metric interpretation and byte-accounting scope¶
Input-byte accounting is currently recorded where the runtime has an explicit payload boundary: accepted cloud payload jobs (once when queued, not once per processing retry), Kafka record values, successfully read size-checked Azure Blob payloads, accepted benchmark HTTP request bodies, and raw syslog messages processed by the syslog listener. The labels identify the source; measurements can represent payloads that later fail parsing, routing or processing. They are not a count of successfully delivered event bytes.
Egress accounting is connector-specific and attempt/response based. slimstream_emitter_request_body_bytes_total measures serialized HTTP request-body bytes submitted to the HTTP client for generic HTTP JSON, Sentinel DCR, Splunk HEC and Elasticsearch Bulk NDJSON—not headers, TLS framing, all network traffic, or vendor-billed bytes. slimstream_emitter_payload_bytes_total measures the serialized JSON object payload passed to S3, GCS and Azure Blob upload SDKs per attempt, and the serialized data bytes of each Pub/Sub message passed to Publish; it excludes Pub/Sub attributes, protocol framing, SDK/network overhead, and provider-billed bytes. Both counters include each submitted retry attempt, and fan-out is counted independently for each destination. Elasticsearch counts the complete serialized NDJSON body on each attempt, then accounts accepted and failed items from Bulk response item statuses. On a partial Bulk response, accepted items are counted immediately while the batch-level retry continues to resend the whole batch; a later successful response can therefore count already accepted items again. For S3/GCS/Azure Blob, accepted events increment only when the corresponding upload SDK operation completes successfully. Pub/Sub acceptance and explicit failure are recorded per message only after that message's PublishResult.Get resolves; the existing batch retry resends the whole batch when any result fails, so previously successful messages can be published and counted as accepted again. None of these acceptance counts establishes downstream consumption or end-to-end delivery.
SQS accounting uses only service response entry IDs: successful entries increment slimstream_emitter_accepted_events_total, and explicitly failed entries increment slimstream_emitter_failed_events_total. Partial batch retries contain only the failed subset. An event accepted on retry can be counted again. Entries with missing/ambiguous results are retried conservatively but are not classified as explicit service failures. SQS request bytes are intentionally not reported because the SDK boundary does not expose a reliable serialized-request size. These counters do not guarantee unique ingestion or durable delivery. slimstream_emitter_acknowledged_events_total remains Splunk HEC-specific and requires a positive protocol ACK. Pub/Sub message payload bytes and per-message publish results are now included, but SQL transaction and JSONL local-file accounting remain outside this scope and are not network-egress metrics. S3, GCS and Azure Blob object payloads are accounted separately by slimstream_emitter_payload_bytes_total. The reserved slimstream_bytes_emitted_total is unsuitable for cross-emitter byte or billing estimates. slimstream_bytes_saved_total remains a compact-JSON transformation-size estimate and does not account for destination envelopes, retries, fan-out, or billing rules. Validate actual vendor billing with destination usage records.
What these metrics are NOT¶
- Not wire bytes: request-body and object-payload bytes exclude headers/metadata, TLS framing, transport overhead, and other network traffic; SQS request bytes are not measured.
- Not billing bytes: destination-side compression, service accounting, retries, and vendor billing rules can differ. Use the destination’s billing/usage records for cost reconciliation.
- Not proof of exactly-once or durable delivery: accepted counts record successful attempts or response items, and a retry can count an event again. They do not establish unique ingestion or durable downstream processing.
Dashboard examples:
- Accepted rises faster than ingested: investigate retries and partial responses. A later successful retry can increment accepted again for an event already accepted by an earlier attempt; the accepted counter is attempt/response-based, not a unique-event total.
- Request-body bytes spike while event throughput is flat: check retry rates and destination fan-out. Each submitted HTTP attempt is counted, and fan-out records a separate request per destination.
- Failed events rise without a matching transport-error spike: inspect Elasticsearch Bulk item statuses, SQS entry failures, or Pub/Sub publish results. The failed counter records explicit per-item/per-entry service failures, while transport errors and ambiguous results are not classified as item failures.
- Blob accepted events rise, but consumers report missing data: the counter records successful upload SDK completion only. Verify object visibility, downstream discovery/processing, and consumer-side acknowledgements independently.
PromQL examples¶
Matched event throughput:
sum(rate(slimstream_events_ingested_total[5m]))
Processing p95 and p99:
histogram_quantile(
0.95,
sum by (le) (rate(slimstream_processing_duration_seconds_bucket[10m]))
)
histogram_quantile(
0.99,
sum by (le) (rate(slimstream_processing_duration_seconds_bucket[10m]))
)
Emitter HTTP request-body bytes, object payload bytes, accepted events and explicitly failed items/entries:
sum by (emitter_name, destination_type) (rate(slimstream_emitter_request_body_bytes_total[5m]))
sum by (emitter_name, destination_type) (rate(slimstream_emitter_payload_bytes_total[5m]))
sum by (emitter_name, destination_type) (rate(slimstream_emitter_accepted_events_total[5m]))
sum by (emitter_name, destination_type) (rate(slimstream_emitter_failed_events_total[5m]))
sum by (emitter_name, destination_type) (rate(slimstream_emitter_acknowledged_events_total[5m]))
These are attempt/response rates; successful retries may increase acceptance counts more than once for an event. The request-body byte rate has no SQS series because SQS request bytes are not measured. Object payload bytes apply to S3, GCS, Azure Blob and Pub/Sub; for Pub/Sub they measure serialized message data per Publish call, not attributes, wire bytes or billing bytes.
Connector errors and retries:
sum by (emitter) (rate(slimstream_emitter_errors_total[10m]))
sum by (emitter) (rate(slimstream_emitter_retries_total[10m]))
DLQ state:
slimstream_dlq_depth
slimstream_dlq_size_bytes
Rollout failures:
sum(rate(slimstream_manager_rollouts_total{status="failed"}[15m]))
Kubernetes¶
The Helm chart keeps Prometheus Operator resources disabled by default. Run the example from the repository root and enable it only when the cluster has the Prometheus Operator CRDs installed. Create the logstrm-hot-reload Kubernetes Secret first, as described in the Helm deployment guide; rendering now fails if authentication is not configured:
helm upgrade --install logstrm slimstream-helm \
--namespace logstrm \
--create-namespace \
--set dataplane.hotReload.existingSecret=logstrm-hot-reload \
--set observability.serviceMonitor.enabled=true \
--set observability.prometheusRule.enabled=true \
--set observability.serviceMonitor.labels.release=kube-prometheus-stack \
--set observability.prometheusRule.labels.release=kube-prometheus-stack
The chart creates two ServiceMonitor resources:
- Data Plane selects the Data Plane Service and scrapes its
apiport; - Manager selects the Manager Service and scrapes its
httpport.
Both scrape /metrics every 30 seconds by default. The slimstream_component target label distinguishes the two components. Alert rules cover endpoint availability, stopped ingestion, high p99 latency, connector errors/retries, DLQ growth and failed rollouts. Additional rule groups can be supplied through observability.prometheusRule.groups.
Import the grafana/logstrm-overview.json dashboard from the Logstrm source distribution into Grafana and select the Prometheus datasource. The dashboard includes ingestion EPS, processing p95/p99, connector health, DLQ depth/bytes, rollout status, node status and scrape availability.
VM¶
Use the deploy/vm/prometheus-scrape.yml example from the Logstrm source distribution as the starting point for a Prometheus instance on the VM. It scrapes:
127.0.0.1:8090/metricsaslogstrm-dataplane;127.0.0.1:8091/metricsaslogstrm-manager.
The loopback-only endpoints avoid exposing administrative metrics outside the host. If Prometheus runs elsewhere, expose the endpoints only through an authenticated, restricted private path and update the targets accordingly.
Journald and log correlation¶
Prometheus does not collect logs. Forward journald through the approved host log agent or SIEM connector and preserve these fields:
unit=logstrm.servicewithslimstream_component=dataplane;unit=logstrm-manager.servicewithslimstream_component=manager;- VM hostname, environment and journal cursor/timestamp.
Useful local checks:
journalctl -u logstrm -u logstrm-manager -f
journalctl -u logstrm --since '1 hour ago' -o json
Do not use credentials, complete log payloads or unbounded event values as metric or log labels. Keep journald retention bounded and monitor the filesystem containing the journal.
Operational checks¶
After deployment, verify both metric endpoints and scrape health:
curl -fsS http://127.0.0.1:8090/metrics | grep logstrm_
curl -fsS http://127.0.0.1:8091/metrics | grep slimstream_manager_
In Prometheus, confirm up == 1 for both targets and validate that histogram series include _bucket, _sum, and _count. Alert thresholds in the chart are starting operational defaults; tune them against measured workload baselines rather than treating them as performance guarantees.