Skip to content

Data Plane delivery contract

This page is the first-release contract for Data Plane delivery. It describes the behavior supported by the current implementation; it is not an exactly-once or high-availability claim.

Guarantees

Ingestion and handoff

A successful volatile enqueue or in-process handoff means only that the Data Plane accepted the event into its current process path. It is not durable acceptance. An event that has not reached a destination or the durable DLQ can be lost if the process, node, or underlying runtime fails.

Delivery is at-least-once-oriented after an event is handed to an emitter: retries and recovery can cause the destination to receive duplicates. Delivery order is not guaranteed; batching, concurrent processing, retries, and replay can reorder events.

Emitter retries

Emitter delivery attempts are bounded by the configured emitter retry.max_attempts (with emitter-specific defaults where configuration is absent). Transient failures use exponential backoff with a configured maximum and bounded deterministic jitter. A retry is not proof that the destination did not receive the previous attempt; downstream consumers must tolerate duplicates.

When the bounded emitter attempts are exhausted, the event is written to the DLQ when DLQ storage is configured and the write, file flush, and sync succeed. The DLQ record receives a stable generated id if the emitter did not provide one, and records emitter, event, failure metadata, and attempt limits.

If the DLQ writer is unavailable or its write/sync fails, the event is reported as neither delivered nor durably queued. The Data Plane records the failure and does not claim recovery for that event.

DLQ storage and recovery

The DLQ is local JSONL storage under the configured dead_letter.directory. Writes use an active file, append a record, flush it with fsync, and rotate it into a sealed file. On restart, a non-empty active file is sealed and recovered; incomplete temporary rewrite files are removed. Malformed lines are moved to a mode-0600 quarantine file rather than silently replayed.

For Kubernetes, a PVC-backed DLQ is optional and must be explicitly enabled in the Helm values. The supported safe topology for a PVC-backed DLQ is one Data Plane replica per DLQ volume. A default ReadWriteOnce claim must not be shared by multiple Data Plane replicas. The chart does not establish shared-volume coordination or active/active DLQ ownership. emptyDir or other ephemeral storage does not survive Pod replacement.

DLQ retention is bounded by the configured sealed-file limit, but capacity exhaustion and storage failures remain operational loss conditions. Operators must monitor slimstream_dlq_depth, slimstream_dlq_size_bytes, DLQ write failures, and emitter errors, and must back up or otherwise protect the PVC according to their recovery requirements.

Replay

The scheduled DLQ reader replays sealed records through the configured process function at its scan interval. A failed replay increments the record attempt and rewrites the record; records at or above the configured dead_letter.max_retries remain in the DLQ and are not retried further. A successful replay removes the record after the rewrite/removal path completes. The record id is retained across failed replay rewrites and is logged and exposed in replay-related diagnostics.

An authenticated Manager endpoint also supports explicit manual replay of up to 50 stable IDs per request. It does not return payloads, supports no replay-all or predicate mode, requires OIDC/RBAC authorization, records an intent audit event before dispatch and a completion audit event after success, rejection, timeout, or audit failure, and uses verified HTTPS to reach the Data Plane. Replay responses are bounded by a deadline and require a complete, validated status/count result.

Replay remains at-least-once-oriented: a process, transport, or storage failure around destination acknowledgement and DLQ removal can result in a duplicate. A timeout or transport failure is reported as unknown and must be reconciled using audit and destination evidence rather than retried blindly. Concurrent scans and manual replay are serialized within one process, but this is not a multi-replica coordination mechanism.

Explicit non-guarantees and limitations

The first release does not claim:

  • exactly-once delivery, exactly-once processing, or duplicate suppression;
  • ordering or preservation of source order;
  • durable acceptance at volatile enqueue or before a successful DLQ write and sync;
  • a durable source offset/ack transaction coupled atomically to emitter delivery;
  • high availability, active/active Data Plane replicas, or automatic failover for a DLQ volume;
  • Kubernetes restart, rescheduling, or PVC attach guarantees beyond what the selected storage class and operator configuration provide;
  • exactly-once manual replay, duplicate suppression, or a guarantee that a timeout means the operation did not execute;
  • successful recovery when the DLQ is disabled, unavailable, full, corrupted, or outside the retention limit;
  • proof that every configured emitter has identical retry defaults or failure classification.

The implementation and tests validate local retry, DLQ, replay identity, restart recovery, malformed-record quarantine, bounded manual replay responses, audit ordering, and lifecycle behavior. They do not prove node-failure durability, Kubernetes restart semantics, destination-side acknowledgement durability, or multi-replica DLQ coordination.

Operational expectations

  1. Configure a PVC-backed DLQ before relying on recovery from Pod replacement, and run one Data Plane replica per claim/volume.
  2. Treat an ingest response or volatile enqueue as non-durable until the downstream system or DLQ has acknowledged the event according to its own contract.
  3. Make destinations idempotent or deduplicate by an application event identity when duplicates are unacceptable.
  4. Monitor and alert on emitter retry/error metrics, DLQ depth/bytes, DLQ write failures, and quarantined records.
  5. Bound replay work operationally, verify the destination before replay, preserve DLQ and audit evidence, and expect replay to be repeatable rather than exactly once.
  6. Test the actual storage class, Pod replacement, backup, restore, and capacity behavior in the target Kubernetes environment; Helm rendering alone is not evidence of those guarantees.