Skip to main content

Runtime Topology: YU12-aeron-cluster

Entrypoints​

EntrypointTransportConsumer
gateway REST portHTTP REST/UIunchanged inherited clients
gateway FIX portFIX 4.4 over TCPunchanged inherited FIX initiators
cluster ingress/egress UDPAeron Cluster client protocolgateway tier and feed adapter
member consensus UDP portsAeron Cluster consensus/log/catch-upcluster members
archive control/replay UDPAeron Archive protocolmember recovery and snapshot retrieval

Components​

  • order-matcher cluster member (3 StatefulSet replicas): one pod runs the Media Driver, Archive, Consensus Module, and the clustered service container hosting the inherited MatchingEngine and two-tier risk core. Per-pod PVC holds the consensus log and snapshots. Stable StatefulSet ordinals provide member identity; the blp-pool dedicated-core pinning applies to the single service thread.
  • fix-gateway tier: terminates counterparty FIX sessions and REST connections, screens admission against control-feed state, forwards through the Aeron Cluster client, and re-points on leader change without dropping counterparty sessions. Scales out horizontally; each replica holds its own cluster session + owner thread, so N replicas = NΓ— parallel ingress; the order-matcher-gw Service round-robins REST (no affinity) and the separate order-matcher-gw-fix Service pins FIX with sessionAffinity: ClientIP; k8s affinity is per-Service, not per-port, so one combined Service would pin REST too (ADR-047).
  • feed adapter: consumes inherited NATS pricing/control subjects and publishes conflated ticks and policy updates as cluster ingress.
  • trade-egress bridge (TradeNatsPublisher, ADR-048): on the leader, republishes every booked trade from the deterministic apply stream to NATS /trades (non-blocking enqueue off the apply thread), so cluster fills reach trade-processor β†’ the SQL DB β†’ the UI blotter/position feeds.
  • NATS/JetStream: inherited non-replication roles only; pricing, control feeds, output distribution (incl. the trade-egress bridge), EOD gating. No replication leg, no witness bucket.
  • trade-processor, MariaDB, position-service, UI, downstream services: unchanged inherited CQRS topology, now fed by committed cluster outputs via the trade-egress bridge; trade-processor consumes /trades, persists Trade + Position to MariaDB, and republishes /accounts/*/trades + /positions to the UI's NATS websocket.

Networking​

  • Cluster members exchange consensus, log, and catch-up traffic over dedicated cluster-internal UDP ports between stable StatefulSet ordinal DNS names on a headless Service with publishNotReadyAddresses: true.
  • The gateway and feed adapter reach members over the cluster ingress/egress ports; a namespace-scoped NetworkPolicy restricts every Aeron port to the participating pods.
  • No Aeron port uses ingress-nginx, LoadBalancer, NodePort, IP multicast, or host mappings.
  • Kind uses a dedicated named multi-node cluster with three schedulable workers and required anti-affinity; the shared single-node cluster is not modified.
  • GKE required anti-affinity keeps one member per blp-pool node.

Startup / Health Order​

  1. Each member opens its Aeron directory, validates the Archive catalog and cluster mark file, and recovers: newest valid snapshot loaded, committed log applied strictly after the snapshot position, generator assertion passed.
  2. Members complete Raft election; a majority elects exactly one leader.
  3. A wiped replacement member retrieves the latest snapshot and replays the committed log tail before reporting follower readiness.
  4. The feed adapter connects and sequences control/pricing ingress; gateway control-feed admission state becomes valid.
  5. The gateway opens counterparty admission only when cluster readiness and admission-state readiness both hold.

Degraded Behavior​

ConditionBehavior
One member lost (of three)Majority holds; commit and admission continue; the replacement rejoins via snapshot retrieval + log replay.
Leader lostRaft re-election among the majority; the gateway re-points on the leader signal; counterparty sessions stay connected.
Partition minorityThe minority cannot elect a leader, extend the log, or admit orders; it rejoins and truncates uncommitted entries on heal.
Two members lost (of three)No majority: commit and admission stop; state is preserved on the surviving log/snapshot volumes.
Gateway instance lostCounterparty sessions drop to ordinary reconnect; cluster state is unaffected; REST routing resumes on the replacement.
Feed adapter lostNo new ticks/control updates are sequenced; order flow continues against last-applied state; adapter restart resumes ingress.
Snapshot/log disk pressureMembers surface archive/log disk state through health; recording refuses before unsafe exhaustion.
Generator assertion failure on recoveryThe member refuses readiness and does not serve or vote leadership with invalid state.