Skip to main content

Non-Functional Delta: YU01-lmax-sequencer

Parent state: 009-order-management-matcher

Document NFR changes introduced by this state.

Runtime / Operations​

  • Keep the LGTM stack from 007 (Grafana, Prometheus, Loki, Tempo, OTel Collector, Promtail, Blackbox Exporter) and all 009 runtime components; ingress routing and UI serving are unchanged.
  • Rebuild order-matcher internals as the LMAX hot-path node hosting the sequencer, input ring, journaler/replicator/un-marshaller, BLP, and output ring (same service identity, port, and health surface as 009).
  • trade-service additionally plays the Gateway/Receptionist role: edge validation from in-memory replicas, ticker -> securityId mapping, fixed-point conversion, SBE encode, sequence submission. Its REST/WS contract is unchanged.
  • New durable runtime state: one journal directory (journal.path, default ./data/journal) holding the append-only input-events.journal, the periodic snapshot.dat checkpoint, and the symbols.tab ticker map (mounted volume order_matcher_journal -> /opt/app/data in containers, durable across recreate). These must survive restarts/recreate. (A standalone projection-checkpoint file is deferred; the projector projectedSeq watermark is in-memory, advanced on a committed DB flush.)
  • Run profiles (runtime.profile): demo (default for C2 containers; BlockingWaitStrategy, no pinning, no hugepages, single replica, replication loopback/stub), perf (bare metal; busy-spin on BLP/Journaler, pinned isolated cores via isolcpus/Affinity, ZGC/Shenandoah, large pages, NUMA, replicas + DR), noGcTest (CI; Epsilon GC small fixed heap). JVM flag matrix in requirements/no-gc-conformance.md.
  • Operability: recovery is snapshot + journal-tail replay on every restart (recovery.source=db warm-start
    • verify, or journal no-DB cutover with output.projector.db.enabled=false). A cron-scheduled nightly bounce (nogc.bounce.cron) and a JIT warm-up replay before going live (nogc.warmup.events) remain aspirational.
  • Startup order and health gating per system/runtime-topology.md (node is "ready" once recovery completes; no warm-up gate yet).

Security / Compliance​

  • No auth/RBAC change; admin view remains local-dev demonstration scope, as in 009.
  • Operational actions (cancel/force-fill) remain auditable: they are now journaled sequenced events, strictly ordered and replayable, in addition to structured logs (off-hot-path/async).
  • The journal and snapshots contain trading data and must live on the same trust footing as the database volume; no secrets in journal, snapshot, or deployment-bundle artifacts.
  • As convergence level C2, container build/publish CI with namespace ghcr.io/finos/traderx-c2/<component>, immutable commit-SHA tags plus latest, GHCR run bundle, and runtime/deploy/ bundle obligations carry forward from 009 unchanged.
  • New dependencies (Disruptor, Agrona, SBE, Chronicle Queue/Aeron, Affinity, HdrHistogram, JMH) are pinned to latest CVE-clean releases and subject to the repo dependency CVE gate.

Performance / Scalability​

Latency budgets (performance profile; demo/C2 profile is exempt from budgets but not from the allocation gate):

StageTypicalp99 budget
Gateway validate + encode + submit2–5 Β΅s< 20 Β΅s
Sequencer + input ring claim< 1 Β΅s< 5 Β΅s
Journal (durable append)5–20 Β΅s< 50 Β΅s
Replication ack (LAN)30–80 Β΅s< 150 Β΅s
BLP business logic (match + book + position + emit)1–5 Β΅s< 25 Β΅s
Output ring + marshal2–5 Β΅s< 20 Β΅s
In-node compute (Gateway -> output emit, excl. network)~15–40 Β΅s< 150 Β΅s
End-to-end incl. durable + replicated ack~0.3–1 ms< 3 ms
  • NATS fan-out and read-model projection are off the order-acknowledgement path and carry no ack-path budget.
  • Throughput headroom: single-threaded BLP capacity (LMAX reference: ~6M events/s/thread) is orders of magnitude beyond demo load; backpressure is bounded by ring capacity, never unbounded queues.
  • Batching: consumers drain to the highest available sequence and amortize per-batch costs (journal flush, output flush) on endOfBatch; the projector enqueues on the ring and its separate drain thread batches the DB writes off it.
  • All latency reporting is full-distribution (HdrHistogram p50/p99/p99.9/max) with jHiccup separating JVM/OS pauses from application latency.

Reliability / Observability​

  • Recovery: periodic snapshot.dat + bounded journal-tail replay to the last journaled sequence (recovery.source=db warm-start + verify, or journal no-DB cutover); restart target < 1 minute. Read-model loss degrades to projector rebuild, not trading-state loss. (Warm-up replay before going live is deferred.)
  • Failover: follower BLPs at the same sequence with output suppressed; promotion-based failover without cold replay (perf profile; contract-level check in demo profile).
  • Decoupling: DB or NATS outage must not stop matching; affected output handlers lag within the bounded ring and catch up (FR-09B24).
  • Order-management components keep Prometheus metrics and /health; mandatory scrape coverage per 009 NFR-01308 continues to apply to every metrics-capable service.
  • Required metrics (in addition to all retained 009 order metric families):
MetricTypeMeaning
traderx_disruptor_input_remaining_capacitygaugeFree input-ring slots (backpressure headroom).
traderx_input_published_seqgaugePublisher cursor.
traderx_input_gating_seqgaugemin(journaler, replicator, unmarshaller).
traderx_input_seq_laggaugepublished βˆ’ BLP consumed.
traderx_input_events_total{type=...}counterPer-type ingest counts.
traderx_input_backpressure_events_totalcounterProducer waits for a free slot.
traderx_journal_write_latency_secondshistogramJournaler append latency.
traderx_replication_ack_latency_secondshistogramReplicator ack latency.
traderx_blp_event_latency_secondshistogramonEvent processing latency (real measurement).
traderx_blp_book_depth{security=...}gaugeResting orders per security.
traderx_blp_positions_totalgaugeDistinct in-memory positions.
traderx_blp_cache_miss_total{cache=...}counterRequest/response events emitted for misses.
traderx_blp_snapshot_secondshistogramSnapshot duration (planned; snapshot/recovery currently logged, not metered).
traderx_blp_replay_secondsgaugeLast recovery replay duration (planned; see JOURNAL-REPLAY VERIFY / LIVE RECOVERY log lines).
traderx_output_publish_latency_secondshistogramTrue end-to-end (now βˆ’ ingressNanos) at egress.
traderx_output_remaining_capacitygaugeOutput ring headroom.
traderx_output_events_total{kind=...}counterPer-kind egress counts.
traderx_output_nats_errors_totalcounterNATS bridge publish failures.
traderx_projector_lag_seqgaugeBLP seq βˆ’ last projected seq.
traderx_projector_batch_sizehistogramRows per projector flush.
traderx_hotpath_alloc_bytes_total{node=...}counterSteady-state allocation (must stay ~0).
traderx_jvm_gc_pause_secondshistogramGC pause distribution (expect empty/sub-ms).
traderx_jit_warmup_secondsgaugeWarm-up duration at startup.
traderx_nightly_bounce_secondsgaugeLast bounce (restart + replay) duration.
  • Grafana dashboard additions: ring headroom + sequence lag, journal/replication latency percentiles, BLP event latency, true end-to-end egress latency, projector lag/batch size, allocation-rate panel alerting if > 0 in steady state, GC-pause panel (expect flat), warm-up/bounce durations. Existing 009 order dashboards continue to work unchanged.
  • Smoke tests assert metrics endpoint availability and non-empty response for the required metric families above.