Systemic architectural invariants for Raven.Bench

#13 · open · 0 comments

View on GitHub ↗

redknightlois

## Purpose Canonical cross-cutting architectural invariants for **Raven.Bench** — the RavenDB knee-finding benchmark harness. This issue describes the rules the system enforces, derived from code, tests, and runtime behavior. ## Scope of this issue This is not an implementation task. It is the authoritative description of the rules that govern load generation, latency measurement, knee detection, transport abstraction, workload construction, metrics collection, and result comparison. ## Systemic invariants ### SI-01 — Closed-loop load generation bounds request rate by actual completions The closed-loop generator uses a bounded channel with capacity equal to the target concurrency. The producer blocks when all workers are busy. No work is queued ahead of completions. `TargetThroughput` is always `null` for closed-loop generators. - **Why:** Unbounded queuing hides the real server limit by shifting latency into the client queue, producing artificially high throughput numbers that do not reflect sustainable capacity. - **Implication:** Latency measurements under closed-loop reflect actual server response time, not client-side queue depth. ### SI-02 — The knee is the last safe step, not the degraded step When `KneeFinder` detects degradation at step `i`, it returns step `i - 1`. The quality metric is `throughput / p999` — higher is better. The finder never declares a knee before the danger zone (p50 >= 100ms) and defers the decision if the next step recovers quality by more than 3%. - **Why:** The knee represents the last operating point where the system delivers reliable throughput without latency blow-up. Returning the degraded step would recommend an unstable operating point. - **Implication:** CI gates that use knee throughput are asserting a sustainable value, not a peak. Single-step runs return that step with reason "single-step"; end-of-range returns the last step if no degradation is found. ### SI-03 — Transports return errors as strings in TransportResult, never throw from ExecuteAsync Both `RawHttpTransport` and `RavenClientTransport` follow the same contract: `TaskCanceledException` from an external token is not an error (benchmark ending), timeouts and HTTP failures produce a non-null `ErrorDetails` string, and `IsSuccess` is defined as `ErrorDetails == null`. Exceptions do not escape `ExecuteAsync`. - **Why:** The latency recorder and metrics pipeline depend on every `ExecuteAsync` call returning a result. An unhandled exception corrupts the HDR histogram state and breaks the step's metrics. - **Implication:** Error rate is computed from `TransportResult.ErrorDetails != null`, not from exception counts. ### SI-04 — Every step carries both raw and normalized latency percentiles `StepResult` always populates both `Raw` and `Normalized` percentile structs. Normalization subtracts the 5th-percentile calibration RTT (measured from 32 startup requests) from all raw percentiles. Calibration failure degrades gracefully — normalized percentiles fall back to raw values. - **Why:** Raw latency includes network overhead that varies by deployment. Normalized latency isolates load-induced queuing from baseline RTT, making results comparable across different network topologies. - **Implication:** All output formats (JSON, CSV, console table) present both latency flavors. The 5th percentile (not median) is used as baseline to capture best-case network latency. ### SI-05 — Percentiles are monotonically non-decreasing: P50 <= P95 <= P99 <= P999 <= P9999 <= PMax This holds for both raw and normalized percentiles. The integration test suite explicitly asserts this chain. HDR Histogram guarantees it mathematically; the test is a regression guard against post-processing that could break ordering. - **Why:** Percentile inversion (e.g., P95 < P50) indicates a bug in latency normalization or histogram snapshot logic. Any consumer of the results (reporting, CI gates) assumes this ordering. - **Implication:** Any transformation applied to percentiles (subtraction, scaling) preserves monotonicity or clamps to non-negative values. ### SI-06 — WorkloadMix reads, writes, and updates sum to exactly 100 The `WorkloadMix` constructor rejects any triple that does not sum to 100. The `FromWeights` factory normalizes arbitrary non-negative weights to integer percentages using the largest-remainder method, guaranteeing the sum is exactly 100. - **Why:** The probabilistic operation dispatch in `MixedProfileWorkload` uses `rng.Next(0, 100)` with range boundaries. A sum != 100 produces either unreachable operations or an out-of-bounds read. - **Implication:** Adding a new operation type requires updating the normalization and the dispatch ranges. ### SI-07 — Read and query workloads fail fast on an empty keyspace `ReadWorkload` and `QueryWorkload` throw `InvalidOperationException` at construction if `initialKeyspace <= 0`. `MixedProfileWorkload` falls through to a write when a read or update is selected but the keyspace is empty. - **Why:** A read-only workload against an empty database produces 100% errors that look like server failures when they are actually client misconfiguration. Failing at construction surfaces the problem before any measurement begins. - **Implication:** Dataset-based profiles (StackOverflow, ClinicalWords) satisfy this by importing data before constructing the workload. ### SI-08 — SNMP metrics take priority over fallback server metrics When SNMP is enabled, the console table and CSV output show SNMP-sourced CPU, memory, and IO metrics. Fallback metrics (from RavenDB admin endpoints) are hidden. A discrepancy warning is emitted when client-measured and SNMP-reported request rates differ by more than 10%. - **Why:** SNMP data comes from the OS/server level and reflects actual resource consumption. Fallback metrics are approximations from application-level counters that can diverge under load. - **Implication:** SNMP and fallback metrics are never mixed for the same dimension in any output. ### SI-09 — Cross-run comparison requires matching profile, dataset, and query profile `RunCompatibilityChecker` blocks comparison of runs with different workload profiles, datasets, or query profiles. Transport, HTTP version, and compression are explicitly allowed to differ — comparing these is the purpose of the tool. - **Why:** Comparing a write benchmark against a read benchmark is meaningless. Comparing the same workload over raw HTTP vs. the RavenDB client is the primary use case. - **Implication:** The comparison model aligns steps by concurrency level. Missing levels in either run produce null entries, not interpolated values. ### SI-10 — StepPlan guarantees monotonic forward progress `StepPlan.Next()` always advances by at least 1: `Math.Max(current * factor, current + 1)` followed by `Math.Max(next, current + 1)`. `StepPlan.IsValid` requires `Start > 0`, `End >= Start`, `Factor > 0`. `Normalize()` clamps degenerate inputs. - **Why:** A step plan that can return the same concurrency level twice produces an infinite loop in the benchmark runner. The double-max guarantees termination regardless of factor value. - **Implication:** Factor values below 1.0 still produce forward progress (+1 per step). The step loop always terminates. ### SI-11 — Latency values exceeding 60 seconds are treated as errors The HDR Histogram range is 1 microsecond to 60 seconds. Values outside this range throw `InvalidOperationException`, which `LoadGeneratorExecution` catches and converts to an error (incrementing error count, skipping histogram recording). Latency is always >= 1 microsecond via `Math.Max(1, ...)`. - **Why:** A 60+ second operation represents a completely unusable response time. Recording it as latency would distort the histogram tail; recording it as an error correctly reflects that the operation failed to meet any reasonable SLA. - **Implication:** The 60s ceiling interacts with the 30s per-operation transport timeout. Only coordinated omission correction backfills can produce values in the 30-60s range. ### SI-12 — Write key generation is atomic across all concurrent workers All write paths (`WriteWorkload`, `BulkWriteWorkload`, `MixedProfileWorkload`) use `Interlocked.Increment(ref _maxKey)` for key generation. Keys follow the format `"bench/{i:D8}"` (zero-padded 8-digit). `RateLoadGenerator` additionally protects `Random` with a lock since it is not thread-safe. - **Why:** Non-atomic key generation under concurrent workers produces duplicate document IDs, causing spurious 409 conflict errors that appear as server failures in the benchmark results. - **Implication:** StackOverflow workloads use pre-existing dataset IDs and do not go through this path. ### SI-13 — The benchmark self-terminates on extreme degradation The step loop stops early when error rate exceeds `max(MaxErrorRate, 5%)`, when rate-mode throughput drops more than 30% below target or regresses more than 30% from the previous step, or when p99.99 exceeds 30 seconds. - **Why:** Continuing to ramp concurrency after the system is clearly saturated produces unreliable data and wastes time. Early termination bounds benchmark duration and prevents result corruption. - **Implication:** The max-error threshold is configurable via `--max-errors`, but the 30-second p99.99 ceiling and 30% regression thresholds are fixed. ### SI-14 — Vector search profiles require the Corax search engine `WorkloadProfiles.SupportsEngine()` returns false for vector search profiles when the engine is not Corax. This validation runs at startup before dataset import or calibration. - **Why:** RavenDB's HNSW vector index implementation is Corax-specific. Running vector search against a Lucene-backed index produces zero results or errors that look like data problems. - **Implication:** The engine validation is fail-fast — the benchmark does not start rather than producing misleading results. ## Epic structure rule Each epic inherits these invariants. Epic-local invariants describe additional constraints specific to that subsystem. Specs within epics define concrete behavior and acceptance criteria.

Comments