Every claim traces to a measurement.

The numbers on our homepage are experimental results, and this page is the full dossier: what we ran, what we measured, what we found, and — just as deliberately — what this evidence does not yet establish.

Why validate against real failure

Every evaluation of a dependency model that is derived from the model's own graph is circular: it can confirm the arithmetic is self-consistent, and nothing more. Our validation breaks the circle. The measurement is the cluster's observed behaviour under real, injected failure — it does not use the graph, the CEI weights, or any inference the system makes. Prediction and measurement are independent instruments; agreement between them means something.

Blast radius: per-experiment results

Six experiments on a live 3-node cluster: Chaos Mesh pod-failure, 90-second windows, one target workload at a time, full recovery verified between runs. The topology has real runtime dependencies — each service's readiness probe checks TCP reachability of its upstreams, so losing a dependency genuinely makes dependents unready, the same mechanism by which real services degrade.

Target killedPredicted affectedMeasured affectedPrecisionRecall
dbapi, reporting, web, workerapi, reporting, web, worker1.001.00
authapi, webapi0.501.00
apiwebweb1.001.00
web——n/an/a
worker——n/an/a
reporting——n/an/a
0.98
Spearman rank correlation, predicted vs. measured blast-radius size
1.00
Mean recall — zero false negatives across all six experiments
0.83
Mean precision — one false positive, a timing artifact (explained below)
n=6
Experiments. Small by design; the protocol is published so you can grow n on your own cluster

Reading the result honestly

  • Zero false negatives is the number that matters most. A false negative is a real dependency the graph never saw — a missing edge, which would make every analysis built on the graph wrong. None occurred.
  • The one false positive is a timing artifact, not a wrong edge. Killing auth was predicted to reach web transitively (web → api → auth). The during-experiment snapshot caught the cascade mid-propagation: api had lost readiness, web's probe had not yet failed. The dependency is real; the instrument sampled before the second hop landed. This is the documented direction in which the method understates agreement.
  • Rank correlation is the honest headline because the product's claim is ordinal — “these are the workloads that can take your system down, worst first.” Predicting four failures where three occur is a good result if the order holds, and it held.
  • Ground truth paid for itself. The first run measured zero impact everywhere, which was impossible given the kubectl output. Cause: Kubernetes omits readyReplicas when zero pods are ready, and the collector treated the missing field as healthy. Found only because an independent measurement disagreed with the graph — which is the argument for this whole method.

Sample size, disclosed — and the protocol, published

This is six controlled experiments on a synthetic six-service topology; we publish the protocol so you can run it on your own cluster. The correlation is real but the sample is small — the method is the deliverable as much as the number.

  1. 1
    Predict first, freeze forever

    Blast radius is computed for every workload from the agent's snapshot before any experiment runs. Predictions are never revised after results arrive.

  2. 2
    Perturb

    Chaos Mesh PodChaos with action: pod-failure, mode: all, 90-second duration, selector scoped to one workload's labels in one namespace. pod-failure rather than pod-kill — a killed pod reschedules in seconds, faster than a dependent's readiness probe can observe, which reads as zero impact for a dependency that is entirely real.

  3. 3
    Measure

    A second agent snapshot during the window. Affected = was ready at baseline AND not ready during the experiment, excluding the target itself. Workloads unhealthy before the experiment are excluded — uncertainty may suppress a measurement, never fabricate one.

  4. 4
    Recover

    The experiment is deleted; the next run begins only after every deployment reports full readiness again.

What this validation does not establish

  • External validity to arbitrary estates. Six experiments on one synthetic topology on one cluster cannot, by themselves, tell you the correlation holds on your estate — different topologies, meshes, and failure modes are untested here. That is exactly why the protocol is reproducible: the roadmap item is on-your-cluster validation, running the same predict-freeze-perturb-measure loop against your staging environment so the result is about your topology, not ours.
  • Resilience-mechanism behaviour. Pod-level failure is the gentlest failure mode. Dependents with retries, caches, or circuit breakers may ride out an outage the graph correctly predicted — so measured agreement is a lower bound, but the untested modes are still untested.
  • Impact beyond readiness. Readiness is the only impact signal in this protocol. Latency degradation, error-rate rises, and partial brownouts are invisible to it.

Recovery curves: the metastable-failure question

The harness also measures what happens after the trigger clears. In a live run, a 45-second pod-failure on the database (with readiness deliberately delayed 30 seconds) was followed by snapshots every 6 seconds for 84 seconds after the experiment ended. The entire dependent cascade stayed degraded for exactly 30 seconds after the trigger was gone, then recovered together — no flapping, no metastable residue, matching the injected delay exactly. The instrument correctly distinguishes “slow but clean recovery” from the metastable signature, which is the distinction the measurement exists for.

Root-cause diagnosis, ground-truth tested

The diagnosis engine was validated the same way: against injected faults on the same six-service topology, where the root cause is known because we caused it. With a single injected root (the database, snapshot taken mid-cascade), the engine reported exactly one root at high confidence, attributed the downstream services as collateral requiring no separate investigation, and blamed no collateral workload. With two simultaneous independent roots injected at once (full 6/6 cascade), it reported exactly the two injected roots and nothing else — shared collateral was attributed to both rather than promoted to a phantom third incident. Both runs: correct roots, zero false accusations. The property under test is the one that costs on-call engineers 60–80% of MTTR — separating the first broken dependency path from the loudest alerts.

CEI: the Criticality-Entropy Index

CEI scores how dangerous each dependency concentration is by combining graph centrality (how many things need this), dependency entropy (how concentrated the need is), and operational risk signals. It is the number that says which service is quietly becoming the single point of failure nobody chose. Patent pending, USPTO App. No. 19/641,446, with additional filings in preparation covering blast-radius computation and related components.

Risk in dollars, not adjectives

Structural risk is priced by Monte-Carlo availability modeling against your downtime cost rate. Every recommendation carries its net annual benefit: risk reduction minus the cost of the change. High, medium, and low are not numbers; dollars per year are.