FAILURE MAP
← Case archive

FA-74796 / Experiment statistics / Open access

Always-valid sequential p-value: The prefactor ignores accumulated sample size · case 01

Likelihood ratios are too large for long streams, overstating significance.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The square-root prefactor uses sigma2 + tau2 instead of sigma2 + n tau2.

VERIFIED REPAIR

Use sqrt(sigma2 / (sigma2 + n tau2)).

Unsuccessful approach: Dropping the square root over-penalises the prefactor.

Case contract

diffs is a stream of paired differences with known variance sigma2; the normal-mixture likelihood ratio after n observations with running mean m is sqrt(sigma2 / (sigma2 + n tau2)) exp(n^2 tau2 m^2 / (2 sigma2 (sigma2 + n tau2))). The always-valid p-value starts at 1 and is the running minimum of 1 / ratio. Return the p-value after each observation rounded to 6.

Why this case matters

Always-valid p-values let teams monitor continuously without inflating false positives.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(diffs, sigma2, tau2):
    p = 1.0
    total = 0.0
    out = []
    for n, d in enumerate(diffs, 1):
        total += d
        mean = total / n
        lam = math.sqrt(sigma2 / (sigma2 + tau2)) * math.exp(n * n * tau2 * mean * mean / (2 * sigma2 * (sigma2 + n * tau2)))
        p = min(p, 1 / lam)
        out.append(round(p, 6))
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 1',
   [[0.0, 1.5, 0.0, -1.0, 0.0, 0.5, -1.0], 2.0, 0.5],
   [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 2',
   [[0.0, 2.0, -1.0, 1.0, 0.5], 1.0, 1.0],
   [1.0, 0.889265, 0.889265, 0.889265, 0.889265]),
  ('difference stream sample 3',
   [[1.0, -1.0, -1.0, 0.5, -1.0, 0.5, 1.5], 1.0, 0.25],
   [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 6', [[0.0, -1.0, 0.0, -1.0, 0.5], 4.0, 0.25], [1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 8', [[1.0, 1.0], 2.0, 1.0], [1.0, 1.0]),
  ('difference stream sample 13', [[2.0], 2.0, 0.5], [0.915369])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 11', [[2.0], 4.0, 1.0], [1.0]),
  ('difference stream sample 22', [[1.0, 0.5, 1.5], 2.0, 0.25], [1.0, 1.0, 0.955693]),
  ('difference stream sample 29', [[1.5], 2.0, 0.5], [0.999072])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 16', [[1.0], 1.0, 0.25], [1.0]),
  ('difference stream sample 40', [[2.0, 0.5], 1.0, 1.0], [0.52026, 0.52026]),
  ('difference stream sample 56', [[2.0, 2.0, 1.0, 2.0], 4.0, 0.5], [1.0, 0.915369, 0.882617, 0.735148])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 4',
   [[0.0, 2.0, 1.0, 0.0, 1.0, 2.0, 1.0], 1.0, 1.0],
   [1.0, 0.889265, 0.649305, 0.649305, 0.645678, 0.202205, 0.132287]),
  ('difference stream sample 21', [[0.0, -1.0, 1.0, 2.0, 1.5], 2.0, 0.25], [1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 48',
   [[2.0, 1.5, 1.5, -1.0, 2.0, 1.5, 2.0], 1.0, 0.25],
   [0.749441, 0.441269, 0.221816, 0.221816, 0.203003, 0.094955, 0.02742])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
p-value never increases after a reversal[0.52026, 0.098264, 0.098264, 0.098264][0.52026, 0.120349, 0.120349, 0.120349]Failed
first observation[1.0][1.0]Passed
null-looking stream stays at one[1.0, 1.0, 1.0][1.0, 1.0, 1.0]Passed
steady effect[1.0, 0.946395, 0.8107, 0.678122, 0.558292][1.0, 1.0, 0.959234, 0.857764, 0.749028]Failed
noisy stream[1.0, 1.0, 0.999965, 0.971365][1.0, 1.0, 1.0, 1.0]Failed
difference stream sample 1[1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0][1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]Passed
difference stream sample 2[1.0, 0.726081, 0.726081, 0.726081, 0.726081][1.0, 0.889265, 0.889265, 0.889265, 0.889265]Failed
difference stream sample 3[1.0, 1.0, 1.0, 1.0, 0.986662, 0.986662, 0.986662][1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]Failed

SHA-256 / bbeaddf5d9fa6a1840e2fc6d8f501d0195409b99c322ce8ec71fb8eb50db747c

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(diffs, sigma2, tau2):
    p = 1.0
    total = 0.0
    out = []
    for n, d in enumerate(diffs, 1):
        total += d
        mean = total / n
        lam = (sigma2 / (sigma2 + n * tau2)) * math.exp(n * n * tau2 * mean * mean / (2 * sigma2 * (sigma2 + n * tau2)))
        p = min(p, 1 / lam)
        out.append(round(p, 6))
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 1',
   [[0.0, 1.5, 0.0, -1.0, 0.0, 0.5, -1.0], 2.0, 0.5],
   [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 2',
   [[0.0, 2.0, -1.0, 1.0, 0.5], 1.0, 1.0],
   [1.0, 0.889265, 0.889265, 0.889265, 0.889265]),
  ('difference stream sample 3',
   [[1.0, -1.0, -1.0, 0.5, -1.0, 0.5, 1.5], 1.0, 0.25],
   [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 6', [[0.0, -1.0, 0.0, -1.0, 0.5], 4.0, 0.25], [1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 8', [[1.0, 1.0], 2.0, 1.0], [1.0, 1.0]),
  ('difference stream sample 13', [[2.0], 2.0, 0.5], [0.915369])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 11', [[2.0], 4.0, 1.0], [1.0]),
  ('difference stream sample 22', [[1.0, 0.5, 1.5], 2.0, 0.25], [1.0, 1.0, 0.955693]),
  ('difference stream sample 29', [[1.5], 2.0, 0.5], [0.999072])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 16', [[1.0], 1.0, 0.25], [1.0]),
  ('difference stream sample 40', [[2.0, 0.5], 1.0, 1.0], [0.52026, 0.52026]),
  ('difference stream sample 56', [[2.0, 2.0, 1.0, 2.0], 4.0, 0.5], [1.0, 0.915369, 0.882617, 0.735148])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 4',
   [[0.0, 2.0, 1.0, 0.0, 1.0, 2.0, 1.0], 1.0, 1.0],
   [1.0, 0.889265, 0.649305, 0.649305, 0.645678, 0.202205, 0.132287]),
  ('difference stream sample 21', [[0.0, -1.0, 1.0, 2.0, 1.5], 2.0, 0.25], [1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 48',
   [[2.0, 1.5, 1.5, -1.0, 2.0, 1.5, 2.0], 1.0, 0.25],
   [0.749441, 0.441269, 0.221816, 0.221816, 0.203003, 0.094955, 0.02742])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
p-value never increases after a reversal[0.735759, 0.20845, 0.20845, 0.20845][0.52026, 0.120349, 0.120349, 0.120349]Failed
first observation[1.0][1.0]Passed
null-looking stream stays at one[1.0, 1.0, 1.0][1.0, 1.0, 1.0]Passed
steady effect[1.0, 1.0, 1.0, 1.0, 1.0][1.0, 1.0, 0.959234, 0.857764, 0.749028]Failed
noisy stream[1.0, 1.0, 1.0, 1.0][1.0, 1.0, 1.0, 1.0]Passed
difference stream sample 1[1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0][1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]Passed
difference stream sample 2[1.0, 1.0, 1.0, 1.0, 1.0][1.0, 0.889265, 0.889265, 0.889265, 0.889265]Failed
difference stream sample 3[1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0][1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]Passed

SHA-256 / e29f4d63a04b9e392a0688220e1be99119a693faca06c9e6c8bec9d8422500ba

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(diffs, sigma2, tau2):
    p = 1.0
    total = 0.0
    out = []
    for n, d in enumerate(diffs, 1):
        total += d
        mean = total / n
        lam = math.sqrt(sigma2 / (sigma2 + n * tau2)) * math.exp(n * n * tau2 * mean * mean / (2 * sigma2 * (sigma2 + n * tau2)))
        p = min(p, 1 / lam)
        out.append(round(p, 6))
    return out
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 1',
   [[0.0, 1.5, 0.0, -1.0, 0.0, 0.5, -1.0], 2.0, 0.5],
   [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 2',
   [[0.0, 2.0, -1.0, 1.0, 0.5], 1.0, 1.0],
   [1.0, 0.889265, 0.889265, 0.889265, 0.889265]),
  ('difference stream sample 3',
   [[1.0, -1.0, -1.0, 0.5, -1.0, 0.5, 1.5], 1.0, 0.25],
   [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 6', [[0.0, -1.0, 0.0, -1.0, 0.5], 4.0, 0.25], [1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 8', [[1.0, 1.0], 2.0, 1.0], [1.0, 1.0]),
  ('difference stream sample 13', [[2.0], 2.0, 0.5], [0.915369])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 11', [[2.0], 4.0, 1.0], [1.0]),
  ('difference stream sample 22', [[1.0, 0.5, 1.5], 2.0, 0.25], [1.0, 1.0, 0.955693]),
  ('difference stream sample 29', [[1.5], 2.0, 0.5], [0.999072])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 16', [[1.0], 1.0, 0.25], [1.0]),
  ('difference stream sample 40', [[2.0, 0.5], 1.0, 1.0], [0.52026, 0.52026]),
  ('difference stream sample 56', [[2.0, 2.0, 1.0, 2.0], 4.0, 0.5], [1.0, 0.915369, 0.882617, 0.735148])],
 [('p-value never increases after a reversal',
   [[2.0, 2.0, -1.0, -1.0], 1.0, 1.0],
   [0.52026, 0.120349, 0.120349, 0.120349]),
  ('first observation', [[1.0], 1.0, 1.0], [1.0]),
  ('null-looking stream stays at one', [[0.0, 0.0, 0.0], 1.0, 0.5], [1.0, 1.0, 1.0]),
  ('steady effect', [[1.0, 1.0, 1.0, 1.0, 1.0], 2.0, 0.5], [1.0, 1.0, 0.959234, 0.857764, 0.749028]),
  ('noisy stream', [[1.5, -1.0, 2.0, 0.5], 4.0, 1.0], [1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 4',
   [[0.0, 2.0, 1.0, 0.0, 1.0, 2.0, 1.0], 1.0, 1.0],
   [1.0, 0.889265, 0.649305, 0.649305, 0.645678, 0.202205, 0.132287]),
  ('difference stream sample 21', [[0.0, -1.0, 1.0, 2.0, 1.5], 2.0, 0.25], [1.0, 1.0, 1.0, 1.0, 1.0]),
  ('difference stream sample 48',
   [[2.0, 1.5, 1.5, -1.0, 2.0, 1.5, 2.0], 1.0, 0.25],
   [0.749441, 0.441269, 0.221816, 0.221816, 0.203003, 0.094955, 0.02742])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
p-value never increases after a reversal[0.52026, 0.120349, 0.120349, 0.120349][0.52026, 0.120349, 0.120349, 0.120349]Passed
first observation[1.0][1.0]Passed
null-looking stream stays at one[1.0, 1.0, 1.0][1.0, 1.0, 1.0]Passed
steady effect[1.0, 1.0, 0.959234, 0.857764, 0.749028][1.0, 1.0, 0.959234, 0.857764, 0.749028]Passed
noisy stream[1.0, 1.0, 1.0, 1.0][1.0, 1.0, 1.0, 1.0]Passed
difference stream sample 1[1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0][1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]Passed
difference stream sample 2[1.0, 0.889265, 0.889265, 0.889265, 0.889265][1.0, 0.889265, 0.889265, 0.889265, 0.889265]Passed
difference stream sample 3[1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0][1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]Passed

SHA-256 / f48077d486ea03e18e4400cf44be2d8be60a4f390d34b47145a91ebfaa848398

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:49:00.142122+00:00.

Case digest / 063a80ca8b7da26c704b03c93458c3205862173ff26122f500065c39d4aee854