FA-74386 / Experiment statistics / Open access
Conversion rate z-test: Zero-variance data is reported as significant · case 01
Arms with no conversions at all are reported with p = 0.
ROOT CAUSE
The se = 0 branch returns [0.0, 0.0].
THE FAILURE
The se = 0 branch returns [0.0, 0.0].
Unsuccessful approach: Returning None makes a valid if uninformative experiment look malformed.
Case contract
Pooled two-proportion z-test: pool = (conv_a + conv_b) / (n_a + n_b), se = sqrt(pool(1 - pool)(1/n_a + 1/n_b)), z = (p_b - p_a) / se, two-sided p = erfc(|z| / sqrt 2). Nonpositive n -> None; se = 0 -> [0.0, 1.0]. Return [round(z, 6), round(p, 6)].
Why this case matters
Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.
1 / The failure
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(conv_a, n_a, conv_b, n_b):
if n_a <= 0 or n_b <= 0:
return None
pa, pb = conv_a / n_a, conv_b / n_b
pool = (conv_a + conv_b) / (n_a + n_b)
se = math.sqrt(pool * (1 - pool) * (1 / n_a + 1 / n_b))
if se == 0:
return [0.0, 0.0]
z = (pb - pa) / se
p = math.erfc(abs(z) / math.sqrt(2))
return [round(z, 6), round(p, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('conversion sample 1', [88, 1000, 75, 500], [3.637163, 0.000276]),
('conversion sample 2', [446, 4000, 43, 2500], [-14.022988, 0.0]),
('conversion sample 3', [86, 500, 18, 100], [0.192927, 0.847016])],
[('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('conversion sample 6', [455, 4000, 121, 1000], [0.642295, 0.520682]),
('conversion sample 7', [16, 500, 457, 2500], [8.446636, 0.0]),
('conversion sample 8', [70, 500, 206, 1000], [3.109778, 0.001872])],
[('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('empty arm', [0, 0, 5, 10], None),
('conversion sample 11', [263, 4000, 78, 1000], [1.374446, 0.169303]),
('conversion sample 12', [5, 100, 238, 1000], [4.320765, 1.6e-05]),
('conversion sample 13', [83, 1000, 484, 2500], [8.022534, 0.0])],
[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('empty arm', [0, 0, 5, 10], None),
('conversion sample 16', [200, 1000, 1, 100], [-4.687832, 3e-06]),
('conversion sample 17', [17, 100, 12, 100], [-1.004125, 0.315318]),
('conversion sample 18', [24, 100, 97, 500], [-1.046545, 0.295309])],
[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('empty arm', [0, 0, 5, 10], None),
('conversion sample 21', [20, 100, 104, 500], [0.180358, 0.856871]),
('conversion sample 22', [1, 100, 105, 500], [4.787119, 2e-06])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal arm sizes weight the pooled rate | [-2.98032, 0.002879] | [-2.98032, 0.002879] | Passed |
| treatment better gives positive z | [2.102741, 0.035488] | [2.102741, 0.035488] | Passed |
| treatment worse gives negative z | [-2.102741, 0.035488] | [-2.102741, 0.035488] | Passed |
| no conversions anywhere | [0.0, 0.0] | [0.0, 1.0] | Failed |
| all conversions everywhere | [0.0, 0.0] | [0.0, 1.0] | Failed |
| conversion sample 1 | [3.637163, 0.000276] | [3.637163, 0.000276] | Passed |
| conversion sample 2 | [-14.022988, 0.0] | [-14.022988, 0.0] | Passed |
| conversion sample 3 | [0.192927, 0.847016] | [0.192927, 0.847016] | Passed |
SHA-256 / 4d67beb996afb9a091a54399a244257b2cbba37b2db4e7c774d7f88c8dc7640d
2 / The unsuccessful fix
Exit 1"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(conv_a, n_a, conv_b, n_b):
if n_a <= 0 or n_b <= 0:
return None
pa, pb = conv_a / n_a, conv_b / n_b
pool = (conv_a + conv_b) / (n_a + n_b)
se = math.sqrt(pool * (1 - pool) * (1 / n_a + 1 / n_b))
if se == 0:
return None
z = (pb - pa) / se
p = math.erfc(abs(z) / math.sqrt(2))
return [round(z, 6), round(p, 6)]
def check(label, actual, expected):
observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('conversion sample 1', [88, 1000, 75, 500], [3.637163, 0.000276]),
('conversion sample 2', [446, 4000, 43, 2500], [-14.022988, 0.0]),
('conversion sample 3', [86, 500, 18, 100], [0.192927, 0.847016])],
[('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('conversion sample 6', [455, 4000, 121, 1000], [0.642295, 0.520682]),
('conversion sample 7', [16, 500, 457, 2500], [8.446636, 0.0]),
('conversion sample 8', [70, 500, 206, 1000], [3.109778, 0.001872])],
[('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('empty arm', [0, 0, 5, 10], None),
('conversion sample 11', [263, 4000, 78, 1000], [1.374446, 0.169303]),
('conversion sample 12', [5, 100, 238, 1000], [4.320765, 1.6e-05]),
('conversion sample 13', [83, 1000, 484, 2500], [8.022534, 0.0])],
[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('empty arm', [0, 0, 5, 10], None),
('conversion sample 16', [200, 1000, 1, 100], [-4.687832, 3e-06]),
('conversion sample 17', [17, 100, 12, 100], [-1.004125, 0.315318]),
('conversion sample 18', [24, 100, 97, 500], [-1.046545, 0.295309])],
[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
('empty arm', [0, 0, 5, 10], None),
('conversion sample 21', [20, 100, 104, 500], [0.180358, 0.856871]),
('conversion sample 22', [1, 100, 105, 500], [4.787119, 2e-06])]]
for label, args, expected in fixtures[N - 1]:
check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
| Boundary fixture | Actual | Expected | Outcome |
|---|---|---|---|
| unequal arm sizes weight the pooled rate | [-2.98032, 0.002879] | [-2.98032, 0.002879] | Passed |
| treatment better gives positive z | [2.102741, 0.035488] | [2.102741, 0.035488] | Passed |
| treatment worse gives negative z | [-2.102741, 0.035488] | [-2.102741, 0.035488] | Passed |
| no conversions anywhere | None | [0.0, 1.0] | Failed |
| all conversions everywhere | None | [0.0, 1.0] | Failed |
| conversion sample 1 | [3.637163, 0.000276] | [3.637163, 0.000276] | Passed |
| conversion sample 2 | [-14.022988, 0.0] | [-14.022988, 0.0] | Passed |
| conversion sample 3 | [0.192927, 0.847016] | [0.192927, 0.847016] | Passed |
SHA-256 / c66173e63974f61eebfcfef6295f08b4973b1b0573567988cfae51eb86b1f685
HELD IN THE MEMBER ARCHIVE
The verified repair and its recorded checks are member-only.
This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.
Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.
Member access is invitation-based. Sign in with your invited account to inspect the repair.
Sign in to the archive ↗Verification & scope
A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.
Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.400745+00:00.
Case digest / f8bc647e0b2b047b2a1ca9b9d95a3755c4431d98ace3071df58b864c7461dffd