FAILURE MAP
← Case archive

FA-74386 / Experiment statistics / Open access

Conversion rate z-test: Zero-variance data is reported as significant · case 01

Arms with no conversions at all are reported with p = 0.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The se = 0 branch returns [0.0, 0.0].

THE FAILURE

The se = 0 branch returns [0.0, 0.0].

Unsuccessful approach: Returning None makes a valid if uninformative experiment look malformed.

Case contract

Pooled two-proportion z-test: pool = (conv_a + conv_b) / (n_a + n_b), se = sqrt(pool(1 - pool)(1/n_a + 1/n_b)), z = (p_b - p_a) / se, two-sided p = erfc(|z| / sqrt 2). Nonpositive n -> None; se = 0 -> [0.0, 1.0]. Return [round(z, 6), round(p, 6)].

Why this case matters

Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(conv_a, n_a, conv_b, n_b):
    if n_a <= 0 or n_b <= 0:
        return None
    pa, pb = conv_a / n_a, conv_b / n_b
    pool = (conv_a + conv_b) / (n_a + n_b)
    se = math.sqrt(pool * (1 - pool) * (1 / n_a + 1 / n_b))
    if se == 0:
        return [0.0, 0.0]
    z = (pb - pa) / se
    p = math.erfc(abs(z) / math.sqrt(2))
    return [round(z, 6), round(p, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('conversion sample 1', [88, 1000, 75, 500], [3.637163, 0.000276]),
  ('conversion sample 2', [446, 4000, 43, 2500], [-14.022988, 0.0]),
  ('conversion sample 3', [86, 500, 18, 100], [0.192927, 0.847016])],
 [('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('conversion sample 6', [455, 4000, 121, 1000], [0.642295, 0.520682]),
  ('conversion sample 7', [16, 500, 457, 2500], [8.446636, 0.0]),
  ('conversion sample 8', [70, 500, 206, 1000], [3.109778, 0.001872])],
 [('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 11', [263, 4000, 78, 1000], [1.374446, 0.169303]),
  ('conversion sample 12', [5, 100, 238, 1000], [4.320765, 1.6e-05]),
  ('conversion sample 13', [83, 1000, 484, 2500], [8.022534, 0.0])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 16', [200, 1000, 1, 100], [-4.687832, 3e-06]),
  ('conversion sample 17', [17, 100, 12, 100], [-1.004125, 0.315318]),
  ('conversion sample 18', [24, 100, 97, 500], [-1.046545, 0.295309])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 21', [20, 100, 104, 500], [0.180358, 0.856871]),
  ('conversion sample 22', [1, 100, 105, 500], [4.787119, 2e-06])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal arm sizes weight the pooled rate[-2.98032, 0.002879][-2.98032, 0.002879]Passed
treatment better gives positive z[2.102741, 0.035488][2.102741, 0.035488]Passed
treatment worse gives negative z[-2.102741, 0.035488][-2.102741, 0.035488]Passed
no conversions anywhere[0.0, 0.0][0.0, 1.0]Failed
all conversions everywhere[0.0, 0.0][0.0, 1.0]Failed
conversion sample 1[3.637163, 0.000276][3.637163, 0.000276]Passed
conversion sample 2[-14.022988, 0.0][-14.022988, 0.0]Passed
conversion sample 3[0.192927, 0.847016][0.192927, 0.847016]Passed

SHA-256 / 4d67beb996afb9a091a54399a244257b2cbba37b2db4e7c774d7f88c8dc7640d

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(conv_a, n_a, conv_b, n_b):
    if n_a <= 0 or n_b <= 0:
        return None
    pa, pb = conv_a / n_a, conv_b / n_b
    pool = (conv_a + conv_b) / (n_a + n_b)
    se = math.sqrt(pool * (1 - pool) * (1 / n_a + 1 / n_b))
    if se == 0:
        return None
    z = (pb - pa) / se
    p = math.erfc(abs(z) / math.sqrt(2))
    return [round(z, 6), round(p, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('conversion sample 1', [88, 1000, 75, 500], [3.637163, 0.000276]),
  ('conversion sample 2', [446, 4000, 43, 2500], [-14.022988, 0.0]),
  ('conversion sample 3', [86, 500, 18, 100], [0.192927, 0.847016])],
 [('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('conversion sample 6', [455, 4000, 121, 1000], [0.642295, 0.520682]),
  ('conversion sample 7', [16, 500, 457, 2500], [8.446636, 0.0]),
  ('conversion sample 8', [70, 500, 206, 1000], [3.109778, 0.001872])],
 [('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 11', [263, 4000, 78, 1000], [1.374446, 0.169303]),
  ('conversion sample 12', [5, 100, 238, 1000], [4.320765, 1.6e-05]),
  ('conversion sample 13', [83, 1000, 484, 2500], [8.022534, 0.0])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 16', [200, 1000, 1, 100], [-4.687832, 3e-06]),
  ('conversion sample 17', [17, 100, 12, 100], [-1.004125, 0.315318]),
  ('conversion sample 18', [24, 100, 97, 500], [-1.046545, 0.295309])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 21', [20, 100, 104, 500], [0.180358, 0.856871]),
  ('conversion sample 22', [1, 100, 105, 500], [4.787119, 2e-06])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal arm sizes weight the pooled rate[-2.98032, 0.002879][-2.98032, 0.002879]Passed
treatment better gives positive z[2.102741, 0.035488][2.102741, 0.035488]Passed
treatment worse gives negative z[-2.102741, 0.035488][-2.102741, 0.035488]Passed
no conversions anywhereNone[0.0, 1.0]Failed
all conversions everywhereNone[0.0, 1.0]Failed
conversion sample 1[3.637163, 0.000276][3.637163, 0.000276]Passed
conversion sample 2[-14.022988, 0.0][-14.022988, 0.0]Passed
conversion sample 3[0.192927, 0.847016][0.192927, 0.847016]Passed

SHA-256 / c66173e63974f61eebfcfef6295f08b4973b1b0573567988cfae51eb86b1f685

HELD IN THE MEMBER ARCHIVE

The verified repair and its recorded checks are member-only.

This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.

Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.

Member access is invitation-based. Sign in with your invited account to inspect the repair.

Sign in to the archive ↗

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.400745+00:00.

Case digest / f8bc647e0b2b047b2a1ca9b9d95a3755c4431d98ace3071df58b864c7461dffd