FAILURE MAP
← Case archive

FA-74381 / Experiment statistics / Open access

Conversion rate z-test: The z statistic has the control-minus-treatment sign · case 01

Dashboards show winning treatments as losing.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

z is computed as (p_a - p_b) / se.

VERIFIED REPAIR

Compute z as (p_b - p_a) / se.

Unsuccessful approach: Taking the absolute value loses the direction entirely.

Case contract

Pooled two-proportion z-test: pool = (conv_a + conv_b) / (n_a + n_b), se = sqrt(pool(1 - pool)(1/n_a + 1/n_b)), z = (p_b - p_a) / se, two-sided p = erfc(|z| / sqrt 2). Nonpositive n -> None; se = 0 -> [0.0, 1.0]. Return [round(z, 6), round(p, 6)].

Why this case matters

Online experiment readouts drive launch decisions; a silent formula slip flips conclusions.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(conv_a, n_a, conv_b, n_b):
    if n_a <= 0 or n_b <= 0:
        return None
    pa, pb = conv_a / n_a, conv_b / n_b
    pool = (conv_a + conv_b) / (n_a + n_b)
    se = math.sqrt(pool * (1 - pool) * (1 / n_a + 1 / n_b))
    if se == 0:
        return [0.0, 1.0]
    z = (pa - pb) / se
    p = math.erfc(abs(z) / math.sqrt(2))
    return [round(z, 6), round(p, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('conversion sample 1', [88, 1000, 75, 500], [3.637163, 0.000276]),
  ('conversion sample 2', [446, 4000, 43, 2500], [-14.022988, 0.0]),
  ('conversion sample 3', [86, 500, 18, 100], [0.192927, 0.847016])],
 [('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('conversion sample 5', [14, 100, 17, 100], [0.586154, 0.557772]),
  ('conversion sample 6', [455, 4000, 121, 1000], [0.642295, 0.520682]),
  ('conversion sample 19', [211, 1000, 115, 1000], [-5.811653, 0.0])],
 [('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 11', [263, 4000, 78, 1000], [1.374446, 0.169303]),
  ('conversion sample 12', [5, 100, 238, 1000], [4.320765, 1.6e-05]),
  ('conversion sample 36', [20, 100, 49, 500], [-2.918697, 0.003515])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 16', [200, 1000, 1, 100], [-4.687832, 3e-06]),
  ('conversion sample 19', [211, 1000, 115, 1000], [-5.811653, 0.0]),
  ('conversion sample 45', [646, 4000, 407, 4000], [-7.903693, 0.0])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 21', [20, 100, 104, 500], [0.180358, 0.856871]),
  ('conversion sample 26', [128, 4000, 22, 100], [9.890895, 0.0]),
  ('conversion sample 56', [869, 4000, 282, 4000], [-18.69958, 0.0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal arm sizes weight the pooled rate[2.98032, 0.002879][-2.98032, 0.002879]Failed
treatment better gives positive z[-2.102741, 0.035488][2.102741, 0.035488]Failed
treatment worse gives negative z[2.102741, 0.035488][-2.102741, 0.035488]Failed
no conversions anywhere[0.0, 1.0][0.0, 1.0]Passed
all conversions everywhere[0.0, 1.0][0.0, 1.0]Passed
conversion sample 1[-3.637163, 0.000276][3.637163, 0.000276]Failed
conversion sample 2[14.022988, 0.0][-14.022988, 0.0]Failed
conversion sample 3[-0.192927, 0.847016][0.192927, 0.847016]Failed

SHA-256 / 5ae22321b2b9e4e9dca3f036435d32bacd6d198ae580da2b9bbf0f5f4fd3ed9e

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(conv_a, n_a, conv_b, n_b):
    if n_a <= 0 or n_b <= 0:
        return None
    pa, pb = conv_a / n_a, conv_b / n_b
    pool = (conv_a + conv_b) / (n_a + n_b)
    se = math.sqrt(pool * (1 - pool) * (1 / n_a + 1 / n_b))
    if se == 0:
        return [0.0, 1.0]
    z = abs(pb - pa) / se
    p = math.erfc(abs(z) / math.sqrt(2))
    return [round(z, 6), round(p, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('conversion sample 1', [88, 1000, 75, 500], [3.637163, 0.000276]),
  ('conversion sample 2', [446, 4000, 43, 2500], [-14.022988, 0.0]),
  ('conversion sample 3', [86, 500, 18, 100], [0.192927, 0.847016])],
 [('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('conversion sample 5', [14, 100, 17, 100], [0.586154, 0.557772]),
  ('conversion sample 6', [455, 4000, 121, 1000], [0.642295, 0.520682]),
  ('conversion sample 19', [211, 1000, 115, 1000], [-5.811653, 0.0])],
 [('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 11', [263, 4000, 78, 1000], [1.374446, 0.169303]),
  ('conversion sample 12', [5, 100, 238, 1000], [4.320765, 1.6e-05]),
  ('conversion sample 36', [20, 100, 49, 500], [-2.918697, 0.003515])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 16', [200, 1000, 1, 100], [-4.687832, 3e-06]),
  ('conversion sample 19', [211, 1000, 115, 1000], [-5.811653, 0.0]),
  ('conversion sample 45', [646, 4000, 407, 4000], [-7.903693, 0.0])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 21', [20, 100, 104, 500], [0.180358, 0.856871]),
  ('conversion sample 26', [128, 4000, 22, 100], [9.890895, 0.0]),
  ('conversion sample 56', [869, 4000, 282, 4000], [-18.69958, 0.0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal arm sizes weight the pooled rate[2.98032, 0.002879][-2.98032, 0.002879]Failed
treatment better gives positive z[2.102741, 0.035488][2.102741, 0.035488]Passed
treatment worse gives negative z[2.102741, 0.035488][-2.102741, 0.035488]Failed
no conversions anywhere[0.0, 1.0][0.0, 1.0]Passed
all conversions everywhere[0.0, 1.0][0.0, 1.0]Passed
conversion sample 1[3.637163, 0.000276][3.637163, 0.000276]Passed
conversion sample 2[14.022988, 0.0][-14.022988, 0.0]Failed
conversion sample 3[0.192927, 0.847016][0.192927, 0.847016]Passed

SHA-256 / 2c68ccfe773aa7c3649210a9d7160dba91ac2d4c2a1218d99a47ff898bfe37a2

3 / The verified repair

Exit 0
"""Failure Map reference implementation. Python standard library only."""
import json
import math
N = 1
observations = []
def solve(conv_a, n_a, conv_b, n_b):
    if n_a <= 0 or n_b <= 0:
        return None
    pa, pb = conv_a / n_a, conv_b / n_b
    pool = (conv_a + conv_b) / (n_a + n_b)
    se = math.sqrt(pool * (1 - pool) * (1 / n_a + 1 / n_b))
    if se == 0:
        return [0.0, 1.0]
    z = (pb - pa) / se
    p = math.erfc(abs(z) / math.sqrt(2))
    return [round(z, 6), round(p, 6)]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('conversion sample 1', [88, 1000, 75, 500], [3.637163, 0.000276]),
  ('conversion sample 2', [446, 4000, 43, 2500], [-14.022988, 0.0]),
  ('conversion sample 3', [86, 500, 18, 100], [0.192927, 0.847016])],
 [('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('conversion sample 5', [14, 100, 17, 100], [0.586154, 0.557772]),
  ('conversion sample 6', [455, 4000, 121, 1000], [0.642295, 0.520682]),
  ('conversion sample 19', [211, 1000, 115, 1000], [-5.811653, 0.0])],
 [('treatment worse gives negative z', [130, 1000, 100, 1000], [-2.102741, 0.035488]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 11', [263, 4000, 78, 1000], [1.374446, 0.169303]),
  ('conversion sample 12', [5, 100, 238, 1000], [4.320765, 1.6e-05]),
  ('conversion sample 36', [20, 100, 49, 500], [-2.918697, 0.003515])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('no conversions anywhere', [0, 500, 0, 500], [0.0, 1.0]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 16', [200, 1000, 1, 100], [-4.687832, 3e-06]),
  ('conversion sample 19', [211, 1000, 115, 1000], [-5.811653, 0.0]),
  ('conversion sample 45', [646, 4000, 407, 4000], [-7.903693, 0.0])],
 [('unequal arm sizes weight the pooled rate', [50, 1000, 90, 3000], [-2.98032, 0.002879]),
  ('treatment better gives positive z', [100, 1000, 130, 1000], [2.102741, 0.035488]),
  ('all conversions everywhere', [200, 200, 300, 300], [0.0, 1.0]),
  ('identical rates', [40, 400, 80, 800], [0.0, 1.0]),
  ('empty arm', [0, 0, 5, 10], None),
  ('conversion sample 21', [20, 100, 104, 500], [0.180358, 0.856871]),
  ('conversion sample 26', [128, 4000, 22, 100], [9.890895, 0.0]),
  ('conversion sample 56', [869, 4000, 282, 4000], [-18.69958, 0.0])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
unequal arm sizes weight the pooled rate[-2.98032, 0.002879][-2.98032, 0.002879]Passed
treatment better gives positive z[2.102741, 0.035488][2.102741, 0.035488]Passed
treatment worse gives negative z[-2.102741, 0.035488][-2.102741, 0.035488]Passed
no conversions anywhere[0.0, 1.0][0.0, 1.0]Passed
all conversions everywhere[0.0, 1.0][0.0, 1.0]Passed
conversion sample 1[3.637163, 0.000276][3.637163, 0.000276]Passed
conversion sample 2[-14.022988, 0.0][-14.022988, 0.0]Passed
conversion sample 3[0.192927, 0.847016][0.192927, 0.847016]Passed

SHA-256 / 8cc7766c1ec7f290d849b7efd9a697b96cb569d818e746e866721ea34d6521ef

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:56.393980+00:00.

Case digest / 3ff02930b785bc85bcf8aba2ac58ab07164ad4fb32077b24b66978217143d61f