{"abstract":"Posterior means are too low because every trial is counted as a failure.","category":"Experiment statistics","checks":8,"contract":"With an integer Beta(prior_a, prior_b) prior, arm posteriors are Beta(prior_a + s, prior_b + n - s). P(B > A) uses the exact integer-parameter sum over i < alpha_B of exp(lnB(alpha_A + i, beta_A + beta_B) - ln(beta_B + i) - lnB(1 + i, beta_B) - lnB(alpha_A, beta_A)). Return [posterior mean A, posterior mean B, P(B > A)] rounded to 6.","evaluation_group":"w2-experiment-statistics-beta-binomial","failed_approach":"Dropping the prior from beta makes zero-failure arms degenerate and ignores prior strength.","family":"w2-experiment-statistics-beta-binomial-failure-count","id":"FA-74581","implementations":{"attempt":{"sha256":"18c3b6a2496f2a09fa97bfdd319733f9c585b7043b1e1d29e4c72f4ff66737d9","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\nimport math\nN = 1\nobservations = []\ndef solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):\n    aa, ba = prior_a + a_succ, max(a_n - a_succ, 1)\n    ab, bb = prior_a + b_succ, max(b_n - b_succ, 1)\n    def lbeta(x, y):\n        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)\n    total = 0.0\n    for i in range(ab):\n        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))\n    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],\n [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),\n  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),\n  ('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),\n  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814]),\n  ('posterior sample 12', [16, 20, 9, 20, 2, 1], [0.782609, 0.478261, 0.013376])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),\n  ('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),\n  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),\n  ('posterior sample 22', [7, 10, 18, 20, 1, 5], [0.5, 0.730769, 0.937656]),\n  ('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"broken":{"sha256":"bca2aff71569193262e33db2ca3eec720b7033b5d47c923f5400d80e244a5b04","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\nimport math\nN = 1\nobservations = []\ndef solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):\n    aa, ba = prior_a + a_succ, prior_b + a_n\n    ab, bb = prior_a + b_succ, prior_b + b_n\n    def lbeta(x, y):\n        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)\n    total = 0.0\n    for i in range(ab):\n        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))\n    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],\n [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),\n  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),\n  ('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),\n  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814]),\n  ('posterior sample 12', [16, 20, 9, 20, 2, 1], [0.782609, 0.478261, 0.013376])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),\n  ('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),\n  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),\n  ('posterior sample 22', [7, 10, 18, 20, 1, 5], [0.5, 0.730769, 0.937656]),\n  ('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"fixed":{"sha256":"4effd8f61813ac918ca6968a9095fedd2d529ddeb4dacfcdb0787d8a90fee7a9","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\nimport math\nN = 1\nobservations = []\ndef solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):\n    aa, ba = prior_a + a_succ, prior_b + a_n - a_succ\n    ab, bb = prior_a + b_succ, prior_b + b_n - b_succ\n    def lbeta(x, y):\n        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)\n    total = 0.0\n    for i in range(ab):\n        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))\n    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],\n [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0]),\n  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),\n  ('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 9', [7, 10, 3, 5, 1, 1], [0.666667, 0.571429, 0.33872]),\n  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814]),\n  ('posterior sample 12', [16, 20, 9, 20, 2, 1], [0.782609, 0.478261, 0.013376])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),\n  ('posterior sample 17', [9, 10, 8, 10, 1, 2], [0.769231, 0.692308, 0.320203]),\n  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),\n  ('posterior sample 22', [7, 10, 18, 20, 1, 5], [0.5, 0.730769, 0.937656]),\n  ('posterior sample 23', [18, 20, 0, 5, 1, 2], [0.826087, 0.125, 7.7e-05])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"}},"limitations":"A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.","method":"Deterministic executable model with adversarial boundary fixtures.","provenance":{"created_by":"Failure Map","dependencies":"Python standard library","family":"w2-experiment-statistics-beta-binomial-failure-count","generated_at":"2026-09-29T14:48:58.156183+00:00","license":"CC0-1.0","python":"3.12.14","seed":1,"split":"open-access"},"relevance":"Bayesian dashboards report a \"chance to beat control\" that product teams act on directly.","repair":"Add only the failures n - s to the beta parameter.","root_cause":"beta is prior_b + n rather than prior_b + n - successes.","sha256":"88fa9bcb248b9a6692083c1c88fc57ca86c7fd84908f2fb2eaeb6da7a7e6a989","title":"Bayesian conversion comparison: Posterior beta ignores observed successes · case 01","variant":1,"variant_policy":"Five numbered records share a model and may reuse boundary fixtures.","verification":{"attempt":{"elapsed_ms":37.956,"exit_code":1,"observations":[{"actual":[0.363636,0.636364,0.910552],"check":"uniform prior small counts","expected":[0.333333,0.583333,0.90081],"passed":false},{"actual":[0.285714,0.428571,0.727273],"check":"informative prior shifts means","expected":[0.166667,0.25,0.706767],"passed":false},{"actual":[0.454545,0.454545,0.5],"check":"equal data gives one half","expected":[0.416667,0.416667,0.5],"passed":false},{"actual":[0.761905,0.190476,4.4e-05],"check":"treatment clearly worse","expected":[0.727273,0.181818,6.3e-05],"passed":false},{"actual":[0.272727,0.090909,0.105263],"check":"zero successes in treatment","expected":[0.25,0.083333,0.107143],"passed":false},{"actual":[0.146341,0.363636,0.934789],"check":"unequal sample sizes","expected":[0.139535,0.307692,0.902022],"passed":false},{"actual":[0.434783,0.615385,0.859124],"check":"posterior sample 1","expected":[0.357143,0.444444,0.724072],"passed":false},{"actual":[0.121951,0.916667,1.0],"check":"posterior sample 2","expected":[0.119048,0.916667,1.0],"passed":false}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"uniform prior small counts\", \"actual\": [0.363636, 0.636364, 0.910552], \"expected\": [0.333333, 0.583333, 0.90081], \"passed\": false}, {\"check\": \"informative prior shifts means\", \"actual\": [0.285714, 0.428571, 0.727273], \"expected\": [0.166667, 0.25, 0.706767], \"passed\": false}, {\"check\": \"equal data gives one half\", \"actual\": [0.454545, 0.454545, 0.5], \"expected\": [0.416667, 0.416667, 0.5], \"passed\": false}, {\"check\": \"treatment clearly worse\", \"actual\": [0.761905, 0.190476, 4.4e-05], \"expected\": [0.727273, 0.181818, 6.3e-05], \"passed\": false}, {\"check\": \"zero successes in treatment\", \"actual\": [0.272727, 0.090909, 0.105263], \"expected\": [0.25, 0.083333, 0.107143], \"passed\": false}, {\"check\": \"unequal sample sizes\", \"actual\": [0.146341, 0.363636, 0.934789], \"expected\": [0.139535, 0.307692, 0.902022], \"passed\": false}, {\"check\": \"posterior sample 1\", \"actual\": [0.434783, 0.615385, 0.859124], \"expected\": [0.357143, 0.444444, 0.724072], \"passed\": false}, {\"check\": \"posterior sample 2\", \"actual\": [0.121951, 0.916667, 1.0], \"expected\": [0.119048, 0.916667, 1.0], \"passed\": false}], \"passed\": false}\n"},"broken":{"elapsed_ms":38.097,"exit_code":1,"observations":[{"actual":[0.266667,0.388889,0.782399],"check":"uniform prior small counts","expected":[0.333333,0.583333,0.90081],"passed":false},{"actual":[0.166667,0.230769,0.670807],"check":"informative prior shifts means","expected":[0.166667,0.25,0.706767],"passed":false},{"actual":[0.3125,0.3125,0.5],"check":"equal data gives one half","expected":[0.416667,0.416667,0.5],"passed":false},{"actual":[0.432432,0.16,0.008505],"check":"treatment clearly worse","expected":[0.727273,0.181818,6.3e-05],"passed":false},{"actual":[0.214286,0.083333,0.141304],"check":"zero successes in treatment","expected":[0.25,0.083333,0.107143],"passed":false},{"actual":[0.125,0.25,0.866026],"check":"unequal sample sizes","expected":[0.139535,0.307692,0.902022],"passed":false},{"actual":[0.285714,0.347826,0.68942],"check":"posterior sample 1","expected":[0.357143,0.444444,0.724072],"passed":false},{"actual":[0.108696,0.5,0.999788],"check":"posterior sample 2","expected":[0.119048,0.916667,1.0],"passed":false}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"uniform prior small counts\", \"actual\": [0.266667, 0.388889, 0.782399], \"expected\": [0.333333, 0.583333, 0.90081], \"passed\": false}, {\"check\": \"informative prior shifts means\", \"actual\": [0.166667, 0.230769, 0.670807], \"expected\": [0.166667, 0.25, 0.706767], \"passed\": false}, {\"check\": \"equal data gives one half\", \"actual\": [0.3125, 0.3125, 0.5], \"expected\": [0.416667, 0.416667, 0.5], \"passed\": false}, {\"check\": \"treatment clearly worse\", \"actual\": [0.432432, 0.16, 0.008505], \"expected\": [0.727273, 0.181818, 6.3e-05], \"passed\": false}, {\"check\": \"zero successes in treatment\", \"actual\": [0.214286, 0.083333, 0.141304], \"expected\": [0.25, 0.083333, 0.107143], \"passed\": false}, {\"check\": \"unequal sample sizes\", \"actual\": [0.125, 0.25, 0.866026], \"expected\": [0.139535, 0.307692, 0.902022], \"passed\": false}, {\"check\": \"posterior sample 1\", \"actual\": [0.285714, 0.347826, 0.68942], \"expected\": [0.357143, 0.444444, 0.724072], \"passed\": false}, {\"check\": \"posterior sample 2\", \"actual\": [0.108696, 0.5, 0.999788], \"expected\": [0.119048, 0.916667, 1.0], \"passed\": false}], \"passed\": false}\n"},"fixed":{"elapsed_ms":39.013,"exit_code":0,"observations":[{"actual":[0.333333,0.583333,0.90081],"check":"uniform prior small counts","expected":[0.333333,0.583333,0.90081],"passed":true},{"actual":[0.166667,0.25,0.706767],"check":"informative prior shifts means","expected":[0.166667,0.25,0.706767],"passed":true},{"actual":[0.416667,0.416667,0.5],"check":"equal data gives one half","expected":[0.416667,0.416667,0.5],"passed":true},{"actual":[0.727273,0.181818,6.3e-05],"check":"treatment clearly worse","expected":[0.727273,0.181818,6.3e-05],"passed":true},{"actual":[0.25,0.083333,0.107143],"check":"zero successes in treatment","expected":[0.25,0.083333,0.107143],"passed":true},{"actual":[0.139535,0.307692,0.902022],"check":"unequal sample sizes","expected":[0.139535,0.307692,0.902022],"passed":true},{"actual":[0.357143,0.444444,0.724072],"check":"posterior sample 1","expected":[0.357143,0.444444,0.724072],"passed":true},{"actual":[0.119048,0.916667,1.0],"check":"posterior sample 2","expected":[0.119048,0.916667,1.0],"passed":true}],"passed":true,"stderr":"","stdout":"{\"observations\": [{\"check\": \"uniform prior small counts\", \"actual\": [0.333333, 0.583333, 0.90081], \"expected\": [0.333333, 0.583333, 0.90081], \"passed\": true}, {\"check\": \"informative prior shifts means\", \"actual\": [0.166667, 0.25, 0.706767], \"expected\": [0.166667, 0.25, 0.706767], \"passed\": true}, {\"check\": \"equal data gives one half\", \"actual\": [0.416667, 0.416667, 0.5], \"expected\": [0.416667, 0.416667, 0.5], \"passed\": true}, {\"check\": \"treatment clearly worse\", \"actual\": [0.727273, 0.181818, 6.3e-05], \"expected\": [0.727273, 0.181818, 6.3e-05], \"passed\": true}, {\"check\": \"zero successes in treatment\", \"actual\": [0.25, 0.083333, 0.107143], \"expected\": [0.25, 0.083333, 0.107143], \"passed\": true}, {\"check\": \"unequal sample sizes\", \"actual\": [0.139535, 0.307692, 0.902022], \"expected\": [0.139535, 0.307692, 0.902022], \"passed\": true}, {\"check\": \"posterior sample 1\", \"actual\": [0.357143, 0.444444, 0.724072], \"expected\": [0.357143, 0.444444, 0.724072], \"passed\": true}, {\"check\": \"posterior sample 2\", \"actual\": [0.119048, 0.916667, 1.0], \"expected\": [0.119048, 0.916667, 1.0], \"passed\": true}], \"passed\": true}\n"}},"verified":true,"visibility":"public"}