{"abstract":"Probabilities can exceed one or fall far below the truth.","category":"Experiment statistics","checks":8,"contract":"With an integer Beta(prior_a, prior_b) prior, arm posteriors are Beta(prior_a + s, prior_b + n - s). P(B > A) uses the exact integer-parameter sum over i < alpha_B of exp(lnB(alpha_A + i, beta_A + beta_B) - ln(beta_B + i) - lnB(1 + i, beta_B) - lnB(alpha_A, beta_A)). Return [posterior mean A, posterior mean B, P(B > A)] rounded to 6.","evaluation_group":"w2-experiment-statistics-beta-binomial","failed_approach":"Mixing alpha_A with beta_B in the normaliser is still the wrong beta function.","family":"w2-experiment-statistics-beta-binomial-normaliser","id":"FA-74596","implementations":{"attempt":{"sha256":"8f1f314be111bcec8a365c3ea13461b12d4b8754f13d858289711154edb17971","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\nimport math\nN = 1\nobservations = []\ndef solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):\n    aa, ba = prior_a + a_succ, prior_b + a_n - a_succ\n    ab, bb = prior_a + b_succ, prior_b + b_n - b_succ\n    def lbeta(x, y):\n        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)\n    total = 0.0\n    for i in range(ab):\n        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, bb))\n    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],\n [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 3', [8, 10, 23, 40, 1, 5], [0.5625, 0.521739, 0.384139]),\n  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),\n  ('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814]),\n  ('posterior sample 12', [16, 20, 9, 20, 2, 1], [0.782609, 0.478261, 0.013376]),\n  ('posterior sample 13', [7, 10, 2, 5, 3, 1], [0.714286, 0.555556, 0.212934])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),\n  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848]),\n  ('posterior sample 19', [12, 40, 7, 40, 1, 1], [0.309524, 0.190476, 0.098949])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),\n  ('posterior sample 25', [10, 10, 5, 5, 2, 1], [0.923077, 0.875, 0.368421]),\n  ('posterior sample 28', [0, 5, 4, 5, 3, 1], [0.333333, 0.777778, 0.97972])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"broken":{"sha256":"0938038dc045b7b6e80fc9dac20479478ee7f09ca0a118ceef4eb26ed9c2b40d","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\nimport math\nN = 1\nobservations = []\ndef solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):\n    aa, ba = prior_a + a_succ, prior_b + a_n - a_succ\n    ab, bb = prior_a + b_succ, prior_b + b_n - b_succ\n    def lbeta(x, y):\n        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)\n    total = 0.0\n    for i in range(ab):\n        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(ab, bb))\n    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],\n [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 3', [8, 10, 23, 40, 1, 5], [0.5625, 0.521739, 0.384139]),\n  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),\n  ('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814]),\n  ('posterior sample 12', [16, 20, 9, 20, 2, 1], [0.782609, 0.478261, 0.013376]),\n  ('posterior sample 13', [7, 10, 2, 5, 3, 1], [0.714286, 0.555556, 0.212934])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),\n  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848]),\n  ('posterior sample 19', [12, 40, 7, 40, 1, 1], [0.309524, 0.190476, 0.098949])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),\n  ('posterior sample 25', [10, 10, 5, 5, 2, 1], [0.923077, 0.875, 0.368421]),\n  ('posterior sample 28', [0, 5, 4, 5, 3, 1], [0.333333, 0.777778, 0.97972])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"},"fixed":{"sha256":"1bd5852be89ee07433721684af5ede4a59413f84ba80ac439c7b6fcafee20bb9","source":"\"\"\"Failure Map reference implementation. Python standard library only.\"\"\"\nimport json\nimport math\nN = 1\nobservations = []\ndef solve(a_succ, a_n, b_succ, b_n, prior_a, prior_b):\n    aa, ba = prior_a + a_succ, prior_b + a_n - a_succ\n    ab, bb = prior_a + b_succ, prior_b + b_n - b_succ\n    def lbeta(x, y):\n        return math.lgamma(x) + math.lgamma(y) - math.lgamma(x + y)\n    total = 0.0\n    for i in range(ab):\n        total += math.exp(lbeta(aa + i, ba + bb) - math.log(bb + i) - lbeta(1 + i, bb) - lbeta(aa, ba))\n    return [round(aa / (aa + ba), 6), round(ab / (ab + bb), 6), round(total, 6)]\ndef check(label, actual, expected):\n    observations.append({\"check\": label, \"actual\": actual, \"expected\": expected, \"passed\": actual == expected})\nfixtures = [[('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 1', [7, 20, 5, 10, 3, 5], [0.357143, 0.444444, 0.724072]),\n  ('posterior sample 2', [4, 40, 10, 10, 1, 1], [0.119048, 0.916667, 1.0])],\n [('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 3', [8, 10, 23, 40, 1, 5], [0.5625, 0.521739, 0.384139]),\n  ('posterior sample 6', [13, 20, 0, 10, 2, 2], [0.625, 0.142857, 0.000906]),\n  ('posterior sample 7', [2, 5, 2, 5, 2, 1], [0.5, 0.5, 0.5])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 11', [5, 5, 30, 40, 1, 1], [0.857143, 0.738095, 0.1814]),\n  ('posterior sample 12', [16, 20, 9, 20, 2, 1], [0.782609, 0.478261, 0.013376]),\n  ('posterior sample 13', [7, 10, 2, 5, 3, 1], [0.714286, 0.555556, 0.212934])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('treatment clearly worse', [15, 20, 3, 20, 1, 1], [0.727273, 0.181818, 6.3e-05]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 16', [7, 10, 1, 20, 1, 2], [0.615385, 0.086957, 0.000212]),\n  ('posterior sample 18', [8, 40, 14, 20, 1, 2], [0.209302, 0.652174, 0.999848]),\n  ('posterior sample 19', [12, 40, 7, 40, 1, 1], [0.309524, 0.190476, 0.098949])],\n [('uniform prior small counts', [3, 10, 6, 10, 1, 1], [0.333333, 0.583333, 0.90081]),\n  ('informative prior shifts means', [0, 5, 1, 5, 2, 5], [0.166667, 0.25, 0.706767]),\n  ('equal data gives one half', [4, 10, 4, 10, 1, 1], [0.416667, 0.416667, 0.5]),\n  ('zero successes in treatment', [2, 10, 0, 10, 1, 1], [0.25, 0.083333, 0.107143]),\n  ('unequal sample sizes', [5, 40, 3, 10, 1, 2], [0.139535, 0.307692, 0.902022]),\n  ('posterior sample 21', [6, 10, 16, 40, 1, 1], [0.583333, 0.404762, 0.132122]),\n  ('posterior sample 25', [10, 10, 5, 5, 2, 1], [0.923077, 0.875, 0.368421]),\n  ('posterior sample 28', [0, 5, 4, 5, 3, 1], [0.333333, 0.777778, 0.97972])]]\nfor label, args, expected in fixtures[N - 1]:\n    check(label, solve(*args), expected)\nprint(json.dumps({\"observations\": observations, \"passed\": all(x[\"passed\"] for x in observations)}, ensure_ascii=False))\nraise SystemExit(0 if all(x[\"passed\"] for x in observations) else 1)\n"}},"limitations":"A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.","method":"Deterministic executable model with adversarial boundary fixtures.","provenance":{"created_by":"Failure Map","dependencies":"Python standard library","family":"w2-experiment-statistics-beta-binomial-normaliser","generated_at":"2026-09-29T14:48:58.282518+00:00","license":"CC0-1.0","python":"3.12.14","seed":1,"split":"open-access"},"relevance":"Bayesian dashboards report a \"chance to beat control\" that product teams act on directly.","repair":"Normalise each term by the control posterior beta function.","root_cause":"The last term subtracts lnB(alpha_B, beta_B) instead of lnB(alpha_A, beta_A).","sha256":"1e658ee5fbded2075edd6d48651eb0e3206aa90f769d48c2f3fc1445c3ce1eae","title":"Bayesian conversion comparison: The sum is normalised by the treatment posterior · case 01","variant":1,"variant_policy":"Five numbered records share a model and may reuse boundary fixtures.","verification":{"attempt":{"elapsed_ms":39.126,"exit_code":1,"observations":[{"actual":[0.333333,0.583333,0.191081],"check":"uniform prior small counts","expected":[0.333333,0.583333,0.90081],"passed":false},{"actual":[0.166667,0.25,0.578264],"check":"informative prior shifts means","expected":[0.166667,0.25,0.706767],"passed":false},{"actual":[0.416667,0.416667,0.5],"check":"equal data gives one half","expected":[0.416667,0.416667,0.5],"passed":true},{"actual":[0.727273,0.181818,3.638359],"check":"treatment clearly worse","expected":[0.727273,0.181818,6.3e-05],"passed":false},{"actual":[0.25,0.083333,0.185714],"check":"zero successes in treatment","expected":[0.25,0.083333,0.107143],"passed":false},{"actual":[0.139535,0.307692,0.000516],"check":"unequal sample sizes","expected":[0.139535,0.307692,0.902022],"passed":false},{"actual":[0.357143,0.444444,0.007929],"check":"posterior sample 1","expected":[0.357143,0.444444,0.724072],"passed":false},{"actual":[0.119048,0.916667,1e-06],"check":"posterior sample 2","expected":[0.119048,0.916667,1.0],"passed":false}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"uniform prior small counts\", \"actual\": [0.333333, 0.583333, 0.191081], \"expected\": [0.333333, 0.583333, 0.90081], \"passed\": false}, {\"check\": \"informative prior shifts means\", \"actual\": [0.166667, 0.25, 0.578264], \"expected\": [0.166667, 0.25, 0.706767], \"passed\": false}, {\"check\": \"equal data gives one half\", \"actual\": [0.416667, 0.416667, 0.5], \"expected\": [0.416667, 0.416667, 0.5], \"passed\": true}, {\"check\": \"treatment clearly worse\", \"actual\": [0.727273, 0.181818, 3.638359], \"expected\": [0.727273, 0.181818, 6.3e-05], \"passed\": false}, {\"check\": \"zero successes in treatment\", \"actual\": [0.25, 0.083333, 0.185714], \"expected\": [0.25, 0.083333, 0.107143], \"passed\": false}, {\"check\": \"unequal sample sizes\", \"actual\": [0.139535, 0.307692, 0.000516], \"expected\": [0.139535, 0.307692, 0.902022], \"passed\": false}, {\"check\": \"posterior sample 1\", \"actual\": [0.357143, 0.444444, 0.007929], \"expected\": [0.357143, 0.444444, 0.724072], \"passed\": false}, {\"check\": \"posterior sample 2\", \"actual\": [0.119048, 0.916667, 1e-06], \"expected\": [0.119048, 0.916667, 1.0], \"passed\": false}], \"passed\": false}\n"},"broken":{"elapsed_ms":39.44,"exit_code":1,"observations":[{"actual":[0.333333,0.583333,1.576417],"check":"uniform prior small counts","expected":[0.333333,0.583333,0.90081],"passed":false},{"actual":[0.166667,0.25,3.180451],"check":"informative prior shifts means","expected":[0.166667,0.25,0.706767],"passed":false},{"actual":[0.416667,0.416667,0.5],"check":"equal data gives one half","expected":[0.416667,0.416667,0.5],"passed":true},{"actual":[0.727273,0.181818,5e-06],"check":"treatment clearly worse","expected":[0.727273,0.181818,6.3e-05],"passed":false},{"actual":[0.25,0.083333,0.002381],"check":"zero successes in treatment","expected":[0.25,0.083333,0.107143],"passed":false},{"actual":[0.139535,0.307692,5.7e-05],"check":"unequal sample sizes","expected":[0.139535,0.307692,0.902022],"passed":false},{"actual":[0.357143,0.444444,0.001669],"check":"posterior sample 1","expected":[0.357143,0.444444,0.724072],"passed":false},{"actual":[0.119048,0.916667,3e-06],"check":"posterior sample 2","expected":[0.119048,0.916667,1.0],"passed":false}],"passed":false,"stderr":"","stdout":"{\"observations\": [{\"check\": \"uniform prior small counts\", \"actual\": [0.333333, 0.583333, 1.576417], \"expected\": [0.333333, 0.583333, 0.90081], \"passed\": false}, {\"check\": \"informative prior shifts means\", \"actual\": [0.166667, 0.25, 3.180451], \"expected\": [0.166667, 0.25, 0.706767], \"passed\": false}, {\"check\": \"equal data gives one half\", \"actual\": [0.416667, 0.416667, 0.5], \"expected\": [0.416667, 0.416667, 0.5], \"passed\": true}, {\"check\": \"treatment clearly worse\", \"actual\": [0.727273, 0.181818, 5e-06], \"expected\": [0.727273, 0.181818, 6.3e-05], \"passed\": false}, {\"check\": \"zero successes in treatment\", \"actual\": [0.25, 0.083333, 0.002381], \"expected\": [0.25, 0.083333, 0.107143], \"passed\": false}, {\"check\": \"unequal sample sizes\", \"actual\": [0.139535, 0.307692, 5.7e-05], \"expected\": [0.139535, 0.307692, 0.902022], \"passed\": false}, {\"check\": \"posterior sample 1\", \"actual\": [0.357143, 0.444444, 0.001669], \"expected\": [0.357143, 0.444444, 0.724072], \"passed\": false}, {\"check\": \"posterior sample 2\", \"actual\": [0.119048, 0.916667, 3e-06], \"expected\": [0.119048, 0.916667, 1.0], \"passed\": false}], \"passed\": false}\n"},"fixed":{"elapsed_ms":41.278,"exit_code":0,"observations":[{"actual":[0.333333,0.583333,0.90081],"check":"uniform prior small counts","expected":[0.333333,0.583333,0.90081],"passed":true},{"actual":[0.166667,0.25,0.706767],"check":"informative prior shifts means","expected":[0.166667,0.25,0.706767],"passed":true},{"actual":[0.416667,0.416667,0.5],"check":"equal data gives one half","expected":[0.416667,0.416667,0.5],"passed":true},{"actual":[0.727273,0.181818,6.3e-05],"check":"treatment clearly worse","expected":[0.727273,0.181818,6.3e-05],"passed":true},{"actual":[0.25,0.083333,0.107143],"check":"zero successes in treatment","expected":[0.25,0.083333,0.107143],"passed":true},{"actual":[0.139535,0.307692,0.902022],"check":"unequal sample sizes","expected":[0.139535,0.307692,0.902022],"passed":true},{"actual":[0.357143,0.444444,0.724072],"check":"posterior sample 1","expected":[0.357143,0.444444,0.724072],"passed":true},{"actual":[0.119048,0.916667,1.0],"check":"posterior sample 2","expected":[0.119048,0.916667,1.0],"passed":true}],"passed":true,"stderr":"","stdout":"{\"observations\": [{\"check\": \"uniform prior small counts\", \"actual\": [0.333333, 0.583333, 0.90081], \"expected\": [0.333333, 0.583333, 0.90081], \"passed\": true}, {\"check\": \"informative prior shifts means\", \"actual\": [0.166667, 0.25, 0.706767], \"expected\": [0.166667, 0.25, 0.706767], \"passed\": true}, {\"check\": \"equal data gives one half\", \"actual\": [0.416667, 0.416667, 0.5], \"expected\": [0.416667, 0.416667, 0.5], \"passed\": true}, {\"check\": \"treatment clearly worse\", \"actual\": [0.727273, 0.181818, 6.3e-05], \"expected\": [0.727273, 0.181818, 6.3e-05], \"passed\": true}, {\"check\": \"zero successes in treatment\", \"actual\": [0.25, 0.083333, 0.107143], \"expected\": [0.25, 0.083333, 0.107143], \"passed\": true}, {\"check\": \"unequal sample sizes\", \"actual\": [0.139535, 0.307692, 0.902022], \"expected\": [0.139535, 0.307692, 0.902022], \"passed\": true}, {\"check\": \"posterior sample 1\", \"actual\": [0.357143, 0.444444, 0.724072], \"expected\": [0.357143, 0.444444, 0.724072], \"passed\": true}, {\"check\": \"posterior sample 2\", \"actual\": [0.119048, 0.916667, 1.0], \"expected\": [0.119048, 0.916667, 1.0], \"passed\": true}], \"passed\": true}\n"}},"verified":true,"visibility":"public"}