FAILURE MAP
← Case archive

FA-74696 / Experiment statistics / Open access

Complier average causal effect: Control crossover is ignored in compliance · case 01

CACE is understated when control users can reach the feature through another path.

Verified by executionVariant 1 · 8 checks per implementationDownload source bundle ↓JSON ↗

ROOT CAUSE

The compliance difference is just the treatment uptake rate.

THE FAILURE

The compliance difference is just the treatment uptake rate.

Unsuccessful approach: Subtracting raw counts over the treatment size ignores unequal arm sizes.

Case contract

Outcome sums cover every assigned user. ITT = outcome_sum_t / assigned_t - outcome_sum_c / assigned_c. Compliance difference = took_t / assigned_t - took_c / assigned_c (control users can cross over). CACE = ITT / compliance difference when that difference is positive, else None. Nonpositive assignment counts -> None. Return [ITT, compliance difference, CACE] rounded to 6.

Why this case matters

Opt-in features have partial uptake; the effect on adopters needs an instrumental-variable estimate.

1 / The failure

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(assigned_c, took_c, outcome_sum_c, assigned_t, took_t, outcome_sum_t):
    if assigned_c <= 0 or assigned_t <= 0:
        return None
    itt = outcome_sum_t / assigned_t - outcome_sum_c / assigned_c
    comp = took_t / assigned_t
    cace = round(itt / comp, 6) if comp > 0 else None
    return [round(itt, 6), round(comp, 6), cace]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 1', [200, 22, 428, 400, 23, 389], [-1.1675, -0.0525, None]),
  ('uptake sample 2', [100, 22, 279, 400, 165, 505], [-1.5275, 0.1925, -7.935065]),
  ('uptake sample 3', [100, 11, 241, 400, 240, 17], [-2.3675, 0.49, -4.831633])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 5', [200, 62, 357, 300, 290, 266], [-0.898333, 0.656667, -1.36802]),
  ('uptake sample 6', [400, 10, 585, 300, 223, 613], [0.580833, 0.718333, 0.808585]),
  ('uptake sample 12', [100, 3, 98, 400, 319, 1150], [1.895, 0.7675, 2.469055])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 11', [400, 10, 635, 100, 93, 173], [0.1425, 0.905, 0.157459]),
  ('uptake sample 12', [100, 3, 98, 400, 319, 1150], [1.895, 0.7675, 2.469055]),
  ('uptake sample 19', [200, 30, 146, 300, 293, 69], [-0.5, 0.826667, -0.604839])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 16', [100, 6, 298, 300, 104, 511], [-1.276667, 0.286667, -4.453488]),
  ('uptake sample 19', [200, 30, 146, 300, 293, 69], [-0.5, 0.826667, -0.604839]),
  ('uptake sample 29', [400, 11, 731, 100, 70, 53], [-1.2975, 0.6725, -1.929368])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 21', [100, 18, 78, 100, 98, 114], [0.36, 0.8, 0.45]),
  ('uptake sample 26', [200, 20, 354, 300, 180, 2], [-1.763333, 0.5, -3.526667]),
  ('uptake sample 39', [200, 59, 406, 400, 358, 158], [-1.635, 0.6, -2.725])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
crossover in control reduces compliance difference[0.6, 0.6, 1.0][0.6, 0.5, 1.2]Failed
unequal arm sizes[0.5, 0.5, 1.0][0.5, 0.5, 1.0]Passed
no uptake difference has no CACE[0.2, 0.5, 0.4][0.2, 0.0, None]Failed
low but positive compliance still yields CACE[0.3, 0.3, 1.0][0.3, 0.3, 1.0]Passed
negative compliance difference[-0.1, 0.4, -0.25][-0.1, -0.2, None]Failed
uptake sample 1[-1.1675, 0.0575, -20.304348][-1.1675, -0.0525, None]Failed
uptake sample 2[-1.5275, 0.4125, -3.70303][-1.5275, 0.1925, -7.935065]Failed
uptake sample 3[-2.3675, 0.6, -3.945833][-2.3675, 0.49, -4.831633]Failed

SHA-256 / 4d4328a897fa475ea8b8fde21b5acf24404b9656c86eb6287a0d4abae558f08a

2 / The unsuccessful fix

Exit 1
"""Failure Map reference implementation. Python standard library only."""
import json

N = 1
observations = []
def solve(assigned_c, took_c, outcome_sum_c, assigned_t, took_t, outcome_sum_t):
    if assigned_c <= 0 or assigned_t <= 0:
        return None
    itt = outcome_sum_t / assigned_t - outcome_sum_c / assigned_c
    comp = (took_t - took_c) / assigned_t
    cace = round(itt / comp, 6) if comp > 0 else None
    return [round(itt, 6), round(comp, 6), cace]
def check(label, actual, expected):
    observations.append({"check": label, "actual": actual, "expected": expected, "passed": actual == expected})
fixtures = [[('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 1', [200, 22, 428, 400, 23, 389], [-1.1675, -0.0525, None]),
  ('uptake sample 2', [100, 22, 279, 400, 165, 505], [-1.5275, 0.1925, -7.935065]),
  ('uptake sample 3', [100, 11, 241, 400, 240, 17], [-2.3675, 0.49, -4.831633])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 5', [200, 62, 357, 300, 290, 266], [-0.898333, 0.656667, -1.36802]),
  ('uptake sample 6', [400, 10, 585, 300, 223, 613], [0.580833, 0.718333, 0.808585]),
  ('uptake sample 12', [100, 3, 98, 400, 319, 1150], [1.895, 0.7675, 2.469055])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 11', [400, 10, 635, 100, 93, 173], [0.1425, 0.905, 0.157459]),
  ('uptake sample 12', [100, 3, 98, 400, 319, 1150], [1.895, 0.7675, 2.469055]),
  ('uptake sample 19', [200, 30, 146, 300, 293, 69], [-0.5, 0.826667, -0.604839])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 16', [100, 6, 298, 300, 104, 511], [-1.276667, 0.286667, -4.453488]),
  ('uptake sample 19', [200, 30, 146, 300, 293, 69], [-0.5, 0.826667, -0.604839]),
  ('uptake sample 29', [400, 11, 731, 100, 70, 53], [-1.2975, 0.6725, -1.929368])],
 [('crossover in control reduces compliance difference', [200, 20, 300, 200, 120, 420], [0.6, 0.5, 1.2]),
  ('unequal arm sizes', [100, 0, 150, 400, 200, 800], [0.5, 0.5, 1.0]),
  ('no uptake difference has no CACE', [100, 50, 100, 100, 50, 120], [0.2, 0.0, None]),
  ('low but positive compliance still yields CACE', [100, 0, 100, 100, 30, 130], [0.3, 0.3, 1.0]),
  ('negative compliance difference', [100, 60, 100, 100, 40, 90], [-0.1, -0.2, None]),
  ('uptake sample 21', [100, 18, 78, 100, 98, 114], [0.36, 0.8, 0.45]),
  ('uptake sample 26', [200, 20, 354, 300, 180, 2], [-1.763333, 0.5, -3.526667]),
  ('uptake sample 39', [200, 59, 406, 400, 358, 158], [-1.635, 0.6, -2.725])]]
for label, args, expected in fixtures[N - 1]:
    check(label, solve(*args), expected)
print(json.dumps({"observations": observations, "passed": all(x["passed"] for x in observations)}, ensure_ascii=False))
raise SystemExit(0 if all(x["passed"] for x in observations) else 1)
Boundary fixtureActualExpectedOutcome
crossover in control reduces compliance difference[0.6, 0.5, 1.2][0.6, 0.5, 1.2]Passed
unequal arm sizes[0.5, 0.5, 1.0][0.5, 0.5, 1.0]Passed
no uptake difference has no CACE[0.2, 0.0, None][0.2, 0.0, None]Passed
low but positive compliance still yields CACE[0.3, 0.3, 1.0][0.3, 0.3, 1.0]Passed
negative compliance difference[-0.1, -0.2, None][-0.1, -0.2, None]Passed
uptake sample 1[-1.1675, 0.0025, -467.0][-1.1675, -0.0525, None]Failed
uptake sample 2[-1.5275, 0.3575, -4.272727][-1.5275, 0.1925, -7.935065]Failed
uptake sample 3[-2.3675, 0.5725, -4.135371][-2.3675, 0.49, -4.831633]Failed

SHA-256 / e6673d13f4063be5b27b8db4aaffd53d951149388f3258bda408cc710b10b75e

HELD IN THE MEMBER ARCHIVE

The verified repair and its recorded checks are member-only.

This mechanism has 8 recorded checks per implementation. The open-access tier publishes the failure and the unsuccessful fix; the repaired source that passes every check, and the observations that prove it, are available to members.

Every case sharing this mechanism uses the same contract and the same repair, so this one record is held back for all of them.

Member access is invitation-based. Sign in with your invited account to inspect the repair.

Sign in to the archive ↗

Verification & scope

A deterministic toy experiment-analysis model with a stipulated contract; results are rounded and are not a substitute for a validated statistics package. This reproducer isolates one failure mechanism. Results cover the supplied fixtures. Variants within a family share a test contract and should remain grouped when constructing evaluation splits. Related mechanisms with a shared evaluation_group must also remain together; these controlled models are not independent production incidents.

Observations recorded using Python 3.12.14 at 2026-09-29T14:48:59.157804+00:00.

Case digest / 60859a83e4ad9602fc870a73476b3c3415badf3f250e62c926e4423c608a87a0