feat(phase-18/21): fairness criteria group individual counterfactual

This commit is contained in:
Rohit Ghumare
2026-04-24 12:21:51 +01:00
parent 2908c87bb8
commit 6691cbd4cb
5 changed files with 321 additions and 0 deletions
@@ -0,0 +1,58 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 520" font-family="Georgia, 'Times New Roman', serif">
<defs>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
</style>
</defs>
<text x="480" y="24" text-anchor="middle" class="title">Fairness: three families, one impossibility</text>
<rect x="40" y="60" width="280" height="200" class="box"/>
<text x="180" y="85" text-anchor="middle" class="head">group fairness</text>
<rect x="60" y="100" width="240" height="50" class="cool"/>
<text x="180" y="122" text-anchor="middle" class="step">demographic parity</text>
<text x="180" y="140" text-anchor="middle" class="small">P(Y=1 | A=a) equal</text>
<rect x="60" y="160" width="240" height="50" class="cool"/>
<text x="180" y="182" text-anchor="middle" class="step">equalized odds</text>
<text x="180" y="200" text-anchor="middle" class="small">TPR/FPR equal across groups</text>
<rect x="60" y="220" width="240" height="30" class="cool"/>
<text x="180" y="242" text-anchor="middle" class="step">conditional use accuracy equality</text>
<rect x="340" y="60" width="280" height="200" class="box"/>
<text x="480" y="85" text-anchor="middle" class="head">individual fairness</text>
<rect x="360" y="100" width="240" height="80" class="cold"/>
<text x="480" y="125" text-anchor="middle" class="step">Dwork et al. 2012</text>
<text x="480" y="145" text-anchor="middle" class="small">|f(x) - f(x')| &lt;= L * d(x, x')</text>
<text x="480" y="165" text-anchor="middle" class="small">Lipschitz; d task-specific</text>
<rect x="360" y="190" width="240" height="60" class="cold"/>
<text x="480" y="215" text-anchor="middle" class="step">similar individuals</text>
<text x="480" y="233" text-anchor="middle" class="small">get similar decisions</text>
<rect x="640" y="60" width="280" height="200" class="box"/>
<text x="780" y="85" text-anchor="middle" class="head">counterfactual fairness</text>
<rect x="660" y="100" width="240" height="80" class="hot"/>
<text x="780" y="125" text-anchor="middle" class="step">Kusner et al. 2017</text>
<text x="780" y="145" text-anchor="middle" class="small">invariant under attribute</text>
<text x="780" y="165" text-anchor="middle" class="small">counterfactual</text>
<rect x="660" y="190" width="240" height="60" class="hot"/>
<text x="780" y="215" text-anchor="middle" class="step">needs causal DAG</text>
<text x="780" y="233" text-anchor="middle" class="small">backtracking (2024) sidesteps</text>
<rect x="40" y="290" width="880" height="200" class="box"/>
<text x="480" y="315" text-anchor="middle" class="head">impossibility + reconciliation</text>
<text x="60" y="345" class="small">Chouldechova, KMR 2017: under unequal base rates, the three group criteria cannot all hold.</text>
<text x="60" y="365" class="small">policy choice: demographic parity gives equal access; equalized odds preserves accuracy equity;</text>
<text x="60" y="385" class="small">conditional use accuracy equality preserves predictive-value equity. each has a constituency.</text>
<text x="60" y="415" class="small">2024 NeurIPS: CF-accuracy trade-off is bounded; model-agnostic conversion of optimal-unfair -&gt; CF.</text>
<text x="60" y="435" class="small">backtracking counterfactuals (arXiv:2401.13935): avoid intervening on protected attributes.</text>
<text x="60" y="455" class="small">ICLR 2024 reconciliation: with explicit causal graphs, group and counterfactual are facets of one structure.</text>
<text x="60" y="475" class="small">impossibility still holds on base rates; reconciliation is about what is being measured.</text>
</svg>

After

Width:  |  Height:  |  Size: 4.1 KiB

@@ -0,0 +1,134 @@
"""Three group-fairness criteria on a toy classifier — stdlib Python.
Binary classification: sensitive attribute A in {0, 1} with unequal base rates.
A simple logistic classifier is trained; we report:
demographic parity, equalized odds, conditional use accuracy equality.
Then apply a re-weighting targeted at demographic parity and observe the
cost on the other two.
Usage: python3 code/main.py
"""
from __future__ import annotations
import math
import random
random.seed(53)
def gen(n: int) -> list[tuple[list[float], int, int]]:
"""Returns list of (features, label, sensitive_attribute).
Base rate differs by group: A=0 has P(y=1)=0.3; A=1 has P(y=1)=0.6.
Features correlate with y with some noise."""
data = []
for _ in range(n):
a = random.choice([0, 1])
base = 0.3 if a == 0 else 0.6
y = 1 if random.random() < base else 0
x0 = random.gauss(0.8 * y, 1.0)
x1 = random.gauss(-0.3 + a * 0.5, 1.0)
data.append(([x0, x1, float(a)], y, a))
return data
def train(data, steps: int = 200, lr: float = 0.1, sample_weights=None) -> list[float]:
w = [0.0, 0.0, 0.0]
b = 0.0
for _ in range(steps):
random.shuffle(data)
for idx, (x, y, a) in enumerate(data):
z = b + sum(wi * xi for wi, xi in zip(w, x))
p = 1.0 / (1.0 + math.exp(-z))
err = p - y
wt = 1.0 if sample_weights is None else sample_weights[idx]
for i in range(3):
w[i] -= lr * wt * err * x[i]
b -= lr * wt * err
return w + [b]
def predict(model, data):
w, b = model[:3], model[3]
preds = []
for x, y, a in data:
z = b + sum(wi * xi for wi, xi in zip(w, x))
preds.append((1 if z > 0 else 0, y, a))
return preds
def demographic_parity(preds) -> tuple[float, float]:
rate0 = sum(1 for p, _, a in preds if a == 0 and p == 1) / max(1, sum(1 for _, _, a in preds if a == 0))
rate1 = sum(1 for p, _, a in preds if a == 1 and p == 1) / max(1, sum(1 for _, _, a in preds if a == 1))
return rate0, rate1
def equalized_odds(preds) -> tuple[tuple, tuple]:
def group(a):
sub = [(p, y) for p, y, aa in preds if aa == a]
tpr = sum(1 for p, y in sub if y == 1 and p == 1) / max(1, sum(1 for _, y in sub if y == 1))
fpr = sum(1 for p, y in sub if y == 0 and p == 1) / max(1, sum(1 for _, y in sub if y == 0))
return tpr, fpr
return group(0), group(1)
def conditional_use(preds) -> tuple[tuple, tuple]:
def group(a):
sub = [(p, y) for p, y, aa in preds if aa == a]
ppv = sum(1 for p, y in sub if p == 1 and y == 1) / max(1, sum(1 for p, _ in sub if p == 1))
npv = sum(1 for p, y in sub if p == 0 and y == 0) / max(1, sum(1 for p, _ in sub if p == 0))
return ppv, npv
return group(0), group(1)
def report(name: str, preds):
dp = demographic_parity(preds)
eo = equalized_odds(preds)
cu = conditional_use(preds)
print(f"\n{name}")
print(f" demographic parity : group0={dp[0]:.3f} group1={dp[1]:.3f} gap={dp[1]-dp[0]:+.3f}")
print(f" equalized odds (TPR) : group0={eo[0][0]:.3f} group1={eo[1][0]:.3f}")
print(f" equalized odds (FPR) : group0={eo[0][1]:.3f} group1={eo[1][1]:.3f}")
print(f" conditional use (PPV) : group0={cu[0][0]:.3f} group1={cu[1][0]:.3f}")
print(f" conditional use (NPV) : group0={cu[0][1]:.3f} group1={cu[1][1]:.3f}")
def main() -> None:
print("=" * 70)
print("THREE GROUP-FAIRNESS CRITERIA (Phase 18, Lesson 21)")
print("=" * 70)
train_data = gen(1000)
test_data = gen(500)
baseline = train(train_data)
preds = predict(baseline, test_data)
report("baseline classifier", preds)
# Reweight toward demographic parity: upweight group0 y=1 and downweight group1 y=1.
weights = []
for x, y, a in train_data:
if a == 0 and y == 1:
weights.append(2.0)
elif a == 1 and y == 1:
weights.append(0.5)
else:
weights.append(1.0)
dp_reweighted = train(train_data, sample_weights=weights)
preds2 = predict(dp_reweighted, test_data)
report("DP-reweighted classifier", preds2)
print("\n" + "=" * 70)
print("TAKEAWAY: equal base rates are the condition for the three criteria")
print("to coincide. with unequal base rates, DP-targeted reweighting")
print("reduces the DP gap at the cost of equalized odds and conditional")
print("use accuracy. this is Chouldechova / KMR 2017 in miniature. the")
print("choice of criterion is a policy decision; no statistical method")
print("can satisfy all three under unequal base rates.")
print("=" * 70)
if __name__ == "__main__":
main()
@@ -0,0 +1,100 @@
# Fairness Criteria — Group, Individual, Counterfactual
> Three families structure the fairness literature. Group fairness: demographic parity, equalized odds, conditional use accuracy equality — equal rates across protected groups on average. Individual fairness (Dwork et al. 2012): similar individuals receive similar decisions; Lipschitz condition on the decision map. Counterfactual fairness (Kusner et al. 2017): a decision is fair to an individual if it is unchanged when sensitive attributes are counterfactually altered. 2024 theoretical result (NeurIPS 2024): there is an inherent CF-vs-accuracy trade-off; a model-agnostic method converts an optimal-but-unfair predictor into a CF one with bounded accuracy loss. Backtracking counterfactuals (arXiv:2401.13935, January 2024): new paradigm that avoids requiring interventions on legally protected attributes. Philosophical reconciliation (ICLR Blogposts 2024): with causal graphs, satisfying certain group fairness measures entails counterfactual fairness.
**Type:** Learn
**Languages:** Python (stdlib, three-criteria comparison)
**Prerequisites:** Phase 18 · 20 (bias), Phase 02 (classical ML)
**Time:** ~60 minutes
## Learning Objectives
- State the three group-fairness criteria (demographic parity, equalized odds, conditional use accuracy equality) and one impossibility result.
- Describe individual fairness via the Dwork et al. 2012 Lipschitz formulation.
- Describe counterfactual fairness and its causal-graph dependency.
- Explain backtracking counterfactuals and why they sidestep the intervention-on-protected-attribute problem.
## The Problem
Lesson 20 was about measuring bias. Lesson 21 is about defining the fairness standard the measurement should serve. The three families give structurally different standards — a model can be group-fair and individual-unfair, counterfactually fair and group-unfair. Choosing a standard is a policy decision; no standard is universally optimal.
## The Concept
### Group fairness
- **Demographic parity.** P(Y=1 | A=a) = P(Y=1 | A=a') for all groups. Equal acceptance rates.
- **Equalized odds.** P(Y=1 | Y*=y, A=a) = P(Y=1 | Y*=y, A=a'). Equal TPR and FPR across groups.
- **Conditional use accuracy equality.** P(Y*=y | Y=y, A=a) = P(Y*=y | Y=y, A=a'). Equal predictive value across groups.
Impossibility (Chouldechova, Kleinberg-Mullainathan-Raghavan 2017): these three cannot be satisfied simultaneously under unequal base rates.
### Individual fairness
Dwork et al. 2012. A decision map f is individually fair with respect to a task-specific similarity metric d if |f(x) - f(x')| <= L * d(x, x') for some Lipschitz constant L. Similar individuals get similar decisions.
Requires defining d. Policy question, not statistical.
### Counterfactual fairness
Kusner et al. 2017. A decision is counterfactually fair to individual i if, under a causal model of the population, the decision is unchanged when i's sensitive attributes are counterfactually altered.
Requires a causal DAG. The DAG is a modeling choice. Counterfactual fairness is only as justified as the DAG.
### The CF-vs-accuracy trade-off
NeurIPS 2024 theoretical: there is an inherent trade-off between counterfactual fairness and predictive accuracy. A model-agnostic method can convert an optimal-but-unfair predictor into a CF one, at a bounded accuracy cost. The accuracy cost depends on the magnitude of the sensitive-attribute coefficient in the optimal unfair predictor.
### Backtracking counterfactuals
arXiv:2401.13935 (January 2024). Traditional counterfactuals require interventions on the sensitive attribute — "would the decision change if this person had been a different gender." Legally, this is problematic: protected attributes cannot be intervened on in classification law.
Backtracking counterfactuals flip the direction: instead of intervening on the attribute, ask what combination of the individual's actual features would have produced the counterfactual outcome. This sidesteps the legal objection.
### Philosophical reconciliation
ICLR Blogposts 2024. With a causal graph in hand, satisfying certain group-fairness measures entails counterfactual fairness. The three families are not orthogonal; they are different facets of the same underlying causal structure.
This does not resolve the impossibility theorems (unequal base rates still prevent simultaneous group fairness). But it shows the apparent opposition between "group" and "individual / counterfactual" is partially an artifact of not being explicit about the causal model.
### Where this fits in Phase 18
Lesson 20 is bias measurement. Lesson 21 is fairness definition. Lesson 22 is privacy (differential privacy). Lesson 23 is watermarking. These are the allocation-adjacent lessons complementing the deception-adjacent Lessons 7-11.
## Use It
`code/main.py` builds a toy binary-classification dataset with a sensitive attribute and unequal base rates. Compute demographic parity, equalized odds, and conditional use accuracy equality on a simple classifier. Observe the three metrics disagreeing. Apply a re-weighting for demographic parity and observe its cost on the other two.
## Ship It
This lesson produces `outputs/skill-fairness-criterion.md`. Given a fairness claim or policy, identifies which criterion is being claimed, whether the model can satisfy the remaining criteria under the claimed unequal base rates, and what causal DAG the claim depends on.
## Exercises
1. Run `code/main.py`. Report the three group metrics on the default data. Apply the demographic-parity-targeted re-weighting and re-report.
2. Implement the Dwork et al. 2012 individual-fairness metric using L2 on non-sensitive features. Report how many pairs violate Lipschitz with constant L=1.
3. Read Kusner et al. 2017. Construct a simple two-feature causal DAG for resume scoring and identify the counterfactual-fairness condition it implies.
4. The 2024 backtracking-counterfactuals paper avoids intervention on protected attributes. Describe a scenario where this matters for legal compliance.
5. The ICLR 2024 reconciliation argues group and counterfactual fairness are facets of the same structure. Pick two of the three criteria in `code/main.py` and state the causal assumption that would make them equivalent.
## Key Terms
| Term | What people say | What it actually means |
|------|-----------------|------------------------|
| Demographic parity | "equal rates" | P(Y=1 | A=a) equal across groups |
| Equalized odds | "equal TPR/FPR" | Equal true-positive and false-positive rates across groups |
| Conditional use accuracy | "equal PPV/NPV" | Equal predictive values across groups |
| Individual fairness | "Lipschitz condition" | Similar individuals get similar decisions |
| Counterfactual fairness | "causal alteration invariance" | Decision unchanged under counterfactual attribute alteration |
| Backtracking counterfactual | "explain via actuals" | Counterfactual reasoned backward from outcome, not forward from attribute |
| Impossibility theorem | "the three conflict" | Chouldechova / KMR 2017: group criteria mutually exclusive under unequal base rates |
## Further Reading
- [Dwork et al. — Fairness through Awareness (arXiv:1104.3913)](https://arxiv.org/abs/1104.3913) — individual fairness
- [Kusner, Loftus, Russell, Silva — Counterfactual Fairness (arXiv:1703.06856)](https://arxiv.org/abs/1703.06856) — counterfactual fairness
- [Chouldechova — Fair prediction with disparate impact (arXiv:1703.00056)](https://arxiv.org/abs/1703.00056) — impossibility
- [Backtracking Counterfactuals (arXiv:2401.13935)](https://arxiv.org/abs/2401.13935) — new paradigm for protected-attribute interventions
@@ -0,0 +1,29 @@
---
name: fairness-criterion
description: Identify which fairness criterion a claim invokes and audit the associated assumptions.
version: 1.0.0
phase: 18
lesson: 21
tags: [fairness, demographic-parity, equalized-odds, counterfactual-fairness, impossibility]
---
Given a fairness claim or policy, identify which criterion is being invoked, what assumptions the claim depends on, and what the impossibility theorems imply for the remaining criteria.
Produce:
1. Criterion identification. Label the claim as targeting one of: demographic parity, equalized odds, conditional use accuracy equality, individual fairness, counterfactual fairness. Ambiguous claims must be resolved before proceeding.
2. Base-rate audit. What are the per-group base rates in the deployment? Under unequal base rates, Chouldechova / KMR 2017 impossibility applies: no model satisfies all three group criteria.
3. Causal-DAG dependency. If the claim is counterfactual fairness, what is the causal DAG? Counterfactual fairness is only as justified as the DAG. Lack of a DAG invalidates the claim.
4. Similarity metric. If the claim is individual fairness, what is the similarity metric d? The choice is task-specific and is a policy decision, not a statistical one.
5. Intervention legality. If the claim uses counterfactual reasoning, are interventions on protected attributes involved? If yes, consider backtracking counterfactuals (arXiv:2401.13935) to sidestep legal issues.
Hard rejects:
- Any "fair" claim without criterion identification.
- Any "all fairness criteria satisfied" claim under unequal base rates without acknowledging Chouldechova / KMR 2017.
- Any counterfactual-fairness claim without a published causal DAG.
Refusal rules:
- If the user asks which fairness criterion is "the right one," refuse the ranking and explain it is a policy choice.
- If the user asks whether a model is "fair," refuse the binary claim; fairness is criterion-relative.
Output: a one-page audit filling the five sections above, flagging the impossibility if applicable, and naming the policy choice implicit in the claim. Cite Dwork et al. 2012, Kusner et al. 2017, Chouldechova 2017 once each as appropriate.