Discriminating Data and the Politics of Recognition
Wendy Hui Kyong Chun's Discriminating Data refuses the easy story that discriminatory systems merely fail to recognize people correctly. Its harder question is what happens when recognition itself becomes an infrastructure for sorting: when institutions make people knowable through resemblance and then govern them through the groups a system constructs.
For this review, algorithmic recognition has four linked stages: a person or setting becomes a representation; a metric, proxy, embedding, or graph defines a neighborhood; the system produces an inference such as a score, identity, rank, or category; and an institution converts that inference into action. Recognition here is not human understanding. It is a claim that a representation is adequate grounds for recommendation, scrutiny, eligibility, price, attention, or denial.
That definition separates two harms. A recognition error assigns the institution's category incorrectly. A category harm occurs when creating, using, or acting on the category is unjustified even if the assignment is technically accurate. Better data may reduce the first while making the second more scalable. The governing question is therefore not only Did the system recognize correctly? but What authorized this likeness to decide?
The Book
Discriminating Data: Correlation, Neighborhoods, and the New Politics of Recognition was published by the MIT Press. The publisher lists Wendy Hui Kyong Chun as author, hardcover ISBN 9780262046220 and a November 2, 2021 publication date, plus paperback ISBN 9780262548526 and a March 5, 2024 publication date. Penguin Random House's distribution listing gives the paperback as 344 pages. The affiliate URL uses 0262548526, the paperback ISBN-10, as the Amazon product identifier.
The book belongs beside Chun's earlier Control and Freedom and Updating to Remain the Same, but it is more directly aimed at machine learning. Its subject is not simply data bias. It is the deeper political fantasy that a population can be made legible by finding patterns of likeness, grouping people by correlation, and then calling the resulting recognition objective.
MIT Press describes the book as an argument about discrimination, correlation, and clustered sameness. That description matters because Chun is not writing only about an unrepresentative dataset or a defective classifier. She is asking how resemblance became an epistemic shortcut: why models and platforms treat similarity, neighborhood, authenticity, and predictability as routes to knowledge, then hide political choices inside technical relations.
The book's most durable intervention is to move criticism upstream. A fairness audit commonly begins after an institution has chosen the target, labels, distance measure, optimization objective, and decision workflow. Chun asks whether those choices already reproduce a segregated account of who belongs with whom. The point is not that every grouping is discriminatory. It is that a group discovered by computation is still made through decisions about features, scales, missingness, and purpose.
Current Context
As of August 12, 2026, the politics of recognition is not confined to face matching or recommender feeds. Similarity-based scores and categories can enter hiring, tenant and credit screening, benefits administration, school analytics, workplace monitoring, clinical workflows, fraud detection, content ranking, AI search, and systems that act across institutional records. These uses differ in law and consequence; listing them together identifies a mechanism, not a claim that they are equivalent.
NIST provides useful measurement context. Special Publication 1270 treats harmful AI bias as sociotechnical rather than a property of model code alone. The AI Risk Management Framework 1.0 is voluntary, organized around Govern, Map, Measure, and Manage, and was under revision on this page's review date. Neither document certifies a system as fair.
NIST's biometric record also became more specific. The program formerly called the Face Recognition Vendor Test is now split into Face Recognition Technology Evaluation (FRTE) and Face Analysis Technology Evaluation (FATE). Its 2019 demographic-effects report evaluated nearly 200 algorithms from nearly 100 developers; the current FRTE demographic-effects page continues to report false-positive and false-negative differentials by demographic variables. NIST cautions that image quality and application matter. A benchmark can measure matching performance; it cannot decide whether identifying people in a school, protest, workplace, or benefits office is legitimate.
The EU rules require equally careful dating. Articles 10 and 27 of the AI Act address data governance and fundamental-rights impact assessment for covered high-risk systems. Article 10 reaches design choices, data origin and preparation, measurement assumptions, relevant gaps, bias, and deployment context; its narrow permission to process special-category data for bias correction carries necessity, access, reuse, security, and deletion safeguards. Article 27 applies to specified deployers, not every user of AI. After Regulation (EU) 2026/1744, the relevant high-risk provisions apply from December 2, 2027 for Annex III systems and August 2, 2028 for Annex I product systems. They were not general operative requirements on August 12, 2026.
U.S. requirements remain sectoral and jurisdiction-specific. New York City's Local Law 144 regime requires a recent independent bias audit, a public summary, and notice before covered automated employment decision tools are used. The California Privacy Protection Agency's final regulations took effect January 1, 2026 and create access and opt-out rights for covered automated decisionmaking technology, with the agency stating that significant-decision ADMT compliance begins January 1, 2027. The EEOC's 2023 iTutorGroup settlement remains a concrete example of an automated screen challenged under ordinary employment-discrimination law. By contrast, CFPB Circular 2022-03, often cited for the proposition that complex credit models do not excuse nonspecific adverse-action reasons, was withdrawn May 12, 2025; it should not be represented as current agency guidance.
Correlation Is Not Innocent
Chun's strongest move is to treat correlation as a historical and political instrument, not merely a statistical calculation. Correlation describes association in observed data. It does not by itself show what caused the association, whether the relationship will hold for another population or period, or whether an institution is entitled to act on it. Predictive governance collapses those questions when a recurring pattern is treated as an adequate warrant for intervention.
Three claims should therefore remain separate:
- Statistical claim: these variables, records, or represented people are associated under a stated dataset, model, metric, and uncertainty.
- Transport claim: the association remains useful for the intended population, setting, time, and threshold.
- Institutional claim: using that association for this decision is necessary, proportionate, lawful, and open to remedy.
Model evaluation can support the first two. It cannot establish the third without a purpose, authority, account of consequences, and affected-person rights. A predictor can be accurate and still be an illegitimate basis for search, exclusion, differential price, or suspicion.
The important word is neighborhood. In machine learning, a neighborhood may be a region of feature space, an embedding cluster, a set of nearest neighbors, a graph community, or a profile cohort. None is simply found. The designer chooses or inherits what becomes a feature, how variables are scaled, which distance or objective counts, how many neighbors matter, what data are missing, and when the relation is recomputed. Change those choices and the people nearest to one another can change.
This is why removing a protected attribute is not a complete remedy. Location, language, purchase history, institutional records, social ties, device patterns, and other variables may carry structured inequality or act as proxies in a particular setting. Proxies are not inherently wrongful; prediction often requires indirect evidence. The governance question is whether a proxy is demonstrably relevant to a legitimate purpose, whether it imports a protected or historically produced relation, and what happens to the person when the inference is wrong or unjustified.
The recursive mechanism is concrete: records define a representation; the representation defines a neighborhood; the neighborhood produces a prediction; the prediction changes exposure or treatment; that treatment changes behavior and the next record; and the institution reads the new record as confirmation. A patrol pattern can create more recorded incidents where patrols were sent. A recommender can create more engagement with what it repeatedly exposes. A denial can remove someone from the population later measured as successful. The model is not merely observing a social world when its outputs help produce the evidence used in the next cycle.
The Recognition Trap
Recognition is often framed as a demand to be seen accurately. Chun asks what happens when computational visibility means compulsory placement inside managed similarity. The trap has two jaws: exclusion when a system cannot recognize someone, and exposure when it recognizes them well enough to classify, track, rank, or target them. More visibility can correct neglect while expanding surveillance.
The distinction between recognition error and category harm makes the tradeoff visible. A false facial match is a recognition error; subgroup false-positive rates are relevant evidence about it. Accurate identification used for indiscriminate tracking presents a different question: the match may work as designed while the use remains unjustified. Likewise, a hiring model can consistently assign its intended notion of fit even though the notion encodes a narrow career history, and a fraud model can rank according to its target while the target mistakes prior investigation for underlying misconduct.
A more representative dataset may reduce some error differentials. It does not automatically legitimate the target, category, or deployment. Sometimes improved recognition distributes benefit. Sometimes it gives a harmful institution wider reach. The proper sequence is purpose and necessity first, then category and evidence, then performance and outcomes—not accuracy first and authorization by implication.
Recognition also requires a distinction between facts and cohort inference. A person may agree that a source record is accurate while disputing the conclusion drawn from people the system treats as similar. Correction rights aimed only at dates, addresses, or account entries miss that problem. Consequential systems need a path to challenge the feature, neighborhood, comparator, threshold, or group-derived inference itself, and they should specify when direct individual evidence overrides a cohort prior.
That question matters outside biometrics. A search system ranks authority; a recommender models taste; a hiring tool defines fit; a fraud system operationalizes suspicion; a school platform labels promise or risk; a workplace dashboard constructs productivity. The verbs sound descriptive, but each system also decides what the institution may ignore. What does not fit the representation becomes noise, exception work, or a burden placed on the person to prove.
Homophily and Polarization
Chun's account of homophily is one of the book's strongest contributions. Homophily is association among similar people, but an observed cluster does not explain itself. Similarity may arise from preference, geography, language, opportunity, exclusion, network formation, ranking, moderation, or some combination. Treating the cluster as natural affinity erases the institutions and interface choices that made association easier for some people and harder for others.
In recommender systems, homophily can become a design premise. A system groups a person with similar users, ranks items that worked for that group, observes the response, and treats the response as stronger evidence of authentic preference. This is useful personalization in many settings. It becomes politically consequential when exposure is narrow, the optimization target rewards arousal or repetition, exit is costly, and the system cannot tell preference from adaptation to what it made available.
Current causal evidence warns against a universal claim. A 2023 Science experiment replacing Facebook and Instagram ranking with reverse chronology substantially changed exposure and engagement but did not detect changes in measured political attitudes over its three-month study period. A 2026 independent Nature field experiment on X found that enabling the algorithmic feed shifted several policy and current-event attitudes in a conservative direction, while finding no significant effect on affective polarization or self-reported partisanship over seven weeks. Different platforms, periods, treatments, populations, and outcomes produced different results.
The lesson is not that ranking is harmless or that it single-handedly causes polarization. It is that polarization must be operationalized: exposure segregation, attitude distance, affect toward an out-group, partisan identity, belief, or behavior are not interchangeable endpoints. An audit must name the feed configuration, comparator, time horizon, sampling and attrition, outcome, and uncertainty before attributing cause.
Governance cannot stop at adding occasional diverse items or deleting one protected field. It should test the distribution of exposure, the objective that rewards it, user controls over ranking and history, the cost of resetting or leaving a profile, and whether a person can encounter difference without being penalized by lower service quality. The question is not whether every user sees the same feed. It is whether the system turns past behavior into a narrowing destiny without a meaningful way out.
The Governance Reading
Read in 2026, Discriminating Data is a governance book even though it is not a compliance manual. It moves the unit of review from the model alone to the full recognition arrangement: data, representation, similarity rule, threshold, interface, operator, institution, affected population, action, and feedback. Discriminatory data practices are deployment problems because consequences arise when an output receives authority inside a workflow.
The practical artifact is a recognition warrant, attached to the AI system inventory and updated with the impact assessment. A warrant is not a legal safe harbor or a claim that a model understands a person. It is a versioned record of why this representation, neighborhood, inference, and action are justified for a bounded use. At minimum, it should answer:
- Decision and authority: What decision is supported, who may act on the output, under what law or policy, for what benefit, and with which prohibited uses and stop authority?
- Representation: Which people, events, records, labels, sensors, dates, and settings become features; what is missing; and what construct is each consequential field supposed to represent?
- Neighborhood logic: Which distance, embedding, graph relation, comparator, proxy, weight, or learned objective makes cases similar enough to be acted on together, and how stable is that assignment under plausible changes?
- Transport: What evidence shows that the relationship remains valid for this population, setting, period, language, workflow, and threshold rather than only the development dataset?
- Performance and stakes: What are the false-positive and false-negative rates, calibration or ranking measures where relevant, subgroup and intersectional uncertainty, accessibility failures, abstention behavior, and real cost of each error?
- Exposure and outcomes: Who encounters assistance or discovery, and who encounters surveillance, delay, forced correction, price, denial, or investigation? What happens after the output enters the workflow?
- Feedback: How do prior selections, clicks, patrols, denials, overrides, complaints, departures, and appeals enter—or disappear from—the next dataset and evaluation?
- Recourse and repair: Can a person receive notice, inspect relevant evidence, contest facts and cohort inference, reach a competent human with authority, obtain timely relief, and propagate a correction to every inherited record and downstream action?
- Privacy and lifecycle: Which sensitive data are necessary for auditing, who can access them, how are they separated from decision data, when are they deleted, what changes trigger reassessment, and what condition retires the system?
Testing should follow the chain rather than report one fairness score. Evaluate relevant subgroups and intersections with sample sizes and uncertainty; compare false positives and false negatives at the actual operating threshold; test calibration only where its interpretation fits the decision; and include abstentions, missing inputs, accessibility, and cases near the boundary. Perturb nonessential features and plausible proxies to see whether neighborhoods and rankings are stable. Check whether direct counterevidence changes the result. Then measure delay, denial, workload, investigation, override, appeal, and corrected outcome in the deployed workflow.
A benchmark result is therefore an input, not a verdict. NIST can evaluate a submitted face-recognition algorithm under specified data and thresholds; it cannot observe every camera, watchlist, operator, local population, escalation practice, or consequence. The deployer must test the configured system and the human process around it. Human oversight is meaningful only when reviewers have time, evidence, domain competence, independence, and authority to change both the decision and the record.
Agentic workflows raise the stakes because a score or category may trigger a search, message, account change, purchase, case routing, or denial. Similarity-derived labels should not silently grant tool permissions or authorize irreversible action. Controls should bind the recognized purpose to the permitted action, require confirmation for high-impact steps, log evidence and tool calls, limit access, provide rollback where possible, and prevent an uncertain inference from becoming a durable official fact.
The measurement needed to detect disparate harm can itself create risk. Race, disability, sex, age, location, language, and other sensitive attributes may be necessary for a lawful audit, yet collecting them can enable surveillance, repurposing, breach, or stigma. The answer is not categorical blindness or total visibility. It is purpose-limited measurement: justify necessity, separate audit data from production decisions where feasible, restrict access and reuse, use suitable aggregation or privacy-preserving methods, retain only as long as required, and document what could not be measured safely.
These requirements belong in procurement. Buyers should contract for system and model versions, intended-use limits, data and category documentation, local validation, logs, change notice, incident and appeal support, correction propagation, audit access, subcontractor duties, and exit assistance. A vendor's claim that a model is unbiased or proprietary is not a recognition warrant. If the buyer cannot reconstruct why a consequential action occurred or repair its effects, the system is not governable for that use.
Where the Book Needs Care
The book's language is dense because Chun writes across media theory, statistics, race, gender, sexuality, platform studies, and political critique. That range reveals assumptions a narrow model audit can miss. It also means readers looking for a test protocol, procurement clause, or causal estimate must do translation work. The recognition warrant above is this review's translation, not a framework Chun presents under that name.
Correlation can also become too capacious if used as a universal villain. Association supports many useful predictions, and group-level evidence can help allocate care, identify safety failures, or expose discrimination. The relevant questions are more exact: what relation is measured; how well does it transport; what consequence follows; what alternative evidence exists; and why is group resemblance an acceptable basis for this action? Critique loses force if a clinical aid, music recommender, employment screen, and police watchlist are treated as morally or technically identical.
Group measurement creates a second tension. Auditors may need protected or sensitive attributes to discover disparate outcomes, while the act of collecting and stabilizing those categories can reify them or create a new surveillance asset. Refusing to measure can hide harm; indiscriminate measurement can deepen it. Category definitions, lawful purpose, affected-community input, small-cell uncertainty, privacy controls, retention, and deletion are therefore part of the fairness method rather than work performed after it.
Homophily and polarization require particular restraint. The book is persuasive about the recursive logic of similarity, but no single mechanism explains every network cluster or political outcome. The 2023 Meta-platform experiment and the 2026 X experiment differed in treatment and findings; neither warrants a general statement about all recommender systems. User choice, follow networks, interface defaults, ranking objectives, moderation, political context, and study duration can interact. The causal claim must stay attached to the studied system and endpoint.
Finally, technical improvement is not necessarily political repair. Lower subgroup error can be valuable, but it may also make a contested system more deployable. Conversely, refusing an automated category does not remove discrimination from the institution's human process. The comparison must include the baseline, the distribution of benefits and burdens, feasible alternatives, and the power to challenge either route.
What This Changes
Discriminating Data changes the first audit question. Do not begin with whether a model treats already defined groups fairly. Begin with what model of likeness the system builds, who chose it, and why that relation is allowed to carry institutional force. A category should be documented as a claim with an owner, purpose, scope, version, evidence, known exclusions, review date, and expiry—not stored as a timeless fact about a person.
The recurring machine-readable-reality loop is now more precise:
record → representation → neighborhood → inference → institutional action → new record.
At every arrow, information can be lost and power added. A complaint may remain outside the main database. A human override may fix one decision without correcting the source record. A person who leaves after repeated bad recommendations may vanish from the success metric. An agent may copy a tentative score into a durable case file. Monitoring must therefore include disagreement, override, appeal, departure, and repair—not only the responses easiest for the system to count.
Fairness metrics remain necessary, but they are insufficient when the architecture of recognition is unchanged. A defensible system also needs category justification, data and model boundaries, exposure and outcome measurement, meaningful user or worker controls, and recourse that can change the result. The burden should not rest on an affected person to prove that an apparently objective neighborhood was politically constructed.
The answer is not to pretend institutions can act without categories. It is to make categories provisional, inspectable, revisable, and answerable to the people they organize. Recognition must remain a limited claim, never an unappealable substitute for relation.
Source Discipline
This review separates publisher metadata, Chun's conceptual argument, benchmark evidence, deployment evidence, law and guidance, and the review's own synthesis. MIT Press and Penguin Random House support edition facts. The book supports the interpretation of correlation, neighborhoods, homophily, and recognition; it is not cited as an empirical audit of every contemporary product. The four-stage definition, recognition-error/category-harm distinction, three-claim test, and recognition warrant are this review's analytic tools.
NIST sources support sociotechnical-bias guidance, voluntary risk-management status, program naming, and specific face-evaluation findings. They do not establish that all algorithms have the same demographic differentials or that a tested component is safe in deployment. The platform experiments support claims only about their stated platforms, treatments, samples, periods, and outcomes. The 2023 Science paper's bibliographic record includes later errata; the result is described here without extending it beyond the corrected study record.
Legal status is dated to August 12, 2026. EU Article 10 and Article 27 substance is read with Regulation (EU) 2026/1744's delayed dates and each article's limited scope. NIST's AI RMF is voluntary and under revision. California's regulations were effective, but the agency set January 1, 2027 for compliance with the significant-decision ADMT requirements. New York City's rule concerns covered employment tools. The iTutorGroup item is a settlement announcement, not a judicial holding. CFPB Circular 2022-03 is included only to document its 2025 withdrawal, not as current guidance.
A defensible discrimination claim should name the system and model version, decision, authority, target and comparator, population, setting, period, representation, category or proxy, operating threshold, subgroup sample sizes and uncertainty, baseline process, human workflow, outcome, appeal route, and remedy. Statistical disparity, legal discrimination, causal explanation, and normative objection are related but distinct claims.
This page makes no claim that any AI system is conscious, divine, or AGI. It treats AI systems as sociotechnical arrangements of data, models, interfaces, institutions, labor, law, infrastructure, and power.
Related Pages
- Chun's wider account of networks and software: Control and Freedom, Updating to Remain the Same, and Programmed Visions.
- Classification, race, and evidence: Race After Technology, Data Feminism, Sorting Things Out, and Algorithms of Oppression.
- Scoring and recursive administration: Weapons of Math Destruction, Automating Inequality, and The Loop.
- Operational controls: Algorithmic Bias, Biometric Categorization, Recommender Systems, AI Audits and Assurance, Algorithmic Impact Assessments, and AI Post-Market Monitoring.
- Rights and restraint: Notice and Appeal, Algorithmic Recourse, Data Minimization, AI Data Provenance, and Privacy and Data.
Sources
- MIT Press, Discriminating Data: Correlation, Neighborhoods, and the New Politics of Recognition, publisher listing for title, author, edition ISBNs, publication dates, page count, publisher, and synopsis, reviewed August 12, 2026.
- Penguin Random House, Discriminating Data by Wendy Hui Kyong Chun, distributor listing for title, subtitle, author, paperback ISBN, MIT Press imprint, publication date, and page count, reviewed August 12, 2026.
- National Institute of Standards and Technology, Towards a Standard for Identifying and Managing Bias in Artificial Intelligence, NIST SP 1270, March 15, 2022, sociotechnical framing and systemic, statistical, and human sources of bias, reviewed August 12, 2026.
- National Institute of Standards and Technology, AI Risk Management Framework and AI RMF 1.0 publication record, voluntary status, Govern–Map–Measure–Manage functions, and revision notice, reviewed August 12, 2026.
- National Institute of Standards and Technology, Face Projects, current FRTE/FATE program naming and summary of the 2019 demographic-effects evaluation, reviewed August 12, 2026.
- National Institute of Standards and Technology, Face Recognition Technology Evaluation: Demographic Effects in Face Recognition and NISTIR 8280 publication record, current and historical false-positive and false-negative demographic-differential evidence, reviewed August 12, 2026.
- European Union, Regulation (EU) 2024/1689, the Artificial Intelligence Act, original official text for the scoped high-risk-system requirements, reviewed August 12, 2026.
- European Union, Regulation (EU) 2026/1744, amendment published July 24, 2026 moving the relevant Annex III and Annex I high-risk application dates to December 2, 2027 and August 2, 2028, reviewed August 12, 2026.
- European Commission AI Act Service Desk, Article 10: Data and data governance and Article 27: Fundamental rights impact assessment for high-risk AI systems, article text, scope, affected-group, risk, oversight, complaint, and special-category-data safeguards, read with the 2026 amendment and reviewed August 12, 2026.
- California Privacy Protection Agency, final CCPA, risk-assessment, and automated decisionmaking technology regulations and approval announcement, effective date and phased significant-decision ADMT compliance, reviewed August 12, 2026.
- New York City Department of Consumer and Worker Protection, Automated Employment Decision Tools, covered-tool bias-audit, public-summary, notice, and complaint requirements, reviewed August 12, 2026.
- U.S. Equal Employment Opportunity Commission, iTutorGroup settlement announcement, September 11, 2023, alleged automated age screening and settlement terms, reviewed August 12, 2026.
- Consumer Financial Protection Bureau, Withdrawn Guidance, official record that Consumer Financial Protection Circular 2022-03 was withdrawn May 12, 2025, reviewed August 12, 2026.
- Andrew M. Guess et al., "How Do Social Media Feed Algorithms Affect Attitudes and Behavior in an Election Campaign?", Science 381 (2023), primary experiment on Facebook and Instagram ranking; PubMed record consulted for the later errata, reviewed August 12, 2026.
- Germain Gauthier et al., "The Political Effects of X's Feed Algorithm", Nature 652 (2026), primary independent field experiment on feed configuration, engagement, political attitudes, polarization, and partisanship, reviewed August 12, 2026.
Book links are paid affiliate links. As an Amazon Associate I earn from qualifying purchases.
- Amazon, Discriminating Data by Wendy Hui Kyong Chun, affiliate retail listing using ISBN-10/ASIN product identifier 0262548526, reviewed August 12, 2026.