Invisible Women and the Data Gap Under AI
Caroline Criado Perez's Invisible Women is not narrowly an AI book. It is an account of how institutions mistake evidence about men for evidence about people, then convert that mistake into medicine, transport, work, products, public space, and policy. Its AI relevance begins before model training: a system cannot recover a population that the institution never observed, described with the wrong construct, or averaged out of view.
For this review, the evidence population is the people, bodies, events, and settings the records actually describe; the decision population is everyone a system will classify, assist, rank, route, or govern. A data gap is a decision-relevant mismatch between the two, or a measurement that fails to represent the thing the decision requires. Absence is only one form. Selection, category, proxy, aggregation, transfer, and feedback gaps can all leave a dataset apparently complete.
The governing principle is precise: representation parity is not performance parity, and performance parity is not outcome parity. More records do not fix a bad category; equal error rates do not guarantee equal stakes; a benchmark does not establish fitness in a new setting. Every automated action therefore needs an evidence boundary, an outcome check, and a correction path.
The Book
Invisible Women: Data Bias in a World Designed for Men was published in the United States by Abrams Press on March 12, 2019. Abrams lists Caroline Criado Perez as the author, 448 pages, and ISBN-13 9781419729072 for this hardcover edition. The site's affiliate URL uses the corresponding ISBN-10, 1419729071.
The book's argument is simple and severe: when institutions collect, classify, and design around incomplete evidence, a missing population does not remain outside the spreadsheet. People are required to live inside systems calibrated for someone else. Criado Perez moves through work, unpaid care, medicine, transport, product design, disaster response, public space, and technology. Across those settings, the male case is repeatedly allowed to stand for the general case while women become deviations requiring accommodation.
The distinctive method is accumulation. Instead of presenting one model of discrimination and testing it in a single domain, the book assembles cases until the default becomes visible as infrastructure. An LSE Review of Books essay usefully groups its central absences around female physiology, unequal care burdens, and violence. That range gives the argument force, but it also creates a limit: cases with different evidence, institutions, and causal pathways cannot all be governed by the same remedy.
The book belongs beside Data Feminism and All Data Are Local because all three reject data as a neutral mirror. Criado Perez adds a particularly practical insight: the population treated as normal at collection time becomes the body, schedule, household, route, symptom, voice, and risk profile around which later systems are optimized. Unfairness can be built before anyone calls the process automated.
Current Context
As of August 12, 2026, one of the book's best-known design concerns has a useful public record. A 2023 U.S. Government Accountability Office review found that the crash-test dummies used by the National Highway Traffic Safety Administration represented a limited range of body sizes, did not reflect some physiological differences between females and males, and lacked lower-leg sensors relevant to known injury differences. GAO did not say that one dummy explained every disparity. It asked NHTSA for a comprehensive plan. GAO now marks that recommendation implemented after NHTSA issued a 2024 plan covering field data, computer simulation, dummy development, and milestones. A data gap can be a safety defect, and correction requires an institutional program rather than a larger spreadsheet alone.
Research policy supplies a second boundary. The National Institutes of Health expects sex as a biological variable to be addressed in the design, analysis, and reporting of NIH-funded vertebrate-animal and human studies, with strong justification for studying only one sex. That is a rule about scientific rigor and generalizability, not permission to treat sex as a universal proxy for gender, identity, anatomy, exposure, behavior, or care roles. The construct must still match the question.
The EU AI Act names data origin, preparation, measurement assumptions, bias, relevant gaps, representativeness, and deployment setting in Article 10. But scope and timing matter. Article 10 sits in the high-risk regime; after Regulation (EU) 2026/1744, the relevant rules apply from December 2, 2027 for Annex III systems and August 2, 2028 for systems tied to Annex I products. They were not a general, operative requirement for every AI dataset on this page's review date.
General-purpose models follow a different track. Article 53's training-content-summary obligation has applied to providers placing covered models on the EU market since August 2, 2025, and Commission enforcement began August 2, 2026; earlier models have a transition until August 2, 2027. The Commission's mandatory template asks for content types, sources, and processing information. That public summary improves source visibility, but it is not a subgroup evaluation, consent audit, or finding that a model is fit for a hiring, health, education, credit, or benefits decision.
U.S. federal procurement is narrower again. OMB Memorandum M-25-22, issued April 3, 2025, directs covered agencies acquiring AI to address fit for purpose, performance tracking, cross-functional review, privacy, government data, documentation, interoperability, and vendor dependence. It is federal acquisition guidance with exclusions and phased applicability, not a general guarantee of demographic validity. Its practical value is contractual: evidence requirements must be set before the buyer loses leverage.
The Default Is a Decision
The book turns default into a decision word. A default body, worker, route, voice, schedule, household, or risk profile appears neutral only after the institution forgets whom it used as the reference class. The default then determines equipment dimensions, survey questions, service hours, labels, benchmark examples, thresholds, and interface paths. It is an active design condition, not an empty starting point.
The error has two layers. A reference-class error occurs when evidence describing one group is generalized to a wider population without adequate validation. A construct error occurs when the recorded variable does not represent the concept the institution needs: employment history stands in for ability, uninterrupted availability for commitment, recorded diagnosis for need, or a binary administrative field for a more complex biological or social question. More rows cannot repair the second error because the wrong thing is being measured more consistently.
Defaults also allocate adaptation work. People outside the assumed body, schedule, name, voice, caregiving pattern, symptom profile, or documentation history must translate themselves, seek exceptions, repeat evidence, or find a human workaround. When the workaround is not recorded as system labor, the dashboard reports a functioning service while affected people absorb its failure. The interface looks universal because the cost of non-fit has been moved off the institution's books.
Automation can hide the choice further. The inherited default appears as a field, embedding, threshold, classifier, benchmark, or synthetic user rather than a manager's explicit judgment. No hostile intention is required. Once the output can allocate attention, rank a case, deny access, or trigger another tool, a partial description becomes authority.
The Gap Is a Pipeline
A data gap is best audited as a pipeline, because different failures require different remedies:
- Coverage and selection gap: relevant people or events are absent, under-sampled, or less likely to leave a record because of access, cost, language, reporting, or institutional exclusion.
- Measurement and construct gap: the instrument, sensor, survey, label, or proxy records the wrong thing, works differently across groups, or treats a contested category as natural.
- Aggregation gap: an overall result conceals subgroup, intersectional, geographic, or high-severity failure; small groups vanish inside a strong average.
- Transfer gap: evidence collected for one purpose, population, place, or period is deployed in another without showing that the relationship still holds.
- Outcome gap: evaluation stops at prediction quality and never measures delay, denial, workload, injury, cost, access, or other consequences after the output enters a workflow.
- Feedback and remedy gap: complaints, overrides, exceptions, and appeals do not update the dataset, threshold, process, or decision record, so the same omission returns as new evidence.
Adding data directly addresses only some coverage gaps. It can worsen a measurement gap by collecting an invalid proxy more comprehensively, worsen a privacy risk by expanding sensitive records, or worsen a feedback loop by giving an unjust decision more historical weight. The first governance question is therefore not How much data? but Evidence for which decision, about whom, produced how?
Averages are especially weak where a group is small or the consequence is severe. A system can score well overall while a rare but foreseeable case carries unacceptable physical, economic, or civil-rights harm. Population-proportional testing may then be inadequate; evaluation needs risk-based cases, subgroup uncertainty, error costs, and enough qualitative investigation to understand why the failure occurs. Disaggregation is a diagnostic, not a ritual.
AI components add handoffs at which the boundary can be lost. Pretraining coverage does not establish decision validity. Retrieval can rank sources that repeat the old default. Synthetic data can reproduce a mistaken construct in cleaner form. Fine-tuning can narrow behavior without fixing the source population. A downstream workflow can convert a tentative output into an official record. Each handoff needs its own claim and evidence.
The AI Reading
Foundation models are often described through scale, but coverage is not validation. A model may contain text about many populations and still lack reliable evidence for a specific clinical, employment, benefits, education, or safety decision. It may reproduce language used about a group without being calibrated for outcomes affecting that group. The model has seen examples is therefore not the same claim as the system was tested for this use.
Fluency makes the boundary easy to miss. A generated summary can remove caveats, merge incompatible reference classes, or present an absent subgroup finding as a general answer. A retrieval system can cite authoritative documents while failing to retrieve the source that describes an exception. An evaluator can then score style or factual overlap while leaving the decision population untested. The answer feels complete because uncertainty was compressed out of the interface.
Agentic workflows turn description into administration. Scheduling, triage, routing, purchasing, screening, and case-management tools act through forms, calendars, identity records, policies, and permissions that already contain defaults. A generic worker may mean no care interruption; a generic household may mean one documentation pattern; a generic patient may mean one symptom baseline. These are hazard scenarios, not claims that every product behaves this way. The control is to test the actual system, version, tool path, population, and consequence.
The feedback loop is the deeper risk. A partial record trains or prompts a system; the system changes access or behavior; the changed behavior produces the next record; and the institution reads that record as confirmation. People who stop using an inaccessible service disappear from usage data. Appeals resolved outside the main database do not change the model's apparent accuracy. Extra work performed by staff or families remains invisible. A gap can therefore manufacture evidence for its own continuation.
Governance and Safety
The practical artifact is a data-gap register maintained beside the AI system inventory and impact assessment. It should be versioned for each consequential use, not copied from a general model card. At minimum, it should record:
- Decision and authority: the decision supported, system and model version, deployer, vendor, accountable owner, legal or policy authority, intended benefit, prohibited uses, and people who can pause or retire the workflow.
- Population boundary: the decision population, evidence population, source settings and dates, inclusion and exclusion mechanisms, sample sizes, missingness, and the groups for whom validity is unknown.
- Construct and category warrant: what each consequential field, label, target, and proxy is supposed to represent; who defined it; alternatives considered; and why the category is necessary for this decision.
- Evaluation: baseline process, overall and relevant subgroup results, intersectional tests where lawful and sufficiently supported, uncertainty and small-cell limits, false-positive and false-negative costs, accessibility, calibration where applicable, and stress cases selected by severity as well as frequency.
- Operational outcomes: access, delay, denial, workload, wages or cost, safety events, overrides, exceptions, complaints, appeal results, downstream corrections, and drift by deployment setting.
- Privacy and retention: lawful basis and necessity for sensitive attributes, access tiers, separation of audit data from decision data where appropriate, aggregation or privacy-preserving methods, retention limits, and deletion or correction procedures.
- Remedy and change: notice, a reason specific enough to contest, a competent human reviewer with authority, repair of inherited records, vendor obligations, remediation deadlines, and stop conditions.
The register changes the status of unknowns. Not measured must not silently become no difference; too few cases to estimate must not become safe; and vendor data unavailable must not become acceptable proprietary detail. In a high-impact system, an unresolved evidence boundary should narrow the use, trigger additional study, preserve a human process, or stop deployment.
NIST supplies useful but limited support. Special Publication 1270 distinguishes systemic, statistical, and human bias and treats bias as sociotechnical. The voluntary AI Risk Management Framework 1.0 organizes work through Govern, Map, Measure, and Manage and was under revision on this page's review date. The Generative AI Profile adds harmful bias and homogenization, privacy, information integrity, and value-chain risks. These are risk-management resources, not certification, law, or proof that a deployment is fair.
Procurement must bind the evidence to the contract. Buyers should require population and setting limits, data and evaluation documentation, change notice, access to performance and incident records, support for explanations and corrections, subcontractor flow-down, monitoring cooperation, and exit assistance. A vendor's overall benchmark or public training summary cannot establish local validity. If the buyer cannot test the relevant population or reconstruct a harmful decision, the system is not auditable for that use.
Human oversight has to include power over the record. Reviewers need time, relevant domain knowledge, access to the evidence, authority to depart from the output, protection from pressure to agree, and a way to correct every downstream file or action that inherited the error. Otherwise the person at the end of the workflow merely absorbs liability for a data boundary they cannot change.
Safety also requires restraint. Sex, gender, race, disability, pregnancy, caregiving, household structure, and other sensitive attributes may be necessary to discover disparate harm, but collecting them can expose people to surveillance, stigma, breach, or repurposing. The answer to invisibility is not total visibility. It is purpose-limited, access-controlled measurement with affected-person notice where required, independent scrutiny, and a defined end to collection.
Where the Book Needs Care
The book's evidentiary abundance is persuasive and sometimes analytically loose. Cases about drug research, snow clearing, unpaid care, protective equipment, public toilets, speech technology, and political representation demonstrate a pattern, but they do not share one causal pathway or one remedy. A catalogue can establish that the default recurs; it cannot by itself tell a hospital, city, employer, or model developer which construct to measure, which intervention caused an outcome, or what error threshold is acceptable.
The tempting remedy is collect more data. Better coverage can expose a problem, but it does not force an institution to change, repair an invalid category, or distribute benefits fairly. More complete records can also become a new surveillance channel for people already scrutinized by medical, workplace, welfare, border, or platform systems. Data quality and institutional power are separate questions.
The terms women, female, sex, and gender must not be treated as interchangeable fields. Some failures concern anatomy or physiology; some concern pregnancy; some concern gendered care roles, violence, labor markets, identity, or administrative treatment; many involve interactions among them. A responsible study names the construct needed for the decision, allows for people who do not fit a binary category, and explains why collection is necessary, proportionate, and safe.
Intersectionality is not solved by an endless table of slices. A system can improve for women on average while still failing disabled women, Black women, trans women, migrant women, older women, pregnant people, low-income caregivers, or speakers of lower-resource languages and dialects. Yet very small cells can produce unstable estimates or privacy risk. Governance needs a combination of quantitative uncertainty, targeted testing, affected-community evidence, and protected reporting rather than a demand that one metric expose every interaction.
Nor should the book be frozen in 2019. The GAO record now includes NHTSA's corrective plan; NIH policy and AI law have developed; model supply chains have changed. Current review means checking whether a documented gap remains, narrowed, shifted, or acquired a new remedy. Repeating an old statistic after the underlying program changes would reproduce the very evidence failure the book teaches readers to notice.
What This Changes
Invisible Women changes the burden of proof. The person who does not fit a system should not have to establish from scratch that the default was partial. The institution choosing the category, dataset, benchmark, threshold, and workflow should show that its evidence fits the people and consequences within scope.
The recursive sequence is concrete: a population is under-observed; the incomplete record defines a default; the system allocates service or risk through that default; people adapt, leave, appeal, or perform extra work; only machine-readable responses enter the next dataset; and the resulting record appears to validate the system. Complaints, overrides, workarounds, and departures are therefore not peripheral support data. They are evidence about model and institutional validity.
The audit questions follow. What decision is being made? Who will live with it? Who is in the evidence? What construct is measured? Which group or setting lacks validation? What does the aggregate hide? What harm follows each kind of error? What sensitive data are truly necessary? Who can inspect, override, correct, compensate, and stop? What happened after deployment, not only on the test set?
The book's lasting force is its refusal to treat exclusion as an accident at the margin. A model does not repair a partial world by optimizing over its records. It can formalize the gap, distribute its costs, and generate new records that make the formalization look normal. Every prediction carries an evidence boundary; every consequential action needs a correction path.
Source Discipline
This review separates publisher metadata, Criado Perez's argument, independent reception, official empirical findings, current policy, law, voluntary guidance, and this page's synthesis. Abrams supports the U.S. edition facts and publisher synopsis. The author page supports Criado Perez's framing, not independent verification of every case. The LSE essay documents critical reception and a useful organization of the book. GAO and NIH support the specific safety and research-policy claims. NIST, OMB, EUR-Lex, and European Commission materials support current governance claims.
The data-gap taxonomy and register are this review's tools, not terms presented as such in the book. The AI reading applies the book's mechanism to models and automated workflows; it does not claim that a 2019 book anticipated every feature of foundation models, retrieval systems, synthetic data, or agents.
Current legal and policy sources are dated to August 12, 2026. Article 10's substance must be read with its delayed high-risk application dates. A general-purpose-model training summary is not a complete provenance or validity audit. NIST's AI RMF is voluntary and under revision. OMB M-25-22 governs covered U.S. federal acquisitions rather than all procurement. NIH's sex-as-a-biological-variable policy should not be cited as a rule for every gender-data question.
A defensible data-gap claim names the population, setting, variable or construct, collection mechanism, source and period, decision, comparator, missingness, uncertainty, harm, and status of any remedy. A fairness claim should say whether it concerns representation, performance, allocation, access, outcome, privacy, or recourse. Without those boundaries, bias becomes a slogan and more data becomes a risky cure-all.
Related Pages
- Power, setting, and documentation: Data Feminism, All Data Are Local, and The Data Sheet Becomes the Supply Chain.
- Systemic classification and feedback: More than a Glitch, Sorting Things Out, and Weapons of Math Destruction.
- Evidence controls: Training Data, AI Data Provenance, Model Cards and System Cards, Algorithmic Impact Assessments, and AI Post-Market Monitoring.
- Rights and restraint: AI in Healthcare, AI in Employment, Human Oversight, Notice and Appeal, Data Minimization, and Privacy and Data.
Sources
- Abrams Books, Invisible Women: Data Bias in a World Designed for Men, U.S. publisher listing for title, author, imprint, 448-page count, publication date, ISBN-13, and synopsis, reviewed August 12, 2026.
- Caroline Criado Perez, Invisible Women author page, author framing and book overview, reviewed August 12, 2026.
- Mariel McKone Leonard, "Exposing the Costs of Uncounting", LSE Review of Books, November 4, 2020, independent reception, thematic organization, and limits of the book's accumulative method, reviewed August 12, 2026.
- U.S. Government Accountability Office, Vehicle Safety: DOT Should Take Additional Actions to Improve the Information Obtained from Crash Test Dummies, GAO-23-105595, March 8, 2023, findings, recommendation, and updated implementation status, reviewed August 12, 2026.
- National Institutes of Health Office of Research on Women's Health, Sex as a Biological Variable, current NIH policy expectation for research design, analysis, and reporting, reviewed August 12, 2026.
- National Institute of Standards and Technology, Towards a Standard for Identifying and Managing Bias in Artificial Intelligence, NIST SP 1270, March 15, 2022, sociotechnical framing and systemic, statistical, and human bias categories, reviewed August 12, 2026.
- National Institute of Standards and Technology, AI Risk Management Framework, voluntary status, Govern–Map–Measure–Manage lifecycle, and revision notice, reviewed August 12, 2026.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, July 26, 2024, harmful-bias, homogenization, privacy, information-integrity, and value-chain risks, reviewed August 12, 2026.
- European Union, Regulation (EU) 2024/1689, the Artificial Intelligence Act, Articles 10 and 53 on high-risk-system data governance and general-purpose-model training-content summaries, read with the 2026 amendment, reviewed August 12, 2026.
- European Union, Regulation (EU) 2026/1744, July 24, 2026 official text amending Article 10 and moving the relevant high-risk application dates to December 2, 2027 and August 2, 2028, reviewed August 12, 2026.
- European Commission, training-content-summary template and explanatory notice and official questions and answers, template scope, application, enforcement, transition, update, and disclosure details, reviewed August 12, 2026.
- White House Office of Management and Budget, M-25-22: Driving Efficient Acquisition of Artificial Intelligence in Government, April 3, 2025, covered federal acquisitions, fit-for-purpose performance, privacy, data, documentation, competition, interoperability, and lifecycle requirements, reviewed August 12, 2026.
Book links are paid affiliate links. As an Amazon Associate I earn from qualifying purchases.
- Amazon, Invisible Women by Caroline Criado Perez, affiliate listing, reviewed August 12, 2026.