Blog · Review Essay · Modified August 12, 2026 · Last reviewed August 12, 2026

What Computers Still Can't Do and the Background of Intelligence

Hubert Dreyfus's What Computers Still Can't Do: A Critique of Artificial Reason should not be read as a scorecard of tasks that machines will never perform. Its durable challenge is sharper: intelligence is not only producing an acceptable result but selecting what matters within an embodied, social, and consequential practice. Modern machine learning defeats any simple claim that broad performance requires hand-written rules; it does not establish that a fluent output carries the practical judgment of the role in which it is used.

Background intelligence, in this review, means the learned capacity to notice salient features, revise the frame when the situation changes, and act under the norms and stakes of a practice. Context is the information made available to a system; background is what makes some of that information relevant; judgment is the accountable decision about what to do. Collapsing those three turns a larger prompt into a theory of understanding and a plausible answer into permission to act.

Three claims must therefore remain separate: success on a bounded task, fitness within a changing practice, and legitimate authority to impose consequences. The institutional risk is a feedback loop, not a mystical deficit. Once generated text becomes a record, score, workflow state, or policy input, omissions can harden into the evidence used for the next decision. The practical question is where evidence, local correction, refusal, and responsibility remain after work has been made machine-readable.

The Book

What Computers Still Can't Do was published by MIT Press in 1992 as a revised version of Dreyfus's earlier What Computers Can't Do. MIT Press lists the paperback as 408 pages, published October 30, 1992, with ISBN 9780262540674. Internet Archive's catalog record identifies the book as a 1992 MIT Press title on artificial intelligence, revised from the 1979 edition, with bibliographical references and an index.

The book sits inside a longer argument. RAND's record for Dreyfus's 1965 paper Alchemy and Artificial Intelligence describes it as an examination of the difficulty of simulating cognitive processes on computers, focused on forms of human information processing that resist translation into digital computer language. Berkeley's obituary for Dreyfus describes him as a scholar of Heidegger and European philosophy who taught at UC Berkeley for nearly fifty years and challenged expectations for artificial intelligence as early as the 1960s.

The book's architecture is more precise than its title. Part II examines four assumptions behind what Dreyfus called persistent optimism: a biological assumption about the brain as a digital information processor, a psychological assumption about thought as formal operations on representations, an epistemological assumption that intelligent activity can be formalized, and an ontological assumption that the relevant world can be decomposed into context-free facts. Part III turns to the body, the situation, and human needs. The argument is therefore not one prediction but a linked challenge to a model of brain, mind, knowledge, and world.

Those claims should not be treated as equally settled. Contemporary neural systems do not depend on a programmer writing every rule, and useful computation need not reproduce human cognition. Dreyfus is strongest where he shows that a formal description has already selected a world: objects, features, goals, normal cases, and success criteria. The missing step is not always another fact. It is the practical ability to determine which fact matters now.

Artificial Reason

The title misleads when read as a capability scoreboard. Systems can perform tasks that were once offered as signs of intelligence, and learned representations have produced broad competence without the explicit symbolic rules Dreyfus targeted. That is evidence against a static list of impossibilities. It is not evidence that benchmark success, verbal fluency, human-like cognition, and institutional fitness are the same claim.

Dreyfus's stronger target is the relevance problem: a finite actor must select what matters without first searching every possible fact, interpretation, and rule. A representation problem asks whether the needed information can be encoded. A relevance problem asks whether the right feature becomes salient in this situation. A responsibility problem asks who is entitled to turn that selection into action. An AI system can improve on the first, perform impressively on the second, and still be assigned too much authority on the third.

This gives the review a sharper vocabulary. Task success is performance on a specified input, output, and metric. Situated fitness is reliability across the exceptions, handoffs, conflicts, and changing conditions of an actual practice. Decision legitimacy concerns mandate, reasons, due process, and who bears the consequence. A model result can support the first claim without establishing either of the others; no single benchmark should silently stand in for all three.

Consider a correct summary that omits the one exception a benefits officer must notice, or a medically accurate answer offered to a patient for whom the normal recommendation is contraindicated. More information may repair omission; it may not repair misframing or unauthorized use. The quality of the output and the legitimacy of the action must be evaluated separately.

That is why the book belongs beside Computer Power and Human Reason, Understanding Computers and Cognition, Human-Machine Reconfigurations, and The Experience Machine. Together they ask what is lost when an institution treats a representation as the world, a plan as the action, or an output as the judgment.

The Background Problem

A person does not enter a task as a blank problem solver facing a database of facts. A person arrives with bodily habits, learned equipment, purposes, emotions, memories, institutional roles, and expectations about what normally matters. Background intelligence is this practical orientation: not a hidden encyclopedia, but a capacity formed through repeated participation, correction, and consequence.

The background is distributed across people and arrangements. It can reside in a clinician's comparison with a patient's baseline, a classroom's rhythm, a shop floor's sound, a courtroom's procedure, a codebase's local conventions, a family's history, a deadline, or a worker's knowledge of how the official process actually fails. A model may receive traces of these things. The input schema, retrieval index, camera angle, sensor package, and permission boundary decide which traces can count.

This yields a useful three-way distinction. Model context consists of supplied instructions, retrieved documents, examples, logs, images, tool descriptions, and memory. Practical background supplies the norms and salience by which those items are interpreted. Institutional authority determines what consequences may follow. A larger context window can improve coverage; it cannot by itself decide whose account is authoritative, which exception is just, or whether the system should be allowed to act.

The background is not ineffable or automatically good. It can be investigated through observation, interviews, incident records, disagreement, and tests in use. It can also carry prejudice, professional blind spots, coercive custom, and unequal power. The lesson is not to protect tacit practice from scrutiny. It is to document enough of the setting to expose both what formalization omits and what local practice should no longer be allowed to hide.

Nor does the demand for background license total collection. A system could gain more signals by retaining private conversations, location histories, health details, or workplace behavior and still become less trustworthy because the information is excessive, stale, repurposed, or available to the wrong actor. The relevant design question is the minimum evidence needed for the declared role. Data minimization, purpose limits, retention rules, and access controls are therefore part of epistemic quality as well as privacy: indiscriminate memory can make both intrusion and false salience more likely.

The governance question follows. If a model drafts a medical note, which observations came from the clinician and which were inferred? If an answer engine compresses a dispute, where are unresolved sources and minority interpretations? If an agent chooses a next step, whose authority does the credential represent? If a score reorganizes public service, can the affected person introduce evidence that the form never requested?

The danger is not that a machine lacks a soul. It is that an institution may act as if the background work has already been done. A fluent response can conceal omitted evidence; a dashboard can turn a contested category into an apparent fact; an efficient workflow can push exception handling onto the person with the least power, time, or access to contest it.

The Background Test

Dreyfus becomes useful as a deployment test when five failures are separated:

Test each failure at the level of the whole workflow. Hold the nominal task constant while varying local conditions, source order, missing records, conflicting duties, rare but consequential exceptions, user roles, changed policies, and adversarial instructions. Then test handoffs: draft to reviewer, reviewer to decision maker, decision to record, and record to later retrieval. Report task performance separately from unsupported-action rate, reviewer detection, reliable escalation, correction latency, and whether amendments propagate downstream.

The practical artifact is a background register. Its observation section states what the system sees, what it does not see, how sources were selected, and which conditions invalidate the output. Its relevance and norms section names contested categories, known exceptions, governing rules, affected viewpoints, and who resolves conflicts. Its authority section states permitted uses, prohibited inferences, irreversible actions, approval thresholds, and the accountable decision owner. Its repair section states what is logged, who can correct the record, how an affected person appeals, whether corrections reach derived records, what triggers incident review, and when the system must be suspended or retired.

A deployment passes only when missing background has somewhere effective to go: visible uncertainty, source provenance, local amendment, meaningful override, incident learning, and recourse after the action. The register does not convert every skill or norm into a rule. It makes the boundary of the formal system inspectable and assigns responsibility for what remains outside it.

After Symbolic AI

The 1992 edition already addressed connectionism and neural networks in a new introduction. Dreyfus later revisited embodied and Heideggerian approaches in a 2007 article in Artificial Intelligence. A 2024 AAAI Spring Symposium paper returned to his claims in the context of deep learning, large language models, and hybrid AI. That history rules out a simple contest between a 1992 prediction and a 2026 product demo.

Learned systems weaken the claim that broad performance must be built from explicit, hand-coded rules. They can absorb statistical regularities from language, images, demonstrations, feedback, and interaction. Those regularities can function as useful proxies for parts of a practice. Calling all such performance "mere symbol manipulation" would miss the change in method and capability.

But a proxy for background is not the same as an institution proving fitness for a role. Training data can contain how people usually speak without preserving why a particular source is authoritative. Retrieval can supply a policy without identifying the exceptional case in which it should not control. Tool use can connect a response to the world while magnifying the cost of a mistaken frame. Human feedback can shape acceptable behavior while concealing whose judgments were treated as normal.

The current test is therefore empirical and operational. Does the system remain reliable when the situation departs from its evaluation distribution? Can it identify missing evidence, conflicting duties, and limits on its authority? Do users detect plausible but unsupported claims? Does the workflow preserve correction when an output is factually fluent but pragmatically unsafe? None of those questions requires a claim that the system is conscious, unconscious, human-like, or destined for general intelligence.

Current Context

As of August 12, 2026, the relevant object is often not a stand-alone model but a system assembled from prompts, retrieval sources, memory, tools, identities, policy layers, human review, and product defaults. NIST's February 2026 draft concept paper describes agentic architectures as systems that obtain additional context and may take action; it focuses on identification, authentication, authorization, delegation, logging, provenance, and prompt-injection controls. The draft is a project proposal, not a completed standard.

NIST's AI Risk Management Framework 1.0 remains voluntary and is being revised. Its 2024 Generative AI Profile already locates risk at model, system, use-case, and ecosystem levels and names confabulation, human-AI configuration, information integrity, and component integration among the relevant risks. That framing supports the book's strongest lesson: performance cannot be assessed apart from the human and institutional arrangement that turns an output into a consequence.

NIST launched an AI Agent Standards Initiative in February 2026 around interoperability, security, and identity. It is an initiative for planned guidance, standards work, community protocols, and research—not a certification or a finished agent standard. OWASP's agentic Top 10, released as a community security taxonomy rather than law or a conformity standard, identifies goal hijacking, tool misuse, identity and privilege abuse, memory and context poisoning, insecure inter-agent communication, cascading failures, and exploitation of human trust. These are not evidence for a philosophical impossibility. They show that connecting models to memory, credentials, and tools turns mistakes about relevance into operational security failures.

"Context engineering" is therefore an engineering practice, not a synonym for understanding. A larger context window, richer retrieval index, or persistent memory store can improve task performance. Each also creates a selection, privacy, and security boundary: who chose the sources, which version was retrieved, what personal data persists, what instructions can override the task, what authority the credential carries, and what record proves that the boundary held? Supplying more background can reduce one omission while creating a new attack surface or an unjustified inference.

The Institutional Reading

Read in 2026, the book is less a metaphysical verdict on machine minds than an audit of delegated judgment. The practical question is not "Can this system think?" It is "What did the workflow remove from the task, who repairs that loss, and who may challenge the resulting action?"

In a private notebook, missing background may be a manageable inconvenience. In a school, hospital, court, benefits office, newsroom, workplace, police department, or military system, it becomes institutional risk. A summary can become the record; the record can become the next model's retrieval source; repeated outputs can become a metric or policy default. The omission then feeds back as evidence. This is why model memory is an attack surface and why a benchmark can become a curriculum: systems do not merely describe a world; their records and measures help produce the next version of it.

The repair labor is easy to hide. A worker finds the missing document, recognizes the local exception, rewrites the generated note, explains the decision to the affected person, or absorbs the anger when no appeal route exists. If evaluation counts only speed or output acceptance, that labor appears as system performance. Measure review time, correction frequency, reopened cases, escalation load, and whose work supplies the missing context. Otherwise the institution has not removed ambiguity; it has reassigned ambiguity to people whose corrections may never enter the official record.

This is where Dreyfus connects to Tools for Thought. A good cognitive tool expands contact with evidence: more sources, alternatives, inspectable uncertainty, and power to revise. A bad one narrows the field while making the narrowed version feel complete. The difference is not whether AI appears in the tool. It is whether the user can reopen the frame.

The same test applies to agents. An agent can select tools, read private context, write to systems, trigger workflows, and act under delegated credentials. If its frame is wrong, an answer error becomes an authority error. Safe assistance therefore requires both epistemic controls around sources and operational controls around permissions, approvals, logs, rollback, and stop conditions.

Governance and Safety

The unit of governance is the decision system: model, data, retrieval, interface, tools, people, policy, vendor, and appeal process. The first task is to locate the missing background before output becomes action: domain assumptions, local exceptions, source limits, uncertainty, affected people, human authority, and the cost of being wrong. A model evaluation cannot establish that the surrounding institution is safe.

Public frameworks support that system-level view, but their force and maturity differ. NIST AI RMF 1.0 is voluntary and currently under revision; its govern, map, measure, and manage functions are a lifecycle framework, not certification. U.S. OMB Memorandum M-25-21 applies to federal agencies. For agency uses whose output is a principal basis for decisions with significant effects on rights or safety, it requires practices including pre-deployment testing, impact assessment, ongoing monitoring, trained human oversight, remedies or appeals, feedback, and discontinuation of non-compliant uses.

The EU position changed after this article's original publication. As of August 12, 2026, Article 50 transparency obligations apply from August 2, subject to their scope, exceptions, and a limited transition for some pre-existing generative systems. Regulation (EU) 2026/1744, published in July and now in force, delayed Sections 1–3 of the AI Act's high-risk chapter to December 2, 2027 for Annex III systems and August 2, 2028 for systems tied to regulated products under Annex I. Consequently, high-risk requirements such as Article 12 record-keeping, Article 14 human oversight, and the Article 27 impact assessment apply only to covered actors and on the amended timetable; they should not be described as universal duties already governing every AI deployment.

The background register should connect to the AI system inventory, system card, audit trail, and recourse process. Record the decision owner, intended role, source boundaries, groups and environments tested, known exceptions, uncertainty shown to users, authority limits, approval points, override events, incidents, complaints, and retirement trigger. Treat disagreement and overrides as evidence about the system, not as reviewer failure to be optimized away.

For consequential uses, preserve a decision trace from source and observation through selection, recommendation, approval, action, and later correction. The trace should distinguish supplied evidence from inference, identify the policy and credential that authorized each action, and show whether an appeal changed every derived record. Keep audit evidence proportionate: logging everything indefinitely can reproduce the very surveillance and stale-context risks the register is meant to expose. Link retention and access decisions to the site's privacy and data commitments.

For AI agents, add distinct non-human identities, least-privilege credentials, per-tool scopes, approval gates before irreversible actions, prompt-injection and poisoned-context testing, tamper-evident action traces, rollback, revocation, rate limits, and a tested stop path. Every delegated action should be attributable to a system identity, a human or institutional principal, and the policy that authorized it.

This is why human oversight must be designed as authority, not presence. A reviewer needs time, domain competence, source access, independence, power to pause or override, and a channel that changes the record. A person who can only confirm a machine-framed choice is not restoring background judgment; the interface is borrowing their signature. Oversight also needs post-deployment monitoring and meaningful notice and appeal, so recurring corrections can change the system rather than disappear case by case.

Where the Book Needs Friction

What Computers Still Can't Do is strongest as a critique of assumptions and weakest as a timeless boundary around capability. Evidence from a particular architecture can defeat a technical forecast; it cannot establish that every future approach will fail. Conversely, a new capability does not erase the philosophical problem merely because the earlier example has been automated. The title encourages both errors.

The book also risks treating "the" human background as more unified and benign than it is. Backgrounds conflict. A worker, manager, patient, insurer, resident, regulator, and vendor may perceive different facts as salient because they bear different risks and powers. Tacit expertise can protect care and safety; it can also shield bias, gatekeeping, and custom from challenge. Any contemporary use of Dreyfus needs a politics of whose background enters the system and whose testimony can revise it.

Formalization is not the enemy. Forms, rules, models, checklists, and records can constrain arbitrary power, coordinate large institutions, and make decisions appealable. The problem begins when a partial representation is treated as exhaustive or when the cost of its omissions is shifted onto people who cannot correct it. The right response is accountable formalization: named scope, visible uncertainty, documented exceptions, and revision when practice changes.

Dreyfus does not supply a governance program. Readers still need work on procurement, discrimination, labor, surveillance, evaluation, security, law, and platform power. His contribution is a prior question those fields cannot skip: what practical world has been selected for the system, and who remains able to reopen that selection?

What This Changes

The book changes the evaluation question. Do not ask only whether a model produces the expected answer. Ask what the institution needs the world to look like for that answer to count as correct, and what happens when the world refuses the frame.

The feedback test comes last. Follow an output after it leaves the screen. If it becomes a note, score, ranking, retrieval source, training example, performance metric, or policy premise, preserve the original evidence, the human amendments, and the dissent. Otherwise yesterday's compressed judgment becomes tomorrow's background, and the system will appear consistent because it has erased the conditions under which people disagreed.

Dreyfus's critique survives not as a ban on machine intelligence but as a discipline for bounded authority. Use systems where their role is explicit and their failures are observable, reversible, and contestable. Where decisions affect rights, care, work, money, bodies, reputation, or public memory, require a named decision owner and a repair path that remains outside the model's permission.

The enduring issue is the world behind the output: bodies, tools, histories, institutions, and consequences. A responsible system does not need to contain that entire world. It must not conceal the boundary, and the institution using it must remain answerable for what the boundary excludes.

Source Discipline

This review separates publication facts, Dreyfus's argument, later scholarship, current technical guidance, and law. MIT Press, the book catalog, RAND, Berkeley, the 2007 article, and the AAAI paper establish bibliographic history and the changing scholarly debate. NIST and NCCoE establish voluntary federal guidance and draft technical work; OWASP supplies a community security taxonomy. OMB supplies binding executive-branch direction within its stated federal scope. EUR-Lex and the European Commission establish current EU dates and obligations. The background register and five-failure deployment test are this review's synthesis, not claims that Dreyfus or any cited institution endorses them.

Source labels matter. A draft concept paper is not a standard. A voluntary framework is not certification. A security taxonomy is not law. A regulation can impose different duties on providers and deployers, and only within defined scope and application dates. A benchmark result does not establish situated understanding; a failure example does not prove impossibility; a legal requirement does not prove a deployment safe.

Capability claims should identify the model or service version, evaluation date, task, metric, population, and conditions. Deployment claims also need data and retrieval sources, available tools, permission boundary, human role, failure modes, downstream use, appeal path, and accountable owner. Legal claims need jurisdiction, covered actor, provision, effective date, and exceptions. Keeping those claim types separate prevents a product result from becoming a theory of mind, or a policy recommendation from being mistaken for current law. This article makes no claim that an AI system is conscious, divine, or artificial general intelligence.

Sources

Book links are paid affiliate links. As an Amazon Associate I earn from qualifying purchases.


Return to Blog · Return to Books