The Theory Loop Becomes the Cognitive Scientist
A theory loop is a bounded workflow that links a verbal mechanism, executable model, discriminating experiment, behavior source, evaluation rule, arbitration step, and revision. AutoCog is a consequential demonstration of that workflow in two online decision-making settings, not evidence that software has become a general scientist.
The defensible claim is narrower: an agentic system can search over cognitive models, collect new human data, and produce a candidate that survives a prospective preregistered test. Credibility still depends on separating discovery from confirmation, model fit from explanation, and automated execution from human responsibility.
The Theory Enters the Loop
Akshay K. Jagadish, Younes Strittmatter, Nori Jacoby, George Kachergis, Eric Schulz, Nathaniel Daw, Suyog H. Chandramouli, and Thomas L. Griffiths introduce AutoCog in Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist. arXiv lists version 1 as submitted June 24, 2026, under Neurons and Cognition (q-bio.NC), with Artificial Intelligence (cs.AI) as a secondary category. The record describes a 44-page preprint with nine figures.
The paper calls AutoCog a "fully autonomous" agentic system. That phrase needs a boundary. Researchers first specify the task domain and instructions, experiment design space, model design space, two seed theories, and behavior source. Within that envelope, the software can propose and validate experiments, recruit participants, collect responses, compare theories, and generate revisions without another researcher action. The autonomy is operational autonomy inside a human-authored protocol, not open-ended scientific authority.
The title of this essay is therefore institutional shorthand. AutoCog can occupy several workflow roles conventionally performed by cognitive scientists. It is not a claim that the system is conscious, a person, or competent across science.
What AutoCog Automates
Each theory slot contains two linked objects: a natural-language account and executable prediction and choice functions with admissible parameter ranges. In each cycle, two LLM agents propose experiments and computable metrics intended to discriminate their incumbent theories. Programmatic simulation checks whether a proposal separates the models at the planned sample size before the experiment is committed.
The committed experiment is rendered as a browser task. In the human runs, the paper reports that SweetBean and AutoRA handled deployment, Prolific supplied participants, and Google Firebase stored trial-level responses. The observed metric is compared with distributions generated by each executable theory. An LLM arbiter interprets those results, and a separate revision stage replaces or repairs the weaker account. Later candidates are tested against the accumulated experiment–metric record, not only the newest study.
The reported runs used Google Gemini 3.1 Pro Preview for the generative stages, with prompts, structured outputs, retry budgets, model identifiers, and run products recorded. The authors' public repository contains code, tests, committed result trees, and analysis scripts. Those artifacts make inspection possible; they do not by themselves establish that an outside team can recreate the findings.
The Evidence Ladder
The evidence has three distinct levels, and they should not be collapsed.
- Synthetic recovery. AutoCog was given behavior generated by known decision rules and asked to recover the hidden rule from different seeds. The paper reports five-cycle recovery for canonical strategies and several deliberately unconventional strategies, with harder single-cue and anti-majority cases requiring 20 cycles and remaining less accurate. This tests search and recovery under known ground truth; it does not test whether the resulting account explains human cognition.
- Adaptive human-data discovery. In each of two decision-task settings, the reported loop ran for five cycles, launching two studies per cycle with 25 participants per study. Surfaced models improved on the seeds over the accumulated studies. In the binary-feature setting, the selected model also fit a previously published held-out dataset better than the named baselines. This is stronger than in-sample fit but remains a narrow task and comparison set.
- Prospective test. The cardinal-feature loop surfaced Diminishing Returns WADD, which applies a concave transform before weighting feature values by cue validity. The team then froze stimuli, model definitions, power simulations, exclusion rules, and confirmatory code in an OSF preregistration before collecting new data from 50 participants in a model-discrimination experiment and 100 in a value-curvature experiment.
The prospective results deserve precision. The authors report that choices favored Diminishing Returns WADD over three alternatives, with all three planned comparisons surviving Holm correction, and that participants preferred an equal-sized advantage in the lower-value region above chance. A third level-shift effect was in the preregistered direction but small: the reported difference was 0.029 with effect size d = 0.182; its two-sided 95 percent confidence interval, [−0.003, 0.061], included zero, while the preregistered one-tailed test gave p = .036. The fair summary is prospective support of unequal strength, not final confirmation of a universal mechanism.
Current Context
AutoCog did not arrive alone. Twenty-two minutes later on the same June 24 arXiv submission day, Ben Prystawski and colleagues posted auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation. That system used nested model-critique and experiment-design loops, programmatic Prolific recruitment, and a different testbed: judgments of randomness in coin-flip sequences.
The parallel preprint is important context, not replication. It used different investigators, cognitive-model formalisms, tasks, analyses, and participant runs; it does not reproduce AutoCog's diminishing-returns result. Together, the two version-1 preprints show that closed-loop computational cognitive science became a live research pattern by mid-2026. They also make broad priority language fragile. The durable question is which claims survive frozen tests, external populations, different model backends, and independent teams.
Why the Loop Matters
The governance issue is not whether a model "had an idea." It is whether the loop preserves the difference among four objects: a verbal theory, its executable implementation, its predictive performance, and a claim about the mechanism people actually use. Good fit can support the third object without proving the fourth. Compiling code proves only that the implementation runs.
AutoCog makes that distinction unusually visible. Its empirical comparison is programmatic, but an LLM arbiter interprets the record, and a second Gemini model scores mechanism similarity in the synthetic study. The paper also reports manual inspection for some exact-mechanism judgments. These are useful evaluation layers, not independent ground truth. A credible record should expose the numerical score, arbiter rationale, judge prompt and samples, manual rubric, and any disagreement among them.
The loop is also adaptive: later theories and experiments are chosen after earlier outcomes are known. That is appropriate for discovery, but accumulated performance inside the loop should not silently become confirmatory evidence. Center for Open Science guidance distinguishes exploratory search from a prespecified confirmatory test. AutoCog's separate, time-stamped follow-up is therefore a central strength of the paper, not a procedural footnote.
Where the Evidence Is Narrow
The human evidence concerns forced choices between fictitious products described by expert ratings. Participants came from one online platform, and the discovery studies used population-level choice proportions from 25 people per committed experiment. The held-out and prospective tests remain close to the same multi-attribute task family. This is a useful proof of workflow in a tractable domain, not evidence of generality across psychological constructs, cultures, laboratories, or consequential interventions.
The search is bounded by human choices: seeds, allowed inputs and outputs, parameter ranges, experiment variables, score construction, sample budget, and cycle count. The paper explicitly notes that parsimony, unification, and interpretability are not all direct objectives. A loop can therefore find a better predictor within its envelope while missing a simpler account, an excluded variable, an individual-difference structure, or a mechanism outside its vocabulary.
There is also a translation failure mode acknowledged by the authors. AutoCog checks whether generated code compiles and predicts the accumulated data, but not whether that code faithfully realizes its paired prose. Theory-to-code governance needs behavioral invariants, counterexample tests, code review, and a signed mapping from each verbal mechanism to the operations that implement it.
Finally, model and infrastructure versions are part of the experiment. The reported agents depended on a preview model, cloud APIs, a browser-task stack, Prolific, and Firebase. A later replay can differ because an endpoint, dependency, recruitment pool, or platform policy changed even when the prompt did not. Reproducibility therefore requires pinned artifacts and environment records, not a model name alone.
The Human Role Moves Upstream
Human participants are not an interchangeable data API. The paper reports a Princeton University IRB protocol, adults aged 18–55, an informed-consent screen, voluntary withdrawal, disclosed data handling, and compensation of US$0.80 for sessions with a reported median duration of about six to eight minutes. Those details matter because the loop launched real studies and wrote trial-level responses to cloud infrastructure.
In the United States, the Common Rule sets basic provisions for institutional review boards and informed consent, while the Belmont Report frames human-subject research around respect for persons, beneficence, and justice. An automated loop does not weaken those duties. Its approved action space must stay inside the protocol: allowed stimuli and variables, participant eligibility, burden and payment rules, privacy controls, recruitment and spending ceilings, withdrawal handling, incident response, and a human stop authority.
This is where responsibility concentrates. Named investigators and institutions remain accountable for the research question, participant protections, statistical claims, data stewardship, and publication decision. An arbiter can recommend a successor model; it cannot accept ethical responsibility or sign off on an inference.
Governance Standard
An automated scientific loop needs two gates and one durable receipt.
- Discovery gate. Before launch, record the responsible investigators, ethics determination and protocol version, permitted design and model spaces, prohibited manipulations, seed theories, model and dependency versions, prompts, metrics, participant and cost ceilings, stopping rules, and who can halt the run.
- Confirmation gate. Treat adaptive-loop results as exploratory. Freeze the selected theory and code, hypotheses, stimuli, sampling plan, exclusions, analysis, decision thresholds, and multiplicity treatment before new outcomes are observed. Report deviations and null results, and require independent replication before making broad mechanism or deployment claims.
- Discovery receipt. Preserve every proposed and committed experiment, metric, validation outcome, consent version, recruitment event, raw-to-analysis data lineage, exclusion, failed candidate, retry, code hash, parameter range, numerical score, arbiter verdict, human intervention, final claim, and sign-off. The W3C PROV-O Recommendation supplies a general vocabulary for linking entities, activities, and responsible agents; a domain-specific receipt can build on that model.
The decisive separation is: proposed, implemented, predicted, prospectively supported, and independently replicated. Those are stages, not synonyms. A system that exposes them can contribute to cumulative science. A system that merges them into one polished theory becomes a persuasion interface.
Source Discipline
This review treats the AutoCog manuscript as arXiv:2606.26448v1 and attributes its empirical results to the authors. It inspected the arXiv abstract and HTML, public repository, and linked OSF preregistration; it did not rerun the agent workflow or reproduce the statistical analyses. The repository increases inspectability but remains a mutable project location unless a specific commit or archival release is cited.
"First," "novel," "fully autonomous," and "confirmed" are claims that need scope. Priority is complicated by parallel same-day work. Autonomy begins after human specification. Novelty is evaluated relative to the paper's literature and search space. The prospective study is stronger than adaptive fit, but its three reported result families are not equally strong and are not an independent laboratory replication.
Related Pages
For the broader category, see AI Scientists and AI in Science. For the recordkeeping and validation layer, continue to The Lab Notebook Becomes the Discovery Engine, The Equation Search Becomes the Closed-Loop Instrument, The Open Artifact Becomes the Reproducibility Receipt, AI Audit Trails, and The Peer Reviewer Becomes the Model Referee. Site-level commitments appear in Research and Editorial Integrity and Privacy and Data.
Sources
- Akshay K. Jagadish, Younes Strittmatter, Nori Jacoby, George Kachergis, Eric Schulz, Nathaniel Daw, Suyog H. Chandramouli, and Thomas L. Griffiths, Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist, arXiv:2606.26448v1 [q-bio.NC], submitted June 24, 2026. Primary versions checked: HTML and PDF.
- AutoCog primary artifacts: public code, data, results, tests, and analysis repository; preregistered Diminishing Returns WADD follow-up.
- Ben Prystawski, Kushin Mukherjee, Daniel Wurgaft, Linas Nasvytis, Michael Y. Li, Noah D. Goodman, and Michael C. Frank, auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation, arXiv:2606.26460v1 [cs.AI], submitted June 24, 2026.
- U.S. Department of Health and Human Services, Office for Human Research Protections: Federal Policy for the Protection of Human Subjects (Common Rule) and The Belmont Report.
- Center for Open Science, exploratory and confirmatory research in preregistration; World Wide Web Consortium, PROV-O: The PROV Ontology, W3C Recommendation, April 30, 2013.
- Internal context: AI Scientists, AI in Science, AI Audit Trails, and Research and Editorial Integrity.