Can AI Scientists Do Independent Research Yet?
A 2026 evidence review of AI scientists, benchmark gaps, reproducibility limits, and the role human researchers still play.

In 2025, a manuscript produced by an autonomous research system passed peer review at an ICLR workshop. That milestone shows that an AI scientist pipeline can complete a research cycle. It does not show that the system can conduct reliable, independent science. The evidence in 2026 supports a narrower conclusion: end-to-end automation is real, while reproducible independent research still requires human judgment.
What "end to end" actually means
An AI scientist pipeline can propose an idea, write code, run computational experiments, analyze results, draft a manuscript, and review its own work. Sakana AI's AI Scientist-v2 did that without human-authored experiment templates, and one of its manuscripts was accepted at a workshop.
Completing every stage is an engineering achievement. Completing each stage correctly, consistently, and with evidence that another researcher can inspect is the scientific standard. Those are separate claims.
The evidence: impressive, but roughly half the expert score
The clearest gap appears on end-to-end evaluations. In Stanford HAI's 2026 AI Index science chapter, the leading PaperArena agent scored 38.8%, compared with 83.5% for PhD experts—roughly half the expert score. On ReplicationBench, frontier models completed fewer than 20% of paper-scale astrophysics replication tasks. ChemBench tells the same mixed story: leading models exceeded average human performance across more than 2,700 chemistry questions while still making basic errors.

Figure 1. End-to-end research performance remains well below the expert reference score, even as models improve on isolated tasks.
The pattern matters. AI systems can be strong at literature search, code generation, and analysis yet weak at coordinating those steps into research that is correct, auditably documented, and repeatable.
The strongest AI scientist results still include researchers
Recent successes are collaborations, not evidence that people have left the loop. FutureHouse's Robin searched literature, nominated candidate molecules, and selected assays for dry age-related macular degeneration; human researchers ran the experiments and returned the results. The team estimated that this collaboration reduced the project timeline by roughly 200-fold, as reported in Nature.
Google's AI Co-Scientist was evaluated with scientists in three biomedical areas, including drug repurposing for acute myeloid leukemia, targets for liver fibrosis, and an antimicrobial-resistance mechanism. The study reports independent in vitro validation, with experts framing, guiding, and checking the work.

Figure 2. Current high-value systems divide the loop: AI expands search and generation; researchers frame questions, run decisive experiments, and judge the evidence.
The useful question is therefore not whether an AI can produce a plausible target. It is whether the assumptions, evidence, methods, and experimental decisions remain visible enough for a researcher to challenge.
The real bottleneck: reproducibility and auditability
Independent research requires more than a polished result. Computational steps should record code, data, parameters, environment, and provenance so another researcher can reconstruct what happened. Stochastic or physical experiments may not replay bit for bit, but their methods and uncertainty still need enough documentation for an independent test.

Figure 3. An audit trail and controlled execution support reproducibility; scientific validity still depends on review of the evidence and method.
The low ReplicationBench result does not prove that one architectural defect explains every failure. It does show that success on question answering is not enough: a research system must coordinate long chains of execution and preserve the record needed to assess them.
What people get wrong
- "It wrote a paper, so it can do research independently." Workshop acceptance shows that a pipeline can produce a reviewable manuscript. It does not establish that its findings will replicate.
- "Question-answering benchmarks settle the issue." Isolated answers and paper-scale execution test different capabilities.
- "Humans are only slowing the loop down." In the strongest biomedical examples, human framing, experimentation, and verification are part of what made the result credible.
FAQ
Can an AI scientist do research without human involvement? Not reliably today. End-to-end systems exist, but current paper-scale and end-to-end evaluations remain far below expert performance, and validated laboratory results still involve researchers.
What is the difference between answering scientific questions and doing research? Question answering tests an isolated response. Research connects hypothesis, method, execution, analysis, and replication, with evidence that supports each transition.
Where is AI already useful in scientific research? AI can accelerate literature search, candidate generation, coding, data analysis, and drafting. Researchers still own the framing, decisive experiments, and final interpretation.
What would make AI research more independent? Better end-to-end reliability, reproducible execution, inspectable provenance, and robust validation would reduce how much supervision each stage requires.
What should teams evaluate in an AI research platform? Look for project continuity and paper-reproduction workflows, then test whether the system preserves the sources, settings, intermediate results, and failures needed to verify its output.
Related reading: Reproducible Research: Closing the Gap in Scientific AI examines the execution standard, while AI for Materials Science shows how generation, simulation, and experiments can form a testable loop.
Where this leaves us
AI scientists can now complete research-shaped workflows, but "independent" remains too strong for reliable scientific practice. The practical opportunity is collaboration: let AI compress search, drafting, coding, and iteration while researchers define the question and verify the chain of evidence.
Mira Science: Mira—presented on its website as AI Scientist Mira—positions itself as an AI Research Partner, lists Paper Reproduction among its core workflows, and says project experience can settle into the platform as work proceeds. Use a bounded task to evaluate that workflow against your own verification standard. Start researching with Mira →

