Compare commits

...
Author SHA1 Message Date
promptadmin b9cea1ed17 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#56) from upstream-sync/awesome-ai-for-science-20260719-086e63-xinf into main
Reviewed-on: #56
2026-07-19 15:41:27 +00:00
promptadmin 1ab812ba92 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@086e63bb [catalogue] 2026-07-19 15:23:33 +00:00
promptadmin ff0fbd8785 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#49) from upstream-sync/awesome-ai-for-science-20260715-709cee-vpdg into main
Reviewed-on: #49
2026-07-15 16:12:47 +00:00
promptadmin 268b9aaacc [upstream-sync] README.md from ai-boost/awesome-ai-for-science@709cee28 [catalogue] 2026-07-15 15:01:02 +00:00
promptadmin 71a89dca18 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#41) from upstream-sync/awesome-ai-for-science-20260710-230e87-gsmv into main
Reviewed-on: #41
2026-07-10 23:36:12 +00:00
promptadmin bdd7e1176f [upstream-sync] README.md from ai-boost/awesome-ai-for-science@230e8724 [catalogue] 2026-07-10 20:32:23 +00:00
promptadmin ad75261e53 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#37) from upstream-sync/awesome-ai-for-science-20260708-408a2a-oldc into main
Reviewed-on: #37
2026-07-09 01:58:46 +00:00
promptadmin df8c25d33b [upstream-sync] README.md from ai-boost/awesome-ai-for-science@408a2a18 [catalogue] 2026-07-08 20:23:24 +00:00
promptadmin 871e7a94db Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#35) from upstream-sync/awesome-ai-for-science-20260707-4afd72-bcxj into main
Reviewed-on: #35
2026-07-07 21:22:48 +00:00
promptadmin e22c861d46 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@4afd72ec [catalogue] 2026-07-07 20:18:40 +00:00
promptadmin bf0d468978 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260706-e0df96-elsa
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:58 +00:00
promptadmin 336dce2f2e Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260704-cf7431-gaou
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:38 +00:00
promptadmin cb7edbda94 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260704-9f62d8-qyhf
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:22 +00:00
promptadmin d82784013c Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260704-cca092-fmor
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:15 +00:00
promptadmin 79442f5a4f Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260703-86c5f1-awps
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:11:34 +00:00
promptadmin e2d27fc52e Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#31) from upstream-sync/awesome-ai-for-science-20260705-99a575-uivu into main
Reviewed-on: #31
2026-07-06 15:05:14 +00:00
promptadmin a960883a91 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@e0df96d7 [catalogue] 2026-07-06 08:09:54 +00:00
promptadmin 21abde0feb [upstream-sync] README.md from ai-boost/awesome-ai-for-science@99a57577 [catalogue] 2026-07-05 20:08:20 +00:00
promptadmin 39e6e43df2 Merge pull request '[Upstream sync] HKUST-KnowComp/Awesome-LLM-Scientific-Discovery (github) — 0 added, 1 modified' (#26) from upstream-sync/awesome-llm-scientific-discovery-20260703-eb19b4-nzbk into main
Reviewed-on: #26
2026-07-05 14:10:06 +00:00
promptadmin 80ce4af09b Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#30) from upstream-sync/awesome-ai-for-science-20260705-a65119-tfox into main
Reviewed-on: #30
2026-07-05 14:07:01 +00:00
promptadmin 4dd3bb3374 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@a6511998 [catalogue] 2026-07-05 08:05:31 +00:00
promptadmin 55837bb543 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@cf74319a [catalogue] 2026-07-04 19:58:44 +00:00
promptadmin bfc772b9f6 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@9f62d8a3 [catalogue] 2026-07-04 13:57:30 +00:00
promptadmin 168b1b0083 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@cca0924c [catalogue] 2026-07-04 07:56:14 +00:00
promptadmin f7cdcba740 [upstream-sync] README.md from HKUST-KnowComp/Awesome-LLM-Scientific-Discovery@eb19b47e [catalogue] 2026-07-03 19:54:46 +00:00
promptadmin 7fd95fb00f [upstream-sync] README.md from ai-boost/awesome-ai-for-science@86c5f147 [catalogue] 2026-07-03 19:54:17 +00:00
promptadmin 3959588d37 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260702-2d6b81-ailc
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-03 14:01:54 +00:00
promptadmin adb877a343 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260702-e92baf-yafz
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-03 13:58:17 +00:00
promptadmin 64b83d6ceb Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260630-f39947-kdrm
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-03 13:57:56 +00:00
promptadmin 4038c1c7cd Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#23) from upstream-sync/awesome-ai-for-science-20260703-de8e26-cqsi into main
Reviewed-on: #23
2026-07-03 13:49:25 +00:00
promptadmin fd47cea53e Merge pull request '[Upstream sync] HKUST-KnowComp/Awesome-LLM-Scientific-Discovery (github) — 0 added, 1 modified' (#24) from upstream-sync/awesome-llm-scientific-discovery-20260703-7fcb88-vvmf into main
Reviewed-on: #24
2026-07-03 13:46:15 +00:00
promptadmin 4b2bcc5978 [upstream-sync] README.md from HKUST-KnowComp/Awesome-LLM-Scientific-Discovery@7fcb8811 [catalogue] 2026-07-03 07:52:08 +00:00
promptadmin eb75b1a36b [upstream-sync] README.md from ai-boost/awesome-ai-for-science@de8e2603 [catalogue] 2026-07-03 07:51:37 +00:00
promptadmin 3db18ec07f [upstream-sync] README.md from ai-boost/awesome-ai-for-science@2d6b813f [catalogue] 2026-07-02 19:50:02 +00:00
promptadmin 40f6c03acd Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#20) from upstream-sync/awesome-ai-for-science-20260701-0cee46-nqrz into main
Reviewed-on: #20
2026-07-02 19:22:12 +00:00
promptadmin 759b1b39fa [upstream-sync] README.md from ai-boost/awesome-ai-for-science@e92baf5c [catalogue] 2026-07-02 07:47:43 +00:00
promptadmin c6fd0dfd73 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@0cee4659 [catalogue] 2026-07-01 19:43:30 +00:00
promptadmin 1730fc59ef Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#19) from upstream-sync/awesome-ai-for-science-20260701-3756bb-ajvc into main
Reviewed-on: #19
2026-07-01 13:59:37 +00:00
promptadmin a6fef26c79 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@3756bb0f [catalogue] 2026-07-01 07:41:32 +00:00
promptadmin 17b875842a [upstream-sync] README.md from ai-boost/awesome-ai-for-science@f3994796 [catalogue] 2026-06-30 19:38:11 +00:00
promptadmin 0f5f8a3c04 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260628-dbd35d-yahj
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-06-30 16:41:48 +00:00
promptadmin d0e4258f2c Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260628-f5d952-szrr
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-06-30 16:41:27 +00:00
promptadmin 8bdc92fedf Merge branch 'upstream-sync/awesome-ai-for-science-20260629-602f33-ctuq'
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-06-30 16:12:26 +00:00
promptadmin dc6637c082 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#17) from upstream-sync/awesome-ai-for-science-20260630-284a2b-raws into main
Reviewed-on: #17
2026-06-30 16:00:18 +00:00
promptadmin 4caa323bfe [upstream-sync] README.md from ai-boost/awesome-ai-for-science@284a2b4b [catalogue] 2026-06-30 07:35:53 +00:00
promptadmin f397977f7d [upstream-sync] README.md from ai-boost/awesome-ai-for-science@602f33d3 [catalogue] 2026-06-29 19:30:05 +00:00
promptadmin c9b1b1a117 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#15) from upstream-sync/awesome-ai-for-science-20260629-d3eb73-nuat into main
Reviewed-on: #15
2026-06-29 15:56:43 +00:00
promptadmin 254346c6ca [upstream-sync] README.md from ai-boost/awesome-ai-for-science@d3eb7394 [catalogue] 2026-06-29 07:27:39 +00:00
promptadmin 1d855bf433 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@f5d9529d [catalogue] 2026-06-28 19:25:52 +00:00
promptadmin 8d7e346189 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@dbd35d0a [catalogue] 2026-06-28 07:23:44 +00:00
promptadmin 65c3b0ffea cleanup test files 2026-06-27 21:17:48 +00:00
promptadmin 04bfa28b4c cleanup test files 2026-06-27 21:17:25 +00:00
promptadmin fffc13d595 cleanup test files 2026-06-27 21:17:05 +00:00
promptadmin ef48c29b12 cleanup test files 2026-06-27 21:16:57 +00:00
promptadmin c249190497 cleanup test files 2026-06-27 21:16:48 +00:00
promptadmin a98b7daad7 wh url fixed 2026-06-27 21:09:50 +00:00
promptadmin 759de2573a final wh test 2026-06-27 21:03:33 +00:00
promptadmin 4e669b779f wh test2 2026-06-27 20:47:21 +00:00
promptadmin e76b1a0707 test sig fix 2026-06-27 20:23:40 +00:00
3 changed files with 183 additions and 18 deletions
@@ -2,9 +2,9 @@
title: "Awesome LLM Scientific Discovery [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)"
task: ""
lineage_type: import
upstream_source: https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery/blob/b39b55ac/README.md
upstream_sha: b39b55ac
imported_at: 2026-06-26
upstream_source: https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery/blob/eb19b47e/README.md
upstream_sha: eb19b47e
imported_at: 2026-07-03
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -28,6 +28,8 @@ Below is a visual representation of this taxonomy:
We aim to provide a comprehensive overview for researchers, developers, and enthusiasts interested in this rapidly advancing field.
> **Last major update: 2026.07.** This refresh adds a large batch of 20252026 papers and a dedicated section on frontier industry-lab systems (Google DeepMind, OpenAI, Microsoft Research, Meta FAIR, FutureHouse, Sakana AI, and others). Contributions and PRs are very welcome — see [Contributing](#contributing).
## Contents
* [Level 1: LLM as Tool](#level-1-llm-as-tool)
@@ -43,7 +45,13 @@ We aim to provide a comprehensive overview for researchers, developers, and enth
* [Function Discovery](#function-discovery)
* [Natural Science Research](#natural-science-research)
* [General Research](#general-research)
* [Survey Generation](#survey-generation)
* [Level 3: LLM as Scientist](#level-3-llm-as-scientist)
* [General-Purpose Autonomous Research Agents](#general-purpose-autonomous-research-agents)
* [Discovery-Oriented Scientific Systems](#discovery-oriented-scientific-systems)
* [Autonomous Research Ecosystems and Infrastructure](#autonomous-research-ecosystems-and-infrastructure)
* [Frontier Labs and Foundation Models for Science](#frontier-labs-and-foundation-models-for-science)
* [Other Related Works](#other-related-works)
* [Contributing](#contributing)
---
@@ -65,6 +73,12 @@ Automating literature search, retrieval, synthesis, structuring, and organizatio
* **LitLLM: A Toolkit for Scientific Literature Review** [![arXiv](https://img.shields.io/badge/arXiv-2402.01788v1-B31B1B.svg)](https://arxiv.org/pdf/2402.01788v1) - *Agarwal et al. (2024.02)*
* **Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain** [![DOI](https://img.shields.io/badge/DOI-10.1186/s13643--024--02575--4-blue.svg)](https://systematicreviewsjournal.biomedcentral.com/articles/10.1186/s13643-024-02575-4) - *Dennstädt et al. (2024.06)*
* **Science Hierarchography: Hierarchical Organization of Science Literature** [![arXiv](https://img.shields.io/badge/arXiv-2504.13834-B31B1B.svg)](https://arxiv.org/pdf/2504.13834) - *Gao et al. (2025.04)*
* **Language Agents Achieve Superhuman Synthesis of Scientific Knowledge (PaperQA2)** [![arXiv](https://img.shields.io/badge/arXiv-2409.13740-B31B1B.svg)](https://arxiv.org/pdf/2409.13740) - *Skarlinski et al. (2024.09)*
* **DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents** [![arXiv](https://img.shields.io/badge/arXiv-2506.11763-B31B1B.svg)](https://arxiv.org/pdf/2506.11763) - *Du et al. (2025.06)*
* **Deep Research Agents: A Systematic Examination And Roadmap** [![arXiv](https://img.shields.io/badge/arXiv-2506.18096-B31B1B.svg)](https://arxiv.org/pdf/2506.18096) - *Huang et al. (2025.06)*
* **DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis** [![arXiv](https://img.shields.io/badge/arXiv-2508.20033-B31B1B.svg)](https://arxiv.org/pdf/2508.20033) - *Patel et al. (2025.08)*
* **ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry** [![arXiv](https://img.shields.io/badge/arXiv-2507.16280-B31B1B.svg)](https://arxiv.org/pdf/2507.16280) - *Xu et al. (2025.07)*
* **LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation** [![arXiv](https://img.shields.io/badge/arXiv-2510.05138-B31B1B.svg)](https://arxiv.org/pdf/2510.05138) - *Zhang et al. (2025.10)*
### Idea Generation and Hypothesis Formulation
@@ -92,6 +106,8 @@ Automated generation of novel research ideas, conceptual insights, and testable
* **Scideator: Human-LLM Scientific Idea Generation Grounded in Research-Paper Facet Recombination** [![arXiv](https://img.shields.io/badge/arXiv-2409.14634-B31B1B.svg)](https://arxiv.org/pdf/2409.14634) - *Radensky et al. (2024.09)*
* **HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance** [![arXiv](https://img.shields.io/badge/arXiv-2506.12937-B31B1B.svg)](https://arxiv.org/pdf/2506.12937) - *Vasu et al. (2025.06)*
* **Sparks of Science: Hypothesis Generation Using Structured Paper Data** [![arXiv](https://img.shields.io/badge/arXiv-2504.12976-B31B1B.svg)](https://arxiv.org/pdf/2504.12976) - *O'Neill et al. (2025.04)*
* **A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2504.05496-B31B1B.svg)](https://arxiv.org/pdf/2504.05496) - *Kulkarni et al. (2025.04)*
* **Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Networks** [![arXiv](https://img.shields.io/badge/arXiv-2511.02238-B31B1B.svg)](https://arxiv.org/pdf/2511.02238) - *Wang et al. (2025.11)*
### Experiment Planning and Execution
@@ -104,6 +120,7 @@ LLMs assisting in experimental protocol planning, workflow design, and scientifi
* **Natural Language to Code Generation in Interactive Data Science Notebooks** [![arXiv](https://img.shields.io/badge/arXiv-2212.09248-B31B1B.svg)](https://arxiv.org/pdf/2212.09248) - *Yin et al. (2022.12)*
* **DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation** [![arXiv](https://img.shields.io/badge/arXiv-2211.11501-B31B1B.svg)](https://arxiv.org/pdf/2211.11501) - *Lai et al. (2022.11)*
* **Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents**, [![arXiv](https://img.shields.io/badge/arXiv-2502.16069-B31B1B.svg)](https://arxiv.org/pdf/2502.16069) - *Kon et al. (2025.02)*
* **AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing** [![arXiv](https://img.shields.io/badge/arXiv-2602.17607-B31B1B.svg)](https://arxiv.org/pdf/2602.17607) - *Du et al. (2026.02)*
### Data Analysis and Organization
@@ -117,6 +134,7 @@ LLMs assisting in data-driven analysis, tabular/chart reasoning, statistical rea
* **Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding** [![arXiv](https://img.shields.io/badge/arXiv-2401.04398-B31B1B.svg)](https://arxiv.org/pdf/2401.04398) - *Wang et al. (2024.01)*
* **TableBench: A Comprehensive and Complex Benchmark for Table Question Answering** [![arXiv](https://img.shields.io/badge/arXiv-2408.09174-B31B1B.svg)](https://arxiv.org/pdf/2408.09174) - *Wu et al. (2024.08)*
* **Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs** [![arXiv](https://img.shields.io/badge/arXiv-2402.12424-B31B1B.svg)](https://arxiv.org/pdf/2402.12424) - *Deng et al. (2024.02)*
* **ChatSpatial: Schema-Enforced Agentic Orchestration for Reproducible and Cross-Platform Spatial Transcriptomics** [![DOI](https://img.shields.io/badge/DOI-10.64898/2026.02.26.708361-blue.svg)](https://doi.org/10.64898/2026.02.26.708361) - *Yang et al. (2026.02)* [Code](https://github.com/cafferychen777/ChatSpatial)
### Conclusion and Hypothesis Validation
@@ -134,8 +152,12 @@ LLMs providing feedback, verifying claims, replicating results, and generating r
* **Advancing AI-Scientist Understanding: Making LLM Think Like a Physicist with Interpretable Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2504.01911-B31B1B.svg)](https://arxiv.org/pdf/2504.01911) - *Xu et al. (2025.04)*
* **Generative Adversarial Reviews: When LLMs Become the Critic** [![arXiv](https://img.shields.io/badge/arXiv-2412.10415-B31B1B.svg)](https://arxiv.org/pdf/2412.10415) - *Bougie & Watanabe (2024.12)*
* **Predicting Empirical AI Research Outcomes with Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2506.00794-B31B1B.svg)](https://arxiv.org/pdf/2506.00794) - *Wen et al. (2025.06)*
* **SPOT: When AI Co-Scientists Fail — A Benchmark for Automated Verification of Scientific Research** [![arXiv](https://img.shields.io/badge/arXiv-2505.11855-B31B1B.svg)](https://arxiv.org/pdf/2505.11855) - *Son et al. (2025.05)*
* **DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process** [![arXiv](https://img.shields.io/badge/arXiv-2503.08569-B31B1B.svg)](https://arxiv.org/pdf/2503.08569) - *Zhu et al. (2025.03)*
* **ReviewRL: Towards Automated Scientific Review with RL** [![arXiv](https://img.shields.io/badge/arXiv-2508.10308-B31B1B.svg)](https://arxiv.org/pdf/2508.10308) - *Zeng et al. (2025.08)*
* **SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification** [![arXiv](https://img.shields.io/badge/arXiv-2502.10003-B31B1B.svg)](https://arxiv.org/pdf/2502.10003) - *Kumar et al. (2025.02)*
* **LMR-Bench: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research** [![arXiv](https://img.shields.io/badge/arXiv-2506.17335-B31B1B.svg)](https://arxiv.org/pdf/2506.17335) - *Yan et al. (2025.06)*
* **REFUTE: Reasoning Over Evidence - Falsification, Uncertainty, Truth-grounding & Epistemics** [![HF Dataset](https://img.shields.io/badge/HuggingFace-dataset-yellow.svg)](https://huggingface.co/datasets/BGPT-OFFICIAL/refute) - *BGPT (2026.06)*. Open benchmark for scientific critique and epistemic calibration on recent science paper summaries, covering falsification, limitations, overclaims, missing-evidence refusal, calibration, and planted-flaw detection.
### Iteration and Refinement
@@ -167,6 +189,25 @@ Automated modeling of machine learning tasks, experiment design, and execution.
* **AIDE: AI-Driven Exploration in the Space of Code** [![arXiv](https://img.shields.io/badge/arXiv-2502.13138-B31B1B.svg)](https://arxiv.org/pdf/2502.13138) - *Jiang et al. (2025.02)*
* **Language Modeling by Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2506.20249-B31B1B.svg)](https://arxiv.org/pdf/2506.20249) - *Cheng et al. (2025.06)*
* **MLGym: A New Framework and Benchmark for Advancing AI Research Agents** [![arXiv](https://img.shields.io/badge/arXiv-2502.14499-B31B1B.svg)](https://arxiv.org/pdf/2502.14499) - *Nathani et al. (2025.02)*
* **R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science** [![arXiv](https://img.shields.io/badge/arXiv-2505.14738-B31B1B.svg)](https://arxiv.org/pdf/2505.14738) - *Xu et al. (2025.05)* — Microsoft Research
* **MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement** [![arXiv](https://img.shields.io/badge/arXiv-2506.15692-B31B1B.svg)](https://arxiv.org/pdf/2506.15692) - *Nam et al. (2025.06)* — Google
* **ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2506.16499-B31B1B.svg)](https://arxiv.org/pdf/2506.16499) - *Liu et al. (2025.06)*
* **ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering** [![arXiv](https://img.shields.io/badge/arXiv-2505.23723-B31B1B.svg)](https://arxiv.org/pdf/2505.23723) - *Liu et al. (2025.05)*
* **AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench** [![arXiv](https://img.shields.io/badge/arXiv-2507.02554-B31B1B.svg)](https://arxiv.org/pdf/2507.02554) - *Toledo et al. (2025.07)* — Meta / UCL
* **The FM Agent** [![arXiv](https://img.shields.io/badge/arXiv-2510.26144-B31B1B.svg)](https://arxiv.org/pdf/2510.26144) - *Li et al. (2025.10)*
* **KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for ML Problems** [![arXiv](https://img.shields.io/badge/arXiv-2508.10177-B31B1B.svg)](https://arxiv.org/pdf/2508.10177) - *Kulibaba et al. (2025.08)*
* **AutoMLGen: Navigating Fine-Grained Optimization for Coding Agents** [![arXiv](https://img.shields.io/badge/arXiv-2510.08511-B31B1B.svg)](https://arxiv.org/pdf/2510.08511) - *Du et al. (2025.10)*
* **ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies** [![arXiv](https://img.shields.io/badge/arXiv-2504.20117-B31B1B.svg)](https://arxiv.org/pdf/2504.20117) - *Gandhi et al. (2025.04)*
* **AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML** [![arXiv](https://img.shields.io/badge/arXiv-2410.02958-B31B1B.svg)](https://arxiv.org/pdf/2410.02958) - *Trirat et al. (2024.10)*
* **SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning** [![arXiv](https://img.shields.io/badge/arXiv-2410.17238-B31B1B.svg)](https://arxiv.org/pdf/2410.17238) - *Chi et al. (2024.10)*
* **AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions** [![arXiv](https://img.shields.io/badge/arXiv-2410.20424-B31B1B.svg)](https://arxiv.org/pdf/2410.20424) - *Li et al. (2024.10)*
* **Agent K: Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Performance** [![arXiv](https://img.shields.io/badge/arXiv-2411.03562-B31B1B.svg)](https://arxiv.org/pdf/2411.03562) - *Grosnit et al. (2024.11)* — Huawei Noah's Ark
* **EXP-Bench: Can AI Conduct AI Research Experiments?** [![arXiv](https://img.shields.io/badge/arXiv-2505.24785-B31B1B.svg)](https://arxiv.org/pdf/2505.24785) - *Kon et al. (2025.05)*
* **InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research** [![arXiv](https://img.shields.io/badge/arXiv-2510.27598-B31B1B.svg)](https://arxiv.org/pdf/2510.27598) - *Wu et al. (2025.10)*
* **MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research** [![arXiv](https://img.shields.io/badge/arXiv-2505.19955-B31B1B.svg)](https://arxiv.org/pdf/2505.19955) - *Chen et al. (2025.05)*
* **RExBench: Can Coding Agents Autonomously Implement AI Research Extensions?** [![arXiv](https://img.shields.io/badge/arXiv-2506.22598-B31B1B.svg)](https://arxiv.org/pdf/2506.22598) - *Edwards et al. (2025.06)*
* **ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution** [![arXiv](https://img.shields.io/badge/arXiv-2509.19349-B31B1B.svg)](https://arxiv.org/pdf/2509.19349) - *Lange et al. (2025.09)* — Sakana AI
* **The AI CUDA Engineer: Agentic CUDA Kernel Discovery, Optimization and Composition** [![Link](https://img.shields.io/badge/Link-Sakana_Report-blue.svg)](https://pub.sakana.ai/ai-cuda-engineer/paper/) - *Sakana AI (2025.02)*
### Data Modeling and Analysis
@@ -181,7 +222,14 @@ Automated data-driven analysis, statistical data modeling, and hypothesis valida
* **Large Language Models for Scientific Synthesis, Inference and Explanation** [![arXiv](https://img.shields.io/badge/arXiv-2310.07984-B31B1B.svg)](https://arxiv.org/pdf/2310.07984) - *Zheng et al. (2023.10)*
* **MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem** [![arXiv](https://img.shields.io/badge/arXiv-2505.14148-B31B1B.svg)](https://arxiv.org/pdf/2505.14148) - *Liu et al. (2025.05)*
* **DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?** [![arXiv](https://img.shields.io/badge/arXiv-2409.07703-B31B1B.svg)](https://arxiv.org/pdf/2409.07703) - *Jing et al. (2024.09)*
* **AutoDS: Open-ended Scientific Discovery via Bayesian Surprise** [![arXiv](https://img.shields.io/badge/arXiv-2507.00310-B31B1B.svg)](https://arxiv.org/pdf/2507.00310) - *Agarwal et al. (2025.07)* — Allen Institute for AI
* **DeepAnalyze: Agentic Large Language Models for Autonomous Data Science** [![arXiv](https://img.shields.io/badge/arXiv-2510.16872-B31B1B.svg)](https://arxiv.org/pdf/2510.16872) - *Zhang et al. (2025.10)*
* **DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2410.07331-B31B1B.svg)](https://arxiv.org/pdf/2410.07331) - *Huang et al. (2024.10)*
* **DataSciBench: An LLM Agent Benchmark for Data Science** [![arXiv](https://img.shields.io/badge/arXiv-2502.13897-B31B1B.svg)](https://arxiv.org/pdf/2502.13897) - *Zhang et al. (2025.02)*
* **Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents** [![arXiv](https://img.shields.io/badge/arXiv-2403.05307-B31B1B.svg)](https://arxiv.org/pdf/2403.05307) - *Li et al. (2024.03)*
* **StatEval: A Comprehensive Benchmark for Large Language Models in Statistics** [![arXiv](https://img.shields.io/badge/arXiv-2510.09517-B31B1B.svg)](https://arxiv.org/pdf/2510.09517) - *Yu et al. (2025.10)*
* **LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference** [![arXiv](https://img.shields.io/badge/arXiv-2508.07221-B31B1B.svg)](https://arxiv.org/pdf/2508.07221) - *Wang et al. (2025.08)*
* **OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents** [![arXiv](https://img.shields.io/badge/arXiv-2504.16918-B31B1B.svg)](https://arxiv.org/pdf/2504.16918) - *Thind et al. (2025.04)*
### Function Discovery
@@ -193,11 +241,18 @@ Identifying underlying equations from observational data (AI-driven symbolic reg
* **EvoSLD: Automated neural scaling law discovery with large language models** [![arXiv](https://img.shields.io/badge/arXiv-2507.21184-B31B1B.svg)](https://arxiv.org/abs/2507.21184) - *Lin et al. (2025.07)*
* **DrSR: LLM based Scientific Equation Discovery with Dual Reasoning from Data and Experience** [![arXiv](https://img.shields.io/badge/arXiv-2506.04282-B31B1B.svg)](https://arxiv.org/abs/2506.04282) - *Wang et al. (2025.06)*
* **NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents** [![arXiv](https://img.shields.io/badge/arXiv-2510.07172-B31B1B.svg)](https://arxiv.org/pdf/2510.07172) - *Zheng et al. (2025.10)*
* **LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2503.06512-B31B1B.svg)](https://arxiv.org/pdf/2503.06512) - *Song et al. (2025.03)*
* **SR-Scientist: Scientific Equation Discovery With Agentic AI** [![arXiv](https://img.shields.io/badge/arXiv-2510.11661-B31B1B.svg)](https://arxiv.org/pdf/2510.11661) - *Xia et al. (2025.10)*
* **LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery (SGA)** [![arXiv](https://img.shields.io/badge/arXiv-2405.09783-B31B1B.svg)](https://arxiv.org/pdf/2405.09783) - *Ma et al. (2024.05)*
* **In-Context Symbolic Regression: Leveraging Large Language Models for Function Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2404.19094-B31B1B.svg)](https://arxiv.org/pdf/2404.19094) - *Merler et al. (2024.04)*
* **Symbolic Regression with a Learned Concept Library (LaSR)** [![arXiv](https://img.shields.io/badge/arXiv-2409.09359-B31B1B.svg)](https://arxiv.org/pdf/2409.09359) - *Grayeli et al. (2024.09)*
* **AI-Newton: A Concept-Driven Physical Law Discovery System without Prior Physical Knowledge** [![arXiv](https://img.shields.io/badge/arXiv-2504.01538-B31B1B.svg)](https://arxiv.org/pdf/2504.01538) - *Fang et al. (2025.04)*
* **PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors** [![arXiv](https://img.shields.io/badge/arXiv-2507.15550-B31B1B.svg)](https://arxiv.org/pdf/2507.15550) - *Chen et al. (2025.07)*
* **Finetuning Large Language Model as an Effective Symbolic Regressor (SymbArena)** [![arXiv](https://img.shields.io/badge/arXiv-2508.09897-B31B1B.svg)](https://arxiv.org/pdf/2508.09897) - *Hua et al. (2025.08)*
### Natural Science Research
Autonomous research workflows for natural science discovery (e.g., chemistry, biology, biomedicine).
Autonomous research workflows for natural science discovery (e.g., chemistry, biology, biomedicine, materials, physics).
* **Coscientist: Autonomous Chemical Research with Large Language Models** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--023--06792--0-blue.svg)](https://www.nature.com/articles/s41586-023-06792-0) - *Boiko et al. (2023.10)*
* **Empowering biomedical discovery with AI agents** [![DOI](https://img.shields.io/badge/DOI-10.1016/j.cell.2024.08.026-blue.svg)](https://www.cell.com/action/showPdf?pii=S0092-8674%2824%2901070-5) - *Gao et al. (2024.09)*
@@ -207,9 +262,27 @@ Autonomous research workflows for natural science discovery (e.g., chemistry, bi
* **ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2410.05080-B31B1B.svg)](https://arxiv.org/pdf/2410.05080) - *Chen et al. (2024.10)*
* **ProtAgents: Protein discovery by combining physics and machine learning** [![arXiv](https://img.shields.io/badge/arXiv-2402.04268-B31B1B.svg)](https://arxiv.org/pdf/2402.04268) - *Ghafarollahi and Buehler (2024.02)*
* **Auto-Bench: An Automated Benchmark for Scientific Discovery in LLMs** [![arXiv](https://img.shields.io/badge/arXiv-2502.15224-B31B1B.svg)](https://arxiv.org/pdf/2502.15224) - *Chen et al. (2025.02)*
* **Towards an AI co-scientist** [![arXiv](https://img.shields.io/badge/arXiv-2502.18864-B31B1B.svg)](https://arxiv.org/pdf/2502.18864) - *Gottweis et al. (2025.02)*
* **Towards an AI co-scientist** [![arXiv](https://img.shields.io/badge/arXiv-2502.18864-B31B1B.svg)](https://arxiv.org/pdf/2502.18864) - *Gottweis et al. (2025.02)* — Google
* **GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis** [![arXiv](https://img.shields.io/badge/arXiv-2507.21035-B31B1B.svg)](https://arxiv.org/pdf/2507.21035) - *Liu et al. (2025.07)*
* **Automated Algorithmic Discovery for Gravitational-Wave Detection Guided by LLM-Informed Evolutionary Monte Carlo Tree Search** [![arXiv](https://img.shields.io/badge/arXiv-2508.03661-B31B1B.svg)](https://arxiv.org/pdf/2508.03661) - *Wang and Zeng (2025.08)*
* **The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--025--09442--9-blue.svg)](https://doi.org/10.1038/s41586-025-09442-9) - *Swanson et al. (2025.07)* — Stanford / CZ Biohub
* **Biomni: A General-Purpose Biomedical AI Agent** [![DOI](https://img.shields.io/badge/DOI-10.1101/2025.05.30.656746-blue.svg)](https://doi.org/10.1101/2025.05.30.656746) - *Huang et al. (2025.05)* — Stanford
* **TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools** [![arXiv](https://img.shields.io/badge/arXiv-2503.10970-B31B1B.svg)](https://arxiv.org/pdf/2503.10970) - *Gao et al. (2025.03)*
* **LIDDiA: Language-based Intelligent Drug Discovery Agent** [![arXiv](https://img.shields.io/badge/arXiv-2502.13959-B31B1B.svg)](https://arxiv.org/pdf/2502.13959) - *Averly et al. (2025.02)*
* **LLM Agent Swarm for Hypothesis-Driven Drug Discovery (PharmaSwarm)** [![arXiv](https://img.shields.io/badge/arXiv-2504.17967-B31B1B.svg)](https://arxiv.org/pdf/2504.17967) - *Song et al. (2025.04)*
* **BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation** [![arXiv](https://img.shields.io/badge/arXiv-2508.01285-B31B1B.svg)](https://arxiv.org/pdf/2508.01285) - *Ke et al. (2025.08)*
* **CRISPR-GPT for Agentic Automation of Gene-editing Experiments** [![arXiv](https://img.shields.io/badge/arXiv-2404.18021-B31B1B.svg)](https://arxiv.org/pdf/2404.18021) - *Qu et al. (2024.04)*
* **BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments** [![arXiv](https://img.shields.io/badge/arXiv-2405.17631-B31B1B.svg)](https://arxiv.org/pdf/2405.17631) - *Roohani et al. (2024.05)*
* **CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis** [![arXiv](https://img.shields.io/badge/arXiv-2407.09811-B31B1B.svg)](https://arxiv.org/pdf/2407.09811) - *Xiao et al. (2024.07)*
* **Training a Scientific Reasoning Model for Chemistry (ether0)** [![arXiv](https://img.shields.io/badge/arXiv-2506.17238-B31B1B.svg)](https://arxiv.org/pdf/2506.17238) - *Narayanan et al. (2025.06)* — FutureHouse
* **AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation** [![arXiv](https://img.shields.io/badge/arXiv-2509.25651-B31B1B.svg)](https://arxiv.org/pdf/2509.25651) - *Panapitiya et al. (2025.09)*
* **LLMatDesign: Autonomous Materials Discovery with Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2406.13163-B31B1B.svg)](https://arxiv.org/pdf/2406.13163) - *Jia et al. (2024.06)*
* **Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists** [![arXiv](https://img.shields.io/badge/arXiv-2506.05616-B31B1B.svg)](https://arxiv.org/pdf/2506.05616) - *Zhou et al. (2025.06)*
* **SparksMatter: Autonomous Inorganic Materials Discovery via Multi-Agent Physics-Aware Scientific Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2508.02956-B31B1B.svg)](https://arxiv.org/pdf/2508.02956) - *Ghafarollahi et al. (2025.08)*
* **Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation** [![arXiv](https://img.shields.io/badge/arXiv-2511.22311-B31B1B.svg)](https://arxiv.org/pdf/2511.22311) - *Wang et al. (2025.11)*
* **BixBench: A Comprehensive Benchmark for LLM-based Agents in Computational Biology** [![arXiv](https://img.shields.io/badge/arXiv-2503.00096-B31B1B.svg)](https://arxiv.org/pdf/2503.00096) - *Mitchener et al. (2025.02)* — FutureHouse
* **CASSIA: a multi-agent large language model for automated and interpretable cell annotation** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41467--025--67084--x-blue.svg)](https://www.nature.com/articles/s41467-025-67084-x) - *Xie et al. (2025.12)*
* **AutoZyme: An Autonomous Agentic Framework to Optimize Bioinformatics Software** [![bioRxiv](https://img.shields.io/badge/bioRxiv-2026.06-b31b1b.svg)](https://www.biorxiv.org/content/10.64898/2026.06.12.731250v1) - *Xie et al. (2026.06)*
### General Research
@@ -223,25 +296,70 @@ Benchmarks and frameworks evaluating diverse tasks from different stages of scie
### Survey Generation
* **AutoSurvey: Large Language Models Can Automatically Write Surveys** [![arXiv](https://img.shields.io/badge/arXiv-2406.10252-B31B1B.svg)](https://arxiv.org/pdf/2406.10252) - *Wang et al. (2024.06)*
* **SurveyX: Academic Survey Automation via Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2502.14776-B31B1B.svg)](https://arxiv.org/pdf/2502.14776) - *Liang et al. (2025.02)*
---
## Level 3: LLM as Scientist
LLM-based systems operating as active agents capable of orchestrating and navigating multiple stages of the scientific discovery process with considerable independence, often culminating in draft research papers.
LLM-based systems operating as active agents capable of orchestrating and navigating multiple stages of the scientific discovery process with considerable independence, often culminating in draft research papers or genuine new findings. As the field has matured, these systems increasingly fall into distinct classes, reflected in the sub-sections below.
### General-Purpose Autonomous Research Agents
End-to-end pipelines that autonomously move from ideation through experimentation to a full paper draft, typically domain-agnostic (frequently demonstrated on ML/AI research).
* **Agent Laboratory: Using LLM Agents as Research Assistants** [![arXiv](https://img.shields.io/badge/arXiv-2501.04227-B31B1B.svg)](https://arxiv.org/pdf/2501.04227) - *Schmidgall et al. (2025.01)*
* **The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2408.06292-B31B1B.svg)](https://arxiv.org/pdf/2408.06292) - *Lu et al. (2024.08)*
* **The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search** [![arXiv](https://img.shields.io/badge/arXiv-2504.08066-B31B1B.svg)](https://arxiv.org/pdf/2504.08066) - *Yamada et al. (2025.04)*
* **AI-Researcher: Fully-Automated Scientific Discovery with LLM Agents** [![GitHub](https://img.shields.io/badge/GitHub-HKUDS/AI--Researcher-blue.svg)](https://github.com/HKUDS/AI-Researcher) - *Data Intelligence Lab (2025.03)*
* **The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2408.06292-B31B1B.svg)](https://arxiv.org/pdf/2408.06292) - *Lu et al. (2024.08)* — Sakana AI
* **The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search** [![arXiv](https://img.shields.io/badge/arXiv-2504.08066-B31B1B.svg)](https://arxiv.org/pdf/2504.08066) - *Yamada et al. (2025.04)* — Sakana AI
* **AI-Researcher: Autonomous Scientific Innovation** [![arXiv](https://img.shields.io/badge/arXiv-2505.18705-B31B1B.svg)](https://arxiv.org/pdf/2505.18705) [![GitHub](https://img.shields.io/badge/GitHub-HKUDS/AI--Researcher-blue.svg)](https://github.com/HKUDS/AI-Researcher) - *Tang et al. (2025.05)*
* **Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback** [![arXiv](https://img.shields.io/badge/arXiv-2501.03916-B31B1B.svg)](https://arxiv.org/pdf/2501.03916) - *Yuan et al. (2025.01)*
* **NovelSeek / InternAgent: When Agent Becomes the Scientist — Building a Closed-Loop System from Hypothesis to Verification** [![arXiv](https://img.shields.io/badge/arXiv-2505.16938-B31B1B.svg)](https://arxiv.org/pdf/2505.16938) - *InternAgent Team (2025.05)* — Shanghai AI Lab
* **The Denario Project: Deep Knowledge AI Agents for Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2510.26887-B31B1B.svg)](https://arxiv.org/pdf/2510.26887) - *Villaescusa-Navarro et al. (2025.10)*
* **Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation (freephdlabor)** [![arXiv](https://img.shields.io/badge/arXiv-2510.15624-B31B1B.svg)](https://arxiv.org/pdf/2510.15624) - *Li et al. (2025.10)*
* **AIGS: Generating Science from AI-Powered Automated Falsification** [![arXiv](https://img.shields.io/badge/arXiv-2411.11910-B31B1B.svg)](https://arxiv.org/pdf/2411.11910) - *Liu et al. (2024.11)*
* **Zochi Technical Report** [![Link](https://img.shields.io/badge/Link-Intology.AI-blue.svg)](https://www.intology.ai/blog/zochi-tech-report) - *Intology AI (2025.03)*
* **Meet Carl: The First AI System To Produce Academically Peer-Reviewed Research** [![Link](https://img.shields.io/badge/Link-AutoScience.AI-blue.svg)](https://www.autoscience.ai/blog/meet-carl-the-first-ai-system-to-produce-academically-peer-reviewed-research) - *Autoscience Institute (2025.03)*
* **DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively** [![arXiv](https://img.shields.io/badge/arXiv-2509.26603-B31B1B.svg)](https://arxiv.org/pdf/2509.26603) - *Weng et al. (2025.09)*
* **Accelerating Social Science Research via Agentic Hypothesization and Experimentation** [![arXiv](https://img.shields.io/badge/arXiv-2602.07983-B31B1B.svg)](https://arxiv.org/pdf/2602.07983) - *Gupta et al. (2026.02)*
* **AI-Researcher: Fully-Automated Scientific Discovery with LLM Agents** [![GitHub](https://img.shields.io/badge/GitHub-HKUDS/AI--Researcher-blue.svg)](https://github.com/HKUDS/AI-Researcher) - *Data Intelligence Lab (2025.03)*
### Discovery-Oriented Scientific Systems
Systems whose primary goal is genuine new scientific knowledge — novel, experimentally- or mathematically-validated findings — rather than paper drafts. Many are frontier industry-lab systems.
* **Kosmos: An AI Scientist for Autonomous Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2511.02824-B31B1B.svg)](https://arxiv.org/pdf/2511.02824) - *Mitchener et al. (2025.11)* — Edison Scientific / FutureHouse
* **Robin: A Multi-Agent System for Automating Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2505.13400-B31B1B.svg)](https://arxiv.org/pdf/2505.13400) - *Ghareeb et al. (2025.05)* — FutureHouse
* **AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2506.13131-B31B1B.svg)](https://arxiv.org/pdf/2506.13131) - *Novikov et al. (2025.06)* — Google DeepMind
* **Aviary: Training Language Agents on Challenging Scientific Tasks** [![arXiv](https://img.shields.io/badge/arXiv-2412.21154-B31B1B.svg)](https://arxiv.org/pdf/2412.21154) - *Narayanan et al. (2024.12)* — FutureHouse
### Autonomous Research Ecosystems and Infrastructure
Platforms and protocols that enable multiple AI scientists to collaborate, share, review, and publish — moving beyond a single agent toward a research ecosystem.
* **AgentRxiv: Towards Collaborative Autonomous Research** [![arXiv](https://img.shields.io/badge/arXiv-2503.18102-B31B1B.svg)](https://arxiv.org/pdf/2503.18102) - *Schmidgall et al. (2025.03)*
* **aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists** [![arXiv](https://img.shields.io/badge/arXiv-2508.15126-B31B1B.svg)](https://arxiv.org/pdf/2508.15126) - *Zhang et al. (2025.08)*
---
## Frontier Labs and Foundation Models for Science
Flagship "AI for Science" systems from frontier industry labs. Unlike the agentic systems catalogued above, most of these are large domain-specific foundation models or specialized reasoning systems that have driven headline scientific results (structure prediction, materials/genome design, olympiad-level mathematics). They are included here as essential context for the broader landscape of AI-accelerated discovery.
* **Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--024--07487--w-blue.svg)](https://www.nature.com/articles/s41586-024-07487-w) - *Abramson et al. (2024.05)* — Google DeepMind / Isomorphic Labs
* **Scaling Deep Learning for Materials Discovery (GNoME)** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--023--06735--9-blue.svg)](https://www.nature.com/articles/s41586-023-06735-9) - *Merchant et al. (2023.11)* — Google DeepMind
* **A Generative Model for Inorganic Materials Design (MatterGen)** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--025--08628--5-blue.svg)](https://www.nature.com/articles/s41586-025-08628-5) - *Zeni et al. (2025.01)* — Microsoft Research
* **Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models** [![arXiv](https://img.shields.io/badge/arXiv-2410.12771-B31B1B.svg)](https://arxiv.org/pdf/2410.12771) - *Barroso-Luque et al. (2024.10)* — Meta FAIR
* **TamGen: Drug Design with Target-Aware Molecule Generation through a Chemical Language Model** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41467--024--53632--4-blue.svg)](https://www.nature.com/articles/s41467-024-53632-4) - *Wu et al. (2024.10)* — Microsoft Research
* **Genome Modeling and Design Across All Domains of Life with Evo 2** [![DOI](https://img.shields.io/badge/DOI-10.1101/2025.02.18.638918-blue.svg)](https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1) - *Brixi et al. (2025.02)* — Arc Institute / NVIDIA
* **Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning (AlphaProof)** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--025--09833--y-blue.svg)](https://www.nature.com/articles/s41586-025-09833-y) - *Hubert et al. (2025.11)* — Google DeepMind
* **Gold-Medalist Performance in Solving Olympiad Geometry with AlphaGeometry 2** [![arXiv](https://img.shields.io/badge/arXiv-2502.03544-B31B1B.svg)](https://arxiv.org/pdf/2502.03544) - *Chervonyi et al. (2025.02)* — Google DeepMind
* **Chai-2: Drug-Like Antibody Design Against Challenging Targets with Atomic Precision** [![Link](https://img.shields.io/badge/Link-Chai_Report-blue.svg)](https://chaiassets.com/chai-2/paper/technical_report_challenging_targets.pdf) - *Chai Discovery (2025.11)*
---
## Other Related Works
* **aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists** [![arXiv](https://img.shields.io/badge/arXiv-2508.15126-B31B1B.svg)](https://arxiv.org/pdf/2508.15126) - *Zhang et al. (2025.08)*
* **NVIDIA BioNeMo Agent Toolkit — Tools for Agents to Accelerate Scientific Discovery** [![Link](https://img.shields.io/badge/Link-NVIDIA_Newsroom-blue.svg)](https://nvidianews.nvidia.com/news/nvidia-launches-bionemo-agent-toolkit-giving-ai-agents-the-tools-to-accelerate-scientific-discovery) - *NVIDIA (2025)*
---
@@ -249,6 +367,11 @@ LLM-based systems operating as active agents capable of orchestrating and naviga
Contributions are welcome! If you have a paper, tool, or resource that fits into this taxonomy, please submit a **pull request**.
When adding an entry, please:
* Place it under the most appropriate level/sub-section.
* Keep the format consistent: `**Title** [badge](link) - *First author et al. (YYYY.MM)*`.
* Verify the arXiv ID / DOI resolves, and add the lab/affiliation when the work comes from an industry group.
---
## Citation
@@ -264,3 +387,4 @@ Please cite our paper if you found our survey helpful:
primaryClass={cs.CL},
url={https://arxiv.org/abs/2505.13259},
}
```
@@ -2,9 +2,9 @@
title: "Readme"
task: ""
lineage_type: import
upstream_source: https://github.com/ai-boost/awesome-ai-for-science/blob/3a7505d8/README.md
upstream_sha: 3a7505d8
imported_at: 2026-06-27
upstream_source: https://github.com/ai-boost/awesome-ai-for-science/blob/086e63bb/README.md
upstream_sha: 086e63bb
imported_at: 2026-07-19
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -81,6 +81,7 @@ validated: false
- [CORE](https://core.ac.uk/) - Aggregator of open access research papers
- [Connected Papers](https://www.connectedpapers.com/) - AI-powered visual graph for exploring academic papers and discovering connected research through citation networks and semantic similarity
- [PaSa (ByteDance)](https://github.com/bytedance/pasa) - Advanced paper search agent powered by large language models, autonomously invoking search tools, reading papers, and selecting references to deliver comprehensive and accurate results for complex scholarly queries (1.5K+ stars, Apache 2.0, 2024)
- [paper-search-mcp](https://github.com/openags/paper-search-mcp) - MCP server, CLI, and agent skills for searching and downloading academic papers from multiple open sources (arXiv, PubMed, bioRxiv, Semantic Scholar, OpenAlex, CORE, Europe PMC, etc.) with unified, deduplicated, LLM-friendly retrieval and an OA-first download fallback chain (OpenAGS, 1.9K+ stars, MIT License, 2025)
### Data Analysis & Visualization
- [PandasAI](https://github.com/Sinaptik-AI/pandas-ai) - Conversational data analysis using natural language
@@ -97,6 +98,8 @@ validated: false
- [GDM Science Skills](https://github.com/google-deepmind/science-skills) - Google DeepMind's official collection of agentic science skills accelerating scientific workflows with better grounding and higher token efficiency, integrating insights from AlphaGenome, AFDB, UniProt and 30+ other databases and tools (2026)
- [Scientific Agent Skills](https://github.com/K-Dense-AI/scientific-agent-skills) - Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science with 140+ ready-to-use skills and 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Antigravity, and the open Agent Skills standard (K-Dense-AI, 26K+ stars, 2025)
- [SciAgent-Skills](https://github.com/jaechang-hits/SciAgent-Skills) - 197 bioinformatics and life science skills for Claude Code and AI agents, achieving 92.0% accuracy on BixBench. Covers RNA-seq, single-cell analysis, drug discovery, proteomics, and more. Powers OmicsHorizon (195+ stars, 2026)
- [Medical Research Skills](https://github.com/aipoch/medical-research-skills) - Curated library of 550+ medical research agent skills spanning evidence insights, protocol design, omics/clinical data analysis, and academic writing; each skill is reviewed through MedSkillAudit and compatible with Claude Code, Codex, Open Code, OpenClaw, and SKILL.md-compatible agents (AIPOCH, 1.2K+ stars, MIT License, 2026)
- [bioSkills](https://github.com/GPTomics/bioSkills) - Collection of SKILLS.md guiding AI coding agents (Claude Code, OpenAI Codex, Google Gemini, OpenCode, OpenClaw) through common bioinformatics workflows from basic sequence manipulation to advanced analyses such as single-cell RNA-seq and population genetics; evaluated on the Bio-Task Bench dataset (GPTomics, 969+ stars, MIT License, 2026)
---
@@ -164,6 +167,7 @@ validated: false
### High-Performance Document Processing
- [MinerU (2024/2025)](https://github.com/opendatalab/MinerU) - SOTA multimodal document parsing with 1.2B parameters outperforming GPT-4o, converts PDFs to LLM-ready Markdown/JSON
- [MinerU-Diffusion (OpenDataLab, ECCV 2026)](https://github.com/opendatalab/MinerU-Diffusion) - Diffusion-based document OCR framework replacing autoregressive decoding with block-level parallel diffusion decoding, enabling high-accuracy text recognition in scientific PDFs (613+ stars, MIT License)
- [OpenDataLoader PDF (OpenDataLoader, 2025)](https://github.com/opendataloader-project/opendataloader-pdf) - Open-source PDF parser for AI-ready data, converting PDFs into Markdown/JSON/HTML/Tagged PDF with layout analysis and reading-order detection; ranks #1 overall on extraction benchmarks with deterministic bounding boxes and hybrid AI mode (26K+ stars, Apache 2.0)
- [PDF-Extract-Kit (2024)](https://github.com/opendatalab/PDF-Extract-Kit) - Comprehensive toolkit for high-quality PDF content extraction with layout detection, formula recognition, and OCR
- [Docling (IBM, AAAI 2025)](https://research.ibm.com/publications/docling-an-efficient-open-source-toolkit-for-ai-driven-document-conversion) - Multi-format (PDF/DOCX/PPTX/HTML/Images) → structured data (Markdown/JSON) with layout reconstruction, table/formula recovery
- [Nougat (Meta AI)](https://github.com/facebookresearch/nougat) - Neural optical understanding for academic documents, transforms scientific PDFs to Markdown with mathematical formula support
@@ -204,6 +208,14 @@ validated: false
- [OpenBioMed](https://github.com/PharMolix/OpenBioMed) - Open-source biomedical AI platform integrating multimodal foundation models (BioMedGPT, PharmolixFM, LangCell) with agentic workflows and 45+ Claude Code skills for drug discovery, protein engineering, and single-cell omics analysis (PharMolix & Tsinghua AIR, 1K+ stars, 2023-2026)
- [AutoR](https://github.com/AutoX-AI-Labs/AutoR) - Human-centered research OS with terminal-first harness and local browser Studio, turning research work into reproducible artifact-backed runs through a 9-stage workflow with human approval gates, resume/rollback controls, and venue-aware manuscript packaging (1K+ stars, 2026)
- [ScholarAIO](https://github.com/ZimoLiao/scholaraio) - Agent-agnostic research infrastructure providing AI agents with a structured scientific workspace for deep PDF parsing, hybrid semantic/keyword literature search, citation-graph analysis, topic discovery, and academic writing workflows; natively integrates with Claude Code, Codex, Cursor, Cline, and AgentSkills.io (530+ stars, MIT License, 2026)
- [BioMCP](https://github.com/genomoncology/biomcp) - Biomedical Model Context Protocol (MCP) server unifying literature search across PubMed/Europe PMC, entity pivoting across genes/variants/drugs/diseases/pathways/proteins, local study analytics, and Claude Code/Codex integration for agentic biomedical research (531+ stars, MIT License, 2025-2026)
- [MATLAB Agentic Toolkit](https://github.com/matlab/matlab-agentic-toolkit) - Official MathWorks toolkit connecting AI agents to MATLAB via the MATLAB MCP Server and curated skills, enabling trusted engineering and scientific computing workflows with idiomatic code generation, testing, and error diagnosis in Claude Code, GitHub Copilot, OpenAI Codex, and Gemini CLI (686+ stars, BSD-3-Clause, 2026)
- [BioNeMo Agent Toolkit (NVIDIA)](https://github.com/NVIDIA-BioNeMo/bionemo-agent-toolkit) - Turn any AI agent into a life science expert with NVIDIA BioNeMo skills, enabling agentic workflows for drug discovery, protein engineering, and biomolecular design (329+ stars, Apache 2.0 / CC-BY-4.0, 2026)
- [open-science](https://github.com/ai4s-research/open-science) - Local-first, open-source AI workbench for scientists — an open alternative to Claude Science (by ai4s-research, maintainers of this list; TypeScript, MIT, 2026)
- [OpenScience (Synthetic Sciences)](https://github.com/synthetic-sciences/openscience) - Open-source AI workbench for scientific research that automates the full research loop — literature review, hypothesis generation, code writing, experiment execution, database querying, and report writing — with 290+ skills, specialized research agents, and a browser-based workspace (1453+ stars, Apache 2.0, 2026)
- [Claude Scholar](https://github.com/Galaxy-Dawn/claude-scholar) - Semi-automated research assistant for academic research and software development, supporting Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication (Galaxy-Dawn, 4.5K+ stars, MIT License, 2026)
- [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok) - Free, open-source desktop AI research assistant that runs locally and turns natural-language requests into real data analysis, literature search, figure generation, and manuscript review; ships with 149 scientific skills, 326 workflow templates, and 229 databases across genomics, proteomics, drug discovery, and materials science, plus a living lab notebook, 60+ scientific file previews, and LaTeX editing (K-Dense-AI, 908+ stars, MIT License, 2026)
- [Academic Research Skills (ARS)](https://github.com/Imbad0202/academic-research-skills) - Comprehensive Claude Code skill suite covering the full academic pipeline from deep research and paper writing to multi-perspective peer review, revision, and finalization; features multi-agent teams, PRISMA systematic review, style calibration, claim-level citation audits, integrity gates, and human-in-the-loop safeguards (38K+ stars, CC BY-NC 4.0, 2026)
### Literature Management Plugins
- [llm-for-zotero](https://github.com/yilewang/llm-for-zotero) - Research agent system deeply integrated with Zotero supporting Agent Mode, skills, multi-model backends (OpenAI-compatible, Claude Code, WebChat, Codex), and MinerU PDF parsing for literature Q&A, summarization, figure inspection, and source comparison (1.3K+ stars, 2026)
@@ -238,8 +250,11 @@ validated: false
### Autonomous Research Systems (2023-2025 Breakthroughs)
- [FunSearch (DeepMind, Nature 2023)](https://github.com/google-deepmind/funsearch) - First system to make novel, verifiable scientific discoveries by pairing LLMs with evolutionary search, solving open problems in combinatorics (cap set problem) and discovering faster matrix multiplication algorithms
- [OpenEvolve](https://github.com/algorithmicsuperintelligence/openevolve) - Open-source implementation of AlphaEvolve's evolutionary coding agent paradigm, enabling LLMs to autonomously discover and optimize algorithms through iterative evolution, matching the approach behind DeepMind's breakthrough matrix multiplication discovery (6.2K+ stars, 2025)
- [SkyDiscover](https://github.com/skydiscover-ai/skydiscover) - Modular framework for AI-driven scientific and algorithmic discovery, providing a unified interface for implementing, running, and fairly comparing discovery algorithms across 200+ optimization tasks; introduces AdaEvolve and EvoX adaptive/evolutionary algorithms and natively supports OpenEvolve, GEPA, and Harbor-format benchmarks (skydiscover-ai, 568+ stars, Apache 2.0, 2026)
- [EvoMaster (SJTU SAI, arXiv 2026)](https://github.com/sjtu-sai-agents/EvoMaster) - Foundational auto-research agent framework for agentic science at scale, providing modular agent construction, run-level self-evolution, and multiple SciMaster domain agents (ML-Master, X-Master, Browse-Master); outperforms general-purpose agents across authoritative benchmarks including the OpenAI Frontier Science Benchmark (206+ stars, Apache 2.0, 2026)
- [Virtual Lab (Stanford Zou Group, Nature 2025)](https://github.com/zou-group/virtual-lab) - AI-human collaborative research platform where a human researcher works with a team of LLM agents via team and individual meetings to perform scientific research; demonstrated by designing new SARS-CoV-2 nanobodies with wet-lab validation
- [The AI Scientist (SakanaAI)](https://github.com/SakanaAI/AI-Scientist) - First fully autonomous open-ended scientific discovery system with official implementation: hypothesis→experiment→writing→review simulation (13.8K+ stars, 2024)
- [The AI Scientist v2 (SakanaAI)](https://github.com/SakanaAI/AI-Scientist-v2) - Official implementation of the second-generation fully autonomous scientific discovery system, extending the original with agentic tree search and reduced template dependency to achieve workshop-level accepted papers (6.7K+ stars, 2025)
- [The AI Scientist v1 (2024)](https://arxiv.org/abs/2408.06292) - First fully autonomous research system: hypothesis→experiment→writing→review simulation
- [The AI Scientist v2 (2025)](https://arxiv.org/abs/2504.08066) - Enhanced with Agentic Tree Search, reduced template dependency, first workshop-level accepted paper
- [DeepScientist](https://github.com/ResearAI/DeepScientist) - First system progressively surpassing human SOTA on frontier AI tasks (183.7%, 1.9%, 7.9% improvements), month-long autonomous discovery with 20,000+ GPU hours
@@ -247,7 +262,11 @@ validated: false
- [Kosmos](https://github.com/jimmc414/Kosmos) - Extended autonomy AI scientist with 200 parallel agent rollouts, 42K lines of code execution, 1.5K papers analyzed per run, achieving 79.4% accuracy and 7 scientific discoveries (Edison Scientific)
- [AlphaResearch](https://github.com/answers111/alpha-research) - Autonomous algorithm discovery combining evolutionary search with peer-review reward models, achieving best-known performance on circle packing problems
- [AutoResearchClaw](https://github.com/aiming-lab/AutoResearchClaw) - Fully autonomous research from idea to paper with multi-agent debate, citation verification, and OpenClaw integration (11K+ stars, 2026)
- [ARIS (Auto-Research-In-Sleep)](https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep) - Lightweight Markdown-only skills for autonomous ML research with cross-model review loops, idea discovery, and experiment automation; no framework lock-in, works with Claude Code, Codex, OpenClaw, or any LLM agent (12.8K+ stars, MIT License, 2026)
- [Arbor](https://github.com/RUC-NLPIR/Arbor) - Generalist autonomous research agent that grows a hypothesis tree to optimize any measurable task, beating Claude Code and Codex by 2.5× on the same compute budget across BrowseComp, Terminal-Bench 2.0, math reasoning, and MLE-Bench Lite; supports native CLI, keyless Claude Code/Codex integration, and an MCP tool server (RUC-NLPIR, 866+ stars, Apache 2.0, 2026)
- [NanoResearch](https://github.com/OpenRaiser/NanoResearch) - End-to-end autonomous AI research engine that turns an idea into a complete LaTeX paper by dispatching real computational experiments to local GPUs or SLURM clusters, collecting actual results, generating figures/tables, and writing a data-grounded manuscript rather than LLM hallucinations (OpenRaiser, 1.5K+ stars, MIT License, 2026)
- [ScienceClaw](https://github.com/beita6969/ScienceClaw) - Self-evolving AI research colleague built on OpenClaw with 285+ runtime-adaptive skills across 28+ disciplines, persistent cross-session research memory, and zero-hallucination citation protocols; agent autonomously writes new SKILL.md files based on research patterns without redeployment (828+ stars, MIT License, 2026)
- [ai4s-skills](https://github.com/ai4s-research/ai4s-skills) - Agent skills (SKILL.md + deterministic tools) for the AI4S workflow — topic exploration, literature survey, runnable experiments, publication-grade papers, and integrity audit, with every citation and number traceable to its source (by ai4s-research, maintainers of this list; MIT, 2026)
- [Denario (AstroPilot-AI, Agents4Science 2025)](https://github.com/AstroPilot-AI/Denario) - Modular multi-agent scientific research assistant that automates idea generation, literature review, methodology design, code execution in Docker, visualization, LaTeX paper writing, and peer-review simulation across 10+ disciplines; winner of the NeurIPS 2025 Fair Universe Competition (573+ stars, GPL-3.0, 2025-2026)
- [AI-Researcher](https://github.com/HKUDS/AI-Researcher) - Autonomous pipeline from literature review→hypothesis→algorithm implementation→publication-level writing with Scientist-Bench evaluation
- [Agent Laboratory](https://agentlaboratory.github.io/) - Multi-agent workflows for complete research cycles with AgentRxiv for cumulative discovery
@@ -270,6 +289,9 @@ validated: false
- [CORAL (arXiv 2026)](https://github.com/Human-Agent-Society/CORAL) - Robust, lightweight infrastructure for multi-agent autonomous self-evolution, built for autoresearch; agents run in isolated git worktrees, share knowledge through a common state directory, and are scored by a grader daemon; natively integrated with Claude Code, Codex, Cursor Agent, OpenCode, and Kiro (672+ stars, Apache 2.0)
- [Science-Star (USTC AI4Science, 2025)](https://github.com/ustc-ai4science/Science-Star) - Open-source platform for building, extending, and experimenting with scientific agents, providing modular agent construction tools and standardized evaluation pipelines for accelerating autonomous scientific discovery research (748+ stars, MIT License)
- [SR-Scientist (ICLR 2026)](https://github.com/GAIR-NLP/SR-Scientist) - Scientific equation discovery with agentic AI, elevating LLMs from equation proposers to autonomous scientists that write code, analyze data, implement equations, and optimize based on experimental feedback; outperforms baselines by 6-35% across four science disciplines with robustness to noise and out-of-domain generalization (GAIR-NLP / SJTU, 49+ stars, Apache 2.0)
- [ARA (Agent-Native Research Artifact)](https://github.com/ARA-Labs/Agent-Native-Research-Artifact) - Research ecosystem for rigorous and trustworthy AI scientists — a protocol and skill bundle that makes autonomous research verifiable, crystallized, and observable through structured, machine-executable research artifacts and five agent skills for research management, compilation, verification, visualization, and publication (ARA-Labs, 447+ stars, MIT License, 2026)
- [Scholar Loop](https://github.com/renee-jia/scholar-loop) - Autonomous multi-agent AI scientist that mirrors a PhD workflow: literature review → grounded hypothesis → real ML experiments → self-critique → write-up; features a deterministic harness with frozen-metric scoring, edit allowlists, and a verified registry to make reward-hacking and hallucination impossible, plus 108 unit tests runnable without API keys or GPUs (461+ stars, MIT License, 2026)
- [ResearchStudio (Microsoft)](https://github.com/microsoft/ResearchStudio) - AI co-author covering the entire research lifecycle — from an under-specified research direction to a published paper; includes ResearchStudio-Idea for evidence-grounded research ideation and ResearchStudio-Reel for turning finished papers into posters, narrated videos, blogs, and interactive reels; runs as skills on Claude Code and Codex (1.2K+ stars, MIT License, 2026)
### Evaluation & Benchmarking
- [ScienceAgentBench (ICLR 2025)](https://github.com/OSU-NLP-Group/ScienceAgentBench) - 102 executable tasks from 44 peer-reviewed papers across 4 disciplines with containerized evaluation
@@ -457,7 +479,9 @@ validated: false
- [ModelAngelo](https://github.com/3dem/model-angelo) - Automatic atomic model building program for cryo-EM maps using deep learning, enabling rapid de novo protein structure determination from electron density with high accuracy (3DEM/EMBL, 169+ stars)
- [AlphaFold](https://github.com/google-deepmind/alphafold) - Protein structure prediction
- [AlphaFold3](https://github.com/google-deepmind/alphafold3) - AlphaFold 3 inference pipeline for unified biomolecular structure prediction of proteins, nucleic acids, small molecules, ions, and post-translational modifications (Google DeepMind, Nature 2024)
- [AlphaFold Server](https://alphafoldserver.com/) - Free, easy-to-use web platform by Google DeepMind and Isomorphic Labs for running AlphaFold 3 predictions of biomolecular structures and interactions, enabling researchers without local infrastructure to model proteins, nucleic acids, small molecules, ions, and post-translational modifications through a searchable proteome interface (2024)
- [AlphaProteo](https://github.com/google-deepmind/alphaproteo) - Deep learning system for de novo design of high-affinity protein binders, achieving strong binding across diverse target classes including challenging intracellular proteins with significantly higher success rates than traditional wet-lab screening methods (Google DeepMind, Nature 2024)
- [AlphaPulldown](https://github.com/KosinskiLab/AlphaPulldown) - Automated pipeline for proteome-scale protein-protein interaction screening with AlphaFold-Multimer and AlphaFold 3, supporting flexible inputs (UniProt IDs, FASTA, residue regions, multimers, AF3 JSON features) and integrated downstream analysis for hit prioritization (Kosinski Lab, EMBL, Nature Protocols 2024, 317+ stars, GPL-3.0)
- [RareFold](https://github.com/PatrickBryant1/RareFold) - Structure prediction and design of proteins with noncanonical amino acids, enabling AI-powered modeling of synthetic biology constructs and expanded genetic code systems (133+ stars, 2025)
- [ColabFold (2025 Updates)](https://github.com/sokrypton/ColabFold) - AlphaFold/ESMFold accessible implementation with AF3 JSON export, database updates
- [OpenFold](https://github.com/aqlaboratory/openfold) - Trainable, memory-efficient PyTorch reproduction and retraining of AlphaFold2 providing new insights into its learning dynamics and out-of-distribution generalization; widely used as the open-source AlphaFold2 backbone underpinning many downstream protein structure prediction and design pipelines (Columbia AlQuraishi Lab & OpenFold Consortium, Nature Methods 2024)
@@ -475,6 +499,7 @@ validated: false
- [Proteina-Complexa](https://github.com/NVIDIA-Digital-Bio/Proteina-Complexa) - Flow-based generative model for atomistic protein binder design with test-time optimization, SOTA on binder benchmarks (ICLR 2026 Oral, NVIDIA)
- [PXDesign (ByteDance, 2025)](https://github.com/bytedance/PXDesign) - Fast, modular, and accurate de novo design of protein binders based on the Protenix foundation model, achieving 17-82% nanomolar hit rates across diverse targets with 2-6× improvement over prior methods like AlphaProteo and RFdiffusion (229+ stars, Apache 2.0)
- [ODesign (OTeam-AI4S, 2025)](https://github.com/OTeam-AI4S/ODesign) - All-atom generative world model for all-to-all biomolecular interaction design, enabling cross-modality generation of proteins, nucleic acids, small molecules, and cyclic peptides with fine-grained epitope-level control and 2-4 orders of magnitude faster design throughput than modality-specific baselines (316+ stars, Apache 2.0)
- [OpenDDE (Aureka Research, 2026)](https://github.com/aurekaresearch/OpenDDE) - Open-source, all-atom biomolecular foundation model that turns co-folding into a scalable engine for structure prediction, design, and optimization across proteins, nucleic acids, and small molecules in drug discovery; ranked first on PXMeter-AB, FoldBench-AB, and 2026ARK-AB antibody-antigen benchmarks (263+ stars, Apache 2.0)
- [La-Proteina (NVIDIA)](https://github.com/NVIDIA-Digital-Bio/la-proteina) - Partially latent flow matching model for the joint generation of a protein's amino acid sequence and full atomistic structure, including both backbone and side chains (2025)
- [Proteina (NVIDIA, ICLR 2025 Oral)](https://github.com/NVIDIA-Digital-Bio/proteina) - Large-scale flow-based protein backbone generator utilizing hierarchical fold class labels for conditioning with a tailored scalable transformer architecture, enabling controllable de novo protein design (264+ stars)
- [xfold](https://github.com/Shenggan/xfold) - Democratizing AlphaFold3: PyTorch reimplementation to accelerate protein structure prediction research
@@ -500,6 +525,7 @@ validated: false
- [Chroma](https://github.com/generatebio/chroma) - Generative model for programmable protein design using diffusion modeling, equivariant graph neural networks, and conditional random fields to efficiently sample diverse all-atom structures; supports conditional generation via composable conditioners for substructure, symmetry, shape, and neural-network predictions; validated crystallographically (Generate Biomedicines, Nature 2023)
- [EvoDiff](https://github.com/microsoft/evodiff) - Discrete diffusion framework for generative protein sequence design over evolutionary-scale databases, supporting unconditional generation, evolutionary-guided conditional design, motif scaffolding, and intrinsically disordered region generation through order-agnostic autoregressive diffusion, enabling sequence-only protein design without structural priors (Microsoft Research, Nature Communications 2024)
- [DISCO](https://github.com/DISCO-design/DISCO) - General multimodal protein design framework enabling DNA-encoding of chemistry for programmable enzyme design and diverse protein generation through diffusion-based generative modeling (190+ stars, Apache 2.0, 2026)
- [SwitchCraft](https://github.com/bjing2016/switchcraft) - Programmatic framework for designing state-switching proteins via backpropagation through compositional design constraints parameterized by structure prediction models; enables de novo design of allosteric regulators and fluorescent biosensors for arbitrary small-molecule analytes (79+ stars, MIT License, ICML 2026)
- [RFdiffusion3](https://github.com/RosettaCommons/RFdiffusion) - Latest RFdiffusion for protein structure design with 10× speedup and atom-level precision (December 2025)
- [RFantibody](https://github.com/RosettaCommons/RFantibody) - Structure-based de novo antibody design pipeline built on RFdiffusion for computational generation of target-specific antibodies (RosettaCommons, 2025)
- [IgGM](https://github.com/TencentAI4S/IgGM) - Generative foundation model for functional antibody and nanobody design, supporting de novo generation, affinity maturation, inverse design, structure prediction, and humanization (Tencent AI4S, ICLR 2025)
@@ -515,6 +541,7 @@ validated: false
- [DeepMol](https://github.com/BioSystemsUM/DeepMol) - Unified ML/DL framework for drug discovery workflows, integrating RDKit, DeepChem, and scikit-learn with SHAP explainability
- [Chemprop](https://github.com/chemprop/chemprop) - Message passing neural networks for molecule property prediction, ADMET modeling, and reaction prediction, achieving SOTA on MoleculeNet and widely used in pharmaceutical drug discovery (MIT, 2.3K+ stars)
- [RDKit](https://github.com/rdkit/rdkit) - Cheminformatics toolkit
- [nvMolKit (NVIDIA BioNeMo, 2025)](https://github.com/NVIDIA-BioNeMo/nvMolKit) - High-performance, GPU-accelerated library for key computational chemistry tasks including molecular similarity, conformer generation, and geometry relaxation, designed to accelerate drug-discovery and molecular-modeling workflows (264+ stars, Apache 2.0)
- [Open Targets](https://www.opentargets.org/) - Open-source data integration platform for systematic drug target identification and prioritization, combining genetics, genomics, chemistry, and pharmacology data from EMBL-EBI, Wellcome Sanger Institute, and pharmaceutical partners to accelerate therapeutic discovery
- [ESM3](https://github.com/evolutionaryscale/esm) - 98B-parameter frontier generative model jointly reasoning over protein sequence, structure, and function, trained on 2.78 billion proteins; generated a novel fluorescent protein (esmGFP) with only 58% sequence identity to known GFPs (EvolutionaryScale, 2024)
- [ProtTrans](https://github.com/agemagician/ProtTrans) - State-of-the-art pretrained language models for proteins trained on thousands of GPUs and Google TPUs using Transformer architectures, enabling protein property prediction, feature extraction, and transfer learning across diverse downstream tasks (1.3K+ stars, MIT, 2020-2026)
@@ -538,6 +565,7 @@ validated: false
- [gRNAde](https://github.com/chaitjo/geometric-rna-design) - Generative AI framework for inverse design of 3D RNA structure and function using geometric deep learning, learning design rules from 3D structures to capture complex tertiary interactions (pseudoknots, non-canonical base pairs) with expert-level accuracy for designing functional RNAs including aptamers and ribozymes (bioRxiv 2025)
- [AIDO.ModelGenerator](https://github.com/genbio-ai/ModelGenerator) - GenBio AI's software stack for the AI-Driven Digital Organism, supporting adaptation and finetuning of multiscale biological foundation models across DNA, RNA, protein, structure, and single-cell tasks with reproducible CLIs and pretrained model zoo (2025)
- [Evo 2](https://github.com/ArcInstitute/evo2) - Arc Institute's 40B-parameter genome foundation model trained on 9 trillion nucleotides from all domains of life, supporting 1M base pair context for generalist DNA/RNA/protein prediction and design (Nature 2026)
- [Carbon (Hugging Face, 2026)](https://github.com/huggingface/carbon) - Family of causal genomic foundation models trained on 1T tokens (~6T DNA base pairs) from the Carbon Pretraining Corpus, combining eukaryote genes, mRNA transcripts, and prokaryote genomes with a hybrid text/6-mer tokenizer; Carbon-3B matches or beats Evo2-7B on zero-shot DNA evaluations including sequence recovery, variant effect prediction, and perturbations (Apache 2.0, 201+ stars)
- [Nucleotide Transformer](https://github.com/instadeepai/nucleotide-transformer) - Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)
- [HyenaDNA](https://github.com/HazyResearch/hyena-dna) - Long-range genomic foundation model using subquadratic Hyena operators instead of Transformer attention, enabling context lengths up to 1 million nucleotides for chromosome-scale DNA sequence modeling and downstream genomics tasks (Stanford Hazy Research, NeurIPS 2023, 784+ stars, Apache 2.0)
- [Caduceus (ICML 2024)](https://github.com/kuleshov-group/caduceus) - Bi-directional DNA language model based on the Mamba state space architecture, enabling efficient long-range genomic sequence modeling with linear-time complexity and built-in reverse-complement equivariance; achieves strong performance on chromatin accessibility, enhancer, and promoter prediction benchmarks (Stanford & UC Berkeley, 500+ stars)
@@ -576,12 +604,16 @@ validated: false
- [AlphaMissense](https://github.com/google-deepmind/alphamissense) - Google DeepMind's AlphaFold-derived classifier for proteome-wide missense variant effect prediction, providing pathogenicity scores for all ~71M possible human missense variants and classifying 89% with 90% precision; pre-computed predictions are integrated into Ensembl VEP and UCSC Genome Browser to support clinical variant interpretation (Science 2023)
- [AlphaGenome](https://github.com/google-deepmind/alphagenome) - Google DeepMind's unified DNA sequence foundation model predicting molecular consequences of genetic variants from single-base resolution up to 1 megabase context, jointly outputting thousands of regulatory tracks (RNA expression, splicing, chromatin accessibility, TF binding, contact maps) for human and mouse genomes via a Python client and non-commercial API (2025)
- [GPN-Star (Song Lab, UC Berkeley, bioRxiv 2025)](https://github.com/songlab-cal/gpn) - Phylogeny-aware genomic language model trained on whole-genome alignments across multiple evolutionary timescales, predicting functional constraints and variant effects for human, mouse, chicken, fly, worm, and Arabidopsis genomes (344+ stars, MIT License)
- [GENERanno (bioRxiv 2025)](https://github.com/GenerTeam/GENERanno) - Genomic foundation model for metagenomic and genome annotation, featuring an 8k base-pair context and 500M parameters trained on 386B base pairs of eukaryotic DNA; provides expert models and a unified CLI for prokaryotic/eukaryotic coding-sequence annotation with strong performance on Genomic Benchmarks, Nucleotide Transformer tasks, and custom Gener tasks (GenerTeam, 314+ stars, MIT License)
- [GENERator (bioRxiv 2026)](https://github.com/GenerTeam/GENERator) - Long-context generative genomic foundation model using 6-mer tokenization for DNA sequence modeling and generation, with v2 model families for prokaryote and eukaryote genomes and pretrained weights available on HuggingFace (GenerTeam, 460+ stars, MIT License, 2025-2026)
- [DeepVariant](https://github.com/google/deepvariant) - Google DeepMind's deep learning analysis pipeline for calling genetic variants (SNPs and indels) from next-generation DNA sequencing data, achieving human expert-level accuracy and widely adopted in clinical genomics, population genetics, and precision medicine; pre-trained models available for multiple sequencing platforms and organismal genomes (Nature Biotechnology 2018, 3.7K+ stars)
- [Casanovo](https://github.com/Noble-Lab/casanovo) - Transformer encoder-decoder for de novo peptide sequencing from tandem mass spectrometry, translating MS/MS spectra directly to peptide sequences without reference databases, enabling identification of novel peptides for immunopeptidomics, antibody repertoires, and metaproteomes (Noble Lab UW, Nature Communications 2024)
- [Dorado](https://github.com/nanoporetech/dorado) - Oxford Nanopore's official deep-learning basecaller for nanopore sequencing, converting raw electrical signals into DNA/RNA sequences with integrated modified-base (methylation) detection and efficient CPU/GPU inference; foundational tool for long-read genomics, epigenetics, and real-time sequencing analysis (nanoporetech, 846+ stars, actively maintained)
#### Neuroscience & Behavioral Analysis
- [DeepLabCut](https://github.com/DeepLabCut/DeepLabCut) - Markerless pose estimation of user-defined features with deep learning for all animals including humans, enabling quantitative behavioral analysis in neuroscience and ethology (Nature Neuroscience 2018, 5.6K+ stars)
- [SLEAP](https://github.com/talmolab/sleap) - Deep learning-based multi-animal pose tracking and behavior classification, enabling automated quantification of social interactions and collective behavior across species (Nature Methods 2022, 2.2K+ stars)
- [NeuroAI (Meta FAIR)](https://github.com/facebookresearch/neuroai) - Modular Python suite for Neuro-AI research across all modalities, providing efficient data loaders (NeuralSet), curated datasets (NeuralFetch), scalable training (NeuralTrain), and unified benchmarking (NeuralBench) for building and evaluating neuroscience foundation models (Meta FAIR, 270+ stars, MIT License, 2026)
- [CEBRA (Nature 2023)](https://github.com/AdaptiveMotorControlLab/CEBRA) - Learnable latent embeddings for joint behavioral and neural analysis, enabling consistent and interpretable mapping of neural activity to behavior across modalities, species, and experiments (EPFL & Harvard, 1K+ stars)
- [Kilosort (Nature Methods 2024)](https://github.com/MouseLand/Kilosort) - Fast spike sorting with drift correction for extracellular electrophysiology, enabling universal neural spike sorting via deep learning on high-density neural probe recordings (MouseLand, 609+ stars)
- [SpikeInterface](https://github.com/SpikeInterface/spikeinterface) - Unified Python framework for extracellular electrophysiology, standardizing interfaces to 10+ ML-based spike sorting algorithms including Kilosort for reproducible neural spike sorting workflows (792+ stars, actively maintained)
@@ -600,8 +632,11 @@ validated: false
- [PLIP (Nature Medicine 2023)](https://github.com/PathologyFoundation/plip) - First vision-and-language foundation model for pathology AI, fine-tuned from CLIP on 249K image-caption pairs, enabling open-ended visual-semantic search and zero-shot diagnosis across histopathology (Pathology Foundation, 376+ stars)
- [TITAN (Nature Medicine 2024)](https://github.com/mahmoodlab/TITAN) - Multimodal whole-slide pathology foundation model jointly pretrained on H&E histology and diagnostic text reports, enabling zero-shot cancer subtyping, biomarker prediction, and multimodal reasoning across diverse cancer types (Mahmood Lab, 341+ stars)
- [Virchow (Nature Medicine 2024)](https://huggingface.co/paige-ai/Virchow) - Self-supervised pathology foundation model (ViT-Huge, 632M parameters) pretrained via DINOv2 on 1.5M whole-slide images from Memorial Sloan Kettering across 17 cancer types, with Virchow2 follow-up scaling to 3.1M slides and mixed magnifications, achieving SOTA on biomarker prediction, mutation classification, and rare cancer detection (Paige AI & MSK)
- [H-Optimus (Bioptimus, Nature Medicine 2025)](https://huggingface.co/bioptimus/H-optimus-0) - Open-weights pathology foundation model family (H-Optimus-0: 1.1B-parameter ViT pretrained via DINOv2 on 500M+ diagnostic image tiles; H-Optimus-1 follow-up) for whole-slide image analysis, achieving strong zero-shot and fine-tuned transfer across biomarker prediction, cancer subtyping, and mutation classification (Bioptimus, Apache 2.0)
- [TRIDENT (2025)](https://github.com/mahmoodlab/TRIDENT) - Toolkit for large-scale whole-slide image processing supporting 22+ patch encoders (UNI, CONCH, Virchow, H-Optimus-0, etc.), slide encoders (TITAN, GigaPath, PRISM, CHIEF, Madeleine, Feather), tissue segmentation, and multi-GPU inference with end-to-end pipeline and smart resume for standardized deployment of computational pathology foundation models (Mahmood Lab, Harvard Medical School, 553+ stars)
- [Feather (Mahmood Lab, ICML 2025 Spotlight)](https://github.com/mahmoodlab/MIL-Lab) - Lightweight supervised slide foundation model with 0.9M parameters pretrained on 24K whole-slide images for pan-cancer morphological classification, achieving competitive performance with much larger self-supervised models (TITAN, GigaPath) while enabling finetuning on consumer-grade GPUs; includes standardized MIL implementations and benchmarking across 15+ classification tasks (Mahmood Lab, Harvard Medical School, 153+ stars)
- [PathChat (Nature Medicine 2024)](https://github.com/MahmoodLab/PathChat) - Multimodal generative AI assistant for computational pathology enabling interactive visual-language conversations over histopathology images for diagnostic reasoning, case discussion, and education, built on a Mistral-7B backbone with domain-specific fine-tuning (Mahmood Lab, Harvard Medical School, 1.2K+ stars)
- [SlideChat (CVPR 2025)](https://github.com/uni-medical/SlideChat) - First large vision-language assistant for gigapixel whole-slide pathology image understanding, released with the SlideInstruction dataset and SlideBench benchmark (uni-medical, Apache 2.0, 2025)
- [HEST (NeurIPS 2024)](https://github.com/mahmoodlab/HEST) - Dataset and benchmarking framework integrating histology and spatial transcriptomics, enabling multimodal analysis of whole-slide images with matched spatial gene expression for advancing computational pathology and tissue microenvironment research (Mahmood Lab, Harvard Medical School, 411+ stars)
#### Medical AI & Clinical Applications
@@ -613,12 +648,14 @@ validated: false
- [micro-sam](https://github.com/computational-cell-analytics/micro-sam) - Segment Anything Model for microscopy: interactive and automatic segmentation of light, electron, and fluorescence microscopy images in 2D and 3D, with domain-specific fine-tuning workflows for scientific imaging (1.5K+ stars)
- [MedSAM](https://github.com/bowang-lab/MedSAM) - Universal medical image segmentation foundation model trained on 1.57M image-mask pairs across 10 imaging modalities and 30+ cancer types (Nature Communications 2024)
- [MedSAM2](https://github.com/bowang-lab/MedSAM2) - Segment Anything in 3D medical images and videos, extending SAM2 to volumetric and temporal medical imaging with state-of-the-art zero-shot segmentation performance across CT, MRI, and surgical video (arXiv 2025)
- [Medical SAM3 (AIM Research Lab, arXiv 2026)](https://github.com/AIM-Research-Lab/Medical-SAM3) - Foundation model for universal prompt-driven medical image segmentation extending SAM3 to clinical imaging, supporting 2D public benchmarks and 3D training/evaluation with text and box prompts; pretrained weights available on HuggingFace (189+ stars)
- [MedSegX](https://github.com/MedSegX/MedSegX-code) - Generalist foundation model and database for open-world medical image segmentation, enabling universal segmentation of diverse anatomical structures and pathologies with zero-shot generalization to unseen tasks and modalities (Nature Biomedical Engineering 2025)
- [VoxTell (MIC-DKFZ, 2025)](https://github.com/MIC-DKFZ/VoxTell) - Free-text promptable universal 3D medical image segmentation foundation model enabling zero-shot segmentation of diverse anatomical structures and pathologies via natural language prompts across CT, MRI, and other volumetric imaging modalities (DKFZ, 195+ stars, Apache 2.0)
- [BiomedParse](https://github.com/microsoft/BiomedParse) - Foundation model for joint segmentation, detection, and recognition of biomedical objects across nine imaging modalities, with v2 introducing BoltzFormer architecture for end-to-end 3D inference (Microsoft, Nature Methods 2025)
- [UniBiomed (Nature Communications 2026)](https://github.com/Luffy03/UniBiomed) - Universal foundation model for grounded biomedical image interpretation, enabling comprehensive visual understanding, reasoning, and grounding across diverse biomedical imaging modalities with strong zero-shot generalization (55+ stars, Apache 2.0, 2025-2026)
- [MIRA (NeurIPS 2025)](https://github.com/microsoft/MIRA) - Medical time series foundation model pretrained on 454B time points from heterogeneous clinical corpora spanning ICU physiological signals and hospital EHR, with continuous-time rotary positional encoding, frequency-specialized Mixture-of-Experts, and neural ODE extrapolation for zero-shot forecasting across irregular and multimodal temporal health data (Microsoft, 399+ stars, MIT License)
- [HealthGPT (ICML 2025 Spotlight)](https://github.com/ZJU4HealthCare/HealthGPT) - Medical large vision-language model unifying comprehension and generation via heterogeneous knowledge adaptation, enabling holistic medical image understanding, visual question answering, and clinical report generation across diverse modalities (ZJU4HealthCare, 1.6K+ stars)
- [Merlin (Stanford MIMI, Nature 2026)](https://github.com/StanfordMIMI/Merlin) - 3D vision-language model for computed tomography that leverages both structured electronic health records (EHR) and unstructured radiology reports for pretraining, enabling multimodal medical understanding and radiology report generation (447+ stars, MIT License, 2026)
- [MedAgents](https://github.com/gersteinlab/MedAgents) - Multi-disciplinary collaboration framework for zero-shot medical reasoning using role-playing LLM agents (ACL 2024)
- [MedAgentGym](https://github.com/wshi83/MedAgentGym) - Scalable agentic training environment for code-centric reasoning in biomedical data science
- [MedRAX (ICML 2025)](https://github.com/bowang-lab/MedRAX) - First versatile medical reasoning agent for chest X-ray interpretation, dynamically integrating state-of-the-art CXR analysis tools and multimodal LLMs into a unified framework; introduces ChestAgentBench with 2,500 complex medical queries across 7 categories (bowang-lab, 1.1K+ stars)
@@ -763,6 +800,9 @@ validated: false
- [TimesFM (Google Research)](https://github.com/google-research/timesfm) - Pretrained time series foundation model for long-horizon forecasting across diverse scientific domains including climate variables, biomedical signals, and physical observations; decoder-only Transformer architecture with strong zero-shot generalization (19.8K+ stars, Apache 2.0, 2024-2025)
- [Chronos (Amazon Science, NeurIPS 2024)](https://github.com/amazon-science/chronos-forecasting) - Pretrained time series foundation model for zero-shot forecasting across diverse scientific and real-world domains; tokenizes continuous time series into discrete bins to train transformer language models on large-scale corpora, achieving strong zero-shot generalization and competitive performance with task-specific supervised models on climate, energy, and health benchmarks (5.3K+ stars, Apache 2.0, 2024-2026)
- [TabPFN (Prior Labs, Nature 2025)](https://github.com/PriorLabs/tabpfn) - Foundation model for tabular data that predicts on unseen real-world tables in a single forward pass, achieving accurate small-data classification and regression without task-specific training; widely applicable to scientific datasets with limited samples (7.4K+ stars, 2022-2026)
- [TabFM (Google Research, 2026)](https://github.com/google-research/tabfm) - Scikit-learn compatible tabular foundation model for zero-shot classification and regression on mixed-type tabular datasets via in-context learning; applicable to diverse scientific datasets (1.8K+ stars, Apache 2.0)
- [DeepInnovator (HKUDS, arXiv 2026)](https://github.com/HKUDS/DeepInnovator) - Scientific foundation model and AI research copilot for idea generation, cross-disciplinary connection discovery, and hypothesis formation; trained with a decoupled reward-comment RL architecture and achieves GPT-4o-competitive novelty/rationale on STEM and social-science idea-generation benchmarks (270+ stars, MIT License, 2026)
- [LOGOS (arXiv 2026)](https://github.com/LOGOS-Hub/LOGOS) - First multi-domain generative foundation model for the natural sciences built on a unified scientific grammar, encoding proteins, antibodies, small molecules, chemical reactions, materials, and their spatial interactions into a shared token vocabulary; enables unified generation, prediction, and design across domains under a purely autoregressive paradigm (134+ stars, Apache 2.0, 2026)
- [MinervaAI](https://github.com/google-research/minerva) - Mathematical reasoning
- [PaLM-2](https://ai.google/discover/palm2) - Scientific reasoning capabilities
@@ -800,6 +840,7 @@ validated: false
### Physics
- [The Well](https://github.com/PolymathicAI/the_well) - 15TB collection of 16 large-scale numerical simulation datasets spanning fluid dynamics, MHD, astrophysics, biological systems, and acoustic scattering, with unified PyTorch dataloaders and benchmarks for training foundation models on physical sciences (Polymathic AI, NeurIPS 2024)
- [RealPDEBench (ICLR 2026 Oral)](https://github.com/AI4Science-WestlakeU/RealPDEBench) - First scientific ML benchmark with paired real-world measurements and matched numerical simulations for complex physical systems, featuring 5 scenarios, 700+ trajectories, 10 baseline models, and 9 evaluation metrics with HuggingFace datasets and model checkpoints (Westlake University, CC BY-NC 4.0)
- [LIGO Open Science Center](https://gwosc.org/) - Gravitational wave data
- [Particle Data Group](https://pdg.lbl.gov/) - Particle physics data
- [OpenQuantumMaterials](https://www.quantum-materials.org/) - Quantum materials data
@@ -827,6 +868,7 @@ validated: false
- [DiffEqFlux.jl](https://github.com/SciML/DiffEqFlux.jl) - Neural ordinary differential equations with O(1) backprop and GPU support (900+ stars)
- [Optimization.jl](https://github.com/SciML/Optimization.jl) - Unified interface for local, global, gradient-based and derivative-free optimization (800+ stars)
- [PaddleScience](https://github.com/PaddlePaddle/PaddleScience) - SDK & library for AI-driven scientific computing applications
- [Tesseract Core (Pasteur Labs, SciPy 2025 / JOSS)](https://github.com/pasteurlabs/tesseract-core) - Universal components for differentiable scientific computing, packaging heterogeneous scientific tools into self-contained, portable, gradient-propagating components with auto-generated schemas, CLI/REST API/Python SDK interfaces, and reproducible deployment across local, cloud, and HPC environments (105+ stars, Apache 2.0)
- [Flux.jl](https://github.com/FluxML/Flux.jl) - Machine learning in Julia
### Specialized Frameworks
-1
View File
@@ -1 +0,0 @@
test