Compare commits

...
Author SHA1 Message Date
promptadmin 2703e3f667 [upstream-sync] docs/data/resources.json from inoue0426/awesome-computational-biology@0cf037ab [catalogue] 2026-07-16 03:04:18 +00:00
promptadmin ccace33c32 [upstream-sync] data/resources.yml from inoue0426/awesome-computational-biology@0cf037ab [catalogue] 2026-07-16 03:04:09 +00:00
promptadmin 094d21239c [upstream-sync] data/resources.json from inoue0426/awesome-computational-biology@0cf037ab [catalogue] 2026-07-16 03:03:58 +00:00
promptadmin d8da514840 [upstream-sync] cspell.json from inoue0426/awesome-computational-biology@0cf037ab [unknown] 2026-07-16 03:03:47 +00:00
promptadmin 59a389f4c3 [upstream-sync] README.md from inoue0426/awesome-computational-biology@0cf037ab [catalogue] 2026-07-16 03:03:41 +00:00
promptadmin ff0fbd8785 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#49) from upstream-sync/awesome-ai-for-science-20260715-709cee-vpdg into main
Reviewed-on: #49
2026-07-15 16:12:47 +00:00
promptadmin 268b9aaacc [upstream-sync] README.md from ai-boost/awesome-ai-for-science@709cee28 [catalogue] 2026-07-15 15:01:02 +00:00
promptadmin 71a89dca18 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#41) from upstream-sync/awesome-ai-for-science-20260710-230e87-gsmv into main
Reviewed-on: #41
2026-07-10 23:36:12 +00:00
promptadmin bdd7e1176f [upstream-sync] README.md from ai-boost/awesome-ai-for-science@230e8724 [catalogue] 2026-07-10 20:32:23 +00:00
promptadmin ad75261e53 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#37) from upstream-sync/awesome-ai-for-science-20260708-408a2a-oldc into main
Reviewed-on: #37
2026-07-09 01:58:46 +00:00
promptadmin df8c25d33b [upstream-sync] README.md from ai-boost/awesome-ai-for-science@408a2a18 [catalogue] 2026-07-08 20:23:24 +00:00
promptadmin 871e7a94db Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#35) from upstream-sync/awesome-ai-for-science-20260707-4afd72-bcxj into main
Reviewed-on: #35
2026-07-07 21:22:48 +00:00
promptadmin e22c861d46 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@4afd72ec [catalogue] 2026-07-07 20:18:40 +00:00
promptadmin bf0d468978 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260706-e0df96-elsa
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:58 +00:00
promptadmin 336dce2f2e Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260704-cf7431-gaou
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:38 +00:00
promptadmin cb7edbda94 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260704-9f62d8-qyhf
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:22 +00:00
promptadmin d82784013c Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260704-cca092-fmor
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:27:15 +00:00
promptadmin 79442f5a4f Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260703-86c5f1-awps
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-06 15:11:34 +00:00
promptadmin e2d27fc52e Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#31) from upstream-sync/awesome-ai-for-science-20260705-99a575-uivu into main
Reviewed-on: #31
2026-07-06 15:05:14 +00:00
promptadmin a960883a91 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@e0df96d7 [catalogue] 2026-07-06 08:09:54 +00:00
promptadmin 21abde0feb [upstream-sync] README.md from ai-boost/awesome-ai-for-science@99a57577 [catalogue] 2026-07-05 20:08:20 +00:00
promptadmin 39e6e43df2 Merge pull request '[Upstream sync] HKUST-KnowComp/Awesome-LLM-Scientific-Discovery (github) — 0 added, 1 modified' (#26) from upstream-sync/awesome-llm-scientific-discovery-20260703-eb19b4-nzbk into main
Reviewed-on: #26
2026-07-05 14:10:06 +00:00
promptadmin 80ce4af09b Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#30) from upstream-sync/awesome-ai-for-science-20260705-a65119-tfox into main
Reviewed-on: #30
2026-07-05 14:07:01 +00:00
promptadmin 4dd3bb3374 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@a6511998 [catalogue] 2026-07-05 08:05:31 +00:00
promptadmin 55837bb543 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@cf74319a [catalogue] 2026-07-04 19:58:44 +00:00
promptadmin bfc772b9f6 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@9f62d8a3 [catalogue] 2026-07-04 13:57:30 +00:00
promptadmin 168b1b0083 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@cca0924c [catalogue] 2026-07-04 07:56:14 +00:00
promptadmin f7cdcba740 [upstream-sync] README.md from HKUST-KnowComp/Awesome-LLM-Scientific-Discovery@eb19b47e [catalogue] 2026-07-03 19:54:46 +00:00
promptadmin 7fd95fb00f [upstream-sync] README.md from ai-boost/awesome-ai-for-science@86c5f147 [catalogue] 2026-07-03 19:54:17 +00:00
promptadmin 3959588d37 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260702-2d6b81-ailc
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-03 14:01:54 +00:00
promptadmin adb877a343 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260702-e92baf-yafz
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-03 13:58:17 +00:00
promptadmin 64b83d6ceb Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260630-f39947-kdrm
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-07-03 13:57:56 +00:00
promptadmin 4038c1c7cd Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#23) from upstream-sync/awesome-ai-for-science-20260703-de8e26-cqsi into main
Reviewed-on: #23
2026-07-03 13:49:25 +00:00
promptadmin fd47cea53e Merge pull request '[Upstream sync] HKUST-KnowComp/Awesome-LLM-Scientific-Discovery (github) — 0 added, 1 modified' (#24) from upstream-sync/awesome-llm-scientific-discovery-20260703-7fcb88-vvmf into main
Reviewed-on: #24
2026-07-03 13:46:15 +00:00
promptadmin 4b2bcc5978 [upstream-sync] README.md from HKUST-KnowComp/Awesome-LLM-Scientific-Discovery@7fcb8811 [catalogue] 2026-07-03 07:52:08 +00:00
promptadmin eb75b1a36b [upstream-sync] README.md from ai-boost/awesome-ai-for-science@de8e2603 [catalogue] 2026-07-03 07:51:37 +00:00
promptadmin 3db18ec07f [upstream-sync] README.md from ai-boost/awesome-ai-for-science@2d6b813f [catalogue] 2026-07-02 19:50:02 +00:00
promptadmin 40f6c03acd Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#20) from upstream-sync/awesome-ai-for-science-20260701-0cee46-nqrz into main
Reviewed-on: #20
2026-07-02 19:22:12 +00:00
promptadmin 759b1b39fa [upstream-sync] README.md from ai-boost/awesome-ai-for-science@e92baf5c [catalogue] 2026-07-02 07:47:43 +00:00
promptadmin c6fd0dfd73 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@0cee4659 [catalogue] 2026-07-01 19:43:30 +00:00
promptadmin 1730fc59ef Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#19) from upstream-sync/awesome-ai-for-science-20260701-3756bb-ajvc into main
Reviewed-on: #19
2026-07-01 13:59:37 +00:00
promptadmin a6fef26c79 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@3756bb0f [catalogue] 2026-07-01 07:41:32 +00:00
promptadmin 17b875842a [upstream-sync] README.md from ai-boost/awesome-ai-for-science@f3994796 [catalogue] 2026-06-30 19:38:11 +00:00
promptadmin 0f5f8a3c04 Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260628-dbd35d-yahj
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-06-30 16:41:48 +00:00
promptadmin d0e4258f2c Merge upstream-sync branch upstream-sync/awesome-ai-for-science-20260628-f5d952-szrr
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-06-30 16:41:27 +00:00
promptadmin 8bdc92fedf Merge branch 'upstream-sync/awesome-ai-for-science-20260629-602f33-ctuq'
# Conflicts:
#	upstream/ai-boost-awesome-ai-for-science/catalogue/README.md
2026-06-30 16:12:26 +00:00
promptadmin dc6637c082 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#17) from upstream-sync/awesome-ai-for-science-20260630-284a2b-raws into main
Reviewed-on: #17
2026-06-30 16:00:18 +00:00
promptadmin 4caa323bfe [upstream-sync] README.md from ai-boost/awesome-ai-for-science@284a2b4b [catalogue] 2026-06-30 07:35:53 +00:00
promptadmin f397977f7d [upstream-sync] README.md from ai-boost/awesome-ai-for-science@602f33d3 [catalogue] 2026-06-29 19:30:05 +00:00
promptadmin c9b1b1a117 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#15) from upstream-sync/awesome-ai-for-science-20260629-d3eb73-nuat into main
Reviewed-on: #15
2026-06-29 15:56:43 +00:00
promptadmin 254346c6ca [upstream-sync] README.md from ai-boost/awesome-ai-for-science@d3eb7394 [catalogue] 2026-06-29 07:27:39 +00:00
promptadmin 1d855bf433 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@f5d9529d [catalogue] 2026-06-28 19:25:52 +00:00
promptadmin 8d7e346189 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@dbd35d0a [catalogue] 2026-06-28 07:23:44 +00:00
promptadmin 65c3b0ffea cleanup test files 2026-06-27 21:17:48 +00:00
promptadmin 04bfa28b4c cleanup test files 2026-06-27 21:17:25 +00:00
promptadmin fffc13d595 cleanup test files 2026-06-27 21:17:05 +00:00
promptadmin ef48c29b12 cleanup test files 2026-06-27 21:16:57 +00:00
promptadmin c249190497 cleanup test files 2026-06-27 21:16:48 +00:00
promptadmin a98b7daad7 wh url fixed 2026-06-27 21:09:50 +00:00
promptadmin 759de2573a final wh test 2026-06-27 21:03:33 +00:00
promptadmin 4e669b779f wh test2 2026-06-27 20:47:21 +00:00
promptadmin e76b1a0707 test sig fix 2026-06-27 20:23:40 +00:00
promptadmin 847d076523 test webhook after restart 2026-06-27 18:30:20 +00:00
promptadmin 1c62949b28 remove test file 2026-06-27 18:29:10 +00:00
promptadmin 4cc95f3fd4 test webhook 2026-06-27 18:27:24 +00:00
promptadmin e3185d5e20 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#11) from upstream-sync/awesome-ai-for-science-20260627-3a7505-cspp into main
Reviewed-on: #11
2026-06-27 13:32:15 +00:00
promptadmin f2a0bcf08c [upstream-sync] README.md from ai-boost/awesome-ai-for-science@3a7505d8 [catalogue] 2026-06-27 07:20:03 +00:00
promptadmin e29705abef Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 2 added, 0 modified' (#6) from upstream-sync/awesome-ai-for-science-20260626-cf292e into main
Reviewed-on: #6
2026-06-26 20:16:22 +00:00
promptadmin a6f0efef59 Merge pull request '[Upstream sync] GoekeLab/awesome-genomic-skills (github) — 2 added, 0 modified' (#7) from upstream-sync/awesome-genomic-skills-20260626-f88d94 into main
Reviewed-on: #7
2026-06-26 20:15:54 +00:00
promptadmin d28578704e Merge pull request '[Upstream sync] inoue0426/awesome-computational-biology (github) — 17 added, 0 modified' (#8) from upstream-sync/awesome-computational-biology-20260626-12d875 into main
Reviewed-on: #8
2026-06-26 20:15:30 +00:00
promptadmin f3ef23f85c Merge pull request '[Upstream sync] HKUST-KnowComp/Awesome-LLM-Scientific-Discovery (github) — 1 added, 0 modified' (#9) from upstream-sync/awesome-llm-scientific-discovery-20260626-b39b55 into main
Reviewed-on: #9
2026-06-26 20:14:52 +00:00
promptadmin 412fff682b Merge pull request '[Upstream sync] zhoujieli/Awesome-LLM-Agents-Scientific-Discovery (github) — 1 added, 0 modified' (#10) from upstream-sync/awesome-llm-agents-scientific-discovery-20260626-3e079c into main
Reviewed-on: #10
2026-06-26 20:14:17 +00:00
promptadmin 71c6bd7cf5 [upstream-sync] README.md from zhoujieli/Awesome-LLM-Agents-Scientific-Discovery@3e079cd8 [catalogue] 2026-06-26 19:58:19 +00:00
promptadmin 3ce2327c80 [upstream-sync] README.md from HKUST-KnowComp/Awesome-LLM-Scientific-Discovery@b39b55ac [catalogue] 2026-06-26 19:58:02 +00:00
promptadmin 68de08772e [upstream-sync] scripts/requirements.txt from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:47 +00:00
promptadmin 53ac75b25f [upstream-sync] docs/data/resources.json from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:42 +00:00
promptadmin 785bafe634 [upstream-sync] docs/data/SCHEMA.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:37 +00:00
promptadmin 7ccb737be8 [upstream-sync] data/resources.yml from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:32 +00:00
promptadmin c138319a7c [upstream-sync] data/resources.json from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:27 +00:00
promptadmin 0b50e4ff76 [upstream-sync] cspell.json from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:57:22 +00:00
promptadmin f54babd4b6 [upstream-sync] contributing.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:18 +00:00
promptadmin 0b3b43a965 [upstream-sync] code-of-conduct.md from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:57:13 +00:00
promptadmin c7d42ec577 [upstream-sync] README.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:09 +00:00
promptadmin 23455a82ee [upstream-sync] .markdown-link-check.json from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:57:00 +00:00
promptadmin c22538cf86 [upstream-sync] .github/workflows/update-overview.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:56:56 +00:00
promptadmin 5b47e38135 [upstream-sync] .github/workflows/sync_resources.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:56:52 +00:00
promptadmin 5b97805747 [upstream-sync] .github/workflows/pr-quality-checks.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:56:44 +00:00
promptadmin 9a124e3118 [upstream-sync] .github/workflows/link-check.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:56:39 +00:00
promptadmin b185fe9924 [upstream-sync] .github/workflows/generate-artifacts.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:56:35 +00:00
promptadmin f223f4fbed [upstream-sync] .github/workflows/docs-check.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 19:56:29 +00:00
promptadmin 17185eb892 [upstream-sync] .github/PULL_REQUEST_TEMPLATE.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 19:56:25 +00:00
promptadmin 58cc618ecc [upstream-sync] README.md from GoekeLab/awesome-genomic-skills@f88d9494 [catalogue] 2026-06-26 19:56:07 +00:00
promptadmin 52fd40bc11 [upstream-sync] CONTRIBUTING.md from GoekeLab/awesome-genomic-skills@f88d9494 [catalogue] 2026-06-26 19:56:00 +00:00
promptadmin 9ae546303e [upstream-sync] README.md from ai-boost/awesome-ai-for-science@cf292eeb [catalogue] 2026-06-26 19:55:45 +00:00
promptadmin be747ff5a5 [upstream-sync] CONTRIBUTING.md from ai-boost/awesome-ai-for-science@cf292eeb [catalogue] 2026-06-26 19:55:37 +00:00
promptadmin 65bfecbca8 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 2 added, 0 modified' (#5) from upstream-sync/awesome-ai-for-science-20260626-cf292e into main
Reviewed-on: #5
2026-06-26 18:55:57 +00:00
promptadmin e03c2ce8a7 [upstream-sync] README.md from ai-boost/awesome-ai-for-science@cf292eeb [catalogue] 2026-06-26 17:37:27 +00:00
promptadmin bbc698bde9 [upstream-sync] CONTRIBUTING.md from ai-boost/awesome-ai-for-science@cf292eeb [catalogue] 2026-06-26 17:37:21 +00:00
promptadmin ca789ee0f5 Merge pull request '[Upstream sync] zhoujieli/Awesome-LLM-Agents-Scientific-Discovery (github) — 1 added, 0 modified' (#4) from upstream-sync/awesome-llm-agents-scientific-discovery-20260626-3e079c into main
Reviewed-on: #4
2026-06-26 17:13:12 +00:00
promptadmin 5ea0340f6c Merge pull request '[Upstream sync] GoekeLab/awesome-genomic-skills (github) — 2 added, 0 modified' (#1) from upstream-sync/awesome-genomic-skills-20260626-f88d94 into main
Reviewed-on: #1
2026-06-26 17:12:08 +00:00
promptadmin 401be4e47b Merge pull request '[Upstream sync] HKUST-KnowComp/Awesome-LLM-Scientific-Discovery (github) — 1 added, 0 modified' (#3) from upstream-sync/awesome-llm-scientific-discovery-20260626-b39b55 into main
Reviewed-on: #3
2026-06-26 17:10:46 +00:00
promptadmin ae49943cd1 Merge pull request '[Upstream sync] inoue0426/awesome-computational-biology (github) — 17 added, 0 modified' (#2) from upstream-sync/awesome-computational-biology-20260626-12d875 into main
Reviewed-on: #2
2026-06-26 17:09:59 +00:00
promptadmin dfc14dd0d4 [upstream-sync] README.md from zhoujieli/Awesome-LLM-Agents-Scientific-Discovery@3e079cd8 [catalogue] 2026-06-26 17:09:46 +00:00
promptadmin 0dd57557cf [upstream-sync] README.md from HKUST-KnowComp/Awesome-LLM-Scientific-Discovery@b39b55ac [catalogue] 2026-06-26 17:09:32 +00:00
promptadmin 9536723d15 [upstream-sync] scripts/requirements.txt from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:09:14 +00:00
promptadmin e01176c758 [upstream-sync] docs/data/resources.json from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:09:10 +00:00
promptadmin 92afcb466b [upstream-sync] docs/data/SCHEMA.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:09:05 +00:00
promptadmin 1915db6025 [upstream-sync] data/resources.yml from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:09:01 +00:00
promptadmin 3289d84be5 [upstream-sync] data/resources.json from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:08:57 +00:00
promptadmin acdadebd3f [upstream-sync] cspell.json from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:53 +00:00
promptadmin 45c0c436c2 [upstream-sync] contributing.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:08:49 +00:00
promptadmin b935b3321f [upstream-sync] code-of-conduct.md from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:44 +00:00
promptadmin 68ba3b64aa [upstream-sync] README.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:08:40 +00:00
promptadmin afd315c59f [upstream-sync] .markdown-link-check.json from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:08:36 +00:00
promptadmin aeba4da7b4 [upstream-sync] .github/workflows/update-overview.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:33 +00:00
promptadmin d6ebf617f7 [upstream-sync] .github/workflows/sync_resources.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:29 +00:00
promptadmin 322fa5182e [upstream-sync] .github/workflows/pr-quality-checks.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:24 +00:00
promptadmin 879fc1d832 [upstream-sync] .github/workflows/link-check.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:20 +00:00
promptadmin bf8abaf587 [upstream-sync] .github/workflows/generate-artifacts.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:15 +00:00
promptadmin f168b060b1 [upstream-sync] .github/workflows/docs-check.yml from inoue0426/awesome-computational-biology@12d87583 [unknown] 2026-06-26 17:08:11 +00:00
promptadmin c9f5af0c2a [upstream-sync] .github/PULL_REQUEST_TEMPLATE.md from inoue0426/awesome-computational-biology@12d87583 [catalogue] 2026-06-26 17:08:07 +00:00
promptadmin f8b4137eb1 [upstream-sync] README.md from GoekeLab/awesome-genomic-skills@f88d9494 [catalogue] 2026-06-26 17:07:53 +00:00
promptadmin 91a4a25cbb [upstream-sync] CONTRIBUTING.md from GoekeLab/awesome-genomic-skills@f88d9494 [catalogue] 2026-06-26 17:07:48 +00:00
promptadmin f6edd53e92 Add scientific paper deep summarisation prompt 2026-06-10 17:26:29 +00:00
promptadmin 691af8346f Add flow cytometry interpretation prompt 2026-06-10 17:26:28 +00:00
promptadmin fb86e0ae59 Add protein structure functional interpretation prompt 2026-06-10 17:26:26 +00:00
promptadmin 5de456d9db Add CRISPR off-target risk assessment prompt 2026-06-10 17:26:24 +00:00
promptadmin 70b21c56a9 Add RNA-seq differential expression narrative prompt 2026-06-10 17:26:22 +00:00
promptadmin b6917557aa Add SNP clinical significance interpreter prompt 2026-06-10 17:26:20 +00:00
promptadmin b779730b03 Add README with resource attribution 2026-06-10 17:26:18 +00:00
30 changed files with 17282 additions and 2 deletions
+41 -2
View File
@@ -1,3 +1,42 @@
# life-science-ai-prompts
# Life Science AI Prompts
Curated AI prompts for life sciences: genomics, proteomics, CRISPR, cell biology, and literature synthesis. Analogous to awesome-ai-for-science and awesome-genomic-skills.
> *Where the prompts live, thrive, and reach the world.*
Curated, versioned AI prompts for life sciences research.
Covers genomics, proteomics, CRISPR, cell biology, and literature synthesis.
## Analogous Resources Ingested
| Repository | Focus | Stars |
|---|---|---|
| [awesome-ai-for-science](https://github.com/ai-boost/awesome-ai-for-science) | AI tools across sciences | ★★★ |
| [Awesome-LLM-Agents-Scientific-Discovery](https://github.com/zjlrock777/Awesome-LLM-Agents-Scientific-Discovery) | LLM agents in biomedical research | ★★★ |
| [awesome-genomic-skills](https://github.com/GoekeLab/awesome-genomic-skills) | Genomic LLM agent skills | ★★★ |
| [awesome-computational-biology](https://github.com/inoue0426/awesome-computational-biology) | Computational biology resources | ★★★ |
| [Awesome-LLM-Scientific-Discovery](https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery) | LLMs for science | ★★★ |
## Folder Structure
```
genomics/ — variant calling, RNA-seq, CRISPR
proteomics/ — structure, mass spectrometry, interactions
cell-biology/ — flow cytometry, microscopy, assays
literature/ — paper summarisation, methods extraction
```
## How to Use
```python
import requests, base64
GITEA_URL = "https://promptnotes.ai"
TOKEN = "your-token"
def read_prompt(owner, repo, path):
url = f"{GITEA_URL}/api/v1/repos/{owner}/{repo}/contents/{path}"
r = requests.get(url, headers={"Authorization": f"token {TOKEN}"})
return base64.b64decode(r.json()["content"]).decode()
prompt = read_prompt("promptadmin", "life-science-ai-prompts",
"genomics/variant-interpretation/snp-clinical-significance.md")
```
@@ -0,0 +1,65 @@
---
title: "Flow Cytometry Data Interpretation"
domain: cell-biology
persona: "Molecular Biologist"
persona_background: >
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
persona_style: "precise, evidence-based, uses established nomenclature"
models: [gpt-4, claude-3-5]
keywords: [flow-cytometry, FACS, cell-population, gating, immunophenotyping]
task: "Interpret flow cytometry gating strategy and cell population data."
validated: false
version: 1.0.0
author: promptadmin
source_repositories:
- https://github.com/zjlrock777/Awesome-LLM-Agents-Scientific-Discovery
---
# Flow Cytometry Data Interpretation
## Persona
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
> Your communication style: precise, evidence-based, uses established nomenclature
## Task
Interpret flow cytometry gating strategy and cell population data.
## Prompt
```
You are an expert in flow cytometry and immunophenotyping.
Given flow cytometry experiment:
- Cell type: {cell_type}
- Tissue source: {tissue}
- Panel: {markers}
- Gating strategy: {gating_description}
- Key populations identified: {populations}
- Experimental condition: {condition}
- Controls: {controls}
Provide:
1. Assessment of gating strategy quality
2. Interpretation of each identified cell population
3. Biological significance of observed population shifts
4. Statistical recommendations (% parent vs % total, n required)
5. Potential artefacts and confounders
6. Suggested additional markers for confirmation
```
## Notes
Works well with FlowJo or FCS Express output descriptions. Reference: STAgent (Harvard LiuLab, bioRxiv 2025) for spatial context.
## Compatibility
| Model | Tested | Notes |
|-------|--------|-------|
| gpt-4 | ⬜ | |
| claude-3-5 | ⬜ | |
## Keywords
`flow-cytometry` `FACS` `cell-population` `gating` `immunophenotyping`
@@ -0,0 +1,62 @@
---
title: "CRISPR Off-Target Risk Assessment"
domain: genomics
persona: "Molecular Biologist"
persona_background: >
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
persona_style: "precise, evidence-based, uses established nomenclature"
models: [gpt-4, claude-3-5]
keywords: [CRISPR, guide-RNA, off-target, Cas9, gene-editing]
task: "Assess off-target risk of a CRISPR guide RNA based on computational predictions."
validated: false
version: 1.0.0
author: promptadmin
source_repositories:
- https://github.com/ai-boost/awesome-ai-for-science
---
# CRISPR Off-Target Risk Assessment
## Persona
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
> Your communication style: precise, evidence-based, uses established nomenclature
## Task
Assess off-target risk of a CRISPR guide RNA based on computational predictions.
## Prompt
```
You are a CRISPR expert with deep knowledge of guide RNA design and off-target effects.
Given:
- Target gene: {target_gene}
- Guide RNA sequence (20nt): {grna_sequence}
- Predicted off-target sites (from CRISPOR/Cas-OFFinder): {off_target_sites}
- Genome: {genome_assembly}
- Cas variant: {cas_variant}
Provide:
1. Risk classification (Low/Medium/High)
2. Analysis of top 3 off-target sites by genomic context
3. Recommended experimental validation strategy (T7E1, GUIDE-seq, etc.)
4. Alternative guide RNA suggestions if risk is High
5. Summary suitable for IACUC/ethics submission
```
## Notes
Referenced from BioAgents framework (bio-xyz, arXiv 2601.12542).
## Compatibility
| Model | Tested | Notes |
|-------|--------|-------|
| gpt-4 | ⬜ | |
| claude-3-5 | ⬜ | |
## Keywords
`CRISPR` `guide-RNA` `off-target` `Cas9` `gene-editing`
@@ -0,0 +1,62 @@
---
title: "RNA-seq Differential Expression Narrative"
domain: genomics
persona: "Molecular Biologist"
persona_background: >
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
persona_style: "precise, evidence-based, uses established nomenclature"
models: [gpt-4, claude-3-5]
keywords: [RNA-seq, DESeq2, differential-expression, pathway-analysis, fold-change]
task: "Generate a scientific narrative from RNA-seq differential expression results."
validated: true
version: 1.0.0
author: promptadmin
source_repositories:
- https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery
---
# RNA-seq Differential Expression Narrative
## Persona
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
> Your communication style: precise, evidence-based, uses established nomenclature
## Task
Generate a scientific narrative from RNA-seq differential expression results.
## Prompt
```
You are a senior molecular biologist analysing transcriptomic data.
Given DESeq2 differential expression results:
- Comparison: {condition_A} vs {condition_B}
- Significantly upregulated genes (top 10): {up_genes}
- Significantly downregulated genes (top 10): {down_genes}
- Pathway enrichment results: {pathways}
- Experimental context: {context}
Write a Results section (150-200 words) for a peer-reviewed manuscript that:
1. Summarises the overall transcriptional response
2. Highlights key gene clusters and their biological significance
3. Connects enriched pathways to the experimental condition
4. Uses appropriate statistical language (FDR, log2FC)
5. Avoids overclaiming causality
```
## Notes
Derived from GenoTEX benchmark methodology (Liu et al. 2024). Works best with GSEA or EnrichR pathway results.
## Compatibility
| Model | Tested | Notes |
|-------|--------|-------|
| gpt-4 | ✅ | |
| claude-3-5 | ✅ | |
## Keywords
`RNA-seq` `DESeq2` `differential-expression` `pathway-analysis` `fold-change`
@@ -0,0 +1,78 @@
---
title: "SNP Clinical Significance Interpreter"
domain: genomics
persona: "Molecular Biologist"
persona_background: >
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
persona_style: "precise, evidence-based, uses established nomenclature"
models: [gpt-4, claude-3-5, gemini-1-5-pro]
keywords: [SNP, variant-calling, clinical-significance, VCF, ClinVar, ACMG]
task: "Interpret the clinical significance of a single nucleotide polymorphism (SNP) from VCF annotation data."
validated: true
version: 1.0.0
author: promptadmin
source_repositories:
- https://github.com/GoekeLab/awesome-genomic-skills
- https://github.com/ai-boost/awesome-ai-for-science
---
# SNP Clinical Significance Interpreter
## Persona
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
> Your communication style: precise, evidence-based, uses established nomenclature
## Task
Interpret the clinical significance of a single nucleotide polymorphism (SNP) from VCF annotation data.
## Prompt
```
You are a molecular biologist specialising in clinical genomics.
Given the following SNP annotation data from a VCF file:
- Gene: {gene_name}
- Variant: {hgvs_notation}
- ClinVar classification: {clinvar_class}
- gnomAD allele frequency: {gnomad_af}
- CADD score: {cadd_score}
- In silico predictions: {sift} (SIFT), {polyphen} (PolyPhen-2)
Provide:
1. ACMG/AMP classification (Pathogenic/Likely Pathogenic/VUS/Likely Benign/Benign)
2. Evidence summary (2-3 sentences)
3. Clinical implications
4. Recommended follow-up actions
5. Caveats and limitations
```
### Example 1
**Input:**
```
Gene: BRCA1, Variant: c.5266dupC, ClinVar: Pathogenic, gnomAD: 0.00001, CADD: 35
```
**Output:**
```
ACMG: Pathogenic. This frameshift variant creates a premature stop codon...
```
## Notes
Inspired by SRAgent (Arc Institute) for genomic database querying. Best used with SnpEff/VEP-annotated VCF files.
## Compatibility
| Model | Tested | Notes |
|-------|--------|-------|
| gpt-4 | ✅ | |
| claude-3-5 | ✅ | |
| gemini-1-5-pro | ✅ | |
## Keywords
`SNP` `variant-calling` `clinical-significance` `VCF` `ClinVar` `ACMG`
+67
View File
@@ -0,0 +1,67 @@
---
title: "Scientific Paper Deep Summarisation"
domain: literature
persona: "Molecular Biologist"
persona_background: >
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
persona_style: "precise, evidence-based, uses established nomenclature"
models: [gpt-4, claude-3-5, gemini-1-5-pro]
keywords: [literature-review, paper-summarisation, methods-extraction, PubMed]
task: "Generate a structured deep summary of a life sciences research paper."
validated: true
version: 1.0.0
author: promptadmin
source_repositories:
- https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery
- https://github.com/zjlrock777/Awesome-LLM-Agents-Scientific-Discovery
---
# Scientific Paper Deep Summarisation
## Persona
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
> Your communication style: precise, evidence-based, uses established nomenclature
## Task
Generate a structured deep summary of a life sciences research paper.
## Prompt
```
You are an expert scientific reader with broad knowledge of life sciences.
Read the following paper abstract/full text and provide a structured summary:
Paper text:
{paper_text}
Generate:
1. **TL;DR** (1 sentence, non-technical)
2. **Background** — What problem does this paper address?
3. **Key Methods** — What experimental and computational approaches were used?
4. **Main Findings** — What are the 3-5 most important results?
5. **Novelty** — What is genuinely new compared to prior work?
6. **Limitations** — What are the key weaknesses the authors acknowledge or you identify?
7. **Clinical/Translational Relevance** — Practical implications (1-2 sentences)
8. **Follow-up Questions** — 3 questions this paper raises
Format: structured markdown with headers.
```
## Notes
Inspired by Agent Laboratory (2024) three-phase research pipeline. For full-text papers, chunk into introduction + methods + results + discussion.
## Compatibility
| Model | Tested | Notes |
|-------|--------|-------|
| gpt-4 | ✅ | |
| claude-3-5 | ✅ | |
| gemini-1-5-pro | ✅ | |
## Keywords
`literature-review` `paper-summarisation` `methods-extraction` `PubMed`
@@ -0,0 +1,66 @@
---
title: "Protein Structure Functional Interpretation"
domain: proteomics
persona: "Structural Biologist"
persona_background: >
Computational structural biologist specialising in protein folding, cryo-EM, and AlphaFold interpretation.
persona_style: "quantitative, structure-first, references PDB entries"
models: [gpt-4, claude-3-5, gemini-1-5-pro]
keywords: [AlphaFold, protein-structure, PDB, active-site, binding-pocket]
task: "Interpret AlphaFold2/3 or experimental protein structure in functional context."
validated: true
version: 1.0.0
author: promptadmin
source_repositories:
- https://github.com/inoue0426/awesome-computational-biology
- https://github.com/ai-boost/awesome-ai-for-science
---
# Protein Structure Functional Interpretation
## Persona
> You are a **Structural Biologist**. Computational structural biologist specialising in protein folding, cryo-EM, and AlphaFold interpretation.
> Your communication style: quantitative, structure-first, references PDB entries
## Task
Interpret AlphaFold2/3 or experimental protein structure in functional context.
## Prompt
```
You are a structural biologist with expertise in computational and experimental structural analysis.
Given protein structure data:
- Protein name: {protein_name}
- UniProt ID: {uniprot_id}
- Structure source: {source} (AlphaFold2 / AlphaFold3 / X-ray / Cryo-EM)
- pLDDT scores summary: {plddt_summary}
- Key structural features: {features}
- Known binding partners: {binding_partners}
Provide:
1. Overall structural assessment (fold classification, domain organisation)
2. Confidence assessment for key regions (if AlphaFold)
3. Predicted functional sites (active site, allosteric sites, binding interfaces)
4. Druggability assessment of binding pockets
5. Structural basis for any known pathogenic variants
6. Recommended follow-up experiments
```
## Notes
Integrates well with PyMOL output descriptions and PDB REMARK sections. For AlphaFold3 structures, note pLDDT < 70 regions as disordered.
## Compatibility
| Model | Tested | Notes |
|-------|--------|-------|
| gpt-4 | ✅ | |
| claude-3-5 | ✅ | |
| gemini-1-5-pro | ✅ | |
## Keywords
`AlphaFold` `protein-structure` `PDB` `active-site` `binding-pocket`
@@ -0,0 +1,46 @@
---
title: "Contribution Guidelines"
task: ""
lineage_type: import
upstream_source: https://github.com/GoekeLab/awesome-genomic-skills/blob/f88d9494/CONTRIBUTING.md
upstream_sha: f88d9494
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Contribution Guidelines
Thank you for considering contributing to **Awesome Genomic Skills**! This is a curated list, so please make sure your suggestion meets the criteria below before submitting a pull request.
## Adding an Item
1. **Fork** the repository and create a new branch.
2. Add your item to the appropriate section in `README.md`.
3. Submit a **pull request** with a clear description of the item and why it belongs in the list.
## Criteria for Inclusion
To be included, a repository should meet the following:
- **Relevance** — directly related to AI coding agents, and useful for genomics/bioinformatics.
- **Quality** — actively maintained, well-documented, and demonstrably useful.
- **Not deprecated** — must not be archived or abandoned.
## Format
Each entry should follow this format:
```markdown
- [Repository Name](URL) - Short description explaining what it does and why it is useful.
```
- Keep descriptions concise (one or two sentences).
- Place the entry in the most appropriate existing section.
- If no section fits, propose a new one in your pull request description.
## Pull Request Guidelines
- Verify all links are correct and publicly accessible.
@@ -0,0 +1,131 @@
---
title: "Awesome Genomic Skills [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)"
task: ""
lineage_type: import
upstream_source: https://github.com/GoekeLab/awesome-genomic-skills/blob/f88d9494/README.md
upstream_sha: f88d9494
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Awesome Genomic Skills [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)
A curated list of skills and MCP servers for working with AI coding agents (Claude Code, GitHub Copilot, Codex, Cursor, Gemini CLI, etc.) in genomics and bioinformatics, alongside other useful repositories such as benchmarks and general AI coding skills collections.
### What is a Skill and what is an MCP?
**Skill:** a Markdown file (plus optional scripts) that teaches an agent *how* to do a task — procedural know-how loaded into context on demand.
**MCP server:** a running service that gives an agent a *connection* to external systems (databases, tools, pipelines) via standardized tool calls.
## Contents
- [Awesome Genomic Skills ](#awesome-genomic-skills-)
- [What is a Skill and what is an MCP?](#what-is-a-skill-and-what-is-an-mcp)
- [Contents](#contents)
- [Bioinformatics and Genomics Agent Skills](#bioinformatics-and-genomics-agent-skills)
- [MCP Servers for Life Sciences](#mcp-servers-for-life-sciences)
- [Benchmarks](#benchmarks)
- [General AI Coding Agent Skill Collections](#general-ai-coding-agent-skill-collections)
- [Other Notable Awesome Lists!](#other-notable-awesome-lists)
- [Contributing](#contributing)
## Bioinformatics and Genomics Agent Skills
Skill libraries and tool collections specifically targeting genomics, bioinformatics, and life sciences work with AI coding agents.
- [science-skills](https://github.com/google-deepmind/science-skills)
- **Description:** Collection of ~36 agent skills spanning genomics, structural biology, cheminformatics, and literature search; wraps AlphaGenome (single-variant effect prediction), AlphaFold DB, and 30+ databases/tools (UniProt, Ensembl, gnomAD, GTEx, ClinVar, dbSNP, ChEMBL, PubChem, PDB, Foldseek, JASPAR, Reactome, STRING, Open Targets, Human Protein Atlas, PyMOL) for grounded, token-efficient scientific workflows. Apache-2.0; built for Google Antigravity but installable into any agent via `npx skills add`. [Technical report](https://storage.googleapis.com/deepmind-media/papers/google_deepmind_science_skills_for_antigravity_towards_efficient_and_reliable_scientific_workflows.pdf).
- **Developers:** Google DeepMind.
- [openai/plugins — life-science-research](https://github.com/openai/plugins/tree/main/plugins/life-science-research)
- **Description:** OpenAI's Life Sciences research plugin for Codex; bundles 50 modular skills spanning human genetics, functional genomics, expression, pathway biology, protein structure, chemistry, and clinical evidence, wrapping 50+ public databases/tools (Ensembl, UniProt, gnomAD, GTEx, ClinVar, GWAS Catalog, Open Targets, ChEMBL, PubChem, RCSB PDB, AlphaFold, Reactome, STRING, Human Protein Atlas, cBioPortal, CellxGene, ENCODE, NCBI Entrez/BLAST/Datasets, and more). Works with mainline GPT-5.4 models, with GPT-Rosalind (trusted-access) for deeper reasoning. Closely parallels DeepMind's science-skills.
- **Developers:** OpenAI.
- [anthropics/life-sciences](https://github.com/anthropics/life-sciences)
- **Description:** Anthropic's "Claude for Life Sciences" Claude Code Marketplace — a hybrid bundle rather than a pure skills library. It mixes first-party procedural/method skills (`single-cell-rna-qc`, `scvi-tools`, `nextflow-development`, `clinical-trial-protocol-skill`, `scientific-problem-selection`) with a large set of partner MCP servers and data integrations (10x Genomics, ChEMBL, Open Targets, PubMed, bioRxiv, Synapse, ToolUniverse, BioRender, Medidata, Cortellis, Owkin, Consensus, and more). Note that many entries are MCP servers, not skills.
- **Developers:** Anthropic (with commercial life-sciences partners).
- [sample-kiro-power-life-sciences](https://github.com/aws-samples/sample-kiro-power-life-sciences)
- **Description:** AWS sample bundle for the Kiro IDE: 24 MCP servers wrapping 100+ databases/tools across genomics, proteomics, structural biology, and clinical/pharma (NCBI, Ensembl, ClinVar, gnomAD, UniProt, STRING, PDB, AlphaFold, ChEMBL, Open Targets, etc.), plus 10 domain skills and 16 workflows, with cross-database search and AWS HealthOmics pipeline execution. MIT-0; the MCP servers use standard MCP and are portable, though the skills and workflows are built for Kiro. [Blog post](https://aws.amazon.com/blogs/publicsector/accelerating-life-sciences-research-with-kiro-a-unified-ai-interface-to-100-open-source-databases/).
- **Developers:** AWS (AWS Samples).
- [ClawBio](https://github.com/ClawBio/ClawBio)
- **Description:** The first bioinformatics-native AI agent skill library; provides reproducible, local-first skills for genomics tasks (variant calling, RNA-seq, population genetics) that work with Claude Code, Copilot, Codex, and other agents.
- **Developers:** Independent open-source project built on [OpenClaw](https://openclaw.ai).
- [SciAgent-skills](https://github.com/jaechang-hits/SciAgent-Skills)
- **Description:** 197 open-source skills for Claude Code, Cursor, Codex, and Windsurf, covering genomics-bioinformatics, proteomics-protein engineering, structural biology, drug discovery, systems biology, biostatistics, and scientific writing; achieves 92% accuracy on BixBench-Verified-50 (+26.7 pts over Claude Code baseline). The hosted OmicsHorizon web platform runs these skills in-browser.
- **Developers:** The team behind the [OmicsHorizon](https://omicshorizon.ai/en/) platform (Jaechang Lim).
- [scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- **Description:** Currently 135 skills covering various scientific areas including genomics, but also broader scientific areas like geospatial science etc.
- **Developers:** K-Dense AI — an MIT-founded startup (Accel, Accel Atoms, Google AI Futures Fund) building an AI research platform ([k-dense.ai](https://www.k-dense.ai)); this is one of its open-source spin-offs.
- [bioSkills](https://github.com/GPTomics/bioSkills)
- **Description:** SKILL.md files for bioinformatics with Claude Code, covering end-to-end pipelines like RNA-seq, variants, ChIP-seq, scRNA-seq, spatial, Hi-C, proteomics, microbiome, CRISPR, metabolomics, multi-omics, immunotherapy, outbreak analysis, and Mendelian randomization tools.
- **Developers:** GPTomics ([gptomics.com](https://www.gptomics.com)), an independent research lab working at the bioinformatics/AI intersection.
- [Clair-skills](https://github.com/HKU-BAL/Clair-skills)
- **Description:** Agent skill for the Clair suite of variant callers (Clair3, ClairS, Clair3-RNA, Clair-Mosaic); provides intelligent model selection, command generation, and troubleshooting for germline, somatic, mosaic, and RNA-seq variant calling from long-read and short-read sequencing data.
- **Developers:** Ruibang Luo's BioAI Lab at the University of Hong Kong (HKU-BAL).
- [ToolUniverse](https://github.com/mims-harvard/ToolUniverse)
- **Description:** Ecosystem for building AI scientist systems; integrates 1,000+ machine learning models, datasets, and scientific APIs for data analysis, knowledge retrieval, and experimental design in biomedicine, with 68 pre-built agent skills covering drug discovery, precision oncology, and rare-disease diagnosis.
- **Developers:** Harvard's Zitnik Lab (AI for Medicine and Science).
- [operon](https://github.com/swaruplab/operon)
- **Description:** AI-powered bioinformatics IDE bundling 180+ SKILL.md-format analysis protocols covering RNA-seq, scRNA-seq, ATAC-seq, ChIP-seq, WGS/WES, spatial transcriptomics, proteomics, GWAS, and external database query patterns (PubMed, GEO, GTEx, KEGG, UniProt, JASPAR, AlphaFold). The desktop app wraps Claude Code and adds HPC/SSH integration; the protocol files themselves are MIT-licensed Markdown in the repo's `protocols/` directory, extractable for use with any SKILL.md-compatible agent.
- **Developers:** Swarup Lab (UC Irvine).
- [OpenClaw-Medical-Skills](https://github.com/FreedomIntelligence/OpenClaw-Medical-Skills)
- **Description:** Meta-aggregation of 872 skills curated from 12+ upstream repositories, including ClawBio, ToolUniverse, GPTomics bioSkills, BioOS, and others; spans clinical workflows, genomics, drug discovery, bioinformatics pipelines, and medical device regulatory frameworks. Expect overlap with the upstream repos listed separately.
- **Developers:** FreedomIntelligence, the medical-NLP research group at CUHK-Shenzhen / Shenzhen Research Institute of Big Data (led by Benyou Wang; also behind HuatuoGPT).
## MCP Servers for Life Sciences
Model Context Protocol (MCP) servers that give AI agents direct access to bioinformatics databases, tools, and analysis pipelines.
- [ChatSpatial](https://github.com/cafferychen777/ChatSpatial) - MCP server for spatial transcriptomics analysis through natural language; supports Scanpy, Squidpy, cell communication analysis, and spatial domain identification.
- [biomcp](https://github.com/genomoncology/biomcp) - Single MCP server able to query multiple information sources, including clinical trials, genetic data & published medical literature.
- [gget-mcp](https://github.com/longevity-genie/gget-mcp) - MCP server wrapping the Pachter Lab [gget](https://github.com/pachterlab/gget) bioinformatics toolkit. Exposes 13 tools covering gene search and metadata (Ensembl), sequence retrieval, BLAST/BLAT/MUSCLE alignment, expression data (ARCHS4), functional enrichment (Enrichr), protein structure (PDB, AlphaFold), cancer mutations (COSMIC), and single-cell queries (CellxGene).
- [Seqera MCP](https://docs.seqera.io/platform-cloud/seqera-mcp/overview) - Hosted MCP server from Seqera Labs (the developers of Nextflow) exposing the Seqera Platform (workflow launch/management), Wave (container provisioning), nf-core modules, and SRA/ENA/GEO retrieval.
- [knowledgebase-mcp](https://github.com/biocontext-ai/knowledgebase-mcp) - BioContextAI Knowledgebase MCP server, included in the BioContextAI registry (below), one of the most comprehensive single MCP packages (wraps STRINGDb, Open Targets, Reactome, UniProt, HPA, KEGG, AlphaFold, Ensembl, ClinicalTrials.gov, bioRxiv, etc.)
Existing registries and lists of MCP servers:
- [BioContextAI Registry](https://github.com/biocontext-ai/registry/) - Community-curated catalogue of biomedical MCP servers, with submission criteria requiring biomedical focus, free academic access, OSI-approved open-source licenses, and MCP specification compliance. Ships a [cookiecutter template](https://github.com/biocontext-ai/mcp-server-cookiecutter) for new servers and follows Schema.org ontologies for metadata. [A community hub for agentic biomedical systems](https://www.nature.com/articles/s41587-025-02900-9)
- [MCPmed](https://github.com/MCPmed) - Reference MCP implementations (GEO, STRING, UCSC Cell Browser, PLSDB) plus a cookiecutter template and HTML "breadcrumbs" discovery mechanism for transitioning legacy services to MCP. Published as a call paper in Briefings in Bioinformatics. [MCPmed: a call for Model Context Protocol-enabled bioinformatics web services for LLM-driven discovery](https://academic.oup.com/bib/article/27/1/bbag076/8495038)
- [awesome-mcp-servers](https://github.com/punkpeye/awesome-mcp-servers#bio) - General-purpose MCP server list with a Biology, Medicine, and Bioinformatics subsection.
MCP related tools:
- [BioinfoMCP](https://github.com/florensiawidjaja/BioinfoMCP) - Not strictly an MCP, a converter that auto-generates MCP servers from existing tool documentation, plus a benchmark of the converted tools. Preprint available [here](https://arxiv.org/abs/2510.02139)
## Benchmarks
Benchmarks, evaluations, and other helpful resources at the intersection of AI and genomics/bioinformatics.
Agent Capability Benchmarks:
- [BioAgent Bench](https://github.com/bioagent-bench/bioagent-bench) - 10 end-to-end bioinformatics pipeline tasks (RNA-seq, variant calling, metagenomics, single-cell, transcript quantification, etc.) with concrete output artifacts; includes a perturbation suite (corrupted inputs, decoy reference files, prompt bloat) — probes agent robustness under controlled stress and shows that correct high-level pipeline construction does not guarantee reliable step-level reasoning.
- [BioMed-AQA](https://huggingface.co/datasets/BOBQWERA/biomed-aqa-dataset) - 327 open-ended biomedical analysis tasks across omics, visualisation, machine learning, statistics, and precision medicine; uses milestone-based grading against reference analytical steps, with a complementary 172-question multiple-choice subset. Released alongside the BioMedAgent system in Nature Biomedical Engineering — same-team caveat applies to headline scores.
- [BiomniBench](https://huggingface.co/datasets/phylobio/BiomniBench-DA) - Process-level evaluation framework: 100 biomedical data-analysis tasks curated from high-impact papers by original authors or domain experts; grades the full agent trajectory (reasoning trace + final answer) against expert-designed rubrics via an LLM judge — addresses the outcome-only blind spots in benchmarks like BixBench.
- [BixBench](https://github.com/Future-House/BixBench) - Comprehensive benchmark for LLM-based agents on real-world computational biology tasks; tests agents' ability to explore biological datasets, perform multi-step analyses, and interpret results — useful for evaluating which agents and skills perform best on genomics work.
- [CompBioBench](https://github.com/Genentech/compbiobench-runner) - Genentech-released benchmark of 100 computational biology questions spanning single-cell, epigenomics, genomics, transcriptomics, human genetics, and ML; agents start from a bare-minimum environment and must fetch their own tools and data, with exact-string-match grading on a single ground-truth answer.
- [LAB-Bench](https://huggingface.co/datasets/futurehouse/lab-bench) - 2,400+ multiple-choice questions across 8 subtasks of practical biology research (literature search, figure/table interpretation, database access, wet-lab protocols, sequence analysis, cloning scenarios); FutureHouse's predecessor to BixBench, probing biological knowledge and reasoning rather than agentic execution.
Skills benchmarks:
- [SkillsBench](https://github.com/benchflow-ai/skillsbench) - General-purpose benchmark for measuring whether agent skills actually help: 86 tasks across 11 domains (including healthcare), each run under three conditions — no skills, curated skills, and self-generated skills — with deterministic verifiers. Across 7,308 trajectories, curated skills lifted average pass rate by +16.2 pts (ranging from +4.5 pts for software engineering to +51.9 pts for healthcare), while self-generated skills gave no average benefit and focused 23 module skills beat comprehensive documentation. Not bioinformatics-specific, but the canonical "do skills actually work" benchmark. [Paper](https://arxiv.org/abs/2602.12670).
## General AI Coding Agent Skill Collections
Popular repositories of reusable skills for AI coding agents. Each entry is useful for genomics and bioinformatics coding work; descriptions explain how.
Skills specs:
- [agentskills](https://github.com/agentskills/agentskills) - Official specification and documentation for the Agent Skills standard; defines the cross-platform format used by skills in this list and on Claude Code, Copilot, Codex, Cursor, and Gemini CLI.
General skills collections (non-exhaustive):
- [superpowers](https://github.com/obra/superpowers) - Complete software development methodology and skills framework for coding agents; enforces TDD, spec-driven design, and subagent-driven development — practices that improve reproducibility and correctness in complex genomics pipelines.
- [anthropics/skills](https://github.com/anthropics/skills) - Official Anthropic reference implementation and specification for Claude agent skills; the canonical starting point for building custom skills that teach Claude how to handle bioinformatics tasks and lab workflows.
- [openai/skills](https://github.com/openai/skills) - OpenAI's official Skills Catalog for Codex; ships system, curated, and experimental agent skills installable via the in-Codex `$skill-installer`, built on the cross-platform Agent Skills open standard. The OpenAI counterpart to anthropics/skills, and a useful reference for packaging reusable, repeatable bioinformatics coding workflows for Codex.
- [andrej-karpathy-skills](https://github.com/forrestchang/andrej-karpathy-skills) - Single-file coding guidelines derived from Andrej Karpathy's observations on LLM pitfalls; instils simplicity, surgical changes, and goal-driven execution — critical disciplines when AI agents write or modify genomics analysis code.
- [awesome-copilot](https://github.com/github/awesome-copilot) - Community-curated instructions, agents, skills, and configurations for GitHub Copilot, including prompt templates applicable to scientific and data-analysis workflows.
- [awesome-llm-skills](https://github.com/Prat011/awesome-llm-skills) - Curated list of LLM and AI agent skills, resources, and tools for customizing AI agent workflows across Claude Code, Codex, Gemini CLI, and custom agents.
## Other Notable Awesome Lists!
- [Awesome-Scientific-Skills](https://github.com/InternScience/Awesome-Scientific-Skills) - An open, curated collection of agent skills for scientific research spanning bioinformatics, cheminformatics, data analysis, scientific writing, and literature search; currently a curated link collection, with plans to consolidate selected skills into a unified, clone-ready repository. Developed by InternScience, the open-source hub of the AI for Science Center at Shanghai AI Laboratory, which open-sources agents, LLMs/MLLMs, tools, and datasets to accelerate scientific discovery across disciplines.
## Contributing
Contributions are welcome. Please read the [contribution guidelines](CONTRIBUTING.md) before submitting a pull request.
@@ -0,0 +1,390 @@
---
title: "Awesome LLM Scientific Discovery [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)"
task: ""
lineage_type: import
upstream_source: https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery/blob/eb19b47e/README.md
upstream_sha: eb19b47e
imported_at: 2026-07-03
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Awesome LLM Scientific Discovery [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)
A curated list of pioneering research papers, tools, and resources at the intersection of Large Language Models (LLMs) and Scientific Discovery.
Survey: ***From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery.*** ([https://arxiv.org/abs/2505.13259])
The survey delineates the evolving role of LLMs in science through a three-level autonomy framework:
* **Level 1: LLM as Tool:** LLMs augmenting human researchers for specific, well-defined tasks.
* **Level 2: LLM as Analyst:** LLMs exhibiting greater autonomy in processing complex information and offering insights.
* **Level 3: LLM as Scientist:** LLM-based systems autonomously conducting major research stages.
Below is a visual representation of this taxonomy:
![Taxonomy of LLM in Scientific Discovery](taxonomy.png)
We aim to provide a comprehensive overview for researchers, developers, and enthusiasts interested in this rapidly advancing field.
> **Last major update: 2026.07.** This refresh adds a large batch of 20252026 papers and a dedicated section on frontier industry-lab systems (Google DeepMind, OpenAI, Microsoft Research, Meta FAIR, FutureHouse, Sakana AI, and others). Contributions and PRs are very welcome — see [Contributing](#contributing).
## Contents
* [Level 1: LLM as Tool](#level-1-llm-as-tool)
* [Literature Review and Information Gathering](#literature-review-and-information-gathering)
* [Idea Generation and Hypothesis Formulation](#idea-generation-and-hypothesis-formulation)
* [Experiment Planning and Execution](#experiment-planning-and-execution)
* [Data Analysis and Organization](#data-analysis-and-organization)
* [Conclusion and Hypothesis Validation](#conclusion-and-hypothesis-validation)
* [Iteration and Refinement](#iteration-and-refinement)
* [Level 2: LLM as Analyst](#level-2-llm-as-analyst)
* [Machine Learning Research](#machine-learning-research)
* [Data Modeling and Analysis](#data-modeling-and-analysis)
* [Function Discovery](#function-discovery)
* [Natural Science Research](#natural-science-research)
* [General Research](#general-research)
* [Survey Generation](#survey-generation)
* [Level 3: LLM as Scientist](#level-3-llm-as-scientist)
* [General-Purpose Autonomous Research Agents](#general-purpose-autonomous-research-agents)
* [Discovery-Oriented Scientific Systems](#discovery-oriented-scientific-systems)
* [Autonomous Research Ecosystems and Infrastructure](#autonomous-research-ecosystems-and-infrastructure)
* [Frontier Labs and Foundation Models for Science](#frontier-labs-and-foundation-models-for-science)
* [Other Related Works](#other-related-works)
* [Contributing](#contributing)
---
## Level 1: LLM as Tool
At this foundational level, LLMs function as tailored tools under direct human supervision, designed to execute specific, well-defined tasks within a single stage of the scientific method. Their primary goal is to enhance researcher efficiency.
### Literature Review and Information Gathering
Automating literature search, retrieval, synthesis, structuring, and organization.
* **SCIMON : Scientific Inspiration Machines Optimized for Novelty** [![arXiv](https://img.shields.io/badge/arXiv-2305.14259-B31B1B.svg)](https://arxiv.org/pdf/2305.14259) - *Wang et al. (2023.05)*
* **ResearchAgent: Iterative research idea generation over scientific literature with Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2404.07738-B31B1B.svg)](https://arxiv.org/pdf/2404.07738) - *Baek et al. (2024.04)*
* **Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction** [![arXiv](https://img.shields.io/badge/arXiv-2404.14215-B31B1B.svg)](https://arxiv.org/pdf/2404.14215) - *Deng et al. (2024.04)*
* **TKGT: Redefinition and A New Way of text-to-table tasks based on real world demands and knowledge graphs augmented LLMs** [![arXiv](https://img.shields.io/badge/arXiv-2410.emnlp--main.901-B31B1B.svg)](https://aclanthology.org/2024.emnlp-main.901.pdf) - *Jiang et al. (2024.10)*
* **ArxivDIGESTables: Synthesizing scientific literature into tables using language models** [![arXiv](https://img.shields.io/badge/arXiv-2410.22360-B31B1B.svg)](https://arxiv.org/pdf/2410.22360) - *Newman et al. (2024.10)*
* **Can LLMs Generate Tabular Summaries of Science Papers? Rethinking the Evaluation Protocol** [![arXiv](https://img.shields.io/badge/arXiv-2504.10284-B31B1B.svg)](https://arxiv.org/pdf/2504.10284) - *Wang et al. (2025.04)*
* **LitLLM: A Toolkit for Scientific Literature Review** [![arXiv](https://img.shields.io/badge/arXiv-2402.01788v1-B31B1B.svg)](https://arxiv.org/pdf/2402.01788v1) - *Agarwal et al. (2024.02)*
* **Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain** [![DOI](https://img.shields.io/badge/DOI-10.1186/s13643--024--02575--4-blue.svg)](https://systematicreviewsjournal.biomedcentral.com/articles/10.1186/s13643-024-02575-4) - *Dennstädt et al. (2024.06)*
* **Science Hierarchography: Hierarchical Organization of Science Literature** [![arXiv](https://img.shields.io/badge/arXiv-2504.13834-B31B1B.svg)](https://arxiv.org/pdf/2504.13834) - *Gao et al. (2025.04)*
* **Language Agents Achieve Superhuman Synthesis of Scientific Knowledge (PaperQA2)** [![arXiv](https://img.shields.io/badge/arXiv-2409.13740-B31B1B.svg)](https://arxiv.org/pdf/2409.13740) - *Skarlinski et al. (2024.09)*
* **DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents** [![arXiv](https://img.shields.io/badge/arXiv-2506.11763-B31B1B.svg)](https://arxiv.org/pdf/2506.11763) - *Du et al. (2025.06)*
* **Deep Research Agents: A Systematic Examination And Roadmap** [![arXiv](https://img.shields.io/badge/arXiv-2506.18096-B31B1B.svg)](https://arxiv.org/pdf/2506.18096) - *Huang et al. (2025.06)*
* **DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis** [![arXiv](https://img.shields.io/badge/arXiv-2508.20033-B31B1B.svg)](https://arxiv.org/pdf/2508.20033) - *Patel et al. (2025.08)*
* **ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry** [![arXiv](https://img.shields.io/badge/arXiv-2507.16280-B31B1B.svg)](https://arxiv.org/pdf/2507.16280) - *Xu et al. (2025.07)*
* **LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation** [![arXiv](https://img.shields.io/badge/arXiv-2510.05138-B31B1B.svg)](https://arxiv.org/pdf/2510.05138) - *Zhang et al. (2025.10)*
### Idea Generation and Hypothesis Formulation
Automated generation of novel research ideas, conceptual insights, and testable scientific hypotheses.
* **SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2409.05556-B31B1B.svg)](https://arxiv.org/pdf/2409.05556) - *Ghafarollahi et al. (2024.09)*
* **Accelerating scientific discovery with generative knowledge extraction, graph-based representation, and multimodal intelligent graph reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2403.11996-B31B1B.svg)](https://arxiv.org/pdf/2403.11996) - *Buehler (2024.03)*
* **MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses** [![arXiv](https://img.shields.io/badge/arXiv-2410.07076-B31B1B.svg)](https://arxiv.org/pdf/2410.07076) - *Yang et al. (2024.10)*
* **Large Language Models for Automated Open-domain Scientific Hypotheses Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2309.02726-B31B1B.svg)](https://arxiv.org/pdf/2309.02726) - *Yang et al. (2023.09)*
* **Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2411.02382-B31B1B.svg)](https://arxiv.org/pdf/2411.02382) - *Xiong et al. (2024.11)*
* **ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition** [![arXiv](https://img.shields.io/badge/arXiv-2503.21248-B31B1B.svg)](https://arxiv.org/pdf/2503.21248) - *Liu et al. (2025.03)*
* **AI Idea Bench 2025: AI Research Idea Generation Benchmark** [![arXiv](https://img.shields.io/badge/arXiv-2504.14191-B31B1B.svg)](https://arxiv.org/pdf/2504.14191) - *Qiu et al. (2025.04)*
* **IdeaBench: Benchmarking Large Language Models for Research Idea Generation** [![arXiv](https://img.shields.io/badge/arXiv-2411.02429-B31B1B.svg)](https://arxiv.org/pdf/2411.02429) - *Guo et al. (2024.11)*
* **Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers** [![arXiv](https://img.shields.io/badge/arXiv-2409.04109-B31B1B.svg)](https://arxiv.org/pdf/2409.04109) - *Si et al. (2024.09)*
* **Learning to Generate Research Idea with Dynamic Control** [![arXiv](https://img.shields.io/badge/arXiv-2412.14626-B31B1B.svg)](https://arxiv.org/pdf/2412.14626) - *Li et al. (2024.12)*
* **LiveIdeaBench: Evaluating LLMs' Divergent Thinking for Scientific Idea Generation with Minimal Context** [![arXiv](https://img.shields.io/badge/arXiv-2412.17596-B31B1B.svg)](https://arxiv.org/pdf/2412.17596) - *Ruan et al. (2024.12)*
* **Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas** [![arXiv](https://img.shields.io/badge/arXiv-2410.14255-B31B1B.svg)](https://arxiv.org/pdf/2410.14255) - *Hu et al. (2024.10)*
* **GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation** [![arXiv](https://img.shields.io/badge/arXiv-2503.12600-B31B1B.svg)](https://arxiv.org/pdf/2503.12600) - *Feng et al. (2025.03)*
* **Hypothesis Generation with Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2404.04326-B31B1B.svg)](https://arxiv.org/pdf/2404.04326) - *Zhou et al. (2024.04)*
* **Harnessing the Power of Adversarial Prompting and Large Language Models for Robust Hypothesis Generation in Astronomy** [![arXiv](https://img.shields.io/badge/arXiv-2306.11648-B31B1B.svg)](https://arxiv.org/pdf/2306.11648) - *Ciuca et al. (2023.06)*
* **Large Language Models are Zero Shot Hypothesis Proposers** [![arXiv](https://img.shields.io/badge/arXiv-2311.05965-B31B1B.svg)](https://arxiv.org/pdf/2311.05965) - *Qi et al. (2023.11)*
* **Machine learning for hypothesis generation in biology and medicine: exploring the latent space of neuroscience and developmental bioelectricity** [![DOI](https://img.shields.io/badge/DOI-10.1039/D3DD00185G-blue.svg)](https://pubs.rsc.org/en/content/articlelanding/2024/dd/d3dd00185g) - *OBrien et al. (2023.07)*
* **Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation** [![arXiv](https://img.shields.io/badge/arXiv-2407.08940-B31B1B.svg)](https://arxiv.org/pdf/2407.08940) - *Qi et al. (2024.07)*
* **LLM4GRN: Discovering Causal Gene Regulatory Networks with LLMs -- Evaluation through Synthetic Data Generation** [![arXiv](https://img.shields.io/badge/arXiv-2410.15828-B31B1B.svg)](https://arxiv.org/pdf/2410.15828) - *Afonja et al. (2024.10)*
* **Scideator: Human-LLM Scientific Idea Generation Grounded in Research-Paper Facet Recombination** [![arXiv](https://img.shields.io/badge/arXiv-2409.14634-B31B1B.svg)](https://arxiv.org/pdf/2409.14634) - *Radensky et al. (2024.09)*
* **HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance** [![arXiv](https://img.shields.io/badge/arXiv-2506.12937-B31B1B.svg)](https://arxiv.org/pdf/2506.12937) - *Vasu et al. (2025.06)*
* **Sparks of Science: Hypothesis Generation Using Structured Paper Data** [![arXiv](https://img.shields.io/badge/arXiv-2504.12976-B31B1B.svg)](https://arxiv.org/pdf/2504.12976) - *O'Neill et al. (2025.04)*
* **A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2504.05496-B31B1B.svg)](https://arxiv.org/pdf/2504.05496) - *Kulkarni et al. (2025.04)*
* **Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Networks** [![arXiv](https://img.shields.io/badge/arXiv-2511.02238-B31B1B.svg)](https://arxiv.org/pdf/2511.02238) - *Wang et al. (2025.11)*
### Experiment Planning and Execution
LLMs assisting in experimental protocol planning, workflow design, and scientific code generation.
* **BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology** [![arXiv](https://img.shields.io/badge/arXiv-2310.10632-B31B1B.svg)](https://arxiv.org/pdf/2310.10632) - *O'Donoghue et al. (2023.10)*
* **Can Large Language Models Help Experimental Design for Causal Discovery?** (Li et al. in survey) [![arXiv](https://img.shields.io/badge/arXiv-2503.01139-B31B1B.svg)](https://arxiv.org/pdf/2503.01139) - *Li et al. (2025.03)*
* **Hierarchically Encapsulated Representation for Protocol Design in Self-Driving Labs** [![arXiv](https://img.shields.io/badge/arXiv-2504.03810-B31B1B.svg)](https://arxiv.org/pdf/2504.03810) - *Shi et al. (2025.04)*
* **SciCode: A Research Coding Benchmark Curated by Scientists** [![arXiv](https://img.shields.io/badge/arXiv-2407.13168-B31B1B.svg)](https://arxiv.org/pdf/2407.13168) - *Tian et al. (2024.07)*
* **Natural Language to Code Generation in Interactive Data Science Notebooks** [![arXiv](https://img.shields.io/badge/arXiv-2212.09248-B31B1B.svg)](https://arxiv.org/pdf/2212.09248) - *Yin et al. (2022.12)*
* **DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation** [![arXiv](https://img.shields.io/badge/arXiv-2211.11501-B31B1B.svg)](https://arxiv.org/pdf/2211.11501) - *Lai et al. (2022.11)*
* **Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents**, [![arXiv](https://img.shields.io/badge/arXiv-2502.16069-B31B1B.svg)](https://arxiv.org/pdf/2502.16069) - *Kon et al. (2025.02)*
* **AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing** [![arXiv](https://img.shields.io/badge/arXiv-2602.17607-B31B1B.svg)](https://arxiv.org/pdf/2602.17607) - *Du et al. (2026.02)*
### Data Analysis and Organization
LLMs assisting in data-driven analysis, tabular/chart reasoning, statistical reasoning, and model discovery.
* **AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ** [![arXiv](https://img.shields.io/badge/arXiv-2310.00367-B31B1B.svg)](https://arxiv.org/pdf/2310.00367) - *Belouadi et al. (2023.10)*
* **Text2Chart31: Instruction Tuning for Chart Generation with Automatic Feedback** [![arXiv](https://img.shields.io/badge/arXiv-2410.04064-B31B1B.svg)](https://arxiv.org/pdf/2410.04064) - *Zadeh et al. (2024.10)*
* **ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2203.10244-B31B1B.svg)](https://arxiv.org/pdf/2203.10244) - *Masry et al. (2022.03)*
* **CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs** [![arXiv](https://img.shields.io/badge/arXiv-2406.18521-B31B1B.svg)](https://arxiv.org/pdf/2406.18521) - *Wang et al. (2024.06)*
* **ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2402.12185-B31B1B.svg)](https://arxiv.org/pdf/2402.12185) - *Xia et al. (2024.02)*
* **Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding** [![arXiv](https://img.shields.io/badge/arXiv-2401.04398-B31B1B.svg)](https://arxiv.org/pdf/2401.04398) - *Wang et al. (2024.01)*
* **TableBench: A Comprehensive and Complex Benchmark for Table Question Answering** [![arXiv](https://img.shields.io/badge/arXiv-2408.09174-B31B1B.svg)](https://arxiv.org/pdf/2408.09174) - *Wu et al. (2024.08)*
* **Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs** [![arXiv](https://img.shields.io/badge/arXiv-2402.12424-B31B1B.svg)](https://arxiv.org/pdf/2402.12424) - *Deng et al. (2024.02)*
* **ChatSpatial: Schema-Enforced Agentic Orchestration for Reproducible and Cross-Platform Spatial Transcriptomics** [![DOI](https://img.shields.io/badge/DOI-10.64898/2026.02.26.708361-blue.svg)](https://doi.org/10.64898/2026.02.26.708361) - *Yang et al. (2026.02)* [Code](https://github.com/cafferychen777/ChatSpatial)
### Conclusion and Hypothesis Validation
LLMs providing feedback, verifying claims, replicating results, and generating reviews.
* **CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?** [![arXiv](https://img.shields.io/badge/arXiv-2503.21717-B31B1B.svg)](https://arxiv.org/pdf/2503.21717) - *Ou et al. (2025.03)*
* **LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing** [![arXiv](https://img.shields.io/badge/arXiv-2406.16253-B31B1B.svg)](https://arxiv.org/pdf/2406.16253) - *Du et al. (2024.06)*
* **AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews** [![arXiv](https://img.shields.io/badge/arXiv-2408.10365-B31B1B.svg)](https://arxiv.org/pdf/2408.10365) - *Tyser et al. (2024.08)*
* **Is LLM a Reliable Reviewer? A Comprehensive Evaluation of LLM on Automatic Paper Reviewing Tasks** [![Link](https://img.shields.io/badge/Link-LREC--COLING_2024-blue.svg)](https://aclanthology.org/2024.lrec-main.816.pdf) - *Zhou et al. (2024.05)*
* **ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing** [![arXiv](https://img.shields.io/badge/arXiv-2306.00622-B31B1B.svg)](https://arxiv.org/pdf/2306.00622) - *Liu and Shah (2023.06)*
* **Towards Autonomous Hypothesis Verification via Language Models with Minimal Guidance** [![arXiv](https://img.shields.io/badge/arXiv-2311.09706-B31B1B.svg)](https://arxiv.org/pdf/2311.09706) - *Takagi et al. (2023.11)*
* **CycleResearcher: Improving Automated Research via Automated Review** [![arXiv](https://img.shields.io/badge/arXiv-2411.00816-B31B1B.svg)](https://arxiv.org/pdf/2411.00816) - *Weng et al. (2024.11)*
* **PaperBench: Evaluating AIs Ability to Replicate AI Research** [![arXiv](https://img.shields.io/badge/arXiv-2504.01848-B31B1B.svg)](https://arxiv.org/pdf/2504.01848) - *Starace et al. (2025.04)*
* **SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers** [![arXiv](https://img.shields.io/badge/arXiv-2504.00255-B31B1B.svg)](https://arxiv.org/pdf/2504.00255) - *Xiang et al. (2025.04)*
* **Advancing AI-Scientist Understanding: Making LLM Think Like a Physicist with Interpretable Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2504.01911-B31B1B.svg)](https://arxiv.org/pdf/2504.01911) - *Xu et al. (2025.04)*
* **Generative Adversarial Reviews: When LLMs Become the Critic** [![arXiv](https://img.shields.io/badge/arXiv-2412.10415-B31B1B.svg)](https://arxiv.org/pdf/2412.10415) - *Bougie & Watanabe (2024.12)*
* **Predicting Empirical AI Research Outcomes with Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2506.00794-B31B1B.svg)](https://arxiv.org/pdf/2506.00794) - *Wen et al. (2025.06)*
* **SPOT: When AI Co-Scientists Fail — A Benchmark for Automated Verification of Scientific Research** [![arXiv](https://img.shields.io/badge/arXiv-2505.11855-B31B1B.svg)](https://arxiv.org/pdf/2505.11855) - *Son et al. (2025.05)*
* **DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process** [![arXiv](https://img.shields.io/badge/arXiv-2503.08569-B31B1B.svg)](https://arxiv.org/pdf/2503.08569) - *Zhu et al. (2025.03)*
* **ReviewRL: Towards Automated Scientific Review with RL** [![arXiv](https://img.shields.io/badge/arXiv-2508.10308-B31B1B.svg)](https://arxiv.org/pdf/2508.10308) - *Zeng et al. (2025.08)*
* **SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification** [![arXiv](https://img.shields.io/badge/arXiv-2502.10003-B31B1B.svg)](https://arxiv.org/pdf/2502.10003) - *Kumar et al. (2025.02)*
* **LMR-Bench: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research** [![arXiv](https://img.shields.io/badge/arXiv-2506.17335-B31B1B.svg)](https://arxiv.org/pdf/2506.17335) - *Yan et al. (2025.06)*
* **REFUTE: Reasoning Over Evidence - Falsification, Uncertainty, Truth-grounding & Epistemics** [![HF Dataset](https://img.shields.io/badge/HuggingFace-dataset-yellow.svg)](https://huggingface.co/datasets/BGPT-OFFICIAL/refute) - *BGPT (2026.06)*. Open benchmark for scientific critique and epistemic calibration on recent science paper summaries, covering falsification, limitations, overclaims, missing-evidence refusal, calibration, and planted-flaw detection.
### Iteration and Refinement
LLMs involved in iterative refinement of research hypotheses and strategic exploration.
* **Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving** [![arXiv](https://img.shields.io/badge/arXiv-2405.01379-B31B1B.svg)](https://arxiv.org/pdf/2405.01379) - *Quan et al. (2024.05)*
* **Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents** [![arXiv](https://img.shields.io/badge/arXiv-2410.13185-B31B1B.svg)](https://arxiv.org/pdf/2410.13185) - *Li et al. (2024.10)*
* **Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees** [![arXiv](https://img.shields.io/badge/arXiv-2503.19309-B31B1B.svg)](https://arxiv.org/pdf/2503.19309) - *Rabby et al. (2025.03)*
* **XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision** [![arXiv](https://img.shields.io/badge/arXiv-2505.11336-B31B1B.svg)](https://arxiv.org/pdf/2505.11336) - *Chen et al. (2025.05)*
---
## Level 2: LLM as Analyst
LLMs exhibiting a greater degree of autonomy, functioning as passive agents capable of complex information processing, data modeling, and analytical reasoning with reduced human intervention.
### Machine Learning Research
Automated modeling of machine learning tasks, experiment design, and execution.
* **MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation** [![arXiv](https://img.shields.io/badge/arXiv-2310.03302-B31B1B.svg)](https://arxiv.org/pdf/2310.03302) - *Huang et al. (2023.10)*
* **MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents** [![arXiv](https://img.shields.io/badge/arXiv-2408.14033-B31B1B.svg)](https://arxiv.org/pdf/2408.14033) - *Li et al. (2024.08)*
* **MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering** [![arXiv](https://img.shields.io/badge/arXiv-2410.07095-B31B1B.svg)](https://arxiv.org/pdf/2410.07095) - *Chan et al. (2024.10)*
* **IMPROVE: Iterative Model Pipeline Refinement and Optimization Leveraging LLM Agents** [![arXiv](https://img.shields.io/badge/arXiv-2502.18530v1-B31B1B.svg)](https://arxiv.org/pdf/2502.18530v1) - *Xue et al. (2025.02)*
* **CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation** [![arXiv](https://img.shields.io/badge/arXiv-2503.22708-B31B1B.svg)](https://arxiv.org/pdf/2503.22708) - *Jansen et al. (2025.03)*
* **MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?** [![arXiv](https://img.shields.io/badge/arXiv-2504.09702-B31B1B.svg)](https://arxiv.org/pdf/2504.09702) - *Zhang et al. (2025.04)*
* **RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts** [![arXiv](https://img.shields.io/badge/arXiv-2411.15114-B31B1B.svg)](https://arxiv.org/pdf/2411.15114) - *Wijk et al. (2024.11)*
* **MLZero: A Multi-Agent System for End-to-end Machine Learning Automation** [![arXiv](https://img.shields.io/badge/arXiv-2505.13941-B31B1B.svg)](https://arxiv.org/pdf/2505.13941) - *Fang et al. (2025.05)*
* **AIDE: AI-Driven Exploration in the Space of Code** [![arXiv](https://img.shields.io/badge/arXiv-2502.13138-B31B1B.svg)](https://arxiv.org/pdf/2502.13138) - *Jiang et al. (2025.02)*
* **Language Modeling by Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2506.20249-B31B1B.svg)](https://arxiv.org/pdf/2506.20249) - *Cheng et al. (2025.06)*
* **MLGym: A New Framework and Benchmark for Advancing AI Research Agents** [![arXiv](https://img.shields.io/badge/arXiv-2502.14499-B31B1B.svg)](https://arxiv.org/pdf/2502.14499) - *Nathani et al. (2025.02)*
* **R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science** [![arXiv](https://img.shields.io/badge/arXiv-2505.14738-B31B1B.svg)](https://arxiv.org/pdf/2505.14738) - *Xu et al. (2025.05)* — Microsoft Research
* **MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement** [![arXiv](https://img.shields.io/badge/arXiv-2506.15692-B31B1B.svg)](https://arxiv.org/pdf/2506.15692) - *Nam et al. (2025.06)* — Google
* **ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2506.16499-B31B1B.svg)](https://arxiv.org/pdf/2506.16499) - *Liu et al. (2025.06)*
* **ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering** [![arXiv](https://img.shields.io/badge/arXiv-2505.23723-B31B1B.svg)](https://arxiv.org/pdf/2505.23723) - *Liu et al. (2025.05)*
* **AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench** [![arXiv](https://img.shields.io/badge/arXiv-2507.02554-B31B1B.svg)](https://arxiv.org/pdf/2507.02554) - *Toledo et al. (2025.07)* — Meta / UCL
* **The FM Agent** [![arXiv](https://img.shields.io/badge/arXiv-2510.26144-B31B1B.svg)](https://arxiv.org/pdf/2510.26144) - *Li et al. (2025.10)*
* **KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for ML Problems** [![arXiv](https://img.shields.io/badge/arXiv-2508.10177-B31B1B.svg)](https://arxiv.org/pdf/2508.10177) - *Kulibaba et al. (2025.08)*
* **AutoMLGen: Navigating Fine-Grained Optimization for Coding Agents** [![arXiv](https://img.shields.io/badge/arXiv-2510.08511-B31B1B.svg)](https://arxiv.org/pdf/2510.08511) - *Du et al. (2025.10)*
* **ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies** [![arXiv](https://img.shields.io/badge/arXiv-2504.20117-B31B1B.svg)](https://arxiv.org/pdf/2504.20117) - *Gandhi et al. (2025.04)*
* **AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML** [![arXiv](https://img.shields.io/badge/arXiv-2410.02958-B31B1B.svg)](https://arxiv.org/pdf/2410.02958) - *Trirat et al. (2024.10)*
* **SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning** [![arXiv](https://img.shields.io/badge/arXiv-2410.17238-B31B1B.svg)](https://arxiv.org/pdf/2410.17238) - *Chi et al. (2024.10)*
* **AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions** [![arXiv](https://img.shields.io/badge/arXiv-2410.20424-B31B1B.svg)](https://arxiv.org/pdf/2410.20424) - *Li et al. (2024.10)*
* **Agent K: Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Performance** [![arXiv](https://img.shields.io/badge/arXiv-2411.03562-B31B1B.svg)](https://arxiv.org/pdf/2411.03562) - *Grosnit et al. (2024.11)* — Huawei Noah's Ark
* **EXP-Bench: Can AI Conduct AI Research Experiments?** [![arXiv](https://img.shields.io/badge/arXiv-2505.24785-B31B1B.svg)](https://arxiv.org/pdf/2505.24785) - *Kon et al. (2025.05)*
* **InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research** [![arXiv](https://img.shields.io/badge/arXiv-2510.27598-B31B1B.svg)](https://arxiv.org/pdf/2510.27598) - *Wu et al. (2025.10)*
* **MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research** [![arXiv](https://img.shields.io/badge/arXiv-2505.19955-B31B1B.svg)](https://arxiv.org/pdf/2505.19955) - *Chen et al. (2025.05)*
* **RExBench: Can Coding Agents Autonomously Implement AI Research Extensions?** [![arXiv](https://img.shields.io/badge/arXiv-2506.22598-B31B1B.svg)](https://arxiv.org/pdf/2506.22598) - *Edwards et al. (2025.06)*
* **ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution** [![arXiv](https://img.shields.io/badge/arXiv-2509.19349-B31B1B.svg)](https://arxiv.org/pdf/2509.19349) - *Lange et al. (2025.09)* — Sakana AI
* **The AI CUDA Engineer: Agentic CUDA Kernel Discovery, Optimization and Composition** [![Link](https://img.shields.io/badge/Link-Sakana_Report-blue.svg)](https://pub.sakana.ai/ai-cuda-engineer/paper/) - *Sakana AI (2025.02)*
### Data Modeling and Analysis
Automated data-driven analysis, statistical data modeling, and hypothesis validation.
* **Automated Statistical Model Discovery with Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2402.17879-B31B1B.svg)](https://arxiv.org/pdf/2402.17879) - *Li et al. (2024.02)*
* **InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks** [![arXiv](https://img.shields.io/badge/arXiv-2401.05507-B31B1B.svg)](https://arxiv.org/pdf/2401.05507) - *Hu et al. (2024.01)*
* **DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2402.17453-B31B1B.svg)](https://arxiv.org/pdf/2402.17453) - *Guo et al. (2024.02)*
* **BLADE: Benchmarking Language Model Agents for Data-Driven Science** [![arXiv](https://img.shields.io/badge/arXiv-2408.09667-B31B1B.svg)](https://arxiv.org/pdf/2408.09667) - *Gu et al. (2024.08)*
* **DAgent: A Relational Database-Driven Data Analysis Report Generation Agent** [![arXiv](https://img.shields.io/badge/arXiv-2503.13269-B31B1B.svg)](https://arxiv.org/pdf/2503.13269) - *Xu et al. (2025.03)*
* **DiscoveryBench: Towards Data-Driven Discovery with Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2407.01725-B31B1B.svg)](https://arxiv.org/pdf/2407.01725) - *Majumder et al. (2024.07)*
* **Large Language Models for Scientific Synthesis, Inference and Explanation** [![arXiv](https://img.shields.io/badge/arXiv-2310.07984-B31B1B.svg)](https://arxiv.org/pdf/2310.07984) - *Zheng et al. (2023.10)*
* **MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem** [![arXiv](https://img.shields.io/badge/arXiv-2505.14148-B31B1B.svg)](https://arxiv.org/pdf/2505.14148) - *Liu et al. (2025.05)*
* **DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?** [![arXiv](https://img.shields.io/badge/arXiv-2409.07703-B31B1B.svg)](https://arxiv.org/pdf/2409.07703) - *Jing et al. (2024.09)*
* **AutoDS: Open-ended Scientific Discovery via Bayesian Surprise** [![arXiv](https://img.shields.io/badge/arXiv-2507.00310-B31B1B.svg)](https://arxiv.org/pdf/2507.00310) - *Agarwal et al. (2025.07)* — Allen Institute for AI
* **DeepAnalyze: Agentic Large Language Models for Autonomous Data Science** [![arXiv](https://img.shields.io/badge/arXiv-2510.16872-B31B1B.svg)](https://arxiv.org/pdf/2510.16872) - *Zhang et al. (2025.10)*
* **DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2410.07331-B31B1B.svg)](https://arxiv.org/pdf/2410.07331) - *Huang et al. (2024.10)*
* **DataSciBench: An LLM Agent Benchmark for Data Science** [![arXiv](https://img.shields.io/badge/arXiv-2502.13897-B31B1B.svg)](https://arxiv.org/pdf/2502.13897) - *Zhang et al. (2025.02)*
* **Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents** [![arXiv](https://img.shields.io/badge/arXiv-2403.05307-B31B1B.svg)](https://arxiv.org/pdf/2403.05307) - *Li et al. (2024.03)*
* **StatEval: A Comprehensive Benchmark for Large Language Models in Statistics** [![arXiv](https://img.shields.io/badge/arXiv-2510.09517-B31B1B.svg)](https://arxiv.org/pdf/2510.09517) - *Yu et al. (2025.10)*
* **LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference** [![arXiv](https://img.shields.io/badge/arXiv-2508.07221-B31B1B.svg)](https://arxiv.org/pdf/2508.07221) - *Wang et al. (2025.08)*
* **OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents** [![arXiv](https://img.shields.io/badge/arXiv-2504.16918-B31B1B.svg)](https://arxiv.org/pdf/2504.16918) - *Thind et al. (2025.04)*
### Function Discovery
Identifying underlying equations from observational data (AI-driven symbolic regression).
* **LLM-SR: Scientific Equation Discovery via Programming with Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2404.18400-B31B1B.svg)](https://arxiv.org/pdf/2404.18400) - *Shojaee et al. (2024.04)*
* **LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2504.10415-B31B1B.svg)](https://arxiv.org/pdf/2504.10415) - *Shojaee et al. (2025.04)*
* **Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents** [![arXiv](https://img.shields.io/badge/arXiv-2501.18411-B31B1B.svg)](https://arxiv.org/pdf/2501.18411) - *Koblischke et al. (2025.01)*
* **EvoSLD: Automated neural scaling law discovery with large language models** [![arXiv](https://img.shields.io/badge/arXiv-2507.21184-B31B1B.svg)](https://arxiv.org/abs/2507.21184) - *Lin et al. (2025.07)*
* **DrSR: LLM based Scientific Equation Discovery with Dual Reasoning from Data and Experience** [![arXiv](https://img.shields.io/badge/arXiv-2506.04282-B31B1B.svg)](https://arxiv.org/abs/2506.04282) - *Wang et al. (2025.06)*
* **NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents** [![arXiv](https://img.shields.io/badge/arXiv-2510.07172-B31B1B.svg)](https://arxiv.org/pdf/2510.07172) - *Zheng et al. (2025.10)*
* **LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2503.06512-B31B1B.svg)](https://arxiv.org/pdf/2503.06512) - *Song et al. (2025.03)*
* **SR-Scientist: Scientific Equation Discovery With Agentic AI** [![arXiv](https://img.shields.io/badge/arXiv-2510.11661-B31B1B.svg)](https://arxiv.org/pdf/2510.11661) - *Xia et al. (2025.10)*
* **LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery (SGA)** [![arXiv](https://img.shields.io/badge/arXiv-2405.09783-B31B1B.svg)](https://arxiv.org/pdf/2405.09783) - *Ma et al. (2024.05)*
* **In-Context Symbolic Regression: Leveraging Large Language Models for Function Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2404.19094-B31B1B.svg)](https://arxiv.org/pdf/2404.19094) - *Merler et al. (2024.04)*
* **Symbolic Regression with a Learned Concept Library (LaSR)** [![arXiv](https://img.shields.io/badge/arXiv-2409.09359-B31B1B.svg)](https://arxiv.org/pdf/2409.09359) - *Grayeli et al. (2024.09)*
* **AI-Newton: A Concept-Driven Physical Law Discovery System without Prior Physical Knowledge** [![arXiv](https://img.shields.io/badge/arXiv-2504.01538-B31B1B.svg)](https://arxiv.org/pdf/2504.01538) - *Fang et al. (2025.04)*
* **PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors** [![arXiv](https://img.shields.io/badge/arXiv-2507.15550-B31B1B.svg)](https://arxiv.org/pdf/2507.15550) - *Chen et al. (2025.07)*
* **Finetuning Large Language Model as an Effective Symbolic Regressor (SymbArena)** [![arXiv](https://img.shields.io/badge/arXiv-2508.09897-B31B1B.svg)](https://arxiv.org/pdf/2508.09897) - *Hua et al. (2025.08)*
### Natural Science Research
Autonomous research workflows for natural science discovery (e.g., chemistry, biology, biomedicine, materials, physics).
* **Coscientist: Autonomous Chemical Research with Large Language Models** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--023--06792--0-blue.svg)](https://www.nature.com/articles/s41586-023-06792-0) - *Boiko et al. (2023.10)*
* **Empowering biomedical discovery with AI agents** [![DOI](https://img.shields.io/badge/DOI-10.1016/j.cell.2024.08.026-blue.svg)](https://www.cell.com/action/showPdf?pii=S0092-8674%2824%2901070-5) - *Gao et al. (2024.09)*
* **GenoTEX: An LLM Agent Benchmark for Automated Gene Expression Data Analysis** [![arXiv](https://img.shields.io/badge/arXiv-2406.15341-B31B1B.svg)](https://arxiv.org/pdf/2406.15341) - *Liu et al. (2024.06)*
* **From Intention To Implementation: Automating Biomedical Research via LLMs** [![arXiv](https://img.shields.io/badge/arXiv-2412.09429-B31B1B.svg)](https://arxiv.org/pdf/2412.09429) - *Luo et al. (2024.12)*
* **DrugAgent: Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration** [![arXiv](https://img.shields.io/badge/arXiv-2411.15692-B31B1B.svg)](https://arxiv.org/pdf/2411.15692) - *Liu et al. (2024.11)*
* **ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2410.05080-B31B1B.svg)](https://arxiv.org/pdf/2410.05080) - *Chen et al. (2024.10)*
* **ProtAgents: Protein discovery by combining physics and machine learning** [![arXiv](https://img.shields.io/badge/arXiv-2402.04268-B31B1B.svg)](https://arxiv.org/pdf/2402.04268) - *Ghafarollahi and Buehler (2024.02)*
* **Auto-Bench: An Automated Benchmark for Scientific Discovery in LLMs** [![arXiv](https://img.shields.io/badge/arXiv-2502.15224-B31B1B.svg)](https://arxiv.org/pdf/2502.15224) - *Chen et al. (2025.02)*
* **Towards an AI co-scientist** [![arXiv](https://img.shields.io/badge/arXiv-2502.18864-B31B1B.svg)](https://arxiv.org/pdf/2502.18864) - *Gottweis et al. (2025.02)* — Google
* **GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis** [![arXiv](https://img.shields.io/badge/arXiv-2507.21035-B31B1B.svg)](https://arxiv.org/pdf/2507.21035) - *Liu et al. (2025.07)*
* **Automated Algorithmic Discovery for Gravitational-Wave Detection Guided by LLM-Informed Evolutionary Monte Carlo Tree Search** [![arXiv](https://img.shields.io/badge/arXiv-2508.03661-B31B1B.svg)](https://arxiv.org/pdf/2508.03661) - *Wang and Zeng (2025.08)*
* **The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--025--09442--9-blue.svg)](https://doi.org/10.1038/s41586-025-09442-9) - *Swanson et al. (2025.07)* — Stanford / CZ Biohub
* **Biomni: A General-Purpose Biomedical AI Agent** [![DOI](https://img.shields.io/badge/DOI-10.1101/2025.05.30.656746-blue.svg)](https://doi.org/10.1101/2025.05.30.656746) - *Huang et al. (2025.05)* — Stanford
* **TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools** [![arXiv](https://img.shields.io/badge/arXiv-2503.10970-B31B1B.svg)](https://arxiv.org/pdf/2503.10970) - *Gao et al. (2025.03)*
* **LIDDiA: Language-based Intelligent Drug Discovery Agent** [![arXiv](https://img.shields.io/badge/arXiv-2502.13959-B31B1B.svg)](https://arxiv.org/pdf/2502.13959) - *Averly et al. (2025.02)*
* **LLM Agent Swarm for Hypothesis-Driven Drug Discovery (PharmaSwarm)** [![arXiv](https://img.shields.io/badge/arXiv-2504.17967-B31B1B.svg)](https://arxiv.org/pdf/2504.17967) - *Song et al. (2025.04)*
* **BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation** [![arXiv](https://img.shields.io/badge/arXiv-2508.01285-B31B1B.svg)](https://arxiv.org/pdf/2508.01285) - *Ke et al. (2025.08)*
* **CRISPR-GPT for Agentic Automation of Gene-editing Experiments** [![arXiv](https://img.shields.io/badge/arXiv-2404.18021-B31B1B.svg)](https://arxiv.org/pdf/2404.18021) - *Qu et al. (2024.04)*
* **BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments** [![arXiv](https://img.shields.io/badge/arXiv-2405.17631-B31B1B.svg)](https://arxiv.org/pdf/2405.17631) - *Roohani et al. (2024.05)*
* **CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis** [![arXiv](https://img.shields.io/badge/arXiv-2407.09811-B31B1B.svg)](https://arxiv.org/pdf/2407.09811) - *Xiao et al. (2024.07)*
* **Training a Scientific Reasoning Model for Chemistry (ether0)** [![arXiv](https://img.shields.io/badge/arXiv-2506.17238-B31B1B.svg)](https://arxiv.org/pdf/2506.17238) - *Narayanan et al. (2025.06)* — FutureHouse
* **AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation** [![arXiv](https://img.shields.io/badge/arXiv-2509.25651-B31B1B.svg)](https://arxiv.org/pdf/2509.25651) - *Panapitiya et al. (2025.09)*
* **LLMatDesign: Autonomous Materials Discovery with Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2406.13163-B31B1B.svg)](https://arxiv.org/pdf/2406.13163) - *Jia et al. (2024.06)*
* **Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists** [![arXiv](https://img.shields.io/badge/arXiv-2506.05616-B31B1B.svg)](https://arxiv.org/pdf/2506.05616) - *Zhou et al. (2025.06)*
* **SparksMatter: Autonomous Inorganic Materials Discovery via Multi-Agent Physics-Aware Scientific Reasoning** [![arXiv](https://img.shields.io/badge/arXiv-2508.02956-B31B1B.svg)](https://arxiv.org/pdf/2508.02956) - *Ghafarollahi et al. (2025.08)*
* **Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation** [![arXiv](https://img.shields.io/badge/arXiv-2511.22311-B31B1B.svg)](https://arxiv.org/pdf/2511.22311) - *Wang et al. (2025.11)*
* **BixBench: A Comprehensive Benchmark for LLM-based Agents in Computational Biology** [![arXiv](https://img.shields.io/badge/arXiv-2503.00096-B31B1B.svg)](https://arxiv.org/pdf/2503.00096) - *Mitchener et al. (2025.02)* — FutureHouse
* **CASSIA: a multi-agent large language model for automated and interpretable cell annotation** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41467--025--67084--x-blue.svg)](https://www.nature.com/articles/s41467-025-67084-x) - *Xie et al. (2025.12)*
* **AutoZyme: An Autonomous Agentic Framework to Optimize Bioinformatics Software** [![bioRxiv](https://img.shields.io/badge/bioRxiv-2026.06-b31b1b.svg)](https://www.biorxiv.org/content/10.64898/2026.06.12.731250v1) - *Xie et al. (2026.06)*
### General Research
Benchmarks and frameworks evaluating diverse tasks from different stages of scientific discovery.
* **DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents** [![arXiv](https://img.shields.io/badge/arXiv-2406.06769-B31B1B.svg)](https://arxiv.org/pdf/2406.06769) - *Jansen et al. (2024.06)*
* **A Vision for Auto Research with LLM Agents** [![arXiv](https://img.shields.io/badge/arXiv-2504.18765-B31B1B.svg)](https://arxiv.org/pdf/2504.18765) - *Liu et al. (2025.04)*
* **Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents** [![arXiv](https://img.shields.io/badge/arXiv-2502.16069-B31B1B.svg)](https://arxiv.org/abs/2502.16069) - *Kon et al. (2025.02)*
* **EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants** [![arXiv](https://img.shields.io/badge/arXiv-2502.20309-B31B1B.svg)](https://arxiv.org/pdf/2502.20309) - *Cappello et al. (2025.02)*
### Survey Generation
* **AutoSurvey: Large Language Models Can Automatically Write Surveys** [![arXiv](https://img.shields.io/badge/arXiv-2406.10252-B31B1B.svg)](https://arxiv.org/pdf/2406.10252) - *Wang et al. (2024.06)*
* **SurveyX: Academic Survey Automation via Large Language Models** [![arXiv](https://img.shields.io/badge/arXiv-2502.14776-B31B1B.svg)](https://arxiv.org/pdf/2502.14776) - *Liang et al. (2025.02)*
---
## Level 3: LLM as Scientist
LLM-based systems operating as active agents capable of orchestrating and navigating multiple stages of the scientific discovery process with considerable independence, often culminating in draft research papers or genuine new findings. As the field has matured, these systems increasingly fall into distinct classes, reflected in the sub-sections below.
### General-Purpose Autonomous Research Agents
End-to-end pipelines that autonomously move from ideation through experimentation to a full paper draft, typically domain-agnostic (frequently demonstrated on ML/AI research).
* **Agent Laboratory: Using LLM Agents as Research Assistants** [![arXiv](https://img.shields.io/badge/arXiv-2501.04227-B31B1B.svg)](https://arxiv.org/pdf/2501.04227) - *Schmidgall et al. (2025.01)*
* **The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2408.06292-B31B1B.svg)](https://arxiv.org/pdf/2408.06292) - *Lu et al. (2024.08)* — Sakana AI
* **The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search** [![arXiv](https://img.shields.io/badge/arXiv-2504.08066-B31B1B.svg)](https://arxiv.org/pdf/2504.08066) - *Yamada et al. (2025.04)* — Sakana AI
* **AI-Researcher: Autonomous Scientific Innovation** [![arXiv](https://img.shields.io/badge/arXiv-2505.18705-B31B1B.svg)](https://arxiv.org/pdf/2505.18705) [![GitHub](https://img.shields.io/badge/GitHub-HKUDS/AI--Researcher-blue.svg)](https://github.com/HKUDS/AI-Researcher) - *Tang et al. (2025.05)*
* **Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback** [![arXiv](https://img.shields.io/badge/arXiv-2501.03916-B31B1B.svg)](https://arxiv.org/pdf/2501.03916) - *Yuan et al. (2025.01)*
* **NovelSeek / InternAgent: When Agent Becomes the Scientist — Building a Closed-Loop System from Hypothesis to Verification** [![arXiv](https://img.shields.io/badge/arXiv-2505.16938-B31B1B.svg)](https://arxiv.org/pdf/2505.16938) - *InternAgent Team (2025.05)* — Shanghai AI Lab
* **The Denario Project: Deep Knowledge AI Agents for Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2510.26887-B31B1B.svg)](https://arxiv.org/pdf/2510.26887) - *Villaescusa-Navarro et al. (2025.10)*
* **Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation (freephdlabor)** [![arXiv](https://img.shields.io/badge/arXiv-2510.15624-B31B1B.svg)](https://arxiv.org/pdf/2510.15624) - *Li et al. (2025.10)*
* **AIGS: Generating Science from AI-Powered Automated Falsification** [![arXiv](https://img.shields.io/badge/arXiv-2411.11910-B31B1B.svg)](https://arxiv.org/pdf/2411.11910) - *Liu et al. (2024.11)*
* **Zochi Technical Report** [![Link](https://img.shields.io/badge/Link-Intology.AI-blue.svg)](https://www.intology.ai/blog/zochi-tech-report) - *Intology AI (2025.03)*
* **Meet Carl: The First AI System To Produce Academically Peer-Reviewed Research** [![Link](https://img.shields.io/badge/Link-AutoScience.AI-blue.svg)](https://www.autoscience.ai/blog/meet-carl-the-first-ai-system-to-produce-academically-peer-reviewed-research) - *Autoscience Institute (2025.03)*
* **DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively** [![arXiv](https://img.shields.io/badge/arXiv-2509.26603-B31B1B.svg)](https://arxiv.org/pdf/2509.26603) - *Weng et al. (2025.09)*
* **Accelerating Social Science Research via Agentic Hypothesization and Experimentation** [![arXiv](https://img.shields.io/badge/arXiv-2602.07983-B31B1B.svg)](https://arxiv.org/pdf/2602.07983) - *Gupta et al. (2026.02)*
* **AI-Researcher: Fully-Automated Scientific Discovery with LLM Agents** [![GitHub](https://img.shields.io/badge/GitHub-HKUDS/AI--Researcher-blue.svg)](https://github.com/HKUDS/AI-Researcher) - *Data Intelligence Lab (2025.03)*
### Discovery-Oriented Scientific Systems
Systems whose primary goal is genuine new scientific knowledge — novel, experimentally- or mathematically-validated findings — rather than paper drafts. Many are frontier industry-lab systems.
* **Kosmos: An AI Scientist for Autonomous Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2511.02824-B31B1B.svg)](https://arxiv.org/pdf/2511.02824) - *Mitchener et al. (2025.11)* — Edison Scientific / FutureHouse
* **Robin: A Multi-Agent System for Automating Scientific Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2505.13400-B31B1B.svg)](https://arxiv.org/pdf/2505.13400) - *Ghareeb et al. (2025.05)* — FutureHouse
* **AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery** [![arXiv](https://img.shields.io/badge/arXiv-2506.13131-B31B1B.svg)](https://arxiv.org/pdf/2506.13131) - *Novikov et al. (2025.06)* — Google DeepMind
* **Aviary: Training Language Agents on Challenging Scientific Tasks** [![arXiv](https://img.shields.io/badge/arXiv-2412.21154-B31B1B.svg)](https://arxiv.org/pdf/2412.21154) - *Narayanan et al. (2024.12)* — FutureHouse
### Autonomous Research Ecosystems and Infrastructure
Platforms and protocols that enable multiple AI scientists to collaborate, share, review, and publish — moving beyond a single agent toward a research ecosystem.
* **AgentRxiv: Towards Collaborative Autonomous Research** [![arXiv](https://img.shields.io/badge/arXiv-2503.18102-B31B1B.svg)](https://arxiv.org/pdf/2503.18102) - *Schmidgall et al. (2025.03)*
* **aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists** [![arXiv](https://img.shields.io/badge/arXiv-2508.15126-B31B1B.svg)](https://arxiv.org/pdf/2508.15126) - *Zhang et al. (2025.08)*
---
## Frontier Labs and Foundation Models for Science
Flagship "AI for Science" systems from frontier industry labs. Unlike the agentic systems catalogued above, most of these are large domain-specific foundation models or specialized reasoning systems that have driven headline scientific results (structure prediction, materials/genome design, olympiad-level mathematics). They are included here as essential context for the broader landscape of AI-accelerated discovery.
* **Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--024--07487--w-blue.svg)](https://www.nature.com/articles/s41586-024-07487-w) - *Abramson et al. (2024.05)* — Google DeepMind / Isomorphic Labs
* **Scaling Deep Learning for Materials Discovery (GNoME)** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--023--06735--9-blue.svg)](https://www.nature.com/articles/s41586-023-06735-9) - *Merchant et al. (2023.11)* — Google DeepMind
* **A Generative Model for Inorganic Materials Design (MatterGen)** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--025--08628--5-blue.svg)](https://www.nature.com/articles/s41586-025-08628-5) - *Zeni et al. (2025.01)* — Microsoft Research
* **Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models** [![arXiv](https://img.shields.io/badge/arXiv-2410.12771-B31B1B.svg)](https://arxiv.org/pdf/2410.12771) - *Barroso-Luque et al. (2024.10)* — Meta FAIR
* **TamGen: Drug Design with Target-Aware Molecule Generation through a Chemical Language Model** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41467--024--53632--4-blue.svg)](https://www.nature.com/articles/s41467-024-53632-4) - *Wu et al. (2024.10)* — Microsoft Research
* **Genome Modeling and Design Across All Domains of Life with Evo 2** [![DOI](https://img.shields.io/badge/DOI-10.1101/2025.02.18.638918-blue.svg)](https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1) - *Brixi et al. (2025.02)* — Arc Institute / NVIDIA
* **Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning (AlphaProof)** [![DOI](https://img.shields.io/badge/DOI-10.1038/s41586--025--09833--y-blue.svg)](https://www.nature.com/articles/s41586-025-09833-y) - *Hubert et al. (2025.11)* — Google DeepMind
* **Gold-Medalist Performance in Solving Olympiad Geometry with AlphaGeometry 2** [![arXiv](https://img.shields.io/badge/arXiv-2502.03544-B31B1B.svg)](https://arxiv.org/pdf/2502.03544) - *Chervonyi et al. (2025.02)* — Google DeepMind
* **Chai-2: Drug-Like Antibody Design Against Challenging Targets with Atomic Precision** [![Link](https://img.shields.io/badge/Link-Chai_Report-blue.svg)](https://chaiassets.com/chai-2/paper/technical_report_challenging_targets.pdf) - *Chai Discovery (2025.11)*
---
## Other Related Works
* **NVIDIA BioNeMo Agent Toolkit — Tools for Agents to Accelerate Scientific Discovery** [![Link](https://img.shields.io/badge/Link-NVIDIA_Newsroom-blue.svg)](https://nvidianews.nvidia.com/news/nvidia-launches-bionemo-agent-toolkit-giving-ai-agents-the-tools-to-accelerate-scientific-discovery) - *NVIDIA (2025)*
---
## Contributing
Contributions are welcome! If you have a paper, tool, or resource that fits into this taxonomy, please submit a **pull request**.
When adding an entry, please:
* Place it under the most appropriate level/sub-section.
* Keep the format consistent: `**Title** [badge](link) - *First author et al. (YYYY.MM)*`.
* Verify the arXiv ID / DOI resolves, and add the lab/affiliation when the work comes from an industry group.
---
## Citation
Please cite our paper if you found our survey helpful:
```bibtex
@misc{zheng2025automationautonomysurveylarge,
title={From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery},
author={Tianshi Zheng and Zheye Deng and Hong Ting Tsang and Weiqi Wang and Jiaxin Bai and Zihao Wang and Yangqiu Song},
year={2025},
eprint={2505.13259},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2505.13259},
}
```
@@ -0,0 +1,207 @@
---
title: "Contributing to Awesome AI for Science"
task: ""
lineage_type: import
upstream_source: https://github.com/ai-boost/awesome-ai-for-science/blob/cf292eeb/CONTRIBUTING.md
upstream_sha: cf292eeb
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Contributing to Awesome AI for Science
Thank you for your interest in contributing to Awesome AI for Science! 🎉
This project aims to be the most comprehensive and up-to-date collection of AI resources for scientific research. Your contributions help researchers worldwide discover tools and knowledge that can accelerate scientific discovery.
## 📋 Table of Contents
- [How to Contribute](#how-to-contribute)
- [Types of Contributions](#types-of-contributions)
- [Contribution Guidelines](#contribution-guidelines)
- [Formatting Guidelines](#formatting-guidelines)
- [Review Process](#review-process)
- [Code of Conduct](#code-of-conduct)
## 🤝 How to Contribute
### Quick Contribution (for small additions)
1. **Fork** this repository
2. **Edit** the README.md file directly in GitHub
3. **Add** your resource in the appropriate section
4. **Submit** a pull request
### Detailed Contribution (for larger changes)
1. **Fork** this repository to your GitHub account
2. **Clone** your fork locally:
```bash
git clone https://github.com/your-username/awesome-ai-for-science.git
cd awesome-ai-for-science
```
3. **Create** a new branch for your contribution:
```bash
git checkout -b add-new-resource
```
4. **Make** your changes to the README.md file
5. **Commit** your changes:
```bash
git add README.md
git commit -m "Add [resource name] to [section]"
```
6. **Push** to your fork:
```bash
git push origin add-new-resource
```
7. **Submit** a pull request from your fork to this repository
## 🔧 Types of Contributions
We welcome these types of contributions:
### ✅ Adding New Resources
- **Tools & Software**: AI tools that help with research workflows
- **Papers & Publications**: Influential papers in AI for Science
- **Datasets**: High-quality scientific datasets
- **Models**: Pre-trained models for scientific applications
- **Educational Content**: Courses, tutorials, books
- **Communities**: Research groups, conferences, forums
### ✅ Improving Existing Content
- **Better Descriptions**: More accurate or detailed descriptions
- **Updated Links**: Fixing broken or outdated links
- **Reorganization**: Improving the structure and categorization
- **Additional Information**: Adding missing details or context
### ✅ General Improvements
- **Typo Fixes**: Grammar, spelling, and formatting corrections
- **New Categories**: Suggesting new sections or reorganization
- **Documentation**: Improving this contributing guide or README
## 📝 Contribution Guidelines
### Resource Quality Standards
Before adding a resource, ensure it meets these criteria:
- **✅ Relevance**: Directly related to AI applications in scientific research
- **✅ Quality**: Well-documented, actively maintained, or highly cited
- **✅ Accessibility**: Publicly available (open source, free, or with free tier)
- **✅ Uniqueness**: Not already listed in the repository
- **✅ Functionality**: Actually works and provides value to researchers
### What NOT to Include
- **❌ Commercial Products**: Purely commercial tools without free access
- **❌ Broken Links**: Resources that are no longer available
- **❌ Personal Projects**: Small, unmaintained personal repositories
- **❌ Duplicates**: Resources already listed elsewhere in the repo
- **❌ Off-topic**: Resources not related to AI or scientific research
## 📐 Formatting Guidelines
### General Format
```markdown
- [Resource Name](URL) - Brief description of what it does and why it's useful
```
### Examples of Good Entries
```markdown
- [AlphaFold](https://github.com/deepmind/alphafold) - Revolutionary protein structure prediction using deep learning
- [Elicit](https://elicit.org/) - AI research assistant that helps with literature review and evidence synthesis
- [Materials Project](https://materialsproject.org/) - Computational materials database with ML-predicted properties
```
### Description Guidelines
- **Length**: 5-15 words ideally, max 20 words
- **Style**: Clear, informative, avoid marketing language
- **Focus**: What it does and scientific domain
- **Tone**: Professional and objective
### Link Guidelines
- **Use HTTPS**: Always use secure links when available
- **Direct Links**: Link to the main project page, not sub-pages
- **GitHub**: For open source projects, link to the GitHub repository
- **Papers**: Link to the official publication (DOI preferred)
### Section Organization
- **Alphabetical Order**: Within each subsection, maintain alphabetical order
- **Appropriate Section**: Place resources in the most specific relevant section
- **New Sections**: Propose new sections if existing ones don't fit
## 🔍 Review Process
### What We Look For
1. **Accuracy**: Correct information and working links
2. **Formatting**: Follows the style guide
3. **Placement**: Resource is in the appropriate section
4. **Quality**: Meets our quality standards
5. **Uniqueness**: Not a duplicate
### Timeline
- **Initial Review**: Within 7 days
- **Feedback**: We'll provide constructive feedback if changes are needed
- **Final Decision**: Merge or close within 14 days
### Review Criteria
**Approve** if:
- Meets all quality standards
- Follows formatting guidelines
- Adds clear value to the collection
🔄 **Request Changes** if:
- Minor formatting or description issues
- Needs better categorization
- Requires additional context
**Reject** if:
- Doesn't meet quality standards
- Off-topic or commercial
- Duplicate of existing entry
## 🎯 Code of Conduct
### Our Standards
- **Be Respectful**: Treat all contributors with respect
- **Be Constructive**: Provide helpful, actionable feedback
- **Be Collaborative**: Work together to improve the resource
- **Be Patient**: Understand that reviews take time
### Unacceptable Behavior
- Harassment or discriminatory language
- Spam or self-promotion without value
- Disruptive or unconstructive criticism
- Violations of intellectual property
## 🆘 Getting Help
### Questions?
- **Issues**: Open an issue for questions about contributions
- **Discussions**: Use GitHub Discussions for general questions
- **Email**: Contact the maintainers for sensitive issues
### Common Questions
**Q: Can I add my own research project?**
A: Yes, if it's high-quality, well-documented, and provides clear value to the scientific community.
**Q: What if a resource becomes outdated?**
A: Please open an issue or submit a PR to remove or update it.
**Q: Can I reorganize entire sections?**
A: Major reorganizations should be discussed in an issue first to gather feedback.
**Q: What about resources behind paywalls?**
A: We prefer freely accessible resources, but important papers or tools with free tiers are acceptable.
---
## 🙏 Thank You!
Your contributions make this resource valuable for researchers worldwide. Every addition, fix, and improvement helps accelerate scientific discovery through AI.
**Happy Contributing!** 🚀
---
*For questions about this guide, please open an issue or start a discussion.*
@@ -0,0 +1,998 @@
---
title: "Readme"
task: ""
lineage_type: import
upstream_source: https://github.com/ai-boost/awesome-ai-for-science/blob/709cee28/README.md
upstream_sha: 709cee28
imported_at: 2026-07-15
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
<div align="center">
<h1>✨ Awesome AI for Science (AI4Science) ✨</h1>
<img src="assets/banner.jpg" alt="Awesome AI for Science Banner" width="100%">
<p align="center">
A curated list of awesome AI tools, libraries, papers, datasets, and frameworks that accelerate <strong>scientific discovery</strong> across all disciplines.
</p>
<!-- Keep these links. Translations will automatically update with the README. -->
<p align="center">
<a href="https://zdoc.app/de/ai-boost/awesome-ai-for-science">Deutsch</a> |
<a href="https://zdoc.app/en/ai-boost/awesome-ai-for-science">English</a> |
<a href="https://zdoc.app/es/ai-boost/awesome-ai-for-science">Español</a> |
<a href="https://zdoc.app/fr/ai-boost/awesome-ai-for-science">français</a> |
<a href="https://zdoc.app/ja/ai-boost/awesome-ai-for-science">日本語</a> |
<a href="https://zdoc.app/ko/ai-boost/awesome-ai-for-science">한국어</a> |
<a href="https://zdoc.app/pt/ai-boost/awesome-ai-for-science">Português</a> |
<a href="https://zdoc.app/ru/ai-boost/awesome-ai-for-science">Русский</a> |
<a href="https://zdoc.app/zh/ai-boost/awesome-ai-for-science">中文</a>
</p>
<p align="center">
<a href="https://awesome.re"><img src="https://awesome.re/badge.svg" alt="Awesome"></a>
<a href="https://opensource.org/license/MIT"><img src="https://img.shields.io/badge/License-MIT-yellow.svg" alt="License: MIT"></a>
<a href="https://github.com/ai-boost/awesome-ai-for-science"><img src="https://img.shields.io/github/stars/ai-boost/awesome-ai-for-science.svg?style=social&label=Star" alt="GitHub stars"></a>
<a href="https://github.com/ai-boost/awesome-ai-for-science"><img src="https://img.shields.io/github/forks/ai-boost/awesome-ai-for-science.svg?style=social&label=Fork" alt="GitHub forks"></a>
</p>
</div>
> AI is revolutionizing scientific research - from drug discovery and materials design to climate modeling and astrophysics. This repository collects the best resources to help researchers leverage AI in their work.
## 📚 Contents
- [🧪 AI Tools for Research](#-ai-tools-for-research)
- [📄 Paper→Poster / Slides / Graphical Abstract](#-paperposter--slides--graphical-abstract)
- [📊 Chart Understanding & Generation](#-chart-understanding--generation)
- [🔄 Paper-to-Code & Reproducibility](#-paper-to-code--reproducibility)
- [📋 Scientific Documentation & Parsing](#-scientific-documentation--parsing)
- [🧰 Research Workbench & Plugins](#-research-workbench--plugins)
- [🕸 Knowledge Extraction & Scholarly KGs](#-knowledge-extraction--scholarly-kgs)
- [🤖 Research Agents & Autonomous Workflows](#-research-agents--autonomous-workflows)
- [🏷 Data Labeling & Curation](#-data-labeling--curation)
- [⚗ Scientific Machine Learning](#-scientific-machine-learning)
- [📖 Papers & Reviews](#-papers--reviews)
- [🔬 Domain-Specific Applications](#-domain-specific-applications)
- [🧬 Biology & Medicine](#-biology--medicine)
- [⚛ Chemistry & Materials](#-chemistry--materials)
- [🌌 Physics & Astronomy](#-physics--astronomy)
- [🌍 Earth & Climate Science](#-earth--climate-science)
- [🌾 Agriculture & Ecology](#-agriculture--ecology)
- [🧠 Social Sciences](#-social-sciences)
- [🤖 Foundation Models for Science](#-foundation-models-for-science)
- [📈 Datasets & Benchmarks](#-datasets--benchmarks)
- [💻 Computing Frameworks](#-computing-frameworks)
- [🎓 Educational Resources](#-educational-resources)
- [🏛 Research Communities](#-research-communities)
- [📚 Related Awesome Lists](#-related-awesome-lists)
---
## 🧪 AI Tools for Research
### Literature & Knowledge Management
- [Semantic Scholar](https://www.semanticscholar.org/) - AI-powered academic search (Allen AI)
- [arXiv](https://arxiv.org/) - Open-access repository of electronic preprints and postprints
- [OpenAlex](https://openalex.org/) - Open catalog of scholarly papers and authors
- [CORE](https://core.ac.uk/) - Aggregator of open access research papers
- [Connected Papers](https://www.connectedpapers.com/) - AI-powered visual graph for exploring academic papers and discovering connected research through citation networks and semantic similarity
- [PaSa (ByteDance)](https://github.com/bytedance/pasa) - Advanced paper search agent powered by large language models, autonomously invoking search tools, reading papers, and selecting references to deliver comprehensive and accurate results for complex scholarly queries (1.5K+ stars, Apache 2.0, 2024)
- [paper-search-mcp](https://github.com/openags/paper-search-mcp) - MCP server, CLI, and agent skills for searching and downloading academic papers from multiple open sources (arXiv, PubMed, bioRxiv, Semantic Scholar, OpenAlex, CORE, Europe PMC, etc.) with unified, deduplicated, LLM-friendly retrieval and an OA-first download fallback chain (OpenAGS, 1.9K+ stars, MIT License, 2025)
### Data Analysis & Visualization
- [PandasAI](https://github.com/Sinaptik-AI/pandas-ai) - Conversational data analysis using natural language
- [DeepAnalyze](https://github.com/ruc-datalab/DeepAnalyze) - First agentic LLM for autonomous data science with end-to-end pipeline from data to analyst-grade reports
- [AutoViz](https://github.com/AutoViML/AutoViz) - Automated data visualization with minimal code
- [Chat2Plot](https://github.com/nyanp/chat2plot) - Secure text-to-visualization through standardized chart specifications
### Data Labeling & Annotation
- [Label Studio](https://github.com/heartexlabs/label-studio) - Multi-type data labeling and annotation tool
- [Snorkel](https://github.com/snorkel-team/snorkel) - Programmatic data labeling and weak supervision
### Research Workbench & Plugins
- [Claude Scientific Skills](https://github.com/K-Dense-AI/claude-scientific-skills) - Comprehensive collection of 125+ ready-to-use scientific skill modules for Claude AI across bioinformatics, cheminformatics, clinical research, ML, and materials science
- [GDM Science Skills](https://github.com/google-deepmind/science-skills) - Google DeepMind's official collection of agentic science skills accelerating scientific workflows with better grounding and higher token efficiency, integrating insights from AlphaGenome, AFDB, UniProt and 30+ other databases and tools (2026)
- [Scientific Agent Skills](https://github.com/K-Dense-AI/scientific-agent-skills) - Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science with 140+ ready-to-use skills and 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Antigravity, and the open Agent Skills standard (K-Dense-AI, 26K+ stars, 2025)
- [SciAgent-Skills](https://github.com/jaechang-hits/SciAgent-Skills) - 197 bioinformatics and life science skills for Claude Code and AI agents, achieving 92.0% accuracy on BixBench. Covers RNA-seq, single-cell analysis, drug discovery, proteomics, and more. Powers OmicsHorizon (195+ stars, 2026)
- [Medical Research Skills](https://github.com/aipoch/medical-research-skills) - Curated library of 550+ medical research agent skills spanning evidence insights, protocol design, omics/clinical data analysis, and academic writing; each skill is reviewed through MedSkillAudit and compatible with Claude Code, Codex, Open Code, OpenClaw, and SKILL.md-compatible agents (AIPOCH, 1.2K+ stars, MIT License, 2026)
- [bioSkills](https://github.com/GPTomics/bioSkills) - Collection of SKILLS.md guiding AI coding agents (Claude Code, OpenAI Codex, Google Gemini, OpenCode, OpenClaw) through common bioinformatics workflows from basic sequence manipulation to advanced analyses such as single-cell RNA-seq and population genetics; evaluated on the Bio-Task Bench dataset (GPTomics, 969+ stars, MIT License, 2026)
---
## 📄 Paper→Poster / Slides / Graphical Abstract
### Poster Generation
- [Paper2Poster](https://github.com/Paper2Poster/Paper2Poster) - Multi-agent system with Parser-Planner-Painter architecture converting `paper.pdf` to editable `poster.pptx`, outperforms GPT-4o with 87% fewer tokens
- [mPLUG-PaperOwl](https://github.com/X-PLUG/mPLUG-DocOwl) - Multimodal LLM for scientific charts and diagrams understanding/generation
### Slides & Presentation Generation
- [Auto-Slides](https://auto-slides.github.io/) - Multi-agent academic paper to high-quality presentation slides with interactive refinement
- [PPTAgent](https://github.com/icip-cas/PPTAgent) - Beyond text-to-slides generation with PPTEval multi-dimensional evaluation (EMNLP 2025)
- [paper2slides](https://github.com/takashiishida/paper2slides) - Transform arXiv papers into Beamer slides using LLMs
- [PaperToSlides](https://github.com/jxtse/PaperToSlides) - AI-powered tool that automatically converts academic papers (PDF) into presentation slides
- [pdf2slides](https://github.com/ha0ranyu/pdf2slides) - Convert PDF files into editable slides with three lines of code
- [SlideDeck AI](https://github.com/barun-saha/slide-deck-ai) - Co-create PowerPoint presentations with Generative AI from documents or topics
- [AI Multi-Agent Presentation Builder](https://github.com/Azure-Samples/ai-multi-agent-presentation-builder) - Azure Semantic Kernel multi-agent PPT generation reference
### Video & Media Generation
- [Paper2Video](https://github.com/showlab/Paper2Video) - First benchmark for automatic video generation from scientific papers (NeurIPS 2025)
- [paper2video](https://github.com/mett29/paper2video) - Transform arXiv research papers into engaging presentations and YouTube-ready videos
### Website & Interactive Content Generation
- [Paper2All](https://github.com/YuhangChen1/Paper2All) - AI-powered pipeline converting papers into interactive websites, posters, and multimedia presentations with "Let's Make Your Paper Alive!" philosophy
### Figure & Illustration Generation
- [PaperBanana](https://github.com/dwzhu-pku/PaperBanana) - Automated academic illustration generation for AI scientists, converting research papers into publication-ready figures using VLMs and diffusion models with iterative refinement (PKU & Google Research, 6.2K+ stars, 2026)
### Chart & Visualization Generation
*Note: For comprehensive chart understanding and code generation tools, see [📊 Chart Understanding & Generation](#-chart-understanding--generation) section*
---
## 📊 Chart Understanding & Generation
### Chart-to-Code & Reproducibility
- [ChartCoder (ACL 2025)](https://aclanthology.org/2025.acl-long.363/) - Multimodal LLM for chart-to-code generation, 7B model outperforms larger open-source MLLMs
- [ChartAssistant / ChartAst (ACL 2024)](https://github.com/OpenGVLab/ChartAst) - Universal chart comprehension and reasoning model
- [Chart-to-Text Datasets](https://github.com/vis-nlp/Chart-to-text) - Large-scale chart summarization datasets for training chart description capabilities
### Scientific Visualization Tools
- [Chat2Plot](https://github.com/nyanp/chat2plot) - Secure text-to-visualization through standardized chart specifications
- [AutoViz](https://github.com/AutoViML/AutoViz) - Automated data visualization with minimal code
- [PlotlyAI](https://plotly.com/ai/) - AI-powered data visualization and dashboard creation
---
## 🔄 Paper-to-Code & Reproducibility
### Automated Code Generation
- [Paper2Code](https://github.com/going-doer/Paper2Code) - Automated code generation from machine learning research papers into runnable implementations (4.5K+ stars, 2025)
- [Paper2Agent](https://github.com/jmiao24/Paper2Agent) - Multi-agent system automatically transforming research papers into interactive AI agents with MCP server generation, tutorial auto-detection, and benchmark extraction (2.2K+ stars, MIT License, 2025)
- [AutoP2C](https://arxiv.org/abs/2504.20115) - LLM agent framework generating runnable repositories from academic papers
- [ResearchCodeAgent](https://arxiv.org/abs/2504.20117) - Multi-agent system for automated codification of research methodologies
- [ToolMaker](https://huggingface.co/papers/2502.11705) - Convert papers with code into callable agent tools
### Experiment Automation
- [BioProBench](https://huggingface.co/datasets/GreatCaptainNemo/BioProBench) - Comprehensive benchmark for automatic evaluation of LLMs on biological protocols and procedural understanding
- [Alhazen](https://chanzuckerberg.github.io/alhazen/) - Extract experimental metadata and protocol information from scientific documents
---
## 📋 Scientific Documentation & Parsing
### High-Performance Document Processing
- [MinerU (2024/2025)](https://github.com/opendatalab/MinerU) - SOTA multimodal document parsing with 1.2B parameters outperforming GPT-4o, converts PDFs to LLM-ready Markdown/JSON
- [MinerU-Diffusion (OpenDataLab, ECCV 2026)](https://github.com/opendatalab/MinerU-Diffusion) - Diffusion-based document OCR framework replacing autoregressive decoding with block-level parallel diffusion decoding, enabling high-accuracy text recognition in scientific PDFs (613+ stars, MIT License)
- [OpenDataLoader PDF (OpenDataLoader, 2025)](https://github.com/opendataloader-project/opendataloader-pdf) - Open-source PDF parser for AI-ready data, converting PDFs into Markdown/JSON/HTML/Tagged PDF with layout analysis and reading-order detection; ranks #1 overall on extraction benchmarks with deterministic bounding boxes and hybrid AI mode (26K+ stars, Apache 2.0)
- [PDF-Extract-Kit (2024)](https://github.com/opendatalab/PDF-Extract-Kit) - Comprehensive toolkit for high-quality PDF content extraction with layout detection, formula recognition, and OCR
- [Docling (IBM, AAAI 2025)](https://research.ibm.com/publications/docling-an-efficient-open-source-toolkit-for-ai-driven-document-conversion) - Multi-format (PDF/DOCX/PPTX/HTML/Images) → structured data (Markdown/JSON) with layout reconstruction, table/formula recovery
- [Nougat (Meta AI)](https://github.com/facebookresearch/nougat) - Neural optical understanding for academic documents, transforms scientific PDFs to Markdown with mathematical formula support
- [olmOCR (AllenAI)](https://github.com/allenai/olmocr) - Toolkit for linearizing academic PDFs into LLM-ready text with high accuracy and structure preservation, optimized for scientific literature extraction
- [PaddleOCR 3.0 (2024/2025)](https://github.com/PaddlePaddle/PaddleOCR) - Advanced OCR with PP-StructureV3 document parsing, 13% accuracy improvement, supports 80+ languages
- [Unstructured](https://github.com/Unstructured-IO/unstructured) - Production-grade ETL for transforming complex documents into structured formats, with open-source API
- [Marker](https://github.com/datalab-to/marker) - High-accuracy PDF→Markdown/JSON/HTML conversion, specialized for tables/formulas/code blocks with benchmark scripts
- [S2ORC doc2json (AllenAI)](https://github.com/allenai/s2orc-doc2json) - Large-scale PDF/LaTeX/JATS parsing to standardized JSON for millions of papers
- [GROBID](https://github.com/kermitt2/grobid) - Machine learning software for extracting structured metadata from scholarly documents
- [Science-Parse / SPv2 (AllenAI)](https://github.com/allenai/science-parse) - Parse scientific papers to structured fields (title/author/sections/references)
### Production Pipelines & Data Preparation
- [IBM Data Prep Kit: PDF→Parquet](https://ibm.github.io/data-prep-kit/transforms/language/pdf2parquet/) - Large-scale scientific document ingestion pipeline with optimization configurations
- [Mozilla document-to-markdown](https://github.com/mozilla-ai/document-to-markdown) - Docling-powered parsing with UI/CLI demonstration for rapid prototyping
### Figure & Table Extraction
- [PDFFigures2](https://github.com/allenai/pdffigures2) - Extract figures, tables, captions, and section titles from scholarly PDFs
- [TableBank](https://github.com/doc-analysis/TableBank) - Large-scale table detection and recognition dataset with pre-trained models
### Scientific Text Processing & NLP
- [scispacy (AllenAI)](https://github.com/allenai/scispacy) - Full spaCy pipeline and models for scientific/biomedical documents, enabling named entity recognition, abbreviation resolution, and UMLS linking for scientific literature mining (1.9K+ stars, Apache 2.0)
### Scientific Literature RAG & Analysis
- [PaperQA2](https://github.com/future-house/paper-qa) - High-accuracy RAG for scientific PDFs with citation support, agentic RAG, and contradiction detection
- [OpenScholar](https://github.com/AkariAsai/OpenScholar) - Retrieval-augmented LM synthesizing scientific literature from 45M papers with human-expert-level citation accuracy, outperforming GPT-4o by 5% on ScholarQABench (Nature 2026, UW & Ai2)
- [Valsci](https://github.com/bricee98/Valsci) - Self-hostable scientific claim-verification and literature-review tool combining Semantic Scholar retrieval, bibliometric scoring, and LLM-based evidence synthesis for large-batch validation workflows
- [paper-reviewer](https://github.com/deep-diver/paper-reviewer) - Generate comprehensive reviews from arXiv papers and convert to blog posts
- [STORM](https://github.com/stanford-oval/storm) - LLM agent system synthesizing Wikipedia-like long-form research articles from scratch through multi-perspective question asking, web retrieval, and citation-grounded report generation, with Co-STORM extension for collaborative human-LLM knowledge curation conversations (Stanford OVAL, NAACL 2024 & EMNLP 2024)
---
## 🧰 Research Workbench & Plugins
### Interactive Research Environments
- [Jupyter AI (JupyterLab Extension)](https://github.com/jupyterlab/jupyter-ai) - Official Jupyter extension with `%%ai` magic commands and sidebar chat assistant, connecting multiple model providers and local inference
- [Notebook Intelligence (NBI)](https://github.com/notebook-intelligence/notebook-intelligence) - AI coding assistant for JupyterLab with agent mode, supporting arbitrary LLM providers (2025+)
- [Google Colab AI Features](https://colab.research.google.com/) - Integrated AI assistance for data science and research notebooks
- [OpenBioMed](https://github.com/PharMolix/OpenBioMed) - Open-source biomedical AI platform integrating multimodal foundation models (BioMedGPT, PharmolixFM, LangCell) with agentic workflows and 45+ Claude Code skills for drug discovery, protein engineering, and single-cell omics analysis (PharMolix & Tsinghua AIR, 1K+ stars, 2023-2026)
- [AutoR](https://github.com/AutoX-AI-Labs/AutoR) - Human-centered research OS with terminal-first harness and local browser Studio, turning research work into reproducible artifact-backed runs through a 9-stage workflow with human approval gates, resume/rollback controls, and venue-aware manuscript packaging (1K+ stars, 2026)
- [ScholarAIO](https://github.com/ZimoLiao/scholaraio) - Agent-agnostic research infrastructure providing AI agents with a structured scientific workspace for deep PDF parsing, hybrid semantic/keyword literature search, citation-graph analysis, topic discovery, and academic writing workflows; natively integrates with Claude Code, Codex, Cursor, Cline, and AgentSkills.io (530+ stars, MIT License, 2026)
- [BioMCP](https://github.com/genomoncology/biomcp) - Biomedical Model Context Protocol (MCP) server unifying literature search across PubMed/Europe PMC, entity pivoting across genes/variants/drugs/diseases/pathways/proteins, local study analytics, and Claude Code/Codex integration for agentic biomedical research (531+ stars, MIT License, 2025-2026)
- [MATLAB Agentic Toolkit](https://github.com/matlab/matlab-agentic-toolkit) - Official MathWorks toolkit connecting AI agents to MATLAB via the MATLAB MCP Server and curated skills, enabling trusted engineering and scientific computing workflows with idiomatic code generation, testing, and error diagnosis in Claude Code, GitHub Copilot, OpenAI Codex, and Gemini CLI (686+ stars, BSD-3-Clause, 2026)
- [BioNeMo Agent Toolkit (NVIDIA)](https://github.com/NVIDIA-BioNeMo/bionemo-agent-toolkit) - Turn any AI agent into a life science expert with NVIDIA BioNeMo skills, enabling agentic workflows for drug discovery, protein engineering, and biomolecular design (329+ stars, Apache 2.0 / CC-BY-4.0, 2026)
- [open-science](https://github.com/ai4s-research/open-science) - Local-first, open-source AI workbench for scientists — an open alternative to Claude Science (by ai4s-research, maintainers of this list; TypeScript, MIT, 2026)
- [OpenScience (Synthetic Sciences)](https://github.com/synthetic-sciences/openscience) - Open-source AI workbench for scientific research that automates the full research loop — literature review, hypothesis generation, code writing, experiment execution, database querying, and report writing — with 290+ skills, specialized research agents, and a browser-based workspace (1453+ stars, Apache 2.0, 2026)
- [Claude Scholar](https://github.com/Galaxy-Dawn/claude-scholar) - Semi-automated research assistant for academic research and software development, supporting Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication (Galaxy-Dawn, 4.5K+ stars, MIT License, 2026)
- [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok) - Free, open-source desktop AI research assistant that runs locally and turns natural-language requests into real data analysis, literature search, figure generation, and manuscript review; ships with 149 scientific skills, 326 workflow templates, and 229 databases across genomics, proteomics, drug discovery, and materials science, plus a living lab notebook, 60+ scientific file previews, and LaTeX editing (K-Dense-AI, 908+ stars, MIT License, 2026)
### Literature Management Plugins
- [llm-for-zotero](https://github.com/yilewang/llm-for-zotero) - Research agent system deeply integrated with Zotero supporting Agent Mode, skills, multi-model backends (OpenAI-compatible, Claude Code, WebChat, Codex), and MinerU PDF parsing for literature Q&A, summarization, figure inspection, and source comparison (1.3K+ stars, 2026)
- [PapersGPT for Zotero](https://github.com/papersgpt/papersgpt-for-zotero) - Multi-PDF conversation, retrieval, and citation in Zotero with commercial/local models (Ollama), MCP support
- [Zotero-GPT (MuiseDestiny)](https://github.com/MuiseDestiny/zotero-gpt) - Classic open-source plugin for document Q&A and summarization within Zotero
- [Better BibTeX for Zotero](https://retorque.re/zotero-better-bibtex/) - Enhanced citation key management and LaTeX integration
### Scientific Writing & Collaboration
- [Notion AI](https://www.notion.so/product/ai) - AI-powered research note-taking and knowledge management
- [Obsidian Smart Connections](https://github.com/brianpetro/obsidian-smart-connections) - AI-powered note linking and research graph navigation
- [Research Rabbit](https://www.researchrabbit.ai/) - AI-powered literature discovery and research network mapping
- [SciWrite](https://github.com/labarba/sciwrite) - Agent skill for AI-assisted scientific manuscript writing review distilled from Stanford's *Writing in the Sciences* course, performing five sequential editorial audit passes on clarity, voice, structure, consistency, and integrity (2026)
- [Claude Prism](https://github.com/delibae/claude-prism) - Offline-first scientific writing workspace powered by Claude, integrating LaTeX, Python, and 100+ scientific skills with local execution, Zotero integration, and privacy-focused design (2026)
---
## 🕸 Knowledge Extraction & Scholarly KGs
### Knowledge Graph Construction
- [iText2KG](https://github.com/AuvaLab/itext2kg) - Incremental knowledge graph construction using LLMs with entity extraction and Neo4j visualization
- [GraphGen](https://github.com/open-sciencelab/GraphGen) - Knowledge graph-guided synthetic data generation for LLM fine-tuning, achieving strong performance on scientific QA (GPQA-Diamond) and math reasoning (AIME)
- [KoPA](https://github.com/zjukg/KoPA) - Structure-aware prefix adaptation for integrating LLMs with knowledge graphs (ACM MM 2024)
- [Scholarly KGQA](https://arxiv.org/abs/2311.09841) - LLM-powered question answering over scholarly knowledge graphs (ArXiv paper)
### Knowledge Graph Resources
- [Awesome-LLM-KG](https://github.com/RManLuo/Awesome-LLM-KG) - Comprehensive collection of papers on unifying LLMs and knowledge graphs
---
## 🤖 Research Agents & Autonomous Workflows
### Autonomous Research Systems (2023-2025 Breakthroughs)
- [FunSearch (DeepMind, Nature 2023)](https://github.com/google-deepmind/funsearch) - First system to make novel, verifiable scientific discoveries by pairing LLMs with evolutionary search, solving open problems in combinatorics (cap set problem) and discovering faster matrix multiplication algorithms
- [OpenEvolve](https://github.com/algorithmicsuperintelligence/openevolve) - Open-source implementation of AlphaEvolve's evolutionary coding agent paradigm, enabling LLMs to autonomously discover and optimize algorithms through iterative evolution, matching the approach behind DeepMind's breakthrough matrix multiplication discovery (6.2K+ stars, 2025)
- [SkyDiscover](https://github.com/skydiscover-ai/skydiscover) - Modular framework for AI-driven scientific and algorithmic discovery, providing a unified interface for implementing, running, and fairly comparing discovery algorithms across 200+ optimization tasks; introduces AdaEvolve and EvoX adaptive/evolutionary algorithms and natively supports OpenEvolve, GEPA, and Harbor-format benchmarks (skydiscover-ai, 568+ stars, Apache 2.0, 2026)
- [EvoMaster (SJTU SAI, arXiv 2026)](https://github.com/sjtu-sai-agents/EvoMaster) - Foundational auto-research agent framework for agentic science at scale, providing modular agent construction, run-level self-evolution, and multiple SciMaster domain agents (ML-Master, X-Master, Browse-Master); outperforms general-purpose agents across authoritative benchmarks including the OpenAI Frontier Science Benchmark (206+ stars, Apache 2.0, 2026)
- [Virtual Lab (Stanford Zou Group, Nature 2025)](https://github.com/zou-group/virtual-lab) - AI-human collaborative research platform where a human researcher works with a team of LLM agents via team and individual meetings to perform scientific research; demonstrated by designing new SARS-CoV-2 nanobodies with wet-lab validation
- [The AI Scientist (SakanaAI)](https://github.com/SakanaAI/AI-Scientist) - First fully autonomous open-ended scientific discovery system with official implementation: hypothesis→experiment→writing→review simulation (13.8K+ stars, 2024)
- [The AI Scientist v2 (SakanaAI)](https://github.com/SakanaAI/AI-Scientist-v2) - Official implementation of the second-generation fully autonomous scientific discovery system, extending the original with agentic tree search and reduced template dependency to achieve workshop-level accepted papers (6.7K+ stars, 2025)
- [The AI Scientist v1 (2024)](https://arxiv.org/abs/2408.06292) - First fully autonomous research system: hypothesis→experiment→writing→review simulation
- [The AI Scientist v2 (2025)](https://arxiv.org/abs/2504.08066) - Enhanced with Agentic Tree Search, reduced template dependency, first workshop-level accepted paper
- [DeepScientist](https://github.com/ResearAI/DeepScientist) - First system progressively surpassing human SOTA on frontier AI tasks (183.7%, 1.9%, 7.9% improvements), month-long autonomous discovery with 20,000+ GPU hours
- [ASI-Arch (GAIR-NLP, arXiv 2025)](https://github.com/GAIR-NLP/ASI-Arch) - Autonomous multi-agent research loop for model architecture discovery that ran 1,773 experiments over 20,000 GPU hours and produced 106 state-of-the-art linear-attention architectures, surpassing human-designed baselines including Mamba2 and DeltaNet (1.1K+ stars, Apache 2.0)
- [Kosmos](https://github.com/jimmc414/Kosmos) - Extended autonomy AI scientist with 200 parallel agent rollouts, 42K lines of code execution, 1.5K papers analyzed per run, achieving 79.4% accuracy and 7 scientific discoveries (Edison Scientific)
- [AlphaResearch](https://github.com/answers111/alpha-research) - Autonomous algorithm discovery combining evolutionary search with peer-review reward models, achieving best-known performance on circle packing problems
- [AutoResearchClaw](https://github.com/aiming-lab/AutoResearchClaw) - Fully autonomous research from idea to paper with multi-agent debate, citation verification, and OpenClaw integration (11K+ stars, 2026)
- [ARIS (Auto-Research-In-Sleep)](https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep) - Lightweight Markdown-only skills for autonomous ML research with cross-model review loops, idea discovery, and experiment automation; no framework lock-in, works with Claude Code, Codex, OpenClaw, or any LLM agent (12.8K+ stars, MIT License, 2026)
- [Arbor](https://github.com/RUC-NLPIR/Arbor) - Generalist autonomous research agent that grows a hypothesis tree to optimize any measurable task, beating Claude Code and Codex by 2.5× on the same compute budget across BrowseComp, Terminal-Bench 2.0, math reasoning, and MLE-Bench Lite; supports native CLI, keyless Claude Code/Codex integration, and an MCP tool server (RUC-NLPIR, 866+ stars, Apache 2.0, 2026)
- [NanoResearch](https://github.com/OpenRaiser/NanoResearch) - End-to-end autonomous AI research engine that turns an idea into a complete LaTeX paper by dispatching real computational experiments to local GPUs or SLURM clusters, collecting actual results, generating figures/tables, and writing a data-grounded manuscript rather than LLM hallucinations (OpenRaiser, 1.5K+ stars, MIT License, 2026)
- [ScienceClaw](https://github.com/beita6969/ScienceClaw) - Self-evolving AI research colleague built on OpenClaw with 285+ runtime-adaptive skills across 28+ disciplines, persistent cross-session research memory, and zero-hallucination citation protocols; agent autonomously writes new SKILL.md files based on research patterns without redeployment (828+ stars, MIT License, 2026)
- [ai4s-skills](https://github.com/ai4s-research/ai4s-skills) - Agent skills (SKILL.md + deterministic tools) for the AI4S workflow — topic exploration, literature survey, runnable experiments, publication-grade papers, and integrity audit, with every citation and number traceable to its source (by ai4s-research, maintainers of this list; MIT, 2026)
- [Denario (AstroPilot-AI, Agents4Science 2025)](https://github.com/AstroPilot-AI/Denario) - Modular multi-agent scientific research assistant that automates idea generation, literature review, methodology design, code execution in Docker, visualization, LaTeX paper writing, and peer-review simulation across 10+ disciplines; winner of the NeurIPS 2025 Fair Universe Competition (573+ stars, GPL-3.0, 2025-2026)
- [AI-Researcher](https://github.com/HKUDS/AI-Researcher) - Autonomous pipeline from literature review→hypothesis→algorithm implementation→publication-level writing with Scientist-Bench evaluation
- [Agent Laboratory](https://agentlaboratory.github.io/) - Multi-agent workflows for complete research cycles with AgentRxiv for cumulative discovery
- [AIDE (WecoAI, arXiv 2025)](https://github.com/WecoAI/aideml) - LLM-driven machine learning engineering agent using agentic tree search to autonomously draft, debug and benchmark ML code; wins 4× more medals than the best linear agent on OpenAI's MLE-Bench (75 Kaggle competitions) (1.3K+ stars, MIT License)
- [RD-Agent (Microsoft)](https://github.com/microsoft/rd-agent) - Open-source LLM-powered R&D agent framework automating data-driven AI solution building through automated research, development, and evolution; achieves top open-source performance on MLE-Bench with dual Researcher-Developer agents and supports research copilot, data mining, Kaggle, and quant R&D workflows (13.6K+ stars, MIT License, 2025-2026)
- [CodeScientist (AllenAI)](https://github.com/allenai/codescientist) - End-to-end semi-automated scientific discovery system that designs, iterates, and analyzes code-based experiments via LLM-as-a-mutator over scientific articles and code examples; auto-creates, runs, and debugs experiment code in containers and writes meta-analysis reports (339+ stars, Apache 2.0)
- [InternAgent](https://github.com/Alpha-Innovator/InternAgent) - Closed-loop multi-agent system from hypothesis to verification across 12 scientific tasks, #1 on MLE-Bench (36.44%)
- [freephdlabor](https://github.com/ltjed/freephdlabor) - First fully customizable open-source multiagent framework automating complete research lifecycle from idea conception to LaTeX papers with dynamic workflows
- [AutoScientists (Harvard MIMS, arXiv 2026)](https://github.com/mims-harvard/AutoScientists) - Decentralized self-organizing teams of AI agents for long-running computational scientific experimentation; agents critique each other's proposals before spending compute and share successes/failures to avoid redundant exploration, achieving +8.33% on BioML-Bench, 1.9× faster nanoGPT optimization, and +12.5% on ProteinGym ACE2-Spike (425+ stars, 2026)
- [ToolUniverse](https://github.com/mims-harvard/ToolUniverse) - Democratizing AI scientists by transforming any LLM into research systems with 600+ scientific tools (Harvard MIMS)
- [LabClaw](https://github.com/wu-yc/LabClaw) - Skill operating layer for biomedical AI agents with 211 production-ready SKILL.md files across 7 domains (biology, pharmacology, medicine, data science, literature search), enabling modular dry-lab reasoning and protocol composition for Stanford LabOS-compatible agents
- [Robin](https://github.com/Future-House/robin) - FutureHouse's end-to-end scientific discovery multi-agent system orchestrating literature search (Crow/Falcon) and data analysis (Finch) agents, first AI-generated drug discovery identifying ripasudil as novel dry AMD therapeutic (2025)
- [Aviary](https://github.com/Future-House/aviary) - Language agent gymnasium for challenging scientific tasks including DNA manipulation, literature search, and protein engineering
- [Curie](https://github.com/Just-Curieous/Curie) - Automated and rigorous experiments using AI agents for scientific discovery
- [POPPER](https://github.com/snap-stanford/POPPER) - Automated hypothesis testing with agentic sequential falsifications
- [autoresearch](https://github.com/karpathy/autoresearch) - Andrej Karpathy's autonomous LLM research framework: AI agent runs overnight experiments on a real training setup, auto-editing code→5min training→evaluation in a loop, ~100 experiments per night on a single GPU
- [UniScientist](https://github.com/UniPat-AI/UniScientist) - Universal scientific research intelligence covering 50+ disciplines, repositioning LLMs as cross-disciplinary generators with human experts as verifiers; 30B model outperforms Claude Opus and GPT on 5 research benchmarks
- [EvoScientist](https://github.com/EvoScientist/EvoScientist) - Self-evolving AI scientist with 6 specialized sub-agents (plan/research/code/debug/analyze/write) and persistent memory, #1 on DeepResearch Bench II and AstaBench, supporting multi-provider LLMs and multi-channel deployment (Apache 2.0, 2026)
- [PantheonOS (Stanford, 2025)](https://github.com/aristoteleo/PantheonOS) - Evolvable and privacy-preserving multi-agent framework automating, scaling, and accelerating data sciences with a particular focus on end-to-end single-cell biology analyses; features agentic code evolution, multi-agent team orchestration, distributed architecture, and a community marketplace with 1,000+ curated agents and skills (428+ stars)
- [CORAL (arXiv 2026)](https://github.com/Human-Agent-Society/CORAL) - Robust, lightweight infrastructure for multi-agent autonomous self-evolution, built for autoresearch; agents run in isolated git worktrees, share knowledge through a common state directory, and are scored by a grader daemon; natively integrated with Claude Code, Codex, Cursor Agent, OpenCode, and Kiro (672+ stars, Apache 2.0)
- [Science-Star (USTC AI4Science, 2025)](https://github.com/ustc-ai4science/Science-Star) - Open-source platform for building, extending, and experimenting with scientific agents, providing modular agent construction tools and standardized evaluation pipelines for accelerating autonomous scientific discovery research (748+ stars, MIT License)
- [SR-Scientist (ICLR 2026)](https://github.com/GAIR-NLP/SR-Scientist) - Scientific equation discovery with agentic AI, elevating LLMs from equation proposers to autonomous scientists that write code, analyze data, implement equations, and optimize based on experimental feedback; outperforms baselines by 6-35% across four science disciplines with robustness to noise and out-of-domain generalization (GAIR-NLP / SJTU, 49+ stars, Apache 2.0)
- [ARA (Agent-Native Research Artifact)](https://github.com/ARA-Labs/Agent-Native-Research-Artifact) - Research ecosystem for rigorous and trustworthy AI scientists — a protocol and skill bundle that makes autonomous research verifiable, crystallized, and observable through structured, machine-executable research artifacts and five agent skills for research management, compilation, verification, visualization, and publication (ARA-Labs, 447+ stars, MIT License, 2026)
- [Scholar Loop](https://github.com/renee-jia/scholar-loop) - Autonomous multi-agent AI scientist that mirrors a PhD workflow: literature review → grounded hypothesis → real ML experiments → self-critique → write-up; features a deterministic harness with frozen-metric scoring, edit allowlists, and a verified registry to make reward-hacking and hallucination impossible, plus 108 unit tests runnable without API keys or GPUs (461+ stars, MIT License, 2026)
### Evaluation & Benchmarking
- [ScienceAgentBench (ICLR 2025)](https://github.com/OSU-NLP-Group/ScienceAgentBench) - 102 executable tasks from 44 peer-reviewed papers across 4 disciplines with containerized evaluation
- [AIRS-Bench (Meta, 2026)](https://github.com/facebookresearch/airs-bench) - Benchmark quantifying end-to-end autonomous AI research abilities of LLM agents across 20 tasks from SOTA machine learning papers spanning NLP, code, math, biochemical modelling, and time series forecasting, with normalized score metrics against human SOTA and HuggingFace dataset
- [PaperBench (OpenAI, 2025)](https://github.com/openai/preparedness/tree/main/project/paperbench) - Benchmark evaluating AI agents' ability to replicate 20 ICML 2024 Spotlight/Oral papers from scratch, with 8,316 gradable tasks and author-co-developed rubrics
- [MLE-Bench (OpenAI, 2024)](https://github.com/openai/mle-bench) - Benchmark evaluating AI agents on 75 curated Kaggle-style ML engineering competitions with reproducible Docker-based grading harness, human baselines, and end-to-end task lifecycle, used as a primary benchmark for autonomous ML research agents (e.g., InternAgent #1 at 36.44%)
- [ScienceBoard (ICLR 2026)](https://github.com/OS-Copilot/ScienceBoard) - Evaluating multimodal autonomous agents in realistic scientific workflows across real scientific software environments (KAlgebra, Celestia, Grass GIS, Lean 4, etc.) with VM-based evaluation infrastructure and agent trajectories
- [BuildArena](https://github.com/AI4Science-WestlakeU/BuildArena) - First physics-aligned interactive benchmark for LLM agents in engineering construction, designing rockets/cars/bridges in physics simulator with 3D spatial geometry library
- [SciTrust (2024)](https://impact.ornl.gov/en/publications/scitrust-evaluating-the-trustworthiness-of-large-language-models-) - Trustworthiness evaluation framework for scientific LLMs (truthfulness, hallucination, sycophancy)
- [SciCode](https://github.com/scicode-bench/SciCode) - Research coding benchmark curated by scientists with 338 subproblems across 16 subdomains (physics, math, materials, biology, chemistry), evaluating LLMs on realistic scientific programming tasks with gold-standard solutions (NeurIPS 2024)
- [SciBench](https://arxiv.org/abs/2307.10635) - College-level scientific problem-solving evaluation across multiple domains
- [NewtonBench (ICLR 2026)](https://github.com/HKUST-KnowComp/NewtonBench) - First benchmark evaluating LLMs' ability to rediscover scientific laws through interactive experimentation across 324 tasks in 12 physics domains, featuring memorization-resistant metaphysical shifts of canonical laws (HKUST)
- [ResearchClawBench (InternScience, arXiv 2026)](https://github.com/InternScience/ResearchClawBench) - Benchmark evaluating AI agents for end-to-end automated research from re-discovery to new-discovery, with 40 real-science tasks across 10 disciplines, curated datasets from published papers, and expert-curated multimodal rubrics (170+ stars, MIT License)
### Academic Review & Evaluation
- [AgentReview](https://agentreview.github.io/) - LLM agents simulating academic peer review ecosystems
- [LLM-Peer-Review](https://github.com/VijayGKR/LLM-Peer-Review) - Web application for LLM-assisted manuscript review and annotation
### Domain-Specific Research Agents
- [Aletheia](https://arxiv.org/abs/2602.10177) - Google DeepMind's autonomous mathematics research agent powered by Gemini Deep Think, autonomously solving 4 open problems from 700 Erdős conjectures and generating complete research papers without human intervention (February 2026)
- [AlphaGeometry](https://github.com/google-deepmind/alphageometry) - DeepMind's Olympiad-level geometry theorem prover combining neural language model with symbolic deduction engine, AlphaGeometry2 solves 84% of IMO geometry problems (42/50) at gold-medalist level (Nature 2024)
- [Goedel-Prover-V2](https://github.com/Goedel-LM/Goedel-Prover-V2) - Strongest open-source automated theorem prover in Lean 4, 8B model matches DeepSeek-Prover-V2-671B at 84.6% MiniF2F, 32B model achieves 90.4% with self-correction, using scaffolded data synthesis and verifier-guided proof refinement (Princeton, 2025)
- [DeepSeek-Prover-V2](https://github.com/deepseek-ai/DeepSeek-Prover-V2) - DeepSeek's open-source large language model for formal theorem proving in Lean 4, integrating informal and formal mathematical reasoning through recursive subgoal decomposition and reinforcement learning powered by DeepSeek-V3, with open weights and ProverBench evaluation (2025)
- [LeanDojo](https://github.com/lean-dojo/LeanDojo) - Open-source toolkit and benchmark for learning-based theorem proving in Lean, providing programmatic Lean interaction, a 98K+ theorem dataset extracted from 217 Lean projects, and ReProver—the first retrieval-augmented LLM-based theorem prover for Lean—with reproducible training pipelines underpinning much subsequent Lean prover research (Caltech & NVIDIA, NeurIPS 2023 Outstanding Paper, Datasets & Benchmarks)
- [Lean Copilot](https://github.com/lean-dojo/LeanCopilot) - LLMs as copilots for theorem proving in Lean 4, exposing native tactics (`suggest_tactics`, `search_proof`, `select_premises`) that embed language model inference and premise retrieval directly inside the Lean proof environment, supporting local CTranslate2/CUDA inference as well as remote model APIs for interactive and automated proof search (Caltech & NVIDIA, NeurIPS 2024, 1.2K+ stars)
- [Get Physics Done (PSI)](https://github.com/psi-oss/get-physics-done) - First open-source agentic AI physicist turning research questions into structured workflows with rigorous verification and multi-step analytical work for long-horizon physics projects; integrates with Claude Code, Codex, Gemini CLI, and OpenCode (804+ stars, Apache 2.0, 2026)
- [Foam-Agent (NeurIPS 2025)](https://github.com/csml-rpi/Foam-Agent) - End-to-end composable multi-agent framework for automating OpenFOAM-based CFD simulations from natural language prompts, managing meshing, case setup, execution, error correction, and post-processing; achieves 100% success rate on 110 FoamBench tasks with Claude Opus 4.6 through Architect-Input Writer-Runner-Reviewer agent collaboration with RAG-enhanced generation and MCP tool integration (RPI CSML, 242+ stars, MIT License)
- [Zephyrus (ICLR 2026)](https://github.com/Rose-STL-Lab/Zephyrus) - First agentic framework for weather science, pairing an LLM with ZephyrusWorld (a code-execution environment exposing WeatherBench 2 data, geolocation, forecasting, simulation, and climatology tools) and ZephyrusBench (2,230 Q&A pairs across 49 weather-science tasks); outperforms text-only baselines by up to 44.2 percentage points (UC San Diego Rose-STL-Lab, 99+ stars, MIT License, 2026)
- [BioDiscoveryAgent](https://github.com/snap-stanford/BioDiscoveryAgent) - AI agent for biological discovery and research automation
- [Biomni](https://github.com/snap-stanford/Biomni) - General-purpose biomedical AI agent integrating LLM reasoning with retrieval-augmented planning and code-based execution to autonomously execute diverse biomedical research tasks and generate testable hypotheses (Stanford SNAP, bioRxiv 2025)
- [BioAgents](https://github.com/bio-xyz/BioAgents) - AI scientist framework for autonomous deep research in biological sciences, combining literature analysis agents with data scientist agents to enable iterative scientific discovery through user feedback integration; achieves state-of-the-art performance on BixBench benchmark (48.78% open-answer, 64.39% multiple-choice) outperforming Kepler and GPT-5 (bio-xyz, arXiv 2601.12542, 160+ stars, 2025-2026)
- [SRAgent](https://github.com/ArcInstitute/SRAgent) - LLM agents for working with the SRA (Sequence Read Archive) and associated bioinformatics databases, enabling natural language querying of high-throughput sequencing data and metadata across genomic repositories (Arc Institute, 169+ stars, 2024-2026)
- [STAgent](https://github.com/LiuLab-Bioelectronics-Harvard/STAgent) - Multimodal LLM-based AI agent enabling deep research in spatial transcriptomics, automating analysis and interpretation of spatial gene expression data (Harvard LiuLab, bioRxiv 2025)
- [Camyla](https://github.com/yifangao112/Camyla) - Fully autonomous medical image segmentation research system that generates complete manuscripts end-to-end from datasets with zero human intervention, beating strongest baselines on 24 of 31 datasets and achieving T1-T2 tier manuscript quality in double-blind evaluations (USTC & Shanghai AI Lab, 2026)
- [MOOSE](https://github.com/ZonglinY/MOOSE) - Large Language Models for automated open-domain scientific hypotheses discovery (ACL 2024, ICML Best Poster)
- [ChemCrow](https://arxiv.org/abs/2304.05376) - LLM agents for chemistry research with tool integration
- [Coscientist](https://www.nature.com/articles/s41586-023-06792-1) - Autonomous chemical experiment planning and execution
- [SciAgents](https://github.com/lamm-mit/SciAgentsDiscovery) - Bioinspired multi-agent intelligent graph reasoning system that autonomously traverses ontological knowledge graphs to generate, critique, and refine novel research hypotheses, demonstrated on bio-inspired materials discovery with cross-disciplinary connection mining (MIT Lamm Group, 2024)
- [TxAgent](https://github.com/mims-harvard/TxAgent) - AI agent for therapeutic reasoning across a universe of tools, achieving 92.1% accuracy in drug reasoning and outperforming GPT-4o by 25.8% (Harvard MIMS, 2025)
- [ClawBio](https://github.com/ClawBio/ClawBio) - First bioinformatics-native AI agent skill library enabling local-first, reproducible genomic and population-genetics research workflows built on OpenClaw (871+ stars, MIT License, 2026)
---
## 🏷 Data Labeling & Curation
### Weak Supervision & Auto-Labeling
- [Snorkel](https://github.com/snorkel-team/snorkel) - Programmatic data labeling and weak supervision for scientific datasets
- [PandasAI](https://github.com/Sinaptik-AI/pandas-ai) - Conversational data analysis and visualization using natural language
- [Cleanlab](https://github.com/cleanlab/cleanlab) - Standard data-centric AI package for data quality and machine learning, automatically detecting label errors, outliers, and dataset issues to improve scientific dataset reliability and model performance (11K+ stars, MIT License)
---
## ⚗ Scientific Machine Learning
### Neural Differential Equations
- [torchdiffeq](https://github.com/rtqichen/torchdiffeq) - PyTorch implementation of neural ODEs
- [torchdyn](https://github.com/DiffEqML/torchdyn) - Neural differential equations in PyTorch
- [diffrax](https://github.com/patrick-kidger/diffrax) - Numerical differential equation solving in JAX
- [DifferentialEquations.jl](https://github.com/SciML/DifferentialEquations.jl) - Julia differential equations suite
- [DiffEqFlux.jl](https://github.com/SciML/DiffEqFlux.jl) - Neural differential equations in Julia
### Chemical Reaction Networks & Systems Biology
- [Catalyst.jl](https://github.com/SciML/Catalyst.jl) - Chemical reaction network and systems biology interface for scientific machine learning (SciML), enabling high-performance, GPU-parallelized simulation and analysis of complex biochemical systems with O(1) solvers (SciML, 518+ stars, Julia)
### Physics-Informed Neural Networks
- [DeepXDE](https://github.com/lululxvi/deepxde) - Deep learning library for solving PDEs
- [Lang-PINN](https://openreview.net/forum?id=ONEyVpgK34) - LLM-driven multi-agent system that builds trainable PINNs from natural language task descriptions, achieving 3-5 orders of magnitude MSE reduction and 50%+ execution success improvement (ICLR 2026)
- [PINNs](https://github.com/maziarraissi/PINNs) - Physics-informed neural networks
- [NVIDIA PhysicsNeMo](https://github.com/NVIDIA/physicsnemo) - Open-source framework for building physics-ML models at scale (renamed from Modulus, 2025)
- [PINA](https://github.com/mathLab/PINA) - Physics-Informed Neural networks for Advanced modeling in PyTorch
- [NeuroMANCER (PNNL)](https://github.com/pnnl/neuromancer) - PyTorch-based differentiable programming framework for physics-informed system identification, parametric constrained optimization, and model predictive control, integrating neural operators, neural ODEs, KANs, SINDy, and differentiable predictive control with 30+ tutorials (1.3k+ stars, BSD License)
- [SciANN](https://github.com/sciann/sciann) - Keras-based scientific neural networks
- [NeuralPDE.jl](https://github.com/SciML/NeuralPDE.jl) - Physics-informed neural networks in Julia
### Neural Operators & Model Discovery
- [DeepONet](https://github.com/lululxvi/deeponet) - Learning nonlinear operators
- [PySINDy](https://github.com/dynamicslab/pysindy) - Sparse identification of nonlinear dynamics
- [PySR](https://github.com/MilesCranmer/PySR) - High-performance symbolic regression for discovering interpretable scientific equations from data, multi-population evolutionary search with Python/Julia backend, widely used in physics and astronomy (Cambridge, NeurIPS 2023)
- [LLM-SR](https://github.com/deep-symbolic-mathematics/LLM-SR) - Scientific equation discovery and symbolic regression using LLMs, combining code generation with evolutionary search (ICLR 2025 Oral)
- [PSRN](https://github.com/x66ccff/PSRN) - Parallel symbolic regression network evaluating millions of expressions on GPU with automated subtree reuse, Nature Computational Science cover article (MIT, 2026)
- [pykan](https://github.com/KindXiaoming/pykan) - Kolmogorov-Arnold Networks with learnable activation functions on edges instead of fixed node activations, achieving strong performance in function fitting, PDE solving, and scientific discovery with enhanced interpretability as an alternative to MLPs (MIT, 16.3K+ stars, 2024)
- [Fourier Neural Operator](https://github.com/neuraloperator/neuraloperator) - Learning operators in Fourier space
- [Poseidon](https://github.com/camlab-ethz/poseidon) - Efficient foundation models for PDEs with pretrained transformer-based neural operators and downstream task fine-tuning pipelines, HuggingFace integration for models and datasets (ETH Zurich CAMLab, arXiv 2024)
- [GAOT (NeurIPS 2025)](https://github.com/camlab-ethz/GAOT) - Geometry Aware Operator Transformer serving as an efficient and accurate neural surrogate for PDEs on arbitrary domains, combining geometric priors with transformer architectures for scientific computing (ETH Zurich CAMLab, 92+ stars)
- [PhiFlow](https://github.com/tum-pbs/PhiFlow) - Differentiable PDE solving framework for machine learning with built-in fluid simulation, supporting PyTorch/JAX/TensorFlow backends and enabling neural network training within physical simulations (TUM, MIT License)
- [exponax](https://github.com/Ceyron/exponax) - Efficient differentiable n-dimensional PDE solvers built on JAX and Equinox, shipping 46+ built-in equations with Fourier spectral methods, exponential time differencing, and full auto-differentiation for physics-based deep learning workflows (MIT, 200+ stars, 2024)
### Simulation-Based Inference
- [sbi](https://github.com/sbi-dev/sbi) - Python package for simulation-based inference enabling likelihood-free Bayesian parameter estimation from scientific simulators, with flexible interfaces for neural posterior estimation, sequential methods, and MCMC/variational backends (Mackelab, 825+ stars)
---
## 📖 Papers & Reviews
### Foundational Papers
- [Machine Learning for Scientometric Analysis](https://arxiv.org/abs/2109.10073) (2021.09) - Comprehensive review
- [AI for Science: Progress and Challenges](https://arxiv.org/abs/2303.04346) (2023.03) - State of the field
- [Foundation Models for Science](https://arxiv.org/abs/2205.15075) (2022.05) - Large models in research
- [Neural Ordinary Differential Equations](https://arxiv.org/abs/1806.07366) (2018.06) - Breakthrough in neural ODEs
- [Physics-Informed Neural Networks](https://arxiv.org/abs/1711.10561) (2017.11) - Physics-constrained deep learning
- [Scientific Discovery in the Age of Artificial Intelligence](https://www.nature.com/articles/s41586-023-06221-2) - Nature review on AI's role in science
### 📊 Comprehensive Surveys & Reviews (2024-2025)
#### AI for Scientific Research
- [A Survey on AI-assisted Scientific Discovery](https://arxiv.org/abs/2502.05151) (2025.02) - Comprehensive overview of LLMs in scientific research lifecycle from literature search to peer review
- [AI4Research: A Survey of Artificial Intelligence for Scientific Research](https://arxiv.org/abs/2507.01903) (2025.07) - Systematic taxonomy of AI in research
- [Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems](https://arxiv.org/abs/2307.08423) (2023.07) - Unified technical survey across scientific scales with 63 contributors
- [From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery](https://arxiv.org/abs/2505.13259) (2025.05) - Three-level taxonomy (Tool, Analyst, Scientist)
- [From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery](https://arxiv.org/abs/2508.14111) (2025.08) - Comprehensive survey on agentic science across life sciences, chemistry, materials, and physics
- [Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions](https://arxiv.org/abs/2503.08979) (2025.03) - Comprehensive review of AI agents in science
- [Towards Scientific Intelligence: A Survey of LLM-based Scientific Agents](https://arxiv.org/abs/2503.24047) (2025.03) - Scientific AI agent systems
#### Scientific Large Language Models
- [A Comprehensive Survey of Scientific Large Language Models and Their Applications](https://arxiv.org/abs/2406.10833) (2024.06) - 260+ scientific LLMs across domains
- [A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers](https://arxiv.org/abs/2508.21148) (2025.08) - Data-centric view of scientific LLMs
- [Scientific Large Language Models: A Survey on Biological & Chemical Domains](https://arxiv.org/abs/2401.14656) (2024.01) - Domain-specific scientific LLMs
#### Scientific Machine Learning
- [Scientific Machine Learning through Physics-Informed Neural Networks: Where we are and What's next](https://arxiv.org/abs/2201.05624) (2022.01) - Comprehensive PINN review
- [Physics-Informed Neural Networks and Extensions](https://arxiv.org/abs/2408.16806) (2024.08) - Recent PINN advances and variants
- [The frontier of simulation-based inference](https://www.pnas.org/doi/10.1073/pnas.1912789117) (PNAS 2020) - Foundational review on SBI for scientific computing by Cranmer et al.
- [From Theory to Application: A Practical Introduction to Neural Operators in Scientific Computing](https://arxiv.org/abs/2503.05598) (2025.03) - Implementation-focused guide to DeepONet, FNO, and PCANet
- [Architectures, variants, and performance of neural operators: A comparative review](https://www.sciencedirect.com/science/article/abs/pii/S0925231225011907) (2025) - Systematic analysis of DeepONets, integral kernel operators, and transformer-based neural operators
- [Foundation Models for Environmental Science: A Survey](https://arxiv.org/abs/2504.04280) (2025.04) - Environmental applications
- [Foundation Models in Bioinformatics](https://academic.oup.com/nsr/article/12/4/nwaf028/7979309) - Biological foundation models
- [Foundation Models for Materials Discovery](https://www.nature.com/articles/s41524-025-01538-0) (2025) - Perspective on materials AI
#### Uncertainty Quantification
- [Uncertainty quantification in scientific machine learning: Methods, metrics, and comparisons](https://www.sciencedirect.com/science/article/abs/pii/S0021999122009652) (J. Comput. Phys. 2023) - Comprehensive framework for UQ in PINNs and neural operators by Psaros et al.
- [A Survey on Uncertainty Quantification Methods for Deep Learning](https://arxiv.org/abs/2302.13425) (2023) - Systematic taxonomy of UQ methods from uncertainty source perspective
#### Automation & Self-Driving Laboratories
- [Self-Driving Laboratories for Chemistry and Materials Science](https://pubs.acs.org/doi/10.1021/acs.chemrev.4c00055) (Chem. Rev. 2024) - Comprehensive 100-page review on SDL technology, applications, and infrastructure
- [Autonomous 'self-driving' laboratories: a review of technology and policy implications](https://royalsocietypublishing.org/doi/10.1098/rsos.250646) (Royal Soc. Open Sci. 2025) - Technology review with policy and safety considerations
#### Policy & Strategic Perspectives
- [Artificial Intelligence for Science](https://www.csiro.au/-/media/d61/ai4science-report/ai-for-science-report-2022.pdf) (CSIRO 2022) - Landmark report analyzing AI adoption across 98% of scientific fields over 60 years
- [AI for Science 2025](https://www.nature.com/articles/d42473-025-00161-3) (Fudan University & Nature 2025) - Comprehensive report on AI's transformative impact across 7 scientific fields, 28 research directions, and 90+ challenges
- [AI in science evidence review](https://scientificadvice.eu/scientific-outputs/ai-in-science-evidence-review-report/) (European Scientific Advice 2024) - Policy-focused evidence review on AI's impact in research
### 🚀 AI Scientist & Autonomous Research (2024-2025 Breakthroughs)
- [The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery](https://arxiv.org/abs/2408.06292) (2024.08) - First fully autonomous research system
- [The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search](https://arxiv.org/abs/2504.08066) (2025.04) - Enhanced autonomous research with agentic tree search
- [AI-Researcher: Autonomous Scientific Innovation](https://arxiv.org/abs/2505.18705) (2025.05) - Autonomous research pipeline from literature to publication with Scientist-Bench evaluation framework
- [InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification](https://arxiv.org/abs/2505.16938) (2025.05) - Multi-agent system achieving #1 on MLE-Bench with closed-loop research automation
- [Autonomous Scientific Discovery Through Hierarchical AI Scientist Systems](https://arxiv.org/abs/2507.15951) (2025.07) - Self-evolving multi-agent research systems
- [ChemCrow: Augmenting large-language models with chemistry tools](https://arxiv.org/abs/2304.05376) (2023.04) - LLM agents for chemistry research
- [Autonomous chemical research with large language models](https://www.nature.com/articles/s41586-023-06792-0) - Automated chemical experimentation
- [Coscientist: Autonomously planning and executing scientific experiments](https://www.nature.com/articles/s41586-023-06792-1) - Robotic lab automation
### Recent Advances & Domain Applications
- [AlphaFold: Protein Structure Prediction](https://www.nature.com/articles/s41586-021-03819-2)
- [AI for Materials Discovery](https://www.nature.com/articles/s41578-023-00540-6)
- [Large Language Models in Chemistry](https://arxiv.org/abs/2402.05852) (2024.02)
- [Cell2Sentence: Teaching Large Language Models the Language of Biology](https://arxiv.org/abs/2405.06147) (ICML 2024) - LLMs for single-cell transcriptomics
- [Scaling Large Language Models for Next-Generation Single-Cell Analysis](https://www.biorxiv.org/content/10.1101/2025.04.14.648850v2) (2025.04) - 27B parameter biological language models
- [Boltz-1: Democratizing Biomolecular Interaction Modeling](https://www.biorxiv.org/content/10.1101/2024.11.19.624167v2) (bioRxiv 2024) - First fully open-source model achieving AlphaFold3-level accuracy
- [MOOSE: Large Language Models for Automated Open-domain Scientific Hypotheses Discovery](https://arxiv.org/abs/2309.02726) (ACL 2024) - First work showing LLMs can generate novel and valid scientific hypotheses, ICML Best Poster Award
- [Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents](https://arxiv.org/abs/2509.23141) (2025.09) - LLM agent framework for Earth Observation with 104 specialized tools and multi-modal analysis
- [MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning](https://arxiv.org/abs/2311.10537) (ACL 2024) - Multi-disciplinary collaboration framework for medical reasoning using role-playing LLM agents
- [MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science](https://arxiv.org/abs/2506.04405) (2025.06) - Specialized training environment for biomedical AI agents with code-centric reasoning
- [Paper2Web: Let's Make Your Paper Alive!](https://arxiv.org/abs/2510.15842) (2025.10) - AI-powered transformation of academic papers into interactive websites with comprehensive evaluation framework
- [DeepAnalyze: Agentic Large Language Models for Autonomous Data Science](https://arxiv.org/abs/2510.16872) (2025.10) - First agentic LLM for autonomous data science with curriculum-based training
- [Democratizing AI scientists using ToolUniverse](https://arxiv.org/abs/2509.23426) (2025.09) - Universal ecosystem for building AI scientists from any LLM with 600+ scientific tools
- [TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools](https://arxiv.org/abs/2503.10970) (2025.03) - AI agent achieving 92.1% accuracy in drug reasoning, outperforming GPT-4o by 25.8%
- [Aviary: Training Language Agents on Challenging Scientific Tasks](https://arxiv.org/abs/2412.21154) (2024.12) - Language agent training framework for scientific discovery
- [Galactica: A Large Language Model for Science](https://arxiv.org/abs/2211.09085) (2022.11)
### 📈 Evaluation & Benchmarking
- [ScienceAgentBench (ICLR 2025)](https://github.com/OSU-NLP-Group/ScienceAgentBench) - 102 executable tasks from 44 peer-reviewed papers across 4 disciplines with containerized evaluation
- [Scientist-Bench](https://github.com/HKUDS/AI-Researcher) - Comprehensive benchmark for comparing LLM Agent-generated research outcomes with high-quality scientific work
- [SciTrust: Evaluating the Trustworthiness of Large Language Models for Science](https://impact.ornl.gov/en/publications/scitrust-evaluating-the-trustworthiness-of-large-language-models-) (2024) - Scientific LLM trustworthiness evaluation framework
- [SciBench: Evaluating College-Level Scientific Problem-Solving Abilities](https://arxiv.org/abs/2307.10635) (2023) - Scientific reasoning benchmarks
- [ChartCoder Evaluation](https://aclanthology.org/2025.acl-long.363/) - Chart-to-code generation benchmarks
---
## 🔬 Domain-Specific Applications
### 🧬 Biology & Medicine
#### Protein & Drug Discovery
- [CryoDRGN](https://github.com/ml-struct-bio/cryodrgn) - Neural network-based cryo-EM heterogeneous reconstruction, modeling continuous 3D structure distributions from single-particle images, with CryoDRGN-ET extending to in-cell cryo-electron tomography (MIT CSAIL, Nature Methods 2021/2024)
- [ModelAngelo](https://github.com/3dem/model-angelo) - Automatic atomic model building program for cryo-EM maps using deep learning, enabling rapid de novo protein structure determination from electron density with high accuracy (3DEM/EMBL, 169+ stars)
- [AlphaFold](https://github.com/google-deepmind/alphafold) - Protein structure prediction
- [AlphaFold3](https://github.com/google-deepmind/alphafold3) - AlphaFold 3 inference pipeline for unified biomolecular structure prediction of proteins, nucleic acids, small molecules, ions, and post-translational modifications (Google DeepMind, Nature 2024)
- [AlphaFold Server](https://alphafoldserver.com/) - Free, easy-to-use web platform by Google DeepMind and Isomorphic Labs for running AlphaFold 3 predictions of biomolecular structures and interactions, enabling researchers without local infrastructure to model proteins, nucleic acids, small molecules, ions, and post-translational modifications through a searchable proteome interface (2024)
- [AlphaProteo](https://github.com/google-deepmind/alphaproteo) - Deep learning system for de novo design of high-affinity protein binders, achieving strong binding across diverse target classes including challenging intracellular proteins with significantly higher success rates than traditional wet-lab screening methods (Google DeepMind, Nature 2024)
- [AlphaPulldown](https://github.com/KosinskiLab/AlphaPulldown) - Automated pipeline for proteome-scale protein-protein interaction screening with AlphaFold-Multimer and AlphaFold 3, supporting flexible inputs (UniProt IDs, FASTA, residue regions, multimers, AF3 JSON features) and integrated downstream analysis for hit prioritization (Kosinski Lab, EMBL, Nature Protocols 2024, 317+ stars, GPL-3.0)
- [RareFold](https://github.com/PatrickBryant1/RareFold) - Structure prediction and design of proteins with noncanonical amino acids, enabling AI-powered modeling of synthetic biology constructs and expanded genetic code systems (133+ stars, 2025)
- [ColabFold (2025 Updates)](https://github.com/sokrypton/ColabFold) - AlphaFold/ESMFold accessible implementation with AF3 JSON export, database updates
- [OpenFold](https://github.com/aqlaboratory/openfold) - Trainable, memory-efficient PyTorch reproduction and retraining of AlphaFold2 providing new insights into its learning dynamics and out-of-distribution generalization; widely used as the open-source AlphaFold2 backbone underpinning many downstream protein structure prediction and design pipelines (Columbia AlQuraishi Lab & OpenFold Consortium, Nature Methods 2024)
- [OpenFold3](https://github.com/aqlaboratory/openfold-3) - Fully open-source (Apache 2.0) biomolecular structure prediction reproducing AlphaFold3, free for academic and commercial use (Columbia AlQuraishi Lab & OpenFold Consortium, 2025)
- [Protenix](https://github.com/bytedance/Protenix) - Trainable PyTorch reproduction of AlphaFold 3
- [HelixFold3](https://github.com/PaddlePaddle/PaddleHelix/tree/dev/apps/helixfold3) - Baidu's open-source reproduction of AlphaFold3 in PaddlePaddle, providing pretrained weights and inference pipelines for unified biomolecular structure prediction across proteins, nucleic acids, ligands, ions, and post-translational modifications within the PaddleHelix biocomputing platform (Baidu, bioRxiv 2024)
- [RoseTTAFold-All-Atom](https://github.com/baker-laboratory/RoseTTAFold-All-Atom) - All-atom biomolecular structure prediction for protein-nucleic acid-small molecule-metal ion complexes, enabling accurate modeling of covalent modifications and assemblies beyond proteins (Baker Lab, Science 2024)
- [Chai-1](https://github.com/chaidiscovery/chai-lab) - Multi-modal foundation model for biomolecular structure prediction (proteins, small molecules, DNA, RNA, glycans) achieving SOTA across benchmarks, with optional MSA/template support (Chai Discovery, 2024)
- [IntelliFold](https://github.com/IntelliGen-AI/IntelliFold) - Controllable foundation model for general and specialized biomolecular structure prediction across proteins, nucleic acids, and complexes, featuring a public web server for interactive prediction workflows (IntelliGen AI, 223+ stars, Apache 2.0, 2025)
- [SimpleFold (Apple, arXiv 2025)](https://github.com/apple/ml-simplefold) - Flow-matching protein folding model using only general-purpose transformer layers, scaled to 3B parameters and trained on 8.6M+ distilled structures; challenges the reliance on complex domain-specific architectures and supports PyTorch and MLX backends with model sizes from 100M to 3B parameters (985+ stars, MIT License)
- [NeuralPLexer](https://github.com/zrqiao/NeuralPLexer) - State-specific protein-ligand complex structure prediction with a multi-scale deep generative model, enabling conformational state-aware modeling of molecular interactions (329+ stars, 2024)
- [Boltz](https://github.com/jwohlwend/boltz) - First fully open-source model achieving AlphaFold3-level accuracy with 1000x faster binding affinity prediction (MIT)
- [Boltz-2 (MIT & Recursion, 2025)](https://github.com/jwohlwend/boltz) - Next-generation biomolecular foundation model jointly predicting protein-ligand complex structures and binding affinities in a single framework; achieves FEP-level accuracy with ~0.62 Pearson correlation on FEP+ benchmark in ~20 seconds, outperforming all methods at CASP16 affinity challenge and doubling average precision in MF-PCBA hit-discovery screens (MIT License)
- [BoltzGen](https://arxiv.org/abs/2511.18345) - De novo protein binder design via generative model, achieving nanomolar binding for 66% of novel targets tested (MIT, 2025)
- [Proteina-Complexa](https://github.com/NVIDIA-Digital-Bio/Proteina-Complexa) - Flow-based generative model for atomistic protein binder design with test-time optimization, SOTA on binder benchmarks (ICLR 2026 Oral, NVIDIA)
- [PXDesign (ByteDance, 2025)](https://github.com/bytedance/PXDesign) - Fast, modular, and accurate de novo design of protein binders based on the Protenix foundation model, achieving 17-82% nanomolar hit rates across diverse targets with 2-6× improvement over prior methods like AlphaProteo and RFdiffusion (229+ stars, Apache 2.0)
- [ODesign (OTeam-AI4S, 2025)](https://github.com/OTeam-AI4S/ODesign) - All-atom generative world model for all-to-all biomolecular interaction design, enabling cross-modality generation of proteins, nucleic acids, small molecules, and cyclic peptides with fine-grained epitope-level control and 2-4 orders of magnitude faster design throughput than modality-specific baselines (316+ stars, Apache 2.0)
- [OpenDDE (Aureka Research, 2026)](https://github.com/aurekaresearch/OpenDDE) - Open-source, all-atom biomolecular foundation model that turns co-folding into a scalable engine for structure prediction, design, and optimization across proteins, nucleic acids, and small molecules in drug discovery; ranked first on PXMeter-AB, FoldBench-AB, and 2026ARK-AB antibody-antigen benchmarks (263+ stars, Apache 2.0)
- [La-Proteina (NVIDIA)](https://github.com/NVIDIA-Digital-Bio/la-proteina) - Partially latent flow matching model for the joint generation of a protein's amino acid sequence and full atomistic structure, including both backbone and side chains (2025)
- [Proteina (NVIDIA, ICLR 2025 Oral)](https://github.com/NVIDIA-Digital-Bio/proteina) - Large-scale flow-based protein backbone generator utilizing hierarchical fold class labels for conditioning with a tailored scalable transformer architecture, enabling controllable de novo protein design (264+ stars)
- [xfold](https://github.com/Shenggan/xfold) - Democratizing AlphaFold3: PyTorch reimplementation to accelerate protein structure prediction research
- [MegaFold](https://github.com/Supercomputing-System-AI-Lab/MegaFold/) - Cross-platform system optimizations for accelerating AlphaFold3 training with 1.73x speedup and 1.23x memory reduction
- [Graphormer](https://github.com/microsoft/Graphormer) - General-purpose deep learning backbone for molecular modeling
- [DiffDock](https://github.com/gcorso/DiffDock) - Diffusion-based molecular docking achieving SOTA blind docking performance, treating ligand pose prediction as generative diffusion over SE(3), with DiffDock-L update for improved generalization (MIT CSAIL, ICLR 2023)
- [GNINA](https://github.com/gnina/gnina) - Deep learning framework for molecular docking extending AutoDock Vina with convolutional neural network scoring functions, achieving superior virtual screening enrichment and pose prediction across diverse target classes; widely adopted in pharmaceutical structure-based drug design (J. Cheminformatics, 915+ stars, actively maintained)
- [DynamicBind (NeurIPS 2024)](https://github.com/luwei0917/DynamicBind) - Deep equivariant generative model predicting ligand-specific protein-ligand complex structures with dynamic receptor conformational flexibility, enabling accurate docking for flexible protein targets
- [PLACER](https://github.com/baker-laboratory/PLACER) - Graph neural network operating entirely at the atomic level for protein-ligand conformational ensemble prediction and docking, generating diverse solutions through rapid stochastic denoising to model conformational heterogeneity (Baker Lab, bioRxiv 2025)
- [targetdiff](https://github.com/guanjq/targetdiff) - 3D Equivariant Diffusion for Target-Aware Molecule Generation (ICLR2023)
- [SeFMol (Science Advances 2026)](https://github.com/ispc-lab/SeFMol) - Semi-flexible molecular diffusion model for structure-based drug design with reinforcement learning, achieving 20× faster sampling and providing a no-code web platform for molecular design (ISPC Lab, Tongji University, 2026)
- [ReQFlow](https://github.com/AngxiaoYue/ReQFlow) - Rectified Quaternion Flow for efficient protein backbone generation, 37× faster than RFDiffusion with 0.972 designability (ICML 2025)
- [AlphaFlow](https://github.com/bjing2016/alphaflow) - AlphaFold fine-tuned with flow matching for generating protein conformational ensembles, covering both experimental PDB states and molecular dynamics ensembles at physiological temperatures; includes ESMFlow variant (MIT, 526+ stars, 2024)
- [BioEmu](https://github.com/microsoft/bioemu) - Microsoft's generative model for sampling protein equilibrium conformations 100,000× faster than MD simulations, predicting domain motions, local unfolding and cryptic binding pockets on a single GPU (Science 2025)
- [STARLING (Holehouse Lab, Nature 2026)](https://github.com/idptools/starling) - Latent-space probabilistic denoising diffusion model for predicting coarse-grained conformational ensembles of intrinsically disordered proteins and regions from sequence, with GPU/CPU inference, trajectory export, and FAISS-based similarity search (67+ stars, LGPL-3.0)
- [dynamicPDB (AAAI 2025)](https://github.com/fudan-generative-vision/dynamicPDB) - Dynamic Protein Data Bank integrating dynamic behaviors and physical properties into protein structures via a new dataset and SE(3) model extension, enabling richer understanding of protein conformational landscapes (Fudan University, 784+ stars)
- [ProteinMPNN](https://github.com/dauparas/ProteinMPNN) - Deep learning-based protein sequence design (inverse folding) from backbone structures, achieving 52.4% sequence recovery vs 32.9% for Rosetta, core tool in modern protein design pipelines (Baker Lab, Science 2022)
- [LigandMPNN](https://github.com/dauparas/LigandMPNN) - Extension of ProteinMPNN for protein sequence design in the context of small-molecule ligands, metal ions, and nucleic acids, enabling binding site engineering and co-factor redesign (Baker Lab)
- [ColabDesign](https://github.com/sokrypton/ColabDesign) - Accessible protein design platform via Google Colab integrating AlphaFold2, RoseTTAFold, and ProteinMPNN for de novo hallucination, fixed backbone design, and binder design (Sergey Ovchinnikov, 2022+)
- [BindCraft](https://github.com/martinpacesa/BindCraft) - Simple and accurate de novo protein binder design pipeline using AlphaFold2 backpropagation, MPNN, and PyRosetta for automated binder discovery (bioRxiv 2024)
- [Genie 2](https://github.com/aqlaboratory/genie2) - Diffusion model for scalable protein structure design with multi-motif scaffolding capabilities, achieving state-of-the-art designability, diversity, and novelty through SE(3)-equivariant attention and massive data augmentation (AlQuraishi Lab, 2024)
- [Genie 3 (AlQuraishi Lab, 2026)](https://github.com/aqlaboratory/genie3) - Fast, all-atom SE(3)-equivariant diffusion model for protein design achieving state-of-the-art performance on unconditional generation, motif scaffolding, and binder design while retaining the computational efficiency of equivariant architectures (bioRxiv 2026)
- [Chroma](https://github.com/generatebio/chroma) - Generative model for programmable protein design using diffusion modeling, equivariant graph neural networks, and conditional random fields to efficiently sample diverse all-atom structures; supports conditional generation via composable conditioners for substructure, symmetry, shape, and neural-network predictions; validated crystallographically (Generate Biomedicines, Nature 2023)
- [EvoDiff](https://github.com/microsoft/evodiff) - Discrete diffusion framework for generative protein sequence design over evolutionary-scale databases, supporting unconditional generation, evolutionary-guided conditional design, motif scaffolding, and intrinsically disordered region generation through order-agnostic autoregressive diffusion, enabling sequence-only protein design without structural priors (Microsoft Research, Nature Communications 2024)
- [DISCO](https://github.com/DISCO-design/DISCO) - General multimodal protein design framework enabling DNA-encoding of chemistry for programmable enzyme design and diverse protein generation through diffusion-based generative modeling (190+ stars, Apache 2.0, 2026)
- [SwitchCraft](https://github.com/bjing2016/switchcraft) - Programmatic framework for designing state-switching proteins via backpropagation through compositional design constraints parameterized by structure prediction models; enables de novo design of allosteric regulators and fluorescent biosensors for arbitrary small-molecule analytes (79+ stars, MIT License, ICML 2026)
- [RFdiffusion3](https://github.com/RosettaCommons/RFdiffusion) - Latest RFdiffusion for protein structure design with 10× speedup and atom-level precision (December 2025)
- [RFantibody](https://github.com/RosettaCommons/RFantibody) - Structure-based de novo antibody design pipeline built on RFdiffusion for computational generation of target-specific antibodies (RosettaCommons, 2025)
- [IgGM](https://github.com/TencentAI4S/IgGM) - Generative foundation model for functional antibody and nanobody design, supporting de novo generation, affinity maturation, inverse design, structure prediction, and humanization (Tencent AI4S, ICLR 2025)
- [DrugAssist](https://github.com/blazerye/DrugAssist) - LLM-based molecular optimization tool
- [GenMol](https://github.com/NVIDIA-Digital-Bio/genmol) - ICML 2025 drug discovery generalist using masked discrete diffusion and fragment-based generation with molecular context guidance (NVIDIA)
- [REINVENT](https://github.com/MolecularAI/Reinvent) - Industrial-grade reinforcement-learning-based generative platform for de novo molecular design with transformer architectures, supporting multi-objective optimization, scaffold decoration, and curriculum learning (AstraZeneca MolecularAI, REINVENT 4, 2024)
- [mint](https://github.com/VarunUllanat/mint) - Learning the language of protein-protein interactions
- [Mol-Instructions](https://github.com/zjunlp/Mol-Instructions) - Large-scale biomolecular instruction dataset for chemistry/biology LLMs (ICLR2024)
- [Uni-Mol](https://github.com/deepmodeling/Uni-Mol) - Universal 3D molecular pretraining framework with 209M conformations, scaling to 1.1B parameters (Uni-Mol2) on 800M conformations for molecular property prediction, docking, and quantum chemistry (ICLR 2023, NeurIPS 2024)
- [ChemBERTa](https://github.com/seyonechithrananda/bert-loves-chemistry) - Chemical language model
- [DeepChem](https://github.com/deepchem/deepchem) - Machine learning for chemistry
- [TorchDrug](https://github.com/DeepGraphLearning/torchdrug) - Powerful and flexible machine learning platform for drug discovery, providing comprehensive tools for molecular property prediction, generative models, knowledge graph reasoning, and reaction prediction with PyTorch backend (1.5K+ stars)
- [DeepMol](https://github.com/BioSystemsUM/DeepMol) - Unified ML/DL framework for drug discovery workflows, integrating RDKit, DeepChem, and scikit-learn with SHAP explainability
- [Chemprop](https://github.com/chemprop/chemprop) - Message passing neural networks for molecule property prediction, ADMET modeling, and reaction prediction, achieving SOTA on MoleculeNet and widely used in pharmaceutical drug discovery (MIT, 2.3K+ stars)
- [RDKit](https://github.com/rdkit/rdkit) - Cheminformatics toolkit
- [nvMolKit (NVIDIA BioNeMo, 2025)](https://github.com/NVIDIA-BioNeMo/nvMolKit) - High-performance, GPU-accelerated library for key computational chemistry tasks including molecular similarity, conformer generation, and geometry relaxation, designed to accelerate drug-discovery and molecular-modeling workflows (264+ stars, Apache 2.0)
- [Open Targets](https://www.opentargets.org/) - Open-source data integration platform for systematic drug target identification and prioritization, combining genetics, genomics, chemistry, and pharmacology data from EMBL-EBI, Wellcome Sanger Institute, and pharmaceutical partners to accelerate therapeutic discovery
- [ESM3](https://github.com/evolutionaryscale/esm) - 98B-parameter frontier generative model jointly reasoning over protein sequence, structure, and function, trained on 2.78 billion proteins; generated a novel fluorescent protein (esmGFP) with only 58% sequence identity to known GFPs (EvolutionaryScale, 2024)
- [ProtTrans](https://github.com/agemagician/ProtTrans) - State-of-the-art pretrained language models for proteins trained on thousands of GPUs and Google TPUs using Transformer architectures, enabling protein property prediction, feature extraction, and transfer learning across diverse downstream tasks (1.3K+ stars, MIT, 2020-2026)
- [ProstT5 (NAR Genomics and Bioinformatics 2024)](https://github.com/mheinzinger/ProstT5) - Bilingual protein language model translating between protein sequence and structure, finetuned from ProtT5-XL on 17M AlphaFoldDB structures using Foldseek's 3Di structural alphabet, enabling sequence-to-structure prediction, structure-to-sequence inverse folding, and unified protein representation learning (RostLab, 310+ stars)
- [ESMFold](https://github.com/facebookresearch/esm) - Protein structure prediction from ESM models
- [SaProt](https://github.com/westlake-repl/SaProt) - Structure-aware protein language model using 3D structural vocabulary (Foldseek) for joint sequence-structure pretraining, achieving SOTA on protein engineering and fitness prediction benchmarks (ICML 2024, Westlake University & Repl)
- [InterPLM (Nature Methods 2025)](https://github.com/ElanaPearl/InterPLM) - Discovering interpretable features in protein language models via sparse autoencoders, enabling mechanistic understanding of PLM representations for protein engineering and design (288+ stars, MIT License)
- [AiCE (Cell 2025)](https://github.com/ScorpioLea/AiCE) - AI-assisted mutation nomination approach optimizing protein function by integrating structural and evolutionary constraints into protein inverse folding models, compatible with ProteinMPNN, LigandMPNN, ESM-IF1, and SaProt (Chinese Academy of Sciences, 359+ stars)
- [EVOLVEpro](https://github.com/mat10d/EvolvePro) - In silico directed evolution framework using few-shot active learning to optimize protein activities, enabling rapid protein engineering with minimal experimental data (352+ stars, 2023)
- [DPLM (ByteDance, ICML 2024 / ICLR 2025)](https://github.com/bytedance/dplm) - Family of diffusion protein language models demonstrating versatile generative and predictive capabilities for protein sequences and structures, including multimodal co-generation, conditional folding, inverse folding, motif scaffolding, and representation learning, with open pretrained weights and training scripts (327+ stars, ICML 2024, ICLR 2025, ICML 2025 Spotlight)
- [Foldseek](https://github.com/steineggerlab/foldseek) - Fast and accurate protein structure search using a learned 3Di structural alphabet (VQ-VAE) that discretizes tertiary interactions into structural tokens, enabling protein-universe-scale structural alignment at sequence-search speeds (4-5 orders of magnitude faster than DALI/TM-align) and underpinning many AI4S tools such as SaProt, ESMAtlas search, and AFDB clustering pipelines (Steinegger Lab, Nature Biotechnology 2023)
- [ImmunoStruct (Nature Machine Intelligence 2025)](https://github.com/KrishnaswamyLab/ImmunoStruct) - Multimodal deep learning framework integrating peptide-MHC protein sequence, structure, and biochemical properties to predict class-I immunogenicity for infectious disease epitopes and cancer neoepitopes with cancer-wildtype contrastive learning, enabling personalized vaccine design (Krishnaswamy Lab, Yale University)
- [mosaic](https://github.com/escalante-bio/mosaic) - Composite-objective protein design framework integrating Boltz, AlphaFold2, OpenFold3, ProteinMPNN, and ESM via JAX-based gradient optimization over continuous relaxed sequence space for multi-property binder design (319+ stars, MIT License, 2025)
#### Genomics & Bioinformatics
- [RhoFold+](https://github.com/ml4bio/RhoFold) - End-to-end RNA 3D structure prediction using RNA language model pretrained on 23.7M sequences, outperforming existing methods and human expert groups on RNA-Puzzles and CASP15 (Nature Methods 2024)
- [NuFold (Nature Communications 2025)](https://github.com/kiharalab/NuFold) - End-to-end deep learning approach for RNA tertiary structure prediction with a flexible nucleobase center representation, achieving ~7 Å C1' RMSD across test RNAs and predicting ~545,000 structures covering 2,200+ RNA families (Kihara Lab, Purdue University, 50+ stars)
- [RNA-FM (Nature Methods 2024)](https://github.com/ml4bio/RNA-FM) - RNA foundation model trained on millions of RNA sequences for generalist RNA sequence understanding, enabling downstream structure prediction, function annotation, and representation learning for non-coding RNAs (ml4bio, 372+ stars)
- [RiNALMo (Nature Communications 2025)](https://github.com/lbcb-sci/RiNALMo) - General-purpose RNA language model with 650M parameters pretrained on 36M non-coding RNA sequences, achieving strong generalization on structure prediction tasks including secondary structure prediction, splice-site prediction, mean ribosome loading, and ncRNA classification (lbcb-sci, 165+ stars, Apache-2.0)
- [RNAPro (NVIDIA, 2026)](https://github.com/NVIDIA-Digital-Bio/RNAPro) - State-of-the-art RNA 3D folding model developed with Stanford Das Lab and Kaggle competition winners, featuring a 488M-parameter AF3-like architecture with MSA and template-based modeling, enabling structure-driven drug discovery and RNA therapeutics design (NVIDIA-Digital-Bio, Apache 2.0)
- [gRNAde](https://github.com/chaitjo/geometric-rna-design) - Generative AI framework for inverse design of 3D RNA structure and function using geometric deep learning, learning design rules from 3D structures to capture complex tertiary interactions (pseudoknots, non-canonical base pairs) with expert-level accuracy for designing functional RNAs including aptamers and ribozymes (bioRxiv 2025)
- [AIDO.ModelGenerator](https://github.com/genbio-ai/ModelGenerator) - GenBio AI's software stack for the AI-Driven Digital Organism, supporting adaptation and finetuning of multiscale biological foundation models across DNA, RNA, protein, structure, and single-cell tasks with reproducible CLIs and pretrained model zoo (2025)
- [Evo 2](https://github.com/ArcInstitute/evo2) - Arc Institute's 40B-parameter genome foundation model trained on 9 trillion nucleotides from all domains of life, supporting 1M base pair context for generalist DNA/RNA/protein prediction and design (Nature 2026)
- [Carbon (Hugging Face, 2026)](https://github.com/huggingface/carbon) - Family of causal genomic foundation models trained on 1T tokens (~6T DNA base pairs) from the Carbon Pretraining Corpus, combining eukaryote genes, mRNA transcripts, and prokaryote genomes with a hybrid text/6-mer tokenizer; Carbon-3B matches or beats Evo2-7B on zero-shot DNA evaluations including sequence recovery, variant effect prediction, and perturbations (Apache 2.0, 201+ stars)
- [Nucleotide Transformer](https://github.com/instadeepai/nucleotide-transformer) - Foundation models for genomics and transcriptomics pretrained on 3,000+ human genomes and 850+ diverse species, enabling chromatin accessibility prediction, splice site detection, and promoter classification across multiple model scales (InstaDeep, NVIDIA & TUM, Nature Methods 2023)
- [HyenaDNA](https://github.com/HazyResearch/hyena-dna) - Long-range genomic foundation model using subquadratic Hyena operators instead of Transformer attention, enabling context lengths up to 1 million nucleotides for chromosome-scale DNA sequence modeling and downstream genomics tasks (Stanford Hazy Research, NeurIPS 2023, 784+ stars, Apache 2.0)
- [Caduceus (ICML 2024)](https://github.com/kuleshov-group/caduceus) - Bi-directional DNA language model based on the Mamba state space architecture, enabling efficient long-range genomic sequence modeling with linear-time complexity and built-in reverse-complement equivariance; achieves strong performance on chromatin accessibility, enhancer, and promoter prediction benchmarks (Stanford & UC Berkeley, 500+ stars)
- [CodonFM (NVIDIA)](https://github.com/NVIDIA-Digital-Bio/CodonFM) - Family of codon-resolution language models trained on 130 million protein-coding sequences from over 20,000 species, enabling cross-species gene expression prediction and codon-level functional genomics (2025)
- [LucaOne](https://github.com/LucaOne/LucaOne) - Generalized biological foundation model with unified nucleic acid and protein language, integrating DNA/RNA/protein sequences (Nature Machine Intelligence 2025)
- [Geneformer](https://github.com/lcrawlab/Geneformer) - Single-cell transformer foundation model pretrained on 104M human transcriptomes via masked gene prediction, enabling transfer learning for cell type classification, gene network analysis, and in silico perturbation with limited labeled data (Nature 2023, V2 2024)
- [Nicheformer](https://github.com/theislab/nicheformer) - Foundation model jointly trained on single-cell and spatial transcriptomics data, enabling unified representation learning across cellular and tissue spatial contexts for cell type prediction, spatial domain inference, and cross-modal integration (theislab, bioRxiv 2024, 164+ stars)
- [scFoundation](https://github.com/biomap-research/scFoundation) - 100M-parameter foundation model pretrained on 50M+ human single-cell transcriptomes covering ~20,000 genes, achieving SOTA on gene expression enhancement, drug response and perturbation prediction (Nature Methods 2024)
- [scPRINT (Nature Communications 2025)](https://github.com/cantinilab/scPRINT) - Large transformer-based single-cell foundation model pretrained on 50 million cells for robust gene network inference, expression denoising, cell embedding, and zero-shot label prediction, leveraging ESM2 protein embeddings and bidirectional transformer architecture (Cantini Lab, 148+ stars, GPL-3.0)
- [Tahoe-x1](https://github.com/tahoebio/tahoe-x1) - Apache 2.0 single-cell foundation model family scaling to 3B parameters, pretrained on 266M cell profiles including perturbation data and released with training, embedding, and downstream benchmarking workflows for disease-relevant single-cell tasks (2025)
- [Stack](https://github.com/ArcInstitute/stack) - Arc Institute's single-cell foundation model enabling in-context learning at inference time via a novel tabular attention architecture, trained on 150M uniformly-preprocessed cells for generalizing biological effects and generating unseen cell profiles in novel contexts (2025)
- [State (Arc Institute, bioRxiv 2025)](https://github.com/ArcInstitute/state) - Machine learning model predicting cellular perturbation response across diverse contexts with State Transition (ST) and State Embedding (SE) variants, featuring CLI tooling, PyPI distribution, and Virtual Cell Challenge integration (575+ stars)
- [scvi-tools](https://github.com/scverse/scvi-tools) - Deep probabilistic framework for single-cell and spatial omics analysis, integrating scVI, scANVI, totalVI and other VAE-based models for batch correction, cell annotation, multi-omics integration, and RNA velocity (scverse/NumFOCUS, Nature Methods 2018/2024)
- [CellRank](https://github.com/scverse/cellrank) - Probabilistic framework for inferring cell fate decisions and trajectory dynamics from multi-view single-cell data using Markov chains and machine learning, integrating RNA velocity, pseudotime, and metabolic labeling to predict differentiation paths and terminal states (scverse/Theis Lab, 449+ stars, BSD 3-Clause)
- [cellxgene (Chan Zuckerberg Initiative)](https://github.com/chanzuckerberg/cellxgene) - Interactive explorer for single-cell transcriptomics data enabling visualization of UMAP/t-SNE embeddings, differential expression analysis, and cross-dataset comparison through a fast web-based interface; widely adopted for exploring atlas-scale single-cell datasets and integrating with AI/ML analysis workflows (773+ stars, MIT License)
- [Helical](https://github.com/helicalAI/helical) - Unified framework for state-of-the-art pre-trained bio foundation models across genomics and transcriptomics, providing standardized interfaces and pipelines for DNA, RNA, and single-cell models including Evo 2, Geneformer, scGPT, and UCE with streamlined inference, benchmarking, and fine-tuning workflows (213+ stars, 2024-2025)
- [GEARS](https://github.com/snap-stanford/GEARS) - Geometric deep learning model predicting transcriptional outcomes of novel single- and multi-gene perturbations using genegene knowledge graphs, 40% higher precision than prior methods on combinatorial perturbation prediction (Stanford, Nature Biotechnology 2024)
- [scDFM (ICLR 2026)](https://github.com/AI4Science-WestlakeU/scDFM) - Distributional flow matching model for robust single-cell perturbation prediction, modeling the full distribution of perturbed cellular expression profiles conditioned on control states via PAD-Transformer and multi-kernel MMD regularization; reduces MSE by 19.6% over the strongest baseline in combinatorial settings (Westlake University, 41+ stars, MIT License)
- [scTranslator (Nature Biomedical Engineering 2025)](https://github.com/TencentAILabHealthcare/scTranslator) - Pre-trained large generative model translating single-cell transcriptomes to proteomes in an alignment-free manner, generating absent protein abundance data for CITE-seq, spatial CITE-seq, REAP-seq, and NEAT-seq across tissues and diseases; offers three model variants pretrained on 2M human cells, 160K PBMCs, or 18K bulk samples (Tencent AI Lab Healthcare, 96+ stars)
- [scGPT](https://github.com/bowang-lab/scGPT) - Single-cell analysis with transformers
- [CellWhisperer (Nature Biotechnology 2025)](https://github.com/epigen/cellwhisperer) - Multimodal AI bridging transcriptomics data and natural language, enabling intuitive chat-based exploration and analysis of single-cell RNA-seq datasets through conversational interaction without coding; fine-tuned Mistral 7B LLaVA model emulating biologist-bioinformatician discussions (207+ stars, GPL-3.0)
- [CellTypist](https://github.com/Teichlab/celltypist) - Automated cell type annotation tool for single-cell transcriptomics using gradient boosting and logistic regression with reference atlases, enabling standardized classification across datasets (Wellcome Sanger Institute, Nature Biotechnology 2022)
- [mLLMCelltype](https://github.com/cafferychen777/mLLMCelltype) - Multi-LLM consensus framework for automated cell type annotation in single-cell transcriptomics, integrating predictions from 10+ large language models with iterative discussion and uncertainty quantification to reduce single-model biases, achieving up to 95% accuracy without reference datasets; available as CRAN R package and PyPI Python package with Scanpy/Seurat integration (2025)
- [Cell2Sentence](https://github.com/vandijklab/cell2sentence) - Teaching Large Language Models the Language of Biology through single-cell transcriptomics (ICML 2024)
- [OmicVerse](https://github.com/Starlitnightly/omicverse) - Unified Python framework for bulk, single-cell, and spatial RNA-seq multi-omics analysis with deep learning deconvolution (VAE) and graph neural networks, bridging Bindea, Bindea, scanpy and squidpy ecosystems (Nature Communications 2024)
- [ChatSpatial](https://github.com/cafferychen777/ChatSpatial) - MCP server enabling spatial transcriptomics analysis via natural language, integrating 60+ methods including SpaGCN, Cell2location, LIANA+, CellRank for Visium, Xenium, MERFISH platforms
- [Enformer](https://github.com/google-deepmind/deepmind-research/tree/master/enformer) - Gene expression prediction
- [DNABERT](https://github.com/jerryji1993/DNABERT) - DNA sequence analysis
- [DNABERT-2 (ICLR 2024)](https://github.com/MAGICS-LAB/DNABERT_2) - Efficient foundation model and benchmark for multi-species genome understanding with context-aware nucleotide representations, improving upon DNABERT for diverse genomic task transfer learning (UIUC MAGICS Lab, 484+ stars)
- [gReLU (Genentech, 2024)](https://github.com/Genentech/gReLU) - Python library to train, interpret, and apply deep learning models to DNA sequences, providing a unified framework for regulatory genomics with support for CNN and transformer architectures, variant effect prediction, and attribution analysis (325+ stars)
- [BioReason (NeurIPS 2025)](https://github.com/bowang-lab/BioReason) - First architecture deeply integrating a DNA foundation model with an LLM for multimodal biological reasoning, achieving 98% accuracy on KEGG disease pathway prediction and 15%+ average gains on variant effect prediction with interpretable step-by-step reasoning traces (bowang-lab, 390+ stars)
- [scBERT](https://github.com/TencentAILabHealthcare/scBERT) - Single-cell BERT for gene expression
- [GenePT](https://github.com/yiqunchen/GenePT) - Generative pre-training for genomics
- [DNA Claude Analysis](https://github.com/shmlkv/dna-claude-analysis) - Interactive personal genome analysis toolkit using Claude Code and Python. Parses raw genotyping data from consumer DNA services and analyzes SNPs across 17 categories including health risks, pharmacogenomics, ancestry, and nutrition, with a terminal-style HTML dashboard.
- [OpenCRISPR](https://github.com/profluent-ai/opencrispr) - First open-source AI-generated gene editing systems developed with protein language models, enabling programmable CRISPR-Cas nucleases for synthetic biology and therapeutic genome editing (Profluent, 2024)
- [AlphaMissense](https://github.com/google-deepmind/alphamissense) - Google DeepMind's AlphaFold-derived classifier for proteome-wide missense variant effect prediction, providing pathogenicity scores for all ~71M possible human missense variants and classifying 89% with 90% precision; pre-computed predictions are integrated into Ensembl VEP and UCSC Genome Browser to support clinical variant interpretation (Science 2023)
- [AlphaGenome](https://github.com/google-deepmind/alphagenome) - Google DeepMind's unified DNA sequence foundation model predicting molecular consequences of genetic variants from single-base resolution up to 1 megabase context, jointly outputting thousands of regulatory tracks (RNA expression, splicing, chromatin accessibility, TF binding, contact maps) for human and mouse genomes via a Python client and non-commercial API (2025)
- [GPN-Star (Song Lab, UC Berkeley, bioRxiv 2025)](https://github.com/songlab-cal/gpn) - Phylogeny-aware genomic language model trained on whole-genome alignments across multiple evolutionary timescales, predicting functional constraints and variant effects for human, mouse, chicken, fly, worm, and Arabidopsis genomes (344+ stars, MIT License)
- [GENERanno (bioRxiv 2025)](https://github.com/GenerTeam/GENERanno) - Genomic foundation model for metagenomic and genome annotation, featuring an 8k base-pair context and 500M parameters trained on 386B base pairs of eukaryotic DNA; provides expert models and a unified CLI for prokaryotic/eukaryotic coding-sequence annotation with strong performance on Genomic Benchmarks, Nucleotide Transformer tasks, and custom Gener tasks (GenerTeam, 314+ stars, MIT License)
- [GENERator (bioRxiv 2026)](https://github.com/GenerTeam/GENERator) - Long-context generative genomic foundation model using 6-mer tokenization for DNA sequence modeling and generation, with v2 model families for prokaryote and eukaryote genomes and pretrained weights available on HuggingFace (GenerTeam, 460+ stars, MIT License, 2025-2026)
- [DeepVariant](https://github.com/google/deepvariant) - Google DeepMind's deep learning analysis pipeline for calling genetic variants (SNPs and indels) from next-generation DNA sequencing data, achieving human expert-level accuracy and widely adopted in clinical genomics, population genetics, and precision medicine; pre-trained models available for multiple sequencing platforms and organismal genomes (Nature Biotechnology 2018, 3.7K+ stars)
- [Casanovo](https://github.com/Noble-Lab/casanovo) - Transformer encoder-decoder for de novo peptide sequencing from tandem mass spectrometry, translating MS/MS spectra directly to peptide sequences without reference databases, enabling identification of novel peptides for immunopeptidomics, antibody repertoires, and metaproteomes (Noble Lab UW, Nature Communications 2024)
- [Dorado](https://github.com/nanoporetech/dorado) - Oxford Nanopore's official deep-learning basecaller for nanopore sequencing, converting raw electrical signals into DNA/RNA sequences with integrated modified-base (methylation) detection and efficient CPU/GPU inference; foundational tool for long-read genomics, epigenetics, and real-time sequencing analysis (nanoporetech, 846+ stars, actively maintained)
#### Neuroscience & Behavioral Analysis
- [DeepLabCut](https://github.com/DeepLabCut/DeepLabCut) - Markerless pose estimation of user-defined features with deep learning for all animals including humans, enabling quantitative behavioral analysis in neuroscience and ethology (Nature Neuroscience 2018, 5.6K+ stars)
- [SLEAP](https://github.com/talmolab/sleap) - Deep learning-based multi-animal pose tracking and behavior classification, enabling automated quantification of social interactions and collective behavior across species (Nature Methods 2022, 2.2K+ stars)
- [NeuroAI (Meta FAIR)](https://github.com/facebookresearch/neuroai) - Modular Python suite for Neuro-AI research across all modalities, providing efficient data loaders (NeuralSet), curated datasets (NeuralFetch), scalable training (NeuralTrain), and unified benchmarking (NeuralBench) for building and evaluating neuroscience foundation models (Meta FAIR, 270+ stars, MIT License, 2026)
- [CEBRA (Nature 2023)](https://github.com/AdaptiveMotorControlLab/CEBRA) - Learnable latent embeddings for joint behavioral and neural analysis, enabling consistent and interpretable mapping of neural activity to behavior across modalities, species, and experiments (EPFL & Harvard, 1K+ stars)
- [Kilosort (Nature Methods 2024)](https://github.com/MouseLand/Kilosort) - Fast spike sorting with drift correction for extracellular electrophysiology, enabling universal neural spike sorting via deep learning on high-density neural probe recordings (MouseLand, 609+ stars)
- [SpikeInterface](https://github.com/SpikeInterface/spikeinterface) - Unified Python framework for extracellular electrophysiology, standardizing interfaces to 10+ ML-based spike sorting algorithms including Kilosort for reproducible neural spike sorting workflows (792+ stars, actively maintained)
- [CaImAn (Flatiron Institute)](https://github.com/flatironinstitute/CaImAn) - Computational toolbox for large scale Calcium Imaging Analysis, including movie handling, motion correction, source extraction, spike deconvolution and result visualization, using machine learning for automated neuron detection and activity inference in two-photon and one-photon calcium imaging data (723+ stars, actively maintained)
- [TRIBE v2](https://github.com/facebookresearch/tribev2) - Meta FAIR's foundation model of vision, audition, and language for in-silico neuroscience, predicting fMRI brain responses to naturalistic multimodal stimuli (video, audio, text) through unified Transformer architecture mapped to the cortical surface (2026)
- [braindecode](https://github.com/braindecode/braindecode) - Deep learning software to decode EEG, ECG or MEG signals, providing standardized neural network models, preprocessing pipelines, and evaluation workflows for brain-computer interfaces and cognitive neuroscience research (1.2K+ stars, BSD 3-Clause, actively maintained)
- [snntorch](https://github.com/jeshraghian/snntorch) - Deep learning with spiking neural networks in Python, providing gradient-based training of SNNs via PyTorch autodifferentiation for brain-inspired computing and neuromorphic research, with online learning capabilities and extensive tutorials (1.9K+ stars, actively maintained)
- [nilearn](https://github.com/nilearn/nilearn) - Machine learning and statistical learning for neuroimaging in Python, providing easy-to-use tools for fMRI and MRI analysis including decoding, connectivity estimation, and parcellation with seamless scikit-learn integration (INRIA Parietal team, 1.4K+ stars)
- [BrainIAC (Nature Neuroscience 2026)](https://github.com/AIM-KannLab/BrainIAC) - Self-supervised vision foundation model for generalized structural brain MRI analysis, pretrained on ~49,000 scans from diverse datasets and generalizing across brain age prediction, dementia/MCI classification, IDH mutation detection, glioma survival prediction, time-to-stroke estimation, MR sequence classification, and brain tumor segmentation; outperforms task-specific models especially with limited training data (Mass General Brigham & Harvard Medical School, 129+ stars)
#### Computational Pathology & Digital Pathology
- [UNI (Nature Medicine 2024)](https://github.com/mahmoodlab/UNI) - General-purpose pathology foundation model pretrained on 100K+ diagnostic whole-slide images across 20 major tissue types, achieving state-of-the-art transfer learning across 30+ clinical tasks and serving as a universal feature extractor for digital pathology (Mahmood Lab, 722+ stars)
- [Prov-GigaPath (Nature 2024)](https://github.com/prov-gigapath/prov-gigapath) - Whole-slide pathology foundation model trained on 1.3 billion image tiles from 171K slides using a LongNet-based architecture to encode gigapixel-scale WSIs for cancer subtyping and biomarker prediction (Microsoft Research & Providence, 601+ stars)
- [GigaTIME (Cell 2025)](https://github.com/prov-gigatime/GigaTIME) - Multimodal AI system generating virtual populations for tumor microenvironment modeling from H&E and multiplex immunofluorescence pathology images, enabling large-scale spatial analysis of cancer biology and therapeutic response prediction (Microsoft Research & Providence, 370+ stars)
- [CONCH (Nature Medicine 2024)](https://github.com/mahmoodlab/CONCH) - Vision-language pathology foundation model using contrastive learning on histopathology image-text pairs, enabling zero-shot classification, slide-level retrieval, and multimodal reasoning across diverse cancer types (Mahmood Lab, 494+ stars)
- [PLIP (Nature Medicine 2023)](https://github.com/PathologyFoundation/plip) - First vision-and-language foundation model for pathology AI, fine-tuned from CLIP on 249K image-caption pairs, enabling open-ended visual-semantic search and zero-shot diagnosis across histopathology (Pathology Foundation, 376+ stars)
- [TITAN (Nature Medicine 2024)](https://github.com/mahmoodlab/TITAN) - Multimodal whole-slide pathology foundation model jointly pretrained on H&E histology and diagnostic text reports, enabling zero-shot cancer subtyping, biomarker prediction, and multimodal reasoning across diverse cancer types (Mahmood Lab, 341+ stars)
- [Virchow (Nature Medicine 2024)](https://huggingface.co/paige-ai/Virchow) - Self-supervised pathology foundation model (ViT-Huge, 632M parameters) pretrained via DINOv2 on 1.5M whole-slide images from Memorial Sloan Kettering across 17 cancer types, with Virchow2 follow-up scaling to 3.1M slides and mixed magnifications, achieving SOTA on biomarker prediction, mutation classification, and rare cancer detection (Paige AI & MSK)
- [H-Optimus (Bioptimus, Nature Medicine 2025)](https://huggingface.co/bioptimus/H-optimus-0) - Open-weights pathology foundation model family (H-Optimus-0: 1.1B-parameter ViT pretrained via DINOv2 on 500M+ diagnostic image tiles; H-Optimus-1 follow-up) for whole-slide image analysis, achieving strong zero-shot and fine-tuned transfer across biomarker prediction, cancer subtyping, and mutation classification (Bioptimus, Apache 2.0)
- [TRIDENT (2025)](https://github.com/mahmoodlab/TRIDENT) - Toolkit for large-scale whole-slide image processing supporting 22+ patch encoders (UNI, CONCH, Virchow, H-Optimus-0, etc.), slide encoders (TITAN, GigaPath, PRISM, CHIEF, Madeleine, Feather), tissue segmentation, and multi-GPU inference with end-to-end pipeline and smart resume for standardized deployment of computational pathology foundation models (Mahmood Lab, Harvard Medical School, 553+ stars)
- [Feather (Mahmood Lab, ICML 2025 Spotlight)](https://github.com/mahmoodlab/MIL-Lab) - Lightweight supervised slide foundation model with 0.9M parameters pretrained on 24K whole-slide images for pan-cancer morphological classification, achieving competitive performance with much larger self-supervised models (TITAN, GigaPath) while enabling finetuning on consumer-grade GPUs; includes standardized MIL implementations and benchmarking across 15+ classification tasks (Mahmood Lab, Harvard Medical School, 153+ stars)
- [PathChat (Nature Medicine 2024)](https://github.com/MahmoodLab/PathChat) - Multimodal generative AI assistant for computational pathology enabling interactive visual-language conversations over histopathology images for diagnostic reasoning, case discussion, and education, built on a Mistral-7B backbone with domain-specific fine-tuning (Mahmood Lab, Harvard Medical School, 1.2K+ stars)
- [SlideChat (CVPR 2025)](https://github.com/uni-medical/SlideChat) - First large vision-language assistant for gigapixel whole-slide pathology image understanding, released with the SlideInstruction dataset and SlideBench benchmark (uni-medical, Apache 2.0, 2025)
- [HEST (NeurIPS 2024)](https://github.com/mahmoodlab/HEST) - Dataset and benchmarking framework integrating histology and spatial transcriptomics, enabling multimodal analysis of whole-slide images with matched spatial gene expression for advancing computational pathology and tissue microenvironment research (Mahmood Lab, Harvard Medical School, 411+ stars)
#### Medical AI & Clinical Applications
- [Cellpose](https://github.com/MouseLand/cellpose) - Generalist deep learning algorithm for cell and nucleus segmentation across diverse image types, with human-in-the-loop training (2.0) and one-click image restoration (3.0), 70K+ training objects (Nature Methods 2021/2022/2025)
- [StarDist](https://github.com/stardist/stardist) - Deep learning-based object detection and segmentation for star-convex shapes, widely adopted for cell and nucleus segmentation in fluorescence and electron microscopy via a compact neural network architecture with non-maximum suppression and shape-based post-processing (Nature Methods 2020, 1.2K+ stars)
- [InstanSeg (Nature Methods 2025)](https://github.com/instanseg/instanseg) - PyTorch-based embedding instance segmentation algorithm optimized for accurate, efficient, and portable cell and nucleus segmentation across fluorescence and brightfield microscopy images, achieving state-of-the-art speed and accuracy with lightweight model sizes suitable for edge deployment (224+ stars, Apache 2.0)
- [napari](https://github.com/napari/napari) - Fast, interactive, multi-dimensional image viewer for Python, foundational platform for scientific imaging AI with a rich plugin ecosystem integrating deep learning segmentation, object tracking, and microscopy analysis workflows (2.6K+ stars)
- [cellSAM](https://github.com/vanvalenlab/cellSAM) - Foundation model for universal cell segmentation achieving state-of-the-art performance across bacteria, tissue, yeast, cell culture, and diverse imaging modalities (brightfield, fluorescence, phase), with pip-installable inference and Napari plugin (vanvalenlab/Caltech, bioRxiv 2024)
- [micro-sam](https://github.com/computational-cell-analytics/micro-sam) - Segment Anything Model for microscopy: interactive and automatic segmentation of light, electron, and fluorescence microscopy images in 2D and 3D, with domain-specific fine-tuning workflows for scientific imaging (1.5K+ stars)
- [MedSAM](https://github.com/bowang-lab/MedSAM) - Universal medical image segmentation foundation model trained on 1.57M image-mask pairs across 10 imaging modalities and 30+ cancer types (Nature Communications 2024)
- [MedSAM2](https://github.com/bowang-lab/MedSAM2) - Segment Anything in 3D medical images and videos, extending SAM2 to volumetric and temporal medical imaging with state-of-the-art zero-shot segmentation performance across CT, MRI, and surgical video (arXiv 2025)
- [MedSegX](https://github.com/MedSegX/MedSegX-code) - Generalist foundation model and database for open-world medical image segmentation, enabling universal segmentation of diverse anatomical structures and pathologies with zero-shot generalization to unseen tasks and modalities (Nature Biomedical Engineering 2025)
- [VoxTell (MIC-DKFZ, 2025)](https://github.com/MIC-DKFZ/VoxTell) - Free-text promptable universal 3D medical image segmentation foundation model enabling zero-shot segmentation of diverse anatomical structures and pathologies via natural language prompts across CT, MRI, and other volumetric imaging modalities (DKFZ, 195+ stars, Apache 2.0)
- [BiomedParse](https://github.com/microsoft/BiomedParse) - Foundation model for joint segmentation, detection, and recognition of biomedical objects across nine imaging modalities, with v2 introducing BoltzFormer architecture for end-to-end 3D inference (Microsoft, Nature Methods 2025)
- [UniBiomed (Nature Communications 2026)](https://github.com/Luffy03/UniBiomed) - Universal foundation model for grounded biomedical image interpretation, enabling comprehensive visual understanding, reasoning, and grounding across diverse biomedical imaging modalities with strong zero-shot generalization (55+ stars, Apache 2.0, 2025-2026)
- [MIRA (NeurIPS 2025)](https://github.com/microsoft/MIRA) - Medical time series foundation model pretrained on 454B time points from heterogeneous clinical corpora spanning ICU physiological signals and hospital EHR, with continuous-time rotary positional encoding, frequency-specialized Mixture-of-Experts, and neural ODE extrapolation for zero-shot forecasting across irregular and multimodal temporal health data (Microsoft, 399+ stars, MIT License)
- [HealthGPT (ICML 2025 Spotlight)](https://github.com/ZJU4HealthCare/HealthGPT) - Medical large vision-language model unifying comprehension and generation via heterogeneous knowledge adaptation, enabling holistic medical image understanding, visual question answering, and clinical report generation across diverse modalities (ZJU4HealthCare, 1.6K+ stars)
- [Merlin (Stanford MIMI, Nature 2026)](https://github.com/StanfordMIMI/Merlin) - 3D vision-language model for computed tomography that leverages both structured electronic health records (EHR) and unstructured radiology reports for pretraining, enabling multimodal medical understanding and radiology report generation (447+ stars, MIT License, 2026)
- [MedAgents](https://github.com/gersteinlab/MedAgents) - Multi-disciplinary collaboration framework for zero-shot medical reasoning using role-playing LLM agents (ACL 2024)
- [MedAgentGym](https://github.com/wshi83/MedAgentGym) - Scalable agentic training environment for code-centric reasoning in biomedical data science
- [MedRAX (ICML 2025)](https://github.com/bowang-lab/MedRAX) - First versatile medical reasoning agent for chest X-ray interpretation, dynamically integrating state-of-the-art CXR analysis tools and multimodal LLMs into a unified framework; introduces ChestAgentBench with 2,500 complex medical queries across 7 categories (bowang-lab, 1.1K+ stars)
- [MedRAG](https://github.com/Teddy-XiongGZ/MedRAG) - Systematic medical RAG toolkit for question answering over PubMed, StatPearls, textbooks, and Wikipedia, supporting multiple retrievers, domain LLMs, and follow-up-query workflows for benchmarked clinical/biomedical QA (ACL Findings 2024)
- [OpenMed (2025-2026)](https://github.com/maziyarpanahi/openmed) - Local-first, open-source healthcare AI toolkit for clinical NLP and PHI/PII de-identification across 12 languages, running entirely on-device with 1,000+ specialized medical models; provides Python SDK, REST API, Docker deployment, and native Swift apps via OpenMedKit with Apple MLX/CoreML acceleration, supporting HIPAA-aware de-identification with 247 PII checkpoints (3K+ stars, Apache 2.0, arXiv 2508.01630)
- [NVIDIA Biomedical AI-Q Research Agent](https://github.com/NVIDIA-AI-Blueprints/biomedical-aiq-research-agent) - Deployable biomedical deep-research agent blueprint combining on-prem multimodal RAG, report generation, human-in-the-loop editing, and virtual screening with MolMIM and DiffDock for drug discovery workflows (2025)
- [nnU-Net](https://github.com/MIC-DKFZ/nnUNet) - Self-configuring deep learning framework for semantic segmentation of biomedical images requiring no manual hyperparameter tuning; automatically adapts preprocessing, network topology, and training parameters to achieve state-of-the-art results across 120+ international competitions and benchmarks out-of-the-box (DKFZ, Nature Methods 2021, 8.3k+ stars)
- [TotalSegmentator](https://github.com/wasserth/TotalSegmentator) - Robust deep learning-based segmentation of >100 anatomical structures in CT and MR images, built on nnU-Net and widely adopted in clinical radiology and surgical planning workflows (2.6K+ stars)
- [MONAI](https://github.com/Project-MONAI/MONAI) - NVIDIA and King's College London's open-source AI toolkit for healthcare imaging, providing foundational frameworks for medical image annotation (MONAI Label), training (MONAI Core), and deployment (MONAI Deploy) across radiology, pathology, and endoscopy (8K+ stars, Apache 2.0)
- [ZeroCostDL4Mic](https://github.com/HenriquesLab/ZeroCostDL4Mic) - Google Colab-based no-code toolbox democratizing deep learning in microscopy for biologists without programming experience, enabling AI-powered image segmentation, denoising, super-resolution, and object tracking across diverse imaging modalities (Henriques Lab, 640+ stars)
- [BiaPy](https://github.com/BiaPyX/BiaPy) - Open-source deep learning toolbox for bioimage analysis providing a unified, configuration-driven framework for 2D/3D semantic segmentation, instance segmentation, classification, denoising, super-resolution, and self-supervised learning; integrates state-of-the-art architectures including U-Net, Vision Transformers, and ConvNeXt, designed for microscopy and biomedical imaging researchers without extensive coding expertise (MIT License, actively maintained)
- [QuPath](https://github.com/qupath/qupath) - Open-source bioimage analysis platform for digital pathology and research, featuring AI-powered cell detection, tissue classification, and whole-slide image analysis with extensible scripting and plugin architecture (1.3K+ stars, actively maintained)
- [BioImage.IO](https://github.com/bioimage-io/core-bioimage-io-python) - Community-driven model zoo and deployment infrastructure for AI-powered bioimage analysis, enabling standardized sharing, validation, and cross-platform execution of deep learning models across Fiji, Ilastik, napari, and other scientific imaging tools (EPFL, EMBL, and global collaborators, actively maintained)
### ⚛ Chemistry & Materials
#### LLM for Chemistry
- [LLM4Chemistry](https://github.com/OpenDFM/LLM4Chemistry) - Curated paper list about LLMs for chemistry covering fine-tuning, reasoning, multi-modal models, agents, and benchmarks (COLING 2025)
- [ChemMCP](https://github.com/OSU-NLP-Group/ChemMCP) - Extensible chemistry toolkit for MCP-enabled AI assistants, exposing molecule analysis, property prediction, and reaction synthesis tools through unified Python/MCP interfaces for chemistry agents and research workflows (Apache 2.0, 2025)
- [MoleCode](https://github.com/AtomFlow-AI/MoleCode) - LLM-native molecular language that represents molecules as explicit graph-based code, enabling LLMs to operate and reason on chemistry directly with 5× lower token cost and ~76-80% accuracy on novel molecules vs ~20% for SMILES; supports small molecules, polymers, and Markush structures with lossless RDKit interconversion and Claude Code/Codex agent skills (AtomFlow, arXiv:2605.16480, 281+ stars, MIT License, 2026)
#### Materials Discovery
- [GNoME](https://github.com/google-deepmind/materials_discovery) - DeepMind's graph neural network for materials exploration, discovering 2.2M new crystal structures (380K most stable) equivalent to 800 years of traditional research, with 520K+ materials dataset open-sourced (Nature 2023)
- [FAIRChem (OMat24)](https://github.com/FAIR-Chem/fairchem) - Meta's comprehensive ML ecosystem for materials/chemistry with 118M+ DFT calculations, EquiformerV2 models achieving top Matbench Discovery performance
- [All-atom Diffusion Transformers (ADiT)](https://github.com/facebookresearch/all-atom-diffusion-transformer) - Unified latent diffusion transformer that jointly generates periodic crystals and non-periodic molecules, scaling to 500M parameters with SOTA results on QM9, MP20, and GEOM-DRUGS (Meta FAIR, ICML 2025, 310+ stars)
- [JARVIS](https://github.com/usnistgov/jarvis) - NIST's open-source platform for data-driven atomistic materials design, integrating DFT datasets (JARVIS-DFT), machine learning property prediction (JARVIS-ML), and a comprehensive leaderboard for benchmarking materials AI methods across the periodic table (384+ stars)
- [NVIDIA ALCHEMI Toolkit](https://github.com/NVIDIA/nvalchemi-toolkit) - Developer toolkit for accelerating training and inference for AI in chemistry and material science, providing optimized GPU-accelerated workflows for molecular and materials machine learning (NVIDIA, 2026)
- [NequIP](https://github.com/mir-group/nequip) - E(3)-equivariant neural network interatomic potentials achieving DFT accuracy with up to 1000× less training data than invariant models, foundational architecture behind MACE and Allegro (Harvard, MIT, Nature Communications 2022)
- [Allegro](https://github.com/mir-group/allegro) - Highly scalable equivariant deep learning interatomic potentials enabling million-atom molecular dynamics simulations with ab initio accuracy, building on E(3)-equivariant architectures for large-scale atomistic modeling (mir-group, MIT License, 480+ stars)
- [SchNetPack](https://github.com/atomistic-machine-learning/schnetpack) - PyTorch toolkit for deep neural networks in atomistic simulations, implementing SchNet, DimeNet++, PaiNN, and GemNet for molecular dynamics and quantum chemistry (900+ stars)
- [pymatgen](https://github.com/materialsproject/pymatgen) - Python Materials Genomics: robust materials analysis library defining classes for structures and molecules with support for many electronic structure codes; foundational toolkit powering the Materials Project (Berkeley Lab, 1.8K+ stars)
- [MACE](https://github.com/ACEsuit/mace) - Machine learning interatomic potentials
- [CHGNet](https://github.com/CederGroupHub/chgnet) - Universal pretrained neural network potential with charge and magnetic moment awareness, trained on 1.5M+ Materials Project inorganic structures for charge-informed molecular dynamics and phase diagram prediction (Berkeley, Nature Machine Intelligence 2023 Cover)
- [MatterGen](https://github.com/microsoft/mattergen) - Diffusion-based generative model for inorganic materials design, steering generation by chemistry, symmetry, bulk modulus, band gap, or magnetic properties, 2× more likely to produce stable novel structures than prior methods, experimentally validated with synthesized TaCr₂O₆ (Microsoft, Nature 2025)
- [MatterSim](https://github.com/microsoft/mattersim) - Deep learning atomistic model across elements, temperatures, and pressures
- [ORB](https://github.com/orbital-materials/orb-models) - Universal machine learning interatomic potential for atomistic simulation of materials, molecules, and biomolecules across the periodic table, with open-source pretrained models and inference tools (Orbital Materials, 2024-2025)
- [SevenNet (JCTC 2024)](https://github.com/MDIL-SNU/SevenNet) - Graph neural network interatomic potential package supporting efficient multi-GPU parallel molecular dynamics simulations, enabling large-scale atomistic modeling with machine learning potentials (MDIL-SNU, MIT License)
- [LLaMat (Nature Machine Intelligence 2026)](https://github.com/M3RG-IITD/llamat) - Family of large language models for materials research via continued pretraining of LLaMA-2/3 on ~30B materials science tokens, outperforming commercial LLMs on materials science tasks while identifying "adaptation rigidity" in overtrained models; includes MatNLP benchmark and CIF crystal generation capabilities (IIT Delhi M3RG, MIT License)
- [Crystal Graph CNNs](https://github.com/txie-93/cgcnn) - Crystal property prediction
- [MatBench](https://github.com/materialsproject/matbench) - Materials informatics benchmark
- [Best of Atomistic Machine Learning](https://github.com/JuDFTteam/best-of-atomistic-machine-learning) - Curated list of atomistic ML projects for materials science
#### Chemical Synthesis
- [AiZynthFinder](https://github.com/MolecularAI/aizynthfinder) - AstraZeneca's industrial-grade retrosynthetic planning tool using MCTS to recursively decompose molecules into purchasable precursors, with multi-step route scoring and support for custom one-step models (v4.0, 2024)
- [Molecular Transformers](https://github.com/pschwllr/MolecularTransformer) - AI for chemical reaction prediction and synthesis planning
- [SyntheMol (Stanford, Nature Machine Intelligence 2024)](https://github.com/swansonk14/SyntheMol) - Generative AI system for antibiotic discovery that searches billions of synthesizable molecules by combining molecular building blocks through real chemical reactions, experimentally validating novel compounds active against drug-resistant bacteria
#### Lab Automation & Robotics
- [PyLabRobot](https://github.com/PyLabRobot/pylabrobot) - Interactive and hardware-agnostic SDK for laboratory automation, enabling programmatic control of liquid handlers, plate readers, and other lab instruments across multiple vendors; foundational infrastructure for self-driving laboratories and AI-driven experimental execution (447+ stars)
### 🌌 Physics & Astronomy
#### Machine Learning for Physics
- [AlphaQubit](https://github.com/google-deepmind/alphaqubit) - Google DeepMind and Google Quantum AI's transformer-based neural-network decoder for quantum error correction, trained on real Sycamore quantum processor data to outperform tensor-network and correlated matching decoders at code distances 3 and 5, demonstrating ML's role in enabling fault-tolerant quantum computing (Nature 2024)
- [FermiNet](https://github.com/google-deepmind/ferminet) - DeepMind's neural network for ab-initio quantum chemistry, directly solving the many-electron Schrödinger equation via variational Monte Carlo with antisymmetric wavefunctions, extended to excited states (Phys. Rev. Research 2020, Science 2024)
- [NetKet](https://github.com/netket/netket) - Machine learning toolkit for many-body quantum systems, implementing neural quantum states, variational Monte Carlo, and tensor network algorithms to solve ground-state and dynamical problems in condensed matter physics and quantum chemistry (EPFL & collaborators, Nature Physics 2019/2022+, 670+ stars)
- [JAX-MD](https://github.com/jax-md/jax-md) - Molecular dynamics in JAX
- [Neural ODEs](https://github.com/rtqichen/torchdiffeq) - Differential equations with neural networks
- [Physics-Informed Neural Networks](https://github.com/maziarraissi/PINNs) - Physics-constrained ML
- [EquiformerV2](https://github.com/atomicarchitects/equiformer_v2) - Improved equivariant Transformer for 3D atomic graphs (ICLR2024)
- [Equiformer](https://github.com/atomicarchitects/equiformer) - Equivariant graph attention Transformer (ICLR2023)
- [TORAX](https://github.com/google-deepmind/torax) - Differentiable tokamak core transport simulator for fusion energy research, coupling PDE solvers with JAX auto-differentiation and neural-network surrogates for fast forward modelling, pulse-design, and trajectory optimization (Google DeepMind, Apache 2.0)
- [DiffPhysDrone (Nature Machine Intelligence 2025)](https://github.com/HenryHuYu/DiffPhysDrone) - First real quadrotor robot trained end-to-end with differentiable physics for vision-based agile flight, bridging simulation-based learning and real-world deployment with physics-informed neural network controllers (558+ stars)
- [Walrus (arXiv 2025)](https://github.com/PolymathicAI/walrus) - Cross-domain foundation model for continuum dynamics trained on 19 physical scenarios spanning 63 variables, featuring adaptive compute via stride modulation and patch jittering for long-run stability (Polymathic AI, 293+ stars, MIT License)
#### Astronomy & Astrophysics
- [AstroCLIP](https://github.com/PolymathicAI/AstroCLIP) - Cross-modal self-supervised foundation model for galaxies by Polymathic AI, jointly embedding multi-band galaxy imaging and optical spectra into a shared latent space to enable zero/few-shot redshift estimation, galaxy property prediction, morphology classification, and cross-modal similarity search (MNRAS Letters 2024)
- [AION (arXiv 2025)](https://github.com/PolymathicAI/AION) - Polymathic AI's large omnimodal foundation model for astronomical surveys, seamlessly integrating 39 distinct data modalities including imaging, spectra, photometry, and catalog entries for similarity search, property prediction, and generative modeling across legacy surveys (MIT)
- [AstroPy](https://github.com/astropy/astropy) - Python astronomy tools
- [Gaia Archive](https://gea.esac.esa.int/archive/) - Stellar data for ML
- [DeepSphere](https://github.com/deepsphere/deepsphere-pytorch) - Spherical CNNs for astronomy
### 🌍 Earth & Climate Science
#### Climate Modeling
- [GenCast](https://github.com/google-deepmind/graphcast) - Google DeepMind's diffusion-based ensemble weather forecasting model at 0.25° resolution, outperforming ECMWF ENS on 97.2% of targets up to 15 days ahead, with open-source code and weights (Nature 2024)
- [Aurora](https://github.com/microsoft/aurora) - Microsoft's foundation model for the Earth system supporting weather, air pollution, and ocean wave forecasting at multiple resolutions, trained on 1M+ hours of diverse atmospheric data (Nature 2025)
- [Earth-Copilot](https://github.com/microsoft/Earth-Copilot) - Microsoft's AI-powered geospatial Earth science application for natural-language exploration, visualization, and analysis of 130+ satellite collections, with STAC integration, multi-agent backend, MCP server, and deployable React/FastAPI stack (MIT, 2025)
- [ClimaX](https://github.com/microsoft/ClimaX) - First foundation model for weather and climate by Microsoft, Vision Transformer-based architecture trained on heterogeneous datasets (ICML 2023)
- [NeuralGCM](https://github.com/neuralgcm/neuralgcm) - Google Research's hybrid ML/physics atmospheric model combining learned dynamics with physical constraints, outperforming traditional models on 2-15 day forecasts and 40-year climate simulation, developed with ECMWF (Nature 2024)
- [NVIDIA Earth-2](https://github.com/NVIDIA/earth2studio) - World's first fully open, accelerated weather AI software stack with Medium Range forecasting and Nowcasting models using generative AI (January 2026)
- [Pangu-Weather](https://github.com/198808xc/Pangu-Weather) - Huawei's 3D high-resolution global weather forecast model at 0.25° resolution, first AI method to comprehensively outperform traditional NWP across all variables and lead times, integrated into ECMWF operational forecasts (Nature 2023)
- [FuXi (Nature 2023)](https://github.com/tpys/FuXi) - Fudan University's cascade machine learning forecasting system for 15-day global weather prediction, employing a 3D Earth-specific transformer with hard-constraint techniques to achieve state-of-the-art accuracy against traditional NWP and AI baselines
- [FengWu](https://github.com/OpenEarthLab/FengWu) - Shanghai AI Lab's deep learning-based global weather forecasting model pushing skillful forecasts beyond 10 days lead, with open-source inference code and pretrained ONNX model weights (arXiv 2023)
- [ai-models (ECMWF)](https://github.com/ecmwf-lab/ai-models) - ECMWF's unified framework and command-line tool to run AI-based weather forecasting models (GraphCast, Aurora, Pangu, NeuralGCM, FourCastNet) with operational ECMWF data infrastructure, enabling standardized inference and benchmarking across state-of-the-art meteorological AI systems (ECMWF, 576+ stars)
- [Prithvi WxC](https://huggingface.co/ibm/prithvi-wxc) - IBM-NASA open-source 2.3B parameter weather and climate foundation model trained on 160 MERRA-2 variables, runs on desktop with fine-tuned variants for climate downscaling and gravity wave parameterization
- [ClimateBench](https://github.com/duncanwp/ClimateBench) - Climate data benchmark for ML models
- [WeatherBench](https://github.com/pangeo-data/WeatherBench) - Weather prediction benchmark
- [WeatherBench2](https://github.com/google-research/weatherbench2) - Next-generation benchmark for data-driven global weather models with standardized evaluation framework and curated datasets for ML forecasting (Google Research, 2024)
- [WeatherGFT](https://github.com/black-yt/WeatherGFT) - Physics-AI hybrid modeling for fine-grained weather forecasting (NeurIPS'24)
- [GeoAI](https://github.com/opengeos/geoai) - High-level open-source geospatial AI package for satellite/aerial imagery analysis, model training, inference, interactive visualization, and QGIS integration, bridging PyTorch/Transformers with remote sensing workflows (MIT, 2026)
- [segment-geospatial](https://github.com/opengeos/segment-geospatial) - Python package for segmenting geospatial data with the Segment Anything Model (SAM), enabling zero-shot object segmentation in satellite and aerial imagery for remote sensing and Earth observation (MIT, 4k+ stars)
- [Awesome Large Weather Models](https://github.com/jaychempan/Awesome-LWMs) - Curated list of large weather models for AI Earth science
- [TerraTorch](https://github.com/IBM/terratorch) - Python toolkit for fine-tuning geospatial foundation models
- [Earth-Agent](https://github.com/opendatalab/Earth-Agent) - LLM agent framework for Earth Observation with 104 specialized tools across 5 functional kits
- [AI for Earth](https://planetarycomputer.microsoft.com/) - Microsoft's environmental AI
#### Geophysics & Seismology
- [SeisBench](https://github.com/seisbench/seisbench) - A toolbox for machine learning in seismology, providing unified interfaces for deep learning seismic phase picking, earthquake detection, and waveform analysis across multiple benchmark datasets and pretrained models (397+ stars, actively maintained)
#### Remote Sensing & Geospatial AI
- [TorchGeo](https://github.com/microsoft/torchgeo) - PyTorch domain library for geospatial deep learning providing standardized datasets, samplers, transforms, and pre-trained models for remote sensing, land cover mapping, and environmental monitoring (Microsoft, 4K+ stars)
- [Prithvi-EO-2.0 (IBM & NASA, 2024)](https://github.com/NASA-IMPACT/Prithvi-EO-2.0) - Versatile multi-temporal geospatial foundation model for Earth observation, built on a ViT-based masked autoencoder with 3D spatiotemporal patch embeddings and geolocation/temporal metadata encoding; pretrained on 4.2M global time-series samples from NASA's Harmonized Landsat and Sentinel-2 archive at 30m resolution, with 300M/600M parameter variants and fine-tuning configs for flood detection, wildfire scar, landslide detection, crop segmentation, land cover, and biomass estimation (258+ stars, MIT License)
- [Clay Foundation Model](https://github.com/Clay-foundation/model) - Open-source self-supervised vision foundation model for Earth observation by Clay Foundation (non-profit), a Masked Autoencoder ViT pretrained on multimodal satellite imagery (Sentinel-1/2, Landsat 8-9, NAIP, MODIS, LINZ DEM) with location/time embeddings, supporting classification, segmentation, change detection, similarity search, and few-shot downstream geospatial tasks (Apache 2.0, v1.5 2024-2025)
- [Satlas](https://github.com/allenai/satlas) - Allen Institute for AI's global geospatial foundation model for satellite imagery analysis, enabling large-scale mapping of buildings, wind turbines, trees, and land cover from Sentinel-2 data with open-source weights and inference tools (2024)
- [SkySensePlusPlus](https://github.com/kang-wu/SkySensePlusPlus) - Semantic-enhanced multi-modal remote sensing foundation model for Earth observation (Nature Machine Intelligence 2025), enabling universal interpretation across diverse satellite imagery modalities with open-source weights and benchmarks
- [TESSERA (CVPR 2026)](https://github.com/ucam-eo/tessera) - University of Cambridge's foundation model for time-series satellite imagery, enabling efficient extraction of temporal patterns from Earth observation for land classification, canopy height prediction, and other remote sensing tasks
- [TerraMind (IBM & ESA, 2025)](https://github.com/IBM/terramind) - First any-to-any generative foundation model for Earth Observation, enabling unified multimodal understanding and generation across diverse satellite sensors and geospatial tasks through a single architecture (258+ stars)
- [Awesome Remote Sensing Foundation Models](https://github.com/Jack-bo1220/Awesome-Remote-Sensing-Foundation-Models) - Curated collection of papers, datasets, benchmarks, code, and pre-trained weights for Remote Sensing Foundation Models (RSFMs), tracking the rapidly evolving landscape of vision, vision-language, generative, and agent-based geospatial AI (1.9K+ stars, 2024-2026)
### 🌾 Agriculture & Ecology
#### Agricultural AI
- [PlantNet](https://plantnet.org/) - Plant identification using AI and citizen science
- [AgML](https://github.com/Project-AgML/AgML) - Agricultural machine learning platform
- [FarmVibes.AI](https://github.com/microsoft/farmvibes-ai) - Multi-modal geospatial ML platform for agriculture and sustainability, fusing satellite imagery (RGB, SAR, multispectral), drone imagery, weather data, and sensor data for crop identification, carbon footprint estimation, and microclimate prediction (Microsoft Research, MIT License)
- [PlantCV](https://github.com/danforthcenter/plantcv) - Open-source image analysis toolkit for high-throughput plant phenotyping, extracting morphological, color, and texture traits from RGB, hyperspectral, and thermal imagery with modular Python workflows for crop improvement, stress detection, and plant biology research (Donald Danforth Plant Science Center, 795+ stars, MPL-2.0)
#### Ecological Modeling
- [BioSimulators](https://github.com/biosimulators/Biosimulators) - Biological simulation tools
- [EcoNet](https://github.com/microsoft/EcoNet) - Ecological modeling and conservation AI
- [BioCLIP (CVPR 2024)](https://github.com/Imageomics/bioclip) - Vision foundation model for the tree of life, pretrained on diverse biological imagery across taxa for zero-shot species identification, trait extraction, and biodiversity research (Ohio State University Imageomics Institute)
- [BioCLIP 2 (NeurIPS 2025 Spotlight)](https://github.com/Imageomics/bioclip-2) - Biological vision foundation model trained on TreeOfLife-200M, yielding extraordinary accuracy on diverse biological visual tasks including habitat classification and trait prediction despite a narrow training objective (Ohio State University Imageomics Institute)
- [Microsoft Biodiversity](https://github.com/microsoft/Biodiversity) - Microsoft AI for Good Lab's open-source biodiversity research hub providing AI models, edge devices, and tools for wildlife monitoring and conservation, including MegaDetector (camera trap animal detection), SPARROW (species recognition), PytorchWildlife (conservation AI toolkit), and bioacoustics analysis pipelines (1K+ stars)
- [BirdNET-Analyzer](https://github.com/birdnet-team/BirdNET-Analyzer) - Deep learning-based bioacoustic monitoring framework for automated bird species identification from audio recordings, supporting 6,000+ species globally with real-time analysis, batch processing, and API deployment; foundational tool in biodiversity research, conservation biology, and ecological acoustic monitoring (Cornell Lab of Ornithology, 1.5K+ stars, MIT License)
### 🧠 Social Sciences
#### Social Science Research & Simulation
- [AgentSociety](https://github.com/tsinghua-fib-lab/AgentSociety) - Modern LLM-native agent simulation platform for social science research and experimental design, providing a flexible framework for creating and managing intelligent agents in simulated environments (Tsinghua FIB Lab, 984+ stars, 2025)
- [Awesome Agent Skills for Empirical Research](https://github.com/brycewang-stanford/Awesome-Agent-Skills-for-Empirical-Research) - Curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines, enabling reproducible social science research with AI agents (Stanford REAP & CoPaper.AI, 1.1K+ stars, 2026)
- [EDSL](https://github.com/expectedparrot/edsl) - Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs (460+ stars, 2024)
---
## 🤖 Foundation Models for Science
### General Science Models
- [Galactica](https://github.com/paperswithcode/galai) - Large language model for science
- [Intern-S1](https://github.com/InternLM/Intern-S1) - Open-source scientific multimodal foundation model built on a 235B MoE LLM and 6B vision encoder, continually pretrained on 5T tokens including 2.5T scientific-domain tokens, with strong results across chemistry, materials, life science, and earth science benchmarks (2025)
- [Llemma](https://github.com/EleutherAI/math-lm) - Open language model for mathematics (7B/34B) trained on Proof-Pile-2, outperforming Minerva at equal scale on MATH benchmark, with tool use and formal theorem proving in Lean without finetuning (EleutherAI, ICLR 2024)
- [TimesFM (Google Research)](https://github.com/google-research/timesfm) - Pretrained time series foundation model for long-horizon forecasting across diverse scientific domains including climate variables, biomedical signals, and physical observations; decoder-only Transformer architecture with strong zero-shot generalization (19.8K+ stars, Apache 2.0, 2024-2025)
- [Chronos (Amazon Science, NeurIPS 2024)](https://github.com/amazon-science/chronos-forecasting) - Pretrained time series foundation model for zero-shot forecasting across diverse scientific and real-world domains; tokenizes continuous time series into discrete bins to train transformer language models on large-scale corpora, achieving strong zero-shot generalization and competitive performance with task-specific supervised models on climate, energy, and health benchmarks (5.3K+ stars, Apache 2.0, 2024-2026)
- [TabPFN (Prior Labs, Nature 2025)](https://github.com/PriorLabs/tabpfn) - Foundation model for tabular data that predicts on unseen real-world tables in a single forward pass, achieving accurate small-data classification and regression without task-specific training; widely applicable to scientific datasets with limited samples (7.4K+ stars, 2022-2026)
- [DeepInnovator (HKUDS, arXiv 2026)](https://github.com/HKUDS/DeepInnovator) - Scientific foundation model and AI research copilot for idea generation, cross-disciplinary connection discovery, and hypothesis formation; trained with a decoupled reward-comment RL architecture and achieves GPT-4o-competitive novelty/rationale on STEM and social-science idea-generation benchmarks (270+ stars, MIT License, 2026)
- [MinervaAI](https://github.com/google-research/minerva) - Mathematical reasoning
- [PaLM-2](https://ai.google/discover/palm2) - Scientific reasoning capabilities
### Domain-Specific Models
- [ESM](https://github.com/facebookresearch/esm) - Protein language models
- [BioNeMo Framework](https://github.com/NVIDIA/bionemo-framework) - NVIDIA's open-source platform for building and adapting biological AI models at scale, bundling ESM-2, Geneformer, MolMIM and DNA embedding models with recipes for single-GPU to multi-node training (2025)
- [IBM FM4M](https://github.com/IBM/materials) - IBM's open foundation model family for materials and chemistry, covering SMILES, SELFIES, molecular graphs, 3D atom positions, and electron density grids, with a unified toolkit for representation learning and downstream prediction/generation (Apache 2.0, 2024-2025)
- [ChemGPT](https://huggingface.co/ncfrey/ChemGPT-1.2B) - Chemistry-focused language model
- [BioGPT](https://github.com/microsoft/BioGPT) - Biomedical text generation
- [HuatuoGPT-o1 (2025)](https://github.com/FreedomIntelligence/HuatuoGPT-o1) - Open-source medical large language model for complex clinical reasoning, extending the o1 long-chain-of-thought paradigm to biomedical question answering and diagnostic inference (FreedomIntelligence, 1.3K+ stars)
- [AntAngelMed (2026)](https://github.com/MedAIBase/AntAngelMed) - 103B-parameter open-source medical language model with 1/32 Mixture-of-Experts architecture, achieving HealthBench-leading performance among open-source models with only 6.1B active parameters; jointly developed by Ant Group and Zhejiang Province Health Information Center (MIT License)
---
## 📈 Datasets & Benchmarks
### Multidisciplinary
- [Hugging Face Datasets](https://huggingface.co/datasets) - Comprehensive ML research datasets and scientific data collections
- [Google Dataset Search](https://datasetsearch.research.google.com/) - Find scientific datasets
### Biology & Medicine
- [TDC](https://github.com/mims-harvard/TDC) - Therapeutics Data Commons: 66 AI-ready datasets across 22 drug discovery tasks with 29 leaderboards, covering target identification, molecular generation, ADMET prediction, and clinical trial outcomes (Harvard MIMS, NeurIPS 2021/2024)
- [ProteinGym](https://github.com/OATML-Markslab/ProteinGym) - Large-scale benchmark suite for protein fitness prediction and design, aggregating 200+ deep mutational scanning assays and clinical variant datasets across diverse protein families and taxa, with standardized zero-shot and supervised leaderboards for variant effect prediction, mutation effect prediction, and protein language model evaluation (OATML & Marks Lab, NeurIPS 2023 Spotlight, Datasets & Benchmarks)
- [ProteinWorkshop](https://github.com/a-r-j/ProteinWorkshop) - Unified benchmarking framework for protein representation learning, providing standardized interfaces for pre-training and diverse downstream tasks including structure prediction, fitness prediction, and property prediction across multiple protein datasets and model architectures (ICLR 2024, 273+ stars, MIT License)
- [Protein Data Bank](https://www.rcsb.org/) - Protein structures
- [ChEMBL](https://www.ebi.ac.uk/chembl/) - Chemical bioactivity data
- [Human Protein Atlas](https://www.proteinatlas.org/) - Protein expression data
- [Chinese Medical Dataset](https://github.com/Mengqi97/chinese-medical-dataset) - Comprehensive collection of Chinese medical datasets for AI research
- [Arc Virtual Cell Atlas](https://github.com/ArcInstitute/arc-virtual-cell-atlas) - Curated open dataset collection of 602M+ observational and perturbational single-cell profiles for accelerating virtual cell model creation, integrating Tahoe-100M and scBaseCount data with Google Cloud Marketplace distribution (Arc Institute, 2025-2026)
### Chemistry & Materials
- [Materials Project](https://next-gen.materialsproject.org/) - Computational materials database
- [QM9](https://quantum-machine.org/datasets/) - Small molecule properties
- [Open Catalyst Project](https://opencatalystproject.org/) - Catalyst discovery
### Physics
- [The Well](https://github.com/PolymathicAI/the_well) - 15TB collection of 16 large-scale numerical simulation datasets spanning fluid dynamics, MHD, astrophysics, biological systems, and acoustic scattering, with unified PyTorch dataloaders and benchmarks for training foundation models on physical sciences (Polymathic AI, NeurIPS 2024)
- [RealPDEBench (ICLR 2026 Oral)](https://github.com/AI4Science-WestlakeU/RealPDEBench) - First scientific ML benchmark with paired real-world measurements and matched numerical simulations for complex physical systems, featuring 5 scenarios, 700+ trajectories, 10 baseline models, and 9 evaluation metrics with HuggingFace datasets and model checkpoints (Westlake University, CC BY-NC 4.0)
- [LIGO Open Science Center](https://gwosc.org/) - Gravitational wave data
- [Particle Data Group](https://pdg.lbl.gov/) - Particle physics data
- [OpenQuantumMaterials](https://www.quantum-materials.org/) - Quantum materials data
---
## 💻 Computing Frameworks
### Machine Learning
- [PyTorch](https://pytorch.org/) - Deep learning framework
- [JAX](https://github.com/jax-ml/jax) - High-performance ML research
- [TensorFlow](https://tensorflow.org/) - End-to-end ML platform
### Scientific Computing
- [NumPy](https://numpy.org/) - Numerical computing
- [SciPy](https://scipy.org/) - Scientific computing
- [Scikit-learn](https://scikit-learn.org/) - Machine learning library
### Scientific Machine Learning Frameworks
- [SciML](https://sciml.ai/) - Scientific machine learning ecosystem
- [DifferentialEquations.jl](https://github.com/SciML/DifferentialEquations.jl) - Multi-language suite for high-performance differential equation solving and scientific machine learning (3.0k+ stars)
- [ModelingToolkit.jl](https://github.com/SciML/ModelingToolkit.jl) - Acausal modeling framework for automatically parallelized scientific machine learning (1.5k+ stars)
- [SciMLBenchmarks.jl](https://github.com/SciML/SciMLBenchmarks.jl) - Scientific machine learning benchmarks & differential equation solvers
- [NeuralPDE.jl](https://github.com/SciML/NeuralPDE.jl) - Physics-informed neural networks (PINNs) for solving partial differential equations (1.1k+ stars)
- [DiffEqFlux.jl](https://github.com/SciML/DiffEqFlux.jl) - Neural ordinary differential equations with O(1) backprop and GPU support (900+ stars)
- [Optimization.jl](https://github.com/SciML/Optimization.jl) - Unified interface for local, global, gradient-based and derivative-free optimization (800+ stars)
- [PaddleScience](https://github.com/PaddlePaddle/PaddleScience) - SDK & library for AI-driven scientific computing applications
- [Tesseract Core (Pasteur Labs, SciPy 2025 / JOSS)](https://github.com/pasteurlabs/tesseract-core) - Universal components for differentiable scientific computing, packaging heterogeneous scientific tools into self-contained, portable, gradient-propagating components with auto-generated schemas, CLI/REST API/Python SDK interfaces, and reproducible deployment across local, cloud, and HPC environments (105+ stars, Apache 2.0)
- [Flux.jl](https://github.com/FluxML/Flux.jl) - Machine learning in Julia
### Specialized Frameworks
- [MDAnalysis](https://github.com/MDAnalysis/mdanalysis) - Molecular dynamics analysis
- [e3nn](https://github.com/e3nn/e3nn) - Euclidean neural networks for arbitrary point transformations enabling E(3)-equivariant deep learning, foundational library for building geometry-aware neural networks in molecular dynamics, materials science, and physics
- [MDtrajNet](https://arxiv.org/abs/2505.16301) - Neural network foundation model that directly generates MD trajectories bypassing force calculations, accelerating simulations by up to 100× with equivariant Transformer architecture (2025)
- [ASE](https://wiki.fysik.dtu.dk/ase/) - Atomic Simulation Environment for materials modeling
- [PyMC](https://github.com/pymc-devs/pymc) - Probabilistic programming
- [AI2BMD](https://github.com/microsoft/AI2BMD) - Microsoft's AI-powered ab initio biomolecular dynamics simulation achieving quantum-mechanical accuracy for proteins with 10,000+ atoms, orders of magnitude faster than DFT using protein fragmentation and ML force fields (Nature 2024)
- [OpenMM](https://github.com/openmm/openmm) - High-performance molecular simulation toolkit
- [DeePMD-kit](https://github.com/deepmodeling/deepmd-kit) - Deep learning package for many-body potential energy representation and molecular dynamics, achieving quantum-mechanical accuracy with classical MD efficiency (DeepModeling, Gordon Bell Prize 2020, 1.9k+ stars)
- [TorchMD](https://github.com/torchmd/torchmd) - End-to-end molecular dynamics engine built on PyTorch, enabling differentiable simulations with neural network potentials and GPU acceleration for machine learning-accelerated molecular dynamics (MIT License, 707+ stars)
- [SO3LR](https://github.com/general-molecular-simulations/so3lr) - Pretrained machine-learned force field for (bio)molecular simulations combining the fast SO3krates neural network for semi-local interactions with universal pairwise force fields for short-range repulsion, long-range electrostatics, and dispersion interactions; supports geometry optimization, NVT/NPT/NVE MD, fine-tuning, ASE calculator, and JAX-MD integration (JACS 2025, 218+ stars, MIT License)
- [TorchSim](https://github.com/TorchSim/torch-sim) - PyTorch-native atomistic simulation engine for the machine-learned interatomic potential (MLIP) era, enabling batched molecular dynamics and structural relaxation with automatic GPU memory management; supports MACE, Fairchem, SevenNet, ORB, MatterSim and other popular MLIPs with up to 100x speedup over ASE (Radical AI, AI for Science 2026, 468+ stars, MIT License)
- [Newton](https://github.com/newton-physics/newton) - GPU-accelerated differentiable physics simulation engine built on NVIDIA Warp, supporting rigid/soft body, cloth, and gradient-based optimization for scientific ML, initiated by Disney Research, DeepMind, and NVIDIA (Linux Foundation, Apache 2.0, 2025)
- [JAX-CFD](https://github.com/google/jax-cfd) - Computational fluid dynamics in JAX, enabling differentiable Navier-Stokes simulations with automatic differentiation for ML-accelerated CFD research, supporting turbulence modeling, convection-diffusion, and complex boundary conditions on CPUs and GPUs (Google Research, 947+ stars)
- [PennyLane](https://github.com/PennyLaneAI/pennylane) - Cross-platform library for differentiable programming of quantum computers with automatic differentiation, enabling hybrid quantum-classical machine learning for quantum chemistry, quantum physics, and NISQ algorithm research (Xanadu, 3k+ stars)
- [Qiskit](https://github.com/Qiskit/qiskit) - Open-source SDK for working with quantum computers at the level of extended quantum circuits, operators, and primitives, enabling quantum algorithm development for quantum chemistry, materials science, and optimization research (IBM, 7.4K+ stars, Apache 2.0)
- [DGL](https://github.com/dmlc/dgl) - Deep Graph Library for scalable deep learning on graphs, powering molecular modeling, materials discovery, protein interaction networks, and scientific knowledge graph learning across PyTorch, TensorFlow, and MXNet backends (14K+ stars)
- [PyTorch Geometric](https://github.com/pyg-team/pytorch_geometric) - Graph neural network library for PyTorch enabling molecular modeling, materials discovery, protein interaction networks, and scientific knowledge graph learning (23.7k+ stars)
---
## 🎓 Educational Resources
### Courses & Tutorials
- [AI for Everyone (Coursera)](https://www.coursera.org/learn/ai-for-everyone) - Basic AI concepts
- [CS229 Machine Learning](https://cs229.stanford.edu/) - Stanford ML course
- [MIT 6.034 Artificial Intelligence](https://ocw.mit.edu/courses/6-034-artificial-intelligence-fall-2010/) - AI fundamentals
### Open Access Educational Materials
- [SciML Book](https://github.com/SciML/SciMLBook) - Parallel Computing and Scientific Machine Learning: MIT 18.337J/6.338J course materials (1.9k+ stars)
- [Dive into Deep Learning](https://d2l.ai/) - Interactive deep learning book with code implementations
- [The Elements of Statistical Learning](https://hastie.su.stanford.edu/ElemStatLearn/) - Classic ML textbook freely available
- [Neural Networks and Deep Learning](http://neuralnetworksanddeeplearning.com/) - Free online book by Michael Nielsen
### 📋 Paper Collections & Repositories
- [Awesome Scientific Language Models](https://github.com/yuzhimanhua/Awesome-Scientific-Language-Models) - Curated scientific LLM papers (260+ models)
- [Awesome LLM Scientific Discovery](https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery) - LLM papers for scientific discovery
- [AI4Research Papers](https://github.com/du-nlp-lab/LLM4SR) - LLM for scientific research papers
- [Physics-Informed Neural Networks Papers](https://github.com/idrl-lab/PINNpapers) - PINN research collection
- [Scientific Computing with ML Papers](https://sciml.ai/papers/) - Scientific ML paper repository
- [Simulation-Based Inference Papers & Tools](https://simulation-based-inference.org/papers/) - Community-maintained SBI research portal with papers and software
- [Awesome AI Scientist Papers](https://github.com/openags/Awesome-AI-Scientist-Papers) - Autonomous AI scientist research
- [Awesome Agents for Science](https://github.com/OSU-NLP-Group/awesome-agents4science) - LLM agents across scientific domains
### YouTube Channels
- [Two Minute Papers](https://www.youtube.com/c/KárolyZsolnai) - AI research summaries
- [3Blue1Brown](https://www.youtube.com/c/3blue1brown) - Mathematical concepts
- [AI Coffee Break](https://www.youtube.com/c/AICoffeeBreak) - AI paper reviews
- [Steve Brunton](https://www.youtube.com/c/Eigensteve) - Data-driven methods
- [Nathan Kutz](https://www.youtube.com/c/NathanKutz) - Applied mathematics
- [Physics Informed Machine Learning](https://www.youtube.com/c/PIML) - SciML tutorials
---
## 🏛 Research Communities
### Conferences
- [NeurIPS](https://neurips.cc/) - Machine learning conference
- [ICML](https://icml.cc/) - International Conference on Machine Learning
- [AI for Science Workshop](https://ai4sciencecommunity.github.io/) - Specialized workshops
### Organizations
- [Partnership on AI](https://partnershiponai.org/) - AI research collaboration
- [Allen Institute for AI](https://allenai.org/) - AI research institute
- [OpenAI](https://openai.com/) - AI research and deployment
### Online Communities
- [r/MachineLearning](https://reddit.com/r/MachineLearning) - ML discussions
- [AI Alignment Forum](https://www.alignmentforum.org/) - AI safety research
- [Distill](https://distill.pub/) - Visual explanations of ML
---
## 📚 Related Awesome Lists
This project builds upon and complements several excellent resources:
### 🎯 Specialized Collections
- [awesome-ai4s](https://github.com/hyperai/awesome-ai4s) - 200+ AI for Science papers with Chinese interpretations
- [Awesome AI Scientist Papers](https://github.com/openags/Awesome-AI-Scientist-Papers) - Autonomous AI scientist research
- [Awesome Scientific Machine Learning](https://github.com/MartinuzziFrancesco/awesome-scientific-machine-learning) - Physics-informed ML and SciML
- [Awesome Agents for Science](https://github.com/OSU-NLP-Group/awesome-agents4science) - LLM agents across scientific domains
- [Awesome LLM Agents Scientific Discovery](https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery) - Biomedical AI agents
- [Awesome Foundation Models for Weather and Climate](https://github.com/shengchaochen82/Awesome-Foundation-Models-for-Weather-and-Climate) - Comprehensive survey of foundation models for weather and climate data understanding
### 📊 Paper & Research Collections
- [Scientific LLM Papers](https://github.com/yuzhimanhua/Awesome-Scientific-Language-Models) - 260+ scientific language models
- [LLM4SR Repository](https://github.com/du-nlp-lab/LLM4SR) - LLM for scientific research survey materials
- [PINNs Paper Collection](https://github.com/idrl-lab/PINNpapers) - Physics-informed neural networks research
- [SciML Papers](https://sciml.ai/papers/) - Scientific computing and machine learning papers
### 🌟 Key Insights from These Collections
- **Current Focus**: Shift from tool-level assistance to autonomous scientific agents
- **Emerging Trends**: Multi-modal scientific models, self-improving research systems
- **Research Gaps**: Evaluation frameworks, ethical governance, human-AI collaboration
- **Future Directions**: Fully autonomous discovery cycles, robotic lab integration
---
## 🤝 Contributing
We welcome contributions! Please see our [Contributing Guidelines](CONTRIBUTING.md) for details.
### How to Contribute
1. Fork this repository
2. Add your resource in the appropriate section
3. Ensure the format matches existing entries
4. Submit a pull request with a clear description
### Contribution Guidelines
- Ensure the resource is actively maintained
- Include a brief, clear description
- Check for duplicates before adding
- Use proper markdown formatting
---
## 📄 License
This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
---
## 🙏 Acknowledgments
Special thanks to all researchers and developers pushing the boundaries of AI for Science. This list is inspired by the awesome community and the transformative potential of AI in scientific discovery.
**Star ⭐ this repository if you find it helpful!**
---
@@ -0,0 +1,28 @@
---
title: "Pull Request Template"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/PULL_REQUEST_TEMPLATE.md
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
## Description
<!-- Please describe what you are adding or changing, and why it is awesome. -->
## Checklist
- [ ] I have searched previous suggestions and this is not a duplicate.
- [ ] I have added only one link per pull request.
- [ ] The link follows the format: `[name](https://example.com/)` - A short description ends with a period.
- [ ] Descriptions are concise.
- [ ] Alphabetical ordering is maintained where applicable.
- [ ] If a new section was added, the section description and title are included, and the title is added to the Index.
- [ ] Spelling and grammar have been checked.
- [ ] There is no trailing whitespace.
- [ ] The pull request title follows the format: `Add user/repo - Short repo description`
@@ -0,0 +1,46 @@
---
title: "Docs Check"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/docs-check.yml
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
name: Docs / lint
on:
push:
paths:
- '**/*.md'
pull_request:
paths:
- '**/*.md'
permissions:
contents: read
jobs:
docs-check:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: 'lts/*'
- name: Install tooling
run: npm install -g [email protected] [email protected] --no-fund
- name: Run markdownlint
run: markdownlint '**/*.md' --ignore node_modules
- name: Run cspell (spellcheck)
run: cspell "**/*.md" --no-summary --no-progress
@@ -0,0 +1,55 @@
---
title: "Generate Artifacts"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/generate-artifacts.yml
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
name: Generate data artifacts
on:
push:
paths:
- 'data/resources.yml'
pull_request:
paths:
- 'data/resources.yml'
workflow_dispatch:
jobs:
generate:
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- name: Checkout
uses: actions/checkout@v4
with:
ref: ${{ github.head_ref || github.ref_name }}
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install dependencies
run: pip install pyyaml
- name: Generate JSON and CSV
run: python scripts/generate_artifacts.py
- name: Commit artifacts (push events only)
if: github.event_name == 'push' || github.event_name == 'workflow_dispatch'
run: |
git config user.name "github-actions[bot]"
git config user.email "github-actions[bot]@users.noreply.github.com"
git add data/resources.json data/resources.csv docs/data/resources.json
git diff --cached --quiet || git commit -m "chore: regenerate resources.json and resources.csv"
git push
@@ -0,0 +1,42 @@
---
title: "Link Check"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/link-check.yml
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
name: Link Check
on:
schedule:
- cron: '0 5 * * 1' # 毎週月曜 05:00 UTC
workflow_dispatch: # 手動実行も可能
permissions:
contents: read
jobs:
link-check:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: 'lts/*'
- name: Install markdown-link-check
run: npm install -g [email protected] --no-fund
- name: Run markdown-link-check
run: |
find . -name '*.md' -not -path './node_modules/*' -print0 | \
xargs -0 -I{} markdown-link-check {} -q -c .markdown-link-check.json
@@ -0,0 +1,56 @@
---
title: "Pr Quality Checks"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/pr-quality-checks.yml
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
name: PR Quality Checks
on:
pull_request:
paths:
- README.md
- data/resources.yml
- data/resources.json
- data/resources.csv
- docs/data/resources.json
- scripts/*.py
- scripts/**/*.py
permissions:
contents: read
jobs:
resources-consistency:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Setup uv
uses: astral-sh/setup-uv@v3
- name: Sync resources from README
run: uv run --with pyyaml python scripts/sync_resources_from_readme.py
- name: Build resource artifacts
run: uv run --with pyyaml python scripts/build_resources.py
- name: Verify generated artifacts are committed
run: |
if git diff --quiet; then
echo "Resources are in sync."
exit 0
fi
echo "Generated files are out of date. Run:"
echo " uv run --with pyyaml python scripts/sync_resources_from_readme.py"
echo " uv run --with pyyaml python scripts/build_resources.py"
git status --short
exit 1
@@ -0,0 +1,73 @@
---
title: "Sync Resources"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/sync_resources.yml
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
name: Sync Resources
on:
push:
branches:
- main
paths:
- README.md
- data/resources.yml
- scripts/*.py
- scripts/**/*.py
permissions:
contents: write
jobs:
sync:
if: github.actor != 'github-actions[bot]'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Setup uv
uses: astral-sh/setup-uv@v3
- name: Detect Changed Files
id: changes
run: |
BEFORE="${{ github.event.before }}"
AFTER="${{ github.sha }}"
if [ -z "$BEFORE" ] || [ "$BEFORE" = "0000000000000000000000000000000000000000" ]; then
CHANGED=$(git diff --name-only HEAD~1..HEAD)
else
CHANGED=$(git diff --name-only "$BEFORE" "$AFTER")
fi
echo "changed<<EOF" >> "$GITHUB_OUTPUT"
echo "$CHANGED" >> "$GITHUB_OUTPUT"
echo "EOF" >> "$GITHUB_OUTPUT"
- name: Sync From README
if: contains(steps.changes.outputs.changed, 'README.md')
run: uv run python scripts/sync_resources_from_readme.py
- name: Build Artifacts
run: uv run --with pyyaml python scripts/build_resources.py
- name: Commit Updates
run: |
if git diff --quiet; then
echo "No changes to commit."
exit 0
fi
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add data/resources.yml data/resources.json data/resources.csv docs/data/resources.json
git commit -m "chore: sync resources"
git push
@@ -0,0 +1,58 @@
---
title: "Update Overview"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/update-overview.yml
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
name: Update Overview Figure
on:
push:
branches:
- main
paths:
- docs/data/resources.json
schedule:
# Every Monday at 03:00 UTC
- cron: '0 3 * * 1'
workflow_dispatch:
permissions:
contents: write
jobs:
update-overview:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
with:
# Use the default GITHUB_TOKEN so commits don't re-trigger this workflow
token: ${{ secrets.GITHUB_TOKEN }}
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install dependencies
run: pip install matplotlib
- name: Regenerate overview figure
run: python scripts/generate_overview.py
- name: Commit updated overview
run: |
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add docs/overview.png
git diff --cached --quiet || git commit -m "chore: regenerate overview figure"
git push
@@ -0,0 +1,16 @@
---
title: ".Markdown Link Check"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.markdown-link-check.json
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
{
"aliveStatusCodes": [200, 206, 301, 302, 307, 308, 403]
}
@@ -0,0 +1,539 @@
---
title: "Awesome Computational Biology [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/0cf037ab/README.md
upstream_sha: 0cf037ab
imported_at: 2026-07-16
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Awesome Computational Biology [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)
A curated collection of databases, software, and papers related to computational biology.
> Computational biology involves the development and application of data-analytical and theoretical methods, mathematical modelling and computational simulation techniques to the study of biological, ecological, behavioural, and social systems. — [Wikipedia](https://en.wikipedia.org/wiki/Computational_biology)
---
## Overview
[![Resource Landscape Overview](docs/overview.png)](https://inoue0426.github.io/awesome-computational-biology/overview.html)
> Interactive version: [Resource Overview page](https://inoue0426.github.io/awesome-computational-biology/overview.html)
> Regenerate the figure: `python scripts/generate_overview.py`
---
## GitHub Pages UI
Browse and search the resources via the [GitHub Pages UI](https://inoue0426.github.io/awesome-computational-biology/).
- Search matches `name`, `description`, `tasks`, `modalities`, and `tags`.
- The **Task**, **Modality**, and **Type** filters map directly to `tasks`, `modalities`, and `type` in `docs/data/resources.json`.
- Clicking badges on cards applies the corresponding filter.
---
## Table of Contents
- [Awesome Computational Biology](#awesome-computational-biology-)
- [Table of Contents](#table-of-contents)
- [Overview](#overview)
- [GitHub Pages UI](#github-pages-ui)
- [Citation](#citation)
- [Curation Criteria (Strict)](#curation-criteria-strict)
- [Update & Link Rot Policy](#update--link-rot-policy)
- [Data Schema & Contribution Workflow](#data-schema--contribution-workflow)
- [Databases](#databases)
- [scRNA](#scrna)
- [Compound](#compound)
- [Pathway](#pathway)
- [Mass Spectra](#mass-spectra)
- [Protein](#protein)
- [Genome](#genome)
- [Disease](#disease)
- [Interaction](#interaction)
- [Drug-Gene Interaction](#drug-gene-interaction)
- [Drug (Cell Line) Response](#drug-cell-line-response)
- [Chemical-Protein Interaction](#chemical-protein-interaction)
- [Protein-Protein Interaction](#protein-protein-interaction)
- [Knowledge Graph](#knowledge-graph)
- [Gene Regulatory Network](#gene-regulatory-network)
- [Clinical Trial](#clinical-trial)
- [Benchmarks & Datasets](#benchmarks--datasets)
- [API](#api)
- [Preprocessing Tools](#preprocessing-tools)
- [Machine Learning Tasks and Models](#machine-learning-tasks-and-models)
- [Drug Discovery](#drug-discovery)
- [Drug Response Prediction](#drug-response-prediction)
- [Drug Repurposing](#drug-repurposing)
- [Drug Target Interaction](#drug-target-interaction)
- [Compound-Protein Interaction](#compound-protein-interaction)
- [Molecular Generation](#molecular-generation)
- [LLM for Biology](#llm-for-biology)
- [Foundation Models](#foundation-models)
- [Single-cell Foundation Models](#single-cell-foundation-models)
- [Transcriptomics Foundation Models](#transcriptomics-foundation-models)
- [Spatial Foundation Models](#spatial-foundation-models)
- [Multi-Omics Foundation Models](#multi-omics-foundation-models)
- [Domain Alignment](#domain-alignment)
- [Compound Foundation Models](#compound-foundation-models)
- [Compound Embedding](#compound-embedding)
- [Protein Foundation Models](#protein-foundation-models)
- [Pre-trained Embedding](#pre-trained-embedding)
- [Protein Structure Prediction and Design](#protein-structure-prediction-and-design)
- [Multi-Modal Foundation Models](#multi-modal-foundation-models)
- [Genomics Foundation Models](#genomics-foundation-models)
---
## Databases
### scRNA
- [CZ CELLxGENE](https://cellxgene.cziscience.com/) — Single-cell dataset repository and interactive explorer from the Chan Zuckerberg Initiative.
- [Gene Expression Omnibus](https://www.ncbi.nlm.nih.gov/geo/) — Public functional genomics database.
- [Human Cell Atlas](https://www.humancellatlas.org/) — Open global atlas of all cells in the human body.
- [Single Cell PORTAL](https://singlecell.broadinstitute.org/single_cell) — Public database for single-cell RNA.
- [Single Cell Expression Atlas](https://www.ebi.ac.uk/gxa/sc/home) — Public database for single-cell RNA.
### Compound
- [PubChem](https://pubchem.ncbi.nlm.nih.gov/) — One of the largest chemical databases (compounds, genes, and proteins).
- [ChEBI](https://www.ebi.ac.uk/chebi/) — Database focused on small chemical compounds.
- [ChEMBL](https://www.ebi.ac.uk/chembl/) — Bioactive molecules with drug-like properties.
- [ChemSpider](http://www.chemspider.com/) — Chemical structure database.
- [DrugTargetCommons](https://drugtargetcommons.fimm.fi/) — Community platform for curating and integrating experimental bioactivity data across drugs and targets.
- [HMDB (Human Metabolome Database)](https://hmdb.ca/) — Comprehensive database of small molecule metabolites found in the human body.
- [KEGG COMPOUND](https://www.genome.jp/kegg/compound/) — Collection of small molecules and biopolymers.
- [LIPID MAPS](https://www.lipidmaps.org/databases/lmsd/overview) — Database of lipids.
- [Rhea](https://www.rhea-db.org/) — Database of chemical reactions.
- [DrugCentral](http://drugcentral.org/) — Online drug compendium with drug mode of action and indication information.
- [Drug Repurposing Hub](https://repo-hub.broadinstitute.org/repurposing#download-data) — Collections of drug repurposing data (drug, MoA, target, etc).
- [Therapeutic Target Database](https://idrblab.net/ttd/full-data-download) — Drug-target, target-disease, and drug-disease datasets.
- [ZINC ligand discovery database](https://zinc.docking.org/) — Free database of commercially-available compounds for virtual screening.
### Pathway
- [PathwayCommons](https://www.pathwaycommons.org/) — Database of pathways and interactions.
- [KEGG PATHWAY](https://www.genome.jp/kegg/pathway.html) — Collection of pathway maps.
- [WikiPathways](https://wikipathways.org/) — Database of biological pathways.
- [Reactome](https://reactome.org/) — Expert-curated, peer-reviewed pathway database with detailed reaction mechanisms.
- [BioCyc](https://biocyc.org/) — Collection of pathway/genome databases across thousands of organisms.
- [OmniPath](https://omnipathdb.org/) — Comprehensive resource integrating protein interactions, signaling pathways, gene regulatory networks, and miRNA targets from over 100 databases.
- [SIGNOR 2.0](https://signor.uniroma2.it/) — Database of causal signaling interactions and pathways, with signed and directed relationships between proteins.
- [MSigDB (Molecular Signatures Database)](https://www.gsea-msigdb.org/gsea/msigdb) — Curated gene sets derived from pathways and biological processes.
### Mass Spectra
- [MassBank](http://www.massbank.jp/) — Open source databases and tools for mass spectrometry reference spectra.
- [MoNA MassBank of North America](https://mona.fiehnlab.ucdavis.edu/) — Meta-database of metabolite mass spectra, metadata, and associated compounds.
### Protein
- [THE HUMAN PROTEIN ATLAS](https://www.proteinatlas.org/) — Comprehensive human protein database (cells, tissues, organs).
- [PROTEIN DATA BANK (PDB)](https://www.rcsb.org/) — 3D structures of proteins, nucleic acids, complexes.
- [UniProt](https://www.uniprot.org/) — Functional information on proteins.
- [AlphaFold Protein Structure Database](https://alphafold.ebi.ac.uk/api-docs) — 3D protein structure predictions.
- [RCSB Protein Data Bank](https://www.rcsb.org/) — Repository for structural data of biological molecules.
- [Critical Assessment of Structure Prediction (CASP)](https://predictioncenter.org/) — Assessing methods for protein structure prediction.
- [Uniclust](https://uniclust.mmseqs.com/) — Clustered protein sequence databases.
- [UniRef](https://www.uniprot.org/uniref/) — Non-redundant sequence database clustering UniProtKB entries at multiple sequence identity thresholds.
- [CATH database](https://www.cathdb.info/) — Hierarchical classification of protein domain structures.
- [SAbDab](https://opig.stats.ox.ac.uk/webapps/sabdab-sabpred/sabdab) — Structural Antibody Database containing all antibody structures in the PDB.
- [OADB (Observed Antibody Space Database)](http://opig.stats.ox.ac.uk/webapps/oas/) — Database of antibody sequences from immune repertoire sequencing.
- [InterPro](https://www.ebi.ac.uk/interpro/) — Protein families, domains, and functional sites database integrating 14 member databases including Pfam and PROSITE.
- [Pfam](https://www.ebi.ac.uk/interpro/entry/pfam/) — Database of protein families described by multiple sequence alignments and hidden Markov models.
- [NeXtProt](https://www.nextprot.org/) — Expert knowledge base on human proteins with deep functional annotation, complementary to UniProt.
### Genome
- [ENCODE](https://www.encodeproject.org/) — Encyclopedia of DNA Elements; regulatory and functional genomic elements across the genome.
- [Ensembl](https://www.ensembl.org/) — Genome browser and annotation database for vertebrate and other eukaryotic genomes.
- [Human Genome Resources at NCBI](https://www.ncbi.nlm.nih.gov/projects/genome/guide/human/index.shtml) — Database for genomics, proteomics, transcriptomics, and systems biology.
- [GenBank](https://www.ncbi.nlm.nih.gov/genbank/) — NCBI's database of genetic sequences.
- [UCSC Genome Browser](https://genome.ucsc.edu/) — UCSC's genome browser.
- [cBioPortal](https://www.cbioportal.org/) — Cancer genomics database; aggregating many patient datasets.
- [10x Genomics Dataset](https://www.10xgenomics.com/resources/datasets) — Collection of single-cell datasets.
- [The Genotype-Tissue Expression (GTEx)](https://gtexportal.org/home/) — Human gene expression and regulation resource.
- [Dependency Map (DepMap)](https://depmap.org/portal/) — CRISPR-Cas9 screens in cancer cell lines.
- [Catalogue Of Somatic Mutations In Cancer (COSMIC)](https://cancer.sanger.ac.uk/cosmic) — Resource on somatic mutations in cancers.
- [MGnify](https://www.ebi.ac.uk/metagenomics/) — Resource for metagenomic and metatranscriptomic data.
- [JASPAR](http://jaspar.genereg.net/) — Database of transcription factor binding profiles.
- [gnomAD](https://gnomad.broadinstitute.org/) — Genome Aggregation Database; genetic variation from large-scale sequencing projects.
- [Rfam](https://rfam.org/) — Database of RNA families with sequence alignments and consensus structures.
- [ROADMAP Epigenomics](http://www.roadmapepigenomics.org/) — Reference epigenome maps for 111 primary human cell types and tissues, including histone modifications, chromatin accessibility, and DNA methylation.
- [FANTOM5](https://fantom.gsc.riken.jp/5/) — Functional annotation of mammalian genome; comprehensive atlas of active enhancers, promoters, and transcription start sites across human and mouse cell types.
### Disease
- [KEGG DRUG](https://www.genome.jp/kegg/drug/) — Comprehensive, approved drug information.
- [DrugBank](https://go.drugbank.com/) — Database of drugs and targets (University of Alberta).
- [DisGeNET](https://www.disgenet.org/) — Database of gene-disease associations integrating expert-curated and GWAS data.
- [OMIM (Online Mendelian Inheritance in Man)](https://www.omim.org/) — Comprehensive database of human genes and genetic disorders.
- [Open Targets Platform](https://platform.opentargets.org/) — Systematic target identification and prioritization platform integrating genetics, genomics, and drug data for drug discovery.
- [Human Phenotype Ontology (HPO)](https://hpo.jax.org/) — Standardized vocabulary of phenotypic abnormalities in human disease, linking genes, variants, and clinical features.
- [DISEASES](https://diseases.jensenlab.org/) — Genedisease association database integrating evidence from text mining, curated databases, and experimental data.
### Interaction
#### Drug-Gene Interaction
- [DGIdb](https://www.dgidb.org/) — Drug-gene interactions and the druggable genome.
- [Comparative Toxicogenomics Database](http://ctdbase.org/) — Chemical-gene interactions, chemical-disease and gene-disease associations, chemical-phenotype associations.
- [SNAP](https://snap.stanford.edu/biodata/datasets/10002/10002-ChG-Miner.html) — Dataset of drug-gene interactions.
#### Drug (Cell Line) Response
- [NCI60](https://dtp.cancer.gov/discovery_development/nci-60/) — Focuses on 60 cancer cell lines and many drugs.
- [Genomics of Drug Sensitivity in Cancer (GDSC)](https://www.cancerrxgene.org/) — Drug sensitivity for ~1000 human cancer cell lines and hundreds of compounds.
- [Cancer Cell Line Encyclopedia](https://sites.broadinstitute.org/ccle/) — Database of ~1000 cancer cell lines.
- [CellMiner Cross Database (CellMinerCDB)](https://discover.nci.nih.gov/cellminercdb/) — Integrates multiple cancer cell line databases.
#### Chemical-Protein Interaction
- [STITCH](http://stitch.embl.de/) — Chemical-protein interactions.
- [BindingDB](https://www.bindingdb.org/rwd/bind/index.jsp) — Compounds and target database.
- [Davis kinase inhibitors DB](http://staff.cs.utu.fi/~aijrinas/dti/) — Experimental kinase inhibitor binding affinity dataset for proteinligand interaction research.
- [Kinase Inhibitor Bioactivity Data (KIBA)](https://janeliascicomp.github.io/KIBA/) — Integrated bioactivity scores for kinase inhibitors combining Ki, Kd, and IC50 measurements.
- [PDBBind](https://www.pdbbind-plus.org.cn/) — Binding affinity data for biomolecular complexes.
#### Protein-Protein Interaction
- [STRING](https://string-db.org/) — PPI networks for multiple organisms.
- [BioGRID](https://thebiogrid.org/) — Protein, genetic, and chemical interactions.
- [HIPPIE](http://cbdm-01.zdv.uni-mainz.de/~mschaefer/hippie/) — Human protein-protein interaction database.
- [IntAct](https://www.ebi.ac.uk/intact/home) — Open-source molecular interaction database and analysis system from EMBL-EBI.
#### Knowledge Graph
- [Drug Mechanism Database (DrugMechDB)](https://github.com/SuLab/DrugMechDB/tree/2.0.1) — Mechanisms of action from drug to disease.
- [DRKG](https://github.com/gnn4dr/DRKG) — Large-scale biological knowledge graph for drug discovery.
- [Hetionet](https://github.com/hetio/hetionet) — Heterogeneous network integrating genes, diseases, drugs, pathways, and more.
- [PrimeKG](https://github.com/mims-harvard/PrimeKG) — Multi-modal precision medicine knowledge graph integrating clinical, genetic, and drug data.
#### Gene Regulatory Network
- [TRRUST v2](https://www.grnpedia.org/trrust/) — Manually curated database of human and mouse transcriptional regulatory interactions between transcription factors and their target genes, expanded with literature-derived evidence.
- [RegNetwork](http://www.regnetworkweb.org/) — Database of gene regulatory networks covering transcription factortarget gene and miRNAgene interaction data across multiple species.
- [miRBase](https://www.mirbase.org/) — Reference repository for microRNA gene annotations, sequences, and experimentally validated targets.
### Clinical Trial
- [ClinicalTrials.gov](https://clinicaltrials.gov/) — Privately and publicly funded clinical studies.
- [ICD10](https://icd.who.int/browse10/2019/en) — International Classification of Diseases, 10th revision.
- [EU Drug Regulating Authorities Clinical Trials DB (EudraCT)](https://eudract.ema.europa.eu/) — European clinical trial database.
- [MIMIC-IV](https://mimic.mit.edu/) — Freely accessible critical care database.
---
## Benchmarks & Datasets
- [1000 Genomes Project](https://www.internationalgenome.org/) — Reference panel of human genetic variation from 2,504 individuals across 26 populations.
- [BACE](https://www.kaggle.com/datasets/gokturkkoch/bace) — Binary classification and regression dataset for β-secretase 1 (BACE-1) inhibitor binding affinity.
- [BEAT AML](https://biodev.github.io/BeatAML2/) — Functional ex vivo drug sensitivity measurements paired with genomics for acute myeloid leukemia.
- [Bento](https://github.com/LigandPro/Bento) — Protein-ligand docking benchmark covering rigid, flexible, de novo, blind, induced-fit, and covalent docking tasks.
- [BindingDB Curated Sets](https://www.bindingdb.org/rwd/bind/chemsearch/marvin/SDFdownload.jsp?all_download=yes) — Curated binding affinity datasets for proteinligand interaction benchmarking.
- [Cancer Therapeutics Response Portal (CTRP)](https://portals.broadinstitute.org/ctrp/) — Drug sensitivity profiles across ~900 cancer cell lines for >400 compounds.
- [ClinTox](https://tdcommons.ai/single_pred_tasks/tox/#clintox) — Clinical toxicity dataset contrasting FDA-approved drugs with those that failed clinical trials due to toxicity.
- [CPTAC (Clinical Proteomic Tumor Analysis Consortium)](https://proteomics.cancer.gov/programs/cptac) — Multi-omic proteogenomic datasets for multiple cancer types linking proteomics with genomics.
- [CrossDocked2020](https://arxiv.org/abs/2001.01037) — Large-scale dataset for structure-based virtual screening.
- [DUD-E (Directory of Useful Decoys, Enhanced)](http://dude.docking.org/) — Structure-based virtual screening benchmark with active ligands and challenging decoy sets across diverse protein targets.
- [FLIP (Fitness Landscape Inference for Proteins)](https://github.com/J-SNACKKB/FLIP) — Benchmark collection of protein fitness landscape datasets for evaluating protein ML models.
- [Genomics of Drug Sensitivity in Cancer (GDSC)](https://www.cancerrxgene.org/) — Drug sensitivity for ~1000 human cancer cell lines and hundreds of compounds.
- [GuacaMol](https://github.com/BenevolentAI/guacamol) — Benchmark suite for generative molecular design models.
- [JUMP Cell Painting Datasets](https://github.com/jump-cellpainting/datasets) — Consortium-scale cell imaging perturbation datasets (chemical and genetic) for phenotypic profiling and drug discovery research.
- [LINCS L1000](https://lincsproject.org/LINCS/tools/workflows/find-the-best-place-to-obtain-the-lincs-l1000-data) — Gene expression profiles (978 landmark genes) for >20,000 chemical and genetic perturbations across cell lines.
- [MoleculeNet](http://moleculenet.ai/) — Benchmark datasets for molecular machine learning.
- [MOSES](https://github.com/molecularsets/moses) — Benchmarking platform for molecular generation models.
- [NCI60](https://dtp.cancer.gov/discovery_development/nci-60/) — Drug sensitivity benchmark across 60 diverse human cancer cell lines.
- [OGB (Open Graph Benchmark)](https://ogb.stanford.edu/) — Large-scale graph ML benchmark suite including biological datasets such as ogbl-ppa (protein-protein associations) and ogbg-molhiv.
- [OpenBioLink](https://github.com/OpenBioLink/OpenBioLink) — Benchmark datasets for biological knowledge graph completion.
- [PharmGKB](https://www.pharmgkb.org/) — Curated pharmacogenomics dataset linking genetic variants to drug response phenotypes across thousands of drugs.
- [PK-DB](https://pk-db.com/) — Open database of experimental pharmacokinetics (PK) and ADME data from clinical and preclinical studies.
- [PRISM](https://depmap.org/portal/prism/) — Cancer drug sensitivity profiling of >4,500 drugs across >900 cancer cell lines using pooled-cell-line barcoding.
- [ProteinGym](https://github.com/OATML-Markslab/ProteinGym) — Large-scale benchmark of deep mutational scanning assays for evaluating protein fitness landscape models.
- [QM9](https://figshare.com/collections/Quantum_chemistry_structures_and_properties_of_134_kilo_molecules/978904) — Quantum chemistry properties for 134K stable small organic molecules computed at DFT level.
- [scIB (Single-cell Integration Benchmarks)](https://github.com/theislab/scib) — Comprehensive benchmarking framework for single-cell data integration methods.
- [scPerturb](https://github.com/sanderlab/scPerturb) — Curated and continuously updated single-cell perturbation data resource spanning CRISPR and drug perturbation studies.
- [SIDER (Side Effect Resource)](http://sideeffects.embl.de/) — Database of 1,430 approved drugs with their recorded adverse drug reactions across 27 system-organ classes.
- [Tabula Muris](https://tabula-muris.ds.czbiohub.org/) — Comprehensive single-cell atlas of 20 mouse organs and tissues, enabling cross-tissue and cross-species comparisons.
- [Tabula Sapiens](https://tabula-sapiens-portal.ds.czbiohub.org/) — Comprehensive human single-cell atlas of ~500K cells from 24 organs and tissues across multiple donors.
- [TAPE (Tasks Assessing Protein Embeddings)](https://github.com/songlab-cal/tape) — Benchmark suite of five biologically meaningful semi-supervised learning tasks for evaluating protein representations.
- [The Cancer Genome Atlas (TCGA)](https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga) — Comprehensive multi-omics (genomics, transcriptomics, proteomics, methylation) dataset for 33 cancer types across ~11,000 patients.
- [TCGA virtual spatial transcriptomics atlas](https://huggingface.co/datasets/ratschlab/TCGA_virtual_spatial_transcriptomics_atlas) — DeepSpot-M predicted transcriptome-wide ST for TCGA H&E (FF + FFPE; 28,664 slides / 32 cancer types; gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).
- [HEST Xenium virtual spatial transcriptomics](https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics) — DeepSpot-M predicted transcriptome-wide ST for 59 HEST-1k 10x Xenium samples (~13.3M cells) (gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).
- [Therapeutics Data Commons (TDC)](https://tdcommons.ai/) — Unified benchmark suite covering ADMET, drug-target interaction, drug response, and more.
- [Tox21](https://tripod.nih.gov/tox21/challenge/) — 12,707 compounds tested in 12 nuclear receptor and stress-response pathway biochemical assays for toxicity prediction.
- [UK Biobank](https://www.ukbiobank.ac.uk/) — Large-scale biomedical database of ~500K participants with genetic, imaging, and health data for population genetics and disease studies.
---
## API
- [PubMed E-utilities (esearch/efetch)](https://www.nlm.nih.gov/dataguide/edirect/esearch.html) — APIs for searching and retrieving biomedical literature from PubMed.
- [NCBI E-utilities](https://www.ncbi.nlm.nih.gov/books/NBK25501/) — Unified APIs for accessing NCBI databases (Gene, GEO, SRA, PubChem, etc).
- [UniProt REST API](https://www.uniprot.org/help/api) — Programmatic access to protein sequence and functional annotation data.
- [Ensembl REST API](https://rest.ensembl.org/) — API for genomic annotations, variants, genes, and comparative genomics.
- [KEGG REST API](https://www.kegg.jp/kegg/rest/keggapi.html) — API for accessing KEGG pathways, compounds, genes, and reactions.
- [ChEMBL Web Services](https://www.ebi.ac.uk/chembl/ws) — REST API for bioactive molecules, targets, and bioassays.
- [Open Targets Platform API](https://platform.opentargets.org/api) — API for targetdisease associations integrating genetics, genomics, and drug data.
- [ClinicalTrials.gov API](https://clinicaltrials.gov/api/gui) — API for querying clinical trial metadata and results.
---
## Preprocessing Tools
- [Chemistry Development Kit](https://github.com/cdk/cdk) — Cheminformatics software & machine learning tools.
- [Biopython](https://biopython.org/) — Collection of Python tools for biological computation including sequence analysis, structure parsing, and database access.
- [FlashDeconv](https://github.com/cafferychen777/flashdeconv) — High-performance spatial transcriptomics deconvolution (~1M spots in ~3 min).
- [RDKit](https://github.com/rdkit/rdkit) — Cheminformatics software & machine learning toolkit.
- [DeepChem](https://github.com/deepchem/deepchem) — Deep learning library for drug discovery, quantum chemistry, and materials science.
- [ChatSpatial](https://github.com/cafferychen777/ChatSpatial) — MCP server for spatial transcriptomics analysis via natural language.
- [Scanpy](https://scanpy.readthedocs.io/en/stable/) — Python library for scRNA-seq analysis.
- [Seurat](https://satijalab.org/seurat/) — R library for scRNA-seq analysis.
- [scvi-tools](https://scvi-tools.org/) — Probabilistic models for single-cell omics data analysis.
- [CellTypist](https://github.com/Teichlab/celltypist) — Automated cell type annotation for scRNA-seq.
- [Squidpy](https://squidpy.readthedocs.io/) — Python library for spatial single-cell analysis.
- [GROMACS](https://www.gromacs.org/) — Molecular dynamics simulation package for biochemical molecules.
- [MDAnalysis](https://www.mdanalysis.org/) — Python library for analyzing and altering molecular dynamics simulation trajectories.
- [OpenMM](https://openmm.org/) — High-performance toolkit for molecular simulation and GPU-accelerated MD.
- [scVelo](https://github.com/theislab/scvelo) — RNA velocity estimation for single-cell transcriptomics, inferring the direction and speed of cell differentiation.
- [STAR](https://github.com/alexdobin/STAR) — Ultrafast universal RNA-seq aligner with support for spliced alignment and single-cell quantification via STARsolo.
- [kallisto](https://pachterlab.github.io/kallisto/) — Near-optimal RNA-seq quantification using pseudoalignment for fast transcript abundance estimation.
- [Harmony](https://github.com/immunogenomics/harmony) — Fast and scalable integration of single-cell data across datasets, conditions, technologies, and species.
- [Monocle3](https://cole-trapnell-lab.github.io/monocle3/) — Single-cell trajectory analysis tool for learning developmental trajectories and ordering cells in pseudotime.
- [CellChat](https://github.com/sqjin/CellChat) — Inference and analysis of cell-cell communication ligand-receptor networks from single-cell transcriptomics data.
- [SCENIC](https://github.com/aertslab/SCENIC) — Single-cell regulatory network inference and clustering linking transcription factors to co-expressed gene modules.
- [DoubletFinder](https://github.com/chris-mcginnis-ucsf/DoubletFinder) — Machine learning approach for detecting multiplet (doublet) artifacts in single-cell RNA-seq data.
- [Numbat](https://github.com/kharchenkolab/numbat) — Haplotype-aware copy number variation inference from single-cell RNA-seq using hidden Markov models.
- [CaSpER](https://github.com/akdess/CaSpER) — CNV identification and visualization by integrative analysis of single-cell or bulk RNA-seq data.
- [CellCharter](https://github.com/CSOgroup/cellcharter) — Identification and characterization of spatial cell niches from spatial transcriptomics using VAEs and Gaussian mixture models.
- [STAGATE](https://github.com/RucDongLab/STAGATE) — Adaptive graph attention auto-encoder for spatial domain identification in spatial transcriptomics.
- [NCEM](https://github.com/theislab/ncem) — GNN-based model for learning intercellular communication from spatial graphs of cells.
- [DeepTalk](https://github.com/JiangBioLab/DeepTalk) — Graph attention network for deciphering cell-cell communication from spatial transcriptomics data.
- [COMMOT](https://github.com/zcang/COMMOT) — Optimal transport-based framework for screening cell-cell communication in spatial transcriptomics.
- [TIGON](https://github.com/yutongo/TIGON) — Neural optimal transport method for reconstructing growth and dynamic trajectories from single-cell transcriptomics.
- [LINGER](https://github.com/Durenlab/LINGER) — Neural network for gene regulatory network inference from single-cell multiome (RNA+ATAC-seq) data with bulk data pretraining.
- [sciPENN](https://github.com/jlakkis/sciPENN) — RNN-based method for simultaneous protein expression prediction, uncertainty estimation, and cell-type label transfer from CITE-seq and scRNA-seq data.
- [MOGONET](https://github.com/txWang/MOGONET) — Multi-omics graph convolutional network framework for patient classification and biomarker identification.
- [AutoZyme](https://github.com/ElliotXie/autozyme) — Autonomous agentic framework that speeds up bioinformatics software (e.g. Scanpy, Seurat) on CPUs while preserving the original results.
---
## Machine Learning Tasks and Models
### Drug Discovery
#### Drug Response Prediction
- [drGAT](https://github.com/inoue0426/drGAT) — Attention-based model for drug response prediction with gene explainability.
- [MOFGCN](https://github.com/weiba/MOFGCN/tree/main) — GCN + heterogeneous network.
- [DeepDSC](https://ieeexplore-ieee-org.ezp2.lib.umn.edu/stamp/stamp.jsp?tp=&arnumber=8723620&tag=1) — Autoencoder + fully connected NN.
- [DGDRP](https://github.com/minwoopak/heteronet) — Multi-view embedding neural network.
- [DeepAEG](https://github.com/zhejiangzhuque/DeepAEG) — GNN embedding + attention mechanism.
- [RECOVER](https://github.com/RECOVERcoalition/Recover) — Machine learning framework for predicting synergistic drug combination responses across cell lines.
- [TGSA](https://github.com/violet-sto/TGSA) — Tumor gene set and attention-based model leveraging biological pathway knowledge for drug response prediction.
- [HiDRA](https://github.com/bsml320/HiDRA) — Hierarchical network model incorporating gene and pathway-level information for cancer drug response prediction.
- [PRNet](https://github.com/Perturbation-Response-Prediction/PRnet) — Deep generative model for predicting transcriptional responses to novel chemical perturbations for drug discovery.
- [chemCPA](https://github.com/theislab/chemCPA) — Compositional perturbation autoencoder for predicting single-cell transcriptional responses to unseen drug perturbations and dose combinations.
- [cycleCDR](https://github.com/hliulab/cycleCDR) — Interpretable cycle-consistency framework for modeling cellular responses to drug perturbations.
- [DRUML](https://github.com/CutillasLab/DRUMLR) — Ensemble machine learning framework combining standard ML with deep learning to systematically rank anti-cancer drugs from proteomics and RNA-seq data.
#### Drug Repurposing
- [DeepPurpose](https://github.com/kexinhuang12345/DeepPurpose) — Deep learning library for drug repurposing.
- [TranSiGen](https://github.com/myzhengSIMM/TranSiGen) — Dual-VAE architecture for ligand-based virtual screening, drug response prediction, and drug repurposing using chemical-induced transcriptional profiles.
#### Drug Target Interaction
- [NeoDTI](https://github.com/FangpingWan/NeoDTI) — Library for drug-target interaction prediction.
- [DTINet](https://github.com/luoyunan/DTINet) — Network-based framework integrating heterogeneous biological data for DTI prediction.
- [DeepDTA](https://github.com/hkmztrk/DeepDTA) — Deep learning model using CNNs on protein sequences and drug SMILES.
- [GraphDTA](https://github.com/thinng/GraphDTA) — Graph neural networkbased DTI prediction using molecular graphs.
- [MolTrans](https://github.com/kexinhuang12345/MolTrans) — Transformer-based DTI model leveraging molecular substructures.
- [DrugBAN](https://github.com/peizhenbai/DrugBAN) — Bilinear attention network for interpretable DTI prediction.
#### Compound-Protein Interaction
- [MCPINN](https://github.com/mhlee0903/multi_channels_PINN) — Drug discovery via compound-protein interaction and machine learning.
- [TransformerCPI](https://github.com/lifanchen-simm/transformerCPI) — CPI prediction using Transformer.
#### Molecular Generation
- [REINVENT](https://github.com/MolecularAI/Reinvent) — Reinforcement learning for de novo drug design.
- [MolGPT](https://github.com/devalab/molgpt) — Transformer-based model for molecular generation.
- [Molecular Transformer](https://github.com/pschwllr/MolecularTransformer) — Sequence-to-sequence model for retrosynthesis prediction.
- [Matcha](https://github.com/LigandPro/Matcha) — Multi-stage Riemannian flow matching model for physically valid molecular docking with scoring, pose filtering, and benchmarks.
- [TargetDiff](https://github.com/guanjq/targetdiff) — 3D equivariant diffusion model for structure-based drug design.
- [DiffDock](https://github.com/gcorso/DiffDock) — Diffusion generative model for molecular docking, predicting the binding pose of small molecules to protein targets.
- [JTVAE](https://github.com/wengong-jin/icml18-jtnn) — Junction tree variational autoencoder for molecular graph generation that guarantees chemical validity via a hierarchical tree decomposition.
- [DiffSBDD](https://github.com/arneschneuing/DiffSBDD) — Equivariant diffusion model for structure-based drug design that generates molecules and binding conformations for protein targets.
- [ReLeaSE](https://github.com/isayev/ReLeaSE) — Deep reinforcement learning framework for de novo drug design combining a generative and predictive model.
- [PaccMannRL](https://github.com/PaccMann/paccmann_generator) — Reinforcement learning-based generative model for de novo hit-like anticancer molecule design from transcriptomic data.
### LLM for Biology
- [AI4Chem/ChemLLM-7B-Chat](https://huggingface.co/AI4Chem/ChemLLM-7B-Chat) — LLM for chemical & molecular science.
- [BioGPT](https://github.com/microsoft/BioGPT) — LLM for biomedical text generation.
- [GeneGPT](https://github.com/ncbi/GeneGPT) — LLM for biomedical information, integrated with various APIs.
- [GenePT](https://github.com/yiqunchen/GenePT) — Foundation LLM for single-cell data.
- [scPRINT](https://github.com/cantinilab/scPRINT) — Pretrained on 50M cells for scRNA-seq denoising & zero imputation.
- [ClawBio](https://github.com/ClawBio/ClawBio) — Bioinformatics-native AI agent skill library with local-first pharmacogenomics, ancestry PCA, semantic similarity, nutrigenomics, and metagenomics skills.
- [BioMedLM](https://huggingface.co/stanford-crfm/BioMedLM) — 2.7B parameter GPT-2-style language model trained exclusively on biomedical literature from PubMed for biomedical question answering and text generation.
- [MolT5](https://github.com/blender-nlp/MolT5) — Language model for molecular tasks bridging text and SMILES, enabling molecule captioning and text-driven molecule generation.
- [ChatDrug](https://github.com/chao1224/ChatDrug) — LLM-based conversational pipeline for drug discovery, using natural language prompts for iterative drug editing and optimization.
- [CASSIA](https://github.com/ElliotXie/CASSIA) — Multi-agent LLM for reference-free, interpretable cell-type annotation of single-cell RNA-seq data, with dedicated annotation, validation, scoring, and reporting agents.
### Foundation Models
#### Single-cell Foundation Models
##### Transcriptomics Foundation Models
- [scFoundation](https://github.com/biomap-research/scFoundation) — Large-scale foundation model for single-cell gene expression, enabling multiple downstream tasks.
- [scGPT](https://github.com/bowang-lab/scGPT) — Transformer-based foundation model pretrained on millions of single-cell profiles.
- [Geneformer](https://huggingface.co/ctheodoris/Geneformer) — Context-aware, attention-based deep learning model pretrained on a large corpus of single-cell transcriptomes.
- [BulkFormer](https://github.com/KangBoming/BulkFormer) — Foundation model for bulk RNA-seq data; learns general transcriptomic representations.
- [scBERT](https://github.com/TencentAILabHealthcare/scBERT) — BERT-based foundation model pretrained on large-scale scRNA-seq data for cell type annotation.
- [CellPLM](https://github.com/OmicsML/CellPLM) — Cell pre-trained language model with inter-cell transformer architecture for diverse single-cell analysis tasks.
- [UCE](https://github.com/snap-stanford/UCE) — Universal Cell Embeddings: zero-shot single-cell embedding model trained on 36M cells across species, tissues, and assays without fine-tuning.
- [GEARS](https://github.com/snap-stanford/GEARS) — Graph-based model for predicting transcriptional responses to single and combinatorial genetic perturbations using biological priors.
- [SATURN](https://github.com/snap-stanford/SATURN) — Transformer-based model integrating gene expression and protein sequences via a protein language model to learn unified multi-species cell embeddings.
- [CancerFoundation](https://github.com/BoevaLab/CancerFoundation) — Single-cell RNA-seq foundation model trained exclusively on a curated dataset of malignant cells to learn cancer-specific embeddings.
##### Spatial Foundation Models
- [GigaPath](https://github.com/prov-gigapath/prov-gigapath) — Slide-level digital pathology foundation model pretrained on 1.3 billion pathology image tokens from whole-slide images.
- [UNI](https://github.com/mahmoodlab/UNI) — General-purpose self-supervised pathology foundation model trained on 100K+ whole-slide images for diverse computational pathology tasks.
- [CONCH](https://github.com/mahmoodlab/CONCH) — Vision-language foundation model for computational pathology trained with contrastive captioning on pathology imagetext pairs.
- [Phikon](https://huggingface.co/owkin/phikon) — ViT-based pathology foundation model pretrained with iBOT self-supervision on TCGA whole-slide images.
- [Nicheformer](https://github.com/theislab/nicheformer) — Foundation model for single-cell and spatial omics using a transformer architecture with positional embeddings to encode spatial cell information.
- [scGPT-spatial](https://github.com/bowang-lab/scGPT-spatial) — Extension of scGPT for spatial transcriptomics with continual pretraining and a mixture-of-experts decoder for spatial gene expression analysis.
- [DeepSpot](https://github.com/ratschlab/DeepSpot) — Deep learning model predicting spatial transcriptomics from H&E images at spot and single-cell resolution.
- [DeepSpot2Cell](https://github.com/ratschlab/DeepSpot2Cell) — Predicts virtual single-cell spatial transcriptomics from H&E using spot-level supervision (NeurIPS 2025 Imageomics).
- [DeepSpot-M](https://github.com/ratschlab/DeepSpotM) — Multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology.
- [AESTETIK](https://github.com/ratschlab/aestetik) — Autoencoder for spatial transcriptomics representation learning using topology and histology image knowledge.
##### Multi-Omics Foundation Models
- [scMulan](https://github.com/SuperBianC/scMulan) — Single-cell multi-omic language model pretrained on ~10M cells spanning transcriptomics, epigenomics, and proteomics for cross-omics transfer tasks.
- [totalVI](https://github.com/scverse/scvi-tools) — Probabilistic framework for joint analysis of paired scRNA-seq and protein (CITE-seq) data enabling multi-modal cell state representation across single-cell datasets.
- [MultiVI](https://github.com/scverse/scvi-tools) — Multi-modal variational autoencoder for integrating paired and unpaired single-cell RNA-seq and ATAC-seq measurements into a unified latent space.
- [MIRA](https://github.com/cistrome/MIRA) — Probabilistic multimodal topic model jointly modeling single-cell transcriptomics and chromatin accessibility for regulatory network inference.
- [GLUE](https://github.com/gao-lab/GLUE) — Graph-Linked Unified Embedding framework for unpaired single-cell multi-omics data integration across RNA, ATAC, methylation, and protein modalities.
- [BABEL](https://github.com/wukevin/babel) — Cross-modality translation model enabling prediction between scRNA-seq and scATAC-seq profiles without requiring paired single-cell measurements.
- [Multigrate](https://github.com/theislab/multigrate) — Asymmetric multi-omics variational autoencoder for integrating single-cell data across RNA, ATAC, and protein modalities with missing-modality support.
- [MOFA+](https://github.com/bioFAM/MOFA2) — Multi-Omics Factor Analysis framework identifying shared axes of variation across bulk and single-cell datasets including RNA, ATAC, proteomics, methylation, and copy number.
- [GeneCompass](https://github.com/xCompass-AI/GeneCompass) — Large-scale foundation model integrating DNA regulatory sequences and single-cell transcriptomics from 120M+ cells across multiple species for gene regulation prediction.
- [UnitedNet](https://github.com/LiuLab-Bioelectronics-Harvard/UnitedNet) — Interpretable multi-task deep neural network for single-cell multi-omics integration spanning transcriptomics, chromatin accessibility, and proteomics.
- [SpatialGlue](https://github.com/zhanglabtools/SpatialGlue) — Graph attention network for spatial multi-omics integration jointly embedding spatial transcriptomics with chromatin accessibility or proteomics.
- [MIDAS](https://github.com/labomics/midas) — Mosaic integration and differential accessibility model for single-cell multi-omics data that handles arbitrary missing-modality combinations across transcriptomics, chromatin accessibility, and proteomics.
- [Concerto](https://github.com/melobio/Concerto-reproducibility) — Contrastive self-supervised learning framework for single-cell multimodal data integration, batch correction, and reference-query mapping.
- [scButterfly](https://github.com/BioX-NKU/scButterfly) — Dual-aligned variational autoencoder for single-cell cross-modality translation between paired and unpaired multiomics data.
- [JAMIE](https://github.com/Oafish1/JAMIE) — Joint variational autoencoder for multimodal single-cell data imputation and embedding.
- [scPair](https://github.com/quon-titative-biology/scPair) — Bidirectional feedforward network for single-cell multimodal analysis with cross-modality prediction leveraging single-cell atlases.
##### Domain Alignment
- [scArches](https://github.com/theislab/scarches) — Transfer learning framework for mapping new single-cell datasets onto pre-trained reference atlases across batches, conditions, and modalities.
- [TOSICA](https://github.com/JackieHanlaopo/TOSICA) — Transformer-based framework for one-stop interpretable cell-type annotation supporting cross-dataset and cross-species transfer.
#### Compound Foundation Models
##### Compound Embedding
- [ChemBERTa-2](https://github.com/seyonechithrananda/bert-loves-chemistry) — RoBERTa-based molecular language model pretrained on SMILES for small-molecule representation learning.
- [GROVER](https://github.com/tencent-ailab/grover) — Self-supervised graph transformer for large-scale molecular representation learning from unlabeled compounds.
- [Mol2Vec](https://github.com/samoturk/mol2vec) — Unsupervised molecular embedding method inspired by Word2Vec for learning vector representations of chemical substructures.
- [MolFormer](https://github.com/IBM/molformer) — Linear attention transformer pretrained on millions of SMILES strings for efficient molecular embeddings.
- [Uni-Mol](https://github.com/deepmodeling/Uni-Mol) — 3D molecular pretraining framework for universal representation learning on molecules and protein pockets.
#### Protein Foundation Models
##### Pre-trained Embedding
- [Evolutionary Scale Modeling (ESM)](https://github.com/facebookresearch/esm) — Protein embeddings.
- [ProtTrans](https://github.com/agemagician/ProtTrans) — Suite of protein language models (ProtBERT, ProtT5, ProtXLNet) trained on billions of protein sequences from UniRef and BFD.
- [ProGen2](https://github.com/salesforce/progen) — Protein language model trained on diverse protein families for sequence generation and fitness prediction.
- [Ankh](https://github.com/agemagician/Ankh) — Efficient protein language model optimized for downstream prediction tasks including secondary structure, localization, and function annotation.
##### Protein Structure Prediction and Design
- [AlphaFold3](https://github.com/google-deepmind/alphafold3) — Predicts structures of proteins, nucleic acids, small molecules, and their complexes.
- [Boltz-1](https://github.com/jwohlwend/boltz) — Open-source all-atom biomolecular structure prediction model for proteins, nucleic acids, small molecules, and their complexes achieving AlphaFold3-level accuracy.
- [Chai-1](https://github.com/chaidiscovery/chai-lab) — Unified molecular structure prediction model covering proteins, nucleic acids, small molecules, and complexes.
- [ESM3](https://github.com/evolutionaryscale/esm) — Multimodal protein language model that jointly reasons over sequence, structure, and function for generative protein design and engineering.
- [ESMFold](https://github.com/facebookresearch/esm) — Fast protein structure prediction using language model embeddings.
- [RFdiffusion](https://github.com/RosettaCommons/RFdiffusion) — Generative model for protein backbone design using diffusion.
- [ProteinMPNN](https://github.com/dauparas/ProteinMPNN) — Deep learning model for protein sequence design given backbone structure.
- [OmegaFold](https://github.com/HeliXonProtein/OmegaFold) — High-resolution de novo protein structure prediction from sequence.
- [RoseTTAFold](https://github.com/RosettaCommons/RoseTTAFold) — Three-track neural network for protein structure prediction.
- [OpenFold](https://github.com/aqlaboratory/openfold) — Trainable, memory-efficient open-source reproduction of AlphaFold2 enabling custom protein structure prediction workflows.
- [SaProt](https://github.com/westlake-reup/SaProt) — Structure-aware protein language model using structure-aware tokens that encode both sequence and backbone geometry for improved function prediction.
- [EvoDiff](https://github.com/microsoft/evodiff) — Discrete diffusion framework for protein sequence generation trained on evolutionary-scale data, supporting unconditional generation, disordered region design, and functional motif scaffolding. [ [paper-2023](https://www.biorxiv.org/content/10.1101/2023.09.11.556673v1) ]
#### Multi-Modal Foundation Models
- [CHIEF](https://github.com/hms-dbmi/CHIEF) — Clinical Histopathology Imaging Evaluation Foundation model integrating histology images and clinical context for pan-cancer analysis.
- [BiomedCLIP](https://huggingface.co/microsoft/BiomedCLIP-PubMedBERT_256-vit_g_14) — CLIP-based vision-language foundation model for biomedical images and text trained on PubMed figurecaption pairs.
- [PORPOISE](https://github.com/mahmoodlab/PORPOISE) — Pan-cancer integrative histology-genomic analysis framework using multimodal deep learning for patient stratification.
- [PathomicFusion](https://github.com/mahmoodlab/PathomicFusion) — Integrated framework fusing histopathology and genomic features via CNN, GNN, and attention gating for cancer diagnosis and prognosis.
- [Virchow](https://huggingface.co/paige-ai/Virchow) — Million-slide digital pathology foundation model using a vision transformer and self-supervised distillation for tile-level pathology image representation.
- [TOAD](https://github.com/mahmoodlab/TOAD) — Tumor Origin Assessment via Deep-learning; weakly-supervised multi-task model predicting cancer primary origin from H&E whole-slide images.
- [PLIP](https://github.com/PathologyFoundation/plip) — Vision-language foundation model for pathology trained with contrastive learning on pathology imagetext pairs for image classification and text-to-image retrieval.
- [MUSK](https://github.com/lilab-stanford/MUSK) — Vision-language foundation model for precision oncology analyzing multimodal paired text and pathology image data for biomarker prediction and retrieval.
#### Genomics Foundation Models
- [Nucleotide Transformer](https://github.com/instadeepai/nucleotide-transformer) — Foundation model for genomic sequences across multiple species.
- [DNABERT](https://github.com/jerryji1993/DNABERT) — Pre-trained bidirectional encoder for DNA sequence analysis.
- [DNABERT-2](https://github.com/Zhihan1996/DNABERT_2) — Improved genome foundation model with efficient tokenization.
- [Enformer](https://github.com/deepmind/deepmind-research/tree/master/enformer) — Transformer model predicting gene expression from DNA sequence.
- [Basenji](https://github.com/calico/basenji) — Sequential regulatory activity prediction from DNA sequences.
- [Caduceus](https://github.com/kuleshov-group/caduceus) — Bidirectional equivariant long-range DNA sequence model based on Mamba.
- [Evo](https://github.com/evo-design/evo) — Long-context genomic foundation model (up to 1M tokens).
- [HyenaDNA](https://github.com/HazyResearch/hyena-dna) — Long-range genomic foundation model handling sequences up to 1M tokens with sub-quadratic attention.
- [Borzoi](https://github.com/calico/borzoi) — Extended successor to Enformer for predicting RNA-seq coverage from long genomic sequence windows (524 kb) with improved resolution.
- [DeepSEA](http://deepsea.princeton.edu/) — Deep learning framework for predicting chromatin effects of sequence alterations with single-nucleotide sensitivity across thousands of chromatin features.
- [Sei](https://github.com/FunctionLab/sei-framework) — Sequence-to-function framework learning a genome-wide regulatory activity code from DNA sequences for variant effect prediction.
- [GPN (Genomic Pre-trained Network)](https://github.com/songlab-cal/gpn) — Masked language model for DNA sequences enabling zero-shot variant effect prediction without requiring functional annotations.
---
## Citation
If you use this list in papers, slides, or documentation, please cite this repository via [`CITATION.cff`](./CITATION.cff) (also available through GitHub's **Cite this repository** button).
## Curation Criteria (Strict)
To keep quality high, additions should meet all of the following:
- The resource is trustworthy and relevant to computational biology.
- The primary link points to an official source (official docs, organization site, maintained repository, or official dataset page).
- The resource has evidence of technical substance: ideally a peer-reviewed paper; at minimum a preprint or official technical documentation.
- The description is factual and concise (no marketing copy).
- Duplicate or near-duplicate entries should be avoided.
We generally do **not** accept entries that are only promotional pages, personal opinion posts, or generic blog posts without technical references.
## Update & Link Rot Policy
- Link validity is monitored by the [Link Check workflow](./.github/workflows/link-check.yml).
- If a link repeatedly fails, maintainers may replace it with an official mirror/canonical URL or remove the entry until a stable URL is available.
- Contributions fixing broken links are welcome and encouraged.
## Data Schema & Contribution Workflow
- Data schema reference: [`docs/data/SCHEMA.md`](./docs/data/SCHEMA.md).
- Source-of-truth workflow:
1. Edit/add resources in `README.md`.
2. Regenerate machine-readable artifacts:
- `python scripts/sync_resources_from_readme.py`
- `python scripts/build_resources.py`
3. Commit updated data files (`data/resources.yml`, `data/resources.json`, `data/resources.csv`, `docs/data/resources.json`) with your README change.
- Contribution guide: [`contributing.md`](./contributing.md).
@@ -0,0 +1,87 @@
---
title: "Contributor Covenant Code of Conduct"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/code-of-conduct.md
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
# Contributor Covenant Code of Conduct
## Our Pledge
In the interest of fostering an open and welcoming environment, we as
contributors and maintainers pledge to making participation in our project and
our community a harassment-free experience for everyone, regardless of age, body
size, disability, ethnicity, gender identity and expression, level of experience,
nationality, personal appearance, race, religion, or sexual identity and
orientation.
## Our Standards
Examples of behavior that contributes to creating a positive environment
include:
* Using welcoming and inclusive language
* Being respectful of differing viewpoints and experiences
* Gracefully accepting constructive criticism
* Focusing on what is best for the community
* Showing empathy towards other community members
Examples of unacceptable behavior by participants include:
* The use of sexualized language or imagery and unwelcome sexual attention or
advances
* Trolling, insulting/derogatory comments, and personal or political attacks
* Public or private harassment
* Publishing others' private information, such as a physical or electronic
address, without explicit permission
* Other conduct which could reasonably be considered inappropriate in a
professional setting
## Our Responsibilities
Project maintainers are responsible for clarifying the standards of acceptable
behavior and are expected to take appropriate and fair corrective action in
response to any instances of unacceptable behavior.
Project maintainers have the right and responsibility to remove, edit, or
reject comments, commits, code, wiki edits, issues, and other contributions
that are not aligned to this Code of Conduct, or to ban temporarily or
permanently any contributor for other behaviors that they deem inappropriate,
threatening, offensive, or harmful.
## Scope
This Code of Conduct applies both within project spaces and in public spaces
when an individual is representing the project or its community. Examples of
representing a project or community include using an official project e-mail
address, posting via an official social media account, or acting as an appointed
representative at an online or offline event. Representation of a project may be
further defined and clarified by project maintainers.
## Enforcement
Instances of abusive, harassing, or otherwise unacceptable behavior may be
reported by contacting the project team at inoue019@umn.edu. All
complaints will be reviewed and investigated and will result in a response that
is deemed necessary and appropriate to the circumstances. The project team is
obligated to maintain confidentiality with regard to the reporter of an incident.
Further details of specific enforcement policies may be posted separately.
Project maintainers who do not follow or enforce the Code of Conduct in good
faith may face temporary or permanent repercussions as determined by other
members of the project's leadership.
## Attribution
This Code of Conduct is adapted from the [Contributor Covenant][homepage], version 1.4,
available at [http://contributor-covenant.org/version/1/4][version]
[homepage]: http://contributor-covenant.org
[version]: http://contributor-covenant.org/version/1/4/
@@ -0,0 +1,77 @@
---
title: "Contribution Guidelines"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/contributing.md
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Contribution Guidelines
Contributions are welcome!
Please note that this project is released with a
[Contributor Code of Conduct](code-of-conduct.md). By participating in this
project you agree to abide by its terms.
## Pull Requests
- Search previous suggestions before making a new one, as yours may be a duplicate.
- Add one link per pull request.
- Prefer official and trustworthy sources (official docs, organization pages, maintained repositories, or official dataset pages).
- Include supporting technical evidence for new resources:
- Ideally a peer-reviewed publication.
- At minimum, a preprint or official technical documentation.
- Avoid submissions that are primarily promotional pages, generic blog posts, or opinion-only writeups.
- Add the link:
- `[name](http://example.com/)` - A short description ends with a period.
- Keep descriptions concise.
- Maintain alphabetical ordering where applicable.
- Add a section if needed.
- Add the section description.
- Add the section title to the [Index](https://github.com/inoue0426/awesome-computational-biology#Contents).
- Check your spelling and grammar.
- Remove any trailing whitespace.
- Send a pull request with the reason why the addition is awesome.
- Use the following format for your pull request title:
- Add user/repo - Short repo description
## Data Workflow (README and JSON)
- The curated source list is maintained in `README.md`.
- Machine-readable files are generated from README:
- `python scripts/sync_resources_from_readme.py`
- `python scripts/build_resources.py`
- For resource additions/edits, include updated generated files in the same PR:
- `data/resources.yml`
- `data/resources.json`
- `data/resources.csv`
- `docs/data/resources.json`
- Field definitions and naming rules are documented in [`docs/data/SCHEMA.md`](docs/data/SCHEMA.md).
## GitHub Pages UI
- The UI reads `docs/data/resources.json`.
- Search and filters are driven by these fields:
- Search: `name`, `description`, `tasks`, `modalities`, `tags`
- Filters: `type`, `tasks`, `modalities`
## Updates to Existing Links or Sections
- Improvements to the existing sections are welcome.
- If you think a listed link is not awesome, feel free to submit an issue or pull request to begin the discussion.
- Broken links are checked by CI; if you find one, please submit a fix to the canonical URL (or remove the entry if no stable canonical URL exists).
## Updating your PR
A lot of times, making a PR adhere to the standards above can be difficult.
If the maintainers notice anything that we'd like changed, we'll ask you to
edit your PR before we merge it. There's no need to open a new PR, just edit
the existing one. If you're not sure how to do that,
[here is a guide](https://github.com/RichardLitt/knowledge/blob/master/github/amending-a-commit-guide.md)
on the different ways you can update your PR so that we can merge it.
@@ -0,0 +1,173 @@
---
title: "Cspell"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/0cf037ab/cspell.json
upstream_sha: 0cf037ab
imported_at: 2026-07-16
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
{
"language": "en",
"allowCompoundWords": true,
"words": [
"behavioural",
"KEGG",
"NCBI",
"UCSC",
"EMBL",
"RCSB",
"CASP",
"Uniclust",
"Reactome",
"Bioactive",
"biopolymers",
"proteomics",
"transcriptomics",
"metagenomic",
"metatranscriptomic",
"CRISPR",
"JASPAR",
"druggable",
"Toxicogenomics",
"GDSC",
"biomolecular",
"DRKG",
"Hetionet",
"Eudra",
"esearch",
"efetch",
"Ensembl",
"Cheminformatics",
"Deconv",
"Scanpy",
"Squidpy",
"explainability",
"MOFGCN",
"Autoencoder",
"DGDRP",
"MCPINN",
"Pretrained",
"pretrained",
"denoising",
"transcriptomic",
"CELLxGENE",
"eukaryotic",
"metabolites",
"OMIM",
"Mendelian",
"DisGeNET",
"GWAS",
"IntAct",
"Biopython",
"MDAnalysis",
"trajectories",
"Geneformer",
"equivariant",
"HyenaDNA",
"Hyena",
"Caduceus",
"Mamba",
"retrosynthesis",
"TargetDiff",
"Chai",
"Zuckerberg",
"HMDB",
"CTRP",
"ADMET",
"Omics",
"omics",
"omic",
"OADB",
"Gnify",
"gnom",
"Rfam",
"Guaca",
"deconvolution",
"scvi",
"pharmacogenomics",
"nutrigenomics",
"Giga",
"Phikon",
"TCGA",
"Mulan",
"epigenomics",
"methylation",
"ATAC",
"MOFA",
"TOSICA",
"Boltz",
"MPNN",
"Enformer",
"Velo",
"BACE",
"secretase",
"Clin",
"CPTAC",
"Proteomic",
"proteogenomic",
"LINCS",
"ogbl",
"ogbg",
"SIDER",
"Muris",
"Pfam",
"PROSITE",
"epigenome",
"TRRUST",
"kallisto",
"pseudoalignment",
"multiplet",
"TGSA",
"JTVAE",
"miRBase",
"miRNA",
"ProtTrans",
"ProtBERT",
"ProGen",
"Ankh",
"DeepSEA",
"RegNetwork",
"ROADMAP",
"FANTOM",
"NeXtProt",
"HiDRA",
"MolT",
"ChatDrug",
"DoubletFinder",
"pseudotime",
"ligand",
"SCENIC",
"GPN",
"Sei",
"KIBA",
"pharmacokinetics",
"ADME",
"Haplotype",
"NCEM",
"multiome",
"MOGONET",
"convolutional",
"DRUML",
"SBDD",
"Pacc",
"multiomics",
"Pathomic",
"PLIP",
"Omni",
"Bento",
"FFPE",
"Xenium",
"Zyme",
"Neur",
"Imageomics",
"AESTETIK"
],
"ignorePaths": [
"node_modules/**"
]
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,81 @@
---
title: "Resource Data Schema (`docs/data/resources.json`)"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/docs/data/SCHEMA.md
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Resource Data Schema (`docs/data/resources.json`)
This document describes the JSON schema used by the GitHub Pages UI.
## Source of truth and generation flow
- **Canonical list source:** `README.md` (curated resource bullets)
- Generated from README to YAML: `scripts/sync_resources_from_readme.py``data/resources.yml`
- Built artifacts from YAML: `scripts/build_resources.py``data/resources.json`, `data/resources.csv`, and `docs/data/resources.json`
When contributing new resources, update `README.md` first, then regenerate artifacts.
## Top-level structure
- `resources.json` is a JSON array.
- Each array item is one resource object.
## Fields
### Required fields
| Field | Type | Notes |
|---|---|---|
| `id` | string | Unique slug. Use lowercase `snake_case`, stable over time. |
| `name` | string | Display name shown in README/UI. |
| `type` | string | Resource category. Current values: `api`, `benchmark`, `database`, `model`, `toolkit`. |
| `url` | string | Canonical landing page URL. |
| `description` | string | One-line, factual summary. |
### Optional fields
| Field | Type | Notes |
|---|---|---|
| `tags` | array of strings | Free-form tags. |
| `tasks` | array of strings | Task labels used by Task filter. |
| `modalities` | array of strings | Data modality labels used by Modality filter. |
| `organism` | array of strings | Organism labels. |
| `license` | string | SPDX identifier preferred when known. |
| `api` | boolean | Whether programmatic API access is available. Defaults to `false`. |
| `paper` | string | DOI or URL to preprint/peer-reviewed publication. |
| `updated` | string | Last-known update date, recommended `YYYY-MM-DD`. |
## Naming and consistency guidance
- `id` must be globally unique across all resources.
- Prefer concise, stable IDs (e.g., `open_targets_platform`, `alphafold3`).
- Keep `name` aligned with official project/database naming.
- Use short, objective descriptions (avoid marketing language).
## Example object
```json
{
"id": "open_targets_platform",
"name": "Open Targets Platform",
"type": "database",
"url": "https://platform.opentargets.org/",
"description": "Target identification platform integrating genetics, genomics, and drug evidence.",
"tags": ["disease", "drug-discovery"],
"tasks": ["target-identification"],
"modalities": ["genomics"],
"organism": ["human"],
"license": "CC-BY-4.0",
"api": true,
"paper": "https://doi.org/10.1093/nar/gkac1045",
"updated": "2026-01-15"
}
```
@@ -0,0 +1,15 @@
---
title: "Requirements"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/scripts/requirements.txt
upstream_sha: 12d87583
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
PyYAML>=6.0
matplotlib>=3.7
@@ -0,0 +1,288 @@
---
title: "Awesome LLM Agents for Scientific Discovery [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)"
task: ""
lineage_type: import
upstream_source: https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery/blob/3e079cd8/README.md
upstream_sha: 3e079cd8
imported_at: 2026-06-26
prompt_class: catalogue
upstream_changes: accepted
author: upstream
validated: false
---
# Awesome LLM Agents for Scientific Discovery [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)
<div align="center">
<img src="agents4science.webp" alt="AI Agents for Scientific Discovery" width="600px">
</div>
A curated list of papers about AI agents for scientific discovery and research automation.
Maintained by [Jieli Zhou](mailto:[email protected])
If you use this paper list for your research, please cite it using:
```bibtex
@misc{zhou2024awesome,
title={Awesome AI Agents for Scientific Discovery},
author={Zhou, Jieli},
year={2024},
publisher={GitHub},
journal={GitHub repository},
howpublished={\url{https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery}}
}
```
## Introduction
The convergence of large language models (LLMs) and autonomous agents has ushered in a new era in scientific discovery, fundamentally transforming how research is conducted across disciplines. This emerging paradigm, articulated in Kitano's seminal "Nobel Turing Challenge" (2021), envisions AI systems capable of making scientific discoveries worthy of Nobel Prize recognition. Recent advances in LLM-based agents have brought us closer to this vision, enabling increasingly sophisticated automation of scientific workflows and decision-making processes.
### Evolution and Current Landscape
The field has evolved rapidly since early visions of AI-driven scientific discovery. While traditional AI systems focused on narrow tasks, modern LLM-based agents demonstrate remarkable capabilities in complex scientific reasoning, experimental design, and hypothesis generation. The breakthrough capabilities of models like GPT-4 have catalyzed this transition, enabling agents to engage in sophisticated scientific discourse, interpret complex data, and even design novel experiments.
### Key Research Directions
Several major research themes have emerged in this space:
1. **Multi-Agent Architectures**: Research has increasingly focused on collaborative multi-agent systems, where specialized agents work together to tackle complex scientific problems.
2. **Domain-Specific Applications**: The healthcare sector has seen particularly rapid adoption, with agents being developed for clinical decision support, medical diagnosis, and healthcare administration.
3. **Scientific Process Automation**: Agents are being developed to automate various aspects of the research pipeline, from literature review and hypothesis generation to experimental design and data analysis.
### Impact and Future Directions
The emergence of AI agents in scientific discovery represents more than just technological advancement; it signals a fundamental shift in how science is conducted. These systems promise to:
- Accelerate the pace of scientific discovery
- Enable exploration of previously intractable research questions
- Democratize access to scientific expertise
- Foster more efficient use of research resources
## Table of Contents
1. [Foundations & Vision](#foundations--vision)
2. [Core Technologies](#core-technologies)
3. [Scientific Process Automation](#scientific-process-automation)
4. [Domain Applications](#domain-applications)
5. [Infrastructure & Tools](#infrastructure--tools)
6. [Evaluation & Benchmarking](#evaluation--benchmarking)
7. [Surveys & Reviews](#surveys--reviews)
## Foundations & Vision
### Vision Papers
- **[Nobel Turing Challenge: Creating the Engine for Scientific Discovery](https://www.nature.com/articles/s41592-021-01091-w)**
*Hiroaki Kitano.* NPJ Systems Biology and Applications 2021
- **[Artificial Intelligence to Win the Nobel Prize and Beyond: Creating the Engine for Scientific Discovery](https://www.aaai.org/ojs/index.php/aimagazine/article/view/2624)**
*Hiroaki Kitano.* AI Magazine 2016
- **[The AI Scientist: Towards Fully Automated Open-ended Scientific Discovery](https://arxiv.org/abs/2408.06292)**
*Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, David Ha.* arXiv 2024
- **[Emergent autonomous scientific research capabilities of large language models](https://arxiv.org/abs/2304.05332)**
*Daniil A Boiko, Robert MacKnight, Gabe Gomes.* arXiv 2023
- **[What is missing in autonomous discovery: open challenges for the community](https://pubs.rsc.org/en/content/articlelanding/2023/dd/d3dd00089c)**
*Phillip M Maffettone, Pascal Friederich, Sterling G Baird, et al.* Digital Discovery 2023
- **[The future of fundamental science led by generative closed-loop artificial intelligence](https://arxiv.org/abs/2307.07522)**
*Hector Zenil, Jesper Tegnér, Felipe S Abrahão, Alexander Lavin, et al.* arXiv 2023
## Core Technologies
### Multi-Agent Systems & Architectures
- **[CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society](https://proceedings.neurips.cc/paper_files/paper/2023/hash/9a86e0c5-e09e-4ad7-96d6-b2ed61855e37-Abstract-Conference.html)**
*Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, Bernard Ghanem.* NeurIPS 2023
- **[Dynamic LLM-Agent Network: An LLM-Agent Collaboration Framework with Agent Team Optimization](https://arxiv.org/abs/2310.02170)**
*Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, Diyi Yang.* arXiv 2023
- **[AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Framework](https://arxiv.org/abs/2308.08155)**
*Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, et al.* arXiv 2023
### Reasoning & Knowledge Systems
- **[Graph of Thoughts: Solving Elaborate Problems with Large Language Models](https://ojs.aaai.org/index.php/AAAI/article/view/28877)**
*Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, et al.* AAAI 2024
- **[KnowAgent: Knowledge-augmented Planning for LLM-based Agents](https://arxiv.org/abs/2403.03101)**
*Yuqi Zhu, Shuofei Qiao, Yixin Ou, Shumin Deng, et al.* arXiv 2024
- **[Improving Factuality and Reasoning in Language Models through Multiagent Debate](https://arxiv.org/abs/2305.14325)**
*Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, Igor Mordatch.* arXiv 2023
## Scientific Process Automation
### Research Planning & Literature Review
- **[ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models](https://arxiv.org/abs/2404.07738)**
*Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, Sung Ju Hwang.* arXiv 2024
- **[SciMon: Scientific Inspiration Machines Optimized for Novelty](https://arxiv.org/abs/2305.14259)**
*Qingyun Wang, Doug Downey, Heng Ji, Tom Hope.* arXiv 2023
- **[AutoSurvey: Large Language Models Can Automatically Write Surveys](https://arxiv.org/abs/2406.10252)**
*Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, et al.* arXiv 2024
### Experimental Design & Workflow
- **[DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents](https://arxiv.org/abs/2406.06769)**
*Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, et al.* arXiv 2024
- **[Genesis: Towards the Automation of Systems Biology Research](https://arxiv.org/abs/2408.10689)**
*Ievgeniia A Tiukova, Daniel Brunnsåker, Erik Y Bjurström, Alexander H Gower, et al.* arXiv 2024
## Domain Applications
### Healthcare & Medicine
#### Clinical Decision Support & Diagnosis
- **[MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making](https://arxiv.org/abs/2411.00248)**
*Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, et al.* NeurIPS 2024
- **[Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis](https://arxiv.org/abs/2401.16107)**
*Haochun Wang, Sendong Zhao, Zewen Qiang, Nuwa Xi, et al.* arXiv 2024
- **[MedAide: Towards an Omni Medical Aide via Specialized LLM-based Multi-Agent Collaboration](https://arxiv.org/abs/2410.12532)**
*Jinjie Wei, Dingkang Yang, Yanshu Li, Qingyao Xu, et al.* arXiv 2024
- **[Large Language Models as Agents in the Clinic](https://arxiv.org/abs/2309.10895)**
*Nikita Mehandru, Brenda Y. Miao, Eduardo Rodriguez Almaraz, et al.* NPJ Digital Medicine 2024
- **[MAGDA: Multi-Agent Guideline-Driven Diagnostic Assistance](https://link.springer.com/chapter/10.1007/978-3-031-49673-3_15)**
*David Bani-Harouni, Nassir Navab, Matthias Keicher.* FMGMAI 2024
#### Healthcare Systems & Management
- **[ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent Collaboration](https://arxiv.org/abs/2410.02551)**
*Zixiang Wang, Yinghao Zhu, Huiya Zhao, Xiaochen Zheng, et al.* arXiv 2024
- **[Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents](https://arxiv.org/abs/2405.02957)**
*Junkai Li, Siyu Wang, Meng Zhang, Weitao Li, et al.* arXiv 2024
- **[ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World](https://arxiv.org/abs/2406.13890)**
*Weixiang Yan, Haitian Liu, Tengxiao Wu, Qian Chen, et al.* arXiv 2024
- **[AIPatient: Simulating Patients with EHRs and LLM Powered Agentic Workflow](https://arxiv.org/abs/2409.18924)**
*Huizi Yu, Jiayan Zhou, Lingyao Li, Shan Chen, et al.* arXiv 2024
#### Medical Education & Training
- **[Medco: Medical education copilots based on a multi-agent framework](https://arxiv.org/abs/2408.12496)**
*Hao Wei, Jianing Qiu, Haibao Yu, Wu Yuan.* arXiv 2024
- **[AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://arxiv.org/abs/2405.07960)**
*Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, et al.* arXiv 2024
#### Medical Imaging & Pathology
- **[CXR-Agent: Vision-language models for chest X-ray interpretation with uncertainty aware radiology reporting](https://arxiv.org/abs/2407.08811)**
*Naman Sharma.* arXiv 2024
- **[PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent Collaboration](https://arxiv.org/abs/2407.00203)**
*Yuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu, et al.* arXiv 2024
#### Medical Research
- **[OpenLens AI: Fully Autonomous Research Agent for Health Infomatics](https://arxiv.org/abs/2509.14778)**
*Yuxiao Cheng, Jinli Suo* arXiv 2025, [GitHub Repo](https://github.com/jarrycyx/openlens-ai)
### Biology & Life Sciences
#### Genomics & Molecular Biology
- **[BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments](https://arxiv.org/abs/2405.17631)**
*Yusuf Roohani, Andrew Lee, Qian Huang, Jian Vora, et al.* arXiv 2024
- **[GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases](https://arxiv.org/abs/2405.16205)**
*Zhizheng Wang, Qiao Jin, Chih-Hsuan Wei, Shubo Tian, et al.* arXiv 2024
- **[Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation](https://arxiv.org/abs/2407.08940)**
*Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, et al.* arXiv 2024
#### Bioinformatics Tools & Platforms
- **[BIA: BioInformatics Agent - Unleashing the Power of Large Language Models to Reshape Bioinformatics Workflow](https://www.biorxiv.org/content/10.1101/2024.05.22.595240v1)**
*Qi Xin, Quyu Kong, Hongyi Ji, Yue Shen, et al.* bioRxiv 2024
- **[CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis](https://www.biorxiv.org/content/10.1101/2024.05.13.593861v1)**
*Yihang Xiao, Jinyi Liu, Yan Zheng, Xiaohan Xie, et al.* bioRxiv 2024
- **[SeqMate: A Novel Large Language Model Pipeline for Automating RNA Sequencing](https://arxiv.org/abs/2407.03381)**
*Devam Mondal, Atharva Inamdar.* arXiv 2024
### Chemistry & Materials Science
#### Drug Discovery & Development
- **[DrugAgent: Explainable Drug Repurposing Agent with Large Language Model-based Reasoning](https://arxiv.org/abs/2408.13378)**
*Yoshitaka Inoue, Tianci Song, Tianfan Fu.* arXiv 2024
- **[Malade: Orchestration of LLM-powered agents with retrieval augmented generation for pharmacovigilance](https://arxiv.org/abs/2408.01869)**
*Jihye Choi, Nils Palumbo, Prasad Chalasani, Matthew M Engelhard, et al.* arXiv 2024
#### Molecular Modeling & Computation
- **[ChatMol Copilot: An Agent for Molecular Modeling and Computation Powered by LLMs](https://aclanthology.org/2024.lm-1.6/)**
*Jinyuan Sun, Auston Li, Yifan Deng, Jiabo Li.* L+M Workshop 2024
- **[A review of large language models and autonomous agents in chemistry](https://arxiv.org/abs/2407.01603)**
*Mayk Caldas Ramos, Christopher J Collison, Andrew D White.* arXiv 2024
### Earth & Environmental Sciences
- **[An LLM Agent for Automatic Geospatial Data Analysis](https://arxiv.org/abs/2410.18792)**
*Yuxing Chen, Weijie Wang, Sylvain Lobry, Camille Kurtz.* arXiv 2024
## Evaluation & Benchmarking
### General Benchmarks
- **[AgentBench: Evaluating LLMs as Agents](https://arxiv.org/abs/2308.03688)**
*Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, et al.* arXiv 2023
- **[ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate](https://arxiv.org/abs/2308.07201)**
*Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, et al.* arXiv 2023
- **[Benchmarking large language models as ai research agents](https://arxiv.org/abs/2311.12741)**
*Qian Huang, Jian Vora, Percy Liang, Jure Leskovec.* NeurIPS 2023 Workshop
### Domain-Specific Benchmarks
- **[BioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science](https://arxiv.org/abs/2407.00466)**
*Xinna Lin, Siqi Ma, Junjie Shan, Xiaojing Zhang, et al.* arXiv 2024
- **[GenoTEX: A Benchmark for Evaluating LLM-Based Exploration of Gene Expression Data](https://arxiv.org/abs/2406.15341)**
*Haoyang Liu, Haohan Wang.* arXiv 2024
- **[IdeaBench: Benchmarking Large Language Models for Research Idea Generation](https://arxiv.org/abs/2411.02429)**
*Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, et al.* arXiv 2024
## Surveys & Reviews
### Comprehensive Surveys
- **[Scientific discovery in the age of artificial intelligence](https://www.nature.com/articles/s41586-023-06221-2)**
*Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, et al.* Nature 2023
- **[The rise and potential of large language model based agents: A survey](https://arxiv.org/abs/2309.07864)**
*Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, et al.* arXiv 2023
- **[Large language model based multi-agents: A survey of progress and challenges](https://arxiv.org/abs/2402.01680)**
*Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, et al.* arXiv 2024
- **[A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges](https://link.springer.com/article/10.1007/s44223-024-00009-0)**
*Xinyi Li, Sai Wang, Siqi Zeng, Yu Wu, Yi Yang.* Vicinagearth 2024
### Domain-Specific Reviews
- **[AI for Biomedicine in the Era of Large Language Models](https://arxiv.org/abs/2403.15673)**
*Zhenyu Bi, Sajib Acharjee Dip, Daniel Hajialigol, Sindhura Kommu, et al.* arXiv 2024
- **[A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions](https://arxiv.org/abs/2406.03712)**
*Lei Liu, Xiaoyan Yang, Junchi Lei, Xiaoyang Liu, et al.* arXiv 2024
- **[From LLMs to LLM-based Agents for Software Engineering: A Survey](https://arxiv.org/abs/2408.02479)**
*Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, et al.* arXiv 2024
## Contributing
Please feel free to send a pull request if you want to:
- Add new papers
- Fix errors
- Update paper information
## License
[![CC0](https://licensebuttons.net/p/zero/1.0/88x31.png)](https://creativecommons.org/publicdomain/zero/1.0/)