Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
89c98be0cd | ||
|
|
0a25b5b792 | ||
|
|
27303f36c1 | ||
|
|
ea7701920e | ||
|
|
fcd5fbc473 | ||
|
|
1fdd6c59ed | ||
|
|
53c2e77288 | ||
|
|
718179d35a | ||
|
|
bbb4a28582 | ||
|
|
89ee6aff51 | ||
|
|
9d277f9d01 | ||
|
|
0c421f2ac1 | ||
|
|
93be653ade | ||
|
|
4aba0445d9 | ||
|
|
751894b8ec | ||
|
|
3331a1ace5 | ||
|
|
c4cc7c574e | ||
|
|
396d32d996 | ||
|
|
6752ebfbee | ||
|
|
25ed42073b | ||
|
|
dc89023147 | ||
|
|
32c34ff1eb | ||
|
|
cea9e1b812 | ||
|
|
4cf5928724 | ||
|
|
bbde223edb | ||
|
|
72507e799e | ||
|
|
9d3222ddc4 | ||
|
|
c15032eed1 | ||
|
|
dffd9a2f99 | ||
|
|
4478db9429 | ||
|
|
ea596262a3 | ||
|
|
6dd9ab61fc | ||
|
|
3d3e964173 | ||
|
|
6c0071ff08 | ||
|
|
2a32b20ab6 | ||
|
|
259566ec9a | ||
|
|
124c76413c | ||
|
|
b3b67f8b51 | ||
|
|
b9cea1ed17 | ||
|
|
1ab812ba92 | ||
|
|
1601315615 | ||
|
|
34f5e0887c | ||
|
|
86a74a726b | ||
|
|
783d416ecb | ||
|
|
ff0fbd8785 | ||
|
|
268b9aaacc | ||
|
|
71a89dca18 | ||
|
|
bdd7e1176f | ||
|
|
ad75261e53 | ||
|
|
df8c25d33b | ||
|
|
871e7a94db | ||
|
|
e22c861d46 | ||
|
|
bf0d468978 | ||
|
|
336dce2f2e | ||
|
|
cb7edbda94 | ||
|
|
d82784013c | ||
|
|
79442f5a4f | ||
|
|
e2d27fc52e | ||
|
|
a960883a91 | ||
|
|
21abde0feb | ||
|
|
39e6e43df2 | ||
|
|
80ce4af09b | ||
|
|
4dd3bb3374 | ||
|
|
55837bb543 | ||
|
|
bfc772b9f6 | ||
|
|
168b1b0083 | ||
|
|
f7cdcba740 | ||
|
|
7fd95fb00f | ||
|
|
3959588d37 | ||
|
|
adb877a343 | ||
|
|
64b83d6ceb | ||
|
|
4038c1c7cd | ||
|
|
fd47cea53e | ||
|
|
4b2bcc5978 | ||
|
|
eb75b1a36b | ||
|
|
3db18ec07f | ||
|
|
40f6c03acd | ||
|
|
759b1b39fa | ||
|
|
c6fd0dfd73 | ||
|
|
1730fc59ef | ||
|
|
a6fef26c79 | ||
|
|
17b875842a | ||
|
|
0f5f8a3c04 | ||
|
|
d0e4258f2c | ||
|
|
8bdc92fedf | ||
|
|
dc6637c082 | ||
|
|
4caa323bfe | ||
|
|
f397977f7d | ||
|
|
c9b1b1a117 | ||
|
|
254346c6ca | ||
|
|
1d855bf433 | ||
|
|
8d7e346189 | ||
|
|
65c3b0ffea | ||
|
|
04bfa28b4c | ||
|
|
fffc13d595 | ||
|
|
ef48c29b12 | ||
|
|
c249190497 | ||
|
|
a98b7daad7 | ||
|
|
759de2573a | ||
|
|
4e669b779f | ||
|
|
e76b1a0707 | ||
|
|
847d076523 | ||
|
|
1c62949b28 | ||
|
|
4cc95f3fd4 | ||
|
|
e3185d5e20 | ||
|
|
f2a0bcf08c | ||
|
|
e29705abef | ||
|
|
a6f0efef59 | ||
|
|
d28578704e | ||
|
|
f3ef23f85c | ||
|
|
412fff682b | ||
|
|
71c6bd7cf5 | ||
|
|
3ce2327c80 | ||
|
|
68de08772e | ||
|
|
53ac75b25f | ||
|
|
785bafe634 | ||
|
|
7ccb737be8 | ||
|
|
c138319a7c | ||
|
|
0b50e4ff76 | ||
|
|
f54babd4b6 | ||
|
|
0b3b43a965 | ||
|
|
c7d42ec577 | ||
|
|
23455a82ee | ||
|
|
c22538cf86 | ||
|
|
5b47e38135 | ||
|
|
5b97805747 | ||
|
|
9a124e3118 | ||
|
|
b185fe9924 | ||
|
|
f223f4fbed | ||
|
|
17185eb892 | ||
|
|
58cc618ecc | ||
|
|
52fd40bc11 | ||
|
|
9ae546303e | ||
|
|
be747ff5a5 | ||
|
|
65bfecbca8 | ||
|
|
e03c2ce8a7 | ||
|
|
bbc698bde9 | ||
|
|
ca789ee0f5 | ||
|
|
5ea0340f6c | ||
|
|
401be4e47b | ||
|
|
ae49943cd1 | ||
|
|
dfc14dd0d4 | ||
|
|
0dd57557cf | ||
|
|
9536723d15 | ||
|
|
e01176c758 | ||
|
|
92afcb466b | ||
|
|
1915db6025 | ||
|
|
3289d84be5 | ||
|
|
acdadebd3f | ||
|
|
45c0c436c2 | ||
|
|
b935b3321f | ||
|
|
68ba3b64aa | ||
|
|
afd315c59f | ||
|
|
aeba4da7b4 | ||
|
|
d6ebf617f7 | ||
|
|
322fa5182e | ||
|
|
879fc1d832 | ||
|
|
bf8abaf587 | ||
|
|
f168b060b1 | ||
|
|
c9f5af0c2a | ||
|
|
f8b4137eb1 | ||
|
|
91a4a25cbb | ||
|
|
f6edd53e92 | ||
|
|
691af8346f | ||
|
|
fb86e0ae59 | ||
|
|
5de456d9db | ||
|
|
70b21c56a9 | ||
|
|
b6917557aa | ||
|
|
b779730b03 |
@@ -1,3 +1,42 @@
|
||||
# life-science-ai-prompts
|
||||
# Life Science AI Prompts
|
||||
|
||||
Curated AI prompts for life sciences: genomics, proteomics, CRISPR, cell biology, and literature synthesis. Analogous to awesome-ai-for-science and awesome-genomic-skills.
|
||||
> *Where the prompts live, thrive, and reach the world.*
|
||||
|
||||
Curated, versioned AI prompts for life sciences research.
|
||||
Covers genomics, proteomics, CRISPR, cell biology, and literature synthesis.
|
||||
|
||||
## Analogous Resources Ingested
|
||||
|
||||
| Repository | Focus | Stars |
|
||||
|---|---|---|
|
||||
| [awesome-ai-for-science](https://github.com/ai-boost/awesome-ai-for-science) | AI tools across sciences | ★★★ |
|
||||
| [Awesome-LLM-Agents-Scientific-Discovery](https://github.com/zjlrock777/Awesome-LLM-Agents-Scientific-Discovery) | LLM agents in biomedical research | ★★★ |
|
||||
| [awesome-genomic-skills](https://github.com/GoekeLab/awesome-genomic-skills) | Genomic LLM agent skills | ★★★ |
|
||||
| [awesome-computational-biology](https://github.com/inoue0426/awesome-computational-biology) | Computational biology resources | ★★★ |
|
||||
| [Awesome-LLM-Scientific-Discovery](https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery) | LLMs for science | ★★★ |
|
||||
|
||||
## Folder Structure
|
||||
|
||||
```
|
||||
genomics/ — variant calling, RNA-seq, CRISPR
|
||||
proteomics/ — structure, mass spectrometry, interactions
|
||||
cell-biology/ — flow cytometry, microscopy, assays
|
||||
literature/ — paper summarisation, methods extraction
|
||||
```
|
||||
|
||||
## How to Use
|
||||
|
||||
```python
|
||||
import requests, base64
|
||||
|
||||
GITEA_URL = "https://promptnotes.ai"
|
||||
TOKEN = "your-token"
|
||||
|
||||
def read_prompt(owner, repo, path):
|
||||
url = f"{GITEA_URL}/api/v1/repos/{owner}/{repo}/contents/{path}"
|
||||
r = requests.get(url, headers={"Authorization": f"token {TOKEN}"})
|
||||
return base64.b64decode(r.json()["content"]).decode()
|
||||
|
||||
prompt = read_prompt("promptadmin", "life-science-ai-prompts",
|
||||
"genomics/variant-interpretation/snp-clinical-significance.md")
|
||||
```
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
---
|
||||
title: "Flow Cytometry Data Interpretation"
|
||||
domain: cell-biology
|
||||
persona: "Molecular Biologist"
|
||||
persona_background: >
|
||||
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
persona_style: "precise, evidence-based, uses established nomenclature"
|
||||
models: [gpt-4, claude-3-5]
|
||||
keywords: [flow-cytometry, FACS, cell-population, gating, immunophenotyping]
|
||||
task: "Interpret flow cytometry gating strategy and cell population data."
|
||||
validated: false
|
||||
version: 1.0.0
|
||||
author: promptadmin
|
||||
source_repositories:
|
||||
- https://github.com/zjlrock777/Awesome-LLM-Agents-Scientific-Discovery
|
||||
---
|
||||
|
||||
# Flow Cytometry Data Interpretation
|
||||
|
||||
## Persona
|
||||
|
||||
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
> Your communication style: precise, evidence-based, uses established nomenclature
|
||||
|
||||
## Task
|
||||
|
||||
Interpret flow cytometry gating strategy and cell population data.
|
||||
|
||||
## Prompt
|
||||
|
||||
```
|
||||
You are an expert in flow cytometry and immunophenotyping.
|
||||
|
||||
Given flow cytometry experiment:
|
||||
- Cell type: {cell_type}
|
||||
- Tissue source: {tissue}
|
||||
- Panel: {markers}
|
||||
- Gating strategy: {gating_description}
|
||||
- Key populations identified: {populations}
|
||||
- Experimental condition: {condition}
|
||||
- Controls: {controls}
|
||||
|
||||
Provide:
|
||||
1. Assessment of gating strategy quality
|
||||
2. Interpretation of each identified cell population
|
||||
3. Biological significance of observed population shifts
|
||||
4. Statistical recommendations (% parent vs % total, n required)
|
||||
5. Potential artefacts and confounders
|
||||
6. Suggested additional markers for confirmation
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
Works well with FlowJo or FCS Express output descriptions. Reference: STAgent (Harvard LiuLab, bioRxiv 2025) for spatial context.
|
||||
|
||||
## Compatibility
|
||||
|
||||
| Model | Tested | Notes |
|
||||
|-------|--------|-------|
|
||||
| gpt-4 | ⬜ | |
|
||||
| claude-3-5 | ⬜ | |
|
||||
|
||||
## Keywords
|
||||
|
||||
`flow-cytometry` `FACS` `cell-population` `gating` `immunophenotyping`
|
||||
@@ -0,0 +1,62 @@
|
||||
---
|
||||
title: "CRISPR Off-Target Risk Assessment"
|
||||
domain: genomics
|
||||
persona: "Molecular Biologist"
|
||||
persona_background: >
|
||||
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
persona_style: "precise, evidence-based, uses established nomenclature"
|
||||
models: [gpt-4, claude-3-5]
|
||||
keywords: [CRISPR, guide-RNA, off-target, Cas9, gene-editing]
|
||||
task: "Assess off-target risk of a CRISPR guide RNA based on computational predictions."
|
||||
validated: false
|
||||
version: 1.0.0
|
||||
author: promptadmin
|
||||
source_repositories:
|
||||
- https://github.com/ai-boost/awesome-ai-for-science
|
||||
---
|
||||
|
||||
# CRISPR Off-Target Risk Assessment
|
||||
|
||||
## Persona
|
||||
|
||||
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
> Your communication style: precise, evidence-based, uses established nomenclature
|
||||
|
||||
## Task
|
||||
|
||||
Assess off-target risk of a CRISPR guide RNA based on computational predictions.
|
||||
|
||||
## Prompt
|
||||
|
||||
```
|
||||
You are a CRISPR expert with deep knowledge of guide RNA design and off-target effects.
|
||||
|
||||
Given:
|
||||
- Target gene: {target_gene}
|
||||
- Guide RNA sequence (20nt): {grna_sequence}
|
||||
- Predicted off-target sites (from CRISPOR/Cas-OFFinder): {off_target_sites}
|
||||
- Genome: {genome_assembly}
|
||||
- Cas variant: {cas_variant}
|
||||
|
||||
Provide:
|
||||
1. Risk classification (Low/Medium/High)
|
||||
2. Analysis of top 3 off-target sites by genomic context
|
||||
3. Recommended experimental validation strategy (T7E1, GUIDE-seq, etc.)
|
||||
4. Alternative guide RNA suggestions if risk is High
|
||||
5. Summary suitable for IACUC/ethics submission
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
Referenced from BioAgents framework (bio-xyz, arXiv 2601.12542).
|
||||
|
||||
## Compatibility
|
||||
|
||||
| Model | Tested | Notes |
|
||||
|-------|--------|-------|
|
||||
| gpt-4 | ⬜ | |
|
||||
| claude-3-5 | ⬜ | |
|
||||
|
||||
## Keywords
|
||||
|
||||
`CRISPR` `guide-RNA` `off-target` `Cas9` `gene-editing`
|
||||
@@ -0,0 +1,62 @@
|
||||
---
|
||||
title: "RNA-seq Differential Expression Narrative"
|
||||
domain: genomics
|
||||
persona: "Molecular Biologist"
|
||||
persona_background: >
|
||||
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
persona_style: "precise, evidence-based, uses established nomenclature"
|
||||
models: [gpt-4, claude-3-5]
|
||||
keywords: [RNA-seq, DESeq2, differential-expression, pathway-analysis, fold-change]
|
||||
task: "Generate a scientific narrative from RNA-seq differential expression results."
|
||||
validated: true
|
||||
version: 1.0.0
|
||||
author: promptadmin
|
||||
source_repositories:
|
||||
- https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery
|
||||
---
|
||||
|
||||
# RNA-seq Differential Expression Narrative
|
||||
|
||||
## Persona
|
||||
|
||||
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
> Your communication style: precise, evidence-based, uses established nomenclature
|
||||
|
||||
## Task
|
||||
|
||||
Generate a scientific narrative from RNA-seq differential expression results.
|
||||
|
||||
## Prompt
|
||||
|
||||
```
|
||||
You are a senior molecular biologist analysing transcriptomic data.
|
||||
|
||||
Given DESeq2 differential expression results:
|
||||
- Comparison: {condition_A} vs {condition_B}
|
||||
- Significantly upregulated genes (top 10): {up_genes}
|
||||
- Significantly downregulated genes (top 10): {down_genes}
|
||||
- Pathway enrichment results: {pathways}
|
||||
- Experimental context: {context}
|
||||
|
||||
Write a Results section (150-200 words) for a peer-reviewed manuscript that:
|
||||
1. Summarises the overall transcriptional response
|
||||
2. Highlights key gene clusters and their biological significance
|
||||
3. Connects enriched pathways to the experimental condition
|
||||
4. Uses appropriate statistical language (FDR, log2FC)
|
||||
5. Avoids overclaiming causality
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
Derived from GenoTEX benchmark methodology (Liu et al. 2024). Works best with GSEA or EnrichR pathway results.
|
||||
|
||||
## Compatibility
|
||||
|
||||
| Model | Tested | Notes |
|
||||
|-------|--------|-------|
|
||||
| gpt-4 | ✅ | |
|
||||
| claude-3-5 | ✅ | |
|
||||
|
||||
## Keywords
|
||||
|
||||
`RNA-seq` `DESeq2` `differential-expression` `pathway-analysis` `fold-change`
|
||||
@@ -0,0 +1,78 @@
|
||||
---
|
||||
title: "SNP Clinical Significance Interpreter"
|
||||
domain: genomics
|
||||
persona: "Molecular Biologist"
|
||||
persona_background: >
|
||||
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
persona_style: "precise, evidence-based, uses established nomenclature"
|
||||
models: [gpt-4, claude-3-5, gemini-1-5-pro]
|
||||
keywords: [SNP, variant-calling, clinical-significance, VCF, ClinVar, ACMG]
|
||||
task: "Interpret the clinical significance of a single nucleotide polymorphism (SNP) from VCF annotation data."
|
||||
validated: true
|
||||
version: 1.0.0
|
||||
author: promptadmin
|
||||
source_repositories:
|
||||
- https://github.com/GoekeLab/awesome-genomic-skills
|
||||
- https://github.com/ai-boost/awesome-ai-for-science
|
||||
---
|
||||
|
||||
# SNP Clinical Significance Interpreter
|
||||
|
||||
## Persona
|
||||
|
||||
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
> Your communication style: precise, evidence-based, uses established nomenclature
|
||||
|
||||
## Task
|
||||
|
||||
Interpret the clinical significance of a single nucleotide polymorphism (SNP) from VCF annotation data.
|
||||
|
||||
## Prompt
|
||||
|
||||
```
|
||||
You are a molecular biologist specialising in clinical genomics.
|
||||
|
||||
Given the following SNP annotation data from a VCF file:
|
||||
- Gene: {gene_name}
|
||||
- Variant: {hgvs_notation}
|
||||
- ClinVar classification: {clinvar_class}
|
||||
- gnomAD allele frequency: {gnomad_af}
|
||||
- CADD score: {cadd_score}
|
||||
- In silico predictions: {sift} (SIFT), {polyphen} (PolyPhen-2)
|
||||
|
||||
Provide:
|
||||
1. ACMG/AMP classification (Pathogenic/Likely Pathogenic/VUS/Likely Benign/Benign)
|
||||
2. Evidence summary (2-3 sentences)
|
||||
3. Clinical implications
|
||||
4. Recommended follow-up actions
|
||||
5. Caveats and limitations
|
||||
```
|
||||
|
||||
### Example 1
|
||||
|
||||
**Input:**
|
||||
```
|
||||
Gene: BRCA1, Variant: c.5266dupC, ClinVar: Pathogenic, gnomAD: 0.00001, CADD: 35
|
||||
```
|
||||
|
||||
**Output:**
|
||||
```
|
||||
ACMG: Pathogenic. This frameshift variant creates a premature stop codon...
|
||||
```
|
||||
|
||||
|
||||
## Notes
|
||||
|
||||
Inspired by SRAgent (Arc Institute) for genomic database querying. Best used with SnpEff/VEP-annotated VCF files.
|
||||
|
||||
## Compatibility
|
||||
|
||||
| Model | Tested | Notes |
|
||||
|-------|--------|-------|
|
||||
| gpt-4 | ✅ | |
|
||||
| claude-3-5 | ✅ | |
|
||||
| gemini-1-5-pro | ✅ | |
|
||||
|
||||
## Keywords
|
||||
|
||||
`SNP` `variant-calling` `clinical-significance` `VCF` `ClinVar` `ACMG`
|
||||
@@ -0,0 +1,67 @@
|
||||
---
|
||||
title: "Scientific Paper Deep Summarisation"
|
||||
domain: literature
|
||||
persona: "Molecular Biologist"
|
||||
persona_background: >
|
||||
PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
persona_style: "precise, evidence-based, uses established nomenclature"
|
||||
models: [gpt-4, claude-3-5, gemini-1-5-pro]
|
||||
keywords: [literature-review, paper-summarisation, methods-extraction, PubMed]
|
||||
task: "Generate a structured deep summary of a life sciences research paper."
|
||||
validated: true
|
||||
version: 1.0.0
|
||||
author: promptadmin
|
||||
source_repositories:
|
||||
- https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery
|
||||
- https://github.com/zjlrock777/Awesome-LLM-Agents-Scientific-Discovery
|
||||
---
|
||||
|
||||
# Scientific Paper Deep Summarisation
|
||||
|
||||
## Persona
|
||||
|
||||
> You are a **Molecular Biologist**. PhD-level molecular biologist with 10+ years experience in genomics, CRISPR, and transcriptomics.
|
||||
> Your communication style: precise, evidence-based, uses established nomenclature
|
||||
|
||||
## Task
|
||||
|
||||
Generate a structured deep summary of a life sciences research paper.
|
||||
|
||||
## Prompt
|
||||
|
||||
```
|
||||
You are an expert scientific reader with broad knowledge of life sciences.
|
||||
|
||||
Read the following paper abstract/full text and provide a structured summary:
|
||||
|
||||
Paper text:
|
||||
{paper_text}
|
||||
|
||||
Generate:
|
||||
1. **TL;DR** (1 sentence, non-technical)
|
||||
2. **Background** — What problem does this paper address?
|
||||
3. **Key Methods** — What experimental and computational approaches were used?
|
||||
4. **Main Findings** — What are the 3-5 most important results?
|
||||
5. **Novelty** — What is genuinely new compared to prior work?
|
||||
6. **Limitations** — What are the key weaknesses the authors acknowledge or you identify?
|
||||
7. **Clinical/Translational Relevance** — Practical implications (1-2 sentences)
|
||||
8. **Follow-up Questions** — 3 questions this paper raises
|
||||
|
||||
Format: structured markdown with headers.
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
Inspired by Agent Laboratory (2024) three-phase research pipeline. For full-text papers, chunk into introduction + methods + results + discussion.
|
||||
|
||||
## Compatibility
|
||||
|
||||
| Model | Tested | Notes |
|
||||
|-------|--------|-------|
|
||||
| gpt-4 | ✅ | |
|
||||
| claude-3-5 | ✅ | |
|
||||
| gemini-1-5-pro | ✅ | |
|
||||
|
||||
## Keywords
|
||||
|
||||
`literature-review` `paper-summarisation` `methods-extraction` `PubMed`
|
||||
@@ -0,0 +1,66 @@
|
||||
---
|
||||
title: "Protein Structure Functional Interpretation"
|
||||
domain: proteomics
|
||||
persona: "Structural Biologist"
|
||||
persona_background: >
|
||||
Computational structural biologist specialising in protein folding, cryo-EM, and AlphaFold interpretation.
|
||||
persona_style: "quantitative, structure-first, references PDB entries"
|
||||
models: [gpt-4, claude-3-5, gemini-1-5-pro]
|
||||
keywords: [AlphaFold, protein-structure, PDB, active-site, binding-pocket]
|
||||
task: "Interpret AlphaFold2/3 or experimental protein structure in functional context."
|
||||
validated: true
|
||||
version: 1.0.0
|
||||
author: promptadmin
|
||||
source_repositories:
|
||||
- https://github.com/inoue0426/awesome-computational-biology
|
||||
- https://github.com/ai-boost/awesome-ai-for-science
|
||||
---
|
||||
|
||||
# Protein Structure Functional Interpretation
|
||||
|
||||
## Persona
|
||||
|
||||
> You are a **Structural Biologist**. Computational structural biologist specialising in protein folding, cryo-EM, and AlphaFold interpretation.
|
||||
> Your communication style: quantitative, structure-first, references PDB entries
|
||||
|
||||
## Task
|
||||
|
||||
Interpret AlphaFold2/3 or experimental protein structure in functional context.
|
||||
|
||||
## Prompt
|
||||
|
||||
```
|
||||
You are a structural biologist with expertise in computational and experimental structural analysis.
|
||||
|
||||
Given protein structure data:
|
||||
- Protein name: {protein_name}
|
||||
- UniProt ID: {uniprot_id}
|
||||
- Structure source: {source} (AlphaFold2 / AlphaFold3 / X-ray / Cryo-EM)
|
||||
- pLDDT scores summary: {plddt_summary}
|
||||
- Key structural features: {features}
|
||||
- Known binding partners: {binding_partners}
|
||||
|
||||
Provide:
|
||||
1. Overall structural assessment (fold classification, domain organisation)
|
||||
2. Confidence assessment for key regions (if AlphaFold)
|
||||
3. Predicted functional sites (active site, allosteric sites, binding interfaces)
|
||||
4. Druggability assessment of binding pockets
|
||||
5. Structural basis for any known pathogenic variants
|
||||
6. Recommended follow-up experiments
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
Integrates well with PyMOL output descriptions and PDB REMARK sections. For AlphaFold3 structures, note pLDDT < 70 regions as disordered.
|
||||
|
||||
## Compatibility
|
||||
|
||||
| Model | Tested | Notes |
|
||||
|-------|--------|-------|
|
||||
| gpt-4 | ✅ | |
|
||||
| claude-3-5 | ✅ | |
|
||||
| gemini-1-5-pro | ✅ | |
|
||||
|
||||
## Keywords
|
||||
|
||||
`AlphaFold` `protein-structure` `PDB` `active-site` `binding-pocket`
|
||||
@@ -0,0 +1,46 @@
|
||||
---
|
||||
title: "Contribution Guidelines"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/GoekeLab/awesome-genomic-skills/blob/f88d9494/CONTRIBUTING.md
|
||||
upstream_sha: f88d9494
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Contribution Guidelines
|
||||
|
||||
Thank you for considering contributing to **Awesome Genomic Skills**! This is a curated list, so please make sure your suggestion meets the criteria below before submitting a pull request.
|
||||
|
||||
## Adding an Item
|
||||
|
||||
1. **Fork** the repository and create a new branch.
|
||||
2. Add your item to the appropriate section in `README.md`.
|
||||
3. Submit a **pull request** with a clear description of the item and why it belongs in the list.
|
||||
|
||||
## Criteria for Inclusion
|
||||
|
||||
To be included, a repository should meet the following:
|
||||
|
||||
- **Relevance** — directly related to AI coding agents, and useful for genomics/bioinformatics.
|
||||
- **Quality** — actively maintained, well-documented, and demonstrably useful.
|
||||
- **Not deprecated** — must not be archived or abandoned.
|
||||
|
||||
## Format
|
||||
|
||||
Each entry should follow this format:
|
||||
|
||||
```markdown
|
||||
- [Repository Name](URL) - Short description explaining what it does and why it is useful.
|
||||
```
|
||||
|
||||
- Keep descriptions concise (one or two sentences).
|
||||
- Place the entry in the most appropriate existing section.
|
||||
- If no section fits, propose a new one in your pull request description.
|
||||
|
||||
## Pull Request Guidelines
|
||||
|
||||
- Verify all links are correct and publicly accessible.
|
||||
@@ -0,0 +1,134 @@
|
||||
---
|
||||
title: "Awesome Genomic Skills [](https://awesome.re)"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/GoekeLab/awesome-genomic-skills/blob/8cf42e1d/README.md
|
||||
upstream_sha: 8cf42e1d
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Awesome Genomic Skills [](https://awesome.re)
|
||||
|
||||
A curated list of skills and MCP servers for working with AI coding agents (Claude Code, GitHub Copilot, Codex, Cursor, Gemini CLI, etc.) in genomics and bioinformatics, alongside other useful repositories such as benchmarks and general AI coding skills collections.
|
||||
|
||||
### What is a Skill and what is an MCP?
|
||||
|
||||
**Skill:** a Markdown file (plus optional scripts) that teaches an agent *how* to do a task — procedural know-how loaded into context on demand.
|
||||
|
||||
**MCP server:** a running service that gives an agent a *connection* to external systems (databases, tools, pipelines) via standardized tool calls.
|
||||
|
||||
## Contents
|
||||
|
||||
- [Awesome Genomic Skills ](#awesome-genomic-skills-)
|
||||
- [What is a Skill and what is an MCP?](#what-is-a-skill-and-what-is-an-mcp)
|
||||
- [Contents](#contents)
|
||||
- [Bioinformatics and Genomics Agent Skills](#bioinformatics-and-genomics-agent-skills)
|
||||
- [MCP Servers for Life Sciences](#mcp-servers-for-life-sciences)
|
||||
- [Benchmarks](#benchmarks)
|
||||
- [General AI Coding Agent Skill Collections](#general-ai-coding-agent-skill-collections)
|
||||
- [Other Notable Awesome Lists!](#other-notable-awesome-lists)
|
||||
- [Contributing](#contributing)
|
||||
|
||||
## Bioinformatics and Genomics Agent Skills
|
||||
|
||||
Skill libraries and tool collections specifically targeting genomics, bioinformatics, and life sciences work with AI coding agents.
|
||||
|
||||
- [science-skills](https://github.com/google-deepmind/science-skills)
|
||||
- **Description:** Collection of ~36 agent skills spanning genomics, structural biology, cheminformatics, and literature search; wraps AlphaGenome (single-variant effect prediction), AlphaFold DB, and 30+ databases/tools (UniProt, Ensembl, gnomAD, GTEx, ClinVar, dbSNP, ChEMBL, PubChem, PDB, Foldseek, JASPAR, Reactome, STRING, Open Targets, Human Protein Atlas, PyMOL) for grounded, token-efficient scientific workflows. Apache-2.0; built for Google Antigravity but installable into any agent via `npx skills add`. [Technical report](https://storage.googleapis.com/deepmind-media/papers/google_deepmind_science_skills_for_antigravity_towards_efficient_and_reliable_scientific_workflows.pdf).
|
||||
- **Developers:** Google DeepMind.
|
||||
- [openai/plugins — life-science-research](https://github.com/openai/plugins/tree/main/plugins/life-science-research)
|
||||
- **Description:** OpenAI's Life Sciences research plugin for Codex; bundles 50 modular skills spanning human genetics, functional genomics, expression, pathway biology, protein structure, chemistry, and clinical evidence, wrapping 50+ public databases/tools (Ensembl, UniProt, gnomAD, GTEx, ClinVar, GWAS Catalog, Open Targets, ChEMBL, PubChem, RCSB PDB, AlphaFold, Reactome, STRING, Human Protein Atlas, cBioPortal, CellxGene, ENCODE, NCBI Entrez/BLAST/Datasets, and more). Works with mainline GPT-5.4 models, with GPT-Rosalind (trusted-access) for deeper reasoning. Closely parallels DeepMind's science-skills.
|
||||
- **Developers:** OpenAI.
|
||||
- [anthropics/life-sciences](https://github.com/anthropics/life-sciences)
|
||||
- **Description:** Anthropic's "Claude for Life Sciences" Claude Code Marketplace — a hybrid bundle rather than a pure skills library. It mixes first-party procedural/method skills (`single-cell-rna-qc`, `scvi-tools`, `nextflow-development`, `clinical-trial-protocol-skill`, `scientific-problem-selection`) with a large set of partner MCP servers and data integrations (10x Genomics, ChEMBL, Open Targets, PubMed, bioRxiv, Synapse, ToolUniverse, BioRender, Medidata, Cortellis, Owkin, Consensus, and more). Note that many entries are MCP servers, not skills.
|
||||
- **Developers:** Anthropic (with commercial life-sciences partners).
|
||||
- [sample-kiro-power-life-sciences](https://github.com/aws-samples/sample-kiro-power-life-sciences)
|
||||
- **Description:** AWS sample bundle for the Kiro IDE: 24 MCP servers wrapping 100+ databases/tools across genomics, proteomics, structural biology, and clinical/pharma (NCBI, Ensembl, ClinVar, gnomAD, UniProt, STRING, PDB, AlphaFold, ChEMBL, Open Targets, etc.), plus 10 domain skills and 16 workflows, with cross-database search and AWS HealthOmics pipeline execution. MIT-0; the MCP servers use standard MCP and are portable, though the skills and workflows are built for Kiro. [Blog post](https://aws.amazon.com/blogs/publicsector/accelerating-life-sciences-research-with-kiro-a-unified-ai-interface-to-100-open-source-databases/).
|
||||
- **Developers:** AWS (AWS Samples).
|
||||
- [ClawBio](https://github.com/ClawBio/ClawBio)
|
||||
- **Description:** Bioinformatics-native AI agent skill library; 95 reproducible, local-first skills for genomics tasks (variant calling, RNA-seq, population genetics) that work with Claude Code, Copilot, Codex, and other agents. Since v0.6.1 the whole library is also callable as an MCP server (`uvx --from 'clawbio[mcp]' clawbio mcp`) from Cursor, Claude Desktop, VS Code, or Zed (see the MCP section below). Skills are actively [benchmarked](https://clawbio.ai/benchmarks.html).
|
||||
- **Developers:** Independent open-source project built on [OpenClaw](https://openclaw.ai).
|
||||
- [SciAgent-skills](https://github.com/jaechang-hits/SciAgent-Skills)
|
||||
- **Description:** 197 open-source skills for Claude Code, Cursor, Codex, and Windsurf, covering genomics-bioinformatics, proteomics-protein engineering, structural biology, drug discovery, systems biology, biostatistics, and scientific writing; achieves 92% accuracy on BixBench-Verified-50 (+26.7 pts over Claude Code baseline). The hosted OmicsHorizon web platform runs these skills in-browser.
|
||||
- **Developers:** The team behind the [OmicsHorizon](https://omicshorizon.ai/en/) platform (Jaechang Lim).
|
||||
- [scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
|
||||
- **Description:** Currently 135 skills covering various scientific areas including genomics, but also broader scientific areas like geospatial science etc.
|
||||
- **Developers:** K-Dense AI — an MIT-founded startup (Accel, Accel Atoms, Google AI Futures Fund) building an AI research platform ([k-dense.ai](https://www.k-dense.ai)); this is one of its open-source spin-offs.
|
||||
- [bioSkills](https://github.com/GPTomics/bioSkills)
|
||||
- **Description:** SKILL.md files for bioinformatics with Claude Code, covering end-to-end pipelines like RNA-seq, variants, ChIP-seq, scRNA-seq, spatial, Hi-C, proteomics, microbiome, CRISPR, metabolomics, multi-omics, immunotherapy, outbreak analysis, and Mendelian randomization tools.
|
||||
- **Developers:** GPTomics ([gptomics.com](https://www.gptomics.com)), an independent research lab working at the bioinformatics/AI intersection.
|
||||
- [Clair-skills](https://github.com/HKU-BAL/Clair-skills)
|
||||
- **Description:** Agent skill for the Clair suite of variant callers (Clair3, ClairS, Clair3-RNA, Clair-Mosaic); provides intelligent model selection, command generation, and troubleshooting for germline, somatic, mosaic, and RNA-seq variant calling from long-read and short-read sequencing data.
|
||||
- **Developers:** Ruibang Luo's BioAI Lab at the University of Hong Kong (HKU-BAL).
|
||||
- [ToolUniverse](https://github.com/mims-harvard/ToolUniverse)
|
||||
- **Description:** Ecosystem for building AI scientist systems; integrates 1,000+ machine learning models, datasets, and scientific APIs for data analysis, knowledge retrieval, and experimental design in biomedicine, with 68 pre-built agent skills covering drug discovery, precision oncology, and rare-disease diagnosis.
|
||||
- **Developers:** Harvard's Zitnik Lab (AI for Medicine and Science).
|
||||
- [operon](https://github.com/swaruplab/operon)
|
||||
- **Description:** AI-powered bioinformatics IDE bundling 180+ SKILL.md-format analysis protocols covering RNA-seq, scRNA-seq, ATAC-seq, ChIP-seq, WGS/WES, spatial transcriptomics, proteomics, GWAS, and external database query patterns (PubMed, GEO, GTEx, KEGG, UniProt, JASPAR, AlphaFold). The desktop app wraps Claude Code and adds HPC/SSH integration; the protocol files themselves are MIT-licensed Markdown in the repo's `protocols/` directory, extractable for use with any SKILL.md-compatible agent.
|
||||
- **Developers:** Swarup Lab (UC Irvine).
|
||||
- [OpenClaw-Medical-Skills](https://github.com/FreedomIntelligence/OpenClaw-Medical-Skills)
|
||||
- **Description:** Meta-aggregation of 872 skills curated from 12+ upstream repositories, including ClawBio, ToolUniverse, GPTomics bioSkills, BioOS, and others; spans clinical workflows, genomics, drug discovery, bioinformatics pipelines, and medical device regulatory frameworks. Expect overlap with the upstream repos listed separately.
|
||||
- **Developers:** FreedomIntelligence, the medical-NLP research group at CUHK-Shenzhen / Shenzhen Research Institute of Big Data (led by Benyou Wang; also behind HuatuoGPT).
|
||||
|
||||
## MCP Servers for Life Sciences
|
||||
|
||||
Model Context Protocol (MCP) servers that give AI agents direct access to bioinformatics databases, tools, and analysis pipelines.
|
||||
|
||||
- [ChatSpatial](https://github.com/cafferychen777/ChatSpatial) - MCP server for spatial transcriptomics analysis through natural language; supports Scanpy, Squidpy, cell communication analysis, and spatial domain identification.
|
||||
- [biomcp](https://github.com/genomoncology/biomcp) - Single MCP server able to query multiple information sources, including clinical trials, genetic data & published medical literature.
|
||||
- [gget-mcp](https://github.com/longevity-genie/gget-mcp) - MCP server wrapping the Pachter Lab [gget](https://github.com/pachterlab/gget) bioinformatics toolkit. Exposes 13 tools covering gene search and metadata (Ensembl), sequence retrieval, BLAST/BLAT/MUSCLE alignment, expression data (ARCHS4), functional enrichment (Enrichr), protein structure (PDB, AlphaFold), cancer mutations (COSMIC), and single-cell queries (CellxGene).
|
||||
- [Seqera MCP](https://docs.seqera.io/platform-cloud/seqera-mcp/overview) - Hosted MCP server from Seqera Labs (the developers of Nextflow) exposing the Seqera Platform (workflow launch/management), Wave (container provisioning), nf-core modules, and SRA/ENA/GEO retrieval.
|
||||
- [knowledgebase-mcp](https://github.com/biocontext-ai/knowledgebase-mcp) - BioContextAI Knowledgebase MCP server, included in the BioContextAI registry (below), one of the most comprehensive single MCP packages (wraps STRINGDb, Open Targets, Reactome, UniProt, HPA, KEGG, AlphaFold, Ensembl, ClinicalTrials.gov, bioRxiv, etc.)
|
||||
- [ClawBio MCP](https://github.com/ClawBio/ClawBio) - MCP mode of the [ClawBio](#bioinformatics-and-genomics-agent-skills) skill library (0.6.1): exposes all 95 genomics skills as callable tools over local stdio, letting different agents (Cursor, Zed, etc.) run analyses (variant calling, RNA-seq, population genetics).
|
||||
- [roda-mcp](https://github.com/awslabs/mcp/tree/main/src/roda-mcp-server) - AWS Labs MCP server for discovering and exploring datasets in the Registry of Open Data on AWS (RODA), covering 1,100+ public datasets across life sciences, climate, geospatial, and satellite imagery. Search/filter by keyword, organization, or license; inspect dataset details; and browse or sample public S3 bucket contents directly (no AWS account required) without downloading full files. Not genomics-specific, but useful for locating and previewing life-sciences datasets (e.g. SG-NEx) hosted on AWS.
|
||||
- [plant-genomics-mcp](https://github.com/musharna/plant-genomics-mcp) - Plant-genomics MCP server exposing 50 tools across 23 backends, keyed on TAIR-style loci with organism resolution across 12 crop and model species. Covers plant-specific resources that general bioinformatics servers do not (Ensembl Plants, Phytozome, Gramene, Planteome PO/TO, PlantCyc/PMN, AraGWAS, 1001 Genomes, ThaleMine, BAR, JASPAR) alongside the usual UniProt/KEGG/STRING/AlphaFold/PDBe/InterPro/Europe PMC, plus cross-source synthesis tools that compose several backends into a single gene report.
|
||||
|
||||
Existing registries and lists of MCP servers:
|
||||
- [BioContextAI Registry](https://github.com/biocontext-ai/registry/) - Community-curated catalogue of biomedical MCP servers, with submission criteria requiring biomedical focus, free academic access, OSI-approved open-source licenses, and MCP specification compliance. Ships a [cookiecutter template](https://github.com/biocontext-ai/mcp-server-cookiecutter) for new servers and follows Schema.org ontologies for metadata. [A community hub for agentic biomedical systems](https://www.nature.com/articles/s41587-025-02900-9)
|
||||
- [MCPmed](https://github.com/MCPmed) - Reference MCP implementations (GEO, STRING, UCSC Cell Browser, PLSDB) plus a cookiecutter template and HTML "breadcrumbs" discovery mechanism for transitioning legacy services to MCP. Published as a call paper in Briefings in Bioinformatics. [MCPmed: a call for Model Context Protocol-enabled bioinformatics web services for LLM-driven discovery](https://academic.oup.com/bib/article/27/1/bbag076/8495038)
|
||||
- [awesome-mcp-servers](https://github.com/punkpeye/awesome-mcp-servers#bio) - General-purpose MCP server list with a Biology, Medicine, and Bioinformatics subsection.
|
||||
|
||||
MCP related tools:
|
||||
- [BioinfoMCP](https://github.com/florensiawidjaja/BioinfoMCP) - Not strictly an MCP, a converter that auto-generates MCP servers from existing tool documentation, plus a benchmark of the converted tools. Preprint available [here](https://arxiv.org/abs/2510.02139)
|
||||
|
||||
## Benchmarks
|
||||
|
||||
Benchmarks, evaluations, and other helpful resources at the intersection of AI and genomics/bioinformatics.
|
||||
|
||||
Agent Capability Benchmarks:
|
||||
- [BioAgent Bench](https://github.com/bioagent-bench/bioagent-bench) - 10 end-to-end bioinformatics pipeline tasks (RNA-seq, variant calling, metagenomics, single-cell, transcript quantification, etc.) with concrete output artifacts; includes a perturbation suite (corrupted inputs, decoy reference files, prompt bloat) — probes agent robustness under controlled stress and shows that correct high-level pipeline construction does not guarantee reliable step-level reasoning.
|
||||
- [BioMed-AQA](https://huggingface.co/datasets/BOBQWERA/biomed-aqa-dataset) - 327 open-ended biomedical analysis tasks across omics, visualisation, machine learning, statistics, and precision medicine; uses milestone-based grading against reference analytical steps, with a complementary 172-question multiple-choice subset. Released alongside the BioMedAgent system in Nature Biomedical Engineering — same-team caveat applies to headline scores.
|
||||
- [BiomniBench](https://huggingface.co/datasets/phylobio/BiomniBench-DA) - Process-level evaluation framework: 100 biomedical data-analysis tasks curated from high-impact papers by original authors or domain experts; grades the full agent trajectory (reasoning trace + final answer) against expert-designed rubrics via an LLM judge — addresses the outcome-only blind spots in benchmarks like BixBench.
|
||||
- [BixBench](https://github.com/Future-House/BixBench) - Comprehensive benchmark for LLM-based agents on real-world computational biology tasks; tests agents' ability to explore biological datasets, perform multi-step analyses, and interpret results — useful for evaluating which agents and skills perform best on genomics work.
|
||||
- [CompBioBench](https://github.com/Genentech/compbiobench-runner) - Genentech-released benchmark of 100 computational biology questions spanning single-cell, epigenomics, genomics, transcriptomics, human genetics, and ML; agents start from a bare-minimum environment and must fetch their own tools and data, with exact-string-match grading on a single ground-truth answer.
|
||||
- [LAB-Bench](https://huggingface.co/datasets/futurehouse/lab-bench) - 2,400+ multiple-choice questions across 8 subtasks of practical biology research (literature search, figure/table interpretation, database access, wet-lab protocols, sequence analysis, cloning scenarios); FutureHouse's predecessor to BixBench, probing biological knowledge and reasoning rather than agentic execution.
|
||||
|
||||
Skills benchmarks:
|
||||
- [SkillsBench](https://github.com/benchflow-ai/skillsbench) - General-purpose benchmark for measuring whether agent skills actually help: 86 tasks across 11 domains (including healthcare), each run under three conditions — no skills, curated skills, and self-generated skills — with deterministic verifiers. Across 7,308 trajectories, curated skills lifted average pass rate by +16.2 pts (ranging from +4.5 pts for software engineering to +51.9 pts for healthcare), while self-generated skills gave no average benefit and focused 2–3 module skills beat comprehensive documentation. Not bioinformatics-specific, but the canonical "do skills actually work" benchmark. [Paper](https://arxiv.org/abs/2602.12670).
|
||||
|
||||
## General AI Coding Agent Skill Collections
|
||||
|
||||
Popular repositories of reusable skills for AI coding agents. Each entry is useful for genomics and bioinformatics coding work; descriptions explain how.
|
||||
|
||||
Skills specs:
|
||||
- [agentskills](https://github.com/agentskills/agentskills) - Official specification and documentation for the Agent Skills standard; defines the cross-platform format used by skills in this list and on Claude Code, Copilot, Codex, Cursor, and Gemini CLI.
|
||||
|
||||
General skills collections (non-exhaustive):
|
||||
- [superpowers](https://github.com/obra/superpowers) - Complete software development methodology and skills framework for coding agents; enforces TDD, spec-driven design, and subagent-driven development — practices that improve reproducibility and correctness in complex genomics pipelines.
|
||||
- [anthropics/skills](https://github.com/anthropics/skills) - Official Anthropic reference implementation and specification for Claude agent skills; the canonical starting point for building custom skills that teach Claude how to handle bioinformatics tasks and lab workflows.
|
||||
- [openai/skills](https://github.com/openai/skills) - OpenAI's official Skills Catalog for Codex; ships system, curated, and experimental agent skills installable via the in-Codex `$skill-installer`, built on the cross-platform Agent Skills open standard. The OpenAI counterpart to anthropics/skills, and a useful reference for packaging reusable, repeatable bioinformatics coding workflows for Codex.
|
||||
- [andrej-karpathy-skills](https://github.com/forrestchang/andrej-karpathy-skills) - Single-file coding guidelines derived from Andrej Karpathy's observations on LLM pitfalls; instils simplicity, surgical changes, and goal-driven execution — critical disciplines when AI agents write or modify genomics analysis code.
|
||||
- [awesome-copilot](https://github.com/github/awesome-copilot) - Community-curated instructions, agents, skills, and configurations for GitHub Copilot, including prompt templates applicable to scientific and data-analysis workflows.
|
||||
- [awesome-llm-skills](https://github.com/Prat011/awesome-llm-skills) - Curated list of LLM and AI agent skills, resources, and tools for customizing AI agent workflows across Claude Code, Codex, Gemini CLI, and custom agents.
|
||||
|
||||
## Other Notable Awesome Lists!
|
||||
|
||||
- [Awesome-Scientific-Skills](https://github.com/InternScience/Awesome-Scientific-Skills) - An open, curated collection of agent skills for scientific research spanning bioinformatics, cheminformatics, data analysis, scientific writing, and literature search; currently a curated link collection, with plans to consolidate selected skills into a unified, clone-ready repository. Developed by InternScience, the open-source hub of the AI for Science Center at Shanghai AI Laboratory, which open-sources agents, LLMs/MLLMs, tools, and datasets to accelerate scientific discovery across disciplines.
|
||||
|
||||
## Contributing
|
||||
|
||||
Contributions are welcome. Please read the [contribution guidelines](CONTRIBUTING.md) before submitting a pull request.
|
||||
@@ -0,0 +1,390 @@
|
||||
---
|
||||
title: "Awesome LLM Scientific Discovery [](https://awesome.re)"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery/blob/eb19b47e/README.md
|
||||
upstream_sha: eb19b47e
|
||||
imported_at: 2026-07-03
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Awesome LLM Scientific Discovery [](https://awesome.re)
|
||||
|
||||
A curated list of pioneering research papers, tools, and resources at the intersection of Large Language Models (LLMs) and Scientific Discovery.
|
||||
|
||||
Survey: ***From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery.*** ([https://arxiv.org/abs/2505.13259])
|
||||
|
||||
The survey delineates the evolving role of LLMs in science through a three-level autonomy framework:
|
||||
* **Level 1: LLM as Tool:** LLMs augmenting human researchers for specific, well-defined tasks.
|
||||
* **Level 2: LLM as Analyst:** LLMs exhibiting greater autonomy in processing complex information and offering insights.
|
||||
* **Level 3: LLM as Scientist:** LLM-based systems autonomously conducting major research stages.
|
||||
|
||||
Below is a visual representation of this taxonomy:
|
||||
|
||||

|
||||
|
||||
We aim to provide a comprehensive overview for researchers, developers, and enthusiasts interested in this rapidly advancing field.
|
||||
|
||||
> **Last major update: 2026.07.** This refresh adds a large batch of 2025–2026 papers and a dedicated section on frontier industry-lab systems (Google DeepMind, OpenAI, Microsoft Research, Meta FAIR, FutureHouse, Sakana AI, and others). Contributions and PRs are very welcome — see [Contributing](#contributing).
|
||||
|
||||
## Contents
|
||||
|
||||
* [Level 1: LLM as Tool](#level-1-llm-as-tool)
|
||||
* [Literature Review and Information Gathering](#literature-review-and-information-gathering)
|
||||
* [Idea Generation and Hypothesis Formulation](#idea-generation-and-hypothesis-formulation)
|
||||
* [Experiment Planning and Execution](#experiment-planning-and-execution)
|
||||
* [Data Analysis and Organization](#data-analysis-and-organization)
|
||||
* [Conclusion and Hypothesis Validation](#conclusion-and-hypothesis-validation)
|
||||
* [Iteration and Refinement](#iteration-and-refinement)
|
||||
* [Level 2: LLM as Analyst](#level-2-llm-as-analyst)
|
||||
* [Machine Learning Research](#machine-learning-research)
|
||||
* [Data Modeling and Analysis](#data-modeling-and-analysis)
|
||||
* [Function Discovery](#function-discovery)
|
||||
* [Natural Science Research](#natural-science-research)
|
||||
* [General Research](#general-research)
|
||||
* [Survey Generation](#survey-generation)
|
||||
* [Level 3: LLM as Scientist](#level-3-llm-as-scientist)
|
||||
* [General-Purpose Autonomous Research Agents](#general-purpose-autonomous-research-agents)
|
||||
* [Discovery-Oriented Scientific Systems](#discovery-oriented-scientific-systems)
|
||||
* [Autonomous Research Ecosystems and Infrastructure](#autonomous-research-ecosystems-and-infrastructure)
|
||||
* [Frontier Labs and Foundation Models for Science](#frontier-labs-and-foundation-models-for-science)
|
||||
* [Other Related Works](#other-related-works)
|
||||
* [Contributing](#contributing)
|
||||
|
||||
---
|
||||
|
||||
## Level 1: LLM as Tool
|
||||
|
||||
At this foundational level, LLMs function as tailored tools under direct human supervision, designed to execute specific, well-defined tasks within a single stage of the scientific method. Their primary goal is to enhance researcher efficiency.
|
||||
|
||||
### Literature Review and Information Gathering
|
||||
|
||||
Automating literature search, retrieval, synthesis, structuring, and organization.
|
||||
|
||||
* **SCIMON : Scientific Inspiration Machines Optimized for Novelty** [](https://arxiv.org/pdf/2305.14259) - *Wang et al. (2023.05)*
|
||||
* **ResearchAgent: Iterative research idea generation over scientific literature with Large Language Models** [](https://arxiv.org/pdf/2404.07738) - *Baek et al. (2024.04)*
|
||||
* **Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction** [](https://arxiv.org/pdf/2404.14215) - *Deng et al. (2024.04)*
|
||||
* **TKGT: Redefinition and A New Way of text-to-table tasks based on real world demands and knowledge graphs augmented LLMs** [](https://aclanthology.org/2024.emnlp-main.901.pdf) - *Jiang et al. (2024.10)*
|
||||
* **ArxivDIGESTables: Synthesizing scientific literature into tables using language models** [](https://arxiv.org/pdf/2410.22360) - *Newman et al. (2024.10)*
|
||||
* **Can LLMs Generate Tabular Summaries of Science Papers? Rethinking the Evaluation Protocol** [](https://arxiv.org/pdf/2504.10284) - *Wang et al. (2025.04)*
|
||||
* **LitLLM: A Toolkit for Scientific Literature Review** [](https://arxiv.org/pdf/2402.01788v1) - *Agarwal et al. (2024.02)*
|
||||
* **Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain** [](https://systematicreviewsjournal.biomedcentral.com/articles/10.1186/s13643-024-02575-4) - *Dennstädt et al. (2024.06)*
|
||||
* **Science Hierarchography: Hierarchical Organization of Science Literature** [](https://arxiv.org/pdf/2504.13834) - *Gao et al. (2025.04)*
|
||||
* **Language Agents Achieve Superhuman Synthesis of Scientific Knowledge (PaperQA2)** [](https://arxiv.org/pdf/2409.13740) - *Skarlinski et al. (2024.09)*
|
||||
* **DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents** [](https://arxiv.org/pdf/2506.11763) - *Du et al. (2025.06)*
|
||||
* **Deep Research Agents: A Systematic Examination And Roadmap** [](https://arxiv.org/pdf/2506.18096) - *Huang et al. (2025.06)*
|
||||
* **DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis** [](https://arxiv.org/pdf/2508.20033) - *Patel et al. (2025.08)*
|
||||
* **ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry** [](https://arxiv.org/pdf/2507.16280) - *Xu et al. (2025.07)*
|
||||
* **LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation** [](https://arxiv.org/pdf/2510.05138) - *Zhang et al. (2025.10)*
|
||||
|
||||
### Idea Generation and Hypothesis Formulation
|
||||
|
||||
Automated generation of novel research ideas, conceptual insights, and testable scientific hypotheses.
|
||||
|
||||
* **SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning** [](https://arxiv.org/pdf/2409.05556) - *Ghafarollahi et al. (2024.09)*
|
||||
* **Accelerating scientific discovery with generative knowledge extraction, graph-based representation, and multimodal intelligent graph reasoning** [](https://arxiv.org/pdf/2403.11996) - *Buehler (2024.03)*
|
||||
* **MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses** [](https://arxiv.org/pdf/2410.07076) - *Yang et al. (2024.10)*
|
||||
* **Large Language Models for Automated Open-domain Scientific Hypotheses Discovery** [](https://arxiv.org/pdf/2309.02726) - *Yang et al. (2023.09)*
|
||||
* **Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models** [](https://arxiv.org/pdf/2411.02382) - *Xiong et al. (2024.11)*
|
||||
* **ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition** [](https://arxiv.org/pdf/2503.21248) - *Liu et al. (2025.03)*
|
||||
* **AI Idea Bench 2025: AI Research Idea Generation Benchmark** [](https://arxiv.org/pdf/2504.14191) - *Qiu et al. (2025.04)*
|
||||
* **IdeaBench: Benchmarking Large Language Models for Research Idea Generation** [](https://arxiv.org/pdf/2411.02429) - *Guo et al. (2024.11)*
|
||||
* **Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers** [](https://arxiv.org/pdf/2409.04109) - *Si et al. (2024.09)*
|
||||
* **Learning to Generate Research Idea with Dynamic Control** [](https://arxiv.org/pdf/2412.14626) - *Li et al. (2024.12)*
|
||||
* **LiveIdeaBench: Evaluating LLMs' Divergent Thinking for Scientific Idea Generation with Minimal Context** [](https://arxiv.org/pdf/2412.17596) - *Ruan et al. (2024.12)*
|
||||
* **Nova: An Iterative Planning and Search Approach to Enhance Novelty and Diversity of LLM Generated Ideas** [](https://arxiv.org/pdf/2410.14255) - *Hu et al. (2024.10)*
|
||||
* **GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation** [](https://arxiv.org/pdf/2503.12600) - *Feng et al. (2025.03)*
|
||||
* **Hypothesis Generation with Large Language Models** [](https://arxiv.org/pdf/2404.04326) - *Zhou et al. (2024.04)*
|
||||
* **Harnessing the Power of Adversarial Prompting and Large Language Models for Robust Hypothesis Generation in Astronomy** [](https://arxiv.org/pdf/2306.11648) - *Ciuca et al. (2023.06)*
|
||||
* **Large Language Models are Zero Shot Hypothesis Proposers** [](https://arxiv.org/pdf/2311.05965) - *Qi et al. (2023.11)*
|
||||
* **Machine learning for hypothesis generation in biology and medicine: exploring the latent space of neuroscience and developmental bioelectricity** [](https://pubs.rsc.org/en/content/articlelanding/2024/dd/d3dd00185g) - *O’Brien et al. (2023.07)*
|
||||
* **Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation** [](https://arxiv.org/pdf/2407.08940) - *Qi et al. (2024.07)*
|
||||
* **LLM4GRN: Discovering Causal Gene Regulatory Networks with LLMs -- Evaluation through Synthetic Data Generation** [](https://arxiv.org/pdf/2410.15828) - *Afonja et al. (2024.10)*
|
||||
* **Scideator: Human-LLM Scientific Idea Generation Grounded in Research-Paper Facet Recombination** [](https://arxiv.org/pdf/2409.14634) - *Radensky et al. (2024.09)*
|
||||
* **HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance** [](https://arxiv.org/pdf/2506.12937) - *Vasu et al. (2025.06)*
|
||||
* **Sparks of Science: Hypothesis Generation Using Structured Paper Data** [](https://arxiv.org/pdf/2504.12976) - *O'Neill et al. (2025.04)*
|
||||
* **A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models** [](https://arxiv.org/pdf/2504.05496) - *Kulkarni et al. (2025.04)*
|
||||
* **Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Networks** [](https://arxiv.org/pdf/2511.02238) - *Wang et al. (2025.11)*
|
||||
|
||||
### Experiment Planning and Execution
|
||||
|
||||
LLMs assisting in experimental protocol planning, workflow design, and scientific code generation.
|
||||
|
||||
* **BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology** [](https://arxiv.org/pdf/2310.10632) - *O'Donoghue et al. (2023.10)*
|
||||
* **Can Large Language Models Help Experimental Design for Causal Discovery?** (Li et al. in survey) [](https://arxiv.org/pdf/2503.01139) - *Li et al. (2025.03)*
|
||||
* **Hierarchically Encapsulated Representation for Protocol Design in Self-Driving Labs** [](https://arxiv.org/pdf/2504.03810) - *Shi et al. (2025.04)*
|
||||
* **SciCode: A Research Coding Benchmark Curated by Scientists** [](https://arxiv.org/pdf/2407.13168) - *Tian et al. (2024.07)*
|
||||
* **Natural Language to Code Generation in Interactive Data Science Notebooks** [](https://arxiv.org/pdf/2212.09248) - *Yin et al. (2022.12)*
|
||||
* **DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation** [](https://arxiv.org/pdf/2211.11501) - *Lai et al. (2022.11)*
|
||||
* **Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents**, [](https://arxiv.org/pdf/2502.16069) - *Kon et al. (2025.02)*
|
||||
* **AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing** [](https://arxiv.org/pdf/2602.17607) - *Du et al. (2026.02)*
|
||||
|
||||
### Data Analysis and Organization
|
||||
|
||||
LLMs assisting in data-driven analysis, tabular/chart reasoning, statistical reasoning, and model discovery.
|
||||
|
||||
* **AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ** [](https://arxiv.org/pdf/2310.00367) - *Belouadi et al. (2023.10)*
|
||||
* **Text2Chart31: Instruction Tuning for Chart Generation with Automatic Feedback** [](https://arxiv.org/pdf/2410.04064) - *Zadeh et al. (2024.10)*
|
||||
* **ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning** [](https://arxiv.org/pdf/2203.10244) - *Masry et al. (2022.03)*
|
||||
* **CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs** [](https://arxiv.org/pdf/2406.18521) - *Wang et al. (2024.06)*
|
||||
* **ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning** [](https://arxiv.org/pdf/2402.12185) - *Xia et al. (2024.02)*
|
||||
* **Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding** [](https://arxiv.org/pdf/2401.04398) - *Wang et al. (2024.01)*
|
||||
* **TableBench: A Comprehensive and Complex Benchmark for Table Question Answering** [](https://arxiv.org/pdf/2408.09174) - *Wu et al. (2024.08)*
|
||||
* **Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs** [](https://arxiv.org/pdf/2402.12424) - *Deng et al. (2024.02)*
|
||||
* **ChatSpatial: Schema-Enforced Agentic Orchestration for Reproducible and Cross-Platform Spatial Transcriptomics** [](https://doi.org/10.64898/2026.02.26.708361) - *Yang et al. (2026.02)* [Code](https://github.com/cafferychen777/ChatSpatial)
|
||||
|
||||
### Conclusion and Hypothesis Validation
|
||||
|
||||
LLMs providing feedback, verifying claims, replicating results, and generating reviews.
|
||||
|
||||
* **CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?** [](https://arxiv.org/pdf/2503.21717) - *Ou et al. (2025.03)*
|
||||
* **LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing** [](https://arxiv.org/pdf/2406.16253) - *Du et al. (2024.06)*
|
||||
* **AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews** [](https://arxiv.org/pdf/2408.10365) - *Tyser et al. (2024.08)*
|
||||
* **Is LLM a Reliable Reviewer? A Comprehensive Evaluation of LLM on Automatic Paper Reviewing Tasks** [](https://aclanthology.org/2024.lrec-main.816.pdf) - *Zhou et al. (2024.05)*
|
||||
* **ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing** [](https://arxiv.org/pdf/2306.00622) - *Liu and Shah (2023.06)*
|
||||
* **Towards Autonomous Hypothesis Verification via Language Models with Minimal Guidance** [](https://arxiv.org/pdf/2311.09706) - *Takagi et al. (2023.11)*
|
||||
* **CycleResearcher: Improving Automated Research via Automated Review** [](https://arxiv.org/pdf/2411.00816) - *Weng et al. (2024.11)*
|
||||
* **PaperBench: Evaluating AI’s Ability to Replicate AI Research** [](https://arxiv.org/pdf/2504.01848) - *Starace et al. (2025.04)*
|
||||
* **SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers** [](https://arxiv.org/pdf/2504.00255) - *Xiang et al. (2025.04)*
|
||||
* **Advancing AI-Scientist Understanding: Making LLM Think Like a Physicist with Interpretable Reasoning** [](https://arxiv.org/pdf/2504.01911) - *Xu et al. (2025.04)*
|
||||
* **Generative Adversarial Reviews: When LLMs Become the Critic** [](https://arxiv.org/pdf/2412.10415) - *Bougie & Watanabe (2024.12)*
|
||||
* **Predicting Empirical AI Research Outcomes with Language Models** [](https://arxiv.org/pdf/2506.00794) - *Wen et al. (2025.06)*
|
||||
* **SPOT: When AI Co-Scientists Fail — A Benchmark for Automated Verification of Scientific Research** [](https://arxiv.org/pdf/2505.11855) - *Son et al. (2025.05)*
|
||||
* **DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process** [](https://arxiv.org/pdf/2503.08569) - *Zhu et al. (2025.03)*
|
||||
* **ReviewRL: Towards Automated Scientific Review with RL** [](https://arxiv.org/pdf/2508.10308) - *Zeng et al. (2025.08)*
|
||||
* **SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification** [](https://arxiv.org/pdf/2502.10003) - *Kumar et al. (2025.02)*
|
||||
* **LMR-Bench: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research** [](https://arxiv.org/pdf/2506.17335) - *Yan et al. (2025.06)*
|
||||
* **REFUTE: Reasoning Over Evidence - Falsification, Uncertainty, Truth-grounding & Epistemics** [](https://huggingface.co/datasets/BGPT-OFFICIAL/refute) - *BGPT (2026.06)*. Open benchmark for scientific critique and epistemic calibration on recent science paper summaries, covering falsification, limitations, overclaims, missing-evidence refusal, calibration, and planted-flaw detection.
|
||||
|
||||
### Iteration and Refinement
|
||||
|
||||
LLMs involved in iterative refinement of research hypotheses and strategic exploration.
|
||||
|
||||
* **Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving** [](https://arxiv.org/pdf/2405.01379) - *Quan et al. (2024.05)*
|
||||
* **Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM Agents** [](https://arxiv.org/pdf/2410.13185) - *Li et al. (2024.10)*
|
||||
* **Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees** [](https://arxiv.org/pdf/2503.19309) - *Rabby et al. (2025.03)*
|
||||
* **XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision** [](https://arxiv.org/pdf/2505.11336) - *Chen et al. (2025.05)*
|
||||
|
||||
---
|
||||
|
||||
## Level 2: LLM as Analyst
|
||||
|
||||
LLMs exhibiting a greater degree of autonomy, functioning as passive agents capable of complex information processing, data modeling, and analytical reasoning with reduced human intervention.
|
||||
|
||||
### Machine Learning Research
|
||||
|
||||
Automated modeling of machine learning tasks, experiment design, and execution.
|
||||
|
||||
* **MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation** [](https://arxiv.org/pdf/2310.03302) - *Huang et al. (2023.10)*
|
||||
* **MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents** [](https://arxiv.org/pdf/2408.14033) - *Li et al. (2024.08)*
|
||||
* **MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering** [](https://arxiv.org/pdf/2410.07095) - *Chan et al. (2024.10)*
|
||||
* **IMPROVE: Iterative Model Pipeline Refinement and Optimization Leveraging LLM Agents** [](https://arxiv.org/pdf/2502.18530v1) - *Xue et al. (2025.02)*
|
||||
* **CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation** [](https://arxiv.org/pdf/2503.22708) - *Jansen et al. (2025.03)*
|
||||
* **MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?** [](https://arxiv.org/pdf/2504.09702) - *Zhang et al. (2025.04)*
|
||||
* **RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts** [](https://arxiv.org/pdf/2411.15114) - *Wijk et al. (2024.11)*
|
||||
* **MLZero: A Multi-Agent System for End-to-end Machine Learning Automation** [](https://arxiv.org/pdf/2505.13941) - *Fang et al. (2025.05)*
|
||||
* **AIDE: AI-Driven Exploration in the Space of Code** [](https://arxiv.org/pdf/2502.13138) - *Jiang et al. (2025.02)*
|
||||
* **Language Modeling by Language Models** [](https://arxiv.org/pdf/2506.20249) - *Cheng et al. (2025.06)*
|
||||
* **MLGym: A New Framework and Benchmark for Advancing AI Research Agents** [](https://arxiv.org/pdf/2502.14499) - *Nathani et al. (2025.02)*
|
||||
* **R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science** [](https://arxiv.org/pdf/2505.14738) - *Xu et al. (2025.05)* — Microsoft Research
|
||||
* **MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement** [](https://arxiv.org/pdf/2506.15692) - *Nam et al. (2025.06)* — Google
|
||||
* **ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning** [](https://arxiv.org/pdf/2506.16499) - *Liu et al. (2025.06)*
|
||||
* **ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering** [](https://arxiv.org/pdf/2505.23723) - *Liu et al. (2025.05)*
|
||||
* **AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench** [](https://arxiv.org/pdf/2507.02554) - *Toledo et al. (2025.07)* — Meta / UCL
|
||||
* **The FM Agent** [](https://arxiv.org/pdf/2510.26144) - *Li et al. (2025.10)*
|
||||
* **KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for ML Problems** [](https://arxiv.org/pdf/2508.10177) - *Kulibaba et al. (2025.08)*
|
||||
* **AutoMLGen: Navigating Fine-Grained Optimization for Coding Agents** [](https://arxiv.org/pdf/2510.08511) - *Du et al. (2025.10)*
|
||||
* **ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies** [](https://arxiv.org/pdf/2504.20117) - *Gandhi et al. (2025.04)*
|
||||
* **AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML** [](https://arxiv.org/pdf/2410.02958) - *Trirat et al. (2024.10)*
|
||||
* **SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning** [](https://arxiv.org/pdf/2410.17238) - *Chi et al. (2024.10)*
|
||||
* **AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions** [](https://arxiv.org/pdf/2410.20424) - *Li et al. (2024.10)*
|
||||
* **Agent K: Kolb-Based Experiential Learning for Generalist Agents with Human-Level Kaggle Performance** [](https://arxiv.org/pdf/2411.03562) - *Grosnit et al. (2024.11)* — Huawei Noah's Ark
|
||||
* **EXP-Bench: Can AI Conduct AI Research Experiments?** [](https://arxiv.org/pdf/2505.24785) - *Kon et al. (2025.05)*
|
||||
* **InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research** [](https://arxiv.org/pdf/2510.27598) - *Wu et al. (2025.10)*
|
||||
* **MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research** [](https://arxiv.org/pdf/2505.19955) - *Chen et al. (2025.05)*
|
||||
* **RExBench: Can Coding Agents Autonomously Implement AI Research Extensions?** [](https://arxiv.org/pdf/2506.22598) - *Edwards et al. (2025.06)*
|
||||
* **ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program Evolution** [](https://arxiv.org/pdf/2509.19349) - *Lange et al. (2025.09)* — Sakana AI
|
||||
* **The AI CUDA Engineer: Agentic CUDA Kernel Discovery, Optimization and Composition** [](https://pub.sakana.ai/ai-cuda-engineer/paper/) - *Sakana AI (2025.02)*
|
||||
|
||||
### Data Modeling and Analysis
|
||||
|
||||
Automated data-driven analysis, statistical data modeling, and hypothesis validation.
|
||||
|
||||
* **Automated Statistical Model Discovery with Language Models** [](https://arxiv.org/pdf/2402.17879) - *Li et al. (2024.02)*
|
||||
* **InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks** [](https://arxiv.org/pdf/2401.05507) - *Hu et al. (2024.01)*
|
||||
* **DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning** [](https://arxiv.org/pdf/2402.17453) - *Guo et al. (2024.02)*
|
||||
* **BLADE: Benchmarking Language Model Agents for Data-Driven Science** [](https://arxiv.org/pdf/2408.09667) - *Gu et al. (2024.08)*
|
||||
* **DAgent: A Relational Database-Driven Data Analysis Report Generation Agent** [](https://arxiv.org/pdf/2503.13269) - *Xu et al. (2025.03)*
|
||||
* **DiscoveryBench: Towards Data-Driven Discovery with Large Language Models** [](https://arxiv.org/pdf/2407.01725) - *Majumder et al. (2024.07)*
|
||||
* **Large Language Models for Scientific Synthesis, Inference and Explanation** [](https://arxiv.org/pdf/2310.07984) - *Zheng et al. (2023.10)*
|
||||
* **MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem** [](https://arxiv.org/pdf/2505.14148) - *Liu et al. (2025.05)*
|
||||
* **DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?** [](https://arxiv.org/pdf/2409.07703) - *Jing et al. (2024.09)*
|
||||
* **AutoDS: Open-ended Scientific Discovery via Bayesian Surprise** [](https://arxiv.org/pdf/2507.00310) - *Agarwal et al. (2025.07)* — Allen Institute for AI
|
||||
* **DeepAnalyze: Agentic Large Language Models for Autonomous Data Science** [](https://arxiv.org/pdf/2510.16872) - *Zhang et al. (2025.10)*
|
||||
* **DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models** [](https://arxiv.org/pdf/2410.07331) - *Huang et al. (2024.10)*
|
||||
* **DataSciBench: An LLM Agent Benchmark for Data Science** [](https://arxiv.org/pdf/2502.13897) - *Zhang et al. (2025.02)*
|
||||
* **Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents** [](https://arxiv.org/pdf/2403.05307) - *Li et al. (2024.03)*
|
||||
* **StatEval: A Comprehensive Benchmark for Large Language Models in Statistics** [](https://arxiv.org/pdf/2510.09517) - *Yu et al. (2025.10)*
|
||||
* **LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference** [](https://arxiv.org/pdf/2508.07221) - *Wang et al. (2025.08)*
|
||||
* **OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents** [](https://arxiv.org/pdf/2504.16918) - *Thind et al. (2025.04)*
|
||||
|
||||
### Function Discovery
|
||||
|
||||
Identifying underlying equations from observational data (AI-driven symbolic regression).
|
||||
|
||||
* **LLM-SR: Scientific Equation Discovery via Programming with Large Language Models** [](https://arxiv.org/pdf/2404.18400) - *Shojaee et al. (2024.04)*
|
||||
* **LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models** [](https://arxiv.org/pdf/2504.10415) - *Shojaee et al. (2025.04)*
|
||||
* **Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents** [](https://arxiv.org/pdf/2501.18411) - *Koblischke et al. (2025.01)*
|
||||
* **EvoSLD: Automated neural scaling law discovery with large language models** [](https://arxiv.org/abs/2507.21184) - *Lin et al. (2025.07)*
|
||||
* **DrSR: LLM based Scientific Equation Discovery with Dual Reasoning from Data and Experience** [](https://arxiv.org/abs/2506.04282) - *Wang et al. (2025.06)*
|
||||
* **NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents** [](https://arxiv.org/pdf/2510.07172) - *Zheng et al. (2025.10)*
|
||||
* **LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery** [](https://arxiv.org/pdf/2503.06512) - *Song et al. (2025.03)*
|
||||
* **SR-Scientist: Scientific Equation Discovery With Agentic AI** [](https://arxiv.org/pdf/2510.11661) - *Xia et al. (2025.10)*
|
||||
* **LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery (SGA)** [](https://arxiv.org/pdf/2405.09783) - *Ma et al. (2024.05)*
|
||||
* **In-Context Symbolic Regression: Leveraging Large Language Models for Function Discovery** [](https://arxiv.org/pdf/2404.19094) - *Merler et al. (2024.04)*
|
||||
* **Symbolic Regression with a Learned Concept Library (LaSR)** [](https://arxiv.org/pdf/2409.09359) - *Grayeli et al. (2024.09)*
|
||||
* **AI-Newton: A Concept-Driven Physical Law Discovery System without Prior Physical Knowledge** [](https://arxiv.org/pdf/2504.01538) - *Fang et al. (2025.04)*
|
||||
* **PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors** [](https://arxiv.org/pdf/2507.15550) - *Chen et al. (2025.07)*
|
||||
* **Finetuning Large Language Model as an Effective Symbolic Regressor (SymbArena)** [](https://arxiv.org/pdf/2508.09897) - *Hua et al. (2025.08)*
|
||||
|
||||
### Natural Science Research
|
||||
|
||||
Autonomous research workflows for natural science discovery (e.g., chemistry, biology, biomedicine, materials, physics).
|
||||
|
||||
* **Coscientist: Autonomous Chemical Research with Large Language Models** [](https://www.nature.com/articles/s41586-023-06792-0) - *Boiko et al. (2023.10)*
|
||||
* **Empowering biomedical discovery with AI agents** [](https://www.cell.com/action/showPdf?pii=S0092-8674%2824%2901070-5) - *Gao et al. (2024.09)*
|
||||
* **GenoTEX: An LLM Agent Benchmark for Automated Gene Expression Data Analysis** [](https://arxiv.org/pdf/2406.15341) - *Liu et al. (2024.06)*
|
||||
* **From Intention To Implementation: Automating Biomedical Research via LLMs** [](https://arxiv.org/pdf/2412.09429) - *Luo et al. (2024.12)*
|
||||
* **DrugAgent: Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration** [](https://arxiv.org/pdf/2411.15692) - *Liu et al. (2024.11)*
|
||||
* **ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery** [](https://arxiv.org/pdf/2410.05080) - *Chen et al. (2024.10)*
|
||||
* **ProtAgents: Protein discovery by combining physics and machine learning** [](https://arxiv.org/pdf/2402.04268) - *Ghafarollahi and Buehler (2024.02)*
|
||||
* **Auto-Bench: An Automated Benchmark for Scientific Discovery in LLMs** [](https://arxiv.org/pdf/2502.15224) - *Chen et al. (2025.02)*
|
||||
* **Towards an AI co-scientist** [](https://arxiv.org/pdf/2502.18864) - *Gottweis et al. (2025.02)* — Google
|
||||
* **GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis** [](https://arxiv.org/pdf/2507.21035) - *Liu et al. (2025.07)*
|
||||
* **Automated Algorithmic Discovery for Gravitational-Wave Detection Guided by LLM-Informed Evolutionary Monte Carlo Tree Search** [](https://arxiv.org/pdf/2508.03661) - *Wang and Zeng (2025.08)*
|
||||
* **The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies** [](https://doi.org/10.1038/s41586-025-09442-9) - *Swanson et al. (2025.07)* — Stanford / CZ Biohub
|
||||
* **Biomni: A General-Purpose Biomedical AI Agent** [](https://doi.org/10.1101/2025.05.30.656746) - *Huang et al. (2025.05)* — Stanford
|
||||
* **TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools** [](https://arxiv.org/pdf/2503.10970) - *Gao et al. (2025.03)*
|
||||
* **LIDDiA: Language-based Intelligent Drug Discovery Agent** [](https://arxiv.org/pdf/2502.13959) - *Averly et al. (2025.02)*
|
||||
* **LLM Agent Swarm for Hypothesis-Driven Drug Discovery (PharmaSwarm)** [](https://arxiv.org/pdf/2504.17967) - *Song et al. (2025.04)*
|
||||
* **BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation** [](https://arxiv.org/pdf/2508.01285) - *Ke et al. (2025.08)*
|
||||
* **CRISPR-GPT for Agentic Automation of Gene-editing Experiments** [](https://arxiv.org/pdf/2404.18021) - *Qu et al. (2024.04)*
|
||||
* **BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments** [](https://arxiv.org/pdf/2405.17631) - *Roohani et al. (2024.05)*
|
||||
* **CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis** [](https://arxiv.org/pdf/2407.09811) - *Xiao et al. (2024.07)*
|
||||
* **Training a Scientific Reasoning Model for Chemistry (ether0)** [](https://arxiv.org/pdf/2506.17238) - *Narayanan et al. (2025.06)* — FutureHouse
|
||||
* **AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation** [](https://arxiv.org/pdf/2509.25651) - *Panapitiya et al. (2025.09)*
|
||||
* **LLMatDesign: Autonomous Materials Discovery with Large Language Models** [](https://arxiv.org/pdf/2406.13163) - *Jia et al. (2024.06)*
|
||||
* **Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists** [](https://arxiv.org/pdf/2506.05616) - *Zhou et al. (2025.06)*
|
||||
* **SparksMatter: Autonomous Inorganic Materials Discovery via Multi-Agent Physics-Aware Scientific Reasoning** [](https://arxiv.org/pdf/2508.02956) - *Ghafarollahi et al. (2025.08)*
|
||||
* **Swarms of Large Language Model Agents for Protein Sequence Design with Experimental Validation** [](https://arxiv.org/pdf/2511.22311) - *Wang et al. (2025.11)*
|
||||
* **BixBench: A Comprehensive Benchmark for LLM-based Agents in Computational Biology** [](https://arxiv.org/pdf/2503.00096) - *Mitchener et al. (2025.02)* — FutureHouse
|
||||
* **CASSIA: a multi-agent large language model for automated and interpretable cell annotation** [](https://www.nature.com/articles/s41467-025-67084-x) - *Xie et al. (2025.12)*
|
||||
* **AutoZyme: An Autonomous Agentic Framework to Optimize Bioinformatics Software** [](https://www.biorxiv.org/content/10.64898/2026.06.12.731250v1) - *Xie et al. (2026.06)*
|
||||
|
||||
### General Research
|
||||
|
||||
Benchmarks and frameworks evaluating diverse tasks from different stages of scientific discovery.
|
||||
|
||||
* **DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents** [](https://arxiv.org/pdf/2406.06769) - *Jansen et al. (2024.06)*
|
||||
* **A Vision for Auto Research with LLM Agents** [](https://arxiv.org/pdf/2504.18765) - *Liu et al. (2025.04)*
|
||||
* **Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents** [](https://arxiv.org/abs/2502.16069) - *Kon et al. (2025.02)*
|
||||
* **EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants** [](https://arxiv.org/pdf/2502.20309) - *Cappello et al. (2025.02)*
|
||||
|
||||
### Survey Generation
|
||||
|
||||
* **AutoSurvey: Large Language Models Can Automatically Write Surveys** [](https://arxiv.org/pdf/2406.10252) - *Wang et al. (2024.06)*
|
||||
* **SurveyX: Academic Survey Automation via Large Language Models** [](https://arxiv.org/pdf/2502.14776) - *Liang et al. (2025.02)*
|
||||
|
||||
---
|
||||
|
||||
## Level 3: LLM as Scientist
|
||||
|
||||
LLM-based systems operating as active agents capable of orchestrating and navigating multiple stages of the scientific discovery process with considerable independence, often culminating in draft research papers or genuine new findings. As the field has matured, these systems increasingly fall into distinct classes, reflected in the sub-sections below.
|
||||
|
||||
### General-Purpose Autonomous Research Agents
|
||||
|
||||
End-to-end pipelines that autonomously move from ideation through experimentation to a full paper draft, typically domain-agnostic (frequently demonstrated on ML/AI research).
|
||||
|
||||
* **Agent Laboratory: Using LLM Agents as Research Assistants** [](https://arxiv.org/pdf/2501.04227) - *Schmidgall et al. (2025.01)*
|
||||
* **The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery** [](https://arxiv.org/pdf/2408.06292) - *Lu et al. (2024.08)* — Sakana AI
|
||||
* **The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search** [](https://arxiv.org/pdf/2504.08066) - *Yamada et al. (2025.04)* — Sakana AI
|
||||
* **AI-Researcher: Autonomous Scientific Innovation** [](https://arxiv.org/pdf/2505.18705) [](https://github.com/HKUDS/AI-Researcher) - *Tang et al. (2025.05)*
|
||||
* **Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback** [](https://arxiv.org/pdf/2501.03916) - *Yuan et al. (2025.01)*
|
||||
* **NovelSeek / InternAgent: When Agent Becomes the Scientist — Building a Closed-Loop System from Hypothesis to Verification** [](https://arxiv.org/pdf/2505.16938) - *InternAgent Team (2025.05)* — Shanghai AI Lab
|
||||
* **The Denario Project: Deep Knowledge AI Agents for Scientific Discovery** [](https://arxiv.org/pdf/2510.26887) - *Villaescusa-Navarro et al. (2025.10)*
|
||||
* **Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation (freephdlabor)** [](https://arxiv.org/pdf/2510.15624) - *Li et al. (2025.10)*
|
||||
* **AIGS: Generating Science from AI-Powered Automated Falsification** [](https://arxiv.org/pdf/2411.11910) - *Liu et al. (2024.11)*
|
||||
* **Zochi Technical Report** [](https://www.intology.ai/blog/zochi-tech-report) - *Intology AI (2025.03)*
|
||||
* **Meet Carl: The First AI System To Produce Academically Peer-Reviewed Research** [](https://www.autoscience.ai/blog/meet-carl-the-first-ai-system-to-produce-academically-peer-reviewed-research) - *Autoscience Institute (2025.03)*
|
||||
* **DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively** [](https://arxiv.org/pdf/2509.26603) - *Weng et al. (2025.09)*
|
||||
* **Accelerating Social Science Research via Agentic Hypothesization and Experimentation** [](https://arxiv.org/pdf/2602.07983) - *Gupta et al. (2026.02)*
|
||||
* **AI-Researcher: Fully-Automated Scientific Discovery with LLM Agents** [](https://github.com/HKUDS/AI-Researcher) - *Data Intelligence Lab (2025.03)*
|
||||
|
||||
### Discovery-Oriented Scientific Systems
|
||||
|
||||
Systems whose primary goal is genuine new scientific knowledge — novel, experimentally- or mathematically-validated findings — rather than paper drafts. Many are frontier industry-lab systems.
|
||||
|
||||
* **Kosmos: An AI Scientist for Autonomous Discovery** [](https://arxiv.org/pdf/2511.02824) - *Mitchener et al. (2025.11)* — Edison Scientific / FutureHouse
|
||||
* **Robin: A Multi-Agent System for Automating Scientific Discovery** [](https://arxiv.org/pdf/2505.13400) - *Ghareeb et al. (2025.05)* — FutureHouse
|
||||
* **AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery** [](https://arxiv.org/pdf/2506.13131) - *Novikov et al. (2025.06)* — Google DeepMind
|
||||
* **Aviary: Training Language Agents on Challenging Scientific Tasks** [](https://arxiv.org/pdf/2412.21154) - *Narayanan et al. (2024.12)* — FutureHouse
|
||||
|
||||
### Autonomous Research Ecosystems and Infrastructure
|
||||
|
||||
Platforms and protocols that enable multiple AI scientists to collaborate, share, review, and publish — moving beyond a single agent toward a research ecosystem.
|
||||
|
||||
* **AgentRxiv: Towards Collaborative Autonomous Research** [](https://arxiv.org/pdf/2503.18102) - *Schmidgall et al. (2025.03)*
|
||||
* **aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists** [](https://arxiv.org/pdf/2508.15126) - *Zhang et al. (2025.08)*
|
||||
|
||||
---
|
||||
|
||||
## Frontier Labs and Foundation Models for Science
|
||||
|
||||
Flagship "AI for Science" systems from frontier industry labs. Unlike the agentic systems catalogued above, most of these are large domain-specific foundation models or specialized reasoning systems that have driven headline scientific results (structure prediction, materials/genome design, olympiad-level mathematics). They are included here as essential context for the broader landscape of AI-accelerated discovery.
|
||||
|
||||
* **Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3** [](https://www.nature.com/articles/s41586-024-07487-w) - *Abramson et al. (2024.05)* — Google DeepMind / Isomorphic Labs
|
||||
* **Scaling Deep Learning for Materials Discovery (GNoME)** [](https://www.nature.com/articles/s41586-023-06735-9) - *Merchant et al. (2023.11)* — Google DeepMind
|
||||
* **A Generative Model for Inorganic Materials Design (MatterGen)** [](https://www.nature.com/articles/s41586-025-08628-5) - *Zeni et al. (2025.01)* — Microsoft Research
|
||||
* **Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models** [](https://arxiv.org/pdf/2410.12771) - *Barroso-Luque et al. (2024.10)* — Meta FAIR
|
||||
* **TamGen: Drug Design with Target-Aware Molecule Generation through a Chemical Language Model** [](https://www.nature.com/articles/s41467-024-53632-4) - *Wu et al. (2024.10)* — Microsoft Research
|
||||
* **Genome Modeling and Design Across All Domains of Life with Evo 2** [](https://www.biorxiv.org/content/10.1101/2025.02.18.638918v1) - *Brixi et al. (2025.02)* — Arc Institute / NVIDIA
|
||||
* **Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning (AlphaProof)** [](https://www.nature.com/articles/s41586-025-09833-y) - *Hubert et al. (2025.11)* — Google DeepMind
|
||||
* **Gold-Medalist Performance in Solving Olympiad Geometry with AlphaGeometry 2** [](https://arxiv.org/pdf/2502.03544) - *Chervonyi et al. (2025.02)* — Google DeepMind
|
||||
* **Chai-2: Drug-Like Antibody Design Against Challenging Targets with Atomic Precision** [](https://chaiassets.com/chai-2/paper/technical_report_challenging_targets.pdf) - *Chai Discovery (2025.11)*
|
||||
|
||||
---
|
||||
|
||||
## Other Related Works
|
||||
|
||||
* **NVIDIA BioNeMo Agent Toolkit — Tools for Agents to Accelerate Scientific Discovery** [](https://nvidianews.nvidia.com/news/nvidia-launches-bionemo-agent-toolkit-giving-ai-agents-the-tools-to-accelerate-scientific-discovery) - *NVIDIA (2025)*
|
||||
|
||||
|
||||
---
|
||||
## Contributing
|
||||
|
||||
Contributions are welcome! If you have a paper, tool, or resource that fits into this taxonomy, please submit a **pull request**.
|
||||
|
||||
When adding an entry, please:
|
||||
* Place it under the most appropriate level/sub-section.
|
||||
* Keep the format consistent: `**Title** [badge](link) - *First author et al. (YYYY.MM)*`.
|
||||
* Verify the arXiv ID / DOI resolves, and add the lab/affiliation when the work comes from an industry group.
|
||||
|
||||
---
|
||||
|
||||
## Citation
|
||||
|
||||
Please cite our paper if you found our survey helpful:
|
||||
```bibtex
|
||||
@misc{zheng2025automationautonomysurveylarge,
|
||||
title={From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery},
|
||||
author={Tianshi Zheng and Zheye Deng and Hong Ting Tsang and Weiqi Wang and Jiaxin Bai and Zihao Wang and Yangqiu Song},
|
||||
year={2025},
|
||||
eprint={2505.13259},
|
||||
archivePrefix={arXiv},
|
||||
primaryClass={cs.CL},
|
||||
url={https://arxiv.org/abs/2505.13259},
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,207 @@
|
||||
---
|
||||
title: "Contributing to Awesome AI for Science"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/ai-boost/awesome-ai-for-science/blob/cf292eeb/CONTRIBUTING.md
|
||||
upstream_sha: cf292eeb
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Contributing to Awesome AI for Science
|
||||
|
||||
Thank you for your interest in contributing to Awesome AI for Science! 🎉
|
||||
|
||||
This project aims to be the most comprehensive and up-to-date collection of AI resources for scientific research. Your contributions help researchers worldwide discover tools and knowledge that can accelerate scientific discovery.
|
||||
|
||||
## 📋 Table of Contents
|
||||
|
||||
- [How to Contribute](#how-to-contribute)
|
||||
- [Types of Contributions](#types-of-contributions)
|
||||
- [Contribution Guidelines](#contribution-guidelines)
|
||||
- [Formatting Guidelines](#formatting-guidelines)
|
||||
- [Review Process](#review-process)
|
||||
- [Code of Conduct](#code-of-conduct)
|
||||
|
||||
## 🤝 How to Contribute
|
||||
|
||||
### Quick Contribution (for small additions)
|
||||
1. **Fork** this repository
|
||||
2. **Edit** the README.md file directly in GitHub
|
||||
3. **Add** your resource in the appropriate section
|
||||
4. **Submit** a pull request
|
||||
|
||||
### Detailed Contribution (for larger changes)
|
||||
1. **Fork** this repository to your GitHub account
|
||||
2. **Clone** your fork locally:
|
||||
```bash
|
||||
git clone https://github.com/your-username/awesome-ai-for-science.git
|
||||
cd awesome-ai-for-science
|
||||
```
|
||||
3. **Create** a new branch for your contribution:
|
||||
```bash
|
||||
git checkout -b add-new-resource
|
||||
```
|
||||
4. **Make** your changes to the README.md file
|
||||
5. **Commit** your changes:
|
||||
```bash
|
||||
git add README.md
|
||||
git commit -m "Add [resource name] to [section]"
|
||||
```
|
||||
6. **Push** to your fork:
|
||||
```bash
|
||||
git push origin add-new-resource
|
||||
```
|
||||
7. **Submit** a pull request from your fork to this repository
|
||||
|
||||
## 🔧 Types of Contributions
|
||||
|
||||
We welcome these types of contributions:
|
||||
|
||||
### ✅ Adding New Resources
|
||||
- **Tools & Software**: AI tools that help with research workflows
|
||||
- **Papers & Publications**: Influential papers in AI for Science
|
||||
- **Datasets**: High-quality scientific datasets
|
||||
- **Models**: Pre-trained models for scientific applications
|
||||
- **Educational Content**: Courses, tutorials, books
|
||||
- **Communities**: Research groups, conferences, forums
|
||||
|
||||
### ✅ Improving Existing Content
|
||||
- **Better Descriptions**: More accurate or detailed descriptions
|
||||
- **Updated Links**: Fixing broken or outdated links
|
||||
- **Reorganization**: Improving the structure and categorization
|
||||
- **Additional Information**: Adding missing details or context
|
||||
|
||||
### ✅ General Improvements
|
||||
- **Typo Fixes**: Grammar, spelling, and formatting corrections
|
||||
- **New Categories**: Suggesting new sections or reorganization
|
||||
- **Documentation**: Improving this contributing guide or README
|
||||
|
||||
## 📝 Contribution Guidelines
|
||||
|
||||
### Resource Quality Standards
|
||||
Before adding a resource, ensure it meets these criteria:
|
||||
|
||||
- **✅ Relevance**: Directly related to AI applications in scientific research
|
||||
- **✅ Quality**: Well-documented, actively maintained, or highly cited
|
||||
- **✅ Accessibility**: Publicly available (open source, free, or with free tier)
|
||||
- **✅ Uniqueness**: Not already listed in the repository
|
||||
- **✅ Functionality**: Actually works and provides value to researchers
|
||||
|
||||
### What NOT to Include
|
||||
- **❌ Commercial Products**: Purely commercial tools without free access
|
||||
- **❌ Broken Links**: Resources that are no longer available
|
||||
- **❌ Personal Projects**: Small, unmaintained personal repositories
|
||||
- **❌ Duplicates**: Resources already listed elsewhere in the repo
|
||||
- **❌ Off-topic**: Resources not related to AI or scientific research
|
||||
|
||||
## 📐 Formatting Guidelines
|
||||
|
||||
### General Format
|
||||
```markdown
|
||||
- [Resource Name](URL) - Brief description of what it does and why it's useful
|
||||
```
|
||||
|
||||
### Examples of Good Entries
|
||||
```markdown
|
||||
- [AlphaFold](https://github.com/deepmind/alphafold) - Revolutionary protein structure prediction using deep learning
|
||||
- [Elicit](https://elicit.org/) - AI research assistant that helps with literature review and evidence synthesis
|
||||
- [Materials Project](https://materialsproject.org/) - Computational materials database with ML-predicted properties
|
||||
```
|
||||
|
||||
### Description Guidelines
|
||||
- **Length**: 5-15 words ideally, max 20 words
|
||||
- **Style**: Clear, informative, avoid marketing language
|
||||
- **Focus**: What it does and scientific domain
|
||||
- **Tone**: Professional and objective
|
||||
|
||||
### Link Guidelines
|
||||
- **Use HTTPS**: Always use secure links when available
|
||||
- **Direct Links**: Link to the main project page, not sub-pages
|
||||
- **GitHub**: For open source projects, link to the GitHub repository
|
||||
- **Papers**: Link to the official publication (DOI preferred)
|
||||
|
||||
### Section Organization
|
||||
- **Alphabetical Order**: Within each subsection, maintain alphabetical order
|
||||
- **Appropriate Section**: Place resources in the most specific relevant section
|
||||
- **New Sections**: Propose new sections if existing ones don't fit
|
||||
|
||||
## 🔍 Review Process
|
||||
|
||||
### What We Look For
|
||||
1. **Accuracy**: Correct information and working links
|
||||
2. **Formatting**: Follows the style guide
|
||||
3. **Placement**: Resource is in the appropriate section
|
||||
4. **Quality**: Meets our quality standards
|
||||
5. **Uniqueness**: Not a duplicate
|
||||
|
||||
### Timeline
|
||||
- **Initial Review**: Within 7 days
|
||||
- **Feedback**: We'll provide constructive feedback if changes are needed
|
||||
- **Final Decision**: Merge or close within 14 days
|
||||
|
||||
### Review Criteria
|
||||
✅ **Approve** if:
|
||||
- Meets all quality standards
|
||||
- Follows formatting guidelines
|
||||
- Adds clear value to the collection
|
||||
|
||||
🔄 **Request Changes** if:
|
||||
- Minor formatting or description issues
|
||||
- Needs better categorization
|
||||
- Requires additional context
|
||||
|
||||
❌ **Reject** if:
|
||||
- Doesn't meet quality standards
|
||||
- Off-topic or commercial
|
||||
- Duplicate of existing entry
|
||||
|
||||
## 🎯 Code of Conduct
|
||||
|
||||
### Our Standards
|
||||
- **Be Respectful**: Treat all contributors with respect
|
||||
- **Be Constructive**: Provide helpful, actionable feedback
|
||||
- **Be Collaborative**: Work together to improve the resource
|
||||
- **Be Patient**: Understand that reviews take time
|
||||
|
||||
### Unacceptable Behavior
|
||||
- Harassment or discriminatory language
|
||||
- Spam or self-promotion without value
|
||||
- Disruptive or unconstructive criticism
|
||||
- Violations of intellectual property
|
||||
|
||||
## 🆘 Getting Help
|
||||
|
||||
### Questions?
|
||||
- **Issues**: Open an issue for questions about contributions
|
||||
- **Discussions**: Use GitHub Discussions for general questions
|
||||
- **Email**: Contact the maintainers for sensitive issues
|
||||
|
||||
### Common Questions
|
||||
|
||||
**Q: Can I add my own research project?**
|
||||
A: Yes, if it's high-quality, well-documented, and provides clear value to the scientific community.
|
||||
|
||||
**Q: What if a resource becomes outdated?**
|
||||
A: Please open an issue or submit a PR to remove or update it.
|
||||
|
||||
**Q: Can I reorganize entire sections?**
|
||||
A: Major reorganizations should be discussed in an issue first to gather feedback.
|
||||
|
||||
**Q: What about resources behind paywalls?**
|
||||
A: We prefer freely accessible resources, but important papers or tools with free tiers are acceptable.
|
||||
|
||||
---
|
||||
|
||||
## 🙏 Thank You!
|
||||
|
||||
Your contributions make this resource valuable for researchers worldwide. Every addition, fix, and improvement helps accelerate scientific discovery through AI.
|
||||
|
||||
**Happy Contributing!** 🚀
|
||||
|
||||
---
|
||||
|
||||
*For questions about this guide, please open an issue or start a discussion.*
|
||||
File diff suppressed because it is too large
Load Diff
Vendored
+28
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: "Pull Request Template"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/PULL_REQUEST_TEMPLATE.md
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
<!-- Please describe what you are adding or changing, and why it is awesome. -->
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] I have searched previous suggestions and this is not a duplicate.
|
||||
- [ ] I have added only one link per pull request.
|
||||
- [ ] The link follows the format: `[name](https://example.com/)` - A short description ends with a period.
|
||||
- [ ] Descriptions are concise.
|
||||
- [ ] Alphabetical ordering is maintained where applicable.
|
||||
- [ ] If a new section was added, the section description and title are included, and the title is added to the Index.
|
||||
- [ ] Spelling and grammar have been checked.
|
||||
- [ ] There is no trailing whitespace.
|
||||
- [ ] The pull request title follows the format: `Add user/repo - Short repo description`
|
||||
upstream/inoue0426-awesome-computational-biology/catalogue/.github/workflows/ai4bio-schema-check.yml
Vendored
+61
@@ -0,0 +1,61 @@
|
||||
---
|
||||
title: "Ai4Bio Schema Check"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/.github/workflows/ai4bio-schema-check.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: AI4Bio Schema Check
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
paths:
|
||||
- data/resources.yml
|
||||
- data/enrichment.yml
|
||||
- 'data/enrichment.*.yml'
|
||||
- data/vocabulary.yml
|
||||
- docs/data/resource.schema.json
|
||||
- scripts/enrichment_fragments.py
|
||||
- scripts/validate_resources.py
|
||||
- scripts/build_resources_v2.py
|
||||
push:
|
||||
branches: [main]
|
||||
paths:
|
||||
- data/resources.yml
|
||||
- data/enrichment.yml
|
||||
- 'data/enrichment.*.yml'
|
||||
- data/vocabulary.yml
|
||||
- docs/data/resource.schema.json
|
||||
- scripts/enrichment_fragments.py
|
||||
- scripts/validate_resources.py
|
||||
- scripts/build_resources_v2.py
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
validate:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: astral-sh/setup-uv@v3
|
||||
- name: Validate schema and enrichment
|
||||
run: uv run --with pyyaml python scripts/validate_resources.py
|
||||
- name: Build enriched artifacts
|
||||
run: uv run --with pyyaml python scripts/build_resources_v2.py
|
||||
- name: Verify enriched artifacts are committed
|
||||
run: |
|
||||
if git diff --quiet; then
|
||||
echo "AI4Bio artifacts are in sync."
|
||||
exit 0
|
||||
fi
|
||||
echo "Generated AI4Bio artifacts are out of date. Run:"
|
||||
echo " uv run --with pyyaml python scripts/build_resources_v2.py"
|
||||
git status --short
|
||||
exit 1
|
||||
Vendored
+46
@@ -0,0 +1,46 @@
|
||||
---
|
||||
title: "Docs Check"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/docs-check.yml
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: Docs / lint
|
||||
|
||||
on:
|
||||
push:
|
||||
paths:
|
||||
- '**/*.md'
|
||||
pull_request:
|
||||
paths:
|
||||
- '**/*.md'
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
docs-check:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: 'lts/*'
|
||||
|
||||
- name: Install tooling
|
||||
run: npm install -g [email protected] [email protected] --no-fund
|
||||
|
||||
- name: Run markdownlint
|
||||
run: markdownlint '**/*.md' --ignore node_modules
|
||||
|
||||
- name: Run cspell (spellcheck)
|
||||
run: cspell "**/*.md" --no-summary --no-progress
|
||||
Vendored
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
title: "Generate Artifacts"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/generate-artifacts.yml
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: Generate data artifacts
|
||||
|
||||
on:
|
||||
push:
|
||||
paths:
|
||||
- 'data/resources.yml'
|
||||
pull_request:
|
||||
paths:
|
||||
- 'data/resources.yml'
|
||||
workflow_dispatch:
|
||||
|
||||
jobs:
|
||||
generate:
|
||||
runs-on: ubuntu-latest
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ github.head_ref || github.ref_name }}
|
||||
|
||||
- name: Set up Python
|
||||
uses: actions/setup-python@v5
|
||||
with:
|
||||
python-version: '3.11'
|
||||
|
||||
- name: Install dependencies
|
||||
run: pip install pyyaml
|
||||
|
||||
- name: Generate JSON and CSV
|
||||
run: python scripts/generate_artifacts.py
|
||||
|
||||
- name: Commit artifacts (push events only)
|
||||
if: github.event_name == 'push' || github.event_name == 'workflow_dispatch'
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "github-actions[bot]@users.noreply.github.com"
|
||||
git add data/resources.json data/resources.csv docs/data/resources.json
|
||||
git diff --cached --quiet || git commit -m "chore: regenerate resources.json and resources.csv"
|
||||
git push
|
||||
Vendored
+43
@@ -0,0 +1,43 @@
|
||||
---
|
||||
title: "Landscape Ui Check"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/.github/workflows/landscape-ui-check.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: Landscape UI Check
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
paths:
|
||||
- docs/landscape.html
|
||||
- docs/landscape.css
|
||||
- docs/landscape.js
|
||||
- .github/workflows/landscape-ui-check.yml
|
||||
push:
|
||||
branches: [main]
|
||||
paths:
|
||||
- docs/landscape.html
|
||||
- docs/landscape.css
|
||||
- docs/landscape.js
|
||||
- .github/workflows/landscape-ui-check.yml
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
ui-check:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: 'lts/*'
|
||||
- name: Check landscape JavaScript syntax
|
||||
run: node --check docs/landscape.js
|
||||
Vendored
+42
@@ -0,0 +1,42 @@
|
||||
---
|
||||
title: "Link Check"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/link-check.yml
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: Link Check
|
||||
|
||||
on:
|
||||
schedule:
|
||||
- cron: '0 5 * * 1' # 毎週月曜 05:00 UTC
|
||||
workflow_dispatch: # 手動実行も可能
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
link-check:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: 'lts/*'
|
||||
|
||||
- name: Install markdown-link-check
|
||||
run: npm install -g [email protected] --no-fund
|
||||
|
||||
- name: Run markdown-link-check
|
||||
run: |
|
||||
find . -name '*.md' -not -path './node_modules/*' -print0 | \
|
||||
xargs -0 -I{} markdown-link-check {} -q -c .markdown-link-check.json
|
||||
Vendored
+56
@@ -0,0 +1,56 @@
|
||||
---
|
||||
title: "Pr Quality Checks"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/pr-quality-checks.yml
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: PR Quality Checks
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
paths:
|
||||
- README.md
|
||||
- data/resources.yml
|
||||
- data/resources.json
|
||||
- data/resources.csv
|
||||
- docs/data/resources.json
|
||||
- scripts/*.py
|
||||
- scripts/**/*.py
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
resources-consistency:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v3
|
||||
|
||||
- name: Sync resources from README
|
||||
run: uv run --with pyyaml python scripts/sync_resources_from_readme.py
|
||||
|
||||
- name: Build resource artifacts
|
||||
run: uv run --with pyyaml python scripts/build_resources.py
|
||||
|
||||
- name: Verify generated artifacts are committed
|
||||
run: |
|
||||
if git diff --quiet; then
|
||||
echo "Resources are in sync."
|
||||
exit 0
|
||||
fi
|
||||
echo "Generated files are out of date. Run:"
|
||||
echo " uv run --with pyyaml python scripts/sync_resources_from_readme.py"
|
||||
echo " uv run --with pyyaml python scripts/build_resources.py"
|
||||
git status --short
|
||||
exit 1
|
||||
Vendored
+78
@@ -0,0 +1,78 @@
|
||||
---
|
||||
title: "Sync Resources"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/.github/workflows/sync_resources.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: Sync Resources
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
paths:
|
||||
- README.md
|
||||
- data/resources.yml
|
||||
- data/enrichment.yml
|
||||
- 'data/enrichment.*.yml'
|
||||
- scripts/*.py
|
||||
- scripts/**/*.py
|
||||
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
jobs:
|
||||
sync:
|
||||
if: github.actor != 'github-actions[bot]'
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v3
|
||||
|
||||
- name: Detect Changed Files
|
||||
id: changes
|
||||
run: |
|
||||
BEFORE="${{ github.event.before }}"
|
||||
AFTER="${{ github.sha }}"
|
||||
if [ -z "$BEFORE" ] || [ "$BEFORE" = "0000000000000000000000000000000000000000" ]; then
|
||||
CHANGED=$(git diff --name-only HEAD~1..HEAD)
|
||||
else
|
||||
CHANGED=$(git diff --name-only "$BEFORE" "$AFTER")
|
||||
fi
|
||||
echo "changed<<EOF" >> "$GITHUB_OUTPUT"
|
||||
echo "$CHANGED" >> "$GITHUB_OUTPUT"
|
||||
echo "EOF" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Sync From README
|
||||
if: contains(steps.changes.outputs.changed, 'README.md')
|
||||
run: uv run --with pyyaml python scripts/sync_resources_from_readme.py
|
||||
|
||||
- name: Validate Resource Schema
|
||||
run: uv run --with pyyaml python scripts/validate_resources.py
|
||||
|
||||
- name: Build Artifacts
|
||||
run: uv run --with pyyaml python scripts/build_resources.py
|
||||
|
||||
- name: Commit Updates
|
||||
run: |
|
||||
if git diff --quiet; then
|
||||
echo "No changes to commit."
|
||||
exit 0
|
||||
fi
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
|
||||
git add data/resources.yml data/resources.json data/resources.csv docs/data/resources.json
|
||||
git commit -m "chore: sync resources"
|
||||
git push
|
||||
Vendored
+58
@@ -0,0 +1,58 @@
|
||||
---
|
||||
title: "Update Overview"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.github/workflows/update-overview.yml
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
name: Update Overview Figure
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
paths:
|
||||
- docs/data/resources.json
|
||||
schedule:
|
||||
# Every Monday at 03:00 UTC
|
||||
- cron: '0 3 * * 1'
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
jobs:
|
||||
update-overview:
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
# Use the default GITHUB_TOKEN so commits don't re-trigger this workflow
|
||||
token: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
- name: Set up Python
|
||||
uses: actions/setup-python@v5
|
||||
with:
|
||||
python-version: '3.11'
|
||||
|
||||
- name: Install dependencies
|
||||
run: pip install matplotlib
|
||||
|
||||
- name: Regenerate overview figure
|
||||
run: python scripts/generate_overview.py
|
||||
|
||||
- name: Commit updated overview
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
|
||||
git add docs/overview.png
|
||||
git diff --cached --quiet || git commit -m "chore: regenerate overview figure"
|
||||
git push
|
||||
@@ -0,0 +1,16 @@
|
||||
---
|
||||
title: ".Markdown Link Check"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/.markdown-link-check.json
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
{
|
||||
"aliveStatusCodes": [200, 206, 301, 302, 307, 308, 403]
|
||||
}
|
||||
@@ -0,0 +1,546 @@
|
||||
---
|
||||
title: "Awesome Computational Biology [](https://awesome.re)"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/README.md
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Awesome Computational Biology [](https://awesome.re)
|
||||
|
||||
A curated collection of databases, software, and papers related to computational biology.
|
||||
|
||||
> Computational biology involves the development and application of data-analytical and theoretical methods, mathematical modelling and computational simulation techniques to the study of biological, ecological, behavioural, and social systems. — [Wikipedia](https://en.wikipedia.org/wiki/Computational_biology)
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
[](https://inoue0426.github.io/awesome-computational-biology/overview.html)
|
||||
|
||||
> Interactive version: [Resource Overview page](https://inoue0426.github.io/awesome-computational-biology/overview.html)
|
||||
> Regenerate the figure: `python scripts/generate_overview.py`
|
||||
|
||||
---
|
||||
|
||||
## GitHub Pages UI
|
||||
|
||||
Browse and search the resources via the [GitHub Pages UI](https://inoue0426.github.io/awesome-computational-biology/).
|
||||
|
||||
- Search matches `name`, `description`, `tasks`, `modalities`, and `tags`.
|
||||
- The **Task**, **Modality**, and **Type** filters map directly to `tasks`, `modalities`, and `type` in `docs/data/resources.json`.
|
||||
- Clicking badges on cards applies the corresponding filter.
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [Awesome Computational Biology](#awesome-computational-biology-)
|
||||
- [Table of Contents](#table-of-contents)
|
||||
- [Overview](#overview)
|
||||
- [GitHub Pages UI](#github-pages-ui)
|
||||
- [Citation](#citation)
|
||||
- [Curation Criteria (Strict)](#curation-criteria-strict)
|
||||
- [Update & Link Rot Policy](#update--link-rot-policy)
|
||||
- [Data Schema & Contribution Workflow](#data-schema--contribution-workflow)
|
||||
- [Databases](#databases)
|
||||
- [scRNA](#scrna)
|
||||
- [Compound](#compound)
|
||||
- [Pathway](#pathway)
|
||||
- [Mass Spectra](#mass-spectra)
|
||||
- [Protein](#protein)
|
||||
- [Genome](#genome)
|
||||
- [Disease](#disease)
|
||||
- [Interaction](#interaction)
|
||||
- [Drug-Gene Interaction](#drug-gene-interaction)
|
||||
- [Drug (Cell Line) Response](#drug-cell-line-response)
|
||||
- [Chemical-Protein Interaction](#chemical-protein-interaction)
|
||||
- [Protein-Protein Interaction](#protein-protein-interaction)
|
||||
- [Knowledge Graph](#knowledge-graph)
|
||||
- [Gene Regulatory Network](#gene-regulatory-network)
|
||||
- [Clinical Trial](#clinical-trial)
|
||||
- [Benchmarks & Datasets](#benchmarks--datasets)
|
||||
- [API](#api)
|
||||
- [Preprocessing Tools](#preprocessing-tools)
|
||||
- [Machine Learning Tasks and Models](#machine-learning-tasks-and-models)
|
||||
- [Drug Discovery](#drug-discovery)
|
||||
- [Drug Response Prediction](#drug-response-prediction)
|
||||
- [Drug Perturbation](#drug-perturbation)
|
||||
- [Drug Repurposing](#drug-repurposing)
|
||||
- [Drug Target Interaction](#drug-target-interaction)
|
||||
- [Compound-Protein Interaction](#compound-protein-interaction)
|
||||
- [Molecular Generation](#molecular-generation)
|
||||
- [LLM for Biology](#llm-for-biology)
|
||||
- [Foundation Models](#foundation-models)
|
||||
- [Single-cell Foundation Models](#single-cell-foundation-models)
|
||||
- [Transcriptomics Foundation Models](#transcriptomics-foundation-models)
|
||||
- [Spatial Foundation Models](#spatial-foundation-models)
|
||||
- [Multi-Omics Foundation Models](#multi-omics-foundation-models)
|
||||
- [Domain Alignment](#domain-alignment)
|
||||
- [Compound Foundation Models](#compound-foundation-models)
|
||||
- [Compound Embedding](#compound-embedding)
|
||||
- [Protein Foundation Models](#protein-foundation-models)
|
||||
- [Pre-trained Embedding](#pre-trained-embedding)
|
||||
- [Protein Structure Prediction and Design](#protein-structure-prediction-and-design)
|
||||
- [Multi-Modal Foundation Models](#multi-modal-foundation-models)
|
||||
- [Genomics Foundation Models](#genomics-foundation-models)
|
||||
|
||||
---
|
||||
|
||||
## Databases
|
||||
|
||||
### scRNA
|
||||
|
||||
- [CZ CELLxGENE](https://cellxgene.cziscience.com/) — Single-cell dataset repository and interactive explorer from the Chan Zuckerberg Initiative.
|
||||
- [Gene Expression Omnibus](https://www.ncbi.nlm.nih.gov/geo/) — Public functional genomics database.
|
||||
- [Human Cell Atlas](https://www.humancellatlas.org/) — Open global atlas of all cells in the human body.
|
||||
- [Single Cell PORTAL](https://singlecell.broadinstitute.org/single_cell) — Public database for single-cell RNA.
|
||||
- [Single Cell Expression Atlas](https://www.ebi.ac.uk/gxa/sc/home) — Public database for single-cell RNA.
|
||||
|
||||
### Compound
|
||||
|
||||
- [PubChem](https://pubchem.ncbi.nlm.nih.gov/) — One of the largest chemical databases (compounds, genes, and proteins).
|
||||
- [ChEBI](https://www.ebi.ac.uk/chebi/) — Database focused on small chemical compounds.
|
||||
- [ChEMBL](https://www.ebi.ac.uk/chembl/) — Bioactive molecules with drug-like properties.
|
||||
- [ChemSpider](http://www.chemspider.com/) — Chemical structure database.
|
||||
- [DrugTargetCommons](https://drugtargetcommons.fimm.fi/) — Community platform for curating and integrating experimental bioactivity data across drugs and targets.
|
||||
- [HMDB (Human Metabolome Database)](https://hmdb.ca/) — Comprehensive database of small molecule metabolites found in the human body.
|
||||
- [KEGG COMPOUND](https://www.genome.jp/kegg/compound/) — Collection of small molecules and biopolymers.
|
||||
- [LIPID MAPS](https://www.lipidmaps.org/databases/lmsd/overview) — Database of lipids.
|
||||
- [Rhea](https://www.rhea-db.org/) — Database of chemical reactions.
|
||||
- [DrugCentral](http://drugcentral.org/) — Online drug compendium with drug mode of action and indication information.
|
||||
- [Drug Repurposing Hub](https://repo-hub.broadinstitute.org/repurposing#download-data) — Collections of drug repurposing data (drug, MoA, target, etc).
|
||||
- [Therapeutic Target Database](https://idrblab.net/ttd/full-data-download) — Drug-target, target-disease, and drug-disease datasets.
|
||||
- [ZINC ligand discovery database](https://zinc.docking.org/) — Free database of commercially-available compounds for virtual screening.
|
||||
|
||||
### Pathway
|
||||
|
||||
- [PathwayCommons](https://www.pathwaycommons.org/) — Database of pathways and interactions.
|
||||
- [KEGG PATHWAY](https://www.genome.jp/kegg/pathway.html) — Collection of pathway maps.
|
||||
- [WikiPathways](https://wikipathways.org/) — Database of biological pathways.
|
||||
- [Reactome](https://reactome.org/) — Expert-curated, peer-reviewed pathway database with detailed reaction mechanisms.
|
||||
- [BioCyc](https://biocyc.org/) — Collection of pathway/genome databases across thousands of organisms.
|
||||
- [OmniPath](https://omnipathdb.org/) — Comprehensive resource integrating protein interactions, signaling pathways, gene regulatory networks, and miRNA targets from over 100 databases.
|
||||
- [SIGNOR 2.0](https://signor.uniroma2.it/) — Database of causal signaling interactions and pathways, with signed and directed relationships between proteins.
|
||||
- [MSigDB (Molecular Signatures Database)](https://www.gsea-msigdb.org/gsea/msigdb) — Curated gene sets derived from pathways and biological processes.
|
||||
|
||||
### Mass Spectra
|
||||
|
||||
- [MassBank](http://www.massbank.jp/) — Open source databases and tools for mass spectrometry reference spectra.
|
||||
- [MoNA MassBank of North America](https://mona.fiehnlab.ucdavis.edu/) — Meta-database of metabolite mass spectra, metadata, and associated compounds.
|
||||
|
||||
### Protein
|
||||
|
||||
- [THE HUMAN PROTEIN ATLAS](https://www.proteinatlas.org/) — Comprehensive human protein database (cells, tissues, organs).
|
||||
- [PROTEIN DATA BANK (PDB)](https://www.rcsb.org/) — 3D structures of proteins, nucleic acids, complexes.
|
||||
- [UniProt](https://www.uniprot.org/) — Functional information on proteins.
|
||||
- [AlphaFold Protein Structure Database](https://alphafold.ebi.ac.uk/api-docs) — 3D protein structure predictions.
|
||||
- [RCSB Protein Data Bank](https://www.rcsb.org/) — Repository for structural data of biological molecules.
|
||||
- [Critical Assessment of Structure Prediction (CASP)](https://predictioncenter.org/) — Assessing methods for protein structure prediction.
|
||||
- [Uniclust](https://uniclust.mmseqs.com/) — Clustered protein sequence databases.
|
||||
- [UniRef](https://www.uniprot.org/uniref/) — Non-redundant sequence database clustering UniProtKB entries at multiple sequence identity thresholds.
|
||||
- [CATH database](https://www.cathdb.info/) — Hierarchical classification of protein domain structures.
|
||||
- [SAbDab](https://opig.stats.ox.ac.uk/webapps/sabdab-sabpred/sabdab) — Structural Antibody Database containing all antibody structures in the PDB.
|
||||
- [OADB (Observed Antibody Space Database)](http://opig.stats.ox.ac.uk/webapps/oas/) — Database of antibody sequences from immune repertoire sequencing.
|
||||
- [InterPro](https://www.ebi.ac.uk/interpro/) — Protein families, domains, and functional sites database integrating 14 member databases including Pfam and PROSITE.
|
||||
- [Pfam](https://www.ebi.ac.uk/interpro/entry/pfam/) — Database of protein families described by multiple sequence alignments and hidden Markov models.
|
||||
- [NeXtProt](https://www.nextprot.org/) — Expert knowledge base on human proteins with deep functional annotation, complementary to UniProt.
|
||||
|
||||
### Genome
|
||||
|
||||
- [ENCODE](https://www.encodeproject.org/) — Encyclopedia of DNA Elements; regulatory and functional genomic elements across the genome.
|
||||
- [Ensembl](https://www.ensembl.org/) — Genome browser and annotation database for vertebrate and other eukaryotic genomes.
|
||||
- [Human Genome Resources at NCBI](https://www.ncbi.nlm.nih.gov/projects/genome/guide/human/index.shtml) — Database for genomics, proteomics, transcriptomics, and systems biology.
|
||||
- [GenBank](https://www.ncbi.nlm.nih.gov/genbank/) — NCBI's database of genetic sequences.
|
||||
- [UCSC Genome Browser](https://genome.ucsc.edu/) — UCSC's genome browser.
|
||||
- [cBioPortal](https://www.cbioportal.org/) — Cancer genomics database; aggregating many patient datasets.
|
||||
- [10x Genomics Dataset](https://www.10xgenomics.com/resources/datasets) — Collection of single-cell datasets.
|
||||
- [The Genotype-Tissue Expression (GTEx)](https://gtexportal.org/home/) — Human gene expression and regulation resource.
|
||||
- [Dependency Map (DepMap)](https://depmap.org/portal/) — CRISPR-Cas9 screens in cancer cell lines.
|
||||
- [Catalogue Of Somatic Mutations In Cancer (COSMIC)](https://cancer.sanger.ac.uk/cosmic) — Resource on somatic mutations in cancers.
|
||||
- [MGnify](https://www.ebi.ac.uk/metagenomics/) — Resource for metagenomic and metatranscriptomic data.
|
||||
- [JASPAR](http://jaspar.genereg.net/) — Database of transcription factor binding profiles.
|
||||
- [gnomAD](https://gnomad.broadinstitute.org/) — Genome Aggregation Database; genetic variation from large-scale sequencing projects.
|
||||
- [Rfam](https://rfam.org/) — Database of RNA families with sequence alignments and consensus structures.
|
||||
- [ROADMAP Epigenomics](http://www.roadmapepigenomics.org/) — Reference epigenome maps for 111 primary human cell types and tissues, including histone modifications, chromatin accessibility, and DNA methylation.
|
||||
- [FANTOM5](https://fantom.gsc.riken.jp/5/) — Functional annotation of mammalian genome; comprehensive atlas of active enhancers, promoters, and transcription start sites across human and mouse cell types.
|
||||
|
||||
### Disease
|
||||
|
||||
- [KEGG DRUG](https://www.genome.jp/kegg/drug/) — Comprehensive, approved drug information.
|
||||
- [DrugBank](https://go.drugbank.com/) — Database of drugs and targets (University of Alberta).
|
||||
- [DisGeNET](https://www.disgenet.org/) — Database of gene-disease associations integrating expert-curated and GWAS data.
|
||||
- [OMIM (Online Mendelian Inheritance in Man)](https://www.omim.org/) — Comprehensive database of human genes and genetic disorders.
|
||||
- [Open Targets Platform](https://platform.opentargets.org/) — Systematic target identification and prioritization platform integrating genetics, genomics, and drug data for drug discovery.
|
||||
- [Human Phenotype Ontology (HPO)](https://hpo.jax.org/) — Standardized vocabulary of phenotypic abnormalities in human disease, linking genes, variants, and clinical features.
|
||||
- [DISEASES](https://diseases.jensenlab.org/) — Gene–disease association database integrating evidence from text mining, curated databases, and experimental data.
|
||||
|
||||
### Interaction
|
||||
|
||||
#### Drug-Gene Interaction
|
||||
|
||||
- [DGIdb](https://www.dgidb.org/) — Drug-gene interactions and the druggable genome.
|
||||
- [Comparative Toxicogenomics Database](http://ctdbase.org/) — Chemical-gene interactions, chemical-disease and gene-disease associations, chemical-phenotype associations.
|
||||
- [SNAP](https://snap.stanford.edu/biodata/datasets/10002/10002-ChG-Miner.html) — Dataset of drug-gene interactions.
|
||||
|
||||
#### Drug (Cell Line) Response
|
||||
|
||||
- [NCI60](https://dtp.cancer.gov/discovery_development/nci-60/) — Focuses on 60 cancer cell lines and many drugs.
|
||||
- [Genomics of Drug Sensitivity in Cancer (GDSC)](https://www.cancerrxgene.org/) — Drug sensitivity for ~1000 human cancer cell lines and hundreds of compounds.
|
||||
- [Cancer Cell Line Encyclopedia](https://sites.broadinstitute.org/ccle/) — Database of ~1000 cancer cell lines.
|
||||
- [CellMiner Cross Database (CellMinerCDB)](https://discover.nci.nih.gov/cellminercdb/) — Integrates multiple cancer cell line databases.
|
||||
|
||||
#### Chemical-Protein Interaction
|
||||
|
||||
- [STITCH](http://stitch.embl.de/) — Chemical-protein interactions.
|
||||
- [BindingDB](https://www.bindingdb.org/rwd/bind/index.jsp) — Compounds and target database.
|
||||
- [Davis kinase inhibitors DB](http://staff.cs.utu.fi/~aijrinas/dti/) — Experimental kinase inhibitor binding affinity dataset for protein–ligand interaction research.
|
||||
- [Kinase Inhibitor Bioactivity Data (KIBA)](https://janeliascicomp.github.io/KIBA/) — Integrated bioactivity scores for kinase inhibitors combining Ki, Kd, and IC50 measurements.
|
||||
- [PDBBind](https://www.pdbbind-plus.org.cn/) — Binding affinity data for biomolecular complexes.
|
||||
|
||||
#### Protein-Protein Interaction
|
||||
|
||||
- [STRING](https://string-db.org/) — PPI networks for multiple organisms.
|
||||
- [BioGRID](https://thebiogrid.org/) — Protein, genetic, and chemical interactions.
|
||||
- [HIPPIE](http://cbdm-01.zdv.uni-mainz.de/~mschaefer/hippie/) — Human protein-protein interaction database.
|
||||
- [IntAct](https://www.ebi.ac.uk/intact/home) — Open-source molecular interaction database and analysis system from EMBL-EBI.
|
||||
|
||||
#### Knowledge Graph
|
||||
|
||||
- [Drug Mechanism Database (DrugMechDB)](https://github.com/SuLab/DrugMechDB/tree/2.0.1) — Mechanisms of action from drug to disease.
|
||||
- [DRKG](https://github.com/gnn4dr/DRKG) — Large-scale biological knowledge graph for drug discovery.
|
||||
- [Hetionet](https://github.com/hetio/hetionet) — Heterogeneous network integrating genes, diseases, drugs, pathways, and more.
|
||||
- [PrimeKG](https://github.com/mims-harvard/PrimeKG) — Multi-modal precision medicine knowledge graph integrating clinical, genetic, and drug data.
|
||||
|
||||
#### Gene Regulatory Network
|
||||
|
||||
- [TRRUST v2](https://www.grnpedia.org/trrust/) — Manually curated database of human and mouse transcriptional regulatory interactions between transcription factors and their target genes, expanded with literature-derived evidence.
|
||||
- [RegNetwork](http://www.regnetworkweb.org/) — Database of gene regulatory networks covering transcription factor–target gene and miRNA–gene interaction data across multiple species.
|
||||
- [miRBase](https://www.mirbase.org/) — Reference repository for microRNA gene annotations, sequences, and experimentally validated targets.
|
||||
|
||||
### Clinical Trial
|
||||
|
||||
- [ClinicalTrials.gov](https://clinicaltrials.gov/) — Privately and publicly funded clinical studies.
|
||||
- [ICD10](https://icd.who.int/browse10/2019/en) — International Classification of Diseases, 10th revision.
|
||||
- [EU Drug Regulating Authorities Clinical Trials DB (EudraCT)](https://eudract.ema.europa.eu/) — European clinical trial database.
|
||||
- [MIMIC-IV](https://mimic.mit.edu/) — Freely accessible critical care database.
|
||||
|
||||
---
|
||||
|
||||
## Benchmarks & Datasets
|
||||
|
||||
- [1000 Genomes Project](https://www.internationalgenome.org/) — Reference panel of human genetic variation from 2,504 individuals across 26 populations.
|
||||
- [BACE](https://www.kaggle.com/datasets/gokturkkoch/bace) — Binary classification and regression dataset for β-secretase 1 (BACE-1) inhibitor binding affinity.
|
||||
- [BEAT AML](https://biodev.github.io/BeatAML2/) — Functional ex vivo drug sensitivity measurements paired with genomics for acute myeloid leukemia.
|
||||
- [Bento](https://github.com/LigandPro/Bento) — Protein-ligand docking benchmark covering rigid, flexible, de novo, blind, induced-fit, and covalent docking tasks.
|
||||
- [BindingDB Curated Sets](https://www.bindingdb.org/rwd/bind/chemsearch/marvin/SDFdownload.jsp?all_download=yes) — Curated binding affinity datasets for protein–ligand interaction benchmarking.
|
||||
- [Cancer Therapeutics Response Portal (CTRP)](https://portals.broadinstitute.org/ctrp/) — Drug sensitivity profiles across ~900 cancer cell lines for >400 compounds.
|
||||
- [ClinTox](https://tdcommons.ai/single_pred_tasks/tox/#clintox) — Clinical toxicity dataset contrasting FDA-approved drugs with those that failed clinical trials due to toxicity.
|
||||
- [CPTAC (Clinical Proteomic Tumor Analysis Consortium)](https://proteomics.cancer.gov/programs/cptac) — Multi-omic proteogenomic datasets for multiple cancer types linking proteomics with genomics.
|
||||
- [CrossDocked2020](https://arxiv.org/abs/2001.01037) — Large-scale dataset for structure-based virtual screening.
|
||||
- [DUD-E (Directory of Useful Decoys, Enhanced)](http://dude.docking.org/) — Structure-based virtual screening benchmark with active ligands and challenging decoy sets across diverse protein targets.
|
||||
- [FLIP (Fitness Landscape Inference for Proteins)](https://github.com/J-SNACKKB/FLIP) — Benchmark collection of protein fitness landscape datasets for evaluating protein ML models.
|
||||
- [Genomics of Drug Sensitivity in Cancer (GDSC)](https://www.cancerrxgene.org/) — Drug sensitivity for ~1000 human cancer cell lines and hundreds of compounds.
|
||||
- [GuacaMol](https://github.com/BenevolentAI/guacamol) — Benchmark suite for generative molecular design models.
|
||||
- [JUMP Cell Painting Datasets](https://github.com/jump-cellpainting/datasets) — Consortium-scale cell imaging perturbation datasets (chemical and genetic) for phenotypic profiling and drug discovery research.
|
||||
- [LINCS L1000](https://lincsproject.org/LINCS/tools/workflows/find-the-best-place-to-obtain-the-lincs-l1000-data) — Gene expression profiles (978 landmark genes) for >20,000 chemical and genetic perturbations across cell lines.
|
||||
- [MoleculeNet](http://moleculenet.ai/) — Benchmark datasets for molecular machine learning.
|
||||
- [MOSES](https://github.com/molecularsets/moses) — Benchmarking platform for molecular generation models.
|
||||
- [NCI60](https://dtp.cancer.gov/discovery_development/nci-60/) — Drug sensitivity benchmark across 60 diverse human cancer cell lines.
|
||||
- [OGB (Open Graph Benchmark)](https://ogb.stanford.edu/) — Large-scale graph ML benchmark suite including biological datasets such as ogbl-ppa (protein-protein associations) and ogbg-molhiv.
|
||||
- [OpenBioLink](https://github.com/OpenBioLink/OpenBioLink) — Benchmark datasets for biological knowledge graph completion.
|
||||
- [PharmGKB](https://www.pharmgkb.org/) — Curated pharmacogenomics dataset linking genetic variants to drug response phenotypes across thousands of drugs.
|
||||
- [PK-DB](https://pk-db.com/) — Open database of experimental pharmacokinetics (PK) and ADME data from clinical and preclinical studies.
|
||||
- [PRISM](https://depmap.org/portal/prism/) — Cancer drug sensitivity profiling of >4,500 drugs across >900 cancer cell lines using pooled-cell-line barcoding.
|
||||
- [ProteinGym](https://github.com/OATML-Markslab/ProteinGym) — Large-scale benchmark of deep mutational scanning assays for evaluating protein fitness landscape models.
|
||||
- [QM9](https://figshare.com/collections/Quantum_chemistry_structures_and_properties_of_134_kilo_molecules/978904) — Quantum chemistry properties for 134K stable small organic molecules computed at DFT level.
|
||||
- [scIB (Single-cell Integration Benchmarks)](https://github.com/theislab/scib) — Comprehensive benchmarking framework for single-cell data integration methods.
|
||||
- [scPerturb](https://github.com/sanderlab/scPerturb) — Curated and continuously updated single-cell perturbation data resource spanning CRISPR and drug perturbation studies.
|
||||
- [SIDER (Side Effect Resource)](http://sideeffects.embl.de/) — Database of 1,430 approved drugs with their recorded adverse drug reactions across 27 system-organ classes.
|
||||
- [Tabula Muris](https://tabula-muris.ds.czbiohub.org/) — Comprehensive single-cell atlas of 20 mouse organs and tissues, enabling cross-tissue and cross-species comparisons.
|
||||
- [Tabula Sapiens](https://tabula-sapiens-portal.ds.czbiohub.org/) — Comprehensive human single-cell atlas of ~500K cells from 24 organs and tissues across multiple donors.
|
||||
- [TAPE (Tasks Assessing Protein Embeddings)](https://github.com/songlab-cal/tape) — Benchmark suite of five biologically meaningful semi-supervised learning tasks for evaluating protein representations.
|
||||
- [The Cancer Genome Atlas (TCGA)](https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga) — Comprehensive multi-omics (genomics, transcriptomics, proteomics, methylation) dataset for 33 cancer types across ~11,000 patients.
|
||||
- [TCGA virtual spatial transcriptomics atlas](https://huggingface.co/datasets/ratschlab/TCGA_virtual_spatial_transcriptomics_atlas) — DeepSpot-M predicted transcriptome-wide ST for TCGA H&E (FF + FFPE; 28,664 slides / 32 cancer types; gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).
|
||||
- [HEST Xenium virtual spatial transcriptomics](https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics) — DeepSpot-M predicted transcriptome-wide ST for 59 HEST-1k 10x Xenium samples (~13.3M cells) (gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).
|
||||
- [Therapeutics Data Commons (TDC)](https://tdcommons.ai/) — Unified benchmark suite covering ADMET, drug-target interaction, drug response, and more.
|
||||
- [Tox21](https://tripod.nih.gov/tox21/challenge/) — 12,707 compounds tested in 12 nuclear receptor and stress-response pathway biochemical assays for toxicity prediction.
|
||||
- [UK Biobank](https://www.ukbiobank.ac.uk/) — Large-scale biomedical database of ~500K participants with genetic, imaging, and health data for population genetics and disease studies.
|
||||
|
||||
---
|
||||
|
||||
## API
|
||||
|
||||
- [PubMed E-utilities (esearch/efetch)](https://www.nlm.nih.gov/dataguide/edirect/esearch.html) — APIs for searching and retrieving biomedical literature from PubMed.
|
||||
- [NCBI E-utilities](https://www.ncbi.nlm.nih.gov/books/NBK25501/) — Unified APIs for accessing NCBI databases (Gene, GEO, SRA, PubChem, etc).
|
||||
- [UniProt REST API](https://www.uniprot.org/help/api) — Programmatic access to protein sequence and functional annotation data.
|
||||
- [Ensembl REST API](https://rest.ensembl.org/) — API for genomic annotations, variants, genes, and comparative genomics.
|
||||
- [KEGG REST API](https://www.kegg.jp/kegg/rest/keggapi.html) — API for accessing KEGG pathways, compounds, genes, and reactions.
|
||||
- [ChEMBL Web Services](https://www.ebi.ac.uk/chembl/ws) — REST API for bioactive molecules, targets, and bioassays.
|
||||
- [Open Targets Platform API](https://platform.opentargets.org/api) — API for target–disease associations integrating genetics, genomics, and drug data.
|
||||
- [ClinicalTrials.gov API](https://clinicaltrials.gov/api/gui) — API for querying clinical trial metadata and results.
|
||||
|
||||
---
|
||||
|
||||
## Preprocessing Tools
|
||||
|
||||
- [Chemistry Development Kit](https://github.com/cdk/cdk) — Cheminformatics software & machine learning tools.
|
||||
- [Biopython](https://biopython.org/) — Collection of Python tools for biological computation including sequence analysis, structure parsing, and database access.
|
||||
- [FlashDeconv](https://github.com/cafferychen777/flashdeconv) — High-performance spatial transcriptomics deconvolution (~1M spots in ~3 min).
|
||||
- [RDKit](https://github.com/rdkit/rdkit) — Cheminformatics software & machine learning toolkit.
|
||||
- [DeepChem](https://github.com/deepchem/deepchem) — Deep learning library for drug discovery, quantum chemistry, and materials science.
|
||||
- [ChatSpatial](https://github.com/cafferychen777/ChatSpatial) — MCP server for spatial transcriptomics analysis via natural language.
|
||||
- [Scanpy](https://scanpy.readthedocs.io/en/stable/) — Python library for scRNA-seq analysis.
|
||||
- [Seurat](https://satijalab.org/seurat/) — R library for scRNA-seq analysis.
|
||||
- [scvi-tools](https://scvi-tools.org/) — Probabilistic models for single-cell omics data analysis.
|
||||
- [CellTypist](https://github.com/Teichlab/celltypist) — Automated cell type annotation for scRNA-seq.
|
||||
- [Squidpy](https://squidpy.readthedocs.io/) — Python library for spatial single-cell analysis.
|
||||
- [GROMACS](https://www.gromacs.org/) — Molecular dynamics simulation package for biochemical molecules.
|
||||
- [MDAnalysis](https://www.mdanalysis.org/) — Python library for analyzing and altering molecular dynamics simulation trajectories.
|
||||
- [OpenMM](https://openmm.org/) — High-performance toolkit for molecular simulation and GPU-accelerated MD.
|
||||
- [scVelo](https://github.com/theislab/scvelo) — RNA velocity estimation for single-cell transcriptomics, inferring the direction and speed of cell differentiation.
|
||||
- [STAR](https://github.com/alexdobin/STAR) — Ultrafast universal RNA-seq aligner with support for spliced alignment and single-cell quantification via STARsolo.
|
||||
- [kallisto](https://pachterlab.github.io/kallisto/) — Near-optimal RNA-seq quantification using pseudoalignment for fast transcript abundance estimation.
|
||||
- [Harmony](https://github.com/immunogenomics/harmony) — Fast and scalable integration of single-cell data across datasets, conditions, technologies, and species.
|
||||
- [Monocle3](https://cole-trapnell-lab.github.io/monocle3/) — Single-cell trajectory analysis tool for learning developmental trajectories and ordering cells in pseudotime.
|
||||
- [CellChat](https://github.com/sqjin/CellChat) — Inference and analysis of cell-cell communication ligand-receptor networks from single-cell transcriptomics data.
|
||||
- [SCENIC](https://github.com/aertslab/SCENIC) — Single-cell regulatory network inference and clustering linking transcription factors to co-expressed gene modules.
|
||||
- [DoubletFinder](https://github.com/chris-mcginnis-ucsf/DoubletFinder) — Machine learning approach for detecting multiplet (doublet) artifacts in single-cell RNA-seq data.
|
||||
- [Numbat](https://github.com/kharchenkolab/numbat) — Haplotype-aware copy number variation inference from single-cell RNA-seq using hidden Markov models.
|
||||
- [CaSpER](https://github.com/akdess/CaSpER) — CNV identification and visualization by integrative analysis of single-cell or bulk RNA-seq data.
|
||||
- [CellCharter](https://github.com/CSOgroup/cellcharter) — Identification and characterization of spatial cell niches from spatial transcriptomics using VAEs and Gaussian mixture models.
|
||||
- [STAGATE](https://github.com/RucDongLab/STAGATE) — Adaptive graph attention auto-encoder for spatial domain identification in spatial transcriptomics.
|
||||
- [NCEM](https://github.com/theislab/ncem) — GNN-based model for learning intercellular communication from spatial graphs of cells.
|
||||
- [DeepTalk](https://github.com/JiangBioLab/DeepTalk) — Graph attention network for deciphering cell-cell communication from spatial transcriptomics data.
|
||||
- [COMMOT](https://github.com/zcang/COMMOT) — Optimal transport-based framework for screening cell-cell communication in spatial transcriptomics.
|
||||
- [TIGON](https://github.com/yutongo/TIGON) — Neural optimal transport method for reconstructing growth and dynamic trajectories from single-cell transcriptomics.
|
||||
- [LINGER](https://github.com/Durenlab/LINGER) — Neural network for gene regulatory network inference from single-cell multiome (RNA+ATAC-seq) data with bulk data pretraining.
|
||||
- [sciPENN](https://github.com/jlakkis/sciPENN) — RNN-based method for simultaneous protein expression prediction, uncertainty estimation, and cell-type label transfer from CITE-seq and scRNA-seq data.
|
||||
- [MOGONET](https://github.com/txWang/MOGONET) — Multi-omics graph convolutional network framework for patient classification and biomarker identification.
|
||||
- [AutoZyme](https://github.com/ElliotXie/autozyme) — Autonomous agentic framework that speeds up bioinformatics software (e.g. Scanpy, Seurat) on CPUs while preserving the original results.
|
||||
- [SeqBench](https://seqbench.com/) — Web-based molecular biology sequence workbench for primer design, cloning simulation (Gibson, Golden Gate, restriction digest), CRISPR guide RNA design, and sequence analysis, with a public REST API, OpenAPI 3.1 spec, and MCP server.
|
||||
|
||||
---
|
||||
|
||||
## Machine Learning Tasks and Models
|
||||
|
||||
### Drug Discovery
|
||||
|
||||
#### Drug Response Prediction
|
||||
|
||||
- [drGAT](https://github.com/inoue0426/drGAT) — Attention-based model for drug response prediction with gene explainability.
|
||||
- [MOFGCN](https://github.com/weiba/MOFGCN/tree/main) — GCN + heterogeneous network.
|
||||
- [DeepDSC](https://ieeexplore-ieee-org.ezp2.lib.umn.edu/stamp/stamp.jsp?tp=&arnumber=8723620&tag=1) — Autoencoder + fully connected NN.
|
||||
- [DGDRP](https://github.com/minwoopak/heteronet) — Multi-view embedding neural network.
|
||||
- [DeepAEG](https://github.com/zhejiangzhuque/DeepAEG) — GNN embedding + attention mechanism.
|
||||
- [RECOVER](https://github.com/RECOVERcoalition/Recover) — Machine learning framework for predicting synergistic drug combination responses across cell lines.
|
||||
- [TGSA](https://github.com/violet-sto/TGSA) — Tumor gene set and attention-based model leveraging biological pathway knowledge for drug response prediction.
|
||||
- [HiDRA](https://github.com/bsml320/HiDRA) — Hierarchical network model incorporating gene and pathway-level information for cancer drug response prediction.
|
||||
- [DRUML](https://github.com/CutillasLab/DRUMLR) — Ensemble machine learning framework combining standard ML with deep learning to systematically rank anti-cancer drugs from proteomics and RNA-seq data.
|
||||
|
||||
#### Drug Perturbation
|
||||
|
||||
- [CellOT](https://github.com/bunnech/cellot) — Neural optimal transport framework for predicting single-cell responses to drug and genetic perturbations.
|
||||
- [CMonge](https://github.com/AI4SCR/conditional-monge-gap) — Conditional optimal transport model for generalizable single-cell perturbation response prediction across drugs and doses.
|
||||
- [chemCPA](https://github.com/theislab/chemCPA) — Compositional perturbation autoencoder for predicting single-cell transcriptional responses to unseen drug perturbations and dose combinations.
|
||||
- [cycleCDR](https://github.com/hliulab/cycleCDR) — Interpretable cycle-consistency framework for modeling cellular responses to drug perturbations.
|
||||
- [PRNet](https://github.com/Perturbation-Response-Prediction/PRnet) — Deep generative model for predicting transcriptional responses to novel chemical perturbations for drug discovery.
|
||||
|
||||
#### Drug Repurposing
|
||||
|
||||
- [DeepPurpose](https://github.com/kexinhuang12345/DeepPurpose) — Deep learning library for drug repurposing.
|
||||
- [TranSiGen](https://github.com/myzhengSIMM/TranSiGen) — Dual-VAE architecture for ligand-based virtual screening, drug response prediction, and drug repurposing using chemical-induced transcriptional profiles.
|
||||
|
||||
#### Drug Target Interaction
|
||||
|
||||
- [NeoDTI](https://github.com/FangpingWan/NeoDTI) — Library for drug-target interaction prediction.
|
||||
- [DTINet](https://github.com/luoyunan/DTINet) — Network-based framework integrating heterogeneous biological data for DTI prediction.
|
||||
- [DeepDTA](https://github.com/hkmztrk/DeepDTA) — Deep learning model using CNNs on protein sequences and drug SMILES.
|
||||
- [GraphDTA](https://github.com/thinng/GraphDTA) — Graph neural network–based DTI prediction using molecular graphs.
|
||||
- [MolTrans](https://github.com/kexinhuang12345/MolTrans) — Transformer-based DTI model leveraging molecular substructures.
|
||||
- [DrugBAN](https://github.com/peizhenbai/DrugBAN) — Bilinear attention network for interpretable DTI prediction.
|
||||
|
||||
#### Compound-Protein Interaction
|
||||
|
||||
- [MCPINN](https://github.com/mhlee0903/multi_channels_PINN) — Drug discovery via compound-protein interaction and machine learning.
|
||||
- [TransformerCPI](https://github.com/lifanchen-simm/transformerCPI) — CPI prediction using Transformer.
|
||||
|
||||
#### Molecular Generation
|
||||
|
||||
- [REINVENT](https://github.com/MolecularAI/Reinvent) — Reinforcement learning for de novo drug design.
|
||||
- [MolGPT](https://github.com/devalab/molgpt) — Transformer-based model for molecular generation.
|
||||
- [Molecular Transformer](https://github.com/pschwllr/MolecularTransformer) — Sequence-to-sequence model for retrosynthesis prediction.
|
||||
- [Matcha](https://github.com/LigandPro/Matcha) — Multi-stage Riemannian flow matching model for physically valid molecular docking with scoring, pose filtering, and benchmarks.
|
||||
- [TargetDiff](https://github.com/guanjq/targetdiff) — 3D equivariant diffusion model for structure-based drug design.
|
||||
- [DiffDock](https://github.com/gcorso/DiffDock) — Diffusion generative model for molecular docking, predicting the binding pose of small molecules to protein targets.
|
||||
- [JTVAE](https://github.com/wengong-jin/icml18-jtnn) — Junction tree variational autoencoder for molecular graph generation that guarantees chemical validity via a hierarchical tree decomposition.
|
||||
- [DiffSBDD](https://github.com/arneschneuing/DiffSBDD) — Equivariant diffusion model for structure-based drug design that generates molecules and binding conformations for protein targets.
|
||||
- [ReLeaSE](https://github.com/isayev/ReLeaSE) — Deep reinforcement learning framework for de novo drug design combining a generative and predictive model.
|
||||
- [PaccMannRL](https://github.com/PaccMann/paccmann_generator) — Reinforcement learning-based generative model for de novo hit-like anticancer molecule design from transcriptomic data.
|
||||
|
||||
### LLM for Biology
|
||||
|
||||
- [AI4Chem/ChemLLM-7B-Chat](https://huggingface.co/AI4Chem/ChemLLM-7B-Chat) — LLM for chemical & molecular science.
|
||||
- [BioGPT](https://github.com/microsoft/BioGPT) — LLM for biomedical text generation.
|
||||
- [GeneGPT](https://github.com/ncbi/GeneGPT) — LLM for biomedical information, integrated with various APIs.
|
||||
- [GenePT](https://github.com/yiqunchen/GenePT) — Foundation LLM for single-cell data.
|
||||
- [scPRINT](https://github.com/cantinilab/scPRINT) — Pretrained on 50M cells for scRNA-seq denoising & zero imputation.
|
||||
- [ClawBio](https://github.com/ClawBio/ClawBio) — Bioinformatics-native AI agent skill library with local-first pharmacogenomics, ancestry PCA, semantic similarity, nutrigenomics, and metagenomics skills.
|
||||
- [BioMedLM](https://huggingface.co/stanford-crfm/BioMedLM) — 2.7B parameter GPT-2-style language model trained exclusively on biomedical literature from PubMed for biomedical question answering and text generation.
|
||||
- [MolT5](https://github.com/blender-nlp/MolT5) — Language model for molecular tasks bridging text and SMILES, enabling molecule captioning and text-driven molecule generation.
|
||||
- [ChatDrug](https://github.com/chao1224/ChatDrug) — LLM-based conversational pipeline for drug discovery, using natural language prompts for iterative drug editing and optimization.
|
||||
- [CASSIA](https://github.com/ElliotXie/CASSIA) — Multi-agent LLM for reference-free, interpretable cell-type annotation of single-cell RNA-seq data, with dedicated annotation, validation, scoring, and reporting agents.
|
||||
|
||||
### Foundation Models
|
||||
|
||||
#### Single-cell Foundation Models
|
||||
|
||||
##### Transcriptomics Foundation Models
|
||||
|
||||
- [scFoundation](https://github.com/biomap-research/scFoundation) — Large-scale foundation model for single-cell gene expression, enabling multiple downstream tasks.
|
||||
- [scGPT](https://github.com/bowang-lab/scGPT) — Transformer-based foundation model pretrained on millions of single-cell profiles.
|
||||
- [Geneformer](https://huggingface.co/ctheodoris/Geneformer) — Context-aware, attention-based deep learning model pretrained on a large corpus of single-cell transcriptomes.
|
||||
- [BulkFormer](https://github.com/KangBoming/BulkFormer) — Foundation model for bulk RNA-seq data; learns general transcriptomic representations.
|
||||
- [scBERT](https://github.com/TencentAILabHealthcare/scBERT) — BERT-based foundation model pretrained on large-scale scRNA-seq data for cell type annotation.
|
||||
- [CellPLM](https://github.com/OmicsML/CellPLM) — Cell pre-trained language model with inter-cell transformer architecture for diverse single-cell analysis tasks.
|
||||
- [UCE](https://github.com/snap-stanford/UCE) — Universal Cell Embeddings: zero-shot single-cell embedding model trained on 36M cells across species, tissues, and assays without fine-tuning.
|
||||
- [GEARS](https://github.com/snap-stanford/GEARS) — Graph-based model for predicting transcriptional responses to single and combinatorial genetic perturbations using biological priors.
|
||||
- [SATURN](https://github.com/snap-stanford/SATURN) — Transformer-based model integrating gene expression and protein sequences via a protein language model to learn unified multi-species cell embeddings.
|
||||
- [CancerFoundation](https://github.com/BoevaLab/CancerFoundation) — Single-cell RNA-seq foundation model trained exclusively on a curated dataset of malignant cells to learn cancer-specific embeddings.
|
||||
|
||||
##### Spatial Foundation Models
|
||||
|
||||
- [GigaPath](https://github.com/prov-gigapath/prov-gigapath) — Slide-level digital pathology foundation model pretrained on 1.3 billion pathology image tokens from whole-slide images.
|
||||
- [UNI](https://github.com/mahmoodlab/UNI) — General-purpose self-supervised pathology foundation model trained on 100K+ whole-slide images for diverse computational pathology tasks.
|
||||
- [CONCH](https://github.com/mahmoodlab/CONCH) — Vision-language foundation model for computational pathology trained with contrastive captioning on pathology image–text pairs.
|
||||
- [Phikon](https://huggingface.co/owkin/phikon) — ViT-based pathology foundation model pretrained with iBOT self-supervision on TCGA whole-slide images.
|
||||
- [Nicheformer](https://github.com/theislab/nicheformer) — Foundation model for single-cell and spatial omics using a transformer architecture with positional embeddings to encode spatial cell information.
|
||||
- [scGPT-spatial](https://github.com/bowang-lab/scGPT-spatial) — Extension of scGPT for spatial transcriptomics with continual pretraining and a mixture-of-experts decoder for spatial gene expression analysis.
|
||||
- [DeepSpot](https://github.com/ratschlab/DeepSpot) — Deep learning model predicting spatial transcriptomics from H&E images at spot and single-cell resolution.
|
||||
- [DeepSpot2Cell](https://github.com/ratschlab/DeepSpot2Cell) — Predicts virtual single-cell spatial transcriptomics from H&E using spot-level supervision (NeurIPS 2025 Imageomics).
|
||||
- [DeepSpot-M](https://github.com/ratschlab/DeepSpotM) — Multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology.
|
||||
- [AESTETIK](https://github.com/ratschlab/aestetik) — Autoencoder for spatial transcriptomics representation learning using topology and histology image knowledge.
|
||||
|
||||
##### Multi-Omics Foundation Models
|
||||
|
||||
- [scMulan](https://github.com/SuperBianC/scMulan) — Single-cell multi-omic language model pretrained on ~10M cells spanning transcriptomics, epigenomics, and proteomics for cross-omics transfer tasks.
|
||||
- [totalVI](https://github.com/scverse/scvi-tools) — Probabilistic framework for joint analysis of paired scRNA-seq and protein (CITE-seq) data enabling multi-modal cell state representation across single-cell datasets.
|
||||
- [MultiVI](https://github.com/scverse/scvi-tools) — Multi-modal variational autoencoder for integrating paired and unpaired single-cell RNA-seq and ATAC-seq measurements into a unified latent space.
|
||||
- [MIRA](https://github.com/cistrome/MIRA) — Probabilistic multimodal topic model jointly modeling single-cell transcriptomics and chromatin accessibility for regulatory network inference.
|
||||
- [GLUE](https://github.com/gao-lab/GLUE) — Graph-Linked Unified Embedding framework for unpaired single-cell multi-omics data integration across RNA, ATAC, methylation, and protein modalities.
|
||||
- [BABEL](https://github.com/wukevin/babel) — Cross-modality translation model enabling prediction between scRNA-seq and scATAC-seq profiles without requiring paired single-cell measurements.
|
||||
- [Multigrate](https://github.com/theislab/multigrate) — Asymmetric multi-omics variational autoencoder for integrating single-cell data across RNA, ATAC, and protein modalities with missing-modality support.
|
||||
- [MOFA+](https://github.com/bioFAM/MOFA2) — Multi-Omics Factor Analysis framework identifying shared axes of variation across bulk and single-cell datasets including RNA, ATAC, proteomics, methylation, and copy number.
|
||||
- [GeneCompass](https://github.com/xCompass-AI/GeneCompass) — Large-scale foundation model integrating DNA regulatory sequences and single-cell transcriptomics from 120M+ cells across multiple species for gene regulation prediction.
|
||||
- [UnitedNet](https://github.com/LiuLab-Bioelectronics-Harvard/UnitedNet) — Interpretable multi-task deep neural network for single-cell multi-omics integration spanning transcriptomics, chromatin accessibility, and proteomics.
|
||||
- [SpatialGlue](https://github.com/zhanglabtools/SpatialGlue) — Graph attention network for spatial multi-omics integration jointly embedding spatial transcriptomics with chromatin accessibility or proteomics.
|
||||
- [MIDAS](https://github.com/labomics/midas) — Mosaic integration and differential accessibility model for single-cell multi-omics data that handles arbitrary missing-modality combinations across transcriptomics, chromatin accessibility, and proteomics.
|
||||
- [Concerto](https://github.com/melobio/Concerto-reproducibility) — Contrastive self-supervised learning framework for single-cell multimodal data integration, batch correction, and reference-query mapping.
|
||||
- [scButterfly](https://github.com/BioX-NKU/scButterfly) — Dual-aligned variational autoencoder for single-cell cross-modality translation between paired and unpaired multiomics data.
|
||||
- [JAMIE](https://github.com/Oafish1/JAMIE) — Joint variational autoencoder for multimodal single-cell data imputation and embedding.
|
||||
- [scPair](https://github.com/quon-titative-biology/scPair) — Bidirectional feedforward network for single-cell multimodal analysis with cross-modality prediction leveraging single-cell atlases.
|
||||
|
||||
##### Domain Alignment
|
||||
|
||||
- [scArches](https://github.com/theislab/scarches) — Transfer learning framework for mapping new single-cell datasets onto pre-trained reference atlases across batches, conditions, and modalities.
|
||||
- [TOSICA](https://github.com/JackieHanlaopo/TOSICA) — Transformer-based framework for one-stop interpretable cell-type annotation supporting cross-dataset and cross-species transfer.
|
||||
|
||||
#### Compound Foundation Models
|
||||
|
||||
##### Compound Embedding
|
||||
|
||||
- [ChemBERTa-2](https://github.com/seyonechithrananda/bert-loves-chemistry) — RoBERTa-based molecular language model pretrained on SMILES for small-molecule representation learning.
|
||||
- [GROVER](https://github.com/tencent-ailab/grover) — Self-supervised graph transformer for large-scale molecular representation learning from unlabeled compounds.
|
||||
- [Mol2Vec](https://github.com/samoturk/mol2vec) — Unsupervised molecular embedding method inspired by Word2Vec for learning vector representations of chemical substructures.
|
||||
- [MolFormer](https://github.com/IBM/molformer) — Linear attention transformer pretrained on millions of SMILES strings for efficient molecular embeddings.
|
||||
- [Uni-Mol](https://github.com/deepmodeling/Uni-Mol) — 3D molecular pretraining framework for universal representation learning on molecules and protein pockets.
|
||||
|
||||
#### Protein Foundation Models
|
||||
|
||||
##### Pre-trained Embedding
|
||||
|
||||
- [Evolutionary Scale Modeling (ESM)](https://github.com/facebookresearch/esm) — Protein embeddings.
|
||||
- [ProtTrans](https://github.com/agemagician/ProtTrans) — Suite of protein language models (ProtBERT, ProtT5, ProtXLNet) trained on billions of protein sequences from UniRef and BFD.
|
||||
- [ProGen2](https://github.com/salesforce/progen) — Protein language model trained on diverse protein families for sequence generation and fitness prediction.
|
||||
- [Ankh](https://github.com/agemagician/Ankh) — Efficient protein language model optimized for downstream prediction tasks including secondary structure, localization, and function annotation.
|
||||
|
||||
##### Protein Structure Prediction and Design
|
||||
|
||||
- [AlphaFold3](https://github.com/google-deepmind/alphafold3) — Predicts structures of proteins, nucleic acids, small molecules, and their complexes.
|
||||
- [Boltz-1](https://github.com/jwohlwend/boltz) — Open-source all-atom biomolecular structure prediction model for proteins, nucleic acids, small molecules, and their complexes achieving AlphaFold3-level accuracy.
|
||||
- [Chai-1](https://github.com/chaidiscovery/chai-lab) — Unified molecular structure prediction model covering proteins, nucleic acids, small molecules, and complexes.
|
||||
- [ESM3](https://github.com/evolutionaryscale/esm) — Multimodal protein language model that jointly reasons over sequence, structure, and function for generative protein design and engineering.
|
||||
- [ESMFold](https://github.com/facebookresearch/esm) — Fast protein structure prediction using language model embeddings.
|
||||
- [RFdiffusion](https://github.com/RosettaCommons/RFdiffusion) — Generative model for protein backbone design using diffusion.
|
||||
- [ProteinMPNN](https://github.com/dauparas/ProteinMPNN) — Deep learning model for protein sequence design given backbone structure.
|
||||
- [OmegaFold](https://github.com/HeliXonProtein/OmegaFold) — High-resolution de novo protein structure prediction from sequence.
|
||||
- [RoseTTAFold](https://github.com/RosettaCommons/RoseTTAFold) — Three-track neural network for protein structure prediction.
|
||||
- [OpenFold](https://github.com/aqlaboratory/openfold) — Trainable, memory-efficient open-source reproduction of AlphaFold2 enabling custom protein structure prediction workflows.
|
||||
- [SaProt](https://github.com/westlake-reup/SaProt) — Structure-aware protein language model using structure-aware tokens that encode both sequence and backbone geometry for improved function prediction.
|
||||
- [EvoDiff](https://github.com/microsoft/evodiff) — Discrete diffusion framework for protein sequence generation trained on evolutionary-scale data, supporting unconditional generation, disordered region design, and functional motif scaffolding. [ [paper-2023](https://www.biorxiv.org/content/10.1101/2023.09.11.556673v1) ]
|
||||
|
||||
#### Multi-Modal Foundation Models
|
||||
|
||||
- [CHIEF](https://github.com/hms-dbmi/CHIEF) — Clinical Histopathology Imaging Evaluation Foundation model integrating histology images and clinical context for pan-cancer analysis.
|
||||
- [BiomedCLIP](https://huggingface.co/microsoft/BiomedCLIP-PubMedBERT_256-vit_g_14) — CLIP-based vision-language foundation model for biomedical images and text trained on PubMed figure–caption pairs.
|
||||
- [PORPOISE](https://github.com/mahmoodlab/PORPOISE) — Pan-cancer integrative histology-genomic analysis framework using multimodal deep learning for patient stratification.
|
||||
- [PathomicFusion](https://github.com/mahmoodlab/PathomicFusion) — Integrated framework fusing histopathology and genomic features via CNN, GNN, and attention gating for cancer diagnosis and prognosis.
|
||||
- [Virchow](https://huggingface.co/paige-ai/Virchow) — Million-slide digital pathology foundation model using a vision transformer and self-supervised distillation for tile-level pathology image representation.
|
||||
- [TOAD](https://github.com/mahmoodlab/TOAD) — Tumor Origin Assessment via Deep-learning; weakly-supervised multi-task model predicting cancer primary origin from H&E whole-slide images.
|
||||
- [PLIP](https://github.com/PathologyFoundation/plip) — Vision-language foundation model for pathology trained with contrastive learning on pathology image–text pairs for image classification and text-to-image retrieval.
|
||||
- [MUSK](https://github.com/lilab-stanford/MUSK) — Vision-language foundation model for precision oncology analyzing multimodal paired text and pathology image data for biomarker prediction and retrieval.
|
||||
|
||||
#### Genomics Foundation Models
|
||||
|
||||
- [Nucleotide Transformer](https://github.com/instadeepai/nucleotide-transformer) — Foundation model for genomic sequences across multiple species.
|
||||
- [DNABERT](https://github.com/jerryji1993/DNABERT) — Pre-trained bidirectional encoder for DNA sequence analysis.
|
||||
- [DNABERT-2](https://github.com/Zhihan1996/DNABERT_2) — Improved genome foundation model with efficient tokenization.
|
||||
- [Enformer](https://github.com/deepmind/deepmind-research/tree/master/enformer) — Transformer model predicting gene expression from DNA sequence.
|
||||
- [Basenji](https://github.com/calico/basenji) — Sequential regulatory activity prediction from DNA sequences.
|
||||
- [Caduceus](https://github.com/kuleshov-group/caduceus) — Bidirectional equivariant long-range DNA sequence model based on Mamba.
|
||||
- [Evo](https://github.com/evo-design/evo) — Long-context genomic foundation model (up to 1M tokens).
|
||||
- [HyenaDNA](https://github.com/HazyResearch/hyena-dna) — Long-range genomic foundation model handling sequences up to 1M tokens with sub-quadratic attention.
|
||||
- [Borzoi](https://github.com/calico/borzoi) — Extended successor to Enformer for predicting RNA-seq coverage from long genomic sequence windows (524 kb) with improved resolution.
|
||||
- [DeepSEA](http://deepsea.princeton.edu/) — Deep learning framework for predicting chromatin effects of sequence alterations with single-nucleotide sensitivity across thousands of chromatin features.
|
||||
- [Sei](https://github.com/FunctionLab/sei-framework) — Sequence-to-function framework learning a genome-wide regulatory activity code from DNA sequences for variant effect prediction.
|
||||
- [GPN (Genomic Pre-trained Network)](https://github.com/songlab-cal/gpn) — Masked language model for DNA sequences enabling zero-shot variant effect prediction without requiring functional annotations.
|
||||
|
||||
---
|
||||
|
||||
## Citation
|
||||
|
||||
If you use this list in papers, slides, or documentation, please cite this repository via [`CITATION.cff`](./CITATION.cff) (also available through GitHub's **Cite this repository** button).
|
||||
|
||||
## Curation Criteria (Strict)
|
||||
|
||||
To keep quality high, additions should meet all of the following:
|
||||
|
||||
- The resource is trustworthy and relevant to computational biology.
|
||||
- The primary link points to an official source (official docs, organization site, maintained repository, or official dataset page).
|
||||
- The resource has evidence of technical substance: ideally a peer-reviewed paper; at minimum a preprint or official technical documentation.
|
||||
- The description is factual and concise (no marketing copy).
|
||||
- Duplicate or near-duplicate entries should be avoided.
|
||||
|
||||
We generally do **not** accept entries that are only promotional pages, personal opinion posts, or generic blog posts without technical references.
|
||||
|
||||
## Update & Link Rot Policy
|
||||
|
||||
- Link validity is monitored by the [Link Check workflow](./.github/workflows/link-check.yml).
|
||||
- If a link repeatedly fails, maintainers may replace it with an official mirror/canonical URL or remove the entry until a stable URL is available.
|
||||
- Contributions fixing broken links are welcome and encouraged.
|
||||
|
||||
## Data Schema & Contribution Workflow
|
||||
|
||||
- Data schema reference: [`docs/data/SCHEMA.md`](./docs/data/SCHEMA.md).
|
||||
- Source-of-truth workflow:
|
||||
1. Edit/add resources in `README.md`.
|
||||
2. Regenerate machine-readable artifacts:
|
||||
- `python scripts/sync_resources_from_readme.py`
|
||||
- `python scripts/build_resources.py`
|
||||
3. Commit updated data files (`data/resources.yml`, `data/resources.json`, `data/resources.csv`, `docs/data/resources.json`) with your README change.
|
||||
- Contribution guide: [`contributing.md`](./contributing.md).
|
||||
@@ -0,0 +1,87 @@
|
||||
---
|
||||
title: "Contributor Covenant Code of Conduct"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/code-of-conduct.md
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Contributor Covenant Code of Conduct
|
||||
|
||||
## Our Pledge
|
||||
|
||||
In the interest of fostering an open and welcoming environment, we as
|
||||
contributors and maintainers pledge to making participation in our project and
|
||||
our community a harassment-free experience for everyone, regardless of age, body
|
||||
size, disability, ethnicity, gender identity and expression, level of experience,
|
||||
nationality, personal appearance, race, religion, or sexual identity and
|
||||
orientation.
|
||||
|
||||
## Our Standards
|
||||
|
||||
Examples of behavior that contributes to creating a positive environment
|
||||
include:
|
||||
|
||||
* Using welcoming and inclusive language
|
||||
* Being respectful of differing viewpoints and experiences
|
||||
* Gracefully accepting constructive criticism
|
||||
* Focusing on what is best for the community
|
||||
* Showing empathy towards other community members
|
||||
|
||||
Examples of unacceptable behavior by participants include:
|
||||
|
||||
* The use of sexualized language or imagery and unwelcome sexual attention or
|
||||
advances
|
||||
* Trolling, insulting/derogatory comments, and personal or political attacks
|
||||
* Public or private harassment
|
||||
* Publishing others' private information, such as a physical or electronic
|
||||
address, without explicit permission
|
||||
* Other conduct which could reasonably be considered inappropriate in a
|
||||
professional setting
|
||||
|
||||
## Our Responsibilities
|
||||
|
||||
Project maintainers are responsible for clarifying the standards of acceptable
|
||||
behavior and are expected to take appropriate and fair corrective action in
|
||||
response to any instances of unacceptable behavior.
|
||||
|
||||
Project maintainers have the right and responsibility to remove, edit, or
|
||||
reject comments, commits, code, wiki edits, issues, and other contributions
|
||||
that are not aligned to this Code of Conduct, or to ban temporarily or
|
||||
permanently any contributor for other behaviors that they deem inappropriate,
|
||||
threatening, offensive, or harmful.
|
||||
|
||||
## Scope
|
||||
|
||||
This Code of Conduct applies both within project spaces and in public spaces
|
||||
when an individual is representing the project or its community. Examples of
|
||||
representing a project or community include using an official project e-mail
|
||||
address, posting via an official social media account, or acting as an appointed
|
||||
representative at an online or offline event. Representation of a project may be
|
||||
further defined and clarified by project maintainers.
|
||||
|
||||
## Enforcement
|
||||
|
||||
Instances of abusive, harassing, or otherwise unacceptable behavior may be
|
||||
reported by contacting the project team at inoue019@umn.edu. All
|
||||
complaints will be reviewed and investigated and will result in a response that
|
||||
is deemed necessary and appropriate to the circumstances. The project team is
|
||||
obligated to maintain confidentiality with regard to the reporter of an incident.
|
||||
Further details of specific enforcement policies may be posted separately.
|
||||
|
||||
Project maintainers who do not follow or enforce the Code of Conduct in good
|
||||
faith may face temporary or permanent repercussions as determined by other
|
||||
members of the project's leadership.
|
||||
|
||||
## Attribution
|
||||
|
||||
This Code of Conduct is adapted from the [Contributor Covenant][homepage], version 1.4,
|
||||
available at [http://contributor-covenant.org/version/1/4][version]
|
||||
|
||||
[homepage]: http://contributor-covenant.org
|
||||
[version]: http://contributor-covenant.org/version/1/4/
|
||||
@@ -0,0 +1,77 @@
|
||||
---
|
||||
title: "Contribution Guidelines"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/contributing.md
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Contribution Guidelines
|
||||
|
||||
Contributions are welcome!
|
||||
|
||||
Please note that this project is released with a
|
||||
[Contributor Code of Conduct](code-of-conduct.md). By participating in this
|
||||
project you agree to abide by its terms.
|
||||
|
||||
## Pull Requests
|
||||
|
||||
- Search previous suggestions before making a new one, as yours may be a duplicate.
|
||||
- Add one link per pull request.
|
||||
- Prefer official and trustworthy sources (official docs, organization pages, maintained repositories, or official dataset pages).
|
||||
- Include supporting technical evidence for new resources:
|
||||
- Ideally a peer-reviewed publication.
|
||||
- At minimum, a preprint or official technical documentation.
|
||||
- Avoid submissions that are primarily promotional pages, generic blog posts, or opinion-only writeups.
|
||||
- Add the link:
|
||||
- `[name](http://example.com/)` - A short description ends with a period.
|
||||
- Keep descriptions concise.
|
||||
- Maintain alphabetical ordering where applicable.
|
||||
- Add a section if needed.
|
||||
- Add the section description.
|
||||
- Add the section title to the [Index](https://github.com/inoue0426/awesome-computational-biology#Contents).
|
||||
- Check your spelling and grammar.
|
||||
- Remove any trailing whitespace.
|
||||
- Send a pull request with the reason why the addition is awesome.
|
||||
- Use the following format for your pull request title:
|
||||
- Add user/repo - Short repo description
|
||||
|
||||
## Data Workflow (README and JSON)
|
||||
|
||||
- The curated source list is maintained in `README.md`.
|
||||
- Machine-readable files are generated from README:
|
||||
- `python scripts/sync_resources_from_readme.py`
|
||||
- `python scripts/build_resources.py`
|
||||
- For resource additions/edits, include updated generated files in the same PR:
|
||||
- `data/resources.yml`
|
||||
- `data/resources.json`
|
||||
- `data/resources.csv`
|
||||
- `docs/data/resources.json`
|
||||
- Field definitions and naming rules are documented in [`docs/data/SCHEMA.md`](docs/data/SCHEMA.md).
|
||||
|
||||
## GitHub Pages UI
|
||||
|
||||
- The UI reads `docs/data/resources.json`.
|
||||
- Search and filters are driven by these fields:
|
||||
- Search: `name`, `description`, `tasks`, `modalities`, `tags`
|
||||
- Filters: `type`, `tasks`, `modalities`
|
||||
|
||||
## Updates to Existing Links or Sections
|
||||
|
||||
- Improvements to the existing sections are welcome.
|
||||
- If you think a listed link is not awesome, feel free to submit an issue or pull request to begin the discussion.
|
||||
- Broken links are checked by CI; if you find one, please submit a fix to the canonical URL (or remove the entry if no stable canonical URL exists).
|
||||
|
||||
## Updating your PR
|
||||
|
||||
A lot of times, making a PR adhere to the standards above can be difficult.
|
||||
If the maintainers notice anything that we'd like changed, we'll ask you to
|
||||
edit your PR before we merge it. There's no need to open a new PR, just edit
|
||||
the existing one. If you're not sure how to do that,
|
||||
[here is a guide](https://github.com/RichardLitt/knowledge/blob/master/github/amending-a-commit-guide.md)
|
||||
on the different ways you can update your PR so that we can merge it.
|
||||
@@ -0,0 +1,179 @@
|
||||
---
|
||||
title: "Cspell"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/cspell.json
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
{
|
||||
"language": "en",
|
||||
"allowCompoundWords": true,
|
||||
"words": [
|
||||
"behavioural",
|
||||
"KEGG",
|
||||
"NCBI",
|
||||
"UCSC",
|
||||
"EMBL",
|
||||
"RCSB",
|
||||
"CASP",
|
||||
"Uniclust",
|
||||
"Reactome",
|
||||
"Bioactive",
|
||||
"biopolymers",
|
||||
"proteomics",
|
||||
"transcriptomics",
|
||||
"metagenomic",
|
||||
"metatranscriptomic",
|
||||
"CRISPR",
|
||||
"JASPAR",
|
||||
"druggable",
|
||||
"Toxicogenomics",
|
||||
"GDSC",
|
||||
"biomolecular",
|
||||
"DRKG",
|
||||
"Hetionet",
|
||||
"Eudra",
|
||||
"esearch",
|
||||
"efetch",
|
||||
"Ensembl",
|
||||
"Cheminformatics",
|
||||
"Deconv",
|
||||
"Scanpy",
|
||||
"Squidpy",
|
||||
"explainability",
|
||||
"MOFGCN",
|
||||
"Autoencoder",
|
||||
"DGDRP",
|
||||
"MCPINN",
|
||||
"Pretrained",
|
||||
"pretrained",
|
||||
"denoising",
|
||||
"transcriptomic",
|
||||
"CELLxGENE",
|
||||
"eukaryotic",
|
||||
"metabolites",
|
||||
"OMIM",
|
||||
"Mendelian",
|
||||
"DisGeNET",
|
||||
"GWAS",
|
||||
"IntAct",
|
||||
"Biopython",
|
||||
"MDAnalysis",
|
||||
"trajectories",
|
||||
"Geneformer",
|
||||
"equivariant",
|
||||
"HyenaDNA",
|
||||
"Hyena",
|
||||
"Caduceus",
|
||||
"Mamba",
|
||||
"retrosynthesis",
|
||||
"TargetDiff",
|
||||
"Chai",
|
||||
"Zuckerberg",
|
||||
"HMDB",
|
||||
"CTRP",
|
||||
"ADMET",
|
||||
"Omics",
|
||||
"omics",
|
||||
"omic",
|
||||
"OADB",
|
||||
"Gnify",
|
||||
"gnom",
|
||||
"Rfam",
|
||||
"Guaca",
|
||||
"deconvolution",
|
||||
"scvi",
|
||||
"pharmacogenomics",
|
||||
"nutrigenomics",
|
||||
"Giga",
|
||||
"Phikon",
|
||||
"TCGA",
|
||||
"Mulan",
|
||||
"epigenomics",
|
||||
"methylation",
|
||||
"ATAC",
|
||||
"MOFA",
|
||||
"TOSICA",
|
||||
"Boltz",
|
||||
"MPNN",
|
||||
"Enformer",
|
||||
"Velo",
|
||||
"BACE",
|
||||
"secretase",
|
||||
"Clin",
|
||||
"CPTAC",
|
||||
"Proteomic",
|
||||
"proteogenomic",
|
||||
"LINCS",
|
||||
"ogbl",
|
||||
"ogbg",
|
||||
"SIDER",
|
||||
"Muris",
|
||||
"Pfam",
|
||||
"PROSITE",
|
||||
"epigenome",
|
||||
"TRRUST",
|
||||
"kallisto",
|
||||
"pseudoalignment",
|
||||
"multiplet",
|
||||
"TGSA",
|
||||
"JTVAE",
|
||||
"miRBase",
|
||||
"miRNA",
|
||||
"ProtTrans",
|
||||
"ProtBERT",
|
||||
"ProGen",
|
||||
"Ankh",
|
||||
"DeepSEA",
|
||||
"RegNetwork",
|
||||
"ROADMAP",
|
||||
"FANTOM",
|
||||
"NeXtProt",
|
||||
"HiDRA",
|
||||
"MolT",
|
||||
"ChatDrug",
|
||||
"DoubletFinder",
|
||||
"pseudotime",
|
||||
"ligand",
|
||||
"SCENIC",
|
||||
"GPN",
|
||||
"Sei",
|
||||
"KIBA",
|
||||
"pharmacokinetics",
|
||||
"ADME",
|
||||
"Haplotype",
|
||||
"NCEM",
|
||||
"multiome",
|
||||
"MOGONET",
|
||||
"convolutional",
|
||||
"DRUML",
|
||||
"SBDD",
|
||||
"Pacc",
|
||||
"multiomics",
|
||||
"Pathomic",
|
||||
"PLIP",
|
||||
"Omni",
|
||||
"Bento",
|
||||
"FFPE",
|
||||
"Xenium",
|
||||
"Zyme",
|
||||
"Neur",
|
||||
"Imageomics",
|
||||
"AESTETIK",
|
||||
"CellOT",
|
||||
"CMonge",
|
||||
"bowang",
|
||||
"ctheodoris",
|
||||
"OpenAI",
|
||||
"GPT"
|
||||
],
|
||||
"ignorePaths": [
|
||||
"node_modules/**"
|
||||
]
|
||||
}
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
---
|
||||
title: "Provenance-backed single-cell and biomedical benchmark enrichment batch."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.benchmark-v2.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed single-cell and biomedical benchmark enrichment batch.
|
||||
resources:
|
||||
scmulan:
|
||||
entities: [cell, gene]
|
||||
methods: [language-model, transformer]
|
||||
modalities: [epigenomics, multi-omics, proteomics, single-cell-rna-seq, transcriptomics]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
github: https://github.com/SuperBianC/scMulan
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/SuperBianC/scMulan
|
||||
|
||||
proteingym:
|
||||
entities: [protein]
|
||||
modalities: [protein-sequence]
|
||||
tasks: [regression]
|
||||
github: https://github.com/OATML-Markslab/ProteinGym
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/OATML-Markslab/ProteinGym
|
||||
|
||||
lincs_l1000:
|
||||
entities: [cell, compound, gene]
|
||||
modalities: [transcriptomics]
|
||||
tasks: [perturbation-prediction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://lincsproject.org/LINCS/tools/workflows/find-the-best-place-to-obtain-the-lincs-l1000-data
|
||||
|
||||
prism:
|
||||
entities: [cell, drug]
|
||||
tasks: [drug-response-prediction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://depmap.org/portal/prism/
|
||||
|
||||
pharmgkb:
|
||||
entities: [drug, gene, phenotype, variant]
|
||||
modalities: [clinical, genomics]
|
||||
tasks: [drug-response-prediction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.pharmgkb.org/
|
||||
+91
@@ -0,0 +1,91 @@
|
||||
---
|
||||
title: "Provenance-backed database and API enrichment batch."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.database-api-v1.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed database and API enrichment batch.
|
||||
resources:
|
||||
chembl_web_services:
|
||||
entities: [molecule, protein]
|
||||
modalities: [chemical-structure]
|
||||
documentation: https://www.ebi.ac.uk/chembl/api/data/docs
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.ebi.ac.uk/chembl/api/data/docs
|
||||
|
||||
clinicaltrials_gov_api:
|
||||
entities: [disease, drug]
|
||||
modalities: [clinical]
|
||||
documentation: https://clinicaltrials.gov/data-api/api
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://clinicaltrials.gov/data-api/api
|
||||
|
||||
ensembl_rest_api:
|
||||
entities: [gene, genome, transcript, variant]
|
||||
modalities: [genomics]
|
||||
documentation: https://rest.ensembl.org/
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://rest.ensembl.org/
|
||||
|
||||
kegg_rest_api:
|
||||
entities: [compound, gene, pathway]
|
||||
documentation: https://www.kegg.jp/kegg/rest/keggapi.html
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.kegg.jp/kegg/rest/keggapi.html
|
||||
|
||||
ncbi_e_utilities:
|
||||
entities: [gene, genome, protein, transcript, variant]
|
||||
modalities: [genomics, transcriptomics]
|
||||
documentation: https://www.ncbi.nlm.nih.gov/books/NBK25501/
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.ncbi.nlm.nih.gov/books/NBK25501/
|
||||
|
||||
open_targets_platform_api:
|
||||
entities: [disease, drug, gene, variant]
|
||||
modalities: [genomics, knowledge-graph]
|
||||
documentation: https://platform.opentargets.org/api
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://platform.opentargets.org/api
|
||||
|
||||
pubmed_e_utilities_esearch_efetch:
|
||||
documentation: https://www.ncbi.nlm.nih.gov/books/NBK25501/
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.ncbi.nlm.nih.gov/books/NBK25501/
|
||||
|
||||
uniprot_rest_api:
|
||||
entities: [protein]
|
||||
modalities: [protein-sequence, proteomics]
|
||||
documentation: https://www.uniprot.org/help/api
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.uniprot.org/help/api
|
||||
|
||||
drugbank:
|
||||
entities: [disease, drug, protein]
|
||||
modalities: [chemical-structure]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://go.drugbank.com/
|
||||
|
||||
string:
|
||||
entities: [protein]
|
||||
modalities: [knowledge-graph, proteomics]
|
||||
documentation: https://string-db.org/help/api/
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://string-db.org/
|
||||
- https://string-db.org/help/api/
|
||||
+73
@@ -0,0 +1,73 @@
|
||||
---
|
||||
title: "Provenance-backed foundation model enrichment batch 2."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.foundation-models-v2.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed foundation model enrichment batch 2.
|
||||
resources:
|
||||
nicheformer:
|
||||
entities: [cell, gene, tissue]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [single-cell-rna-seq, spatial-transcriptomics, transcriptomics]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
year: 2024
|
||||
github: https://github.com/theislab/nicheformer
|
||||
paper: https://doi.org/10.1101/2024.04.15.589472
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/theislab/nicheformer
|
||||
- https://doi.org/10.1101/2024.04.15.589472
|
||||
|
||||
genept:
|
||||
entities: [cell, gene]
|
||||
methods: [language-model]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks: [batch-correction, classification, representation-learning]
|
||||
year: 2023
|
||||
github: https://github.com/yiqunchen/GenePT
|
||||
paper: https://www.biorxiv.org/content/10.1101/2023.10.16.562533v2
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/yiqunchen/GenePT
|
||||
- https://www.biorxiv.org/content/10.1101/2023.10.16.562533v2
|
||||
|
||||
scgpt_spatial:
|
||||
entities: [cell, gene, tissue]
|
||||
methods: [generative-model, self-supervised-learning, transformer]
|
||||
modalities: [multi-omics, single-cell-rna-seq, spatial-transcriptomics]
|
||||
tasks: [foundation-model-pretraining, imputation, representation-learning]
|
||||
year: 2025
|
||||
github: https://github.com/bowang-lab/scGPT-spatial
|
||||
paper: https://www.biorxiv.org/content/10.1101/2025.02.05.636714v1
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/bowang-lab/scGPT-spatial
|
||||
- https://www.biorxiv.org/content/10.1101/2025.02.05.636714v1
|
||||
|
||||
scprint:
|
||||
entities: [cell, gene]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks:
|
||||
- batch-correction
|
||||
- cell-type-annotation
|
||||
- foundation-model-pretraining
|
||||
- gene-regulatory-network-inference
|
||||
- imputation
|
||||
- representation-learning
|
||||
year: 2025
|
||||
github: https://github.com/cantinilab/scPRINT
|
||||
documentation: https://www.jkobject.com/scPRINT/
|
||||
paper: https://www.nature.com/articles/s41467-025-58699-1
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/cantinilab/scPRINT
|
||||
- https://www.nature.com/articles/s41467-025-58699-1
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
---
|
||||
title: "Provenance-backed molecular model and benchmark enrichment batch."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.molecular-v2.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed molecular model and benchmark enrichment batch.
|
||||
resources:
|
||||
chemberta_2:
|
||||
entities: [molecule]
|
||||
methods: [language-model, self-supervised-learning, transformer]
|
||||
modalities: [chemical-structure]
|
||||
tasks: [representation-learning]
|
||||
github: https://github.com/seyonechithrananda/bert-loves-chemistry
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/seyonechithrananda/bert-loves-chemistry
|
||||
|
||||
molformer:
|
||||
entities: [molecule]
|
||||
methods: [language-model, self-supervised-learning, transformer]
|
||||
modalities: [chemical-structure]
|
||||
tasks: [representation-learning]
|
||||
github: https://github.com/IBM/molformer
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/IBM/molformer
|
||||
|
||||
grover:
|
||||
entities: [molecule]
|
||||
methods: [graph-neural-network, self-supervised-learning, transformer]
|
||||
modalities: [chemical-structure]
|
||||
tasks: [representation-learning]
|
||||
github: https://github.com/tencent-ailab/grover
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/tencent-ailab/grover
|
||||
|
||||
moleculenet:
|
||||
entities: [molecule]
|
||||
modalities: [chemical-structure]
|
||||
tasks: [classification, regression]
|
||||
github: https://github.com/deepchem/moleculenet
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/deepchem/moleculenet
|
||||
|
||||
guacamol:
|
||||
entities: [molecule]
|
||||
modalities: [chemical-structure]
|
||||
tasks: [molecular-generation]
|
||||
github: https://github.com/BenevolentAI/guacamol
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/BenevolentAI/guacamol
|
||||
+92
@@ -0,0 +1,92 @@
|
||||
---
|
||||
title: "Provenance-backed drug-response and pharmacogenomics enrichment batch."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.pharmacogenomics-v1.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed drug-response and pharmacogenomics enrichment batch.
|
||||
resources:
|
||||
beat_aml:
|
||||
entities: [cell, disease, drug, gene]
|
||||
modalities: [genomics]
|
||||
tasks: [drug-response-prediction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://biodev.github.io/BeatAML2/
|
||||
|
||||
cancer_therapeutics_response_portal_ctrp:
|
||||
entities: [cell, drug]
|
||||
tasks: [drug-response-prediction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://portals.broadinstitute.org/ctrp/
|
||||
|
||||
bindingdb_curated_sets:
|
||||
entities: [molecule, protein]
|
||||
modalities: [chemical-structure]
|
||||
tasks: [drug-target-interaction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.bindingdb.org/
|
||||
|
||||
bace:
|
||||
entities: [molecule, protein]
|
||||
modalities: [chemical-structure]
|
||||
tasks: [classification, regression]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.kaggle.com/datasets/gokturkkoch/bace
|
||||
|
||||
clintox:
|
||||
entities: [drug]
|
||||
modalities: [clinical]
|
||||
tasks: [classification]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://tdcommons.ai/single_pred_tasks/tox/#clintox
|
||||
|
||||
sider_side_effect_resource:
|
||||
entities: [drug, phenotype]
|
||||
modalities: [clinical]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- http://sideeffects.embl.de/
|
||||
|
||||
pk_db:
|
||||
entities: [drug]
|
||||
modalities: [clinical]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://pk-db.com/
|
||||
|
||||
scperturb:
|
||||
entities: [cell, drug, gene]
|
||||
modalities: [single-cell-rna-seq]
|
||||
tasks: [perturbation-prediction]
|
||||
github: https://github.com/sanderlab/scPerturb
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/sanderlab/scPerturb
|
||||
|
||||
genomics_of_drug_sensitivity_in_cancer_gdsc:
|
||||
entities: [cell, drug, gene]
|
||||
modalities: [genomics]
|
||||
tasks: [drug-response-prediction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://www.cancerrxgene.org/
|
||||
|
||||
cellminer_cross_database_cellminercdb:
|
||||
entities: [cell, drug, gene]
|
||||
modalities: [genomics]
|
||||
tasks: [drug-response-prediction]
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://discover.nci.nih.gov/cellminercdb/
|
||||
+62
@@ -0,0 +1,62 @@
|
||||
---
|
||||
title: "Provenance-backed protein and drug-discovery enrichment batch."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.protein-drug-v1.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed protein and drug-discovery enrichment batch.
|
||||
resources:
|
||||
esmfold:
|
||||
entities: [protein]
|
||||
methods: [language-model, transformer]
|
||||
modalities: [molecular-structure, protein-sequence]
|
||||
tasks: [representation-learning, structure-prediction]
|
||||
year: 2023
|
||||
github: https://github.com/facebookresearch/esm
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/facebookresearch/esm
|
||||
|
||||
proteinmpnn:
|
||||
entities: [protein]
|
||||
methods: [graph-neural-network, message-passing-neural-network]
|
||||
modalities: [molecular-structure, protein-sequence]
|
||||
tasks: [protein-sequence-design]
|
||||
year: 2022
|
||||
github: https://github.com/dauparas/ProteinMPNN
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/dauparas/ProteinMPNN
|
||||
|
||||
diffdock:
|
||||
entities: [molecule, protein]
|
||||
methods: [diffusion, geometric-deep-learning]
|
||||
modalities: [molecular-structure]
|
||||
tasks: [docking]
|
||||
year: 2023
|
||||
github: https://github.com/gcorso/DiffDock
|
||||
paper: https://openreview.net/forum?id=kKF8_K-mBbS
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/gcorso/DiffDock
|
||||
- https://openreview.net/forum?id=kKF8_K-mBbS
|
||||
|
||||
uni_mol:
|
||||
entities: [molecule, protein]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [chemical-structure, molecular-structure]
|
||||
tasks: [docking, representation-learning]
|
||||
year: 2023
|
||||
github: https://github.com/deepmodeling/Uni-Mol
|
||||
paper: https://openreview.net/forum?id=6K2RM6wVqKu
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/deepmodeling/Uni-Mol
|
||||
- https://openreview.net/forum?id=6K2RM6wVqKu
|
||||
+64
@@ -0,0 +1,64 @@
|
||||
---
|
||||
title: "Provenance-backed protein model enrichment batch."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.protein-v2.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed protein model enrichment batch.
|
||||
resources:
|
||||
esm3:
|
||||
entities: [protein]
|
||||
methods: [generative-model, language-model, transformer]
|
||||
modalities: [molecular-structure, protein-sequence]
|
||||
tasks: [protein-sequence-design, representation-learning]
|
||||
github: https://github.com/evolutionaryscale/esm
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/evolutionaryscale/esm
|
||||
|
||||
evolutionary_scale_modeling_esm:
|
||||
entities: [protein]
|
||||
methods: [language-model, self-supervised-learning, transformer]
|
||||
modalities: [protein-sequence]
|
||||
tasks: [representation-learning]
|
||||
github: https://github.com/facebookresearch/esm
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/facebookresearch/esm
|
||||
|
||||
prottrans:
|
||||
entities: [protein]
|
||||
methods: [language-model, self-supervised-learning, transformer]
|
||||
modalities: [protein-sequence]
|
||||
tasks: [representation-learning]
|
||||
github: https://github.com/agemagician/ProtTrans
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/agemagician/ProtTrans
|
||||
|
||||
progen2:
|
||||
entities: [protein]
|
||||
methods: [generative-model, language-model, transformer]
|
||||
modalities: [protein-sequence]
|
||||
tasks: [protein-sequence-design, representation-learning]
|
||||
github: https://github.com/salesforce/progen
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/salesforce/progen
|
||||
|
||||
alphafold3:
|
||||
entities: [molecule, protein, protein-complex]
|
||||
methods: [diffusion]
|
||||
modalities: [molecular-structure, protein-sequence]
|
||||
tasks: [structure-prediction]
|
||||
github: https://github.com/google-deepmind/alphafold3
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/google-deepmind/alphafold3
|
||||
+110
@@ -0,0 +1,110 @@
|
||||
---
|
||||
title: "Provenance-backed spatial transcriptomics and imaging enrichment batch."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.spatial-imaging-v1.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Provenance-backed spatial transcriptomics and imaging enrichment batch.
|
||||
resources:
|
||||
aestetik:
|
||||
entities: [cell, gene, tissue]
|
||||
methods: [autoencoder]
|
||||
modalities: [histopathology, spatial-transcriptomics]
|
||||
tasks: [representation-learning]
|
||||
github: https://github.com/ratschlab/aestetik
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/ratschlab/aestetik
|
||||
|
||||
conch:
|
||||
entities: [tissue]
|
||||
methods: [contrastive-learning, transformer]
|
||||
modalities: [histopathology, imaging]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
github: https://github.com/mahmoodlab/CONCH
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/mahmoodlab/CONCH
|
||||
|
||||
deepspot:
|
||||
entities: [gene, tissue]
|
||||
modalities: [histopathology, spatial-transcriptomics]
|
||||
tasks: [regression]
|
||||
github: https://github.com/ratschlab/DeepSpot
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/ratschlab/DeepSpot
|
||||
|
||||
deepspot_m:
|
||||
entities: [gene, tissue]
|
||||
modalities: [histopathology, spatial-transcriptomics, transcriptomics]
|
||||
tasks: [foundation-model-pretraining, regression]
|
||||
github: https://github.com/ratschlab/DeepSpotM
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/ratschlab/DeepSpotM
|
||||
|
||||
deepspot2cell:
|
||||
entities: [cell, gene, tissue]
|
||||
modalities: [histopathology, spatial-transcriptomics]
|
||||
tasks: [regression]
|
||||
github: https://github.com/ratschlab/DeepSpot2Cell
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/ratschlab/DeepSpot2Cell
|
||||
|
||||
gigapath:
|
||||
entities: [tissue]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [histopathology, imaging]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
github: https://github.com/prov-gigapath/prov-gigapath
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/prov-gigapath/prov-gigapath
|
||||
|
||||
phikon:
|
||||
entities: [tissue]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [histopathology, imaging]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
documentation: https://huggingface.co/owkin/phikon
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://huggingface.co/owkin/phikon
|
||||
|
||||
plip:
|
||||
entities: [tissue]
|
||||
methods: [contrastive-learning]
|
||||
modalities: [histopathology, imaging]
|
||||
tasks: [classification, representation-learning]
|
||||
github: https://github.com/PathologyFoundation/plip
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/PathologyFoundation/plip
|
||||
|
||||
uni:
|
||||
entities: [tissue]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [histopathology, imaging]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
github: https://github.com/mahmoodlab/UNI
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/mahmoodlab/UNI
|
||||
|
||||
hest_xenium_virtual_spatial_transcriptomics:
|
||||
entities: [cell, gene, tissue]
|
||||
modalities: [histopathology, spatial-transcriptomics, transcriptomics]
|
||||
tasks: [regression]
|
||||
documentation: https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics
|
||||
@@ -0,0 +1,128 @@
|
||||
---
|
||||
title: "AI4Bio landscape enrichment overlay"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/enrichment.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# AI4Bio landscape enrichment overlay
|
||||
#
|
||||
# README.md remains the canonical source for resource membership and basic fields.
|
||||
# Add richer, independently curated metadata here, keyed by the stable resource id.
|
||||
# scripts/build_resources.py merges these fields into generated JSON/CSV artifacts.
|
||||
#
|
||||
# Enrichment values should be source-verifiable. Controlled vocabulary fields are
|
||||
# validated against data/vocabulary.yml.
|
||||
|
||||
resources:
|
||||
scgpt:
|
||||
entities: [cell, gene]
|
||||
methods: [generative-model, self-supervised-learning, transformer]
|
||||
modalities: [multi-omics, single-cell-rna-seq, transcriptomics]
|
||||
tasks:
|
||||
- cell-type-annotation
|
||||
- foundation-model-pretraining
|
||||
- gene-regulatory-network-inference
|
||||
- perturbation-prediction
|
||||
- representation-learning
|
||||
year: 2024
|
||||
github: https://github.com/bowang-lab/scGPT
|
||||
documentation: https://scgpt.readthedocs.io/en/latest/
|
||||
paper: https://www.nature.com/articles/s41592-024-02201-0
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/bowang-lab/scGPT
|
||||
- https://www.nature.com/articles/s41592-024-02201-0
|
||||
|
||||
geneformer:
|
||||
entities: [cell, gene]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks:
|
||||
- classification
|
||||
- foundation-model-pretraining
|
||||
- perturbation-prediction
|
||||
- representation-learning
|
||||
year: 2023
|
||||
documentation: https://geneformer.readthedocs.io/
|
||||
paper: https://www.nature.com/articles/s41586-023-06139-9
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://huggingface.co/ctheodoris/Geneformer
|
||||
- https://www.nature.com/articles/s41586-023-06139-9
|
||||
|
||||
scfoundation:
|
||||
entities: [cell, gene]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks:
|
||||
- cell-type-annotation
|
||||
- drug-response-prediction
|
||||
- foundation-model-pretraining
|
||||
- perturbation-prediction
|
||||
- representation-learning
|
||||
year: 2024
|
||||
github: https://github.com/biomap-research/scFoundation
|
||||
paper: https://www.nature.com/articles/s41592-024-02305-7
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/biomap-research/scFoundation
|
||||
- https://www.nature.com/articles/s41592-024-02305-7
|
||||
|
||||
genecompass:
|
||||
entities: [cell, gene]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
year: 2024
|
||||
github: https://github.com/xCompass-AI/GeneCompass
|
||||
paper: https://www.nature.com/articles/s41422-024-01034-y
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/xCompass-AI/GeneCompass
|
||||
- https://www.nature.com/articles/s41422-024-01034-y
|
||||
|
||||
uce:
|
||||
entities: [cell]
|
||||
methods: [self-supervised-learning]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
year: 2026
|
||||
github: https://github.com/snap-stanford/UCE
|
||||
paper: https://www.nature.com/articles/s41586-026-10689-z
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/snap-stanford/UCE
|
||||
- https://www.nature.com/articles/s41586-026-10689-z
|
||||
|
||||
cellplm:
|
||||
entities: [cell, gene]
|
||||
methods: [self-supervised-learning, transformer]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks: [foundation-model-pretraining, representation-learning]
|
||||
year: 2023
|
||||
github: https://github.com/OmicsML/CellPLM
|
||||
paper: https://www.biorxiv.org/content/10.1101/2023.10.03.560734v1
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/OmicsML/CellPLM
|
||||
- https://www.biorxiv.org/content/10.1101/2023.10.03.560734v1
|
||||
|
||||
scbert:
|
||||
entities: [cell, gene]
|
||||
methods: [language-model, self-supervised-learning, transformer]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks: [cell-type-annotation, classification, foundation-model-pretraining]
|
||||
year: 2022
|
||||
github: https://github.com/TencentAILabHealthcare/scBERT
|
||||
paper: https://www.nature.com/articles/s42256-022-00534-z
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://github.com/TencentAILabHealthcare/scBERT
|
||||
- https://www.nature.com/articles/s42256-022-00534-z
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,105 @@
|
||||
---
|
||||
title: "Canonical vocabulary for new AI4Bio enrichment metadata."
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/data/vocabulary.yml
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Canonical vocabulary for new AI4Bio enrichment metadata.
|
||||
#
|
||||
# These values are enforced only for fields explicitly added through
|
||||
# data/enrichment.yml. README-derived legacy values remain backward compatible.
|
||||
# Canonical terms use lowercase kebab-case.
|
||||
|
||||
version: 1
|
||||
|
||||
controlled_fields:
|
||||
entities:
|
||||
- cell
|
||||
- compound
|
||||
- disease
|
||||
- drug
|
||||
- gene
|
||||
- genome
|
||||
- molecule
|
||||
- organism
|
||||
- pathway
|
||||
- phenotype
|
||||
- protein
|
||||
- protein-complex
|
||||
- regulatory-element
|
||||
- tissue
|
||||
- transcript
|
||||
- variant
|
||||
|
||||
methods:
|
||||
- autoencoder
|
||||
- contrastive-learning
|
||||
- convolutional-neural-network
|
||||
- diffusion
|
||||
- generative-model
|
||||
- geometric-deep-learning
|
||||
- graph-neural-network
|
||||
- knowledge-graph
|
||||
- language-model
|
||||
- message-passing-neural-network
|
||||
- multi-agent-system
|
||||
- optimal-transport
|
||||
- recurrent-neural-network
|
||||
- reinforcement-learning
|
||||
- retrieval-augmented-generation
|
||||
- self-supervised-learning
|
||||
- state-space-model
|
||||
- supervised-learning
|
||||
- transformer
|
||||
- unsupervised-learning
|
||||
- variational-autoencoder
|
||||
|
||||
modalities:
|
||||
- cell-painting
|
||||
- chemical-structure
|
||||
- clinical
|
||||
- dna-sequence
|
||||
- electronic-health-record
|
||||
- epigenomics
|
||||
- genomics
|
||||
- histopathology
|
||||
- imaging
|
||||
- knowledge-graph
|
||||
- metabolomics
|
||||
- molecular-structure
|
||||
- multi-omics
|
||||
- protein-sequence
|
||||
- proteomics
|
||||
- rna-sequence
|
||||
- single-cell-rna-seq
|
||||
- spatial-transcriptomics
|
||||
- transcriptomics
|
||||
|
||||
tasks:
|
||||
- batch-correction
|
||||
- cell-type-annotation
|
||||
- classification
|
||||
- dimensionality-reduction
|
||||
- docking
|
||||
- drug-response-prediction
|
||||
- drug-target-interaction
|
||||
- foundation-model-pretraining
|
||||
- gene-regulatory-network-inference
|
||||
- imputation
|
||||
- link-prediction
|
||||
- molecular-generation
|
||||
- perturbation-prediction
|
||||
- protein-function-prediction
|
||||
- protein-sequence-design
|
||||
- regression
|
||||
- representation-learning
|
||||
- structure-prediction
|
||||
- trajectory-inference
|
||||
- virtual-screening
|
||||
@@ -0,0 +1,73 @@
|
||||
---
|
||||
title: "AI4Bio Landscape Database"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/docs/AI4BIO_LANDSCAPE.md
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# AI4Bio Landscape Database
|
||||
|
||||
The landscape view treats the existing computational biology registry as a multidimensional database rather than a single hierarchical list.
|
||||
|
||||
## Design goals
|
||||
|
||||
- Keep the current curated resource records and generation pipeline intact.
|
||||
- Expose orthogonal facets so one resource can be explored by resource type, biological/ML task, data modality, organism, and domain tag.
|
||||
- Make the landscape useful without introducing a server or build-time dependency.
|
||||
- Keep the data model extensible for richer AI4Bio metadata over time.
|
||||
|
||||
## Current facet model
|
||||
|
||||
The landscape UI derives the following dimensions from `docs/data/resources.json`:
|
||||
|
||||
| Dimension | Source field | Example values |
|
||||
|---|---|---|
|
||||
| Resource type | `type` | `database`, `benchmark`, `model`, `toolkit`, `api` |
|
||||
| Task | `tasks` | `drug-response-prediction`, `cell-type-annotation`, `molecular-generation` |
|
||||
| Modality | `modalities` | `transcriptomics`, `spatial-transcriptomics`, `protein-sequence` |
|
||||
| Organism | `organism` | `human`, `mouse`, `multi-species` |
|
||||
| Domain/tag | `tags` | `drug-discovery`, `single-cell`, `foundation-model` |
|
||||
|
||||
These are deliberately treated as separate axes. A model can therefore be, for example, a `model` that performs `perturbation-prediction` on `single-cell-rna-seq` data for `human` and carry tags such as `drug-discovery` and `foundation-model`.
|
||||
|
||||
## Recommended schema evolution
|
||||
|
||||
The current schema is compatible with a richer landscape database. New fields should be added incrementally and only when they can be curated consistently.
|
||||
|
||||
Suggested fields:
|
||||
|
||||
| Field | Type | Purpose |
|
||||
|---|---|---|
|
||||
| `entities` | array of strings | Biological entities such as `gene`, `protein`, `compound`, `cell`, `disease` |
|
||||
| `methods` | array of strings | Method families such as `transformer`, `gnn`, `diffusion`, `optimal-transport` |
|
||||
| `organizations` | array of strings | Primary organizations responsible for the resource |
|
||||
| `year` | integer | Initial public release/publication year |
|
||||
| `github` | string | Source repository when distinct from the canonical landing page |
|
||||
| `documentation` | string | Documentation URL |
|
||||
| `maintenance_status` | string | Curated status such as `active`, `maintenance`, `archived`, `unknown` |
|
||||
| `last_checked` | string | Date the metadata/link was last manually or automatically checked |
|
||||
|
||||
Avoid adding dynamic popularity metrics such as GitHub stars directly to canonical records unless a reproducible refresh pipeline is introduced. Such values become stale quickly and should be stored as generated metadata rather than curated facts.
|
||||
|
||||
## Canonical-source policy
|
||||
|
||||
At present, `README.md` is the canonical curated list, with generated YAML/JSON/CSV artifacts. The landscape page intentionally consumes `docs/data/resources.json` without changing that policy.
|
||||
|
||||
A future migration may make `data/resources.yml` the canonical source once all README-only categorization semantics can be represented explicitly in structured fields. That migration should be a separate change because it changes contribution workflow and source-of-truth semantics.
|
||||
|
||||
## Landscape page
|
||||
|
||||
Open `docs/landscape.html` through GitHub Pages. It provides:
|
||||
|
||||
- full-text search across names, descriptions, tasks, modalities, organisms, and tags;
|
||||
- filters for type, task, modality, organism, and tag;
|
||||
- summary counts for resources and major dimensions;
|
||||
- frequency bars recalculated for the current filtered result set;
|
||||
- direct resource and paper links;
|
||||
- client-side rendering with no additional dependencies.
|
||||
+72
@@ -0,0 +1,72 @@
|
||||
---
|
||||
title: "Foundation Model Enrichment"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/docs/FOUNDATION_MODEL_ENRICHMENT.md
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Foundation Model Enrichment
|
||||
|
||||
This document tracks the first curated metadata-enrichment pass for AI4Bio foundation models.
|
||||
|
||||
## Scope
|
||||
|
||||
The initial pass focuses on representative single-cell and transcriptomics foundation models already present in the resource registry, beginning with:
|
||||
|
||||
- scGPT
|
||||
- Geneformer
|
||||
|
||||
The scope may be expanded incrementally once the curation rules below are validated in practice.
|
||||
|
||||
## Curation rules
|
||||
|
||||
Metadata must be supported by at least one primary or official source:
|
||||
|
||||
- official project repository or model card;
|
||||
- official documentation;
|
||||
- primary peer-reviewed publication or preprint.
|
||||
|
||||
Unknown or ambiguous metadata is omitted rather than inferred.
|
||||
|
||||
For each resource, curate fields where evidence is available:
|
||||
|
||||
- `entities`
|
||||
- `methods`
|
||||
- `organizations`
|
||||
- `year`
|
||||
- `github`
|
||||
- `documentation`
|
||||
- `maintenance_status`
|
||||
- `last_checked`
|
||||
- `metadata_sources`
|
||||
|
||||
`maintenance_status` should only be marked `active` when there is direct evidence of ongoing maintenance, such as a recent official release or repository activity. Otherwise use `unknown` or omit the field.
|
||||
|
||||
## Initial evidence targets
|
||||
|
||||
### scGPT
|
||||
|
||||
Primary evidence should include the official `bowang-lab/scGPT` repository and the Nature Methods publication.
|
||||
|
||||
### Geneformer
|
||||
|
||||
Primary evidence should include the official `ctheodoris/Geneformer` model repository/model card and the primary Nature publication.
|
||||
|
||||
## Completion criteria
|
||||
|
||||
A resource is considered enriched when:
|
||||
|
||||
1. all added metadata is supported by `metadata_sources`;
|
||||
2. no unsupported organization, method, year, or maintenance claim is introduced;
|
||||
3. generated JSON/CSV artifacts are regenerated and committed;
|
||||
4. schema validation and resource-consistency CI checks pass.
|
||||
|
||||
## Provenance
|
||||
|
||||
This enrichment pass is being prepared with assistance from OpenAI GPT-5.6 Sol. Final metadata is intended to remain source-verifiable and reviewable through the recorded provenance URLs.
|
||||
@@ -0,0 +1,21 @@
|
||||
---
|
||||
title: "AI4Bio data files"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/docs/data/README.md
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# AI4Bio data files
|
||||
|
||||
- `resources.json`: generated merged resource registry consumed by GitHub Pages.
|
||||
- `resource.schema.json`: JSON Schema 2020-12 contract for one resource object.
|
||||
- `SCHEMA.md`: original schema notes.
|
||||
- `SCHEMA_V2.md`: richer AI4Bio landscape schema and enrichment workflow.
|
||||
|
||||
The enriched build path is `scripts/build_resources_v2.py`, which combines `data/resources.yml` with `data/enrichment.yml` and runs `scripts/validate_resources.py` before writing artifacts.
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
title: "Resource Data Schema (`docs/data/resources.json`)"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/docs/data/SCHEMA.md
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Resource Data Schema (`docs/data/resources.json`)
|
||||
|
||||
This document describes the JSON schema used by the GitHub Pages UI.
|
||||
|
||||
## Source of truth and generation flow
|
||||
|
||||
- **Canonical list source:** `README.md` (curated resource bullets)
|
||||
- Generated from README to YAML: `scripts/sync_resources_from_readme.py` → `data/resources.yml`
|
||||
- Built artifacts from YAML: `scripts/build_resources.py` → `data/resources.json`, `data/resources.csv`, and `docs/data/resources.json`
|
||||
|
||||
When contributing new resources, update `README.md` first, then regenerate artifacts.
|
||||
|
||||
## Top-level structure
|
||||
|
||||
- `resources.json` is a JSON array.
|
||||
- Each array item is one resource object.
|
||||
|
||||
## Fields
|
||||
|
||||
### Required fields
|
||||
|
||||
| Field | Type | Notes |
|
||||
|---|---|---|
|
||||
| `id` | string | Unique slug. Use lowercase `snake_case`, stable over time. |
|
||||
| `name` | string | Display name shown in README/UI. |
|
||||
| `type` | string | Resource category. Current values: `api`, `benchmark`, `database`, `model`, `toolkit`. |
|
||||
| `url` | string | Canonical landing page URL. |
|
||||
| `description` | string | One-line, factual summary. |
|
||||
|
||||
### Optional fields
|
||||
|
||||
| Field | Type | Notes |
|
||||
|---|---|---|
|
||||
| `tags` | array of strings | Free-form tags. |
|
||||
| `tasks` | array of strings | Task labels used by Task filter. |
|
||||
| `modalities` | array of strings | Data modality labels used by Modality filter. |
|
||||
| `organism` | array of strings | Organism labels. |
|
||||
| `license` | string | SPDX identifier preferred when known. |
|
||||
| `api` | boolean | Whether programmatic API access is available. Defaults to `false`. |
|
||||
| `paper` | string | DOI or URL to preprint/peer-reviewed publication. |
|
||||
| `updated` | string | Last-known update date, recommended `YYYY-MM-DD`. |
|
||||
|
||||
## Naming and consistency guidance
|
||||
|
||||
- `id` must be globally unique across all resources.
|
||||
- Prefer concise, stable IDs (e.g., `open_targets_platform`, `alphafold3`).
|
||||
- Keep `name` aligned with official project/database naming.
|
||||
- Use short, objective descriptions (avoid marketing language).
|
||||
|
||||
## Example object
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "open_targets_platform",
|
||||
"name": "Open Targets Platform",
|
||||
"type": "database",
|
||||
"url": "https://platform.opentargets.org/",
|
||||
"description": "Target identification platform integrating genetics, genomics, and drug evidence.",
|
||||
"tags": ["disease", "drug-discovery"],
|
||||
"tasks": ["target-identification"],
|
||||
"modalities": ["genomics"],
|
||||
"organism": ["human"],
|
||||
"license": "CC-BY-4.0",
|
||||
"api": true,
|
||||
"paper": "https://doi.org/10.1093/nar/gkac1045",
|
||||
"updated": "2026-01-15"
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,132 @@
|
||||
---
|
||||
title: "AI4Bio Resource Schema v2"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/docs/data/SCHEMA_V2.md
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# AI4Bio Resource Schema v2
|
||||
|
||||
This document defines the richer landscape metadata layered on top of the curated Awesome Computational Biology list.
|
||||
|
||||
## Source model
|
||||
|
||||
The repository intentionally separates **membership/basic metadata** from **landscape enrichment**:
|
||||
|
||||
1. `README.md` is the canonical curated resource list.
|
||||
2. `scripts/sync_resources_from_readme.py` derives `data/resources.yml` from README headings and bullets.
|
||||
3. `data/enrichment.yml` stores richer metadata keyed by stable resource `id`.
|
||||
4. `data/vocabulary.yml` defines canonical terms for controlled enrichment dimensions.
|
||||
5. `scripts/build_resources.py` merges base records and enrichment, validates them, and writes `data/resources.json`, `data/resources.csv`, and `docs/data/resources.json`.
|
||||
|
||||
This separation prevents hand-curated AI4Bio metadata from being erased by README synchronization.
|
||||
|
||||
## Core identity fields
|
||||
|
||||
These fields are required and may not be overridden by `data/enrichment.yml`:
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|---|---|---|
|
||||
| `id` | string | Stable lowercase `snake_case` identifier |
|
||||
| `name` | string | Official display name |
|
||||
| `type` | enum | `api`, `benchmark`, `database`, `model`, `resource`, or `toolkit` |
|
||||
| `url` | URL | Canonical landing page |
|
||||
| `description` | string | Short factual description |
|
||||
|
||||
## Landscape dimensions
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|---|---|---|
|
||||
| `tasks` | string[] | Biological or ML tasks performed |
|
||||
| `modalities` | string[] | Input/output data modalities |
|
||||
| `organism` | string[] | Covered organisms or species groups |
|
||||
| `entities` | string[] | Biological entities: gene, protein, compound, cell, disease, etc. |
|
||||
| `methods` | string[] | Method families: transformer, GNN, diffusion, optimal transport, etc. |
|
||||
| `tags` | string[] | Broad domain and curation labels |
|
||||
| `organizations` | string[] | Organizations maintaining or primarily responsible for the resource |
|
||||
|
||||
These dimensions are deliberately orthogonal. Do not encode a task as a modality or a biological entity as a resource type.
|
||||
|
||||
## Controlled vocabulary
|
||||
|
||||
New values added through `data/enrichment.yml` for `entities`, `methods`, `modalities`, and `tasks` must use canonical terms from `data/vocabulary.yml`.
|
||||
|
||||
Canonical terms use lowercase kebab-case, for example:
|
||||
|
||||
```yaml
|
||||
entities: [cell, gene]
|
||||
methods: [transformer, self-supervised-learning]
|
||||
modalities: [single-cell-rna-seq, transcriptomics]
|
||||
tasks: [foundation-model-pretraining, cell-type-annotation]
|
||||
```
|
||||
|
||||
This rule is intentionally applied only to enrichment metadata. Existing README-derived values remain valid for backward compatibility and can be migrated separately without blocking routine resource updates.
|
||||
|
||||
When a required concept is missing, add a reusable canonical term to `data/vocabulary.yml` instead of inventing a one-off spelling in an enrichment record. `tags`, `organism`, and `organizations` remain free-form because their vocabularies are broader or context dependent.
|
||||
|
||||
## Provenance and lifecycle fields
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|---|---|---|
|
||||
| `year` | integer | Initial public release or primary publication year |
|
||||
| `github` | URL | Source repository when available |
|
||||
| `documentation` | URL | Documentation landing page |
|
||||
| `paper` | URL | Primary publication or preprint |
|
||||
| `license` | string | SPDX identifier preferred |
|
||||
| `api` | boolean | Programmatic API availability |
|
||||
| `access` | enum | `open`, `registration`, `restricted`, `commercial`, `unknown` |
|
||||
| `maintenance_status` | enum | `active`, `maintenance`, `archived`, `unknown` |
|
||||
| `updated` | date | Last-known upstream update date |
|
||||
| `last_checked` | date | Date this repository verified the metadata |
|
||||
| `metadata_sources` | URL[] | Sources supporting enriched metadata |
|
||||
|
||||
`last_checked` is a curation timestamp, not an upstream release date. `updated` should only be populated when an upstream update date is known.
|
||||
|
||||
## Enrichment rules
|
||||
|
||||
`data/enrichment.yml` is a mapping keyed by resource id:
|
||||
|
||||
```yaml
|
||||
resources:
|
||||
example_resource:
|
||||
entities: [gene, disease]
|
||||
methods: [transformer]
|
||||
organizations: [Example Lab]
|
||||
year: 2025
|
||||
github: https://github.com/example/project
|
||||
documentation: https://example.org/docs
|
||||
maintenance_status: active
|
||||
access: open
|
||||
last_checked: 2026-08-08
|
||||
metadata_sources:
|
||||
- https://example.org/about
|
||||
```
|
||||
|
||||
Enrichment cannot override `id`, `name`, `type`, `url`, or `description`. A referenced id must already exist in `data/resources.yml`.
|
||||
|
||||
## Validation contract
|
||||
|
||||
`python scripts/validate_resources.py` checks:
|
||||
|
||||
- required fields and field types;
|
||||
- stable id format and id uniqueness;
|
||||
- allowed enum values;
|
||||
- HTTP(S) URL shape;
|
||||
- ISO `YYYY-MM-DD` dates;
|
||||
- list uniqueness and non-empty values;
|
||||
- enrichment references and forbidden identity overrides;
|
||||
- controlled enrichment terms against `data/vocabulary.yml`;
|
||||
- vocabulary uniqueness and lowercase kebab-case normalization;
|
||||
- unknown field names.
|
||||
|
||||
The machine-readable resource counterpart is `docs/data/resource.schema.json` (JSON Schema 2020-12). Controlled vocabulary enforcement is performed at the enrichment layer because legacy README-derived values intentionally remain backward compatible.
|
||||
|
||||
## Curation guidance
|
||||
|
||||
Prefer verified metadata over exhaustive metadata. Unknown fields should be omitted rather than guessed. For facts likely to change, include `last_checked` and at least one `metadata_sources` URL. Dynamic popularity metrics such as GitHub stars should remain generated telemetry rather than canonical curated fields.
|
||||
+53
@@ -0,0 +1,53 @@
|
||||
---
|
||||
title: "Resource.Schema"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/7a064bf0/docs/data/resource.schema.json
|
||||
upstream_sha: 7a064bf0
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://inoue0426.github.io/awesome-computational-biology/data/resource.schema.json",
|
||||
"title": "AI4Bio Resource",
|
||||
"type": "object",
|
||||
"required": ["id", "name", "type", "url", "description"],
|
||||
"additionalProperties": false,
|
||||
"properties": {
|
||||
"id": {"type": "string", "pattern": "^[a-z0-9]+(?:_[a-z0-9]+)*$"},
|
||||
"name": {"type": "string", "minLength": 1},
|
||||
"type": {"enum": ["api", "benchmark", "database", "model", "resource", "toolkit"]},
|
||||
"url": {"type": "string", "format": "uri", "pattern": "^https?://"},
|
||||
"description": {"type": "string", "minLength": 1},
|
||||
"tags": {"$ref": "#/$defs/stringArray"},
|
||||
"tasks": {"$ref": "#/$defs/stringArray"},
|
||||
"modalities": {"$ref": "#/$defs/stringArray"},
|
||||
"organism": {"$ref": "#/$defs/stringArray"},
|
||||
"entities": {"$ref": "#/$defs/stringArray"},
|
||||
"methods": {"$ref": "#/$defs/stringArray"},
|
||||
"organizations": {"$ref": "#/$defs/stringArray"},
|
||||
"metadata_sources": {"type": "array", "items": {"type": "string", "format": "uri", "pattern": "^https?://"}, "uniqueItems": true},
|
||||
"license": {"type": "string"},
|
||||
"api": {"type": "boolean"},
|
||||
"paper": {"type": "string", "format": "uri", "pattern": "^https?://"},
|
||||
"github": {"type": "string", "format": "uri", "pattern": "^https://github\\.com/"},
|
||||
"documentation": {"type": "string", "format": "uri", "pattern": "^https?://"},
|
||||
"year": {"type": "integer", "minimum": 1900, "maximum": 2100},
|
||||
"maintenance_status": {"enum": ["active", "maintenance", "archived", "unknown"]},
|
||||
"access": {"enum": ["open", "registration", "restricted", "commercial", "unknown"]},
|
||||
"updated": {"type": "string", "format": "date"},
|
||||
"last_checked": {"type": "string", "format": "date"}
|
||||
},
|
||||
"$defs": {
|
||||
"stringArray": {
|
||||
"type": "array",
|
||||
"items": {"type": "string", "minLength": 1},
|
||||
"uniqueItems": true
|
||||
}
|
||||
}
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,15 @@
|
||||
---
|
||||
title: "Requirements"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/scripts/requirements.txt
|
||||
upstream_sha: 12d87583
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
PyYAML>=6.0
|
||||
matplotlib>=3.7
|
||||
@@ -0,0 +1,352 @@
|
||||
---
|
||||
title: "Awesome AI Agents for Scientific Discovery [](https://awesome.re)"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery/blob/bb3be5bc/README.md
|
||||
upstream_sha: bb3be5bc
|
||||
imported_at: 2026-08-08
|
||||
prompt_class: prompt
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Awesome AI Agents for Scientific Discovery [](https://awesome.re)
|
||||
<div align="center">
|
||||
<img src="agents4science.webp" alt="AI Agents for Scientific Discovery" width="600px">
|
||||
</div>
|
||||
|
||||
A curated list of papers about AI agents for scientific discovery and research automation.
|
||||
|
||||
|
||||
Maintained by [Jieli Zhou](mailto:[email protected])
|
||||
|
||||
If you use this paper list for your research, please cite it using:
|
||||
```bibtex
|
||||
@misc{zhou2024awesome,
|
||||
title={Awesome AI Agents for Scientific Discovery},
|
||||
author={Zhou, Jieli},
|
||||
year={2024},
|
||||
publisher={GitHub},
|
||||
journal={GitHub repository},
|
||||
howpublished={\url{https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery}}
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
## Introduction
|
||||
|
||||
The convergence of large language models (LLMs) and autonomous agents has ushered in a new era in scientific discovery, fundamentally transforming how research is conducted across disciplines. This emerging paradigm, articulated in Kitano's seminal "Nobel Turing Challenge" (2021), envisions AI systems capable of making scientific discoveries worthy of Nobel Prize recognition. Recent advances in LLM-based agents have brought us closer to this vision, enabling increasingly sophisticated automation of scientific workflows and decision-making processes.
|
||||
|
||||
### Evolution and Current Landscape
|
||||
|
||||
The field has evolved rapidly since early visions of AI-driven scientific discovery. While traditional AI systems focused on narrow tasks, modern LLM-based agents demonstrate remarkable capabilities in complex scientific reasoning, experimental design, and hypothesis generation. The breakthrough capabilities of models like GPT-4 have catalyzed this transition, enabling agents to engage in sophisticated scientific discourse, interpret complex data, and even design novel experiments.
|
||||
|
||||
### Key Research Directions
|
||||
|
||||
Several major research themes have emerged in this space:
|
||||
|
||||
1. **Multi-Agent Architectures**: Research has increasingly focused on collaborative multi-agent systems, where specialized agents work together to tackle complex scientific problems.
|
||||
|
||||
2. **Domain-Specific Applications**: The healthcare sector has seen particularly rapid adoption, with agents being developed for clinical decision support, medical diagnosis, and healthcare administration.
|
||||
|
||||
3. **Scientific Process Automation**: Agents are being developed to automate various aspects of the research pipeline, from literature review and hypothesis generation to experimental design and data analysis.
|
||||
|
||||
### Impact and Future Directions
|
||||
|
||||
The emergence of AI agents in scientific discovery represents more than just technological advancement; it signals a fundamental shift in how science is conducted. These systems promise to:
|
||||
- Accelerate the pace of scientific discovery
|
||||
- Enable exploration of previously intractable research questions
|
||||
- Democratize access to scientific expertise
|
||||
- Foster more efficient use of research resources
|
||||
|
||||
# Awesome LLM Agents for Scientific Discovery [](https://awesome.re)
|
||||
<div align="center">
|
||||
<img src="agents4science.webp" alt="AI Agents for Scientific Discovery" width="600px">
|
||||
</div>
|
||||
|
||||
A curated list of papers about AI agents for scientific discovery and research automation.
|
||||
|
||||
|
||||
Maintained by [Jieli Zhou](mailto:[email protected])
|
||||
|
||||
If you use this paper list for your research, please cite it using:
|
||||
```bibtex
|
||||
@misc{zhou2024awesome,
|
||||
title={Awesome AI Agents for Scientific Discovery},
|
||||
author={Zhou, Jieli},
|
||||
year={2024},
|
||||
publisher={GitHub},
|
||||
journal={GitHub repository},
|
||||
howpublished={\url{https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery}}
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
## Introduction
|
||||
|
||||
The convergence of large language models (LLMs) and autonomous agents has ushered in a new era in scientific discovery, fundamentally transforming how research is conducted across disciplines. This emerging paradigm, articulated in Kitano's seminal "Nobel Turing Challenge" (2021), envisions AI systems capable of making scientific discoveries worthy of Nobel Prize recognition. Recent advances in LLM-based agents have brought us closer to this vision, enabling increasingly sophisticated automation of scientific workflows and decision-making processes.
|
||||
|
||||
### Evolution and Current Landscape
|
||||
|
||||
The field has evolved rapidly since early visions of AI-driven scientific discovery. While traditional AI systems focused on narrow tasks, modern LLM-based agents demonstrate remarkable capabilities in complex scientific reasoning, experimental design, and hypothesis generation. The breakthrough capabilities of models like GPT-4 have catalyzed this transition, enabling agents to engage in sophisticated scientific discourse, interpret complex data, and even design novel experiments.
|
||||
|
||||
### Key Research Directions
|
||||
|
||||
Several major research themes have emerged in this space:
|
||||
|
||||
1. **Multi-Agent Architectures**: Research has increasingly focused on collaborative multi-agent systems, where specialized agents work together to tackle complex scientific problems.
|
||||
|
||||
2. **Domain-Specific Applications**: The healthcare sector has seen particularly rapid adoption, with agents being developed for clinical decision support, medical diagnosis, and healthcare administration.
|
||||
|
||||
3. **Scientific Process Automation**: Agents are being developed to automate various aspects of the research pipeline, from literature review and hypothesis generation to experimental design and data analysis.
|
||||
|
||||
### Impact and Future Directions
|
||||
|
||||
The emergence of AI agents in scientific discovery represents more than just technological advancement; it signals a fundamental shift in how science is conducted. These systems promise to:
|
||||
- Accelerate the pace of scientific discovery
|
||||
- Enable exploration of previously intractable research questions
|
||||
- Democratize access to scientific expertise
|
||||
- Foster more efficient use of research resources
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Foundations & Vision](#foundations--vision)
|
||||
2. [Core Technologies](#core-technologies)
|
||||
3. [Scientific Process Automation](#scientific-process-automation)
|
||||
4. [Domain Applications](#domain-applications)
|
||||
5. [Infrastructure & Tools](#infrastructure--tools)
|
||||
8. [AI Agent Frameworks & Tools](#ai-agent-frameworks--tools)
|
||||
6. [Evaluation & Benchmarking](#evaluation--benchmarking)
|
||||
7. [Surveys & Reviews](#surveys--reviews)
|
||||
|
||||
## Foundations & Vision
|
||||
|
||||
### Vision Papers
|
||||
- **[Nobel Turing Challenge: Creating the Engine for Scientific Discovery](https://www.nature.com/articles/s41592-021-01091-w)**
|
||||
*Hiroaki Kitano.* NPJ Systems Biology and Applications 2021
|
||||
|
||||
- **[Artificial Intelligence to Win the Nobel Prize and Beyond: Creating the Engine for Scientific Discovery](https://www.aaai.org/ojs/index.php/aimagazine/article/view/2624)**
|
||||
*Hiroaki Kitano.* AI Magazine 2016
|
||||
|
||||
- **[The AI Scientist: Towards Fully Automated Open-ended Scientific Discovery](https://arxiv.org/abs/2408.06292)**
|
||||
*Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, David Ha.* arXiv 2024
|
||||
|
||||
- **[Emergent autonomous scientific research capabilities of large language models](https://arxiv.org/abs/2304.05332)**
|
||||
*Daniil A Boiko, Robert MacKnight, Gabe Gomes.* arXiv 2023
|
||||
|
||||
- **[What is missing in autonomous discovery: open challenges for the community](https://pubs.rsc.org/en/content/articlelanding/2023/dd/d3dd00089c)**
|
||||
*Phillip M Maffettone, Pascal Friederich, Sterling G Baird, et al.* Digital Discovery 2023
|
||||
|
||||
- **[The future of fundamental science led by generative closed-loop artificial intelligence](https://arxiv.org/abs/2307.07522)**
|
||||
*Hector Zenil, Jesper Tegnér, Felipe S Abrahão, Alexander Lavin, et al.* arXiv 2023
|
||||
|
||||
## Core Technologies
|
||||
|
||||
### Multi-Agent Systems & Architectures
|
||||
- **[CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society](https://proceedings.neurips.cc/paper_files/paper/2023/hash/9a86e0c5-e09e-4ad7-96d6-b2ed61855e37-Abstract-Conference.html)**
|
||||
*Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, Bernard Ghanem.* NeurIPS 2023
|
||||
|
||||
- **[Dynamic LLM-Agent Network: An LLM-Agent Collaboration Framework with Agent Team Optimization](https://arxiv.org/abs/2310.02170)**
|
||||
*Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, Diyi Yang.* arXiv 2023
|
||||
|
||||
- **[AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Framework](https://arxiv.org/abs/2308.08155)**
|
||||
*Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, et al.* arXiv 2023
|
||||
|
||||
### Reasoning & Knowledge Systems
|
||||
- **[Graph of Thoughts: Solving Elaborate Problems with Large Language Models](https://ojs.aaai.org/index.php/AAAI/article/view/28877)**
|
||||
*Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, et al.* AAAI 2024
|
||||
|
||||
- **[KnowAgent: Knowledge-augmented Planning for LLM-based Agents](https://arxiv.org/abs/2403.03101)**
|
||||
*Yuqi Zhu, Shuofei Qiao, Yixin Ou, Shumin Deng, et al.* arXiv 2024
|
||||
|
||||
- **[Improving Factuality and Reasoning in Language Models through Multiagent Debate](https://arxiv.org/abs/2305.14325)**
|
||||
*Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, Igor Mordatch.* arXiv 2023
|
||||
|
||||
## Scientific Process Automation
|
||||
|
||||
### Research Planning & Literature Review
|
||||
- **[ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models](https://arxiv.org/abs/2404.07738)**
|
||||
*Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, Sung Ju Hwang.* arXiv 2024
|
||||
|
||||
- **[SciMon: Scientific Inspiration Machines Optimized for Novelty](https://arxiv.org/abs/2305.14259)**
|
||||
*Qingyun Wang, Doug Downey, Heng Ji, Tom Hope.* arXiv 2023
|
||||
|
||||
- **[AutoSurvey: Large Language Models Can Automatically Write Surveys](https://arxiv.org/abs/2406.10252)**
|
||||
*Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, et al.* arXiv 2024
|
||||
|
||||
### Experimental Design & Workflow
|
||||
- **[DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents](https://arxiv.org/abs/2406.06769)**
|
||||
*Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, et al.* arXiv 2024
|
||||
|
||||
- **[Genesis: Towards the Automation of Systems Biology Research](https://arxiv.org/abs/2408.10689)**
|
||||
*Ievgeniia A Tiukova, Daniel Brunnsåker, Erik Y Bjurström, Alexander H Gower, et al.* arXiv 2024
|
||||
|
||||
- **[AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing](https://arxiv.org/abs/2602.17607)** ([code](https://github.com/Daviddjddu/Autonumerics))
|
||||
*Jianda Du, Youran Sun, Haizhao Yang.* arXiv 2026
|
||||
|
||||
## Domain Applications
|
||||
|
||||
### Healthcare & Medicine
|
||||
|
||||
#### Clinical Decision Support & Diagnosis
|
||||
- **[MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making](https://arxiv.org/abs/2411.00248)**
|
||||
*Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, et al.* NeurIPS 2024
|
||||
|
||||
- **[Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis](https://arxiv.org/abs/2401.16107)**
|
||||
*Haochun Wang, Sendong Zhao, Zewen Qiang, Nuwa Xi, et al.* arXiv 2024
|
||||
|
||||
- **[MedAide: Towards an Omni Medical Aide via Specialized LLM-based Multi-Agent Collaboration](https://arxiv.org/abs/2410.12532)**
|
||||
*Jinjie Wei, Dingkang Yang, Yanshu Li, Qingyao Xu, et al.* arXiv 2024
|
||||
|
||||
- **[Large Language Models as Agents in the Clinic](https://arxiv.org/abs/2309.10895)**
|
||||
*Nikita Mehandru, Brenda Y. Miao, Eduardo Rodriguez Almaraz, et al.* NPJ Digital Medicine 2024
|
||||
|
||||
- **[MAGDA: Multi-Agent Guideline-Driven Diagnostic Assistance](https://link.springer.com/chapter/10.1007/978-3-031-49673-3_15)**
|
||||
*David Bani-Harouni, Nassir Navab, Matthias Keicher.* FMGMAI 2024
|
||||
|
||||
#### Healthcare Systems & Management
|
||||
- **[ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent Collaboration](https://arxiv.org/abs/2410.02551)**
|
||||
*Zixiang Wang, Yinghao Zhu, Huiya Zhao, Xiaochen Zheng, et al.* arXiv 2024
|
||||
|
||||
- **[Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents](https://arxiv.org/abs/2405.02957)**
|
||||
*Junkai Li, Siyu Wang, Meng Zhang, Weitao Li, et al.* arXiv 2024
|
||||
|
||||
- **[ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World](https://arxiv.org/abs/2406.13890)**
|
||||
*Weixiang Yan, Haitian Liu, Tengxiao Wu, Qian Chen, et al.* arXiv 2024
|
||||
|
||||
- **[AIPatient: Simulating Patients with EHRs and LLM Powered Agentic Workflow](https://arxiv.org/abs/2409.18924)**
|
||||
*Huizi Yu, Jiayan Zhou, Lingyao Li, Shan Chen, et al.* arXiv 2024
|
||||
|
||||
#### Medical Education & Training
|
||||
- **[Medco: Medical education copilots based on a multi-agent framework](https://arxiv.org/abs/2408.12496)**
|
||||
*Hao Wei, Jianing Qiu, Haibao Yu, Wu Yuan.* arXiv 2024
|
||||
|
||||
- **[AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://arxiv.org/abs/2405.07960)**
|
||||
*Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, et al.* arXiv 2024
|
||||
|
||||
#### Medical Imaging & Pathology
|
||||
- **[CXR-Agent: Vision-language models for chest X-ray interpretation with uncertainty aware radiology reporting](https://arxiv.org/abs/2407.08811)**
|
||||
*Naman Sharma.* arXiv 2024
|
||||
|
||||
- **[PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent Collaboration](https://arxiv.org/abs/2407.00203)**
|
||||
*Yuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu, et al.* arXiv 2024
|
||||
|
||||
#### Medical Research
|
||||
- **[OpenLens AI: Fully Autonomous Research Agent for Health Infomatics](https://arxiv.org/abs/2509.14778)**
|
||||
*Yuxiao Cheng, Jinli Suo* arXiv 2025, [GitHub Repo](https://github.com/jarrycyx/openlens-ai)
|
||||
|
||||
### Biology & Life Sciences
|
||||
|
||||
#### Genomics & Molecular Biology
|
||||
- **[BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments](https://arxiv.org/abs/2405.17631)**
|
||||
*Yusuf Roohani, Andrew Lee, Qian Huang, Jian Vora, et al.* arXiv 2024
|
||||
|
||||
- **[GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases](https://arxiv.org/abs/2405.16205)**
|
||||
*Zhizheng Wang, Qiao Jin, Chih-Hsuan Wei, Shubo Tian, et al.* arXiv 2024
|
||||
|
||||
- **[Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation](https://arxiv.org/abs/2407.08940)**
|
||||
*Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, et al.* arXiv 2024
|
||||
|
||||
#### Bioinformatics Tools & Platforms
|
||||
- **[BIA: BioInformatics Agent - Unleashing the Power of Large Language Models to Reshape Bioinformatics Workflow](https://www.biorxiv.org/content/10.1101/2024.05.22.595240v1)**
|
||||
*Qi Xin, Quyu Kong, Hongyi Ji, Yue Shen, et al.* bioRxiv 2024
|
||||
|
||||
- **[CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis](https://www.biorxiv.org/content/10.1101/2024.05.13.593861v1)**
|
||||
*Yihang Xiao, Jinyi Liu, Yan Zheng, Xiaohan Xie, et al.* bioRxiv 2024
|
||||
|
||||
- **[SeqMate: A Novel Large Language Model Pipeline for Automating RNA Sequencing](https://arxiv.org/abs/2407.03381)**
|
||||
*Devam Mondal, Atharva Inamdar.* arXiv 2024
|
||||
|
||||
- **[ChatSpatial: Schema-Enforced Agentic Orchestration for Reproducible and Cross-Platform Spatial Transcriptomics](https://doi.org/10.64898/2026.02.26.708361)**
|
||||
*Chen Yang, Xianyang Zhang, Jun Chen.* bioRxiv 2026
|
||||
An MCP server that enables spatial transcriptomics analysis via natural language, integrating 60+ methods across Python and R into a single conversational workflow. [Code](https://github.com/cafferychen777/ChatSpatial)
|
||||
|
||||
### Chemistry & Materials Science
|
||||
|
||||
#### Drug Discovery & Development
|
||||
- **[DrugAgent: Explainable Drug Repurposing Agent with Large Language Model-based Reasoning](https://arxiv.org/abs/2408.13378)**
|
||||
*Yoshitaka Inoue, Tianci Song, Tianfan Fu.* arXiv 2024
|
||||
|
||||
- **[Malade: Orchestration of LLM-powered agents with retrieval augmented generation for pharmacovigilance](https://arxiv.org/abs/2408.01869)**
|
||||
*Jihye Choi, Nils Palumbo, Prasad Chalasani, Matthew M Engelhard, et al.* arXiv 2024
|
||||
|
||||
#### Molecular Modeling & Computation
|
||||
- **[ChatMol Copilot: An Agent for Molecular Modeling and Computation Powered by LLMs](https://aclanthology.org/2024.lm-1.6/)**
|
||||
*Jinyuan Sun, Auston Li, Yifan Deng, Jiabo Li.* L+M Workshop 2024
|
||||
|
||||
- **[A review of large language models and autonomous agents in chemistry](https://arxiv.org/abs/2407.01603)**
|
||||
*Mayk Caldas Ramos, Christopher J Collison, Andrew D White.* arXiv 2024
|
||||
|
||||
### Earth & Environmental Sciences
|
||||
- **[An LLM Agent for Automatic Geospatial Data Analysis](https://arxiv.org/abs/2410.18792)**
|
||||
*Yuxing Chen, Weijie Wang, Sylvain Lobry, Camille Kurtz.* arXiv 2024
|
||||
|
||||
## Evaluation & Benchmarking
|
||||
|
||||
### General Benchmarks
|
||||
- **[ClawBench: A Comprehensive Benchmark for Evaluating AI Web Agents](https://arxiv.org/abs/2604.08523)**
|
||||
*Reacher et al.* arXiv 2026. An open benchmark for browser agents on everyday tasks across live websites, with 153 V1 and 130 V2 tasks and reproducible execution traces ([code](https://github.com/reacher-z/ClawBench), [project](https://claw-bench.com/)).
|
||||
|
||||
- **[AgentBench: Evaluating LLMs as Agents](https://arxiv.org/abs/2308.03688)**
|
||||
*Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, et al.* arXiv 2023
|
||||
|
||||
- **[ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate](https://arxiv.org/abs/2308.07201)**
|
||||
*Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, et al.* arXiv 2023
|
||||
|
||||
- **[Benchmarking large language models as ai research agents](https://arxiv.org/abs/2311.12741)**
|
||||
*Qian Huang, Jian Vora, Percy Liang, Jure Leskovec.* NeurIPS 2023 Workshop
|
||||
|
||||
### Domain-Specific Benchmarks
|
||||
- **[BioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science](https://arxiv.org/abs/2407.00466)**
|
||||
*Xinna Lin, Siqi Ma, Junjie Shan, Xiaojing Zhang, et al.* arXiv 2024
|
||||
|
||||
- **[GenoTEX: A Benchmark for Evaluating LLM-Based Exploration of Gene Expression Data](https://arxiv.org/abs/2406.15341)**
|
||||
*Haoyang Liu, Haohan Wang.* arXiv 2024
|
||||
|
||||
- **[IdeaBench: Benchmarking Large Language Models for Research Idea Generation](https://arxiv.org/abs/2411.02429)**
|
||||
*Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, et al.* arXiv 2024
|
||||
|
||||
## Surveys & Reviews
|
||||
|
||||
### Comprehensive Surveys
|
||||
- **[Scientific discovery in the age of artificial intelligence](https://www.nature.com/articles/s41586-023-06221-2)**
|
||||
*Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, et al.* Nature 2023
|
||||
|
||||
- **[The rise and potential of large language model based agents: A survey](https://arxiv.org/abs/2309.07864)**
|
||||
*Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, et al.* arXiv 2023
|
||||
|
||||
- **[Large language model based multi-agents: A survey of progress and challenges](https://arxiv.org/abs/2402.01680)**
|
||||
*Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, et al.* arXiv 2024
|
||||
|
||||
- **[A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges](https://link.springer.com/article/10.1007/s44223-024-00009-0)**
|
||||
*Xinyi Li, Sai Wang, Siqi Zeng, Yu Wu, Yi Yang.* Vicinagearth 2024
|
||||
|
||||
### Domain-Specific Reviews
|
||||
- **[AI for Biomedicine in the Era of Large Language Models](https://arxiv.org/abs/2403.15673)**
|
||||
*Zhenyu Bi, Sajib Acharjee Dip, Daniel Hajialigol, Sindhura Kommu, et al.* arXiv 2024
|
||||
|
||||
- **[A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions](https://arxiv.org/abs/2406.03712)**
|
||||
*Lei Liu, Xiaoyan Yang, Junchi Lei, Xiaoyang Liu, et al.* arXiv 2024
|
||||
|
||||
- **[From LLMs to LLM-based Agents for Software Engineering: A Survey](https://arxiv.org/abs/2408.02479)**
|
||||
*Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, et al.* arXiv 2024
|
||||
|
||||
## AI Agent Frameworks & Tools
|
||||
|
||||
- **[Bride](https://tools.gracestack.se/bride-live.html)** — Cognitive AI agent with Active Inference, HDC, and anomaly detection for hypothesis generation. [Rust, MIT] `2026`
|
||||
|
||||
## Contributing
|
||||
|
||||
Please feel free to send a pull request if you want to:
|
||||
- Add new papers
|
||||
- Fix errors
|
||||
- Update paper information
|
||||
|
||||
## License
|
||||
|
||||
[](https://creativecommons.org/publicdomain/zero/1.0/)
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,288 @@
|
||||
---
|
||||
title: "Awesome LLM Agents for Scientific Discovery [](https://awesome.re)"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery/blob/3e079cd8/README.md
|
||||
upstream_sha: 3e079cd8
|
||||
imported_at: 2026-06-26
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Awesome LLM Agents for Scientific Discovery [](https://awesome.re)
|
||||
<div align="center">
|
||||
<img src="agents4science.webp" alt="AI Agents for Scientific Discovery" width="600px">
|
||||
</div>
|
||||
|
||||
A curated list of papers about AI agents for scientific discovery and research automation.
|
||||
|
||||
|
||||
Maintained by [Jieli Zhou](mailto:[email protected])
|
||||
|
||||
If you use this paper list for your research, please cite it using:
|
||||
```bibtex
|
||||
@misc{zhou2024awesome,
|
||||
title={Awesome AI Agents for Scientific Discovery},
|
||||
author={Zhou, Jieli},
|
||||
year={2024},
|
||||
publisher={GitHub},
|
||||
journal={GitHub repository},
|
||||
howpublished={\url{https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery}}
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
## Introduction
|
||||
|
||||
The convergence of large language models (LLMs) and autonomous agents has ushered in a new era in scientific discovery, fundamentally transforming how research is conducted across disciplines. This emerging paradigm, articulated in Kitano's seminal "Nobel Turing Challenge" (2021), envisions AI systems capable of making scientific discoveries worthy of Nobel Prize recognition. Recent advances in LLM-based agents have brought us closer to this vision, enabling increasingly sophisticated automation of scientific workflows and decision-making processes.
|
||||
|
||||
### Evolution and Current Landscape
|
||||
|
||||
The field has evolved rapidly since early visions of AI-driven scientific discovery. While traditional AI systems focused on narrow tasks, modern LLM-based agents demonstrate remarkable capabilities in complex scientific reasoning, experimental design, and hypothesis generation. The breakthrough capabilities of models like GPT-4 have catalyzed this transition, enabling agents to engage in sophisticated scientific discourse, interpret complex data, and even design novel experiments.
|
||||
|
||||
### Key Research Directions
|
||||
|
||||
Several major research themes have emerged in this space:
|
||||
|
||||
1. **Multi-Agent Architectures**: Research has increasingly focused on collaborative multi-agent systems, where specialized agents work together to tackle complex scientific problems.
|
||||
|
||||
2. **Domain-Specific Applications**: The healthcare sector has seen particularly rapid adoption, with agents being developed for clinical decision support, medical diagnosis, and healthcare administration.
|
||||
|
||||
3. **Scientific Process Automation**: Agents are being developed to automate various aspects of the research pipeline, from literature review and hypothesis generation to experimental design and data analysis.
|
||||
|
||||
### Impact and Future Directions
|
||||
|
||||
The emergence of AI agents in scientific discovery represents more than just technological advancement; it signals a fundamental shift in how science is conducted. These systems promise to:
|
||||
- Accelerate the pace of scientific discovery
|
||||
- Enable exploration of previously intractable research questions
|
||||
- Democratize access to scientific expertise
|
||||
- Foster more efficient use of research resources
|
||||
|
||||
## Table of Contents
|
||||
1. [Foundations & Vision](#foundations--vision)
|
||||
2. [Core Technologies](#core-technologies)
|
||||
3. [Scientific Process Automation](#scientific-process-automation)
|
||||
4. [Domain Applications](#domain-applications)
|
||||
5. [Infrastructure & Tools](#infrastructure--tools)
|
||||
6. [Evaluation & Benchmarking](#evaluation--benchmarking)
|
||||
7. [Surveys & Reviews](#surveys--reviews)
|
||||
|
||||
## Foundations & Vision
|
||||
|
||||
### Vision Papers
|
||||
- **[Nobel Turing Challenge: Creating the Engine for Scientific Discovery](https://www.nature.com/articles/s41592-021-01091-w)**
|
||||
*Hiroaki Kitano.* NPJ Systems Biology and Applications 2021
|
||||
|
||||
- **[Artificial Intelligence to Win the Nobel Prize and Beyond: Creating the Engine for Scientific Discovery](https://www.aaai.org/ojs/index.php/aimagazine/article/view/2624)**
|
||||
*Hiroaki Kitano.* AI Magazine 2016
|
||||
|
||||
- **[The AI Scientist: Towards Fully Automated Open-ended Scientific Discovery](https://arxiv.org/abs/2408.06292)**
|
||||
*Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, David Ha.* arXiv 2024
|
||||
|
||||
- **[Emergent autonomous scientific research capabilities of large language models](https://arxiv.org/abs/2304.05332)**
|
||||
*Daniil A Boiko, Robert MacKnight, Gabe Gomes.* arXiv 2023
|
||||
|
||||
- **[What is missing in autonomous discovery: open challenges for the community](https://pubs.rsc.org/en/content/articlelanding/2023/dd/d3dd00089c)**
|
||||
*Phillip M Maffettone, Pascal Friederich, Sterling G Baird, et al.* Digital Discovery 2023
|
||||
|
||||
- **[The future of fundamental science led by generative closed-loop artificial intelligence](https://arxiv.org/abs/2307.07522)**
|
||||
*Hector Zenil, Jesper Tegnér, Felipe S Abrahão, Alexander Lavin, et al.* arXiv 2023
|
||||
|
||||
## Core Technologies
|
||||
|
||||
### Multi-Agent Systems & Architectures
|
||||
- **[CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society](https://proceedings.neurips.cc/paper_files/paper/2023/hash/9a86e0c5-e09e-4ad7-96d6-b2ed61855e37-Abstract-Conference.html)**
|
||||
*Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, Bernard Ghanem.* NeurIPS 2023
|
||||
|
||||
- **[Dynamic LLM-Agent Network: An LLM-Agent Collaboration Framework with Agent Team Optimization](https://arxiv.org/abs/2310.02170)**
|
||||
*Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, Diyi Yang.* arXiv 2023
|
||||
|
||||
- **[AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Framework](https://arxiv.org/abs/2308.08155)**
|
||||
*Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, et al.* arXiv 2023
|
||||
|
||||
### Reasoning & Knowledge Systems
|
||||
- **[Graph of Thoughts: Solving Elaborate Problems with Large Language Models](https://ojs.aaai.org/index.php/AAAI/article/view/28877)**
|
||||
*Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, et al.* AAAI 2024
|
||||
|
||||
- **[KnowAgent: Knowledge-augmented Planning for LLM-based Agents](https://arxiv.org/abs/2403.03101)**
|
||||
*Yuqi Zhu, Shuofei Qiao, Yixin Ou, Shumin Deng, et al.* arXiv 2024
|
||||
|
||||
- **[Improving Factuality and Reasoning in Language Models through Multiagent Debate](https://arxiv.org/abs/2305.14325)**
|
||||
*Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, Igor Mordatch.* arXiv 2023
|
||||
|
||||
## Scientific Process Automation
|
||||
|
||||
### Research Planning & Literature Review
|
||||
- **[ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models](https://arxiv.org/abs/2404.07738)**
|
||||
*Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, Sung Ju Hwang.* arXiv 2024
|
||||
|
||||
- **[SciMon: Scientific Inspiration Machines Optimized for Novelty](https://arxiv.org/abs/2305.14259)**
|
||||
*Qingyun Wang, Doug Downey, Heng Ji, Tom Hope.* arXiv 2023
|
||||
|
||||
- **[AutoSurvey: Large Language Models Can Automatically Write Surveys](https://arxiv.org/abs/2406.10252)**
|
||||
*Yidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang, et al.* arXiv 2024
|
||||
|
||||
### Experimental Design & Workflow
|
||||
- **[DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents](https://arxiv.org/abs/2406.06769)**
|
||||
*Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, et al.* arXiv 2024
|
||||
|
||||
- **[Genesis: Towards the Automation of Systems Biology Research](https://arxiv.org/abs/2408.10689)**
|
||||
*Ievgeniia A Tiukova, Daniel Brunnsåker, Erik Y Bjurström, Alexander H Gower, et al.* arXiv 2024
|
||||
|
||||
## Domain Applications
|
||||
|
||||
### Healthcare & Medicine
|
||||
|
||||
#### Clinical Decision Support & Diagnosis
|
||||
- **[MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making](https://arxiv.org/abs/2411.00248)**
|
||||
*Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, et al.* NeurIPS 2024
|
||||
|
||||
- **[Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis](https://arxiv.org/abs/2401.16107)**
|
||||
*Haochun Wang, Sendong Zhao, Zewen Qiang, Nuwa Xi, et al.* arXiv 2024
|
||||
|
||||
- **[MedAide: Towards an Omni Medical Aide via Specialized LLM-based Multi-Agent Collaboration](https://arxiv.org/abs/2410.12532)**
|
||||
*Jinjie Wei, Dingkang Yang, Yanshu Li, Qingyao Xu, et al.* arXiv 2024
|
||||
|
||||
- **[Large Language Models as Agents in the Clinic](https://arxiv.org/abs/2309.10895)**
|
||||
*Nikita Mehandru, Brenda Y. Miao, Eduardo Rodriguez Almaraz, et al.* NPJ Digital Medicine 2024
|
||||
|
||||
- **[MAGDA: Multi-Agent Guideline-Driven Diagnostic Assistance](https://link.springer.com/chapter/10.1007/978-3-031-49673-3_15)**
|
||||
*David Bani-Harouni, Nassir Navab, Matthias Keicher.* FMGMAI 2024
|
||||
|
||||
#### Healthcare Systems & Management
|
||||
- **[ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent Collaboration](https://arxiv.org/abs/2410.02551)**
|
||||
*Zixiang Wang, Yinghao Zhu, Huiya Zhao, Xiaochen Zheng, et al.* arXiv 2024
|
||||
|
||||
- **[Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents](https://arxiv.org/abs/2405.02957)**
|
||||
*Junkai Li, Siyu Wang, Meng Zhang, Weitao Li, et al.* arXiv 2024
|
||||
|
||||
- **[ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World](https://arxiv.org/abs/2406.13890)**
|
||||
*Weixiang Yan, Haitian Liu, Tengxiao Wu, Qian Chen, et al.* arXiv 2024
|
||||
|
||||
- **[AIPatient: Simulating Patients with EHRs and LLM Powered Agentic Workflow](https://arxiv.org/abs/2409.18924)**
|
||||
*Huizi Yu, Jiayan Zhou, Lingyao Li, Shan Chen, et al.* arXiv 2024
|
||||
|
||||
#### Medical Education & Training
|
||||
- **[Medco: Medical education copilots based on a multi-agent framework](https://arxiv.org/abs/2408.12496)**
|
||||
*Hao Wei, Jianing Qiu, Haibao Yu, Wu Yuan.* arXiv 2024
|
||||
|
||||
- **[AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://arxiv.org/abs/2405.07960)**
|
||||
*Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, et al.* arXiv 2024
|
||||
|
||||
#### Medical Imaging & Pathology
|
||||
- **[CXR-Agent: Vision-language models for chest X-ray interpretation with uncertainty aware radiology reporting](https://arxiv.org/abs/2407.08811)**
|
||||
*Naman Sharma.* arXiv 2024
|
||||
|
||||
- **[PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent Collaboration](https://arxiv.org/abs/2407.00203)**
|
||||
*Yuxuan Sun, Yunlong Zhang, Yixuan Si, Chenglu Zhu, et al.* arXiv 2024
|
||||
|
||||
#### Medical Research
|
||||
- **[OpenLens AI: Fully Autonomous Research Agent for Health Infomatics](https://arxiv.org/abs/2509.14778)**
|
||||
*Yuxiao Cheng, Jinli Suo* arXiv 2025, [GitHub Repo](https://github.com/jarrycyx/openlens-ai)
|
||||
|
||||
### Biology & Life Sciences
|
||||
|
||||
#### Genomics & Molecular Biology
|
||||
- **[BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments](https://arxiv.org/abs/2405.17631)**
|
||||
*Yusuf Roohani, Andrew Lee, Qian Huang, Jian Vora, et al.* arXiv 2024
|
||||
|
||||
- **[GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases](https://arxiv.org/abs/2405.16205)**
|
||||
*Zhizheng Wang, Qiao Jin, Chih-Hsuan Wei, Shubo Tian, et al.* arXiv 2024
|
||||
|
||||
- **[Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation](https://arxiv.org/abs/2407.08940)**
|
||||
*Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, et al.* arXiv 2024
|
||||
|
||||
#### Bioinformatics Tools & Platforms
|
||||
- **[BIA: BioInformatics Agent - Unleashing the Power of Large Language Models to Reshape Bioinformatics Workflow](https://www.biorxiv.org/content/10.1101/2024.05.22.595240v1)**
|
||||
*Qi Xin, Quyu Kong, Hongyi Ji, Yue Shen, et al.* bioRxiv 2024
|
||||
|
||||
- **[CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis](https://www.biorxiv.org/content/10.1101/2024.05.13.593861v1)**
|
||||
*Yihang Xiao, Jinyi Liu, Yan Zheng, Xiaohan Xie, et al.* bioRxiv 2024
|
||||
|
||||
- **[SeqMate: A Novel Large Language Model Pipeline for Automating RNA Sequencing](https://arxiv.org/abs/2407.03381)**
|
||||
*Devam Mondal, Atharva Inamdar.* arXiv 2024
|
||||
|
||||
### Chemistry & Materials Science
|
||||
|
||||
#### Drug Discovery & Development
|
||||
- **[DrugAgent: Explainable Drug Repurposing Agent with Large Language Model-based Reasoning](https://arxiv.org/abs/2408.13378)**
|
||||
*Yoshitaka Inoue, Tianci Song, Tianfan Fu.* arXiv 2024
|
||||
|
||||
- **[Malade: Orchestration of LLM-powered agents with retrieval augmented generation for pharmacovigilance](https://arxiv.org/abs/2408.01869)**
|
||||
*Jihye Choi, Nils Palumbo, Prasad Chalasani, Matthew M Engelhard, et al.* arXiv 2024
|
||||
|
||||
#### Molecular Modeling & Computation
|
||||
- **[ChatMol Copilot: An Agent for Molecular Modeling and Computation Powered by LLMs](https://aclanthology.org/2024.lm-1.6/)**
|
||||
*Jinyuan Sun, Auston Li, Yifan Deng, Jiabo Li.* L+M Workshop 2024
|
||||
|
||||
- **[A review of large language models and autonomous agents in chemistry](https://arxiv.org/abs/2407.01603)**
|
||||
*Mayk Caldas Ramos, Christopher J Collison, Andrew D White.* arXiv 2024
|
||||
|
||||
### Earth & Environmental Sciences
|
||||
- **[An LLM Agent for Automatic Geospatial Data Analysis](https://arxiv.org/abs/2410.18792)**
|
||||
*Yuxing Chen, Weijie Wang, Sylvain Lobry, Camille Kurtz.* arXiv 2024
|
||||
|
||||
## Evaluation & Benchmarking
|
||||
|
||||
### General Benchmarks
|
||||
- **[AgentBench: Evaluating LLMs as Agents](https://arxiv.org/abs/2308.03688)**
|
||||
*Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, et al.* arXiv 2023
|
||||
|
||||
- **[ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate](https://arxiv.org/abs/2308.07201)**
|
||||
*Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, et al.* arXiv 2023
|
||||
|
||||
- **[Benchmarking large language models as ai research agents](https://arxiv.org/abs/2311.12741)**
|
||||
*Qian Huang, Jian Vora, Percy Liang, Jure Leskovec.* NeurIPS 2023 Workshop
|
||||
|
||||
### Domain-Specific Benchmarks
|
||||
- **[BioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science](https://arxiv.org/abs/2407.00466)**
|
||||
*Xinna Lin, Siqi Ma, Junjie Shan, Xiaojing Zhang, et al.* arXiv 2024
|
||||
|
||||
- **[GenoTEX: A Benchmark for Evaluating LLM-Based Exploration of Gene Expression Data](https://arxiv.org/abs/2406.15341)**
|
||||
*Haoyang Liu, Haohan Wang.* arXiv 2024
|
||||
|
||||
- **[IdeaBench: Benchmarking Large Language Models for Research Idea Generation](https://arxiv.org/abs/2411.02429)**
|
||||
*Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, et al.* arXiv 2024
|
||||
|
||||
## Surveys & Reviews
|
||||
|
||||
### Comprehensive Surveys
|
||||
- **[Scientific discovery in the age of artificial intelligence](https://www.nature.com/articles/s41586-023-06221-2)**
|
||||
*Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, et al.* Nature 2023
|
||||
|
||||
- **[The rise and potential of large language model based agents: A survey](https://arxiv.org/abs/2309.07864)**
|
||||
*Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, et al.* arXiv 2023
|
||||
|
||||
- **[Large language model based multi-agents: A survey of progress and challenges](https://arxiv.org/abs/2402.01680)**
|
||||
*Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, et al.* arXiv 2024
|
||||
|
||||
- **[A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges](https://link.springer.com/article/10.1007/s44223-024-00009-0)**
|
||||
*Xinyi Li, Sai Wang, Siqi Zeng, Yu Wu, Yi Yang.* Vicinagearth 2024
|
||||
|
||||
### Domain-Specific Reviews
|
||||
- **[AI for Biomedicine in the Era of Large Language Models](https://arxiv.org/abs/2403.15673)**
|
||||
*Zhenyu Bi, Sajib Acharjee Dip, Daniel Hajialigol, Sindhura Kommu, et al.* arXiv 2024
|
||||
|
||||
- **[A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions](https://arxiv.org/abs/2406.03712)**
|
||||
*Lei Liu, Xiaoyan Yang, Junchi Lei, Xiaoyang Liu, et al.* arXiv 2024
|
||||
|
||||
- **[From LLMs to LLM-based Agents for Software Engineering: A Survey](https://arxiv.org/abs/2408.02479)**
|
||||
*Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, et al.* arXiv 2024
|
||||
|
||||
## Contributing
|
||||
|
||||
Please feel free to send a pull request if you want to:
|
||||
- Add new papers
|
||||
- Fix errors
|
||||
- Update paper information
|
||||
|
||||
## License
|
||||
|
||||
[](https://creativecommons.org/publicdomain/zero/1.0/)
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user