Compare commits

...
Author SHA1 Message Date
promptadmin 0c2cf61fbc [upstream-sync] README.md from ai-boost/awesome-ai-for-science@22846384 [catalogue] 2026-08-08 21:50:31 +00:00
promptadmin 4478db9429 Merge pull request '[Upstream sync] inoue0426/awesome-computational-biology (github) — 0 added, 4 modified' (#53) from upstream-sync/awesome-computational-biology-20260717-478be8-wkjf into main
Reviewed-on: #53
2026-08-08 15:42:14 +00:00
promptadmin ea596262a3 Merge pull request '[Upstream sync] ai-boost/awesome-ai-for-science (github) — 0 added, 1 modified' (#63) from upstream-sync/awesome-ai-for-science-20260808-34bc16-dqxe into main
Reviewed-on: #63
2026-08-08 15:34:08 +00:00
promptadmin 6dd9ab61fc Merge pull request '[Upstream sync] GoekeLab/awesome-genomic-skills (github) — 0 added, 1 modified' (#64) from upstream-sync/awesome-genomic-skills-20260808-8cf42e-kipo into main
Reviewed-on: #64
2026-08-08 15:33:28 +00:00
promptadmin 3d3e964173 Merge pull request '[Upstream sync] zhoujieli/Awesome-LLM-Agents-Scientific-Discovery (github) — 0 added, 1 modified' (#65) from upstream-sync/awesome-llm-agents-scientific-discovery-20260808-bb3be5-vxhe into main
Reviewed-on: #65
2026-08-08 15:32:08 +00:00
promptadmin 2a32b20ab6 [upstream-sync] README.md from GoekeLab/awesome-genomic-skills@8cf42e1d [catalogue] 2026-08-08 15:30:17 +00:00
promptadmin 259566ec9a [upstream-sync] README.md from ai-boost/awesome-ai-for-science@34bc16f8 [catalogue] 2026-08-08 15:29:51 +00:00
promptadmin 1601315615 [upstream-sync] docs/data/resources.json from inoue0426/awesome-computational-biology@478be843 [catalogue] 2026-07-17 21:14:17 +00:00
promptadmin 34f5e0887c [upstream-sync] data/resources.yml from inoue0426/awesome-computational-biology@478be843 [catalogue] 2026-07-17 21:14:07 +00:00
promptadmin 86a74a726b [upstream-sync] data/resources.json from inoue0426/awesome-computational-biology@478be843 [catalogue] 2026-07-17 21:13:57 +00:00
promptadmin 783d416ecb [upstream-sync] README.md from inoue0426/awesome-computational-biology@478be843 [catalogue] 2026-07-17 21:13:48 +00:00
6 changed files with 379 additions and 19 deletions
@@ -2,9 +2,9 @@
title: "Awesome Genomic Skills [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)"
task: ""
lineage_type: import
upstream_source: https://github.com/GoekeLab/awesome-genomic-skills/blob/f88d9494/README.md
upstream_sha: f88d9494
imported_at: 2026-06-26
upstream_source: https://github.com/GoekeLab/awesome-genomic-skills/blob/8cf42e1d/README.md
upstream_sha: 8cf42e1d
imported_at: 2026-08-08
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -50,7 +50,7 @@ Skill libraries and tool collections specifically targeting genomics, bioinforma
- **Description:** AWS sample bundle for the Kiro IDE: 24 MCP servers wrapping 100+ databases/tools across genomics, proteomics, structural biology, and clinical/pharma (NCBI, Ensembl, ClinVar, gnomAD, UniProt, STRING, PDB, AlphaFold, ChEMBL, Open Targets, etc.), plus 10 domain skills and 16 workflows, with cross-database search and AWS HealthOmics pipeline execution. MIT-0; the MCP servers use standard MCP and are portable, though the skills and workflows are built for Kiro. [Blog post](https://aws.amazon.com/blogs/publicsector/accelerating-life-sciences-research-with-kiro-a-unified-ai-interface-to-100-open-source-databases/).
- **Developers:** AWS (AWS Samples).
- [ClawBio](https://github.com/ClawBio/ClawBio)
- **Description:** The first bioinformatics-native AI agent skill library; provides reproducible, local-first skills for genomics tasks (variant calling, RNA-seq, population genetics) that work with Claude Code, Copilot, Codex, and other agents.
- **Description:** Bioinformatics-native AI agent skill library; 95 reproducible, local-first skills for genomics tasks (variant calling, RNA-seq, population genetics) that work with Claude Code, Copilot, Codex, and other agents. Since v0.6.1 the whole library is also callable as an MCP server (`uvx --from 'clawbio[mcp]' clawbio mcp`) from Cursor, Claude Desktop, VS Code, or Zed (see the MCP section below). Skills are actively [benchmarked](https://clawbio.ai/benchmarks.html).
- **Developers:** Independent open-source project built on [OpenClaw](https://openclaw.ai).
- [SciAgent-skills](https://github.com/jaechang-hits/SciAgent-Skills)
- **Description:** 197 open-source skills for Claude Code, Cursor, Codex, and Windsurf, covering genomics-bioinformatics, proteomics-protein engineering, structural biology, drug discovery, systems biology, biostatistics, and scientific writing; achieves 92% accuracy on BixBench-Verified-50 (+26.7 pts over Claude Code baseline). The hosted OmicsHorizon web platform runs these skills in-browser.
@@ -83,6 +83,9 @@ Model Context Protocol (MCP) servers that give AI agents direct access to bioinf
- [gget-mcp](https://github.com/longevity-genie/gget-mcp) - MCP server wrapping the Pachter Lab [gget](https://github.com/pachterlab/gget) bioinformatics toolkit. Exposes 13 tools covering gene search and metadata (Ensembl), sequence retrieval, BLAST/BLAT/MUSCLE alignment, expression data (ARCHS4), functional enrichment (Enrichr), protein structure (PDB, AlphaFold), cancer mutations (COSMIC), and single-cell queries (CellxGene).
- [Seqera MCP](https://docs.seqera.io/platform-cloud/seqera-mcp/overview) - Hosted MCP server from Seqera Labs (the developers of Nextflow) exposing the Seqera Platform (workflow launch/management), Wave (container provisioning), nf-core modules, and SRA/ENA/GEO retrieval.
- [knowledgebase-mcp](https://github.com/biocontext-ai/knowledgebase-mcp) - BioContextAI Knowledgebase MCP server, included in the BioContextAI registry (below), one of the most comprehensive single MCP packages (wraps STRINGDb, Open Targets, Reactome, UniProt, HPA, KEGG, AlphaFold, Ensembl, ClinicalTrials.gov, bioRxiv, etc.)
- [ClawBio MCP](https://github.com/ClawBio/ClawBio) - MCP mode of the [ClawBio](#bioinformatics-and-genomics-agent-skills) skill library (0.6.1): exposes all 95 genomics skills as callable tools over local stdio, letting different agents (Cursor, Zed, etc.) run analyses (variant calling, RNA-seq, population genetics).
- [roda-mcp](https://github.com/awslabs/mcp/tree/main/src/roda-mcp-server) - AWS Labs MCP server for discovering and exploring datasets in the Registry of Open Data on AWS (RODA), covering 1,100+ public datasets across life sciences, climate, geospatial, and satellite imagery. Search/filter by keyword, organization, or license; inspect dataset details; and browse or sample public S3 bucket contents directly (no AWS account required) without downloading full files. Not genomics-specific, but useful for locating and previewing life-sciences datasets (e.g. SG-NEx) hosted on AWS.
- [plant-genomics-mcp](https://github.com/musharna/plant-genomics-mcp) - Plant-genomics MCP server exposing 50 tools across 23 backends, keyed on TAIR-style loci with organism resolution across 12 crop and model species. Covers plant-specific resources that general bioinformatics servers do not (Ensembl Plants, Phytozome, Gramene, Planteome PO/TO, PlantCyc/PMN, AraGWAS, 1001 Genomes, ThaleMine, BAR, JASPAR) alongside the usual UniProt/KEGG/STRING/AlphaFold/PDBe/InterPro/Europe PMC, plus cross-source synthesis tools that compose several backends into a single gene report.
Existing registries and lists of MCP servers:
- [BioContextAI Registry](https://github.com/biocontext-ai/registry/) - Community-curated catalogue of biomedical MCP servers, with submission criteria requiring biomedical focus, free academic access, OSI-approved open-source licenses, and MCP specification compliance. Ships a [cookiecutter template](https://github.com/biocontext-ai/mcp-server-cookiecutter) for new servers and follows Schema.org ontologies for metadata. [A community hub for agentic biomedical systems](https://www.nature.com/articles/s41587-025-02900-9)
@@ -2,9 +2,9 @@
title: "Readme"
task: ""
lineage_type: import
upstream_source: https://github.com/ai-boost/awesome-ai-for-science/blob/be91bc08/README.md
upstream_sha: be91bc08
imported_at: 2026-07-25
upstream_source: https://github.com/ai-boost/awesome-ai-for-science/blob/22846384/README.md
upstream_sha: 22846384
imported_at: 2026-08-08
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -216,7 +216,14 @@ validated: false
- [OpenScience (Synthetic Sciences)](https://github.com/synthetic-sciences/openscience) - Open-source AI workbench for scientific research that automates the full research loop — literature review, hypothesis generation, code writing, experiment execution, database querying, and report writing — with 290+ skills, specialized research agents, and a browser-based workspace (1453+ stars, Apache 2.0, 2026)
- [Claude Scholar](https://github.com/Galaxy-Dawn/claude-scholar) - Semi-automated research assistant for academic research and software development, supporting Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication (Galaxy-Dawn, 4.5K+ stars, MIT License, 2026)
- [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok) - Free, open-source desktop AI research assistant that runs locally and turns natural-language requests into real data analysis, literature search, figure generation, and manuscript review; ships with 149 scientific skills, 326 workflow templates, and 229 databases across genomics, proteomics, drug discovery, and materials science, plus a living lab notebook, 60+ scientific file previews, and LaTeX editing (K-Dense-AI, 908+ stars, MIT License, 2026)
- [Science Superpowers (K-Dense-AI)](https://github.com/K-Dense-AI/science-superpowers) - Composable computational-science methodology skills for AI research agents emphasizing pre-registration, reproducible workspaces, and red-team review to guard against p-hacking and HARKing; zero third-party dependencies and runs with any agent harness plus a POSIX shell (281+ stars, MIT License, 2026)
- [Wisp Science](https://github.com/xuzhougeng/wisp-science) - Open-source, local-first desktop AI research workbench for scientific computing with Python/R, MCP bioinformatics tools, SSH/WSL/GPU runtimes, and OpenAI/Anthropic models (857+ stars, 2026)
- [Academic Research Skills (ARS)](https://github.com/Imbad0202/academic-research-skills) - Comprehensive Claude Code skill suite covering the full academic pipeline from deep research and paper writing to multi-perspective peer review, revision, and finalization; features multi-agent teams, PRISMA systematic review, style calibration, claim-level citation audits, integrity gates, and human-in-the-loop safeguards (38K+ stars, CC BY-NC 4.0, 2026)
- [Qinyan Academic Skills](https://github.com/LeonChaoX/qinyan-academic-skills) - Curated, multilingual library of 182 installable AI agent skills for end-to-end academic research spanning literature discovery, scientific writing, grant development, bioinformatics, drug discovery, clinical research, machine learning, and data analysis (779+ stars, MIT License, 2026)
- [SkillOpt (Microsoft, 2026)](https://github.com/microsoft/SkillOpt) - Text-space optimizer that treats agent skill documents as trainable parameters for frozen LLMs, using scored rollouts and held-out validation gates to iteratively improve reusable natural-language skills; includes SkillOpt-Sleep for nightly self-evolution and improves accuracy across Claude Code, Codex, Copilot, and direct-chat harnesses, making it a meta-tool for evolving scientific agent skill workflows (15.5K+ stars, MIT License, PyPI)
- [Open Science (AIPOCH)](https://github.com/aipoch/open-science) - Open-source, local-first, model-agnostic AI research workbench for reproducible scientific discovery; runs Python/R notebooks, searches the web, calls scientific data connectors, and produces inspectable reports, tables, and figures in a self-hosted desktop workspace (1.5K+ stars, Apache 2.0, 2026)
- [OmicsClaw](https://github.com/TianGzlab/OmicsClaw) - Local-first, conversational AI research partner for multi-omics analysis with CLI, desktop app, and 95+ reproducible skills; keeps raw data local while routing natural-language requests to Python/R/CLI tools with persistent memory, autonomous analysis paths, and multi-method consensus workflows (TianGzlab, 155+ stars, Apache 2.0, 2026)
- [MedgeClaw](https://github.com/xjtulyc/MedgeClaw) - Open-source AI research assistant for biomedicine — chat to run RNA-seq, drug discovery, clinical analysis, and more; built on OpenClaw and Claude Code with 140 K-Dense scientific skills, real-time dashboard, and RStudio/JupyterLab integration (xjtulyc, 669+ stars, 2026)
### Literature Management Plugins
- [llm-for-zotero](https://github.com/yilewang/llm-for-zotero) - Research agent system deeply integrated with Zotero supporting Agent Mode, skills, multi-model backends (OpenAI-compatible, Claude Code, WebChat, Codex), and MinerU PDF parsing for literature Q&A, summarization, figure inspection, and source comparison (1.3K+ stars, 2026)
@@ -240,6 +247,7 @@ validated: false
- [GraphGen](https://github.com/open-sciencelab/GraphGen) - Knowledge graph-guided synthetic data generation for LLM fine-tuning, achieving strong performance on scientific QA (GPQA-Diamond) and math reasoning (AIME)
- [KoPA](https://github.com/zjukg/KoPA) - Structure-aware prefix adaptation for integrating LLMs with knowledge graphs (ACM MM 2024)
- [Scholarly KGQA](https://arxiv.org/abs/2311.09841) - LLM-powered question answering over scholarly knowledge graphs (ArXiv paper)
- [SciAtlas](https://github.com/zjunlp/SciAtlas) - Large-scale knowledge graph and pip-installable client for literature-grounded automated scientific research, connecting papers, authors, institutions, venues, keywords, citations, and a four-level research taxonomy across medicine, social sciences, engineering, computer science, materials science, and more (ZJU NLP, arXiv 2026, 136+ stars, MIT License)
### Knowledge Graph Resources
- [Awesome-LLM-KG](https://github.com/RManLuo/Awesome-LLM-KG) - Comprehensive collection of papers on unifying LLMs and knowledge graphs
@@ -254,10 +262,13 @@ validated: false
- [SkyDiscover](https://github.com/skydiscover-ai/skydiscover) - Modular framework for AI-driven scientific and algorithmic discovery, providing a unified interface for implementing, running, and fairly comparing discovery algorithms across 200+ optimization tasks; introduces AdaEvolve and EvoX adaptive/evolutionary algorithms and natively supports OpenEvolve, GEPA, and Harbor-format benchmarks (skydiscover-ai, 568+ stars, Apache 2.0, 2026)
- [EvoMaster (SJTU SAI, arXiv 2026)](https://github.com/sjtu-sai-agents/EvoMaster) - Foundational auto-research agent framework for agentic science at scale, providing modular agent construction, run-level self-evolution, and multiple SciMaster domain agents (ML-Master, X-Master, Browse-Master); outperforms general-purpose agents across authoritative benchmarks including the OpenAI Frontier Science Benchmark (206+ stars, Apache 2.0, 2026)
- [Virtual Lab (Stanford Zou Group, Nature 2025)](https://github.com/zou-group/virtual-lab) - AI-human collaborative research platform where a human researcher works with a team of LLM agents via team and individual meetings to perform scientific research; demonstrated by designing new SARS-CoV-2 nanobodies with wet-lab validation
- [AI Co-Scientist (Google DeepMind, Nature Medicine 2026)](https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/) - Multi-agent AI research partner that generates, reviews, ranks, and evolves research hypotheses alongside scientists, with experimental validation in biomedicine and other domains (2026)
- [Hyra (Tencent Hunyuan, 2026)](https://hy.tencent.com/research/hyra) - Hunyuan Research Agent for autonomous open-ended discovery across AI4Science, mathematics, and engineering, releasing reproducible solution artifacts for autocorrelation constants, Erdős problems, PARP1 docking, qubit routing, and record-breaking packing problems ([Hyra-results](https://github.com/Tencent-Hunyuan/Hyra-results), 112+ stars, Apache 2.0)
- [The AI Scientist (SakanaAI)](https://github.com/SakanaAI/AI-Scientist) - First fully autonomous open-ended scientific discovery system with official implementation: hypothesis→experiment→writing→review simulation (13.8K+ stars, 2024)
- [The AI Scientist v2 (SakanaAI)](https://github.com/SakanaAI/AI-Scientist-v2) - Official implementation of the second-generation fully autonomous scientific discovery system, extending the original with agentic tree search and reduced template dependency to achieve workshop-level accepted papers (6.7K+ stars, 2025)
- [The AI Scientist v1 (2024)](https://arxiv.org/abs/2408.06292) - First fully autonomous research system: hypothesis→experiment→writing→review simulation
- [The AI Scientist v2 (2025)](https://arxiv.org/abs/2504.08066) - Enhanced with Agentic Tree Search, reduced template dependency, first workshop-level accepted paper
- [FAROS (OpenNSWM-Lab)](https://github.com/OpenNSWM-Lab/FAROS) - Foundation AutoResearch Operating System: blueprint-driven runtime for orchestrating AI research workflows from idea generation and experiments to paper writing and peer review (OpenNSWM-Lab, 2.4K+ stars, 2026)
- [DeepScientist](https://github.com/ResearAI/DeepScientist) - First system progressively surpassing human SOTA on frontier AI tasks (183.7%, 1.9%, 7.9% improvements), month-long autonomous discovery with 20,000+ GPU hours
- [ASI-Arch (GAIR-NLP, arXiv 2025)](https://github.com/GAIR-NLP/ASI-Arch) - Autonomous multi-agent research loop for model architecture discovery that ran 1,773 experiments over 20,000 GPU hours and produced 106 state-of-the-art linear-attention architectures, surpassing human-designed baselines including Mamba2 and DeltaNet (1.1K+ stars, Apache 2.0)
- [Kosmos](https://github.com/jimmc414/Kosmos) - Extended autonomy AI scientist with 200 parallel agent rollouts, 42K lines of code execution, 1.5K papers analyzed per run, achieving 79.4% accuracy and 7 scientific discoveries (Edison Scientific)
@@ -299,6 +310,7 @@ validated: false
- [ScienceAgentBench (ICLR 2025)](https://github.com/OSU-NLP-Group/ScienceAgentBench) - 102 executable tasks from 44 peer-reviewed papers across 4 disciplines with containerized evaluation
- [AIRS-Bench (Meta, 2026)](https://github.com/facebookresearch/airs-bench) - Benchmark quantifying end-to-end autonomous AI research abilities of LLM agents across 20 tasks from SOTA machine learning papers spanning NLP, code, math, biochemical modelling, and time series forecasting, with normalized score metrics against human SOTA and HuggingFace dataset
- [PaperBench (OpenAI, 2025)](https://github.com/openai/preparedness/tree/main/project/paperbench) - Benchmark evaluating AI agents' ability to replicate 20 ICML 2024 Spotlight/Oral papers from scratch, with 8,316 gradable tasks and author-co-developed rubrics
- [PaperGuru (AutoTrustAI, 2026)](https://github.com/AutoTrustAI/PaperGuru-Benchmark) - Lifecycle-Aware Memory (LAM) primitive and benchmark for long-horizon research agents, achieving 65.95% mean reproduction on PaperBench and 94.66% on SurveyBench through Capital Chunk Memory (CCM) with versioned content, structural multi-hop relevance, and provenance-grounded composition; 10 peer-reviewed acceptances at FSE/ICML/TOSEM/AEI/ICoGB (1.3K+ stars)
- [MLE-Bench (OpenAI, 2024)](https://github.com/openai/mle-bench) - Benchmark evaluating AI agents on 75 curated Kaggle-style ML engineering competitions with reproducible Docker-based grading harness, human baselines, and end-to-end task lifecycle, used as a primary benchmark for autonomous ML research agents (e.g., InternAgent #1 at 36.44%)
- [ScienceBoard (ICLR 2026)](https://github.com/OS-Copilot/ScienceBoard) - Evaluating multimodal autonomous agents in realistic scientific workflows across real scientific software environments (KAlgebra, Celestia, Grass GIS, Lean 4, etc.) with VM-based evaluation infrastructure and agent trajectories
- [BuildArena](https://github.com/AI4Science-WestlakeU/BuildArena) - First physics-aligned interactive benchmark for LLM agents in engineering construction, designing rockets/cars/bridges in physics simulator with 3D spatial geometry library
@@ -315,11 +327,13 @@ validated: false
### Domain-Specific Research Agents
- [Aletheia](https://arxiv.org/abs/2602.10177) - Google DeepMind's autonomous mathematics research agent powered by Gemini Deep Think, autonomously solving 4 open problems from 700 Erdős conjectures and generating complete research papers without human intervention (February 2026)
- [AlphaProof Nexus (Google DeepMind, arXiv 2026)](https://github.com/google-deepmind/alphaproof-nexus-results) - LLM-driven formal proof search system that pairs large language models with Lean verification to solve open mathematics problems; autonomously resolved 9 of 353 Erdős problems and 44 of 492 OEIS conjectures, with proofs and natural-language prose released for combinatorics, optimization, graph theory, algebraic geometry, and quantum optics collaborations (282+ stars, Apache 2.0)
- [AlphaGeometry](https://github.com/google-deepmind/alphageometry) - DeepMind's Olympiad-level geometry theorem prover combining neural language model with symbolic deduction engine, AlphaGeometry2 solves 84% of IMO geometry problems (42/50) at gold-medalist level (Nature 2024)
- [Goedel-Prover-V2](https://github.com/Goedel-LM/Goedel-Prover-V2) - Strongest open-source automated theorem prover in Lean 4, 8B model matches DeepSeek-Prover-V2-671B at 84.6% MiniF2F, 32B model achieves 90.4% with self-correction, using scaffolded data synthesis and verifier-guided proof refinement (Princeton, 2025)
- [DeepSeek-Prover-V2](https://github.com/deepseek-ai/DeepSeek-Prover-V2) - DeepSeek's open-source large language model for formal theorem proving in Lean 4, integrating informal and formal mathematical reasoning through recursive subgoal decomposition and reinforcement learning powered by DeepSeek-V3, with open weights and ProverBench evaluation (2025)
- [LeanDojo](https://github.com/lean-dojo/LeanDojo) - Open-source toolkit and benchmark for learning-based theorem proving in Lean, providing programmatic Lean interaction, a 98K+ theorem dataset extracted from 217 Lean projects, and ReProver—the first retrieval-augmented LLM-based theorem prover for Lean—with reproducible training pipelines underpinning much subsequent Lean prover research (Caltech & NVIDIA, NeurIPS 2023 Outstanding Paper, Datasets & Benchmarks)
- [Lean Copilot](https://github.com/lean-dojo/LeanCopilot) - LLMs as copilots for theorem proving in Lean 4, exposing native tactics (`suggest_tactics`, `search_proof`, `select_premises`) that embed language model inference and premise retrieval directly inside the Lean proof environment, supporting local CTranslate2/CUDA inference as well as remote model APIs for interactive and automated proof search (Caltech & NVIDIA, NeurIPS 2024, 1.2K+ stars)
- [MathCode](https://github.com/math-ai-org/mathcode) - Terminal AI coding assistant with a built-in math formalization engine that converts plain-language math problems into Lean 4 theorems and attempts formal proofs; bundles a local Lean toolchain and WebUI for interactive mathematical reasoning (math-ai-org, 582+ stars, 2026)
- [Get Physics Done (PSI)](https://github.com/psi-oss/get-physics-done) - First open-source agentic AI physicist turning research questions into structured workflows with rigorous verification and multi-step analytical work for long-horizon physics projects; integrates with Claude Code, Codex, Gemini CLI, and OpenCode (804+ stars, Apache 2.0, 2026)
- [Foam-Agent (NeurIPS 2025)](https://github.com/csml-rpi/Foam-Agent) - End-to-end composable multi-agent framework for automating OpenFOAM-based CFD simulations from natural language prompts, managing meshing, case setup, execution, error correction, and post-processing; achieves 100% success rate on 110 FoamBench tasks with Claude Opus 4.6 through Architect-Input Writer-Runner-Reviewer agent collaboration with RAG-enhanced generation and MCP tool integration (RPI CSML, 242+ stars, MIT License)
- [Zephyrus (ICLR 2026)](https://github.com/Rose-STL-Lab/Zephyrus) - First agentic framework for weather science, pairing an LLM with ZephyrusWorld (a code-execution environment exposing WeatherBench 2 data, geolocation, forecasting, simulation, and climatology tools) and ZephyrusBench (2,230 Q&A pairs across 49 weather-science tasks); outperforms text-only baselines by up to 44.2 percentage points (UC San Diego Rose-STL-Lab, 99+ stars, MIT License, 2026)
@@ -794,6 +808,7 @@ validated: false
- [AgentSociety](https://github.com/tsinghua-fib-lab/AgentSociety) - Modern LLM-native agent simulation platform for social science research and experimental design, providing a flexible framework for creating and managing intelligent agents in simulated environments (Tsinghua FIB Lab, 984+ stars, 2025)
- [Auto-Empirical-Research-Skills](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills) - Curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines, enabling reproducible social science research with AI agents (Stanford REAP & CoPaper.AI, 3K+ stars, 2026)
- [EDSL](https://github.com/expectedparrot/edsl) - Design, conduct and analyze results of AI-powered surveys and experiments. Simulate social science and market research with large numbers of AI agents and LLMs (460+ stars, 2024)
- [GABRIEL (OpenAI, 2026)](https://github.com/openai/GABRIEL) - Generalized Attribute Based Ratings Information Extraction Library; official OpenAI toolkit that turns messy qualitative corpora into analysis-ready datasets for social scientists and data scientists, measuring quantitative attributes in text, images, or audio using the GPT API. See the [official blog post](https://openai.com/index/scaling-social-science-research/) and [NBER working paper](http://www.nber.org/papers/w34834) (413+ stars, Apache 2.0)
---
@@ -880,6 +895,7 @@ validated: false
### Specialized Frameworks
- [MDAnalysis](https://github.com/MDAnalysis/mdanalysis) - Molecular dynamics analysis
- [Nano World Model](https://github.com/simchowitzlabpublic/nano-world-model) - Minimalist, batteries-included repository for training video world models with diffusion-forcing, supporting long-horizon rollouts, 3D point-cloud generation, and model-predictive control with pretrained checkpoints (Simchowitz Lab, 700+ stars, MIT License, 2026)
- [e3nn](https://github.com/e3nn/e3nn) - Euclidean neural networks for arbitrary point transformations enabling E(3)-equivariant deep learning, foundational library for building geometry-aware neural networks in molecular dynamics, materials science, and physics
- [MDtrajNet](https://arxiv.org/abs/2505.16301) - Neural network foundation model that directly generates MD trajectories bypassing force calculations, accelerating simulations by up to 100× with equivariant Transformer architecture (2025)
- [ASE](https://wiki.fysik.dtu.dk/ase/) - Atomic Simulation Environment for materials modeling
@@ -959,6 +975,7 @@ This project builds upon and complements several excellent resources:
- [awesome-ai4s](https://github.com/hyperai/awesome-ai4s) - 200+ AI for Science papers with Chinese interpretations
- [Awesome AI Scientist Papers](https://github.com/openags/Awesome-AI-Scientist-Papers) - Autonomous AI scientist research
- [Awesome Scientific Machine Learning](https://github.com/MartinuzziFrancesco/awesome-scientific-machine-learning) - Physics-informed ML and SciML
- [Awesome Scientific Skills](https://github.com/InternScience/Awesome-Scientific-Skills) - Curated collection of agent skills for scientific research (InternScience, 493+ stars, 2026)
- [Awesome Agents for Science](https://github.com/OSU-NLP-Group/awesome-agents4science) - LLM agents across scientific domains
- [Awesome LLM Agents Scientific Discovery](https://github.com/zhoujieli/Awesome-LLM-Agents-Scientific-Discovery) - Biomedical AI agents
- [Awesome Foundation Models for Weather and Climate](https://github.com/shengchaochen82/Awesome-Foundation-Models-for-Weather-and-Climate) - Comprehensive survey of foundation models for weather and climate data understanding
@@ -2,9 +2,9 @@
title: "Awesome Computational Biology [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/README.md
upstream_sha: 12d87583
imported_at: 2026-06-26
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/478be843/README.md
upstream_sha: 478be843
imported_at: 2026-07-17
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -265,6 +265,8 @@ Browse and search the resources via the [GitHub Pages UI](https://inoue0426.gith
- [Tabula Sapiens](https://tabula-sapiens-portal.ds.czbiohub.org/) — Comprehensive human single-cell atlas of ~500K cells from 24 organs and tissues across multiple donors.
- [TAPE (Tasks Assessing Protein Embeddings)](https://github.com/songlab-cal/tape) — Benchmark suite of five biologically meaningful semi-supervised learning tasks for evaluating protein representations.
- [The Cancer Genome Atlas (TCGA)](https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga) — Comprehensive multi-omics (genomics, transcriptomics, proteomics, methylation) dataset for 33 cancer types across ~11,000 patients.
- [TCGA virtual spatial transcriptomics atlas](https://huggingface.co/datasets/ratschlab/TCGA_virtual_spatial_transcriptomics_atlas) — DeepSpot-M predicted transcriptome-wide ST for TCGA H&E (FF + FFPE; 28,664 slides / 32 cancer types; gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).
- [HEST Xenium virtual spatial transcriptomics](https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics) — DeepSpot-M predicted transcriptome-wide ST for 59 HEST-1k 10x Xenium samples (~13.3M cells) (gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).
- [Therapeutics Data Commons (TDC)](https://tdcommons.ai/) — Unified benchmark suite covering ADMET, drug-target interaction, drug response, and more.
- [Tox21](https://tripod.nih.gov/tox21/challenge/) — 12,707 compounds tested in 12 nuclear receptor and stress-response pathway biochemical assays for toxicity prediction.
- [UK Biobank](https://www.ukbiobank.ac.uk/) — Large-scale biomedical database of ~500K participants with genetic, imaging, and health data for population genetics and disease studies.
@@ -320,6 +322,7 @@ Browse and search the resources via the [GitHub Pages UI](https://inoue0426.gith
- [sciPENN](https://github.com/jlakkis/sciPENN) — RNN-based method for simultaneous protein expression prediction, uncertainty estimation, and cell-type label transfer from CITE-seq and scRNA-seq data.
- [MOGONET](https://github.com/txWang/MOGONET) — Multi-omics graph convolutional network framework for patient classification and biomarker identification.
- [AutoZyme](https://github.com/ElliotXie/autozyme) — Autonomous agentic framework that speeds up bioinformatics software (e.g. Scanpy, Seurat) on CPUs while preserving the original results.
- [SeqBench](https://seqbench.com/) — Web-based molecular biology sequence workbench for primer design, cloning simulation (Gibson, Golden Gate, restriction digest), CRISPR guide RNA design, and sequence analysis, with a public REST API, OpenAPI 3.1 spec, and MCP server.
---
@@ -412,6 +415,10 @@ Browse and search the resources via the [GitHub Pages UI](https://inoue0426.gith
- [Phikon](https://huggingface.co/owkin/phikon) — ViT-based pathology foundation model pretrained with iBOT self-supervision on TCGA whole-slide images.
- [Nicheformer](https://github.com/theislab/nicheformer) — Foundation model for single-cell and spatial omics using a transformer architecture with positional embeddings to encode spatial cell information.
- [scGPT-spatial](https://github.com/bowang-lab/scGPT-spatial) — Extension of scGPT for spatial transcriptomics with continual pretraining and a mixture-of-experts decoder for spatial gene expression analysis.
- [DeepSpot](https://github.com/ratschlab/DeepSpot) — Deep learning model predicting spatial transcriptomics from H&E images at spot and single-cell resolution.
- [DeepSpot2Cell](https://github.com/ratschlab/DeepSpot2Cell) — Predicts virtual single-cell spatial transcriptomics from H&E using spot-level supervision (NeurIPS 2025 Imageomics).
- [DeepSpot-M](https://github.com/ratschlab/DeepSpotM) — Multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology.
- [AESTETIK](https://github.com/ratschlab/aestetik) — Autoencoder for spatial transcriptomics representation learning using topology and histology image knowledge.
##### Multi-Omics Foundation Models
@@ -2,9 +2,9 @@
title: "Resources"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/data/resources.json
upstream_sha: 12d87583
imported_at: 2026-06-26
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/478be843/data/resources.json
upstream_sha: 478be843
imported_at: 2026-07-17
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -292,6 +292,20 @@ validated: false
"organism": [],
"api": false
},
{
"id": "hest_xenium_virtual_spatial_transcriptomics",
"name": "HEST Xenium virtual spatial transcriptomics",
"type": "benchmark",
"url": "https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics",
"description": "DeepSpot-M predicted transcriptome-wide ST for 59 HEST-1k 10x Xenium samples (~13.3M cells) (gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).",
"tags": [
"benchmarks-and-datasets"
],
"tasks": [],
"modalities": [],
"organism": [],
"api": false
},
{
"id": "jump_cell_painting_datasets",
"name": "JUMP Cell Painting Datasets",
@@ -530,6 +544,20 @@ validated: false
"organism": [],
"api": false
},
{
"id": "tcga_virtual_spatial_transcriptomics_atlas",
"name": "TCGA virtual spatial transcriptomics atlas",
"type": "benchmark",
"url": "https://huggingface.co/datasets/ratschlab/TCGA_virtual_spatial_transcriptomics_atlas",
"description": "DeepSpot-M predicted transcriptome-wide ST for TCGA H&E (FF + FFPE; 28,664 slides / 32 cancer types; gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).",
"tags": [
"benchmarks-and-datasets"
],
"tasks": [],
"modalities": [],
"organism": [],
"api": false
},
{
"id": "the_cancer_genome_atlas_tcga",
"name": "The Cancer Genome Atlas (TCGA)",
@@ -2095,6 +2123,27 @@ validated: false
"organism": [],
"api": false
},
{
"id": "aestetik",
"name": "AESTETIK",
"type": "model",
"url": "https://github.com/ratschlab/aestetik",
"description": "Autoencoder for spatial transcriptomics representation learning using topology and histology image knowledge.",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "ai4chem_chemllm_7b_chat",
"name": "AI4Chem/ChemLLM-7B-Chat",
@@ -2667,6 +2716,69 @@ validated: false
"organism": [],
"api": false
},
{
"id": "deepspot",
"name": "DeepSpot",
"type": "model",
"url": "https://github.com/ratschlab/DeepSpot",
"description": "Deep learning model predicting spatial transcriptomics from H&E images at spot and single-cell resolution.",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "deepspot_m",
"name": "DeepSpot-M",
"type": "model",
"url": "https://github.com/ratschlab/DeepSpotM",
"description": "Multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology.",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "deepspot2cell",
"name": "DeepSpot2Cell",
"type": "model",
"url": "https://github.com/ratschlab/DeepSpot2Cell",
"description": "Predicts virtual single-cell spatial transcriptomics from H&E using spot-level supervision (NeurIPS 2025 Imageomics).",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "dgdrp",
"name": "DGDRP",
@@ -4913,6 +5025,22 @@ validated: false
"organism": [],
"api": false
},
{
"id": "seqbench",
"name": "SeqBench",
"type": "toolkit",
"url": "https://seqbench.com/",
"description": "Web-based molecular biology sequence workbench for primer design, cloning simulation (Gibson, Golden Gate, restriction digest), CRISPR guide RNA design, and sequence analysis, with a public REST API, OpenAPI 3.1 spec, and MCP server.",
"tags": [
"preprocessing-tools"
],
"tasks": [
"Preprocessing"
],
"modalities": [],
"organism": [],
"api": false
},
{
"id": "seurat",
"name": "Seurat",
@@ -2,9 +2,9 @@
title: "Awesome Computational Biology - machine-readable resource list"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/data/resources.yml
upstream_sha: 12d87583
imported_at: 2026-06-26
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/478be843/data/resources.yml
upstream_sha: 478be843
imported_at: 2026-07-17
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -248,6 +248,17 @@ resources:
organism: []
api: false
- id: hest_xenium_virtual_spatial_transcriptomics
name: "HEST Xenium virtual spatial transcriptomics"
type: benchmark
url: https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics
description: "DeepSpot-M predicted transcriptome-wide ST for 59 HEST-1k 10x Xenium samples (~13.3M cells) (gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1)."
tags: [benchmarks-and-datasets]
tasks: []
modalities: []
organism: []
api: false
- id: jump_cell_painting_datasets
name: "JUMP Cell Painting Datasets"
type: benchmark
@@ -435,6 +446,17 @@ resources:
organism: []
api: false
- id: tcga_virtual_spatial_transcriptomics_atlas
name: "TCGA virtual spatial transcriptomics atlas"
type: benchmark
url: https://huggingface.co/datasets/ratschlab/TCGA_virtual_spatial_transcriptomics_atlas
description: "DeepSpot-M predicted transcriptome-wide ST for TCGA H&E (FF + FFPE; 28,664 slides / 32 cancer types; gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1)."
tags: [benchmarks-and-datasets]
tasks: []
modalities: []
organism: []
api: false
- id: the_cancer_genome_atlas_tcga
name: "The Cancer Genome Atlas (TCGA)"
type: benchmark
@@ -1491,6 +1513,17 @@ resources:
organism: []
api: false
- id: aestetik
name: "AESTETIK"
type: model
url: https://github.com/ratschlab/aestetik
description: "Autoencoder for spatial transcriptomics representation learning using topology and histology image knowledge."
tags: [foundation-models, single-cell-foundation-models, spatial-foundation-models]
tasks: [Foundation Model]
modalities: [Single Cell, Spatial Transcriptomics]
organism: []
api: false
- id: ai4chem_chemllm_7b_chat
name: "AI4Chem/ChemLLM-7B-Chat"
type: model
@@ -1810,6 +1843,39 @@ resources:
organism: []
api: false
- id: deepspot
name: "DeepSpot"
type: model
url: https://github.com/ratschlab/DeepSpot
description: "Deep learning model predicting spatial transcriptomics from H&E images at spot and single-cell resolution."
tags: [foundation-models, single-cell-foundation-models, spatial-foundation-models]
tasks: [Foundation Model]
modalities: [Single Cell, Spatial Transcriptomics]
organism: []
api: false
- id: deepspot_m
name: "DeepSpot-M"
type: model
url: https://github.com/ratschlab/DeepSpotM
description: "Multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology."
tags: [foundation-models, single-cell-foundation-models, spatial-foundation-models]
tasks: [Foundation Model]
modalities: [Single Cell, Spatial Transcriptomics]
organism: []
api: false
- id: deepspot2cell
name: "DeepSpot2Cell"
type: model
url: https://github.com/ratschlab/DeepSpot2Cell
description: "Predicts virtual single-cell spatial transcriptomics from H&E using spot-level supervision (NeurIPS 2025 Imageomics)."
tags: [foundation-models, single-cell-foundation-models, spatial-foundation-models]
tasks: [Foundation Model]
modalities: [Single Cell, Spatial Transcriptomics]
organism: []
api: false
- id: dgdrp
name: "DGDRP"
type: model
@@ -3097,6 +3163,17 @@ resources:
organism: []
api: false
- id: seqbench
name: "SeqBench"
type: toolkit
url: https://seqbench.com/
description: "Web-based molecular biology sequence workbench for primer design, cloning simulation (Gibson, Golden Gate, restriction digest), CRISPR guide RNA design, and sequence analysis, with a public REST API, OpenAPI 3.1 spec, and MCP server."
tags: [preprocessing-tools]
tasks: [Preprocessing]
modalities: []
organism: []
api: false
- id: seurat
name: "Seurat"
type: toolkit
@@ -2,9 +2,9 @@
title: "Resources"
task: ""
lineage_type: import
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/12d87583/docs/data/resources.json
upstream_sha: 12d87583
imported_at: 2026-06-26
upstream_source: https://github.com/inoue0426/awesome-computational-biology/blob/478be843/docs/data/resources.json
upstream_sha: 478be843
imported_at: 2026-07-17
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -292,6 +292,20 @@ validated: false
"organism": [],
"api": false
},
{
"id": "hest_xenium_virtual_spatial_transcriptomics",
"name": "HEST Xenium virtual spatial transcriptomics",
"type": "benchmark",
"url": "https://huggingface.co/datasets/ratschlab/HEST_Xenium_virtual_spatial_transcriptomics",
"description": "DeepSpot-M predicted transcriptome-wide ST for 59 HEST-1k 10x Xenium samples (~13.3M cells) (gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).",
"tags": [
"benchmarks-and-datasets"
],
"tasks": [],
"modalities": [],
"organism": [],
"api": false
},
{
"id": "jump_cell_painting_datasets",
"name": "JUMP Cell Painting Datasets",
@@ -530,6 +544,20 @@ validated: false
"organism": [],
"api": false
},
{
"id": "tcga_virtual_spatial_transcriptomics_atlas",
"name": "TCGA virtual spatial transcriptomics atlas",
"type": "benchmark",
"url": "https://huggingface.co/datasets/ratschlab/TCGA_virtual_spatial_transcriptomics_atlas",
"description": "DeepSpot-M predicted transcriptome-wide ST for TCGA H&E (FF + FFPE; 28,664 slides / 32 cancer types; gated). Paper: [DeepSpot-M](https://www.medrxiv.org/content/10.64898/2026.06.19.26356060v1).",
"tags": [
"benchmarks-and-datasets"
],
"tasks": [],
"modalities": [],
"organism": [],
"api": false
},
{
"id": "the_cancer_genome_atlas_tcga",
"name": "The Cancer Genome Atlas (TCGA)",
@@ -2095,6 +2123,27 @@ validated: false
"organism": [],
"api": false
},
{
"id": "aestetik",
"name": "AESTETIK",
"type": "model",
"url": "https://github.com/ratschlab/aestetik",
"description": "Autoencoder for spatial transcriptomics representation learning using topology and histology image knowledge.",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "ai4chem_chemllm_7b_chat",
"name": "AI4Chem/ChemLLM-7B-Chat",
@@ -2667,6 +2716,69 @@ validated: false
"organism": [],
"api": false
},
{
"id": "deepspot",
"name": "DeepSpot",
"type": "model",
"url": "https://github.com/ratschlab/DeepSpot",
"description": "Deep learning model predicting spatial transcriptomics from H&E images at spot and single-cell resolution.",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "deepspot_m",
"name": "DeepSpot-M",
"type": "model",
"url": "https://github.com/ratschlab/DeepSpotM",
"description": "Multimodal foundation model for transcriptome-wide virtual spatial transcriptomics from histology.",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "deepspot2cell",
"name": "DeepSpot2Cell",
"type": "model",
"url": "https://github.com/ratschlab/DeepSpot2Cell",
"description": "Predicts virtual single-cell spatial transcriptomics from H&E using spot-level supervision (NeurIPS 2025 Imageomics).",
"tags": [
"foundation-models",
"single-cell-foundation-models",
"spatial-foundation-models"
],
"tasks": [
"Foundation Model"
],
"modalities": [
"Single Cell",
"Spatial Transcriptomics"
],
"organism": [],
"api": false
},
{
"id": "dgdrp",
"name": "DGDRP",
@@ -4913,6 +5025,22 @@ validated: false
"organism": [],
"api": false
},
{
"id": "seqbench",
"name": "SeqBench",
"type": "toolkit",
"url": "https://seqbench.com/",
"description": "Web-based molecular biology sequence workbench for primer design, cloning simulation (Gibson, Golden Gate, restriction digest), CRISPR guide RNA design, and sequence analysis, with a public REST API, OpenAPI 3.1 spec, and MCP server.",
"tags": [
"preprocessing-tools"
],
"tasks": [
"Preprocessing"
],
"modalities": [],
"organism": [],
"api": false
},
{
"id": "seurat",
"name": "Seurat",