[Upstream sync] K-Dense-AI/scientific-agent-skills (github) — 5 added, 2 modified #48

Open
promptadmin wants to merge 7 commits from upstream-sync/scientific-agent-skills-20260818-9e8b0c-fcfk into main
7 changed files with 1183 additions and 29 deletions
@@ -2,9 +2,9 @@
title: "Scientific Agent Skills" title: "Scientific Agent Skills"
task: "" task: ""
lineage_type: import lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/991bd993/README.md upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9e8b0cb0/README.md
upstream_sha: 991bd993 upstream_sha: 9e8b0cb0
imported_at: 2026-08-08 imported_at: 2026-08-18
prompt_class: catalogue prompt_class: catalogue
upstream_changes: accepted upstream_changes: accepted
author: upstream author: upstream
@@ -14,34 +14,26 @@ validated: false
# Scientific Agent Skills # Scientific Agent Skills
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE.md) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE.md)
[![Version](https://img.shields.io/badge/Version-2.62.0-blue.svg)](pyproject.toml) [![Version](https://img.shields.io/badge/Version-2.64.0-blue.svg)](pyproject.toml)
[![Skills](https://img.shields.io/badge/Skills-159-brightgreen.svg)](#-whats-included) [![Skills](https://img.shields.io/badge/Skills-163-brightgreen.svg)](#-whats-included)
[![Databases](https://img.shields.io/badge/Databases-100%2B-orange.svg)](#-whats-included) [![Databases](https://img.shields.io/badge/Databases-100%2B-orange.svg)](#-whats-included)
[![Agent Skills](https://img.shields.io/badge/Standard-Agent_Skills-blueviolet.svg)](https://agentskills.io/) [![Agent Skills](https://img.shields.io/badge/Standard-Agent_Skills-blueviolet.svg)](https://agentskills.io/)
[![Agent Plugins](https://img.shields.io/badge/Standard-Agent_Plugins-0A7A72.svg)](https://agent-plugins.org/)
[![Security Scan](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/security-scan.yml/badge.svg)](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/security-scan.yml) [![Security Scan](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/security-scan.yml/badge.svg)](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/security-scan.yml)
[![Skill Tests](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/skill-tests.yml/badge.svg)](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/skill-tests.yml) [![Skill Tests](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/skill-tests.yml/badge.svg)](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/skill-tests.yml)
[![Works with](https://img.shields.io/badge/Works_with-Cursor_|_Claude_Code_|_Codex_|_Google_Antigravity-blue.svg)](#-getting-started) [![Works with](https://img.shields.io/badge/Works_with-Cursor_|_Claude_Code_|_Codex_|_Google_Antigravity-blue.svg)](#-getting-started)
[![X](https://img.shields.io/badge/Follow_on_X-%40k__dense__ai-000000?logo=x)](https://x.com/k_dense_ai) [![X](https://img.shields.io/badge/Follow_on_X-%40k__dense__ai-000000?logo=x)](https://x.com/k_dense_ai)
[![LinkedIn](https://img.shields.io/badge/LinkedIn-K--Dense_Inc.-0A66C2?logo=linkedin)](https://www.linkedin.com/company/k-dense-inc) [![LinkedIn](https://img.shields.io/badge/LinkedIn-K--Dense_Inc.-0A66C2?logo=linkedin)](https://www.linkedin.com/company/k-dense-inc)
[![YouTube](https://img.shields.io/badge/YouTube-K--Dense_Inc.-FF0000?logo=youtube)](https://www.youtube.com/@K-Dense-Inc) [![YouTube](https://img.shields.io/badge/YouTube-K--Dense_Inc.-FF0000?logo=youtube)](https://www.youtube.com/@K-Dense-Inc)
[![Reddit](https://img.shields.io/badge/Reddit-u%2F--k--dense---FF4500?logo=reddit&logoColor=white)](https://www.reddit.com/user/-k-dense-/)
## Star History
<a href="https://www.star-history.com/?repos=K-Dense-AI%2Fscientific-agent-skills&type=date&legend=top-left">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/chart?repos=K-Dense-AI/scientific-agent-skills&type=date&theme=dark&legend=top-left&sealed_token=rL_5GLS9f4Fbyr1_VYZLGMF-8Rr6ZlWNaYNecajc52QSQq6KL7HrzSea_tGQGy1mBMXgVvAUMSIYAc0w39si9v5Up1RIw74-UDGZg_9HvH_chiyS0Njf-5tebtPh1LJjXTG6mH5Iv2pMJNivgfPsyB-oOgbaIV3uSc7DzSeZFCTE4WOcHX4y2BR76k5g" />
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/chart?repos=K-Dense-AI/scientific-agent-skills&type=date&legend=top-left&sealed_token=rL_5GLS9f4Fbyr1_VYZLGMF-8Rr6ZlWNaYNecajc52QSQq6KL7HrzSea_tGQGy1mBMXgVvAUMSIYAc0w39si9v5Up1RIw74-UDGZg_9HvH_chiyS0Njf-5tebtPh1LJjXTG6mH5Iv2pMJNivgfPsyB-oOgbaIV3uSc7DzSeZFCTE4WOcHX4y2BR76k5g" />
<img alt="Star History Chart" src="https://api.star-history.com/chart?repos=K-Dense-AI/scientific-agent-skills&type=date&legend=top-left&sealed_token=rL_5GLS9f4Fbyr1_VYZLGMF-8Rr6ZlWNaYNecajc52QSQq6KL7HrzSea_tGQGy1mBMXgVvAUMSIYAc0w39si9v5Up1RIw74-UDGZg_9HvH_chiyS0Njf-5tebtPh1LJjXTG6mH5Iv2pMJNivgfPsyB-oOgbaIV3uSc7DzSeZFCTE4WOcHX4y2BR76k5g" />
</picture>
</a>
> **🔔 Claude Scientific Skills is now Scientific Agent Skills.** Same skills, broader compatibility — now works with any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, not just Claude. > **🔔 Claude Scientific Skills is now Scientific Agent Skills.** Same skills, broader compatibility — now works with any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, not just Claude.
> **New: [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok)** — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 159 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via [Modal](https://modal.com/) for heavy workloads. [Get started here.](https://github.com/K-Dense-AI/k-dense-byok) > **New: [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok)** — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 161 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via [Modal](https://modal.com/) for heavy workloads. [Get started here.](https://github.com/K-Dense-AI/k-dense-byok)
> **Stay up to date:** Follow K-Dense on [X](https://x.com/k_dense_ai), [LinkedIn](https://www.linkedin.com/company/k-dense-inc), and [YouTube](https://www.youtube.com/@K-Dense-Inc) for new skills, release announcements, walkthroughs, research workflow demos, and examples you can use with your own AI agent. > **Stay up to date:** Follow K-Dense on [X](https://x.com/k_dense_ai), [LinkedIn](https://www.linkedin.com/company/k-dense-inc), [YouTube](https://www.youtube.com/@K-Dense-Inc), and [Reddit](https://www.reddit.com/user/-k-dense-/) for new skills, release announcements, walkthroughs, research workflow demos, and examples you can use with your own AI agent.
A comprehensive collection of **159 ready-to-use scientific and research skills** (covering cancer genomics, individual-level 1000 Genomes queries, hosted regulatory-sequence prediction, live pathogen-variant surveillance, analytical method validation, PK/PD modelling and dose selection, full-text biomedical and regulatory literature retrieval, drug-target binding, molecular dynamics, RNA velocity, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond. A comprehensive collection of **163 ready-to-use scientific and research skills** (covering cancer genomics, individual-level 1000 Genomes queries, hosted regulatory-sequence prediction, live pathogen-variant surveillance, analytical method validation, PK/PD modelling and dose selection, full-text biomedical and regulatory literature retrieval, drug-target binding, bounded biomedical knowledge graph search, molecular dynamics, RNA velocity, microbiome foundation models, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). The repository is also a portable [Agent Plugins](https://agent-plugins.org/) package (`plugin.json` + `skills/`), so plugin-capable clients can load the whole collection as one plugin. Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
> ⭐ **Help make AI for science easier to discover:** If Scientific Agent Skills saves you time, teaches your agent a workflow, or helps your lab move faster, please [star this repository](https://github.com/K-Dense-AI/scientific-agent-skills). A star is a public signal that these open, reusable research skills are worth maintaining: it helps scientists, engineers, and open-source contributors find the project, shows which agent-skill standards are gaining real adoption, and gives us a clear reason to keep expanding the collection for the community. > ⭐ **Help make AI for science easier to discover:** If Scientific Agent Skills saves you time, teaches your agent a workflow, or helps your lab move faster, please [star this repository](https://github.com/K-Dense-AI/scientific-agent-skills). A star is a public signal that these open, reusable research skills are worth maintaining: it helps scientists, engineers, and open-source contributors find the project, shows which agent-skill standards are gaining real adoption, and gives us a clear reason to keep expanding the collection for the community.
@@ -76,9 +68,9 @@ These skills enable your AI agent to seamlessly work with specialized scientific
## 📦 What's Included ## 📦 What's Included
This repository provides **159 scientific and research skills** organized into the following categories: This repository provides **163 scientific and research skills** organized into the following categories:
- **100+ Scientific & Financial Databases** - A unified database-lookup skill provides deterministic, provenance-rich access to 78 public databases (PubChem, ChEMBL, UniProt, COSMIC, ClinicalTrials.gov, FRED, USPTO, and more), plus dedicated skills for DepMap, Imaging Data Commons, PrimeKG, U.S. Treasury Fiscal Data, Hugging Science, OneKGPd, and Genomic Intelligence. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage - **100+ Scientific & Financial Databases** - A unified database-lookup skill provides deterministic, provenance-rich access to 78 public databases (PubChem, ChEMBL, UniProt, COSMIC, ClinicalTrials.gov, FRED, USPTO, and more), plus dedicated skills for DepMap, Imaging Data Commons, PrimeKG, NCATS ARAX, U.S. Treasury Fiscal Data, Hugging Science, OneKGPd, and Genomic Intelligence. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage
- **70+ Optimized Python Package Skills** - Explicitly defined, version-aware workflows for RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, PathML, pydicom, NeuroKit2, PufferLib, QuTiP, GeoPandas, pymatgen, BioPython, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), and others. The agent can still use *any* Python package; these skills provide stronger, safer guidance for the packages listed - **70+ Optimized Python Package Skills** - Explicitly defined, version-aware workflows for RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, PathML, pydicom, NeuroKit2, PufferLib, QuTiP, GeoPandas, pymatgen, BioPython, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), and others. The agent can still use *any* Python package; these skills provide stronger, safer guidance for the packages listed
- **9 Scientific Integration Skills** - Explicitly defined skills for Benchling, DNAnexus, LatchBio, OMERO, Protocols.io, Open Notebook, Ginkgo Cloud Lab, LabArchives, and Opentrons. Again, the agent is not limited to these — any API or platform reachable from Python is fair game; these skills are the optimized, pre-documented paths - **9 Scientific Integration Skills** - Explicitly defined skills for Benchling, DNAnexus, LatchBio, OMERO, Protocols.io, Open Notebook, Ginkgo Cloud Lab, LabArchives, and Opentrons. Again, the agent is not limited to these — any API or platform reachable from Python is fair game; these skills are the optimized, pre-documented paths
- **30+ Analysis & Communication Tools** - Literature review, evidence-traceable scientific writing, confidential peer review, document processing, Paperclip (full-text papers, FDA/PMDA/EMA filings, and trial registries with line-pinned citations), Paperzilla, Exa Search, macro-free PPTX posters, slides, schematics, infographics, Mermaid diagrams, and more - **30+ Analysis & Communication Tools** - Literature review, evidence-traceable scientific writing, confidential peer review, document processing, Paperclip (full-text papers, FDA/PMDA/EMA filings, and trial registries with line-pinned citations), Paperzilla, Exa Search, macro-free PPTX posters, slides, schematics, infographics, Mermaid diagrams, and more
@@ -123,7 +115,7 @@ Each skill includes:
- **Multi-Step Workflows** - Execute complex pipelines with a single prompt - **Multi-Step Workflows** - Execute complex pipelines with a single prompt
### 🎯 **Comprehensive Coverage** ### 🎯 **Comprehensive Coverage**
- **159 Skills** - Extensive coverage across all major scientific domains - **161 Skills** - Extensive coverage across all major scientific domains
- **100+ Databases** - Unified access to 78+ databases via database-lookup, plus dedicated data access skills and multi-database packages like BioServices, BioPython, and gget - **100+ Databases** - Unified access to 78+ databases via database-lookup, plus dedicated data access skills and multi-database packages like BioServices, BioPython, and gget
- **70+ Optimized Python Package Skills** - Current, version-scoped guidance for packages including RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, pydicom, PufferLib, QuTiP, GeoPandas, pymatgen, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), scVelo, and TimesFM (the agent can use any Python package; these are the pre-documented paths) - **70+ Optimized Python Package Skills** - Current, version-scoped guidance for packages including RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, pydicom, PufferLib, QuTiP, GeoPandas, pymatgen, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), scVelo, and TimesFM (the agent can use any Python package; these are the pre-documented paths)
@@ -178,7 +170,7 @@ Pin to a specific release tag or commit SHA for reproducible installs:
```bash ```bash
# Pin to a release tag # Pin to a release tag
gh skill install K-Dense-AI/scientific-agent-skills --pin v2.62.0 gh skill install K-Dense-AI/scientific-agent-skills --pin v2.64.0
# Pin to a commit SHA # Pin to a commit SHA
gh skill install K-Dense-AI/scientific-agent-skills --pin abc123def gh skill install K-Dense-AI/scientific-agent-skills --pin abc123def
@@ -194,6 +186,27 @@ gh skill update
gh skill update --all gh skill update --all
``` ```
### Option 3: Agent Plugins (Cursor, Codex, and other plugin clients)
This repository is a valid [Agent Plugins](https://agent-plugins.org/) 1.0.0 package: root [`plugin.json`](plugin.json) plus Agent Skills under `skills/`. Clients that support the standard discover every immediate child of `skills/` that contains a `SKILL.md`.
**Cursor** — symlink or copy the repo into the local plugins directory, then reload:
```bash
mkdir -p ~/.cursor/plugins/local
ln -s "$(pwd)" ~/.cursor/plugins/local/scientific-agent-skills
```
Restart Cursor or run **Developer: Reload Window**, then confirm the plugin and its skills appear under **Customize**. See [Cursor plugins](https://cursor.com/docs/plugins).
**Codex** — install from a local checkout (confirm the current CLI flag names in Codex docs):
```bash
codex plugins install .
```
Compatible clients (Cursor, Codex, GitHub Copilot, VS Code, Kiro, and others listed at [agent-plugins.org](https://agent-plugins.org/compatible-clients)) share the same package layout; installation UX stays client-specific.
### Other Agent Skills hosts (OpenClaw, NemoClaw, Pi, Hermes, …) ### Other Agent Skills hosts (OpenClaw, NemoClaw, Pi, Hermes, …)
Agent hosts differ in install paths, discovery settings, and support for optional frontmatter fields. `npx skills add` (Option 1) commonly installs into the `~/.agents/skills/` convention, with project-scoped installs under `.agents/skills/`; confirm both paths against your host's current documentation. To install manually on a host configured to scan one of those locations: Agent hosts differ in install paths, discovery settings, and support for optional frontmatter fields. `npx skills add` (Option 1) commonly installs into the `~/.agents/skills/` convention, with project-scoped installs under `.agents/skills/`; confirm both paths against your host's current documentation. To install manually on a host configured to scan one of those locations:
@@ -209,7 +222,7 @@ For Hermes versions that support skill taps, add the repository as a tap:
hermes skills tap add K-Dense-AI/scientific-agent-skills hermes skills tap add K-Dense-AI/scientific-agent-skills
``` ```
Every `SKILL.md` has YAML frontmatter, but legacy and community skills vary in `metadata` formatting (block or flow style) and optional extension fields. Repository updates must keep `metadata.version` as a quoted numeric string and pass canonical `skills-ref validate ./skills/<skill-name>` checks. Hosts may interpret optional metadata and credential prompts differently, so verify behavior on the target host. Because 159 skills add up to a lot of standing context, consider installing a topical subset rather than the whole collection. Every `SKILL.md` has YAML frontmatter, but legacy and community skills vary in `metadata` formatting (block or flow style) and optional extension fields. Repository updates must keep `metadata.version` as a quoted numeric string and pass canonical `skills-ref validate ./skills/<skill-name>` checks. Hosts may interpret optional metadata and credential prompts differently, so verify behavior on the target host. Because 161 skills add up to a lot of standing context, consider installing a topical subset rather than the whole collection.
> **NemoClaw note:** NemoClaw runs agents inside NVIDIA OpenShell with default-deny outbound networking. Skills are discovered and loaded normally, but any skill that needs the network — package installs via `uv`, or API calls (Exa, Parallel, Benchling, NCBI, Materials Project, …) — only works once the operator pre-approves the relevant domains in the OpenShell TUI. > **NemoClaw note:** NemoClaw runs agents inside NVIDIA OpenShell with default-deny outbound networking. Skills are discovered and loaded normally, but any skill that needs the network — package installs via `uv`, or API calls (Exa, Parallel, Benchling, NCBI, Materials Project, …) — only works once the operator pre-approves the relevant domains in the OpenShell TUI.
@@ -448,13 +461,13 @@ networks, and search GEO for similar patterns.
## 📚 Available Skills ## 📚 Available Skills
This repository contains **159 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools. This repository contains **163 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
### Skill Categories ### Skill Categories
> **Note:** The Python package and integration skills listed below are *explicitly defined* skills — curated with documentation, examples, and best practices for stronger, more reliable performance. They are not a ceiling: the agent can install and use *any* Python package or call *any* API, even without a dedicated skill. The skills listed simply make common workflows faster and more dependable. > **Note:** The Python package and integration skills listed below are *explicitly defined* skills — curated with documentation, examples, and best practices for stronger, more reliable performance. They are not a ceiling: the agent can install and use *any* Python package or call *any* API, even without a dedicated skill. The skills listed simply make common workflows faster and more dependable.
#### 🧬 **Bioinformatics & Genomics** (26 skills) #### 🧬 **Bioinformatics & Genomics** (27 skills)
- RNA-seq pipelines: Bulk RNA-seq (end-to-end FASTQ -> counts -> DE -> enrichment orchestrator) - RNA-seq pipelines: Bulk RNA-seq (end-to-end FASTQ -> counts -> DE -> enrichment orchestrator)
- Sequence analysis: BioPython, pysam, scikit-bio, BioServices - Sequence analysis: BioPython, pysam, scikit-bio, BioServices
- Single-cell analysis: Scanpy, AnnData, scvi-tools, scVelo (RNA velocity), Arboreto, Cellxgene Census - Single-cell analysis: Scanpy, AnnData, scvi-tools, scVelo (RNA velocity), Arboreto, Cellxgene Census
@@ -464,6 +477,7 @@ This repository contains **159 scientific and research skills** organized across
- Differential expression: PyDESeq2 - Differential expression: PyDESeq2
- Functional enrichment: Pathway Enrichment (ORA, GSEA/preranked, ssGSEA via gseapy + g:Profiler; GO, KEGG, Reactome, WikiPathways, MSigDB) - Functional enrichment: Pathway Enrichment (ORA, GSEA/preranked, ssGSEA via gseapy + g:Profiler; GO, KEGG, Reactome, WikiPathways, MSigDB)
- Phylogenetics: ETE Toolkit, Phylogenetics (MAFFT, IQ-TREE 2, FastTree) - Phylogenetics: ETE Toolkit, Phylogenetics (MAFFT, IQ-TREE 2, FastTree)
- Microbiome foundation models: Waypoint (Outpost Bio's open Waypoint-6m/45m/170m checkpoints, the Atlas 539k-sample MGnify pretraining corpus, and the eight-task Compass benchmark — embedding, fine-tuning, benchmarking, and pretraining on taxonomic abundance profiles, with MetaPhlAn/Kraken2/QIIME 2 conversion)
#### 🧪 **Cheminformatics & Drug Discovery** (10 skills) #### 🧪 **Cheminformatics & Drug Discovery** (10 skills)
- Molecular manipulation: RDKit, Datamol, Molfeat - Molecular manipulation: RDKit, Datamol, Molfeat
@@ -512,7 +526,8 @@ This repository contains **159 scientific and research skills** organized across
- Astronomy: Astropy - Astronomy: Astropy
- Quantum computing: Cirq, PennyLane, Qiskit, QuTiP 5.3 - Quantum computing: Cirq, PennyLane, Qiskit, QuTiP 5.3
#### ⚙️ **Engineering & Simulation** (5 skills) #### ⚙️ **Engineering & Simulation** (6 skills)
- Lab hardware CAD: parametric build123d 0.11.1 models for microfluidic chips and molds, optomechanical mounts, microplate and cuvette adapters, and behavior rigs, checked against ANSI/SLAS and optical-table dimensional standards and reviewed with mandatory multi-view renders
- Numerical computing: proprietary MATLAB R2026a and distinct GNU Octave 11.3 planning/review workflows - Numerical computing: proprietary MATLAB R2026a and distinct GNU Octave 11.3 planning/review workflows
- Computational fluid dynamics: bounded FluidSim 0.9 simulations with numerical-validity and HPC checks - Computational fluid dynamics: bounded FluidSim 0.9 simulations with numerical-validity and HPC checks
- Experimental flow measurement: OpenPIV (velocity fields from PIV image pairs, interrogation-window cross-correlation, spurious-vector validation, vorticity/strain-rate/turbulence statistics) - Experimental flow measurement: OpenPIV (velocity fields from PIV image pairs, interrogation-window cross-correlation, spurious-vector validation, vorticity/strain-rate/turbulence statistics)
@@ -564,12 +579,13 @@ This repository contains **159 scientific and research skills** organized across
- Citations: Citation Management, pyzotero - Citations: Citation Management, pyzotero
- Illustration: Generate Image (AI image generation with FLUX.2 Pro and Gemini 3.1 Flash Image / Nano Banana 2) - Illustration: Generate Image (AI image generation with FLUX.2 Pro and Gemini 3.1 Flash Image / Nano Banana 2)
#### 🔬 **Scientific Databases & Data Access** (10 skills → 100+ databases total) #### 🔬 **Scientific Databases & Data Access** (11 skills → 100+ databases total)
> A unified database-lookup skill provides deterministic REST API access to 78 public databases across all domains, with retrieval contracts, pagination/count reconciliation, and endpoint provenance. Dedicated skills cover specialized data platforms. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage. > A unified database-lookup skill provides deterministic REST API access to 78 public databases across all domains, with retrieval contracts, pagination/count reconciliation, and endpoint provenance. Dedicated skills cover specialized data platforms. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage.
- Unified access: Database Lookup (78 databases spanning chemistry, genomics, clinical, pathways, patents, economics, and more — PubChem, ChEMBL, UniProt, PDB, AlphaFold, KEGG, Reactome, STRING, ClinVar, COSMIC, ClinicalTrials.gov, FDA, FRED, USPTO, SEC EDGAR, and dozens more — with auditable filters and provenance) - Unified access: Database Lookup (78 databases spanning chemistry, genomics, clinical, pathways, patents, economics, and more — PubChem, ChEMBL, UniProt, PDB, AlphaFold, KEGG, Reactome, STRING, ClinVar, COSMIC, ClinicalTrials.gov, FDA, FRED, USPTO, SEC EDGAR, and dozens more — with auditable filters and provenance)
- Cancer genomics: DepMap (cancer cell line dependencies, drug sensitivity, gene effect profiles) - Cancer genomics: DepMap (cancer cell line dependencies, drug sensitivity, gene effect profiles)
- Cancer imaging: Imaging Data Commons (NCI radiology & pathology datasets via idc-index) - Cancer imaging: Imaging Data Commons (NCI radiology & pathology datasets via idc-index)
- Knowledge graph: PrimeKG (precision medicine knowledge graph — genes, drugs, diseases, phenotypes) - Knowledge graph: PrimeKG (precision medicine knowledge graph — genes, drugs, diseases, phenotypes)
- Biomedical knowledge graph search: [NCATS ARAX](skills/ncats-arax/) (bounded, Biolink-constrained one-hop and endpoint-pinned two-hop queries over knowledge graphs with up to five explicitly selected NCATS Translator providers, with provenance preservation)
- Fiscal data: U.S. Treasury Fiscal Data (national debt, Treasury statements, auctions, exchange rates) - Fiscal data: U.S. Treasury Fiscal Data (national debt, Treasury statements, auctions, exchange rates)
- Scientific ML resource catalog: Hugging Science (curated index of datasets, models, blog posts, and interactive Spaces across 17 scientific domains — astronomy, biology, chemistry, climate, genomics, materials science, medicine, physics, scientific reasoning, and more — with usage patterns for `datasets`, `transformers`, and `gradio_client`) - Scientific ML resource catalog: Hugging Science (curated index of datasets, models, blog posts, and interactive Spaces across 17 scientific domains — astronomy, biology, chemistry, climate, genomics, materials science, medicine, physics, scientific reasoning, and more — with usage patterns for `datasets`, `transformers`, and `gradio_client`)
- Individual-level population genomics: OneKGPd (3,202-person high-coverage 1000 Genomes cohort queries) - Individual-level population genomics: OneKGPd (3,202-person high-coverage 1000 Genomes cohort queries)
@@ -620,6 +636,7 @@ Deep dives, benchmarks, and guides from the [K-Dense blog](https://www.k-dense.a
- **[Agent Skills: The Final Piece for AI-Powered Scientific Research](https://www.k-dense.ai/blog/agent-skills-final-piece-for-ai-powered-research)** — What Agent Skills are, why curated domain guidance beats raw model capability, and an introduction to this repository. - **[Agent Skills: The Final Piece for AI-Powered Scientific Research](https://www.k-dense.ai/blog/agent-skills-final-piece-for-ai-powered-research)** — What Agent Skills are, why curated domain guidance beats raw model capability, and an introduction to this repository.
- **[K-Dense Web vs Scientific Agent Skills: Why We Built Both (And Which One You Should Use)](https://www.k-dense.ai/blog/k-dense-web-vs-scientific-agent-skills)** — When the open-source skills are the right tool, and when a hosted platform with managed compute makes more sense. - **[K-Dense Web vs Scientific Agent Skills: Why We Built Both (And Which One You Should Use)](https://www.k-dense.ai/blog/k-dense-web-vs-scientific-agent-skills)** — When the open-source skills are the right tool, and when a hosted platform with managed compute makes more sense.
- **[AI Co-Scientists, Answered: 20 Questions from a Live Session with a University Research Center](https://www.k-dense.ai/blog/ai-co-scientists-answered-20-questions)** — Practical questions from a research center evaluating AI co-scientists: what stays open source and MIT-licensed, how local and desktop deployments work, how data is handled, and how to choose between the hosted platform and the BYOK setup that runs these skills.
### Skill benchmarks and deep dives ### Skill benchmarks and deep dives
@@ -631,6 +648,13 @@ Deep dives, benchmarks, and guides from the [K-Dense blog](https://www.k-dense.a
- **[Benchmarking Nano Banana 2 Lite for Scientific Image Generation](https://www.k-dense.ai/blog/benchmarking-nano-banana-2-lite-scientific-image-model)** — A 240-image comparison of scientific-diagram models, useful when choosing a backend for [generate-image](skills/generate-image/): 3.8 s median latency for Nano Banana 2 Lite against 49 s for GPT Image 2, with a quality tradeoff. - **[Benchmarking Nano Banana 2 Lite for Scientific Image Generation](https://www.k-dense.ai/blog/benchmarking-nano-banana-2-lite-scientific-image-model)** — A 240-image comparison of scientific-diagram models, useful when choosing a backend for [generate-image](skills/generate-image/): 3.8 s median latency for Nano Banana 2 Lite against 49 s for GPT Image 2, with a quality tradeoff.
- **[Benchmarking NVIDIA BioNeMo Agent Toolkit Skills for NIM microservices](https://www.k-dense.ai/blog/benchmarking-nvidia-bionemo-nim-skill)** — A separate NVIDIA skill set rather than one of these, but the findings generalize: skills help most with routing to non-obvious endpoints and with weak-model reliability, and do not improve the underlying scientific model's accuracy. - **[Benchmarking NVIDIA BioNeMo Agent Toolkit Skills for NIM microservices](https://www.k-dense.ai/blog/benchmarking-nvidia-bionemo-nim-skill)** — A separate NVIDIA skill set rather than one of these, but the findings generalize: skills help most with routing to non-obvious endpoints and with weak-model reliability, and do not improve the underlying scientific model's accuracy.
### Why the workflow layer matters
- **[The Model Is No Longer the Bottleneck](https://www.k-dense.ai/blog/the-model-is-no-longer-the-bottleneck)** — The case for why a repository like this one exists: frontier models now match specialized scientific software on raw capability (±0.079 ppm on NMR hydrogen shift prediction), so the limiting factor has moved to the workflow around the model — data access, code execution, verification, and auditable output.
- **[The AI Co-Scientist Is Here. The Bottleneck Is Verification.](https://www.k-dense.ai/blog/ai-co-scientist-verification-bottleneck)** — A 10-point checklist for evaluating a research agent, built around exposing sources, code, data provenance, and intermediate work rather than a polished final answer — the same reasoning behind the provenance and retrieval-contract requirements in skills like [database-lookup](skills/database-lookup/) and [scientific-writing](skills/scientific-writing/).
- **[Reproduction, Not Generation, Is AI's Killer App for Science](https://www.k-dense.ai/blog/reproduction-not-generation-ai-for-science)** — Why re-running published analyses is the highest-value use of an agent: 78% of papers and 93% of individual analysis tasks reproduced across a 221-study benchmark, because a reproduction can be checked against known numbers while a generated claim cannot.
- **[Introducing K-Bench 01: Nine Frontier Models, 178 Real Scientific Tasks, and a Lot of Confident Wrong Answers](https://www.k-dense.ai/blog/introducing-k-bench-01-internal-benchmark)** — Nine frontier models on 178 real user tasks, with overclaiming in 40% of runs. Useful calibration for what to check when an agent reports success, and context for the verification boundaries written into the clinical, regulatory, and research-methodology skills above.
### Security and safe deployment ### Security and safe deployment
- **[Security in the Science Agent Era: What Every Lab Needs to Know Before Installing Skills](https://www.k-dense.ai/blog/skill-security-before-you-install)** — The practical review checklist behind this repo's [Security Disclaimer](#%EF%B8%8F-security-disclaimer): read the full `SKILL.md` and `scripts/`, scan before installing, and pin versions instead of tracking a branch. - **[Security in the Science Agent Era: What Every Lab Needs to Know Before Installing Skills](https://www.k-dense.ai/blog/skill-security-before-you-install)** — The practical review checklist behind this repo's [Security Disclaimer](#%EF%B8%8F-security-disclaimer): read the full `SKILL.md` and `scripts/`, scan before installing, and pin versions instead of tracking a branch.
@@ -817,7 +841,7 @@ Need help? Here's how to get support:
- 📖 **Documentation**: Check the relevant `SKILL.md` and `references/` folders - 📖 **Documentation**: Check the relevant `SKILL.md` and `references/` folders
- 🐛 **Bug Reports**: [Open an issue](https://github.com/K-Dense-AI/scientific-agent-skills/issues) - 🐛 **Bug Reports**: [Open an issue](https://github.com/K-Dense-AI/scientific-agent-skills/issues)
- 💡 **Feature Requests**: [Submit a feature request](https://github.com/K-Dense-AI/scientific-agent-skills/issues/new) - 💡 **Feature Requests**: [Submit a feature request](https://github.com/K-Dense-AI/scientific-agent-skills/issues/new)
- 📣 **Updates and demos**: Follow [X](https://x.com/k_dense_ai), [LinkedIn](https://www.linkedin.com/company/k-dense-inc), and [YouTube](https://www.youtube.com/@K-Dense-Inc) to keep up with new skills, tutorials, and Scientific Agent Skills releases - 📣 **Updates and demos**: Follow [X](https://x.com/k_dense_ai), [LinkedIn](https://www.linkedin.com/company/k-dense-inc), [YouTube](https://www.youtube.com/@K-Dense-Inc), and [Reddit](https://www.reddit.com/user/-k-dense-/) to keep up with new skills, tutorials, and Scientific Agent Skills releases
- 💼 **Enterprise Support**: Contact [K-Dense](https://k-dense.ai/) for commercial support - 💼 **Enterprise Support**: Contact [K-Dense](https://k-dense.ai/) for commercial support
--- ---
@@ -842,7 +866,7 @@ Recommended practice:
title = {Scientific Agent Skills: A Comprehensive Collection of Scientific Tools for AI Agents}, title = {Scientific Agent Skills: A Comprehensive Collection of Scientific Tools for AI Agents},
year = {2026}, year = {2026},
url = {https://github.com/K-Dense-AI/scientific-agent-skills}, url = {https://github.com/K-Dense-AI/scientific-agent-skills},
note = {159 skills covering databases, packages, integrations, and analysis tools} note = {161 skills covering databases, packages, integrations, and analysis tools}
} }
``` ```
@@ -905,3 +929,13 @@ See [LICENSE.md](LICENSE.md) for full terms.
### Individual Skill Licenses ### Individual Skill Licenses
> ⚠️ **Important**: Each skill has its own license specified in the `license` metadata field within its `SKILL.md` file. These licenses may differ from the repository's MIT License and may include additional terms or restrictions. **Users are responsible for reviewing and adhering to the license terms of each individual skill they use.** > ⚠️ **Important**: Each skill has its own license specified in the `license` metadata field within its `SKILL.md` file. These licenses may differ from the repository's MIT License and may include additional terms or restrictions. **Users are responsible for reviewing and adhering to the license terms of each individual skill they use.**
## Star History
<a href="https://star-history.dera.page/#K-Dense-AI/scientific-agent-skills">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://star-history.dera.page/svg?repos=K-Dense-AI/scientific-agent-skills&theme=dark" />
<source media="(prefers-color-scheme: light)" srcset="https://star-history.dera.page/svg?repos=K-Dense-AI/scientific-agent-skills" />
<img alt="Star History Chart" src="https://star-history.dera.page/svg?repos=K-Dense-AI/scientific-agent-skills" />
</picture>
</a>
@@ -0,0 +1,36 @@
---
title: "Plugin"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9e8b0cb0/plugin.json
upstream_sha: 9e8b0cb0
imported_at: 2026-08-18
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
{
"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
"name": "scientific-agent-skills",
"version": "2.64.0",
"description": "Ready-to-use scientific and research Agent Skills for biology, chemistry, medicine, and related workflows.",
"author": {
"name": "K-Dense Inc.",
"url": "https://k-dense.ai"
},
"homepage": "https://github.com/K-Dense-AI/scientific-agent-skills",
"repository": "https://github.com/K-Dense-AI/scientific-agent-skills",
"license": "MIT",
"keywords": [
"agent-skills",
"science",
"research",
"bioinformatics",
"cheminformatics",
"biology",
"chemistry",
"medicine"
]
}
@@ -0,0 +1,279 @@
---
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9e8b0cb0/skills/waypoint-bio/SKILL.md
upstream_sha: 9e8b0cb0
imported_at: 2026-08-18
prompt_class: catalogue
upstream_changes: accepted
name: waypoint-bio
description: Use when working with Outpost Bio's open microbiome foundation models - the Waypoint checkpoints (Waypoint-6m, Waypoint-45m, Waypoint-170m), the Atlas pretraining corpus, the Compass eight-task benchmark, or the `waypoint` CLI from the `waypoint-bio` package. Covers embedding microbiome samples, fine-tuning on taxonomic abundance data, benchmarking a checkpoint on Compass, pretraining a GPT-2 model on taxonomic abundance profiles, and converting MetaPhlAn, Kraken2, QIIME 2, or MGnify abundance tables into waypoint format.
license: MIT
compatibility: Requires Python 3.10+ with `waypoint-bio` (pulls torch, transformers, datasets, peft, scikit-learn). Needs network access and a Hugging Face token with access granted to the gated outpost-bio repos. A GPU is strongly recommended for pretraining and benchmarking.
metadata:
version: "1.0"
skill-author: K-Dense Inc.
upstream-version: "waypoint-bio 1.0.2 (PyPI); GitHub main 1.0.4"
last-reviewed: "2026-08-17"
openclaw:
primaryEnv: HF_TOKEN
envVars:
- name: HF_TOKEN
required: true
description: Hugging Face read token with access to the gated outpost-bio/Waypoint-*, outpost-bio/Atlas, and outpost-bio/Compass repos.
---
# Waypoint: Outpost Bio's Open Microbiome Foundation Models
## Overview
Outpost Bio open-sourced three artefacts under Apache 2.0, described in
[Treloar et al., bioRxiv 2026.05.02.722381](https://www.biorxiv.org/content/10.64898/2026.05.02.722381v2):
| Artefact | What it is | Hugging Face |
| --- | --- | --- |
| **Waypoint** | GPT-2-style causal LMs over taxonomic tokens, 6M170M params | `outpost-bio/Waypoint-6m`, `-45m`, `-170m` |
| **Atlas** | 539,308 microbiome samples scraped from MGnify (485,377 pretrain / 53,931 benchmark) | `outpost-bio/Atlas` |
| **Compass** | Eight downstream tasks over four studies | `outpost-bio/Compass` |
The unifying idea: a microbiome sample is a *sentence*. Each taxon is one token, tokens are ordered
by descending abundance z-score, and the model is trained with next-token prediction. A pretrained
checkpoint then supplies sample-level embeddings or a fine-tuning backbone for prediction tasks.
All of it is driven by one CLI, `waypoint`, with five subcommands: `prepare-dataset`, `embed`,
`finetune`, `benchmark`, `pretrain`.
## When to use
- Embedding 16S/shotgun taxonomic profiles into fixed-size vectors for clustering, visualisation, or
a downstream classifier.
- Fine-tuning a Waypoint checkpoint to predict a phenotype, treatment, or continuous readout from
community composition.
- Scoring your own microbiome model against Compass so the number is comparable to the paper.
- Pretraining a taxonomic language model on Atlas or on your own corpus.
- Converting profiler output (MetaPhlAn, Kraken2/Bracken, QIIME 2, MGnify TSVs) into the input format
these tools expect.
**Do not reach for this** when you have fewer than ~1,000 labelled samples — see
[Scientific caveats](#scientific-caveats). A random forest on relative abundances is the better tool
there, and the paper says so.
## Setup
```bash
pip install waypoint-bio # installs the `waypoint` command
```
Atlas, Compass, and every Waypoint checkpoint are **gated**. Access is auto-approved, but you must
click through once per repo and then authenticate:
1. Request access on each repo page you need: [Waypoint-6m](https://huggingface.co/outpost-bio/Waypoint-6m),
[Waypoint-45m](https://huggingface.co/outpost-bio/Waypoint-45m),
[Waypoint-170m](https://huggingface.co/outpost-bio/Waypoint-170m),
[Atlas](https://huggingface.co/datasets/outpost-bio/Atlas),
[Compass](https://huggingface.co/datasets/outpost-bio/Compass).
2. Authenticate locally:
```bash
hf auth login # or: export HF_TOKEN=hf_...
```
A 401/403 from any subcommand almost always means access was never requested on that specific repo —
a token alone is not enough. Use a read-scoped token. The tokenizer loads via
`trust_remote_code=True`, so pin a `revision` if you need the remote code fixed across runs.
## The waypoint data format
Everything except `prepare-dataset` consumes **waypoint format**: a `.parquet` / `.csv` / `.tsv`
whose rows are samples, with two aligned list-columns plus any label columns you need.
| Column | Type | Notes |
| --- | --- | --- |
| `Taxa` | `list[str]` | Full lineage strings, `;`-separated: `k__Bacteria; p__Firmicutes; ...; g__Lactobacillus` |
| `Relative Abundances` | `list[float]` | Same length as `Taxa`, same order |
| *(any)* | scalar | Targets, covariates, or a `Split` column |
Prefer parquet. CSV/TSV stores the lists as `repr` strings and round-trips through `ast.literal_eval`.
**Give full lineages, not bare names.** The tokenizer extracts the genus segment (`g__`) from each
lineage and falls back to the most specific higher rank when genus is missing. Bare names disable
that fallback entirely.
## Workflow
### 1. Get your data into waypoint format
If you already have a sample × taxa (or taxa × sample) abundance matrix with lineage labels:
```bash
waypoint prepare-dataset \
--input abundance_matrix.tsv \
--metadata sample_labels.csv \
--output dataset.parquet
```
Orientation is auto-detected from the first column header (`taxonomy`, `lineage`, `taxon`, `otu`,
`#otu id` ⇒ taxa-as-rows); override with `--orientation`. Rows are normalised to sum to 1 unless you
pass `--no_normalize`, and zeros are dropped unless you pass `--keep_zeros`.
`prepare-dataset` cannot read profiler output directly — MetaPhlAn uses `|` separators, Kraken2
reports encode the hierarchy as indentation, and QIIME 2/SILVA prefixes the domain `d__` instead of
`k__` (which the tokenizer silently ignores). Use the bundled converter for those:
```bash
python scripts/profiler_to_waypoint.py \
--input merged_metaphlan.tsv --format metaphlan \
--output dataset.parquet
python scripts/profiler_to_waypoint.py \
--input reports/*.kreport --format kraken \
--output dataset.parquet
python scripts/profiler_to_waypoint.py \
--input feature-table.tsv --format qiime2 \
--output dataset.parquet
```
See `references/data-preparation.md` for every input layout, rank handling, and the `d__`/`|` gotchas.
### 2. Check vocabulary coverage before anything else
Waypoint's vocabulary is fixed at pretraining time from Atlas. Taxa absent from it become `<unk>` and
are **silently dropped** by `waypoint embed`; the paper names this as the models' main limitation. A
sample whose taxa are all out-of-vocabulary yields a degenerate `[BOS][EOS]` embedding.
```bash
python scripts/vocab_coverage.py --model outpost-bio/Waypoint-6m --data dataset.parquet
```
It reports per-sample and abundance-weighted coverage and flags samples below a threshold. Treat
median abundance-weighted coverage under ~0.8 as a reason to re-examine your taxonomy labels before
trusting any downstream number.
### 3. Embed samples
```bash
waypoint embed \
--model outpost-bio/Waypoint-6m \
--data dataset.parquet \
--output embeddings.parquet
```
Output is indexed by sample ID with columns `dim_0 … dim_{H-1}` (`H` = 256 for 6m, 512 for 45m,
768 for 170m). Defaults: `--pooling last_token`, `--batch_size 32`, `--max_length 512`, device
auto-detected (`cuda` → `mps` → `cpu`).
Keep `--pooling last_token` unless you have a reason to change it: it matches how the checkpoints
were pretrained and how `benchmark` and `finetune` pool. `mean` is a reasonable alternative for
unsupervised use; `first_token`/`cls_token` return the BOS position and carry little signal in a
causal LM.
### 4. Fine-tune on your labels
```bash
# classification
waypoint finetune \
--model outpost-bio/Waypoint-45m \
--data dataset.parquet \
--output_dir outputs/ft_disease \
--task_type classification \
--target "Disease Status" \
--config configs/finetune_classification.yaml
# regression, with a categorical covariate one-hot appended to the pooled embedding
waypoint finetune \
--model outpost-bio/Waypoint-45m \
--data dataset.parquet \
--output_dir outputs/ft_degradation \
--task_type regression \
--target "Degradation Rate" \
--covariate_column Drug \
--config configs/finetune_regression.yaml
```
Config paths resolve against the bundled `waypoint_bio/configs/` tree, so `configs/...` works from
any directory without cloning.
Defaults worth overriding for small datasets: `warmup_steps: 1000` (drop to ~50 so warmup finishes
before early stopping), `num_epochs: 1` in the shipped configs (raise it — early stopping on
validation loss is what actually terminates training), and `use_lora: true` when VRAM is tight
(~1% of parameters trained; adapters are merged back before saving, so the checkpoint stays a plain
`AutoModel`).
Splits default to a random 80/10/10. **Set `split_column` to a `Split` column whenever samples are
correlated** — repeated measures, one donor sampled over time, technical replicates — or a random
split leaks and the test score is meaningless.
Outputs land in `--output_dir`: `best_model/` (loadable by `embed`/`benchmark`),
`test_metrics.json`, `training_log.csv` + `.html`, and `finetune_results.json`.
### 5. Benchmark on Compass
```bash
waypoint benchmark --model outpost-bio/Waypoint-6m --output_dir outputs/benchmark
waypoint benchmark --model outputs/pretrain/best_model --tasks 1 6 --output_dir outputs/smoke
```
Fine-tunes a fresh head per task and writes `benchmark_results.json`. Classification tasks score
macro-F1; the one regression task scores R² clamped to [0, 1]; `final_score` is the unweighted mean
across tasks. Full task table, metric keys, and result-file schema: `references/compass-benchmark.md`.
### 6. Pretrain
```bash
waypoint pretrain \
--model_config configs/models/gpt2-45m.yaml \
--pretrain_config configs/pretraining.yaml \
--output_dir outputs/pretrain_45m
```
Downloads Atlas, builds a taxonomic tokenizer from the corpus, computes per-token abundance
mean/std for z-score ordering, then trains with next-token prediction and early stopping. Add
`--data my_corpus.parquet` to pretrain on your own waypoint-format corpus instead, and
`--max_samples N` for a smoke test.
Nine architectures ship, from `gpt2-6m.yaml` (8 layers, 256 hidden) to `gpt2-170m.yaml` (24 layers,
768 hidden); per-head dimension is fixed at 64 throughout. `references/cli-reference.md` has the
full table and every config key.
## Scientific caveats
These are load-bearing. Ignoring them produces numbers that look fine and mean nothing.
- **Below ~1,000 labelled examples, Waypoint underperforms a random forest on raw abundances.** The
paper's crossover against the RF baseline sits near **10,000** training examples. Fit the baseline
first; only adopt the transformer if it wins on your data.
- **Out-of-vocabulary taxa are dropped, not flagged.** Every Compass dataset carries some. Run
`scripts/vocab_coverage.py` and report the coverage alongside your results.
- **45M, not 170M, was the best benchmark model.** Pretraining loss keeps falling with scale, but
downstream Compass score does not — start at 6m or 45m and only scale up if it demonstrably helps.
- **Genus-level tokenisation is the default**, so species-level distinctions are collapsed. Changing
`taxon_rank` requires re-pretraining, not just re-tokenising.
- **Compositional data.** Relative abundances are constrained to sum to 1; differences in one taxon
induce apparent changes in others. This affects interpretation of any per-taxon attribution.
- **Batch and study effects dominate microbiome data.** Atlas spans MGnify pipelines v1.0v5.0 and
four sequencing modalities. Never let a study or run boundary coincide with your label boundary.
- **Not a clinical or diagnostic tool.** The model cards state this explicitly.
## References
- `references/cli-reference.md` — every subcommand flag, every config key, the model-size table.
- `references/compass-benchmark.md` — the eight tasks, filters, metrics, `benchmark_results.json` schema.
- `references/data-preparation.md` — waypoint format, profiler conversions, taxonomy string rules.
- `references/python-api.md` — using the tokenizer, datasets, heads, and checkpoints from Python.
## Scripts
- `scripts/profiler_to_waypoint.py` — MetaPhlAn / Kraken2 / QIIME 2 / generic lineage tables → waypoint format.
- `scripts/vocab_coverage.py` — tokenizer coverage report for a waypoint-format file.
## Upstream
Code [github.com/Outpost-Bio/waypoint](https://github.com/Outpost-Bio/waypoint) ·
package `waypoint-bio` ·
paper [bioRxiv 2026.05.02.722381](https://www.biorxiv.org/content/10.64898/2026.05.02.722381v2) ·
community [Waypoint Slack](https://join.slack.com/t/outpostbio-waypoint/shared_invite/zt-3w6ivgtba-WJOCkdxiISxQpwVq9ZZxTA) ·
contact `waypoint@outpost.bio`.
Cite Treloar, N. J., Ur-Rehman, S., Yang, J., & Outpost Bio (2026). *Learning the Language of the
Microbiome with Transformers.* bioRxiv. Per-artefact DOIs are listed at
[outpost.bio/citations](https://www.outpost.bio/citations).
@@ -0,0 +1,223 @@
---
title: "`waypoint` CLI reference"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9e8b0cb0/skills/waypoint-bio/references/cli-reference.md
upstream_sha: 9e8b0cb0
imported_at: 2026-08-18
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
# `waypoint` CLI reference
Targets `waypoint-bio` 1.0.2 (PyPI) / 1.0.4 (GitHub main, commit `f45eee6`, 2026-07-16).
```
waypoint {pretrain,benchmark,finetune,embed,prepare-dataset} ...
```
Config paths are resolved first against the working directory, then against the bundled
`waypoint_bio/configs/` tree inside the installed wheel. So `--config configs/benchmark.yaml`
works from anywhere without cloning the repo. The same fallback applies to the bundled example
data (`examples/abundance_matrix.tsv`, `examples/finetune_classification.parquet`, …).
---
## `waypoint prepare-dataset`
Converts a sample × taxa abundance matrix into waypoint format.
| Flag | Default | Notes |
| --- | --- | --- |
| `--input` | *required* | `.csv` / `.tsv` abundance matrix. |
| `--output` | *required* | `.parquet` recommended; `.csv` supported. |
| `--orientation` | `auto` | `auto`, `samples_as_rows`, `taxa_as_rows`. |
| `--taxonomy_format` | `full` | `full` for lineage strings; a rank name (`genus`, `species`, …) to prefix bare names. |
| `--no_normalize` | off | Skip row-normalisation to relative abundances. |
| `--keep_zeros` | off | Keep zero-abundance entries in each sample's lists. |
| `--metadata` | none | CSV/TSV/parquet of per-sample metadata, indexed by sample ID, merged in as extra columns. |
`auto` treats the file as taxa-as-rows when the first column header is `taxonomy`, `lineage`,
`taxon`, `otu`, or `#otu id` (case-insensitive); otherwise samples-as-rows with the first column
as the sample ID.
`--taxonomy_format genus` prefixes bare column names with `g__`. It disables higher-rank fallback,
because a bare name carries no lineage to fall back to — prefer real lineage strings.
---
## `waypoint embed`
One fixed-size vector per sample from a pretrained checkpoint. No fine-tuning, no labels needed.
| Flag | Default | Notes |
| --- | --- | --- |
| `--model` | `outpost-bio/Waypoint-6m` | Hub id or local checkpoint directory. |
| `--data` | *required* | Waypoint-format `.parquet` / `.csv` / `.tsv`. |
| `--output` | *required* | `.parquet`, or `.csv` if the path ends in `.csv`. |
| `--pooling` | `last_token` | `last_token`, `mean`, `first_token`, `cls_token`. |
| `--batch_size` | `32` | |
| `--max_length` | `512` | Truncates after ordering, so the least informative taxa are lost first. |
| `--device` | auto | `cuda`, `mps`, or `cpu`; auto-detects in that order. |
Output columns are `dim_0 … dim_{H-1}`, indexed by sample ID. Hidden size `H` is 256 (6m),
512 (45m), 768 (170m).
**Behaviour worth knowing:** tokens that map to `<unk>` are *dropped* before ordering, not encoded.
A row with no in-vocabulary taxa still produces an output row, but its sequence is `[BOS][EOS]` and
the embedding is meaningless. Run `scripts/vocab_coverage.py` first.
Ordering: by descending abundance z-score when `token_std_means.parquet` is present (it ships with
every published checkpoint and with `waypoint pretrain` output), otherwise by descending raw
relative abundance.
---
## `waypoint finetune`
Fine-tunes a checkpoint on your own labelled waypoint-format data.
| Flag | Default | Notes |
| --- | --- | --- |
| `--model` | *required* | Hub id or local checkpoint. |
| `--data` | *required* | Waypoint-format file containing `--target`. |
| `--output_dir` | *required* | |
| `--task_type` | *required* | `classification` or `regression`. |
| `--target` | *required* | Target column name. |
| `--covariate_column` | none | Categorical column, one-hot encoded and concatenated to the pooled embedding before the head. |
| `--config` | task default | Flat YAML; defaults to the bundled classification/regression config. |
### Fine-tuning config keys
```yaml
split_column: null # column holding train/validation/test; null = random split
val_fraction: 0.1
test_fraction: 0.1
max_length: 512 # must match the checkpoint's pretraining context
pooling_strategy: last_token
filter_unk_taxa: true # drop out-of-vocabulary taxa rather than feed <unk>
seed: 42
learning_rate: 0.00003
num_epochs: 1 # raise this; early stopping is what should terminate training
batch_size: 64
warmup_steps: 1000 # lower to ~50 for small datasets
weight_decay: 0.001
eval_strategy: steps
eval_steps: 400
logging_steps: 5
patience: 5 # eval steps without improvement before early stopping
save_total_limit: 1
use_lora: false
lora_r: 8
lora_alpha: 16 # convention: 2 * r
lora_dropout: 0.05
lora_target_modules: [c_attn, c_proj] # GPT-2 fused QKV and output projection
lora_bias: none
lora_fan_in_fan_out: true # required for GPT-2 Conv1D layouts
```
`num_epochs: 1` in the shipped configs is tuned for the large Compass tasks. On a few-thousand-row
dataset one epoch is a handful of optimizer steps and the model barely moves — raise `num_epochs`
and let `patience` stop it. Likewise `eval_steps: 400` may never fire; lower it so early stopping
and best-checkpoint selection can actually work.
LoRA adapters are merged back into the base transformer before saving, so `best_model/` loads with
a plain `AutoModel.from_pretrained` and works with `waypoint embed` and `waypoint benchmark`.
### Outputs
| Path | Contents |
| --- | --- |
| `best_model/` | Fine-tuned base transformer in standard HF format, plus tokenizer and `token_std_means.parquet`. |
| `best_model/finetuned_model_state.pt` | Full torch state dict: transformer + head + covariate embedding. |
| `validation_metrics.json`, `test_metrics.json` | Per-split scores, benchmark-equivalent. |
| `training_log.csv`, `training_log.html` | Every row of `trainer.state.log_history`; the HTML is an interactive plotly line plot. |
| `finetune_results.json` | Run config, label maps, covariate map, val/test scores. |
---
## `waypoint benchmark`
| Flag | Default | Notes |
| --- | --- | --- |
| `--model` | `outpost-bio/Waypoint-6m` | Hub id or local checkpoint. |
| `--config` | bundled `configs/benchmark.yaml` | Shared by all eight tasks. |
| `--output_dir` | `outputs/benchmark` | |
| `--tasks` | all 8 | Space-separated task numbers, e.g. `--tasks 1 6`. |
| `--seed` | `42` | |
| `--max_samples` | none | Caps each split; use for smoke tests only, never for a reported score. |
`configs/benchmark.yaml` is the fine-tuning config applied identically to every task:
`learning_rate: 3e-5`, `num_epochs: 1`, `batch_size: 64`, `warmup_steps: 1000`,
`weight_decay: 0.001`, `patience: 5`, `pooling_strategy: last_token`, `eval_steps: 400`,
`filter_unk_taxa: true`, `seed: 42`. Change it and your score is no longer comparable to the paper.
The paper reports means over three independent runs. A single run is noisy; vary `--seed` and
report the spread.
---
## `waypoint pretrain`
| Flag | Default | Notes |
| --- | --- | --- |
| `--model_config` | `configs/models/gpt2-6m.yaml` | Architecture YAML. |
| `--pretrain_config` | `configs/pretraining.yaml` | Hyperparameter YAML. |
| `--output_dir` | `outputs/pretrain` | Best checkpoint written to `<output_dir>/best_model/`. |
| `--max_samples` | none | Limit training samples for a quick test. |
| `--data` | none | Local waypoint-format corpus instead of downloading Atlas. |
Steps: download the Atlas `pretrain` split → build a taxonomic tokenizer from the corpus →
compute per-token abundance mean/std for z-score ordering → train GPT-2 with next-token prediction
and early stopping → save `best_model/`.
### `configs/pretraining.yaml`
```yaml
training_type: next_token_prediction
taxon_rank: genus # tokenization rank; changing it means re-pretraining
fallback_to_higher_rank: true # use the most specific higher rank when genus is absent
max_length: 512
learning_rate: 0.001
warmup_steps: 1000
weight_decay: 0.001
batch_size: 32
num_epochs: 100
patience: 10
eval_steps: 3261
save_steps: 3261
logging_steps: 100
val_split: 0.1
seed: 42
```
### Architectures
All share `model_type: gpt2`, `n_positions: 512`, and a fixed per-head dimension of 64.
| Config | Layers | Hidden | Heads | ~Params |
| --- | --- | --- | --- | --- |
| `gpt2-6m.yaml` | 8 | 256 | 4 | 6M |
| `gpt2-6m-mgm.yaml` | 8 | 256 | 8 | 6M — matches the MGM baseline architecture |
| `gpt2-10m.yaml` | 8 | 320 | 5 | 10M |
| `gpt2-18m.yaml` | 10 | 384 | 6 | 18M |
| `gpt2-29m.yaml` | 12 | 448 | 7 | 29M |
| `gpt2-45m.yaml` | 14 | 512 | 8 | 45M |
| `gpt2-79m.yaml` | 16 | 640 | 10 | 79M |
| `gpt2-85m-gpt-small.yaml` | 12 | 768 | 12 | 85M — GPT-2 small geometry |
| `gpt2-170m.yaml` | 24 | 768 | 12 | 170M |
Only 6m, 45m, and 170m are published as checkpoints. The rest exist so the paper's scaling study is
reproducible; `gpt2-6m-mgm` isolates the effect of head count against the MGM baseline.
Parameter counts exclude token and positional embeddings, so the Hub's reported sizes are larger
(the 6m checkpoint reports ~10.1M, the 45m ~51.8M).
Pretraining Atlas end to end is a multi-GPU-day job. Validate the pipeline with
`--max_samples 5000` before committing to a full run.
@@ -0,0 +1,137 @@
---
title: "Compass: the eight-task microbiome benchmark"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9e8b0cb0/skills/waypoint-bio/references/compass-benchmark.md
upstream_sha: 9e8b0cb0
imported_at: 2026-08-18
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
# Compass: the eight-task microbiome benchmark
`outpost-bio/Compass` on the Hugging Face Hub — gated, Apache 2.0, ~605 MB, ~62.8k rows across four
Hub configurations. Eight tasks are derived from those four configurations by filtering and by
choosing different target columns.
Every configuration exposes `train` / `validation` / `test` splits and carries a `Split` column
recording the same assignment.
```python
from datasets import load_dataset
ds = load_dataset("outpost-bio/Compass", "mgnify-biomes") # requires access + HF_TOKEN
```
## The four source datasets
| Config | Source | Rows (train/val/test) | Extra columns |
| --- | --- | --- | --- |
| `mgnify-biomes` | MGnify metagenomic profiles across gut, skin, oral, marine, freshwater, soil, engineered systems | 33,121 / 4,139 / 4,139 | `Biome 1``Biome 5`, `Run Accession`, `Data Type`, `Sequencing Method`, `Pipeline Version`, `Study Accession` |
| `handuo` | Han, Duo et al. — 16S amplicon study of drugmicrobiome interactions in stool-derived communities | 3,168 / 396 / 396 | `SIC Name`, `Control`, `ATC Class`, `Sample ID` |
| `mastrorilli` | Mastrorilli et al. — drug degradation by gut communities | 9,282 / 3,084 / 3,053 | `Degradation Rate`, `Drug`, `Sample ID` |
| `roswall` | Roswall et al. — longitudinal infant gut cohort | 2,031 total | `Timepoint`, `Delivery Mode`, `Sample ID` |
All configs carry `Taxa` and `Relative Abundances` as aligned list columns.
## The eight tasks
As defined in `waypoint_bio/benchmark.py`:
| # | Internal id | Config | Targets | Type | Pre-filter |
| --- | --- | --- | --- | --- | --- |
| 1 | `1_biome` | `mgnify-biomes` | `Biome 1``Biome 5` | classification (5 outputs) | none |
| 2 | `2_biome_gut` | `mgnify-biomes` | `Biome 4`, `Biome 5` | classification (2 outputs) | `Biome 3 == "Digestive system"` |
| 3 | `3_sic` | `handuo` | `SIC Name` | classification | `SIC Name` starts with `SIC`, excludes `control` and `seed` |
| 4 | `4_drug_non_drug` | `handuo` | `Control` | binary classification | none |
| 5 | `5_drug_class` | `handuo` | `ATC Class` | classification | `ATC Class` not null |
| 6 | `6_drug_degradation` | `mastrorilli` | `Degradation Rate` | regression | none; `Drug` used as covariate |
| 7 | `7_infant_age` | `roswall` | `Timepoint` | classification | none |
| 8 | `8_birth_mode` | `roswall` | `Delivery Mode` | binary classification | none |
What each asks, in plain terms:
1. **Biome classification** — predict all five levels of the MGnify biome ontology at once
(e.g. `root → Host-associated → Human → Digestive system → Large intestine`).
2. **Gut biome classification** — same, restricted to digestive-system samples, predicting only the
two finest levels. Harder: the easy environmental separations are gone.
3. **SIC classification** — identify which stool-derived in-vitro community a drug-perturbed sample
came from.
4. **Drug vs. control** — did this community receive a drug?
5. **Drug class** — recover the ATC class of the applied drug from the resulting composition.
6. **Drug degradation** — regress the degradation rate from composition plus drug identity. The
`Drug` covariate is one-hot encoded and concatenated to the pooled embedding.
7. **Infant age** — predict the sampling timepoint from an infant gut sample.
8. **Birth mode** — vaginal vs. caesarean delivery.
## Scoring
- **Classification:** macro-averaged F1 — F1 per class, averaged with equal weight. Chosen so the
metric is not dominated by majority classes. Where a task has several target columns (1 and 2),
the per-target macro-F1s are averaged.
- **Regression (task 6):** R², clamped to `[0, 1]` so it shares a scale with the F1 scores. A
negative R² therefore reads as `0.0`, not as "worse than the mean".
- **Final score:** unweighted arithmetic mean of the eight task scores.
Supplementary metrics are computed and stored but do not enter the score: one-vs-one macro ROC-AUC,
macro PR-AUC (pairwise average precision over the same OVO pairs), balanced accuracy, plain
accuracy; and MSE, Pearson, Spearman for regression.
## `benchmark_results.json`
```
benchmark_results.json
├── model string — the value passed to --model
├── final_score number — mean of every results[].score
└── results array, one object per task
├── task string — "1_biome", "6_drug_degradation", ...
├── task_type "classification" | "regression"
├── score number — macro F1, or R² clamped to [0,1]
└── metrics object — keys depend on task_type
```
`metrics` keys are suffixed with the target column name:
| Task type | Keys |
| --- | --- |
| `classification` | `accuracy_<target>`, `balanced_accuracy_<target>`, `f1_macro_<target>`; with probabilities, binary `roc_auc_<target>` / `pr_auc_<target>` or multiclass `roc_auc_macro_ovo_<target>` / `pr_auc_macro_ovo_<target>`. Means: `f1_macro_mean`, optionally `roc_auc_mean`, `pr_auc_mean`. |
| `regression` | `mse_<target>`, `r2_<target>`, usually `pearson_<target>` and `spearman_<target>`. Mean: `r2_mean`. |
Example:
```json
{
"model": "outpost-bio/Waypoint-6m",
"final_score": 0.71,
"results": [
{"task": "1_biome", "task_type": "classification", "score": 0.65,
"metrics": {"f1_macro_mean": 0.65, "roc_auc_mean": 0.81, "pr_auc_mean": 0.74}},
{"task": "6_drug_degradation", "task_type": "regression", "score": 0.42,
"metrics": {"mse_Degradation Rate": 0.019, "r2_Degradation Rate": 0.44, "r2_mean": 0.44}}
]
}
```
The numbers above are the illustrative values from the upstream README, not measured results.
## Interpreting a benchmark run
**Baselines matter more than the absolute score.** The paper compares Waypoint against classical
baselines (random forest and logistic regression on relative abundances) and against MGM, the prior
microbiome foundation model. Two findings shape how a Compass number should be read:
- Waypoint beats the random-forest baseline from roughly **10,000 training examples upward**, and
*loses* to it below about 1,000. Report the training-set size next to any score.
- Baselines can use every taxon; the transformer sees only its fixed vocabulary. The paper's fair
comparison is the `(no unk)` baseline, with out-of-vocabulary taxa stripped from the baseline's
input too. Compare against that, not against a baseline given the full table.
**Scale does not monotonically help.** Pretraining loss falls all the way to 170M, but the best
Compass score in the paper came from the **45M** model. Non-pretrained transformers get *worse* as
they grow — the gain from scale is a property of pretraining, not of capacity.
**Reproducibility.** Use the bundled `configs/benchmark.yaml` unchanged, do not pass `--max_samples`,
and run at least three seeds. Comparing a run that changed the learning rate or capped splits against
published numbers is not a comparison.
@@ -0,0 +1,213 @@
---
title: "Preparing data for Waypoint"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9e8b0cb0/skills/waypoint-bio/references/data-preparation.md
upstream_sha: 9e8b0cb0
imported_at: 2026-08-18
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
# Preparing data for Waypoint
## Waypoint format
Rows are samples. Two aligned list-columns, plus whatever labels you need.
| Column | Type | Required |
| --- | --- | --- |
| `Taxa` | `list[str]` — full lineage strings | yes |
| `Relative Abundances` | `list[float]` — same length and order as `Taxa` | yes |
| `Split` | `str``train` / `validation` / `test` | only when using `split_column` |
| *(any)* | scalar targets and covariates | as needed |
The DataFrame index holds the sample ID and is preserved through `embed`.
Use `.parquet`. CSV/TSV works but stores each list as its Python `repr`, parsed back with
`ast.literal_eval` — brittle and large.
```python
import pandas as pd
df = pd.DataFrame(
{
"Taxa": [["k__Bacteria; p__Firmicutes; c__Bacilli; o__Lactobacillales; f__Lactobacillaceae; g__Lactobacillus",
"k__Bacteria; p__Bacteroidota; c__Bacteroidia; o__Bacteroidales; f__Bacteroidaceae; g__Bacteroides"]],
"Relative Abundances": [[0.41, 0.59]],
"Group": ["Case"],
},
index=pd.Index(["sample_001"], name="sample_id"),
)
df.to_parquet("dataset.parquet")
```
## How taxonomy strings are read
`TaxonomicTokenizer` splits each lineage on `;`, strips whitespace, and inspects each segment's
three-character prefix:
| Prefix | Rank |
| --- | --- |
| `s__` | species |
| `g__` | genus |
| `f__` | family |
| `o__` | order |
| `c__` | class |
| `p__` | phylum |
| `k__` | kingdom |
With `taxon_rank: genus` and `fallback_to_higher_rank: true` (the published defaults), each lineage
becomes one token:
1. If a `g__` segment exists, that segment *including the prefix* is the token — `g__Lactobacillus`.
2. Otherwise the **most specific higher rank** present is used — a lineage stopping at
`f__Lactobacillaceae` tokenises to `f__Lactobacillaceae`.
3. If nothing matches, the token is `<unk>`.
Consequences that bite:
- **Any prefix outside that table is invisible.** QIIME 2 / SILVA / Greengenes2 write the domain as
`d__Bacteria`; `d__` is not in the table, so such a segment is skipped entirely. A lineage
truncated at domain becomes `<unk>`. Rewrite `d__` to `k__`.
- **A `s__` species segment does not help by itself.** Species is *more* specific than genus, so
fallback (which only goes up) cannot use it. A lineage with `s__` but no `g__` tokenises to
whatever higher rank is present — or `<unk>` if none is. Keep the full lineage, not just the tip.
- **Separator is `;`, not `|`.** A `|`-joined MetaPhlAn lineage is one unsplittable segment. Its
first three characters are `k__`, so it matches at kingdom rank and the *entire pipe-joined
string* is returned as a single token — which is not in the vocabulary, so it becomes `<unk>`.
Verified against `TaxonomicTokenizer` 1.0.2:
`k__Bacteria|p__Firmicutes|g__Lactobacillus` extracts to itself, while the `;`-separated form
extracts to `g__Lactobacillus`.
- **Bare names never tokenise.** `Lactobacillus` has no prefix. Use `prepare-dataset
--taxonomy_format genus` to prefix them, accepting the loss of fallback.
## Token ordering and truncation
Samples are encoded as `[BOS] + ordered_token_ids + [EOS]`, padded to `max_length` (512).
Ordering is by **descending abundance z-score** — `(ra - mean) / std` per token, using
`token_std_means.parquet` from the checkpoint. This puts taxa that are unusually abundant *for that
taxon* first, rather than merely abundant. Without that file, ordering falls back to raw descending
abundance.
Because truncation is applied after ordering, a sample with more than 510 in-vocabulary taxa loses
its least distinctive ones. That is the intended behaviour, but it means `max_length` interacts with
how deeply you profiled.
## Out-of-vocabulary taxa
The vocabulary is frozen at pretraining time from the Atlas corpus. During `waypoint embed`, tokens
resolving to `<unk>` are dropped before ordering; during fine-tuning and benchmarking,
`filter_unk_taxa: true` does the same. Neither warns you.
Every Compass dataset carries out-of-vocabulary taxa, and the paper names this the models' key
limitation. Measure it before drawing conclusions:
```bash
python scripts/vocab_coverage.py --model outpost-bio/Waypoint-6m --data dataset.parquet
```
If coverage is poor, the usual causes are, in order: a different taxonomy database (SILVA vs. NCBI
vs. GTDB naming), the `d__` prefix problem, `|` separators, and genuinely novel environments.
## Converting profiler output
`waypoint prepare-dataset` reads a plain abundance matrix whose labels are already `;`-separated
lineages. `scripts/profiler_to_waypoint.py` handles the formats it cannot.
### MetaPhlAn
Merged tables from `merge_metaphlan_tables.py`: rows are clades with `|`-separated lineages, columns
are samples, values are **percentages**, and the table is cumulative — every rank appears as its own
row.
```bash
python scripts/profiler_to_waypoint.py \
--input merged_abundance_table.txt --format metaphlan \
--rank species --output dataset.parquet
```
The converter drops `#` comment lines and the `NCBI_tax_id` / `clade_taxid` column, keeps only rows
whose deepest rank equals `--rank` (default `species`, which avoids double-counting parents),
rewrites `|` to `; `, and renormalises each sample to sum to 1. Strain rows (`t__`) are always
excluded.
### Kraken2 / Bracken
Kraken2 reports are per-sample and encode the hierarchy as two-space indentation, with no lineage
string. Pass one report per sample:
```bash
python scripts/profiler_to_waypoint.py \
--input reports/*.kreport --format kraken \
--rank species --output dataset.parquet
```
The converter walks the indentation to rebuild each lineage, maps Kraken rank codes to prefixes
(`D`/`K` → `k__`, `P` → `p__`, `C` → `c__`, `O` → `o__`, `F` → `f__`, `G` → `g__`, `S` → `s__`),
skips sub-ranks (`D1`, `S1`, …) and unclassified rows, takes clade-level read counts at the target
rank, and normalises. Sample IDs come from the filenames. Both the 6-column and the 8-column
(`--report-minimizer-data`) layouts are handled.
Bracken's own `.bracken` output carries no lineage at all — use the Kraken-style report Bracken
writes with `-o`/`--report`, not the tabular abundance file.
### QIIME 2 / biom TSV
Exported feature tables with a `taxonomy` column (or `#OTU ID` rows already labelled by lineage):
```bash
python scripts/profiler_to_waypoint.py \
--input feature-table.tsv --format qiime2 \
--taxonomy-column taxonomy --output dataset.parquet
```
The converter strips the `# Constructed from biom file` banner, uses the taxonomy column as the
lineage, rewrites `d__` to `k__`, and normalises counts to relative abundances. Features whose
taxonomy is `Unassigned` are dropped.
### MGnify
MGnify amplicon abundance TSVs are taxa-as-rows with a `taxonomy` first column and `;`-separated
lineages — the native layout. `waypoint prepare-dataset --orientation auto` reads them directly; no
conversion needed. This is the format Atlas itself was built from.
### Anything else
If you already have a sample × taxa table with lineage labels, use `--format generic`, which applies
only the separator and prefix normalisation:
```bash
python scripts/profiler_to_waypoint.py \
--input my_table.tsv --format generic --orientation taxa_as_rows \
--output dataset.parquet
```
## Attaching labels
Either merge them at conversion time —
```bash
waypoint prepare-dataset --input matrix.tsv --metadata labels.csv --output dataset.parquet
python scripts/profiler_to_waypoint.py --input ... --metadata labels.csv --output dataset.parquet
```
— where `labels.csv` is indexed by sample ID, or join afterwards in pandas. Sample IDs must match
exactly; the converters do an inner-style alignment and will silently produce `NaN` targets for
unmatched rows, which then fail at fine-tuning time.
## Splits
`waypoint finetune` defaults to a random 80/10/10 split. Add a `Split` column and set
`split_column: Split` in the config whenever samples are not independent:
- longitudinal cohorts (the Roswall infant data is exactly this shape),
- technical or biological replicates,
- multiple communities derived from one donor,
- multiple drugs applied to the same starting community.
Grouping by subject or study when you build `Split` is the difference between a generalisation
estimate and a memorisation estimate.
@@ -0,0 +1,232 @@
---
title: "Using Waypoint from Python"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9e8b0cb0/skills/waypoint-bio/references/python-api.md
upstream_sha: 9e8b0cb0
imported_at: 2026-08-18
prompt_class: unknown
upstream_changes: accepted
author: upstream
validated: false
---
# Using Waypoint from Python
The CLI covers the standard paths. Drop to Python when you need a custom training loop, a different
head, or embeddings inside a larger pipeline.
## Package surface
`waypoint_bio` lazily re-exports:
```python
from waypoint_bio import (
TaxonomicTokenizer, # the tokenizer class
load_tokenizer, # load one from a Hub id or local dir
MicrobiomePretrainingDataset, # causal-LM dataset
MicrobiomeBenchmarkDataset, # supervised dataset with targets/covariates
load_waypoint_dataframe, # read waypoint-format parquet/csv/tsv
load_abundance_matrix, # read a sample x taxa matrix
matrix_to_waypoint_df, # matrix -> waypoint format
)
```
Imports are deferred, so `import waypoint_bio` does not pull in torch.
## Loading a checkpoint directly with transformers
The tokenizer is custom and ships as remote code, so `trust_remote_code=True` is required for it.
The model itself is a stock GPT-2 and does not need it.
```python
from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("outpost-bio/Waypoint-45m", trust_remote_code=True)
model = AutoModel.from_pretrained("outpost-bio/Waypoint-45m") # gated: needs HF_TOKEN
```
`trust_remote_code=True` executes the tokenizer code stored in the repo. Pin a revision when that
matters to you, so the code cannot change under a later run:
```python
tok = AutoTokenizer.from_pretrained(
"outpost-bio/Waypoint-45m", trust_remote_code=True, revision="1664ab5"
)
```
`AutoModelForCausalLM` also works if you want the LM head for likelihood scoring or generation —
generation samples taxa, which is occasionally useful for probing what the model learned about
co-occurrence, but is not a validated use.
## Tokenizing by hand
```python
from waypoint_bio import load_tokenizer
tok = load_tokenizer("outpost-bio/Waypoint-6m")
lineage = "k__Bacteria; p__Firmicutes; c__Bacilli; o__Lactobacillales; f__Lactobacillaceae; g__Lactobacillus"
print(tok.tokenize(lineage)) # ['g__Lactobacillus']
print(tok.convert_tokens_to_ids(["g__Lactobacillus"]))
# One sample = newline-separated lineages
sample = "\n".join([lineage, "k__Bacteria; p__Bacteroidota; g__Bacteroides"])
print(tok(sample)["input_ids"])
```
Checking whether a taxon is in vocabulary:
```python
vocab = tok.get_vocab()
"g__Lactobacillus" in vocab # True for anything seen in Atlas
tok.convert_tokens_to_ids("g__Nonesuch") == tok.unk_token_id
```
`tok._extract(lineage)` applies the rank extraction and higher-rank fallback and returns the token
string, or `None`. It is private but stable across 1.0.x and is what the datasets and
`scripts/vocab_coverage.py` use.
## Building a dataset
```python
import pandas as pd
from waypoint_bio import MicrobiomePretrainingDataset, load_tokenizer, load_waypoint_dataframe
from waypoint_bio.dataset import try_load_token_std_means
df = load_waypoint_dataframe("dataset.parquet")
tok = load_tokenizer("outpost-bio/Waypoint-6m")
stats = try_load_token_std_means("outpost-bio/Waypoint-6m") # None if absent
ds = MicrobiomePretrainingDataset(df, tok, max_length=512, token_std_means=stats)
ds[0]["input_ids"].shape # torch.Size([512])
```
Each item is `[BOS] + z-score-ordered token ids + [EOS]`, right-padded.
Computing the ordering statistics for a corpus of your own:
```python
from waypoint_bio.dataset import compute_token_std_means
stats = compute_token_std_means(df, tok, show_progress=True)
stats.to_parquet("token_std_means.parquet") # index name "token", columns mean/std
```
Drop that file next to a checkpoint and `embed`, `finetune`, and `benchmark` will pick it up.
## Embeddings without the CLI
```python
import torch
from transformers import AutoModel
from waypoint_bio.dataset import load_waypoint_dataframe, try_load_token_std_means
from waypoint_bio.embed import tokenize_for_embedding
from waypoint_bio.models import _pool
from waypoint_bio.tokenizer import load_tokenizer
model_id = "outpost-bio/Waypoint-45m"
df = load_waypoint_dataframe("dataset.parquet")
tok = load_tokenizer(model_id)
model = AutoModel.from_pretrained(model_id).eval()
samples = tokenize_for_embedding(df, tok, max_length=512,
token_std_means=try_load_token_std_means(model_id))
input_ids = torch.stack([s["input_ids"] for s in samples])
attn = torch.stack([s["attention_mask"] for s in samples])
with torch.no_grad():
hidden = model(input_ids=input_ids, attention_mask=attn).last_hidden_state
emb = _pool(hidden, attn, "last_token") # [n_samples, hidden_size]
```
`tokenize_for_embedding` preserves one output row per input row even when a row has no
in-vocabulary taxa, so `emb` stays aligned with `df.index`. Those rows encode as `[BOS][EOS]` and
their embeddings should be discarded, not interpreted.
## Custom heads
`waypoint_bio.models` provides the two heads used by `finetune` and `benchmark`:
```python
from waypoint_bio.models import ClassificationModel, RegressionModel
head = ClassificationModel(
base_model=model,
tokenizer=tok,
label_dims=[3], # one entry per target column
pooling_strategy="last_token",
covariate_dim=0, # width of the one-hot covariate block
class_weights=None, # list[torch.Tensor], one per target
)
```
Both pool `last_hidden_state`, concatenate the one-hot covariate block if present, and apply one
`nn.Linear` per target column. Multi-target classification masks label `-100` per target, so targets
with missing values in some rows are handled without dropping the row.
Pooling strategies: `mean` (mask-weighted average), `last_token` (last non-padding position — the
default and what the checkpoints were tuned for), `first_token` / `cls_token` (position 0, the BOS
token; weak in a causal LM).
## Loading Atlas and Compass
```python
from datasets import load_dataset
atlas = load_dataset("outpost-bio/Atlas", split="pretrain") # 485,377 rows
atlas_bench = load_dataset("outpost-bio/Atlas", split="benchmark") # 53,931 held out
compass = load_dataset("outpost-bio/Compass", "mastrorilli")
compass["train"], compass["validation"], compass["test"]
```
Atlas is ~5.6 GB. Stream it if you are only inspecting:
```python
atlas = load_dataset("outpost-bio/Atlas", split="pretrain", streaming=True)
first = next(iter(atlas))
```
Atlas rows carry `Taxa`, `Relative Abundances`, `Run Accession`, `Data Type`, `Sequencing Method`,
`Pipeline Version`, `Study Accession`. Filtering by `Data Type` or `Sequencing Method` before
pretraining is a reasonable way to build a modality-specific model; filtering by `Study Accession` is
how you would hold out whole studies.
Provenance: scraped from MGnify across pipeline versions v1.0v5.0 and four modalities (16S amplicon,
whole-genome shotgun, metagenomic assembly, and metatranscriptomic), then filtered to a minimum
relative abundance of 1e-4 and a minimum of 10 taxa per sample. The pretrain/benchmark split is
random with `seed=42` — it is *not* a study-level holdout, so the Atlas `benchmark` split shares
studies with `pretrain`.
## Fine-tuning programmatically
There is no stable public function for the whole loop; `waypoint_bio.finetune` is written as a CLI
module. Two workable options:
1. Call the CLI with `subprocess` and read `finetune_results.json` — what the upstream webinar
notebooks do.
2. Assemble it yourself from `MicrobiomeBenchmarkDataset` + `ClassificationModel`/`RegressionModel`
and a `transformers.Trainer`, mirroring `benchmark.py`. Reuse `waypoint_bio.scoring.score_task`
and `predictions_to_arrays` so your metrics match the published definitions.
```python
import json, subprocess
subprocess.run([
"waypoint", "finetune",
"--model", "outpost-bio/Waypoint-45m",
"--data", "dataset.parquet",
"--output_dir", "outputs/ft",
"--task_type", "classification",
"--target", "Group",
], check=True)
results = json.loads(open("outputs/ft/finetune_results.json").read())
print(results["test_score"], results["test_metrics"])
```
The upstream repo's `examples/webinar/` carries two worked notebooks — a regression walkthrough on
Compass task 6 and a classification walkthrough on task 8 that also plots PCA / t-SNE projections of
the embeddings against a logistic-regression baseline. Shared helpers live in `webinar_utils.py`.