Compare commits

..
Author SHA1 Message Date
promptadmin 5a8f84424a [upstream-sync] docs/skills.md from K-Dense-AI/scientific-agent-skills@0807ddbc [catalogue] 2026-06-30 01:34:48 +00:00
promptadmin f8762e8a4d [upstream-sync] docs/open-source-sponsors.md from K-Dense-AI/scientific-agent-skills@0807ddbc [catalogue] 2026-06-30 01:34:26 +00:00
promptadmin 782566b017 [upstream-sync] README.md from K-Dense-AI/scientific-agent-skills@0807ddbc [catalogue] 2026-06-30 01:34:05 +00:00
promptadmin 8a4cab56d2 [upstream-sync] skills/tamarind/references/workflows.md from K-Dense-AI/scientific-agent-skills@0807ddbc [prompt] 2026-06-30 01:33:43 +00:00
promptadmin b6e0b442af [upstream-sync] skills/tamarind/references/tool_catalog.md from K-Dense-AI/scientific-agent-skills@0807ddbc [prompt] 2026-06-30 01:33:24 +00:00
promptadmin e0e50b11d5 [upstream-sync] skills/tamarind/references/examples.md from K-Dense-AI/scientific-agent-skills@0807ddbc [prompt] 2026-06-30 01:33:07 +00:00
promptadmin 639afa462e [upstream-sync] skills/tamarind/references/api_reference.md from K-Dense-AI/scientific-agent-skills@0807ddbc [prompt] 2026-06-30 01:32:52 +00:00
promptadmin fbc1c1ca8b [upstream-sync] skills/tamarind/SKILL.md from K-Dense-AI/scientific-agent-skills@0807ddbc [prompt] 2026-06-30 01:32:38 +00:00
promptadmin 4dd1a7c253 [upstream-sync] skills/onekgpd/references/onekgpd_commands.md from K-Dense-AI/scientific-agent-skills@0807ddbc [unknown] 2026-06-30 01:32:24 +00:00
promptadmin 584aef6aeb [upstream-sync] skills/onekgpd/references/annotation_vocabularies.md from K-Dense-AI/scientific-agent-skills@0807ddbc [unknown] 2026-06-30 01:32:11 +00:00
promptadmin a5d0e304ae [upstream-sync] skills/onekgpd/SKILL.md from K-Dense-AI/scientific-agent-skills@0807ddbc [unknown] 2026-06-30 01:32:00 +00:00
11 changed files with 994 additions and 30 deletions
@@ -2,9 +2,9 @@
title: "Scientific Agent Skills"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9c9bd2e9/README.md
upstream_sha: 9c9bd2e9
imported_at: 2026-06-26
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/README.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -14,8 +14,8 @@ validated: false
# Scientific Agent Skills
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE.md)
[![Version](https://img.shields.io/badge/Version-2.50.0-blue.svg)](pyproject.toml)
[![Skills](https://img.shields.io/badge/Skills-147-brightgreen.svg)](#-whats-included)
[![Version](https://img.shields.io/badge/Version-2.53.0-blue.svg)](pyproject.toml)
[![Skills](https://img.shields.io/badge/Skills-148-brightgreen.svg)](#-whats-included)
[![Databases](https://img.shields.io/badge/Databases-100%2B-orange.svg)](#-whats-included)
[![Agent Skills](https://img.shields.io/badge/Standard-Agent_Skills-blueviolet.svg)](https://agentskills.io/)
[![Security Scan](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/security-scan.yml/badge.svg)](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/security-scan.yml)
@@ -30,11 +30,11 @@ validated: false
> **🔔 Claude Scientific Skills is now Scientific Agent Skills.** Same skills, broader compatibility — now works with any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, not just Claude.
> **New: [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok)** — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 147 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via [Modal](https://modal.com/) for heavy workloads. [Get started here.](https://github.com/K-Dense-AI/k-dense-byok)
> **New: [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok)** — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 148 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via [Modal](https://modal.com/) for heavy workloads. [Get started here.](https://github.com/K-Dense-AI/k-dense-byok)
> **Stay up to date:** Follow K-Dense on [X](https://x.com/k_dense_ai), [LinkedIn](https://www.linkedin.com/company/k-dense-inc), and [YouTube](https://www.youtube.com/@K-Dense-Inc) for new skills, release announcements, walkthroughs, research workflow demos, and examples you can use with your own AI agent.
A comprehensive collection of **147 ready-to-use scientific and research skills** (covering cancer genomics, drug-target binding, molecular dynamics, RNA velocity, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
A comprehensive collection of **148 ready-to-use scientific and research skills** (covering cancer genomics, drug-target binding, molecular dynamics, RNA velocity, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
> ⭐ **Help make AI for science easier to discover:** If Scientific Agent Skills saves you time, teaches your agent a workflow, or helps your lab move faster, please [star this repository](https://github.com/K-Dense-AI/scientific-agent-skills). A star is a public signal that these open, reusable research skills are worth maintaining: it helps scientists, engineers, and open-source contributors find the project, shows which agent-skill standards are gaining real adoption, and gives us a clear reason to keep expanding the collection for the community.
@@ -68,12 +68,12 @@ These skills enable your AI agent to seamlessly work with specialized scientific
## 📦 What's Included
This repository provides **147 scientific and research skills** organized into the following categories:
This repository provides **148 scientific and research skills** organized into the following categories:
- **100+ Scientific & Financial Databases** - A unified database-lookup skill provides deterministic, provenance-rich access to 78 public databases (PubChem, ChEMBL, UniProt, COSMIC, ClinicalTrials.gov, FRED, USPTO, and more), plus dedicated skills for DepMap, Imaging Data Commons, PrimeKG, U.S. Treasury Fiscal Data, and Hugging Science (curated catalog of scientific datasets, models, and demos across 17 scientific domains on Hugging Face). Multi-database packages like BioServices (~40 bioinformatics services), BioPython (38 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage
- **70+ Optimized Python Package Skills** - Explicitly defined skills for RDKit, Scanpy, PyTorch Lightning, scikit-learn, BioPython, pyzotero, BioServices, PennyLane, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), scVelo, TimesFM, and others — with curated documentation, examples, and best practices. Note: the agent can write code using *any* Python package, not just these; these skills simply provide stronger, more reliable performance for the packages listed
- **9 Scientific Integration Skills** - Explicitly defined skills for Benchling, DNAnexus, LatchBio, OMERO, Protocols.io, Open Notebook, Ginkgo Cloud Lab, LabArchives, and Opentrons. Again, the agent is not limited to these — any API or platform reachable from Python is fair game; these skills are the optimized, pre-documented paths
- **30+ Analysis & Communication Tools** - Literature review, scientific writing, peer review, document processing, Paperzilla, PACSOMATIC, Exa Search, posters, slides, schematics, infographics, Mermaid diagrams, and more
- **30+ Analysis & Communication Tools** - Literature review, scientific writing, peer review, document processing, Paperzilla, Exa Search, posters, slides, schematics, infographics, Mermaid diagrams, and more
- **10+ Research & Clinical Tools** - Hypothesis generation, grant writing, clinical decision support, treatment plans, BIDS, regulatory compliance, scenario analysis, and workflow-derived skill drafting with Autoskill
Each skill includes:
@@ -113,7 +113,7 @@ Each skill includes:
- **Multi-Step Workflows** - Execute complex pipelines with a single prompt
### 🎯 **Comprehensive Coverage**
- **147 Skills** - Extensive coverage across all major scientific domains
- **148 Skills** - Extensive coverage across all major scientific domains
- **100+ Databases** - Unified access to 78+ databases via database-lookup, plus dedicated data access skills and multi-database packages like BioServices, BioPython, and gget
- **70+ Optimized Python Package Skills** - RDKit, Scanpy, PyTorch Lightning, scikit-learn, BioServices, PennyLane, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), scVelo, TimesFM, and others (the agent can use any Python package; these are the pre-documented, higher-performing paths)
@@ -198,7 +198,7 @@ git clone https://github.com/K-Dense-AI/scientific-agent-skills.git .agents/skil
hermes skills tap add K-Dense-AI/scientific-agent-skills
```
These skills stay portable across all of them: `metadata` is single-line JSON (so OpenClaw's line-based reader parses it), credentialed skills declare a top-level `required_environment_variables` field (so Hermes prompts for keys), and unknown fields are ignored everywhere else. Because 147 skills add up to a lot of standing context, consider installing a topical subset rather than the whole collection.
These skills stay portable across all of them: `metadata` is single-line JSON (so OpenClaw's line-based reader parses it), credentialed skills declare a top-level `required_environment_variables` field (so Hermes prompts for keys), and unknown fields are ignored everywhere else. Because 148 skills add up to a lot of standing context, consider installing a topical subset rather than the whole collection.
> **NemoClaw note:** NemoClaw runs agents inside NVIDIA OpenShell with default-deny outbound networking. Skills are discovered and loaded normally, but any skill that needs the network — package installs via `uv`, or API calls (Exa, Parallel, Benchling, NCBI, Materials Project, …) — only works once the operator pre-approves the relevant domains in the OpenShell TUI.
@@ -423,7 +423,7 @@ networks, and search GEO for similar patterns.
## 📚 Available Skills
This repository contains **147 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
This repository contains **148 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
### Skill Categories
@@ -511,19 +511,20 @@ This repository contains **147 scientific and research skills** organized across
- Multi-omics: HypoGeniC
- Data management: LaminDB
#### 🧬 **Protein Engineering & Design** (3 skills)
#### 🧬 **Protein Engineering & Design** (4 skills)
- Protein language models: ESM
- Glycoengineering: Glycoengineering (N/O-glycosylation prediction, therapeutic antibody optimization)
- Cloud laboratory platform: Adaptyv (automated protein testing and validation)
- Cloud structure & design platform: Tamarind (managed-GPU access to AlphaFold, Boltz, Chai, ESMFold, RFdiffusion, ProteinMPNN, BoltzGen, antibody/nanobody design, DiffDock/Vina docking, binding affinity, and MSA generation via REST API or MCP)
#### 📚 **Scientific Communication** (27 skills)
#### 📚 **Scientific Communication** (26 skills)
- Literature: Paper Lookup (PubMed, PMC, bioRxiv, medRxiv, arXiv, OpenAlex, Crossref, Semantic Scholar, CORE, Unpaywall), Literature Review, Paperzilla
- Advanced paper search: BGPT Paper Search (25+ structured fields per paper — methods, results, sample sizes, quality scores — from full text, not just abstracts)
- Web search: Parallel Web, Exa Search, and Research Lookup
- Research notebooks: Open Notebook (self-hosted NotebookLM alternative — PDFs, videos, audio, web pages; 16+ AI providers; multi-speaker podcast generation)
- Writing: Scientific Writing, Peer Review
- Document processing: LiteParse, PDF, DOCX, PPTX, XLSX, and MarkItDown
- Publishing and paper workflows: Venue Templates, PACSOMATIC
- Publishing and paper workflows: Venue Templates
- Presentations: Scientific Slides, LaTeX Posters, PPTX Posters
- Diagrams: Scientific Schematics, Markdown & Mermaid Writing
- Infographics: Infographics (10 types, 8 styles, colorblind-safe palettes)
@@ -539,10 +540,11 @@ This repository contains **147 scientific and research skills** organized across
- Fiscal data: U.S. Treasury Fiscal Data (national debt, Treasury statements, auctions, exchange rates)
- Scientific ML resource catalog: Hugging Science (curated index of datasets, models, blog posts, and interactive Spaces across 17 scientific domains — astronomy, biology, chemistry, climate, genomics, materials science, medicine, physics, scientific reasoning, and more — with usage patterns for `datasets`, `transformers`, and `gradio_client`)
#### 🔧 **Infrastructure & Platforms** (9 skills)
#### 🔧 **Infrastructure & Platforms** (11 skills)
- Cloud compute: Modal
- GPU acceleration: Optimize for GPU (CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, RAFT)
- Genomics platforms: DNAnexus, LatchBio
- Workflow engines: Nextflow (build/run/debug Nextflow & nf-core pipelines — DSL2 modules, executors/containers, HPC/cloud scaling) and pacsomatic (operator toolkit for the nf-core/pacsomatic tumor-normal somatic variant-calling workflow)
- Microscopy: OMERO
- Automation: Opentrons
- Resource detection: Get Available Resources
@@ -750,7 +752,7 @@ Recommended practice:
title = {Scientific Agent Skills: A Comprehensive Collection of Scientific Tools for AI Agents},
year = {2026},
url = {https://github.com/K-Dense-AI/scientific-agent-skills},
note = {147 skills covering databases, packages, integrations, and analysis tools}
note = {148 skills covering databases, packages, integrations, and analysis tools}
}
```
@@ -2,9 +2,9 @@
title: "Support the Open Source Projects We Depend On"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9c9bd2e9/docs/open-source-sponsors.md
upstream_sha: 9c9bd2e9
imported_at: 2026-06-26
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/docs/open-source-sponsors.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -13,7 +13,7 @@ validated: false
# Support the Open Source Projects We Depend On
Scientific Agent Skills is built on the shoulders of giants. The 146 skills in this repository leverage dozens of incredible open source projects created and maintained by dedicated developers and research communities around the world.
Scientific Agent Skills is built on the shoulders of giants. The 148 skills in this repository leverage dozens of incredible open source projects created and maintained by dedicated developers and research communities around the world.
**If you find value in these skills, please consider supporting the underlying open source projects that make them possible.**
@@ -2,9 +2,9 @@
title: "Scientific Skills"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9c9bd2e9/docs/skills.md
upstream_sha: 9c9bd2e9
imported_at: 2026-06-26
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/docs/skills.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -40,6 +40,8 @@ validated: false
### Workflow Platforms & Cloud Execution
- **LatchBio Integration** - Integration with the Latch platform for building, deploying, and executing bioinformatics workflows. Provides comprehensive support for creating serverless bioinformatics pipelines using Python decorators, deploying Nextflow/Snakemake pipelines, managing cloud data (LatchFile, LatchDir) and structured Registry (Projects, Tables, Records), configuring computational resources (CPU, GPU, memory, storage), and using pre-built Latch Verified workflows (RNA-seq, AlphaFold, DESeq2, single-cell analysis, CRISPR editing). Enables automatic containerization, UI generation, workflow versioning, and execution on scalable cloud infrastructure with comprehensive data management
- **Nextflow** - Build, run, and debug Nextflow data pipelines and nf-core workflows end to end. Covers writing and testing DSL2 modules/subworkflows (processes, channels, operators, nf-test), running community pipelines (nf-core/rnaseq, nf-core/sarek), authoring samplesheets and `nextflow.config`, configuring executors and containers (Docker, Singularity/Apptainer, Conda, Wave), scaling to HPC/SLURM or cloud (AWS Batch, Google Batch, Azure, Kubernetes), and debugging failed or `-resume` runs. Use for any reproducible scientific/bioinformatics workflow work and for authoring nf-core-compliant pipelines, modules, configs, and linting
- **pacsomatic** - Operator toolkit for running the nf-core/pacsomatic matched tumor-normal (somatic variant calling) workflow from BAM inputs. Validates run inputs, generates pacsomatic-compliant samplesheets, prepares reproducible Nextflow launch artifacts, runs locally or submits to schedulers (LSF/Slurm/PBS/SGE), and triages execution and startup failures. Use to prepare launch commands/scripts, perform dry-run checks, or troubleshoot pipeline and scheduler submission errors
### Microscopy & Bio-image Data
- **OMERO Integration** - Toolkit for interacting with OMERO microscopy data management systems using Python. Provides comprehensive access to microscopy images stored in OMERO servers, including dataset and screening data retrieval, pixel data analysis, annotation and metadata management, regions of interest (ROIs) creation and analysis, batch processing, OMERO.scripts development, and OMERO.tables for structured data storage. Essential for researchers working with high-content screening data, multi-dimensional microscopy datasets, or collaborative image repositories
@@ -115,6 +117,7 @@ validated: false
- **ESM (Evolutionary Scale Modeling)** - Protein language models from EvolutionaryScale/Biohub for protein design, structure prediction, and representation learning. Covers current `esm` SDK workflows for local ESM3/ESMC open models, Forge-hosted ESM3 and ESMC inference with `ESM_API_KEY`, Biohub-hosted ESMC embeddings, and ESMFold2 all-atom structure prediction. Use cases: novel protein design, sequence/structure co-design, protein embeddings, function annotation, variant generation, and directed evolution workflows
- **Glycoengineering** - Analyze and engineer protein glycosylation. Scan sequences for N-glycosylation sequons (N-X-S/T), predict O-glycosylation hotspots, and access curated glycoengineering tools (NetOGlyc, GlycoShield, GlycoWorkbench). Use for glycoprotein engineering, therapeutic antibody optimization, and vaccine design
- **Molecular Dynamics** - Run and analyze molecular dynamics simulations with OpenMM and MDAnalysis. Set up protein/small molecule systems, define force fields, run energy minimization and production MD, analyze trajectories (RMSD, RMSF, contact maps, free energy surfaces). Use for structural biology, drug binding studies, and biophysics research
- **Tamarind** - Run a large catalog of open-source molecular design and structural biology tools on the Tamarind Bio managed-GPU cloud via its REST API (`x-api-key` header) or MCP server — no local GPUs required. Covers structure prediction (AlphaFold, Boltz-2, Chai-1, ESMFold2), protein/binder/de novo design (RFdiffusion, ProteinMPNN/LigandMPNN, BoltzGen, BindCraft), antibody and nanobody design, humanization and developability, protein-ligand docking (DiffDock, AutoDock Vina) and binding-affinity prediction, MSA generation, and molecular dynamics — all through one uniform job API with single and batch submission and tool chaining (design → fold → score). Discovers tools and schemas live (`GET /tools`, MCP `getAvailableTools`/`getJobSchema`) and reads the key from `TAMARIND_API_KEY`. There is no official Python SDK — use plain `requests` or the MCP server. Use cases: cloud structure prediction and protein design without provisioning hardware, high-throughput batch characterization of sequences/designs, and chaining design-fold-score pipelines
### Machine Learning & Deep Learning
- **aeon** - Comprehensive scikit-learn compatible Python toolkit for time series machine learning providing state-of-the-art algorithms across 7 domains: classification (13 algorithm categories including ROCKET variants, deep learning with InceptionTime/ResNet/FCN, distance-based with DTW/ERP/LCSS, shapelet-based, dictionary methods like BOSS/WEASEL, and hybrid ensembles HIVECOTE), regression (9 categories mirroring classification approaches), clustering (k-means/k-medoids with temporal distances, deep learning autoencoders, spectral methods), forecasting (ARIMA, ETS, Theta, Threshold Autoregressive, TCN, DeepAR), anomaly detection (STOMP/MERLIN matrix profile, clustering-based CBLOF/KMeans, isolation methods, copula-based COPOD), segmentation (ClaSP, FLUSS, HMM, binary segmentation), and similarity search (MASS algorithm, STOMP motif discovery, approximate nearest neighbors). Includes 40+ distance metrics (elastic: DTW/DDTW/WDTW/Shape-DTW, edit-based: ERP/EDR/LCSS/TWE/MSM, lock-step: Euclidean/Manhattan), extensive transformations (ROCKET/MiniRocket/MultiRocket for features, Catch22/TSFresh for statistics, SAX/PAA for symbolic representation, shapelet transforms, wavelets, matrix profile), 20+ deep learning architectures (FCN, ResNet, InceptionTime, TCN, autoencoders with attention mechanisms), comprehensive benchmarking tools (UCR/UEA archives with 100+ datasets, published results repository, statistical testing), and performance-optimized implementations using numba. Features progressive model complexity from fast baselines (MiniRocket: <1 second training, 0.95+ accuracy on many benchmarks) to state-of-the-art ensembles (HIVECOTE V2), GPU acceleration support, and extensive visualization utilities. Use cases: physiological signal classification (ECG, EEG), industrial sensor monitoring, financial forecasting, change point detection, pattern discovery, activity recognition from wearables, predictive maintenance, climate time series analysis, and any sequential data requiring specialized temporal modeling beyond standard ML
@@ -153,7 +156,6 @@ validated: false
- **GeoPandas** - Python library extending pandas for working with geospatial vector data including shapefiles, GeoJSON, and GeoPackage files. Provides GeoDataFrame and GeoSeries data structures combining geometric data with tabular attributes for spatial analysis. Key features include: reading/writing spatial file formats (Shapefile, GeoJSON, GeoPackage, PostGIS, Parquet) with Arrow acceleration for 2-4x faster I/O, geometric operations (buffer, simplify, centroid, convex hull, affine transformations) through Shapely integration, spatial analysis (spatial joins with predicates like intersects/contains/within, nearest neighbor joins, overlay operations for union/intersection/difference, dissolve for aggregation, clipping), coordinate reference system (CRS) management (setting CRS, reprojecting between coordinate systems, UTM estimation), and visualization (static choropleth maps with matplotlib, interactive maps with folium, multi-layer mapping, classification schemes with mapclassify). Supports spatial indexing for performance, filtering during read operations (bbox, mask, SQL WHERE), and integration with cartopy for cartographic projections. Use cases: spatial data manipulation, buffer analysis, spatial joins between datasets, dissolving boundaries, calculating areas/distances in projected CRS, reprojecting coordinate systems, creating choropleth maps, converting between spatial file formats, PostGIS database integration, and geospatial data analysis workflows
- **Matplotlib** - Comprehensive Python plotting library for creating publication-quality static, animated, and interactive visualizations. Provides extensive customization options for creating figures, subplots, axes, and annotations. Key features include: support for multiple plot types (line, scatter, bar, histogram, contour, 3D, and many more), extensive customization (colors, fonts, styles, layouts), multiple backends (PNG, PDF, SVG, interactive backends), LaTeX integration for mathematical notation, and integration with NumPy and pandas. Includes specialized modules (pyplot for MATLAB-like interface, artist layer for fine-grained control, backend layer for rendering). Supports complex multi-panel figures, color maps, legends, and annotations. Use cases: scientific figure creation, data visualization, exploratory data analysis, publication graphics, and any application requiring high-quality plots
- **NetworkX** - Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs. Supports four graph types (Graph, DiGraph, MultiGraph, MultiDiGraph) with nodes as any hashable objects and rich edge attributes. Provides 100+ algorithms including shortest paths (Dijkstra, Bellman-Ford, A*), centrality measures (degree, betweenness, closeness, eigenvector, PageRank), clustering (coefficients, triangles, transitivity), community detection (modularity-based, label propagation, Girvan-Newman), connectivity analysis (components, cuts, flows), tree algorithms (MST, spanning trees), matching, graph coloring, isomorphism, and traversal (DFS, BFS). Includes 50+ graph generators for classic (complete, cycle, wheel), random (Erdős-Rényi, Barabási-Albert, Watts-Strogatz, stochastic block model), lattice (grid, hexagonal, hypercube), and specialized networks. Supports I/O across formats (edge lists, GraphML, GML, JSON, Pajek, GEXF, DOT) with Pandas/NumPy/SciPy integration. Visualization capabilities include 8+ layout algorithms (spring/force-directed, circular, spectral, Kamada-Kawai), customizable node/edge appearance, interactive visualizations with Plotly/PyVis, and publication-quality figure generation. Use cases: social network analysis, biological networks (protein-protein interactions, gene regulatory networks, metabolic pathways), transportation systems, citation networks, knowledge graphs, web structure analysis, infrastructure networks, and any domain involving pairwise relationships requiring structural analysis or graph-based modeling
- **Plotly** - Interactive scientific and statistical data visualization library for Python with 40+ chart types. Provides both high-level API (Plotly Express) for quick visualizations and low-level API (graph objects) for fine-grained control. Key features include: comprehensive chart types (scatter, line, bar, histogram, box, violin, heatmap, contour, 3D plots, geographic maps, financial charts, statistical distributions, hierarchical charts), interactive features (hover tooltips, pan/zoom, legend toggling, animations, rangesliders, buttons/dropdowns), publication-quality output (static images in PNG/PDF/SVG via Kaleido, interactive HTML with embeddable figures), extensive customization (templates, themes, color scales, fonts, layouts, annotations, shapes), subplot support (multi-plot figures with shared axes), and Dash integration for building analytical web applications. Plotly Express offers one-line creation of complex visualizations with automatic color encoding, faceting, and trendlines. Graph objects provide precise control for specialized visualizations (candlestick charts, 3D surfaces, sankey diagrams, gauge charts). Supports pandas DataFrames, NumPy arrays, and various data formats. Use cases: scientific data visualization, statistical analysis, financial charting, interactive dashboards, publication figures, exploratory data analysis, and any application requiring interactive or publication-quality visualizations
- **Polars** - High-performance DataFrame library written in Rust with Python bindings for fast data manipulation, ETL, analytics, and pandas migration. Provides expression-based transformations, lazy query optimization, automatic parallel execution, streaming out-of-core processing, Arrow interoperability, and optional GPU execution. Supports common data sources and formats including CSV, Parquet, JSON/NDJSON, Excel, Arrow IPC, cloud object storage, databases, and BigQuery. Use cases: large-scale data processing, memory-conscious analytical pipelines, feature engineering, and high-performance DataFrame workflows
- **Seaborn** - Statistical data visualization with dataset-oriented interface, automatic confidence intervals, publication-quality themes, colorblind-safe palettes, and comprehensive support for exploratory analysis, distribution comparisons, correlation matrices, regression plots, and multi-panel figures
- **Vaex** - High-performance Python library for lazy, out-of-core DataFrames to process and visualize tabular datasets larger than available RAM. Processes over a billion rows per second through memory-mapped files (HDF5, Apache Arrow), lazy evaluation, and virtual columns (zero memory overhead). Provides instant file opening, efficient aggregations across billions of rows, interactive visualizations without sampling, machine learning pipelines with transformers (scalers, encoders, PCA), and seamless integration with pandas/NumPy/Arrow. Includes comprehensive ML framework (vaex.ml) with feature scaling, categorical encoding, dimensionality reduction, and integration with scikit-learn/XGBoost/LightGBM/CatBoost. Supports distributed computing via Dask, asynchronous operations, and state management for production deployment. Use cases: processing gigabyte to terabyte datasets, fast statistical aggregations on massive data, visualizing billion-row datasets, ML pipelines on big data, converting between data formats, and working with astronomical, financial, or scientific large-scale datasets
@@ -163,7 +165,6 @@ validated: false
- **Phylogenetics** - Build and analyze phylogenetic trees using MAFFT (multiple alignment), IQ-TREE 2 (maximum likelihood), and FastTree (fast NJ/ML). Visualize with ETE3 or FigTree. Use for evolutionary analysis, microbial genomics, viral phylodynamics, protein family analysis, and molecular clock studies
### Multi-omics & AI Agent Frameworks
- **Denario** - Multiagent AI system for scientific research assistance that automates complete research workflows from data analysis through publication. Built on AG2 and LangGraph frameworks, orchestrates specialized agents for hypothesis generation, methodology development, computational analysis, and LaTeX paper writing. Supports multiple LLM providers (Google Vertex AI, OpenAI) with flexible pipeline stages allowing manual or automated inputs. Key features include: end-to-end research automation (data description → idea generation → methodology → results → paper), journal-specific formatting (APS and others), GUI interface via Streamlit, Docker deployment with LaTeX environment, reproducible research with version-controlled outputs, literature search integration, and integration with scientific Python stack (pandas, sklearn, scipy). Provides both programmatic Python API and web-based interface. Use cases: automated hypothesis generation from datasets, research methodology development, computational experiment execution with visualization, publication-ready manuscript generation, time-series analysis research, machine learning experiment automation, and accelerating the complete scientific research lifecycle from ideation to publication
- **Pi Agent** - Build with and use Pi, the minimal terminal coding harness. Covers installing and configuring Pi, authenticating providers, adding custom models and providers, creating Pi skills/extensions/packages/themes/prompt templates, embedding Pi through the Node/TypeScript SDK, integrating over RPC or JSON event streams, parsing session JSONL files, and building custom TUI components. Use cases: building custom agent UIs, exposing Pi from another process or language, wiring private model gateways, packaging reusable Pi workflows, and helping agents operate Pi as a platform rather than only a CLI
- **HypoGeniC** - Automated hypothesis generation and testing using large language models to accelerate scientific discovery. Provides three frameworks: HypoGeniC (data-driven hypothesis generation from observational data), HypoRefine (synergistic approach combining literature insights with empirical patterns through an agentic system), and Union methods (mechanistic combination of literature and data-driven hypotheses). Features iterative refinement that improves hypotheses by learning from challenging examples, Redis caching for API cost reduction, and customizable YAML-based prompt templates. Includes command-line tools for generation (hypogenic_generation) and testing (hypogenic_inference). Research applications have demonstrated 14.19% accuracy improvement in AI-content detection and 7.44% in deception detection. Use cases: deception detection in reviews, AI-generated content identification, mental stress detection, exploratory research without existing literature, hypothesis-driven analysis in novel domains, and systematic exploration of competing explanations
@@ -178,7 +179,6 @@ validated: false
- **Infographics** - Create professional infographics using Nano Banana Pro AI with smart iterative refinement. Uses Gemini 3 Pro for quality review. Integrates research-lookup and web search for accurate data. Supports 10 infographic types, 8 industry styles, and colorblind-safe palettes
- **LaTeX Posters** - Create professional research posters in LaTeX using beamerposter, tikzposter, or baposter. Support for conference presentations, academic posters, and scientific communication with layout design, color schemes, multi-column formats, figure integration, and poster-specific best practices. Features compliance with conference size requirements (A0, A1, 36×48"), complex multi-column layouts, and integration of figures, tables, equations, and citations. Use cases: conference poster sessions, thesis defenses, symposia presentations, and research group templates
- **Market Research Reports** - Generate comprehensive market research reports (50+ pages) in the style of top consulting firms (McKinsey, BCG, Gartner). Features professional LaTeX formatting, extensive visual generation, deep integration with research-lookup for data gathering, and multi-framework strategic analysis including Porter's Five Forces, PESTLE, SWOT, TAM/SAM/SOM, and BCG Matrix. Use cases: investment decisions, strategic planning, competitive landscape analysis, market sizing, and market entry evaluation
- **Paper-2-Web** - Autonomous pipeline for transforming academic papers into multiple promotional formats using the Paper2All system. Converts LaTeX or PDF papers into: (1) Paper2Web - interactive, layout-aware academic homepages with responsive design, interactive figures, and mobile support; (2) Paper2Video - professional presentation videos with slides, narration, cursor movements, and optional talking-head generation using Hallo2; (3) Paper2Poster - print-ready conference posters with custom dimensions, professional layouts, and institution branding. Supports GPT-4/GPT-4.1 models, batch processing, QR code generation, multi-language content, and quality assessment metrics. Use cases: conference materials, video abstracts, preprint enhancement, research promotion, poster sessions, and academic website creation
- **PPTX Posters** - Create professional research posters using PowerPoint/HTML formats for researchers who prefer WYSIWYG tools over LaTeX. Features design principles, layout templates, quality checklists, and export guidance for poster sessions. Use cases: conference posters when LaTeX is not preferred, quick poster creation, and collaborative poster design
- **Scientific Schematics** - Create publication-quality scientific diagrams using Nano Banana Pro AI with smart iterative refinement. Uses Gemini 3 Pro for quality review with document-type-specific thresholds (journal: 8.5/10, conference: 8.0/10, poster: 7.0/10). Specializes in neural network architectures, system diagrams, flowcharts, biological pathways, and complex scientific visualizations. Features natural language input, automatic quality assessment, and publication-ready output. Use cases: creating figures for papers, generating workflow diagrams, visualizing experimental designs, and producing graphical abstracts
- **Scientific Slides** - Build slide decks and presentations for research talks using PowerPoint and LaTeX Beamer. Features slide structure, design templates, timing guidance, and visual validation. Emphasizes visual engagement with minimal text, research-backed content with proper citations, and story-driven narrative. Use cases: conference presentations, academic seminars, thesis defenses, grant pitches, and professional talks
@@ -197,6 +197,7 @@ validated: false
- **PyLabRobot** - Hardware-agnostic, pure Python SDK for automated and autonomous laboratories. Provides unified interface for controlling liquid handling robots (Hamilton STAR/STARlet, Opentrons OT-2, Tecan EVO), plate readers (BMG CLARIOstar), heater shakers, incubators, centrifuges, pumps, and scales. Key features include: modular resource management system for plates, tips, and containers with hierarchical deck layouts and JSON serialization; comprehensive liquid handling operations (aspirate, dispense, transfer, serial dilutions, plate replication) with automatic tip and volume tracking; backend abstraction enabling hardware-agnostic protocols that work across different robots; ChatterboxBackend for protocol simulation and testing without hardware; browser-based visualizer for real-time 3D deck state visualization; cross-platform support (Windows, macOS, Linux, Raspberry Pi); and integration capabilities for multi-device workflows combining liquid handlers, analytical equipment, and material handling devices. Use cases: automated sample preparation, high-throughput screening, serial dilution protocols, plate reading workflows, laboratory protocol development and validation, robotic liquid handling automation, and reproducible laboratory automation with state tracking and persistence
### Tool Discovery & Computational Resources
- **Autoskill** - Observe the user's screen via the local screenpipe daemon, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon running locally on port 3030; all detection runs locally and only redacted cluster summaries reach the LLM
- **Get Available Resources** - Detect available computational resources and generate strategic recommendations for scientific computing tasks at the start of any computationally intensive scientific task. Automatically identifies CPU capabilities, GPU availability (NVIDIA CUDA, AMD ROCm, Apple Silicon Metal), memory constraints, and disk space. Creates JSON file with resource information and recommendations for parallel processing (joblib, multiprocessing), out-of-core computing (Dask, Zarr), GPU acceleration (PyTorch, JAX), or memory-efficient strategies. Use cases: determining optimal computational approaches before data analysis, model training, or large file operations
### Research Methodology & Proposal Writing
@@ -2,7 +2,7 @@
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/onekgpd/SKILL.md
upstream_sha: 0807ddbc
imported_at: 2026-06-29
imported_at: 2026-06-30
prompt_class: unknown
upstream_changes: accepted
name: onekgpd
@@ -4,7 +4,7 @@ task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/onekgpd/references/annotation_vocabularies.md
upstream_sha: 0807ddbc
imported_at: 2026-06-29
imported_at: 2026-06-30
prompt_class: unknown
upstream_changes: accepted
author: upstream
@@ -4,7 +4,7 @@ task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/onekgpd/references/onekgpd_commands.md
upstream_sha: 0807ddbc
imported_at: 2026-06-29
imported_at: 2026-06-30
prompt_class: unknown
upstream_changes: accepted
author: upstream
@@ -0,0 +1,283 @@
---
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/tamarind/SKILL.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: prompt
upstream_changes: accepted
name: tamarind
description: Access a collection of open-source molecular design and structural biology tools on the Tamarind Bio platform, via its REST API or MCP server — no local GPUs required. Tamarind bundles popular open-source models for structure prediction (AlphaFold, Boltz, Chai, ESMFold), protein, binder, and de novo design (RFdiffusion, ProteinMPNN, BoltzGen), antibody and nanobody design and developability, protein-ligand docking (DiffDock, Autodock Vina), binding-affinity prediction, MSA generation, and molecular dynamics. Use when the user mentions Tamarind or tamarind.bio, wants to run any of these open-source tools in the cloud, references app.tamarind.bio/api or the x-api-key header, or needs to submit batches of sequences for structural or biophysical characterization.
license: MIT
compatibility: Requires Python 3.10+, a Tamarind Bio account, and an API key from app.tamarind.bio. Uses the `requests` library against the public REST API (no dedicated Python SDK exists). Network access required. Optional MCP server at mcp.tamarind.bio/mcp for agent hosts.
metadata: {"version": "1.0", "skill-author": "Tamarind Bio", "trigger-keywords": "protein structure prediction, AlphaFold, Boltz, Chai, ESMFold, protein design, binder design, de novo design, antibody design, nanobody, protein-ligand docking, DiffDock, Autodock Vina, binding affinity, MSA generation, inverse folding, ProteinMPNN, RFdiffusion, BoltzGen, cloud GPU biology, structure prediction API, x-api-key, developability, adme, enzyme, peptide, protein language models, molecular design", "openclaw": {"primaryEnv": "TAMARIND_API_KEY", "envVars": [{"name": "TAMARIND_API_KEY", "required": true, "description": "Tamarind Bio API key sent as the x-api-key header."}]}}
required_environment_variables: [{"name": "TAMARIND_API_KEY", "prompt": "Tamarind Bio API key (sent as the x-api-key header).", "required_for": "full functionality"}]
---
# Tamarind Bio
Tamarind Bio is a cloud platform that runs computational biology tools — structure prediction, protein and antibody design, docking, binding-affinity, MSA generation, and molecular dynamics — on managed GPUs. Users submit sequences or structures and get back predicted structures, designs, and biophysical scores, without provisioning their own hardware. It exposes hundreds of tools (AlphaFold, Boltz-2, Chai-1, RFdiffusion, ProteinMPNN, BoltzGen, ESMFold2, DiffDock, Autodock Vina, and many more) through one uniform job API.
**Official docs:** [app.tamarind.bio/api-docs](https://app.tamarind.bio/api-docs) · platform UI at [app.tamarind.bio](https://app.tamarind.bio)
## Canonical sources — fetch these, don't rely on a stale copy
Tamarind publishes live, machine-readable sources. Prefer fetching them at runtime over trusting any hardcoded list — tool names, schemas, and endpoints change frequently:
- **`https://app.tamarind.bio/llms.txt`** — LLM index: links to the spec, API docs, and MCP guide.
- **`https://app.tamarind.bio/openapi.yaml`** — OpenAPI 3.0 spec for the 8 core job endpoints (submit-job/-batch, jobs, result, upload, files, delete-job/-file; auth `ApiKeyAuth`). Fetch it for those exact shapes. Discovery/management endpoints (`/tools`, `/usage-statistics`, pipelines, …) aren't in it — use the MCP/REST discovery tools for those.
- **`https://docs.tamarind.bio/llms.txt`** — documentation index; every page has a `.md` form (e.g. `docs.tamarind.bio/tamarind/batch.md`, `/tamarind/api.md`, `/tamarind/pipelines.md`).
- **Live tool discovery**`GET /tools` (REST) or MCP `getAvailableTools` + `getJobSchema(jobType)` are the source of truth for what tools exist and their parameters.
This skill teaches the surface + the non-obvious behaviors those sources don't spell out (see the reference files). When in doubt about a shape, fetch `openapi.yaml`.
## When to use this skill
Use Tamarind when the user wants to:
- **Predict structure** of a protein, complex, or protein-ligand system (AlphaFold, Boltz-2, Chai-1, ESMFold2, Chai/Boltz cofolding)
- **Design proteins or binders** (RFdiffusion, BoltzGen, BindCraft, ProteinMPNN/LigandMPNN inverse folding)
- **Design or characterize antibodies/nanobodies** (sequence generation, humanization, developability, immunogenicity)
- **Dock small molecules** to a protein (DiffDock, Autodock Vina) or predict **binding affinity**
- **Generate MSAs** for downstream folding
- **Run molecular dynamics** or other biophysical workflows on managed GPUs
- **Batch-screen** many sequences or designs through the same tool
- **Chain tools** into pipelines (e.g. design → fold → score) using the output of one job as the input of the next
This skill is the right fit when the work should run on Tamarind's managed cloud rather than on a local install. For purely local cheminformatics or one-off sequence I/O, use a local library (RDKit, BioPython) instead.
## Access and authentication
1. Sign in at [app.tamarind.bio](https://app.tamarind.bio) and create an API key from the account/API settings.
2. Authenticate every REST request with the `x-api-key` header.
3. **Never hardcode the key.** Read it from the `TAMARIND_API_KEY` environment variable or a `.env` file (use `python-dotenv`). Never commit keys to source control.
**Pricing:** Every user gets **10 free jobs**. For larger usage, contact [[email protected]](mailto:[email protected]) to purchase a subscription.
```bash
export TAMARIND_API_KEY="your_api_key"
# List available tools
curl https://app.tamarind.bio/api/tools \
-H "x-api-key: $TAMARIND_API_KEY"
```
**Base URL:** `https://app.tamarind.bio/api/`
There is **no official Python SDK** — the PyPI package named `tamarind` is an unrelated Neo4j tool. Do not `pip install tamarind`. Write plain `requests` calls against the REST API (the endpoint shapes are in `openapi.yaml`), or use the MCP server for agent hosts.
## Two ways to call Tamarind
### MCP server (best for AI agents)
Tamarind hosts an MCP server at `https://mcp.tamarind.bio/mcp` (API-key auth via the `X-API-Key` header). When your agent host supports MCP, prefer it — the tools mirror the REST API with agent-friendly schemas:
- `listModalities()` / `listTags()` — the live filter vocabulary (molecule type / function) with labels + tool counts; call these to learn valid `modality`/`function` values instead of hardcoding
- `getAvailableTools(modality?, function?, search?, custom?)` — discover tools (`category`/`tag` are deprecated aliases still honored)
- `getJobSchema(jobType)` — exact parameter schema for a tool, plus an `exampleJob` starting payload (validate it before submitting)
- `validateJob(jobName, type, settings)` — dry-run validation before submitting
- `submitJob(jobName, type, settings)` / `submitBatch(batchName, type, settings[], jobNames[])`
- `getJobs(jobName?, batch?, limit?, includeSequences?)` — list/inspect jobs and statuses (the bulky per-job input blob is omitted by default; pass `includeSequences=true` to keep it)
- `getJobLogs(jobName)` — fetch output logs for debugging
- `listJobFiles(jobName)` — list output files (returns `s3Path` for chaining)
- `getResult(jobName, fileName?)` — download results
- `uploadFile(filename)` — presigned upload URL; or `uploadFileContent(filename, content, encoding?)` to send file content through MCP when the host can't reach S3 (sandboxed agents)
Scope note: MCP query tools (`getJobs`, `getResult`, `listJobFiles`, …) are scoped to the authenticated account.
### REST API (universal)
Use plain HTTP with `requests` — the endpoint shapes are in `openapi.yaml`. The core loop is below; `references/workflows.md` has full recipes.
## Core workflow
Always follow discover → schema → validate → submit → poll → results. Do not hardcode tool names or settings — the catalog changes frequently.
```python
import os, time, requests
BASE = "https://app.tamarind.bio/api"
HEADERS = {"x-api-key": os.environ["TAMARIND_API_KEY"]}
# 1. Discover tools. REST /tools returns the full list; filter client-side.
tools = requests.get(f"{BASE}/tools", headers=HEADERS).json()
alphafold = next(t for t in tools if t["name"] == "alphafold")
# 2. Get the exact schema for the chosen tool.
# REST: each /tools entry already includes its inline `settings` schema
# (parameter list) — find the entry whose name == your job type.
# MCP: getJobSchema(jobType) returns the same per-tool detail.
# 3. Submit a job. `settings` is tool-specific — match the schema exactly.
payload = {
"jobName": "my-alphafold-run", # ^[a-zA-Z0-9_-]+$, <=100 chars, unique
"type": "alphafold",
"settings": {
"sequence": "MKTVRQERLKSIVRILERSKEPVSGAQLAEELSVSRQVIVQDIAYLRSLGYNIVATPRGYVLAGG",
"numRecycles": 3,
},
}
resp = requests.post(f"{BASE}/submit-job", headers=HEADERS, json=payload)
resp.raise_for_status() # 200 ok; 400 bad request; 403 budget exceeded; 401 unauthorized
# 4. Poll for completion.
# NOTE the response shape: GET /jobs?jobName=<name> returns the job ROW
# directly (no "jobs" wrapper); the list query (no jobName) returns
# {"jobs": [...]}. Don't index ["jobs"][0] on the by-name response.
while True:
job = requests.get(f"{BASE}/jobs", headers=HEADERS,
params={"jobName": "my-alphafold-run"}).json()
if job["JobStatus"] in ("Complete", "Stopped", "Deleted"):
break
time.sleep(30)
# 5. Retrieve results. POST /result returns a presigned URL *string*;
# GET that URL to download the actual results zip (two-step).
url = requests.post(f"{BASE}/result", headers=HEADERS,
json={"jobName": "my-alphafold-run"}).text.strip('"')
open("my-alphafold-run.zip", "wb").write(requests.get(url).content)
```
For the agentic version of this loop using MCP tools, and for richer examples, see `references/workflows.md`.
## Discovering tools
The catalog has hundreds of tools. Always enumerate at runtime — never rely on a hardcoded list.
**REST** `GET /tools` returns the **full list** (it does not filter server-side); each item is `{name, displayName, github, paper, description, settings}` where `settings` is that tool's inline parameter schema. Filter client-side:
```python
tools = requests.get(f"{BASE}/tools", headers=HEADERS).json() # a list
boltz = [t for t in tools if "boltz" in t["name"].lower()]
```
Note: both surfaces return one row per tool name — REST `/tools` and MCP `getAvailableTools` are both deduplicated (the MCP keeps the newest tool version), so a name match returns a single row.
**MCP** `getAvailableTools(search=..., modality=..., function=...)` filters server-side and adds `categories`/`tags` per tool (`category`/`tag` are deprecated aliases of `modality`/`function`, still honored). Don't hardcode the vocabulary — it drifts. Get the live values from `listModalities()` / `listTags()` (each returns `value`, `label`, `description`, and `toolCount`), or read the `availableCategories` / `availableTags` facet arrays returned on every `getAvailableTools` response. Modalities are molecule types (protein, antibody, peptide, small-molecule, nucleic-acid, …); functions are what a tool does (structure-prediction, binder-design, protein-ligand-docking, …).
A representative set of widely-used tools (verify with `/tools`): `alphafold`, `boltz` (Boltz-2), `chai` (Chai-1), `esmfold` / `esmfold2`, `rfdiffusion`, `proteinmpnn`, `ligandmpnn`, `boltzgen`, `bindcraft`, `diffdock`. See `references/tool_catalog.md` for the full category/tag map and how to read tool metadata.
## Choosing the right tool
The catalog has many tools per task; **don't hardcode a favorite — filter by `function` (and `modality`), then read each candidate's `description` and match it to the user's actual goal** (input you have, output you need, constraints like speed or "no MSA"). The `description` and `tags` fields are the public "what it's for" signal; let them, plus `validateJob`, drive the pick. Quick orientation by task:
- **Fold a single protein / complex** (`function=structure-prediction`): the AlphaFold3-class reproductions — `boltz`/`chai`/`openfold`/`protenix`/`intfold` — are the accurate default for **everything**, including protein-only systems; they also handle **nucleic-acid + small-molecule complexes**, so reach for them whenever a ligand/RNA/DNA is part of the system (and `boltz` adds binding-affinity). `alphafold` (AF2) remains a solid choice for monomers + multimers (join chains with `:`). `esmfold` is single-sequence (no MSA) and fast — reach for it when you want speed and have no MSA; `esmfold2` is newer and conditions on an MSA by default (its `model` setting offers a faster single-sequence mode). Specialized folders exist for antibodies (`abodybuilder`, `immunebuilder`), cyclic peptides (`highfold`), and conformational ensembles (`afcluster`, `alphaflow`) — filter and read descriptions.
- **Design a binder** (`function=binder-design`): `bindcraft` (de novo miniprotein binders) and `boltzgen` (binders for protein **and** small-molecule targets, incl. nanobodies/antibodies/peptides) are the go-to de novo binder tools; `rfdiffusion` also does binder design and is the pick for **motif scaffolding** / diversifying an existing backbone. Antibody-specific generators live under `function=antibody-design`.
- **Design sequence for a known backbone** (`function=inverse-folding`): `proteinmpnn` (general), `ligandmpnn` (ligand-aware), plus thermostable/soluble/antibody MPNN variants. Inverse folding takes a **structure** and emits **sequences** — fold them back to verify (see chaining).
- **Dock a small molecule** (`function=protein-ligand-docking`): prefer `boltz`/`chai` — they co-fold the ligand into the complex and predict the bound structure rather than docking into a fixed receptor; reach for `autodock-vina` when you need fast, large-scale screening against a known pocket.
- **Predict binding affinity** (`function=binding-affinity`) or **generate an MSA** (search `msa`) — filter and read.
When the user names a specific tool, evaluate that one **and** sanity-check the alternatives in its `tag` group — a faster or more appropriate sibling often exists. When unsure, `getJobSchema`/`validateJob` to confirm a candidate actually accepts the input you have before committing.
## Job settings, schemas, and validation
Each tool has its own `settings` schema. Fetch it before submitting:
- **REST** `/tools` entry: each `settings` param is a **trimmed** dict. Only `name` and `required` are always present; `type`, `default`, `description`, `options` appear only when relevant (≈60% have `type`) — so use `param.get("type")`, not `param["type"]`. The advanced gating keys (`exclude`, `conditionals`) are **NOT in the REST response** at all.
- **MCP** `getJobSchema(jobType)`: the **full** schema, including `exclude`, `conditionals`, and bounds. Use MCP when you need to reason about those gating keys. (`restrictOrgs` is stripped on both surfaces — an org-gated param you can't use is simply omitted; see `references/api_reference.md`.)
**Always `validateJob` (MCP) before submitting** — it's the reliable guard. It runs the same validation as `/submit-job` without submitting, and surfaces the first missing/invalid field. Don't try to hand-derive which fields to strip from the schema keys (over REST you can't see them anyway) — let `validateJob` tell you. (The response may include a `source` field, e.g. `"static-fallback"` — an internal note on which schema source validated; `valid: true/false` is the signal you act on.)
`validateJob` echoes a `normalized` view of your settings with defaults filled in. Submit the same clean `settings` you validated; treat `normalized` as informational (it can carry defaults you didn't set, and for some tools platform-managed fields), so build your submit from your own settings rather than the normalized blob.
**Sequences:** amino-acid string; separate chains of a multimer with a colon (`:`), e.g. `"MVLS...:EVQL..."`. Note that some tools (e.g. `boltz`, `chai`) require more than `sequence``boltz` also requires `inputFormat` (and accepts `yamlFile`/`molecules`). Always `getJobSchema`/`validateJob` to learn a tool's required fields; don't assume `sequence` alone suffices.
**Platform-internal fields** — never set these yourself; the platform owns them: `submit_method`, `monomer_msa`, `msa`. See `references/api_reference.md` for the full field-handling rules.
**Surface consequential choices before submitting, don't default silently.** When the request fully specifies what to run, proceed. But when it's open-ended, or when a setting materially changes the results, runtime, or cost (model/variant, number of samples or seeds, MSA on/off, GPU tier, batch size), present the meaningful options plus the default you'd otherwise apply and let the user pick **before** you submit — rather than choosing silently and reporting it after the job is queued. `getJobSchema` and `validateJob`'s `normalized` show exactly which knobs you're filling in on the user's behalf, so you can flag the few worth a quick confirm. This matters most for **batches**, where one shared-settings choice multiplies across every job.
## File inputs (PDB, CIF, SDF, …)
Tools with file parameters accept input three ways:
1. **Upload first, then reference by bare filename.** `PUT /upload/{filename}`, or MCP `uploadFile` → presigned URL → `curl -X PUT -T file "<url>"`. If your host can't reach S3 (a sandboxed agent with no outbound network), use MCP `uploadFileContent(filename, content, encoding?)` to send the file's content through the MCP channel instead — text by default, `encoding="base64"` for binary. The object lands at the S3 key `{email}/{filename}`, **but you reference it in `settings` by the bare `filename` only** (e.g. `"targetFile": "GLP1R_ECD.pdb"`) — the platform scopes it to your account automatically. **Do NOT prefix the email**: passing `{email}/{filename}` double-prefixes the lookup and `submit-job` 400s with `"The following files have not been uploaded: <email>/<file>"`. Confirm the exact name the store registered with MCP `getFiles(search=...)` / REST `GET /files` (a flat list of bare names).
2. **Reference a prior job's output** by its path: `JobName/path/to/file.ext` (this is how you chain jobs — see below).
3. **Inline content.** Send the file's text content directly as the field value.
**Foot-gun:** for a file-typed parameter, a **plain string value is treated as inline file content**, not as a path to an existing object. To point at an already-uploaded file, use the bare `filename` (not the `{email}/...` S3 key) or, for a prior job's output, the `JobName/...` path form — not a bare string you expect to resolve to new content.
**`validateJob` notes.** The response may carry a `source` field (e.g. `"static-fallback"`) — it labels how the tool's *schema* was resolved (built-in tools always report `static-fallback`), **not** whether the validator was reachable, so act on `valid`, not `source`. For file params: reference an uploaded file by its **bare filename** (above) — a bare name resolves to your account-scoped object, whereas an email-prefixed string can be read as inline content and fail the file-type check (`"... must contain ATOM records"`). And passing **inline** file content makes `validateJob` upload it synchronously before validating, which can be slow; prefer referencing an uploaded file by name (above). If a dry-run is slow, skip it and let `submit-job` validate.
## Chaining jobs into pipelines
A finished job's output becomes the next job's input — no download/re-upload. **Match the input type the next tool actually wants:** a sequence-design tool (ProteinMPNN) emits *sequences*, so you fold them by passing each as a `sequence`; a tool that takes a *file* parameter takes a path.
The cleanest design→fold chain is the MCP `submitBatch(fromJob=...)`, which reads a completed design job's generated sequences and folds each as one job:
```
# ProteinMPNN designs sequences -> fold every one with AlphaFold, one call:
submitBatch(batchName="verify-designs", type="alphafold", fromJob="my-proteinmpnn-job")
```
For a **file** input (e.g. a tool that takes a `.pdb`/`.cif`), reference a prior job's output by the path form `JobName/path/to/file.ext` in that file parameter. Two cautions, both confirmed by validation: (1) match the parameter's required **file type** — e.g. AlphaFold's `templateFiles` accepts only `.cif` and is a list, and is gated behind `templateMode: "custom"`; (2) `templateFiles` is for *structural templates*, not for "fold this designed sequence" — to fold a sequence, pass `sequence`. Always `getJobSchema`/`validateJob` to confirm a file param's type/conditions before chaining into it.
To discover a job's exact output paths, use MCP `listJobFiles(job1)` — it returns each file's `s3Path`, usable directly in the next `submitJob`. (The REST `GET /files` lists your account's *uploaded* files as a flat name list; it does not enumerate a job's outputs.) Tamarind also supports saved **pipelines**: build one in the UI, then drive it with `/run-pipeline` (`{pipelineName, initialInputs, inputs}`) or define `stages[]` inline via `/submit-pipeline` (each stage names a `task` + `toolSettings`, using `"pdbFile": "pipe"` to thread one stage's output into the next). See `references/workflows.md`.
## Batch submission
Submit many jobs of the **same tool** in one call. The Python form uses parallel `settings[]` and `jobNames[]` arrays (same length, up to 100):
```python
requests.post(f"{BASE}/submit-batch", headers=HEADERS, json={
"batchName": "egfr-binder-screen",
"type": "alphafold",
"jobNames": ["seq1", "seq2", "seq3"],
"settings": [{"sequence": "..."}, {"sequence": "..."}, {"sequence": "..."}],
# optional: "maxRuntimeSeconds": 3600, "weightedHoursBudget": 100,
# (some accounts also accept an optional "gpuType" — confirm with support)
})
```
**Poll the batch *parent* on `batchStatus`, not subjob `JobStatus`.** A batch creates a parent job (`Type: "batch"`) plus subjobs. Subjobs flip to `Complete` as soon as they finish computing, but the batch then spends a few minutes **aggregating** results into the final downloadable output. Fetch the parent by name and watch `batchStatus`:
```python
import time
while True:
# ?jobName= returns the parent ROW directly (no "jobs" wrapper)
parent = requests.get(f"{BASE}/jobs", headers=HEADERS,
params={"jobName": "egfr-binder-screen"}).json()
bs = parent.get("batchStatus")
if bs == "Complete":
break
if bs in ("Stopped", "AggregationFailed"):
raise RuntimeError(parent.get("AggregationError", bs))
time.sleep(15) # Running / Aggregating -> keep waiting
# When Complete, the parent carries a presigned `resultUrl` and a `statuses`
# subjob tally ({Complete, Running, In Queue, Stopped}).
open("batch.zip", "wb").write(requests.get(parent["resultUrl"]).content)
```
Add `includeSubjobs=true` to `GET /jobs?batch=<name>` to list per-subjob rows.
## Job status lifecycle
Single jobs report `JobStatus`; batch parents report `batchStatus` (poll that for batches — see above).
| Status | Meaning |
|---|---|
| `In Queue` | Accepted, waiting for capacity |
| `Running` | Executing on a worker |
| `Complete` | Finished successfully — results available |
| `Stopped` | Stopped (failure, timeout, manual stop, or budget) |
| `Deleted` | Job was deleted out-of-band |
| `Aggregating` | (batch parent only) subjobs done; building the final output |
| `AggregationFailed` | (batch parent only) aggregation step failed |
Completed jobs carry a `Score` (tool-specific metrics, e.g. pLDDT/pTM/ipTM for folding) and `WeightedHours`. Treat `Complete`/`Stopped`/`Deleted` (and `AggregationFailed` for batches) as terminal; poll on a 15-30s interval. **Break your poll loop on any terminal status, not just `Complete`/`Stopped`** — a job that goes `Deleted` mid-poll would otherwise loop forever. For a `Stopped` job, fetch `getJobLogs(jobName)` to see why. `WeightedHours` is the usage unit billed per job; cap a batch with `weightedHoursBudget`, and a `403` on submit means a budget was hit (see `references/api_reference.md` and the `/usage-statistics` endpoint).
## Error handling
| Code | Meaning | Action |
|---|---|---|
| 400 | Bad request / invalid settings | Re-check against the schema; run `validateJob` first |
| 401 | Unauthorized | Check `x-api-key` |
| 403 | Budget exceeded (org/team) | Lower scope or raise the budget |
| 429 | Rate limited | Back off and retry |
| 500 | Server error | Retry; if persistent, contact support |
## Reference files
The `openapi.yaml` spec is the source of truth for endpoint shapes; these files add the behaviors and gotchas the spec doesn't spell out:
- `references/examples.md`**validated** `settings` payloads per common tool (alphafold/boltz/diffdock/autodock-vina/proteinmpnn/batch), a copy-paste self-check, the "what fails and the exact error" list, and output-shape notes. Start here for a working payload.
- `references/api_reference.md` — endpoint quick-reference + the non-obvious shapes: `/jobs` by-name returns a bare row (not `{jobs:[...]}`), `/result` is a two-step download, batch parents poll on `batchStatus`, `/files` is a flat name list, the `settings` field-handling rules.
- `references/tool_catalog.md` — category/tag map, how to read tool + parameter metadata, common tool families.
- `references/workflows.md` — end-to-end recipes: fold a sequence, validate-before-submit, upload + reference a file, design→fold chaining, batch screen with aggregation polling, usage stats, pagination, and the non-blocking submit-now/check-later pattern for long jobs.
@@ -0,0 +1,178 @@
---
title: "Tamarind Bio REST API reference"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/tamarind/references/api_reference.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: prompt
upstream_changes: accepted
author: upstream
validated: false
---
# Tamarind Bio REST API reference
**Spec:** the OpenAPI spec at `https://app.tamarind.bio/openapi.yaml` (3.0, auth `ApiKeyAuth`) covers the 8 **core job endpoints** (`/submit-job`, `/submit-batch`, `/jobs`, `/result`, `/upload/{filename}`, `/files`, `/delete-job`, `/delete-file`) — fetch it for those exact shapes. It does **not** include the discovery/management endpoints (`/tools`, `/usage-statistics`, `/submit-pipeline`, `/run-pipeline`, `/stop-job`) — for those, use this file + the live MCP `getAvailableTools`/`getJobSchema`/`getJobs`. This file also adds the behaviors no spec spells out (response-shape-by-query, two-step result download, batch aggregation polling, REST-vs-MCP field differences).
Base URL: `https://app.tamarind.bio/api/`
Authentication: `x-api-key: <YOUR_KEY>` header on every request.
Interactive docs: [app.tamarind.bio/api-docs](https://app.tamarind.bio/api-docs) · markdown docs at [docs.tamarind.bio](https://docs.tamarind.bio)
There is no official Python SDK. Call the API with `requests` (Python) or `curl`. An MCP server (`https://mcp.tamarind.bio/mcp`, `X-API-Key` header) exposes the same operations with agent-friendly schemas.
## Endpoints
| Method | Path | Purpose |
|---|---|---|
| GET | `/tools` | List available tools and their inline parameter schemas. Returns the **full list** (no server-side filtering — filter client-side). |
| POST | `/submit-job` | Submit one job. Body: `jobName`, `type`, `settings` (+ optional `projectTag`). |
| POST | `/submit-batch` | Submit many jobs of the same tool. See payload shapes below. |
| GET | `/jobs` | List/inspect jobs. Query: `jobName`, `batch`, `limit`, `startKey`, `organization`, `includeSubjobs`, `jobEmail`. |
| POST | `/result` | Get a presigned download URL for job results (two-step — see below). Body: `jobName` (+ optional `fileName`, `pdbsOnly`, `jobEmail`). |
| POST | `/stop-job` | Stop a running or queued job. Body: `jobName`. |
| DELETE | `/delete-job` | Delete a job and its data. Body: `jobName`. |
| PUT | `/upload/{filename}` | Upload a file (`--data-binary`; add `?folder=` to file it). Or get a presigned URL via MCP `uploadFile`. |
| GET | `/files` | List your account's uploaded files as a flat array of filename strings. Query: `folder`, `includeFolders=true`. Does **not** enumerate a specific job's outputs — use MCP `listJobFiles` for that. |
| DELETE | `/delete-file` | Remove a file/folder. Query: `filePath` or `folder`. |
| POST | `/submit-pipeline` | Run a multi-step pipeline defined inline via `stages[]`. |
| POST | `/run-pipeline` | Run a pipeline saved in the UI. Body: `pipelineName`, `initialInputs`/`inputs`. |
| GET | `/usage-statistics` | Usage/billing. Query: `statistic` (`weighted_hours`/`jobs`), `scope` (`user`/org). |
## Request shapes
### GET /tools
Returns a JSON **array**. Each element:
```json
{
"name": "alphafold",
"displayName": "AlphaFold",
"description": "Accurate and quick protein structure prediction ...",
"github": "https://github.com/...",
"paper": "https://...",
"settings": [ { "name": "sequence", "type": "sequence", "required": true, "description": "..." }, ... ]
}
```
In each `settings` param, only `name` and `required` are guaranteed; `type`, `default`, `description`, `options` are present only when applicable (about 60% of params carry `type`). Read them with `param.get("type")`, not `param["type"]`.
`settings` is the tool's inline parameter schema — read it directly, no separate schema endpoint over REST. The REST list is not filtered by query params; filter client-side on `name`/`displayName`/`description`. (The MCP `getAvailableTools` wraps the list as `{"totalTools", "tools":[...]}` and adds `categories`/`tags` per tool plus server-side `search`/`category`/`tag` filtering.)
### POST /submit-job
```json
{
"jobName": "my-protein-analysis",
"type": "alphafold",
"settings": { "sequence": "MKT...", "numRecycles": 3 },
"projectTag": "proj_xxxxxxxx"
}
```
- `jobName` — unique, `^[a-zA-Z0-9_-]+$`, 1-100 chars.
- `type` — a tool name from `/tools`. The list changes often; never hardcode.
- `settings` — tool-specific; match the schema from `/tools` (or MCP `getJobSchema`).
- `projectTag` — optional `proj_...` ProjectId to file the job under a project.
Response (200): a confirmation string like `myJobName submitted to queue.`
### POST /submit-batch
Two payload shapes appear in the official docs — the **Python** form uses parallel arrays; the **curl** form uses a `jobs[]` array of objects with a `tool` key. The parallel-array form matches the MCP `submitBatch` and is the recommended one:
```json
{
"batchName": "egfr-screen",
"type": "alphafold",
"jobNames": ["seq1", "seq2"],
"settings": [{ "sequence": "..." }, { "sequence": "..." }],
"maxRuntimeSeconds": 3600,
"weightedHoursBudget": 100
}
```
curl-form alternative (same endpoint): `{ "tool": "<type>", "batchName": ..., "jobs": [{ "jobName": ..., "settings": {...} }, ...] }`.
- `jobNames` and `settings` are parallel arrays, same length, 1-100 items, all using the same tool.
- `maxRuntimeSeconds` — optional per-job timeout. `weightedHoursBudget` — optional budget cap.
- The MCP `submitBatch` schema exposes `maxRuntimeSeconds` + `weightedHoursBudget`. Some accounts/tools may accept an optional `gpuType` (seen in the docs UI), but it isn't in `openapi.yaml` or the MCP schema — treat it as unverified and confirm with support before relying on it.
### GET /jobs
**Response shape depends on the query:**
- **List / batch query** (no `jobName`, or `?batch=`/`?organization=`) → `{ "jobs": [...], "startKey": "...", "statuses": {...} }`.
- **By-name** (`?jobName=<name>`) → the **job row object directly** (no `jobs` wrapper). Don't index `["jobs"][0]` on this response.
Each job row includes `JobName`, `Type`, `JobStatus`, `Created`, `Started`, `Completed`, `Settings` (JSON string), `Score` (JSON string, tool metrics), `WeightedHours`. Use `startKey` for pagination past the `limit` (default 1000). Only top-level jobs return by default; add `includeSubjobs=true` for batch subjobs.
**Batch parent rows** have `Type: "batch"` and carry `batchStatus`. Fetched by name (`?jobName=<batchName>`), a complete batch parent also includes `resultUrl` (presigned download). `batchStatus` transitions: `Running``Aggregating``Complete` (or `AggregationFailed`, with `AggregationError`). Poll the parent's `batchStatus`, not subjob `JobStatus` — subjobs go `Complete` before the aggregated output is ready.
**Discriminate batch vs single by `Type == "batch"` (or presence of `batchStatus`), not by `statuses`.** A by-name response can carry a `statuses` tally even for a single (non-batch) job, so `statuses` presence is not a reliable batch signal.
### POST /result (two-step download)
POST returns a presigned URL as a **bare string** (not JSON). Fetch that URL with a second GET to download the results zip:
```python
url = requests.post(f"{BASE}/result", headers=H, json={"jobName": "myJob"}).text.strip('"')
open("myJob.zip", "wb").write(requests.get(url).content)
```
Optional body fields: `fileName` (one file instead of the zip), `pdbsOnly: true` (PDB outputs only), `jobEmail` (a teammate's job, if permitted).
## Status codes
| Code | Meaning |
|---|---|
| 200 | Success |
| 400 | Bad request — invalid parameters/settings |
| 401 | Unauthorized — invalid/missing `x-api-key` |
| 403 | Budget exceeded (org/team) |
| 429 | Rate limited |
| 404 | Not found (e.g. unknown job) |
| 500 | Server error |
## Field-handling rules (important)
**The REST and MCP schemas expose different fields.** The REST `/tools` entry
gives a trimmed per-param view — `{name, type, required, default, description, options}`.
The advanced gating keys `exclude` and `conditionals` appear **only in MCP
`getJobSchema`**, not in REST `/tools` (`restrictOrgs` is no longer returned by
either surface — see below). So don't try to hand-derive what to strip from REST
schema keys — they aren't there. The reliable guard on
both surfaces is **`validateJob`** (MCP): it runs `/submit-job`'s exact validation
without submitting and returns the first error.
- **Build your submit from your own settings, not `validateJob`'s `normalized` output.**
`normalized` is informational (defaults filled in, sometimes platform-managed
fields). Submit the same clean settings you validated, not the normalized echo.
- **Platform-internal routing fields**`submit_method`, `monomer_msa`, `msa` are
set by the platform. Never pass them.
- **`restrictOrgs`** — org-gated parameters. `getJobSchema` no longer returns this
key (it's stripped server-side): a parameter your account isn't authorized for is
dropped from the schema entirely, and any param you do see is one you may set. So
you won't encounter `restrictOrgs` in a response — don't look for it.
- **`conditionals`** (MCP schema only) — a field only applies when another field
has a given value (e.g. `pairMode` applies only when `useMSA` is `true`). Don't
send conditioned fields when their condition isn't met.
- **`exclude: [...]`** (MCP schema only) — marks a field as UI/pipeline-only for a
surface. Treat it as advisory; `validateJob` is the authority on what a given
submission accepts.
- **`required: true`** — must be present. Some tools require more than `sequence`
(e.g. `boltz` requires `inputFormat`). Run `validateJob` to get the first
missing/invalid field before submitting.
- **File-typed fields with a plain string value are treated as INLINE CONTENT**,
not a path. To reference an **uploaded file**, use its **bare filename**
(`target.pdb`) — the platform scopes it to your account, so do NOT email-prefix
it. The `{email}/{filename}` form is the underlying S3 key, and passing it makes
`submit-job` 400 with `"The following files have not been uploaded: <email>/<file>"`.
To reference a **prior job's output**, use `JobName/path/to/file.ext`. Confirm the
exact registered name with `getFiles` / `GET /files` (a flat list of bare names).
## Authentication and secrets
- Read the key from `TAMARIND_API_KEY` (env or `.env`); never hardcode or commit it.
- The same key authenticates REST (`x-api-key`) and the MCP server (`X-API-Key`).
- Query operations are scoped to the authenticated account (and, with `organization=true`/`jobEmail`, to your org if permitted).
@@ -0,0 +1,145 @@
---
title: "Tamarind Bio — validated examples & output shapes"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/tamarind/references/examples.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: prompt
upstream_changes: accepted
author: upstream
validated: false
---
# Tamarind Bio — validated examples & output shapes
**The freshest example for any tool is the `exampleJob` field MCP `getJobSchema(<tool>)`
now returns** — an `{jobName, type, settings}` built from each param's example/default
(with an `exampleJobNote`; org-gated params you can't use are omitted, file params get
placeholder filenames). It's the best starting point, but **run `validateJob` on it
before submitting** — it's assembled from per-param examples, not a guaranteed-valid
payload, so a given tool's `exampleJob` can need a tweak. The payloads below are a
`validateJob`-confirmed fallback for REST callers or when you want a worked example.
Tool schemas evolve — if one stops validating, re-fetch with `getJobSchema(<tool>)` /
`GET /tools`. Sequences here are illustrative; swap your own.
**File params (`proteinFile`, `pdbFile`, `ligandFile`, …) need a real file value** —
either the **bare filename** of an uploaded file (`target.pdb` — NOT email-prefixed),
a prior-job output **path** (`JobName/out/x.pdb`), or
**inline PDB/SDF-format text** (multi-line `ATOM`/`HETATM` records). The
`<...>` placeholders below are NOT valid as written — replace them. **Do not put an
amino-acid sequence in a file param** — `validateJob` rejects it with
`File ... must be of types: ["pdb"]`. (A sequence goes in `sequence`, a structure
goes in a file param.)
`BASE = "https://app.tamarind.bio/api"`, `HEADERS = {"x-api-key": <key>}`.
## Self-check (run this first to confirm the skill works for you)
Read-only + dry-run, no submission, no cost. Confirms the discover → schema →
validate loop end-to-end:
```python
import os, requests
BASE, HEADERS = "https://app.tamarind.bio/api", {"x-api-key": os.environ["TAMARIND_API_KEY"]}
# 1. discovery reachable?
tools = requests.get(f"{BASE}/tools", headers=HEADERS).json()
assert isinstance(tools, list) and any(t["name"] == "alphafold" for t in tools), "tools endpoint"
# 2. validate a known-good payload (MCP validateJob; or skip if REST-only)
# expect {"valid": true, ...}
```
With the MCP server: `validateJob(jobName="selfcheck", type="alphafold",
settings={"sequence": "MKTAYIAKQRQISFVKSHFSRQLEERLGLIE"})` → `valid: true`.
## Validated input payloads
### AlphaFold — monomer
```json
{ "sequence": "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALPDAQFEVVHSLAKWKR",
"numModels": "1", "numRecycles": 3 }
```
Only `sequence` is required; everything else has a default. `numModels` is a string
dropdown (`"1"``"5"`).
### AlphaFold — multimer (colon-separated chains)
```json
{ "sequence": "MKTAYIAKQRQISFVKSHFSRQLEERLGLIE:DIQMTQSPSSLSASVGDRVTITCRASQSISSYLN" }
```
Join chains with `:`. No separate "multimer" flag — chain count drives it.
### Boltz-2 — sequence mode
```json
{ "inputFormat": "sequence",
"sequence": "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALP" }
```
`inputFormat` is **required** (`"sequence"` / `"list"` / `"molecules"` / `"yaml"`).
Omitting it fails — see "What fails" below.
### DiffDock — protein + SMILES ligand
```json
{ "ligandFormat": "SMILES",
"ligandSmiles": "CC(=O)Oc1ccccc1C(=O)O",
"proteinFile": "<uploaded-path-or-inline-PDB-text>" }
```
`ligandFormat` chooses the conditional field: `"SMILES"``ligandSmiles`;
`"sdf/mol2 file"``ligandFile`. `proteinFile` is a file param — pass an uploaded
file's bare filename (`target.pdb`, not email-prefixed), a prior-job path
(`JobName/...`), or inline PDB text (see file-input rules in `api_reference.md`).
### Autodock Vina — protein + SMILES ligand (classical docking into a pocket)
```json
{ "receptorFile": "receptor.pdb",
"ligandFormat": "smiles",
"ligandSmiles": "CC(=O)Oc1ccccc1C(=O)O",
"boxX": 15.19, "boxY": 53.903, "boxZ": 16.917,
"width": 20, "height": 20, "depth": 20 }
```
Unlike DiffDock, Autodock Vina docks into a **fixed pocket**, so it requires a bounding
box (`boxX/Y/Z` center + `width/height/depth`, all required) and the receptor in
`receptorFile` (not `proteinFile`). Its `ligandFormat` enum is **lowercase**
(`"smiles"` / `"sdf"`) — different from DiffDock's `"SMILES"` / `"sdf/mol2 file"`, so
don't copy DiffDock's value across. `exhaustiveness` (default 8) is optional. `validateJob`-confirmed.
### ProteinMPNN — design residues on a backbone
```json
{ "pdbFile": "<uploaded-path-or-inline-PDB-text>",
"designedResidues": { "A": "1 2 3 4 5" },
"numSequences": 4, "modelType": "proteinmpnn" }
```
Requires `pdbFile` + `designedResidues` (per-chain, space-separated resnums).
`modelType``proteinmpnn`/`ligandmpnn`/`solublempnn`/`hypermpnn`/`abmpnn`.
Note `designedChains` is `exclude:["api"]` — don't send it over the API.
### Batch (same tool, many jobs)
```json
{ "batchName": "screen-1", "type": "alphafold",
"jobNames": ["s1", "s2"],
"settings": [ { "sequence": "MKT..." }, { "sequence": "AVF..." } ] }
```
## What fails (and the exact error) — confirmed live
- **Boltz without `inputFormat`**`valid:false`, `Missing required boltz field "inputFormat"`. Always check required fields with `getJobSchema` first; `sequence` alone is not enough for boltz/chai.
- **Building a submit from `validateJob`'s `normalized` blob**`normalized` is informational (defaults filled in, sometimes platform-managed fields). Submit the clean `settings` you validated, not the normalized echo.
- **File param given a bare string that isn't a real path** → treated as INLINE file content (uploaded as `<email>/<jobname>-<param>.<ext>`), not a reference. To point at an existing uploaded file use its **bare filename** (`target.pdb` — do NOT email-prefix it; `{email}/{filename}` is the S3 key and 400s as not-uploaded), or `JobName/...` for a prior job's output. Referencing a path that doesn't exist → `File ... has not been uploaded`.
## Output shapes (describe, don't expect exact values)
Outputs are non-deterministic (seed/model/MSA) — reason about the *shape*, not
golden numbers.
- **Job row `Score`** (JSON string on completed jobs): tool-family dependent.
- Folding (alphafold/boltz/chai/esmfold): `plddt`, `ptm`, and for complexes
`iptm` plus interface metrics (`ipSAE_*`, `pDockQ_*`). Higher pLDDT/pTM = more
confident; iptm/ipSAE gauge interface quality.
- Other families carry their own metrics — read the keys, don't assume.
- **Results zip** (`POST /result` → presigned URL → GET): per-tool, typically the
structure files (`rank_*.pdb` / `*.cif`), a scores CSV, and logs. Use
`listJobFiles(jobName)` (MCP) to enumerate exact filenames before downloading.
- **`WeightedHours`** on the row is the billing unit (see `usage-statistics`).
To learn a specific tool's exact outputs, run one small job and `listJobFiles` it —
don't hardcode filenames, which vary by tool and version.
@@ -0,0 +1,79 @@
---
title: "Tamarind Bio tool catalog"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/tamarind/references/tool_catalog.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: prompt
upstream_changes: accepted
author: upstream
validated: false
---
# Tamarind Bio tool catalog
Tamarind exposes hundreds of tools through one uniform job API. The catalog changes frequently — **always enumerate at runtime** with `GET /tools` (or MCP `getAvailableTools`) rather than hardcoding names. This file is a map for interpreting what you get back.
## How to discover
**REST** `GET /tools` returns the **full list** (it does not filter server-side). Filter client-side:
```python
tools = requests.get(f"{BASE}/tools", headers=HEADERS).json() # a list
docking = [t for t in tools if "vina" in t["name"].lower()]
```
Each REST tool entry carries: `name` (the `type` you submit), `displayName`, `description`, `github`, `paper`, and `settings` (the inline parameter schema). REST entries do **not** include `categories`/`tags`.
**MCP** `getAvailableTools(search=..., modality=..., function=...)` filters server-side and returns entries with `categories` and `tags` (`category`/`tag` are deprecated aliases of `modality`/`function`, still honored).
## Modalities and functions (the two filter axes)
Don't hardcode the filter vocabulary — it drifts as tools are added. Fetch it live: `listModalities()` returns the molecule-type axis (protein, antibody, enzyme, peptide, nucleic-acid, small-molecule, small-molecule-binding-protein, cryoem, …); `listTags()` returns the function axis (structure-prediction, protein-design, binder-design, protein-ligand-docking, binding-affinity, inverse-folding, developability, molecular-dynamics, finetuning, …). Each entry carries `value`, `label`, `description`, and a live `toolCount`. Every `getAvailableTools` response also includes `availableCategories` / `availableTags` arrays computed from the current catalog. Filter with `getAvailableTools(modality=..., function=...)`.
## Representative tool families
Verify exact names and availability with `/tools` — these are common anchors, not an exhaustive or guaranteed list.
**Structure prediction / folding**
- `alphafold` — AlphaFold; monomer + multimer, MSA + templates, recycles, relaxation.
- `boltz` — Boltz-2; structure + affinity, biomolecular complexes incl. ligands.
- `chai` — Chai-1; complex structure prediction with optional MSA.
- `esmfold` / `esmfold2` — fast single-sequence folding.
**Protein / binder design**
- `rfdiffusion` — protein/binder design and motif scaffolding.
- `boltzgen` — generative design.
- `bindcraft` — binder design.
- `proteinmpnn` / `ligandmpnn` — inverse folding (sequence given backbone; ligand-aware variant).
**Docking / affinity**
- `boltz` / `chai` — co-fold the ligand into the complex (predict the bound structure); the default for protein-small-molecule docking.
- `autodock-vina` — classical docking into a known pocket; the pick for fast, large-scale screening.
- Boltz/affinity tools — binding-affinity prediction.
**Antibody**
- Antibody language models and generators, humanization, developability, immunogenicity scoring.
**MSA / utilities**
- MSA generation tools feed downstream folding; utilities cover format conversion, scoring, and analysis.
## Reading a tool schema
`getJobSchema(jobType)` (MCP) or the `/tools` entry returns a `parameters` list. Each parameter has:
- `name`, `type` (`sequence`, `number`, `boolean`, `dropdown`, file types like `pdb`/`cif`/`sdf`, …)
- `descr`, `displayName`
- `required`, `default`
- `options` / `optionsDescr` (for dropdowns), `lowerBound` / `upperBound` / `lengthLimit`
- `conditionals` — applies only when another field has a given value
- `exclude` (`["api"]` / `["batch"]`) — omit on that surface
- `list: true` — accepts multiple values/files
- `example` — a sample value
(Org-gated parameters are filtered server-side: `getJobSchema` drops a param your account isn't authorized for and never returns the old `restrictOrgs` key.)
Top-level tool metadata also includes a `hint`, and `getJobSchema` returns an `exampleJob` built from each parameter's example/default — start from that (then `validateJob` it) rather than hand-building `settings`.
Always read the schema before constructing `settings`, and run `validateJob` to confirm before `submitJob`.
@@ -0,0 +1,276 @@
---
title: "Tamarind Bio workflow recipes"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/0807ddbc/skills/tamarind/references/workflows.md
upstream_sha: 0807ddbc
imported_at: 2026-06-30
prompt_class: prompt
upstream_changes: accepted
author: upstream
validated: false
---
# Tamarind Bio workflow recipes
End-to-end examples using plain `requests`. All use `BASE = "https://app.tamarind.bio/api"` and
`HEADERS = {"x-api-key": os.environ["TAMARIND_API_KEY"]}`. For exact request/response
shapes, fetch the spec at `https://app.tamarind.bio/openapi.yaml`.
The canonical loop is always: **discover → schema → validate → submit → poll → results**.
## 1. Fold a single sequence (AlphaFold)
```python
import os, time, requests
BASE = "https://app.tamarind.bio/api"
HEADERS = {"x-api-key": os.environ["TAMARIND_API_KEY"]}
# discover + confirm the tool exists (REST returns the full list; filter client-side)
tools = requests.get(f"{BASE}/tools", headers=HEADERS).json()
assert any(t["name"] == "alphafold" for t in tools)
job = {
"jobName": "ubiquitin-fold",
"type": "alphafold",
"settings": {
"sequence": "MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG",
"numModels": "5",
"numRecycles": 3,
"useMSA": True,
},
}
requests.post(f"{BASE}/submit-job", headers=HEADERS, json=job).raise_for_status()
# poll. GET /jobs?jobName= returns the job ROW directly (no "jobs" wrapper).
while True:
row = requests.get(f"{BASE}/jobs", headers=HEADERS,
params={"jobName": "ubiquitin-fold"}).json()
if row["JobStatus"] in ("Complete", "Stopped", "Deleted"):
break
time.sleep(30)
print("status:", row["JobStatus"], "score:", row.get("Score"))
# results download is two-step: POST /result returns a presigned URL *string*,
# then GET that URL for the zip.
url = requests.post(f"{BASE}/result", headers=HEADERS,
json={"jobName": "ubiquitin-fold"}).text.strip('"')
open("ubiquitin-fold.zip", "wb").write(requests.get(url).content)
```
## 2. Multimer / complex (colon-separated chains)
For AlphaFold, a multimer is just one `sequence` with chains joined by `:`.
```python
job = {
"jobName": "ab-ag-complex",
"type": "alphafold",
"settings": {
# heavy:light:antigen — separate chains with ":"
"sequence": "EVQLVESGGG...:DIQMTQSPSS...:MKTAYIAKQR...",
},
}
requests.post(f"{BASE}/submit-job", headers=HEADERS, json=job).raise_for_status()
```
Other folding tools need more fields — `boltz`/`chai` require `inputFormat`
(`"sequence"`/`"list"`/`"molecules"`/`"yaml"`), e.g. boltz sequence-mode is
`{"inputFormat": "sequence", "sequence": "...:..."}`. **Always check required
fields with `getJobSchema`/`validateJob` first** — don't assume `sequence` alone
is enough.
## 3. Validate before submitting (MCP)
When your agent host has the Tamarind MCP server, dry-run first to catch errors
without spending a submission. Validate and submit **your own clean settings**
build the submit from `my_settings`, not `verdict["normalized"]` (normalized is
informational: defaults filled in, sometimes platform-managed fields).
```
getJobSchema(jobType="boltz") # learn required fields first
my_settings = {"inputFormat": "sequence", "sequence": "...:..."}
verdict = validateJob(jobName="x", type="boltz", settings=my_settings)
# verdict.valid == True -> good; submit my_settings (NOT verdict.normalized)
# verdict.valid == False -> verdict.error is the first problem to fix
if verdict["valid"]:
submitJob(jobName="x", type="boltz", settings=my_settings)
```
## 4. Upload a structure, then submit a job that uses it
```python
# REST: PUT the file to /upload/{filename}
with open("target.pdb", "rb") as fh:
requests.put(f"{BASE}/upload/target.pdb", headers=HEADERS, data=fh).raise_for_status()
# the object's S3 key is "{your-email}/target.pdb", but you reference it by the
# BARE filename — the platform scopes it to your account. Do NOT email-prefix it.
job = {
"jobName": "dock-run",
"type": "diffdock",
"settings": {
"proteinFile": "target.pdb", # bare filename, NOT inline content, NOT email-prefixed
"ligandFormat": "SMILES", # required; gates ligandSmiles vs ligandFile
"ligandSmiles": "CC(=O)Oc1ccccc1C(=O)O",
},
}
requests.post(f"{BASE}/submit-job", headers=HEADERS, json=job).raise_for_status()
```
MCP variant: `uploadFile("target.pdb")` returns a presigned `uploadUrl`; then
`curl -X PUT -T target.pdb "<uploadUrl>"`.
**Reminder:** a bare *non-filename* string in a file-typed field is uploaded as inline content.
To point at an existing uploaded file, use its **bare filename** (`target.pdb`) — NOT the
`{email}/{filename}` S3-key form, which `submit-job` 400s as `"... has not been uploaded"`.
Confirm the registered name with `getFiles`/`GET /files`. For a prior job's output, use `JobName/...`.
For `autodock-vina` instead of DiffDock, the same upload-then-reference flow applies, but the
settings differ: it docks into a fixed pocket, so it needs `receptorFile` + a bounding box
(`boxX/Y/Z`, `width/height/depth`) and a **lowercase** `ligandFormat` (`"smiles"`/`"sdf"`).
Run `getJobSchema("autodock-vina")` for the full shape; see `examples.md` for a worked payload.
## 5. Chain jobs: design → fold
A sequence-design tool (ProteinMPNN) emits **sequences**, so you fold them by
passing each as a `sequence` — NOT via a template/file field. The cleanest way is
the MCP `submitBatch(fromJob=...)`, which reads the design job's generated
sequences and folds each as one job in a single call:
```
# Step 1: design sequences for a backbone
submitJob(jobName="design-step", type="proteinmpnn", settings={...}) # poll to Complete
# Step 2: fold every designed sequence (MCP reads them from the design job)
submitBatch(batchName="fold-designs", type="alphafold", fromJob="design-step")
```
Doing it over plain REST instead: read the design job's output sequences (MCP
`listJobFiles("design-step")``s3Path`, or download the FASTA via `/result`),
then submit one fold per sequence with `settings={"sequence": "<designed seq>"}`.
**Don't chain a designed sequence through a file/template field.** A file
parameter wants a *file of the right type*, and a template field is for structural
homology, not "fold this sequence." Example of the trap: AlphaFold's
`templateFiles` accepts only `.cif`, must be a **list**, and is gated behind
`templateMode: "custom"` — so `{"templateFiles": "design-step/out/x.pdb"}` fails
validation three ways and isn't how you fold a design anyway. When a chain really
does feed a file (e.g. a PDB into a docking tool), `getJobSchema`/`validateJob`
first to confirm the param's type and conditions.
For reusable multi-step flows, build a saved pipeline with `/submit-pipeline`
and run it with `/run-pipeline`.
## 6. Batch screen many sequences through one tool
Submit, then poll the batch **parent** on `batchStatus` (not subjob `JobStatus`)
— the batch aggregates results after subjobs finish computing.
```python
seqs = ["MKT...", "AVF...", "GEV..."]
requests.post(f"{BASE}/submit-batch", headers=HEADERS, json={
"batchName": "binder-screen",
"type": "alphafold",
"jobNames": [f"cand-{i}" for i in range(len(seqs))],
"settings": [{"sequence": s} for s in seqs],
"weightedHoursBudget": 50, # optional budget cap
}).raise_for_status()
# poll the parent until the aggregated output is ready
# (?jobName= returns the parent ROW directly — no "jobs" wrapper)
while True:
parent = requests.get(f"{BASE}/jobs", headers=HEADERS,
params={"jobName": "binder-screen"}).json()
bs = parent.get("batchStatus")
if bs == "Complete":
break
if bs in ("Stopped", "AggregationFailed"):
raise RuntimeError(parent.get("AggregationError", bs))
time.sleep(15) # Running / Aggregating -> keep waiting
print(parent["statuses"]) # e.g. {"Complete": 3, "Running": 0, "In Queue": 0, "Stopped": 0}
open("binder-screen.zip", "wb").write(requests.get(parent["resultUrl"]).content)
# Per-subjob rows (e.g. to read each candidate's Score):
subjobs = requests.get(f"{BASE}/jobs", headers=HEADERS,
params={"batch": "binder-screen", "includeSubjobs": "true"}).json()
```
## 7. Debug a stopped job
```python
# REST: pull results/log path; MCP gives logs directly
logs = getJobLogs("binder-screen-cand-2") # MCP: last N lines of output log
# Inspect the tail for the failure reason (bad input, OOM, timeout, budget).
```
A `Stopped` status with no `Score` usually means a failure — read the log tail.
A `403` at submit means a budget cap was hit.
## 8. List every job (paginate past the 1000 limit)
The list query returns `{"jobs": [...], "startKey": ...}`; pass `startKey` back
until it's absent.
```python
jobs, params = [], {"limit": 1000}
while True:
resp = requests.get(f"{BASE}/jobs", headers=HEADERS, params=params).json()
jobs += resp["jobs"]
if "startKey" not in resp:
break
params["startKey"] = resp["startKey"]
print(len(jobs))
```
## 9. Submit now, check back later (non-blocking)
Bio jobs run for minutes to hours — you don't have to hold a blocking poll loop
open. Jobs are addressable by `jobName` from any process, so submit, **persist the
names**, and reconnect in a separate session/process to collect results. This is the
right pattern for long campaigns or fire-and-forget pipelines.
```python
# --- Session 1: submit and save the job names ---
import os, json, requests
BASE = "https://app.tamarind.bio/api"
HEADERS = {"x-api-key": os.environ["TAMARIND_API_KEY"]}
seqs = {"cand-a": "MKT...", "cand-b": "AVF...", "cand-c": "GEV..."}
for name, seq in seqs.items():
requests.post(f"{BASE}/submit-job", headers=HEADERS,
json={"jobName": name, "type": "alphafold",
"settings": {"sequence": seq}}).raise_for_status()
json.dump(list(seqs), open("pending_jobs.json", "w")) # persist to disk/db
print("submitted; check back later")
```
```python
# --- Session 2 (later, fresh process): collect whatever is done ---
import os, json, requests
BASE = "https://app.tamarind.bio/api"
HEADERS = {"x-api-key": os.environ["TAMARIND_API_KEY"]}
names = json.load(open("pending_jobs.json"))
done, pending = [], []
for name in names:
row = requests.get(f"{BASE}/jobs", headers=HEADERS,
params={"jobName": name}).json() # bare row, by-name
(done if row["JobStatus"] in ("Complete", "Stopped", "Deleted") else pending).append(name)
print(f"{len(done)} terminal, {len(pending)} still running")
for name in done:
url = requests.post(f"{BASE}/result", headers=HEADERS,
json={"jobName": name}).text.strip('"')
open(f"{name}.zip", "wb").write(requests.get(url).content)
```
Re-run session 2 until `pending` is empty. For a server-driven variant, poll a batch
parent's `batchStatus` (recipe 6) instead of looping job-by-job.
## Notes
- **Polling cadence:** 15-30s. `Complete` and `Stopped` are terminal.
- **Scores:** completed folding jobs return pLDDT / pTM / ipTM (and interface
metrics like ipSAE / pDockQ for complexes) in the `Score` field.