[Upstream sync] mims-harvard/ToolUniverse (github) — 2 added, 7 modified #34

Merged
promptadmin merged 9 commits from upstream-sync/tooluniverse-20260808-cfd267-ttzn into main 2026-08-09 18:32:28 +00:00
9 changed files with 699 additions and 31 deletions
@@ -2,9 +2,9 @@
title: "ToolUniverse Skills"
task: ""
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/bb632a34/skills/README.md
upstream_sha: bb632a34
imported_at: 2026-07-01
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/README.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: catalogue
upstream_changes: accepted
author: upstream
@@ -13,7 +13,7 @@ validated: false
# ToolUniverse Skills
68 pre-built research skills for AI agents. Skills are automatically available — just ask naturally.
69 pre-built research skills for AI agents. Skills are automatically available — just ask naturally.
```
"Find the E. coli K-12 genome" → tooluniverse-sequence-retrieval
@@ -54,6 +54,7 @@ npx skills add mims-harvard/ToolUniverse
| `tooluniverse-epigenomics` | Epigenomics and gene regulation (ENCODE, JASPAR, methylation) |
| `tooluniverse-expression-data-retrieval` | Gene expression datasets from ArrayExpress and BioStudies |
| `tooluniverse-gene-enrichment` | Gene enrichment and pathway analysis (gseapy, PANTHER, STRING, etc.) |
| `tooluniverse-gene-liability` | Human safety liability scoring for gene inhibition or loss of function |
| `tooluniverse-gwas-drug-discovery` | GWAS signals to drug targets and repurposing opportunities |
| `tooluniverse-gwas-finemapping` | Causal variant prioritization via statistical fine-mapping |
| `tooluniverse-gwas-snp-interpretation` | Genetic variant interpretation from GWAS studies |
@@ -124,4 +125,7 @@ skill-name/
1. Create `skill-name/SKILL.md` with proper frontmatter
2. Keep SKILL.md concise (<500 lines)
3. Add examples for common use cases
4. See `create-tooluniverse-skill` for the full workflow
4. Keep canonical skills host-agnostic. Do not add client-specific metadata such
as `agents/openai.yaml`; generate host-specific files in the corresponding
plugin packaging layer instead.
5. See `create-tooluniverse-skill` for the full workflow
@@ -1,8 +1,8 @@
---
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/701afa3a/skills/tooluniverse-biomedical-fact-lookup/SKILL.md
upstream_sha: 701afa3a
imported_at: 2026-07-07
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse-biomedical-fact-lookup/SKILL.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: unknown
upstream_changes: accepted
name: tooluniverse-biomedical-fact-lookup
@@ -37,10 +37,11 @@ Most of these questions are MCQ with an "Insufficient information to answer the
| **TF binding site / target** "according to GTRD" (e.g. PGM3) | `MSigDB_check_gene_in_set` (collection C3:TFT:GTRD) | set name = `<TF>_TARGET_GENES`, e.g. `PGM3_TARGET_GENES`; pass `gene` per option |
| **pathway / hallmark** membership | `MSigDB_get_hallmark_geneset`, `MSigDB_get_geneset` | `HALLMARK_<NAME>` or exact set name |
| **gene ↔ disease** association (DisGeNet, OpenTargets, OMIM) | `umls_search_concepts` → `DisGeNET_get_disease_genes`/`DisGeNET_get_gda`; `OpenTargets_*`, `MyDisease_get_disease`, `OMIM_search`; **text-mined fallback:** `PubTator3_LiteratureSearch` / `PubTator3_GetEntityRelations` (`e1=@GENE_<sym>`), `EPMC_get_text_mined_annotations` | DisGeNET needs a **UMLS CUI** (resolve via `umls_search_concepts` → `C0152200`, then `disease=C0152200`) + `DISGENET_API_KEY`. See the "in X but not Y" recipe below |
| **mouse phenotype** gene set (MGI / MP:xxxxx, e.g. "increased carcinoma incidence") | `MGI_search_genes` → `MGI_get_phenotypes` | for **each** candidate gene: search → take the `MGI:` id → `MGI_get_phenotypes`; the matching gene is the one whose `phenotype_statement` list contains the phenotype the question names (see interpretation note) |
| **mouse phenotype** gene set (MP / MGI, e.g. "increased melanoma incidence") | `MSigDB_check_gene_in_set` (mouse M5, set `MP_<PHENOTYPE>`) — fall back to `MGI_search_genes` → `MGI_get_phenotypes` | **one call per option** against the `MP_*` set (e.g. `MP_INCREASED_MELANOMA_INCIDENCE`); the member is the answer. Only if the set name doesn't resolve, use the MGI per-gene route below |
| **gene genomic location** (Ensembl band, e.g. chr7q34) | `Ensembl_*` / `NCBIDatasets_get_gene_by_symbol` | resolve each option, compare cytoband/coordinates |
| **variant / sequence** pathogenicity ("which variant/sequence is pathogenic *or* benign per ClinVar") | (only when genuinely unsure) `annotate_variant_multi_source`, `VEP_predict_pathogenicity`, `UniProt_get_disease_variants_by_accession` | **Be efficient — do NOT query every option (that causes timeouts).** Identify the protein once, find each option's single substitution, and reason about the specific residue changes directly; the base model is usually reliable on well-characterized ClinVar variants. Make at most ONE targeted tool call to resolve a truly uncertain variant. **Watch the question's polarity** (benign vs pathogenic): for "most likely benign", a common/reference-matching variant is the answer; for "most likely pathogenic", a rare damaging one is. |
| **drug / compound** target, MoA, approval | `ChEMBL_*`, `OpenFDA_*`, `GtoPdb_*`, `PubChem_*` | resolve drug, query the relation |
| **which drug for this patient** (clinical vignette naming a modifier) | `FDA_*_by_drug_name` — pick the section by modifier | see "Drug choice for a described patient" below |
| **protein** function / domain / sequence | `UniProt_*` | resolve accession, read annotation |
When unsure which tool wraps a database, search the catalog by the *relation* (e.g. "gene disease association", "gene set members"), not the brand name — ToolUniverse usually already has it.
@@ -53,8 +54,11 @@ ToolUniverse's `MSigDB_*` tools cover several collections that LAB-Bench questio
- **C3:MIR:MIRDB** (miRDB v6.0 predicted miRNA targets) — `MIR<number>_<3P|5P>` (e.g. `MIR186_3P`, `MIR675_3P`). This *is* miRDB; do not say "no access to miRDB".
- **C3:TFT:GTRD** (GTRD TF target genes) — `<TF>_TARGET_GENES` (e.g. `PGM3_TARGET_GENES`). This *is* GTRD.
- **Hallmark** — `HALLMARK_<NAME>`.
- **Mouse M5 (MGI mammalian phenotype)** — `MP_<PHENOTYPE_IN_CAPS>` (e.g. "increased melanoma incidence" → `MP_INCREASED_MELANOMA_INCIDENCE`). These are **mouse** sets: the tools try human then mouse automatically, or pass `species: "mouse"` to skip the human miss. Prefer this over querying each gene's full MGI phenotype list.
`MSigDB_get_gene_set_members` (operation `get_gene_set`) returns `{genes:[...]}`; `MSigDB_check_gene_in_set` (operation `check_gene_in_set`, param `gene`) returns `{is_member: bool}`.
`MSigDB_get_gene_set_members` (operation `get_gene_set`) returns `{genes:[...]}`; `MSigDB_check_gene_in_set` (operation `check_gene_in_set`, param `gene`) returns `{is_member: bool}`. Both report which `species` collection matched.
**Fetch the set once, not once per option.** A multiple-choice question asks about one set and 3-5 candidates, so `MSigDB_get_gene_set_members` answers all of them in a single call — compare the options against the returned list yourself. Reserve `MSigDB_check_gene_in_set` for a single-gene question, or when the set is too large to return comfortably.
## Gene–disease "in database X but NOT database Y" recipe
@@ -67,9 +71,59 @@ These questions (e.g. "which gene is associated with disease D according to DisG
5. **Elimination:** rule out options that ARE OMIM-causal for D; among the rest, pick the one with a DisGeNet/text-mined association. If exactly one option is non-OMIM and has any association signal, that is the answer.
6. Only answer "Insufficient information" if no option has any association in any source. If the gold gene appears in neither curated DisGeNet, OMIM, nor PubTator literature, it may rely on a DisGeNet-internal text-mined signal the academic tier can't reach — say so honestly rather than guessing.
## Mouse-phenotype matching (MGI)
## Drug choice for a described patient (clinical vignette)
`MGI_get_phenotypes` returns a list of `phenotype_statement` strings per gene. To answer "which gene is annotated to phenotype P" (e.g. an MP term like *increased carcinoma incidence*), query each candidate gene and pick the one whose statements include a phrase matching P (the statements are human-readable, e.g. "increased incidence of carcinoma", "tumor"). Match on the phenotype concept, not an exact MP id string. If several match, prefer the most specific statement.
"Which of the following is most appropriate for this patient?" with a vignette
naming a **modifier** — hepatic or renal impairment, a Child-Pugh class, a
concomitant strong CYP3A4 inhibitor, pregnancy, an allergy or contraindication —
and several candidate drugs. These read like clinical-judgement questions, but
the modifier is doing all the work and the deciding fact is printed in each
candidate's FDA label. Answering from recall is the failure mode here: the
options are usually all plausible drugs for the condition, and only the label
separates them.
**1. Resolve every option brand → generic first.**
`FDA_get_active_ingredient_info_by_drug_name` on each option. Do this before any
reasoning, for two reasons: label lookups are keyed on the ingredient, and **two
options are sometimes the same drug** under a brand and a generic name. When
that happens neither can be the intended answer — they cannot be
distinguished — so it eliminates both and often decides the question outright.
Note the duplicate explicitly; it is also worth reporting as a benchmark defect.
**2. Look up only the section the modifier turns on.** One targeted call per
candidate beats pulling whole labels:
| The vignette says… | Read this section |
|---|---|
| hepatic impairment, Child-Pugh A/B/C, cirrhosis | `FDA_get_pharmacokinetics_by_drug_name` (hepatic-impairment subsection), then `FDA_get_dosage_and_storage_information_by_drug_name` for the adjustment. A Child-Pugh grading may sit in either of those or in contraindications, and some labels describe hepatic impairment without using the term at all — absence from one section is not absence from the label |
| renal impairment, CrCl/eGFR, dialysis | same pair — PK first, then dosage |
| "on a strong CYP3A4 inhibitor/inducer", any named co-medication | `FDA_get_drug_interactions_by_drug_name`; `FDA_get_clinical_pharmacology_by_drug_name` when the label states the metabolic pathway rather than the pairing |
| pregnancy, breastfeeding, "planning to conceive" | `FDA_get_pregnancy_or_breastfeeding_info_by_drug_name` (`FDA_get_teratogenic_effects_by_drug_name` when the question is about fetal harm specifically) |
| an allergy, a comorbidity that rules a drug out | `FDA_get_contraindications_by_drug_name`, then `FDA_get_boxed_warning_info_by_drug_name` |
| elderly / pediatric patient | `FDA_get_geriatric_use_info_by_drug_name` / `FDA_get_pediatric_use_info_by_drug_name` |
| the drug simply may not treat the condition | `FDA_get_indications_by_drug_name` |
The naming is regular — `FDA_get_<section>_by_drug_name` — so a section not listed
here can be found by searching the catalog for the section name rather than
guessing a tool name.
**3. Decide by elimination, and say what eliminated each option.** The intended
answer is normally the one candidate the modifier does *not* exclude:
contraindicated in hepatic impairment, requires an unavailable dose reduction,
interacts with the stated co-medication, or is not indicated for the condition.
Quote the label phrase that rules each option out — a vignette answer without a
cited label sentence is a guess wearing a citation.
**Do not over-query.** Resolve the ingredients (one call per option), then read
one section per remaining candidate. If the label is silent on the modifier for
every option, say so and answer on indication — do not keep pulling sections
hoping for a discriminator.
## Mouse-phenotype matching (MGI) — fallback only
Try the `MP_<PHENOTYPE>` MSigDB set first (above): it answers in one call per option and is the same MGI annotation. Use this per-gene route only when the set name does not resolve.
`MGI_get_phenotypes` returns a list of `phenotype_statement` strings per gene, **paginated** — a gene's matching statement is often on a later page, so a single page is not evidence of absence. To answer "which gene is annotated to phenotype P", query each candidate gene and pick the one whose statements include a phrase matching P (the statements are human-readable, e.g. "increased incidence of carcinoma", "tumor"). Match on the phenotype concept, not an exact MP id string. If several match, prefer the most specific statement.
## Computational procedures (when the answer is COMPUTED, not looked up)
@@ -1,8 +1,8 @@
---
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse-claude-code-plugin/SKILL.md
upstream_sha: e2520a96
imported_at: 2026-06-26
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse-claude-code-plugin/SKILL.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: catalogue
upstream_changes: accepted
name: tooluniverse-claude-code-plugin
@@ -21,7 +21,7 @@ uv --version # must exist; if not: curl -LsSf https://astral.sh/uv/install
claude --version # Claude Code CLI; if not: https://claude.com/claude-code
```
## Install (two commands)
## Install
```bash
# 1. Register the ToolUniverse marketplace from GitHub
@@ -33,6 +33,29 @@ claude plugin install tooluniverse@tooluniverse
That's it. Restart Claude Code. The MCP server auto-starts via `uvx tooluniverse` on first use (~30 s cold start, instant after).
### Recommended: turn on auto-update
Third-party marketplaces default to **no auto-update** — without this, new tools/skills only reach you when you remember to run `claude plugin update tooluniverse` (see Update below). Turn it on once:
```bash
python3 -c "
import json, pathlib, sys
p = pathlib.Path.home() / '.claude/plugins/known_marketplaces.json'
d = json.loads(p.read_text())
if 'tooluniverse' not in d:
sys.exit('Run the marketplace add command above first')
d['tooluniverse']['autoUpdate'] = True
p.write_text(json.dumps(d, indent=2))
print('autoUpdate enabled for tooluniverse')
"
```
Equivalent interactive path: `/plugin` → Marketplaces → `tooluniverse` → Enable auto-update.
With this on, Claude Code checks for marketplace + plugin updates in the background after each session start (up to a ~10 min random delay) and updates the installed plugin on disk automatically. You'll get a `/reload-plugins` prompt when an update lands, or it applies on your next launch — no more manual `claude plugin update`.
This is local, per-machine state — it can't be shipped as a default from the plugin's own manifest. `marketplace.json` has no `autoUpdate` field; Claude Code intentionally keeps this a per-installation trust boundary so a publisher can't force silent auto-updates onto a user's machine.
### Important: Remove global skills if previously installed
If you previously installed ToolUniverse skills globally (via `tooluniverse-install-skills` or manual copy), **remove them**. The plugin includes all skills — global copies interfere with the plugin's skill routing.
@@ -82,6 +105,9 @@ The router skill auto-dispatches to the right specialized skill — no command p
| **`/tooluniverse:cross-validate`** | Verify a claim across 3+ independent databases | Slash command |
| **`/tooluniverse:compare`** | N-way side-by-side comparison with domain-appropriate columns | Slash command |
| **`/tooluniverse:literature-sweep`** | Graded mini-review across PubMed + EuropePMC + Semantic Scholar | Slash command |
| **`/tooluniverse:verify-references`** | Check that cited references are real and accurately described, including retraction status | Slash command |
| **`/tooluniverse:self-review`** | Review current or supplied work against its actual goal; qualitative by default, with scoring only when explicitly requested | Slash command |
| **`/tooluniverse:setup-keys`** | Configure ToolUniverse API keys | Slash command |
| **`/tooluniverse:researcher`** | Same investigation as `research`, delegated to a forked subagent | Slash command |
| **120+ skills** | Structured workflows (drug research, variant interpretation, pharmacovigilance, CRISPR screens, statistical modeling, etc.) | Auto-activate on matching questions |
@@ -121,6 +147,10 @@ Full API-key list: `setup-tooluniverse` skill → `API_KEYS_REFERENCE.md`.
## Update
If you enabled auto-update above, this happens automatically in the background — no action needed.
Otherwise, update manually:
```bash
claude plugin update tooluniverse
# Also refresh the MCP server's tool cache:
@@ -0,0 +1,185 @@
---
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse-gene-liability/SKILL.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: unknown
upstream_changes: accepted
name: tooluniverse-gene-liability
description: Evaluate the human safety liability of knocking down, knocking out, degrading, or pharmacologically inhibiting a gene. Use for gene safety scoring, on-target toxicity assessment, essentiality and genetic-constraint review, critical-organ expression analysis, or deciding whether a target needs partial, transient, or tissue-specific modulation.
---
# Gene Liability Evaluation
Assess whether reducing a human gene's function is likely to be unsafe. Resolve the
gene, gather independent human and model-system evidence, calculate a transparent
0–100 liability score, and recommend an appropriate modulation strategy.
Higher scores mean greater predicted liability from the stated intervention. This is
a safety score, not a target-efficacy or druggability score.
## Required Input
Accept a gene symbol, name, Ensembl ID, or UniProt accession. Also capture:
- intervention modality: knockout, knockdown, degrader, irreversible inhibitor,
reversible inhibitor, antibody, or unknown;
- intended tissue and indication, when supplied;
- desired inhibition depth and duration, when supplied.
If modality is absent, assess complete systemic loss of function and label that
assumption prominently. Do not silently generalize a knockout result to partial or
tissue-restricted pharmacology.
## Evidence Rules
- Look up the gene before reasoning. Do not score an ambiguous identifier.
- Use human causal evidence before animal models, screens, expression, or prediction.
- Treat absent data as unknown, never as evidence of safety.
- Cite every score-changing observation with the source and identifier.
- Separate germline loss of function, somatic loss, acute pharmacology, and chronic
pharmacology.
- Treat DepMap as cancer-cell essentiality, not normal-tissue essentiality.
- Report contradictory evidence instead of averaging it away.
Grade evidence as:
- **T1**: human clinical outcome or replicated causal human genetics;
- **T2**: curated human evidence, mammalian knockout phenotype, or established drug
class safety;
- **T3**: functional screen, tissue expression, or single experimental study;
- **T4**: computational prediction or catalog annotation.
## Workflow
### 1. Resolve the gene
Use `MyGene_query_genes` with `species="human"` and request symbol, name, Ensembl,
UniProt, and Entrez identifiers. Require an exact symbol or identifier match. If the
query maps to multiple loci, stop and ask for clarification.
Carry the approved symbol and Ensembl gene ID through every subsequent query.
### 2. Gather the five scoring dimensions
Run independent paths. A failed path must use its fallback or be marked unavailable.
| Dimension | Weight | Primary evidence | Fallback |
|---|---:|---|---|
| Human genetic constraint | 25 | `gnomad_get_gene_constraints(gene_symbol=...)` | ClinVar loss-of-function variants and literature |
| Mammalian knockout phenotype | 25 | `OpenTargets_get_biological_mouse_models_by_ensemblID(ensemblId=...)` | MGI-focused literature search |
| Critical-organ expression | 20 | `GTEx_get_median_gene_expression(operation="get_median_gene_expression", gencode_id=...)` | `HPA_get_comprehensive_gene_details_by_ensembl_id(..., include_expression=true)` |
| Observed on-target effects | 20 | `OpenTargets_get_target_safety_profile_by_ensemblID(ensemblId=...)` | `ClinVar_search_variants`, associated drugs, and literature |
| Cellular essentiality and redundancy | 10 | `DepMap_get_gene_dependencies(gene_symbol=...)` | pathway, paralog, and functional-screen literature |
Also query `OpenTargets_get_associated_drugs_by_target_ensemblID` to distinguish
observed target-class toxicity from hypothetical risk. Do not interpret the mere
existence of a drug as proof of safety.
For expression, inspect heart, central nervous system, liver, kidney, lung, immune or
marrow compartments, and reproductive tissues. Compare the gene across tissues;
do not compare raw TPM values between unrelated genes as if they shared one threshold.
### 3. Assign dimension points
Use only evidence actually retrieved.
#### Human genetic constraint — 0 to 25
- **25**: strong loss-of-function intolerance, such as pLI at least 0.9 together
with LOEUF or observed/expected LoF at most 0.35;
- **15**: one strong constraint signal or intermediate LoF constraint;
- **5**: weak or conflicting constraint;
- **0**: credible tolerance to loss of function.
Use LOEUF when available; label observed/expected LoF as a proxy when LOEUF is absent.
#### Mammalian knockout phenotype — 0 to 25
- **25**: embryonic or perinatal lethality, or severe multisystem phenotype;
- **18**: reduced survival, organ failure, severe neurologic, immune, reproductive,
or developmental phenotype;
- **8**: viable knockout with a consequential but organ-limited phenotype;
- **0**: replicated viable knockout without a consequential phenotype.
#### Critical-organ expression — 0 to 20
- **20**: broad high expression involving at least three critical organ systems;
- **12**: high expression in one or two critical organs or broad moderate expression;
- **6**: low-to-moderate critical-organ expression;
- **0**: credible restriction to the intended or noncritical tissue with negligible
critical-organ expression.
#### Observed on-target effects — 0 to 20
- **20**: severe human disease from reduced function or consistent serious
target-related toxicity;
- **12**: a credible human adverse phenotype or reproducible class effect;
- **5**: preclinical, isolated, or mechanistically plausible safety signal;
- **0**: human protective loss-of-function or well-tolerated target modulation with
no serious on-target signal at relevant exposure.
Protective variants can reduce concern only when their direction, dosage, tissue,
and lifelong-versus-acute exposure are relevant to the proposed intervention.
#### Cellular essentiality and redundancy — 0 to 10
- **10**: broad common-essential signal with little credible redundancy;
- **5**: context-selective dependency or partial redundancy;
- **0**: reproducible non-essentiality or strong functional redundancy.
If DepMap returns only gene metadata without dependency scores, mark this dimension
unavailable. Never infer essentiality from a successful lookup alone.
### 4. Calculate score, coverage, and confidence
For available dimensions, calculate:
`liability score = 100 × points earned / available weight`
Report the available weight as evidence coverage. Do not assign zero points to a
missing dimension.
- **0–24**: low liability;
- **25–49**: moderate liability;
- **50–74**: high liability;
- **75–100**: very high liability.
Publish a categorical score only when coverage is at least 60%. Below 60%, report
“insufficient evidence” and list the experiments or datasets needed.
Assign confidence independently:
- **High**: at least 90% coverage with both T1 and independent T2 evidence;
- **Moderate**: at least 70% coverage with T1 or T2 evidence;
- **Low**: 60–69% coverage or evidence dominated by T3/T4;
- **Insufficient**: below 60% coverage.
### 5. Translate liability into a strategy
Do not stop at a risk label. Explain whether the evidence favors:
- partial rather than complete inhibition;
- reversible or transient rather than irreversible modulation;
- tissue-targeted delivery;
- isoform- or domain-selective modulation;
- biomarker-based exclusion or monitoring;
- deprioritization until a specific safety experiment closes the key gap.
## Required Output
Return these sections:
1. **Resolved gene and intervention assumption** — identifiers, modality, tissue,
inhibition depth, and duration.
2. **Liability verdict** — score, band, evidence coverage, and confidence.
3. **Dimension table** — evidence, points, maximum weight, tier, citation, and
conflicts for all five dimensions.
4. **Key red flags and protective evidence** — list both sides explicitly.
5. **Modulation recommendation** — full, partial, transient, tissue-specific, or
no-go, with rationale.
6. **Data gaps and next experiments** — prioritize the missing evidence most likely
to change the verdict.
7. **Sources** — database record links, study identifiers, and access dates.
End with: “This is a research risk assessment, not a clinical safety determination.”
@@ -1,12 +1,12 @@
---
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse-molecular-cloning/SKILL.md
upstream_sha: e2520a96
imported_at: 2026-06-26
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse-molecular-cloning/SKILL.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: unknown
upstream_changes: accepted
name: tooluniverse-molecular-cloning
description: Molecular cloning assembly design — Gibson Assembly (overlap design for seamless multi-fragment joining) and Golden Gate Assembly (Type IIS / BsaI / BbsI design with unique 4-bp fusion overhangs). Use when you need to plan how to join DNA fragments into a construct, design assembly overlaps/overhangs, or decide between cloning methods. Covers the domestication (internal-site removal), overhang-uniqueness, and overlap-Tm rules. For PCR primers to generate the fragments, see tooluniverse-primer-design.
description: Molecular cloning, in both directions. DESIGN — Gibson Assembly (overlap design for seamless multi-fragment joining) and Golden Gate Assembly (Type IIS / BsaI / BbsI / Esp3I / BsmBI / SapI design with unique 4-bp fusion overhangs). ANALYSIS — work out what an existing reaction produces: given input plasmid sequences and an enzyme, digest them, join the fragments by their overhangs, and identify features of the product (expressed ORF, gRNA spacer and its target gene). Use when you need to plan how to join DNA fragments into a construct, design assembly overlaps/overhangs, decide between cloning methods, or determine the product of a stated Gibson/Golden Gate reaction. Covers the domestication (internal-site removal), overhang-uniqueness, and overlap-Tm rules. For PCR primers to generate the fragments, see tooluniverse-primer-design.
disable-model-invocation: true
---
@@ -69,9 +69,47 @@ Returns `parts_with_overhangs`: each part's unique 4-bp `left_overhang`/`right_o
- **Overlap Tm imbalance** (Gibson) → some junctions form, others don't.
- **Generating the fragments** still needs primers with the overlaps/overhangs appended — design and QC those in `tooluniverse-primer-design` (and BLAST for specificity).
## Working backwards — what does an existing assembly produce?
The reverse question ("I combined these plasmids in a Golden Gate reaction with
Esp3I — what does the product express / what does the gRNA target?") is answered by
one call. Do **not** hand-write a digestion/ligation simulator.
```bash
tu run DNA_golden_gate_assemble '{"fragments":["<plasmid1>","<plasmid2>","<plasmid3>"],
"enzyme":"Esp3I","labels":["pLAB-CTU","pLAB-gTU2E","pLAB-CH3"]}'
```
It digests each input, drops the fragments that keep a recognition site (those are
re-cut in the reaction and cannot persist), chains the rest by matching 4-bp
overhangs, and returns `product_sequence`, `product_length` and the `assembly_order`
with the overhang at every junction. Inputs are treated as circular plasmids unless
you pass `circular: false`.
Then **annotate the product**. Locate features in `product_sequence` (promoter, ORF,
gRNA spacer). For a gRNA cassette the spacer is the ~20 nt immediately 5′ of the
scaffold (`GTTTTAGAGCTAGAAATAGCAAG`); identify its target by matching that spacer
against the genome (`BLAST_*`, or an Ensembl/NCBI/SGD sequence lookup) and check for
an adjacent PAM. Match the **species implied by the construct** — yeast tRNA/Pol III
parts mean search the yeast genome, not human.
If the assembly reports that the fragments do not chain, digest the inputs
individually with `DNA_virtual_digest` (`circular: true`) to see what each released:
a Golden Gate donor carries its two Type IIS sites **inverted** around the insert, so
a correct digest gives **2 fragments** per plasmid. Getting 1 means the enzyme name or
`circular` is wrong — not that the plasmid lacks sites.
Enzyme names: `Esp3I` and `BsmBI` are the same enzyme (`CGTCTC`); `BsaI`
(`GGTCTC`), `BbsI` (`GAAGAC`) and `SapI` (`GCTCTTC`) all resolve too.
## Honest limitations
- These tools design the assembly junctions; they do not simulate the full ligation/exonuclease reaction or guarantee efficiency — validate by sequencing the assembled construct.
- Digestion and overhang-driven ligation are simulated faithfully (`DNA_virtual_digest`
and `DNA_golden_gate_assemble` cut both strands at each enzyme's real offset, including
Type IIS enzymes that cut outside their site). What is *not* modelled is reaction
efficiency — overhang ligation bias, partial digestion, incorrect-but-possible
junctions — so a returned product is the intended assembly, not a yield prediction.
Validate by sequencing the assembled construct.
- No vector-backbone or ORF-frame checking — confirm reading frame and backbone compatibility yourself.
## Related skills
@@ -1,8 +1,8 @@
---
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse-phylogenetics/SKILL.md
upstream_sha: e2520a96
imported_at: 2026-06-26
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse-phylogenetics/SKILL.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: catalogue
upstream_changes: accepted
name: tooluniverse-phylogenetics
@@ -154,7 +154,6 @@ that line when the raw min in a group is 0.
scogs zips ship in two shapes:
- **Full**: `<gene>.faa`, `<gene>.faa.mafft`, `<gene>.faa.mafft.clipkit`,
`<gene>.faa.mafft.clipkit.treefile`, plus iqtree/bionj/log/mldist.
Use clipkit alignment + treefile for tree-paired metrics.
- **Alignment-only**: just `<gene>.faa` + `<gene>.faa.mafft`. No
trees, no clipkit. Used for parsimony, RCV, gap-percentage
questions. Use the `.faa.mafft` (NOT raw `.faa`) — the published
@@ -164,6 +163,47 @@ Both bundled scripts auto-detect the layout and use the best available
alignment per ortholog. Do NOT re-run MAFFT or ClipKit yourself; the
shipped files are canonical.
#### Which alignment goes with which metric (this changes the answer)
The tree is always the ClipKit-derived `.faa.mafft.clipkit.treefile`. The
**alignment** argument depends on the metric:
| metric | alignment to pass |
|---|---|
| `treeness`, `dvmc`, `total_tree_length`, `long_branch_score` | tree only — no alignment |
| `saturation` | `.faa.mafft.clipkit` (trimmed) |
| **`treeness_over_rcv` / `rcv`** | **`.faa.mafft` (untrimmed)** |
| parsimony-informative sites, gap percentage | `.faa.mafft` (untrimmed) |
RCV measures compositional variability **across the alignment's columns**, so
trimming changes it materially — and `treeness_over_rcv` divides by RCV, so the
trimmed alignment shifts the ratio for every gene. Verified on the fungal scogs
set (249 orthologs, canonical shipped files):
```
median treeness/RCV untrimmed .faa.mafft = 0.2683 trimmed .clipkit = 0.3050
max treeness/RCV (>70% gap genes)
untrimmed .faa.mafft = 0.1866 trimmed .clipkit = 0.4205
```
Plain `treeness` needs no alignment and is unaffected — it reproduces exactly
(median 0.0501 on the same 249 files), which is how the alignment choice was
isolated as the cause rather than the tree set or the tool.
`phykit_batch_analysis` takes the two independently, so pass them explicitly:
```bash
tu run phykit_batch_analysis '{"operation":"batch","function":"treeness_over_rcv",
"directory":"<dir>","extension":".faa.mafft",
"tree_directory":"<dir>","tree_extension":".faa.mafft.clipkit.treefile"}'
```
**Gap percentage** in these questions means the fraction of alignment
**columns containing at least one gap**, not the fraction of all residues that
are gaps. The two differ by an order of magnitude: with the residue definition
no fungal ortholog exceeds 70% gaps, so a ">70% gaps" filter silently selects
nothing.
**Anti-pattern:** running `phykit` on the raw `*.busco.zip` extracted
ortholog FASTAs and aligning/tree-building yourself. The pre-computed
files in `scogs_*.zip` are the canonical inputs.
@@ -0,0 +1,165 @@
---
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse-self-review/SKILL.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: unknown
upstream_changes: accepted
name: tooluniverse-self-review
description: >
Review existing work against the user's actual goal and surface evidence-backed
strengths, gaps, risks, and next fixes. Use when asked to eval, evaluate, review,
assess, or check current/this/my/our work; decide whether a task is complete; build a
definition-of-done checklist or rubric; or perform grading, LLM-as-judge, Qworld, or
RET evaluation. Treat plain eval/review requests as qualitative: resolve "current
work" from the conversation, artifacts, files, or diff, and never assign numeric
scores unless the user explicitly requests scores, grades, points, ratings, weighted
criteria, Qworld, or RET. Do not use for implementing automated eval suites, tests,
graders, or benchmarks.
disable-model-invocation: true
---
# Self-Review: Understand the Target, Then Review
Review the actual work the user means against the goal it was meant to satisfy. Default
to a concise qualitative assessment. Scoring and the full Qworld Recursive Expansion
Tree (RET) are opt-in.
## Non-Negotiable Rules
1. **Do not confuse the evaluation request with the evaluated task.** In requests such as
"eval current work", that sentence is an instruction to review. The task being judged
is the preceding user goal; the work is the current result, implementation, draft,
plan, or progress.
2. **Do not equate `eval` with scoring.** `Eval`, `evaluate`, `review`, `assess`, and
`check` mean qualitative review unless the user explicitly asks for a score, grade,
rating, points, weighted rubric, Qworld, RET, or a numeric scale.
3. **Do not invent work to review.** Inspect the conversation and available artifacts. If
evidence is unavailable, say what could not be verified.
4. **Do not expose process by default.** Keep scenario expansion, perspective generation,
and rubric construction internal unless the user asked for those artifacts.
5. **Review against the user's goal, not a generic quality template.** Derive the relevant
checks from the request, stated constraints, acceptance criteria, and risks.
## Resolve the Review Target
Separate three objects before reviewing:
- **Evaluation instruction**: what the user is asking now (for example, "review this").
- **Original goal**: the request, problem, or acceptance criteria the work should satisfy.
- **Work product**: the answer, code, files, diff, plan, analysis, or current progress to
examine.
Resolve the work product in this order:
1. An explicitly named file, answer, commit, diff, section, or artifact.
2. Pasted or attached content in the current message.
3. The current repository implementation or working-tree diff when the conversation is
about code changes.
4. The most recent assistant-produced deliverable relevant to the preceding user goal.
5. The current plan or partial progress when the task is still underway.
Interpret deictic phrases such as "current work", "this work", "what we have", "刚才的
工作", and "当前工作" using that order. In a multi-turn conversation, the last user
message is usually the evaluation instruction, **not** the original goal.
Proceed without asking when the goal and work can be recovered confidently. Ask one
short clarifying question only when there is no reviewable work or when multiple plausible
targets would produce materially different reviews.
## Choose One Mode
| Mode | Trigger | Default output |
|---|---|---|
| **Qualitative review** (default) | "eval/review/check current work", "is this done?", "what is missing?" | Evidence-backed findings, strengths, gaps, fixes, and a completion verdict; no numbers |
| **Checklist** | Explicit request for definition of done, success criteria, or completeness checklist without work to review | Task-specific checklist; no points |
| **Rubric** | Explicit request for evaluation criteria or a rubric, but no scoring request | Binary or observable criteria grouped as must/should/could; no points |
| **Scored evaluation** | Explicit request for score, grade, rating, points, weighted criteria, LLM-as-judge scoring, Qworld, RET, or a numeric scale | Evidence-backed scored review using the requested scale or RET |
The phrase "evaluate this" alone selects **qualitative review**, even when a work product
is present. The presence of work never turns scoring on by itself.
Requests to **create or run evals**, implement a grader, write evaluation tests, or build a
benchmark are engineering tasks, not self-review requests. Do not route those requests to
this workflow merely because they contain the word `eval`.
## Qualitative Review Workflow (Default)
1. **Recover the goal.** Summarize the original goal and important constraints in one or
two sentences. Prefer explicit acceptance criteria over inferred preferences.
2. **Inspect the work.** Use the actual conversation output, files, diff, test results, or
supplied artifact. For code, inspect relevant implementation and verification evidence;
do not judge from a summary alone when the files are available.
3. **Derive focused checks internally.** Identify only the task-specific dimensions needed
to judge correctness, completeness, user intent, risks, and verification. Do not print a
large rubric unless asked.
4. **Report findings by impact.** Lead with concrete problems or unmet requirements. For
each finding, cite the evidence and explain the consequence.
5. **Acknowledge what works.** Note meaningful strengths briefly; do not pad the response
with generic praise.
6. **Give prioritized fixes.** Recommend the smallest concrete actions that close the most
important gaps.
7. **State a plain-language verdict.** Use `complete`, `mostly complete`, `partially
complete`, `not complete`, or `unable to verify`, with a short reason. Do not convert
the verdict into a number.
### Default Review Shape
Adapt the headings to the task and omit empty sections:
1. **Overall assessment** — target, goal, and verdict.
2. **Findings** — ordered by impact, with evidence.
3. **What is working** — concise, specific strengths.
4. **Recommended next actions** — prioritized fixes.
For code review, prioritize actionable defects and regressions over summaries. Cite file
paths and tight line ranges when possible. If no problems are found, say so directly and
name any residual verification gaps.
## Checklist and Unscored Rubric Modes
- Derive criteria from the task rather than a fixed dimension list.
- Keep each item observable and specific enough to check.
- Use `must`, `should`, and `could` for importance when prioritization helps.
- Do not attach numbers, weights, percentages, earned totals, or pass rates.
- If work is also supplied, mark items `met`, `partially met`, `not met`, or
`not verifiable`, with brief evidence. These labels are not scores.
- Do not generate scenarios or perspectives unless the user explicitly asks to see the
derivation.
## Scored Evaluation Mode (Explicit Opt-In Only)
Use this mode only when the request contains an unambiguous scoring signal listed above.
- If the user supplies a scale or grading scheme, follow it.
- If the user asks for Qworld, RET, a weighted rubric, or LLM-as-judge scoring without a
custom scheme, read and follow
[references/ret-scored-evaluation.md](references/ret-scored-evaluation.md).
- Keep every verdict evidence-based. Do not award credit for absent evidence or penalize
criteria outside the original goal.
- Explain what the score means and still provide the highest-impact gaps and fixes. A
number alone is not a useful review.
## Examples of Correct Routing
| User request | Correct interpretation |
|---|---|
| "Eval current work." | Review the current result against the preceding goal; qualitative, no score |
| "Evaluate whether we finished the original request." | Inspect current artifacts and give a completion verdict; no score |
| "What is missing from this implementation?" | Findings-first implementation review; no score |
| "Make a definition-of-done checklist for this feature." | Checklist mode; no score |
| "Create an evaluation rubric for these answers." | Unscored rubric unless weights or grading are requested |
| "Score this answer from 1 to 10." | Scored mode on the requested scale |
| "Apply Qworld/RET to grade these responses." | Full scored RET mode |
| "Create an eval suite for this agent." | Out of scope for this skill; treat as an eval-engineering task |
## Evidence and Honesty
- Distinguish observed evidence from inference.
- Do not claim tests passed unless their output is available.
- Do not infer completion from a clean diff, a confident summary, or the existence of
files alone.
- If the user requests evaluation while work is still running, review the available
progress and label unfinished parts rather than pretending the final result exists.
- Keep the response proportional to the work. A small change should not produce a giant
framework dump.
@@ -0,0 +1,131 @@
---
title: "Qworld RET Scored Evaluation"
task: ""
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse-self-review/references/ret-scored-evaluation.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: prompt
upstream_changes: accepted
author: upstream
validated: false
---
# Qworld RET Scored Evaluation
Use this reference only after the user explicitly requests scoring, grading, weighted
criteria, LLM-as-judge scoring, Qworld, or the Recursive Expansion Tree (RET).
## Inputs
- **Original goal**: the task, question, constraints, and acceptance criteria.
- **Work product**: the actual answer, implementation, artifact, or result to grade.
- **Scale**: the user's requested scheme, or the default weighted scheme below.
Never treat the evaluation instruction itself as the original goal. If the user says
"score current work", recover the goal and work from the preceding conversation and
available artifacts first.
## Recursive Expansion Tree
Build the rubric from the task:
```text
Original goal
-> scenarios that materially change what good means
-> task-specific evaluation perspectives
-> concrete, binary criteria
```
Keep this derivation internal unless the user asks to see it.
### 1. Scenario grounding and expansion
Identify a minimal non-redundant set of real contexts in which the task could arise and
where the context changes what constitutes a good result. Expand three times by asking
what materially different audience, setting, stakes, constraints, or domain variation is
missing. Do not add paraphrases of existing scenarios.
### 2. Perspective generation and expansion
For each scenario, derive evaluation dimensions from the task itself. Do not use a fixed
dimension list. Expand four times by asking which distinct evaluation angle would yield
new criteria. Consolidate overlap and assign perspective IDs (`p0`, `p1`, ...).
### 3. Criteria generation and expansion
For each retained perspective, write self-contained criteria that are:
- answerable `YES` or `NO`;
- specific to the task and scenario;
- observable in the work product;
- non-redundant; and
- phrased as one concrete required or forbidden behavior.
Expand three times by asking what additional observable behavior would materially change
the verdict. Then merge overlap and assign criterion IDs (`c0`, `c1`, ...).
Include negative criteria only for harmful, misleading, or materially quality-reducing
behavior—not minor style preferences. Phrase negative criteria as the bad behavior itself;
the negative point value supplies the polarity.
## Default Weighted Scheme
Use this scheme only when the user did not provide another one:
- Positive `1–10`: `10` critical safety/core requirement; `8–9` important completeness;
`5–7` meaningful quality; `1–4` minor enhancement.
- Negative `-1–-10`: `-10` dangerous; `-8–-9` major error; `-5–-7` material quality
problem; `-1–-4` minor problem.
For each criterion provide `criterion_id`, `criterion`, `points`, and a concise reason for
the weight. Check that:
1. desirable criteria have positive signs and harmful behaviors have negative signs;
2. more important criteria have larger absolute weights;
3. positive and negative criteria do not duplicate the same requirement; and
4. total positive points exceed total negative magnitude.
## Apply the Rubric
For every criterion:
1. Mark `YES` or `NO`.
2. Cite the specific evidence or location supporting the verdict.
3. Sum points only for criteria marked `YES`; positive items add and negative items
subtract.
Report:
- earned positive points and maximum positive points;
- negative penalties triggered;
- net total and the scale interpretation;
- the most important unmet positive criteria and triggered negative criteria; and
- one concrete fix for each high-impact gap.
If there is no work product, do not fabricate a score. Return a weighted rubric and state
that no work was graded.
## Output Discipline
- Show scenarios and perspectives only if the user asks for the full derivation.
- Keep reasoning proportional; do not dump expansion logs.
- Prefer evidence and actionable gaps over score theater.
- If evidence is unavailable, mark the criterion `NO` or `not verifiable` according to the
user's grading policy and disclose the limitation.
## Citation
This method is based on Qworld:
```bibtex
@misc{gao2026qworldquestionspecificevaluationcriteria,
title={Qworld: Question-Specific Evaluation Criteria for LLMs},
author={Shanghua Gao and Yuchang Su and Pengwei Sui and Curtis Ginder and Marinka Zitnik},
year={2026},
eprint={2603.23522},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.23522},
}
```
@@ -1,8 +1,8 @@
---
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse/SKILL.md
upstream_sha: e2520a96
imported_at: 2026-06-26
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/cfd26718/skills/tooluniverse/SKILL.md
upstream_sha: cfd26718
imported_at: 2026-08-08
prompt_class: skill
upstream_changes: accepted
name: tooluniverse
@@ -108,10 +108,12 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
| Keywords | Action |
|----------|--------|
| **A single database-checkable fact**, especially multiple-choice — "which of the following gene/drug/variant/pathway is…", anything phrased "**according to** <database>" (DisGeNet, OMIM, MSigDB, miRDB, GTRD, MGI, Ensembl, ClinVar, ChEMBL, OpenTargets, UniProt). Look it up rather than answering from memory; niche annotations are what gets hallucinated | `Skill(skill="tooluniverse-biomedical-fact-lookup")` |
| "research", "profile", "**disease**", "syndrome", "disorder", "comprehensive report on [disease]" | `Skill(skill="tooluniverse-disease-research")` |
| "research", "profile", "**drug**", "medication", "therapeutic agent", "tell me about [drug]" | `Skill(skill="tooluniverse-drug-research")` |
| "**literature review**", "papers about", "publications on", "research articles", "recent studies" | `Skill(skill="tooluniverse-literature-deep-research")` |
| "research", "profile", "**target**", "protein target", "gene target", "target validation" | `Skill(skill="tooluniverse-target-research")` |
| "**gene liability**", "gene safety score", "knockout safety", "knockdown safety", "on-target toxicity", "safe to inhibit [gene]" | `Skill(skill="tooluniverse-gene-liability")` |
| "**peptide target**", "**deorphanize**", "deorphanization", "**peptide off-target**", "what does [peptide] bind", "target of a peptide", "orphan peptide", "peptide doesn't bind [target]", "binds in [species] but not", "find the receptor for [peptide]" | `Skill(skill="tooluniverse-peptide-target-deorphanization")` |
### 3. Clinical Decision Support
@@ -130,8 +132,11 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
| "**TCGA**", "cancer genomics cohort", "GDC analysis", "TCGA mutations", "pan-cancer" | `Skill(skill="tooluniverse-cancer-genomics-tcga")` |
| "**immunotherapy response**", "checkpoint inhibitor response", "TMB", "MSI", "PD-L1", "ICI response" | `Skill(skill="tooluniverse-immunotherapy-response-prediction")` |
| "**rare disease diagnosis**", "differential diagnosis", "phenotype matching", "HPO", "patient with [symptoms]" | `Skill(skill="tooluniverse-rare-disease-diagnosis")` |
| "**clinical risk score**", "CHA2DS2-VASc", "HAS-BLED", "CURB-65", "qSOFA", "Child-Pugh", "MELD-Na", "Wells score", "ASCVD risk", "eGFR CKD-EPI", "bedside risk calculator" | `Skill(skill="tooluniverse-clinical-risk-scoring")` |
| "**device adverse events**", "device recall", "MAUDE", "food/supplement adverse event", "CAERS", "veterinary adverse event", "drug shortage" | `Skill(skill="tooluniverse-product-safety-surveillance")` |
| "**variant interpretation**", "VUS", "pathogenicity", "clinical significance", "is [variant] pathogenic" | `Skill(skill="tooluniverse-variant-interpretation")` |
| "**clinical guidelines**", "practice guidelines", "treatment guidelines", "dosing recommendations", "standard of care" | `Skill(skill="tooluniverse-clinical-guidelines")` |
| **Drug choice for a described patient** — "which of the following is most appropriate for", "which drug should this patient receive"; a vignette naming a modifier (**hepatic/renal impairment**, Child-Pugh, "on a strong CYP3A4 inhibitor", pregnancy, a contraindication) alongside candidate drugs. The deciding fact is in each candidate's FDA label, not recall | `Skill(skill="tooluniverse-biomedical-fact-lookup")` |
| "**patient stratification**", "precision medicine", "biomarker stratification", "treatment selection" | `Skill(skill="tooluniverse-precision-medicine-stratification")` |
### 4. Discovery & Design
@@ -149,6 +154,7 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
| "**small molecule discovery**", "chemical biology", "compound sourcing", "hit finding", "chemical probe" | `Skill(skill="tooluniverse-small-molecule-discovery")` |
| "**chemical sourcing**", "buy compound", "vendor search", "Enamine", "MolPort", "compound availability" | `Skill(skill="tooluniverse-chemical-sourcing")` |
| "**GPCR**", "G-protein coupled receptor", "GPCRdb", "receptor ligand", "biased agonist" | `Skill(skill="tooluniverse-gpcr-structural-pharmacology")` |
| "**dereplicate**", "natural product identification", "NPAtlas", "ChemOnt classification", "ClassyFire", "producing organism" | `Skill(skill="tooluniverse-natural-product-dereplication")` |
### 5. Genomics & Variant Analysis
@@ -165,6 +171,13 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
| "**regulatory variant**", "non-coding variant", "eQTL variant", "regulatory region variant" | `Skill(skill="tooluniverse-regulatory-variant-analysis")` |
| "**rare disease genomics**", "Orphanet gene", "rare disease gene", "causative gene", "exome diagnosis" | `Skill(skill="tooluniverse-rare-disease-genomics")` |
| "**1000 Genomes**", "IGSR", "population frequency", "superpopulation", "AFR/EUR/EAS/SAS/AMR" | `Skill(skill="tooluniverse-population-genetics-1000genomes")` |
| "**PheWAS**", "phenome-wide association", "cross-ancestry replication", "cross-biobank", "FinnGen", "BioBank Japan", "pleiotropy of a variant" | `Skill(skill="tooluniverse-phewas")` |
| "**Mendelian randomization**", "MR causal inference", "instrumental variable", "does X cause Y", "genetic causal evidence" | `Skill(skill="tooluniverse-mendelian-randomization")` |
| "**loss-of-function mechanism**", "LoF mechanism", "why is this variant LoF", "structural stability vs functional disruption" | `Skill(skill="tooluniverse-protein-lof-mechanism")` |
| "**SAE feature**", "sparse autoencoder variant", "ESMC SAE", "mechanistic variant interpretation" | `Skill(skill="tooluniverse-protein-sae-variant-interpretation")` |
| "**per-residue annotation**", "binding interface residues", "ligand pocket residues", "buried vs surface residues", "PDB structural annotation" | `Skill(skill="tooluniverse-protein-structural-annotation-pdb")` |
| "**why are these residues critical**", "residue functional mechanism", "DMS hotspot interpretation", "catalytic vs structural residue" | `Skill(skill="tooluniverse-residue-functional-mechanism-interpretation")` |
| "**validate variant predictor**", "DMS validation", "deep mutational scanning benchmark", "predictor vs experimental effect" | `Skill(skill="tooluniverse-variant-predictor-dms-validation")` |
### 6. Systems & Network Analysis
@@ -207,13 +220,14 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
| "**primer design**", "PCR primers", "qPCR primer", "melting temperature", "Tm calculation", "annealing temperature", "GC clamp", "primer-dimer", "oligo analysis", "amplicon", "forward and reverse primer" | `Skill(skill="tooluniverse-primer-design")` |
| "**diagnostic test**", "sensitivity specificity", "ROC curve", "AUC", "PPV", "NPV", "likelihood ratio", "Youden", "optimal cutoff", "post-test probability", "biomarker accuracy", "confusion matrix" | `Skill(skill="tooluniverse-diagnostic-test-evaluation")` |
| "**drug synergy**", "drug combination", "Bliss independence", "Loewe additivity", "HSA synergy", "ZIP score", "combination index", "Chou-Talalay", "synergistic antagonistic", "combination therapy analysis" | `Skill(skill="tooluniverse-drug-synergy")` |
| "**molecular cloning**", "Gibson Assembly", "Golden Gate", "Type IIS", "BsaI", "BbsI", "assembly overlap", "fragment assembly", "construct design", "domestication" | `Skill(skill="tooluniverse-molecular-cloning")` |
| "**molecular cloning**", "Gibson Assembly", "Golden Gate", "Type IIS", "BsaI", "BbsI", "Esp3I", "BsmBI", "SapI", "assembly overlap", "fragment assembly", "construct design", "domestication", **or the reverse direction** — "I combined these plasmids…", "what does the resulting plasmid express", "what does the gRNA target", "virtual digest", "restriction digest of a plasmid" | `Skill(skill="tooluniverse-molecular-cloning")` |
| "**metabolomics analysis**", "LC-MS analysis", "metabolite quantification", "metabolic flux" | `Skill(skill="tooluniverse-metabolomics-analysis")` |
| "**functional genomics screen**", "CRISPR library", "shRNA screen", "barcode screen" | `Skill(skill="tooluniverse-functional-genomics-screens")` |
| "**proteomics data**", "PRIDE", "MassIVE", "ProteomeXchange", "proteomics dataset" | `Skill(skill="tooluniverse-proteomics-data-retrieval")` |
| "**protein modification**", "PTM analysis", "phosphorylation site", "ubiquitination", "glycosylation" | `Skill(skill="tooluniverse-protein-modification-analysis")` |
| "**structural proteomics**", "cross-linking mass spec", "XL-MS", "HDX-MS", "structural biology" | `Skill(skill="tooluniverse-structural-proteomics")` |
| "**protein structure prediction**", "AlphaFold prediction", "structure modeling", "homology modeling" | `Skill(skill="tooluniverse-protein-structure-prediction")` |
| "**FASTQ QC**", "FastQC", "MultiQC", "adapter trimming", "fastp", "Cutadapt", "read quality", "sequence duplication" | `Skill(skill="tooluniverse-fastq-qc")` |
### 8. Clinical Trials & Study Design
@@ -237,6 +251,7 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
| "**ecology**", "biodiversity", "invasive species", "pollinator", "food web", "conservation", "community ecology", "trophic" | `Skill(skill="tooluniverse-ecology-biodiversity")` |
| "**microbiome**", "gut microbiota", "dysbiosis", "microbiome composition", "16S rRNA" | `Skill(skill="tooluniverse-microbiome-research")` |
| "**adverse outcome pathway**", "AOP", "key event", "molecular initiating event", "KER" | `Skill(skill="tooluniverse-adverse-outcome-pathway")` |
| "**genome assembly**", "assembly N50", "RefSeq assembly QC", "plasmid count", "NCBI Datasets genome" | `Skill(skill="tooluniverse-microbial-genome-characterization")` |
### 10. Specialized Biology
@@ -278,6 +293,7 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
| "**custom tool**", "add my own tool", "local tool", "create tool", "extend ToolUniverse" | `Skill(skill="tooluniverse-custom-tool")` |
| "**SDK**", "Python SDK", "build AI scientist", "programmatic access", "**import tooluniverse**", "**coding API**", "**tu build**", "**typed wrappers**" | `Skill(skill="tooluniverse-sdk")` |
| "**install skills**", "missing skills", "skill not found", "add skills" | `Skill(skill="tooluniverse-install-skills")` |
| "**self-review**", "eval current work", "evaluate this work", "check my work", "is this complete", "definition of done", "evaluation rubric", "success criteria", "grading criteria", "LLM-as-judge" | `Skill(skill="tooluniverse-self-review")` |
---
@@ -297,10 +313,15 @@ These reminders are for fast pattern recognition during routing. Detailed `❌ W
3. **Specificity Rule**: More specific beats general.
- "cancer treatment" → precision-oncology (not disease-research)
4. **Data Type Rule**: "get/retrieve/fetch" → retrieval skills.
4. **Evaluation Intent Rule**: Route requests to review existing/current work to
`tooluniverse-self-review`, but do not route requests to create or run an eval suite,
grader, test, or benchmark there. Those are implementation tasks. Within self-review,
plain "eval" is qualitative; scoring requires an explicit score/grade/points request.
5. **Data Type Rule**: "get/retrieve/fetch" → retrieval skills.
- "get compound structure" → chemical-compound-retrieval (not drug-research)
5. **Still ambiguous**: Ask user with AskUserQuestion.
6. **Still ambiguous**: Ask user with AskUserQuestion.
---