Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7ed86c83fb | ||
|
|
2b4e62829f |
@@ -2,9 +2,9 @@
|
||||
title: "Scientific Agent Skills"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/991bd993/README.md
|
||||
upstream_sha: 991bd993
|
||||
imported_at: 2026-08-08
|
||||
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/13385c7c/README.md
|
||||
upstream_sha: 13385c7c
|
||||
imported_at: 2026-08-13
|
||||
prompt_class: catalogue
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
@@ -14,10 +14,11 @@ validated: false
|
||||
# Scientific Agent Skills
|
||||
|
||||
[](LICENSE.md)
|
||||
[](pyproject.toml)
|
||||
[](#-whats-included)
|
||||
[](pyproject.toml)
|
||||
[](#-whats-included)
|
||||
[](#-whats-included)
|
||||
[](https://agentskills.io/)
|
||||
[](https://agent-plugins.org/)
|
||||
[](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/security-scan.yml)
|
||||
[](https://github.com/K-Dense-AI/scientific-agent-skills/actions/workflows/skill-tests.yml)
|
||||
[](#-getting-started)
|
||||
@@ -25,23 +26,13 @@ validated: false
|
||||
[](https://www.linkedin.com/company/k-dense-inc)
|
||||
[](https://www.youtube.com/@K-Dense-Inc)
|
||||
|
||||
## Star History
|
||||
|
||||
<a href="https://www.star-history.com/?repos=K-Dense-AI%2Fscientific-agent-skills&type=date&legend=top-left">
|
||||
<picture>
|
||||
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/chart?repos=K-Dense-AI/scientific-agent-skills&type=date&theme=dark&legend=top-left&sealed_token=rL_5GLS9f4Fbyr1_VYZLGMF-8Rr6ZlWNaYNecajc52QSQq6KL7HrzSea_tGQGy1mBMXgVvAUMSIYAc0w39si9v5Up1RIw74-UDGZg_9HvH_chiyS0Njf-5tebtPh1LJjXTG6mH5Iv2pMJNivgfPsyB-oOgbaIV3uSc7DzSeZFCTE4WOcHX4y2BR76k5g" />
|
||||
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/chart?repos=K-Dense-AI/scientific-agent-skills&type=date&legend=top-left&sealed_token=rL_5GLS9f4Fbyr1_VYZLGMF-8Rr6ZlWNaYNecajc52QSQq6KL7HrzSea_tGQGy1mBMXgVvAUMSIYAc0w39si9v5Up1RIw74-UDGZg_9HvH_chiyS0Njf-5tebtPh1LJjXTG6mH5Iv2pMJNivgfPsyB-oOgbaIV3uSc7DzSeZFCTE4WOcHX4y2BR76k5g" />
|
||||
<img alt="Star History Chart" src="https://api.star-history.com/chart?repos=K-Dense-AI/scientific-agent-skills&type=date&legend=top-left&sealed_token=rL_5GLS9f4Fbyr1_VYZLGMF-8Rr6ZlWNaYNecajc52QSQq6KL7HrzSea_tGQGy1mBMXgVvAUMSIYAc0w39si9v5Up1RIw74-UDGZg_9HvH_chiyS0Njf-5tebtPh1LJjXTG6mH5Iv2pMJNivgfPsyB-oOgbaIV3uSc7DzSeZFCTE4WOcHX4y2BR76k5g" />
|
||||
</picture>
|
||||
</a>
|
||||
|
||||
> **🔔 Claude Scientific Skills is now Scientific Agent Skills.** Same skills, broader compatibility — now works with any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, not just Claude.
|
||||
|
||||
> **New: [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok)** — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 159 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via [Modal](https://modal.com/) for heavy workloads. [Get started here.](https://github.com/K-Dense-AI/k-dense-byok)
|
||||
> **New: [K-Dense BYOK](https://github.com/K-Dense-AI/k-dense-byok)** — A free, open-source AI co-scientist that runs on your desktop, powered by Scientific Agent Skills. Bring your own API keys, pick from 40+ models, and get a full research workspace with web search, file handling, 100+ scientific databases, and access to all 161 skills in this repo. Your data stays on your computer, and you can optionally scale to cloud compute via [Modal](https://modal.com/) for heavy workloads. [Get started here.](https://github.com/K-Dense-AI/k-dense-byok)
|
||||
|
||||
> **Stay up to date:** Follow K-Dense on [X](https://x.com/k_dense_ai), [LinkedIn](https://www.linkedin.com/company/k-dense-inc), and [YouTube](https://www.youtube.com/@K-Dense-Inc) for new skills, release announcements, walkthroughs, research workflow demos, and examples you can use with your own AI agent.
|
||||
|
||||
A comprehensive collection of **159 ready-to-use scientific and research skills** (covering cancer genomics, individual-level 1000 Genomes queries, hosted regulatory-sequence prediction, live pathogen-variant surveillance, analytical method validation, PK/PD modelling and dose selection, full-text biomedical and regulatory literature retrieval, drug-target binding, molecular dynamics, RNA velocity, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
|
||||
A comprehensive collection of **161 ready-to-use scientific and research skills** (covering cancer genomics, individual-level 1000 Genomes queries, hosted regulatory-sequence prediction, live pathogen-variant surveillance, analytical method validation, PK/PD modelling and dose selection, full-text biomedical and regulatory literature retrieval, drug-target binding, bounded biomedical knowledge graph search, molecular dynamics, RNA velocity, geospatial science, time series forecasting, scientific ML resource discovery via Hugging Science, 78+ scientific databases, and more) for any AI agent that supports the open [Agent Skills](https://agentskills.io/) standard, created by [K-Dense](https://k-dense.ai). The repository is also a portable [Agent Plugins](https://agent-plugins.org/) package (`plugin.json` + `skills/`), so plugin-capable clients can load the whole collection as one plugin. Works with **Cursor, Claude Code, Codex, Google Antigravity, and more**. Transform your AI agent into a research assistant capable of executing complex multi-step scientific workflows across biology, chemistry, medicine, and beyond.
|
||||
|
||||
> ⭐ **Help make AI for science easier to discover:** If Scientific Agent Skills saves you time, teaches your agent a workflow, or helps your lab move faster, please [star this repository](https://github.com/K-Dense-AI/scientific-agent-skills). A star is a public signal that these open, reusable research skills are worth maintaining: it helps scientists, engineers, and open-source contributors find the project, shows which agent-skill standards are gaining real adoption, and gives us a clear reason to keep expanding the collection for the community.
|
||||
|
||||
@@ -76,9 +67,9 @@ These skills enable your AI agent to seamlessly work with specialized scientific
|
||||
|
||||
## 📦 What's Included
|
||||
|
||||
This repository provides **159 scientific and research skills** organized into the following categories:
|
||||
This repository provides **161 scientific and research skills** organized into the following categories:
|
||||
|
||||
- **100+ Scientific & Financial Databases** - A unified database-lookup skill provides deterministic, provenance-rich access to 78 public databases (PubChem, ChEMBL, UniProt, COSMIC, ClinicalTrials.gov, FRED, USPTO, and more), plus dedicated skills for DepMap, Imaging Data Commons, PrimeKG, U.S. Treasury Fiscal Data, Hugging Science, OneKGPd, and Genomic Intelligence. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage
|
||||
- **100+ Scientific & Financial Databases** - A unified database-lookup skill provides deterministic, provenance-rich access to 78 public databases (PubChem, ChEMBL, UniProt, COSMIC, ClinicalTrials.gov, FRED, USPTO, and more), plus dedicated skills for DepMap, Imaging Data Commons, PrimeKG, NCATS ARAX, U.S. Treasury Fiscal Data, Hugging Science, OneKGPd, and Genomic Intelligence. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage
|
||||
- **70+ Optimized Python Package Skills** - Explicitly defined, version-aware workflows for RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, PathML, pydicom, NeuroKit2, PufferLib, QuTiP, GeoPandas, pymatgen, BioPython, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), and others. The agent can still use *any* Python package; these skills provide stronger, safer guidance for the packages listed
|
||||
- **9 Scientific Integration Skills** - Explicitly defined skills for Benchling, DNAnexus, LatchBio, OMERO, Protocols.io, Open Notebook, Ginkgo Cloud Lab, LabArchives, and Opentrons. Again, the agent is not limited to these — any API or platform reachable from Python is fair game; these skills are the optimized, pre-documented paths
|
||||
- **30+ Analysis & Communication Tools** - Literature review, evidence-traceable scientific writing, confidential peer review, document processing, Paperclip (full-text papers, FDA/PMDA/EMA filings, and trial registries with line-pinned citations), Paperzilla, Exa Search, macro-free PPTX posters, slides, schematics, infographics, Mermaid diagrams, and more
|
||||
@@ -123,7 +114,7 @@ Each skill includes:
|
||||
- **Multi-Step Workflows** - Execute complex pipelines with a single prompt
|
||||
|
||||
### 🎯 **Comprehensive Coverage**
|
||||
- **159 Skills** - Extensive coverage across all major scientific domains
|
||||
- **161 Skills** - Extensive coverage across all major scientific domains
|
||||
- **100+ Databases** - Unified access to 78+ databases via database-lookup, plus dedicated data access skills and multi-database packages like BioServices, BioPython, and gget
|
||||
- **70+ Optimized Python Package Skills** - Current, version-scoped guidance for packages including RDKit, Scanpy, PyTorch Lightning, scikit-learn, PyTDC, pydicom, PufferLib, QuTiP, GeoPandas, pymatgen, Qiskit, Molecular Dynamics (OpenMM/MDAnalysis), scVelo, and TimesFM (the agent can use any Python package; these are the pre-documented paths)
|
||||
|
||||
@@ -178,7 +169,7 @@ Pin to a specific release tag or commit SHA for reproducible installs:
|
||||
|
||||
```bash
|
||||
# Pin to a release tag
|
||||
gh skill install K-Dense-AI/scientific-agent-skills --pin v2.62.0
|
||||
gh skill install K-Dense-AI/scientific-agent-skills --pin v2.63.0
|
||||
|
||||
# Pin to a commit SHA
|
||||
gh skill install K-Dense-AI/scientific-agent-skills --pin abc123def
|
||||
@@ -194,6 +185,27 @@ gh skill update
|
||||
gh skill update --all
|
||||
```
|
||||
|
||||
### Option 3: Agent Plugins (Cursor, Codex, and other plugin clients)
|
||||
|
||||
This repository is a valid [Agent Plugins](https://agent-plugins.org/) 1.0.0 package: root [`plugin.json`](plugin.json) plus Agent Skills under `skills/`. Clients that support the standard discover every immediate child of `skills/` that contains a `SKILL.md`.
|
||||
|
||||
**Cursor** — symlink or copy the repo into the local plugins directory, then reload:
|
||||
|
||||
```bash
|
||||
mkdir -p ~/.cursor/plugins/local
|
||||
ln -s "$(pwd)" ~/.cursor/plugins/local/scientific-agent-skills
|
||||
```
|
||||
|
||||
Restart Cursor or run **Developer: Reload Window**, then confirm the plugin and its skills appear under **Customize**. See [Cursor plugins](https://cursor.com/docs/plugins).
|
||||
|
||||
**Codex** — install from a local checkout (confirm the current CLI flag names in Codex docs):
|
||||
|
||||
```bash
|
||||
codex plugins install .
|
||||
```
|
||||
|
||||
Compatible clients (Cursor, Codex, GitHub Copilot, VS Code, Kiro, and others listed at [agent-plugins.org](https://agent-plugins.org/compatible-clients)) share the same package layout; installation UX stays client-specific.
|
||||
|
||||
### Other Agent Skills hosts (OpenClaw, NemoClaw, Pi, Hermes, …)
|
||||
|
||||
Agent hosts differ in install paths, discovery settings, and support for optional frontmatter fields. `npx skills add` (Option 1) commonly installs into the `~/.agents/skills/` convention, with project-scoped installs under `.agents/skills/`; confirm both paths against your host's current documentation. To install manually on a host configured to scan one of those locations:
|
||||
@@ -209,7 +221,7 @@ For Hermes versions that support skill taps, add the repository as a tap:
|
||||
hermes skills tap add K-Dense-AI/scientific-agent-skills
|
||||
```
|
||||
|
||||
Every `SKILL.md` has YAML frontmatter, but legacy and community skills vary in `metadata` formatting (block or flow style) and optional extension fields. Repository updates must keep `metadata.version` as a quoted numeric string and pass canonical `skills-ref validate ./skills/<skill-name>` checks. Hosts may interpret optional metadata and credential prompts differently, so verify behavior on the target host. Because 159 skills add up to a lot of standing context, consider installing a topical subset rather than the whole collection.
|
||||
Every `SKILL.md` has YAML frontmatter, but legacy and community skills vary in `metadata` formatting (block or flow style) and optional extension fields. Repository updates must keep `metadata.version` as a quoted numeric string and pass canonical `skills-ref validate ./skills/<skill-name>` checks. Hosts may interpret optional metadata and credential prompts differently, so verify behavior on the target host. Because 161 skills add up to a lot of standing context, consider installing a topical subset rather than the whole collection.
|
||||
|
||||
> **NemoClaw note:** NemoClaw runs agents inside NVIDIA OpenShell with default-deny outbound networking. Skills are discovered and loaded normally, but any skill that needs the network — package installs via `uv`, or API calls (Exa, Parallel, Benchling, NCBI, Materials Project, …) — only works once the operator pre-approves the relevant domains in the OpenShell TUI.
|
||||
|
||||
@@ -448,7 +460,7 @@ networks, and search GEO for similar patterns.
|
||||
|
||||
## 📚 Available Skills
|
||||
|
||||
This repository contains **159 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
|
||||
This repository contains **161 scientific and research skills** organized across multiple domains. Each skill provides comprehensive documentation, code examples, and best practices for working with scientific libraries, databases, and tools.
|
||||
|
||||
### Skill Categories
|
||||
|
||||
@@ -564,12 +576,13 @@ This repository contains **159 scientific and research skills** organized across
|
||||
- Citations: Citation Management, pyzotero
|
||||
- Illustration: Generate Image (AI image generation with FLUX.2 Pro and Gemini 3.1 Flash Image / Nano Banana 2)
|
||||
|
||||
#### 🔬 **Scientific Databases & Data Access** (10 skills → 100+ databases total)
|
||||
#### 🔬 **Scientific Databases & Data Access** (11 skills → 100+ databases total)
|
||||
> A unified database-lookup skill provides deterministic REST API access to 78 public databases across all domains, with retrieval contracts, pagination/count reconciliation, and endpoint provenance. Dedicated skills cover specialized data platforms. Multi-database packages like BioServices (~40 bioinformatics services), BioPython (39 NCBI sub-databases via Entrez), and gget (20+ genomics databases) add further coverage.
|
||||
- Unified access: Database Lookup (78 databases spanning chemistry, genomics, clinical, pathways, patents, economics, and more — PubChem, ChEMBL, UniProt, PDB, AlphaFold, KEGG, Reactome, STRING, ClinVar, COSMIC, ClinicalTrials.gov, FDA, FRED, USPTO, SEC EDGAR, and dozens more — with auditable filters and provenance)
|
||||
- Cancer genomics: DepMap (cancer cell line dependencies, drug sensitivity, gene effect profiles)
|
||||
- Cancer imaging: Imaging Data Commons (NCI radiology & pathology datasets via idc-index)
|
||||
- Knowledge graph: PrimeKG (precision medicine knowledge graph — genes, drugs, diseases, phenotypes)
|
||||
- Biomedical knowledge graph search: [NCATS ARAX](skills/ncats-arax/) (bounded, Biolink-constrained one-hop and endpoint-pinned two-hop queries over knowledge graphs with up to five explicitly selected NCATS Translator providers, with provenance preservation)
|
||||
- Fiscal data: U.S. Treasury Fiscal Data (national debt, Treasury statements, auctions, exchange rates)
|
||||
- Scientific ML resource catalog: Hugging Science (curated index of datasets, models, blog posts, and interactive Spaces across 17 scientific domains — astronomy, biology, chemistry, climate, genomics, materials science, medicine, physics, scientific reasoning, and more — with usage patterns for `datasets`, `transformers`, and `gradio_client`)
|
||||
- Individual-level population genomics: OneKGPd (3,202-person high-coverage 1000 Genomes cohort queries)
|
||||
@@ -620,6 +633,7 @@ Deep dives, benchmarks, and guides from the [K-Dense blog](https://www.k-dense.a
|
||||
|
||||
- **[Agent Skills: The Final Piece for AI-Powered Scientific Research](https://www.k-dense.ai/blog/agent-skills-final-piece-for-ai-powered-research)** — What Agent Skills are, why curated domain guidance beats raw model capability, and an introduction to this repository.
|
||||
- **[K-Dense Web vs Scientific Agent Skills: Why We Built Both (And Which One You Should Use)](https://www.k-dense.ai/blog/k-dense-web-vs-scientific-agent-skills)** — When the open-source skills are the right tool, and when a hosted platform with managed compute makes more sense.
|
||||
- **[AI Co-Scientists, Answered: 20 Questions from a Live Session with a University Research Center](https://www.k-dense.ai/blog/ai-co-scientists-answered-20-questions)** — Practical questions from a research center evaluating AI co-scientists: what stays open source and MIT-licensed, how local and desktop deployments work, how data is handled, and how to choose between the hosted platform and the BYOK setup that runs these skills.
|
||||
|
||||
### Skill benchmarks and deep dives
|
||||
|
||||
@@ -631,6 +645,13 @@ Deep dives, benchmarks, and guides from the [K-Dense blog](https://www.k-dense.a
|
||||
- **[Benchmarking Nano Banana 2 Lite for Scientific Image Generation](https://www.k-dense.ai/blog/benchmarking-nano-banana-2-lite-scientific-image-model)** — A 240-image comparison of scientific-diagram models, useful when choosing a backend for [generate-image](skills/generate-image/): 3.8 s median latency for Nano Banana 2 Lite against 49 s for GPT Image 2, with a quality tradeoff.
|
||||
- **[Benchmarking NVIDIA BioNeMo Agent Toolkit Skills for NIM microservices](https://www.k-dense.ai/blog/benchmarking-nvidia-bionemo-nim-skill)** — A separate NVIDIA skill set rather than one of these, but the findings generalize: skills help most with routing to non-obvious endpoints and with weak-model reliability, and do not improve the underlying scientific model's accuracy.
|
||||
|
||||
### Why the workflow layer matters
|
||||
|
||||
- **[The Model Is No Longer the Bottleneck](https://www.k-dense.ai/blog/the-model-is-no-longer-the-bottleneck)** — The case for why a repository like this one exists: frontier models now match specialized scientific software on raw capability (±0.079 ppm on NMR hydrogen shift prediction), so the limiting factor has moved to the workflow around the model — data access, code execution, verification, and auditable output.
|
||||
- **[The AI Co-Scientist Is Here. The Bottleneck Is Verification.](https://www.k-dense.ai/blog/ai-co-scientist-verification-bottleneck)** — A 10-point checklist for evaluating a research agent, built around exposing sources, code, data provenance, and intermediate work rather than a polished final answer — the same reasoning behind the provenance and retrieval-contract requirements in skills like [database-lookup](skills/database-lookup/) and [scientific-writing](skills/scientific-writing/).
|
||||
- **[Reproduction, Not Generation, Is AI's Killer App for Science](https://www.k-dense.ai/blog/reproduction-not-generation-ai-for-science)** — Why re-running published analyses is the highest-value use of an agent: 78% of papers and 93% of individual analysis tasks reproduced across a 221-study benchmark, because a reproduction can be checked against known numbers while a generated claim cannot.
|
||||
- **[Introducing K-Bench 01: Nine Frontier Models, 178 Real Scientific Tasks, and a Lot of Confident Wrong Answers](https://www.k-dense.ai/blog/introducing-k-bench-01-internal-benchmark)** — Nine frontier models on 178 real user tasks, with overclaiming in 40% of runs. Useful calibration for what to check when an agent reports success, and context for the verification boundaries written into the clinical, regulatory, and research-methodology skills above.
|
||||
|
||||
### Security and safe deployment
|
||||
|
||||
- **[Security in the Science Agent Era: What Every Lab Needs to Know Before Installing Skills](https://www.k-dense.ai/blog/skill-security-before-you-install)** — The practical review checklist behind this repo's [Security Disclaimer](#%EF%B8%8F-security-disclaimer): read the full `SKILL.md` and `scripts/`, scan before installing, and pin versions instead of tracking a branch.
|
||||
@@ -842,7 +863,7 @@ Recommended practice:
|
||||
title = {Scientific Agent Skills: A Comprehensive Collection of Scientific Tools for AI Agents},
|
||||
year = {2026},
|
||||
url = {https://github.com/K-Dense-AI/scientific-agent-skills},
|
||||
note = {159 skills covering databases, packages, integrations, and analysis tools}
|
||||
note = {161 skills covering databases, packages, integrations, and analysis tools}
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
+328
@@ -0,0 +1,328 @@
|
||||
---
|
||||
title: "Smart Iterative Refinement Workflow"
|
||||
task: ""
|
||||
lineage_type: import
|
||||
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/13385c7c/skills/scientific-schematics/references/iterative_refinement.md
|
||||
upstream_sha: 13385c7c
|
||||
imported_at: 2026-08-13
|
||||
prompt_class: unknown
|
||||
upstream_changes: accepted
|
||||
author: upstream
|
||||
validated: false
|
||||
---
|
||||
|
||||
# Smart Iterative Refinement Workflow
|
||||
|
||||
How the generate-review-refine loop works: the initial generation, the quality review,
|
||||
the decision to continue or stop, subsequent iterations, and the review log. Then the
|
||||
advanced generation options (Python API, command-line options, prompt engineering) and
|
||||
four worked examples.
|
||||
|
||||
## Smart Iterative Refinement Workflow
|
||||
|
||||
The AI generation system uses **smart iteration** - it only regenerates if quality is below the threshold for your document type:
|
||||
|
||||
### How Smart Iteration Works
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────┐
|
||||
│ 1. Generate image with Nano Banana 2 │
|
||||
│ ↓ │
|
||||
│ 2. Review quality with Gemini 3.6 Flash │
|
||||
│ ↓ │
|
||||
│ 3. Score >= threshold? │
|
||||
│ YES → DONE! (early stop) │
|
||||
│ NO → Improve prompt, go to step 1 │
|
||||
│ ↓ │
|
||||
│ 4. Repeat until quality met OR max iterations │
|
||||
└─────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Iteration 1: Initial Generation
|
||||
**Prompt Construction:**
|
||||
```
|
||||
Scientific diagram guidelines + User request
|
||||
```
|
||||
|
||||
**Output:** `diagram_v1.png`
|
||||
|
||||
### Quality Review by Gemini 3.6 Flash
|
||||
|
||||
Gemini 3.6 Flash evaluates the diagram on:
|
||||
1. **Scientific Accuracy** (0-2 points) - Correct concepts, notation, relationships
|
||||
2. **Clarity and Readability** (0-2 points) - Easy to understand, clear hierarchy
|
||||
3. **Label Quality** (0-2 points) - Complete, readable, consistent labels
|
||||
4. **Layout and Composition** (0-2 points) - Logical flow, balanced, no overlaps
|
||||
5. **Professional Appearance** (0-2 points) - Publication-ready quality
|
||||
|
||||
**Example Review Output:**
|
||||
```
|
||||
SCORE: 8.0
|
||||
|
||||
STRENGTHS:
|
||||
- Clear flow from top to bottom
|
||||
- All phases properly labeled
|
||||
- Professional typography
|
||||
|
||||
ISSUES:
|
||||
- Participant counts slightly small
|
||||
- Minor overlap on exclusion box
|
||||
|
||||
VERDICT: ACCEPTABLE (for poster, threshold 7.0)
|
||||
```
|
||||
|
||||
### Decision Point: Continue or Stop?
|
||||
|
||||
| If Score... | Action |
|
||||
|-------------|--------|
|
||||
| >= threshold | **STOP** - Quality is good enough for this document type |
|
||||
| < threshold | Continue to next iteration with improved prompt |
|
||||
|
||||
**Example:**
|
||||
- For a **poster** (threshold 7.0): Score of 7.5 → **DONE after 1 iteration!**
|
||||
- For a **journal** (threshold 8.5): Score of 7.5 → Continue improving
|
||||
|
||||
### Subsequent Iterations (Only If Needed)
|
||||
|
||||
If quality is below threshold, the system:
|
||||
1. Extracts specific issues from Gemini 3.6 Flash's review
|
||||
2. Enhances the prompt with improvement instructions
|
||||
3. Regenerates with Nano Banana 2
|
||||
4. Reviews again with Gemini 3.6 Flash
|
||||
5. Repeats until threshold met or max iterations reached
|
||||
|
||||
### Review Log
|
||||
All iterations are saved with a JSON review log that includes early-stop information:
|
||||
```json
|
||||
{
|
||||
"user_prompt": "CONSORT participant flow diagram...",
|
||||
"doc_type": "poster",
|
||||
"quality_threshold": 7.0,
|
||||
"iterations": [
|
||||
{
|
||||
"iteration": 1,
|
||||
"image_path": "figures/consort_v1.png",
|
||||
"score": 7.5,
|
||||
"reviewed": true,
|
||||
"review_error": null,
|
||||
"needs_improvement": false,
|
||||
"critique": "SCORE: 7.5\nSTRENGTHS:..."
|
||||
}
|
||||
],
|
||||
"final_score": 7.5,
|
||||
"final_reviewed": true,
|
||||
"early_stop": true,
|
||||
"early_stop_reason": "Quality score 7.5 meets threshold 7.0 for poster"
|
||||
}
|
||||
```
|
||||
|
||||
**Note:** With smart iteration, you may see only 1 iteration instead of the full 2 if quality is achieved early!
|
||||
|
||||
### When the review does not run
|
||||
|
||||
The reviewer is a second model call, and it can fail on its own — a rate limit, a content filter,
|
||||
an answer in a shape the parser cannot read. The image is generated first and is kept regardless;
|
||||
what is missing in that case is the *measurement*, so the log says so rather than substituting a
|
||||
number:
|
||||
|
||||
```json
|
||||
{
|
||||
"iterations": [
|
||||
{
|
||||
"iteration": 1,
|
||||
"image_path": "figures/consort_v1.png",
|
||||
"score": null,
|
||||
"reviewed": false,
|
||||
"review_error": "the review model returned no choices",
|
||||
"needs_improvement": false,
|
||||
"critique": "Review unavailable: the review model returned no choices."
|
||||
}
|
||||
],
|
||||
"final_score": null,
|
||||
"final_reviewed": false
|
||||
}
|
||||
```
|
||||
|
||||
The run exits 0 — the diagram is real — and prints
|
||||
`Review unavailable — image kept, quality not verified`. It does **not** regenerate: a reviewer that
|
||||
did not answer says nothing about the diagram, so another generation would be guesswork. Look at the
|
||||
image yourself, and re-run if you want a score; these failures are usually transient.
|
||||
|
||||
`final_reviewed` is the field to check in automation. `final_score` alone cannot distinguish
|
||||
"scored 7.5" from "never scored".
|
||||
|
||||
## Advanced AI Generation Usage
|
||||
|
||||
### Python API
|
||||
|
||||
```python
|
||||
from scripts.generate_schematic_ai import ScientificSchematicGenerator
|
||||
|
||||
# Initialize generator
|
||||
generator = ScientificSchematicGenerator(
|
||||
api_key="your_openrouter_key",
|
||||
verbose=True
|
||||
)
|
||||
|
||||
# Generate with iterative refinement (max 2 iterations)
|
||||
results = generator.generate_iterative(
|
||||
user_prompt="Transformer architecture diagram",
|
||||
output_path="figures/transformer.png",
|
||||
iterations=2
|
||||
)
|
||||
|
||||
# Access results
|
||||
print(f"Final score: {results['final_score']}/10")
|
||||
print(f"Final image: {results['final_image']}")
|
||||
|
||||
# Review individual iterations
|
||||
for iteration in results['iterations']:
|
||||
print(f"Iteration {iteration['iteration']}: {iteration['score']}/10")
|
||||
print(f"Critique: {iteration['critique']}")
|
||||
```
|
||||
|
||||
### Command-Line Options
|
||||
|
||||
```bash
|
||||
# Basic usage (default threshold 7.5/10)
|
||||
python scripts/generate_schematic.py "diagram description" -o output.png
|
||||
|
||||
# Specify document type for appropriate quality threshold
|
||||
python scripts/generate_schematic.py "diagram" -o out.png --doc-type journal # 8.5/10
|
||||
python scripts/generate_schematic.py "diagram" -o out.png --doc-type conference # 8.0/10
|
||||
python scripts/generate_schematic.py "diagram" -o out.png --doc-type poster # 7.0/10
|
||||
python scripts/generate_schematic.py "diagram" -o out.png --doc-type presentation # 6.5/10
|
||||
|
||||
# Custom max iterations (1-2)
|
||||
python scripts/generate_schematic.py "complex diagram" -o diagram.png --iterations 2
|
||||
|
||||
# Verbose output (see all API calls and reviews)
|
||||
python scripts/generate_schematic.py "flowchart" -o flow.png -v
|
||||
|
||||
# Provide API key via flag
|
||||
python scripts/generate_schematic.py "diagram" -o out.png --api-key "sk-or-v1-..."
|
||||
|
||||
# Combine options
|
||||
python scripts/generate_schematic.py "neural network" -o nn.png --doc-type journal --iterations 2 -v
|
||||
```
|
||||
|
||||
### Setup and Cost
|
||||
|
||||
```bash
|
||||
# Get a key at https://openrouter.ai/keys
|
||||
export OPENROUTER_API_KEY='sk-or-v1-your_key_here'
|
||||
|
||||
# Or persist it in the shell profile
|
||||
echo 'export OPENROUTER_API_KEY="sk-or-v1-your_key"' >> ~/.zshrc
|
||||
|
||||
# Or drop it in a .env file at the project root
|
||||
echo "OPENROUTER_API_KEY=sk-or-v1-..." >> .env
|
||||
|
||||
# The only Python dependency
|
||||
uv pip install requests
|
||||
```
|
||||
|
||||
Each iteration costs **two API calls**: one image generation and one vision review. A diagram that
|
||||
passes on the first try therefore costs two calls, and the maximum for any single run is four. The
|
||||
image model dominates the bill. Check current per-token pricing for
|
||||
`google/gemini-3.1-flash-image` and `google/gemini-3.7-flash` on OpenRouter — it changes, and any
|
||||
figure written here would go stale.
|
||||
|
||||
### Prompt Engineering Tips
|
||||
|
||||
**1. Be Specific About Layout:**
|
||||
```
|
||||
✓ "Flowchart with vertical flow, top to bottom"
|
||||
✓ "Architecture diagram with encoder on left, decoder on right"
|
||||
✓ "Circular pathway diagram with clockwise flow"
|
||||
```
|
||||
|
||||
**2. Include Quantitative Details:**
|
||||
```
|
||||
✓ "Neural network with input layer (784 nodes), hidden layer (128 nodes), output (10 nodes)"
|
||||
✓ "Flowchart showing n=500 screened, n=150 excluded, n=350 randomized"
|
||||
✓ "Circuit with 1kΩ resistor, 10µF capacitor, 5V source"
|
||||
```
|
||||
|
||||
**3. Specify Visual Style:**
|
||||
```
|
||||
✓ "Minimalist block diagram with clean lines"
|
||||
✓ "Detailed biological pathway with protein structures"
|
||||
✓ "Technical schematic with engineering notation"
|
||||
```
|
||||
|
||||
**4. Request Specific Labels:**
|
||||
```
|
||||
✓ "Label all arrows with activation/inhibition"
|
||||
✓ "Include layer dimensions in each box"
|
||||
✓ "Show time progression with timestamps"
|
||||
```
|
||||
|
||||
**5. Mention Color Requirements:**
|
||||
```
|
||||
✓ "Use colorblind-friendly colors"
|
||||
✓ "Grayscale-compatible design"
|
||||
✓ "Color-code by function: blue for input, green for processing, red for output"
|
||||
```
|
||||
|
||||
## AI Generation Examples
|
||||
|
||||
### Example 1: CONSORT Flowchart
|
||||
```bash
|
||||
python scripts/generate_schematic.py \
|
||||
"CONSORT participant flow diagram for randomized controlled trial. \
|
||||
Start with 'Assessed for eligibility (n=500)' at top. \
|
||||
Show 'Excluded (n=150)' with reasons: age<18 (n=80), declined (n=50), other (n=20). \
|
||||
Then 'Randomized (n=350)' splits into two arms: \
|
||||
'Treatment group (n=175)' and 'Control group (n=175)'. \
|
||||
Each arm shows 'Lost to follow-up' (n=15 and n=10). \
|
||||
End with 'Analyzed' (n=160 and n=165). \
|
||||
Use blue boxes for process steps, orange for exclusion, green for final analysis." \
|
||||
-o figures/consort.png
|
||||
```
|
||||
|
||||
### Example 2: Neural Network Architecture
|
||||
```bash
|
||||
python scripts/generate_schematic.py \
|
||||
"Transformer encoder-decoder architecture diagram. \
|
||||
Left side: Encoder stack with input embedding, positional encoding, \
|
||||
multi-head self-attention, add & norm, feed-forward, add & norm. \
|
||||
Right side: Decoder stack with output embedding, positional encoding, \
|
||||
masked self-attention, add & norm, cross-attention (receiving from encoder), \
|
||||
add & norm, feed-forward, add & norm, linear & softmax. \
|
||||
Show cross-attention connection from encoder to decoder with dashed line. \
|
||||
Use light blue for encoder, light red for decoder. \
|
||||
Label all components clearly." \
|
||||
-o figures/transformer.png --iterations 2
|
||||
```
|
||||
|
||||
### Example 3: Biological Pathway
|
||||
```bash
|
||||
python scripts/generate_schematic.py \
|
||||
"MAPK signaling pathway diagram. \
|
||||
Start with EGFR receptor at cell membrane (top). \
|
||||
Arrow down to RAS (with GTP label). \
|
||||
Arrow to RAF kinase. \
|
||||
Arrow to MEK kinase. \
|
||||
Arrow to ERK kinase. \
|
||||
Final arrow to nucleus showing gene transcription. \
|
||||
Label each arrow with 'phosphorylation' or 'activation'. \
|
||||
Use rounded rectangles for proteins, different colors for each. \
|
||||
Include membrane boundary line at top." \
|
||||
-o figures/mapk_pathway.png
|
||||
```
|
||||
|
||||
### Example 4: System Architecture
|
||||
```bash
|
||||
python scripts/generate_schematic.py \
|
||||
"IoT system architecture block diagram. \
|
||||
Bottom layer: Sensors (temperature, humidity, motion) in green boxes. \
|
||||
Middle layer: Microcontroller (ESP32) in blue box. \
|
||||
Connections to WiFi module (orange box) and Display (purple box). \
|
||||
Top layer: Cloud server (gray box) connected to mobile app (light blue box). \
|
||||
Show data flow arrows between all components. \
|
||||
Label connections with protocols: I2C, UART, WiFi, HTTPS." \
|
||||
-o figures/iot_architecture.png
|
||||
```
|
||||
|
||||
---
|
||||
Reference in New Issue
Block a user