[Upstream sync] K-Dense-AI/scientific-agent-skills (github) — 3 added, 12 modified #38

Open
promptadmin wants to merge 15 commits from upstream-sync/scientific-agent-skills-20260811-d661d2-etzt into main
Showing only changes of commit 160ac9e438 - Show all commits
@@ -2,9 +2,9 @@
title: "Direct Parquet Access Guide for IDC"
task: ""
lineage_type: import
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/9c9bd2e9/skills/imaging-data-commons/references/parquet_access_guide.md
upstream_sha: 9c9bd2e9
imported_at: 2026-06-27
upstream_source: https://github.com/K-Dense-AI/scientific-agent-skills/blob/d661d27e/skills/imaging-data-commons/references/parquet_access_guide.md
upstream_sha: d661d27e
imported_at: 2026-08-11
prompt_class: prompt
upstream_changes: accepted
author: upstream
@@ -13,20 +13,23 @@ validated: false
# Direct Parquet Access Guide for IDC
**Tested with:** idc-index-data 23.10.1, DuckDB 1.x
**Tested with:** idc-index-data 24.2.2 (IDC data version v24), DuckDB 1.5
All idc-index metadata tables are published as Parquet files to a public GCS bucket with unrestricted CORS access. This enables metadata queries with DuckDB or pandas without installing idc-index — useful for quick exploration or environments where pip install is unavailable.
All idc-index metadata tables are published as Parquet files to a public GCS bucket with unrestricted CORS access. This enables metadata queries with DuckDB or pandas without installing idc-index.
**Limitation:** download helpers (`download_from_selection()`), viewer URLs (`get_viewer_URL()`), and citation generation require the idc-index client and are not available from raw Parquet files.
**This is not the first no-install option to reach for.** It still needs DuckDB installed, and the per-collection clinical tables are not published here — only the `clinical_index` dictionary. For ad-hoc metadata with nothing installed, the REST API (`rest_api_guide.md`) needs no install at all and reaches `clinical.<table>` through `POST /sql`.
## When to Use This Guide
Load this guide when you need to:
- Query IDC metadata without installing idc-index
- Run ad-hoc DuckDB queries against the latest index files
- Access `volume_geometry_index` or `rtstruct_index` for geometry validation or RT structure queries
- Pin queries to a specific IDC data version (see *Pinning to a Specific Version* below) rather than whatever the hosted API currently serves
- Return more rows than the REST `/sql` ceiling of 10 000
- Run heavy or repeated local DuckDB analysis without driving the hosted API
- Query IDC metadata where DuckDB is available but idc-index is not
For full API access (downloads, viewer, citations), use idc-index as documented in the main SKILL.md.
For downloads, viewer URLs, and citations, use idc-index as documented in the main SKILL.md.
## URL Pattern
@@ -51,16 +54,17 @@ https://storage.googleapis.com/idc-index-data-artifacts/current/release_artifact
| `collections_index.parquet` | — | Collection-level metadata |
| `analysis_results_index.parquet` | — | Derived dataset metadata |
| `clinical_index.parquet` | ~0.2 MB | Clinical data column dictionary |
| `ct_index.parquet` | — | CT acquisition/reconstruction parameters |
| `mr_index.parquet` | — | MR sequence/acquisition parameters |
| `pt_index.parquet` | — | PET acquisition/radiopharmaceutical parameters |
| `prior_versions_index.parquet` | — | Series from previous IDC releases |
**Note:** the main index file is named `idc_index.parquet`, not `index.parquet`. Reference it with an alias in SQL queries (e.g., `FROM read_parquet(...) AS index`).
## Prerequisites
```bash
pip install duckdb
# or: uv add duckdb
```
Install the Python `duckdb` package, using whatever installer manages the environment you are
running in.
DuckDB reads Parquet directly from HTTPS URLs using HTTP range requests — no GCS client library or authentication required.