12 KiB
12 KiB
title, task, lineage_type, upstream_source, upstream_sha, imported_at, prompt_class, upstream_changes, author, validated
| title | task | lineage_type | upstream_source | upstream_sha | imported_at | prompt_class | upstream_changes | author | validated |
|---|---|---|---|---|---|---|---|---|---|
| ToolUniverse SDK Reference | import | https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse-sdk/REFERENCE.md | e2520a96 | 2026-06-26 | prompt | accepted | upstream | false |
ToolUniverse SDK Reference
Detailed reference for ToolUniverse Python SDK features and advanced usage.
Installation Details
Package Options
# Base installation (~200-300 MB)
pip install tooluniverse
# With embedding search (+1-2 GB for model weights)
pip install tooluniverse[embedding]
# With ML tools (+2-5 GB for model weights)
pip install tooluniverse[ml]
# All features
pip install tooluniverse[all]
# Minimal client for HTTP API access
pip install tooluniverse[client]
Performance Metrics
- First
load_tools(): 5-10 seconds (1000++ tools) - Subsequent loads: 2-3 seconds (cached)
- Specific categories: <2 seconds
- Memory: 200-300 MB base, +1-2 GB with embeddings
Initialization Options
tu = ToolUniverse(
use_cache=True, # Global caching
hooks_enabled=True, # Auto-summarization
log_level="INFO", # Logging level
tool_files={}, # Custom tool configs
keep_default_tools=True # Include defaults
)
# Load options
tu.load_tools(
categories=["proteins", "drugs"], # Specific categories
tool_config_files={...} # Custom configs
)
Tool Discovery Methods
Method 1: Keyword Search
Best for: Fast searches, exact matches, no API key
tools = tu.run({
"name": "Tool_Finder_Keyword",
"arguments": {
"description": "protein structure",
"limit": 10
}
})
# Always check structure
if isinstance(tools, dict) and 'tools' in tools:
for tool in tools['tools']:
print(f"{tool['name']}: {tool['description']}")
Method 2: LLM Search
Best for: Complex queries, natural language, best matches
Requires: OPENAI_API_KEY
tools = tu.run({
"name": "Tool_Finder_LLM",
"arguments": {
"description": "find genetic variants associated with Alzheimer's",
"limit": 5
}
})
Method 3: Embedding Search
Best for: Semantic similarity, concept matching
Requires: GPU for embedding model
tools = tu.run({
"name": "Tool_Finder",
"arguments": {
"description": "protein-protein interaction networks",
"limit": 10,
"return_call_result": False # Only return info
}
})
Advanced Caching
Cache Configuration
import os
from pathlib import Path
cache_dir = Path.home() / ".tooluniverse" / "cache"
cache_dir.mkdir(parents=True, exist_ok=True)
os.environ["TOOLUNIVERSE_CACHE_PATH"] = str(cache_dir / "cache.sqlite")
os.environ["TOOLUNIVERSE_CACHE_ENABLED"] = "true"
os.environ["TOOLUNIVERSE_CACHE_PERSIST"] = "true"
tu = ToolUniverse(use_cache=True)
tu.load_tools()
Cache Management
# Get statistics
stats = tu.get_cache_stats()
print(f"Hits: {stats['hits']}, Misses: {stats['misses']}")
# Inspect cache
for entry in tu.dump_cache():
print(f"Tool: {entry['tool_name']}")
print(f"Hit count: {entry['hit_count']}")
# Clear cache
tu.clear_cache()
# Always close connections
tu.close()
When to Cache
✅ Good candidates:
- ML model predictions (deterministic)
- Database queries (stable data)
- Protein structure predictions
- Literature searches (stable results)
❌ Avoid:
- Real-time data
- Time-sensitive queries
- User-specific data
- Rapidly changing data
Hooks System
Basic Hooks
# Enable default summarization hook
tu = ToolUniverse(hooks_enabled=True)
tu.load_tools()
result = tu.tools.OpenTargets_get_target_gene_ontology_by_ensemblID(
ensemblId="ENSG00000012048"
)
# Check if hook applied
if isinstance(result, dict) and "summary" in result:
print(f"Original: {result['original_length']} chars")
print(f"Summary: {len(result['summary'])} chars")
Custom Hook Configuration
hook_config = {
"exclude_tools": [
"Tool_RAG",
"ToolFinderEmbedding",
"CustomTool_*" # Wildcard pattern
],
"hooks": [{
"name": "summarization_hook",
"type": "SummarizationHook",
"enabled": True,
"conditions": {
"output_length": {"operator": ">", "threshold": 5000}
},
"hook_config": {
"max_tokens": 2000,
"summary_style": "concise"
}
}]
}
tu = ToolUniverse(hooks_enabled=True, hook_config=hook_config)
tu.load_tools()
File Save Hook
import tempfile
hook_config = {
"hooks": [{
"name": "file_save_hook",
"type": "FileSaveHook",
"enabled": True,
"conditions": {
"output_length": {"operator": ">", "threshold": 10000}
},
"hook_config": {
"temp_dir": tempfile.gettempdir(),
"file_prefix": "tool_output",
"include_metadata": True
}
}]
}
tu = ToolUniverse(hooks_enabled=True, hook_config=hook_config)
tu.load_tools()
result = tu.tools.some_large_output_tool(param="value")
if isinstance(result, dict) and "file_path" in result:
print(f"Saved to: {result['file_path']}")
Remote HTTP API
Deploy ToolUniverse as a server for remote access.
Server Setup
# Install on server
pip install tooluniverse
# Start server
tooluniverse-http-api --host 0.0.0.0 --port 8080
Client Usage
# Install minimal client
pip install tooluniverse[client]
from tooluniverse import ToolUniverseClient
client = ToolUniverseClient("http://server:8080")
# Use like local ToolUniverse
client.load_tools(tool_type=['uniprot', 'ChEMBL'])
result = client.run_one_function({
"name": "UniProt_get_entry_by_accession",
"arguments": {"accession": "P05067"}
})
Benefits:
- Minimal client dependencies (
requests+pydantic) - All computation on server
- Automatic method updates
- Shared instance across users
Complex Workflows
Multi-Step Pipeline
def analyze_disease_targets(disease_efo_id):
"""Complete disease-to-drug pipeline"""
tu = ToolUniverse(use_cache=True, hooks_enabled=True)
tu.load_tools()
try:
# Step 1: Get targets
targets = tu.tools.OpenTargets_get_associated_targets_by_disease_efoId(
efoId=disease_efo_id
)
if not targets or 'data' not in targets:
return {"error": "No targets found"}
# Step 2: Get target info (batch)
target_ids = [t['target']['id'] for t in targets['data'][:5]]
target_calls = [
{"name": "UniProt_get_entry_by_accession",
"arguments": {"accession": tid}}
for tid in target_ids
]
target_info = tu.run_batch(target_calls)
# Step 3: Find compounds (batch)
compound_calls = [
{"name": "ChEMBL_search_molecule_by_target",
"arguments": {"target_id": tid, "limit": 10}}
for tid in target_ids
]
compounds = tu.run_batch(compound_calls)
# Step 4: ADMET predictions
all_smiles = []
for comp_list in compounds:
if comp_list and 'molecules' in comp_list:
all_smiles.extend([m['smiles'] for m in comp_list['molecules'][:3]])
admet_calls = [
{"name": "ADMETAI_predict_admet", "arguments": {"smiles": s}}
for s in all_smiles
]
admet_results = tu.run_batch(admet_calls)
# Step 5: Literature search
gene_names = [t.get('gene_name', '') for t in target_info if t]
lit_calls = [
{"name": "PubMed_search_articles",
"arguments": {"query": f"{gene} drug therapy", "max_results": 10}}
for gene in gene_names[:3] if gene
]
literature = tu.run_batch(lit_calls)
return {
"targets": targets,
"target_info": target_info,
"compounds": compounds,
"admet": admet_results,
"literature": literature
}
except Exception as e:
return {"error": str(e)}
finally:
tu.close()
Error Handling Wrapper
from tooluniverse.exceptions import ToolError, ToolUnavailableError
def safe_tool_executor(tu, tool_name, arguments, fallback_tool=None):
"""Execute tool with comprehensive error handling"""
try:
return tu.run({
"name": tool_name,
"arguments": arguments
})
except ToolUnavailableError:
if fallback_tool:
print(f"⚠️ {tool_name} unavailable, trying {fallback_tool}")
return tu.run({
"name": fallback_tool,
"arguments": arguments
})
return {"error": f"Tool {tool_name} unavailable"}
except ToolError as e:
print(f"❌ Tool error: {e}")
return {"error": str(e)}
except Exception as e:
print(f"❌ Unexpected error: {e}")
return {"error": str(e)}
# Usage
result = safe_tool_executor(
tu,
"UniProt_get_entry_by_accession",
{"accession": "P05067"},
fallback_tool="AlternativeTool"
)
Environment Variables
API Keys
export OPENAI_API_KEY="sk-..." # Required for LLM tools
export NCBI_API_KEY="..." # Optional, higher rate limits
export USPTO_API_KEY="..." # For patent tools
Cache Configuration
export TOOLUNIVERSE_CACHE_ENABLED="true"
export TOOLUNIVERSE_CACHE_PERSIST="true"
export TOOLUNIVERSE_CACHE_PATH="/path/to/cache.sqlite"
Performance Tuning
export TOOLUNIVERSE_LIGHT_IMPORT="1" # Minimal imports
export TOOLUNIVERSE_LOAD_TIMEOUT="30" # Load timeout (seconds)
export TOOLUNIVERSE_TIMEOUT="120" # Execution timeout (seconds)
Common Issues & Solutions
Issue: Tool Not Loading
# Check tool registry
from tooluniverse.tool_registry import get_tool_registry
registry = get_tool_registry()
print(f"Registered: {len(registry)}")
# Enable debug logging
import logging
logging.basicConfig(level=logging.DEBUG)
tu = ToolUniverse()
tu.load_tools()
Issue: Slow Performance
# Solution 1: Load specific categories
tu = ToolUniverse()
tu.load_tools(categories=["proteins", "drugs"])
# Solution 2: Enable caching
tu = ToolUniverse(use_cache=True)
# Solution 3: Use batch execution
results = tu.run_batch(calls)
# Solution 4: Disable validation (after testing)
result = tu.tools.tool_name(param="value", validate=False)
Issue: High Memory Usage
# Use light import mode
import os
os.environ["TOOLUNIVERSE_LIGHT_IMPORT"] = "1"
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools(categories=["proteins"])
# Clear cache periodically
tu.clear_cache()
# Always close
tu.close()
Issue: API Key Errors
import os
# Check if set
api_key = os.environ.get("OPENAI_API_KEY")
if not api_key:
print("⚠️ Set: export OPENAI_API_KEY='sk-...'")
else:
print(f"✅ Key set: {api_key[:10]}...")
# Test key
try:
tools = tu.run({
"name": "Tool_Finder_LLM",
"arguments": {"description": "test", "limit": 1}
})
print("✅ API key valid")
except Exception as e:
print(f"❌ API key error: {e}")
Best Practices Summary
- Always call
load_tools()after initialization - Check result structures (tool finders return nested dicts)
- Use caching for expensive, deterministic operations
- Use batch execution for parallel tasks
- Handle errors with try/except blocks
- Close connections with
tu.close()when done - Validate parameters before execution (check tool schema)
- Use appropriate tool finder (keyword for speed, LLM for accuracy)
- Enable hooks for large outputs
- Document workflows clearly
Additional Resources
- Main Documentation: https://zitniklab.hms.harvard.edu/ToolUniverse/
- API Reference: https://zitniklab.hms.harvard.edu/ToolUniverse/api/modules.html
- Tutorials: https://zitniklab.hms.harvard.edu/ToolUniverse/guide/scientific_workflows.html
- GitHub Examples: https://github.com/mims-harvard/ToolUniverse/tree/main/examples
- Community Slack: https://join.slack.com/t/tooluniversehq/shared_invite/zt-3dic3eoio-5xxoJch7TLNibNQn5_AREQ