Files
drug-discovery-prompts/upstream/mims-harvard-ToolUniverse/skills/tooluniverse-gene-enrichment/references/troubleshooting.md

637 lines
14 KiB
Markdown

---
title: "Troubleshooting Guide"
task: ""
lineage_type: import
upstream_source: https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse-gene-enrichment/references/troubleshooting.md
upstream_sha: e2520a96
imported_at: 2026-06-26
prompt_class: prompt
upstream_changes: accepted
author: upstream
validated: false
---
# Troubleshooting Guide
Complete troubleshooting guide for gene enrichment analysis issues.
---
## Common Issues Index
1. [No Significant Enrichment Found](#no-significant-enrichment-found)
2. [Gene Not Found Errors](#gene-not-found-errors)
3. [gseapy Library Not Found](#gseapy-library-not-found)
4. [PANTHER Returns Empty Results](#panther-returns-empty-results)
5. [STRING Returns All Categories](#string-returns-all-categories)
6. [Reactome Identifiers Not Found](#reactome-identifiers-not-found)
7. [enrichr_gene_enrichment_analysis Returns Wrong Format](#enrichr_gene_enrichment_analysis-returns-wrong-format)
8. [Results Don't Match Between Tools](#results-dont-match-between-tools)
9. [GSEA Gives Unstable Results](#gsea-gives-unstable-results)
10. [Memory Errors with Large Gene Lists](#memory-errors-with-large-gene-lists)
---
## No Significant Enrichment Found
### Symptoms
```python
go_bp_result = gseapy.enrichr(gene_list=genes, gene_sets='GO_Biological_Process_2021', ...)
go_bp_sig = go_bp_result.results[go_bp_result.results['Adjusted P-value'] < 0.05]
print(len(go_bp_sig)) # Output: 0
```
### Possible Causes & Solutions
#### 1. Gene Symbols Are Invalid
**Diagnosis**:
```python
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
# Check if genes are recognized
mapped = tu.tools.STRING_map_identifiers(
protein_ids=genes,
species=9606
)
print(f"Mapped: {len(mapped['data'])}/{len(genes)}")
```
**Solution**:
- Convert IDs using MyGene: `tu.tools.MyGene_batch_query(gene_ids=genes, fields='symbol')`
- Use official HGNC symbols
- Remove version suffixes from Ensembl IDs
#### 2. Library Version Too Old/New
**Solution**:
```python
# Try different library versions
for lib_version in ['GO_Biological_Process_2021', 'GO_Biological_Process_2023', 'GO_Biological_Process_2025']:
result = gseapy.enrichr(gene_list=genes, gene_sets=lib_version, ...)
sig = result.results[result.results['Adjusted P-value'] < 0.05]
print(f"{lib_version}: {len(sig)} significant terms")
```
#### 3. Significance Threshold Too Stringent
**Solution**:
```python
# Try relaxing threshold
for cutoff in [0.05, 0.1, 0.15, 0.2]:
sig = go_bp_result.results[go_bp_result.results['Adjusted P-value'] < cutoff]
print(f"Cutoff {cutoff}: {len(sig)} significant terms")
```
#### 4. Wrong Statistical Method
**Solution**:
```python
# If ORA fails, try GSEA
import pandas as pd
# Create ranked list (even if you don't have fold-changes)
# Use presence/absence scoring: 1 for genes in list, 0 for others
ranked_dict = {gene: 1 for gene in genes}
ranked_series = pd.Series(ranked_dict).sort_values(ascending=False)
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
...
)
```
#### 5. Organism Mismatch
**Solution**:
```python
# Verify organism parameter
result = gseapy.enrichr(
gene_list=genes,
gene_sets='GO_Biological_Process_2021',
organism='human', # NOT 'homo_sapiens' or 9606
...
)
```
#### 6. Gene List Too Small or Too Large
**Diagnosis**:
```python
print(f"Gene list size: {len(genes)}")
# Small (<5): insufficient power
# Large (>500): too non-specific
```
**Solution**:
- Small lists (<5 genes): Use PANTHER/STRING (better for small lists) or switch to gene-level annotation
- Large lists (>500 genes): Use custom background or stricter DEG cutoff
#### 7. No True Enrichment
**Valid Result**: Sometimes there genuinely is no enrichment. Document this:
```markdown
## Results
No significant enrichment was found at FDR < 0.05.
- Gene list size: 45 genes
- Libraries tested: GO BP, GO MF, GO CC, KEGG, Reactome
- Alternative cutoff (FDR < 0.1): 3 weak hits (exploratory only)
```
---
## Gene Not Found Errors
### Symptoms
```
Error: Gene 'TP-53' not found
Warning: 15 out of 50 genes not mapped
```
### Solutions
#### 1. Check ID Type
```python
# Detect ID type
sample = genes[0]
if sample.startswith('ENSG'):
id_type = 'ensembl'
elif sample.isdigit():
id_type = 'entrez'
elif len(sample) == 6 and sample[0].isalpha():
id_type = 'uniprot'
else:
id_type = 'symbol'
print(f"Detected ID type: {id_type}")
```
#### 2. Convert IDs
```python
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
# Convert Ensembl -> Symbol
if id_type == 'ensembl':
result = tu.tools.MyGene_batch_query(
gene_ids=genes,
fields='symbol,entrezgene,ensembl.gene'
)
symbol_map = {hit['query']: hit.get('symbol', hit['query'])
for hit in result.get('results', [])}
gene_symbols = [symbol_map.get(g, g) for g in genes]
```
#### 3. Remove Version Suffixes
```python
# ENSG00000141510.16 -> ENSG00000141510
genes_clean = [g.split('.')[0] for g in genes]
```
#### 4. Fix Common Typos
```python
# Common issues
fixes = {
'TP-53': 'TP53',
'BRCA 1': 'BRCA1',
'p53': 'TP53',
'egfr': 'EGFR', # case-sensitive
}
genes_fixed = [fixes.get(g, g) for g in genes]
```
#### 5. Try Aliases
```python
# Use MyGene to find aliases
for gene in unmapped_genes:
result = tu.tools.MyGene_query_genes(query=gene)
if result.get('hits'):
official_symbol = result['hits'][0].get('symbol')
print(f"{gene} -> {official_symbol}")
```
---
## gseapy Library Not Found
### Symptoms
```
ValueError: gene_set 'GO_Biological_Process' not found
```
### Solutions
#### 1. List Available Libraries
```python
import gseapy as gp
all_libs = gp.get_library_name(organism='human')
print(f"Total libraries: {len(all_libs)}")
# Search for library
matching = [lib for lib in all_libs if 'GO_Biological' in lib]
print("Available GO BP libraries:", matching)
```
#### 2. Use Exact Library Name
```python
# Wrong (will fail)
gene_sets='GO_Biological_Process'
# Correct
gene_sets='GO_Biological_Process_2021' # or 2023, 2025
```
#### 3. Check Case Sensitivity
```python
# Library names are case-sensitive
gene_sets='KEGG_2021_Human' # NOT 'kegg_2021_human'
```
#### 4. Organism-Specific Libraries
```python
# Wrong for mouse
gene_sets='KEGG_2021_Human' # Will fail for mouse genes
# Correct for mouse
gene_sets='KEGG_2019_Mouse'
```
---
## PANTHER Returns Empty Results
### Symptoms
```python
panther_result = tu.tools.PANTHER_enrichment(gene_list=genes, organism=9606, ...)
print(panther_result.get('data', {}).get('enriched_terms', [])) # Output: []
```
### Solutions
#### 1. Check Input Format
```python
# Wrong (array)
gene_list=["TP53", "BRCA1", "EGFR"]
# Correct (comma-separated string)
gene_list="TP53,BRCA1,EGFR"
```
#### 2. Check Organism ID
```python
# Wrong
organism='human' # PANTHER uses taxonomy IDs
# Correct
organism=9606 # NCBI Taxonomy ID for human
```
#### 3. Check Annotation Dataset
```python
# Valid datasets
annotation_datasets = [
'GO:0008150', # Biological Process
'GO:0003674', # Molecular Function
'GO:0005575', # Cellular Component
'ANNOT_TYPE_ID_PANTHER_PATHWAY',
'ANNOT_TYPE_ID_PANTHER_REACTOME_PATHWAY',
]
# Wrong
annotation_dataset='biological_process'
# Correct
annotation_dataset='GO:0008150'
```
#### 4. Gene Symbols Not Recognized
```python
# PANTHER uses official gene symbols
# Convert using STRING first
mapped = tu.tools.STRING_map_identifiers(
protein_ids=genes,
species=9606
)
official_symbols = [item['preferredName'] for item in mapped.get('data', [])]
gene_list = ','.join(official_symbols)
```
---
## STRING Returns All Categories
### Symptoms
```python
string_result = tu.tools.STRING_functional_enrichment(
protein_ids=genes,
species=9606,
category='Process' # Only want GO BP
)
# But result includes Process, Function, Component, KEGG, Reactome, etc.
```
### Solution
**This is expected behavior**. STRING always returns all categories. Filter after receiving results:
```python
string_data = string_result.get('data', [])
# Filter by category
string_go_bp = [d for d in string_data if d.get('category') == 'Process']
string_go_mf = [d for d in string_data if d.get('category') == 'Function']
string_go_cc = [d for d in string_data if d.get('category') == 'Component']
string_kegg = [d for d in string_data if d.get('category') == 'KEGG']
string_reactome = [d for d in string_data if d.get('category') == 'Reactome']
string_wp = [d for d in string_data if d.get('category') == 'WikiPathways']
print(f"GO BP terms: {len(string_go_bp)}")
print(f"KEGG pathways: {len(string_kegg)}")
```
---
## Reactome Identifiers Not Found
### Symptoms
```python
reactome_result = tu.tools.ReactomeAnalysis_pathway_enrichment(identifiers=genes, ...)
identifiers_not_found = reactome_result.get('data', {}).get('identifiers_not_found', [])
print(f"Not found: {len(identifiers_not_found)}") # High number
```
### Solutions
#### 1. Use Space-Separated String
```python
# Wrong (array)
identifiers=["TP53", "BRCA1", "EGFR"]
# Correct (space-separated string)
identifiers="TP53 BRCA1 EGFR"
# or
identifiers=' '.join(genes)
```
#### 2. Use Official Gene Symbols
```python
# Use STRING to get official names
mapped = tu.tools.STRING_map_identifiers(
protein_ids=genes,
species=9606
)
official_symbols = [item['preferredName'] for item in mapped.get('data', [])]
identifiers = ' '.join(official_symbols)
```
#### 3. Try UniProt IDs
```python
# Convert to UniProt IDs
mygene_result = tu.tools.MyGene_batch_query(
gene_ids=genes,
fields='uniprot.Swiss-Prot'
)
uniprot_ids = [hit['uniprot']['Swiss-Prot']
for hit in mygene_result.get('results', [])
if 'uniprot' in hit]
identifiers = ' '.join(uniprot_ids)
```
---
## enrichr_gene_enrichment_analysis Returns Wrong Format
### Symptoms
```python
result = tu.tools.enrichr_gene_enrichment_analysis(genes=genes, ...)
# Returns: connected_paths JSON string, NOT standard enrichment results
```
### Solution
**DO NOT USE `enrichr_gene_enrichment_analysis`** for standard enrichment analysis.
This ToolUniverse tool returns path analysis, not standard ORA enrichment.
**Use `gseapy.enrichr()` directly instead**:
```python
import gseapy
# Correct approach
result = gseapy.enrichr(
gene_list=genes,
gene_sets='GO_Biological_Process_2021',
organism='human',
outdir=None,
no_plot=True,
)
# Returns standard enrichment DataFrame
```
---
## Results Don't Match Between Tools
### Symptoms
```python
# gseapy finds 50 significant terms
# PANTHER finds 30 significant terms
# Only 10 overlap
```
### Explanation & Solutions
#### This is Normal
Different tools use:
- Different gene set annotations (versions, curation)
- Different backgrounds (genome-wide vs database-specific)
- Different statistical methods (Fisher's exact vs hypergeometric)
- Different multiple testing corrections
#### Focus on Consensus
```python
# Extract GO IDs from each source
import re
gseapy_terms = set()
for term in go_bp_sig['Term']:
match = re.search(r'(GO:\d+)', term)
if match:
gseapy_terms.add(match.group(1))
panther_terms = set(t['term_id'] for t in panther_bp_terms if t.get('fdr', 1) < 0.05)
string_terms = set(d['term'] for d in string_go_bp if d.get('fdr', 1) < 0.05)
# Consensus terms (in 2+ sources)
all_terms = gseapy_terms | panther_terms | string_terms
consensus = []
for term in all_terms:
sources = []
if term in gseapy_terms: sources.append('gseapy')
if term in panther_terms: sources.append('PANTHER')
if term in string_terms: sources.append('STRING')
if len(sources) >= 2:
consensus.append((term, sources))
print(f"Consensus terms: {len(consensus)}")
```
#### Report Both
```markdown
| GO Term | gseapy FDR | PANTHER FDR | STRING FDR | Consensus |
|---------|-----------|-------------|-----------|-----------|
| GO:0051726 | 3.4e-06 | 4.5e-05 | 3.2e-05 | 3/3 ✓ |
| GO:0006281 | 1.2e-05 | 2.1e-04 | - | 2/3 |
| GO:0007049 | 5.6e-06 | - | 8.9e-06 | 2/3 |
```
---
## GSEA Gives Unstable Results
### Symptoms
```python
# Run 1: 50 significant terms
# Run 2: 45 significant terms (different terms!)
```
### Solutions
#### 1. Set Random Seed
```python
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
seed=42, # CRITICAL for reproducibility
permutation_num=1000,
...
)
```
#### 2. Increase Permutation Number
```python
# More permutations = more stable p-values
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
seed=42,
permutation_num=5000, # instead of 1000
...
)
```
#### 3. Check for Ties
```python
# Check for many genes with same score
print(ranked_series.value_counts().head())
# If many ties, add small noise
import numpy as np
np.random.seed(42)
ranked_series = ranked_series + np.random.normal(0, 0.001, len(ranked_series))
```
---
## Memory Errors with Large Gene Lists
### Symptoms
```
MemoryError: Unable to allocate array
```
### Solutions
#### 1. Run Libraries Sequentially
```python
# Wrong (runs all at once)
result = gseapy.enrichr(
gene_list=genes,
gene_sets=['GO_Biological_Process_2021', 'GO_Molecular_Function_2021', ...], # 10+ libraries
...
)
# Correct (one at a time)
results = {}
for lib in ['GO_Biological_Process_2021', 'GO_Molecular_Function_2021', 'KEGG_2021_Human']:
results[lib] = gseapy.enrichr(gene_list=genes, gene_sets=lib, ...)
```
#### 2. Filter Gene Sets by Size
```python
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
min_size=10, # instead of 5
max_size=200, # instead of 500
...
)
```
#### 3. Use Smaller Libraries
```python
# Instead of GO_Biological_Process_2021 (17,000 terms)
# Use GO Slim (high-level terms only)
gene_sets='ANNOT_TYPE_ID_PANTHER_GO_SLIM_BP' # via PANTHER
```
---
## Getting Help
If you encounter an issue not covered here:
1. **Check gseapy documentation**: https://gseapy.readthedocs.io/
2. **Check tool parameter documentation**: `references/tool_parameters.md`
3. **Test with known good example**:
```python
# Known working example
genes = ["TP53", "BRCA1", "EGFR", "MYC", "AKT1", "PTEN"]
result = gseapy.enrichr(
gene_list=genes,
gene_sets='GO_Biological_Process_2021',
organism='human',
outdir=None,
no_plot=True,
)
```
---
See also:
- ora_workflow.md - Complete ORA examples
- gsea_workflow.md - Complete GSEA examples
- tool_parameters.md - Complete parameter reference