14 KiB
14 KiB
title, task, lineage_type, upstream_source, upstream_sha, imported_at, prompt_class, upstream_changes, author, validated
| title | task | lineage_type | upstream_source | upstream_sha | imported_at | prompt_class | upstream_changes | author | validated |
|---|---|---|---|---|---|---|---|---|---|
| Troubleshooting Guide | import | https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse-gene-enrichment/references/troubleshooting.md | e2520a96 | 2026-06-26 | prompt | accepted | upstream | false |
Troubleshooting Guide
Complete troubleshooting guide for gene enrichment analysis issues.
Common Issues Index
- No Significant Enrichment Found
- Gene Not Found Errors
- gseapy Library Not Found
- PANTHER Returns Empty Results
- STRING Returns All Categories
- Reactome Identifiers Not Found
- enrichr_gene_enrichment_analysis Returns Wrong Format
- Results Don't Match Between Tools
- GSEA Gives Unstable Results
- Memory Errors with Large Gene Lists
No Significant Enrichment Found
Symptoms
go_bp_result = gseapy.enrichr(gene_list=genes, gene_sets='GO_Biological_Process_2021', ...)
go_bp_sig = go_bp_result.results[go_bp_result.results['Adjusted P-value'] < 0.05]
print(len(go_bp_sig)) # Output: 0
Possible Causes & Solutions
1. Gene Symbols Are Invalid
Diagnosis:
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
# Check if genes are recognized
mapped = tu.tools.STRING_map_identifiers(
protein_ids=genes,
species=9606
)
print(f"Mapped: {len(mapped['data'])}/{len(genes)}")
Solution:
- Convert IDs using MyGene:
tu.tools.MyGene_batch_query(gene_ids=genes, fields='symbol') - Use official HGNC symbols
- Remove version suffixes from Ensembl IDs
2. Library Version Too Old/New
Solution:
# Try different library versions
for lib_version in ['GO_Biological_Process_2021', 'GO_Biological_Process_2023', 'GO_Biological_Process_2025']:
result = gseapy.enrichr(gene_list=genes, gene_sets=lib_version, ...)
sig = result.results[result.results['Adjusted P-value'] < 0.05]
print(f"{lib_version}: {len(sig)} significant terms")
3. Significance Threshold Too Stringent
Solution:
# Try relaxing threshold
for cutoff in [0.05, 0.1, 0.15, 0.2]:
sig = go_bp_result.results[go_bp_result.results['Adjusted P-value'] < cutoff]
print(f"Cutoff {cutoff}: {len(sig)} significant terms")
4. Wrong Statistical Method
Solution:
# If ORA fails, try GSEA
import pandas as pd
# Create ranked list (even if you don't have fold-changes)
# Use presence/absence scoring: 1 for genes in list, 0 for others
ranked_dict = {gene: 1 for gene in genes}
ranked_series = pd.Series(ranked_dict).sort_values(ascending=False)
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
...
)
5. Organism Mismatch
Solution:
# Verify organism parameter
result = gseapy.enrichr(
gene_list=genes,
gene_sets='GO_Biological_Process_2021',
organism='human', # NOT 'homo_sapiens' or 9606
...
)
6. Gene List Too Small or Too Large
Diagnosis:
print(f"Gene list size: {len(genes)}")
# Small (<5): insufficient power
# Large (>500): too non-specific
Solution:
- Small lists (<5 genes): Use PANTHER/STRING (better for small lists) or switch to gene-level annotation
- Large lists (>500 genes): Use custom background or stricter DEG cutoff
7. No True Enrichment
Valid Result: Sometimes there genuinely is no enrichment. Document this:
## Results
No significant enrichment was found at FDR < 0.05.
- Gene list size: 45 genes
- Libraries tested: GO BP, GO MF, GO CC, KEGG, Reactome
- Alternative cutoff (FDR < 0.1): 3 weak hits (exploratory only)
Gene Not Found Errors
Symptoms
Error: Gene 'TP-53' not found
Warning: 15 out of 50 genes not mapped
Solutions
1. Check ID Type
# Detect ID type
sample = genes[0]
if sample.startswith('ENSG'):
id_type = 'ensembl'
elif sample.isdigit():
id_type = 'entrez'
elif len(sample) == 6 and sample[0].isalpha():
id_type = 'uniprot'
else:
id_type = 'symbol'
print(f"Detected ID type: {id_type}")
2. Convert IDs
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
# Convert Ensembl -> Symbol
if id_type == 'ensembl':
result = tu.tools.MyGene_batch_query(
gene_ids=genes,
fields='symbol,entrezgene,ensembl.gene'
)
symbol_map = {hit['query']: hit.get('symbol', hit['query'])
for hit in result.get('results', [])}
gene_symbols = [symbol_map.get(g, g) for g in genes]
3. Remove Version Suffixes
# ENSG00000141510.16 -> ENSG00000141510
genes_clean = [g.split('.')[0] for g in genes]
4. Fix Common Typos
# Common issues
fixes = {
'TP-53': 'TP53',
'BRCA 1': 'BRCA1',
'p53': 'TP53',
'egfr': 'EGFR', # case-sensitive
}
genes_fixed = [fixes.get(g, g) for g in genes]
5. Try Aliases
# Use MyGene to find aliases
for gene in unmapped_genes:
result = tu.tools.MyGene_query_genes(query=gene)
if result.get('hits'):
official_symbol = result['hits'][0].get('symbol')
print(f"{gene} -> {official_symbol}")
gseapy Library Not Found
Symptoms
ValueError: gene_set 'GO_Biological_Process' not found
Solutions
1. List Available Libraries
import gseapy as gp
all_libs = gp.get_library_name(organism='human')
print(f"Total libraries: {len(all_libs)}")
# Search for library
matching = [lib for lib in all_libs if 'GO_Biological' in lib]
print("Available GO BP libraries:", matching)
2. Use Exact Library Name
# Wrong (will fail)
gene_sets='GO_Biological_Process'
# Correct
gene_sets='GO_Biological_Process_2021' # or 2023, 2025
3. Check Case Sensitivity
# Library names are case-sensitive
gene_sets='KEGG_2021_Human' # NOT 'kegg_2021_human'
4. Organism-Specific Libraries
# Wrong for mouse
gene_sets='KEGG_2021_Human' # Will fail for mouse genes
# Correct for mouse
gene_sets='KEGG_2019_Mouse'
PANTHER Returns Empty Results
Symptoms
panther_result = tu.tools.PANTHER_enrichment(gene_list=genes, organism=9606, ...)
print(panther_result.get('data', {}).get('enriched_terms', [])) # Output: []
Solutions
1. Check Input Format
# Wrong (array)
gene_list=["TP53", "BRCA1", "EGFR"]
# Correct (comma-separated string)
gene_list="TP53,BRCA1,EGFR"
2. Check Organism ID
# Wrong
organism='human' # PANTHER uses taxonomy IDs
# Correct
organism=9606 # NCBI Taxonomy ID for human
3. Check Annotation Dataset
# Valid datasets
annotation_datasets = [
'GO:0008150', # Biological Process
'GO:0003674', # Molecular Function
'GO:0005575', # Cellular Component
'ANNOT_TYPE_ID_PANTHER_PATHWAY',
'ANNOT_TYPE_ID_PANTHER_REACTOME_PATHWAY',
]
# Wrong
annotation_dataset='biological_process'
# Correct
annotation_dataset='GO:0008150'
4. Gene Symbols Not Recognized
# PANTHER uses official gene symbols
# Convert using STRING first
mapped = tu.tools.STRING_map_identifiers(
protein_ids=genes,
species=9606
)
official_symbols = [item['preferredName'] for item in mapped.get('data', [])]
gene_list = ','.join(official_symbols)
STRING Returns All Categories
Symptoms
string_result = tu.tools.STRING_functional_enrichment(
protein_ids=genes,
species=9606,
category='Process' # Only want GO BP
)
# But result includes Process, Function, Component, KEGG, Reactome, etc.
Solution
This is expected behavior. STRING always returns all categories. Filter after receiving results:
string_data = string_result.get('data', [])
# Filter by category
string_go_bp = [d for d in string_data if d.get('category') == 'Process']
string_go_mf = [d for d in string_data if d.get('category') == 'Function']
string_go_cc = [d for d in string_data if d.get('category') == 'Component']
string_kegg = [d for d in string_data if d.get('category') == 'KEGG']
string_reactome = [d for d in string_data if d.get('category') == 'Reactome']
string_wp = [d for d in string_data if d.get('category') == 'WikiPathways']
print(f"GO BP terms: {len(string_go_bp)}")
print(f"KEGG pathways: {len(string_kegg)}")
Reactome Identifiers Not Found
Symptoms
reactome_result = tu.tools.ReactomeAnalysis_pathway_enrichment(identifiers=genes, ...)
identifiers_not_found = reactome_result.get('data', {}).get('identifiers_not_found', [])
print(f"Not found: {len(identifiers_not_found)}") # High number
Solutions
1. Use Space-Separated String
# Wrong (array)
identifiers=["TP53", "BRCA1", "EGFR"]
# Correct (space-separated string)
identifiers="TP53 BRCA1 EGFR"
# or
identifiers=' '.join(genes)
2. Use Official Gene Symbols
# Use STRING to get official names
mapped = tu.tools.STRING_map_identifiers(
protein_ids=genes,
species=9606
)
official_symbols = [item['preferredName'] for item in mapped.get('data', [])]
identifiers = ' '.join(official_symbols)
3. Try UniProt IDs
# Convert to UniProt IDs
mygene_result = tu.tools.MyGene_batch_query(
gene_ids=genes,
fields='uniprot.Swiss-Prot'
)
uniprot_ids = [hit['uniprot']['Swiss-Prot']
for hit in mygene_result.get('results', [])
if 'uniprot' in hit]
identifiers = ' '.join(uniprot_ids)
enrichr_gene_enrichment_analysis Returns Wrong Format
Symptoms
result = tu.tools.enrichr_gene_enrichment_analysis(genes=genes, ...)
# Returns: connected_paths JSON string, NOT standard enrichment results
Solution
DO NOT USE enrichr_gene_enrichment_analysis for standard enrichment analysis.
This ToolUniverse tool returns path analysis, not standard ORA enrichment.
Use gseapy.enrichr() directly instead:
import gseapy
# Correct approach
result = gseapy.enrichr(
gene_list=genes,
gene_sets='GO_Biological_Process_2021',
organism='human',
outdir=None,
no_plot=True,
)
# Returns standard enrichment DataFrame
Results Don't Match Between Tools
Symptoms
# gseapy finds 50 significant terms
# PANTHER finds 30 significant terms
# Only 10 overlap
Explanation & Solutions
This is Normal
Different tools use:
- Different gene set annotations (versions, curation)
- Different backgrounds (genome-wide vs database-specific)
- Different statistical methods (Fisher's exact vs hypergeometric)
- Different multiple testing corrections
Focus on Consensus
# Extract GO IDs from each source
import re
gseapy_terms = set()
for term in go_bp_sig['Term']:
match = re.search(r'(GO:\d+)', term)
if match:
gseapy_terms.add(match.group(1))
panther_terms = set(t['term_id'] for t in panther_bp_terms if t.get('fdr', 1) < 0.05)
string_terms = set(d['term'] for d in string_go_bp if d.get('fdr', 1) < 0.05)
# Consensus terms (in 2+ sources)
all_terms = gseapy_terms | panther_terms | string_terms
consensus = []
for term in all_terms:
sources = []
if term in gseapy_terms: sources.append('gseapy')
if term in panther_terms: sources.append('PANTHER')
if term in string_terms: sources.append('STRING')
if len(sources) >= 2:
consensus.append((term, sources))
print(f"Consensus terms: {len(consensus)}")
Report Both
| GO Term | gseapy FDR | PANTHER FDR | STRING FDR | Consensus |
|---------|-----------|-------------|-----------|-----------|
| GO:0051726 | 3.4e-06 | 4.5e-05 | 3.2e-05 | 3/3 ✓ |
| GO:0006281 | 1.2e-05 | 2.1e-04 | - | 2/3 |
| GO:0007049 | 5.6e-06 | - | 8.9e-06 | 2/3 |
GSEA Gives Unstable Results
Symptoms
# Run 1: 50 significant terms
# Run 2: 45 significant terms (different terms!)
Solutions
1. Set Random Seed
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
seed=42, # CRITICAL for reproducibility
permutation_num=1000,
...
)
2. Increase Permutation Number
# More permutations = more stable p-values
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
seed=42,
permutation_num=5000, # instead of 1000
...
)
3. Check for Ties
# Check for many genes with same score
print(ranked_series.value_counts().head())
# If many ties, add small noise
import numpy as np
np.random.seed(42)
ranked_series = ranked_series + np.random.normal(0, 0.001, len(ranked_series))
Memory Errors with Large Gene Lists
Symptoms
MemoryError: Unable to allocate array
Solutions
1. Run Libraries Sequentially
# Wrong (runs all at once)
result = gseapy.enrichr(
gene_list=genes,
gene_sets=['GO_Biological_Process_2021', 'GO_Molecular_Function_2021', ...], # 10+ libraries
...
)
# Correct (one at a time)
results = {}
for lib in ['GO_Biological_Process_2021', 'GO_Molecular_Function_2021', 'KEGG_2021_Human']:
results[lib] = gseapy.enrichr(gene_list=genes, gene_sets=lib, ...)
2. Filter Gene Sets by Size
gsea_result = gseapy.prerank(
rnk=ranked_series,
gene_sets='GO_Biological_Process_2021',
min_size=10, # instead of 5
max_size=200, # instead of 500
...
)
3. Use Smaller Libraries
# Instead of GO_Biological_Process_2021 (17,000 terms)
# Use GO Slim (high-level terms only)
gene_sets='ANNOT_TYPE_ID_PANTHER_GO_SLIM_BP' # via PANTHER
Getting Help
If you encounter an issue not covered here:
- Check gseapy documentation: https://gseapy.readthedocs.io/
- Check tool parameter documentation:
references/tool_parameters.md - Test with known good example:
# Known working example genes = ["TP53", "BRCA1", "EGFR", "MYC", "AKT1", "PTEN"] result = gseapy.enrichr( gene_list=genes, gene_sets='GO_Biological_Process_2021', organism='human', outdir=None, no_plot=True, )
See also:
- ora_workflow.md - Complete ORA examples
- gsea_workflow.md - Complete GSEA examples
- tool_parameters.md - Complete parameter reference