8.8 KiB
title, task, lineage_type, upstream_source, upstream_sha, imported_at, prompt_class, upstream_changes, author, validated
| title | task | lineage_type | upstream_source | upstream_sha | imported_at | prompt_class | upstream_changes | author | validated |
|---|---|---|---|---|---|---|---|---|---|
| Troubleshooting Guide | import | https://github.com/mims-harvard/ToolUniverse/blob/e2520a96/skills/tooluniverse-single-cell/references/troubleshooting.md | e2520a96 | 2026-06-26 | prompt | accepted | upstream | false |
Troubleshooting Guide
Common errors and solutions for single-cell analysis.
Installation Issues
ModuleNotFoundError: scanpy
pip install scanpy anndata
ModuleNotFoundError: leidenalg
Leiden clustering requires leidenalg package:
pip install leidenalg
ModuleNotFoundError: umap
UMAP requires umap-learn:
pip install umap-learn
ModuleNotFoundError: harmonypy
Batch correction requires Harmony:
pip install harmonypy
Data Loading Issues
Wrong Matrix Orientation
Problem: Genes and samples/cells swapped
Solution: Check and transpose
# AnnData expects: cells as rows, genes as columns
if df.shape[0] > df.shape[1] * 5:
print("Transposing: genes were rows")
df = df.T
adata = ad.AnnData(df)
Index Mismatch
Problem: Metadata doesn't align with expression
Solution: Find common indices
common = adata.obs_names.intersection(meta.index)
adata = adata[common].copy()
for col in meta.columns:
adata.obs[col] = meta.loc[common, col]
Gene Name Mismatch
Problem: Gene names in different formats (Ensembl vs Symbol)
Solution: Convert IDs
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
result = tu.tools.MyGene_batch_query(
gene_ids=['ENSG00000141510', 'ENSG00000139618'],
fields='symbol,ensembl.gene'
)
# Extract symbols from result
Sparse Matrix Issues
TypeError: can't multiply sequence
Problem: Operating on sparse matrix incorrectly
Solution: Convert to dense
from scipy.sparse import issparse
X = adata.X.toarray() if issparse(adata.X) else adata.X
Memory Error
Problem: Dataset too large to convert to dense
Solution: Use sparse operations
# Good: Works on sparse
mean_expr = np.array(adata.X.mean(axis=0)).flatten()
# Bad: Converts entire matrix to dense
# mean_expr = adata.X.toarray().mean(axis=0)
QC and Filtering Issues
No mitochondrial genes found
Problem: Gene names don't start with "MT-"
Solution: Check prefix
# Check gene name format
print(adata.var_names[:10])
# Try different prefixes
adata.var['mt'] = adata.var_names.str.startswith(('MT-', 'mt-', 'Mt-'))
# Or manually specify
mt_genes = ['MT-CO1', 'MT-CO2', 'MT-ND1', ...] # Add all MT genes
adata.var['mt'] = adata.var_names.isin(mt_genes)
All cells filtered out
Problem: QC thresholds too stringent
Solution: Adjust thresholds
# Check distributions first
print(adata.obs['n_genes_by_counts'].describe())
print(adata.obs['pct_counts_mt'].describe())
# Adjust thresholds
sc.pp.filter_cells(adata, min_genes=100) # Lower from 200
adata = adata[adata.obs['pct_counts_mt'] < 25].copy() # Higher from 20
Clustering Issues
ValueError: n_neighbors must be less than n_samples
Problem: Too few cells for neighbor graph
Solution: Reduce n_neighbors
# Default is 15, reduce for small datasets
sc.pp.neighbors(adata, n_neighbors=min(15, adata.n_obs - 1))
Only one cluster
Problem: Resolution too low
Solution: Increase resolution
# Default is 1.0, increase for more clusters
sc.tl.leiden(adata, resolution=1.5) # or 2.0
Too many clusters
Problem: Resolution too high
Solution: Decrease resolution
sc.tl.leiden(adata, resolution=0.3) # or 0.5
Differential Expression Issues
ValueError: Not enough cells
Problem: < 3 cells per condition
Solution: Skip cell types with insufficient cells
n_treat = (adata_ct.obs['condition'] == 'treatment').sum()
n_ctrl = (adata_ct.obs['condition'] == 'control').sum()
if n_treat < 3 or n_ctrl < 3:
print(f"Skipping {cell_type}: insufficient cells")
continue
NaN values in results
Problem: Gene not expressed or zero variance
Solution: Filter before analysis
# Filter lowly expressed genes
sc.pp.filter_genes(adata, min_cells=3)
# Remove genes with zero variance
from sklearn.feature_selection import VarianceThreshold
selector = VarianceThreshold(threshold=0)
X_filtered = selector.fit_transform(X)
Statistical Issues
NaN in correlation
Problem: NaN values in data
Solution: Filter NaN before computing
valid = ~np.isnan(gene_lengths) & ~np.isnan(mean_expr)
r, p = stats.pearsonr(gene_lengths[valid], mean_expr[valid])
Inf values in log transform
Problem: log(0) = -inf
Solution: Use log1p or pseudocount
# Good: log1p(x) = log(x + 1)
sc.pp.log1p(adata)
# Good: log(x + pseudocount)
X_log = np.log10(X + 1)
# Bad: log(x) directly
# X_log = np.log10(X) # Will produce -inf for zeros
Performance Issues
Memory error on large datasets
Problem: Dataset too large to fit in memory
Solution: Subsample or use HVG
# Option 1: Subsample cells
sc.pp.subsample(adata, fraction=0.5)
# Option 2: Use only highly variable genes
sc.pp.highly_variable_genes(adata, n_top_genes=2000)
adata = adata[:, adata.var['highly_variable']].copy()
Slow clustering
Problem: Large dataset (>100k cells)
Solution: Reduce PCs or subsample
# Use fewer PCs
sc.pp.neighbors(adata, n_pcs=20) # instead of 30
# Or subsample for exploration
adata_sample = sc.pp.subsample(adata, fraction=0.1, copy=True)
# Analyze adata_sample first
ToolUniverse Integration Issues
API timeout
Problem: ToolUniverse API call times out
Solution: Reduce query size or retry
# Split large queries
batch_size = 50
for i in range(0, len(gene_list), batch_size):
batch = gene_list[i:i+batch_size]
result = tu.tools.MyGene_batch_query(gene_ids=batch)
ID conversion fails
Problem: Gene IDs not recognized
Solution: Try different ID types
# Try Ensembl ID
result = tu.tools.MyGene_query_genes(query="ENSG00000141510")
# Try gene symbol
result = tu.tools.MyGene_query_genes(query="TP53")
# Try with species
result = tu.tools.ensembl_lookup_gene(
gene_id="ENSG00000141510",
species="homo_sapiens"
)
OmniPath / Cell Communication Issues
No interactions found
Problem: Gene names don't match database
Solution: Check gene name format
# OmniPath uses gene symbols (not Ensembl IDs)
# Convert first if needed
result = tu.tools.STRING_map_identifiers(
protein_ids=['ENSG00000141510'],
species=9606
)
# Use preferredName from result
Empty communication matrix
Problem: Expression thresholds too stringent
Solution: Reduce thresholds
expressed_lr = filter_expressed_lr_pairs(
adata, lr_pairs,
min_frac=0.02, # Lower from 0.05
min_mean=0.01 # Lower from 0.05
)
Batch Correction Issues
Harmony not converging
Problem: Batch effects too strong
Solution: Adjust Harmony parameters
# Increase max iterations
ho = harmonypy.run_harmony(
adata.obsm['X_pca'],
adata.obs,
'batch',
max_iter_harmony=20 # Increase from default 10
)
Overcorrection
Problem: Biological variation removed
Solution: Use fewer PCs or weaker correction
# Use fewer PCs
ho = harmonypy.run_harmony(
adata.obsm['X_pca'][:, :20], # Instead of :30
adata.obs,
'batch'
)
File Format Issues
h5ad version mismatch
Problem: "Cannot read h5ad file"
Solution: Update anndata
pip install --upgrade anndata
10X format changed
Problem: Features file instead of genes file
Solution: Specify file names
# CellRanger v3+
adata = sc.read_10x_mtx(
"filtered_feature_bc_matrix/",
var_names='gene_symbols',
cache=True
)
Tips
- Always check data shape: Cells vs genes orientation
- Check for NaN/Inf: Before statistical tests
- Visualize before filtering: Check QC metric distributions
- Save intermediate results: After time-consuming steps
- Use try/except: For robust per-cell-type analysis
Debugging Checklist
- Data loaded correctly (check shape, obs, var)
- Matrix oriented correctly (cells as rows)
- Metadata aligned with expression data
- No NaN/Inf values in expression
- QC thresholds appropriate for dataset
- Sufficient cells per condition (>= 3)
- Gene names match between data and annotations
- Sparse matrix handled correctly
- Random seed set for reproducibility
See Also
- scanpy_workflow.md - Standard pipeline
- clustering_guide.md - Clustering troubleshooting
- cell_communication.md - OmniPath API issues