tooluniverse-microbiome-research
Microbiome research using MGnify, GTDB, ENA, OLS (ENVO biomes), EuropePMC, BacDive (strain phenotypes), MediaDive (growth media recipes), and GMrepo (gut microbiome disease-association counts). Covers study discovery, taxonomic profiling, host-microbe interaction analysis, biome-by-condition queries, strain culturing requirements, and phenotype/disease prevalence lookups. Use for microbiome study selection, organism-environment associations, clinical-microbiome literature review, and "how do I grow this organism" / "how prevalent is this species in condition X" questions. Distinct from analytical workflow (use tooluniverse-metagenomics-analysis for that).
How do I install this agent skill?
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-microbiome-researchIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill facilitates microbiome research by querying reputable scientific databases via the ToolUniverse library. It is generally safe but carries a low risk of indirect prompt injection due to the ingestion of external data (abstracts, metadata) without explicit boundary markers.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Microbiome Research with ToolUniverse
Comprehensive microbiome analysis using MGnify (EBI metagenomics), GTDB (genome taxonomy), ENA (sequencing data), OLS (ontology lookup for ENVO biomes), and EuropePMC (literature).
Core Tools
| Tool | Purpose | Auth |
|---|---|---|
| MGnify_search_studies | Find metagenomics studies by biome/keyword | None |
| MGnify_get_study_detail | Study metadata, abstract, sample counts | None |
| MGnify_list_analyses | List taxonomic/functional analysis outputs for a study | None |
| MGnify_get_taxonomy | Taxonomic composition from an analysis | None |
| MGnify_get_go_terms | GO functional annotations from an analysis | None |
| MGnify_get_interpro | InterPro protein domain annotations | None |
| MGnify_list_biomes | Browse MGnify biome hierarchy | None |
| MGnify_search_genomes | Search metagenome-assembled genomes (MAGs) | None |
| MGnify_get_genome | Genome quality metrics (completeness, contamination) | None |
| GTDB_search_genomes | Search bacterial/archaeal genomes by taxonomy | None |
| GTDB_get_species | Species cluster details from GTDB | None |
| GTDB_get_taxon_info | Taxonomic rank info in GTDB hierarchy | None |
| GTDB_search_taxon | Search taxa by partial name across all ranks | None |
| ENAPortal_search_studies | Find sequencing studies in ENA. Query format: description="keyword" | None |
| ENAPortal_search_samples | Find samples with environmental metadata | None |
| ols_search_terms | Search ENVO ontology for biome/environment terms | None |
| EuropePMC_search_articles | Find microbiome publications | None |
| PubMed_search_articles | Literature search (different coverage than EuropePMC) | None |
| BacDive_search_by_taxon | List BacDive strain IDs for a genus (+ optional species) | None |
| BacDive_get_strain | Curated strain phenotype: morphology, Gram stain, culture temp/media, oxygen tolerance, isolation source | None |
| MediaDive_search_media | Find a cultivation medium by name substring | None |
| MediaDive_get_medium | Full recipe (ingredients + amounts + prep steps) for a medium | None |
| MediaDive_search_ingredients / MediaDive_get_ingredient | Look up a medium ingredient's chemical identifiers (CAS/ChEBI/PubChem) | None |
| GMrepo_search_species | Gut-microbiome sample counts and phenotype breadth for a species (96,000+ curated samples) | None |
| GMrepo_get_phenotypes | List/filter the 350+ health conditions GMrepo has gut-microbiome data for | None |
For drug-microbiome studies, also use:
PubChem_get_CID_by_compound_name/PubChem_get_compound_properties_by_CID— drug identityCTD_get_chemical_gene_interactions— drug-gene interactions (e.g., metformin affects 1,175+ genes)kegg_search_pathway/kegg_get_pathway_info— microbial metabolic pathways (butanoate, propanoate)ReactomeAnalysis_pathway_enrichment— host pathway enrichment for drug-affected genesdrugbank_vocab_search— drug mechanism and targets
MGnify tip: Use concise single-keyword searches (e.g., "metformin") — multi-word queries may timeout. The MGnify API can be slow for broad searches.
Quick Start
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()
# 1. Search for gut microbiome studies
studies = tu.run_one_function({
'name': 'MGnify_search_studies',
'arguments': {'search': 'gut microbiome', 'size': 5}
})
# 2. Get study details
detail = tu.run_one_function({
'name': 'MGnify_get_study_detail',
'arguments': {'study_accession': 'MGYS00006860'}
})
# 3. List analyses for a study
analyses = tu.run_one_function({
'name': 'MGnify_list_analyses',
'arguments': {'study_accession': 'MGYS00006860', 'size': 5}
})
# 4. Get taxonomic profile from an analysis
taxonomy = tu.run_one_function({
'name': 'MGnify_get_taxonomy',
'arguments': {'analysis_id': 'MGYA00793746'}
})
# 5. Get functional annotations
go_terms = tu.run_one_function({
'name': 'MGnify_get_go_terms',
'arguments': {'analysis_id': 'MGYA00793746'}
})
Common Workflows
Workflow 1: Study Discovery by Environment
Find studies for a specific biome using MGnify's biome hierarchy:
# Browse biome hierarchy
biomes = tu.run_one_function({
'name': 'MGnify_list_biomes',
'arguments': {'lineage': 'root:Host-associated:Human', 'depth': 3}
})
# Search studies in a specific biome
studies = tu.run_one_function({
'name': 'MGnify_search_studies',
'arguments': {'biome': 'root:Host-associated:Human:Digestive system', 'size': 10}
})
# Look up ENVO ontology terms for environment metadata
envo = tu.run_one_function({
'name': 'ols_search_terms',
'arguments': {'query': 'human gut', 'ontology': 'envo', 'rows': 5}
})
Workflow 2: Taxonomic Profiling
Get the microbial composition of a metagenomics sample:
# Get analyses for a study
analyses = tu.run_one_function({
'name': 'MGnify_list_analyses',
'arguments': {'study_accession': 'MGYS00006860', 'size': 3}
})
# Get taxonomy for a specific analysis
taxonomy = tu.run_one_function({
'name': 'MGnify_get_taxonomy',
'arguments': {'analysis_id': 'MGYA00793746'}
})
# Returns organisms with lineage, abundance counts, and taxonomy rank
Workflow 3: Genome Quality Assessment
Evaluate metagenome-assembled genomes (MAGs):
# Search for genomes from a specific taxon
genomes = tu.run_one_function({
'name': 'MGnify_search_genomes',
'arguments': {'taxonomy': 'Faecalibacterium prausnitzii', 'page_size': 5}
})
# Get quality metrics for a genome
genome = tu.run_one_function({
'name': 'MGnify_get_genome',
'arguments': {'genome_id': 'MGYG000000001'}
})
# Returns completeness, contamination, N50, genome length, taxonomy
# Cross-reference with GTDB taxonomy
gtdb = tu.run_one_function({
'name': 'GTDB_search_genomes',
'arguments': {'operation': 'search_genomes', 'query': 'Faecalibacterium', 'items_per_page': 5}
})
Workflow 4: Functional Annotation
Discover functional potential of a metagenome:
# GO terms from an analysis
go_terms = tu.run_one_function({
'name': 'MGnify_get_go_terms',
'arguments': {'analysis_id': 'MGYA00793746'}
})
# InterPro domains
interpro = tu.run_one_function({
'name': 'MGnify_get_interpro',
'arguments': {'analysis_id': 'MGYA00793746'}
})
Workflow 5: Literature Integration
Combine metagenomics data with published research:
# Find relevant publications
papers = tu.run_one_function({
'name': 'EuropePMC_search_articles',
'arguments': {'query': 'gut microbiome AND Faecalibacterium AND (IBD OR "Crohn")', 'limit': 10}
})
# Find sequencing data in ENA
ena_studies = tu.run_one_function({
'name': 'ENAPortal_search_studies',
'arguments': {'query': 'description="gut microbiome 16S"', 'limit': 5}
})
Workflow 6: Strain Phenotype, Culturing Requirements, and Disease-Association Prevalence
Answer "what is this organism actually like and how prevalent is it in condition X" — a question MGnify/GTDB (composition and taxonomy only) can't answer:
# 1. Find BacDive strain IDs for a taxon (genus required, species optional)
strains = tu.run_one_function({
'name': 'BacDive_search_by_taxon',
'arguments': {'genus': 'Faecalibacterium', 'species': 'prausnitzii', 'limit': 3}
})
# -> [{'bacdive_id': 159475}, {'bacdive_id': 159476}, {'bacdive_id': 159477}]
# 2. Get the curated phenotype for one strain
strain = tu.run_one_function({
'name': 'BacDive_get_strain',
'arguments': {'bacdive_id': 159475}
})
# -> anaerobe, mesophilic (37C), isolated from human feces, grows on
# "YCFA-MEDIUM (MODIFIED) (DSMZ Medium 1611)" and chopped-meat medium
# 3. Cross-reference the named medium in MediaDive for the actual recipe
medium = tu.run_one_function({
'name': 'MediaDive_search_media',
'arguments': {'query': 'YCFA'}
})
# -> medium_id 1611 matches the BacDive culture_media entry exactly
recipe = tu.run_one_function({
'name': 'MediaDive_get_medium',
'arguments': {'medium_id': 1611}
})
# -> full solution-by-solution ingredient list with amounts/units and prep steps
# 4. Check GMrepo for how prevalent/associated the species is in human gut studies
prevalence = tu.run_one_function({
'name': 'GMrepo_search_species',
'arguments': {'query': 'Faecalibacterium prausnitzii'}
})
# -> present in 21,286 of ~96,000 samples (22.0%), spanning 88 phenotypes
conditions = tu.run_one_function({
'name': 'GMrepo_get_phenotypes',
'arguments': {'query': 'Crohn'}
})
# -> Crohn Disease (MeSH D003424): 5,636 samples across 1,440 species, 602 genera
Reading BacDive's culture_media field: it names the medium (e.g. "YCFA-MEDIUM (MODIFIED) (DSMZ Medium 1611)") but gives no recipe — always follow up in MediaDive by searching the medium name to find its medium_id, then MediaDive_get_medium for the actual formulation. oxygen_tolerance can legitimately be an empty array (not curated for that strain) — don't infer aerobe/anaerobe status from an empty list.
GMrepo scope note: presented_samples/all_samples are curated sample COUNTS from public studies, not a claim about true population prevalence — a species absent from GMrepo's phenotype list for a condition may simply not have been studied there yet, not proven absent from that condition's gut microbiome.
MGnify Biome Hierarchy
Key biome lineages (use MGnify_list_biomes to discover others):
- Human gut:
root:Host-associated:Human:Digestive system - Human oral/skin:
root:Host-associated:Human:Oral/root:Host-associated:Human:Skin - Soil:
root:Environmental:Terrestrial:Soil - Ocean/Freshwater:
root:Environmental:Aquatic:Marine/root:Environmental:Aquatic:Freshwater - Wastewater:
root:Engineered:Wastewater
Key Identifiers
MGnify: studies=MGYS*, analyses=MGYA*, genomes=MGYG*. ENA studies=PRJEB*. GTDB genomes=GCA_*. ENVO terms=ENVO:* (e.g. ENVO:00002041). BacDive strain IDs and MediaDive medium/ingredient IDs are plain integers (occasionally suffixed, e.g. "1a") with no fixed prefix — always resolve via search first. GMrepo phenotypes use MeSH IDs (e.g. D003424 for Crohn Disease).
Reasoning Framework
Starting Point: Define the Question First
Microbiome analysis starts with: what is the question? LOOK UP DON'T GUESS — always check the study type and sequencing method before interpreting results.
Decision tree for data type:
- Community composition (who is there?) → 16S/ITS amplicon → alpha/beta diversity, differential abundance
- Functional potential (what can they do?) → Shotgun metagenomics → MGnify GO terms, InterPro, KEGG pathways
- Active function (what are they doing now?) → Metatranscriptomics → specialized pipelines (not MGnify/GTDB alone)
Before calling any tool, determine which data type the user has via MGnify_get_study_detail — the pipeline type (amplicon vs shotgun) determines which analyses are valid. Do not apply 16S diversity metrics to metagenomic data or vice versa.
Dysbiosis Assessment Strategy
Dysbiosis (microbial imbalance) is context-dependent — there is no universal "healthy" microbiome. LOOK UP DON'T GUESS — compare to study-matched controls, not general population references.
- Check alpha diversity: Reduced Shannon index relative to controls suggests dysbiosis. Use
MGnify_get_taxonomyto get community profiles, then assess richness and evenness. - Identify keystone taxa shifts: Loss of known beneficial taxa (e.g., Faecalibacterium, Roseburia in gut) or bloom of pathobionts (e.g., Enterobacteriaceae). LOOK UP taxa roles with
GTDB_get_speciesand literature viaEuropePMC_search_articles. - Functional consequences: Does taxonomic shift correlate with loss/gain of metabolic pathways? Check
MGnify_get_go_termsandMGnify_get_interprofor the affected samples. - Confounders: Antibiotics, diet, age, and geography all affect microbiome composition. A dysbiosis claim requires controlling for these factors or acknowledging them as limitations.
Taxonomic vs Functional Analysis: When to Use Each
- Taxonomic analysis alone is sufficient when the question is "which organisms are present?" or "does community composition differ between groups?" Use
MGnify_get_taxonomy+GTDB_search_genomes. - Functional analysis is needed when the question is "what metabolic capabilities differ?" or "why does a taxonomic shift matter?" Use
MGnify_get_go_terms+MGnify_get_interpro+kegg_search_pathway. - Both together when linking organisms to functions (e.g., "which taxa drive butyrate production in healthy vs IBD gut?"). Cross-reference taxonomic profiles with functional annotations from the same MGnify analysis.
Evidence Grading
| Tier | Description | Example |
|---|---|---|
| T1 | Replicated finding across multiple cohorts with consistent effect | Reduced Faecalibacterium in IBD (>10 independent studies) |
| T2 | Single well-powered study (n > 100) with appropriate controls | Metformin-associated Akkermansia enrichment in a controlled trial |
| T3 | Pilot study or observational association, small sample size | Taxonomic shift in n=15 case-control, no validation cohort |
| T4 | Computational prediction or single-sample observation | Novel MAG with predicted function, no culture confirmation |
Interpretation Guidance
Alpha diversity (within-sample): Shannon index measures richness and evenness. Higher Shannon (>3.0 for gut) suggests a stable community. Reduced alpha diversity is associated with dysbiosis (IBD, antibiotics). Always compare to study-matched controls — diversity varies by body site, sequencing depth, and geography.
Beta diversity (between-sample): Bray-Curtis (abundance-based) or UniFrac (phylogenetic). PERMANOVA p < 0.05 with R-squared > 0.05 indicates condition-driven clustering. Low R-squared (<0.02) even with significant p suggests the effect is small relative to inter-individual variation. Choose weighted UniFrac when abundant taxa matter most; unweighted when rare taxa are important.
Taxonomic composition: Relative abundance at phylum level (Firmicutes/Bacteroidetes ratio) is a coarse indicator; genus- or species-level resolution is preferred. A taxon present at >1% relative abundance in multiple samples is reliably detected. Taxa at <0.1% may be noise or sequencing artifacts. GTDB taxonomy may reclassify NCBI names (e.g., Firmicutes split into multiple phyla).
Functional profiling: GO terms and InterPro domains from MGnify reflect the metabolic potential (not necessarily activity) of the community. Enrichment of specific pathways (e.g., butyrate production, LPS biosynthesis) should be interpreted alongside taxonomic data to identify which organisms contribute the functions.
Synthesis Questions
A complete microbiome report should answer:
- How does alpha diversity compare between conditions, and is the difference significant?
- Does beta diversity analysis show condition-driven clustering (PERMANOVA)?
- Which taxa are differentially abundant, and are they known commensals or pathobionts?
- What functional pathways are enriched, and which taxa likely drive them?
- How do findings compare to published studies for the same biome/condition (literature context)?
Tips
- MGnify study accessions start with
MGYS, analyses withMGYA, genomes withMGYG - Use
MGnify_list_biomesfirst to find the correct biome lineage string MGnify_get_taxonomyreturns phylum-level to species-level composition- GTDB provides standardized bacterial/archaeal taxonomy (differs from NCBI in some lineages)
- For 16S amplicon studies, taxonomy is the primary output; for shotgun metagenomics, both taxonomy and functional annotations are available
- The
sizeparameter in MGnify tools controls results per page (max 100)
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/mims-harvard/tooluniverse/tooluniverse-microbiome-research">View tooluniverse-microbiome-research on skillZs</a>