This skill should be used when users need to perform known motif enrichment analysis on ChIP-seq, ATAC-seq, or other genomic peak files using HOMER (Hypergeometric Optimization of Motif EnRichment)...
This skill enables comprehensive known motif enrichment analysis using HOMER tools for genomic peak files. It identifies enrichment of known transcription factor binding motifs from established databases in genomic regions.
Use this skill when you need to uncover the enrichment of a certain motif in the promoter regions of a set of genes, or directly from a set of genomic regions, such as peaks from ChIP-seq or ATAC-seq, with prior assumptions about which transcription factors are involved. Typical use cases include:
Input files should be in one of the following formats:
- BED files: Standard genomic interval format
- narrowPeak: narrow peak format
- broadPeak: broad peak format
- gene list: A list of genes provided by user or generated in previous analysis. May end with .txt, .tsv, .csv, etc.
${sample}_known_motif_enrichment/
results/
homerResults.html # De novo motif discovery results
seq.autonorm.tsv # Sequence composition statistics
motifFindingParameters.txt # Parameters used for analysis
homerMotifs.all.motifs
homerMotifs.motifs12
homerMotifs.motifs10
homerMotifs.motifs8
nonRedundant.motifs
homerResults/
motif1.similar1.motif
motif1.info.html
motif1.logo.svg
motif1.motif
motif1.similar.html
motif1.similar2.motif
motif1.similar3.motif
motif1.similar4.motif
motif1RV.logo.svg
motif1RV.motif
# ...
logs/ # analysis logs
motif.log
# ...
Before calling any tool, ask the user:
sample): used as prefix and for the output directory ${sample}_known_motif_enrichment.genome): e.g. hg38, mm10, danRer11. Call:
mcp__project-init-tools__project_initwith:
sample: the user-provided sample nametask: de_novo_motif_discoveryThe tool will:
${sample}_known_motif_enrichment directory.${sample}_known_motif_enrichment directory, which will be used as ${proj_dir}.Call:
mcp__homer-tools__check_genome_installationWith:
genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11The tool will:
This step is optional. Only perform this step if the input file is a BED file. If the input file is a gene list, skip this step.
From 1 format to chr1 format
From MT format to chrM format
Call:
mcp__file-format-tools__standardize_bed_chrom_nameswith:
input_bed: the user-provided BED fileoutput_bed: the path to save the standardized BED fileThe tool will:
If the user provides a TF name instead of a motif file, locate the motif file for the TF.
Call:
mcp__homer-tools__locate_motif_fileWith:
TF_name: the user-provided TF namemotif_type: Typically do not need to specify for model organisms. If the user provides data in "insects", "plants", "rna", "worms", "yeast", choose one as the appropriate motif type.The tool will:
Call:
mcp__homer-tools__find_motifsWith:
sample: the user-provided sample nameproj_dir: directory to save the de novo motif discovery results. In this skill, it is the full path of the ${sample}_known_motif_enrichment directory returned by mcp__project-init-tools__project_initinput_file: the user-provided file containing genome regions or gene list. May end with .bed, .narrowPeak, .broadPeak, .txt, .tsv, .csv, etc.genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11size: region size for motif finding for genome regions, typically 200-500bp for transcription factors (default: 200). If the input file is a gene list, set to None.mask: mask repeat regions for cleaner motif analysis (default: True)threads: number of processors to use (default: 4)num_motifs: number of motifs to find (default: 25)lengths: motif lengths to search (default: 8,10,12)nomotif: must set as True.mknown: motif file to use for enrichment analysis. May be the motif file returned by mcp__homer-tools__locate_motif_file.mcheck: motif file to check the enrichment of. May be the motif file returned by mcp__homer-tools__locate_motif_file.The tool will:
${proj_dir}/results/ directory.