@vorontsovie.bsky.social

(1/13) Excited to share the outcome of the IBIS Challenge! The IBIS challenge united dozens of teams across the world in tackling the problem of modeling transcription factor (TF) binding specificity using a diverse collection of experimental datasets for understudied human TFs.

Bild

(8/12) With that many PWMs at hand, we also demonstrate that several motifs can be easily combined into a better model with ArChIPelago: regression or Random Forest on top of PWM scans.

Bild

(7/12) Comparison of different tools yielded many surprises. On the one hand, underused motif discovery tools such as Dimont excel across different types of experimental data. On the other hand, no single tool or platform is enough to get the most for each TF and each data type.

Bild

(6/12) In total, we processed data from 4,237 experiments and generated 219,939 motifs. 164,350 motifs (for 236 TFs) from curated and approved datasets were taken for benchmarking. The motifs and datasets are available at ZENODO zenodo.org/records/1018....

Codebook Motif Explorer Supplementary Dataset

This is the supplementary dataset of the GRECO-BIT/Codebook Motif Explorer (MEX),which is a detailed online catalog of DNA motifs built using a huge collection of experimental data produced by the Cod...

zenodo.org

(5/12) Testing all existing tools is hardly realistic. Focusing on position weight matrices (PWMs), we tested the most popular tools (e.g. MEME, HOMER), less popular yet powerful tools (e.g. ChIPMunk, Dimont), and a few advanced methods still able to yield PWMs (e.g. ExplaiNN, ProBound).

Bild

(4/12) We aimed to tackle both problems by running multiple motif discovery software against a large set of newly generated DNA on human TF-DNA interactions from Codebook [https://dx.doi.org/10.1101/2024.11.11.622097].

Bild

(3/12) There are a multitude of high-throughput methods for assessing the DNA binding specificity of transcription factors. The variety of existing motif discovery tools is even greater. Which one to use?

Bild

(2/12) The human genome encodes ~1.5 thousand transcription factors (TFs), and hundreds of them still lack "DNA motifs", i.e., compact human- and machine-readable representations of the TF-DNA binding specificity.

Bild

(5/7) Collectively the uncharacterized TFs bind directly to tens of thousands of conserved sequences, providing biochemical functions for these sites. Intriguingly, many of these sites are in genomic “dark matter”. Explored further in doi.org/10.1101/2024...

Bild

(3/7) Of the 332 uncharacterized TFs, PWMs were identified for 177, the vast majority of which have unique and previously unseen motifs.

Bild

(2/4) v13 covers >1100 of ~1600 human TFs with >1600 primary motifs and subtypes. Since v12 we also provide a reduced non-redundant set of motifs, which are often shared between TFs with similar DBDs.

Bild