Ming Tommy Tang

@tommytang.bsky.social

Director of bioinformatics at AstraZeneca. subscribe to my youtube channel @chatomics. On my way to helping 1 million people learn bioinformatics. Educator, Biotech, single cell. Also talks about leadership. tommytang.bio.link

Bioinformatics is hard before you even write a single line of code. Here's why. 1/ You haven’t started your DNA-seq analysis. You haven’t aligned a read. And yet you’ve already hit a wall. Which human genome to use?

Bild

Good tutorial: Removing tumour purity, library size and batch effects from the TCGA breast cancer RNA-seq data using RUV-III-PRPS https://htmlpreview.github.io/?https://github.com/RMolania/TCGA_PanCancer_UnwantedVariation/blob/master/Vigettes/TCGA_BRCA_RNAseq_Vignette.html

Bild

Your data is lying to you. Here’s how technical artifacts distort biology—and how to see the truth. 👇 1/ Beautiful t-SNE? Shiny heatmap? Look closer. Technical artifacts can fake whole cell types. Here’s where the ghosts hide.

Bild

1/You know the feeling. You open a CSV or log file in the terminal— and it’s chaos. Wrapped lines. Misaligned columns. Impossible to read. Here’s how to turn that mess into clarity:

Bild

Fix your data before you implement AI. We all know this: garbage in, garbage out. How many of us are actually doing it? It takes money and resources, but that's the dirty work no one wants to do. AI won't fix your dirty data.

Almost every fusion transcript people report in plants is not real. A new Genome Biology paper looked at rice with long-read RNA-seq and found the vast majority are technical artifacts. Not biology. Not novel. Just noise your pipeline confidently labels as a fusion.

Bild