Paul Harrison

@paulfharrison.bsky.social

Bioinformatician at Monash University, Melbourne, Australia. I also use mastodon: @pfh@mastondon.online https://mastodon.online/@pfh My homepage is: https://logarithmic.net/pfh/ On Twitter I was: @paulfharrison

Minor theme this year: There's a rough right number of files to split datasets into. For my data, on the order of 100s of files. Too few, can't process in parallel. Want to be writing in parallel now! Too many, filesystems go slow or break, especially network/cloud storage. HPC team becomes sad.

I'm trying using background workers in Shiny. Pattern: Cache results on disk. On a cache miss, launch-and-forget a background worker, tell Shiny to invalidate later, and throw an error for Shiny to display. Workers use file locks to avoid doubling up work. App remains responsive! #R #Shiny

Pondering k-Nearest Neighbor density estimation. There's some subtlety making the density smooth and integrate to 1. Here is a simple scheme: - Assign each point a radius from the distance to its kth nearest neighbor. - The density is the sum of a set of Gaussian splats with those radii.

grug brain bioinformatician not trust maximum a posteriori estimate. big brained bioinformatics shaman develop map estimate. danger! noise demon hide deeper in data! grug prefer count matrix. grug know what to do when have count matrix.

📢 PostDoc opportunity in our Bioinformatics & Cellular Genomics lab at SVI! 🧬 You’d join a welcoming, supportive, and brilliant team. Why not spend a few years in Melbourne and be part of something exciting? Apply here: www.seek.com.au/job/84737876 #ScienceCareers #PostDoc #Bioinformatics

Research Officer - Bioinformatics Job in Fitzroy, Melbourne VIC - SEEK

Seeking a Postdoc to develop computational toolkits to enable large-scale studies of single-cell and spatial 'omics and statistical genetics

seek.com.au

I continue to be astounded at the number of compositional data analysis packages that will happily report differential abundance of individual species. Did they not understand the concept of compositional data? How is it possible to publish methods with this premise?

Normalization and log transformation of log count data. Pseudocounts, library size adjustment, Centered Log Ratios (CLR), Variance Stabilizing Transformation, and all that. Many variations on a similar task. Here's something I haven't seen done:

An experimental design I've seen this three times in different contexts in the last couple of years. The RNA-Seq DE analysis turns out to be non-obvious: You apply treatment X to one group of subjects and treatment Y to another group. You have samples from before and after the treatment.

🧵 PCA is everywhere in bioinformatics—but did you know it’s just SVD in disguise? 1/ If you've done bioinformatics, you've likely used PCA. But did you know Singular Value Decomposition (SVD) is at its core? Let’s break it down. 👇

Bild

Comparison of optimization and sampling from a distribution defined by an energy function. I use a continuous version of the Ising model spin lattice energy. First, optimization from a random initial state using gradient descent with momentum, using the SGD optimizer in PyTorch.