Falvyu

@falvyu.bsky.social

PhD | French | Hardware-aware Algorithm design | Image Processing | HPC SIMD-friends: #SSE, #AVX512, #NEON, #RVV (Opinions are my own)

The most underutilized #AVX512 instruction, VPSHUFBITQMB, is actually very useful for compactly generating a mask with arbitrary criteria based on b[5:0], e.g. to select special characters for parsing: if the second operand is a broadcasted 64-bit vector, where the ones represent the searched values

BildBildBild
nietras@nietras.bsky.social · last yr.

New blog post "Sep 0.10.0 - 21 GB/s CSV Parsing Using SIMD on AMD 9950X 🚀" 📈 Sep #performance from 7 GB/s to 21 GB/s over last two years 🧑‍💻 #csharp #SIMD and #x64 assembly on #dotnet 9.0 🛠️ Tweaks and new #AVX512-to-256 parser 🔢 Lots of benchmarks 👇 nietras.com/2025/05/09/s... Please repost 🙏

Sep performance progression.

Oh shit. I just realised no - it really is. I thought I was joking. This is going to require a long and complex backstory. Sorry absolutely not sorry it'll be funny maybe but also you'll learn something.

So long and thanks for all the fish! I just observed my last star this morning... #GaiaDPAC and @esa.int Gaia teams will do their best to make some great data releases from the data I gathered! #GaiaDR4 in 2026, #GaiaDR5 around the end of the decade. Keep posted for our news today at 10:00 CET.

rotating european space agency GIF - Find & Share on GIPHY

Discover & share this rotating european space agency GIF with everyone you know. GIPHY is how you search, share, discover, and create GIFs.

giphy.com

As Harold mentions, Prefix-Or is just v | -v. Segment-scan-or can also be simplified, for example uint64_t s = (v | m) >> 1; return s ^ (s - v) (4 instructions instead of 7).