HPC Guru

@hpcguru.bsky.social

Not a HPC Guru, but I play one on social media

Observation from #ISC25: None of the Big 3 cloud service providers have booths this year. Don’t recognize any of the big GPUaaS providers either, not even the ones who debuted on Top500 today. Plenty of quantum though. What’s this say about the state of AI in Europe? And cloud HPC for that matter.

Bild

Dynamic Use of Resources for Affordable Exascale and Beyond – BoF at #ISC25 Take aways for our sysadmin team – we need (rather sooner than later): 👉 Monitoring 👉 Al-Based Data Analytics 👉 Dynamic Scheduling & Resource Mgmt We‘ll definitely need to break some long standing dogmas as Martin says:

Dynamic Use of Resources for Affordable Exascale and Beyond – BoF at ISC25:

Martin Schulz (TUM)

Implementing all these ideas means breaking long standing dogmas in HPC

One Node = On Job
—> On-Node Co-Scheduling
—> Requires deep understanding and dynamic tracking of topology
==> Effective use of complementary accelerators.

Jobs have static resource usage from start to end
—> Dynamic Process Allocation
—> Requires new resource management interfaces and management approaches
==> Adapt to external conditions
(e.g., contention)

Worst case and fixed power distribution
—> PowerStack / SEANERGYS
—> Requires continuous monitoring and live feedback & interactions with the grid
==> Adaptive power steering on over-provisioned nodes

Panelist view from the full-house "ML/#AI for Workload Analysis in #HPC" BoF at #ISC25! Great talks and interaction - we went into "flipped classroom" and started asking questions to the audience. I very much enjoyed the invite from Kadida Konate (LBL) and Sergio Iserte (BSC).

Bild

At the #ISC25 vendor showdown, I think Bruno from Eviden captured a key difference between HPC and AI businesses. Paraphrasing, HPC people like to fuss with the system. AI customers don’t care about the system. Suggests why NVIDIA has been so successful. Meet the customers where they are.

Though not the intent of this slide, Erich pointed out that a bunch of GPUaaS/neocloud providers (Nebius, Core42, etc) are now submitting HPL numbers. These systems are still modest in size, but if the bigger ones like CoreWeave ever jumped aboard, the significance of Top500 might skew. #ISC25

Bild

If you define hyperscale to be the size of the largest #HPC datacenters, Top500 systems qualify as hyperscale. I’m not sure what the value of this proclamation is though. In real hyperscale, the qualifier is how many of these 50-200 MW sites you have, not just the fact that you have one. #ISC25

Bild

One thing I love about ISC is that Chinese researchers are given top billing. I was surprised to see focus on applying FPGAs for reverse time migration. I’ve not heard of others leading in this field thinking about that. I also didn’t realize China cared this much about #HPC for oil & gas. #ISC25

Bild

Rupak Biswas and I are ready to pose our questions to eight #ISC25 sponsors: Eviden, Broadcom, VAST Data, IBM, Starfish Storage, Cornelis Networks, Samsung, HPE

HPC-AI Leadership Organization (HALO)@hpcaileadershiporg.bsky.social · last yr.

Join @addisonsnell.bsky.social as he moderates the #ISC25 Vendor Showdown on Tues, 6/10 from 1:00-3:00 pm in Hall Z and have your chance to ask the difficult questions of the participating vendors! You’ll even be given the opportunity to rate the vendors. #ConnectingtheDots

The power wall facing #HPC seems increasingly like the Great Manure Crisis of 1984. Exascale in 20 MW turned out to be an arbitrary and false barrier. 1 GW for a zettaflop will be too. It will happen regardless of if people write PhDs on it or not. #ISC25

Bild