Underfox

@underfox3.bsky.social

Physicist, Telecom Engineering lover, HPC Enthusiast. Prog Rock/Metal fan. --- Independent tech analyst focused on semiconductors, patent analysis and emerging technologies.

Researchers have conducted a systematic characterization of LLM kernel memory access patterns through the lens of multi-partition NUMA effects, introducing a memory trace analysis methodology to derive workgroup-level data access and sharing behavior. arxiv.org/pdf/2607.28824

BildBildBild

In this paper, Huawei researchers have proposed the first end-to-end FP4 reinforcement learning post-training framework for LLMs, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. arxiv.org/pdf/2607.26515

BildBildBild

In this paper is presented the Marvell Photonic Fabric™ (PF™), a photonic-CXL hybrid architecture that replaces electrical switches with a passive fiber shuffle to deliver 32 TB of shared memory across 16 hosts via a switch-free full-crossbar topology. arxiv.org/pdf/2607.27187

BildBildBildBild

In this paper is presented a cross-gen cost model for warp divergence in Ampere, Hopper and Blackwell GPUs, showing that divergence serializes linearly with a small constant per path and no super-linear reconvergence penalty up to a full 32-way split. arxiv.org/pdf/2607.23402

BildBildBildBild

Today, GlobalFoundries announced that it has entered into a letter of intent (LOI) with the U.S. Department of Commerce to accelerate R&D in next-generation silicon photonics, under which GF is expected to receive $300 million in funding. Press Release: gf.com/news-and-eve...

GlobalFoundries signs letter of intent with the U.S. Department of Commerce for a $300 million award to accelerate U.S. silicon photonics leadership | GlobalFoundries

GF has entered into an LOI with the U.S. Department of Commerce to accelerate research & development of next-generation silicon photonics.

gf.com

Base Computer researchers have developed a new family of hand-written Metal 4 cooperative-tensor kernels that drive the Apple M5's per-core GPU Neural Accelerators, covering dense GEMM, MoE expert GEMM, fused expert projections, and prefill attention. arxiv.org/pdf/2607.19438

BildBild

In this paper, researchers implemented a silicon undercut process for lithium-tantalate-on-insulator (LTOI) Mach-Zehnder modulators (MZMs), showing that the suspended electrode architecture provides a significant improvement in electro-optic bandwidth. arxiv.org/pdf/2607.17436

Bild