Researchers have conducted a systematic characterization of LLM kernel memory access patterns through the lens of multi-partition NUMA effects, introducing a memory trace analysis methodology to derive workgroup-level data access and sharing behavior. arxiv.org/pdf/2607.28824
Underfox
@underfox3.bsky.social
Physicist, Telecom Engineering lover, HPC Enthusiast. Prog Rock/Metal fan. --- Independent tech analyst focused on semiconductors, patent analysis and emerging technologies.
In this paper, Huawei researchers have proposed the first end-to-end FP4 reinforcement learning post-training framework for LLMs, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. arxiv.org/pdf/2607.26515
In this paper are proposed two low-overhead instruction-reuse mechanisms for the NEORV32 RISC-V processor: a dynamic short-backward-branch-based loop cache and a software-managed static hot-code buffer with program-counter range matching. arxiv.org/pdf/2607.22792
In this paper, researchers have demonstrated the first monolithic non-volatile photonics platform on lithium tantalate-on-insulator (LTOI) for intrinsic, ferroelectric-domain-based non-volatile phase control without an added state-retentive material. arxiv.org/pdf/2607.23247
In this paper is presented the Marvell Photonic Fabric™ (PF™), a photonic-CXL hybrid architecture that replaces electrical switches with a passive fiber shuffle to deliver 32 TB of shared memory across 16 hosts via a switch-free full-crossbar topology. arxiv.org/pdf/2607.27187
In this paper is presented a cross-gen cost model for warp divergence in Ampere, Hopper and Blackwell GPUs, showing that divergence serializes linearly with a small constant per path and no super-linear reconvergence penalty up to a full 32-way split. arxiv.org/pdf/2607.23402
In this paper, Intel researchers have described the implementation of OpenMP Unified Shared Memory, explaining how to program with USM, and provided some initial performance results from adopting USM in a number of OpenMP applications on Intel GPUs. arxiv.org/pdf/2607.26584
Today, GlobalFoundries announced that it has entered into a letter of intent (LOI) with the U.S. Department of Commerce to accelerate R&D in next-generation silicon photonics, under which GF is expected to receive $300 million in funding. Press Release: gf.com/news-and-eve...
GlobalFoundries signs letter of intent with the U.S. Department of Commerce for a $300 million award to accelerate U.S. silicon photonics leadership | GlobalFoundries
GF has entered into an LOI with the U.S. Department of Commerce to accelerate research & development of next-generation silicon photonics.
gf.com
Researchers have demonstrated an approach toward the creation of all-vdW on-chip polarization optical components, providing a foundation for nanoscale photonic architectures based entirely on layered materials. arxiv.org/pdf/2607.21748
In this paper is presented SPDP, a unified sparse-inference framework that integrates unstructured static pruning with input-adaptive dynamic pruning for efficient LLM inference on GPUs. arxiv.org/pdf/2607.21985
In this paper, Meta researchers have introduced PRISM, a benchmark suite that reproduces representative AI research workloads to assess and qualify POSIX storage systems along both usability and performance dimensions on GPU clusters. arxiv.org/pdf/2607.21746
Researchers have developed an agentic scheduler that pairs an LLM agent with an algorithmic runtime monitor, where the monitor expands what the LLM can observe without ever prescribing which mapping to adopt. arxiv.org/pdf/2607.22242
In this paper is proposed TileSight, a tile-centric GPU performance-modeling tool that leverages the tile as both a programming primitive and an analysis primitive. arxiv.org/pdf/2607.22432
In this paper, researchers have proposed an open-source framework to enable complete on-device training on resource-constrained single-core RISC-V systems by leveraging the standard float16 Zfh and Zvfh RISC-V extensions. arxiv.org/pdf/2607.21130
In this paper, researchers have demonstrated the first on-chip radio-frequency maser operating at room temperature, exploiting optically pumped triplet states of pentacene. arxiv.org/pdf/2607.21002
In this paper, researchers have demonstrated an all-van der Waals high-β nanobeam laser based on a WS2/MoSe2/WS2 heterostructure, with the MoSe2 monolayer directly integrated into the WS2-based optical resonator for optimal gain-mode overlap. arxiv.org/pdf/2607.21566
In this paper, researchers have demonstrated the linear superposition of independently generated spin waves propagating in a CoFeB waveguide using a fully electrical excitation and detection scheme. arxiv.org/pdf/2607.20102
In this paper is presented DGNA, a framework to uncover the NUMA architecture within the GPU memory hierarchy through microbenchmarking and data analysis. arxiv.org/pdf/2607.19922
In this paper, Qualcomm researchers have established the first formally grounded treatment of Known Good Reliable Die screening for chiplet-based AI SoCs. arxiv.org/pdf/2607.20141
In this paper is presented HijackKV, the first framework that hijacks KV reuse without requiring system-level privileges, revealing a previously unknown vulnerability in modern LLM serving systems. arxiv.org/pdf/2607.19957
Base Computer researchers have developed a new family of hand-written Metal 4 cooperative-tensor kernels that drive the Apple M5's per-core GPU Neural Accelerators, covering dense GEMM, MoE expert GEMM, fused expert projections, and prefill attention. arxiv.org/pdf/2607.19438
In this paper is presented a kernel-centric characterization of Huawei Ascend 910A/910B/910C for scientific computing, showing how representative kernels interact with Cube Units, Vector Units, and the multi-tier memory hierarchy. arxiv.org/pdf/2607.20120
From time to time it becomes necessary to return to some fundamental discussions. Excerpt taken from "Resiliency in Numerical Algorithm Design for Extreme Scale Simulations", Agullo et al., Arxiv, 2020. arxiv.org/pdf/2010.13342
In this paper is presented a great review of quantum-enabled extreme sub-wavelength spintronic transmitting antennas, showing how these unprecedented features can write a new chapter in the fundamental science of antennas. arxiv.org/pdf/2607.18469
[Basic Concepts] In this paper is presented a selection of experiments that illustrate some of the key ways in which nanoscale systems reveal new aspects of thermodynamics which would not be apparent or even accessible in larger systems. arxiv.org/pdf/2607.16861
In this paper is presented ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure to learn controllable world dynamics. arxiv.org/pdf/2607.19191
In this paper is presented a scalable log analysis workflow for extracting, detecting, and interpreting fault-related patterns in large-scale HPC environments. arxiv.org/pdf/2607.19143
In this paper is presented a review on the emerging role of ferroionic 2D materials as a platform for programmable photonics, outlining how these material dynamics can be exploited to realize multi-level photonic states and nonvolatile phase control. arxiv.org/pdf/2607.18061
In this paper, researchers implemented a silicon undercut process for lithium-tantalate-on-insulator (LTOI) Mach-Zehnder modulators (MZMs), showing that the suspended electrode architecture provides a significant improvement in electro-optic bandwidth. arxiv.org/pdf/2607.17436
In this paper is presented Adaptive Mamba Neural Operator, the first neural operator that explicitly incorporates Takenaka-Malmquist systems and Fourier-based methods into the Mamba structure. arxiv.org/pdf/2607.18043