@hochreitersepp.bsky.social

xLSTM for three-dimensional bin packing problem: arxiv.org/abs/2607.28257 xLSTM mixes item and candidate tokens with a recurrent action mixer (LRAM). xLSTM based OPAL has mean space utilization of 0.49 (15.1% improvement), while being fast. Promising for high-fidelity robotics.

Bild

xLSTM for Brain MRI Analysis: arxiv.org/abs/2607.17782 Three-dimensional Bi-Directional xLSTM-UNet achieved second place overall and first in the meningioma segmentation task on FOMO 2025 leaderboard. xLSTM models long-range dependencies throughout the volumetric brain MRI.

Bild

1/5 🚀 Introducing TiRex-2 — our next-generation time series foundation model. Time series forecasting in the real world is streaming: • new observations continuously arrive, • variables interact, • some covariates are known into the future, • models must update predictions efficiently.

Comparison of sub-quadratic architectures xLSTM, Mamba-2, and Gated DeltaNet: arxiv.org/abs/2606.12364 Comparison of xLSTM, Mamba-2, and Gated DeltaNet on code pre-training, distillation, and time-series. xLSTM outperforms the others due to its gating scheme and state tracking.

Bild

xLSTM is more expressive than Transformer, Mamba: arxiv.org/abs/2603.03612 *nonlinear RNNs: sLSTM, LSTM *DLPR linear RNNs: mLSTM, RWKV-7, DeltaNet *Non PNC1-complete: Mamba, Transformer “fundamental expressivity gaps between linear and nonlinear RNNs.” World models require nonlinear RNNs.

BildBild

xLSTM Distillation: arxiv.org/abs/2603.15590 Near-lossless distillation of quadratic Transformer LLMs into linear-time xLSTM architectures enables cost- and energy-efficient alternatives without sacrificing performance. Efficient xLSTM variants of instruction-tuned Llama, Qwen, and Olmo models.

BildBild

xLSTM for Financial Time Series: arxiv.org/abs/2603.01820 "VLSTM achieved the highest overall Sharpe ratio" "VxLSTM and LPatchTST exhibited superior downside-adjusted characteristics" “xLSTM achieves the highest portfolio-level cost buffer" xLSTM excels in financial time series as TiRex does.

Bild

xLSTM for Lensed Gravitational Waves: arxiv.org/abs/2512.21370 sLSTM models fine-grained temporal structures, while mLSTM finds large-scale global patterns. xLSTM achieves AUC beyond 0.99, a TPR above 98% a FPR below 1% and is robust against noise, lens type, lens mass. Cool xLSTM application.

Bild

xLSTM for Real-Time DNS Tunnel Detection: arxiv.org/abs/2512.09565 DNS-HyXNet = xLSTM for DNS tunnels. DNS-HyXNet has 99.99% accuracy, with F1-scores exceeding 99.96%, and per-sample detection latency of just 0.041 ms, confirming its scalability and real-time readiness. wow!

Bild

xLSTM for PINNs that learn PDEs: arxiv.org/abs/2511.12512 “Across four PDEs under matched size and budget, xLSTM-PINN consistently reduces MSE, RMSE, MAE, and MaxAE with markedly narrower error bands.” “cleaner boundary transitions with attenuated high-frequency ripples”

BildBild

xLSTM for Vehicle Trajectory Prediction: arxiv.org/abs/2511.00266 X-TRACK based on xLSTM achieves SOTA. “Compared to state-of-the-art baselines, X-TRACK achieves performance improvement by 79% at the 1-second prediction and 20% at the 5-second prediction in the case of highD” Again xLSTM excels.

BildBild

xLSTM for robotic manipulation systems via diffusion-based imitation learning: arxiv.org/abs/2510.20406 PMP leverages xLSTM to denoise actions for robotics. “PMP not only achieves state-of-the-art performance but also offers significantly faster training and inference.” xLSTM excels in robotics.

Bild

xLSTM for Toxic Comment Classification: arxiv.org/abs/2510.17018 “On the Jigsaw Toxic Comment benchmark, xLSTM attains 96.0% accuracy and 0.88 macro-F1, outperforming BERT by 33% on threat and 28% on identity_hate categories, with 15× fewer parameters and <50 ms inference latency.” xLSTM is fast!

BildBild

gLSTM extends xLSTM to a graph neural network architecture: arxiv.org/abs/2510.08450 "gLSTM mitigates sensitivity over-squashing and capacity over-squashing." "gLSTM achieves comfortably state of the art results on the Diameter and Eccentricity Graph Property Prediction tasks"

Bild

xLSTM for long-term context using short sliding windows: arxiv.org/abs/2509.24552 "SWAX, a hybrid consisting of sliding-window attention and xLSTM." "SWAX trained with stochastic window sizes significantly outperforms regular window attention both on short and long-context problems."

Bild

xLSTM shines as an Electrocardiogram (ECG) foundation model: arxiv.org/abs/2509.10151 "xECG achieves superior performance over earlier approaches, defining a new baseline for future ECG foundation models." xLSTM is perfectly suited for time series prediction as shown by TiRex.

BildBild

xLSTM excels in time series forecasting: arxiv.org/abs/2509.01187 . Introduces "stochastic xLSTM" (StoxLSTM). "StoxLSTM consistently outperforms state-of-the-art baselines with better robustness and stronger generalization ability." We know that xLSTM is king at time series from our TiRex.

BildBildBild

xLSTM for Aspect-based Sentiment Analysis: arxiv.org/abs/2507.01213 Another success story of xLSTM. MEGA: xLSTM with Multihead Exponential Gated Fusion. Experiments on 3 benchmarks show that MEGA outperforms state-of-the-art baselines with superior accuracy and efficiency”

Bild