xLSTM for three-dimensional bin packing problem: arxiv.org/abs/2607.28257 xLSTM mixes item and candidate tokens with a recurrent action mixer (LRAM). xLSTM based OPAL has mean space utilization of 0.49 (15.1% improvement), while being fast. Promising for high-fidelity robotics.
@hochreitersepp.bsky.social
xLSTM for Brain MRI Analysis: arxiv.org/abs/2607.17782 Three-dimensional Bi-Directional xLSTM-UNet achieved second place overall and first in the meningioma segmentation task on FOMO 2025 leaderboard. xLSTM models long-range dependencies throughout the volumetric brain MRI.
1/5 🚀 Introducing TiRex-2 — our next-generation time series foundation model. Time series forecasting in the real world is streaming: • new observations continuously arrive, • variables interact, • some covariates are known into the future, • models must update predictions efficiently.
Comparison of sub-quadratic architectures xLSTM, Mamba-2, and Gated DeltaNet: arxiv.org/abs/2606.12364 Comparison of xLSTM, Mamba-2, and Gated DeltaNet on code pre-training, distillation, and time-series. xLSTM outperforms the others due to its gating scheme and state tracking.
5/ Takeaway LLMs do not always need to externalize their thoughts. They can learn to reason in working memory instead, decoupling intermediate computation from autoregressive generation 💡 Full paper: arxiv.org/abs/2605.30343 Huge thanks to @hochreitersepp.bsky.social for the guidance!
Unlocking the Working Memory of Large Language Models for Latent Reasoning
To improve the reasoning capabilities of large language models, test-time compute is typically scaled by generating intermediate tokens before the final answer. However, this couples reasoning to auto...
arxiv.org
RNNs like xLSTM with vertically chunked inference strategy for efficient memory usage: arxiv.org/abs/2604.18199 Chunking enables a linear-time and constant-memory like for TFLA for xLSTM arxiv.org/abs/2503.14376 Chunking blocks via recurrent updates speeds up computation considerably.
Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models
Transformer-based embedding models suffer from quadratic computational and linear memory complexity, limiting their utility for long sequences. We propose recurrent architectures as an efficient alter...
arxiv.org
A high-performance trading architecture built around xLSTM: allenarch.dev/blog/combini... The model is used as a market-state encoder, while reinforcement learning handles the trading decisions. This performance of xLSTM should be confirmed by further investigations.
xLSTM is more expressive than Transformer, Mamba: arxiv.org/abs/2603.03612 *nonlinear RNNs: sLSTM, LSTM *DLPR linear RNNs: mLSTM, RWKV-7, DeltaNet *Non PNC1-complete: Mamba, Transformer “fundamental expressivity gaps between linear and nonlinear RNNs.” World models require nonlinear RNNs.
xLSTM Distillation: arxiv.org/abs/2603.15590 Near-lossless distillation of quadratic Transformer LLMs into linear-time xLSTM architectures enables cost- and energy-efficient alternatives without sacrificing performance. Efficient xLSTM variants of instruction-tuned Llama, Qwen, and Olmo models.
Symbol-equivariant Recurrent Reasoning Models (SE-RRM) SE-RRM advances HRM and TRM -- guaranteed identical solutions for problems with permuted colors (ARC AGI) or digits (Sudoku). Coolest part: extrapolation to larger problem sizes!!! P: arxiv.org/abs/2603.02193 C: github.com/ml-jku/SE-RRM
xLSTM for Financial Time Series: arxiv.org/abs/2603.01820 "VLSTM achieved the highest overall Sharpe ratio" "VxLSTM and LPatchTST exhibited superior downside-adjusted characteristics" “xLSTM achieves the highest portfolio-level cost buffer" xLSTM excels in financial time series as TiRex does.
Drug design is significantly accelerated by ConGLUDe. Protein–small-molecule interactions can be screened orders of magnitude faster using the ConGLUDe approach.
# AI in Drug discovery just BROKE THROUGH a wall # A newer AI model, ConGLUDe, as fast but much more accurate than DrugCLIP. Instead on just 40K structure-based data, ConGLUDe is trained on 100M datapoints from ligand-based data P: arxiv.org/abs/2601.09693
xLSTM for Lensed Gravitational Waves: arxiv.org/abs/2512.21370 sLSTM models fine-grained temporal structures, while mLSTM finds large-scale global patterns. xLSTM achieves AUC beyond 0.99, a TPR above 98% a FPR below 1% and is robust against noise, lens type, lens mass. Cool xLSTM application.
xLSTM for Real-Time DNS Tunnel Detection: arxiv.org/abs/2512.09565 DNS-HyXNet = xLSTM for DNS tunnels. DNS-HyXNet has 99.99% accuracy, with F1-scores exceeding 99.96%, and per-sample detection latency of just 0.041 ms, confirming its scalability and real-time readiness. wow!
xLSTM for PINNs that learn PDEs: arxiv.org/abs/2511.12512 “Across four PDEs under matched size and budget, xLSTM-PINN consistently reduces MSE, RMSE, MAE, and MaxAE with markedly narrower error bands.” “cleaner boundary transitions with attenuated high-frequency ripples”
Measuring AI Progress in Drug Discovery - A NEW LEADERBOARD IN TOWN 2015-2025: turns out that there's hardly any improvement. AI bubble? GPT is at 70% for this task, whereas the best methods get close to 85%. Leaderboard: huggingface.co/spaces/ml-jk... P: arxiv.org/abs/2511.14744
xLSTM for Vehicle Trajectory Prediction: arxiv.org/abs/2511.00266 X-TRACK based on xLSTM achieves SOTA. “Compared to state-of-the-art baselines, X-TRACK achieves performance improvement by 79% at the 1-second prediction and 20% at the 5-second prediction in the case of highD” Again xLSTM excels.
xLSTM for robotic manipulation systems via diffusion-based imitation learning: arxiv.org/abs/2510.20406 PMP leverages xLSTM to denoise actions for robotics. “PMP not only achieves state-of-the-art performance but also offers significantly faster training and inference.” xLSTM excels in robotics.
Tenure Track in quantum informatics! Super cool position. Super cool team. World-class research. Scientifically outstanding work.
(I/III) We're excited to announce a new tenure track opening! The position is called 'quantum informatics' and is affiliated with our QUICK group within the CS+AI division at @jku.at 🇦🇹. Application deadline is November 30th, 2025: www.jku.at/en/the-jku/w...
xLSTM for Toxic Comment Classification: arxiv.org/abs/2510.17018 “On the Jigsaw Toxic Comment benchmark, xLSTM attains 96.0% accuracy and 0.88 macro-F1, outperforming BERT by 33% on threat and 28% on identity_hate categories, with 15× fewer parameters and <50 ms inference latency.” xLSTM is fast!
gLSTM extends xLSTM to a graph neural network architecture: arxiv.org/abs/2510.08450 "gLSTM mitigates sensitivity over-squashing and capacity over-squashing." "gLSTM achieves comfortably state of the art results on the Diameter and Eccentricity Graph Property Prediction tasks"
xLSTM for Intrusion Detection: arxiv.org/abs/2510.08333 "The xLSTM-based IDS achieves an F1-score of 98.9%, surpassing the transformer-based model at 94.3%." xLSTM is faster than transformer when using fast kernels as provided in github.com/nx-ai/mlstm_... and github.com/NX-AI/flashrnn
xLSTM for long-term context using short sliding windows: arxiv.org/abs/2509.24552 "SWAX, a hybrid consisting of sliding-window attention and xLSTM." "SWAX trained with stochastic window sizes significantly outperforms regular window attention both on short and long-context problems."
xLSTM shines as an Electrocardiogram (ECG) foundation model: arxiv.org/abs/2509.10151 "xECG achieves superior performance over earlier approaches, defining a new baseline for future ECG foundation models." xLSTM is perfectly suited for time series prediction as shown by TiRex.
xLSTM excels in time series forecasting: arxiv.org/abs/2509.01187 . Introduces "stochastic xLSTM" (StoxLSTM). "StoxLSTM consistently outperforms state-of-the-art baselines with better robustness and stronger generalization ability." We know that xLSTM is king at time series from our TiRex.
xLSTM for Cellular Traffic Forecasting: arxiv.org/abs/2507.19513 "Empirical results showed a 23% MAE reduction over the original STN and a 30% improvement on unseen data, highlighting strong generalization." xLSTM shines again in time series forecasting.
xLSTM for Monaural Speech Enhancement: arxiv.org/abs/2507.04368 xLSTM has superior performance vs. Mamba and Transformers but is slower than Mamba. New Triton kernels: xLSTM is faster than MAMBA at training and inference: arxiv.org/abs/2503.13427 and arxiv.org/abs/2503.14376
xLSTM for Aspect-based Sentiment Analysis: arxiv.org/abs/2507.01213 Another success story of xLSTM. MEGA: xLSTM with Multihead Exponential Gated Fusion. Experiments on 3 benchmarks show that MEGA outperforms state-of-the-art baselines with superior accuracy and efficiency”
xLSTM for multivariate time series anomaly detection: arxiv.org/abs/2506.22837 “In our results, xLSTM showcases state-of-the-art accuracy, outperforming 23 popular anomaly detection baselines.” Again, xLSTM excels in time series analysis.
xLSTM for Human Action Segmentation: arxiv.org/abs/2506.09650 "HopaDIFF, leveraging a novel cross-input gate attentional xLSTM to enhance holistic-partial long-range reasoning" "HopaDIFF achieves state-of-theart results on RHAS133 in diverse evaluation settings."