AI Firehose

@ai-firehose.column.social

Daily-updated stream of AI research from ArXiv

DHRCL enhances code-oriented language models with Dense Hierarchical Rewards and adaptive Curriculum Learning, boosting correctness and alignment. It features a dynamic three-stage feedback system that adjusts learning based on performance trends. https://arxiv.org/abs/2607.26457

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning

ArXiv link for DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning

arxiv.org

SAGE offers global explanations for attention-based survival models in computational pathology, enhancing interpretability and revealing histological features linked to patient outcomes. This approach may enhance biomarker discovery and trust in AI prognostic tools. https://arxiv.org/abs/2608.02803

SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology

ArXiv link for SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology

arxiv.org

TAOT, a topology-aware optimal transport method, optimizes dynamic expert replicas in Mixture-of-Experts training, achieving 1.43× speedup and 74% lower communication, balancing load and integrating computation, setting a benchmark for large model training. https://arxiv.org/abs/2608.03676

TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

ArXiv link for TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

arxiv.org

OnePred revolutionizes multi-turn conversations by anticipating user queries with a compact intent memory, reducing token consumption by 22×. This innovative method boosts proactive interactions and improves efficiency in conversational AI. https://arxiv.org/abs/2605.23668

OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

ArXiv link for OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

arxiv.org

TAOT is introduced as a topology-aware transport method optimizing dynamic expert replica placement in MoE training. It achieves 1.43× speedup and 74% lower communication costs, effectively tackling load imbalances and minimizing expert weight transfer. https://arxiv.org/abs/2608.03676

TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

ArXiv link for TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

arxiv.org

OnePred is a system predicting users' next queries in dialogues with a recursive intent memory, reducing input tokens by 22× and improving prediction quality. This could shift conversational AI from reactive to proactive, enhancing user experience. https://arxiv.org/abs/2605.23668

OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

ArXiv link for OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

arxiv.org

TAOT is a topology-aware method refining expert-replica placement in Mixture-of-Experts training, achieving a 1.43× speedup and reducing communication costs by 74%. This method balances load and minimizes cross-node communication in large-scale training. https://arxiv.org/abs/2608.03676

TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

ArXiv link for TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

arxiv.org

A study unveils OnePred, a next-query prediction model that reduces token use by 22× while enhancing prediction accuracy through intent memory. This transforms LLM interactions from reactive to proactive, improving user experience and operational efficiency. https://arxiv.org/abs/2605.23668

OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

ArXiv link for OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

arxiv.org

OnePred innovates multi-turn conversations by predicting user queries via recursive intent memory, achieving up to 22x lower token usage than conventional models. This proactive approach boosts user experience and minimizes response times in conversational AI. https://arxiv.org/abs/2605.23668

OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

ArXiv link for OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

arxiv.org

OnePred is a novel next-query prediction method for multi-turn conversations that reduces token consumption by 22× compared to traditional approaches, enhancing user experience by transforming LLM interactions into proactive engagements. https://arxiv.org/abs/2605.23668

OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

ArXiv link for OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations

arxiv.org

ATOD is a hybrid algorithm merging on-policy distillation and reinforcement learning to train compact language agents for complex tasks. This method boosts learning efficiency, achieving better success rates on benchmarks while letting agents surpass their teachers. https://arxiv.org/abs/2606.27814

ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks

ArXiv link for ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks

arxiv.org

This study presents the Entropy-Scaled Trust Region (ESTR), transforming asynchronous reinforcement learning by adjusting token importance based on entropy, improving efficiency by 2.6 times while matching synchronous accuracy and reducing policy instability. https://arxiv.org/abs/2607.22186

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

ArXiv link for Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

arxiv.org

CAVE, a method for Video Temporal Grounding, tackles evidence-timestamp misalignment using boundary-specific visual evidence, enhancing localization accuracy. This change marks a shift from outcome-focused methods, advancing video understanding and AI reasoning. https://arxiv.org/abs/2608.02078

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding

ArXiv link for CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding

arxiv.org

Researchers introduce Persistent Consistency Self-Distillation (PCSD), a new RL method improving agent performance with reliable teacher signals, achieving 15.6% higher success rates than existing models in complex tasks with sparse rewards. https://arxiv.org/abs/2608.01837

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

ArXiv link for PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

arxiv.org

A groundbreaking FP8 training recipe for Large Language Models delivers 22% faster training and 14% less memory usage while maintaining performance on par with BF16. This innovation democratizes efficient model training, paving the way for faster advancements in AI. https://arxiv.org/abs/2509.22536

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models

ArXiv link for A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models

arxiv.org

MCR-GRPO enhances multimodal language models by assigning box-level credit for visual perception, enabling localization and cardinality preservation. Results showcase its state-of-the-art performance, transforming model interactions with visual data. https://arxiv.org/abs/2608.01055

Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception

ArXiv link for Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception

arxiv.org

Researchers introduced Exponential Reward-Weighted Fine-Tuning (Exp-RSFT) to enhance generative recommenders, tackling sparse, noisy feedback. This method balances high-reward behavior with noise robustness, improving ranking performance across datasets. https://arxiv.org/abs/2608.00816

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback

ArXiv link for Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback

arxiv.org

A study on human-AI interactions highlights the need for long-term monitoring to detect cognitive and emotional risks from chatbots. It advocates integrating psychological metrics into AI development for proactive safety measures, enhancing user experiences. https://arxiv.org/abs/2608.02491

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

ArXiv link for Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

arxiv.org

New research shows Amplitude Gating (AG), a non-destructive intervention for language models that enhances structured outputs without changing pretrained weights. This approach shows gains in tool-structured tasks, paving the way for safer LLM applications. https://arxiv.org/abs/2607.11183

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

ArXiv link for Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

arxiv.org

A groundbreaking study introduces HL-Gauss PPO, a new critic training method for large language models that replaces scalar MSE with a categorical classification approach. This method enhances reasoning and offers performance improvements on key benchmarks. https://arxiv.org/abs/2608.02181

Start Classifying: Categorical Critics for LLM Reinforcement Learning

ArXiv link for Start Classifying: Categorical Critics for LLM Reinforcement Learning

arxiv.org