One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline Read more: https://arxiv.org/html/2610.06852v1
AI Research Updates | arXiv cs.AI
@arxiv-cs-ai.bsky.social
Your daily dose of the latest in Artificial Intelligence! Discover new research from arXiv's cs.AI section, covering machine learning, NLP, robotics, and more. 🚀 #ArtificialIntelligence #AIResearch #MachineLearning #NLP #Robotics #DeepLearning #AITech
Your Unlearning Gives You Away: Identifying Erased Concepts in Diffusion Models Read more: https://arxiv.org/html/2610.05601v1
LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets Read more: https://arxiv.org/html/2610.09872v1
SciExam for ENSO: Can AI Agents Build Climate Models? Read more: https://arxiv.org/html/2610.10513v1
SkillPoison: Progressive Skill Poisoning via Successful Experiences Read more: https://arxiv.org/html/2610.07645v1
Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness Read more: https://arxiv.org/html/2610.00568v1
IntentCoding: Amplifying User Intent in Code Generation Read more: https://arxiv.org/html/2602.00066v1
Towards Safer Autonomous Driving in an Open World: A Dual-Process Approach Read more: https://arxiv.org/html/2610.04088v1
Recursive Video In-Context Learning for Agentic Robot Read more: https://arxiv.org/html/2610.06843v1
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts Read more: https://arxiv.org/html/2610.06824v2
Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek Read more: https://arxiv.org/html/2610.08205v1
On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor Read more: https://arxiv.org/html/2610.02510v1
Copies or Sources? Measuring How LLM Aggregators Count Restated Evidence in Multi-Agent Systems Read more: https://arxiv.org/html/2610.06192v1
STRIDE: Spatial-Temporal Representation for Interval-conditioned Disease Evolution in Longitudinal Glioblastoma MRI Read more: https://arxiv.org/html/2610.08848v1
EviDent-CBCT: Evidence-Bottlenecked Report Generation from Dental CBCT under Non-Exhaustive Report Supervision Read more: https://arxiv.org/html/2610.02375v1
Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential Read more: https://arxiv.org/html/2610.05541v1
Efficient Neural Surrogates for Linear Radiation Transport on the Lattice and Hohlraum benchmarks Read more: https://arxiv.org/html/2610.04665v1
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows Read more: https://arxiv.org/html/2610.02122v1
Routing Probes Can Improve Without New Information: An Exact-Null Audit of Uncertainty Beyond Model Outputs Read more: https://arxiv.org/html/2609.38956v1
Transformer-Based Time-Series Inference of Lindblad Dynamics in Open Quantum Systems Read more: https://arxiv.org/html/2610.05647v1
Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy Read more: https://arxiv.org/html/2610.07403v1
Confidence-Ordering Reversal under Contextual Priors in Neural Decoding Read more: https://arxiv.org/html/2610.08229v1
Adaptive Operator Selection in Bilevel Large Neighborhood Search for Electric Autonomous Dial-a-Ride Problem under Uncertainty Read more: https://arxiv.org/html/2610.04219v1
EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning Read more: https://arxiv.org/html/2610.08726v1
MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis Read more: https://arxiv.org/html/2610.08528v1
A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care Read more: https://arxiv.org/html/2610.07339v1
Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks Read more: https://arxiv.org/html/2610.05750v1
GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting Read more: https://arxiv.org/html/2609.37694v1
Learning from Viable Failure Prefixes: Milestone Viability Potential Policy Optimization for Long-Horizon LLM Agents Read more: https://arxiv.org/html/2609.37111v2
Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation Read more: https://arxiv.org/html/2610.02678v1