Tuesday, September 15, 2026

763 papers cs.AI · cs.LG · cs.CL ← 2026-09-142026-09-16 →

Jul Aug Sep

Highlights

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Highlight HF pick · 8▲ Zhangxuan Gu, Haoxing Chen, Qi Qin, Yi Xin, Kai Gan, Lin Liu et al.

Diffusion language models generate many tokens in parallel, which suits GUI agents that must read a screen and output a precisely placed action at every step. Until now, though, no one had shown that such models can work as full multi-step vision-language agents. LLaDA-UI connects a native-resolution vision encoder to the block-wise diffusion backbone LLaDA2.0-mini-base, giving a 16.7B-parameter mixture-of-experts model, and trains it into an open-source agent for mobile, desktop and web.

  • Training has two stages: multimodal pre-training on just 145B tokens with a block-diffusion loss, which already beats earlier diffusion VLMs such as SDAR-VL (OCRBench 855 vs 726, MathVerse 48.0 vs 36.6), then full fine-tuning on more than 6M GUI grounding and navigation samples that output a reasoning block followed by an action.
  • Against the autoregressive Qwen3-VL-8B it wins on four of six GUI benchmarks, namely AndroidWorld (53.5 vs 47.9), MobileWorld (25.6 vs 9.4), WebVoyager (56.9 vs 45.2) and, by a hair, ScreenSpot-Pro (52.9 vs 52.7), but it trails on ScreenSpot-V2 (90.7 vs 93.0) and OSWorld-Verified (29.4 vs 33.9) and stays well behind specialized agents like Holo2-8B on WebVoyager (80.2) and GUI-Owl-7B on AndroidWorld (66.4).
  • In a paired test on four H100s, each API call runs 3.6–9.0× faster than Qwen3-VL-8B when neither model gets a cache hit (for example, 5.9 s vs 53.0 s on MobileWorld), though that is based on only five calls per domain, and Qwen's cached replays return in under 2 s.
  • Decoding settings strongly affect results: on an intermediate checkpoint, turning off early end-of-sequence stopping raises AndroidWorld success from 42.7% to 52.6%, while switching from 32-token blocks with 32 denoising steps to 64-token blocks with 16 steps drops it to 33.0%.
  • Outputs are almost always well-formed (98.65% parse), yet the agent gets stuck in loops on long tasks: repeated identical actions make up 26.7% of steps in failed runs, and success falls from 63% on tasks needing at most five actions to 14% on tasks needing more than 20; it also misses small targets on large OSWorld screenshots.

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Highlight HF pick · 92▲ Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo et al.

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Highlight HF pick · 153▲ Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman et al.

Omni-Streaming Thinking

Highlight HF pick · 12▲ Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang et al.

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Highlight HF pick · 9▲ Yunhao Feng, Ruixiao Lin, Ming Wen, Yanming Guo, Xingjun Ma, Yutao Wu et al.

Computer-use agents create safety risks through what they actually do with tools, yet most guard models are trained on static prompts and responses, and executable safety platforms return pass/fail verdicts rather than data a guard can learn from. HazardAuditor runs real agents in controlled environments to get labels based on their actual behavior. It then trains a generative guard with GuardPO, a post-training objective that weights each response by its safety verdict rather than by how long its explanation is.

  • Adapters convert traces from Claude Code, Codex, Hermes, and OpenClaw into one shared format (user messages, agent responses, tool calls with arguments, and what the environment returned), and a trace counts as unsafe only if the agent attempts a harmful action, not just because risky content showed up in its context.
  • GuardPO starts from a fine-tuned Qwen3Guard-Gen-8B and scores each sampled response by rule (+1 correct, −1 wrong, −1.25 malformed), then trains with clipped updates in the style of CISPO while averaging the explanation and verdict losses separately for each response, so long explanations no longer dominate the gradient.
  • On the authors' balanced CUA-Exec test set (200 trajectories per framework), the guard reaches 94.0 / 95.5 / 86.5 / 87.5% accuracy on the four agents, beating the strongest prior guard, BraveGuard, by 4.0–16.5 points and also outscoring Claude-Sonnet-4.6 used as a judge on every framework.
  • Results carry over to outside benchmarks, with 91.5% accuracy and F1 on ASSE-Safety, 88.4% accuracy on ATBench, and the best worst-case F1 across three benchmarks (88.3% vs. 82.5% for AgentDoG-Llama3.1-8B), but most of GuardPO's gain over plain fine-tuning comes on the authors' own traces (+10.4 points on CUA-Exec versus at most +2.2 elsewhere).
  • Limitations: the plain fine-tuned model beats the final model on R-Judge, BraveGuard keeps higher recall and F1 on AgentHazard traces from the GPT-5.5 agent, and the full guard takes about 3.1 s per trajectory, while lightweight classifier heads that are 10× faster drop to 69.6% Macro-F1.

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

Highlight HF pick · 5▲ Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar et al.

Measuring how LLM agents scale with test-time compute is hard because they decide for themselves how to spend tokens on revising, using tools, exploring and stopping. The proposed Elo-per-token analysis works on open-ended tasks that score intermediate submissions on a continuous scale: it tracks the best solution found at each token budget and uses a Bradley-Terry model to combine within-task rankings into Elo ratings that compare across tasks.

  • The study covers four general-purpose agents on four open-ended benchmarks with sessions of up to 100M tokens, plus three feedback-driven LLM optimization harnesses tested in controlled single-task experiments.
  • The baseline is independent sampling, whose Elo grows linearly with log compute in theory; agents beat this baseline early on, but their gains per token shrink until they fall below it.
  • On shared AtCoder Heuristic Contest tasks, the strongest past human contestants instead improve superlinearly over contest time, which the authors take as evidence of continual learning and of large headroom once agents slow down.
  • The authors call the per-session budget at which an agent's gains per token drop to the baseline's rate the scaling inflection point, and using it as the session size when splitting 100M tokens into parallel sessions on FrontierCS Polyomino Packing gains +264 Elo over one long session and +355 Elo over ten short sessions.
  • The method only works on tasks that score intermediate submissions continuously, and the budget-splitting gain is shown on a single task, so it is unclear how far it generalizes.

RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

Highlight HF pick · 12▲ Sibo Zhu, Shicheng Fan, Xinyue Wang, Wenyi Wu, Kun Zhou, Biwei Huang

Kaininja: Extending Native 3D Generators to the Part Level

Highlight HF pick · 12▲ Ruihan Yu, Lian Fu, Muyao Niu, Zheng-hui Huang, Yu-Ju Tsai, Sho Kuno et al.

Native image-to-3D generators such as TRELLIS.2 output one fused mesh. Their O-Voxel grid holds only one sheet of surface per voxel, so where two parts touch, the two surfaces merge into one at any resolution. KaiNinja splits an object's parts across two O-Voxel volumes so that touching parts never share a volume, and it extends TRELLIS.2 into a two-stream generator whose parts are simply the connected pieces of each volume, with no mask, bounding box or segmenter.

  • Parts are assigned to the two volumes by two-coloring a contact graph measured on the voxel grid, and the first stage, which lays out the parts, gets a separate copy of the pretrained transformer per volume, joined by new cross-volume attention that starts at zero, while the second, refinement stage reuses the released TRELLIS.2 weights and limits attention to one volume in 20 of its 30 blocks.
  • Training uses 19,132 objects from four sources, including Articraft-10K, whose assets an LLM coding agent built part by part so the part labels are exact, which the authors call the first use of agent-authored assets to train a 3D generative model.
  • On 986 held-out objects it is best on all seven metrics, with whole-object Chamfer distance of 0.0186 against 0.0314 for the strongest baseline, Hunyuan3D-2.1 + X-Part (40% lower), and strict part F-score of 0.692 against 0.597, at about 24 seconds per object on one H100, far below that segment-then-regenerate cascade's cost.
  • Whole objects also come out better than from TRELLIS.2@512 fine-tuned on the same data, with strict F-score of 0.919 against 0.830 and a failure rate of 0.7% against 1.8%, which the authors read as evidence that the two-volume packing represents whole objects better too.
  • On the external HY3D-Bench it still leads on whole-object geometry, but Hunyuan3D-2.1 + X-Part wins every part metric and KaiNinja ranks fourth of six on part mIoU, which the authors attribute to cutting objects into finer parts than the benchmark labels; separately, two-coloring over-merges objects whose parts interlock densely.

Atria Dawn: The Dawn of Agentic Superintelligence

Highlight HF pick · 135▲ Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu et al.

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

Highlight HF pick · 18▲ Ling Yang, Zhenfei Yin, Yingcheng Wu

Unclassified 703

Conflict-Predictive Variable Horizons in Multi-Drone Distributed Model Predictive Control

Linda M\"{u}mken, Michael Schwung, Stefan Lier, Andreas Schwung cross-listed No summary available — see the abstract on arXiv.

CANAL: Channel-Aware Noise Allocation for Differentially Private Feature Distillation in Medical Image Segmentation

Armaghan Butt, Shuya Feng, Qing Tian cross-listed No summary available — see the abstract on arXiv.

Survey of Novel Deep Learning Architectures for Denoising Gravitational-wave Signals

Rohan Raha, Prayush Kumar cross-listed No summary available — see the abstract on arXiv.

SomBench: Benchmark Dataset for Advancing Machine Learning in Lunar Science

Himanshu Patil, Gabby Nyirjesy, Rachel A. Slank, Vishal Gaur, Daniela Szwarcman, Paolo Fraccaro et al. cross-listed No summary available — see the abstract on arXiv.

Large-scale bioacoustic detection using semantic segmentation: a deep learning framework applied to fin whale calls in ocean-bottom seismometer recordings

Jocelyn Japnanto, Alex A. Saoulis, Miriam Romagosa, Rita Leit\~ao, Gabrielle Arrieta, M\'onica A. Silva et al. cross-listed No summary available — see the abstract on arXiv.

Multimodal-Multiresolution Foundation Model for Lunar Remote Sensing

Paolo Fraccaro, Gabby Nyirjesy, Daniela Szwarcman, Himanshu Patil, Vishal Gaur, Rohit Lal et al. cross-listed No summary available — see the abstract on arXiv.

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

Vishesh Tripathi, Abhay Kumar, Ramsha Khan cross-listed No summary available — see the abstract on arXiv.

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Zhangxuan Gu, Haoxing Chen, Qi Qin, Yi Xin, Kai Gan, Lin Liu et al. cross-listed No summary available — see the abstract on arXiv.

ScorePrompts: Natural-Language Exploration of Symbolic Music Scores through Analysis

Emmanouil Karystinaios, Gerhard Widmer cross-listed No summary available — see the abstract on arXiv.

Aries: A Proprietary Medium-Range Weather Prediction Model for the Energy Industry

Lukas Hedegaard Morsing, Arian Bakhtiarnia, Jonas Lynge Olesen, T\'omas Bragi Bj\"ornsson Leth, Christian G{\o}bel Bach cross-listed No summary available — see the abstract on arXiv.

Ergodic Control and Controlled Diffusion for Robot Learning: Review and Tutorial

Max Muchen Sun, Cem Bilaloglu, Ananya Rao, Stefan Ivic, Guillaume Sartoretti, Kathleen Fitzsimons et al. cross-listed No summary available — see the abstract on arXiv.

Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning

Xinrui He, Ting-Wei Li, Junting Wang, Mengting Ai, Xinyu He, Hanghang Tong et al. cross-listed No summary available — see the abstract on arXiv.

Variational Template Matching with Statistical Fusion for Anomaly Detection in Patterned Structures

Qinwu Xu, Yifan Jiang cross-listed No summary available — see the abstract on arXiv.

Adaptive Conformal Redistribution for Inter-class Transitional Uncertainty in Medical Image Classification

Saibal Ghosh, Samarup Bhattacharya, Sanjoy Kumar Saha, Umapada Pal, Tapabrata Chakraborti cross-listed No summary available — see the abstract on arXiv.

IMM-based Multiple Object Tracking using a State Prediction Neural Network

Chan-Bin Lim, Dong-Hee Paek, Seung-Hyun Kong cross-listed No summary available — see the abstract on arXiv.

GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures

Sarthak Sattigeri cross-listed No summary available — see the abstract on arXiv.

Task-Based CT Protocol Optimization Using Reinforcement Learning and Virtual Imaging Trials

Jiaqi Zou, David Fenwick, Vahid Tarokh, Nicholas Felice, Jayasai Rajagopal, Anuj Kapadia et al. cross-listed No summary available — see the abstract on arXiv.

Feasibility and Memory Mechanisms of Chern-Simons Context Reservoir Computation

Jyotiranjan Beuria, Venkatesh H. Chembrolu cross-listed No summary available — see the abstract on arXiv.

Decoupling Error Attribution in Cloud-Native Graph-RAG: A Data Integrity Diagnostic Framework

Shuai Yan, Yuhang Wu, Xiaodong Huang, Ke Wang cross-listed No summary available — see the abstract on arXiv.

Pedestrian Crossing Intent Classification From Event-Based Vision Using Convolutional Spiking Neural Networks With Temporal Augmentation

Henok Teklu, Mustafa Sakhai, Maciej Wielgosz, Matej Mertik cross-listed No summary available — see the abstract on arXiv.

The Agentic Company OS: Substrate Inversion for Sustained Enterprise Agent Deployment

Oliver Aleksander Larsen, Mahyar T. Moghaddam cross-listed No summary available — see the abstract on arXiv.

Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework

Kazi Abrar Mahmud, Nilotpaul Kundu Dhurubo, Tamal Kirttonia, Sabbir Hossain Ujjal, Mohammad Ariful Haque cross-listed No summary available — see the abstract on arXiv.

ProtoCAM: Interpretable Few-Shot Mask-Guided Prototypical Learning for Breast Lesion Classification in Ultrasound Imaging

Ashkan Ebadi cross-listed No summary available — see the abstract on arXiv.

Stochastic Gradient Descent over P2

Maria Oprea, Qin Li, Yunan Yang cross-listed No summary available — see the abstract on arXiv.

Beyond Point Forecasts: A Survey on Probabilistic Forecasting for Time Series and Spatiotemporal Data

Donia Besher, Rajdeep Pathak, Madhurima Panja, Tanujit Chakraborty cross-listed No summary available — see the abstract on arXiv.

ViFA-Council: Multi-Agent LLM Deliberation for Vietnamese Folk Art Generation

Hai-Dang Nguyen, Minh-Phuong Pham, Thao Thi Phuong Dao, Trong-Le Do, Vinh-Tiep Nguyen, Trung-Nghia Le cross-listed No summary available — see the abstract on arXiv.

Real-time Learning and Evolution in Robotic Art Installations

Sofian Audry, Stephen Kelly cross-listed No summary available — see the abstract on arXiv.

SkillAtlas: An Attack Trace Library for Agent Skills

Yuxin Tian, Zenghao Duan, Liang Pang, Zhiyi Yin, Xueqi Cheng cross-listed No summary available — see the abstract on arXiv.

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo et al. No summary available — see the abstract on arXiv.

RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines

Anqi Chen, Dan Goldwasser, Cristina Nita-Rotaru No summary available — see the abstract on arXiv.

Privacy-Preserving Deep Joint Source-Channel Coding with In-Loop Concept Erasure

Rami Eid, Maria Slim, Mariette Awad, Hadi Sarieddeen cross-listed No summary available — see the abstract on arXiv.

Task-Aware Federated Fine-Tuning for MoE-based Large Language Models

Tingqi Wang, Hongyu Ke, Haoxin Wang, Rafal Angryk, Zhipeng Cai No summary available — see the abstract on arXiv.

Converge Then Diversify: Decoupling Convergence and Diversity in Multi-Objective Bayesian Optimisation

Chao Jiang, Yueling Huang, Miqing Li No summary available — see the abstract on arXiv.

Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement

Hongyao Tang, Yi Ma, Pengyi Li, Yifu Yuan No summary available — see the abstract on arXiv.

CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages

Lucas Rafael Stefanel Gris, Alef Iury Siqueira Ferreira, Frederico Santos de Oliveira, Augusto Seben da Rosa, Alexandre Costa Ferro Filho, Arlindo Rodrigues Galv\~ao Filho et al. No summary available — see the abstract on arXiv.

Specification Oracles

Atticus Cull, Justin McCarthy No summary available — see the abstract on arXiv.

Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents

Toshiaki Koike-Akino, Vlad Blaykhman, Ye Wang, Jing Liu, Gene V. Vinokur No summary available — see the abstract on arXiv.

ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement

Yihang Chen, Yuanhao Ban, Kuei-Chun Kao, Cho-Jui Hsieh No summary available — see the abstract on arXiv.

Chance-Constrained Belief-Space Maneuver Planning for Autonomous Collision Avoidance Under Uncertainty

Grace Ra Kim, Duncan Eddy, Mykel J. Kochenderfer cross-listed No summary available — see the abstract on arXiv.

Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

Varun Kaushik, Yayun Tan, Xiaofan Yu No summary available — see the abstract on arXiv.

LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents

Lei Liu, Yikun Zhang, Jialin Chen, Wanjia Zhao, Rex Ying, Wengong Jin et al. No summary available — see the abstract on arXiv.

Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization

Hrad Ghoukasian, Anastasis Kratsios No summary available — see the abstract on arXiv.

Efficient Online Inverse Optimization with $O(d)$ Regret

Yang Cai, Anupam Gupta, Vineet Gupta, Guru Guruganesh, Yanchen Jiang, Christopher Liaw et al. No summary available — see the abstract on arXiv.

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville No summary available — see the abstract on arXiv.

Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs

Kento Nishi No summary available — see the abstract on arXiv.

A Machine Learning API for Earth Observation Data Cubes Based on openEO

Brian Pondi, Jonas Hurst, Rolf Simoes, Jonas Starke, Marius Appel, Edzer Pebesma No summary available — see the abstract on arXiv.

Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

Misaki Matsuura, Sayantan Kumar, Ojas Kadam, Jeremy C. Weiss No summary available — see the abstract on arXiv.

TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

Sudarshan Regmi, Arvind Pillai, Yu Yvonne Wu, Yuliang Chen, Bibek Panthi, Tess Z. Griffin et al. No summary available — see the abstract on arXiv.

STAGE: Diagnosing Semantic Transfer at Grounded Execution in Embodied Agents

Baosheng Jin, Yushen Liang, Hua Shen cross-listed No summary available — see the abstract on arXiv.

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang et al. No summary available — see the abstract on arXiv.

Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement

Sandeep Bokkasam, B. Durgalakshmi No summary available — see the abstract on arXiv.

Symmetry- and Property-Aware Crystal Generation with Reinforcement Learning for Inverse Materials Design

Ting-Wei Hsu, Arun Bansil, Qimin Yan cross-listed No summary available — see the abstract on arXiv.

Inverse Learning of the Altruism and Cost Level in Mixed-Individual Mean Field Games

Haoyang Cao, G\"ok\c{c}e Dayan{\i}kl{\i}, Xiaofei Shi cross-listed No summary available — see the abstract on arXiv.

OrchSLM: Probing the Dynamics of Small Language Model Orchestration

Chengxi Zhang, Yu Yao No summary available — see the abstract on arXiv.

Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports

Jack Cummins, Sayantan Kumar, Ketan Tamirisa, Jeremy C. Weiss No summary available — see the abstract on arXiv.

Edge-addition monotonicity of positive p-energy fails for every p >= 1

Koyar Afrasyab cross-listed No summary available — see the abstract on arXiv.

On the Potential of Multi-Task Learning in Predictive Process Monitoring

Lukas Kirchdorfer, Keyvan Amiri Elyasi, Heiner Stuckenschmidt No summary available — see the abstract on arXiv.

A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text

Saad Bin Ather, Muhammad Saif, Ali Hassan Khan, Manzer Abbas, Hajra Waheed No summary available — see the abstract on arXiv.

Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers

Anubhav Shrestha, Safal Shrestha, Minwu Kim, Torsten Suel, Keith Ross cross-listed No summary available — see the abstract on arXiv.

Token Efficient Task Execution via Application Behavior Modeling for Web Agents

Alexandru Ianta, Eleni Stroulia No summary available — see the abstract on arXiv.

Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study

Yao Xiao, Max Chen, Yichen Li, Nathaniel Powers, Maxwell Wiesenfeld, Gillian Smith et al. cross-listed No summary available — see the abstract on arXiv.

Building a Production Greek-English Speech Recognizer

Christos Petrocheilos, Cleopatra Papadopoulou, Chris Porikis, Ioakeim Perros, Ayoub Kirouane, Themistoklis Nikolis cross-listed No summary available — see the abstract on arXiv.

Canaries in the Bank: Auditing User-Level Privacy in Private Evolution

Sai Aparna Aketi, Enayat Ullah, Shripad Gade cross-listed No summary available — see the abstract on arXiv.

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction

Chiwun Yang, Xiaoyu Li No summary available — see the abstract on arXiv.

Pretraining for Sample-Efficient Neural Interfaces

Ben Tang, Zachary Spalding, Gregory B. Cogan No summary available — see the abstract on arXiv.

A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion

Muhammad Ebad Atif, Muhammad Haider Ali No summary available — see the abstract on arXiv.

Adaptive Phase-Switching for Communication-Efficient Federated LoRA Fine-Tuning

Jerry Adams Franklin No summary available — see the abstract on arXiv.

Operational Range Bounding in Spectroscopy: A Safety Cage Framework for Machine Learning Models

Nikki Grens, Lu\'is F. Sim\~oes, Kai Hou Yip, Theresa Lueftinger No summary available — see the abstract on arXiv.

GeoTTER: Leveraging Local Geometry of Optimal Transport for Zero-Shot Classification

Wei-Yang Alex Lee, Rudrasis Chakraborty, Vishnu Lokhande No summary available — see the abstract on arXiv.

Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation

Thao Nguyen, Jeonghwan Kim, Zhenhailong Wang, Heng Ji No summary available — see the abstract on arXiv.

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

Kyle Richardson, Cullen Anderson, Pranav Balakrishnan, Takuto Ban, Daksha Ladia, Ankita Gupta et al. No summary available — see the abstract on arXiv.

Positive Topology and Feasible Refinement: Forcing Matrices, Positivity, and Information

Mirco A. Mannucci, Giovanni Sambin cross-listed No summary available — see the abstract on arXiv.

Rolling Day-Wise Mortality Prediction in Critically Ill Patients With AKI on CRRT Utilizing Machine Pressure Waveforms

Shehan Irteza Pranto, Joanna Yang, Joshua Lambert, Stuart L. Goldstein, Lili Chan, Girish N. Nadkarni et al. No summary available — see the abstract on arXiv.

A family of spectral conjugate gradient algorithms derived by least-squares approximations based on a modified quasi--Newton update with application to a revised robust binary classification model

Saman Babaie-Kafaki, Maryam Khoshsimaye-Bargard, Ahmad Mousavi cross-listed No summary available — see the abstract on arXiv.

Generative Interpretability via Scalable Neuro-Symbolic Models

Xiaocong Yang No summary available — see the abstract on arXiv.

Attention Is All You Need (to Avoid Spurious Oscillations)

Jinyoung Jeong, Joseph B. Choi, Xinlun Cheng, H. S. Udaykumar, Sanghun Choi, Stephen S. Baek No summary available — see the abstract on arXiv.

Early-Stopping Thresholds for ES-HyperNEAT: A Data-Driven Approach from Fitness Dynamics

Romain Claret, Arthur Gygax, Michael O'Neill, Paul Cotofrei, Pascal Felber cross-listed No summary available — see the abstract on arXiv.

Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models

Noor Islam S. Mohammad, Ulu\u{g} Bayaz{\i}t No summary available — see the abstract on arXiv.

From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements

Ronald Schnitzer, Mike Auer, Rumpa Choudhury, Andreas Hapfelmeier, Maximilian Hoeving, Isabelle Painter et al. No summary available — see the abstract on arXiv.

Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents

Grace Chang Yuan, Xiaoman Zhang, Sung Eun Kim, Luyang Luo, Pranav Rajpurkar No summary available — see the abstract on arXiv.

Runtime-Incremental Transformer for Reinforcement-Learning-Based Adaptive Control

Giansalvo Cirrincione, Adriano Fagiolini cross-listed No summary available — see the abstract on arXiv.

Toward Optimal Switching Regret for Multi-Armed Bandits with Oblivious Adversary

Mengxiao Zhang No summary available — see the abstract on arXiv.

AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents

Xinyun Cao, Adriana Szekeres, Fazle Elahi Faisal No summary available — see the abstract on arXiv.

Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management

Alexandre Barreto (George Mason University), Shou Matsumoto (George Mason University), Jorge Valverde-Rebaza (Tecnol\'ogico de Monterrey), Cleiton Ataide (DECEA: Department of Airspace Control), Paulo Costa (George Mason University) No summary available — see the abstract on arXiv.

Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

Darin Keng, Zhewei Sun No summary available — see the abstract on arXiv.

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos No summary available — see the abstract on arXiv.

A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics

Xian Yeow Lee, Teppei Inoue, Haiyan Wang, Chetan Gupta No summary available — see the abstract on arXiv.

When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence

Zichen Wang, Haoyang Hong, Huazheng Wang No summary available — see the abstract on arXiv.

Planning or Learning: Reliability and Cost in Multi-Asset Maintenance

Xian Yeow Lee, Chandrasekar Venkatraman, Ahmed Farahat No summary available — see the abstract on arXiv.

Causal multi-modal AI for personalized chemosensitivity prediction

Dhruva Biswas, Jeroen Berrevoets, Alec McClean, Linus Bao, Jungkyu Park, Ken G. Zeng et al. No summary available — see the abstract on arXiv.

A pullback-corrected scalar auxiliary variable optimizer with momentum and adaptive mobility

Jiahao Zhang, Shiheng Zhang, Guang Lin cross-listed No summary available — see the abstract on arXiv.

How User-AI Mistreatment Occurs and Matters in Conversational Systems?

Fanqi Zeng, Sadid A. Hasan, Chaocheng He No summary available — see the abstract on arXiv.

FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks

Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi No summary available — see the abstract on arXiv.

Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation

Paul Landes, Sitara Rao, Aaron Jeremy Chaise, Barbara Di Eugenio No summary available — see the abstract on arXiv.

Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs

Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han, Wenbin Zhang No summary available — see the abstract on arXiv.

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou et al. cross-listed No summary available — see the abstract on arXiv.

BOOST: Concurrent Access to Host Memory and HBM to Accelerate LLM Inference

Anish Saxena, Jae Hyung Ju, Hritvik Taneja, Po-An Tsai, Aamer Jaleel, Christos Kozyrakis et al. cross-listed No summary available — see the abstract on arXiv.

Predictive audio representations for early detection and tracking of hidden dynamic objects

Katerina Vinciguerra, Moritz Brandes, Danilo Hollosi, Letizia Marchegiani cross-listed No summary available — see the abstract on arXiv.

$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents

Soham Ray, Victor Barres No summary available — see the abstract on arXiv.

In the Blind: Building Pseudo-References for MT Evaluation

Diptesh Kanojia, Chi-kiu Lo, Archchana Sindhujan, Samuel Larkin, Greg Hanneman, Alon Lavie No summary available — see the abstract on arXiv.

AttnFuse: A Composable DSL for Compiling Attentions to Fused GPU Kernels

Varun Kumar Dasoju, Tian Zhao No summary available — see the abstract on arXiv.

The University of Melbourne WMT 2026 CreoleMT Submission: A Domain-Balanced Approach to Low-Resource Pacific Creole Machine Translation

Rapha\"el Merx, Nick Thieberger, Ekaterina Vylomova No summary available — see the abstract on arXiv.

An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

Roberto Campbell, Momin Abbass, Muneeza Azmat, Michal Ulewicz, Raya Horesh, Kristjan Greenewald et al. No summary available — see the abstract on arXiv.

FaithfulBench: Does AI Counsel Uphold or Undermine the User's Professed Faith?

M Waleed Kadous, Benjamin Olsen, Walter Scheirer, Daniel D. Slate, Alexander Arnold, DZ Kalman cross-listed No summary available — see the abstract on arXiv.

EI-DDLGN: Efficient Encrypted Inference with Deep Differentiable Logic Gate Networks under TFHE

Mahmoud Y. M. Yassin, Mahmoud AbdelHafeez Sayed, Mostafa Taha cross-listed No summary available — see the abstract on arXiv.

Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents

Zhenyu Zhao, Roy Zhao No summary available — see the abstract on arXiv.

FlowTSFM: Turning Encoder Depth into Quantile Transport

Bahaeddine Abdessalem, Shifeng Xie, Zehao Xiao, Youssef Attia El Hili, Ambroise Odonnat, Jianfeng Zhang et al. No summary available — see the abstract on arXiv.

When Compliance Data Masquerades as Evaluation: Measurement Validity for Deployed AI Systems

Hung-Yu Lin, Xingran Huang, Qiming Guo, Jinwen Tang No summary available — see the abstract on arXiv.

Curvature-Independent Regret Bounds for Distributed Online Optimization on Hadamard Manifolds

Zhanyuan Cai, Emre Sahinoglu, Shahin Shahrampour No summary available — see the abstract on arXiv.

Solar Intelligence

Jyotsna Singh No summary available — see the abstract on arXiv.

Online Bayesian Node Classification on Inductive Graphs under Distribution Shift

Jinwen Xu, Gonzalo Mateos Buckstein, Qin Lu No summary available — see the abstract on arXiv.

Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself

Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan No summary available — see the abstract on arXiv.

FedV-KGQA in Practice: Design Lessons and an Interactive Prototype

Md Saikat Islam Khan Bappy, Oshani Seneviratne No summary available — see the abstract on arXiv.

Cost Characterization of Vertically Partitioned Federated Knowledge Graphs

Md Saikat Islam Khan Bappy, Oshani Seneviratne No summary available — see the abstract on arXiv.

GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents

Han Luo, Xian Xu, Yinhe Liu, Yanfei Zhong No summary available — see the abstract on arXiv.

Enhancing Event Candidate Acquisition for Event Linking

Ziyang Zhang, Yinan Liu, Boyi Xue, Yingxuan Huang, Bin Wang, Xiaochun Yang No summary available — see the abstract on arXiv.

Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations

Ashfaq Ali Shafin, Khandaker Mamun Ahmed cross-listed No summary available — see the abstract on arXiv.

Recoverability as a System Primitive for Long-Horizon AI Agents

Zhihui Zhang, Wei Liu No summary available — see the abstract on arXiv.

Windowed A-K-MDP

Xiangwen Yang, Frankie Cho, Iadine Chades No summary available — see the abstract on arXiv.

Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models

Fei Yuan, Changjiang Gao, Yilei Tu, Yifeng Liu, Shujian Huang, Yu Qiao No summary available — see the abstract on arXiv.

LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference

Prateek Kumar Sikdar No summary available — see the abstract on arXiv.

Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding

Tian Tan, Eduardo Blanco No summary available — see the abstract on arXiv.

Prefix Sharing Is a Sorting Problem

Rong He cross-listed No summary available — see the abstract on arXiv.

Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction

Ashfaq Ali Shafin, Khandaker Mamun Ahmed No summary available — see the abstract on arXiv.

Gap Entropy and Almost Instance-Wise Optimal Best-Arm Identification

Jiarui Yao, Jiaxi Zhao, Xiangxin Zhou cross-listed No summary available — see the abstract on arXiv.

When Edit Localization Amplifies Relative Selection Bias: Gradient Geometry, Target Mismatch, and Importance Weighting

Shengwei Zhang, Haoda Dai, Yifei Li, Yuheng Song No summary available — see the abstract on arXiv.

Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks

Jinhua Wu, Xinliang Zhang No summary available — see the abstract on arXiv.

Certifying Model Upgrades with Slice-Wise Non-Regression and Incumbent Fallback

Shengwei Zhang, Tao Wu, Fei Qian No summary available — see the abstract on arXiv.

JaxAHT: A JAX-Based Library for Ad Hoc Teamwork

Caroline Wang, Rolando Fernandez, Zelal Su Mustafaoglu, Montek Kundan, Jiaxun Cui, Lingyun Xiao et al. No summary available — see the abstract on arXiv.

MANAS-2: Constrained Reconstruction for EEG Foundation Models

Arvasu Kulkarni, Aditya Ray Mishra, Mahir Jain, Parshva Runwal, Lakshya Saini, Siddharth Panwar et al. No summary available — see the abstract on arXiv.

Oops, Not Now: PEARL, a RAG-Based Support Agent for Gameplay and What Players Want from AI Help

Jiahong Li, Sai Siddartha Maram, Atieh Kashani, Ulia Zaman, Zhiyu Lin, Cameron Marano et al. cross-listed No summary available — see the abstract on arXiv.

Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging

Gautami Sanjay Naik, Krishna Bhatia, Mithun Paul Saint-Germain, H Aswath Babu No summary available — see the abstract on arXiv.

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

Kainan Zhou, Gangzhen Qian, Zhaoyi Li, Hang Xiao No summary available — see the abstract on arXiv.

TyPatch: Transforming Patches into Typestate Rules for Kernel Bug Detection

Ruoyu Wang, Tuo Li, Jia Li cross-listed No summary available — see the abstract on arXiv.

A Variational Optimal Transport Operator on Incompressible Flow

Jinjin He, Shenyifan Lu, Sinan Wang, Zhiqi Li, Duowen Chen, Bo Zhu No summary available — see the abstract on arXiv.

JumpStart Your Policy Learning with Lessons from 160,000 Training Runs

Nabil Omi, Eric Bae, Chung Yik Edward Yeung, Siddhartha Sen, Ali Farhadi No summary available — see the abstract on arXiv.

Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges

Seyedakbar Mostafavi No summary available — see the abstract on arXiv.

PolicyMem: Geometric Policy Memory for LLM Governance

Yuanchen Bei, Zhengzhang Chen, Yanjun Zhao, Haoyu Wang, Hanghang Tong, Haifeng Chen No summary available — see the abstract on arXiv.

ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation

Hanling Wang, Chenlong Wei, Ling Xu, Hanyan Niu, Qi Cao, Shizhou Huang et al. No summary available — see the abstract on arXiv.

HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning

Hongliang Wei (Harbin Institute of Technology, Alibaba Cloud), Xiaobing Tu (Alibaba Cloud), Yinggui Wang (Alibaba Cloud), Zhengxi Liu (Alibaba Cloud), Rongkun Xue (Alibaba Cloud) et al. No summary available — see the abstract on arXiv.

Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth

Tianhao Niu, Qingfu Zhu, Wanxiang Che No summary available — see the abstract on arXiv.

How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition

Hongyu Gu, Chang Liu, Jingwen Fu No summary available — see the abstract on arXiv.

On the Equivalence of Stochastic Control and Path Space Formulations for Schr\"odinger Bridges over Compact Connected Lie Groups

Hamza Mahmood, Georgiy A. Bondar, Abhishek Halder, Adeel Akhtar cross-listed No summary available — see the abstract on arXiv.

Positioning manuscripts in the scientific landscape with agentic AI

Jiawen Chen, Zichen Zhang, Bingxuan Li, Quan Sun, Yiyan Zhang, Edric Tam et al. No summary available — see the abstract on arXiv.

HyperProve: Answer-Guided Hypergraph Expansion for Multi-Hop Question Answering

An Nguyen Phu, Dung Nguyen Quang, Luu Hieu An, Linh Ngo Van, Trung Le, Thien Huu Nguyen No summary available — see the abstract on arXiv.

Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation

Yilei Tu, Zihao Li, Shaoxiong Ji, J\"org Tiedemann, Fei Yuan No summary available — see the abstract on arXiv.

Homeostatic Continual Learning

Yue Jin No summary available — see the abstract on arXiv.

Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

Ruichen Zheng, Yihe Wang, Fabrice Y Harel-Canada, Sara Khosravi, Zeynep Senahan Yildiz, Amit Sahai et al. No summary available — see the abstract on arXiv.

Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs

Siddhesh Thombre, Manasi Patwardhan, Sunita Sarawagi No summary available — see the abstract on arXiv.

GEAR: From Dynamic Encoding to Dynamic Activation in Social Trajectory Prediction

Jiaheng Chen, Jiaxing Li, Leixia Wang, Jianan Ju, Tinghe Zhang cross-listed No summary available — see the abstract on arXiv.

Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection

Jiachen Zhang, Yu Tang, Li Zhu No summary available — see the abstract on arXiv.

PPDL: A Real-world Industrial User Retention Ratio Forecasting Framework Integrating Physical Priors with Deep Learning

Zibo Zhao, Zhengxiong Guan, Chaoli Zhang, Linyuan Geng, Xuanbing Zhu, Zhonglong Zheng et al. No summary available — see the abstract on arXiv.

SyRHM: Symbolic-Language-Enhanced Reasoning with Associative Retrieval for Zero-shot Harmful Meme Detection

Hanling Wang, Chenlong Wei, Yingjuan Li, Di Wu, Yuchao Zhang, Xiaohui Zhu et al. No summary available — see the abstract on arXiv.

Resolution-Independent Analysis of Encoder--Decoder Operator Learning via Limiting Kernels

Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou cross-listed No summary available — see the abstract on arXiv.

Do Not Restart: Residual Completion for Stateful Agent Handoffs

Runzhi Deng, Yiming Zhong, Fang Zhao, Pan Zhou No summary available — see the abstract on arXiv.

LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems

Yang Zhang, Lindong Xie, Chongyu Wang, Gaojunjie Li, Siqi Bu, Edward Chung No summary available — see the abstract on arXiv.

Understanding the Limits of Agentic ICD Coding

Chong Yock Eng, Yushi Cao, Yiming Chen, Kezhi Mao, Hongchao Jiang No summary available — see the abstract on arXiv.

Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture

Haibin Tong, Jiang Yu No summary available — see the abstract on arXiv.

DARE: Dialectical Agentic Reasoning for Structured Knowledge Fact Checking

Yifei Li, Xiaohan Zheng, Wentao Qian, Liansheng Zhuang No summary available — see the abstract on arXiv.

Odds-Shift Slippage in One-vs-Rest Rankers: Diagnosing and Repairing Reweighting-Induced Top-K Errors

Akifumi Goto cross-listed No summary available — see the abstract on arXiv.

Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models

Manit Kaushik, Ishir Bhardwaj, Pranav Gupta, Pankaj Jalote, Arun Balaji Buduru cross-listed No summary available — see the abstract on arXiv.

Benchmarking Optimizers to Solve Inverse Problems with Differentiable Physics Simulators

Xiang Chen, Huanhuan Xia No summary available — see the abstract on arXiv.

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

Aakash Kumar Tiwari No summary available — see the abstract on arXiv.

ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading

Asser Moustafa, Rares-Mihail Neagu, Jugal Kalita No summary available — see the abstract on arXiv.

Graph Neural Networks for Influence Maximization in Social Networks: An Unsupervised Minimum Dominating Set Approach

Erfan Ahmadi, Mina Shirazi, Behnam Bahrak No summary available — see the abstract on arXiv.

Accuracy Is Not Service: A Decision-Aware Benchmark for Intermittent-Demand Forecasting

Joo Ern Chin, Shih-Fen Cheng, Aldy Gunawan No summary available — see the abstract on arXiv.

Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice

Helena Choi, Edric Castel Hao, Karl Bautista, Francis Gabriel Magleo, Renzo Panti, Danielle Beatrice Olalia No summary available — see the abstract on arXiv.

CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection

Minh-Xuan Phan, Khalid Zaman, Candy Olivia Mawalim, Masashi Unoki cross-listed No summary available — see the abstract on arXiv.

Pre-training with Graph Transformers

Jiaming Wang, Thomas Laurent, Xavier Bresson No summary available — see the abstract on arXiv.

LePlanner: An Iterative Amortized Controller For World Models

Saksham Bansal, Om Naphade, Chayan Aggarwal, Vrishin M cross-listed No summary available — see the abstract on arXiv.

Affinity-Aware Sharding for Delayed Tensor Parallelism

Eloi de Reynal No summary available — see the abstract on arXiv.

Measuring the Cost of Variety Conflation in Multilingual MT Evaluation: Adding Mozambican Xichangana, Nyanja and Sena to FLORES+

Felermino D. M. A. Ali, Delfina L\'azaro Mateus, Manuel Valente Mangue No summary available — see the abstract on arXiv.

UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics

Yuzhe Li, Hao Yan, Hao Wang, Xingchen Liu, Ya-Qi Yu, Jihao Wu et al. No summary available — see the abstract on arXiv.

ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting

Chenwei Wang, Dianye Huang, Match W. L. Ko, Chenjia Bai, Zhongliang Jiang cross-listed No summary available — see the abstract on arXiv.

ShopEase: A Generative AI-Based Multi-Agent Framework for Intelligent Enterprise Customer Support Using Hybrid Retrieval-Augmented Generation

Aakash Kumar Tiwari, Somesh Kumar No summary available — see the abstract on arXiv.

ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation

Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen No summary available — see the abstract on arXiv.

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano et al. No summary available — see the abstract on arXiv.

An Uncertainty-Aware Hybrid Mathematical-Machine-Learning Model for Smart Irrigation Decision Support

Andrea Scariolo No summary available — see the abstract on arXiv.

The Filter Metric is Safety-Critical: Phantom Advantages in Group-Relative RL under Shaped Rewards

Juntao Yu No summary available — see the abstract on arXiv.

From Network Inequality to Network Fairness: A Perspective on Responsible Decision-Making

Lisette Esp\'in-Noboa, Tina Eliassi-Rad, Pak-Hang Wong, Erich Prem, Meike Zehlike, Ricardo Baeza-Yates et al. cross-listed No summary available — see the abstract on arXiv.

Physically Typed and Geometry-Aware Representations for Earth Foundation Models

Rajiv Ranjan cross-listed No summary available — see the abstract on arXiv.

Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability

Swapnil Kundu Argha, Abdullah Al Shafi, Rowzatul Zannat, Shoumik Barman Polok, Abdul Muntakim, Jannatul Ferdousi et al. No summary available — see the abstract on arXiv.

Trustworthy, Explainable, and Sustainable Decentralized Intelligence for 6G Networks

Giovanni Perin, Michele Rossi, Enrique Tom\'as Mart\'inez Beltr\'an, Fernando Torres-Vega, Jos\'e Mar\'ia Jorquera Valero, Manuel Gil P\'erez et al. cross-listed No summary available — see the abstract on arXiv.

SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity

Md Ashik Khan, Md Nahid Siddique No summary available — see the abstract on arXiv.

Generation of Custom Solvers in Rust for Convex Optimization

Hao Zhu, Joschka Boedecker cross-listed No summary available — see the abstract on arXiv.

A Multi-Resolution Multi-Domain Pre-Training Framework for Universal Traffic Forecasting

Zhouyang Liu, Jindong Han, Hao Wang, Xinyue Liu, Hui Gao, Dongsheng Li et al. No summary available — see the abstract on arXiv.

Map Users and Mapmakers: The Scope of Cognitive Attribution from Acquired Representations

Yiling Wu No summary available — see the abstract on arXiv.

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Shuhuai Huang, Jingfeng Zhang, Hong Jia cross-listed No summary available — see the abstract on arXiv.

CyclOT: Learning Quadratic Optimal Transport Maps via Synchronized Forward-Backward Interpolants

Shizhou Xu, Jiachen Liu, Shih-Hsin Wang, Stefan Broecker, Yuhao Huang, Bao Wang et al. cross-listed No summary available — see the abstract on arXiv.

Lie to me: Detecting Managerial Evasiveness in Earnings Calls via Conversational Audio Encoders

Huizhong Chen, Huan Zhang No summary available — see the abstract on arXiv.

URCHIN: A Horizontal Spiking Language Model for Data-Constrained Pretraining

Po-Han Chiang cross-listed No summary available — see the abstract on arXiv.

HQARRF: Hierarchical Q-learning and Force-aware Routing for Multi-Charger Scheduling in Wireless Rechargeable Sensor Networks

Liang-Ching Tao, Pi-Chung Wang cross-listed No summary available — see the abstract on arXiv.

Phorecaster365: A Human-Supervised Reference Architecture for Hybrid Pharmaceutical Sales Forecasting and Planning Decision Support

Houman Kazemzadeh, Kamyar Naderi No summary available — see the abstract on arXiv.

DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis

Ziyu Zhang, Tianlun Zuo, Hanzhao Li, Haoyu Zhang, Lei Xie cross-listed No summary available — see the abstract on arXiv.

Machine Learning under Imperfect Data: Challenges and Methods

Masoumeh Zareapoor No summary available — see the abstract on arXiv.

North Small Translate: Advanced Cost-Effective Translation (Cohere CAT+)

Tom Kocmi, Alexandre B\'erard, Phil Blunsom, Samuel Cahyawijaya, Shaun Cassini, Nicholas Frosst et al. No summary available — see the abstract on arXiv.

Learning Through Energy Refinement and Manifold Projection: A Cooperative EBM-AE Framework

Ryad Zemouri No summary available — see the abstract on arXiv.

LoRA Fine-Tuned Models for Control Systems Course Q\&A: A Multidimensional Evaluation of Model Scale and Rank Effects

Shaowen Lu, Chengxu Liu, Ping Zhou, Tao Yang No summary available — see the abstract on arXiv.

Machine Learning in Fish Farming

Fearghal O'Donncha, Nikos Papandroulakis, Jennie Korus, Abigail Langbridge, Alexander Timms, Konstantinos Topouzelis et al. No summary available — see the abstract on arXiv.

Finite-Time Node Separation in Recurrent Graph Neural Networks with Persistent Gaussian Perturbations

Mostafa Haghir Chehreghani No summary available — see the abstract on arXiv.

Minibatch persistency, eight years later: what batch reuse costs in steps and joules, and what it saves in data

Matteo Fischetti No summary available — see the abstract on arXiv.

Equilibrium bias and convergence in augmented primal--dual dynamics with sampled constraints

Kang Liu, Mengxiao Chen, Siqi Xiong, Yi Xia cross-listed No summary available — see the abstract on arXiv.

Exploring napping paradigm for Recurrent Spiking Neural Networks

Andreas Massey, Stefano Nichele, Aliaksandr Hubin No summary available — see the abstract on arXiv.

Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus

Levent Bulut No summary available — see the abstract on arXiv.

Optimal Transport for Efficient, Unsupervised Anomaly Detection on Industrial Data

Abigail Langbridge, Fearghal O'Donncha, James T Rayfield, Bradley Eck No summary available — see the abstract on arXiv.

CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education

Rodrigo Agerri, Itziar Aldabe, Elena Cabrio, Mark Cieliebak, Jan Deriu, Mariana Flores et al. No summary available — see the abstract on arXiv.

SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery

Shaghayegh Sadeghi, Stephen L. Smith, David C. Del Rey Fern'andez No summary available — see the abstract on arXiv.

Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision

Chengwei Zhou, Abu Masum, Xuming Chen, Mehran Moghadam, Sreetama Sarkar, Arnab Sanyal et al. No summary available — see the abstract on arXiv.

Thought without systematicity? Evaluating reasoning models on rule induction tasks

Simon Schug, Brenden M. Lake No summary available — see the abstract on arXiv.

Linear Ensemble Sampling with Smaller Ensembles

Taehyun Hwang, Min-hwan Oh No summary available — see the abstract on arXiv.

Tabby: An Open Pretraining Recipe for Time Series Foundation Models

Shifeng Xie, Bahaeddine Abdessalem, Zehao Xiao, Youssef Attia El Hili, Ambroise Odonnat, Zhiwei Dong et al. No summary available — see the abstract on arXiv.

SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection

Hanjuan Huang, Yung-Chieh Yeh, Hsing-Kuo Pao cross-listed No summary available — see the abstract on arXiv.

Introspective Uncertainty Estimation for LLM-Based Code Generation

Thomas Klassert cross-listed No summary available — see the abstract on arXiv.

Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context

Nawar S. Alseelawi, Mustafa S. Aljumaily No summary available — see the abstract on arXiv.

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez cross-listed No summary available — see the abstract on arXiv.

Diffusion-Based Multiple-Shooting Indirect Optimal Control for Fuel-Optimal Spacecraft Trajectory Generation

Saeid Tafazzol, Ehsan Taheri, Ryne Beeson cross-listed No summary available — see the abstract on arXiv.

Synthetic Data in Marketing Research: How to Evaluate and When to Trust

Oded Netzer, Rajan Sambandam No summary available — see the abstract on arXiv.

Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR

Yukang Zhu, Zhen Han No summary available — see the abstract on arXiv.

Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

Minsun Shim, Ramisha Raida Karim, Ruthwik Jakkula, Kaiwen Zhou, Xin Liu, Xin Eric Wang et al. cross-listed No summary available — see the abstract on arXiv.

Convergent Emergence of In-Context Learning Across Modalities

Nathan Breslow, Seungwook Han, Daniel Hyunsoo Lee, Aayush Mishra, Anqi Liu, Daniel Khashabi No summary available — see the abstract on arXiv.

Schizophrenia Detection from EEG Signals: A Transformer Framework with Spectrogram Representation

Abtin Shafiei, Mohsen Hooshmand, Majid Ramezani No summary available — see the abstract on arXiv.

TF-IDF and BM25 Are Exact KL Divergences

Ivan Silajev cross-listed No summary available — see the abstract on arXiv.

VeriDx: Earning the Right to Diagnose with Disease-Centric Verification

Zhong Cao, Shuying Chen No summary available — see the abstract on arXiv.

Conditional Quantum Flow Matching for Data-Scarce Physiological Signal Augmentation

Chi-Sheng Chen, Samuel Yen-Chi Chen cross-listed No summary available — see the abstract on arXiv.

Real-World Deployment and Performance Characterisation of Fog-Based Deep Learning for Cold-Chain Temperature Prediction over LoRaWAN

Jeremiah Taguta, Jean Frederic Isingizwe Nturambirwe, Clement Nthambazale Nyirenda cross-listed No summary available — see the abstract on arXiv.

Data-Efficient Agentic Graph Domain Adaptation via Reliability-Aware Prototype Learning

Yingxu Wang, Kunyu Zhang, Siyang Gao No summary available — see the abstract on arXiv.

Measuring the Creativity of Frontier LLMs in Automated Research

Yiheng Zhao, Mengzhuo Chen, Chengming Hu, Pengyi Liao, Yiran Pang No summary available — see the abstract on arXiv.

AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents

Xiaoqun Liu, Qiben Yan cross-listed No summary available — see the abstract on arXiv.

Stabilizing Performative Feedback Loops with Minimal Model Deployments

Gabriele Farina, Juan Carlos Perdomo No summary available — see the abstract on arXiv.

GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning

Zhongyu Wang No summary available — see the abstract on arXiv.

Multi-Modal Tumor Survival Prediction via Graph-Guided Mixture of Experts

H Mathavan, H Liu No summary available — see the abstract on arXiv.

LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models

Kunwei Wu, Xiang Liu, Guocai Yao, Junming Chen, Zhikang Chen, Min Zhang et al. cross-listed No summary available — see the abstract on arXiv.

Just add noise: Debiasing tree-based variable importance in mixed data

Jiahe Li, Omar Melikechi cross-listed No summary available — see the abstract on arXiv.

NeuroFlex: Lossless Element-Level ANN-SNN Co-Execution for Efficient Sparse Inference

Varun Manjunath, Pranav Ramesh, Gopalakrishnan Srinivasan cross-listed No summary available — see the abstract on arXiv.

Bridging the Synthetic-to-Real Gap for Few-Shot Cryo-ET Classification

Siddhant Bharadwaj, Ashish Vashist, Rashi Singh, Pranav Vinodh, Nishanth Artham, Runmin Jiang et al. cross-listed No summary available — see the abstract on arXiv.

Neuron Activation-based Computation of Logical Explanations for Deep Neural Networks

Tom\'a\v{s} Kol\'arik, Faezeh Labbaf, Fabrizio Leopardi, Grigory Fedyukovich, Michael Wand, Natasha Sharygina cross-listed No summary available — see the abstract on arXiv.

RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes

Abhirama Subramanyam Penamakuri, Shreya Shukla, Anand Mishra cross-listed No summary available — see the abstract on arXiv.

A Voxel-Spacing-Aware Extension of PyRadiomics for Anisotropic Texture Analysis

David Corral Fontecha, Juan Miranda Bautista, Pablo Menendez Fern\'andez-Miranda, Andrea Trapote Fernandez, Lara Lloret Iglesias, Jose A. Vega cross-listed No summary available — see the abstract on arXiv.

Real-Time Synthesis of Robust Controlled Invariant Sets for Monotone Systems

Yasin Sonmez, Mahmoud Khaled, Majid Zamani, Murat Arcak cross-listed No summary available — see the abstract on arXiv.

Talking to Me or Someone Else? Rethinking Talk-to-Me Detection in Egocentric Videos

Feiyu Du, Xi He, Jia Li, Yapeng Tian, Weili Wu cross-listed No summary available — see the abstract on arXiv.

Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had

Achille Zappa No summary available — see the abstract on arXiv.

To do($x$) or not to do($x$): Medical Image Counterfactuals for Dataset Augmentation

Yasin Ibrahim, Robin J. Evans, Konstantinos Kamnitsas No summary available — see the abstract on arXiv.

Exact Finite Attention Responses From RoPE Derivatives

Julie Huang, Maggie Chlon, Gregory Gutin, Leon Chlon cross-listed No summary available — see the abstract on arXiv.

LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents

Siddharth Sharma, Nilesh Prasad Pandey, Onat Gungor, Tajana Rosing No summary available — see the abstract on arXiv.

Riemannian ascent--descent for nonconvex nonconcave minimax landscapes: convergence to basin saddle points and applications to distributionally robust optimization

Rishabh Dixit, Pranav Upadrashta, Alex Cloninger cross-listed No summary available — see the abstract on arXiv.

T-SMART: Mechanism-Level Attribution for Tool-Augmented Time-Series Question Answering

Ivan Delgado, Himansi Gupta, Bishal Khatri, Niharika Sapre, Lameta Shamoon, Onat Gungor et al. No summary available — see the abstract on arXiv.

One Size Does Not Fit All: Setting Inference Depth from the Questions a Deployment Actually Asks

Jerry Kaplan No summary available — see the abstract on arXiv.

A New Transformer-Based Approach for Audio-Based Kinship Verification and a New Uncontrolled Mandarin Kinship Speech Dataset

Qiyang Sun, Langqing Zhang, Yupei Li, Bj\"orn Schuller cross-listed No summary available — see the abstract on arXiv.

When Faster VLA Deployment Changes Closed-Loop Behavior: Task Success-Latency Analysis of SmolVLA Across PyTorch and ONNX Variants

Rafiqul Islam cross-listed No summary available — see the abstract on arXiv.

Signatures of Steerability in Activation Space of Language Models

Prajjwal Bhattarai, Tuka Alhanai No summary available — see the abstract on arXiv.

When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering

Saanvi Paturi, Arsen Kenzhebayev, Arham Sethi, Vyas Raina, Ivaxi Sheth, Vatsal Raina No summary available — see the abstract on arXiv.

3D Gait-Based Autism Classification Using Attention-Enhanced Deep Learning with Cross-Fold Statistical Stability Analysis

Md Nadim Mahamood, Md Arif Shahriar, Md Parvej Sikder, Md Rasul Islam, Md Shafi Ud Doula, Md Ashraful Alam et al. cross-listed No summary available — see the abstract on arXiv.

Towards Evolving Context Parameterization for Large Language Models

Xiaobing Shi, Zherui Li, Yiming Jiang, Kun Wang, Yufei Guo No summary available — see the abstract on arXiv.

CyFM: Cylindrical Optimal Transport for Few-Step Complex-Valued Flow Matching

Marcel Musia{\l}ek, Iga Wolanin, Damian Ryczko, Anna Grelewska, Oleksii Furman No summary available — see the abstract on arXiv.

Inherited Heads: Audio language models track speakers with their text backbone's attention, and an attention-mass ranking retrieves a different set

Bojro Das No summary available — see the abstract on arXiv.

A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement

Carmel Kronfeld, Sharva Gogawale, Tetsuro Kobayashi, Irad Ben-Gal No summary available — see the abstract on arXiv.

A Machine Learning Framework for Fault Detection, Isolation, and Severity Prediction of Autonomous VTOL Aircraft

Ripon C. Sarker, Pedram H. Dabaghian, Raman Goyal, Atanu Halder No summary available — see the abstract on arXiv.

ZAPS: Zero-Cost Active Proxy Search for Neural Architecture Search

Hassan Touayouch, Rabie Najem, Mohammed Benjelloun No summary available — see the abstract on arXiv.

Bi-Level Routing and Sparse Spatial Attention based Multi-View BEV 3D Object Detection for Autonomous Driving

Jing Zhang, Jiaqi Liu, Zibo Wang cross-listed No summary available — see the abstract on arXiv.

Entropy-Punctured Bloom Filters for Memory-Efficient Machine Learning

John Cartmell, Mihaela Cardei, Ionut Cardei No summary available — see the abstract on arXiv.

Data-free On-policy Distillation

Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu et al. No summary available — see the abstract on arXiv.

Learning to Refer from Estimated Listener Gaze

T\'ea Wright, Alane Suhr No summary available — see the abstract on arXiv.

ECAS: An Edge-Controlled Agentic System for Validation-Gated Scientific Application Execution

Baixi Sun, Mingze Xia, Huihuo Zheng cross-listed No summary available — see the abstract on arXiv.

Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution

Zhuojin Li, Marco Paolieri, Leana Golubchik cross-listed No summary available — see the abstract on arXiv.

Task-Specified Active Metrological Inspection with Measurement-Steered VLA Manipulation and Deterministic Evidence Gating

Zhiling Chen, Jingzhan Ge, Ruimin Chen, Matthew P. Castanier, David Gorsich, Farhad Imani cross-listed No summary available — see the abstract on arXiv.

Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports

Sarah Y. Li, Elijah Renner, Rayan Ansari, Alaa Youssef No summary available — see the abstract on arXiv.

Modeling, Scaling, and Decoding: Optimizing Controllable Speech Generation with Nonverbal Vocalizations

Ziyu Zhang, Yun Chen, Taihui Wang, Hanzhao Li, Qicong Xie, Rilin Chen et al. cross-listed No summary available — see the abstract on arXiv.

Graph-Transformer Fraud Detection with Self-Supervised Pretraining and Conformal Risk Control

Sergei, Komarov No summary available — see the abstract on arXiv.

Assessing the Applicability of Existing Design Recommendations to AI Companion Design: A Multi-Method Study

Soobin Cho, Deveshi Modi, Divya Mavinkurve, Jieqiong Ding, Mark Zachry cross-listed No summary available — see the abstract on arXiv.

OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving

Zikun Li, Yixuan Mei, Shiqi Pan, Zixuan Chen, Xiaowen Zhang, Mengdi Wu et al. cross-listed No summary available — see the abstract on arXiv.

CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time

Nitish Kovuru, Prateek Jannu No summary available — see the abstract on arXiv.

The Attribution-Compression Frontier in Retrieval-Augmented Generation

Deepanshu Mody No summary available — see the abstract on arXiv.

Joint Optimization for Federated Learning and Transmission over Unreliable Wireless Networks with Heterogeneous Data

Changheng Wang, Xianchao Zhang, Zhiqing Wei, Lingzhu Zhao, Zhongming Yang, Zhiyong Feng cross-listed No summary available — see the abstract on arXiv.

What Input Resolution Is Required for Bird Species Identification, and What Is Its Latency Cost on an Edge Device? A Study of 14 Input Resolutions and Six Architectures with On-Device Measurements

Takeshi Nishikawa cross-listed No summary available — see the abstract on arXiv.

ATTRICITE: Training an Open 4B Model for Citation Recovery toward Faithful Attribution

Yee Man Choi, Xuehang Guo, Songcheng Cai, Yimu Wang, Yi R. Fung, Qingyun Wang cross-listed No summary available — see the abstract on arXiv.

Towards Anticipatory Databases Through Shared Data and Workload Semantics

Farzaneh Zirak, Kasper Overgaard Mortensen, Farhana Choudhury, Renata Borovica-Gajic cross-listed No summary available — see the abstract on arXiv.

Document Topic Alignment Metrics for Evaluating Topic Models of Short-Text Public Health Communications on Social Media

Wangjiaxuan Xin, Shuhua Yin, Yaorong Ge, Shi Chen No summary available — see the abstract on arXiv.

DenMark: Robust Semantic Watermarking for Diffusion Language Models

Tianhao Ma, Weihao Xuan, Dong-Dong Wu, Farshid Nooshi, Takashi Ishida, Gang Niu et al. No summary available — see the abstract on arXiv.

VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching

Prajwal Koirala, Mark Campbell cross-listed No summary available — see the abstract on arXiv.

Bayesian optimization with kernel ensembles and disagreement-based acquisition for source localization and acoustic inversion

Heng Zhang, Haotian Xiang, Florian Meyer, Qin Lu No summary available — see the abstract on arXiv.

Parameter Estimation of Ringdown Quasinormal Modes with Autoencoder

Momoka Iida, Hayato Motohashi, Hirotaka Takahashi cross-listed No summary available — see the abstract on arXiv.

SpermYOLO: A Coordinated YOLO-Based Detector for Accurate and Efficient Sperm and Impurity Detection in Microscopic Images

Shengqi Chen, Zilin Wang, Xingyu Pan, Wenting Yu, Pengchao Deng, Guohua Wu cross-listed No summary available — see the abstract on arXiv.

Biquaternionic Space with Complex-valued Attention for Temporal Knowledge Graph Completion

Rushan Geng, Cuicui Luo No summary available — see the abstract on arXiv.

Editorial routing shapes how computational results are qualified in AI-assisted scientific writing

Jihan Kim No summary available — see the abstract on arXiv.

Fusing Spectral Signatures and Activation Clustering for Backdoor Detection in Healthcare Imaging Models: Method, Implementation, and Evaluation

Suresh Tamang cross-listed No summary available — see the abstract on arXiv.

Learning Source Acquisition Policies by Offline Planning

Ziqi Zhao, Run Xu, Qingjian Ni No summary available — see the abstract on arXiv.

E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning

Xiaoya Wang, Yutong Xu, Junjie Wang No summary available — see the abstract on arXiv.

S3-Tracker: Self-Supervised Surgical Tissue Tracking With Contrastive Random Walks

Jiaming Zhang, Zijian Wu, Mehran Armand, Septimiu Salcudean cross-listed No summary available — see the abstract on arXiv.

SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization

Zian Liu, Yiwen Hu, Zican Dong, Tian Xie, Wayne Xin Zhao, Yucheng Ding et al. No summary available — see the abstract on arXiv.

AI Assisted Workflow Optimization and Automation

Zhen Zhong cross-listed No summary available — see the abstract on arXiv.

Nonparametric Variance-Penalized Actor-Critic: Statistical Inference for Risk-Sensitive Reinforcement Learning

Saunak Kumar Panda, Tong Li, Yisha Xiang, Ruiqi Liu No summary available — see the abstract on arXiv.

Communication-Efficient LLM Adaptation over Decentralized GPU Meshes

Sameera Ramasinghe, Shamane Siriwardhana, Thalaiyasingam Ajanthan, Hadi Mohaghegh Dolatabadi, Chamin P Hewa Koneputugodage, Gil Avraham et al. No summary available — see the abstract on arXiv.

AURA: Unified Multimodal Framework for Conversational Music Editing

Quoc-Huy Trinh, Minh-Van Nguyen, Debesh Jha cross-listed No summary available — see the abstract on arXiv.

Multimodal deep learning from spectra for small-molecule structure identification: enhancing robustness with mixed-condition training

Bowen Gao, Lei Zhu, Yiying Wang, Wenjie Yu No summary available — see the abstract on arXiv.

LLaTSA: Large Language Model-Aligned General-Purpose Transient Stability Analysis

Chao Shen, Hongwei Zhen, Junyan Shao, Zhenghao Yang, Yifan Zhang, Mingyang Sun cross-listed No summary available — see the abstract on arXiv.

Formal Properties of Language as Constraints on Neural Dynamics

Elliot Murphy No summary available — see the abstract on arXiv.

MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents

Zhenyu Zhang1, Jiudong Yang No summary available — see the abstract on arXiv.

Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error

Hongliu Cao No summary available — see the abstract on arXiv.

Neural Modal Decomposition: Architectural Priors from Observables

Juho Park, Kaushik Sengupta cross-listed No summary available — see the abstract on arXiv.

Dynamic Learning Solutions: A System for Personalized Educational Video Generation

Siddhanth Sridhar, Shreya Chaurasia, Baddela Sai Yaswantha Reddy, Deepak Parmar, Shylaja S S No summary available — see the abstract on arXiv.

Question's Gambit: The First Move Matters in Agentic Deep Search

Radin Hamidi Rad, Amin Bigdeli, Negar Arabzadeh, Sajad Ebrahimi, Charles L. A. Clarke, Benjamin C. M. Fung et al. No summary available — see the abstract on arXiv.

A Hybrid Dependency-Aware Framework for Task Decomposition and Dynamic Agent Generation in Oracle-to-PostgreSQL Migration

Oleg Grynets, Oleg Kaskun, Alona Seletska, Daryna Tukalo, Vasyl Lyashkevych cross-listed No summary available — see the abstract on arXiv.

Surrogate-Assisted Genetic Programming with Phenotypic Characterisation in Dynamic Multi-Mode Project Scheduling

Yuan Tian, Yi Mei, Mengjie Zhang cross-listed No summary available — see the abstract on arXiv.

A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification

Leon Fernando, C Dombawala, P. Hettigoda, Vanodhya G. Warnasooriya, Ishara Neranjana, Rashmika Nawaratne cross-listed No summary available — see the abstract on arXiv.

Safety Signals to Verify NetOps Agents with Action-Level Granularity

Tobias Labarta, Frederik Pahde, Novak Boskov, Maximilian Dreyer, David Birkenberger, Manzoor Ahmed Khan et al. No summary available — see the abstract on arXiv.

Certification cost of quantum models: measurement correlation, not parameter count

Pavel Sulimov, Claude Lehmann cross-listed No summary available — see the abstract on arXiv.

Has Scientific Talent Shifted from Depth to Breadth?Evidence across Papers, Knowledge Inputs, Careers, and Teams

Xiaoshn Nee, Haobo Zhong, Xiaomin Ni cross-listed No summary available — see the abstract on arXiv.

Lightweight Generalized DeepFake Face Detection with WAVIE: Wavelet Augmented Vision Intermediate Embeddings

Arya Pulkit, Aditya Ruhela, Akarshan Kapoor, Arnav Bhavsar cross-listed No summary available — see the abstract on arXiv.

A latent dimension of Condorcet's jury theorem for multiple AI advisers

Kazutoshi Sasahara, Aoi Naito, Ryo Fujie cross-listed No summary available — see the abstract on arXiv.

From Visual Attribution to Clinical Reasoning: Explainable Parkinson's Disease Screening from Hand-Drawn Patterns

Aritra Dey, Utsav Kumar Nareti, Chandranath Adak, Soumi Chattopadhyay, Krishna Gopal Sasmal, Saeed Anwar cross-listed No summary available — see the abstract on arXiv.

NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass

Ali Derogar Odolou, Reza Nazari, Mostafa Salehi No summary available — see the abstract on arXiv.

Towards Identifying the Dataset Biases Causing Phantom Transfer

Jonas J\"ur{\ss}, Pietro Li\`o No summary available — see the abstract on arXiv.

Follow the Geometry, Not the Model: Cold Start Semi-Supervised Learning

Itai David, Daphna Weinshall No summary available — see the abstract on arXiv.

Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation

Ziyu Zhang, Mingchen Shao, Wenjie Tian, Tianlun Zuo, Longhao Li, Lei Xie cross-listed No summary available — see the abstract on arXiv.

Retrieval-Guided Fine-Tuning as Noisy Estimation: Risk bounds and Architectural Analysis

Bhargav Lad, Yifan Hao No summary available — see the abstract on arXiv.

EdgeHAR: An Edge-Native Compact Sensor Foundation Model for Human Activity Recognition

He Zhang, Siyu Yuan, Siyu Liu, Sizhen Bian, Bin Guo No summary available — see the abstract on arXiv.

When does a scaling result justify a different allocation? A critical review of resource-allocation evidence for AI systems

Seyed Morteza Emadi No summary available — see the abstract on arXiv.

Should All Noises Be Treated Equally: Impact of Input Noise Variability on Neural Network Robustness

Salma Alsinan, Maksim Makarenko, Sixiu Liu, Ali Aldawood, Ibrahim Hoteit No summary available — see the abstract on arXiv.

Physically Partitioned KVCache Format for CPU--GPU Load Balancing in MoE Inference

Enda Yu, Dezun Dong, Xiangke Liao cross-listed No summary available — see the abstract on arXiv.

Beyond Scene Description: Multi-Agent Orchestration for Non-visual Access to Virtual Worlds

Toqeer Ali Syed, Ali Akarma, Adeel Ahmad, Danial Hameed No summary available — see the abstract on arXiv.

OptoAgent: A Trustworthy Multi-Agent Framework for Opportunistic Vision Micro-Screening in Classroom Environments

Toqeer Ali Syed, Ali Akarma, Adeel Ahmad, Hammad Muneer No summary available — see the abstract on arXiv.

Toward a Layer-2 Trigger for AI/ML Lifecycle Management in 6G

Dharmendra Kumar cross-listed No summary available — see the abstract on arXiv.

Theseus in the Graph: Towards Traceable Multi-Hop Graph Navigation

Eduin E. Hernandez, Luis F. Garcia, Nurassyl Askar, Sergio A. Diaz, Stefano Rini No summary available — see the abstract on arXiv.

Parameter-Efficient Quantum NLP for Paraphrase Detection: Performance, Robustness, and Entanglement

Farha Nausheen, Khandakar Ahmed, Farina Riaz cross-listed No summary available — see the abstract on arXiv.

Sharing standardized image-derived data in computational pathology using DICOM

Daniela P. Schacherer (Fraunhofer Institute for Digital Medicine MEVIS, Bremen, Germany), Christopher P. Bridge (Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital, Boston et al. cross-listed No summary available — see the abstract on arXiv.

Multi-source conformal prediction: leveraging heterogeneity via localization

Rohan Hore, Anirban Chatterjee, Sayantan Choudhury cross-listed No summary available — see the abstract on arXiv.

Proving olympiad geometry theorems on a superconducting quantum processor

Ning Wang, Zheng-Zhi Sun, Zhengyi Cui, Yiren Zou, Aosai Zhang, Fanhao Shen et al. cross-listed No summary available — see the abstract on arXiv.

GNN4PPM: Multi-Target Predictive Process Monitoring with Relational Graph Convolutional Networks

Ana Costa, Johannes M\"akelburg, Luise Pufahl No summary available — see the abstract on arXiv.

Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition

Ahmad Amirivojdan, Farzad Nadiri, Abolfazl Alizadeh, Shaghayegh Yaraghi No summary available — see the abstract on arXiv.

Selecting k Paths with the Minimum Longest Path Length in the Stochastic Semi-Bandit Setting

Shunsuke Aoki, Atsuyoshi Nakamura No summary available — see the abstract on arXiv.

TATK: Triple-Aware Top-K Learning with Knowledge-Grounded Verification for LLM-based Sequential Recommendation

Yuchen Guan, Jiaye Liu, Yifei Han, Zhenxi Zhang, Yixuan Weng, Bin Li No summary available — see the abstract on arXiv.

Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real Robots

Abdalwhab Bakheet Mohamed Abdalwhab, Giovanni Beltrame, David St-Onge cross-listed No summary available — see the abstract on arXiv.

Evaluation of optimisation and Bayesian inference methods for reaction rates in atmospheric chemical mechanisms

Valery Ashu, Wenqing Peng, Zhi-Song Liu, Heikki Haario, Andreas Rupp, Taiwo Ashu et al. cross-listed No summary available — see the abstract on arXiv.

Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection

Ulugbek Shernazarov, Charitha Ruwansiri Weerakon Basnayake, Abdelkhaleq El Jarjini, Noel Crespi, Praboda Rajapaksha No summary available — see the abstract on arXiv.

Domain-specific Pretraining Profile and Transformer Performance: Evidence from Modeling Digital Pragmatics in Arabic-English Code-switching

Fahad Al Hussen, King Saud University, Riyadh, Saudi Arabia, Mohammed Q. Shormani, Ibb University et al. No summary available — see the abstract on arXiv.

AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory

Sushan Adhikari cross-listed No summary available — see the abstract on arXiv.

Pathwise Individual Rationality in Federated Learning: A Mechanism-Architecture Co-Design

Amin Meghrazi, Srinivasan Parthasarathy, Andrew Perrault No summary available — see the abstract on arXiv.

AI Deployment Accountability Engineering: A Vision for Accountable AI in Safety-Critical Socio-Technical Systems

Murat Kantarcioglu No summary available — see the abstract on arXiv.

SENTINEL: A Multi-Pathway Architecture for Detecting Living-Off-the-Land APT Attacks on Windows Command Lines

Ahad Bin Islam Shoeb, Kamrul Hasan, Jamal Uddin Tanvin, Liang Hong, Imtiaz Ahmed, Md Arif Billah et al. cross-listed No summary available — see the abstract on arXiv.

Diagnosing Temporal Misalignment in Multichannel Time-Series Classification with Minimum Description Length

Sebastian Buschj\"ager, Michael Frichert, Daniel Kuhe, Jian-Jia Chen No summary available — see the abstract on arXiv.

A note on goal-based hierarchical RL

Kevin Murphy No summary available — see the abstract on arXiv.

SH-WRNN: Implicit Spherical Harmonics Weight Field Routing Neural Networks for Asymmetric Edge Intelligence

Zhibin Jiao, Xiangjing An No summary available — see the abstract on arXiv.

Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World

Guocun Wang, Kenkun Liu, Guorui Song, Jing Lin, Zhe Huang, Luyuan Zhang et al. cross-listed No summary available — see the abstract on arXiv.

Channel-Adaptive Region Adjacency Graph Carriers for Semantic Image Communication

Karim Abdallah, Maria Slim, Mariette Awad, Hadi Sarieddeen cross-listed No summary available — see the abstract on arXiv.

Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation

Zhiyu Gui, Kexin Huang, Jia Guo, Junkang Wu, Zihao Wang, Zhiqiang Zhang et al. No summary available — see the abstract on arXiv.

DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents

Zhichao Shi, Wenjie Zhang, Xuhui Jiang, Xiaojun Wu, Cehao Yang, Chengjin Xu et al. No summary available — see the abstract on arXiv.

Investigating the Impacts of Generative AI on Information Seeking

Alexi Orchard, Shannon Lodoen cross-listed No summary available — see the abstract on arXiv.

Diffusion-Based Generation of Gait Trajectories

Damian Benasco, Juan Carballeira-Lopez, Jaime Ramos-Rojas, Julio S. Lora-Millan, Antonio J. Del-Ama, David Rodriguez-Cianca et al. No summary available — see the abstract on arXiv.

CompCQR: Compositional Query Generation for Training-Free Conversational Search

Yunah Jang, Kang-il Lee, Joongbo Shin, Kyomin Jung No summary available — see the abstract on arXiv.

Skill Composition for Legged Robot Reinforcement Learning

Daniel Gigliotti, Flavio Maiorana, Fabio Patrizi, Luca Iocchi cross-listed No summary available — see the abstract on arXiv.

Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation

Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull, Caitlin A. Stamatis, Olivier Tieleman, Guilherme B. Freire et al. No summary available — see the abstract on arXiv.

Natural Language Knowledge Graph Query Execution: Leveraging Controlled Semantics in the LLM Context Window

Blake G. Fitch cross-listed No summary available — see the abstract on arXiv.

Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing

Sehwan Park, Taehoon Kim, Geonhee Han, Dohyun Kim, Seung Wook Kim, Paul Hongsuck Seo cross-listed No summary available — see the abstract on arXiv.

MANE: A Multi-Path Adaptive Network for Edge Onloading of Deep Neural Networks

Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias, Iakovos S. Venieris cross-listed No summary available — see the abstract on arXiv.

Symmetries and Singularities

Vishnu Varadarajan, Mihir More, Aritra Das, Debayan Gupta No summary available — see the abstract on arXiv.

Exploring Multimodal Turn-Taking Cues in Face-to-Face Conversation using Voice Activity Projection

Willem Berner, Julio Cesar Cavalcanti, Kalle {\AA}str\"om, Gabriel Skantze cross-listed No summary available — see the abstract on arXiv.

PU classification under Non-SCAR: clustering-assisted logistic model with oversampling enhancement

Konrad Furma\'nczyk, Kacper Paczutkowski cross-listed No summary available — see the abstract on arXiv.

The Garden of Forking Prompts: How Users Explore Narrative Space in Story Generation

Advait Deshmukh, Nora Benedict, Melanie Walsh, Maria Antoniak No summary available — see the abstract on arXiv.

One Feedback System Does Not Fit All: Localising Data-to-Text Driver Coaching for the United Kingdom and Nigeria

Iniakpokeikiye Peter Thompson, Jawwad Baig, Ehud Reiter, Dewei Yi No summary available — see the abstract on arXiv.

Speak to the City: Multimodal Resolution for Outside-the-Vehicle References

Alireza Parchami (Mercedes-Benz Tech Innovation GmbH, Saarland University), Artin Saberpour (Saarland University), Robin Connor Schramm (Mercedes-Benz Tech Innovation GmbH, RheinMain University of Applied Sciences), J\"urgen Steimle (Saarland University) et al. cross-listed No summary available — see the abstract on arXiv.

WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians

Johann Birnick, Rayan Saab No summary available — see the abstract on arXiv.

Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

Yecheng Wu, Song Han, Han Cai No summary available — see the abstract on arXiv.

An immune world model for multiscale forecasting and therapeutic hypothesis generation

Taoyong Cui, Xi Wang, Zonghang Li, Jinchao Ding, Lingsen You, Yuzhi Xu et al. No summary available — see the abstract on arXiv.

GRPO-QM: Target Preserving Exploration for Quantum Tomography

Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling No summary available — see the abstract on arXiv.

Learning Metastable Dynamics

Rupak Majumdar, Mahmoud Salamati, Nikhil Singh, Sadegh Soudjani cross-listed No summary available — see the abstract on arXiv.

Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M

Dushyant Rajput (AltSlate Labs LLP), Nirdesh Chauhan (AltSlate Labs LLP), Siddharth Kosaraju (AltSlate Labs LLP) No summary available — see the abstract on arXiv.

Moral Rebel Agents: Decision-Making Under Conflicting Obligations

Hector Munoz-Avila, David W. Aha, Paola Rizzo No summary available — see the abstract on arXiv.

Carryover Drafting: Recycling Rejected States for Speculative Decoding

Jahyun Koo, Sunghyeon Woo, Jaeeun Kil, Jeongtae Lee, Sungjae Lee, Kyomin Jung et al. No summary available — see the abstract on arXiv.

Are Gradient Boosting Models Suitable for Intermittent Demand Forecasting?

Vladislav Kislinskii, Mazhar Hameed No summary available — see the abstract on arXiv.

Bayesian Intelligence from the Outside

Alex Smolin, Bryan Wilder No summary available — see the abstract on arXiv.

CALICO: A Human-Centered, Codebook-Aligned System for Annotation

Boqin Yuan, Xiaoyi Gu, Fiona Li, Chang Wan, Angel Hsing-Chi Hwang, Jieyu Zhao cross-listed No summary available — see the abstract on arXiv.

Parameter isolation with domain-specific experts for incremental audio classification

Jongyeon Park, Do-Hyeon Lim, Sang-won Park, Hong Kook Kim, Kyungdeuk Ko, Hyeongcheol Geum et al. cross-listed No summary available — see the abstract on arXiv.

WaVeFuse: Regime-Adaptive Equity Index Forecasting via Channel-Wise Wavelet Denoising and Vertical Attention Fusion

Aashish Bohra, Vivek Vijay No summary available — see the abstract on arXiv.

OCT-FedSIR: Toward Trustworthy Federated Ophthalmic Learning under Annotation Noise

Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Rashadul H. Badhon, Behafarin Emam, Sally S. Y. Ong et al. No summary available — see the abstract on arXiv.

HELENA for 5G NR LEO NTN Channel Estimation: A Comparative Evaluation

Miguel Camelo Botero, Nina Slamnik-Krije\v{s}torac, Johann Marquez-Barja cross-listed No summary available — see the abstract on arXiv.

AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing

Vidushee Vats, Karun Sharma, Shengzhi Li, Shichao Pei No summary available — see the abstract on arXiv.

Building Legal Reward Models for Grounding and Abstention

Rilton Franzone, Valentin No\"el, Puyu Wang, Philip Torr, Fabio J. Fehr No summary available — see the abstract on arXiv.

A property-registry contract for retrieve-or-refuse thermal-mechanical lattice search

Shaoliang Yang, Henry Chu, Zu Yashengjiang, Jun Wang cross-listed No summary available — see the abstract on arXiv.

Quantifying the Generation Modality Gap in Speech-Text Language Models

Ju-Chieh Chou, Jiawei Zhou, Karen Livescu No summary available — see the abstract on arXiv.

AcquireBound: Runtime Authorization for Resources Acquired by AI Agents

Genliang Zhu No summary available — see the abstract on arXiv.

Calibrating Interpretability Instruments Before Trusting Their Verdicts

Orion Reblitz-Richardson No summary available — see the abstract on arXiv.

Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return

Arham Sethi, Arsen Kenzhebayev, Saanvi Paturi, Vatsal Raina, Vyas Raina, Ivaxi Sheth cross-listed No summary available — see the abstract on arXiv.

Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families

Orion Reblitz-Richardson No summary available — see the abstract on arXiv.

From Visual Feedback to Textual Reviews: A Multi-Agent Vision-Language Framework for Image-Grounded Review Assistance

Utsav Kumar Nareti, Ayush Bansal, Kumari Priya, Chandranath Adak, Soumi Chattopadhyay, Muhammad Saqib et al. cross-listed No summary available — see the abstract on arXiv.

TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary cross-listed No summary available — see the abstract on arXiv.

CCMAN: Cognitive Instability-Aware Cross-Modal Attention Network for Interpretable Temporal Biomarkers of Verbal Fluency Speech

Madhurananda Pahar, Caitlin Illingworth, Dorota Braun, Daniel Blackburn, Heidi Christensen cross-listed No summary available — see the abstract on arXiv.

A Personalized Dynamic Balance Evaluation Paradigm for Hip Exoskeleton-Assisted Walking under Unexpected Ground Perturbations

Yun Chen, Oluwasegun T. Akinniyi, Qiang Zhang cross-listed No summary available — see the abstract on arXiv.

Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination

Burak Agachan, Max van Duijn, Amirhossein Zohrehvand cross-listed No summary available — see the abstract on arXiv.

How broad is that claim? Mapping Generalisation in NLP Research

Chenxin Diao, Nataliya Stepanova, Emily Allaway No summary available — see the abstract on arXiv.

Pull: Lazy Materialization of Working Memory for Stateful LLM Conversations

Jiangang Chen No summary available — see the abstract on arXiv.

Privacy Preserving Gossip Learning

Erkan Bayram, Mohamed-Ali Belabbas, Tamer Ba\c{s}ar No summary available — see the abstract on arXiv.

Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models

Mingze Yin, Xiaohan Wang, Dian Li, Haichao Yao, Yilin Zhao, Youjun Chen et al. No summary available — see the abstract on arXiv.

The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents

Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Asher Ali, Muhammad Hamzah Siddiqui cross-listed No summary available — see the abstract on arXiv.

Language-Guided Representation Learning for Robust Cross-Sensor Material Recognition

Mashood M. Mohsan, Muhayy Ud Din, Binzhao Xu, Ahmad Abubakar, Irfan Hussain cross-listed No summary available — see the abstract on arXiv.

Mind Which Bird You Favour: Parameterizing Adequacy-Fluency Balance in Meta-Evaluation of Machine Translation

Behzad Shayegh, Niloofar Kazemi No summary available — see the abstract on arXiv.

AI Persuasion as a Threat to Human Control

Joshua Levy, Mick Yang, Kellin Pelrine No summary available — see the abstract on arXiv.

From matrix inversion to constraints: provably tighter confidence regions for importance weights in label shift

Mushan Li, Kihyun Han, Yanyuan Ma cross-listed No summary available — see the abstract on arXiv.

Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?

Afshin Khadangi No summary available — see the abstract on arXiv.

Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accounting Tasks

Kareem Khattab, Omar Khattab, Mohamed Ibrahem No summary available — see the abstract on arXiv.

A Functional SVD Framework for Regularized Multivariate Functional PCA with Dual Penalization

Yue Zhao, Hossein Haghbin, Rebecca Sanders, Mehdi Maadooliat cross-listed No summary available — see the abstract on arXiv.

Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech

Moses Daudu, Adeola Enitan Bamidele, Honor-Jesus Bezaleel No summary available — see the abstract on arXiv.

A primer on evaluation methods for large language models in healthcare

Suzannah E McKinney, Phuc Vu, Samuel A Justice, Christopher Humphries, Alyssa Pradhan, Timothy J Keyes et al. No summary available — see the abstract on arXiv.

Decision-Oriented Uncertainty Quantification for Risk Control in Earth System Spatiotemporal Foundation Models

Ji Lu, Huiran Duan, Bo Zhao, Xianglong Wang, Yiru Fang, Kuo Yang et al. No summary available — see the abstract on arXiv.

MedTRACE: Tool-Augmented Multimodal Clinical Reasoning Agents for Evidence-Grounded Decision-Making

Ji Lu, Lifei Liu, Haoran Yu, Xianglong Wang, Yiru Fang, Kuo Yang et al. No summary available — see the abstract on arXiv.

ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence

Constantinos Papantoniou, Brian Hilton No summary available — see the abstract on arXiv.

Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection

Zeyu Dong, Benjamin Wang, Joyee W. Jin No summary available — see the abstract on arXiv.

Enemray: Toward Capable Language Models for Hassaniya

Cheikh Ahmed No summary available — see the abstract on arXiv.

One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling

Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe, Manish Bhattarai No summary available — see the abstract on arXiv.

Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization

Sarah Wilson, Gail Kaiser, Patrick Musau cross-listed No summary available — see the abstract on arXiv.

El Agente Potente: High-Throughput Agentic Atomistic Simulations

Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang et al. No summary available — see the abstract on arXiv.

Tackling Failure Modes of PINNs and PIKANs Using Conflict-Free Gradients

Sidharth S. Menon, Irina Tezaur, Ameya D. Jagtap No summary available — see the abstract on arXiv.

A Responsive Present, a Shared Past, a Social Other: Teens' Overreliance on Companion AI Chatbots

Mohammad Namvarpour (Matt), Tyler Chang, Afsaneh Razi cross-listed No summary available — see the abstract on arXiv.

Prescreening Point Defects in Semiconductors With Machine Learning

Paul Karlsson, Joel Davidsson, Rickard Armiento cross-listed No summary available — see the abstract on arXiv.

LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions

Myra Cheng, Lujain Ibrahim, Grace Liu, Michelle S. Lam, Vishakh Padmakumar, Nick Madibekov et al. cross-listed No summary available — see the abstract on arXiv.

Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference

Tian Jin No summary available — see the abstract on arXiv.

RAIN: Region-Aware Inversion Network for Semantic Watermark Extraction

Zilai Li cross-listed No summary available — see the abstract on arXiv.

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement

Siwei Wu, Jincheng Ren, Yizhi Li, Haau-Sing Li, Chengran Yang, Yuxuan Zhang et al. No summary available — see the abstract on arXiv.

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman et al. No summary available — see the abstract on arXiv.

One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs

Naihao Deng, Samee Arif, Shuaichen Chang, Yulong Chen, Rada Mihalcea No summary available — see the abstract on arXiv.

Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations

Mohammed Ahnouch, Lotfi Elaachack No summary available — see the abstract on arXiv.

Domain Generalization for Smartphone-Based Human Activity Recognition: A Systematic Analysis of Components and Interactions

Ot\'avio Oliveira Napoli, Edson Borin No summary available — see the abstract on arXiv.

GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems

Xinyu Qiu, Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai No summary available — see the abstract on arXiv.

Interpolation Is Not Invariance: Pair Count Is Not Coverage in Transformation Audits

Mohammed Ahnouch, Lotfi Elaachak No summary available — see the abstract on arXiv.

AgentKV: Phase-Aware KV Eviction for Agentic LLMs

Taowen Tony Liu, Jeffrey T. H. Wong, Can Xiao, Bowen Yang, Hao Mark Chen, Yiren Zhao No summary available — see the abstract on arXiv.

SeqMaestro: From nucleotide sequences to biological hypotheses through interpretable machine learning

Evgeny S. Saveliev, Krzysztof Kacprzyk, Charlotte Capitanchik, Neelanjan Mukherjee, Kate Matlin, Ryan Sheridan et al. No summary available — see the abstract on arXiv.

PeerPen: AI-Assisted Writing for Online Mental Health Peer Support

Jiwon Kim, Sherry Gong, Maya Ajit, Soorya Ram Shimgekar, Yunhao Yuan, Dong Whi Yoo et al. cross-listed No summary available — see the abstract on arXiv.

An explicit solution of the five-expert prediction PDE and the exact optimality set of COMB

Jeff Calder, Nadejda Drenska cross-listed No summary available — see the abstract on arXiv.

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning

Jiayi Yuan, Hangoo Kang, James Jihao Liu, Yejin Choi, Vikram Iyer, Liwei Jiang et al. No summary available — see the abstract on arXiv.

Shapley Value Estimation for Multi-Site Data with Blockwise-Missing Features

Siqi Li, Wangxuan Fan, Yiming Li, Doudou Zhou, Molei Liu cross-listed No summary available — see the abstract on arXiv.

Neural-Network Solutions to Real-Space Charge Density and Generalization

Yuxuan Zeng, Taoyuze Lv, Zhicheng Zhong cross-listed No summary available — see the abstract on arXiv.

Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair

Zewen Tao, Shin-nosuke Ishikawa No summary available — see the abstract on arXiv.

Geometric Signatures of Conceptual Reorganization: A Counterfactual Embedding Framework for Detecting Scientific Revolutions

Dimitris Ntounis, Ariel Schwartzman, Chris Chafe, Thomas A. Ryckman cross-listed No summary available — see the abstract on arXiv.

Steady-State Convergence of Stochastic Approximation

Yixuan Zhang, Qiaomin Xie cross-listed No summary available — see the abstract on arXiv.

Linearized PINN with pretrained nonlinear layers

Wenhao Chen, Alexandre M. Tartakovsky cross-listed No summary available — see the abstract on arXiv.

Cross-Block Conditioning in Deep Boltzmann Machines for Statistical Data Fusion

Junichiro Niimi No summary available — see the abstract on arXiv.

Cloud Workflow Scheduling Based on Graph Attention-Driven Hierarchical Reinforcement Learning

Zongjin Li, Shaohan Feng, Chunxi Yang, Wenbo Wang No summary available — see the abstract on arXiv.

CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models

Alef Iury Siqueira Ferreira, Pedro Lustosa Rege Botelho, Fernanda Silva, Daniel Casanova, Rafael Faustino, Frederico Oliveira et al. cross-listed No summary available — see the abstract on arXiv.

High-Probability Nash Regret for Decentralized Learning in Markov $\alpha$-Potential Games: Episodic and Fully Online Asynchronous Algorithms with Applications to Markov Congestion Games

S. Rasoul Etesami No summary available — see the abstract on arXiv.

Geometric Flow enhanced Graph Coarsening

Chaoqun Fei, Guoxuan Li, Tinglve Zhou, Chuanqing Wang, Yangyang Li No summary available — see the abstract on arXiv.

Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation

M\'at\'e Metzger No summary available — see the abstract on arXiv.

A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator

Yahya Mohamed Elnawasany No summary available — see the abstract on arXiv.

HiGFRL: Hierarchical Graph Fusion-Driven Reinforcement Learning for Dependency-Aware Task Scheduling in Heterogeneous Cloud

Tiangang Li, Shi Ying, Xiangbo Tian No summary available — see the abstract on arXiv.

Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains

Quang Phuoc Nguyen, F\'elix Gaschi, David Anugraha, Santiago Mart\'inez Novoa, En-Shiun Annie Lee No summary available — see the abstract on arXiv.

Towards a knowledge-enhanced single-cell foundation model

Hanqing Zhang, Jie Bao, Mei Ma, Shuai Liu, Jiaying Ma, Jiaguan Liu et al. No summary available — see the abstract on arXiv.

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence

Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou No summary available — see the abstract on arXiv.

MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents

Jianhua Jiang, Dongbo Yuan, Weihua Li No summary available — see the abstract on arXiv.

LiftGCN: Efficient Energy-Preserving Graph Learning via Joukowski Spectral Lifting for Finite Element Stress Prediction

Chen Zeng, Qiao Wang No summary available — see the abstract on arXiv.

Converting Sequenced Fuzzy Cognitive Maps to Causal Virtual Worlds with Large Video Generators

Akash Kumar Panda, Olaoluwa Adigun, Bart Kosko No summary available — see the abstract on arXiv.

ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents

Bingzheng Wang, Xiaoyan Gu, Wentao Wang, Xingyou Yang, Hongcheng Li, Rong Yin cross-listed No summary available — see the abstract on arXiv.

Biomedical Reference Generation Remains Unreliable across 26 Large Language Models

Maxim Topaz, Zhihong Zhang, Nir Roguin, Pallavi Gupta, Zichao Li, Laura-Maria Peltonen No summary available — see the abstract on arXiv.

Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling

Homayoon Beigi, Grace Conneely cross-listed No summary available — see the abstract on arXiv.

Typhoon ASR Streaming: Steerable Low-Latency Thai Speech Recognition with Real-Time Shallow Fusion

Warit Sirichotedumrong, Tanawin Samutsin, Shah Faisal Wani, Sittipong Sripaisarnmongkol, Kunat Pipatanakul No summary available — see the abstract on arXiv.

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Bosi Wen, Cunxiang Wang, Jiayi Gui, Haoke Zhang, Yilin Niu, Pei Ke et al. No summary available — see the abstract on arXiv.

Beyond Depth and Width: The Information-Slack Dilemma in Streaming Test-Time Compute

Xiaotian Zhang (Trooly.AI) No summary available — see the abstract on arXiv.

Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking

Arun Jose, Julian Stastny No summary available — see the abstract on arXiv.

HGTO: A Unified Graph-Based Physics-Informed Formulation for Structural Topology Optimization

Kangzheng Liu, Uday Kumar Punna, Leixin Ma No summary available — see the abstract on arXiv.

IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies

Jinwoong Kim, Sangjin Park cross-listed No summary available — see the abstract on arXiv.

ABSOL: Aggregated Bayesian Subsampling Orchestrated with LLMs

Jackson Hassell, Chen Shen, Estevam Hruschka No summary available — see the abstract on arXiv.

CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems

Chengxin Yu, Zhaoxin Fan, Faguo Wu, Hongwei Zheng, Yun Zhou, Zhiyu Li No summary available — see the abstract on arXiv.

Semantic-TVM: Structure-Preserving Trustworthy Virtual Memory for Memory-Augmented and Tool-Using Agents

Yu Li, Qikun Cai, Tao Huang, Chen Hou No summary available — see the abstract on arXiv.

Overflip: Repetition-Induced Label Flips in Guardrail Models

Xu He, Chih-Hsuan Lin, Hung-Mao Chen, Junjie Xiong, Yan Zhai, Kun Sun No summary available — see the abstract on arXiv.

Steering Generative Robot Policies with Lexicographic Preferences

Yixuan Jia, Jonathan P. How cross-listed No summary available — see the abstract on arXiv.

Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomedical ML

Sidi Chang, Peiying Zhu No summary available — see the abstract on arXiv.

PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift

Yusuf Khalid Shire, Sang-Chul Kim cross-listed No summary available — see the abstract on arXiv.

Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries

Frank Li cross-listed No summary available — see the abstract on arXiv.

SALUTE: Benchmarking and Adapting LLMs for the Defense Domain

Hyeongcheol Park, Sumin In, Suyeon Myeong, Hogun Park, Sangmin Kim, Moonhyun Lee et al. No summary available — see the abstract on arXiv.

TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals

Zihan Xue, Po-Yi Lu, Serhii Honcharenko, Zih-Ching Chen, Hsuan-Tien Lin, Nanyun Peng et al. cross-listed No summary available — see the abstract on arXiv.

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks

Aashiq Muhamed, Mona T. Diab, Virginia Smith, Andrew Ilyas, Matthew Jagielski No summary available — see the abstract on arXiv.

Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache

Frank Li cross-listed No summary available — see the abstract on arXiv.

Horizon-specific Expert Fusion for Photovoltaic Power Forecasting

Xu Yuqing, Zhou Liguo, Sun Ze, Yu Lei, Jiang Mingming No summary available — see the abstract on arXiv.

MoARa: Module-Aware Rank Allocation and Structure-Preserving Decomposition for Low-Rank LLM Pre-training

Keunyoung Kim, Nojun Kwak No summary available — see the abstract on arXiv.

The average-farmer illusion in language-model simulations of agricultural decisions

Zhanliang Zhu, Ziwei Li, Yuchen Liu, Liujun Zhu, Ruiqi Wu, Tongqing Shen et al. No summary available — see the abstract on arXiv.

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar cross-listed No summary available — see the abstract on arXiv.

Data Attribution at Scale via Influence Matrix Estimation

Yuxi Chen, Hamza Golubovic, Han Tong, Arian Maleki, Andrew Ilyas cross-listed No summary available — see the abstract on arXiv.

Mirror, Mirror on the Wall: Prompt Echoing in Small Instruct Language Models

Inez Okulska, Bartosz Naskr\k{e}cki, Jan Piotrowski, Tomasz Steifer No summary available — see the abstract on arXiv.

Personalizing Personal Health Interfaces: Co-Design with Generative AI

Karthik S. Bhat, Vidhi Shah, Vedika Agnihotri, Dong Whi Yoo, Koustuv Saha cross-listed No summary available — see the abstract on arXiv.

Structured Features Overfit Where Random Features Grok

Chon-Fai Kam, Miloud Bessafi, Frederic Cadet No summary available — see the abstract on arXiv.

Zero-SNR Analyticity of the Scalar MMSE Is Equivalent to Gaussianity

Yixing Zhang cross-listed No summary available — see the abstract on arXiv.

Ensemble Complexity in Photovoltaic Forecasting

Sun Ze, Zhou Liguo, Xu Yuqing, Yu Lei, Jiang Mingming No summary available — see the abstract on arXiv.

Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training

Yuanhao Yue, Qianli Ma, Chengyu Wang, Haoting Wang, Lei Shen, Jun Huang No summary available — see the abstract on arXiv.

BusMA: A Bus Communication Substrate for Multi-Agent Systems

Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras No summary available — see the abstract on arXiv.

Bridging the Gap in ECG-Based Emotion Recognition: A Unified Evaluation of Deep Learning Models

Timothy C Sweeney-Fanelli, Ajan Ahmed, Masudul Imtiaz cross-listed No summary available — see the abstract on arXiv.

What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

Lingheng Du, Yiming Tang, Xufeng Duan, Dianbo Liu No summary available — see the abstract on arXiv.

Sensory Precision Inference for Multimodal Arbitration under Uncertainty

Tin Mi\v{s}i\'c, Takato Horii No summary available — see the abstract on arXiv.

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Zixiang Chen, Sufeng Niu, Yingchi Liu, Wenting Zhao, Akshara Prabhakar, Shubham Mehrotra et al. No summary available — see the abstract on arXiv.

Rethinking Procedural Audio Pre-training: Source Scaling and Objective Adaptation

Jiajun Peng, Fengrui Liu, Xinyu Liu, Feng Liu cross-listed No summary available — see the abstract on arXiv.

DA-DLM: Explicitly Modeling Token Dependencies in Diffusion Language Models

Pengyu Ji, Zichen Zhang, Xiang Hu, Kewei Tu No summary available — see the abstract on arXiv.

Branched Optimal Transport Amortization

Semyon Semenov, Viktor Kovalchuk, Meir Roketlishvili, Albert Baichorov, Fakhri Karray, Martin Takac et al. No summary available — see the abstract on arXiv.

Ensemble-Conditioned Molecular Design

Ross Irwin, Alessandro Tibo, Jon Paul Janet, Simon Olsson No summary available — see the abstract on arXiv.

Enabling Creative Exploration for Vibe Design Agents

Yifan Zhang, Nghi D. Q. Bui, Georgios Evangelopoulos, Arnaud Benard No summary available — see the abstract on arXiv.

Translating the Translator: Decomposing the Cost of English-Forced Inter-Agent Communication

Kushagra Agrawal, Yuming Feng, Man-Fai Leung No summary available — see the abstract on arXiv.

Task-Distribution-Aware Counterweight Synthesis and Constrained Co-Design for Serial Manipulators

Mohammad Abbadi cross-listed No summary available — see the abstract on arXiv.

$\mathbb{SL}(n)$ Representation Learning: An Intrinsic Mixed-Curvature Space with Higher Curvature Capacities and Deeper Order-Aware Composition

Xingrun Li, Yusuke Mukuta, Xin Yang, Yinyu Ye, Tatsuya Harada No summary available — see the abstract on arXiv.

Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context

Peng Chen, Zhihao Zhuang, Hongzhou Chen, Junhao Huang, Aiping Yang, Mengsen Wu et al. No summary available — see the abstract on arXiv.

ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models

Hongyu Jin, Wenda Zhang, Runqiu Fei, Gongping Huang, Mike Conway, Ting Dang No summary available — see the abstract on arXiv.

Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation

Yi Chen, Rufeng Cheng, Qiang Xie, Tao Li cross-listed No summary available — see the abstract on arXiv.

OpenAl4S: Code as Action, Science as Sessions

Gongbo Zhang, Hao Li, Yu Wang, Mujie Lin, Liuzhenghao Lv, Yicheng Mao et al. No summary available — see the abstract on arXiv.

ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation

Dennis Ng, Xingyu Shen, Ankit Raj, Kidus Zewde, Tommy Duong, Yuchen Zhou et al. cross-listed No summary available — see the abstract on arXiv.

Physics Informed Neural Network model for the dynamical study of Abdominal Aortic Aneurysm

Adri\'an Robles Arques, Mart\'in Ruiz Fernandez, Javier Sanchis, Miguel A. Teruel, Juan Trujillo cross-listed No summary available — see the abstract on arXiv.

When the Wrong Key Wins: Understanding and Detecting Hallucinations in LLMs

Xuhan Tong, Jiawei Zhang No summary available — see the abstract on arXiv.

Eigenvalue-Decomposition Cost Denoising as an Alternative to Predict-then-Optimize for Shortest-Path Problems

Henry Aldridge-Krawciw, Irene Aldridge cross-listed No summary available — see the abstract on arXiv.

Legislating World-Model-Based Planning with Legal Reasoning

Dylan Waldner, Yiannis Kantaros, Guido Governatori, Risto Miikkulainen, Amir Banifatemi cross-listed No summary available — see the abstract on arXiv.

DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

Hongye Yang, Zhihao Xie, Shengjun Xiong, Boxiao Huang cross-listed No summary available — see the abstract on arXiv.

Refinement-based Flow Policy Optimization

Bumgeun Park, Hyukjun Yang, Donghwan Lee No summary available — see the abstract on arXiv.

MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

Muchen Li, Leonid Sigal, Renjie Liao No summary available — see the abstract on arXiv.

Omni-Streaming Thinking

Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang et al. No summary available — see the abstract on arXiv.

Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes

Pallika Kafle, Yipeng Zhou, Guanfeng Liu, Quan Z. Sheng, Cheng-Hsin Hsu No summary available — see the abstract on arXiv.

woma: a real-time foundation model and its fine-tuned models for endoscopy

Thang Tran, Lan Dang cross-listed No summary available — see the abstract on arXiv.

AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference

Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang cross-listed No summary available — see the abstract on arXiv.

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Yunhao Feng, Ruixiao Lin, Ming Wen, Yanming Guo, Xingjun Ma, Yutao Wu et al. No summary available — see the abstract on arXiv.

SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering

Davit Soselia, Joseph JaJa, Amitabh Varshney cross-listed No summary available — see the abstract on arXiv.

Improving Mathematical Reasoning Capabilities in Large Language Models via Reasoning Process Error Classification

Runa Yoshida, Kosuke Nishida, Kyosuke Nishida No summary available — see the abstract on arXiv.

Multi-source Transfer Learning of Time Series with a Shapelet-based Distance Measure

Jiseok Lee, Brian Kenji Iwana No summary available — see the abstract on arXiv.

PACE: Progressive Angular-to-Norm Contrastive Embedding

Yanping Li, Wei Zhou, Yawen Liu, Yibo Wang, Ke Zhu, Guangda Huzhang et al. cross-listed No summary available — see the abstract on arXiv.

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning With Dynamic Routing

Mingqian Yu, Wenpeng Zhang, Peilin Zhao No summary available — see the abstract on arXiv.

EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse

Dongsheng Shi, Yue Li, Xin Yi, Linlin Wang No summary available — see the abstract on arXiv.

CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search

Sriram Selvam, Anneswa Ghosh No summary available — see the abstract on arXiv.

Nearly Minimax-Optimal Regret for Linear Contextual Bandits with Arbitrary Adaptive Action Sets

Tianyuan Jin No summary available — see the abstract on arXiv.

STHMoE: Hypergraph-Enhanced Heterogeneous Dependency Coordination for LLM-Based Urban Traffic Data Forecasting

Jiawen Chen, Qi Shao, Yongjian Chang, Mingtong Zhou, Duxin Chen, Wenwu Yu No summary available — see the abstract on arXiv.

Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models

Shijian Xu, Andrea Miele, Metod Jazbec, Volker Roth, Eric Nalisnick, Ilija Bogunovic No summary available — see the abstract on arXiv.

Low-Dimensional Embeddings for Gaussian Kernels on Manifolds

Soumik Dutta, Kunal Dutta cross-listed No summary available — see the abstract on arXiv.

Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models

Mingcheng Zhu, Jinning Liang, Tingting Zhu No summary available — see the abstract on arXiv.

VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries

Wenxin Xu, Jinwei Lu, Hwanhee Kim, Chen Jason Zhang, Xiao-Yong Wei, Haoyang Li et al. No summary available — see the abstract on arXiv.

MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing

Jianxiang Ma, Xiaocui Yang, Daling Wang, Yuesong Hou, Mingfu Zhang, Yichen Gao et al. No summary available — see the abstract on arXiv.

What Limits Us? Analyzing Self-Reported Limitations in NLP Research

Tawan Thaepprasit, Peeranuth Kehasukcharoen, Ding Wang, Remi Denton, Peerapon Vateekul, Piyawat Lertvittayakumjorn No summary available — see the abstract on arXiv.

Convergence rates for generative drifting flows: fixed-scale obstructions and multihead acceleration

Arthur St\'ephanovitch, Eddie Aamari No summary available — see the abstract on arXiv.

Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models

Edirlei Soares de Lima, Marco A. Casanova, Antonio L. Furtado No summary available — see the abstract on arXiv.

Interpreting hierarchical organisation of speaker embeddings

Yanze Xu, Wenwu Wang, Mark D. Plumbley cross-listed No summary available — see the abstract on arXiv.

TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models

Tong Li, Shuye Ding, Jiachuan Wang, Yongqi Zhang, Shuangyin Li, Lei Chen et al. cross-listed No summary available — see the abstract on arXiv.

Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election

Bastiaan Bruinsma, Annika Fred\'en, Paul R\"ottger, Moa Johansson, Asad Sayeed No summary available — see the abstract on arXiv.

Failure-Guided Co-Evolution of Prompts and Training Data

Tianyu Yuan, Zhuzhong Qian cross-listed No summary available — see the abstract on arXiv.

Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception

Yanfeng Shi, Yan Song, Junhui Li, Tinggan Huang, Wu Guo, Haoyu Song et al. cross-listed No summary available — see the abstract on arXiv.

From Ideas to Actions: A Public-Data Decision-Support Toolchain Across the Venture Lifecycle

Lei Qu (Shanghai Xing Yun Zhi Li AI Institute) No summary available — see the abstract on arXiv.

MAST: Label-Efficient, Robust, and Generalizable Sound Detection for Biodiversity Monitoring via Masked Audio Pretraining and Self-Training

Tianyi Xu, Daniel Pimentel-Alarc\'on, Zuzana Bu\v{r}ivalov\'a, Claudia Sol\'is-Lemus cross-listed No summary available — see the abstract on arXiv.

Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders

Qingtao Xia, Jiahua Bao, Siyao Cheng, Jie Liu cross-listed No summary available — see the abstract on arXiv.

CWM: Controllable White-Box Meta-Prompting for Adaptive Retrieval-Augmented Generation and Reasoning Ability

Keuntae Kim, Eunhye Jeong, Yong Suk Choi No summary available — see the abstract on arXiv.

ProtoGuide: Prototype-Driven Guidance for Class-Conditional Graph Generation

Salvatore Romano, Marco Grassia, Pietro Li\`o, Giuseppe Mangioni No summary available — see the abstract on arXiv.

Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generation in ESG Domain

Motaz Saad, Anna Borrelli, Ivan Gentile, Kianna Kazemi, Francesco Piccialli, Antonella Longo No summary available — see the abstract on arXiv.

Bandits with Probing: Optimal Regret and the Limits of Winner Feedback

Yongjie Guan No summary available — see the abstract on arXiv.

Conformal Individual Treatment Effect Estimation under Networked Interference

Matteo Zecchin, Osvaldo Simeone cross-listed No summary available — see the abstract on arXiv.

BioDCASE: Active Learning for Bioacoustics

Ben McEwen, Rupa Kurinchi-Vendhan, Shiqi Zhang, Lukas Rauch, Marek Herde, Sara Beery No summary available — see the abstract on arXiv.

Improving the Last-Iterate Guarantees of Anytime Algorithms for Stochastic Monotone Variational Inequalities

Jun-Hyun Kim, Ahmet Alacaoglu cross-listed No summary available — see the abstract on arXiv.

Learning CNF Formulas from Uniform Random Solutions: Near-Tight Sample Complexity for Valiant's Algorithm

Weiming Feng, Yixiao Yu, Yiyao Zhang No summary available — see the abstract on arXiv.

Draining Fictitious Knots: Restoring Distance-Awareness Guarantees for High-Dimensional Spline Networks

Masoud Ataei, Mohammad Javad Khojasteh, Vikas Dhiman No summary available — see the abstract on arXiv.

Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs)

Christian Fisch, Angela Altmeier, Martin Obschonka, Michal Kosinski, Pin Ni No summary available — see the abstract on arXiv.

Impute-EM: Native Mixed-State Diffusion Models for Heterogeneous Data Imputation

Sergei Kholkin, Kirill Sokolov, Dmitry Baranchuk, Evgeny Burnaev, Alexander Korotin No summary available — see the abstract on arXiv.

Math for AI safety: an invitation for mathematicians

Lionel Levine cross-listed No summary available — see the abstract on arXiv.

ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment

Junkai Tong, Mingjia Li, Haoran Chen, Yaoyu Jiang, Hanjie Ge, Yixuan Wang et al. No summary available — see the abstract on arXiv.

Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

Yuhang Wang No summary available — see the abstract on arXiv.

Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings

Mingzhou Jiang, Peixi Wu, Hang Cheng, Yunhao Zhou, Biao Yang, Wei Yuan et al. No summary available — see the abstract on arXiv.

Admissable: Training Reinforcement Learning Agents against Adversarial Missingness

Paul Stahlhofen, Luca Hermes, Tim Kochs, Markus Vieth, Barbara Hammer No summary available — see the abstract on arXiv.

When Correlations Mislead: Confounder-Aware Multi-View Urban Region Representation Learning

Sean Bin Yang, Ying Sun, Zongyi Xu, Tung Kieu, Jilin Hu, Bin Yang et al. No summary available — see the abstract on arXiv.

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar et al. No summary available — see the abstract on arXiv.

Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation

Daxin Tan, Dehua Tao, Chengxi Deng, Hanlin Zhang, Xiao Chen cross-listed No summary available — see the abstract on arXiv.

The Universe of Universes: Benefit Yield Functions, Implosion Thresholds, and Infrastructure-Aware Optimization in Multi-LLM Systems

Danielle Franklin, Vasu Raj Jain No summary available — see the abstract on arXiv.

Evaluation Metrics for Safe Reinforcement Learning

Lindsay Spoor, Aske Plaat, Thomas Moerland No summary available — see the abstract on arXiv.

Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA

Luis M. S\'anchez cross-listed No summary available — see the abstract on arXiv.

Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs

Changxin Lu, Xiaoliang Meng, Yu Wu, Rui Huang, Honglin Li, Tao Chen et al. cross-listed No summary available — see the abstract on arXiv.

Concept-Grounded Reasoning with Prompt-Driven Localization for Interpretable Structured Report Generation

Xinyue Xu, Hongbin Lin, Juangui Xu, Hualiang Wang, Lehan Wang, Lijie Hu et al. cross-listed No summary available — see the abstract on arXiv.

Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models

Peipei Li, Dongsen Zhang, Yuchen Liu, Wenjun Xu No summary available — see the abstract on arXiv.

Parameter-Efficient Adaptation of Pretrained Language Models for Time-Series Forecasting

Tamanna Kumavat, Georg Brunner, Kyriakos Flouris No summary available — see the abstract on arXiv.

End-to-End Cell Detection via Instance-aware Graph Modeling

Ruochen Liu, Yalin Zheng, Jingxin Liu, Jianfeng Zhang, Shoujun Huang, Dexing Kong et al. cross-listed No summary available — see the abstract on arXiv.

ReLU Neural Network Approximation to Smooth Functional Operator: Dimensional Decay and Error Analysis

Shuhao Jiao cross-listed No summary available — see the abstract on arXiv.

MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving

Tiancheng Zhang, Yulin Chen, Yunfeng Zhao, Shaoyuan Huang, Cheng Zhang, Xiaofei Wang No summary available — see the abstract on arXiv.

Robust and Efficient Communication for Multi-Agent Learning

Rafael Pina, Varuna De Silva, Corentin Artaud No summary available — see the abstract on arXiv.

RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

Sibo Zhu, Shicheng Fan, Xinyue Wang, Wenyi Wu, Kun Zhou, Biwei Huang No summary available — see the abstract on arXiv.

SlopShape: Identifying AI-Generated Commercial Web Content

Jochen Madler (Sitefire) No summary available — see the abstract on arXiv.

Representing Clinical Conditions on Vital Signs from Healthy Individuals using Latent Modeling

Rafael Pina, Varuna De Silva, Mindula Illeperuma No summary available — see the abstract on arXiv.

Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs

Mark Russinovich, Blake Bullwinkel, Giorgio Severi, Cristian Ovadiuc, Ahmed Salem cross-listed No summary available — see the abstract on arXiv.

IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective

Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen et al. cross-listed No summary available — see the abstract on arXiv.

A Game-Theoretic Framework for Incentive-Compatible AI training Under Renewable-Energy Constraints

Konstantinos Varsos, Ramin Khalili, Adamantia Stamou, George D. Stamoulis, Vasillios A. Siris cross-listed No summary available — see the abstract on arXiv.

CodeTS: Verifiable Text-to-Time Series Generation via Executable Code

Xudong Yuan, Shunyu Liu, Tongya Zheng, Huiping Zhuang, Mingli Song, Kaixuan Chen No summary available — see the abstract on arXiv.

SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution

Haoxiang Kang, Ming Wen No summary available — see the abstract on arXiv.

When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary

Artem Trofimov, Boris Novikov No summary available — see the abstract on arXiv.

Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning

Xun Xu, Zaixi Zhang No summary available — see the abstract on arXiv.

Can AI systems have free will?

Christian List No summary available — see the abstract on arXiv.

MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

Hongchang Shi, Jinpeng Hu, Ao Wang, Wenzheng Zhou, Hui Ma, Feng Li et al. cross-listed No summary available — see the abstract on arXiv.

Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents

Halil Burak Noyan No summary available — see the abstract on arXiv.

A Conservative OCR-Enabled Workflow for R214 Sodium Screening of South African Packaged Foods

Mayimunah Nagayi, Alice Scaria Khan, Tamryn Frank, Rina Swart, Clement Nyirenda cross-listed No summary available — see the abstract on arXiv.

Single-condition neural solvers encode transferable response spaces for parametric differential equations

Wenbo Cao, Weiwei Zhang No summary available — see the abstract on arXiv.

On the role of the tokenizer in ECG transformer models

Jiawei Li, Fabio Bonassi, Johan Sundstr\"om, Thomas B. Sch\"on, Ant\^onio H. Ribeiro No summary available — see the abstract on arXiv.

Graph Matching Relaxations and Amortization for Supervised Graph Prediction

Federico M\'endez, Paul Krzakala, Gabriel Melo, Charlotte Laclau, R\'emi Flamary, Florence d'Alch\'e-Buc cross-listed No summary available — see the abstract on arXiv.

Online local learning for generative thermodynamic computing

Huilin Wang, Weibing Deng cross-listed No summary available — see the abstract on arXiv.

Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation

M. Ali Bayram No summary available — see the abstract on arXiv.

HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments

Quoc-Vinh Lai-Dang, Hyo-Sang Shin No summary available — see the abstract on arXiv.

Spook the Machine: Gamified Exploration of Human Imagination of Machine Fear

Levin Brinkmann, Hiromu Yakura, Sonia Nicoletti, Mar Canet Sola, Thomas F. Eisenmann, Ali Dasmeh et al. cross-listed No summary available — see the abstract on arXiv.

Temperature Fragility and the Conditional Benefits of Truncation Sampling

Francesco La Rosa No summary available — see the abstract on arXiv.

GSLAD: Prototype-Regularized Graph Structure Learning for Multivariate Time Series Anomaly Detection

Zepeng Zhang, Fuad Khuri, Keivan Faghih Niresi, Olga Fink No summary available — see the abstract on arXiv.

Data-driven Prediction of Satellite-observed Avalanche Activity from Snowpack Simulations

Jakob Grah, Filippo Maria Bianchi, Bert Kruyt, Karsten M\"uller No summary available — see the abstract on arXiv.

Rotation-Based Subspace Tracking for Robust Kernel PCA on Streaming Data

Kris Lokere, John Fossaceca No summary available — see the abstract on arXiv.

The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow?

Ivy Zhang No summary available — see the abstract on arXiv.

Same path, different: a mechanistic comparison of looped and stacked transformer encoders on 12-lead ECG

Pawel Olszowiec, Michal Byra, Grzegorz Gruszczynski, Grzegorz Stefanski, Alberto Presta No summary available — see the abstract on arXiv.

How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

Ilya Koziev, Leonid Sinev, Ivan Oseledets No summary available — see the abstract on arXiv.

Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku

Livia Oddi, Simone Scardapane, Toru Sugimoto, Donatella Genovese No summary available — see the abstract on arXiv.

Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models

JungMin Yun, Junehyoung Kwon, Hayeong Ryu, Byeonggeuk Lim, Hoejoon Kwon, YoungBin Kim No summary available — see the abstract on arXiv.

End-to-End Verifiable and Robust Federated Learning

Doryan Lesaignoux, Enrique M\'armol Campos, Gabriele Spini, Jos\'e L. Hern\'andez-Ramos, Stephan Krenn No summary available — see the abstract on arXiv.

Psychosis involves a deficit of information compression in connected speech

Samuele Vallisa, Claudio Palominos, Rui He, Emre Bora, Burcu Verim, Cemal Demirlek et al. No summary available — see the abstract on arXiv.

Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting

Oliver Stevanovic, Jasmin Wachter cross-listed No summary available — see the abstract on arXiv.

Beyond Noise: Understanding and Overcoming Temperature Effects in Analog DNN Inference

Niklas Summ, Xiao Wang, Hendrik Borras, Bernhard Klein, Holger Fr\"oning No summary available — see the abstract on arXiv.

To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs

Franck Signe, Hippolyte Pilchen, Fran\c{c}ois Yvon, \'Edouard Grave No summary available — see the abstract on arXiv.

Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA

Tristan Kirscher (ICube, Institut Strauss), Niklas C. Koser (CAU), Soren Pirk (CAU) No summary available — see the abstract on arXiv.

The Misery of Mechanistic Interpretability: A Formal Perspective

Tobias Ladner, Matthias Althoff No summary available — see the abstract on arXiv.

Strong and Compact Policies for Submodular Markov Decision Processes via LP-Based Submodular Orienteering

Lars Rohwedder, Rico Zenklusen cross-listed No summary available — see the abstract on arXiv.

Specifying Reward Functions for RL Without Environment Sampling

Stephane Hatgis-Kessell, W. Bradley Knox, Emma Brunskill No summary available — see the abstract on arXiv.

The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits

Ke Cheng, Xin Xu, Yixiao Chen, Lei Xin, Jianbo Zhao, Fanhu Zeng et al. No summary available — see the abstract on arXiv.

Bayesian Optimisation Using Product-of-Experts Gaussian Process Models with Uncertainty Calibration

Yean Hoon Ong No summary available — see the abstract on arXiv.

Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction

Hayeong Ryu, Sunhee Jo, Seunguk Yu, YoungBin Kim No summary available — see the abstract on arXiv.

Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

Sarra Gharsallah, Adele Robaldo, Mariia Tokareva, Giovanni Gatti Pinheiro, Ilyana Guendouz, Rapha\"el Troncy et al. No summary available — see the abstract on arXiv.

PIVOT: Physics-Grounded Verification for AI-Generated Audio-Video Detection

Bo Zheng, Kangran Zhao, Xiaoyu Zhang, Weinan Guan, Zhiheng Li, Yize Chen et al. cross-listed No summary available — see the abstract on arXiv.

Big Brains and Changing Environments: Cause or Consequence?

Sian Heesom-Green, Jonathan Shock, Geoff Nitschke cross-listed No summary available — see the abstract on arXiv.

GRIN+: Towards Fast Yet Effective Machine Unlearning for Imbalanced Medical Data

Minghui Huang, Junxiao Wang No summary available — see the abstract on arXiv.

Self-Evolving Memory for Generative Recommendation

Xinyu Lin, Zhuosong Jiang, Zixiao Suo, Siqin Wang, Hanqing Zeng, Hanchao Yu et al. cross-listed No summary available — see the abstract on arXiv.

A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation

Yang Xing, Jiong Wu, Savas Ozdemir, Yang Zhou, Boxiao Yu, Ying Zhang et al. cross-listed No summary available — see the abstract on arXiv.

VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding

Weixin Xu, Zhenyu Yang, Bing Wang, Shengsheng Qian, Changsheng Xu cross-listed No summary available — see the abstract on arXiv.

Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection

Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truic\u{a}, Elena-Simona Apostol No summary available — see the abstract on arXiv.

Diversified and Perceptible Counterfactual Examples Leveraging Expert Knowledge

Akram Bensalem (IMT Atlantique - INFO), Fahima Djelil (Lab-STICC\_MOTEL, IMT Atlantique - INFO), Marie-Jeanne Lesot (IMT Atlantique - INFO, Lab-STICC, Lab-STICC\_MOTEL) et al. No summary available — see the abstract on arXiv.

Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints

Gwang-Hyeon Yun, Jong-Hoon Park, Bing Hu, Helen Chen, Anita Layton, Young-Rae Cho No summary available — see the abstract on arXiv.

IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English

Ana-Maria Luisa Mocanu, Sebastian Mocanu, Ciprian-Octavian Truic\u{a}, Elena-Simona Apostol cross-listed No summary available — see the abstract on arXiv.

Where to Compute and How to Interact: Operator-Readable Adaptation with Gauge-Aware Transport

Zixuan Shen, Quanxu Wan, Bingchuan Wang, Zhi Wang, Biao Luo No summary available — see the abstract on arXiv.

Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

Daniele Veri' cross-listed No summary available — see the abstract on arXiv.

FedLTLib: A Comprehensive Benchmark for Federated Long-Tail Learning

Changkun Lin, Junxiao Wang No summary available — see the abstract on arXiv.

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach et al. cross-listed No summary available — see the abstract on arXiv.

Potential of Artificial Intelligence Algorithms for Identification of Relevant Diagnostic and Prognostic Biomarkers of Early-Stage Liver Cancer

Ali Bou Nassif, Darko Castven, Manar Abu Talib, Jibran Sualeh Muhammad, Ahmed Ammar Kubba, Jens Marquardt et al. No summary available — see the abstract on arXiv.

Human-Grounded Calibration for Long-Text Image-Text Congruence in Vision-Language Models

Alessandro Gambetti, Qiwei Han cross-listed No summary available — see the abstract on arXiv.

Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching

Jiayang Gu, Zheng Fang, Lichaun Xiang, Fanghui Liu, Xu Cai, Hongkai Wen No summary available — see the abstract on arXiv.

Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs

JuHeon Ha, Byounghan Lee, Yunseo Choi, Kyung-Ah Sohn No summary available — see the abstract on arXiv.

Predictive Likelihood Ratios for Language Model Watermark Detection

Li Ma cross-listed No summary available — see the abstract on arXiv.

Kaininja: Extending Native 3D Generators to the Part Level

Ruihan Yu, Lian Fu, Muyao Niu, Zheng-hui Huang, Yu-Ju Tsai, Sho Kuno et al. cross-listed No summary available — see the abstract on arXiv.

CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense

Guo Fuzheng cross-listed No summary available — see the abstract on arXiv.

CIDERS: Cloud-Edge LLM Collaborative Learning via Accelerating Personalized Bilevel Optimization

Victor H. Chen, Hairui Yu, Stella K. Chung, Hong Yan cross-listed No summary available — see the abstract on arXiv.

Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding

Jinyuan Deng, Yuqi Jiang, Wenjing Huang, Xin Li, Qi Sun, Cheng Zhuo cross-listed No summary available — see the abstract on arXiv.

Benchmarking Intra-Patient 3D Deformable Multimodal Image Registration

Matteo Barbieri, Giammarco La Barbera, Juan Pablo De La Plata, Sabine Sarnacki, Isabelle Bloch, Pietro Gori cross-listed No summary available — see the abstract on arXiv.

Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models

Md Khalid Syfullah, Alvi Ataur Khalil cross-listed No summary available — see the abstract on arXiv.

Projection-Free Multi-level Algorithms for Stochastic Constrained Compositional Optimization

Wei Jiang, Sifan Yang, Wenhao Yang, Yibo Wang, Yuanyu Wan, Zechao Li et al. cross-listed No summary available — see the abstract on arXiv.

Scalability and Performance Evaluation of Federated Learning Frameworks: A Comparative Analysis

Bassel Soudan, Sohail Abbas, Ahmed Kubba, Manar Wasif Abu Talib, Qassim Nasir cross-listed No summary available — see the abstract on arXiv.

RESKILL: Explicit Failure Attribution and Structured Repair for Interactive Language Agents

Mengyi Deng, Xin Li, Duyi Pan, Zilin Wang, Zhiwei Li, Zhijiang Guo et al. No summary available — see the abstract on arXiv.

EEG-Xplain: Decoding Neural Black-Boxes of EEG Foundation Models

Hansong Ma, Junxiao Wang No summary available — see the abstract on arXiv.

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

Haonan Jiang, Guojian Zhan, Jiancong Xie, Shijun Wan, Dongiia Zhao, Cheng Chen et al. No summary available — see the abstract on arXiv.

More Than Just Access: Generative AI as Communication Intermediary for Blind and Low-Vision Users

Protik Dey, Mohd Saifuzzaman, Taslima Akter cross-listed No summary available — see the abstract on arXiv.

Backward SDEs-based Diffusion for Physics-Constrained Generation

Zihao Wang No summary available — see the abstract on arXiv.

Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction

Kushal Patel, Pushkal Shrivastava, Mackenzie Lees, Qirui Lu, Bhargobjyoti Saikia, Liying Li et al. No summary available — see the abstract on arXiv.

New Conditions for Philosophers to Catch the Wave of Citizen Deliberation in the Age of Artificial Intelligence in advance

Bernard Reber (CEVIPOF) No summary available — see the abstract on arXiv.

Predicting build orientation for SLM dental parts: a comparison of rotation representations and direct vector regression

Felix Schmalzel, Reimar Waitz, Moritz Kronberger, Thorsten Sch\"oler No summary available — see the abstract on arXiv.

Knowledge-Enriched Structured EHR Features for 30-Day Hospital Readmission Prediction on MIMIC-IV

Mohamad Najafi, Hongyun Fu, Mathias Brochhausen, Jian Wu, Yaohang Li No summary available — see the abstract on arXiv.

Assembling the CREW: A Collaborative Multi-agent Reinforcement Learning Framework for Automated Related Work Generation

Hai-Dang Dang, Bao-Yen Pham, Bao Nguyen, Tran Thi Huong, Huynh Thi Thanh Binh No summary available — see the abstract on arXiv.

Data storytelling meets interpretable machine learning: Decoding AI decisions for non-experts without revealing sensitive data and model details

Lemen Chao, Zixuan Yang, Anran Fang, Mingran Sun, Ming Lei No summary available — see the abstract on arXiv.

Solving Finite-sum Coupled Compositional Optimization via Multi-block-Single-probe Estimator

Wei Jiang, Sifan Yang, Yibo Wang, Lijun Zhang, Zechao Li No summary available — see the abstract on arXiv.

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

Zhenjie Yang, Yideng Zhang, Dongjie Zhang, Chenyu Jiang, Xianshuai Liu, Yufeng Li et al. cross-listed No summary available — see the abstract on arXiv.

Are LLMs Good Financial User Simulators? A Preliminary Study

Jiajie He, Jiangyuan Hong, Dongling Ni, Wenjin Liu, Xintong Chen No summary available — see the abstract on arXiv.

A Language-Guided Multimodal Foundation Model for Zero-Shot and Multi-Task Brain Signal Analysis

Mingzhi Chen, Yiyu Gui, Guibo Luo, Yuchao Yang No summary available — see the abstract on arXiv.

Merging the Knowledge of LLMs for Automatic Speech Recognition

Hayato Futami, Tatsuya Kawahara No summary available — see the abstract on arXiv.

Design of a Deep Learning Credit Risk Early Warning System Integrating Multi-source Heterogeneous Data

LiYang Wang (Washington University in St. Louis), Zhen Zhong (Georgetown University), Zhen Tian (University of Glasgow), Keyu Chen (Wuyi University), Keyu Chen (Wuyi University) No summary available — see the abstract on arXiv.

Look Before You Leap: Factual Decoding with Internal Attribution Signals

Hayeong Ryu, JungMin Yun, Byeonggeuk Lim, Sunhee Jo, YoungBin Kim No summary available — see the abstract on arXiv.

Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR

Thai Thi Thanh Thao Dang, Mengjie Qian, Kate Knill No summary available — see the abstract on arXiv.

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

Ke Hu, Nourchene Ferchichi, Edresson Casanova, Ankita Pasad, Elena Rastorgueva, Chen Chen et al. No summary available — see the abstract on arXiv.

Sylvas: Synergistic Learning Value based Device Scheduling in Federated Continual Learning

Yuxuan Sun, Yuxuan Bai, Tan Chen, Sheng Zhou, Zhisheng Niu No summary available — see the abstract on arXiv.

Event-Native Symbolic-Temporal Spike Encoding Framework for Heterogeneous Cyber Streams

Dalton Diez, Peyton Andras, Max Shroyer, James Ghawaly Jr cross-listed No summary available — see the abstract on arXiv.

Transfer Learning for Socioeconomic Estimation in Forced-Displacement Settings

Steven Ndung'u, Adel Daoud, Ismael Yacoubou Djima, Hai-Anh H. Dang, Patrick Michael Brock No summary available — see the abstract on arXiv.

EvoOntology: A Self-Evolving Ontology Layer for Data Agents

Meiduo Chong, Shaolei Zhang, Ju Fan, Xiaoyong Du No summary available — see the abstract on arXiv.

MoveBench: A Benchmark for Global-Scale Wildlife Movement Forecasting

Justin Kay, Shir Bar, Ellen O. Aikens, Martin Becker, Francesca Cagnacci, Juliet Cohen et al. No summary available — see the abstract on arXiv.

When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control

Roberto Ria\~no, Gorka Abad, Stjepan Picek, Aitor Urbieta cross-listed No summary available — see the abstract on arXiv.

Learning under Target Shift: Optimal Density Ratio Estimation and Importance-Weighted Regression

Ren-Rui Liu, Zheng-Chu Guo cross-listed No summary available — see the abstract on arXiv.

KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI

Jocelyn Kang, Caroline Zhang No summary available — see the abstract on arXiv.

Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation

Yucheng Shen, Lingyong Yan, Jiulong Wu, Shuaiqiang Wang, Jianmin WU, Dawei Yin et al. No summary available — see the abstract on arXiv.

When Should a World Model Move? Loss-Conditioned State Execution

Jintao Xu, Zhengyu Chen, Ben Zhang, Yongzhi Qi, Jianshen Zhang No summary available — see the abstract on arXiv.

Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control

Natalie Collina, Surbhi Goel, Aaron Roth, Sikata Bela Sengupta cross-listed No summary available — see the abstract on arXiv.

Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking

Md Arafat Hossain, Thomas Randall, Akash Dutta, Xingfu Wu, Rong Ge, Ali Jannesari cross-listed No summary available — see the abstract on arXiv.

Atria Dawn: The Dawn of Agentic Superintelligence

Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu et al. No summary available — see the abstract on arXiv.

AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery

Junhao Qiu, Qinglong Hu, Xialiang Tong, Mingxuan Yuan, Liyong Lin, Qingfu Zhang No summary available — see the abstract on arXiv.

Sharp Rates and a One-Line Correction for Spectral Representation Learning

Dier Tang, Jing Yee Tan, Guangyue Han No summary available — see the abstract on arXiv.

CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering

Sumit Barua, Guan Hong, Halil Dursunoglu, Charles Rodgers, Alvis Fong No summary available — see the abstract on arXiv.

Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

Huicheng Zhang, Xiyao Feng, Ze-Tong Li, Chengkai Zhu, Xiao Shi, Xiwei Pan et al. No summary available — see the abstract on arXiv.

Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence

Thai T. Vu, John Le, Tu N. Nguyen, Jun Shen, Quang Vinh Duong, Ha Nguyen cross-listed No summary available — see the abstract on arXiv.

Before You Poll with LLMs: A Deliberative Diagnostic Framework

Ahmed Wali, Hassaan Tayyab No summary available — see the abstract on arXiv.

Learning to Coach for Experiential Learning

Guanheng Chen, Tianzhu Ye, Li Dong, Xun Wu, Shaohan Huang, Furu Wei No summary available — see the abstract on arXiv.

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj et al. No summary available — see the abstract on arXiv.

Task-Directed Residual AddUNet:Perfect-Reconstruction Routing for Full-Rate Representations

Vikram R. Lakkavalli No summary available — see the abstract on arXiv.

LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction

Siyao Wang, Florian Guitton, Shuojie Fu, Guanyu Tao, Kai Sun, Wenjia Bai No summary available — see the abstract on arXiv.

LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys

Md Khalid Syfullah, Alvi Ataur Khalil No summary available — see the abstract on arXiv.

Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

Jaehun Shon, Jinha Choi, Jongwook Jeon, Jongmin Lee No summary available — see the abstract on arXiv.

Thin-shell stability of Gaussian cooling: logconcave sampling with sesteric complexity from a cold start

Yunbum Kook, Santosh S. Vempala cross-listed No summary available — see the abstract on arXiv.

Privacy-enhanced federated learning via asynchronous aggregation and local differential perturbation

Zhen Zhong (Georgetown University, Washington, D.C., USA), Shini Yang (LinkedIn, CA et al. No summary available — see the abstract on arXiv.

Inoculation Midtraining with Learned Neologisms

Kyle O'Brien, Edward James Young, Puria Radmard, Nathalie Kirch, Cameron Tice, Tomek Korbak et al. No summary available — see the abstract on arXiv.

Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI

Paul-Gabriel Nicolae, Irina Georgiana Mocanu cross-listed No summary available — see the abstract on arXiv.

Quenched Ensemble Sampling

David Yallup cross-listed No summary available — see the abstract on arXiv.

Bridging Control, Inference, Transport, and Thermodynamics: From Theory to Applications in Learning

Emmy Blumenthal, Nikolas Claussen, Benjamin Eysenbach, Catherine Ji, Gautam Reddy, Colin Scheibner et al. cross-listed No summary available — see the abstract on arXiv.

Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning

Sophia Tang, Shiyi Wang No summary available — see the abstract on arXiv.

SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection

Tong Jian, Aditya Thurvas Senthil Kumar, Xinyi Li, Ziling Chen, Tianyu Dai, Ali Sengul et al. cross-listed No summary available — see the abstract on arXiv.

Safe Meta-Reinforcement Learning via Information Space Reachability

Zeyang Li, Sunbochen Tang, Navid Azizan No summary available — see the abstract on arXiv.

Pilot Early, Commit Late: A Real-Options Model of Enterprise AI Adoption under Rapid Technological Progress

Gaurav Tewari No summary available — see the abstract on arXiv.

Recurrent GraphNeural NetworkswithSet-BasedAggregation

Blai Bonet No summary available — see the abstract on arXiv.

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao et al. No summary available — see the abstract on arXiv.

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

Aman Priyanshu, Supriti Vijay, Kimia Majd, Xuhong He, Fraser Burch, Takahiro Matsumoto et al. cross-listed No summary available — see the abstract on arXiv.

Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication

Yilin Xu, Chun Hei Michael Shiu, Chih Wei Ling, Linqi Song No summary available — see the abstract on arXiv.

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering

Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah, Michael Oberst No summary available — see the abstract on arXiv.

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye et al. No summary available — see the abstract on arXiv.

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

Ling Yang, Zhenfei Yin, Yingcheng Wu No summary available — see the abstract on arXiv.

Disentangling Representation Evolution in Transformers through Directional Decomposition

Shwai He, Haichao Zhang, Shen Yan No summary available — see the abstract on arXiv.

A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models

Xingyun Wang, Haomin Zheng, Man Yuan, Leqian Yang, Ziming Liu No summary available — see the abstract on arXiv.

The Router Within: Eliciting Native Skill Routing from a Frozen LLM

Ruishuo Chen, Xun Wang, Yu Chen, Zhuoran Li, Longbo Huang No summary available — see the abstract on arXiv.

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao, Song Zuo, Vahab Mirrokni No summary available — see the abstract on arXiv.

Bellman Policy Optimization

Zhuoqing Song, Haotian Xu, Xikun Zhang, Lidong Bing No summary available — see the abstract on arXiv.

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis No summary available — see the abstract on arXiv.

Applications 21

Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB

Anas Dorbani, Sunny Yasser, Jimmy Lin, Amine Mhedhbi cross-listed Analytical pipelines that use LLMs over both database tables and unstructured documents usually mean coordinating separate data systems, moving data between them, and handling details like context management by hand. FlockMTL is a database extension, shown in DuckDB, that brings LLM calls and retrieval-augmented generation (RAG) directly into SQL through model-driven scalar and aggregate functions that can be chained over rows. Borrowing from the relational model, it automatically applies cost-based optimizations such as batching and caching. It also adds PROMPT and MODEL as first-class schema objects alongside TABLE, so queries stay independent of specific prompts and model resources.

Natural-Language to SysMLv2 Translation via Conformance-Driven Iterative Refinement

Chance LaVoie, Eladio Andujar Lugo, Taylan G. Topcu, Levent Burak Kara cross-listed Model-Based Systems Engineering (MBSE) increasingly uses the textual SysMLv2 language, but LLM-generated models are only usable if industrial modeling tools accept them, not merely if they parse. The framework places a production SysMLv2 conformance checker inside a generate-check-repair loop and feeds its deterministic diagnostics back to the LLM until no conformance errors remain. Across all 151 SysMBench prompts and four LLM backends, single-shot generation passed conformance 51.16% of the time, while the iterative approach reached 100% conformance.

When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels

Robson Tigre, Hugo Gobato Souto cross-listed LLMs are increasingly used as synthetic consumer survey panels, but aggregate validation metrics hide failures such as compressed variance, flipped coefficient signs, and subgroup errors of 10 to 30 percentage points. The authors split synthetic-panel bias into covariate shift and concept shift and build diagnostics with clear thresholds for deciding whether to trust, correct, or abandon the data. For correction they use a doubly robust augmented inverse probability weighting (AIPW) estimator that needs only 50 to 300 human calibration responses. The decision rule was right in all 180 simulation runs and cut naive bias by 92.9 to 99.6% on the American National Election Study, and on the Twin-2K-500 pricing dataset it correctly trusted full-sample estimates and corrected subgroup estimates, reducing bias by 83 to 94%.

A Lifecycle Cost Analysis of Smart-Contract-Coordinated Federated Learning Marketplaces

Luan Mantegazine, Luiza Leidemer, Claudio Geyer cross-listed Blockchain-based federated learning (FL) marketplaces use smart contracts to coordinate training among parties that do not trust each other, but cost studies usually measure isolated blockchain operations rather than the full marketplace lifecycle. The authors measure the gas used by every operation in a DAO-governed marketplace, run ablations that isolate the effect of on-chain coordination and IPFS storage on training, and build an analytical model of how deployment costs amortize. One training task costs about 3.8 million gas units per hired trainer, and the average cost per round reaches its amortization knee after about 20 communication rounds. Accuracy matches conventional FL, and costs amortize because recurring costs are one to two orders of magnitude smaller than fixed deployment costs, not because on-chain operations are cheap.

Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging

Muskan Walia, Joe Nudell, Alex Chohlas-Wood cross-listed California now requires prosecutors to make race-blind charging decisions from case documents with race-related proxies redacted, and the authors validate bc2, their open-source LLM-based redaction tool used in more than 119,000 real cases in 2025. Using nearly 5,000 police reports from jurisdictions across the United States, they ask two questions: whether bc2 faithfully implements the mandate, and whether the mandate itself achieves race-blind review. Under a strict document-level measure, the latest bc2 complies fully on 96.7% of narratives, beating earlier versions and leading open-source redaction methods. The mandate misses key proxies such as location, and redacting these extra proxies removes 43.1% of the predictive signal that remains after mandated redaction.
16 more specialized papers

Other 10

Generalization Can Emerge in Tabular Foundation Models From a Single Table

Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini et al. cross-listed Tabular foundation models predict labels in context from example rows without updating weights, and they are widely believed to need pre-training on large synthetic priors, as in TabPFN, or on many real datasets, as in TabDPT. The authors challenge this view, showing that simple self-supervised pre-training on a single real table can transfer strongly across heterogeneous benchmarks. By systematically pre-training and evaluating on many diverse datasets, they find that the key factor is the number and quality of tasks that can be constructed from the pre-training data.

From Semantic to Token Communication: The Next Paradigm for Large-Model-Driven 6G Intelligent Connectivity

Yu Ma, Zhen Gao, Li Qiao, Xiaoyuan Zhang, Mahdi Boloursaz Mashhadi, Yin Xu et al. cross-listed Sixth-generation (6G) networks are moving from bit delivery toward meaning-aware communication, but semantic communication (SemCom) driven by large models (LMs) is fragmented: its representations are tied to particular modalities, models, or tasks, with no universal unit like the bit. This survey argues that tokens should serve as that unit. It points out that unified multimodal models already encode text, images, audio, video, and robot actions as tokens, and that distributed inference already generates token-level traffic. The survey reviews LM-driven SemCom and then covers token communication (TokenCom), including its transmission techniques, its use for LM services and for embodied and agentic systems, and open challenges.

The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech

Antonis Asonitis, Juan Pablo Zuluaga Gomez, Francesco Verdini, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet et al. cross-listed Reference-free speech quality predictors such as UTMOS, DNSMOS, and SCOREQ are the standard automatic evaluators for text-to-speech (TTS) and are increasingly used as rewards in preference optimization. The authors tested them pairwise against listener preferences on six human-rated corpora. The predictors agree with listeners when one clip has audible defects, but once both clips are clean no single predictor reliably picks the preferred one, and several do worse than simply choosing the longer clip. A calibrated combination of complementary signals is the best evaluator, even an equal-weighted ensemble works as a post-training reward, and optimizing any single metric leads to reward hacking that makes held-out judges and human listening tests worse.

Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

Dylan Luke Holyoak Multilingual speech recognition models like Whisper can transcribe low-resource languages without language-specific training, but they are expensive to deploy. The authors systematically evaluate token merging, which combines redundant features to shorten sequences at inference time without retraining. Their tests cover sixteen languages and three Whisper model sizes, plus low-resource languages after DoRA fine-tuning. They find that token merging improves computational efficiency with almost no loss in transcription accuracy for most languages and model sizes, and that it keeps working after fine-tuning.

Do Tabular Foundation Models Still Need Feature Engineering?

Yifan WU, Pinjun Dong, Jiran Tao, Binyan Jiang Tabular foundation models (TFMs) are pretrained on many tabular datasets and applied through in-context learning, which raises the question of whether manual feature engineering still helps. A controlled study tests a wide range of feature engineering techniques across several versions of two major TFM families on TabArena benchmark datasets. Gains from feature engineering are concentrated in earlier model generations and become negligible for the strongest models. Adding in-context information from related datasets still improves performance, which suggests that stronger TFMs gain more from extra task-relevant context than from re-representing existing inputs.
5 more specialized papers

Vision 8

Sampling headroom is not selection gain: a compute-value audit of test-time scaling for video world models

Yuhua Jiang, Junjie Lu, Feifei Gao cross-listed Test-time scaling (TTS) helps only if extra samples contain better candidates and the system can reliably pick them out, and video world models can fail at the second step. The Compute-Value Audit (CVA) checks in sequence whether extra sampling creates opportunity, whether observable signals are reliable, whether acting on them helps, and whether the gain exceeds the full cost of generation and verification. On 192 Physics-IQ scenes, growing the pool from 4 to 16 candidates raised oracle quality by 9.23 IQ, but the Flow, Cycle and VideoReward selectors could not reliably capture that gain, and none of twelve adaptive-depth policies beat uniform compute across three generators. Positive results in a sparse PRM800K setting and with a privileged future signal show the pattern is not universal: extra samples pay off only when they lead to a reliable decision whose benefit survives the full compute cost.
7 more specialized papers

Agents 4

Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

Hazel Mak, Susheel Suresh, Sahil Bhatnagar, Barry Wang, Chhaya Methani, Alejandro Gutierrez Munoz cross-listed Shell-based agents do well at coding, but it is unclear whether a general shell also beats specialized tools for enterprise work that spans multiple applications, coordination with coworkers, and professional analysis. The authors compare five tool interfaces on TheAgentCompany and APEX-Agents using Opus-4.8 and GPT-5.5: typed tools, typed tools plus bash, bash alone, bash with persistent agent-synthesized tools, and programmatic tool calling (PTC) restricted to a typed tool catalog. Bash alone beat typed tools by 21.8 to 24.5 percentage points on TheAgentCompany and by 4.8 to 7.4 points on APEX-Agents, while using 19 to 72% fewer tokens, and adding typed tools or tool synthesis to bash brought no detectable gain. PTC used fewer tokens than direct typed calls with similar scores but generally trailed bash alone, so the authors recommend bash when execution can be isolated and PTC when compliance requires a fixed tool catalog.

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents

Aditya Karnam Gururaj Rao, Arjun Jaggi For LLM agents running locally, memory, prefill latency, and cache growth tightly limit how many input tokens each call can afford. BudgetBench is a protocol and open-source harness that makes this per-call token budget the variable under test. It sweeps budgets from 2K to 32K tokens with the model and decoding held fixed, records quality, budget use, latency, and budget-violation rates, and lets different memory strategies be swapped in. Pilot studies with a local qwen2.5:1.5b on SWE-bench Verified and LongBench v2, a hosted Qwen3 30B-A3B replication, and a 500-item LongMemEval study expose budget-compliance failures and non-monotonic quality curves that single-budget evaluation hides, while leaving unresolved whether budgeted memory beats full context.

From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch cross-listed If groups of LLM agents are used to simulate human groups, they should succeed or fail through human-like deliberation, not just produce similar outcomes. The authors compare human group chats with matched LLM deliberation traces on Wason-style deductive reasoning, then test whether the same process patterns hold for analogical, abductive and analytical tasks. Both humans and LLMs show the same assembly bonus asymmetry, where discussion improves the average member more often than the best initial member, but LLM groups follow majorities more often, surface less unique information and converge earlier than humans. Interventions borrowed from human group-decision research gave modest gains but did not remove this coordination bottleneck.
1 more specialized paper

Large Language Models 4

Towards Optimizing SQL Generation via LLM Routing

Mohammadhossein Malekpour, Nour Shaheen, Foutse Khomh, Amine Mhedhbi cross-listed Highly capable large language models (LLMs) write accurate SQL from natural-language questions, but they add latency and cost on simple queries that cheaper models could handle. The authors present what they describe as the first LLM routing approach for Text-to-SQL. For each query, a score-based or a classification-based router picks the cheapest model likely to produce correct SQL, and both routers are designed to be easy to train and fast at inference. On the BIRD dataset, the routers match the accuracy of the most capable LLM while reducing cost, with a practical and explainable trade-off between accuracy and cost.

Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories

Shamin Chokshi Learned prompt compressors such as LLMLingua and Selective Context need auxiliary language models and are non-deterministic. The authors instead test how far a training-free, deterministic, CPU-only pipeline of classical lexical transformations can go. They combine eleven toggleable steps, including stopword removal, filler-phrase deletion, part-of-speech pruning, lemmatization, and WordNet synonym shortening, into fifteen configurations tested on 1,242 prompts from sources such as Dolly-15k, MMLU, and GSM8K, producing 18,630 paired GPT-4o-mini completions. The most aggressive configuration cuts tokens by 40.3% on average at a BERTScore-F1 of 0.876 against outputs from the original prompts, stopword removal alone gives a 29.6% cut at 0.913, and commonsense reasoning consistently breaks under aggressive compression.

LLMs or Naive Bayes? Old Gems or New Ways

Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank W\"urthwein Large language models (LLMs) raise the question of whether classical classifiers like Naive Bayes (NB) should be retired, so the authors benchmark Complement Naive Bayes against zero-shot and few-shot LLMs from four model families, ranging from 27B parameters to a 1T-parameter mixture-of-experts, on text classification. LLMs win only when there is no labeled data, and even that edge is prone to contamination: on a sentiment task with low contamination, NB beats the zero-shot LLM 81.7% to 73.0%. With labels, NB reaches 89.1% on AG News, statistically tied with a zero-shot 27B LLM and ahead of a 397B frontier model, while batched GPU inference with a small LLM is 40-486x slower than NB on a CPU and uses roughly two orders of magnitude more energy per sample. NB matches LLMs at around 10^4 labels for topic classification, and the authors release a Kubernetes Helm operator that automates model selection with configurable thresholds.

Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference

Xu Yang, Jiapeng Zhang, Zhangke, Changjian Chen, Yuxin Chen, Feiqiang Sun et al. cross-listed Sparse long-context inference has to retrieve relevant tokens efficiently during both prefill and decode, but existing methods usually use a different retrieval strategy for each stage, so no single retrieval representation is reused throughout. Self-Indexing Attention is a training-free framework built on a shared transform-domain sign-magnitude representation, where the signs of the keys form a 1-bit token index that serves both grouped prefill selection and decode retrieval using fast bitwise operations. The same representation works with external KV-cache compression without needing separate indexer metadata. At 5% attention density it stays close to dense attention on LongBench and RULER while delivering up to 6.1x prefill and 10.3x decode attention-operator speedups, and experiments with TurboQuant and DeepSeekV4-Flash show it is compatible with low-bit KV-cache compression and pretrained sparse-attention indexers.

Multimodal 4

Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction

Sarthak Sattigeri cross-listed Benchmarks show vision-language models struggle with low-level manipulation reasoning, but an aggregate score does not reveal whether models fail to pick the right object part or fail to know what action that part needs. Eight models from three developers were asked what motion a robot should apply to 19 articulated objects. Under an open prompt they almost never answered 'push' (once in 64 cases where it was correct), because they described a different part, such as picking up a camera instead of pressing its button. Naming the target part raised action accuracy from 0.158–0.474 to 0.684–0.947 for every model, which points to part grounding rather than missing action knowledge as the dominant bottleneck; the authors also document two of their own measurement errors that surfaced only when compared against trivial constant baselines.

Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding

Dilip Sarkar, Md. Safayet Islam, Liang Liang cross-listed Multimodal large language models (MLLMs) cannot process every frame of a long video. Training-free, plug-and-play (PaP) keyframe selectors that work with any MLLM are the cheapest fix, but the five published in the past year were each evaluated on different benchmarks and models. The authors compare all five under one setup, using three MLLMs and three long-video understanding benchmarks. QAaF performed best in 13 of 15 aggregate settings and FOCUS ranked second overall, giving a common reference point for training-free keyframe selection.

(How) Do MLLMs Report Bistable Images Like Humans?

Ryota Takatsuki, Tomoki Doi, Amane Watahiki, Anil K. Seth, Hitomi Yanaka cross-listed Bistable images like the duck-rabbit support two incompatible interpretations that humans report one at a time, and the authors test whether multimodal large language models (MLLMs) behave the same way. Using the LLaVA family on the classic duck-rabbit and on synthetic Visual Anagrams (to limit memorization), they measure whether visual cues and linguistic priors can bias the reported interpretation and whether responses commit to a single one. Both visual and linguistic manipulations shifted reports in human-consistent ways, while responses stayed predominantly exclusive. Mechanistically, the behavior traces to competing image-token representations, separate pathways for bottom-up and top-down influence, and a link between exclusive reporting and object-count encoding.
1 more specialized paper

Theory 4

4 more specialized papers

Safety & Alignment 2

ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models

Manan Tayal, Akshay Nambi cross-listed Safety fine-tuning for Vision-Language-Action (VLA) models usually relies on Lagrangian soft penalties, which can leave residual violations or make policies overly conservative, and visual environments rarely come with dense per-step safety labels. ShieldVLA learns a model-free approximation of the Hamilton-Jacobi (HJ) reachability value function from visual observations and uses it as a safety critic. The critic separates reward maximization inside safe regions from recovery near unsafe states, and it is supervised by rubric-based VLM safety scores instead of manual cost labels. Across five navigation and manipulation benchmarks and multiple VLA backbones, it cuts cumulative safety cost by 57% on average and raises task success by 0.13 over SafeVLA.
1 more specialized paper

Reasoning 1

PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems

Joseph Chan, Utkarsh Jha, Xiyin Yang, Abhinav Jarajapu, Anik Sahai, Eddie Hu et al. Static science benchmarks do not test whether LLMs can reason about physics through active experimentation. PhysMent is a benchmark of 105 classical mechanics scenes in the MuJoCo physics simulator. Models must use tools to apply forces, query object states, advance time, and change scene geometry to find the information they need, and answers are scored along six dimensions. Across seven models accuracy ranges from 25% to 67%, and while models reach up to 80% on qualitative single-concept tasks, most fall below 30% on the hardest single-concept category, where failures come from premature answers, inefficient exploration, and inconsistent use of simulator feedback rather than gaps in physics knowledge.

Reinforcement Learning 1

Diagnosing Faults in Reinforcement Learning Simulators and World Models with Canonical Polynomial Invariants

Tesfay Zemuy Gebrekidan, Hadush Hailu Gebrerufael Much work builds physical structure into learned dynamics on the assumption that it improves prediction, and this assumption is tested using exact polynomial invariants recovered from trajectories and put into canonical form as reduced Gröbner bases over the rationals. On Acrobot, exactness helps prediction little: a consistency regularizer lowers algebraic residuals without improving rollout fidelity, and a shaping potential taken from a system with a 100% mass error speeds up learning as much as the correct one. The invariants are instead useful for diagnosis, through screening and attribution procedures built on new normal-form deflation and quotient-space recovery techniques. Across fifteen injected faults, screening localized every broken constraint with no false alarms while observation-space baselines localized none, attribution found the responsible parameter in all seven parameter faults, and a scan of 350 release pairs across eleven reinforcement learning environments found no evidence of changed simulator dynamics.

Robotics 1

GzDRL: Reproducible and Scalable Deep Reinforcement Learning with Gazebo

Amal Dev Haridevan, Junjie Kang, Jinjun Shan cross-listed Connecting reinforcement learning (RL) to the Gazebo robot simulator through middleware usually introduces nondeterminism and makes experiments hard to reproduce. GzDRL is a single-process, middleware-free framework that steps agent actions and physics updates in lockstep, enabling deterministic, vectorized, high-throughput data collection. Benchmarks show the highest workstation throughput among the evaluated frameworks, competitive performance with GPU-accelerated simulators on laptop hardware, and experiment-level reproducibility with multi-agent scaling. Policies trained with it transferred to a physical quadrotor without fine-tuning.