Tuesday, September 15, 2026
Highlights
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Diffusion language models generate many tokens in parallel, which suits GUI agents that must read a screen and output a precisely placed action at every step. Until now, though, no one had shown that such models can work as full multi-step vision-language agents. LLaDA-UI connects a native-resolution vision encoder to the block-wise diffusion backbone LLaDA2.0-mini-base, giving a 16.7B-parameter mixture-of-experts model, and trains it into an open-source agent for mobile, desktop and web.
- Training has two stages: multimodal pre-training on just 145B tokens with a block-diffusion loss, which already beats earlier diffusion VLMs such as
SDAR-VL(OCRBench855 vs 726,MathVerse48.0 vs 36.6), then full fine-tuning on more than 6M GUI grounding and navigation samples that output a reasoning block followed by an action. - Against the autoregressive
Qwen3-VL-8Bit wins on four of six GUI benchmarks, namelyAndroidWorld(53.5 vs 47.9),MobileWorld(25.6 vs 9.4),WebVoyager(56.9 vs 45.2) and, by a hair,ScreenSpot-Pro(52.9 vs 52.7), but it trails onScreenSpot-V2(90.7 vs 93.0) andOSWorld-Verified(29.4 vs 33.9) and stays well behind specialized agents likeHolo2-8Bon WebVoyager (80.2) andGUI-Owl-7Bon AndroidWorld (66.4). - In a paired test on four H100s, each API call runs 3.6–9.0× faster than
Qwen3-VL-8Bwhen neither model gets a cache hit (for example, 5.9 s vs 53.0 s on MobileWorld), though that is based on only five calls per domain, and Qwen's cached replays return in under 2 s. - Decoding settings strongly affect results: on an intermediate checkpoint, turning off early end-of-sequence stopping raises AndroidWorld success from 42.7% to 52.6%, while switching from 32-token blocks with 32 denoising steps to 64-token blocks with 16 steps drops it to 33.0%.
- Outputs are almost always well-formed (98.65% parse), yet the agent gets stuck in loops on long tasks: repeated identical actions make up 26.7% of steps in failed runs, and success falls from 63% on tasks needing at most five actions to 14% on tasks needing more than 20; it also misses small targets on large
OSWorldscreenshots.
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Omni-Streaming Thinking
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Computer-use agents create safety risks through what they actually do with tools, yet most guard models are trained on static prompts and responses, and executable safety platforms return pass/fail verdicts rather than data a guard can learn from. HazardAuditor runs real agents in controlled environments to get labels based on their actual behavior. It then trains a generative guard with GuardPO, a post-training objective that weights each response by its safety verdict rather than by how long its explanation is.
- Adapters convert traces from
Claude Code,Codex,Hermes, andOpenClawinto one shared format (user messages, agent responses, tool calls with arguments, and what the environment returned), and a trace counts as unsafe only if the agent attempts a harmful action, not just because risky content showed up in its context. GuardPOstarts from a fine-tunedQwen3Guard-Gen-8Band scores each sampled response by rule (+1 correct, −1 wrong, −1.25 malformed), then trains with clipped updates in the style ofCISPOwhile averaging the explanation and verdict losses separately for each response, so long explanations no longer dominate the gradient.- On the authors' balanced
CUA-Exectest set (200 trajectories per framework), the guard reaches 94.0 / 95.5 / 86.5 / 87.5% accuracy on the four agents, beating the strongest prior guard,BraveGuard, by 4.0–16.5 points and also outscoringClaude-Sonnet-4.6used as a judge on every framework. - Results carry over to outside benchmarks, with 91.5% accuracy and F1 on
ASSE-Safety, 88.4% accuracy onATBench, and the best worst-case F1 across three benchmarks (88.3% vs. 82.5% forAgentDoG-Llama3.1-8B), but most ofGuardPO's gain over plain fine-tuning comes on the authors' own traces (+10.4 points onCUA-Execversus at most +2.2 elsewhere). - Limitations: the plain fine-tuned model beats the final model on
R-Judge,BraveGuardkeeps higher recall and F1 onAgentHazardtraces from the GPT-5.5 agent, and the full guard takes about 3.1 s per trajectory, while lightweight classifier heads that are 10× faster drop to 69.6% Macro-F1.
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Measuring how LLM agents scale with test-time compute is hard because they decide for themselves how to spend tokens on revising, using tools, exploring and stopping. The proposed Elo-per-token analysis works on open-ended tasks that score intermediate submissions on a continuous scale: it tracks the best solution found at each token budget and uses a Bradley-Terry model to combine within-task rankings into Elo ratings that compare across tasks.
- The study covers four general-purpose agents on four open-ended benchmarks with sessions of up to 100M tokens, plus three feedback-driven LLM optimization harnesses tested in controlled single-task experiments.
- The baseline is independent sampling, whose Elo grows linearly with log compute in theory; agents beat this baseline early on, but their gains per token shrink until they fall below it.
- On shared
AtCoder Heuristic Contesttasks, the strongest past human contestants instead improve superlinearly over contest time, which the authors take as evidence of continual learning and of large headroom once agents slow down. - The authors call the per-session budget at which an agent's gains per token drop to the baseline's rate the scaling inflection point, and using it as the session size when splitting 100M tokens into parallel sessions on
FrontierCSPolyomino Packing gains +264 Elo over one long session and +355 Elo over ten short sessions. - The method only works on tasks that score intermediate submissions continuously, and the budget-splitting gain is shown on a single task, so it is unclear how far it generalizes.
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Kaininja: Extending Native 3D Generators to the Part Level
Native image-to-3D generators such as TRELLIS.2 output one fused mesh. Their O-Voxel grid holds only one sheet of surface per voxel, so where two parts touch, the two surfaces merge into one at any resolution. KaiNinja splits an object's parts across two O-Voxel volumes so that touching parts never share a volume, and it extends TRELLIS.2 into a two-stream generator whose parts are simply the connected pieces of each volume, with no mask, bounding box or segmenter.
- Parts are assigned to the two volumes by two-coloring a contact graph measured on the voxel grid, and the first stage, which lays out the parts, gets a separate copy of the pretrained transformer per volume, joined by new cross-volume attention that starts at zero, while the second, refinement stage reuses the released
TRELLIS.2weights and limits attention to one volume in 20 of its 30 blocks. - Training uses 19,132 objects from four sources, including
Articraft-10K, whose assets an LLM coding agent built part by part so the part labels are exact, which the authors call the first use of agent-authored assets to train a 3D generative model. - On 986 held-out objects it is best on all seven metrics, with whole-object Chamfer distance of 0.0186 against 0.0314 for the strongest baseline,
Hunyuan3D-2.1 + X-Part(40% lower), and strict part F-score of 0.692 against 0.597, at about 24 seconds per object on one H100, far below that segment-then-regenerate cascade's cost. - Whole objects also come out better than from
TRELLIS.2@512fine-tuned on the same data, with strict F-score of 0.919 against 0.830 and a failure rate of 0.7% against 1.8%, which the authors read as evidence that the two-volume packing represents whole objects better too. - On the external
HY3D-Benchit still leads on whole-object geometry, butHunyuan3D-2.1 + X-Partwins every part metric andKaiNinjaranks fourth of six on part mIoU, which the authors attribute to cutting objects into finer parts than the benchmark labels; separately, two-coloring over-merges objects whose parts interlock densely.
Atria Dawn: The Dawn of Agentic Superintelligence
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence
Unclassified 703
Conflict-Predictive Variable Horizons in Multi-Drone Distributed Model Predictive Control
No summary available — see the abstract on arXiv.
CANAL: Channel-Aware Noise Allocation for Differentially Private Feature Distillation in Medical Image Segmentation
No summary available — see the abstract on arXiv.
Survey of Novel Deep Learning Architectures for Denoising Gravitational-wave Signals
No summary available — see the abstract on arXiv.
SomBench: Benchmark Dataset for Advancing Machine Learning in Lunar Science
No summary available — see the abstract on arXiv.
Large-scale bioacoustic detection using semantic segmentation: a deep learning framework applied to fin whale calls in ocean-bottom seismometer recordings
No summary available — see the abstract on arXiv.
Multimodal-Multiresolution Foundation Model for Lunar Remote Sensing
No summary available — see the abstract on arXiv.
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
No summary available — see the abstract on arXiv.
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
No summary available — see the abstract on arXiv.
ScorePrompts: Natural-Language Exploration of Symbolic Music Scores through Analysis
No summary available — see the abstract on arXiv.
Aries: A Proprietary Medium-Range Weather Prediction Model for the Energy Industry
No summary available — see the abstract on arXiv.
Ergodic Control and Controlled Diffusion for Robot Learning: Review and Tutorial
No summary available — see the abstract on arXiv.
Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning
No summary available — see the abstract on arXiv.
Variational Template Matching with Statistical Fusion for Anomaly Detection in Patterned Structures
No summary available — see the abstract on arXiv.
Adaptive Conformal Redistribution for Inter-class Transitional Uncertainty in Medical Image Classification
No summary available — see the abstract on arXiv.
IMM-based Multiple Object Tracking using a State Prediction Neural Network
No summary available — see the abstract on arXiv.
GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures
No summary available — see the abstract on arXiv.
Task-Based CT Protocol Optimization Using Reinforcement Learning and Virtual Imaging Trials
No summary available — see the abstract on arXiv.
Feasibility and Memory Mechanisms of Chern-Simons Context Reservoir Computation
No summary available — see the abstract on arXiv.
Decoupling Error Attribution in Cloud-Native Graph-RAG: A Data Integrity Diagnostic Framework
No summary available — see the abstract on arXiv.
Pedestrian Crossing Intent Classification From Event-Based Vision Using Convolutional Spiking Neural Networks With Temporal Augmentation
No summary available — see the abstract on arXiv.
The Agentic Company OS: Substrate Inversion for Sustained Enterprise Agent Deployment
No summary available — see the abstract on arXiv.
Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework
No summary available — see the abstract on arXiv.
ProtoCAM: Interpretable Few-Shot Mask-Guided Prototypical Learning for Breast Lesion Classification in Ultrasound Imaging
No summary available — see the abstract on arXiv.
Stochastic Gradient Descent over P2
No summary available — see the abstract on arXiv.
Beyond Point Forecasts: A Survey on Probabilistic Forecasting for Time Series and Spatiotemporal Data
No summary available — see the abstract on arXiv.
ViFA-Council: Multi-Agent LLM Deliberation for Vietnamese Folk Art Generation
No summary available — see the abstract on arXiv.
Real-time Learning and Evolution in Robotic Art Installations
No summary available — see the abstract on arXiv.
SkillAtlas: An Attack Trace Library for Agent Skills
No summary available — see the abstract on arXiv.
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
No summary available — see the abstract on arXiv.
RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines
No summary available — see the abstract on arXiv.
Privacy-Preserving Deep Joint Source-Channel Coding with In-Loop Concept Erasure
No summary available — see the abstract on arXiv.
Task-Aware Federated Fine-Tuning for MoE-based Large Language Models
No summary available — see the abstract on arXiv.
Converge Then Diversify: Decoupling Convergence and Diversity in Multi-Objective Bayesian Optimisation
No summary available — see the abstract on arXiv.
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
No summary available — see the abstract on arXiv.
CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages
No summary available — see the abstract on arXiv.
Specification Oracles
No summary available — see the abstract on arXiv.
Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
No summary available — see the abstract on arXiv.
ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement
No summary available — see the abstract on arXiv.
Chance-Constrained Belief-Space Maneuver Planning for Autonomous Collision Avoidance Under Uncertainty
No summary available — see the abstract on arXiv.
Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?
No summary available — see the abstract on arXiv.
LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents
No summary available — see the abstract on arXiv.
Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization
No summary available — see the abstract on arXiv.
Efficient Online Inverse Optimization with $O(d)$ Regret
No summary available — see the abstract on arXiv.
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
No summary available — see the abstract on arXiv.
Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs
No summary available — see the abstract on arXiv.
A Machine Learning API for Earth Observation Data Cubes Based on openEO
No summary available — see the abstract on arXiv.
Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment
No summary available — see the abstract on arXiv.
TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models
No summary available — see the abstract on arXiv.
STAGE: Diagnosing Semantic Transfer at Grounded Execution in Embodied Agents
No summary available — see the abstract on arXiv.
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
No summary available — see the abstract on arXiv.
Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement
No summary available — see the abstract on arXiv.
Symmetry- and Property-Aware Crystal Generation with Reinforcement Learning for Inverse Materials Design
No summary available — see the abstract on arXiv.
Inverse Learning of the Altruism and Cost Level in Mixed-Individual Mean Field Games
No summary available — see the abstract on arXiv.
OrchSLM: Probing the Dynamics of Small Language Model Orchestration
No summary available — see the abstract on arXiv.
Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports
No summary available — see the abstract on arXiv.
Edge-addition monotonicity of positive p-energy fails for every p >= 1
No summary available — see the abstract on arXiv.
On the Potential of Multi-Task Learning in Predictive Process Monitoring
No summary available — see the abstract on arXiv.
A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text
No summary available — see the abstract on arXiv.
Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers
No summary available — see the abstract on arXiv.
Token Efficient Task Execution via Application Behavior Modeling for Web Agents
No summary available — see the abstract on arXiv.
Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study
No summary available — see the abstract on arXiv.
Building a Production Greek-English Speech Recognizer
No summary available — see the abstract on arXiv.
Canaries in the Bank: Auditing User-Level Privacy in Private Evolution
No summary available — see the abstract on arXiv.
One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction
No summary available — see the abstract on arXiv.
Pretraining for Sample-Efficient Neural Interfaces
No summary available — see the abstract on arXiv.
A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion
No summary available — see the abstract on arXiv.
Adaptive Phase-Switching for Communication-Efficient Federated LoRA Fine-Tuning
No summary available — see the abstract on arXiv.
Operational Range Bounding in Spectroscopy: A Safety Cage Framework for Machine Learning Models
No summary available — see the abstract on arXiv.
GeoTTER: Leveraging Local Geometry of Optimal Transport for Zero-Shot Classification
No summary available — see the abstract on arXiv.
Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
No summary available — see the abstract on arXiv.
From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models
No summary available — see the abstract on arXiv.
Positive Topology and Feasible Refinement: Forcing Matrices, Positivity, and Information
No summary available — see the abstract on arXiv.
Rolling Day-Wise Mortality Prediction in Critically Ill Patients With AKI on CRRT Utilizing Machine Pressure Waveforms
No summary available — see the abstract on arXiv.
A family of spectral conjugate gradient algorithms derived by least-squares approximations based on a modified quasi--Newton update with application to a revised robust binary classification model
No summary available — see the abstract on arXiv.
Generative Interpretability via Scalable Neuro-Symbolic Models
No summary available — see the abstract on arXiv.
Attention Is All You Need (to Avoid Spurious Oscillations)
No summary available — see the abstract on arXiv.
Early-Stopping Thresholds for ES-HyperNEAT: A Data-Driven Approach from Fitness Dynamics
No summary available — see the abstract on arXiv.
Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models
No summary available — see the abstract on arXiv.
From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements
No summary available — see the abstract on arXiv.
Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
No summary available — see the abstract on arXiv.
Runtime-Incremental Transformer for Reinforcement-Learning-Based Adaptive Control
No summary available — see the abstract on arXiv.
Toward Optimal Switching Regret for Multi-Armed Bandits with Oblivious Adversary
No summary available — see the abstract on arXiv.
AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
No summary available — see the abstract on arXiv.
Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management
No summary available — see the abstract on arXiv.
Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models
No summary available — see the abstract on arXiv.
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
No summary available — see the abstract on arXiv.
A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics
No summary available — see the abstract on arXiv.
When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence
No summary available — see the abstract on arXiv.
Planning or Learning: Reliability and Cost in Multi-Asset Maintenance
No summary available — see the abstract on arXiv.
Causal multi-modal AI for personalized chemosensitivity prediction
No summary available — see the abstract on arXiv.
A pullback-corrected scalar auxiliary variable optimizer with momentum and adaptive mobility
No summary available — see the abstract on arXiv.
How User-AI Mistreatment Occurs and Matters in Conversational Systems?
No summary available — see the abstract on arXiv.
FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks
No summary available — see the abstract on arXiv.
Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation
No summary available — see the abstract on arXiv.
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
No summary available — see the abstract on arXiv.
mKernel: Fast Multi-GPU, Multi-Node Fused Kernels
No summary available — see the abstract on arXiv.
BOOST: Concurrent Access to Host Memory and HBM to Accelerate LLM Inference
No summary available — see the abstract on arXiv.
Predictive audio representations for early detection and tracking of hidden dynamic objects
No summary available — see the abstract on arXiv.
$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
No summary available — see the abstract on arXiv.
In the Blind: Building Pseudo-References for MT Evaluation
No summary available — see the abstract on arXiv.
AttnFuse: A Composable DSL for Compiling Attentions to Fused GPU Kernels
No summary available — see the abstract on arXiv.
The University of Melbourne WMT 2026 CreoleMT Submission: A Domain-Balanced Approach to Low-Resource Pacific Creole Machine Translation
No summary available — see the abstract on arXiv.
An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS
No summary available — see the abstract on arXiv.
FaithfulBench: Does AI Counsel Uphold or Undermine the User's Professed Faith?
No summary available — see the abstract on arXiv.
EI-DDLGN: Efficient Encrypted Inference with Deep Differentiable Logic Gate Networks under TFHE
No summary available — see the abstract on arXiv.
Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
No summary available — see the abstract on arXiv.
FlowTSFM: Turning Encoder Depth into Quantile Transport
No summary available — see the abstract on arXiv.
When Compliance Data Masquerades as Evaluation: Measurement Validity for Deployed AI Systems
No summary available — see the abstract on arXiv.
Curvature-Independent Regret Bounds for Distributed Online Optimization on Hadamard Manifolds
No summary available — see the abstract on arXiv.
Solar Intelligence
No summary available — see the abstract on arXiv.
Online Bayesian Node Classification on Inductive Graphs under Distribution Shift
No summary available — see the abstract on arXiv.
Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself
No summary available — see the abstract on arXiv.
FedV-KGQA in Practice: Design Lessons and an Interactive Prototype
No summary available — see the abstract on arXiv.
Cost Characterization of Vertically Partitioned Federated Knowledge Graphs
No summary available — see the abstract on arXiv.
GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents
No summary available — see the abstract on arXiv.
Enhancing Event Candidate Acquisition for Event Linking
No summary available — see the abstract on arXiv.
Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations
No summary available — see the abstract on arXiv.
Recoverability as a System Primitive for Long-Horizon AI Agents
No summary available — see the abstract on arXiv.
Windowed A-K-MDP
No summary available — see the abstract on arXiv.
Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models
No summary available — see the abstract on arXiv.
LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference
No summary available — see the abstract on arXiv.
Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding
No summary available — see the abstract on arXiv.
Prefix Sharing Is a Sorting Problem
No summary available — see the abstract on arXiv.
Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction
No summary available — see the abstract on arXiv.
Gap Entropy and Almost Instance-Wise Optimal Best-Arm Identification
No summary available — see the abstract on arXiv.
When Edit Localization Amplifies Relative Selection Bias: Gradient Geometry, Target Mismatch, and Importance Weighting
No summary available — see the abstract on arXiv.
Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks
No summary available — see the abstract on arXiv.
Certifying Model Upgrades with Slice-Wise Non-Regression and Incumbent Fallback
No summary available — see the abstract on arXiv.
JaxAHT: A JAX-Based Library for Ad Hoc Teamwork
No summary available — see the abstract on arXiv.
MANAS-2: Constrained Reconstruction for EEG Foundation Models
No summary available — see the abstract on arXiv.
Oops, Not Now: PEARL, a RAG-Based Support Agent for Gameplay and What Players Want from AI Help
No summary available — see the abstract on arXiv.
Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging
No summary available — see the abstract on arXiv.
IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
No summary available — see the abstract on arXiv.
TyPatch: Transforming Patches into Typestate Rules for Kernel Bug Detection
No summary available — see the abstract on arXiv.
A Variational Optimal Transport Operator on Incompressible Flow
No summary available — see the abstract on arXiv.
JumpStart Your Policy Learning with Lessons from 160,000 Training Runs
No summary available — see the abstract on arXiv.
Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges
No summary available — see the abstract on arXiv.
PolicyMem: Geometric Policy Memory for LLM Governance
No summary available — see the abstract on arXiv.
ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
No summary available — see the abstract on arXiv.
HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning
No summary available — see the abstract on arXiv.
Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth
No summary available — see the abstract on arXiv.
How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition
No summary available — see the abstract on arXiv.
On the Equivalence of Stochastic Control and Path Space Formulations for Schr\"odinger Bridges over Compact Connected Lie Groups
No summary available — see the abstract on arXiv.
Positioning manuscripts in the scientific landscape with agentic AI
No summary available — see the abstract on arXiv.
HyperProve: Answer-Guided Hypergraph Expansion for Multi-Hop Question Answering
No summary available — see the abstract on arXiv.
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
No summary available — see the abstract on arXiv.
Homeostatic Continual Learning
No summary available — see the abstract on arXiv.
Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging
No summary available — see the abstract on arXiv.
Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs
No summary available — see the abstract on arXiv.
GEAR: From Dynamic Encoding to Dynamic Activation in Social Trajectory Prediction
No summary available — see the abstract on arXiv.
Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection
No summary available — see the abstract on arXiv.
PPDL: A Real-world Industrial User Retention Ratio Forecasting Framework Integrating Physical Priors with Deep Learning
No summary available — see the abstract on arXiv.
SyRHM: Symbolic-Language-Enhanced Reasoning with Associative Retrieval for Zero-shot Harmful Meme Detection
No summary available — see the abstract on arXiv.
Resolution-Independent Analysis of Encoder--Decoder Operator Learning via Limiting Kernels
No summary available — see the abstract on arXiv.
Do Not Restart: Residual Completion for Stateful Agent Handoffs
No summary available — see the abstract on arXiv.
LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems
No summary available — see the abstract on arXiv.
Understanding the Limits of Agentic ICD Coding
No summary available — see the abstract on arXiv.
Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture
No summary available — see the abstract on arXiv.
DARE: Dialectical Agentic Reasoning for Structured Knowledge Fact Checking
No summary available — see the abstract on arXiv.
Odds-Shift Slippage in One-vs-Rest Rankers: Diagnosing and Repairing Reweighting-Induced Top-K Errors
No summary available — see the abstract on arXiv.
Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models
No summary available — see the abstract on arXiv.
Benchmarking Optimizers to Solve Inverse Problems with Differentiable Physics Simulators
No summary available — see the abstract on arXiv.
When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings
No summary available — see the abstract on arXiv.
ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading
No summary available — see the abstract on arXiv.
Graph Neural Networks for Influence Maximization in Social Networks: An Unsupervised Minimum Dominating Set Approach
No summary available — see the abstract on arXiv.
Accuracy Is Not Service: A Decision-Aware Benchmark for Intermittent-Demand Forecasting
No summary available — see the abstract on arXiv.
Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice
No summary available — see the abstract on arXiv.
CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection
No summary available — see the abstract on arXiv.
Pre-training with Graph Transformers
No summary available — see the abstract on arXiv.
LePlanner: An Iterative Amortized Controller For World Models
No summary available — see the abstract on arXiv.
Affinity-Aware Sharding for Delayed Tensor Parallelism
No summary available — see the abstract on arXiv.
Measuring the Cost of Variety Conflation in Multilingual MT Evaluation: Adding Mozambican Xichangana, Nyanja and Sena to FLORES+
No summary available — see the abstract on arXiv.
UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics
No summary available — see the abstract on arXiv.
ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting
No summary available — see the abstract on arXiv.
ShopEase: A Generative AI-Based Multi-Agent Framework for Intelligent Enterprise Customer Support Using Hybrid Retrieval-Augmented Generation
No summary available — see the abstract on arXiv.
ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation
No summary available — see the abstract on arXiv.
ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
No summary available — see the abstract on arXiv.
An Uncertainty-Aware Hybrid Mathematical-Machine-Learning Model for Smart Irrigation Decision Support
No summary available — see the abstract on arXiv.
The Filter Metric is Safety-Critical: Phantom Advantages in Group-Relative RL under Shaped Rewards
No summary available — see the abstract on arXiv.
From Network Inequality to Network Fairness: A Perspective on Responsible Decision-Making
No summary available — see the abstract on arXiv.
Physically Typed and Geometry-Aware Representations for Earth Foundation Models
No summary available — see the abstract on arXiv.
Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability
No summary available — see the abstract on arXiv.
Trustworthy, Explainable, and Sustainable Decentralized Intelligence for 6G Networks
No summary available — see the abstract on arXiv.
SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity
No summary available — see the abstract on arXiv.
Generation of Custom Solvers in Rust for Convex Optimization
No summary available — see the abstract on arXiv.
A Multi-Resolution Multi-Domain Pre-Training Framework for Universal Traffic Forecasting
No summary available — see the abstract on arXiv.
Map Users and Mapmakers: The Scope of Cognitive Attribution from Acquired Representations
No summary available — see the abstract on arXiv.
When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents
No summary available — see the abstract on arXiv.
CyclOT: Learning Quadratic Optimal Transport Maps via Synchronized Forward-Backward Interpolants
No summary available — see the abstract on arXiv.
Lie to me: Detecting Managerial Evasiveness in Earnings Calls via Conversational Audio Encoders
No summary available — see the abstract on arXiv.
URCHIN: A Horizontal Spiking Language Model for Data-Constrained Pretraining
No summary available — see the abstract on arXiv.
HQARRF: Hierarchical Q-learning and Force-aware Routing for Multi-Charger Scheduling in Wireless Rechargeable Sensor Networks
No summary available — see the abstract on arXiv.
Phorecaster365: A Human-Supervised Reference Architecture for Hybrid Pharmaceutical Sales Forecasting and Planning Decision Support
No summary available — see the abstract on arXiv.
DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis
No summary available — see the abstract on arXiv.
Machine Learning under Imperfect Data: Challenges and Methods
No summary available — see the abstract on arXiv.
North Small Translate: Advanced Cost-Effective Translation (Cohere CAT+)
No summary available — see the abstract on arXiv.
Learning Through Energy Refinement and Manifold Projection: A Cooperative EBM-AE Framework
No summary available — see the abstract on arXiv.
LoRA Fine-Tuned Models for Control Systems Course Q\&A: A Multidimensional Evaluation of Model Scale and Rank Effects
No summary available — see the abstract on arXiv.
Machine Learning in Fish Farming
No summary available — see the abstract on arXiv.
Finite-Time Node Separation in Recurrent Graph Neural Networks with Persistent Gaussian Perturbations
No summary available — see the abstract on arXiv.
Minibatch persistency, eight years later: what batch reuse costs in steps and joules, and what it saves in data
No summary available — see the abstract on arXiv.
Equilibrium bias and convergence in augmented primal--dual dynamics with sampled constraints
No summary available — see the abstract on arXiv.
Exploring napping paradigm for Recurrent Spiking Neural Networks
No summary available — see the abstract on arXiv.
Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus
No summary available — see the abstract on arXiv.
Optimal Transport for Efficient, Unsupervised Anomaly Detection on Industrial Data
No summary available — see the abstract on arXiv.
CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education
No summary available — see the abstract on arXiv.
SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery
No summary available — see the abstract on arXiv.
Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision
No summary available — see the abstract on arXiv.
Thought without systematicity? Evaluating reasoning models on rule induction tasks
No summary available — see the abstract on arXiv.
Linear Ensemble Sampling with Smaller Ensembles
No summary available — see the abstract on arXiv.
Tabby: An Open Pretraining Recipe for Time Series Foundation Models
No summary available — see the abstract on arXiv.
SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection
No summary available — see the abstract on arXiv.
Introspective Uncertainty Estimation for LLM-Based Code Generation
No summary available — see the abstract on arXiv.
Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context
No summary available — see the abstract on arXiv.
Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration
No summary available — see the abstract on arXiv.
Diffusion-Based Multiple-Shooting Indirect Optimal Control for Fuel-Optimal Spacecraft Trajectory Generation
No summary available — see the abstract on arXiv.
Synthetic Data in Marketing Research: How to Evaluate and When to Trust
No summary available — see the abstract on arXiv.
Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR
No summary available — see the abstract on arXiv.
Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control
No summary available — see the abstract on arXiv.
Convergent Emergence of In-Context Learning Across Modalities
No summary available — see the abstract on arXiv.
Schizophrenia Detection from EEG Signals: A Transformer Framework with Spectrogram Representation
No summary available — see the abstract on arXiv.
TF-IDF and BM25 Are Exact KL Divergences
No summary available — see the abstract on arXiv.
VeriDx: Earning the Right to Diagnose with Disease-Centric Verification
No summary available — see the abstract on arXiv.
Conditional Quantum Flow Matching for Data-Scarce Physiological Signal Augmentation
No summary available — see the abstract on arXiv.
Real-World Deployment and Performance Characterisation of Fog-Based Deep Learning for Cold-Chain Temperature Prediction over LoRaWAN
No summary available — see the abstract on arXiv.
Data-Efficient Agentic Graph Domain Adaptation via Reliability-Aware Prototype Learning
No summary available — see the abstract on arXiv.
Measuring the Creativity of Frontier LLMs in Automated Research
No summary available — see the abstract on arXiv.
AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents
No summary available — see the abstract on arXiv.
Stabilizing Performative Feedback Loops with Minimal Model Deployments
No summary available — see the abstract on arXiv.
GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning
No summary available — see the abstract on arXiv.
Multi-Modal Tumor Survival Prediction via Graph-Guided Mixture of Experts
No summary available — see the abstract on arXiv.
LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models
No summary available — see the abstract on arXiv.
Just add noise: Debiasing tree-based variable importance in mixed data
No summary available — see the abstract on arXiv.
NeuroFlex: Lossless Element-Level ANN-SNN Co-Execution for Efficient Sparse Inference
No summary available — see the abstract on arXiv.
Bridging the Synthetic-to-Real Gap for Few-Shot Cryo-ET Classification
No summary available — see the abstract on arXiv.
Neuron Activation-based Computation of Logical Explanations for Deep Neural Networks
No summary available — see the abstract on arXiv.
RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
No summary available — see the abstract on arXiv.
A Voxel-Spacing-Aware Extension of PyRadiomics for Anisotropic Texture Analysis
No summary available — see the abstract on arXiv.
Real-Time Synthesis of Robust Controlled Invariant Sets for Monotone Systems
No summary available — see the abstract on arXiv.
Talking to Me or Someone Else? Rethinking Talk-to-Me Detection in Egocentric Videos
No summary available — see the abstract on arXiv.
Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had
No summary available — see the abstract on arXiv.
To do($x$) or not to do($x$): Medical Image Counterfactuals for Dataset Augmentation
No summary available — see the abstract on arXiv.
Exact Finite Attention Responses From RoPE Derivatives
No summary available — see the abstract on arXiv.
LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents
No summary available — see the abstract on arXiv.
Riemannian ascent--descent for nonconvex nonconcave minimax landscapes: convergence to basin saddle points and applications to distributionally robust optimization
No summary available — see the abstract on arXiv.
T-SMART: Mechanism-Level Attribution for Tool-Augmented Time-Series Question Answering
No summary available — see the abstract on arXiv.
One Size Does Not Fit All: Setting Inference Depth from the Questions a Deployment Actually Asks
No summary available — see the abstract on arXiv.
A New Transformer-Based Approach for Audio-Based Kinship Verification and a New Uncontrolled Mandarin Kinship Speech Dataset
No summary available — see the abstract on arXiv.
When Faster VLA Deployment Changes Closed-Loop Behavior: Task Success-Latency Analysis of SmolVLA Across PyTorch and ONNX Variants
No summary available — see the abstract on arXiv.
Signatures of Steerability in Activation Space of Language Models
No summary available — see the abstract on arXiv.
When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering
No summary available — see the abstract on arXiv.
3D Gait-Based Autism Classification Using Attention-Enhanced Deep Learning with Cross-Fold Statistical Stability Analysis
No summary available — see the abstract on arXiv.
Towards Evolving Context Parameterization for Large Language Models
No summary available — see the abstract on arXiv.
CyFM: Cylindrical Optimal Transport for Few-Step Complex-Valued Flow Matching
No summary available — see the abstract on arXiv.
Inherited Heads: Audio language models track speakers with their text backbone's attention, and an attention-mass ranking retrieves a different set
No summary available — see the abstract on arXiv.
A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement
No summary available — see the abstract on arXiv.
A Machine Learning Framework for Fault Detection, Isolation, and Severity Prediction of Autonomous VTOL Aircraft
No summary available — see the abstract on arXiv.
ZAPS: Zero-Cost Active Proxy Search for Neural Architecture Search
No summary available — see the abstract on arXiv.
Bi-Level Routing and Sparse Spatial Attention based Multi-View BEV 3D Object Detection for Autonomous Driving
No summary available — see the abstract on arXiv.
Entropy-Punctured Bloom Filters for Memory-Efficient Machine Learning
No summary available — see the abstract on arXiv.
Data-free On-policy Distillation
No summary available — see the abstract on arXiv.
Learning to Refer from Estimated Listener Gaze
No summary available — see the abstract on arXiv.
ECAS: An Edge-Controlled Agentic System for Validation-Gated Scientific Application Execution
No summary available — see the abstract on arXiv.
Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution
No summary available — see the abstract on arXiv.
Task-Specified Active Metrological Inspection with Measurement-Steered VLA Manipulation and Deterministic Evidence Gating
No summary available — see the abstract on arXiv.
Corpus Characterization and Inverse Constitutional Fine-Tuning for Style-Aware Radiology Reports
No summary available — see the abstract on arXiv.
Modeling, Scaling, and Decoding: Optimizing Controllable Speech Generation with Nonverbal Vocalizations
No summary available — see the abstract on arXiv.
Graph-Transformer Fraud Detection with Self-Supervised Pretraining and Conformal Risk Control
No summary available — see the abstract on arXiv.
Assessing the Applicability of Existing Design Recommendations to AI Companion Design: A Multi-Method Study
No summary available — see the abstract on arXiv.
OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving
No summary available — see the abstract on arXiv.
CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time
No summary available — see the abstract on arXiv.
The Attribution-Compression Frontier in Retrieval-Augmented Generation
No summary available — see the abstract on arXiv.
Joint Optimization for Federated Learning and Transmission over Unreliable Wireless Networks with Heterogeneous Data
No summary available — see the abstract on arXiv.
What Input Resolution Is Required for Bird Species Identification, and What Is Its Latency Cost on an Edge Device? A Study of 14 Input Resolutions and Six Architectures with On-Device Measurements
No summary available — see the abstract on arXiv.
ATTRICITE: Training an Open 4B Model for Citation Recovery toward Faithful Attribution
No summary available — see the abstract on arXiv.
Towards Anticipatory Databases Through Shared Data and Workload Semantics
No summary available — see the abstract on arXiv.
Document Topic Alignment Metrics for Evaluating Topic Models of Short-Text Public Health Communications on Social Media
No summary available — see the abstract on arXiv.
DenMark: Robust Semantic Watermarking for Diffusion Language Models
No summary available — see the abstract on arXiv.
VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching
No summary available — see the abstract on arXiv.
Bayesian optimization with kernel ensembles and disagreement-based acquisition for source localization and acoustic inversion
No summary available — see the abstract on arXiv.
Parameter Estimation of Ringdown Quasinormal Modes with Autoencoder
No summary available — see the abstract on arXiv.
SpermYOLO: A Coordinated YOLO-Based Detector for Accurate and Efficient Sperm and Impurity Detection in Microscopic Images
No summary available — see the abstract on arXiv.
Biquaternionic Space with Complex-valued Attention for Temporal Knowledge Graph Completion
No summary available — see the abstract on arXiv.
Editorial routing shapes how computational results are qualified in AI-assisted scientific writing
No summary available — see the abstract on arXiv.
Fusing Spectral Signatures and Activation Clustering for Backdoor Detection in Healthcare Imaging Models: Method, Implementation, and Evaluation
No summary available — see the abstract on arXiv.
Learning Source Acquisition Policies by Offline Planning
No summary available — see the abstract on arXiv.
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
No summary available — see the abstract on arXiv.
S3-Tracker: Self-Supervised Surgical Tissue Tracking With Contrastive Random Walks
No summary available — see the abstract on arXiv.
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
No summary available — see the abstract on arXiv.
AI Assisted Workflow Optimization and Automation
No summary available — see the abstract on arXiv.
Nonparametric Variance-Penalized Actor-Critic: Statistical Inference for Risk-Sensitive Reinforcement Learning
No summary available — see the abstract on arXiv.
Communication-Efficient LLM Adaptation over Decentralized GPU Meshes
No summary available — see the abstract on arXiv.
AURA: Unified Multimodal Framework for Conversational Music Editing
No summary available — see the abstract on arXiv.
Multimodal deep learning from spectra for small-molecule structure identification: enhancing robustness with mixed-condition training
No summary available — see the abstract on arXiv.
LLaTSA: Large Language Model-Aligned General-Purpose Transient Stability Analysis
No summary available — see the abstract on arXiv.
Formal Properties of Language as Constraints on Neural Dynamics
No summary available — see the abstract on arXiv.
MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents
No summary available — see the abstract on arXiv.
Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error
No summary available — see the abstract on arXiv.
Neural Modal Decomposition: Architectural Priors from Observables
No summary available — see the abstract on arXiv.
Dynamic Learning Solutions: A System for Personalized Educational Video Generation
No summary available — see the abstract on arXiv.
Question's Gambit: The First Move Matters in Agentic Deep Search
No summary available — see the abstract on arXiv.
A Hybrid Dependency-Aware Framework for Task Decomposition and Dynamic Agent Generation in Oracle-to-PostgreSQL Migration
No summary available — see the abstract on arXiv.
Surrogate-Assisted Genetic Programming with Phenotypic Characterisation in Dynamic Multi-Mode Project Scheduling
No summary available — see the abstract on arXiv.
A Generative AI Integrated Multimodal Framework for Low-Latency Multi-Camera Person Re-Identification
No summary available — see the abstract on arXiv.
Safety Signals to Verify NetOps Agents with Action-Level Granularity
No summary available — see the abstract on arXiv.
Certification cost of quantum models: measurement correlation, not parameter count
No summary available — see the abstract on arXiv.
Has Scientific Talent Shifted from Depth to Breadth?Evidence across Papers, Knowledge Inputs, Careers, and Teams
No summary available — see the abstract on arXiv.
Lightweight Generalized DeepFake Face Detection with WAVIE: Wavelet Augmented Vision Intermediate Embeddings
No summary available — see the abstract on arXiv.
A latent dimension of Condorcet's jury theorem for multiple AI advisers
No summary available — see the abstract on arXiv.
From Visual Attribution to Clinical Reasoning: Explainable Parkinson's Disease Screening from Hand-Drawn Patterns
No summary available — see the abstract on arXiv.
NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass
No summary available — see the abstract on arXiv.
Towards Identifying the Dataset Biases Causing Phantom Transfer
No summary available — see the abstract on arXiv.
Follow the Geometry, Not the Model: Cold Start Semi-Supervised Learning
No summary available — see the abstract on arXiv.
Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation
No summary available — see the abstract on arXiv.
Retrieval-Guided Fine-Tuning as Noisy Estimation: Risk bounds and Architectural Analysis
No summary available — see the abstract on arXiv.
EdgeHAR: An Edge-Native Compact Sensor Foundation Model for Human Activity Recognition
No summary available — see the abstract on arXiv.
When does a scaling result justify a different allocation? A critical review of resource-allocation evidence for AI systems
No summary available — see the abstract on arXiv.
Should All Noises Be Treated Equally: Impact of Input Noise Variability on Neural Network Robustness
No summary available — see the abstract on arXiv.
Physically Partitioned KVCache Format for CPU--GPU Load Balancing in MoE Inference
No summary available — see the abstract on arXiv.
Beyond Scene Description: Multi-Agent Orchestration for Non-visual Access to Virtual Worlds
No summary available — see the abstract on arXiv.
OptoAgent: A Trustworthy Multi-Agent Framework for Opportunistic Vision Micro-Screening in Classroom Environments
No summary available — see the abstract on arXiv.
Toward a Layer-2 Trigger for AI/ML Lifecycle Management in 6G
No summary available — see the abstract on arXiv.
Theseus in the Graph: Towards Traceable Multi-Hop Graph Navigation
No summary available — see the abstract on arXiv.
Parameter-Efficient Quantum NLP for Paraphrase Detection: Performance, Robustness, and Entanglement
No summary available — see the abstract on arXiv.
Sharing standardized image-derived data in computational pathology using DICOM
No summary available — see the abstract on arXiv.
Multi-source conformal prediction: leveraging heterogeneity via localization
No summary available — see the abstract on arXiv.
Proving olympiad geometry theorems on a superconducting quantum processor
No summary available — see the abstract on arXiv.
GNN4PPM: Multi-Target Predictive Process Monitoring with Relational Graph Convolutional Networks
No summary available — see the abstract on arXiv.
Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition
No summary available — see the abstract on arXiv.
Selecting k Paths with the Minimum Longest Path Length in the Stochastic Semi-Bandit Setting
No summary available — see the abstract on arXiv.
TATK: Triple-Aware Top-K Learning with Knowledge-Grounded Verification for LLM-based Sequential Recommendation
No summary available — see the abstract on arXiv.
Learning Multi-Agent Task Assignment and Navigation in the Factory: from Simulation to Real Robots
No summary available — see the abstract on arXiv.
Evaluation of optimisation and Bayesian inference methods for reaction rates in atmospheric chemical mechanisms
No summary available — see the abstract on arXiv.
Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection
No summary available — see the abstract on arXiv.
Domain-specific Pretraining Profile and Transformer Performance: Evidence from Modeling Digital Pragmatics in Arabic-English Code-switching
No summary available — see the abstract on arXiv.
AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory
No summary available — see the abstract on arXiv.
Pathwise Individual Rationality in Federated Learning: A Mechanism-Architecture Co-Design
No summary available — see the abstract on arXiv.
AI Deployment Accountability Engineering: A Vision for Accountable AI in Safety-Critical Socio-Technical Systems
No summary available — see the abstract on arXiv.
SENTINEL: A Multi-Pathway Architecture for Detecting Living-Off-the-Land APT Attacks on Windows Command Lines
No summary available — see the abstract on arXiv.
Diagnosing Temporal Misalignment in Multichannel Time-Series Classification with Minimum Description Length
No summary available — see the abstract on arXiv.
A note on goal-based hierarchical RL
No summary available — see the abstract on arXiv.
SH-WRNN: Implicit Spherical Harmonics Weight Field Routing Neural Networks for Asymmetric Edge Intelligence
No summary available — see the abstract on arXiv.
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
No summary available — see the abstract on arXiv.
Channel-Adaptive Region Adjacency Graph Carriers for Semantic Image Communication
No summary available — see the abstract on arXiv.
Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation
No summary available — see the abstract on arXiv.
DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents
No summary available — see the abstract on arXiv.
Investigating the Impacts of Generative AI on Information Seeking
No summary available — see the abstract on arXiv.
Diffusion-Based Generation of Gait Trajectories
No summary available — see the abstract on arXiv.
CompCQR: Compositional Query Generation for Training-Free Conversational Search
No summary available — see the abstract on arXiv.
Skill Composition for Legged Robot Reinforcement Learning
No summary available — see the abstract on arXiv.
Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation
No summary available — see the abstract on arXiv.
Natural Language Knowledge Graph Query Execution: Leveraging Controlled Semantics in the LLM Context Window
No summary available — see the abstract on arXiv.
Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing
No summary available — see the abstract on arXiv.
MANE: A Multi-Path Adaptive Network for Edge Onloading of Deep Neural Networks
No summary available — see the abstract on arXiv.
Symmetries and Singularities
No summary available — see the abstract on arXiv.
Exploring Multimodal Turn-Taking Cues in Face-to-Face Conversation using Voice Activity Projection
No summary available — see the abstract on arXiv.
PU classification under Non-SCAR: clustering-assisted logistic model with oversampling enhancement
No summary available — see the abstract on arXiv.
The Garden of Forking Prompts: How Users Explore Narrative Space in Story Generation
No summary available — see the abstract on arXiv.
One Feedback System Does Not Fit All: Localising Data-to-Text Driver Coaching for the United Kingdom and Nigeria
No summary available — see the abstract on arXiv.
Speak to the City: Multimodal Resolution for Outside-the-Vehicle References
No summary available — see the abstract on arXiv.
WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians
No summary available — see the abstract on arXiv.
Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
No summary available — see the abstract on arXiv.
An immune world model for multiscale forecasting and therapeutic hypothesis generation
No summary available — see the abstract on arXiv.
GRPO-QM: Target Preserving Exploration for Quantum Tomography
No summary available — see the abstract on arXiv.
Learning Metastable Dynamics
No summary available — see the abstract on arXiv.
Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M
No summary available — see the abstract on arXiv.
Moral Rebel Agents: Decision-Making Under Conflicting Obligations
No summary available — see the abstract on arXiv.
Carryover Drafting: Recycling Rejected States for Speculative Decoding
No summary available — see the abstract on arXiv.
Are Gradient Boosting Models Suitable for Intermittent Demand Forecasting?
No summary available — see the abstract on arXiv.
Bayesian Intelligence from the Outside
No summary available — see the abstract on arXiv.
CALICO: A Human-Centered, Codebook-Aligned System for Annotation
No summary available — see the abstract on arXiv.
Parameter isolation with domain-specific experts for incremental audio classification
No summary available — see the abstract on arXiv.
WaVeFuse: Regime-Adaptive Equity Index Forecasting via Channel-Wise Wavelet Denoising and Vertical Attention Fusion
No summary available — see the abstract on arXiv.
OCT-FedSIR: Toward Trustworthy Federated Ophthalmic Learning under Annotation Noise
No summary available — see the abstract on arXiv.
HELENA for 5G NR LEO NTN Channel Estimation: A Comparative Evaluation
No summary available — see the abstract on arXiv.
AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing
No summary available — see the abstract on arXiv.
Building Legal Reward Models for Grounding and Abstention
No summary available — see the abstract on arXiv.
A property-registry contract for retrieve-or-refuse thermal-mechanical lattice search
No summary available — see the abstract on arXiv.
Quantifying the Generation Modality Gap in Speech-Text Language Models
No summary available — see the abstract on arXiv.
AcquireBound: Runtime Authorization for Resources Acquired by AI Agents
No summary available — see the abstract on arXiv.
Calibrating Interpretability Instruments Before Trusting Their Verdicts
No summary available — see the abstract on arXiv.
Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return
No summary available — see the abstract on arXiv.
Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
No summary available — see the abstract on arXiv.
From Visual Feedback to Textual Reviews: A Multi-Agent Vision-Language Framework for Image-Grounded Review Assistance
No summary available — see the abstract on arXiv.
TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
No summary available — see the abstract on arXiv.
CCMAN: Cognitive Instability-Aware Cross-Modal Attention Network for Interpretable Temporal Biomarkers of Verbal Fluency Speech
No summary available — see the abstract on arXiv.
A Personalized Dynamic Balance Evaluation Paradigm for Hip Exoskeleton-Assisted Walking under Unexpected Ground Perturbations
No summary available — see the abstract on arXiv.
Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination
No summary available — see the abstract on arXiv.
How broad is that claim? Mapping Generalisation in NLP Research
No summary available — see the abstract on arXiv.
Pull: Lazy Materialization of Working Memory for Stateful LLM Conversations
No summary available — see the abstract on arXiv.
Privacy Preserving Gossip Learning
No summary available — see the abstract on arXiv.
Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models
No summary available — see the abstract on arXiv.
The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents
No summary available — see the abstract on arXiv.
Language-Guided Representation Learning for Robust Cross-Sensor Material Recognition
No summary available — see the abstract on arXiv.
Mind Which Bird You Favour: Parameterizing Adequacy-Fluency Balance in Meta-Evaluation of Machine Translation
No summary available — see the abstract on arXiv.
AI Persuasion as a Threat to Human Control
No summary available — see the abstract on arXiv.
From matrix inversion to constraints: provably tighter confidence regions for importance weights in label shift
No summary available — see the abstract on arXiv.
Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
No summary available — see the abstract on arXiv.
Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accounting Tasks
No summary available — see the abstract on arXiv.
A Functional SVD Framework for Regularized Multivariate Functional PCA with Dual Penalization
No summary available — see the abstract on arXiv.
Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech
No summary available — see the abstract on arXiv.
A primer on evaluation methods for large language models in healthcare
No summary available — see the abstract on arXiv.
Decision-Oriented Uncertainty Quantification for Risk Control in Earth System Spatiotemporal Foundation Models
No summary available — see the abstract on arXiv.
MedTRACE: Tool-Augmented Multimodal Clinical Reasoning Agents for Evidence-Grounded Decision-Making
No summary available — see the abstract on arXiv.
ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence
No summary available — see the abstract on arXiv.
Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection
No summary available — see the abstract on arXiv.
Enemray: Toward Capable Language Models for Hassaniya
No summary available — see the abstract on arXiv.
One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling
No summary available — see the abstract on arXiv.
Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization
No summary available — see the abstract on arXiv.
El Agente Potente: High-Throughput Agentic Atomistic Simulations
No summary available — see the abstract on arXiv.
Tackling Failure Modes of PINNs and PIKANs Using Conflict-Free Gradients
No summary available — see the abstract on arXiv.
A Responsive Present, a Shared Past, a Social Other: Teens' Overreliance on Companion AI Chatbots
No summary available — see the abstract on arXiv.
Prescreening Point Defects in Semiconductors With Machine Learning
No summary available — see the abstract on arXiv.
LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions
No summary available — see the abstract on arXiv.
Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference
No summary available — see the abstract on arXiv.
RAIN: Region-Aware Inversion Network for Semantic Watermark Extraction
No summary available — see the abstract on arXiv.
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
No summary available — see the abstract on arXiv.
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
No summary available — see the abstract on arXiv.
One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs
No summary available — see the abstract on arXiv.
Semantic Fibers and Cross-Gram Interference: A Calculus of Safety Drift in Overcomplete Representations
No summary available — see the abstract on arXiv.
Domain Generalization for Smartphone-Based Human Activity Recognition: A Systematic Analysis of Components and Interactions
No summary available — see the abstract on arXiv.
GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems
No summary available — see the abstract on arXiv.
Interpolation Is Not Invariance: Pair Count Is Not Coverage in Transformation Audits
No summary available — see the abstract on arXiv.
AgentKV: Phase-Aware KV Eviction for Agentic LLMs
No summary available — see the abstract on arXiv.
SeqMaestro: From nucleotide sequences to biological hypotheses through interpretable machine learning
No summary available — see the abstract on arXiv.
PeerPen: AI-Assisted Writing for Online Mental Health Peer Support
No summary available — see the abstract on arXiv.
An explicit solution of the five-expert prediction PDE and the exact optimality set of COMB
No summary available — see the abstract on arXiv.
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
No summary available — see the abstract on arXiv.
Shapley Value Estimation for Multi-Site Data with Blockwise-Missing Features
No summary available — see the abstract on arXiv.
Neural-Network Solutions to Real-Space Charge Density and Generalization
No summary available — see the abstract on arXiv.
Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair
No summary available — see the abstract on arXiv.
Geometric Signatures of Conceptual Reorganization: A Counterfactual Embedding Framework for Detecting Scientific Revolutions
No summary available — see the abstract on arXiv.
Steady-State Convergence of Stochastic Approximation
No summary available — see the abstract on arXiv.
Linearized PINN with pretrained nonlinear layers
No summary available — see the abstract on arXiv.
Cross-Block Conditioning in Deep Boltzmann Machines for Statistical Data Fusion
No summary available — see the abstract on arXiv.
Cloud Workflow Scheduling Based on Graph Attention-Driven Hierarchical Reinforcement Learning
No summary available — see the abstract on arXiv.
CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models
No summary available — see the abstract on arXiv.
High-Probability Nash Regret for Decentralized Learning in Markov $\alpha$-Potential Games: Episodic and Fully Online Asynchronous Algorithms with Applications to Markov Congestion Games
No summary available — see the abstract on arXiv.
Geometric Flow enhanced Graph Coarsening
No summary available — see the abstract on arXiv.
Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation
No summary available — see the abstract on arXiv.
A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator
No summary available — see the abstract on arXiv.
HiGFRL: Hierarchical Graph Fusion-Driven Reinforcement Learning for Dependency-Aware Task Scheduling in Heterogeneous Cloud
No summary available — see the abstract on arXiv.
Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains
No summary available — see the abstract on arXiv.
Towards a knowledge-enhanced single-cell foundation model
No summary available — see the abstract on arXiv.
Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence
No summary available — see the abstract on arXiv.
MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
No summary available — see the abstract on arXiv.
LiftGCN: Efficient Energy-Preserving Graph Learning via Joukowski Spectral Lifting for Finite Element Stress Prediction
No summary available — see the abstract on arXiv.
Converting Sequenced Fuzzy Cognitive Maps to Causal Virtual Worlds with Large Video Generators
No summary available — see the abstract on arXiv.
ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents
No summary available — see the abstract on arXiv.
Biomedical Reference Generation Remains Unreliable across 26 Large Language Models
No summary available — see the abstract on arXiv.
Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling
No summary available — see the abstract on arXiv.
Typhoon ASR Streaming: Steerable Low-Latency Thai Speech Recognition with Real-Time Shallow Fusion
No summary available — see the abstract on arXiv.
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
No summary available — see the abstract on arXiv.
Beyond Depth and Width: The Information-Slack Dilemma in Streaming Test-Time Compute
No summary available — see the abstract on arXiv.
Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
No summary available — see the abstract on arXiv.
HGTO: A Unified Graph-Based Physics-Informed Formulation for Structural Topology Optimization
No summary available — see the abstract on arXiv.
IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies
No summary available — see the abstract on arXiv.
ABSOL: Aggregated Bayesian Subsampling Orchestrated with LLMs
No summary available — see the abstract on arXiv.
CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems
No summary available — see the abstract on arXiv.
Semantic-TVM: Structure-Preserving Trustworthy Virtual Memory for Memory-Augmented and Tool-Using Agents
No summary available — see the abstract on arXiv.
Overflip: Repetition-Induced Label Flips in Guardrail Models
No summary available — see the abstract on arXiv.
Steering Generative Robot Policies with Lexicographic Preferences
No summary available — see the abstract on arXiv.
Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomedical ML
No summary available — see the abstract on arXiv.
PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift
No summary available — see the abstract on arXiv.
Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries
No summary available — see the abstract on arXiv.
SALUTE: Benchmarking and Adapting LLMs for the Defense Domain
No summary available — see the abstract on arXiv.
TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals
No summary available — see the abstract on arXiv.
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
No summary available — see the abstract on arXiv.
Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache
No summary available — see the abstract on arXiv.
Horizon-specific Expert Fusion for Photovoltaic Power Forecasting
No summary available — see the abstract on arXiv.
MoARa: Module-Aware Rank Allocation and Structure-Preserving Decomposition for Low-Rank LLM Pre-training
No summary available — see the abstract on arXiv.
The average-farmer illusion in language-model simulations of agricultural decisions
No summary available — see the abstract on arXiv.
SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing
No summary available — see the abstract on arXiv.
Data Attribution at Scale via Influence Matrix Estimation
No summary available — see the abstract on arXiv.
Mirror, Mirror on the Wall: Prompt Echoing in Small Instruct Language Models
No summary available — see the abstract on arXiv.
Personalizing Personal Health Interfaces: Co-Design with Generative AI
No summary available — see the abstract on arXiv.
Structured Features Overfit Where Random Features Grok
No summary available — see the abstract on arXiv.
Zero-SNR Analyticity of the Scalar MMSE Is Equivalent to Gaussianity
No summary available — see the abstract on arXiv.
Ensemble Complexity in Photovoltaic Forecasting
No summary available — see the abstract on arXiv.
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
No summary available — see the abstract on arXiv.
BusMA: A Bus Communication Substrate for Multi-Agent Systems
No summary available — see the abstract on arXiv.
Bridging the Gap in ECG-Based Emotion Recognition: A Unified Evaluation of Deep Learning Models
No summary available — see the abstract on arXiv.
What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track
No summary available — see the abstract on arXiv.
Sensory Precision Inference for Multimodal Arbitration under Uncertainty
No summary available — see the abstract on arXiv.
Salesforce Koa: An Enterprise Language Model for Agentic Tool Use
No summary available — see the abstract on arXiv.
Rethinking Procedural Audio Pre-training: Source Scaling and Objective Adaptation
No summary available — see the abstract on arXiv.
DA-DLM: Explicitly Modeling Token Dependencies in Diffusion Language Models
No summary available — see the abstract on arXiv.
Branched Optimal Transport Amortization
No summary available — see the abstract on arXiv.
Ensemble-Conditioned Molecular Design
No summary available — see the abstract on arXiv.
Enabling Creative Exploration for Vibe Design Agents
No summary available — see the abstract on arXiv.
Translating the Translator: Decomposing the Cost of English-Forced Inter-Agent Communication
No summary available — see the abstract on arXiv.
Task-Distribution-Aware Counterweight Synthesis and Constrained Co-Design for Serial Manipulators
No summary available — see the abstract on arXiv.
$\mathbb{SL}(n)$ Representation Learning: An Intrinsic Mixed-Curvature Space with Higher Curvature Capacities and Deeper Order-Aware Composition
No summary available — see the abstract on arXiv.
Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context
No summary available — see the abstract on arXiv.
ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models
No summary available — see the abstract on arXiv.
Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation
No summary available — see the abstract on arXiv.
OpenAl4S: Code as Action, Science as Sessions
No summary available — see the abstract on arXiv.
ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
No summary available — see the abstract on arXiv.
Physics Informed Neural Network model for the dynamical study of Abdominal Aortic Aneurysm
No summary available — see the abstract on arXiv.
When the Wrong Key Wins: Understanding and Detecting Hallucinations in LLMs
No summary available — see the abstract on arXiv.
Eigenvalue-Decomposition Cost Denoising as an Alternative to Predict-then-Optimize for Shortest-Path Problems
No summary available — see the abstract on arXiv.
Legislating World-Model-Based Planning with Legal Reasoning
No summary available — see the abstract on arXiv.
DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?
No summary available — see the abstract on arXiv.
Refinement-based Flow Policy Optimization
No summary available — see the abstract on arXiv.
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
No summary available — see the abstract on arXiv.
Omni-Streaming Thinking
No summary available — see the abstract on arXiv.
Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes
No summary available — see the abstract on arXiv.
woma: a real-time foundation model and its fine-tuned models for endoscopy
No summary available — see the abstract on arXiv.
AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference
No summary available — see the abstract on arXiv.
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
No summary available — see the abstract on arXiv.
SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering
No summary available — see the abstract on arXiv.
Improving Mathematical Reasoning Capabilities in Large Language Models via Reasoning Process Error Classification
No summary available — see the abstract on arXiv.
Multi-source Transfer Learning of Time Series with a Shapelet-based Distance Measure
No summary available — see the abstract on arXiv.
PACE: Progressive Angular-to-Norm Contrastive Embedding
No summary available — see the abstract on arXiv.
T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning With Dynamic Routing
No summary available — see the abstract on arXiv.
EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse
No summary available — see the abstract on arXiv.
CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search
No summary available — see the abstract on arXiv.
Nearly Minimax-Optimal Regret for Linear Contextual Bandits with Arbitrary Adaptive Action Sets
No summary available — see the abstract on arXiv.
STHMoE: Hypergraph-Enhanced Heterogeneous Dependency Coordination for LLM-Based Urban Traffic Data Forecasting
No summary available — see the abstract on arXiv.
Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models
No summary available — see the abstract on arXiv.
Low-Dimensional Embeddings for Gaussian Kernels on Manifolds
No summary available — see the abstract on arXiv.
Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models
No summary available — see the abstract on arXiv.
VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries
No summary available — see the abstract on arXiv.
MUSE: A Theory-Harnessed Story Engine for Vibe Narrativizing
No summary available — see the abstract on arXiv.
What Limits Us? Analyzing Self-Reported Limitations in NLP Research
No summary available — see the abstract on arXiv.
Convergence rates for generative drifting flows: fixed-scale obstructions and multihead acceleration
No summary available — see the abstract on arXiv.
Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models
No summary available — see the abstract on arXiv.
Interpreting hierarchical organisation of speaker embeddings
No summary available — see the abstract on arXiv.
TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models
No summary available — see the abstract on arXiv.
Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election
No summary available — see the abstract on arXiv.
Failure-Guided Co-Evolution of Prompts and Training Data
No summary available — see the abstract on arXiv.
Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception
No summary available — see the abstract on arXiv.
From Ideas to Actions: A Public-Data Decision-Support Toolchain Across the Venture Lifecycle
No summary available — see the abstract on arXiv.
MAST: Label-Efficient, Robust, and Generalizable Sound Detection for Biodiversity Monitoring via Masked Audio Pretraining and Self-Training
No summary available — see the abstract on arXiv.
Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders
No summary available — see the abstract on arXiv.
CWM: Controllable White-Box Meta-Prompting for Adaptive Retrieval-Augmented Generation and Reasoning Ability
No summary available — see the abstract on arXiv.
ProtoGuide: Prototype-Driven Guidance for Class-Conditional Graph Generation
No summary available — see the abstract on arXiv.
Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generation in ESG Domain
No summary available — see the abstract on arXiv.
Bandits with Probing: Optimal Regret and the Limits of Winner Feedback
No summary available — see the abstract on arXiv.
Conformal Individual Treatment Effect Estimation under Networked Interference
No summary available — see the abstract on arXiv.
BioDCASE: Active Learning for Bioacoustics
No summary available — see the abstract on arXiv.
Improving the Last-Iterate Guarantees of Anytime Algorithms for Stochastic Monotone Variational Inequalities
No summary available — see the abstract on arXiv.
Learning CNF Formulas from Uniform Random Solutions: Near-Tight Sample Complexity for Valiant's Algorithm
No summary available — see the abstract on arXiv.
Draining Fictitious Knots: Restoring Distance-Awareness Guarantees for High-Dimensional Spline Networks
No summary available — see the abstract on arXiv.
Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs)
No summary available — see the abstract on arXiv.
Impute-EM: Native Mixed-State Diffusion Models for Heterogeneous Data Imputation
No summary available — see the abstract on arXiv.
Math for AI safety: an invitation for mathematicians
No summary available — see the abstract on arXiv.
ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment
No summary available — see the abstract on arXiv.
Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
No summary available — see the abstract on arXiv.
Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
No summary available — see the abstract on arXiv.
Admissable: Training Reinforcement Learning Agents against Adversarial Missingness
No summary available — see the abstract on arXiv.
When Correlations Mislead: Confounder-Aware Multi-View Urban Region Representation Learning
No summary available — see the abstract on arXiv.
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
No summary available — see the abstract on arXiv.
Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation
No summary available — see the abstract on arXiv.
The Universe of Universes: Benefit Yield Functions, Implosion Thresholds, and Infrastructure-Aware Optimization in Multi-LLM Systems
No summary available — see the abstract on arXiv.
Evaluation Metrics for Safe Reinforcement Learning
No summary available — see the abstract on arXiv.
Clean Scores, Buried Evidence, and Confident Wrong: A Receipt-Based Audit of Frontier Agentic QA
No summary available — see the abstract on arXiv.
Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs
No summary available — see the abstract on arXiv.
Concept-Grounded Reasoning with Prompt-Driven Localization for Interpretable Structured Report Generation
No summary available — see the abstract on arXiv.
Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models
No summary available — see the abstract on arXiv.
Parameter-Efficient Adaptation of Pretrained Language Models for Time-Series Forecasting
No summary available — see the abstract on arXiv.
End-to-End Cell Detection via Instance-aware Graph Modeling
No summary available — see the abstract on arXiv.
ReLU Neural Network Approximation to Smooth Functional Operator: Dimensional Decay and Error Analysis
No summary available — see the abstract on arXiv.
MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving
No summary available — see the abstract on arXiv.
Robust and Efficient Communication for Multi-Agent Learning
No summary available — see the abstract on arXiv.
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
No summary available — see the abstract on arXiv.
SlopShape: Identifying AI-Generated Commercial Web Content
No summary available — see the abstract on arXiv.
Representing Clinical Conditions on Vital Signs from Healthy Individuals using Latent Modeling
No summary available — see the abstract on arXiv.
Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
No summary available — see the abstract on arXiv.
IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective
No summary available — see the abstract on arXiv.
A Game-Theoretic Framework for Incentive-Compatible AI training Under Renewable-Energy Constraints
No summary available — see the abstract on arXiv.
CodeTS: Verifiable Text-to-Time Series Generation via Executable Code
No summary available — see the abstract on arXiv.
SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution
No summary available — see the abstract on arXiv.
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary
No summary available — see the abstract on arXiv.
Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning
No summary available — see the abstract on arXiv.
Can AI systems have free will?
No summary available — see the abstract on arXiv.
MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
No summary available — see the abstract on arXiv.
Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents
No summary available — see the abstract on arXiv.
A Conservative OCR-Enabled Workflow for R214 Sodium Screening of South African Packaged Foods
No summary available — see the abstract on arXiv.
Single-condition neural solvers encode transferable response spaces for parametric differential equations
No summary available — see the abstract on arXiv.
On the role of the tokenizer in ECG transformer models
No summary available — see the abstract on arXiv.
Graph Matching Relaxations and Amortization for Supervised Graph Prediction
No summary available — see the abstract on arXiv.
Online local learning for generative thermodynamic computing
No summary available — see the abstract on arXiv.
Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation
No summary available — see the abstract on arXiv.
HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments
No summary available — see the abstract on arXiv.
Spook the Machine: Gamified Exploration of Human Imagination of Machine Fear
No summary available — see the abstract on arXiv.
Temperature Fragility and the Conditional Benefits of Truncation Sampling
No summary available — see the abstract on arXiv.
GSLAD: Prototype-Regularized Graph Structure Learning for Multivariate Time Series Anomaly Detection
No summary available — see the abstract on arXiv.
Data-driven Prediction of Satellite-observed Avalanche Activity from Snowpack Simulations
No summary available — see the abstract on arXiv.
Rotation-Based Subspace Tracking for Robust Kernel PCA on Streaming Data
No summary available — see the abstract on arXiv.
The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow?
No summary available — see the abstract on arXiv.
Same path, different: a mechanistic comparison of looped and stacked transformer encoders on 12-lead ECG
No summary available — see the abstract on arXiv.
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
No summary available — see the abstract on arXiv.
Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku
No summary available — see the abstract on arXiv.
Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models
No summary available — see the abstract on arXiv.
End-to-End Verifiable and Robust Federated Learning
No summary available — see the abstract on arXiv.
Psychosis involves a deficit of information compression in connected speech
No summary available — see the abstract on arXiv.
Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting
No summary available — see the abstract on arXiv.
Beyond Noise: Understanding and Overcoming Temperature Effects in Analog DNN Inference
No summary available — see the abstract on arXiv.
To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
No summary available — see the abstract on arXiv.
Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA
No summary available — see the abstract on arXiv.
The Misery of Mechanistic Interpretability: A Formal Perspective
No summary available — see the abstract on arXiv.
Strong and Compact Policies for Submodular Markov Decision Processes via LP-Based Submodular Orienteering
No summary available — see the abstract on arXiv.
Specifying Reward Functions for RL Without Environment Sampling
No summary available — see the abstract on arXiv.
The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits
No summary available — see the abstract on arXiv.
Bayesian Optimisation Using Product-of-Experts Gaussian Process Models with Uncertainty Calibration
No summary available — see the abstract on arXiv.
Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction
No summary available — see the abstract on arXiv.
Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation
No summary available — see the abstract on arXiv.
PIVOT: Physics-Grounded Verification for AI-Generated Audio-Video Detection
No summary available — see the abstract on arXiv.
Big Brains and Changing Environments: Cause or Consequence?
No summary available — see the abstract on arXiv.
GRIN+: Towards Fast Yet Effective Machine Unlearning for Imbalanced Medical Data
No summary available — see the abstract on arXiv.
Self-Evolving Memory for Generative Recommendation
No summary available — see the abstract on arXiv.
A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation
No summary available — see the abstract on arXiv.
VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
No summary available — see the abstract on arXiv.
Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection
No summary available — see the abstract on arXiv.
Diversified and Perceptible Counterfactual Examples Leveraging Expert Knowledge
No summary available — see the abstract on arXiv.
Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints
No summary available — see the abstract on arXiv.
IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English
No summary available — see the abstract on arXiv.
Where to Compute and How to Interact: Operator-Readable Adaptation with Gauge-Aware Transport
No summary available — see the abstract on arXiv.
Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use
No summary available — see the abstract on arXiv.
FedLTLib: A Comprehensive Benchmark for Federated Long-Tail Learning
No summary available — see the abstract on arXiv.
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
No summary available — see the abstract on arXiv.
Potential of Artificial Intelligence Algorithms for Identification of Relevant Diagnostic and Prognostic Biomarkers of Early-Stage Liver Cancer
No summary available — see the abstract on arXiv.
Human-Grounded Calibration for Long-Text Image-Text Congruence in Vision-Language Models
No summary available — see the abstract on arXiv.
Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching
No summary available — see the abstract on arXiv.
Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs
No summary available — see the abstract on arXiv.
Predictive Likelihood Ratios for Language Model Watermark Detection
No summary available — see the abstract on arXiv.
Kaininja: Extending Native 3D Generators to the Part Level
No summary available — see the abstract on arXiv.
CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense
No summary available — see the abstract on arXiv.
CIDERS: Cloud-Edge LLM Collaborative Learning via Accelerating Personalized Bilevel Optimization
No summary available — see the abstract on arXiv.
Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
No summary available — see the abstract on arXiv.
Benchmarking Intra-Patient 3D Deformable Multimodal Image Registration
No summary available — see the abstract on arXiv.
Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
No summary available — see the abstract on arXiv.
Projection-Free Multi-level Algorithms for Stochastic Constrained Compositional Optimization
No summary available — see the abstract on arXiv.
Scalability and Performance Evaluation of Federated Learning Frameworks: A Comparative Analysis
No summary available — see the abstract on arXiv.
RESKILL: Explicit Failure Attribution and Structured Repair for Interactive Language Agents
No summary available — see the abstract on arXiv.
EEG-Xplain: Decoding Neural Black-Boxes of EEG Foundation Models
No summary available — see the abstract on arXiv.
NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities
No summary available — see the abstract on arXiv.
More Than Just Access: Generative AI as Communication Intermediary for Blind and Low-Vision Users
No summary available — see the abstract on arXiv.
Backward SDEs-based Diffusion for Physics-Constrained Generation
No summary available — see the abstract on arXiv.
Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction
No summary available — see the abstract on arXiv.
New Conditions for Philosophers to Catch the Wave of Citizen Deliberation in the Age of Artificial Intelligence in advance
No summary available — see the abstract on arXiv.
Predicting build orientation for SLM dental parts: a comparison of rotation representations and direct vector regression
No summary available — see the abstract on arXiv.
Knowledge-Enriched Structured EHR Features for 30-Day Hospital Readmission Prediction on MIMIC-IV
No summary available — see the abstract on arXiv.
Assembling the CREW: A Collaborative Multi-agent Reinforcement Learning Framework for Automated Related Work Generation
No summary available — see the abstract on arXiv.
Data storytelling meets interpretable machine learning: Decoding AI decisions for non-experts without revealing sensitive data and model details
No summary available — see the abstract on arXiv.
Solving Finite-sum Coupled Compositional Optimization via Multi-block-Single-probe Estimator
No summary available — see the abstract on arXiv.
Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands
No summary available — see the abstract on arXiv.
Are LLMs Good Financial User Simulators? A Preliminary Study
No summary available — see the abstract on arXiv.
A Language-Guided Multimodal Foundation Model for Zero-Shot and Multi-Task Brain Signal Analysis
No summary available — see the abstract on arXiv.
Merging the Knowledge of LLMs for Automatic Speech Recognition
No summary available — see the abstract on arXiv.
Design of a Deep Learning Credit Risk Early Warning System Integrating Multi-source Heterogeneous Data
No summary available — see the abstract on arXiv.
Look Before You Leap: Factual Decoding with Internal Attribution Signals
No summary available — see the abstract on arXiv.
Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR
No summary available — see the abstract on arXiv.
Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models
No summary available — see the abstract on arXiv.
Sylvas: Synergistic Learning Value based Device Scheduling in Federated Continual Learning
No summary available — see the abstract on arXiv.
Event-Native Symbolic-Temporal Spike Encoding Framework for Heterogeneous Cyber Streams
No summary available — see the abstract on arXiv.
Transfer Learning for Socioeconomic Estimation in Forced-Displacement Settings
No summary available — see the abstract on arXiv.
EvoOntology: A Self-Evolving Ontology Layer for Data Agents
No summary available — see the abstract on arXiv.
MoveBench: A Benchmark for Global-Scale Wildlife Movement Forecasting
No summary available — see the abstract on arXiv.
When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control
No summary available — see the abstract on arXiv.
Learning under Target Shift: Optimal Density Ratio Estimation and Importance-Weighted Regression
No summary available — see the abstract on arXiv.
KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
No summary available — see the abstract on arXiv.
Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
No summary available — see the abstract on arXiv.
When Should a World Model Move? Loss-Conditioned State Execution
No summary available — see the abstract on arXiv.
Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
No summary available — see the abstract on arXiv.
Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking
No summary available — see the abstract on arXiv.
Atria Dawn: The Dawn of Agentic Superintelligence
No summary available — see the abstract on arXiv.
AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery
No summary available — see the abstract on arXiv.
Sharp Rates and a One-Line Correction for Spectral Representation Learning
No summary available — see the abstract on arXiv.
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering
No summary available — see the abstract on arXiv.
Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression
No summary available — see the abstract on arXiv.
Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence
No summary available — see the abstract on arXiv.
Before You Poll with LLMs: A Deliberative Diagnostic Framework
No summary available — see the abstract on arXiv.
Learning to Coach for Experiential Learning
No summary available — see the abstract on arXiv.
K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
No summary available — see the abstract on arXiv.
Task-Directed Residual AddUNet:Perfect-Reconstruction Routing for Full-Rate Representations
No summary available — see the abstract on arXiv.
LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction
No summary available — see the abstract on arXiv.
LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys
No summary available — see the abstract on arXiv.
Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
No summary available — see the abstract on arXiv.
Thin-shell stability of Gaussian cooling: logconcave sampling with sesteric complexity from a cold start
No summary available — see the abstract on arXiv.
Privacy-enhanced federated learning via asynchronous aggregation and local differential perturbation
No summary available — see the abstract on arXiv.
Inoculation Midtraining with Learned Neologisms
No summary available — see the abstract on arXiv.
Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI
No summary available — see the abstract on arXiv.
Quenched Ensemble Sampling
No summary available — see the abstract on arXiv.
Bridging Control, Inference, Transport, and Thermodynamics: From Theory to Applications in Learning
No summary available — see the abstract on arXiv.
Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning
No summary available — see the abstract on arXiv.
SlipSense: Multimodal Tactile Learning for Low-Latency and Generalized Slip Detection
No summary available — see the abstract on arXiv.
Safe Meta-Reinforcement Learning via Information Space Reachability
No summary available — see the abstract on arXiv.
Pilot Early, Commit Late: A Real-Options Model of Enterprise AI Adoption under Rapid Technological Progress
No summary available — see the abstract on arXiv.
Recurrent GraphNeural NetworkswithSet-BasedAggregation
No summary available — see the abstract on arXiv.
HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
No summary available — see the abstract on arXiv.
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale
No summary available — see the abstract on arXiv.
Privacy-Aligned Personalized Federated Learning with Compact Adaptation and Variable-Length Gaussian Communication
No summary available — see the abstract on arXiv.
Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
No summary available — see the abstract on arXiv.
Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
No summary available — see the abstract on arXiv.
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence
No summary available — see the abstract on arXiv.
Disentangling Representation Evolution in Transformers through Directional Decomposition
No summary available — see the abstract on arXiv.
A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models
No summary available — see the abstract on arXiv.
The Router Within: Eliciting Native Skill Routing from a Frozen LLM
No summary available — see the abstract on arXiv.
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
No summary available — see the abstract on arXiv.
Bellman Policy Optimization
No summary available — see the abstract on arXiv.
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
No summary available — see the abstract on arXiv.
Applications 21
Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB
Analytical pipelines that use LLMs over both database tables and unstructured documents usually mean coordinating separate data systems, moving data between them, and handling details like context management by hand. FlockMTL is a database extension, shown in DuckDB, that brings LLM calls and retrieval-augmented generation (RAG) directly into SQL through model-driven scalar and aggregate functions that can be chained over rows. Borrowing from the relational model, it automatically applies cost-based optimizations such as batching and caching. It also adds PROMPT and MODEL as first-class schema objects alongside TABLE, so queries stay independent of specific prompts and model resources.
Natural-Language to SysMLv2 Translation via Conformance-Driven Iterative Refinement
Model-Based Systems Engineering (MBSE) increasingly uses the textual SysMLv2 language, but LLM-generated models are only usable if industrial modeling tools accept them, not merely if they parse. The framework places a production SysMLv2 conformance checker inside a generate-check-repair loop and feeds its deterministic diagnostics back to the LLM until no conformance errors remain. Across all 151 SysMBench prompts and four LLM backends, single-shot generation passed conformance 51.16% of the time, while the iterative approach reached 100% conformance.
When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels
LLMs are increasingly used as synthetic consumer survey panels, but aggregate validation metrics hide failures such as compressed variance, flipped coefficient signs, and subgroup errors of 10 to 30 percentage points. The authors split synthetic-panel bias into covariate shift and concept shift and build diagnostics with clear thresholds for deciding whether to trust, correct, or abandon the data. For correction they use a doubly robust augmented inverse probability weighting (AIPW) estimator that needs only 50 to 300 human calibration responses. The decision rule was right in all 180 simulation runs and cut naive bias by 92.9 to 99.6% on the American National Election Study, and on the Twin-2K-500 pricing dataset it correctly trusted full-sample estimates and corrected subgroup estimates, reducing bias by 83 to 94%.
A Lifecycle Cost Analysis of Smart-Contract-Coordinated Federated Learning Marketplaces
Blockchain-based federated learning (FL) marketplaces use smart contracts to coordinate training among parties that do not trust each other, but cost studies usually measure isolated blockchain operations rather than the full marketplace lifecycle. The authors measure the gas used by every operation in a DAO-governed marketplace, run ablations that isolate the effect of on-chain coordination and IPFS storage on training, and build an analytical model of how deployment costs amortize. One training task costs about 3.8 million gas units per hired trainer, and the average cost per round reaches its amortization knee after about 20 communication rounds. Accuracy matches conventional FL, and costs amortize because recurring costs are one to two orders of magnitude smaller than fixed deployment costs, not because on-chain operations are cheap.
Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging
California now requires prosecutors to make race-blind charging decisions from case documents with race-related proxies redacted, and the authors validate bc2, their open-source LLM-based redaction tool used in more than 119,000 real cases in 2025. Using nearly 5,000 police reports from jurisdictions across the United States, they ask two questions: whether bc2 faithfully implements the mandate, and whether the mandate itself achieves race-blind review. Under a strict document-level measure, the latest bc2 complies fully on 96.7% of narratives, beating earlier versions and leading open-source redaction methods. The mandate misses key proxies such as location, and redacting these extra proxies removes 43.1% of the predictive signal that remains after mandated redaction.
16 more specialized papers
- Approximating neutron-star radii using gravitational-wave only measurements with symbolic regression Micha{\l} Bejger
- Physically Aware Radiomics Without Interpolation: Disentangling Voxel Geometry and Signal Modification in CT and MRI David Corral Fontecha, Juan Miranda Bautista, Pablo Menendez Fern\'andez-Miranda et al.
- Exploring new directions in enhancing the ACTS parameter optimization suite Chance LaVoie, Qi Bin Lei, Rocky Bala Garg et al.
- Multilingual Agent System for Inclusive Wildfire Evacuation Guidance Shruti Kulkarni, Lynn Tong, Aditi Namboodiripad et al.
- Land Art as a Big-Data Climate Sensor Alev Cinbarci, Sean Kalaycioglu
- Early Prediction of Satellite Collision Probability Using a Hybrid TCN-Transformer Model for a CDM-Based Conjunction Analysis Framework Rabia T\"uylek Tok, Burak Ya\u{g}l{\i}o\u{g}lu, Enes Da\u{g} et al.
- Evaluating LLM-Generated Rules for Heart Disease Prediction Feisal Alaswad, Batoul Aljaddouh, Maher Alrahhal et al.
- PRISM-UDE: Physics-Regularized Iterative Symbolic Modeling of 3nm FinFETs via Universal Differential Equation Pranavanath Balamurali, Prathamesh Dinesh Joshi, Raj Abhijit Dandekar et al.
- Scalable partial information decomposition for symptom networks via supervised embeddings Cillian Hourican, Eric Dignum, Rick Quax et al.
- From objective discovery to prediction of global ocean eco-provinces: A pathway for trustworthy learning Makayla McDevitt, Maike Sonnewald, Stephanie Dutkiewicz
- Chemical and geometric representation fidelity improves drug--target affinity prediction Yixiao Li, Yining Qian, Yefan Chen et al.
- Occlusal Geometry in Closed Form for Orthodontic Report Generation Ajo Babu George, Govind Arun, Sidharth N Krishna et al.
- Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation Ajo Babu George, Govind Arun, Sidharth N Krishna et al.
- Calibrating subgrid parametrizations of single-column ocean models via simulation-based inference Luben M. C. Cabezas, Sacha Wendling, Aur\`ele Gallard et al.
- Machine learning-assisted calibration of Agent-based Models: surrogate-based optimization with Genetic Algorithm and Particle Swarm Optimization Duguma Yeshitla Habtemariam, Jihwan Lee
- Forward-Facing Near-Infrared Adds Little to Colour for Farm-Machinery Traversability: A Site-Disjoint Evaluation of Sensor-Dependent Spatial Leakage Sungwoo Kang
Other 10
Generalization Can Emerge in Tabular Foundation Models From a Single Table
Tabular foundation models predict labels in context from example rows without updating weights, and they are widely believed to need pre-training on large synthetic priors, as in TabPFN, or on many real datasets, as in TabDPT. The authors challenge this view, showing that simple self-supervised pre-training on a single real table can transfer strongly across heterogeneous benchmarks. By systematically pre-training and evaluating on many diverse datasets, they find that the key factor is the number and quality of tasks that can be constructed from the pre-training data.
From Semantic to Token Communication: The Next Paradigm for Large-Model-Driven 6G Intelligent Connectivity
Sixth-generation (6G) networks are moving from bit delivery toward meaning-aware communication, but semantic communication (SemCom) driven by large models (LMs) is fragmented: its representations are tied to particular modalities, models, or tasks, with no universal unit like the bit. This survey argues that tokens should serve as that unit. It points out that unified multimodal models already encode text, images, audio, video, and robot actions as tokens, and that distributed inference already generates token-level traffic. The survey reviews LM-driven SemCom and then covers token communication (TokenCom), including its transmission techniques, its use for LM services and for embodied and agentic systems, and open challenges.
The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech
Reference-free speech quality predictors such as UTMOS, DNSMOS, and SCOREQ are the standard automatic evaluators for text-to-speech (TTS) and are increasingly used as rewards in preference optimization. The authors tested them pairwise against listener preferences on six human-rated corpora. The predictors agree with listeners when one clip has audible defects, but once both clips are clean no single predictor reliably picks the preferred one, and several do worse than simply choosing the longer clip. A calibrated combination of complementary signals is the best evaluator, even an equal-weighted ensemble works as a post-training reward, and optimizing any single metric leads to reward hacking that makes held-out judges and human listening tests worse.
Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning
Multilingual speech recognition models like Whisper can transcribe low-resource languages without language-specific training, but they are expensive to deploy. The authors systematically evaluate token merging, which combines redundant features to shorten sequences at inference time without retraining. Their tests cover sixteen languages and three Whisper model sizes, plus low-resource languages after DoRA fine-tuning. They find that token merging improves computational efficiency with almost no loss in transcription accuracy for most languages and model sizes, and that it keeps working after fine-tuning.
Do Tabular Foundation Models Still Need Feature Engineering?
Tabular foundation models (TFMs) are pretrained on many tabular datasets and applied through in-context learning, which raises the question of whether manual feature engineering still helps. A controlled study tests a wide range of feature engineering techniques across several versions of two major TFM families on TabArena benchmark datasets. Gains from feature engineering are concentrated in earlier model generations and become negligible for the strongest models. Adding in-context information from related datasets still improves performance, which suggests that stronger TFMs gain more from extra task-relevant context than from re-representing existing inputs.
5 more specialized papers
- QSTAR: Quantum Selective Transfer with Adaptive Routing Saim Rehman, Nouhaila Innan, Muhammad Shafique
- A Cross Community Agenda for Speech AI Maria Teleki, Kimi Wenzel, Anna Seo Gyeong Choi et al.
- A derivative-fidelity failure mode in physics-informed neural networks: strengthened benchmark evidence from function-value training Koji Koyamada
- GradRepair-ODE: Certified Gradient Repair for Neural ODE Training Ziqian Bi, Xin Liang Chia
- From Masking to Merging: Rethinking SpecAugment for Efficient Audio Spectrogram Transformer Minhee Park, Hyowon Ahn, Chanwoo Kim
Vision 8
Sampling headroom is not selection gain: a compute-value audit of test-time scaling for video world models
Test-time scaling (TTS) helps only if extra samples contain better candidates and the system can reliably pick them out, and video world models can fail at the second step. The Compute-Value Audit (CVA) checks in sequence whether extra sampling creates opportunity, whether observable signals are reliable, whether acting on them helps, and whether the gain exceeds the full cost of generation and verification. On 192 Physics-IQ scenes, growing the pool from 4 to 16 candidates raised oracle quality by 9.23 IQ, but the Flow, Cycle and VideoReward selectors could not reliably capture that gain, and none of twelve adaptive-depth policies beat uniform compute across three generators. Positive results in a sparse PRM800K setting and with a privileged future signal show the pattern is not universal: extra samples pay off only when they lead to a reliable decision whose benefit survives the full compute cost.
7 more specialized papers
- SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation Sehyun Lee, Dahee Kwon, Damin Lee et al.
- Discovering and Preserving Category Correlation Knowledge via Adaptive Reciprocal Knowledge Distillation Dawen Jiang, Zhishu Shen, Zeyu Liu et al.
- Abstract-LoRA: Unlocking Single-Image Style Transfer through Targeted U-Net Block Training Xinglin Hu
- Interpretable Temporal Video Reasoning with EventGraph and EventField Durgendra Narayan Singh
- TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On Xueheng Li, Yong Liu, Xiaolong Fu et al.
- BEACON: Behavior and Appearance Control for Subject-Specific Video Generation Pokrzywa Baptiste, Nabyl Quignon, Yara Bahram et al.
- Frame-Synchronous Hand Gesture Detection by Projected Winding Order Amey Thakur
Agents 4
Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents
Shell-based agents do well at coding, but it is unclear whether a general shell also beats specialized tools for enterprise work that spans multiple applications, coordination with coworkers, and professional analysis. The authors compare five tool interfaces on TheAgentCompany and APEX-Agents using Opus-4.8 and GPT-5.5: typed tools, typed tools plus bash, bash alone, bash with persistent agent-synthesized tools, and programmatic tool calling (PTC) restricted to a typed tool catalog. Bash alone beat typed tools by 21.8 to 24.5 percentage points on TheAgentCompany and by 4.8 to 7.4 points on APEX-Agents, while using 19 to 72% fewer tokens, and adding typed tools or tool synthesis to bash brought no detectable gain. PTC used fewer tokens than direct typed calls with similar scores but generally trailed bash alone, so the authors recommend bash when execution can be isolated and PTC when compliance requires a fixed tool catalog.
BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
For LLM agents running locally, memory, prefill latency, and cache growth tightly limit how many input tokens each call can afford. BudgetBench is a protocol and open-source harness that makes this per-call token budget the variable under test. It sweeps budgets from 2K to 32K tokens with the model and decoding held fixed, records quality, budget use, latency, and budget-violation rates, and lets different memory strategies be swapped in. Pilot studies with a local qwen2.5:1.5b on SWE-bench Verified and LongBench v2, a hosted Qwen3 30B-A3B replication, and a 500-item LongMemEval study expose budget-compliance failures and non-monotonic quality curves that single-budget evaluation hides, while leaving unresolved whether budgeted memory beats full context.
From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration
If groups of LLM agents are used to simulate human groups, they should succeed or fail through human-like deliberation, not just produce similar outcomes. The authors compare human group chats with matched LLM deliberation traces on Wason-style deductive reasoning, then test whether the same process patterns hold for analogical, abductive and analytical tasks. Both humans and LLMs show the same assembly bonus asymmetry, where discussion improves the average member more often than the best initial member, but LLM groups follow majorities more often, surface less unique information and converge earlier than humans. Interventions borrowed from human group-decision research gave modest gains but did not remove this coordination bottleneck.
1 more specialized paper
Large Language Models 4
Towards Optimizing SQL Generation via LLM Routing
Highly capable large language models (LLMs) write accurate SQL from natural-language questions, but they add latency and cost on simple queries that cheaper models could handle. The authors present what they describe as the first LLM routing approach for Text-to-SQL. For each query, a score-based or a classification-based router picks the cheapest model likely to produce correct SQL, and both routers are designed to be easy to train and fast at inference. On the BIRD dataset, the routers match the accuracy of the most capable LLM while reducing cost, with a practical and explainable trade-off between accuracy and cost.
Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories
Learned prompt compressors such as LLMLingua and Selective Context need auxiliary language models and are non-deterministic. The authors instead test how far a training-free, deterministic, CPU-only pipeline of classical lexical transformations can go. They combine eleven toggleable steps, including stopword removal, filler-phrase deletion, part-of-speech pruning, lemmatization, and WordNet synonym shortening, into fifteen configurations tested on 1,242 prompts from sources such as Dolly-15k, MMLU, and GSM8K, producing 18,630 paired GPT-4o-mini completions. The most aggressive configuration cuts tokens by 40.3% on average at a BERTScore-F1 of 0.876 against outputs from the original prompts, stopword removal alone gives a 29.6% cut at 0.913, and commonsense reasoning consistently breaks under aggressive compression.
LLMs or Naive Bayes? Old Gems or New Ways
Large language models (LLMs) raise the question of whether classical classifiers like Naive Bayes (NB) should be retired, so the authors benchmark Complement Naive Bayes against zero-shot and few-shot LLMs from four model families, ranging from 27B parameters to a 1T-parameter mixture-of-experts, on text classification. LLMs win only when there is no labeled data, and even that edge is prone to contamination: on a sentiment task with low contamination, NB beats the zero-shot LLM 81.7% to 73.0%. With labels, NB reaches 89.1% on AG News, statistically tied with a zero-shot 27B LLM and ahead of a 397B frontier model, while batched GPU inference with a small LLM is 40-486x slower than NB on a CPU and uses roughly two orders of magnitude more energy per sample. NB matches LLMs at around 10^4 labels for topic classification, and the authors release a Kubernetes Helm operator that automates model selection with configurable thresholds.
Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference
Sparse long-context inference has to retrieve relevant tokens efficiently during both prefill and decode, but existing methods usually use a different retrieval strategy for each stage, so no single retrieval representation is reused throughout. Self-Indexing Attention is a training-free framework built on a shared transform-domain sign-magnitude representation, where the signs of the keys form a 1-bit token index that serves both grouped prefill selection and decode retrieval using fast bitwise operations. The same representation works with external KV-cache compression without needing separate indexer metadata. At 5% attention density it stays close to dense attention on LongBench and RULER while delivering up to 6.1x prefill and 10.3x decode attention-operator speedups, and experiments with TurboQuant and DeepSeekV4-Flash show it is compatible with low-bit KV-cache compression and pretrained sparse-attention indexers.
Multimodal 4
Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction
Benchmarks show vision-language models struggle with low-level manipulation reasoning, but an aggregate score does not reveal whether models fail to pick the right object part or fail to know what action that part needs. Eight models from three developers were asked what motion a robot should apply to 19 articulated objects. Under an open prompt they almost never answered 'push' (once in 64 cases where it was correct), because they described a different part, such as picking up a camera instead of pressing its button. Naming the target part raised action accuracy from 0.158–0.474 to 0.684–0.947 for every model, which points to part grounding rather than missing action knowledge as the dominant bottleneck; the authors also document two of their own measurement errors that surfaced only when compared against trivial constant baselines.
Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding
Multimodal large language models (MLLMs) cannot process every frame of a long video. Training-free, plug-and-play (PaP) keyframe selectors that work with any MLLM are the cheapest fix, but the five published in the past year were each evaluated on different benchmarks and models. The authors compare all five under one setup, using three MLLMs and three long-video understanding benchmarks. QAaF performed best in 13 of 15 aggregate settings and FOCUS ranked second overall, giving a common reference point for training-free keyframe selection.
(How) Do MLLMs Report Bistable Images Like Humans?
Bistable images like the duck-rabbit support two incompatible interpretations that humans report one at a time, and the authors test whether multimodal large language models (MLLMs) behave the same way. Using the LLaVA family on the classic duck-rabbit and on synthetic Visual Anagrams (to limit memorization), they measure whether visual cues and linguistic priors can bias the reported interpretation and whether responses commit to a single one. Both visual and linguistic manipulations shifted reports in human-consistent ways, while responses stayed predominantly exclusive. Mechanistically, the behavior traces to competing image-token representations, separate pathways for bottom-up and top-down influence, and a link between exclusive reporting and object-count encoding.
1 more specialized paper
Theory 4
4 more specialized papers
- Algorithmic Information Dynamics of Learning: A Certified, Differentiable Complexity Controller for Grokking Luan Ozelim, Hector Zenil
- Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies Ghurumuruhan Ganesan
- Planning as Dynamics Relaxation: Hippocampal Recurrent Network Realizes Optimal Goal-Directed Navigation Yuhang He, Junfeng Zuo, Tianhao Chu et al.
- An Evolutionary Computation Framework for Multi-Agent Q-Learning with Mean-Field Environmental Feedback Lichen Wang, Shijia Hua, Linjie Liu