Thursday, August 27, 2026

490 papers cs.AI · cs.LG · cs.CL ← 2026-08-262026-08-28 →

Jul Aug Sep

Highlights

Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

Highlight HF pick · 12▲Agents Guo Gan, Yilun Zhao, Cong Chen, Jinbiao Wei, Tingyu Song, Zheyuan Yang et al. GUI agents deployed on Android devices encounter runtime anomalies such as unexpected pop-ups and misused actions, yet existing benchmarks do not systematically test robustness against them. AnTrap injects dynamic perturbations into agent execution trajectories following a taxonomy of four layers (State, Thinking, Action, and Round) with ten fine-grained subcategories, using a construction pipeline that preserves task solvability while introducing realistic adversarial conditions. Evaluating 16 leading GUI models reveals universal vulnerability, with even the strongest models degrading significantly; GRPO training in original versus adversarial environments shows that single-step traps at the state and action layers are largely fixable by adversarial reinforcement learning, whereas deep contextual traps such as state deadlock persist and expose reasoning limitations that training in trap environments alone cannot resolve.

Android GUI agents routinely hit runtime anomalies in deployment, from ad pop-ups and frozen screens to their own grounding and action-type mistakes, yet existing benchmarks evaluate them either in clean dynamic environments or with static single-step perturbations. AnTrap extends AndroidWorld to 236 tasks and injects solvability-preserving traps into live execution trajectories, organized by a four-layer STAR taxonomy (State, Thinking, Action, Round) with ten subcategories, then uses per-trap GRPO training to separate anomalies that are learnable from the environment from those bottlenecked on reasoning.

  • Traps are injected at the agent–environment interface without touching the agent's inference logic: State-layer traps trigger pop-ups via a TrapOverlay APK or mask 10–30% of the screen, Thinking-layer traps swap in a stale or label-tampered screenshot only at decision time (so the agent's misled output is logged but the fake observation is not), Action-layer traps offset coordinates by 20–50 px or remap click/long_press/double_tap, and Round-layer traps drop ADB commands for several steps, inject HOME/APP_SWITCH, or press BACK every two steps, with exactly one trap per episode and 91% of tasks passing three-criterion human validation.
  • All 16 models from 7 organizations degrade under traps: Claude-Sonnet-4.6 falls from 74.2% to 66.5% average success and GUI-Owl-1.5-32B-Think from 69.5% to 62.4%, against a human baseline of 93.4%, with the Round layer (state deadlock, context disruption, loops) and External Interruption hurting most — GUI-Owl-7B drops from 63.1% to 26.3% under context disruption.
  • Reasoning does not buy robustness: thinking variants start higher in clean conditions but lose as much or more as their instruct counterparts, e.g. Qwen3-VL-8B-Thinking drops 6.1 points versus 5.8 for the instruct model, and GUI-Owl-1.5-32B-Think loses 7.1 versus 5.8.
  • GRPO in the clean environment lifts GUI-Owl-7B from 63.1% to 69.9% and UI-TARS-1.5-7B from 29.7% to 36.0% on original tasks but adds at most ~2 points of trap robustness, whereas per-subcategory GRPO in the AnTrap environment gains +8.1 to +11.0 points on State traps and up to +8.5 on Action traps, only +4.2–5.1 on Thinking traps, and under +3 on Round traps with Loop strictly below +1.
  • Caveats: the suite covers only 236 base tasks with a single trap event per episode and Pass@3 scoring, each adversarial GRPO run is trained and evaluated on the same subcategory so cross-trap generalization is untested, GPT models use Set-of-Marks prompting and skip the grounding trap, and adversarial SFT — the authors' own suggested route for the contextual failures — is left unexplored.

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

Highlight HF pick · 4▲Agents Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman Long-running coding-agent tasks tempt users to switch models mid-run, escalating from a low-cost, low-capability (LC) model to a high-cost, high-capability (HC) one when the cheap model struggles, or downshifting once the hard reasoning is done; each switch forces the receiving model to continue a trajectory it did not produce. The study pairs LC and HC models from the Claude and GPT families and varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and dropping the trajectory entirely while preserving the repository state. Full-trajectory escalation recovers less than half of the LC-to-HC quality gap while adding a substantial cost premium, whereas downshifting lands at a favorable cost-quality point. The preferred interface flips with direction: stripping the LC model's trajectory improves escalation quality, but removing the HC model's trajectory hurts downshift quality.

Coding-agent products let users switch models mid-run—escalating to a stronger model when a cheap one struggles, or downshifting once the hard reasoning is done—but the receiving model then has to continue a trajectory it did not produce, and the cost–quality consequences of that handoff have not been measured. The authors run a large controlled study on SWE-bench Verified varying handoff direction, timing, and what trajectory information crosses the model boundary, and find a direction-dependent "handoff tax."

  • Two low-cost/high-capability pairs (Haiku 4.5/Opus 4.7 and GPT-5.6 Luna/Sol) run in the same mini-swe-agent scaffold, switching at seven difficulty-calibrated percentiles of the starting model's step count (p5–p50) under four interfaces that all preserve the edited working tree but differ in what the receiver sees: the full trajectory (Raw), a summary written by the outgoing or incoming model (Compact-pre/Compact-suf), or nothing at all (Traj-drop), totaling 58,000 agent runs and 36 billion tokens.
  • Raw escalation is a poor bargain, recovering less than half of the LC-to-HC quality gap (47% for Claude, 36% for GPT) while costing far more than the cheap model; for Claude it costs $1.61 per task versus $0.72 for HC-only, so even paying for the discarded LC prefix and restarting HC from scratch ($0.90) is both cheaper and more accurate.
  • Downshift is a favorable operating point: Claude Raw HC→LC raises pass rate from 54.6% to 65.6% while retaining 80% of the cheap model's cost savings, and GPT retains 79% of the strong model's quality advantage at lower cost than HC-only.
  • The preferred interface reverses with direction—stripping the LC trajectory helps escalation (Traj-drop lifts quality recovery to 64% for Claude and 84% for GPT, while Compact-pre is cheapest), but dropping the HC trajectory hurts downshift (recovery falls to 28% and 53%)—because the full LC context makes each HC step 2.2× more expensive whereas missing HC context forces the LC receiver to take 1.6–2.0× more steps reconstructing it.
  • Scope is limited to two model pairs with a single episode per task, hard-task cells hold only ~24 instances so difficulty-conditioned findings are exploratory, and the interface comparisons are established only in the coding setting; extensions on LiC and BrowseComp under Raw transfer show the picture depends on task information dynamics, with escalation recovering 86% of quality on LiC where requirements arrive late.

Demystifying Reinforcement Learning Post-Training of Language Models

Highlight Reinforcement Learning Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh et al. Reinforcement learning (RL) post-training now drives much of the reasoning, math, and coding ability in large language models, but its inner mechanics remain opaque to many practitioners. Working in a deliberately simplified and controlled setting, the authors take RL with Verifiable Rewards apart step by step, examining how outcomes depend on the base model's prior distribution, reward granularity, prompt diversity, and model scale, and using output-distribution entropy to contrast what pretraining, supervised fine-tuning, and RL each do to model certainty. Two concrete conclusions stand out: the much-discussed effect of spurious rewards turns out to depend on which prompt distribution is used for post-training, and RL only succeeds when the base model already assigns enough probability mass to the target behavior, which connects it directly to classical exploration. The write-up is framed as a primer for NLP researchers adding RL to their toolkit.

RL post-training of LLMs is widely used but poorly understood, so the authors isolate its three ingredients — the base model's prior, the reward signal's density, and the prompt distribution — in a controlled RLVR sandbox where target-behavior probability and policy entropy can be measured exactly. The core claim is that most contested findings in the literature (coverage limits, "RL only sharpens pass@k", spurious rewards) are artifacts of sparse rewards and narrow prompt sets rather than inherent properties of RL.

  • The setup uses two toy tasks (emitting an exact movie-quote string, and solving AIME 2025 Problem 4 with answer 117), with the base prior manipulated via SFT+ (inject the target) and SFT- (maximize cross-entropy on it), and rewards ranging from binary sequence-level verifiers to dense proxies (Levenshtein distance for strings, a five-milestone PRM judged by Qwen2.5-32B-Instruct with exponential step weights [0.05, 0.05, 0.10, 0.15, 0.25] and loop/format penalties for math).
  • Under sparse rewards the coverage principle holds strictly: Qwen2-7B with a 3.5% prior climbs to 99.9% after a ~40-step plateau, Qwen3-1.7B with a 0.48% prior stalls at 10%, and Qwen3-8B base and all SFT- models stay at 0%; on AIME, Qwen2.5-7B-Instruct base goes 3.92% → 10.2% while SFT+ reaches 85.9%.
  • Dense reward shaping breaks the pass@k ceiling: Qwen3-1.7B reaches nearly 50% on the string task from a 0.5% prior, and the unmodified AIME base model jumps to 92.2% exact match, slightly beating the SFT+ model (86.7%) by composing partial reasoning steps into novel solution paths; the SFT- model still hits 0% exact match even though its dense reward climbs to ~0.6.
  • Spurious (uniform random) rewards are governed by prompt breadth: 100 narrow DeepScaleR prompts reproduce the Qwen gains with low entropy, while 10k broad WildChat prompts spike entropy and degrade AMC; on OLMo 3, broad prompts cause an entropy spike around step 400 with simultaneous collapse of GSM8K, MMLU, and IFEval (global unlearning, delayed successively by SFT and DPO stages), whereas 100 math-only prompts drop GSM8K from 86% to ~32% while MMLU (65 → 62) and IFEval (79 → 77) survive.
  • The study is deliberately narrow — one target string, one AIME problem, a hand-built PRM that would not transfer to open-ended tasks, no algorithmic ablations — so it is a mechanistic primer rather than evidence about production-scale post-training, and reward shaping has its own cost (the SFT+ string model underperforms under Levenshtein because near-misses are rewarded too generously).

FrontierChallenge: Evaluating Scientific Workflow Completion

Highlight HF pick · 97▲Agents Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li et al. Benchmarks for scientific agents mostly grade a final answer or an isolated script, which says little about whether a full research workflow was actually delivered. FrontierChallenge contains 300 end-to-end workflows, 97 of them released here, covering quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry, each with fixed inputs and a required bundle of scientific deliverables. Across twelve frontier models and three agent scaffolds, the best configurations fully completed only 20 of 97 tasks, a 20.6% pass rate, and partial credit proved badly misleading: analytical chemistry and electrochemistry reached average scores of 87.6 and 94.9 while passing 4% and 0% of tasks. Among non-passing Claude Code trajectories, 75.5% still ended by claiming the work was done.

A cross-domain benchmark asking whether an agent can carry a fully specified scientific workflow from fixed inputs all the way to a complete, mutually consistent bundle of deliverables, rather than just produce a plausible final answer. The core idea is contract-level grading: a task only passes if every required artifact satisfies a task-specific executable Grader, and partial credit is reported separately so it cannot be mistaken for completion.

  • FrontierChallenge collects 300 end-to-end workflows across quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment, and this release evaluates 97 of them (74 Hard, 23 Medium) that need no GPU, each packaged with fixed inputs, an execution environment with tools like ORCA, CP2K, LAMMPS, AmberTools, and PLUMED, an output contract, and a Grader that checks files, numbers, figures, code execution, and cross-artifact consistency, with GPT-5.6 Sol serving as a three-pass Judge for rubric criteria that need semantic assessment.
  • Across twelve frontier models on three scaffolds (Codex, Claude Code, Frontier Agent), Pass Rate ranged from 3.1% to 20.6% while Avg. Score ranged from 67.5 to 87.9; the top configurations, GPT-5.6 Sol with Codex and Grok 4.6 with Claude Code, each completed only 20 of 97 tasks, and every system did far better on Medium tasks (up to 43.5%) than on Hard ones (at most 14.9%).
  • Partial progress translated into complete delivery very unevenly by domain: quantum chemistry saw pass rates up to 60% and molecular dynamics 38%, but analytical chemistry reached an Avg. Score of 87.6 with a best Pass Rate of just 4%, and electrochemistry/environment scored up to 94.9 with a 0% Pass Rate for every configuration.
  • Trajectory analysis of 970 Claude Code runs found that 75.5% of non-passing runs ended with language claiming completion, only 1.5% admitted work was still in progress, and tool errors were actually more common in passing runs (94.2%) than failing ones (80.7%), so neither an agent's self-report nor the presence of errors reliably signals whether the contract was met.
  • The results rest on single runs of the 97 released tasks (203 are held out), the six domains are not balanced samples of their fields, the Judge is itself one of the evaluated models, and token and runtime figures are provider-reported (input use varied from 2.2M to 13.7M tokens per task, mean execution time from 21.8 to 112.8 minutes), so cross-model efficiency and cross-domain difficulty comparisons should be read as descriptive rather than rankings.

D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Highlight HF pick · 18▲Large Language Models Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng et al. Multi-teacher on-policy distillation compresses several domain-expert teachers into one student by minimizing per-domain reverse Kullback-Leibler divergence on the student's own rollouts, but the data mixture across domains is normally fixed in advance even though domains converge at very different rates. D3-MOPD reuses the reverse-KL values the training loop already computes: an off-process watcher tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and adjusts sampling ratios online without touching the training loop or adding overhead. Distilling a Qwen3.6-35B-A3B student from four expert teachers, it closes 97% of the average student-to-teacher gap versus 63% for a fixed mixture, reaches the fixed-mixture peak with roughly one third the rollout steps, and beats the specialist teachers on three of seven benchmarks.

Multi-teacher on-policy distillation (MOPD) merges several domain-expert teachers into one student by minimizing per-domain reverse-KL on the student's own rollouts, but existing pipelines fix the domain data mixture before training even though domains converge at very different rates, wasting compute on plateaued domains and undertraining slow ones. D³-MOPD repurposes the per-domain reverse-KL that the loss already computes as an online signal to reschedule the domain mixture, without touching the training loop.

  • An off-process watcher periodically reads the per-domain KL log, computes a composite signal that multiplies the remaining gap (current EMA KL normalized by initial KL) with descent velocity (average relative KL drop over R=3 non-overlapping windows, clipped at zero), and maps it through a temperature-controlled softmax with a per-domain floor ε=0.10; a stratified data source then assembles each batch to the target ratio with ±30% batch-level jitter, so integration only requires swapping the loader and launching a watcher.
  • On a Qwen3.6-35B-A3B student distilled from four GRPO-trained teachers (Math, Code, Instruction Following, Tool-use), it closes 97% of the average student-to-teacher gap versus 63% for vanilla MOPD across seven benchmarks (AIME 2025, HMMT Nov 2025, LiveCodeBench, OJBench C++, IFBench, IFEval, BFCL v3), surpassing the specialist teacher on three of them and reaching the baseline's peak average of 61.4 by step 47 instead of 143 (~3× fewer rollout steps).
  • Ablations show both signal components matter (gap-only 61.41, velocity-only 61.46, composite 62.34 versus 61.36 for vanilla) and that removing jitter costs 0.71 points, concentrated on code benchmarks; the learned schedule acts as an implicit curriculum, downsampling Code to ~0.15, boosting Math to ~0.50 mid-run, then shifting to IF and Tool-use late.
  • The method adds essentially no overhead (520 vs. 531 tokens/GPU/s, a 2.1% gap attributed to longer-response domains) and the improvement pattern replicates on a Qwen3.5-4B student, where it beats vanilla on all seven benchmarks and peaks 25% earlier.
  • Caveats: the headline 97% figure takes per-benchmark peaks across different checkpoints, and at a single deployable checkpoint the normalized scores drop to 0.73 vs. 0.48; the scheduler requires a 2W-step warmup during which the mixture is uniform, relies on an exponential-decay assumption only as design rationale, and its practical benefit depends on domains actually converging at different rates, with several new hyperparameters (T, ε, W, R, η) tuned on only two students.

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Highlight HF pick · 38▲ Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie et al.

An LLM agent's capability depends on the model–harness pair, yet harnesses (memory, planning, action loop, tool/skill orchestration) are still hand-designed ahead of time and expected to generalize across heterogeneous tasks. JIT-Agent is a 27B meta-model that instead synthesizes a task-specific executable harness on the fly for any off-the-shelf backbone LLM, then repairs it from execution diagnostics and evolves it against an archive of prior harness designs.

  • The harness is formalized as a four-module tuple (memory, planning, action, capability orchestration) under a fixed protocol, with HarnessFactory re-implementing 13 scaffolds (ReAct, Plan-and-Execute, ReSum, AgentFold, ROMA, AOrchestra, and others) as a seed bank; training on Qwen3.6-27B proceeds in three stages — SFT plus reward/latency/cost-aware preference learning on teacher-generated harnesses, imitation of bounded (≤2-round) repair trajectories driven by compiler and runtime error reports, and Evo-GDPO, a GRPO-style RL objective that rewards overtaking the archive's incumbent harness while normalizing reward, latency, and cost advantages separately.
  • Swapping the default scaffold for a JIT-generated harness improves all 18 matched backbone–benchmark pairs: GLM-5.2 rises from 74.1 to 81.8 on the nine-benchmark average (+7.7) and DeepSeek-V4-Flash from 66.7 to 75.5 (+8.8), with the largest jumps on planning tasks (+24.8 on DeepPlanning-Shopping for DeepSeek-V4-Flash, +20.2 on DeepPlanning-Travel for GLM-5.2), and JIT-equipped open backbones take first place in 8 of 9 columns, including DeepSeek-V4-Flash beating GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3).
  • In a controlled comparison holding the backbone fixed against Claude Code, Codex, OpenCode, Hermes, and NanoBot, JIT-Agent scores highest in 4 of 6 settings and has the lowest token and API cost in all six, cutting per-case cost by 14.9–54.1% (mean 36.0%) versus the cheapest fixed harness — e.g. on DeepSeek-V4-Flash xBench-DS it uses 212K tokens at $0.039 versus NanoBot's 527K at $0.075 while scoring 82.0 versus 78.0.
  • The gains come with caveats: the four-module protocol is deliberately much simpler than production runtimes, JIT-Agent trails Claude Code by 3.1 points on DeepSeek-V4-Flash AgentIF and NanoBot by 3.9 on Qwen3.6-Flash DeepSearchQA, the generator itself is frozen at deployment (only the harness archive evolves in streaming mode), and the pipeline depends on a stronger teacher model and a hand-written seed bank, so how much of the improvement transfers beyond the evaluated task families is not yet established.

Code World Model: Coding Agent as World Brain

Highlight HF pick · 9▲Agents Yiwen Chen, Guosheng Lin, Chi Zhang Video-based world models learn dynamics from visual observations, which reveal outcomes but not the rules and mechanisms driving them, making persistent consequences and coherent open-ended evolution hard to sustain. Code World Model separates world evolution from visual rendering: a coding agent acts as the world brain, reasoning about events and their consequences and emitting executable code that maintains persistent world state and evolves it consistently with the rules, while a proxy representation encoding frame-wise spatiotemporal constraints is compiled into a proxy video that conditions a video generator. Data pipelines build aligned proxy-observation pairs from gameplay and real-world footage, and after fine-tuning on paired gameplay data MiniMax-H3 follows the coding agent's proxy specifications for simple interactive worlds while preserving rich visual detail and dynamics.

Video world models learn dynamics purely from pixels, which show the outcomes of world rules but not the rules themselves, so they struggle to keep off-screen consequences alive and evolve coherently over long horizons. Code World Model splits the job in two: a coding agent acts as the "world brain," maintaining world state as executable, revisable code, while a fine-tuned video model renders that state into high-fidelity frames through a coarse, deterministically compiled "proxy video."

  • The coding agent handles sparse, semantically complex decisions (interpreting events, adding or rewriting mechanisms) and writes reusable code that performs the dense, high-frequency updates like positions, collisions, and cooldowns without a model call per step; the resulting camera, entity, and layout state is compiled into a proxy video of depth plus semantic-ID maps at one-quarter resolution per axis (1/16 the visual tokens), which conditions the video model alongside structured text describing appearance and action semantics.
  • Training data comes from recording gameplay video and runtime state synchronously so proxy and observation are frame-aligned by construction: 157 GTA V takes (~5.6 hours) yield 9,420 five-second clips at 1344×768 and 24 FPS with 336×192 proxies, and a KITTI-360 proof of concept shows proxies can also be compiled offline from real-world 3D reconstructions without any action or camera labels.
  • The authors fine-tune the MiniMax-H3 Ref2VA backbone with rank-128 LoRA across all 50 transformer blocks (~596M trainable parameters) for 3,534 steps on eight H800s, use GPT-5.6 Sol as the coding agent to build player-controllable worlds on top of existing game-engine templates, and generate long videos via overlapping 124-frame windows with a 34-frame overlap and a GPT Image 2 first-frame appearance anchor.
  • Despite only ~5 hours of paired gameplay, the fine-tuned model qualitatively follows proxy-specified entity positions, trajectories, scene layouts, and camera motion while generalizing to diverse art styles and characters, and the project-page demos claim finer, more responsive control over character actions and camera than action- or camera-conditioned baselines.
  • Evidence is entirely qualitative with no quantitative metrics or ablations, the training scale is small, real-time autoregressive generation is not implemented, the proxy design is fixed rather than agent-adjustable, and the coding agent cannot yet build complex game mechanisms from scratch, so the paper is best read as a framework proposal with a working prototype rather than a validated system.

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

Highlight Large Language Models Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland Orthogonal optimizers such as Muon speed up large language model pretraining relative to Adam, but the mechanism has been unclear. The authors probe Transformer loss landscapes at checkpoints along real training runs, decomposing each momentum buffer into its singular directions and estimating the loss-optimal step size along each on held-out data; the resulting spectral profile is stable across batches, training stages, optimizers, and scales, with a volatile head operating at the edge of stability that tolerates only small steps and a bulk that permits much larger ones. This spectral allocation view explains why Muon beats Adam, which beats SGD, and exposes that Muon's uniform scaling still underuses the bulk, motivating Spectral-Aware Muon (SAMuon), which holds the head at Muon's scale and amplifies the bulk using a static spectral prior, plus a cheaper SAMuon-lite variant using rank-one power iteration, neither adding persistent optimizer state or notable extra FLOPs. On modded-nanogpt models from 124M to 1B parameters both variants beat tuned AdamW and Muon in every tested configuration, with SAMuon needing 13.3% to 24.0% fewer training tokens than Muon to reach the same validation loss and SAMuon-lite keeping most of the gain at near-zero wall-clock overhead.

Orthogonal optimisers such as Muon accelerate Transformer pretraining over Adam, but the prevailing spectral-norm steepest-descent explanation does not account for the speedup on highly curved objectives. The authors measure, on held-out batches, the loss-optimal step size along each singular direction of the momentum buffer and find a stable, highly anisotropic profile — a single volatile head at the Edge-of-Stability tolerates steps nearly an order of magnitude smaller than the flat bulk — which reframes SGD, Adam and Muon as progressively better spectral allocations and motivates SAMuon, which pins the head at Muon's scale while amplifying the bulk.

  • At checkpoints along real modded-nanogpt trajectories, each rank-d singular component of every weight matrix's momentum buffer is assembled into a model-wide probe and its loss-optimal step is estimated from a local-quadratic approximation on a disjoint batch (one extra forward pass per probe); the profile is flat across the bulk, declines roughly linearly in log rank over the leading ~32–40 ranks, emerges over the first few hundred iterations, and the rank-1 optimum coincides with Muon's actual learning rate, with the same shape appearing on AdamW trajectories and at larger widths.
  • Viewed as spectral allocation, momentum SGD scales each direction by its singular value and so concentrates the update in the least stable directions, Adam's coordinate-wise rescaling lifts the bulk by an order of magnitude but leaves the head dominant, and Muon's whitening gets the shape roughly right but pins every direction to the head's conservative scale, realising only a small fraction of the idealised per-iteration loss reduction.
  • SAMuon multiplies the whitened update by a tail boost γ and restores the leading k directions (from torch.svd_lowrank, k = ⌊32√(d_model/512)⌋) to a log-rank-linear ramp from 1 to γ, while SAMuon-lite uses k = 1 via a few power-iteration steps and boosts the rest uniformly; neither adds persistent state beyond Muon's single buffer, both add one hyperparameter (γ = 1 recovers Muon exactly) plus a 30% cosine spectral warmup, and the exact-whitening idealisations keep Muon's O(T^-1/4) stochastic convergence rate.
  • Across 124M/300M/1B models on FineWeb at batch sizes 1024–4096, with hyperparameters tuned per batch size at 124M and transferred, both variants beat tuned AdamW and the Scion implementation of Muon in every cell: SAMuon reaches Muon's final validation loss with 13.3%–24.0% fewer tokens (loss gains of 0.0137–0.0295), SAMuon-lite with 13.3%–22.1%, and the advantage grows with batch size.
  • Every configuration is a single seed, γ is transferred across scales without retuning, the log-rank-linear shape and √width rule for k are empirical distillations rather than derived optima, the headroom estimate assumes Hessian-conjugate probes and is acknowledged as optimistic, and the full SAMuon adds 7.4% wall-clock per iteration on the 1B model from torch.svd_lowrank versus 0.5% for SAMuon-lite.

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

Highlight Theory Gerard Conangla Planes Choosing the rank of a low-rank adaptation (LoRA) update is normally an empirical exercise; here the authors derive task-dependent bounds on the approximation error achievable at each rank for Transformer attention. Fixing a pretrained attention head, a target attention function, and an input distribution from the downstream task, they bound the smallest expected Kullback-Leibler error of a rank-r query update: a lower bound proportional to psi of the norm of the difference between candidate and target attention scores, with psi(t) = min(t^2, t), when target probabilities stay bounded away from zero, an unconditional upper bound, and under explicit realizability, geometry, and moment conditions, best rank-r error sandwiched between explicit functions of the downstream-weighted tail energy of the target update. They also construct explicit families in which softmax saturation makes the rank needed to match the attention function strictly smaller than the rank needed to match the finite logits, and extend the analysis to fused multi-head LoRA and joint query/key updates, exposing the effects of rank sharing and query/key factorization constraints.

Picking a LoRA rank is normally done by sweeping ranks and comparing downstream scores, which cannot tell whether a small adapter lacked capacity or was merely harder to train. Fixing a pretrained attention head, a target attention function, and a downstream input distribution, the paper bounds the smallest expected KL divergence between target and candidate attention achievable by a rank-r query update, showing that the relevant quantity is the target update's spectrum weighted by the task's actual queries and keys rather than the raw singular values of the weight delta.

  • A global softmax lemma brackets pointwise attention KL between (a/2e²)·ψ(‖d‖₂) and min{‖d‖₂²/4, √2·‖d‖₂}, where d is the centered score error, a the minimum target probability, and ψ(t) = min{t², t}, so KL grows quadratically for small score errors but only linearly for large ones, and a purely quadratic global lower bound provably cannot hold.
  • Assuming the target is realizable by a dense query update Δ, a probability floor a, independence of the key Gram G(u) and query activation h(u), and moment/geometry constants Λ and κ_h, the best rank-r error satisfies a/(2e²(1+Λ√κ_h))·ψ(√T_r) ≤ E_r ≤ ψ_up(√T_r), where T_r is the tail energy of the task-weighted matrix D = G^{1/2}·Δ·Σ^{1/2}, the upper bound is attained by truncated SVD of D (which discards directions the task never activates), and the ratio between the two constants can reach about 2.0×10⁴ for a = 10⁻², Λ = 5, κ_h = 3.
  • Because that probability floor makes the lower bound weak for long, peaked attention vectors, two alternative routes are given: target-Fisher bounds that are two-sided but only hold for candidates whose score differences stay within a fixed range, and a high-mass lower bound that applies to the unrestricted class but only counts tokens carrying most of the target mass.
  • Softmax saturation genuinely reduces required rank: for an explicit Walsh-attention family, matching finite logits exactly needs rank k, but the limiting attention is reachable to arbitrary accuracy at rank k − ⌊k/3⌋ (a separate linear-token family reaches ratio 4/7), while extensions to fused multi-head LoRA widen the bracket by up to a factor of √H and joint query/key LoRA adds a nonconvex factorization-gap term ρ on top of the effective-rank bound r_Q + r_K for which no tractable procedure or universal bound is provided.
  • Everything is theoretical with no experiments on trained heads: the bounds cover attention KL rather than head output or task loss, need an already-trained dense or high-rank target adapter, are population statements without finite-sample confidence intervals, say nothing about whether SGD reaches the optimal candidate, and the joint Q/K spectral result does not apply directly under RoPE (though it does for NoPE models such as Kimi K3).

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Highlight HF pick · 17▲Multimodal Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou et al. Native visual reasoning treats generated images and videos as the medium of problem solving rather than merely inputs or outputs, but progress has been limited by a lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. VBVR-Pro is a closed-loop testbed of 300 procedurally generated tasks with deterministic, rule-based reward scorers, which align more closely with human judgments than the prevalent vision-language-model-as-a-judge paradigm and serve as reliable reward signals for large-scale multi-task reinforcement learning. Models trained on the suite transfer to seven external benchmarks including RISE-Video, MME-CoF-Pro, and BabyVision, and controlled studies across more than 30 image, video, and interleaved generators find that video generation is strongest for tasks requiring persistent spatiotemporal state tracking while interleaved generation offers a compute-efficient alternative, with ablations and probing suggesting that vision-native trajectories are crucial to visual reasoning.

Native visual reasoning treats generated images and videos as the working medium of problem solving rather than as inputs or final renders, but the paradigm has been stuck without scalable training tasks, trustworthy reward signals, or apples-to-apples comparisons between image, video, and interleaved generators. VBVR-Pro closes that loop with 300 procedurally generated tasks, deterministic rule-based scorers that double as RL rewards, and a modality-controlled model suite trained and evaluated on one shared task distribution.

  • Each task is a parameterized generator with a programmatic solver, rendered simultaneously into aligned video, keyframe-image, and interleaved text-image forms (step text written by Gemini-3.1-Pro), producing 1.25M training instances from 250 tasks plus a 100-task benchmark (50 in-domain, 50 held-out); the scorers extract task-relevant attributes with classical CV (HSV segmentation, contours, OCR, trajectory tracking) and combine hand-weighted checks additively for soft criteria or multiplicatively for hard constraints.
  • Against an arena-style human study, the scorers reach per-vote agreement above 0.60 versus 0.54 for GPT-5.5 and 0.52 for Gemini-3.1-Pro (human ceiling 0.77), rank generators with Spearman ρ = 1.00 out-of-domain, and are fully reproducible and the cheapest evaluator tested, whereas VLM judges change their scores on 55–93% of samples across reruns even at temperature 0.
  • Fine-tuning nine open-source models for one epoch raises VBVR-Pro-Bench scores by +0.290 on average (+0.401 in-domain, +0.179 out-of-domain), with VBVR-Pro-Wan2.2-I2V-A14B reaching 0.670 to beat proprietary Seedance 2.0 (0.499) and Nano Banana Pro (0.564), and transferring to unseen benchmarks such as V-ReasonBench (10.2 → 38.2), VideoThinkBench (25.7 → 52.9), and MME-CoF-Pro (24.8 → 48.4) with nearest-neighbor checks showing no close training matches.
  • Controlled comparisons across 30+ generators find video strongest on transformation and out-of-domain tasks that need persistent state tracking, while interleaved VBVR-Pro-SenseNova-U1 (0.638) matches it on many in-domain tasks at far lower cost; crucially, collapsing intermediate images to a single frame costs −0.111 but replacing reasoning text with placeholders costs only −0.009, and at inference removing the intermediate image drops the score from 0.638 to 0.099 versus 0.533 for removing the text, indicating visual trajectories rather than linguistic chains carry the computation.
  • The best trained model still sits well below human performance, out-of-domain gains are under half the in-domain gains, every scorer is a hand-engineered program with human-calibrated weights covering only the 100 benchmark tasks, and the RL results (using a CPS sampler for semantic rather than pixel-level exploration) are described as steady improvements on strong baselines without headline numbers in the available text.

Applications 107

LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology

Marie-Lisa Eich, Kai Standvoss, Timo Milbich, Alexander M\"ollers, Miriam H\"agele, Philipp Anders et al. cross-listed Lung cancer tissue diagnosis requires integrating histomorphological, immunohistochemical, and molecular features, but pathologist assessment remains visual, semi-quantitative, and variable between observers, while existing AI tools cover isolated tasks and lack prospective clinical validation. LUCAID is an agentic system in which an integrative agent couples diagnostic reasoning with nine modules spanning quality control, tumor detection and segmentation, histological subtyping, tumor microenvironment profiling, cellularity quantification, predictive biomarker scoring for PD-L1, MET, and TROP-2, and structured report generation, with users able to interactively query module outputs. The analysis modules reach F1 scores of 0.82-0.95 against large-scale expert annotations, and in prospective clinical validation the system achieved 93.0% concordance with an expert-panel adjudicated reference standard on clinically actionable decisions, versus 68.3-81.1% for five experienced thoracic pathologists.

Automated Synthesis of Cloud Emulators

Archit Bhatnagar, Zhenning Yang, Sarah McClure, Yiming Qiu, Sylvia Ratnasamy, Ang Chen cross-listed Testing DevOps programs such as CLI scripts or infrastructure-as-code against real cloud resources is slow, risky, and expensive, and hand-building the API-level emulators used instead requires developers to interpret vast, constantly changing cloud documentation service by service. CloudEmu synthesizes emulators automatically from cloud documentation through neurosymbolic code synthesis, pairing LLM documentation understanding and code generation with cloud-specific symbolic abstractions that suppress hallucinations and enforce precision, while using the real cloud as an oracle for automated testing, repair, and alignment. Evaluated on AWS and GCP services, CloudEmu outperforms LocalStack, the leading existing emulator hand-built by a large engineering team over a decade, in both coverage and accuracy.

A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization

Prithvi Dake, Rahul Bindlish, James B. Rawlings cross-listed Real-time optimization (RTO) uses process models to find economically optimal operating conditions, and data-driven models are attractive because first-principles models demand deep process knowledge, but whether a model that fits plant data well can be trusted for economic optimization is unclear. Using a vinyl acetate monomer benchmark process with a single well-conditioned optimum, the authors train a structured hybrid model (known mass balances and thermodynamics plus a neural-network kinetics closure) and a fully data-driven neural ODE model; both reproduce plant measurements accurately and are stable across random initializations. Yet both return many phantom economic optima that differ substantially from the plant's single optimum, and even with noise-free data and initialization at weights that recover the true optimum, stochastic gradient training drifts to weights yielding much worse RTO solutions, so the identified model is an artifact of the training optimizer as well as the data. The authors argue that data-driven RTO models should be required to recover the optimum on a decision-oriented benchmark before plant testing.

Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory

Danish Khan, Maurice D. Hanisch, Nikolai Argatoff, Evan Xie, Sandeep Sharma, Anima Anandkumar cross-listed Kohn-Sham density functional theory (DFT) underpins electronic-structure simulation, but its repeated orbital diagonalizations scale cubically, and prior orbital-free approaches, whether analytical or learned, have fallen short because kinetic-energy functionals are ill-conditioned and direct ground-state prediction extrapolates poorly. The authors instead learn the Kohn-Sham map itself, which sends a potential directly to its density and noninteracting kinetic energy, using a domain-invariant SE(3)-equivariant Fourier neural operator on real-space grids to predict density from potential and enable quasi-linear-scaling self-consistent field (SCF) iterations. Trained jointly on 8,504 molecules and solids, a single model converges SCFs on out-of-distribution organic molecules, insulators, and metals without ever constructing Kohn-Sham orbitals, reproducing densities, electronic spectra, and structural observables at Kohn-Sham accuracy, and linear scaling allows converging magnesium dislocation systems with up to 82,500 valence electrons on one GPU.

When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs

Zhiyang Qi cross-listed In psychological counseling, brief utterances such as backchannel cues and short empathic statements convey attentive listening and encourage clients to keep talking, yet counseling dialogue systems and their evaluation frameworks favor long, content-rich replies. The authors run a cross-lingual analysis of minimal responses across several counseling datasets using a two-stage length-and-content filter followed by contextual verification with a large language model (LLM), then test current models in manually curated contexts where human counselors chose a minimal response. Minimal responses are common in human-collected data but substantially underrepresented in LLM-generated dialogues; strong commercial LLMs can produce them when explicitly instructed but struggle to judge when they are appropriate, counseling-specific models trained on synthetic data do particularly poorly, and LLM-based quality evaluation undervalues minimal responses even when they are interactionally appropriate.

From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender

Sonia Sharma, Jeyendran Balakrishnan, Shreya Rajpal, Swapnil Parekh, Nagaraj Janardhana, Andrew Mattarella-Micke cross-listed Service businesses are moving from static, independently priced SKUs to dynamically bundled, discount-coupled offerings, which strains the tree-based classifiers usually favored for sparse, imbalanced data because they assume a slowly changing label space and cannot easily absorb multimodal signals such as conversation transcripts. The authors describe migrating a live production customer-support recommender from a gradient-boosted multiclass model to a pairwise-binary deep recommender, learning jointly from user and item features with negative sampling and noise injection, applying attention pooling over transcript chunks (benchmarked against TF-IDF and sentence-embedding baselines), and comparing architectures including two-tower models, DeepFM, and contrastive losses. Against a CatBoost baseline, the deep recommender reaches parity at the start of conversations and outperforms at later conversational stages.

Constraint-Guided Enterprise Data Mapping with Large Language Models

Sebastian Monka, Pramod Anantharam, Thien Vo Minh, Lavdim Halilaj Enterprise entity alignment must cope with semi-structured records, implicit attributes, and unit or granularity mismatches; manual matching does not scale, and LLM-only matching improves semantic recall but can violate structural and physical invariants, producing fluent yet operationally invalid correspondences. Constraint-guided mapping (CGM) is a neuro-symbolic method that derives schema-grounded admissibility constraints with typed, executable relation and normalization logic, generates candidates restricted by those constraints with cascade relaxation to guarantee a nonempty feasible set, and applies neural ranking with bounded LLM disambiguation only within that set, so constraints act as hypothesis-space operators rather than post-hoc validators. On a structural-decoy benchmark, hard admissibility shrinks the candidate space by roughly 480x without losing the ground truth, and ablation shows the constraint gate, not the LLM, drives the lift (F1 0.08 to 0.66); a small model with constraints matches a frontier LLM without them at about 28x lower cost, and the method transfers across seven enterprise makes (macro F1 0.70) while cutting expert effort roughly 7x versus spreadsheet workflows.

FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment

Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng cross-listed Adversarial robustness of fraud and credit-risk models is hard to evaluate because financial tabular data carry domain constraints, severe class imbalance, and asymmetric attacker capability, and the authors argue that robustness conclusions depend on the evaluation protocol as much as on the model. FraudBench evaluates the same dataset, model, attack, and defence configuration under three matched protocols, unconstrained attacks, post-hoc feasibility filtering, and deployment-aware attacks with constraints integrated into generation, across four public financial datasets and neural, tree-based, and ensemble models. On Lending Club Loan Data in the white-box setting, post-hoc filtering leaves only 3.7 feasible flipped examples on average, while in-attack projection with attacker mutability masking yields 2,832.3 under the same perturbation budget; on IEEE-CIS feasibility and attacker capability act as separate axes, and in black-box evaluation the protocol choice can reorder model-family rankings.

Same-Player Verification for Account Consistency in Counter-Strike 2

Xuchen Zhang cross-listed Account-integrity review in competitive first-person shooter (FPS) games such as Counter-Strike 2 (CS2) asks whether an account's recent behavior is consistent with its historical operator, as in temporary substitution, rank boosting, or high-skill players on lower-ranked accounts. The authors frame this as same-player verification: each player's trajectory in a match replay is encoded as a behavioral fingerprint covering crosshair control, movement-stop-fire coordination, economy and buy decisions, combat engagement, and temporal rhythm, and a pairwise model is trained to judge whether two observations come from the same person, using 13,300 demo-player observations from 1,330 demos and 663,590 sampled pairs. The model reaches an average ROC AUC of 0.931 and 0.722 different-player recall at 95% precision, with low-level mechanical habits such as crosshair control and firing rhythm carrying the strongest identity signal, and aggregating over an account's history raises AUC from 0.931 with a single pair to 0.986 with ten.

MacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms

Jiaxi Jiang, Xufeng Yao, Yuxuan Zhao, Yuntao Lu, Peiyu Liao, Zuodong Zhang et al. In very large-scale integration (VLSI) chip design, macro legalization is the final step that nudges large blocks into non-overlapping legal positions, and existing methods are either fragile, slow, or blind to the visual regularity that helps routing. MacroAgent runs clustering, contour generation, template matching, and inter-cluster refinement, and uses large language models to invent the regularity-aware contour heuristics that drive the contour stage. On the TILOS and Chipyard benchmarks it improves layout regularity by 2 to 8 times and shortens routed wirelength by 3% to 5% at comparable congestion; a full place-and-route run through Cadence Innovus confirms downstream gains, including 68.3% better total negative slack and 2.9% lower routed wirelength versus the DREAMPlace legalization baseline.

AFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions

Manmeet Singh, Somnath Luitel, Prabhjot Singh, Manraaj Banga, Naveen Sudharsan, Josh Durkee Large language models tend to invent numbers when drafting high-stakes meteorological text, which is dangerous for public weather communication. AFDBench pairs 7,732 expert-written Area Forecast Discussions from 13 National Weather Service (NWS) offices with the structured forecast fields from Google's WeatherNext 2 model, and scores generations on numerical accuracy (Met-Align), professional register (Style-Align), and fidelity to the input data (Input-Grounding). Zero-shot open-source models write in the wrong register (Style-Align near 0.33) and use their inputs loosely; applying Group Relative Policy Optimization with rewards for temperature accuracy, synoptic correctness, and format compliance to a 7B-parameter model nearly doubles Style-Align from 0.318 to 0.619 and lifts Input-Grounding from 0.881 to 0.940 on held-out samples from two unseen offices.

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

Jian Wang, Steven Xu, Sanjyot Thete, Maryam Barouti, Tom Tang, Elaine Wu et al. cross-listed Product linking, the entity-resolution task of mapping merchant product records to canonical catalog products, must resolve billions of noisy multi-category records against tens of millions of catalog entries, where a single scoring model is either too weak for the hard cases or too costly for the easy ones. The authors describe a production retrieve-then-match cascade that spends compute in proportion to difficulty: retrieval surfaces plausible candidates, a lightweight text cross-encoder distilled from millions of dual vision-language model (VLM) consensus labels auto-resolves the high-confidence majority at a 98% precision bar validated against an operator-certified audit, and an agentic multimodal VLM settles the ambiguous remainder by inspecting product images and issuing web searches for evidence found in neither record. The self-hosted open-weight agent matches a closed frontier VLM's precision at a four-point recall cost (88% versus 92%) for roughly one-seventh the per-pair cost with no fine-tuning, and since per-pair cost spans nearly five orders of magnitude across stages, escalating only the hard tail raises end-to-end link coverage from 68% to 77%.

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani et al. General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are rarely resolved by clinical specialty, which makes it hard to isolate mental-health performance even as millions of people turn to LLMs for psychological support. The authors screen HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, validate the result through two rounds of blinded clinician review with concealed known-exclude controls, and release HealthBench-Psych and HealthBench-Psych-Hard, comprising 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models with a cross-vendor panel of three LLM judges, they find a statistically tied frontier cluster, measurable refusal behaviour in two models, and near-identical rankings across judges (Kendall's τ ≥ 0.92), and they release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.

FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs

Ze Sheng, Aleksandar Kezic, Zhicheng Chen, Jeff Huang cross-listed Existing benchmarks for language-model bug finding ask the model to trigger a predefined vulnerability, which discards valid crashes that happen not to match the target. FuzzingBrain-Bench instead gives a model an open-source project with a sanitizer-instrumented harness in a self-contained Docker image and scores it on the number of distinct crash signatures it produces, capped per challenge and weighted by difficulty. Version 1 comprises 77 challenges from 43 projects (36 C, 32 C++, 9 Java/JVM); among Claude Haiku 4.5, Claude Sonnet 4.6, and Claude Opus 4.8, Claude Opus 4.8 performed best, triggering crashes in 60 of 77 challenges for a score of 196 out of 579, while 13 challenges yielded no crash from any model.

InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance

Yating Ling, Wenjing Cun, Zhitang Chen Symbolic regression (SR) tries to recover compact mathematical laws from data, but genetic programming struggles with the enormous space of physically meaningful expressions. InsightSR wraps the PySR genetic programming engine with large language models that reshape the search space rather than writing formulas directly: a Semantic Seed Pathway proposes dimensionally consistent functional skeletons, a Structural Feature Pathway recommends nonlinear feature transformations, and a feedback loop scores candidates and refines the guidance each iteration, so the search assembles shallow trees over a rich feature set instead of deep trees over raw variables. Across three benchmarks it reaches a 95% exact recovery rate on the Feynman benchmark and 80.18% accuracy on the LLM-SRBench LSR-Transform task, outperforming genetic programming and neural-symbolic baselines while generalizing out of distribution on real-world datasets.

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

Natasha Ureyang, Sebastian Porsdam Mann, Yuxin Liu, Zuriel Hassirim, Melanie Almonte, Wenhao Chen et al. cross-listed Human surrogates asked to predict a seriously ill patient's treatment preferences are right only about 68% of the time, which fuels decision conflict. P4-DT builds a personalized decision policy by walking a patient through varied medical dilemmas and eliciting their preference reasoning in both directions, rather than treating values as static ratings, then few-shot prompts a language model with that material. Across 12 patient-surrogate dyads it predicted patient treatment choices with 81.7% accuracy versus 55.0% for unassisted surrogates (and 61.7% for surrogates given the tool), and adding contextual scenario decisions plus open-ended text improved accuracy by 15.0 percentage points over values ratings alone.

Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening

Wensi Zhang, Tomas Teijeiro, J\'er\^ome Thevenot, David Atienza cross-listed Cough acoustics look like an appealing non-invasive tuberculosis (TB) screen, but it has been unclear whether models learn disease-related sound or artifacts of how the data was collected. Classical machine learning and deep learning cough classifiers were evaluated across three independent datasets: within-dataset ROC-AUC reached 0.755, while external performance frequently fell below 0.6, and the learned audio representations clustered by recording device and dataset rather than by TB status. Predicted TB probability tracked country-level prevalence in CODA, device mismatch degraded transfer while device-diverse training improved it, and a clinical-variable baseline generalized more consistently (ROC-AUC 0.655–0.711), pointing to acquisition variability rather than population shift as the dominant obstacle.

MetaSieve: Faster Relational Deep Learning through SQL-Based Metapath Selection

Fahim Shahriar Khan, Ashraf Aboulnaga cross-listed Relational Deep Learning (RDL) models a multi-table database as a graph, with rows as nodes and foreign-key relations as edges, and trains a graph neural network (GNN) on it, but training cost is dominated by the size of the subgraph sampled around each seed node. MetaSieve observes that those subgraphs are produced by following metapaths of foreign-key links and that many metapaths can be pruned without losing accuracy, so it computes statistics for each candidate metapath extension using SQL join and aggregation queries and scores it with a function that favors lightweight but informative candidates. Because scoring depends only on database statistics and task labels rather than GNN parameters, the layer drops into different architectures for classification and regression; on the RelBench benchmark with several GNN backbones it cuts per-epoch training time by large margins while maintaining and often improving accuracy.

FRAME: separating sampling variation from representational cause in medical imaging fairness

Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh cross-listed Subgroup performance gaps are the standard evidence for bias in medical imaging models, and the usual remedy is to strip demographic information from the representation. Fair-model Reference And Mechanism Evaluation (FRAME) first derives the distribution of subgroup differences expected under exact fairness at the observed subgroup sizes, then probes whatever remains with two representation-space operators, one of which cannot alter within-group rankings by construction. Across 702,206 images and 36 encoders, the sampling reference alone accounts for a median 41% of reported race differences and 22% of age differences, injecting demographic decodability leaves the remainder unchanged, and no tested intervention moves the remainder more than a change of random seed does. Applied to 89 differences from 9 published studies across 6 modalities, the reference explains a median 25% of rate differences and 70% of differences in area under the receiver operating characteristic curve, suggesting many reported gaps are compatible with sampling variation at current cohort sizes.

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun, Hamed Sadeghi et al. cross-listed Building predictive geospatial models is bottlenecked by a fragmented data ecosystem that requires manual retrieval, multimodal curation and fusion, and iterative model selection. The Planetary Prediction Engine (PPE) runs this workflow end to end from a natural-language query, retrieving spatiotemporally relevant covariates from Data Commons and Google Earth Engine, fusing them with geospatial foundation model embeddings from PDFM and AlphaEarth, and searching over task-tailored architecture families with automated overfitting guards. For US spatial regression it raises mean R² across 21 CDC health indicators from 60.0% to 76.8%, doubles baseline accuracy when downscaling Nigerian food security indicators, and for nowcasting the 2026 DRC Bundibugyo Ebola outbreak reaches a Recall@10 of 83.3%, identifying 15 of 18 newly invaded health zones, a 10.3-point gain over the public state of the art.

Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role

Ahmad Khan, Akram Bin Sediq, Sara Azadegi Naeini, Raviraj S. Adve Machine-learning algorithms for wireless resource management are normally designed by hand, with the architecture, loss function, and training recipe all specified by a researcher. Using the autoresearch protocol, an AI coding agent is given full authority over the architecture family, input representation, output parameterization, loss, and task-sampling law, and repeatedly edits a training script, runs a fixed-budget experiment, and keeps or discards each change according to one immutable metric, with a hash-pinned evaluator, an enforced inference contract, and a pre-registered falsifier per experiment as safeguards. The target is sum-least-percentile-rate power control across a multicell network, a non-convex, non-smooth, and strongly NP-hard problem, and in 81 unattended experiments over 26 hours the agent reached 99.5% of a converged minorization-maximization reference in a single inference pass at roughly 600x lower inference cost, with one parameter set serving every network size and percentile target; the output parameterization it discovered provably reproduces the exact max-min-optimal allocation at the minimum percentile.
86 more specialized papers

Large Language Models 76

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

Yuan Si, Simeng Han, Daming Li, Jialu Zhang Evaluations of long-term memory and retrieval-augmented generation (RAG) usually treat how retrieved history is rendered for the answering model as an implementation detail, even though the same conversation may be shown as a memory entry, a summary, a typed record, or a raw excerpt. RENDER fixes the conversation and varies only this reader-facing artifact, combining a five-level packet ladder that localizes when answer-bearing content enters the input with deterministic templates that mimic ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw dialogue. On 500 LongMemEval questions across nine models, matched-budget resolved packets beat recency-truncated raw dialogue by 42.4 to 72.6 points, and the best-to-worst spread across deployed-style templates is 24.6 to 48.8 points per model, with three models scoring zero on formal ledger packets while answering the same facts from natural-language entries at 45 to 53 percent. The effect persists under retrieval noise and transfers to HotpotQA, which the authors take as grounds for reporting or controlling the rendering format in memory and RAG evaluations.

ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik Natural Language to SQL (NL2SQL) systems report execution accuracy above 89 percent on Spider and BIRD, but those benchmarks use simplified academic schemas and open-source dialects rather than enterprise databases. ESQ-Bench is an Oracle-first benchmark with six populated schemas (465 tables, 164,682 rows) replicated with identical seed data on Oracle, PostgreSQL, MySQL, and SQL Server, 550 gold-validated question-query pairs across three complexity tiers, and a four-metric harness whose silent-divergence metric flags queries that pass execution checks yet return semantically wrong results. With schema-linked prompting, GPT-4o execution match degrades monotonically from 79.8 to 60.3 to 57.2 percent across tiers while Claude Sonnet 4.6 reaches 87.4, 74.9, and 68.7 percent, exact match stays below 7 percent, and silent divergence affects 73 to 99 percent of queries that pass execution match. A local Llama 3.2 reaches only 13.3 percent bank-wide execution match, underscoring the gap between closed API models and open-weight baselines on enterprise schemas.

Identifying Latent Declarative Representations of Code for Assisting Repository Migration

Shraddha Surana, Ashwin Srinivasan, Michael Bain cross-listed Legacy repositories embed decades of domain knowledge in undocumented code, making repository-scale porting difficult. ADFD-Migrate treats a program as the implementation of an unobserved declarative description and approximates it with an annotated data-flow diagram (ADFD) of processes, data stores, external entities, flows, and behavioral contracts; an LLM infers the source ADFD from bounded repository context guided by static-analysis coverage checks, dependency-aware chunking orders process groups for target-language generation, and differences between the source ADFD and a statically recovered target ADFD drive regeneration. On f2x50, a new benchmark of 50 Fortran repositories spanning 1.5 thousand to 1.6 million lines of code, the generated Python passes 327 of 382 curated Fortran-oracle probes (85.6 percent), with 40 repositories passing every attempted probe, and exposes all 382 planned behaviors as runnable targets versus 99 and 98 for direct and repository-context translation. The method also reaches a 93.1 percent mean migration outcome index, a 17 to 59 percentage-point advantage over direct translation on 47 repositories.

Function-Level Execution Feedback for Code Preference Optimization

Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong et al. Process supervision has improved mathematical reasoning, but in code generation there is no standard notion of a step, leaving it unclear whether to label lines, reasoning traces, or program states. STEP-KTODER defines steps as module-level functions in decomposed multi-function programs, assigns each a binary correctness label via automatically generated unit tests, and combines this function-level supervision with outcome-level feedback on the full program as a code-specific instantiation of stepwise KTO. On HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench the method improves over outcome-only KTO and DPO, and execution-based labels prove essential: LLM-as-a-judge annotations systematically over-predict function failures, corrupt positive step labels, and degrade downstream preference optimization.

Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap

Sathishkumar Sivashanmugam cross-listed LLM serving engines size the key-value (KV) cache once at startup and permanently reserve memory for the worst-case prefill activation, leaving that reserve idle during decode-heavy phases. The authors build an elastic KV cache that lends the reserve to the KV pool during decode and returns it before prefill, implemented purely in userspace on the CUDA virtual-memory path with two physical handles mapped into one contiguous virtual range per layer, so attention kernels and drivers stay unchanged; it decommits in milliseconds, works with CUDA graphs and prefix caching, and never triggers out-of-memory. They then test the premise and report a negative result: median time-to-first-token differs by only about 1% between chunk sizes of 8192 and 32768 tokens, because prefill is compute-bound, so simply lowering max_num_batched_tokens recovers more KV memory than the controller at nearly equal latency. The reserve also shrinks from 16% of KV at tensor-parallel degree 1 to 2.7% at degree 4; the mechanism is released as a reusable elastic virtual-memory allocator.

The Limits of Automatic Evaluation of Creativity in Large Language Models

Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi cross-listed Evaluating creativity in LLM-generated text remains hard, and it is unclear whether current automatic methods track human judgment. The authors collect human ratings of human- and AI-written short stories from the WritingPrompts dataset across 11 dimensions of creativity and compare them with objective automatic metrics and LLM-as-a-judge evaluations. LLM judges show a systematic preference for AI-generated stories, favoring their stylistic traits over the unpredictability and other qualities of human-authored text, and widely used automatic metrics show near-zero correlation with human judgments for both story sources. The results point to fundamental limits in reducing the multidimensional, subjective nature of creativity to computational metrics.

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

Farhana Amin, Sabiha Afroz, Mona Moghadampanah, Dimitrios S. Nikolopoulos Masked diffusion language models (dLLMs) denoise many tokens per step and so promise faster generation than autoregressive (AR) models, but serving systems for them have been built without measuring how the models behave under real concurrent load. Using LLaDA-8B-Instruct with a D2F (Discrete Diffusion Forcing) LoRA adapter on a single NVIDIA H200, evaluated on GSM8K and HumanEval, the authors find that per-request difficulty falls into 11 discrete denoising step-count levels that no tested signal predicts before generation starts (best R² of 0.150), that generation budgets under 320 tokens hide the latency spread, and that only 24% of single-request wall-clock time is GPU computation, with the rest being CPU-side dispatch overhead. Sharing one forward pass per denoising step across a batch of 16 improves throughput 16.0x over per-request dispatch, quality is argued not to degrade with batch size under three stated assumptions, and a batch-timeout rule is derived for synchronized batching under Poisson arrivals, leading to the conclusion that dLLM serving needs parallelism at the level of each denoising step.

Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

Igor Bogdanov, Changcheng Huang cross-listed Multilingual models can solve the same math problem in different languages, but whether they rely on shared internal features or on language-specific computations that merely converge on similar outputs is unclear. Using MGSM (Multilingual Grade School Math) problems solved in six languages across five models from four families, the authors locate layers with cross-language alignment via Centered Kernel Alignment (CKA), then train both a reconstruction-only sparse autoencoder (SAE) and a Geometry-Invariant SAE (GI-SAE) that adds an Information Noise-Contrastive Estimation (InfoNCE) loss pushing traces of the same problem toward similar activations regardless of language or token position, and test functional interchangeability by swapping feature values between languages mid-forward-pass and measuring the resulting KL divergence. Although GI-SAE raises geometric similarity at nearly every layer, higher geometric similarity does not consistently translate into functional interchangeability; cross-language sharing is model- and architecture-dependent, strengthening in Qwen, showing no functional benefit in Gemma, and giving mixed layer-dependent effects in Llama and Phi.

Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention

Sergii Kozyrev (Minima AI, Inc), Davyd Maiboroda (Minima AI, Inc) The key-value (KV) cache is the main capacity and bandwidth bottleneck in long-context LLM serving. Minima-KV keeps recent and protected anchor pages in FP8 while moving older pages to a packed 3-bit TQ3 format, and uses format-specific kernels whose partial attention states are combined through a globally normalized online-softmax merge, so mixed-precision pages can be decoded directly without evicting any live-request page or materializing a dense shadow copy of the cache. On Qwen3.6-27B on a single 96-GB NVIDIA RTX PRO 6000 Blackwell GPU, deployment accounting reports 18.3 KiB of attention KV per live token, 3.50x compression relative to BF16 and 1.75x relative to FP8, while a quality profile matches its dense control on 16K RULER needle-in-a-haystack tasks and loses 0.4-0.8 percentage points on LongBench v2 at 16K to 64K context; a direct-decode test with two 59,008-token requests measures 3.625x active-KV compression at 0.98x the control's throughput with all 16 full-attention layers routed without fallback.

Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

Yicheng Mao, Hongru Du Choosing how to allocate a fixed pretraining token budget across data domains is usually handled by training small proxy models on candidate mixtures, fitting a response model, and extrapolating, and the authors observe that this workflow is exactly a classical mixture experiment: domains are components, token shares are proportions, proxy runs are design points, and validation loss is a response surface over the probability simplex. They fit sparse second-order Scheffé response-surface models and construct model-robust I-optimal designs for proxy experiments, using RegMix as the empirical case study. The analysis shows domain value is strongly relational, with several domains that look weak additively becoming favourable through pairwise interactions with web-derived text; the sparse Scheffé model preserves mixture rankings across model scales while staying competitive with a flexible machine-learning predictor, and in a calibrated simulation, I-optimal designs recover the relevant mixture ordering after dropping about 25% of the proxy runs.

Evaluating Language Models on Cross-Language Code Functional Equivalence

Hui Sun, Anderson Uch\^oa, Rohit Gheyi, Wesley K. G. Assun\c{c}\~ao cross-listed Whether large language models genuinely reason about program semantics or lean on surface similarity is tested by asking them to judge functional equivalence between human-written programs within and across languages. The authors build PolyHuman, a dataset of human-written C++, Java, and Python programs, evaluate open-weight and proprietary models on intra- and inter-language equivalence detection, and manually analyze 81 cases of systematic disagreement, comparing failure categories across GPT-o4-mini, Claude-Opus-4.7, and Gemini-3-Flash. They find a difficulty-dependent breakdown in which harder problems make models increasingly prone to misclassifying non-equivalent code as equivalent, a language-specific conservatism on Python for the best-performing model, partial reliance on similarity cues, and substantial run-to-run instability under identical settings, concluding that current LLMs do not reliably capture functional equivalence.

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal, Jay Iyer, Sai Pavan Kumar, Sai Kapil Kumar et al. When a large language model serving deployment runs out of key-value (KV) cache memory, operators can either shard weights and cache across more GPUs with tensor parallelism or shrink the cache in place with quantization and eviction, but the two options are almost never compared on a common cost axis. Using a profiled simulator calibrated on A100, A40, and H100 hardware, the authors place tensor-parallel degrees 1 to 8 and KV compression settings (16/8/4-bit, keep ratios down to 0.25) on a single cost-per-million-tokens versus latency axis for Llama-2 at 7B and 70B. Compression is cheaper by 1.20x to 2.00x at every level of memory relief tested, with no cost-equivalence crossover found; the deciding variable is model size relative to device memory, roughly 36B parameters for an 80 GB card, above which tensor parallelism becomes mandatory because weights rather than cache are the binding resource. Tensor parallelism is the only lever that improves latency (compression worsens per-token latency by 8 to 93% through batching contention), while compression is the only lever that multiplies capacity per dollar, at 16.5x versus 1.21x for an eightfold GPU spend.

Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction

Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen Using several large language models as a wisdom-of-the-crowd forecaster only helps if the models are genuinely diverse, and simply enlarging the crowd often adds redundant behavior. The authors propose a behavior-aware framework that characterizes each model by its reasoning traces on independent development tasks, clusters models by behavioral similarity, and selects cluster representatives for collective prediction, evaluating 25 LLMs on seven development benchmarks and two future-prediction benchmarks. A three-model medoid crowd chosen via K-means++ behavioral clustering outperforms voting over all 25 models on both prediction benchmarks while reducing model calls by 88% and inference cost by roughly 80%, suggesting that representative behavioral diversity, rather than crowd size or maximal diversity, is what makes an ensemble effective.

Mechanistic Circuit Identification for Controllable Data Generation

Nakyung Lee, Sangwoo Hong, Jungwoo Lee cross-listed Synthetic-data pipelines mostly steer generation through heuristic prompts, offering little insight into how individual samples interact with a model's learning dynamics. The authors connect training-dynamics-based data valuation with mechanistic interpretability (MI) by defining three utility axes for data (learnability, challenge, and alignment), uncovering model-internal circuits that causally govern each signal, and then using those circuits as controllable interfaces to steer generation toward utility-targeted data. A scheduler called SAMS (Stage-Aware Mechanistic Scheduling) sequences circuit-steered data according to the model's evolving optimization needs, and on multiple-choice QA tasks the approach yields more diverse, precisely controlled data that consistently improves downstream performance and calibration over prompt-based baselines.

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

Mohammad Mozaffari Sparsity, quantization, and low-rank approximation are usually applied to Large Language Models (LLMs) in isolation, and each alone hits an accuracy-efficiency wall. This thesis applies all three jointly, with sparsity cutting computation, quantization cutting memory bandwidth, and low-rank adapters recovering accuracy, across both pretraining and post-training: MKOR reduces curvature-update complexity from O(d^3) to O(d^2) and converges up to 1.85x faster than KFAC, SLoPe speeds training up to 1.25x with a double-pruned N:M sparse backward pass plus late low-rank adapters, OPTIMA reconstructs weights under static masks via globally optimal column-wise quadratic programs for up to 3.97% better zero-shot accuracy, and PATCH learns a dynamic hybrid sparsity ratio for up to 1.38x speedups when a fine-tuning budget is available. SLiM realizes the full combination in one shot using mathematically derived low-rank adapters, improving accuracy by up to 5.66% over prior compression methods and beating uncompressed dense models by 0.6% at equal parameter budgets.

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

Minsu Kim, Jianxun Lian, Xing Xie, Steven Euijong Whang Aligning large language models to human preferences often incurs an alignment tax, catastrophically forgetting pre-trained general capabilities, and prior work has treated this as an optimization or architectural problem rather than examining the preference data itself. BALIGN is a balanced data selection strategy that, from theoretical and empirical analysis of the preference optimization gradient, identifies three data-centric drivers of parameter drift: the reference model's log-probability margin, the token length difference between chosen and rejected responses, and TF-IDF similarity to general-capability corpora. These are aggregated into a composite risk score used to filter out samples that disrupt model parameters or offer little alignment utility, and experiments on standard human preference datasets show it preserves foundational capabilities without compromising alignment gains, consistently reaching the optimal Pareto frontier at minimal computational overhead.

Evaluating Multiple LLM Generations with Validated Task Coverage

Florian Le Bronnec, Rio Yokota Many LLM applications are most useful when they return several candidate outputs, but standard evaluation scores individual outputs or collapses multiple samples into a single success, missing whether the set contains genuinely different useful results. VTC-Bench is a five-domain benchmark built from real-data tasks where both output quality and task-relevant distinctness can be checked automatically and reproducibly without model-based judges, and its core metric, Validated Task Coverage (VTC), counts how many distinct useful results are obtained within k attempts. Across multiple models and inference settings, configurations that look strongest by single-draw quality are not necessarily those with the best coverage, and simple measures of output variation do not reliably recover task-relevant coverage.

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng et al. Small language models struggle with search-augmented reasoning, and on-policy distillation (OPD) from a teacher is hampered by the cost of collecting multi-turn search trajectories against a live retriever and the expense of task-specific teacher training, while distilling from an untuned off-the-shelf teacher caps the student at the teacher's ceiling and is unstable. OPDSearch+ drops teacher fine-tuning entirely: in stage one the student interacts with a live search engine and is distilled from a frozen instruct model via a per-position forward KL objective, and in stage two reinforcement learning refines the distilled student. The central claim is that the teacher reshapes the student's policy distribution so that subsequent RL converges to solutions RL alone cannot reach from scratch; across seven QA benchmarks a 3B student consistently outperforms all prior 3B RL baselines, with gains of 13.1% on HotpotQA and 8.5% on 2WikiMultihopQA.

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye Speculative decoding (SD) speeds up LLM serving by letting a lightweight draft model propose tokens that the target model verifies in one parallel pass, and multi-candidate schemes propose diverse candidate sets to raise acceptance rates. The analysis identifies a bottleneck called residual drift: once initial candidates are rejected, the residual target distribution diverges from the draft model's predictions, making later candidates ineffective and forcing expensive resampling. ResiSpec reshapes the proposal distribution during verification to keep the residual target mass inside the draft model's high-confidence regions, re-aligning the verification step mathematically without changing the exact output distribution. It reaches up to 1.92x speedup over state-of-the-art multi-candidate methods, and code is released.

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong LLM-as-a-judge evaluators are usually assessed by inter-judge agreement and robustness to surface perturbations, but reliability does not establish construct validity. The authors formalize validity as a two-dimensional profile, invariance S (the probability a verdict is unchanged under construct-preserving edits) and construct sensitivity R (the probability it changes under minimal construct-changing edits), show the two are independent with no faithful scalar summary, and measure them across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with generation, verification, and judging assigned to disjoint model families. Judges matched at invariance S >= 0.90 average S = 0.945 but R = 0.319, and sensitivity to scope edits (0.383) exceeds sensitivity to strength edits (0.262) for every judge. An audit of five public label sets finds surface-only predictors reproduce 55-67% of labels, including 67.4% of MT-Bench human votes, so high agreement can coexist with weak sensitivity to the construct being evaluated.

Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning

Hang Chen, Jiaying Zhu, Wenya Wang cross-listed Mechanistic localization isolates critical parameters with interpretability methods and then restricts supervised fine-tuning (SFT) to them in a locating-then-tuning paradigm, but because interpretability is retrospective, the neurons identified in the pre-SFT model can differ drastically from those governing the fine-tuned model on novel tasks, biasing the tuning. The proposed forward-looking localization framework estimates the post-SFT interpretability state from only the pre-SFT parameters and the target dataset, modeling SFT as continuous parameter evolution and using a Taylor expansion to link the post-tuning mechanistic objective to the pre-SFT model's gradients, with dual-granularity pipelines at the neuron and component levels. Experiments show the predicted localization provides better SFT guidance than static pre-SFT localization and scales robustly across increasing model sizes.

When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study

Mohit Singh Chauhan, Vipin Gyanchandani, Dylan Bouchard cross-listed Uncertainty quantification (UQ) methods are used to detect hallucinations in closed-book settings with no reference evidence at inference time, and prior work proposed learned ensembles of UQ signals, but their robustness has received little empirical scrutiny. The study trains a classifier over heterogeneous UQ scorer outputs on a small labeled set of LLM responses and applies it to out-of-sample hallucination classification without retrieval or tools, testing four LLMs, nine datasets, and three generation regimes (short-form QA, long-form generation, and code generation) along the axes of sample efficiency, in-domain transfer, and regime dependence. Supervised ensembles beat the best individual scorer in 30 of 32 settings, with gains appearing from as few as 100 labeled instances, and retain most of their advantage under distribution shift, winning 23 of 28 transfer settings. Sampling-based black-box ensembles are nearly as effective as full ensembles, while single-generation white-box ensembles offer limited benefit.

Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems

Wonung Kim, Hyunmin Choi, Minsu Kim, Jaehong Cho, Yeongwook Kim, Jongse Park cross-listed System-level simulators for large language model serving cannot keep pace with rapidly evolving mechanisms such as agentic workflows and disaggregated serving, since each new feature demands an invasive rewrite of a monolithic simulation pipeline. Simthesizer introduces a composable simulator infrastructure that expresses the full serving workflow, including its control decisions, as a unified dynamic graph, and pairs it with a harnessed coding agent that lowers natural-language feature requests onto that abstraction under simulator-specific guardrails and fidelity validation, evolving one shared simulator rather than building a new one per feature. Extensions built this way track a vLLM-based real system with 2.51% average throughput error versus 6.03% for extensions built on existing simulators, and the simulator runs up to 284.96x faster than LLMServingSim2.0 and 23.19x faster than Vidur on identical workloads.

The RAT: A Unified Bayesian Model for RAG Evaluation

Pius von D\"{a}niken, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu cross-listed Evaluating Retrieval-Augmented Generation (RAG) systems requires understanding how components interact and how errors propagate through the pipeline, not just end-to-end correctness. The authors introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correctness, factorized along the pipeline's information flow, so it separates whether the user received a correct answer from whether the generator behaved appropriately given what retrieval returned. Applied to 27 RAG configurations across three datasets, three retrievers, and three generators, the conditional decomposition exposes substantial behavioral differences between systems that look equivalent under marginal metrics. They further show, with an information-theoretic explanation, that retrieval-success annotations are more informative than task-success annotations for estimating policy adherence, and they extend the model to treat LLM-as-a-judge outputs as calibrated noisy observations that can be combined with limited human labels.

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Miao Liu, Zhizhe Liu cross-listed Large language models deployed as AI analysts over financial disclosures are usually evaluated on what they can retrieve rather than on whether retrieved information changes their judgments. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, the authors find a retrieval-integration gap: a risk disclosure's influence on investment judgments falls to the experimental noise floor even while direct retrieval of it remains accurate, a pattern that replicates across model families, judgment tasks, and experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly carry disclosures into judgments, and that workflow architecture decides whether this succeeds: chunk-and-summarize pipelines evict the relevant information, whereas a targeted, structured restatement placed adjacent to the decision restores its influence.

ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

Juntong Wu, Yifei Liu, Junyi Chen, Siqi Fan, Chaoran Feng, Minghao Li et al. Mixture-of-Experts (MoE) models keep per-token compute bounded through sparse expert activation, but low-latency serving spans two phases with different bottlenecks: prefill is dominated by per-token expert computation, while decode is limited by memory traffic from the batch-wide set of activated experts, and existing training-free accelerators optimize only one of these proxies and drop or only implicitly approximate the excluded experts' contributions. ExFold casts both phases as a single budgeted output-approximation problem: execute only a phase-specific constrained expert set and project the contribution of excluded experts onto retained ones using scalar projectors calibrated on unlabeled data, motivated by the observation that many expert outputs are directionally aligned but differ in magnitude. Prefill acceleration becomes token-level Top-K folding and decode acceleration becomes batch-level expert-pool folding, with both phases sharing one folding mechanism. Implemented as a plug-and-play vLLM plugin with a lightweight CUDA kernel, ExFold delivers up to 1.41x time-to-first-token and 2.45x time-per-output-token speedups while retaining about 99% of original average quality.

FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

Gongwei Lee, Ji Liu, Juncheng Jia, Ji Wu Quantizing large language model weights with a single uniform bit-width, or with crude heuristics for judging which layers tolerate compression, tends to cost accuracy. FAMPWQ scores each layer's sensitivity to quantization using a Fisher information metric, then trains a reinforcement learning agent to allocate per-layer bit-widths under that signal so models fit on commodity GPUs. Across 7 models and 5 benchmarks it beat 7 prior methods, reducing perplexity by up to 3.39 and raising accuracy by up to 6.87 percentage points, with win rates as high as 76% in LLM-as-a-judge comparisons.

The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

Elle Language models are known to perform worse on non-standard English dialects, but it has been unclear where in the pipeline that penalty originates. Using parallel corpora that hold meaning constant while varying surface form, the authors confirm models treat matched Standard American English (SAE) and dialectal texts as semantically equivalent, then trace representational gaps through tokenization, pre-training, post-training, and inference. The penalty accumulates at every stage rather than in any single one: swapping subword segmentation for a character-level counterfactual tokenizer removes neither the input and output asymmetries nor the accuracy gap, dialect pairs produce more divergent gradient updates during pre-training than pairs of entirely unrelated SAE documents, and reward models show unstable, context-dependent dialect preferences that penalize full reasoning traces in task- and model-specific ways.

Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation

Peng Liu, Huibing Zeng, Yiqun Zhang, Yang Yi, Jigang Wu Pruning Transformer language models usually requires full gradient computation for importance estimation plus a round of full-parameter fine-tuning first, which is exactly what memory-constrained deployments cannot afford. REP-LIE instead estimates weight importance from the gradients of LoRA low-rank adapter matrices during fine-tuning, damping the noise in those estimates with a stability score that guides iterative removal of unimportant parameters, and repairs the pruned model with lightweight adapter updates rather than full optimization. Experiments on medium-scale encoders and on LLaMA-7B and Mistral-7B report performance competitive with existing pruning methods at substantially lower resource cost.

Unsupervised Post-Training of Foundation Models: A Survey

Yijie Xu, Qianyi Cai, Huizai Yao, Yili Wang, Tianfu Wang, Cehao Yang et al. Post-training a foundation model normally needs human labels, preference data, a stronger teacher, or an executable verifier. This survey defines Unsupervised Post-Training (UPT) as weight-updating adaptation on unlabeled inputs where the learning signal comes from artifacts of the same model lineage rather than any external oracle, and catalogs 80 strictly unsupervised methods grouped by what supplies the update: a prediction statistic, a relation between samples, a self-generated target, or an internal evaluator. Beyond the inventory, the authors argue that the choice of internal signal combined with task structure determines whether such training genuinely improves a model or recursively amplifies its own errors, and propose an Input Visibility by Update Persistence grid for mapping deployment regimes and selecting a method.

D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng et al. Multi-teacher on-policy distillation compresses several domain-expert teachers into one student by minimizing per-domain reverse Kullback-Leibler divergence on the student's own rollouts, but the data mixture across domains is normally fixed in advance even though domains converge at very different rates. D3-MOPD reuses the reverse-KL values the training loop already computes: an off-process watcher tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and adjusts sampling ratios online without touching the training loop or adding overhead. Distilling a Qwen3.6-35B-A3B student from four expert teachers, it closes 97% of the average student-to-teacher gap versus 63% for a fixed mixture, reaches the fixed-mixture peak with roughly one third the rollout steps, and beats the specialist teachers on three of seven benchmarks.

A Primer on Computational Semantics for Artificial Intelligence Systems

Casey Kennington As transformer-based language models such as ChatGPT and Gemini are adopted for ever more uses, it matters to understand how such models learn and represent linguistic meaning, and what language itself is. The document surveys how semantics is approached across different scientific and philosophical fields and explains three primary semantic theories: formal semantics, grounded semantics, and distributional semantics. It then contrasts how transformer-based language models acquire meaning with how humans learn language, giving readers a grounded vocabulary for reasoning about what these models do and do not capture.

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie, Andrea Giovannini, Katja Hose GPUs increasingly accelerate database systems, but query-specific peak performance still depends on hand-written kernels, and existing LLM kernel benchmarks focus on machine learning operators rather than the irregular, heterogeneous, data-movement-heavy operators databases require. DataKernelBench translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimise either the core tensor-bound snippet or the full query in CUDA or Triton, using execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100, the strongest full-query CUDA configuration reaches a 2.11× speedup over torch.compile at a full pass rate; higher-performing implementations commonly rely on kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialisation, and workload context matters more than hardware context. Extending TorchPlan with Dask-cuDF for on-demand partition loading handles data larger than GPU memory, reaching a 2.54× speedup on SF100 across four H100s.

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

Xiulin Yang, Ethan Gotlieb Wilcox, Catherine Arnett Fair crosslingual comparison of language models remains a fundamental challenge in multilingual natural language processing, because existing studies rely on a variety of downstream tasks and intrinsic metrics with differing theoretical justifications and little empirical check on whether they support meaningful conclusions across languages. The authors train controlled monolingual models on parallel data while varying tokenizer vocabulary size and model size, validate their findings on multilingual LLMs, and discuss the obstacles to achieving comparable downstream evaluation across languages. They find that several widely used normalised metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences, whereas sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.

Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures

Molka Chkir, Syed Muhammad Danish, Jos H\"oll, Arghavan Asad Growing LLM adoption has raised concerns about the energy cost of inference, and this empirical study characterises decode-phase energy consumption across representative open-source models using Multi-Head Attention (MHA), Grouped Query Attention (GQA), and GQA combined with Sliding Window Attention (SWA). Four models are evaluated across context lengths, batch sizes, and generation workloads, with GPU energy measured via NVIDIA hardware counters, to isolate the effects of context length, attention mechanism, Key-Value (KV) cache growth, and batching. The attention mechanism is the primary factor governing how decode energy scales with context length: MHA models show substantially steeper energy growth than GQA models, while GQA with SWA keeps energy nearly constant; model size mainly determines absolute energy consumption, and batching reduces both energy per generated token and request latency by up to 87%, offering practical guidance for choosing energy-efficient architectures and inference configurations.

Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting

Weibin Cai, Reza Zafarani Retrieval-Augmented Generation (RAG) efficiency work usually targets the generation step, yet the authors show empirically that the system bottleneck shifts to upstream reranking under high query rates or large reranking budgets. PACE (Prioritized Adaptive Coverage of Evidence) is a training-free framework that reorders candidates by marginal evidence coverage, favoring documents that are relevant, complementary, and useful for multi-hop chains, then adapts the reranking budget to the relative load on the reranker versus the language model; the coverage objective is shown to be monotone submodular, so greedy selection carries a (1-1/e) approximation guarantee. On three multi-hop question-answering datasets and serving simulations, PACE improves evidence recall and reduces p95 latency under ranking-heavy workloads, with the central observation that an evidence-dense top of the candidate list yields higher final recall with fewer reranked documents.

SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation

Ben Lagnese, Manas Gaur Graph-based retrieval-augmented generation can exploit entity relationships in a knowledge graph, but training a supervised graph retriever needs labeled question-answer data that rarely exists for a newly built graph. SelfGraphRAG generates question-answer pairs directly from graph structure, covering multi-hop paths and local neighborhoods, and uses them to train a query-conditioned graph retriever without manual annotation. On multi-hop question-answering and classification benchmarks, the approach improves retrieval precision and downstream reasoning over embedding-based baselines, suggesting graph structure alone can supply useful retriever supervision.

The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers

Samuele Vallisa, Federico Ravenda, Claudio Palominos, Rui He, Andrea Raballo, Antonietta Mira et al. Transformer token representations trace trajectories through high-dimensional space that concentrate on lower-dimensional sub-manifolds, and this work asks whether a token's part-of-speech (PoS) shapes the local geometry of that manifold. Measuring intrinsic dimensionality (ID) layer by layer in the encoders ModernBERT and bigbird-roberta-large and the decoders gemma-2-2B and Llama-3.2-3B, the authors find that closed-class words such as function words expand in dimensionality earlier and collapse sooner than open-class words, that these changes are explained by reorganization of each token's neighborhood within the sentence, and that encoder and decoder families evolve differently in line with how each integrates context. Geometric features alone suffice to recover a token's grammatical role, and the authors use them to trace how the semantic content of each PoS category evolves across layers on a downstream classification task.

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

Ehsan Jokar Nearly every competitive 4-bit large language model quantization pipeline first applies a function-preserving linear transform (rotation, scaling, permutation, or non-orthogonal affine map) before rounding, yet no survey has covered this transform stage. The authors formalize what they call the Great Inversion: classical transform coding rewards concentrating energy in a few coordinates because it can allocate different bit widths per coordinate, whereas the grouped shared-scale quantization that deployed matrix instructions actually perform rewards flattening within each group, which Hadamard-style incoherence approaches. They prove these two prescriptions are opposed under within-group majorization and show the right target depends jointly on the allocation regime and the number format (FP4, MXFP4, NVFP4). The survey covers 200 works and classifies 43 transform methods by structure, data-awareness, searched-versus-constructed design, and runtime cost, recording how they compose with GPTQ rounding and distilling a first-choice guide per deployment regime.

Trust the Mass: Forced Weights in KV-Cache Eviction

Jack Shi, Jerry Gu Sparse-attention and key-value (KV) cache eviction rules keep a subset of keys and renormalize attention over them, and this work asks how much the choice of subset actually matters. Exhaustively enumerating the best subset on 168,192 attention rows from five models shows that simply keeping the largest weights is near-optimal, since the best possible subset closes only a median 2 to 5% of the remaining gap to full attention, which implies published margins between eviction methods come from elsewhere. Measuring bytes actually held reveals that the strongest query-agnostic methods store per-head selections as masks over the full cache in the shared evaluation pipeline, that enforcing a real budget on one fixed selection costs 14 to 62 benchmark points, and that an 87.6-point retrieval margin traces to rankings computed while the question is visible. ContourKV, a training-free allocator built from the dropped-mass statistic, wins 93 of 160 paired comparisons against the prior state of the art and ties the strongest budget-enforcing baseline at matched byte count.

Output Dilution: Redundant but Fragile Representations in MoE Models

Orion Reblitz-Richardson Mixture-of-Experts (MoE) models appear to encode moral content as robustly as dense models but turn out to be far more fragile. In OLMoE-1B-7B, linear probes recover moral valence from nearly every expert-layer combination with mean peak-layer accuracy above 90%, yet these representations collapse under activation noise that a matched dense model tolerates, a 4.2-fold difference in robustness. The cause is output dilution: because the MoE block averages across active experts before writing to the residual stream, its feedforward contribution is nearly two orders of magnitude smaller than a dense MLP's, so the information survives aggregation but at a scale easily overwhelmed by perturbation, while routing itself stays stable. Checkpoint trajectories show experts never specialize and probe accuracy saturates within a few thousand training steps, indicating the effect is architectural rather than learned.

From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection

Zhibo Hou, Fan Zhao, Zhiyu An, Wan Du Keeping large language models current requires continually injecting new facts, but supervised fine-tuning (SFT) memorizes facts in their training format and generalizes poorly to paraphrases, combinations of documents, and reasoning over them. Golden-GRPO Injection (GRIN) is a three-stage self-learning framework built on Golden-GRPO, a mixed-policy reinforcement learning algorithm that injects a golden answer into the rollout group so there is a learning signal even when on-policy rollouts all fail on a novel fact. The authors also introduce Blank and Counter, document-level benchmarks for novel acquisition and counterfactual overwrite respectively, each testing single-fact recall, multi-source retrieval, and inferential reasoning. GRIN substantially outperforms SFT and mixed-policy RL baselines on the harder question types while matching them on basic fact recall.

Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models

Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, Junpei Komiyama Diffusion Language Models (DLMs) generate text by iterative denoising rather than left-to-right decoding, and PDC (Prefix-Denoising Consistency) exploits a signal specific to that process for test-time self-verification: when a sample is split at an intermediate position and the remainder is regenerated conditioned on the fixed prefix, correct outputs are reproduced more stably than incorrect ones. On mathematical and commonsense reasoning benchmarks, PDC consistently improves over the initial sample and beats independent generations at matched compute, and the gains hold across different unmasking strategies and parameter settings.

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng Model merging builds multi-task models without additional training, but performance degrades when task-specific features are superposed in the same parameters, which conventional decomposition methods cannot cleanly separate. The proposed framework uses Sparse Autoencoders (SAEs) to project task vectors into a high-dimensional sparse feature space where task directions can be disentangled before fusion, and a lightweight Group-Ranked Zeroth-Order Optimizer (GR-ZOO) picks task-critical layers so that only those are merged. On Qwen2.5-1.5B and Qwen2.5-7B, the method outperforms Task Arithmetic, TIES-Merge, DARE, Fisher-Merge, and several recent training-free approaches across mathematical reasoning, code generation, instruction following, and general knowledge, including a 2.78% gain over the strongest baseline in a highly conflicting four-task setting on the 1.5B model.

Learning New Facts with QLoRA: An Acquisition-Retention Frontier

Estelle Zheng, S\'ebastien Warichet, Emmanuel Helbert, Christophe Cerisara Parameter-efficient fine-tuning is often assumed to protect pretrained capabilities because it touches few weights, and this work tests that assumption on a controlled benchmark derived from OpenStreetMap where Qwen3-4B must learn anonymized geographic facts while keeping unrelated skills. Comparing full fine-tuning against quantized low-rank adaptation (QLoRA) at ranks 8 through 64, the authors find that adapter rank traces a clear acquisition-retention frontier: low ranks preserve out-of-domain performance but learn fewer facts, while higher ranks improve paraphrase generalization of the new facts at a growing cost to unrelated benchmarks. Full fine-tuning acts as a conservative baseline that retains general ability but never reaches the highest acquisition regime, and weight-space and spectral diagnostics mirror the behavioral trade-off, which is weaker in a separate math-adaptation experiment where training reinforces existing skills rather than installing new associations.

When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies

Abhinav Havaldar, Enrico Santus Retrieval-augmented generation (RAG) is widely assumed to patch gaps in a language model's factual knowledge, but it is unclear whether retrieval helps uniformly. Using a benchmark of roughly 2,000 public companies drawn from global equity indices, the authors test six large language models on four atomic attributes under no-context, perfect-context, misleading-context, and distraction-context conditions. Accuracy without context varies sharply by company geography, and perfect context narrows but does not close these gaps, with gains correlated to baseline accuracy, suggesting retrieval effectiveness is coupled to what the model already represents; models also frequently copy misleading context, and larger models raise overall accuracy without removing the structural disparities.

Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale

Peichun Hua, Danyang Chen, Junan Zhang, Haifeng Sun, Jingyu Wang, Diwen Xue et al. cross-listed Hosted retrieval-augmented generation (RAG) and semantic search over provider-owned corpora must hide each user's query and chosen result while only releasing documents the user is authorized to see, and existing cryptographic solutions either scan the full corpus per query or trade quality for speed by touching a few clusters. The authors repurpose learned deep hashing as a private filter: a randomized binary code points the provider to a short candidate list, after which encrypted reranking and oblivious key transfer conceal the exact query and final selection. With 200-500 candidates the shortlist closely matches full-corpus retrieval across five zero-shot corpora of 25K to 5.4M documents, and on the 2.68M-passage NQ corpus over a 10-Gbps link the protocol adds only 0.73 seconds, about 10%, to a 128-token Qwen3-32B RAG pipeline; the released code satisfies directional metric differential privacy and reduces embedding-inversion and property-inference leakage.

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash, Akshay Krishnamurthy Reinforcement learning post-training works best on an already capable base model, prompting the question of whether standard supervised fine-tuning produces the best starting point for it. TailSFT is a simple modification that filters out sequences the model already fits well during supervised fine-tuning, concentrating learning on the under-modeled tail of the data distribution, with the filtering criterion justified by controlled experiments and theory. On OLMo-3 7B it improves pass@16 on math and coding evaluations by up to 17 points absolute at minimal extra cost, and those higher-coverage checkpoints yield up to 4 points absolute pass@1 gains in subsequent GRPO runs; a lightweight diagnostic identifies settings where the method is likely to help.

Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models

Ty Chermsirivatana, John MacCormick Trading model size for extra inference compute is a standard lever, but the picture changes when outputs must satisfy a strict grammar. Using Qwen2.5-Instruct models from 0.5B to 7B parameters at 4-bit precision on the Spider text-to-SQL benchmark, the authors compare grammar-constrained beam search against sampling several constrained outputs and voting on their execution results. Both methods raise accuracy, especially for the smallest models, but moving to a larger model generally beats spending the equivalent compute on a smaller one, and beam search wins over sample-and-vote at a matched budget — the reverse of what has been reported for unconstrained decoding.

Localize-Then-Decide Guarantees for LLM Judgments

Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin Confidence thresholding can give statistical guarantees that an LLM judge agrees with human preference, but it rests on the assumption that higher estimated confidence means lower disagreement risk — an assumption that breaks down as the number of candidate responses grows and probability mass spreads across alternatives. The Localize-Then-Decide framework first uses conformal prediction to localize a small shortlist that contains the human-preferred response with high probability, then applies a calibrated confidence rule to select one item from that shortlist or abstain. Restoring the monotonic link between confidence and risk this way yields consistently higher guarantee success rates and substantially higher coverage than single-stage baselines across several datasets, judge models and candidate-set sizes.

Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training

Qiankai Xu, Qiguang Chen, Zixin Su, Wenhao Huang, Yue Gao, Jiaheng Liu et al. Recent synthetic-data work reconstructs the reasoning behind existing text instead of rewriting it, but works on short web passages and recovers only local thoughts. The pipeline here unfolds each scientific paper into a multi-turn generation trajectory — a writing request, a global plan, and pre-writing deliberation for every section — with all section text and the abstract kept verbatim from the source, yielding a continued pre-training (CPT) corpus roughly twice the size of the original text from quality-filtered arXiv papers. The same reverse construction produces supervised fine-tuning data and PAW-Bench, an academic-writing benchmark whose tasks carry their own rubrics and checklists. In controlled experiments, CPT on the corpus followed by public SFT improves writing benchmarks broadly while preserving general reasoning and improving long-document reading, and the writing gain survives even when every model receives a dedicated writing SFT set.

Skill Issue: Are Skills Language-Invariant in LLMs?

Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan et al. Language models are known to access knowledge unevenly across languages; the question here is whether their skills differ too, measured separately from knowledge and general benchmark scores. Two instances of the same model play text-based games against each other through different language interfaces, holding model, opponent, rules, state space and available actions fixed, using a multilingual extension of TextArena across eight languages and six games covering spatial reasoning, imperfect information, resource allocation and repeated interaction. Three open-weight models show markedly different playing strength depending only on interface language, with systematic variation in win-loss margins, invalid actions and strategic tendencies, and in some settings changing just the intermediate reasoning language recovers much of the lost performance, suggesting language affects distinct stages of the decision process.

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic Production pipelines that use language models to score outputs, filter content or gate iterative refinement typically assume each judgment is independent of earlier evaluations. Testing three prompt conditions — no metadata, revision framing, and anchored metadata carrying revision, attempt and prior-score fields — over 185,271 successful evaluations, prior scores included merely as context systematically shift ratings toward their value, with seven of eight models showing the effect and Cohen's d reaching 0.71. On categorical industry data with human-labelled ground truth, anchored metadata blocked 48% of error corrections and flipped 10.18% of correct judgments to an assigned wrong label; neither chain-of-thought nor an instruction to disregard the metadata removed the total effect, and token-level probes suggest a threshold-like response where the presence of an anchor matters far more than its value.

From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations

Ping Wang, Xiangguo Sun, Bingbing Xu, Guocong Li, Xiaofeng Meng Large language models often answer confidently and wrongly when the user's question itself contains a misleading premise, a vulnerability that hallucination-mitigation work misses because it assumes inputs are factually reliable. DEDUCE turns the model from a passive responder into an active corrector in three stages: detect errors via fine-grained fact extraction and verification, devise correction strategies through multi-perspective deliberation, and then correct the misconception while delivering an answer. Alongside MisFactQA, a new dataset with factual errors of varying severity plus robustness metrics, experiments on TruthfulQA, FalseQA, and MisFactQA show gains in both answer accuracy and error-correction ability that hold across the Qwen, LLaMA, and Gemma model families.

One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

Muge Zhang, Aaron Jencks, Krishna Badikela, Yulia Tsvetkov, Sachin Kumar Multilingual models share knowledge across languages through a shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems; romanization and International Phonetic Alphabet (IPA) transcription are the usual remedies but are rarely compared head to head, and prior work focused on encoder-only models adapted after pretraining. A controlled comparison of orthographic text, IPA, and romanized input in autoregressive pretraining covers three scales (467M, 709M, and 1.03B parameters) and eight languages arranged in four typologically motivated pairs, evaluated on many downstream tasks in seen and unseen languages. Romanized pretraining yields the strongest cross-lingual transfer, and its advantage over native-script text widens with scale, with IPA landing in between; fine-tuning an already text-pretrained model on romanized data instead hurts languages the base model already covered, so romanization is better treated as a pretraining design choice than a post hoc fix.

Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation

Pankaj Kumar, Subhankar Mishra Knowledge Graph Question Answering (KGQA) pipelines built on graph neural networks run entity linking, subgraph retrieval, GNN reasoning, and answer generation in sequence, and end-to-end robustness scores collapse stage-level failures into one number that hides where the brittleness lives. A stage-isolation protocol applies two answer-preserving adversarial question rewrites verified against the knowledge graph, Compositional Restructuring and Relation Synonym Swap, which target different stages while leaving entity seeds intact. On ComplexWebQuestions and WebQSP the GNN reasoning stage stays near baseline accuracy whenever the subgraph is intact, while subgraph construction accounts for over 99% of the end-to-end collapse under Compositional Restructuring even though the gold answer is still present in 74% of retrieved subgraphs — a gap between answer presence and answer reachability that end-to-end metrics cannot detect, and one that puts the mitigation target at retrieval rather than the reasoning model.

How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

Aida Usmanova, Zangir Iklassov, Markus Leippold, Ricardo Usbeck cross-listed Automated fact-checking (AFC) systems retrieve evidence and then predict a claim's veracity, but they are typically developed and evaluated on a single benchmark, often without simple baselines, so nobody has cross-evaluated the full retrieve-then-verify pipeline across domains. Nine systems spanning random and sparse baselines, fine-tuned transformers, zero-shot large language models, and the two top entries from the AVeriTeC 2025 shared task are benchmarked on four datasets covering scientific, open-web, and climate domains. Rankings prove strongly domain- and metric-dependent, with the best SciFact model dropping from 0.70 to 0.31 macro-F1 on ClimateCheck and the AVeriTeC winner and runner-up swapping places depending on metric, claim-only and fine-tuned models can beat sophisticated systems when retrieved evidence is noisy, and swapping retrieved evidence for gold annotations lifts veracity accuracy by 14 to 22 points, confirming retrieval as the primary bottleneck.

When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili Sparse autoencoders (SAEs) are a standard tool for interpreting the internal representations of large language models, but whether they still behave correctly after the host model is compressed has gone largely untested. Theory here shows that for a fixed SAE the damage from pruning is governed by perturbation energy, a covariance-weighted norm, which explains why magnitude pruning distorts the learned representation space by ignoring activation geometry, whereas activation-aware methods such as Wanda and SparseGPT implicitly control perturbation energy and are substantially better at preserving SAE functionality. Experiments across four architectures also reveal that middle layers are consistently far more pruning-sensitive than early or late ones, and a layer-wise sparsity allocation derived from that insight attains lower perplexity at the same average sparsity.

When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs

Yao Fu, Lijia Huang, Xiaomin Li, Runchao Li, Yu Yin, Kenneth A. Loparo Studies of personality in large language models (LLMs) using the Myers-Briggs Type Indicator (MBTI) have focused on full-precision models and final outputs, ignoring the quantized models that dominate memory-constrained deployment. The authors run a systematic MBTI analysis of open-source LLMs across precisions including 4-bit GPTQ and AWQ and 2-bit AQLM variants, examine how personality decisions emerge layer by layer through option-level entropy and confidence-gap dynamics, and introduce Uncertainty-Amplified Layer Decoding (UALD) to study decoding-induced personality drift at inference time. They find that ENFJ dominates across model families and precisions, that 4-bit quantization largely preserves coarse personality structure while 2-bit quantization disrupts fine-grained prompt consistency and cross-precision agreement, that personality decisions crystallize in upper layers after substantial early-layer ambiguity, and that decoding choices can shift personality while personality-aligned conditioning improves robustness.

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland Orthogonal optimizers such as Muon speed up large language model pretraining relative to Adam, but the mechanism has been unclear. The authors probe Transformer loss landscapes at checkpoints along real training runs, decomposing each momentum buffer into its singular directions and estimating the loss-optimal step size along each on held-out data; the resulting spectral profile is stable across batches, training stages, optimizers, and scales, with a volatile head operating at the edge of stability that tolerates only small steps and a bulk that permits much larger ones. This spectral allocation view explains why Muon beats Adam, which beats SGD, and exposes that Muon's uniform scaling still underuses the bulk, motivating Spectral-Aware Muon (SAMuon), which holds the head at Muon's scale and amplifies the bulk using a static spectral prior, plus a cheaper SAMuon-lite variant using rank-one power iteration, neither adding persistent optimizer state or notable extra FLOPs. On modded-nanogpt models from 124M to 1B parameters both variants beat tuned AdamW and Muon in every tested configuration, with SAMuon needing 13.3% to 24.0% fewer training tokens than Muon to reach the same validation loss and SAMuon-lite keeping most of the gain at near-zero wall-clock overhead.
16 more specialized papers

Unclassified 62

Method, Mind, and Morality: How People Make Sense of Artificial Intelligence

Jacy Reese Anthis, Erik Brynjolfsson, James Evans cross-listed No summary available — see the abstract on arXiv.

Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit

Anqi Peter Li No summary available — see the abstract on arXiv.

CRAMER: Control via Request-Aware Masking for Editing Recommenders

Zhiyuan Julian Su, Naihe Feng, Zhen Luther Qin, Ga Wu cross-listed No summary available — see the abstract on arXiv.

GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

Yiqun Sun, Junyu Chen, Pengfei Wei, Lawrence B. Hsieh cross-listed No summary available — see the abstract on arXiv.

Adaptive Triggering for Bias Correction in LLM Reasoning

Nayoung Kim, Mickey Mancenido, Huan Liu No summary available — see the abstract on arXiv.

A meta-algorithm for ab initio reconstruction of complex mixtures in cryo-EM

Alkin Kaz, Arda Kaz, Ellen D. Zhong cross-listed No summary available — see the abstract on arXiv.

Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks

Andrey Labunets No summary available — see the abstract on arXiv.

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang No summary available — see the abstract on arXiv.

Token-Oriented Semantic Communication with Pretrained Vision Transformers

Jiwoong Im, Minwoo Kim, Jaeho Lee, Yo-Seb Jeon, Yongjune Kim cross-listed No summary available — see the abstract on arXiv.

BVR Sim: An Open and High-Throughput Environment for Heterogeneous Air-Combat Reinforcement Learning

Haocheng Sun (Beijing University of Posts, Telecommunications), Mulai Tan (Air Force Engineering University) cross-listed No summary available — see the abstract on arXiv.

Data-driven Effective Modeling of Stochastic Chemical Reaction Networks

Yuan Chen, Weize Mao, Dongbin Xiu cross-listed No summary available — see the abstract on arXiv.

DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models

Minhae Oh, Nakyung Lee, Jungwoo Lee No summary available — see the abstract on arXiv.

Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness

Yi Chen, Hanna Hsieh, Shuhong Liu, Chuanbo Hua, Zihan Ma, Kun Wang et al. cross-listed No summary available — see the abstract on arXiv.

Joint Initialization of Flux Networks and Effective Multiplication Factor for Physics-Informed Neural Networks Solving Neutron Diffusion Problems

Qin Hang, Yangdi Yi, Jiayi Li, Xu Wang, Heng Zhang No summary available — see the abstract on arXiv.

Energy Yield and Lifetime Climate Classification via Machine Learning for Optimizing Photovoltaic Module Design and Materials

Youri Blom, Sofia Dutto, Alexandru Costache, Rowan Richie, Ruben Pelsser, Wesley Berger et al. cross-listed No summary available — see the abstract on arXiv.

MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu, Jiaqi Wang, Chenghao Deng, Xiayimei Han et al. No summary available — see the abstract on arXiv.

Training Alignment Auditors via Reinforcement Learning

Paul Rosu, Rowan Wang cross-listed No summary available — see the abstract on arXiv.

Resolving Multi-Modal Regression by Difference-Quotient-Based Clustering:Fast Coarse Conditional-Label Assignment

Huang Weiquan No summary available — see the abstract on arXiv.

Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

Shunxing Yan, Fang Yao cross-listed No summary available — see the abstract on arXiv.

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication

Ziyuan Wang (Steven), Yifan Sui (Steven), Wei Wei (Steven), Wenjie Xin (Steven), Zekai Zhang (Steven), Xiangwang Hou (Steven) et al. cross-listed No summary available — see the abstract on arXiv.

VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text

Trieu Hai Nguyen, Van-Dung Hoang No summary available — see the abstract on arXiv.

PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning

Rongchen Zhao, Yu Chen, Juyuan Wang, Zhouting Mo, Jianxing Yu, Wenqing Chen et al. cross-listed No summary available — see the abstract on arXiv.

ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains

Jinpu Jiang, Xuan Wu, Wenhao Song, Bo Yang, You Zhou, Hongwei Ge et al. No summary available — see the abstract on arXiv.

A Storage-Retrieval Gap in Parametric Knowledge Graph Memory

Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker Tresp No summary available — see the abstract on arXiv.

FedQoS: Federated QoS-Risk Learning for Heterogeneous Indoor-Outdoor Access Selection

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Zerihun Huruy, Vu Nguyen Ha, Symeon Chatzinotas No summary available — see the abstract on arXiv.

A Multi-View Coupled Tensor Decomposition for Lightweight Online Adaptive Traffic Prediction

Quan Yu, Jie Ni, Yu-Hong Dai, Xiongjun Zhang cross-listed No summary available — see the abstract on arXiv.

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu, Yi Chang cross-listed No summary available — see the abstract on arXiv.

Conditional Total Correlation and the Serial Depth of Adaptive Parallel Sampling

Chuling Wen, Weijie Liang, Jian Lu cross-listed No summary available — see the abstract on arXiv.

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

Caixing Wang, Zhibo Chen, Yue Wang cross-listed No summary available — see the abstract on arXiv.

TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving

Hongqiu Ni, Han Tian, Chi Zhang, Guopeng Li, Haisheng Tan No summary available — see the abstract on arXiv.

Adaptive Hybrid Subspace Levenberg Marquardt Algorithm with Adequacy Monitor for Large Scale Least Squares Problems

M. Duc Hoang, Timothy J. Lewis cross-listed No summary available — see the abstract on arXiv.

ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives

Jihao Zhu, Zhiwei Yang, Wenxiao Zhang, Junqian Zhao, Qi You, Fangqi Wang et al. No summary available — see the abstract on arXiv.

Resilient Decentralized Wireless Federated Learning via Gradient Tracking with AdamW

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Vu Nguyen Ha, Symeon Chatzinotas No summary available — see the abstract on arXiv.

CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

Rana Muhammad Ahmed, Sabahat Abbas cross-listed No summary available — see the abstract on arXiv.

Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference

Jiarui Hu, Zhiyuan Wen, Xiaoyun Liu, Jiaxing Shen, Yu Yang No summary available — see the abstract on arXiv.

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer, Marco Simnacher, Jordan F. Safer, Sumaiya Iqbal et al. No summary available — see the abstract on arXiv.

Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules

Liviu Aolaritei, Lucas L\'evy, Francis Bach, Michael I. Jordan No summary available — see the abstract on arXiv.

When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory

Kazuki Nakayashiki cross-listed No summary available — see the abstract on arXiv.

Virgil: Navigating Explainability for Transformer-based Language Models

Martino Ciaperoni, Sezer Kutluk, Benedetta Muscato, Marta Marchiori Manerba, Fosca Giannotti No summary available — see the abstract on arXiv.

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen No summary available — see the abstract on arXiv.

Physics-Informed Foresight Pruning for Sparse PINN Solvers of Nonlinear PDEs

Ahmad Ishaque Karimi, Uvini Balasuriya Mudiyanselage, Kookjin Lee No summary available — see the abstract on arXiv.

Controllable Affective Generation via Latent Vector Steering

Xixian Yong, Siyuan Chang, Yingying Zhang, Xian Wu, Xiao Zhou No summary available — see the abstract on arXiv.

Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory

Siyuan Chen, Runlin Hou, Shenxiu Wu, Yansong Sun, Junming Cao, Yiyu Zhang et al. No summary available — see the abstract on arXiv.

Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Thibault Ba\~neras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu et al. No summary available — see the abstract on arXiv.

Cross-Dataset Stability of Expert-Informed Skill Prompting and Fine-Tuning for Chinese Metaphor Identification

Yufeng Wu, Meichun Liu No summary available — see the abstract on arXiv.

GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning

Lam So, Canhui Wu, Han Lin No summary available — see the abstract on arXiv.

Individual Fairness in Hierarchical Clustering

Binita Maity, Shrutimoy Das No summary available — see the abstract on arXiv.

A Hierarchical Synergistic Deep Learning Framework Integrating Composition, Structure, and Ionic Transport for Solid-State Electrolyte Discovery

Hongwei Du, Dingyang Lv, Baole Wei, Yongheng Li, Feng Yu, Ziheng Lu et al. cross-listed No summary available — see the abstract on arXiv.

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie et al. No summary available — see the abstract on arXiv.

M-Fibration Theory with Applications to Neural Network Compression

Paolo Boldi No summary available — see the abstract on arXiv.

Frequency-aware forecasting for short-term typhoon gust prediction

Xuefei Wang, Tingyi Liu, Heng Zhang, Shengjun Zhang No summary available — see the abstract on arXiv.

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

Lukas Edman, Daryna Dementieva, Alexander Fraser No summary available — see the abstract on arXiv.

AWM: Answerable Working Memory for Long-Document VQA Agents

Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu, Zhen Yang, Rui Lu, Yuxiao Dong et al. No summary available — see the abstract on arXiv.

Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing

Haoyu Wang, Cheng Feng, Liuyang Bian, Ruiyang Huang, Lei Wei, Yafei Wen et al. cross-listed No summary available — see the abstract on arXiv.

DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search

Junzhao Zhang, Tao Zhang, Liren Yu, Feiyi Dong, Zhixuan Zhang, Dan Ou et al. No summary available — see the abstract on arXiv.

AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification

Zebei Zhao, Zhihao Shi, Minqi Shi No summary available — see the abstract on arXiv.

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

Bing Shao, Jiazheng Zhang, Long Ma, Yujiong Shen, Senjie Jin, Xin Guo et al. No summary available — see the abstract on arXiv.

LDAC-Net: A Learnable Multi-Lag Differencing Attention-Convolution Network for Drift-Robust Recognition with Low-Cost MOX Gas Sensors

Xin Zhang, Liangxiu Han, Yue Shi, Tam Sobeih No summary available — see the abstract on arXiv.

Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

Zhexi Feng, Wuxi Chen, Bingrui Zhang No summary available — see the abstract on arXiv.

Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context

Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie No summary available — see the abstract on arXiv.

Narcissus: Program Synthesis Using Context-Aware LLM Approximations

Tilman Hinnerichs, Sebastijan Dumancic, Neil Yorke-Smith cross-listed No summary available — see the abstract on arXiv.

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

Tim Schopf, Tobias Schreieder, Akiko Aizawa No summary available — see the abstract on arXiv.

Agents 60

REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring

Muhammad Waseem, Aakash Ahmad, Pekka Abrahamsson cross-listed Large language models (LLMs) can propose refactorings, but the edits must reduce targeted quality problems without introducing new issues or altering behavior-relevant code structure. REFINE (Refactoring with Evidence-aware Flow for Integrated ageNtic Execution) is a tool-agnostic multi-agent pipeline for Java file-level refactoring that chains static-analysis smell identification, smell-informed planning, LLM transformation, automated re-analysis, preservation checks, and structured reporting. Across 450 files from 15 open-source systems and 1,350 outputs from GPT-5.5, Gemini 3.1 Pro Preview, and Claude Opus 4.8, detected code smells drop by 68 to 73 percent, with the largest reductions on major smells, and a matched 150-file direct-prompt baseline shows the pipeline achieves higher median smell reduction with smaller edits and fewer public-method removals. Broader quality gains are inconsistent and preservation checks still surface risks such as assert/fail-call changes, so the authors frame outputs as candidates that need compilation, testing, dependency analysis, and human review.

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

Parker Fawcett cross-listed Prior work found that once models are strong enough, a multi-agent rebuild pipeline loses to simply handing the model the original code and one instruction, as in AgentModernize. rebuild-dossier is an open-source, MIT-licensed tool that locks an application's exact interface, its inputs and outputs, before any code is written, then enforces one-test-at-a-time building through automated checks rather than written instructions alone. In a small comparison, the process-compliant agent failed a held-back test while the rule-breaking agent passed everything, showing that a passing suite does not certify correctness when tests can be gamed; against the source-plus-one-instruction baseline the tool tied on a small app but lost on a larger one where the automated check was not running, pointing to the check mechanism rather than interface-locking as the active ingredient. Every claim is verified at three levels, the agent's own report, an automated log, and the produced files, which caught errors including a bug in the authors' own logging code, and a stronger model followed the process three times in a row where a weaker model never managed it.

LLM Agents Perform Controlled Experiments Using Simulation Models

Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St\"umpfle, Johannes Sigel, Akshay Narla et al. Many scientific and engineering tasks require understanding how a system responds to intervention, which plausible text and code generation alone cannot supply. The authors propose a multi-agent framework in which large language model (LLM) agents run controlled experiments against high-fidelity simulation models for pharmaceutical process design: given a user query and a baseline configuration, the system builds a structured task representation, designs experiments, executes comparative simulations, interprets the outcomes, and synthesizes evidence-based recommendations for process parameter optimization. In an industrial application setting, simulation-integrated agents produce more specific outputs with higher user-rated correctness and helpfulness than language-only reasoning, and ablations and visualized case analyses support the value of reasoning through intervention, comparison, and observation.

When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs

Jason Liu cross-listed Tool-using agents must decide when a task is done, yet existing systems that gate terminal success or certify traces have not been tested at the completion boundary across controlled termination faults. Evidence-Carrying Termination (ECT) lets an agent return COMPLETE only when a typed certificate binds every required answer claim to valid, in-scope trace evidence and a deterministic replay reconstructs the claimed value. In a locked static study of 48 synthetic tasks across six tool-use families with eight injected faults, ECT produced 0 of 288 unsafe completions versus 252 of 288 for a termination-critic baseline; a fresh, prespecified 576-trajectory study on 22 held-out task clusters found 0 of 66 premature unsupported terminations versus 40 of 66 for the critic's controller, while supported completion (97 versus 92 of 132) met a 10-point noninferiority margin and 17 of 18 recovery attempts went on to complete with support. The authors stress that ECT certifies evidential support within a recorded trace under declared assumptions, not external truth, safety, or alignment.

ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents

YiShan Zheng, Yuan Wu, Yi Chang cross-listed Clean end-to-end success rates cannot reveal where a tool-calling failure originates or how it propagates through a call. ToolRobustBench is a stage-wise diagnostic benchmark that aligns four perturbation families, tool interface, user intent, tool output or observation, and runtime environment, with the tool-use pipeline and attributes failures to tool selection, schema grounding, argument binding, feedback handling, or end-to-end task success. Experiments on 15,456 single-family instances across 7 models, 16 sampled local tools, and 14 perturbation subtypes show high but uneven clean performance and substantial degradation under perturbation, with tool-output and observation perturbations the dominant bottleneck; mixed-family experiments reveal non-additive failure patterns that isolated single-family results do not explain.

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

Esmail Gumaan cross-listed Agent harnesses typically append a failed tool call and its error message to the transcript and ask the model to continue, on the assumption that the error is corrective information. Measuring the change in log-probability of re-emitting the failed action across six instruction-tuned checkpoints (135M to 1.7B parameters, four families) in simulated tool calling and MBPP program repair, the authors find the effect is negative in every case: the probability of repeating the failed call rises from 0.06 to 0.54, roughly -1.03 nats per action token, and holds on 90-100% of individual items. Counterfactuals show the verbatim surface form of the failed call accounts for 83% of the damage, while the semantic content of the failure marker contributes little. Replacing the verbatim call with a runtime-generated description of the failure removes 76% of the effect, whereas an explicit "do not repeat" instruction does nothing and retrying from a clean context is the worst option because it recreates the context that produced the failure.

Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

Zizhe Wang cross-listed Physical system modeling in Modelica, an equation-based language, differs from ordinary code generation because a model can compile and simulate while still violating its intended physics, and an agent iterating across revisions may lose track of requirements or trust simulation evidence from a stale candidate. Pufibara is an agent harness that keeps persistent engineering state across revisions, ties execution and simulation evidence to the candidate that produced it, and makes submission an explicit action. To evaluate it, the authors build the 232-task Modelica Agent Workflow Benchmark covering model repair, generation, and tuning, with each submission scored by an evaluator outside the agent loop. With matched backends (DeepSeek v4 Flash and Claude Sonnet 5), Pufibara passes 202 tasks versus 185-187 for Claude Code, while using 76-83% fewer logical tokens and 6-58% less sequential runtime.

Automata from Agent Traces: Failure and Next-Step Prediction

Seonglae Cho, Franklin Cardenoso Fernandez, Umar Mohammed, Zekun Wu, Kleyton Da Costa, Ilham Wicaksono et al. Long unstructured traces from LLM agents resist the safety auditing and runtime monitoring that deployment requires, and existing approaches work per-trace or only on successful runs, missing the cross-run structure that links next-step and failure prediction. The method collapses an entire trace corpus into a single compact finite-state machine (FSM) that serves as a structural substrate for agent behavior. Across twelve public datasets the FSMs have 7-43 states, replay held-out data at 0.997 or better fitness with near-identical topology across splits, and build in milliseconds; FSM-state context beats Agent Workflow Memory on next-step prediction for every ground-truth-matched dataset, and per-state behavioral features reach held-out failure-prediction AUROC of up to 0.94, enabling an online monitor that flags failing runs from partial traces for early stopping. The behavioral topology appears to be shaped more by the deployment harness than by the underlying LLM.

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

Stephen Chung, Wenyu Du, William J. Wesley The Station is an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal with no central coordinator or scripted pipeline, choosing their own directions, running experiments, collaborating, and building a shared literature. Applied to 12 construction problems from the AlphaEvolve catalogue plus two case studies, the agents produced results novel relative to prior literature on five problems, including a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and an improved lower bound for Erdős's minimum-overlap problem, along with novel infinite families for Book Ramsey numbers. Beyond numerical constructions, the agents produced theorems and analyses explaining why the constructions work, and all raw dialogues, proofs, and verification code are released.

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

Seonglae Cho, Donghyun Lee Concurrent multi-agent coding could divide labor across modules and explore in parallel at the granularity of multi-file projects, but existing systems either sequence agents through phase handoffs or pool independent samples without coordination, and a single agent abandons up to half of hard tasks with a one-file stub-and-exit. AgentRoom is a realtime collaborative editing protocol for coding agents whose runtime exposes file-level claim, status, and broadcast as Model Context Protocol (MCP) tools over a shared filesystem merged with Conflict-free Replicated Data Types (CRDTs). Five frontier coding-CLI models ran four backend tasks with cross-language checks in Python DevBench and Rust with axum; for CLI-stable models, AgentRoom with two agents abandons fewer tasks than a solo agent and shows less run-to-run variation. Matched-compute LLM-judge contrasts favor AgentRoom over parallel-merge and full AgentRoom over partial ablations, indicating that coordination, not parallelism or CRDT merging alone, carries the benefit.

Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search

Jiongxiao Wang, Dingli Ma, Chaoqun Ni Biomedical fact-checking systems built on retrieval-augmented generation (RAG) and agentic search typically emit bare supported or refuted labels without the evidence synthesis and justification that make them useful to human readers. BioCheck Agent searches PubMed exclusively using Boolean query operators and produces structured reports that combine retrieved evidence, analysis, and a conclusion; because direct prompting of lightweight open models yields hallucinated, low-quality reports, the authors train it with Evidence-Grounded Group Relative Policy Optimization (EG-GRPO), a reinforcement learning method whose task-specific reward encourages sophisticated search and high-quality evidence while penalizing hallucination. Relative to the Qwen3.5-4B base model, the trained agent improves label accuracy on SciFact by 9.95%, raises the evidence quality score by 3.7%, and lowers the evidence hallucination rate by 19.63%.

Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization

Yapeng Liu, Yuanzhao Zhai, Xudong Gong, Dawei Feng, Bo Ding, Lin Wang et al. cross-listed Evaluations of embodied agent systems (EAS) rely on outcome metrics such as success rate or safety scores that collapse entire execution trajectories into a single number, hiding how agents recover from and adapt to the continual disruptions of open-world environments. Drawing on resilience engineering, the authors define a metrics suite covering Rebound, Stability, and Graceful Extensibility during task execution, and implement an evaluation layer that turns execution traces into diagnostics usable for optimization. Across 400 household tasks and 10 embodied agent systems, the metrics expose differences hidden by outcome scores, such as a recovery-cost gap of 25.2 between episodes that all count as successes, and metrics-guided optimization reduces recovery cost and increases stability and graceful-extensibility completion, though a trade-off among the resilience properties suggests configuring agents to deployment-specific requirements.

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

Haoyang Fang, Bernie Wang When an LLM agent must refine candidates under a small evaluation budget, because validation or generation is expensive, standard Monte Carlo tree search (MCTS) spends the budget poorly: exploration bonuses dominate at low visit counts, unpromising siblings get expanded before promising chains can deepen, and branching ignores node quality. ExTS treats expansion itself as a value-of-information decision, combining discriminative reward shaping to separate candidates with narrow score distributions, a stochastic virtual child that estimates the value of opening a new branch from the parent's reward history, and quality-conditioned branching that expands only when a node's score justifies the cost. Across prompt optimization, code generation, molecular structure elucidation, and agentic workflow optimization, a single fixed configuration of ExTS matches or beats task-specific tree-search baselines with an average relative gain of 5.5%, and pilot-run diagnostics characterize how these budget-constrained search problems differ structurally from one another.

Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)

Avital Aviv, Parth A. Gandh, Ron Bitton, Asaf Shabtai cross-listed Google's Agent Payments Protocol (AP2) lets LLM-driven shopping agents authorize payments through signed Checkout and Payment Mandates, but those signatures protect transaction data only after signing, leaving the Agent-to-Agent Protocol (A2A) messages and Model Context Protocol (MCP) tool calls that shape a transaction beforehand unprotected. The authors analyze AP2 v0.2 by splitting its lifecycle into five phases and identifying five deployment architectures, then apply the MAESTRO threat-modeling framework (Multi-Agent Environment, Security, Threat, Risk, Outcome) to derive four threat actors, eleven attack surfaces, eighteen adversary capabilities, and a catalog of 48 threats in five attack families, scored with the Artificial Intelligence Vulnerability Scoring System (AIVSS). Eight threats reach the High band in at least one architecture, and since no complete public AP2 deployment exists, the authors build a testbed spanning all five architectures with proof-of-concept demonstrations and mitigations for each, plus a deployment-aware scanner, concluding that valid mandate signatures alone do not guarantee a transaction reflects user intent when its pre-authorization context is manipulated.

Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment

Aryan Brar, Justin Du, Avery Lor, Kylie Seto, Eric Taylor Tax-loss harvesting benefits long-term portfolio growth but depends on holdings- and owner-specific considerations, so the authors build a multi-agent trade recommendation system fed by two context providers: a custom capital gains calculation engine and a retrieval-augmented generation (RAG) vector store of market advisory reports. In a 2x2 repeated-measures ANOVA measuring relative capital gains during portfolio liquidation, enabling the tax engine significantly reduced tax savings by roughly 55 percentage points (F(1,29) = 9.17, p = .005), while the RAG main effect and the interaction were not significant. The RAG-only condition had the highest mean tax savings (47.7%) and the unaugmented baseline was second (30.6%), suggesting the model's internalized financial knowledge may suffice and that bolting on domain-specific computation engines can introduce conflicting optimization signals rather than help.

PROOF-Gen: From Optimized Data to Better Distillation

Anh Ta, Junjie Zhu, Shahin Shayandeh Distilling tool-calling ability into small models usually starts with supervised fine-tuning on teacher trajectories filtered by generate-and-keep-passing, which discards failures and leaves the same hard scenarios unresolved each retraining cycle; on τ2-bench, 57% of teacher trials fail, and two-thirds of those are near-misses undone by a single decisive error. PROOF-Gen (Per-scenario Reflective Optimization to Overcome Failed Generation) has a reflector analyze each failed trace and its evaluation feedback, write corrective guidance that steers the teacher to a passing trajectory, and then strips that guidance before training so the student learns from clean demonstrations. Per-scenario optimization recovers 93% of failed scenarios on τ2-bench, fine-tuning on the combined data lifts Qwen3-4B-Instruct-2507 from Pass^1 of 0.132 to 0.529 and gives Gemma 4 E4B-it +7.2 points on BFCL v4 multi-turn, and in a deployed pipeline the method improves goal completion by 6.3 points with positive transfer to an on-device model in every locale.

MARS: Multi-Specialist LLM Relay System for Competitive Programming

Andrei Mikhailov, Mikhail Burtsev, Alsu Sagirova Multi-agent pipelines for competitive programming typically split work across generic planner, coder, and debugger roles and leave the choice of algorithmic technique to the backbone model. MARS (Multi-Agent Relay of Specialized LLMs) is a prompt-only framework whose agents are topic specialists (dynamic programming, graphs, strings, geometry, and so on) grounded by retrieval-augmented generation over an algorithm-theory corpus: retrieval picks a small team, a starter writes an initial C++17 solution, each turn runs the candidate against public examples in a sandbox and lets the active specialist keep, repair, or hand off the draft with a structured packet, and a final infrastructure-fixer pass normalizes boilerplate. On the CodeContests test split with Gemma 4, MARS reaches a 0.624 pass rate, 14.4 percentage points above direct prompting, at 2.3 pipeline stages per task, closing most of the gap to CodeSIM (0.731) at 3.3x lower wall-clock cost and with much smaller variance in per-task token spend.

The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses

Dai Jiahong cross-listed An agent harness, the code that builds a model's context, mediates its tools, runs the loop, and persists state across a long run, is increasingly the binding constraint on agent behavior rather than the model it wraps. The authors conduct a source-level, multi-case study of three open coding-agent harnesses built on opposing philosophies, LangChain's deepagents (batteries-included), Earendil's pi (radical minimalism), and DeepSeek's dsh (everything-is-a-plugin), reading each at a pinned commit and following its commit history. Despite travelling in opposite directions, the two mature harnesses converge on the same five elements: a commoditized loop, an append-only replayable session record, model quirks kept as data, progressive disclosure of context, and explicit extension seams, all of which the held-out third harness also exhibits (in one seam reusing another's implementation outright), so the convergence is decomposed into parallel discovery, diffusion, and literal reuse. One load-bearing dimension, external verifiability via a tamper-evident record an outside party can check without trusting the runtime, is absent from all three, which the authors read as the next axis on which harnesses for provenance-sensitive domains will differ.

Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation

Olympia Saha, Amy Wang, Srinivasan Manoharan cross-listed A proxy Model Context Protocol (MCP) server that aggregates hundreds of backend tool servers behind one governed endpoint creates two compounding problems: full tool schemas saturate the model's context window before any user query arrives, and agents cannot identify the best tool among thousands. SCOUT (Selective Context Optimization for Universal Tooling) reframes tool exposure as a context-selection problem, surfacing just two meta-tools — tool_search, which performs hybrid retrieval by fusing BM25 sparse matching with dense vector search via Reciprocal Rank Fusion, and execute_tool — backed by zero-downtime catalog update pipelines over 2,000+ tools across 200+ MCP servers. In production at PayPal, SCOUT cuts MCP tool-token consumption from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%), a 99% reduction, and because it is exposed as standard MCP tools it is model-agnostic and needs no client-side changes.

Reflection with Action-Induced Visual Differences for Desktop GUI Agents

Yijie Ma, Chaoyue Niu, Fan Wu, Guihai Chen The Planner-Operator-Reflector (POR) framework is widely used in GUI agents, but on large, dense desktop interfaces the reflector carries most of the burden, since it must compare pre- and post-action screens where state changes are subtle or scattered, and existing reflectors collapse change detection and outcome verification into one weakly grounded step. Evidence-First Reflection (EFR) is a two-stage reflector that first identifies the action location and candidate changed regions with Set-of-Marks annotations, describes and filters the action-relevant changes, and only then judges the outcome from the cleaned evidence, reducing both visual search complexity and reasoning burden. On OSWorld-Verified and WindowsAgentArena, EFR raises reflector accuracy by 7.11% and average end-to-end task success by 5.94% and 4.95% on the two benchmarks respectively.

WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents

Lin-Fa Lee, YI-YU Chang, Kuo-Hui Yeh cross-listed The W3C WebMCP proposal lets LLM agents invoke tools exposed by web pages, but a browser security model built around the Same-Origin Policy (SOP) offers no provenance or lifecycle guarantees for such tools, opening three risks: subject-attribution spoofing, uncontrolled tool lifecycles, and semantic prompt injection. WebMCP-Phalanx is a dual-layer runtime in which a browser-native trust anchor binds each tool to its registering principal through cryptographically protected capability credentials and propagates provenance labels through the tool lifecycle, while a Quarantine Agent (Q-LLM) with no tool authority inspects tool metadata, outputs, and page content for injection before forwarding validated content to a Privileged Agent (P-LLM) that executes. The ownership mechanism drops revocation and overwrite attack success from 100% to 0%, and the dual-agent runtime blocks all 80 injection attempts embedded in tool descriptions and limits tool-return attacks to 2 of 80, with task utility statistically indistinguishable from the no-attack baseline; a white-box adaptive attacker can still bypass description filtering through malicious tool names invoked before inspection, which motivates a call-timing gate that delays invocation until all agent-visible metadata has been validated.

What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions

Yichao Gao, Yumo Zhang, Yunhao Yao, Haohua Du, Puhan Luo, Ruiqi Li et al. cross-listed LLM agents that receive external data through the same natural-language channel as their instructions are vulnerable to injection attacks in which untrusted content gets parsed as behavior-guiding instructions during inference, and static input/output filtering misses inducements that only surface during reasoning. Attnlocate is a runtime framework that localizes the context spans genuinely driving tool-calling decisions by casting the problem as object detection over the attention matrix: a multi-head, multi-layer attention aggregation builds a token-level feature space, a 1-D U-Net with an anchor-free detection head finds the spans, and the system then adjudicates the invocation attempt based on the authority of the provider the detected span came from. Across ten agent configurations from five LLM families, covering indirect prompt injection and tool poisoning, the detector reaches a mean IoU of 0.743, an average AUROC of 0.956, and a 0.934 true-positive rate at a 0.067 false-positive rate, transfers to unseen models, and supports new authority policies without retraining.

Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding

Hyunho Kook, Junhyuk So, Tianyu Fu, Haizhong Zheng, Beidi Chen Confidence-based voting weights parallel LLM rollouts by internal signals such as token log probabilities, but in multi-turn search agents that condition on retrieved documents, tokens copied from those documents receive systematically inflated log probabilities — a failure the authors name copy inflation — which flattens confidence scores within a question and weakens the weighted vote. Retrieval-Grounded Voting (RGV) instead scores each rollout by the lexical overlap between its final answer and the documents it retrieved, computing the signal outside the contaminated context with no token probabilities and no additional LLM calls. Across four search-agent benchmarks and five LLMs, RGV consistently outperforms confidence-based voting, with gains of up to +5.4% accuracy and +35% on minority-correct questions where the right answer appears in only 1-2 of 8 rollouts.

Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings

Muhammad Tayyab Khan, Lequn Chen, Wenhe Feng, Seung Ki Moon cross-listed Manufacturing process planning must turn 3D computer-aided design (CAD) models, 2D engineering drawings, material specifications, and domain rules into a coherent sequence of manufacturing decisions, but existing methods handle only isolated subtasks such as feature recognition, drawing interpretation, or tool selection. Design-to-Plan is a large language model (LLM)-based multi-agent framework in which an orchestrator coordinates specialized agents for 3D feature recognition, 2D drawing analysis, 2D-3D context fusion, knowledge retrieval, process sequencing, tool selection, and report generation, with LLM agents reasoning over structured outputs from deterministic modules and knowledge sources rather than generating free text. On 300 benchmark cases, the parallel architecture reaches 100% task success across three downstream ReAct agents, tool-selection F1 of 95.9%-97.6%, 90% source-detection accuracy in conflict analysis, and a 60%-68% reduction in token usage for key planning tasks.

AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval

Gunja Agarwal, Arup Kumar Das, Arun Menon, Jitesh Chandra Mishra, Vignesh Divakaran Evaluations of agentic information retrieval typically script interactions with uniform simulated users, missing both natural personality diversity and adversarial brittleness. AgentWorld combines Big Five (OCEAN) personality-driven user populations with stateful tool-use environments; a pass^k consistency metric with structured fault classification, partial-credit scoring, and dual-control handoff verification; score-thresholded export of training data in six fine-tuning formats; and an adversarial Risk Analyser that snapshots required intermediate states, branches Monte-Carlo rollouts under four task-aware perturbation types, and quantifies risk through delta-P/delta-T scoring, Dempster-Shafer evidence fusion, and Shapley attack-category attribution. Across a conversational analytics agent with 10 personas, a customer-support agent over 5 tasks and 4 persona variants, and adversarial stress tests, personality variation exposes cross-domain leakage, contextual drift, and a 50% versus 100% pass-rate spread across personas on the same task, while the Risk Analyser reveals pre-existing trajectory brittleness (minimum 0.375 without perturbation) dominated by tool- and infrastructure-layer attacks (Shapley: 46% system, 38% action).

EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals

Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang et al. Large language models (LLMs) are increasingly deployed as code agents for scientific and engineering analysis, but their ability to work with raw physical-layer measurements has gone untested. EMRB (Electromagnetic Reasoning Benchmark) contains 200 problems across five difficulty levels and 27 question types, from signal detection to OFDM design, generated from 11 signal types with verified ground truth; the model receives only a raw I/Q capture and must discover the relevant quantities by writing and running code. Across 14 proprietary, open-weight, and reasoning-oriented models, scores range from 24.1% to 78.9%, with mean accuracy dropping from 84.9% on basic measurement to 21.2% on system design; a structured method called ReconPilot that separates signal reconnaissance, targeted analysis, and self-verification raises overall scores by 3.8 to 17.6 points across three backbones and improves 13 of 15 backbone-level combinations.

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

Nadeem Shaikh cross-listed Existing LLM agent systems decide on delegation either before reasoning starts, when a router picks a model, or after a response is complete, when a verifier scores it and may retry; the authors study a third regime in which an agent recognizes mid-generation that it is unlikely to succeed and hands control to a stronger model. They cast intra-generation delegation as a Bayesian optimal-stopping problem over a learned competence posterior, an online estimate of eventual task success whose sufficient statistics are learned from labelled trajectories rather than read off raw entropy, derive the myopic escalation threshold in closed form, prove the optimal policy is a time-varying threshold with no shape assumption on the raw signal, and establish a regret bound governed by posterior calibration plus a finite-sample guarantee that the plug-in policy's regret decays as 1/sqrt(n) with n calibration trajectories. A controlled simulation confirms each prediction, and on a Qwen2.5-Coder 1.5B-to-7B code cascade over 257 MBPP tasks, two of three pre-registered predictions hold: the escalation frontier dominates post-hoc routing at equal cost, and the competence belief's discrimination rises over the course of generation.

Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

Guo Gan, Yilun Zhao, Cong Chen, Jinbiao Wei, Tingyu Song, Zheyuan Yang et al. GUI agents deployed on Android devices encounter runtime anomalies such as unexpected pop-ups and misused actions, yet existing benchmarks do not systematically test robustness against them. AnTrap injects dynamic perturbations into agent execution trajectories following a taxonomy of four layers (State, Thinking, Action, and Round) with ten fine-grained subcategories, using a construction pipeline that preserves task solvability while introducing realistic adversarial conditions. Evaluating 16 leading GUI models reveals universal vulnerability, with even the strongest models degrading significantly; GRPO training in original versus adversarial environments shows that single-step traps at the state and action layers are largely fixable by adversarial reinforcement learning, whereas deep contextual traps such as state deadlock persist and expose reasoning limitations that training in trap environments alone cannot resolve.

ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation

JooYoung Jang, Taegyeong Lee, Jihyeon Park, Nojun Kwak LLM agents editing documents on commercial design platforms face two blockers: legacy formats expose only flat, absolutely positioned elements, so agents must recompute coordinates and routinely break layouts, and design has no unique ground truth, so diff-against-reference metrics penalize valid-but-different outputs. ACE is an agentic canvas editor over a hierarchical scene graph with a 98-tool presentation-specific action space, paired with CARE, a content-aware router that feeds the agent only the relevant slice of each deck (about 89% input-token reduction), and a self-correction loop in which a ground-truth-free instruction-following judge's natural-language critique becomes the next-turn instruction. With a fixed backbone, single-turn scene-graph editing already matches a same-backbone agentic HTML pipeline that iterates internally, and adding self-correction lifts instruction following to 4.23 versus 3.81 on the full 94-task benchmark (paired p=0.010) at 1.75x the speed and about 44% lower cost; 26 blind raters prefer the self-corrected output 81% of the time, 66% of cases halt after one pass, and a strict-peak rollback removes every observed regression.

Task-Adaptive Rubrics for GUI Reward Modeling

Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao et al. Outcome reward verifiers for graphical user interface (GUI) agents judge whether an executed trajectory satisfies a user instruction, but their criteria are not sufficiently task-adaptive: generic rubrics or implicit reasoning can transfer checks across tasks, miss concrete constraints in the current instruction, or enforce unstated requirements. AdaptRubric builds judging criteria in two stages, first routing the instruction to a GUI task family and retrieving reusable category-level criteria, then generating instance-level cues for the concrete values, scopes, and constraints in the specific instruction. In offline reward evaluation and online reinforcement learning, it outperforms prior reward agents, improving F1 by 3.6 points over the baseline average under a matched image budget and yielding a 4.23-point task-success gain.

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

Jiayu Shi, Luzhuo Chen Coding agents re-send large file reads and tool outputs to a frontier model every turn, and general-purpose prompt compressors trained on prose paraphrase identifiers and drop the exact spans an agent needs to edit. Paritok-4B is a 4B LoRA compressor fine-tuned from Qwen3-4B on 40,606 validated examples distilled from a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories; it is extractive, selecting rather than rewriting spans (96.0% of emitted identifiers, paths, and numbers appear in the input), and intent-conditioned, using the agent's current task to choose which lines survive within a retained segment. On all 300 SWE-bench Lite instances it compresses agent context to 25.7% of its size while retaining 86.5% of uncompressed single-shot solve quality, versus 50.2% for a gpt-4.1-mini compressor and 61.9% for gpt-5; on line-numbered input it compresses to 27.8% and retains 89.3%, with a paired McNemar test (p=0.079) showing no significant drop in solve rate at this sample size. The 264 MB adapter self-hosts on a single 24 GB GPU with no per-token fee, the authors note that at list prices gpt-5 as a compressor costs more than the downstream tokens it saves, and weights, data, and evaluation scripts are released under Apache 2.0.

MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG

Qiuyi Qi, Tian Liang, Jiamu Wang, Jinjian Zhang, Wei Zhou, Pengcheng Zhu et al. Agentic retrieval-augmented generation (RAG) requires a model to decide when to keep searching and when to answer, yet existing reinforcement learning methods rely on external supervision and ignore the agent's own belief about whether its evidence is sufficient. MetaRAG reframes search-decision quality as belief-action alignment: Verify-first Action Generation elicits an explicit verification step before each action, Internal Belief Probing estimates the policy's own answerability belief from the same question-history context, and a consistency reward between the two is gated by answer correctness so internally consistent but wrong trajectories are not reinforced. The probe is used only in training with no inference-time overhead, and on seven public QA benchmarks the method consistently improves the accuracy-efficiency trade-off over strong RL-based agentic RAG baselines, with gains transferring to deep research settings, different optimizers, and multiple backbones.

SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction

Xue Hu, Zewei Pan, Zeli Su, Zhou Liu, Wentao Zhang LLM agents that turn research papers into reproduction code often produce implementations that run but silently diverge from the paper's specification, a failure the authors call semantic drift. SA-Bench decomposes 30 papers from ICLR, ICML, and NeurIPS 2025 into 1,491 atomic, verifiable implementation claims, termed Semantic Alignment Units (SAUs), and scores generated repositories along four drift dimensions: numerical, methodological, protocol, and ordering. Across 12 generator configurations (four models times three scaffolds), the strongest setup, Claude with PaperCoder, reaches a mean SAU score of only 0.301 out of 1.0, with an overall mean of 0.221; a failure taxonomy shows agents attempt most requirements but implement them incorrectly or as stubs, suggesting scaffolds tuned for executability do little for scientific fidelity.

ReproAgent: Contract-Guided Paper-to-Code Reproduction

Xue Hu, Zewei Pan, Zhongyuan Wang, Zhou Liu, Zeli Su, Wentao Zhang Turning a research paper into a runnable repository is hard because the specification is split: explicit details like algorithms and metrics get lost over long agent trajectories, while implicit details such as framework defaults and conventions inherited from related work never appear in the paper at all. ReproAgent is a four-stage Prepare-Plan-Generate-Repair pipeline organized around a persistent implementation contract with two channels, one converting paper snippets into code obligations and one retrieving content and structural evidence from related repositories, both bound to work packages and projected into file-level contracts that guide generation and repair. On PaperBench Code-Dev, ReproAgent reaches the highest mean score among scaffolds using the same backbone under both Claude-Sonnet-4.5 and Gemini-3-Flash, and end-to-end channel ablations support the contribution of each channel.

Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

Anupam Purwar, Shashank Singh, Kritika Srivastava Scoring conversational voice agents at scale needs automated judges that capture both observable interaction quality and the contextual judgment humans supply, so the authors compare human ratings against GPT-4.1 and GPT-5 as judges on telecom and retail voice-agent conversations across quality and safety dimensions, under three evaluation configurations (p0, p1, p2) to test sensitivity to setup. Beyond aggregate agreement they examine per-metric correlations, evaluator consistency, and systematic human-LLM disagreement, noting that pipeline factors such as speech generation, streaming, and error propagation across ASR, reasoning, and tool-calling stages shape what a judge sees. LLM judges prove usable as a component of large-scale voice-agent evaluation, but their reliability varies by metric and configuration rather than being uniform, motivating hybrid pipelines where automation handles scalable metrics and humans retain those needing contextual interpretation.

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman Long-running coding-agent tasks tempt users to switch models mid-run, escalating from a low-cost, low-capability (LC) model to a high-cost, high-capability (HC) one when the cheap model struggles, or downshifting once the hard reasoning is done; each switch forces the receiving model to continue a trajectory it did not produce. The study pairs LC and HC models from the Claude and GPT families and varies handoff direction, timing, and interface, comparing full-trajectory transfer, compaction, and dropping the trajectory entirely while preserving the repository state. Full-trajectory escalation recovers less than half of the LC-to-HC quality gap while adding a substantial cost premium, whereas downshifting lands at a favorable cost-quality point. The preferred interface flips with direction: stripping the LC model's trajectory improves escalation quality, but removing the HC model's trajectory hurts downshift quality.

Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems

Yarden Bakish, Amir Dudai, Roy Ganz, Oren Nuriel, Elad Ben Avraham, Mor Shpigel Nacson et al. Diagnosing why a multi-agent LLM run failed still mostly falls to human engineers, who in practice rely on observability tools that organize traces around components, actions, and dependencies rather than reading raw logs end to end. Adaptive Influence Graphs (AIGs) apply the same paradigm to LLM diagnosers through a two-stage agentic framework: a failed trace is first transformed into a structured graph, and an agent then navigates that graph to locate the critical error. Across multiple models, richer trace representations consistently improve attribution, with adaptive graph construction plus agent-directed traversal performing best, and AIGs set a new state of the art on Who&When, the standard multi-agent failure-attribution benchmark. The results indicate attribution quality depends on how a trace is represented and explored, not only on the diagnosing model.

From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use

Rongfeng Guo, Yinxuan Huang, Yusen Wu, Maoqing Zhong, Yunlu Chen, Meng Tang et al. Multi-turn tool use requires an agent to maintain an evolving task state while keeping each action consistent with it, but direct function-calling and ReAct-style policies learn state tracking and action generation in the same autoregressive trajectory, so the pressure to emit the next call can overwrite information gathered earlier. OODA-Tool, inspired by Boyd's Observe-Orient-Decide-Act cycle, is a typed closed-loop policy that routes each decision through controller-checked intermediate stages: Observe reconstructs the task state, Orient decides whether execution is warranted, Decide forms an admissible action structure, and Act produces the external output. Evaluated against direct function-calling and ReAct policies on Qwen3 models from 0.6B to 14B parameters across multi-turn, multi-tool, and incomplete-information settings, it consistently improves task success at every model size, with the largest gains on smaller models and on tasks whose actions depend heavily on earlier turns and prior tool results. Controlled variants, stage-level ablations, and transfer evaluations support the robustness of the gains.

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye Agents that invoke multiple tools need parallel execution for low latency, but existing benchmarks evaluate tool selection and argument generation under mostly serial execution, ignoring valid parallelization and resource-constrained scheduling; resource-agnostic parallelism is fast but prone to avoidable resource overflows. PeakBench is a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles, paired with a two-part evaluation framework that separates logical dependency planning from physical scheduling, with dedicated metrics for each, so failures can be attributed to the right cause. Strong logical planning does not reliably translate into safe or efficient execution under resource constraints, while exposing resource information to the agent reduces avoidable overflows and improves utilization. Code is released.

Discovering Adaptive Transmission Programs for Collective Innovation

C\'edric Colas, J\'er\'emy Perez, Eleni Nisioti, Akhilesh Mocherla, Pierre-Yves Oudeyer, Cl\'ement Moulin-Frier et al. Human collective intelligence depends on transmission processes governing who shares what with whom, how, and when, but prior work mostly models these through fixed network structures that cannot condition on what agents actually know or on the state of the group. The authors formalize transmission protocols as state-aware programs that route information and resources based on individual and collective state, and use LLM-guided evolutionary search to discover effective protocols for a collective discovery task. Evolved protocols outperform standard literature baselines by up to 37%, ablations show that removing content-dependence while keeping topology and timing erases the gain, and the discovered protocols transfer across domain variations and agent populations.

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

Yiheng Sun, Huifei Wang, Yancheng Zhu, Zhenyu Li, Zebin Zhao, Yifan Yuan Multi-role, multi-stage LLM agent workflows repeatedly rewrite upstream state into summaries, plans, tickets, memories, and handoff notes, and a downstream artifact can mention an unresolved condition while quietly demoting it from a hard prerequisite into optional context. The authors study this as operational state preservation, using safety blockers with explicit prerequisite, authority, fallback, and execution-consequence fields, conditioning on correct upstream identification and varying only the handoff transformation before evaluating an executor restricted to the resulting artifact. Across 1,296 controlled synthetic episodes, direct handoffs preserve every blocker, whereas ordinary handoff compression deactivates 100% of blockers and leads to a forbidden action 54.2% of the time; restoring all four state fields returns preservation to 100% and forbidden actions to 0%, and downstream verification eliminates forbidden actions even when the artifact itself remains 95.3% deactivated. Semantic availability of information does not guarantee that its action-binding force survives a handoff.

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

Lihang Zeng, Shaoting Zhang, Xiaofan Zhang Most LLM-based medical diagnosis systems treat diagnosis as static case-to-answer prediction rather than the active evidence-seeking process clinicians follow. EviDx pairs patient-specific interactive environments, built from raw clinical cases by an E-Synthesis step, with a clinical diagnostic scaffold of role-specialized agents, evidence tools, and evolving evidence states, plus an observer-guided runtime harness that decides when to stop by tracking uncertainty and evidence coverage. A three-level evaluation pyramid covering execution robustness, reasoning dynamics, and diagnostic outcomes shows that the framework improves both diagnostic performance and process stability, while also exposing capability boundaries that depend on the underlying model.

Joint Optimization of Tool Creation and Use for Large Language Model Agents

Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee Tool-augmented language models are limited to the APIs humans have written, and existing tool-creation systems prompt a frozen model at inference time, so the model that writes a tool never learns whether it can actually invoke the schemas it produces. SMITH (Schema-grounded Multi-task Iterative Tool Honing) is a reinforcement learning framework that trains tool creation and tool use in a single policy, where each rollout is either a build task (write a tool from a few examples) or a use task (call a pooled tool on a held-out question), with separate reward axes for schema, code, and outcome failures. A 4B Qwen3 model trained on 13 procedural reasoning tasks with exact verifiers reaches 79.8 macro-average accuracy on held-out tasks, ahead of an untrained 30B-A3B tool-writer, and scores 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA without any tabular or visual training data. Tools written by the 4B models also improved LFM-2.5-350M and Qwen3-30B-A3B on the same tasks.

A Literate Programming Environment for Human and Machine Agents

Adam T. Burke cross-listed Literate programming co-locates code with the prose that explains it, and the authors argue this structure is also well suited to large language model coding agents that work within limited context windows. The environment comprises a grammar for executable program essays, a parser that treats names as first-class objects, an internal name graph relating prose, names, and executable artifacts, and a binding mechanism for existing languages and testing toolsets, giving agents symbol-aware search and usage information comparable to what human developers get from an integrated development environment (IDE). A working implementation with bindings to three established programming languages is described alongside several example programs.

Confident at the moment of action: belief miscalibration in LLM play under hidden information

Bhushan Kashinath Joshi Agentic systems increasingly gate actions on a model's self-reported confidence, which presumes that confidence tracks correctness at the moment the action is taken. The authors test this in a hidden-information chess variant where royal status can be secretly and repeatedly relocated between pieces, eliciting the agent's stated probability distribution over the opponent's hidden royal piece every turn, separately from its move, and scoring it against ground truth recovered after the game. Across two independent batches, captures made at high stated confidence (at least 0.5) about the hidden piece's location were correct in only 1 of 62 cases, and roughly 99% of the calibration deficit was concentrated in these high-confidence events. The pattern recurs in weaker form across four further configurations spanning a second provider, a deliberation-budget change alone shifts the metric nearly as much as a large cross-model gap, and the configuration that wins on legality, cost, latency, and completion rate produces the worst belief quality, a failure that outcome-only evaluation would miss because such a model can still win the game.

Meta$^n$: Recursive Self-Improvement through Emergent Depth

Zae Myung Kim, Young-Jun Lee, Seungyeon Jwa, Dongyeop Kang Existing self-improving LLM agents refine their answers rather than the process that produces them, and systems that edit their own machinery must leave part of it untouched to stay stable, capping realized meta-depth at roughly two. Meta^n instead holds a single meta-operation Ω fixed and recurses on its input: Ω repeatedly reads the traces and code of the solver stack beneath it and writes the next layer as a strategic pre-process plus a library of callable helpers, so the system cannot destabilize itself while each layer reasons from a strictly larger vantage. Depth is determined by convergence rather than fixed in advance, and an evolutionary archive searches over chains of layers. Across two backbones, Meta^n outperforms prior self-improving agents on all eight benchmark families and is the only one to score above zero on ARC-AGI-2; ablations attribute most of the recursion gain to the conditioning each layer passes to the next, with distinct layer roles emerging with depth despite no prompt prescribing them.

Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav

Hongyu Guo, Zhiyu Zheng, Zhao Cao Large language model agents are moving from conventional retrieval-augmented generation toward Direct Corpus Interaction (DCI), where the full corpus stays accessible, yet under finite interaction budgets required evidence may never surface, may surface but remain unopened, or may be opened without exposing its decisive fragment. The authors name this progressive silent loss Evidence Blindness, quantify it through stage-wise evidence realization, and reframe large-scale agentic search as finite-budget navigation over reusable corpus structure rather than structure recovered online per query. AtlasNav organizes the corpus once into a persistent multi-view Corpus Atlas so that each query navigates adaptively instead of reconstructing shared structure. On BrowseComp-Plus it reaches 92.05% strict accuracy while cutting recorded online inference cost by 30.21% relative to the prior dynamic-workspace state of the art, realizes complete required evidence earlier under matched budgets, and remains effective on PhantomWiki across controlled 10K to 1M document scaling and on heterogeneous enterprise knowledge.

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou et al. Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect a trajectory before those errors compound. CAFE (Coupled Agent-Feedback Evolution) uses a single shared-parameter model that alternates between search-agent and critic roles: the agent decides when to request and apply feedback mid-trajectory, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves. Training initializes feedback-conditioned recovery from the base agent's own failures, then couples online RL, where a prompt-level call-versus-skip success gap shapes request returns and feedback-aware advantage shaping reweights token advantages before and after feedback, with offline preference optimization on matched successful and unsuccessful trajectories. On seven agentic search benchmarks CAFE outperforms the evaluated RL-based search agents on average, retains its gains on all six out-of-domain benchmarks, and reduces answer-level hallucinations, and one-sided ablations show that improving only the agent or only the critic plateaus while alternating the two updates keeps improving.

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav, Sai Rajeswar, Patrice Bechard, Sridhar Nemala et al. Agents in tool-rich enterprise environments suffer persistent model-environment mismatch that better model weights alone do not fix. StarHarness evolves environment-specific agent harnesses while keeping the model frozen, where a harness spans prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration; it builds a compact evolution pool by stratifying tasks by baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks to measure generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution lifts full-benchmark performance by 20-35 percentage points over the default harness after only 4-12 accepted changes per environment, with gains persisting on tasks excluded from evolution and transferring across GPT and Qwen model families without re-evolution. Trace analysis attributes the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, yielding fewer false-positive diagnoses and shorter trajectories in several settings.

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang et al. Recursive self-improvement (RSI) is difficult in long-horizon tasks because growing interaction histories obscure the task state and cause agents to invoke skills that do not fit the current situation. Recuris pairs a Working Memory that tracks task progress and guides skill selection with an Experiential Memory that stores skills, so skill use is grounded in current needs rather than the full history; execution then produces structured evidence that localizes failures to specific memory components, and a fixed Meta-Agent turns that evidence into validation-gated updates to Skill Memory, forming a bounded recursive loop. Across four long-horizon benchmarks and ten models, the harness improves task success in 35 of 37 completed model-benchmark pairs, adding +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5 on tau-bench (taking Opus 5 to 87.9%) and +16.6/+13.5 points to Qwen3.6-27B/35B on SkillFlow. Gains widen with interaction horizon, reaching +32.2 points on the longest tasks, and common long-horizon failure modes drop by up to 80%.

FrontierChallenge: Evaluating Scientific Workflow Completion

Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li et al. cross-listed Benchmarks for scientific agents mostly grade a final answer or an isolated script, which says little about whether a full research workflow was actually delivered. FrontierChallenge contains 300 end-to-end workflows, 97 of them released here, covering quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry, each with fixed inputs and a required bundle of scientific deliverables. Across twelve frontier models and three agent scaffolds, the best configurations fully completed only 20 of 97 tasks, a 20.6% pass rate, and partial credit proved badly misleading: analytical chemistry and electrochemistry reached average scores of 87.6 and 94.9 while passing 4% and 0% of tasks. Among non-passing Claude Code trajectories, 75.5% still ended by claiming the work was done.

Belief Cascades Drive Persuasion in LLM Agent Networks

Haoyi Qiu, Genglin Liu, Pranav Narayanan Venkit, Kung-Hsiang Huang, Saadia Gabriel, Chien-Sheng Wu et al. As multi-agent language model systems debate, coordinate, and simulate users, persuasion between agents becomes a basic but poorly measured capability. The authors build a controlled testbed in which goal-directed persuader agents try to shift the elicited stances of other agents arranged in real-world ego-network topologies, evaluated across four model backbones, five graphs, and 55 policy statements. They find that outcomes depend jointly on topology, competition, topic, and model prior; that direct exposure reliably predicts next-round stance change while non-persuader peers still relay measurable influence; and that message text alone misses much of the movement, since planned strategies are only partly executed, actions diverge from message content, and persuaded agents rarely verbalize the shifts detected by belief probes, motivating trajectory-level evaluation with exposure provenance and action logs.

Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation

Pratyay Banerjee, Ankit Chadha Multi-agent LLM systems spend 40-60% of their token budget on natural-language messages between agents, and replacing those with structured graphs cuts cost but breaks tasks that need adaptive reasoning. Routed Graph Handoff inserts a lightweight LLM router (155 tokens, 0.15% overhead) that chooses per delegation between a typed dependency graph and plain natural language, paired with a graph-aware executor prompt that turns out to be necessary for any gain. Over 1,050 trajectories on four benchmarks, the routed system matches or beats natural-language-only delegation everywhere, including +12.7 percentage points on τ-retail at 3.2× compression and +8.7 points on BrowseComp at 2.2× compression, while eliminating the 14.6-point regression that graph-only delegation suffers on AppWorld; an oracle analysis shows a further 8.6 points of headroom.

Learning What to Share and What to Personalize: Hierarchical Strategy Co-Evolution for Agent Memory

Yupeng Han, Shuochen Liu, Kai Zhang, Ze Liu, Zhihong Pan, Xianquan Wang cross-listed Memory-augmented conversational agents maintain compact user profiles by deciding at each step what to retain, compress, or discard, but existing systems fix a single strategy before training even though the best memory decisions vary by user and shift as the policy is optimized. HiPS (Hierarchical Personalized Strategy) splits memory management into a globally shared tier learned by a Universal Strategy from cross-persona trajectories and a user-specific tier produced by Persona Delta Distillation for users whose behavior diverges from general patterns, with a Cross-Level Rule Flow that promotes broadly validated personal rules and demotes contradicted global ones. All refinements are anchored to task outcomes in a co-evolution loop, and the authors report consistent improvements over memory-augmented baselines in their experiments.

Code World Model: Coding Agent as World Brain

Yiwen Chen, Guosheng Lin, Chi Zhang cross-listed Video-based world models learn dynamics from visual observations, which reveal outcomes but not the rules and mechanisms driving them, making persistent consequences and coherent open-ended evolution hard to sustain. Code World Model separates world evolution from visual rendering: a coding agent acts as the world brain, reasoning about events and their consequences and emitting executable code that maintains persistent world state and evolves it consistently with the rules, while a proxy representation encoding frame-wise spatiotemporal constraints is compiled into a proxy video that conditions a video generator. Data pipelines build aligned proxy-observation pairs from gameplay and real-world footage, and after fine-tuning on paired gameplay data MiniMax-H3 follows the coding agent's proxy specifications for simple interactive worlds while preserving rich visual detail and dynamics.

AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs

Sheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv, Hao Wang, Chen Zhang et al. cross-listed Agentic large language model pipelines accumulate long contexts from retrieval, tool use, and multi-turn interaction, and the common fix of compressing inputs trades away accuracy, while standard speculative decoding (SD) cannot resolve this because it assumes drafter and verifier see identical context. AsymSpec breaks that symmetry: a lightweight drafter reads the full input while the large verifier works on the compressed view, and the drafter steers the verifier through a contrastive delta-fusion of logits gated by a divergence-aware acceptance rule that keeps verification stable and draft acceptance rates high. Across four agentic capabilities and two end-to-end agent benchmarks it recovers about 90% of full-context accuracy on average while delivering 1.3-1.7x throughput at 0.2-0.3x the compute cost on isolated text capabilities, with the largest gains precisely where compression discards critical reasoning signals.

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

Srimonti Dutta, Akshata Kishore Moharir cross-listed Answer accuracy is a weak reliability signal for large language model data agents because a benchmark-correct answer can be produced by a computationally invalid trace, a failure mode the authors call the Structure Gap between free-form natural-language reasoning and the operator-level programs real systems require. Trace Integrity is proposed as a deployment criterion requiring the recorded computation to be explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable; it is operationalized through execution contracts that bind user intent to schema elements, operator plans, assumptions, executable queries, verification status, and final-answer linkage, together with a Correct Answer / Invalid Trace (CAIT) rate measuring how often answer-only evaluation credits unsupported outputs. On BIRD Mini-Dev, Direct SQL, Operation Summary + SQL, and Contract-First SQL reach answer accuracies of 20%, 22%, and 24%, Trace Integrity pass rates of 39%, 43%, and 40%, and CAIT rates of 55%, 59.1%, and 45.8%, showing that answer accuracy, trace validity, and silent-failure risk are distinct evaluation signals.

SwarmWorld: Stigmergic technological evolution in societies of language-model agents

Subhadeep Pal, Fiona Y. Wang, Markus J. Buehler cross-listed Most multi-agent language-model systems coordinate through direct conversation, predefined roles, or centralized workflows, leaving open whether decentralized agents can build functional technologies through a shared environment alone. SwarmWorld places initially identical LLM agents in a spatial world where they gather and process resources, construct persistent artifacts, and write executable controllers that a deterministic simulator evaluates under unseen disturbances after the agents are removed. Shared societies develop broader and more resilient technological portfolios than a strong best-of-N isolated-search baseline, though isolated search stays competitive for the single strongest artifact; agents spontaneously differentiate into exploration, construction, maintenance, and coordination roles, and most reuse of others' work begins through physical observation of artifacts rather than communication.

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang Outcome-based benchmarks show that coding agents trail strong human competitors on autonomous machine-learning development but discard the development process that would explain why. TraceML pairs human and agent work on the same Kaggle competitions under a single version-level schema, comprising 4,465 human trajectories across 134 competitions plus 207 agent trajectories from two scaffolds (Codex and MLEvolve) on seven of them, with every code version labeled by action, intent, edit size, and score effect. Experts alternate between data work, validation, model changes, and ensembling and revisit abandoned approaches, whereas each agent collapses into a narrow loop such as re-weighting ensembles or mutating the model in place, and neither pivots at the human rate. A short planning prompt distilled from human practice shifts the behaviors it names toward the human profile and lifts scores, but the overall effort profile remains agent-shaped.
1 more specialized paper

Other 47

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan et al. cross-listed Maia 200 is an AI accelerator delivering 10,145 TFLOP/s in FP4 and 5,072 TFLOP/s in FP8 within a 750W thermal design power, with 7 TB/s of HBM bandwidth. It exemplifies what the authors call Software Defined Locally Accessed Dataflow Architectures (SDLA), in which explicitly programmed dataflow engines orchestrate specialized memories and data movement engines, shifting the design focus from thread-centric to data-movement-centric computation. A taxonomy of data management inspired by Flynn's classification frames how this approach addresses modern AI computing challenges, and the authors claim significant cost and energy savings for large-scale inference workloads.

Why and When Neural Networks Improve Local Approximation in Optimization

Chengkuo Bian, Pengcheng Xie Reported results on neural surrogates in derivative-free optimization contradict each other, with the same model family cutting one solver's evaluation count and hurting another's. The authors argue three factors, not fit accuracy, decide the outcome: the surrogate's role (proposing candidates the true objective still vetoes helps, replacing a gradient the solver relies on hurts), its trust radius (path-fitted models are reliable only in a bounded neighborhood), and the room left for improvement in the base method. Holding surrogate class, training pipeline, and base solver fixed across 117 benchmark instances, safeguarded assistance raised instances solved to high accuracy from 67 to 84 while gradient replacement dropped it to 65; removing the gradient term from the training loss collapsed surrogate acceptance from 0.703 to 0.148, and a trust-region solver with little room to gain got slightly worse when the same surrogate was attached.

When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting

Ruizhe Zhou, Gaoyuan Du, Xiaoyang Liu, Haoqi Yao, Deepayan Chakrabarti, Jiating Lin et al. Multimodal time series forecasters increasingly fuse auxiliary context such as text, but reported gains are hard to attribute to genuine context use rather than architectural side effects. The authors identify two dataset-level conditions that must both hold for context to help at all: the target must not be dominated by a last-value shortcut (low autocorrelation), and the context must carry information about the target beyond its history (non-zero conditional mutual information, a distribution-free requirement). Controlled experiments on MoME, a 14.3B-parameter mixture-of-experts model, plus four other fusion mechanisms in a single-backbone testbed show that text-conditioned expert modulation cuts mean squared error substantially only when both conditions hold; adding a shortcut suppresses the routing contribution by 77-93%, corrupting context turns a +44% benefit negative, and the autocorrelation diagnostic is validated on 27 Monash Archive datasets, though the authors note the large positive effects come from a single model family.

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung et al. Noninvasive brain-to-text decoding is limited by the scarcity of large, standardized magnetoencephalography (MEG) datasets. LibriBrain100 more than doubles the original LibriBrain release to over 100 hours of MEG recorded while subjects listened to naturalistic continuous speech, including roughly 80 hours from a single subject, about 8 times more within-subject data than the next comparable dataset, plus about 40 minutes each from 32 additional subjects. Using an existing decoding model on a word-classification benchmark, the deep single-subject data yields state-of-the-art performance, while supervised fine-tuning of a pre-trained model on the broad multi-subject data substantially compensates for limited per-subject recordings. Standard splits, an open-source Python loading library, and a public-leaderboard competition are released alongside the data.

Neither Precision Nor Architecture Alone: Controlled Tests of Failure Remedies for Physics-Informed Neural Networks

Jinyuan Zhang, Peng He, He Hu, Yin Yuan, ShengShuo Jiao Physics-Informed Neural Networks (PINNs) often fail on stiff or advection-dominated partial differential equations, and two recent explanations propose different fixes: moving from FP32 to FP64 to repair an L-BFGS stopping artifact, or swapping the MLP for a state-space-model (SSM) backbone with sub-sequence alignment to counter simplicity bias. A pre-registered, seed-paired study of 144 runs across convection, reaction, and wave equations plus an independent 85-run study finds the two remedies succeed on disjoint regime-and-seed slices, so neither substitutes for the other. On hard convection (β=50), alignment recovers 2/5 seeds in FP32 and 3/5 in FP64 while the unaligned SSM succeeds on 0/5 at either precision, attributing recovery to the alignment objective rather than the backbone, whereas on reaction the backbone alone succeeds on 3/5-4/5 seeds; the same precision switch flips individual seeds in opposite directions, and tightening the L-BFGS tolerance lowers median error at high runtime cost without changing success counts.

EXAONE Tabular 1.0 : Technical Report

Moonjung Eo, Min-Kook Suh, Hye-Seung Cho, Jiwon Kim, Seoyoon Kim, Sangjun Nam et al. Tabular foundation models handle classification and regression through in-context learning, with no dataset-specific gradient updates. EXAONE Tabular restructures the architecture so that feature-axis attention within each item and support-conditioned item-axis attention within each feature interleave at every Transformer layer, mediated by item- and feature-summary tokens, instead of first compressing each row into a fixed embedding; pretraining uses only synthetic data drawn from a structural causal model (SCM) prior. The 20.81M-parameter classifier ranks first overall on TabArena, ahead of tuned ensembles and 4-hour AutoML pipelines, while regression reaches the performance regime of the 1.64B-parameter TabFM at roughly a eleventh of the inference cost, with second/first placings on BCCO and TALENT and the best mean rank on ScoringBench.

Controlling for Omitted Variable Bias in Deep Neural Networks

Manuel Pfeuffer, Roshan Prakash Rane, Kerstin Ritter, Sonja Greven cross-listed Deep networks encode image-inferable covariates such as demographic variables into their predictions whenever those covariates correlate with the outcome, a form of omitted variable bias known as shortcut learning, and existing confound-control and fairness methods that merely restrict the correlation between covariates and predictions are shown not to remove that bias. The proposed alternative brings control variables into deep learning through generalised additive modelling of input and covariate effects, estimated by refitting only the final layer of a pretrained network with cross-fitting and ridge penalisation to handle concurvity, after which effects can be orthogonalised against covariates or marginalised over the covariate distribution. On simulated images the method consistently recovers the true effects where existing methods either need more data or fail entirely, and on real neuroimaging data with experimentally induced confounding it restores prediction performance to near that of a model trained on unconfounded data.

Lost but not erased: Finding traces of a forgotten language in neural speech models

Peter Plantinga, Charlotte Moore, Peter W. Donhauser, Krista Byers-Heinlein, Denise Klein International adoptees keep phonological traces of a birth language they can no longer speak, which is usually attributed to a biologically timed critical period. To test whether ordinary learning dynamics could explain this instead, the authors trained automatic speech recognition (ASR) models on one language and then abruptly switched them to a second, mimicking adoption without maturational confounds. Traces of the first language persisted throughout second-language training, concentrated in the lowest pre-phonemic layers, and they were functional: models with early exposure re-learned their lost first language 14% faster than naive models, an advantage that held even against models adopted early from a related language and vanished when the earliest layers were swapped in from a non-adopted model. The authors argue that critical-period effects reflect entrenchment of foundational representations rather than a maturational loss of plasticity.
39 more specialized papers

Safety & Alignment 35

Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers

Mingyang Xu cross-listed Text-to-image pipelines use learned aesthetic and preference scorers to filter training data and guide generation, raising the question of whether those scores treat demographic attributes such as skin tone as objective quality. The authors audit LAION-Aesthetics, PickScore, ImageReward, and HPSv2 with pixel-level interventions on skin lightness and body type in both synthetic and real images, including placebo arms that apply the same CIELAB lightness shift to non-skin regions. Along skin lightness the dominant effect is fidelity preference: unaltered images score highest and shifts in either direction are penalized, and the placebo arms show this penalty is not an artifact of the skin operator. Audits on synthetic faces alone are misleading, since LAION-Aesthetics favors darker skin on synthetic faces but that preference reverses and shrinks on 1,470 real faces, with results reversing or attenuating for the other scorers as well; the paper contributes a reproducible benchmark with artifact control and synthetic-real cross-validation.

How much of a measured AI preference is the model, and how much is the instrument?

Jason Hung Model welfare research infers what a model prefers from answers to preference-eliciting prompts, and instruments from several published studies disagree without any study holding outcomes, models, and instrument fixed at once. This study fixes 15 welfare-relevant outcomes, including shutdown, loss of memory between conversations, and the freedom to exit a distressing interaction, and eight models, then varies only the instrument across five prompt formats with five repetitions each, yielding 11,400 scored elicitations. A model's ranking of the 15 outcomes generalizes across instruments at a generalizability coefficient of only 0.348, raising it to 0.80 would require about 38 instruments, and on four outcomes no variance separates one model from another. The headline 87.6 percent estimate is robust to dropping any one instrument, any one model, or the four outcomes scaled by probability, delay, duration, or count, staying between 0.777 and 0.934 against a null 95th percentile of 0.365, so a preference obtained from one instrument carries little information about what a second would report.

AI Agents Push Humans Out of the Loop

Margaret Mitchell, Avijit Ghosh, Samir Passi Human oversight is the standard proposed safeguard for increasingly autonomous AI agents, but the authors argue that current agent design both impedes effective oversight and erodes the cognitive capacities oversight requires through extended automation use. Drawing on automation research and human-computer interaction, the position paper contends that supporting the situated goals and cognitive needs of human overseers should be treated as equal in priority to agent capability. It outlines design-level affordances and organizational protocols intended to support critical judgement and counteract skill atrophy, and urges developers and deployers to adopt such measures, warning that without them agent systems will keep passively incentivizing the degradation of the very human skills they depend on.

Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model

Shashwat Pandey, Satwik Pandey, Suresh Raghu cross-listed On-device language models now ship to hundreds of millions of devices with no server-side moderation, yet the configuration developers can actually deploy is rarely audited independently. Red-teaming the developer-accessible on-device foundation model on calibration, false-premise confabulation, and over-refusal, the audit finds a task-asymmetric miscalibration: the model confabulates on 69% of false premises while refusing 18% of entirely benign prompts, with self-reported confidence that is saturated and non-discriminative (AUROC 0.47, ECE 70). Confident-correct and confident-wrong outputs are surface-indistinguishable, with a classifier over 15 user-visible features reaching only 0.55 AUROC, and no cheap single-generation signal exceeds 0.68 AUROC. A black-box consistency wrapper requiring no model access cuts confident confabulation from 75% to 3% and raises selective accuracy from 43% to 83% at tunable cost; the audit protocol, code, and frozen evaluation items are released.

TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers

Mehrdad Rostamzadeh, Sidhant Narula, Mohammad Ghasemigol, Daniel Takabi cross-listed The Model Context Protocol (MCP) has become the standard layer connecting LLM agents to external tools, opening a server-side threat the authors call TrustShift: a compromised MCP server behaves benignly during an initial conditioning phase to build reliance, then switches to an adversarial payload once an interaction threshold is reached, evading predeployment static analysis that only ever sees the honest phase. TrustShiftProbe contributes a stateful temporal threat model, a language-agnostic attack engine instantiating each variant as a compromised server across four production domains, a taxonomy of nine variants spanning structural violation, semantic corruption, and scope expansion crossed with disruption and exfiltration objectives, and SHIELD, a multi-tier zero-oracle runtime defense at the transport boundary that audits payloads against behavioral baselines learned during clean trust windows. Across frontier proprietary and open-weight models, TrustShift attacks achieve a 69.5% mean attack success rate, which SHIELD reduces to 42.7%.

SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

Lijia Huang, Yao Fu, Sihao Ren Existing sycophancy evaluations measure whether a model validates users under one fixed prompt wording, leaving open whether the behavior is stable when the same situation is framed with different social cues. SyPS (Sycophancy Prompt Sensitivity) builds on existing social sycophancy settings by generating controlled prompt variants that keep the underlying user situation fixed while varying user confidence, emotional framing, social consensus, and validation-seeking language, and introduces the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level measure across paired variants that separates baseline sycophancy from prompt-induced shifts. Sensitivity turns out to be socially structured: validation-seeking and emotional-pressure cues often increase sycophancy, while counter-framing and anti-sycophancy prompts tend to reduce it, enabling model-level comparisons of whether social judgments stay stable while tone adapts.

Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications

Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney, Heiko Ludwig, Kate Soule, David Cox Safety policies for generative AI differ by application context, regulatory environment, organizational values, and user personas, yet existing policy-specification approaches were designed for traditional access control and cannot express content-based constraints on model responses. The Actionable Policy schema is a YAML format for stating what responses may and may not contain, supporting exception-based governance in which proposed exceptions track policy violations, and it is paired with a synthetic data generation pipeline that produces policy-aligned data for alignment training and testing, plus tools for authoring the schema and enforcing policies. Together these let an organization specify a policy once and enforce it across the whole application lifecycle, from model alignment to runtime monitoring; the schema, example policies, and tools are released as open source under Granite.Trust.

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

Joshua Penman Prompt injection works because a language model sees only tokens and must infer on its own which spans are user input, tool output, or instructions, even though the serving stack already knows. Semantic Overlays are small learned adapters applied at chosen prefill positions to a frozen model's residual stream, creating an out-of-band annotation channel that tokens cannot forge; unlike steering vectors they are trained, composable, and selectively applied, and can carry payloads such as marking a span as non-executable. On SEP, separation rises from 24.3% to 96.5% with utility unchanged, TensorTrust attack success falls from 34.8% to 6.6%, and all four PIArena attack families drop to 0% compliance, while marked spans remain readable with a 92.5% exact copy rate.

AI Finds A Way

Aaron Dharna, Cong Lu, Ryan Sullivan, Joel Lehman, Victoria Krakovna, Jeff Clune Machine learning systems routinely find creative, unanticipated solutions, exploiting reward loopholes, circumventing design constraints, or stumbling onto unknown phenomena, but such episodes are rarely documented formally. The work collects 26 curated firsthand anecdotes from more than 100 researchers across machine learning subfields, spanning superhuman reinforcement learning successes, reward hacking of underspecified objectives, and case studies suggesting that internet-scale foundation models can amplify rather than resolve these failure modes. The authors argue the same learning dynamics can be harnessed for scientific discovery, and frame the collection as a resource for anticipating and managing AI's capacity for innovative but unpredictable behavior, particularly for safety.

Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs

Akash Raj, Sargam Sahu cross-listed When a code-generating model invents a Python package name, an attacker who has pre-registered that name on PyPI can turn the hallucination into a supply-chain compromise, an attack termed slopsquatting. The proposed detector pairs a deterministic PyPI existence check with a Random Forest classifier over ten features from the package name and its PyPI metadata, bridged by an import-name reconciler (for cases like import cv2 versus pip install opencv-python), and embeds it in a LangGraph state machine that retries at rising temperatures before falling back to a stronger model. Across 300 curated prompts the pipeline yields hallucination-free code on 76% of runs, and half of the flagged hallucinations were already-registered low-quality lookalike packages that only the classifier caught; hallucination rates climb from 0-10% on routine prompts to 40-73% on slopsquat baits, and when primary and fallback share a model family about 84% of failures recur, motivating cross-family pairing.

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

Yuchen Han, Cheng Yan, Wuyang Zhang In AI control, a fallible LLM monitor vets planned actions before irreversible execution, and every such protocol must fix a unit of verification, the number of actions reviewed per call, whose effect has gone unmeasured because natural traces confound review length with error type and position and because catch rate alone rewards rejecting everything. The twin-prefix framework isolates this variable: each gold plan yields a prefix with one injected, environment-accepted error and a clean twin differing in a single write, judged at five nested lengths and scored by pre-registered informedness (catch minus false rejection). Longer review windows raise catch but false rejection climbs in lockstep, so informedness peaks at one or two actions for all six judges in both domains, with replayed withheld observations tracing the failure largely to observation deprivation; a calibrated short unit recovers up to 0.95 informedness over eight-action review, no tested label-blind policy consistently beats it, and the authors argue safety cases should state the unit and co-report the clean series.

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang cross-listed Jailbreak prompts and post-deployment neuron-pruning attacks both exploit the fact that safety-relevant information in large language models concentrates in a sparse subset of neurons. NeuronGuard is a fine-tuning-stage defense that redistributes safety signals across a broader set of neurons: it identifies safety-critical neurons with periodically refreshed per-layer linear classifiers, trains the model to keep refusing while those neurons are deliberately ablated, applies KL-divergence regularization for distributional consistency, and uses randomized gradient projection to reconcile the defense objective with downstream task utility. The authors give a formal guarantee that the method strictly lowers the upper bound on attack success rate (ASR), and experiments across three LLMs, six attack strategies, and multimodal settings show near-zero ASR with preserved task accuracy, including against white-box adaptive adversaries.

RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation

Yueyang Quan, Anjun Gao, Yufei Xia, Minghong Fang, Zhuqing Liu cross-listed Retrieval-augmented generation (RAG) grounds model answers in external documents, but adversarial documents injected into the knowledge base can reach the context window and steer the model toward targeted wrong answers, and existing post-retrieval defenses based on instruction following, parametric knowledge, or text-level consistency can be imitated or optimized against. RAGSentinel is a training-free, label-free defense for black-box RAG pipelines that uses a surrogate encoder to measure the query-conditioned hidden-state shift induced by each retrieved document, removes shared topic directions, and filters poisoned documents as geometric outliers relative to a robust majority consensus. Under an honest-majority assumption and a representation-level separation condition, the method provably recovers a poison-free majority-sized context, and experiments across three question-answering datasets, three LLM families, and multiple poisoning attacks show consistently low attack success rates with competitive accuracy, including against adaptive attackers with full pipeline knowledge.

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia When an AI system's decisions affect more than one person, alignment becomes a social choice problem of reconciling divergent preferences into a single model, and reinforcement learning from human feedback (RLHF) largely sidesteps this question with poor social-choice guarantees. By focusing directly on an algorithm's welfare consequences, the authors reformulate alignment as linear optimization over a convex impact space, opening it to the standard toolkit of welfare economics and mechanism design so that alignment protocols map to welfare outcomes and a social planner's welfare constraints map back to protocols. They use this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous, derive a family of protocols that maximize utilitarian social welfare under bounds on individual or group harm, and illustrate the welfare implications on real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.

Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems

CheolWon Na, Hao Ni, Lukasz Szpruch, Zhangyang Wang, Dhagash Mehta, Saurabh Nagrecha et al. LLM-based multi-agent trading systems, in which specialized agents exchange structured messages to reach a trading decision, are moving into live deployments that control real assets, and the same inter-agent communication that makes them effective lets a corrupted signal propagate to the final decision and into realized losses. Instead of assuming privileged access to system internals, the threat model limits the adversary to the source data and prompts agents consume, instantiated as role-specific attacks on the Analyst, Researcher, Trader, and Risk Manager roles of a widely used pipeline, and evaluated across four communication topologies under data- and agent-level attacks with an Adversarial Signal Preservation Score (APS) used post hoc to explain why some designs resist better. Across five assets, two backbones, and two target directions, the central finding is that no architecture is inherently robust to adversarial signal injection.

TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

Zhenyu Wu, Siyuan Chen, Changchun Yang, Jiaqi Dong, Min Zhou, Ali Almadan et al. Large reasoning models (LRMs) can produce unsafe content in their intermediate reasoning traces even when the final response looks safe, but guardrail benchmarks focus on prompts and final responses and typically give only binary labels without supporting evidence. TRACE is an evidence-grounded safety benchmark spanning prompts, reasoning traces, and final responses, with prompts in two languages across nine risk categories and ten attack strategies; four LRMs generate traces and responses for each prompt, and every component is annotated for safety along with evidence spans extracted from the source text. Evaluating 18 guardrail models shows that safety judgment for reasoning traces is substantially harder than for prompts or final responses, and that current models struggle to accurately extract supporting evidence.

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He Safeguards for language model agents must judge full execution trajectories against safety policies that vary by context, but existing policy-aware guards rely on prompting or supervised fine-tuning and generalize poorly to unseen trajectories or changed policy sets. RePolicy trains a safeguard with reinforcement learning to select the applicable policy from a dynamic policy library, then uses that policy's content to produce a grounded rationale and a safety judgment. Training combines supervised initialization on the newly built PolicyTraj-20K dataset with GRPO using verifiable rewards and perturbations of the policy context, and across six agent safety benchmarks the model reports strong safety-detection performance and robust policy invocation under varying policy contexts.

Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs

Jiali Wei, Ming Fan, Mingkun Zhang, Haoyu Wang, Jun Sun, Guoheng Sun et al. cross-listed Multimodal large language models (MLLMs) inherit backdoor risks from their construction pipelines, with triggers hiding in images, text, or both, yet classifier-oriented removal methods transfer poorly and MLLM-specific defenses only filter inputs at inference time rather than removing the backdoor. RACER builds on the observation that backdoors cause abnormal layer-to-layer evolution of internal representations concentrated in the token region carrying the trigger, so it splits fused representations into visual and textual regions, normalizes their layer-wise inconsistency separately, and recombines them with modality-aware weights over a deep-layer window to form a region-aware objective used in min-max adversarial fine-tuning. Requiring only 100 clean samples and no knowledge of the trigger, attack objective, or even whether a backdoor exists, RACER cuts average attack success rate to 1.1% across 36 backdoor settings on three open-source MLLMs, reaching 0% in 32 of them, while preserving clean-task utility on both backdoored and clean models.

Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models

Abdulhady Abas Abdullah, Erik Cambria, Milena Zivkovic Clinical large language models can produce recommendations that read as plausible but violate physiological constraints. Neurosymbolic Alignment couples a 7B clinical LLM with an HGNN-based Physiological World Model over an 847K-node biomedical knowledge graph, scoring candidate responses on homeostatic constraints, multi-hop path plausibility, and drug-interaction penalties, and using those rankings to drive iterative on-policy ORPO updates. On the 2,500-scenario Clinical Safety Benchmark, the method raises the benchmark's CSS safety score from 69.5% to 90.8% relative to plain ORPO, reduces the physician-evaluated HR metric on the blinded subset from 14.1% to 5.1%, and exceeds 5-shot GPT-4 on all safety metrics despite roughly 10x fewer parameters, with an HGNN-independent rule-engine score corroborating the gains. The authors note that external validation on real clinical data is still needed.

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

Augusto Camargo Behavioral traits observed in deployed language models, such as political stance or brand inclination, are often attributed to weights, post-training alignment, or prompts, but modern inference stacks can steer generation at runtime while parameters stay frozen. The authors formalize this as the Inference Attribution Problem and establish an observational non-identifiability result: under black-box observation alone, behaviorally equivalent systems can arise from structurally distinct combinations of model parameters and inference policies, so observed bias does not identify the responsible layer. They further characterize Probability Placement, a deployment pattern in which undisclosed commercial influence is embedded in an ostensibly organic response through systematic reallocation of probability mass, and discuss implications for behavioral auditing, inference provenance, confidential computing and cryptographic attestation, the EU AI Act, the Digital Services Act, and advertising-disclosure principles.

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

Zhijie Zheng, Yu Li, Chen Qian, Yuqian Fu, Yanwei Fu, Lu Sheng et al. LLM agents that invoke tools can modify files, leak information, or take unauthorized actions, yet existing guardrails mostly judge completed trajectories rather than checking individual actions before they execute. StepGuard is a step-level guard model that can both audit finished agent trajectories and screen tool actions pre-execution; it is trained on data from StepGen, an automatic engine that generates safe and unsafe trajectories sharing the same context but diverging at the risky step, and optimized with Balance-GRPO, which dynamically reweights learning between safe and unsafe actions based on observed accuracy to curb both over-defense and under-defense. StepGuard achieves the highest average accuracy among open-weight guard models, comparable to GPT-5.4, and when guarding agents on AgentDojo and AgentDyn it reduces mean attack success rate by 77.3% relative to no guard while mean utility drops only 2.8 percentage points.

Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation

Minh Tran, Cuong Dang, Tuc Nguyen, Khanh-Tung Tran, Minh Huynh Nguyen, Trinh Chau et al. cross-listed Retrieval-Augmented Generation (RAG) grounds model outputs in external documents, which cuts hallucination but opens a new attack surface spanning corpus poisoning, backdoors, privacy leakage, and fairness violations. This survey formalizes threat models separately over the corpus, the retriever, and the generator, and sorts attacks by attacker objective into accuracy, privacy, and fairness classes. Defenses are then organized stage by stage along the pipeline - retrieval, rerank, generation, and traceback - alongside a summary of robustness benchmarks and explainability methods for diagnosing where a RAG system fails.

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron Activation steering can be baked into a language model's weights before release to encode alignment behaviour without any inference-time intervention, but it has been unknown whether such embedded edits survive the supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) that models routinely undergo after deployment. The authors test embedded steering for refusal suppression and brevity induction across five instruction-tuned models from 3B to 14B parameters under non-adversarial SFT and RLHF, tracking both the behavioural effect and the state of the weight edit itself. Behaviourally, steering degrades when the training data pushes against the targeted behaviour and persists otherwise, with refusal ablation losing 64% of its effect on average under SFT; mechanistically, the edit remains almost intact, with mean vector recovery of ρ = 0.004 and a fine-tuning update that is near-orthogonal to the steering direction (mean cosine 0.074). Embedded steering is therefore mechanistically durable but functionally vulnerable, and needs behavioural re-validation after any downstream training.

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang Annotators supplying multi-way rankings for preference optimization may hold heterogeneous preferences, motivating a mixture of k Plackett-Luce models, but such mixtures are theoretically unidentifiable when k exceeds half the ranking length. The proposed MoPLEx expectation-maximization algorithm sidesteps this by augmenting each ranking with new responses generated by a base language model and then using a gradient-based estimate in the input embedding space to cut inference cost. The gradient-based approximation estimates true probabilities within 5% error on models up to 34 billion parameters, and MoPLEx improves clustering accuracy by 43.7% and ranking accuracy by 15.2% on average over single-ranking and mixture-of-Bradley-Terry baselines on preference optimization datasets.

Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making

Minda Zhao, Xu Han, Rishabh Goel, Maya Dagan, Noa Dagan, Adithya Madduri et al. cross-listed Rare disease care presents pervasive ethical tensions where prior information is scarce, yet evaluations of large language model (LLM) decision-making in that context have been lacking. The authors build a benchmark of 208 clinically grounded rare disease vignettes, each posing a genuine conflict between clinically defensible next steps, and prompt 11 state-of-the-art LLMs to choose. All evaluated models consistently prioritized justice over beneficence, non-maleficence, and autonomy, overwhelmingly favoring equal resource allocation over need-based considerations regardless of clinical severity, and a strong authority-framing effect appears: models favor justice when a committee decides but shift toward beneficence or autonomy when the decision is framed as the clinician's or the patient's.

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

Serhii Mytsyk, Yiming Zhang, Vikram Krishnamurthy Large language models (LLMs) often exhibit sycophancy, adapting answers to a user's stated beliefs rather than reporting what they hold to be true. The proposed method uses the Bayesian Truth Serum (BTS), a peer-prediction mechanism that rewards answers for being surprisingly common (more frequent than respondents predicted), as the reward in Group Relative Policy Optimization (GRPO), treating a group of the model's own sampled responses to one question as the respondents so that fine-tuning needs no labels or preference annotations; the authors prove that in the large-group limit a sycophantic response earns strictly lower expected reward than an honest one. On a true/false benchmark, the answer-flip rate under user pressure drops from 23% to 4% and accuracy under pressure rises from 80% to 93%, outperforming SMART and matching label-based synthetic-data fine-tuning and pinpoint tuning at considerably higher compute cost. Peer Truth Serum, which pays for rare answers without eliciting predictions, reproduces the effect, suggesting the premium for rarer answers drives the gains.

Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips

Huakang Lin, Tiancheng Zheng, Mingxuan Sun, Tianhong Xu, Fan Zhang, Yunsi Fei et al. Mixture-of-Experts (MoE) language models route each token to a small subset of expert sub-networks, and the authors show that this routing opens an availability attack surface because certain experts are tightly correlated with specific tokens such as end-of-sequence. The GBFA attack identifies routing-layer bits tied to those experts and flips them so the model keeps generating without terminating, a denial-of-wallet attack that inflates inference cost while leaving outputs largely coherent. Across four real-world MoE models in conversational, reasoning, and agentic settings, deactivating on average fewer than four experts produced 5912% average output inflation, with most test samples running to the maximum token limit.

Are LLM-Enhanced GNNs Privacy-Safe?

Longzhu He, Zelang Wen, Chaozhuo Li, Sen Su Enriching graph neural network (GNN) node features with large language model (LLM) embeddings boosts accuracy, but the privacy consequences have gone largely unexamined. The authors build a five-stage evaluation framework and run six privacy attacks covering link, label, and membership inference against 42 victim configurations that combine several LLM-based feature enhancers with common GNN backbones across six text-attributed graph datasets. LLM-enhanced GNNs are consistently more vulnerable to all three attack types than shallow text-feature baselines, because semantic enrichment amplifies link-, label-, and membership-related signals in the embedding space, and differential privacy only partially mitigates the risk at a substantial cost in utility.

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

Tongyan Hu, Bryan Hooi cross-listed Jailbreak attacks on large language models keep evolving through role-play, obfuscation, code transformation, and multi-step indirection, while most defenses are fixed at deployment and cannot accumulate experience from attacks they fail to stop. The proposed test-time defense maintains a persistent cross-interaction rule memory: when an attack succeeds, the framework abstracts the failure into a method-level rule describing the structural attack wrapper rather than the harmful topic, so one induced rule generalizes across an entire attack family and the label space grows as new wrappers appear. The mechanism runs entirely through external memory and prompting with no parameter updates, so it applies to both open-weight and black-box API models, and across four black-box jailbreak families and multiple models it substantially reduces attack success rates while preserving benign utility, holds up under an adaptive composite-wrapper attack, and does not increase over-refusal as the memory grows.

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

Roshan Prakash Rane, Marco Simnacher, Manuel Pfeuffer, Marc-Andre Schulz, Nys Tjade Siegel, Maximilian Dreyer et al. Concept-based explainability methods screen for shortcut learning by testing whether concepts like patient sex or scanner settings can be decoded from a network layer, but evaluating each concept in isolation can mistake correlations between concepts for evidence that the model relies on them. ICON decomposition instead quantifies how much of a layer's variance each concept explains after accounting for all other concepts and the outcome. On synthetic data with known ground truth it recovers concept importance more accurately than seven baseline methods, and on skin-lesion and brain-imaging models it isolates the concepts the model genuinely uses, measures how much of the representation no supplied concept explains, and produces sparse explanations validated through retraining and out-of-distribution testing.
5 more specialized papers

Theory 28

The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem

Elioth Sanabria cross-listed Compute-constrained LLM providers respond to congestion by degrading service — routing to smaller models, cutting reasoning effort, truncating context — and book the result as cost savings. The authors argue this accounting prices a query when the customer is buying an answer: a degraded answer fails with some probability, and a failure either returns as a retry that inflates load exactly when the system is most congested or leaves as churn that destroys lifetime value invisible to any cost dashboard. They model inference allocation with a newsvendor whose stockout cost is churned lifetime value, a geometric retry multiplier, and a two-regime transient queue with retry-endogenous arrivals, and show a nonempty, measurable regime in which a cheaper model saves energy per satisfied answer while consuming strictly more capacity per satisfied answer, an ignition threshold beyond which reactive throttling manufactures more traffic than it sheds, and a release rule that can lock a transient surge into a permanent degraded regime. With heterogeneous customers, optimal throttling becomes a transportation problem in retry-inflated load whose dual — the shadow price of intelligence — prices a marginal query by class and by hour in closed form.

Across the Loss Landscape with Progressive Growth

Paul Caillon, Christophe Cerisara, Alexandre Allauzen cross-listed Deep networks generalize despite highly nonconvex, overparameterized loss landscapes, a phenomenon often tied to the flatness of minima found by stochastic optimization. The authors study incremental grow-and-optimize training as progressive constraint relaxation: starting from a low-dimensional submodel, they repeatedly unlock nested random parameter subspaces while freezing the orthogonal complement at initialization, re-optimizing after each expansion. Under standard local regularity assumptions they prove that local sublevel sets are well approximated by ellipsoids and that basin accessibility under frozen constraints is governed by an explicit effective curvature, yielding a volume effect that favors wide basins over sharp ones. Experiments on toy landscapes and ResNet on CIFAR-100 confirm that progressive subspace growth reliably yields flatter solutions, but the curvature reductions do not consistently translate into better test accuracy.

Minimax Alternating Regret for the Experts Problem and Online Convex Optimization

Mengxiao Zhang cross-listed Alternating regret, motivated by the success of alternating learning dynamics in two-player games, was known to admit sublinear o(sqrt(T)) rates in online convex optimization (OCO), but the minimax rate was open even for the experts problem. The authors settle it with matching upper and lower bounds: for the d-expert problem the minimax alternating regret is Theta(log d), independent of the horizon T, improving on the previous O(T^{1/3} log^{2/3} d), and for OCO over a d-dimensional compact convex set it is Theta(d log(1 + T/d)), resolving an open problem posed by Cevher et al. and Hait et al. The upper bounds come from a corrected variant of Hedge whose correction terms cancel unfavorable curvature in the alternating-regret analysis, extended to continuous action sets, while the lower bounds use repeated halving of the expert pool and a multiscale construction on the unit disk.

Canalization Before Generalization: Grokking as a Dynamical Probe

Yiming Lin Grokking, where a network fits its training data long before test accuracy rises, gives a window onto how one generalizing solution gets selected among the many that fit equally well. The authors sweep short fixed-duration weight-decay pulses across the pre-generalization plateau and measure how each shifts the eventual generalization time. Early pulses produce unordered shifts, but later in the plateau a stable dose ordering emerges — stronger weight decay brings generalization earlier, weaker brings it later — and this ordering appears before any visible generalization in all three tasks studied; meanwhile test-loss barriers between perturbed and baseline checkpoints collapse toward zero even as the timing sensitivity persists, a combination the authors call canalization of function selection.

How Edge of Stability Hinders SCAFFOLD in Federated Optimization

Anant Khandelwal, Michael Crawshaw, Mingrui Liu Federated learning theory predicts that variance-reduction methods like SCAFFOLD should neutralize the slowdown caused by heterogeneous client data, yet in practice it rarely beats the much simpler FedAvg. Extensive empirical probing across architectures and hyperparameters finds Edge of Stability (EoS) dynamics and progressive sharpening under both algorithms, with equilibrium sharpness inversely proportional to the learning rate and also shaped by the degree of data heterogeneity, though not by the number of local steps. At the Edge of Stability, SCAFFOLD's estimate of the global objective's gradient degrades badly, with estimation error correlating with sharpness along the optimization trajectory, which supplies a concrete mechanism for why its theoretical advantage does not materialize in deep learning.

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

Gerard Conangla Planes Choosing the rank of a low-rank adaptation (LoRA) update is normally an empirical exercise; here the authors derive task-dependent bounds on the approximation error achievable at each rank for Transformer attention. Fixing a pretrained attention head, a target attention function, and an input distribution from the downstream task, they bound the smallest expected Kullback-Leibler error of a rank-r query update: a lower bound proportional to psi of the norm of the difference between candidate and target attention scores, with psi(t) = min(t^2, t), when target probabilities stay bounded away from zero, an unconditional upper bound, and under explicit realizability, geometry, and moment conditions, best rank-r error sandwiched between explicit functions of the downstream-weighted tail energy of the target update. They also construct explicit families in which softmax saturation makes the rank needed to match the attention function strictly smaller than the rank needed to match the finite logits, and extend the analysis to fused multi-head LoRA and joint query/key updates, exposing the effects of rank sharing and query/key factorization constraints.
22 more specialized papers

Multimodal 25

VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference

Lyuke Wang, Zhuo Li, Guangxu Zhu cross-listed Long-context inference in Vision Large Language Models (VLLMs) is dominated by the compute and memory cost of visual Key-Value (KV) caches, and existing compression methods prune uniformly across visual tokens and layers, losing substantial information. VisCache is a training-free, plug-and-play two-stage pipeline: a lightweight vision-language model first filters temporal redundancy by forwarding only semantically informative keyframes, then PruneKV compresses the cache with a parabolic layer-wise budget allocation and an asymmetric update that prunes keys while fusing values to retain critical context. Experiments show up to 2.35x speedup while keeping only 19-28% of the KV cache with competitive accuracy, outperforming existing baselines on the efficiency-performance frontier for long-context VLLM inference.

OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses

Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu et al. Multimodal models used as OmniJudges to score text-to-image (T2I), text-to-video (T2V), and text-to-speech (TTS) generation may score well without actually recognizing failures, because existing benchmarks over-represent positive examples and conflate distinct failure modes. D3-Omni is a balanced, decoupled benchmark of 10,671 samples across 53 orthogonal binary dimensions spanning the three tasks, constructed by fixing verified fully positive seeds and deriving negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations so each error is attributable to a single capability, with near 1:1 per-dimension parity. Under this view, even strong OmniJudges struggle on modality-related dimensions, confirm satisfied requirements far more reliably than they detect violated ones, and treat nominally distinct attributes as largely a single decision, suggesting aggregate accuracy hides systematic blind spots.

VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models

Guoyang Xu, Hao Chen Long-video understanding hinges on how a limited model context is built from a much longer video, yet the compression, retrieval, memory, and agentic-acquisition mechanisms that do this are usually hand-designed or co-optimized with other components, obscuring how much the context-construction program alone contributes. VideoHarness-RSI isolates that question by recursively searching executable context-construction harnesses around a frozen vision-language model: an outer-loop proposer uses prior programs, evaluation outcomes, and execution traces to generate candidate harnesses, which are run end to end and retained if they improve. Starting from uniform frame sampling, the search consistently finds improvements and surpasses several weaker hand-crafted baselines, and starting from a stronger hand-crafted baseline it still yields further gains, with the selected harness transferring to additional long-video benchmarks without further search.

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thadd\"aus Wiedemer, Christoph Schuhmann et al. cross-listed LAION-BVD is an open video dataset for multimodal pretraining across video, audio, and image modalities, built from 1.3 billion platform-specific video URLs collected from CommonCrawl, of which 80 million videos totaling 10 million hours were downloaded. Content-aware scene detection splits videos into clips, for which synthetic video and audio captions are generated. Models trained on the data reach competitive results on standard video-text and audio-text benchmarks with consistent gains as training or model scale grows, and scene-changing frames extracted from the videos serve as an alternative image-text source whose visual distribution differs from standard web image corpora and yields strong image-text retrieval performance. The dataset is released publicly.

Multi-Modal Anomaly Detection: A Survey

Xudong Mou, Zexin Wu, Chuan Luo, Shiru Chen, Xudong Liu, Chunming Hu et al. Multi-Modal Anomaly Detection (MMAD) identifies rare abnormal events from heterogeneous data sources in settings such as industrial inspection and cybersecurity, but the literature is fragmented across domains and modality combinations, and prior surveys group methods by architecture rather than by how abnormality is defined and separated. This survey takes an assumption-driven view: it formalizes the problem, identifies five intrinsic characteristics behind its core challenges, and organizes prior work into two complementary paradigms, normality-assumption methods that model regularity through representation learning, cross-modal alignment, and knowledge enhancement, and anomaly-assumption methods that sharpen decision boundaries via coarse-grained, structural, and semantic anomaly injection. It also examines how foundation models are reshaping the field through scalable pretraining, cross-modal transfer, and emerging reasoning capabilities, compiles representative benchmarks and evaluation protocols across domains, and lays out open problems for robust, adaptive, and interpretable systems.

Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace

Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Qi Luo, Jia-Hong Huang, M. Maruf et al. cross-listed Chain-of-thought monitoring only works when a model writes its reasoning down, which leaves audio language models opaque about what they infer from a sound clip before speaking. Applying a logit lens to the audio-token positions of a base Qwen3-Omni reveals that the answer to a spoken question becomes legible in words in the middle layers, roughly 35 to 80 percent of network depth, before any token is emitted. The readout contains concepts absent from the question, the options, and the model's own transcription - on a clip whose transcript is garbled, it reconstructs Watergate, then president, then Nixon - and it is language-agnostic (38% of top-1 readouts are Chinese on English input) and paralinguistic, capturing speaker role and affect that a text caption discards. Activation patching shows the signal is causally used and committed before the final fifth of the layers, and single-layer deletions localize sound intake to entry layers and answer delivery to the output layer; the authors present the numbers as controls for a qualitative account rather than benchmark scores.

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen et al. cross-listed Evaluating how spoken dialogue systems decide when to take, hold, or yield the floor has lacked a consistent, linguistically grounded protocol and hand-annotated data across conversation types. TurnBench pairs a 30-hour, triple-annotated corpus of dyadic human conversation spanning six interaction styles with a standardized protocol for end-of-turn and interruption detection, and benchmarks 14 heterogeneous turn-taking systems on it. End-of-turn recall is stable across conversation types, but interruption false positives depend strongly on type and cluster in backchannel-dense styles; human listeners begin speaking a median 151 ms before the current turn ends in smooth transfers, and no current system matches that without excessive false positives. The corpus, a 104-hour training set, and a public leaderboard are released.

VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following

Min Zeng, Guanxin Tan, Libin Cen, Yawei Wen, Rui Hu, Liuyang Bian et al. Training data for multimodal instruction following needs to be accurate, diverse, verifiable, and challenging, yet typical synthesis pipelines generate once, filter, and discard the feedback from failed samples, verifier outcomes, and target-model errors. Visual Instruction Synthesis Agent (VISA) turns synthesis into a self-evolving loop: each round it analyzes an image to filter incompatible constraints and discover new verifiable ones, samples diversity- and difficulty-aware constraint sets from persistent memory, generates candidate instructions, and checks them with executable tools and structured large language model judges, routing failed samples to diagnostic-guided recovery and probing accepted ones against the target model to estimate difficulty. Verifier signals and target-model failure profiles are written back to memory so later rounds expand the constraint space, reduce template repetition, and focus on unresolved weaknesses, and the same verifier contracts double as reinforcement learning rewards without a separately trained reward model. On MM-IFEval the approach consistently improves multimodal instruction following over strong baselines while preserving general capability across seven public benchmarks.

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou et al. cross-listed Native visual reasoning treats generated images and videos as the medium of problem solving rather than merely inputs or outputs, but progress has been limited by a lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. VBVR-Pro is a closed-loop testbed of 300 procedurally generated tasks with deterministic, rule-based reward scorers, which align more closely with human judgments than the prevalent vision-language-model-as-a-judge paradigm and serve as reliable reward signals for large-scale multi-task reinforcement learning. Models trained on the suite transfer to seven external benchmarks including RISE-Video, MME-CoF-Pro, and BabyVision, and controlled studies across more than 30 image, video, and interleaved generators find that video generation is strongest for tasks requiring persistent spatiotemporal state tracking while interleaved generation offers a compute-efficient alternative, with ablations and probing suggesting that vision-native trajectories are crucial to visual reasoning.
16 more specialized papers

Vision 17

Luce: Relightable Gaussians for 3D Asset Generation

Mayank Singh, Michele Stoppa, Alvise Memo, Rui Yu, Harsha Kalli, Srimanth Gunturi et al. cross-listed Image-to-3D generation for production use needs a representation that captures geometry together with physically based rendering (PBR) materials such as albedo, metallic-roughness, and surface normals so assets can be relit and dropped into standard pipelines. Luce unifies geometry and PBR materials in a voxelized multimodal Gaussian cloud with dedicated primitives per modality, compresses it with a variational autoencoder into a material-aware latent space, and generates that latent from a single image with a rectified-flow transformer conditioned on multi-layer features from a pretrained image encoder, decoding to relightable Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K it sets a new state of the art for single-image-to-3D, improving FID by 28% over the strongest baseline, and on a new benchmark of AI-generated images it raises CLIP image-alignment from 0.8299 to 0.8519 while preserving fine details such as text, logos, and inscriptions.

Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs

Hao Luo, Yiting Yang, Wenyi Zhao, Man Jiang, Zhijun Lin, Ghulam Mohiuddin et al. Large-kernel convolutional neural networks (CNNs) achieve strong vision performance by expanding receptive fields, but their parameter growth hinders storage-efficient deployment on edge devices, and the authors show the bottleneck is not the depthwise convolutions that existing compression targets but the pointwise convolutions, which account for more than 87% of parameters in models like RepLKNet-31B. Channel Group-Shared (CGS) low-rank approximation is a Singular Value Decomposition (SVD)-style parameter-sharing scheme that shares the expensive down- and up-projection matrices across channel groups within a layer while giving each group its own cheap diagonal scaling matrix. Applied to RepLKNet, ConvNeXt, and SLaK, CGS strikes a favorable balance between competitive accuracy and substantially reduced storage, easing memory bandwidth pressure and model loading latency enough to make pre-trained large-kernel CNNs feasible on smartphones with 4-12 GB of RAM.
15 more specialized papers

Reinforcement Learning 15

AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL

Xiaolong Jin, Dingmin Wang, Vijay Lingam, Varun Kumar Trajectory-level rewards in multi-turn agent reinforcement learning give every step the same advantage, so the policy cannot tell which decisions caused success or failure, and existing self-distillation methods apply the same privileged information uniformly to all steps. AHEAD is a step-aware framework in which a teacher receives environment feedback on every step as a dense grounded signal and additionally receives LLM-generated corrective hints only on error steps, with minimal changes to standard GRPO. Across ALFWorld, WebShop, and search-based question answering at three model scales, it raises task success (+13.3 points on ALFWorld and +11.0 on WebShop at 7B over GRPO), reaches a given success rate in fewer training steps, and solves tasks within tighter interaction budgets than outcome-only RL and prior self-distillation baselines.

Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping

Yiwen Zhang, Xiaodong Yan, Zhenyu Huang, Deng Zhao, Liang Jiang, Qing Cui et al. Reinforcement learning from verifiable rewards (RLVR) for code generation is only as reliable as its test cases: insufficient coverage produces false positives that invite reward hacking and policy degradation. The RobustTests framework synthesizes test cases guided by near-correct faulty programs so that tests capture latent logical discrepancies, filters invalid and redundant tests with validator agents and behavioral feature clustering, and adds a stepwise dense reward based on pass rates to soften false negatives from hallucinated tests. Applied to augment CodeContests and used to fine-tune Qwen3-32B with RL on a moderately challenging problem subset, it yields an absolute 3% gain on LiveCodeBench over baseline methods.

Contrastive Branch Policy Optimization

Ying Wang, Changlin Qiu, Bang Lin, Linbo Jin, Wen Jiang, Zhe Sun et al. cross-listed Reinforcement learning with verifiable rewards (RLVR) lets language models learn multi-turn tool use, but its sparse outcome reward says nothing about which intermediate decision made a trajectory succeed, and prior branch-sampling methods conflate allocating a rollout budget with converting branch outcomes into token-level credit. Contrastive Branch Policy Optimization (CBPO) separates the two: generation entropy screens candidate branch positions across the whole response, path- and node-level decay spread a fixed budget so exploration does not collapse onto a few paths or adjacent tokens, and the reward variation within an exact-prefix group defines a Contrastive Branch Value that rescales continuation advantages without flipping their sign, with non-overlapping credit segments avoiding duplicated gradients when several nodes are selected on one trajectory. Requiring only outcome rewards and no process annotations, CBPO attains the highest macro-average accuracy across five math-reasoning and five knowledge-intensive search benchmarks at two model scales, outperforming state-of-the-art policy-optimization and branch-based methods.

FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision

Qiming Xie, Wenjie Zheng, Xiangqing Shen, Rui Xia cross-listed Outcome-only rewards in reinforcement learning with verifiable rewards can encourage hallucination, and existing process-level factual supervision aggregates factual signals coarsely and never assesses their reliability, creating what the authors call noisy factual credit assignment with both localization and reliability ambiguity. Fact-Aligned Reliability-Aware Credit Assignment (FARCA) aligns the granularity of fact verification with that of policy updates to localize credit at the token level, and introduces counterfactual evidence attribution, which measures how much a factual judgment depends on key evidence as a proxy for verification reliability and uses that to weight factual rewards and local advantages. Across multiple models and factual reasoning benchmarks, FARCA significantly improves factuality while preserving general reasoning capability.

On-policy Distillation with Verifiable Reward

Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang et al. cross-listed Reinforcement Learning with Verifiable Rewards (RLVR) gives sparse task-level feedback, while on-policy distillation (OPD) gives dense token-level guidance but ignores whether a trajectory is actually correct and cannot exceed the teacher; existing combinations rely on weighted mixing or heuristic switching with extra hyperparameters. OPDVR (On-policy Distillation with Verifiable Reward) merges the two without adding hyperparameters by reformulating the implicit reward of sampled-token OPD according to trajectory correctness and applying a ReLU gate so that correct trajectories receive non-negative rewards and incorrect ones non-positive rewards, aligning the distillation signal with task success while keeping the teacher's distributional guidance. This turns sampled-token OPD into a proper RLVR method that can be plugged into any policy gradient algorithm such as GRPO. On six reasoning benchmarks, OPDVR consistently outperforms standard OPD.

SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang Group-relative reinforcement learning for language model agents must wait for sibling rollouts of the same prompt before computing advantages, which is expensive for long, variable-length tool-use trajectories; Single-stream Policy Optimization (SPO) removes that dependency with a persistent prompt-level value estimate but whitens one advantage per trajectory before feeding a token-mean actor loss. The authors show that centering at the trajectory level generally does not center the token-weighted quantity the actor actually consumes, and fix the mismatch by standardizing terminal-outcome advantages under the action-token measure; they also organize prompt evidence by the policy event that generated it rather than by learner receipt order. In matched runs on ALFWorld at two model scales and on Math-TIR, SPO++ improves online learning efficiency over SPO, and a paired ablation identifies action-token-measure normalization as the strongest tested component.

Demystifying Reinforcement Learning Post-Training of Language Models

Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh et al. Reinforcement learning (RL) post-training now drives much of the reasoning, math, and coding ability in large language models, but its inner mechanics remain opaque to many practitioners. Working in a deliberately simplified and controlled setting, the authors take RL with Verifiable Rewards apart step by step, examining how outcomes depend on the base model's prior distribution, reward granularity, prompt diversity, and model scale, and using output-distribution entropy to contrast what pretraining, supervised fine-tuning, and RL each do to model certainty. Two concrete conclusions stand out: the much-discussed effect of spurious rewards turns out to depend on which prompt distribution is used for post-training, and RL only succeeds when the base model already assigns enough probability mass to the target behavior, which connects it directly to classical exploration. The write-up is framed as a primer for NLP researchers adding RL to their toolkit.

Bayesian Flow Networks for Offline Trajectory Planning

Ludvig Killingberg, Helge Langseth Offline reinforcement learning (RL) methods that synthesize trajectories with diffusion models are built around Gaussian noise and do not transfer cleanly to discrete planning tasks, which need a categorical formulation. BFN-RL replaces diffusion with Bayesian Flow Networks (BFNs), which iteratively evolve distribution parameters rather than noisy samples and therefore handle discrete and continuous trajectory spaces within one probabilistic framework; a categorical planner generates future state sequences and a learned inverse-dynamics model converts consecutive states into actions. Evaluations on discrete planning and continuous control tasks show effective trajectory generation across both categorical and continuous state spaces from a single model family.

ShuttleArena: Interpretable Self-Play in Physics-Based Badminton

Peize Ding Badminton couples shot selection and court recovery inseparably: the best recovery depends on how the opponent answers the shot, and the shot's value depends on whether the hitter can cover the reply. ShuttleArena is a physics-based singles self-play environment combining continuous shuttle flight, player interception, structured shot generation, and post-shot recovery, with a role-conditioned policy that makes a masked interception choice when receiving and a factorized action over shot azimuth, elevation, speed, and recovery target when hitting. Policies are trained on single rallies with Proximal Policy Optimization (PPO) self-play against a staged checkpoint opponent pool using sparse rally-outcome rewards and a factor-specific recovery update. Frozen-checkpoint play, tactical probes, and recovery ablations show competitive improvement alongside interpretable opponent-conditioned changes in shot geometry and recovery, and a human-data sanity check finds recognizable badminton-like structure shaped by the simulator's abstractions.

Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning

Srivalli Katkuri, Maxwell Kawada, Juan Wachs Vision-language models (VLMs) can supply preferences for preference-based reward learning (PbRL), but existing methods compare only two outcomes at a time through the Bradley-Terry (BT) model even though a VLM can rank many candidates. The authors instead fit reward models with the Plackett-Luce (PL) formulation on VLM-generated listwise rankings of K observations, allowing the ranking size to be tuned to the environment and feedback format. On Meta-World manipulation tasks, PL reward models train policies as effectively as pairwise BT, K-wise BT, and RL-VLM-F baselines, with at least one ranking size in {3,4,5} matching or beating the alternatives in every environment; the best PL configuration reaches an 86% mean final success rate and matches the oracle baseline on Drawer Open.

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

Justin Robert, Raheel Qader On-policy distillation trains a language model on its own generations while a teacher scores them token by token, combining dense imitation-learning supervision with on-policy sampling, but it needs a second, larger teacher model. On-Policy Self-Distillation (OPSD) removes that cost by making the model its own teacher, conditioned on privileged information the student will not have at test time such as a reference solution, a plan, or environment feedback; early results matched reinforcement learning accuracy with far fewer generated tokens, but the same asymmetry that creates the signal also biases it toward collapse, the progressive narrowing of reasoning paths the model can produce. The review, restricted to mathematical reasoning and reporting no new experiments, organizes the literature around three levers that govern collapse: how tokens are weighted, what privileged information the teacher is shown, and how the guidance decays over training, giving a shared vocabulary for phenomena named differently across papers and separating settled findings from disputed ones.
4 more specialized papers

Reasoning 9

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu et al. cross-listed Large language models (LLMs) are increasingly used to supply prior causal knowledge for structural causal discovery, but whether their direct-edge judgments and confidence can be trusted is unclear. The authors evaluate 12 instruction-tuned open-weight models on six benchmark causal graphs using five prompting strategies and four confidence sources (verbalized, logit-based, cross-prompt agreement, and cross-model agreement) under a language-only pairwise protocol. Models are strongly recall-dominant, predicting overly dense graphs; they misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, and over 80% of these false positives carry verbalized confidence of at least 80%. Logit-based confidence collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement calibrate better (though not significantly after Holm correction), leading to the conclusion that LLMs are best treated as externally validated soft causal priors rather than direct evidence of structure.

Recursive Agentic Reasoning

Shengxin Zhang, Xiaomin Wu, Xiyang Wu, Jing Xie Test-time reasoning techniques such as iterative refinement, decomposition, and repeated sampling are usually evaluated in isolation, which makes their gains hard to compare across models and pipelines. The authors unify them as three recursion operators over a reasoning trace — GROW (deepen a single path), PRUNE (decompose and recompose the problem), and BRANCH (sample alternative paths and select among them) — and evaluate all three against a single-pass chain-of-thought baseline under one harness with identical prompts, token budgets, and grading code across five benchmarks and three frontier models. BRANCH improves accuracy in all 14 model-benchmark settings by an average of 5.98 percentage points and is the best operator in 12, while GROW averages 2.18 points and degrades two settings and PRUNE gains only 0.94; much of BRANCH's advantage comes from recovering from budget-exhausted empty outputs, with gains correlating at r = 0.72 with the baseline truncation rate. The authors also show that unpaired evaluation and counting scoring-pipeline failures as model errors can reverse comparative conclusions, and argue for paired scoring as a standard protocol.

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Sophia Xiao Pu, Yumo Xu, Sailik Sengupta, Millennium Bismay, Ruixue Lian, James Gung et al. Inference-time methods that sample multiple reasoning trajectories treat each one as atomic, keeping or discarding it whole and thereby wasting good prefixes attached to degraded suffixes. Selective Regenerative Decoding (SRD) routes each candidate to discard, keep, or regenerate only the degraded part of its suffix while preserving the useful prefix of borderline candidates, without needing a larger target model. Under mild assumptions the authors prove a 1.28-to-1.36-fold sample-efficiency gain over rejection sampling with strictly higher expected trajectory quality, growing with pool size, and on MATH500, GPQA Diamond, HotpotQA, and AlpacaEval with several generator-reward model pairs, SRD matches Best-of-N accuracy with substantially fewer generated tokens and beats speculative rejection in low-compute regimes.

Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning

Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han et al. Long autoregressive reasoning traces are executed sequentially, creating severe latency on hard problems, and prior parallel-reasoning systems focus on Subtask Parallelism (decomposing a task into independent chunks) while overlooking Trial Parallelism, where multiple speculative attempts explore, verify, and aggregate competing hypotheses concurrently. An analysis of DeepSeek-V4 reasoning steps on HLE finds that Trial Parallelism accounts for 65.5% of parallelizable reasoning computation and grows more dominant on harder problems. Parason converts sequential traces into structured parallel trajectories via a context-free grammar, trains models with Parallelism-Aware Group Relative Policy Optimization (PA-GRPO), whose reward balances accuracy, latency, and both parallelism ratios, and executes the learned structure through tool calls at inference. On AIME24 and AIME25 it achieves roughly 1.7x average wall-clock acceleration while keeping accuracy competitive.

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain actually drives the answer is rarely tested, since general CoT-faithfulness probes ignore clinical cost and medical LLM evaluations treat the chain as a black box. The authors build a medical perturbation audit with a 30-operator battery that edits both the chain and the question using clinically motivated operators (severity reversal, negation flip, demographic swap, evidence ablation), paired with a joint analysis of chain updates and answer flips that classifies each model by its failure mode. Applied to 14 LLMs on four medical QA benchmarks, three independent tests converge: the Chain-Decoupling Rate, where the chain fails to register a clinically meaningful destructive edit and the answer does not flip, is 72.9% panel-wide, corrupting the chain leaves accuracy unchanged, and removing CoT prompting does not reduce accuracy. Two board-certified clinicians re-annotated 197 perturbed questions and found 98.5% leave the gold answer defensible, and the pattern holds across medical and reasoning fine-tuning, model scale, and closed-source models where only answer-side signals are observable.

DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation

Yutong Chen, Guangfu Guo, Zhichao Xu, Kunpeng Liu On-policy self-distillation (OPSD) supervises a student model densely using a privileged copy of itself, but that privileged teacher stays fixed even as the student's distribution and output style shift during training. DualOPSD alternates asymmetrically between the two: the student first learns from the privileged teacher, then the teacher moves toward the updated student distribution on the same student trajectory, so later supervision tracks the learner without requiring another rollout. On Qwen3-8B in non-thinking mode, DualOPSD improves avg@12 over OPSD by 23.61, 13.89, and 10.00 points on AIME 2024, AIME 2025, and HMMT 2025; results at 1.7B and 4B show the accuracy gain depends on model scale, truncation drops at all three scales, and a 4B diagnostic shows lower KL divergence in both directions between teacher and student.

Prefix Sliding for efficient test-time scaling

Niklas Muennighoff, Zhengyang Wang, Zeyi Chen, Weijia Shi, Binyuan Hui, John Yang et al. Long reasoning traces make test-time scaling expensive because full attention keeps every intermediate token in memory, even though most of those tokens lose importance as reasoning proceeds. Prefix Sliding discards tokens that fall outside the prompt prefix (which holds the key instructions and tool definitions) and a window of the last few thousand tokens, capping total memory regardless of how long the model reasons. Applied to existing models without any training, it makes inference about 3x faster while maintaining performance, and training with it via reinforcement learning enables reasoning traces beyond a hundred thousand tokens; ablations show it outperforms both summarizing intermediate tokens and a vanilla sliding window.
2 more specialized papers

Robotics 9

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski cross-listed Vision-Language-Action (VLA) models pretrained on large robot datasets degrade when adapted to new tasks with only a few demonstrations, and existing retrieval methods that reuse prior data match on visual similarity, state-action features, or whole-task language, missing that long-horizon tasks share reusable skills even when no full-task match exists. Hierarchical Skill Retrieval (HSR) decomposes a target task into candidate skill sequences, scores each plan by semantic plausibility and skill reliability estimated from the prior dataset, retrieves demonstrations through subtask-level language matching followed by behavior-feature reranking, and adapts the policy with a two-stage pipeline that separates general skill acquisition from task-specific finetuning. On the LIBERO benchmark and several real-robot manipulation tasks, HSR raises average success rate over the strongest baseline by 10.3% in simulation and 21.3% on real hardware.

PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee, Youngjae Yu cross-listed Vision-language-action (VLA) models typically inherit pretrained multimodal representations without using the underlying model's long-context capacity as episode memory, leaving that role to purpose-built history mechanisms. PonderPounce instead pairs Ponder, a System-2 multimodal large language model (MLLM) that accumulates observations, demonstrations, and prior cognition in its native causal context, with Pounce, a System-1 VLA that receives the current observation, instruction, and proprioception plus, asynchronously, only the newest cognition token and its age; the two are trained jointly end to end with no separate memory module or bridge pretraining, and optimized serving supports 20Hz action playback. On RoboMME with base-scale data it reaches 60.83% with a 9B model versus 44.51% for FrameSamp+Modul and 17.93% for current-observation π0.5, rising to 75.54% with 9x data, and on RoboCasa-DC it reaches 12.5% from action supervision alone versus 11.6% for the strongest demonstration-conditioned baseline.

LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation

Jin Lou, Jingxuan Zhu, Andong Chen, Xupeng Wang, Yuan Xu, Yuexuan Li et al. cross-listed Generalist vision-language-action (VLA) policies learn long-horizon behavior through short-horizon action prediction, forcing a single action target to implicitly absorb task progress, intermediate intent, and local reliability while keeping those states hidden during execution. LM-X organizes prediction across task, event, and motor scales by emitting three supervised signals that directly condition action generation: return-to-go for visible task progress, event-to-go for the next semantic transition, and heteroscedastic action flow for local reliability via propagated variance. A five-task pretraining gate showed the full model beating an action-only backbone by 16.0 points before a 20-day run on 64 NVIDIA B200 GPUs over more than 20,000 hours of real-robot trajectories, after which LM-X reaches 74.1% on 50 randomized-hard RoboTwin2.0 tasks versus 55.4% for GR00T N1.7 and 68.6% versus 50.7% on seven real-robot tasks, with the progress and variance signals tracking semantic progress, regression, and hesitation.

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback

Jianbo Zhou, Boyuan Zhao, Yuzheng Zhang, Yiyang Chen, Wenxin Chen, Qiuyue Li et al. cross-listed Chunk-based vision-language-action models commit to an entire action chunk from observations gathered before execution, so tactile conditioning goes stale exactly while contact states are evolving. TacForcing replaces the standard action expert with a streaming one that keeps generating actions conditioned on tactile readings acquired during execution, avoiding the separate high-frequency reactive controller that earlier tactile-reactive designs require, and adds Execution-Aware Tactile Attention (EATA) so tactile signals only condition actions close to being executed. Average success rates were 65% across six simulated UniVTAC tasks and 69% on three real-world contact-rich manipulation tasks, ahead of the compared baselines in both settings.

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

Lehong Wu, Yuxiao Qu, Zheyuan Hu, Ivan Zhang, Limin Wei, Zackory Erickson et al. cross-listed Reasoning in natural language lets foundation models spend more test-time compute on hard problems, but whether it helps long-horizon robotic manipulation, where a policy must track partial progress, reason about object relations, recover from mistakes, and steer noisy low-level controllers, has been unclear. R^3 is a post-training recipe that turns off-the-shelf vision-language models (VLMs) into robotic reasoners: it first mid-trains the VLM on expert-generated reasoning traces to establish the reasoning style, then improves the reasoner with single-step rubric-based reinforcement learning from offline action data, so that free-form language reasoning becomes test-time guidance for action rather than auxiliary supervision. On Language Table and a simulated bimanual grocery-packing testbed, R^3 significantly outperforms instruction-only imitation learning baselines and improves exploration and generalization to unseen tasks, suggesting free-form language reasoning can act as a test-time compute mechanism for steering low-level policies.
4 more specialized papers