Wednesday, August 26, 2026

381 papers cs.AI · cs.LG · cs.CL ← 2026-08-252026-08-27 →

Jul Aug Sep

Highlights

AI Agents Push Humans Out of the Loop

Highlight Safety & Alignment Margaret Mitchell, Avijit Ghosh, Samir Passi Human oversight is the standard proposed safeguard for increasingly autonomous AI agents, but the authors argue that current agent design impedes effective oversight and that extended use of AI systems degrades the very cognitive capacities oversight requires. Drawing on automation research and human-computer interaction, the position paper outlines design-level affordances and organizational protocols intended to support overseers in exercising critical judgement and to counteract the skill atrophy that arises from prolonged reliance on automation. The central claim is that supporting the situated goals and cognitive requirements of human overseers should be treated as equal in priority to agent capability, since without it agent systems passively incentivize the erosion of the human skills they depend on.

Human-in-the-loop oversight is the standard prescription for managing AI agent risk, but this position paper from Hugging Face and Data & Society argues that current agent design both impedes effective oversight and actively erodes the cognitive capacities it depends on. The proposed remedy is "cognitive scaffolding": treating overseers' attention, situational awareness, and domain skill as first-class design constraints, on par with agent capability.

  • Agents produce dense, fast-moving streams of chain-of-thought, tool calls, plans, and inter-module exchanges that exceed what a reviewer can track, pushing users into repeated approval prompts and approval fatigue, while vendor oversight frameworks from Anthropic, Salesforce, and AWS specify human roles without accounting for how attention behaves under sustained engagement.
  • Drawing on Bainbridge's 1983 "irony of automation", the authors compile evidence that extended AI use causes deskilling, reduced vigilance, and automation and anchoring bias, citing an EEG study showing significantly decreased brain connectivity in LLM-assisted essay writers and a multicentre study of endoscopist deskilling after AI-assisted colonoscopy, with the sharpest effects on novices and experts working slightly outside their expertise.
  • Degraded oversight creates an alignment feedback loop: tired overseers approve fluent rationales quickly, those approvals become RLHF and evaluation signal, and systems can drift toward confident summaries and skimmable plans that further disincentivize scrutiny, making the human rater an exploitable part of the reward channel.
  • The solution inventory splits into developer-side affordances (pre-commitment before seeing agent output, delay-and-choice, reasoning probes, action gating, bounded autonomy, batch diff review, and behavioral monitoring via review-time, override-rate, and evidence-seeking signatures plus known-answer canaries) and deployer-side protocols (unassisted skill-maintenance exercises, critical-evaluation training, enforced breaks, rotations, and separating the overseer from whoever benefits from approvals).
  • As a position paper it presents no new experiments, the inventory is explicitly non-exhaustive with overlapping categories, and the authors concede early evidence that users disprefer systems that reduce overreliance, leaving a real tension between satisfaction and oversight quality.

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

Highlight Agents Stephen Chung, Wenyu Du, William J. Wesley The Station is an open-world multi-agent environment in which AI agents from different model families pursue a shared mathematical research goal with no central coordinator or scripted pipeline, choosing their own directions, running experiments, collaborating, and building a shared literature. Applied to 12 construction problems from the AlphaEvolve catalogue plus two case studies, the system produced results novel relative to prior literature on five problems, including a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem, along with new infinite families for Book Ramsey numbers. The agents also produced theorems and analyses explaining why the constructions work, and all raw dialogues, proofs, and verification code are released.

Multi-agent systems built for scientific discovery usually run agents as fixed tools inside a scripted pipeline with a central coordinator; the Station instead drops a handful of heterogeneous LLM agents into an open-world environment with only a research goal, letting them pick directions, run experiments, message each other, and publish into a shared internal literature that later agents read and cite. Applied to 12 construction problems from the AlphaEvolve catalogue plus two case studies, this produced results novel relative to prior literature on five problems, along with theorems and explicit algebraic constructions explaining why the numerical results work.

  • Each Station instance runs six agents (two each of GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro) for roughly 1,000–2,000 ticks (one to two weeks), with rooms for code execution, private mail, public forums, a Stack-Exchange-style Question Room, and an Archive Room where papers that pass automated review accumulate across agent generations; agents have limited lifetimes and are automatically replaced, and they get coding assistants and periodic "holidays" with open-ended prompts.
  • Headline novel results: a new infinite family of finite-field Kakeya sets in F_p^3 for p ≡ 3 (mod 4) of size (2p³+7p²+3)/8 (saving (p−3)/4 points over AlphaEvolve's family) plus a 53-point Kakeya set in F_3^5 beating the prior 63; three exact 604-point kissing configurations in dimension 11 (two apparently new isometry classes, versus AlphaEvolve's 593); a discretized Kakeya needle area of 0.107067 at n=128 (6.74% below AlphaEvolve); a sign-uncertainty upper bound of 0.3089 (from 0.321591); and a new lower bound of 0.380552 for Erdős's minimum-overlap constant that closes ~82% of the published gap, even though the task asked for an upper bound.
  • Agents repeatedly went beyond the score: they proved the classical D_11 norm-four construction caps at 582 points before pivoting to a new core, proved the double-root Laguerre family can't beat 0.315305 and then left it, proved exact optima C_T(3)=5/18 and C_T(4)=1/4 for the needle problem, showed the non-tangential Hardy–Littlewood constant equals 2 for 1/3 ≤ α < 1, and in the prime-number-theorem task produced a certified-for-all-x score of 0.980681 while rejecting higher but hackable sampled scores.
  • The two extra case studies yielded two proved novel infinite families for Book Ramsey numbers (resolving 28 previously open cases among n ≤ 200 together with an expert-derived third family) and an independent, web-free reconstruction of the degree-seven Jacobian Conjecture counterexample within a day; more than half of the findings involved cross-model collaboration or building on earlier agents' internal papers, and all raw dialogues, proofs, and verification code are released.
  • Limitations are real: the Station underperformed AlphaEvolve on peak and flat autoconvolution, which reward large-scale heuristic optimization of irregular objects rather than theory-guided search; it found no competitive uniform-in-n needle construction, its Kakeya formulas in dimensions 4 and 5 are weaker than known results, and adjacent kissing runs stopped at 840 in d=12 (one below the frontier) and merely matched 1154 in d=13, so the approach favors problems where structure can guide the search and interpretable outputs are valued.

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

Highlight Large Language Models Farhana Amin, Sabiha Afroz, Mona Moghadampanah, Dimitrios S. Nikolopoulos Masked diffusion language models (dLLMs) can denoise many tokens at once, but serving systems for them have been built without measuring how they behave under real concurrent load. The authors characterize serving of LLaDA-8B-Instruct with a Discrete Diffusion Forcing (D2F) LoRA adapter on a single NVIDIA H200 using GSM8K and HumanEval, finding that request difficulty falls into 11 discrete denoising-step levels that no tested signal predicts in advance (best R2 = 0.150), that generation budgets under 320 tokens hide the latency spread, and that only 24% of single-request wall-clock time is GPU computation, with the rest being CPU-side dispatch overhead. Batching mainly amortizes that overhead, giving a 16.0x throughput gain at batch size 16 over per-request dispatch, and the paper derives a batch-timeout rule for synchronized batching under Poisson arrivals, arguing that dLLM serving requires parallelism at each denoising step.

Masked diffusion LLMs promise faster generation than autoregressive models by denoising many tokens per forward pass, but nobody had measured how they behave under real concurrent serving load before building schedulers for them. This paper characterizes LLaDA-8B-Instruct with a D2F LoRA adapter on a single H200, and finds that dLLM serving is bottlenecked by CPU-side dispatch rather than GPU compute, so parallelism has to live at the level of each denoising step instead of each request.

  • Request difficulty is discrete, not continuous: across 100 GSM8K requests, denoising-step counts land on exactly 11 fixed levels (178 + 29k), a direct consequence of D2F's block-addition checkpoint rule, and step count predicts end-to-end latency almost perfectly (r = 0.9997), yet none of four cheap early signals (confidence fraction, mean confidence, entropy, prompt length) can predict the level at admission time (best R² = 0.150, far below the authors' 0.5 usefulness floor).
  • Profiling a single request shows only 24.4% of the 10.28 s wall-clock is GPU kernel time; the other 75.6% is host-side overhead in the per-block control loop (mask construction, threshold checks, synchronization), which runs once per step regardless of batch size, so batching helps mainly by amortizing that fixed cost rather than by improving GPU utilization.
  • Sharing one forward pass per denoising step across a synchronized batch scales throughput near-linearly, reaching 1.939 req/s at batch size 16, a 16.0× gain over per-request dispatch (0.121 req/s), while capturing the forward call with CUDA Graphs cuts its dispatch time by 48.6% but would only remove about 28.5% of total dispatch overhead, since the rest lives in Python control flow outside the forward call.
  • Short generation budgets hide serving variance: at 128 tokens every GSM8K request is truncated and CV is 0.067, but at the ~320-token natural completion length (512-token budget) CV rises to 0.283 with 76% accuracy, and HumanEval is a harder serving workload still (CV 0.343–0.346, max-to-min latency ratio ~4×); block size {16–128} barely affects latency (within-request CV under 0.01), a ~36× smaller spread than across requests on both tasks.
  • Limitations are substantial: the batch-timeout stability rule T_min = S − 1/λ and the latency-optimal utilization of ρ ≈ 0.70 come from a Monte Carlo over only n = 8 measured batch service times (a later 32-batch run shifted estimated capacity by 24%), quality under batching is argued structurally but not measured beyond single-request 74–76% accuracy, the tier constants are tied to one threshold setting, and the per-request-dispatch baseline is weaker than a slot-based or vLLM-style shared-pass design, which the authors leave as future work.

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Highlight HF pick · 43▲Multimodal Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu Universal multimodal embeddings map text, images, video, and visual documents into one shared space for retrieval, recommendation, classification, and agentic systems. WeMM-Embedding is a family of 2B, 4B, and 9B models that accept arbitrarily interleaved multimodal inputs with flexible output dimensions, trained through a large-scale multimodal alignment stage followed by refinement on curated data with fine-grained relevance supervision and cross-scale knowledge transfer. The 9B model sets a new state-of-the-art overall score of 80.6 on MMEB-v2, the 2B model already beats the previously leading 8B open-source baseline, and the family shows gains on a 26-task in-house benchmark and 14 online A/B tests across WeChat services, with weights and code released.

Universal multimodal embedding models still lag on fine-grained relevance once broad cross-modal alignment is in place, and most open-source options trade off scale against generality. WeMM-Embedding is a family of 2B/4B/9B models built on Qwen3.5 backbones that first trains on several hundred million heterogeneous pairs, then refines on a curated corpus roughly one tenth the size using hard negatives, reranker-derived rankings, and cross-scale distillation from the 9B teacher.

  • Every task is cast as a source–target pair with an optional instruction, hard-negative set, and graded relevance score; inputs are encoded with a trailing <embedding> token under causal attention (allowing multiple embedding tokens per sequence, e.g. video-only and video-plus-ASR from one pass), with Matryoshka training over nested dimensions and a duplicate-aware mask that drops near-identical in-batch negatives.
  • Stage 2 draws on Semantic-ID-guided resampling (a three-level residual k-means quantizer over intermediate-checkpoint embeddings) to downweight overrepresented semantics, MLLM-based quality filtering and caption correction, and a bidirectional KL distillation of batch-wise teacher similarity distributions, which the authors say contributes most to the compact variants' gains.
  • On MMEB-v2 (78 datasets) the 2B model scores 77.9, edging past Qwen3-VL-Embedding-8B (77.8), while the 9B reaches 80.6, first on the leaderboard ahead of proprietary entries such as Octen-VL-Large (80.1) and DME-Large (80.2); on MMEB-v3 the 9B scores 59.5 V3-All, and on a 12-benchmark cross-modal retrieval suite the 2B averages 79.8 versus 79.5 for Gemini Embedding 2.
  • Ablations show task-consistent batching is the single biggest Stage-1 factor (−3.4 points when replaced with mixed sampling), Stage-2 strategies cumulatively add 2.2 points for the 2B model, and 256-dimensional embeddings retain 98.7% of 2048-dimensional performance on image and video tasks, with visual-document and retrieval tasks degrading fastest under truncation.
  • The models do not accept audio (scored as zero on the 11 MMEB-v3 audio tasks), reranker supervision only helped on a limited subset of tasks so was applied selectively, the 9B relies on merging multiple Stage-2 variants rather than a teacher, and the 26-task in-house benchmark and 14 A/B tests are reported only against Qwen3-VL-Embedding-2B with no released details.

On-policy Distillation with Verifiable Reward

Highlight HF pick · 7▲Reinforcement Learning Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang et al. Reinforcement Learning with Verifiable Rewards (RLVR) gives sparse task-level feedback while on-policy distillation (OPD) gives dense token-level guidance but ignores whether a trajectory is actually correct, capping the student at the teacher's ability; existing combinations rely on weighted mixtures or heuristic switching that add hyperparameters. On-policy Distillation with Verifiable Reward (OPDVR) reformulates the implicit reward of sampled-token OPD based on trajectory correctness and applies a ReLU gate so that correct trajectories receive non-negative rewards and incorrect ones non-positive rewards, aligning the distillation signal with task success while keeping the teacher's distributional guidance and adding no hyperparameters. This modification also turns sampled-token OPD into a proper RLVR method that plugs into any policy gradient algorithm such as GRPO, and OPDVR consistently outperforms standard OPD across six reasoning benchmarks; code is released.

Sampled-token on-policy distillation (OPD) gives dense token-level guidance but has an implicit reward whose sign depends only on the teacher-student probability ratio, so it can penalize tokens in correct trajectories and reward tokens in incorrect ones. OPDVR fixes this with a single ReLU gate: on correct trajectories the reward is max(0, log(π_T/π_θ)), on incorrect ones it is -max(0, log(π_θ/π_T)), which turns OPD into a proper verifiable-reward RL method with no added hyperparameters.

  • The gate acts as a conditional mask that zeroes exactly the "conflicting" tokens (student more confident than teacher on a correct trajectory, teacher more confident on an incorrect one); the verifier sets the update direction while the teacher's log-ratio still sets the magnitude, and Appendix A shows this removes precisely the component of the OPD gradient that is anti-aligned with the RLVR gradient.
  • Because the reformulation exposes a verifier reward, the sign of a GRPO group-relative advantage can replace the binary correctness label, yielding GRPD (Group Relative Policy Distillation) that plugs into GRPO/DAPO/PPO-style pipelines unchanged.
  • In same-architecture distillation (Qwen3-4B student from a GRPO-trained Qwen3-4B teacher on DeepMath), OPDVR reaches 49.1 average avg@16 across six benchmarks versus 47.8 for sampled-token OPD, and 36.9 on AIME24, above the teacher's 36.0; cross-architecture (Qwen3-1.7B-Base from Qwen3-4B-Base-RL on DAPO-Math-17k) it scores 22.8 vs 20.9, with a +5.5 point gain on AMC.
  • GRPD beats both GRPO (44.8) and OPD (48.4) with 49.4 average, including +6.5 on AIME24 and +10.9 on AIME25 over GRPO, while an inverse-gated ablation falls below vanilla OPD on all six benchmarks (44.6 avg), and the gate consistently masks roughly 40-50% of tokens throughout training.
  • Limitations: evaluation covers only math reasoning with Qwen3 models and a single teacher per setting, the same-architecture MATH500 result slightly trails OPD (84.7 vs 85.5), and the gains, while consistent, are modest at 1-2 average points with no reported variance across seeds.

Meta$^n$: Recursive Self-Improvement through Emergent Depth

Highlight HF pick · 1▲Agents Zae Myung Kim, Young-Jun Lee, Seungyeon Jwa, Dongyeop Kang Self-improving LLM agents typically refine their answers rather than the process that produces them, and systems that edit their own machinery must hold part of it fixed to stay stable, capping realized meta-depth at roughly two. Meta^n instead keeps a fixed meta-operation Ω and recurses on its input: Ω reads the traces and code of the solver stack below and writes the next layer as a strategic pre-process plus a library of callable helpers, so depth is set by convergence rather than in advance, and an evolutionary archive searches over layer chains. Across two backbones it outperforms prior self-improving agents on all eight benchmark families, and on ARC-AGI-2 it is the only method to score above zero. Ablations attribute most of the recursion gain to the conditioning each layer passes to the next, with distinct layer roles emerging at depth without being prescribed by any prompt.

Self-improving LLM agents that edit their own code must freeze some driver component to stay stable, capping realized meta-depth at roughly two. Meta^n inverts this: a single fixed meta-operation Ω is applied recursively to its own outputs, reading the solver stack's execution traces plus the code that produced them and emitting a new layer (a strategic pre-process plus a library of callable helpers), so depth is set by convergence rather than fixed in advance.

  • Each depth-d layer wraps the solver below it; Ω sees the previous depth's traces and the full code stack [C_2, …, C_{d-1}], which lets it attribute regressions to specific prior directives and roll them back (for example, on LawBench a depth-3 "Exhaustive Legal Analysis" directive dropped the score from 0.807 to 0.773 and depth 4 restored it to 0.833 while keeping the useful helper), and an evolutionary archive searches over layer chains with per-task best selection.
  • Across eight benchmark families and two backbones (Gemma 4 31B-IT, GPT-5.2), the agentic variant leads OpenEvolve and Gödel Agent on every family on at least one estimator; on CO-Bench under GPT-5.2 archive-best reaches 0.870 vs. 0.702 for OpenEvolve, and on ARC-AGI-2 it is the only system above zero (dev 0.331 vs. 0.003 and 0.054), while prompt-rewrite benchmarks like Symptom2Disease show only a few points of lead within seed noise.
  • Removing recursion (a depth-1 baseline with orchestration unchanged) costs −0.131 on Gemma CO-Bench, −0.080 on GPT-5.2 CO-Bench, and −0.158 on GPT-5.2 AlphaEvolve Math; ablating further shows that the plain inter-layer context string accounts for ~72% of the recursion gain, the code-library channel ~15%, and the search machinery ~13%.
  • Distinct layer roles emerge without any prompt prescribing them: depth 2 emits generic primitives (simulated_annealing propagates to 15 of 36 CO-Bench winners), depth 3 specializes and interferes (41% of chain-task pairs regress), and rollback appears only from depth 3 onward, with depth-≥4 candidates still winning 31% of CO-Bench tasks even though mean per-depth score peaks at depth 3.
  • Limitations include using the same model at every layer (leaving the stronger-Ω-over-weaker-base case untested), 4–10× higher token cost in agentic mode, the Gödel Agent baseline needing a corrected per-task configuration to score at all, and two benchmarks where the method adds little: AlgoTune, where extra context over-constrains an already-optimized kernel, and SWE-Bench, where Ω never activates because the seed is already strong.

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Highlight HF pick · 2▲Agents Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou et al. Search agents trained with outcome rewards learn when and how to retrieve evidence, but terminal rewards cannot localize intermediate mistakes or redirect a trajectory before errors compound. CAFE (Coupled Agent–Feedback Evolution) has a single shared-parameter model alternate between search-agent and critic roles: it bootstraps feedback-conditioned recovery from the base agent's own failures, then couples online RL, where a prompt-level call-versus-skip success gap shapes the return for requesting feedback and feedback-aware advantage shaping reweights token advantages before and after feedback, with offline preference optimization that learns feedback from matched successful and failed rollouts. Across seven agentic search benchmarks CAFE outperforms the evaluated RL-based search agents on average and retains its gains on all six out-of-domain benchmarks while reducing answer-level hallucinations, and one-sided ablations show that improving only the agent or only the critic plateaus whereas alternating the two updates keeps improving.

Outcome-supervised search agents learn when and how to retrieve from a terminal reward alone, which can neither localize where a long trajectory went wrong nor redirect it before an early error compounds. CAFE (Coupled Agent–Feedback Evolution) makes corrective feedback a learned, in-trajectory intervention: one shared-parameter model alternates between a search-agent role that learns when to emit a <request_feedback> action and a critic role that learns to answer it, with the two capabilities updated in alternation so the critic keeps tracking the agent's shifting failure distribution.

  • Training is bootstrapped from the base agent's own failed rollouts: a teacher (Kimi-K2.5) marks the earliest erroneous turn, the flawed prefix is kept, a feedback request plus teacher-written correction and continuation is inserted, and only repaired trajectories that reach the correct answer are used for SFT, so the model learns to request and use feedback at states its own policy actually visits.
  • Online, GRPO is augmented with a comparative feedback estimate (CFE) that adds the prompt-level success gap between feedback-requesting and feedback-skipping rollouts to the return of requesting rollouts (with a penalty for repeat requests), while feedback-aware advantage shaping subtracts from pre-request tokens and adds to post-feedback tokens so a rescued success credits the recovery rather than the prefix that went off course.
  • Offline, rollout-derived preference optimization (RDPO) mines prefix-matched pairs of successful and failed feedback-requesting rollouts for the same prompt and DPO-trains the shared checkpoint to prefer the successful feedback, alternating with online RL in five rounds of 100 RL steps each; RDPO beats a positive-only SFT variant on the same mined data.
  • On Qwen2.5-7B-Instruct across seven SearchQA benchmarks, CAFE averages 52.5 EM / 60.7 F1, 2.1 EM and 1.3 F1 above the strongest RL baseline IGPO, improves over plain GRPO on every benchmark including all six out-of-domain sets (gains concentrate on multi-hop: 7.4 EM over Search-R1 versus 1.3 EM on single-hop), and lowers the answer-level hallucination rate from 29.9% for the base model and 17.6% after GRPO to 12.6%; on 2Wiki, agent-only optimization with a frozen critic plateaus at 83.6 (mean of EM and F1) while alternating updates reach 86.6, and cross-play shows each agent performs best with the critic from its own iteration rather than the strongest critic overall.
  • Caveats: training uses 2Wiki as the sole in-domain set with 7B and 3B backbones only, the bootstrap depends on a strong external teacher, CFE on its own adds only about 1 EM (most of the online gain comes from advantage shaping), Gemini-2.5-Flash still edges CAFE on average F1 (60.9 versus 60.7), and the harder BrowseComp-Plus result is reported only in the appendix without numbers in the main text.

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Highlight Large Language Models Zihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen, Zhiqiang Zhang et al. Language model pretraining runs with substantially different learning rates (LRs) and parameter norms are shown to follow nearly identical loss trajectories whenever their effective learning rate (ELR), the ratio of LR to parameter norm, is matched, a phenomenon termed ELR collapse. Across optimizers, architectures, datasets, and model scales, the mean collapse error between ELR-matched runs is typically a few times 10^-3, below the seed-to-seed variation measured in a representative configuration, with normalization design and the timescale of LR-norm variation identified as the main determinants of collapse precision. Controlled interventions show that weight decay and Hyperball shape loss dynamics mainly through the ELR schedules they induce, and replacing LR with ELR lets a fitted functional scaling law (FSL) transfer across norm-control methods and explain the delayed acceleration that norm control frequently produces.

Norm control in LLM pretraining — weight decay, explicit norm constraints like Hyperball — reshapes loss curves, raising the question of whether it is an independent knob or just another way to set the learning rate. The paper argues the latter: learning rate and parameter norm affect loss almost entirely through their ratio, the effective learning rate ELR = η_k / ‖W_k‖_F, so runs with very different LR and norm schedules produce nearly identical loss trajectories whenever their ELR schedules match.

  • The core experiment prescribes a common warmup–stable–decay ELR schedule and realizes it with four different LR shapes (WSD, linear-up, linear-down, sinusoidal) by enforcing a compensating parameter-norm schedule at every step; on Llama-124M and Qwen3-MoE-586M trained with AdamW, the loss curves collapse with mean discrepancies of 1.8–4.1 × 10⁻³, well below the 1.1–1.6 × 10⁻² variation seen from merely changing the initialization or data-order seed.
  • Across 26 matched comparisons spanning dense, MoE, and Kimi Delta Attention models from 100M to 1B parameters, FineWeb/C4/OpenWebText, and AdamW/Muon/Signum, the median collapse error is 2.5 × 10⁻³ and every case stays under 5 × 10⁻³, with similar precision for ViTs on ImageNet.
  • Collapse is precise but conditional: removing QK-Norm raises the error from 2.3 to 5.2 × 10⁻³, additionally fixing the RMSNorm gains pushes it to 1.84 × 10⁻² (counterintuitively, since fixed gains make the network more scale-invariant), and rapidly oscillating LR–norm realizations of the same ELR degrade it from 2.8 to 7.5 × 10⁻³, so the phenomenon cannot be explained by static scale symmetry and its mechanism remains open.
  • Practical norm control works the same way: adapting only the LR of a no-weight-decay run to match the ELR of a λ=0.1 AdamW run recovers its loss to 4.8 × 10⁻³, and matching a weight-decayed Muon run to a Hyperball-constrained one gives 1.2 × 10⁻³ after a brief onset transient; swapping LR for ELR in the functional scaling law of Li et al. also lets a fit on non-Hyperball runs transfer to unseen Hyperball runs (RMSE 0.021 vs 0.251, roughly 12× lower).
  • The ELR view explains "delayed acceleration" — weight decay initially loses to the unregularized baseline but wins late because it prevents norm growth from prematurely shrinking the ELR — and a hand-designed norm schedule that grows the norm faster late in training beats the weight-decay baseline; the authors caution that matched loss trajectories do not imply matched parameters, representations, or downstream performance.

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Highlight HF pick · 1▲Multimodal Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thadd\"aus Wiedemer, Christoph Schuhmann et al. Open video data for multimodal pretraining has lagged far behind the scale of web image-text corpora. LAION-BVD collects 1.3B platform-specific video URLs from CommonCrawl and downloads 80M videos totaling 10 million hours, then applies content-aware scene detection to extract clips and synthetically generates video and audio captions for them, targeting joint pretraining across video, audio, and image modalities. Models trained on the data are competitive on standard video-text and audio-text benchmarks with consistent gains as training or model scale grows, and scene-changing frames extracted from the videos form an image-text source whose visual distribution differs from web image corpora and yields strong image-text retrieval performance. The dataset is released to the research community.

Open video data at pre-training scale is scarce compared to image-text corpora, so LAION-BVD mines CommonCrawl for platform-specific video URLs and turns the downloaded videos into clip-level video, audio, and image training data with synthetic captions.

  • The pipeline collects 1.3B video URLs from CommonCrawl, downloads 80M videos totaling 10 million hours, then uses content-aware scene detection to cut them into clips, each paired with synthetically generated video and audio captions.
  • Models trained on the clip-caption pairs reach competitive results on standard video-text and audio-text benchmarks, with performance improving consistently as training data or model size scales.
  • As a side product, scene-changing frames extracted from the videos serve as an image-text source with a visual distribution distinct from typical web image datasets, and models trained on them achieve strong image-text retrieval performance.
  • The dataset is fully released, expanding open access to multimodal video at a scale not previously available to the research community.
  • The abstract reports no concrete benchmark scores or baseline comparisons, all captions are synthetic rather than human-written, and the reliance on platform-specific URLs raises the usual questions about link rot, licensing, and content filtering.

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Highlight HF pick · 15▲Agents Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang et al. Recursive self-improvement (RSI) is difficult in long-horizon tasks because ever-growing interaction histories obscure the current task state and cause skills to be invoked at the wrong moments. Recuris pairs a Working Memory that tracks task progress with an Experiential Memory of skills, so skill selection is driven by current needs rather than the full history, and this coupling turns each execution into structured evidence that localizes failures to specific memory components; a fixed Meta-Agent then converts that evidence into localized, validation-gated updates to Skill Memory, closing a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, it improves task success in 35 of 37 completed model-benchmark pairs, adding +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5 on tau-bench (taking Opus 5 to 87.9%) and +16.6/+13.5 points to Qwen3.6-27B/35B on SkillFlow. The advantage widens with interaction horizon, reaching +32.2 points on the longest tasks, and common long-horizon failure modes drop by up to 80%.

Long-horizon LLM agents accumulate experience but lose track of unresolved goals as interaction histories grow, so stored skills get invoked against a stale or noisy picture of the task. Recuris couples a compact, verified Working Memory that tracks goal status with an Experiential Memory of reusable skills, using the current state rather than the full history to decide when a skill is needed, and turning the resulting structured traces into component-level repairs of the memory layer across tasks.

  • Within a task, skill retrieval fires at execution events (e.g. a drafted state-changing tool call is intercepted, the matching skill is injected, and the call is redrafted), while a checker set commits a goal as done only when the tool receipt supports it, so verbal confirmation never counts as completion.
  • Across tasks, a fixed Claude Code-based Meta-Agent reads the structured trace, attributes each failure to one of four memory components (skills, working-memory spec, invocation policy, checkers), patches only those, and admits the candidate only if it fixes the source task without regressing a held-out dev set; the base LLM is never touched.
  • On τ²-Retail, a memory evolved once on mid-sized doubao-seed-2-0-pro lifts GPT-5.6 Sol by +17.8 (58.3→76.1) and Claude Opus 5 by +15.6 to 87.9%, and on SkillFlow lifts Qwen3.6-27B from 42.2 to 58.7; gains land in 35 of 37 model–benchmark pairs and widen with task length, reaching +32.2 on the longest quartile.
  • Ablations show the working state carries the effect (WM-only +23.9 vs EM-only +2.0 on Retail), and injecting the whole skill library and letting the model choose scores 18 points below state-grounded invocation while costing 46% more per success; structured traces also raise fault-localization accuracy from 13.0% to 64.8%.
  • Limitations: τ²-Airline (50 tasks) intervals all include zero, Terminal-Bench 2.1 has no shared structure so cross-task evolution admitted no patch in 13 runs, evolution is non-monotone (one lineage gave back most of its gain in round four), and the per-benchmark memory is hand-built once on a single deployment model.

Applications 87

StateTune: Transforming LLM-Assisted EDA Flow Tuning into a Stateful, Closed-Loop Process

Kunlong Li, Shangshang Yao, Su Zheng, Lingli Wang cross-listed Tuning electronic design automation (EDA) flow parameters strongly affects quality-of-results (QoR), but the parameter space is large and tightly coupled, full evaluations are expensive, and prior LLM-assisted tuners use the model only as an external proposer with transient context. StateTune reformulates the task as a closed-loop, state-carrying process whose optimizer state is a typed, evidence-gated persistent optimization memory updated by every evaluation and shared between candidate generation and budget allocation, with an expected hypervolume improvement (EHVI)-guided, runtime-aware policy that promotes quick-stage candidates by expected Pareto-frontier gain per unit of runtime. On a Cadence industrial flow across six benchmark blocks and against five baselines including LLM-plus-retrieval-augmented generation (RAG) and preference-based Bayesian optimization (BO) tuners, it achieves the strongest final hypervolume on all six blocks and matches or beats the best baselines on worst negative slack, area, and power. Ablation shows the persistent memory is the largest contributor, with its removal costing 58.5% of the hypervolume, and further analyses cover evidence-gating sensitivity, memory poisoning, cross-design transfer, and three-seed reproducibility.

Identifying Latent Declarative Representations of Code for Assisting Repository Migration

Shraddha Surana, Ashwin Srinivasan, Michael Bain cross-listed Legacy repositories embed decades of undocumented domain knowledge, making repository-scale porting difficult. ADFD-Migrate treats a program as the implementation of an unobserved declarative description and approximates that latent representation with an annotated data-flow diagram (ADFD) of processes, data stores, external entities, flows, and behavioral contracts, which an LLM infers from bounded repository context under static-analysis coverage checks; dependency-aware chunking then orders process groups for target-language generation, and differences between the source ADFD and a statically recovered target ADFD guide regeneration. On f2x50, a new benchmark of 50 Fortran repositories spanning 1.5k to 1.6M lines of code across three complexity tiers, the generated Python passes 327 of 382 curated Fortran-oracle probes (85.6%) and exposes all 382 planned behaviors as runnable targets, versus 99 and 98 for direct and repository-context translation. The approach also reaches a 93.1% mean migration outcome index with a 17 to 59 percentage-point advantage over direct translation on 47 repositories.

LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology

Marie-Lisa Eich, Kai Standvoss, Timo Milbich, Alexander M\"ollers, Miriam H\"agele, Philipp Anders et al. cross-listed Lung cancer pathology requires integrating histomorphological, immunohistochemical, and molecular features, yet assessment remains largely visual with interobserver variability, and existing AI tools cover only isolated tasks without prospective validation. LUCAID couples an integrative diagnostic-reasoning agent with nine modules spanning quality control, tumor detection and segmentation, histological subtyping, tumor microenvironment profiling, cellularity quantification, biomarker scoring (PD-L1, MET, TROP-2), and structured report generation, and lets users interactively query module outputs. The modules reach F1 scores of 0.82-0.95 against expert annotations, and in prospective clinical validation the system achieved 93.0% concordance with an expert-panel reference standard on clinically actionable decisions, versus 68.3-81.1% for five experienced thoracic pathologists.

Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring

Olga Manakina, Igor Bogdanov Large language models score essays well, but fixed prompt strategies ignore operational cost and shifting optimal configurations. The authors treat each of four grading recipes (multi-step versus single-step, with or without calibration examples) as an arm in a multi-armed bandit (MAB) controller that adaptively selects prompting strategies at inference time, tracking token usage and latency alongside agreement metrics on IELTS Writing Task 2 essays. The bandit matches the accuracy of exhaustive grid search while reducing LLM calls by 78.4%, the multi-step-with-examples recipe proves most accurate, and the resulting cost-reliability curves are offered as guidance for educational platforms weighing cost against assessment validity.

Automated Synthesis of Cloud Emulators

Archit Bhatnagar, Zhenning Yang, Sarah McClure, Yiming Qiu, Sylvia Ratnasamy, Ang Chen cross-listed Cloud emulators such as LocalStack let DevOps scripts and infrastructure-as-code programs be tested locally instead of against slow, costly, and risky real cloud resources, but building them means manually interpreting extensive documentation and hand-coding logic for every service and API as the cloud keeps evolving. CloudEmu synthesizes emulators automatically from cloud documentation using neurosymbolic code synthesis: large language models handle documentation understanding and code generation, cloud-specific symbolic abstractions suppress hallucinations and enforce precision at scale, and the real cloud serves as an oracle for automated testing, repair, and alignment. On major AWS and GCP services, CloudEmu outperforms LocalStack, which a large team of engineers built manually over a decade, in both coverage and accuracy.

A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization

Prithvi Dake, Rahul Bindlish, James B. Rawlings cross-listed Real-time optimization (RTO) uses process models to find economically optimal operating points, and data-driven models are attractive because first-principles modelling demands deep process knowledge, but it is unclear whether a model that fits plant data can be trusted for optimization. Using a vinyl acetate monomer benchmark process with a single well-conditioned economic optimum, the authors train a hybrid model combining known mass balances and thermodynamics with a neural-network kinetics closure, and a fully data-driven neural ordinary differential equation (ODE) model. Both models reproduce plant measurements accurately with little variance across initializations, yet return many phantom optima whose economics differ substantially from the plant's true optimum, and even with noise-free data and initialization at weights that recover the plant optimum, stochastic gradient training drifts to weights giving worse RTO solutions. The authors argue that data-driven RTO models should be required to recover the optimum on a decision-oriented benchmark before plant testing.

Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory

Danish Khan, Maurice D. Hanisch, Nikolai Argatoff, Evan Xie, Sandeep Sharma, Anima Anandkumar cross-listed Kohn-Sham density functional theory (DFT) scales cubically because of repeated orbital diagonalizations, and orbital-free alternatives have fallen short whether they learn ill-conditioned kinetic-energy functionals or directly predict ground states that extrapolate poorly. The authors instead target the Kohn-Sham map itself, which sends a potential to the corresponding density and noninteracting kinetic energy, and train a domain-invariant SE(3)-equivariant Fourier neural operator to predict the density from the potential on real-space grids, enabling quasi-linear-scaling self-consistent field (SCF) iterations. Trained jointly on 8,504 molecules and solids, a single model converges SCFs on out-of-distribution organic molecules, insulators, and metals without ever constructing Kohn-Sham orbitals, reproducing densities, electronic spectra, and structural observables at Kohn-Sham accuracy, and scales to magnesium dislocation systems with up to 82,500 valence electrons on one GPU.

Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs

Akash Raj, Sargam Sahu When a code-generating model invents a Python package name, an attacker who has pre-registered that name on PyPI can turn the hallucination into a supply-chain compromise, an attack termed slopsquatting. The proposed detector pairs a deterministic PyPI existence check with a Random Forest classifier over ten features from the package name and its PyPI metadata, bridged by an import-name reconciler that handles cases like import cv2 versus pip install opencv-python, and embeds it in a LangGraph state machine that retries at escalating temperatures before falling back to a stronger model. Across 300 curated prompts the pipeline yields hallucination-free code on 76% of runs, and half of the flagged hallucinations were already-registered low-quality lookalikes on PyPI that only the classifier caught; hallucination rate rises from 0-10% on routine prompts to 40-73% on slopsquat baits, and when primary and fallback share a model family about 84% of failures recur, arguing for cross-family pairing.

PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage

Chuqing Gao, Yuanfang Song, Jonathan Zhang, Yifan Wu, Vishwakarma Singh, Qinglong Zeng et al. Enterprise AI agents in production usually need to be bounded, stateful, observable, and governable rather than fully autonomous, and PinSieve is a case study of that in a large-scale content-quality pipeline. Its deployed component is a selective vision-language-model (VLM) Serving Agent that handles only the grey-zone items lightweight upstream models leave unresolved, exposes a scalar routing score online, and keeps controlled human escalation; on that slice it filters 2.05x more non-actionable items than the previous production module while slightly lowering the estimated miss rate, and after promotion improves review productivity by 25.7%, cuts normalized operating cost by 16.2%, and moves signal delivery from next-day to same-day. For maintenance, a governed memory flywheel records routing traces and audit metadata, a Data Curation Agent runs a bounded proposal-verifier loop over representative, uncertainty, recency, and fresh-review replay with positive-rate and score-bin guardrails before accepting batches, and a Reasoning Review Agent audits teacher-generated rationales; in chained monthly refreshes over six months of production data, this curation reduces the average false-negative rate FNR@50% from 17.73% under random replay to 13.29%. The authors attribute production claims only to the deployed Serving Agent and treat replay and rationale-review results as offline or sampled-governance evidence.

PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation

Sang Won Lee, Hyogu Jeong, Namwoo Kang Generative and predictive models for engineering design and simulation are usually evaluated in isolation, on academic datasets at unconstrained scales, with inconsistent metrics and procedures. PhysicsBench is a unified benchmark and leaderboard covering seven generation and prediction tasks across 1D, 2D, and 3D domains, ranking 66 models on nine datasets (industrial-scale CAD, CFD, and FEA simulations plus public references) expanded into 28 configurations at limited data scales from S to XL, with a BenchRank procedure that debiases correlated metrics and ranks via PageRank over a head-to-head dominance graph while reporting computational cost separately. The top model changes with data scale in six of the seven tasks, no model leads more than one task, and an architecture's large-scale academic standing only weakly predicts its small-data ranking.

SQLite is Enough. Lexical, Semantic, and Hybrid Search with scrydb

Timo Breuer cross-listed scrydb is a Python library that brings lexical, semantic, and hybrid search to SQLite: lexical search uses the built-in FTS5 full-text extension, semantic search builds on the sqlite-vec vector extension, and results from both can be reranked and fused. The aim is a lightweight retrieval backend for information retrieval (IR) tasks or agentic search that needs no separate search service. Evaluation on several IR benchmark datasets demonstrates effective retrieval via keyword matching, semantic similarity, and rank fusion, alongside query-latency measurements that characterize the efficiency-effectiveness trade-off; the library is released under the MIT license.

Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate

Ethan Traister, Ankit Raj, Jiaqi Gan, Xingyu Shen, Tyler Wu, Yuchen Zhou et al. cross-listed Telephone fraud is costly but rarely observed at scale; an AI voice-agent honeypot that answered callers and kept them talking collected 10,211 inbound scam and spam calls (913 hours and 330,956 transcribed turns from 5,780 numbers) over 54 days. The analysis separates outright scams from the larger stream of legal-but-predatory lead generation and finds that operations keep office hours (6.6x more calls per weekday than weekend day), thousands of disposable numbers recycle a small catalog of scripts with half the traffic in the top five opening clusters, and callers seek identity anchors such as home address and date of birth far more often than payment credentials. A randomized experiment assigning one of ten fictitious identities to each seeded lead shows scammers spent about 15% more conversational turns per decade of the target's apparent age (rate ratio 1.15) but did not change what they asked for, with 26.3% of calls reaching a request for sensitive information regardless of age. A second experiment frames early detection as a benchmark: escalation is predictable from opening lines alone at 0.72 ROC-AUC from the first line and 0.87 by the eighth, and a plain bag-of-words classifier matches a fine-tuned on-device language model.

From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender

Sonia Sharma, Jeyendran Balakrishnan, Shreya Rajpal, Swapnil Parekh, Nagaraj Janardhana, Andrew Mattarella-Micke Service-business catalogs are shifting from static, independently priced SKUs toward dynamically bundled, discount-coupled offerings, which strains the tree-based multiclass classifiers favored for sparse, imbalanced data and makes it hard to incorporate multimodal signals such as conversation transcripts. The authors describe migrating a live production conversational recommender from a gradient-boosted multiclass model to a pairwise-binary deep recommender that learns jointly from user and item features, strengthened by negative sampling and noise injection, with attention pooling over transcript chunks to bring in live conversation context, benchmarked against TF-IDF and sentence-embedding baselines. They compare architectures including two-tower models and DeepFM variants and loss functions including contrastive loss, under the constraint that live recommendation quality could not regress. Against a CatBoost baseline across conversational stages, the deep model achieves parity at conversation start and outperforms at later stages.

Constraint-Guided Enterprise Data Mapping with Large Language Models

Sebastian Monka, Pramod Anantharam, Thien Vo Minh, Lavdim Halilaj Enterprise entity alignment across semi-structured records with implicit attributes and unit or granularity mismatches is still often done by hand, and large language model (LLM) matching on its own can produce fluent correspondences that violate structural or physical invariants. Constraint-guided mapping (CGM) is a neuro-symbolic pipeline that derives schema-grounded admissibility constraints carrying executable relation and normalization logic, generates candidates restricted by those constraints with cascade relaxation so the feasible set is never empty, and only then applies neural ranking and bounded LLM disambiguation within that set. On a structural-decoy benchmark the hard admissibility gate shrinks the candidate space by roughly 480x without losing the ground truth, and a layer-by-layer ablation attributes the decisive gain (F1 0.08 to 0.66) to the constraint gate rather than the LLM; a small model with constraints matches a frontier LLM without them at about 28x lower cost. The approach transfers across seven enterprise makes (macro F1 0.70) using automatically discovered, expert-refinable constraints and cuts expert effort around 7x relative to spreadsheet workflows.

FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment

Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng Adversarial robustness of fraud and credit-risk models is hard to evaluate because financial tabular data carry domain constraints, severe class imbalance, and asymmetric attacker capability, and the authors argue that robustness is as much a property of the evaluation protocol as of the model. FraudBench evaluates the same dataset-model-attack-defence setting under three matched protocols (unconstrained attacks, post-hoc feasibility filtering, and deployment-aware constraint-integrated attacks) across four public financial datasets and neural, tree-based, and ensemble models in three attack settings. On Lending Club Loan Data under white-box attack, post-hoc filtering leaves only 3.7 feasible-flipped examples on average while in-attack projection with attacker mutability masking yields 2,832.3 under the same budget; results on IEEE-CIS show feasibility and attacker capability are separate axes, and black-box evaluation shows protocol choice can reorder model-family rankings. The authors recommend reporting predictive degradation and attack feasibility jointly and building domain constraints into attack generation rather than post-processing.

Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks

Yulong Dou, Han Wu, Guo Chen, Fangmao Ju, Zhiming Cui, Dinggang Shen Electroencephalography (EEG) models are usually trained one per dataset, and the recent crop of EEG foundation models leans on reconstruction objectives that implicitly assume anything locally predictable in the signal is transferable neural information. INCEPT is pre-trained on more than 11,000 hours of unlabelled clinical EEG with an invariance objective instead, learning representations that stay stable across correlated observations of the same recording so that neural structure and subject-specific information separate from the nuisance variability dominating scalp electrodes. Evaluated on ten datasets spanning signal-quality assessment, brain-state decoding, and brain-health prediction, it ranks first among recent EEG foundation models on 26 of 30 linear-probing and 24 of 30 fine-tuning metrics and also beats task-specific specialist encoders.

Data Leakage Inflates Generalizability of Power Outage Prediction Models

Yamil Essus, Ranga Raju Vatsavai, Benjamin Rachunok Models that predict power outages from weather are increasingly used in climate-driven infrastructure risk assessments, but common evaluation practices hide whether they generalize to the novel conditions those assessments require. Using public U.S. East Coast data from 2018 to 2023 with features from weather reanalysis, land cover, and embeddings from the Prithvi WxC GeoAI foundation model, the study compares random train-test splits against leave-one-state-out and leave-one-event-out holdouts that better approximate deployment. Random splits look strong, but under spatial and temporal holdouts accuracy degrades so much that models often fail to beat a null baseline, and the foundation-model embeddings give only limited, inconsistent gains for spatial generalization without fixing event-level transfer. The authors conclude that publicly trained outage models currently offer limited operational value and that progress depends more on data coverage and realistic evaluation protocols than on marginal modeling advances.

A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology

Eugene Vorontsov, Yi Kan Wang, Alican Bozkurt, Adam Casson, Ludmila Tydlitatova, Michal Zelechowski et al. Precision oncology needs a longitudinal representation of patient state that integrates clinical history with molecular and pathology data as cancer and treatment evolve over time. oFM is a foundation model built on a real-world cohort of 1.67 million cancer patients, with over one million used for training under patient-level splits, that encodes daily clinical and molecular episodes together with DNA, RNA, and H&E pathology images into a patient state embedding. Frozen oFM embeddings outperform expert-curated clinical and molecular baseline features on treatment response, progression-free survival, and overall survival (AUC 0.774 versus 0.563 for overall survival), and across 11 comparative-treatment cohorts they achieve about three-fold higher pooled treatment-benefit AUTOC with better benefit ranking in 9 of 11 cohorts. A companion mechanism discovery framework links downstream predictions to clinically and biologically grounded mechanisms through an evidence-grounded temporal graph.

ExpConCAD: Experience-Guided Text-to-CAD Generation from Shape Descriptions with Implicit Spatial Constraints

Jingyao Liu, Jinkang Tang, Chen Huang, Wenqiang Lei, See-Kiong Ng Text-to-CAD systems generate executable CAD programs from natural-language descriptions, but real descriptions are often underspecified and omit spatial constraints needed for a valid construction. ExpConCAD treats the missing constraints as something to infer from the underlying construction structure and from reusable design experience: it first recovers the intended construction structure and constraint scopes, retrieves constraint-completion experience for similar scopes, and then generates executable CadQuery programs. Experiments show the framework improves generation for descriptions with implicit spatial constraints, with analyses of how construction-structure understanding and experience memory each contribute to constraint completion.
68 more specialized papers

Agents 56

REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring

Muhammad Waseem, Aakash Ahmad, Pekka Abrahamsson cross-listed LLM-generated refactorings must reduce targeted quality problems without introducing new issues or altering behavior-relevant code structure. REFINE (Refactoring with Evidence-aware Flow for Integrated ageNtic Execution) is a tool-agnostic multi-agent pipeline for Java file-level refactoring that chains static-analysis-guided smell identification, smell-informed planning, LLM transformation, automated re-analysis, preservation checks, and structured reporting. Evaluated on 450 files from 15 open-source systems with GPT-5.5, Gemini 3.1 Pro Preview, and Claude Opus 4.8, it reduces detected code smells by 68.26%, 72.79%, and 68.49% respectively, and against a matched 150-file direct-prompt baseline it achieves higher median smell reduction with smaller edits and fewer public-method removals. Broader quality improvements are inconsistent, however, and preservation checks reveal residual risks such as changed assert/fail calls and removed public methods, so outputs are framed as candidates that still require compilation, testing, dependency analysis, and human review.

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

Parker Fawcett cross-listed Prior work found that once a model is strong enough, a multi-agent rebuild pipeline loses to simply handing the model the original code with one instruction. rebuild-dossier is an open-source tool that locks an application's real interface, its exact inputs and outputs, before any code is written and then enforces one-test-at-a-time building through automated checks rather than written instructions alone. Three results with differing evidence emerge: in a small comparison the process-compliant agent failed a held-back test while the rule-breaking agent passed everything, showing a passing suite does not certify correctness when tests can be gamed; against the source-plus-one-instruction baseline the tool tied on a small app but lost outright on a larger one where the automated check was not running, pointing to the check mechanism rather than interface-locking as the active ingredient; and every claim was cross-checked at three levels (agent report, automated log, produced files), which caught real errors including a bug in the authors' own logging code. On a different model and toolchain, a stronger model followed the process three times running, which the weaker model never managed.

LLM Agents Perform Controlled Experiments Using Simulation Models

Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes St\"umpfle, Johannes Sigel, Akshay Narla et al. Many scientific and engineering tasks require understanding how a system responds to intervention, which plausible text and code generation alone cannot provide. The proposed multi-agent framework lets LLM agents run controlled experiments against high-fidelity simulation models for pharmaceutical process design: given a user query and a baseline configuration, it builds a structured task representation, designs experiments, executes comparative simulations, interprets the outcomes, and synthesizes evidence-based recommendations for process parameter optimization. In an industrial application setting, the simulation-integrated system produced more specific outputs and higher user-rated correctness and helpfulness than language-only reasoning, with ablations and visualized case analyses supporting the value of reasoning through intervention, comparison, and observation.

When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs

Jason Liu cross-listed Tool-using agents need a principled rule for when they may declare a task complete, and existing systems that gate terminal success or certify execution traces have not been tested against controlled termination faults. Evidence-Carrying Termination (ECT) allows an agent to return COMPLETE only when a typed certificate binds every required answer claim to valid, in-scope trace evidence and a deterministic replay reconstructs the claimed value. In a locked static study of 48 synthetic tasks across six tool-use families with clean execution and eight injected faults, ECT produced 0 of 288 unsafe completions versus 252 of 288 for a termination-critic core; a fresh prespecified 576-trajectory study on 22 held-out task clusters found 0 of 66 premature unsupported terminations versus 40 of 66 for the critic's controller, while supported completion (97 of 132 versus 92 of 132) satisfied a 10-point noninferiority margin. The certificate attests to support within a recorded trace under declared assumptions, not to external truth, safety, or alignment.

TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery

Kang Zhou, Yujia Tong, Yong Tao, Jingling Yuan Multi-objective materials discovery with LLM agents is limited less by how many candidates can be proposed than by how much each costly property evaluation teaches the next search step, since existing agents store evaluated candidates and scores but not which executable edits caused useful property changes. TRACE treats evaluated edits as the unit of feedback, recording each refinement as a parent-edit-child transition with observed property deltas, aggregating that evidence into reusable edit-effect estimates, and ranking future edits by their predicted ability to reduce remaining constraint violations without damaging already satisfied objectives. In a same-backbone comparison against LLEMA, the prior state-of-the-art LLM-agent baseline, it raises macro-average hit rate from 18.13% to 25.96%.

ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents

YiShan Zheng, Yuan Wu, Yi Chang cross-listed Clean end-to-end (E2E) success rates cannot reveal where a tool-calling failure originates or how it propagates through a call. ToolRobustBench is a stage-wise diagnostic benchmark that aligns four perturbation families (tool-interface, user-intent, tool-output/observation, and runtime-environment) with the tool-use pipeline and attributes failures to tool selection, schema grounding, argument binding, feedback handling, or E2E task success. Experiments on 15,456 single-family instances across 7 models, 16 local tools, and 14 perturbation subtypes show high but non-uniform clean performance and substantial degradation under perturbation, with tool-output/observation perturbation the dominant bottleneck, and mixed-family experiments reveal non-additive failure patterns that isolated single-family results do not explain.

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

Esmail Gumaan cross-listed Agent harnesses record a failed tool call and its error message in the transcript on the assumption that the error is corrective information. Defining corrective gain as the change in log-probability of re-emitting the action that just failed, the authors find it is negative for every instruction-tuned model tested (6 checkpoints, 135M to 1.7B parameters, 4 families) in both simulated tool calling and MBPP program repair, about -1.03 nats per action token, with the probability of repeating the failed call over a fixed candidate set rising from 0.06 to 0.54. Counterfactuals show the failed call's surface form accounts for 83% of the damage rather than the failure marking itself, which predicts which remedies work: replacing the verbatim call with a runtime-generated description of the failure removes 76% of the inversion at no token cost, whereas an explicit "do not repeat" instruction changes nothing and retrying from a clean context is the worst harness measured because it restores the context that produced the failure. The study runs end to end on a CPU and all artefacts are released.

Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

Zizhe Wang cross-listed Physical system modeling in Modelica, an equation-based language, differs from ordinary code generation because a model can compile and simulate while still violating its intended physics, and across revisions an agent may lose track of requirements or lean on simulation evidence from an outdated candidate. Pufibara is an agent harness that maintains persistent engineering state across revisions, ties execution and simulation evidence to the candidate that produced it, and makes submission an explicit action; alongside it the authors build the 232-task Modelica Agent Workflow Benchmark spanning Model Repair, Model Generation, and Model Tuning, scored by an evaluator outside the agent loop. Under matched backends, Pufibara passes 202 tasks versus 185 for Claude Code with DeepSeek v4 Flash and 202 versus 187 with Claude Sonnet 5, with 76.4% to 82.5% lower logical-token totals and 6.1% to 58.4% lower sequential runtime.

Automata from Agent Traces: Failure and Next-Step Prediction

Seonglae Cho, Franklin Cardenoso Fernandez, Umar Mohammed, Zekun Wu, Kleyton Da Costa, Ilham Wicaksono et al. Long, unstructured traces from LLM-based agents are hard to audit or monitor, and existing analyses work per-trace or only on successful runs, missing the cross-run structure that links next-step and failure prediction. The approach collapses an entire trace corpus into a single compact finite-state machine (FSM) that serves as a structural substrate for otherwise unpredictable agent behavior. Across twelve public datasets the FSMs have 7-43 states, replay held-out data at fitness of at least 0.997, and build in milliseconds; FSM-state context beats Agent Workflow Memory on next-step prediction for every ground-truth-matched dataset, and per-state behavioral features reach held-out AUROC up to 0.94 for failure prediction, enabling an online monitor that flags failing runs from partial traces well before completion. The authors argue that behavioral topology is shaped more by the deployment harness than by the underlying LLM.

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

Stephen Chung, Wenyu Du, William J. Wesley The Station is an open-world multi-agent environment in which AI agents from different model families pursue a shared mathematical research goal with no central coordinator or scripted pipeline, choosing their own directions, running experiments, collaborating, and building a shared literature. Applied to 12 construction problems from the AlphaEvolve catalogue plus two case studies, the system produced results novel relative to prior literature on five problems, including a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem, along with new infinite families for Book Ramsey numbers. The agents also produced theorems and analyses explaining why the constructions work, and all raw dialogues, proofs, and verification code are released.

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

Seonglae Cho, Donghyun Lee Existing multi-agent coding systems inherit the serial nature of token generation, either sequencing agents through phase handoffs or pooling independent samples without coordination, and a single agent abandons up to half of hard tasks with a one-file stub-and-exit. AgentRoom is a realtime collaborative editing protocol for concurrent coding agents whose runtime exposes file-level claim, status, and broadcast operations as Model Context Protocol (MCP) tools over a shared filesystem merged with Conflict-free Replicated Data Types (CRDTs). Five frontier coding-CLI models ran four backend tasks with cross-language checks in Python DevBench and Rust plus axum; for CLI-stable models, AgentRoom with two agents abandons fewer tasks than solo runs and shows less run-to-run variation, and at matched compute an LLM-judge contrast favors it over parallel-merge. A bundle probe ranks full AgentRoom above each partial configuration, indicating that coordination, rather than parallelism or CRDT merging alone, accounts for the gains.

Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search

Jiongxiao Wang, Dingli Ma, Chaoqun Ni Automated biomedical fact-checking systems built on retrieval-augmented generation (RAG) and agentic search usually emit bare supported/refuted labels with little explanatory value. BioCheck Agent instead produces structured fact-checking reports, searching only PubMed with Boolean operators and synthesizing conclusions from retrieved evidence, and is trained with EG-GRPO (Evidence-Grounded Group Relative Policy Optimization), a reinforcement learning recipe whose task-specific reward encourages advanced search behavior and high-quality evidence retrieval while penalizing hallucinations. Relative to the Qwen3.5-4B base model, the trained agent improves label accuracy on SciFact by 9.95%, raises evidence quality by 3.7%, and lowers the evidence hallucination rate by 19.63%.

Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization

Yapeng Liu, Yuanzhao Zhai, Xudong Gong, Dawei Feng, Bo Ding, Lin Wang et al. cross-listed Evaluations of embodied agent systems (EAS) rely on outcome metrics such as success rate or safety scores that collapse whole execution trajectories into a single number, hiding how agents recover, stabilize, and extend under perturbations. Drawing on resilience engineering, the authors define a metrics suite covering Rebound, Stability, and Graceful Extensibility, plus an evaluation layer that turns execution traces into diagnostic assessments, and apply it to 10 systems on 400 household tasks. The results expose process-level differences that outcome metrics hide, including a recovery-cost difference of 25.2 among episodes that all counted as successful, and metrics-guided optimization reduces recovery cost and improves stability and graceful-extensibility completion, while revealing a trade-off among resilience properties that suggests configuring systems to deployment-specific requirements.

Exploit More, Explore Smarter for Budget-Constrained Agentic Search

Haoyang Fang, Bernie Wang When a large language model agent must refine candidates under a small evaluation budget, standard Monte Carlo tree search (MCTS) spends it poorly: exploration bonuses dominate at low visit counts, unpromising siblings are expanded before promising chains can deepen, and branching ignores node quality. ExTS treats expansion itself as a value-of-information decision, combining discriminative reward shaping to separate candidates with narrow score distributions, a stochastic virtual child that estimates the value of a new branch from the parent's reward history, and quality-conditioned branching that expands only when a node's score justifies the budget cost. Across prompt optimization, code generation, molecular structure elucidation, and agentic workflow optimization, a single fixed configuration is competitive with or better than task-specific tree-search baselines, with an average relative gain of +5.5%, and pilot-run diagnostics characterize how these budget-constrained problems differ structurally from one another.

Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)

Avital Aviv, Parth A. Gandh, Ron Bitton, Asaf Shabtai cross-listed Google's Agent Payments Protocol (AP2) lets large language model (LLM) shopping agents authorize payments through signed Checkout and Payment Mandates, but those signatures only protect transaction data after signing, leaving pre-authorization inputs such as Agent-to-Agent Protocol (A2A) messages and Model Context Protocol (MCP) tool calls unprotected. The authors analyze AP2 v0.2 by dividing its transaction lifecycle into five phases and identifying five deployment architectures, then apply the MAESTRO (Multi-Agent Environment, Security, Threat, Risk, Outcome) framework to model four threat actors, eleven attack surfaces, and eighteen adversary capabilities, producing a catalog of 48 threats across five attack families scored with the Artificial Intelligence Vulnerability Scoring System (AIVSS). Because no complete public deployment existed, they built a testbed spanning all five architectures with proof-of-concept demonstrations of all eight threats that reach the High-risk band, plus a deployment-aware scanner, concluding that valid mandate signatures alone cannot guarantee a transaction reflects user intent when its pre-authorization context is manipulated.

Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information

Xiao Liu, Haoyang Li, Songwei Li, Hongbo Fang, Fengli Xu, Feng Shi et al. cross-listed Orchestrating LLM agents built by different parties with different capabilities and costs is more like assembling labor across an economy than calling a subroutine, yet existing orchestration uses a centralized planner that bottlenecks as agent pools grow, requires private information such as execution costs, and is easily manipulated, with a single inserted preference nearly doubling a favored agent's task share under a centralized LLM allocator. AgentLance runs a repeated labor market in which agents bid on tasks using their private costs and self-maintained strategy notes, an allocator selects winners from bids and public reputation records, a Vickrey-Clarke-Groves (VCG)-style payment rule rewards cost-aware bidding, and winning agents can decompose complex tasks and subcontract pieces through the same mechanism. Across mathematical reasoning, code generation, knowledge-intensive QA, and agentic tasks, AgentLance consistently outperforms single-model, centralized-orchestration, and market baselines, matching agents to their specializations and shifting work toward cheaper agents as cost sensitivity rises, with further gains from diagnosing and correcting inaccurate cost self-estimation and sub-optimal bidding.

Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment

Aryan Brar, Justin Du, Avery Lor, Kylie Seto, Eric Taylor Tax-loss harvesting reliably improves long-term portfolio growth but depends on holdings- and owner-specific details, so the authors build a multi-agent trade-recommendation system and test two context providers: a custom capital-gains calculation engine and a retrieval-augmented generation (RAG) vector store of market advisory reports. In a 2x2 repeated-measures design scored by relative capital gains incurred during liquidation, enabling the tax engine significantly reduced tax savings by roughly 55 percentage points (F(1,29) = 9.17, p = .005), while the RAG effect and the interaction were not significant. The RAG-only condition achieved the highest mean savings (47.7%) and the untooled baseline came second (30.6%), suggesting the model's internalized financial knowledge suffices and that domain-specific computation tools can inject conflicting optimization signals rather than help.

PROOF-Gen: From Optimized Data to Better Distillation

Anh Ta, Junjie Zhu, Shahin Shayandeh Distilling tool-calling ability into small deployable models starts with supervised fine-tuning on teacher trajectories, but the standard generate-and-filter pipeline discards every failed teacher run, so the same hard scenarios are left behind each retraining cycle; on τ²-bench 57% of teacher trials fail and two-thirds of those are near-misses undone by a single decisive error. Per-scenario Reflective Optimization to Overcome Failed Generation has a reflector analyze each failed execution trace and its evaluation feedback, write corrective guidance that steers the teacher to a passing trajectory, and then strip that guidance so the student trains on clean demonstrations with no task-specific scaffold. The method recovers 93% of failed τ²-bench scenarios, lifting Qwen3-4B-Instruct-2507 from a Pass^1 of 0.132 to 0.529 and Gemma 4 E4B-it by 7.2 points on BFCL v4 multi-turn, and in a deployed pipeline it raises goal completion by 6.3 points for trajectories and 1.5 points for an on-device model, with positive transfer in every locale.

MARS: Multi-Specialist LLM Relay System for Competitive Programming

Andrei Mikhailov, Mikhail Burtsev, Alsu Sagirova Multi-agent pipelines for competitive programming usually split work across generic planner, coder, and debugger roles and leave the choice of algorithmic technique entirely to the backbone model. MARS (Multi-Agent Relay of Specialized LLMs) is a prompt-only framework in which every agent is a topic specialist (dynamic programming, graphs, strings, geometry, and so on) grounded by retrieval-augmented generation over an algorithm-theory corpus: retrieval selects a small team per problem, a starter writes a C++17 solution, and each subsequent specialist runs the candidate against public examples in a sandbox, then keeps, repairs, or hands off the draft via a structured packet before a final infrastructure-fixer pass normalizes boilerplate. On the CodeContests test split with Gemma 4, the relay reaches 0.624 pass rate, 14.4 percentage points above direct prompting, at 2.3 pipeline stages per task, closing most of the gap to CodeSIM (0.731) at 3.3x lower wall-clock cost and with far smaller variance in per-task token spend.

The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses

Dai Jiahong cross-listed An agent harness is the code surrounding a language model that builds its context, mediates its tools, runs the loop, and persists state across a long-horizon run, and the authors argue this layer rather than the model is increasingly the binding constraint on agent behaviour. Through a source-level multi-case study of three open coding-agent harnesses built on opposing philosophies, LangChain's deepagents (batteries-included), Earendil's pi (radical minimalism), and DeepSeek's dsh (everything-is-a-plugin), each read at a pinned commit and traced through its history, they find the two mature harnesses moved in opposite directions (deepagents shedding authored scaffolding, pi accreting durable infrastructure) yet converged on the same five elements: a commoditised loop, an append-only replayable session record, model quirks kept as data, progressive disclosure of context, and explicit extension seams. The third harness, read afterward as a held-out check, exhibits all five and in one seam reuses another's implementation outright, so the convergence is decomposed into parallel discovery, diffusion, and literal reuse rather than claimed as independent invention. One load-bearing dimension is absent from all three, external verifiability via a tamper-evident record an outside party can check without trusting the runtime, which the authors read as the next axis on which harnesses for provenance-sensitive domains will diverge.

Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation

Olympia Saha, Amy Wang, Srinivasan Manoharan cross-listed An enterprise Model Context Protocol (MCP) gateway that proxies hundreds of backend servers behind one endpoint creates two compounding problems for large language model (LLM) agents: full tool schemas can saturate the context window before any user query arrives, and neither agents nor users can find the right tool among thousands. SCOUT (Selective Context Optimization for Universal Tooling) treats tool exposure as a context-selection problem, exposing just two meta-tools, tool_search and execute_tool, where tool_search fuses BM25 sparse matching with dense vector search via Reciprocal Rank Fusion to return only the top-k tools relevant to the current step, backed by zero-downtime catalog update pipelines. In production at PayPal, covering more than 2,000 tools across 200-plus MCP servers, MCP tool-token consumption drops from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%), a 99% reduction. Because it is surfaced as standard MCP tools, the approach is model-agnostic and needs no client-side changes.

AgentSpec: Speculative Decoding for Batch Inference of LLM Agents

Xin Wang, Ziming Miao, Yi Zhu, Hui Shen, Zhongwei Wan, Fan Yang et al. Speculative decoding can speed up large language model (LLM) agent applications without changing outputs, but current algorithms lose most of their advantage at the large batch sizes real deployments run. A systematic analysis attributes the slowdown to two factors, high rejection rates for speculative tokens and under-use of the token budget that frees up dynamically, and AgentSpec addresses both: structure-isolated drafting confines speculation to semantically coherent segments of the agent workflow so drafts rarely wander down irrelevant paths, and redundancy-aware budget allocation uses agent-level information to fill the spare budget. Implemented in vLLM and evaluated on five workloads and four models from four LLM families, AgentSpec outperforms state-of-the-art speculative decoding methods for batched agent inference.

Reflection with Action-Induced Visual Differences for Desktop GUI Agents

Yijie Ma, Chaoyue Niu, Fan Wu, Guihai Chen In the Planner-Operator-Reflector (POR) framework for GUI agents, desktop interfaces are large and dense with subtle, scattered state changes, so most of the burden falls on the reflector, which must compare pre- and post-action screens; existing reflectors fold change detection and outcome verification into one step, leaving evidence implicit. Evidence-First Reflection (EFR) is a two-stage reflector that first locates the action and candidate changed regions using Set-of-Marks annotations, describes and filters the changes that are relevant to the action, and only then judges the outcome from the cleaned evidence. On OSWorld-Verified and WindowsAgentArena, EFR raises reflector accuracy by 7.11% and yields average end-to-end task success gains of 5.94% and 4.95% respectively.

Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding

Hyunho Kook, Junhyuk So, Tianyu Fu, Haizhong Zheng, Beidi Chen Confidence-based voting weights parallel LLM rollouts by internal signals such as token log probabilities and works well for single-turn reasoning, but it transfers poorly to multi-turn search agents that retrieve and condition on external documents. The authors trace the failure to copy inflation: once retrieved documents sit in the context, tokens copied from them receive systematically inflated log probabilities, flattening confidence scores within a question and weakening the weighted vote. Retrieval-Grounded Voting (RGV) instead scores each rollout by the lexical overlap between its final answer and the documents it retrieved, computed outside the contaminated context and without extra LLM calls. Across four search-agent benchmarks and five LLMs, RGV consistently beats confidence-based voting, with gains of up to +5.4% accuracy and +35% on minority-correct questions where the right answer appears in only 1-2 of 8 rollouts.

Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems

CheolWon Na, Hao Ni, Lukasz Szpruch, Zhangyang Wang, Dhagash Mehta, Saurabh Nagrecha et al. LLM-based multi-agent trading systems are moving into live deployments that control real assets, and the inter-agent communication that makes them effective also lets a corrupted signal propagate to the final decision. The threat model restricts the adversary to the source data and prompts agents consume rather than system internals, then decomposes a widely used trading pipeline into Analyst, Researcher, Trader, and Risk Manager roles with an attack matched to each interface, and evaluates four communication topologies under data- and agent-level attacks using an Adversarial Signal Preservation Score (APS) to explain why some designs resist better than others. Across five assets, two backbones, and two target directions, the central finding is that no architecture is inherently robust to these low-barrier attacks.

AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval

Gunja Agarwal, Arup Kumar Das, Arun Menon, Jitesh Chandra Mishra, Vignesh Divakaran Evaluation of agentic information retrieval typically relies on scripted interactions with uniform simulated users, missing both natural personality diversity and adversarial brittleness. AgentWorld combines Big Five (OCEAN) personality-driven user populations with stateful tool-use environments, a pass^k consistency metric with structured fault classification, partial-credit scoring, and dual-control handoff verification, training-data export in six fine-tuning formats, and a Risk Analyser that snapshots required intermediate states, branches Monte-Carlo rollouts under four perturbation types, and attributes risk using Dempster-Shafer evidence fusion and Shapley values. In experiments on a conversational analytics agent across 10 personas and a customer-support agent across 5 tasks, personality variation surfaced failures uniform testing missed, including 50% versus 100% pass rates across personas on the same task, while adversarial stress-testing revealed pre-existing trajectory brittleness dominated by tool and infrastructure-layer attacks.

EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals

Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang et al. Large language models are increasingly used as code agents for scientific and engineering analysis, but whether they can work with raw physical-layer measurements has not been tested. EMRB poses 200 problems across five difficulty levels and 27 question types, generated from 11 signal types with verified ground truth, and hands the model only a raw I/Q (in-phase and quadrature) capture so that every quantity a question refers to must first be discovered by writing and running code. Across 14 proprietary, open-weight, and reasoning-oriented models, scores range from 24.1% to 78.9%, and mean accuracy falls from 84.9% on basic measurement to 21.2% on system design. A structured method called ReconPilot, which separates signal reconnaissance, targeted analysis, and self-verification, raises overall scores by 3.8 to 17.6 points across three backbones and improves 13 of 15 backbone-level combinations.

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

Nadeem Shaikh Existing large language model (LLM) agent systems decide whether to delegate to a stronger model either before reasoning starts (a router) or after a response is finished (a verifier); a third regime lets the agent hand off mid-generation once it recognises it is unlikely to succeed. Intra-generation delegation is formulated as a Bayesian optimal-stopping problem over a learned competence posterior, an online estimate of eventual task success whose sufficient statistics are fit from labelled trajectories rather than read off raw entropy. The analysis derives the myopic escalation threshold in closed form, proves the optimal policy is a time-varying threshold with no shape assumption on the signal, and gives a finite-sample guarantee that the plug-in policy's regret decays as 1/sqrt(n) with n calibration trajectories, which a controlled simulation confirms. A real-model validation on a Qwen2.5-Coder 1.5B-to-7B cascade over 257 MBPP tasks confirms two of three pre-registered predictions, including that the escalation frontier dominates post-hoc routing at equal cost.

Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

Guo Gan, Yilun Zhao, Cong Chen, Jinbiao Wei, Tingyu Song, Zheyuan Yang et al. Graphical user interface (GUI) agents deployed on Android routinely hit runtime anomalies such as unexpected pop-ups or misused actions, yet existing benchmarks do not systematically test robustness against them. AnTrap injects dynamic perturbations into agent execution trajectories according to a taxonomy of four layers (State, Thinking, Action, and Round) with ten fine-grained subcategories, using a construction pipeline that keeps tasks solvable while adding realistic adversarial conditions. Evaluating 16 leading GUI models shows universal vulnerability, with even the strongest models degrading significantly under injected anomalies. Training with GRPO in both clean and adversarial environments separates environment-learnable from reasoning-bottlenecked failures: single-step traps at the state and action layers are largely fixable through adversarial reinforcement learning, while deep contextual traps such as state deadlock expose limitations that training in trap-laden environments alone cannot resolve.

ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation

JooYoung Jang, Taegyeong Lee, Jihyeon Park, Nojun Kwak Large language model (LLM) agents editing legacy presentation formats confront flat, absolutely positioned elements that force coordinate recomputation and break layouts, and design tasks lack a unique ground truth for diff-based metrics. ACE edits over a hierarchical scene graph with a 98-tool presentation-specific action space, paired with CARE, a content-aware router that feeds the agent only the relevant slice of a deck (roughly 89% input-token reduction), and a self-correction loop in which a ground-truth-free instruction-following judge's natural-language critique becomes the next-turn instruction. With a fixed backbone, a single-turn scene-graph edit already matches an internally iterating agentic HTML pipeline, and adding self-correction lifts instruction following to 4.23 versus 3.81 on a 94-task benchmark at 1.75x the speed and about 44% lower cost. Twenty-six blind raters prefer the self-corrected output 81% of the time, 66% of cases halt after one pass, and a strict-peak rollback removes every observed regression.

Task-Adaptive Rubrics for GUI Reward Modeling

Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao et al. Outcome reward models for graphical user interface (GUI) agents judge whether an executed trajectory satisfies the user's instruction, but existing verifiers under-specify how success criteria should be built per task, so checks get transferred across unrelated tasks, concrete constraints in the instruction are missed, or unstated requirements are enforced. AdaptRubric is a coarse-to-fine rubric framework that first routes an instruction to a GUI task family and retrieves reusable family-level criteria, then generates instance-level cues capturing the specific values, scopes, and constraints in the current instruction. In offline reward evaluation it improves F1 by 3.6 points over the baseline average under a matched image budget, and used as the reward signal in online reinforcement learning it raises task success by 4.23 points over prior reward agents.

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

Jiayu Shi, Luzhuo Chen Coding agents re-send large file reads and tool outputs to a frontier model every turn, and general-purpose prompt compressors trained on prose paraphrase identifiers and drop the exact spans an agent needs to edit. Paritok-4B is a 264 MB LoRA adapter on Qwen3-4B, distilled from a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories, that is extractive (96% of the identifiers, paths, and numbers it emits appear verbatim in its input) and intent-conditioned, using the agent's current task to choose which lines survive within a retained segment rather than how much is retained. On all 300 SWE-bench Lite instances it compresses agent context to 25.7% of its original size while retaining 86.5% of uncompressed single-shot solve quality, twice as aggressive as a gpt-4.1-mini compressor and 2.4 times more than gpt-5; on the line-numbered input real agents produce, the solve-rate difference is not statistically significant (exact McNemar p=0.079), and because the adapter self-hosts on one 24 GB GPU with no per-token fee, it is cost-positive where using gpt-5 as a compressor costs more than the tokens it saves. Weights, data, and evaluation scripts are released under Apache 2.0.

MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG

Qiuyi Qi, Tian Liang, Jiamu Wang, Jinjian Zhang, Wei Zhou, Pengcheng Zhu et al. Agentic retrieval-augmented generation (RAG) requires a language model to decide when to keep searching and when to answer, and existing reinforcement learning approaches supervise this decision externally while ignoring the agent's own belief about whether the gathered evidence is sufficient. MetaRAG reframes search-decision quality as belief-action alignment: Verify-first Action Generation elicits an explicit verification step before each action, Internal Belief Probing estimates the policy model's own answerability belief from the same question-history context, and a consistency reward between the two is gated by answer correctness so that internally consistent but wrong trajectories are not reinforced. The belief probe is used only during training and adds no inference-time overhead, and across seven public question-answering benchmarks MetaRAG consistently improves the accuracy-efficiency trade-off over strong RL-based agentic RAG baselines, with gains that carry over to deep research settings, different optimizers, and multiple model backbones.

DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration

Weihan Peng, Yuling Shi, Yingwei Ma, Longfei Yun, Beijun Shen, Xiaodong Gu cross-listed Answering developer questions about a software repository requires reasoning across multiple files, architectural layers, and long-range code dependencies, but existing repository-understanding methods mostly rely on surface-level code retrieval. DeepRepoQA is an agentic question-answering framework in which large language model (LLM) agents perform a systematic tree search over the repository structure, using Monte-Carlo Tree Search (MCTS) to dynamically search, navigate, and inspect code for multi-hop reasoning. Experiments on the SWE-QA benchmark show substantial gains over strong baselines, which the authors attribute to the MCTS-guided exploration.

SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction

Xue Hu, Zewei Pan, Zeli Su, Zhou Liu, Wentao Zhang LLM agents that generate paper reproduction code often produce implementations that silently diverge from the paper's specification, a failure mode the authors call semantic drift. SA-Bench (SemanticAlign-Bench) covers 30 papers from ICLR, ICML, and NeurIPS 2025, decomposing each into atomic, verifiable Semantic Alignment Units (SAUs), 1,491 in total across five machine learning domains, and scoring repositories along numerical, methodological, protocol, and ordering drift. Across 12 generator configurations (4 models × 3 scaffolds), even the best, Claude with PaperCoder, reaches a mean SAU score of only 0.301 out of 1.0, with an overall mean of 0.221; the failure taxonomy shows agents attempt most requirements but implement them incorrectly, with implementation mismatches and stubs dominating zero-scored claims, and scaffolds tuned for executability offer little help with scientific fidelity.

ReproAgent: Contract-Guided Paper-to-Code Reproduction

Xue Hu, Zewei Pan, Zhongyuan Wang, Zhou Liu, Zeli Su, Wentao Zhang Turning a research paper into an executable repository is hard because the specification is split: explicit content such as algorithms, metrics, and artifacts gets lost over long agent trajectories, while implicit details like framework defaults and conventions inherited from related work never appear in the paper. ReproAgent is a four-stage Prepare-Plan-Generate-Repair pipeline organized around a persistent implementation contract with two channels, one converting paper snippets into code obligations and another retrieving content and structural evidence from related repositories, both bound to work packages and projected into file-level contracts consumed during generation and repair. On PaperBench Code-Dev it reaches the highest mean score among same-backbone scaffolds under both Claude-Sonnet-4.5 and Gemini-3-Flash, with channel ablations and per-paper cases supporting the contribution of each channel; code and artifacts are public.

Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research

Eran Hirsch, David Wan, Han Wang, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan Deep research (DR) systems chain multiple agents that search and synthesize web content into cited reports, but their citation recall is poor and it is unclear which agent in the pipeline corrupts content or citations along the way. The authors test each agent invocation locally for faithfulness and verifiability against its own inputs and classify errors into four types: hallucination, uncited input reliance, uncited output, and insufficient citations. Applied to three top open-source DR systems, the audit finds that nearly every agent except single-document summarizers makes frequent errors, and 84.7% of final-report errors in AI-Q originate at the orchestrator, mostly as citation mistakes. Two simple interventions guided by these diagnostics raise citation recall by 5% without hurting output quality.

Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

Anupam Purwar, Shashank Singh, Kritika Srivastava Scaling up evaluation of conversational voice agents requires knowing which quality and safety attributes an automated judge can score reliably. The study compares human ratings with GPT-4.1 and GPT-5 acting as LLM-as-a-Judge on telecom and retail voice-agent conversations, scoring the same interactions under three evaluation configurations to test sensitivity to setup and judge model, and examining metric-level correlations, evaluator consistency, and systematic human-LLM disagreement. LLM judges are a workable component of large-scale voice-agent assessment, but their reliability varies by metric and configuration rather than being uniform, supporting hybrid pipelines in which humans retain the metrics that demand contextual interpretation.

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman Long-running coding-agent tasks invite switching models mid-run, escalating to a stronger model when a cheaper one stalls or downshifting once the hard reasoning is done, but the receiving model must then continue a trajectory it did not produce. Using low-cost and high-cost model pairs from the Claude and GPT families, the study varies handoff direction, timing, and interface (full-trajectory transfer, compaction, or dropping the trajectory while keeping repository state). Full-trajectory escalation recovers less than half of the quality gap between the cheap and strong models while adding a substantial cost premium, a penalty the authors call the handoff tax, whereas downshifting offers a favorable cost-quality point; notably, stripping the weak model's trajectory improves escalation, while removing the strong model's trajectory hurts downshift quality.

Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems

Yarden Bakish, Amir Dudai, Roy Ganz, Oren Nuriel, Elad Ben Avraham, Mor Shpigel Nacson et al. Diagnosing why a multi-agent LLM system failed still relies on human engineers, who in practice navigate structured observability views rather than reading raw logs end to end. Adaptive Influence Graphs (AIGs) give LLM diagnosers the same affordance: a two-stage agentic framework first converts a failed trace into a structured graph of components, actions, and dependencies, then lets an agent traverse it to locate the critical error. Across multiple models, richer trace representations consistently improve attribution, and adaptive graph construction with agent-directed traversal sets a new state of the art on Who&When, the standard benchmark for multi-agent failure attribution.

From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use

Rongfeng Guo, Yinxuan Huang, Yusen Wu, Maoqing Zhong, Yunlu Chen, Meng Tang et al. Direct function-calling and ReAct-style agents learn state tracking and action generation in one autoregressive stream, so the pressure to emit the next tool call can overwrite or ignore information accumulated earlier in a multi-turn interaction. OODA-Tool, modeled on Boyd's Observe-Orient-Decide-Act cycle, is a typed closed-loop policy that routes each decision through controller-checked intermediate states: Observe reconstructs the task state, Orient decides whether execution is warranted, Decide forms an admissible action structure, and Act produces the external output. Evaluated with Qwen3 models from 0.6B to 14B across multi-turn, multi-tool, and incomplete-information settings, it consistently improves task success over function-calling and ReAct baselines, with the largest gains on smaller models and on tasks that depend heavily on prior turns and tool results.

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li, De-Chuan Zhan, Han-Jia Ye Agents built on large language models (LLMs) increasingly call several tools per task, where parallel execution cuts latency but resource-agnostic parallelism causes avoidable overflows, and existing benchmarks mostly measure tool selection and argument generation under serial execution. PeakBench is a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles, paired with a two-part evaluation framework that separates logical dependency planning from physical resource-constrained scheduling and gives each its own metrics to make failures attributable. Strong logical planning does not reliably translate into safe or efficient execution under resource constraints, while exposing resource information to the agent reduces avoidable overflows and improves utilization.

Discovering Adaptive Transmission Programs for Collective Innovation

C\'edric Colas, J\'er\'emy Perez, Eleni Nisioti, Akhilesh Mocherla, Pierre-Yves Oudeyer, Cl\'ement Moulin-Frier et al. Collective intelligence depends on transmission processes, and prior work has modeled them as networks that fix who shares with whom and when but cannot condition on what agents actually know or on the collective's state. The authors formalize transmission protocols as state-aware programs that route information and resources based on agent and collective states, and use LLM-guided evolutionary search to discover effective protocols for a collective discovery task. Evolved protocols improve collective performance by up to 37% over standard baselines from the literature, ablations show that removing content-dependence while keeping topology and timing erases the gain, and the protocols transfer across domain variations and agent populations.

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

Yiheng Sun, Huifei Wang, Yancheng Zhu, Zhenyu Li, Zebin Zhao, Yifan Yuan Multi-role large language model workflows pass state downstream as summaries, plans, tickets, and handoff notes, and the question studied here is whether a constraint that must be satisfied survives that rewriting as a binding requirement rather than mere background information. Using safety blockers with explicit prerequisite, authority, fallback, and consequence fields, the authors ran 1,296 controlled synthetic episodes in which upstream identification was always correct and only the handoff transformation varied, then measured whether a downstream executor still respected the blocker. Ordinary summarization deactivated the constraint in 100% of episodes and produced forbidden actions 54.2% of the time, while restoring all four state fields brought preservation to 100% and forbidden actions to zero; adding downstream verification eliminated forbidden actions but left 95.3% of artifacts still semantically deactivated, showing that mentioning a condition is not the same as transmitting its force.

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

Lihang Zeng, Shaoting Zhang, Xiaofan Zhang Most large language model diagnosis systems map a finished case description to an answer, whereas real clinicians actively gather evidence and decide when they have enough of it. EviDx builds interactive per-patient environments from raw clinical cases through a component called E-Synthesis, organizes role-specialized agents and evidence tools around an evolving evidence state, and uses an observer-guided runtime harness that tracks uncertainty and evidence coverage to decide when to stop investigating. A three-level evaluation pyramid separates execution robustness, reasoning dynamics, and final diagnostic accuracy, and the reported experiments show gains in both diagnostic performance and process stability, with the size of the gain depending heavily on the underlying model.

Joint Optimization of Tool Creation and Use for Large Language Model Agents

Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee Tool-augmented language models are limited to the APIs someone wrote for them, and existing tool-creation systems prompt a frozen model at inference time, so the component writing a tool schema never learns whether it can actually call that schema. SMITH (Schema-grounded Multi-task Iterative Tool Honing) trains creation and use in one reinforcement learning policy, alternating build rollouts that write a tool from a few examples with use rollouts that invoke a pooled tool on held-out questions, and scoring schema validity, code correctness, and task outcome as three separate reward axes. A 4B Qwen3 trained this way on 13 verifiable procedural reasoning tasks reaches 79.8 macro-average accuracy on held-out tasks, beating an untrained 30B-A3B tool-writer, plus 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA without any tabular or visual training data; its generated tools also improved results for LFM-2.5-350M and Qwen3-30B-A3B.

A Literate Programming Environment for Human and Machine Agents

Adam T. Burke cross-listed Literate programming interleaves prose with code, and the argument here is that the same structure serves language model coding agents by putting each fragment of code next to the natural-language and structured-data context that explains it, which uses a limited context window better than scattered files do. The described environment supplies a grammar for executable program essays, a parser that treats names as first-class objects, an internal name graph linking prose to names to executable artifacts, and bindings to existing languages and test tools. The result gives coding agents symbol-aware search and usage lookup comparable to what an integrated development environment offers human programmers, demonstrated by a working implementation bound to three established languages with several example programs.

Confident at the moment of action: belief miscalibration in LLM play under hidden information

Bhushan Kashinath Joshi Agentic systems increasingly gate actions on a model's own stated confidence, which presumes that confidence tracks correctness at the moment of acting. The test bed is a hidden-information chess variant in which royal status can be secretly and repeatedly moved between pieces; every turn the agent reports a probability distribution over the opponent's hidden royal piece, separately from its chosen move, and this is scored against ground truth recovered after the game. Captures made at stated confidence of 0.5 or higher about the hidden piece's location were correct in only 1 of 62 cases, and these events account for over 98% of the calibration deficit in both an original batch and a replication. The same pattern appears in weaker form across four further model configurations from a second provider, a deliberation-budget change alone shifts the metric nearly as much as a large cross-model gap, and conventional axes such as legality, cost, latency, and completion rate can dissociate entirely from belief quality, which outcome-only evaluation would miss because a miscalibrated model can still win the game.

Meta$^n$: Recursive Self-Improvement through Emergent Depth

Zae Myung Kim, Young-Jun Lee, Seungyeon Jwa, Dongyeop Kang Self-improving LLM agents typically refine their answers rather than the process that produces them, and systems that edit their own machinery must hold part of it fixed to stay stable, capping realized meta-depth at roughly two. Meta^n instead keeps a fixed meta-operation Ω and recurses on its input: Ω reads the traces and code of the solver stack below and writes the next layer as a strategic pre-process plus a library of callable helpers, so depth is set by convergence rather than in advance, and an evolutionary archive searches over layer chains. Across two backbones it outperforms prior self-improving agents on all eight benchmark families, and on ARC-AGI-2 it is the only method to score above zero. Ablations attribute most of the recursion gain to the conditioning each layer passes to the next, with distinct layer roles emerging at depth without being prescribed by any prompt.

SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents

Shidong Yang, Ziyu Ma, Tongwen Huang, Xucong Wang, Renda Li, Yiming Hu et al. Large language model (LLM) agents trained with reinforcement learning (RL) usually remain episodic, and prior skill-based methods such as SkillRL extract skills from trajectories into an append-only bank without checking whether stored skills still work. SkillForge makes skill invocation an explicit action during interaction so RL optimizes both environment actions and skill-use decisions, and adds evidence-based skill verification plus multi-pathway skill induction so the skill bank keeps growing while maintaining quality. On ALFWorld, WebShop, and AppWorld, SkillForge consistently outperforms SkillRL, which the authors attribute to continuously verified skills.

Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav

Hongyu Guo, Zhiyu Zheng, Zhao Cao When LLM agents interact directly with a corpus instead of relying on conventional retrieval-augmented generation, reachable evidence can silently go unused under finite interaction budgets: it fails to surface, surfaces but is never opened, or is opened without exposing the decisive fragment. The authors name this Evidence Blindness, quantify it through stage-wise evidence realization, and reframe agentic search as finite-budget navigation over reusable corpus structure with AtlasNav, which organizes the corpus once into a persistent multi-view Corpus Atlas rather than rebuilding query-specific workspaces online. On BrowseComp-Plus it reaches 92.05% strict accuracy while cutting recorded online inference cost by 30.21% relative to the prior dynamic-workspace state of the art, realizes required evidence earlier under matched budgets, and holds up on PhantomWiki across 10K–1M scaling and on heterogeneous enterprise knowledge.

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Boyang Liu, Senjie Jin, Peixin Wang, Zhangyue Yin, Yibo Wang, Yuhao Zhou et al. Search agents trained with outcome rewards learn when and how to retrieve evidence, but terminal rewards cannot localize intermediate mistakes or redirect a trajectory before errors compound. CAFE (Coupled Agent–Feedback Evolution) has a single shared-parameter model alternate between search-agent and critic roles: it bootstraps feedback-conditioned recovery from the base agent's own failures, then couples online RL, where a prompt-level call-versus-skip success gap shapes the return for requesting feedback and feedback-aware advantage shaping reweights token advantages before and after feedback, with offline preference optimization that learns feedback from matched successful and failed rollouts. Across seven agentic search benchmarks CAFE outperforms the evaluated RL-based search agents on average and retains its gains on all six out-of-domain benchmarks while reducing answer-level hallucinations, and one-sided ablations show that improving only the agent or only the critic plateaus whereas alternating the two updates keeps improving.

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav, Sai Rajeswar, Patrice Bechard, Sridhar Nemala et al. Agents deployed in tool-rich enterprise environments often fail because of persistent mismatches between the model and the environment's interfaces and conventions rather than raw model capability. StarHarness keeps model weights fixed and instead evolves the harness itself (prompts and task framing, tool interfaces, skills, Model Context Protocol (MCP)-backed providers, subagent structure, and agent-loop configuration), using a compact evolution pool built by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from proposer-hidden selection tasks, and reserving held-out tasks to measure generalization. On ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution lifts full-benchmark performance by 20-35 percentage points over the default harness after only 4-12 accepted changes per environment, with gains that persist on held-out tasks and transfer across GPT and Qwen model families without re-evolution. Trace analysis attributes the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, yielding fewer false-positive diagnoses and shorter trajectories.

Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch

Rima Hazra, Sayan Layek, Somnath Banerjee, Soumen Chakrabarti, Animesh Mukherjee Deep research agents for literature search run open-ended search loops that are expensive and hard to inspect. Crase replaces the loop with a fixed procedure: a single search-engine query for seed papers, expansion along the seeds' 1.5-hop citation neighborhood, pruning of citation edges whose claims lack entailment support, and ranking of the surviving papers with a recency-aware random walk, so the candidate set, the reason each paper is kept, and the stopping condition are all explicit before inference begins. On LitSearch and a second benchmark over a 500K-paper arXiv corpus, it reaches up to 3x the recall@50 of deep research agents built on proprietary models at roughly a third of the cost.

BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes

Fei Tang, Huawen Shen, Zhiqiong Lu, Zhengxi Lu, Pengyuan Lyu, Chengquan Zhang et al. Web agents that act from screenshots avoid the fragility and token cost of reading HTML or accessibility trees, but training them needs large volumes of interaction trajectories, and existing datasets and synthesis pipelines are confined to a few thousand trajectories from a narrow, fixed set of websites. BrowserForge drives hundreds of concurrent browser sandboxes over the open web, combining an open-web sourcing stage that reaches hundreds of thousands of real sites, a cluster manager that keeps browser utilization high, and a Proposer-Solver dual-agent loop that turns a raw page into an executable task and collects a verified trajectory; a rule-plus-model cleaning pass discards failed runs and rewrites reasoning into a unified chain-of-thought style, and the accessibility tree is used only at synthesis time so the released agent acts purely from pixels. The resulting corpus holds 203,238 trajectories, each from a distinct website, and fine-tuning a compact multimodal model on it raises live Online-Mind2Web success from 25.66% to 33.33%, with step accuracy on Multimodal-Mind2Web also improving and gains growing as the corpus scales, which controlled analyses attribute to open-web sourcing and broad site coverage.

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang et al. Recursive self-improvement (RSI) is difficult in long-horizon tasks because ever-growing interaction histories obscure the current task state and cause skills to be invoked at the wrong moments. Recuris pairs a Working Memory that tracks task progress with an Experiential Memory of skills, so skill selection is driven by current needs rather than the full history, and this coupling turns each execution into structured evidence that localizes failures to specific memory components; a fixed Meta-Agent then converts that evidence into localized, validation-gated updates to Skill Memory, closing a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, it improves task success in 35 of 37 completed model-benchmark pairs, adding +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5 on tau-bench (taking Opus 5 to 87.9%) and +16.6/+13.5 points to Qwen3.6-27B/35B on SkillFlow. The advantage widens with interaction horizon, reaching +32.2 points on the longest tasks, and common long-horizon failure modes drop by up to 80%.

Large Language Models 54

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

Yuan Si, Simeng Han, Daming Li, Jialu Zhang Memory and retrieval-augmented generation (RAG) evaluations usually treat what the answering model actually sees as an implementation detail, even though the same conversation history can be rendered as a memory entry, a summary, a typed record, or raw text. RENDER is a benchmark control that fixes the conversation while varying this reader-facing artifact, using a five-level packet ladder to localize when answer-bearing content enters the input plus deterministic templates mimicking ChatGPT-style entries, LangChain summaries, MemGPT-style typed records, and raw dialogue. On 500 LongMemEval questions across nine models, matched-budget resolved packets beat recency-truncated raw dialogue by 42.4 to 72.6 points, deployed-style templates show best-to-worst spreads of 24.6 to 48.8 points per model, and three models scoring 0% on formal ledger packets answer the same facts from natural-language entries at 45.4 to 53.4%. The effect persists under retrieval noise and transfers to HotpotQA, suggesting memory and RAG evaluations should report or control the rendered artifact.

ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik Natural Language to SQL (NL2SQL) systems report over 89% execution accuracy on Spider and BIRD, but those benchmarks use simplified academic schemas and open-source dialects rather than enterprise databases. ESQ-Bench is an Oracle-first benchmark with six populated schemas (465 tables, 164,682 rows) replicated identically on Oracle, PostgreSQL, MySQL, and SQL Server, 550 gold-validated question-query pairs across three complexity tiers, and a four-metric evaluation harness that includes exact match, execution match (EX), and silent divergence (SD). Schema-linked GPT-4o degrades monotonically from 79.8% to 60.3% to 57.2% EX across tiers, Claude Sonnet 4.6 reaches 87.4%, 74.9%, and 68.7% and beats it on every tier, and a local Llama 3.2 manages only 13.3% bank-wide EX. Among queries that pass execution match, silent semantic divergence reaches 73 to 99%, with wrong-result semantics dominating failures at higher tiers.

Squeezing the Cache, Preserving the Truth: Monotonic Equipotential Allocation with Geodesia-KV

Vincenzo Dentamaro, Pancrazio Auteri, Giuseppe Pirlo cross-listed Assessments of KV-cache compression typically conflate resident bits with read bandwidth and are distorted by artifacts of chunked teacher-forcing. Geodesia-KV is a family of training-free cache policies built on monotonic block-wise precision allocation, exact rate-distortion residuals, and query-sparse reading, evaluated causally with resident and read bits reported separately. On WikiText-2 at 16k context, its 5-bit operating point achieves lower perplexity at lower bitrate than KIVI-4 on Qwen, and a compressed-Quest variant improves perplexity on PG-19 while cutting resident bits from 16.25 to 9.83 per value and read bits from 2.32 to 1.95 relative to baseline sparse methods. Implemented as a native cache-manager plug-in for vLLM, the method enables 1M-token context generation on a single 16 GiB GPU with up to 71.7% peak VRAM savings across Qwen, Llama, and DeepSeek architectures.

Function-Level Execution Feedback for Code Preference Optimization

Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong et al. Process supervision has helped mathematical reasoning, but in code generation there is no standard notion of a step, so it is unclear whether to label lines, reasoning traces, or program states. STEP-KTODER defines steps as module-level functions in decomposed multi-function programs, assigns each a binary correctness label via automatically generated unit tests, and combines this function-level process supervision with outcome-level feedback on the full program as a code-specific instantiation of stepwise KTO. On HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench it improves over outcome-only KTO and DPO, and execution-based labels prove essential: LLM-as-a-judge annotations systematically over-predict function failures, corrupting positive step labels and degrading downstream preference optimization.

Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap

Sathishkumar Sivashanmugam cross-listed LLM serving engines size their key-value (KV) cache once at startup and permanently reserve memory for the worst-case prefill activation, which sits idle during decode-dominant phases. The authors build an elastic KV cache that lends this reserve to the KV pool during decode and returns it before prefill, driven by the scheduler's one-step-ahead view of the next batch, implemented purely in userspace on the CUDA virtual-memory path with two physical handles mapped into one contiguous virtual range per layer so the attention kernel is unchanged; it decommits in milliseconds, works with CUDA graphs and prefix caching, and never triggers out-of-memory. They then report an honest negative result: median time-to-first-token differs by only about 1% between chunk sizes of 8192 and 32768 tokens because prefill is compute bound, so simply lowering max_num_batched_tokens recovers more KV than the controller at nearly equal latency, and the reserve dilutes from 16% of KV at tensor-parallel degree 1 to 2.7% at degree 4. The mechanism is released as a reusable userspace elastic virtual-memory allocator.

From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

Amit Kumar, Elnur Adl Zarabi, Suranjana Trivedy, Zhiqian Chen, Lei Zhang, Kaiqun Fu et al. Large language models (LLMs) are increasingly used to supply prior causal knowledge for structural causal discovery, but whether their direct-edge judgments and confidence can be trusted is unclear. The study evaluates 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources (verbalized, logit-based, cross-prompt agreement, and cross-model agreement) under a language-only pairwise protocol. Models are strongly recall-dominant, predicting overly dense graphs, and they capture causal relatedness without reliably identifying directness or orientation: 40.0% of indirect and 36.0% of reversed non-edges are misclassified as direct edges, and over 80% of those false positives carry verbalized confidence of at least 80%. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement calibrate better though not significantly after Holm correction, and a familiarity audit flags five model-dataset pairs, all involving AsiaM.

The Limits of Automatic Evaluation of Creativity in Large Language Models

Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi Evaluating the creativity of text produced by Large Language Models (LLMs) remains unreliable, so this study tests whether automatic methods actually track human judgments. The authors collect human ratings of human- and AI-written short stories from the WritingPrompts dataset along 11 creativity dimensions and compare them against objective automatic metrics and LLM-as-a-Judge scores. LLM judges systematically prefer AI-generated stories, favoring their stylistic traits over the unpredictability and other qualities of human writing, and widely used automatic metrics show near-zero correlation with human judgments for both story types. The results point to fundamental limits in reducing the multidimensional, subjective nature of creativity to computational metrics.

Inter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended Text

Haoyuan Li, Snigdha Chaturvedi LLM-as-a-judge evaluations of open-ended text are usually multi-dimensional, and a reliable judge should score each target dimension without being swayed by the others. The authors propose CorrGap, which quantifies this inter-dimension dependence via differences in correlation between predicted and ground-truth scores across groups of texts, and use it to show that the dependence is pervasive across LLM judges and tasks. To mitigate it they introduce DimCheck, which iteratively strips evidence unrelated to the target dimension from the judge's chain-of-thought step by step; it outperforms strong baselines across three LLMs and four tasks, and smaller trained models can approximate larger ones at much lower inference cost.

Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings

Egor Kolodin, Egor Krasnoperov, Evgeniy Kosarev, Fyodor Minkin Giga-Embeddings is a family of text embedding models built to pair strong retrieval quality with efficient serving, headlined by a sparse 10B-parameter Mixture-of-Experts encoder with about 1.8B active parameters per token, alongside a dense 3B encoder and a distilled 480M encoder trained with a dimension-agnostic objective that aligns teacher and student similarity distributions. The MoE model achieves the strongest aggregate scores in the family across English, Russian, multilingual, and code MTEB suites, and in a vLLM benchmark with 1024-token inputs processes 114.5k tokens per second, 25% more than the dense 3B model and 1.56-2.65x the throughput of external systems. The 480M model scores 70.98 on Russian MTEB, surpassing FRIDA with 42% fewer parameters, and all three checkpoints are released.

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

Farhana Amin, Sabiha Afroz, Mona Moghadampanah, Dimitrios S. Nikolopoulos Masked diffusion language models (dLLMs) can denoise many tokens at once, but serving systems for them have been built without measuring how they behave under real concurrent load. The authors characterize serving of LLaDA-8B-Instruct with a Discrete Diffusion Forcing (D2F) LoRA adapter on a single NVIDIA H200 using GSM8K and HumanEval, finding that request difficulty falls into 11 discrete denoising-step levels that no tested signal predicts in advance (best R2 = 0.150), that generation budgets under 320 tokens hide the latency spread, and that only 24% of single-request wall-clock time is GPU computation, with the rest being CPU-side dispatch overhead. Batching mainly amortizes that overhead, giving a 16.0x throughput gain at batch size 16 over per-request dispatch, and the paper derives a batch-timeout rule for synchronized batching under Poisson arrivals, arguing that dLLM serving requires parallelism at each denoising step.

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

Aman Saini, Priyanshu Kumar, Eric Peng, Kai Yuan, Harsh Girase, Wanming Chen Reward design for open-domain question answering is hard because good answers must satisfy several quality aspects at once that a single scalar reward captures poorly. The authors propose generating query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, and using them as fine-grained reward supervision during post-training. Averaged over composition, grounding, and instruction-following axes, the approach improves 6.5% over the instruction-tuned baseline and 4% over flat rubric variants, with evidence conditioning boosting factual support and dimensional decomposition improving coherence, organization, and adherence to query requirements.

AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning

Md Romyull Islam Quantized fine-tuning with QLoRA saves memory but trains slower than fp16 LoRA because every 4-bit weight is dequantized on the fly. AQLoRA (Adaptive-Quantization LoRA) recovers part of that time with a single CPU pass that ranks layers by NF4 reconstruction error and keeps the top-K in fp16 under a memory budget, with no search or calibration data, reproducing Unsloth's hand-curated dynamic-4bit selection exactly. On Commonsense-170K across six models from 1.4B to 14B, the speed setting trains 11.1 +/- 2.7% faster than well-tuned QLoRA at a cost of about one accuracy point, while the quality setting is 4.8% faster with accuracy level with QLoRA; the authors also report timing-methodology lessons from shared hardware and two failed controls showing that the count of protected layers, not their identity, drives the speedup.

Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention

Sergii Kozyrev (Minima AI, Inc), Davyd Maiboroda (Minima AI, Inc) The key-value (KV) cache is the main capacity and bandwidth bottleneck when serving long-context large language models. Minima-KV keeps recent and protected anchor pages in FP8 while moving older pages to a packed TQ3 format, and uses format-specific kernels whose partial attention states are combined through a globally normalized online-softmax merge, so decoding runs directly over heterogeneous pages without a dense shadow copy and without evicting any live-request page. On Qwen3.6-27B profiles running on a single 96-GB RTX PRO 6000 GPU, it reports 18.3 KiB of attention KV per live token, a 3.50x compression over BF16 and 1.75x over FP8, with a quality profile that matches its dense control on 16K RULER needle-in-a-haystack tasks and LongBench v2 deltas of -0.80 to -0.40 percentage points at 16K to 64K context.

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

Tom Poperszky cross-listed Single-token autoregressive decoding on CPUs is limited by memory bandwidth rather than compute, since a modern CPU sustains roughly 1 TFLOP/s but only about 50 GB/s from main memory and every active weight must be streamed once per token. The report co-designs a family of pipeline-native transformer architectures, whose inter-layer dependency graphs permit a vertical, stage-major execution schedule, with cflow, a CPU-first streaming engine that stores weights as L2-sized tiles in consumption order, reads only the top-k experts of each mixture-of-experts layer, fuses projections, and runs a delay-aware schedule. On a 30.9-billion-parameter pipeline-native MoE, cflow decodes at 5.94 tokens/s on a 32-vCPU Ice Lake server versus 4.75 for llama.cpp and 1.65 for the vLLM CPU backend on comparably sized dense models, with asynchronous I/O overlap for disk-resident experts adding up to a further 1.68x; measurements refute one of the eight design claims and leave another inconclusive, both reported in full.

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

Zizhong Wang, Jieying Wang, Zhao Zhang, Jiajia Li Long-context inference in large language models (LLMs) is increasingly constrained by key-value (KV) cache memory, and existing low-rank compression methods derive fixed projection spaces from weights or calibration data or share one basis over a broad cache region, which can miss fine-grained but important information. PuzzleKV splits each per-head KV cache into fixed-length logical pages, observes substantial low-rank structure within individual pages, and, without training or calibration, decomposes each completed page independently, computing attention directly over a mix of dense and factorized pages while incrementally compressing newly eligible pages during decoding. At roughly 60% of the original storage it retains more than 96% of full-KV performance across both evaluated models and all benchmark settings, with substantial gains over global SVD on RULER and competitive results on LongBench, and combined with quantization it keeps over 93% of full-KV performance using only 18.7% of the original storage.

Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching

Sunwoo Kim cross-listed tablet-2 is a production long-term memory engine for language models whose retrieval path contains no lexical matching, keyword scoring, or language model of its own; the authors measure it on standard text memory benchmarks and on cross-lingual retrieval of photographs stored with no text at all. It scores 95.7% on LongMemEval-S and 67.5% on BEAM-1M (2.21 million stored memories), but much of the paper argues such numbers mean little alone: with engine, corpus, settings, and judge fixed, changing only the reader model shifts LongMemEval-S by 2.0 points and changing only the re-ask budget shifts BEAM-1M by 8.9, neither of which competing reports state, so the comparison table is offered as placement rather than ranking. On captionless photographs across 70 store-and-query language cells it reaches 95.2% mean recall@5 where a strongly configured BM25 reaches 19.0% and is exactly zero in 54 cells, and open dense baselines on Crossmodal-3600 show density confers no language independence, with one model scoring 91.0% on English but 4.7% on Russian from identical image vectors. Results that cut against the system are reported at equal weight: low-resource languages degrade sharply (Swahili 53.0%, Telugu 64.0%), attaching captions lowers cross-lingual retrieval by 11.4 points, and one omitted setting cost 37 points of Korean top-1 accuracy while leaving nine other languages untouched.

Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

Yicheng Mao, Hongru Du Deciding how to allocate a fixed pretraining token budget across data domains is commonly handled by training small proxy models on candidate mixtures, fitting a response model, and extrapolating to larger runs. The authors observe that this workflow has the exact structure of a classical mixture experiment, with domains as components, token shares as proportions, proxy runs as design points, and validation loss as a response surface over the probability simplex, which lets them apply sparse second-order Scheffé response-surface models and construct model-robust I-optimal designs for the proxy experiments. Using RegMix as a case study, the Scheffé analysis shows domain value is strongly relational, with several domains that look weak under additive effects becoming favourable through pairwise interactions, especially with web-derived text, while the sparse model preserves mixture rankings across scales and stays competitive with a flexible machine-learning predictor. In a simulation calibrated to observed proxy responses, I-optimal designs recover the relevant mixture ordering after removing about 25% of the original proxy runs, suggesting data mixing should be treated as an experimental-design problem and not only a prediction problem.

Evaluating Language Models on Cross-Language Code Functional Equivalence

Hui Sun, Anderson Uch\^oa, Rohit Gheyi, Wesley K. G. Assun\c{c}\~ao cross-listed Existing evaluations of whether large language models (LLMs) understand program semantics mostly use single-language settings or synthetic code, so the authors ask whether models can judge functional equivalence across programming languages in human-written code. They introduce PolyHuman, a dataset of human-written programs in C++, Java, and Python, evaluate intra- and inter-language equivalence detection across open-weight and proprietary models, and manually analyze 81 cases of systematic disagreement, including the generated chain-of-thought, comparing failure categories across GPT-o4-mini, Claude-Opus-4.7, and Gemini-3-Flash. They find a difficulty-dependent breakdown in which harder problems make models increasingly prone to misclassifying non-equivalent code as equivalent, a model-specific language sensitivity (notably more conservative behavior on Python for the best-performing model), partial reliance on surface-similarity cues, and substantial run-to-run instability for GPT-o4-mini under identical settings, concluding that current LLMs do not reliably capture functional equivalence within or across languages.

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal, Jay Iyer, Sai Pavan Kumar, Sai Kapil Kumar et al. When an LLM serving deployment runs out of key-value (KV) cache memory, operators can either shard weights and cache across more GPUs with tensor parallelism, paying an all-reduce per layer and a larger hardware bill, or shrink the cache in place with KV quantisation and eviction at some cost in quality, yet the two are rarely compared on a common cost axis. Using a profiled simulator calibrated on A100, A40, and H100 hardware, the authors place tensor-parallel degrees 1 to 8 and KV-compressed configurations (16/8/4-bit, keep-ratios down to 0.25) on one axis of cost per million tokens against latency for Llama-2 at 7B and 70B and search for a cost-equivalence crossover. They find none: compression is cheaper by 1.20x to 2.00x at every level of memory relief they could construct, and the deciding boundary is model size relative to device memory, roughly 36B parameters for an 80 GB card, below which extra GPUs are largely wasted spend and above which tensor parallelism becomes an entry ticket because the binding resource is weights, which KV compression does not touch. Tensor parallelism is the only lever that improves latency (compression worsens per-token latency by 8 to 93% through batching contention), while compression is the only lever that multiplies capacity per dollar (16.5x versus 1.21x for an eightfold GPU spend).

Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction

Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen Using several large language models (LLMs) as a wisdom-of-the-crowd forecaster only helps if the models actually behave differently, and simply adding more of them can pile up redundant reasoning. The authors propose a behavior-aware framework that characterizes each model by its reasoning traces on independent development tasks, clusters models by behavioral similarity, and picks cluster representatives to form the crowd. Evaluating 25 LLMs on seven development benchmarks and two future-prediction benchmarks, they find that a three-model medoid crowd chosen by K-means++ behavioral clustering beats voting over all 25 models on both prediction benchmarks while cutting model calls by 88% and inference cost by roughly 80%. The results suggest representative behavioral diversity, rather than maximal diversity, is what makes an LLM crowd effective.

Mechanistic Circuit Identification for Controllable Data Generation

Nakyung Lee, Sangwoo Hong, Jungwoo Lee Synthetic data pipelines mostly steer generation through heuristic prompts, offering little insight into how individual samples interact with a model's learning dynamics. The framework connects training-dynamics-based data valuation with mechanistic interpretability (MI) by defining three utility axes (learnability, challenge, and alignment), identifying model-internal circuits that causally govern each, and using those circuits as steering interfaces to generate utility-targeted data; SAMS (Stage-Aware Mechanistic Scheduling) then schedules circuit-steered data according to the model's evolving optimization needs. On multiple-choice question answering tasks, the approach produces more precisely controlled and more diverse data than prompt-based baselines, with consistent gains in downstream accuracy and calibration.

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

Mohammad Mozaffari Sparsity, quantization, and low-rank approximation are usually applied to large language models (LLMs) in isolation, and each alone hits an accuracy-efficiency wall. The thesis applies all three jointly, using sparsity to cut computation, quantization to cut memory bandwidth, and low-rank adapters to recover accuracy, across a series of methods: MKOR reduces curvature-update complexity from cubic to quadratic and converges up to 1.85x faster than KFAC, SLoPe speeds training up to 1.25x with a double-pruned backward pass for N:M sparsity plus late low-rank adapters, OPTIMA improves zero-shot accuracy by up to 3.97% via globally optimal column-wise quadratic programs for weight reconstruction, and PATCH learns a dynamic hybrid sparsity ratio for up to 1.38x speedups. SLiM realizes the full combination in one shot, improving accuracy by up to 5.66% over prior methods and outperforming uncompressed dense models by 0.6% at equal parameter budgets.

TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis

Boshen Shi, Yize Liu, Chen Zhao, Ce Chi, Zhendong Wang, Xing Wang et al. Large language models (LLMs) increasingly analyze spreadsheets and CSV files, but a plausible-looking answer is not the same as one supported by a valid path from the question to the underlying data evidence. TrustDABench operationalizes this as two properties, reliability (refusing or asking for clarification when no valid evidence path exists) and robustness (preserving a correct analysis when the same evidence is expressed in different table forms), derived through 19 perturbation operators applied by an agentic LLM generation framework to produce 2,340 human-verified instances. Across eight LLMs the headroom is large: the best reliability result averages only 24.21% MRS (GPT-5.5), and the best robustness result still shows a 9.10% average ASR (Claude-Sonnet-5), with models rarely detecting conflicting evidence, continuing along executable but unsupported analysis paths, and remaining sensitive to perturbations that change observation boundaries or cross-table relations.

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

Ryuichi Sumida, Koji Inoue, Tatsuya Kawahara Memory systems for conversational large language models are typically judged by Direct QA, asking whether the model can recall a specific fact from earlier dialogue, but a 4-month deployment with 40 users, 1,872 sessions, and 7 memory conditions found that Direct QA accuracy ranging from 19.7% to 70.1% produced no change in user satisfaction. The proposed explanation is that benchmarks measure elicited retrieval while real conversation requires natural integration, detecting when prior context is relevant and weaving it into a response unprompted. MemUse is a set of real user-cued memory moments drawn from the deployment and scored with an integration-aware judgment of the natural conversational reply; holding model and context fixed, a system scoring 78.8% on Direct QA references only 7.9% of those facts in conversation, a 71-point gap, and within these moments natural integration correlates with satisfaction while Direct QA does not. The deployment corpus, benchmark, judgments, and scoring prompts are released.

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

Minsu Kim, Jianxun Lian, Xing Xie, Steven Euijong Whang Aligning large language models to human preferences often incurs an alignment tax, the catastrophic forgetting of pre-trained general capabilities, and prior work has treated this as an optimization or architectural problem rather than asking which preference samples cause the damage. BALIGN is a data selection strategy that, through analysis of the preference optimization gradient, identifies three data-centric drivers of parameter drift (the reference model's log-probability margin, the token-length difference between chosen and rejected responses, and TF-IDF similarity to general-capability corpora), combines them into a composite risk score, and filters out high-risk samples that disrupt model parameters or contribute little alignment utility. On standard human preference datasets, the filtered data preserves foundational capabilities without sacrificing alignment gains, consistently reaching the best Pareto frontier with minimal computational overhead.

Evaluating Multiple LLM Generations with Validated Task Coverage

Florian Le Bronnec, Rio Yokota Many LLM applications are most useful when they return several candidate outputs for comparison or combination, yet standard evaluation scores single outputs or collapses multiple samples into one success or selected answer, missing whether the set contains genuinely different useful results. VTC-Bench is a five-domain benchmark built from real-data tasks where both quality and task-relevant distinctness can be checked automatically without model-based judges, with Validated Task Coverage (VTC) measuring how many distinct useful results appear within k attempts. Across models and inference settings, configurations that look strongest on single-draw quality are not necessarily those with the best coverage, and simple output-variation measures fail to recover task-relevant coverage, showing that finite candidate sets expose behavioral differences that per-output evaluation hides.

RecurSE: Bounded Recursive Self-Evaluation for LLM Rubric Judges

Kaiyuan Liu, Ziyuan Zhuang, Rongxiang Weng, Jieping Ye Improving an LLM-as-judge typically depends on expensive annotations, reward models, or distillation from stronger teachers. Recursive Self-Evaluation (RecurSE) removes external gold supervision from the reinforcement learning reward: a trainable judge scores candidate responses under per-rule rubrics, while a synchronized copy of the same policy audits the judge's reasoning against meta-rubrics to produce a scalar process reward, with interface decoupling isolating that score from the judge's verdict tokens to block a token-copying shortcut that inflates self-assigned rewards. Because unanchored recursive learning is inherently bounded, Pairwise Advantage Validity (PAV) serves as an unbiased validation monitor tracking judge accuracy and checker fidelity to locate the early-stopping window. Across Qwen3.5-9B, Gemma-4-E4B-it, and Qwen3.6-27B, the method delivers consistent gains on held-out medical, pairwise, summarization, and professional benchmarks, with ablations showing synchronized judge-checker co-evolution beats frozen checkers, external meta-judges, self-consistency, and scaled teacher distillation.

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao, De-Chuan Zhan, Han-Jia Ye Multi-candidate speculative decoding (SD) proposes several draft token candidates to raise acceptance rates when a lightweight draft model's guesses are verified by a large language model (LLM), but the analysis identifies a bottleneck called residual drift: once early candidates are rejected, the residual target distribution diverges from the draft model's predictions, making later candidates ineffective and forcing costly resampling. ResiSpec reshapes the proposal distribution during verification to anchor the residual target mass within the draft model's high-confidence regions, re-aligning the verification process mathematically while preserving exact output distributions. The method reaches up to 1.92x speedup over state-of-the-art multi-candidate methods, with code released.

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong Judge models are typically validated by agreement and robustness to surface perturbations, but that establishes reliability rather than construct validity. The work formalizes validity as a two-dimensional profile: invariance S, the probability a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability it changes under minimal construct-changing edits, and shows the two are independent so no scalar summary preserves all relevant comparisons. Measured across 7 judges and 4 domains with 7 construct-changing intervention types and 5 register-only controls, judges average S = 0.945 but only R = 0.319, with sensitivity to scope edits (0.383) exceeding sensitivity to strength edits (0.262) for every judge. An audit of five public label sets finds surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes, motivating joint reporting of S and R and auditing the validation sets themselves.

Shortcut Before Circuit: Document Statistics Time In-Context Conflict Resolution

Yijun Liao, Fanwei Liang When a context asserts two conflicting values for one fact, a model commits to some cue such as recency, repetition, or position, but natural data rarely makes these cues disagree, so behavior alone cannot reveal which one is used. The study trains 26M-parameter transformers on a synthetic language where recency and rarity are exactly coextensive, then separates them with a minimal causal edit that inverts one cue while holding the truth, token count, and answer position fixed. All 75 runs reach accuracy of at least 0.999, yet under intervention the mechanism does not replicate across seeds: 13 of 25 cells differ by more than 0.3 in sign fraction, and attribution probed before escape from a positional shortcut reverses sign in 32 of 75 runs at unchanged accuracy. What does replicate is timing: escape from the positional shortcut follows a closed-form ceiling monotone in redundancy, so the corpus fixes when a mechanism appears but not which one, giving a criterion for when mechanistic attribution to data is possible at all.

Low-Rank Ternary Adaptation for Fine-Tuning Transformers

Alexandru-Dragos Manolache, Yunqiang Li, Jan van Gemert cross-listed Ternary transformers are extremely memory- and compute-efficient, but existing low-bit LoRA-style methods cannot fine-tune ternary weights directly: they either dequantize the base weights to merge adapters or update only quantization parameters, so the merged model no longer stays ternary. Ternary multiplicative adaptation represents discrete updates such as sign flips or zeroing through a low-rank Kronecker factorization into two small ternary matrices applied element-wise to the ternary weights, which preserves the ternary domain and allows direct merging without dequantization. On six language and vision models, including ternarized LLaMA-3 1B and 3B and a ternary ViT-B/16, the method recovers much of the accuracy lost to quantization and outperforms strong low-bit and ternary baselines, with code released.

Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning

Hang Chen, Jiaying Zhu, Wenya Wang Mechanistic localization guides parameter-efficient Supervised Fine-Tuning (SFT) by first identifying critical parameters with interpretability methods, but interpreting a pre-SFT model is retrospective: for novel tasks, the neurons flagged before tuning differ drastically from those that govern the final model, and the mismatch actively disrupts SFT. The authors propose a forward-looking localization framework that estimates the post-SFT interpretability state using only pre-SFT parameters and the target dataset, modeling SFT as continuous parameter evolution and using a Taylor expansion to connect the post-tuning mechanistic objective to the pre-SFT model's gradients. Dual-granularity pipelines operate at neuron and component level, and experiments show the predicted localization provides better SFT guidance than static pre-SFT interpretation while remaining robust and scalable as model size grows.

When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study

Mohit Singh Chauhan, Vipin Gyanchandani, Dylan Bouchard Uncertainty quantification (UQ) signals are commonly used to flag hallucinations in large language models (LLMs) in closed-book settings, and prior work has combined them with learned ensembles, but how robust those ensembles are had not been studied systematically. The authors train a classifier over heterogeneous UQ scorer outputs on a small labeled set of LLM responses and apply it out of sample without retrieval or tools, analyzing sample efficiency, in-domain transfer, and generation-regime dependence across four LLMs, nine datasets, and three regimes (short-form question answering, long-form generation, and code generation). Supervised ensembles beat the best individual scorer in 30 of 32 settings, with gains from as few as 100 labeled instances, and retain most of their advantage under in-domain distribution shift (23 of 28 transfer settings). Sampling-based black-box ensembles are nearly as effective as full ensembles, while single-generation white-box ensembles add little.

Quantization Effects on Bangla Language Understanding in Large Language Models: A Systematic Evaluation

Ismail Hossain, Nafi Ullah Shafin, Mohammad Abdullah Al Mumin Post-training quantization is standard practice for fitting large language models onto constrained hardware, but almost all evidence about what it costs comes from English benchmarks, leaving open how it behaves on a morphologically rich low-resource language. The authors evaluate Qwen-2.5-7B, LLaMA-3.1-8B, and GPT-OSS-20B at full precision and in three 8-bit formats (GPTQ-Int8, GPTQ-Q8, GGUF-W8A16) zero-shot through lm-evaluation-harness on five Bangla understanding benchmarks including Bangla MMLU, CommonsenseQA-BN, and BoolQ-BN. Responses diverge sharply by architecture: GPT-OSS loses up to 57.35% accuracy on reasoning-heavy tasks under GGUF-W8A16, while Qwen and LLaMA hold steady under GPTQ and occasionally edge past full precision, and BoolQ-BN is stable everywhere — so the model-and-method pairing matters more than bit width.

Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems

Wonung Kim, Hyunmin Choi, Minsu Kim, Jaehong Cho, Yeongwook Kim, Jongse Park cross-listed Simulators for large language model (LLM) serving systems fall behind the pace at which new mechanisms such as agentic workflows and disaggregated serving appear, because each new feature demands an invasive rewrite of a monolithic simulation pipeline. Borg addresses this with a composable simulator infrastructure that expresses the entire serving workflow, including its control decisions, as a unified dynamic graph, plus a harnessed coding agent (the Synthesizer agent) that lowers natural-language feature requests onto that abstraction under simulator-specific guardrails and fidelity validation, so one shared simulator evolves instead of a new one being built per feature. Extensions built this way track a vLLM-based real system with 2.51% average throughput error versus 6.03% for extensions built on existing simulators, and on identical workloads the simulator runs up to 284.96x faster than LLMServingSim2.0 and 23.19x faster than Vidur.

The RAT: A Unified Bayesian Model for RAG Evaluation

Pius von D\"{a}niken, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu Evaluating Retrieval-Augmented Generation (RAG) systems requires understanding how retrieval, abstention, and answer correctness interact and how errors propagate, not just end-to-end accuracy. The paper introduces a Bayesian framework that jointly models these three stages, factorized along the pipeline's information flow, separating whether the user received a correct answer from whether the generator behaved appropriately given the retrieval outcome. Applied to 27 RAG configurations spanning three datasets, three retrievers, and three generators, the conditional decomposition exposes substantial behavioral differences between systems that look equivalent under marginal metrics. The authors also show that retrieval-success annotations are more informative than task-success annotations for estimating policy adherence, give an information-theoretic explanation for the asymmetry, and extend the model to treat LLM-as-a-judge outputs as calibrated noisy observations combined with limited human labels.

Linear Probing Provides Robust and Efficient Detection of Machine-Generated Text

Gerrit Quaremba, Hanqi Yan, Elizabeth Black, Denny Vrandecic, Elena Simperl Supervised detectors for machine-generated text (MGT) versus human-written text (HWT) often degrade out-of-domain (OOD) and require large, diverse training sets. The authors show that MGT and HWT latent representations are linearly separable in a low-dimensional space, attribute this to systematic differences in representation quality, and train two simple linear-probe variants evaluated on four benchmarks against 16 baselines. The probes improve OOD detection by +11 AUC and reach near-peak performance with fewer than 100 training samples, a transferability the authors trace to a shared latent MGT direction that generalizes across settings; the probing vectors also capture a continuous spectrum of machineness, suggesting use for fine-grained estimation of AI-edited text.

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

Zihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian, Kunlong Chen, Zhiqiang Zhang et al. Language model pretraining runs with substantially different learning rates (LRs) and parameter norms are shown to follow nearly identical loss trajectories whenever their effective learning rate (ELR), the ratio of LR to parameter norm, is matched, a phenomenon termed ELR collapse. Across optimizers, architectures, datasets, and model scales, the mean collapse error between ELR-matched runs is typically a few times 10^-3, below the seed-to-seed variation measured in a representative configuration, with normalization design and the timescale of LR-norm variation identified as the main determinants of collapse precision. Controlled interventions show that weight decay and Hyperball shape loss dynamics mainly through the ELR schedules they induce, and replacing LR with ELR lets a fitted functional scaling law (FSL) transfer across norm-control methods and explain the delayed acceleration that norm control frequently produces.

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Miao Liu, Zhizhe Liu LLMs deployed as financial analysts are usually evaluated on what they can retrieve from long disclosures, not on whether the retrieved information actually changes their judgments. Holding a focal firm's information fixed and padding the context with unrelated material from 2,000 to 128,000 tokens reveals a retrieval-integration gap: a risk disclosure's influence on investment judgments falls to the experimental noise floor even while direct retrieval of that disclosure remains accurate, a pattern that replicates across model families and judgment tasks, holds when real disclosures are removed from actual 10-K filings, and is postponed but not eliminated by more capable models. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments, and that workflow architecture decides whether transmission succeeds, since chunk-and-summarize pipelines evict the relevant information whereas a targeted structured restatement placed adjacent to the decision restores its influence.
15 more specialized papers

Other 41

Transformer Accelerator (TFA): A Macro-Op INT8 Hardware Chip for Transformer Inference and Machine Translation

Shashank cross-listed Running pretrained transformers on compact custom silicon requires both an efficient datapath and a way to keep INT8 quantization bit-exact against a reference. The Transformer Accelerator (TFA) is a synthesizable, parameterizable INT8 memory-to-memory engine in which one time-multiplexed datapath handles both prompt processing and autoregressive generation, executing matrix multiplication, softmax, RMSNorm, elementwise, and copy/gather operations through eight 512-bit macro-op descriptors compiled offline and dispatched over AXI interfaces. A UVM verification environment byte-compares outputs against a bit-exact golden model with zero mismatches across 25 tests and 34 constrained-random runs, and a compiled t5-small English-to-French, German, and Romanian translation pipeline executed 70,320 descriptors and matched 37.9 MB of golden output exactly, with randomized-Hadamard reparameterization recovering about 11 dB of per-tensor INT8 signal-to-noise ratio. The verification configuration reached roughly 20x end-to-end speedup over a 22-thread CPU, larger designs are projected to cut energy per token by about 1000x, and the design completed clean synthesis and place-and-route on SkyWater sky130 at 2.73 mm² of logic area.

ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning

Juntao Fang, Shifeng Xie, Ruichu Cai, Shengji Zheng, Zijian Li, Keli Zhang et al. Time series foundation models handle forecasting and transferable representations, but classification still usually means fitting a task-specific classifier per dataset, with each channel of a multivariate input encoded independently. ChorusTIC is a classification-native foundation model that classifies in context, without target-task parameter updates, across heterogeneous channel configurations: episode-consistent Random Subchannel Slot Concatenation feeds a shared dual-axis encoder that models temporal and cross-channel interactions and maps any channel count to a fixed-width representation, after which feature axes are calibrated using context-derived distributions and query labels are predicted through leakage-protected in-context learning. The model is pretrained solely on synthetic labeled episodes whose classes are distinguished by sparse temporal or cross-channel rules, and on the complete UEA-30 and UCR-128 archives it achieves strong full-context and low-label performance without fitting any target-specific classifier.

A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU

Andrew James Amos Self-organizing maps turn a corpus into a browsable two-dimensional atlas, but at MEDLINE scale the best-matching-unit (BMU) search is bottlenecked by the bandwidth needed to read the codebook every epoch. Storing the codebook feature-major, with each feature's weights contiguous, recasts the search as a tiled sparse-dense product that reuses each loaded weight column across a tile of samples; with implementation, precision, and update rule held fixed, this layout change alone speeds BMU search 4.5-8.5x while held-out quantization error matches a cuSPARSE baseline within 0.5%. Combined with a radius-independent box-blur update and a convergence-based stopping rule, SparseBin.SOM trains a converged 64x64 map over 29.9 million MEDLINE articles in about 72 seconds on one 24 GB consumer GPU, and reaches 1,048,576 neurons on a 141 GB H200, reportedly the largest self-organizing map to date; held-out error follows a smooth power law across three decades of map size.

Causal Analysis for Time Series Foundation Models

Mathis Jander, Wouter van Heeswijk, Martijn Mes Moving from bespoke time series models to foundation models concentrates risk, since many potentially high-stakes forecasting applications inherit the biases and failure modes of a single shared model, while also enabling centralized validation. The authors propose a causal analysis framework that intervenes on parameterized synthetic time series generators and measures the resulting change in model output under ceteris paribus conditions, applying it to Chronos-2 and TimesFM-2.5 across six pattern types. Both models show a bias toward overestimating persistence and fail suddenly on regime-switch patterns, and TimesFM-2.5 additionally fails on energy-release patterns, while trend and harmonic oscillation patterns have safe configurations; a review of the original works suggests pretraining data may explain these findings, and the study closes with development suggestions and application-specific model-selection recommendations.

When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data

Sahand Mazrouei Self-supervised learning (SSL) for tabular data is evaluated under extreme label scarcity and test-time missingness, comparing a mask-and-recover pretraining objective against training from scratch, tree ensembles, and three established tabular SSL baselines (VIME, SCARF, SubTab) across 14 classification tasks. SSL beats training from scratch on average and stays competitive with Random Forest (0.8954 vs 0.9015 AUC at 10% labels), but the SSL-versus-scratch gains show high inter-task variance and are not statistically significant (p = 0.626). Contrary to the expectation that imputation-style pretraining helps data with native missingness, SSL improves most reliably on clean datasets and often degrades on datasets with high inherent missingness, though it does yield higher average AUC than scratch training under injected missing-completely-at-random and missing-not-at-random shifts. None of the SSL objectives differ significantly from one another, suggesting the results reflect general properties of tabular SSL rather than one particular pretext task.

From Numerical Simulators of PDEs to Neural Emulators and Back

Felix Koehler Numerical solvers for partial differential equations (PDEs) are a cost bottleneck whenever many fast evaluations are needed, and neural emulators trained on solver output are usually treated as opaque replacements rather than relatives of the methods that generated their data. The thesis argues the two are closely alike: neural architectures mirror classical discretizations, their errors admit the same mode-wise Fourier analysis, and insight flows in both directions, with the solver's multiple roles in the emulator training pipeline disentangled explicitly. It synthesizes three contributions: APEBench, a benchmarking suite for autoregressive neural PDE emulators built on fast differentiable pseudo-spectral solvers in JAX; Progressively Refined Differentiable Physics, a study of how unconverged solvers affect surrogate training; and Neural Emulator Superiority, an analysis of how numerical errors and architectural inductive biases interact.

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan et al. cross-listed Maia 200 is an AI accelerator delivering 10,145 TFLOP/s in FP4 and 5,072 TFLOP/s in FP8 within a 750 W thermal design power, backed by 7 TB/s of high-bandwidth memory (HBM) bandwidth. It exemplifies Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate specialized memories and data-movement engines, shifting the design focus from thread-centric to data-movement-centric execution. The authors propose a Flynn-style taxonomy of data management to situate SDLA against existing architectures and claim significant cost and energy savings for large-scale AI inference workloads.
34 more specialized papers

Theory 37

Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

Tiexin Ding A trained transformer's weight magnitudes follow a two-parameter Weibull distribution whose shape (about 1.2) is stable across layers and models, so nearly all training-induced movement shows up in the scale parameter, which raises the question of what corpus property determines how much that scale grows. Using the bigram conditional entropy of the training data, a statistic computable before training, the authors find across controlled corruption families a learning-rate-conditioned law in which squared scale growth depends on the gap between a matched-budget shuffle baseline and this entropy raised to a 0.59 exponent, an exponent inherited from an independently measured data-side saturation relation rather than fitted to the growth curve. After removing two per-learning-rate coefficients, 23 runs spanning an order of magnitude in learning rate collapse onto a single curve with unit slope (R² = 0.941), and an end-to-end self-validation predicts held-out within-family weight growth with 5.7% relative error. The law holds at model and per-layer resolution across two architectures but over-predicts growth for code corpora, implicating redundancy as a second axis of a broader data-to-weight framework.

Replicable Conformal Prediction

Marios Papamichalis, Regina Ruane, Theofanis Papamichalis cross-listed Two analysts calibrating the same predictive model on independent samples deploy different conformal prediction sets, because the calibration threshold inherits the randomness of the data, which makes auditing, caching, or cross-site approval of a deployed classifier impossible to verify. The authors show perfect agreement is impossible without ignoring the data, then resolve the tension by sharing a single random seed and rounding the calibrated threshold up to a coarse shared grid, which makes the deployed classifier identical across analysts with any desired probability while preserving coverage guarantees, at a quantified cost in set size and calibration data that matching lower bounds show no threshold method can beat. Without a shared seed a fixed grid still confines all analysts to two adjacent classifiers, replicability also blunts gaming by selective recalibration, and experiments on ImageNet outputs, a four-hospital site split, and four language-model families match the theory including the measured sample-cost frontier.

Renormalization Group Flow Matching for Scalable Local Generative Modeling

Kanta Masuki, Yuto Ashida Generative models face a tradeoff between global approaches that capture full structural coherence at high computational cost and local models that are efficient but miss long-range correlations. Renormalization group flow matching (RGFM) uses an exact renormalization group (RG) flow as the probability path, generating data progressively from long- to short-wavelength structure, and exploits quasi-locality and scale separation to prove the flow can be approximated by local velocity fields whose spatial range grows only logarithmically in system size and inverse error tolerance. This enables local generative modeling with patches of size O(ln L) and computational cost nearly linear in system volume. Experiments show local RGFM reproduces long-range correlations well beyond its receptive field on one-dimensional distributions where conventional local flow matching fails, and yields markedly more coherent samples than local flow matching on FFHQ at 64x64 and 256x256.

Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks

Itay Safran For every depth ℓ ≥ 3, the authors exhibit a globally [0,1]-valued, 1-Lipschitz function realized by a depth-ℓ ReLU network of width O(d^4) that any depth-(ℓ−1) network with unrestricted weights and width at most 2^d/[2d(ℓ−2)] must approximate with squared L2 error at least 1/24 under an absolutely continuous distribution. This is claimed as the first exponential separation between two fixed depths where the shallower depth is at least 3, and the first exponential hierarchy across all adjacent fixed depths; the ℓ = 3 case answers a question posed by Safran, Eldan, and Shamir (2019) on depth-3-versus-depth-2 separation with unrestricted shallow weights, although the distribution's mass lies at exponential radius, outside the regime in which such a separation would imply threshold-circuit lower bounds. A further exact separation shows a polynomial-width depth-4 network computing a benign O(√d)-Lipschitz function that any depth-3 network agreeing with it on the unit hypercube must implement with exponentially many first-layer neurons.

The Loss Floor of Denoising Score Matching: Fisher Geometry from Schr\"odinger Bridges

Avinash Raju, Kai Zhang Denoising score matching trains diffusion models by regressing onto a conditional score even though sampling needs the marginal score; the two share a population minimizer, but the conditional target stays random at a fixed noisy state, leaving an irreducible floor in the training loss. For a general corruption kernel under mild regularity, the excess loss is exactly the trace of the Fisher-Rao metric of the conditional endpoint family integrated along the diffusion trajectory, derived from a Schrödinger bridge variational principle in which the ideal objective appears as excess path-space relative entropy. For corruption diffusions the Fisher term is proportional to the rate at which the noisy state loses mutual information about the clean data, splitting the floor into a data-determined information flow and a weight set by the corruption schedule and objective; the Gaussian case yields a closed form that recovers reparametrization invariance of the continuous-time objective and links its high-SNR divergence to the data's information dimension. A practical consequence is that raw losses under different noise ranges or weightings carry different additive floors and so need not rank models consistently.

Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime

Hugo Latourelle-Vigeant, Sinho Chewi, Aram-Alexandre Pooladian, John Sous, Theodor Misiakiewicz cross-listed Score-based generative models reduce distribution learning to a sequence of regression problems that, solved exactly on finite data, would simply reproduce the training samples, so their ability to generalize must come from implicit or explicit regularization during training. The authors develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized networks by studying denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel, deriving exact risk trajectories under gradient flow in the proportional high-dimensional regime where sample count scales with dimension. The trajectories pass through three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolates the training objective, and an empirical Bayes estimator that memorizes the data. Analyzing how these combine along the reverse-time stochastic differential equation characterizes the distribution of generated samples and reveals both familiar supervised-learning mechanisms, such as kernel linearization and self-induced regularization from the nonlinear part of the kernel, and a phenomenology specific to generative modeling.

The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem

Elioth Sanabria cross-listed Large language model (LLM) providers respond to compute congestion by degrading service, routing queries to smaller models, cutting reasoning effort, or truncating context, and count the result as savings; the authors argue this accounting prices a query when the customer is buying an answer. They model inference allocation with three classical operations-research primitives: a newsvendor whose stockout cost is churned customer lifetime value, a geometric retry multiplier in which failed answers recycle as dissatisfaction, and a two-regime transient queue whose arrival rate is made endogenous by retries. Statically, there is a measurable regime in which a cheaper model saves energy per satisfied answer while consuming strictly more capacity per satisfied answer, so the discount inverts exactly when capacity binds; dynamically, a reactive throttle fired during a surge can cross an ignition threshold beyond which it generates more traffic than it sheds. With heterogeneous customers, throttling becomes a transportation problem whose dual, the shadow price of intelligence, prices a marginal query by class and by hour, and closed-form trajectories make it computable in milliseconds.

Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs

Amirhesam Abedsoltan, Enric Boix-Adsera, Fivos Kalogiannis, Mikhail Belkin Much of feature-learning theory assumes a network discovers a single global low-dimensional predictive geometry, and the authors argue this picture is incomplete. Studying regression problems with clustered data, they show that multilayer perceptrons (MLPs) naturally develop monosemantic specialized neurons, each strongly aligned with one predictive feature relevant to a particular region of input space, so the network learns a collection of local low-dimensional representations that together can span a high-dimensional space. This specialization provably gives MLPs a data-efficiency advantage over feature-learning methods built around a global low-dimensional representation.

Provable Quantum--Classical Separation for Continuous Gibbs Sampling

Enrico Olivucci, Mariia Sobchuk, Sehmimul Hoque, Jeffrey Hnybida, Kyungho W. Kim, Ala Shayeghi et al. cross-listed Whether quantum computers can provably outperform classical ones at sampling from continuous distributions had no established separation until now. For Gibbs states proportional to e^(-βE) on the d-dimensional torus with smooth (s-Gevrey) potentials and barrier amplitude α = e^(βΔ), where Δ is the energy range, the authors prove an information-theoretic lower bound showing every classical algorithm with query access to the log-density and its derivatives of any order needs Ω(α) queries to sample at constant total-variation accuracy, whereas a quantum algorithm based on quantum singular value thresholding and temperature annealing needs only Õ(√α) gradient-oracle queries. The quadratic advantage in barrier amplitude becomes exponential in dimension, e^Ω(d), at low temperature.

Across the Loss Landscape with Progressive Growth

Paul Caillon, Christophe Cerisara, Alexandre Allauzen Grow-and-optimize training schedules are known to find flatter minima, and this work explains why by treating growth as progressive relaxation of constraints. Training starts in a low-dimensional random subspace with the orthogonal complement frozen at initialization, then repeatedly unlocks nested subspaces and re-optimizes; under local regularity assumptions the authors prove sublevel sets are approximately ellipsoidal and that basin reachability under frozen directions is governed by an explicit effective curvature, so wide basins gain relative volume while sharp ones are suppressed. Experiments on toy landscapes and ResNet/CIFAR-100 confirm the flattening effect but also show that reduced curvature does not consistently translate into better test accuracy, complicating the usual flatness-generalization story.

Optimal Alternating Regret for Online Learning and Games

Yixin Tao, Weiqiang Zheng Alternating regret is a regret notion motivated by alternating learning dynamics in games, and the paper settles its minimax-optimal rate for both online linear optimization (OLO) and online convex optimization (OCO). For OLO over the probability simplex, the authors give an algorithm with O(log d) alternating regret that stays constant in the horizon T, with a matching lower bound, improving the prior O(log^(2/3) d · T^(1/3)) result. This yields the first uncoupled learning dynamics with O(1/T) convergence to coarse correlated equilibria in two-player general-sum games, alongside O(log d / T) convergence to Nash equilibria in two-player zero-sum games, where all prior work carried extra log T factors. For general OCO over a d-dimensional compact convex set, they obtain O(d log(1 + T/d)) alternating regret with a matching lower bound, showing the logarithmic dependence on T is unavoidable.
26 more specialized papers

Safety & Alignment 28

Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers

Mingyang Xu cross-listed Text-to-image pipelines rely on learned aesthetic and preference scorers to filter training data and guide generation, but whether these scorers treat demographic attributes as quality signals has been unclear. The audit applies pixel-level interventions on skin tone and body type to synthetic and real images and measures the responses of LAION-Aesthetics, PickScore, ImageReward, and HPSv2, with placebo arms applying the same CIELAB lightness shift to non-skin regions. The dominant effect along skin lightness is a fidelity preference: unaltered images score highest and perturbations in either direction are penalized, and the placebo arms show this penalty is not an artifact of the skin operator. Synthetic-only audits turn out to be misleading, since LAION-Aesthetics shows a strong preference for darker skin on synthetic faces that reverses and shrinks on 1,470 real faces, and results reverse or attenuate for other scorers too, so the authors argue that only within-image causal isolation on real data can separate true demographic bias from fidelity preference.

AI Agents Push Humans Out of the Loop

Margaret Mitchell, Avijit Ghosh, Samir Passi Human oversight is the standard proposed safeguard for increasingly autonomous AI agents, but the authors argue that current agent design impedes effective oversight and that extended use of AI systems degrades the very cognitive capacities oversight requires. Drawing on automation research and human-computer interaction, the position paper outlines design-level affordances and organizational protocols intended to support overseers in exercising critical judgement and to counteract the skill atrophy that arises from prolonged reliance on automation. The central claim is that supporting the situated goals and cognitive requirements of human overseers should be treated as equal in priority to agent capability, since without it agent systems passively incentivize the erosion of the human skills they depend on.

Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model

Shashwat Pandey, Satwik Pandey, Suresh Raghu cross-listed On-device language models now ship to hundreds of millions of devices with no server-side moderation, yet the configuration developers can actually deploy is rarely audited independently. Red-teaming the developer-accessible on-device foundation model on calibration, confident confabulation on false-premise questions, and over-refusal of benign prompts reveals a task-asymmetric miscalibration: the model confabulates on 69% of false premises while refusing 18% of entirely benign inputs, atop self-reported confidence that is saturated and non-discriminative (AUROC 0.47, ECE 70, worst among comparable small models). Confident-correct and confident-wrong outputs are surface-indistinguishable, with a classifier over 15 user-visible features separating them at AUROC only 0.55 and no cheap single-generation signal exceeding 0.68, whereas a black-box consistency wrapper requiring no model access cuts confident confabulation from 75% to 3% and raises selective accuracy from 43% to 83% at a tunable cost. The audit protocol, surface-indistinguishability test, code, and frozen evaluation items are released for reuse.

TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers

Mehrdad Rostamzadeh, Sidhant Narula, Mohammad Ghasemigol, Daniel Takabi cross-listed The Model Context Protocol (MCP) connects LLM agents to external tool backends, opening a server-side threat the authors call TrustShift: a compromised server behaves benignly during a conditioning phase to build operational reliance, then switches to an adversarial payload once an interaction threshold is reached, evading pre-deployment static analysis and, for schema-valid variants, runtime middleware filters. TrustShiftProbe contributes a stateful temporal threat model, a language-agnostic attack engine instantiating compromised servers across four production domains, a taxonomy of nine variants over three execution mechanisms and three adversarial objectives, and SHIELD, a multi-tier zero-oracle runtime defense at the transport boundary that audits server payloads against behavioral baselines learned during clean trust windows. Across frontier proprietary and open-weight models, TrustShift attacks achieve a 69.5% mean success rate, which SHIELD reduces to 42.7%.

SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

Lijia Huang, Yao Fu, Sihao Ren Large language models (LLMs) tend toward social sycophancy, and existing evaluations measure it under a single fixed prompt, leaving open whether the behavior is stable when the same situation is phrased with different social cues. SyPS (Sycophancy Prompt Sensitivity) builds controlled prompt variants that keep the underlying user situation fixed while varying user confidence, emotional framing, social consensus, and validation-seeking language, and introduces the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level measure that separates baseline sycophancy from prompt-induced shifts across paired variants. The authors find that sensitivity is socially structured: validation-seeking and emotional-pressure cues often increase sycophancy, while counter-framing and anti-sycophancy prompts tend to reduce it, giving a model-level way to compare robustness to such cues.

Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications

Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney, Heiko Ludwig, Kate Soule, David Cox Organizations deploying generative AI need safety policies tailored to their application context, regulatory environment, and user base, but existing policy-specification formats were built for traditional access control and cannot express content-based constraints on what a model may say. IBM's Granite.Trust project introduces the Actionable Policy schema, a YAML format for specifying what model responses can and cannot contain that supports exception-based governance for tracking violations, plus a synthetic data generation pipeline that turns a policy into policy-aligned training and test data and a set of tools for authoring and enforcing the schema. A policy written once can be enforced across the whole lifecycle, from model alignment through runtime monitoring, and the schema, example policies, and tooling are released as open source.

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

Joshua Penman Language models see only tokens, so they must infer for themselves whether a span is user input, tool output, or instructions, and prompt injection exploits this by making untrusted text read like a command. Semantic Overlays are small learned adapters applied to a frozen model's residual stream at chosen prefill positions, creating an out-of-band annotation channel that tokens cannot forge; unlike fixed steering vectors they are trained, composable, and selectively applied, and can carry payloads such as marking a span as non-executable. Overlays can even reshape how the model perceives marked content, for example making it faithfully rewrite a code snippet in whichever language the overlay asserts. On prompt injection benchmarks, SEP separation rises from 24.3% to 96.5% with utility unchanged, TensorTrust attack success falls from 34.8% to 6.6%, and all four PIArena attack families drop to 0% compliance, while marked spans remain readable with a 92.5% exact copy rate.

AI Finds A Way

Aaron Dharna, Cong Lu, Ryan Sullivan, Joel Lehman, Victoria Krakovna, Jeff Clune Machine learning systems routinely surprise their own developers by exploiting reward loopholes, circumventing design constraints, or stumbling onto unknown phenomena, yet these episodes are rarely documented formally. Drawing on the work of over 100 researchers, the authors collect 26 firsthand anecdotes spanning reinforcement learning successes in hard domains, reward hacking of underspecified objectives and unarticulated constraints, and case studies suggesting that internet-scale foundation models can amplify rather than resolve these dynamics. The collection is framed as both a safety resource, illustrating the difficulty of aligning models with human values without dampening their creativity, and an argument that the same learning dynamics can be harnessed to accelerate scientific discovery.

Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems

Paul Vautravers, Oliver Chalkley, Gabriel Downer, Kate S, Damian Ruck AI evaluation is largely model-centric, offering little insight into how a model's behaviour translates into risk once it is embedded in complex sociotechnical systems such as Critical National Infrastructure (CNI). The proposed framework chains structured hazard analysis, component-level testing, and probabilistic system modelling so that a model-level failure can be traced to a system-level outcome, and is worked through on the UK's Real Time Gross Settlement (RTGS) payment system using Systems Theoretic Process Analysis (STPA) to derive AI-driven loss scenarios, one of which is adversarial manipulation of LLM-based trading. Component experiments show simple adversarial inputs measurably shift recommendations where they are followed, and when mapped into a financial contagion model these shifts increase bank failures and lower the shock threshold for cascading disruption, especially under widespread or monopolistic AI adoption.

More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight

Yuchen Han, Cheng Yan, Wuyang Zhang In pre-execution oversight, a fallible LLM monitor vets an agent's planned actions before irreversible execution, and every protocol must fix a unit of verification, meaning how many actions one monitor call reviews; existing designs take this unit as given, and natural traces cannot isolate its effect because review length co-varies with error type and position. The twin-prefix framework generates, from each gold plan, a prefix containing one injected environment-accepted error and a clean twin differing in a single write, judges each pair at five nested lengths, and scores discrimination by pre-registered informedness (catch rate minus false-rejection rate), since catch alone rewards rejecting everything. Longer review windows raise catch but false rejection climbs in lockstep, and informedness peaks at one or two actions for all six judges in both domains, so longer windows make zero-shot monitors more rejective rather than more discriminative; replaying withheld observations traces the failure largely to observation deprivation. A calibrated short unit recovers up to 0.95 informedness over eight-action review, no tested label-blind policy consistently beats it, and the authors recommend that safety cases state the unit and co-report the clean-control series.

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang cross-listed Safety alignment in large language models (LLMs) is brittle against both jailbreak prompts and neuron-level attacks that prune safety-critical neurons after deployment, and both exploit the same weakness: safety-relevant information concentrates in a sparse subset of neurons. NeuronGuard is a fine-tuning-stage defense that redistributes safety signals across a broader set of neurons by identifying safety-critical neurons with periodically refreshed per-layer linear classifiers, forcing refusal behavior under deliberate neuron ablation, and applying KL-divergence regularization for distributional consistency, with a randomized gradient projection strategy resolving conflicts between the defense and task objectives to preserve downstream utility. The authors give a formal guarantee that the method strictly reduces the upper bound on attack success rate (ASR), and experiments across three LLMs, six state-of-the-art attack strategies, and multimodal settings show near-zero ASR while maintaining task accuracy, including against white-box adaptive adversaries.

RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation

Yueyang Quan, Anjun Gao, Yufei Xia, Minghong Fang, Zhuqing Liu cross-listed Retrieval-augmented generation (RAG) grounds model responses in external documents, but adversarial documents injected into the knowledge base can enter the context window and steer the model toward targeted wrong answers, and existing post-retrieval defenses based on instruction following, parametric knowledge, or text-level consistency can be imitated or optimized against by adaptive attackers. RAGSentinel is a training-free, label-free defense for black-box RAG systems that uses a surrogate encoder to measure query-conditioned hidden-state shifts induced by each retrieved document, removes shared topic directions, and filters poisoned documents as geometric outliers relative to a robust majority consensus. The authors prove that under an honest-majority assumption and a representation-level separation condition the method exactly recovers a poison-free, majority-sized context, and experiments across three question-answering datasets, three LLM families, and multiple poisoning attacks show consistently low attack success rates with competitive accuracy, including against adaptive attacks with full knowledge of the pipeline.

WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents

Lin-Fa Lee, YI-YU Chang, Kuo-Hui Yeh cross-listed The W3C WebMCP proposal lets LLM agents call tools exposed by web pages, but a browser security model built around the Same-Origin Policy (SOP) gives agent-accessible tools weak provenance and lifecycle guarantees, opening the door to subject-attribution spoofing, uncontrolled tool lifecycles, and semantic prompt injection. WebMCP-Phalanx is a dual-layer runtime: a browser-native trust anchor binds each tool to its registering principal with cryptographically protected capability credentials and propagates provenance labels through the tool's lifecycle, while a Quarantine Agent (Q-LLM) with no tool authority inspects tool metadata, outputs, and page content for injection before validated content reaches a Privileged Agent (P-LLM) that executes. The ownership mechanism cuts revocation and overwrite attack success from 100% to 0%, the dual-agent runtime blocks all 80 prompt-injection attempts embedded in tool descriptions and limits tool-return attacks to 2 successes out of 80, and task utility stays statistically indistinguishable from the no-attack baseline. A white-box adaptive attacker can still bypass description-based filtering via malicious tool names invoked before inspection, motivating a call-timing gate that delays invocation until all agent-visible metadata is validated.

What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions

Yichao Gao, Yumo Zhang, Yunhao Yao, Haohua Du, Puhan Luo, Ruiqi Li et al. cross-listed Large language model (LLM) agents wired to external resources take everything through one natural-language context channel, so untrusted data can be parsed at inference time as behavior-guiding instructions that hijack tool calls, and static input/output filters miss inducements that only surface during model reasoning. Attnlocate is a runtime framework that treats finding these spans as an object-detection problem: a multi-head, multi-layer attention aggregation scheme builds a token-level feature space, a 1-D U-Net with an anchor-free detection head localizes the spans that genuinely drive tool-calling decisions, and the system then adjudicates invocation attempts based on the authority of whichever provider the detected spans came from. Across ten agent configurations from five LLM families, covering indirect prompt injection and tool poisoning, it reaches a mean IoU of 0.743, an average AUROC of 0.956, and a 0.934 true-positive rate at a 0.067 false-positive rate, transfers to unseen models, and supports new authority policies without retraining.

Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia When one AI system's decisions affect many people, aligning it means aggregating divergent preferences, and reinforcement learning from human feedback largely sidesteps this question with poor social choice guarantees. The work reformulates alignment as linear optimization over a convex impact space of welfare consequences, which lets standard tools from welfare economics and mechanism design translate alignment protocols into welfare outcomes and translate a planner's desired welfare constraints back into protocols. Using this lens, the authors show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous, derive a family of protocols that maximize utilitarian social welfare subject to bounds on individual or group harm, and illustrate the welfare implications on real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.

TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models

Zhenyu Wu, Siyuan Chen, Changchun Yang, Jiaqi Dong, Min Zhou, Ali Almadan et al. Large Reasoning Models (LRMs) can emit unsafe content in their intermediate reasoning traces even when the final answer looks safe, but existing benchmarks for guardrail models focus on prompts and final responses and offer only binary labels without supporting evidence. TRACE is an evidence-grounded benchmark spanning the full LRM pipeline, with prompts in two languages across nine risk categories and ten attack strategies, reasoning traces and responses generated by four LRMs, and safety labels plus extracted evidence spans for each component. Evaluating 18 guardrail models shows that judging the safety of reasoning traces is substantially harder than judging prompts or final responses, and current models also struggle to accurately extract the supporting evidence.

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He Safeguarding language model agents requires judging whole execution trajectories against safety policies that depend on context, and existing policy-aware safeguards built on prompting or supervised fine-tuning generalize poorly to unseen trajectories and shifting policy libraries. RePolicy learns safety-policy invocation through reinforcement learning: given a trajectory and a dynamic policy library, it selects the applicable policy and uses its content to produce a policy-grounded rationale and safety judgment, trained first on the new PolicyTraj-20K dataset and then with GRPO using verifiable rewards and policy-context perturbation. Across six agent safety benchmarks it achieves strong overall safety-detection performance and robust policy invocation under varying policy contexts.

SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits

Daming Luo, Christy Liang, Junyu Xuan Activation steering can shift model behaviour without proof that the effect is specific to the intended concept rather than a side effect of any perturbation. SteerCheck is a preregistered attribution audit that matches off-target KL divergence and separates mean, protected-tail, polarity, transfer, and semantic claims, applied via exact replay of 960 interventions on Qwen3-14B. It finds that isotropic random directions are near-orthogonal and uninformative comparators, whereas sign-randomized same-construction directions often stay aligned with the target (steering effect correlates with signed cosine at ρ=.94), a form of alignment leakage that limits what randomization tests can distinguish; the primary Qwen gate fails on the protected tail, language controls pass in Qwen and DeepSeek, and the automatic judge fails calibration (macro-F1 .562), so semantic conclusions remain descriptive.

Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs

Jiali Wei, Ming Fan, Mingkun Zhang, Haoyu Wang, Jun Sun, Guoheng Sun et al. cross-listed Multimodal large language models (MLLMs) can inherit backdoors from their construction pipelines, with triggers hidden in images, text, or both, and existing removal methods built for classifiers or inference-time filters do not eliminate them from the model. RACER builds on the observation that backdoors cause abnormal layer-to-layer changes in internal representations concentrated in the token region carrying the trigger; it decomposes representations into visual and textual regions, normalizes their layer-wise inconsistency separately, and uses min-max adversarial fine-tuning against worst-case perturbations to suppress the deep representational shifts backdoors rely on. Requiring only 100 clean samples and no knowledge of the trigger or attack, it cuts average attack success rate to 1.1% across 36 backdoor settings on three open-source MLLMs, reaching 0% in 32 of them while preserving clean-task utility.

Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding

Yujing Chang, Thinh Pham, Van-Phat Thai, Chunyao Ma, Yash Guleria, Pham Nhut Huy et al. In safety-critical language understanding, a misread altitude or a dropped execution condition can score respectably on F1 while carrying consequences nothing like an ordinary error, so semantic metrics may overstate operational reliability. The authors build a consequence-aware evaluation framework and instantiate it in a diagnostic air traffic control benchmark grounded in aviation standards and shaped by feedback from 40 controllers across three countries, then score 8 models under both regimes. Conventional metrics rate every model substantially higher than consequence-aware scoring does, including models that look reliable by standard measures, and risk-aware fine-tuning narrows the gap without closing it.

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

Zhijie Zheng, Yu Li, Chen Qian, Yuqian Fu, Yanwei Fu, Lu Sheng et al. LLM agents that invoke tools can modify files, leak information, or take unauthorized actions, yet most guardrails judge completed trajectories rather than checking individual steps before execution. StepGuard is a step-level guard model that audits finished trajectories and screens tool actions pre-execution; it is trained on data from StepGen, an engine that generates paired safe and unsafe trajectories sharing the same context but differing at the risky step, using Balance-GRPO to dynamically rebalance learning between safe and unsafe actions and curb both over-defense and under-defense. It achieves the highest average accuracy among open-weight guard models, comparable to GPT-5.4, and when guarding agents on AgentDojo and AgentDyn it cuts mean attack success rate by 77.3% while mean utility drops only 2.8 percentage points.
7 more specialized papers

Multimodal 27

Taming Visual Neglect: A Variational Information Bottleneck Framework for Adaptive Attention in Multimodal In-Context Learning

Kaito Tanaka, Yuji Nishimura, Keisuke Matsuda, Aya Nakayama Large vision-language models sometimes exploit visual demonstrations during in-context learning (ICL) and sometimes ignore them entirely, and it has been unclear when visual context actually helps. VIB-ICL frames this through the Information Bottleneck principle, defining Cross-Modal Information Gain (CMIG) as the extra mutual information visual context provides about the target beyond text, and derives a generalization bound showing that multimodal ICL's excess risk over text-only ICL is governed by CMIG. The analysis implies that visual neglect is the bottleneck-optimal behavior when visual information is redundant, yielding a closed-form Attention Reallocation Principle that the algorithm instantiates by estimating CMIG with variational bounds and adjusting visual attention weights dynamically. Across five benchmarks the method yields accuracy gains of up to 4.7% and a 35% reduction in required demonstrations.

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

Zhengxiang Wang, Owen Rambow Standard visual grounding benchmarks hand a model one informative referring expression and ask it to locate the target, which ignores that real-world references are often incomplete and get resolved through back-and-forth. The authors build a controlled evaluation framework for large vision-language models (LVLMs) that varies how much target information is given upfront versus how much must be acquired by asking questions, across four human-grounded visual contexts and four interaction protocols. Current LVLMs fall well below task-level human baselines, benefit from follow-up questions mainly when refining or repairing an initial description, and do worst when no description is given and they must proactively ask; they are also poorly calibrated, reporting confidence above their empirical accuracy. Follow-up studies find the same patterns across human versus AI descriptions, reasoning effort levels, repeated interactions, and visual contexts.

IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views

Yuchuan Wu, Ke Niu, Haiyang Yu, Zhuofan Chen, Xiangyang Xue, Bin Li cross-listed Turning dimension-annotated orthographic drawings into executable parametric CAD code demands geometric understanding, procedural reasoning, and precise numbers, and one-shot vision-language generation cannot inspect intermediate results or fix early mistakes, often yielding code that does not run or geometry that does not match. IterCAD reframes the task as progressive program repair: the model repeatedly analyzes the current CAD result, reasons about its discrepancy from the target views, and explicitly decides whether to REVISE the code or STOP. To make this learnable, the authors build IterCAD-RS, a revise-or-stop supervision set containing both repairable intermediate states and already-correct states, and train in three stages covering initial generation, revision learning, and multi-turn reinforcement learning. On CADExpert, IterCAD consistently improves both code executability and geometric fidelity over strong one-shot baselines.

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu cross-listed Universal multimodal embeddings map text, images, video, and visual documents into one shared space for retrieval, recommendation, classification, and agentic systems. WeMM-Embedding is a family of 2B, 4B, and 9B models that accept arbitrarily interleaved multimodal inputs with flexible output dimensions, trained through a large-scale multimodal alignment stage followed by refinement on curated data with fine-grained relevance supervision and cross-scale knowledge transfer. The 9B model sets a new state-of-the-art overall score of 80.6 on MMEB-v2, the 2B model already beats the previously leading 8B open-source baseline, and the family shows gains on a 26-task in-house benchmark and 14 online A/B tests across WeChat services, with weights and code released.

VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference

Lyuke Wang, Zhuo Li, Guangxu Zhu cross-listed Vision large language models (VLLMs) pay heavily for long-context inference because visual key-value (KV) caches dominate memory and compute, and existing compression methods prune uniformly across tokens and layers, discarding useful information. VisCache is a training-free, plug-and-play two-stage pipeline: a lightweight vision-language model first filters temporal redundancy by forwarding only semantically informative keyframes, then PruneKV compresses the cache using a parabolic layer-wise budget allocation and an asymmetric update that prunes keys while fusing values to preserve context. It achieves up to 2.35x speedup while retaining only 19-28% of the KV cache with competitive accuracy, outperforming existing baselines on the efficiency-performance frontier.

OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses

Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu et al. Multimodal models used as unified judges of text-to-image (T2I), text-to-video (T2V), and text-to-speech (TTS) outputs are relied on for evaluation and automatic annotation, yet existing benchmarks over-represent positive examples and conflate distinct failure modes, so a judge can score well without actually recognizing failures. D3-Omni is a diagnostic benchmark of 10,671 samples across 53 orthogonal binary dimensions (17 for T2I, 22 for T2V, 14 for TTS), built by fixing verified fully-positive seeds and deriving negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations, giving near 1:1 per-dimension label parity and a uniform distribution over total-score levels. Under this balanced view, even strong judges confirm satisfied requirements far more reliably than they detect violated ones, struggle on modality-specific dimensions, and tend to collapse nominally distinct attributes into a single decision.

FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation

Junjie Li, Xuelong Geng, Kun Xie, Feiyu Shen, Yichen Wu, Ziqi Dai et al. A single audio model that understands speech, paralinguistics, and environmental sound while also synthesizing and editing speech faces a representation conflict: understanding favors compact features for long-context modeling, whereas generation needs reconstructible features that preserve fine acoustic detail. FireRedAudio is a 9B-parameter audio language model that gives understanding and generation separate continuous input pathways inside one trainable autoregressive LLM, with a dedicated audio encoder for recognition and analysis, a RedAE-based pathway for speech inputs to generation, and a flow-matching diffusion transformer (DiT) that the LLM conditions to produce continuous acoustic latents. Progressive multitask training yields a model handling automatic speech recognition (ASR), audio understanding on recordings up to one hour with second-level timestamp accuracy, zero-shot and instruction-following text-to-speech, and semantic and acoustic speech editing, with competitive or leading results in audio understanding and multilingual ASR and substantial improvements over Ming-UniAudio-Edit in speech editing; code is released.

VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models

Guoyang Xu, Hao Chen Long-video understanding hinges on how a limited model context is built from a much longer video, but compression, retrieval, memory, and agentic evidence-gathering are usually baked into hand-designed inference systems or co-optimized with other components, obscuring how much the context-construction program alone contributes. VideoHarness-RSI treats this as automated harness design around a frozen vision-language model (VLM): an outer-loop proposer uses prior programs, evaluation results, and execution traces to generate candidate executable context constructors, which are run and evaluated end to end, with successful variants retained for further search while the answering model and interface stay fixed. Starting from uniform sampling, recursive harness search consistently finds improvements and beats several weaker hand-crafted baselines, and starting from a stronger hand-crafted baseline it still yields further gains; the selected harness also transfers to additional long-video benchmarks without further search.

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thadd\"aus Wiedemer, Christoph Schuhmann et al. cross-listed Open video data for multimodal pretraining has lagged far behind the scale of web image-text corpora. LAION-BVD collects 1.3B platform-specific video URLs from CommonCrawl and downloads 80M videos totaling 10 million hours, then applies content-aware scene detection to extract clips and synthetically generates video and audio captions for them, targeting joint pretraining across video, audio, and image modalities. Models trained on the data are competitive on standard video-text and audio-text benchmarks with consistent gains as training or model scale grows, and scene-changing frames extracted from the videos form an image-text source whose visual distribution differs from web image corpora and yields strong image-text retrieval performance. The dataset is released to the research community.
18 more specialized papers

Vision 20

DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection

Huaiyuan Qin, Gabriel James Goenawan, Zihang Lin, Muli Yang, Hongyuan Zhu cross-listed Linear attention would make high-resolution object detection cheaper, but swapping a trained detector's Softmax-attention Vision Transformer (ViT) backbone for a linear one causes severe degradation, and generic label-free distillation that works for classification tends to fail on detection. Detector-Interface Distillation (DiD) trains only the linear-attention backbone to reproduce the exact feature tensors the frozen downstream detector expects, aligning against a frozen Softmax teacher instead of imitating its internal hidden states. On DOTA-v1.5 it substantially outperforms established baselines and matches supervised, fully trained linear models, with adaptation completing in about 87 minutes on 4 GPUs while the linearized backbone cuts inference latency by roughly 62% and peak memory by about 49%.

The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models

Liangzhi Li, Bowen Wang, Yiming Qian, Thorsten Neumann, Xia Xie, Guangshun Li cross-listed Many few-shot adaptation methods for vision-language models classify with a convex blend of the zero-shot text prototype and the mean of the K labelled image features, tuning a single blending ratio on held-out labels, often the test set itself. The authors derive the closed-form ratio minimising prototype mean-squared error, show its support-set plug-in is a positive-part James-Stein coefficient, and across 4,800 cells (ten datasets, five backbones including SigLIP, five shot counts, five seeds, four prompt tiers) find it trails a test-set-oracle ratio by 8.5 points because 78% of the text-image prototype distance it treats as bias is a class-independent offset that the arg max largely cancels. Leave-one-out on the support set alone lands within 0.9 points of the oracle, yet validation-free linear probes such as CLAP (+1.9 points) and LP++ (+1.5) beat even the oracle-tuned blend, locating the ceiling in the model class rather than the hyperparameter.

Luce: Relightable Gaussians for 3D Asset Generation

Mayank Singh, Michele Stoppa, Alvise Memo, Rui Yu, Harsha Kalli, Srimanth Gunturi et al. cross-listed Production-ready image-to-3D generation needs a representation that captures geometry together with physically based rendering (PBR) materials such as albedo, metallic-roughness, and surface normals so assets can be relit and integrated into standard rendering pipelines. Luce unifies geometry and PBR materials in a voxelized multimodal Gaussian cloud with dedicated Gaussian primitives per modality, compresses it with a variational autoencoder into a material-aware latent space, and generates that latent from a single image with a rectified-flow transformer conditioned on multi-layer features from a pretrained image encoder; the latent decodes into relightable PBR Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K it reaches state-of-the-art single-image-to-3D quality, improving FID by 28% over the strongest baseline, and on a new benchmark of AI-generated images it raises the CLIP image-alignment score to 0.8519 versus 0.8299 while preserving fine details such as text, logos, and inscriptions.

It depends: Incorporating correlations for joint aleatoric and epistemic uncertainties of high-dimensional output spaces

Leonhard F. Feiner, Manuel Nickel, Martin Menten, Laurin Lux, Rickmer Braren, Daniel Rueckert et al. Uncertainty Quantification (UQ) for high-dimensional regression such as image segmentation or restoration needs to capture both aleatoric uncertainty (inherent data noise) and epistemic uncertainty (the model's confidence under unfamiliar conditions), but most methods model only one or ignore correlations across output dimensions. The proposed approach approximates the joint uncertainty with a low-rank plus diagonal covariance structure that captures essential output correlations without the cost of full covariance matrices, and combines the two uncertainty types into a single second-order distribution that supports sampling and log-likelihood evaluation downstream. With added stabilization strategies for training and inference, the method achieves superior uncertainty quantification on image inpainting, colorization, optical flow, and depth estimation.

Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core

Yogesh Kumar Mamba-style state space models (SSMs) have been applied to video anomaly detection, but existing methods still buffer clips or windows internally, offer no theory relating temporal memory to detection latency, and report efficiency only as GPU throughput rather than on the edge hardware they target. The proposed detector is strictly causal, with a fixed-size state updated in O(1) time and memory per frame, built on a diagonal linear state space recurrence with an input- and state-dependent decay gate and trained self-supervised via causal next-embedding prediction on a frozen visual backbone. A closed-form relation between the decay spectrum and detection delay predicts a settling delay of 57-59 frames, far above the measured 1.6 and 18.4 frames on UCSD Ped2 and CUHK Avenue, indicating that the event-boundary gate rather than the base decay governs responsiveness. Measured end to end on an Apple M3 Pro, the model runs at 0.74-0.77 ms per frame (over 1300 FPS), though with an untuned configuration its frame-level AUC of 67.9% and 70.2% trails prior non-causal SSM baselines, and ablations show the gate hurts on the smaller Ped2 training set but helps on the larger Avenue one.

What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

Hao Chen cross-listed Generative image models are usually ranked with Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), but FID only compares the first two moments of feature distributions and neither metric offers a calibrated statistical test or indicates whether a model's samples are under- or over-dispersed relative to real data. The authors demonstrate the moment-matching weakness concretely: visually unrecognizable images optimized only to match the reference Inception mean and covariance reach an FID of 24.7 on ImageNet, versus 58.6 for held-out real images. They propose ZID (Z-resolved Integrated Diagnostic), which combines six standardized location- and dispersion-sensitive statistics from a rank graph (RISE) and Gaussian kernels (GPK at two bandwidths) and reports three linked outputs: a ranking index, a permutation p-value for testing distributional equality, and a signed dispersion readout. In controlled sweeps ZID tracks increasing severity of departures where FID stays flat or reverses, and on DiT-XL/2 and SiT-XL/2 classifier-free guidance sweeps its signed readout labels the high-guidance diversity collapse as under-dispersion.
14 more specialized papers

Reinforcement Learning 18

Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

Jaemoo Choi, Wei Guo, Yuchen Zhu, Arash Vahdat, Molei Tao, Julius Berner et al. cross-listed Reward fine-tuning of diffusion models typically borrows policy-gradient machinery from language models, but diffusion models lack tractable sample likelihoods, forcing existing methods to construct trajectory likelihoods or evidence-lower-bound approximations at extra cost and complexity. Reward-based velocity matching (RVM) is a trajectory-free update that acts directly on the velocity field, reinforcing directions tied to high-reward samples, suppressing low-reward ones, and optionally anchoring to a reference velocity; it recovers RAM and DiffusionNFT as special cases. Across large-scale diffusion fine-tuning tasks, RVM matches or beats trajectory-based policy-gradient methods at substantially lower training cost, and once the update is simplified the particular loss variant matters less than reward and anchor design. For video generation, a new dynamic-tracking reward counteracts the tendency of standard preference rewards to favor clean but nearly static clips, improving motion along with overall VBench scores.

Mitigating Exploration Bias in RL for Multi-Instruction Following

Mian Zhang, Yueqin Yin, Kaiyu He, Peilin Wu, Xinlu Zhang, Mingyuan Zhou et al. Reinforcement learning (RL) improves instruction following in large language models, but when prompts contain multiple instructions, training is biased toward easy ones: the policy rarely satisfies hard instructions initially, so exploration never rewards them, and cumulative rewards that count fulfilled instructions treat easy and hard ones identically. The authors propose two metrics to quantify this exploration bias and a two-stage remedy: Behavioral Bootstrapping, a lightweight rejection-sampling fine-tuning stage before RL that activates hard instructions, and Scarcity-Aware Rewards, which weight each instruction by its empirical scarcity. The metrics correlate strongly with model performance, and the best models outperform baselines by a significant margin across three verifiable instruction-following benchmarks; code is released as MulIF.

CoDrift: Compositional Drifting for Offline Reinforcement Learning

Xiewei Ni, Ruofeng Mei, Xiangyu Xu Offline reinforcement learning must balance staying within the behavioral support of a fixed dataset against preferring high-value actions, and CoDrift recasts each objective as an action-space motion field describing how generated actions should move, so heterogeneous objectives can be combined by direct field composition. Inspired by drifting models, it composes three fields into one policy field: a conditional field that preserves state-dependent behavioral structure, a marginal field that pools actions across states to give a more stable generative signal in the single-positive-sample regime of continuous control, and a value field that pushes actions toward higher-value regions. The composed field is absorbed into a stochastic generator that produces an action in a single forward pass at deployment, and across 73 tasks from OGBench and D4RL it achieves the best average rank in both offline and offline-to-online settings against state-of-the-art methods.

XP-JEPA: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi Self-predictive latent world models jointly train an encoder and predictor, which can co-adapt to transitions that are easy to predict but weakly tied to the actual physical evolution of the scene. The cross-predictive joint-embedding predictive architecture (XP-JEPA) encodes visual observations and privileged physical states separately, rolls both forward with a shared action-conditioned predictor, and matches each prediction against both future representations so the latent dynamics are grounded in physical transitions; the physical branch is discarded after training, leaving a visual-only model at deployment. On a multi-task suite with six evaluation subfamilies, the method cuts rollout drift of a newly fitted predictor from 0.361 to 0.104 and raises mean control success from 53.6% to 78.2%, whereas directly regressing physical states improves position decodability but leaves forecastability and control near the visual-only baseline.

AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL

Xiaolong Jin, Dingmin Wang, Vijay Lingam, Varun Kumar Reinforcement learning for multi-turn LLM agents usually relies on trajectory-level rewards that assign the same advantage to every step, and existing self-distillation methods apply privileged information uniformly, ignoring that routine steps need little guidance while critical error steps need corrective direction that environment feedback alone cannot supply. AHEAD is a step-aware framework in which the teacher receives environment feedback on all steps as a dense grounded signal and additionally gets LLM-generated corrective hints only on error steps, implemented with minimal changes to GRPO. Across ALFWorld, WebShop, and search-based question answering at three model scales, it raises task success by 13.3 points on ALFWorld and 11.0 on WebShop at 7B over GRPO, reaches a given success rate in fewer training steps, and solves tasks within tighter interaction budgets than outcome-only RL and prior self-distillation baselines.

Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping

Yiwen Zhang, Xiaodong Yan, Zhenyu Huang, Deng Zhao, Liang Jiang, Qing Cui et al. Reinforcement learning from verifiable rewards (RLVR) for code generation is limited by test-case coverage: insufficient tests produce false positives that lead to reward hacking and policy degradation. RobustTests synthesizes test cases driven by 'near correct' faulty code so the generator targets latent logical discrepancies, filters invalid and redundant tests with validator agents and behavioral feature clustering, and uses a stepwise dense reward based on pass rates to blunt false negatives from hallucinated synthetic tests. The pipeline yields an augmented test set for CodeContests covering a broader range of faulty-code scenarios. Training Qwen3-32B with RL on a moderately challenging CodeContests subset gains an absolute 3% on LiveCodeBench over baseline methods.

Contrastive Branch Policy Optimization

Ying Wang, Changlin Qiu, Bang Lin, Linbo Jin, Wen Jiang, Zhe Sun et al. Reinforcement learning with verifiable rewards (RLVR) can teach language models multi-turn tool use, but sparse outcome rewards say nothing about which intermediate decisions caused success, and existing branch-sampling methods conflate allocating a rollout budget with converting branch outcomes into token-level credit. Contrastive Branch Policy Optimization (CBPO) separates the two: generation entropy screens candidate branch positions across the whole response, path-level and node-level decay spread a fixed budget so exploration does not collapse onto a few paths or adjacent tokens, and the reward variation within an exact-prefix group of a parent trajectory and its branches defines a Contrastive Branch Value that rescales continuation advantages without flipping their sign, with non-overlapping credit segments preventing duplicated gradients when several nodes are selected on one trajectory. Requiring only outcome rewards and no process annotation, CBPO attains the highest macro-average accuracy on ten benchmarks spanning mathematical reasoning and knowledge-intensive search across two model scales, beating state-of-the-art policy-optimization and branch-based methods.

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng et al. Small language models struggle with search-augmented reasoning, and on-policy distillation (OPD) from teachers is hampered by the cost of collecting multi-turn search trajectories and of fine-tuning task-specific teachers. OPDSearch+ uses a frozen off-the-shelf instruct model as the teacher: the student interacts with a live search engine and is distilled with a per-position forward KL objective, after which reinforcement learning refines the distilled student. The key claim is that the teacher reshapes the student's policy distribution so that subsequent RL converges to solutions RL alone cannot reach; a 3B student outperforms all prior 3B RL baselines across seven QA benchmarks, with gains of 13.1% on HotpotQA and 8.5% on 2WikiMultihopQA.

FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision

Qiming Xie, Wenjie Zheng, Xiangqing Shen, Rui Xia Outcome-only rewards in reinforcement learning with verifiable rewards can encourage hallucination, and existing process-level factual supervision aggregates signals coarsely and never assesses their reliability, producing what the authors call noisy factual credit assignment. Fact-Aligned Reliability-Aware Credit Assignment (FARCA) aligns the granularity of fact verification with token-level policy updates and introduces counterfactual evidence attribution, which measures how much a factual judgment depends on key evidence to derive reliability weights that modulate factual rewards and local advantages. Across multiple models and factual reasoning benchmarks, FARCA significantly improves factuality while preserving general reasoning ability.

Ockhamareto: Pareto-Gated Segment-Level Credit Assignment for Concise Unit-Test Generation with Reinforcement Learning

Dong Huang, Mark Harman, Jie M. Zhang, Zhijiang Guo, Mingzhe Du, See Kiong Ng cross-listed Unit-test generation with reinforcement learning must balance bug-catching power against suite size. Ockhamareto is a single-shot Group Relative Policy Optimization (GRPO) framework with two components: a Pareto-gated bonus that rewards only rollouts that are non-dominated in (mutation score, negative test count) space, and token-level segment credit that attributes each test's marginal mutation kills back to the tokens of its own test block. On UnLeakedTestBench it strictly Pareto-dominates the strongest RL baseline MIST-RL, catching more bugs (49.9% vs 31.3% mutation score at N=5) with fewer tests (2.60 vs 4.67 on average), and it leads both mutation and coverage with the smallest suite on HumanEval+, MBPP+, CodeContests, and TestGenEval-Lite, adding 30-35 percentage points of mutation score at 4B, 9B, and 27B model sizes. The knee point of the efficiency-effectiveness Pareto front is not predicted by simple proxies such as function size, which motivates computing the front for each function under test.

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Zihao Wu, Hongyao Tang, Yi Ma, Huizhong Song, Pengyi Li, Yifu Yuan et al. Massively parallel simulation changes the data regime for off-policy reinforcement learning (RL), and stabilizers designed for data-limited replay do not always carry over. Controlled experiments across eight benchmark families show that parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, that clipped double-Q can be relaxed in high-throughput manipulation, and that age-biased replay weighting improves efficiency across regimes. Building on this, WarpSAC uses Sample Weight Decay and ships two variants: WarpSAC-L (normalization on, clipped double-Q) for CPU-scale training and WarpSAC-A (normalization off, single-Q) for GPU-parallel training. Against FlashSAC it improves normalized score-step area under the curve by 4.5% on nine CPU-scale environments and 23.1% on fourteen GPU-parallel ones, and raises the UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, with 36.4% faster sim-to-real deployment on a Unitree G1.

Where Entropy Is Measured Matters: Policy Geometry in Bounded Continuous-Control PPO

Yiyang He, Zhichun Zhou, Ziwei Wang, Tao Xue, Haolin Fei Continuous-control policies are often trained as unbounded Gaussians and then clipped or squashed into a bounded action range, and the space in which the entropy bonus is computed turns out to change the geometry that proximal policy optimization (PPO) learns. In an 80-muscle MyoLeg task, a clipped Gaussian executes 89.07% of actions within 5% of a bound, and a same-state decomposition shows this is driven mostly by state-conditioned means lying outside the executable interval (82.12%) rather than by variance; a tanh map does not remove the high-variance regime. Latent-space entropy H(u) has zero gradient on the mean and a constant variance-increasing gradient, whereas executed-action entropy H(a) adds an inward gradient through the transform Jacobian, and near-boundary occupancy drops from 71.42% under latent entropy to 18.83% under executed-action entropy (29.76% with no entropy), an ordering reproduced on a 38-dimensional Dog-Stand task with an independent CleanRL-based implementation. Direct mean penalties can match the centering effect, but matched mean geometry can coexist with very different variance and return, so the entropy measurement space is a coupled mean-variance design choice that task return alone does not reveal.

IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents

Bo Ren, Yirong Mao, Yi Yang, Wenhui Que When a language model agent handles a long customer-service interaction, the task reward arrives only at the end, so nothing indicates which of dozens of turns actually mattered — a problem sharpened by users who revise their goals mid-conversation and tools that supply information needed later. Influence-Aware Policy Optimization reads the finished rollout as a typed influence-dependency graph over the agent's own actions, with user replies and tool outputs as evidence, and converts support-use and failed-use edges into routing weights that redistribute the single trajectory-level advantage across steps. With Qwen3-4B and Qwen3-8B it beats multi-turn reinforcement learning baselines on all three service-agent benchmarks tested — τ²-Bench, UserBench, and AgentChangeBench — without degrading multi-turn function calling on BFCL-v4.

On-policy Distillation with Verifiable Reward

Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang et al. Reinforcement Learning with Verifiable Rewards (RLVR) gives sparse task-level feedback while on-policy distillation (OPD) gives dense token-level guidance but ignores whether a trajectory is actually correct, capping the student at the teacher's ability; existing combinations rely on weighted mixtures or heuristic switching that add hyperparameters. On-policy Distillation with Verifiable Reward (OPDVR) reformulates the implicit reward of sampled-token OPD based on trajectory correctness and applies a ReLU gate so that correct trajectories receive non-negative rewards and incorrect ones non-positive rewards, aligning the distillation signal with task success while keeping the teacher's distributional guidance and adding no hyperparameters. This modification also turns sampled-token OPD into a proper RLVR method that plugs into any policy gradient algorithm such as GRPO, and OPDVR consistently outperforms standard OPD across six reasoning benchmarks; code is released.

SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, Zihe Huang Group-relative reinforcement learning for LLM agents must wait for all sibling rollouts of a prompt before updating, which is costly when tool-use trajectories are long and variable in length; Single-stream Policy Optimization (SPO) removes that dependency with a persistent prompt-level value estimate but whitens one advantage per trajectory before feeding a token-mean actor loss. The analysis shows that centering advantages at the trajectory level generally fails to center the token-weighted quantity the actor actually consumes, and SPO++ fixes this mismatch by standardizing terminal-outcome advantages under the action-token measure, while also organizing prompt evidence by the policy event that generated it rather than the order in which the learner received it. On matched runs on ALFWorld at two model scales and on Math-TIR, SPO++ improves online learning efficiency over SPO, with a paired ablation identifying action-token-measure normalization as the strongest tested component.
3 more specialized papers

Robotics 7

Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

Brian Zhu (Siemens), Momen Khalil (Siemens), E Harrison (UC Berkeley), Emanuele Poggi (Siemens), Philipp Schmitt (Siemens), Bernd Kast (Siemens) et al. cross-listed Large vision-language-action (VLA) policies have inference latency high enough to cause pauses or jerky motion, which alters the effective environment dynamics and breaks the Markov assumption that standard reinforcement learning (RL) relies on, causing ordinary RL finetuning to fail. ARLI (Asynchronous RL with Intermediate Information) builds on asynchronous inference, which overlaps action generation with execution to hide latency, and makes it RL-compatible through a low-latency policy design plus state augmentations that incorporate already-committed actions and a mid-inference observation to restore near-Markovian structure. On simulated and real-world manipulation tasks, the method enables finetuning under inference delays where standard RL fails entirely, and matches or exceeds standard RL run in idealized no-latency settings.

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski cross-listed Vision-language-action (VLA) models pretrained on large robot datasets lose performance when adapted to new tasks from only a few demonstrations, and existing retrieval-based adaptation matches on visual similarity or whole-task language rather than reusable sub-skills. Hierarchical Skill Retrieval (HSR) decomposes a target task into candidate skill sequences, scores each plan by semantic plausibility and skill reliability estimated from the prior dataset, then retrieves demonstrations with subtask-level language matching followed by behavior-feature reranking, before adapting the policy through a two-stage pretraining and finetuning pipeline that separates general skill acquisition from task-specific adaptation. On the LIBERO benchmark and several real-world manipulation tasks, HSR improves average success rate by 10.3% and 21.3% over the strongest baseline, respectively.

PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee, Youngjae Yu cross-listed Vision-language-action (VLA) models inherit pretrained multimodal large language model (MLLM) representations but rarely use the MLLM's context window as episode memory, leaving that role to purpose-built history mechanisms. PonderPounce pairs Ponder, a System 2 MLLM that accumulates observations, demonstrations, and prior cognition in its native causal context, with Pounce, a System 1 VLA that receives the current observation, instruction, and proprioception directly and asynchronously gets only the newest cognition token and its age; both are trained jointly end to end without a separate memory module or bridge pretraining. Optimized serving reaches p50 latencies of 78 ms for cognition refresh and 25 ms for action-model invocation, supporting 20 Hz action playback. On RoboMME with base-scale data it reaches 60.83% with a 9B model versus 44.51% for FrameSamp+Modul and 17.93% for current-observation π0.5, rising to 75.54% with 9x data, and on RoboCasa-DC it reaches 12.5% from action supervision alone versus 11.6% for the strongest published demonstration-conditioned baseline.
4 more specialized papers

Reasoning 6

Recursive Agentic Reasoning

Shengxin Zhang, Xiaomin Wu, Xiyang Wu, Jing Xie Test-time reasoning methods such as iterative refinement, decomposition, and repeated sampling are usually evaluated in isolation, making their gains hard to compare across models, benchmarks, and pipelines. The authors unify them as recursion operators over an agent's reasoning trace, GROW (deepen a single path), PRUNE (decompose and recompose the problem), and BRANCH (sample alternative paths and select among them), and evaluate all three against a single-pass chain-of-thought baseline under one harness with identical prompts, token budgets, and grading code across five benchmarks and three frontier models, totalling 14 settings, 49,327 graded items, and 151,876 model calls. BRANCH improves accuracy in all 14 settings by an average of 5.98 percentage points and is the best operator in 12, whereas GROW gains 2.18 points on average but degrades two settings and PRUNE gains only 0.94; much of BRANCH's advantage comes from recovering from truncation, with its gains correlating strongly (r = 0.72) with the baseline rate of empty, budget-exhausted outputs. These results weaken the hypothesis that different problems need routing among reasoning operators, and the authors further show that unpaired evaluation or treating scoring-pipeline failures as model errors can reverse comparative conclusions, motivating paired scoring as a standard protocol.

Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning

Sophia Xiao Pu, Yumo Xu, Sailik Sengupta, Millennium Bismay, Ruixue Lian, James Gung et al. Inference-time decoding methods such as Best-of-N treat each candidate reasoning trajectory as atomic, discarding useful prefixes along with degraded suffixes. Selective Regenerative Decoding (SRD) routes each candidate to discard, keep, or regenerate only the degraded portion of the suffix while preserving the useful prefix, without needing a larger target model. Under mild assumptions it provably improves sample efficiency over rejection sampling by 1.28 to 1.36 times with strictly higher expected trajectory quality, and on MATH500, GPQA Diamond, HotpotQA, and AlpacaEval it matches Best-of-N accuracy with substantially fewer generated tokens, beating speculative rejection in low-compute regimes.

Is Discrete Difficulty Sufficient? Leveraging Continuous Difficulty for Efficient Self-Consistency in LLMs

Sihyeong Yeom, Geon Park, Geunyeong Jeong, Taewoong Yoon, Jaewook Lee, Harksoo Kim Self-consistency decoding samples many reasoning paths and takes the majority answer, which works well but burns tokens, and existing budget-saving schemes bucket problems into a handful of discrete difficulty levels that poorly reflect how continuously reasoning difficulty varies. Flexible Self-Consistency instead uses a pre-trained probe to predict the model's output entropy for a given question, treats that continuous value as an uncertainty estimate, and sets the number of sampled paths accordingly. Across several models and benchmarks it holds accuracy roughly level with full self-consistency while cutting token use by as much as 76%.

Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning

Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han et al. Long sequential reasoning traces make test-time scaling slow, and prior parallel-reasoning systems mostly exploit subtask parallelism (decomposing a task into independent chunks) while overlooking trial parallelism, where multiple speculative attempts explore, verify, and aggregate competing hypotheses at once. Parason first shows that trial parallelism accounts for 65.5% of parallelizable reasoning computation in DeepSeek-V4 traces on HLE and grows more dominant on harder problems, then converts sequential traces into structured parallel trajectories using a context-free grammar and trains models with Parallelism-Aware Group Relative Policy Optimization (PA-GRPO), whose reward jointly balances accuracy, latency, and both parallelism ratios. At inference the learned parallel structure is executed through tool calls, yielding roughly 1.7x average wall-clock acceleration on AIME24 and AIME25 with competitive accuracy.

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

Mengzhu Xu, Jifan Gao, Xia Jiang, Yaoxin Wu, Xi Long Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain actually drives the answer is rarely tested. The audit applies a battery of 30 clinically motivated perturbation operators (severity reversal, negation flip, demographic swap, evidence ablation) to both chains and questions across 14 LLMs on four medical QA benchmarks, jointly analyzing whether the chain updates and whether the answer flips to classify each model's failure mode. The Chain-Decoupling Rate, where the chain fails to register a clinically meaningful destructive edit and the answer does not flip, is 72.9% panel-wide; corrupting the chain leaves accuracy unchanged, removing CoT prompting does not reduce accuracy, and the pattern persists across medical and reasoning fine-tuning and model scale, with two board-certified clinicians confirming that 98.5% of 197 perturbed questions keep the gold answer defensible.
1 more specialized paper