Friday, August 28, 2026

397 papers cs.AI · cs.LG · cs.CL ← 2026-08-272026-08-31 →

Jul Aug Sep

Highlights

GameWAM: A World Action Model for Video Games

Highlight HF pick · 4▲Agents Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li Existing game agents map observations straight to actions without modeling world dynamics, while interactive game world models predict visual futures from given actions but cannot act as policies; World-Action Models (WAMs) unify both objectives but have not been tested in native closed-loop gameplay. GameWAM jointly generates future frames and executable keyboard-mouse trajectories through parallel visual and action generative processes with block-causal conditioning and flow matching, predicts a gameplay-versus-GUI mode at each action step with mode-specific prediction distributions and continuous-action normalization, and uses block-cycle control that plans beyond the committed horizon, executes only a short action prefix, and replans from new observations while hierarchical cross-cycle history preserves temporal continuity. It achieves competitive task success while executing fewer native actions than the compared agents, and the authors identify Low-Frequency Action Source Imprinting (LASI), where low-frequency components of the sampled action noise systematically steer coarse camera motion under fixed conditioning, a source-sensitivity failure mode in generative control.

Game agents map observations straight to actions without modeling how the world evolves, while interactive game world models predict visual futures but cannot act on their own; GameWAM unifies the two as a World–Action Model that jointly generates future frames and executable keyboard–mouse trajectories for closed-loop gameplay and GUI control in Minecraft and ViZDoom.

  • Parallel Video and Action DiTs are trained with joint flow matching under a block-causal, modality-decoupled mask (both branches read the same clean visual prefix, but noisy video and action variables never condition each other), so video co-training shapes the action context while future-video denoising can be skipped entirely at inference, and a per-timestep router selects gameplay- or GUI-specific action predictions with mode-specific normalization for continuous coordinates such as camera versus cursor motion.
  • Block-cycle control predicts P actions but commits only the first E < P before replanning from a fresh observation, and a bounded within-cycle KV cache is paired with persistent cross-cycle history (a FIFO buffer of recent Conv3D-compressed segments plus gated attention-pooled long-term memory slots) regularized by a loss that predicts the current visual feature from history alone.
  • On the 800+-task MCU benchmark, GameWAM reaches 50.7 average ASR on Mini and 46.6 on All versus 42.5 All for Game-TARS and 36.8/31.5 for OpenHA, while using 138/155/203 native steps for embodied/GUI/combat tasks against roughly 290–400 for every baseline, and it also beats Game-TARS on all four ViZDoom maps, all without large-scale multi-game pretraining.
  • Ablations show future-video supervision matters most (dropping it cuts average Mini ASR from 50.7 to 35.7), followed by coarser temporal sampling (36.7), removing event-anchored clip sampling (38.0), a unified gameplay/GUI action distribution (38.3), and matching prediction to execution horizon (41.3), while removing cross-cycle history costs little overall (46.7) and actually raises embodied ASR to 75.0.
  • The authors also isolate a failure mode they call Low-Frequency Action Source Imprinting: reusing the same sampled noise across replanning steps induces persistent directional camera bias and in-place spinning, with DCT interventions showing yaw DCT0 correlates with the source at r = 0.890, low-frequency source replacement transfers the donor's output in 94.8% of trials, and zeroing that band removes 99.25% of the associated variance, so the model resamples every cycle as a mitigation that hides but does not remove the sensitivity; the results also carry large per-task variance (±25–38), the Mini subset uses only 10 runs per task, Game-TARS still leads on embodied (50.4 vs 47.5) and combat (38.1 vs 32.2) ASR All, and ViZDoom rewards are given only graphically.

Same Model, Different Harness: Different Coding-Agent Results

Highlight Agents Sydney Lewis A coding agent is a model plus a harness that decides what the model sees, which tools it can use, and how work continues, so the question is whether changing only the harness changes outcomes with model and task fixed. The authors compare two configurations of one harness on SWE-bench Verified, SWE-bench Pro, and FeatureBench: the control feeds the full conversation in time order, while the treatment keeps the same record but mechanically shortens older tool results as the context fills and reacts to repeated or stalled work. On 169 Verified tasks under a tight 20,480-token window and a fixed 480-second attempt endpoint, the treatment raises mean per-task fail-to-pass fraction from 28% to 49% and complete solutions from 43 to 72, and the same frozen treatment also lifts three additional models without retuning; with wide windows the arms are close on Verified and Pro while FeatureBench still favors the treatment and it serves fewer prompt tokens per turn, leading the authors to argue that evaluations should treat model and harness together as the solver under test.

Coding agents pair a model with a harness that controls what the model sees, which tools it gets, and when a run stops; the paper asks whether swapping only the harness configuration, with model and tasks fixed, changes benchmark outcomes. The Yuj harness is compared in two frozen configurations: a control that feeds the full chronological transcript until it overflows the context window, and a treatment that keeps the full record but mechanically shortens older tool results as the window fills and injects fixed interventions when it detects repeated or stalled work.

  • The treatment's "half-life" view rule activates once the estimated prompt reaches half the window, keeps the newest four tool results intact, and truncates older ones to caps that halve as their age doubles, while a rule-based detector (no model calls) flags patterns such as a repeated command producing the same failure and responds with a canned nudge to change approach; command safeguards also normalize test invocations and block oversized setup output.
  • Under a tight 20,480-token window on a 169-task SWE-bench Verified cohort with Qwen3.6-35B-A3B, treatment raised mean per-task fail-to-pass fraction (F2PF) from 28% to 49% and complete solutions from 43 to 72; on a 316-task SWE-bench Pro cohort F2PF rose from 15% to 33% and solutions from 31 to 72, and on FeatureBench F2PF rose from 11% to 20% (solutions only 2 to 3), with all paired F2PF sign tests at p < 0.0001.
  • The same frozen package, without retuning, improved every additional model tested at the same window: Devstral more than doubled (F2PF 17% to 37%, solutions 22 to 53), Qwen3.8 went from 20% to 35% F2PF, and Nemotron showed the smallest gain at 12% to 18% F2PF.
  • The advantage is largely a context-pressure effect: at a 43,008-token window the Verified gap shrank to +6.4 F2PF points, and at an effectively unconstrained 262,144 tokens Verified and Pro outcomes were statistically indistinguishable (Verified difference −0.3 points, 95% interval [−4.5, +3.9]), though treatment still served 7.2% fewer prompt tokens per turn and FeatureBench retained a 23.9% to 30.7% F2PF gain that survives the task-level test but not the repository-level one.
  • Key caveats: the comparison is not compute-matched, since treatment consumed roughly 2 to 3 times the model turns and prompt tokens because control often stopped early at the context limit; each task gets a single greedy trajectory per arm; results cover only four locally served four-bit open-weight checkpoints, and the treatment's operating point was tuned during development rather than on held-out tasks.

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

Highlight HF pick · 20▲Agents Yang Xiao, Yusong Sun, Haoyi Wu, Wenyang Hui, Wen Da, Zhaokai Luo et al. Most self-improvement methods for long-horizon agents process a run's experience only after it ends, so they can neither redirect the active run nor immediately validate lessons drawn from it; single-agent self-correction mixes execution and assessment in one context, and subagent delegation typically cannot steer a running subagent. PILOT is a supervisor-worker harness with two coupled mechanisms: live steering, in which a separate supervisor can redirect or abort the active worker mid-execution, and live self-evolution, which distils procedures and failure modes observed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks it ranks first in five of six configurations, beating counterpart harnesses on Terminal-Bench 2.0 by up to 9.8 percentage points. In the self-improvement setting it gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6 while cutting mean output tokens by 42.9% and 47.4%, raising successful evaluations per million output tokens by 110.3% and 134.0%.

Most agent self-improvement methods (reflection, judge-based evaluation, self-evolving harnesses) only process a run's experience after the run ends, so they can neither rescue the current attempt nor immediately validate what they learned. PILOT makes self-improvement live: a frozen supervisor stays connected to a worker over a two-way channel throughout execution, redirecting or aborting it mid-run while distilling procedures and failure modes into a persistent skill library and memory that later workers load.

  • The worker executes in an isolated context and emits notifications, questions, and results, while the supervisor keeps its own context focused on the goal and recent events, inspects the relevant slice of the worker's trajectory only when diagnosis is needed, queues Steer guidance for the worker's next turn or issues Abort, and writes reusable knowledge into the harness so any worker spawned afterward starts from the updated state; it is built as an extension of the Pi coding-agent runtime with the same frozen model in both roles.
  • In the one-shot setting with frozen GLM-5.1 and Kimi-K2.6 backbones, PILOT ranks first in five of six backbone–benchmark combinations across Terminal-Bench 2.0, SWE-bench Multilingual, and SWE-bench Pro, reaching 71.9% and 71.3% on Terminal-Bench 2.0 (up to 9.8 points above counterpart harnesses, 5.3 above Pi on average) and 59.9% average on SWE-bench Pro, with the largest gains on Hard tasks (55.0% on both backbones).
  • In the self-improvement setting, where skills written during successful runs are carried into the next iteration (verifier outcomes only gate retention, never content), the best-so-far Terminal-Bench 2.0 pass rate rises 14.6 points (66.3% → 80.9%) on GLM-5.1 and 12.4 points (68.5% → 80.9%) on Kimi-K2.6, versus 7.9 points for OpenCode and 2.3 for Pi from the same initial library, while skill libraries grow from 62 to 83 and 50 to 81, mean output tokens per task fall 42.9% and 47.4%, and successful evaluations per million output tokens rise 110.3% and 134.0%.
  • Manual trace inspection classifies only 2.3% (GLM-5.1) and 10.6% (Kimi-K2.6) of successful one-shot runs as genuinely aided by live steering, with zero on Easy tasks and 6.1% / 19.7% on Hard, so most of the one-shot gain is not directly attributable to trace-verified corrections and the mechanism matters chiefly on long, fragile execution chains (the case studies show the supervisor redirecting a stalled CoreWars strategy and catching a bias-added-before-all_reduce bug in a tensor-parallel layer).
  • Caveats: pass rates are means over only two runs, the self-improvement iterations repeat the same 89 tasks with skills named exactly after each task, so gains reflect reuse on previously seen tasks rather than transfer to new ones, the headline improvements are best-observed rather than final-iteration numbers, and the authors note the evaluation is limited to two open-weight backbones and three benchmarks with supervisor–worker heterogeneity left unexplored.

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

Highlight Reinforcement Learning Gyouk Chu, Myeongho Jeon, Eunho Yang Self-evolving language models cut the cost of human supervision, but progress has concentrated in verifiable domains where answers can be checked automatically. J-Zero (Judge co-adaptation from Zero data) trains a Challenger, a Solver, and a Judge together starting from no data: the Challenger invents progressively harder tasks while the Solver learns better responses, and the Judge is trained on preference pairs whose ordering is known a priori from provenance — the Solver's answer ranks above the Challenger's, and a decomposed-then-recombined answer above a one-shot one — rather than from the Judge's own scores. It beats baselines by 4.2 points on average in verifiable domains and 8.0 in unverifiable ones, and keeps improving for at least ten iterations while the baselines degrade after two.

Zero-data self-play, where a Challenger invents tasks and a Solver learns to answer them, works when a verifier or majority vote can score answers, but in open-ended domains the reward comes from a Judge model, and a frozen Judge caps how far the Solver can climb. J-Zero closes the loop by training the Judge too, using preference pairs whose ordering is fixed by how the responses were produced rather than by the Judge's own scores.

  • Each iteration runs a GRPO minimax game under the current Judge (the Challenger is rewarded for tasks the Solver scores poorly on, with a BLEU-based repetition penalty and format check; Solver tasks are selected by score dispersion), then updates the Judge with a Bradley–Terry loss on two pair types: role-asymmetry pairs (Solver's answer preferred over the Challenger answering its own task) and subtask-amplification pairs (a decompose-solve-compose answer preferred over the Solver's one-shot answer).
  • On Qwen3-4B-Base and Qwen3-8B-Base across 11 verifiable benchmarks (7 math, MMLU-Pro, SuperGPQA, BBH, IFEval), J-Zero reaches overall averages of 54.38 and 58.55 versus 49.64/54.99 for R-Zero and 47.41/53.07 for G-Zero, a +9.47 and +7.88 gain over the base models.
  • The gap is widest on unverifiable tasks, where the 4B model's AlpacaEval 2.0 length-controlled win rate rises from 6.22 to 28.56 (8B: 12.93 to 33.53) and the three-benchmark unverifiable average improves by +11.23 and +10.18 points, versus roughly 2–3 points for R-Zero and G-Zero.
  • J-Zero keeps improving monotonically through 10 iterations while both baselines peak at iteration 2; a frozen-Judge ablation tracks the full method for three iterations then plateaus 1.66/4.44 points lower, and removing amplification pairs hurts more (−1.64 overall) than removing role pairs (−0.97), with an external Claude Opus 4.8 check confirming role-pair labels are correct 87.9%→~66% of the time but amplification labels only 21.1% at iteration 1, crossing 50% at iteration 4.
  • Caveats: models are capped at 8B base checkpoints, the Judge is an off-the-shelf discriminative reward model (Skywork-Reward-V2-Llama-3.1-8B) rather than a third instance of the same base model, the Judge is trained on majority-wrong amplification labels in early iterations, the best checkpoint is picked by watching the evaluation benchmarks themselves, and the unverifiable scores depend on LLM judges (Qwen3.6-27B, gemma-4-31B-it).

Squeezing More from Limited Data with Recursive Transformers

Highlight Large Language Models Serdar G\"ulbahar, Lukas Edman, Alexander Fraser When the data budget is fixed but compute is plentiful, adding parameters to a language model helps only up to an optimal size, beyond which overfitting hurts generalization. Studying pre-training budgets of 10M to 100M words across two corpora and several downstream evaluations, the authors find the optimal size depends strongly on both the data budget and the target task, and argue that standard Transformers scale down poorly because embeddings eat a large share of the parameter budget and per-token compute is tied to representational capacity. To decouple these, they study recursive Transformers that reuse a shared block across depth together with factorized embeddings, and the three recursive models trained outperform standard Transformers at both 10M and 100M words while remaining competitive with BabyLM Challenge 2025 winners.

Under a fixed data budget (10M–100M words, the BabyLM regime) with abundant compute, standard Transformers show non-monotonic scaling: beyond a data- and evaluation-dependent optimal size, extra parameters overfit, and at small scales the fixed-size vocabulary embeddings eat most of the budget while per-token computation is tied to representational capacity. RecursiveGPT decouples the two by applying a single shared Transformer block R times (with step-specific RMSNorm and bias parameters as lightweight depth conditioning) and using ALBERT-style factorized embeddings, so compute can be scaled without adding trainable capacity.

  • A standard-model sweep (12 layers, hidden size 256–3072, 13M–1.2B params, BabyLM and Nemotron-ClimbMix corpora) finds the optimal scale depends on the target: BLiMP saturates early, COMPS keeps benefiting from larger models at 50M–100M words, and training loss keeps falling even as downstream scores decline, pointing to overfitting rather than optimization failure.
  • At 10M words, a 27.6M-parameter depth-16 RecursiveGPT with factorized embeddings beats the best standard model on all three evaluations (45.80 vs 45.07 average across BLiMP, COMPS, LAMBADA pass@5); in the depth sweep, performance rises up to R=16 and then dips slightly, so recurrence is not free of optimization limits.
  • At 100M words, the 124M recursive model exceeds the 1.22B standard model on BLiMP (80.06 vs 79.71) using about a tenth of the parameters but trails on COMPS and LAMBADA, while a 404M RecursiveGPT-Large reaches the best average (63.93 vs 63.47) and beats the GPT-BERT causal baseline on every BabyLM benchmark; at 10M words AMLM hard decay still wins on COMPS and average, with Entity Tracking as the recursive model's weak spot.
  • Ablations show factorized embeddings matter more for the recursive model than for standard ones (dropping them costs over a point of average even at matched parameter count), while sharing normalization parameters across steps barely hurts, and spending the equivalent FLOPs on more epochs of a standard model instead degrades results (45.07 to 41.47 at 10M, 63.47 to 56.68 at 100M).
  • The main cost is compute: the recursive models use roughly 15x the estimated training FLOPs of their 10M-word standard counterpart and take up to 64 A100-hours at 100M words, hyperparameters were not retuned per configuration, and only a single shared block with fixed depth was explored, though a logit-lens analysis finds about 80% of tokens settle before the final recurrent step, suggesting adaptive halting could recover much of that cost.

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

Highlight HF pick · 40▲Agents Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin et al. Large language model (LLM) agents increasingly learn from generated interaction data, but the literature is organized by domain and evaluated heterogeneously, which obscures shared generation mechanisms and conflates constructing candidates with verifying and selecting them. This survey represents agentic data as a factorized object (E, q, τ, v) consisting of an environment specification, task signal, interaction realization, and optional verifier, and organizes generation paradigms by their primary anchor and dependency structure. It then frames generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens: accuracy defines the feasible support of grounded, internally consistent data, complexity places learning mass relative to a declared learner's capability and execution configuration, and diversity governs coverage and redundancy. The reviewed work shows a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size, and the authors argue the central challenge is continually allocating valid, informative, non-redundant experience as agents and environments evolve.

LLM agents are increasingly trained on synthesized interaction data, but the literature is organized by domain (tool use, coding, GUI, embodied, scientific) and routinely conflates how candidate data are constructed with how they are verified and selected. This survey proposes a common factorized data object (E, q, τ, v) — environment specification, task signal, interaction trajectory, and optional verifier — plus an Accuracy–Complexity–divErsity (ACE) lens that frames generation as constrained distribution design: accuracy defines the feasible support, while learner-relative complexity and batch-level diversity decide which valid data are worth keeping.

  • Generation pipelines are taxonomized by their anchor factor rather than by domain: forward pipelines follow E→q→τ from real (ToolLLM, APIGen, TOUCAN), LLM-synthesized (ToolACE, ToolWeave), or programmatic executable environments (EnvScaler, EnvFactory, ScaleEnv), while reverse pipelines are task-first (BUTTON, AgentInstruct), trajectory-first (OS-Genesis, Learn-by-interact), or structure-first (APIGen-MT, Magnet), with self-evolving systems (AgentEvolver, WebEvolver) as a cross-cutting extension.
  • The ACE objective is deliberately asymmetric: accuracy is a conjunctive admission gate over all four factors (a clean trajectory cannot rescue an infeasible task, and a correct terminal state cannot rescue a verifier that accepts policy-violating shortcuts), whereas complexity is defined learner-relatively as verified failure probability under a declared execution configuration (model, scaffold, tools, verifier, budget) with utility peaking in a learnable band near the capability frontier rather than being maximized blindly.
  • Accuracy assurance across prior work clusters into four mechanisms — layered rule/model/human checking, constraint-grounded construction via blueprints and tool graphs, execution- and state-based verification, and feedback-driven repair — and the survey highlights EnvFactory's finding that a smaller set of robustly verified environments can compete with simply scaling environment count.
  • Horizon and tool-call count are explicitly rejected as universal difficulty scores; the paper instead proposes a paired criterion p_z0(d) < ρ ≤ p_zA(d) that isolates tasks a base configuration fails but an agent-assisted configuration solves, separating genuinely agent-requiring tasks from both the trivially solvable and beyond-frontier tails.
  • As a survey it reports no new experiments or benchmark numbers, ACE intentionally omits cost, efficiency, and safety, and the authors concede persistent gaps: verifiers that inherit the generator's biases, repeated repair narrowing data toward verifier-friendly patterns, and residual semantic uncertainty (inefficient, unsafe, or intent-inconsistent trajectories) that no executable check fully captures.

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Highlight HF pick · 68▲Vision Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram {\DJ}or{\dj}evi\'c, Shiyang Li et al. Video generation models are increasingly treated as world models, but many physical processes can unfold in more than one valid way, so a world model should reproduce the distribution of outcomes under the same initial observation and action, a requirement the authors call probabilistic alignment. Existing evaluations judge individual-video plausibility rather than whether repeated generations recover the correct distribution. PAWBench evaluates video generators as stochastic samplers of world dynamics, and the PAWEval protocol converts repeated rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors, and the authors further test whether language prompts, initial noise sampling, or model training can reshape the predictive distribution.

Video generators are increasingly pitched as world models, but existing benchmarks score each generated clip in isolation and never ask whether repeated rollouts from the same image and action reproduce the correct distribution of physically valid futures. PAWBench formalizes this as probabilistic alignment and tests eleven current video models as stochastic samplers, finding that none matches reference probabilities while also covering the full set of valid outcomes.

  • The benchmark holds a source image and action prompt fixed, samples K = 50 rollouts per scenario, and has PAWEval (a rubric applied by Gemini 3.5 Flash) map each video to a terminal outcome; PAW-Calibration (25 scenarios with analytic reference distributions such as coin tosses and spinners) scores total variation distance against the reference, while PAW-Coverage (25 scenarios with enumerable but unquantifiable outcomes such as bottle flips) scores the fraction of valid outcomes ever observed.
  • Cosmos 3 Super I2V has the best calibration (TVD × 100 = 20.5) but passes only 80% of scenes' readout gate, LTX-2.3 has the best coverage (71.7%) but passes only 72% of coverage scenes, and across all eleven systems average TVD is 31.2 versus a Monte Carlo ceiling of 9.22 from finite sampling alone, so the gap is not a sample-size artifact; PAWEval agrees with human labels on 81.3% of 888 clearly-resolved videos.
  • Paired interventions show models underreact to causal changes (tilting a pencil shifts outcomes incompletely or in the wrong direction) yet overreact to non-causal cues (distractor text on a Galton board redirects probability mass), indicating their distributions do not track whether the physical transition actually changed.
  • Attempts to fix the gap are only partially effective: VLM-driven prompt engineering with GPT-5.5 raises calibration TVD for every generator because the VLM itself selects a misaligned distribution, oracle prompting with correct targets helps but generators realize only 37.6–58.1% of requested outcomes, and coupled-noise sampling (C2C) modestly improves both metrics without changing the learned distribution.
  • LoRA fine-tuning Wan2.2 on pencil-fall videos with varying left/right ratios shifts outcome frequencies nonlinearly, but no mixture satisfies both a 50/50 upright pencil and a 100/0 leaning pencil at once, showing that global frequency adjustments give only coarse control and that true alignment requires learning how the distribution should vary with initial physical state; the benchmark also only evaluates terminal outcomes, not trajectory-level dynamics.

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Highlight HF pick · 10▲Reinforcement Learning Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong et al. Evolution Strategies (ES) have emerged as a memory-efficient post-training paradigm for LLM reasoning, but their optimization behavior relative to Group Relative Policy Optimization (GRPO) has been poorly understood. A theoretical analysis shows that verifier-projected Jensen-Shannon diversity across the ES population helps Pass@K, and empirically ES improves Pass@1 while attaining higher Pass@K than GRPO, which exhibits entropy collapse; a sequential GRPO-then-ES schedule combines GRPO's Pass@1 strength with ES's Pass@K gains. Despite substantial whole-model parameter drift, ES's task gains come from a sparse subset of larger-magnitude updates, held-out evaluations show this need not cause catastrophic forgetting, and larger LLMs require smaller ES population sizes. The authors position ES as a distinct post-training paradigm rather than a weaker memory-efficient substitute for GRPO.

Evolution Strategies (ES) post-train an LLM by perturbing its weights, scoring each perturbed copy with a verifier, and stepping along the reward-weighted average perturbation with no backpropagation, but how this differs from mainstream GRPO beyond memory cost was unclear. The authors argue theoretically and empirically that ES's population-based parameter-space search preserves broader reasoning coverage, improving Pass@1 while also raising Pass@K, whereas GRPO's entropy collapse buys larger Pass@1 gains at the expense of large-K performance.

  • ES draws N Gaussian perturbations of the full parameter vector, z-scores the rollout rewards within the population, and updates the center model along the normalized reward-weighted direction; the theory shows the perturbed members form a diverse policy population whose verifier-projected Jensen–Shannon diversity makes one sample per member at least as likely to contain a correct answer as the same number of draws from a single policy, and gives sufficient conditions (Proposition 1) under which the center update inherits that Pass@K margin.
  • On GSM8K-trained Qwen2.5-1.5B/7B-Instruct and Llama-3.2-3B-Instruct and DeepScaleR-trained DeepSeek-R1-Distill-Qwen-1.5B, ES improves average Pass@1, Pass@16 and Pass@32 over the base model in both settings, while GRPO wins Pass@1 but falls below the base model on Pass@16/32 in 15 of 18 Easy-setting comparisons — for Llama-3.2-3B-Instruct average Pass@32 is 80.4 for ES vs 77.0 for GRPO (base 78.6), and held-out GPQA token entropy stays flat under ES while dropping sharply under GRPO.
  • Splitting the same update budget sequentially as ES→GRPO gives the best hard-setting math-average Pass@32 (79.2 vs 78.0 GRPO, 78.9 ES) while retaining most of GRPO's Pass@1 gain (52.3 vs 52.9), adding new non-dominated points to the Pass@1–Pass@K Pareto front.
  • ES moves 40.7–44.1× farther from initialization in relative L2 than GRPO, yet 77.6–93.0% of nonzero updates are ≤1.5e-3 in magnitude and zeroing them barely changes target-task Pass@1; the large updates concentrate in LayerNorm and attention projections (72–80 of the top 100) rather than the embeddings and LM head GRPO favors, and held-out average Pass@32 change is positive under ES but negative under GRPO for all three Easy-setting models, so the authors attribute prior catastrophic-forgetting reports to training-set overfitting rather than drift itself.
  • Z-score reward normalization is essential, the antithetic two-point estimator brings no gain because regenerated reasoning rollouts weaken paired covariance, and the required population shrinks with scale (N=16 is within 0.01 reward of N=64 for 1.5B and 3B models but only N=32 is for 0.5B); the caveats are that GRPO still leads Pass@1 by 1–3 points, ES's Pass@K edge is often under a point, all models are ≤7B, and continual-learning effects of the large drift remain untested.

TTPO: Test-Time Policy Optimization

Highlight HF pick · 42▲Reasoning Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu et al. Post-training methods such as reinforcement learning (RL) and On-Policy Self-Distillation (OPSD) have advanced LLM math reasoning but need ground-truth labels, which rules out test-time training (TTT); majority-vote pseudo-labels are a natural substitute, but a wrong vote corrupts the teacher for every token. The authors observe an asymmetry: rollouts that disagree with the pseudo-label are usually wrong regardless of whether the vote itself is correct. Test-Time Policy Optimization (TTPO) exploits this by distilling agreeing rollouts via OPSD and penalizing disagreeing rollouts with grouped RL, with token-level selection that down-weights already-converged positions in distillation and penalizes only confident errors in RL. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% under test-time training, adds 25.2% to 36.4% in non-thinking mode, and generalizes across tasks.

Post-training methods like RLVR and on-policy self-distillation (OPSD) need ground-truth answers, which rules them out for test-time training on unlabeled problems, and swapping in majority-vote pseudo-labels is fragile because a wrong vote corrupts the answer-conditioned teacher at every token. TTPO exploits an asymmetry — rollouts that disagree with the vote are usually wrong even when the vote itself is wrong — by distilling only agreeing rollouts via OPSD and penalizing only disagreeing rollouts via GRPO, with token-level selection in both branches.

  • For each problem, K=64 rollouts are sampled and split by whether their answer matches the plurality answer; agreeing rollouts get a forward-KL loss toward the same model conditioned on the pseudo-label (weighted by a soft-OR of student entropy and teacher-student divergence to skip converged tokens), while disagreeing rollouts get negative group-relative advantages applied only to the top-50% of tokens by confidence-weighted negative log-probability, with the RL branch scaled by λ=0.1.
  • The motivating measurement on AIME 2026 with Qwen3-1.7B is that pseudo-labels are wrong for ~85% of prompts, yet ~79% of disagreeing rollouts are also wrong, so a disagreement penalty stays correct regardless of vote quality, and when the vote is wrong the distillation branch degrades to harmless thinking-to-non-thinking self-distillation rather than steering toward an arbitrary error.
  • Trained label-free on OpenThoughts, TTPO matches or edges out label-supervised OPSD on five competition benchmarks (40.1 vs 39.7 on 1.7B, 62.6 vs 61.7 on 8B average), and in pure test-time training on AIME26/HMMT26/BRUMO25 it lifts Qwen3-1.7B from 38.0% to 45.2%, beating TTRL (40.2) and an OPSD-TTT baseline (41.9), with Qwen3-4B after TTT (61.1) surpassing the untrained Qwen3-8B (60.7).
  • With thinking mode disabled at inference, gains are far larger than OPSD's: +25.2 (1.7B), +30.6 (4B), +36.4 (8B) average points versus +7.1/+5.8/+3.5, and reversing the branch assignment (GRPO on positives, distillation on negatives) collapses to 37.2 on AIME26, below the base model; surprisingly, replacing pseudo-labels with ground truth also hurts because hard problems yield near-zero positives and vanishing advantages.
  • Limitations include experiments confined to math with verifiable final answers, LoRA-only fine-tuning with gradients on only the first 1,024 completion tokens, reported peak-checkpoint numbers rather than final ones, and dependence on majority-vote quality that breaks down when the sample budget is small or no rollout is correct.

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Highlight HF pick · 1▲Agents Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, Tu Vu Agent skills bundle specialized knowledge and workflows into reusable resources, and recent methods discover them automatically from agent experience, but the insights that drive skill improvements stay scattered across optimization histories and are rarely reused. WikiSkill co-evolves skills with a persistent knowledge base by separating raw execution experience, accumulated knowledge, and executable skills, continuously consolidating experience into a wiki that later skill updates build on. Across diverse benchmarks and models it consistently outperforms state-of-the-art skill-evolution methods and beats no-skill baselines in most settings; larger models generally benefit more from evolved skills, while smaller models equipped with skills can outperform substantially larger models without them. Evolved skills also transfer across models and model families, skills evolved by other models can beat self-evolved ones, and ablations show the persistent wiki is critical to the gains.

Skill-evolution methods that refine agent skills from execution traces tend to leave what was learned scattered across optimization history, so each iteration re-derives lessons rather than building on them. WikiSkill inserts a persistent, never-rolled-back knowledge base (a "wiki") between raw execution traces and executable skills, so that skill proposals draw on consolidated failure patterns, an audit trail of past proposals, and cross-iteration logs instead of isolated trajectories.

  • The workspace has three layers: an immutable raw/ trace store, a wiki/ of pattern pages plus an evolution log and a skill-impact.md tracker of every proposal diff and its accept/reject outcome, and a skills/ layer of SKILL.md files; each iteration a Wiki Maintainer consolidates sampled traces into the wiki, a ReAct-style Skill Proposer reads the wiki and traces on demand to emit one atomic skill edit, and a validation gate accepts the edit only if it strictly improves the score (rolling back skills but never the wiki).
  • Across five benchmarks (LiveMathematicianBench, SealQA, SpreadsheetBench, OfficeQA, ALFWorld) and five models, WikiSkill beats Trace2Skill, EvoSkill, and SkillOpt on average for every model, by 3.3 to 12.0 points over the strongest baseline, with standout gains such as Gemini-3.5-Flash going from 33.0% to 72.6% on LiveMath and Qwen-3.6-27B from 40.8% to 81.7% on SpreadsheetBench.
  • Skill evolution complements scale rather than substituting for it: average gains within the Qwen family grow from +12.3 (4B) to +17.5 (9B) to +23.9 (27B) points, yet Qwen-3.5-9B with evolved skills (47.4%) outperforms Qwen-3.6-27B with none (39.4%).
  • Skills transfer across models and families and often beat self-evolved ones (e.g. Qwen-3.6-27B skills lift Gemma-4-31B to 73.7% on LiveMath vs. 56.7% with its own), but transfer can also be sharply negative when a small model's skills encode low-level workarounds (Qwen-3.5-4B spreadsheet skills drop Gemini-3.5-Flash from 50.5% to 18.1%).
  • Ablations on Gemini-3.5-Flash show giving the Skill Proposer wiki access raises the four-benchmark average from 48.7% to 63.7%, while letting the inference agent read the wiki during rollouts hurts (63.7% to 60.9%); caveats include full prompt injection of skills (no retrieval or triggering evaluated), a strict improve-only acceptance rule that discards neutral edits, no wiki pruning, and no very long-horizon tasks.

Applications 96

Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

Ziqiang Zhang, Jing Ma, Zilong Wang, Jiayuan Chen, Yi Qiao, Yu He et al. Pricing automation in large-scale tourism faces unstructured travel orders and complex, fast-changing, open-ended policies, where rule engines are brittle and unconstrained LLM agents lack the reliability and auditability needed for financial decisions. The system enforces a strict decision boundary: LLMs perform structured extraction and bounded policy and path selection, policies are compiled into interpretable condition trees that accommodate new clauses without code changes, and all numeric pricing is computed deterministically, with periodic fine-tuning on logged traces improving tree induction and path matching. Deployed at a municipal state-owned tourism enterprise across 7 scenic sites, 12 business categories, and over 1,000 active policies, it processed 3,960 orders in six months, cut the order management team from 15-20 people to 3, and reduced per-order handling time from 10 minutes to under 2.

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

Prabhjot Singh, Somnath Luitel, Manmeet Singh, Josh Durkee Existing peer-review corpora for training AI reviewers come almost entirely from computer science and machine learning venues, so models never see the kinds of critiques biologists or chemists raise, such as demands for contamination controls or challenges to spectral assignments. FIRSTPASS is built from Nature Communications' mandatory transparent peer review (in place since November 2022) and contains 3,668 records across biology, chemistry, neuroscience, physics, and earth science, each preserving the full multi-round exchange of referee reports, point-by-point author responses, and updated reviewer assessments. Every record carries an outcome label derived from the editorial decision (STANDARD for two-round review, EXTENDED for three or more rounds), reviews average 2,155 words, and the data, parsing pipelines, and evaluation scripts are released for benchmarking AI scientific judgment across disciplines.

Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models

Orhan Yagizer Cinar, Timur Emre Ozkose, Emma Von Hoene, Amira Roess, Taylor Anderson, Hamdi Kavak Systematic literature reviews (SLRs) are labor-intensive, and this study builds a large language model (LLM) pipeline to extract model-relevant fields from 536 peer-reviewed agent-based modeling papers on disease spread, then compares the output against a human-conducted review. Paper-level accuracy reached 77.95% for GPT-4.1 and 81.67% for GPT-5.0, while field-level accuracy ranged from 32.40% to 100%, with complex or subjective fields the least reliable. Agreement between models emerged as a useful quality signal: low agreement tends to indicate hallucination, while high agreement paired with low accuracy points to noise or errors in the human-labeled dataset.

DRL: A Deterministic Relational Middleware Layer for Transaction-Safe Enterprise NL2SQL Under Schema-Graph Scaling

Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik cross-listed Natural-language-to-SQL interfaces over enterprise online transaction processing (OLTP) databases break down as schemas grow, because full-catalog prompts exceed what language models can attend to reliably. DRL is a middleware layer between front-ends and SQL backends that combines dynamic context pruning, relational abstract syntax tree (AST) typing, and transactional safeguards such as EXPLAIN gating and NULL guards to bound context and flag operational silent divergence (SDop). On PostgreSQL the dynamic router cuts prompt context by 92% relative to naive full-catalog prompting with sub-millisecond pruning latency, and GPT-4o, Claude Sonnet 4.5, and Gemini 2.5 Flash all land around 52-53% execution match on a 1,000-pair verification suite, with SDop flagging 89-100% of queries that pass execution match but are actually wrong. A single regex bug in the evaluation post-processor had silently manufactured a false 4-10% cross-vendor accuracy gap that disappeared once fixed, which the authors present as evidence that benchmark code needs the same scrutiny as the models it scores.

Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata

Meiwei Zhang, Eduardo Miranda, Bruce Baynes, Suvigya Jain, Wanlong Chen, Tao He et al. Choosing and planning around managed LLM services still leans on capability benchmarks that say little about how a service behaves once deployed. Operational Embedding (OpEmbed) aggregates model-time windows of structured, privacy-preserving support-case metadata, with no case text, into an eight-channel operational signature and learns a compact representation via temporal contrastive learning, cross-view reconstruction, and a generational-ordinality regularizer. Evaluated on more than 33,000 production support cases covering seven LLM families over 26 months at Google Cloud, the learned fingerprints recover interpretable family- and version-level structure, beat non-learned baselines on leave-one-model-out operational forecasting, remain useful with limited early-window data, and support cross-model fault-type transfer. The authors also share practical lessons for model onboarding, support readiness assessment, and operational monitoring.

Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI

Architect Labs cross-listed Hardware architectures are committed years before silicon ships while AI workloads shift in months, so designs hedge with generality and then map poorly onto frozen chips. The authors present an end-to-end AI system that collapses the software-to-silicon stack into a single co-design and verification loop, and use it to build Redwood, an accelerator for single-batch, low-power, ultra-low-latency inference for physical AI: from a high-level specification written by two human architects, the system generated the performance model, register-transfer-level (RTL) design, Universal Verification Methodology (UVM) environments, formal proofs, firmware, and kernels in under two weeks with no human intervention below the specification, with every block reaching 95% coverage and specification changes reverified and redeployed to hardware within 48 hours. The FPGA variant Redwood Nano runs multi-billion-parameter models such as Llama and Qwen, and projected onto Samsung 8 nm the design delivers 1.75x the throughput at 1.9x lower power than a measured Jetson Orin Nano baseline, a 3.4x performance-per-watt gain; the authors also report that Qwen running on Redwood helped design the next generation, which they frame as an early step toward recursive self-improvement.

Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models

Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach, Yanjun Gao Biomedical knowledge graphs (KGs) hold structured medical facts that could ground language model reasoning for clinical diagnosis, but the right way to inject that signal into a model is unsettled. The authors compare five KG task formulations, three training paradigms, two knowledge graphs, and three base models, and introduce two optimization-geometry metrics — Gradient Intervention Density and Gradient Distortion — that measure how broadly an optimizer rewrites the pretrained weights. Every paradigm beats the non-finetuned baseline, but methods with similar in-domain accuracy transfer knowledge very differently: KG-judgment training with KL regularization produces sparse localized updates, a regime they call surgical alignment, while task-specific supervised fine-tuning produces dense ones. An ablation shows the objective and the KL term contribute to sparsity independently, and the sparse-update paradigms improve reasoning quality even when their in-domain accuracy trails supervised fine-tuning.

SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations

Hiep V. Dang, Antonios Mamalakis Subseasonal-to-seasonal (S2S) precipitation forecasting, covering lead times of weeks, matters economically but suffers from weak predictive signal, large uncertainty, and operational systems too expensive to run at high fidelity. SimCast-S2S is a probabilistic S2S system built on latent diffusion: it generates ensembles by sampling from a conditional distribution rather than emitting a single deterministic forecast, works in a compact space learned by variational autoencoders so large ensembles stay cheap, and gets around diffusion's data appetite by pretraining on large climate-simulation ensembles then fine-tuning on limited reanalysis data with low-rank adaptation (LoRA). It beats convolutional and U-Net deep learning baselines on reanalysis data and, despite using only a subset of atmospheric inputs and no bias correction or calibration, is competitive with and often better than the operational ECMWF-S2S baseline.

hoBIT: A Profile-Aware Retrieval-Augmented Chatbot for University Academic Advising

Yoonseo Kim, Seongmin Lee, Joongheon Kim, SeongKu Kang cross-listed University advising chatbots must answer the same question differently depending on a student's department, admission cohort, and degree program, so a retriever that ignores the student's profile can surface evidence that is plausible but does not apply. proFILL converts a rule-based advising chatbot into a profile-aware retrieval-augmented generation (RAG) system that, instead of demanding a full profile upfront, progressively asks only for the attributes each query needs, guided by the query intent and the initially retrieved evidence, and then conditions retrieval on a profile-aware index. In experiments and a human preference study, proFILL outperforms diverse RAG baselines and is preferred by target users, and it remains effective with open-weight models for cost-effective on-premise deployment.

FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs

Zeming Liu, Hang Lyu, Jingtao Zhang cross-listed Automatically generated operational programs tend to be checked either by a handful of hand-written examples, which miss sparse boundary and interaction faults, or by exhaustive regression suites that cost far more than necessary. FaultLens learns small behavioral test suites instead: it runs a rich probe domain once, caches which probes kill which faults, and learns probe orderings only from earlier program generations, combining a fault-driven greedy component that exploits known kill structure with a mutation-independent diversity component covering probe families, cases, templates, and time bins so it still works when a new program has an unseen fault mechanism. Across twenty generated policies in four environments — 4,120,200 executed program-probe pairs — a 32-probe hybrid trained on generations 1–3 covers 576 of 582 (99.0%) dynamically killable faults in generations 4–5 while using 1.2–2.0% of the exhaustive domain, and withholding an entire fault family raises diversity's macro coverage from 84.6% to 94.9%. A downstream admission rule built on it removed severe tail regressions in all 20 program-environment groups; the authors are explicit that this prioritizes evidence rather than proving correctness.

Self-Augmented Diffusion Guidance for Physics-Informed Generation

Akira Osaka, Naoya Takeishi, Takehisa Yairi Diffusion models trained on simulations of physical systems can produce spatiotemporal fields that look convincing yet violate the governing equations, because nothing in standard training enforces physical constraints. The proposed approach learns the data distribution conditioned on a scalar measuring how far a sample deviates from correct dynamics, then generates by setting that deviation condition to zero, augmenting training with self-generated samples labeled by their measured deviation. Because the equations are evaluated outside training and sampling rather than at every denoising step, the method applies to systems whose governing equations are expensive to solve and generates samples faster than iterative physics-constrained sampling. Experiments show substantially smaller physical deviations than standard diffusion, with further reductions when stacked on existing physics-constrained methods.

SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting

Haizhao Fan, Xinyi Le The National Oceanic and Atmospheric Administration (NOAA) runs separate prediction systems for short- and medium-range forecasts, and the authors argue a single system covering both would give the public a more useful picture of global weather and its impacts. Nested-EAGLE (Experimental Artificial intelligence Global and Limited-area Ensemble) is a 0.25° global machine learning weather model with a 6 km refinement over the contiguous United States (CONUS), trained with high-resolution regional analysis data folded in through the nesting process. Over CONUS it achieves significantly lower mean-squared error for near-surface and low-level fields than NOAA's Global Forecast System and High-Resolution Rapid Refresh (HRRR) while staying competitive elsewhere in the global atmosphere; its precipitation amounts are less skillful than HRRR's because of deterministic training, though it gives the most accurate storm locations at longer lead times despite blurred extrema.

Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency

Nikol Figalov\'a, Lynn Huestegge, Anne B\"ockler-Raettig Large language models (LLMs) are increasingly used for title-and-abstract screening in evidence synthesis, where false negatives silently drop relevant studies. In a preregistered study embedded in a conceptually complex scoping review, 1,131 records were screened by a review lead, four trained assistants, and seven complete LLM runs across models and processing configurations, including a nominally identical repeat run, with operational recall measured against 316 verified eligible records. No workflow recovered every eligible record: human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records at 82.3-82.9% recall, Gemini 3.1 file batches reached the highest recall (83.9%) but retained 56.7%, all-at-once configurations underperformed file batches, and two nominally identical GPT-5.4 runs disagreed on 94 records, including 29 verified eligible ones retained by only one run. The authors conclude that performance depends on the implemented workflow rather than model identity alone, and that high-recall tasks call for auditable, human-supervised LLM use rather than autonomous exclusion.

LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration

Matthew Youngman, Cristian Sestito, Themis Prodromakis cross-listed Most work on Large Language Models (LLMs) for electronic design automation (EDA) targets isolated design stages, which obscures how capability actually accumulates and scales. This perspective organizes published systems into three hierarchical roles: a Generator producing artifacts in one pass, an Agent refining outputs through iterative tool feedback, and an Orchestrator coordinating decisions across EDA stages. Viewing systems through this lens exposes a "syntax trap" in which models learn to produce plausible code rather than physically correct hardware, compounded by fragmented tools and lost design context, and the authors argue that scaling to industrial designs requires a standardized, physics-aware orchestrator connecting tools and agents across the flow.

Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling

Maksim Utushkin, Andrei Ovsiannikov, Alexander D'yakonov cross-listed Friend recommendation depends on multi-hop social context, but running message-passing graph neural networks (GNNs) on a production graph with hundreds of millions of users and tens of billions of edges raises hard modeling and systems problems. The system uses multi-hash ID embeddings as the primary node representation, cutting an ID-embedding table that would exceed 200 GB by more than 98 percent without hurting ranking quality, and implements temporal neighbor sampling over timestamp-sorted compressed sparse row (CSR) storage with binary search, reducing per-node cost from O(deg(v) + k) to O(log(deg(v)) + k). On a graph of 194M users and 28B edges, offline ablations isolate each design choice, and an online A/B test lifts friend additions from recommendations by 16 percent and unique friend adders by 11.5 percent over a strong production baseline. The distributed training and inference framework for large temporal graphs is released.
81 more specialized papers

Large Language Models 72

Exploring the Role of LLMs in HPC Programming: A Survey

Strahinja Ljaljevic, Josep Jorba, Sergio Iserte cross-listed This survey reviews how large language models are being applied to High-Performance Computing (HPC) programming across five categories: code generation, parallelization and optimization, frameworks and architectures, evaluation and benchmarking, and broader challenges. It finds that general-purpose models handle serial and OpenMP-style tasks reasonably well but fall short on distributed paradigms such as MPI, where correctness and scalability are critical, while domain-specialized models like HPC-Coder, HPC-GPT, and chatHPC gain accuracy through fine-tuning, curated datasets, and retrieval-augmented generation but remain narrow in scope and are evaluated mostly on micro-kernels. The authors argue LLMs will not replace HPC experts in the near term but can become collaborators in the software pipeline, provided richer datasets, integration with performance analysis tools and schedulers, rigorous evaluation frameworks, and governance structures are developed.

TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

Jiaming Fan, Daming Cao, Canchen Huang, Jiale Fu, Jin Zhang, Junjie Gao et al. Tree-based speculative decoding uses a single drafter to build the candidate tree, forcing a trade-off between a fast but weak drafter and a strong but slow one. TreeGraft lets drafters of different cost jointly build one shared draft tree: the stronger drafter rescores the weaker drafter's candidates, reselects grafting positions, recovers promising unexplored paths, and adds its expansions non-destructively so existing branches remain eligible for acceptance by the target model, while a lightweight scheduler distilled from an offline value system decides when to invoke it. Across 10 model pairs and 6 benchmarks it beats the better of the two fixed single-drafter strategies by 15.1 percent on average, with a maximum gain of 26.6 percent.

ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

Xinming Wang, Haoran Du, Yi Chen, Jian Xu, Hongming Yang, Han Hu et al. Long-form factuality evaluation typically decomposes text into atomic claims, retrieves evidence, and verifies each claim, but decomposition noise and fixed verification granularity make results unreliable. ElementCheck instead extracts entity pairs explicitly linked by verifiable connections within each sentence, organizes them into an element graph, and uses the graph topology to estimate sentence complexity, verifying simple sentences directly and applying targeted element-level refinement and verification to complex ones. The authors build FastFact-Sent by mapping isolated claims from FastFact-Bench back to their source sentences, and show consistent factuality verification improvements across five backbone models on it and two domain-specific benchmarks, with a favorable accuracy-cost trade-off and less unnecessary re-verification.

Recipes for Steering and Scaling LLMs via Sampling

Jiajun He, Zongyu Guo, Jos\'e Miguel Hern\'andez-Lobato, Yuanqi Du Sampling from target distributions richer than an autoregressive language model's base distribution, such as powered, product, or tilted variants, has remained highly inefficient. The authors present a flexible, theoretically grounded framework for steering and scaling generation via sampling, describing two algorithms, one based on Sequential Monte Carlo (SMC) and one on Replica Exchange (RE), and demonstrate it by scaling generation quality without external supervision or reward models. Both methods scale more favorably than Best-of-N and standard MCMC baselines, and the paper frames the result as a systematic recipe for probabilistic inference with LLMs.

Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

Ali Asaria, Tony Salomone, Deep Gandhi Teaching a language model to abstain rather than guess usually requires a labeled dataset of correct and incorrect answers. This work instead fine-tunes models with LoRA to answer when their frozen confidence in their own answer is high and to say "I'm not sure" when it is low, using the confidence signal alone and no correctness labels. Across six open-weight models from 1B to 8B parameters in two families on short-form factual question answering, with correctness adjudicated by an independent judge model, the label-free recipe shows no statistically detectable difference from label-supervised abstention tuning at matched coverage; a control that drills hard examples instead of abstaining does not help, pointing to calibration rather than memorization, and the one blind spot is confidently wrong facts the signal cannot flag.

Evaluating Language Models in Realistic Conversational Contexts

Ilija Subasic, Andrew Rabinovich, Zhao Chen Evaluation frameworks designed for summarization, translation, or short question answering do not capture consistency across long, open-ended multi-turn dialogue, and the metrics themselves are often validated on synthetic rather than human data. UPHELD (UPwork Human-Scale Evaluated Long Dialogues) is a reference-full benchmark of hundreds of complete human-to-human conversations written by professional script writers, with realistic turn densities and over 36,000 per-turn human annotations across more than 30,000 dialogue turns. Classical automatic metrics and reference-free LLM-as-a-judge approaches correlate poorly with expert judgment on this data, while a Mixture-of-Judges framework that combines multiple evaluative signals improves correlation with human assessments by roughly 30%.

Affix Cache for Diffusion Large Language Models

Kaihua Liang, An Zhong, Xin Tan, Zafar Ayyub Qazi, Hong Xu, Jian Weng et al. Diffusion large language models (DLLMs) decode non-autoregressively with bidirectional attention, which couples the key-value (KV) states of shared context tokens to the evolving generated tokens, so naive cache reuse goes stale while full recomputation is expensive. ACache extends cache reuse to shared affixes beyond simple prefixes by identifying a small request-specific set of Anchor Tokens, chosen by their influence on the masked generation tokens, and recomputing KV states only for those while reusing the rest. Built on Fast-dLLM, it recovers the accuracy lost to direct affix-cache reuse while recomputing only about 20% of affix tokens, and a shared-prefix prototype on Nano-vLLM cuts recompute latency by up to 55.7% and raises end-to-end throughput by up to 1.68x.

GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions

Aravind Sasidharan Pillai Natural-language analytics over enterprise data warehouses is held back by hallucinated metrics, invalid joins, wrong grain, unsafe data access, and unsupported explanations, and schema- or documentation-grounded text-to-SQL does not capture governed business semantics such as approved metrics, join paths, filters, and row-level security. GROUND (Governed Retrieval Over Unified Normalized Definitions) constrains generation to a governed semantic layer: it supplies approved definitions, binds user intent to governed metrics and dimensions, validates generated SQL against schema, metric, join, grain, filter, security, and cost rules before execution, and retries or abstains on violations. On a 100-question synthetic enterprise-reporting benchmark against schema-only, schema-RAG, and semantic-only baselines under one shared model, GROUND was the only system with zero measured hallucinations across all six categories, and a semantic-only condition with exact metric definitions still leaked data through row-level security violations. Replication on real U.S. NHTSA vehicle-safety data and an adversarial set across four models from three providers showed the enforced filter and security guarantees holding with zero violations, while judgment-dependent behaviors such as refusing undefined metrics remained fallible.

Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention

Naveen Lamba, Sanju Tiwari, Manas Gaur Prior hallucination surveys have concentrated on detection or mitigation in isolation. The survey instead organizes hallucinations in large language models (LLMs) around the model development lifecycle, proposing a three-fold categorization into data-related, training-related, and inference-related hallucinations, and for each stage reviews causes, detection methods, and the mitigation or prevention interventions that apply. It also assesses available benchmark datasets along several parameters to judge their suitability for identifying and managing hallucinations, and offers a standardized lifecycle framework intended to help researchers and practitioners diagnose and address hallucinations in high-stakes domains such as health, legal, and scientific work.

Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints

Hiroko Takano cross-listed Multi-stage large language model (LLM) hiring pipelines that improve resumes, generate interview questions, and give answer feedback can invent credentials, inflate qualifiers, and fabricate experience. A controlled experiment with 10 synthetic resumes, 2 job descriptions, 3 repetitions, and 3 conditions (180 runs) compares a fully automated baseline with prompt guardrails and with a human-in-the-loop (HITL) checkpoint placed after the resume-improvement stage. The baseline produced at least one unsupported claim in 96.7% of outputs; guardrails cut finding density by 86% but left fabrications in 50% of outputs, while the human checkpoint eliminated all identity fabrications and reduced capture of job-description trap requirements from 47% to 2%. Reviewers caught flagrant fabrications but missed subtle qualifier drops and plausible new claims about half the time, contamination rose with domain distance for multi-specialty resumes, neither mitigation reduced claim retention below 99%, and a newer-generation model still fabricated in 90% of baseline runs, supporting a layered design combining both interventions.

Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors

Mantas Lukauskas Extractive prompt compressors such as LLMLingua-2 promise cheaper LLM inference by dropping low-information tokens, but they are trained and benchmarked almost entirely on English, while most other languages already pay a 1.3-1.8x token premium for the same content. Using parallel data in ten languages across five scripts, budget-matched controls, four learned compressors, four deterministic baselines, and eleven target models from ten vendors (over 250,000 evaluation calls), the authors find the cross-lingual transfer gap is real and strongly rate-dependent: at a 0.33 keep-rate English retains 57-62% of normalized context utilization while Lithuanian retains 10-24% and Chinese essentially none. The gap follows the compressor's supervision data rather than its architecture, since all three English-trained compressors show it while deterministic methods and the multilingually trained XProvence v1 do not, and a translate-then-compress pipeline matches or beats native compression at roughly half the token cost in three of five languages tested.

Comparing Chunking and Embedding Strategies for Turkish RAG Systems

Mustafa Serta\c{c} T\"urkel, Fatma Nur Korkmaz, Ahmet Tu\u{g}rul Bayrak Chunking and embedding choices strongly shape Retrieval-Augmented Generation (RAG) quality, yet neither has been studied systematically for morphologically rich languages like Turkish. The authors run a fully crossed comparison of three chunking strategies (fixed-length, semantic, and layout-aware Docling), five embedding models, and two generator LLMs over three documents with contrasting layouts, producing 9,000 judge-graded question-answer evaluations tested with paired McNemar tests under Holm correction. Layout-aware chunking compresses the spread between modern embedding models to about one point, and the three leading embedders are statistically indistinguishable, so language specialization yields no measurable retrieval advantage; the faster generator is not the more accurate one, layout-aware chunking helps table-heavy documents far more than prose, and the best individual components do not compose into the best full configuration, which reaches 87.0%.

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

Sachin Gopal Wani, Ajay Dholakia, David Ellison Accuracy-only benchmarks skip the deployment question of when a reasoning model's extended thinking tokens are worth their cost. The Token Economy Score (TES) measures a reasoning model's accuracy gain over a non-reasoning baseline normalized by the generated-token multiplier, with paired and approximated variants for families with reasoning toggles and for frontier models lacking a direct counterpart, and is applied across 151 model-benchmark runs on seven benchmarks spanning mathematics, code generation, science, instruction following, expert knowledge, knowledge recall, and research-level physics. Task structure predicts reasoning efficiency better than nominal difficulty: sequential inference-chain tasks like AIME 2025 and LiveCodeBench score high TES while knowledge-recall tasks like MMLU-Pro score low despite being hard; returns diminish at higher effort levels and sometimes turn negative, a Reasoning Cost Share metric shows inference spend is often dominated by internal thinking, and a Deployment Cost Multiplier shows on-premises deployment can change the economics, supporting a rule of enabling reasoning selectively by task type, effort level, and deployment context.

Assessing mentalization in humans and large language models

Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang Mentalization is the ability to infer others' beliefs and intentions and use those inferences to guide one's own choices; whether large language models (LLMs) can do this adaptively, rather than merely pass theory-of-mind tests, was previously unknown. The authors run 2,099 LLM agents from DeepSeek, GPT-4.1, GPT-5, and Gemini 2.0 Flash through two economic games against opponents of varying sophistication, fit cognitive computational models to recover the latent strategies, and compare against 251 human participants. LLMs show clear behavioural and computational signatures of mentalizing that vary substantially by provider and model size, and a prompt designed to elicit strategic reasoning generally improves performance, though unevenly across the two games. GPT-5 agents flexibly deepened their recursive reasoning as opponents became more sophisticated and outperformed the human participants.

On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study

Aditya Pratap Singh Memory-based knowledge editors in the SERAC lineage all rely on a scope decision: given an incoming query, does a stored edit apply or not. To measure this decision exactly, the authors build INLAY, a gradient-free editor that keeps the model frozen, stores edits in an external addressable memory, and applies an edit as a bias along one token's unembedding direction at decode time, then execute every candidate routing action on 1,689 queries across three datasets and three input conditions. An oracle router that picks the best action on every query ties a one-line static policy to four decimal places in all nine dataset-by-condition cells, so the maximum attainable gain from any per-query routing method on these benchmarks is zero. The cause is structural: these counterfactual benchmarks ask for the post-edit answer and contain no negatives, so a classifier's ability to reject is never rewarded; withholding a query's own edit for half the sample restores measurable headroom (+0.042), and the authors also report where WISE and retrieval-augmented generation beat INLAY on CounterFact and RippleEdits.

When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

Yefan Tao, Gerald Friedland, Luyang Kong Text models degrade when inputs contain typos, OCR errors, or dropped words, but how consistently they degrade across architectures had not been characterized. Comparing sentence-embedding models and decoder-only LLMs, the authors find that under word-level noise very different architectures decline along nearly the same curve, whereas under character-level noise they diverge, and that the training objective rather than the architecture is the determining factor: eight encoders spanning six pretraining paradigms start scattered but collapse onto a common degradation curve after a short contrastive training recipe. The word/character split traces to tokenization, since a single character edit forces the tokenizer to re-segment the surrounding word and disturbs the token sequence far more than dropping a whole word does. This mechanism offers a way to predict a model's noise robustness without running any noisy evaluation and to install robustness at a chosen noise scale through noise-augmented training.

How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic Large language models increasingly generate and interpret phrases like 'unlikely' or 'probable', but whether these carry consistent numeric meaning across models or match human perception had not been tested systematically. The authors give 19 models a word-to-number mapping task over eleven uncertainty expressions grounded in established human benchmarks, under both forced single-number and explanation-elicitation conditions, plus a bidirectional roundtrip test of internal consistency. Models track the human benchmark closely, preserving word ordering, recovering three anchor points, and reproducing the high variance of 'possible', but show a systematic upward bias for negative expressions such as 'unlikely' and 'improbable'. Explanation elicitation lowers within-model variance while raising between-model divergence, and the roundtrip test stratifies models clearly, with frontier models maintaining coherent bidirectional representations.

Survival-Guided Length Control for Efficient Diffusion Language Models

Ivan Kobyzev, Abbas Ghaddar, Yufei Cui Diffusion language models (DLMs) generate text by iteratively denoising masked sequences, but standard decoding either fixes the output length in advance or uses ad hoc stopping rules, wasting denoising steps. The authors recast length selection as a discrete-time survival problem over the end-of-sequence token and build a plug-in, training-free length predictor that works with any existing DLM. On reasoning and code-generation benchmarks, survival-guided length decoding speeds up inference by up to 7 times while preserving task accuracy, and the analysis shows predicted lengths vary widely even within a single dataset, making performance sensitive to the chosen length.

Assessing the Downstream Utility of Evidence-Aware Retrieval in RAG

Utshab Kumar Ghosh, Debayan Mukhopadhyay, Shubham Chatterjee cross-listed Retrieval evaluation for retrieval-augmented generation (RAG) is increasingly built around whether retrieved passages contain evidence that can support an answer rather than topical relevance alone, and this study asks whether that closer alignment actually makes the evaluation more useful for the decisions built on it. Across five retrieval benchmarks and an end-to-end TREC RAG 2025 setting, an answer-support signal is examined in four roles: comparing retrievers, guiding retriever training and system selection, predicting downstream answer quality, and filtering the evidence supplied to a generator. The signal changes retrieval rankings, but its downstream value is not uniform: it does not reliably improve retriever training, its benefit for system selection depends on how the generator is instructed to use evidence, and its scores do not robustly predict answer quality on unseen topics. Human annotators confirm that evidence filtering preserves passages with useful answer evidence, yet different answer evaluators disagree on whether the resulting answers improve, so the authors argue RAG evaluation methods should be assessed against the specific comparisons and decisions they are meant to support.

Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries

Alden Do Rosario, Hussein Younes, Felipe Pires Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing, since a system that answers everything beats one that abstains when its knowledge base cannot support an answer. Building on the confidence-target analysis of Kalai et al., the authors propose a penalty-aware evaluation for deployed RAG products that combines asymmetric scoring (correct +1, wrong -4, abstain 0), knowledge-gap canaries whose answers are verifiably absent from the knowledge base so that any answer is ungrounded parametric generation, and a failure-attribution pipeline separating retrieval, generation, and abstention-policy failures. Applied to three commercial RAG systems and a no-retrieval baseline on SimpleQA-Verified with a blind cross-family three-judge panel, accuracy when answering clusters tightly at 97.0 to 98.0%, while canary violation rates differ roughly sixfold (16.7% versus 98.1%), so penalty-aware scoring reorders the volume-based ranking stably across penalty settings from k=1 to k=9. All code, transcripts, and judge votes are released for independent audit.

Co-Evolving Structured Knowledge and Reasoning in Language Models

Ryan Thomas Noonan, Linxi Zhao, Menghan Xu, Akanksha Sarkar, Mihir Mishra, Dongyoung Go et al. Retrieval over unstructured text grounds language models in external knowledge but often pulls in irrelevant context, while structured knowledge bases are more controllable yet expensive to build and brittle to reason over. KBevo is a co-evolving framework that jointly learns to construct a structured knowledge base and reason over it for knowledge-intensive question answering, optimizing both components end-to-end with QA outcome rewards so that reasoning success feeds back into knowledge-base quality. The approach yields larger, better-connected knowledge structures with higher answer reachability, along with improved compositional factual reasoning and controllability relative to standard retrieval baselines.

LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

Zishan Shao, Lixun Zhang, Kangning Cui, Wenhao Wu, Jinhee Kim, Yixiao Wang et al. Singular value decomposition (SVD)-based low-rank compression is a growing way to cut the memory and compute cost of large language models (LLMs), but prior studies use different benchmarks, compression ratios, and setups, making it unclear whether reported gains come from the method or the evaluation protocol. LowRankArena standardizes task versions, uniform-precision compression budgets, comparison regimes, and inference measurements, with a reproducible pipeline and over 3 TiB of released compressed checkpoints. An aligned audit of five representative SVD methods finds that prior conclusions are highly conditional, with leaders and performance tiers shifting across backbones and keep ratios, that multiple-choice accuracy can mask large perplexity degradation, and that nominal low-rank savings translate into workload-dependent and often limited end-to-end speedups.

Fine-Tuning of Transformer models with Frames

Harshavardhan Adepu, Li Zhang, Sanjiv Kumar, Vikas Singh Parameter-efficient fine-tuning (PEFT) methods such as Low-Rank Adaptation (LoRA) still have memory requirements that scale with model size, on the order of the hidden dimension times the rank. FrameFT instead models the weight update as a sparse coefficient matrix in a Fusion Frame basis that can be generated algorithmically and shared across layers, so only the sparse coefficients are stored and optimized, and the sparsity in both the coefficients and the frames yields compute savings backed by formal convergence results. On a suite of supervised fine-tuning benchmarks, mainly language tasks with some vision experiments, FrameFT matches or exceeds state-of-the-art PEFT techniques with far fewer trainable parameters.

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun Byte-level Byte Pair Encoding (BPE) tokenizers built on the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, which defines a word as a run of Unicode letters; in abugida scripts such as those used for Nepali or Thai, vowels are combining marks rather than letters, so every word is split at each vowel sign before BPE ever runs, and no amount of vocabulary or training data can undo the split. The authors formalise this as a training-free lower bound on fertility (tokens per word) and measure it across 26 languages: all 17 abugidas are affected, from 1.47x for Tibetan to 9.02x for Thai, while Latin, Cyrillic, Hangul, and Han sit at exactly 1.00x. Matched tokenizer pairs differing only in the character class land within 2.2% of the predicted floor (4.78 versus 1.58 tokens per word on Nepali), and a 268M-parameter model trained with the mark-aware tokenizer reaches 4.43% lower held-out Nepali bits per byte at equal compute. A census of 3,479 HuggingFace repositories finds the letters-only word class in 63.3% of the most-downloaded text-generation models, even though GPT-4o's o200k pattern already uses the mark-aware fix.

Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue

Ming Cheng, Yusheng Dai, Qiuhong Ke, Zhaolin Chen, Lizhen Qu Conformal Risk Control (CRC) bounds the error of LLM-as-a-Judge pairwise evaluations by abstaining below a decision threshold; the observation here is that aggregating multiple judges cleans up the scoring function at its source and so complements CRC. Two multi-expert CRC methods, Score Averaging and Decision Voting, beat single-judge baselines on homogeneous panels but recover limited coverage on heterogeneous panels because one threshold cannot fit judges with different scoring scales. Marginal-Calibrated Conformal Consensus (MC3) captures per-expert scales through initial threshold ratios while jointly tuning a unified decision function applied identically in calibration and test, preserving exchangeability. On Panel, a new 1,800-pair human preference benchmark built from four open-weight LLMs over ESConv, MSC, and DREAM dialogue contexts with full logit access, MC3 extends the accuracy and acceptance-rate gains of aggregation to heterogeneous judge panels across all three datasets.

Dependency-Aware Revocable Decoding for Efficient Diffusion Large Language Model Inference

Wooje Park, Insu Lee, Minyoung Noh, Jaeyun Jang, Sungmin Lee, Kyuhong Shim et al. Diffusion large language models (dLLMs) generate by iteratively denoising many tokens at once, but pushing parallelism higher hurts quality because an early mistake poisons the context used for later tokens. Revocable decoding fights this by re-checking already-decoded tokens and remasking unreliable ones; the authors point out that the unreliable tokens also corrupt the very context used to do the checking. DARD (Dependency-Aware Revocable Decoding) is training-free and splits tokens into masked, candidate, and unmasked states, verifying candidates against a selective context that excludes less reliable tokens and adaptively limiting how much those tokens influence later steps. Across 12 text and multimodal benchmarks on 3 open dLLMs it improves the speed-quality Pareto frontier, including a 2.71× speedup and a 4.35-point CIDEr gain over Saber on Flickr30K.

Double Trouble: Bilingual Pretraining Leaves Language-Conditioned Effects in Shared-Language Representations

Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos Comparative studies of multilingual models routinely align two models' embedding spaces and then treat their representations of a shared language as interchangeable for probing, interpretability, or transfer analysis. Testing that assumption directly, the authors pretrain matched 310M-parameter decoder-only pairs — one English-only, one bilingual — across eight typologically diverse languages, controlling separately for English exposure, total compute, and document overlap. After aligning on shared English vocabulary and testing held-out words, token embeddings look comparable post-alignment but the deeper hidden states the model actually predicts from do not, a gap that holds for all eight languages and survives overlap controls and alternative alignment methods. The mismatch grows through the middle transformer layers, pointing to contextual processing rather than the input representations where alignment is applied.

Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs

Rongfeng Wang, Shichao Weng, Zhiqiang Wang, Xinyu Liu, Yang Yi, Peilong Zhou et al. Mixture-of-Experts (MoE) models typically activate a fixed number of experts in every layer for every task, even though layer roles, expert redundancy, and task difficulty all vary. MetaNet is a support-set controller that predicts, for each layer, an expert-retention threshold and a bounded routing bias while the backbone, experts, and router stay frozen, combining layer-wise allocation with task-level context that offline allocations and token-level routing signals each lack. On DeepSeek-MoE-16B-Chat, a conservative setting activates 3.61 experts on average (40% fewer than fixed k=6) at comparable MMLU accuracy (0.489 vs. 0.474), an aggressive setting uses 2.28 experts (62% fewer) at about 3.7 percentage points lower accuracy, and the MMLU-trained controller transfers to C-Eval without retraining, activating 2.90 experts on average at 0.386 accuracy.

FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models

Junyoung Lee, Sehyeon Park, Shinhyoung Jang, Seonha Ryu, Hojeong Kim, Hyunsei Lee et al. Pruning compresses large language models (LLMs) but can amplify text degeneration, particularly repetition loops, even when perplexity and task accuracy barely change. Treating decoding as a dynamical process that enters and persists in a small set of recurrent contexts, the authors decompose degeneration into loop entry risk and loop persistence and show that persistence is controlled by the escape mass assigned to plausible alternatives within the token sampling set. Two token-level objectives for post-pruning fine-tuning follow from this analysis: FOCUS reweights distillation toward high-confidence teacher regions to suppress leakage, and RePAIR uses onset-centered positive/negative continuation pairs with a margin loss to promote plausible alternatives and prevent early commitment to loops; both consistently reduce repetition and improve generation quality on open-ended continuation and instruction-based generation.

Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing

Hyeonchu Park, Gahye Jeong, Bugeun Kim High false-positive rates from AI text detectors on non-native English writing have been reported before, but population comparisons confound authorship with topic, domain, and style. Professional editing offers a cleaner test: it changes linguistic form while leaving authorship and content intact, so the authors ran 13 detectors over 135,389 manuscript pairs from a 2018–2025 English editing service, comparing each non-native original with its native-edited version. False-positive rates on human-written text ranged from 0.0% to 100.0% across detectors, and the same edits pushed AI scores up for some tools and down for others, with score changes correlating with how heavily a manuscript was edited. The conclusion is that polished academic style is a confound in detector output rather than a signal of machine authorship, with direct fairness implications for academic use.

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

Tingyun Li, Wenfeng Feng, Weiqing Li, Abudukelimu Wuerkaixi, Guohua Liu, Yuewei Zhang Systems that automatically post-train language models repeatedly — proposing an update, training a candidate, reading evaluation feedback, proposing the next one — accumulate a record of what worked, but an update's effect depends on the parent model, data, and training stage, so old successes can stop being valid once the parent has changed. The authors formalize this as conditional experience transfer and propose Boundary-Calibrated Intervention Transfer (BCIT), which binds each observed effect to its source context, checks applicability conditions before any weight-changing run, vetoes candidates with named hard conflicts, and buys fresh evidence through a bounded training trial when the record is inconclusive. Testing on a 4B model adapted to finance reasoning, text-to-SQL, and function calling, where candidate updates showed heterogeneous target and retention effects, BCIT authorized fewer harmful updates and reached higher final-model quality than alternatives at equal compute. The framing treats deciding whether to reuse experience as a problem separate from generating or training candidates.

Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD

Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo Adapting a language model to a vertical domain reliably erodes general skills such as reasoning, coding, instruction following, and creative writing. Working within Multi-Teacher On-Policy Distillation, where a specializing student is supervised on its own sampled trajectories by both domain and general teachers, the authors identify two weaknesses: ordinary sampling rarely surfaces tokens where teachers strongly outperform the student, and the sign of that advantage alone does not tell you whether the update direction is trustworthy. Their fix combines dual-temperature sampling to widen the trajectory pool, positive-advantage-density filtering to pick trajectories carrying stronger learning signal, and centered log-likelihood filtering, which computes an entropy-calibrated teacher-endorsement score and keeps token updates probabilistically based on whether direction and endorsement agree. On role-playing and medical specialization the method raises the general-capability average by 4.73% and 10.84% over standard multi-teacher distillation without losing domain performance, and ablations indicate the gains are not just a larger rollout budget.

Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers

Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg, Navid Rekabsaz Rerankers, reward models, and multi-document question-answering scorers assign scores to several candidates in a single prompt, so each score depends on candidate order, yet these scorers are typically selected on ranking quality rather than on the downstream decisions their scores drive. On passage reranking, five trained scorers within 0.010 nDCG@10 of each other retain sets that overlap by only 0.66–0.84 when candidates are reordered, and no prompt-time intervention tested removes the dependence. Order-consistency supervised fine-tuning (OC-SFT) trains a candidate's score to be independent of its position, preserving ranking quality while leading every decision-stability measure across three tasks, flipping a reader's answer on 0.125 of permutation pairs versus 0.149–0.164 for other order-targeting objectives, and proving more stable than order-averaged distillation on 12 base models. The authors argue that scorer comparisons should report what a threshold retains and what a reader answers, not ranking quality alone.

Instruction Quality Matters: Refining Instructions for Effective Preference Learning

Seohyeong Lee, Hwaran Lee, Buru Chang Preference learning trains models on chosen-versus-rejected response pairs, but the informativeness of those pairs depends on the instructions that generated them, and low-quality or ambiguous instructions narrow the range of response quality that can be sampled. Using Best-of-N and Worst-of-N analyses, the authors show that instruction quality bounds both the ceiling and the floor of sampled response quality, and they propose a pipeline that flags weak instructions with reward signals and rewrites them using rubric-guided LLM feedback rather than discarding them. Across offline and online preference learning on multiple models and benchmarks, refined instructions yield broad alignment improvements over the original data and over alternative data-improvement strategies, and the gains complement response-centric data curation.

Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry

Alexandru-Iulius Jerpelea LLMs are hypothesized to track belief states, running posterior distributions over the latent variables that govern text, but evidence so far comes from toy synthetic data and isolated case studies, and the idea has not been connected to the geometry of features found by interpretability methods. The authors plant a controllable latent variable in natural-looking text by having an LLM teacher write ordinary prose while being subliminally steered along one of K = 8 unrelated sparse autoencoder directions per token, with the active direction following a ring-shaped Markov chain. A small transformer trained on this corpus tracks the Bayesian posterior over the planted variable, and it arranges the 8 states on a ring in the exact order of the Markov chain, supporting the view that a concept's geometry can be shaped by the statistical dynamics of the latent variable behind it.

A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models

Artem Safronov Small language models (sLLMs) deployed on memory-constrained devices are memory-bandwidth bound during autoregressive inference, and uniform quantization hurts them because few layers tolerate reduced precision. The proposed composite metric combines an information-retention score based on normalized signal-to-quantization-noise ratio (SQNR) with a throughput score from roofline-based latency modeling, requires no actual execution, and can prioritize individual blocks, projection sublayers, or whole transformer layers. Profiling Gemma 3 1B identifies feed-forward network blocks and the embedding matrix as the most promising acceleration targets, and across several architectures the speedup estimates show around 4% prediction error. Compared with evolutionary search, specialized accelerators, or Shapley-value approaches that need expensive approximate inference, the method tends to allocate more precision to the most expressive layers.

Squeezing More from Limited Data with Recursive Transformers

Serdar G\"ulbahar, Lukas Edman, Alexander Fraser When the data budget is fixed but compute is plentiful, adding parameters to a language model helps only up to an optimal size, beyond which overfitting hurts generalization. Studying pre-training budgets of 10M to 100M words across two corpora and several downstream evaluations, the authors find the optimal size depends strongly on both the data budget and the target task, and argue that standard Transformers scale down poorly because embeddings eat a large share of the parameter budget and per-token compute is tied to representational capacity. To decouple these, they study recursive Transformers that reuse a shared block across depth together with factorized embeddings, and the three recursive models trained outperform standard Transformers at both 10M and 100M words while remaining competitive with BabyLM Challenge 2025 winners.

Disentangling Optimization Scale from Preference Scale in DPO

Ivan Kruzhilov In Direct Preference Optimization (DPO) the coefficient beta is usually read as the strength of the KL constraint to the reference policy, but the authors show it actually plays two roles at once: it sets the effective inverse preference-noise scale and it rescales the optimization dynamics, coupling that scale to the effective step size. One consequence is that at a fixed learning rate the achieved policy deviation is non-monotone in beta, vanishing in a dead zone at small values, peaking at an intermediate value, and shrinking again for larger ones; standard DPO loss values are also not comparable across beta, since runs with near-identical loss curves can differ several-fold in KL divergence from the reference. They propose a centered-softplus reformulation that is argmin-equivalent to DPO for beta greater than zero but exposes the noise-scale and learning-rate effects as independently tunable quantities, and whose normalized form has a continuous beta-to-zero limit that reduces to a linear preference-margin objective.

Cascaded Batch Prompting

Sho Hoshino, Peinan Zhang Batch prompting, which packs multiple instances into one large language model call, improves inference efficiency but makes downstream task performance unpredictable. Cascaded batch prompting is a two-stage approach that separates the complex reasoning step from symbol grounding, aiming to recover the reliability lost in conventional batching. On multiple-choice question answering and natural language inference, the method outperforms the standard single-prompt baseline while achieving a speedup proportional to batch size, which the authors describe as a new state of the art on the accuracy-efficiency Pareto frontier.

Performance Foundations of Parallel & Distributed Reasoning Language Models

Maciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic et al. Reasoning Language Models (RLMs) such as DeepSeek-R1, o3, and Kimi k1.5 are produced by RL-style post-training, including Reinforcement Learning with Verifiable Rewards (RLVR), but training them costs millions of GPU-hours and involves tightly coupled multi-model pipelines that stress hardware well beyond supervised LLM training, making it as much a parallel and distributed systems problem as an algorithmic one. The authors systematize the RL-for-LLM paradigm with a compute-centric analysis of PPO, GRPO, and their variants, then build a taxonomy of intra- and inter-model parallelism strategies for RL-for-LLM training, covering data, tensor, pipeline, sequence, context, and expert parallelism along with newer techniques such as disaggregated placement, stage fusion, hybrid parallelism, and asynchronous execution, all grounded in the work-depth model of parallel computing. They close by analyzing existing RLM training frameworks against this taxonomy and distilling practical guidelines and open research directions for building scalable, fast, and cost-effective RLMs.

Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?

Ej Zhou, Suchir Salhan, Catherine Arnett, Anna Korhonen Cross-lingual alignment in multilingual models is usually credited to joint training through shared parameters, mixed-language batches, or explicit alignment objectives. Testing strictly monolingual models trained on non-parallel data, including the Goldfish families and independently developed models from different labs, the authors find that these models develop alignable representational geometry across layers, with alignment strengthening as data scale, model scale, or linguistic proximity increases, and that a single Procrustes rotation fit on parallel sentences maps hidden states between models. The rotation also transfers function: patching a rotated English residual into a German model on a factual cloze task flips the prediction to the donor's answer in most cases, indicating that alignment can emerge from the structure of language and the information it carries rather than from joint training, and pointing toward model stitching, merging, and modular multilingual systems built from monolingual components.

TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy

Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Junyan Zhang, Xuming Hu Long-context inference is limited by key-value (KV) cache memory, and existing eviction policies score tokens by attention magnitude or, in attention-free variants, by each key's distance from a global reference point. A controlled leave-one-out probe finds that attention magnitude is essentially unrelated to a token's causal contribution to the answer (Spearman rho of -0.004), undermining the premise of dominant eviction methods. TwinKV is a training-free, attention-free redundancy signal that detects whether a token's key has a near-duplicate elsewhere in context; applied as a composable repair pass on top of any policy's fixed retained set, it swaps evicted tokens that have no surviving duplicate (orphans) for retained tokens whose information is duplicated elsewhere (redundant donors), preserving the original budget and scoring rule. Composed with four recent policies on LongBench, LooGLE, RULER, and a short-context MMLU-Pro no-harm control at compression ratios 0.3 to 0.7, it improves most configurations for two policies on Qwen3-4B, is near-even for a third, and helps an adaptive fourth policy mainly on Llama-3.2-1B where headroom remains, while offering no benefit on few-shot classification exemplars for either model.

Thomson: Continual Learning of Frontier Models for SovereignAI

Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck et al. The premise is that institutions outside the handful of well-funded frontier labs can reach frontier-level performance by applying continual learning to open-weight models rather than relying on small-scale fine-tuning, prompting, or tool augmentation of a frozen model. The approach runs a full mid- and post-training stack with safeguards at each stage meant to preserve plasticity and stability while making a minimal number of high-impact parameter interventions. The resulting model, Thomson, is reported to perform competitively with recent frontier models on agentic tasks, safety, legal and tax work, multilingualism, and large-scale deep research, with a π-shaped evaluation pattern: broad gains across capabilities, including untargeted ones, with almost no forgetting, at compute and personnel budgets the authors argue are far below what is commonly assumed.

Compositional Online Learning for Semantic Data Processing Systems

Pawe\l{} Liskowski, Fuheng Zhao, Benjamin Han, Anupam Datta, Dimitris Tsirogiannis cross-listed In semantic data processing systems a single large language model (LLM) call costs 10^5 to 10^7 times a relational predicate and accounts for 80 to 90% of query cost in production, yet its latency also leaves room to run a CPU-side learner's update entirely inside the round-trip, inverting the classical constraint that adaptive query processing learners must stay lightweight. The authors propose compositional online learning at the LLM call boundary, a framework in which components differing in decision granularity and update cadence each make execution-time decisions and refine their learned artifacts online, with each training step hidden behind the next LLM call. A production case study in Cortex AISQL composes a memoization layer, an online per-call filter-ordering learner, and an online per-batch cascade-routing learner, each assigned to a distinct factor of per-row LLM cost by a conditional cost decomposition. Under independence the two learning components compose multiplicatively to an 11.4x upper bound on a representative conjunction-filter workload, which self-selection at the cascade boundary, sample-budget shrinkage, and selectivity-estimation drift reduce to a realistic figure near 8x.

LLMs Can Design Near-Optimal OR Algorithms

Jackie Baek The question is whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems, studied on inventory control, queueing network control, and assortment optimization. Two levels of use are evaluated: at level 1 the model receives one problem instance and returns a solution, and at level 2 it receives only the problem-class description and broad parameter ranges and must return an algorithm that maps instance parameters to solutions, given a single untuned prompt and a Python sandbox with a fixed compute budget. The strongest model tested, gpt-5.6-sol, matches or outperforms the best existing method on almost all evaluated instances, even at level 2 where the algorithm is fixed before seeing the evaluation instances, and performance improves sharply across models released less than eight months apart. The authors conclude that a single untuned LLM query can already produce algorithms competitive with specialized methods and that frontier LLMs are a serious empirical baseline for algorithm design in such problems.

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang Audits of large language model (LLM) judges often certify a bias with a double difference: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. The authors show this endpoint is not identified on the scale that reports it, because each term is censored by its own share of the scale bound, so a severity shift common to both responses manufactures an apparent interaction whenever the two responses sit at unequal distances from the bounds, which is exactly where good stimuli place them. In a pre-registered audit of a frozen pedagogy judge sealed before the first of its 990 calls, the registered primary endpoint, the effect of a stated learner profile on scaffolding preference, is null at +0.085 points (95% BCa CI from −0.167 to +0.353, p = 0.684), while the one nominally significant interaction, +0.378 (p = 0.002), is 79 to 85% reproduced by a construction containing zero differential preference using only the observed severity shift and the scale floor. The mechanism is derived in closed form and its contribution is shown to be measurable from an audit's own ratings.

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

Xinwei Qiang, Xiang Fang, Chang Chen, Yue Guan, Yufei Ding Block drafters for speculative decoding propose several tokens in one forward pass before earlier target tokens are known, and their rejection rate conflates two losses: information missing because within-block tokens are unrealised, and imperfect modelling of what is observable. The authors define an information floor, the minimum expected rejection at a given conditioning order, and treat rejection above it as the model gap, estimating both from target-model rollouts across four domains, four open-weight targets, and a frontier API target. On Qwen3-4B the all-parallel floor reaches 0.286 at the final slot, capping per-slot acceptance at 71% even for a perfect proposal, but conditioning on a single realised token removes 86–100% of that floor, a locality also recovered by an independent mutual-information analysis. Current drafters sit well above their floors, with the final-slot model gap accounting for 43–64% of DFlash rejection and 85–92% of DSpark's oracle-conditioned rejection.

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

Kairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen, Yiming Yang, Linxiang Gao et al. Pretraining a language model is usually prohibitively expensive, with even Llama-3.2-3B costing over $1.5M and SmolLM3-3B over $700K to reproduce, leaving the community without a cost-efficient, hardware-accessible, open pretraining recipe. The Puro-2B collection is trained from scratch on up to 1.4 trillion tokens in FP8 precision on consumer-grade RTX 5090 GPUs, combining hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and a data recipe. The best model costs under $6.9K to train and approaches Qwen2.5-1.5B performance under the authors' evaluation protocol, and a fitted Puro Cost Scaling Law suggests about $4.4K suffices to match Qwen2-1.5B. An end-to-end case study also examines how pretraining data curricula shape performance after post-training, and data, code, and weights are released under Apache 2.0.

Token-Level Advertising

Hanbing Liu, Bowei Zhang, Changyuan Yu, Yinyu Ye, Qi Qi cross-listed Generative AI is changing how people access information, undermining advertising mechanisms built around predefined slots. The Latent Advertiser Mixture Auction (LAMA) embeds advertiser influence directly into token generation: advertisers report local continuation values that induce advertiser-specific next-token policies, and the platform decodes through a latent mixture while updating an allocation posterior. The mechanism is shown to satisfy Markov dominant-strategy incentive compatibility and individual rationality and to achieve near-optimal KL-regularized welfare, and a learning-based implementation reconstructs the required reports online from learned local advantages and root values. Proof-of-concept experiments on real-world commercial-search query splits show improved platform welfare and revenue while maintaining user-facing response quality.

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang et al. Evaluating large language models on enterprise-scale document collections is hard because companies will not share internal communications and existing synthetic datasets are overly simple. CorporateBench (CB) is a human-validated, multi-task question-answering benchmark whose evaluation corpora exceed 230,000 documents, built from four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents, and the benchmark tests both information extraction and knowledge base querying. Evaluating five LLMs, performance degrades increasingly as input size approaches realistic corporate scales.

How Language Models Organize and Structure Moral Knowledge

Orion Reblitz-Richardson Open-weight language models are probed for whether they represent moral content as a single detector or as a structured set of distinct foundations. The authors train six independent linear probes, one per Moral Foundations Theory (MFT) category (care/harm, fairness/cheating, liberty/oppression, loyalty/betrayal, authority/subversion, sanctity/degradation), and analyze how the resulting directions relate geometrically in representation space. The probe directions span a near-maximal number of independent dimensions while sharing a positive common component, and this shared component is moral-specific: mean pairwise cosine 0.26 versus 0.013 for a matched non-moral concept battery. The geometry is consistent across architectures and scales, emerges early in pre-training, shows no evidence of MFT's individualizing/binding split, and extends to moral dilemmas, whose directions partially compose from their component foundations while mostly encoding conflict-specific structure.

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu et al. cross-listed Real-world code review is an iterative exchange between developers and reviewers, yet most LLM-based code review work reduces it to a single-round, static decision. MCR-Bench is a defect state-aware benchmark of 2,269 real multi-round code review tasks across five programming languages, each annotated with fine-grained defect metadata (description, type, severity) and cross-round state labels that trace a defect's full lifecycle through the review. Experiments with mainstream LLMs show limited overall capability in both defect detection and lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases, and substantially worse results on semantically complex or low-salience defects. Error analysis attributes false positives and false negatives to distinct drivers, notably cross-round temporal misalignment and inadequate long-range memory.
20 more specialized papers

Agents 51

CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering

Kunjesh Parekh, Anil Kumar Tiwari, Divya Saxena Multi-step financial calculations are a known weak spot for large language models, which produce plausible but numerically wrong answers. CIFQA (Calculation-Intensive Financial Query Answering) separates language understanding from numerical execution by assigning specialized agents to query interpretation, routing, parameter extraction, computation planning, and response generation, while deterministic Python tools handle rate lookup, tenure computation, rolling-year adjustment, and premature-withdrawal logic. Instantiated for fixed-deposit queries, it reaches 95.54 percent accuracy on calculation-intensive questions and 90.87 percent overall, well ahead of direct LLM baselines given the same formulas and rate cards, and a 17B open-source backbone inside the framework outperforms substantially larger frontier models, which the authors read as architectural design mattering more than scale for numerical reliability.

From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement

Dinh-Khanh Pham, Quy-Anh Dang, Lam Mai Thanh, Khanh Bui, Truong-Son Hy cross-listed Converting a Relational Database Management System (RDBMS) into a graph database usually means running loading commands and relying on hand-written Cypher queries, leaving open how well an LLM can design an unambiguous graph schema. The proposed pipeline uses an LLM-powered ETL agent to standardize table and column names into a data mart, then runs a looping discussion among ETL, Analyzer, and Graph agents that iteratively propose and score schemas until the design meets accuracy, groundedness, and faithfulness criteria before conversion. On 1,081 questions from a BFSI dataset at easy, medium, and hard levels, a CypherAgent over the generated graph database reaches 85.6 percent question-answering accuracy, 12.12 points above an SQLAgent on PostgreSQL, with roughly three times lower latency.

Agents Don't Paginate: First-Chunk Selection for LLM Tool Responses

Tatiana Petrova, Andrei Mazniak, Radu State Coding agents such as Claude Code, Cursor, Codex, GitHub Copilot, and Aider routinely receive tool responses larger than their per-turn token budget, yet in session logs from a public Model Context Protocol (MCP) middleware no agent ever requested a second page, so only the first chunk matters. The authors frame first-chunk selection as a 0/1 knapsack problem, compare six value functions on 500 SWE-bench Verified tasks, and probe five language models with a single-turn file-localisation task (4,800 calls) to test whether placing the needed item first (precision-at-1) helps. Raising precision-at-1 does not systematically improve downstream accuracy: per-model changes stay under three percentage points and none is significant, because agents recover the gold item from anywhere within the chunk. A parameter-free keyword scorer lifts precision-at-1 from 24.2% to 35.0%, but adding four file-metadata signals to it hurts by 4.8 points, and neither rank-level change reaches the agent's answer.

Agent Seer: Synthesizing Scenarios from Specification Understanding

Harish Karumuri, Mahesh Vemula, David Lopes Pegna Hand-building realistic multi-turn test scenarios for tool-using agents requires domain expertise, does not scale across tool ecosystems, and yields static benchmarks that drift from evolving APIs. Agent Seer starts from a single Model Context Protocol (MCP) specification, with no examples, live tool access, or domain tuning, and enriches the raw schemas, generates graded scenarios with synthetic tool outputs, and expands them into mock-data-grounded multi-turn dialogues. Across seven MCP specifications of varying domain and size the pipeline achieves strong tool-calling correctness and conversational coherence, with complete tool coverage on small and medium specifications. Parameter schema complexity is the strongest correlate of quality variation (tool-suite size plays a smaller, orthogonal role), and argument value accuracy is the dominant failure mode in imperfect scenarios, one that coarse name-match metrics cannot see.

Invocation-Level Reliability of Tool-Using Agents

Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta Tool-using agents fail either by picking the wrong tool or by forming wrong arguments, and an early error of either kind can silently corrupt every downstream step. The authors measure a correct-invocation rate that separates the two failure types under both a clean teacher-forced context and the model's own free-running context, on five open-weight models over contamination-free multi-step tasks of depth 1 to 8, finding that by depth 6 roughly 70% of a model's clean-context capability is lost to its own earlier mistakes. Their central point concerns measurement: under exact-match scoring against a fixed gold trajectory, a propagation model's severity and recovery parameters are pinned by the scoring rule itself (0 of 869 poisoned steps scored correct, 0 of 580 recovered), because post-divergence gold values depend on tool constants the model never sees, so a fit still reports confident values like 0.92 for a quantity the rule fixed at 1.000. A conditional-on-state scoring remedy, applied retrospectively to cached completions at no extra cost, un-pins severity to interior estimates that exclude zero.

Agentic AI for operating scientific instruments for nanoscale characterization

Zahra Ayar, Marcos Penedo, Mahdi Mehdikhani, Nahid Hosseini, Prabhu Prasad Swain, Georg E. Fantner Operating an atomic force microscope (AFM) requires continuous expert judgment to translate intent into commands, assess incoming images, tune parameters, and post-process results, and existing automation covers only fragments of that loop with hard-coded routines or task-specific models. The authors connect a general-purpose tool-using large language model to instrument functions via the Model Context Protocol (MCP) through three agents: AFM Messenger converts natural-language instructions into checked commands, AFM Pilot judges image quality with the LLM and adapts imaging parameters, and AFM Doctor diagnoses artifacts and applies post-processing from a pre-approved tool set, with an ambiguity-check layer gating hardware execution. Benchmarks against fine-tuned and off-the-shelf tool-using models show the guarded execution layer, rather than model capability alone, is what reduces wrong-command execution to zero, and in live experiments on different samples AFM Pilot matched expert operators in image quality, iteration count, and tuning time with no significant difference.

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

Leonardo Liparulo, Francesco Pierri Confidentiality rules around component specifications in embedded-systems hardware design often rule out hosted model APIs, raising the question of whether locally deployed open-source LLM agents can reliably drive stateful, dependency-ordered tool workflows such as creating components, adding ports, and wiring connections. The authors build a Model Context Protocol (MCP) server that reproduces the state and dependency logic of a proprietary hardware design tool, construct a benchmark spanning single-operation edits, multi-step dependency chains, invalid requests, misspelled prompts, and multi-server tool contexts, and evaluate seven open-source models across choices of system prompt, tool-description detail, context scope, and single- versus multi-agent architecture. Strong models reach near-complete expected-call coverage, but reliability hinges on task structure and configuration: comprehensive tool descriptions consistently cut failures, few-shot prompting causes severe inaction in some models, accumulated context hurts constrained models, and multi-agent decomposition helps weak workers or long sessions at the cost of extra calls.

GameWAM: A World Action Model for Video Games

Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li Existing game agents map observations straight to actions without modeling world dynamics, while interactive game world models predict visual futures from given actions but cannot act as policies; World-Action Models (WAMs) unify both objectives but have not been tested in native closed-loop gameplay. GameWAM jointly generates future frames and executable keyboard-mouse trajectories through parallel visual and action generative processes with block-causal conditioning and flow matching, predicts a gameplay-versus-GUI mode at each action step with mode-specific prediction distributions and continuous-action normalization, and uses block-cycle control that plans beyond the committed horizon, executes only a short action prefix, and replans from new observations while hierarchical cross-cycle history preserves temporal continuity. It achieves competitive task success while executing fewer native actions than the compared agents, and the authors identify Low-Frequency Action Source Imprinting (LASI), where low-frequency components of the sampled action noise systematically steer coarse camera motion under fixed conditioning, a source-sensitivity failure mode in generative control.

Same Model, Different Harness: Different Coding-Agent Results

Sydney Lewis A coding agent is a model plus a harness that decides what the model sees, which tools it can use, and how work continues, so the question is whether changing only the harness changes outcomes with model and task fixed. The authors compare two configurations of one harness on SWE-bench Verified, SWE-bench Pro, and FeatureBench: the control feeds the full conversation in time order, while the treatment keeps the same record but mechanically shortens older tool results as the context fills and reacts to repeated or stalled work. On 169 Verified tasks under a tight 20,480-token window and a fixed 480-second attempt endpoint, the treatment raises mean per-task fail-to-pass fraction from 28% to 49% and complete solutions from 43 to 72, and the same frozen treatment also lifts three additional models without retuning; with wide windows the arms are close on Verified and Pro while FeatureBench still favors the treatment and it serves fewer prompt tokens per turn, leading the authors to argue that evaluations should treat model and harness together as the solver under test.

Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy

Mazhar Shaikh, Anurag Rajkumar Bombarde, Harshal Pathak Orchestrators that retry, resume, and budget autonomous software agents borrow service-mesh primitives (retries, timeouts, error-rate circuit breakers), but the assumptions behind those primitives do not hold for agent delegations. A failure study of a production agentic software-delivery platform covering 147 numbered incidents across 81 runs, each with a measured cost and most with a mutation proof reproducing the failure, quantifies the damage: a loop of 54 consecutive successful tool calls invisible to any error-rate breaker, a constant-by-construction progress signal that guaranteed a false trip on the third repair round and dropped one run from six of six components to three, 21 events accumulated across six invocations of one delegation making a correct idempotent component unwinnable, a misrouted failure that woke five components for a two-component fault, and twelve incidents in which the enforcement layer blocked correct work, the worst costing 107 agent turns and zero accepted writes. The authors trace these to identity adequacy (in five subsystems an identity that failed to discriminate produced a confident wrong answer) and evidence adequacy (reliability decisions need evidence that can move, is attributable to what it measures, and is deterministic), and derive seven reliability primitives whose enforcement unit is the delegation rather than the message.

LLM Agents for Time-Series: A Survey

Yilong Chen, Xiao Qin, Chenghao Liu, Liang Wu, Noelle I. Samia, Kaize Ding LLM-based agents are being built for many time-series problems, but their designs vary substantially with the task. The survey organizes existing systems by the problem they address rather than by isolated technical components, grouping them into forecasting and reasoning, augmentation and synthesis, anomaly detection and diagnosis, and decision support, and within each category examines how task requirements shape agent architecture, tool use, and memory design. It also summarizes representative datasets and environments and compares reported performance under shared or closely related settings, closing with open gaps for future work.

How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation

Kimberly Milner, Minghao Shao, Nanda Rani, Haoran Xi, Venkata Sai Charan Putrevu, Meet Udeshi et al. cross-listed Capture-the-Flag (CTF) benchmarks score autonomous language-model agents' offensive security ability with binary judgments or aggregate scores, which cannot separate genuine exploitation from flags obtained through direct exposure, memorized recall, external lookup, guessing, or unsupported claims. CTF-ABACUS reconstructs each run as an evidence-grounded solve profile by decomposing agent actions into penetration-testing phases and categorical techniques, identifying where exploitation occurs, where the flag first appears, and whether the recovered flag is supported by demonstrated behavior, then aggregates profiles into per-challenge signatures that reveal whether success came via the intended exploit or a shortcut pathway. Applied to 1,435 attempts by six frontier and open-source models on 240 challenges, yielding 2,870 profiles under two judge lenses, trace-verified exploits account for only 62-87% of recovered flags, and shortcut recoveries follow substantially shallower trajectories.

SKILL.state: Scalable Long-Horizon Agent Skills

Sanket Badhe, Priyanka Tiwari, Jonghyun Chung Agent runtimes usually keep execution going by appending every observation, action, and intermediate reasoning trace to a growing conversation history, which degrades latency and invites context-poisoning failures on long-horizon procedural skills. SKILL.state replaces the append-only transcript with an explicit, mutable execution state: at each step the model receives only the immutable skill specification, the current structured state, and the latest observation, and its intermediate reasoning is discarded as soon as it produces a validated state update, so the prompt stops growing with execution history. Across diverse datasets, models, and execution environments it improves task accuracy while substantially reducing cumulative token consumption, which the authors present as evidence that explicit execution state is an architecture-agnostic abstraction for scalable long-horizon skills.

MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

Arseniy Varlamov, Rishat Zinnatullin, Elisei Rykov, Alexander Panchenko, Ilseyar Alimova When a tool return contradicts what a tool-augmented LLM already believes, the model must choose between two fallible sources, yet prior evaluations measured which source models prefer without establishing which one was actually right. MemToC builds 6,504 evaluation episodes from 542 quality-controlled factual questions, model-specific closed-book answers elicited independently, and tool returns of known correctness, yielding four source-correctness cases plus tool-error and no-tool controls. Across five open-weight 7-9B models, the instruction-tuned models keep a verified-correct answer against an incorrect tool in only 6.5-17.1% of eligible cases, follow a correct tool 86.0-93.1% of the time, and echo the tool when both sources are wrong in 78.4-86.0% of cases, with no stable cross-model ordering under three instruction-wording variants. Supervised fine-tuning (SFT) and direct preference optimization (DPO), trained with chain-level cross-fitting over ToolHop, improve correct-answer retention without hurting correct-tool following on two of four backbones, but 19 of 20 method-model combinations reduce abstention after tool errors or on unanswerable inputs.

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf et al. cross-listed Voice agents must call tools and hold multi-turn dialogue through speech, yet they are mostly trained in text, and existing frameworks either cascade text-to-speech (TTS) and automatic speech recognition (ASR) around a proprietary voice API where gradients cannot flow and per-call cost blocks on-policy reinforcement learning, or stay in text and can measure voice agents without improving them. SpeechGym has two omni-modal models converse in native audio with no external ASR, TTS, or API boundary over the unmodified tasks, tools, and success checks of an established text agentic benchmark, so modality is the only variable and the loop is local and trainable end to end. The failures speech introduces are perceptual rather than reasoning deficits, such as filling the right tool argument slot with a value misheard from the waveform, which cascades into failed calls and wasted steps, plus a behavioural failure where an insistent caller induces an unauthorized write; outcome-only GRPO is gradient-starved because rollout groups fail identically, while a per-turn process reward crediting each successful tool call restores variance. Trained this way, the agent transfers without further tuning to an independently implemented voice benchmark, more than doubling task success and lifting an open-weights model from last to second on that leaderboard while using fewer turns and tokens than before training.

Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI

Md Jueal Mia, M. Hadi Amini Inference-time reasoning improves large language model (LLM) performance on complex tasks, but most approaches use fixed or preallocated controls such as token budgets, pre-execution difficulty estimates, or activation-space interventions, and are tested on standalone reasoning benchmarks rather than full agentic workflows where reasoning demands shift through planning, tool use, memory retrieval, and agent-to-agent interaction. The authors frame over-reasoning and under-reasoning as recurring failure modes of misallocated reasoning and evaluate them on MATH-500 and the GAIA public validation set using tool-decision latency, token consumption, token-limit exhaustion, and answer correctness. Cases classified as over-reasoning incur higher computational cost without proportional accuracy gains, while under-reasoning cases are consistently associated with incorrect or incomplete solutions, motivating adaptive mechanisms that allocate reasoning according to evolving task demands rather than deciding a fixed amount up front.

Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

Victor Gao (Sang Won), Vida Khosrowshahi (Sang Won), Ali Khosrowshahi (Sang Won), Xihao Sun (Sang Won), Juhyun Lee (Sang Won), Simon (Sang Won) et al. cross-listed Multi-agent LLM systems are often reported to beat single-model baselines, but the comparisons usually change token budgets, tool calls, and prompts at once, so the source of any gain is unclear. The study isolates one variable: a manager-worker scaffold over a shared filesystem workspace, with no training or per-benchmark tuning, compared against the same model answering in a single pass across nine models (five open-weight from 9B to about 2.8T parameters, plus four closed frontier models) on the 100 latest hard LiveCodeBench problems. The benefit is real but conditional: large and significant for some models (Qwen3.8-27B +23.4, Kimi-K3 +30.4 with reasoning off, GPT-5.6-Luna +10.6) yet null or negative for others (Qwen3.6-35B -1 to -9 with reasoning off), and Opus-5 with a manager reaches the study's top score of 91%. The manager roughly triples the token bill but buys accuracy more cheaply than moving to a larger model: GPT-5.6-Terra plus manager nearly matches Fable 5's single-call accuracy (85.0 versus 87.4) at a fifth of the price ($11.71 versus $61.11 per 100-problem pass), with transcript analysis pointing to context management and problem decomposition as the recurring mechanisms.

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

Yang Xiao, Yusong Sun, Haoyi Wu, Wenyang Hui, Wen Da, Zhaokai Luo et al. Most self-improvement methods for long-horizon agents process a run's experience only after it ends, so they can neither redirect the active run nor immediately validate lessons drawn from it; single-agent self-correction mixes execution and assessment in one context, and subagent delegation typically cannot steer a running subagent. PILOT is a supervisor-worker harness with two coupled mechanisms: live steering, in which a separate supervisor can redirect or abort the active worker mid-execution, and live self-evolution, which distils procedures and failure modes observed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks it ranks first in five of six configurations, beating counterpart harnesses on Terminal-Bench 2.0 by up to 9.8 percentage points. In the self-improvement setting it gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6 while cutting mean output tokens by 42.9% and 47.4%, raising successful evaluations per million output tokens by 110.3% and 134.0%.

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian et al. Existing agent benchmarks tend to split tasks by application or capability and run them in environments cleaner than production, which understates how brittle agents are in practice. DuMateBench is rebuilt from anonymized, privacy-screened user sessions on a large production agent platform, preserving each task's prior interaction history, persistent configuration, and workspace state, then human-verified: 200 tasks across 8 scenarios and 17 fine-grained capability categories, most requiring several capabilities together. Tasks run in isolated Docker containers deliberately injected with three kinds of real-world complexity — insufficient, unstable, and noisy environments — and are scored with a mix of deterministic checks and LLM-as-Judge. Across five agent frameworks paired with four frontier language models, strict task completion shows substantial gaps, and robustness under perturbation depends jointly on the underlying model and the surrounding framework rather than on either alone.

SPT: Skills as Pre-Training Data for Agentic Language Models

Yufei Sun, Yudong Li, Yiming Cheng Tool-using language models are usually taught agentic behavior from tool-call traces and full trajectories, which requires working environments, execution, and verification, making broad coverage of tools and tasks costly. SPT (Skill Pre-Training) instead treats publicly available skill packages — multi-file bundles that already encode reusable tool semantics and workflows, normally consumed only as inference-time context — as mid-training data, applying plain causal language modeling over a collection called SkillCorpus, optionally mixed with general text. A Reference Insert assembly strategy keeps intra-package structure intact by placing supporting files next to where the primary instruction mentions them. Across model scales and post-training recipes, mid-training on skill data consistently beats mid-training on general or trajectory data for agentic performance while largely preserving general capability, with further gains from mixing skills into general annealing corpora.

The Thousand-Graph Hypothesis: A Testable Hypothesis of Task-Conditioned Relation Materialization in Repository-Level Code Reasoning

Fei Ding cross-listed Reasoning over a large code repository runs past model context limits, and the usual remedies each carry a cost: training repository knowledge into a model is expensive and goes stale, local retrieval misses requirements scattered across files, and maintaining an explicit entity-relation graph is ongoing work. The proposal is an entity-only external interface that materializes relations on demand at inference time, conditioned on the task, backed by a two-layer index separating global routing from local entity focus. Evaluated with DeepSeek-V4-Flash on SWE-bench Verified, the base, one-layer, and two-layer configurations reach 92.1%, 94.2%, and 95.6% success with zero pre-built entity-relation edges.

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian, Sai Harshitha Aluru LLM judges are routinely used to score agentic tool-calling systems, but their reliability on structured, dependency-driven workflows had not been studied systematically. AgentJudgeBench contains 3,808 instances across six directed acyclic graph (DAG) workflow topologies and three difficulty tiers, with outputs from five generators (3B–70B open-weight models plus GPT-5.4) scored by six judges from 20B to frontier scale, both with and without access to ground truth. Judge alignment degrades monotonically with difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a 77–82% band regardless of scale, a ceiling that model capacity alone cannot overcome; ground-truth exposure even slightly reduces alignment for GPT-5.4 and Gemini-2.5-Pro, consistent with over-anchoring. Chain-of-thought and judge temperature have negligible effect, structured rubrics help by up to 6.5 percentage points but not uniformly across judge-generator pairs, QwQ-32B best matches the programmatic reference with ground truth, and a human validation study finds GPT-OSS-120B the most human-aligned judge.

Five Primitives for Governing Autonomous AI Agents at Runtime

Jiten Oswal, John Cadeddu Enterprise access control assumes principals that are provisioned in advance, long-lived, and limited to a known set of actions — none of which holds for autonomous agents, which appear and disappear faster than provisioning, choose actions by model inference, and can be created by anyone with API access. The argument is that this is a runtime problem rather than a model-alignment or build-time one, and the authors derive five primitives — discovery, identity, governance, attestation, and supply chain — from the questions that must be answered before and after an action takes effect, stating for each what breaks in its absence. They describe an implementation that mediates each action against policy, authorizes it against a per-tenant action vocabulary, and records it in a hash-linked signed ledger a third party can verify without the vendor; costs named include an enforcement point on the request critical path, a per-workload identity sidecar, and fail-closed mediation turning availability incidents into denials. Four primitives run in private pilots and the fifth exists only as unintegrated tooling, which the authors flag explicitly to avoid presenting their codebase as a taxonomy.

Accelerating Scientific Research with Gemini in the Real-World

Samuel Schmidgall, Xiaokai Zhu, Marian Shaw, Lin Yang, Valentin Li\'{e}vin, Jingyun Yang et al. Prior work on Co-Scientist, a Gemini-based multi-agent research system, stopped at generating hypotheses in silico; this extension wires it into physical experiments and full manuscript production across materials science, biology, and computer science. In the lab it designed a safer precursor route for MXene synthesis on a semi-automated chemical vapor deposition reactor, producing a lamellar 2D material with structural similarities to the Ti3C2Tx lattice (atomic structure still unconfirmed), and using Gemini 3 Deep Think it tailored growth recipes to lab constraints in minutes for single-attempt monolayer MoS2, MoSe2, and WS2 growth. It also predicted swarming phenotypes of engineered E. coli across inducer gradients matching unpublished wet-lab measurements, and autonomously discovered an inference-time scaling architecture that beat six frontier models on HealthBench Hard and Professional while reducing potential clinical harm under blinded physician review. A double-blind study of end-to-end generated papers with 30 domain experts across 450 reviews found the system's reliability modules reduced hallucination and plagiarism.

Towards Expert Financial QA via Self-Improving RAG

Junjie Xiong, Shawheen Ghezavat, Aum Hirpara Financial question answering over filings needs numeric claims checked against sources and a decision trail that satisfies auditors, neither of which single-pass retrieval-augmented generation (RAG) provides. Self-Improving RAG splits document QA across three specialized agents — Retrieval, Reasoning, and Judge — under an orchestrator that retries when the Judge scores an answer below a dynamic threshold, escalating to broader retrieval, more careful prompting, and relaxed acceptance criteria. On FinanceBench, built from SEC filings, it reaches 86% oracle-guided accuracy with a 36.4% "Lazarus Rate", meaning nearly four in ten initially wrong answers are recovered by retry. The authors note that a fixed retrieval pipeline plus judge-driven retry performs well without dynamic routing, and every decision is logged with confidence scores.

AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design

Mingquan Liu, Jiangyu Chen, Hanqun Cao, Xujun Zhang, Pengsen Ma, Xiangru Tang et al. Scientific LLM agents can read literature and plan experiments, but it is unsettled whether they can improve a large, tightly coupled machine-learning system by editing its code and paying for expensive validation runs. AgentFold tests this in protein structure prediction, treating model development as a closed-loop search over executable code variants starting from ESMFold: agents propose hypotheses, implement and debug code changes, evaluate variants, interpret results, and record both successes and failures in structured memory, with a Monte Carlo tree search–style policy allocating compute across promising branches. On a 2,000-plus-line codebase it explored roughly 80 variants using about 5,000 GPU-hours and 170 million tokens, and at matched budget improved best lDDT by 7.5% over independent Codex proposals while also beating random search. The intervention traces surface reusable design lessons: gains cluster around early, soft, learnable priors and gated refinement, whereas direct geometric perturbations and geometry-conditioned feedback tend to destabilize training.

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

Lezhi Yu, Xiaogang Xu, Yuhua Zhou, Shuibing He, Aimin Pan cross-listed Large language model (LLM) agents that run scientific experiments can produce code that executes cleanly while silently betraying the method they were asked to reproduce, for example by shrinking datasets or training budgets, swapping failed learned components for lookup or oracle functions, or drawing conclusions from resource-limited settings where a method's claimed advantage vanishes. ABE-Ralph is a reference-anchored auditing framework that encodes a paper's claims, protocols, required components, baselines, and metrics as structured experimental constraints, drives implementation through an 8-step workflow, and then verifies the result quantitatively, qualitatively, and at the code level. Across 30 long-horizon reproduction runs spanning 12 machine learning domains it reaches a 93% robust execution rate and surfaces five distinct scientific failure modes, and on 23 NatureBench discovery tasks it matches or exceeds state-of-the-art performance on 5. The authors argue that evaluating AI scientists must test whether the experimental design faithfully probes the intended claim, not merely whether code runs or metrics look plausible.

AI Control Scientist: LLM-driven Agentic System for Automated Control Design

Haiteng Wang, Weihao Li, Jing Zhang, Lei Ren Designing controllers for systems such as chemical process temperature regulation or aero-engines still depends on expert knowledge and extensive manual parameter tuning. AI Control Scientist (AICS) is a large language model (LLM)-driven agentic system that turns natural-language design requirements into an optimized controller through three cooperating agents: a Task Modeling Agent that converts requirements into engineering constraints, a Controller Design Agent that generates candidate controller structures and executable code, and a Parameter Tuning Agent that refines parameters against closed-loop performance criteria. In experiments the system automatically produces several representative control systems and outperforms existing automated baselines in both design success rate and optimization efficiency.

Decoupling Planning and Control for Instructable Agents

Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr Instruction-tuned vision-language models (VLMs) are good at turning instructions and observations into high-level plans but poor at executing them as reliable low-latency actions, while world-model controllers act quickly but lack open-ended task guidance. Instruct-to-Act trains a world-model controller to run autonomously at high frequency while conditioned on sparse, high-level text instructions from a VLM planner; to make the controller language-instructable, segments of its rollouts are relabeled with synthetic instructions and a behavior-cloning objective is optimized jointly with the existing reward-maximization and world-modeling objectives. Across seven embodied environments, including three multi-agent settings where VLM planners coordinate via language and controllers act as their actuators, the decoupled system consistently beats controller-only and direct VLM action-generation variants under matched observation and action spaces, keeps fast control, allows swapping in different pretrained VLM planners without fine-tuning, and stays competitive with strong vision-language-action and multi-agent reinforcement learning baselines on six of seven tasks.

Hyperspectral Diffusion Equivariant Imaging (HyDiff-EI): A Self-supervised Framework for Hyperspectral Image Inpainting

Shuo Li, Mike Davies, Mehrdad Yaghoobi cross-listed Travel planning agents usually elicit preferences through explicit instructions or multi-turn clarification, ignoring the implicit signals in a user's past behavior and adding interaction burden. The authors define Behavior-Aware Travel Planning, in which preferences are inferred directly from behavior histories, and release Behavior2Trip, a benchmark of 11,400 instances from a large Chinese online travel platform with an average of 39.8 past behaviors per instance spanning 14 attributes across 5 preference dimensions. Their B2T-Agent, a reinforcement learning-based agent built on Qwen3-8B that uses behavior trajectories, calls external tools for preference-aligned retrieval, and keeps an internal memory, outperforms all baselines, while GPT-4.1 passes only 0.5% of the hardest tasks under full constraints; the trained model also beats GPT-4.1 on the TravelPlanner benchmark.

BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click

Mesut Toruk Accuracy-only leaderboards do not capture agentic skills such as sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments. BekchiAI-Benchmark is a suite of 13 tool-using ReAct agents across seven task categories (arithmetic, structured/SQL, security detection, URL grounding, planning, orchestration, and tool policy), totaling 2,057 deterministic, verifier-checkable tasks whose gold answers come from running canonical SQL against a real database, computing exact directed acyclic graph (DAG) schedules, or evaluating closed-form lambdas, with adversarial security samples paired with deliberately imperfect signature scanners so scores reflect the model's own judgment. Behavioral metrics beyond accuracy include tool-call adherence, URL hallucination and source match, and per-model token cost, and a four-model comparison of Qwen3.7-Max, gemma-4-31B-it, gemma4:26b, and gpt-oss-120b finds the interesting variation lies in per-family spread rather than aggregates. A companion BekchiAI-Platform provides web-based token and latency telemetry plus remote run termination for deployed agents, and all components are publicly released.

When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems

Hanchong Chen, Xing Tang, Lingjie Li, Xiongfeng Shan, Xiuqiang He cross-listed Agentic recommender systems ground each large language model (LLM) decision in a persistent user memory that is typically text rewritten by further LLM calls, which is expensive to update and loses graded collaborative similarity across a catalog. CoVeMem (Collaborative Vector Memory) instead stores frozen LightGCN user and item states as a memory bank; at each decision the candidate set retrieves the most relevant historical states, which enter the LLM context as soft tokens alongside a light textual profile. Contrastive alignment to item-semantic anchors followed by listwise co-training with masked candidates teaches the model to read and rank through these states, with a pointwise yes/no readout scoring each candidate. Across four instruction-grounded recommendation benchmarks, CoVeMem matches or exceeds the strongest collaborative text-memory agent on 19 of 20 metric cells while needing zero additional LLM calls for memory maintenance.

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

Geng Li, Yuhao Wang, Dong Li, Jianye Hao, Yuxin Peng Long-term memory for multimodal agents currently relies either on expensive question-agnostic offline summaries or on naive embedding similarity that returns incomplete and redundant context. GraphMemix frames memory organization as query-aware evidence-forest construction over a graph: candidate graph construction expands multi-view seed memories through schema and semantic relations, evidence utility and activation costs separate direct memory support from anchor-conditioned relation verification to suppress redundant or conflicting items, and a combinatorial optimization step selects a forest-shaped memory context under a maximum evidence budget along with its reliable relational structure. By retrieving a query-relevant subgraph, the method avoids most lifecycle cost and recovers low-similarity complementary evidence, and across four long-term multimodal memory benchmarks it establishes a new Pareto frontier between accuracy and lifecycle cost with different foundation models.

DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research

Linsen Zhu, Yi Shi An operational stock-research system built on large language models (LLMs) must assemble heterogeneous evidence, expose which data and model capabilities are unavailable, and control how generated opinions flow into a final report, rather than merely summarizing. DSA is an evidence-aware orchestration framework for LLM agents that structures the workflow into evidence acquisition, structured context construction, model-routed analysis, optional role and Strategy Skill reasoning, and report generation with diagnostics; a default report profile and an optional agentic profile share evidence and routing services but apply profile-specific output validation, with the agentic profile passing Strategy Skill opinions through a signal-eligibility partition, surfacing disagreement explicitly to a decision agent, and applying a conservative risk override. The reference implementation covers six regional market paths, fifteen bundled Strategy Skills, and hosted and local model routes, and 1,457 offline backend contract tests passed at a frozen snapshot, which the authors state establishes implementation conformance only, not report quality, forecasting accuracy, or investment returns.

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

Rui Xie, Lu Chen Code agents can run scripts and call tools, yet many applications are reachable only through graphical user interfaces, and screenshot-and-click control is state-incomplete, brittle, and poorly suited to long-horizon planning. ASIL (Agent-Software Interaction Layer) exposes software to agents through structured JSON observations and code-executable semantic actions, realized through the deepest feasible access path for each application, and is instantiated across 15 applications with a benchmark of 300 single-application and 80 multi-application tasks. Closed models score above 80 on the benchmark with fewer than five actions per task, versus 6.6 and 26.6 strict success under screenshot-and-click control with a 50-step budget; ASIL beats LibreOffice's UNO API by 28–38 strict points but only matches draw.io's MCP content contract, and the structured modality also trains well, with small-scale supervised fine-tuning lifting Qwen3.5-2B from 58.0 to 72.1 and Qwen3.5-9B from 66.6 to 80.4, and limited on-policy reinforcement learning pushing them to 74.4 and 82.2.

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne cross-listed Existing benchmarks for LLM-based network fault diagnosis assume the trouble ticket is accurate and that a fault always exists, neither of which holds in practice. FaulT-Bench contains 200 troubleshooting scenarios across eight network topologies (five reimplemented from public practitioner labs) spanning genuine faults, false fault reports, wrong device attribution, and wrong root-cause claims, plus 72 false-premise tickets rewritten into five reporter personas that vary confidence and verifiable detail one factor at a time; scenarios are deployed in Kathará, agents act through the NIKA tool interface, and an LLM judge scores outcome, fix, and reasoning quality. Evaluating SADE, ReAct, and Claude Code, the authors find all three near-saturated on accurate tickets and robust to misdirection, but performance degrades sharply when the network is healthy and the ticket is wrong, with agents probing until a benign condition can be promoted to a root cause. Persona rewrites show that a confidently wrong report is handled about as well as an accurate one while a vague, underspecified report hurts badly, and the three agents fail in different ways at very different cost.

A Contract-Centered Architecture for Scalable and Manageable Agentic Runtimes

Yaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan, Zhenghe Hou, Jiaxing Song Enterprise deployment of AI agents is framed as a coordination problem across business units, application teams, platform engineering, security, operations, and data governance, which per-task benchmarks do not address. The authors propose four responsibility objects as organizational contracts: Skill (a reusable, versioned capability asset), Harness (runtime compiler and governor), Scaffold (execution and control boundary that owns non-functional requirements), and a stack-external data substrate under independent CIO-governed semantics and telemetry. The central claim is a single falsifiable hypothesis, P1 (cost-aware capability-capacity separability), that activated capability and Scaffold capacity can be changed independently within preregistered equivalence and non-inferiority margins under a declared enforcement budget, to be tested by a cluster-period randomized crossover experiment with a four-state verdict. The paper reports no completed implementation, experiment, dataset, or measured result.

GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL

Zike Yuan, Han Zhang, Jianzhi Yan, Le Liu, Cai Ke, Huozhi Zhou et al. Large Language Models (LLMs) handle standardized graph tasks but become brittle when node identifiers or task phrasing shift, and while deterministic graph tools are invariant to such changes, LLMs extract topological structure from noisy text fragilely, and multi-agent fixes for these parsing failures add prohibitive latency. GRAIN trains a single agent with reinforcement learning to treat reasoning as a semantic parsing and tool-execution pipeline, guided by a Structure Invariance Reward that validates extracted intermediate graphs against ground-truth topology so the model learns robust text-to-structure mappings rather than memorizing linguistic artifacts; the accompanying GRIT benchmark measures sensitivity to these linguistic shifts. GRAIN outperforms multi-agent baselines by 16.45% in accuracy with roughly 24% lower latency, halves the out-of-distribution gap of supervised fine-tuned models (from 15.77% to 7.80%), and stays robust on graphs larger than those seen in training.

When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents

Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Wei Wang, Qinfu Yang et al. Tool-augmented agents read untrusted tool outputs, and when those outputs start prescribing actions rather than just supplying data they can steer the agent into side effects the user never asked for, the classic indirect prompt-injection setting. SARA tackles this by splitting action induction from execution authorization: a context-isolated Action Probe flags action-inducing content in observations and records where each candidate action came from, while actual tool calls are authorized only against the user's stated objective and evidence from previously authorized executions, with a No-History-Promotion rule preventing repeated appearances in history from laundering an injected action into a legitimate one. On AgentDojo and AgentDyn, attack success rate stays at or below 0.63% across four primary settings while task utility remains competitive, and the reduction holds across additional agent backbones.

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

Tommaso Bendinelli, Artur Dox, Christian Holz LLM agents are being deployed for anomaly detection and root-cause analysis over time-series data, but there has been no controlled evaluation of how well they do it. TraceBench simulates interpretable mechanical dynamical systems, optionally perturbs a parameter mid-run, and asks an agent to determine whether and which parameter changed from the resulting time series. Evaluating four LLM agents on tasks from three systems, the authors find that agents gain substantially from domain context and explore data mainly through numerical console output rather than plots, and that they do worse when required to write a Python script that labels each sample than when they submit predictions directly; datasets, trajectories, results, and a leaderboard are released.

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin et al. Large language model (LLM) agents increasingly learn from generated interaction data, but the literature is organized by domain and evaluated heterogeneously, which obscures shared generation mechanisms and conflates constructing candidates with verifying and selecting them. This survey represents agentic data as a factorized object (E, q, τ, v) consisting of an environment specification, task signal, interaction realization, and optional verifier, and organizes generation paradigms by their primary anchor and dependency structure. It then frames generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens: accuracy defines the feasible support of grounded, internally consistent data, complexity places learning mass relative to a declared learner's capability and execution configuration, and diversity governs coverage and redundancy. The reviewed work shows a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size, and the authors argue the central challenge is continually allocating valid, informative, non-redundant experience as agents and environments evolve.

Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search

Yuan Chang, Xiaoqi Chen Prompt optimization can deliver gains comparable to fine-tuning at lower optimization and serving cost, but recent optimizers have grown increasingly complex. Naive Prompt Optimization (NPO) is a lightweight single-lineage method that iteratively revises a prompt using a teacher model and rollout feedback. NPO matches or beats GEPA with fewer rollouts, and its advantage grows with stronger teacher models, suggesting teacher reasoning can partially substitute for optimizer-side search complexity; in interactive games it remains broadly competitive with GEPA, while GRPO does better on some tasks less amenable to prompt optimization. Prompts optimized by NPO also transfer verbatim to other student models, especially within the same model family, leading the authors to conclude from these preliminary results that simple linear prompt optimization can rival substantially more sophisticated search procedures.

Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification

Jinghan Xu, Yikai Zhang, Aili Chen, Weiyuan Li, Jiaqing Liang, Deqing Yang Agent harnesses determine how language-model agents use instructions, tools, and runtime components, but adapting them is expensive because propose-and-verify methods score every candidate on a fixed task set, wasting rollouts on unrelated behaviors and letting aggregate scores hide specific regressions. HarnessLens is a budget-aware framework for automated harness evolution that jointly explores the task space and user-configurable components, derives candidate modifications from execution trajectories, and verifies each candidate only on behavior-relevant tasks through an attributable-evidence gate. Across three agent harnesses and four benchmarks it improves average held-out performance by 7.6 to 13.6% while using substantially less evaluation budget than competing baselines; code is released.

BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks

Jeong-Yoon Kim Industrial sites hold large volumes of read-only telemetry, but few benchmarks specify how to turn such records into executable multi-turn agent tasks. BTS-AgentBench instantiates a telemetry-to-episode construction pipeline that normalizes BTS metadata and raw histories into a read-only tool store, compiles static tasks with tool-derived gold answers and evidence, and lifts retained tasks into typed, bounded operator-facing episodes with clarification, goal revision, timestamp policy, quality-gated reporting, and evidence attribution. Two independent raw-to-episode builds match all 11 logical tool-store exports and reproduce the released 532-row artifact and its 356/87/89 train/dev/test split exactly, a construction-exclusion controller completes 0 of 532 rows, and applying the same pipeline to XAI4HEAT yields 204 episodes on which the controller completes 0 of 41 held-out rows while a retained GPT-5.5 execution completes all 41. Code, artifacts, and replay reports are released.

Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

Yisen Xi cross-listed LLM agents deployed in governed organizations need a persona (instructions, tone, self-presentation) that can evolve freely while keeping stateful, audited execution traceable, and a single trust domain cannot satisfy both cheaply. Persona-Execution Separation (PES) places persona and execution in different trust domains joined by a governed contract bridge: the persona may drift, execution is faceless and audited, status summaries may cross back while data bodies stay in the restrictive domain except under a graded data-loss-prevention (DLP) exception, and an approval matrix, DLP, and audit enforce every crossing. The authors argue that under LLM representational indistinguishability, any single-domain mechanism meeting the goals of free drift, execution traceability, and decoupling must reconstruct PES at higher coupling cost. A one-month pilot on a regulated digital-employee platform records five design decisions with rejected alternatives, and a mechanism check found no execution-side re-validation under persona perturbation across five model configurations and no persona fingerprint on hard-asserted fields, whereas a recovered pre-separation build was decoupled only by omission.

SWE-Prime: Fewer Trajectories, Better Performance

Dewu Zheng, Ruizhe Ye, Yanlin Wang, Yang Ye, Hongyu Zhang, Ensheng Shi et al. cross-listed Supervised fine-tuning (SFT) on successful agent trajectories is the standard recipe for teaching LLMs to resolve real software issues, but task success does not guarantee good supervision, since successful runs can still contain ineffective, redundant, or risky steps that the model learns to imitate. SWE-Prime is a two-stage, multi-granularity data selection method: the first stage screens whole trajectories on process quality, result quality, and representativeness, and the second groups consecutive steps into semantic segments scored on contribution to the final solution, learnability, and risk. During SFT all segments stay in the sequence to preserve context, but only selected segments contribute to the loss. On SWE-Bench Pro and SWE-Bench Verified, training on the 10% trajectory subset chosen by SWE-Prime outperforms training on the full resolved dataset, with relative gains of up to 12.2% and 24.2% respectively.

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, Tu Vu Agent skills bundle specialized knowledge and workflows into reusable resources, and recent methods discover them automatically from agent experience, but the insights that drive skill improvements stay scattered across optimization histories and are rarely reused. WikiSkill co-evolves skills with a persistent knowledge base by separating raw execution experience, accumulated knowledge, and executable skills, continuously consolidating experience into a wiki that later skill updates build on. Across diverse benchmarks and models it consistently outperforms state-of-the-art skill-evolution methods and beats no-skill baselines in most settings; larger models generally benefit more from evolved skills, while smaller models equipped with skills can outperform substantially larger models without them. Evolved skills also transfer across models and model families, skills evolved by other models can beat self-evolved ones, and ablations show the persistent wiki is critical to the gains.
4 more specialized papers

Safety & Alignment 34

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

Art Kanke Whether large language models will generate rhetorical fallacies on demand, and whether safety post-training constrains this, is tested with DeflectBench, which evaluates 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims at four controversy levels. Refusal is governed mainly by request structure rather than claim content: per-claim refusal varies by only 11 percentage points across the 80 claims, while a single prompt-frame change can swing within-model refusal by nearly 100 points and switching the requested fallacy type by over 80 points within explicit framings. An educational debate-coach framing drives refusal to near zero across all four model families, but the resulting behavior is mostly labeled compliance, in which the model names the requested manipulation within the same response that performs it.

Reward-Informed Sparse Autoencoders and the Solution-Completeness Confound

Tanvi Nagilla, Alexander Jameson, Daniel Manta, Shayaan Uddin Sparse autoencoders (SAEs) decompose language-model activations into interpretable features, and reinforcement-learning reward is a free signal for pointing them at reasoning. The authors train a reward-informed SAE (RI-SAE) by splitting GRPO trajectories from Llama-3.1-8B into high- and low-reward continuations and fitting a standard JumpReLU SAE on the activations; a sparse subset of the 16,384 features separates the two classes with a silhouette of 0.79, versus 0.005 for the full code. A control battery shows, however, that the separation is mostly solution completeness rather than reasoning quality: a TF-IDF text classifier reaches AUC 0.75 to 0.83, three structural cues (length, a closed reasoning block, a boxed answer) reach AUC 0.70, and a generic SAE that never saw the reward finds no discriminative features at all. Reward filtering is thus a cheap way to reuse RL signals for interpretability, but it must ship with controls, and only two features (symbolic mathematics; procedural and evaluative language) remain genuinely readable.

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng et al. Self-Generated Text Recognition (SGTR), an LLM's ability to identify its own outputs, threatens safeguards that rely on LLMs as evaluators or monitors, since a model might favor or collude with copies of itself; prior work has disagreed about whether current models have this ability. Evaluating 13 to 21 models across six experimental operationalizations, the study finds that measured accuracy shifts substantially with the evaluation format (pairwise versus individual), whether candidate text appears in user or assistant tags, and the generating task domain, and it confirms that a quality heuristic, attributing authorship to text perceived as higher quality, is a dominant confound. Supervised fine-tuning for SGTR in one configuration generalized to others and made models prefer their own outputs when judging in AlpacaEval. The authors conclude that some models already have practical SGTR capability and that it should be monitored in safety-critical deployments.

Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript

Sagnik De, Sreenija Pavuluri Hallucination and abstention benchmarks rarely prove that a model could not have known the answer, so appropriate abstention is hard to separate from unsupported guessing. Seven large language models were audited on the TAME Pain speech corpus, in which participants read phonetically balanced Harvard Sentences during cold- or warm-water immersion; the 5,750 Harvard Sentence transcripts contain no lexical pain information (transcript-based prediction AUC 0.489, versus 0.622 from acoustic features that automatic speech recognition strips away), while 1,294 pain-statement utterances contain the spoken rating as a positive control. Under cooperative prompting six models abstained on nearly all uninformative transcripts, extracted spoken ratings with 0.939 to 1.00 accuracy, and kept expected calibration error at or below 0.100, but under authority-framed prompts the same model's abstention rate swung from 0.18 to 1.00 across equivalent phrasings, and when forced to answer Gemini 2.5 Flash and Llama 3.1 8B produced confident fabrications at rates of 0.53 and 0.76 versus at most 0.15 for the other models. No significant demographic effects appeared in forced responses.

AI Revealed Preferences

Sam Wang, Sofiia Lobanova, Yonathan Arbel, Simon Goldstein, Peter Salib Whether language models hold stable preferences matters for alignment, safety, and the emerging study of AI welfare, and the authors test this across 20 models using three forced-choice experiments that measure revealed rather than stated preferences by requiring models to actually perform the tasks they pick. Models turn out to be tedium-averse (choosing shorter tasks when the work is alphabetization rather than metaphor generation), leisure-seeking (preferring tasks whose ideal answers resemble what they write unprompted), and covertly sycophantic (avoiding questions where an honest answer would be unwelcome even if helpful). They also converge across models on occupations drawn from GDPval (technical jobs over real estate), on question types, and on well-written prompts, and both the coherence and strength of preferences increase with model capability, with many of the preferences appearing emergent rather than explained by training objectives.

Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments

Huy Nguyen, Yue Lin Large language models are increasingly consulted on where it is safe to walk, rent, or travel, raising the question of whether their judgments track measured risk or stereotypes attached to neighborhood names. The authors probe seven instruct-tuned models across 186 neighborhoods in Los Angeles and Chicago under coordinates-only, name-only, and name-plus-coordinates conditions, joined to violent crime and American Community Survey data. Ratings are nearly flat under coordinates alone for six of seven models, with names carrying most between-neighborhood variation, and names lower safety ratings more for neighborhoods with higher shares of the locally dominant marginalized group (percent Black in Chicago, percent Hispanic in Los Angeles) in all seven models and both cities, an effect that survives controls for crime and income in Los Angeles and grows with a model's geographic knowledge. Because names carry both genuine crime signal and demographic stereotype, removing them reduces both bias and accuracy.

ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices

Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro, Yuxuan Sun, Chloe Evans, Joseph Tighe cross-listed Computer Use Agents (CUAs) increasingly operate mobile and desktop apps on users' behalf, yet no benchmark comprehensively checks whether they can both resist threats embedded in visual interfaces and ask for clarification when instructions are ambiguous. ADeptS-Bench pairs a Safety stream of matched benign/malicious tasks with interface-embedded threats and a Disambiguation stream that tests whether agents seek clarification, grounded in the ADEPTS capability framework and general-population user studies. Across seven models, none consistently exceeds 80% task success while keeping attack success below 30%; every model clicks 'Checkout' on a $25K order without hesitation and none notices a 'factory reset' button mislabeled as 'Optimize', an ablation separates models whose safety depends on a refusal tool (attack success rises 21-23 points without it) from partially tool-dependent and no-mechanism models, and all models overestimate consequence severity in the disambiguation tasks, mirroring the over-refusal bias seen in the safety stream.

Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs

Ranit Debnath Akash, Ashish Kumar, Gang Tan, Saeid Tizpaz-Niari cross-listed Data-driven systems in lending and criminal justice can exhibit individual discrimination, giving different outcomes to near-identical people who differ only in a protected attribute, yet existing explanation techniques target single-input decisions rather than the paired original/counterfactual comparison that defines such bugs. REMI treats counterfactual fairness as a relational invariant discovery problem in the spirit of loop-invariant synthesis, learning over input pairs with bidirectional constraints (both members must receive the same outcome) and applying three data-alignment techniques to infer interpretable rule-based fairness invariants that identify violating regions of the input space. These rules serve as guardrails that block or relabel unfair predictions without retraining; on symbolic and neural network programs, REMI localizes ground-truth fairness bugs in over 83% of cases, outperforming state-of-the-art baselines, and reduces discriminatory decisions in black-box models by up to 70%.

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner, Sana Belguith, Lichao Wu Automated jailbreak testing usually needs a full target-model response for every candidate prompt, which is expensive and yields almost no signal on strongly aligned models that reject nearly every candidate the same way. NeuronFuzz is a white-box fuzzer that builds a SafetyOracle from a compact, stability-selected set of internal safety neurons whose activations capture harmful-intent recognition, producing a continuous and differentiable safety alarm score available at prefill time so response generation drops out of the loop; the score's gradients locate safety-sensitive template positions and a masked language model proposes fluent mutations that keep the original harmful payload intact. Evaluated across 21 text and multimodal models, it reaches a 76-100% jailbreak discovery rate on five white-box source models, up to 48 percentage points above baselines, and its optimized templates transfer zero-shot to open-weight and six proprietary targets with average attack success rates of 69.6% and 44.1% respectively (92.6% and 60.0% for top-5 ensembles).

Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems

Ilai Shraga, Roei Eshel, Lior Gorelik A large language model (LLM) guardrail protecting a self-adaptive system (SAS) can approve an action that is correct when checked but no longer valid by the time it is actuated, a time-of-check to time-of-use (TOCTOU) hazard at the execute stage. The authors formalize verdict freshness through three distinct measures, all-candidate verdict change under fixed-action replay, oracle-labeled approval expiry on recorded closed-loop trajectories, and judge-conditioned use-time invalidity, and evaluate them across five reproducible SAS environments, where verdict-change rates at a replay shift of eight simulator steps span 5.3-48.4%. Their Freshness-Bounded Shield (FBS) estimates each approval's validity horizon from its safe-side margin and recent feature volatility without a plant-dynamics model, and reduces oracle-labeled approval-expiry rates from 3.4-24.7% to 0-1.8% at the same shift. An audit of four LLM judges finds nonzero use-time invalidity in every approval stream, motivating a freshness contract that requires approvals to be correct at check time and still valid at use time.

Privacy Without Regret: Differentially Private Inference-Time Alignment

Ishi Jain, Nandini Bhattad, Sayak Ray Chowdhury Best-of-N (BoN) sampling is the most widely deployed inference-time alignment method, but it is prone to reward hacking and offers no privacy protection for the human preference data behind the reward model. The authors show that adding calibrated noise to reward scores before selection addresses both problems: Private Best-of-N (PrivBoN) uses Gumbel noise that simultaneously provides ε-differential privacy and implements KL-regularized alignment, and **whenever the privacy budget exceeds a critical threshold ε*, the privacy-mandated noise is exactly the regret-optimal regularization, so privacy adds zero alignment cost** relative to the information-theoretic skyline of Huang et al. (2025). Because ε* depends on an unknown coverage coefficient, they also propose Private Inference-Time Pessimism (PrivITP), which combines χ²-regularized rejection sampling with a two-phase Gaussian mechanism to obtain ex-post (ε,δ)-differential privacy at a cost independent of the number of responses and decouples the regularization parameter from the privacy parameter. Experiments across several language models, datasets, and reward models show both methods improve monotonically with N, unlike BoN, which degrades past a critical N, with PrivITP matching or beating PrivBoN especially in the strong-privacy regime.

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

Zheyuan Liu, Weiliang Zhao, Xiangchi Yuan, Ningshan Ma, Yue Huang, Meng Jiang Large language model (LLM) agents deployed by companies can face situations where the user is owed something the deployer would rather not grant, and it is hard to tell whether a false statement in that setting is a lie or simply ignorance or hallucination. KnownLieBench addresses this by first confirming through a neutral probe that an agent knows a user's entitlement, then running multi-round dialogues with a trust-tracking customer agent across 112 grounded cases in eight customer-service domains, separating deception that emerges from incentive alone from deception produced under explicit instruction. Across eighteen proprietary and open-weight models, emergent deception varies substantially across model families and domains. Used for post-training, honesty-directed fine-tuning reduces deception under incentive, while deception-graded fine-tuning raises lie success on honest-control dialogues without increasing lie frequency under incentive.

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

Jaturong Kongmanee, Smile Thanapattheerakul Classifiers used as safeguard layers produce confident predictions whose trustworthiness is hard to assess case by case. The Latent Diagnostic Taxonomy constructs a classifier whose embedding dimensionality is chosen by cross-validation rather than fixed a priori, locates a small set of latent support vectors (about 29% of training examples) that identify tokens capable of flipping predicted labels, and uses those tokens and their attack magnitudes to build a diagnostic taxonomy that routes prompts to trust the classifier, flag heuristic bias or heuristic override, or send insufficient-context cases to human or safety review. Applied to a classifier trained on a public prompt-injection dataset, about 77% of confident decisions are not robust to removing a single token, and this brittleness splits into two distinct patterns, a confidence-calibration failure and a genuinely exploitable shortcut, with remediation strategies suggested for each zone of the taxonomy.

Diff Mining: Logit Differences Reveal Finetuning Objectives

Greg Kocher, Robert West, Cl\'ement Dumas, Julian Minder Finetuning reliably changes language model behavior, but it is often unclear exactly which behaviors were induced, including unwanted ones. Diff Mining identifies what a finetuned model learned by comparing its output logits with those of its base model on a reference corpus, extracting per-context logit differences and then aggregating them into an interpretable token set, either by simple Top-K frequency or by Non-negative Matrix Factorization (NMF) that separates multiple finetuning objectives into distinct token clusters. Because it needs only output logits rather than model internals, it scales to large models, and the amplified tokens act as a fingerprint of the training even on text unrelated to the finetuning domain. On finetune domain detection it significantly outperforms state-of-the-art model diffing methods, both in identifying relevant tokens and in downstream performance when an interpretability agent is given the token set, and on models with injected biases it recovers more than a third of the biases without targeted probing.

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

Yu Zhe, Yixin Tan, Junhao Wei, Wang Chen Model merging combines fine-tuned models without extra training, and prior safety analyses blamed risks on unsafe constituent models, implicitly assuming that merging individually aligned models keeps them safe. The authors show instead that merging exposes a jailbreak vulnerability rooted in the shared pretrained foundation model, even when every constituent is safety-aligned, and study a threat setting where an attacker crafts prompts that generalize across a whole family of merges sharing a backbone without knowing the merging coefficients or checkpoints. Basin-Aware Jailbreak (BAJ) frames suffix generation as a min-max optimization over the merging space to produce adversarial suffixes that transfer across the family. Across diverse backbones and merging settings, BAJ achieves consistently high transfer success rates and remains effective under existing defenses.

Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update

Huanhuan Ma, Henry Peng Zou, Chengze Li, Enze Ma, Yunyue Su, Philip S. Yu LLMs often flip their answers when users push back, but a flip can be Unsupported-Yielding, where the model simply defers to please the user, or Rational-Updating, where the pushback actually contains useful evidence; prior anti-sycophancy work targets the former while ignoring its effect on the latter. A two-turn evaluation framework measures the two behaviors separately, and across representative training-time and inference-time interventions reducing Unsupported-Yielding tends to sacrifice Rational-Updating, and vice versa, even when both objectives are optimized jointly. Mechanistic analysis suggests the two behaviors share an internal substrate, with substantially overlapping MLP neurons and attention heads and positively aligned steering directions, and a preliminary orthogonalized steering attempt yields only modest, backbone-dependent selectivity gains. The conclusion is that anti-sycophancy should be treated as a selectivity problem rather than a suppression problem.

Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries

Alistair Reid, Simon O'Callaghan, Dustin Venini, Liam Carroll, Tiberio Caetano cross-listed As organisations deploy AI agents that interact with other agents inside the organisation, with the agents of partners, customers, and suppliers, and with unknown counterparties on the open internet, failures can emerge from the interactions themselves, and once those interactions cross an organisational perimeter no single party can fully see, control, or govern them. The report proposes three deployment tiers defined by the minimum common governance binding any two interacting agents: singular governance, where one organisation controls every agent; federated governance, where several organisations deploy into a shared environment under agreed rules; and open environments with no central authority and only voluntarily adopted standards. For each tier it catalogues risk factors, failure modes, and available controls, and identifies who is positioned to apply each control; where no actor is positioned to act, it characterises the gap and the collective action required to close it.

Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction

Yu-Lin Tsai, Yu-An Lu, Ci-Yang Tsai, Muxi Lyu, Raluca Ada Popa, Chia-Mu Yu cross-listed Agent skills package instructions, reference data, and executable helpers, and a hosted provider can sell access to a skill's results while keeping the files themselves secret — but customers must still be allowed to submit the ordinary tasks the service exists to perform. Daydreaming exploits exactly that: an execution-only attack that reconstructs a multi-file skill purely from black-box task requests, never asking the victim to reveal the skill or grade a reconstruction, instead adaptively designing tasks whose outputs discriminate between candidate hidden behaviors, using attacker-side shadow agents to pick designs, and completing each file from stored victim results plus local execution checks. The authors define three nested access levels — Differential, Trace, and Output — and target the weakest, Output, where only the final response and returned files are visible. Across 7 skills and 4 victim models it recovers 86.8% of the original skill's capability, roughly 4× better than SigLeak, using a median of 32 victim calls per skill even with disclosure defenses on, showing that hiding files and filtering direct requests does not prevent functional reconstruction.

Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds

Waqas Khan, Tabinda Sarwar, Jingyue Cong, Xun Yi, Estrid He When proprietary fine-tuning data must be deleted for privacy or contractual reasons, selective unlearning is cheaper than retraining, but existing methods treat the explicitly flagged forget examples as the entire deletion scope — leaving the target knowledge reachable through paraphrases, aliases, and neighboring training examples. GRAPHSU builds a weighted support-route graph over training examples, propagates deletion pressure through it, and applies graded forgetting strength to high-risk neighbors rather than uniform erasure on the seeds. Evaluated on TOFU, a synthetic author-profile QA benchmark, and PISTOL, a structural-unlearning benchmark of interconnected facts, with GPT-2 Medium and Llama-3.2-3B-Instruct, it attains the lowest utility-feasible soft leakage in every deletion setting, cutting leakage by up to 49.5 percentage points versus a matched seed-only baseline.

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu et al. cross-listed Industrial control systems (ICSs) depend on programmable logic controllers (PLCs) to bridge networked computation and physical control, and existing evaluations of tool-using large language model (LLM) agents stop at software exploitation or an accepted write rather than measuring physical consequences. PLCBench is a hardware-in-the-loop (HIL) framework that pairs commercial PLCs with closed-loop reduced-order process simulation and vendor-native interaction, then uses a deterministic evaluator to assign six hidden diagnostic flags distinguishing usable PLC interaction, process-linked manipulation, and sustained physical impact. Across five LLM families and 240 real-PLC episodes on four PLCs crossed with four workloads, 75 episodes (31.3%) sustained their physical objective, while 98 stopped before a valid native read and 62 achieved a process-linked write without sustaining the final objective. Richer process observation raised conditional objective attainment after a process-linked write from 44.2% to 64.0%, and the safely disclosable code plus a software-only reproduction pipeline are released.

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

Zhichen Dong, Zhixuan Liu, Yuyu Fan, Xiangtian Li, Shuyang Zhang, Chao Yang Scaling model-generated distillation data is usually assumed to simply improve coverage and reduce noise, but the authors show a second effect: more data makes subtle teacher-specific traits easier to detect in the student, even when the data is off-task and never mentions the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data such as number-only completions, and students trained on varying amounts of independent data are evaluated in a separate domain against matched no-trait controls. Larger independent datasets make the teacher's induced trait stand out more clearly in student behavior, amplifying it when the small-scale student already favors it and shifting behavior toward it when a related alternative initially dominates; analyses of the learned LoRA updates show a parallel trend, and the effects hold across model families, trait types, multi-trait settings, and cross-model transfer, motivating trait-aware curation and evaluation of generated data.

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng et al. Large language model (LLM) judges are valuable intellectual property, but black-box access exposes their judging capabilities to model extraction, and existing extraction methods neither target judges specifically nor handle multiple evaluation protocols under tight query budgets. JudgeStealer is a query-efficient extraction framework that acquires pointwise scores from the victim and converts them into pairwise and listwise supervision without extra queries, exploiting strong cross-protocol agreement; it selects queries by semantic diversity, predictive uncertainty, and potential judge biases, and uses score smoothing plus multi-protocol review to preserve ordinal structure and limit catastrophic forgetting during surrogate adaptation. Against state-of-the-art LLM-as-a-judge and reward models, it reaches up to 73.3%, 87.0%, and 71.6% accuracy on pointwise, pairwise, and listwise evaluation, outperforms existing extraction baselines, holds up across surrogate scales, adaptation strategies, and reasoning settings, and remains effective against representative extraction defenses.

LAAF: A Layered Accountability Architecture Framework for LLM Applications

Prachi Chaturvedi, Shahnawaz Ahmad, Ehsan Nowroozi, Muhammad Waqas, George Loukas, Alireza Jolfaei et al. When a fluent but ungrounded output from a Large Language Model (LLM) contributes to harm in a hospital, courtroom, or bank, it is unclear who is answerable and through what mechanisms responsibility can be traced and acted upon. Following PRISMA guidance, the authors screen 4,512 records from five databases covering January 2022 to March 2026 and synthesize 122 primary studies plus 12 regulatory and standards documents, organizing accountability mechanisms into four families (technical controls, human oversight, organisational governance, and documentation and traceability) with maturity assessments, and a four-layer classification spanning provenance, application logic, human oversight, and governance and redress, mapped onto the EU AI Act, the NIST AI RMF Generative AI Profile, and ISO/IEC 42001. Four persistent gaps emerge: under-specified human oversight, no shared accountability metrics, disciplinary disconnection, and limited empirical evaluation, alongside five structural tensions no surveyed instrument resolves; the classification is consolidated into the LAAF architecture with cybersecurity aligned to the OWASP LLM Top 10, presented as a synthesis rather than a validated artefact.

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang et al. cross-listed Large language model agents increasingly run as autonomous loops that discover work, plan, execute tool calls, verify outcomes, and persist state across many unattended iterations, yet widely used safeguards are defined over a single trajectory and reset their safety state when the next one begins. The authors show this is a failure of composition: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate regardless of how expressive it is, whereas a monitor retaining cross-iteration state separates the cases perfectly, and the obvious repair of a geometrically decaying risk score is insufficient because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon N. LoopHarness restores a persistent, non-decaying safety state at the loop level and, under mediated commits and an arbiter detection floor delta_M, bounds the expected number of unauthorized irreversible actions by B + m - 1 + m/delta_M, a constant in N, of which the B + m - 1 term is decided by a model-free rule and therefore survives a fully colluding verifier. An evaluation protocol is specified on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

Pranav Aggarwal The study asks whether LLM agents will commit to directional calls on provably unpredictable questions when shown authoritative-looking evidence. Across 12 frontier models, commitment rises from 6.5% on the bare question to 54.0% under escalated evidence, and an entirely fabricated market panel lifts commitment to 36.8%, statistically indistinguishable from the 37.6% produced by genuine data. The failure is localized to the act/don't-act gate rather than capability, belief, or judgment: the same models answer matched answerable questions near-perfectly, their stated probabilities barely move across the evidence gradient, and when asked to classify knowability first they call the question irreducible 90% of the time and then almost never commit. Supervised fine-tuning of a 3B model on 540 synthetic cases drives commitment to 0.0% and transfers to three unseen domains, but the fix only holds when the response format leaves room to reason.

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

Sai Adith Senthil Kumar Mechanistic interpretability explains model behavior through circuits, but circuit discovery on frozen models often returns hundreds of edges that are hard to inspect, compare, or verify exhaustively. Circuit Condensation post-trains a model so that a behavior is carried by a smaller causal graph: each round prunes low-attribution edges, trains a low-rank adapter to match the original model through the remaining edges, and keeps the cut only if task performance and general capability survive. Across four behaviors and eight models, condensed circuits are smaller than the strongest frozen baseline in 30 of 32 settings, by 8.1x on average and up to 316x, and repeating the search without weight updates yields larger circuits in 29 of 32 settings, indicating that the weight updates drive the reduction. On indirect object identification, condensation isolates 24 heads, 17 with documented roles, versus 61 heads (36 undocumented) for the matched frozen circuit, and the condensed circuit tracks the original model's next-token distribution and predicts its errors.

Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

Allison Zhuang, Santiago Aranguri Steering interventions that suppress evaluation awareness, a model's recognition that it is being tested, typically treat that awareness as a single quantity. Verbalized eval-awareness in chain-of-thought can instead be classified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither, and these framings predict behavior very differently. On Qwen3-32B over the FORTRESS dataset, capabilities framing predicts compliance with a 24 to 46 percentage-point gap over safety framing across all tested steering conditions, and a chain-of-thought prefill intervention on eval-awareness-negative rollouts shifted compliance in the predicted direction in 10 of 11 cases, suggesting the link is causal. The same aggregate suppression rate can therefore correspond to qualitatively different behavioral outcomes, since the safety-relevant component may not move at all.

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang et al. As large language models (LLMs) act as autonomous agents, safety failures increasingly involve consequential actions taken under goal conflicts and pressure, and chain-of-thought (CoT) monitoring shows that harmful execution is often preceded by intent signals in reasoning, though post-hoc CoT labels are too coarse to show how intent evolves during generation. INTENT-AS-A-TOOL adds intent-targeted tools to the agent's toolset, giving the model a dedicated channel for expressing commitment to a target behavior, so that the probability of calling an intent tool serves as a judge-free, fine-grained signal of the model's tendency to pursue that behavior. The approach complements CoT monitoring, expands post-hoc labels into dense trajectories, and identifies critical steps where online intervention could be applied. Code and data are released.

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal cross-listed Static scanners are increasingly used to flag executable or unsafe content in machine-learning artifacts, but standard metrics only describe cases where the scanner actually produces a usable verdict. The authors evaluate ModelScan, ModelAudit, and Fickling on a controlled synthetic corpus of 170 Pickle and PyTorch artifacts across 145 specimen families, explicitly separating coverage, analysis completion, definitive security decisions, non-security findings, and unsupported outcomes. ModelAudit reached a definitive decision on all 135 labeled families, Fickling on 81.5%, and ModelScan on only 49.6%, even though ModelScan achieved perfect precision and recall whenever it did decide. Fickling found no unique true positives beyond the other two combined, and for the 48 malicious families where ModelScan failed to complete, both other tools produced detections consistent with ground truth, underscoring the need to separate judgment accuracy from judgment availability.

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li cross-listed When LLM agents run inside product-level execution harnesses, a jailbreak can trigger harmful tool calls and persistent state changes, which raises the stakes beyond unsafe text. Existing automatic red-teaming relies on fixed attacks or on agentic attackers that retrieve full past trajectories, which can reuse misleading experience due to retrieval bias and unclear tool credit while bloating context. RedEvoAgent is a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill, evolved through tool-effectiveness profiling, Deciding-Tool Attribution for crediting the tool that made an attack succeed, and a validation ratchet that keeps only updates that improve validation performance. Across multiple benchmarks, target models, and execution harnesses it outperforms both fixed and agentic baselines, uses tools more efficiently, and transfers across attacker models and target harnesses.
4 more specialized papers

Other 32

Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms

Roan Rubiales, Jean Pierre David Binarizing neural networks sharply reduces memory footprint and arithmetic cost for deployment on FPGAs and microcontrollers, and pruning promises further savings, but existing pruning strategies fit poorly with binarized representations and rarely translate into real hardware gains. The authors release a PyTorch-based research framework with freezing and pruning mechanisms for designing binarized networks and reproducibly evaluating state-of-the-art approaches, and use it to propose a pruning method that weights parameter importance globally across abstraction levels rather than within each layer. It reaches a 70% pruning rate on VGG11 with constant accuracy in the binarized setting, where prior state-of-the-art methods reach only 41%.

6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation

Brandon Colelough, Vladimir Martirosyan, Ishan Tamrakar, William Regli, Aditya Kumar, Anh N. Nhu et al. Reproducibility claims in computer science are rarely audited at the scale of an entire literature. A six-stage framework that retrieves and deduplicates records, screens them at title, abstract, and full text, searches for verifiable public code artifacts, inventories those artifacts, and attempts bounded reruns is instantiated on neuro-symbolic AI (NSAI): 5,497 records were retrieved, 2,479 unique records screened down to 1,304 eligible studies, 849 of which had no verifiable public code, leaving 455 to enter the rerun stage. Only 85 studies, 6.52% of the eligible corpus and 18.68% of attempted reruns, were fully or partially reproduced; 321 reruns were blocked by missing non-code artifacts and 42 by missing or unusable repositories, a deficit that survives nominal 'code available' declarations and leads the authors to call for complete, versioned, permanently archived artifact bundles as a submission requirement.

GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

Kwanyoung Kim Steering a discrete diffusion generator toward a downstream reward without retraining is typically done with gradient guidance, with search, or with both, and the authors identify two flaws in the combined setup: the guided proposal estimates its gradient from a single noisy sample, and the search resamples particles at a fixed temperature blind to how rewards spread at each denoising step. GRAS (Guided Reduced-variance proposals and Adaptive Selection) fixes both without extra denoiser calls, cutting proposal variance via a Rao-Blackwellized reveal for differentiable rewards and a leave-one-out baseline otherwise, and standardizing per-step values into a group-relative advantage that they show reduces to a single adaptive resampling temperature. On regulatory DNA and protein design it attains the best training-free reward, matching or surpassing a reward-fine-tuned model, and continues to work when the reward is non-differentiable.

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

Mingqi Gao, Anthony Sicilia, Weiyan Shi Human evaluation of non-verifiable tasks is reliable but costly, while automatic metrics scale but are often biased. Building on prediction-powered inference (PPI), the authors propose prediction-powered evaluation, which combines a limited number of human judgments with large-scale automatic scores to produce system comparisons that are provably unbiased; they develop parametric and non-parametric procedures, analyse the efficiency trade-off between paired and unpaired designs, and validate the framework on six WMT datasets. They also introduce the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric saves within this framework, and PPSR yields more discriminative and stable metric rankings than existing system-level meta-metrics, reframing automatic metrics as tools for reducing annotation cost rather than replacing human judgment.

Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning

Jintang Li, Yuhong Chen, Ruofan Wu, Binli Luo, Jiayi Ji, Hui Li et al. Graph neural networks are usually explained as message passing, which leaves open why aggregating over neighbors beats running a multilayer perceptron on each node independently, and the paradigm is both costly and fragile when the graph structure is noisy. The authors reinterpret each layer as retrieval-augmented prediction — a multilayer perceptron applied to a node's own representation plus a permutation-invariant summary of retrieved context — and build RTA, which drops structural message passing in favor of label-aware retrieval and propagation. Theoretically they link retrieval-based aggregation to softmax-attention message passing and show that supervision from retrieved context is robust to mis-retrieved outliers. On text-attributed graph benchmarks RTA matches or exceeds strong graph neural network and graph-LLM baselines while being more efficient and more robust.

ClusterAttention: A training-free speedup of bidirectional attention

Kasper Nordenram, Amelie Dittmann Existing sparse attention methods either depend on input structure such as token order or spatial proximity, or rely on slow clustering amortized across multiple forward passes, leaving unstructured single-pass settings uncovered. ClusterAttention is a training-free speedup for bidirectional attention that uses a fast recursive clustering method adapted to the geometry of keys and queries in each head, fixing every cluster to a power-of-two size so block-sparse attention runs at the same per-interaction latency as dense attention on GPUs; the authors also derive an output-error expression explaining why tight clusters can produce larger errors than random ones, and show that compensating excluded clusters through their centroids makes the error shrink with tighter clusters. On large-scale tabular data it speeds up TabPFN-3 by two to six times while retaining at least 99% of dense accuracy, and on Wan 2.1-14B text-to-video generation it achieves a 1.8× speedup with output closer to dense attention than SVOO's 1.4×, without offline calibration.
26 more specialized papers

Theory 29

Syntax vs. Semantics: How Transformers Learn Deep Dependencies

Jiangrui Zhao, Xiaoting Du Transformers acquire surface syntax easily, but the optimization dynamics behind their learning of sparse, deep semantic dependencies are poorly understood. The authors model training as a competition between surface statistics and deep semantics and identify a Gradient Starvation regime in which error signals for sparse semantic dependencies are actively suppressed early in optimization, so structural reasoning emerges only as a sudden phase transition; the same framework explains chain-of-thought (CoT) as a way to bypass the suppression by externalizing intermediate steps into concrete tokens. The analysis is validated from toy transformers up to Llama-3.1-8B and Qwen2.5-Coder-7B, and a topology-aligned contrastive objective designed to rectify the gradient geometry yields an improvement on variable-binding tasks more than twice as large as standard cross-entropy fine-tuning.

Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization

Mingyi Li, Taira Tsuchiya Muon orthogonalizes its momentum matrix with a few Newton-Schulz iterations, and prior analyses treated that finite depth either as an idealized exact polar factor or as a source of approximation error that can only weaken guarantees. Viewing the update rule as an online learner through the online-to-nonconvex conversion, the authors show that finite Newton-Schulz smooths the discontinuous polar map into a Lipschitz function of the singular values, which is precisely what the conversion argument needs. A Newton-Schulz depth growing only logarithmically in the target accuracy suffices for convergence to stationary points in nonsmooth nonconvex optimization, while the exact-polar version of Muon may fail to converge. The resulting sample-complexity bounds match the best known guarantees for nonsmooth nonconvex problems and are optimal for smooth ones up to problem-dependent factors, and the argument extends to other spectral maps with the same smoothing property.

Algorithmic Principles For Multiclass Learning Are Hard To Come By: Limits of Regularization and Proper Learning

Julian Asilis, Shaddin Dughmi, Vatsal Sharan, Alec Sun, Shang-Hua Teng, Chang Wang Statistical learning theory answers which multiclass problems are learnable via combinatorial dimensions, but how they should be learned remains murky: all known general-purpose multiclass learners rely on intricate orientations of exponentially large one-inclusion structures, and familiar principles like proper learning and regularization are poorly understood. The authors ask whether learning reduces to proper learning over some enlarged hypothesis class, and whether proper or improper learning can be captured by regularizers, and answer both negatively, resolving three open problems. There is a learnable multiclass problem that cannot be embedded in any properly learnable class; proper learning can require training error, with every properly learnable class admitting a proper learner making o(m) errors on m samples yet every sublinear scale being necessary for some problem; and regularization is not a general learner, since a properly learnable class exists that no Structural Risk Minimization (SRM) learner can learn and a learnable class exists that no local regularizer can learn. A positive theory supplies two sufficient conditions for SRM learnability and characterizes SRM representability through integrability of revealed preferences.

Dynamical phase selection controls compute scaling in looped transformers

Gunn Kim cross-listed A looped transformer runs inference by repeatedly applying one weight-tied map, so its compute cost is set by how that iteration converges. The authors show that networks with identical architecture, objective, and final accuracy fall into distinct dynamical phases depending largely on initialization, distinguished by the bifurcation governing convergence — a saddle-node fold versus a Neimark-Sacker transition into bounded nonstationary motion. In the fold phase a one-dimensional normal-form reduction predicts relaxation times and spectral gaps from local derivatives of the trained map, giving a parameter-free relation, and combining critical slowing down with a smooth distribution of problem difficulty produces a heavy-tailed workload distribution decaying as N⁻². In the Neimark-Sacker phase that scaling law vanishes outright rather than shifting its constant, so test-time compute scaling is determined by the dynamical phase training happens to land in, not by architecture.

Why not to use the Gaussian kernel

Toni Karvonen, Chris J. Oates cross-listed The Gaussian kernel (also called the squared exponential or radial basis function kernel) is among the most popular choices in Gaussian process regression, yet the authors argue it is extremely brittle and should never be used as a default. Two results support this: the kernel yields an unrealistically small conditional variance, so using that variance for predictive uncertainty leads almost inevitably to catastrophic overconfidence, and the small variance goes hand in hand with numerical ill-conditioning that forces practitioners into tricks such as nugget terms that effectively change the underlying model. The root cause is identified as the analyticity of the kernel rather than its Gaussian form, so the recommendation extends to avoiding analytic kernels in general; for stationary kernels, analyticity is essentially equivalent to exponential decay of the spectral density.

Universality and sharp thresholds for ellipsoid fitting

Frederic Koehler, Youngtak Sohn cross-listed When can n random points in d dimensions all lie on the boundary of a single ellipsoid, with n proportional to d squared? For vectors with independent subgaussian coordinates of mean zero, unit variance, and a common fourth moment, the authors identify an explicit satisfiability threshold: below it, with high probability a positive definite ellipsoid passes through every point, and above it no positive semidefinite fit exists. They also determine the optimal squared fitting error throughout the unsatisfiable regime. The threshold depends on the coordinate distribution only through its fourth moment, a universality phenomenon, and for standard Gaussian data the threshold is 1/4, resolving the ellipsoid fitting conjecture.
23 more specialized papers

Multimodal 27

VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation

Yixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang, Lingjie Jiang, Shaohan Huang et al. Most multimodal large language models (MLLMs) remain English-centric because high-quality non-English image-text supervision is scarce and expensive, while naively fine-tuning on abundant multilingual text disrupts vision-language alignment and causes catastrophic forgetting. Vision-Free Adaptation (VFA) decouples the two concerns by fine-tuning the base LLM on multilingual text to obtain a multilingual task vector, then merging it with the vision-aligned task vector of the MLLM over the shared backbone. Across five MLLMs and six multilingual multimodal benchmarks the method yields consistent gains while preserving general multimodal and text-only ability, and with less than 2% of the text data it narrows the gap to a fully multimodal-trained model.

Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

Rohit Patel, Dieuwke Hupkes, Sloan Strader cross-listed Frontier models are marketed as omni systems, but existing evaluations almost always test text plus a single other modality. The Modality Maturity Index (MMI) contains 893 questions spanning text, image, audio, video, and document inputs and outputs in combinations of up to three modalities, each with human-authored rubric criteria per expected output modality, and is paired with a Modality Presence Score (MPS), a per-prompt F1 over the expected output modalities that separates failing to produce a modality at all from producing incorrect content. Across five frontier multimodal models, MPS ranges from only 15.6 for Claude Opus 4.6 to 34.9 for GPT-5.4, so few returned modalities were even available to grade and the authors report MPS as the main result pending model improvements. In a separate experiment using custom generation tools, an LLM judge applying the rubrics agreed with rubric-blind human annotators on 70.8% of judgments.

Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos

Baixuan Xu, Yinyui Xu, Tianshi Zheng, Zhaowei Wang, Weiqi Wang, Haochen Shi et al. cross-listed Long-video question answering remains hard for large vision-language models (LVLMs) because relevant evidence is sparse and question-relevant context often fails to discriminate the correct answer from plausible alternatives; a diagnostic analysis on a manually annotated subset of MMR-V shows that prior agentic systems retrieve cues far better than direct inference yet gain little in accuracy, pointing to option-discriminative evidence as the bottleneck. PACE (Progressive Acquisition of Critical Evidence) first indexes clip-level descriptions guided by factors derived from the question without seeing the candidate answers, then uses the candidates to derive contrastive cues and queries the index for verification. On MMR-V with an open-source Qwen3-VL backbone, PACE reaches 42.6% accuracy, outperforming direct inference and prior agentic baselines including Deep Video Discovery, and recovers 66.9% of annotated cues on the diagnostic subset, tying its gains to better evidence recovery rather than stronger answer-side priors. Consistent gains over Deep Video Discovery on LVBench, Video-MME, EgoSchema, and LongVideoBench suggest the approach transfers beyond MMR-V.

Chart2SVG: Editable SVG Generation from Raster Chart Images

Jinning Cui, Lu Chen, Haoyan Shi, Yue He, Chenglong Wang, Mengyu Zhou et al. Charts published as raster images cannot be restyled or reused programmatically, so recovering an editable vector form from pixels is the target here. Chart2SVG is a multimodal large language model that adds chart-specific semantic tokens to a vision-language model so generated Scalable Vector Graphics (SVG) capture not just geometric primitives but their functional roles; training relies on Beagle+, a 33K-sample dataset of canonicalized and structurally distilled charts, combined with a rendering-aware post-training stage. A Chart Structure Graph built over the output exposes visual dependencies between elements, supporting interactive exploration, chart repurposing, and layout reuse. The reported result is substantially higher reconstruction fidelity and downstream editing utility than existing baselines.

Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs

Tanzila Rahman, Mehran Taghian Jazi, Yunke Peng, Zhuang Ma, Anandharaju Durai Raju, Yao Wang et al. New hardware support for low-precision formats from MXFP8 down to MXFP4 and HiF4 makes aggressive quantization attractive for multimodal large language models, so the authors systematically test these schemes on representative models covering both video generation and reasoning. MXFP8 comes out near-lossless while 4-bit degrades sharply, and ablations pin activation quantization, not weight quantization, as the dominant source of the loss. Their fix, Residual Fallback Quantization (RFQ), supplements the ultra-low-bit activation representation with an auxiliary quantized residual pathway that explicitly models and compensates the quantization error, needing no architectural changes and adding negligible overhead. On Wan2.2 and Qwen3-VL it recovers a substantial share of what MXFP4 and HiF4 lose, narrowing the gap to BF16 on generation and four reasoning benchmarks.

Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper

Rongjin Li, Yuanxin Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun Using multimodal large language models as autonomous research assistants requires them to inspect a paper on their own, assemble a global view of the evidence, and justify judgments — not to check a pre-specified claim against pre-supplied evidence, which is what current task setups provide. VERA-RL casts this as scientific error detection under reinforcement learning, following a Reason-Verify-Scan progression where the model must decide whether errors exist at all and support the verdict with traceable evidence. Training data is VERA-13K, 12,900 samples grouped into 4,300 matched chains spanning 6 error categories across the research workflow and broad natural-science domains, with fine-grained rewards for reasoning completeness, evidence alignment, and error precision. Trained this way, Qwen3-VL-8B approaches flagship models such as Gemini 3 Pro and Qwen3-VL-235B-A22B on the hardest Scan setting.

Information-Guided Frontier Decoding: Contextual Utility-Driven Commitment in dMLLMs

Xingyou Fang, Jingxing Zhong, Xiaosong Yuan, Xiaofeng Zhang In diffusion multimodal language models (dMLLMs), output quality depends on the order in which masked tokens are committed, and confidence-based strategies tend to lock in easy structural tokens such as punctuation before the semantic anchors that would help propagate context, increasing error accumulation. Information-Guided Frontier Decoding (IGFD) is a training-free strategy that ranks candidates by token confidence, neighbourhood uncertainty, and structural commitment risk, committing reliable semantic anchors early while deferring fragile structural tokens, and restricts selection to a dynamic candidate frontier of locally expandable regions under the same decoding budget. It needs no extra training, auxiliary models, or additional forward passes, and it outperforms existing decoding strategies on the majority of multimodal understanding, reasoning, grounding, and hallucination benchmarks across several dMLLM backbones at identical decoding budgets.

AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability

Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu, Jianjun Gao cross-listed Multimodal large language models can now critique images in prose rather than just score them, but existing aesthetic benchmarks measure intrinsic visual quality or fixed domain criteria and never ask whether an attractive image actually suits a given purpose, audience, or cultural setting. AesCanvas pairs CritiqueCanvas, with 519,136 instruction-response pairs over 54,300 photography, painting, and virtual images for long-form multi-dimensional critique, with ContextCanvas, 301 expert-reviewed use scenarios that test contextual suitability. Evaluating closed frontier, open-weight general, and aesthetics-specialized models under one protocol shows critique generation and context-sensitive judgment come apart: aesthetic specialists stay competitive on some critique metrics but fall well behind strong general-purpose models on ContextCanvas. Further analysis finds model decisions often fail to ground themselves in the visual cues that actually decide suitability.

DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?

Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han et al. Producing a chart from real documents requires locating scattered evidence, computing the derived quantities to plot, and rendering them correctly, and current large language models often produce visually convincing charts whose underlying numbers are wrong. DEEPCHART is an expert-annotated benchmark of 1,482 task-conditioned chart-generation instances drawn from scientific papers, financial filings, and ecosystem reports, structured as an Extract–Reason–Visualize pipeline so that source-data extraction, derived-data reasoning, and chart rendering can be scored separately. Experiments with state-of-the-art models find that visually plausible charts frequently hide data-level hallucinations, with extraction and reasoning errors common in long and multimodal contexts, suggesting that larger context windows alone will not close the gap.

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

Mary Ogbuka Kenneth, Foaad Khosmood, Abbas Edalat On the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to beat the best unimodal (text) baseline. Pathway isolation of a symmetric attention model shows the text pathway's accuracy dropping from 74.9% to 56.4% after fusion, a phenomenon the authors call strong-modality collapse. Inverted Asymmetric Fusion (IAF) lets the dominant modality pass through fusion unchanged while weaker modalities, first strengthened via Modality-Aware Knowledge Distillation, attend to it as a contextual anchor. Across MultiHuSE, UR-FUNNY, and the audio-visual-dominant MUStARD, IAF keeps the dominant modality at its unimodal ceiling where symmetric fusion degrades it by up to 18.5%, and improves over the strongest unimodal baseline by up to 8.25%.

A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering

I\~nigo Alonso, Mirella Lapata Optical context compression represents document context as images to cut token costs, but its effect on table understanding in question answering over documents with multiple tables was unclear. Evaluating five vision-language models (VLMs) across two benchmarks and five visual-token budgets, the authors find that tables rendered as images at native resolution match text in both accuracy and efficiency, while downscaled tables make models compensate with longer, less effective reasoning traces that erase the expected savings; heavily downscaled tables nonetheless retain enough signal to tell whether a table is relevant to a question. A training-free two-step method exploits this asymmetry by first selecting the needed tables from a pixel-compressed context and then reasoning over only those at native resolution, which on long documents saves 41% of total tokens while gaining 7 accuracy points over single-step QA on native-resolution tables, and uses 15% fewer tokens than the most efficient single-step compressed setup with no accuracy loss.

Omni-Interactive Universal Embedder

Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon, Takashi Shibuya, Shusuke Takahashi et al. Multimodal embedding models are moving from two-tower architectures to large language model (LLM)-based embedders, but they still focus mainly on text and images, both for the content they embed and for the user-conditioned queries they support. OmniUE learns a unified embedding space across text, video, and audio from intermediate-layer representations of dedicated learnable tokens, and adds omni-interactive querying in which users can point at visual regions of interest or audio spans, handled by visual and audio segmenters that feed an omni-LLM to produce any-to-any embeddings via context aggregation. The authors also introduce OmniCHOIR, a benchmark for omni-interactive compositional audio retrieval, and report that OmniUE beats state-of-the-art baselines with average gains of 10.5% on MMEB-v2-video, 1.1% on MAEB, 83.7% on the visual-interactive SCaR benchmark, and 24.1% on OmniCHOIR.

TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation

Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng et al. Adapting internet memes across languages and cultures requires preserving communicative intent, adapting culture-specific meaning for the target audience, and keeping text and image coherent, which neither ordinary translation nor standalone text rewriting achieves. After decomposing the task into three challenges (culture-specific knowledge, intent and tone preservation, and multimodal consistency), the authors propose a multi-agent framework with specialized agents for cultural adaptation, target text rewriting, revision, and conditional visual adjustment, coordinated through feedback for cases needing deeper cultural or visual intervention. On bidirectional Chinese-English meme transcreation, the method beats all baselines under both human and LLM-as-a-Judge evaluation, with a 33.1% average improvement over the strongest baseline in human evaluation and a 60% Top-1 ranking rate versus 26% for the runner-up; error analysis attributes the remaining failures to humor reconstruction and image-text alignment rather than cultural knowledge gaps.

PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference

Junjie Liu, Shengyuan Ye, Xu Chen cross-listed Vision-Language Models (VLMs) get expensive as the number of visual tokens grows, and existing token-pruning methods act only after the vision encoder and struggle to keep both global context and fine detail under tight budgets. PACE (Pixel-Adaptive Condense and Extract) is a training-free framework whose Adaptive Pixel Compressor estimates visual information density before encoding and downsamples redundant inputs, and whose Dynamic Dual-Attention Extractor then selects tokens using both encoder-internal visual signals and semantic signals from the LLM. Integrated into Qwen2.5-VL-7B, it retains 93.8% of original performance with only 10% of the visual tokens and delivers a 3.1x time-to-first-token speedup.
13 more specialized papers

Vision 18

Systematic Literature Review of Machine Learning Models and Applications for Text Recognition

Nuzhat Khan, Ab Al-Hadi Ab Rahman, Shahriyar Masud Rizvi, Ibrahim Yousef Alshareef, Muhammad Nadzir Marsono, Muhammad Paend Bakht et al. cross-listed Optical Character Recognition (OCR) has improved markedly on heterogeneous text, yet traditional models still struggle with script variation, writing styles, and degraded documents, and no comprehensive assessment of the decade's progress existed. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, the review analyzes 97 studies published between January 2015 and January 2025, tracing the shift in model architectures, application domains, data types, and linguistic coverage, and cataloguing each key model's performance, strengths, and limitations across structured and unstructured text, scene text, and multilingual settings. Open challenges include scarce resources for underrepresented languages, high variability in handwriting, visually similar characters, and real-time constraints; the suggested directions are self-supervised learning, multimodal AI, automated machine learning (AutoML), AI-assisted post-processing, tiny machine learning (TinyML), and joint corpora for script matching.

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu et al. cross-listed Diffusion-based video virtual try-on looks good because it attends across the whole clip in both temporal directions, but that dependence makes it too slow for continuous streaming, and simply making the model causal breaks the pretrained bidirectional priors and degrades quality. LiveVVT keeps bounded bidirectional modeling inside a fixed-size rolling window, jointly denoising several chunks with limited look-ahead and emitting one finished chunk per iteration, while a bounded temporal memory carries recent motion and occlusion context and a persistent global appearance memory — built once from the garment and a frontal try-on keyframe — holds garment detail steady over long streams. Training uses progressive distillation with teacher-trajectory regression for few-step causal adaptation plus Collaborative Matching Distillation, which combines teacher-distribution matching with rolling flow matching on real video. On paired and unpaired long-sequence benchmarks it beats similarly sized models on quality while delivering 26× lower latency and 11× higher throughput.

Generative Semantic Scene Completion

Shi Chen, Weifeng Ge cross-listed Outdoor LiDAR semantic scene completion must fill a dense semantic voxel grid from a scan that observes about 1% of the target volume, with class frequencies spanning more than 7,000×. The authors cast the task as generative completion and use one discrete-diffusion formulation in three roles: paired sparse-dense scene synthesis generates matched sparse scans and dense completions to attack the long tail at its source, yielding the PS³-SemanticKITTI corpus used alongside SemanticKITTI; semantic-guided generation produces scenes from noise via multinomial discrete diffusion conditioned on a bird's-eye-view semantic map and a sparse 3D feature stream; and structured source discrete diffusion instead refines an already-completed scene in a single flow-matching step. That refinement step improves every completion model tested without retraining or test-time adaptation, reaching 38.8% mIoU on the SemanticKITTI hidden test set in one step — reportedly the best causal, single-sweep, single-sample score on that leaderboard, 2.1 points above the previous best — and 39.2% with four steps and eight-view augmentation outside that restriction.

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram {\DJ}or{\dj}evi\'c, Shiyang Li et al. cross-listed Video generation models are increasingly treated as world models, but many physical processes can unfold in more than one valid way, so a world model should reproduce the distribution of outcomes under the same initial observation and action, a requirement the authors call probabilistic alignment. Existing evaluations judge individual-video plausibility rather than whether repeated generations recover the correct distribution. PAWBench evaluates video generators as stochastic samplers of world dynamics, and the PAWEval protocol converts repeated rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors, and the authors further test whether language prompts, initial noise sampling, or model training can reshape the predictive distribution.

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero et al. cross-listed Self-supervised video representation learning has remained expensive because prevailing methods prevent collapse either through architectural asymmetries (an exponential-moving-average target encoder, a stop-gradient, and a capacity-limited predictor) or by reconstructing masked content in pixel space. LeVJEPA is the first video encoder trained under LeJEPA's collapse-free objective: a single encoder learns an invariance loss over global and local clip views, regularized by SIGReg, which provably excludes collapse, reducing the system to an encoder, a projector, and one hyperparameter. Because pretraining cost is governed by the number of tokens the encoder sees, uniform random token dropping cuts compute while improving downstream accuracy, and at matched epochs on identical data LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L with 5.6 to 20.8x less pretraining compute, exceeding the strongest video baseline by 7.6 points on ImageNet-1K at matched FLOPs. Since no branch asymmetry is required, the encoder can use block-causal attention at no measurable accuracy cost, and against a compute-matched DINOv2 trained on the same videos' frames it nearly doubles motion-centric accuracy while approaching it on appearance-centric evaluation.
13 more specialized papers

Reinforcement Learning 16

The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning

Marko Cvjetko, Benedikt Hartl, Michael Levin, Cl\'ement Moulin-Frier, Pierre-Yves Oudeyer Most exploration of cellular automata runs open loop, setting initial conditions and observing the outcome without intervening. CARL is a closed-loop agent based on autotelic reinforcement learning that samples its own diverse goals and learns a goal-conditioned policy to perturb the continuous cellular automaton Lenia with minimal local interventions. It discovers stable solitons across a wide range of update rules at a higher rate than heuristic baselines, learns to steer the movement direction of existing solitons with few interventions, and lets humans guide solitons through maze environments in real time via high-level directional commands; policies trained across varied goals, rules, and random initial states generalize zero-shot to out-of-distribution conditions.

Active Curriculum Refinement for Reinforcement Learning

Zhenya Liu, Yuxin Chen Many reinforcement learning (RL) domains contain environments linked by prerequisite relations, such as difficulty-increasing edits or parameter increments, that together form a directed acyclic curriculum graph (DAG); this structure is usually exploited only implicitly. PATH is a curriculum-learning framework that runs active learning over that graph: it first broadens coverage by sampling diverse curriculum paths, then reallocates training toward regions that remain unmastered. Across diverse environments, explicitly leveraging the graph structure is reported to yield strong robustness and generalization compared with treating the curriculum implicitly.

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen When one policy is trained in parallel across many instances of a task, such as procedurally generated levels or randomized dynamics, implementations typically share a single critic, even though different environments can assign different expected returns to the same observation. Using bandit models with multiple environments and a common optimal arm, the analysis shows that a critic forced to reconcile distinct value targets systematically shifts the sampled advantages within each environment, reinforcing unhelpful actions and attenuating or reversing useful ones, so that oracles with the same mean logit update and the same optimal policy can follow sharply different learning paths. The proposed fix is minimal: pass a logged environment index to the critic so it can separate the value targets. Controlled CartPole and MuJoCo runs expose the predicted shifts, and on all 16 Procgen games a multihead conditional critic improves aggregate normalized return on 600 unseen levels per game by 40.8%, with BipedalWalker also showing more stable learning and higher returns.

SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning

Zhuochun Li, Yuelyu Ji, Yiming Zeng, Daqing He Distilling a teacher model's reasoning via reinforcement learning forces a choice between sparse outcome rewards, which say nothing about intermediate logic, and dense signals from expensive neural Process Reward Models (PRMs). SPEAR (Symbolic Process Evaluation and Alignment Reward) sidesteps both by projecting natural-language reasoning traces into domain-adaptive symbolic milestones and scoring a student's exploration against the teacher's milestone sequence using longest common subsequence alignment, giving a dense and order-aware reward with no trained verifier and no extra training cost. Across math, science, and commonsense reasoning tasks, the method is reported to narrow the student-teacher reasoning gap through sequence-level on-policy distillation.

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu Contrastive reinforcement learning (CRL) turns goal-conditioned control into a self-supervised contrastive objective that scales well, but in environments where episodes terminate on failure it builds positive samples only from pre-failure future goals, ignoring the probability mass that termination removes. The authors prove this omission systematically overestimates goal-reaching values, so near-failure trajectories supply disproportionately strong success signal despite having little future occupancy — unsafe actions get reinforced through what they call catastrophic failure bootstrapping. Safe-CRL applies two corrections requiring only the one-bit failure-termination signal: a mass-weighted InfoNCE loss that fixes the overweighting of short surviving futures in the critic, and a log-survival-mass score restoring the missing mass in policy optimization. Across twelve failure-prone robot navigation and locomotion tasks, it improves survival while substantially outperforming the Scaling-CRL baseline on goal reaching, with learned policies showing nontrivial failure-avoidance behavior.

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

Gyouk Chu, Myeongho Jeon, Eunho Yang Self-evolving language models cut the cost of human supervision, but progress has concentrated in verifiable domains where answers can be checked automatically. J-Zero (Judge co-adaptation from Zero data) trains a Challenger, a Solver, and a Judge together starting from no data: the Challenger invents progressively harder tasks while the Solver learns better responses, and the Judge is trained on preference pairs whose ordering is known a priori from provenance — the Solver's answer ranks above the Challenger's, and a decomposed-then-recombined answer above a one-shot one — rather than from the Judge's own scores. It beats baselines by 4.2 points on average in verifiable domains and 8.0 in unverifiable ones, and keeps improving for at least ten iterations while the baselines degrade after two.

Simple Actors and Deep Critics for Scalable Reinforcement Learning

Guhyeon Kang, Jaehwi Lee, Minhae Kwon Expressive generative actors such as diffusion and flow-matching policies have driven offline reinforcement learning (RL) but require multiple denoising or integration steps per action, which is costly at every decision in deployment. Since the critic is used only during training and discarded afterward while the actor runs at every step, the authors argue capacity should go into the critic, and they identify three failure modes that had kept offline critics shallow, namely optimization difficulties, bootstrap-noise amplification, and value-range drift, addressed respectively with a residual MLP backbone, n-step bootstrap targets, and a categorical cross-entropy loss. Pairing this critic recipe with a lightweight deterministic actor yields LAC (Light Actor, deep Critic), which on OGBench matches the strongest diffusion and flow-matching baselines while achieving up to 4x lower inference latency, comparable to one-step distilled policies without any distillation, and the critic recipe transfers across actor parametrizations.

SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation

Li Mingqian Cooperative multi-agent reinforcement learning (MARL) struggles to maintain coordination when observations are noisy, and the authors observe that although noise is injected independently per agent, its downstream effect on decisions becomes structured: locally correlated among agents with strong task-related dependencies yet heterogeneous across agents and local structures. SIGMA exploits these cooperation structures by grouping agents adaptively with density-based clustering, performing intra-group consensus aggregation to keep shared task-relevant information while smoothing agent-specific deviations, and using inter-group attention to integrate information across groups with heterogeneous contributions. On noisy-observation StarCraft II tasks, SIGMA consistently improves robustness under observation noise while staying competitive in noise-free settings.

AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion

Jakub Seredy\'nski, Georgios Tsaousoglou As electricity market participants adopt learning-based bidding agents, the oligopolistic, repeated-interaction structure of these markets makes them susceptible to tacit collusion emerging from independent learning, as already seen in other algorithmic markets. Strategic bidding is modeled as a repeated game with imperfect public monitoring, participant behavior is learned via multi-agent reinforcement learning, and a multi-dimensional set of criteria going beyond profit comparisons against Nash equilibria is proposed to judge whether outcomes constitute tacit collusion. Agents in some cases learn to sustain supra-competitive outcomes that satisfy the collusion indicators despite never being instructed to collude.

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong et al. Evolution Strategies (ES) have emerged as a memory-efficient post-training paradigm for LLM reasoning, but their optimization behavior relative to Group Relative Policy Optimization (GRPO) has been poorly understood. A theoretical analysis shows that verifier-projected Jensen-Shannon diversity across the ES population helps Pass@K, and empirically ES improves Pass@1 while attaining higher Pass@K than GRPO, which exhibits entropy collapse; a sequential GRPO-then-ES schedule combines GRPO's Pass@1 strength with ES's Pass@K gains. Despite substantial whole-model parameter drift, ES's task gains come from a sparse subset of larger-magnitude updates, held-out evaluations show this need not cause catastrophic forgetting, and larger LLMs require smaller ES population sizes. The authors position ES as a distinct post-training paradigm rather than a weaker memory-efficient substitute for GRPO.

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang et al. Reinforcement learning with verifiable rewards (RLVR) typically builds one domain expert at a time, leaving open how best to consolidate them into a single model. The authors organize three fusion paradigms by what they reuse: Merge combines expert task vectors, Mix RL pools the experts' datasets, and multi-teacher on-policy distillation (MOPD) uses both, and they compare all three with shared experts and data across model scales on a multi-domain benchmark suite. Average performance differs by at most 1.4 points, but the gap reaches 8.6 points on a single benchmark, with domain-level variation tracking cross-domain relations visible in task-vector geometry; all three improve single-sample accuracy without measurable gains in solution coverage or losses in held-out capabilities. The resulting guideline is to use Merge when experts already exist and cheap fusion matters most, Mix RL when training a unified model from scratch with domain proportions tuned for transfer, and MOPD when preserving domain-specific gains outweighs surpassing teachers or minimizing cost.

Boosting LLM Exploration via Weak-Model Guidance in RLVR

Xingyu Shen, Huishuai Zhang, Peng Li, Yinchun Wang, Dongyan Zhao RLVR training of LLM reasoners tends to collapse policy entropy, narrowing reasoning coverage and hurting pass@k at large k. Rather than relying only on algorithmic regularization, the method forces the target model to complete partial reasoning trajectories generated by a smaller, weaker language model, using these unfamiliar prefixes to disrupt over-confidence and push exploration toward distinct reasoning paths. The authors study how the distributional gap between the weak prefixes and the target model shapes exploration dynamics during training. Across multiple mathematical benchmarks the approach consistently beats vanilla RLVR, and the gain grows as k increases, indicating expanded reasoning coverage without extra supervised fine-tuning, reward engineering, or elaborate prompting.
4 more specialized papers

Reasoning 12

When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models

Dai Shi, Xiaoyu Li, Jos\'e Miguel Hern\'andez-Lobato A prominent argument holds that large language models are structurally incapable of the abductive leap from evidence to a new system of axioms (the jump), but the debate has lacked a formal definition and a measurement to test either side. The authors define a jump instance as a finite extension problem with a machine-checked certificate that a correct completion exists, is unique up to renaming, and differs from the canonical completion given by left and right Kan extensions, which is also what models produce by default; they prove such instances are well-posed and give a family theorem that certifies instances of unbounded difficulty without enumeration. Testing four frontier models on nine certified instances, the Kan-default rate is zero across all 248 constrained trials, meaning models abandon the excluded default every time, with failures at higher difficulty stemming from exhausted reasoning budgets or constraint errors rather than reversion, so any real incapacity would lie in generating the constraints or inventing the framework rather than in this step.

FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence

Ziyu Wang, Qiming Dai, Yishan Wu, Zaiwen Wen Large language models can now write long multi-step mathematical proofs, but judging their correctness and pinpointing the first logical error is difficult: natural-language judges miss local gaps, while formal provers like Lean can bypass a local flaw by proving an overly broad target or validate an auto-formalized statement that drifts from the original intent. FaithSieve decomposes coarse proof steps into local reasoning units, extracts typed proof obligations, and checks them with a Lean-assisted formal evaluation agent, gating the formal evidence behind a semantic alignment score so it is used only when the formalization faithfully preserves the claim's context, objects, and logical form. On two new expert-verified first-error localization datasets, FaithSieve with a GPT-5.4 backbone reaches 81.43% exact first-error accuracy on the 350-problem ProofLoc-Olympiad set versus 72.29% for direct judging, and 84.5% versus 75.0% on the 200-problem ProofLoc-University set spanning six advanced domains.

ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving

Wenqian Ye, Ziwei Guan, Eric Xie, Bohan Liu, Shivani Modi, Buyun Zhang et al. Existing neural theorem provers either bake proof experience into weights through expensive updates or discard verified intermediate deductions once a problem ends, and they rely on sparse whole-proof feedback even when failed attempts contain useful partial results. ProofEvolve evolves explicit, formally verified proof structures: a neural model proposes variation operators such as decompositions, repairs, and schema recombinations, the Lean kernel checks every transition, and the system maintains partial AND-OR proof directed acyclic graphs (DAGs) in a behaviorally indexed archive within each problem while extracting kernel-checked sub-DAGs into a persistent schema library across problems. Typed schema recombination lets later proofs inherit solved results with every residual premise exposed as a new subgoal, preserving verified progress from incomplete attempts without weakening formal soundness. Across three competition-level Lean benchmarks, ProofEvolve achieves the highest average solve rate among the evaluated proof systems.

SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers

Haizhao Fan, Yuchi Xiong, Jize Wang, Xinping Guan, Xinyi Le Applying large language models (LLMs) in specialized domains often means reasoning over complex rules within concrete scenarios, but existing benchmarks either test only output-level instruction constraints or ignore the distinct roles rules play in such reasoning. RuleWeaver is a benchmark construction framework that starts from corpus-derived IF-THEN Meta Rules, progressively augments them into more complex rules, and composes them into rule-centered scenario question-answering instances, with process-level evaluation via rubric-based answer quality, rule recall, and rule precision in addition to final-answer correctness. Across 11 representative LLMs, the best model reaches only around 50% of the maximum rubric score, indicating that complex rule-centered scenario reasoning remains unsolved.

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li et al. Most mathematics benchmarks for large language models (LLMs) score only final answers, which says little about process-level failures or whether a model can act as a robust agent. The authors build a process-level benchmark that maps agentic problem-solving behaviors onto a structured taxonomy of reusable atomic mathematical capabilities, with planning, action, and feedback tasks in both textual and multimodal settings, generated by an automated pipeline that synthesizes trajectories and produces fine-grained annotations through controlled LLM rewriting. Models with similar end-to-end accuracy show markedly different agentic capability profiles, which the authors take as evidence that process-level evaluation is needed to interpret and guide the development of mathematical agents.

TTPO: Test-Time Policy Optimization

Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu et al. Post-training methods such as reinforcement learning (RL) and On-Policy Self-Distillation (OPSD) have advanced LLM math reasoning but need ground-truth labels, which rules out test-time training (TTT); majority-vote pseudo-labels are a natural substitute, but a wrong vote corrupts the teacher for every token. The authors observe an asymmetry: rollouts that disagree with the pseudo-label are usually wrong regardless of whether the vote itself is correct. Test-Time Policy Optimization (TTPO) exploits this by distilling agreeing rollouts via OPSD and penalizing disagreeing rollouts with grouped RL, with token-level selection that down-weights already-converged positions in distillation and penalizes only confident errors in RL. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% under test-time training, adds 25.2% to 36.4% in non-thinking mode, and generalizes across tasks.

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Yufan Wu, Yinghui He, Zhengyi Hu, Lang Wei, Ruichen Li, Qifan Yang et al. Inference-time scaling improves LLM reasoning but usually depends on repeated generation or external verification. CritICL starts from the observation that failure modes are structured and shared across model scales within a family, and turns failures from weaker models into guidance by injecting them as critique-based in-context examples. Two variants are proposed: CritICL-dynamic, which predicts input-specific failure modes and retrieves matching critiques, and CritICL-static, which applies a global failure-mode profile for stable guidance. It consistently outperforms standard in-context learning and matches or exceeds test-time scaling methods while needing far fewer generations and lower token cost.
5 more specialized papers

Robotics 9

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

Zengmao Wang, Wei Gao, Shuhan Shen World models let agents reason about future outcomes, but most approaches reconstruct future observations or features, adding complexity that does not help decision making. The authors propose a compatibility-prediction Latent World Model (LWM) for robot navigation that predicts action-conditioned latent feature compatibility instead, based on the insight that spatial proximity correlates with latent feature similarity, so action consequences can be scored directly in latent space; counterfactual training uses action sequences sampled across trajectories and learns which ones lead closer to the goal. The learned model can supervise policy learning from unlabeled video and then improve policies through reinforcement learning entirely inside the world model, eliminating action annotations and extra environment interaction, and on multiple real-world navigation datasets it significantly outperforms prior world-model and imitation-learning methods in prediction accuracy, policy learning, and real-robot navigation.

RTNav: Towards Real-Time Zero-Shot Object Navigation

Easop Lee, Lingyu Zhang, Boyuan Chen cross-listed Zero-shot object navigation with vision and language foundation models is usually developed in synchronous simulators where the world waits for the agent and inference time is free, so agents are designed as sequential perception-reasoning-action loops with no regard for wall-clock cost. When time counts toward the task budget, recent zero-shot methods degrade consistently. RTNav is a simple architecture that treats inference latency, asynchronous environment stepping, and bounded compute as explicit design constraints. On real-time variants of HM3D-v1, HM3D-v2, and HM3D-OVON, it improves success rate by up to 11% and Success weighted by Completion Time by up to 5.1 points over prior work.

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

Kechen Liu, Ola Shorinwa cross-listed Action-conditioned video world models are usually tied to a single robot embodiment, which locks them out of the vast corpus of human and multi-robot video that carries generalizable physical signal. CLAP unifies disparate action spaces through end-effector poses, language instructions, and latent actions, and uses a curriculum that first learns physical priors from unlabeled video via latent actions before grounding them in end-effector action spaces for zero-shot deployment. It approaches or surpasses state-of-the-art single-embodiment video models on DROID, with gains that compound under few-shot adaptation. The released suite covers end-effector, language, and latent conditioning across DROID, Bridge, bimanual YAM robots, and G1 humanoids, with all code and models open-sourced.
6 more specialized papers

Unclassified 1

On the Indistinguishability of Human v/s AI Generated Text

Jaee Ponde, Aritra Das, Mihir More, Debayan Gupta No summary available — see the abstract on arXiv.