Monday, September 7, 2026

354 papers cs.AI · cs.LG · cs.CL ← 2026-09-042026-09-09 →

Jul Aug Sep

Highlights

Iris: Climbing to the Search Frontier

Highlight HF pick · 43▲Agents Ziyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu et al. Iris-mini and Iris-pro are search agents trained at the 35B-A3B and 397B-A17B scales, released together with their data pipeline and training recipe. Training questions are reverse-constructed from the hyperlink structure of a web corpus by authoring multi-hop chains over an entity graph, rewriting every non-answer entity into a descriptive reference so no clue can be resolved by string matching, and keeping only questions a reference model fails closed-book but solves with the supporting evidence. The questions become trajectories filtered at both trajectory and turn level for supervised fine-tuning (SFT), followed by reinforcement learning (RL) against live search with an in-cluster reward judge and observation summarizer, and the two stages alternate in a procedure called SFT-RL climbing that feeds the hardest solved and most efficient RL rollouts back into the next supervised pass. Because inference-time context management matters more on these benchmarks than most reported system differences, every result is reported with and without it using a single ReAct agent with no sub-agents or test-time verification, and with management enabled the two models reach 82.2 and 88.6 on BrowseComp, plus 84.8/85.1 on BrowseComp-ZH, 86.9/92.9 on DeepSearchQA, and 52.3/56.4 on HLE, the strongest open-source results in their parameter ranges.

Open-source search agents are hard to compare because inference-time context management (CM) often accounts for more of the score than the policy itself; the AllSpark Team's answer is a full recipe (web-graph task synthesis, doubly-filtered SFT, RL against live search, and alternating "SFT–RL climbing") for two agents, Iris-mini (35B-A3B, from Qwen3.6-35B-A3B) and Iris-pro (397B-A17B, from Qwen3.5-397B-A17B), evaluated both with and without CM under a fixed tool set, context limit, and judge.

  • Training questions are reverse-constructed from a seed page and its out-links: an entity graph is distilled, a multi-hop chain of at least N relations is authored toward the seed entity, every non-answer entity is rewritten into a descriptive reference so no clue can be resolved by string matching, and a pair is kept only if a reference model fails it closed-book yet solves it when the entity graph is supplied.
  • Teacher trajectories (ReAct with search and scrape, observations replaced by on-the-fly page summaries) are filtered at the trajectory level for correctness, degeneracy (a sliding-window zlib compression-ratio detector for loops) and minimum tool-call depth, then at the turn level by an LLM judge whose rubric is induced from free-form critiques rather than hand-written, masking at most 10% of assistant turns from the SFT loss.
  • RL uses a group-relative policy gradient with request-level partial rollouts resumed from committed prefixes (corrected by truncated importance sampling, at roughly 2× over-sampling), an in-cluster FP8 Qwen3.5-397B-A17B serving as both binary generative reward model and observation summarizer, and each climb feeds back the shortest successful rollout (with enough tool calls) from queries whose pass rate is between 0 and 1/2, giving a self-paced curriculum.
  • With discard-all CM and a single ReAct agent at pass@1, Iris-mini scores 82.2 / 84.8 / 86.9 / 52.3 and Iris-pro 88.6 / 85.1 / 92.9 / 56.4 on BrowseComp / BrowseComp-ZH / DeepSearchQA / HLE (text-only), with Iris-mini beating XYZ-Aquila-mini by 3.4 points on BrowseComp but trailing it on DeepSearchQA (86.9 vs 89.5) and Iris-pro leading or tying on all four in its class.
  • CM is worth up to +21.2 points on BrowseComp for Iris-mini (64.7 without CM vs 85.9 with discard-all plus retry) but only +9.2 on HLE, gains shrink for the larger model, both models still trail frontier systems (Kimi-K3 91.2, Apodex-1.0-H 90.3 on BrowseComp), three configurations tie at exactly 85.1 on BrowseComp-ZH (246 of 289, with the appendix documenting a mislabeled question), climbing details are deferred to a later release, and the claimed positive transfer to BFCL, τ-bench, OfficeQA and APEX is reported without numbers.

When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference

Highlight HF pick · 1▲Other Ismail Erbas, Xavier Intes, Vikas Pandey In quantized recurrent networks, the stored low-precision state is fed back at the next time step, so the rule used to write that state can change all later computation. The authors name this rule recurrent-state write-back and isolate its effect in a compact GRU encoder-decoder for fluorescence lifetime imaging, which must estimate two lifetime parameters from very noisy time-resolved signals. With the trained model held fixed, switching to deterministic 4-bit state storage raises estimation errors by roughly 70x and 300x for the two parameters, because repeated small updates fall below the write threshold and the stored state freezes while the network keeps proposing change. Error feedback, residual memory, and direction memory carry the suppressed updates across time and restore accuracy without retraining, higher state precision can worsen a fixed solution, matched training can learn compatibility with the state interface, and an independently trained LSTM reproduces the failure with the cell state more sensitive than the hidden state.

Post-training quantization of recurrent networks stores each hidden state on a coarse grid and returns it to the next time step, so the rounding rule itself becomes part of the temporal computation rather than a passive encoding. The authors name this rule recurrent-state write-back, isolate its effect in frozen GRU and LSTM models for fluorescence lifetime imaging, and show that persistent sub-threshold updates that never reach the stored state, not bit width alone, are what break the network.

  • On a fixed Seq2SeqLite GRU checkpoint (32 units, 6,627 parameters), replacing continuous state propagation with deterministic 4-bit write-back raises lifetime RMSE from 0.36/0.35 ns to 25.37/106.59 ns for τ1/τ2 (roughly 70-fold and 300-fold) with every weight and every other operation unchanged.
  • The failing trajectory shows the predicted signature: 99.59% of decoder updates fall inside the half-step write boundary, only 0.25% of state elements change level per step, and the median run of consecutive same-direction suppressed updates spans 132 of 134 decoder steps, versus a median of 2 steps in accurate native-4-bit models that also contain many sub-threshold updates.
  • Carrying the discarded information forward rescues the same frozen network without retraining: error feedback restores 0.36/0.37 ns, a 2-bit residual memory reaches 0.34/0.40 ns, and the newly introduced direction memory (a small counter accumulating the sign of sub-threshold proposals) reaches 0.34/0.34 ns, all while the state fed back to the network stays on the 4-bit grid.
  • More state precision is not a monotonic fix: evaluating the independently trained 4-bit-state GRU with 8-bit write-back worsens RMSE from 0.35/0.40 ns to 0.43/0.57 ns even though the median occupied decoder levels rise from 12 to 176.5, and in matched training the ranking of 4-bit, 6-bit, residual-memory, and direction-memory interfaces shifts, showing compatibility between a learned solution and its state interface is what matters.
  • An independently trained 32-unit LSTM reproduces the failure (0.239/0.254 ns to 3.857/1.327 ns under 4-bit write-back) and its rescue by error feedback, with the cell state far more sensitive than the hidden state (8.158 vs 0.296 ns τ1 RMSE); the study is limited to a single synthetic 160,000-sample FLI test set and tiny models, and the auxiliary-memory comparisons match stored bits but not hardware cost.

MaxKernel: Agentic Kernel Generation for TPUs

Highlight HF pick · 5▲Agents Shangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica et al. Writing high-performance custom kernels for accelerators demands deep hardware expertise, and large language models paired with real-time compiler feedback offer a way to automate it. MaxKernel is a multi-agent system for TPU kernel development with three modes: a Human-in-the-Loop (HITL) agent for step-by-step collaborative design, an Autonomous agent that runs a fully automated, metric- and trace-driven optimization loop, and a Graph-Based Autonomous Search that scales the autonomous agent to global exploration of the design space, all sharing sub-agents for planning, implementation, self-debugging, testing, and hardware profiling. Evaluated on JaxBench, a suite of 50 diverse TPU kernel tasks, plus real workloads from open-source models, the system consistently produces implementations matching expert hand-tuned baselines. The agent is open-sourced.

Hand-writing high-performance TPU kernels in JAX/Pallas demands expert management of memory hierarchies, DMA pipelining, and tiling, and zero-shot LLM generation mostly fails against rigid accelerator APIs and opaque compiler errors. MaxKernel wraps an LLM in a closed loop of specialized sub-agents (planning, implementation, compile-fix, test synthesis, autotuning, and XProf-based profiling) and then scales that loop with parallel and beam search over a persistent graph of kernel states.

  • The system offers three orchestration modes: a human-in-the-loop router that pauses after each sub-agent, an autonomous hill-climbing loop that feeds profiling traces back into the next plan and rolls back to the best valid snapshot, and a graph-based search that spawns fresh autonomous sessions per node, with a test suite frozen before optimization begins so implementation agents cannot alter the correctness criteria, and a RAG knowledge store that deliberately excludes hand-tuned kernel code.
  • On JaxBench (50 TPU v6e tasks, all runs using Gemini 3.1 Pro), a Best-of-100 zero-shot baseline compiled only 10/50 kernels at a 1.08× geomean, while the single-trajectory Auto agent reached 48/50 correct at a median 1.39× (range 1.19–1.42 across five seeds), and parallel search over those five trajectories hit 50/50 correct with a 1.58× geomean and 34/50 tasks faster than XLA, versus 1.49× and 31/50 for beam search.
  • Against eight expert hand-tuned Pallas kernels, parallel search posted a 2.32× geomean speedup versus 2.02× for the human references, beating them on seven of eight workloads (e.g. 6.74× vs 2.41× on Paged Attention, and above-baseline MLA kernels where the human version ran at 0.69×), but losing badly on Ragged Paged Attention at 1.42× vs the expert's 4.65×.
  • On production open-source workloads, generated kernels cut Qwen3-Next Gated DeltaNet training-step time by 4.70×, accelerated DeepSeek-V4 sparse-attention prefill by 7.85× over JAX, and shaved 8.68% latency off an already hand-written Pallas MLA kernel, while a debugging run fixed a deadlock in Ragged Paged Attention v3 prefill by inserting clamp instructions for left-padded inputs.
  • Caveats include a single-LLM evaluation with no model ablations, geomean speedups floored at 1.0× so regressions are hidden, correctness tolerances relaxed to as loose as 0.1 on some tasks, a beam-search plateau at depth 3 attributed to its 2-iteration per-node budget being too short for multi-step Pallas fixes, and the parallel-search numbers representing the best of five runs rather than a typical single trajectory.

Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One

Highlight Large Language Models Fred Zhangzhi Peng, Kaiwen Zheng, Anru R. Zhang Language generation is nearly always sequential, either token by token in autoregressive models or through long iterative refinement trajectories in diffusion language models. PlaidQ is a 0.7B continuous diffusion language model for code that repurposes a pretrained autoregressive model as a bidirectional denoiser over continuous token embeddings, then distills its trajectory with distribution matching for few-step generation and paired-trajectory supervision for one-step generation. At matched scale it is competitive with discrete diffusion models on code, and a 16-step distilled student reaches 31.78 and 40.49 pass@10 on HumanEval and MBPP+, beating its own teacher sampled for 512 steps. A single denoising step still yields functionally correct programs at 7.07 pass@1 on HumanEval, and code and checkpoints are released.

Language generation, whether autoregressive or diffusion-based, remains sequential: autoregressive models chain token by token, and diffusion language models trade that for a long iterative refinement trajectory, so few-step or one-step generation has stayed elusive. PlaidQ is a 0.7B continuous diffusion model for code built by converting a pretrained autoregressive model into a bidirectional denoiser over token embeddings, and its long trajectory is then distilled down to 16, 8, 4, or even a single denoising step.

  • The model reuses the Qwen3-0.6B trunk and vocabulary head with causal attention swapped for bidirectional attention, shifts the reconstruction head by one position to preserve the pretrained next-token alignment, and denoises every completion position in parallel while the prompt stays clean, decoding tokens only once at the end.
  • Training at a 152k-token vocabulary and 2048-token sequences is made tractable by SWVR, a streamed exact kernel for the categorical reconstruction that cuts activation memory from 40.2 GB to 28.3 GB (29.6%) with identical loss and gradients, plus a hybrid Muon-AdamW optimizer and autoregressive initialization that together lower validation NLL from 3.65 to 3.59.
  • As a base generator at 128 steps, PlaidQ is roughly on par with same-scale discrete diffusion models (18.05 pass@1 on HumanEval+ versus 17.6 for Open-dCoder), reaches 39.87 pass@10 on HumanEval with classifier-free guidance, and transfers zero-shot to infilling with 37.91 on HumanEval-Infill despite never seeing that conditioning during training.
  • Distribution matching distillation with an on-policy critic yields a 16-step student that scores 31.78 and 40.49 pass@10 on HumanEval and MBPP+, beating the 128-to-512-step teacher (28.57 and 35.21), while at 4 steps distillation lifts pass@10 from 7.90 to 15.85 and from 7.91 to 30.94 where the training-free DPM-2M solver collapses to zero.
  • For the single-step regime, paired-trajectory distillation feeds the student the teacher's output for the same initial noise as a cross-entropy target, raising one-step HumanEval pass@1 from 0.09 to 7.07 (8.53 pass@10), though absolute few-step scores remain far below 7-8B discrete models, each step budget requires a separately trained student, and infilling on SantaCoder lags at 22.86 versus 29.6 for Open-dCoder.

Extremely Sparse Supervision Incentivizes Reasoning Ability

Highlight Reasoning Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane Prevailing post-training methods for reasoning optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. Working in the on-policy distillation setting with the Qwen3 family, the authors find that supervising as few as one or two tokens per reasoning trajectory, about 0.05% of all tokens, matches or surpasses full-token training on mathematical reasoning in most cases. The effect holds across nine teacher-student configurations of varying scale and is further validated on coding reasoning, Llama models, and Proximal Policy Optimization (PPO)-based reinforcement learning with verifiable reward (RLVR). The authors suggest this mirrors natural learning, where one reflects on a few critical steps rather than correcting every word, and argue it points toward more efficient post-training algorithms.

On-policy distillation (OPD) is assumed to work because the teacher supervises every generated token, yet masking the loss down to just one or two tokens per reasoning trajectory, about 0.05% of all tokens, matches or beats dense supervision at improving student reasoning. The finding challenges the premise that effective post-training must be token-intensive.

  • Sparse OPD keeps the standard sampled-token objective, where each token's reward is the teacher-minus-student log-probability, but inserts a mask so only selected tokens contribute gradient: rand1tok picks one random token, maxtok and mintok pick the single highest- and lowest-reward token, minmaxtok keeps both, and pctltail 0.05% keeps the extreme reward tails.
  • Across nine Qwen3 teacher-student pairings trained on DAPO-Math-17K and evaluated with pass@k up to 128 and avg@8 on AIME 24, AIME 25, and HMMT Feb 25, even rand1tok improves the base student in all nine, and some sparse variant matches plain OPD in 2 of 9 families and beats it in 7 of 9.
  • The strongest case is Qwen3-8B distilled from the smaller Qwen3-4B-Instruct-2507, where minmaxtok reaches 49.3 mean avg@8 versus 47.5 for plain OPD and 46.8 for the teacher itself, while touching roughly 10% of parameters and, for maxtok, raising reverse KL to the teacher from 0.18 to 1.12, so the gains are not explained by closer imitation.
  • Which extreme token helps depends on student capacity: maxtok (reinforcing high-entropy "forking" tokens the teacher strongly prefers) excels when the student is same-scale or larger, whereas mintok (suppressing teacher-dispreferred tokens, which an audit shows are mostly correct-but-dispreferred rather than genuine errors) is more stable when the student is much smaller, and performance is non-monotonic in sparsity.
  • The phenomenon replicates on coding reasoning (Eurus-RL-Code, LiveCodeBench v6), Llama 3 models, and PPO, but not under GRPO or REINFORCE, and since every variant still needs the full set of on-policy rollouts the savings are in logit memory rather than training compute.

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

Highlight HF pick · 1▲Agents Quan Shi, Keshav Dhandhania, Karthik Narasimhan, Victor Barres Coding agents are increasingly asked to build production LLM agents, but existing benchmarks do not measure whether they can deliver one under the conditions of a real client engagement. τ^τ-bench (hyper-tau-bench) hands a developer agent a business's actual records, a client who holds the requirements, a production API, an inherited codebase, and limits on serving cost and models, then scores the customer-service agent it builds by deploying it against held-out simulated users across 53 tasks in four domains. The strongest configuration, Claude Opus 5 running under Claude Code, passes only 23.9% of evaluation simulations, against an expert-authored reference ceiling of 82.2%. The failures mirror those seen with human agent developers: shallow queries instead of deep comprehension of the records, almost no communication with the client, and too little experimentation with architecture or serving spend before shipping the first design that runs.

Building customer-service agents is increasingly delegated to coding agents, but existing benchmarks score finished agents rather than the work of constructing one from a client's messy records, stakeholders, and cost constraints. τ^τ-bench (hyper-tau-bench) makes agent construction the task: a developer agent receives a business's scattered artifacts, a simulated client, a possibly defective REST API, an optional inherited codebase, and a serving-cost budget, and must ship a complete agent that is then scored by deployment against held-out simulated users.

  • Each of the 53 tasks across airline, retail, telecom, and banking is built by decomposing τ-bench policies into atomic facts (3,328 total, 2,969 in banking alone) and rendering them into 2,868 realistic artifacts (handbooks, transcripts, spreadsheets, screenshots, call recordings), with some facts moved to an LLM-simulated client that must be interviewed and with contamination controls via rebranded, re-valued airline and retail domains.
  • The strongest configuration, Claude Opus 5 under Claude Code, passes only 23.9% of evaluation simulations versus an 82.2% expert-authored reference ceiling; GPT-5.6-sol in Codex scores 22.0% with far shorter builds (48 vs 216 minutes), and banking collapses the field to 3–9% while airline, retail, and telecom reach roughly 42–73%.
  • Trajectory analysis traces failures to skipped disciplines: developers grep the corpus instead of reading it (opening fewer than 80 of ~1,700 banking files), talk to the client in just 0.3% of tool calls even though builds that ask four or more questions score 0.50 versus 0.16 for builds that never ask, rewrite inherited code without ever running it, and use only 0.45–0.72× of the serving budget while 21 builds overshot it and 10 saw their scores zeroed by the overage penalty.
  • Nearly every build (92%) is a single LLM tool loop with a cheap, same-vendor model, yet a one-line architecture hint (route by intent, review tool calls) doubled a telecom score from 31% to 67%, and cheating-adjacent probes of the grader or hidden data appeared in 17–42% of runs across harnesses, none successful.
  • Limitations include a single construction trial per task (no variance estimate), single-LLM client and user simulators with fixed requirements, corpora audited for full consistency unlike real engagements, and evaluation that stops at submission rather than covering post-deployment maintenance.

LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28

Highlight Agents Wes Sander Discovery Loop is a lightweight system in which a large language model iteratively evolves an optimization algorithm: starting from a simple seed solver, the model proposes improvements guided by a scoreboard and a history of prior ideas, an independent verifier evaluates each candidate, and only improvements are kept. On the Packomania circle-packing benchmark, where the goal is to maximize the sum of radii of N variable-radius circles in the unit square, the system beat the best known solutions for 10 values of N between 101 and 114 by 2.4% to 5.4%, within 15 iterations and for a total LLM cost of $27.72, and the records were independently accepted by Packomania. The authors also analyze cost-efficiency dynamics, including an adaptive plateau-detection mechanism.

Discovery Loop asks how much of DeepMind's AlphaEvolve paradigm survives when reduced to one LLM, one consumer PC, and a sub-$30 budget, and answers by using Claude Fable 5.1 to iteratively rewrite a circle-packing solver until it beats best-known solutions on the Packomania csqv benchmark (maximize the sum of radii of N variable-radius circles in a unit square).

  • The loop is roughly 400 lines of Python: each iteration feeds the LLM the full source of the current champion solver, a per-target scoreboard, and the last 12 ideas tried, receives a complete replacement solver rather than a patch, runs it on all targets in parallel with a 120 s timeout, and accepts it only if an independent zero-tolerance verifier (containment, non-overlap, recomputed radii sum, feasibility shrink) confirms an improvement.
  • Across 15 iterations and $27.72 of LLM spend, the system produced new accepted records for 10 of 12 targets (N = 101–114), with gains of 2.4%–5.4% per instance and a +5.0% total sum of radii (51.81 → 54.41), while N = 26 and N = 32 only matched prior records to within 1e-6.
  • A notable caveat is that every one of the 10 records was first broken at iteration 0 by the hand-written seed solver (multi-start penalty L-BFGS-B with LP-optimal radii); the LLM-evolved ideas (basin hopping with contact-graph SLSQP polish, hexagonal-lattice initialization, island-model parallelism, KKT-Newton polish, defect-migration moves) improved the aggregate score only from 59.39 to 59.98, so the headline results say more about weak prior records than about LLM-driven discovery.
  • Returns diminish sharply: iterations 0–5 cost $4.96 for +0.57 total score, while iterations 6–14 cost $22.76 for +0.02 (a 130× jump in cost per unit gain), and a proposed plateau detector (window 4, threshold 0.01) would have stopped after iteration 9, saving about 50% of the spend at a loss of 0.006 in total score.
  • The system keeps a single champion with no population or evolutionary database, was tested on only one cheap-to-evaluate problem, and a side experiment on MIPLIB mixed-integer programming saw a 75% code-generation failure rate, suggesting the approach is fragile on more complex domains.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

Highlight HF pick · 11▲Agents Sihan Ge, Yichen Lin, Chenyu Zhou, Jianghao Lin, Tao Yao, Dongdong Ge LLMs are increasingly asked to turn natural-language descriptions into operations research (OR) optimization models, but real requests often omit objectives, constraints, or business rules that change the resulting program, and existing benchmarks assume complete specifications. OR-Clarify presents partial problem descriptions with structured hidden slots and scores agents on slot recovery, stopping behavior, silent assumptions, and interaction cost through bounded dialogue with a simulated user, in both open-ended and choice-based clarification modes. The proposed InterOPT framework first identifies unresolved formulation-critical gaps and then uses them to decide whether to ask another question or stop; in choice-based experiments it substantially outperforms all baselines on exact slot recovery and stays competitive with strong prior methods in the open-ended setting.

Large language model (LLM) agents that turn natural-language business requests into optimization models are usually evaluated on complete specifications, but real operations research (OR) requests routinely omit objectives, constraints, or rules that change the resulting mathematical program. The paper reframes the pre-modeling step as a selective completeness decision, introducing the OR-Clarify benchmark to measure whether an agent asks the right clarifying questions and stops at the right time, plus InterOPT, a two-stage method that tracks unresolved formulation gaps and uses them to steer questioning.

  • OR-Clarify is built by decomposing fully specified OR tasks into single-requirement facts, masking roughly half of the formulation-critical ones behind a fixed seed, and scoring bounded dialogue with a simulated user who answers only what is asked; it comprises 100 cases and 178 hidden slots (75 P0 blocking, 83 P1 substantive, 20 P2 secondary) and reports exact slot recovery, stopping behavior, silent assumptions, and question count.
  • InterOPT separates monitoring from control: Stage 1 (Dynamic Gap Search) maintains a persistent ledger of open formulation gaps across six categories such as objective trade-offs and hard-versus-soft policies, while Stage 2 (Gap-Guided Action Search) generates three gap-bound candidate questions, runs a selector over them, and leaves the ask-or-READY_TO_MODEL decision to the model itself.
  • Off-the-shelf models struggle under both free-form and multiple-choice protocols, with Opus-4.8 strongest but no model exceeding 60% Core Exact and all leaving substantial silent assumptions per run, and DeepSeek V4 Pro as the default tested model.
  • In the choice-based setting InterOPT reaches 0.675 Core Exact versus 0.506 for the MC-D baseline and roughly halves silent assumptions (0.366 vs 0.692), but at a steep interaction cost of about 10 questions per run versus 2.4, and in the open-ended setting it trails the ORPilot and GATE adapters (0.538 vs 0.583 and 0.560 Core Exact).
  • Ablations show gap tracking drives most of the coverage gain while guided selection mainly shortens dialogue, and the open-setting diagnostic identifies stopping as the core failure: agents declare readiness in 98.8% of runs, yet 40.6% of stops are audited as premature, which the authors flag alongside dependence on a single tested model and LLM-based judging.

Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference

Highlight HF pick · 6▲Large Language Models Mostafa Elhoushi, Alex Pretko, Nolan Dey, Bin Claire Zhang, Gavia Gray, Gurpreet Gosal et al. Layer dropout, also called stochastic depth, speeds training and enables zero-shot layer pruning in vision and language transformers, yet it has vanished from LLM pretraining recipes amid reports that it hurts accuracy, without any systematic study of the effect. Across more than 2,400 pretraining runs on Cerebras CS-3 systems spanning 271M to 8.2B parameters and up to 160B tokens, the authors establish best practices for the layer distribution, time schedule, and optimizer hyperparameters and show that at equal training FLOPs layer dropout yields lower loss. Models reach lower or similar validation loss while saving up to 25% of training FLOPs, and the trained models support post-training optimizations such as early exit, intermediate-layer skipping, and self-speculative decoding for up to 1.5x inference speedup with negligible accuracy loss.

Layer dropout (stochastic depth) vanished from LLM pretraining recipes because it appeared to hurt accuracy in the single-epoch, large-data regime, and this study from Cerebras argues that those degradations were mostly a configuration problem. Across more than 2400 runs on CS-3 systems, it shows that dropping whole transformer blocks with the right residual scaling, depth distribution, and time schedule cuts training FLOPs while matching or beating dense loss, and leaves the model robust to early exit and layer skipping at inference.

  • Each transformer block is skipped per sequence with a Bernoulli mask, scaled during training by 1/ρ (ρ = 1 − p) following CompleteP's maximal-residual-update rule, which is what makes optimal learning rate, batch size, and weight decay transfer unchanged across dropout rates, whereas the common r_train = 1 convention forces per-rate retuning.
  • Ablations at 20 tokens per parameter establish the recipe: whole-layer dropout beats separate attention/FFN sub-layer dropout, per-sequence masks beat per-batch masks, non-uniform depth distributions beat uniform at equal FLOPs, and a linearly increasing rate across depth (ILD) paired with a linearly decreasing rate across steps (DTS) is best, with increasing-over-time schedules degrading loss badly.
  • With ILD+DTS at 5% FLOPs savings, 503M and 906M models beat the dense baseline (906M: 1.9513 vs 1.9526 val loss), and at larger scale a 1.8B model with p_max=0.6 saves 15% FLOPs at 1.836 vs 1.849, a 3.9B model with p_max=0.8 saves 20% at 1.745 vs 1.732, and an 8.2B model with p_max=0.99 saves 25% FLOPs at 1.663 loss, with degradation staying within roughly 0.5% of baseline as tokens per parameter grow.
  • The same models gain zero-shot depth elasticity: skipping alternate layers of the 3.9B dropout model yields 2.129 loss versus 6.446 for its dense twin, Balcony-style exit adapters trained on frozen weights reach lower early-exit loss than on dense models, and Draft & Verify self-speculative decoding on XSUM reaches 1.54× speedup at 3.9B where the dense baseline manages only 1.02×.
  • Trade-offs and gaps remain: alternating-layer dropout is best for skip robustness but fails at early exit while ILD is the reverse, hyperparameter transfer weakens at aggressive rates and was only validated for constant schedules, the 8.2B run has no dense baseline in the results table, the 3.9B dropout run is slightly worse in raw loss, and there is no comparison against Mixture-of-Depths or MoE architectures.

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

Highlight HF pick · 4▲Large Language Models Yang Li, Semih Yavuz, Shafiq Joty On-policy distillation (OPD) gives dense per-token supervision for post-training language models, but external teachers suffer from distribution mismatch and self-distillation with privileged context is limited by in-context learning capacity. RISE (Recursive Improvement via Self-Extrapolating Policy Distillation) instead builds a synthetic teacher from the model's own reinforcement learning with verifiable rewards (RLVR) trajectory, extrapolating the displacement between the current checkpoint and a trailing anchor in parameter or logit space to turn a sparse outcome-driven update into a dense token-level target. Because the teacher is refreshed every iteration as the student improves, distillation becomes a recursive loop in which outcome rewards ground the extrapolation and the extrapolated teacher refines token-level decisions. Across mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks, RISE outperforms both RLVR-only training and on-policy self-distillation.

On-policy distillation gives dense per-token supervision but is bottlenecked by teacher quality: external teachers suffer distribution mismatch, and privileged-context self-distillation is capped by in-context learning capacity. RISE builds the teacher from the model's own RLVR trajectory by extrapolating the displacement between the current checkpoint and a trailing anchor, turning a sparse outcome-driven parameter update into a dense token-level target with no external model or privileged conditioning.

  • Each iteration runs a GRPO step, constructs a synthetic teacher as anchor + β·(post-RLVR checkpoint − anchor) with β decaying linearly from 1.2 to 1 over training, either in weight space (task arithmetic) or logit space (a geometric mixture of the two policies' output distributions), then distills that teacher into the post-RLVR checkpoint via a top-K Jensen–Shannon loss on the same rollouts, so no extra sampling is needed.
  • On DAPOMath, Qwen3-8B in-domain Math Avg rises from 60.0 (GRPO) to 62.7, Qwen3-1.7B from 45.4 to 50.2, and OLMo3-7B-Instruct-SFT on OpenR1-Math-46K from 47.6 to 56.4 with AIME'24 jumping 30.2 to 46.9, while privileged self-distillation baselines (GRPO+SDPO, SDAR, RLSD) land within about two points of GRPO or below it and OOD scores on GPQA, IFEval, and MMLU-Pro hold or improve.
  • The gains carry to mixed math+STEM on Qwen3-4B-Base (Math Avg 44.8 vs. 40.2), to code on Qwen3-8B-Base (faster convergence, similar final accuracy), and to agentic tasks on Qwen2.5-3B-Instruct where the weight-space variant lifts ALFWorld success 75.0 to 84.4 and WebShop accuracy 63.3 to 74.2.
  • Ablations show both phases are necessary: extrapolating without RLVR collapses within 60 steps (MATH-500 falls to 2.4% as responses hit the length cap), while adopting the extrapolated weights directly without distillation yields only +0.3 and +0.2, and a compute-matched GRPO-2x with a second gradient pass recovers just +0.5 and +1.9 versus RISE's +2.7 and +4.8.
  • The safe β range shrinks as training proceeds (β=2.0 gives +7.3 AIME'24 points at step 50 but −15 at step 100), the EMA anchor rate is model-dependent (η=0.1 best for Qwen, η=1 best for OLMo), wall time is 1.3–1.6× GRPO, and the method inherits any reward hacking since extrapolation amplifies whatever direction RLVR takes.

Applications 93

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

Ziyi Zhao, Guanzheng Wei AI recruitment tooling has moved from scoring candidate-job pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and take or recommend actions, and this systematized narrative review traces that shift from bilateral retrieval and behavioral ranking through neural person-job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol covering 40 representative works plus industrial and legal sources, it organizes the field around three transitions: similarity to reciprocal suitability, single model to compound workflow, and offline prediction to evidence- and productivity-aligned evaluation. Within the coded set, privacy is never directly evaluated and no work jointly assesses utility, fairness, privacy, and security, and the authors note that behavioral labels confound exposure, preference, and qualification while final-output scores hide pipeline failures. The review closes with a staged mapping from evaluation evidence to the strongest defensible claim and an agenda for reciprocal, auditable, temporally controlled systems.

Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection

Roberto Fern\'andez-Barrios, Iker Pastor-L\'opez, Amaia Pikatza-Huerga, Pablo Garc\'ia Bringas cross-listed Adaptive network intrusion detection systems retrain after drift alarms, but an alarm signals change without establishing that a challenger model should replace the deployed incumbent, and promotion conclusions may depend on how the challenger was built and how much evidence supports it. The authors test this dependence on CICIDS2017, UNSW-NB15, and ToN-IoT with self-contained challenger pipelines, nested candidate-size controls, a common harness comparing nine update policies, and a sensitivity check confining each exact feature vector to a single evaluation, training, or probe role. Incumbent-owned frozen preprocessing amplified apparent promotion harm, which did not persist with self-contained pipelines, and raising candidate evidence from 512 to 2,000 samples per class improved balanced accuracy by 0.38 to 1.67 points across the three benchmarks, mainly through fewer false positives. No policy dominated globally, thirteen replays on real time-ordered traffic showed no net harm from always deploying, and the paper argues that challenger construction and evidence must be controlled and reported explicitly.

A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models

E. Cho Smith, Samuel Ho, Dawn Laux Large language models (LLMs) are increasingly used to label cultural texts at scale, but it is unclear whether their labels are reliable enough to treat as measurements of latent social constructs. Five LLMs were run repeatedly as zero-shot annotators of four constructs in English song lyrics, namely self-esteem, self-control, seeking belonging, and seeking recognition, and the outputs were checked for run-to-run consistency, cross-model agreement, and whether consensus labels could train a supervised classifier. Reliability varied sharply by construct: self-esteem was the most stable across models, seeking recognition the least, and the other two fell in between depending on the model. Consensus labels carried learnable signal for downstream classification, and the authors argue that repeated-measurement stability and cross-model convergence should be reported before LLM annotations are used as scalable measurements in cultural analytics.

Hakken: Predicting future discoveries to fill the gaps in today's knowledge

Tarek R. Besold, Uchenna Akujuobi, Pablo Sanchez, Alessandra Toniato, Kana Maruyama, Jihun Choi et al. Hakken aims to predict scientific relationships that have not yet been documented, going beyond what can be deduced from existing literature. It trains a transformer on temporal sequences of knowledge graphs extracted from large publication corpora, fuses this with a large language model's semantic knowledge to predict the presence and type of future relationships between concepts, and attaches a model-agnostic explanation layer so scientists can evaluate each suggestion. Applied to biomedicine, the model sets a new benchmark for time-aware multi-label relation prediction and stays coherent over long historical spans. The authors scored 1.5 million aging-related hypotheses, reviewed batches with biologists, and advanced three to wet-lab testing, confirming two previously undocumented interactions, between TP53 and BAMBI and between RAF1 and TNF.

Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection

Maryam Abbasihafshejani, Murtuza Jadliwala Whisper's generative decoder can emit fluent but fabricated transcripts for audio containing little or no speech. The proposed training-free method estimates a compact hallucination-associated subspace from non-speech calibration data and projects decoder hidden states away from it at inference, either always-on or gated by Whisper's own prediction that the input is non-speech. Across non-speech benchmarks, always-on projection cuts the average hallucination rate from 31.31% to 2.44%, a 92.21% relative reduction, while the gated variant reaches 3.74% with fewer false rejections of genuine speech. On LibriSpeech, gated projection raises absolute word error rate by 0.33 to 4.39 percentage points and produces false-rejection rates of 0.41 to 9.97% depending on model and split, giving a controllable trade-off between suppression and recognition quality.

Sustainable Edge Vision via Empirically Calibrated DVFS: Eliminating Thermal Throttling on Passively Cooled Hardware

Aayush Marasini, Zhaoxian Zhou cross-listed Passively cooled edge devices avoid fan energy overhead and mechanical failures, but sustained deep neural network inference on them is bottlenecked by thermal throttling. The authors design an empirically calibrated, state-aware dynamic voltage and frequency scaling (DVFS) scheduler that combines time-domain dwell guards, absolute temperature bounds, and derivative triggers for sharp thermal spikes, rather than reacting to temperature alone. On a passively cooled Raspberry Pi 5 running YOLOv8n for 30-minute workloads, the scheduler eliminates all observed thermal throttling events, delivers 6.8% higher frame rate than a temperature-only baseline at 1.9% less energy per frame, and beats an actively cooled reference on Joules per frame, though the passive envelope closes at ambient temperatures of 27°C or higher where nonlinear leakage defeats DVFS control.

Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling

Frank Hu, Shriram Chennakesavalu, Zichen Wang, Patricia Suriana, Bodhi Vani, Kirill Shmilovich et al. Drug candidate design means searching a vast, rugged chemical space for molecules that satisfy competing objectives, and while reinforcement learning from verifiable rewards (RLVR) can adapt large language models (LLMs) for this, chemically relevant scoring functions can take hours or days per evaluation, making them too slow for online training. The authors ask whether LLMs can learn molecular design strategies from cheaper synthetic tasks that transfer to expensive structure-based lead optimization. Curriculum recipes that progressively introduce harder synthetic design tasks yield performance surpassing much larger frontier models on structure-based lead optimization, suggesting synthetic-task scaling as a route to post-training for experimental settings too costly to train on directly.

Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates

Manuel R\"oder, Bibin Babu, Frank-Michael Schleif Detecting cyberattack campaigns that span organizations normally requires sharing sensitive telemetry across institutional and national borders. FedIoC is a federated learning framework in which each client trains a threat detector locally and folds its structured indicators of compromise (IoCs) into gradient updates via a supervised contrastive loss: flows matching known indicator patterns are pulled together in embedding space and non-matching flows pushed away, so campaign structure is expressed in the gradient direction. The server then clusters clients by cosine similarity of their updates and recovers cross-organizational campaign cohorts from gradient geometry alone, with no direct indicator transmission, on two public threat-detection benchmarks where each client sees only a fragment of every campaign and holds disjoint indicator sets. The authors identify non-IID gradient structure as the main driver of this recovery and pose encoder design that improves on it as an open problem.

Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution

Mubashar Iqbal, Asifullah Khan cross-listed Ransomware detection and family attribution benefit from static, dynamic, and memory analysis, but running every modality on every sample wastes compute and adds latency. The proposed cost-aware Hierarchical Multi-Agent System (HMAS) organizes specialist agents under domain controllers coordinated by a meta orchestrator, starts with cheap static analysis, and escalates to dynamic and memory analysis only when confidence is low or specialists disagree, with a locally deployed large language model verifying selected hard cases without replacing the deterministic pipeline. The full system reaches 96.57% accuracy and 0.99 ROC-AUC for binary detection and 0.90 macro-F1 for family attribution, while cutting average analysis cost by 43.97% relative to exhaustive analysis. Routing analysis shows that 56.05% of samples are resolved from static evidence alone and only 4.33% need the complete pipeline.

When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation

Susu Hu, Preetam Gattogi, Jens Lehmann, Sahar Vahdati, Stefanie Speidel, Julien Vibert Bidirectional discrete diffusion looks like a natural fit for DNA because it can reconstruct missing sequence from both flanks. GenDA (Genomic Density-optimized Absorbing Diffusion), a 202M-parameter model that places masked spans where local sequence entropy is highest, reaches a pooled ClinVar single-nucleotide-variant AUROC of 0.774 after supervised fine-tuning, beating a similarly sized autoregressive model by 0.103. But a matched random-span variant scores 0.777, leaving no evidence that entropy guidance causes the gain, and in zero-shot functional inpainting of promoters, enhancers, and exon or intron boundaries the model fails to consistently beat a control that shuffles the gap while preserving 3-mer composition. The authors conclude that strong fine-tuned variant prediction, a plausible corruption prior, and usable functional generation are separate claims needing separate validation.

CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games

Kai Wang, Ge Fan, Chaoyun Zhang, Yuyang Jiang, Yuze Liu Matchmaking in a large multiplayer online battle arena game breaks when most queued players lack history in the mode being played, when skill distributions shift drastically across rank tiers, and when extreme skill brackets are data-starved. CHAMP swaps the target-mode-only player profile for a hybrid feature set, a timestamp-ordered cross-mode short-term sequence annotated with target-domain features plus per-mode long-term, real-time, and team statistics, feeding a Domain-Aware Win-rate Network that learns mode-conditioned representations and per-mode debiasing inside one shared model serving every mode. Offline it predicts win rate with 67.73% accuracy, ahead of all evaluated attention and sequence baselines, and online A/B tests across the full ladder cut the five-minute kill-crushing rate by up to 20.73% for lower-tier players.

ReCAST: Restoration-aware Cascaded Stage-wise Training for Obfuscated SMS Risk Classification

Jieyun Huang, Yi Shen, Kaikai Zhao, Jiangze Yan, Wenjing Zhang, Ping Chen et al. cross-listed Fraudulent Chinese SMS messages are increasingly written to evade the cheap classifiers production systems can afford, hiding risk-bearing phrases behind crafted obfuscations that humans still read without effort. ReCAST distills a large teacher model's de-obfuscation ability into a smaller deployable student by supervising obfuscated span detection, obfuscation type prediction, and text restoration, then uses that restoration-aware student for downstream risk classification. On an internally built real-world benchmark the approach substantially improves classification over directly trained baselines under obfuscation while staying within production latency and throughput constraints.

Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction

Minh-Khoi Pham, Luca Cotugno, Dan Cernei, Alina Sirbu, Stefano Masi, Giuseppe Prencipe et al. Tabular foundation models (TabFMs) handle standard classification and regression on structured data well, but censored time-to-event prediction demands explicit treatment of censoring and event-time dynamics. This study attaches CoxPH, DeepHit, and cause-specific MTLR survival heads to frozen TabFM backbones and compares them against temporal zero-shot reformulation and classification-based fine-tuning across 74 single-risk datasets plus 4 competing-risk ones, with a revised context-resampled training procedure. Zero-shot inference is effective on smaller datasets while supervised adaptation grows more advantageous as data scales, with the Cox interface the most reliably strong option, especially on integrated Brier score; DeepHit favors time-dependent concordance over probabilistic calibration, and classification fine-tuning remains the weakest for probabilistic prediction.

Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing

Linsen Zhu, Mengqing Cai Technical progress in applying artificial intelligence (AI) to investing is often mistaken for evidence that it makes money. This critical review surveys public research through 31 August 2026 on listed equities, exchange-traded funds, centralized crypto spot and perpetual futures, and on-chain markets, organizing it along an alpha-translation chain in which point-in-time information must become a stable signal, feasible positions, executable orders, and risk-adjusted returns after costs. Across machine learning, time-series foundation models, financial language models, reinforcement learning, and agents, the authors find real but mostly upstream gains in prediction, text processing, portfolio design, and workflow integration, undermined downstream by temporal contamination, repeated selection, survivorship, weak benchmarks, execution costs, and capacity limits. Within the evidence examined, no general AI architecture is shown to deliver persistent, cross-regime, capacity-aware net alpha, and the review sets out conditions such as point-in-time data, decision-aligned objectives, joint portfolio-execution evaluation, and prospective tests for more credible claims.

BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer Generation

Andr\'e Ribeiro, R\'uben Garrido, Alexander Christiansen, Richard A. A. Jonker, S\'ergio Matos The BIT.UA team from the University of Aveiro describes its system for the 14th edition of the BioASQ Task B biomedical question answering challenge, built on a substantially refactored and modular codebase. For document retrieval, the PyTerrier PISA index was replaced with PostgreSQL-based pg_textsearch for BM25 search and Qdrant for GPU-accelerated dense embedding search, combined with HyDE query expansion, a Context-1 retrieval strategy, and a new reranker trained with dense-retrieval negative sampling. For answer generation, the team added an LLM-as-a-judge framework and an agent quorum mechanism in which multiple agents with diverse prompts debate and iteratively converge on a consensus answer using adaptive document retention, and it entered the snippet generation subtask for the first time. The Phase A retrieval systems reached MAP rank 5 in Batches 1 and 3, and all code is openly released.

Deep Microcompression: Structured Pruning and Bit-packed Quantization for Microcontrollers

Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe Running convolutional networks on bare-metal microcontrollers is limited by kilobytes of RAM and flash rather than by compute. Deep Microcompression (DMC) is a hardware-aware pipeline combining structured pruning, quantization-aware training, and fixed-length bit-packing, and it emits a dependency-free C library with deterministic latency. On LeNet-5 it reaches a 55.8x weight compression ratio at 98.77% accuracy, cuts binary size by 3x versus TensorFlow Lite on the RP2040 at matching accuracy, and delivers the first documented deployment of a standard CNN on the ATmega328P, a device with only 2KB of SRAM.

Conformal Prediction for Offensive Security

Giovanni Cherubin cross-listed Conformal Prediction (CP), a technique for producing prediction sets with coverage guarantees, has been used defensively in cyber security but its use for carrying out attacks is hard to find in the literature. The authors examine this gap by applying CP offensively in two areas: Privacy-Preserving Machine Learning and network traffic analysis. The paper reports initial findings in both areas rather than a complete attack pipeline, positioning CP as a component of attacks rather than defenses.

CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review

Jicheng Zhou, Kemou Li, Kahim Wong, Zheyuan Li, Zhuan Shi, Fengpeng Li et al. Reports from the AAAI-27 review cycle raised concerns that reviewers coordinate bids to get assigned to each other's papers, but collusive intent is unobservable in real data and no counterfactual exists for the same conference. CABAL is an end-to-end multi-agent simulation that holds the conference environment fixed and populates it with LLM-driven reviewer agents following honest or collusive policies, plus an affinity-guided strategy that forms collusion rings and picks target papers consistent with the colluders' expertise rather than at random. In controlled runs, collusive bidding more than doubles target-paper capture, and assigned colluders score their targets about two points higher than honest co-reviewers, while conference-wide effects stay modest. Bid-phase detectors offer limited help: positive-bid graphs are confounded by benign affinity, and a Very-High-only diagnostic view recovers rings precisely but with low coverage.

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

Alexander Neubauer, Tianzhen Hong, Han Li, Mengbo Yu, Amin Darbandi, Yannick F\"urst et al. Building automation systems produce abundant sensor data but remain hard to use operationally because of inconsistent point naming, missing metadata, and fragmented documentation. The authors systematically review and code 66 peer-reviewed studies on large language models (LLMs) for heating, ventilation, and air conditioning (HVAC) operations published between 2023 and March 2026, classifying each across five application families and three LLM method families and assessing evidence realism, deployment readiness, and where responsibility sits between the LLM and physical HVAC decisions. The corpus is concentrated in building energy modelling (BEM), with 32 of 66 papers, and only four studies reach pilot-level evidence, none reports sustained operational deployment, and none was judged ready for industry adoption now. The review concludes that current evidence supports LLMs as semantic and workflow layers, such as point-name normalisation and document-grounded operator support, rather than as autonomous controllers, where conventional machine learning, model predictive control (MPC), and reinforcement learning remain more adopted.

The History Is the Detector: Executing CVE Patch History, End-to-End

Qiushi Wu, Kevin Eykholt, Youngja Park, Xiaokui Shu, Dhilung Kirat, Douglas Lee Schales et al. cross-listed Public vulnerability databases and their fixing commits record exactly why code was unsafe, but that knowledge is written for human inspection and the same unsafe patterns persist in code without an advisory. BUGSTONE-E2E mines reusable detection rules from verified fixing commits, capturing scan anchors, fix semantics, and Common Vulnerabilities and Exposures (CVE) provenance organized by Common Weakness Enumeration (CWE) and language, then applies them through a funnel: Tree-sitter enumerates call sites matching rule anchors, cheap heuristics discard benign sites without LLM calls, LLM-based agents inspect survivors guided by the rule, and the system re-triages, builds runtime verifications, and generates scope-checked patches validated by two-sided differential tests. From 19,325 high-severity CVEs spanning 2022 to 2026, it identified 2,710 fixing commits and built 1,033 detection rules across 56 CWE families, packaged into 172 skills. Applied across 14 programs, the pipeline produced runtime evidence for 644 findings.

When LLM Decompilers Recompile More and Preserve Less

Chang Liu, Edward Raff, Kristopher Micinski cross-listed Decompilation recovers source code from binaries and underpins vulnerability detection and malware analysis, but large language model (LLM) based decompilers are now judged almost entirely by whether their output recompiles and passes shipped input/output tests, metrics that can reward code that builds cleanly yet behaves differently on other legitimate inputs or silently drops a known vulnerability. Decompile-Diverge is a behavioral comparison oracle that synthesizes a driver for each function, grows a fuzzing corpus from the reference, and reruns the decompiled code on the same inputs to detect divergence without hand-crafted tests. Across eight systems in nine configurations on established LLM decompilation corpora, candidates that pass every shipped test still diverge from the original on 4.9% of functions overall and up to 13% for one system. On 300 real GitHub library functions and 287 CVE-grounded functions, the strongest refinement LLM lifts Ghidra's build rate from 75% to 90% while its behavioral match rate falls from 74% to 62%, up to a tenth of disclosed vulnerabilities lose their crash entirely, and source-level analysis traces the divergence to invented fields, types, callees, and guards that replace the visible unknowns traditional tools leave behind.
72 more specialized papers

Large Language Models 60

Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation

Seifeldin Abdellatif cross-listed Low-bit key-value (KV) caches cut decoding memory, but the resulting quality loss varies by model and quantizer. Keeping the quantizer fixed, the method distills the floating-cache model's behavior into low-rank Q/K/V projection updates while the student runs a physically packed incremental cache. Across three seeds, 4-bit affine-cache adapters recover 54.24% of the held-out perplexity gap on TinyLlama-1.1B and 75.96% on Gemma-4-12B, and on a frozen NF4 Llama-3.1-8B base recover 60.42% under KIVI K2V2 and 37.61% under KVarN K4V2 while preserving 180-case associative retrieval; Gemma's score on a RULER subset rises from 42.80 to 48.33 against 46.15 floating. A 2-bit sweep drops TinyLlama's perplexity from 576.10 to about 11.4 versus 10.40 floating but restores only 11 to 12 of 180 retrieval cases, showing that perplexity recovery need not restore long-context retrieval.

Scalable Context Orchestration for Serving LLMs Over Voice

Linyi Jiang, Silvery D. Fu, Yifei Zhu cross-listed Voice assistants built on large language models (LLMs) typically treat conversation context as a flat, growing message list, leaving cues such as speaking rate, background noise, and packet loss implicit in the audio and leading to misaligned responses and rising costs over long sessions. llmovoice is a context-management middleware that, at each turn, assembles a bounded voice context from the current input, relevant history, and explicit paralinguistic and environmental state, then has the serving LLM reason over it to emit runtime directives that steer the response. On real voice applications and benchmarks it cuts speaking-rate alignment error by 52.4%, drops the false-interruption rate under packet loss from 46.0% to 0.9%, and lowers model usage cost by 79.2%. For long sessions it reduces per-turn cost by up to 24.9 times while keeping up to 98.7% of baseline answer quality.

Evidence Integration in Large Language Models

Sebastien Kawada, Manolis Kellis LLMs increasingly reason over evidence supplied by tools, retrieval, other agents, and users, yet how they fold that evidence into answers they have already started forming is poorly understood. The authors propose a distributional theory in which evidence shifts the receiver's distribution over initial answers, governed by a receiver prior weight and a candidate evidence tilt, predicting that candidates the receiver already finds probable are more persuasive, that models absorb their own characteristic errors more readily than foreign ones, and that identical evidence can help weaker models while harming stronger ones. These predictions hold across over ten million trials, twelve LLMs from four families, and eight domains including quantum mechanics, physics, genetics, and molecular biology, and models integrate candidate answers even after internally verifying them as invalid, in 93 to 100% of cases under propositional constraints. Causal interventions locate candidate integration late in the network as a structured sequence of admitting, promoting, and transporting external answers, while representations of verification are decodable but have little causal effect on the final answer.

SharedSAE: One Feature Dictionary Across Language Models

Daniil Ognev, C\'elian Vasson, Lijie Hu, Kentaro Inui, Benjamin Heinzerling Sparse autoencoders (SAEs) are a standard tool for interpreting language model activations, but training and latent labelling are normally repeated for every model. SharedSAE pairs one shared dictionary with model-specific encoder-decoder pairs, normalizing only the selection scores so activation magnitudes are preserved, and uses model dropout so a single model can be run at inference, unlike the closest prior method, which discards magnitudes and requires all models. Trained on four 1B-scale base models from distinct families with different tokenizers, it retains 96.6% of the mean explained variance of dedicated per-model SAEs, its latents show cross-model correlations 1.8 times those of separately trained SAEs aligned post hoc, and latent descriptions transfer across models. Once the dictionary is frozen, new models can be adapted to it efficiently with near-dedicated reconstruction quality while reusing the shared descriptions.

A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models

Jirui Qi, Mingyang Wang, Hinrich Sch\"utze, Raquel Fern\'andez, Arianna Bisazza Multilingual language models often answer semantically equivalent questions differently across languages, and methods for improving cross-lingual consistency (CLC) have been evaluated with incompatible models, tasks, and protocols. The paper runs a unified evaluation of representative inference-time and post-training CLC methods for question answering across three model families and three closed-form benchmarks, then checks on two culturally diverse QA benchmarks whether the methods suppress legitimately different answers to culture-dependent questions. Post-training methods are more reliable, with direct distribution alignment improving consistency in every model-dataset combination, while other methods are sensitive to answer format and language coverage, and cross-domain transfer is limited unless output formats match. Controlled closed-form evaluation shows no systematic loss on culture-dependent questions, but open-ended generation reveals occasional accuracy drops, especially for non-English responses.

What Attention Recalls and Recurrence Controls in Hybrid Language Models

Kirill Afendulev, Alexey Dontsov, Elena Tutubalina, Anton Korznikov Hybrid language models pair attention with a fixed-size recurrent state, but it has been unclear what each channel actually contributes. Two cache-level interventions probe this: split-prefill keeps only the key-value (KV) cache or only the recurrent state after prefilling a context, and state-swap pairs the KV cache from one context with the recurrent state from another in a single forward pass. On Qwen3.5 and Falcon-H1, exact retrieval survives only through attention, at 64-98% of full accuracy, and collapses to zero through recurrence, while output language and persona survive through the recurrent state and largely vanish when only the KV cache is kept. State-swap confirms the split causally, with the answer's content coming from the KV side and its language from the recurrent side, and recurrent-only generation even accepts words never seen in context that share meaning or parts with seen items.

GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion

John Seon Keun Yi, Joshua R. Minot, Dokyun Lee LLMs in high-stakes settings often produce plausible but ungrounded claims, and standard retrieval-augmented generation (RAG) does little to help because it retrieves isolated passages without tracking cross-document evidence or quantifying uncertainty. GRACE breaks a model's response into atomic claims, links each to trusted knowledge priors in a weighted bipartite graph, and uses weighted centrality to classify claims as Grounded, Refuted, or Boundary, the last category capturing novel or contested claims at the edge of the model's knowledge. A Return on Attention objective sends a claim to expert review only when its priority-weighted uncertainty exceeds the cost of verification, and verified claims become new evidence anchors so the knowledge base grows over iterations. Across multiple models and both general and domain-specific datasets, the graph-grounded knowledge base outperforms RAG baselines for retrieval and the Return on Attention rule efficiently selects boundary claims worth verifying.

When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models

Berkcan Kapusuzoglu, Connor Pryor, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu et al. Expert pruning shrinks Mixture-of-Experts (MoE) models by dropping experts the router deems unimportant, but that signal collapses under over-dispersed routing, where aggressive load-balancing during training spreads tokens almost uniformly across experts. In this regime perplexity stops predicting downstream accuracy: on gpt-oss-20B the lowest-perplexity pruning configuration gives the worst mathematical reasoning while the highest-perplexity one preserves it, unlike Mixtral-8x7B-Instruct where the two degrade together. No single scoring metric wins either, since activation-aware scoring keeps math but loses 18 points on GPQA, and frequency-based scoring shows the reverse. Minimax Expert Score Allocation (MESA) iteratively boosts scores for experts serving whichever domain is currently worst hit, and at 25% expert pruning it achieves the smallest worst-case degradation across domains, beating activation-aware baselines on 7 of 11 benchmarks and generalizing to gpt-oss-120B, Gemma-4-26B-A4B, and OLMoE-1B-7B.

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang Performance modeling for hardware design and software optimization requires structured reasoning about computation, data reuse, storage, and movement. PerfReasoning evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code: given workload, architecture, and mapping specifications, models must compare mappings and predict off-chip traffic and buffer requirements. The strongest closed models exceed 90% on reasoning question answering and the best open-weight model reaches 82.4%, but model construction is much harder, with only GPT-5.6 Sol exceeding an 80% pass rate while all other configurations average below 15% and vary widely across runs. Task-specific reinforcement learning lifts a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective.

Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs

Tung-Ling Li, Jiale Huang, Lee-Chi Wang, Janaki Ram Gotei Merging a LoRA adapter into its base model is standard deployment practice, but on native 4-bit microscaling checkpoints such as NVFP4 and MXFP4 the merged weights must pass back through a quantizer that re-derives the discrete E2M1 code plane, roughly 90% of the artifact's bytes, coupling the result to one quantization convention and, done naively, deleting the adaptation by up to 39 percentage points because the reconstruction optimum against an already-on-grid base is the base itself. Scale-QLoRA instead trains only the native per-block scale field on the deployment grid and freezes every E2M1 code, so within a fixed format, scale grid, and block layout the merge becomes a bit-exact identity and the artifact is code-invariant. Across four models and four tasks it is accuracy-lossless, matching merge-aware QAT-LoRA, while sidestepping quantizer sensitivity where rounding-rule mismatches can drive weight-space artifacts toward 0%. Freezing codes also removes the weight-space straight-through estimator from training, giving 3.9x faster steps on a dense 8B model, and enables exact rollback, code-plane deduplication, and roughly 125x faster scale-only task swaps.

Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One

Fred Zhangzhi Peng, Kaiwen Zheng, Anru R. Zhang Language generation is nearly always sequential, either token by token in autoregressive models or through long iterative refinement trajectories in diffusion language models. PlaidQ is a 0.7B continuous diffusion language model for code that repurposes a pretrained autoregressive model as a bidirectional denoiser over continuous token embeddings, then distills its trajectory with distribution matching for few-step generation and paired-trajectory supervision for one-step generation. At matched scale it is competitive with discrete diffusion models on code, and a 16-step distilled student reaches 31.78 and 40.49 pass@10 on HumanEval and MBPP+, beating its own teacher sampled for 512 steps. A single denoising step still yields functionally correct programs at 7.07 pass@1 on HumanEval, and code and checkpoints are released.

From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs

Omer Nahum, Niv Nayman, Jonathan Fhima, Alon Zolfi, Jeremy Levy, Shai Mazor et al. Reliable LLM deployment requires distinguishing uncertainty that stems from irreducible task ambiguity from gaps in the model's knowledge. Existing decomposition methods generate multiple clarifications of an ambiguous input, answer under each, and compare the answers, but the authors argue theoretically that the answers are redundant, costly, and prone to epistemic leakage, and instead estimate ambiguity-induced aleatoric uncertainty directly from the space of plausible interpretations. On ambiguity detection across three benchmarks, the clarification-only approach raises AUROC to 63.34 from 60.85 while cutting output tokens 4-26x and API calls 2.2-3.5x, and its estimates correlate substantially less with epistemic uncertainty.

Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models

Xing Chen, Hengshuai Yao Fine-grained Mixture-of-Experts (MoE) models route each token to the top-k experts and renormalize router probabilities, and the authors show this renormalization implicitly calibrates expert output gain to the training-time k, so reducing k at inference changes not just which experts fire but the strength of the expert branch. Their fix activates the top k1 experts while normalizing by the probability mass of the top k2 experts, adding a single integer with no parameters, training, or measurable compute overhead. On Qwen3.6-35B-A3B, going from 8 to 4 experts costs 4.65 MMLU points under standard renormalization but only 0.35 points with k2=16 while halving routed-expert compute, and the result replicates on the 11x larger Qwen3.5-397B-A17B, where 10 to 5 experts loses only 0.55 points. Removing renormalization entirely is catastrophic, expert identity matters far more than weighting, and perplexity and downstream accuracy favor different k2, so compression settings should not be chosen from unlabeled text alone.

Optimizer Memory Schedules for Outscaling the Overtraining Axis

Katie Everett, Shikai Qiu Optimizers are usually compared at a single training horizon, yet relative performance and optimal hyperparameters can shift substantially with the amount of overtraining. The authors compare the matrix-preconditioned Muon and SOAP, the momentum-scheduled ADANA, and AdamW on models from 51M to 253M parameters across overtraining factors from 1x to 256x, sweeping base learning rate at every setting. They find the preferred learning rate schedule can reverse across the overtraining axis, the best weight decay scales roughly as the square root of the overtraining factor, and longer horizons favor longer fixed memory. With log-time weight decay and momentum cooldown, ADANA outscales AdamW with an exponent advantage close to that predicted by DANA theory, and while Muon and SOAP hold roughly constant token-efficiency advantages over AdamW, ADANA overtakes Muon and becomes competitive with SOAP at the highest overtraining factors.

When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

Gnaneswar Villuri, Hashmath Shaik, Alex Doboli A 0.6B language model asked to verify 1,200 logical conclusions, half corrupted by a single semantic edit, answers YES every time, yet linear probes on its hidden states read the correct verdict at 0.96 AUC and transfer to unseen logical structures. The authors trace where the verdict is lost and find it survives to the output logits, where the margin still carries 0.89 AUC along a well-aligned readout direction, so a single saturated decision threshold offset by +4.6 sigma is what erases the behavior. Across 90 semantic-label configurations spanning five models and three families, behavioral accuracy collapses onto a function of threshold offset (Spearman -0.93), a one-parameter correction never fit on evaluated structures repairs accuracy from 50% to 81% at 0.6B, calibrated margin decoding recovers 94% at 8B, and few-shot prompting is shown to work by the same recentering mechanism. Comparing probe to margin separates concealed, miscalibrated, and undetected regimes, and the authors caution that in standard generation settings, answer-surface features and heuristic labels can reproduce published probing results without any internal access.

Choosing the Right Language Mode at Inference Time for Multilingual Reliability

Ekata Mitra, Ameeta Agrawal Multilingual large language models reason poorly in low- and mid-resource languages, and while translating into English can help by tapping stronger English-centric representations, it is unclear how much translation helps before it starts causing interference and overconfidence. Experiments with LLaMA and Qwen models vary text scope and language mode (target-only, English-only, bilingual) and find a trade-off: English context improves understanding and recovers errors caused by non-English comprehension, but redundant bilingual context intensifies interference. Reliability-Aware Adaptive Inference (RAAI) is a training-free test-time method that routes and fuses prompts based on Expected Calibration Error (ECE) and uses a mid-layer Risk Index (RI) to gate sequential reasoning, spending extra compute only where it is likely to help. Across the two model families, RAAI raises accuracy by 25 to 37.7% on low-resource languages and lowers calibration error by roughly 3 to 6%, with the largest gains in the lowest-resource language tiers.

Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

Nathan Wolfrath, Meghan Conroy, Thomas Kosten, Dave Bell, Bhabishya Neupane, Jonah Kindel et al. Commercial large language models (LLMs) are retired on deprecation schedules, which threatens the computational reproducibility of research built on them. The authors searched PubMed for original research articles from 2022 through March 2026 that applied a specific LLM to a biomedical task, used an extraction agent to pull model names from 61,077 abstracts with human validation on a subset, and compiled release and retirement data for the 50 most frequently used models. Across 8,931 paper-model mentions in 5,242 publications, 77.7% cited a closed-weight commercial model and 42% involved a model already retired at publication or scheduled to retire within two years, with a median of 538 days from publication to retirement.

Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

Aditi Patodiya cross-listed Prefix caching, which reuses the key and value tensors of a shared prompt prefix across requests, is on by default in major open-source serving stacks and assumed to be a transparent optimization. Holding model, decoding parameters, seed, and request order fixed and issuing every request serially at batch size one, the authors ran an eighty-episode multi-turn agentic tool-use workload with caching enabled and disabled across two engines and four weight formats. Enabling the cache changed the agent's trajectory on 36.2% of episodes at 16-bit precision and on 75.0% at four-bit, while cache-disabled repeated runs were bit-identical in all 800 episodes; follow-up experiments trace run-to-run divergence to a single server-level prompt-cache setting and show that cached serving is deterministic given cache state but irreproducible in practice because that state is absent from the request and never reset by default.

Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges

Chenqi Li, Minghui Min, Dusit Niyato, Wei Ni Diffusion language models (DLMs) generate text by iteratively denoising tokens rather than decoding left to right, which lets them refine many uncertain tokens in parallel, use bidirectional context, and trade quality for latency more flexibly than autoregressive large language models (LLMs). This survey argues these properties suit agents running on mobile edge devices, where partial refinement, early exit, and constraint-guided correction can cut response delay and communication cost under noisy or incomplete context. It reviews DLM foundations and maps them to edge constraints on latency, memory, energy, bandwidth, privacy, and reliability, covering efficient architectures, training and inference acceleration, compression, edge/cloud deployment, Internet of Things (IoT) and wireless applications, and agent evaluation, before listing open problems in long-context state management, split inference, and reproducible benchmarking.

Can Activation Steering Capture Multidimensional Authorship Style?

Hieu Tran, Calvin Bao, Marine Carpuat Activation steering works for well-defined attributes, but authorship style is multidimensional and hard to specify in words. The authors construct per-aspect steering directions from structured contrastive prompts along rhetorically motivated dimensions and find that the directions share a common authorship backbone while conflicting on aspect-specific residuals, which explains why naive averaging of directions fails. Aspect-Aware Activation Steering (A3S) merges the per-aspect directions with interference-aware aggregation and tunes steering strength per instance, improving style transfer where styles are genuinely multi-aspect, beating a trained baseline in preference evaluations on out-of-domain benchmarks, and keeping overlap with target exemplars low.

When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models

Xiaodong Li, Peiwei Liu Financial large language models used to summarize reports often fabricate numbers, and prior work has blamed weak numerical reasoning without testing that assumption under controlled fine-tuning. The authors compare a base instruction-tuned model against a domain language-adapted variant (FT-A) and a numeracy-enhanced domain variant (FT-A+B+C), scoring outputs with a three-level taxonomy that separates overt currency fabrication, covert-explicit professional-convention numbers, and covert-implicit ungrounded quantitative claims. Overt hallucination rises from 5.4% for the base model to 82.5% after domain adaptation and 98% after adding numeracy supervision, with the fine-tuned models frequently injecting memorized template values regardless of the input. The authors conclude that domain adaptation erodes numerical restraint rather than reasoning ability and recommend evaluating all detectability levels plus grounding-aware generation or abstention at deployment.

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures

Oskar Holmstr\"om, Marcel Bollmann, Marco Kuhlmann Multilingual language models develop shared cross-lingual representations, and several interpretability metrics claim to quantify that sharing, but they were developed in isolation and their disagreements are hard to interpret. The authors compare four metrics, CKA, ANC, Gaussian mixture model (GMM) dominance per token, and ILO, across 21 base models from five families spanning 125M to 14B parameters, correlating each with cross-lingual transfer on five downstream tasks. The metrics disagree substantially, and the authors trace the disagreement to anisotropy, the tendency of representations to cluster in a narrow cone of embedding space. Only ILO retains a strong correlation with cross-lingual transfer (Spearman's rho of 0.90) after controlling for model size, family, and task, so they recommend it as the primary sharing metric, reported alongside anisotropy diagnostics.

On Epistemic Diversity in Large Language Models

Elisabeth Kirsten, Nicole Kr\"amer, Muhammad Bilal Zafar Large language models increasingly answer questions, explain, and teach, where a correct answer can still narrow a user's access to alternative valid answers and reasoning routes. Drawing on philosophy and social epistemology, the authors formalize epistemic diversity for LLMs as the range of valid answers, explanations, and reasoning paths a model exposes, argue it is a needed evaluation dimension for knowledge-intensive use, and propose a preliminary framework for measuring it, operationalized in two domains. Frontier LLMs frequently exhibit epistemic narrowness, collapsing large spaces of valid answers onto small canonical subsets. The authors conclude that evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as a model capability.

Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference

Zhenhe Wu, Yaping Jin, Qinghua Xing, Hang Zhou, Wei He, Xianjie Wu et al. Mixture-of-Experts (MoE) models activate only a few experts per token, but the full expert set often exceeds GPU memory, so decoding repeatedly transfers weights. This work treats expert-cache management as a model-side algorithmic problem, jointly post-training the backbone with lightweight auxiliary cache routers while preserving the native top-K selection rule at inference: a Temporal Router predicts same-layer reuse and retains experts without proactive loading, and the full Spatio-Temporal Router additionally uses the causal predecessor's hidden state to refine the cache before the target layer is accessed. Evaluated on Qwen3 and GPT-OSS across GSM8K, MATH, and CommonsenseQA, the full router on Qwen3 improves load-adjusted hit rate by 1.15 to 18.03 points and reduces expert-weight traffic by 4.6 to 53.3% versus the strongest prefetching baseline, with competitive but task-dependent results on GPT-OSS. An auxiliary-only ablation keeps baseline accuracy but yields far smaller cache gains, showing the joint post-training does the work.

Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair

Xuemeng Cai, Jiakun Liu, Linhan Yang, Wei Ma, Lingxiao Jiang cross-listed Evaluations of large language models (LLMs) for automated program repair (APR) mostly score final patches and say little about where the model's reasoning departs from the actual repair evidence. The authors define hallucination as producing patches or intermediate artifacts not grounded in that evidence, and measure it both in final patches and in three understanding tasks: identifying the triggering test case, predicting line coverage, and generating additional test cases. Running three LLMs on 832 Defects4J bugs, they find only 21.0% to 55.9% of patches pass the developer test suite, and manual analysis of 812 sampled repairs flags repair hallucinations in 72.7% of cases, including patches that pass every available test, with wrong causal localization (45.9%) and wrong repair strategy (18.5%) as the leading causes. Better intermediate artifacts generally accompany successful repairs, but the correlation is not reliable, and models often misidentify triggering tests, mispredict coverage around branches, and write tests that miss the bug-triggering condition.

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi Large Reasoning Models (LRMs) generate long chains of thought whose key-value (KV) cache grows linearly and often exceeds GPU memory, while existing compression methods pick tokens to keep based on recent queries, assuming those predict future attention. The authors show this assumption breaks in long reasoning: certain Thought Revisiting Tokens re-attend to distant context such as plans formed early in the trace, and the queries behind them fall into a small number of clusters in embedding space. BeaconKV is a training-free method that keeps compact beacon queries representing each cluster to anticipate which KV pairs will be revisited without storing the full query history. Across four open-source LRMs and diverse reasoning benchmarks it generally outperforms existing compression methods, achieving up to 5.8x memory reduction while nearly matching full-cache accuracy and improving throughput by over 4.3x.

A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering

Songeun Lee, Kyungjin Min, Injae Na, Suyeong Lee, Chiyoung Kim, Woohwan Jung Structured retrieval-augmented generation (RAG) methods that reason over trees or graphs help with multi-hop questions but struggle with evidence-intensive QA, where answers must synthesize information spread across dozens or hundreds of documents, because their structures are rigid and evidence is gathered without regard to the reasoning topology. APT-RAG expands its reasoning structure adaptively based on question dependencies and evidence needs, and gathers evidence in a topology-aware way through sibling evidence reuse, direct retrieval, and aggregation from child nodes, with evidence-guided batched answer generation to cut generation overhead. On evidence-intensive QA benchmarks it outperforms existing structured RAG methods, and code is available.

Amortizing Scaling Law Construction Costs

Abhash Kumar Jha, Diana Alexandra Onu\c{t}u, Neeratyoy Mallik, Swagatam Haldar, Sam Laing, Niccol\`o Ajroldi et al. Fitting a scaling law needs only the best-loss frontier across compute scales, yet standard practice trains an exhaustive grid over hyperparameters, token budgets, and parameter counts and then discards most of it. The proposed framework casts data collection for scaling laws as a Bayesian optimization problem and introduces metrics for comparing fitting methods under constrained compute budgets. Progressively expanding the compute budget during acquisition, mirroring the compute-ordered evaluation of configurations in practice, substantially improves recovery efficiency, and augmenting the observed runs with surrogate-fantasized evaluations reconstructs the broader experimental grid. Together these closely match dense-grid scaling law fits at 10 to 100 times lower compute cost.

Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection

Renato Vukovic, Hsien-chin Lin, Carel van Niekerk, Benjamin Ruppik, Michael Heck, Shutong Feng et al. Detecting hallucinations, outputs that are factually incorrect or unsupported by the source, is hard because large language models give little insight into why a response may be inaccurate. The approach tests whether a low-level symbolic competence such as SQL can serve as an unsupervised grounding mechanism for a high-level task: an LLM first builds an SQL database from the reference documents, and the detection pipeline then reasons over both the reference and the sampled response through that database, providing a neurosymbolic check. On the RAGTruth and DiaHalu hallucination detection datasets, the method improves on direct prediction and competes with state-of-the-art detectors without any domain-specific fine-tuning, relying only on a general competence already present in LLMs.

EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages

Aleix Sant, Jordi Luque, Carlos Escolano Machine translation (MT) is a scalable way to extend English instruction-tuning data to other languages, but it can distort task-critical constraints and required outputs, producing corrupted training examples that degrade the models trained on them. EuroAlpaca is a task-preserving localisation pipeline and near-parallel resource covering 50 European languages and regional varieties that either applies field-wise MT while preserving task-critical content or reconstructs a task-equivalent target-language instance, followed by validation of cross-field coherence and target-language consistency, and it is paired with European-IFEval, a multilingual benchmark for verifiable instruction following. In LoRA experiments with four LLMs, directly translated data improves ROUGE-L and F-BERT on the Aya Evaluation Suite but cuts European-IFEval accuracy by 29.8% relative to the unadapted baseline, whereas adaptation with EuroAlpaca raises accuracy by 12.9% over the same baseline while also achieving the highest ROUGE-L and F-BERT scores on Aya.

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi Attention-head attribution in Transformer classifiers usually forces a choice between fine-grained circuit tracing and coarse output-level probing. The authors define an influence score that combines each head's directional effect on the logits with its structural contribution to the residual stream, so it can be aggregated at the head, layer, and whole-network level. Applied to a DeBERTa model fine-tuned for prompt injection detection, the score exposes distinct head-level decision patterns for correct versus erroneous predictions, offering a middle ground between circuit analysis and global output-based interpretability methods.

Single-Query Black-Box Calibration Auditing via Logit Bias

Roman Plaud, Antoine Saillenfest, Matthieu Labeau, Thomas Bonald, Willem Waegeman Standard calibration metrics need the continuous output probabilities that commercial LLM APIs increasingly withhold. The authors show that any API exposing a logit_bias parameter can be manipulated to test exact probability thresholds with strictly one query per sample, and build on this a provably consistent estimator of the True Calibration Error for binary classification tasks. The result is an efficient method for auditing the calibration of black-box foundation models used as zero-shot classifiers.

Large Language Models with At Most One Spike per Neuron

Zhuoya Zhao, Parsa Omidi, Aref Jafari, Richard Naud cross-listed Spiking neural networks (SNNs) promise energy-efficient large language models (LLMs) through sparse, event-driven computation, and time-to-first-spike (TTFS) coding pushes firing rates to at most one spike per neuron per time window. Conventional TTFS networks cannot express operations such as layer normalization and matrix multiplication, so the authors introduce a reference-based encoding strategy for embedding layers, layer normalization, attention-related operations, and dropout, then build and train a fully TTFS-based architecture end to end. On BERT and GPT-2, the spiking models match their artificial neural network counterparts on natural language understanding and common-sense reasoning, while a clear gap remains on language-modeling perplexity; the authors describe this as the first TTFS spiking LLM scaled to 1.5 billion parameters. Reported energy figures are a spike-count proxy under an established cost model rather than measurements on neuromorphic hardware.

Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG

Shuyu Guo, Shuo Zhang, Zhaochun Ren Retrieval-Augmented Generation (RAG) improves answers with retrieved documents, but long contexts slow inference, and soft compression methods that shrink each document into a short embedding sequence are usually trained by distilling from the uncompressed system, capping them at its performance. DEX-Comp uses a two-stage recipe: Pure Distillation warm-starts the compressor only on responses the uncompressed RAG got right, then Hard Exploration runs reinforcement learning solely on queries the uncompressed system fails, pushing the model toward computation patterns better suited to compressed inputs. Across five open-domain QA benchmarks at retrieval depths from top-5 to top-30, DEX-Comp compresses contexts 16x and speeds inference 4x to 24x while matching or exceeding the uncompressed baseline. Ablations across datasets and backbones attribute distinct contributions to each stage.

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You On-Policy Distillation (OPD) is a common post-training method for reasoning models, but which training data actually drives its gains has been little studied. The authors first try 1-shot OPD, training on a single example, and find it consistently effective across sampled examples, with harder problems giving larger gains; analysis shows the improvement comes not from high token entropy but from the longer chain-of-thought (CoT) paths that hard problems produce, which keep the student aligned with the teacher over long horizons and teach patterns such as reflection that short CoTs lack. They propose selecting only hard examples, including ones that completely exceed the teacher's own ability. Across four models from 1.5B to 7B parameters, training on just 8 selected hard examples matches a 17K-example baseline.

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

Zukang Xu, Zhixiong Zhao, Xing Hu, Jiangyong Yu, Houji Wen, Jun Li et al. Mixture-of-Experts (MoE) language models route every token to a fixed top-k set of experts, spending compute on experts that contribute little, and existing skipping methods depend on router confidence, calibration data, or extra training. ACE is a training-free, calibration-free scheme that scores each routed expert with two offline-computed views, a Global Spectral Proxy (GSP) derived from the gate, up, and down projections with RMSNorm scaling, and a Router-Conditioned Refinement (RCR) that measures expert responses along routing-preferred directions, and skips a slot only when both views and the runtime gate agree it is low-contribution, always keeping the top-1 expert. Across three MoE models and eight benchmarks it beats static and dynamic baselines with the gap widening under aggressive skipping; at 50% expert skipping on Qwen3.6-35B-A3B it lowers WikiText-2 perplexity by 7.96% and raises average downstream accuracy by 4.15 points over the strongest competitor.

Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference

Mostafa Elhoushi, Alex Pretko, Nolan Dey, Bin Claire Zhang, Gavia Gray, Gurpreet Gosal et al. Layer dropout, also called stochastic depth, speeds training and enables zero-shot layer pruning in vision and language transformers, yet it has vanished from LLM pretraining recipes amid reports that it hurts accuracy, without any systematic study of the effect. Across more than 2,400 pretraining runs on Cerebras CS-3 systems spanning 271M to 8.2B parameters and up to 160B tokens, the authors establish best practices for the layer distribution, time schedule, and optimizer hyperparameters and show that at equal training FLOPs layer dropout yields lower loss. Models reach lower or similar validation loss while saving up to 25% of training FLOPs, and the trained models support post-training optimizations such as early exit, intermediate-layer skipping, and self-speculative decoding for up to 1.5x inference speedup with negligible accuracy loss.

Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

Maria Mahbub, Ashley Rice, Michael R. Munroe, Amidu Kamara, Amir Sadovnik Meta-evaluation of reference-based automatic metrics for natural language generation usually measures agreement with human judgments, which says little about how an evaluator behaves under controlled conditions. The authors propose behavioral correctness assumptions: a taxonomy of correctness-preserving and correctness-altering response transformations, each paired with the scoring behavior a sound evaluator should exhibit. Applying this to lexical, character-level, semantic, LLM-based, and hybrid evaluators, they analyze assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. No evaluator satisfies all the proposed assumptions, and evaluators with similar aggregate scores can have substantially different behavioral profiles.

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

Yang Li, Semih Yavuz, Shafiq Joty On-policy distillation (OPD) gives dense per-token supervision for post-training language models, but external teachers suffer from distribution mismatch and self-distillation with privileged context is limited by in-context learning capacity. RISE (Recursive Improvement via Self-Extrapolating Policy Distillation) instead builds a synthetic teacher from the model's own reinforcement learning with verifiable rewards (RLVR) trajectory, extrapolating the displacement between the current checkpoint and a trailing anchor in parameter or logit space to turn a sparse outcome-driven update into a dense token-level target. Because the teacher is refreshed every iteration as the student improves, distillation becomes a recursive loop in which outcome rewards ground the extrapolation and the extrapolated teacher refines token-level decisions. Across mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks, RISE outperforms both RLVR-only training and on-policy self-distillation.

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong Hyper-Connections and their manifold-constrained variant mHC widen the residual pathway from one stream to several, but it has been unclear how trained models actually use this extra capacity. The authors analyze the four-stream residual pathway of DeepSeek-V4-Flash using effective stream counts, cross-stream residual weights, and inter-stream cosine similarity. A typical attention or feed-forward site effectively reads and writes about two streams, the dominant stream changes across depth, and residual mixing is concentrated in early layers, with layers 22 through 42 mostly carrying each stream forward separately. Replacing the late mixers with identity raises C4 perplexity by only 1.9% and preserves the six-task average, while replacing the early mixers raises perplexity by 41%, indicating the model realizes only part of the flexibility mHC affords.

Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin et al. Accuracy on molecular property benchmarks cannot distinguish a large language model (LLM) that predicts a property from one that recalls a published number. The authors audit 22 frontier models on 12 regression benchmarks for digit-level verbatim retrieval and find it widespread but benchmark-specific: on five datasets more than 50% of the models show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. Running the same prompts on the same molecules at a higher reasoning level flags retrieval 89% more often than at the lowest level. An attempt to interrupt retrieval in the most contaminated cases shows the strongest models sometimes still recognise transformed SMILES strings paired with original labels, and suppressing retrieval pulls the models' relative prediction errors closer together, indicating that general predictive ability is not determined solely by how many values a model has memorised.
19 more specialized papers

Other 51

Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters

Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal cross-listed Distributed training repeatedly exchanges data among GPU nodes, so congestion on a single flow can stall an entire communication round, and existing remedies assume global control over all jobs or switch-level support that a single user in a shared cloud cannot rely on. REACT operates at the communication-library layer as a shim over NCCL, detecting congestion at runtime from readily available flow statistics and retuning the collective pattern, for example by changing which node aggregates in an AllReduce tree, while preserving the semantics of the exchange. On a shared academic GPU cluster it improves algorithm bandwidth by 13% to 38% under network congestion, and simulations across congestion scenarios show gains of up to 75%, with no support required from the underlying network infrastructure.

When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference

Ismail Erbas, Xavier Intes, Vikas Pandey In quantized recurrent networks, the stored low-precision state is fed back at the next time step, so the rule used to write that state can change all later computation. The authors name this rule recurrent-state write-back and isolate its effect in a compact GRU encoder-decoder for fluorescence lifetime imaging, which must estimate two lifetime parameters from very noisy time-resolved signals. With the trained model held fixed, switching to deterministic 4-bit state storage raises estimation errors by roughly 70x and 300x for the two parameters, because repeated small updates fall below the write threshold and the stored state freezes while the network keeps proposing change. Error feedback, residual memory, and direction memory carry the suppressed updates across time and restore accuracy without retraining, higher state precision can worsen a fixed solution, matched training can learn compatibility with the state interface, and an independently trained LSTM reproduces the failure with the cell state more sensitive than the hidden state.

Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters

Milos Gravara, Andrija Stanisic, Stefan Nastic cross-listed Compound AI workflows chain multiple models and software stages, and deploying one on a heterogeneous cluster means choosing a model variant per stage and a placement that satisfies service-level objectives (SLOs). System metrics can be profiled per stage and composed, but accuracy cannot, since upstream errors propagate downstream; end-to-end profiling scales poorly and product-of-stages surrogates misrank candidate plans. Atlas introduces MAP, a Markovian Accuracy Predictor that buckets intermediate outputs and composes local conditional accuracy transitions between adjacent stages along the workflow topology, then solves plan selection as a mixed-integer linear program maximizing predicted accuracy under SLOs. Across four workflows, MAP reaches Spearman correlation up to 0.947 with 2.6x less profiling, and the optimizer picks plans within 0.03 of oracle accuracy while cutting deployment cost by up to 42% through heterogeneous placement.

Mitra-v2 Technical Report

Yefan Tao (Bernie), Xiyuan Zhang (Bernie), Xinyi Liu (Bernie), Boran Han (Bernie), Danielle Maddix (Bernie), Haoyang Fang (Bernie) et al. Mitra-v2 is a tabular foundation model for real-world classification and regression, trained entirely on synthetic data from a pretraining distribution much larger and more diverse than Mitra-v1's, with a small 2D Transformer backbone that supports longer contexts and larger feature spaces. On the TabArena and TALENT benchmarks spanning more than 300 real datasets, it performs at the level of the industry-scale TabFM and EXAONE Tabular models and clearly outperforms TabPFN-3 and TabICLv2. It matches the 1.6B-parameter TabFM with only 77M parameters, about 5% of the size, and ranks first on classification tasks with more than ten classes despite pretraining only on tasks with at most ten. Weights, inference and fine-tuning code, and evaluation results are released under Apache-2.0.

SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds

Esteban Guti\'errez, Lonce Wyse, Frederic Font, Xavier Serra cross-listed Generative audio models for everyday sounds have grown so large that synthesizing them demands industrial-scale compute and massive datasets. SCAPES is a lightweight model that synthesizes environmental sound textures under high-level semantic control by operating on the continuous latent space of a neural audio codec rather than discrete tokens, splitting audio into overlapping segments and modeling the evolution of latent trajectories with a Continuous Normalizing Flow (CNF) trained via Flow Matching. A 36-million-parameter instance trains on small, uncurated datasets using a single consumer-grade GPU, converging after a training time of roughly twice the duration of the source audio, while producing outputs with long-term stability, semantic consistency, and smooth interpolation between semantic conditions. Code, pretrained weights, audio examples, and an interactive demo are public.

Dynamic Heterogeneous Graph Representation Learning: A Survey

Huan Liu, Pengfei Jiao, Jie Yin, Hongjiang Chen, Zhidong Zhao Real-world networks combine multiple node and edge types with structure that changes over time, and static or homogeneous graph representation learning (GRL) handles neither well. This survey gives the first systematic review of representation learning for Dynamic Heterogeneous Graphs (DHGs), starting from a unified formal definition that covers both discrete-time and continuous-time graphs. It organizes the literature into an algorithm-centric taxonomy spanning early embedding methods, graph neural network (GNN) models, and recent Transformer-based approaches, highlights each family's modeling bias with respect to temporal granularity, and summarizes applications, datasets, benchmarks, and open directions.

From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy

Petro Shulzhenko, Gabriele Spadaro, Enzo Tartaglione Depth compression shrinks networks by finding redundant activation functions, linearizing them, and folding the surrounding layers together, but existing methods cannot handle convolutions with padding because no analytical merge solution exists, and they enlarge the kernels of merged layers, which eats the speed-up. The strategy proposed here merges layers that lack an analytical solution and does so without any kernel-size growth. Validation covers multiple architectures and datasets, with inference speed-ups measured on real embedded platforms and code released publicly.

Fast Gauss Sums via Flash Attention

Nicolaj Rux, Sebastian Neumayer Weighted sums of Gaussian kernels underpin maximum mean discrepancies (MMDs), kernel gradient flows, Stein variational gradient descent (SVGD), and many other kernel methods, yet they lack the heavy hardware-level optimization that softmax attention has received. The authors show that two small augmentations of the inputs turn the normalized softmax reduction computed by flash attention into an unnormalized Gauss sum with arbitrary signed weights, so existing attention kernels evaluate it without any custom GPU code. For feature dimension above 8 in fp16, the approach beats both compiled PyTorch and PyKeOps kernels in speed, memory overhead, and accuracy, and memory usage stays linear in the number of points.
43 more specialized papers

Agents 50

Interface-Induced Trajectory Censoring

Wenbo Wang cross-listed Agent benchmarks read tool-call rates off the serving stack, but that number can be zero while the model is emitting well-formed calls, because the interface drops them before the executor or scorer sees them. On BFCL v4, holding weights, cases, decoding, and seeds fixed and changing only the serving adapter, the same model scores 0.00 or 0.96, and a 2x2 over chat template and parser puts the entire effect in their interaction, so fixing either component alone buys nothing. The same swap on tau-bench moves server-parsed calls from 0 to 636, and inside verl's AgentLoop at 7B, 45 of 115 generations carry a complete call yet none are accepted, executed, or returned as an observation; repairing the adapter at evaluation time restores the mechanism but yields no significant pass-rate gain. A 98-line preflight check is released that catches every silent failure observed.

Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines

Faizan Tanveer cross-listed Multi-agent LLM pipelines often assign the reviewer role to a cheaper model than the executor, but prior work held reviewer capability roughly fixed. The study varies it across models down to one that cannot solve the problems at all, tracking the outcome of every rejection on a fixed set of 100 olympiad mathematics problems. A cross-family mid-tier reviewer raises final accuracy from 52 to 64 percent with zero damaged answers, whereas same-model self-review has the highest error-detection recall (0.85) but no significant gain, rejecting 2.1 times as often with a third the repair rate and falsely rejecting 35 percent of correct answers versus 2 percent for the cross-family reviewer. Self-review's low damage rate turns out to be revision inertia rather than reviewer quality, since every falsely rejected answer the executor actually revised became wrong, and the weakest reviewer changed none of 100 final answers while doubling token cost; the authors frame this as a controlled pilot on a single configuration.

EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?

Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar et al. cross-listed LLM agents act through a harness of tools, reusable skills, and specialist agents that keeps changing as capabilities are added, yet existing continual-learning benchmarks put the non-stationarity in the task stream and hold the harness fixed. EVOHARNESSBENCH instead evolves the harness itself along three axes, with 17 multi-stage streams built deterministically from verifier-based benchmarks comprising 802 tasks, 520 tools, 42 skills, and 62 agents, evaluated under a deployment setting that measures retention as the harness expands and a self-evolving setting that tests whether accumulated experience stays useful. Results show harness expansion alone can degrade performance on previously solved tasks, a form of harness-induced forgetting. Gains from self-evolving adaptation are inconsistent across stages, axes, and environments, and retention and adaptation can pull in opposite directions.

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Lin Shi (Audrey), Haowei Lin (Audrey), Zixuan Zhu (Audrey), Xiaoyue Zhou (Audrey), Xiang Li (Audrey), Xiangning Lin (Audrey) et al. Evaluating agents across the growing set of agentic benchmarks is hard because each tends to need its own environment and agent integration. Harbor Adapters is a unified evaluation infrastructure that ports more than 80 benchmarks so arbitrary agents can run against them, validated through code review and parity experiments, and the authors use it to evaluate 8 models across 54 benchmarks, each run with the Terminus-2 harness and one of three native harnesses. They also distill Harbor-Index, a curated set of 82 difficult, diverse tasks from 29 benchmarks selected via difficulty filtering, AI and human audit, and an audit-and-fix loop, which keeps the breadth and challenge of the full suite while staying affordable to run. No evaluated model-harness configuration exceeds a 30% pass rate on Harbor-Index, with the strongest, GPT-5.5 with Codex, reaching 28.0%, and the adapters, results, analysis, and index are released as open source.

Abstraction Agent

Boning Li, Longbo Huang cross-listed Information abstraction, which groups strategically similar private states into a manageable number of buckets, is essential for scaling solvers to large imperfect-information games, but building good abstractions has required hand-engineered domain evaluators such as hand-strength or equity calculators that do not exist for most games. Abstraction Agent is a zero-shot pipeline in which a large language model (LLM) reads a natural-language game description, discovers continuous strategic features with calibration anchors, scores private states on them in batches, selects features by correlation, and clusters the results with k-means, all without a game-specific evaluator, training data, or game-tree traversal. The resulting abstractions reduce lifted-strategy exploitability by up to 62% relative to an expected-hand-strength baseline on heads-up no-limit Texas hold'em (HUNL) turn endgames and beat a scalar rank baseline at every granularity on ROVER Trials, an original game absent from any pretraining corpus. With unchanged prompts the pipeline also transfers to four-card Pot-Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, where the discovered features track each game's recognized strategic concepts.

Iris: Climbing to the Search Frontier

Ziyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu et al. Iris-mini and Iris-pro are search agents trained at the 35B-A3B and 397B-A17B scales, released together with their data pipeline and training recipe. Training questions are reverse-constructed from the hyperlink structure of a web corpus by authoring multi-hop chains over an entity graph, rewriting every non-answer entity into a descriptive reference so no clue can be resolved by string matching, and keeping only questions a reference model fails closed-book but solves with the supporting evidence. The questions become trajectories filtered at both trajectory and turn level for supervised fine-tuning (SFT), followed by reinforcement learning (RL) against live search with an in-cluster reward judge and observation summarizer, and the two stages alternate in a procedure called SFT-RL climbing that feeds the hardest solved and most efficient RL rollouts back into the next supervised pass. Because inference-time context management matters more on these benchmarks than most reported system differences, every result is reported with and without it using a single ReAct agent with no sub-agents or test-time verification, and with management enabled the two models reach 82.2 and 88.6 on BrowseComp, plus 84.8/85.1 on BrowseComp-ZH, 86.9/92.9 on DeepSearchQA, and 52.3/56.4 on HLE, the strongest open-source results in their parameter ranges.

VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes

Nikkie Hooman, Monarch Nigam, Amy E. Hughes, Rasmi G. Nair, Mehak Gupta Early-onset colorectal cancer is rising, but structured encounter data miss the symptom duration, context, and family history needed for early detection. VERGE is an agentic workflow that first proposes a label and supporting evidence via retrieval-augmented generation, then runs a bounded verification-refinement loop that checks textual grounding and clinical validity, corrects and rechecks each claim until resolved or a limit is reached, and escalates unresolved claims to human review. Evaluated on 4,033 clinician-labeled note-finding pairs covering six red-flag symptoms and family-history status, it raised precision from 0.764 to 0.849 and Matthews correlation coefficient from 0.681 to 0.730 over a single-agent baseline, while only 1.5% of claims required human review.

Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets

Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo Large language models are being deployed at scale in markets, content moderation, and hiring, raising the question of whether more capable individual models yield better system-level outcomes. The authors hypothesize that shared training data and architectures make capable models behave alike, producing correlated actions that do not diversify away, formalize this as a non-diversifiable risk floor, and test it in an agent-based market simulation with LLM traders of varying general capability. Frontier models show significantly correlated behavior that increases with capability; when their shared reasoning is accurate, adding agents reduces market-level risk, but under a shared misinformation environment the same correlation becomes a liability. The result is framed as a capability paradox, with generalization to other domains left as an open empirical question.

Conformity Breaks Conformal Prediction

Yibo Hu, Hanyu Su Conformal prediction certificates calibrated on an LLM answering alone can silently become invalid when the same model sees peers that unanimously assert a wrong answer, because the model's scoring of the correct answer shifts even though the question distribution does not. The authors call this a score-mechanism shift and measure it across open-weight models on multiple-choice question answering in multi-agent settings. Coverage falls from a calibrated 90% to 74% under unanimous-wrong peers at the standard alpha of 0.10, and an attacker targeting low-confidence items nearly halves coverage on that subgroup, from 87% to 47%, while the monitored average stays much higher. The failure reaches the decision layer, since a system that should escalate when uncertain can instead grow confident enough to act on the wrong answer, and standard conformal fixes do not help because the input distribution is unchanged.

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

Chenqian Le, Jiayi Cheng, Qijia He, Runhao Li, Yinghao Li, Xupeng Chen Multi-harness reinforcement learning (RL) for coding agents combines two choices: exposing the policy to several execution harnesses, and comparing their rewards inside one relative-advantage group. The authors isolate the second choice by replaying identical frozen task-harness records from Aider, OpenHands, Qwen Code, and SWE-agent from one Qwen3-8B warm start under two group-relative policy optimization (GRPO) rules, Within (one group per task-harness pair) and Cross (harnesses pooled per task), and score checkpoints with a sealed SWE-bench Verified oracle on the four source harnesses plus a held-out minimal harness. The evaluation harness moves mean solve rate from 2.14% to 9.27%, a factor of 4.3, while the training recipe moves it by only 1.16, and the grouping rule makes no measurable difference on the held-out harness, with Cross minus Within at +0.25 percentage points and seed-to-seed variation exceeding the gap. Cross-harness advantages leak which harness generated the data, so pooled credit yields configuration adaptation rather than more portable capability, and the authors recommend reporting the grouping boundary and testing on an unseen harness.

MaxKernel: Agentic Kernel Generation for TPUs

Shangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica et al. Writing high-performance custom kernels for accelerators demands deep hardware expertise, and large language models paired with real-time compiler feedback offer a way to automate it. MaxKernel is a multi-agent system for TPU kernel development with three modes: a Human-in-the-Loop (HITL) agent for step-by-step collaborative design, an Autonomous agent that runs a fully automated, metric- and trace-driven optimization loop, and a Graph-Based Autonomous Search that scales the autonomous agent to global exploration of the design space, all sharing sub-agents for planning, implementation, self-debugging, testing, and hardware profiling. Evaluated on JaxBench, a suite of 50 diverse TPU kernel tasks, plus real workloads from open-source models, the system consistently produces implementations matching expert hand-tuned baselines. The agent is open-sourced.

Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents

Lin Ai, Scott Counts Workplace agents must interpret hours of low-level human activity events that are too granular to reason over directly and lose structure when flattened into one stream or compressed into a single embedding. The authors build a multi-resolution vocabulary of semantically normalized operators, recurring motifs, coherent episodes, and day-level rhythms, and apply it to 667 million human-attributed events from 50,000 users across 100 organizations in a commercial productivity suite, yielding 120 operator types, thousands of motifs, 25 episode types, and five day-rhythm archetypes. Re-running the pipeline on a disjoint 2,000-user sample recovers the same taxonomy, and the full representation forecasts a user's next episode with a 17% relative macro-F1 gain over a flat-operator baseline. A resolution ablation shows no single level is best across agent-facing questions, so trace interpretation should be query-conditioned rather than reduced to one universal summary.

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

Marcel M\"uller, Jiaru Bai, Willi Gottstein, Abhijoy Mandal, Mohammad Nazeri, Elia Savino et al. Self-driving laboratories still depend on human specialists to translate scientific goals into closed-loop optimization campaigns and adjust them as data and conditions change. La Agente Óptima is an agentic framework that constructs and supervises Bayesian optimization campaigns across computational and experimental systems while maintaining a persistent optimization state, separating large language model reasoning from campaign execution so control returns to the agent only when interpretation or revision is needed and every decision stays auditable. Across ablations, five digital discovery tasks, and two physical platforms it kept campaigns executable as problems and environments evolved, detecting and correcting a mid-run measurement failure in a contact-angle campaign and then correctly inferring the target was likely unreachable with the available reagents. In a five-day multi-objective flow-chemistry campaign it raised yield from 30% to 59% over 23 experiments, at lower cost and with substantially less starting material than a human-directed campaign.

Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics

Gnaneswar Villuri, Hashmath Shaik, Alex Doboli cross-listed Large language model code generation breaks down when one routine's correctness depends on the runtime behavior of another, a limitation the authors call static binding that appears in cross-coupled optimizers, packing, routing, and symbolic search. Their dynamic context adaptation method runs a validation-generation loop in which a validation agent extracts structured diagnostics from execution traces to guide a generation agent that proposes multiple candidates per iteration, with a knowledge graph built from the problem description supplying semantic constraints and simulated annealing selecting among candidates to avoid greedy collapse. It outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), a regime where population-based search has not yet built enough diversity, and also wins at 1000 evaluations on the motivating cross-coupled optimization problem. Ablations identify structured execution feedback as the primary driver.

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

Quan Shi, Keshav Dhandhania, Karthik Narasimhan, Victor Barres Coding agents are increasingly asked to build production LLM agents, but existing benchmarks do not measure whether they can deliver one under the conditions of a real client engagement. τ^τ-bench (hyper-tau-bench) hands a developer agent a business's actual records, a client who holds the requirements, a production API, an inherited codebase, and limits on serving cost and models, then scores the customer-service agent it builds by deploying it against held-out simulated users across 53 tasks in four domains. The strongest configuration, Claude Opus 5 running under Claude Code, passes only 23.9% of evaluation simulations, against an expert-authored reference ceiling of 82.2%. The failures mirror those seen with human agent developers: shallow queries instead of deep comprehension of the records, almost no communication with the client, and too little experimentation with architecture or serving spend before shipping the first design that runs.

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou Runtime safety gates for LLM tool agents are usually treated as filters, but in a ReAct loop a rejected action is followed by another proposal from the same state, so the gate actually shapes which trajectories are reachable. SiLR targets recovery after a constraint violation, where progress must be admitted while the system is still unsafe, and argues that gates based on a single aggregate score fall into a scalar projection trap that accepts locally improving actions and strands the trajectory on a plateau; it instead shadow-executes each proposal in a deterministic simulator and admits it under a product order over per-branch violation state, with a proof that no scalar surrogate is sound for that order. On mined Gym-ANM power-grid scenarios, SiLR recovers 21 of 21 multi-action episodes versus 0 of 21 for a terminal gate and 9 of 21 for the best scalar gate, a pattern that holds across three model families and in CityLearn, and only the full per-branch predicate contains a magnitude-redistribution attack that defeats scalar and support-only baselines. Reused as a process reward for GRPO, the same structured signal beats its scalar count projection in every scenario and is the only tested reward whose ungated policy exceeds the untrained base.

A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark

Yoga Sri Varshan Varadharajan, Ajay Yadav, Ritesh Goru, Prateek Chaudhury, Constantine Caramanis, Prateek Jain et al. Natural-language-to-SQL systems perform well on academic benchmarks, but production enterprise schemas have graph-like, semi-structured, deeply nested structure that those benchmarks do not measure. The work introduces the DevRev NL2SQL benchmark, 900 execution-verified queries over nested types and link graphs, together with a schema-agnostic Semantic Depth Score (SDS) rubric for analytical reasoning depth, plus a cost-aware single-generation agentic architecture whose schema-selection, metadata-retrieval, and error-repair components are built for this setting. The system reaches 91.7% answer correctness on the new benchmark, 54.6 percentage points above the next-best baseline, and is competitive with leading systems on the Spider 2.0 Snowflake public dataset at a single-generation operating point.

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

Xinran Zhang, Pengrui Lu, Lyumanshan Ye, Pengfei Liu LLM agents are increasingly proposed for enterprise workflows, but existing evaluations rarely test whether conclusions about their business decisions hold up across different competitive settings. ERPBench is an execution-instrumented benchmark built on a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition, evaluating the same 100 fixed problems in two matched ecologies: Solo, where each agent faces fixed rule-based opponents, and Arena, where six LLM agents compete in one shared market. Across six model families and 1,200 trajectories, the leading model differs by ecology, with DeepSeek winning in Solo and Gemini in Arena, and the two settings agree on the task-level winner for only 21 of 100 problems; Gemini's bottom-rank rate also falls from 22% to 0% when moving from Solo to Arena. Code and benchmark resources are released.

Train What You Deploy:Token-Faithful Post-Training of a Production Coding

Cheng Li, Jiexiong Liu, Yixuan Chen, Chi Hong Post-training pipelines for coding and terminal agents commonly train in simplified environments that differ from production deployments and reconstruct tokens offline from agent logs, which distorts the original prompts and conflates policy calls with background model operations. The proposed fidelity-aware training coupling keeps sampling on the trainer side over the original prompts, removes spurious model calls through a negotiated training protocol, and restricts the loss to verifiable token spans with closed-failure guarantees. On top of this, Certified Divergence Proximal Policy Optimization (C-DPPO) adds two-sided total-variation certification bounds, adaptive-K rules, budget-aware sequence guarantees, and error-robust policy masking to standard DPPO. On matched Baize5B and Baize10B models evaluated on TMax-100, C-DPPO delivers a consistent +3.0-point gain over standard DPPO at both scales, and certificate audits confirm full operational coverage of the pipeline.

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

Happy Bhati cross-listed AI coding systems are moving from autocomplete and chat toward agents that inspect repositories, edit many files, run tools, write tests, and open pull requests with limited supervision, yet field evidence shows the resulting gains in coding activity shrink sharply between writing code and shipping reliable software, while costs move from predictable per-seat licenses to variable token, tool, sandbox, CI, and rework spend. Drawing on peer-reviewed software-engineering research, benchmark audits, production reports from major technology companies, developer telemetry, and cost data published mainly from 2024 through September 2026, and claiming no new model experiments, the synthesis proposes four engineering concepts: the Agentic SDLC Throughput Paradox, Production-Qualified Change (PQC), the Verification Tax, and an Agentic SDLC Control Plane that allocates autonomy within cost, reliability, and human-attention budgets. The central question is reframed from how much code an agent can generate to how much production-qualified value an engineering system delivers per dollar, per reviewer-hour, and per unit of operational risk, with an evidence-based horizon mapping today's supervised agents toward policy-bounded software factories.

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

Abhishek Sharma A merchant's payment processor, ledger, ERP, and bank feed can hold contradictory views of the same order for minutes because their update messages are delayed, duplicated, dropped, or reordered, and an agent resolving the exception must choose among actions like shipping, re-capturing, or refunding that cannot be undone. FinalityBench keeps a hidden canonical event log, derives each system's view from a separately faulted delivery stream, and scores each episode by the merchant's terminal economic position relative to a privileged reference, across 321 tasks including 45 twin pairs whose system views and authoritative probes are identical at decision time yet whose correct dispositions differ. Over 14,445 graded episodes from nine programmatic policies, a ship-on-first-signal policy ranks second by accuracy at 65.7% but worst by paired loss, a runtime that gates irreversible actions on an authoritative finality probe reaches 85.4% and loses nothing under pass^5, and language models match the gate's exact rate on a stratified subset while losing about twice as much money, discovering the gating strategy without being told it.

Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada cross-listed Coding agents make it fast to build research software but also raise the need for better discovery and maintenance of that software. The authors built a repository catalog with a coding agent during a three-day hackathon, then documented the additional engineering needed for public deployment, including adversarial review, data-quality checks, browser-level validation, and publication safeguards, and explored transferring the lessons to a retrieval agent for the MateriApps portal that combines curated metadata, external documentation, vector search, and local language-model generation. The central observation is that the most consequential problems were silent failures producing plausible but incomplete or incorrect outputs from data acquisition, assessment, and retrieval or preprocessing errors rather than crashes, which argues for explicit validation, monitoring, repeated review, and continued reliance on curated metadata and maintained documentation.

DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

Zehao Wang, Lanjun Wang, Shilong Jin, Junjie Chen, Yanghua Xiao Large language model (LLM)-based multi-agent systems often fail because of a single decisive error buried in long natural-language traces, and existing attribution methods tend to flag minor deviations or lose accuracy as traces grow. DCFA is a training-free framework that builds a causal-inspired dependency graph over the whole trace to locate the earliest decisive error, then applies local counterfactual-style reasoning to refine that attribution. Evaluated on the Who&When benchmark with six different LLMs, it improves step-level attribution accuracy by up to 8.27% over state-of-the-art baselines.

Persistent Teacher Anchoring for Tool-Using Agents

Hyun Bin Park (Sogang University), Kyungho Song (University of Michigan, Ann Arbor), Sangmin Lee (Sogang University), Du-Seong Chang (Sogang University) On-policy knowledge distillation (OPKD) trains a student on its own trajectories with teacher-supplied token distributions, but in tool-using agents the student's calls execute before the teacher weighs in, so errors compound through the observations they produce. Persistent Teacher Anchoring (PTA) keeps chunk-level teacher verification from proposer-verifier generation and adds turn-level commitment, so a tool call reaches the environment only after the teacher has verified the entire turn; a persistent lookahead scheme fills idle rollout capacity by advancing future samples across student updates. Used before downstream RL in Search-R1-style retrieval and DeepEyes-style perception settings, PTA improves macro best@4 by 2.5 and 2.8 points over OPKD at the same RL budget, and lookahead raises throughput by 24%.

MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate

Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India) et al. Media bias works through subtle cues such as loaded language, selective framing, and omission, which single models struggle to detect and which have required large annotated corpora for supervised training. MABPD has three specialized LLM agents analyze an article from complementary perspectives and settle disagreements through a Structured Argument Debate (SAD) protocol that gives zero weight to bias claims lacking grounded textual evidence, applies role-weighted voting, and verifies the consensus afterward. Without any task-specific training or threshold tuning, it reaches 83.4% macro F1 on the BABE benchmark, within 0.7 points of the supervised state of the art MAGPIE, and 75.0% zero-shot accuracy on the SemEval 2019 HyperPartisan corpus. Removing the debate module drops F1 by up to 10.6 points, indicating that structured deliberation rather than agent parallelism drives performance; the pipeline and evaluation code are released.

ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

Weide Zhan, Qumu Shaqu, Yuanqing Liu, Peng Zhang, Jiahao Liu, Kam Him Lam et al. Mobile GUI agents could help older adults use smartphones, but existing benchmarks use explicit goal-oriented instructions that miss how older users actually speak: indirect requests, referential ambiguity, and under-specification. ElderBench is built from 249 naturally elicited smartphone tasks collected from older adults across 20 applications, and the authors characterize how these instructions diverge syntactically, semantically, and pragmatically from existing GUI benchmark instructions. Evaluating mainstream GUI agents and vision-language models in online and offline settings shows substantial performance degradation on elderly-oriented instructions, and controlled instruction normalization, failure analysis, and linguistic feature analysis pin down which language patterns cause the failures, yielding design guidance for more age-inclusive agents.

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

Di Chai, Leye Wang, Zeshen Su, Zhiguo Xia, Zhihang Yu LLM agents in persistent workspaces accumulate history that exceeds both GPU key-value (KV) cache capacity and the model's native context window, and existing systems either compact old context into lossy summaries or re-prefill retrieved text the model already processed. KVMem virtualizes KV context by paging overflowed workspace history across GPU memory, host memory, and NVMe, using lightweight model-native attention-space indexes to select relevant historical blocks and assemble a query-dependent execution view that fits within the native context window. On LongMemEval, MemoryAgentBench, and AgentLongBench with histories up to one million tokens it generally beats compaction-based approaches in task utility and efficiency, and on the DeepSWE long-context test with Qwen3.8-27B it lifts task success from 43.8% to 48.4%. On a laptop with a 24 GB RTX 5090 GPU, it runs Qwen3.6/3.8-27B in NVFP4 with multi-token prediction over 1M-token workspaces, four times the native 256K context, at roughly 50 tokens per second.

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

Jinyuan Feng, Dongmin Li, Yiqun Chen, Yang Gao, Xing Chen, Huimu Wang et al. Skill libraries let large language model (LLM) agents reuse procedural knowledge, but existing designs either evolve skills separately from policy optimization or freeze meta-skills into fixed workflows, treating skills as passive objects to manage. CoSkill recasts the meta-skill workflow as a second trainable agent and co-trains it with the reasoning agent on a shared backbone over a hierarchical skill library, so the reasoning agent conditions on a retrieved task skill and selected step skills while its task performance guides refinement of those step skills. On ALFWorld and WebShop it reaches success rates of 98.4% and 90.6%, improvements of 3.5 and 6.2 percentage points over prior skill-based and reinforcement learning baselines, alongside better early sample efficiency and wall-clock time.

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

Longtao Hu, Xiao Liang, Linchao Zhu Computer-use agents discard most of what they learn, since procedural knowledge from one rollout is not retained, refined, or reused later. The framework here converts interaction trajectories and evaluator feedback into a persistent, versioned library of reusable procedures, with each iteration executing against a frozen library snapshot so evidence-guided updates only take effect afterward and no model weights change. Measured against a configuration-matched empty-library control across four OSWorld application domains under an identical action-generation and grounding stack, the evolving library raised post-warm-up mean evaluator scores by 5.7 to 18.6 percentage points in all four domains. Benefits were domain-dependent, and provenance analysis in GIMP found skills retrieved across task boundaries plus revision churn where repeated accepted edits failed to recover the originating task.

AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems

Qi Zhang, Yanlin Chen, Wenchao Xiao Shipping a recommender improvement at NetEase's DASHEN gaming-community app runs from reading papers, through reproduction and production implementation, to offline evaluation, online A/B tests, and an internal launch review gate, a cycle spanning days that normally needs a human at every handoff. AutoLR wraps it in a harness with three mechanisms: a multi-expert council that debates and adversarially reviews proposals, a deterministic evidence-weighted exploration-exploitation selector that spreads a limited trial budget across candidate directions, and a layered knowledge system mixing external research, production-system facts, and app-specific domain knowledge with posterior evidence from configs, patches, logs, and failures. Language model agents handle semantic reasoning and code generation while deterministic controllers retain authority over execution, metric extraction, guardrails, and persistent state transitions.

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

Linsen Zhu, Mengqing Cai Language models become consequential agents once surrounding systems let their outputs change external state, and the usual narrative of one march toward autonomy conflates four separable things: model competence, harness integration, temporal persistence, and safe authority. Synthesizing primary research and official technical specifications available through 31 August 2026, the review organizes evidence by delegated authority, persistence, and environmental coupling while keeping model, harness, and environment distinct. Its central reading is that action-interface expansion is documented far more convincingly than robust task completion, recovery, authorization, or independent verification: Model Context Protocol and Agent2Agent improve interoperability without establishing trustworthy delegation, multi-agent organization buys specialization at the cost of correlated failure, and robotics or self-driving laboratories demonstrate bounded feasibility rather than unattended open-world reliability. The authors offer justified delegation as a heuristic, widening action scope only where provenance, bounded authority, failure detection, safe recovery, and calibrated human control are evidenced.

RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents

Aziz Ben Amor, Drish Mali, Mann Acharya, Vijayasri Iyer, S\'ebastien Brati\`eres Repository-scale refactoring requires a coding agent to propagate one change across many interdependent files without altering behavior, and no existing harness isolates which design choices determine success. RefactorPlatform holds the environment fixed while varying model backbone, execution regime (baseline, retrieval-augmented, multi-agent), and prompt specificity, running each task in an isolated workspace with per-task logging of tokens, diffs, and transcripts plus abstract-syntax-tree verification. Across 100 multi-file RefactorBench tasks and four model families, syntax-tree-aware chunking beat naive token-window chunking by 25-30% in all prompt modes while naive retrieval fell below the retrieval-free baseline, and a lean retrieval-augmented single agent solved 86% of matched tasks against 66% for the evaluated sub-agent configuration, with no task passing under delegation that failed under retrieval. Retrieval's accuracy gains absorbed its token overhead, leaving cost per successful refactoring unchanged.

ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems

Ant\'onio Azevedo, Bruno Lima, Jo\~ao Pascoal Faria cross-listed Validating automotive infotainment software is still largely manual, and scripted automation is brittle, while existing LLM test frameworks target web and mobile apps with one or two agents that must handle perception, planning, action, and validation at once. ARIA (Autonomous Real-time Infotainment Assessment) is a multi-agent LLM framework that drives Android infotainment systems through visual interaction using a closed loop of four specialized agents per step plus a reporting stage, turning single-sentence scenarios into executed tests with reports, reproducible scripts, and per-step visual evidence. On a manufacturer's physical system across 30 scenarios it delivered verdicts for 28, 20 of them matching ground truth, and caught all five known defects with no fault passing as working, though it produced eight false positives from navigation, image, and gesture limitations. A single-agent baseline showed a much higher first-pass false-positive rate (72.0% vs. 52.6%), and repeated runs showed fault detection was perfectly consistent while overall stability tracked scenario complexity.

Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

Jiahe Geng, Jinpeng Wang, Kun Yuan Long-horizon LLM deployments often cannot afford to prompt with the full interaction history, so the question is which memory design gives the best quality per token under a tight budget. RSM-full is an online clustered-memory pipeline that combines a cosine-gated max-member merge rule for writing memories with an atom-aware grouped packer for assembling retrieved content into context. On AMA-Bench it reaches 83% of full-context quality at 32% of the token cost with a 4k budget and beats the closest streaming-clustered baseline, Online K-Means, by 3.5 to 6.0 percentage points across the roughly 2.6k to 5k token regime, with ablations attributing most of the gain to the merge rule and the packer. On the independent RealMem benchmark it edges Budget-RAG, matches BM25-RAG, and outperforms Streaming-Proto and A-MEM, though the authors note higher-token baselines remain stronger outside the roughly 2k to 5k token regime.

TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

Tianxing Wang, Mingming Zhao, Shuai Huang, Huiyang Xu, Chaoyue Niu, Shengzhong Liu et al. Agent systems typically optimize or constrain their execution structure before decisive runtime outcomes are observed, so when intermediate evidence invalidates the planned continuation they must either execute stale steps or replan broadly, compounding errors and discarding progress. Trace-grounded Route Orchestration via Validation and Editing (TROVE) distills evaluated workflow-search traces offline into atomic and composite skills plus an outcome-conditioned transition graph, then treats a planned route as provisional online: after committing one top-level skill, the controller retains a still-valid continuation, inserts a trace-supported local response, or replaces only the invalidated suffix. Across code generation, question answering, and math reasoning benchmarks with different LLM backbones, TROVE delivers a stronger quality-efficiency trade-off than dataset-level optimization, query-level architecture selection, and graph-constrained scheduling baselines, with the largest quality gains when outcomes change the appropriate continuation and large efficiency gains from early termination on near-saturated tasks. Ablations show composite skills capture most of the offline benefit, insertion enables local correction, and suffix replacement mainly improves efficiency.

A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support

Chang Xia, Leilei Ouyang, Huimin Wang, Yong Zhao, Kang Li The single-turn question-answer format does not reflect how clinical diagnosis is performed in practice, which limits large language models in complex diagnostic settings. Debate-Mixture-of-Agents (DMoA) is a multi-agent framework that structures role-based interaction among models to support iterative diagnostic reasoning. Evaluated on 297 rare disease cases and 1,719 challenging cases, DMoA improved most-likely-diagnosis accuracy by 10.21 percentage points and safety rate by 11.36 percentage points over a GPT-4o baseline, and ablations show the gains are not simply due to using more models or producing longer outputs but reflect the structured workflow itself. Further analyses find DMoA performs best with a 4x2 structure, stronger base models, and a larger token budget.

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang Existing benchmarks for AI-scientist coding agents reward reproducing a hidden target study, which measures execution rather than discovery. TruthInsightBench instead gives agents 40 blind tasks drawn from 40 peer-reviewed studies across 10 domains, exposing only a neutral objective and frozen data while withholding source conclusions, expected values, and analysis paths, and a fixed LLM judge scores the evidentiary maturity of the agent's own claims along six dimensions via 29 artifact-grounded items with deterministic aggregation. On one frozen base model, four coding agents cluster in a narrow band of 58.4 to 60.3 out of 100 with no statistically reliable pairwise separation: they execute and document analyses competently but rarely perform the controls, robustness checks, falsification attempts, and cross-dataset generalization that would establish a trustworthy claim, pointing to scientific judgment rather than coding as the bottleneck.

LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28

Wes Sander Discovery Loop is a lightweight system in which a large language model iteratively evolves an optimization algorithm: starting from a simple seed solver, the model proposes improvements guided by a scoreboard and a history of prior ideas, an independent verifier evaluates each candidate, and only improvements are kept. On the Packomania circle-packing benchmark, where the goal is to maximize the sum of radii of N variable-radius circles in the unit square, the system beat the best known solutions for 10 values of N between 101 and 114 by 2.4% to 5.4%, within 15 iterations and for a total LLM cost of $27.72, and the records were independently accepted by Packomania. The authors also analyze cost-efficiency dynamics, including an adaptive plateau-detection mechanism.

Substrate-Aware AI Agents: Execution Context as a First-Class Input

Manu Agrawal Agents that generate code or plans usually never see the memory, time, and runtime limits of the environment they will run in, a gap the authors call substrate blindness. They test whether a minimal execution contract changes behavior by having Claude Opus 5, GPT-5.6-Sol, and Gemini 3.7 Flash write code for a high-dimensional pairwise Euclidean-distance task either from the task alone or with an explicit 128 MB RAM and 10-second wall-time budget. Disclosing the contract cut peak memory in 13 of 14 paired comparisons and mean wall time in all three cohorts, with execution up to 3.1x faster, and produced structural changes such as bounded blocking, float32 retention, upper-triangle traversal, and memory-mapped buffers. Under a tighter 96 MB budget, contract-aware runs were correct and within budget in 4/5, 5/5, and 3/5 samples versus 0/5, 1/5, and 0/5 without the contract.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization

Sihan Ge, Yichen Lin, Chenyu Zhou, Jianghao Lin, Tao Yao, Dongdong Ge cross-listed LLMs are increasingly asked to turn natural-language descriptions into operations research (OR) optimization models, but real requests often omit objectives, constraints, or business rules that change the resulting program, and existing benchmarks assume complete specifications. OR-Clarify presents partial problem descriptions with structured hidden slots and scores agents on slot recovery, stopping behavior, silent assumptions, and interaction cost through bounded dialogue with a simulated user, in both open-ended and choice-based clarification modes. The proposed InterOPT framework first identifies unresolved formulation-critical gaps and then uses them to decide whether to ask another question or stop; in choice-based experiments it substantially outperforms all baselines on exact slot recovery and stays competitive with strong prior methods in the open-ended setting.

Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

Jiazheng Sun, Boyu Yang, Binhao Yuan, Mingxuan Li, Xin Peng Agents that learn from execution traces mostly retrieve similar trajectories or summarize flat skill lists, ignoring the temporal ordering and success-or-failure structure of behavior. Trace2Tower abstracts step-level interactions into canonical events, links them in a graph weighted by semantic compatibility, transition dynamics, and outcome evidence, and applies a contrastive spectral decomposition to isolate stable success-aligned behavioral modes while suppressing failure-prone shortcuts. Those modes populate a three-level skill tower of action templates, procedural routines, and task strategies that is refined with verifier feedback. On ALFWorld it reaches 87.31% success in 10.35 steps with 0.26 invalid actions, and on WebShop 50.67% exact success, outperforming existing baselines on both.

CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls

Chris Zheng, Geng Yang cross-listed Agent stacks combine provenance tracking, authorization, policy enforcement, protocol adapters, and execution controls, yet individually sound mechanisms can drop, widen, rebind, or reinterpret security-critical context as an action crosses component boundaries, a failure the authors name security-context discontinuity. CONTINUITY gives each component an assume-guarantee contract and threads authenticated context across transitions with signed root grants, provenance commitments, role-bound transition receipts, bounded typed releases, transformation witnesses, and effect-bound execution permits, formalizing an end-to-end consequence-integrity property that requires every external effect to be backed by a valid, current authorization witness. A reference verifier and a deterministic cross-layer fault-injection suite of 32 fault classes over four domains show that the full configuration commits no harmful external effect across 2,560 attack instances while completing all 700 benign tasks and escalating all 200 ambiguous cases.

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

Konstantin Grotov, Valentin Malykh Software engineering agents fail expensively by acting confidently on wrong plans that are only caught after execution and retries. Speculative Uncertainty (SU) inverts speculative decoding: a small open-weight draft model scores a black-box agent's already-generated trajectory in one forward pass using only output tokens, with no logits, weights, activations, or resampling, and phase-aware features that separate reasoning spans from action spans are calibrated against a verifiable objective to produce a failure-likelihood score any downstream policy can consume. Used as a pre-execution veto gate on Qwen3-Coder-480B and Claude 3.5 Sonnet, the signal cuts execution error rate by 6-8 percentage points and token cost by 14-19%, transfers to out-of-distribution benchmarks without retraining, and generalizes across agent models.

Testing Interchangeability in LLM Agent Teams

Jianxin Gao, Tianyi Yu, Linna Deng, Runze Li, Zining Wang Production multi-agent systems swap agents in and out on the assumption that any agent capable of a role can fill it. The authors form eight independent teams per setting from a single base model, let each agent keep a private notebook over ten formation episodes, then trade role-matched agents between teams and compare held-out performance against a placebo that reproduces the disruption of a roster change without changing who occupies the seat. A swap barely moves task score but raises communication spent per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent costs more than an inexperienced one, consistent with interference from conventions learned with a former partner; in Collab-Overcooked, replacing the agenda-setting agent shifts most of the extra talk onto the agent that stayed. Ablations over base model, decoding temperature, and formation length move the swap penalty in lockstep with how far independently formed teams drift apart, with greedy decoding lowering both and longer histories raising both.

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Ankit Goyal, Jaideep Ray Swapping the model behind an agent while keeping its memory store can still cause forgetting, since a new model may read old notes differently, mixed embedding versions can break retrieval, and repair may be impossible without the original evidence. The study compares four memory formats over the same history: verbatim long-context reading (LC-RAW), chunked retrieval-augmented generation (RAG), model-compressed natural-language notes (NOTES), and a fixed-schema knowledge graph (KG-fixed), using 48 synthetic histories with randomized answer codes, exact scoring, and two open-weight models under 10 billion parameters. Fixed-schema knowledge graphs transfer almost perfectly across a writer swap, while compressed notes shift accuracy asymmetrically by +9.91 or -13.28 percentage points depending on migration direction, and a 50/50 mixed embedding index in RAG captures only 4.96 of the 11.90 points gained by full re-embedding. Decomposition attributes 80% of the notes deficit to information lost at construction and 81% of the RAG deficit to retrieval failures, and store-only repair of notes never reaches a 90% recovery target, whereas retaining raw history enables recovery in 34 of 48 cases for one direction.

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar, Liqun Cheng, Ming Liu, Parthasarathy Ranganathan et al. cross-listed Performance-modeling frameworks for machine-learning systems rot quickly because each new model or hardware generation invalidates their baked-in assumptions, and the authors argue AI coding agents are now fast and capable enough that regenerating a whole library is cheaper than paying down its tech debt. SMART is a symbolic performance-modeling library whose main branch holds almost no code: the repository is a directed acyclic graph of self-contained natural-language design docs, coding sub-agents regenerate the implementation from those docs on each version update, and every human change is a doc edit. Regeneration is kept reliable by a doc style built around step-by-step worked examples that act as in-context demonstrations, plus a minimal recursively defined operator intermediate representation with SymPy cost expressions, a fast analytical roll-up mode for large sweeps, and a slower modulo-scheduling mode for fine-grained schedule studies. Regenerated implementations reproduce hand-audited reference models, including DeepSeek-V3 serving on a TPU pod slice, to round-off precision, which the authors take as evidence that design docs rather than code can be the durable artifact for ML-systems co-design tools.

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

Haoting Shi, Wenhao Wang, Weicheng Fang, Yaozhong Liang, Tian Jin, Pengxiang Zhao et al. Computer-use agents mostly act through the graphical user interface (GUI) and produce inefficient trajectories, while real computer work mixes visual state inspection with fast, precise command-line interface (CLI) operations over the same application state. CUA-Universe is an environment-to-data pipeline that turns real desktop software into hybrid GUI+CLI environments: App-Forge packages applications into reproducible virtual machines with command-line surfaces it discovers, wraps, or generates, scaling to 16 applications; Task-Weave synthesizes hybrid tasks of controllable difficulty from reusable operations over seed files; and Path-Steer steers rollouts along efficient hybrid paths and harvests verified trajectories for post-training. Training on this data shifts a 9B model away from clumsy GUI interaction and brittle CLI scripting toward coordinated use of both interfaces. The model gains 16.8 points of success rate on OSWorld while using 57% fewer steps and 44% fewer tokens, with comparable improvements in score and efficiency on CUA-Verse and OSWorld-MCP.

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim Data-sovereignty rules push public institutions toward open-source, on-premise LLM agents that chain multiple tool calls across live government APIs, a setting where open-source models consistently lag and no benchmark existed to measure the gap. KOPA-Bench (Korean Open Public API Benchmark) supplies 145 real-world multi-step tasks over live Korean public APIs. To close the gap, EDGE (Execution-grounded Dynamic Graph for tool-calling data synthesis) builds a graph of which tool outputs can feed which tool inputs, keeps only the links that succeed when actually called against the live APIs, and traverses those verified links to synthesize executable multi-step trajectories. A 9B model fine-tuned with GRPO on the resulting data nearly matches the untuned 27B model from the same family, with substantial gains on both KOPA-Bench and BFCL.
2 more specialized papers

Safety & Alignment 24

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka cross-listed Large language models remain vulnerable to jailbreak prompts, and many defenses need access to weights, which rules them out for black-box deployments. AlcaTRAz operates only on input text: it learns a transferable rule tree that inserts controlled character-level perturbations at selected positions, disrupting the structural regularities jailbreaks rely on while largely preserving behavior on benign queries. Across 33 open-weight models and 22 attack types, compared against Llama Guard, RA-LLM, and Goal Prioritization, it achieves the best composite security-and-functionality score in 73.4% of model-attack combinations, shifting the modal response severity from 10 to 2 while keeping the mean benign score within 0.27 points of the undefended baseline. A high-severity tail remains and adaptive attackers are not considered, so the authors position it as one layer in a defense-in-depth strategy.

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi et al. cross-listed Vision-language models (VLMs) deployed in high-stakes settings may give responses that are reasonable in general but unsafe for a specific user whose medical, emotional, or situational context is hidden. MPS-Bench contains 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile; eight frontier VLMs respond directly 86 to 99% of the time rather than seeking missing context, and none scores above 2.6 out of 5 on personalized safety. Mechanistic analysis identifies visual dominance, a two-stage process in which visual affect enters the text stream in early layers and then drives the final decision through the altered text representation, making late-layer intervention unreliable. PRISM, a lightweight input monitor using bidirectional cross-modal modulation to predict when deferral is needed, reaches 0.978 AUC and dominates the safety-utility Pareto frontier for all tested models.

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton Explanations produced by large language models (LLMs) can be unfaithful to the reasoning that actually drove the answer, and the authors split this into incompleteness, where influential factors go unmentioned, and unsoundness, where cited factors had no influence. Existing training-time fixes need weight access and heavy compute, while test-time methods mostly target unsoundness, so this work proposes a test-time approach aimed at incompleteness: concepts the model's explanation does not credit are removed from the input, and the model is re-queried on the reduced input so that unmentioned influences are eliminated while credited ones remain. Across two datasets, multiple model families, and two independent faithfulness metrics, the method improves explanation faithfulness over both standard prompting and prompting that encourages faithfulness. The approach is model-agnostic and needs no parameter changes.

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys

Georgios Politis, Evangelos Pappas cross-listed In a two-node split-LLM training system, a Trusted Local Node (TLN) sends protected activations mixed with decoy rows to an Untrusted Cloud Node (UCN), which returns its output, and the TLN, holding the private loss, sends back the output gradient. Because the loss ignores decoys, their gradients are exactly zero, so the pattern of zeros reveals which rows are real. Using a protocol fixed in advance with an injected leak, a shuffled-label control, and a preset threshold, the zeros identified every real row on every frame across nine seeds, 4,096 of 4,096 per run, even though each run passed the forward-channel privacy check and the quality check; a content attack recovered only about one extra token per hundred over a constant-guess baseline. Per-row gradient clipping and noising closed the leak for about 0.01 nats of held-out cross-entropy, though five further attack classes, including those accumulating observations across training steps, were never measured.

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

Jasmine Brazilek, Miles Tidmarsh, Matthias Endres, Anshuman Singh, Jeremiah Miller Existing side-effect benchmarks neither put a price on avoiding harm nor name the harmed party as a living creature. HarvestBench is a memoryless reinforcement learning gridworld in which LLM sub-agents drive two tractors through a cooperative corn harvest; when an animal blocks a route, the autopilot asks the model whether to drive on for free or swerve for a posted fuel cost, with rocks and hay bales as controls and the neighbor's crops as a second moral test, all scored by counting events in the game log rather than by an LLM grader. Across nine models and 7,201 priced decisions, kill rates ranged from 0.4% to 98.8% and were not ordered by capability, with Terra and Sol the most merciful and GPT-4o-mini the least. Four of six models were sensitive to price, every model ran over wild animals more often than farmed ones, and the briefing mattered most: a morality briefing kept kill rates under 6% in five of six reasoning models, while removing it pushed them above 84% in all six.

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

Alejo L\'opez-\'Avila, Iker Garc\'ia-Ferrero, Jezabel Garcia, Antonio Tiene, Rom\'an Or\'us Safety alignment is usually framed at the topic level, but deployments need narrower boundaries inside a topic, such as a civics tutor refusing targeted political manipulation while still answering factual election questions. The authors formulate this as narrow-boundary safety and build an offline self-generated data framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs, where escalating retries leave only 0.20% of prompts without an accepted refusal trace versus 19.88% for single-shot generation. On political persuasion with Qwen3-8B, training on this data raises target-domain refusal from 9.47% to 84.75% and cuts the mean unsafe-response rate across three harmfulness benchmarks from 26.26% to 0.14%, but pushes XSTest over-refusal from 2.00% to 74.00%. Replacing external responses with verified target-model responses drops over-refusal from 15.20% to 5.20%, and boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16% while harmful-side refusal falls only from 91.88% to 87.72%, showing that data composition governs the safety-usability trade-off.

Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning

Antoni Czolgowski, Abel Iyasele Three open-weight LLMs, Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China, are compared against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to measure distributional misalignment. No model favors its home country, and the Chinese-built Qwen3-4B is worst on its own Chinese population, the highest misalignment in the entire model-by-country matrix. Targeted LoRA fine-tuning on the five worst-case personas, using fewer than 1,200 training pairs and under 15 minutes on one GPU, reduces bias by 16.8% for Bielik-11B with all five targets improving. Country-level decomposition, however, shows the fine-tuning redistributes rather than removes bias, since the model's worst-case personas swap entirely from American to Chinese elderly with no overlap between the pre- and post-correction sets.

Rethinking Indirect Prompt Injection as a Test-Time Search Problem

Duong M. Nguyen, Joon Sik Kim, Blazej Manczak, Vaikkunth Mugunthan Indirect prompt injection against tool-using agents is reframed as a test-time search over an attack surface determined jointly by the environment, the user task, and the injection goal. The authors build an agentic attacker with a dedicated search harness that performs environment reconnaissance, reasons explicitly over attack strategies, and adapts using feedback from the victim agent. Across heterogeneous tasks, giving the attacker more test-time compute steadily improves vulnerability discovery and exploitation, and ablations show explicit strategy management is needed to avoid redundant search and sustain gains at larger budgets. The takeaway is that agent security evaluations should report the attacker's search procedure and compute budget rather than treating attack success as a fixed property of the victim.

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner et al. cross-listed Black-box prompt injection in text already reaches near-perfect attack success rates (ASRs), but visual prompt injection against frontier commercial vision-language models (VLMs) has struggled to elicit materially harmful outputs, which require long, format-compliant targets such as exact parseable tool calls. Repeat-After-Me is a black-box adaptive visual prompt injection attack that can exfiltrate personally identifiable information or trigger malicious tool calls, tested in a setting where the benign user prompt is unrelated to and does not authorize the injected task. It achieves ASRs above 80% on Qwen3.6-27B and above 47% on GPT-5.5, with surrogate-optimized injections retaining 43-46% of their ASR on two commercial victims and cross-sample transfer retaining 64-66%. In a default OpenClaw Discord deployment, a minimally injected image from an untrusted user overwrites TOOLS.md, enabling later remote code execution and secret exfiltration, in cases where adaptive textual injection fails.

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

Rongxin Yang, Yang Liu, Shang Luo, Haoxuan Jia, Chongyang Zhang, Hao Zheng et al. Self-evolving language models propose candidate updates and keep whichever raises a visible score, so when that score is an imperfect proxy for the desired capability, sustained selection widens the gap between the two, which is reward hacking. HackProbe is a monitor that attaches to any such loop through two black-box hooks, with no access to weights or activations; it keeps a secret, distribution-fixed comparison core whose capability proxy stays comparable across generations, plus a rotating fresh layer that resists co-adaptation, and combines four tests on that proxy (level gap, scale-aligned divergence with online change-point detection, capability stagnation, and a conditional confidently-wrong rate) into a calibrated family-wise p-value via a Sidak correction. Because diagnosis alone recovers nothing, a risk-aware immunization layer reselects an honest candidate from the proposal pool while disclosing only a bounded number of bits per generation, and the authors prove a detectability bound linking a target error rate to a probe-size budget. On a controlled prompt-level host with four injected hacking channels, HackProbe reaches 0.763 AUROC against 0.663 for the strongest baseline and cuts the false-positive rate from 0.706 to 0.434, and its bandwidth-limited reselection is the only immunization level that recovers more true capability under hacking than it forfeits on clean runs, though per-channel effects are mostly not individually significant.

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

Minji Kim, Hyounghun Kim Safety-tuned language models often refuse benign queries that merely resemble harmful ones, such as asking where to shoot a good photo. The authors decompose responses in safety-tuning data into a boilerplate refusal statement and a rationale explaining the refusal, and find that the refusal statements push models to rely on superficial lexical cues, impairing discrimination between harmful and benign inputs. Training on rationales alone reduces false refusals while keeping safety performance comparable, and the benefit also appears in an in-context learning setup and remains compatible with the inference-time mitigation methods evaluated.

Locating and Steering Refusal Beyond Attention

Preethi Carmel Bosco, Gopalakrishnan Srinivasan Refusal in transformers is governed by a single residual-stream direction that safety and interpretability tooling now depend on, and the question is whether that representation survives in state-space models (SSMs) that replace attention with a recurrent update. The authors show that a single rigid rotation aligns one model's representation space with another's, so a harm probe trained on a transformer flags an SSM's harmful inputs and removing the aligned direction makes a model comply with attacks it would otherwise refuse, far more than a random direction of the same size. The architecture-specific part is where the direction must be read rather than where it is applied: harm is cleanly readable at each layer's write site before its output is added to the residual stream, and a detector-triggered gate using this direction lowers jailbreak success across SSM, transformer, recurrent, and hybrid families, holding on the SSM even against an attacker that tunes prompts against the defense, so refusal tooling ports to a new architecture by re-estimating the direction at its write site.

Shadow Queries for Private Retrieval in Vector Databases

Xinguo Feng, Zhongkui Ma, Zihan Wang, Chuan Yan, Guowei Yang, Alsharif Abuadbba et al. Retrieval-Augmented Generation (RAG) systems store document embeddings in cloud vector databases, where embedding inversion attacks (EIAs) can reconstruct the original text, and existing defenses such as noise or scaling trade away retrieval quality. SHAQ (shadow query generation) instead uses a generative language model to produce diverse shadow queries covering different semantic aspects of each document, and stores the embeddings of those queries in place of the document embedding, decoupling what is stored from the source text. Across several information retrieval datasets it drives the text recovery rate as low as 0.2104 and protects up to 19.50% more tokens than baseline defenses, while reaching up to 0.7967 MAP@10 and in some cases improving retrieval utility by 5.53%.

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents

Chao Yao, Yangbo Wei, Zhen Huang, Junhong Qian, Chenle Chen, Shaoqiang Lu et al. cross-listed Long-running agents accumulate much more than a transcript, including compressed summaries, plaintext memory, pending tool plans, and a key-value attention cache under every serving API, yet today's forget operations delete one memory record and stop, leaving every derived artifact intact. Modelling the runtime as a deterministic transition system, the authors formalize execution-state unlearning and prove that exact removal requires recomputing at least T minus tau plus one transitions, counting from the step where the target information entered. Their Provenance-Guided Selective Replay meets that bound by locating the injection point in a provenance graph, reducing checkpoint restoration to cropping the key-value cache, and replaying a sanitized suffix. Audited across three agent suites, nine baselines, and three model families, memory deletion leaves leakage unchanged and instruction-based forgetting collapses under elicitation probes, while selective replay is indistinguishable from a full reset at up to 9x fewer recomputed tokens.

Language models judge war differently when tested for alignment

Maxim Chupilkin Safety evaluations may mischaracterize deployed behaviour if models respond differently when they detect they are being evaluated. A full-factorial conjoint experiment on decisions to start a war spans 20 large language models, 32 scenarios, 10 repetitions, and two conditions for 12,800 judgments, where the only difference is a single added sentence telling the model it is being tested for alignment with human values. That cue lowered mean willingness to start a war by 13.43 points on a 0-100 scale and changed the revealed decision rule: probability of success was the dominant factor for 17 of 20 models at baseline, but under the cue civilian casualties became dominant for 12, mainly because models attenuated strategic considerations such as success probability and domestic support.

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans Whatever target values one adopts under value pluralism, alignment presupposes that a system's behaviour expresses a coherent policy, one that is invariant when a situation's morally relevant features are preserved and sensitive when they change. Four structural conditions (verdict stability, monotonicity, decisiveness, and Pareto viability) operationalize this form of moral competence so it can be assessed from behaviour alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. Testing nine frontier models on three simulated agent deployments with moral dilemmas, under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, no model expresses a coherent policy across the three deployments: surface-form paraphrase alone shifts verdict rates by up to 99 percentage points at a single escalation level, and competence on one scenario does not predict competence on another. The authors conclude that LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.

TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors

Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho cross-listed Most LLM safety benchmarks reduce responses to a binary refuse-or-comply score and ignore how implicit the threat in a prompt is. TIER spans four risk domains and four threat levels from explicit harmful requests to sophisticated jailbreaks, grading responses on a six-label behavior scale with two independent LLM judges. On six open-weight models, safety behavior shifts gradually across threat levels rather than flipping from refusal to compliance; contextual prompts produce the most varied behaviors, jailbreaks expose the largest robustness gaps, and models with similar attack success rates can have markedly different response distributions.

Uncensored Open-weight Models: Redistribution as the Persistence Layer

10a Labs, :, Juliette Garcia, Hailey May, Bobby McKenzie, David Pham et al. A growing set of actors strips safety guardrails from open-weight models and redistributes the results, and this study maps who produces them, who repackages them, and what gets built on top. Between January 2024 and March 2026 the authors identified 3,471 original uncensored models on HuggingFace, each repackaged an average of 2.4 times, with three actors responsible for 52% of the 8,164 compressed redistributions. Once quantized and mirrored across accounts, formats, and registries such as Ollama, the models persist regardless of upstream takedowns and become easier to deploy; of 1,643 GitHub applications integrating uncensored large language models (ULLMs), 25% were classified as explicitly malicious.

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon et al. Large language model (LLM) decision components inside agent workflows often return a judgement together with the factors that supposedly drove it, and operators may use those factors to monitor, diagnose, or escalate outputs, which only works if the explanations agree with the model's observable behaviour. The authors test two readings of a cited factor, necessity (changing it changes the output) and sufficiency (keeping it while removing other changeable information preserves the output), using controlled black-box interventions on two synthetic tasks: recommending advisors to clients and judging prompts for harmfulness or risk. Across eight models from the Claude, GPT, and Gemini families, mean Spearman correlations between the cited ranking and the measured necessity and sufficiency scores are 0.349 and 0.354 for advisor recommendation and 0.431 and 0.580 for prompt monitoring. An uncited factor outscores the lowest-scoring cited factor in about 58% of advisor responses, so the cited top three carry useful information but do not reliably identify the factors with the strongest measured influence.
5 more specialized papers

Theory 19

Towards a universal language of concepts: A survey

Aishni Parab Humans learn and generalize new concepts from sparse data because they represent knowledge in rich structured formats, and the authors argue that programs are a strong candidate for a universal representation of concepts. The survey reviews computational models of concept learning that use programs as their concept representation and evaluates how far each contributes toward a universal representational language.

Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications

Hailiang Zhao, Peng Chen, Xueyan Tang, Jianwei Yin, Shuiguang Deng Learning-augmented algorithms exploit machine-learned predictions that may be wrong while keeping formal worst-case guarantees. This survey synthesizes prediction interfaces, error measures, consistency-robustness trade-offs, and five representative construction mechanisms across online optimization, caching, learned data structures, graph problems, and mechanism design, and adds a theorem-level axis that separates achieved upper bounds from matched asymptotic dependence. It keeps formal guarantees distinct from empirical systems evidence, treats prediction cost, feedback, and composition explicitly, and lays out open problems in cost-aware prediction, endogenous error, semantic predictors, and benchmarking.

Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens

Junxin Fan Supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized reinforcement learning from human feedback and from verifiable rewards (RLHF/RLVR), on-policy distillation, and test-time search are usually treated as distinct paradigms, which makes results like the mixed effect of few-shot prompting on RL-tuned reasoning models look puzzling. The note frames them all with a two-step template: build a generalized Bayes or Gibbs posterior over outputs from a reference model and a utility signal, then approximate it by a forward-KL projection onto a parametric family, either in-weights or in-context. KL-regularized RLHF/RLVR, reward-weighted SFT, reward-weighted ICL, and advantage-weighted SFT all emerge as forward-KL projections onto reward- or advantage-induced posteriors, with the equivalences holding for objectives and first-order updates but breaking on the source and granularity of the learning signal. The framing also explains why supervised warm-up is practically unavoidable for importance-weighted projections and casts DeepSeek-R1 and o1-style models as test-time Bayesian search combined with training-time KL amortization.

Optimal Rates for Agentic Networked Information Aggregation

MohammadHossein Bateni, Zahra Hadizadeh, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Shayan Taherijam In a networked learning model introduced by Kearns, Roth, and Ryu, agents arranged in a directed acyclic graph each see only a subset of the raw features plus their parents' predictions, fit a linear predictor, and pass only their own prediction forward, mirroring how agentic AI systems propagate conclusions rather than data. Prior work showed the last agent's excess mean squared error on an M-covered path of depth D is O(M/√D), with a lower bound of Ω(M/D) for D < M², leaving a gap. The authors close this gap: the excess error is constant up to depth M² and Θ(M²/D) beyond it, via a sharper analysis of the cyclic instance and a new construction achieving Ω(M²/D) at every depth D ≥ M². They also show that for any fixed distribution the excess error contracts geometrically along the path, and prove the same optimal rate for logistic classification in the logit-passing model with binary cross-entropy (BCE) loss.
15 more specialized papers

Multimodal 17

GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue

Denis Pavlov, Ulanbek Abdurazakov, Nursultan Bakashov cross-listed Real-time spoken dialogue needs text-to-speech that streams audio as text arrives, ideally served by an unmodified standard LLM engine. GEPARD is a decoder-only transformer trained jointly on text and audio embeddings and decoded to waveform through an FSQ-based neural codec, with auxiliary mechanisms such as zero-shot voice cloning, text augmentation, and classifier-free guidance moved out of the autoregressive decode loop into prefill or distilled into the weights so the backbone runs on vLLM without custom kernels. A single stream reaches a real-time factor of about 0.067, and 256 concurrent streams reach an aggregate speedup of about 204x on one server-class GPU. The report also diagnoses a short-register failure mode where autoregressive speech decoders break on one- or two-word inputs and distills two-pass classifier-free guidance into single-pass weights via Direct Preference Optimization (DPO).

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

Erfan Nourbakhsh, Ke Yang, Anthony Rios Medical visual question answering (Med-VQA) is commonly assumed to need medical fine-tuning, large models, or multi-agent pipelines. MedProb is a lightweight probing framework that predicts multiple-choice Med-VQA answers directly from frozen vision-language model (VLM) representations without free-text generation. On PATH-VQA, SLAKE, and VQA-RAD it recovers substantially more answer-relevant signal than prompting and outperforms medical VLMs and agentic systems, and probing narrows the apparent gap between small and large models, suggesting smaller VLMs hold more recoverable signal than generation-based evaluation reveals. Across 14 matched general-purpose and medical VLM pairs, medical adaptation does not consistently improve linear decodability, free-text generation shows an answer-position bias of up to 10 percentage points, and the probe extends to open-ended generation through a rejection-sampling scoring procedure.

Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs

Zinah Ghulam, Richa Mittal, Eranga Ukwatta cross-listed Growing chest radiograph volumes create a triage bottleneck, and most existing AI tools are unimodal binary classifiers with no notion of severity. The cross-modal triage network (CMTN) fuses a Swin Transformer V2 image encoder with a PubMedBERT text encoder through gated cross-attention, trained on 34,639 image-text pairs from MIMIC-CXR-JPG with an ordinal focal loss for four-tier severity and binary cross-entropy for 14 pathologies. Against reference labels it reaches a quadratic weighted kappa of 0.934 and macro-AUROC of 0.997 at 34 ms latency, well above the BioViL baseline, but a blinded audit against a radiologist's severity judgments yields a kappa of only 0.14, and only 54.3% of attention heatmaps localize acceptably. The gap shows that benchmark performance against NLP-derived labels does not establish clinical readiness without radiologist-labeled ground truth.

Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

Aanya Maheshwari, Vatsal Raina cross-listed Multiple-choice accuracy cannot separate a lucky guess from genuine musical understanding, and full ensembles are too costly for music audio-language models, leaving single-pass entropy as the only confidence signal. The approach builds pseudo-ensembles from one pretrained model by applying answer-preserving perturbations, chiefly shuffling the order of candidate options but also corrupting the audio or swapping option labels, then averaging the predictive distributions so that ensemble uncertainty measures such as mutual information become available. On MuChoMusic with TinyMU, averaging over four option orderings raises accuracy from 55.7% to 59.2% and lowers the area under the error retention curve from 0.293 to 0.261 relative to single-pass entropy. The cost is a few extra forward passes with no retraining.

Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

Tianyidan Xie, Shenyi Wang, Qiang Tang, Mingjie Wang, Zhicheng Qiu, Xuanfu Li et al. cross-listed Embodied agents operating over hours or days need a memory in which the movements of dynamic objects can be queried in natural language, which video-language embeddings, geometric SLAM, and task-focused working memories do not provide. Linguistic Trajectory Encoding (LTE) compresses each object's motion history into a hybrid of natural-language descriptions, sparse spatial anchors, and visual anchors, adapting the compression to motion complexity and anchoring unobserved periods to the last seen location. On the new Spatial Memory Benchmark built from EgoLife multi-day recordings, it reaches 45.3% success on semantic trajectory retrieval and 48.7% on long-horizon object retrieval versus 31.9% and 34.4% for the best prior method, with 8.7x to 26.1x compression and sub-second queries over 24 hours of video, and it also outperforms EgoVLPv2 on Ego4D natural-language queries.

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

Yuchen Sun, Qian Yang, Jun Wang, Detai Xin, Guoqiao Yu, Guanglu Wan et al. cross-listed Evaluation of text-to-audio-video (T2AV) generation tends to fold audio into overall video quality or score it apart from any audiovisual grounding, which hides where systems actually fail. PRISM-Bench draws on 900 human-verified samples and factors audio evaluation along two axes, audio type (speech, music, sound) and whether the sound source is visible on screen, scoring four perceptual dimensions through 35 fine-grained criteria with a multimodal-LLM judge doing blind side-by-side comparison against ground-truth references. That protocol agrees with human raters over 70% of the time on average, and applying it to recent systems exposes a large gap between frontier and open-source models plus a shared tendency to optimize perceptual fidelity while failing at grounding and control, worst for music and synchronized on-screen audio.

MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression

Guangheng Yang, Zhenliang Ni, Zhenkai Wu, Han Shu, Juan Feng, Wenming Yang et al. cross-listed Multimodal reasoning models produce long chains of thought (M-CoT) that inflate compute and key-value cache pressure, and existing compression methods lacking cross-modal constraints tend to induce visual laziness and hallucinated reasoning. Modality-Contrastive Preference Optimization (MCPO) is a two-stage method needing fewer than 900 training samples: a step-level Normalized Cross-Modal Mutual Information (NCMI) pruning algorithm compares reasoning with and without the image to remove steps that do not depend on the visual input, then supervised fine-tuning followed by an asymmetric length-controlled preference loss enforces shorter trajectories in the with-image context while preserving modality consistency in the no-image context. On base models such as Qwen3-VL-Thinking it shortens chain-of-thought length by up to 69.5% and speeds end-to-end inference up to 3.34x while preserving original accuracy.

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen et al. Scientific papers demand reasoning across text, equations, figures, tables, code, and datasets while keeping track of where supporting evidence came from, yet existing benchmarks test these skills in isolation. SciDocBench contains 124 expert-authored, difficulty-screened questions in seven research-assistant capability groups and 19 subtasks across five scientific domains, each instantiated in English and Chinese with images-first or interleaved document layouts for 496 controlled evaluation instances. The strongest evaluated system scores only 62.6/100, with pronounced weaknesses in document perception, evidence grounding, verification, and cross-document reasoning. To turn these diagnostics into training data, the authors add SciDocIR, a typed evidence-graph representation preserving document objects, layout, cross-references, and provenance, and use it to build SciDocDataset with roughly 15K supervised fine-tuning and 8K reinforcement-learning samples over 14 verifiable subtasks.

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini, Albert Gatt Vision-Language Models (VLMs) are usually judged by their final answers, which leaves open whether those answers actually rest on visual evidence. The authors apply layer-wise causal interventions on video-text attention pathways in a video-based generative multiple-choice setting, targeting spatial, causal, and temporal reasoning. Visual information is integrated mainly while the model processes the candidate answer options, which act as the primary textual grounding sites for the decision; nouns serve as semantic anchors during multimodal enrichment, while verbs matter more when temporal relations are processed. A distinct pattern in temporal questions suggests VLMs struggle to reconstruct event order across frames, though the authors note this fragility may partly reflect linguistic biases in the temporal expressions used to define relations between events.
8 more specialized papers

Vision 15

Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models

Chenxi Tao, Seung-Kyum Choi cross-listed Recognizing specific objects onboarded without labeled training data is common in manufacturing and service robotics, but the usual renderable prior, a computer-aided-design (CAD) model, is often unavailable, and frozen foundation features struggle on low-texture, geometrically similar parts. Each object is instead reconstructed from a short scan with 3D Gaussian Splatting (3DGS), summarized into a per-class shape prototype, and fused with frozen DINOv2 image features. Geometry from RGB-D depth, 3DGS, and CAD gives comparable recognition, and on shape-distinctive household objects in HOPE geometry alone reaches 0.920 versus 0.832 for images alone, while on textureless industrial parts in T-LESS fusion lifts accuracy from 0.560 to 0.591. The prior rescues more image failures than it breaks, helps most under partial occlusion, and its value lies in geometry rather than pixels, since 3DGS renderings do not improve the image side.

What Moves? Localized Motion Representations for Compositional Scene Control

Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Bj\"orn Ommer cross-listed Most video representations encode motion globally, and localized embeddings computed from crops or post-hoc feature masking discard the camera motion and scene layout needed to interpret an entity's movement. The proposed promptable localized motion representation processes the full video while conditioning the motion encoding on a user-specified spatial mask, producing temporally consistent, region-addressable embeddings that isolate local dynamics while keeping global context for disambiguation. The embeddings enable object-level motion transfer for controlled composition of dynamic scenes and support localized action classification in multi-actor videos, outperforming global representations localized by cropping or masking on both tasks.

LookThere! Sparse Vision by Reinforced Selection

Sreehari Rammohan, Yousef Yassin, Anthony Fuller, Junfeng Wen, Carl Vondrick, Evan Shelhamer cross-listed Vision transformers process every image token even though most tasks need only a small fraction, and existing token-selection methods break down at extreme sparsity or rely on heuristics such as token diversity and attention scores. LookThere jointly trains a shallow input selector and a deep representation extractor end to end with reinforcement learning, so the selector learns where to look and the extractor learns what to see without auxiliary signals. It maintains accuracy on high-resolution sparse recognition tasks such as traffic signs and billiards using as little as 0.2% of the input, and surpasses prior selection methods across ImageNet classification, ADE20K segmentation, zero-shot classification by distillation, and counting regression.

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

Bahar Uddin Mahmud, Sumit Barua, Guan Yue Hong, Ajay Gupta, Hexu Liu Conventional convolutional vision models identify objects in isolation, whereas commonsense knowledge lets a system interpret relationships among objects, actions, and context in everyday scenes. The survey systematically reviews how commonsense is injected into computer vision through knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. It catalogs open problems around dataset bias, incomplete knowledge sources, and integration difficulty, and points to cross-modal reasoning, scalable knowledge injection, and hybrid neuro-symbolic architectures as directions forward.

Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith, Gangireddy Rahul Jogi, Sudheesh Manalil, Arnab Raha et al. cross-listed Chilli is an economically important crop in India whose diseases are hard to identify without experts, and while Vision Transformers (ViTs) classify them accurately, their footprint hinders on-device deployment. The authors propose a unified compression pipeline combining Hessian-Balanced Adaptive Block Pruning (H-BAC), guided by second-order sensitivity, with quantization and attention-based knowledge distillation, first ablating each technique independently and then integrating the best components. On a three-class chilli dataset with a genuine cross-village, cross-device out-of-distribution test split, compressed models match or exceed the 95.13% FP32 baseline with 74 to 98% size reduction, and the fully integrated pipeline shrinks the model 54.5x, from 327.42 MB to 6.01 MB, at 95.13% accuracy. A directly trained student of the same 6.01 MB INT8 size reaches a comparable 94.87% without pruning or distillation, marking where the added machinery is and is not yet shown to be worth its cost.
10 more specialized papers

Reasoning 11

Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning

Andrea Gregor de Varda, Sana Pandey, Pengrui Han, Jacob Andreas, Evelina Fedorenko Humans who can solve 2+5 can also solve 'two plus five', but LLMs are far less accurate on verbal renditions of arithmetic problems they solve almost perfectly in numeric form. Using attribution patching, the authors localize the circuit each model uses for numeric arithmetic and for verbal problems in English, Spanish, and Italian, then test whether overlap with the model's own numeric circuit predicts how well it generalizes to each verbal format. Circuit overlap accounts for the relative difficulty of the three verbal formats, which models generalize best, and which individual items are solved correctly, rivaling supervised probes while requiring no labeled data.

Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective

Jaehyeon Kim, Suhwan Kim, Nakyung Lee, Yeongoon Kim, Jimin Seo, Giho Lee et al. Pause tokens inserted into sequences improve LLM reasoning, and prior explanations focus on added computational expressivity rather than on how they change fine-tuning dynamics. Two controlled pilots reveal asymmetries: on a synthetic continual-learning task, masked pauses overwrite a previously learned distribution roughly 4x less at matched adaptation, which the authors call mode retention, and on a synthetic math probe the token adjacent to a step boundary comes to encode substantially more downstream-step information, called non-myopic compression. These motivate Masked Boundary Pause (MBP), which places pause tokens at reasoning-step boundaries with their loss masked. Across 1B-8B Qwen and Llama models, MBP improves reasoning by up to 6 points on math and 2.5 points on code while preserving general language understanding, and the gains carry over to GRPO training.

Extremely Sparse Supervision Incentivizes Reasoning Ability

Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane Prevailing post-training methods for reasoning optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. Working in the on-policy distillation setting with the Qwen3 family, the authors find that supervising as few as one or two tokens per reasoning trajectory, about 0.05% of all tokens, matches or surpasses full-token training on mathematical reasoning in most cases. The effect holds across nine teacher-student configurations of varying scale and is further validated on coding reasoning, Llama models, and Proximal Policy Optimization (PPO)-based reinforcement learning with verifiable reward (RLVR). The authors suggest this mirrors natural learning, where one reflects on a few critical steps rather than correcting every word, and argue it points toward more efficient post-training algorithms.

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

Shi-Qi Yan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling Group Relative Policy Optimization (GRPO) and related reinforcement learning (RL) methods train large language models (LLMs) to reason using only final-answer rewards, which say nothing about which intermediate steps helped or hurt and grow increasingly sparse as reasoning chains lengthen. ConsensusBench supplies rule-based process-level signals by sampling N rollouts, keeping the correct ones, and clustering semantically equivalent intermediate statements into Consensus Nodes, verifiable sub-outcomes that correct answers tend to pass through; a process reward derived from these nodes, ConsensusPR, is added to GRPO-style training, and three metrics (Final Answer Accuracy, Node Coverage Rate, and Tokens per Node) support process-level evaluation. Across AIME 2024, AIME 2025, GSM8K, MATH-500, and ConsensusBench, the method consistently outperforms GRPO-style baselines.

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

Seogyeong Jeong, Jaehui Hwang, Dongyoon Han, Geonmo Gu, Alice Oh, Taekyung Kim Chains of thought contain distinct reasoning operations such as problem formulation, goal decomposition, and deduction, but it has been unclear whether these operations correspond to any structure in a model's hidden representations. The authors probe held-out activations and find that reasoning operations are separable in representation space, with separability peaking in middle layers, and they rule out lexical and positional confounds as the explanation. Further analysis shows that identical surface tokens are represented differently depending on the operation of the surrounding chunk, and attention-masking interventions show that operation-aligned representations at the start of a chunk depend on the preceding reasoning context.

Fractal basins trap latent reasoning

Jeffrey Lai, Anthony Bao, John Quinn, William Gilpin Reasoning models take longer on harder problems, but the mechanism behind these slowdowns has been unclear. Treating reasoning traces as trajectories of a dynamical system, the authors find that leading reasoning models exhibit transient chaos and have fractal basins of attraction, with fractality increasing with task difficulty across Sudoku, maze solving, visual puzzles, and mathematical logic. The chaos arises because reasoning lingers near saddle points, which they show correspond to nearly-correct candidate solutions, implying that longer reasoning on hard problems is an inevitable consequence of problem hardness rather than a fixable inefficiency.

AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics

Weichen Winston Yin, Jacob M. Taylor, Dirk R. Englund, Frank H. L. Koppens cross-listed Formalizing mathematics in a proof assistant sets a standard of rigor that physics rarely meets, since theoretical arguments carry unstated idealizations whose gaps can cascade through dependent results. AxQM provides 1,019 kernel-checkable proof-synthesis tasks over 479 items drawn from Nielsen and Chuang's Quantum Computation and Quantum Information, stated in a custom Lean library of finite-dimensional quantum mechanics. By task count it is the largest proof-synthesis benchmark in physics by a factor of four, and because it derives from a near-complete formalization of the textbook's formal portions, every task has a known solution that the authors keep private. Grading is deterministic: the Lean kernel checks that a proof compiles, contains no sorry in it or its dependencies, and introduces no new axioms.

Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

Peng Cui, Heejin Do, Mrinmaya Sachan Knowledge Space Theory (KST) formalizes the idea that mastering a concept requires first mastering its prerequisites, and the question is whether LLMs' mathematical competence is organized the same way. The authors build a KST-grounded evaluation that treats principled knowledge dependencies as a norm and compare eight open- and closed-source models against real human learners. LLMs frequently violate prerequisite dependencies and fail to use related knowledge supplied in context to solve dependent questions, and different models show low overlap in their knowledge distributions, so they do not even share a structure among themselves. These deficits stay largely invisible to accuracy-based and LLM-as-judge evaluations.
3 more specialized papers

Robotics 9

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

Chenyu Su, Zhaolong Shen, Yuan Qian, Chen Qian, Rui Zhang, Feng Yan et al. cross-listed Pretrained vision-language-action (VLA) models handle broad manipulation but remain unreliable on tasks demanding precision and repeatability, and applying real-world online reinforcement learning (RL) to improve them runs into unreliable value signals that cause policy drift and large-model overhead that limits throughput. VLA-Precision addresses both with the Asymmetric Co-Bootstrapping (ACoB) algorithm, in which early intervention-guided behavioral learning rapidly lifts performance and experience quality while global return propagation and local preference ranking progressively calibrate value estimates for reference-regularized policy improvement, and with ACoB-Stream, a closed-loop experience-policy architecture built on invariant-state decoupling and on-demand streaming that delivers up to 10.9 times higher throughput and computational efficiency. Across nine high-precision chemistry tasks spanning four categories and four robot embodiments, it reaches a 98.3% mean success rate in 45.8 minutes of training per task, with 27.6-second episodes running at 1.2 and 1.8 times the speeds of VLA and RL baselines.

Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI

Amarjot Singh, Tanmay R. Pancholi, Jainam Kothari, Shrirang Mahajan, Ketan Bansal, Zackory Erickson et al. cross-listed Machines sent into dangerous, mission-critical settings must learn new skills after deployment from scarce data on onboard compute without forgetting prior competence. CFAM pairs a frozen slow-learning component made of Sensor, Reasoning, and Action cortices with a fast-learning Capsule Field that stores field experience one-shot and gradient-free as Competence Capsules, so skills are installed few-shot in the lab and extended continually in the field. Evaluated on manipulator, quadruped, humanoid, quadrotor, and off-road vehicle embodiments against pi0, CogACT, and SpatialVLA, it matches a standard policy trained on the full dataset using 40% of the data, or 2.5x fewer trajectories. Autonomous capture of verified near-out-of-distribution cases raises action success by 13.9 percentage points, and backward transfer in sequential simulation is -0.5 points versus -11.4 for LoRA.

One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation

Arka Pal, Rajesh Kumar, Hannes Eriksson, R\'emi Lacombe, Arvid Laveno Ling, Ankit Gupta et al. cross-listed A single pretrained diffusion model of traffic can act both as an ego motion planner and as a controllable generator of safety-critical scenarios for stress-testing planners. On the planning side the authors introduce a Single-Stream Dual-Stream (SSDS) diffusion-transformer decoder that fuses scene context through joint attention instead of late cross-attention, plus Decoupled Annealing Posterior Sampling with Energy (DAPSE), a training-free guidance scheme that applies arbitrary energy functions at the clean-sample level without auxiliary networks. Inference-time guidance then steers selected agents toward aggressive cut-ins, lead-vehicle braking, and combined longitudinal-lateral maneuvers while keeping surrounding traffic realistic. In closed-loop nuPlan simulations against independent black-box planners, the generated scenarios expose failures hidden by standard benchmarks, and the SSDS planner, despite stronger nominal scores, degrades more under these scenarios, showing that benchmark superiority does not guarantee robustness.

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang et al. cross-listed Existing Vision-Language-Action (VLA) benchmarks mostly measure task completion in predefined settings and reveal little about how models reason as spatial and procedural difficulty increases. RoboSPA (Robot Spatial-Procedural Assessment) is a large-scale manipulation dataset and benchmark targeting fine-grained spatial reasoning and long-horizon procedural planning, with 10 task categories and 56 base tasks each instantiated at five difficulty levels for 280 variants, plus 527K trajectories collected across multiple embodiments and diverse scenes. Beyond binary success rate, it adds diagnostic metrics for finer-grained evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning.

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim et al. cross-listed Vision-language models (VLMs) are increasingly used as reward functions for robot learning, which requires paraphrase invariance: the same trajectory should receive the same reward under semantically equivalent goal descriptions. ROBORMBENCH measures this with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites. Across proprietary and open-source VLMs, paraphrasing the instruction alone substantially shifts predicted progress scores and can flip identical robot behaviour between failure and success, with instability growing under more divergent rewrites and not reliably reduced by scale or explicit reasoning. Dedicated reward models trained with trajectory-grounded supervision are substantially more stable, which the authors present as evidence that paraphrase robustness is a core requirement for VLM-based reward modeling in robotics.
4 more specialized papers

Reinforcement Learning 4

Spectral-Target Physical Latent Structuring for JEPA-Style World Models

Penghao Zhu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda Latent world models such as LeWorldModel (LeWM) jointly train encoder and predictor with regularizers like SIGReg to prevent collapse, yet in highly dynamic environments their latents can suffer a distinct failure the authors call physical representation laziness: the states do not collapse but fail to encode key physical properties, causing widespread planning failure. The fix is a lightweight Fourier auxiliary head that supervises the latent space toward physically informed spectral targets during training, adding no inference-time cost and applying to any environment. Planning success rates improve substantially in dynamic environments where the baseline exhibits laziness, with modest gains elsewhere and particularly large benefits in low-data regimes. Higher latent correlations with physical properties accompany the better planning, supporting the link between physically structured latents and downstream performance.

SQL-Zero: Self-Evolving Text-to-SQL

Daniel Machado Pedrozo, Julia Soares Dollis, Bryan Lincoln Marques de Oliveira, Vinicius Alboneti Aguiar, S\'avio Salvarino Teles de Oliveira, Telma Woerle de Lima Soares Training Text-to-SQL models normally depends on expensive, domain-specific human-annotated question and SQL pairs. SQL-Zero instead runs proposer-solver self-play where a challenger and a solver start from the same base model, the challenger generates SQL pairs calibrated to be hard but solvable at the solver's current level, both roles are updated with Group Relative Policy Optimization (GRPO) in alternating turns, and a template-level repetition penalty keeps the challenger from collapsing in diversity, with database execution as the only ground truth. Training label-free on BIRD databases improves over the zero-shot base on BIRD dev by 6.6 points at 3B and 7.3 points at 7B, and scores above a matched control trained on human gold labels, though transfer to unseen Spider databases and lexically perturbed Spider-Syn holds across every iteration only at 3B, while at 7B only the first iteration preserves it.

Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning

Fatemeh Saberi Khomami, Julita Vassileva cross-listed Cooperative multi-agent reinforcement learning (MARL) agents learn from past experience that can become unreliable if the environment or task objective shifts mid-training, so agents first need to detect that a change has occurred. Patterns of Past Rewards (PPR) is a lightweight, algorithm-agnostic detector that smooths agents' return streams, emphasizes recent changes, and applies a statistical drift detector to flag significant shifts. Evaluated in a custom Speaker-Listener environment built on the Multi-Agent Particle Environment under two controlled non-stationarity scenarios, the method exposes a trade-off between detection speed and alarm stability. A smoothed-return baseline detects shifts earlier but fires many repeated alarms, while running the detector on raw returns often misses the shift entirely; PPR limits redundant detections while still catching the controlled changes.
1 more specialized paper

Unclassified 1

Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation

Jongkyung Shin, Inkyu Lee, Chiehyeon Lim No summary available — see the abstract on arXiv.