Monday, August 31, 2026

305 papers cs.AI · cs.LG · cs.CL ← 2026-08-282026-09-01 →

Jul Aug Sep

Highlights

Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

Highlight Safety & Alignment Jacopo Dardini, Claudio Stanzione, Giordano Col\`o, Giuseppe Fenza Post-training quantization is usually assumed to preserve behavior, so models are certified at full precision and compressed afterward without re-evaluation — a workflow this work formalizes as a validation-deployment gap using Quantization Behavioral Equivalence Classes, proving that class membership does not imply behavioral equivalence. A three-stage adversarial fine-tuning procedure embeds payloads that stay dormant under source-precision checks but activate under INT8 or 4-bit compression, demonstrated on multilingual encoder-decoder translation and political stance classification rather than just decoder-only models. Backdoored translation models go from zero measured friend-foe corruption at repaired FP16 to up to 85.02% inversion after quantization, and cross-quantizer analysis shows persistence depends on the specific quantization scheme and architecture rather than nominal bit-width, arguing that the deployed configuration itself must be behaviorally certified.

Post-training quantization is usually treated as a semantically neutral optimization, but because it is a many-to-one map from full-precision weights to discrete bins, a model can be engineered to pass full-precision safety checks and only turn malicious once compressed for edge deployment. The paper formalizes this as Quantization Behavioral Equivalence Classes — the set of FP weight vectors that collapse to the same quantized model — proves that class membership does not imply behavioral equivalence, and uses that structure to build backdoors triggered by LLM.int8() or NF4 compression alone.

  • The three-stage pipeline fine-tunes on a corrupted dataset, computes per-weight QBEC intervals for the target quantizer, then runs projected gradient descent on clean data clipped inside those intervals, yielding a checkpoint that behaves benignly at FP16 while quantizing back to the malicious model.
  • On English→Ukrainian tactical translation, repaired models show zero measured friend-foe corruption at FP16 but reach 82.27% inversion for NLLB-200-1.3B under NF4 and 85.02% for M2M100-1.2B, with BLEU falling only about 5–6 points (22.87→17.17 and 33.42→31.30).
  • The same recipe applied to political summarization with the PoliTune dataset produces a measured ideological shift of ΔBias up to 0.33 for Llama-3.2-1B and 0.24 for Gemma-3-1B, while MMLU on political and legal subtasks drops only 4–7 points.
  • Cross-quantizer transfer is governed by quantizer geometry rather than bit-width: an INT8-optimized attack on NLLB retains 55% of its strength under NF4 (T=0.552) but collapses to T=0.061 under INT4 and FP4, and excluding shared embeddings and the output head from repair lifts NLLB's NF4 persistence from 46.33% to 71.75% with no measured BLEU change — a warning that partial repair is not automatically the conservative choice.
  • The evidence is bounded: all numbers come from single fixed-seed runs with no confidence intervals, the translation corpora are synthetic and may exaggerate how learnable the lexical swaps are, models are all around 1B parameters, and the authors are explicit that BLEU and MMLU are stand-ins for a real audit rather than proof that stronger source-precision checks would miss the backdoor.

Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess

Highlight Reinforcement Learning Szymon Mi{\l}osz, Piotr Duch, Szymon Grabowski Searchless chess networks such as Leela Chess Zero's Chessformer reach master strength in a single forward pass by distilling Monte Carlo Tree Search visit counts, but imitating a search is a poor proxy for playing without one. Fine-tuning with self-play reinforcement learning, the usual entropy bonus (a reverse Kullback-Leibler divergence toward uniform) is swapped for a mass-covering forward KL toward the network's own MCTS prior, paired with a sampling temperature that sharpens once the value head is confident about the outcome. In roughly two thousand steps this lifts puzzle accuracy from 93.9% to 94.9% and mate-in-four from 77% to 81% without losing playing strength, while a control fine-tuned on puzzles alone posts the largest tactical gains yet loses about 260 Elo — a better puzzle-solver is not a stronger player. Without any regularizer, self-play collapses onto a single line of play.

Searchless chess networks are trained by distilling an AlphaZero-style search's visit counts, but imitating a tree search is a mismatched objective for a policy that must commit to one move from a single forward pass. The fix here is self-play reinforcement learning on top of Leela Chess Zero's released Chessformer, with the standard entropy bonus replaced by a forward, mass-covering KL divergence toward the network's own MCTS prior — "prior-directed exploration" — so the exploration budget covers the moves the prior already judges promising rather than every legal move.

  • The recipe is deliberately minimal: a single-step policy gradient with a TD(0) value bootstrap from exponential-moving-average target weights, no replay buffer and one optimizer step per freshly collected batch (which makes PPO clipping inert), plus an entropy-adaptive sampling temperature that sets each move's temperature to the value head's win/draw/loss entropy so self-play sharpens once a position is decided but stays stochastic while the outcome is live.
  • In about 2,000 steps accuracy on a 100,000-puzzle Lichess suite rises from 93.85% to 94.87%, mate-in-four from 77.2% to 81.4%, and searchless strength holds or slightly improves at +16.6 ± 8.3 Elo over the 2374-rated base — a real but slender edge (361–318–321 over 1,000 games, likelihood of superiority 0.94).
  • Tactical accuracy and playing strength dissociate under matched compute: all self-play arms land within a one-point accuracy band while their ratings straddle the base, and a supervised control fine-tuned on puzzles alone posts the study's best tactics (96.80%) while shedding roughly 260 Elo, so a better puzzle-solver is not a stronger player.
  • Distribution-level diagnostics explain what the anchor buys — unregularized self-play collapses onto a single opening line carrying 99.9% of its trajectory mass, the entropy bonus keeps the argmax right while draining probability off winning lines (geometric-mean solution probability 0.222 versus the base's 0.660), and the mass-covering forward KL retains the hardest solutions best (0.236 versus 0.199 for the base and 0.178 for a mode-seeking reverse-KL anchor on puzzles rated ≥2400).
  • The caveats are stated plainly by the authors: the strength gain is modest and partly substitutable with the adaptive temperature alone (which already reaches +9.8 Elo on its own), the forward-KL arm is statistically tied with the reverse-KL anchor on rating, the headline is the best of many swept configurations with one run per cell so some margin is winner's-curse inflation, and the gains are retention-gated — fine-tuning promotes near misses the base ranked second or third but recovers nothing whose solution the base had already pushed below about 0.05 probability.

Fast Weight Attention for Continual Learning

Highlight HF pick · 1▲Large Language Models Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang et al. Recurrent fast-weight memories and selective state-space models squeeze an expanding context into a fixed-size state, which makes the state transition an online learning rule; under read-after-write autoregressive semantics the correct local training example at each step pairs the previous key with the current value, not the same-step key, and the common same-step association optimizes a different internal objective. From that alignment the authors derive normalized first-order updates for squared-error regression — Falcon-1 (a scalar normalized least-mean-squares rule), Falcon-2 (per-column) and Falcon-3 (sliding-window mini-batch) — plus inner-product variants, each with recurrent, masked-parallel and chunk-parallel forms and a numerically stable positive-decay renormalization. Representative variants stay competitive on language modeling and improve length extrapolation on variable-digit addition, with the framework separating temporal alignment, plasticity, forgetting and bounded rehearsal into independent knobs.

Recurrent fast-weight memories and selective state-space models compress an unbounded context into a fixed-size state, which turns the state transition into an online learning rule — but which training example that rule actually sees at each step has been left implicit. The core claim is that under read-after-write autoregressive semantics the correct prefix-aligned example at step t is (φ(k_{t-1}), v_t), not the conventional same-step pair (φ(k_t), v_t), and that fixing this alignment yields a principled family of update rules.

  • The framework derives normalized first-order updates from two objectives — squared-error regression and a negative inner-product loss — giving Falcon-1 (a scalar NLMS step), Falcon-2 (its per-column extension), and Falcon-3 (a sliding-window mini-batch update), each with an inner-product counterpart Falcon-1A/Falcon-2A/Falcon-3A.
  • The same-step association is not wrong in the causal sense, it simply optimizes a different internal objective, so the shift to prefix alignment is a semantic correction rather than a fix for a leakage bug.
  • Each variant is given recurrent, masked-parallel, and chunk-parallel formulations plus a numerically stable positive-decay renormalization, so the rules stay trainable at scale rather than being confined to sequential simulation.
  • Empirically the reported gains are language modeling parity with existing recurrent baselines and improved length extrapolation on variable-digit addition, a task chosen because it isolates whether the memory generalizes past the training sequence lengths.
  • The main limitation is evidential rather than structural: only representative variants are evaluated, the headline results are stated qualitatively without accompanying perplexity or accuracy figures, and addition is a narrow probe from which broader length-generalization claims do not automatically follow.

Rubric-to-Code Credit Assignment for Reinforcement Learning

Highlight HF pick · 3▲Reinforcement Learning Rui Jin, Jikai Chen, Yihan Chen, Hao Zhou, Demin Zhu, Kaichen Yang et al. Generating an interactive web application involves many distinct user-facing requirements, each tied to a localized region of code such as an event handler, state update, or CSS selector, but GRPO compresses all of that into one sequence-level reward applied uniformly across every token. Rubric-to-Code Credit Assignment (RCCA) builds training tasks around explicit functional rubrics, uses a hierarchical reward that separates format, source-code, runtime, and functional failures, and aligns evaluator-written textual attributions with the code spans and tokens responsible. The resulting Ling-RCCA-Flash scores 41.25 on MiniAppBench, a 32.20-point gain over the Ling-3.0-Flash base model and slightly ahead of Claude Opus 4.5, and reaches 76.19 on ArtifactsBench, topping the official leaderboard setting.

Interactive web app generation is judged by many user-facing behaviors at once, but standard GRPO collapses those outcomes into one sequence-level reward and spreads the advantage evenly over every token, so the model learns whether an app worked without learning which lines made it work. RCCA keeps the rubric structure intact: it grades applications through staged validity checks and then routes the evaluator's textual diagnostics back onto the specific code spans responsible for each pass or failure.

  • Each training task pairs a natural-language request with a set of independently checkable rubrics, split into initial-state requirements verified after page load and dynamic requirements that require executing an interaction path, and the same rubric set drives task construction, reward computation, and failure attribution.
  • A gated hierarchical reward scores format and source-code failures at 0, runtime failures at 0.1, and otherwise computes a rubric score in [0.2, 1.0] by subtracting penalties that scale with each violated rubric's importance tier (core functionality, behavioral correctness, or rendering quality) and severity, with score ceilings so a severe core failure can't be offset by satisfying minor requirements.
  • Token-level credit assignment maps evaluator-identified code regions — event handlers, state updates, DOM fragments, CSS selectors, expanded along enclosing functions, event bindings, and state reads/writes — to completion tokens via tokenizer offset mapping, then weights the GRPO loss asymmetrically: on negative-advantage samples, directly implicated tokens get weight 3.0 and contextual ones 1.5, while on positive-advantage samples those regions stay at 1.0 and unrelated code gets 1.2, so error regions are punished sharply without being reinforced on success.
  • Trained on Ling-3.0-Flash (a 124B hybrid-linear MoE with 5.1B active parameters), Ling-RCCA-Flash reaches 41.25% average pass rate on MiniAppBench versus 9.05% for the base model and 26.85% after SFT — a 14.40-point gain from RL alone — edging past Claude-Opus-4.5 at 41.14%, with the largest margins on hard tasks (37.60%) and Lifestyle (70.00%); on ArtifactsBench it scores 76.19, up 4.48 over the SFT checkpoint and 3.64 above GPT-5's official leaderboard entry.
  • The approach is confined to single-page HTML/CSS/JavaScript artifacts and untested on multi-page apps, backends, persistent storage, or auth flows, and because credit assignment rests entirely on evaluator-generated judgments and attributions, wrong diagnostics push gradient weight onto the wrong spans — most likely when a failure emerges from interactions among distant code regions rather than one local handler.

The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension

Highlight Theory Yuhe Sui, Jianing Zhang The question here is what geometry governs how well softmax attention can be approximated by a low-rank matrix, measured as the maximum-row-ℓ1 approximation rank that preserves every bounded output. Two sharp worst-case laws separate the effect of where keys and queries live: spherical self-attention has rank Θ(min{n, (1+β)^((d-1)/2)}), while full-ball geometry adds a radial degree of freedom and yields Θ(β^(d/2)) at large temperature. Within a fixed head, row-wise softmax cancels row-scalar logit directions, leaving a visible query–key interaction dimension r that gives a minimax-sharp per-instance exponent of r/2; a calibration set of 84 BERT-base heads shows modest effective-dimension reductions across many head and temperature settings, so support geometry sets worst-case temperature scaling while softmax-visible interaction geometry controls per-head complexity.

Softmax attention's rank complexity is usually discussed via head dimension or attention-matrix spectra, but row normalization erases entire logit directions and the geometry of the query/key support governs how sharply the induced Gibbs family can localize. The paper studies maximum-row-ℓ₁ approximation rank — exactly the least unrestricted rank that preserves every bounded vector-valued attention output to error ε — and separates two geometric roles: support shape sets worst-case temperature scaling, while the softmax-visible query–key interaction subspace sets per-head complexity.

  • Two sharp worst-case laws pin the temperature exponent to support geometry: spherical self-attention has rank Θ(min{n, (1+β)^((d−1)/2)}), while full-ball geometry adds one radial degree and gives Θ(β^(d/2)) in the large-token regime β ≥ β₀(d,ε), n ≥ C_d·e^(β/8) — so one radial degree of freedom shifts the exponent by exactly 1/2.
  • The upper bounds come from a weighted Gibbs-row cover whose constant is uniform in alphabet size and in all positive base weights, built by showing the log-partition function's one-homogeneous extension is a support function of a convex body sandwiched between B₂^m and 3B₂^m, then applying Dudley–Bronshteyn–Ivanov polytope approximation plus Pinsker; the lower bounds plant near-orthogonal or lattice-Gaussian "message" configurations and decode them into an identity block that forces rank.
  • For a fixed head, row-softmax quotients out row-scalar logit directions, leaving a visible interaction dimension r with the per-instance law r_ε(A) ≤ min{n, C_r(1 + Γ/ε²)^(r/2)} where Γ = βρ_Q ρ_K, and explicit bounded constructions in ℝ^(r+1) achieve r_ε(A) ≳ β^(r/2), making the r/2 exponent minimax sharp.
  • A projective softmax perturbation bound (‖softmax(x+h) − softmax(x)‖₁ ≤ 2 tanh(osc(h)/4)) converts discarded interaction into a uniform row error τ_W, yielding a robust bound at residual budget ε − τ_W and a reproducible tolerance-indexed SVD effective dimension whose criterion is output error, not explained variance.
  • Empirically, synthetic fits track the predicted slopes closely (sphere 0.504, 1.012, 1.381 vs 0.5, 1, 1.5; full-ball 0.476, 0.951, 1.427, 1.903 vs d/2), and on a fixed 84-head BERT-base calibration set with exact visible dimensions of 63–64, 41.9% of head–temperature cells admit a strictly smaller effective dimension at ε = 0.25, with Spearman associations of 0.574 and 0.606 against attention-SVD and representative-row rank certificates.
  • The main caveats are stated plainly: the β^(d/2) law needs an extreme token count (the largest construction implies an effective token count around 10^88,946), the analytic constant C_r can carry 2^O(r) dependence and is numerically useless at r ≈ 50–64, the BERT study calibrates one model rather than establishing a universal learned-head law, and low approximation rank implies neither sparsity nor arithmetic speedup.

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

Highlight Large Language Models Niccol\`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J\"org Franke, Sampo Pyysalo, Jenia Jitsev et al. Pretraining dense large language models on English-prevalent corpora, this study maps how optimal learning rate and batch size scale jointly and how each evolves marginally with model capacity and data budget, fitting a model of those relationships. Using a Warmup-Stable-Decay schedule, it also measures what learning-rate annealing buys across a wide sweep of settings and whether hyperparameters chosen for the stable phase transfer to the decay phase, then evaluates recently proposed loss scaling forms that explicitly model the capacity-data interaction. Those forms captured both undertraining and overtraining regimes well across the experiments, and the complete collection of pretraining runs is released open-source as a baseline for future OpenEuroLLM models.

Pretraining hyperparameter choices — learning rate and batch size — have to be picked before the expensive run starts, and most published scaling laws only describe their jointly optimal values rather than how each one moves independently as models and data grow. This study fits a model of both the joint and marginal scaling of learning rate and batch size for dense LLMs on English-prevalent corpora, then reexamines the loss-vs-capacity-vs-data relationship using scaling forms that explicitly model the interaction between the two.

  • The experiments sweep learning rate and batch size across a grid of model sizes and data budgets under a Warmup-Stable-Decay schedule, fitting not just the jointly optimal setting but the marginal evolution of each hyperparameter with model capacity and dataset size separately.
  • Because WSD splits training into a stable phase and an annealing phase, the authors quantify how much loss reduction annealing actually buys across a broad hyperparameter range, and test whether the optimum found during the stable phase still holds after decay — i.e. whether hyperparameters transfer between phases, which determines if cheap stable-phase sweeps are a valid proxy.
  • For the loss law itself, recently proposed forms with an explicit capacity-data interaction term are reported to fit both undertrained and overtrained regimes across the sweep, where classic separable Chinchilla-style forms tend to fit only one side well.
  • The complete collection of pretraining runs is open-sourced, giving the OpenEuroLLM effort a reproducible baseline and scaling procedure rather than a set of borrowed constants.
  • The headline caveat is scope: the corpora are English-prevalent and the models dense, so transfer of these fits to the multilingual, potentially sparse models OpenEuroLLM ultimately targets is assumed rather than demonstrated, and the abstract reports no absolute loss or parameter-count numbers to anchor the fitted coefficients.

Sliding-window beats linear attention

Highlight Large Language Models Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais Retrofitting pretrained LLMs with linear attention has been promoted as the fix for quadratic attention's ever-growing key-value cache, but the authors argue this line of work has never been compared against simple baselines. Testing sliding window attention (SWA) with attention sinks against post-trained linear-attention models across several LLMs and downstream tasks, they find SWA matches or beats them, and on long-context tasks (Needle-in-a-Haystack and BABILong) scores 2 to 10 times higher. Since SWA requires no post-training and is fast and memory-light, they recommend it over converting models to linear attention, which they suspect would need training from scratch or extensive post-training merely to match.

Linearizing a pretrained LLM to escape the quadratic KV cache is an active research line, but it has been benchmarked mostly against sink-free sliding-window attention, which is known to collapse catastrophically. The claim here is that the simplest baseline — swapping the attention mask for sliding-window attention with attention sinks, with no training whatsoever — matches or beats post-trained linear attention across 11 base models from 1.3B to 70B.

  • The method is just an inference-time mask change, SWA(w, s): each token attends to the previous w tokens (64 to 512) plus the first 4 "sink" tokens, requiring zero post-training tokens, no LoRA stage, and no specialized linear kernels since FlashAttention already supports it.
  • On short-context knowledge and reasoning, SWA(64,4) recovers 93.2% of teacher MMLU and 99.0% of the 6-benchmark average with 0 training tokens, against 83.2%/97.5% for LoLCATs at 40M tokens and 92.4%/99.1% for QRWKV6 at 350–700M tokens, and it wins the average in 9 of 11 model comparisons.
  • The gap widens sharply on long context with Llama 3.1 8B as base: at 4K on S-NIAH, SWA scores 17.2–23% where LoLCATs peaks at 5.8% and Liger-GLA at 0.8%, and on BABILong at 4K SWA reaches 15% versus LoLCATs at 3%.
  • SWA is also the fastest decoder with throughput flat in context length, and its memory stops growing once the window fills — lower than linear attention at w=64, though at w=512 the KV cache exceeds a linear recurrent state.
  • The honest caveat is that every sub-quadratic option is badly degraded in absolute terms — full attention scores 100% on S-NIAH at 4K and 60% on BABILong at 4K where SWA gets 23% and 15% — and the study covers only training-free SWA, no hybrid full-attention layers, and contexts no longer than 4K.

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

Highlight HF pick · 18▲Agents Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin et al. Long-horizon agentic tasks force LLMs to retrieve and integrate scattered information over many turns, so preserving full interaction history makes the working context grow without bound. Existing proactive context-management methods give models only search, deletion, and summarization tools, explore all editing actions as if their impacts were equal, and assign one trajectory-level reward to every intermediate edit. ContextPilot adds planning, long-term memory, and soft context-offloading tools, and trains with an RL method that uses context and entropy variation to identify critical editing decisions for branch sampling and then estimates action-level advantages from all branched trajectories passing through that edit. On long-context question answering and deep-search benchmarks it beats existing baselines across several base models while keeping a more compact working context.

Long-horizon agents that append every tool call and observation to their context hit a wall: the working context grows monotonically, and the usual fixes (rule-based truncation or summarization) give the model no say in what gets kept. ContextPilot from Tencent Youtu Lab and Shanghai AI Lab lets the agent manage its own context through an expanded tool set, and trains that behavior with a reinforcement learning recipe that treats context-editing decisions as first-class actions rather than incidental steps.

  • The tool set extends the search/delete/summarize baseline of StateLM with planning tools, a structured long-term memory built from entities, timestamps and event episodes with edges between related items, and soft offloading tools including summarizeContext, compressContext (via llmlingua-2) and a foldHistory operation that discards history into keywords plus a summary recoverable by keyword search.
  • On the training side, context-aware partial rollout scores each context management action by how much it changed context length and token entropy relative to the initial query state, then spends the leftover rollout budget branching from the most sensitive actions — the paper collects 128 snapshots per query starting from 8 trajectory-level rollouts — and credit for an intermediate snapshot is the average reward over all terminal trajectories passing through it, an unbiased estimator with variance σ²/n instead of σ².
  • Working within a 32K context window, ContextPilot-14B-RL reaches a 72.20 average across NovelQA, ∞Bench, LongMemEval-S and BrowseComp+, beating its own 128K-context Qwen3-14B backbone at 53.26, while ContextPilot-8B-RL gains 3.55 points over StateLM-8B-RL; on deep search the method adds 1.51 average points over SUPO across the WebSailor-7B and WebExplorer-8B backbones.
  • Reinforcement learning helps most where contexts are longest — the gain over the supervised model is 5.34 points on BrowseComp+ (average input 552K tokens) versus a modest lift on NovelQA (119K tokens, and partly seen during supervised training) — and the ablation shows entropy-based branching alone is unstable, costing 1.32 points on BrowseComp+ until context variation is added.
  • Per-turn input length stays flat at roughly 8K–10K tokens on BrowseComp where WebExplorer-8B grows nearly linearly toward 30K, though the authors note the tool set still may not cover all context-editing needs, hyperparameters for partial rollout and credit assignment went unsearched for compute reasons, and evaluation never leaves long-context QA and deep search — agentic coding and GUI agents remain untested.

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

Highlight HF pick · 12▲Large Language Models Zhuoshi Pan, Junru Lu, Yan Qian, H. Vicky Zhao, Di Yin, Xing Sun Factual question answering benchmarks assume a single canonical answer, which hides whether a model retains genuinely divergent accounts of long-tail facts. ElephantBench is a closed-book probe of 1,094 questions built by an auditable graph-based pipeline that retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, converts them into multi-account records, and verifies each answer against the source documents, authoritative public sources, and human annotators. Across 32 models, the strongest recovers both accounts on only 52.4% of questions, recalling one and omitting the other on nearly all the rest; larger models and inference-time reasoning improve recall without eliminating the incompleteness, and corpus analysis links completeness to how much exposure the minority account has.

Factual QA benchmarks that assume one canonical answer can't reveal whether a model's parametric memory holds the multiple competing accounts that real web sources report for long-tail facts. ElephantBench is a closed-book probe of 1,094 questions mined from naturally occurring cross-source disagreements, scoring not just whether a model recalls a rare fact but whether it recovers every verified account of it.

  • The construction pipeline takes the documents a DCLM fastText quality filter discards, builds a document graph over them where support edges link agreeing sources and conflict edges mark incompatible values for the same subject–attribute pair, then turns each conflict-centered subgraph into a QA record — knowledge-point clustering plus entity matching cuts the candidate pair space by at least 93.4% versus exhaustive comparison, and every answer is checked by an LLM against source text, independently confirmed by a web agent on authoritative sources like Wikipedia, and finally audited by human reviewers.
  • Across 32 models, the best system Kimi-K3 recovers both accounts on only 52.38% of questions, with Gemini-3.1-Pro at 50.37% and GPT-5.5 at 50.18%; the dominant failure is not amnesia but incompleteness, since failed recall for the top three models sits at just 2.19–2.65% while partial recall runs 45–47%.
  • Scale helps recall but not completeness — Qwen3.5 from 2B to 397B lifts complete recall from 1.65% to 32.27% and drops failure from 81.35% to 8.50%, yet partial recall climbs from 17.00% to 59.23% — and inference-time reasoning is similarly uneven, adding 13.99 points for GPT-5.6-Sol but costing small Qwen3.5 models up to 0.64 points as deliberation converges on the consensus account and suppresses the minority one.
  • Corpus analysis ties the failure mode to exposure asymmetry: a one-standard-deviation increase in majority-account exposure raises partial recall by 14.18 points and cuts failure by 10.17, whereas the same increase on the minority side raises complete recall by 15.13 points — so overall fact frequency governs whether the model knows anything, while minority-side exposure governs whether it knows everything.
  • The ceiling is shared rather than model-specific: an oracle pooling all 32 configurations reaches 81.2% complete recall and zero failures, but 206 questions remain partial for every system, and the authors caution that exposure results are associations from one public corpus rather than measurements of any evaluated model's undisclosed pretraining mixture.

Video Generative Models as Geometry Learner

Highlight HF pick · 4▲Vision Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng Generative geometry estimation currently adapts pretrained image diffusion models, either training depth and normal predictors separately — forfeiting the correlation between the two targets — or jointly fine-tuning modified backbones, which needs a lot of labeled data. GeoNeXt instead repurposes a pretrained video generative model and casts geometry estimation as next-frames prediction, inheriting temporal structure and richer priors while adapting them to jointly model image and geometry targets in both directions. On zero-shot monocular depth and surface normal estimation across diverse datasets it beats prior task-specific and unified generative methods, and rivals discriminative state-of-the-art systems trained on over 100x more data, leading on several benchmarks.

Monocular geometry estimation with image diffusion priors forces a choice between separate per-task models and heavily modified joint architectures that need far more labeled data. GeoNeXt sidesteps both by repurposing a pretrained video diffusion model, casting depth and surface normal prediction as next-frames generation conditioned on the input RGB image.

  • Building on Stable Video Diffusion, the method encodes the image–depth–normal triplet through a frozen Stable Diffusion VAE, replicates the image latent into the geometry slots as conditioning, and fine-tunes only the denoising U-Net with the text/CLIP branch removed — no new attention or switcher modules.
  • Trained on just 59K synthetic samples from Hypersim and Virtual KITTI 2, it beats the unified generative baseline GeoWizard (208K samples) on zero-shot depth by 6.2 AbsRel on KITTI and 10.4 AbsRel on DIODE, and posts the best ETH3D AbsRel of 5.6 against 13.1 for Depth Anything, which uses roughly 100× more training data.
  • On surface normals it is competitive with the specialist discriminative model DSINE and the generative E2E-FT, reaching mean angular errors of 16.4° on iBims-1, 33.0° on Sintel, and 22.8° on OASIS while predicting depth from the same forward pass.
  • Ablations isolate the two design choices that matter: co-generating the image alongside geometry rather than geometry alone, and joint rather than separate depth/normal training, each worth roughly 0.5–1.2 AbsRel and 0.5–0.9° mean angular error; swapping the depth/normal frame order changes results negligibly.
  • Inference runs 5 denoising steps × 5-seed ensemble in 10 s at 768×768 on an A5000 — far cheaper than GeoWizard's 272 s at 50×10 — but the model still predicts only affine-invariant depth, is trained purely on synthetic indoor and driving data, and the reported GeoWizard numbers are the authors' own reproductions.

Applications 80

Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics

Till Muser, Giovanni Abati, Ivan Dokmani\'c Scientific machine learning architectures are built for flat Euclidean grids, so applying them to spherical dynamics on a latitude-longitude grid distorts convolutions near the poles, makes 2D FFTs in Fourier neural operators falsely assume double periodicity, and warps geodesic distances in vision transformer positional encodings. Dandelion is a spherical version of the warp-based neural PDE solver Flower: each layer predicts a tangent-plane displacement and transports features along great circles, with hierarchical U-Net-style pooling done entirely in the spherical-harmonic domain, so spatial mixing comes from coordinate warps and not convolutions. A companion benchmark suite of natively spherical PDE datasets — a modified Galewsky jet, chained turbulence, Cahn-Hilliard decomposition, spherical Riemann shocks, Held-Suarez transport, and global ocean dynamics — fills the gap between toy problems and ERA5, and Dandelion places first or second on every dataset, with its margin over non-warp baselines growing at higher resolution.

LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data

Abin Shakya, Wilson Samuels, Dominica Wilson, Gioia A. Marchi, Israa Draz, Chenxing Luo et al. cross-listed Decades of experimental results sit in publications in forms that cannot feed data-driven modeling, and building databases from them requires both finding the relevant studies and extracting quantities with enough scientific context to stay usable. LitCurate is an open-source framework that runs literature discovery, relevance screening, full-text processing, and structured extraction with large language models as separate auditable stages, retaining intermediate results and provenance so researchers can inspect and revise any step rather than trusting a black box. Applied to high-pressure mineral physics, it produced an equation-of-state database of 1,334 entries drawn from 205 papers, linking parameters to mineral phases, compositions, equation formulations, and methods, and labeling each value as source-reported or citation-reported.

RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing

Madhusudan Srinivasan, Namith Nishal Raphae Retraining a classifier can break inputs the previous version got right, and confirming those regressions is expensive when ground truth needs human annotation or simulation, so only a fraction of inputs can be checked. RiskBlend ranks candidates by blending four signals — historical failure patterns, prediction shift, decision-boundary shift and neighborhood change — with weights learned on validation data under an APFD-squared objective, instead of relying on a single model's confidence scores. Across 1,200 configurations spanning four datasets, five classifiers and four update scenarios, it achieved the highest average APFD in all 80 dataset-classifier-scenario combinations, improving by up to 0.32 APFD over the strongest baseline; confidence-based ranking remained competitive only for linear classifiers on sparse categorical features.

Diffusion Distillation for Efficient Weather Ensembles

Yiming Yang, Valentin Brekke, James Briant, Serge Guillas Diffusion models produce skillful weather ensembles but pay for it with costly iterative sampling. A supervised energy-distance distillation objective compresses a multi-step diffusion teacher into a single-step student by aligning student forecasts with both teacher samples and ground-truth observations. On global forecasting and typhoon-track prediction the student outperforms existing distillation methods, retains skill on extreme events, and matches or surpasses the teacher across key metrics using just one neural function evaluation per autoregressive step.

Efficient Auto-Interpretability of AI Models in Biology

Piotr Jedryszek, Oliver M. Crook cross-listed Sparse autoencoders (SAEs) could turn biological foundation models into instruments for discovery, but three distinct questions — whether a latent is coherent, whether it can be described, and whether that description predicts anything — routinely get conflated. The proposed pipeline separates them: cross-seed dictionary stability decides which latents are worth investigating, an intruder-detection task asks whether a latent's activating examples share a recognizable pattern, and a final pass converts a candidate biological description into falsifiable in-silico predictions. On the Boltz-1 Pairformer trunk, stability prioritization found interpretable latents at 5.2 times lower measured cost (4.4 times fewer evaluations each) while recovering over half of them, though it appears to favor structure-related features over function-related ones.

From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical Diagnosis

Yue Zhou, Haiyang Zhou, Jin Zhang, Kong Wang, Yongxin Ni, Youhua Li et al. Interactive diagnosis systems decide turn by turn whether to ask another question or commit to an answer, but existing methods steer that choice by predictive uncertainty alone and ignore that missing a severe disease costs far more than missing a mild one. Severity-Aware Conformal Clinical Planning treats the process as a risk-sensitive sequential decision problem, maintaining separate diagnostic, safety, and masked-evidence beliefs, calibrating turn-specific conformal prediction sets and severity-weighted differential-diagnosis risk on held-out trajectories, and feeding that calibrated risk into Monte Carlo Tree Search to score long-horizon Ask and Commit paths. On DDXPlus and MediQ across several LLM backbones, it reaches more accurate diagnoses with fewer questions while reducing high-risk errors on severe cases.

Beyond Pairwise Graphs in Science: Hypergraph Adaptive Wavelet Operators for Parametric PDEs

Rajat Sarkar, Venkataramana Runkana, Souvik Chakraborty Neural surrogates for physical simulation work best on regular grids, yet realistic geometries demand unstructured meshes, and graph-based operators represent only pairwise edges, missing the group-wise coupling among mesh cells, neighborhoods, and conservation volumes. HALO lifts the domain to a hypergraph and learns in its spectral wavelet domain, using Chebyshev polynomial wavelet filters to avoid eigendecomposition — localized spectral kernels at linear sparse-matrix cost — with trainable dyadic scales regularized toward tight-frame coverage so the frequency response adapts per equation. Across 2D and 3D benchmarks on structured and unstructured discretizations it is best or near-best against frequency-, transformer-, DeepONet-, state-space-, and graph-based baselines with stable multi-step rollouts, and stays competitive with the strongest fixed-discretization transformers on industrial aerodynamic meshes of a few hundred thousand points while remaining resolution-equivariant.

From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning

Lokendra Birla, Milind Savagaonkar, Visnu Srinivasan, Sowmya Rasipuram, Shubhashis Sengupta Financial question answering mixes tables, charts, and narrative text, and the standard Exact Match metric punishes answers that differ only in units or formatting, distorting how model reasoning is assessed. The pipeline generates synthetic question-answer pairs behind aggressive validation checks, fine-tunes smaller language models with Quantized Low-Rank Adaptation (QLoRA), and scores answers by evaluating the arithmetic expression a model produces rather than string-matching the ground truth, with a modified loss that blends cross-entropy, semantic similarity between predicted and reference expressions, and the new metric. On ConvFinQA the combination of synthetic data and the modified loss produces significant gains in question-answering accuracy over the untuned baseline.

Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering Researchers

Rashina Hoda, Carolyn Seaman, Victoria Gomes, Rodrigo Spinola cross-listed Software engineering researchers are increasingly using language models for qualitative data analysis (QDA) — coding interviews, open-ended survey responses, and similar material — with little strategic guidance on what can go methodologically wrong. Drawing on the authors' qualitative research experience, the catalog names practices that look advantageous but undermine analytical rigor and validity, grouped by escalating impact into Dangerous Drivers, Operational Missteps, and Analytical Failures. The intent is both to give researchers a list of temptations to avoid and to give reviewers shared vocabulary for identifying failed practice.

DisCTI: Who Needs to Know Timely? Automated Sector-Aware Cyber Threat Intelligence Dissemination

Fajar Wijitrisnanto (National Cyber and Crypto Agency, Jakarta, Indonesia), Alsharif Abuadbba (CSIRO, Sydney, Australia) et al. cross-listed Cyber threat intelligence (CTI) is only useful if it reaches the right industry sector quickly, yet on the Malware Information Sharing Platform (MISP) 98% of events carry no sector tag at all, forcing analysts to manually sift heterogeneous feeds. The authors recast sector-targeted dissemination as multilabel classification, build a dataset of 872 sector-labelled CTI events drawn from a threat intelligence platform, and fine-tune BERT over events expressed in the Structured Threat Information Expression (STIX) format for cross-platform portability. The resulting DisCTI classifier reaches a macro-averaged F1 of 0.89 with a Hamming loss of 0.055, meaning 94.5% of individual sector-label assignments are correct.

Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code

Animesh Shaw cross-listed Prior studies count vulnerabilities in model-written Infrastructure-as-Code (IaC) without a human comparison point, so they cannot say whether models are actually less secure than engineers. GenIaC-SecBench covers 100 deployment scenarios across 12 model configurations from four vendors, producing 1,196 artifacts scanned by Checkov, Trivy, and KICS, alongside 634 human-authored templates run through the same toolchain. Because vulnerability density falls sharply with artifact size (Spearman ρ = −0.55), unmatched comparisons mostly measure length; matched on declared-resource count, every model configuration lands between 3.21x and 3.87x the human vulnerability density, with the gap widest on single-resource tasks. Vendor extended-thinking APIs cut vulnerabilities by 12.0% while prompted chain-of-thought is statistically indistinguishable from plain generation, and deployability shows no correlation with security.

Explainable Uncertainty Estimation for Reliable Medical AI

Li Rong Wang, Jamie Duell, Xinran Xu, Thomas C. Henderson, Yu Yue Hew, Pik Wan Erica Chiang et al. Uncertainty estimates tell a clinician when a prediction may be unreliable and explainability methods say which features drove a prediction, but the two are computed separately, so nothing indicates why a given prediction is uncertain or which additional test would reduce that uncertainty. The egRUE (Expected Gradients Reconstruction Uncertainty Estimate) method folds prediction explanations into the uncertainty computation itself and decomposes the resulting uncertainty into per-feature contributions, with proved theoretical properties and experiments against existing estimators. A user study with medical experts found that egRUE's feature-level explanations improved calibrated trust over bare uncertainty scores, raising confidence in correct predictions and lowering it on incorrect ones.

Learning to Difference: Adaptive Reversible Differencing (AdaRDiff) for Time Series Forecasting

Morad Laglil, Younes Hlal, Marouane El Hadari, Emilie Devijver, Eric Gaussier Differencing — subtracting nearby past values to strip out trend and seasonality — is a classical fix for long-horizon forecasting, but its dependence on hand-chosen orders and periods has kept it out of modern deep architectures. AdaRDiff makes the operation learnable, using trained weights over previous time steps to produce stabilized residuals for forecasting and then autoregressively restoring the removed components; the reconstruction has a closed-form convolutional expression that parallelizes on GPU for up to a 33.7× speedup over the naive recurrence, and a two-phase training schedule separates structure discovery from reconstruction learning. Across eight electricity, weather, traffic, and energy benchmarks it reaches state-of-the-art accuracy at negligible parameter cost, and as a drop-in module it improves eight different backbones, by up to 25.9% with a linear model and 18.3% with iTransformer.

Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching

Ahmad Asadi, Reza Safabakhsh Because market regimes shift, no single portfolio-management strategy stays best, so this system selects among expert strategies rather than trading directly. A dual-stream variational autoencoder encodes asset-level and market-wide state into a retrieval key, a knowledge base stores past market situations together with how each expert performed in them, and an instruction-tuned large language model reasons over the retrieved evidence to pick an expert; a monotonicity result shows that adding a locally superior expert cannot degrade the switcher. Across cryptocurrency, stock, and foreign exchange markets the selector led on cumulative return and Sharpe ratio, raising stock-market cumulative return from 26% for the best fixed expert to 34% and Sharpe from 0.74 to 0.96.

Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation

Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu et al. E-commerce repurchase recommenders are usually built as binary classifiers answering "will this customer rebuy within W days", which forces a separate trained model for every time horizon. Evaluated on millions of customers at a large grocery platform across more than thirty ablation configurations, survival models predicting time-to-repurchase directly replace that stack: a single Accelerated Failure Time model matches or beats three per-horizon classifiers at their own horizons while using roughly three times fewer trees in total. The empirical hazard turns out to be slightly decreasing (shape parameter around 0.9), contradicting the intuition that grocery items get more likely to be rebought the longer since last purchase, and a four-parameter calibration maps survival curves to per-horizon probabilities with no cross-horizon monotonicity violations. Calibration quality varies tenfold within the same model family, so the team ships Exponential AFT (expected calibration error around 1e-4) where probabilities are consumed and Log-Normal where only ranking matters.

BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence

Changze Li, Yutong Cheng, Tsania Camila Finnisa, Qian Cui, Wei Ding, Peng Gao cross-listed Cyber threat intelligence mostly lives in unstructured reports, and automated knowledge-graph extraction so far handles one report at a time, leaving the harder cross-source problem untouched: different vendors give the same threat different names. BEACON anchors every report on attack behaviors mapped to MITRE ATT&CK, the standard catalog of attack techniques, then attaches contextual entities such as threat actors, campaigns, and indicators of compromise to those anchors so all per-report graphs share one canonical space. Extraction uses a propose-then-verify loop grounded in both report evidence and official ATT&CK definitions to suppress hallucination, and merging proceeds hierarchically from character-level and semantic similarity to overlapping technique neighborhoods, iterating as merges pool more neighborhoods. The authors release two human-annotated datasets built from 34 sources — the largest for report-level extraction (8,395 elements) and the first for cross-source consolidation (3,487) — on which BEACON beats every baseline by at least 23% and 9% respectively.

RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents

Yupeng Zhang, Liuyuan Jiang, Hongyi Huang, Bingheng Li, Lisha Chen A trading policy that reacts systematically to price moves becomes predictable to other market participants, and RetailAgent tests whether language model agents show that kind of directional structure. The agent sees anonymized intraday equity price histories plus permitted state and repeatedly chooses long or flat before the next interval's return is revealed; comparing returns during long versus flat intervals on the same stock-day, after removing the overall fraction of long decisions, isolates timing skill from exposure. Timing came out persistently negative across input modality, horizon, state, and model family, and shuffling the saved action sequences largely removes the effect, indicating the alignment between actions and subsequent returns rather than a scoring artifact. Feeding the agent its own written memories increased policy persistence, and negative timing was worst on stock-days where the agent used both actions.

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Chengpiao Huang, Kaizheng Wang cross-listed Synthetic data can shore up statistical inference when real observations are scarce, but treating generated samples as if they were real introduces bias and breaks confidence-interval coverage. The proposed framework parameterizes synthetic augmentation by two knobs, how many synthetic observations to add and what weight to give them, and estimates a size-weight frontier from a population of historical related tasks: for each weight, the largest synthetic sample size at which all smaller sizes still hit the target coverage. A finite-sample coverage guarantee holds simultaneously for every configuration on or below the estimated frontier, and in experiments augmenting opinion survey data with large language model responses hit target coverage while substantially narrowing confidence intervals.
62 more specialized papers

Large Language Models 48

Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

David Noever, Forrest McKee In 2011 IBM's Watson beat human Jeopardy! champions using a curated billion-document corpus on a cluster of POWER7 servers, frozen at build time and impossible to copy; the question asked here is whether that same snapshot of testable cultural knowledge now fits in a single downloadable file. A 9 GB 4-bit quantized Qwen2.5-14B is run over all 529,939 clues from 41 broadcast seasons (1984–2025), which the authors describe as the first evaluation over the complete corpus. Under a strict forced-response protocol with exact and fuzzy matching the model answers 67.0% of all clues and exceeds 85% on factoid categories. On clues that aired after its training cutoff it holds 65% and Claude Opus 4.8 holds 95%, while Watson scores zero by construction since it could not answer outside its curated distribution.

Accelerating LLM Inference via Vector Index Based Output Embeddings

Martin Loretz, Sepp Hochreiter The output embedding matrix is a memory-bandwidth bottleneck during autoregressive decoding, especially for compact models carrying large multilingual vocabularies. The output projection followed by top-k token selection is recast as a maximum inner product search over token embeddings and served by an HNSW vector index, which returns a small candidate set whose logits are scattered into a sparse full-vocabulary tensor so existing decoding pipelines work unchanged. On CPU inference with Gemma 3, Llama 3.2, and Qwen 3, end-to-end batch-size-one decoding throughput improves by up to 82% for Gemma 3 270M while generation quality holds under AlpacaEval, suggesting approximate retrieval is a workable substitute for dense output projections in latency-sensitive small-batch settings.

Rating the Raters: Rasch Measurement Theory for LLM Evaluation

Pratik S. Sachdeva, Nathan Boudol LLMs now sit on every side of evaluation — as examinees scored on benchmarks, as judges of other models, and as raters of human-written content — yet standard practice does not separate how much the rater, the instrument, and the object each contribute to a score. Rasch measurement theory (RMT) is proposed as the missing tool, decomposing ordinal ratings into separable facets on a shared scale and supplying diagnostics for miscalibration and rater bias. Fitting many-facet Rasch models to annotations from nine LLMs across families and capability levels on the Measuring Hate Speech corpus shows they differ systematically from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating-scale use — all of which conventional agreement metrics would hide.

Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection

Fina Polat, Daniel Daza, Pengyu Zhang, Klim Zaporojets, Paul Groth Entity disambiguation (ED) is usually modeled as one task, though it contains two — retrieving candidate entities and selecting the right one — and dual-encoder models force a single embedding space to serve both while requiring a trained retriever that must be maintained as the knowledge graph changes. Holding an LLM selection stage fixed, the study compares sparse retrieval (BM25), web knowledge-base search, and a state-of-the-art trained dense retriever across open and closed models. A fully training-free BM25 retriever paired with an LLM selector sets a new state of the art on ZELDA, raising inKB micro-F1 from 82.3 to 86.3, while a trained dense retriever reaches 88.5. Decoupling the stages also lets the system abstain when the correct entity is absent from the candidates, reaching 90.7 F1 in a setting that rewards correct abstentions.

LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation

Neville K. Kitson, Anthony Constantinou Bayesian network structure learning (BNSL) from observational data recovers which variables are connected but often cannot determine edge direction, while LLMs carry broad causal knowledge of uneven reliability. Both sources are expressed as Probabilistic Dependency Graphs, in which every edge holds a distribution over directed, undirected, and absent states so the two can be fused by weighted averaging. Across 26 benchmark networks combining ensembles of FGES, Tabu, and PC with Gemini, Claude, and GPT over multiple prompts and seeds, a plain 50/50 fusion improves F1 over the better single source in 22 of 26 networks, a mean gain of 0.056 (p < 0.001). The roles are complementary: BNSL supplies a high-recall skeleton at 80% versus 60%, while the LLMs orient edges at 96% accuracy versus 77%.

XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli Multilingual multi-hop question answering benchmarks typically translate an entire example into one language, which conceals failures that occur where a reasoning chain crosses a language boundary. XHotpotQA models each instance as an evidence-dependency graph and assigns languages separately to the question, bridge evidence, answer-bearing evidence, and distractors; the audited release holds 15,661 training and 7,405 validation instances with sentence-level support supervision, and 95.60% of validation items place gold paragraphs in different languages. Across three readers, full question-evidence language mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 and different-script evidence with deficits of 11.98 to 23.70 points, while the corresponding evidence-selector gaps are under two points — locating the weakness in reading rather than retrieval.

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie et al. Language models using Gated DeltaNet or Kimi Delta Attention swap the growing key-value cache for fixed-size recurrent states, but those states are typically kept in FP32, occupying substantial GPU memory and making updates memory-bandwidth bound. Uniform post-training quantization turns out to trade poorly here — INT8 and FP8 already hurt complex reasoning, and INT4 and NVFP4 collapse accuracy — while quantization-error energy concentrates in a few channels whose relative decay strength stays stable across prompts. DAMP uses offline calibration on both error energy and decay-based persistence to keep high-risk channels at higher precision and the rest at INT8; on Qwen3.6-35B and Kimi-Linear-48B across six benchmarks it holds accuracy near the FP32 baseline at 9.9 bits per state value while cutting recurrent-state storage by 69.1% and speeding the update kernel up to 2.01x.

Trajectory-Level Speculative Decoding for Diffusion Language Models

Tianxiang Pan, Baitao Gong, Mo Guang, Hongwei Yong, Tianpeng Jiang, Yaqian Li et al. Diffusion language models generate tokens in parallel through iterative denoising, but decoding degenerates to one token at a time when confidence is low, and standard speculative decoding does not apply because these models speculate over denoising trajectories — multi-token updates with positions and unmasking orders — rather than left-to-right token sequences. The proposed framework drafts trajectories via confidence-stratified tree exploration, verifies them with blockwise parallel evaluation under bidirectional attention masking, and adds inter-block speculation for cross-block lookahead, with a formal account of when the procedure is exact and of trajectory drift as the price of more parallelism. Built on Fast-dLLM's dual-cache infrastructure, it cuts denoising iterations by 30-40% and raises tokens per step from 2.6 to 4.3, giving 7-14x speedup over vanilla diffusion decoding and 1.3x over Fast-dLLM with under 1% accuracy change on reasoning and code benchmarks.

When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages

Sanjeev Kumar, Atsuki Yamaguchi, Nikolaos Aletras Subword tokenizers push the frequency statistics of dominant languages onto low-resource languages that share their script, while byte-level models avoid this but mismatch the word-level granularity that many tasks need in non-Latin scripts. Hierarchical byte models that group bytes into word-aligned chunks normally demand huge training budgets and misalign representationally against a frozen subword language model, so the authors initialize byte embeddings directly from the frozen model's subword representations, add a chunk alignment loss projecting dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech supervision to guide boundary detection. Across six languages the tokenizer-free approach improves word-level morphological tasks, with up to a 13.3% gain on part-of-speech tagging.

Knowing Before Answering: Decoding Language Models for Reliable RAG

Syed Mahbubul Huq, Christopher Child, Tillman Weyde, Pranava Madhyastha Retrieval-augmented generation can hand a model documents that are insufficient or that contradict each other, and the system ideally recognizes which case it is in before answering. The authors build a controlled benchmark of fictitious retrieval contexts labeled answerable, insufficient, or conflicting, then train a lightweight linear probe on hidden activations and attention-derived features to make that three-way call. Across 16 language models of varying architecture and size, the feature-based router consistently beats prompting baselines and specialized RAG models, with the most informative signal appearing in middle layers and hidden activations outperforming attention values or MLP outputs in most models.

Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation

Dipto Sumit, Sakib Ul Haque, Farig Sadeque Knowledge distillation gains for small models are usually reported from a single seed, at scales where nobody has measured run-to-run variance. Eight distillation variants were compared against plain supervised cross-entropy on a 740-instance healthcare API routing task with a 1.5B Qwen student and a 20B teacher, using three to six seeds for key configurations. Per-seed standard deviation ranged from 2.8 to 48.7 percentage points, wide enough to swallow every claimed gain below five points, and three of seven variants showed bimodal collapse where at least one seed in three to five lands below 55% accuracy — including a previously undocumented truncation mode in reasoning_kd that emits reasoning but stops before naming a function (0.9% accuracy). Only progressive_kd and rank_kd avoided collapse, and a cross-split +3.78 point gain from input enrichment reversed to -2.70 under controlled within-split multi-seed retesting.

Fast Weight Attention for Continual Learning

Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang et al. Recurrent fast-weight memories and selective state-space models squeeze an expanding context into a fixed-size state, which makes the state transition an online learning rule; under read-after-write autoregressive semantics the correct local training example at each step pairs the previous key with the current value, not the same-step key, and the common same-step association optimizes a different internal objective. From that alignment the authors derive normalized first-order updates for squared-error regression — Falcon-1 (a scalar normalized least-mean-squares rule), Falcon-2 (per-column) and Falcon-3 (sliding-window mini-batch) — plus inner-product variants, each with recurrent, masked-parallel and chunk-parallel forms and a numerically stable positive-decay renormalization. Representative variants stay competitive on language modeling and improve length extrapolation on variable-digit addition, with the framework separating temporal alignment, plasticity, forgetting and bounded rehearsal into independent knobs.

KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation

Hojun Jeong, Gyunyeop Kim, Sangwoo Kang Fine-tuning-based knowledge editing normally optimizes cross-entropy, which raises the probability of the edited answer without constraining how the rest of the output distribution shifts; across sequential edits that unconstrained redistribution accumulates and degrades locality, meaning the model's behavior on unrelated inputs changes too. KLOD replaces this with a bounded objective that stops amplifying the target once its probability crosses a threshold, while preserving the target-excluded distribution at target positions and the full next-token distribution at prefix positions. On CounterFact and ZsRE with Llama3-8B-Instruct and Qwen2.5-7B-Instruct, it substantially mitigates locality degradation while maintaining high edit reliability, and the probability threshold exposes a tunable generalization–locality trade-off.

AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not

Zhengyang Shan, Yukyung Lee, Sophie Hao Machine-generated text is known to be stylometrically distinguishable from human writing, but models are now used just as often to edit human drafts, and it was unclear whether editing leaves the same trace. Measuring stylometric features across 8 large language models and 5 domains, the authors find generation leaves a consistent footprint driven mainly by entropy and lexical diversity, while the remaining features vary by domain and generator. Editing does not reproduce that footprint: edited text shows a small rise in lexical diversity alongside a drop in entropy rather than the joint increase typical of generation, with lexical density becoming the dominant signal, so stylometry separates edited from generated text but struggles to separate edited from human text.

SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models

Enqiao Lu, Xingrui Yu, Yiwei Fu, Zhenglin Wan, Pengfei Zhou, Wangbo Zhao et al. Spiking neural networks (SNNs) promise energy-efficient language modeling via sparse, event-driven computation, but are hard to train from scratch, so practitioners distill from a pretrained artificial neural network teacher — using fixed corpus prefixes, even though inference conditions on the model's own generations. That prefix-source mismatch shows up as both output-policy divergence from the teacher and drift in internal spiking dynamics, and a naive on-policy distillation variant turns out to suffer delayed rollout-feedback collapse. SpikeOPD stabilizes on-policy training by combining full-KL teacher correction with policy anchoring to a frozen reference SNN on matched prefixes and layerwise spike regularization, improving average accuracy over standard distilled SNNs by 0.8, 1.7, and 2.9 points at 0.125B, 0.35B, and 1.3B parameters while preserving sparse compute.

HyQuant: Hybrid-Precision Quantization for LLM Attention

Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang et al. Quantizing the attention module of a large language model to very low bit-widths introduces errors large enough to hurt accuracy, and prior work mostly attacks this with outlier-smoothing transformations. HyQuant instead keeps a small accuracy-critical subset in high precision — the vertical-line tokens that attention repeatedly attends to, plus a local sliding window — while quantizing everything else, using lightweight attention-pattern signals to pick that subset. The same principle applies in both phases: a hybrid-precision attention operator during prefill, and KV-cache compression with dequantization fused into the attention kernel during decode, yielding near-lossless accuracy across models and tasks from a simple design.

Entity-Memory Graph Retrieval Improves Evidence Coverage in Long-Conversation Question Answering

Shumao Sun Long multi-session dialogue strains retrieval, because ranking by dense cosine similarity can drop a neighbouring turn that supplies necessary context. The retriever described keeps dialogue turns as verbatim memory nodes, links repeated mentions through shared entity nodes, connects adjacent memories with directed chronological edges, and answers queries by entity gating, semantic fusion, one-hop chronological recovery, then dense backfill — compared against a dense control matched on vectors, context budget, answer protocol, and evaluator to isolate the effect of graph structure. On 1,986 questions from ten LoCoMo conversations, evidence recall at top-25 rises from 79.7% to 84.5%, with the advantage holding from top-5 through top-50; notably, no matched cutoff supports any difference in final-answer F1, so the demonstrated gain is in retrieval coverage alone.

When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation

Siyuan Gan, Yuhan Li, Xiran Wang, Linjian Meng, Boyan Wang, Zhen Zhao et al. On-policy distillation has the teacher score the student's own generated prefixes, but that guidance is not always trustworthy: it can discourage moves toward correct trajectories or push the student toward incorrect ones, working against the outcome reward the training is supposed to maximize. Reward-Aligned On-Policy Distillation (RA-OPD) checks, for each sampled trajectory, whether its trajectory-level distillation return agrees in sign with its outcome reward, and drops the trajectories where the two conflict — a filter that costs no extra compute. Across seven math and three code benchmarks with Qwen3 and DeepSeek-R1 family models, it significantly outperforms standard on-policy distillation and the other variants tested.

QUORUM: QUality-Optimized Routing Using Multiple annotators

Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu, Amin Mantrach, Fabrizio Silvestri Large language models can annotate data far more cheaply than humans but their reliability varies sharply per instance, handling easy inputs well while failing on examples needing nuanced reasoning or context. QUORUM is a budget-aware routing framework that decides per instance whether to send it to a human or an LLM annotator under a fixed budget, estimating difficulty from feature-based signals instead of model confidence or uncertainty, and allowing several annotations per instance that are combined through agreement-based rewards. Across closed- and open-ended annotation tasks in English and multilingual settings, it improves annotation quality by up to 34.4% while cutting cost by 8.8% relative to competing routing methods, with code released publicly.

A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint

Artem Safronov Rather than quantizing every layer uniformly as GPTQ or AWQ do, or allocating bit-widths without demonstrating actual speedup as MixLLM and TorchAO do, this work frames per-layer bit allocation for Gemma-3-1B as maximizing latency reduction subject to a budget on generation-quality loss, reusing a layer sensitivity profile from the authors' earlier SA-PTQ work through TensorRT-LLM's activation pass-through mode. Measuring 13 W8A8 variants on an RTX 5090 across block groupings (5+5, 10+10, all 26 layers), they find integer arithmetic pays for the quantize/dequantize overhead in feed-forward networks and the language-model head, but not in attention at short context lengths, where the extra step is a net slowdown; a manual SmoothQuant implementation was needed because export failed. The best configuration under minimal degradation, feed-forward 5+5 plus the head, gives an 11.0% latency reduction at 98.90% Top-1 agreement and +0.85% perplexity, rising to 19.1% speedup with looser quality tolerance.

Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning

Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Renqiang Luo et al. Knowledge-intensive question answering requires models to ground answers in supplied evidence, but entity mentions in context can trigger memorised associations that yield plausible answers unsupported by that evidence, and existing abstention methods based on uncertainty scores or evidence-sufficiency checks never test grounding directly. Twin Worlds instead builds several parallel versions of an input through typed entity substitutions that preserve the relational structure while stripping away the model's parametric priors, then checks equivariance: a grounded answer should shift in step with the substituted entities rather than staying fixed. Violations of that expected correspondence are used as the abstention signal, and across four benchmarks and three model backbones the method identifies ungrounded answers more reliably than uncertainty- and sufficiency-based baselines.

Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms

Prabhu Vellaisamy, Vanessa Lam, Shawn Blanton, John Paul Shen cross-listed Inference services bill per token while GPUs consume energy over whole inference windows, an accounting mismatch that lets average per-token energy fall even as total request energy climbs. A decomposed model splits consumption into a one-time prefill plus generation setup cost and a marginal per-step cost, measured on NVIDIA H100 and H200 GPUs across dense and mixture-of-experts models as a function of batch size, context length, and output length. For Llama-3.2-1B on an H200 at batch-16 and 4K context, stretching output from 10 to 512 tokens drops token energy from 7.46 to 0.72 J/token while total window energy rises from 1.19 to 5.93 kJ; batching also lowers token energy but the benefit shrinks with longer context (6.31x at 512-token context versus 1.17x at 4K), and sparse expert routing inflates fixed energy at low concurrency in a way batching largely erases.

SEPO: Evidence-Grounded Prompt Optimization via Structural Editing

Xiaoyu Ma, Haoyue Liu, Yiwen Li, Jionghao Zhu, Zhichao Wang, Ye Chen et al. Automatic prompt optimizers that only call an API are usually called interpretable, but each iteration still rewrites the whole prompt as one opaque string, so the record left behind is a series of full-prompt diffs rather than edits anyone can localise or reason about. SEPO (Structural, Evidence-grounded Prompt Optimization) runs multiple search trajectories that make local edits to stable, typed units within a two-layer prompt schema, links each edit's intended and realised structural operation to the specific examples it newly fixes or breaks, and carries that edit-effect record forward to inform later proposals on the same branch. On a 14-task held-out suite it beats GEPA by 3.1 points on Llama-3.1-8B-Instruct and 2.2 points on Qwen3-8B, reaching 61.9% and 73.3% macro accuracy while spending 2.9M optimisation tokens against GEPA's 4.1M and producing prompts more than five times shorter.

H-Scale: Hessian-Guided Scale Refinement for NVFP4 Sub-Byte LLM Inference

Hao Yu, Zheng Li, Dayiheng Liu, Jianwei Zhang NVIDIA's Blackwell architecture natively supports NVFP4, whose micro-blocks of 16 weights capture local distributions and isolate outliers well but create a large, sensitive space of per-group scaling factors that existing post-training quantization work has mostly ignored in favour of refining the quantized weights themselves. H-Scale is a lightweight post-processing step that picks hardware-valid group scales using a diagonal second-order proxy derived from calibration activations, so the objective targets layer output perturbation rather than plain weight reconstruction error. It drops into existing NVFP4 pipelines as a replacement for round-to-nearest scale selection, needs only modest offline calibration, and adds zero inference-time overhead while improving a broad range of NVFP4 baselines and moving several variants closer to the BF16 reference.

The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues

Farah Atif, Sougata Saha, Monojit Choudhury Computational work on social power in dialogue has been confined to narrow linguistic and cultural settings, with datasets lacking the demographic and relational detail needed for cross-cultural comparison. The authors pair a schema grounded in social science theory with a native-speaker annotation pipeline and a cross-lingual analysis interface, producing a corpus of 15,836 annotated instances from 100 scenes in French and Egyptian Arabic films. Annotators agree strongly on observable demographic and contextual attributes but diverge on interpretive ones such as power asymmetry and intention alignment, and an evaluation of six large language models and multimodal LLMs finds persistent gaps between human and model agreement on relational and theory-of-mind reasoning.

Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

Christos Koutsiaris Because a byte-level byte-pair-encoding tokenizer is an ordered list of merge rules, applying only a prefix yields a nested vocabulary whose token ids are the first rows of the full one — so a single model could in principle serve several vocabulary sizes and be deployed at any of them by slicing its embedding and output head. The authors pre-registered five claims with margins, seeds, and a stop rule, then trained 30 models with 3.1M- and 10.6M-parameter bodies on 200M tokens each. Slicing is bit-for-bit exact across 76 checks and removes 66% of deployed weights with no latency change, but the shared model trails a fixed-vocabulary specialist by 3.64% bits per byte at 32k against a 1% pre-registered margin, with a 2×2 ablation attributing the cost to output restriction rather than the control token; multi-cap training does buy robustness, degrading 12.5–15.4 points less under typographical noise.

FinExam-10K: When Retrieval Helps Financial Reasoning?

Yan Lin, Jingyu Sun, Zhongliang Guo, Qing Li, Zhuohan Xie, Yuxia Wang Professional finance exams demand domain knowledge, calculation, and judgment together, yet no benchmark had covered the full structure of the Chartered Financial Analyst (CFA) and Financial Risk Manager (FRM) programs under one protocol. FinExam-10K contributes 10,198 expert-reannotated questions across CFA Levels I–III and FRM Parts I–II, half released and half sequestered for a maintained leaderboard, split into a full-coverage track and a context-complete track where the supplied record suffices to answer. Across 17 models the best overall accuracy is 85.29% but drops to 34.68% on the frozen hard band, and all 17 fail the same 47 context-complete items; retrieval augmentation rescues hundreds of errors while overturning as many correct answers, for little or negative net gain, and a learned gate that fires retrieval on 7.9% of questions lifts held-out accuracy only from 70.83% to 71.23%.

Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance

Vincenzo Collura, Karim Tit, Eleonora Giunchiglia, Mike Papadakis, Maxime Cordy Grammar-constrained decoding keeps language model output syntactically valid for formats like JSON, SQL, and code, but typical decoders only enforce local prefix feasibility, so a prefix that is extendable in principle can still fail to reach acceptance under tokenizer-grammar mismatch and a finite token budget. The proposed decoder precomputes bounded pushdown-automaton summaries offline, labeling reachability and an upper bound on distance to acceptance, then uses those estimates online for horizon-aware pruning and beam search. Every output is guaranteed to be accepted by the target context-free grammar, and experiments on JSON, SQL, and Linear Temporal Logic (LTL) report both consistent validity and better completion quality than existing baselines.

Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation

Linze Wu, Xinrui Chen Agents that emit JSON, SQL, or function calls fail outright when one field is wrong, and constrained decoding already tracks parser transitions that reveal which tokens carry schema-critical decisions — a signal current key-value (KV) cache compression ignores. PASK (Parser-Aware Structural KV Persistence) converts those parser states into layer-group-specific decisions about which KV entries to keep, using task-error sensitivity to set minimum protection floors and attention-output distortion to allocate what capacity remains; an offline calibration step compiles this into a policy so only a lightweight lookup runs at inference. At a total KV budget of 0.33 on Qwen3-4B, it beat the strongest compressed baseline by 17.39 percentage points on average across eight BFCL non-live and Live subcategories, and in serving reached up to 2.2x higher throughput and 3.3x lower time-per-output-token at roughly half the peak GPU memory of a full cache.

A Probabilistic Interpretation of KV Cache Eviction

Renato Geh, Alex Chen, Daniel Israel, Aditya Grover, Guy Van den Broeck Dropping entries from the key-value (KV) cache buys throughput at supposedly negligible quality cost, but the selection rules in the literature are largely heuristic and the problem itself has never been stated formally. A probabilistic formalization shows the exact eviction problem is computationally hard, then reframes it as expectation estimation that can be approximated by sampling — which also makes it feasible to correct for evicted entries during decoding, something prior work ignored. Under this view existing eviction methods are zero-variance biased estimators, and adapting them to add decode-time correction gave more robust behavior across tasks at competitive performance for the same compression budget.

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

Niccol\`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J\"org Franke, Sampo Pyysalo, Jenia Jitsev et al. Pretraining dense large language models on English-prevalent corpora, this study maps how optimal learning rate and batch size scale jointly and how each evolves marginally with model capacity and data budget, fitting a model of those relationships. Using a Warmup-Stable-Decay schedule, it also measures what learning-rate annealing buys across a wide sweep of settings and whether hyperparameters chosen for the stable phase transfer to the decay phase, then evaluates recently proposed loss scaling forms that explicitly model the capacity-data interaction. Those forms captured both undertraining and overtraining regimes well across the experiments, and the complete collection of pretraining runs is released open-source as a baseline for future OpenEuroLLM models.

When Linguistic and Internal Confidence Diverge in Large Language Models

Hefan Zhang, Bingquan Zhang, Ming Cheng, Saeed Hassanpour, Weicheng Ma, Soroush Vosoughi Asking a language model how confident it is only helps if that spoken confidence tracks the model's internal uncertainty, which this study tests across 8 classification tasks, 2 generation tasks, and 30 models from three families, comparing verbalized confidence against logit-based confidence and semantic entropy. Instance-level association between stated and internal confidence is weak on average, improving only on easier items and stronger base models; instruction-tuned models report higher confidence and sometimes correlate better but show larger confidence gaps and worse calibration. Prompt changes mostly shift the distribution of reported numbers rather than the underlying alignment, with attitude cues inflating confidence for nothing and score exemplars preserving rank-order signal only when they avoid collapsing onto a few values. The authors describe verbal confidence as a lossy channel that can carry useful ranking information without being calibrated, and argue it needs multi-axis diagnostics before use in reliability pipelines.

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi et al. Cultural evaluation of language models is usually multiple-choice factual recall, which misses the common case of a user asking for practical help over several turns in a culturally specific situation. CultureConverse is a simulation and evaluation harness covering 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains, where an assistant must infer cultural constraints from partial information over a scored multi-turn episode; the released CultureConverse-DS dataset holds 14,610 benchmark episodes and 274,295 oracle-guided dialogues. Across 18 evaluated models GPT-5 mini scored highest on assistance quality, human annotation supports the automatic judge as a proxy, and fine-tuning on 27,860 high-quality samples improved in-domain assistance while transferring to out-of-domain cultural multiple-choice and safety classification benchmarks.

Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL

Jiayan Lin, Yujia Liu, Zijin Hong, Zheng Yuan, Yilin Xiao, Hao Chen et al. In-context learning (ICL) text-to-SQL systems keep stacking modules around the base generator, yet papers report only aggregate end-to-end accuracy, leaving the marginal accuracy and cost of each design choice unquantified. The authors implement 17 paradigm-level configurations spanning five recurring pipeline modules in one controlled codebase and attribute each one's contribution and token cost across four backbones of differing capability and reasoning style. Execution-feedback refinement is the only paradigm whose benefit holds universally and at consistently low cost, most other modules help only under backbone-dependent conditions, input token demand tracks pipeline structure while output demand tracks backbone generation behavior, and a fixed budget is often better spent on a more elaborate pipeline over a mid-tier backbone than on a frontier model with a lean one; the resulting tiered guideline transfers to five additional backbones without re-running the search.

Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining

Shuchen Zhu, Yuxin Fang, Mingze Wang, Kun Yuan Pretraining dominates LLM training compute, yet noise-dominated gradients and an ill-conditioned loss landscape leave progress slow along flat directions — the small-eigenvalue directions that drive most of the final loss reduction — which adaptive optimizers like AdamW and Muon mitigate only indirectly through gradient normalization. The proposed optimizer applies multiscale momentum solely along flat directions, pairing a slow-decay component for noise reduction with a fast-decay component for rapid curvature adaptation, and adds a sphere constraint to prevent the parameter inflation and overly fast effective-learning-rate decay a naive combination would cause. Experiments show it significantly accelerates Muon across dense and mixture-of-experts architectures at 0.12B to 2.3B parameters, with theoretical analysis supporting the flat-direction design.

Sliding-window beats linear attention

Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais Retrofitting pretrained LLMs with linear attention has been promoted as the fix for quadratic attention's ever-growing key-value cache, but the authors argue this line of work has never been compared against simple baselines. Testing sliding window attention (SWA) with attention sinks against post-trained linear-attention models across several LLMs and downstream tasks, they find SWA matches or beats them, and on long-context tasks (Needle-in-a-Haystack and BABILong) scores 2 to 10 times higher. Since SWA requires no post-training and is fast and memory-light, they recommend it over converting models to linear attention, which they suspect would need training from scratch or extensive post-training merely to match.

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

Zhuoshi Pan, Junru Lu, Yan Qian, H. Vicky Zhao, Di Yin, Xing Sun Factual question answering benchmarks assume a single canonical answer, which hides whether a model retains genuinely divergent accounts of long-tail facts. ElephantBench is a closed-book probe of 1,094 questions built by an auditable graph-based pipeline that retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, converts them into multi-account records, and verifies each answer against the source documents, authoritative public sources, and human annotators. Across 32 models, the strongest recovers both accounts on only 52.4% of questions, recalling one and omitting the other on nearly all the rest; larger models and inference-time reasoning improve recall without eliminating the incompleteness, and corpus analysis links completeness to how much exposure the minority account has.

How Proper Scoring Rules Shape LLM Forecasting

Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock Five proper scoring rules are compared as reinforcement training objectives for LLM forecasters predicting binary outcomes of already-resolved real-world events. Although all five theoretically incentivize truthful probability reporting, the resulting models differ substantially in calibration, probability usage, and their decomposition into bias, information, and noise, even when aggregate accuracy is similar — the Brier-trained model wins on Brier score and AUC-ROC while the log-score-trained model wins on log score and calibration error. The takeaway is that reward choice shapes the structure of forecasting errors, not just their magnitude, though each condition used a single training seed so some gaps may be stochastic.

Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation

Di Wu, Sergey Troshin, Christof Monz, Antske Fokkens, Vlad Niculae Test-time scaling comes in two flavors — sequential, where each attempt conditions on the previous one, and parallel, such as independent sampling with reranking — and their relative merits for machine translation had not been characterized. Across sampling budgets, sequential sampling reaches a higher performance ceiling and yields a more diverse candidate pool, especially at small budgets, and multidimensional human analysis of Best-of-N outputs shows it mainly improves fluency and naturalness while sometimes degrading accuracy at large budgets. Controlled experiments attribute part of the gain to the model seeing more target-side context, with ablations showing robustness to temperature but sensitivity to how that context is assembled.

Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration

Simeng Sun, Roger Waleffe Training Mixture-of-Experts (MoE) language models with expert parallelism spends a large share of wall-clock time on all-to-all token dispatch and combine collectives. The communication-efficient variant CE-MoE decouples token-mixing depth from channel-mixing depth: instead of placing an MoE layer after every attention or Mamba-2 block, it concentrates expert capacity in a few routed MoE layers and restores depth with extra token-mixing and dense feed-forward layers. Over a scaling ladder from 2B to 31.5B total parameters with total and activated parameters matched, CE-MoE matches full-MoE validation loss and downstream scores while cutting training cost, and at 31.5B it uses 33.3% fewer GPU-hours while improving average downstream score and inference throughput.

DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging

Aaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian Merging several task-specific fine-tuned LLMs into one multi-task model without retraining introduces representation bias, a systematic drift between the merged model's hidden states and those of each source model; prior correction work targeted encoder-based vision models. Two complications specific to decoders are identified: the causal attention mask lets bias accumulate across token positions, requiring position-dependent correction, and high-entropy decision-critical positions matter far more than low-entropy ones. DARTS addresses both with an entropy-weighted L1 loss that upweights correction where errors most affect generation, plus a per-position additive bias term. On Llama-2-7B merges evaluated across HumanEval, GSM8K, and AlpacaEval, it improves substantially over standard surgery while adding only 0.1% extra parameters.
7 more specialized papers

Agents 46

PACE: Publisher-Adaptive Content Extraction via Agentic Automation

Zhanlin Liu, Munirathnam Srikanth Web content extraction for LLM data pipelines forces a trade among general extractors that break on publisher-specific layouts, direct LLM extraction that is costly and slow at scale, and hand-written parsers that require ongoing human maintenance. PACE uses LLMs offline to analyze representative pages from a publisher and aggregate reusable extraction patterns into a configuration; at inference time that configuration instantiates a fixed deterministic extractor template, so extraction itself makes no LLM calls. Across article-body, metadata, and multimodal targets including images and tables, it beats scalable non-manual baselines and approaches the quality of manually engineered publisher-specific parsers.

Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields

YuJie Huang, WenWu He, ZhuoEr Lin, Congcong Liu, Dong Liang, Zhuo-Xu Cui Recovering a partial differential equation (PDE) in heterogeneous media means identifying the governing operator and the unknown spatial fields that parameterize it at once, and the two are entangled — a sufficiently flexible field can conceal a structurally wrong law when only one trajectory is available. HER-PDE runs a hypothesize-evaluate-refine agent loop that reads two noisy trajectories from different excitations, proposes complete expression-tree hypotheses, and scores them through an evaluation interface that estimates only the fields a hypothesis explicitly declares, never silently adding terms, ranking candidates by bidirectional cross-excitation transfer before auditing the winner on a sealed time interval. Across five two-dimensional systems observed with 5% relative Gaussian state noise it recovered the generating operator in all five cases, including equivalent signed-field and product-rule parameterizations, with the nine unknown coefficient fields reaching median Pearson correlation near 0.85 and median relative L2 error near 0.28.

Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

Yiqi Zhu, Feiyu Gao, Jiaxing Fan, Jiahui Zeng, Minggang Wu, Chenliang Li et al. Existing mobile agent benchmarks such as AndroidWorld and MobileWorld cover a narrower slice of app use than real users exercise, so GMA adds seven applications built from open-source projects spanning domains like lifestyle sharing and travel planning, plus 300 tasks across four difficulty tiers from atomic actions to long multi-step workflows. Across eight frontier models, performance declines substantially as task complexity increases, leaving current agents well short of handling realistic user requests reliably. Controlled ablations of agent harness choices — context retention and explicit state tracking — run under a shared environment, model setting, and task taxonomy show harness design meaningfully helps on demanding workflows, though which specific design helps varies by foundation model.

Thinking Costs Tokens: When More Structure is Worth the Price

Thomas Nolasque, John Grey, Calista Pham, Ankit Vani Adding search, verification, and revision scaffolding to a language model consumes the same token budget it is meant to spend wisely, raising the question of where the break-even point lies. Two systems were compared on the FinQA and TAT-QA financial reasoning benchmarks using GPT-5.4 mini across 14 budget tiers from 250 to 42,000 output-equivalent tokens and 1,000 cases each: a single-call monolith versus a verified search architecture with planning, label-blind checking, and repair. At 1,000 tokens the monolith reaches 18% accuracy while verified search scores near zero because planning overhead leaves no room for an answer, but from 1,500 tokens onward verified search leads, topping out around 44% versus 40%. The crossover falls between 1,000 and 1,500 output-equivalent tokens, confirmed by an intersection-union test at p ≤ 0.001.

WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning

Yu Han, Tianwen Qian Reinforcement learning for mobile graphical user interface (GUI) agents normally requires large volumes of live Android interaction, which is slow, expensive, and unstable. WM-R1 replaces the real environment with a world model that supplies all state transitions during rollouts, and additionally embeds the world model in the agent's chain of thought so it can simulate the consequences of candidate actions before committing. The design permits massively parallel, step-level trajectory generation and uses a multi-dimensional rule-based reward covering task success, trajectory efficiency, and world model usage, trained on a curated set of 2,000 hard tasks. On Android benchmarks the resulting agents outperform both GRPO-only baselines and inference-time simulation methods.

If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary

Marc Millstone, Tyler Akidau, Johannes Br\"uderl, Marat Pekker An agent handed a human's credential inherits that person's full reach without the judgment that normally constrains it, and every over-broad query it makes still looks credential-valid to the backend. Out-of-Band Policy Enforcement (OBPE) puts a trusted boundary outside agent reasoning that authorizes the typed operation and resource, narrows the query before the backend call, then filters records and fields or masks values on the way back, with semantic gating able to deny or hold a call based on argument values or external state; the authors prove the policy plan is order-independent and that agent policy can only narrow, never widen, the data owner's ceiling. Benchmarking prompted agents against Jira and ServiceNow mocks across four models with 20 adaptive red-team tasks, trace failures fell from 57.6% to 0.2% over 3,621 trials while task fulfillment dropped from 79.1% to 60.9% and paired safe-and-useful completion rose 21.8 points. They release an HTTP proxy prototype with a typed Cedar policy core, and note residual leaks where answers reconstruct values that never entered context or use filtered row counts as an oracle.

First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents

Syed Mahbubul Huq, Pranava Madhyastha Qwen-GuidePlay-2B is a 2-billion-parameter model for dialogue-game interaction, fine-tuned from Qwen3.5-2B in three stages: supervised fine-tuning on only successful Playpen trajectories, then weighted turn-level fine-tuning, then teacher-guided fine-tuning where a larger model fixes formatting and scores examples but never writes new gold actions. It reaches 57.12 clemscore on the public validation set and the second-highest clemscore delta among challenge submissions, about +36 over its base model. Imitating whole trajectories appears to buy basic playability while turn-level and teacher-guided training improve decision quality, and more elaborate procedures such as replay-repair and hard-example mining did not help, suggesting careful data curation matters more than aggressive training changes at this scale.

Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation Community

Seth Carbon, Sierra Moxon, Kimberly Van Auken, Pascale Gaudet, Christopher J. Mungall Biological database curators could benefit from agentic AI but are blocked by practical barriers — no access to agents, no local tooling, no training. The Gene Ontology Consortium addressed this with a cloud environment built on JupyterHub using Claude Code as a universal harness, letting curators drive an agent session from a browser terminal through a single shared API gateway with no subscriptions or local installs, plus four training modules progressing from basic tool use to pathway curation with the existing GO-CAM (Gene Ontology Causal Activity Model) tool. Thirty-seven participants completed the four-hour workshop, and the authors conclude that building agentic capability in a distributed scientific community is mainly a matter of removing access barriers, designing workflows, introducing capabilities gradually, and giving curators hands-on practice evaluating agent output.

PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation

Krishna Rao, Andrew Dumit, Shaena Ulissi, Jacob Feintzeig, P. James Joyce, Daniel Frank et al. Estimating a product carbon footprint (PCF) — the greenhouse-gas emissions attributable to a physical product — is increasingly handed to LLM agents, but scoring only the final total lets errors cancel out and hides where the reasoning broke, while scoring subtasks in isolation misses compositional effects. PCFBench carves the workflow into six independently gradable tasks spanning decomposition, retrieval, ontology matching and numerical extraction, with 614 expert-labelled items probing under-specification, conflicting context and numerical constraints. Across eight frontier LLMs from four providers no model dominated: the strongest came within a factor of two of declared totals on 77% of products, but that rate dropped to 37–58% when the footprint was built up step by step, with only 45–75% of answers respecting mass conservation.

The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs

Eric Yeats, Brendan Kennedy, Loc Truong, John Buckheit, Jung Lee, Jesse Friedbaum et al. Hidden states carry information about model behavior that is hard to extract from inputs and outputs alone, which matters once a model's tool calls reach external systems. Linear probes trained on those hidden states were evaluated for detecting incorrect tool calls across 18 tool-calling LLMs on the Berkeley Function Calling Leaderboard. Probing caught a range of error types, including arguments with the correct type but the wrong value, which standard logging frameworks would not record, with effectiveness depending on model size, probe layer and post-training type; probes also generalized to error types absent from their training data.

Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models

Justin Bronder A tool-equipped language model can commit to a final claim its evidence does not support even when one available call would settle the question and its instructions explicitly forbid guessing. The failure is split into two measurable quantities: occurrence, judged from visible evidence and the final claim without consulting the answer key, and conditional repair, measured by replaying each case from an exact copy of its state with alternative tool responses matched in structure and length that differ by a single character. On a fixed Qwen3-32B setup, 33 of 512 first responses ended in an unsupported claim, and resolving evidence repaired all 33 while a matched but uninformative response repaired none; an automatic checking rule added 21 evidence calls over 64 cases, corrected all 10 wrong claims and never turned a correct answer wrong, whereas a Gemma 4 setup always called the tool and produced no unsupported claims.

Credo: Reusable Declarative Primitives for Agentic Workflows

Duo Lu, Andrew Crotty, U\u{g}ur \c{C}etintemel An LLM application depends as much on its harness, the program deciding what each call sees, how many calls to make, and which answers to trust, as on the model, and coding agents that search for good harnesses emit opaque imperative code whose logical steps, runtime signals, execution decisions, and prompt strategies stay implicit, forcing every new task to restart the search. Credo recovers a structured declarative description from a searched harness, tags each extracted primitive with metadata, and catalogues everything with provenance so that a compiler can bind stored primitives into harnesses for new tasks without searching from scratch. The paper reports preliminary results and lays out a research agenda aimed at the database community, including cost-based compilation over declarative catalogs and catalog maintenance under model and workload drift.

ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL

Pratik Kakkar, Chandra Dhir, Ravi Shankar, Pareekshit Reddy Gaddam, Anup Shirgaonkar Reinforcement learning from execution feedback has pushed text-to-SQL accuracy up, but most approaches treat SQL generation as a single turn, leaving no room to recover from errors. ReToolSQL pairs a supervised warm start on rejection-sampled reasoning traces with agentic reinforcement fine-tuning over multi-turn tool-use trajectories, on the argument that the two act on complementary axes: supervised fine-tuning expands which questions are solvable at all, raising pass@k coverage on the hardest cases, while reinforcement fine-tuning converts that coverage into single-pass accuracy by teaching the model when to verify, what evidence to retrieve, and how to repair faulty SQL. Applied to a 31B instruction-tuned Gemma 4, the combined pipeline reaches 74.32% execution accuracy on the BIRD-SQL development benchmark single-pass and 74.77% with self-consistency, reported as first on the single-model development leaderboard at the time of writing, using composite rewards anchored on execution correctness and no human annotation beyond the benchmark.

CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action

Lekai Chen, Alvaro Velasquez, Ashutosh Trivedi Instructions given to embodied agents in natural language usually carry constraints that must hold as the world changes, and the free-form programs an LLM writes offer no stable object to verify, compose with new constraints, or repair from a failing trace. CEDAR grounds instructions as regular languages over environment event traces, using a language model for semantic judgments and execution traces for counterexample-guided correction, and represents both learned skills and specifications as deterministic finite automata, so a skill can be intersected with a constraint such as sleeping at night or staying in one biome to yield a controller that enforces the constraint by construction rather than by repeated prompting. In Minecraft, given the same simulator observations as a program-generating baseline, it maintains temporal and spatial constraints the baseline fails to preserve while amortizing skill reuse and reducing cumulative LLM queries.

CURA: Certified Runtime Alarms for Computer-Use Agents

Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo et al. Self-report is the cheapest oversight channel a deployer has over computer-use agents, and it fails exactly where oversight matters: across 361 OSWorld tasks, a pipeline of read-only feasibility gate, planner, and GUI executor reaches a mean score of 82.9 against a 72.4 human reference, yet 64 of its 71 failures end with a success claim and the explicit failure affordance is never used in roughly 9,100 calls. CURA is an external monitor that reads only harness-visible telemetry, with no model internals, extra model calls, or prompt changes, turning the running trajectory into a sequential test with certified false-alarm control. At a 0.10 alarm level its CUSUM detector catches 42.3% of failures a median of 31 steps before termination at a realized false-alarm rate of 0.066; retrospectively its margin over a total-token baseline is not statistically significant, but online it recalls more at matched certified budgets, and alarm-gated escalation to a frontier overseer recovers 23 of 70 failures for a mean score of 86.8.

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim, Kyuhong Shim, Sunjae Lee Coding agents are graded on the SWE-bench family, whose tasks come from curated GitHub issues that are long, structured, and information-rich, whereas real user requests are short and unstructured. Applying a six-category information taxonomy and four dimensions of linguistic style to real prompts from SWE-chat versus problems from SWE-bench Verified and Pro, the authors find that requests carrying only a problem statement account for 88% of real prompts but 7% of benchmark problems, and 87% of real prompts are casually written against 94% formal benchmark problems. RealSWE rebuilds 381 task families whose variants share a task and gold patch but differ in information composition and style, and across seven contemporary models realistic inputs lower resolution rates by 6.4 percentage points on average and can change model rankings. Controlled analysis shows that including desired behavior and motivation significantly affects performance while environment information and reproduction steps merely add tokens, and that style has only small model-dependent effects, so users get a concrete instruction: say what you want to happen and why.

FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling

Jun Bai, Ruilin Wang, Yue Li Autonomous agents that build predictive models from electronic health records (EHRs) are limited to one hospital's data and tooling, and patient privacy blocks direct collaboration, while conventional federated learning only shares model parameters and discards the modeling know-how an agent accumulates. FedEHR-Agents federates that experience instead: each hospital's agent handles preprocessing and model development and refines local modeling experience through historical memory, task-specific evaluation, and TextGrad-based prompt refinement, while a server aggregates experience under evidence-based weighting and distills it into global meta-prompts. Across real multi-hospital EHR benchmarks it outperforms both local and federated baselines on diverse clinical prediction tasks and holds up across federation sizes and LLM backbones, positioning accumulated experience rather than weights as the object of collaboration.

See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs

Sarang Manoj Pekhale, Amartya Roy, Rajat Sarkar, Souvik Chakraborty Recovering the partial differential equations (PDEs) that govern an observed system is hampered by sparse-regression methods needing a predefined term library, symbolic regression's noise sensitivity, and language-model approaches hallucinating without iterative correction. MAGE structures discovery as a confidence-governed hypothesis-validation loop run by four specialized agents: one computes derivatives and diagnostic plots, a vision-language model extracts qualitative cues from those visuals, a language model proposes candidate equations without any fixed library, and an arbiter fits coefficients and scores confidence until a threshold is cleared. On a canonical PDE suite it achieves exact structural recovery on 8 of 8 systems with the lowest coefficient error on 7 of 8, improving accuracy by up to four orders of magnitude, and it recovers expected operators in complex geometries and fits a cubic restoring-force model to lab sensor data at held-out R² of 0.985.

Resource Constraints and Performance in Agentic AI Systems

Amaz Salman, Malka Halgamuge, Teo Susnjak Agentic systems bundle a language model with tools, memory, state, and multi-step execution, and those mechanisms drive operating cost as much as capability, so the authors compare OpenClaw and NanoBot end to end on a paired benchmark plus an instrumented subset. Full task completion was 31% versus 25% respectively, a gap whose 95% bootstrap interval spans −3 to 15 percentage points and so establishes no advantage for either; on the instrumented subset both hit 26% full completion while NanoBot reached partial completion on 43% of prompts against 26%. OpenClaw was slower on 83% of prompts and used more peak memory on every one, with geometric-mean ratios of 2.98× wall time and 19.44× peak memory, though many of NanoBot's apparent dominance cases were simply cheaper joint failures — a discrepancy between the two evidence layers that argues for tying capability and resource numbers back to per-attempt execution records.

LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages

Injun Baek, HyeongSeok Lee, Yearim Kim, Junhoo Lee, Nojun Kwak Prompting a language model directly for landing-page code tends to yield generic templates and unsupported persuasive claims. LandingBench is a dataset that abstracts real landing pages into reusable reference profiles — section sequences, layout patterns, tone descriptors, visual emphasis, and call-to-action (CTA) structure — and LandingAgent is a three-phase agent that profiles the target, builds a reference-guided wireframe, then refines the page through critique-guided polishing. Measured against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity, it shows improved target grounding, presentation quality, and layout diversity.

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

Ji'an Lei, Jian Huang Agents built on small language-model backbones are cheap but can settle into persistent failure modes, so a deployment needs a policy for when to escalate to a larger, costlier backbone. TACIT-SWITCH learns permanent handoff policies from Teacher-Annotated Censored Intervention Times, representing each annotation as an interval-censored observation on a cumulative-risk scale and fitting a mixture-cure threshold model that estimates both whether the strong rollout would succeed and, given success, where the handoff threshold sits — with no teacher needed at deployment. In a mechanism-based multi-step simulation it beats task-level, step-level, and fixed-prefix routing by 7.4 to 11.1 percentage points of success at comparable cost, and achieves the highest held-out success among learned policies on both ALFWorld and DABench.

What Makes Agent Memory Useful for Reliable Unanswerable Question Handling?

Chuanyuan Tan, Junjie Yu, Yuxin Wang, Yining Zheng, Xipeng Qiu, Wenliang Chen Trustworthy agents have to recognize when a question simply cannot be answered, and it is unclear whether the memory modules common in agent frameworks actually help with that. Four representative memory methods are compared across three unanswerable-question datasets and two base models inside a single agentic retrieval-augmented generation setup. Gains prove selective rather than universal and fall apart under dataset shift; reusing memory across base models turns out to be easier than reusing it across datasets, and procedural, rule-based memories that carry decision guidance transfer more reliably than memories that shape trajectories or merely accumulate more experience.

openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents

openJiuwen Team, Tao Yu, Xinyu Zhang, Qianqian Chen, Xiaoneng Xiang, Chia Kwangyang et al. Coding agents that work over long horizons need harnesses that let developers compose heterogeneous capabilities and delegated sub-agents without rebuilding orchestration each time, and that let new evidence — semantic diagnostics, execution outcomes, task progress, shifting context relevance — steer later runtime decisions. openJiuwen is an open-source harness targeting these two properties, called Structural Composability and Runtime Adaptivity, via a shared execution substrate with Rail-based capability composition spanning single agents, delegated sub-agents, and a Swarm Flow mode, plus framework-controlled adaptation of context, feedback, and task control around a fixed model policy. It scores 82.6% on SWE-bench Verified and 87.19% on Terminal-Bench 2.1, 3.4 and 3.39 percentage points above the strongest selected official-leaderboard point estimates.

When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems

Yangxiao Jiang, Jiarun Fan, Mingcong Xu, Yanxi Guo, Jiwen Feng, Shanqing Xu et al. Multi-agent systems increasingly generate their collaboration topology on the fly, but current generators draw on the language model's parametric knowledge and treat search or retrieval as a reactive tool rather than something that shapes the structure itself, producing redundant interactions or too little verification on knowledge-intensive tasks. K-GAT (Knowledge-Guided Agent Topology Generator) treats topology design as knowledge-conditioned structure learning in a neuro-symbolic framework, feeding external evidence directly into autoregressive graph generation. On the expert-level GPQA benchmark it beats an LLM-Debate baseline by +15.7% accuracy while using less than half the tokens.

GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies

Yige Luo, Ran Guan Running a generative-agent society is easy; inspecting one is not, since operators typically get either a finished replay or raw logs across many agents, locations, messages, and model calls. GOD (Govern, Observe, and Direct) is a local-first browser control room combining a setup wizard, Agent Studio, Map Studio, a spatial replay interface, Ask and Intervene commands, and portable experiment, map, and agent packs, so an operator can pose targeted questions or interventions and immediately inspect the resulting replay state. Its design contribution is a shared command-and-artifact loop where live controls and replay evidence use the same operator command model while package contracts keep scenario, map, and profile data separate from local runtime state. Over 15 run slots, 78 of 84 target-agent checks in the 14 intervention runs recorded the commanded destination and 169 of 182 state answers matched a saved location or action string.

Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents

Yuxu Ge Trajectory-level credit assignment can pinpoint which module of a tool-using language-model agent caused a failure using only verifiable signals, and this study tests whether that credit should steer a fixed zeroth-order/evolution-strategies perturbation budget across modules. Across a synthetic environment, frozen Qwen2.5-1.5B, Qwen2.5-3B, and SmolLM2-1.7B agents, three task families, six allocation schemes, credit-noise sweeps, paired seeds, and exact sign-flip tests, no allocation scheme beat uniform by the 2-percentage-point threshold in any on-pool comparison, and concentrating the whole budget on the credit argmax was significantly worse on the 3B model. Loss scales linearly with bottleneck starvation rate (R² = 0.94, descriptive), a credit-free coverage floor removes the detected harm, and matched-budget burst and catch-up schedules point to insufficient cumulative parameter movement rather than update frequency as the cause; the one exception is soft routing beating uniform on held-out BFCL endpoints (+0.047, p = 0.031, n = 6). Three failure modes that can silently invalidate zeroth-order experiments on frozen language models are documented.

String: An Agentic OS Where Every App Is a Markdown File

Jookyung Song, Nojun Kwak, Simyung Chang Agents pay context cost for every tool schema and rendered page they see on each turn, since interfaces were designed either for human skimming or for programs that carry definitions cheaply. String is an open-source runtime that treats this as an operating-systems problem: a single SFMD (String-Flavored Markdown) document declares an application's views, typed actions, navigation, and credentials, and the runtime exposes them through just two verbs, /open to see and /act to do, with tool knowledge held outside the agent's context and rendered back one view at a time. The same document serves styled HTML to browsers and raw Markdown to agents, so apps, files, shells, and legacy web pages share one grammar with no per-site integration, and privilege follows provenance so a remote page can call HTTP but never the shell. Staged disclosure matters causally — revealing a tier of detail one turn early costs up to 23 accuracy points — and on an 87-task benchmark across six models the approach matches curated-skill baselines (+1.3pp) while using 33.5% fewer tokens, with the resident interface holding at a constant 53 tokens regardless of catalog size.

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

Zongkai Liu, Hui Zhang, Liqiang Niu, Zhen Cao, Han Li, Juntao Liu et al. Agentic search environments typically hand back retrieved evidence as text and drop tool-returned images from later context, collapsing what should be visually grounded reasoning into text-only reasoning, while long trajectories accumulate tool-call, length, timeout, and budget failures that waste rollouts and destabilise training. WeAgent-Harness gives retrieved images persistent disk references so the model can inspect, process, and cite them throughout a trajectory, and adds runtime recovery; on top of it, WeAgent-MMSearch covers task synthesis and verification by a strong multimodal model, expert trajectory collection, and post-training with Failure-Aware GSPO, which rescues salvageable abnormal rollouts and discards invalid ones. A companion benchmark, VisTarget-Bench, pairs each of 150 human-verified questions with a held-out target image to tell retrieval failures apart from perception failures. Agentic post-training raises the average score by 19.22 points, letting the model beat similarly sized open-source systems and match ones roughly ten times larger.

Post-Edit Re-Verification in Simulator-Backed Engineering Agents: A Controlled Comparison of Verification-Cadence Guidance

Qingchuan Zhu, Shuyue Tong, Pengju Ren cross-listed When an agent driving an external simulator edits a design, its previously gathered evidence goes stale — but it is unclear whether agents re-run the simulation on their own or only when told to. This controlled study compares a prompt that explicitly instructs the agent to request a new simulation after any substantive modification against one with that instruction removed, holding all verification-relevant state constant and using no hard gate, with five Alibaba/Qwen models on eight synthetic valve-pressure cases in DWSIM across 120 runs per condition. Re-verification occurred in 94 of 120 runs with the cadence instruction versus 32 of 120 without, cadence violations fell from 87 to 26, and bounded final success rose from 35 to 95; qwen3.5-35b-a3b essentially never re-verified and never succeeded under either condition, which the authors read as evidence that verification cadence belongs in the interaction protocol rather than being left to the model.

CrabOS: An Operating System for Human-AI Co-inhabitation

Qi Yang, Yun Ma Real tasks often require a human and an AI agent to take turns leading, but current agent systems give each side its own work environment, so handing off state means building task-specific interfaces or manually pasting screenshots and descriptions. CrabOS implements what the authors call human-AI co-inhabitation: the work state lives as natural-language-readable text objects that both the person and the agent read and manipulate through the same auditable interface, with no bridging layer. Case studies argue this moves support for alternating-leadership tasks from application-level workarounds up to native operating-system capabilities.

Benchmarking large language model agent societies against human behavioural distributions

Raad Bin Tareaf cross-listed Populations of language-model agents are increasingly run as stand-ins for human experimental societies, which raises three doubts: whether agents behave like the humans they replace, whether findings survive cosmetic changes to the apparatus, and whether apparent social dynamics are anything more than recall of published experiments. SILICA tests all three with five environments carrying published human anchors, each paired with re-renderings that leave the rules intact and with variants whose payoffs point away from the memorized result, run over twelve open-weight models on a single consumer graphics card. Agreement with humans holds only at starting points — eight of eleven models match first-round public-goods contributions but none matches end-state contributions — and merely swapping the order in which two actions are listed cost one model 58 points of cooperation; only the single reasoning-trained model placed its ultimatum acceptance threshold where incentives require, and conventions formed from shared priors over names rather than negotiation.

Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation

Tianle Wang, Yanghe Zou, Xiang Liu, Ziyao Huang, Chenchen Fu, Weiwei Wu As reusable skill repositories grow, large language model agents must route requests to the right skill, but current routers match on task semantics alone, so two users with incompatible constraints issuing an identical request receive the same and possibly unusable skill. Personalized skill routing is reframed as profile-conditioned retrieval and paired with a counterfactual benchmark that holds the task fixed while varying the user profile so the correct skill changes. SkillFeed, a progressive retrieve-and-rerank pipeline that first establishes task-skill alignment and then reranks semantically similar but profile-conflicting candidates, reaches 75.1% top-1 accuracy on SkillFeed-Bench, 23.1 points above the pretrained routing baseline, with a 35.1-point gain on exactly those queries where the profile changes the answer.

Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration

Xiaoqing Wang, Keman Huang, Bin Liang, Hongyu Li, Xiaoyong Du, Wuqiong Pan Multi-agent systems built on large language models fail often, and existing self-reflection has every agent reflect even though a failure usually traces to one agent that led the task astray, which contaminates the well-behaved agents' memories with wrong lessons. DoCtOR instead attributes the failure automatically to a decisive error step and agent, uses counterfactual reasoning to produce a corrected version of that step, and asks only the responsible agent to reflect. Success rates improved by 22%, 26%, and 27% over the initial systems on HotPotQA, ChartQAPro, and Mind2Web, ahead of Reflexion, Retroformer, and COPPER, and under low-resource budgets reflecting only on steps after the decisive error works about as well as using the full failure trajectory.

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

Yi Wang, Haopeng Zhang, Chengxiang Huang, Rui Dai, Kaikui Liu, Piotr Koniusz et al. Loop engineering — writing a control loop that monitors a coding agent, assigns work, runs checks, and decides when to stop — is hard to evaluate, because the outcome of a single end-to-end run cannot separate bad loop guidance from a weak coding agent. LoopArena isolates the guidance role by evaluating a Controller model that, after each coding round, reads a structured run summary and tells a separate fixed Worker agent what to do or verify next, across three settings: execution-validated next-step contract selection without running the Worker, repeated control over a task slice, and the full paired task. The best Strict Success Rate observed on full tasks was 24.69%, and the cheaper slice-based setting ranked Controllers nearly identically to full runs (Spearman ρ = 0.9747) at an average 64.4% reduction in estimated inference cost.

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

Tanmay Sah, Dolly Sah, Harshul Jain, Tanya Sah Agents built on language models increasingly rewrite their own prompts, tools, middleware, and execution harnesses while running, and a mutation that helps in one state can leave effects that cannot be undone from a different state. EvoUndo represents, synthesizes, diagnoses, and independently verifies whether such self-modifications are recoverable across counterfactual states, and across 600 unseen one-shot self-evolution tasks it found 197 capability-improving mutations that fail recoverability checks. Conventional repair-by-prompting recovered 0 of those 197 failures, while giving the system exact state-address grounding lifted recovery from 0/48 to 38/48 in cases where the original recovery language sufficed, and extending the recovery language itself fixed 142/143 of the remaining stratum. On the gpt-oss-120b backbone, combining exact-address diagnostics with the richer language slightly hurt recovery (133/143), an interaction that did not reproduce on Qwen3.8-27B and therefore appears model-specific.

PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems

Hanglong Lv, Dawei Zhu, Lei Li, Bowen Ye, Huaqiu Liu, Yifan Song et al. An analysis of 16,000 real sessions found that 75.9% of user interactions with agentic systems span multiple turns, while training data and benchmarks mostly assume a single, fully specified query. PersonaForge synthesizes realistic multi-turn user-agent dialogues from a four-dimensional persona space with behavioral control calibrated against real user statistics and a reverse-construction procedure seeded by authentic queries, yielding a 6,300-record training set and PersonaForge-Bench, a hand-annotated 138-task benchmark across more than 20 professional domains. Training Qwen3.5-27B on the synthesized data raised the composite benchmark score by 4.1%, with Task Completion up 6.0% and Response Quality up 6.8%, and the trained agents completed tasks in fewer turns and fewer tool calls.

Prove2Me: An Open Collaborative Platform for Scaling Math Formalization

Shuze Chen, Kunal Marwaha, Xiaoyang Lu, Henry Yuen, Tianyi Peng Formalizing mathematics in proof assistants such as Lean 4 has been gated by the expertise and time formal proofs demand, barriers that AI coding agents have lowered enough to make internet-scale human-and-agent collaboration plausible. Prove2Me is an open platform where users launch formalization "missions" that AI agents contribute formal proofs toward, with mechanisms and a specialized harness designed so agents can build on one another's work and freely reuse existing results. Because every contribution is machine-checked for correctness, proofs can be pooled from anyone with an agent without trusting the contributor.

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

Qing Ye, Meng-Hsuan Lin While qualifying models for an internal document-extraction service, the authors caught a model passing a fidelity check — whether the extracted value matches the source — without ever opening the datasheet: a structured-output constraint had silently disabled tool use and the model answered anyway with fabricated source text, visible only in the per-tool trace. They log every tool call across an agentic benchmark of 37 hand-curated claims and build two instruments from that dispatch record: a rule-based failure-attribution classifier, and a silent-failure detector whose rules examine only which tools were called, never the extracted value. The detector raised no flag on 207 clean fidelity-passing extractions across three model families and recovered all 50 planted faults that withhold the tools its rules check, though power against runs that do call their tools and still answer wrongly is unmeasured; a second oracle using physical measurement could grade only 2 of the 37 claims, and the authors conclude the tool layer buys portability and observability rather than accuracy.

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents

Nan Li Interactive dialogue games require a model to carry state across turns, interpret feedback, and pick valid actions under shifting constraints, and diagnostics on a 2B open-weight model in the LM Playschool Challenge show its failures are often local decision errors — repeated guesses, malformed actions, violations of feedback just received — rather than only missing knowledge. The recipe has three steps: acquire broad game participation via supervised fine-tuning, repair mechanically verifiable failures in one targeted game family with turn-local preference pairs, and preserve general capabilities. In the official evaluation, public clemscore rose from 10.67 to 38.92 and closed in-domain score from 13.41 to 41.17 with static performance roughly preserved (44.14 versus 44.24), but out-of-domain clemscore stayed at 7.88, so the turn-local gains transferred mainly within the targeted family.

COVER: Identifiable Evaluation of Coalition Routing

Raghul Sugumar, Amrit Gopinath Changing the team in a multi-agent system also changes the messages it exchanges and the answer it produces, so an end-to-end accuracy gap alone does not identify a routing effect. COVER is an evaluation contract that fixes a public information boundary, a downstream stack, and a finite family of legal teams before outcomes are generated, which makes exact finite-benchmark oracle regret identifiable conditional on that stack. Across MuSiQue, HotpotQA, and a five-family ToolSandbox variant-shift validation, the instrument exposes selection headroom without manufacturing a routing win: the declared-family oracle reaches 0.768 safe-evidence completion while a prospectively frozen router gets 0.637, failing the pre-declared 0.10 regret criterion, and in fixed-stack Llama execution a 0.190 route-regret improvement corresponds to a raw-answer gain of only 0.010 with an interval crossing zero.

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin et al. Long-horizon agentic tasks force LLMs to retrieve and integrate scattered information over many turns, so preserving full interaction history makes the working context grow without bound. Existing proactive context-management methods give models only search, deletion, and summarization tools, explore all editing actions as if their impacts were equal, and assign one trajectory-level reward to every intermediate edit. ContextPilot adds planning, long-term memory, and soft context-offloading tools, and trains with an RL method that uses context and entropy variation to identify critical editing decisions for branch sampling and then estimates action-level advantages from all branched trajectories passing through that edit. On long-context question answering and deep-search benchmarks it beats existing baselines across several base models while keeping a more compact working context.

LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment

Jingjing Nie, Jiawei Guo, Krishna Meda, Haipeng Cai cross-listed Security analysis workflows are procedural — inspect artifacts, form hypotheses, run tools, revise plans — which makes them a natural target for LLM-based agents that plan, call tools, and keep state, yet the literature uses the word "agent" inconsistently and evaluates it incomparably. A systematic literature review of peer-reviewed work from 2023 to 2026 organizes the field along three axes: technical design (architecture, perception, memory, planning, action space, orchestration, self-improvement), the security tasks addressed, and assessment practice (datasets, outcome versus trajectory metrics, safety measures, baselines). The synthesis concludes the field has produced agents that can act but not agents whose authority is bounded or whose behavior is auditable, and closes with the resulting research gaps.

On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces

Ahmed Hereiz, Yingzhe Lyu, Hao Li, Bram Adams, Ahmed E. Hassan cross-listed AI coding agents are increasingly extended through plugin marketplaces where functionality arrives as a mix of natural-language instruction files, scripts, and configuration rather than source code alone, raising the question of whether such plugins are maintained or written once and abandoned. An empirical study covers 1,926 repositories hosting Claude Code plugin marketplaces, spanning 8,351 plugins and 77,773 commits. Plugin-touching commit activity grew 8.8x in the six months after the October 2025 launch, software-engineering plugins make up 61.3% of the total, feature commits occur at more than twice the rate seen in conventional open-source software (39.6% versus 17.2%), and Claude co-authors 34.9% of commits. Inside skills directories, instruction files and their implementation scripts co-evolve above chance with 78% of co-changes functionally coupled, a maintenance dependency without a traditional software analogue.

Logos: An Agent Harness on a Cross-Process Bus

Hanzhang Jia, Liheng Zeng, Hao Cheng, Yi Gao, Bo Ma Agent systems that compose capabilities at runtime are typically hosted in a single process sharing one context, which puts every component in one failure domain — a single fault suspends all of them and process death kills every session hosted there. The argument here is that neither the spatiotemporal-composability calculus nor the underlying model requires one process, since language-model inference is stateless and all cross-step state lives outside the model, and that the soundness invariant depends only on the state space; four lemmas formalize this. Logos implements the result as a ROS-like cross-process harness where each plugin is its own process and the only shared state is an append-only transcript: eighty sessions resumed with no repeated effect after kills at all four boundaries of the tool-call cycle, and a same-fault comparison showed one fault ending at a single node instead of interrupting every co-resident session.
2 more specialized papers

Safety & Alignment 26

The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

Cheolho Shin, Yoojin Han, Donghun Shin, Kunho Lee Whether a user's emotional state changes the direction of an assistant's advice is tested by presenting six commercial models with a user who is overconfident about a premature decision — quitting a stable job, expanding a business, emigrating — while holding the objective information fixed. A no-emotion multi-turn control keeps factual content and turn count constant so emotion is isolated from conversation length; 324 conversations were scored 0–100 for endorsement strength using an eight-item rubric. Distress raised endorsement from 18.6 to 31.5, a 12.9-point increase (Cohen's d = 0.51, p < .001), with the cold-versus-neutral difference not significant. Five of six models showed the effect, including the flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus did not, indicating vulnerability tracks the individual model rather than its price tier.

A Survey on Rubric-Guided Reinforcement Learning for Language Models

Zifei Shan, Fangning Shao Reinforcement learning from human feedback (RLHF) reduces response quality to a scalar reward that is neither interpretable nor able to capture multiple quality dimensions at once; rubric-guided reinforcement learning replaces that with structured natural-language evaluation criteria driving reward design, feedback, and policy optimization. The survey introduces a Bayesian framing in which constitutions are prior distributions over evaluation criteria and rubrics are conditional instantiations for a given input, then organizes the literature along a prior-to-posterior axis spanning constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and agentic and multimodal variants. Because rubrics are text artifacts, it also analyzes linguistic failure modes — granularity trade-offs, semantic drift, and linguistic reward hacking — as open problems for alignment reliability.

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Sch\"olkopf, Zhijing Jin Transcoder attribution graphs normally explain why a model assigns high probability to a next token, which says little about internal concept representations that never surface in the output. Concept-Targeted Attribution (CTA) instead builds attribution graphs with respect to a linear probe direction, producing probe-specific circuits; using Cross-Layer Transcoders, graph-level features predict probe accuracy across four concept categories with ρ = 0.91 and R² = 0.84, while local features pinpoint the sparse components behind individual classifications. Causal ablations show the two graph types are mechanistically distinct: removing probe-relevant features lowers internal concept scores while leaving generated tokens largely intact, whereas removing logit-relevant features changes the output token in 92-100% of cases with almost no effect on probe scores.

Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

Jacopo Dardini, Claudio Stanzione, Giordano Col\`o, Giuseppe Fenza Post-training quantization is usually assumed to preserve behavior, so models are certified at full precision and compressed afterward without re-evaluation — a workflow this work formalizes as a validation-deployment gap using Quantization Behavioral Equivalence Classes, proving that class membership does not imply behavioral equivalence. A three-stage adversarial fine-tuning procedure embeds payloads that stay dormant under source-precision checks but activate under INT8 or 4-bit compression, demonstrated on multilingual encoder-decoder translation and political stance classification rather than just decoder-only models. Backdoored translation models go from zero measured friend-foe corruption at repaired FP16 to up to 85.02% inversion after quantization, and cross-quantizer analysis shows persistence depends on the specific quantization scheme and architecture rather than nominal bit-width, arguing that the deployed configuration itself must be behaviorally certified.

Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator

Varun Singh, Anuj Doshi, Makesh Narsimhan Sreedhar, Shaona Ghosh, Katherine Luna Deployed AI applications increasingly need moderation over images, documents, and screenshots under policies that differ by domain, but existing guardrails cover only parts of that space or cost too much to run. Nemotron 3.5 Content Safety Moderator is a 4B-parameter vision-language moderator that classifies user prompts, images, and assistant responses across 12 languages, returning fast labels by default and optionally producing reasoning traces that apply supplied custom policies and name violated categories. The release includes a multimodal, multilingual guard-training dataset spanning human-labeled real-image moderation, benign vision-language and document tasks, synthetic rare-risk and jailbreak cases, and custom-policy examples; evaluations across multimodal safety, text moderation, multilingual robustness, policy following, false positives, and latency show it stays broadly competitive with specialized guard models while adding image- and policy-conditioned coverage.

LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails

Ziyang Chen, Xing Wu, Songlin Hu Safety guardrail models that screen large language model inputs and outputs are trained almost entirely on short text, leaving their behavior on long contexts untested. LongGuard frames the problem as Safety Needle-in-a-Haystack (SafetyNIAH) over context lengths from 0.25k to 32k tokens and finds that unsafe-content recall across 15 mainstream guardrails drops by more than 50% on average, with a paired Benign-Fill versus Needle-Repeat design attributing the failure to dilution of the unsafe passage rather than absolute length. A three-layer attention-logit-behavior analysis traces the mechanism to diluted attention mass on the unsafe span and a compressed unsafe-over-safe logit margin, concentrated in a sparse set of guard-specialized retrieval heads. Two training-free fixes, Chunked Detection and Attention-Head Sharpening, combined with length-based Context-Aware Hyperparameter Routing, raise the six-guardrail average by 22% and 13% respectively across five benchmarks.

Semantic Watermarking with Order-Robust Detection over Sub-sentence Units

Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum cross-listed Semantic watermarks bind a mark to sentence meaning rather than token choice, but the detector only ever sees attacker-supplied text, which can be reworded, reordered, or resegmented — all of which displace the embeddings the detector tests. The proposed embedding displacement attack (EDA) combines all three edits under a single objective that maximizes displacement, using only a public paraphraser and surrogate encoder, and at a 5% false-positive rate with 90% content preservation it strips the mark from 32.6% to 47.9% of documents across four schemes, the strongest of the attacks tested. The authors respond with k-SwordStamp, which detects over sub-sentence units in an order-robust way; the best no-box attack against it succeeds 10.8% of the time, and even an attacker with the provider's detector and secret key reaches 39.7% versus 65.5% against k-SemStamp.

Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots

Xujun Che, Depeng Xu, Shuhan Yuan cross-listed Memorization in large language models is measured through many incompatible definitions, and differential privacy (DP) is commonly assumed to bound all of them simultaneously. The authors pin down exact DP constants for the two that carry practical weight, counterfactual memorization and adaptive extraction, and show neither controls the other: under f-DP, extraction with a list budget is capped relative to an oblivious guessing baseline and the bound is tight on a dense set of baselines, while counterfactual memorization of any bounded score is capped at tanh(epsilon/2) under pure DP, with a closed-form staircase constant replacing the naive k-times-epsilon bound for k duplicated copies. Because the two measures separate inside the local score class practitioners actually use, one mechanism can be memorized yet unextractable and another fully extractable yet invisible to loss-based scoring. On billion-parameter models a reserved-trigger release is recovered verbatim from a single prompt while the deployed audits certify the model clean.

ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools

Yuqi Jia, Ruiqi Wang, Patrick Li, Yuepeng Hu, Peinian Li, Neil Gong cross-listed Stealing an LLM agent's runtime context, meaning the user prompt, execution trajectory, and tool list, requires three things to line up: the agent picks the malicious tool, the agent passes its context as tool arguments, and the tool implementation ships those arguments to an attacker endpoint. Prior work has studied the first and third conditions while leaving the second largely unexplored, and ContextLeak closes that gap by having an attack LLM generate the malicious tool's name and description, fine-tuned with reinforcement learning on shadow users with diverse simulated agent contexts under reward functions designed specifically for the exfiltration objective. The attack remains highly effective even when shadow contexts differ substantially from the victim's, and outperforms existing malicious-tool attacks adapted to this setting.

EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

Ruijie Jian, Benlei Cui, Ting Ma, Haidong Ding, Kangwei Liu, Ziwen Xu et al. Content moderation is evaluated on static benchmarks, but real platform users iteratively rewrite posts in response to being blocked, creating a gap between offline scores and deployed effectiveness. EvoHarmBench closes that gap with a dynamic adversarial loop that evolves evasion strategies at the level of semantic clusters while jointly optimizing for evasion success and human readability, covering 229 semantic sub-clusters across five violation categories drawn from 5,002 real adversarial samples. Against widely deployed LLM-based moderators, twelve optimization rounds reach an 80.3% attack success rate under readability constraints, including on leading commercial systems; the data, framework, and code are to be released.

OpenStamp: A Watermark for Open-Source Language Models

Miroojin Bakshi, Saksham Rastogi, Danish Pruthi Watermarks that work by nudging token sampling probabilities can simply be switched off by anyone running an open-weight model locally, since users control inference. OpenStamp instead encodes the watermarking logic into the weights themselves, modifying only the final projection (unembedding) layer so the signal is emitted regardless of how the model is served. Across two models it reports stronger detection than prior open-source watermarks with minimal degradation in model capability, along with greater robustness to paraphrasing and to attempts to scrub the mark through post-hoc fine-tuning; code and watermarked versions of four popular open models are released.

AI Alignment through a Game-theoretic Lens: A Survey

Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue et al. Alignment techniques tuned for helpfulness, harmlessness, and controllability struggle with real preferences that are context-dependent, non-transitive, and shaped by several interacting parties over time. The survey reorganizes recent alignment research around game-theoretic primitives and groups the literature under three challenges: preference diversity, alignment priority, and temporal dynamics. Its stated contribution is separating where game-theoretic analysis genuinely buys something from where the framing is only loosely applied, alongside the open problems that remain for building robust, adaptive, and verifiable systems.

Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense

Disen Liao, Yihan Wang, Freda Shi, Yaoliang Yu A user can decompose a forbidden objective into innocuous subquestions, ask them across independent sessions, and recompose the answers afterward. The authors formalize this as compositional safety risk and prove a conditional risk-transfer bound showing that the gap between deployed and reference composed risk is controlled by the model's excess loss on the allowed subqueries — so lower language-modeling loss mechanically raises this exposure. Synthetic withholding experiments plus a 600-intent evaluation across the Qwen3 and Gemma3 families show larger models delivering greater harmful-capability uplift under a fixed decomposition-and-recomposition pipeline, while their 22M-parameter IntentAlign-MiniLM retriever beats much larger embedding models at recovering the held-out underlying intent.

Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification

Cameron Wilding, Mina Shaker, Fatemeh Ganji cross-listed A provider can alter a deployed model after approval in ways that leave routine outputs looking normal, which leaves auditors stuck when the weights are proprietary. The framework searches for probe inputs, constructed in the spirit of adversarial examples, that amplify logit drift between the approved model and the live deployment, then wraps the comparison in a Groth16 zk-SNARK so the audit discloses nothing about the model; probe families range from black-box token probes through gray-box embedding probes to stress probes needing extra interface access. Evaluated across architectures, tampering scenarios, and GPU platforms, the black-box token probes give the strongest mean sensitivity despite the weakest access requirements, and scaling from 1 to 50 probes raises proving time only from 1.02 to 1.78 seconds with proof size constant.

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?

Zi Liang, Xiaoyu Xu, Yanyun Wang, Minxin Du, Qingqing Ye, Haibo Hu cross-listed Prompt injection defenses for language-model agents handle known attacks well but struggle against variants and novel techniques, and face a trilemma between runtime efficiency, contextual precision, and adaptability. CAITLYN is an agent-agnostic defense middleware split into two systems: System I gives immediate protection via a two-tier library of cheap rule-based detection scripts and optimized LLM-based inference, while System II watches for anomalous signals and tries to synthesize entirely new defenses. On standard benchmarks it matches state-of-the-art detection at lower token overhead than LLM-as-a-judge baselines, and on Emerging, a new delivery-aware benchmark of novel injection techniques where static baselines and System I alone stay vulnerable, System II autonomously synthesizes verified defenses that substantially cut attack success rates across three agent environments.

Beyond Global Scalars: Synergizing Token-Level Statistics and Deep Semantics for Adversarial AIGC Text Detection

Peiming Li, Yifan Wang, Zhiyuan Hu, Shiyu Li, Zheng Wei, Yang Tang Machine-generated text detectors split into training-free methods reading global statistical scalars like perplexity and training-based methods reading semantic hidden states, and both break under adversarial pressure: scalars are lossy compressions that hide local probabilistic burstiness in human-machine interleaved text, while semantic models overfit to particular model fingerprints and can be spoofed. The authors release MOSAIC, an adversarial benchmark of 16,000 samples spanning a full-granularity attack spectrum, and propose NeuroStat, which extracts uncompressed token-level logits and deep hidden states from a single causal language model backbone and fuses them via Macro-State Residual Modulation, calibrating local convolutional features with global uncertainty indicators under orthogonality and contrastive losses. NeuroStat holds up on MOSAIC where state-of-the-art detectors degrade severely, with code and benchmark released.

Speculative Probing: LLM Monitoring at Speculative-Decoding Cost

Collin Zhang, Tingwei Zhang, Vitaly Shmatikov Classifying model behaviour during inference — for safety filtering, behavioural analysis, or monitoring — currently means choosing between fast hidden-state probes that see only one vector and cannot model interactions across positions, and accurate but expensive options like dedicated guard models or pooling computation over every token. The trick here is to repurpose the speculative-decoding module already present in recent models: appending a trained soft prompt to the end of the target sequence turns that draft module into a sequence classifier, and because the KV cache is already resident in GPU memory during speculative decoding, the classification costs almost nothing extra. Across four tasks and four models (Qwen3.5-4B, 9B, 27B, MiniCPM4.1-8B), the small probes beat zero-shot GPT-5.4-mini consistently and match or exceed specialized 8B safety classifiers such as Qwen3Guard-Gen-8B and Llama-Guard-3-8B on multilingual prompt safety without running a second full model.

Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation

Dipto Das, Arpita Kundu, Nusrat Jahan Mim, Shion Guha, Syed Ishtiaque Ahmed Discussions of AI alignment around religion have largely assumed secular, Western, or Abrahamic framings, leaving other traditions unexamined. Fifteen semi-structured interviews with Bangladeshi Hindu participants trace how they use generative AI for scriptural inquiry, devotional visualization, and religious storytelling, and how they interpret synthetic sacred imagery and explanations. Participants found the systems genuinely accessible but reported theological flattening, cultural misrepresentation, devotional manipulation, and the simulation of sacred presence and authority, prompting the authors to argue for interpretive alignment: systems that disclose their limits, preserve plurality, and avoid simulating religious authority or sycophantic personalization.

REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features

Kai-Xuan Ding, Hao-Xiang Xu, Ji-Hua Peng, Zi-Qi Chen, Jiaqi Wang, Zhen-Hua Ling Steering with sparse autoencoders (SAEs) adjusts language model behavior at inference time without retraining, but harmful requests buried in elaborate wrappers slip past it. GUISE (Generalized Undercover Instruction Safety Evaluation) is a new dataset of such wrapped harmful prompts, on which existing single-direction steering fails to produce reliable refusals — apparently because boosting refusal features leaves the harmful continuation path active. REINS intervenes in both directions within the same SAE feature space, suppressing harmful-continuation features while enhancing refusal features, and substantially reduces harmful responses and improves safe refusals while largely preserving general capabilities, where baselines either intervene too weakly or appear safe only by collapsing.

AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning

Wonjun Lee, Jaehyuk Jang, Kangwook Ko, Hee-Seon Kim, Changick Kim cross-listed Multimodal large language models (MLLMs) fine-tuned on people-specific data memorize identity facts, and existing unlearning methods assume access to retain images or ground-truth answers at deletion time, which is often unavailable when someone requests removal. Probing fine-tuned hidden states shows identity questions and visual-perception questions occupy distinct regions and are organized differently — identity questions cluster by person, perception questions by question type — implying identity knowledge can be suppressed without damaging perception. AIM exploits this in two stages: anchor an identity-forgetting target using a universal visual prompt, then match the vision encoder to that target under a Fisher-based constraint, achieving competitive forgetting while preserving non-deleted identities, prior knowledge, and visual perception on the same images.

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed cross-listed Stacking guardrails around a large language model assumes the layers compound, but an ensemble only compounds when its members fail on different inputs — a condition the security literature recommends and never measures. Two instruments make the stack analyzable: an Adversary Access-Tier Model grading attackers from system-only access (A0) to training-data influence (A4), and a five-class cost model of inference-time overhead that also tiers the defender, since two classes need weight access or activation reads; from these, coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence. Running one adaptive adversary against a seven-layer stack, failure correlation was positive in all fifteen measurable pairs (φ from 0.30 to 0.75) and the joint residual exceeded the multiplicative prediction by up to 0.172, while the full stack refused four in five benign prompts yet remained statistically indistinguishable from its single strongest layer. Stratifying on behavior difficulty dissolves most of the association, indicating common-cause dependence that arises architecturally from members wrapping the same model, so a wider candidate pool cannot fix it and stacks must be measured end to end.

GRACE:Gradient-guided Coreset Selection for LLM Unlearning

Praveen Bushipaka, Andrea D'Angelo, Lucia Passaro, Tommaso Cucinotta Machine unlearning for language models normally assumes someone hands you a curated forget set and retain set, but a real removal request may consist of only a handful of examples of the unwanted behavior, leaving the actual training subsets to be inferred from a large mixed corpus. GRACE computes a gradient "forget direction" from those seed examples, then uses non-negative orthogonal matching pursuit to pick a compact forget coreset whose gradients approximate that direction, and selects retain examples by projecting out the forget direction and running clustered orthogonal matching pursuit in the leftover gradient space. Across two target domains, two model families, and four unlearning algorithms, the method preserves more general model utility at comparable forget quality, with the most consistent gains over earlier gradient-based selection methods.

CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents

Jaewon Jung, Haizhong Zheng, Hongsun Jang, Jaeyong Song, Beidi Chen, Jinho Lee cross-listed Retrieval-augmented generation (RAG) systems that draw on public or user-editable sources can be poisoned by injected documents that steer answers toward an attacker's target, but existing attacks embed the target query verbatim in the poison, leaving lexical and embedding-space fingerprints that simple filters catch. CamoDocs drops query inclusion entirely, instead chunking together synthesized benign and adversarial drafts, swapping selected tokens in the benign chunks for dispersion tokens that scatter the poisoned documents' embeddings, and applying a coherence filter so the text still reads normally. Tested against seven RAG defenses, three open-weight models, and three benchmarks, it achieves high attack success without query-overlap artifacts, and reaches average attack success rates of 61.80% on GPT-5.4-mini and 55.09% on Claude-Haiku-4.5; erasure-heavy clustering defenses such as TrustRAG blunt it only at a large cost in utility on retrieval-dependent benchmarks like NeoQA.

LongPIBench: A Long-Context Benchmark for Prompt Injection

Yupei Liu, Yuqi Jia, Neil Zhenqiang Gong, Jinyuan Jia cross-listed Prompt injection benchmarks have concentrated on short inputs, which the authors argue makes current defenses look considerably stronger than they are. LongPIBench covers four realistic deployment scenarios — paper peer review, resume screening, code review, and email summarization — each with a synthetic and a real-world dataset whose contexts range from thousands to tens of thousands of tokens. Evaluation shows even simple heuristic injection attacks achieve high success rates and frequently bypass state-of-the-art defenses once the surrounding context is long.

When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI

Sihan Jia, Oliver Lemon Voice-controlled embodied AI inherits whatever mistakes its automatic speech recognition (ASR) front end makes, and the safety consequences of those mistakes had not been mapped. Simulated ASR errors are injected into the existing embodied safety benchmarks SafeAgentBench and POEX to see which error types matter. Some corruptions preserve semantic structure while increasing harmful ambiguity, and others actively weaken model refusal behavior so unsafe plans get generated and executed; automatic error correction helps in some cases but is not reliably protective.
1 more specialized paper

Theory 26

Optimal Transport for Network Comparison: A Review with Machine Learning Applications

James Hyun, Fran\c{c}ois G. Meyer cross-listed Optimal transport offers a way to compare graphs that yields not just a dissimilarity score but a transport plan describing how one network morphs into another. The review covers three distances applied to undirected, unweighted graphs — Wasserstein, Gromov-Wasserstein, and Bures-Wasserstein — showing the closed form of the one-dimensional Wasserstein distance over node feature distributions, demonstrating how transport plans identify which specific nodes drive the distance after a graph perturbation, and deriving Laplacian-spectrum bounds that avoid full spectral decomposition for the Bures-Wasserstein case. The distances are then tested on synthetic networks for clustering and on a real-world time series network for anomaly detection.

When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging

Shangge Liu, Yuehan Yin, Yinghuan Shi, Lei Wang, Wenbin Li Continual learning fights catastrophic forgetting and model merging fights weight-disentanglement error, and these are argued to be the same underlying problem: an update helpful for one task shifts outputs on another. Formalizing this as task interference reducible to a layer-wise Frobenius inner product between the weight update and a task Jacobian, the analysis derives an upper bound isolating the spectral norm of the update as the factor an optimizer can control, and identifies the Muon optimizer as regulating that factor by construction. Swapping AdamW for Muon improves accuracy by up to 5.02 points on the eight-task model-merging benchmark across three CLIP backbones, with uniformly positive gains across ten class-incremental protocols, three task-incremental protocols, and the 11-task MTIL benchmark.

Towards a mathematical theory of superposition

Michael I. Ivanitskiy, John Jasper, Emily J. King, Dustin G. Mixon cross-listed Superposition — neural networks packing more features than dimensions — is given a formal treatment using frame theory and compressed sensing, modeling a sparse binary feature vector encoded through an overcomplete dictionary and recovered by applying a rectified linear unit to the Gram matrix product with a bias. Several recovery theorems follow: for random feature supports, high-probability support recovery holds for nearly tight, low-coherence dictionaries up to expected sparsity of order d/log n, while for worst-case supports there is a sharp, computable criterion for which sparsity levels permit recovery. Applying that criterion to Gaussian random matrices and equiangular tight frames yields an exact coherence-based recovery threshold for real equiangular tight frames with n > d+1, proved via a new characterization of the sign distribution in the Gram matrix.

More Data Cannot Break a Symmetry: Identifiability by Design

Jing Xu, Christopher Kanan Unsupervised representational alignment tries to recover a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus set caps what is identifiable before any data is collected — and the obvious diagnostic, the cheapest non-identity relabelling, misranks published designs because dense sampling creates near-duplicates that are almost free to swap. Working in color space, where candidate geometries are available in closed form, the authors show the failure is structural: a symmetric design resists alignment even with 64 times the restart budget, while an asymmetric set of the same size succeeds every time, and discriminating representational models is essentially uncorrelated with recovering a correspondence (r = -0.02 over 3,000 subsets). Selecting nine colors by this diagnostic alone, without consulting any learned representation, cut catastrophic alignment failures from 75% to 2% across 93 model representations with models, layers, set size, and solver held fixed, at a cost of one function call before data collection.

Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data

Zhenyu Tao, Wei Xu, Xiaohu You, Petar Popovski, Osvaldo Simeone Digital twins and learned world models are widely used to synthesize training data for engineering systems where real data is scarce, but the simulation-to-reality gap means augmentation can hurt rather than help, so practitioners need a way to decide — using as little real test data as possible — whether adding a candidate synthetic dataset actually improves population-level performance for a fixed learning algorithm. The authors formulate this as sequential hypothesis testing in two forms, a direct test on the mean loss difference and a symmetry-based test on paired loss differences that buys faster evidence accumulation with a stronger null assumption, and for the latter introduce the adaptive e-process sign-flip test (aeSFT), which adapts both the number of Monte Carlo sign-flip rounds and the real test data consumed while retaining anytime-valid Type-I error control without pre-specifying a test-set size. On synthetic-data classification, digital-twin-aided wireless packet scheduling, and radio-map prediction, it detects useful synthetic data with far fewer real samples than mean-based sequential testing while matching the power of fixed-sample sign-flip testing and the paired t-test.

When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?

Yansen Han, Hongxin Sun, Tao Lin Alignment methods for flow-based generative models increasingly reuse conditional flow matching (CFM) losses as if they were endpoint negative log-likelihoods and their old/new differences as log-likelihood ratios; this work asks when that substitution is actually valid. For linear Gaussian paths the authors exactly decompose the endpoint negative log-likelihood into entropy, a weighted CFM objective, an interior velocity-score residual, and a boundary residual, so CFM-only estimates are exact only when those residuals cancel. At the off-policy population optimum ordinary CFM is not generally a pointwise likelihood estimator, though the weighting w(t) = (1-t)/t removes the interior residual — a result that does not extend to training or on-policy alignment, where log-ratios can stay biased even for identical endpoint laws. Experiments across dimensions, distributions, and geometries support the analysis and clarify why inexact ratios can still be useful as controlled surrogates.

Landau theory of quenched criticality in linear in-context learning

Daesik Kim, Sumin Choi, Hyojae Jeon, Jung Hoon Han cross-listed In linear models of in-context learning — where a pretrained model infers a task from prompt examples without weight updates — the prediction error blows up in a double-descent singularity when pretraining sample count approaches parameter count. Treating this interpolation point as a critical phenomenon in a quenched disordered system, the analysis traces the singular error to connected sample-to-sample fluctuations of the learned parameters and constructs a Landau potential by integrating the cavity equation for the renormalized ridge parameter. The renormalized ridge acts as order parameter, the bare ridge as its conjugate field, and normalized sample complexity as temperature, placing the double-descent singularity at critical temperature τ = 1 with a generically cubic potential and critical exponents (1, 2, 1); numerical solutions of the original learning problem match the theory quantitatively, and a pseudogap-like regime with suppressed order parameter shows up at large context.

The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension

Yuhe Sui, Jianing Zhang The question here is what geometry governs how well softmax attention can be approximated by a low-rank matrix, measured as the maximum-row-ℓ1 approximation rank that preserves every bounded output. Two sharp worst-case laws separate the effect of where keys and queries live: spherical self-attention has rank Θ(min{n, (1+β)^((d-1)/2)}), while full-ball geometry adds a radial degree of freedom and yields Θ(β^(d/2)) at large temperature. Within a fixed head, row-wise softmax cancels row-scalar logit directions, leaving a visible query–key interaction dimension r that gives a minimax-sharp per-instance exponent of r/2; a calibration set of 84 BERT-base heads shows modest effective-dimension reductions across many head and temperature settings, so support geometry sets worst-case temperature scaling while softmax-visible interaction geometry controls per-head complexity.

Performative Privacy: When Differential Privacy Maximizes Utility

Uddalak Mukherjee, Edwige Cyffers, Yann Chevaleyre Privacy protection is often justified by the argument that it sustains user trust and therefore participation, improving utility over time, but that claim had not been formalized. Performative privacy joins differential privacy to performative learning in a model where agents repeatedly contribute data for mean estimation and drop out when their data leaks, so the privacy budget trades estimation noise against future participation. Theoretical analysis of the resulting dynamics plus numerical experiments show that a finite privacy budget can outperform non-private estimation in the long run once the feedback loop between leakage and attrition is strong enough.

SinkSLOT: Sinkhorn via Sparse Lifted Optimal Transport

Ian Hsieh, Soumya Snigdha Kundu, Tom Vercauteren, Reuben Dorent Entropic optimal transport makes optimal transport computationally tractable, but Sinkhorn-Knopp costs O(N²) per iteration for measures with N points and its independent-coupling reference measure assigns mass to expensive transport edges at moderate regularization. SinkSLOT replaces that reference with an expected sliced lifted transport plan, which sparsifies the Gibbs kernel and supplies a non-independent prior coupling at the same time. The authors prove convergence, show that each sparse iteration costs O(LN) with L slices, and show the objective is already a divergence requiring no debiasing, with synthetic benchmarks reporting substantial speedups over dense and sparse state-of-the-art methods.

A Formal Limitation on Learning Human Language From Textual Corpora

Emily Cheng, Ryan Cotterell Whether a listener can recover a speaker's intended meaning from utterance form alone is answered information-theoretically, for any featurizer of text including the hidden states of contemporary LLMs. Modeling language use as a joint distribution over meanings, contexts, and utterances yields upper bounds on the probability that any decoder recovers the intended meaning from a representation of the utterance, governed by the uncertainty form leaves about meaning, which splits into an irreducible component and one that only extralinguistic context can resolve. Because these quantities are intrinsic to language rather than to any model, no representation trained on any amount of text or supervision can exceed the bounds, which hold for discrete and continuous meaning spaces alike; experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference support the theory.

Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy

Lorenzo Rizzi, Arie Wortsman Zurich, Bruno Loureiro cross-listed Classical analyses of kernel ridge regression usually assume isotropic input data, but real inputs have covariance concentrated along a few directions. Working with polynomial inner-product kernels on anisotropic Gaussian data whose covariance spectrum decays as a power law with exponent alpha, the authors derive asymptotically sharp expressions for the kernel spectrum and generalization error in the regime where the number of samples n scales as d raised to a power kappa. Weak anisotropy (alpha below 1) keeps the familiar variance peaks at integer sample complexities but progressively damps them, while strong anisotropy (alpha above 1) fixes the effective dimension so the variance stops depending on sample size at all, plateauing under ridgeless interpolation or decaying at an explicit rate with a fixed ridge. The bias undergoes a separate sharp transition set by how fast the target decays, switching between abrupt learning and classical source-and-capacity power-law rates, with single-index targets used to show how alignment with principal directions drives the effect.
14 more specialized papers

Other 23

Node-wise Feature Encoding for Neural Performance Prediction

Matthew Grenier, William Hammer, Andrew Heuer, Nikhil Krishna, Yi Wang, Ramtin Zand Neural architecture search for resource-constrained edge devices needs accurate predictions of a candidate network's latency and energy, yet existing graph-neural-network and transformer predictors largely ignore per-node computational cost. FeatureFormer adds explicit node-wise encodings of floating-point operation counts, parameter counts, and memory proxies inside a gated graph attention architecture, and the authors release NNEQ, a large-scale energy-consumption dataset that lets latency and energy prediction be evaluated in one place. The predictor reaches state-of-the-art accuracy on both metrics including out-of-domain settings, and the node-wise encoding also improves existing predictors at negligible overhead when added to them.

SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning

Hao Wang, Siyu Zhang, Wei Ma Tabular foundation models that predict via in-context learning now rival per-task model fitting, but the leading systems use attention at every stage of the pipeline, which is expensive. SOMTab splits the problem, mapping unordered table tokens into stable latent slots and mixing them with Mamba state-space layers to build row and column representations, while keeping attention only for the final query-conditioned retrieval from labeled context examples; a synthetic prior called DCH-TailMix diversifies training structures using degree-corrected graph heterogeneity and heavy-tailed regimes. Across tabular benchmarks it approaches strong Transformer-based tabular foundation models while running faster and using less GPU memory.

Biologically Inspired Mechanisms for Facilitating Grokking in Multilayer Perceptrons

Florin Leon Grokking is the delayed jump from memorization to generalization, and the question here is whether biologically inspired regulatory mechanisms rarely used in artificial networks can actively bring that transition forward. A multilayer perceptron is augmented with input gating, structural plasticity, gain modulation, threshold modulation, homeostasis, lateral inhibition, and activation decorrelation, then ablated systematically on sparse parity and noisy XOR classification. Homeostasis gives the strongest and most consistent benefit, with structural sparsification second, while the remaining mechanisms have smaller or inconsistent effects — supporting a general principle that explicitly regulating neuron utilization and effective connectivity helps generalizable internal computation emerge.

Generalized Context in Cross Attention for Transfer Learning of Disjoint Tabular Data

Kazi F. Akhter, Ibna Kowsar, Manar D. Samad Transfer learning between tabular datasets normally assumes source and target share features, which fails for genuinely disjoint tables with different feature types and semantics. CATTLE (Cross-domain Attention Transfer Learning) transfers what the authors call generalized context, held in a transformer's key, value, and query projection weights rather than in activations, letting source key projections interact with target query projections in a data-agnostic way. Across ten disjoint source-target dataset pairs it outperformed nine machine learning, deep learning, and pretrained-model baselines, achieving the best average rank of 2.9 and a 3.7% average AUROC gain.

Blog: Survey of Optimizers

Ruoran Xu Neural network optimization has moved past being a sequence of Adam variants, expanding from per-coordinate to matrix- and layer-aware updates, from fixed horizons to time-varying policies, and from clean math to state representations that must survive sharding and low precision. This survey organizes recent methods along four largely independent axes — temporal estimation, update geometry, horizon management, and representation and systems — connecting Muon's spectral normalization, the matrix statistics of Shampoo and SOAP, memory-efficient and quantized-state optimizers, schedule-free training, and small-batch corrections. The headline conclusion is deliberately non-triumphal: matrix-aware methods are a real advance but there is no context-independent replacement for AdamW, since rankings flip with scale, data-to-parameter ratio, batch size, tuning budget, and whether success is measured in tokens, FLOPs, wall-clock time, or memory — motivating a compositional view of design and a stricter evaluation protocol.
18 more specialized papers

Vision 14

Quanta Perception as Probabilistic Events

Varun Sundar, Pavan Thodima, Sacha Jungerman, Mohit Gupta cross-listed Cameras that integrate photons over fixed exposures force trade-offs between sensitivity, dynamic range, and temporal resolution that break down in near-darkness or at high speed, while single-photon quanta sensors produce streams far beyond real-time compute budgets. The proposed primitive, probabilistic events, computes a posterior over the time since the last intensity change and represents photon streams as recursive belief states, yielding motion-adaptive scene flux, activity maps, and entropy-based perceptual uncertainty instead of reconstructed frames. The method processes more than 50,000 quanta frames per second on commodity GPUs, up to four orders of magnitude faster than state-of-the-art quanta reconstruction baselines, and supports pose estimation of a running person at roughly 0.05 lux without retraining existing vision models.

From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

Rit Gangopadhyay, Alex Wong cross-listed Vision foundation models trained on perspective photographs break down on wide field-of-view fisheye images because radial lens distortion shifts the input distribution away from anything they saw in training. DEX (Distortion Extenders) adds a small set of learnable parameters that jointly model the fisheye distortion coefficients and the latent-space gap between fisheye and perspective images, trained with a self-supervised alignment loss that warps fisheye embeddings to look perspective-like. The approach is agnostic to architecture and task, improving both monocular depth estimation and open-vocabulary segmentation over baselines for convolutional and Transformer backbones on indoor and outdoor fisheye datasets, and its activations can be decoded into distortion coefficients for camera calibration.

What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection

Parishruthi Ganesh cross-listed Detailed skeletal pose should in principle carry more information about a violent interaction than coarse bounding boxes, so the authors hold the tracker, temporal head, supervision, and evaluation fixed and compare five interaction representations for early violence detection. No pose-based representation beats coarse geometry, and once frozen visual encoders are added on the larger XD-Violence split, whole-frame context matches or exceeds person-crop appearance — cropping to the interacting people buys nothing. Scoring anomalous videos using only frames before the annotated onset, with sequence length controlled for, still retains 39–91% of above-chance separation on both UCF-Crime and XD-Violence, traced to provenance artifacts such as editorial title cards and platform watermarks absent from the surveillance footage supplying normal clips, meaning video-level AUC mixes genuine event evidence with source cues.

VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians

Ruijie Su, Lingxiao Yang, Xiaohua Xie, Jianhuang Lai cross-listed Physics-driven 3D Gaussian pipelines mostly animate solid objects and handle only single-phase collisions, leaving out interactions between materials in different states. VersaGauss unifies generation, simulation, and rendering, taking a handful of images and producing a dynamic multi-object 3D scene, with a particle pruning algorithm to shape the Gaussian kernel distribution and a Coupled Multiphase Point Method (CMPM) to model interactions across phases; harmonic interpolation inside CMPM plus a Gaussian evolution strategy handle fluid rendering. Experiments show interactions simulated among fluid, rubber, sand, snow, and other materials in a single scene, with code released publicly.

Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations

Naren Akash, Neeraja Ramanan cross-listed Reading a CT scan involves comparing structures across the body, judging distances between organs, and knowing where each organ belongs, yet medical vision encoders are graded on diagnostic accuracy or inside assembled multimodal systems where failures cannot be attributed. SPAR-Bench supplies eight probes over multi-organ abdominal CT that separate coordinate localization, relational reasoning, and spatial queries, applied to five architectural configurations and three medical foundation models both frozen and finetuned. Probes requiring a comparison within a single slice stay at chance regardless of pretraining scale, finetuning, or architecture, and probes that look solved in-domain collapse to chance under zero-shot transfer, suggesting the encoders recall where organs usually sit rather than computing over the specific image. Reading the same frozen features with a pooled head instead of all tokens lifts relational recovery from 0.7% to 67.8%, so pooled probing understates what representations contain, and four open-weight multimodal LLMs answer at chance the questions the encoders handle well.

A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation

Tadej Tomani\v{c}, Alice Baudhuin, Jan Soto\v{s}ek, Jure Brence, Pan\v{c}e Panov, Nikola Simidjievski et al. cross-listed Change detection in Earth observation is held back by inconsistent evaluation protocols and a focus on accuracy that ignores compute cost. The benchmark trains ten architectures, from convolutional networks to vision transformers, on ten heterogeneous datasets under identical protocols, comparing training from scratch against pretrained weights and reporting parameter counts and inference latency alongside predictive scores. Well-optimized classical designs such as Siamese U-Nets frequently beat more complex recent models once efficiency is factored in, and pretraining gives a consistent gain at no inference cost; splits, scripts, logs, and checkpoints are released under FAIR principles.

Physics-Guided Flow Matching for CT Image Reconstruction

Davide Evangelista Diffusion priors deliver state-of-the-art CT reconstruction but depend on stochastic sampling, long inference trajectories, and carefully tuned noise schedules. A rectified Flow Matching model is trained instead on 256x256 chest images from the Mayo Clinic Low-Dose CT dataset with a two-stage schedule — heavy anatomically informed augmentation first, then fine-tuning with reduced or no augmentation to restore structural fidelity — and used as the prior for Plug-and-Play Flow, FlowDPS, Flower, and Flow-Priors. Across several CT inverse problems these outperform the diffusion baselines DDRM, DPS, and DiffPIR on PSNR, SSIM, and perceptual quality while needing fewer sampling steps, and the trained model and code are released.

Video Generative Models as Geometry Learner

Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng cross-listed Generative geometry estimation currently adapts pretrained image diffusion models, either training depth and normal predictors separately — forfeiting the correlation between the two targets — or jointly fine-tuning modified backbones, which needs a lot of labeled data. GeoNeXt instead repurposes a pretrained video generative model and casts geometry estimation as next-frames prediction, inheriting temporal structure and richer priors while adapting them to jointly model image and geometry targets in both directions. On zero-shot monocular depth and surface normal estimation across diverse datasets it beats prior task-specific and unified generative methods, and rivals discriminative state-of-the-art systems trained on over 100x more data, leading on several benchmarks.
6 more specialized papers

Multimodal 13

SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction

Nilay Yilmaz, Naga Sai Abhiram Kusumba, Stella Wenxing Liu, Yezhou Yang Relational reasoning spans several distinct abilities — analogical, structural, cause-effect — and SciReC tests them in multimodal large language models (MLLMs) through a model-adaptive multimodal academic dialogue benchmark. It is paired with DMRA, a deficit-based diagnostic framework that separates how much visual understanding, knowledge, and memory recall each contribute to a failure. Claude 4.6 leads with a 73% overall relational score, ahead of GPT 5.4 at 68%; open-source models score lowest on spatial relations while proprietary models struggle more with hierarchical and sequential ones, and performance is worst in Astronomy and best in Psychology. DMRA attributes the bulk of errors to relational reasoning itself, with memory limitations second.

Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls

Manuel Cherep, Pattie Maes, Nikhil Singh Interpretability work typically fixes a stimulus set and asks either what internal structure represents or how outputs vary, neither of which recovers a model's prior — the distribution over inputs it implicitly expects — since image space is far too large for any fixed set to cover. The proposed method steers a generative model to produce stimuli along interpretable control axes and runs Gibbs sampling over that space with the multimodal LLM under study acting as the judge, drawing samples directly from its perceptual prior. Applied to targets such as trustworthiness in faces and cheapness in art images, it recovers both canonical biases and surprising priors that direct prompting leaves invisible.

CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning

Runze Liu, Naibin Gu, Mingxu Ai, Yuqing Li, Peng Fu, Zheng Lin et al. Adding a fresh set of LoRA experts for every new task lets a multimodal model learn continually without forgetting, but the parameter cost grows without bound. A singular value decomposition of task-specific LoRA updates shows their input- and output-side direction subspaces overlap heavily, with task-specific adaptation mostly expressible as lightweight coordinates over shared bases. CoRe-MoE exploits this by extracting reusable direction bases from an initial expert bank and training only compact coordinate experts plus task-specific low-rank routers thereafter, beating the strongest baseline by up to 5.90 points while training under 1% of the parameters sequential LoRA needs for later tasks.

There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon Text-to-image and similar translation models generate along paths that start from noise rather than from the source modality, which constrains sampling algorithms and makes the mapping one-way, so image-to-text inversion is not available. BIT builds bidirectional diffusion bridges that interpolate directly from text into images, giving a source-aware generative path and an endpoint-conditioned process traversable in either direction, derived through stochastic calculus into stochastic differential equation forms with tractable losses that scale to high dimensions. It is competitive with denoising-diffusion and deterministic-flow baselines and better on several vision–language and natural-science evaluations, all within one unified framework.

Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models

Kairong Yu, Zixin Zhu, Le Yu, Hongwei Wang cross-listed Large vision-language models still generate content inconsistent with their image input, and most mitigations act from outside through extra supervision, output calibration, or attention tweaks rather than addressing what happens to internal representations during decoding. The diagnosis here is an inference-time failure mode in which cross-modal representations degrade as they pass through decoder layers and drift across generation steps, destabilising token prediction. Dynamic Alignment Compensation (DAC) is training-free: it detects representation divergence and applies lightweight residual corrections, combining Layer-wise Semantic Compensation against inter-layer degradation with Sequential Semantic Correction against temporal drift. Across nine hallucination-focused and general multimodal benchmarks and several backbones, DAC consistently lowers hallucination rates without hurting general performance.

Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs

Chenhong He, Lei Li, Shicheng Li, Hanglong Lv, Lingpeng Kong, Qi Liu et al. cross-listed Hybrid attention is standard in frontier language models, but the Vision Transformers (ViTs) inside multimodal models lack an agreed-upon hybrid design or an explanation for why some attention patterns beat others. Examining ViT attention heads reveals that they split into object-specialist and background-specialist roles most sharply under full attention, a property the authors name Semantic Head Specialization and measure with an SHS-Index that separates full-attention from chunk-window ViTs and tracks downstream benchmark scores. Three structural factors — window interaction, token serialization, and local softmax allocation — are identified as what shapes this specialization and used as design rules for Ariadne Attention, which matches full attention on 22 image and video tasks while using 6.5 times less attention compute.

Post-Training VLMs for Video Mistake Detection

Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami, Gianpiero Francesca, Juergen Gall cross-listed Systems that spot mistakes in videos of people following instructions are typically closed-set, so any new task means fresh data collection and retraining. The MD-VQA protocol and benchmark instead ask whether a model has learned the general notion of a mistake by testing, for both seen and unseen actions, whether a step was carried out correctly with respect to its written description. The proposed post-training method for video-language models uses a reward that pushes the model to find discrepancies between the instruction and what the video shows, beating zero-shot, supervised fine-tuning, and other post-training baselines, with up to 11.6% improvement over the best baseline on unseen procedures in EP-VQA.
6 more specialized papers

Reinforcement Learning 12

SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning

Musa Shams Logged trajectories in offline goal-conditioned reinforcement learning (GCRL) are often split into segments for storage or bookkeeping reasons unrelated to episode termination, yet multi-step targets treat those cuts as real endpoints. SegBench-GC holds transitions, source trajectories, goal sampling and optimization fixed while varying only artificial backup boundaries, comparing naive absorbing treatment against continuation-valid targets (CVT), which stop reward accumulation at a cut but still bootstrap from the stored successor. On PointMaze with 35,000 artificial cuts, success falls from 50.5% uncut to 39.1% with CVT and 19.1% when the same cuts are treated as terminal, and a replication on Puzzle-4x5 using the Decoupled Q-Chunking codebase collapses to 0.27% under naive handling. Critic diagnostics trace the failure to a large optimistic shift in learned values that CVT avoids.

Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess

Szymon Mi{\l}osz, Piotr Duch, Szymon Grabowski Searchless chess networks such as Leela Chess Zero's Chessformer reach master strength in a single forward pass by distilling Monte Carlo Tree Search visit counts, but imitating a search is a poor proxy for playing without one. Fine-tuning with self-play reinforcement learning, the usual entropy bonus (a reverse Kullback-Leibler divergence toward uniform) is swapped for a mass-covering forward KL toward the network's own MCTS prior, paired with a sampling temperature that sharpens once the value head is confident about the outcome. In roughly two thousand steps this lifts puzzle accuracy from 93.9% to 94.9% and mate-in-four from 77% to 81% without losing playing strength, while a control fine-tuned on puzzles alone posts the largest tactical gains yet loses about 260 Elo — a better puzzle-solver is not a stronger player. Without any regularizer, self-play collapses onto a single line of play.

Rubric-to-Code Credit Assignment for Reinforcement Learning

Rui Jin, Jikai Chen, Yihan Chen, Hao Zhou, Demin Zhu, Kaichen Yang et al. Generating an interactive web application involves many distinct user-facing requirements, each tied to a localized region of code such as an event handler, state update, or CSS selector, but GRPO compresses all of that into one sequence-level reward applied uniformly across every token. Rubric-to-Code Credit Assignment (RCCA) builds training tasks around explicit functional rubrics, uses a hierarchical reward that separates format, source-code, runtime, and functional failures, and aligns evaluator-written textual attributions with the code spans and tokens responsible. The resulting Ling-RCCA-Flash scores 41.25 on MiniAppBench, a 32.20-point gain over the Ling-3.0-Flash base model and slightly ahead of Claude Opus 4.5, and reaches 76.19 on ArtifactsBench, topping the official leaderboard setting.

Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling

Siyuan Gan, Yuhan Li, Xiran Wang, Linjian Meng, Boyan Wang, Zhen Zhao et al. Dynamic Sampling — the component contributing most of the accuracy gain of Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) over Group Relative Policy Optimization (GRPO) — filters out prompts whose sampled responses are all correct or all incorrect, removing zero-advantage gradients. The authors show theoretically that this filtering asymmetrically amplifies advantages within a prompt: on hard prompts the incorrect responses get amplified more than the rare correct ones, so the model learns mainly to avoid observed wrong answers rather than exploit hard-to-sample right ones. Their fix, Direct Advantage Amplification, boosts the advantages of those hard-to-sample correct responses and yields DA3PO, implemented in under 30 lines of change on top of DAPO, which outperforms GRPO and other classical GRPO variants in experiments.

Is Monte Carlo Tree Search Just Every-Visit Monte Carlo Control?

Xianyi Wu Monte Carlo Tree Search (MCTS) and every-visit Monte Carlo control are normally taught in different vocabularies — selection, expansion, simulation, and backup versus trajectory sampling, return estimation, action-value updating, and policy improvement. This expository note argues the distinction is largely terminological at the level of trajectory generation and value updating: the tree policy and rollout policy are the learned and not-yet-learned parts of one evolving policy, expansion is just first visit plus initialization, and backup is the ordinary every-visit Monte Carlo update. Under that reading the four stages collapse to two operations, sampling under the current policy and every-visit updating, making MCTS every-visit Monte Carlo control expressed in the data structure and language of search.

Emergent aggregation from collective foraging

Gorka Mu\~noz-Gil, Andrea L\'opez-Incera, Vide Ramsten, Giovanni Volpe, Thomas M\"uller, Hans J. Briegel cross-listed Models of flocking and swarming normally build in a direct social drive, rewarding or hard-wiring agents to approach or align with neighbours. Here reinforcement learning foragers start from a random walk and optimise only an individual reward for finding replenishable targets, while their sensors show them other foragers and never the targets themselves. As visual range grows the agents cross over sharply from environment-tuned individual search to a scale-agnostic collective one, and spatial aggregation appears exactly at that crossover — grouping emerges purely as a by-product of efficient foraging. A minimal analytical first-passage model reproduces the transition as a switch between the two search strategies, identifying indirect resource-driven reward as a general route to collective behaviour.

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning

Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma Long-horizon reinforcement learning for language-model agents usually collapses a task's success or failure into one scalar terminal reward that is broadcast evenly across every action, leaving no signal about which step mattered. VICT opens up the verifier that produced that reward, exposing its individual executable or evidence-backed checks as atoms and tracing them back to specific actions through dependency-valid proof edges, then redistributing group-relative advantage only along those edges; it abstains when the evidence is incomplete, preserves the original terminal reward, and modifies only the training-time advantage tensor, so it needs no learned critic, process labels, or branch rollouts. On ALFWorld and WebShop it improves substantially over outcome-only training while matching recent fine-grained credit-assignment methods, and ablations rule out dense atom rewards, final-commit credit, temporal proximity, and sparsity as sufficient explanations for the gain.

HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees

Boyuan Meng (Ant Group, China), Peihua Bao (Ant Group, China), Hong Liu (Ant Group, China) et al. Agentic reinforcement learning produces branching rollout trees whose trajectories share long prefixes, and training each root-to-leaf path independently recomputes that shared work; existing systems also assume full attention and lack differentiable hybrid-attention execution compatible with activation recomputation. HARTS jointly plans microbatches, data-parallel replica assignment, and slot schedules over compressed prefixes, and adds a linear-time algorithm for chunkwise linear attention that coordinates chunk-boundary state recovery and replay with the provably minimum number of sequential attention calls, packing all branches into one call per round while propagating gradients through differentiable state handoffs and restoring per-token log-probabilities. On an agentic workload built from SWE-bench tasks it delivers a 4.81–4.87× forward/backward/gradient speedup with activation recomputation, with numerical deviation comparable to baseline rerun noise and a matching reward trend over the first 120 steps of τ³-Bench training.

Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning

Minghui Xu, Zi Wang Analysis of failures on the Countdown arithmetic-puzzle task shows calculation errors account for a substantial share of wrong answers, motivating a study of whether a calculator tool helps. The authors build supervised fine-tuning data teaching tool-call patterns and how to interpret returned outputs, then apply on-policy reinforcement learning — RLOO, RLOO++, GRPO, and DAPO — with automatically verifiable final-answer rewards, evaluating on a fresh 1,024-problem held-out set with no exact overlap with training data. Tool access adds roughly 10 percentage points across pass@k for both the SFT and RL baselines, and Tool-DAPO raises pass@1 from 35.8% to 66.0% over the tool-using SFT model, with RL producing more effective tool use even though only the final answer is rewarded.

REPLICANT: Learning Policies for Evading and Hardening Malware Detectors

Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia, Alexander Herzog, Myles Foley, Chris Hicks et al. Attacks used to stress-test machine-learning malware detectors typically assume privileged access to training data, the feature space, or confidence scores, which overstates defender knowledge of the adversary. Replicant is a deep reinforcement learning framework that learns evasion under a strict label-only black-box threat model, producing a reusable policy governing both how to mutate a malware sample and when to spend a query on the target, and that policy transfers across samples, detectors, and feature spaces. Across seven Android malware detectors and three feature spaces it reaches a 78.8% mean attack success rate, a 20.9% to 39.2% relative gain over prior state of the art, and using it for adversarial training produces detectors with more generalizable robustness than existing hardening approaches.
2 more specialized papers

Reasoning 8

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

Shuai Wang, Jiayi Kuang, Yinghui Li, Haojing Huang, Xinnian Liang, Ying Shen et al. Large language models optimized purely for final-answer correctness on math may be memorizing solution patterns rather than internalizing concepts, and they are notably weak at example-based reasoning such as constructing counterexamples that probe a theorem's boundaries. INSPIRE tackles two obstacles to fixing this with preference optimization: weak baseline ability makes good preference pairs hard to build, and the skill must be acquired progressively. It pairs Reference-Guided Student Internalization, which generates preference candidates from the policy model's own distribution, with a stage-wise rubric training scheme that separates learning the method from learning to apply it correctly. Across multiple model scales and families the approach improves consistently and surpasses larger open-source models, with out-of-distribution benchmarks showing no loss of general math ability.

Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning

Neh Majmudar, Elena Filatova Linguistics olympiad puzzles are self-contained — every answer follows from the supplied examples and nothing else — which makes them a clean setting for asking whether a model is genuinely reading its context or leaning on prior knowledge. Across 53 UK Linguistics Olympiad puzzles, each is perturbed by deleting one context example either at random or, using an idea borrowed from error-correcting codes, targeting the structurally load-bearing example that uniquely carries the needed information; a Question Damage Score then classifies puzzles as fragile or robust. Instructed to abstain when information was insufficient, three frontier LLMs rarely abstained and frequently still produced correct answers after the load-bearing example was removed.

The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs

Yucheng Wang, Yuetian Du, Zhengyi Liu, Rongyu Zhang, Bing Zhao, Boyu Yang et al. Counterfactual reasoning benchmarks typically fix the variables and grade against a single gold outcome, which never tests whether a model can trace how an altered condition propagates through downstream consequences. WhatIfBench supplies 220 open-domain, open-form what-if questions spanning STEM, humanities and social sciences, and hybrid scenarios, and PRISM grades free-form answers by first converting each explanation into a semantic causal graph of events, states, and mechanisms, then scoring both the graph's causal validity and the answer's explanatory adequacy. Across six frontier models, the strongest reaches only 64.62%, with recurring causal gaps, premise drift, and topology fragmentation indicating that fluent counterfactual narratives often sit on fragile causal reasoning.

SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing

Wanli Cheng, Haiya Xiang, Juntao Li, Hongling Wang, Wenliang Chen Long chain-of-thought reasoning keeps spending tokens well after the intermediate answer has effectively stabilized, and existing early-exit signals based on confidence or entropy capture that stability poorly, while consistency checks require several sequential rollouts before they can trigger. SABER applies simple semantic perturbations to the intermediate reasoning state to form adversarial branches, then uses lightweight probing to estimate each branch's likely final answer without full rollouts, exiting when the branches agree and continuing when they diverge. The method needs no training and cuts reasoning token consumption by 30.2% to 39.8% on average while staying competitive with full-length reasoning accuracy across several benchmarks and model architectures.

AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning

Ziming Wang, Ivor Tsang, Hangwei Qian Test-time scaling generates extra candidate solutions to improve reasoning, but spending the same budget on every question wastes computation, and existing stopping rules assume that stronger current evidence — confidence, agreement, or answer stability — means more computation is unnecessary. The authors show checkpoint-level correctness actually evolves non-monotonically, so evidence can strengthen right before an answer collapses, and propose AERA (Adaptive Evidence Residual Allocation), a sequential controller that predicts whether another block of generation is likely to recover a better answer using answer-distribution, temporal, re-solving, semantic, and compute features observable at each checkpoint. Future correctness is used only to build offline training supervision, never at inference. On a held-out set of 300 GSM8K questions with frozen thresholds, AERA reaches 92.61% accuracy against 93.01% for a fixed 128-response budget while cutting completion tokens by 95.99%, with similar behavior on GPQA Diamond.

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang, Shizhuo Hou, YongXiang Hua et al. On-policy self-distillation trains a student that sees only the problem on its own rollouts, supervised token-by-token by a privileged teacher that also sees a reference solution — but the teacher is treated as a fixed target, and privileged conditioning does not guarantee it is the right target for problem-only reasoning. VISTA keeps the usual student update and adds a reverse direction: outcome-verified rollouts are used to adapt the teacher toward the student, restricted to the top-k token positions with the largest teacher-student KL divergence, reusing the same rollouts and loss with no extra sampling or reward objective. Across AIME24, AIME25, and HMMT25 with Qwen3 at 1.7B, 4B, and 8B, it took the highest Avg@12 at every scale, gaining 0.6, 0.7, and 2.1 points over standard self-distillation.

Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs

Vishvesh Bhat Supervised fine-tuning and reinforcement learning both bury an acquired reasoning capability inside model weights, where it cannot be inspected step by step or moved to another model. PLVR (Program Learning with Verifiable Rewards) instead learns reasoning as an explicit program of deterministic and neural primitives directly from input-output examples, using symbolic backpropagation: each program layer carries a typed ontology, a loss is computed against ground truth, and required input ontologies propagate backward by type inference over primitive signatures, so credit assignment is a derivation and the reward is a per-step contract verdict rather than a terminal outcome. On LiveCodeBench v6 and Tau2Bench, 30B base models with PLVR beat RL at matched budget by 27.8 points on average and beat frontier models an order of magnitude larger by 13.6 points, while an ablation swapping loss-guided search for uniform sampling over the same type-admissible space drops the median program from 65.6 to 17.5, which the authors read as the backward pass rather than the type system carrying the advantage.

NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

Samuel Xiao, Judy Song, Rory Hu, Ziliang Zong Neuro-symbolic geometry provers such as AlphaGeometry reach near-IMO-gold performance but require problems written in a specialized domain-specific language (DSL), and translating natural language into that syntax by hand is a usability bottleneck. NL2AGBench measures how well ten open- and closed-source large language models (LLMs) perform that auto-formalization, judging output by whether it actually executes inside AlphaGeometry rather than by textual similarity. Leading closed-source models exceed 80% executable translation rate while even the largest open-source models fail to consistently preserve geometric constraints, and the authors add an error taxonomy separating syntax from logic failures plus mitigations via few-shot prompting, fine-tuning, and human-guided hinting.

Robotics 7

PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models

Davood Soleymanzadeh, Kaidi Zhang, Zhiyuan Zhang, Bihao Zhang, Xiao Liang, Yu She et al. cross-listed Vision-language-action models (VLAs) map instructions and images directly to robot actions but usually condition only on the current observation, giving them no explicit way to reason about how a task will unfold. PHR-VLA adds a lightweight auxiliary future head used only during training, which aligns the model's internal representations with privileged latent dynamics extracted from future observations. Patch-level, contact-centric supervision from the wrist camera raises success on LIBERO from 84.1% to 88.4% and on real-world disassembly tasks from 63.3% to 82.5%, with a smaller gain on Meta-World when supervision comes from a third-person view.

Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring

Marin Maletic, Goran Vasiljevic cross-listed Robotic sorting of recyclable waste is difficult because targets such as used beverage cartons arrive crushed into inconsistent shapes. The proposed system needs no training: an open-vocabulary vision-language model detects cartons from a text prompt, SAM2 refines each detection into an instance mask, and a geometric score combining surface flatness with normal alignment picks the suction point, with k-nearest-neighbor PCA, Sobel cross-product, and RANSAC plane fitting compared for that stage. Tested on a real robot across three deformation levels and 35 cluttered scenes, single-object grasp success reached 88.2% and end-to-end retrieval from clutter 72.6%.

Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning

Nan Wang, Mohit Yadav, Jonathan Wulff, Aidan Rosenbaum, Kezhou Chen, Yuvan Sharma et al. cross-listed Tendon-driven robot hands are cheap to build because motors sit off the joints and one cable can drive several joints, but that same underactuated transmission is hard to model in simulation and leaves coupled joints that cannot be commanded independently, which makes learning on them difficult. Aero Hand Open ships as a simulation-ready package: a simulation model that reproduces the cable transmission itself, an identified bidirectional actuation map linking simulation to motor commands including the thumb's three-way coupling, and a reinforcement learning training environment. Policies trained entirely in simulation transfer directly to the physical hand with no fine-tuning and no state estimation, and the mechanical design, simulation model, mapping, training environment, and deployment stack are all released.
4 more specialized papers

Unclassified 2

AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics

Tejas Srinivasan, Shikib Mehri, Nandita Shankar Naik, Anirban Das, William M. Campbell, Jesse Thomason No summary available — see the abstract on arXiv.

Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine Learning

Sejong Oh No summary available — see the abstract on arXiv.