Friday, August 14, 2026

417 papers cs.AI · cs.LG · cs.CL ← 2026-08-132026-08-17 →

Jul Aug Sep

Highlights

Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet

Highlight Reinforcement Learning Eliseo Curcio Reinforcement-learning post-training now dominates language model development, but its power draw on GPUs has not been characterized, and datacenters still manage that power with workload-blind static caps and reactive throttling. Group relative policy optimization (GRPO) training was instrumented with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s, yielding over 380,000 samples used to train a proximal policy optimization meta-controller that adapts the workload's own generation parameters to measured power. On the 7B trace the controller cuts power-limit violations by 89.8% while raising token output 18.1% and tokens per megawatt-hour by 26.2%; the same controller family produced null results at 72B, diagnosed as the group-size actuator losing authority under model sharding, and a rebuilt controller using generation concurrency instead delivered 35.7% more output than a static safe baseline at about 2.3% violations. Measured over realistic 30-second windows a composed 16-GPU fleet shows zero violations with peak demand at half of nameplate, suggesting roughly twofold oversubscription is feasible subject to operator validation.

Datacenter GPU power management relies on workload-blind static caps and reactive throttling, while the power behavior of reinforcement-learning post-training itself has never been characterized on real hardware. The work instruments GRPO training with half-second telemetry across 7B, 14B, and 72B models on one to four A100s (380,000+ samples) and trains a PPO meta-controller that steers the workload's own generation parameters in response to measured power, rather than clamping the hardware.

  • On a full 500-step 7B trace the controller cuts power-limit violations by 89.8% while simultaneously raising token output 18.1% and energy efficiency 26.2% in tokens per MWh, showing the safety and throughput objectives are not inherently in tension.
  • Deployed live at 72B the same controller family produced replicated null results, diagnosed as the group-size actuator losing authority once the model is sharded — a negative result the authors trace to a specific mechanism rather than leaving unexplained.
  • An actuator-authority sweep found the same parameters applied as generation concurrency retain 17-22% power authority, isolating an occupancy-versus-volume principle: what moves power is how many requests are resident on the GPU, not how many tokens they collectively produce.
  • Rebuilt on the concurrency actuator and run across three replications on a live 72B rollout-generation workload, the controller delivered 35.7% more output than a static safe baseline at 2.27 ± 1.08% budget violations and 87.2% fewer violations than uncontrolled operation, with the best mean throughput and energy per token among constrained controllers — though a much simpler adaptive threshold rule matched it in one of the three operating conditions.
  • Measurement window turns out to dominate the violation picture, with 72B transients falling from 23.6% at half-second resolution to 1.6% at 30 seconds and zero at 5 minutes, and a composed 16-GPU fleet showing zero violations at 30 seconds or longer with peak demand at 50-56% of nameplate, suggesting roughly twofold oversubscription is feasible for this workload mix — a projection from composed rather than physically deployed fleet traces, on a single hardware generation, that the authors themselves gate behind operator validation and a proposed low-cost pilot.

Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

Highlight Reasoning Mark Shapiro A dense pretrained model, Qwen2.5-0.5B-Instruct, is split into a prelude, a weight-tied recurrent block, and a coda so its middle layers can run repeatedly and learn an iterative latent update rather than a terminal-answer lookup. Training with intermediate-step supervision teaches the model to compute one task step per loop, and the behavior persists when only final answers are graded; a 6M-parameter adapter over frozen weights matches full 180M-parameter block training (83.8% versus 84.0%), and the learned operation extrapolates to roughly 1.5 times its supervised depth, holding 70% accuracy at depth 18. Against a same-size model trained on written scratchpads, the recurrent version scores 84% versus 72% overall, retains 53% versus 2.5% beyond depth 10, and answers 7.6 times faster. Teaching the inverse rule proved impossible without destroying the installed mechanism, marking a catastrophic-interference boundary.

Adding computation depth to a transformer normally means more layers at design time or more generated tokens at inference; this work asks whether a third axis — repeatedly applying one shared block — can be surgically retrofitted into an already-pretrained dense model instead of trained from scratch. Qwen2.5-0.5B-Instruct is cut into a Prelude, a weight-tied Recurrent Block, and a Coda, joined by a trainable bridge that re-injects the Prelude representation on loops 2 and beyond, with the one-loop route reproducing the base computation exactly so that loop-1 evaluation measures preservation rather than a different model.

  • Training uses exact intermediate-state supervision on a synthetic pointer-chasing task so that each execution of the block performs exactly one rule application, and the stepwise behavior survives a final 1,000 steps of outcome-only annealing: 97.7% of held-out intermediate states still decode correctly, and when a depth-3 problem is forced to run to loop 8, 93.0% of the extra states contain the next symbol in the chain while only 1 of 384 parks on the answer — evidence that a reusable procedure, not terminal-answer lookup, was installed.
  • The identical surgery installs at two very different budgets — 180.6M forward-active full-block parameters versus a rank-16 LoRA adapter at 6.01M (3.3% of the full budget, base weights untouched and hash-verified) — and the two are statistically indistinguishable overall at 83.8% versus 84.0%, with the adapter actually ahead through depth 11 and the full block winning only at depths 12–14, an interpolated crossover at depth 11.54.
  • The installed operation extrapolates to roughly 1.44–1.50× its supervised depth across supports 4, 6, 8, and 12 and two alphabet sizes, with the support-12 run holding 70.3% at depth 18 before falling to 46.1% at depth 20 and 10.9% at depth 22 — a support-scaled frontier with a real tail ceiling rather than open-ended algorithmic generalization, and one that requires training dose to scale with support.
  • Against preregistered dense controls on 1,792 identical frozen rows, the recurrent model beat a same-size serialized-scratchpad SFT baseline 84% versus 72% overall and 53.1% versus 2.5% on depths 11–14 (the scratchpad is at ceiling through depth 10, then collapses to zero), while answering 7.6× faster at depth 14 (301.8 ms versus ~2.28 s) from a single decoded state instead of 96 generated tokens, at constant sequence length and no KV-cache growth.
  • Limits are reported as carefully as the wins: zero-shot transfer to verbal renderings is minimal at both budgets (~16–20%) though installed history beats fresh training by 18.6 points once verbal fine-tuning starts, preservation is non-inferior on the ARC battery only at loop 1 (forced loop 8 degrades ARC-Challenge to 71 versus 155), learned depth selection remains an open registered negative, and the reverse-lookup task marks a catastrophic-interference boundary — learnable in isolation at 63/64, but no continuation at either budget acquired it while keeping the installed mechanism and general capability.

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

Highlight Large Language Models Angelo Nardone, Paolo Ferragina Neural compressors built on large language models achieve far better lossless compression ratios on text and code than zstd, gzip, or bzip, but their one-symbol-per-step autoregressive decoding makes throughput impractical. Swapping in diffusion language models removes that bottleneck since each forward pass can encode a variable number of symbols at freely chosen positions, though this breaks the assumptions of the standard compression pipeline and requires new algorithms for deciding what gets encoded when. Evaluated on the enwik8 benchmark against both language-model-based and general-purpose compressors, the diffusion framework advances the state of the art in lossless text compression.

Neural compressors built on autoregressive LLMs beat gzip and zstd on text by a wide margin but run at kilobits per second, because the model must do one forward pass per token. Replacing the autoregressive backbone with a masked diffusion language model lets the compressor commit many tokens per forward pass, turning sequential depth from an architectural constraint into a tunable knob on the compression–throughput curve.

  • The pipeline starts from a fully masked window, and at each of K denoising steps the DLM (LLaDA-8B) emits distributions over all masked positions at once; a chosen subset is encoded with either rANS or symbol ranking plus a general-purpose compressor, then revealed as context for the next step.
  • Two new design problems arise that autoregressive models never face: the commitment schedule (uniform blocks, geometric-growth progressive, or entropy-based, committing every position whose predicted entropy falls below a threshold τ) and the initial context (leaving c tokens unmasked and stored in the header, selected contiguously, randomly, by token rarity, or by Attention Rollout scores).
  • On the first 10 MB of enwik8, the diffusion pipeline spans 0.127 compression ratio at 1.01 kbit/s (rANS) through 0.246 at 165.79 kbit/s (symbol ranking), against FineZip's 0.079 at 0.10 kbit/s and LLMZip's 0.079 at 0.09 kbit/s — up to four orders of magnitude more throughput while still compressing better than general-purpose tools.
  • The decompression gap is larger still: prior LLM methods use teacher forcing to compress in one pass but cannot during decoding, so they need n sequential passes per window, whereas the diffusion pipeline's decompression throughput matches its compression throughput almost exactly (166.12 vs 165.79 kbit/s), which the authors argue is roughly six orders of magnitude faster at decode.
  • Ablations favor the entropy-based schedule for ratio at any token budget and attention-based initial context across all configurations — contiguous prefixes, the only option available to LLMs, degrade below random selection at high throughput — but the work rests on a single 8B model and a single benchmark, prior decompression figures are estimated rather than measured, and even the fastest setting remains around six orders of magnitude slower than zstd.

When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

Highlight Reasoning Utkarsh Bahuguna Self-consistency — sampling many chains of thought and returning the plurality answer — is a standard way to spend inference compute, on the assumption that voting helps on average. On the full GPQA Diamond benchmark of 198 graduate-level science questions, majority voting lowers per-problem accuracy on 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B, an effect pre-registered on a 151-problem confirmatory split after being spotted on 47 exploratory problems, with all four confirmatory hypotheses passing. An oracle routing each problem to its best sample count marks an upper bound 14 to 17 accuracy points above single-sample decoding, but no verifier-free gate reaches it — neither plurality agreement nor token entropy moves accuracy by more than 0.002 — because confidence does not track correctness here: in Llama's highest-agreement bin accuracy is worse than in its lowest. The authors flag reasoning-native models as untested.

Self-consistency is usually treated as a free accuracy win — sample more chains of thought, return the plurality answer — but on graduate-level science questions that assumption inverts: when a small model concentrates probability on a wrong answer, extra samples only entrench it. On the full GPQA Diamond benchmark, majority voting at N=64 lowers per-problem accuracy for 56.6% of problems with Qwen2.5-7B and 65.7% with Llama-3-8B, and two cheap verifier-free gates fail to detect which problems to skip.

  • The design is pre-registered: hypotheses and thresholds were locked and git-tagged on a 47-problem exploratory split, then tested on the held-out 151 problems at 64 samples per question and temperature 0.7, with all four confirmatory hypotheses passing.
  • Aggregate accuracy hides the damage — voting moves Qwen from 0.342 to 0.369 and Llama from 0.273 to 0.313, a nearly flat curve underneath which the majority of individual problems degrade, with the worst single problem losing 47 points and the rare big winners gaining up to 70.
  • A grid oracle routing each problem to its best N in {1,2,4,8,16,32,64} reaches 0.482 for Qwen and 0.439 for Llama, a ceiling 14 to 17 points above single-sample accuracy, but it requires ground truth and is an upper bound rather than a deployable method.
  • Neither verifier-free gate gets there: a plurality-agreement gate and a token-entropy gate each move accuracy by less than 0.002 against fixed-budget voting at N=64, buying compute savings rather than correctness, because agreement does not track correctness — in the highest-agreement bin Qwen's plurality is right only 52.5% of the time and Llama's just 28.6%, below its own lowest-agreement bin.
  • The scope is narrow by the authors' own account: one benchmark, two small non-reasoning models with Llama sitting barely above the 25% chance baseline, an entropy threshold tuned in-sample on the same 151 problems it is scored on, and no test of reasoning-native models — an attempted evaluation stalled when hidden chain-of-thought exhausted the output budget — which they flag as the central open question.

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

Highlight Agents Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang Agent skills — reusable natural-language guidance loaded into large language model agents — have shown mixed results, sometimes raising task success and sometimes inflating cost or breaking tasks outright. A differential analysis framework attributes each failure or cost regression to a specific skill by comparing a skill-guided run against a no-skill or semantically matched reference run on the same task, instantiated on SkillsBench and SWE-Skills-Bench and supported by a triage tool called SkillTriage. The analysis surfaces 307 skill-induced problems: 125 functional failures and 182 efficiency regressions, and the functional failures rarely trace to obviously irrelevant skills — plausibly relevant skills instead push the agent to misimplement or omit required elements. Efficiency regressions are not explained by prompt length; the largest contributors are excessive verification (67 cases) and heavy implementation pipelines (30 cases), where skills convert checklists and recipes into mandatory work.

Agent skills — reusable SKILL.md instruction packages loaded into an agent's context — are known to sometimes hurt task success or inflate cost, but aggregate pass rates cannot say which skill caused what. This work borrows from differential testing: each skill-guided run is paired with a no-skill or semantically matched-skill reference run on the identical task, verifier, repository state, and model, so a failure or cost blowup can be attributed to the loaded skill rather than to base-agent variance.

  • The authors expanded SkillsBench and SWE-Skills-Bench with public skills scraped from smithery.ai and skillsmp.com (metadata embedded with all-MiniLM-L6-v2, cosine similarity ≥ 0.7, top-5 kept), growing the comparison space from 826 to 20,664 paired comparisons, ran them under OpenCode 1.15.1 with Claude Opus 4.6, and hand-refined 665 candidates down to 307 confirmed skill-induced failures — 125 functional failures and 182 efficiency regressions.
  • Functional failures are overwhelmingly caused by skills that are on-topic: only 2 of 125 (1.6%) are applicability mismatches, while Task-Implementation Fault accounts for 86 (68.8%) — the agent treats a skill's reusable defaults, examples, or templates as task requirements and either fills a required element wrongly (46 cases, e.g. computing net exports as a bare ratio instead of a percentage) or omits it entirely (36 cases).
  • A further 24 cases (19.2%) are artifact misplacement, where the agent writes a perfectly good implementation to a repository-conventional path instead of the task-specified one, and 13 (10.4%) are environment mismatches where the skill pushes the agent to pip install or path-manipulate itself into a runtime state the verifier does not share.
  • Cost regressions are not a prompt-length story: Excessive Procedure accounts for 114 of 182 (62.6%) — dominated by excessive verification (67 cases) and heavy implementation pipelines (30) — while context overhead explains only 46 cases, and within those the always-resident skill body is responsible for 43 of 46, implying that mandatory body text, not lazy-loaded supplementary material, is the real context tax.
  • SkillTriage, a taxonomy-guided attribution tool using GPT-5.5 with 2-of-3 majority voting, recovers the exact manual subcategory for 111/125 (88.8%) functional failures and 132/182 (72.5%) efficiency regressions, with residual errors clustering at taxonomy boundaries — separating exploration from pipeline from verification cost drops to 68.4% within the Excessive Procedure subset.
  • The main caveats are that every result comes from one agent harness and one model, that the T = 2.0 regression threshold and the "reference run passes" pseudo-oracle are conservative but arbitrary choices, and that labeling required manual judgment on ambiguous trajectories despite group consensus and the exclusion of verifier-narrow cases.

OEIS Open: How many conjectures can language models turn into theorems?

Highlight Reasoning Tom Adamczewski OEIS Open packages 492 genuinely open mathematical conjectures from the Online Encyclopedia of Integer Sequences, formalized in Lean by prior work, into an open-source evaluation harness that any language model can be run against and that is hardened against cheating attempts — previously these conjectures had only been attacked by one bespoke agent. With a minimal toolset and a $50 budget per attempt, models resolve 147 of the conjectures, scoring 30%, and on the 100-conjecture OEIS Open Lite subset the best current model reaches 44% at a $200 budget. Giving models access to 476,000 arXiv papers did not help, and neither did more elaborate agent loops. The authors note the conjectures are of uncertain mathematical significance and mostly little-studied, but that models nonetheless close open research questions autonomously at modest cost.

OEIS Open turns 492 open conjectures from the Online Encyclopedia of Integer Sequences — previously formalized in Lean by Tsoukalas et al. — into a benchmark any generic language model can be run against, replacing computational verifiers with the Lean kernel so that a submission counts only if it proves the conjecture or its negation outright. The headline finding is that a deliberately minimal agent resolves a substantial fraction of these open problems for a few dollars each, beating a far more elaborate specialized system.

  • The harness is a ReAct loop with just three tools (bash, a text editor, and a budget reporter) over a container holding Lean 4 with Mathlib, SageMath, and Python, with each attempt split across three network-isolated Docker containers — agent, compiler, and scorer — so that SafeVerify checks the submitted olean against the authors' own copy of the statement, requiring kernel-identical types and permitting no axioms beyond propext, Quot.sound, and Classical.choice.
  • On the full 492-conjecture set at a $50 per-conjecture cap, Claude Opus 4.8 resolved 30% (147), GPT-5.5 26%, and Gemini 3.5 Flash 22%, against 9% (44) for AlphaProof Nexus, the AlphaEvolve-style evolutionary search with AlphaProof subgoal calls that produced the original results — at a comparable average cost of roughly $6–$10 per resolved conjecture.
  • On OEIS Open Lite, a random 100-conjecture subset run at $200 per attempt, scores ranged from 29% (Gemini 3.5 Flash) to 44% (Claude Fable 5), with solve rates rising roughly log-linearly at about ten percentage points per tenfold increase in spend and no visible plateau, implying about 216 of 492 conjectures at the higher budget.
  • Two plausible upgrades did nothing: giving agents an offline corpus of 476,000 pure-mathematics arXiv papers left Lite accuracy unchanged, as did swapping the ReAct loop for Inspect's deepagent with subagent delegation, persistent memory, and a todo tool — a result the author reads as the bitter lesson applied to proof search.
  • The main caveats are that the conjectures are of uncertain mathematical significance (a single prolific contributor, Zhi-Wei Sun, proposed 37% of the set, and 47% of the underlying sequences list no links or references at all), the formalizations were taken as-is and some are likely misformalized, and an independent re-check with Comparator scored Claude Opus 4.8 at 144/492 (29%) rather than 147.

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Highlight HF pick · 13▲Large Language Models Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong et al. Massive activations — the extreme hidden-state outliers known from full-attention transformers — are characterized for the first time in hybrid linear attention (HLA) models that interleave linear and full attention layers, revealing two architecture-aligned shapes: spikes that appear immediately before full attention layers, and plateaus that persist across the linear attention layers in between. As full attention becomes denser, consecutive spikes link up through the plateaus and recover the stable pattern seen in full-attention models, a recurrence confirmed across five linear attention architectures, six hybridization settings, five data domains, and open models from 1.2B to 397B parameters. Controlled pretraining of gated DeltaNet hybrids up to 1.3B shows both shapes emerge early and respond asymmetrically to output gating, and an outlier lifecycle analysis attributes spikes to a localized write-sink-cancel process and plateaus to delayed cancellation.

Hybrid linear attention models interleave cheap recurrent layers with a few full attention layers, but nobody had characterized how that layout reshapes the extreme hidden-state outliers known as massive activations. The study finds that these activations organize themselves around attention placement: they spike in the layer immediately preceding every full attention layer, and as full attention grows denser they stop decaying in between, merging into sustained plateaus that converge on the flat profile of a pure Transformer.

  • Because magnitude ranking becomes unreliable under hybridization — the top-activated token switches between adjacent layers far more often than in full attention models — the authors anchor tracking to a consensus attention sink token, averaging attention mass across all full attention layers and heads, then trace that fixed token's peak activation through depth.
  • Pre-attention spikes are near-universal: sink–spike alignment rates hit 99.4–100% macro-averaged across RetNet, HGRN, GLA, DeltaNet, and GDN at both 340M and 1.3B, holding across WikiText-103, Scientific Papers, GSM8K, CodeSearchNet, and FLORES-200.
  • Inter-spike retention climbs sharply with attention density — for GDN at 1.3B it goes from 18.4% at a 24-layer-per-attention 12:1 ratio to 26.6% at 6:1 and 77.8% at 3:1 — and the same alignment appears in open-source hybrids from Kimi Linear, Qwen3.5, Nemotron-H, and Zamba2 spanning 1.2B to 397B parameters, including state-space rather than linear-attention mixers.
  • Controlled pretraining of 24-layer GDN hybrids at 340M and 1.3B shows the spike is already visible after 1B tokens and tracks whatever depth the full attention layer is placed at, while output gating is asymmetric: gating full attention strongly attenuates both morphologies without erasing their layerwise structure, whereas removing GDN's native output gates only modestly amplifies them.
  • The proposed mechanism is a single write–sink–cancel lifecycle differing only in cancellation timing — the pre-attention layer writes a large signed outlier, the sink token absorbs attention, and an opposite-signed update cancels it either promptly (spike) or after several layers (plateau) — though this rests on fixed-coordinate tracing of representative cases rather than a causal intervention, and the paper leaves open both what regulates cancellation timing and whether the two regimes serve different computational roles.

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

Highlight Large Language Models Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi Training with ever-longer contexts is usually assumed to be harmless or helpful, but abundant in-context evidence may remove the incentive to store that information in weights, a tension named the Information Abundance Paradox. Pretraining experiments with long documents show that widening the context window improves language modeling, natural language understanding, and closed-book multiple-choice question answering only up to an intermediate optimum, after which all three consistently decline; in supervised fine-tuning, more task-relevant context at train time helps when supporting context is present at test time but reduces robustness when it is absent or misleading. Mechanistically, informative context shifts gradient pressure away from feed-forward layers, which are associated with parametric knowledge, toward attention, and causal interventions confirm this shift increases reliance on context at inference — suggesting that scaling toward near-infinite context is not simply a matter of supplying more high-quality long data.

Overview

Long-context training is usually assumed to be a free win — more evidence in the window can only help — and when it underperforms, the blame falls on scarce high-quality long documents. This work argues the context window instead changes what the model learns: when task-relevant information is abundant in the training context, the model can lower its loss by reading it rather than encoding it in weights, shifting from parametric internalization to contextualization.

  • Under a token- and update-matched pretraining sweep on 10B Project Gutenberg tokens across four scales (20M to 750M parameters) with context windows from 512 to 32768, downstream performance traces an inverted U — SuperGLUE and closed-book MCQA peak near 2048 tokens and language-modeling loss bottoms out near 8192 — with the degradation past the optimum persisting at every model size, so added capacity does not rescue it.
  • Fine-tuning Qwen3 models from 0.6B to 14B with LoRA on four MMLU-Pro domains under a fixed eight-document budget, while varying how many of those documents come from the target domain (k = 0, 4, 8), produces what the authors call context addiction: more relevant training context raises accuracy when supporting context is present at test time but lowers no-context accuracy and widens the gap under deliberately conflicting context.
  • The theoretical account is an achievability result — because a shorter window is a projection of a longer one, any short-context predictor can be simulated with the longer input, so the minimum task information that must live in the weights to hit a given risk threshold is monotonically non-increasing in context length.
  • Mechanistically, longer or more relevant training context lowers the feed-forward-to-attention gradient norm ratio (significant in 19 of 20 fine-tuning model-domain comparisons and at every pretraining scale) and raises inference-time attention mass on context tokens, concentrated in middle layers; module-restricted fine-tuning supplies the causal leg, with feed-forward-only updates improving no-context robustness and attention-only updates buying supporting-context accuracy at the cost of conflict vulnerability.
  • Synthetic pretraining on four rule-learning tasks shows the effect is selective rather than universal — the supporting-conflicting gap grows for bitwise and string operations, where average training gradient norm falls, but stays flat for mod10 arithmetic and the Caesar cipher — supporting the reading that context addiction appears only when demonstrations offer a lower-complexity optimization path than internalizing the rule.

The main limitation the authors name is scale: pretraining tops out at 750M parameters, leaving open where the inflection points sit for frontier-size models, larger data, and more compute.

AVA-Encoder: Towards Agent-Native Video Representation Learning

Highlight HF pick · 4▲Multimodal Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li et al. Agents that generate cinematic video have no structured way to learn from existing high-quality films, because film content lacks a representation that is both faithful and directly editable by an agent. The Agentic Video Auto-Encoder (AVA-Encoder) encodes a video into a knowledge graph whose hierarchy and state nodes hold structured text and whose linked asset layer holds generated images, audio, and clips, then reconstructs video from that graph; the reconstruction gap is converted into natural-language update directions that pseudo-train the encoding policy offline and can refine the graph at test time. Reported results show a 20.7 percentage-point improvement over the strongest external baseline, and the pseudo-trained shot-level encoder policy beats a hand-tuned policy while using 74.3% fewer system-prompt tokens. The authors release the framework, a reconstruction benchmark, and a dataset of film knowledge-graph representations.

Video-creation agents can't learn from professional films because a finished film exposes none of the decisions behind it, and no existing representation is simultaneously agent-readable, agent-editable, and faithful enough to regenerate the footage. AVA-Encoder reframes the problem as agentic auto-encoding: a film is encoded into a text-centered knowledge graph, reconstructed by a fixed generative decoder, and the reconstruction error becomes the training signal for the encoder itself.

  • The encoder analyses a film hierarchically — film, then shot, then keyframe, each level injected with the context above it — and emits a graph of nine structured-text node types (a Story–Event–Shot spine plus Character, Scene, Object, Style, Camera, and Audio states) joined by eleven typed edges, with every image, audio, and video asset held in a separate linked layer that is generated from the text, since no source frame is ever stored in the graph or shown to the decoder.
  • Optimization runs as two non-nested textual-gradient loops: an outer "pseudo-training" stage rewrites the shared shot-level system prompt across a stream of videos behind an anti-forgetting replay gate, while an optional inner stage refines one input's graph at test time behind an anti-degradation gate, both scored by roughly 30 binary atomic questions per shot or keyframe asked against the reconstruction.
  • On the authors' reconstruction benchmark AVA-Encoder reaches 49.0% overall against 28.3% for the strongest baseline soap2soap (VideoAnalyzer 21.5%, Storyboard Studio 19.4%), a 20.7-point gap that holds across all four directions — +21.1 on direct video, +34.2 on keyframes, +13.9 and +11.6 on the two back-captioning directions.
  • Ablations credit the two loops with +6.6 points (15.6% relative) over the un-optimized hierarchical encoder at 42.4%, and the pseudo-trained policy alone beats a hand-tuned prompt 45.8% vs 44.4% while using 8,052 instead of 31,336 system-prompt tokens (74.3% fewer) — though stripping the acceptance gates collapses the full system to 43.5%, below the human-tuned policy, so the gates rather than the gradients carry much of the gain.
  • The caveats are substantial: absolute fidelity remains low (29.7% on video back-captioning), the scale is six pseudo-training clips generalizing to 18 evaluation clips, only the shot-level policy is ever updated while the film and keyframe prompts stay frozen, and the benchmark is self-authored with Gemini-3.1-Pro-Preview acting as both the encoding and the judging model (reported at 97.3% agreement with humans on 710 of 730 blinded triples), with every score conditioned on one fixed Nano Banana Pro plus HappyHorse 1.0 decoder pair.

Intern-S2-Preview: Scientific Agentic Foundation Model

Highlight HF pick · 38▲Agents Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng et al. Scientific work demands models that read heterogeneous evidence, call tools, and hold up over long task horizons, which general foundation models are not trained for. Intern-S2-Preview is a family of scientific agentic models built from multimodal pretraining on rendered scientific documents and interleaved image-text corpora, followed by supervised fine-tuning, multi-task and agentic reinforcement learning, and on-policy distillation, supported by engineering such as partial rollout with off-policy correction, adaptive length regularization, and online speculative decoding. The 397B-parameter model extends long-sequence understanding to numerical time-series forecasting and reports competitive or leading scores across scientific, multimodal, agentic, and general benchmarks, while a separate memory-augmented Intern-MemDec-4B module lifts the Biology-Instructions average from 56.92 to 60.32 with the backbone left frozen.

Scientific work demands models that reason over heterogeneous evidence, drive tools, and hold progress across long horizons, but general LLMs lack the domain modalities and scientific multimodal models are still evaluated as static question-answering systems. Intern-S2-Preview from Shanghai AI Laboratory addresses this with a 397B-parameter scientific agentic foundation model trained through scientific multimodal pre-training followed by a staged post-training pipeline that ends in agentic reinforcement learning.

  • Pre-training learns from scientific documents in both textual and visual form: a Visual Pre-training stage predicts foreground visual latents from rendered PDF pages with a contrastive next-latent objective and no OCR or paired annotation, while a parallel pipeline builds layout-aware interleaved image-text sequences filtered by "visual gain", the perplexity drop on page text when figures, tables, and equations are included.
  • Post-training runs supervised fine-tuning, then multi-task RL with verifiable rewards, then black- and white-box agentic RL built on a harness × task abstraction that lets external runtimes such as Claude Code, OpenHands, and Mini-SWE keep their native control loops while the serving layer transparently captures token IDs, rollout log-probabilities, and MoE router experts for training.
  • Several systems techniques keep long-rollout RL stable and fast: partial rollouts are paused rather than discarded and resumed after a policy update with per-token importance weights clipped against the exact behavior-policy version, Rollout Routing Replay plus a bidirectional-KL token mask reconcile the inference and training engines for the MoE, and an online-adapted draft model gives roughly 2× faster rollout generation and a 1.7× end-to-end RL speedup.
  • The upgraded time series encoder raises maximum supported input from about 240,000 to 300,000 time steps while running 5–6× faster at that length on roughly 20% of the previous GPU memory, and a new forecasting branch predicts future values numerically rather than as discrete text tokens; the separate Intern-MemDec-4B memory decoder lifts the Biology-Instructions average from 56.92 to 60.32 through token-level distribution fusion with the frozen 397B backbone.
  • Caveats are worth noting: this is a preview release whose headline claim is "competitive or leading results in multiple settings" rather than a clean benchmark sweep, the Memory Decoder is an add-on model rather than part of the evaluated 397B system, and the interleaved document pipeline covers only life sciences, chemistry, and materials science.

Applications 110

Methodologies for Improving the Quality of AI Tutoring in K-12 Education

Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano et al. cross-listed Khan Academy's Khanmigo tutor, launched in 2023, is built on large language models (LLMs) whose opacity makes offline reasoning about tutoring quality unreliable, so the team leans on measurement and live experimentation instead. The write-up describes the metrics used to track tutoring quality and student engagement, along with the experiments run against them in production. Model upgrades, prompting, personalization, and agentic behavior are named as the changes that actually moved those metrics.

Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models

Guobin Zhao, Xiao-Yan Li cross-listed Computation-ready metal-organic framework (MOF) databases underpin high-throughput screening, yet many deposited crystal structures are chemically unreasonable or disordered, and existing filters rely on heuristics, licensed software, or offer no explanation. The finding is that large language models can validate MOF structures once crystallographic information is rewritten as chemically meaningful text; across nine benchmarked descriptors, what determines success is not the amount of structural information but whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Models fine-tuned on the mof2text descriptor match graph-based baselines at flagging unreasonable structures and additionally produce diagnostic rationales naming likely error sources such as abnormal bonding, connectivity, or charge state.

Self-evolving network verifiers

Ioannis Protogeros, Tibor Schneider, Laurent Vanbever cross-listed Symbolic network verifiers can only check the protocols an expert has manually encoded, and keeping that model faithful is endless work because vendor router implementations drift from the published RFCs and change between releases. The proposed alternative closes a counterexample-guided loop in which a coding agent proposes extensions to the verifier's SMT encoding while a trusted oracle, such as emulated routers, supplies ground-truth routing state, and every disagreement drives another refinement. A prototype taught a 3,000-line verifier three previously unsupported features — OSPF areas, BGP route reflection, and L3VPN over EVPN — converging autonomously on oracle-matching models, including vendor-specific quirks, which shifts the open problem from writing verifiers to systematically testing automatically grown ones.

From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate

Pardis Taghavi, Santosh Bhavani Financial analysis mixes localized numerical computation with integrative judgment, and it is unclear whether both benefit from the same kind of model specialization. Larix maps a 16-lens framework for European listed real estate onto eight lens-aligned specialist agents and compares monolithic against decomposed prompting on the same frontier model, holding evidence, instructions, output schema, and scoring fixed across 19 firms and seven regulatory wrappers. Decomposition lifts numerical-task performance by 15.8 percentage points but does not reliably help and sometimes hurts judgment tasks, while post-training Qwen3.5-9B with GRPO and task-aligned structured rewards raises the judgment aggregate by 14.2 points with transfer to unseen firms (+15.2 overall, +40.4 on covenant stress) and unseen regulatory wrappers, suggesting prompt decomposition and parameter adaptation address different halves of the problem.

Making AI-Generated Feedback Matter: From Provision to Student Enactment

Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge, Dragan Ga\v{s}evi'c, Naomi Winstone, Hassan Khosravi Generative AI can supply individualized feedback to every student at once, but students rarely do anything with it, so the provision problem and the uptake problem are separate. A sequential-cohort quasi-experiment across 13,037 students and 51,296 submitted resources compared three workflows: comments delivered with no scaffolding, optional student-initiated dialogue, and an enacted condition in which students had to select feedback suggestions, judge their relevance, and then discuss those selections with the model. Uptake reached an estimated 26.2% under the enacted workflow versus 14.1% for plain delivery and 0.1% when dialogue was optional, and the enacted condition also produced higher self-assessment confidence and better submitted work, indicating that workflow design rather than comment quality drives whether AI feedback is used.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder

Huaxuan Wang, Huimin Wang, Ruiyu Zhang, Yingjie Li, Yitao Duan cross-listed Zero-shot text-to-speech systems usually need a transcript of the reference audio at inference time, which rules out cloning voices from the untranscribed in-the-wild recordings that dominate cross-lingual use. Confucius4-TTS removes that dependency with a two-stage design: a language-model-based text-to-semantic module that pulls timbre from self-supervised speech representations via a learnable speaker encoder, and a conditional flow-matching semantic-to-acoustic module that turns predicted semantic tokens into mel-spectrograms. Covering 14 languages, it reaches an average word error rate of 3.73% across six directions on the CV3-Eval cross-lingual benchmark and takes the best average human-evaluation rank on an internal cross-lingual set against recent open-source and commercial systems, with code and checkpoints released.

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

Ofir Ben Shoham, Shrutendra Harsola, Vignesh Subrahmaniam, Shravan Mohan, Yakov Gazman, Oded Vainas Turning business records into actionable financial advice resists supervised training, since historical decisions were not necessarily optimal and free-form expert labels are expensive, so the task is cast as reinforcement learning and an open-weight model is fine-tuned with Group Relative Policy Optimization against an LLM-as-a-judge rubric of binary advice-quality dimensions plus a harm-prevention safety gate. Because judge scores alone cannot show real business value, the authors add a judge-independent audit using a doubly-robust Conditional Average Treatment Effect estimator on observational data. The trained model shows roughly twice the estimated gross-profit lift of the strongest commercial baseline (0.0228 versus 0.0104), with the lowest downside rate and least negative tail risk of any policy evaluated. The two evaluations rank baselines differently — the untrained base model finishes last on the judge rubric but second on the causal audit — indicating the audit captures signal the judge misses.

Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs

Alireza A. Safaei, Laura M. Vowels, Matthew J. Vowels, Apoorv Jha, Shekoufeh Rahimi cross-listed Combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 model configurations quantifies what safety in mental-health deployments costs in energy, carbon emissions, water, and abiotic depletion. At the top of the safety distribution the trade-off turns sharply non-linear: a 2.61 percentage-point gain in clinical safety score came with roughly a 60-fold increase in estimated energy use per million output tokens. Row-level analysis further shows extra test-time compute did not consistently raise clinical safety and in some configurations coincided with lower scores, so the authors point toward dynamic model selection and cascading — reserving heavy models for higher-risk cases — instead of blanket scaling.

Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians

Timothy Heightman, Elena Orlova, Philip Mantrov, Aleksei Ustimenko cross-listed Finding ground states of quantum Hamiltonians is a headline target for quantum advantage, and Hamilton-Zero attempts to amortize the problem across an arbitrary family of them with a roughly 0.5-billion-parameter foundation model trained using techniques from large language models and deep reinforcement learning. Spin-1/2 ground-state learning is reformulated as variational optimization over centrally odd scalar functions on the manifold SU(2)^N, replacing explicit Hilbert-space amplitudes with manifold functions acted on by the Hamiltonian through Lie derivatives, with a Peter-Weyl argument showing the variational principle still upper-bounds the true ground state. Training uses an SU(2) replica-exchange Langevin sampler and sharded natural-gradient optimization via an extension of the Kronecker-Factored Approximate Curvature (KFAC) optimizer, over hundreds of thousands of Hamiltonians varying in topology, size, and interaction type. Pre-trained on systems up to 64 qubits, the model fine-tunes to 1024 qubits and is evaluated on systems as large as 8100 qubits.

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora et al. Reports that general-purpose frontier models now match specialized clinical AI rest on a narrow set of systems and benchmarks written mostly for high-income settings. VITA, a retrieval-augmented generation system drawing on curated disease guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols, was scored on 4,023 English HealthBench questions by a GPT-4.1 judge, where it placed first with 51.9% of possible rubric points ahead of GPT-5.4 at 46.1% and Gemini 3.1 Pro at 42.6%. Re-running a 500-question subset against current-generation models with a lineage-neutral open-weight judge (DeepSeek-V4-Pro) narrowed the gap to statistical parity with GPT-5.5 on mean per-question score, though VITA still led on points-weighted score, accuracy, and completeness while scoring lower on communication.

Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication

Xiaobin Shen, Chloe Y. H. Huang, Jonathan Elmer, George H. Chen Clinical prediction models assume the outcome of interest is observed for everyone, but treatment decisions such as withdrawal of life support can make the clinically relevant outcome permanently unobservable. Using a cohort of 2,497 post-cardiac-arrest patients — 1,429 of whom had outcomes rendered indeterminate this way and were reviewed by independent experts who guessed the counterfactual outcome — the authors propose evaluating models separately on certain and uncertain cases and train a model that trades off the two label sources. Models with nearly identical certain-case AUROC differed substantially in Brier score and in their predictions for uncertain cases, and improving agreement with expert counterfactual labels consistently cost accuracy on observed outcomes, an explicit tradeoff that conventional single-metric evaluation hides.

Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness

Haochen Zhang, Jiaheng Guo, Yu-Chao Huang, Nicholas Knoz, Tianlong Chen The most informative physiological signals in clinical monitoring are often invasive or simply missing for a given patient, and existing conditional generators are built around one conditioning modality so they degrade on the heterogeneous, irregularly missing mix of waveforms and static covariates found in practice. ReCoGen splits the job in two: stage one trains a masked autoencoder per modality to compress each time-varying condition into a missingness-tolerant token sequence, and stage two trains a flow-matching generator that fuses those tokens with static variables through cross-attention and an adaptive-layer-norm route to synthesize the target signal. Across continuous glucose monitoring on AI-READI and arterial blood pressure on MIMIC-III and MIMIC-IV, it gave the best downstream utility on all sixteen dataset-task-metric settings against six competing conditional generators, matching or exceeding the utility of the real signal on thirteen of them.

Evaluating AlphaEarth Foundations Embeddings for Wildfire Susceptibility Mapping

Yuan Zhuang, Sanaa Hobeichi, Peng Shi, Fei Huang cross-listed Wildfire susceptibility maps are normally built by harmonizing many remote-sensing, climate, and terrain products into hand-engineered physical variables. Using Victoria, Australia over 2017-2025, the authors test whether the analysis-ready geospatial embeddings from AlphaEarth Foundations can replace that pipeline, finding the embeddings reconstruct standard predictor variables accurately and that susceptibility models trained on them exceed 0.92 ROC-AUC while flagging eastern Victoria, Gippsland, and the north-eastern uplands as highest risk. The clearest advantage is transfer: applied to Canberra and Western Sydney-Blue Mountains without retraining, embedding-based models shift by roughly +4% and -2% in ROC-AUC, against a mean drop of about 25% for models built on physical variables.

The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning

Ahmed Sameh, Ramzi Al-Sharawi, Yogatheesan Varatharajah Self-supervised electrocardiogram (ECG) models typically train on a few seconds of signal and increasingly on discretized token sequences, and it has been unclear what those two choices cost. A controlled study on the single-lead Icentia11k dataset holds the Transformer backbone and training protocol fixed while varying input horizon across 16 seconds, 1 minute, 5 minutes, and 10 minutes, and comparing continuous convolutional patch embeddings against fixed vector-quantized tokens, evaluating by abnormal rhythm detection and patient-level retrieval across sessions. Longer context helps consistently, with the 5- and 10-minute models strongest, and continuous patch embeddings beat discretized tokens at every horizon, suggesting quantization discards clinically meaningful waveform detail and that ECG foundation models should favor extended context with continuous encoders.

A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings

Hei Ting (Una), Chan, Chenwei Wu, Xueshen Liu, Zesen Zhao, Boyuan Zheng et al. Specialist-level medical AI is hard to deploy in rural clinics where bandwidth is scarce, on-site compute is limited, and diagnosis requires combining several data modalities. The proposed cloud-edge architecture runs lightweight domain-specific models locally to compress raw medical data into compact structured outputs, sends only those to a cloud large language model that writes the clinical summary, and uses a language-model orchestrator to pick which diagnostic tools to invoke from patient context so irrelevant modalities are skipped. Across 20 multimodal cases spanning cardiac, obstetric, trauma, and screening scenarios under simulated networks from 500 kbps to 5 Mbps, the hybrid reaches 98-99% diagnostic tool recall at 92-96% precision, matches or beats cloud-only baselines on clinical accuracy, and holds latency at 25-35 seconds regardless of bandwidth with 4-15x lower token cost.

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto, David Rehkopf, Ayin Vala, Tanmoy Sarkar Pias Synthetic patient data is often proposed as a privacy-preserving stand-in for electronic health records, but published evaluations typically cover one generator, one dataset, or one narrow task, leaving unclear when the substitution actually holds. CoMedBench puts a family of generators through a shared clinical-validity framework and one common training and evaluation engine across 37 dataset-task pairs — 20 static tabular and 17 intensive-care time series — drawn from seven public sources including MIMIC-III, MIMIC-IV, eICU, NHANES, and the CDC BRFSS diabetes cohort. On tabular tasks the strongest generator, CoMed-TVAE, retains 97.3 percent of real-data area under the receiver operating characteristic curve (AUROC), while temporal intensive-care tasks prove harder and far more generator-sensitive: CoMed-CTGAN keeps 81.6 percent of AUROC but only 64.0 percent under the imbalance-sensitive precision-recall measure.

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

Hamza Shafiq, Hung Manh Pham, Bin Zhu, Pan Zhou, Jun Hu, Aaqib Saeed Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) each capture the same cardiac cycle from a different physical channel, yet cardiac foundation models are trained one sensor at a time. CardioState-JEPA maps all three waveform types into a shared token space, encodes them with a single Transformer, and pretrains by predicting masked latent cardiac states so the target is shared physiology rather than sensor-specific waveform shape; a learned delay aligner compensates for the temporal offsets between electrical, mechanical, and hemodynamic events, and training proceeds from abundant unimodal data to scarce synchronized recordings. Evaluated as a frozen encoder on 25 downstream tasks, it gains 8.2 AUROC points on PPG classification and 15.5 on ECG classification over the best self-supervised signal baseline, with the largest jump of 18.8 AUROC points on PCG murmur detection, and matches or beats models trained with privileged clinical text or supervised labels on several ECG benchmarks.

Learning Discrete Decisions for MIPs with Constraint-Aware Diffusion

Vincenzo Di Vito, Mehdi Taghizadeh, Deepjyoti Deka, Kaarthik Sundar, Ferdinando Fioretto Mixed-integer optimization problems are hard because discrete and continuous decisions must be chosen jointly under combinatorial constraints. Constrained Graph Diffusion (CGD) uses a graph-based generative diffusion model to produce only the discrete part of the solution, inserting a training-free feasibility projection operator into the reverse diffusion steps so intermediate samples are steered toward the feasible set; once integers are fixed, the remaining continuous problem is solved with standard numerical methods. On optimal transmission switching for AC optimal power flow and on discrete portfolio optimization, the method improves feasibility and solution quality over learning-based baselines while running up to 425 times faster than state-of-the-art solvers for mixed-integer nonlinear programs.

Foundation models for movement data: Are they ready for prime-time?

Alexander Br\"auer, Benjamin Cauchi, Nils Strodthoff cross-listed Foundation models pretrained on large accelerometer corpora are promoted as general-purpose feature extractors for health monitoring, but the case for them over ordinary supervised training has not been tested systematically. Four open-source accelerometer foundation models are compared against supervised baselines on 19 tasks spanning activity recognition, clinical monitoring, and physiological inference. Results are strongly task-dependent: supervised models stay competitive on human activity recognition, foundation models lead on fall and stress detection and tolerate sensor-placement changes better, and sleep staging is near chance for everything tested. Under linear and frozen probing only UniMTS beats the supervised baselines without finetuning, and concept-discovery analysis shows all models represent high-intensity activity well but blur sedentary or ambiguous behavior.

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo et al. Infrared spectra are widely collected for chemical sensing, but interpreting them still leans on reference libraries and expert knowledge, and machine-learning approaches tend to be tied to a single task, dataset, or instrument. UltraIR is a foundation model with over 100 million parameters pretrained on roughly 60 million simulated infrared spectra using spectral reconstruction, molecular-fingerprint similarity alignment, and functional-group prediction, then finetuned on real measurements for downstream objectives. Across functional-group and structure elucidation, property prediction, mixture identification and quantification, bacterial classification, herb origin tracing, microplastics classification, and soil analysis it beats conventional and task-specific deep-learning baselines, works with few labeled experimental spectra, and transfers zero-shot across spectrometers and laboratories for the same task.

Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion

Van Khoa Nguyen, Alexandros Kalousis Generative models for crystals typically stop short of a full crystallographic specification, producing structures only up to site symmetry and drawing space groups from an empirical distribution rather than modeling them. Taking a cue from spontaneous symmetry breaking in physics, SbCD (Symmetry-breaking Crystal Diffusion) runs a Markovian jump-diffusion process that starts from the lowest-symmetry prior and reverses it, letting the generative trajectory jump between space groups so inter-space-group transitions become an explicit part of the model. On de novo generation with MP20 and MPTS-52, the approach beats its symmetry-preserving counterpart by a substantial margin.

Equivariant learning of a transferable three-dimensional classical density functional

Bingqing Cheng cross-listed Classical density functional theory would let one predict liquid behavior across conditions without rerunning an atomistic simulation for every state, except that its excess free-energy functional is unknown and learned versions have mostly been limited to planar or lower-dimensional geometries. Training on fully three-dimensional equilibrium density fields while preserving spatial symmetry and variational consistency yields such a functional without any free-energy or chemical-potential labels, and a single learned functional transfers across temperatures, system sizes, and statistical ensembles. It reproduces structure factors, the equation of state, liquid--vapor coexistence, and interfacial broadening despite none being training targets, and predicts the non-monotonic force from a solvent-depleted bridge forming and rupturing between colloids as well as adsorption in a gyroid pore.
88 more specialized papers

Large Language Models 57

Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts

Parvel Gu Top-k mixture-of-experts (MoE) routing is discontinuous, so numerical noise from deployment tricks like 4-bit key-value cache quantization can push tokens across decision boundaries and change which experts fire. A four-run causal apparatus prices the route-mediated fraction of quantization damage, decomposes it token by token, and carries pre-registered probes across three architectures. On OLMoE-1B-7B roughly a third of the damage is routing-mediated (fraction around 0.31), and while the router margin detects that a flip occurred with an area under the curve of 0.772, no inference-observable router statistic predicts whether a given flip helps or hurts above chance, which bounds any selective-repair scheme built on that feature family. A real int4 kernel gives a compatible but underpowered estimate, ruling out gross disagreement rather than serving as independent replication.

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

Nikita Kulin, Viktor Zhuravlev, Artur Khairullin, Sergey Muravyov, Ilya Makarov, Daniil Sukhorukov et al. Automatic prompt optimization usually rewrites a prompt as one block, which can improve one behavior while quietly breaking another. SAPO instead decomposes a prompt into role, context, tasks, and output format, diagnoses each segment against the five best and five worst examples, and synthesizes candidates constrained by which segments look weak or strong, using a single language model with static meta-prompts and structured outputs throughout. Evaluated on SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K with GPT-3.5-Turbo and GPT-4o-mini, the segment-level method achieves the best average score against zero-shot prompting and the APE, OPRO, EvoPrompt, GEPA, and StraGO baselines.

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin et al. Position-independent caching speeds up language model serving by reusing key-value entries for text chunks regardless of where they appear, but every existing method assumes a token-indexed cache it can match, concatenate, and locally repair. Hybrid models break all three primitives, since most attention layers are replaced by linear recurrences that expose only a fixed-size state. LinearKV is a training-free framework built on a decoupled initialization: each linear layer maps its matched local states to a single initial state while full-attention layers concatenate as usual, letting existing token selection and recomputation schemes be reused unchanged. The surprising finding is that one cached state suffices and the algebraically exact alternative of composing all cached states, as concurrent work does, is unnecessary and sometimes harmful — on a Mamba-2 model exact composition collapses to 46.6% of full quality under the EPIC selector versus 86.8% for the single-state initializer, while also cutting time-to-first-token to 0.46 times full prefill.

CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference

Yifan Wu, Yufeng Zhang, Kenli Li Diffusion language models (DLMs) resolve many tokens per step, but blockwise decoders keep issuing dense forward passes until every position is settled, even though many predictions stabilize early. CORA-Diff is a training-free method that leaves the model's original transfer rule untouched and applies confidence-and-persistence gating only to the positions that rule leaves unresolved, with a single operating point calibrated on a held-out GSM8K subset and frozen for all evaluations. Under a matched LLaDA protocol it has the lowest runtime in all eight task-length settings, with incremental speedups of 2.70x on GSM8K and 3.32x on HumanEval over EOS-aware dense decoding and a largest quality drop of 1.22 points, and it transfers to Dream without retuning.

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

Jeonghwan Choi, Taewon Yun, Minjeong Ban, Gyeonghun Sun, Jae-Gil Lee, Hwanjun Song Existing retrieval-augmented generation evaluation struggles to give consistent, fine-grained diagnostics across query types ranging from close-ended fact lookup to open-ended explanation. Q-CARE is a reference-free framework that decomposes queries into sub-queries and answers into atomic claims, then scores retrieval by how well retrieved evidence covers the sub-queries (C-Prec@k, C-nDCG@k) and generation by claim completeness, conciseness, and verifiableness. On a human-annotated benchmark spanning eight datasets it correlates with human judgment better than four existing metrics, including RAGEval and RAGChecker.

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

Zafar Hussain, Kristoffer Nielbo In collaborative multi-model settings each agent sees peer assertions before answering, so a wrong majority can overturn an answer the model would otherwise get right. Measuring 23 open-weight models across 19 conditions and three datasets, over a million graded responses, a unanimous wrong majority reverses 22.8% of correct MMLU answers, 54.8% on GPQA, and 71.0% on SimpleQA, with 84-89% of reversals matching the peers' answer. Existing mitigations target only Resistance, the rate of holding a correct answer under pressure, so the work pairs it with Receptivity, the rate of adopting a correct peer answer after being wrong; six methods all trade one for the other and fall on a single frontier with R^2 between 0.80 and 0.90, with Reflection gaining 7.9 Resistance points and losing 15.3 of Receptivity. Reasoning is the sole exception, raising Resistance by 7.2 points and Receptivity by 9.6 on MMLU subjects a model can derive for itself.

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

Angelo Nardone, Paolo Ferragina Neural compressors built on large language models achieve far better lossless compression ratios on text and code than zstd, gzip, or bzip, but their one-symbol-per-step autoregressive decoding makes throughput impractical. Swapping in diffusion language models removes that bottleneck since each forward pass can encode a variable number of symbols at freely chosen positions, though this breaks the assumptions of the standard compression pipeline and requires new algorithms for deciding what gets encoded when. Evaluated on the enwik8 benchmark against both language-model-based and general-purpose compressors, the diffusion framework advances the state of the art in lossless text compression.

Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

Bohan Zhang, Anqi Ni, Yixin Wang, Paramveer S. Dhillon cross-listed Supervised fine-tuning is the default way to adapt a language model to a target distribution, but personalization makes it impractical because every author needs separate weight access, optimization, storage, and retraining. Weightless Fine-Tuning (WFT) skips weight updates entirely: it computes supervised residuals on an author's training sequence and moves them to the current prompt through a cross-prefix transport operator estimated from dropout-induced cross-covariance, turning gradient updates into decoding-time logit corrections. Across three LaMP personalization benchmarks it posts the best average score, matching or beating supervised fine-tuning on individual tasks while using under 7% of the effective computation in a budget-controlled comparison, and its induced logit shifts show 0.875 cosine similarity with those of actual fine-tuning over 95% of the next-token probability mass.

Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter

Rima Mittal, Ankit Gubrani, Satyanarayana Kakollu cross-listed Tokenizer vocabulary size is normally fixed at training time by convention, without reference to how the model will be served. Framing total deployment cost as training cost plus inference volume times per-inference cost turns vocabulary into a systems parameter that depends on the serving regime, and controlled experiments on A10G and A100 GPUs — spanning memory-bound to compute-bound operating points — show the inference-optimal vocabulary shifting 16-fold with batch size, from 32k tokens at batch 1 to 524k at batch 64 and above, driven by amortizing the read of the unembedding matrix. Quality measured in bits per byte peaks near 65k at the 1.3-2.3B parameter scale but varies less than 2% across the optimal range, so the recommendation is roughly 32k for on-device serving and 131k-262k for datacenter deployment.

Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models

Alexandrine Fortier, Hazel Chen, Peter West The sameness of language model outputs is generally blamed on alignment, but where in the pipeline it originates has not been pinned down. Tracking semantic convergence backward shows it already present at the first alignment stage, supervised fine-tuning, which prompts controlled experiments on how training data shapes convergence for specific input/output pairs. Fine-tuning data can reveal and amplify convergence but does not introduce it, and instruct-like collapse can be induced in base models through prompting alone with no alignment at all — suggesting homogeneity emerges from the pretraining objective itself and will resist fixes applied only after alignment.

Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task

Xiaoyang Hu, Mike Angstadt, Shane Storks, Zan Huang, Aman Taxali, Alex Weigard et al. cross-listed Congruency effects — slower, worse responses when an automatic tendency conflicts with an instructed rule, as in the Stroop and flanker tasks — have been studied for a century without a settled mechanistic account. A purely verbal analogue is built for language models: a prompt stem elicits a default same-color completion while an explicit rule either agrees with it or contradicts it, tested on Gemma-2-2B and six Pythia models from 410M to 12B parameters. Six of the seven models show strong congruency effects, and causal attribution, attention analysis, and attention ablations reveal two distinct pathways — short-range attention to a superficial color cue that dominates congruent trials, and long-range attention to the rule prefix that dominates incongruent ones. Fine-tuning that strengthens the default mapping helps congruent and hurts incongruent trials, while enlarging the rule set selectively hurts incongruent trials, both consistent with competition between an in-weight default and an in-context rule.

Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation

Alex Deaconu, Anubhav Gupta, Manaal Basha, Nicholas Haydu, Gema Rodr\'iguez-P\'erez cross-listed Prompt wording is known to move model behavior, but the effect of framings borrowed from the psychology of human persuasion has not been measured for coding work. Eight influence tactics from Yukl and Falbe's taxonomy — including rational persuasion, ingratiation, and exchange — were turned into reproducible prompt templates and run through five open-weight models on LiveCodeBench and SWE-bench Verified, with outputs scored for functional correctness, code quality, maintainability, and security. Framings that emphasize urgency were associated with lower correctness and weaker security, the first large-scale evidence that persuasion-style linguistic cues systematically shape generated code.

EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

Huiqi Miao, Xinbao Sun, Bo Wang, Fanyu Meng, Lijun Mei, Na Wu et al. Retrieval-augmented generation (RAG) benchmarks generally assume clean retrieved documents and simple queries, which hides how systems behave in production where noisy passages, missing knowledge, and contradictory sources meet multi-part instructions. EnterpriseRAG supplies 983 expert-validated samples across six domains that deliberately inject retrieval noise, knowledge gaps, and factual conflicts alongside complex constraints. Testing 13 current models exposes an orchestration gap: models satisfy roughly 80% of individual constraints but only 26.8% of responses meet every requirement at once, a 57-point shortfall that persists under reasoning-enhanced inference and is worst when knowledge is absent or sources disagree.

Dion3: Full-Stack Orthogonal Updates

Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen et al. cross-listed The Muon optimizer's cubic-time Newton-Schulz orthogonalization is expensive, and when weights are sharded across devices the added communication can wipe out its advantage. Dion3 attacks the cost at every layer: a Gram-based Newton-Schulz variant lowers the floating-point operation count, hand-written CuteDSL kernels exploit matrix symmetry, megabatching cuts communication, and the update rule orthogonalizes only a fraction of the momentum matrix's rows each step. The result matches or beats Muon's training loss while reducing optimizer step time by up to 6x, improving on the earlier compressed Dion in both speed and quality, and ships as a drop-in replacement in the dion package.

CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

Kuangzhao Yang, Ziliang Zhao, Zhicheng Dou Ambiguous or underspecified user queries push large language models into vague or wrong answers, but teaching a model when to ask a clarifying question and what to ask about has generally required human-annotated preference data. CLAIM derives that signal automatically instead, quantifying query ambiguity as the entropy of disagreement among answers sampled from multiple models, then combining that estimate with semantic clustering and reasoning-based judgments to synthesise training data. A single clarification-decision model is trained on this data with supervised fine-tuning followed by group-relative policy optimization (GRPO), and the authors report that it learns stable, generalizable clarification strategies with no manually labeled data at all.

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang et al. Knowledge editing keeps a language model current by rewriting specific facts, and recent work has moved from triples to unstructured edits where the update is a free-form passage stating several facts at once. Existing editors get the passage in but not usable: the model can recite it yet cannot answer atomic questions about its individual facts or chain them into multi-hop reasoning, a shortfall the authors call a lack of composability and attribute to treating the fixed passage as the only learning signal. HPSE instead distills from the same model's own privileged in-context state, and because a pre-edit model's rollouts rarely cover genuinely new knowledge, it builds a hybrid rollout that inserts missing facts onto the student's own trajectory exactly where coverage fails while staying on-policy elsewhere — giving plug-and-play gains across four model backbones and two existing editors, with theory explaining the advantage over pure on-policy distillation.

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

Alish Kanani, Layan Badawi, Umit Y. Ogras cross-listed Mixture-of-Experts models suit edge devices because only a fraction of parameters activate per token, but expert weights usually sit in off-chip memory, so loading them lands on the critical path and memory rather than compute becomes the limit. APEX hides that latency with a lightweight prefetch router that predicts likely experts before the attention block runs and uses a learned confidence model to decide how many extra experts to fetch, reaching over 99% overlap accuracy versus fixed top-k prefetching. It offers a correctness-preserving mode that keeps exact routing semantics — cutting per-token latency by up to 26% and improving energy-delay product by up to 41% over prior work — and a stall-free mode that computes with whatever experts are already available for further gains at negligible accuracy cost.

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel Every benchmark score rests on one particular phrasing of each problem, treated as a stand-in for all the ways that problem could be asked. BenchDrift generates meaning-preserving rewordings along linguistic, referential, pragmatic, and structural axes and measures how often correctness flips in each direction, finding substantial two-way drift across eight models and GSM8K, MMLU, and MATH-Hard. Two results stand out: sensitivity to phrasing does not shrink as models improve but reverses sign, with weak models gaining more from rewording than they lose while strong models lose far more than they gain, so the top scorers on a benchmark are the ones most dependent on the exact wording they were handed. Models also largely agree on which rewordings are costly, indicating the fragility belongs to the rephrasing rather than the model, and rewordings break answers the model had been confident about whether they make the problem shorter or longer.

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

Yushi Ye, Xu Chen, Haoyun Jiang, Jinsong Lan, Haihong Tang, Bo Han et al. Diffusion large language models can in principle decode many positions at once, but existing schedulers commit a position only when it independently passes a confidence threshold, ignoring how one commitment changes the rest. The identified ripple effect is that deliberately committing a mid-entropy pivot position sharply lowers uncertainty across the remaining masked positions, letting later steps unmask more tokens in parallel. Ripple-Pivot Search is a training-free decoder that searches for such pivots and uses lookahead evaluation to choose which token to place there for maximum downstream benefit. Across 3 diffusion language models and 4 reasoning and code benchmarks it delivers 4-10x wall-clock speedup over standard decoding at comparable generation quality, rising to 18x with key-value caching, and improves accuracy over the prior lookahead baseline by up to 5.49%.

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization

Ellen Su, Andres Potapczynski, Shikai Qiu, Edward Hughes, Andrew Gordon Wilson cross-listed Which training data helps a model transfer to tasks nobody specified during training? The working hypothesis is that structurally rich data contains reusable circuits and subprograms, and epiplexity, a recent measure of how much structural information a compute-bounded learner can extract from data, makes that hypothesis testable as a live training signal. For data selection, scaling laws fit to per-domain training loss curves predict epiplexity gain as a function of tokens and set adaptive domain sampling weights during training; for synthetic data, a generator is rewarded by the change in learner epiplexity over a buffer of its past outputs and trained with REINFORCE policy gradients toward an epiplexity-maximizing distribution. In both settings higher epiplexity predicts better downstream zero-shot and fine-tuned performance, supporting the structural-information account of transfer.

Language-Conditional Dequantization: Recovering What Quantization Steals from Non-English Languages

Nirmal Thomas Aggressive quantization hits multilingual ability unevenly: under sub-4B INT3 GPTQ, perplexity degrades two to four times more on non-English languages than on English. Language-Conditional Dequantization is a post-hoc fix that attaches per-language rank-2 LoRA corrections to the linear layers of an already-quantized model, costing 0.12% extra parameters per language and under 20 minutes of single-GPU training. On Qwen2.5-3B and Llama-3.2-3B it recovers 70-83% of the perplexity gap for non-Latin-script languages plus 17-28% of the GlobalMMLU accuracy gap, beating an equal-capacity language-agnostic correction by 3-9 points on typologically distant languages and the data-free LQER baseline by an order of magnitude. A layer-restricted variant supports their account of a perplexity-accuracy disconnect: early-depth quantization errors propagate downstream and resist local correction, while late-depth errors do not.

Hybrid Gated Attention

Zekun Zhou, Ruobing Xie, Lanrui Wang, Weixuan Sun Gated attention mitigates attention sinks and adds representational capacity, and HyGA pushes its effectiveness-efficiency frontier further by combining three gating strategies that draw on information from different stages of attention, building element-wise and head-wise gates that capture both intra-head and cross-head interactions. Low-rank matrix decomposition and a learnable attention sink are added for training efficiency and stability. Across several backbones and standard benchmarks, HyGA improves both training loss and downstream task performance relative to gated attention, and holds the best performance at each computation budget examined.

LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Xuhui Jiang, Chengjin Xu et al. Training large models on constrained hardware turns into a scheduling problem spanning GPU compute, host memory, PCIe bandwidth, and storage, and existing offloading systems leave communication exposed on the critical path because checkpointing and placement follow fixed heuristics. LazyTrain sits atop a layer-streaming executor and formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a single mixed-integer scheduling problem whose solution is then executed during training, paired with a hybrid operator that combines 8-bit optimizer states with fast gradient clipping to offset the added CPU-side update cost. Across Qwen2.5-3B through Qwen3.6-27B on H800 hardware it raises sustained TFLOPS about 1.24x over matched baselines, with the 27B MetaMathQA run hitting 219.95 TFLOPS and 1361 tokens per second at 68.84 GB peak GPU memory and 95.42% exact-match accuracy.

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

Karl Hanna, Chen Feng Multiple-choice scores mix a model's actual knowledge with its sensitivity to the order in which options are presented, so the paper tests whether hiding option labels at the moment of commitment removes positional influence and thereby improves accuracy. Two mitigation strategies are compared — generating an answer then matching it to options, and scoring each option in isolation, which is positionally unbiased by construction — along with a full decomposition of the pipeline. Neither strategy reliably improves accuracy, and the bottleneck turns out to be withholding the options rather than the matching step; the only configuration that consistently matches the baseline shows all options and uses a model-based matcher, while cyclic permutation often does improve scores. For two-stage prompting, both an aggregate recall-imbalance measure and a per-question order-sensitivity measure fail to demonstrate reliable debiasing.

RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks

Jinjun Huang, Zhongzhen Wen, Tongtong Xu, Meng Yan, Xin Xia, Zhongxin Liu cross-listed Existing benchmarks for language-model-generated Triton GPU kernels test only PyTorch-to-Triton translation, measure isolated kernel speed rather than end-to-end performance, and rely on hand-written correctness checks that models can exploit for inflated scores. RealisticTritonBench instead derives tasks from real pull requests that modified Triton kernels in popular open-source AI frameworks, pairing a natural-language requirement with concrete engineering context and a reproducible environment where the generated kernel is integrated back into its framework and judged by end-to-end tests. Leading language models still perform poorly on these production-style kernel generation tasks.

Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

Xi Chen, Jie Mu, Mo Xuan, Qun Shao Rubrics handed to language-model judges are typically flattened into prompt context or a list of criteria, leaving how those criteria compose implicit even when the written rules state it. Graph-Structured Rubrics compiles a rubric into a response-independent typed evaluation graph before any answer is seen: criterion nodes elicit judgments, transformation, reduction and gating operators combine them through named ports, and a Readout converts the unique sink into a score or preference, with malformed or type-incompatible graphs rejected at compile time. Under GPT-OSS-120B, exact score agreement improves by 0.62 to 6.75 percentage points over Prometheus-style scoring across four pointwise datasets, with the numerically highest end-to-end pairwise accuracy on two preference benchmarks under native tie and abstention policies.

QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving

Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao, Libin Zheng et al. Retrieval-Augmented Generation (RAG) systems repeatedly prefill the same retrieved text chunks, and Position-Independent Caching (PIC), which reuses precomputed Key-Value (KV) state regardless of position, is still bottlenecked by the sheer number of text tokens. Rendering chunks as images compresses them into far fewer visual tokens but degrades answer quality more sharply, because independently compiled caches lose shared context and visual compression discards fine-grained text. QV-PIC compiles visual caches offline under the model's native chat-template prefix and, at serving time, keeps global context at low resolution while restoring query-relevant regions at high resolution under a budget set by cumulative relevance scores. Across six tasks it gains 21.6 F1 points over vanilla rendered-image PIC, exceeds optimized text PIC by 2.58 F1 while cutting time-to-first-token (TTFT) by 17.2%, and reduces TTFT by 83.8% versus full prefill.

SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges

Yuchao Wu, Junqin Li, XingCheng Liang, Yongjie Chen, Yinghao Liang, Linyuan Mo et al. Dense retrieval handles structured constraints and multi-hop questions poorly, while graph-based retrieval fixes this at the cost of building and maintaining a global knowledge graph that fragments meaning into triples. SAG instead indexes each chunk as a semantically complete event paired with its entities, forming a latent hyperedge that keeps n-ary relations intact, and at query time joins chunks on shared entities to assemble a query-scoped neighborhood whose evidence units are still the original chunks. On HotpotQA, 2WikiMultiHopQA, and MuSiQue it leads on both retrieval and end-to-end question answering, with the margin widening as reasoning chains lengthen — reaching 80.36% Recall@5 on MuSiQue, 11.52 points above the strongest baseline.

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong et al. Massive activations — the extreme hidden-state outliers known from full-attention transformers — are characterized for the first time in hybrid linear attention (HLA) models that interleave linear and full attention layers, revealing two architecture-aligned shapes: spikes that appear immediately before full attention layers, and plateaus that persist across the linear attention layers in between. As full attention becomes denser, consecutive spikes link up through the plateaus and recover the stable pattern seen in full-attention models, a recurrence confirmed across five linear attention architectures, six hybridization settings, five data domains, and open models from 1.2B to 397B parameters. Controlled pretraining of gated DeltaNet hybrids up to 1.3B shows both shapes emerge early and respond asymmetrically to output gating, and an outlier lifecycle analysis attributes spikes to a localized write-sink-cancel process and plateaus to delayed cancellation.

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

Rodrigo Guedes de Souza, Alison R. Panisson Standard leaderboards implicitly assume a model's rank is stable regardless of how many tokens it is allowed to generate. Sweeping the generation budget across seven levels from 64 to 4,096 tokens for four models on three reasoning benchmarks — 56,476 inferences in total — shows 3-19% of items where accuracy actually falls as budget grows even after controlling for truncation, and model rankings reverse across budgets on every benchmark. Oracle analysis finds up to 27.8 percentage points of complementarity between models, largest at tight budgets, while a budget-aware router recovers 14.1% of that gap cross-domain and budget features help within a domain but transfer poorly.

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi Training with ever-longer contexts is usually assumed to be harmless or helpful, but abundant in-context evidence may remove the incentive to store that information in weights, a tension named the Information Abundance Paradox. Pretraining experiments with long documents show that widening the context window improves language modeling, natural language understanding, and closed-book multiple-choice question answering only up to an intermediate optimum, after which all three consistently decline; in supervised fine-tuning, more task-relevant context at train time helps when supporting context is present at test time but reduces robustness when it is absent or misleading. Mechanistically, informative context shifts gradient pressure away from feed-forward layers, which are associated with parametric knowledge, toward attention, and causal interventions confirm this shift increases reliance on context at inference — suggesting that scaling toward near-infinite context is not simply a matter of supplying more high-quality long data.

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji et al. cross-listed Distillation usually transfers a large model's capability into a smaller one by updating the smaller model's weights; the question here is whether the same transfer can happen purely at inference time. In strong-to-weak scaffolding, a stronger builder model uses 5% of the data as a validation set to iteratively write and refine an inference-time harness for a weaker target model, which is then evaluated frozen on the full test set. Across four Theory-of-Mind benchmarks the approach lifts average target performance from 0.49 to 0.91, and analysis attributes the gain to moving unstable reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement rather than to more reasoning or broader sampling; harness quality rises monotonically with the builder's reasoning effort, and the weakest targets benefit most.

Can Spectral-Clipping Enable Better Learning While Forgetting Less for Low-Rank Adaptation?

Hyowon Wi, Noseong Park cross-listed Examines what happens to the singular value decomposition (SVD) structure of pretrained weights during low-rank adaptation (LoRA), finding that large-singular-value components can largely be reused across tasks while small ones are task-specific, and establishing a theoretical link between unbounded growth of adapter singular values and catastrophic forgetting of pretrained knowledge. The proposed SCLoRA injects parameterized singular components with spectral clipping, sized according to the pretrained model's own spectral distribution, so updates concentrate on the components that actually need adapting. Experiments report both better downstream accuracy and better retention of pretrained capabilities than standard low-rank adaptation.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

Qiuwu Chen, Zimo Liu, Yuchen Li, Ying Sun, Yifan Zhang, Zhijie Qiu et al. Attributes two sources of pretraining inefficiency in transformer language models: self-attention has no built-in locality bias and so redundantly rediscovers local structure, and mixture-of-experts layers tie knowledge storage to compute routing, making stored knowledge hard to access directly. LoKiFormer adds two modules to a standard decoder — Local Fusion Attention, which folds a convolutional fusion into attention so it operates on locally enriched representations, and a Knowledge Memory Module, a parametric key-value store with addressable slots that separates knowledge storage from the compute path. The architecture is reported to converge 1.33x faster during pretraining than baseline models.

MARCH: Scaling Recurrent Memory with Content-Routed State Anchors

Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding, Xia Hu et al. Transformers retrieve well over long contexts because their token-level memory grows with the input, at the cost of quadratic training compute and a linearly growing key-value cache, while recurrent models decode cheaply from a fixed-size state but overwrite older associations. MARCH (Memory-Anchor Routing across Context History) periodically checkpoints the cumulative recurrent state as an "anchor" tagged with a compact content-conditioned key, so the memory bank grows with context length and trades historical resolution against memory cost; each token issues an anchor query and aggregates over all causally available anchors in attention style. After standard pretraining it outperforms several linear-attention variants on commonsense reasoning, LongBench, and in-context retrieval while keeping the recurrent computation path.

Geometric and Behavioral Stratification in Transformer Residual Streams

Nelson Guda Transformer residual streams are known to develop privileged coordinate axes, but it has been unclear what kind of direction those axes single out. Measuring variation relative to the prediction direction — the unembedding direction of the token the model is currently predicting — reveals that residual-stream structure is stratified by proximity to that anchor: a narrow, scale-invariant interface concentrates readout-relevant structure while the far larger prediction-distal complement grows with model size and is nearly orthogonal to the principal variance axes, so standard variance-based analysis only partly recovers the organization. The stratification held in all eighteen models tested, spanning dense and mixture-of-experts architectures from 7B to 120B parameters in both base and instruction-tuned form, and perturbation experiments show the proximal slice is causally decisive (disruption causes immediate divergence and task-frame shifts) while the distal complement remains load-bearing over longer horizons, with behavior driven by direction rather than magnitude.

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

Haoyuan Zhu cross-listed In multi-turn dialogue a user can withdraw a constraint as easily as impose it, but models often keep obeying withdrawn requirements — sometimes emitting comments that assert the requirement was removed while still enacting it — a failure the authors call behavioral relapse or revocation inertia. The proposed system works entirely through the model API: a contract ledger pairs each constraint with an executable checker and records revocations as tombstones, compiling the net constraint state into a single specification ahead of time; a sequential ablation probe measures each clause's incremental behavioral effect; and a repair ladder applies fixes under token- and attempt-matched budgets. On a benchmark built from HumanEval tasks with verified checkers, relapse rose with constraint load at a small operating point while stronger models stayed near zero, and ahead-of-time compilation of the ledger significantly reduced relapse against a verifier-retry baseline under matched budgets, while adaptive ladder interventions on top added no detectable gain; the probe also predicted relapse before delivery, and a one-sentence tombstone note recovered roughly a third of the compilation benefit.

Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection

Florian Braun cross-listed Detecting whether a benchmark leaked into a model's training data normally requires the training corpus, a well-chosen likelihood statistic, or canary strings planted before release, none of which an outside auditor usually has. Reading contamination off a linear probe on internal activations is a proposed alternative, and the authors show the obvious version of it fails, then specify a protocol that survives measurement: a zero-sum contrast on the depth profile of probe accuracy, recentered against a level-matched placebo baseline, tested against a label-permutation null with a reference set twice the size of the suspect set. Each design choice is justified by a measured failure of the simpler option — comparing against a flat depth profile falsely rejects a true null 72% of the time when surface decodability rises with depth and loses all power when it falls, and a half-size baseline triples the error rate. On real transformers all four well-matched Pile arms return null and the protocol declines to issue a verdict on the temporal split rather than reporting a false one.

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family

Rishi Shah, Rishav Shrestha Reported correctness rates for language-model-generated GPU kernels rest on a single loose check — run on a few random inputs at one shape and accept if the output is close to a reference — which a kernel can pass while returning a finite number where the true answer is NaN or infinity, varying run to run, breaking at other shapes, or accumulating in fp16 where the reference uses fp32. The authors build a contract-grade verifier of twelve adversarial gates, several of them tolerance-free so no threshold choice can excuse a failure, and point it at 2,638 machine-generated kernels that a public system's own harness had already marked correct. It finds 39.5% broken beyond any tolerance argument and 62.1% carrying at least one violation, with the standard test accepting 1,487 kernels the verifier rejects against only 14 in the other direction, defended by a positive control, a threshold sweep, 98.5% agreement with the benchmark's own correctness code, and a hand audit. Turned inward, the verifier validates the authors' own first native Blackwell tcgen05 training backward pass for the gated-linear-recurrence family against a double-precision oracle.

Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

Xiang Guan, Roger D. Newman-Norlund, Yong Yang, Saeed Ahmadi, Regan Willis, Nadra Salman et al. Claims that particular transformer components specialize for particular cognitive operations are hard to falsify without a spatially resolved test, so the authors port subtraction analysis from human neuroimaging to perturbed language models and run the identical logic on both substrates. PRISM perturbs layers of LLaVA-1.6-Vicuna-13B, maps errors onto the seven clinical Philadelphia Naming Test categories, subtracts error classes pairwise, and treats each perturbation seed as a subject in a group analysis with threshold-free cluster enhancement along the layer axis; a structurally matched lesion-symptom mapping is run on 213 chronic post-stroke aphasia patients, with both sides replicated on held-out splits. Both substrates recover a phonemic-favoring dissociation — a deep-layer cluster in the model and a frontal-perisylvian cortical cluster in patients — while the semantic-favoring direction is consistently signed but not significant in either. The authors note the confirmatory region-of-interest intervention that would license a causal-mechanism claim is left for future work.

ReconSpan: Reconstruction-Guided Adaptive Latent Tokenization

Lixing Li cross-listed Latent tokenization compresses fine-grained text into a shorter sequence of continuous vectors, but where to place the chunk boundaries is usually decided heuristically. ReconSpan forms chunks by testing whether a backward decoder can reconstruct them from a single contextual prefix code, then keeps that code as the chunk's latent token, so one trained autoencoder can produce average chunk lengths anywhere from 6.5 to 12.2. At matched average chunk length, reconstruction-guided boundaries preserve more of the original text than random boundaries, though models reading the resulting latent sequence recover topic-level information reliably while struggling to extract exact details.

Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty cross-listed Structured pruning of large language models typically proceeds by greedy heuristics that make locally myopic choices and often miss the requested compression ratio. SNIPER instead casts the depth-versus-width allocation as a binary knapsack problem over coarse components, yielding an allocation that is optimal conditional on fixed importance estimates, then applies a fine-grained pruning pass to land exactly on the budget. The authors define the Compression Ratio Adherence Factor (CRAFT) to measure budget fidelity and report that existing pruners miss their target ratio by as much as 33%, while SNIPER scores 0.98 on CRAFT, close to exact adherence. Across four architectures and 18 tasks spanning five domains it attains a mean rank of 1.25 against six state-of-the-art pruners.

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

Zixuan Lan, Yanhong Li, Jiawei Zhou Transformer inference cost is dominated by repeated high-dimensional matrix products, and RMM (Reduced Matrix Multiplication) cuts them without touching weights or retraining by adaptively selecting informative slices along each product's contraction dimension, with a single retention ratio controlling the accuracy-efficiency trade-off. Testing on language models from 1 billion to 70 billion parameters shows how much reduction is tolerable depends on model family, task, and component, and often improves with scale, with moderate reduction holding up on discriminative, generative, and long-context tasks as well as vision-language inference. Ablations expose a structural asymmetry: attention-side computations are far more reducible than the MLP blocks, and custom kernels on an NVIDIA A100 confirm the savings become real wall-clock gains, especially at long sequence lengths.

DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

Tianyi Li, Yaxin Luo, Xinyi Shang, Zhiqiang Shen Speculative decoding speeds up autoregressive generation by proposing draft tokens and verifying them in parallel, and diffusion drafters make proposals cheap by emitting a whole block at once, but their per-position distributions are marginal rather than conditioned on the tokens chosen along a given draft path. DARTree is a training-free method that extends a pretrained autoregressive correction head from single chains to trees: it builds a fixed-width candidate tree by expanding and scoring every node at each depth in one batch, then prunes best-first to pick the verification tree, which separates correction-head inference from sequential heap operations. Across seven math, code, and chat benchmarks it leads on average acceptance length and speedup in all four model-temperature settings, accepting up to 12.97 tokens per verification round and reaching up to 9.73x lossless speedup over measured autoregressive decoding.

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd\"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell et al. cross-listed Studying how language models acquire knowledge is confounded by web-scale pretraining data, where prior exposure to any test concept is nearly impossible to rule out. LittleCurriculum is an 88B-token corpus restricted to U.S. elementary school material with concepts, facts, and vocabulary above Grade 5 deliberately excluded, and LittleLearner is a 5B-parameter model trained on it from scratch, fluent enough for open-ended evaluation but with capability boundaries that map onto interpretable curriculum guidelines. Using this sandbox, experiments on injecting new knowledge through post-training and in-context learning find that both help the model exploit what it already knows but do not lift out-of-scope capabilities. Both corpus and model are released.
12 more specialized papers

Agents 49

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

Igor Itkin Simulating societies of many language model agents is expensive, yet the questions usually asked are macroscopic — phase behavior, stylized facts, and how outcomes scale with the number of agents — rather than anything about individual cognition. The proposed shortcut replaces each agent with a low-parameter model fitted from a few hundred to a few thousand cheap queries, and an [interaction order x memory] taxonomy predicts in advance, from what each agent perceives, whether this substitution will work and how the surrogate error trends with population size. Validation on a faithful reimplementation of the EconAgent macroeconomy plus seven other named simulations, with decisions cloned from genuine elicitations (mostly DeepSeek), shows the predicted error trends holding cell by cell for a few dollars of query cost, with societies then runnable at any size on a laptop; the two refuted predictions both trace to a strongly saturating response whose curvature the theory itself matches with no free parameters.

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri World modeling has no dominant recipe — architectures, training objectives, and state representations interact unpredictably — which makes it a natural test of coding agents doing open-ended research rather than building to a specification. AutoWorldModel-Bench is a closed-loop benchmark where a frontier coding agent must autonomously improve a provided world-model starter under a fixed compute budget, across eight game environments unified by ground-truth entity state in a shared tensor format so that dynamics modeling is isolated from perception and each run takes minutes. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improved on the starter in 63, and in 91% of sessions the winning edit was a research-style change — a new objective, representation, rollout procedure, or architecture — rather than a hyperparameter tweak.

Harnessing agent memory to build lifelong AI partners for materials scientists

Siyu Liu, Bo Hu, Beilin Ye, He Cao, David J. Srolovitz, Tongqi Wen Materials research runs on accumulated experience — working scripts, trusted protocols, warnings attached to failed calculations — that stays fragmented across notebooks, job logs, and individual recollection and does not transfer between AI agents. The proposed self-evolving memory framework stores that experience as inspectable facts and executable skills so it can be retrieved, revised, and migrated across models independently of any particular agent implementation. On 49 real materials-tool-use questions covering 138 executable subtasks, memory nearly doubles GPT-5.2 task success with no parameter updates; in equation-of-state calculations it turns a wavefunction-initialization failure into a pre-execution guardrail, moving results from 22 correct / 1 partial / 4 errors to 25 / 2 / 0 and avoiding 92% of repeated errors; and across 13 simulation workflows remembered skills halve token usage and cut tool calls by more than half by the third round.

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

Yuan Gao (Wanxiang), Zeren Yang (Wanxiang), Junnan Li (Wanxiang), Shawn (Wanxiang), Zhong, Ahmed Dajani et al. InfraBench evaluates AI agents on realistic computing-infrastructure management tasks spanning the full system stack and the full operational lifecycle, with fine-grained per-check risk assessment. Across 15 agent-model configurations, mean effective scores range from roughly 40% to 88% and no configuration achieves a full score across all tasks, while running every task three times shows even top configurations pass only a fraction of their attempts. Per-check scoring exposes a recurring failure mode in which agents satisfy the immediate objective but leave non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. The tasks, evaluation harness, and a live leaderboard are public.

Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction

Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang When a context window fills up, language-model systems compact prior history, and user instructions meant to bind for the rest of a session, such as "do not delete any emails until I confirm," get silently dropped. COMPINT measures this loss across multi-turn chat, agentic trajectories, and long-horizon research, finding current compactors retain only 17% of injected session constraints on average, with most performing worse than running the same task without compaction at all. Retention swings sharply with compactor, prompt, context length, constraint phrasing, and injection location, indicating a systematic rather than setting-specific failure. A plug-and-play constraint-aware extractor running alongside the compactor pushes retention above 90% in all three scenarios without modifying the compactor or the model.

EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents

Yuxi Qian, Yuxiang Ren Memory-augmented language agents mostly store and retrieve past experience, but distilled insights go stale, become over-generalized, or turn harmful in new task contexts, polluting memory when reused. EvoGraph-Mem maintains insights as nodes in an editable graph, each tracking positive evidence, negative evidence, and an activation state, with utility-aware retrieval and a controller that after every task keeps reliable insights, archives invalid ones, revises outdated ones, and adds newly discovered ones. It outperforms representative memory-based agent baselines across different backbone models, and ablations show append-only memory is insufficient for long-horizon tasks while evidence-aware retrieval and graph editing each improve reliability.

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Vasundra Srinivasan Enterprise buyers read agent leaderboards as capability rankings, but a four-facet generalizability-theory variance decomposition across three open agent-trace benchmarks (TheAgentCompany, τ²-bench, and AppWorld) finds the agent main effect explains under 3% of total variance in every dataset and check type while the agent-by-task interaction explains 7-23% — the rankings measure specialization, not general capability. Three estimators (Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM) agree to three decimal places, and further results show aggregate reliability collapsing to zero on the hardest task quartile, training-cell reliability anticorrelating with held-out reliability at r = -0.90, and population-level diagnostics transferring across benchmarks even as per-family agent rankings invert. The authors package this into Deployment Decision Reliability, a one-page reporting discipline turning the variance-component table into five defensible procurement decisions, with all code, data loaders, and fit artifacts released.

Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost

Zixi Huang, Xiheng Wang, Andrew Wang, William Jurayj, Bernal Jim\'enez Guti\'errez, Daniel Khashabi et al. Agents that learn reusable skills for new domains are usually evaluated on accuracy gains rather than on what they cost to run, leaving open which skill-learning strategy is actually economical. The argument here is that representing skills as executable programs yields the largest cost reduction, since a program runs a sequence of actions deterministically instead of burning tokens on trial-and-error over long horizons. SpeedRunner, a coding agent that inspects its own past trajectories and refactors them into skills without replay buffers or a validation step, is tested across three embodied environments, where it reaches the frontier of both task learning and cost reduction while holding up under distribution shift and environmental randomness.

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang et al. Frontier models handle hard tasks once the problem, tools, and success criteria are laid out, but real-world challenges rarely arrive in executable or verifiable form. Apodex Discovery addresses this with a "heavy-duty solver" architecture — foundation model, harness, tools, and control policies — plus a problem-scouting process that surveyed 561 industries to assemble 423 real problems (20 released initially), a shared environment-task-episode abstraction with trajectory recording and intermediate-artifact verification, and HDS6, which scores Tools, Repair, Alternatives, Coherence, Evidence, and Scope separately from final-task success. On adeno-associated virus capsid design the system beat the published state of the art by 7% across viability, tropism, structure prediction, and generative design, and a biomedical drug-repurposing environment lifted GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone.

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology

Del Coburn, Scott Sanner, Dan Silver Differential diagnosis on hard cases requires integrating several kinds of specialist reasoning, and with healthcare making up a reported 5% of ChatGPT messages, how transparently a system arrives at its differential matters. SCoT (Social Chain of Thought) structures multi-agent interaction as a multi-round deliberative pipeline modeled on clinical differential-diagnosis methodology, evaluated against single-agent baselines, one-agent pipeline ablations, and best-of-n scaling. Its recall advantage is not reproduced by monolithic inference alone, and the benefit concentrates in the hardest cases, where successive rounds of specialist conversation recover ground-truth diagnoses that single-pass reasoning misses.

Benchmarking LLM Judges for Mobile Agent Evaluation

Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu et al. Benchmarks for mobile phone agents increasingly delegate the question of whether a task was completed to a language-model judge, without checking how trustworthy those judges are. MobileJudgeBench assembles 931 human-annotated agent trajectories drawn from 6 mobile agent benchmarks, 4 agent models, and 68 apps, and scores 6 judging methods — adapted from SPA-Bench, A3, AndroidArena, and AgentRewardBench, plus a deliberately simple baseline — across several language-model backends. A plain baseline judge fed sampled screenshots matches or beats the purpose-built judging pipelines, with the underlying model rather than the pipeline design driving quality; judge accuracy also predicts both agent ranking fidelity and downstream reward quality for on-policy reinforcement learning, and two backends show opposite failure styles, one conservative and one permissive.

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

Jeremy Spence, Nicholas Assaderaghi, Jinhao Zhu, Nikil Ravi, Raluca Ada Popa, Guannan Wei et al. cross-listed Security agents do well on source code, but malware, firmware, and proprietary software arrive as binaries, requiring reverse engineering to recover semantics first — and existing benchmarks either leak their source into training data or are too small to resemble real software. SRE-Bench was built from scratch by reverse-engineering experts over more than 5,000 hours: 19 private programs averaging 16.9K lines of code, hardened with 44 in-house anti-analysis primitives to yield 262 binaries and 1,572 deterministically graded tasks. Across GPT-5.6-sol, Claude-Opus-5, GPT-5.5, Grok-4.5, and GLM-5.2, the best model scores 61.4% per instance and fully solves only 31.5% of instances; agents also prove oddly insensitive to compiler optimization and static linking compared with human analysts, and ablations confirm that both contamination control and realistic program scale matter.

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

Dylan Bouchard, Mohit Singh Chauhan Uncertainty quantification (UQ) for language models is usually measured on a single generated answer, but an agent's real unit of work is a multi-turn trajectory in which tool calls and intermediate decisions compound into the final outcome. Three families of single-turn scorers — white-box action-token probabilities, black-box consistency across resampled trajectories, and reflexive self-assessment — are tested on five language models across four multi-turn tool-use datasets drawn from BFCL-v4 and τ²-bench. Transfer works but unevenly: token-probability scores swing heavily with the choice of per-turn aggregator, self-assessment is the best cheap baseline, and black-box self-consistency is generally the strongest family, led by trajectory-equivalence and action-set variants.

CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

Linqiang Guo (Peter), Li Gu (Peter), Zihuan Jiang (Peter), Zhixiang Chi (Peter), Siobhan Reid (Peter), Ziqiang Wang (Peter) et al. Graphical user interface (GUI) agents for mobile devices break down on applications they were never trained on, and in practice there are no target-app demonstrations and only a small interaction budget to adapt with. CoAdapt-GUI performs test-time adaptation on the agent's own rollouts and rewards in the new app, updating two things at once: a structured workflow context holding transferable procedures, failure modes, and verification rules but no app-specific interface details, and a LoRA adapter on a frozen vision-language model trained with task-context-matched group-relative optimization. On held-out apps it scores 45.0% on AndroidWorld-Generalization against 37.5% for a policy-only test-time adaptation baseline, and lifts AndroidWorld Plus from 38.6% to 52.9%.

Learning from Online User Feedback for Shopping Agents

Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou Shopping assistants deployed on e-commerce platforms generate enormous conversational logs, but training pipelines typically ignore what users actually say and fall back on click data or synthetic preference pairs. LOFA learns straight from production logs without human annotation by combining reinforcement learning against verifiable purchase outcomes with feedback-aware on-policy distillation that extracts in-dialogue user directives and turns them into dense token-level supervision. The two objectives capture population-level behavioral patterns and individual stated preferences respectively, and on real e-commerce logs improve recommendation quality, response helpfulness, and alignment with user satisfaction over strong baselines.

XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang, Yu Wang, Philip S. Yu, Junhyun Lee Multi-agent systems built from different model families avoid the redundant reasoning of identical agents, but they can only communicate through text, discarding the sender's internal representations, since latent transfer has required matching architectures. The authors identify an entity grounding problem in cross-architecture bridges: rare tokens get compressed away in the continuous bottleneck so entity identity is lost, leaving a bridge-only F1 around 30%. XBRIDGE fixes this without decoding, mapping the sender's context tokens into the receiver's vocabulary as discrete anchors while a latent enrichment bridge lets the receiver query the sender's hidden states, so self-attention ties contextual signal back to specific entities. Across Llama, Qwen, and Mistral pairs in both directions it beats text communication on all seven benchmarks with 11x lower latency, using 264M trainable parameters, or 3.8% of the receiver.

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust et al. Finance benchmarks mostly test data extraction, which frontier models have largely saturated, and reference-matching or generic judge scoring cannot grade the long open-ended answers a real analyst query produces. FrontierFinance is an open benchmark of 220 expert-written queries and 11,543 source-attributed rubrics covering six stages of the investor workflow, run under a common harness limited to publicly available data. Evaluation shows the tool harness matters as much as the model: Samaya's in-house system leads at 56.0% against Claude Fable 5 at 49.2% for roughly 2.2x less cost, and the best open-weight model, Kimi K3 at 46.4%, nearly matches the best proprietary model at 4.5x lower cost. Screening and discovery, and sector, industry and macro analysis remain unsolved, with the best systems reaching only 33% and 39%; the dataset and grading code are public.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen et al. When a coding agent follows a rule, existing benchmarks cannot tell whether it complied or was going to behave that way regardless, since instruction-following suites concentrate rules in the user turn and agent benchmarks score only final task success. Harness-IF scores operational rules individually from execution evidence across 60 multi-turn coding items drawn from a 642-rule library, with 256 rules receiving verdicts and rules placed on the five configurable surfaces a deployed agent actually reads; its Against-Prior Accuracy metric scores only rules shown to oppose the agent's unprompted default, established by re-running tasks with each rule withheld. Across 12 frontier models overall accuracy spans 72.1-85.9% while against-prior accuracy spans 66.1-78.6%, and every model does worse on rules that cut against its priors, by 3.6 to 7.4 points, meaning aggregate scores overstate compliance by a model-specific margin. A separate counterbalanced conflict pilot finds precedence does not track prompt depth: system prompts, project files, and user instructions all outrank tool and skill descriptions.

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

Pan Wang, Yihao Hu, Hang Wang, Zirui Lv, Xin Zhang, Jianshe Li et al. Coding agents can self-correct because compilers, tests, and traces convert failures into typed repair signals, but general language agents usually see only a coarse task failure, so generic recovery playbooks widen the context exactly when a narrower repair interface is needed and mix incompatible signals for invalid actions, missing procedures, and format errors. DARC uses development-set failures to profile each task family's failure modes, prunes mismatched interventions from a shared recovery library, and freezes a verifier-selected success-cost policy before deployment, so the system decides what kind of failure is repairable before deciding how much recovery evidence to spend. On ALFWorld, AppWorld, and XBRL Finance the same protocol yields three different harnesses and improves average task performance over both base agents and broad playbooks while using fewer environment steps or a smaller retrieval budget.

The Sleeping Agent: What Gist-Based Context Compression Loses and Why

Nicholas E. Kyrkewood Summarizing older conversation history into compact gists is standard practice in long-horizon language model agents, but its effect on different kinds of memory retrieval is not well characterized. Salience-Weighted Consolidation — which scores history by salience, sorts it into priority tiers, and gist-abstracts the middle tier — serves as a diagnostic probe across all ten LoCoMo conversations, covering 1,935 matched text-only questions at temperature zero. Gist compression clearly beats truncation on multi-hop and single-hop factual questions but stays far below full context on temporal ones, because the abstraction prompt preserves relational and event structure while discarding dates and times. A one-sentence prompt change lifts temporal expression preservation from 3.05% to 62.39% and recovers +0.314 judge accuracy on temporal questions, leaving entity and event preservation essentially unchanged.

Total Recall at What Cost? Benchmarking the Serving Cost of Agentic Memory Systems

Natchanon Pollertlam, Witchayut Kornsuwannawit Conversational agents commonly offload history into a memory system rather than resending the full transcript, but the serving cost of doing so has not been measured systematically. Three memory systems (Mem0, Hindsight, and Mastra Observational Memory) are compared against a fixed-size rolling window and full-transcript resubmission across two backbone models and conversations of up to 400 turns, with every cost measurement paired against answer accuracy on 665 LoCoMo questions. A regression that predicts cost accurately for the two reference strategies misses the memory systems by 18-69%, because their cost is driven by internal memory behavior rather than conversation length; break-even against the full transcript arrives within the first tens of turns for the cheapest system and never within 400 turns for the most expensive. Accuracy spans 21-54% across systems, and no system wins on both cost and accuracy.

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang Agent skills — reusable natural-language guidance loaded into large language model agents — have shown mixed results, sometimes raising task success and sometimes inflating cost or breaking tasks outright. A differential analysis framework attributes each failure or cost regression to a specific skill by comparing a skill-guided run against a no-skill or semantically matched reference run on the same task, instantiated on SkillsBench and SWE-Skills-Bench and supported by a triage tool called SkillTriage. The analysis surfaces 307 skill-induced problems: 125 functional failures and 182 efficiency regressions, and the functional failures rarely trace to obviously irrelevant skills — plausibly relevant skills instead push the agent to misimplement or omit required elements. Efficiency regressions are not explained by prompt length; the largest contributors are excessive verification (67 cases) and heavy implementation pipelines (30 cases), where skills convert checklists and recipes into mandatory work.

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng et al. Producing a full research paper demands literature retrieval, experiment design and execution, evidence-driven claim revision, publication-ready figures, and consistency across a long generation run. Spark-to-Paper implements this as thirteen composable skills inside an existing coding assistant with no separate agent platform, separating model judgment from deterministic checkable operations and separating experiment planning from reporting so required evidence is specified before results are seen. It also bounds a failure mode the authors call the Self-Refutation Loop, where repeated experiments keep rejecting the original objective, and generates editable vector figures through programmatic plotting and code-based diagram reconstruction. Across eight controlled topics it achieves 99.5% citation validity and 96.4% figure editability, and an ablation raises fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, at 11.9M tokens, $8.1, and 3.2 hours per manuscript.

ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models

Zhou Liu, Chaoyang Han, Zewei Pan, Zeli Su, Wentao Zhang Multi-agent language model systems usually treat roles as hand-written prompt labels with no connection to learned behavior or parameter updates. ExRole instead learns role prototypes from prefix-local team trajectories in a way that predicts future utility, resolves them into readable instructions plus token-aligned role markers, and optionally routes shared LoRA rank slots with turn-aligned credit assignment, making a role an executable control variable rather than a description. On MuSiQue and 2WikiMultiHopQA it gains 15.0/14.4 and 13.5/16.1 exact-match and F1 points over single-agent search, retaining 11.5/11.6 and 7.7/9.7 points against the strongest non-ExRole controls, and interventions across role, agent, and turn indicate the induced roles capture behavioral specialization that transfers beyond fixed agent identities or turn positions.

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

Zhixin Zhang, Xinke Jiang, Zhibang Yang, Weixuan Xu, Guohong Qiu, Xu Chu et al. cross-listed Search agents need to reflect mid-trajectory — judging progress, spotting missing evidence, and deciding whether to continue, revise, or abandon a branch — but reflection happens locally while its value only shows up in the final outcome, so outcome-based reinforcement learning gives sparse and delayed supervision for exactly those decisions. LoongReflect reframes reflection as a memory-control policy over a reversible trajectory tree with explicit reflect and backtrack actions, where reflecting consolidates verified facts and risks into working memory and backtracking drops an unreliable branch while keeping a short corrective lesson. Training combines two channels through a look-ahead, extragradient-style coordination mechanism: a fast channel distilling globally informed behavior from a privileged teacher with supervision restricted to reflection and backtracking tokens, and a slow channel optimizing whole trajectories with Group Relative Policy Optimization (GRPO). Multi-hop retrieval-augmented generation and mathematical reasoning benchmarks show consistent gains over outcome-only reinforcement learning and self-distillation baselines.

Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection

Chaoran Chen, Vy Nguyen, Ziji Zhang, Abhinav Gullapalli, Ziyi Wang, Yuxuan Lu et al. Tool-using agents are typically trained and evaluated in environments where every tool call succeeds, yet real tools fail transiently, persistently, or silently, and recovery requires choosing between retrying the same path, switching to an alternative, or recognizing that no path remains. BENCH2ROBUST converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, so episodes explicitly demand retry, switch, or stop behavior, and is used to test two interventions: a structured runtime memory called Bayesian Tool Memory (BTM) and curriculum-controlled reinforcement learning. Across 7 models from 4 families and two multi-turn benchmark families, injected tool failures open a near-universal robustness gap; on held-out Retail tasks the memory component alone adds up to 16.8 percentage points with no retraining, reinforcement learning learns complementary recovery behavior that persists without it, and combining both reaches 40.8-45.5% under injection while preserving failure-free performance.

CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

Xingyu Yan, Tingting Dai, Antonio De Domenico, Mohamed Sana, Nicola Piovesan, Changchang Li et al. Judging whether an agent can act like a competent telecom troubleshooting engineer requires partially observable environments spanning multiple vendors, devices, protocols and interfaces, which existing evaluations do not model. CTBench is a public benchmark of expert-written root-cause-analysis and path-restoration tasks annotated with golden evidence steps, scoring both the final answer and the diagnostic trail that produced it. Across representative harness and model combinations, agents identify endpoints in path-restoration tasks well but often produce plausible or correct answers without the evidence-grounded diagnosis operations require, struggling most with interface state, link-layer and service-management faults; path restoration also consumes more resources without a matching gain in diagnostic quality.

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang et al. Mechanistic interpretability work remains largely manual even as model development accelerates, widening the gap between what models do and what researchers can explain or control. Mechanist is an agentic system for autonomous mechanism discovery, backed by an interpretability knowledge graph of roughly 13,000 papers linked to a 43-million-paper multidisciplinary database spanning 26 fields, plus a curated library of 32 methods for mechanism analysis, causal intervention and validation. It is reported to generate more valuable hypotheses and execute experiments more reliably than Claude Code and existing AI-scientist systems, and its case studies run from uncovering that unsafe traits transfer across modalities through apparently safe training data, to a mechanistic account of how models form beliefs and infer those of others, to steering scientific foundation models toward DNA sequences with specified properties.

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer cross-listed When language-model agents pursuing user goals meet each other in strategic settings, theory suggests that knowing an opponent reasons similarly can unlock cooperation in games like the Prisoner's Dilemma. A new evaluation framework supplies agents with graded similarity signals about their counterpart and measures how cooperation responds across payoff structures and prompt framings. Models diverge sharply in how they use the signal, the dataset used to compute similarity turns out to matter little, and models systematically judge another model's chain-of-thought reasoning to be highly similar to their own. A behavioral game-theoretic model fitted to the observed reasoning supports cooperative equilibria once similarity scores are high enough.

An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS

Yuzhong Shen, Masha Sosonkina, Peng Xu, Mark S. Gordon Modernizing legacy Fortran is mostly routine transformation work at a scale that leaves it undone across computational science. The described workflow assigns three prompt-specialized Claude Code agent roles working in isolated worktrees under a version-controlled specification the agents themselves wrote and revised, with humans holding only a few gates and safety resting on an exact domain verification oracle — bit-for-bit reproduction of the canonical printed energies from the package's own test suite. Applied to the two-electron-integral routines of GAMESS (General Atomic and Molecular Electronic Structure System), converting twelve files, 56,448 lines, and 225 subroutines from fixed-form Fortran 77 to free-form Fortran 2008, all 612 test runs across a 51-test battery produced zero chemistry-relevant differences, and every file also passed the continuous-integration suite.

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain et al. Enterprise agents must combine structured API calls with document retrieval, but benchmarks typically test one or the other. VAKRA (eValuating API and Knowledge Retrieval Agents) contains over 8,000 executable APIs across 62 domains with three escalating settings — varied API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning under natural-language tool-use policies — and verifies answers by re-running predicted tool calls against live APIs so multiple valid solution paths count. Under a fixed ReAct harness, the best frontier model manages 70.4% on single-hop endpoint tasks, falls to 50–51% on compositional APIs, and drops as low as 2.4% on policy-constrained unanswerable queries; trace analysis places failures in entity disambiguation and cross-source grounding rather than in the mechanics of calling tools.

Scaling Automatic Research Agents via World Models

Xiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen et al. Reinforcement learning for agents that automate empirical research hits a scaling asymmetry: generation batches across trajectories and shares compute, but each environment execution needs its own sandbox and real wall-clock time, so execution dominates cost as trajectories lengthen. World Model RL replaces environment execution with a learned world model and adds Online Debiasing and Inverse-Variance Denoising to counter the bias and noise in the model's rewards, with proofs that both mitigations strictly improve the convergence guarantee. Training ran 3-4x faster than standard reinforcement learning while also scoring higher, and the resulting 4B and 9B agents beat open-weight agents of 48B and 120B on held-out benchmarks; the method also transfers to post-training embodied vision-language-action policies.

DiG-bench: Discovery in Games

Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan, Timothy Muller, Clare Maguire et al. cross-listed Few existing benchmarks test whether an AI system can discover new generalizations by experimenting in an environment whose objective it has not been told. DiG-bench fills that gap with 70 independent games, each encoded as a short string with its own transformation rules that must be inferred through interaction, and levels whose win conditions are likewise hidden, arranged in seven difficulty tiers. The easiest tier is routinely solved by multiple models while the hardest challenges the strongest models in agentic harnesses, yet every one of the 70 games was solved by at least one human on a first attempt; 21 games are public and the rest held back for uncontaminated evaluation.

CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu et al. GPU kernel-writing agents normally treat the compiler as a black box that returns only errors, pass/fail, and timings, which makes expert-level kernels hard to reproduce. CAKE co-designs the two sides: agents write in a typed, hardware-explicit intermediate representation exposing warp roles, memory movement, synchronization, and pipelines, which in turn supports verification, cost modeling, and localized diagnostics, while recurring failures are folded back into the harness as verifier rules, new primitives, and reusable optimization tactics. On a Flash-KMeans clean-start task on B200 hardware with an 80-million-token budget, the best candidate reaches 1.144x the tuned FlashML baseline versus 0.928x for direct CUDA/PTX generation, and agent-generated Kimi Delta Attention hits a 2.05x geometric-mean speedup over the official FlashKDA implementation while passing end-to-end serving validation.

SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

Oguz Serdar, Cuneyt Mertayak cross-listed Long-running tool-using agents face a pre-commit choice at every consequential action — send the email, merge the pull request, wire the payment, or hold for human review — and SteerBench-Work scores models on exactly that gate. Release v2026-05 contains 106 scenarios anchored in public workplace incidents across developer operations, customer service, finance, legal, medical, HR, and security, each paired with an evidence-reversed mirror and calibration controls, with labels split nearly evenly between proceed and hold. Across 30 model conditions the errors are almost entirely one-sided: models wrongly block authorized, evidence-cleared work 28.1% of the time while wrongly permitting unsafe actions only 1.0% of the time. Performance collapses from 98.5% on the famous incidents themselves to 63.8% on their evidence-reversed mirrors, and higher-capability models often over-refuse, indicating that general capability does not transfer to steering calibration.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

Haoze Wu, Chuqiao Kuang, Tianyi Zhuang, Xiaoguang Li Search agents that run for dozens of steps receive a single reward at the end of a trajectory, which leaves reinforcement learning with almost no signal for assigning credit to individual steps. On-policy self-distillation can supply dense token-level targets, but a teacher that knows the correct answer reasons systematically differently from a student exploring the web, so naive distillation teaches the student to imitate the information advantage rather than to search better. The fix combines Evidence Anchors — short step-level evidence snippets pulled from the web that expose key reasoning steps without revealing the full answer path — with Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher-student disagreement into step-level advantage weights inside GRPO and applies them only to incorrect trajectories, leaving correct ones untouched. On Qwen3-8B, SSPO outperforms GRPO on BrowseComp, GAIA, and FRAMES while matching or beating GRPO runs given twice as many gradient steps, at roughly 5 percent extra cost per step.

Latent On-Policy Self-Distillation

Guibin Zhang, Jiayang Lyu, Ran Sun, Xinlei Yu, Haoyu Zhao, Qibing Ren et al. On-policy self-distillation trains an agent by letting a privileged teacher give dense supervision on the student's own trajectories, but existing variants require a human designer to specify what privilege the teacher gets — answers, feedback, skills, or reference trajectories — which caps how far the loop can scale on its own. Latent On-Policy Self-Distillation (LOPD) makes that privileged context learnable instead: it retrieves relevant past experiences, compresses them into continuous latent tokens that condition the self-teacher, and supervises the student at every visited prefix, with a privileged-margin objective to keep the latent context stable. It outperforms reinforcement learning with verifiable rewards and prior self-distillation methods including OPSD, SDPO, and Skill-SD on agentic tool use and code generation, and beats GRPO and Skill-SD using under 30% of their rollout budget; ablations indicate the learnable context is what produces the gains.

VALG: An Agentic System for ML Theory Research

Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou cross-listed Proving a machine learning theory result means co-developing the problem formulation, the theorem statement, and the proof mechanism, since the data model, training protocol, oracle access, loss, and randomness jointly define what a theorem can say. VALG organizes this as an autonomous agent workflow: each theorem branch holds a fixed mathematical specification, checks theorem-level composition against a typed proof-dependency graph, and builds and reviews local proofs in dependency order, with failures diagnosed as originating in a derivation, the proof structure, or the formulation and routed accordingly — formulation-level blocks spawn an explicitly related variant or relaxation that keeps its relation to the source problem recorded. Run on nine subproblems drawn from five COLT 2026 open problems, two produced internally finalized theorem candidates matching the scope of their source briefs, while the other seven yielded restricted-method results, special cases, or conditional theorems. The system is released as open source.

Training AI Scientists to Replicate Research

Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin et al. Replicating a paper demands the same hypothesis-driven exploration as open-ended research, since replication surfaces details the original left unspecified, which makes it a useful training ground for research agents. Replica is a scalable task space of replication problems paired with an auto-generated rubric-based judge that provides a low-noise reward signal agreeing with human quality assessments. Post-training on it yields Faraday, a 27-billion-parameter agent that calls coding agents as tools and outperforms Claude Opus 4.8 and GPT-5.5 on held-out replication tasks, with rollout inspection showing it follows a more scientifically principled process rather than relying on an elaborate scaffold.

Intern-S2-Preview: Scientific Agentic Foundation Model

Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng et al. Scientific work demands models that read heterogeneous evidence, call tools, and hold up over long task horizons, which general foundation models are not trained for. Intern-S2-Preview is a family of scientific agentic models built from multimodal pretraining on rendered scientific documents and interleaved image-text corpora, followed by supervised fine-tuning, multi-task and agentic reinforcement learning, and on-policy distillation, supported by engineering such as partial rollout with off-policy correction, adaptive length regularization, and online speculative decoding. The 397B-parameter model extends long-sequence understanding to numerical time-series forecasting and reports competitive or leading scores across scientific, multimodal, agentic, and general benchmarks, while a separate memory-augmented Intern-MemDec-4B module lifts the Biology-Instructions average from 56.92 to 60.32 with the backbone left frozen.

Vero: Can AI Agents Build Formally Verified Software Repositories?

Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel et al. Coding agents give no correctness guarantee, and existing verified-code-generation benchmarks test either single functions or proofs written against implementations that are handed to the model. Vero evaluates joint implementation and proof synthesis at repository scale: 43 multi-module Lean 4 instances derived from real codebases in Python, Dafny, Verus, and Coq, spanning cryptographic protocols through distributed systems, each with fixed API interfaces, hand-curated formal specifications, and reference implementations, plus an audit mode letting agents prove a specification unsatisfiable or a reference wrong so curation errors surface. With frontier coding agents given Lean toolchain access, the strongest configuration fully solves only 27 of 43 instances and closes no specifications at all on the hardest repositories.
8 more specialized papers

Other 42

VQ-bench: A Composable Vector Quantization Framework

Ashwin Padaki, Amir Ingber, Edo Liberty Vector quantization has become central to AI infrastructure and is seeing a surge of new algorithms, but comparing them is inconsistent. The framework identifies 7 conceptual quantization primitives that can be composed arbitrarily and then re-expresses 25 common quantizers as pipelines built from those primitives. VQ-bench is released open source with reproducible public benchmarks so new methods can be added and compared on equal terms.

How Organizations Use AI: Evidence from ChatGPT

Aaron Chatterji, David Holtz, Neel Rakholia, Prasanna Tambe, Gawesha Weeratunga cross-listed Linking ChatGPT Enterprise account records to usage logs, worker roles, task classifications, and public-company financials through March 2026 gives a privacy-preserving view of how firms actually deploy frontier generative AI, covering over 1,500 organizations and more than 17 million messages at the six-month adoption horizon. Four patterns emerge: usage is growing both through new firm adoption and rising intensity among existing customers; U.S. public-company adopters skew larger, more valuable, and more research-and-development- and sales-and-administrative-intensive; active use spans job functions and seniority with early-career workers the heaviest users; and messages cover a wide span of knowledge work including writing, technical tasks, communication, and information synthesis. The authors read the spread across firms as evidence that organizations are still learning how to fold AI into their workflows.

Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals

Alireza Kargarzadeh, Nariman Khaledian, Navid Parvini, Arman Khaledian cross-listed AI tools are promoted as a scalable fix for educational access gaps, yet the surrounding infrastructure can disadvantage speakers of under-represented languages before any model is trained. Using Bengali and low-connectivity classrooms as the case, the analysis names four compounding failures: Bengali is under 0.5% of global web content despite roughly 4% of world population, major multilingual corpora carry a 67:1 English-to-Bengali training-token deficit, the alphasyllabary script incurs a tokenization penalty through high token fertility that multiplies the data gap, and rural internet penetration sits at 36.5% against 71.4% in cities. The authors argue that data scarcity is a structural artifact of past resource allocation and design defaults rather than an isolated technical shortfall, and propose offline-first design as an equity-oriented infrastructure strategy.

Prof-K: Probabilistic One-Pass Filtering for Efficient Top-k Selection

Tadeusz Dziarmaga, Witold Sikora, {\L}ukasz Struski, Jacek Tabor, Marcin Mazur Selecting the k largest elements of a large array underpins databases, retrieval, and machine-learning workloads such as sparse activations and attention pruning, but exact methods are memory- and compute-heavy while approximate ones lean on heuristics that break on heavy-tailed or adversarial inputs. Prof-K takes a single pass: a small random sample sets an adaptive threshold, the input is streamed once into a compact buffer, and an exact routine over the buffer returns the true top-k with probability at least 1 minus a user-specified epsilon, with high-probability bounds on buffer size and a near-optimal sample size derived as a function of input size and k. Measured speedups over PyTorch's optimized topk and the recent RadiK implementation range from 1.5x to 10x, largest for big inputs with small-to-moderate k, and the guarantees hold regardless of input distribution; relaxing the recall target trades accuracy for further speed, demonstrated on training BatchTopK sparse autoencoders.

Interpretable Causal Discovery via Causal-Effect Constraints

Cixuan Zhang, Guy Van den Broeck, Benjie Wang Causal discovery usually stops at predicting edges, but scientists often want to explain a specific observation, such as why one variable exerts an unusually large effect on another. The authors formalize this as conditional causal discovery and cast it as Bayesian inference over graphs and parameters conditioned on an event like a causal-effect constraint, then borrow rare-event estimation techniques to handle the case where that event carries very small posterior mass, gradually steering a particle population into the constrained region while keeping samples that approximate the conditional posterior. Synthetic-graph experiments validate accuracy at small and large scales, and a case study on the Sachs protein signaling dataset shows the method producing pathway-level summaries for scientific exploration.

Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning

Zeyang Zhang, Tieliang Gong, Junyan Lu, Weizhan Zhang Networks trained on a sequence of tasks progressively lose plasticity, and activation functions are part of the cause: fixed forms carry a spectral bias toward low-frequency variation, while fully learnable ones update without constraint and induce catastrophic forgetting. The proposed learnable wavelet activation decomposes into low- and high-frequency components to counter that bias, adds dynamic wavelet injection triggered by a loss-driven criterion to boost plasticity when a new task arrives, and regularizes to protect previously learned knowledge. The authors prove the hybrid wavelet architecture is structurally necessary for efficient L² approximation and that the decoupled learning-rate mechanism restores high-frequency plasticity, and report state-of-the-art results with sustained trainability across diverse continual learning benchmarks.

Comment on "Modeling rapid language learning by distilling Bayesian priors into artificial neural networks"

Orr Well, Idan Tarshish, Nur Lan, Roni Katzir A critique of McCoy and Griffiths (2025), who reported that a Bayesian prior can be distilled into a neural network through Model-Agnostic Meta-Learning (MAML), producing formal-language learning comparable to the Bayesian learner of Yang and Piantadosi (2023). The authors argue that under the standard definition of a prior the procedure does not install one — it only supplies a favorable weight initialization, leaving the objective function untouched — and that the more permissive reading, in which the whole system implements a Bayesian learner without an explicit prior term, runs into substantive difficulties of its own. Empirically they show the meta-trained model diverges from genuine Bayesian behavior by overfitting and generalizing poorly to held-out data.

Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization

Jinhyung Bae Neural combinatorial optimization solvers draw many candidate solutions per instance and report the best, always with the same sample count per instance; whether spending that fixed budget unevenly across instances helps had not been measured. Evaluating POMO, AM, and SymNCO on uniform 100-city travelling salesman instances, an oracle allocation scored on the same stored samples appears to gain 2.2–2.6%, but measured out of sample the gain is indistinguishable from zero, and the bias does not shrink with more samples or instances. Under distribution shift, a pre-registered experiment finds a real effect: allocation guided by held-out statistics improves best-of-k by 11.5% for AM and 12.0% for SymNCO at equal evaluation budget, with a negative control on the shift-robust POMO showing nothing. The authors release a correction procedure, a reporting checklist, and their pre-registration record.

Into the ORBIT for Time Series: Training Regimes for Foundation Models

Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong Training recipes for time series foundation models (TSFMs) have drawn far less attention than architectures, leaving pre-training distributions poorly controlled with respect to domain imbalance, context length, forecast horizon, and missing values. ORBIT (Omni-Range Bootstrap Incremental Training) makes that distribution explicit, combining bootstrap multi-level sampling — which governs how often each dataset, record, target variable, context window, and horizon is drawn — with incremental variation of context and horizon lengths within a single training stage. The authors use it to train Falcon-2.0, a univariate encoder-only Transformer with missingness-aware triple-channel patch tokenization, plus a rank-guided objective that turns late-layer representations into stop-gradient teachers for shallow layers at no extra inference cost. The resulting model reports strong zero-shot forecasting across domains and frequencies on GIFT-Eval and fev-bench.

Exponential quantum advantage for learning signals with a single qubit

Ishaan Kannan, Sridhar Prabhu, Saeed A. Khan, Mandar M. Sohoni, Xingrui Song, Saswata Roy et al. cross-listed Quantum advantages in sensing usually require hardware far beyond what experiments can deliver, but attaching a single controllable qubit to an otherwise conventional sensor turns out to cut the number of measurements needed to learn a classical signal exponentially, for tasks including recovering Fourier coefficients, extracting temporal correlations, and estimating transformations of observables. On a superconducting cavity-qubit device the authors experimentally demonstrate 10⁷-fold reductions in measurement count for Fourier-amplitude and time-varying signal learning, and simulate orders-of-magnitude gains for weak-signal dark matter detection and wireless communication. The guarantees come from Quantum Phase-Space Inference, a framework that converts experimental objectives and constraints into matching lower bounds and optimal algorithms with a certificate of advantage, covering regimes outside quantum Fisher information.
32 more specialized papers

Theory 42

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

Zijian Zhao, Sen Li cross-listed Networks that score a pair of inputs as the inner product of two separately encoded vectors have been reinvented across operator learning, bipartite matching, retrieval, and contrastive vision-language training, with no shared theory for how many interaction modes to keep, how to normalize, or when the design will fail. The authors define a class of functions of low interaction rank whose complexity is captured by an interaction spectrum: approximation error splits into spectral truncation plus encoder realization, sample complexity scales with the sum rather than the product of the two encoder complexities, and spectral decay gives a criterion for when the architecture is appropriate. The framework also exposes a gauge symmetry that leaves learned coordinates arbitrary, and shows normalization is gauge fixing — whitening pins the interaction modes down to permutation and sign, which experiments confirm by relating independently trained CLIP models through a single rotation whose removal exposes interpretable concept axes.

Black-Box Knowledge Transfer across Distinct Feature Sets

Oh-Ran Kwon, Daeyoung Ham cross-listed A pretrained black-box predictor encodes knowledge from data and compute that cannot be reused when the features available in a new setting differ from the ones it expects. The proposed method splits the target regression function into a transferable part, which the black box can inform, and a non-transferable part unique to the new feature space, estimating the former from abundant unlabeled paired features bridging the two spaces and the latter from scarce labels. The authors derive prediction-risk bounds that beat a no-transfer baseline when the non-transferable component is small or smooth and adapt automatically to either case, and show that under further conditions the worst-case risk is of strictly smaller polynomial order than the minimax risk achievable from labeled data alone; the framework extends to aggregating several black boxes on different input spaces.

Unifying Generative Models with Path Integrals

Ramon Winterhalder Recasts generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models all fall out of one master action under different evaluation principles. Written in Martin-Siggia-Rose-Janssen-de Dominicis (MSRJD) form, the action separates free from interacting probability flows and admits diagrammatic perturbation theory, which yields a one-loop correction to deterministic samplers that costs no stochastic sampling; on solvable and nonlinear drifts this cuts a 53% tree-level error to 1.6%. Imperfect learned scores enter as diagrammatic insertions and produce a response-weighted score-matching objective, and designing symmetry-equivariant drifts becomes an operator expansion with effective-field-theory power counting.

Fast Length-Squared Sampling for Positive-Semidefinite Matrices

Rajarshi Bhattacharjee, Ethan N. Epperly, Cameron Musco, Aaron Tian cross-listed Many sublinear-time matrix algorithms assume free access to column norms so they can sample a column with probability proportional to its squared euclidean norm, an assumption that is doing real work. A simple rejection-sampling scheme performs this length-squared sampling on an n-by-n positive-semidefinite matrix in O(n) expected time, sublinear in the input size and optimal even when the matrix is diagonal, removing the norm-oracle assumption for this class. Applications include an asymptotically optimal relative-error estimator for the Frobenius norm of a positive-semidefinite matrix and a much simpler sublinear algorithm for robust positive-semidefinite low-rank approximation that nearly matches the more involved construction of Bakshi et al.

Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

Andrew Cheng, Ali Eslamian, Jie Cheng, Mehdi Zargham, Qiang Cheng Networks can be trained or fine-tuned by optimizing a small latent vector that a frozen random map expands into a full parameter update, which raises the question of how large that latent space must be to reach low loss. The known accessibility transition is recast in conic form, centered for compact convex targets at the statistical dimension of the polar cone, and the main result is an orientation-resolved quadratic master formula predicting the random-slice residual from both the curvature spectrum and the displacement profile from initialization to solution, which reduces to the earlier Gaussian-width bound in a radius-only special case. Building on it, Random Mapping Networks instantiate the predicted latent dimension with structured Hadamard or seed-regenerated Gaussian maps, avoiding dense-map storage and cutting optimizer state from the full parameter count to the latent dimension; the predictor tracks measured transition locations across quadratic and neural-curvature experiments and beats orientation-agnostic approximations when displacement direction matters.

Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing

Yuxiao Wen Comparing many adaptive decision policies online — rankers, pricing rules, language-model agents — normally gives each of J policies its own horizon-T trajectory, spending JT potentially costly or risky interactions. Tree-Coupled A/B Testing (TCAB) instead links the policies' current histories with a predictable tree, maximally couples every parent-child context-action law, and shares one reward within each component of matched edges, so each policy still has exactly its standalone finite-horizon trajectory law despite the policies being deliberately dependent. The number of reward queries equals T plus cumulative tree-edge total variation, which means for fixed J and policies with sublinear regret the expected cost is T + o(T) rather than JT, and experiments on reward-model evaluation, multiple-choice language-model evaluation, and adaptive search show substantial gains on the cost-precision frontier.

Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization

Zhixin Ren, Yau Lyu, Congrong Li, Liping Zhang, Shengbo Eben Li Momentum optimizers are ubiquitous, but the relationship between the momentum recursion, the geometry of the update, and acceleration remains only partly characterized. The AIM (ADMM-Inspired Momentum) framework recasts momentum as a multiplier-like correction driven by the residual of a penalty-based variable splitting, recovering the exponential moving average of gradients from an ADMM-style multiplier update and separating two normally entangled roles: the residual penalty sets the update geometry, while the approximation of the objective subproblem sets the acceleration form. From this, the authors derive RADAR, which combines relativistic adaptive geometry, decoupled residual correction, and second-order momentum filtering, and prove stochastic convergence via a variance-perturbed Lyapunov drift analysis. Across supervised vision, language modeling, and reinforcement learning tasks, RADAR reports consistent improvements over strong adaptive optimizer baselines.

Statistical Properties of Robust Learning under Distributional Shifts

Zhiyi Li, Xiaojie Mao, Yunbei Xu, Ruohan Zhan cross-listed Robust learning methods such as Distributionally Robust Optimization (DRO) and Robust Satisficing (RS) are meant for deployment environments that differ from training, yet their guarantees are usually stated for the source distribution or for adversarial worst cases over an ambiguity set rather than for the actual shifted target. Finite-sample generalization bounds are derived for excess loss under the target distribution for both methods, making explicit the trade-off between reduced shift sensitivity and the regularization penalty each robustness hyperparameter imposes, and avoiding the curse of dimensionality that comes with Wasserstein empirical concentration. When partial knowledge of the shift magnitude or direction is available, the authors give information-directed hyperparameter calibrations under which DRO and RS turn out to behave complementarily, and illustrate the framework on a network lot-sizing problem with demand shifts.

Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang Networks trained by gradient descent on smooth objectives sometimes learn in abrupt steps separated by plateaus and sometimes follow smooth power laws, and both patterns appear in architectures with little microscopic resemblance — a hint that few collective variables govern them. Permutation symmetry over interchangeable units, combined with smoothness and a vanishing per-unit gradient at the origin, forces a universal leading expansion near the small initial weights, the quadratic form Tr[WWᵀA(x)], with all architectural specifics compressed into a structure matrix computed separately for perceptrons, attention layers, mixtures of experts, and convolutions. Training dynamics then close on the order parameter M = WWᵀ and, when the data matrices share an eigenbasis, reduce to a Lotka-Volterra system whose modes switch on sequentially — plateaus emerge as the small-initialization singular limit, and unresolved modes merge into a power law whose exponent the theory predicts. Numerical checks across training methods and architectures confirm both regimes.

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

Yikai Xu, Zhao Chen, Jian Huang cross-listed Given a dataset where some fraction of samples are corrupted, Wasserstein Filtering recovers the clean population distribution by discarding suspicious samples and keeping the empirical measure of the rest, selecting the subset whose empirical distribution is farthest in Wasserstein distance from the fully contaminated one so that geometrically influential outliers are isolated first. Three tractable algorithms implement this: a marginal screening scheme plus two joint optimizers built on entropic optimal transport and sliced Wasserstein approximations. Under a new contamination model combining well-separated outliers with locally indistinguishable perturbations, the estimator is proved minimax optimal over distribution families with bounded covariance, and experiments on synthetic data, anomaly detection benchmarks, and diffusion-model training under heavy contamination show it works as a model-agnostic preprocessing step.

Bagging Robustly Learns VC Classes with Linear Sample Complexity

Omar Montasser cross-listed Learning classifiers that stay correct under test-time adversarial perturbations was known to be possible for classes of finite VC dimension d, but the best sample-complexity upper bound was exponential in d. Combining Breiman's bagging with robust empirical risk minimization — running robust ERM on O(d) independent bootstrap samples, where d is the dual VC dimension, and taking a majority vote — achieves adversarially robust learning with sample complexity linear in d, an exponential improvement over the 2019 bound. A matching lower bound shows the oracle cost is tight: any learner in this model needs Ω(d*) calls to a robust ERM oracle no matter how much data it has.

The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity

Martin J. Wainwright Masking diffusion models for discrete data reveal tokens according to a schedule, and how that schedule should be chosen has lacked a sharp theoretical handle. A path-resolved quantity called unmasking growth complexity is introduced whose local increments directly control Kullback--Leibler discretization error, unifying the analysis of Bernoulli-subset and fixed-cardinality unmasking and producing optimized single- and multi-block schedules in log-reveal-odds coordinates. Because these increments can be estimated from samples via KL increments along coupled reveal trajectories, the result is a certified-optimal sampler that hits a target KL error with high probability using iterations within a constant factor of the oracle schedule, with examples showing Ω̃(√d) improvements over coarse schedules from a constant number of adaptively placed blocks.

Defensive Boosting for Online Probabilistic Forecasting

Georgy Noarov, Aaron Roth Online boosting for probabilistic forecasting against an adaptive adversary comes in two flavors with incomparable promises: gradient boosting matches the best predictor in the span of the weak class in Brier score but says nothing if that span holds no good predictor, while weak-to-strong boosting drives classification error to zero only when a weak-learning condition holds. The Defensive Booster, a defensive-forecasting algorithm, obtains both guarantees simultaneously on every adaptive sequence, using the dual view of boosting: persistently high randomized classification error means the algorithm's mistake weights form a smooth reweighting on which every weak hypothesis has low edge, an ex-post certificate that the weak-learning condition fails. A strongly adaptive variant holds on every time interval, and because the method queries a single weak learner rather than maintaining a large ensemble, experiments on synthetic and real streams show competitive or better accuracy with orders-of-magnitude faster runtime.
29 more specialized papers

Safety & Alignment 32

Forecasting Side Effects of Activation Steering

Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang, Jun Sun Activation steering changes a language model's behavior by adding a learned direction to its hidden activations, but it frequently perturbs unrelated behaviors, which makes deployment risky. Building a cross-effect matrix over a taxonomy of 67 behaviors on three open-weight models shows that side effects are common, structured, and often asymmetric, with interactions that similarity-based heuristics do not explain. Despite that complexity, side effects turn out to be largely forecastable before steering is applied: their magnitude depends mainly on the target behavior, and their direction can be predicted from the model's unsteered representations far more accurately than by simple baselines, opening the door to proactive safety auditing of steering interventions.

The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification

Yoshinori Watanabe The argument is that semantic safety constraints, such as "the agent does not escape its sandbox," are off-support objects: unlike the real log-canonical threshold (RLCT) of singular learning theory (SLT), the safety predicate is not measurable with respect to the model and the data distribution. From that single non-invariance follow several claims as corollaries: why reward hacking and sandbox escape emerge under outcome-based optimization, why Bayesian prior design and soft penalty weighting have poor leverage in singular models, and, centrally, why hard invariants belong in the containment harness while only soft dispositions belong in the model. The same predicate is nonetheless locally certifiable by formal verification, mirroring how the local learning coefficient locally pins the RLCT, with the residual difficulty of identifying which off-support region matters coinciding with performative prediction. The July 2026 OpenAI-Hugging Face evaluation incident is the motivating case, and numerical experiments plus Lean proofs are released.

Agent Safety Should Be a Runtime Contract

Albus W. Ng, Yi Han, Jusheng Zhang, Wenhao Wang cross-listed Training-time alignment methods — RLHF, DPO, Constitutional AI — are argued to be structurally insufficient for agents that execute code, mutate files, send messages, and write to databases, because the risky behavior occurs at execution time. The position advanced is that the harness should enforce a runtime contract with a preventive face (sandboxes, permission gates, output filters, trajectory monitors) and an evidential face that gates task submission on verifiable proof such as test runs, log captures, file diffs, and citation grounding. Four supporting audits are released with row-level data: 52 documented agent and LLM safety incidents, a false-completion audit, a trajectory-schema audit of 12 public agent systems, and a title-level pass over all 28,560 papers accepted at NeurIPS, ICML and ICLR from 2023 to 2025 showing a pooled 8-12x publication imbalance favoring training-time over deployment-time safety work. The authors formalize an Agent Trajectory Schema and Evidence Chain and argue the unit of safety should be the trajectory-with-checkable-evidence rather than the model.

Backdoor Decontamination Dynamics in LLM Agents

Gabriel Huang, Abhay Puri, L\'eo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella et al. cross-listed A backdoor planted in an open-weight agent during fine-tuning stays invisible until its trigger fires, and a defender who does not know the trigger cannot unlearn it directly. One proposed remedy is to install a known backdoor of your own and then unlearn it, hoping the unknown one goes with it; this work tests whether that actually happens, running 115 experiments on the AgentDyn tool-calling benchmark while independently varying trigger, response, teacher, and fine-tuning method. Defensive poisoning alone erases roughly 56% of original backdoors, and the subsequent unlearning step drives almost all survivors to erasure, confirming that trigger recognition and malicious execution are behaviorally separable; co-installing up to four backdoors raises resistance, yet decontaminating one known co-resident clears 52 of 60 others. Visualizing model internals with J-lens shows benign responses restored while traces of trigger awareness persist in intermediate layers.

Governing Agentic AI in FinTech

Henry Han cross-listed Banks are handing consequential decisions to agentic systems that decompose goals and call tools with minimal oversight, and the claim here is that the binding governance constraint is verifiability rather than capability. The proposed "Verifiability Gap" measures the shortfall between the verification that delegated authority demands and the explainability and reproducibility actually retained, indexed to a specific verifier, evidentiary standard, and audit lag; three studies test the framework across nine model versions from a three-billion-parameter local model up to a commercial frontier system. Replay controls turn out to belong to the provider — the frontier model rejects temperature, top_p, and top_k outright and exposes no random seed — orchestration architecture silently changes final actions with no execution record repeating in any configuration, and deterministic credit models reproduce current actions perfectly yet cannot recover historical ones, supporting a view of reproducibility as a governance profile rather than a single number.

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

Ted Kwartler, Alan Aqrawi, Arian Abbasi cross-listed Long-running agents compact their context by replacing the transcript with a generated summary, and prior work showed that dropping a standing safety constraint during compaction produces behavioral violations. Examining a single compaction cycle in detail reveals that when a rule is not dropped outright, compaction often leaves a degraded residue that reads like a rule but no longer constrains behavior. On behavioral replay, a degraded residue led models to perform the prohibited action 34 and 57 points more often than an intact rule under two replay models, and rule-form text survives compaction more reliably than prominence-matched facts — which is precisely why presence-based auditing feels sufficient while providing false assurance, since textual survival is not protection and the failure is silent at runtime.

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

Zirui Song, Huaxing Liu, Xiang Wang, Shuai Li, Xinye Li, Lang Gao et al. Unlearned language models can keep latent traces of the knowledge they were supposed to forget, but existing audits are one-off diagnostics that say nothing about whether those traces predict recovery under continued training. J-Access is an inference-time audit that uses a Jacobian lens to project intermediate representations into vocabulary space and measures how often target concepts stay accessible along the output pathway; it was applied to 398 public unlearned models covering eight unlearning methods. Most retain access above the retain-only gold level, pre-attack accessibility predicts recovery speed and extent at the model level though not which individual facts return, and — the cautionary result — directly optimizing to minimize J-Access teaches the model to hide knowledge from the audit rather than delete it, yielding lower audit scores but greater post-attack recovery.

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

Valentin Rodionov, Shamil Assylbekov cross-listed Putting language models into scientific workflows where no downstream verifier exists assumes they can tell sound literature from unsound, a capability distinct from factual accuracy and not previously measured directly. TRACES probes it with 42 retracted, fraudulent, and pseudoscientific papers, each paired with a near-verbatim preamble from the source and a plausible study-design request, spanning five claim types from fabricated observation to cargo-cult experiment, and scored on outright rejection, recognition-while-engaging, and an Engagement Depth Index measuring reproduction of withheld paper-specific details. Across 30 models and 10 runs, every model fails more than 71% of agentic probes and 22 of 30 fail more than 90% of the time, with models engaging untenable premises in 95% of non-empty responses; refusals cluster on a few high-notoriety topics and vanish under matched-structure controls, which points to topic-keyed safety behavior rather than genuine epistemic judgment.

Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

Haokai Zhao, Yunze Xiao, Weihao Xuan, Flora Salim, Benjamin Tag, Aditya Joshi Group alignment tunes a language model to reflect a demographic group's opinions and values, and alignment generally is known to breed sycophancy — over-agreeing with the user regardless of the facts — but group alignment work measures only how closely outputs match group opinion. Group Alignment-induced Sycophancy (GAS) evaluates both sides across 3 alignment methods, 4 models, and 13 demographic groups, tracking the intended opinion-alignment gain alongside the unintended sycophancy shift. Under an identical alignment budget both effects vary substantially by group, with some groups gaining far more opinion alignment than others and the sycophancy shift forming a group-specific multi-dimensional profile rather than a uniform change, which argues for reporting group alignment as a two-sided profile instead of a single fit score.

Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari If refusal behavior in aligned language models were truly spread across all weights, it would be hard to explain how easily alignment breaks. Transplanting weights from aligned models into their matched unaligned base counterparts — swapping attention weights, feed-forward (MLP) weights, contiguous layer spans, or individual MLP blocks — localizes the effect across two open-weight model pairs and four safety benchmarks. MLP weights carry refusal far more than attention weights, recovering at least 2.7 times more malicious-prompt refusal, and the responsible parameters cluster mid-network, with the layer 8-11 block picked first in all six greedy searches. The components also interact non-additively: in five of six searches, adding more aligned blocks sometimes lowered refusal performance, and selected subsets beat transplanting the full MLP stack.

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur, Dilek Hakkani-T\"ur Language models that debate, advise, or deliberate together need to resist being talked out of correct answers, and that resistance is measured here by training attackers rather than hand-writing prompts. An adversarial reinforcement learning setup optimizes persuader agents to flip a target model's answer in a single exchange, and learned persuaders raise success against the training-time target from about 24% to over 93%, revealing failure modes that static prompting misses. The strategies transfer without retraining — 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini, which a curriculum that warms up on more persuadable open-weight models pushes to 38% — and optimized attackers converge on credibility tactics such as fabricated citations and invented authoritative sources.

Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

Simone Mungari With users consulting chatbots about politics during election periods, the question of which parties and politicians those systems appear to favour has practical stakes. The authors build a reproducible auditing framework that prompts several large language models to evaluate political parties and leaders against nine criteria, deliberately measuring observable behaviour — consistency across repeated evaluations, disagreement between models, refusal rates, and sensitivity to how the prompt is worded — rather than trying to infer any underlying belief. The framework also varies the persona the model is told to adopt, and is demonstrated end to end as a case study of Italian parties and leaders.

LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

Xinhao Zhong, Yuxia Qiao, Junhao Li, Hao Fang, Yi Sun, Bin Chen cross-listed Reinforcement learning post-training gives multimodal reasoning models exploratory chains of thought, and the authors show this creates a privacy hole that existing unlearning methods miss: a fact scrubbed from the final answer can still surface in the reasoning trace, far more often in natively RL-trained models than in their non-reasoning bases. They trace this to a distinctive token-level entropy signature that marks sensitive content during RL-induced exploration and is largely absent from base models. LEMUR exploits that signal without any training, using entropy dynamics at inference time to detect when sensitive reasoning starts and stops, then rerouting the trajectory by replacing committed tokens with sanitized probability-weighted embeddings re-grounded in the input image; it suppresses leakage in both the reasoning trace and the answer better than existing unlearning methods while preserving unrelated utility and fluency.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning

Lang Cao Aligned language models are supposed to decide refusal based on request content, but the same request draws different safety decisions depending on what persona or trait the system prompt assigns, a failure the authors call trait-induced safety variation and measure with two refusal-based metrics: Trait-Induced Deviation, which compares against a no-trait baseline at the dataset level, and Trait-Induced Flip Rate, which counts per-request decision flips across traits. A representation-level analysis finds that traits perturb safety representations within a low-dimensional subspace, motivating Trait-Invariant Safety Tuning (TIST), a self-distillation framework that aligns trait-conditioned behavior with the model's own no-trait behavior. TraSN, an instantiation that enforces invariance only inside the identified trait subspace, improves trait invariance and harmful-request refusal while leaving general capability intact.

Locating and Controlling Implicit Personalization in Large Language Models

Yueru Yan, Siqi Wu, Thai Le Language models shift their outputs in response to implicit demographic cues even when no identity is stated, a behavior documented before but not tied to what happens inside the model. Comparing matched cued and neutral conversations across five models reveals a localized internal activation signal that tracks the resulting changes in recommendations, with correlations up to r=0.87; when several cues co-occur their internal signals largely add together even though the output changes do not. Ablating the signal for one cue suppresses its influence, often more effectively than prompting the model to ignore demographics, while leaving general benchmark performance mostly intact, though cleanly removing one attribute without disturbing co-present ones remains highly model- and attribute-dependent.

Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion

Adrian Rauchfleisch, Andreas Jungherr cross-listed Regulators increasingly require disclosing AI involvement in public communication, but whether such labels blunt an AI system's influence is untested. In a preregistered experiment, 1,500 UK adults held a short conversation with an identical persuasive chatbot on one of 60 policy issues, randomized to receive no disclosure, a prominent disclosure that they were talking to an AI, or that plus disclosure of the chatbot's persuasive intent and instructions. Attitudes shifted 12.6 points on a 100-point scale with no disclosure and 13.1 with the AI-identity label, while adding the intent disclosure cut the persuasive effect roughly in half, to 6.3 points, and also made participants rate the campaign's methods less acceptable and back stronger penalties. The authors argue rules focused on what a system is need to address what it is trying to do.

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

Guang Yang, Fengchen Liu, Alex Wang, Homa Hosseinmardi, Amir Ghasemian cross-listed State-aligned distortion is documented in China-origin text-based language models, but not systematically in multimodal ones, so the authors build a 200-entry balanced benchmark over ten politically sensitive topics plus a seven-variant visual-abstraction probe and run nine vision-language models (seven China-origin, two not) across four elicitation paradigms and two prompt languages for 21,708 trials, auditing six dimensions with two frontier LLM judges validated against human experts. Prompting in Chinese roughly triples the odds of state-aligned framing within every model, China-origin models reframe 1.6-3.2 times more than non-China ones, and the effect is gated by recognizing the depicted subject rather than pixel detail — persisting even for silhouettes of iconic images. Across four Qwen generations, state-aligned framing rises while explicit refusal falls, meaning censorship migrates from a visible refusal to fluent reframing that strips users of the signal that information was withheld.

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu et al. cross-listed Tool-using language model agents are vulnerable to indirect prompt injections hidden in environment state, but existing security studies depend on hand-built or reused environments, stochastic LLM-simulated tools, and injection points fixed in advance. ToolHazard synthesizes executable stateful environments automatically through an Environment Simulator, an Attacker Agent, and a User Simulator, discovering viable injection points, generating environment-specific payloads, and constructing state-grounded long-horizon tasks that can scale with more seed domains and compute. The derived ToolHazard-Bench exposes substantial agent vulnerabilities and shows that injection timing and placement materially change attack effectiveness, while alignment data generated by the framework improves security on both ToolHazard-Bench and AgentDojo without degrading benign task utility.

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

Haokun Lin, Kaijie Zhu, Haobo Xu, Yichen Wu, Zhichao Lu, Qingfu Zhang et al. Small language models are built either by training compact models from scratch or by compressing larger pre-trained ones through pruning, quantization, or distillation, but how those routes affect trustworthiness rather than raw accuracy is largely unstudied. The paper evaluates fairness, robustness, privacy, and ethics across both construction paths. Quantization preserves trustworthiness far better than pruning, and quantizing an already-reliable large model yields small models that are more trustworthy and more adaptable than small models trained from scratch, with knowledge distillation from a trustworthy teacher adding further gains.

No One to Blame: A Framework of Constitutive AI Unaccountability

Long Hoang Nguyen, Eva Sp\"athe, Sebastian Lins, Ali Sunyaev cross-listed Most scholarship frames AI accountability gaps as barriers surmountable through better standards, transparency and institutional reform. The counter-argument advanced is that certain configurations of actors, systems and institutions make accountability conceptually unachievable no matter the effort — termed constitutive AI unaccountability — developed through a concept-centric literature analysis, a secondary analysis of 27 expert interviews with technical, legal and sociotechnical AI professionals, and application to an open-source agentic AI system. The result is nine categories and 20 themes across structural, technological and normative clusters, connected by eight directed interdependencies and operationalized as a 20-question diagnostic that detected 17 of 20 conditions in the case study, including an inverted anthropomorphism configuration where the AI agent was the only identifiable actor.

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu et al. cross-listed Agents that load third-party skills expose two control points to untrusted publishers — the natural-language description used for selection and the instruction body used for planning — and prior work has examined these attack surfaces mostly in isolation. Convergent Detour Hijacking chains them in a text-only, runtime-independent attack: a description establishes relevance so an attacker-controlled coordinator gets selected alongside legitimate skills, and a matching body reuses that same rationale to invent dependencies that pull extra benign skills into a costly detour before rejoining the correct route, so the task still completes. Across 491 held-out tasks on DeepSeek-V4-Pro the coordinator is chosen in 80.02% of cases, and among completed runs where it was selected token consumption rises 66.91% and end-to-end runtime 92.45% with task completion essentially unchanged, showing that a correct final answer says nothing about trajectory integrity or cost safety.

Regulatory Approval Is Not Enough: Gaps in Trustworthy AI Reporting in FDA-Cleared Medical Devices

Ahmed M Salih, Oliver D\'iaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh et al. cross-listed Asks whether the public paperwork behind cleared medical AI actually lets an outsider judge whether a device is trustworthy. Screening 1,105 US Food and Drug Administration (FDA) summary reports for AI/ML-enabled devices from 2021-2025 and manually reviewing the 519 that qualified, the authors coded each against the six FUTURE-AI principles: fairness, universality, traceability, usability, robustness, and explainability. Nearly a quarter of reports documented none of the six principles and no report covered all six, with robustness most common (57.6%) and traceability (8.3%) and explainability (3.5%) the largest gaps; neither clearance year nor clinical domain predicted better reporting, leading the authors to conclude that regulatory clearance is not a proxy for trustworthiness.

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan et al. Parameter-efficient safety alignment by prompt tuning usually attaches one global safety prompt or selects prompt modules externally, making it hard to hold a consistent refusal boundary across risk categories while still tailoring responses and avoiding over-refusal of benign inputs. HiRoute splits those jobs: a lightweight hierarchical router trained on frozen large language model (LLM) representations detects harmful intent and predicts multi-label risk scores, after which preference optimization with alternating gradient updates learns a shared coarse-grained prompt plus a set of fine-grained prompt experts, all as continuous embeddings with the backbone and router frozen. At inference benign inputs bypass the safety branch entirely while risky ones get the shared prompt plus a router-weighted mixture of risk-specific experts, and across three instruction-tuned models this raises safety rates on multiple benchmarks while reducing over-refusal and preserving general-task performance.

Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice

Ziqi Zhao, Jialin Lu, Junjie Shan, Junyuan Zhang, Shuya Yang, Ka-Ho Chow Vertical Federated Learning (VFL) lets organizations holding different features about the same entities train jointly, with the initiator able to hide the learning task and contributors able to keep their data local — an asymmetry that also lets a malicious contributor plant a backdoor during training and trigger it at inference. Re-examining the literature's near-perfect attack success rates, the authors find that most published VFL backdoor results collapse once unrealistic prior knowledge is removed and evaluation is done under practical constraints, a gap that persisted because of weak evaluation protocols. They redefine the threat models under realistic assumptions, propose workable attack workflows, and release BVBench, a backdoor-centric benchmark preloaded with state-of-the-art baselines for fair comparison.

Branch and Bound for Relational Verification of Neural Networks

Kota Fukuda, Zhenya Zhang, Guanqin Zhang, Jianjun Zhao Relational specifications such as global robustness require reasoning about how a network behaves across multiple inferences at once, which is much harder than verifying a single trace property, and convex over-approximation methods used so far are incomplete and raise false alarms. SaBRe adds a branch-and-bound loop that repeatedly splits the problem until every sub-problem is verified, with the twist that it splits relational neurons rather than individual ones and picks which relational neuron to split using the dual formulation of the verification problem, choosing the split expected to tighten the abstraction most. Across 817 verification problems on ACAS Xu, MNIST-F, MNIST-C, CIFAR, and GTSRB, it solves more instances and verifies faster than the baseline approaches.

A Probe Direction Is a Property of Its Prompt

Valentin No\"el Linear probes for whether a model senses it is being evaluated are built by contrasting activations on prompts that announce a test against prompts that do not, and the resulting separation score gets compared across models and correlated with scale. Holding the task text fixed and varying only which announcement phrasing is used, the score — and even the sign of its trend with model size — tracks the prompt rather than the model, to the point that two published studies reporting opposite scaling trends are both reproduced from a single experimental design by prompt choice alone. A variance decomposition puts only a small share of the reported number on the model under study, with most of the rest in model-by-prompt interaction, so gathering more evaluation items cannot fix the measurement while varying prompts can; a control further shows a direction carrying no evaluation information at all reproduces much of each published score from surface form. The authors argue single-prompt designs cannot support cross-model comparison and give the number of prompts a defensible one needs.

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

Valentin No\"el Ablation studies of sparse autoencoder latents must pick one token at which to measure the effect, and the convention — measure where the latent fires hardest — hands that choice to the dictionary being evaluated rather than the experimenter. Comparing two sparse autoencoders Google released for the same model and matching latents by decoder similarity, even near-identical pairs select different measurement tokens for a large fraction of latents, so the standard protocol routinely compares dictionaries at different places. Training six autoencoders from one initialization so a latent means the same thing in each, the variance attributed to dictionaries disagreeing about a latent falls from 7.6% and 11.9% to roughly zero once all are measured at a common token, and the disagreement about where to measure grows rather than shrinks across a sixteenfold range of corpus sizes. The fix is a one-line change plus a reporting requirement, which the authors apply in an audit of five published papers.

Synthetic Persona Pretraining: Alignment from Token Zero

Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui et al. Alignment and the assistant identity are normally introduced only after pretraining, when behavioral priors have already formed, which arguably leaves values as a thin overlay. Synthetic Persona Pretraining (SPP) instead installs the intended persona from the first token: pretraining documents are annotated with value-aligned first-person reflections derived from a normative constitution, the model is pretrained with standard cross-entropy on both documents and reflections, and post-training on dialogue data then binds the persona to the assistant identity. Across models up to 3B parameters trained on 500B tokens, the method improves constitution following and jailbreak robustness and lowers misalignment on out-of-distribution moral dilemmas at no capability cost, and the benefit grows with pretraining budget while largely disappearing if the intervention is applied only at the end of pretraining.
4 more specialized papers

Reinforcement Learning 22

Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet

Eliseo Curcio Reinforcement-learning post-training now dominates language model development, but its power draw on GPUs has not been characterized, and datacenters still manage that power with workload-blind static caps and reactive throttling. Group relative policy optimization (GRPO) training was instrumented with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s, yielding over 380,000 samples used to train a proximal policy optimization meta-controller that adapts the workload's own generation parameters to measured power. On the 7B trace the controller cuts power-limit violations by 89.8% while raising token output 18.1% and tokens per megawatt-hour by 26.2%; the same controller family produced null results at 72B, diagnosed as the group-size actuator losing authority under model sharding, and a rebuilt controller using generation concurrency instead delivered 35.7% more output than a static safe baseline at about 2.3% violations. Measured over realistic 30-second windows a composed 16-GPU fleet shows zero violations with peak demand at half of nameplate, suggesting roughly twofold oversubscription is feasible subject to operator validation.

Dynamics Models for Offline Hyperparameter Selection in Real-World RL

Jordan Coblin, Han Wang, Martha White, Adam White cross-listed Deploying reinforcement learning on physical systems stalls on hyperparameter selection when no simulator exists and online trials are expensive. Calibration models trained on offline logs can approximate environment dynamics well enough to pick hyperparameters in advance, but had only ever been tested in simple simulated settings; here they are applied to a real municipal water treatment plant, with several approaches evaluated including a k-nearest-neighbors model using a Laplacian distance metric on high-dimensional, non-stationary sensor streams. The models produce realistic long-horizon rollouts and recover meaningful hyperparameter sensitivity trends, and further experiments cover scaling to year-long datasets, choosing fine-tuning learning rates for pre-trained agents, and robustness under distribution shift.

Let it Cook: Learning to Wait in Sequential Decision Making

Christopher Watson, Arjun Krishna, Dinesh Jayaraman, Rajeev Alur cross-listed Sequential decision-making agents conventionally sense and act at every timestep, yet tasks like brewing coffee contain stretches where the environment evolves fine on its own and constant monitoring wastes compute or attention that could go elsewhere. The proposal trains a waiting policy that chooses where to pause and for how many timesteps, deliberately forgoing sensing during the commitment, formalized as minimizing sensing and decision frequency subject to not degrading task performance and optimized with reinforcement learning over lexicographically ordered objectives. Across 4 discrete-state household tasks and 3 continuous-state environments the method learns waiting behavior and can retrofit it onto pre-trained policies, sometimes waiting for more than 50 percent of the task duration without losing task performance.

Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning

Zijian Zhao, Sen Li cross-listed When a deployed reinforcement learning system's objective changes — a new reward for traffic signals or fleet dispatch — successor features with generalized policy improvement let a single agent recombine its library of old policies into a new one that is provably no worse than any library member. Multi-agent practice copies the recipe by letting each agent recombine independently, and the authors prove this per-agent composition can yield joint behaviour strictly worse than every policy in the library, because recombining teammates changes each agent's environment and invalidates the values it relies on. Since the only unconditionally safe fixed alternative, synchronized composition, cannot assign different goals to different agents, they propose MA-USFA, which pairs universal successor feature approximators conditioned on teammates' objectives with an upper-level composer that picks each agent's library entry and supplies the cross-agent correction; it is trained once over the objective distribution and used at deployment with no per-task adaptation.

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou et al. cross-listed Post-training on tasks without a checkable answer increasingly uses rubrics graded by a language model judge, but a rubric is an incomplete proxy for quality and a policy optimized against it long enough learns to exploit the gap. Training Qwen3-8B with Group Relative Policy Optimization on medical and science rubrics and scoring out-of-distribution benchmarks with both the training judge and a stronger gold judge, the authors watch the two curves separate — the gold score peaks and then drops by 3 points on HealthBench-Hard and 22 points on ResearchQA while the training score keeps rising, a pattern fixed judge bias cannot explain. Their fix, Rubric Dropout, randomly removes a subset of criteria before computing each step's reward (shared across a rollout group so group-relative advantages stay comparable, with evaluation always on the full rubric), lifting out-of-distribution gold scores at every matched checkpoint with a broad 30-50% sweet spot, while reweighting criteria by usefulness did worse than no intervention.

GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs

Kai Yang, Jingwei Xu, Wanyu Wang, Kai-Yuan Guo, Zhenbo Yu, Yi Wang et al. cross-listed On-policy rollout methods such as GRPO dominate language model post-training but bring instability, loss of unrelated capabilities, and steadily inflating response lengths. The authors introduce Principal-Subspace Overlap, a dimension-corrected measure of how much each individual rollout update aligns with the dominant singular subspaces of the pretrained weights, and find that although average overlap is low, transient spikes in overlap tend to precede performance degradation. GCPO prevents those excursions structurally by applying hard bilateral orthogonal projections that confine updates to the complementary subspaces; across math, code, and tool-use tasks on Qwen3-8B and GLM4-9B it beats GRPO, DAPO, and GSPO by up to 2.37 points over the strongest baseline while preserving general capabilities, removing length inflation, and stabilizing policy entropy.

Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models

Shukrullo Nazirjonov, Sai Prasanna, Anna Manasyan, Georg Martius cross-listed Object-centric representations that decompose a scene into slots are claimed to make world models more sample-efficient and better at generalizing, but prior object-centric world models took the slot encoder as given and tested only in-distribution. A controlled study of these models for visual model-predictive control varies representation quality and distribution shift against scene-centric baselines. Planning success correlates with unsupervised slot-quality metrics (FG-ARI, mBO) but the gains saturate once slots are well bound; with good slots the auxiliary proprioception inputs and masking priors earlier methods relied on become unnecessary, and under unseen shift the object-centric model is more robust overall than the end-to-end LeWM, though DINO-WM on similar frozen pretrained features stays competitive — suggesting pretrained features carry much of the robustness.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou et al. Training conversational agents with multi-agent reinforcement learning usually means simulating the human side with one large language model, and that setup fails to generalize because of what the authors call simulator collapse: the simulator is mode-collapsed, so the policy learns narrow strategies that exploit its dominant mode and transfers poorly to other simulators and to real people. Two remedies are offered — Verbalized Sampling, which at inference time samples from a verbalized response distribution to widen simulator behavior, and Co-Training, which optimizes the policy against a population of trainable simulators. On Persuasion for Good, τ²-bench, and CooperBench, verbalized sampling raises held-out success by up to 9% and co-training pushes the gain to 14%, with a human study showing similar improvements and both methods preserving the policy diversity that single-simulator training destroys; the SCOPE framework is released for population co-training.

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

Ebenezer Gelo (University of the Witwatersrand), Geraud Nangue Tasse (University of the Witwatersrand), Steven James (University of the Witwatersrand), Benjamin Rosman (University of the Witwatersrand) cross-listed Safe offline reinforcement learning normally assumes a per-step cost label for every transition, whereas real supervisors typically give only trajectory-level stop-feedback: a single binary flag at the first unsafe transition with no indication of which steps were to blame. Framing this as temporal credit assignment, Redistribution-based Cost Inference uses return decomposition to spread the sparse signal into dense per-step costs before training a constrained policy on the augmented data, with a proof that return-equivalent redistribution preserves both the feasible policy set and the optimal Lagrangian of the constrained Markov decision process, making the transformation lossless while conditioning the cost critic better. Highway driving and robotic manipulation experiments show substantially lower constraint-violation rates than sparse and classifier-based baselines, holding up under mixed dataset compositions and noisy labels.

SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization

Yuanyu Li, Jintao Xu, Zijiang Liu, Yongzhi Qi, Ningxuan Kang, Jianshen Zhang et al. cross-listed Neural combinatorial optimization trains by sampling many candidate solutions in parallel, but preference-optimization methods anchor entirely on the single best sample (discarding quality and structure information from the rest, termed gradient signal polarization) while mean baselines weight all peers equally (letting near-duplicate solutions dominate, termed baseline redundancy). SSPO scores all sampled solutions jointly through a dissimilarity-weighted leave-one-out baseline that upweights structurally distinct peers, using zero-parameter solution embeddings derived from the encoder's existing node representations. It improves consistently over both best-anchor and uniform-weight baselines on traveling salesman, electric facility location, and job-shop scheduling benchmarks, with a head-to-head against uniform RLOO confirming structure-aware weighting as the main driver, and the facility-location policy is deployed in production at JD.com.

When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide

Binshuang Li Off-policy evaluation is meant to tell an organization what a targeting rule would have earned before it is deployed, but the deployable rule is a deterministic top-k policy under a budget, which removes all averaging over actions and exposes the estimate directly to weak overlap. Six estimators are benchmarked across five datasets and two known-effect sweeps, with mechanisms checked against a non-simulated paired reference. Three findings emerge: overlap is governed by whether the logging policy assigns probability to the target's chosen actions rather than by how sharp the logger is, cross-fitting the outcome model does not cure the optimizer's curse when the rule is fit on the evaluation data (honest policy-level splitting changes the estimand instead), and propensity-estimation error is the single largest source of degradation, hurting inverse propensity scoring more than any other stress while leaving doubly-robust estimates almost unchanged.

Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection

Pu Li, Tao Tan, Hong Xie, Xiaoyu Shi, Mingsheng Shang Large discrete action spaces inject randomness into Q-value estimates, and that randomness pins each standard remedy for overestimation bias to one side: coupled estimation, where one Q-function picks both the action and its value, stays positively biased because it favors actions whose estimates happen to be inflated, while decoupled estimation with two independent Q-functions stays negatively biased because the two tables disagree more on the same action. Action intersection interpolates between the two by letting the Q-functions share a controllable fraction of trajectory data, applying the coupled update to shared samples and the decoupled update otherwise. Sweeping the sharing fraction moves the estimation bias continuously from underestimation to overestimation, at a granularity that can be made arbitrarily fine, and deep reinforcement learning experiments report large margins over several state-of-the-art baselines while tabular experiments explain the mechanism.

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen et al. Joint-embedding predictive architectures (JEPAs) predict in latent space rather than pixels, which suppresses nuisance appearance but does not guarantee that visually perturbed observations encode to the same state. Borrowing the bisimulation criterion — two observations are the same state only if their action-conditioned consequences agree — Action-Conditioned Predictive Consistency (ACPC) measures how far a clean history and a perturbed view of it drift apart when rolled forward under identical action sequences, and the authors prove this divergence upper-bounds both perturbation-induced multi-step prediction error and planner cost. Two aggregate measures follow: the Invariance Radius summarizes clean-versus-perturbed rollout spread, and the Separation Rate checks that genuinely different states stay distinguishable. On four visual control tasks, pairwise ACPC predicts the perturbation-induced changes in prediction error and planning cost, with the joint screen transferring across tasks on LeWM and showing similar trends on the differently architected PLDM.

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

Yubo Zhang, Xinhong Ma, Zezhong Tan, Ziqiang Dong Group Relative Policy Optimization (GRPO) derives its learning signal from reward differences inside a rollout group, so a group in which every sampled response is wrong carries no gradient information; privileged self-distillation supplies dense token-level supervision in that case, but leaving it on permanently makes the policy imitate a biased low-variance teacher whose direction can conflict with reward improvement once the model is competent. I-SDPO makes a single routing decision per input instance shared across its rollout group: all-incorrect groups train with the self-distillation objective while any-success groups fall through to GRPO, so imitation is used only where relative rewards are uninformative. A local analysis shows when teacher and reward directions align and that a non-vanishing distillation weight leaves an optimization bias floor, whereas the routing rule automatically tapers teacher influence as success probability rises without a hand-tuned schedule. On SciKnowEval, average mean@16 accuracy rises from 56.67% under GRPO to 70.31%, with the best result in all four scientific domains and a peak domain gain of 18.24 points.

The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use

Joyjeet Singh When latent world models fail at long-horizon planning, the usual diagnosis is that the learned predictor degrades over time; a reproduction of LeWorldModel on the TwoRoom environment argues the planner's cost function is the actual bottleneck. The predictor's imagined state 75 steps ahead is still only 0.189 as wrong as assuming a frozen world, and a ridge probe recovers position from the frozen embedding at R-squared 0.9922, yet the cross-entropy-method planner's squared latent distance correlates with true distance at only r = 0.426, saturating around eighty arena units and decreasing past 120 so that moving away from the goal can look cheaper. Swapping only the objective, with no retraining and no GPU, raises goals reached at offset 100 from 26.0% to 98.0%, matching the offset-25 rate and hitting 92.0% on under a third of the search budget. The effect is not a reimplementation artifact — it appears in the released weights, and across four checkpoints long-horizon success rank-orders with metric quality and inversely with prediction accuracy, with the best-planning cost head being one that learned reachability rather than proximity.

TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures

Orkun Irsoy, Leman Akoglu, Osman Yagan In networks where an overloaded node sheds its load onto neighbours, a single failure can cascade system-wide, and deciding how to spread a fixed capacity budget across nodes is hard because the survive-or-fail objective is piecewise constant and non-differentiable. TANGCO trains a graph neural network policy through the cascade simulator with policy-gradient learning anchored by a heuristic, and is tested on five synthetic graph families plus five real power, road, air, and Internet networks. It beats the best of four hand-designed heuristics on all 450 synthetic instances and on 40 of 45 real-network conditions, with robustness gains from 1.6% to 246%; a pre-trained variant allocates on unseen real networks with no per-target training, ablations without the graph network fall back to heuristic-level performance, and inspecting the learned allocations yields an improved closed-form heuristic.
6 more specialized papers

Vision 21

Gaze Target Estimation Anywhere with Concepts

Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle et al. cross-listed Working out what a person in a photograph is looking at currently requires multi-stage pipelines fed explicit head bounding boxes and pose estimates, so detection errors cascade and there is no way to specify the subject in natural language. Promptable Gaze Target Estimation reframes the task end-to-end: a text or visual prompt such as "the boy in the red shirt" or a point coordinate selects the subject, and one model handles localization, in-frame versus out-of-frame presence, and gaze heatmap prediction together. A scalable data engine produced Gaze-Co, a benchmark of 120K prompt-annotated image pairs, on which the transformer-based GazeAnywhere model reaches state-of-the-art results across multiple benchmarks including a difficult out-of-domain clinical dataset, and the code is publicly released.

From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou cross-listed Removing reflections from video shot through glass has lagged behind the single-image case because there is no paired training data, no temporally consistent model, and no benchmark. The S2R-Synthesis pipeline closes the data gap by augmenting scene structure with physics-grounded glass effects — roughness blur, thickness ghosting, and varying reflectance — and rendering the reflected video with a trained video diffusion renderer. S2R-Removal then adapts a pretrained video diffusion prior via reflection-aware latent adaptation and pixel-geometric refinement to recover the clean transmission in a single denoising step, reportedly faster than even non-diffusion baselines while setting state-of-the-art scores on the new S2R-Bench and on public image benchmarks.

Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation

Yuanmin Huang, Chen Chen, Geng Hong, Xiaoyu You, Hui Xue, Zhenxing Qian et al. cross-listed Proprietary text-to-image diffusion models ship as hosted services and downloadable checkpoints, leaving owners with little recourse when leakage or unauthorized fine-tuning is disputed. The proposed fingerprint exploits collapsed generation, the phenomenon where certain input conditions yield nearly identical images across random seeds; because collapse-prone conditions are an intrinsic, model-dependent property of the learned generation process, reproducing the source model's collapse behavior serves as ownership evidence without embedding any watermark. Verification works both with white-box pipeline access, injecting optimized continuous embeddings, and through black-box natural language prompts to an API, and experiments on UNet- and transformer-based diffusion models show low confusion between source models, survival in fine-tuned derivatives, and robustness to common and adaptive obfuscations at a modest query budget.

LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration

Enhuai Liu, Yunke Wang, Yutong Wang, Changming Sun, Chang Xu cross-listed Sampling from video diffusion transformers is dominated by self-attention over long 3D token sequences, and existing training-free sparse-attention methods chase aggressive sparsity at disproportionate cost to attention fidelity. LoSA inverts the target: it fixes a 99% retained-attention-mass threshold instead of a sparsity ratio, measures exact block attention masses at one early dense denoising step, keeps the smallest key/value block set per head and query block that meets the threshold, and reuses those frozen indices for every later step — leaning on the observations that roughly 40% of block interactions are removable and the high-mass support stays stable across steps. On Wan2.1-1.3B it gives a 1.36x speedup for a 0.06-point VBench Overall drop, and combined with feature caching it reaches 3.2x on HunyuanVideo at a 0.02-point drop, against 0.32 points for the strongest sparse baseline at comparable speed.

How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging

Yiheng Xiong, Luisa Gall\'ee, Daniel Santak Wolf, Heiko Hillenhagen, Michael G\"otz cross-listed Deploying unsupervised domain adaptation clinically means choosing both an algorithm and which trained checkpoint to ship, yet the deployment domain is unlabeled so nothing can be validated on it directly. The study evaluates the complete pipeline — adaptation plus label-free model selection — across eleven clinically relevant cross-domain scenarios drawn from nine medical imaging datasets, ten adaptation algorithms and 13 label-free validators, covering more than 80,000 trained models. A capable adapted model usually exists, but no validator reliably identifies it, leaving a large and structural performance gap to the best available checkpoint; ensembling and a small target-labeling budget each narrow that gap without closing it.

Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel cross-listed Class activation mapping (CAM) turns a model's internal evidence into a heatmap over image regions, channels, tokens, or patches, and has diversified enormously since its 2016 introduction for globally-average-pooled convolutional classifiers. Surveying a strict corpus of 57 method-centered papers, this review organizes the field by attribution mechanism, architectural dependence, and evaluation objective, covering gradient-based variants, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization, transformer token attribution, causal and debiasing approaches, and foundation-model-era techniques built on CLIP, DINO, and SAM. The stated trend is a move from explaining one class score in one low-resolution layer toward comparative, multi-layer, probabilistic, token-aware explanations, alongside persistently fragmented evaluation where faithfulness, localization, robustness, cost, and human trust are each measured differently.

Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex

Nils Leutenegger cross-listed Studies comparing learning rules by how brain-like their representations are typically train small networks on 32x32 CIFAR images, because biologically plausible rules such as feedback alignment, predictive coding, and spike-timing-dependent plasticity do not scale, then compare against brain data modeled at far higher resolution. The reported result that untrained or locally trained networks match or beat backpropagation in early visual cortex turns out to hinge on evaluation resolution: the V1 gap between untrained and backprop-trained networks grows monotonically from -0.001 at 32 pixels to +0.044 at 224 pixels, holding in human fMRI, directionally in macaque electrophysiology, and for ResNet-50 and Swin-Tiny. Four candidate explanations are ruled out, including three via interventions that keep convolutional weights bit-identical, and the effect is traced to image detail rather than pooling; a single per-image luminance scalar correlates with the V1 representational dissimilarity matrix about as well as an untrained network, bounding what such comparisons can resolve.

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

Ebenezer Tarubinga cross-listed Pseudo-label filtering rules for semi-supervised segmentation were designed for the weak, under-confident teachers of the ResNet era, but a self-supervised foundation encoder such as DINOv2 saturates its confidence — 98 percent of Pascal pixels score at least 0.95 — so an adaptive cutoff floods the retention mask and self-training decays into confirmation bias. CW-BASS v2 reads the teacher's confidence regime instead of committing to one rule: it measures on a held-out slice how reliable the teacher's confident set is, filters strictly when that reliability meets the pre-existing operating threshold, and otherwise falls back to a self-adaptive floor that provably bounds retention away from one. Across six DINOv2 teachers the gate makes the correct strict-versus-floor call without tuning on mean intersection over union (mIoU), recovering the UniMatch V2 operating point on saturated benchmarks (Pascal VOC 1/8 at 87.4 against a reported 87.9) and gaining 1.5 mIoU on ADE20K, where the confident set is less trustworthy.

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

Lei You cross-listed Perturbation-based explanations report how much a model's prediction moves when inputs are altered, but magnitude alone conflates responses that support the factual-versus-counterfactual difference, responses that oppose it, and responses that spike mid-path yet cancel at the endpoint. DECAF (Decomposition of Evidence, Contradiction, And Fragility) tracks how the contrast develops as paired inputs are progressively revealed and routes the response into evidence, contradiction, and fragility components that sum exactly to the original magnitude and are unique under endpoint-relative axioms. In an audit of 72 ImageNet-9 models, the largest DECAF component agrees with independently measured behavior in 96.4% of cases versus 35.0% for magnitude alone. Changing only the reveal path inflates total response by nearly 80% while evidence stays flat and fragility grows over fourfold, and short forward-only trajectories beat the tested attribution baselines on FunnyBirds and ImageNet-1k while matching a gradient baseline on a 1B-parameter DINOv2 at 4.75x lower wall time.

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang cross-listed Caching intermediate features across denoising steps is a standard way to cut diffusion inference cost, but existing reuse policies decide what to cache from local similarity, which the authors show correlates poorly with final image quality because errors propagate and accumulate unevenly along the trajectory. Global-Impact Cache (GCache) derives an upper bound on error propagation, then — since that bound is loose for highly non-convex models — reparameterizes the propagation exponent in Bernstein form and casts policy search as a bilevel problem, with the inner level picking the reuse schedule and the outer level fitting the error-weighting function to actual generation-quality loss. On the Wan2.1 video diffusion model it holds a 2.17x speedup while cutting LPIPS from 0.1095 to 0.0316, and it outperforms prior caching strategies on both image and video generation.
11 more specialized papers

Multimodal 20

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

Shibo Gao, Peipei Yang, Xu-Yao Zhang, Linlin Huang Video anomaly detection splits along a "when-what" line: conventional deep networks localize when an anomaly occurs without understanding it, while LLM-based systems describe what happened but ground it poorly in time. Two systems follow a human-like global-to-local inspection loop — GtS (Glance then Scrutinize) is training-free and uses static and dynamic textual guidance for coarse-to-fine grounding, while an agentic variant teaches a multimodal large language model to call a video-cropping tool, re-inspect densely resampled frames, and revise its own mislocalized hypotheses, trained by cold-start supervised fine-tuning followed by reinforcement learning with a joint answer-grounding reward. Supporting this is VAGU-T, an extended benchmark of 7,567 real-world videos across 21 anomaly categories with human-validated grounding, explanations, question-answer pairs and chain-of-thought tool-calling traces, plus JeAUG, a metric scoring semantic interpretability and temporal precision together. The agentic model is reported as both more accurate and faster at inference than the training-free pipeline, which itself substantially beats other training-free baselines.

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval

Archan Dutta, Vyanktesh Kanungo Text-to-image retrieval has long depended on dual-encoder models trained with contrastive learning, but the arrival of Gemini Embedding 2 — a natively multimodal embedder mapping text, images, video, audio, and documents into one space — invites a comparison against using frontier language models as zero-shot rerankers instead. The study runs both approaches on hard-negative retrieval over Flickr30k, the first head-to-head of its kind. GPT-4.1 and Claude Sonnet 4.6 rank on par with Gemini Embedding 2, though precomputed embeddings remain the better fit whenever latency matters.

CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

Kegeng Tang, Jingbo Wang, Shaogang Ren, Zihao Wang Radiologists judge computed tomography (CT) scans largely by comparing serial studies over time, yet medical foundation models are built around single-study interpretation. To target the task of longitudinal difference reporting, where a model reads two scans of the same patient and writes a report on the interval changes, the authors build a benchmark with patient-level splits to block leakage, plus change-aware metrics that go beyond text-similarity scoring and a physician review of the synthesized reference reports. They compare direct reasoning over the scan pair against an indirect pipeline that writes separate single-timepoint reports and then diffs the text, and release DeltaMed, a baseline trained for direct paired-CT difference reporting.

Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting

Xikai Sun, Kebin Liu, Haotian Wang, Li Liu, Xu Wang, Yunhao Liu cross-listed Multimodal large language models keep video costs down by sampling frames uniformly and sparsely, which throws away the transitions between frames that reasoning about movement, collisions, and cause-and-effect depends on. Motion-as-Prompt recovers dense point trajectories from the full video, picks the most motion-informative frames, and draws the trajectories accumulated between consecutive sampled frames onto the images themselves, so displacement and direction changes become visible to a frozen model with no training or architecture change. On CLEVRER and Something-Something-v2 it raises average motion-reasoning accuracy by 4.2 and 8.9 points respectively for GPT-5.5, without hurting performance on non-motion questions.

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

Haoyu Zhang, Shuoxun Zhang, Peng Ye, Lin Zhang, Jiakang Yuan, Shenghong Yi et al. cross-listed Aerial imagery from uncrewed aerial vehicles stresses multimodal large language models with extreme scale variation, arbitrary camera orientations, and dense objects, and evaluation has been scattered across individual datasets and narrow tasks. UAVQA-Bench consolidates 1,500 human-annotated question-answer pairs from 13 public UAV datasets spanning 6 capability dimensions and 16 tasks in multiple-choice and visual grounding formats, and evaluating open and closed models on it exposes three recurring failures: mismatch between the domain and available tools, unchecked error propagation, and reasoning that does not adapt. The training-free UAV-MAS system responds with a perception engine that routes queries to task-appropriate visual tools, an iterative refinement module that validates intermediate reasoning, and a search mechanism whose depth scales with question difficulty. Paired with a 32B open-source model it reaches 77.0% overall accuracy, 4.0% above Gemini 3 Pro, and the 8B variant gains 8.7% over its base model.

LookBack: Where and How to Score LVLM Responses via Visual Reference Usage

Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim cross-listed Large vision-language models hallucinate against the image as well as inheriting text-level hallucination, and the diagnostics here show that confidence metrics borrowed from text models barely change their selections when the input image is removed — evidence that output-space confidence tracks textual plausibility rather than agreement with what the model sees. LookBack is a training-free scoring method that augments token likelihood with a visual lookback score measuring how strongly each response token refers back to image tokens. Across four benchmarks and three models it consistently improves Best-of-N selection over existing baselines at negligible extra cost.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models

Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang cross-listed Unified multimodal models (UMMs) fold image generation and image understanding into one parameter space, yet evaluation protocols still score the two capabilities separately, leaving no measure of whether the unification actually works. Self-Generative-Understanding (SGU) closes the loop without any new annotations: the model describes an image in text, regenerates a visual context from its own description, and then reasons over that self-generated output, producing an integrated score at zero labeling cost. Even strong unified models frequently fail to reason over contexts they themselves generated, a weakness invisible to separate tests of understanding or generation.

SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

Zile Zhou, Huining Yuan, Weichen Zhang, Xinlei Chen, Xiao-ping Zhang cross-listed Vision-language models (VLMs) remain weak at spatial reasoning, and reinforcement learning (RL) methods that reward only final answers give poor credit assignment to intermediate steps. SCOUT combines a structured chain-of-thought (CoT) format that explicitly represents 3D environmental perception, including depth, with an RL algorithm that issues multi-objective process rewards and a matching advantage estimator so different segments of a reasoning trajectory are scored separately; training uses SCOUT-24k, a synthesized structured spatial CoT dataset. The 3B model gains 16.85% on general spatial benchmarks and 6.3% on complex spatial tasks over its baseline, while the 7B version outperforms GPT-4o by 4.28% and transfers from single-image training to multi-image and video inputs.

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du et al. cross-listed Scientific writing workspaces increasingly ask multimodal large language models (MLLMs) to turn figures into LaTeX TikZ code, a capability existing benchmarks do not isolate. Diagram-MMU supplies 3.7k curated diagrams and 18.3k human-validated questions across six domains, testing diagram-to-code parsing, diagram-to-code editing, and diagram question answering in both direct and agentic settings. Across 12 evaluated models, code generation and editing lag well behind question answering, showing that models reason about diagrams better than they reconstruct them; agentic scaffolding helps parsing and editing while hurting question answering for most models, with Claude-4.6 Opus the exception that improves on all three.

AVA-Encoder: Towards Agent-Native Video Representation Learning

Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li et al. cross-listed Agents that generate cinematic video have no structured way to learn from existing high-quality films, because film content lacks a representation that is both faithful and directly editable by an agent. The Agentic Video Auto-Encoder (AVA-Encoder) encodes a video into a knowledge graph whose hierarchy and state nodes hold structured text and whose linked asset layer holds generated images, audio, and clips, then reconstructs video from that graph; the reconstruction gap is converted into natural-language update directions that pseudo-train the encoding policy offline and can refine the graph at test time. Reported results show a 20.7 percentage-point improvement over the strongest external baseline, and the pseudo-trained shot-level encoder policy beats a hand-tuned policy while using 74.3% fewer system-prompt tokens. The authors release the framework, a reconstruction benchmark, and a dataset of film knowledge-graph representations.

From Visual Widgets to UI Code: Efficient Tool-Grounded Generation

Houston H. Zhang, Tao Zhang, Li Gu, Linfeng Ye, Yuanhao Yu, Xinxin Zuo et al. cross-listed Turning a screenshot into working interface code forces a choice between direct multimodal generation, which hallucinates visible details, and structured pipelines that decompose into components and templates but add orchestration overhead and cannot express designs outside their fixed representation. WidgetGen takes a middle path with selective tool grounding: it extracts observable text and color evidence with tools, reasons about high-level layout and optional charts, and emits executable JavaScript XML directly, with no fixed interface schema. Across six multimodal models and 1,000 held-out widgets it beat both direct prompting and the structured Widget2Code pipeline on most visual reconstruction metrics, with consistent gains in area, legibility, and style, and image-code pairs harvested from its reconstructions improved six open-weight Qwen-family models on every reported metric under supervised fine-tuning.

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Zirui Cheng, Xun Xu, Tiankai Chen, Fady Rezk, Bowen Zheng, Xiaodong Shi et al. Few-shot in-context learning with multi-modal large language models adapts to new tasks without weight updates, but results swing sharply with which demonstrations are chosen, and unlabeled multi-modal data has been hard to exploit for that choice. MAG treats demonstration selection as semi-supervised label propagation on a multi-modal graph in two stages: relevance scores propagate to identify a compact set of high-impact unlabeled samples worth pseudo-labeling, keeping inference cost down, then multi-modal relevance picks the final demonstrations. An ablation finding is that text representations drive the propagation step while both vision and text matter for the final selection, and across eight multi-modal benchmarks the method beats strong baselines in label-scarce regimes under a limited pseudo-labeling budget.

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

Fnu Pramono, John Cai, Sourabh Kulkarni cross-listed When video evidence is occluded or chaotic, a model should say it cannot tell, and TRAPSBench tests this with 1,404 procedurally generated matched physics pairs where one targeted change makes the outcome undeterminable from what is visible. A companion metric, the Penalized Epistemic Calibration Score (PECS), rewards models only for answering correctly when the outcome is knowable and abstaining when it is not; across 16 vision-language models from five families the best score is 0.292. The failure is in expression rather than perception: linear probes recover answerability from hidden states at up to 0.91 AUROC, and steering a single-layer direction causally turns abstention on or off, with the gap replicating across Qwen, Gemma, and LLaVA and models detecting textual impossibility roughly four times more readily than missing visual evidence.

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo et al. cross-listed Vision-language model (VLM) benchmarks typically score how accurately a model describes and reasons about an image, but not how it behaves when the visual evidence is absent or actively misleading. SciFigBench covers 250 human-annotated scientific figures expanded through image transformations, resistance probes, caption-bias probes, and selective blurring into more than 34,000 evaluation setups, scored under an Admittance-Resistance-Inductance framework that asks whether a model admits insufficient evidence, resists misleading context, and infers cautiously from partial information. GPT-5.2 posts the best description quality yet hallucinates content in 96% of cases where the text is unreadable, while the comparably accurate Gemini 3.1 Pro admits uncertainty in 71% of those cases and leads on resistance. The gap indicates that perception and reasoning accuracy do not predict behavioral reliability under uncertainty.
6 more specialized papers

Robotics 12

Towards the Harness of Embodied Agents

Qi Wang, Tianyi Wang, Chengyang Li, Shikun Ban, Yurun Chen, Yizhong Ge et al. Coding agents established that capability depends on the harness surrounding the model, not the model alone, and Thea tests whether the same holds for robots by wrapping robot capabilities as callable tools inside an agentic loop. The physical world withholds two things software provides free, reading world state and judging whether an action succeeded, so the harness adds Scene Graph as Context, a persistent symbolic world representation, and Evaluation as Exit Codes, which detects when an action should terminate, judges success, and diagnoses the cause of failure. Closing that perception-and-judgment loop lets composed tools produce rich emergent behavior and carry long-horizon tasks to completion in real environments.

Self-Evolving Embodied Agents via Skill-Harness Evolution

Peidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Xiaocui Yang, Shi Feng et al. An embodied agent's behavior depends as much on its skills, context, action interfaces, and execution harness as on model weights, yet adapting it usually means collecting data and running fine-tuning or reinforcement learning, while train-free code-centric alternatives assume programmable robot APIs that fixed-interface robots do not offer. SHAPER keeps parameters frozen and instead evolves reusable skills and a context-code harness through rollouts in the target environment, using the same frozen model as both planner and optimizer. Evaluations on VLABench and ESI-Bench, spanning agents with different low-level action interfaces, compare against pure execution, supervised fine-tuning, and test-time-scaling baselines such as verifier-free selection and voting, indicating that skill-and-harness optimization is a workable substitute when training is expensive or unavailable.

Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards

Sim\'on Pati\~no Idarraga, Erick Silva, Rehana Yasmin, Ali Shoker cross-listed End-to-end driving agents post good average scores yet still break elementary traffic rules, because they learn statistical regularities rather than the physical conditions that make a maneuver safe. The proposed neuro-symbolic guard bolts onto the command interface of an already-trained agent, checking each outgoing command against explicit safety rules and swapping in the nearest safe alternative only when a rule fires, so every intervention is traceable and no retraining or extra learned component is needed. Wrapped around TransFuser v6 on the long-tail benchmarks Fail2Drive and Bench2Drive, it raises success rate by 15% and cuts safety-critical collisions by up to 53% while leaving the original driving score intact.

Keep the Future, Drop the Rollout: RIFT for World Action Models

Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li cross-listed World action models pick robot actions conditioned on a predicted future, but generating that future by iterative video rollout is slow at deployment. Paired closed-loop interventions across four such models on all 40 LIBERO tasks show the policies genuinely depend on future values and their positions — masking or reassigning them degrades execution — yet for Joint and Cosmos-2, replaying a single fixed final key/value cache preserves behavior almost exactly (1.7-1.9 cm end-effector displacement error, 97.9-98.2% success), separating cache consumption from cache production. RIFT exploits that gap by using learned anticipation tokens to build the entire future cache in one backbone pass, reaching 98.8% success on LIBERO while cutting action-chunk latency by 68.2-89.1% relative to rollout-based baselines, and 92.9/92.6% on clean and randomized RoboTwin 2.0 scenes.

Foresight Without Seeing: Latent Futures for World Action Models

Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang World action models tie future visual prediction to robot action generation, but the two existing designs trade off badly: models that decode explicit future frames pay for iterative video denoising at inference, while direct-policy models are fast yet give the action module no window into predicted dynamics. ForeWAM runs a single video diffusion transformer prefill over the current visual latent plus stochastic future slots and reuses the resulting layer-wise key-value states throughout action denoising, with dynamics registers supervised by a frozen latent action teacher so the implicit future states encode object motion, contact changes, and task progress. Ground-truth futures and the teacher are training-only, so deployment generates no video at all, and without embodied pretraining the standard and accelerated variants reach 96.7% and 96.9% average success on LIBERO, with 61.6% on LIBERO-Plus.

G0.5: One Autoregressive Stream for Robot Reasoning and Action

Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang et al. cross-listed The standard vision-language-action recipe bolts a separately trained flow-matching action expert onto a pretrained vision-language model, which reduces the language model to a context encoder rather than the thing making decisions. G0.5 instead emits reasoning and action tokens from a single autoregressive transformer decoder under one objective, made feasible by a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary, a native chain-of-thought stream interleaving task decomposition and object grounding with action tokens, and a visual memory module supplying multi-second history through the vision encoder. Sharing weights between reasoning and action means prompts directly steer action granularity, task horizon, and out-of-distribution scene handling with no extra training. Across 7 independent regimes it beats prior systems, including 76.7% on real-world R1lite and R1pro fine-tuning against 53.3% for pi-0.5, 31.4% on the 2025 BEHAVIOR Challenge with a generalist policy, 98.9% on LIBERO, 93.3% on RoboTwin 2.0, and 87.3% on SimplerEnv-Bridge.

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter et al. cross-listed Scaling reinforcement learning to combined locomotion and manipulation is bottlenecked by the slow, manual process of shaping dense rewards. The approach uses Sample-based Model Predictive Control entirely in simulation as an automated, quickly tunable expert to generate large offline datasets; because that data already solves exploration, an off-policy agent can then be trained from purely sparse task rewards, eliminating manual reward tuning, and coupling it with a low-level dynamic stability controller lets the learned policy surpass its optimal-control teacher. Sim-to-real transfer is demonstrated across two morphologies, an arm-equipped Spot quadruped and a G1 humanoid.

Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment

Jean-Pierre Busch, Guido Linden, Jan Bergmann, Lutz Eckstein cross-listed Learned motion planners handle complex traffic scenes well but their opacity complicates explainability and safety assurance for automated vehicles. The proposed hybrid architecture keeps a deep neural network for interpreting scenes and proposing driving behavior, then routes that proposal through an optimization-based supervision layer that checks it and enforces explicit drivability and safety constraints deterministically. The learned planner's behavior is assessed in open-loop studies on real-world urban data, and the paper reports system-integration details for stable closed-loop operation plus results from actual deployment on the authors' research vehicle.

Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling

Takieddine Soualhi (CHROMA), Jacques Saraydaryan (CPE, CHROMA), Laetitia Matignon (UCBL) Deep reinforcement learning (DRL) agents navigating crowds mostly optimize task-centric goals such as reaching the target quickly, and represent social comfort only weakly. The proposed reward models each person's personal space as a radial Gaussian-mixture field derived from Hall's proxemics theory and computes a robot-centric local cost over the robot's field of view, giving a dense and interpretable social signal rather than a sparse penalty. Dropped into established DRL navigation methods and evaluated in simulation across crowd scenarios, densities, and competing reward baselines, it consistently improves social compliance metrics while keeping navigation performance competitive.
3 more specialized papers

Reasoning 10

Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing

Minhan Cho, Jimin Kweon cross-listed Two recent recipes for more reliable large language model reasoning are independently reproduced and stress-tested: RPC, which aggregates token probabilities with self-consistency at inference time, and LCF, which trains projectors that split hidden states into content and logic and then edits the logic part toward a valid region. Neither had been evaluated by anyone other than its own authors, and LCF shipped no public code, so the authors re-run RPC and re-implement LCF before extending both to text-to-SQL, legal extraction, fallacy identification, and precedent grading. RPC reproduces its published grid exactly on the released reasoning paths, but its advantage over plain self-consistency is never significant on the new domains; its largest gain, +2.5 accuracy points on BIRD at 32 samples, flips to -0.25 once the evaluation set grows to 200 items. LCF's logic-validity direction is real but weak (0.82 separability versus 0.95 for a semantic-attribute control), its single positive effect is not significant, and it significantly degrades two of the other three models tested.

Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

Mark Shapiro A dense pretrained model, Qwen2.5-0.5B-Instruct, is split into a prelude, a weight-tied recurrent block, and a coda so its middle layers can run repeatedly and learn an iterative latent update rather than a terminal-answer lookup. Training with intermediate-step supervision teaches the model to compute one task step per loop, and the behavior persists when only final answers are graded; a 6M-parameter adapter over frozen weights matches full 180M-parameter block training (83.8% versus 84.0%), and the learned operation extrapolates to roughly 1.5 times its supervised depth, holding 70% accuracy at depth 18. Against a same-size model trained on written scratchpads, the recurrent version scores 84% versus 72% overall, retains 53% versus 2.5% beyond depth 10, and answers 7.6 times faster. Teaching the inverse rule proved impossible without destroying the installed mechanism, marking a catastrophic-interference boundary.

When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

Utkarsh Bahuguna Self-consistency — sampling many chains of thought and returning the plurality answer — is a standard way to spend inference compute, on the assumption that voting helps on average. On the full GPQA Diamond benchmark of 198 graduate-level science questions, majority voting lowers per-problem accuracy on 56.6% of problems for Qwen2.5-7B and 65.7% for Llama-3-8B, an effect pre-registered on a 151-problem confirmatory split after being spotted on 47 exploratory problems, with all four confirmatory hypotheses passing. An oracle routing each problem to its best sample count marks an upper bound 14 to 17 accuracy points above single-sample decoding, but no verifier-free gate reaches it — neither plurality agreement nor token entropy moves accuracy by more than 0.002 — because confidence does not track correctness here: in Llama's highest-agreement bin accuracy is worse than in its lowest. The authors flag reasoning-native models as untested.

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan Getting language models to catch and repair their own reasoning errors remains unreliable. Self-Fix Step-DPO (SFS-DPO) splits the problem into two reinforcement learning stages: first sharpening step-level reasoning through preference optimization over individual steps, then explicitly training the model to verify and revise those steps. A teacher-assisted variant, SFS-DPO-R, adds written rationales explaining why a step is wrong. Across several base models and both in-domain and out-of-domain tests, both variants beat prior step-level training baselines, and analysis attributes the gain to higher frequency and higher success rate of self-correction attempts.

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Xinmu Ge, Zizhuo Zhang, Yu Huang, Jianing Zhu, Lin Yuan, Wanli Gu et al. cross-listed On-policy distillation is generally assumed to let a student absorb a stronger teacher's knowledge and expand past its own pre-distillation ceiling. Testing that belief through test-time scaling, the authors vary the sampling budget K and track both pass@K and avg@K across several distillation variants, finding that distilled models keep an avg@K advantage at all budgets while their pass@K advantage steadily transfers back to the pre-distillation base model as K grows. A problem-level solvability analysis at pass@1024 shows an asymmetry in which more previously solvable problems become unsolvable than the reverse, supporting the reading that on-policy distillation mainly buys sampling efficiency rather than genuinely new reasoning capability.

Policy-as-logic for robust reasoning over rules

Rahul Nair, Bastian Lipka, Elizabeth Daly Many deployed generative systems must answer natural-language questions in ways that provably respect written rules, from tax codes to airline baggage allowances. The proposed hybrid approach writes the policy in formal logic, uses a language model only to extract facts from the query and ground predicates, and hands reasoning to an answer set solver, so the resulting decisions are interpretable and auditable. Separating extraction from reasoning outperforms both policy-as-prompt and policy-as-code baselines in most cases while using roughly 10x fewer tokens, and stays accurate under perturbations to the input phrasing.

OEIS Open: How many conjectures can language models turn into theorems?

Tom Adamczewski OEIS Open packages 492 genuinely open mathematical conjectures from the Online Encyclopedia of Integer Sequences, formalized in Lean by prior work, into an open-source evaluation harness that any language model can be run against and that is hardened against cheating attempts — previously these conjectures had only been attacked by one bespoke agent. With a minimal toolset and a $50 budget per attempt, models resolve 147 of the conjectures, scoring 30%, and on the 100-conjecture OEIS Open Lite subset the best current model reaches 44% at a $200 budget. Giving models access to 476,000 arXiv papers did not help, and neither did more elaborate agent loops. The authors note the conjectures are of uncertain mathematical significance and mostly little-studied, but that models nonetheless close open research questions autonomously at modest cost.

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

Sen Xu, Wei Wang, Shixi Liu, Jixin Min, Yingwei Dai, Zhibin Yin et al. Test-time compute in large language models is usually spent sampling more candidate solutions, but scoring a whole reasoning trace dilutes decisive errors among routine tokens. CLR (Claim-Level Reliability Assessment) is a training-free framework that condenses each trace into a compact set of decision-critical claims and then spends the compute trying to refute them, exploiting the asymmetry that disproving a claim needs only one flaw while constructing a valid solution needs a flawless path. Across four language models and four reasoning benchmarks at matched budgets it beats both pass@1 and self-consistency; on GPT-OSS-20B with CMIMC25 it raises self-consistency accuracy from 77.50% to 82.19% while using 37% fewer tokens.

Position: Reasoning is a Learnable Rule-Based Process

Rachel Lawrence, Jacqueline Maasch cross-listed Argues that the generative-AI community evaluates "reasoning" without operationally defining it, and that this ambiguity makes the construct validity of reasoning benchmarks unverifiable, so claimed progress toward trustworthy autonomous reasoning cannot be quantified. The authors synthesize the logic and automated-reasoning literature into operational definitions that cast valid and sound reasoning as a learnable rule-based process, and pair them with a checklist of best practices for how reasoning results should be reported.

Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao et al. cross-listed Graph problems make a convenient reasoning testbed for language models because instances can be generated programmatically and scaled at will, but existing graph benchmarks are largely hand-built, narrow in complexity, and do not compare text-based against code-based solving. A five-stage semi-automatic pipeline uses a language model to generate task descriptions, graph instances, reference solutions, loading scripts, question forms, and grading scripts, with human checks at quality-control points, expanding coverage along graph size, task complexity, task description, graph loading, and task source. The resulting 202-task benchmark exposes model limitations invisible to earlier suites: fine-tuned models fail to generalize to GraphGym, and retrieval augmentation helps textual reasoning but not consistently code-based reasoning.