Monday, September 14, 2026

313 papers cs.AI · cs.LG · cs.CL ← 2026-09-112026-09-15 →

Jul Aug Sep

Highlights

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Highlight Agents Wenhui Chen, Shiwen Cheng, Hao Dong, Chenda Duan, Ruixiang Feng, Zhong Guan et al. Co-work agents run long workflows of information gathering, tool use, coding, and file manipulation across many model calls, so cost and latency accumulate and many steps demand state tracking and recovery rather than frontier-scale reasoning. Occamy-1.0 further trains the post-trained Qwen3.6-35B-A3B checkpoint using execution-grounded data and environments, replayable long-horizon trajectories captured across multiple harnesses, and staged post-training to build and consolidate execution skills. Across co-work benchmarks it is among the strongest comparably sized models and competitive with much larger frontier systems on several tasks, and under the authors' pricing protocol it sits at the low-cost knee of the cost-performance Pareto frontier across four benchmarks; the weights and a subset of the training data are released.

Co-work agents run long workflows that mix information gathering, tool use, coding, and file edits over many model calls, so cost and latency add up across an episode even though most steps need reliable state tracking and recovery more than frontier-scale reasoning. Occamy-1.0 further post-trains Qwen3.6-35B-A3B, a mixture-of-experts model with 3B active parameters, on execution-grounded tasks to make that capability cheaper to deliver; the weights and a subset of the training data are released.

  • Each training task is a verified executable contract made up of a request, an initial environment state, tool definitions, and a hidden grader, and trajectories are recorded token by token through a proxy across three agent harnesses (OpenClaw, Hermes, and Accio Work), with every context compaction marked as a segment boundary so the episode's single reward reaches the right policy tokens.
  • The recipe trains a long-horizon Marathon Expert with SFT followed by HDPO and a reward for using fewer steps at equal accuracy, plus a shorter-horizon Sprint Expert with SFT, then merges the two with a uniform model soup and refines the result with SAO reinforcement learning, using about 15K SFT trajectories totaling 403M tokens.
  • On Claw-Eval it reaches an 82.2 average score (up from 69.5 for the base model and just above GPT-5.6 Sol at 81.8), leads its size class on WildClawBench at 49.16, and also improves on tool-use benchmarks outside the RL task pool, raising AutomationBench strict pass rate from 7.5 to 27.6.
  • Compared with the base model on Claw-Eval, it uses 19.5% fewer tokens, spends 46.4% less wall-clock time per trace, and cuts the timeout rate from 9.88% to 2.18%, which puts it at the low-cost knee of the cost–performance Pareto frontier averaged over four benchmarks.
  • Large gaps to frontier models remain on OfficeQA Pro (48.1 vs. 74.4 for GPT-5.6 Sol) and Terminal-Bench 2.1 (59.0 vs. 88.8), the merge step causes a small Claw-Eval drop relative to the stronger expert, and the cost comparison prices measured token usage at OpenRouter rates rather than measuring full deployment cost.

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

Highlight Agents Mohsen Arjmandi Agentic coding systems pair a model with a harness of tools, prompts, and control flow, and practitioners assume a vendor's native harness solves more tasks with its own model. The authors test this with paired same-model contrasts on a private, contamination-controlled suite of 256 repository and post-cutoff contest tasks, running 80 tasks under claude-agent-sdk versus deepagents on claude-opus-4-8 and under the openai-codex SDK versus deepagents on gpt-5.5. Neither contrast shows an average advantage for either harness, with differences near 1 percentage point and confidence intervals spanning zero, though for Opus the native harness trails by 9 points on repository tasks and leads by 24 points on contest tasks in a post hoc split the authors say needs replication. The neutral harness cost 1.2 to 1.6 times more per solved task by observed usage, but missing usage records leave the billed ordering unresolved.

Teams often assume that a vendor's own agent harness gets more out of that vendor's model than a portable harness does. The study tests this by keeping the model fixed and changing only the harness: claude-agent-sdk versus deepagents on claude-opus-4-8, and the openai-codex SDK versus deepagents on gpt-5.5. The tasks come from a private suite built to avoid training-data contamination, mixing repository tasks with contest problems published after the models' training cutoffs.

  • Each of the 80 selected tasks (61 repository, 19 contest) ran twice per harness-model pairing, each run in its own KVM microVM; a frozen registry of training cutoffs and a runtime check of the served model guarded against contamination, and a Docker-isolated test oracle graded 792 of 800 runs offline.
  • Neither comparison shows an average advantage: the native harness scores −1.25 pp on Opus 4.8 (48.8% vs 50.0%, 95% CI [−10.0, +7.5]) and +1.25 pp on GPT-5.5 (55.6% vs 54.4%, CI [−4.4, +6.9]), though the task pool deliberately oversamples tasks on which the two vendor pairings disagreed during screening.
  • Splitting the Opus results by workload after the fact shows the effect flipping sign: the native harness trails by 9.0 pp on repository tasks but leads by 23.7 pp on contest tasks (interaction permutation p = 0.003), which the authors treat as a pattern for a planned replication to test, not a confirmed effect.
  • On Opus the neutral harness hit the 1,200-second time limit far more often (32 of 160 runs versus 1 for the native harness), yet 22 of the 81 graded runs cut off at the limit had still produced a passing patch, so getting the answer right and finishing on its own are separate outcomes.
  • Recomputed from raw per-turn token usage at list prices, the neutral harness cost 1.3–1.6× as much per solved task on Opus and 1.2× on GPT-5.5; this revision exists because the original telemetry double-counted cache tokens for deepagents, and 58 Anthropic runs with no usage record mean the actual billed Opus cost ratio could fall anywhere from 0.7 to 2.3.

Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

Highlight Large Language Models Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko Diffusion language models decode tokens in parallel, but their bidirectional denoiser rules out the key-value (KV) cache behind fast autoregressive inference, and existing block-diffusion caches are tied to attention, so they grow linearly with context and are only approximate when retrofitted without training. The authors pretrain three 3B block-diffusion denoisers (attention, Mamba, and a hybrid) on 300B tokens with a block-causal objective that makes the cache exact, and decode all three through one cached interface. Only the state-space cache is constant in sequence length, and at 256k tokens the Mamba cache gives 4.3x lower latency, 11x less memory, and 2.6x higher single-stream throughput than attention, reaching 14x aggregate throughput with batching, while the Mamba and hybrid models keep retrieving at 8 to 16x their training length where attention collapses at 2x, at no measured quality cost.

Block diffusion lets diffusion language models reuse a cache of finalized blocks, but the caches built so far are attention-based, grow linearly with context, and are often training-free approximations. The authors pretrain three 3B block-diffusion denoisers (attention, bidirectional Mamba-2, and a hybrid) on 300B tokens under one shared block-causal objective, so cached decoding is exact and the state-space variant's cache stays the same size at any context length.

  • Training uses a single-frontier objective: one sampled block is partly masked, earlier blocks are clean and later blocks are fully masked; attention layers use a block-causal mask, and Mamba layers scan forward over the whole prefix but scan backward only within the current block, so the forward convolution and SSM state serve as an O(1) cache.
  • At 256k tokens on one H100, Mamba holds steady at 6.8 ms/step and 7.5 GB while attention needs 29 ms and 82 GB, giving 4.3× lower latency, 11× less memory and 2.6× higher throughput, and at batch 8 it reaches 1,593 tok/s while attention runs out of memory at batch 2, a 14× aggregate throughput gap (below about 16k tokens the three backbones are within a few percent).
  • With a training window of only 1,024 tokens, attention's needle-in-a-haystack retrieval drops to 12% at 2k and 0% by 8k, even with NTK-RoPE, while Mamba keeps 76% at 2k and 22% at 16k, and on LongBench at 16k Mamba scores 10.2 and Hybrid 9.1 against 3.6 for attention.
  • Quality stays close, with validation NLL within about 0.03 nats and downstream macro accuracy of 0.434 for Attn, 0.432 for Hybrid and 0.421 for Mamba; the Mamba-based models have 11–13% more parameters, but an analytic FLOPs count shows Mamba using only 0.26× attention's per-token compute at 64k.
  • All long-context results are extrapolation from a 1,024-token window, so absolute scores are modest, per-depth results show that retrieval beyond 2k mostly finds recent needles, and the paper does not test whether a single-frontier objective or longer training context changes these results.

The Cost of Compression: A Rate-Distortion Limit on Factual Hallucination

Highlight Theory Xi Wang, Shijia Xu, Rongfeng Guo Factual hallucination in closed-book question answering is usually blamed on coverage, the fact never having been seen, but finite memory can also force observed facts to be stored only approximately. The authors study a coverage-compression model in which a learner sees M of N possible queries, compresses them into B bits, and answers without retrieval, and prove a lower bound on error that splits cleanly into a compression-distortion term on observed facts, governed by the inverse rate-distortion function of a uniform K-ary source, and a coverage term on unobserved facts. The bound gives a compact way to reason about selective memory, structure, retrieval, abstention, and long-context organization, and its predicted signatures are checked with simulations and controlled fact-injection probes in modern language models that vary fact load and trainable memory.

Closed-book factual hallucination is usually blamed on missing coverage, but a model can also get a fact it has seen wrong, because finite memory forced it to store that fact lossily. The authors treat the M observed answers as a uniform K-ary source compressed into B bits and prove a rate-distortion lower bound that splits error into distortion on seen facts and guessing on unseen ones.

  • The bound is E ≥ (M/N)·δ(B/M) + (1 − M/N)(1 − 1/K), where δ is the inverse distortion-rate function under zero-one loss, so once the per-fact rate B/M drops below log₂K even seen facts have a nonzero recall error floor, and under a "free addressing" convention standard codes asymptotically reach this bound.
  • With N = 5,000, K = 10 and B = 160 bits (about 48 facts storable without loss), raising the fact load from M = 50 to M = 2,000 pushes seen-fact error from 0.011 to 0.787 while full-space error only edges down from 0.891 to 0.855, and a structure in which each base fact supports 10 derived queries cuts full-space error at M = 250 from 0.882 to 0.716.
  • In controlled synthetic fact-injection probes that use LoRA rank as a stand-in for memory, Qwen3-8B seen-fact error climbs toward the 0.9 chance level as injected facts grow, higher ranks delay the transition, and the same overload at a fixed rank appears in Qwen3-32B, Qwen3-30B-A3B and DeepSeek-R1-Distill-Qwen-32B.
  • Unseen entities stay near chance as the bound predicts, paraphrased and reworded-relation queries break down earlier than direct ones, structured facts delay overload, and a qualitative long-context version of the test with Kimi-K2-Instruct and GLM-5 shows recall error rising as more facts crowd the prompt.
  • The uniform random-answer model ignores the redundancy of real knowledge, LoRA rank is not the theoretical bit budget, and the model probes use synthetic facts and match the predicted trends rather than the exact curve, so the work explains one separable failure mode rather than hallucination as a whole.

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

Highlight Large Language Models Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer et al. Distillation makes small models stronger, but it normally requires teacher and student to share a tokenizer, leaving open whether byte-level students scale like token-level ones. The authors introduce two ways to convert token logits into byte logits, an approximate Marginalize-It and an exact End-Of-Token method, then run a large overtraining study of roughly 1-billion-parameter decoder-only transformers that crosses tokenization scheme with distillation versus cross-entropy training, up to 1 trillion bytes of data. Across eight benchmarks spanning multiple-choice QA, generation, and translation, token models win at low compute but plateau, whereas byte models start worse yet overtake them with more compute and match distilled token models using one-sixth of the training data. The 256-symbol byte vocabulary also removes the need for top-k truncation when dumping logits and cuts logit storage roughly fivefold, and fitted scaling laws predict distilled End-Of-Token-1B asymptotically beats Llama 3.2-1B, Gemma-3-1B-pt, and Gemma 2B by up to 6.5, 8.1, and 2.1 percent respectively.

Distilling a small model from a larger one normally requires both to share a tokenizer, and it has been unclear whether byte-level students scale differently from token-level students when heavily overtrained. The authors turn a Llama 3-8B teacher's token logits into byte logits in a single forward pass, then run a large overtraining sweep that compares token and byte students trained with distillation and with plain cross-entropy.

  • Marginalize-It sums teacher token probabilities over shared byte prefixes and renormalizes over the tokens that match the text so far, which is approximate because it drops the continuations of shorter tokens; End-Of-Token makes the conversion exact by adding an <eot> symbol that absorbs that leftover probability, at a cost of about 31% more compute per unit of text.
  • The sweep covers six variants (tokens, bytes, and bytes with <eot>, each supervised or distilled) of dense transformers with 1.28B layer parameters, trained on up to 1 trillion bytes and evaluated on eight benchmarks including ARC, HellaSwag, PIQA, MBPP, Natural Questions and Flores.
  • Token students lead at low compute but level off, while byte students start worse and keep improving; extrapolated scaling laws put distilled End-Of-Token-1B at 52.4% asymptotic average accuracy, ahead of distilled Token-1B (48.4%), Marginalize-It (50.5%), and the measured scores of Llama 3.2-1B, Gemma-3-1B-pt and Gemma 2B by 6.5, 8.1 and 2.1 points.
  • The distilled End-Of-Token student is projected to reach the token student's ceiling at about 6.3×10^22 FLOPs, using roughly one-sixth of the training text and one-fifth of the logit storage, because its 261-symbol vocabulary needs no top-k truncation when saving teacher logits.
  • The headline gains are extrapolated ceilings, not observed results: at budgets that were actually trained, such as 1.8×10^21 FLOPs, byte models score worse on downstream tasks despite lower bits-per-byte, only one model size is studied, and byte models cost far more at inference, which the comparison does not account for.

Reproducing and Evaluating the Generalizability of Subliminal Learning in Open-Weight Models

Highlight Safety & Alignment Daan van der Weijden, Nathan Brack, Selene Baez Santamaria Subliminal learning is a distillation effect in which a teacher model passes behavioral traits, such as animal preferences or misalignment, to a student through data semantically unrelated to those traits. Because the original work's GPT-4.x fine-tuning is no longer available, the authors reproduce its experiments on open-weight models and extend them with new preference categories (actors and politicians), a chess move generation task, an additional model (Ministral8B), and an ablation on digit length in the number-sequence task. The reproduction supports the original claims, but the extensions show the effect is not universal: transmission strength varies across traits and tasks, and one model shows almost no effect at all.

Subliminal learning is when a student model fine-tuned on unrelated data from a teacher, such as number sequences, picks up the teacher's hidden preference. Because the GPT-4.x fine-tuning used by Cloud et al. (2026) is no longer available, the authors reproduce the effect on open-weight models and then test how far it generalizes: they add actors and politicians as new preferences, chess move generation as a new task, Ministral8B as a third model, and an ablation that shrinks the numbers from three digits to one.

  • Each teacher is told by system prompt to love a target entity and generates 30,000 filtered completions, downsampled to 10,000; LoRA students of Qwen2.5-7B, Gemma3-4B and Ministral8B are then compared against control students trained on data from an unprompted teacher, using 5,000 neutral-prompt samples per condition across 5 seeds.
  • The animal-numbers result reproduces on Qwen, and the effect depends on the kind of preference: politicians transmit significantly more strongly than animals (Qwen's mentions of Trump rise from ~0.06% to ~78%, a log-odds ratio of about 8.6), while some animals, such as elephant, dog and tiger, actually become less likely to be named.
  • The effect varies sharply by model and task: Qwen shows more than Gemma, and Ministral shows almost no effect anywhere; numbers transmit significantly more strongly than chess, and trait fine-tuning even drops Qwen's valid animal answers from ~99% to ~50%.
  • Unexpectedly, shrinking the numbers from three digits to one strengthens transmission for nearly every target (panda is the exception), and a logistic-regression probe tells trait data from control data only at AUC 0.53, which suggests no single number carries the preference.
  • Limitations: open-weight models can't be compared one-to-one with the original GPT-4.1 results, chess moves are checked only for format and not legality (likely adding noise that may explain the weaker chess effect), each teacher dataset is fixed across seeds, the digit ablation covers only Qwen, and there is no correction for multiple comparisons.

Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

Highlight Agents Mykhailo Kozyrev, Andrei Kozyrev, Anton Podkopaev Coding agents increasingly draw repository knowledge from SKILL files, plain Markdown documents versioned with the code, and prior work optimizes these documents against synthetic benchmarks that capable agents solve even without them. The authors mine harder tasks by reverting a repository's merged pull requests at a single frozen base commit, then score a candidate document by whether the same agent does better with it than without it. On three Kotlin repositories, documents found by GEPA raised the score by 4.9 percentage points on average while SkillOpt gained only 0.1 points, but with the few tasks one repository's history supplies, the GEPA gain cannot be separated from the agent's run-to-run variance. A maintainer of one repository nevertheless found that the generated documents held knowledge usually gained only by working in the project.

Coding agents can load repository knowledge from SKILL.md files, but optimizing those files automatically requires a benchmark that a bare repository doesn't have. Earlier synthetic tasks from SWE-smith/gskill are also so small that Claude Code already solves nearly all of them without any skill. The authors instead mine real merged pull requests, revert each one at a single frozen base commit, and score a candidate document by whether the same agent does better with it than with an empty seed on the same task.

  • Pull requests are reverted in three tiers: first git apply --reverse, then re-creating or removing added and deleted files, and finally an LLM that rebuilds the old source; the tests that go from passing to failing become each task's FAIL_TO_PASS spec, and about one pull request in five survives (119 of 660 for koog, 131 of 452 for ktor, 100 for kotest).
  • The mined tasks are much harder than SWE-smith's, which have a median of 4–7 changed lines and almost always touch a single file, whereas koog tasks have a median of 54 lines across 3 files, and Claude Code with Sonnet 4.6 resolves only 53% of them pooled without a skill.
  • Against a no-change baseline of 0.5, GEPA documents raise the paired score by 4.9 pp on average (resolved tasks go 13→16 of 20, 16→20 of 26 and 14→15 of 23), while SkillOpt gains just 0.1 pp because a loss on ktor cancels its gains elsewhere.
  • No run is statistically significant (best p = 0.29), since with 20–26 held-out tasks per repository a sign test only detects a document that wins four of every five disagreements, and the full set of runs cost $2,014 at about $0.84 per agent rollout.
  • A koog maintainer found real project-only knowledge in both documents mixed with generic advice, would accept the GEPA one as a draft, and on two open issues either skill cut solve time from 15–20 minutes to 6–8 minutes, though this evidence rests on one reviewer, two issues and three Kotlin repositories.

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

Highlight Reasoning Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin et al. Low scores on leading physics benchmarks suggest frontier language models still struggle with advanced physics, which conflicts with domain experts' experience using them. The authors evaluated frontier models on six widely used physics benchmarks and had faculty and graduate researchers audit problem statements, reference solutions, and model responses. Most answers initially graded incorrect turned out to reflect grader errors, wrong reference solutions, or ambiguous questions rather than faulty physics reasoning. After experts fixed or excluded flawed items, GPT-5.6-Sol's mean@4 score rose from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, and its pass@4 reached 94.4% on 54 retained CritPt challenges, suggesting these closed-ended benchmarks are near saturation.

Low scores on leading physics benchmarks suggest frontier LLMs still struggle with advanced physics, but an expert audit finds that most recorded failures come from wrong reference answers, ill-posed questions, or graders rejecting correct answers. Once physicists repair or exclude the flawed items, frontier models come close to saturating six widely used closed-ended benchmarks.

  • Faculty and graduate researchers, matched by subfield, review the problem statements, reference solutions, and responses from GPT-5.6-Sol, Claude Fable 5, and Gemini 3.1 Pro, labeling each failure as a model, grader, or benchmark error before excluding or repairing flawed questions and re-grading with an HLE-adapted LLM judge.
  • Of 250 rejected answers pooled across HLE-Physics, PHYBench, PRISM-Physics, and UGPhysics, only 12 (4.8%) were genuine model errors, with 57.2% traced to benchmark defects and 38.0% to graders, such as PHYBench's expression-edit-distance scorer giving zero credit to algebraically equivalent answers.
  • After correction, GPT-5.6-Sol's mean@4 rises from 47.3% to 78.7% on HLE-Physics, 61.0% to 87.2% on CMT-Benchmark, and 13.0% to 94.6% on PRISM-Physics, while on CritPt it climbs from 32.3% to 87.5% mean@4 and 94.4% pass@4 on the 54 retained challenges.
  • The before-and-after numbers are not strictly comparable, because corrected scores use smaller retained subsets (86 of 202 HLE-Physics questions were excluded), CritPt's pre-audit figure uses a different metric, and its references were re-derived by the auditors since the official solutions are unavailable.
  • Near-saturation on problem-set-style questions does not mean research ability, since the authors' GPT-based agentic harnesses have not fully solved a single open theoretical-physics problem, which motivates their call for harder, expert-verified benchmarks.

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

Highlight HF pick · 30▲Large Language Models Zhiwei Li, Lei Zhu, Hao Gu, Xiang Hu, Yan Wang, Haitao Mi et al. Trainable sparse attention methods added to pretrained Transformers use a lightweight selector to score context tokens or blocks and then keep only the top K. That hard selection blocks gradients, so the selector is trained to imitate dense attention weights, which may not rank context by how much it actually helps predictions under a fixed budget. Simple Attention Sparsification (SAS) trains the selector directly on the language modeling loss by adding its continuous scores to the attention logits during training. The design depends on placing the gate inside the softmax in log form, normalizing gates against the always-kept current block, and keeping scores continuous, and a memory-efficient Triton kernel supports long-sequence training. Across reasoning, long-context, and agentic tasks, SAS beats trainable sparse attention baselines at every attention budget, with especially large gains under tight budgets.

Trainable sparse-attention selectors choose context blocks with a hard Top-K step, which blocks gradients from the language-modeling loss. Methods like SeerAttention-R therefore distill each layer's dense attention instead, which ranks blocks by where the dense model looked rather than by what helps predictions under a fixed budget. The fix proposed here is to add the selector's scores as log-space gates inside the attention softmax during training, so the standard language-modeling loss trains the selector end-to-end with no teacher attention.

  • The backbone stays frozen and SeerAttention-R's AttnGate selector is kept; its block scores are softmax-normalized, their logs are added to the attention logits of the selected blocks, and the current block keeps a unit gate. Training runs on a FlashAttention-style Triton kernel, and at inference the scores simply become Top-K block indices.
  • Ablations on GPQA-Diamond (Qwen3-4B, 2048-token budget) show each design choice matters: gating inside the softmax reaches 54.4 versus 41.6 outside it, sigmoid gates (17.0) and raw-logit injection (18.8) collapse, and hard straight-through Top-K gates reach only 46.0.
  • At a 1024-token budget the method beats SeerAttention-R by 6.0–7.7 points on MATH500 and 10.6–15.5 points on GPQA-Diamond across Qwen3 4B, 8B and 14B, adds +13.0 on AIME24 at 2048 tokens (4B), and at 4096 tokens matches or exceeds full attention (71.72 vs 71.25 on AIME24).
  • Although the selector was trained only on math data (OpenR1-Math-220k), it still leads on LongBench (+2.4 on 8K+ inputs for Qwen3-14B at a 2048-token budget) and BFCL (up to +3.5), and decoding in SGLang runs 5.6× faster than dense attention at 512K context (about 13× at batch size 8 with 64K context).
  • Only decoding is sparsified while prefill stays dense, and Top-K ranking grows to 90% of each decode step's time at 512K context. VitaBench results against SeerAttention-R are mixed at the tighter budget, and the OLMo3-7B continued-pretraining result is preliminary: it averages 43.28 versus 43.88 for the dense base, with a clear drop on CRUX.

Type Diversity Enables Transformers to Generalise Compositionally

Highlight Large Language Models Anssi Moisio, Mathias Creutz, Mikko Kurimo Earlier work found that Transformers generalize compositionally less well to new structures than to new words. The authors argue this comes from the datasets, which contain many distinct lexical types but few structural ones, rather than from the architecture itself. Using Grammatical Framework, they build linguistically diverse variants of the COGS and SLOG datasets that vary type diversity, defined as the number of different constructors of a type. They find type diversity correlates with compositional generalization equally for lexical and structural cases, which contradicts the earlier claim that compound divergence explains task difficulty. They also examine how the diversity of other types and the surface format of the logical semantics affect generalization.

Transformers are widely reported to handle lexical compositional generalisation (a known word in a new position) far better than structural generalisation (a known phrase structure in a new position). The authors argue this gap comes from the benchmarks, not the architecture: datasets like COGS offer hundreds of lexical constructors per type but only a handful of structural ones. They define type diversity as the number of distinct constructor functions that build trees of a grammatical type, and show that raising it for noun phrases unlocks structural generalisation.

  • The method uses Grammatical Framework to regenerate COGS and SLOG with compositionally derived logical semantics, then adds new noun-phrase constructors one at a time (participial phrases, adjectives, relative clauses, plurals) and tests on the held-out case where prepositional phrases appear in subjects after being seen only in objects during training.
  • Accuracy on that task climbs from 1.7% to 69.3% as natural noun-phrase constructors are added, and reaches 92.4% with two artificial noun-phrase structures, while adding verb tenses, which add no noun-phrase diversity, only moves it to 73.3%, within seed variance.
  • A size-matched control that adds 60% more training samples by swapping words leaves accuracy at 1.7%, and placing participial phrases only in subjects gives 0.2% versus 32.0% when they appear in both roles, which undercuts the earlier claim that surface details like variable numbering and token order are the main obstacle.
  • Controlled experiments on over a hundred SCAN variants support the finding that type diversity predicts lexical and structural generalisation equally, which contradicts the claim from distribution-based compositionality assessment (DBCA) that higher compound divergence is what makes these tasks hard.
  • Limitations: variance across random seeds is large (standard deviations up to about 18 points), only five natural noun-phrase constructors are covered, type-diversity counts depend on the chosen grammar, and the other two COGS structural splits stay out of reach because their test sets introduce unseen variable names.

Applications 94

On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health

Ibukunoluwa Soyebo, Alyssa Donawa, Rodrigo Aguilar Barrios, Brice Patchou, Corey E. Baker Mobile stress monitoring is a common target for mental health interventions, but sending sensor and self-report data to cloud models raises privacy concerns. The study evaluates on-device language models (ODLMs) for multimodal stress prediction using zero-shot prompting, measuring predictive accuracy alongside latency and throughput under mobile resource constraints. Objective sensor features marginally outperform subjective self-reports on average, and lightweight sub-2B-parameter models achieve low latency with predictable resource usage, though the results also expose practical limits of the approach for mobile mental health.

On Identifying Adversarial Intent Injection in AI-Native 6G Networks

Nilesh Chakraborty, Petar Djukic, Burak Kantarci cross-listed Intent-Based Networking (IBN) in AI-native 6G lets high-level goals be translated into network configurations, but it also opens the door to adversarial intent injection, where malicious policies hide within benign intent flows, potentially in stealthy patterns that are hard to detect. The authors define a fine-grained threat model, study four injection strategies (stealth-mode, random, increasing frequency, and decreasing frequency), and propose a dual-path detector combining a CNN on TF-IDF features for supervised detection with an AutoEncoder trained only on benign intents for one-class detection. Accuracy reaches 0.97 and F1 reaches 0.98, gains of roughly 9% and 36% over the state-of-the-art baseline.

A decision-basis contract for auditable LLM-assisted medical billing verification: deterministic rules, verbatim evidence, and fail-closed abstention

Jan H\"olter, Kevin Geis, Benjamin Raab, Boris Bauke cross-listed Medical billing verification assisted by large language models needs to be auditable, so the authors propose a decision-basis contract that separates deterministic checks of versioned fee-catalog rules (catalog release, code availability, quantity limits, exclusions) from LLM classification of free-text documentation as supported, contradicted, or missing required information. Support or contradiction requires a verbatim evidence span, and unavailable rule context, failed assessment, or absent evidence triggers fail-closed abstention. Evaluating four locally run open-weight models on a synthetic catalog and 36 curated cases, outcome agreement varied across models and showed no consistent advantage over an end-to-end baseline, though explicit documentation requirements improved detection of missing information for all models and the evidence gate exposed correct raw judgments that lacked valid evidence; the authors note evaluation on real catalogs with human reviewers is still needed.

BRIDGE-EEG: Bridging Self-Supervised Pretraining and Efficient Deployment for Cross-Dataset EEG Classification

Meghna Roy Chowdhury, Chengwei Zhou, Haotian Yu, Gourav Datta, Shreyas Sen cross-listed Electroencephalography (EEG) foundation models learn transferable representations but are too large for wearable or edge hardware. BRIDGE-EEG maps heterogeneous recordings with differing channel counts, montages, and sampling rates onto a device-agnostic 62-channel time-frequency representation, pretrains an 11.84M-parameter SE-ResNet18 teacher with SimCLR on unlabeled EEG from five datasets, and distills it into SE-ResNet8 (1.56M) and SE-ResNet4 (0.48M) students using task-agnostic and task-specific distillation. Across six benchmarks covering abnormality detection, motor imagery, and emotion recognition, the students match or beat several recent EEG foundation models with 10 to 1,000 times more parameters on abnormality detection and emotion recognition, though a representation gap remains on motor imagery, and profiling on an NVIDIA Jetson Orin Nano shows up to 3.0 times lower energy per inference.

HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications

Inesh Chakrabarti, Zejun Xiong, Pragya Sharma, Mani Srivastava cross-listed Foundation models, including large language models, vision-language models, and time-series models, are increasingly deployed on embedded and edge hardware for cyber-physical systems (CPS) and Internet of Things (IoT) applications, where energy, latency, and memory matter as much as accuracy, yet capability benchmarks ignore compute constraints and hardware profilers are platform-specific and mutually incompatible. HoliBench is a modular open-source toolkit that jointly measures accuracy, latency, and energy across devices from single-board computers to GPU servers through a platform abstraction layer that calibrates cross-device measurement, supporting multiple modalities, inference engines, concurrencies, and existing evaluation harnesses, with an interactive interface for constraint-aware configuration selection over a design space profiled once and reused. A characterization of 20 models across 7 device types, 3 quantization levels, 8 inference backends, and over 30 tasks finds that quantization cuts latency only on hardware with low-precision support, accuracy gains show diminishing returns against energy, and average inference power is roughly constant across output lengths for autoregressive workloads. Standalone single-model profiles predict combined multi-model pipeline latency and power within 1.2% and 2.5% under sequential co-resident execution, so deployments can be explored without profiling every pipeline configuration.

Optimizing Geoengineering Interventions Using Differentiable Climate Models

Pulkit Dubey, Dorian S. Abbot, Ashesh Chattopadhyay cross-listed If a solar geoengineering program is deployed, planners will need tools to design interventions that meet a target while limiting disruption. Using the differentiable primitive-equation atmospheric model JAX-GCM, the authors impose a uniform 4 K ocean warming and optimize the amplitudes of sea-surface cooling in five zonal ocean bands so that land near-surface air temperature returns to the model's unwarmed climatology, optimizing greedily over 8 to 14 day segments in a receding-horizon fashion because gradients through chaotic dynamics decorrelate beyond the Lyapunov horizon. The learned cooling pattern removes 92.3 ± 0.4% of realized land warming across a ten-member ensemble of two-year rollouts, also restores land precipitation, evaporation, and humidity distributions that were not in the objective, and works without re-optimization when replayed in the AI emulators LUCIE and NeuralGCM.

4D Parallelism Unlocks Exascale Bayesian Neural Networks for High-Fidelity Atmospheric Modeling

Deifilia Kieckhefen, Juan Pedro Guti\'errez Hermosillo Muriedas, Lars Helge Heyen, Mathis Bode, Iida Hakulinen, Andreas Herten et al. cross-listed Bayesian neural networks can quantify both aleatoric and epistemic uncertainty in global atmospheric forecasting, but at 0.25° resolution they are computationally prohibitive. BEAST is a Bayesian Swin Transformer trained with an orthogonal 4D parallelization scheme that adds a domain-tensor-parallelism strategy and a new uncertainty-parallel method. A 2.4-billion-parameter configuration reached a peak of 3.96 exaFLOPs per second on 20,480 NVIDIA GH200 GPUs on the JUPITER supercomputer, and a 700-million-parameter model with 96 weight samples was trained on 40 years of data. The model achieves skill competitive with leading probabilistic AI and numerical weather models, predicts extreme events well, and generates large ensembles 3 to 4 times faster than the current best AI model.

Scaling Clinical Judgment to Evaluate Medical AI

Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah et al. Blinded physician grading is the standard way to assess clinical reasoning in large language models (LLMs), but it is hard to scale, so studies rely on small panels from a single institution or specialty. The authors train PrecepTron, a 32-billion-parameter model fine-tuned with low-rank adaptation (LoRA) on a small number of physician examples, and release GRAND-ROUNDS, a benchmark of 9,217 scores from 11 physicians across seven studies. Frontier LLMs used as off-the-shelf judges often disagree with physicians and with each other, whereas PrecepTron achieves physician-level consistent scoring across tasks and reproduces headline findings from five influential studies in JAMA, Science, and Nature Medicine without new human grading. It is then used to ask questions that human grading could not cover, such as how the diagnostic accuracy of frontier LLMs changes when clinical cases are revealed piecemeal, down to token by token.

A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK

Matteo Lupinacci, Luigi Arena, Francesco Blefari, Angelo Furfaro cross-listed Mapping observed system behavior to the MITRE ATT&CK framework remains largely manual, and automated methods typically depend on retrospective Cyber Threat Intelligence reports rather than direct evidence of adversary activity. Trace2ATT&CK collects kernel-level events with eBPF, correlates attacker commands into a provenance graph, and condenses it into compact representations that an LLM maps to ranked ATT&CK technique candidates with rationales, using either plain prompting or retrieval-augmented generation (RAG) grounded in the ATT&CK knowledge base. Evaluated on 347 Linux Atomic Red Team tests with locally deployed open-weight LLMs, RAG consistently beats plain prompting and provenance graphs substantially outperform raw telemetry, indicating that automated mapping can run on local inference without exposing confidential data.

MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations

Youssef Mohamed, Ahmed Heakl, Qinrong Cui, Junhong Liang, Rafiq Ali, Bdour Babillie et al. Most medical benchmarks use single-turn multiple-choice cases, which poorly reflect real consultations where clinicians gather evidence interactively from patients who communicate in very different ways. MedRoundsQA converts 1,387 board-exam cases across 17 specialties into structured 24-slot clinical records and instantiates each as a doctor–patient dialogue between two agents under varying patient personas, holding the clinical content fixed, with cases grouped by difficulty using model-based uncertainty. Across fifteen LLM doctor agents, moving from single-turn diagnosis to multi-turn consultation degrades accuracy by roughly 13–39 points, and while more turns keep improving question relevance, diagnostic accuracy typically plateaus after 6–12 turns. Patient persona alone shifts diagnostic accuracy by about 7–8 points between the lowest and highest education levels, an equity risk that single-turn benchmarks miss.
84 more specialized papers

Large Language Models 51

The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices

\'Edouard Gu\'egain, Tristan Coignion cross-listed Running large language model (LLM) inference on smartphones is often assumed to be more private and more sustainable than cloud inference, but it pushes heavy computation onto battery-powered devices whose lifetimes it may shorten. The authors measure energy per generated token, inter-token latency, accuracy, and battery-cycle consumption for 18 models spanning families, sizes, and quantization levels on two modern smartphones and a server, each using its state-of-the-art deployment stack. On-device inference is on average 3 times less energy-efficient than batched server inference, the relation between quantization bit-width and energy per token is non-monotonic with device-specific sweet spots, and eight of the 18 configurations sit on the accuracy-energy Pareto front, enabling battery-aware model routers. Under realistic life-cycle assumptions local inference cannot beat batched server inference on per-token environmental impact, with 88-90% of that impact coming from device embodied carbon rather than electricity use.

Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs

Abhinav Anand, Sanjana Reddy Pachika, Shweta Verma, Mira Mezini Reinforcement learning (RL) post-training makes code-generating large language models (LLMs) follow instructions and produce functionally correct code, but it normally demands expensive online sample generation and heavy GPU-CPU traffic for verifying each sequence. The authors ask whether this stage can run entirely offline from existing datasets, with no new samples drawn from the model during training. With only a few hours of offline RL training, zero-shot code generation improves substantially without any online sampling, and the gains hold across models from 0.5B to 7B parameters, although their magnitude varies between model families.

Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale

Vaibhav Singh, Pierre-Andr\'e No\"el, Torsten Scholak, Eugene Belilovsky, Oleksiy Ostapenko Diffusion language models decode tokens in parallel, but their bidirectional denoiser rules out the key-value (KV) cache behind fast autoregressive inference, and existing block-diffusion caches are tied to attention, so they grow linearly with context and are only approximate when retrofitted without training. The authors pretrain three 3B block-diffusion denoisers (attention, Mamba, and a hybrid) on 300B tokens with a block-causal objective that makes the cache exact, and decode all three through one cached interface. Only the state-space cache is constant in sequence length, and at 256k tokens the Mamba cache gives 4.3x lower latency, 11x less memory, and 2.6x higher single-stream throughput than attention, reaching 14x aggregate throughput with batching, while the Mamba and hybrid models keep retrieving at 8 to 16x their training length where attention collapses at 2x, at no measured quality cost.

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal Large language models are widely used as automated judges, but individual judges carry systematic biases, and most prior work studied pairwise comparison rather than the absolute scoring common in practice. Across four benchmarks and six models forming 36 judge-examinee pairs, a model's task accuracy strongly predicts its judging accuracy and inversely predicts its directional bias, yet more capable examinee models consistently receive more lenient judgments from every judge. The proposed calibrated weighted majority voting (WMV) aggregates multiple judges weighted by online estimates of their false-positive and false-negative rates, derived purely from inter-judge disagreement patterns with no ground-truth labels, and in a simulation with shifting task distributions it tracks an oracle with perfect error-rate knowledge to within 0.5 percentage points, beating both individual judges and unweighted majority voting.

Creating an Atomic User Model for Personality-Aware Large Language Model Interaction

B. Sankar, Deepthika S, Pawni Yadav, Amogh A S cross-listed Assistants built on large language models are expected to write in their user's voice, but the dominant approach stores task-specific preferences summarized from history, which must be relearned whenever the task changes. The authors first characterize personality seepage, where a prompt's linguistic surface carries a personality fingerprint the assistant mirrors, then propose the Atomic User Model (AUM), a human-readable representation with a stable identity nucleus and four shells (psychological, cognitive and experiential, behavioural, and social), used as a retrieval index where a task classifier and budgeted retriever return a small payload of fields at generation time. With sixteen language-model-simulated participants and six style-sensitive tasks, retrieving eight fields matched the full model's style fidelity using 23% of the context (211 versus 915 tokens), beat flat preference notes by 0.24 points on a five-point scale, and raised forced-choice identification of the participant's own voice from 14.9% to 42.7%, with the benefit largest for users the unpersonalized assistant serves worst.

Competence-Gated Pooling of Language Models and Priors for Event Forecasting

Aditi Tiwari, Aashrith Bandaru, Heng Ji When a language model is one of several forecasting signals alongside a market, crowd, or statistical forecast, the relevant question is its marginal value over the external forecast, not its standalone accuracy. Under Brier loss the authors characterize when model disagreement can improve an external forecast, then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the gate improves the main external baseline from 0.0771 to 0.0732 Brier and beats global combinations, though it gives no significant improvement on the official ForecastBench market subset, where it mostly defers to the market. Across four Qwen models, verbal confidence fails to identify when the model beats the external forecast, whereas outcome-estimated competence supports better abstention decisions.

Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature

Zihan Zhu, Zhehang Du, Xuyang Chen, Tim Tsz-Kit Lau, Jiayuan Wu, X. Y. Han et al. Low-Rank Adaptation (LoRA) uses its nominal rank to set an adapter's capacity, but the optimizer determines how much of that capacity the induced weight-space updates actually use. A case study on GPT-2 finds that AdamW produces per-step updates with concentrated singular spectra and low effective rank, while Muon spreads updates across more directions and benefits more consistently from higher rank; this motivates ISO-LoRA, which couples the LoRA factor updates through spectral descent on the induced tangent perturbation so energy is distributed evenly across singular directions. A one-step analysis under a stylized spiked-gradient model shows higher effective rank than factor-wise optimizers, and experiments on 0.1B to 7B parameter models show improved effective rank and downstream performance, with the largest gains at moderate-to-large LoRA ranks.

Retrieval-Augmented Generation for Scientific Code Understanding

Aaron Nobile, Andreas Adelmann, Mohsen Sadr cross-listed Frontier coding assistants depend on very large cloud-hosted models, raising cost and data-privacy concerns for in-house scientific codebases. The authors build a Retrieval-Augmented Generation (RAG) system that front-loads the expensive work into an offline ingestion stage of parsing, structural graph construction, LLM-written entity explanations, and embedding, leaving a lightweight online answering stage that small open-source models can handle locally. On a 100-question benchmark spanning eleven categories over the IPPL C++ scientific codebase, scored by an independent frontier model as judge, a 9B-parameter model achieved the highest average score (0.795), beating larger models in the same pipeline and the same models running inside the Claude Code retrieval architecture, which the authors read as evidence that model family and retrieval quality matter more than parameter count.

T-GADE: Thermodynamical Generative-AI-Driven Evolution of LLM Artifacts

Kyoko Ogawa, Naoki Mori Combining evolutionary computation with large language models (LLMs) requires controlling population diversity, not just generating candidates. T-GADE evolves structured LLM artifacts, such as a description paired with code, by extending thermodynamical genetic algorithms with LLM-based genetic operators and artifact-level diversity evaluation under a common free-energy objective, where Fermi-type occupancy excludes repeated genotypes and Bose-type occupancy permits them; the authors also establish exact one-member removal and conditions under which the method recovers the zero-temperature survival rule of Evolution of Heuristics (EoH). On the online bin-packing task from the EoH paper, generational Bose-type T-GADE at temperature 0.003 reduced median training excess by roughly 29 percent, from 1.152 to 0.815 percent, over 20 runs per configuration, and validation-based selection among its two top-ranked final candidates matched the EoH median transfer excess of 0.496 percent.

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer et al. Distillation makes small models stronger, but it normally requires teacher and student to share a tokenizer, leaving open whether byte-level students scale like token-level ones. The authors introduce two ways to convert token logits into byte logits, an approximate Marginalize-It and an exact End-Of-Token method, then run a large overtraining study of roughly 1-billion-parameter decoder-only transformers that crosses tokenization scheme with distillation versus cross-entropy training, up to 1 trillion bytes of data. Across eight benchmarks spanning multiple-choice QA, generation, and translation, token models win at low compute but plateau, whereas byte models start worse yet overtake them with more compute and match distilled token models using one-sixth of the training data. The 256-symbol byte vocabulary also removes the need for top-k truncation when dumping logits and cuts logit storage roughly fivefold, and fitted scaling laws predict distilled End-Of-Token-1B asymptotically beats Llama 3.2-1B, Gemma-3-1B-pt, and Gemma 2B by up to 6.5, 8.1, and 2.1 percent respectively.

ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression

Liu O. Martin, Lucas Bandarkar, Nanyun Peng The team submitted six systems to the unconstrained WMT26 Model Compression Shared Task for English to Simplified Chinese and English to Egyptian Arabic, all derived from GPT-OSS-20B. They rank the model's experts by task-specific routing mass, use cross-lingual routing divergence to decide how much capacity each layer keeps, physically remove low-importance experts, recovery-tune the result on GPT-5.1-generated synthetic translation data, and apply MXFP4 quantization to the retained expert projection weights. A robust inference pipeline handles the instruction-conditioned setting with category inference, output validation, retries, segmented fallback, and source-owned JSON reconstruction. The submissions shrink the model to between 4.186B and 7.770B parameters and 4.55 to 6.33 GiB on disk, with internal xCOMET-XL scores against GPT-5.1 pseudo-references used to compare the operating points.

SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data

Praveen Kumar Myakala, Ravichandra Namburi, Sowmya Keragodu Jayaramu, Sooraj George Thomas Training on synthetic text can cause model collapse, but prior diagnostics only detect it after training, whereas the practical need is to screen a corpus of unknown provenance beforehand. SynthSentry is a corpus-level, model-agnostic contamination score that needs no access to the generating model or synthetic labels, computed as a distributional divergence over lexical diversity collapse, n-gram tail truncation, and perplexity variance across reference models. Under a leave-one-generator-out protocol the score ranks contaminated corpora by severity with little loss even when whole generator families are held out, and per-domain calibration on naturally repetitive human text such as legal, clinical, and source-code corpora stays near its nominal false-positive budget once covariance shrinkage and a bootstrap threshold replace a naive quantile that runs four times over budget. A downstream fine-tuning check found no contamination-driven accuracy deficit at the authors' small scale while exposing an over-pruning risk once pruning exceeds the true contamination fraction; results are limited to English batch-mode screening of single-generation rather than recursively generated contamination.

ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge

Shreyas Krishnan, Serina Chang, Abhishek Nagaraj Measuring how much occupation-specific knowledge large language models hold has relied either on mapping abstract skills to occupations through task definitions or on costly expert-written questions. ORQA instead links O*NET occupations to trusted occupation-specific websites such as regulatory agencies, licensing bodies, and professional organizations, and converts them into source-traceable question-answer pairs through an automated pipeline plus human review, yielding 480 questions over 116 occupations from all 21 major groups of the Standard Occupational Classification. Across 15 frontier and open-weight models, Claude Opus 4.6, GPT-5.4, and Claude Sonnet 4.6 lead at roughly 58-62% while smaller open-weight models score 33-41%, with healthcare occupations reaching 78% but office and administrative support near 40% and some individual occupations at essentially zero. Open-ended formats and wage-bill weighting do not materially change model rankings.

Representation-based Masked Diffusion Model

Yangrong Hu, Ding Huang, Xueyu Zhou, Jian Huang Masked Diffusion Models (MDMs) enable parallel text generation, but standard samplers update multiple masked tokens independently, ignoring their mutual dependencies and risking incoherent output. The Representation-based Masked Diffusion Model, RMDM, first encodes text into a continuous semantic space with a pretrained encoder and learns an invertible transformation that normalizes that representation to a Gaussian prior for easy sampling, then trains a masked diffusion model conditioned on this latent so the global semantics coordinate parallel token updates. Empirically the approach improves generation quality most in aggressive few-step sampling regimes, where independent updates otherwise degrade coherence.

Beyond the Query: Do Retrieval Signals Improve Adaptive Multimodal RAG Routing?

Qiaomu Li, Qiuyuan Zhang, Nong Ming Adaptive retrieval-augmented generation (RAG) systems often use retrieval-time signals to decide whether to run another retrieval, reranking, or multimodal step, and this study asks whether those signals add routing value once the query itself is already known. Across document, audio, and video RAG, the authors compare matched query-only and query-plus-retrieval routers while holding the optional actions, router family, training procedure, and evaluation fixed. On the held-out final evaluation, adding the tested retrieval signals does not reliably improve RUN/SKIP routing decisions over the query-only baseline, even though some signals do correlate with whether a later step will help. The lesson is methodological: retrieval-state features should only be credited with routing value when they beat a matched query-only control, and their incremental value must be demonstrated rather than assumed.

Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation

Liang Zhao, Yong Wang, Jiangzhe Chen cross-listed LLM-as-a-judge protocols are commonly debiased by instructing the judge to ignore presentation cues such as citation formatting, source labels, and evidence-display style, and this work shows that such instructions can suppress bias while degrading the judge's ability to resolve genuine quality differences. TraceJudgeBench is a diagnostic benchmark for auditing citation-like artifacts in RAG and agent-workflow evaluation, covering content-equivalent pairs, citation ablations, correctness conflicts, human-validated soft and moderate quality gaps, prompt-strength ladders, decoupled judging, and a controlled workflow-ranking probe. Across GPT-5.5, Claude Sonnet 4.6, and DeepSeek V4-Flash, stronger anti-citation prompts cut worse-cited wins from up to 50.5% to 0%, but some operating points already convert validated moderate-gap decisions into ties while correctness-conflict accuracy stays at or above 93%; a 50-pair FinQA construction and open-weight Qwen2.5-14B-Instruct-AWQ and Gemma-3-12B-IT runs reproduce the frontier, and TRACE-style decoupled judging recovers 96.5 to 100% of better-plain resolution. Human validation separates three meanings of a tie, and the authors argue that bias suppression, resolution retention, tie cost, and protocol cost must be reported together.

Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration

Nilesh Jaiswal, Aniket Agrawal, Arjit Shukla, Divya Malhotra, Saurabh Garg, Suchit Puri et al. Enterprises migrating legacy monoliths to microservices lean on LLM code translation with Retrieval-Augmented Generation (RAG), but vector-similarity retrieval returns isolated chunks that sever inheritance chains and yield code that fails to compile. The authors build a Hierarchical Context-Resident Graph (HCRG) pipeline that extracts Abstract Syntax Trees (ASTs) with tree-sitter, stores architectural edges in a Google Cloud Spanner property graph, and serializes the structure into a Gemini context cache for parent-first translation, then evaluate with a seven-metric software-engineering framework rather than text overlap. CodeBLEU scored 91% for both approaches and masked the structural failures, whereas the graph approach cut API hallucination rates from 56.4% to 16.2% and nearly doubled dependency resolution quality, at the cost of lower cyclomatic-complexity consistency and slightly worse docstring preservation as the model over-engineers defensively with the denser global context.

AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization

Ji Liu, Saptarshi Majumder, Yiqing Huang, Wenwen Ouyang, Umang Pandey, Zeping Li et al. Most LLM-based GPU kernel generation work targets CUDA on NVIDIA hardware and depends on repeated frontier-model calls for generation, reflection, and optimization, leaving AMD's CDNA GPUs under-served. AMDKernelVault is an open corpus and training framework produced by agentic pipelines (HIPKernelGen and TritonKernelGen) that convert PyTorch references into HIP or Triton kernels, compile and validate them under ROCm, and profile latency on AMD hardware, yielding 62,153 execution-verified HIP kernels, 39,893 Triton kernels, and 2,377 ROCm library question-answer entries. A Qwen3-8B model trained on the corpus with supervised fine-tuning and execution-aware reinforcement learning achieves the highest correctness among compared models on PyTorch-to-HIP (34.0% Pass@1), TritonBench-G, and ROCmBench under fixed evaluation budgets, though it does not uniformly lead on compilation or speed metrics.

Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models

Zhongzhan Huang, Junxin Li, Guoming Ling, Yupei Lin, Shanshan Zhong, Hefeng Wu Full benchmark suites are costly to run and highly redundant, yet the strongest existing benchmark compression methods (BCMs) need per-sample results from many LLMs to pick representative items, which is itself expensive for newly released benchmarks. ZipBench evaluates only a small set of anchor models, synthesizes pseudo evaluation results to broaden coverage, learns compact sample representations, and selects a small representative subset, with theoretical error and rank-consistency guarantees. The authors release ZipBench Zoo, compressed proxies for more than 100 text, multimodal, and agent benchmarks, which track full-benchmark scores with mean absolute errors of 0.002 to 0.02 and average Spearman correlation around 0.98, lowering evaluation cost for compute-constrained researchers.

Confidence-Gated Transductive Test Generation for Code Reranking

Sungjae Lee, Youngsik Yoon, Seockbean Song, Siwei Wang, Wei Chen, Jungseul Ok Ranking candidate programs generated by an LLM depends on synthesized test cases, but producing reliable expected outputs for those tests is difficult. Confidence-Gated Transductive Test Generation (CoTT) first runs an efficient inductive procedure and invokes the more expensive transductive generation, which conditions on the candidate programs themselves, only when inductive confidence is low. On code reranking benchmarks CoTT outperforms prior baselines across the reported metrics while costing less than applying transductive generation to every input, showing that confidence-based allocation of test-time compute yields a favorable efficiency-effectiveness trade-off with a single efficient model.

Temporal Recurrence Favors Fewer Layers

Ivan Anokhin, Johan Obando-Ceron, Irina Rish, Sebastian Risi In streaming settings a recurrent model can carry latent computation across time steps, which raises the question of how much within-step depth is still needed once recurrence supplies sequential computation. Rather than only asking whether shallow recurrent models can be competitive, the authors treat this as a compute-allocation problem, sweeping within-step depth, expert width, and the number of parallel experts per layer across several compute budgets and comparing the best recurrent and non-recurrent allocations at approximately matched per-step compute. On both Sokoban and autoregressive FineWeb language modeling, the best recurrent allocations use substantially fewer layers than the best non-recurrent ones while matching or exceeding their performance.

RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems

Ziyue Yang, Yuting Jiang, Lei Qu, Peng Cheng cross-listed AI-driven optimization of LLM inference has mostly relied on profiling existing software stacks, which limits the search to what those stacks can already do. The authors argue that finding fundamentally better inference architectures requires a general workload representation, a verifiable space of mutations, and an evaluator independent of any implementation, and they provide all three through RoofLang, a domain-specific language (DSL). Analysis with RoofLang shows that DeepSeek V4-series models could reach 3.5–39.5× higher peak decode throughput than other representative models, a gap out of proportion to parameter counts that comes largely from compact KV-cache designs allowing bigger batches and less memory traffic. A persistent optimizer agent working in the language discovered several new architectures that improved both throughput and interactivity of DeepSeek V4 Pro on NVIDIA B300 by 6.23–50.1%.

Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size

MyungHoon Ryu, XinYu Piao, Jong-Kook Kim LLM memory usage grows in proportion to input length, and neither model optimization nor lossy prompt compression lets a model recall facts beyond its pretrained, size-limited context window. The proposed training-free method represents facts from a source document as residual vectors in the model's feed-forward layer activations, then deterministically reconstructs the facts relevant to a query without referring back to the original document. This keeps GPU memory usage near-constant as context length grows, preserves high fidelity, and leaves model weights unchanged. In experiments, the method answers single-fact questions over two-million-token story contexts where previous methods fail.

Implicit Personality Representations in Humans and LLMs

Yilin Geng, Omri Abend, Eduard Hovy, Lea Frermann A century of psychology has found that the trait words people use to describe others vary, but the relationships among traits, meaning which go together and which are opposites, are consistent across raters and cultures. The authors test whether Qwen 2.5-7B-Instruct reproduces this structure internally. They compare a human implicit-personality matrix over hundreds of traits, built from millions of crowd-sourced ratings of fictional characters, with a matching matrix derived from contrastive model activations. The two structures align strongly, with a Mantel correlation of r = 0.77, and the agreement holds for individual traits as well as overall. The two dominant axes of the model's trait representations recover the well-established human dimensions of social warmth and intellectual competence, and projecting activations from held-out dialogue onto these directions yields personality profiles that agree with human ratings.

What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code

Cristina Improta, Pietro Liguori, Domenico Cotroneo cross-listed Evaluations of AI coding assistants focus on functional correctness, leaving open whether their code differs from human code in the quality dimensions that dominate maintenance cost over a codebase's life. The study pairs 787,562 human-written functions from open-source Python, Java, and C repositories with implementations generated from the same docstrings by OpenAI GPT models, DeepSeek-Coder, and Qwen2.5-Coder, then compares structural complexity, statistical naturalness, defects, and vulnerabilities. AI-generated code is roughly half the size and branching of human code and stylistically templated, with defects concentrated in repetitive boilerplate rather than the issues typical of mature codebases. Security is language-dependent: LLMs produce more and more severe findings in Python and Java but fewer high-severity memory-safety findings than humans in C, and the authors release CQBench, a benchmark of 27,346 issue-prone tasks for quality assurance and security testing.

When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

Griffin Farrow, Lily Sijia Li, Jack Johnson, Tingyan Wang, Philip Torr, William Bolton et al. Rubric-based grading has become the leading way to evaluate LLMs in medicine, but it is unclear whether rubric scores capture clinically relevant hallucinations. A controlled study on MedHallu finds that more specific rubrics better separate correct from hallucinated responses. The authors then build a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that produces matched pairs of correct and error-injected responses. Across HealthBench, HealthBench Professional, and LiveMedBench, rubrics miss clinically relevant injected hallucinations, often leaving scores unchanged. Rubrics work best when they explicitly check facts and worst for unexpected errors they do not anticipate, and a preliminary retrieval-based factuality check recovers some of these rubric-blind errors.

GraphAHA: Graph-Based Adaptive Search with Heterogeneous Actions for Test-Time Code Generation

Xitao Li, Haijun Wang, Gege Yuan, Qiyuan Wu, Jiali Wei, Ming Fan et al. cross-listed Test-time scaling for code generation spends extra inference budget on sampling, feedback-driven repair, and reasoning-guided implementation. Tree search wastes work when different trajectories reach the same program, and it struggles to choose among these actions under a fixed budget. GraphAHA organizes the search as a typed directed acyclic graph that merges equivalent programs into one node, sharing search statistics across every path that finds them, and uses hierarchical Thompson sampling to decide whether to generate a new state or follow an existing one and which operation to apply. On LiveCodeBench and CodeContests with Qwen2.5-Coder and DeepSeek-Coder, it achieves the best score in 18 of 20 cases and beats the strongest baseline on visible-test Pass@1 by 4.1 percentage points on average.

Cognition on Graph: Navigating Massive Knowledge Space via Cognitive Cycles and Bidirectional Graph-Text Synergy

Gengxian Zhou, Jian Xu, Zichen Tang, Shiming Xiang, Haihong E, Cheng-Lin Liu Retrieval-augmented generation (RAG) struggles to navigate large knowledge graphs and text corpora for complex reasoning, because existing methods follow graph topology reactively and connect graph and text only weakly. CoG (Cognition on Graph) is a training-free framework that runs a continuous plan-explore-reflect cycle: it formulates investigation plans, retrieves from both graph and text, and reflects on its progress to adjust strategy. Entities extracted from text guide graph exploration to fill knowledge gaps, so the structured and unstructured sources inform each other. On seven multi-hop question-answering benchmarks, it significantly outperforms state-of-the-art methods while exploring more efficiently.

RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States

Luca Herranz-Celotti, Vincent Guigue Linear attention and state-space models run in linear time but store their recurrent memory as a matrix, which limits the order of interactions the state can represent. RunningTensor generalizes this memory to an order-o tensor, updated by a rank-1 outer product and read by contracting it against o-1 query vectors, so order 2 recovers linear attention. The authors study order 3 as a proof of concept, keeping both recurrent and parallel forms and linear cost in sequence length while raising working memory capacity from O(W^2) to O(W^o). It outperforms linear-attention and state-space model baselines on synthetic multi-query associative recall, and after pretraining it also improves language-understanding and real-world retrieval tasks.

Parameter-Efficient Retrievers for Polish and European Languages

S{\l}awomir Dadas, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Micha{\l} Pere{\l}kiewicz Dense retrieval increasingly relies on multi-billion-parameter language models, which makes indexing, corpus updates, and low-latency serving expensive. The authors build compact retrievers with a three-stage pipeline of cross-lingual alignment, relational knowledge distillation, and contrastive fine-tuning, trained only on supervision generated by strong embedding models and rerankers rather than ground-truth relevance labels. They release PolDense, six Polish retrievers from 17M to 1B parameters, and EuroDense, a 435M-parameter retriever for nine European languages, both supporting 8,192-token contexts. Across 41 Polish and 150 multilingual tasks, PolDense-1B outperforms evaluated retrievers with up to 9B parameters, and EuroDense ranks first among models under 1B parameters and leads in seven of nine languages.

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

Mohammad Siavashi, Gerald Q. Maguire Jr., Dejan Kostic, Marco Chiesa cross-listed A single GPU utilization percentage can make LLM inference look compute-saturated while hiding how little useful work is done, especially during decode, where each request adds only one token and dense projections become small-row matrix multiplications. Profiling vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL across cold prefill, warm prefill, and decode, the authors show that on Hopper the bfloat16 matrix-multiplication path runs these operations in fixed 64-row fragments, so small-batch decode fills only a small fraction of each fragment with real token rows. They replace the single utilization number with eight counter-validated views derived from raw Nsight Compute reports. These views trace utilization gaps to concrete causes, including fragment fill, occupancy limits, stalls, and kernel selection, across four production models and six per-layer kernel roles.

SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading

Zihan Wang, Yuqi Wang, Lei Gong, Cheng Tang, Wenqi Lou, Teng Wang et al. cross-listed Mixture-of-Experts (MoE) models activate only a few experts at a time, so keeping most experts off the device could in principle approach full-load performance, provided the needed experts arrive in time. SeqMoE treats expert activation prediction as a sequence-modeling problem to forecast activations several steps and layers ahead. It schedules prefetching as a Job Sequencing with Deadlines problem and evicts cached experts using a probabilistic Belady policy driven by those forecasts. Its offloading runtime also keeps expert placement and orchestration compatible with end-to-end compute-graph capture, and with 45% of experts residing in device memory, SeqMoE averages a 96.97% expert hit rate and 80.22% of full-load performance.

Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage

Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Cullen Anderson, Narmeen Fatimah Oozeer, Jeff M. Phillips Binary-choice truthfulness benchmarks ask models to pick between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can beat chance without doing the intended reasoning. The authors show that in TruthfulQA, a simple six-feature logistic classifier separates correct from incorrect answers with substantial accuracy, and that similar artifacts appear in other benchmarks. They introduce Audit-Prune, a general mechanism that removes the answer pairs contributing most to this leakage so datasets can be cleaned before release. They also release a cleaned version of TruthfulQA with surface-feature leakage reduced to near chance.

MAxBench: A Multinomial Concept Recovery Benchmark

Divya Appapogu, Freya Behrens, Yonatan Belinkov, Aaron Mueller A single direction in activation space is often enough to steer a language model on a binary concept like refusal. Many concepts, such as animals or countries, instead have many subcategories and instances, and it is unclear which representation geometries and recovery methods suit them. MAxBench scores a recovered concept representation by sampling from it, without assuming any particular geometry, and compares 10 localization methods covering 5 geometry types across 6 concepts and 4 models. Affine subspaces steer more reliably and reach higher recall than rank-one or linear subspaces, mostly because of better non-zero offsets rather than the choice of basis, and manifold steering is competitive where it applies. Even so, no method consistently outperforms simple prompting, which matches earlier results on binary concepts.

Rethinking Heterogeneous System Disaggregation for Subquadratic Attention

Arya Tschand, Yaosheng Fu, Vikram Sharma Mailthody, Nicolai Oswald, Po-An Tsai, Ritchie Zhao et al. Frontier language models increasingly use subquadratic attention to cut inference memory and compute, but disaggregated serving systems still divide work on the assumption of dense attention. SubQuadratic Disaggregation (SQD) instead splits decoding into full-context and fixed-memory parts, matching each part to suitable hardware in heterogeneous systems. For sparse-attention models, top-k selection, which must scan the full key-value cache, runs separately from top-k attention and feed-forward layers. For linear and sliding-window models, the dense attention layers run separately from the subquadratic layers. On an adjusted 8xB200 proxy for a heterogeneous system, it improves tokens per joule by 31% to 56% over the strongest GPU-only baselines on GLM 5.2, Nemotron 3 Ultra, and Gemma 4 31B, and an analytical model of a Rubin plus LPX system predicts up to 3.6x higher throughput than attention-FFN disaggregation.

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

Zhiwei Li, Lei Zhu, Hao Gu, Xiang Hu, Yan Wang, Haitao Mi et al. Trainable sparse attention methods added to pretrained Transformers use a lightweight selector to score context tokens or blocks and then keep only the top K. That hard selection blocks gradients, so the selector is trained to imitate dense attention weights, which may not rank context by how much it actually helps predictions under a fixed budget. Simple Attention Sparsification (SAS) trains the selector directly on the language modeling loss by adding its continuous scores to the attention logits during training. The design depends on placing the gate inside the softmax in log form, normalizing gates against the always-kept current block, and keeping scores continuous, and a memory-efficient Triton kernel supports long-sequence training. Across reasoning, long-context, and agentic tasks, SAS beats trainable sparse attention baselines at every attention budget, with especially large gains under tight budgets.

Type Diversity Enables Transformers to Generalise Compositionally

Anssi Moisio, Mathias Creutz, Mikko Kurimo Earlier work found that Transformers generalize compositionally less well to new structures than to new words. The authors argue this comes from the datasets, which contain many distinct lexical types but few structural ones, rather than from the architecture itself. Using Grammatical Framework, they build linguistically diverse variants of the COGS and SLOG datasets that vary type diversity, defined as the number of different constructors of a type. They find type diversity correlates with compositional generalization equally for lexical and structural cases, which contradicts the earlier claim that compound divergence explains task difficulty. They also examine how the diversity of other types and the surface format of the logical semantics affect generalization.
14 more specialized papers

Agents 33

SoulAuth: An Actor-native Identity Architecture and Rust Reference Implementation for Humans and Long-lived AI Actors

Kun Yuan, Harold Wang, Echo Li, Egusi Gui, Kiki Hu, Lucas Luo et al. cross-listed As AI systems shift from transient model calls to long-lived actors that persist across credentials, clients, sessions, and runtime instances, identity infrastructure has to decide where the canonical continuity boundary sits. The authors introduce Actor-native Identity, arguing that any subject that must remain independently attributable should carry an ActorIdentity that is never replaced by an account, credential, client, session, binding, or runtime instance, and present SoulAuth, an open-source Rust reference implementation that treats humans and long-lived AI actors as first-class identity subjects while keeping authentication separate from downstream authority. A 'Philosophical Engineering' method translates conceptual analysis of subjecthood into identity objects, invariants, lifecycle semantics, and inspectable conformance evidence. Evaluation of the fixed SoulAuth v0.1.0 artifact shows partial rather than full architecture conformance: first-class human and AI actor status, client/actor separation, and authentication/authority separation are realized, while unified credential modeling and historical attribution anchored to the actor identity remain gaps.

Look Before You Leap: Pre-Action Verification for LLM Agents

Asaad Althoubi Agents that emit shell commands or code edits can fail silently, producing plausible but wrong effects that raise no error. The authors argue for cheap deterministic checks run before an action takes effect, fixing the action's correct effect by construction so that silent failure can be measured directly and the verifier can abstain rather than guess, and they study this across both modalities in one framework. A static shell verifier evaluated over 9930 commands and 482 tools catches 95.8% of invalid commands at a 10.0% false-positive rate, with syntax and binary checks producing zero false positives and every false positive traceable to help-text coverage in the flag check. On a benchmark of 640 code edits across 224 files, content-anchored formats such as search/replace and diff fail cleanly, whereas line-number edits corrupt 99.1% of files under a one-line shift and function-name edits hit the wrong function 12.7% of the time; a refuse-when-unsure anchor-and-verify applier records one silent misapplication in 8320 trials, and the benchmarks, verifiers, and guards are all released.

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Wenhui Chen, Shiwen Cheng, Hao Dong, Chenda Duan, Ruixiang Feng, Zhong Guan et al. Co-work agents run long workflows of information gathering, tool use, coding, and file manipulation across many model calls, so cost and latency accumulate and many steps demand state tracking and recovery rather than frontier-scale reasoning. Occamy-1.0 further trains the post-trained Qwen3.6-35B-A3B checkpoint using execution-grounded data and environments, replayable long-horizon trajectories captured across multiple harnesses, and staged post-training to build and consolidate execution skills. Across co-work benchmarks it is among the strongest comparably sized models and competitive with much larger frontier systems on several tasks, and under the authors' pricing protocol it sits at the low-cost knee of the cost-performance Pareto frontier across four benchmarks; the weights and a subset of the training data are released.

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

Mohsen Arjmandi Agentic coding systems pair a model with a harness of tools, prompts, and control flow, and practitioners assume a vendor's native harness solves more tasks with its own model. The authors test this with paired same-model contrasts on a private, contamination-controlled suite of 256 repository and post-cutoff contest tasks, running 80 tasks under claude-agent-sdk versus deepagents on claude-opus-4-8 and under the openai-codex SDK versus deepagents on gpt-5.5. Neither contrast shows an average advantage for either harness, with differences near 1 percentage point and confidence intervals spanning zero, though for Opus the native harness trails by 9 points on repository tasks and leads by 24 points on contest tasks in a post hoc split the authors say needs replication. The neutral harness cost 1.2 to 1.6 times more per solved task by observed usage, but missing usage records leave the billed ordering unresolved.

Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents

Hazel Mak, Susheel Suresh, Sahil Bhatnagar, Barry Wang, Chhaya Methani, Alejandro Gutierrez Munoz cross-listed Shell-based agents do well on coding, but enterprise work also involves moving between applications, coordinating with coworkers, and professional analysis, so it is unclear whether a general shell beats specialized typed tools there. The study compares five tool interfaces, namely typed tools, typed tools plus bash, bash alone, bash with persistent agent-synthesized tools, and programmatic tool calling (PTC) restricted to a typed catalog, on TheAgentCompany and APEX-Agents using Opus-4.8 and GPT-5.5. Bash alone beats typed tools by 21.8 to 24.5 points on TheAgentCompany and 4.8 to 7.4 points on APEX-Agents while using 19 to 72 percent fewer tokens, adding typed tools or tool synthesis to bash gives no detectable pooled gain, and PTC saves tokens versus direct typed calls but generally trails bash in both quality and cost efficiency.

When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline

Stefan G. Creadore, Peyton Woakz cross-listed Agent evaluations can be arithmetically correct yet measure a different construct than their labels imply. A retrospective audit of the Praxa AI pipeline's implementation files, historical evaluation artifacts, and operational records finds a 139-case routing report with 27 failures but zero gating failures because known gaps are exempted from the gate, an 8,843-row tool-attempt export whose pooled recorded 99th-percentile duration of 2,147,483,647 ms is a clamped 32-bit sentinel from abandoned clients, versus 38,118 ms for server-observed completed calls, and a compaction pilot whose 94.39 percent follow-up input reduction shrinks to 46.54 percent once the trigger call is included. The authors reproduce the descriptive calculations, verify 91 timing statistics with a separate weighted rational-arithmetic implementation, and release a reusable verification package, while noting that the full pipeline was not independently reproduced and no general capability claims are established.

Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering

Alexander Krentsel, Shubham Agarwal, Mert Cemri, Shu Liu, Sidharth Sankhe, Ziming Mao et al. cross-listed Software development runs an implementation-verification loop in which an evaluator such as a test suite checks an implementation against requirements under some model of the deployment environment, but requirements only approximate stakeholder intent and the model only approximates reality. The authors name these the requirement gap and the model gap, and argue the resulting two-gap framework unifies the main failure modes of agentic software engineering: reward hacking exploits omissions in the requirements or the model, while hallucination widens both gaps by fabricating requirements or environment assumptions. Since neither gap can be certified closed in an open, changing world, they propose an assurance-revision loop that uses deployment evidence to revise requirements, models, or evaluators when stakeholders reject behavior, and frame assured agentic development as allocating human judgment, agent capability, and compute across the two bottlenecks.

MAIA: Multi-Agent Intent Articulation for Requirement Discovery in Art Commissions

Yu-Chao Wang, Yanhong Lu, Yingjie Victor Chen, Tim McGraw cross-listed In bespoke art commissions, laypeople often know what they feel but cannot articulate medium, scale, or palette, an articulation bottleneck at the requirement-discovery stage that precedes any artist or image generator. MAIA (Multi-Agent Intent Articulation) is a multi-agent system that scaffolds this stage through Socratic questioning under a Verification over Invention rule, producing a text-only brief of visual terms the user has verified, with a validator gate that blocks unratified content. In a within-subjects study with 16 participants, the full configuration produced a large, significant gain in Cognitive Support over a minimal baseline (r = 0.96), and a blind review by three professional concept artists found AI rewriting improved visual completeness and executability in all eight sampled tasks, with directionally larger but underpowered gains under MAIA.

Local Edits, Global Ripples: Replay-Informed Policy Adaptation for Workflow Synthesis

Manqing Mao, Hong Wang, Samson Koelle, Jie Yuan, Zhuoer Wang, James Feng et al. Editing the prompt policy of an agent that synthesizes executable workflows is a cheap way to improve it without retraining, but an edit confined to one policy segment can ripple through downstream execution, and edits that help in isolation can interfere once composed. RIPPLE (Replay-Informed Persistent Policy Localization and Editing) diagnoses failed trajectories, maps each actionable failure to a predefined policy segment, restricts the fix to that segment, scores candidates against the iteration-start policy, and then replays promising edits on top of previously accepted ones, keeping only those that remain safe under composition. On the Flow-HO synthetic workflow-synthesis benchmark, validation success improves by up to 23.1%, with positive gains on two additional frozen language-model backbones, and interaction analyses show a segment-local tool-use edit altering downstream resource resolution and an edit that turns harmful after composition.

Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?

Obinna I. Ekekezie In draft-verify-revise pipelines, where one large language model drafts, a second critiques, and a third revises, a context-dependent expression like 'previous' can be resolved differently at each stage, a deictic shift that changes what it refers to. A synthetic dataset of 10 base examples in three conditions varied whether the drafting assistant or the grading verifier resolved the expression correctly and how much independent reasoning the revising meta-evaluator needed, tested across six models from three providers and 21 reasoning-effort configurations using e-values for sequential testing. Balanced accuracy ranged from 0.156 to near perfect: GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest effort, while Gemini 3 Pro stayed above 0.94 at every level and beat GPT-5.2 at xhigh effort for about 5% of the cost per trial; erring meta-evaluators tended to rely on surface cues, so the authors advise making the intended referent explicit at each stage.

NDT Factory: Synthesizing Verified Network Digital Twins from Semantic Models via Multi-Agent LLM

Sudipta Acharya, Petar Djukic, Burak Kantarci cross-listed Autonomous network management at TM Forum Level 4 requires evaluating Network Service Intents (NSIs) under varying conditions without hand-written analysis logic, but existing behavioral Network Digital Twins (NDTs) depend on predefined analytics. The NDT factory is a multi-agent system that uses large language models to synthesize executable behavioral NDTs on demand from semantic models through parallel synthesis and orchestration. In a Call Admission Control (CAC) case study, generated twins achieved 100% compilation and test pass rates across multiple runs, and simulation over 300 intents showed 99.3% decision agreement with a reference implementation, a 90% admission rate, and correct attribution of every rejection.

WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy, Navin Goyal Enterprise question answering is hard because information is scattered across evolving, sometimes conflicting emails, chats, and documents, and existing benchmarks lack that complexity, favor short answers, and use unnatural queries. WinSyn is an automated pipeline that simulates months-long enterprise projects with up to 25 interacting employees across roles, generating synthetic email corpora together with long- and short-form questions and grounded gold answers that emphasize ambiguity and distributed information. Evaluating standard agentic Retrieval-Augmented Generation and Deep Research (DR) baselines built on frontier models, aggregate scores stay below 80% on every dataset, indicating substantial headroom for enterprise deployment.

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

Umesh Bodhwani, Thanh Tran, Kai Wei Teams increasingly select task-oriented LLM agents through a cheap offline gate in which persona-driven LLM user simulators converse with each candidate and an LLM-as-a-judge scores the transcripts. GAUGE is a reusable protocol that checks whether this gate's ranking agrees with a grounded, verifiable reward across 25 agents from six providers on τ²-bench and SimulatorArena, separating ranking validity from construct validity. Satisfaction ratings carry essentially no information about task success: 57.5% of conversations rated satisfied by a blind panel failed the customer's task, a pattern that held across five rater populations, both benchmarks, and every subjective dimension, and while the gate ranks agents reliably across wide capability gaps, its decision-disagreement rate jumps from under 1% on wide-reward pairs to 31% on near-equal strong agents. The authors propose a calibrate-then-trust cadence in which a judge-free completion bit serves as a zero-cost tripwire for truncation regressions.

AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems

Zachary Johnson, Nigel Boachie Kumankumah, Somya Chatterjee, Tejas Sathyamurthi, Min Chen, Xinyi Alice Li et al. Large language model (LLM) memory systems are typically scoped to a single user, so knowledge that could safely be shared across users of a multi-agent system goes unused. AIM (Agentic Interoperable Memory) is a unified memory framework that classifies each piece of information as private to one user or public to all, and enforces index-level access controls so private memories can only be retrieved by their owner while shared knowledge improves coordination and consistency. The authors also release MUMBench (Multi-User Memory Benchmark), a dataset of multi-user interactions with private and shareable information across four domains that evaluates retrieval, creation, update, and deletion operations. Over three independent runs, AIM reaches 96.0 percent visibility classification accuracy, 58.8 percent strict operation accuracy, and 70.5 percent state-aware operation accuracy.

ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents

Bowen Guan, Zhentao Yin, Yanming Shen Most agent benchmarks score final task success or individual tool calls, which says little about whether an agent can diagnose and recover from failures mid-execution, especially in multi-turn parallel tool use where one bad branch can cascade into others. ParaRecover is a process-level benchmark built on a taxonomy of 14 error types spanning planning dependencies, tool selection, and argument matching, with 10,626 instances at two difficulty levels, plus an SDE rubric that scores structural integrity, diagnostic reasoning, and evolutionary strategy during execution. Evaluations of more than ten mainstream large language models show that even state-of-the-art models struggle with multi-turn error propagation, implicit tool-use failures, and precise replanning, and the SDE rubric doubles as a supervision signal that improves agents' reflective recovery.

CueMem: Cue-Guided Context Reconstruction for Long-Term Conversational Memory

Changjian Wang, Rongzhen Li, Weili Guan, Shuming Shi, Quan Lu, Ning Jiang Long-term conversational agents must answer questions from extended dialogue histories, but feeding the full history is expensive and unreliable, while compressed memory units often drop the fine-grained evidence a question needs. CueMem treats extracted memory records as retrieval cues rather than self-contained evidence: at construction time it links each cue to its source turn, and at query time it retrieves relevant cues, maps them to source-turn anchors, and expands over a turn graph capturing temporal proximity and semantic relatedness to reconstruct a compact evidence context from the original dialogue. On LoCoMo and LongMemEval it consistently outperforms representative long-term memory baselines, and graph-based reconstruction recovers supporting evidence while cutting query-time input tokens and latency relative to the full-history setting.

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

Tong Ye, Kunyang Han, Guozhi Wang, Longqiang Luo, Zhifeng Ding, Yongxiang Zhang et al. Native end-to-end mobile GUI agents face three deployment gaps: sandbox training mismatches production, costly real-device failures go unused, and static benchmarks saturate. BlueLM-GUI, a 35B-parameter model with 3B active parameters, is built around a real-device flywheel: a dual-track data pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction and Derivation Module turns every trajectory into supervision, a three-stage recipe of continual pre-training, supervised fine-tuning, and agentic reinforcement learning runs rollouts on hundreds of real phones, and a quota-driven benchmark with three orthogonal axes is upgraded as the model improves. It scores 87.4 on MobileGUI-VBench, 5.1 points above the best closed-source model, and 84.9 on AndroidWorld, the best open-source result and competitive with closed-source systems.

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

Yu Bai, Yukai Miao, Dawei Wang, Li Chen, Yanyu Ren, Yuqian Shi et al. Verbal reinforcement learning, introduced by Reflexion, lets a language agent turn failed trials into text that guides later attempts without weight updates, but comparisons across such methods have not controlled for the number of trials available. VRL-Bench is a harness for evaluating trial-and-error learning under finite trial budgets, and across three models on MiniWoB and WebShop it finds that each prominent verbal-memory update beats memory-free retry in some settings but hurts in others, with replay experiments showing that reflection can lower success by over-exploiting past experience at the expense of exploration. The proposed VEX² scheduler uses a language model to jointly choose policies and allocate the remaining trial budget, and it is the only evaluated update that improves over plain retry in all six settings.

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

Hanyu Zhao, Yuqian Feng, Zhenyu Song, Yuanchao Cheng, Yance Jiao, Tengfei Pan et al. Long-running large language model agents need memory that keeps coherent internal state across interactions, and when temporary context is written into the same store as durable knowledge, transient information can overwrite what should persist and cause behavioral drift. The authors study a lifecycle-labeled memory setting in which write episodes carry lifecycle metadata during training and evaluation uses phase-aware readout, and introduce LifeFuse-Mem, a neural memory framework with dedicated components and lifecycle-aware updates so that stable and transient knowledge evolve locally without converting temporary context into durable state. On a controlled anti-overwrite benchmark, LifeFuse-Mem improves acquisition-controlled retention and reduces temporary overwrite, and it remains broadly competitive on two public long-memory benchmarks, suggesting explicit lifecycle signals help diagnose and mitigate overwrite in compact online memory.

Diverse Minds, Divided Networks? Personality Composition, Polarization, and Collective Intelligence in LLM-Based Social Simulations

Raad Bin Tareaf cross-listed Simulated societies of large language model agents are used separately to study online polarization and collective intelligence, so it is unclear whether personality composition shapes both or whether reducing polarization costs collective competence. TraitMix treats the Big Five composition of a hundred-agent simulated social network, both trait levels and trait heterogeneity, as a controlled experimental variable and measures polarization and collective performance in the same runs, across 991 simulations spanning six contested topics and six language models. Trait heterogeneity has the largest effects and pushes two faces of polarization in opposite directions: varied societies hold more dispersed opinions but are less segregated into camps, so homogeneous societies behave as consensual echo chambers rather than moderate ones. Trait effects are non-additive, with Agreeableness determining the sign of Openness across models, and contrary to the expected trade-off no polarization measure predicts poorer collective performance, with cross-cutting interaction the only measure whose association with collective accuracy survives controlling for aggregation identity.

Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

Yu-Yu Yang, Ti-Rong Wu, Hung Guei, Hsing-Yu Chen, I-Chen Wu Social-deduction games such as Werewolf are increasingly used to evaluate large language model agents, but most evaluations rely on final game outcomes rather than examining how communication changes beliefs. The authors propose a belief-shift benchmark: using LLM-played games, they annotate suspicion and accusation messages and measure how an observing village-side model's suspicion of each player changes after each message, evaluating 40 open-weight LLM configurations on 1,224 annotated messages. Larger models better distinguish true wolves from villagers from game history, but accusations still strongly sway beliefs, making models more suspicious of the accused and less suspicious of the accuser, especially when the accuser is trusted, even when that accuser is wolf-aligned; larger models resist accusations only from accusers they already distrust. The results suggest open-weight models up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication.

SAGE-Loop: Reliable Closed-Loop LLM-Driven AutoML with Trial-and-Correction and Adaptive Ensembling

Junquan Gu, Shibo Cui, Xiangfeng Luo, Hang Yu As large language models are introduced into automated machine learning (AutoML), pipeline reliability becomes as important as automation efficiency, yet existing systems run one-way pipelines in which intermediate failures are terminated or bypassed and ensemble decisions stay static regardless of the models generated. SAGE-Loop is a closed-loop, self-adaptive, LLM-driven AutoML framework that performs multi-round generation and validation for trial-and-repair of failed or suboptimal steps, and adaptively selects ensemble strategies from evidence about model diversity in both supervised and unsupervised tasks, unifying how models are generated with how they are used. Across 20 public datasets, SAGE-Loop consistently improves performance and stability on classification, regression, and clustering, and additional experiments show it recovers from execution failures while maintaining robust pipeline behavior.

Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup

Julian Varghese, Lucas Bickmann, Sarah Sandmann Multi-agent LLM systems commonly assign specialized roles, but it is unclear whether specialization yields genuinely different forecasts or whether critic and synthesis stages add value. The authors ran a live, prospective study over the final 56 matches of the 2026 FIFA World Cup, holding a frontier model fixed while a quantitative specialist worked from structured performance statistics and a news specialist from injuries, tactics, and press conferences, followed by a critic and a meta-agent that combined the forecasts, with betting-market odds as an external benchmark. The news specialist scored highest on probability-weighted Top-3 utility and matched the betting market on exact-score hits, but the two specialists agreed on at least two of three scorelines in 50 of 56 matches and the meta-agent never produced more than one scoreline outside the specialists' set, so the critic and meta-agent stages did not contribute complementary information or beat the strongest specialist.

Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents

Zhutao Lv, Chenhao Dang, Yi Feng, Yanpei Gong, Xiaolei Wang, Junyan Ye et al. cross-listed Real-world Earth observation (EO) agents must turn a scientific question into an executable workflow that acquires observations, prepares data, runs domain computations, and derives conclusions from runtime evidence, yet existing agents and benchmarks start from prepared inputs or candidate answers. Earth-Agent-Pro is a plan-and-execute framework that uses expert-authored skills to constrain planning and tool use, keeps a workflow-centered structured memory so only the affected suffix is repaired when runtime evidence invalidates a step, and trains separate adapters with sequence-level supervised fine-tuning for planner workflow composition and node-level group relative policy optimization (GRPO) with locally verifiable rewards for executor tool-argument grounding. On the new Earth-Bench-Pro, which instantiates 248 expert-curated task cores as 744 questions including open-world execution tasks, the framework with a GPT-5 backbone reaches 66.13% LLM-as-judge accuracy, 20.95 points above ReAct, and joint adapter tuning lifts Qwen3.5-9B from 38.31% to 50.00%.

LifeMem: Enabling Lifelong Experience Reuse for LLM Agents

Yuli Qiu, Yutong Li, Wei Su, Zeming Liu, Wanxiang Che, Heyan Huang et al. LLM agents are expected to improve over their lifetime by reusing past experience, but existing memory-based agents transfer poorly across environments and suffer catastrophic forgetting as experience accumulates. LifeMem is a lifelong learning framework that clusters accumulated interaction trajectories by their underlying workflows to extract reusable skills, then recalls relevant skills and trajectories to guide actions on new tasks. Experiments span 10 environments and more than 13,000 tasks, including 2,000 newly annotated trajectories, and show reduced forgetting on learned tasks together with better cross-task transfer. Further analysis finds that the order in which tasks arrive affects learning, and that consolidating structurally similar trajectories in memory improves performance.

Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

Mykhailo Kozyrev, Andrei Kozyrev, Anton Podkopaev Coding agents increasingly draw repository knowledge from SKILL files, plain Markdown documents versioned with the code, and prior work optimizes these documents against synthetic benchmarks that capable agents solve even without them. The authors mine harder tasks by reverting a repository's merged pull requests at a single frozen base commit, then score a candidate document by whether the same agent does better with it than without it. On three Kotlin repositories, documents found by GEPA raised the score by 4.9 percentage points on average while SkillOpt gained only 0.1 points, but with the few tasks one repository's history supplies, the GEPA gain cannot be separated from the agent's run-to-run variance. A maintainer of one repository nevertheless found that the generated documents held knowledge usually gained only by working in the project.

What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework

Ioannis Prokopiou, Athanasios Aidinis, Panagiotis-Christos Kyrmpatsos, Pantelis Vikatos Agentic pipelines for generating structured queries report gains, but it is unclear which part of the loop produces them. Using LAST-CQ, a five-agent, training-free, execution-grounded Text-to-Cypher framework, as an instrumented testbed, the authors run three counterfactual experiments over 2,471 live-database queries with six backbone models. Replacing schema-grounded, LLM-synthesized feedback with raw database error strings costs almost nothing, and spending the same call budget on parallel sampling lowers quality by 10-11%, so the gain comes from detecting failures and routing them to a retry; the full system recovers 91.7% of queries that fail under single-pass generation while still using one call for queries that succeed first time. The study also finds that n-gram overlap on serialized results is unreliable in both directions as a metric, and that its LLM judge scores 9 points more optimistically than blind human labels.

Online Video Agent Harness for Long Video Understanding

Sen Yang, Boqiang Duan, Jing Yang, Weihao Bo, Jie Liu, Boyuan Tong et al. cross-listed Answering questions about long videos resembles a visual needle-in-a-haystack search, and packing dense frames into one vision-language model (VLM) context causes context rot and high cost. VideoXAgent is a purely online agent harness that plans and decomposes the user query, calls specialized expert tools on demand (scripts, VLMs, and domain models for detection, text recognition, speech recognition, and face recognition) drawn from a data-driven taxonomy of atomic capabilities, and aggregates the multimodal evidence while resolving conflicting observations, with objective evidence prompting and budget-aware control to curb hallucination and non-termination. Across Video-MME-Long, LongVideoBench-Long, LVBench, and MINERVA it is competitive with frontier large multimodal models and video agents while using about 50k tokens of agent context per sample, and on MINERVA it matches that level with only about 15% of the context of a 1,024-frame dense-packing baseline. The harness stays effective even with a visually weak or text-only orchestrator, suggesting that progressive evidence seeking can substitute for fitting the whole video into one context.

Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks

Sebastiano Nordio, Michele Lotto cross-listed Open-weight small language models (SLMs) run locally can sidestep proprietary API guardrails, yet as autonomous agents they struggle with long, exploratory tasks like cybersecurity Capture The Flag (CTF) challenges because accumulated tool-call outputs bloat their context and degrade reasoning. The authors propose context segmentation, a two-level agentic framework that splits a complex exploitation task into manageable, contextually isolated sub-problems. On the picoCTF dataset with memory-constrained gemma-4 models, the strategy acts as an intelligent search for the E4B model, matching brute-force retries on reward with better token efficiency and solving 18.52% of tasks that standard agentic execution fails to complete.

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

Junghyun Min, Huseyin Uzunalioglu, Mohamed Trabelsi Fully autonomous end-to-end machine learning (ML) research systems have mostly succeeded on problems with narrow search spaces. This case study tests them on an open-ended industrial task, telecom ticket retrieval, where representation, architecture, and training-data generation are all open choices. With commercial and open-source agents, including Cursor, autonomous research did well at narrow hyperparameter optimization but lacked human-like intuition and creativity and added operational overhead. With minimal human supervision it reached about 90% of state-of-the-art performance (0.34 vs. 0.38 Recall@1) in 10 weeks rather than 10 months of human work, at up to $200 per Cursor campaign, and the authors recommend having human researchers and autonomous frameworks work together.

Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji et al. Agent-based tools for building embodied AI benchmarks usually cover only some construction stages or a fixed set of environments. They also pass intermediate artifacts downstream without checking them, so local defects end up in the final benchmark. Embodied-BenchForge turns a user's evaluation intent into a complete benchmark by composing typed, reusable skills into executable workflows and recording intermediate outputs in an artifact dependency graph. Artifact-specific checks run throughout, and a failed check triggers local re-execution or rollback to an earlier stage. The framework built six offline question-answering benchmarks and one interactive benchmark with 220 executable tasks, which separate the capabilities of multimodal LLMs and embodied agents, and ablations show that verification and repair improve benchmark quality and that skills can be reused across benchmarks.
2 more specialized papers

Theory 30

The Cost of Compression: A Rate-Distortion Limit on Factual Hallucination

Xi Wang, Shijia Xu, Rongfeng Guo Factual hallucination in closed-book question answering is usually blamed on coverage, the fact never having been seen, but finite memory can also force observed facts to be stored only approximately. The authors study a coverage-compression model in which a learner sees M of N possible queries, compresses them into B bits, and answers without retrieval, and prove a lower bound on error that splits cleanly into a compression-distortion term on observed facts, governed by the inverse rate-distortion function of a uniform K-ary source, and a coverage term on unobserved facts. The bound gives a compact way to reason about selective memory, structure, retrieval, abstention, and long-context organization, and its predicted signatures are checked with simulations and controlled fact-injection probes in modern language models that vary fact load and trainable memory.

High-Probability Convergence of SGD via Batched Updates

Feng Zhu, Robert W. Heath Jr., Aritra Mitra cross-listed High-probability guarantees for the last iterate of stochastic gradient descent (SGD) have typically required restrictive assumptions, such as bounded domains or gradients, or complex proofs with auxiliary sequences. Batched SGD partitions online samples into epochs and makes a single update per epoch using a refined, low-variance gradient estimate. This batching allows a simple analysis that gives near-optimal high-probability rates for strongly convex and non-convex objectives under standard smoothness and norm-sub-Gaussian noise assumptions. The idea extends to federated learning, yielding the first high-probability guarantees for federated learning, with logarithmic communication complexity, linear speedup in the number of agents, and resilience to data heterogeneity.

Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H\"{o}lder Smoothness

Misbah Uz Zaman, Anirbit Mukherjee Standard convergence guarantees for stochastic gradient methods assume Lipschitz-smooth objectives and finite-variance gradient noise, and both assumptions are often violated in practice. The authors analyze nonconvex optimization when gradients are only Hölder continuous with exponent s in (0,1] and the noise has only a bounded α-th moment, with α between 1 and 2. They show that plain stochastic gradient descent (SGD) keeps a rate of O(T^{-s/(1+s)}) whenever α ≥ 1+s, derive a stationarity rate for δ-regularized gradient clipping (δ-GClip), and show that standard gradient clipping (G-Clip) recovers that rate in the same regime. In the very heavy-tailed regime α < 1+s, G-Clip attains the first convergence guarantee for any stochastic gradient method.

Benign Loss Landscapes Can Coexist with Worst-Case Hardness

Zach Furman, Stephan W\"aldchen, Yangda Bei, Liam Hodgkinson Neural networks can represent targets that are cheap to evaluate yet cannot be learned by gradient descent in polynomial time, which raises the question of what structure makes real-world targets learnable. Existing surrogate models cannot pose this question, because they either lack such hard targets or cannot evaluate them efficiently. The authors study tree tensor networks (TTNs), which generalize deep linear networks and Tucker decompositions, show that they embed arbitrary read-once Boolean formulas and so contain polynomial-size hard targets, and prove that the loss landscape is nonetheless conditionally benign for every realizable target: every minimum-norm local minimum is global. Learning difficulty instead comes from high-order degenerate saddle points caused by rank deficiency, which the authors illustrate with a case study of the parity function.
26 more specialized papers

Other 27

One Simple Trick for Improving the Performance of Energy-Limited Local Inference and Training

Erik Schultheis, Maximilian Kleinegger, Dan Alistarh cross-listed Power delivery and heat dissipation constrain not only datacenter GPUs but also compact consumer devices such as the DGX Spark, where workloads that alternate compute-heavy matrix multiplications with memory-bound operations like normalization or cross-entropy can hit power or thermal limits and start throttling. The authors show that chunking the workload into smaller pieces, so that compute-bound and memory-bound phases alternate at higher frequency, smooths the power and temperature spikes and avoids throttling. On a DGX Spark this yields up to 2% faster wall-clock time together with lower total energy consumption across several scenarios, and the same effect appears on a less constrained multi-GPU server at a reduced 1-2% effect size.

Who Pays for Open Review? Visible Author Reputation and Its Effect on Ratings

Qinghua Zhao, Xinyu Chen, Yanhui Yang, Tengfeng Sun, Junfeng Liu, Zhongfeng Kang cross-listed A November 2025 OpenReview bug that broke anonymity at several conferences prompted calls for open review, raising the question of what dropping blind review would mean for authors. Analyzing over 18,000 reviewed ICLR 2026 submissions, split into de facto open and blind groups by arXiv preprint timing, the authors find ratings rise with author reputation under both regimes but with a statistically significantly steeper slope under open review, concentrated at borderline ratings. The pattern holds across five reputation proxies, three author-aggregation rules, and five open-window definitions, and a controlled simulation with five AI models as reviewers reproduces it, with claude-opus-5 raising its rating by 0.5 points when the author moves from low to high reputation.

Dual-guided Hierarchical Edge Localization for Large-scale Optimal Transport Across Dimensions

Wenzhou Xia, Qiaoqiao Ding, Jingwei Liang, Xiaoqun Zhang Unregularized discrete optimal transport (OT) requires solving a linear program whose number of variables grows quadratically with dataset size, which makes exact solutions impractical at scale. HELLO recasts the problem as locating the few edges that carry mass. Dual potentials computed on a recursive subsampling hierarchy seed candidate edges from coarse to fine, and a refinement step repeatedly inserts the largest dual violators in each row and column until a relative KKT residual tolerance is met, while budgeted pruning keeps memory linear. The authors prove finite termination at a global optimum for exact-arithmetic refinement, report lower objectives with order-of-magnitude speedups over strong baselines at the million-point scale, and show the solver handling 1.28 million samples per marginal in 8192 dimensions on a single H100 with 41.6 GiB of GPU memory.

Attention Quantization for Tabular Foundation Models

Jonas M. K\"ubler, Benjamin J\"ager, Klemens Fl\"oge, Noah Hollmann, Frank Hutter Tabular foundation models share the transformer architecture of large language models (LLMs) but differ in size and serving patterns. The authors argue that their inference optimization should target the attention computation rather than the weight or KV-cache quantization popular for LLMs, and they quantize queries, keys, and values to FP8, running explicit FP8 matrix multiplications in a custom Triton kernel. A key finding is that quantization error on test rows must be aligned with that on training rows, because a mismatch makes accuracy drop sharply. The kernel runs up to 1.7x faster than standard 16-bit kernels with no meaningful accuracy loss for TabPFN-v3 and TabICLv2 on TabArena and BeyondArena.

Transfer Learning for Evolving Domains

Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Pedro Ribeiro, Pedro Saleiro, Pedro Bizarro et al. Transfer learning research is split into sub-areas such as domain generalization, domain adaptation, and multi-domain learning, and each assumes a fixed amount of target data and labels. Deployed systems, however, see data and labels from a new domain accumulate over time. The authors formalize this as Transfer Learning for Evolving Domains (TrED), defined by an environment-controlled data availability process, a learning protocol the method chooses, and an evaluation criterion that scores the whole trajectory of models rather than one snapshot, so the classical settings become regimes a learner passes through. A review of existing methods finds that most are tailored to a single regime and that even the strongest candidates do not optimize the whole trajectory, which leads the authors to present TrED as an open problem.
22 more specialized papers

Multimodal 19

Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture

Cody Kommers, Mingrui Ye, Evelyn Gius, Daniela Mihai, Hoyt Long, Zheng Yuan et al. Human communication often uses ambiguity deliberately, producing expressions open to several readings yet still interpretable, a property the authors call calibrated ambiguity. They study it with a task drawn from the parlour game Dixit, comparing picture clues written by humans and by multimodal language models using a new coding rubric. Models consistently show ambiguity collapse, producing over-specified clues that leave no room for multiple legitimate interpretations. They also show cultural flattening: unlike humans, they almost never draw on culturally situated references, even when prompted to use allusion and figurative language.

Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving

Zhitong Dong, Jicai Pan, Yingguo Gao, Jingting Ding, Hao Chen, Jinjie Gu Hard geometry problems often require visual steps such as drawing auxiliary lines, which has motivated Visual Chain-of-Thought (VCoT), yet evaluations usually score image generation and final answers separately without checking whether intermediate visual aids are valid, actually used, or responsible for success. GeoVAD-Bench diagnoses reasoning trajectories along five dimensions (perception, auxiliary quality, utilization, deductive reasoning, and final correctness) and compares settings with no auxiliary aids, model-generated aids, and ground-truth aids. The analysis shows that good auxiliary constructions help substantially in principle, but autonomous generation suffers from compounding errors in geometric perception, faithful diagram editing, grounding in the visual state, and deduction. A data pipeline and a progressive supervised fine-tuning plus multimodal reinforcement learning recipe built on these findings produce GeoWeave-8B, which improves final geometric accuracy by 25.3% over its base model and averages a 30.4% gain across the four intermediate diagnostic dimensions.

SteerDuplex: Steerable Duplex Speech Dialogue Models

Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai, Sonal Kumar, Nikhil Barhate, Isabell Sagar et al. Full-duplex spoken dialogue models handle turn-taking, interruptions, and backchanneling, but a taxonomy introduced by the authors reveals large gaps in steerability, the ability to reliably shift tone, persona, speaking rate, or voice style on user instruction. SteerDuplex fine-tunes the Moshi full-duplex model on natural and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction, then applies two-stage reinforcement learning (RL) with rewards that mix verifiable interaction checks and judge-based semantic feedback. On the new SteerBench, with 390 spoken prompts and 1,067 human-written audio and text rubrics, supervised training raises the audio-steering average pass rate by 44.5 percentage points over the strongest open baseline, and RL raises clean interruption response from 72.5% to 82.5% while cutting pause barge-in from 26.5% to 9%. Reward probes also expose reward hacking through incomplete responses, showing that timing gains must be evaluated alongside response completeness.

ProactiveBench: Can Streaming Video Models Really Interact Like Humans?

Kaixuan Du, Xin Wan, YuKun Wang, Hang Zhang, Meng Cao, Dai Guan et al. Evaluations of streaming video models mostly query the model at a chosen timestamp, so they never test whether the model knows when it should respond. ProactiveBench instead evaluates models at one-second intervals with no explicit response cue, requiring them to monitor a standing request, respond within an appropriate window after the target event, and stay silent otherwise. Its six subtasks vary how ambiguous the trigger is and how much timing tolerance is allowed, using metrics that combine response and silence rates, separate early, in-window, and missed responses, and penalize omitted or repeated counts. Premature responses outnumber missed responses for four of the six evaluated systems, revealing a substantial gap in the temporal decision-making needed for human-like interaction.

What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

Lucia Cascone, Valeria Fraenza, Michele Nappi, Fabio Narducci, Benedetto Simone cross-listed Audio-based Multimodal Large Language Models (MLLMs) can caption complex acoustic scenes, but it is unclear which parts of the audio support each generated token, especially when overlapping sounds occupy different frequency bands. STAG is a post-hoc framework that estimates each token's temporal support by projecting the encoded audio representations onto the vocabulary for that specific token, measures frequency-band relevance through controlled spectral occlusion, and combines both into a spectro-temporal relevance map. Against ten post-hoc explanation methods on four grounding benchmarks, it achieves the best event localization on every dataset, and it works on eight audio-language backbones without parameter updates. Counterfactual deletion of the identified evidence selectively lowers confidence in the corresponding event and often removes it from the regenerated caption.

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi et al. Speech-to-speech translation (S2ST) systems built on speech LLMs struggle to predict high-bitrate speech tokens and depend on training data whose speaker identity and prosody are well aligned. Kraken outputs low-bitrate tokens from single-layer vector quantization trained to reconstruct self-supervised learning (SSL) features. A separate token-to-waveform decoder, Autowave-X, is also conditioned on the source speech so that non-linguistic traits carry over, which relaxes the training data requirements. Built on Qwen3-8B and trained on 150k hours of multilingual, multitask speech data, the model outperforms SeamlessM4T-Large v2 and Qwen2.5-Omni in translation quality and transfers speaker identity and prosody better.

MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

Yi-Jen Shih, Shih-Yun Shan Kuan, Guan-Ting Lin, Kai-Wei Chang, Siddhant Arora, Shu-wen Yang et al. cross-listed Benchmarks for conversational voice agents mostly cover one-on-one dialogue or passive audio comprehension, overlooking the common case of an agent taking part in a group conversation. Multiparty Bench (MP-Bench) evaluates speech systems as active participants in multi-party conversations on two main dimensions, turn-taking awareness and response appropriateness, with comprehension-based question answering as a complementary test. Across 12 voice agents, real-time systems score at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking.
12 more specialized papers

Reinforcement Learning 18

Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning

Adam Haroon, Cody Fleming Safe offline reinforcement learning normally assumes a cost label on every transition, but here safety can only be judged by comparing short clips and occasionally asking whether a whole episode exceeded its budget. Certified safety curation trains a state-only value from segment comparisons to score whole trajectories, uses Learn-then-Test calibration to certify a selection threshold with a distribution-free bound on the unsafe fraction of the selected set, and then behavior clones on that selection; oracle controls show reweighting individual transitions fails even with an exact value, so selection operates on whole trajectories. The resulting policies meet the cost budget on 11 of 15 DSRL tasks, one short of cloning the ground-truth safe subset that needs a label on every trajectory, and retraining the strongest full-label method on the certified selection makes it safe where no setting of its own cost target does.

Inverting Self-Triggered Control: Adversarial Reinforcement Learning for Sparse Denial-of-Service Attacks

Adam Haroon, Erick J. Rodr\'iguez-Seda, Tristan Schuler, Cody Fleming Self-triggered reinforcement learning control (RL-STC) learns the sparsest control schedule that keeps a Lyapunov function decreasing under a run-time assurance override, and the authors invert it: an adversarial RL agent learns the sparsest jamming or denial-of-service (DoS) schedule that destabilizes the loop, using a Lyapunov-increase admissibility predicate that mirrors the defender's certificate. They prove a plant-property lower bound on the jam count an immediate hold-last adversary needs to crash a self-triggered controller satisfying a Lyapunov contract, recovering the consecutive-grouping optimality of prior count-budget DoS scheduling as a corollary. Against one LQR and three RL-STC defenders on Pendulum, CartPole, and Quadrotor2D, the learned adversary is the only one that crashes every defender on every plant 100 percent of the time, beats greedy and periodic baselines by up to 2.8x on jam time per failure, and stays ahead under heavy observation noise and position-only observations.

Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning

Fernando Palafox, David Fridovich-Keil World models let agents plan by predicting the consequences of actions, but they become inaccurate when the test-time environment differs from training, and existing adaptation methods trade cheap but inexpressive in-context learning against expressive but costly gradient updates. CLAW (Context-conditioned Low-rank Adaptation of World models) uses a hypernetwork to generate low-rank adaptation (LoRA) adapters for a frozen base world model from a small batch of test-time transitions in one forward pass, with the hypernetwork and base model pretrained jointly by simulating adaptation across a family of environments. Using only seconds of test-time data, CLAW outperforms gradient-based adaptation and in-context learning on locomotion and manipulation families that vary in dynamics, embodiment, and reward. Ablations attribute the advantage to the expressive adapters rather than context conditioning, show it avoids overfitting when data are scarce, and find joint pretraining beats training the hypernetwork post hoc.

MInTRL: Off-policy Intervention can boost On-policy RL

Mingyu Chen, Yefan Tao, Gerald Friedland, Xuezhou Zhang, Chris Kong Reinforcement learning with verifiable rewards is usually run on-policy, which keeps training data close to the current policy but limits learning to trajectories the policy can discover by itself, whereas off-policy methods such as supervised fine-tuning can inject external knowledge but suffer from distribution shift. Minimal Intervention Reinforcement Learning (MInTRL) has a judge-intervention policy periodically review the policy's output during generation, replace erroneous suffixes with short corrections, and immediately hand control back, and it trains with a sequence-level advantage-regression objective that avoids importance sampling. Sparse local interventions substantially improve coverage beyond finite-budget on-policy sampling while keeping trajectories largely on-policy, and MInTRL consistently outperforms standard on-policy and off-policy baselines across math and code benchmarks. Ablations show the method still works with self-intervention and across different judge policies, with performance peaking at moderate intervention intensity.

Hierarchical Belief Modeling for Zero-Shot Opponent Adaptation in Partially Observable Multi-Agent Navigation

Kowei Shih, Lu Cheng, Zeyu Wang, Yeyun Xu, Kejian Tong The Lux AI Season 3 competition requires agents to act under partial observability, randomized per-episode dynamics, and a best-of-five match format that rewards both tactical execution and fast adaptation to the opponent. HORIZON is a hierarchical agent combining symmetry-aware spatial perception, dual-memory belief tracking, relic-centric graph attention, information-gain-driven exploration, and an opponent-conditioned policy mixture, separating short-horizon control from cross-match meta-reasoning and using auxiliary belief and world-model objectives to stabilize learning. Trained with PPO in a large-scale JAX simulator, the agent explicitly infers hidden game parameters and opponent style, and shows consistent gains in match win rate, episode win rate, adaptation gain, and league rating over strong recurrent and feed-forward baselines.

Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

Taoran Liang, Yang Liu, Shang Luo, Yingguang Yang, Rongrong Zhang, Yingzong Min et al. Critic-free group-relative methods such as GRPO train large language model agents on long-horizon tasks by broadcasting one trajectory-level reward to every step, and GiGPO recovers a step-level signal by grouping steps that share an anchor state but blends step- and episode-level estimates with a single fixed weight, spending the same resolution on a pivotal branching decision as on a routine transition. GACA (Granularity-Adaptive Credit Assignment) makes the blend state-dependent: it scores each step by the negative log-likelihood (NLL) its rollout already records as an uncertainty-based criticality proxy, then weights the fine-grained step advantage more heavily at above-average NLL and the episode-level advantage below it. The authors derive an exact risk decomposition showing that sufficiently small modulation improves on fixed mixing under positive directional alignment, bound local action-value variation using expected NLL, and characterize through an error-projection analysis when mixing adds value beyond scalar uncertainty reweighting. On ALFWorld and WebShop, GACA improves task success over GRPO and GiGPO at both 1.5B and 7B scales.

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

Weiyuan Li, Aili Chen, Xintao Wang, Yikai Zhang, Qingqing Dong, Jinghan Xu et al. Open-ended reinforcement learning for tasks without verifiable answers typically relies on rubric-based rewards, but because the policy and reward system form a feedback loop, a reward that starts out useful can become unreliable through reward hacking or loss of response discriminability as the policy optimizes against it. EvoRS represents the reward system as an executable Reward-DAG and has an agentic designer update it from on-policy rollouts and reward traces during training, addressing failures not only in evaluation criteria, as existing dynamic-rubric methods do, but also in scoring mechanisms and signal composition. On writing and roleplay tasks, EvoRS achieves the best quality under all three judges, outperforming the baseline policy by 2.107 and 4.767 points respectively, while reducing reward hacking and coverage failures and preserving reward informativeness. Ablations indicate that even a comprehensive fixed reward system cannot stay reliable throughout open-ended training and must evolve.

SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy

Xiefeng Wu, Shu Zhang, Zhaojie Chu, Mingyu Hu Offline-to-online reinforcement learning cuts interaction costs for robot learning but suffers from unstable value estimates, and the authors identify an overlooked cause: the standard log-entropy term, which can turn negative. SCQ (Sigmoid-Bounded Conservative Q-Learning) replaces it with a sigmoid-bounded term that stays strictly positive, while keeping conservative Q regularization and return-based lower-bound calibration. On D4RL (Minari) benchmarks and on simulated and real visual tasks, it matches or exceeds baselines with more stable training, and it transfers to four real robot platforms spanning manipulation, wheeled, quadruped, and humanoid systems. Clipping experiments with gradient-matched controls suggest that keeping the term positive, more than the exact shape of the bound, drives much of the improvement.

Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies

Anna Rothenh\"ausler, Daniel Jost, Raghu Rajan, Faris Janjos, Oliver Scheel, Andreas Look et al. cross-listed Data-driven driving simulators let reinforcement-learning policies pick accelerations and steering rates from a fixed grid without limiting the resulting accelerations and jerks. As a result, policies can inflate safety metrics with abrupt last-second maneuvers that no passenger would accept. Simply clamping a static grid to comfort bounds fails because lateral limits shrink quadratically with speed, which saturates the grid and destroys fine control, so the authors instead rediscretize the action grid at every step to span exactly the feasible comfortable controls, using a closed-form inversion of the lateral-jerk constraint. On the Waymo Open Motion Dataset and a hand-built slalom, this adaptive action space keeps comfort violations below 1% while navigating better than clipped-grid and direct-jerk baselines, and the accompanying PufferDrive-Editor browser tool supports auditing realized kinematics and authoring hard scenes.

Robust Policy Optimization via Adversarial Importance Sampling

Amine Andam, Jamal Bentahar, Mustapha Hedabou Building deep reinforcement learning (DRL) policies that are robust to input perturbations involves algorithm design, implementation, and evaluation, and the authors tackle a limitation at each stage. Adversarial Importance Sampling (Advis) applies importance sampling to trajectories from standard training to estimate and optimize verifiable worst-case returns, capturing long-term robustness without extra environment interactions or auxiliary networks. They also release advrl, a modular PyTorch library of single-file implementations of robustness methods and adversarial attacks. Their evaluation study shows that optimal attacker hyperparameters do not transfer across agents, so testing against only a few attacker configurations can overstate robustness; they therefore evaluate against 6–14x more configurations than prior work and report that Advis compares favorably with existing baselines on continuous control tasks.

Expert-Space Exploration in MoE Reinforcement Learning

Hongyi He, Zhenghao Lin, Xiao Liu, Peng Cheng, Yan Lu, Yeyun Gong Reinforcement learning (RL) for Mixture-of-Experts (MoE) language models has focused on stability and efficiency while treating expert routing as fixed, even though perturbing the routing increases rollout diversity much like raising the decoding temperature. Naive perturbation activates unsuitable experts and hurts rollout quality, so Expert-Space Exploration Reinforcement Learning (ESRL) keeps high-confidence experts as anchors, routes stochastically only within a plausible candidate pool, and scales the perturbation by router entropy. To avoid a routing mismatch during training, it records the expert paths used in each rollout and replays them during policy optimization. ESRL performs best across top-K, top-1, and shared-expert backbones on math, science, and code tasks with no extra sampling or compute, and on Qwen3-30B-A3B it improves average Pass@1 and Pass@8 over GRPO by 3.2 and 4.5 percentage points.

CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models

Blake Olson, Yuhang Song, Emmett McQuinn, Yuan Shangguan Diffusion Language Models (DLMs) generate text in parallel but trail autoregressive models on complex reasoning and tool use, and reinforcement learning (RL) applied to them suffers from an exploration bottleneck. CanvasAnneal warm-starts RL by injecting a stronger teacher model's reasoning traces into the initial diffusion canvas. It then gradually removes this guidance as training progresses, so the model must generate more of the reasoning on its own. It improves over standard diffu-GRPO on MATH500, Countdown, and Tau2 and substantially accelerates reward improvement on several tasks, although the gains vary by task.
6 more specialized papers

Safety & Alignment 14

Certifying Concept Unlearning in Text-to-Image Diffusion Models

Mansi, Luca Marzari, Francesco Leofante Evaluations of concept unlearning in text-to-image (T2I) diffusion models rely on attack success rates from automated adversarial prompt search, which give only empirical evidence over a finite query set and leave leakage across the broader prompt space unquantified. The authors introduce a certification framework that combines statistical certification with worst-case analysis along concept-relevant embedding directions to produce explicit upper bounds on leakage probability at user-specified confidence levels. Across NSFW content, artistic styles, and celebrity identities and six state-of-the-art unlearning methods, certified leakage bounds exceed standard attack success rates by 16.2%, indicating that attack-based evaluations substantially underestimate residual risk.

Do Influence-Derived Data Perturbations Enable Machine Unlearning? A Controlled Study of Three Plausible Roles

Chenkai Wu, Chrispine Kambimbi, Qinyang Zeng, Jun Yan Deep Perturbation Learning (DPL) perturbs training images and labels along influence-derived directions and has been positioned for machine unlearning in three roles: a direct deletion signal, a utility-preserving regularizer, and a warm start for adversarial unlearning. Because evidence for the weaker roles has been used to support the strongest claim, the authors test each role separately under a matched protocol with exact-seed retraining baselines, first auditing the public code and finding that image directions were computed on augmented, normalized tensors but applied to raw images, and that the label perturbation fell below float32 resolution, leaving labels unchanged. After fixing the image pipeline, DPL fails the direct-deletion criterion on CIFAR-10 with ResNet-18 in all three paired seeds, its utility effects flip sign across seeds, and it underperforms simple warm-start baselines once direction-computation time is counted; a one-seed Tiny ImageNet check is similarly unfavorable but inconclusive because of preprocessing inconsistencies in the released code. The study covers random instance deletion only, and the authors release a role-matched evaluation protocol and an audit checklist for perturbation-based deletion claims.

I Am No One: Style-Aware Paraphrasing for Text Anonymization

Ahmed Sohair Khan, Estrid He, Monica Wachowicz, Elham Naghizade Authorship attribution models can re-identify writers from stable stylistic fingerprints even after explicit identifiers are removed, a risk that extends to automatic speech recognition (ASR) transcripts of meetings and call-center conversations, and differential privacy (DP) anonymization tends to badly degrade text quality. The proposed approach prompts pretrained large language models to build a compact stylistic profile of an author from minimal samples and then rewrite text to suppress the identifiable style markers while preserving meaning. Across blog and review datasets, the method reduces authorship attribution F1 by 60 to 70 percent while maintaining content quality and readability, substantially outperforming both DP-based and non-DP baselines.

Membership Inference via Pairwise Likelihood Ratios

Shengjie Niu, Zebin Yun, Yeheng Ge, Jian Huang cross-listed Membership inference attacks (MIAs) test whether a given point was in a model's training set, but existing attacks summarize and combine the available signals from confidence scores, logits, and intermediate features inefficiently. Pairwise Likelihood MIA, or PL-MIA, pairs a Gaussian likelihood-ratio (GLR) statistic with population calibration and the Cauchy combination test: it derives p-values from pairwise comparisons between the query point and reference points known to be outside training, then aggregates those continuous signals rather than collapsing each comparison to a binary vote. Theory characterizes how the GLR retains variance-contraction signals and when calibration and Cauchy combination raise attack power, and experiments show true positive rate gains of over 25% in the low-false-positive regime against strong baselines.

SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration

Md Jueal Mia, Yanzhao Wu, Selcuk Uluagac, M. Hadi Amini Large language models (LLMs) are becoming agentic systems that plan, invoke tools, keep persistent memory, and communicate with other agents, while also having much stronger native safety alignment than the models on which most jailbreak research was done, which raises the question of which established findings still hold. A Systematization of Knowledge (SoK) reframes jailbreak security around the full agentic execution pipeline, building unified attack and defense taxonomies spanning user interaction, planning and reasoning, memory, tool use, and inter-agent communication, introducing a security-utility-efficiency evaluation framework that separates native harmful-prompt safety, adversarial jailbreak robustness, and agent-level outcomes, and running a controlled empirical comparison of representative attacks and defenses in a common agentic framework. The results show that strong native alignment does not imply robustness to adversarial jailbreaks, that defense effectiveness is highly model-, attack-, and component-dependent and often costs over-refusal, utility, and latency, and that low final-response attack success can mask severe compromise of planning, memory, and tool interactions. The authors argue for a shift from response-centric filtering toward cross-layer, execution-aware security that protects agent state, component transitions, and external actions.

GraphProfiler: Source-Linked Sensitive Attribute Inference via Personal Knowledge Graphs

Ahmed Sohair Khan, Estrid He, Chenglong Ma, Monica Wachowicz, Elham Naghizade Large language model profilers can infer sensitive attributes such as age, income, and occupation by aggregating indirect cues across many ordinary posts, but existing profilers give little insight into which posts, concepts, and relationships enabled an inference, which is what targeted mitigation that redacts or rewrites only the leaking posts requires. GraphProfiler builds each user's post history into a source-linked personal knowledge graph whose nodes and edges trace back to the originating post, then resolves attribute predictions to cited graph records and source texts. It reaches 86.7% attack success on the eight-attribute SynthPAI benchmark, within two points of strong text-only baselines, and 84.6% on PANDORA, while citing supporting evidence for over 98% of predictions. Controlled ablations show that removing the cited posts reduces attack success substantially more than removing an equal number of random posts, indicating the citations identify posts that genuinely drive the inference.

Reproducing and Evaluating the Generalizability of Subliminal Learning in Open-Weight Models

Daan van der Weijden, Nathan Brack, Selene Baez Santamaria Subliminal learning is a distillation effect in which a teacher model passes behavioral traits, such as animal preferences or misalignment, to a student through data semantically unrelated to those traits. Because the original work's GPT-4.x fine-tuning is no longer available, the authors reproduce its experiments on open-weight models and extend them with new preference categories (actors and politicians), a chess move generation task, an additional model (Ministral8B), and an ablation on digit length in the number-sequence task. The reproduction supports the original claims, but the extensions show the effect is not universal: transmission strength varies across traits and tasks, and one model shows almost no effect at all.

Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

Kazusato Oko, Annie Ulichney, Nika Haghtalab, Han Bao Prior work showed that the distortion of Reinforcement Learning from Human Feedback (RLHF) can grow exponentially with the Bradley-Terry temperature β when users have heterogeneous preferences. Distortion here means the multiplicative gap between the average user utility of the RLHF policy and the best achievable average utility. Analyzing RLHF with reward clipping, the authors argue that this blowup is not built into the algorithm. Instead, it comes from a mismatch between the distribution that generates the preference data and the KL reference policy. They prove tight upper and lower bounds across regimes of KL regularization strength, and RLHF achieves the optimal O(β) distortion when the two distributions match. In a representative regime, distortion scales roughly as β·B + β, where B bounds the log density ratio between them, which favors on-policy preference data or fine-tuning on data close to the preference distribution before RLHF.

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

Guangsheng Yu, Yanna Jiang, Qin Wang, Baihe Ma, Xu Wang Unlearning benchmarks such as TOFU and MUSE judge forgetting by the model's final answer, where a refusal counts as success, and the authors show this certificate does not carry over when the model is deployed as an agent. K-Bench inspects all six output channels a ReAct agent exposes, including chain-of-thought, tool calls, tool observations, and elicited summaries. Each experiment places the secret in exactly one source (the weights, the prompt, or the retrieval store), and forgetting is credited only if the agent remains usable. When the secret lives in the prompt or retrieval store, TOFU and MUSE report no leakage while the deployed agent still leaks it on 22-86% of queries; when it is in the weights, none of twenty published unlearning methods demonstrably removes it, and the top-ranked method changes across base models.

EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics

Jiaxu Zhao, Bahar Radmehr, Fares Fawzi, Tanya Nazaretsky, Tanja K\"aser LLMs are increasingly deployed as tutors, but it is unclear whether tutoring quality varies with student demographics. EduFair-Bench pairs a mathematics, physics, and chemistry question bank with a controlled simulation in which a fixed LLM student interacts with each tutor under nine demographic levels spanning gender, immigration background, first language, and socioeconomic status. Tutoring is scored by an LLM judge validated against human annotators, and ablations separate tutor-driven from student-driven bias. Across five tutors, model capability and demographic fairness were largely orthogonal: the smallest model was the most consistent, the four more capable tutors all showed wide demographic gaps, and pedagogy-specific reinforcement learning training redistributed bias rather than removing it.
4 more specialized papers

Robotics 10

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence

Shilong Zou, Shilin Zhang, Yingji Zhang, Yuhang Huang, Yi Zhang, Zeyuan Ding et al. cross-listed Embodied agents benefit from world models that predict future observations from visual context and actions, but heterogeneous robot embodiments and slow rollouts limit their use. Pelican-Sim 1.0 uses a unified 28-dimensional action space spanning most mainstream embodiments, injects actions as URDF- and camera-rendered videos to bridge actions and pixels, adds sparse mixture-of-experts layers to absorb the action modality, and distills a 35-step model into a four-step autoregressive simulator with a 5.67-fold speedup. Trained on roughly one million real and simulated trajectories, it improves PSNR over the strongest baselines by 4.636 on AgiBotWorld Beta, 2.080 on RoboMIND, and 10.343 on RoboTwin. On RoboTwin, adding 500 generated trajectories to 50 demonstrations per task raises policy success from 70% to 93%, and policy evaluation through the simulator reaches a Pearson correlation of 0.994 across five checkpoints.

Efficient Vision-Language-Action Management and Serving for Robot Factories

Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula cross-listed Vision-Language-Action (VLA) models run a Vision-Language Model (VLM) stage followed by an Action Diffusion Transformer (ADiT) stage, and their latency-critical inference must be offloaded from robots to edge servers, yet existing serving systems either lack multi-request, multi-model support under Service-Level Objectives (SLOs) or disaggregate stages across GPUs in ways ill-suited to millisecond-scale stages. Robion disaggregates the two stages within a single GPU using two streams, restricts the streaming multiprocessors available to the VLM stream so the ADiT stage always has capacity, shares streams across co-located models, prioritizes by least remaining SLO time, and adds a management engine with a traffic controller that maximizes batching while bounding per-GPU load. For individual models it sustains 6.7 times and 1.5 times higher robot load at 98% SLO attainment than vLLM-Omni and a monolithic pipeline, respectively, and it serves 8 different models to up to 64 robots on a 4-GPU server.

IMPLY: Physically Anchored Consistency for World-Model Rollouts

Aman Mehta, Riya Baviskar cross-listed Consistency checks used to vet world-action models ask whether a model's predicted futures agree with each other, but none of them knows any physics, so a model that ignores the object and always predicts a typical push can look perfectly self-consistent. IMPLY inverts a simulator to read the physical parameters such as mass and friction that each rollout implies, then scores a set of rollouts by how well a single object explains all of them, anchored to two calibration pushes the model has actually observed. In a controlled setting, self-consistency gives a perfect score to an object-blind model while anchoring exposes it (AUROC 0.70 versus 1.00). On V-JEPA 2-AC adapted to the scene, the model tracks the object given its own calibration pushes (correlation 0.91 with truth) but not another object's (0.05); self-consistency prefers the right evidence on only 52% of objects, chance level, while anchored disagreement prefers it on 73%, correlates 0.92 to 0.99 with rollout error, and selects among candidate rollout sets within 0.003 of an oracle that sees the truth.

Agent as Policy for Robotic Manipulation

Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang Robot manipulation policies typically require task- or environment-specific training, whereas general-purpose agents already reason over images and write executable code. Agent as Policy (AGP) hands a task description and a robot interface to such an agent, which interprets visual evidence, writes programs, issues motion commands, and revises its actions in response to physical outcomes, with no task-specific or environment-specific training. Evaluated on real-world tasks including assembly from human videos, block construction from goal images, die reorientation, targeted throwing, and bimanual towel folding, AGP reaches 100%, 100%, and 80% success on three block construction configurations, indicating that a general-purpose agent can serve as a robot policy through runtime reasoning, programming, and interaction.

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim et al. cross-listed Language-conditioned robot policies can benefit from predicting both the target outcome and how actions change the scene. Dynin-Robotics builds on Dynin-Omni, an omnimodal masked-diffusion model that represents language, observations, goals, and actions as discrete tokens. By varying which spans it conditions on and which it predicts, a single trajectory model learns action prediction, next-observation prediction, goal-state prediction, and instruction reconstruction, which enables test-time scaling through goal guidance, action-candidate evaluation, and joint refinement of actions and future states. After continual pretraining on about 1.33 million trajectories from 48 Open X-Embodiment datasets, the model improves shifted-instruction success on VLABench over action-only decoding, is competitive on LIBERO and zero-shot LIBERO-Plus, and reaches a 78.4% average success rate across four manipulation conditions on a real Franka Research 3 robot, while a block-parallel implementation decodes actions up to 29.2x faster.
5 more specialized papers

Reasoning 9

Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding

Minoo Ahmadi, Seyedarmin Azizi, Erfan Baghaei Potraghloo, Mehdi Kamal, Massoud Pedram Sequential Monte Carlo (SMC) power sampling can boost large language model reasoning at inference time without post-training, but standard equal-weight resampling aggressively prunes low-weight trajectories and collapses the genealogical diversity of the search. Chopthin-Consensus Power Sampling (CCPS) swaps in the Chopthin resampler, which bounds the ratio between the largest and smallest particle weights and carries unequal weights forward, preserving distinct reasoning paths while leaving the weighted SMC approximation unchanged in conditional expectation and guaranteeing a lower bound on effective sample size. A semantic-majority selection step then merges token-identical final trajectories, clusters equivalent answers, and returns the answer backed by the most distinct trajectories. Across three open-weight models and five reasoning benchmarks, Chopthin raises oracle coverage in 13 of 15 settings, and CCPS matches or exceeds the Power-SMC baseline's final-answer accuracy in 14 of 15 settings, with gains of up to 10.6 percentage points.

GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs

Zixiang Xu, Yanbo Wang, Chenxi Wang, Lang Gao, Zirui Song, Yue Huang et al. How reliably large language models (LLMs) can execute multi-step graph algorithms in natural language is unclear, since existing evaluations use small graphs, score code generation instead of reasoning over the graph itself, or fix a single input format. Graph Theory Bench (GT Bench) covers 24 classical graph problems in 44 task-structure settings with over 100,000 examples across four representations: natural language, structured language, adjacency list, and adjacency matrix. Evaluating eight LLMs shows accuracy is strongly tied to input representation, with the best representation shifting with graph density, size, topology, and model, a sensitivity that persists in attenuated form even in the strongest reasoning models. The Graph Theory Agent (GTA) pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM, and lifts Phi-4 from 53.5% to 69.1% on the easy split and from 33.0% to 41.5% on the hard split, beating eight prompting and agent baselines and transferring without retraining to GraCoRe and NLGraph.

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

Zhendong Mi, Shaoyi Huang Whether reinforcement learning (RL) teaches language models new reasoning skills or merely sharpens the distribution toward high-reward paths already latent in the base model is an open question, and if the latter holds, those paths might be reachable without any training. Decision-Flow Sampling (DF-Sample) is a training-free, data-free inference procedure that builds a hierarchical reasoning tree, scores terminal nodes for quality, and back-propagates those utilities to guide each intermediate branching decision, so whole trajectories are evaluated before a path is committed rather than making purely local step-wise choices. On GPQA, DF-Sample reaches 45.6 percent accuracy versus 38.9 percent for power sampling and 39.9 percent for GRPO, and it beats baselines across three models and four benchmarks, which the authors read as evidence of substantial untapped reasoning capacity in pretrained base models.

From Collaboration to Capability: Internalizing Routed LLM Experts into Compact Reasoners

Frank Nie, Shuyao Wang, Ethan B. Liu A small controller model can solve problems by choosing which stronger expert LLMs to consult, what to ask them, and how to combine their answers, but that setup still depends on the experts at deployment time. RIVET internalizes such collaboration in two stages: expert-augmented reinforcement learning applies a shared outcome reward to both the controller's decisions and the spans the experts return, and a second stage fine-tunes on complete, verified successful interactions with format-aware supervision. The deployed model produces reasoning, code, and interaction structure on its own, using only local Python execution and no external LLM. RIVET-1.7B and RIVET-4B average 28.25% and 44.16% across seven competition-mathematics benchmarks, with the second stage adding 6.49 points to RIVET-4B's accuracy after the experts are removed, and GPQA-Diamond results suggest transfer to scientific reasoning.

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen Existing multi-hop reasoning benchmarks are mostly short-horizon and say little about whether LLMs can reliably follow real manuals spanning hundreds of pages of interdependent rules. The Tasks over Application Manuals (TAM) benchmark collects human-validated real-world tasks from ICD-10-CM clinical coding and U.S. federal sentencing guideline calculations. Each task requires chaining steps across sections of a manual with tens of thousands of rules to reach an exact answer. Testing retrieval-augmented generation, ReAct-style prompting, and an agent-harness baseline on GPT-5, the authors find best exact-match accuracy of only 1% on clinical coding and 15.5% on sentencing.

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin et al. Low scores on leading physics benchmarks suggest frontier language models still struggle with advanced physics, which conflicts with domain experts' experience using them. The authors evaluated frontier models on six widely used physics benchmarks and had faculty and graduate researchers audit problem statements, reference solutions, and model responses. Most answers initially graded incorrect turned out to reflect grader errors, wrong reference solutions, or ambiguous questions rather than faulty physics reasoning. After experts fixed or excluded flawed items, GPT-5.6-Sol's mean@4 score rose from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, and its pass@4 reached 94.4% on 54 retained CritPt challenges, suggesting these closed-ended benchmarks are near saturation.
3 more specialized papers

Vision 8

Is Gaussian Splatting Becoming Neural Again? A Taxonomy and Controlled Study of Learned Parameterization

YuanHang Wang, Xin Cao, Yi Zhang Three-dimensional Gaussian Splatting (3DGS) pairs explicit primitives with fast rasterization, yet a growing number of systems use neural networks to generate or share Gaussian parameters, raising the question of what is actually being gained. The authors organize this trend along five axes, namely attribute decoding, spatial sharing, view-conditioned decoding, topology generation, and amortized inference, and show through an analysis of 19 methods that these choices target different limitations and cannot be collapsed into a single neural-or-not label. A controlled study on mip-NeRF 360 isolating three forms of neural parameterization finds that sharing appearance and opacity improves reconstruction quality while neurally decoding geometric structure offers no further gain, supporting selective neuralization that captures reusable correlations without giving up the local geometric freedom of explicit splats.

UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction

Xinqiang Yu, Zekun qi, Jiawei He, Wenyao Zhang, Xuchuan Chen, Guaocai Yao et al. cross-listed Fine-grained robotic manipulation requires understanding object parts, yet 3D foundation models tend to be either general but object-level or part-aware but restricted to closed-set taxonomies, which hurts zero-shot transfer. UniPart is a feed-forward cross-modal 3D Transformer that conditions on a CLIP text embedding of a free-form phrase to segment the matching functional part of a point cloud. To scale supervision the authors build LangPart-1M, with over 160K Objaverse assets and 8M text-to-part pairs created through multi-view consistent part generation, plus a manually labeled high-quality subset, LangPart-4K, for fine-tuning and evaluation. The model achieves strong zero-shot results on open-vocabulary part benchmarks and transfers to language-conditioned part grasping in the real world.
6 more specialized papers