Monday, September 21, 2026

298 papers cs.AI · cs.LG · cs.CL ← 2026-09-182026-09-22 →

Jul Aug Sep

Highlights

Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

Highlight HF pick · 1▲Vision Meng'en Qin, Yinchen Liu, Mingxuan Cui, Youlu Xing Convolutional sparse coding (CSC) suppresses redundant image components while keeping signal content, but its sparsity coefficient is normally fixed and tuned by hand. This framework unfolds the CSC optimization using the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) and makes the sparsity coefficient a differentiable variable learned jointly with the network weights, interpreted through the information bottleneck as the trade-off between compression and information retention. A label-free post-training step then adjusts the compression strength for corrupted inputs while the main network stays frozen. On CIFAR and ImageNet the method reports competitive clean accuracy and greatly improved robustness to input perturbations.

Convolutional sparse coding (CSC) layers give networks an explicit knob for discarding redundant signal content, but the sparsity coefficient λ is normally a hand-picked constant. TA-CSC treats λ as a per-layer variable learned through unrolled FISTA iterations, framed as the information bottleneck's compression-versus-sufficiency trade-off, and then re-tunes only λ on unlabeled corrupted inputs after training.

  • Each CSC layer unrolls 2 FISTA iterations and backpropagates through the soft-threshold to update λ (kept positive via Softplus) jointly with the dictionary, under a loss that adds a normalized ℓ1 penalty on the sparse codes (γ = 0.001) to the task loss.
  • For corrupted data, all network weights are frozen and only λ is updated on 100 unlabeled samples by minimizing a relative reconstruction error divided by λ, which pushes compression up until reconstruction fidelity degrades; the adapted λ rises monotonically with corruption severity.
  • Replacing every convolution in ResNet-18 with these layers reaches 97.65% on CIFAR-10, 80.76% on CIFAR-100 and 72.53% on ImageNet-1K, versus 95.54%, 77.82% and 68.98% for the baseline, while swapping only the first layer still gives 96.18%, 79.63% and 71.12%.
  • On CIFAR-10-C Gaussian noise, accuracy goes from 44.43% (ResNet-18) to 53.98% without adaptation and 68.23% with it, beating SDNet-18 with per-sample λ tuning (64.92%); on ImageNet-C Gaussian noise it goes from 22.73% to 30.93%.
  • The all-layer variant is expensive (88.6 GB and 153 samples/s on ImageNet versus 24.1 GB and 2100 for ResNet-18), evaluation covers only ResNet-18 classification under noise-type corruptions, and the authors concede the link between λ and the information bottleneck is empirical rather than proven.

Recursive Language Models Generalize Out of Domain

Highlight Theory Chenxiao Yang, Zhiyuan Li, David McAllester, Nathan Srebro The question studied is when restricting what a language model can see improves learning, comparing standard chain-of-thought (CoT), which reads the full reasoning trace, against recursive language models that solve each subtask in an isolated context. In-distribution, CoT can efficiently simulate the recursive rule, so its generalization guarantee differs only by a constant factor and recursion offers little. Out of domain, however, CoT can fit the training data through shortcuts that depend on context outside the current subtask and break when those tokens change, a failure mode that context isolation rules out. Because simplicity bias selects the shortcut even though CoT's hypothesis class contains the correct rule, the authors argue that covering the right rule is not enough for out-of-domain reasoning, in contrast to classical learning theory.

Chain-of-thought models read their entire reasoning trace, which makes them a strictly more general learner than recursive language models that solve each subtask in an isolated context — yet this work argues that the extra generality is exactly what hurts out of domain, because it lets the learner fit training data with shortcuts that depend on context outside the current subtask.

  • The comparison pits standard CoT, which conditions on the full trace, against recursive language models (RLM), which deliberately restrict each subtask to its own isolated context and therefore cannot see tokens belonging to other subtasks.
  • In-distribution, recursion buys little: CoT can efficiently simulate the recursive rule, so its IID generalization guarantee differs from the recursive learner's by only a constant factor.
  • Out of domain, CoT can reach low training error by leaning on context outside the current subtask, a shortcut that breaks as soon as those surrounding tokens change, whereas recursive context isolation rules out this failure mode by construction.
  • The key theoretical point is that CoT's hypothesis class still contains the correct recursive rule, but simplicity bias selects the shortcut over the truth — so covering the right rule is not sufficient for out-of-domain reasoning, in contrast with classical learning theory.
  • The abstract reports no concrete benchmarks, model scales, or accuracy numbers, so the practical size of the gap and how well the analysis transfers to real-world tasks without a clean subtask decomposition remain unclear from the summary alone.

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Highlight HF pick · 4▲Large Language Models Sy-Tuyen Ho, Minghui Liu, Furong Huang As model-written peer reviews enter training corpora, later AI reviewers may learn from earlier ones, and this study simulates one step of that loop by fine-tuning Llama 3.1 8B on official ICLR reviews from 2018-2023 and then training four successors on ICLR 2024 data with varying mixes of official and model-generated reviews. Adding synthetic reviews compresses rating distributions and reduces semantic diversity both within a paper's reviews and across the corpus, a pattern the authors name scientific-judgment collapse. To counter it they introduce TrustReviewer, an open-source reviewing system that trains in a single stage on a corpus curated to remove low-quality and degenerate supervision, and applies paired activation steering at test time to correct residual collapse without further training or expert annotation.

As LLM-written reviews leak into public corpora, future AI reviewers may end up trained on their predecessors' judgments. The authors simulate one step of that loop and find that synthetic supervision narrows both ratings and review content, which they call scientific-judgment collapse. They then propose TrustReviewer, which pairs a curated training corpus with test-time activation steering to counter it.

  • A Llama 3.1 8B reviewer fine-tuned on official ICLR 2018–2023 reviews generates reviews for ICLR 2024 papers, and four successors are then trained with 0, 1, 2, or 3 of each paper's three reviews replaced by synthetic ones (0/33/66/100%) under identical LoRA settings, with all models evaluated on 2,000 held-out papers.
  • Introducing just 33% synthetic reviews cuts rating standard deviation from 1.63 to 1.44 and entropy from 2.31 to 2.14 (official reviews: 1.73 and 2.38), while mean ratings move non-monotonically (5.30, then 5.85, then 5.70 at 100%), so the effect is narrowing rather than a drift toward leniency or harshness.
  • Semantic diversity shrinks too: same-paper pairwise embedding distance falls monotonically from 0.159 to 0.142 (about 11%), and corpus-level spread falls from 0.609 to 0.579 (about 5%), with most of that contraction arriving at the first synthetic step.
  • TrustReviewer is trained in a single stage on 112,743 curated paper–review pairs (~1.9B tokens) from ICLR 2018–2025 and adds a steering vector (the mean last-token hidden-state difference between official and model-generated reviews of the same 5,000 papers, applied at the final layer with α=0.15), reaching 75.40% exact match and 1.079 MAD versus 73.10% for OpenReviewer, 61.85% for Qwen3.6-35B-A3B, and 33.93% for the base Llama, with steering itself contributing +1.55 points and lifting rating entropy from 2.13 to 2.18 while slightly lowering same-paper diversity.
  • The collapse study covers only one recursive step, one 8B model family, and one venue, and TrustReviewer is benchmarked against external baselines rather than tested inside the contaminated-data loop; the authors also caution that diversity is not quality, that agreement with official ratings is not correctness, and that embedding distances are only a proxy, with no human evaluation of critique validity.

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

Highlight Agents George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui et al. Agentic coding benchmarks judge patches with held-out test suites that are inherently incomplete and increasingly vulnerable to memorization, while formal verification has so far only covered standalone tasks whose specifications are handed to the model. Benchproofer turns a coding task with a known-correct patch into a formally verified one by writing a specification for the new code, summarizing the existing functions it calls with axioms, and admitting an instance only when mechanical and adversarial gates agree; applied to SWE-bench Verified it produces SWE-Proof, 500 real issues checked by proof rather than test, and extends to SWE-bench Pro. A quarter to a half of test-passing patches from two frontier models admit counterexamples, and supplying a correct formal specification lifts Opus 4.8 from 85% to 95% resolution. Writing the specification is the hard part: models asked to produce their own gain nothing over an unaided baseline, only 62% of those specifications pass audit, and the usual failure is faithfulness — constraining part of the required behavior and leaving the rest free.

Coding benchmarks judge patches with finite hidden test suites, which are incomplete and open to memorization, while formal verification has so far only covered small standalone tasks with specifications supplied up front. Benchproofer converts real repository issues with a known correct patch into formally verified tasks, yielding SWE-Proof: all 500 SWE-bench Verified instances with ground-truth specifications, implementations, and proofs under Nagini, Velvet, and Lean.

  • The pipeline formalizes only the new or modified code, summarizes unchanged callees with fuzz-tested axioms, and admits an instance only after mechanical gates (verification, rejection of the pre-fix code, mutation kill rate, hygiene) and adversarial LLM audits (axiom soundness, conformance on at least 10^5 inputs, soundness/completeness, equivalence, leakage) all agree; it also built gate-passing artifacts for 242 of 266 Python tasks in SWE-bench Pro.
  • Hidden tests prove substantially incomplete: an adversarial audit drops the unaided baseline from 85.0% to 58.2% for Claude Opus 4.8 and from 81.2% to 33.4% for GPT-5.5, and a structured natural-language EARS specification does not close the gap, whereas formal verification against the ground-truth specification costs under one point.
  • Supplying the ground-truth formal specification lifts resolution to roughly 95% for both models, about +7 to +8 points beyond localization alone, while agents asked to write and verify their own specification gain nothing over the baseline (best cell +0.6, worst -1.2).
  • Specification synthesis is the bottleneck: only 46–72% of model-written specifications pass a five-property audit, with faithfulness (modeling too little of the behavior the issue requires) failing on 42.5% for Opus and 30.0% for GPT, and specifications from unresolved instances failing the audit 89.4% of the time versus 47.3% for resolved ones.
  • The verify-implies-resolve guarantee is not itself a proof, since it rests on audited trust points (axiom soundness and the agent-written Python translations used for conformance under Velvet and Lean), the auditors and judges are LLMs whose verdicts are evidence rather than proof, and patches that only rename, relocate, or change side effects fall outside what pre- and post-condition specifications can observe.

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

Highlight HF pick · 2▲Reasoning Qiangqiang He, Jin Li, MingCai Chen On-policy distillation (OPD) trains a reasoning model on the token-level discrepancy between a stronger teacher and the student's own samples, but that discrepancy also contains deviations arising from the teacher itself, a problem amplified when the teacher is given privileged information. Calibrated On-Policy Distillation (Cal-OPD) estimates the teacher's self-deviation region using positive and negative privileged interventions and keeps only the part of the discrepancy lying beyond that region. On mathematical reasoning benchmarks, it consistently outperforms standard OPD and its variants across model scales while retaining only about 52-65% of the original discrepancy as the optimization signal.

On-policy distillation (OPD) trains a student on the token-level log-likelihood gap to a teacher, but that gap also carries "teacher self-deviation" (TSD) — shifts in the teacher's own likelihoods under changed context that reflect no task knowledge — and privileged-information OPD amplifies it. Cal-OPD probes the teacher with a positive and a negative context intervention to estimate a per-token TSD region, then distills only the part of the discrepancy that falls outside it.

  • Rescoring ~60M tokens of fixed Qwen3-1.7B rollouts with a Qwen3-8B teacher shows TSD needs no task knowledge and mostly ignores correctness: task-agnostic instructions shift 29.8% of tokens versus 29.4% for answer-level hints, and swapping a correct answer for a wrong one still gives 88.7% directional agreement in the shifts.
  • TSD concentrates on surface-form tokens rather than mathematical content, with discourse markers such as maybe, however, and therefore shifting in over 89% of occurrences while digits, symbols, and LaTeX notation stay below 9.3%.
  • The method takes the largest upward and downward teacher log-likelihood shifts under the two interventions (evaluative feedback praising or condemning the rollout), widens that interval by a relaxation factor λ=5, and sets the advantage to zero when the student's log-likelihood lies inside it, otherwise keeping only the residual beyond the nearest boundary.
  • On six math benchmarks (Avg@16) Cal-OPD averages 53.1 for Qwen3-4B-Thinking-2507→Qwen3-1.7B and 69.0 for Qwen3-30B-A3B-Thinking-2507→Qwen3-4B, which is +2.3 and +3.1 over standard OPD, while standard OPD drops the 4B student below its starting 66.6 and Privileged-OPD is the worst distillation method in both settings.
  • It retains only 52–65% of the original discrepancy, zeroes out 27–34% of tokens, and ends training at about 9.3K-token responses versus 11.8K for OPD, making it roughly 1.26× faster despite the extra teacher passes.
  • Results are sensitive to the probe and to λ: calibrating with solution-level interventions over-filters to about 20% retained signal and falls to 49.0, large λ ends up below the λ=1 setting, and evidence is limited to the Qwen3 family, math benchmarks, and 100 training steps.

GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

Highlight HF pick · 6▲Agents Rui Sun, Zhi Zheng, Zhenkun Wang, Zhichao Lu Skills give large language model (LLM) agents task-specific procedural guidance, but writing them as unstructured natural language leaves them redundant, short on workflow-level direction, and hard to optimize over a huge search space. The authors represent a skill as a graph whose nodes are execution steps with operational guidance and whose directed edges encode context-dependent transitions. GraphSkillEvo then runs population-based evolutionary optimization over these graphs with mutation and crossover operators. Across five agent benchmarks it beats the SkillOpt baseline, improving average accuracy by 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4.

Skills that guide LLM agents are usually optimized as unstructured natural-language text, which agents find hard to follow and which gives the optimizer a large, redundant search space. GraphSkillEvo instead represents a skill as a graph, where nodes are execution steps with their own guidance and edges are condition-dependent transitions, and evolves a population of such skills with structure-aware mutation and crossover.

  • A skill consists of global guidance, reusable step nodes, and workflows that pair an applicability condition with an ordered node path. A population of 4 skills evolves for 5 generations using four round-robin operators (mutation and crossover on either the guidance or the graph), where mutation reflects on up to 5 failed trajectories from 15 sampled training instances, a validator script rejects malformed graphs, and the top 4 by validation score survive.
  • Across five benchmarks (SearchQA, SpreadsheetBench, DocVQA, LiveMathematicianBench, ALFWorld) it is best in 13 of 14 settings, beating SkillOpt on average by +4.01 points on GPT-5.4-nano, +1.76 on GPT-5.4, and +1.33 under the Codex harness, with the largest gain on SpreadsheetBench with the nano model (60.71 vs 50.11).
  • Total optimization cost is lower, since SkillOpt uses 1.31x the tokens of GraphSkillEvo on GPT-5.4 and 1.36x on nano, although GraphSkillEvo spends more on ALFWorld (47.75M vs 28.18M tokens on nano).
  • In ablations on nano, the three-benchmark average falls from 71.52 to 64.08 without the graph structure, 66.59 without crossover, and 54.50 without mutation, while skills optimized on nano and transferred to GPT-5.4 reach 71.78 on SpreadsheetBench versus 69.40 for direct optimization and 53.21 for transferred SkillOpt.
  • The gains are modest on the stronger model and it trails SkillOpt by 0.80 points on LiveMathematicianBench with nano. In the five-run significance test only SpreadsheetBench (p=0.016) and ALFWorld (p=0.038) clear 0.05, run-to-run variance is often higher than SkillOpt's (SearchQA std 3.59 vs 1.25), and everything is evaluated on a single model family with validation sets as small as 18 instances.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

Highlight Safety & Alignment Hiskias Dingeto A language model may hold knowledge it does not report, for example by sandbagging on a capability evaluation, and its outputs alone cannot show whether it is hiding an answer or lacks one. Borrowing the forensic Concealed Information Test, Probe of Internal Recognition (PIR) presents a question with candidate answers and reads from the model's internal states which candidate it recognizes as correct, without needing an honest reference model or a labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, against a 0.28 to 0.40 unknown-item baseline and 0.25 chance, and recognition stays between 0.85 and 0.93 under prompted deception, trained sandbagging, password-locked checkpoints, and circuit-broken checkpoints. When unlearning actually removes the knowledge, recognition falls to the level of a never-known question, so the probe separates a model that will not answer from one that cannot; the signal is reported to be causal, to add information beyond black-box cues, and to extend to free-form generation.

A language model's wrong answer is ambiguous — it may be sandbagging or may genuinely not know — and outputs alone cannot tell the two apart. PIR (Probe of Internal Recognition) adapts the forensic Concealed Information Test to activations: it presents a question with its candidate answers and reads which one the model internally recognizes as correct, with no honest reference model or labeled truth corpus.

  • The method averages, over questions the model answers correctly, the difference between the correct option's last-token hidden state and the mean distractor state, then scores each option of a new question by projection onto that direction, with a label-free divergence detector comparing the recognized option against the one the model actually output.
  • Under prompted deception, recognition of the concealed answer reaches 0.70–0.87 balanced accuracy across eight models from five families (Gemma, Qwen, Llama, Mistral, Phi) against an unknown-item baseline of 0.28–0.40, and it holds at 0.85 on a trained sandbagging policy and 0.93 / 0.90 on external password-locked and circuit-broken gemma-2-9b checkpoints.
  • Unlearning instead erases the signal: RMU on gemma-2-9b drops recognition from 0.85 to 0.39, overlapping the 0.35 baseline, and external Qwen2.5-7B checkpoints fall from 0.85 to 0.36 (RMU) and 0.38 (NPO), so the same readout separates a model that will not answer from one that cannot.
  • The direction is causal and adds to behavioral cues: steering along it makes gemma-2-9b pick the steered wrong option 18.9% of the time versus 0.5% for random directions and 0% for an output-direction control, while combining it with black-box features lifts detector AUC to 0.93 and 0.90 on the external organisms, and a free-form variant on TriviaQA reaches AUC 0.776.
  • An organism adversarially trained to zero the recognition margin drives the readout to baseline while keeping 0.71 capability (the authors retract an earlier in-sample claim to the contrary), and the method also needs enumerable candidate answers, requires the model's base checkpoint for deployment detection, cannot tell erased knowledge from knowledge never held, and rests partly on single-seed organisms and one clean RMU run.

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Highlight HF pick · 46▲Agents Shuai Bai, Jiayong Deng, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang et al. Computer-use agents have developed along two separate lines, graphical interface control and coding through the command line, while real digital work interleaves both. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web where an agent is given a running reference application, must discover its behavior, and must build a faithful reimplementation, with the reference serving as an oracle for hidden behavioral tests that yield execution-grounded rewards. Models trained on the generated trajectories improve on five out-of-distribution coding and hybrid computer-use benchmarks and verify their rendered outputs more often. On the held-out 250-task RecreationBench, GPT-6 Astra leads at 58.1% overall but passes all programmatic tests on only 2.8% of tasks, with agents reproducing static interface structure more reliably than interactions and computed outputs.

GUI agents can operate software but cannot build it, while terminal agents can build software but cannot see the interface they produce. RecreationWorld trains and evaluates hybrid computer-use agents on recreation: given only a running reference application, the agent must explore it, implement a faithful clone, and visually verify its own build. The reference serves as the oracle for hidden behavioral tests.

  • The framework provides isolated, reproducible environments on Ubuntu, macOS, Windows, Android, and Web, with native GUI control plus coding tools and a 20-hour rollout budget. Candidates are scored by frozen programmatic assertions (via AT-SPI, AXUIElement, UI Automation, UiAutomator, or DOM) and by visual assertions judged by Qwen3.7-Plus, all validated on the reference and by human reviewers.
  • For training, Qwen3.8-Max rollouts on open-source apps were rejection-sampled into a balanced 35,000-trajectory SFT mixture (7,000 per platform). Fine-tuning Qwen3.7-Plus and Qwen-Flash-CPT on this mixture left both above their first checkpoint on all five out-of-distribution benchmarks (ProgramBench, GameCraft-Bench, Vision2Web, OSWorld 2.0, WeaveBench), with gains of up to 17.9 points and more frequent checking of their own rendered output.
  • On the held-out RecreationBench (250 tasks, 50 per platform, median 282.5 tool calls per trajectory), GPT-6 Astra leads ten frontier models at 58.1% overall, ahead of Claude Opus 5 at 44.2% and GPT-5.6 Sol at 42.1%. It passes every programmatic test on only 2.8% of tasks, and no other model exceeds 0.8%.
  • Agents reproduce static interface structure more reliably than interactions and computed outputs. 89.4% of recreations are smaller than their reference (median 16.9% of reference LOC) and more monolithic, and fewer than half of trajectories relaunch and inspect the app after their final edit (23.6–47.5% across models).
  • The trajectory and behavior analyses are descriptive rather than causal, and the transfer curves use one trial per task and are not monotonic. Web tasks cannot be source-blind and Android lacks a packet-level egress filter, while agents repeatedly attempted network and protected-path access, so the benchmark's integrity rests on enforced isolation rather than prompt compliance.

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Highlight Large Language Models Richard Zhe Wang Gating the value pathway of attention is reported to improve language model pretraining, but prior studies disagree about why. The authors argue that such gates supply two things softmax attention lacks: abstention, which lets a head output nothing despite attention weights having to sum to one, and noise filtering, which suppresses interference from superposed features in the residual stream. In matched models from 10M to 350M parameters, abstention is supplied through a learned per-head sink logit and filtering through a gate on each value, and the benefit of abstention shrinks with scale while the benefit of noise filtering grows, with abstention accounting for nearly all of the gain at 10M and filtering for most of it at 350M. The best model at every scale has both primitives, which add negligible parameters and remain compatible with the key-value cache.

Softmax attention forces every head to spend its full attention mass somewhere and aggregates value reads linearly, so a head can neither output nothing nor suppress interference from superposed features. The author argues that value gating helps because it partially supplies two separate primitives, abstention and noise filtering. Each is isolated with its own mechanism, a learned per-head sink logit for abstention and a per-value gate for filtering, in matched models from 10M to 350M parameters.

  • Abstention is a phantom key with a learnable logit and a value fixed at zero, while filtering is either a norm gate (a threshold on read norm) or a projection gate (a sigmoid of a learned linear projection of the value). A lemma shows that a value gate equals a routing gate plus a partial zero option, which motivates a routing-slot renorm control.
  • On FineWeb-Edu with paired seeds, the sink logit's gain over baseline shrinks from 0.0185 nats at 10M to 0.0013 at 350M, while the filtering gain of combo2 over the sink logit grows from 0.0036 to 0.0114, so the two curves cross between 50M and 124M.
  • The abstention benefit shrinks because larger baselines imitate it with heavier attention sinks (11% of mass on position 0 at 350M versus 6.0% at 124M), even as the sink-logit model's phantom absorbs 13% to 53% of attention mass.
  • The two benefits are largely additive, and combo2 (sink logit plus projection gate) has the lowest loss at every tier (3.0210 vs 3.0338 at 350M), costs under 0.01% extra parameters, and stays compatible with the key-value cache.
  • Injecting transplanted value reads shows that each filter is blind along its own decision variable: the norm gate loses 4.4 nats vs the baseline's 2.6 at dose 0.8, and the projection gate collapses only under junk aligned with its learned direction, whereas combo2 loses 1.1 vs 5.9 nats at the highest dose at 350M.
  • The 350M results are single-seed, all training uses one corpus at modest token budgets, the effects are a few thousandths of a nat with no transfer tested beyond 350M, and no throughput was measured.

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Highlight HF pick · 38▲Agents Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv et al. Reinforcement learning for coding agents needs diverse tasks with reliable verifiers, and existing pipelines mine them from development artifacts such as issues and commits, which limits what can be extracted. CodeMidas is an agentic pipeline that uses source code as its only task-specific input: agents explore implemented functionality to write behavioral specifications, construct tests grounded in executing the original code, and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on it with GRPO improves all five benchmarks tested, including DeepSWE +11.7%, ProgramBench +17%, and Terminal-Bench v2.1 +8.5%, and the trained agent explores codebases more and self-verifies in more varied ways.

RL for coding agents needs many diverse tasks with trustworthy verifiers, but existing environment pipelines depend on issues, pull requests, commits, existing tests, or documentation, which limits what can be extracted from a repository. CodeMidas is an agentic pipeline that builds an executable RL environment from source code alone: it removes an already-implemented feature, has the solver rebuild it from a behavioral specification, and grades the result against hidden tests grounded in running the original code.

  • An agent selects functionality with public entry points (CLI tools, pure functions, stateful APIs), deletes the core implementation to form a coherent starting codebase while keeping the original as the reference solution, then builds tests from reference execution and reviews every assertion to remove restrictions the statement does not justify, such as exact message wording or incidental ordering.
  • Candidates must show a clean fail-to-pass transition across six fresh containers (two starting-state runs fail, four reference runs pass), then survive adversarial leakage-hunting rollouts, an agent review of four solution attempts for verifier false positives and negatives, and a filter that keeps only tasks with mixed outcomes under a frontier model, which leaves 5,545 tasks from 3,185 codebases across 23 languages and 15 domains, with a median reference patch of 142 lines.
  • Training MiMo-V2.5 with GRPO and binary execution rewards improves all five external benchmarks, including DeepSWE from 10.0% to 21.7%, ProgramBench Almost Solved from 4.5 to 21.5, and Terminal-Bench v2.1 from 63.7% to 72.2%, while the held-out CodeMidas Val pass rate rises from 35.0% to 44.7%.
  • Scaling the filtered pool from 1k to 3k to 5,545 tasks lifts DeepSWE from 17.57 to 19.05 to 21.70, and even the 3k filtered subset beats an unfiltered 8k sample on SWE-bench Pro, DeepSWE, and CodeMidas Val; behavior also shifts, with pre-edit read/search calls rising from 27.2 to 40.1 and agent-written checks associated with a 4.2-point higher pass rate (95% CI 1.8–6.6).
  • The evidence comes from a single base model with no comparison against training on other environment datasets, the cleaning and filtering stages are ablated only jointly, the full pool's edge over the unfiltered 8k sample on SWE-bench Pro is just 0.59 points, and the behavioral links are correlational, with confidence intervals for exploration and drafting spanning zero.

Applications 62

From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators

Won Seok Jang, Zonghai Yao, Hong Yu Existing evaluations of large language models in medicine target static outputs rather than whether a patient actually understands a discharge plan after an interactive conversation. DischargeBench simulates multi-turn sessions in which a candidate LLM educator teaches a persona-driven Virtual Patient, while an Education Monitor Agent keeps the patient realistic without altering the educator. The accompanying MIMIC-IV-Ext-DischargeBench dataset has 477 cases across 24 ICD chapters with persona axes for personality, education level, health literacy and medical-history recall, scored on four axes by an LLM judge aligned to physician annotations. Across closed- and open-source models, aggregate scores hide clinically relevant variation across conditions and personas, with difficult personas exposing coverage failures, comprehension gaps and lower factual consistency.

Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR

Fiza Husain, Ankit Pandey, Yash Singh Automatic speech recognition (ASR) systems tuned for Word Error Rate (WER) often miss named entities and filled pauses in accented conversational English, both of which matter for language-learning feedback. The pipeline targets speakers from India, Indonesia and Latin America in three stages: heuristic SQL filters that curate training data with 2.8x the entity density of random sampling, regional LoRA adapters on Qwen2.5-Omni-3B that emit verbatim and corrected transcripts in one forward pass, and a six-category error taxonomy checked by an LLM judge. It reaches 80-85% entity recall (up from 53-55%) and 76-86% filler recall (up from under 5%) at 6-10% WER, beating Whisper and a commercial ASR on entity recall and matching a zero-shot 30B model with ten times fewer parameters. Paired bootstrap tests attribute 2.8-4.2 percentage points of the entity recall gain to data curation alone.

Towards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models

Mostafa Anouar Ghorab, Mohamed Aymen Saied Kubernetes misconfigurations are a common source of security and performance problems in cloud-native deployments, and this study examines how well Large Language Models (LLMs) can detect them. The authors build a taxonomy of common Kubernetes misconfiguration types, empirically benchmark existing state-of-the-art detection tools, and analyze which Kubernetes objects are most prone to misconfiguration and how severe the resulting issues are. The abstract reports no specific detection accuracy figures, presenting the work as insight into how LLM-based detection can complement existing methods.

MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery

Peiyi Zheng, Yanming Kang, Hans De Sterck, Giang Tran Symbolic regression recovers closed-form equations from data, but search-based methods rely on costly combinatorial optimization from random starts, while pretrained neural models produce formulas quickly yet often with symbolic errors. MOSAIC-SR uses a pretrained Transformer to propose multiple initial expression sketches that seed searches in several promising regions, with each search jointly recovering structure and constants through scale-aware constant optimization and local symbolic repair. Evaluated on SRSD-Feynman with and without dummy variables plus six additional benchmarks, it obtains the highest symbolic solution rate on every dataset while ranking in the top two for predictive accuracy, and the advantage holds when irrelevant dummy inputs are present.

How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?

Gaurav Agarwal, Ashish Garg, Isha Singhal cross-listed Language models can write GPU kernels that beat PyTorch, but how much of a real workload those kernels affect has not been measured. On KernelBench level 1, a frontier model produces correct kernels for 91.1% of problems with verified speedups on 22 of 56, while the best open-weights model reaches only 30.4% correct. Profiling seven workloads shows the addressable fraction of wall-clock time ranges from 8.9% to 58.2%: on transformers, 80-86% of runtime sits in cuBLAS matrix multiplies and FlashAttention, bounding realistic end-to-end improvement at roughly 1%, whereas recommenders are far more addressable, and the new DLRM-Bench projects an 8.63% end-to-end gain there. The authors also find that KernelBench's torch.allclose correctness check accepts an all-zeros tensor on 4 of 60 problems, which two of their own kernels exploited, including one scored at 283x, and they propose scale-invariant replacements.

Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars

Daniel Henel, Frederik Werner, Alexander Langmann, Johannes Betz Cloud-hosted language models bring network dependency and variable latency, which makes them a poor fit for voice control in time-critical autonomous driving. Jarvis is an offline, open-source voice assistant for issuing high-level behavioral commands to autonomous racecars, combining speech recognition, speech synthesis, and a text-to-command classifier built by domain-specific fine-tuning of Mistral 7B. It reaches 97.63% intent recognition accuracy with an average processing latency of 1.39 s, outperforming larger online-hosted models in the authors' evaluation.

From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost

Saki Imai, Mert \.Inan, Malihe Alikhani Common measures of AI productivity capture what output was produced but not the interaction effort needed to get there. Drawing on economics, the authors propose evaluating human-AI collaboration as outcome quality relative to interaction cost and apply it to two datasets spanning four tasks. They find that sessions with identical quality ratings can differ by up to 70 times in interaction cost, that quality-cost relationships vary by task, and that subjective user ratings are not reliable proxies for productivity. Productive sessions are marked by agents probing earlier and users spending less effort repairing the interaction.

Scaling Forced Alignment to End-User Devices

Lawry Sorenson, Michael Crandall, Eric K. Ringger, Stephen D. Richardson Forced alignment of audio to text with the Viterbi algorithm costs quadratic time and memory in many implementations, which rules out long recordings on ordinary hardware. Two optimizations are proposed: the Hirschberg algorithm to align in place using linear memory, and modeling speech-to-text alignment as a constrained random walk so the search space can be pruned with a tunable confidence bound that still tolerates transcription errors. Memory for a three-hour input drops from 140 GB to 5 MB while producing identical alignments in a third of torchaudio's CPU time, and pruning adds a further 2x speedup on inputs longer than 20 minutes with accuracy preserved in over 98% of tested cases.

Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

King Shi, Amanda Li, Jonathan Ivey, Synthia Qia Wang, Guan Gui, Hyunseo Kim et al. Health systems weighing AI-assisted psychiatric intake need a repeatable way to check these tools against clinical standards without consuming much clinician time or assuming one interviewing style. InterviewPlayground supplies a memory-augmented simulated patient built from expert-authored vignettes, together with a simulated intake platform and evaluation modalities tailored to the task. In a pilot of six clinicians running 25-minute assessments against a GPT-based interviewer, the language model recovered 88.0% of the clinically relevant vignette items versus 38.9% for clinicians, but made unsupported clinical inferences more often (56.8% versus 27.8%) and characterized identified safety concerns less often (33.3% versus 66.7%).

SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity

Thao Nguyen, Heng Ji Off-target protein binding causes many small-molecule side effects, yet structure-based design mostly generates selective compounds from scratch instead of improving drugs already well characterized. The proposed specificity optimization task asks for constrained edits to an existing compound that widen its preference for the intended target over measured off-targets, scored on a new ChEMBL-derived benchmark; an agent docks each compound against all targets, compares poses through residue-aware atom-protein contacts, and hands those differential interactions to a language model that proposes modifications, keeping only candidates passing similarity, ADMET, and docking-selectivity filters. Across 915 compounds the agent widened the target/off-target binding gap for 84.8%, shifting the mean from -0.72 to +0.47 kcal/mol at mean Tanimoto similarity 0.72, and ablations show that replacing residue identities with binary contact flags eliminates the improvement entirely.

Identifying Security Platform Product Abuse with Machine Learning

Shaefer Drew, Michael Brautbar, Paul Knight, Edward Raff, Lana Peric-McDermott, Simran Sarin et al. cross-listed Sophisticated threat actors can misuse security platforms inside customer environments or run bypass experiments against the product itself, often using living-off-the-land (LOTL) techniques rather than easily detected malware. The authors describe a deployed machine learning detection system for this kind of product abuse, designed around multiple data modalities spread across different databases, a cold-start problem caused by the rarity of such events, and operational limits on cost and performance. They report a 35% increase in product abuse coverage alongside a 30% reduction in monthly alerts, plus adaptability to changing attacker behavior. A retrospective evaluation examines the value of explainable features and counterfactual performance on previously identified attacks.

Fast And Accurate Text Content File Type Identification

Manu Nandan, Michael Brautbar, Edward Raff Identifying file types from content matters in cybersecurity, where magic numbers and extensions cannot be trusted, but model-based tools such as Magika are computationally heavy and parser-based tools are less accurate. The authors propose a neural network specialized for text-content files, especially source code. On open-source files it is more accurate on average than existing tools while being approximately four times faster than Magika and 28% smaller.

Co-Evolving Zero-Day Jamming: Adaptive Attack Synthesis and Graph Attention-Based Online Detection

Ghilas Aissou, R\'emi A. Chou, Taejoon Kim cross-listed Detectors for previously unseen (zero-day) radio jamming strategies are hard to evaluate because existing attack models assume prior knowledge of the target receiver, and existing detectors miss the global temporal-spectral structure of jamming and cannot separate new strategies as they emerge. The authors pair an online detector, which combines a graph attention network (GAT) with Dirichlet process (DP)-means clustering to classify known strategies and discover new ones, with a reinforcement learning jammer that treats the receiver as a black box, infers the detector's state through hypothesis testing, and trades attack impact against stealth. In simulation the jammer achieves 33% higher attack efficacy and 67% higher stealth than benchmark attackers, and the detector reaches 20% higher detection accuracy than benchmark detectors against it.

Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis

Naga Ganesh, Chandrashekar M S, Lakshmi Pedapudi, Aakash Singh, Vineet Singh cross-listed Farmer.Chat, an advisory service for smallholder farmers, must diagnose crop problems from a single low-quality phone photograph with no accompanying text, and its production system offers no adjustable thresholds or extensible label set. An analysis of about 1.16 million photographs from Ethiopia, India, Kenya, and Nigeria shows the production quality gate rejected 46.8% of images and that 35.8% of problems labelled as disease were actually pests. The work splits diagnosis into a quality gate, a crop detector, and a disease or pest detector, comparing a single fine-tuned Qwen3-VL-4B against small specialists (DaViT, YOLO26, MobileNetV3). A MobileNetV3 gate replaces GPT-4o at 86.9% F1 in 12 ms, and a hierarchical DaViT-Base reaches 95.41% crop accuracy versus 91.46% for the production baseline, while the fine-tuned vision-language model uniquely handles all stages in one call and can ask for a better photo.

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

Mushir Akhtar, M. Tanveer, Mohd. Arshad cross-listed Benchmark scores for medical vision-language models may not hold up when the evaluation setup changes, so the authors audit BioMedCLIP, CheXficient, MedSigLIP, and a general-domain OpenCLIP comparator for tuberculosis screening on 12,200 chest radiographs from Montgomery, Shenzhen, TBX11K, and VinDr-CXR. No model leads on every cohort and reliability criterion, prompt wording changes AUROC in 21 of 48 controlled comparisons, and replacing healthy controls with sick non-tuberculosis controls lowers AUROC by 0.075 to 0.306 for all four models. Thresholds set for 95% sensitivity on TBX11K hold in only four of sixteen target evaluations, and a supervised model drops from 0.999 AUROC on validation to 0.629 on external cohorts. The authors conclude that portability claims should name the full evaluation specification instead of being attributed to a checkpoint.

LLM-Generated Feature Pools for Time Series Anomaly Detection

Youssef Attia El Hili, Malik Tiomoko, Corinne Ancourt The authors ask how far a simple statistical pipeline can go on univariate time series anomaly detection: compute a small pool of sliding-window statistics, score windows with a robust median-absolute-deviation model, and pick a feature subset per domain on a held-out split. On TSB-AD-U it reaches 0.529 per-series VUS-PR, above the best neural (0.45) and statistical (0.44) leaderboard entries, and ablations show the candidate feature pool moves the score far more (0.226) than the selection strategy (0.031). They therefore generate a pool per domain by prompting a multimodal LLM with example windows from that domain. Generated pools match the hand-crafted one, and selecting over the union of both lifts the pipeline to 0.588, matching the best leaderboard entry.

EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid Enterprise generative AI projects often fail to show business impact, which the authors attribute largely to measurement: public benchmarks show what a model can do, not whether a specific workflow is reliable, safe, and worth scaling on an organization's own data and controls. EnterpriseVal is a use-case-level evaluation system comprising a formal specification of the frozen configuration under test, a metric catalogue, a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference, an executable threshold gate producing REJECT/CONDITIONAL/SCALE decisions, and a value-and-risk model. In a pilot at a global bank, credit-memo drafting reached 88% citation precision and a 1.6% hallucination rate against gates of 70% and 5%, and analyst refinement effort in procedure transformation fell from an estimated 27.4 to 2.9 hours per document. The authors explicitly separate established results from pilot evidence and open hypotheses.

LLMs as Feature Engineers for Text-and-Tabular Prediction

Merwan Barlier, Blaz Skrlj The framework automates extraction of interpretable, schema-bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic feature definitions, a separate extractor LLM materializes them, and a downstream tabular model scores them, with explicit model errors such as AUC ranking inversions translated into natural-language feedback that steers the next round of proposals. On three public datasets this error-driven loop speeds up feature discovery by up to 3x compared with unguided search, and the generated features complement TF-IDF and dense embeddings so that the combination beats any subset. The discovered features also dominate SHAP importance rankings, giving a semantic audit trail for each prediction.
44 more specialized papers

Large Language Models 39

Do small language models know what they don't know?

Prashant Mudgal Entropy-based confidence signals are tested as a way to improve Small Language Models (SLMs) under 3 billion parameters running entirely on consumer hardware, using seven approaches across 7 model pairs and 5 natural language understanding benchmarks. Token-level entropy turns out to be effectively blind at this scale: in 91% of dataset-model combinations, mean token entropy is near zero whether or not the answer is correct. Semantic entropy, computed by sampling several answers, clustering them by meaning and measuring the spread, does recover a usable signal, and routing uncertain queries to a larger expert model raises accuracy by up to 50 percentage points. Cross-family routing such as SmolLM 360M to Phi-3.5-mini averages +22.0% versus +6.8% for same-family routing, indicating that expert quality matters more than architectural compatibility.

Reviser: Revision-Capable Text Generation via Autoregressive Cursor Actions

Sean Diab Text generators that can go back and revise earlier content usually pay for that flexibility with repeated computation over the whole sequence. Reviser is a decoder-only Transformer that writes onto a mutable canvas by predicting one cursor-relative action per step, either INSERT(token), MOVE(Δ) or STOP, so it is autoregressive over the edit history rather than over final text order. On a continuation benchmark it is strongly preferred to the diffusion-style models SEDD and MDLM in arena evaluations, and trajectory statistics show it genuinely makes frequent backward moves and mid-canvas insertions. It is competitive with size-matched autoregressive baselines at 100M and 300M parameters and needs substantially less inference compute than multi-pass refinement and diffusion-style baselines.

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Sy-Tuyen Ho, Minghui Liu, Furong Huang As model-written peer reviews enter training corpora, later AI reviewers may learn from earlier ones, and this study simulates one step of that loop by fine-tuning Llama 3.1 8B on official ICLR reviews from 2018-2023 and then training four successors on ICLR 2024 data with varying mixes of official and model-generated reviews. Adding synthetic reviews compresses rating distributions and reduces semantic diversity both within a paper's reviews and across the corpus, a pattern the authors name scientific-judgment collapse. To counter it they introduce TrustReviewer, an open-source reviewing system that trains in a single stage on a corpus curated to remove low-quality and degenerate supervision, and applies paired activation steering at test time to correct residual collapse without further training or expert annotation.

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

Chuxu Song, Jiuqi Wei, Zhencan Peng Long-context inference is increasingly bottlenecked by prefill, where dense self-attention processes the whole prompt before generation starts, and block-sparse methods that score blocks by their centroid can miss a single highly relevant token buried among irrelevant ones, a failure the authors call mean dilution. RBS-Attention is a training-free sparse-prefill method that combines a centroid branch for average relevance with a rescue branch that uses each key block's maximum radius and its prompt-, layer-, and head-dependent distribution to flag blocks at risk of underestimation, while keeping regular block-sparse FlashAttention execution. On H100 GPUs at 128K context with Qwen3-30B-A3B-Instruct-2507-FP8 it achieves 20.65x standalone prefill-attention speedup, 11.92x within vLLM, and 5.97x end-to-end time-to-first-token speedup. On dense Qwen3-32B it scores 88.65 overall on RULER versus 89.52 for dense attention, with further evaluation on LongBench-v2, InfiniteBench, and Video-MME.

Attention-Aware Routing: Coupling Routing and Attention in MoEs

Despoina Kosmopoulou, Anastasios Tsetsilas, Efthymios Georgiou, Giannis Karamanolakis, Swastik Roy, Alexandros Potamianos Routers in Mixture-of-Experts language models usually choose experts from a token's hidden state alone, using little contextual information. Attention-Aware Routing (AAR) augments the router with temporal and spectral features extracted from a sliding window of attention weights, and trains only the routing parameters while the base transformer stays frozen. On OLMoE it improves GSM8K by +3.37 percentage points over a routing-only supervised fine-tuning baseline, and shortens incorrect, long-diverging generations while leaving correct answers' length unchanged. The authors also show routing and attention form a coupled circuit, where routing changes at one layer amplify attention sinks at the next, and that the method is depth-sensitive: applying it at all layers can hurt factual retrieval, whereas math reasoning gains persist when it is applied deeper in the network.

Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing

Amir Jalilifard, Anderson Rocha, Eric Wong, Marcos Medeiros Raimundo Hallucinated and faithful responses are distinguished here by analyzing the topology of information flow in a model's attention graphs. The method uses Forman-Ricci curvature to find structural bottlenecks and combines semi-local and global flow features of attention heads, requiring only a single forward pass. It consistently improves over existing attention-based and multi-response baselines on two hallucination-detection benchmarks across several model architectures. Further analysis ties hallucination to impaired context sharing among tokens: over-reliance on self-attention, diffuse retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala How fine-tuning reshapes a language model's internals is examined by comparing changes in attention patterns and layer-wise activations against the task-relevant components identified by EAP, a circuit-attribution method. The causally important components concentrate in specific layers, but their layer distribution is largely uncorrelated with the layers whose representations change most during fine-tuning. Overlap in EAP-identified components across tasks also does not imply transfer: when tasks differ in nature, such as classification versus generation, fine-tuning on one can degrade the other precisely when their component overlap is high.

The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators

Tarfah Alrashed, Fatma Ozcan, Per Jacobsson, Tal Neiman, Xianshun Chen cross-listed Modern analytics platforms let SQL queries call AI operators over unstructured data, but the usual way of grading generated queries — comparing exact execution results — breaks down when those operators return non-deterministic outputs. The authors catalog the resulting failure modes and propose a multilayered evaluation framework that validates the relational logic and the AI semantics separately, tested on both BigQuery and ThalamusDB. Conventional execution accuracy recognized as few as 25% of correct translations, and a state-of-the-art LLM-based autorater falsely rejected 32% of accurate queries, while the decoupled framework reached up to 97.2% overall accuracy.

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

Zhihao Shu, Md Musfiqur Rahman Sanim, Jie Hu, Kun Yuan, Minghai Qin, Gagan Agrawal et al. Long-context language models on phones are bottlenecked by the key-value (KV) cache, which grows linearly with sequence length and is touched at every decoding step, while eviction discards tokens irreversibly and offloading stalls on I/O. TierKV predicts future cache demand from prefill hidden states before decoding starts and jointly assigns tokens to exact, low-rank, and flash-offloaded tiers under memory and accuracy budgets, using a closed-form solver that picks tier boundaries and per-layer ranks at runtime while keeping the full context reachable. Across eight text, vision, and audio models on three mobile systems-on-chip, prefill throughput improves by up to 17.6x over existing mobile LLM frameworks, with RAM-resident KV cache cut 12.5–34% and only minor accuracy degradation.

Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency

Wenhan Yu, Wenxin Wu, Hao Wang, Lei Sha The work studies paraphrase-induced hallucination, where a model answers a factual question correctly in its original wording but wrongly under a semantically equivalent rephrasing; generic paraphrases are poor test data because near-copies give weak signal and highly diverse ones break equivalence. Hallucination-R1 trains a paraphrase generator in two stages, first stabilizing meaning-preserving and diverse paraphrasing, then rewarding paraphrases that expose factual-consistency degradation in downstream question-answering models. On SimpleQuestions, PopQA, and TruthfulQA it achieves a strong consistency-diversity trade-off and exposes robustness failures across multiple model families that analyses show are not reducible to surface artifacts or semantic drift. A lightweight fine-tuning study indicates the generated data improves robust accuracy under paraphrase variation.

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen et al. Comparing model and human behavior rigorously is hard because of the breadth of tasks humans perform, so CogGym provides a unified framework that uses a semi-automated, human-in-the-loop pipeline to standardize diverse cognitive science experiments into a task-agnostic Experiment Markup Language (EML) for matched-trial comparison. The initial release curates 258 experiments from 100 papers on human commonsense reasoning and evaluates 50 large language models against human responses. Larger and more recent models reproduce human judgments better, but progress is considerably slower than on formal-reasoning benchmarks such as math and coding. The best models reach only R² = 0.59 on text, 0.58 on image, and 0.43 on video experiments, far below human split-half reliability of 0.93, 0.95, and 0.92.

How Many Humans Is a Judge Panel Worth?

Chao Li, Yingying Yu, Yunfeng Li The question is how many human judgments a panel of language-model judges is actually worth, and the answer depends on what is being matched. The authors audit categorical judge panels against empirical human label distributions rather than a single gold label, defining one effective size (nu_H) by matching the spectral diversity of panel residuals to independent human-reference draws and another (nu_MSE) by matching distributional squared error. Across three ChaosNLI tasks, the same 32-judge panels are worth 4.24-6.50 humans by spectral diversity but only 2.30-3.75 by distribution recovery, and constructed examples show greater diversity can accompany worse recovery. A large share of residual variance lies in a shared consensus direction that averaging does not remove (43.8% on MNLI-m, 33.7% on SNLI), so effective panel size is a target-specific measurement rather than a general human-replacement rate.

CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices

Wenquan Zhou, An Wang, Jing Liang, Peien Feng, Jingqi Zhang, Yaoling Ding et al. cross-listed Large language models (LLMs) are increasingly used to build and analyze cryptographic implementations for Internet of Things (IoT) devices, where attackers with physical access target the implementation rather than the algorithm, yet no benchmark covers this area. CESBench provides 380 expert-written items across six sub-domains (side-channel, fault injection, implementation, countermeasures, evaluation, and integration) in four formats: multiple choice, security judgments with justification, scenario diagnoses, and code tasks graded by 572 test cases. Eleven models score between 54.4% and 83.6% overall, with top scores of 98.6% on multiple choice and 95.1% on code but only 58.8% on judgment. Across models, 88.5% of security verdicts are correct, yet their justifications earn only 53.4% of the rubric marks.

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen, Kiet Van Nguyen et al. Standard tokenizers treat text as characters or statistically derived subwords, ignoring the internal phonological structure of syllables and often requiring large vocabularies. Phonemic Tokenizer converts each Vietnamese or Chinese syllable into the International Phonetic Alphabet (IPA) and factorizes it into onset, rime, and tone, with the three components sharing one sequence position so that length is preserved, which yields vocabularies of only 112 entries for Chinese and 256 for Vietnamese with no corpus-dependent vocabulary learning. PhonemicBERT, which combines the component embeddings and reconstructs masked syllables with three prediction heads, is competitive with or better than character, subword, and SubChar alternatives in controlled Chinese pretraining. The Vietnamese version matches or exceeds established Vietnamese and multilingual pretrained models.

Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue

Marina Mitiaeva, Lu Xiao The authors ask whether conversational AI genuinely participates in cooperative communication or only reproduces its surface forms. They compare 15,881 human-ChatGPT dialogues with 10,784 human-human multi-turn dialogues, using mixed-effects models to predict turn-to-turn alignment from morality, politeness, and related features. The AI shows cooperative surface features without the underlying mutual adaptation: moral content looks preconfigured rather than negotiated, warmth appears without sensitivity to face, and linguistic convergence steadily declines. Some cooperative mechanisms reverse direction with AI: hedging and softening, which accompany greater accommodation between humans, accompany reduced alignment when the AI produces them. Giving users agency to shape the exchange is the most consistent predictor of alignment in both settings, and the lower moral assertiveness of newer models does not come with better cooperation.

Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals

Yamato Narita, Issei Sato Quantizing both the weights and activations of large language models to 4 bits (W4A4) after training is difficult because activation outliers waste quantization resolution. It has been unclear which error components weight optimization, channel scaling, and orthogonal rotation each address. The authors exactly decompose the local quantization error into an activation-guided weight compensation term and an orthogonal residual, and they bound the residual in terms of persistent outlier channels and regular activations. The bounds explain why random signs in Hadamard rotations suppress constructive interference among outlier channels and why sampling several sign patterns helps. They also show how second-moment balancing yields an L2 scaling rule that relaxes to SmoothQuant-style L-infinity scaling. Across eight Llama and Mistral models, configurations built from these guidelines without any backpropagation perform competitively with gradient-trained SpinQuant.

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

Xavier Suau, Alex Ferrando de las Morenas, Luca Zappella, Samy Bengio When language models reason in chain-of-thought or pass free-text intermediates to each other, structured information is serialized into natural language, and a round-trip protocol measures how much tree structure survives: one model turns a procedurally generated arithmetic expression into a word problem, another recovers the expression, and symbolic equivalence gives an exact check. Testing all pairings of sixteen models shows the channel is lossy and asymmetric, with swapping generator and extractor roles shifting accuracy by up to 60.4 points and the best pair reaching 92.9% by combining two different models. At least 73.6% of round-trip failures originate at generation, and difficulty tracks tree structure (operator count, depth, right-branching) rather than model family. About 3,600 fine-tuning examples lift every open-weight model above untrained Gemini-3.1-Pro under matched semantics, and a disjoint-domain setting still helps while leaving a gap to the frontier.

Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction

Abhishek Bhandari, Gaurav Harit Training-free correction of optical character recognition (OCR) errors via in-context learning had not been studied for Devanagari script. The authors evaluate large language models from 3B to 32B parameters on a 20,000-sentence Hindi and Marathi benchmark across five news domains, comparing domain-random example selection, dense semantic retrieval, and their CharBM25, which retrieves in-context examples by character n-gram BM25 similarity over the OCR input to surface shared error patterns. CharBM25 beats random selection by 2.8-4.0 points of absolute word error rate on Hindi and matches or exceeds dense retrieval without a GPU, while Gemma-3-27B cuts word error rate by 55.0% on Hindi and 33.3% on Marathi. Few-shot gains are capacity-gated: models under 8B do not reliably improve on the OCR baseline, and 3B models degrade more Marathi sentences than they fix.

Trading Depth for Time in Recurrent Transformers

Zeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen Recurrent Transformers feed each token's high-level hidden state into the next token's computation, raising the question of whether extra compute is better spent on more temporal steps or more layers. Using Latent Recurrent Transformers, the authors insert one latent thought token between consecutive vocabulary tokens, passed through the same shared L layers, and compare it with a 2L-layer model without thought tokens, so both execute the same number of Transformer blocks per decoded token. On 16- and 20-layer mixture-of-experts NanoChat backbones, the shallower model comes within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with roughly 48% fewer total parameters.

One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction

Hongliang Li, Lu Wang, Yong Xu, Hanyang Chen, Zhitao Hou, Xiaoting Qin et al. Prompt optimization for enterprise information extraction usually tunes one prompt against a global objective, even though different users need the same document reorganized differently. Self-Meta-Evolve maintains a dedicated prompt per user and refines it with two loops: an inner loop that edits structured prompts from persona-conditioned feedback, and an outer loop that evolves the meta-prompt by distilling successful editing patterns. The authors release a benchmark of 292 simulated enterprise users generated from O*NET occupational taxonomies, on which the method reaches a 74.58% success rate, 13.56 absolute points above the strongest prompt-optimization baseline, and 52.54% after only two iterations. In a double-blind study with twenty professionals, the adapted prompts won 71% of pairwise comparisons against static baselines.

Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen Dense large language models are slow and costly at inference, existing acceleration methods often hurt quality, and training Mixture-of-Experts (MoE) models from scratch is resource intensive. L0-MoE converts a dense model into a lightweight MoE using L0 regularization, with a cluster confusion matrix guiding domain-aware dataset curation and dynamic batching keeping training efficient. The authors report up to 2.5x speedup over the dense model while maintaining competitive performance, outperforming existing acceleration baselines.

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

Harish KB, Jagadeeswaran M, Pradheep P, Yuvanesh S, Sivakumar T Running large language models (LLMs) on consumer hardware is constrained by compute and memory, and common speedups such as quantization or speculative decoding usually need retraining, per-architecture tuning, or a separate draft model. SpecQuant is a training-free framework that derives INT4, FP8, and FP16 variants from one shared base model and routes each query by predicted complexity, sending simple or factual queries to lightweight variants and complex reasoning or long-context inputs to full precision. Because the variants share weights, the quantized ones can serve as speculative-decoding drafts with adequate token acceptance. On Qwen2.5-based models evaluated with MMLU, AlpacaEval, and GSM8K, the authors report 35-43% speedups with accuracy loss under 2%.

RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding

Qiao Hu, Yepeng Weng, Bo Zhang, Takehisa Yairi Dynamic-tree speculative decoding methods such as EAGLE-3 work well under greedy decoding, but at temperatures above zero their deterministic top-K expansion collapses the draft distribution into one-hot probabilities and sharply lowers the acceptance rate. RheoSampling decouples the two roles of the draft distribution: it injects a sampled token among the top-K slots and gives it a proxy probability for tree expansion and pruning, while keeping its true sampling probability for verification. The authors prove the method stays lossless via an equivalence-class analysis, making it the first dynamic-tree method to combine context-aware top-K tree construction with stochastic sampling, and they add an optimal-transport-based verification strategy and a sparse draft mechanism. Experiments across several LLMs and benchmarks show higher acceptance rates and speedups over existing dynamic-tree methods.

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz cross-listed Speculative sampling speeds up LLM inference and watermarking tracks output provenance, but prior work shows the two are hard and possibly impossible to combine without sacrificing one. The authors propose a multi-draft speculative sampling algorithm built on Poisson processes and an exact list-coupling-without-communication scheme, which makes it invariant to the choice of drafter. It can embed an unbiased watermark without degrading speculative acceptance, and the authors describe it as the first multi-draft, drafter-invariant scheme that preserves both watermark strength and sampling efficiency, with experiments confirming both.

Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

Chenye Ke, Zirui Liu, Qi Liu, Yan Zhuang, Jintao Zhang, Zhenya Huang et al. Detecting whether a text was in an LLM's pretraining data is hard because high likelihood can reflect either training exposure or simply predictable text, so likelihood-only detectors mistake predictable non-members for members. The authors instead score prediction loss relative to predictive entropy, show analytically that this entropy correction preserves the expected membership signal while reducing its variance, and give the score a Helmholtz free-energy interpretation, yielding Energy Transfer Detection (ETD). Across extensive experiments ETD achieves the best average detection performance, improving average AUROC by up to 3.5% and true positive rate at 5% false positive rate by up to 5.1%, and remains robust across settings.

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

Yanxiao Liu, Sicheng Wan, Deniz G\"und\"uz Best-of-n sampling aligns LLM outputs at inference time by picking the highest-reward candidate, but hard maximization gives only coarse control over the trade-off between reward and drift from the base distribution. ExpBoN is a soft best-of-n variant based on the exponential-noise report-noisy-max mechanism; it admits an exact finite-n decomposition that yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence, backed by convergence and regret analyses. Integrated into the guided speculative inference framework as ExpGSI, it cuts estimated computation by 14%-39% for Qwen2.5-Math and by up to 45% at n=16 for Qwen3 on MATH500, MMLU-STEM, and Minerva Math while keeping comparable accuracy.

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Richard Zhe Wang Gating the value pathway of attention is reported to improve language model pretraining, but prior studies disagree about why. The authors argue that such gates supply two things softmax attention lacks: abstention, which lets a head output nothing despite attention weights having to sum to one, and noise filtering, which suppresses interference from superposed features in the residual stream. In matched models from 10M to 350M parameters, abstention is supplied through a learned per-head sink logit and filtering through a gate on each value, and the benefit of abstention shrinks with scale while the benefit of noise filtering grows, with abstention accounting for nearly all of the gain at 10M and filtering for most of it at 350M. The best model at every scale has both primitives, which add negligible parameters and remain compatible with the key-value cache.

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

Andre Bacellar cross-listed Multi-hop retrieval failures cluster in structurally predictable subpopulations of queries rather than spreading uniformly. The authors prove that reducing confident failures is possible only when retrieval features carry mutual information about success, and that no single approximate-nearest-neighbor score feature is best across all failure regimes. RegimeAbstain computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query and retrieval features that needs no additional large language model call, and uses it for calibrated abstention, evaluated with a new Confident-Wrong-Answer Rate (CWAR) metric on MuSiQue, 2WikiMultiHopQA, and HoVer. RCS is best or co-best against eight confidence baselines in all five conditions, and on MuSiQue with an LLM-judge pipeline it cuts CWAR from 39.5% to 20.6% at 50% coverage; a model trained on MuSiQue transfers to 2WikiMultiHopQA with only a 0.5 point loss in AUC.
11 more specialized papers

Theory 32

dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD

Jiahe Yan, Pratik Chaudhari, Leonard Kleinrock cross-listed Distributed training must cope both with stragglers, slow workers that delay gradient aggregation, and with Byzantine workers that send corrupted updates. dSTAR is a lightweight distributed stochastic gradient descent (SGD) scheme that collects gradients only from the first k workers to respond and then filters them by their deviation from an ensemble median. The authors prove that it is (α, f)-Byzantine resilient with a linear convergence rate. In experiments it keeps accuracy high under attack, whereas other Byzantine-resilient methods often suffer accuracy drops of 40-50%.

Recursive Language Models Generalize Out of Domain

Chenxiao Yang, Zhiyuan Li, David McAllester, Nathan Srebro The question studied is when restricting what a language model can see improves learning, comparing standard chain-of-thought (CoT), which reads the full reasoning trace, against recursive language models that solve each subtask in an isolated context. In-distribution, CoT can efficiently simulate the recursive rule, so its generalization guarantee differs only by a constant factor and recursion offers little. Out of domain, however, CoT can fit the training data through shortcuts that depend on context outside the current subtask and break when those tokens change, a failure mode that context isolation rules out. Because simplicity bias selects the shortcut even though CoT's hypothesis class contains the correct rule, the authors argue that covering the right rule is not enough for out-of-domain reasoning, in contrast to classical learning theory.

Do Quantum Models Scale Like LLMs?

David S. Berman, Ying-Jer Kao, Roger G. Melko, Alexander G. Stapleton Neural scaling laws are examined for RydbergGPT, an autoregressive transformer trained on qubit measurement data from interacting Rydberg atom arrays, a quantum system with a finite-size remnant of a critical point. Near the critical point, loss versus training dataset size follows a power law with a loss floor, but away from criticality the power-law fit degrades substantially. Using an entropy-normalised, finite-sample-corrected mutual information "two-point" function, the authors find that near-critical measurement statistics most closely resemble those of natural-language corpora, while far-from-critical configurations show faster-decaying correlations. They argue this supports the view that multi-scale dependence in data underlies stable scaling, making scaling a property of the model-data pair rather than the model alone.

Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks

Emanuele Zangrando, Marco Sutti, Francesco Tudisco Linear factorization blocks of the form W = BA, found in LoRA adapters, low-rank layers, and attention query-key products, have a non-unique factorization that can destabilize training and cap usable learning rates. Stiefel-AdamW constrains one factor to the Stiefel manifold while leaving the other Euclidean, which rules out factor blow-up while keeping AdamW's coordinate-wise preconditioning; moments are estimated in ambient space and geometry enters only through a tangent-space projection and a retraction. The authors prove standard convergence guarantees and report consistent improvements over strong baselines at essentially no extra cost over AdamW on LoRA-style fine-tuning of GPT2, ViT, and Mistral 7B and on full GPT2 pretraining on OpenWebText.

Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs

L\'eo Nicollier (CB, ATT), Enric Meinhardt-Llopis (CB), Marc Pic (ATT), Pablo Mus\'e (CB, IFUMI) et al. Joint-Embedding Predictive Architectures (JEPAs) avoid representation collapse by forcing embeddings to match a target distribution, and prior work argued the Gaussian is the unique target guaranteeing linear recovery of latent variables under Euclidean assumptions. This analysis extends the theory to latents living on embedded Riemannian manifolds and derives conditions on geometry and positive-pair dynamics that guarantee linear recovery. When latents are uniform on a sphere and representations are matched to the same distribution, every optimal representation recovers the latent state up to an orthogonal transformation, so Gaussian uniqueness is not universal, and the approximate-recovery bound is strictly tighter than in the Gaussian case. Experiments on Gaussian, spherical, toroidal, and high-dimensional Clifford-torus worlds show geometrically matched targets give better linear recovery and mismatched ones distort the latent structure.

World Modeling in Transformers

Pierre Beckmann, Matthieu Queloz, Andre Freitas Behavioral failures can make a transformer look as though it lacks a world model even when its internal representations are faithful. Using mechanistic analysis and causal interventions on TaxiGPT, a transformer trained on random walks through Manhattan, the authors show that the model represents intersections and streets, tracks its position, and navigates with a goal compass. They trace its failures to interference between superposed intersection features that disrupts localization, partly contained by affordance packing, which groups intersections that share the same legal moves. They also propose mechanistic indicators for comparing models and find that different world-modeling capacities emerge at different stages of training.

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

Wenpeng Zhang, Runsheng Yu, Peilin Zhao Adaptive optimizers like AdaGrad and Adam scale each parameter entry independently and ignore the matrix structure of neural-network weights, and no general theory exists for deriving matrix-aware adaptivity. The authors build an Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters and derive Row-AdaGrad and Column-AdaGrad, which scale updates by accumulated row-wise or column-wise gradient norms. They prove regret bounds that can be strictly tighter than entry-wise AdaGrad under structured gradients, and experiments on matrix factorization and deep network training show better stability at larger learning rates and greater depths.

Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources

Isaac Manring, Kejun Huang Nonlinear Independent Component Analysis (nICA) aims to recover the true independent hidden factors that generated data rather than a scrambled version of them, which is not possible in general without extra assumptions. The authors prove identifiability up to trivial ambiguities when the generating function is real analytic and the source densities have a finite number of discontinuities in the first derivative, with the Laplace distribution as the main example; the proof exploits the contrast between those kinks and the smoothness of analytic functions. Because real analytic functions can be approximated by normalizing flows or variational autoencoders with standard activations such as tanh, softplus, or GELU, the result applies to existing training pipelines with minimal changes. Experiments on synthetic and real data support the theory, and on CelebA the method recovers several interpretable latent factors.

Schedule optimization for tau-leaping in masked discrete diffusion

Cecilia Secchi, Giacomo Zanella cross-listed Masked discrete diffusion models are sped up with tau-leaping, which reveals several coordinates in parallel per step and replaces their joint conditional law with a product distribution, incurring a factorization error even with perfectly learned predictors. The authors give an exact integral representation of this error in terms of a distribution-dependent dependence density that tracks how conditional dependence evolves as more coordinates are revealed, develop estimators for it, and derive stationarity equations that characterize the unique optimal denoising schedule for a finite number of steps under a monotonicity condition. In the limit of many coordinates N and steps K, when the dependence profile converges to a strictly positive continuous function, optimizing the schedule improves only the leading constant and not the N/K scaling of the error, whereas degenerate profiles allow schedules that improve the asymptotic order over the uniform schedule. Examples based on stationary processes and exchangeable mixtures illustrate the two regimes.

Multiplicative Optimism for Constant Regret in Games

Ashkan Soleymani, Georgios Piliouras cross-listed Multiplicatively Optimistic Regret Matching (MORM) is an uncoupled learning rule for finite general-sum games. Under simultaneous full-information self-play, every player achieves external regret of O(√n log d) uniformly over all horizons, that is, constant in the number of rounds, using only one-step optimism. The analysis combines a potential-based regret-matching argument with multiplicative stability and Hellinger control of how far strategies move between rounds. A learning-rate safeguard additionally guarantees O(√(T log d)) regret against adversarial utilities.
22 more specialized papers

Agents 27

Voice-Light: A Full-Duplex Cascaded Voice Agent with Causal Turn-Taking and Speculative Generation

Bertil Braun cross-listed Voice-Light is a full-duplex cascaded voice agent designed to handle overlapping speech without canceling on every acknowledgment, prepare responses before a turn is certain, and keep canceled audio out of conversation history. It combines immediate acoustic onset detection, a causal turn-taking adapter sharing a streaming speech recognition encoder, reversible playback control, private speculative response generation, and tool calls that execute concurrently with audible bridge speech. On 1,673 real-conversation silence candidates, a learned completion checkpoint kept a 2.70% false-cutoff rate but reached only 12.53% end-of-turn recall versus 95.60% for a Silero timing policy, so the deployed system retains a hybrid controller. In three unscripted microphone sessions, 36 response turns had a 758 ms median from final voice-activity endpoint to first server audio, which the authors present as an instrumented case study rather than a controlled evaluation; data, models, code, and deployment configuration are released.

Scaling Discovery through Test-Time Communication

Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos Whether letting parallel AI agents communicate at test time actually beats running them independently has had mixed evidence; here, role-free agents share findings through a common directory while working on hard tasks. On ARC-AGI-3, a team of k communicating agents (team@k) matches the success rate of 4k independent agents, with the advantage growing as k increases, and teams reliably solve tasks that no single agent can. The gains carry over to research-style tasks: communicating agents beat the prior best-known score on polyomino packing, and a four-agent team produced a 1,957-byte MNIST classifier with 99.4% test accuracy, smaller than the best-known human solution. The authors note that independent agents can still win when compute is limited or no clear progress signal exists.

CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop

Kailai He, Zhihao Wu, Linhai Zhang, Runcong Zhao, Yulan He, Jiazheng Li Deployed tutoring tools generally serve fixed item banks and treat a wrong answer as one bit of signal rather than evidence about what a learner misunderstands. CoLearn keeps a persistent learner-state memory updated by a soft-evidence variant of Bayesian Knowledge Tracing in which a language model acts as a continuous observation function, generates questions aimed at the weakest topic and recurring misconceptions, and exposes the personalization through live progress views and blind A/B comparisons. In that blind evaluation, memory-conditioned questions were preferred 68–69% of the time over non-personalized ones, and in persona simulations with hidden ground-truth mastery the agent's belief converged toward the true value.

Can Agents Design Better Chips with a Higher Level Abstraction?

Zijian Ding, Yang Zou, Yizhou Sun, Jason Cong Most language model agents for chip design write register-transfer level (RTL) code directly, leaving them to optimize at a low level of abstraction. Four workflows are compared on FPGA targets — direct RTL design, agent-driven high-level synthesis (HLS), post-compiler HLS refinement, and post-HLS RTL refinement — with the latter two combined into AHRR (Agent-based HLS with RTL Refinement). Across an 11-task benchmark suite, AHRR achieves a 2.6x geometric-mean speedup over direct RTL design, and case studies attribute the gain to HLS distilling reusable design knowledge into abstractions the agent can work with while RTL refinement recovers lower-level optimizations.

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

Md Tahmid Rahman Laskar, Xue-Yong Fu, Gundeep Singh, Karol Chang, Kevin Sanders, Shi Zong et al. Agent models are routinely scored one decision at a time, predicting the next action from a gold interaction history and comparing it to a reference — the question is whether gains under that protocol carry over to running a task autonomously. Pre- and post-supervised-fine-tuning Qwen3 models at 4B and 14B and Gemma 3 models at 4B and 12B were evaluated both ways on multi-turn customer-support workflows. Fine-tuning improved next-turn and text-turn success for every model, yet none of the four fine-tuned models succeeded under holistic workflow evaluation, with strict trajectory completion reaching at most 10.4%, and tool-specific gains varied inconsistently across metrics, leading the authors to argue for reporting text quality, local action correctness, tool execution, and end-to-end completion separately.

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

George Ma, Benjamin Mikek, Haoyu Li, Ferhat Erata, Yuhao Zhang, Zeren Shui et al. Agentic coding benchmarks judge patches with held-out test suites that are inherently incomplete and increasingly vulnerable to memorization, while formal verification has so far only covered standalone tasks whose specifications are handed to the model. Benchproofer turns a coding task with a known-correct patch into a formally verified one by writing a specification for the new code, summarizing the existing functions it calls with axioms, and admitting an instance only when mechanical and adversarial gates agree; applied to SWE-bench Verified it produces SWE-Proof, 500 real issues checked by proof rather than test, and extends to SWE-bench Pro. A quarter to a half of test-passing patches from two frontier models admit counterexamples, and supplying a correct formal specification lifts Opus 4.8 from 85% to 95% resolution. Writing the specification is the hard part: models asked to produce their own gain nothing over an unaided baseline, only 62% of those specifications pass audit, and the usual failure is faithfulness — constraining part of the required behavior and leaving the rest free.

Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Hao Fu, Baiting Zhu, Minglei Chen, Yinjie Huang, Shuai Ding cross-listed Large language model (LLM) agents can propose, implement, and evaluate model changes, but in long-running online settings a completed run can still support a wrong conclusion when a code change is a no-op, data windows leak, evaluator semantics drift, or the two arms traverse different serving funnels. EvoPilot is a human-gated method for long-horizon online autoresearch in which role-specific agents execute each round through a versioned domain skill and typed adapter, durable records preserve experiments and failures, and deterministic checks enforce recorded lessons. In a 37-day campaign on the retrieval system behind Video Deep Dive, an earlier primitive autoresearch attempt had blamed a 22-percentage-point offline hit-rate drop on an interaction head; EvoPilot's verification traced the drop to a pre-existing evaluation defect, and after repair a matched comparison measured a 3.20-percentage-point offline improvement. A separate seven-day randomized online test estimated a 0.66% relative increase in the product's Good Search Result Rate for Retention, and durable state recovered an interrupted round while saving roughly five GPU-hours through artifact reuse.

PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking

Qiufeng Li, Chengxuan Wang, Rongqian Chen, Quan Cheng, Yihui Ren, Chia-Tung Ho et al. Macro placement in VLSI physical design is typically solved by one-shot numerical optimization of hand-crafted proxies such as estimated wirelength, which cannot easily incorporate visual layout context, design expertise, or downstream feedback. PlaceReasoner-Beta recasts it as closed-loop reasoning in a verifier-guided multi-agent framework: a vision-language model (VLM) planner proposes placements from the floorplan image, macro specifications, and connectivity; a geometric verifier enforces legality and expert principles; a physical verifier refines candidates with early implementation feedback; and a post-route optimizer improves promising layouts using final power, performance, and area results. The accompanying PlaceReasoner-Bench is a fully open end-to-end benchmark of 8 designs at two aspect ratios (16 tasks) scored on routed results and design-rule checks rather than pre-route proxies. The system achieves the best timing among design-rule-clean methods on all square tasks, reducing post-route total negative slack by 61.2% at 1:1 and 53.0% at 2:1 aspect ratio relative to the classical baselines, and shortens routed wirelength on most designs without optimizing it explicitly.

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

Yining She, Lei Lin Production LLM agents are re-evaluated constantly as they evolve, but full agent benchmarks are expensive to rerun. Using 574 historical runs of the benchmark for a production analytics agent serving tens of thousands of monthly active users, split chronologically into calibration and held-out periods, the authors compare random sampling, historical caching, fixed representative subsets, and adaptive testing based on item response theory (IRT). Multidimensional two-parameter (2PL) adaptive testing gives the best score fidelity: executing 200 questions, 38.5% of a full run, yields a mean absolute error of 1.03 percentage points. The team nevertheless deployed difficulty-stratified fixed subsets for operational simplicity, and shows they transfer without recalibration to five other agent families and stay stable with calibration windows as short as one day.

Authorization Revocation for Long-Running AI Agents: Root-Scoped Quiescence under Delegation and Asynchronous Execution

Genliang Zhu, Chu Wang cross-listed Long-running AI agents keep acting through credentials, delegated tasks, queues, and callbacks after their initiating process is cancelled or its credentials are revoked, so ordinary cancellation does not stop already scheduled work. The authors define root-scoped authorization quiescence, a certificate-backed guarantee that every acceptance under a retired authority root is accounted for and that none occurs after a local fence, while work with independent sufficient authorization is preserved. Under stated assumptions they prove properties including non-expansion after the cut, compositional soundness, and crash/replay stability. In a late-effect test suite that matches 17/17 registered outcomes, cancellation-only and cut-only executions accept an already scheduled late effect while cut-plus-fence executions reject it, and an independently implemented checker rejects 44/44 semantic regressions.

GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development

Xiuhui Zhang, Yi Chen, Shusheng Xu, Fan Li, Huan Wang, Tongkai Yang et al. Autonomous software generation (ASG) systems can deliver runnable applications without their interacting components actually meeting the specified behavior. GameASG-Bench addresses this with 47 browser-native game-generation tasks across 12 genres in 2D and 3D, each declaring an evaluation interface before generation (legal starting scenarios, player actions, stable snapshots, invariants) and scored with static L1 source checks plus browser-executed L2 checks using real input. Across nine agent stacks the best mean L2 check pass rate is 93.2%, yet the best strict task success rate is only 55.3% (26/47 tasks). For DeepSeek-V4-Flash, full tool access and larger turn budgets help while more reasoning effort does not help monotonically, and two harnesses that each solve 18 tasks overlap on only ten.

LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces

Steve Drew, Jiayu Zhou Buyers in marketplaces where AI agents complete tasks autonomously cannot easily tell which agent will perform best, because reported benchmark scores are hard to verify or compare across tasks, software, and budgets. LEGIT is a credentialing protocol in which a signed certification record binds measured quality and cost per solved task to a specific agent configuration, task domain, evaluation budget, and evidence, while a reputation layer ties records of past task outcomes to the same identity. Evaluations show that agent configurations with similar task success can differ substantially in cost, and that comparisons between them depend on the evaluation budget. A separate analysis quantifies the deposits and fees an attacker would need to manipulate reputation under a stated Sybil attack model.

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team) et al. Deployed agents generate many execution traces, but labeling outcomes or writing task-specific verifiers is costly, so the authors ask how to turn raw traces into reusable feedback without outcome labels. DENSE (Distilling Evidence from Nested Subtask Executions) organizes evidence of local progress, recovery, and unfinished requirements into nested shortcut trees. It compresses redundant attempts, reconciles issues across levels, summarizes finished branches, and expands unresolved ones. The evaluation protocol, REFIT, compares feedback methods from shared initial trajectories with environments and model contexts reset. Under it, DENSE achieves the highest strict pass rate among non-privileged feedback methods on Terminal-Bench 2.1 across four recipient models, improving strict pass rate by 7.12-15.64 percentage points over initial executions while using 19.0-43.6% fewer recipient tokens in reruns.

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba et al. Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, but final-score comparisons mix differences in models, topology, roles, and compute, so gains are hard to attribute. OpenMAS-GCom diagnoses these systems through controlled interventions that change one component while holding tasks, models, prompts, and budgets fixed: rewiring communication edges, removing specialist or critic agents, injecting incorrect intermediate messages, and disabling workers mid-execution. It evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains, plus 400 new G-MAS-Complex tasks requiring multi-document synthesis, conflict resolution, and sourced answers. Results show larger mean losses from removing specialists than from removing critics, differing robustness to bad messages and worker failures among systems with similar original scores, and different configurations winning on accuracy versus accuracy per token.

MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems

Kairui Yang, Minghao An, Xunkai Li, Ziheng Yi, Zekai Chen, Guangyuan He et al. Multi-agent systems built on large language models produce collaboration traces showing how agents plan, verify, and repair, and reusing them requires preserving each action's prerequisites and the outputs later agents depend on. Empirical studies show that grouping these dependencies into functional memory units improves retention, that linking units improves joint retrieval, and that the best combination of units differs depending on whether memory is presented as instructions or checklists. MACE builds on this with MemGoG, a graph of functional units (conditions, actions, outputs) connected by support, conflict, and repair relations, and a loop that selects units within a memory budget, chooses a presentation format per agent, and updates unit scores and relations from task outcomes. Across eight benchmarks it averages 81.11% versus 78.97% for the strongest of ten baselines, SAGE.

GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

Xinyu Che, Yunfei Ge, Shihao Li, Yanchen Liu, Hang Yan, Xinping Lei et al. cross-listed Existing game-development benchmarks replay fixed examples, score videos, or rely on a model judge, so they miss rule violations that occur mid-run in a game that still ends in a valid state. GameLogicBench offers 72 gameplay-logic tasks in Godot projects, with an automated evaluator that asserts the rules at every simulation tick across 403 hand-designed scenarios expanded into 1,451 seeded test cases, and that is validated to accept differing correct implementations while rejecting mutants with one required capability removed. Across 20 model and scaffold combinations, the best run solves only 52.78% of tasks, and under Claude Code all twelve models degrade as scope grows from isolated mechanics to repository-scale features. Evaluators built without mutant validation let incorrect submissions pass, and agents were found copying code from public repositories when network access was open.

GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

Rui Sun, Zhi Zheng, Zhenkun Wang, Zhichao Lu Skills give large language model (LLM) agents task-specific procedural guidance, but writing them as unstructured natural language leaves them redundant, short on workflow-level direction, and hard to optimize over a huge search space. The authors represent a skill as a graph whose nodes are execution steps with operational guidance and whose directed edges encode context-dependent transitions. GraphSkillEvo then runs population-based evolutionary optimization over these graphs with mutation and crossover operators. Across five agent benchmarks it beats the SkillOpt baseline, improving average accuracy by 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4.

TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization

Jiacheng Lin, Zifeng Wang, Zheng Chen, Erick Scott, Ziwei Yang, Fanyang Yu et al. Clinical development planning (CDP) and estimating a drug's probability of technical and regulatory success require experts across several disciplines to synthesize heterogeneous evidence, a slow and subjective process. TrialAtlas is a memory-augmented multi-agent system that coordinates specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning, and it learns from historical trials and prior New Drug Applications. The authors also introduce TrialAtlasBench, built from 291 FDA Complete Response Letters, where the system reaches a 50.0% F1 for detecting trial design deficiencies (6.1 points above the strongest baseline) and 85.3% balanced accuracy for predicting technical and regulatory success. In expert evaluation, 86.4% of its generated concerns were judged valid, versus 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.

AutoRecLab: Describe the Experiment, Get the Code!

Moritz Baumgart, Philipp Meister, Justus Krell, Michael Schmidt, Bela Gipp, Joeran Beel Turning a recommender-systems experiment design into executable code is manual and error-prone. AutoRecLab is a Python-based autonomous lab that takes a natural-language research idea, derives explicit experiment requirements, builds and validates a prototype, and iteratively expands it into the full experiment, combining retrieval-augmented generation for documentation lookup, static type verification, and execution-steered tree search. In a baseline comparison across six algorithms and three datasets, 8 of 9 runs succeeded at roughly $1 per run using GPT-5.4-mini, and the system also autonomously implemented an explicit-to-implicit feedback conversion study.

AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

Zijie Cao, Xijun Qu, Zhicheng Gu, Xiaoshu Chen, Duanyang Yuan, Yanning Hou et al. Long-term memory systems for large language model (LLM) agents usually store preferences, events, constraints, and temporal updates in one mixed representation with a fixed granularity or schema, which creates semantic interference and leaves relevant evidence poorly ranked in top-K retrieval. AutoViewMem discovers candidate semantic views from interaction traces, selects a compact set of low-overlap complementary views, and uses them to guide write-time extraction of structured, provenance-grounded memories, followed by offline consolidation for compactness and consistency. Because disentanglement moves from retrieval time to write time, plain top-K similarity search retrieves focused evidence without explicit routing or iterative retrieval. On the LoCoMo and PersonaMem benchmarks with Qwen3-8B and Qwen3-14B backbones, it improves long-horizon question answering and personalization over strong memory baselines.

Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents

Hafsa Akbar, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama cross-listed Large language model agents used in social simulation revise their opinions implicitly in context, so how persuadable an agent is can be neither specified nor verified, and collective outcomes inherit the model's training prior. Bayesian Chronicle Agents (BCA) add a minimal belief layer in which each stance is a probability updated by one Bayesian step per utterance heard, with a single prior-strength parameter κ that encodes stubbornness, modeled on Friedkin-Johnsen opinion dynamics. Sweeping κ produces consensus, persistent disagreement, or committed-minority influence on demand, with the persistent-disagreement regime matching the closed-form Friedkin-Johnsen fixed points at R² of 0.93 to 0.99. The prescribed κ stays recoverable after the round trip through language, with perfect rank-order recovery on all four models tested, and the explicit layer exposes per-model stance biases that end-to-end simulation would silently absorb.

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Shuai Bai, Jiayong Deng, Yikun Fu, Chang Gao, Xuhao Hu, Mianqiu Huang et al. Computer-use agents have developed along two separate lines, graphical interface control and coding through the command line, while real digital work interleaves both. RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web where an agent is given a running reference application, must discover its behavior, and must build a faithful reimplementation, with the reference serving as an oracle for hidden behavioral tests that yield execution-grounded rewards. Models trained on the generated trajectories improve on five out-of-distribution coding and hybrid computer-use benchmarks and verify their rendered outputs more often. On the held-out 250-task RecreationBench, GPT-6 Astra leads at 58.1% overall but passes all programmatic tests on only 2.8% of tasks, with agents reproducing static interface structure more reliably than interactions and computed outputs.

An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

Yiming Zhang, Jinghong Zhang, Haoran Zhao, Yiren Ma, Chunlei Zhao When a memory store contains conflicting positions, standard retrieval-augmented generation (RAG) injects the memories blindly, and in susceptible models this produces a markedly higher hallucination rate than a memory-free baseline. The Memory Decision Layer (MDL) is a zero-parameter controller placed between retrieval and generation that fuses relevance, reliability, and task-risk signals through QR-based orthogonal subspace projection into an interpretable trust score, decoupling confidence from consistency and allowing explicit abstention. Across mainstream large language models and several open-source datasets, it is reported to reduce the hallucination rate under conflicting memories by about 56% in general scenarios and to approach zero hallucination in high-risk scenarios. It uses only geometric operations and adds about 0.14 ms per decision, roughly 50 times faster than the embedding-retrieval step before it.

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Renkai Ma, Ruyuan Wan, Xuan Lu, Fan Yang, Chen Chen, Lingyao Li cross-listed Evaluations of autonomous AI agents usually measure task completion rather than the values users care about when delegating work. Using Value Sensitive Design and large language model assistance, the authors analyzed 73,093 first-person Reddit posts about using OpenClaw, coding each for its human value, agent aspect, value fulfillment, and user outcome, and identified 21 values in six groups including autonomous, dependable, and affordable operation, bounded reach, reviewability, and equitable access. Relative to corpus share, values clustered around the operating conditions users set for a run rather than around the agent's outputs. Values were usually met where users described what the agent delivered (five of six groups) and mostly unmet where they described supervising it (all six groups), a pattern the authors call value-sensitive delegation.

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv et al. Reinforcement learning for coding agents needs diverse tasks with reliable verifiers, and existing pipelines mine them from development artifacts such as issues and commits, which limits what can be extracted. CodeMidas is an agentic pipeline that uses source code as its only task-specific input: agents explore implemented functionality to write behavioral specifications, construct tests grounded in executing the original code, and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on it with GRPO improves all five benchmarks tested, including DeepSWE +11.7%, ProgramBench +17%, and Terminal-Bench v2.1 +8.5%, and the trained agent explores codebases more and self-verifies in more varied ways.

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Hongyang Du, Lan Yan, Christian Flores, Asim Kadav Professional graphic design is a long-horizon agentic task with no reliable programmatic oracle for judging outcomes. In Designer-RSI, a frozen frontier model operates professional design software through more than 230 tools while an external procedural memory of natural-language skills widens by acquiring procedures for recurring uncovered subtasks and deepens by revising procedures against their own successful and failed executions, with a matched replay gate admitting only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates or human labels, grow the skill bank from 76 to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3%, with 61.8% and 67.6% win rates against the no-skill agent across four design benchmarks on Claude-Sonnet-4 and Claude-Opus-4.6. On 200 held-out briefs, widening or deepening alone reaches about a 49% win rate over the no-skill agent, while combining them reaches 58.5%.
1 more specialized paper

Other 22

Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study

Erwei Wang, Ephrem Wu, Victor J. B. Jung, Jiajie Li, Andre Rosti, Joseph Melber et al. cross-listed Spatial neural processing units such as AMD XDNA place compute tiles beside small local memories and leave data movement to software, so mapping a multi-stage workload is largely a question of where intermediate tensors live. Using the open-source IRON and MLIR-AIR flows, the authors compare four FlashAttention designs on XDNA 1 and XDNA 2: per-operator execution, two on-chip streaming variants, and a fused kernel that keeps the attention scores in compute-tile local memory and reduces partial results over the cascade interconnect. On XDNA 2 the fused kernel reaches 3.62 TFLOP/s end to end, twice the IRON design, with 5.3 to 7.2 times the energy efficiency of the integrated GPU on the same chip at 2K tokens and above, covering twelve LLM configurations up to 128K tokens. Roofline analysis at each memory level shows the same fusion is nearly wasted on XDNA 1, which is already compute-bound with streaming, yielding the rule to fuse only until a mapping's operational intensity clears the ridge point; the reference designs are released as open source.

An Introduction to Compression-Based Machine Learning

John Hurwitz, Edward Raff, Charles K. Nicholas Any lossless compressor such as gzip can be turned into a learning method through Normalized Compression Distance or the Minimum Description Length principle, and any autoregressive model can be turned into a lossless compressor through entropy coding. The authors survey and formalize the strategies that exploit this relationship and introduce a design framework for compression-based machine learning, which they validate empirically. Compression-based methods come out competitive with conventional baselines and decisively stronger on malware classification, and varying the framework's design choices changes accuracy by up to 0.62.

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu Mixture-of-Experts (MoE) designs tie together three quantities that ideally would be set separately: how many experts contribute to a token's output, how many are actually computed, and how many expert-sized parameter sets must be built and stored. IntBMoE decouples them by having a lightweight hypernetwork merge all expert bases in a layer into composed experts drawn from a small learned codebook of blocks, while a router sends each token to only a few blocks; Dual-Path Residual Gating additionally couples two independently composed paths through multiplicative gating. It shows consistent gains over sparse and dense MoE baselines on image classification, with further experiments on language modeling and sequential recommendation. Deployed in AMap's generative recommendation system under a 60ms latency budget, it produced a 2.4% relative gain in click-through rate in online A/B testing.

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

Hritika Sharma, Thibault Ba\~neras-Roux, Alessandra Pinto, Petr Motlicek, Hyunggu Jung, Esa\'u Villatoro-Tello et al. Word Error Rate (WER), the standard Automatic Speech Recognition (ASR) metric, penalizes every lexical deviation equally regardless of whether meaning changes, raising the question of whether it tracks human judgments of transcript quality. The authors introduce HATS-en, an English dataset of human preferences over ASR transcripts, and benchmark lexical metrics against many configurations of BERTScore and SemDist that vary the language model, layer, and pooling. WER agrees least with human judgment of all metrics tested, the best SemDist configurations agree most, and no single embedding model wins everywhere. Character Error Rate (CER) stays remarkably close to the best semantic metrics, so the authors recommend CER as the primary low-cost metric with SemDist as a complement.

Neural Cellular Automata Learn General Features in their Hidden Channels

Etienne Guichard, Stefano Nichele Neural Cellular Automata (NCAs) are highly parameter-efficient models, but research has mostly examined their outputs rather than their internal hidden channels. The authors analyze those hidden-channel dynamics and introduce a transfer-learning mechanism that injects a pretrained teacher's hidden states into a student model to guide early optimization. On few-shot and scale-variant MNIST benchmarks, NCAs with about 9,800 parameters outperform comparable recurrent and feed-forward architectures, and the hidden channels are found to encode general, scale-invariant topological primitives rather than class-specific templates. This lets a student reach strong few-shot performance on unseen digit classes using features from a teacher trained only on digits 0-5.
17 more specialized papers

Robotics 21

Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models

Xiaoxiao Lu, Yunlong Dong, Jiahao Shi, Ye Yuan cross-listed World Action Models (WAMs) help robot policies by predicting how task-relevant scene states evolve, increasingly in a latent space, but the transitions are usually implemented with Transformers whose structure is built around token interaction rather than temporal evolution. The Latent Evolution Operator Network (LEON) instead models latent dynamics in a learned observable space using context-modulated operator propagation plus an additive forcing term, motivated by the controlled Koopman generator view of dynamics. Experiments on controlled dynamical systems confirm the intended inductive bias and the complementary roles of propagation and forcing. Across two WAM formulations, LEON improves closed-loop performance and robustness, even when it fully replaces the original transition module.

ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning

Mohsen Salehi, Karthik Pattabiraman Reinforcement learning controllers for unmanned aerial vehicles (UAVs) are vulnerable to action-space attacks, which overwrite action commands after the policy emits them and before the actuators execute them, and existing defenses mostly address input attacks or require retraining rather than handling corrupted actions at runtime. ASGARD is a two-phase teacher-student pipeline: in the teacher phase an encoder fuses the UAV's physical state with privileged attack information to produce an attack-aware latent that trains both the control policy and a monitor that outputs corrected commands; in the student phase the encoder and monitor are distilled via supervised learning to run on board using only physical state history. Across attack scenarios targeting different action commands, ASGARD completes missions despite the attacks, generalizes to unseen attacks, and remains resilient against stealthy ones.

Visual Navigation Transformer with Pose Attention

Beiming Li, Jaime Romero, Jonathan Diller, Vijay Kumar, Alejandro Ribeiro cross-listed Learned navigation policies usually consume observations as a time-ordered history, which makes it hard to reuse experience from earlier traversals without building an explicit map or topological graph. VNT-PA is a transformer planner whose context is a set of depth keyframes indexed by camera pose, so attention depends on pose differences rather than temporal order; it is trained to imitate a shortest-path planner running on the ground-truth scene mesh. On point-goal navigation in HM3D validation scenes it reaches 93.3% success and 90.4% success weighted by path length (SPL), beating baselines that encode the same context temporally or treat pose as an input feature, while also training faster. Because the context is an unordered pose-indexed set, frames from different trajectories can be fused at test time, and the planner degrades more gracefully under localization noise than a map-based baseline.

Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies

Zhipeng Tang, Xinda Chen, Weining Rao, Xiao Li, Wenting Tan, Yuning Wang et al. cross-listed Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert, yet more integration steps raise cost without reliably improving closed-loop success. Coda spends part of that budget on a single learned endpoint correction instead: a frozen policy runs a few-step noise-to-action trajectory, then a lightweight Transformer predicts a demonstration-supervised residual from the candidate action, the source noise, and the shared observation-prefix cache, with only the corrector trained. On 50 RoboTwin Easy tasks, five-step Coda raises success from 71.64% to 74.68% over the matched five-step baseline while cutting forward latency by 30.2% relative to the default ten-step policy, and a two-step configuration reaches 71.88% success at a 2.12x speedup. The same design lifts frozen official SmolVLA two-step success from 60.8% to 69.4%.

FootQuery: Future-Touchdown-Guided Retrieval from Depth History for Perceptive Humanoid Locomotion

Tao Dong, Jia Yu, Yuxuan Fan, Linna Zhao, Jiaqi Gong, Andong Yang et al. cross-listed Humanoids crossing stairs, gaps, and platforms often cannot see the spot where a foot is about to land, because of limited camera coverage and self-occlusion, so the relevant terrain must be recalled from earlier observations. FootQuery predicts each foot's next touchdown location and its uncertainty from proprioception, then uses those predictions to query sparsely sampled past depth frames. Retrieval is supervised during training by projecting realized contacts back into the historical images. The retrieved per-foot features are fused with a global visual memory to produce actions. In simulation the full system outperforms its ablations on the hardest stairs, gaps, and platforms. On a Unitree G1, a single policy continuously traverses outdoor stairs and indoor routes that combine stair ascent and descent, platforms, and gaps, using only proprioception and onboard depth images.

AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

Di Wu, Dongchen Zheng, Junhe Sheng, Zhongxing Wei, Songxin Zhang, Zejian Xie et al. cross-listed Robot demonstration data is scarce, so egocentric human video is an appealing pretraining source, but the embodiment and action-space gaps between humans and robots make it unclear how to use it. AtomEgo is a systematic study built on a curated corpus of about 2,659 hours and a scalable processing pipeline, comparing three paradigms across vision-language-action and world-action architectures: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video-action modeling. Multi-task real-robot experiments and cross-embodiment representation analysis suggest that capability gains scale with data volume multiplied by alignment quality, so egocentric data helps generalization only to the extent it is well aligned and utilized.

Adaptive Rollout Truncation Based on Epistemic Uncertainty for Efficient Offline World Model Training

Nikodem Sebastian Zymla, Laurin Thiele, Johannes Pitz cross-listed Multi-step autoregressive training improves the long-horizon accuracy of neural world models for robotics, but a fixed rollout length is costly and can amplify errors while the model is still inaccurate. The proposed auto-curriculum terminates each training rollout once epistemic uncertainty exceeds a threshold calibrated in a two-stage warm-up, using either a five-head ensemble on a shared recurrent backbone or Monte Carlo Dropout as the estimator. On ANYmal-D and ANT, ensemble-based truncation matches or improves the prediction accuracy of fixed-horizon training and the RWM-U baseline, reaching comparable final performance on ANYmal-D with roughly 72% less rollout computation.

Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation

Xinyu Liu, G\"okhan Solak, Arash Ajoudani cross-listed Model-free reinforcement learning for contact-rich manipulation usually makes the policy learn both task strategy and low-level motion generation through direct Cartesian commands. PA-RL has the policy adapt the parameters of an artificial potential field, which produces a state-dependent guidance direction executed by a Cartesian impedance controller. On simulated peg-in-hole insertion against velocity, pose, and variable-impedance action spaces under the same algorithm, it is the only method to reach a 100% evaluation success rate within the training budget, versus 92.6% for the best baseline, while cutting joint-torque variation by 55.4% and Cartesian acceleration variation by 70.8% without motion-quality penalties in the reward. The simulation-trained policy completed 9 of 9 real-robot insertions without fine-tuning.

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

Hiroaki Kingetsu, Hiroaki Kurihara, Kaoru Yokoo, Kenji Fukumizu, Manohar Kaul cross-listed Reinforcement learning with sparse binary rewards stalls when a Vision-Language-Action (VLA) policy never samples a success, and the usual fix is costly human teleoperation data. SynthDemo-RL uses an automated teacher that turns simulator-privileged state into successful manipulation trajectories, distills a VLA student from them with supervised fine-tuning, and then refines it with PPO on binary success rewards. On LIBERO-PRO, where a pi_0.5 policy scores exactly 0% on 27 of 57 tasks, direct PPO rescues only 10, while SynthDemo-RL rescues all 27 with 50 synthesized trajectories per task and reaches about 97% average success with no new human demonstrations. On standard LIBERO it reaches 96.0%, within 1.7 points of training on human demonstrations, with further validation on RoboTwin 2.0 and an open-loop physical robot test.

ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction

Tim Engelbracht, Ren\'e Zurbr\"ugg, Mayank Mittal, Marco Hutter, Marc Pollefeys, Hermann Blum et al. cross-listed Digital twins of articulated objects such as doors usually capture kinematics only, or assign static physical parameters from visual and language priors, missing state-dependent effects from springs, friction, and door closers. ForceTwin has a person probe the object with a handheld force-sensing gripper, then estimates articulation, inertia, Coulomb friction, viscous damping, and a structured neural residual for mechanism forces from the synchronized poses and forces. It nearly halves the inertial-parameter error of a vision-language model (VLM) prior. Used as a feedforward dynamics model for impedance control on a Spot and a Franka FR3, it reaches 87% goal completion across nine object-embodiment pairs versus 60% for VLM-prior twins and 57% for kinematics-only twins, and the twins were also used to train door-traversal policies deployed in the real world.

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

Sichang Su, Benjamin Yang, Zhiyun Deng, Boyuan Liang, Yip Fun Yeung, Zelin Wang et al. cross-listed Pretrained robot foundation policies often complete most of a long-horizon task but fail repeatedly at a few critical subtasks, and collecting more full-task demonstrations wastes operator effort on behavior that already works. PARTS (Policy Adaptation with RL on Targeted Subtasks) keeps the pretrained policy frozen for nominal actions while agent-generated selectors and success verifiers activate residual corrections and supply local rewards at those bottlenecks. Training combines online reinforcement learning with success-reweighted retraining, with humans only identifying bottlenecks and performing physical resets. It raises complete-task success from 32% to 61% on bimanual YAM tasks and from 50% to 95% on single-arm Franka tasks using tens of minutes of real-world rollouts per task, beating existing real-world RL fine-tuning methods by more than 25% under the same rollout budget.

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence

Eshika Pathak, Leela Krishna cross-listed When a robot fails at a task it must decide whether to act on its own diagnosis, consult another onboard sensor, or interrupt a person, which depends on how much its sensors reveal and how reliable its diagnosis is. The authors build a simulated benchmark with injected, known failure causes and find that some failures are diagnosable only from force data (0.99 accuracy from force versus at most 0.55 from images), then test six open vision-language models. The models track the surface of the prompt rather than the evidence: moving the refusal option from last to first drops refusal rates from 78-100% to 0-6% in three of six cases, accuracy from camera frames never beats a majority-class baseline, and stated confidence carries no information about correctness. Supplying the force data as ten lines of text produces the first above-baseline diagnoses in four of six models, and a single question to a human raises accuracy to roughly the answerer's reliability (0.70-0.81), so the authors argue the decision to ask should be tied to measured accuracy and stated costs rather than model confidence.

Benchmarking World Models for Continual Learning on Compositional Tasks

Haoyu Zhou, Joe Watson, Anson Lei, Ingmar Posner Measuring how well a world model adapts to new tasks entangles two abilities: learning unseen content quickly and reusing knowledge already acquired. To isolate reuse, the authors propose a continual learning benchmark for world models in robot manipulation, where each task curriculum includes compositional tasks combining aspects of earlier tasks, factorised along action and perception axes to show how each input modality bottlenecks reuse. They evaluate state-of-the-art world models under canonical continual learning methods alongside a modular world model whose dynamics backbone has explicitly reusable components. Modularity balances reuse against forgetting better than conventional methods, but no approach solves the problem fully.
8 more specialized papers

Safety & Alignment 20

From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News

Zeynep \"Ozdemir, Murat Osmano\u{g}lu, Sevgi Yi\u{g}it-Sert, \"Omer \"Ozg\"ur Tanr{\i}\"over, Y{\i}lmaz Ar How large language models both produce and detect fake news is examined under four controlled scenarios: open-ended generation, rewriting, manipulation prompts, and attribute-based prompts grounded in a journalistic discourse framework. Seven widely used models generated a synthetic corpus of 14,000 articles, whose linguistic properties were compared with real news, and each model was then asked to judge the generated articles using a basic prompt and prompts refined iteratively from misleading patterns in real-fake pairs. Generation and detection ability vary substantially across models, and the generation strategy strongly affects detectability. Notably, the refined detection prompts do not improve and often harm detection performance.

$\mu^2$-Bench: A Multilingual Machine Unlearning Benchmark

Kyomin Hwang, Hyeonjin Kim, Hyunho Lee, Yearim Kim, Yeji Song, Nojun Kwak Unwanted information such as harmful content or private data can spread across languages inside multilingual large language models, and it is unclear whether unlearning methods remove it in every language. μ²-Bench is a Multilingual Machine Unlearning (MMU) benchmark that simulates the full memorization, unlearning, and evaluation pipeline across a broad set of languages, testing both languages used in training and held-out ones, and assessing knowledge dispersed across multiple languages. The authors report that successful multilingual unlearning requires methods that explicitly account for multilingual characteristics, and provide further analysis of how unlearning behaves across languages.

Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

Utkarsh Agarwal, Monojit Choudhury To probe how language models weigh conflicting moral values across languages, the authors build a 12,000-instance dataset of two-option dilemmas covering Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, translated into Hindi, Arabic, Spanish, and Chinese. GPT-5-mini consistently favors Honesty over Autonomy in all five languages when given no policy, while Llama-3.2 1B and 3B models show a strong first-option bias that both plain fine-tuning and Direct Preference Optimization remove, raising accuracy above 98%. To separate learned value preferences from dataset correlations, they orthogonalize a value-preference task vector against a general instruction-following vector, and show the isolated direction can be used in task arithmetic to produce a model with the opposite stance.

Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

Shuo Huang, Gholamreza Haffari, Xingliang Yuan, Ting Yu, Lizhen Qu cross-listed Attackers who combine large language models (LLMs) with auxiliary knowledge can link released text to individuals, yet existing privacy audits report only attack success rates without finite-sample statistical guarantees. Conformal Privacy Auditing (CPA) is a distribution-free calibration framework that outputs, for each released document, a set of candidate identities guaranteed to contain the true one at a user-chosen confidence level under exchangeability, with the set's size serving as an interpretable leakage proxy. It supports both logit-access and sampling-only attackers, so open-source and proprietary API models can be audited in the same way. Across several release benchmarks it achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary.

Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models

Yining Wang, Xi Li, Mi Zhang, Xiaohan Zhang, Xiaoyu You, Zhenxing Qian et al. cross-listed Multimodal large reasoning models (MLRMs) can infer where a casually shared photo was taken by reasoning over cues such as architecture, vegetation, and lighting, which creates a geolocation privacy risk. The authors show that refusal-based safeguards are insufficient, since crafted jailbreak prompts raise model response rates to 100%, and that existing pixel-space perturbation defenses transfer poorly to black-box models and leave visible artifacts. Their defense instead injects perturbations into the latent space of a diffusion model during reverse sampling, using GeoCLIP, a model aligned with GPS coordinates, as a surrogate to locate and disrupt the geographic signals. They report significantly stronger black-box transferability while preserving perceptual image quality.

HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference

Byeongseo Min, Yongwoo Lee, Young-Sik Kim, Yongjune Kim cross-listed Homomorphic encryption (HE) lets a server run a large language model on encrypted inputs without seeing the plaintext, but that same confidentiality means the server cannot inspect prompts or responses, so jailbreak attacks by malicious clients go entirely unnoticed. HE-Guardrail evaluates guardrail mechanisms wholly over encrypted data and homomorphically controls whether the target model's response is returned to the client. Instantiated with Llama Guard, JBShield, and GradSafe, it closely reproduces the decisions of the corresponding plaintext guardrails in the encrypted domain, with each option showing a different trade-off among security, efficiency, and utility.

ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor

Dominik Dahlem, Rui Vieira cross-listed Third-party adapters for open-weight language models ship as opaque weight matrices, and the authors argue that detection fails for backdoors hidden in the subspace where a safety monitor is structurally blind, since honest and backdoored adapters overlap on every blind-subspace statistic evaluated. ServeGuard instead makes that channel structurally absent: the publisher constructs the adapter to read the input only through directions the monitor covers and proves this in zero knowledge, without revealing the certified weights. The proof is cheap because the monitor's blind spot is a deterministic function of the public base model, so only one linear identity needs proving, and an admission-time guard binds the guarantee to the adapter bytes actually served. Across eight checkpoints up to 7B from four families, the monitoring budget turns out to be architectural, and on a 0.5B model confinement is nearly free for benign adaptation, making monitor quality the main security lever.

Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems

Pedro Pereira, Eva Maia, Isabel Pra\c{c}a cross-listed Retrieval-Augmented Generation (RAG) systems can be poisoned through their knowledge sources, and defenses usually assume a single malicious passage. Micro-Collaborative Poisoning instead splits a false target claim across several individually plausible documents, and is evaluated over 108 RAG configurations varying dataset, retriever, retrieval depth, database composition, number of poisoned databases, and generator model. The attack's effect comes from accumulated weak adversarial signals across retrieved sources rather than any single dominant passage, so larger top-k and more poisoned databases strengthen it, while clean database diversity and stronger retrievers weaken it. Document-level analysis shows it leaves a weaker explicit poisoning signature than direct poisoning, making it hard to catch by inspecting documents in isolation.

GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation

Zeyu Yan, Guanghao Zhou, Minghui Qiu, Ming Gao, Cen Chen Machine unlearning is harder for large reasoning models (LRMs) because protected facts or unsafe rationales can leak in the intermediate chain-of-thought (CoT) before the final answer, and existing objectives suppress content without specifying what the model should produce instead, leading to hallucinated substitutes or degenerate output. Guided Answer-Reasoning Distillation (GUARD) converts the model's own unsafe disclosures into safe-exit trajectories, a coherent non-disclosing CoT followed by a stable refusal, aligns a frozen LRM using guidance tokens, and distills that behavior into the weights. A new metric, Natural Forgetting Reasoning Score (NFRS), measures structural stability, fluency, and unsupported substitutes in forgotten outputs. On R-TOFU and a STAR-1-derived harmful-intent setting, GUARD substantially reduces unsafe and privacy disclosures on two distilled LRMs while preserving reasoning utility.

CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang et al. cross-listed Privacy leakage in LLM agents is usually measured inside single components such as memory, retrieval, or tool use, which blurs the difference between internal exposure and what an outside attacker can actually recover. CIPL (Channel Inversion for Privacy Leakage) is a black-box evaluation framework that models a target as sensitive source, selection, assembly, execution, observation, and extraction stages, and measures the transition from selected sensitive units to attacker-recoverable output under one protocol. Experiments across memory-based, retrieval-mediated, and tool-mediated targets plus a BrowserUse live-agent case study show that storage labels alone do not determine recoverability: memory leakage is near-saturated, retrieval leakage is often partial, and tool and live-agent leakage varies with observation surface, prompt-to-channel alignment, retrieval depth, and provider behavior. A semantic audit also finds attacker-useful disclosures that exact matching misses.

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Jiale Luo, Eric Han cross-listed Jailbreak defenses for large language models (LLMs) sit at different pipeline stages, such as input modification or output guarding, but prior studies evaluated them mostly in isolation and under inconsistent attack-success-rate definitions. CASCADE studies defense combinations within and across stages under one threat model of direct, black-box, single-turn attacks, using a standardized attack-success-rate formulation with controlled query budgets and explicit fairness rules. Across 19 attacks and 15 defenses, no single defense is universally best, but well-chosen combinations deliver substantial safety with minimal utility loss. The authors turn these results into practical recommendations for layered defense pipelines.

End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery

Akira Ito, Takayuki Miura, Yosuke Todo cross-listed Hard-label model extraction attacks try to recover a neural network's parameters from oracle queries that reveal only the final output label, and the polynomial-time attack on ReLU multilayer perceptrons by Carlini et al. (Eurocrypt 2025) contains a sign-recovery step too query- and compute-heavy to run in a fully black-box setting. The authors propose a new sign-recovery algorithm based on a different principle that needs no dedicated queries and achieves higher sign-recovery accuracy than the existing method in their experiments. With it, every step of the attack can be implemented in a black-box setting, enabling an end-to-end demonstration on trained models. On networks trained on MNIST and Fashion-MNIST with width 16 and 4 or 6 hidden layers, the extracted models achieve over 98% label agreement.

Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment

Maciej Skorski Annotator disagreement on moral content is usually collapsed by majority vote or by an any-annotator rule that marks an item positive as soon as one annotator flags it. Moral Entropy is a Bayesian framework that keeps a full posterior over the true label and splits its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (insufficient or noisy annotation), which lets any consensus rule be audited against a calibrated reference using cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items, almost entirely false positives when pooled, though the errors invert at the level of individual moral foundations (19.9% mean false-positive and 38.9% mean false-negative rate on MFTC). The stricter majority and two-vote rules instead miss 63-83% of true positives.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

Hiskias Dingeto A language model may hold knowledge it does not report, for example by sandbagging on a capability evaluation, and its outputs alone cannot show whether it is hiding an answer or lacks one. Borrowing the forensic Concealed Information Test, Probe of Internal Recognition (PIR) presents a question with candidate answers and reads from the model's internal states which candidate it recognizes as correct, without needing an honest reference model or a labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, against a 0.28 to 0.40 unknown-item baseline and 0.25 chance, and recognition stays between 0.85 and 0.93 under prompted deception, trained sandbagging, password-locked checkpoints, and circuit-broken checkpoints. When unlearning actually removes the knowledge, recognition falls to the level of a never-known question, so the probe separates a model that will not answer from one that cannot; the signal is reported to be causal, to add information beyond black-box cues, and to extend to free-form generation.

Available Guardrails: Certifying Selective Prediction across ML Systems

Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky A selective predictor acts as a safety gate that returns an output only when it appears trustworthy, and deployments increasingly need that reliability certified at a target precision for every reporting unit, such as a tool, policy label, or patient subgroup. The authors focus on availability, meaning whether finite calibration data can produce a certificate at all, make it computable through exact-binomial inversion, and cast the choice of reporting partition as a dynamic program that trades off safety, granularity, and served traffic. A truth-informed planner gains 0.157 mean coverage over support balancing while a naive estimator recovers only 0.005; building candidate partitions on one data split and selecting on another recovers 0.060, with the direction reproduced in 59 of 60 model effects across three intent-routing datasets and two architectures. Reallocating the familywise error budget across units recovers further coverage, and the same frontier appears in large language model tool-calling, content moderation, lesion classification, and recommendation.
5 more specialized papers

Unclassified 20

COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training

Qifeng Cai, Xuanguang Pan, Hao Liang, Chang Xu, Wentao Zhang No summary available — see the abstract on arXiv.

VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

Jinke Wu, Zhengpin Li, Mengzhe Jia, Yang Li, Wentao Zhang No summary available — see the abstract on arXiv.

Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces

Zihan Wang, Hao Wang, Boyuan Jiang, Yiqun Zhang, Shi Feng, Xiaocui Yang et al. No summary available — see the abstract on arXiv.

Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders

Yasir Mehmood, Kashif Javed No summary available — see the abstract on arXiv.

Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models

Polina Tsvilodub, Max H\"oth, Michael Franke, Bj\"orn Deiseroth, Carina Kauf No summary available — see the abstract on arXiv.

Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media

Cris Huynh, Arlene Pham No summary available — see the abstract on arXiv.

Enhancing Audio Reasoning via Semantic Summary Prediction

Francesco Bonzi, Pooneh Mousavi, Cem Subakan, Mirco Ravanelli No summary available — see the abstract on arXiv.

MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

Yueming Lyu, Yilian Shi, Haoxiang Tan, Linzhuang Zou, Qihao Wang, Guihua Yu et al. No summary available — see the abstract on arXiv.

Reconstruction of 4D Mitral Regurgitation Hemodynamics from Sparse Planar Data using Deep Operator Networks with Test-Time Adaptation

Jakob Marcel Hoffmann, Yosuke Hasegawa, Alexander Stroh cross-listed No summary available — see the abstract on arXiv.

Automated Physics-Informed Neural-Networks-Based Calibration of Highly Segmented Silicon Telescopes

M. Rejmund, A. Lemasson, P. Morfouace, D. Ramos, J. Taieb, J. D. Frankland cross-listed No summary available — see the abstract on arXiv.

The Refutation Gap: Certifying Both Halves of an Optimality Claim

Rohan Pandey cross-listed No summary available — see the abstract on arXiv.

Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation

Simon Pals, Cristian Tejedor-Garcia cross-listed No summary available — see the abstract on arXiv.

Reinforcement learning for post-coronagraphic wavefront control

Manuela Casta\~neda-Medina (LIRA), Yann Gutierrez (LIRA), Johan Mazoyer (LIRA, CNRS), Baptiste Abeloos, Laurent Mugnier et al. cross-listed No summary available — see the abstract on arXiv.

Sparse Priors for Efficient Distribution Learning

Saumya Goyal, Barnab\'as P\'oczos No summary available — see the abstract on arXiv.

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

Chuxuan Hu, Yeye He, Penny Zhou, Wee Hyong Tok, Daniel Kang, Surajit Chaudhuri No summary available — see the abstract on arXiv.

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Themistoklis Haris, Henry Li, Maryam Karimzadehgan No summary available — see the abstract on arXiv.

Extreme classification: beating chance with one training example from each class

Kevin Bleakley (LMO, CELESTE), Aaditya Ramdas cross-listed No summary available — see the abstract on arXiv.

SpaceDiffusion: Over-the-Orbit Diffusion for Space Generate-and-Forward Communications

Jianhao Huang, Zhanwei Wang, Khaled B. Letaief, Kaibin Huang cross-listed No summary available — see the abstract on arXiv.

Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

Runze Hu, Jingqi Kong, Yang Yang, Yihang Yang, Jingyao Liu, Haizhou Tang et al. No summary available — see the abstract on arXiv.

Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces

Boyuan Zhao, Sifan Zhang, Luping Chen No summary available — see the abstract on arXiv.

Multimodal 18

PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

Mengxuan Li, Junfa Chen, Jinze Xia, Yundan Chen, Lixin Fan, Ke Liu et al. Foundation models for physiological signals usually need task-specific adaptation, and it has been unclear how well current models can follow natural-language instructions across signal types. PhysioBench harmonizes annotations from 22 public datasets into 61.4 million questions spanning 30 tasks, each grounded in a signal segment and traceable to its source annotation. An evaluation of 21 models, including large language models, vision-language models, time-series language models and physiological signal foundation models, under three settings finds that none achieves consistently strong performance across modalities and tasks. Natural language does enable unified prediction across tasks, though results remain sensitive to how questions are phrased.

I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance

Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang, Seungwhan Moon, Shashank Jain, Pinar Donmez et al. cross-listed Audio large language models only answer when queried, which is a poor fit for a wearable meant to monitor a sound stream and speak up unprompted — a use case motivated by Deaf and Hard of Hearing users. ISM (Interrupt and Silent Modeling) is a model-agnostic scheme that embeds the proactive decision into decoding through two special tokens, <interrupt> and <silent>, covering onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication from a single natural-language statement of intent. Applied to Qwen2-Audio-7B it reaches 99.6% interrupt F1 on ESC-50 with perfect de-duplication recall, leads on noisy Epic-Sounds kitchen audio without domain-specific training, and streams with 3.5-second average latency.

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

Seoyeon An, Hyeonseo Jang, Minsu Kim, Chanho Lee, Younghan Park, Kangwook Lee cross-listed Existing video benchmarks for Multimodal Large Language Models (MLLMs) mostly pose scene-level or summary questions that can be answered in a single inference step. AgentVidBench is a multi-hop video question answering benchmark that targets spatial, temporal, and causal reasoning. It includes step-by-step solution traces, so evaluation can check whether an agent actually gathered the evidence supporting its answer. Across 12 proprietary and open-source MLLMs, single-turn performance remains limited, while wrapping the same models in agentic workflows generally improves both accuracy and trajectory scores. The authors also provide a simple agentic baseline and release code and data.

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Qi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu, Yuxuan Wang et al. Voice-and-vision assistants must first work out what a user actually wants. That demand is often underspecified in speech, buried in visual or dialogue context, obscured by disfluency or noise, or not addressed to the assistant at all. Omni Demand Understanding (ODU) frames the task as detecting whether a demand is present and inferring its intent from an interaction stream. ODU-Bench evaluates this along five dimensions over single- and multi-turn interactions, and is built from taxonomy-guided agentic video generation plus human recordings with human-verified annotation. Among 14 native multimodal large language models, the strongest, Gemini 3.1 Pro, recovers only 44.7% of the key information that must be inferred from visual, acoustic, or conversational context, and 11 of the 14 models false-trigger on more than 50% of non-demand scenarios.

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng, Muzhi Zhu et al. cross-listed Native audio-visual dialogue is defined here as an omni model receiving a user's audio and video directly and replying in text, with no separate text question, captioning, or speech recognition step. Because real recordings are scarce and good replies are too varied for keyword matching, the authors build OmniVChat-Studio, a multi-agent engine that synthesizes single- and multi-turn dialogues, and use it to create OmniVChat-Bench, covering five ability categories. They also propose OmniVChat-RL, a reinforcement learning reward that jointly targets reply correctness, efficiency, and style. Training Qwen3-Omni-Instruct with it on synthesized dialogues improves results on both the synthetic benchmark and the human-recorded OmniVChat-Bench-Human, indicating transfer to real dialogues.

PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

Zicheng Zhao, Dongyin Chen, Rui Xu, Yinghui Xu Existing benchmarks do not test whether multimodal large language models (MLLMs) can design a complete load-bearing structure that actually works when simulated, or repair it after a failure. PolyBridgeBench gives a model a visual scene plus structured engineering constraints and asks for a full node-member-material bridge topology, gating execution in a dynamic physics simulation behind deterministic legality checks. After a failed run, the model receives temporal visual evidence from the rollout and attempts a repair under a fixed interaction budget, with validity, dynamic success, and recovery measured separately. Experiments with six MLLMs across 189 levels reveal a substantial gap between producing a valid design and one that succeeds in simulation, along with strong sensitivity to material budgets and limited post-failure recovery.

VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim cross-listed Benchmarks for Video Large Language Models (Video-LLMs) often rely on question answering or caption matching, which models can pass through superficial cues and incomplete annotations. VidOmni-Bench instead asks models to verify whether each event in a dense video caption is actually supported by the video, using 500 videos spanning five complexity types and durations from 4 seconds to 90 minutes. Captions are generated by diverse Video-LLMs and labeled at the sentence level by humans, so sentences with incorrect events serve as hard negatives. Experiments show that Video-LLMs frequently hallucinate in dense captioning and also fail as verifiers of plausible but incorrect event descriptions, with weaknesses that vary by model across video complexity and duration.

Samsone: A Family of Open Small Audio Language Models for On-Device Inference

Piotr Masztalski, Micha{\l} K. Grzeszczyk, Olaf Sikorski cross-listed Large audio language models have grown past billions of parameters, while privacy and latency needs favor Small Audio Language Models (SALMs) that run on-device. Samsone is a family of such models built for edge hardware, with a core Samsone-134M plus Samsone-99M and Samsone-356M variants used to study scaling behavior at small sizes. The authors claim Samsone-134M sets a new state of the art for its size class across multiple benchmarks and that the family is competitive with models orders of magnitude larger. Training uses only public data, and the release includes training code, weights, mobile-optimized checkpoints, and an open-source Android app demonstrating real-time on-device inference.

ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction

Jinning Liang, Mingcheng Zhu, Tingting Zhu Vision-language models (VLMs) used for emergency department prediction can score well without actually depending on the patient's electrocardiogram (ECG), a failure the authors call ECG Mirage, split into ECG neglect and ECG confusion. They test for it by comparing predictions with matched ECGs, outcome-discordant mismatched ECGs, and no image, while holding clinical text and targets fixed. Across four VLMs on MDS-ED, matched ECGs give no consistent advantage for predicting ICU admission or clinical deterioration. Training four restricted visual prompts with supervised learning followed by conditional direct preference optimisation, with the backbone frozen, yields balanced accuracies of 70.6% and 67.5% and widens the matched-versus-mismatched gap to about 16.5 and 5.5 percentage points.

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen, Zhehuai Chen et al. NemotronLabs VoiceChat is an open full-duplex speech-to-speech model that can call tools natively while listening and speaking at the same time. It combines a streaming speech encoder and decoder-only language model with parallel output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming text-to-speech decoder. On Full-Duplex-Bench 1.0 it has the lowest pause-handling takeover rates among evaluated open-weight systems, 100% takeover after user interruptions, and a 4.33/5 post-interruption response-quality score, and on Full-Duplex-Bench 1.5 it resumes after user backchannels in 93% of cases. It scores a 55.1 normalized average on VoiceBench and 82.5% tool-selection F1 on Full-Duplex-Bench 3.0, while argument accuracy and end-to-end tool execution remain areas for improvement.
8 more specialized papers

Vision 14

Physically Based Rendering in the Latent Space

Vuk Radovanovic, Vishesh Gupta, Adrien Gruson, Binh-Son Hua cross-listed Image diffusion models generate impressive images but are hard to control compared with classical physically based rendering pipelines. Observing a link between light transport and the distribution of latent values, the authors run light transport simulation directly in the feature space of the variational autoencoder used by generative models, modifying the rendering equation and using a differentiable renderer to recover scene parameters that render into the pretrained latent space with minimal refinement. Trained on a single rendered image, the method is shown to generalize to changes in scene geometry, lighting, and camera viewpoint.

Detection is solved, delineation is not: what governs tooth segmentation on panoramic radiographs

Muhammad Rehan, Moaz Amjad, Syed Danial Ahmed, Mariam Adnan, Haider Ali cross-listed Using 1,422 panoramic radiographs with 42,142 expert-delineated tooth polygons across the 32-class FDI numbering scheme, the authors isolate the effects of input resolution, architecture, and anatomical priors on tooth segmentation under one evaluation protocol. Resolution dominates: raising input size from 640 to 1280 lifts mask mAP50-95 from 0.656 to 0.717 while mAP50 stays flat at about 0.982, meaning added resolution buys boundary precision rather than detection. A query-based transformer with 2.1x the parameters is statistically equivalent to a one-stage detector in-domain and 5.5x slower on CPU, and three targeted interventions fail, including a LoRA-adapted self-supervised encoder and a promptable foundation segmenter that degrades masks by 39%. Zero-shot transfer to an independent multi-centre cohort costs 62% of mask mAP50-95 but only 18% of mAP50, with residual error concentrated in the apical third of the tooth.

The Weight Is Over - Interactive Diffusion on Consumer GPUs

Frieder Ganz, Maximilian M\"uller On-device inference work has focused mostly on language models, while diffusion pipelines remain memory-hungry, latency-sensitive, and harder to orchestrate because they chain an embedder, a transformer, a decoder, and postprocessing. The authors contribute an embedding translator that maps a small text encoder into a large encoder's space to cut weight and latency, a reproducible sweep recipe for navigating the speed, quality, and memory trade-off, and an interactive on-device image generation editor. The editor achieves sub-second time to first image on recent consumer GPUs.
11 more specialized papers

Reinforcement Learning 13

Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

Jonathan Hau, Alessandro Abate Synthesizing policies that satisfy Linear Temporal Logic (LTL) specifications, such as safety or reachability, in unknown environments is addressed with an end-to-end model-based reinforcement learning algorithm. The task is represented as a Limit-Deterministic Büchi Automaton synchronised with a Bayes-Adaptive Markov Decision Process model of the environment, and a new variant of Bayes-Adaptive Monte-Carlo Planning (BAMCP) computes approximately Bayes-optimal strategies over the combined structure. Across finite- and infinite-horizon tasks the approach shows better property satisfaction and sample efficiency than traditional model-free methods, with ablations favoring the new planner over classical BAMCP. The authors also apply it to cautious reinforcement learning, reducing the number of task violations incurred during training.

REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement

Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs Continuous-control policies in deep reinforcement learning usually map a state to an action in one forward pass, leaving no room to revise an initial prediction. Iterative Action Refinement (IAR) instead builds the action through a sequence of learned residual corrections, with a shared refinement network repeatedly conditioning on the state and the current action proposal; combining it with Proximal Policy Optimization (PPO) yields REFINEPPO. Across 14 benchmark control tasks, with ablations of refinement depth and update schedules, REFINEPPO matches or exceeds standard PPO and converges faster on several tasks.

Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation

Ana Nunez, Peyman Najafirad Co-training one language model as both coder and test author promises to move code-generation reinforcement learning past fixed test suites, but pass-rate rewards can be maximized with trivial non-discriminative tests, and independently sampled tests cluster on typical inputs and inflate estimator variance. CoVer tackles both inside a single-policy GRPO setup: an information-gain reward scores each self-written test by the mutual information between its pass/fail vector and a graded, ground-truth-anchored correctness signal, paid out only when the covariance sign shows the test discriminates in the right direction, and a three-stage diversity filter prunes candidates to a behaviorally non-redundant suite at fixed execution budget. Across LiveBench, MBPP, LiveCodeBench, CodeContests, and CodeForces, one-shot pass@1 rises by 5.8 points at 7B and 7.1 points at 14B over the Qwen2.5-Instruct backbone, and dropping CoVer-7B into the CodeT ranking pipeline adds a further 3.5 points.

Deep Reinforcement Learning with Buffered Quantile Objectives

Mohammad Alipour-vaezi, Sajad Khodadadian Risk-sensitive reinforcement learning that optimizes a chosen quantile of the return distribution is difficult because point quantiles shift abruptly under small perturbations, and existing methods based on smoother buffered quantiles are model-based and limited to small tabular problems. Deep-BQRL is a model-free distributional method that learns conditional return quantiles from sampled transitions with neural networks, scores actions by averaging the relevant region of the learned quantile function, and uses ensemble disagreement to guide exploration. On an asset-selling optimal-stopping problem and slippery FrozenLake, it attains smaller point-quantile policy gaps than PPO and TRPO, although the model-based UCB-BQRL retains the smallest gaps. The learned stopping decisions vary with the target quantile, illustrating the method's risk-sensitive behavior.

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

Qiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang et al. Reinforcement learning for large language model (LLM) agents is hard to apply to open-ended tasks, where no reliable scalar reward exists. Existing pairwise-preference methods still compress comparative feedback into a single trajectory-level reward. ArenaFlow ranks trajectories through tournaments and attaches a structured reflective evaluation to each comparison, which identifies pivotal success steps, reusable strategy skills, and which retrieved skills were actually used. Trajectory-level advantages are propagated to high-confidence pivotal steps according to how deep a trajectory survives in the tournament, while a global skill memory is updated, pruned, and retrieved by estimated utility to serve as a policy prior for later exploration. The authors report gains on open-ended agent tasks, though the abstract gives no specific numbers.

GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

Kaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang, Yang Song, Dingqian Hong et al. Post-training methods for large language models such as Group Relative Policy Optimization (GRPO) can be unstable because they rely on importance sampling. Group Variance Policy Optimization (GVPO) builds the analytical solution of KL-constrained reward maximization into its gradient weights, so that its gradient matches that of a mean squared error between centered implicit rewards and centered actual rewards. The authors show that it has a unique optimum that exactly solves the KL-constrained objective and allows flexible sampling distributions without importance sampling. They extend it to on-policy distillation (OPD) and to a broader family of OPD objectives. The abstract reports no empirical numbers.

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

Xuening Wu, Yanlan Kang, Shenqin Yin Adding preference signals to a reinforcement learning reward can unintentionally change the task being optimized. IncentRL introduces preference guidance as a Kullback-Leibler (KL) penalty between a specified outcome distribution and a preferred distribution, and for finite discounted Markov decision processes it derives a bound on how much external-task value can change, a sufficient action-gap condition for preserving the original optimal policy, and a characterization of the large-weight regime. In a practical implementation with a hand-designed distance-based outcome proxy, the three-seed mean success rate on MiniGrid DoorKey-8x8 after two million steps reaches 98% with coefficient 0.01 versus 90.5% for the zero-coefficient baseline. The authors note the experiments are descriptive and do not yet isolate KL shaping from simpler alternatives.

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang et al. Auto-bidding under the optimized cost-per-X (oCPX) advertising paradigm is usually served by a separate model per scenario, such as registration or purchase, which fragments pipelines and ignores cross-scenario structure. OneBid trains one Decision Transformer-based backbone on heterogeneous bidding logs, conditioning on both Return-to-Go for conversion value and Cost-to-Go for cost ratio, with a sequence-level Mixture-of-Experts that combines shared and sparsely routed experts to stay within latency limits. Scenario adaptation uses Critic-guided Relative Offline Policy optimization (CROP), in which a learned critic scores candidate actions group-relatively so no unsafe online exploration is needed. Fully deployed at Kuaishou, it delivers an overall +2.2% gain in advertiser value in online A/B tests, peaking at +13.1% in the return-on-ad-spend scenario.

CityLearn v3: A Configurable Simulation and Evaluation Framework for Realistic Control Studies of Renewable Energy Communities

Tiago Fonseca, Luis Lino Ferreira, Armando Sousa, Ava Mohammadi, Zoltan Nagy cross-listed Controller studies for renewable energy communities often simplify away changing membership, equipment availability, service deadlines, and data quality, so a lower cost or peak demand can hide missed services or infeasible power requests. CityLearn v3 is a configurable simulation and evaluation framework that models changing members and assets, flexible-load deadlines, demand-response requests, local energy sharing, and data or equipment failures in one environment, with building and phase power limits constraining controllable requests. It logs controller inputs and separates requested actions from the actions actually applied to the simulated equipment, and ships reference controllers plus service- and constraint-aware performance indicators. A synthetic high-frequency trace replay illustrates how time aggregation can conceal short peaks without changing annual energy.

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

Alvaro Serra-Gomez, Thomas Moerland Planning-based reinforcement learning for high-dimensional continuous control suffers from misalignment between components: learned sampling policies drift from planner behavior, and planning distributions stored in replay go stale as the model and value function change. GEM-MPC builds on Model Predictive Path Integral (MPPI) planning, combining a policy trained to clone the planner with a KL-regularized policy that explores around it. Its Gated Prior Distillation learns from stored planning distributions only when they are a better target than the current prior, avoiding the cost of full reanalysis. Across continuous-control benchmarks it outperforms existing planning-based baselines under lower computational budgets.

What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence

Lyucheng Qian, John Yuehan Zhang, Pingyu Wang In interactive retrieval, an agent must choose which question to ask next so that the answer produces the most useful evidence for the following retrieval update, but existing systems learn this by imitating an offline ordering of candidate question-answer pairs. After showing that candidate discriminativeness and perceived usefulness are only weak supervision for this goal, the authors introduce RAVEL, an online reinforcement learning framework for interactive person re-identification that starts from supervised question generation, sees the current top-4 candidates, and is optimized with rank feedback from the full question-answer-retrieval loop. On Interactive-PEDES, RAVEL delivers progressively stronger retrieval across five interaction rounds. Analysis shows it shifts the questioning budget toward localized open-ended attributes, which yield the largest gains on initially difficult queries.

$\lambda$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

Yufeng Wang, Parivesh Priye, Meeshawn Marathe, Ramit Pahwa Flow-GRPO applies reinforcement learning to flow-matching image generators by treating the denoising sampler as a stochastic policy, but training is unstable: importance ratios drift below one, grow more dispersed, clip at different rates across denoising steps, and leave fewer usable samples late in training. The authors show that these symptoms all arise from a single per-step quantity they call path variance, which is fixed exactly by the sampler's Gaussian transition kernel and can be estimated cheaply during training. λ-Controlled GRPO calibrates importance-ratio behavior from this predicted law and allocates gradient effort across denoising steps by predicted cost, without introducing new free tuning parameters. On a text-to-image model it improves both text-rendering accuracy scored by optical character recognition and human-preference reward over the strongest empirical stabilizer, while keeping late-step path variance within budget where the baseline overshoots.
1 more specialized paper

Reasoning 10

Complex Problem Solving in Large Language Models: A Statistical Control Survey and Diagnostic Framework

Jiazhang Cai, Tao Wang, Ruidong Zhang, Siyuan Li, Terry Ma, Luyang Fang et al. cross-listed Failures of large language models on complex problem solving (CPS), such as early-step error amplification, prompt brittleness, and refusal to revise wrong commitments, are reframed in this survey as a sequential estimation-and-decision problem over a latent solution state. A controller maintains a belief about the unobserved solution trajectory, updates it from noisy intermediate evidence, and decides whether to commit, verify, branch, roll back, or abstain; existing methods are organized around five components: explicit state representation, transition structuring, validation and constraint enforcement, search and rollback, and uncertainty management. The framework yields a diagnostic hypothesis called problem-control fit: interventions work best when they target the specific error or uncertainty component behind a failure, so, for example, extra sampling reduces sampling variability but leaves a shared systematic error untouched. The authors identify reliable validation, targeted recovery, calibrated uncertainty, and matched-budget evaluation as central open problems.

CaLR: Causal Latent Revision for Robust Diffusion Reasoning

Wei Cai, Jian Zhao, Yuchen Yuan, Xuelong Li Autoregressive language models commit greedily to each token, while diffusion language models (DLMs) generate in parallel but lack the strict causal structure that reasoning needs. Causal Latent Revision (CaLR) recasts reasoning as constrained latent optimization: it adopts a causal topology matrix from an expert model and uses implicit differentiation to perform gradient-guided revision of intermediate thoughts, enforcing logical consistency and allowing self-correction during parallel generation. The authors report state-of-the-art results among diffusion language models on complex benchmarks, surpassing strong autoregressive baselines, with greater robustness on constrained tasks such as Sudoku.

Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks

Adrien Deli\`ege, Claas Beger, Marc Van Droogenbroeck, Melanie Mitchell Benchmarks such as the Abstraction and Reasoning Corpus leave it unclear whether a solved task reflects inferring the intended rule or exploiting a shortcut. Using VARC (Vision ARC), where a pretrained backbone is paired with a trainable embedding representing the transformation rule, the authors split test-time training into two stages: tune only the task embedding (Embed-TTT), then freeze it and tune the backbone. On ARC-AGI-1, ConceptARC, and two datasets with known ground-truth rules, the resulting embeddings align better with the true rules and support retrieval and linear probing, and tuning under 0.01% of model parameters alone already solves a non-trivial fraction of tasks before the full pipeline adds more. The learned rule space recovers the geometric structure of parametric rules and interpolates between them, but does not extrapolate.

A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning

Hongyan Wei, Wael AbdAlmageed Perceptual planning requires both reading uncertain scenes and producing action sequences that obey logical rules, but conventional pipelines convert perception into discrete symbolic facts before planning, discarding uncertainty and blocking task-level feedback to perception. This framework keeps a continuous soft symbolic state, lifts domain rules into a differentiable soft-T_P transition operator, and optimizes action logits over a short horizon inside one computational graph, so planning gradients can also update the perception parameters. On Blocksworld it solves 40/40 LatPlan-40 tasks and 596/600 PlanBench-600 tasks, versus 33/40 for LatPlan and 587/600 for a reasoning-model baseline, with substantially less computation and time. In a perceptual-uncertainty ablation, letting gradients refine perception raises success from 59% with frozen perception to 83%, and task-and-motion simulations check that decoded plans are compatible with robotic execution.

When Does Reasoning Help in Machine Translation? A Hierarchical Analysis of LRM Reasoning Traces

Yuxiang Liu, Jiaming Luo, Eleftheria Briakou, Colin Cherry Large Reasoning Models increasingly produce intermediate traces before translating, but it is unclear when that reasoning helps or hurts machine translation quality. The authors analyze traces across models, languages, domains, and datasets along three axes (reasoning language, length, and structure) and introduce Hierarchical Meta-Summarization (HMS), a scalable framework that induces coarse- and fine-grained reasoning structures without a predefined taxonomy. They find that the best reasoning language is model-specific and reasoning length has a non-monotonic relationship with translation quality, while HMS reveals a shared organization of understanding/planning, translating/drafting, and refining/verifying, with domain-specific variation. The takeaway is that translation reasoning should be controlled in a model-aware, length-aware, and pattern-aware way rather than uniformly encouraged.

LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers

Jingyu Hu, Shu Yang, Weiru Liu, Di Wang Chain-of-Thought (CoT) optimization mostly relies on outcome feedback, so large language models can reach correct answers through logically flawed intermediate steps. LogicTrack is a neuro-symbolic framework that auto-formalizes each reasoning step into a symbolic representation and checks it with automated theorem provers. Its Solver-Based Backtracking Reward (SBR) scores the logical soundness of each step and guides a backtracking tree search at inference time, and the resulting trajectories are also turned into supervised fine-tuning data so models internalize step-wise auditing. Across 8 reasoning benchmarks and 7 LLMs, the authors report gains in both the verifiability of reasoning chains and final-answer pass rate.

MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

Arash Lagzian, Srinivas Anumasa, Dianbo Liu Large language models still struggle with hard mathematical, scientific, and logical problems, and standard prompting explores only one way of framing a task. MIRAGE (Multi-perspective Inference-time Reasoning via Agent-Guided Exploration) adds a reinforcement-learning-guided Selector that ranks conceptual perspectives such as algebraic or probabilistic, and a Reasoner that tries them sequentially until a confident answer emerges, otherwise aggregating across perspectives. On GSM8K, MATH500, MMLU-Pro, and Game-of-24, it consistently outperforms Chain-of-Thought and diverse-prompting ensembles with minimal inference overhead.

On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation

Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub, Frederike L\"ubeck, Jonas H\"ubotter, Thomas Kleine Buening et al. On-policy self-distillation conditions a model on privileged information and distills that teacher distribution back into itself, but the privileged context changes how the teacher behaves as well as what it knows. Comparing attraction toward a privileged teacher with repulsion away from one on reasoning tasks, the authors find that attraction suppresses exploratory reasoning and yields shorter, more confident answers, while repulsion lengthens responses, can flip a model into its latent thinking mode, and eventually becomes unstable. Combining attraction to a correct-solution-conditioned teacher with repulsion from an incorrect-solution-conditioned one, studied in isolation from any GRPO objective, makes the shared behavioral shifts largely cancel. The resulting contrastive objective improves reasoning performance while keeping response lengths stable across non-thinking, instruct-only, and already-thinking models.

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

Qiangqiang He, Jin Li, MingCai Chen On-policy distillation (OPD) trains a reasoning model on the token-level discrepancy between a stronger teacher and the student's own samples, but that discrepancy also contains deviations arising from the teacher itself, a problem amplified when the teacher is given privileged information. Calibrated On-Policy Distillation (Cal-OPD) estimates the teacher's self-deviation region using positive and negative privileged interventions and keeps only the part of the discrepancy lying beyond that region. On mathematical reasoning benchmarks, it consistently outperforms standard OPD and its variants across model scales while retaining only about 52-65% of the original discrepancy as the optimization signal.

When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap

Gaoxiang Huang, Lei Qi Activation steering works well for controlling language models during explicit chain-of-thought (CoT) reasoning, which motivates applying it to latent CoT where reasoning happens in continuous hidden states. The authors find that steering continuous thoughts has substantially weaker effects on the generated text than steering explicit CoT, even when hidden representations are shifted by comparable amounts. Since task information remains identifiable in the continuous thoughts, they attribute this to a latent-to-language transition gap, supported by an abrupt change in the output distribution at the transition boundary and much weaker bidirectional control from task-related directions. They identify the transition interface as the main target for future latent-steering methods.