Thursday, September 10, 2026

323 papers cs.AI · cs.LG · cs.CL ← 2026-09-092026-09-11 →

Jul Aug Sep

Highlights

World-Time Compute with Verified Code World Models

Highlight Reasoning James Schwoebel, Ingrida Semenec, Jenia Rousseva, Marcos Ortiz, Collin Overbay, Christopher Klaus et al. Language models generalize within a domain only after seeing many labeled examples, which most domains lack. When a domain's dynamics can be written as code, one template can instantiate many executable, verifiable world models that each yield unlimited exactly-labeled trajectories; fine-tuning on trajectories from many such worlds, which the authors call world-time compute, improves generalization to held-out worlds the model never saw. Gains concentrate where capability is weakest, +29 points for a 0.5B model, while the largest model's lift is within noise; a corrupted-label control shows label exactness rather than task variety drives the effect, and on List Functions one adapter trained on 128 disjoint worlds reaches 40% on held-out worlds versus 6% for the control. The same lever works as per-world test-time training on ARC-AGI, List Functions, and CLRS, fading for long reasoning chains and perception-heavy tasks, with worlds authored and served by the OpenWorld framework.

Large language models only generalize within a domain after seeing many exactly labeled examples, which most domains lack. The paper's answer is "world-time compute": synthesize many code world models (executable, invariant-checked programs over symbolic state) from a domain template, then fine-tune on the endless exactly labeled trajectories they produce, so the model generalizes to worlds it never trained on.

  • An LLM writes each world's transition function from declared rules and the program is accepted only after it parses with the required signature, survives sandboxed runs that preserve declared invariants, and optionally passes a second-model critic, with the zero-dependency OpenWorld framework making authoring and serving these worlds a minutes-scale task.
  • Verified code dynamics are exact on 24/24 twenty-step rollouts and score 100% on 10x out-of-distribution probes, while per-step LLM next-state prediction completes 0/24 and diverges by step 2.3 on average, and a 10,000-transition MLP scores 0% out of distribution.
  • On a synthesized diagnosis-world family (60 train, 20 held-out specialties), fine-tuning lifts held-out accuracy by +29 points at 0.5B (37% to 66%) but only +3 at 32B where the task saturates, and a corrupted-label control shows accuracy falling monotonically from 35% with exact labels to 19% with fully wrong labels, so label exactness rather than task variety drives the gain.
  • On human-authored benchmarks the mechanism transfers: a single adapter trained on 128 disjoint List Functions worlds reaches 40% on 60 held-out worlds versus 6% for a corrupted-label control (+34 points, CI [29, 39]), and per-world test-time training on ARC-AGI, List Functions, and CLRS-Text scales with compute (to 10%, 69%, and 17% respectively) and collapses to the floor under label corruption.
  • The lever fades for long reasoning chains, saturated tasks, and perception-heavy problems (Bongard-RWR stays at chance), cross-world transfer fails when worlds share no skill (held-out CLRS algorithms and cross-language HumanEval-X frames show no lift), a 164-world coding family hurts HumanEval greedy pass@1 (78% to 70%) while helping pass@5, and "verified" only certifies invariant-consistent code rather than agreement with true domain dynamics, with most experiments run single-seed on quantized 7B-class models.

AgenticGen: Reward-Guided Agentic Video Generation for Advertising

Highlight HF pick · 2▲Applications Xingyuan Bu, Chengru Song, Hao Zhou, Tao Zhou, Dong Li, Wei Li et al. Advertising video generation is judged by online business metrics, yet video foundation models do not learn how a product should become an effective ad or improve from online feedback. AgenticGen decomposes the task into two trainable stages, strategy selection and draft generation, learns a performance-based reward from accumulated online feedback plus a rubric-based reward aligned with human quality standards, and optimizes the stage policies with DPO followed by GRPO using process and outcome rewards. Online A/B tests in the TikTok advertising system show CTR up 2.72%, CVR up 2.63%, and advertiser value up 9.61% over a supervised fine-tuned baseline.

Advertising video generation is framed as a product-conditioned reasoning problem judged by online business metrics, not just a synthesis task, and off-the-shelf video foundation models never learn from that feedback. AgenticGen closes the loop by splitting ad generation into two trainable reasoning stages and optimizing them against reward models learned from online performance and human quality rubrics.

  • The pipeline decomposes ad creation into a strategy-selection stage (how a product should be turned into an ad) and a draft-generation stage, both implemented as agentic policies that expose supervisable optimization targets rather than a single opaque generation step.
  • Two reward signals guide training: a performance-based reward fitted on accumulated online business feedback, and a rubric-based reward aligned with human quality standards, which together provide both process and outcome supervision.
  • Policy optimization runs in two passes, with DPO first shifting the policies toward online preferences and GRPO then refining both stages using the process and outcome rewards.
  • In online A/B tests on the TikTok advertising system, the DPO+GRPO model beat the SFT baseline by +2.72% CTR, +2.63% CVR, and +9.61% Advv, with offline experiments separately validating the reward models and the successive optimization stages.
  • Results are tied to a single proprietary ad platform with a learned performance reward, so gains depend on access to large-scale online feedback and may not transfer to settings without such a feedback loop or to reproducible public benchmarks.

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

Highlight HF pick · 13▲Agents Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng AI research agents draw on prior knowledge, web sources, and experimental feedback, which makes it hard to tell whether a reported result is a genuine discovery or a recovery of something already known. The Discovery Certification Protocol (DCP) converts such claims into executable tests: Gate 1 checks for useful improvement on a sealed evaluation, Gate 2 gives matched agents the registered starting information and observed web content while withholding the target research history to see whether they recover the result, and an optional Gate 3 measures the effect of truthful feedback against a neutral policy from a shared checkpoint. Two controlled audits on SQLite optimization and virtual catalyst control produced zero recoveries in 96 episodes each, with an upper bound of 0.0468 on the recovery probability, and paired studies yielded 30 truthful recoveries against zero neutral ones. A deterministic verifier that uses no language model reproduces the certification decisions from frozen evidence.

AI research agents report high scores, but a score on a sealed test says nothing about whether the result depended on the agent's experimental history or was reachable from background knowledge and public sources alone. The Discovery Certification Protocol (DCP) turns that question into executable tests: a sealed utility check, a recovery challenge where matched agents receive the starting information but not the run's experiments, and an optional randomized comparison of truthful versus neutral feedback.

  • Gate 1 requires a lower confidence bound on the gain over a registered baseline of at least a preset minimum, Gate 2 hands fresh matched agents the background, initial observations, and every Web byte the target run saw while withholding its experimental history, and any valid artifact scoring within a tolerance ε of the target counts as a recovery witness that vetoes the Core certificate, which otherwise needs zero recoveries in n episodes, passing positive controls, and a finite-sample bound p_upper = 1 − α^(1/n) below a registered ρ.
  • Two full audits, SQLite partial-index selection with DeepSeek-v4-flash and virtual catalyst control with DeepSeek-v4-pro, each produced 0 recoveries in 96 challenger episodes with 45/45 positive controls and an upper recovery bound of 0.0468, while the best challengers reached only 0.6734 against a 0.8805 recovery line and 0.8146 against 0.95.
  • Gate 3 paired studies from a frozen checkpoint gave 30/30 truthful versus 0/30 neutral recoveries (estimated effect 1.0, exact 99% interval [0.6379, 1.0]) and the separate 60-pair null calibrations showed a contrast of 0 with intervals of ±0.095 inside the registered ±0.17 band, so both audits earned the Evidence decision.
  • Calibration cases cover the rest of the decision space: a multidimensional knapsack challenger found two legal solutions at 0.9363 and 0.9356 above the 0.9329 line and correctly refuted Core, an undersampled affine case returned audit-incomplete, and a deterministic LLM-free verifier replays every decision from frozen bundles.
  • The tasks are small synthetic environments with known optima, each bound holds only for the registered model, budget, and information packet, the released bundles are locally registered with formal certification not yet issued, and each audit cost roughly 435 to 507 model sessions and about $56 to $61.

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

Highlight Agents Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla Agentic workflows have complex failure modes across planning, tool invocation, and environment interaction, making it important to know how confident an agent should be that its actions will succeed. The authors test whether a model's internal representations carry stronger signals of eventual task success than its outputs, introducing Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action decision points. Across three interactive benchmarks (Bash, SQL, Python) and three model families (Qwen14B, Qwen7B, DeepSeek6.7B), both methods consistently outperform surface-level generation and sequence-based calibration baselines, providing a zero-overhead reliability monitor that needs neither prompt changes nor multi-sample rollouts.

Confidence estimation for multi-turn agents is hard because output-level signals like token log-probabilities miss the internal states that precede planning and tool-use failures. The paper shows that reading an agent's residual-stream activations, either as trajectory geometry or as a probe on action-time states, predicts eventual task success better than surface-level calibration baselines.

  • LTD (Latent Trajectory Dynamics) teacher-forces stored trajectories through the generating model, records residual states at observation, reasoning, action, and feedback boundaries, and summarises cosine and relative displacement between adjacent states into 28 kinematic features fed to an L2-regularised logistic model.
  • ARP (Action Representation Probe) mean-pools the final-layer residual state at the last token of each non-submit action span, standardises and projects it onto 64 principal components fit inside each training fold, and trains a logistic probe followed by a monotone Platt calibrator.
  • Across InterCode Bash, SQL, and Python with Qwen2.5-Coder-14B, Qwen2.5-Coder-7B, and DeepSeek-Coder-6.7B, an internal method achieves the best AUROC and Brier score in all nine model-environment settings; for example ARP reaches 0.842 AUROC on SQL and 0.814 on Bash with Qwen-14B, versus 0.743 for the trajectory-level HTC baseline and 0.624 to 0.627 for calibrated log-probabilities.
  • ARP generally dominates LTD, but not uniformly: on DeepSeek Bash LTD leads with 0.727 AUROC against 0.650 for ARP, and on Qwen-7B SQL the external HTC baseline edges out LTD (0.800 vs 0.786), so the gains are consistent in aggregate rather than per-method.
  • Evaluation uses nested 5-fold cross-validation with fold-internal standardisation, PCA, and calibration, and the authors argue deployment cost is negligible (under 0.05 ms per action for a 5,120-dimensional state), though results cover only three small open-weight coder models, a 200-task Bash split, and greedy decoding, and residual-state extraction from optimised serving stacks like vLLM requires custom hooks the paper does not implement.

Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

Highlight Reasoning Yunxiang Mo, Donghao Zhao, Hejia Geng A tempting way to cut reasoning-model inference cost is to repeatedly probe a partial trajectory for its current answer and stop once the probes agree, a rule called self-consensus. A preregistered sweep of 3,520 such rules, replayed on frozen trajectories from two models and three benchmarks, finds that none clears three pre-specified acceptance gates, a result that reproduces on a held-out split and two unseen models, whereas a boundary-confidence control, DEER, run through the same pipeline clears all three. The failure traces to a consensus-termination gap: agreement shows the answer persists under a fixed probing procedure, not that reasoning has finished, so at a rule saving 32% of tokens one stop in nine fires on an answer the trajectory later abandons, usually cutting off a correction, and widening the agreement window only lowers that share to about 7% while shrinking savings to 8%.

Reasoning language models can burn thousands of tokens per answer, and a tempting way to cut that cost is self-consensus: periodically probe a single partial chain of thought for its current answer and stop once recent probes agree. The paper argues this proxy measures the wrong thing, since agreement under a fixed probing procedure only shows the current answer persists, not that reasoning has terminated, which the authors call the consensus–termination gap.

  • The study freezes one trajectory per problem from DeepSeek-R1-Distill-Qwen-7B and Qwen3-8B on MATH 500, AMC23, and AIME24, probes fixed prefixes offline with a boxed-answer suffix, and replays a preregistered grid of 3,520 consensus rules (window size, share threshold, probe schedule, maturity floor, certainty and validity filters) against three acceptance gates fixed in advance, charging probe output tokens against savings.
  • Not one of the 3,520 rules clears any gate on the development split: the only 40 rules that stay within a 1.0 pp accuracy drop save at most 0.2% of tokens, while the first rule to save 10% costs 2.66 pp, 20% costs 6.17 pp, and 30% costs 11.8 pp; the frontier reproduces on the held-out test split (r = 0.98) and on two unseen models (Distill-Qwen-32B and DeepSeek-R1-Distill-Llama-8B), and the shipped CertaIndex (CoT) default loses 56–70 pp while saving 77–90% of tokens.
  • The failure is directional: at a 12-probe window still saving 32% of tokens, one stop in nine commits an answer the trajectory later abandons, and of those 216 stops, 155 cut off a correction the model would have made versus only 16 that bank a correct answer from a trajectory ending wrong; widening the window only flattens the non-terminal share near 7% while savings collapse to 8%.
  • Two diagnostics explain why: re-probing the same frozen prefix with a differently worded query returns a different answer 54% of the time in the first tenth of a trajectory (versus 16% near the end), and a two-annotator taxonomy of 134 wrong stops finds 56.7% commit to an answer the model had not converged on, 18.7% to a probe-format artifact, and only 24.6% to a genuinely settled wrong value.
  • Early exit itself is not the problem: DEER, a boundary-confidence method swept through the identical pipeline and gates, clears all three, losing only 0.33 pp while saving 28.2% on dev and holding on the unseen models, though the results cover only competition math with checkable answers, one probe wording, a single consensus schema, and small AIME/AMC dev splits of 6–8 problems per seed.

Through the Looking Glass: Directly Reading and Writing Transformers

Highlight Large Language Models Mark Oskin Counting transformer components by the absolute value of their logit contribution suggests thousands participate in each token prediction, but the authors argue contributions are signed and largely cancel: across eighteen models the mass pushing away from the predicted token is a median of seven times the mass carrying it. Dividing by the net, on the baseline model 53 components carry ninety percent of a prediction, 13 are indispensable, and 8 suffice to produce it alone, and across twelve models from 124M to 7B parameters the full backward trace touches only one to three percent of the model regardless of size. Everything is read directly from parameters and activations with nothing trained or fitted, naming each component by what it writes and what it reads, with input identification at 58.9 percent above chance. The same access supports edits: a new association installs into one spare unit at a fortieth of the held-out loss cost of a rank-one update, and installed heads can make an edit fire only when a token appeared earlier in context.

Standard attribution says a single token prediction in a transformer rests on thousands to hundreds of thousands of feed-forward units and attention channels, a count too large to explain anything. Oskin argues that number is an accounting artifact: per-component logit contributions are signed and mostly cancel, so dividing by the net rather than the absolute total collapses the same attribution to dozens of components, and a lens built purely from the model's own weights and activations, with nothing trained or fitted, can then name, trace, and edit those components directly.

  • The logit is written as an exact sum over feed-forward units and (head, channel) attention pairs, ranked by signed contribution; on a GPT-2 small-style baseline trained on OpenWebText, reaching 90% of the prediction takes a median of 53 components under the net denominator versus 23,971 under absolute value, a 452x inflation, with opposing mass running 7x the net at the median across 18 models and 20x on the baseline.
  • Necessity and sufficiency are tested by ablation and a repair-then-cut search: removing the top 13 components flips the prediction where random removal does not, and 8 components suffice to reproduce the token with every other component at that position zeroed (0.017% of 46,080), a feat 8 random components achieve on 0 of 115 predictions, while across 12 external models from GPT-2 to Mistral 7B the sufficient set runs from 2 to 16 components and the full dependency closure stays at 1 to 3% of the model regardless of size.
  • The remainder is neither idle nor staging for later tokens: keeping only the circuit at a position costs the same as a random set on each of the next 512 tokens (ratio 0.98 to 0.99), and 74% of what a layer adds to the residual is a fixed linear map of its input (63% a pure rotation), so substituting one layer's output with that map preserves the predicted token on 92% of cases.
  • Components are named on both sides and then edited: a unit's eight strongest inputs are recovered from its weights at 58.9% overlap, 37x chance, and installing a new association into one spare unit with key and value read from the weights moves the target from rank 578 to rank 1 for 0.25% held-out loss, a fortieth of what a rank-one editor costs to reach the top ten, with a two-layer installed circuit and an 86%-through tap into a trained unit showing that edit placement is governed by depth.
  • Limitations are substantial: credit is direct, so components acting only through downstream components are invisible to both the ranking and the sufficient-set search; ablations realize only 42 to 50% of the predicted logit drop because surviving components self-repair; late layers are over-credited by a factor of about 0.53; sufficient-set sizes are heuristic upper bounds whose search cost grows with depth; and parallel-sublayer architectures like Pythia are excluded entirely.

$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

Highlight HF pick · 4▲Large Language Models Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han et al. Existing coding benchmarks test isolated kernels, predefined operators, or fixed optimization targets, so they cannot tell whether large language models (LLMs) can do open-ended, long-horizon engineering on the infrastructure that serves them. Φ-Bench derives tasks from optimization problems studied in frontier research and grounds them in real code repositories, spanning localized kernel-level function completion through long-horizon implementation to end-to-end system optimization across the LLM infrastructure stack. Experiments on frontier LLMs expose substantial remaining limitations in engineering complex LLM infrastructure, and the authors use the results to characterize the challenges on the path to autonomous optimization of AI infrastructure.

Kernel benchmarks like KernelBench and TritonBench hand the model a fixed interface and a named optimization target, so they never test whether an agent can navigate a real training or serving codebase, find the bottleneck, and iterate on a fix. Φ-Bench (the Frontier AI Infrastructure Benchmark) grounds 85 tasks in real systems papers and public repositories, spanning nine infrastructure domains and three formats of increasing open-endedness, and finds that even the best frontier model scores just over a third of the maximum.

  • Task synthesis is guided by a bottom-up taxonomy built from 2,260 papers and 1,852 repository artifacts (410 fine-grained tags clustered into 62 topics and 9 categories), with tasks reconstructed from real PRs and issues, mined by an agent loop that deletes implementation sites and generates coverage-driven tests, or curated by experts from influential papers.
  • The three formats escalate in scope: KFC (55 tasks) completes or optimizes one kernel in a single file with a single submission, LHI (20 tasks) implements an issue-style feature across multiple files, and E2EO (10 tasks) gives only a system-level objective over a fully editable repository, with the latter two allowing up to 16 submissions.
  • Rewards are correctness-gated and proctored against cheating by both rule-based monitors and an inspector agent; implementation tasks score binary pass/fail, while performance tasks use a log-normalized speedup relative to the reference solution that gives zero credit for merely matching the reference and full credit only at the reference speedup squared.
  • Claude Opus 5 leads with 36.53% overall (37.16% KFC, 21.60% LHI, 62.94% E2EO), ahead of Kimi K3 at 28.12% and Qwen3.8 Max at 27.73%, while every model collapses on Hardware & Edge tasks (best is 5.4%) and all score lower on long-horizon implementation than on kernel completion.
  • Trajectory analysis shows the strongest models produce more errors because they persist through trial and repair, and Opus in particular runs cheap local validation experiments, isolates variables against measurement noise, and checks confounders like Inductor cache invalidation before attributing outcomes; the study's weaknesses are a single-task iteration ablation, only ten E2EO problems, and the fact that all models except GPT-5.6 Sol ran inside the Claude Code scaffold.

A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out

Highlight Other Mahdi Naser Moghadasi (BrightMind AI), Faezeh Ghaderi (University of Texas at Arlington) Time-series foundation models are evaluated almost entirely on public archives that predate them, so a strong score cannot be separated from having seen the test data during pretraining. The authors build a contamination-free hold-out: thirteen forecasters (four classical, three trained per dataset, six pretrained) on seven groups from five domains, with every observation published after the last model's release and every dataset rebuildable without an API key. Pretrained models win five of seven groups, lose one to a Theta baseline, and are indistinguishable from seasonal naive on daily exchange rates, yet neither seasonal strength nor spectral entropy of the input window explains the pattern; what tracks the wins is corpus familiarity, with the largest gain (28% lower mean absolute scaled error) landing on weekly Wikipedia pageviews, the domain that dominates TimesFM's pretraining corpus, and the TimesFM family outranking Chronos far more on Wikipedia than elsewhere. The conclusion is that a temporal hold-out removes memorization of a window but not familiarity with a domain, so benchmarks need domain hold-outs stated relative to disclosed corpora and practitioners should ask whether their domain is one the model was raised on.

Time-series foundation models are scored almost entirely on public archives that predate them, so a strong benchmark result cannot be separated from having seen the test data during pretraining. The authors build a hold-out in which every observation postdates every evaluated checkpoint, and find that this removes memorisation of a time window but not the advantage a model gets from having been pretrained on the same kind of series.

  • The benchmark runs 13 forecasters (four classical, three trained per dataset, and six pretrained: Chronos-Bolt small and base, Chronos-2, TimesFM 2.5, TimesFM 3.0, Moirai-2) on seven groups from five domains (Wikipedia pageviews at daily, weekly and monthly granularity, hourly ERA5 weather, hourly CAMS air quality, hourly Danish grid electricity, daily ECB exchange rates), with a fixed forecast origin of 1 January 2026, keyless rebuildable fetchers, and comparison by average rank with Nemenyi intervals and Holm-corrected Wilcoxon tests rather than by mean loss.
  • Pretrained models win 5 of 7 groups, lose electricity to a Theta baseline, and on exchange rates are statistically indistinguishable from every other method including seasonal naive, with the largest gain being 28% lower MASE than the best classical method on weekly Wikipedia pageviews.
  • Neither seasonal strength nor spectral entropy of the input window explains where the advantage appears (seasonal strength is if anything negatively associated), but corpus familiarity does: on identical series, where difficulty cancels, the TimesFM family outranks the Chronos family by 0.53 ranks on Wikipedia versus 0.09 on the other four domains (1,500 vs. 754 series, Mann-Whitney p < 10^-5), and Wikipedia pageviews are the disclosed bulk of TimesFM's pretraining corpus.
  • Two findings the point-accuracy tables hide: every pretrained model's nominal 80% interval undercovers (Moirai-2 worst at 71%) while all three automatic classical methods overcover, and AutoARIMA costs a median 54x more per series than a pretrained forward pass, rising to 2,689x on hourly weather.
  • The corpus-overlap result is an association rather than a demonstrated mechanism, with a modest rank-biserial effect of -0.13, a single forecast origin per group, two groups under 50 series, and only one model whose corpus composition is disclosed well enough to test.

ConvMem: Convolutional Memory for Long-Context Reasoning

Highlight Large Language Models Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang et al. Large language models struggle with extremely long inputs because of fixed context limits, and sequential memory approaches like MemAgent that read text segment by segment incur high latency and require costly reinforcement learning training that can overfit to specific datasets. ConvMem is a training-free framework that reformulates long-context reasoning as a hierarchical convolution: a query-prompted LLM acts as a convolutional kernel that summarizes text segments level by level, turning a linear reasoning chain into a logarithmic-depth tree. Configurable strides and skip connections help capture and propagate evidence, while multi-kernel convolution decomposes complex queries into separate semantic channels, and the whole process parallelizes across both segments and reasoning threads. On RULER-HotpotQA and RULER-2WikiMultiHopQA it outperforms training-free baselines and avoids the overfitting to parametric priors seen in RL-trained models on out-of-distribution tasks.

Sequential memory agents such as MemAgent extend LLM context by streaming text through a fixed-size memory, but the strict step-by-step dependency forces linear latency and the memory policy usually needs RL training that can overfit. ConvMem instead treats a frozen, query-prompted LLM as a convolutional kernel that summarizes overlapping text segments in parallel and stacks those summaries hierarchically, so the reasoning path becomes a logarithmic-depth tree with no training at all.

  • Each kernel call takes a segment plus a sub-question and returns a condensed summary and a relevance score in {0, 1, 2}; segments scored 2 are copied verbatim into a residual buffer that skips straight to the final answer step, while the concatenated summaries are re-sliced and re-summarized layer by layer until they fit in one window (W = 8000 tokens, stride 1600, so every token is scanned 5x; depth is typically 2 to 4 layers up to 1M tokens).
  • Multi-kernel convolution first decomposes the query into 2 to 4 sub-questions, runs an independent kernel and residual buffer per sub-question, answers each sub-question from its own summaries plus raw evidence, and then synthesizes the final answer from the sub-answers.
  • On the out-of-distribution RULER-2WikiMultiHopQA with Qwen2.5-32B-Instruct, ConvMem reaches 72.3 / 59.1 F1 at 28k / 896k tokens versus 60.9 / 58.4 for RL-trained MemAgent and 60.5 / 49.7 for training-free MemAgent-W/O-RL; on in-distribution RULER-HotpotQA it beats every training-free baseline but trails MemAgent (63.1 vs 68.8 F1 at 896k, with a notable dip to 57.8 at 112k).
  • Ablations show single-pass scanning, dropping skip connections, or using one kernel instead of several all hurt, and the paper argues MemAgent's in-distribution edge partly reflects memorized labels, citing a case where it reproduces a dataset typo that contradicts the provided context.
  • Total token consumption is higher than linear scanning since every token is processed five times per layer, results are reported only for 128-example RULER-style multi-hop QA, latency is claimed as O(log N) but not measured against baselines, and the whole pipeline depends on the backbone correctly decomposing the query.

Show-Harness: Just a VLM Agent Can Play Robots

Highlight HF pick · 42▲Robotics Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin et al. Foundation vision-language models (VLMs) hold broad knowledge about the world, but turning that into robot control remains hard. Show-Harness is an embodied harness that exposes discrete semantic action units the VLM can reason over, with embodiment-specific interpreters deterministically grounding them into local robot actions, so the VLM stays responsible for fine-grained physical decisions. Through this interface, closed-source frontier VLMs can control robots zero-shot, and small open-source VLMs can be adapted with only a few GPU-hours of fine-tuning, while the companion GUMI (GUI Manipulation Interface) lets humans and agents collect demonstrations across embodiments without teleoperation hardware. Harness-equipped VLM agents generalize across tasks, embodiments, and environments and outperform representative agentic and vision-language-action (VLA) paradigms, suggesting the interface rather than added model capacity unlocks embodied capability.

Foundation vision-language models know a lot about objects, spatial relations, and task decomposition, but turning that knowledge into robot control usually means either fine-tuning them into opaque vision-language-action (VLA) policies or hiding execution behind skill libraries and controllers. Show-Harness instead exposes a compact, discrete set of semantic action units (single-step moves, incremental rotations, grasp, release, done) that a VLM reasons over directly, with an embodiment-specific interpreter deterministically grounding each unit into a small bounded robot motion.

  • The harness wraps the VLM in a perceive-reason-act loop with pluggable modules for multi-view guidance, proprioception text, subtask planning, situated replanning, adaptive step size, action chunking, action history, and grasp-failure recovery, and the same interface supports both zero-shot control with frontier models (Gemini-3.1 Pro by default) and rank-64 LoRA fine-tuning of Qwen3.5-2B on a few GPU-hours from just 164 real-robot episodes collected via a keyboard GUI called GUMI rather than teleoperation hardware.
  • On ten real pick-and-place tasks across a Franka arm and a dual-arm AgileX rig, both the zero-shot and fine-tuned agents outperform π0.5, GR00T, VLA-centric agents, and code-as-policy baselines across task, environment, and embodiment shifts, and the fine-tuned 2B model transfers sim-to-real using only simulated demonstrations where the VLA baselines fail.
  • Because precision lives in the interpreter rather than the policy, halving the step size from 2 cm to 1 cm without retraining lifts stacking and peg-insertion success from 60% to 82% (zero-shot) and 40% to 65% (fine-tuned), while π0.5 sits at 18% with the same data; incremental 15° rotation units also let the fine-tuned model reach 70% on an unseen 90° carrot orientation versus 20% for π0.5.
  • Ablations show the harness matters: dropping subtask planning cuts success to 60%, dropping failure recovery to 72%, and replacing semantic action names with arbitrary symbols still works when written conventions are supplied but collapses to 1/20 successes when the agent must infer action effects from observation alone.
  • Evaluation is limited to parallel-jaw single- and dual-arm manipulation with 10 trials per task, frontier-model errors concentrate on fine-grained grasp and placement rather than planning, and higher thinking budgets mostly trim redundant steps while raising wall-clock cost up to 3.4x.

Applications 83

Auditable Emergency Triage for Maternal and Newborn Care in India

Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini, Neelima Devadas, Amrita K Prasen, Nikhil Nalin et al. Nurses at Noora Health answer over 50,000 medical queries a month on a WhatsApp service for caregivers, and their most time-critical task is deciding which messages are emergencies. An earlier large language model (LLM) classifier with free-form rationales was opaque and costly to iterate on, so triage was decomposed into two stages: an LLM extracts canonical symptoms and patient context using a clinician-authored vocabulary, and a deterministic rule engine encodes the scenarios that indicate an emergency. The redesign raised recall from 0.565 to 0.810 and F1 from 0.606 to 0.702, with the structured rules driving most of the gain while the decomposition lets clinicians inspect each stage for mistranslation, extraction errors, wrong context, or missing rules, and add rules without regressions or full re-evaluations. Since deployment the system has triaged 152,421 queries, flagging 18.7% as emergencies with a 17.8% over-escalation rate and no increase in missed emergencies, and clinicians have added 48 new rules.

X-amine509: Predicting the Practical Risk Level of Enterprise X.509 Certificates

Cameron Keith, Shubh Patel, JD Kilgallin, Caleb Shorter cross-listed Enterprises holding millions of X.509 certificates cannot afford to run exhaustive deterministic standards checks on every one, so the authors build a two-stage triage system that uses machine learning to rank certificates by predicted risk and route only the riskiest to full analysis. Risk is a composite score from 177 defect checks grounded in CA/Browser Forum Baseline Requirements and NIST guidance, weighted across four severity tiers, computed on 1,027,714 certificates gathered from Fortune 500, .gov, and .edu domains. On a held-out set of 201,976 certificates, an Extra Trees regressor reaches R² of 0.993 with mean absolute error of 2.26, while a Decision Tree scores R² of 0.986 at 3.7 million certificates per second on one machine, with 98.90% recall on critical-tier defects. Retesting on 571,374 certificates collected thirteen months later shows only modest degradation, and feature analysis points to validity period, Extended Key Usage configuration, negative serial number encoding, and self-signed status as the strongest risk predictors.

The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption

Sales G. Aribe Jr., Louie Jay S. Labastida cross-listed Vibe coding, in which developers produce software by conversing with a large language model in natural language, is compared against traditional and AI-assisted coding in a mixed-methods study where 30 professional developers and advanced students completed equivalent tasks under all three conditions. Vibe coding cut task completion time by 27% versus traditional coding and 12% versus AI-assisted coding, but produced lower maintainability indices and more security vulnerabilities. Usability scored well on the SUS scale and workload was moderate on NASA-TLX, while thematic analysis surfaced trust calibration, loss of control, cognitive adaptation, and prompt-engineering strategy as central themes, with perceived loss of control linked to higher security risk. The authors propose a three-pillar framework for responsible adoption built on hybrid human-AI integration, human oversight with transparent accountability, and context-aware deployment.

Robust Industrial Cyber Physical Classification Using Neuromorphic Temporal Embeddings and Hybrid SNN XGBoost Under Machine Unlearning Attacks

Ammar Kamoona, Sajad Koushkbaghi, Mahdi Jalili, Peter McTaggart, Xinghuo Yu Digitalised power-distribution grids need intrusion detection that runs at the edge, yet deep-learning detectors are costly to deploy and their periodic retraining exposes them to machine unlearning attacks, where selective data removal degrades detection. The proposed architecture pairs a Spiking Neural Network (SNN), trained once on clean data and then frozen as a temporal feature extractor, with an XGBoost classifier that is the only component retrained during updates. On two public power-system datasets the hybrid reaches 99.9% accuracy on the Synchrophasor data and 95.0% on the MSU/ORNL data, and under label-flipping attacks it loses only 0.9% macro-F1 at 10% poisoning while delaying target-class collapse from 60% to 70% poisoning compared with unprotected models.

Positional task conditioning for scalable defect detection across product families in large product catalogs

Soham Satyadharma, Gabriel Roccabruna, Suleiman A. Khan Product families in large e-commerce catalogs contain defects such as duplicates and unit mismatches, and asking a large language model to spot many error types across long listings suffers from long-context degradation. Decomposing detection into focused sub-tasks that shrink the context and isolate one error type each raises F1 from 52% to 87%, and Positional Task Conditioning (PTC) then distills that capability into a single smaller model by reinforcing task identity at structural prompt boundaries. PTC beats rationale-based distillation across five models and two architecture families, landing within 1.79% F1 of the frontier model at up to 98% lower cost, and the system is deployed across multiple countries over more than 10 million product families.

Scaling E-Commerce Attribute Extraction with Parallel Decoding

Nikhita Vedula, Dushyanta Dhyani, Bryan Wang, Shervin Malmasi E-commerce catalogs are messy, and standard Attribute Value Extraction (AVE) systems treat every attribute equally, yielding large and inconsistent attribute sets that do not reflect what shoppers actually compare. The proposed two-stage pipeline first uses a large language model to discover a compact ranked schema of purchase-discriminative attributes per product category, then extracts values from catalog text with a fine-tuned Qwen3-4B using Hyper-Parallel Decoding (HPD). Extraction accuracy reaches 85%, matching the foundation model it was distilled from, at 92% lower inference cost, and the resulting category-level structured outputs act as automatically built product knowledge bases for downstream applications.

Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending

Gijs A. F. Niewzwaag, Marijn G. S. Veth, Manuele Massei, Marcos R. Machado Credit-scoring models in peer-to-peer (P2P) lending can be gamed by applicants who alter self-reported inputs, yet most adversarial-robustness evidence comes from images and text and tests one attack against a matching defense. The benchmark trains logistic regression, a feed-forward network, and a tabular transformer on a large Lending Club subset and evaluates them under Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Salt-and-Pepper noise, and DeepFool attacks restricted to applicant-mutable features, plus a mixed regime, with stratified cross-validation. Adversarial training strongly hardens a model against the attack it saw and transfers within the gradient-based family, but transfers weakly to non-gradient corruption, so single-attack defenses overstate real-world resilience; mixed-attack training gives the most balanced robustness while preserving clean-test performance.

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

Stig Hellemans, Tom Stroobants, Elyne Scheurwegs, Pieter Meysman, Philippe G. Jorens, Kris Laukens Clinical notes carry personally identifiable information that blocks reuse for research, and institutions often cannot send that text to an external service for de-identification. MedDeID runs entirely on premises, combining in-house annotation, synthetic note generation, model training, inference, pseudonymisation, and evaluation in one workflow. On an independently adjudicated 300-note Dutch hospital benchmark a compact transformer trained on hospital text found 98.9% of identifying text while over-redacting only 0.24%, and a counterpart trained purely on synthetic notes still reached 96.1%, then beat the hospital-trained model on 100 primary-care notes with 90.3% versus 87.0% recall and greater robustness to identifier-format changes. An English version trained without any real text detected 99.7% and 98.9% of identifier characters on two external synthetic benchmarks, which the authors present as evidence the workflow transfers across languages rather than as validated clinical English performance.

Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations

Soumyadeep Roy cross-listed Voice-cloning fraud often takes the form of surgical injection, where a real recorded conversation has only one or two sentences swapped for synthetic speech, which clip-level detectors cannot localize because they emit a single label. The authors define Temporal Deepfake Localisation in Multi-Speaker Conversations, argue that equal error rate and minimum detection cost become ill-posed when one file contains both real and fake speech, propose temporal metrics instead, and build a training-free five-stage pipeline that wraps any frozen binary detector and converts noisy window scores into intervals using a two-threshold hysteresis state machine. On 180 constructed conversations from ASVspoof 5 the system reaches temporal intersection-over-union of 0.90 with a 0.95 detection rate, with false alarms on genuine speech under 6% and under 2% on real multi-speaker dialogue from AMI. Training a supervised localiser inside the same pipeline improved temporal overlap by only about 0.04, which bounds what supervision buys here.

A Systematic Evaluation of Molecule Generation Models for De Novo Drug Design: From Benchmarks to Practical Insights

Xinrui Xu, Xueer Wang, Dan Luo, Sisi Yuan, Xuan Lin Generative models for de novo drug design have proliferated, but reviews typically cover one model family at a time rather than the whole generation workflow. This survey evaluates 82 methods across five deep generative frameworks — recurrent and Transformer-based sequence models, variational autoencoders (VAEs), generative adversarial networks (GANs), flow-based models, and diffusion models — alongside the molecular representations and benchmarks they are scored on. Its main contribution is a comparative synthesis of reported results across common benchmarks and metrics for both unconditional and pocket-conditioned generation, plus a summary of experimentally validated case studies. The authors flag standardized 3D data, interaction-aware generation, receptor flexibility, and multi-objective design as open gaps, and release all collected benchmarks and references in a public repository.

Are You Learning Biological Signal or Shortcuts? Auditing and Mitigating Bias in Protein-Protein Interaction Datasets

Judith Bernett, Anton Spannagl, Joel {\AA}s, Markus List, David B. Blumenthal cross-listed Machine-learning models for protein-protein interaction (PPI) prediction can score well by exploiting artifacts of how interaction databases were built rather than learning biology. The authors audit HIPPIE, IntAct, STRING, and two datasets derived from Protein Data Bank structures, showing that random train-test splits introduce strong topological shortcuts and that, even after removing protein overlap, shortcuts from self-interactions, taxonomic identity, and functional relatedness remain, with their prevalence depending on the data source. Sampling negatives from high-confidence non-interactors, an intuitively appealing choice, actually amplifies the functional-relatedness shortcut. They release a Nextflow pipeline that performs similarity-aware, data-loss-minimizing splitting and bias-minimizing negative sampling as integer linear programs, an approach they argue extends to any task where negative candidates vastly outnumber positives.

PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving

Lin Huang, Yujuan Tan, Weisheng Li, Lixiang Zeng, Kun Yang, Suihan Xiao cross-listed PACE is a framework for retrieval-augmented dialogue serving that formalizes Perceived Time-to-First-Response (PTFR) as a quality-of-experience objective and minimizes it under quality and cost constraints. Unlike prior cascaded routing, semantic caching, or adaptive retrieval work, it jointly decides which answer source composes the response and what fills the waiting window, combining a load-adaptive cascading router, a joint path-filler controller, and volatility-aware cache admission, deployed on a humanoid-robot sales service. On 75k CarQA requests the cascade halves pure-LLM P95 PTFR (0.29s versus 0.53s at 16 concurrent clients), the adaptive controller reaches 0.41s P95 and outperforms plain retrieval-augmented generation by 2.4 times at high load with equal quality, the filler controller cuts calls by 94% with zero filler-answer conflicts, and volatility-aware admission drops stale answers from 86% to 0%. A gating rule guarantees the controller never does worse than the baseline, with exposure bounded by one hold period.

Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

Ayan Majumdar, Shounak Paul, Pushpdeep Singh, Ines Abdelaziz, Sayeh Jarollahi, Seungeon Lee et al. Content moderation policies are increasingly complex, and it is unclear whether foundation models can apply them consistently. The authors compare two ways of guiding a Vision-Language Model (VLM): an instruction-driven approach where the model reasons from written policy precepts, and an example-driven approach where it generalizes from prior moderation precedents, evaluated on ModerationBench, a new set of 4,000 manually annotated in-the-wild posts from Bluesky. Foundation models substantially outperform Bluesky's deployed moderation system, nearly tripling its F1 score (0.60 vs. 0.22) on random posts, and the two guidance paradigms reach comparable peak effectiveness.
70 more specialized papers

Large Language Models 53

Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation

Ling Team, Ang Li, Ben Liu, Binbin Hu, Bing Li, Bingwei Zeng et al. cross-listed Ling 2.0 is a family of open, reasoning-oriented Mixture-of-Experts (MoE) language models that scales from 16B to one trillion total parameters under the same high-sparsity design, guided by empirical scaling laws. The recipe combines a sparse MoE with multi-token prediction, reasoning-focused pretraining data with mid-training chain-of-thought activation, reinforcement-based fine-tuning (DFT and Evo-CoT), and full FP8 training on fine-grained heterogeneous pipelines. The three instruct models, Ling-mini-2.0, Ling-flash-2.0, and Ling-1T, reach up to 7-fold active-compute efficiency over dense counterparts, and the authors claim the trillion-parameter model sets a new Pareto frontier of reasoning accuracy against compute, serving as the base for the thinking-oriented Ring series.

From Plausible to Actionable: A Position on LLM Self-Explanations

Elize Herrewijnen, Benedetta Muscato, Gizem Gezici, Fosca Giannotti cross-listed Large language models can produce natural-language rationales for their own outputs, but whether these self-explanations faithfully reflect the model's underlying computation remains unresolved. This opinion piece argues from a traditional explainable AI (XAI) perspective that self-explanations can be highly plausible, questionably faithful, and still highly actionable, and it lays out practical guidelines for assessing plausibility and faithfulness given the limits of standard evaluation protocols. The authors propose adding actionability as a third evaluation criterion and highlight uses of model rationalization that support informed decisions across different stakeholders.

X-CoSD: Communication-Efficient Cross-Vocabulary Collaborative Speculative Decoding

Jaeduk Lee, Wan Choi Collaborative speculative decoding (CoSD) lets an on-device small language model (SLM) draft tokens that a server LLM verifies, but existing methods assume a shared vocabulary and exchange full token distributions during residual resampling, which is costly in communication. X-CoSD introduces hybrid resampling that splits the residual resampling step between the common-vocabulary region on the device and the LLM-only region on the server, so distributions are transmitted only for the common region; X-CoSD-E goes further by having the server send only sampled replacement candidates and their probabilities for local verification on the device. Both variants provably preserve the server LLM's output distribution, and experiments show substantially faster token generation with quality comparable to the server model alone.

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

Anirudh Malik, M Sparsh Mehra, Poojith Devan A nominal '1.58-bit' label says nothing about the deployed representation or its execution cost, so the work characterizes scaling an aggressive post-training ternarisation pipeline from Qwen3-4B to Qwen3-8B end to end. The pipeline combines KOTMS rotation, E2M-ATQ adaptive ternarisation, and GPTQ-style error compensation in a weight-only configuration with 16-bit activations, and the contribution is the characterization: an external reproduction gate, matched 4B/8B capability analysis, cross-corpus perplexity, effective-bit accounting, lossless lattice-aware packing, and direct packed execution. The 8B model shows a 1.361x perplexity ratio across WikiText-2, C4, and PTB, and on eight zero-shot tasks it scores 64.6% versus 72.4% for FP16, a 78.5% chance-corrected retention that beats the matched 4B run by 8.9 points. The packed checkpoint occupies 8.24 GiB and runs directly at 15.52 tokens/s in 7.35 GiB, though the preliminary packed matrix-vector kernel is still slower than FP16 cuBLAS.

Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

Dohyeon Kim, Bedionita Soro, Sung Ju Hwang Sparse Mixture-of-Experts (MoE) models normally route every token to a fixed top-k set of experts, and inference-time dynamic routing that activates fewer experts saves compute without retraining but shifts the model away from its training-time configuration. The analysis shows that reducing the number of activated experts consistently inflates the RMS scale and variance of MoE layer outputs, a representation mismatch that degrades downstream performance beyond the loss of expert capacity itself. Layer-wise Distribution Alignment (LDA) is a lightweight correction that uses per-layer calibration statistics to realign reduced-routing representations with the default configuration. Across multiple sparse MoE language models, benchmarks, and routing strategies, LDA recovers much of the performance lost to the distribution shift while keeping sparse-inference efficiency at negligible overhead.

Talking to Itself While Coding: What Makes Comments Help Code Generation?

Dangfeng Pan, Zhensu Sun, Cenyuan Zhang, David Lo, Xiaoning Du cross-listed Large language models (LLMs) frequently write natural-language comments while producing code, and those comments become context for the code that follows, yet which properties of comments actually help remains unclear. On LiveCodeBench, neither comment frequency nor broad comment intent reliably predicts pass@1, so the authors run controlled interventions that prefill weaker recipient models with comment blocks written by stronger source models to separate surface form from conveyed solution content. Comments from source solutions that pass the tests raise recipient pass@1 by 17.2% on average, comments describing failed solutions give no reliable gain, and comments written for a different problem cut pass@1 by 20.8%. Across many models and prompt variants, most recipients cannot recover this gain through prompting alone, with the best case recovering only 24%, indicating that comments help because they carry correct solution content rather than because they are comments.

Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding

Fengxiang Bie, Yuqing Jian, Yifan Yu, Zhongzhu Zhou, Zelei Shao, Ben Athiwaratkun et al. Speculative decoding drafters are usually trained against a narrow distribution for one target model, so their acceptance rates collapse under workload shifts, and pretraining has been hard to apply because existing recipes consume the target's hidden states and distill on its logits. Osprey instead bootstraps drafters from off-the-shelf pretrained small language models, pruning them to a shallow backbone, restoring language-modeling ability with target-agnostic next-token pretraining, and then adapting to each target through vocabulary alignment, zero-initialized QKV expansion, and distillation from the target's output distribution. A single pretrained backbone transfers across targets, improving mean acceptance length by 16.1% for Qwen3-8B, 21.2% for Llama-3.3-70B-Instruct, and 22.7% for the 229B MiniMax-M2.5 with 17.5% higher tokens per second, with the largest gains on out-of-domain and multilingual data.

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

Sanghyeok Park, Minji Kang, Hosung Kwak, Jinhyuk Yun Standard multilingual benchmarks reward selecting correct answers rather than testing whether models genuinely verify facts. Systematic Wikidata-based Object-Relation Distortion (SWORD) generates syntactically well-formed but factually incorrect statements in eight widely spoken languages by perturbing Wikidata triples, ranging from random entity substitutions to semantically plausible property-based swaps, and checks whether models reject them consistently across languages. Models score higher on plausible distortions than on nonsensical random ones, suggesting reliance on distributional familiarity rather than verification, and models with comparable baseline accuracy across languages degrade sharply on East Asian languages when given distorted statements, with cross-lingual gaps reaching 28 percentage points, a 49% relative reduction, in some models.

What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores

Dana Paquin, Riddhiman Jain cross-listed MMLU is widely used to calibrate general AI capability, yet its aggregate score is shown to measure factual retrieval more than reasoning. Calibrating item difficulty with Item Response Theory across 1,000 open-weights models and 14,042 test items, then regressing difficulty on a deterministic text-extractable framework of structural complexity with a joint Wald test and subject-clustered covariances, reveals that the mapping from complexity to difficulty differs between STEM and non-STEM partitions, so a single test conflates separable constructs. Aggregate leaderboard ranks track non-STEM accuracy more closely than STEM accuracy, and selecting a Top-50 model on the aggregate for a reasoning-intensive deployment displaces roughly 22% of the STEM-appropriate choices. After controlling for the multiple-choice guessing floor inside the response model, higher-ability models still degrade more steeply with reasoning depth, and the framework is released as a reproducible auditing instrument alongside a recommendation for disaggregated reporting.

From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls

Hamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi Nia In-vehicle assistants need small language models (SLMs) that map natural-language requests to vehicle function calls on constrained hardware, and a central design choice is whether each function is represented by a dedicated Functional Token (FT) learned during training or by a schema placed directly in the prompt (Schema-in-Prompt, SIP). The authors build a benchmark of 9,822 single-turn requests over 79 functions derived from Android Automotive, including held-out functions and requests that should be refused, and fine-tune four SLMs from 270M to 1.7B parameters under both representations. On functions seen in training, the 270M model matches the 1.7B model and the best overall results come from the 0.6B model, while on held-out functions FT scores zero by construction and SIP generalizes with gains that grow with scale. SIP also refuses out-of-scope requests more reliably, whereas FT can emit a function that is not actually available, at the cost of higher memory use and latency that a theoretical analysis attributes to longer schema contexts.

PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling

Weisi Yang, Stephen Xia Running large language models (LLMs) on mobile and edge devices improves privacy and latency, but the compute load causes heat buildup on fanless hardware that leads to thermal throttling, and existing dynamic voltage and frequency scaling (DVFS) governors, even ones tailored to LLMs, only tune hardware parameters and fall short under thermal constraints. PELM builds on the observation that not every token needs full-depth inference and adds two workload-level knobs to DVFS, speculative decoding and variable verification depth, turning power governance into a multi-dimensional optimization. In evaluations across hardware platforms and datasets it delivers up to 23.1% speedup and 52.4% lower energy consumption versus state-of-the-art power governors while keeping task performance comparable, and the source code is released.

PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations

Hyojeong Yu, Hyukhun Koh, Minsung Kim, Yunah Jang, Kyomin Jung Large language models (LLMs) serving as long-term personal assistants cannot efficiently rely on full interaction histories, which motivates memory systems, yet existing conversational memory evaluations focus on retrieval and factual recall rather than the practical guidance users actually ask for, such as recommendations, planning, and decision support. PRAGMA is a benchmark of curated longitudinal conversation histories with evidence annotations and guidance scenarios grounded in evolving user contexts and incorrect user assumptions, requiring models to integrate information across multiple past conversations and reason about changing preferences. Experiments across retrieval systems, memory systems, and long-context models show that current systems struggle both to recover the relevant conversational evidence and to use it for personalized guidance, pointing to the need for memory architectures that support memory-grounded reasoning beyond recall.

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar, Medina Maloku Large language models are increasingly used to audit documents for errors, but their reliability at this task is largely uncharacterised. The authors plant 450 contaminants of three kinds (typographical corruption, semantic reversal, and absurd out-of-context insertions) into 150 academic papers and test whether Gemini 3.0 Pro can recover them when prompted with a single document, a small batch, or a large batch. Recovery falls from 50% on single documents to 2.8% on large batches, and the failure mode is fabrication rather than abstention: the model confidently reports contaminants that appear in no document, such as a 'telepathic squirrel'. Plausible corruptions were missed far more often than absurd ones, and the authors recommend bounded batch sizes plus mechanical verification of every reported finding against source text.

EFQ-Softmax: Exp-Free Quantization for Softmax

Haohui Han (Xi'an Jiaotong University), Yuming Wan (Huawei Technologies Co., Ltd), Hongni Wang (Shandong University of Finance and Economics), Pengcheng Xie (Huawei Technologies Co., Ltd) et al. Low-bit attention runs the QK-transpose and PV matrix multiplications on FP8 or FP4 engines, but softmax typically still computes exponentials in higher precision, materializes a probability block, and quantizes it afterward, creating a mismatch between the high-precision producer and the low-bit consumer. EFQ-Softmax maps shifted attention scores directly to block-scaled E2M1 codes by choosing an exponent-only scale per microscaling block and applying a single affine rule, using the same operand for both the numerator and denominator updates while leaving the FlashAttention-style running-max and rescaling logic untouched. On Qwen3-8B and Qwen3-VL-8B-Instruct it slightly improves task means over MXFP4, keeps WAN2.2-TI2V-5B video quality comparable to FP16 under VBench, and cuts vector-stage latency of the fused probability-generation kernel by 40.33% on average for sequence lengths from 16K to 128K on the A5 vector unit.

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi Retrieval-Augmented Generation (RAG) concatenates many retrieved chunks into a long prompt, which inflates prefill work and time to first token (TTFT); reusing precomputed key-value (KV) caches per chunk cuts that cost, but it is unclear whether answer quality survives at very long context lengths. The proposed approach both fine-tunes the model to expect concatenated KV caches and selectively recomputes a subset of the caches at inference time. On RULER with 124k-token inputs the combination improves the score by 9.7 points over recomputation alone, while TTFT falls by 80% relative to full attention.

Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format

Touchapon Kraisingkorn, Krittin Pachtrachai, Wachiravit Modecrua Language models fine-tuned on customer behavior can both output a probability and write an explanation, and practitioners often treat the two readouts as interchangeable. Holding checkpoint and prompt content fixed, this study compares probabilities read directly off answer tokens against predictions generated after a written rationale, across 13 model-domain cells covering four retail tasks in three markets. The scored readout ranks outcomes better in 12 of 13 cells, by 1.5 to 14.5 points of area under the receiver operating characteristic curve (AUC), with the gap widening to 13.7 points under rationale-format supervision. Analysis of roughly 9,000 rationales links the deficit to reduced reliance on the dominant predictive feature and convergence on stock phrasings; a third format that elicits a probability before any verdict sharply improves calibration, cutting Brier score from 0.47 to 0.15, leading to a recommendation to keep rationales for explanation but source rankings from the scored head.

Forward-Free LLM Depth Pruning via Weight Redundancy

Vincent-Daniel Yun, Woosang Lim Depth pruning cuts LLM inference cost by deleting whole Transformer blocks, but activation-based selection needs forward passes over calibration data while existing calibration-free methods score each block in isolation and never compare blocks to each other. Weight-Redundancy Pruning (WRP) estimates inter-layer redundancy purely from checkpoint weights, comparing attention output and MLP down-projection matrices across layers and combining pairwise similarities with relative projection-scale information into an all-pairs matrix that guides layer grouping and block selection. Across multiple pruning ratios, model families, and downstream tasks, WRP consistently beats forward-free magnitude pruning and comes close to activation-based methods while requiring no data and no forward passes.

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

Joana Rosa, Pedro Santos, Valdemar Oliveira, Rom\~ao Silva, L. Miguel Silveira, Bruno Martins Translating natural language planning descriptions into Planning Domain Definition Language (PDDL) problem instances is commonly scored by whether the output parses and a planner solves it, which can badly overstate faithfulness since a solvable problem may still misstate the initial state, goal, object structure, or optimization target. The pipeline studied here chains LLM generation with parsing, planning and validation checks, a domain-conformance checker, an LLM critic, and iterative repair, building repair feedback from the domain description, the generated problem, the original natural language text, and operational diagnostics. Post-hoc analysis compares outputs to curated reference problems using renaming-invariant structural matching and semantic equivalence. Across Planetarium, AutoPlanBench, and curated PDDL 2.1 problems, operational success and reference reconstruction diverge substantially, and PDDL 2.1 remains hard to reconstruct faithfully even as operational success improves under structured repair.

Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses

Olli Tuomi Projecting a transformer's hidden state through the unembedding to read it in token space is cheap but untrustworthy, because at intermediate layers the output is dominated by generic tokens the model would predict for nearly any input. Reading the difference between two closely matched prompts cancels that shared component and surfaces only what distinguishes them, an operation equivalent to viewing a representation-engineering or activation-addition steering vector through a logit lens. Built into a training-free tracer that reads every position, sub-layer, and head while averaging over designed baselines, it traces a compound-noun chain from MLP to attention in Phi-2 that activation patching confirms, and reads distinctions such as real versus fictional entity retrieval and metaphor as domain-to-domain mappings. A cross-seed control marks the method's limit: across five networks differing only in initialization, the same distinction surfaces as almost entirely different tokens, with top-10 overlap of 0.08, meaning the token-space appearance of a computation is network-specific even when the distinction it draws is not.

Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning

Zhirayr Hayrapetyan, Andrei Kalmykov, Denis Kokosinskii, Dmitry Stanishevskii, Dmitry Zmitrovich Financial text is plentiful, yet little of it is usable for reasoning-focused post-training because question-answer pairs often lack explicit reasoning, context, or verifiable answers. The described pipeline mines open reasoning traces, distills instruction data, and generates knowledge-graph-guided pairs from textbooks, then uses semantic deduplication and three lightweight classifiers to keep finance-relevant, well-specified items and flag those suitable for reinforcement learning with rule-based verifiers. On FINESSE-Bench, ordinary supervised fine-tuning cost 3.2 to 4.0 accuracy points while self-distilled fine-tuning gained 1.0 to 2.8, equal-weight model merging recovered 3.0 points over its fine-tuned parent, and GRPO added 0.4 points after self-distillation or 3.0 points applied directly to verifiable tasks. The takeaway is that retention-aware adaptation avoids the capability regressions that plain fine-tuning introduces.

ProbPlug: A Plugin Uncertainty Network for Reliable Confidence in LLM Binary Classification

Jianzong Wang, Chuhang Liu, Botao Zhao, Zuheng Kang, Xulong Zhang, Xiaoyang Qu et al. Large language models classify well but give poorly calibrated confidence, which blocks deployment where a wrong answer is costly. ProbPlug is a small add-on that predicts whether a binary classification output is correct, using a self-attention module over internal token representations pulled from a frozen base model, so it slots into the existing inference pipeline without touching the model itself. Across text-only and multimodal tasks it yields better-calibrated confidence estimates and slightly improves classification accuracy at negligible extra cost, and transfers across tasks; code is released publicly.

If It's Not Buggy, Don't Fix It: On the Dynamics of Iterative Bug-fixing with LLMs

Xietao Wang-Lin, Anton Isopoussu, Louis Mahon cross-listed Automated program repair tools driven by language models are now routine in code review, so the authors examine what happens when such a model is applied blindly and repeatedly. Across several models and repair setups, the models consistently report bugs in programs that contain none, and break correct programs more often than they fix broken ones. Run iteratively, the process often settles into a pseudo-bug-fixing cycle in which the same edit is applied and reverted indefinitely. Mechanistic probing uncovers a steering vector controlling editing propensity, suggesting an internal representation of "buggy code" that is spuriously activated, which bears on stopping conditions for fully autonomous repair agents.

CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts

Naibin Gu, Qingyi Si, Chenxu Yang, Chuanyu Qin, Junhao Zhou, Peng Fu et al. On-policy distillation gives dense token-level supervision on the student's own generations and works well when teacher and student share a family, but degrades across families even after tokenizer alignment, with stronger external teachers adding little. Decomposing the cross-family signal shows it mixes an offset between a weak teacher-family reference and the student with the within-family log-likelihood shift from that reference to the strong teacher, and the offset dominates the update direction. CompassOPD drops the offset, transfers only the within-family shift, and anchors updates with a frozen student reference, raising average reasoning accuracy by up to 5.50 points over standard cross-family distillation across three student families and multiple teacher families. For a mixture-of-experts teacher, reducing expert activation yields the reference directly from the teacher checkpoint, removing the need for a separate reference model while retaining a 3.43-point gain.

From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

Christoph Wigbels, Ali Abusaleh, Markus T. Jansen, Alexander Mehler, Markus J. Hofmann Treating a person's browsing history as an individual text corpus, the authors ask whether episodic and semantic memory can be simulated by writing that corpus into a small language model rather than retrieving from it. They crawled the search histories of 515 participants who answered 36 multiple-choice knowledge items, analyzed a stratified subsample of 150, and trained one DoRA adapter per participant on a model whose baseline accuracy sat below the participants' lowest quartile. Each adapter fits its own participant's held-out text better than other participants' text, with an effect size of dz = 1.27 that grows with corpus size, confirming the corpus is genuinely written into the weights. On the general knowledge test, however, the adapter adds knowledge without aligning to the individual: log-loss match improves but bias-corrected answer match does not, and retrieval adds nothing further.

Through the Looking Glass: Directly Reading and Writing Transformers

Mark Oskin Counting transformer components by the absolute value of their logit contribution suggests thousands participate in each token prediction, but the authors argue contributions are signed and largely cancel: across eighteen models the mass pushing away from the predicted token is a median of seven times the mass carrying it. Dividing by the net, on the baseline model 53 components carry ninety percent of a prediction, 13 are indispensable, and 8 suffice to produce it alone, and across twelve models from 124M to 7B parameters the full backward trace touches only one to three percent of the model regardless of size. Everything is read directly from parameters and activations with nothing trained or fitted, naming each component by what it writes and what it reads, with input identification at 58.9 percent above chance. The same access supports edits: a new association installs into one spare unit at a fortieth of the held-out loss cost of a rank-one update, and installed heads can make an edit fire only when a token appeared earlier in context.

$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han et al. Existing coding benchmarks test isolated kernels, predefined operators, or fixed optimization targets, so they cannot tell whether large language models (LLMs) can do open-ended, long-horizon engineering on the infrastructure that serves them. Φ-Bench derives tasks from optimization problems studied in frontier research and grounds them in real code repositories, spanning localized kernel-level function completion through long-horizon implementation to end-to-end system optimization across the LLM infrastructure stack. Experiments on frontier LLMs expose substantial remaining limitations in engineering complex LLM infrastructure, and the authors use the results to characterize the challenges on the path to autonomous optimization of AI infrastructure.

The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs

Arquimedes Canedo A graph retrieval-augmented generation pipeline makes four prompt-construction choices: which triples to include, their syntax, their order, and the grounding instruction telling the model what to do with them. Varying all four across six large language models and two knowledge-graph question answering benchmarks, with subgraphs drawn from gold SPARQL rather than a retriever, the authors find only two choices matter. Removing the answer path costs most of what the graph was worth, while replacing every off-path triple with material from an unrelated entity shifts F1 by only +0.003, implying retrieval budget belongs on recall rather than precision. With an empty context, instructing the model to use only the provided facts drops F1 from 0.299 to 0.035, and applying that instruction to the context arm but not the baseline manufactures a spurious finding that graph context hurts at depth, which the authors found in their own results and retract.

LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented Generation

Daniel Alejandro Coll Tejeda, Pedro Garc\'ia L\'opez, Daniel Barcelona-Pons cross-listed Graph-based retrieval helps multi-hop question answering but often incurs high query-time cost from LLM-driven traversal and produces diffuse, oversized contexts. LiteRAG replaces retrieval-time LLM control with query-conditioned algorithmic graph exploration and reasoning-chain context construction, relying on query-adaptive thresholding and community-aware hub penalization. On DistComp, a multi-hop benchmark over distributed-systems papers, it attains the highest overall quality among evaluated methods at 0.798 while cutting per-query latency by over 100x and cost by over 99% relative to GraphRAG Global and DRIFT, and on UltraDomain it matches LinearRAG on quality with about 14x fewer tokens.

Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation

Ben Merbaum, Mohammad Amin Raeisi, Wenhao Wang, Charalampos Papamanthou, Katerina Sotiraki, Fan Zhang cross-listed Open-source large language models (LLMs) let users keep inputs private by running locally, but large models exceed most local hardware, and outsourcing inference reintroduces privacy and correctness concerns. Maverick delegates matrix-vector multiplication, the dominant LLM operation, to an untrusted server using what the authors describe as the first information-theoretically sound verification protocol for this task with transparent preprocessing, efficient batch verification, and virtually no server overhead, combined with LPN-based pseudorandom masking for input privacy. In an end-to-end prototype on Qwen3-4B with one client thread and a 128-thread CPU server, throughput improves over local inference by up to 17x with online mask generation, 45x with precomputed masks, and 44x with verification alone, and with four client threads the gains are 13x, 18x, and 17x. Client-side microbenchmarks with simulated network delay, where server compute is no longer the bottleneck, show speedups from 12x up to 157x depending on configuration.

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

Xi Shi, Qian Lou LLM serving systems reuse key-value (KV) caches only when the reused text sits at the very start of the prompt, a condition broken by retrieval-augmented generation servers assembling different chunk sets per query and by multi-agent coordinators reading other agents' reports, since a reused cache carries wrong positions, never attended to the other sources, and may have been written by a different checkpoint. KVShareArena benchmarks cache-repair methods from three separate research communities across prompt contexts and model checkpoints, scoring each by the fraction of the gap recovered between no cache and full recomputation while charging compute, memory, and per-request latency. Position correction alone, which needs no recomputation, suffices until a question needs several sources at once; there only methods that re-encode part of the cache or train recover half to two thirds of the gap, and unrepaired caches can be worse than no cache. Cache-compression methods harmless on a single prompt fall well behind position correction on fresh agent reports, caches written by a different checkpoint barely affect training-free methods but degrade a trained adapter, and the harness ships as a pip package with a public leaderboard.

RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding

Fang Li Language models under one million parameters matter for edge deployment and reproducible research, yet at embedding width 128 a two-layer LSTM or Transformer spends roughly a third of its capacity on the output projection matrix. Riemannian Language Models (RiLM) remove that layer entirely: context evolves as a trajectory on a Riemannian manifold and next-token probabilities come from squared geodesic distance between the current state and the vocabulary embeddings, so a single embedding map serves both input and output. Instantiated on flat Euclidean space (Flat RiLM) and the Poincaré ball (HypRiLM) with about 290k parameters and a 2,000-word vocabulary, HypRiLM reaches 54.2 validation perplexity on WikiText-2 versus 87.6 for Flat RiLM and 113 to 147 for tied, parameter-matched LSTM, Transformer, and state-space controls. Results on Penn Treebank and a 10k-vocabulary stress test show geodesic decoding transfers across corpora while hyperbolic curvature helps selectively, the authors characterize boundary collapse in naive hyperbolic recurrence and a Möbius stabilization that restores trainability, and claims are scoped to controlled small-model comparisons rather than full-vocabulary state of the art.

Retrofitting Code Using LLMs to Support Exceptional Behavior

Linghan Zhong, Jiyang Zhang, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric cross-listed Exception Related Code (ERC), meaning throw statements, the guard conditions around them, and try/catch blocks, is essential to robust software but tedious to write across large codebases. The authors define a new task, retrofitting existing code with ERC so that given Exceptional Behavior Tests (EBTs) pass, and build EXCODER, which performs context engineering by feeding static and dynamic program analysis results to a Large Language Model (LLM). On a benchmark of 304 methods from 75 GitHub Java projects with ERC systematically removed, EXCODER with Qwen 2.5 Coder 32b reaches 85.92% pass@1 on developer-written test suites, 12.56 percentage points above the baseline, with pass@5 and pass@10 of 86.18% and 86.51%. Manual inspection of the generated code surfaces remaining limitations and directions for future work.

ConvMem: Convolutional Memory for Long-Context Reasoning

Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang et al. Large language models struggle with extremely long inputs because of fixed context limits, and sequential memory approaches like MemAgent that read text segment by segment incur high latency and require costly reinforcement learning training that can overfit to specific datasets. ConvMem is a training-free framework that reformulates long-context reasoning as a hierarchical convolution: a query-prompted LLM acts as a convolutional kernel that summarizes text segments level by level, turning a linear reasoning chain into a logarithmic-depth tree. Configurable strides and skip connections help capture and propagate evidence, while multi-kernel convolution decomposes complex queries into separate semantic channels, and the whole process parallelizes across both segments and reasoning threads. On RULER-HotpotQA and RULER-2WikiMultiHopQA it outperforms training-free baselines and avoids the overfitting to parametric priors seen in RL-trained models on out-of-distribution tasks.

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan Enterprises deploy full serving systems rather than model checkpoints, yet all 18 benchmarks the authors audited score advertised model identifiers, ignoring how weights, serving route, precision, output contract, and harness jointly determine usable capability. The IB2 protocol treats this as measurement error and makes it reportable through three parts: a gold-blind preflight that verifies a route can execute the evaluation contract, a first-pass scoring rule that keeps failures in the score, and structurally score-blind adjudication, with a sealed reference suite of 128 tasks and 987 assertions over document, spreadsheet, chart, tool, and database work. Across eleven systems, two complete runs on identical weights later failed distinct predicates of the binding gate that the advertised identifier did not expose, four of seven suites saturated within a six-system band so results are reported as interval-backed resolution groups rather than ranks, and switching serving arm moved one declared revision and precision from 77.38 to 82.54. Excluding failed responses from denominators changes the point ordering, so including reliability alters conclusions rather than just wording.

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, Arman Cohan A research idea can be novel and coherent while its method description still omits details that an implementer or coding agent would need, forcing unsupported assumptions during implementation. IdeaAMBIG collects 660 evidence-grounded instances, 163 real gaps mined from reproducibility reports and GitHub issues plus 497 synthetic gaps injected into implementation-ready specifications, and tests three capabilities: judging whether a specification is ready to codify, localizing the defect, and generating a clarifying action once the defect is known. Across 13 large language models, the best model recovers only 9.6% of real-world defects yet succeeds 80.6% of the time at proposing clarifications when told where the defect is, and supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%, pointing to defect localization as the main bottleneck.
18 more specialized papers

Theory 37

What Fixed-Rollout pass@k Evaluations Can Identify

Pranav Singh, Prashant Singh cross-listed Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples actually drawn per problem. Under the pooled conditional-Binomial model, fixed-n success counts identify only the first n moments of the latent per-task success distribution, so direct pass@k is identified only for k <= n, and extrapolated pass@k, tail exponents, and tail constants are not identified beyond n even with unlimited tasks at the same rollout budget. The authors give exact constructions with identical count laws but incompatible extrapolations, compute sharp identified intervals via Hausdorff principal representations, and show on the public 10,000-rollout release of Brown et al. that counterfactual n = 16 evaluations leave failure at k = 1000 ambiguous by factors from 1.5 to over 2,600 across MATH, GSM8K, and CodeContests configurations. They propose a conservative finite-task confidence certificate and a reporting standard that separates direct estimates, identified sets, and model-conditioned forecasts.

Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization

Nour Jamoussi, Marios Kountouris Divergence-based regularization and Sharpness-Aware Minimization (SAM) both aim to improve generalization through robustness to perturbations, but their relationship had not been examined. Using second-order expansions of f-divergences, the analysis shows the two are locally consistent under parameter-space perturbations, with divergence regularization yielding a Fisher-weighted quadratic penalty and SAM penalizing the dominant Hessian eigenvalue, a correspondence that becomes transparent for negative log-likelihood objectives with exponential-family outputs where the Fisher and Gauss-Newton matrices coincide. The same view extends to input-space perturbations, where the regularizer induces a pullback quadratic form on the input space that generalizes standard SAM. Using the asymmetric alpha-skew Jensen-Shannon divergence as a testbed, whose curvature coefficient scales as alpha(1-alpha), accuracy and negative log-likelihood on four benchmark datasets are consistently best near alpha = 1/2, the point of maximal curvature penalization, and loss-landscape visualizations link stronger curvature penalties to flatter minima.

High-probability guarantees for linear accessibility in feature superposition

Enrico Vompa cross-listed Neural networks use superposition to pack more features than dimensions, but interference between simultaneously active features limits how many can be read out linearly. By casting linear accessibility as a compressed sensing problem, the authors derive high-probability bounds for fixed supports under subgaussian noise. The sufficient dimension scales linearly, d = O(k log m), rather than at the quadratic worst-case limit of prior work. The bounds are validated across system parameters with Gaussian-tail approximations and offered as a framework for evaluating sparse autoencoders, compositional generalization, and interpretability claims.

Learning with Synthetic Data via SGD in High-Dimensional Linear Regression

Jichu li, Difan Zou cross-listed Training on synthetic data can cause strong model collapse, where any fixed fraction of synthetic data leaves a non-vanishing excess-risk floor no matter how much data is added. The analysis studies one-pass stochastic gradient descent in high-dimensional linear regression with a shift between the synthetic and real data-generating models, deriving finite-sample risk bounds for mixed training and for two-stage training that uses synthetic data only in a first stage. Mixed training provably induces the collapse floor while two-stage training avoids it, scaling laws under a random sketch model show larger models can amplify synthetic-induced degradation under mixing, and an exact necessary-and-sufficient condition characterizes when two-stage training strictly beats real-only training at the same real-data budget.

When Does Low-Bit Quantization Preserve the Decisions of Vector Search?

Wenxuan Xiao, Xu Cao cross-listed Low-bit quantization of vector embeddings sometimes preserves retrieval recall and sometimes collapses, and neither average distortion nor global rank correlation predicts which. The analysis moves to the level of individual pairwise comparisons that ranking and graph-pruning algorithms actually consume, deriving a distribution-free bound in which comparison flip probability decomposes into exact margins near zero plus a calibrated residual tail, extended with covariance-aware moment identities for residuals sharing a query or graph node. A deterministic coupling theorem shows Vamana neighbour selection replays exactly when all candidate-level pruning actions agree on frozen exact states. Across learned, classical, and synthetic embeddings, standardized exact margins predict held-out flip rates substantially better than global rank correlation, and the framework covers binary codes, RaBitQ, Lucene BBQ, and product quantizers under one decision interface.

A Sharp Barrier for Consistent Submodular Maximization: Any Improvement over $2-\sqrt{2}$ Entails Exponential Queries or Linear Recourse

Shi Fu, Qixin Zhang, Dacheng Tao cross-listed In consistent submodular maximization, elements arrive over time and an algorithm must maintain a set of at most k elements while changing only a constant number after each insertion; Dütting et al. showed a tight 2/3 approximation with unrestricted computation and a polynomial-time 0.51, leaving open whether efficient algorithms can match the offline 1-1/e guarantee. The answer is no: the best approximation achievable with polynomially many value queries and worst-case constant recourse is exactly 2-√2 ≈ 0.5858, reached up to any ε by a randomized algorithm with O(ε^-2) changes per insertion, and any fixed improvement requires exponentially many queries before one critical insertion or Ω(k) recourse at it. The work also determines the exact curvature-dependent threshold 1-(√2-1)ϑ, attains 1-1/e-ε for weighted coverage with O(1/ε) recourse, and separates the existence of universal future-price certificates from their efficient computation.

A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram

Alexis D. Plascencia Sparse autoencoders (SAEs) are used to extract interpretable features from network activations, but systematic feature co-occurrence can cause distinct features to be absorbed or merged, and the MAIS-O43 open problem proposes a controlled synthetic-dictionary experiment to map when true-feature recovery gives way to merging as nesting fraction, sparsity penalty, and dictionary size vary. Running 200 independently initialized fits across ten of the 165 grid cells, plus 3,300 more fits over the full grid with standard minibatch Adam, the authors observe zero full-dictionary recoveries and zero merges. Every run instead lands in a reproducible diffuse phase where reconstruction is nearly perfect but learned atoms stay far from the true features (median best cosine 0.5 to 0.7 against a 0.95 recovery threshold) and learned codes are roughly ten times denser than ground truth. Since the exact sparse-coding objective is known to merge nested features at its global optimum in the two-feature case, the results suggest trained SAEs need not reach those minima and that the phase diagram of trained models may differ fundamentally from that of objective minimizers.

Algorithmic stability via ensembling

Rina Foygel Barber, Richard J. Samworth cross-listed Algorithmic stability describes how insensitive a learning algorithm is to perturbations of its input data, where the relevant kind of perturbation depends on the setting. The authors develop a general framework for quantifying how much any ensembling strategy defined through averaging yields stability guarantees under any type of data perturbation. The main result bounds the stability of the ensembled algorithm in terms of the norm of a covariance operator that describes the ensembling process. Applied to several practically relevant perturbation types, the framework gives interpretable insights and much sharper guarantees than those obtained from privacy-based arguments.

Learning with Covariance Matrices: Principal Component Analysis Meets Learning with Graphs

Saurabh Sihag, Andrea Cavallo, Elvin Isufi, Gonzalo Mateos, Alejandro Ribeiro Graph neural networks (GNNs) are often deployed on graphs built from pairwise statistical dependencies, yet existing GNN theory treats abstract graphs and does not account for the data-driven nature of covariance matrices. The authors give a tutorial overview of coVariance neural networks (VNNs), GNNs that operate on covariance matrices as graphs, covering a conceptual equivalence between VNNs and principal component analysis (PCA)-based information processing, refined stability bounds under finite-sample covariance perturbations, and a characterization of transferability across multiscale datasets. These results are offered as justification for adopting VNNs over PCA-based pipelines wherever covariance matrices describe data structure, illustrated with a case study on estimating brain age gap in neurodegenerative conditions from neuroimaging data.

Characterizing Language Generation in the Limit: Finite Witnesses and a Separation-Width Hierarch

Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao cross-listed Language generation in the limit asks an algorithm to eventually produce valid unseen elements of an unknown infinite language from any exhaustive enumeration of its positive examples. The authors characterize when this is possible for arbitrary families over a countable universe: generation succeeds exactly when each target language can be assigned a finite positive witness such that the targets activated by any finite sample share an infinite common intersection, with necessity following from a normalization that converts any successful generator into one depending only on the observed set. They then define positive separation width, the smallest uniform bound on witness size, and show every level of the resulting hierarchy occurs, from singleton witnesses for countable families to unbounded finite witnesses for unions of families with infinite common cores. Countable-support and finite-profile obstructions explain why local combinatorial data cannot determine generability, and the characterization and full hierarchy are checked in Lean.

A positive resolution of the gap-entropy conjecture

P. M. Aronow, Nathan Kallus, Patrick Lopatto Fixed-confidence best-arm identification asks how many samples are needed to find the arm with the highest mean, with failure probability at most delta, when arms are independent unit-variance Gaussians with means in [0,1] and a unique best arm. The gap-entropy conjecture predicted that the optimal sample complexity depends not only on the usual hardness quantity H, the sum of inverse squared gaps, but also on an entropy term measuring how that hardness is spread across dyadic gap scales. The proof shows that, averaged over permutations of the arm labels, the optimal expected sample count is within absolute constant factors of H times (log(1/delta) plus the gap entropy), and it provides an instance-independent algorithm that matches this bound up to an additive term governed by the smallest gap.
26 more specialized papers

Other 36

Encrypt What Matters: When Selective Homomorphic Inference Is Efficient

Ali Backour, Juan Reyes, Jaime Punyed, Ana Onoprishvili cross-listed Fully homomorphic encryption (FHE) allows inference on private data without exposing it to the server, but evaluating an entire input under FHE is expensive. Selective homomorphic inference encrypts only a sensitive region of interest (ROI) and runs computations independent of that region in plaintext, producing the same output as full FHE on the same model without retraining. The speedup depends on how quickly encrypted dependencies propagate through the network: for small encrypted ROIs, locality-preserving architectures achieve order-of-magnitude homomorphic-evaluation speedups, while architectures with early global mixing gain essentially nothing, identifying locality as the key architectural property.

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

Shivam Singh, Aditya Yadavalli, Catherine Arnett, Alex Warstadt Multilingual speech recognition models such as Whisper underperform on languages that are sparse in their training data, and while monolingual fine-tuning is a known fix, it has only been applied to a handful of languages. BuzzASR scales this to all 102 languages in FLEURS, combining plain fine-tuning with a heavier adaptation recipe that replaces the tokenizer with a monolingual one and augments with text-only fine-tuning. The models beat Whisper-large-v3 on 77 of 102 languages, cutting character error rate by a factor of over 2.8 on average, and set open-source state of the art on 27 languages on the combined FLEURS and Common Voice test set. The tokenizer replacement raises compression by 3.3x on average and up to 21.7x, and all models, code, and results are released.

Distillation of Synthetic Data for Time Series Foundation Models

Niloy Biswas, Noureddine El Karoui cross-listed Time series foundation models (TSFMs) are increasingly pretrained on synthetic trajectories whose generating process is known, yet standard recipes still compare model outputs to the single realized future of each trajectory. Synthetic data distillation (SDD) instead compares outputs to the trajectory's exact conditional forecast distribution, which amounts to a Rao-Blackwellization of the training objective: the expected stochastic gradient is unchanged while its covariance is provably reduced under the Loewner ordering. Validated on a TSFM family from 4M to 2.5B parameters, SDD converges faster at every size, and on Gaussian Process data matches or beats the status-quo loss with 10% to 40% fewer training iterations.

Muon-C: Operator-Aligned Muon for Convolutional Kernels

Jiaxin Qing, Lexin Li The Muon optimizer replaces matrix momentum with an approximately orthogonal polar direction, but its geometry depends on how weights are laid out as matrices, and the standard unfolding of a convolution kernel describes a local patch map rather than the convolution operator itself. Muon-C represents kernel momentum as frequency-wise channel-transfer matrices, polarizes each block independently, and uses a critical Fourier grid to map updates exactly back to the original finite kernel support, and the resulting exact-polar direction is a linear minimization oracle under the critically sampled convolution norm with a worst-case guarantee never weaker than unfolding and strictly stronger for 3x3 kernels. On CIFAR-10 flow matching with matched applied-update RMS, Muon-C reaches 9.87 FID at 40k iterations versus 22.26 for unfolded Muon and 51.31 for Adam, matching their final quality with about 0.62 to 0.64 times the model FLOPs and reaching 3.42 FID under equal tuning budgets. Gains persist across data scales and transfer to classification on several convolutional architectures.

SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference

Ting Liu SymbolicLight V2 is a hybrid neuromorphic language model architecture that mixes sparse event-driven computation with continuous-state processing, extending the previous version with graded signed events and softmax-free local attention. A 194M-parameter model is implemented in fixed-point arithmetic on an Alveo U50C FPGA and with sparse integer execution on ARM CPUs; on the FPGA, active-row weight gathering and valid-state KV loading raise decode throughput from 474.6 to 643.2 tokens per second and cut estimated energy per generated token by 27.6%, with idle power accounting for 82.8% of card energy. Against a recorded RTX 5090 FP32 baseline the FPGA uses 89.1% less estimated energy during short-context decode, and four Cortex-A76 cores on a ROCK 5T reach 65.4 tokens per second at 9.80 W, though the deployed checkpoint's quality trails a same-budget dense model, so the results do not establish equal-quality efficiency.

One Loop, Two Gains: Can Active Learning win the Lottery for Free?

Benedikt Tscheschner, Eduardo Veas, Marc Masana Iterative magnitude pruning, the standard way to find lottery-ticket subnetworks, and pool-based deep active learning both retrain a model from scratch many times, yet the two have been studied separately. Improve & Prune (I&P) folds magnitude pruning into each active learning retraining cycle at practically no extra cost, and the authors ask whether winning tickets can emerge under the non-stationary data regime of active learning. Across multiple acquisition functions, architecture families, and image classification datasets, including an active fine-tuning scenario, I&P produces sparse, deployable models at every acquisition round that match dense accuracy at sparsities up to 95%. These per-round sparse models can relieve the two computational bottlenecks that limit deep active learning on large architectures and large unlabeled pools: retraining each round and scoring the pool for acquisition.

A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out

Mahdi Naser Moghadasi (BrightMind AI), Faezeh Ghaderi (University of Texas at Arlington) Time-series foundation models are evaluated almost entirely on public archives that predate them, so a strong score cannot be separated from having seen the test data during pretraining. The authors build a contamination-free hold-out: thirteen forecasters (four classical, three trained per dataset, six pretrained) on seven groups from five domains, with every observation published after the last model's release and every dataset rebuildable without an API key. Pretrained models win five of seven groups, lose one to a Theta baseline, and are indistinguishable from seasonal naive on daily exchange rates, yet neither seasonal strength nor spectral entropy of the input window explains the pattern; what tracks the wins is corpus familiarity, with the largest gain (28% lower mean absolute scaled error) landing on weekly Wikipedia pageviews, the domain that dominates TimesFM's pretraining corpus, and the TimesFM family outranking Chronos far more on Wikipedia than elsewhere. The conclusion is that a temporal hold-out removes memorization of a window but not familiarity with a domain, so benchmarks need domain hold-outs stated relative to disclosed corpora and practitioners should ask whether their domain is one the model was raised on.
29 more specialized papers

Agents 34

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

Aayam Bansal, Keertan Balaji Benchmarks for autonomous AI scientists score only final outputs, discarding the reasoning process and making it impossible to audit methodology or separate systematic reasoning from lucky guessing. OpenDiscoveryTrace is a public dataset of 558 complete agent trajectories over 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis, recording a nine-field trace per step (thoughts, tool calls, observations, errors, revision triggers, self-reported confidence) for three frontier models (GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro) and four open-weight models. A pilot analysis of 363 LLM-judged trajectories finds the frontier models have comparable success rates of 84 to 89%, yet Claude Opus 4.6 logs 30 times more errors per trajectory than GPT-5.4, mostly tool misuse versus mostly reasoning errors. Five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformers are released alongside the dataset, schema, and harness under CC BY 4.0.

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng cross-listed AI research agents draw on prior knowledge, web sources, and experimental feedback, which makes it hard to tell whether a reported result is a genuine discovery or a recovery of something already known. The Discovery Certification Protocol (DCP) converts such claims into executable tests: Gate 1 checks for useful improvement on a sealed evaluation, Gate 2 gives matched agents the registered starting information and observed web content while withholding the target research history to see whether they recover the result, and an optional Gate 3 measures the effect of truthful feedback against a neutral policy from a shared checkpoint. Two controlled audits on SQLite optimization and virtual catalyst control produced zero recoveries in 96 episodes each, with an upper bound of 0.0468 on the recovery probability, and paired studies yielded 30 truthful recoveries against zero neutral ones. A deterministic verifier that uses no language model reproduces the certification decisions from frozen evidence.

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

Wasu Top Piriyakulkij, Rachel Lawrence, Alicia Curth, Sushrut Karmalkar, Niranjani Prasad Agent skills package reusable knowledge as multi-file bundles of instructions, scripts, and resources that a language model agent loads into its context and follows, but this becomes brittle on long-horizon tasks as accumulated context degrades reasoning. The alternative studied here invokes each skill package as a subagent, spawning a fresh context window dedicated to that subtask instead of inlining the instructions into the main agent's context. Subagent execution outperforms in-context skill execution when skill packages expose clear input-output contracts and their instructions encode the procedural knowledge needed to meet those contracts, at the cost of extra tokens spent coordinating between the main agent and its subagents. The benefit of reusable knowledge therefore depends on how it is organized and invoked, not only on its content.

Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management

Antonino Vaccarella, Lanpei Li, Vincenzo Lomonaco, Massimo Coppola cross-listed Resource management across Internet of Things (IoT), edge, and cloud layers requires continuous context-aware decisions, and while deep reinforcement learning (DRL) suits this and large language models (LLMs) increasingly augment DRL pipelines, the architectural relationship between the two is rarely made explicit. The authors extend Wang et al.'s taxonomy of DRL-based Continuum Orchestration Systems with two dimensions: an AI Augmentation Paradigm capturing how LLMs are exploited, and a Feedback channel capturing whether and through which path execution feedback returns to the LLM to close the MAPE control loop. Applying the taxonomy to six recent architectures, they find that none combines full LLM orchestration with full agent-layer feedback in a cloud continuum setting, and attribute the gap to a missing cross-tier feedback abstraction that would bridge incommensurable per-tier signals and the LLM orchestrator.

The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

Bo Yan, Weikai Lin, Song Wang Agents that act through tools face libraries of thousands of interfaces, so a short ordered tool menu is shown before execution and the agent may only call tools in it. Existing menu constructors rank tools by request relevance, which surfaces the final action but can omit or delay the prerequisite tools that produce its inputs. State-Path Tool Menu treats the menu as an execution prior over pre-execution routes from the observable request state to the goal: an encoder captures which tools can run from the current state, how their outputs satisfy later inputs, and which orders recur in training paths, a retriever covers an executable entry point, missing-input producers, and the final action, and a reranker places producers before consumers. On ToolBench the menu raises online success from 0.737 to 0.898 without changing the agent, covers more complete chains with 32 tools than the official list does with 128, and the gain holds across executor families of different capacities.

Benchmarking Hybrid Deep Research Across Database Querying and Web Search

Ruofan Wu, Peiran Xu, Xiaolong Li, Fan Shu, Soyoung Yoon, Yite Wang et al. Deep-research agents are usually benchmarked on open-web browsing alone, but real analytical tasks often require combining ambiguous unstructured text with precise structured data and preserving constraints across the handoff between them. HybridDeepResearch contains 380 tool-dependent tasks grounded in LiveSQLBench-Base-Lite databases and public web corpora, each requiring both web search and SQL to reach a verifiable answer, validated through automated checks and human review, and covering three reasoning patterns (SQL2S, S2SQL, and Parallel). Even GLM-5.2, Claude-Sonnet-4.6, and GPT-5 reach only about 50 to 54% Pass@8 on the hard subset under various agentic scaffolds. Directional reasoning, where evidence from one system must constrain the query to the other, proves substantially harder than parallel intersection.

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

Priyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla Agentic workflows have complex failure modes across planning, tool invocation, and environment interaction, making it important to know how confident an agent should be that its actions will succeed. The authors test whether a model's internal representations carry stronger signals of eventual task success than its outputs, introducing Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action decision points. Across three interactive benchmarks (Bash, SQL, Python) and three model families (Qwen14B, Qwen7B, DeepSeek6.7B), both methods consistently outperform surface-level generation and sequence-based calibration baselines, providing a zero-overhead reliability monitor that needs neither prompt changes nor multi-sample rollouts.

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

Praphul Singh, Shanu Kumar, Akshat Agarwal, Ganesh Kumar When LLM agents execute procedures, a final answer can look acceptable even though the agent skipped a required check, branch, dependency, or invariant, and neither output-only nor trace-aware judging identifies which obligations were active for a given query. ContractEval represents procedural instructions as query-active obligations and matches them against response or trace evidence, classifying omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output-contract violations as distinct conformance failures. On a controlled suite of audited procedural contracts, output-only and trace-aware LLM judges miss many injected structural failures, while the framework detects and localizes all of them given gold expected and observed graphs. LLM-backed extraction of those graphs preserves much of the signal but is calibration-sensitive, so the authors position the tool as making conformance auditable rather than as a compliance guarantee.

Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization

Yi Wu, Zheng Ren, Zhiyu Hu, Haochen Wang, Daryl Chang, Li Wei et al. Unaided language models fall well short of classical optimizers on low-budget black-box optimization, so the authors ask whether an agent can learn a search strategy by writing and running optimizer programs, then hand that strategy to other models as plain text. During development the agent repeatedly writes and evaluates optimizer code, and the resulting program and practice record are distilled once into a 197-word text prompt, Harness A, which is frozen before evaluation. The harness cuts Gemini Flash regret by 48% in an independent N=30 study, reaches the Gaussian-process Bayesian optimization (GP-BO) performance range on the practice family, lowers mean regret on three held-out BBOB landscapes, and transfers to Claude Sonnet with 43% to 49% regret reductions. An independent replication produces a different program and text at the same performance tier, and the framework also attains the lowest regret on a sealed YouTube reward-tuning production benchmark.

An Efficient and Effective Agentic Group Shilling Attack on Recommender Systems

Quoc Viet Nguyen, Trinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen, Bay Vo et al. cross-listed Recommender systems are vulnerable to shilling attacks in which injected fake profiles distort item rankings, but existing attacks depend on target-specific fine-tuning or fixed profile templates, making them hard to adapt or easy to detect. The Agentic Group Attack System (AGAS) uses a central Coordinator that directs a group of worker agents to promote a target item across different victim recommender families, adjusting strategy when progress stalls or suppression signals rise, while workers switch between active and inactive roles to avoid repetitive patterns. Under matched attack budgets, AGAS consistently outperforms strong baselines in target promotion while better preserving benign recommendation quality, weakening representative detectors, and running more efficiently than prior attacks. The authors argue defenses must handle adaptive campaigns rather than isolated fake-profile injection.

Multi-Agent Agentic Graph Learning via Structural Signatures

Liang Qu, Jianxin Li, Hua Wang Agentic graph learning (AGL) has a large language model agent sample a graph step by step as evidence for a prediction, but existing single-agent and role-based multi-agent approaches share one reasoning policy across the whole graph, which suits graphs with heterogeneous structure poorly. MAAGL partitions the graph into communities and assigns each an independent agent with its own memory, represents structural evidence with a permutation-invariant, fixed-size structural signature that is updated dynamically, and filters semantic evidence to the top-k most relevant nodes to keep context bounded. Agents estimate confidence from past trajectories with similar signatures and trigger debate-style collaboration when uncertain, and the framework outperforms state-of-the-art AGL methods on four benchmark datasets.

CityPlanner: A Sandbox Agent for Executable Urban Planning

Wentao Zhang, Jingyuan Wang, Zetong Zhou, Yifan Yang, Wenrui Wang Urban planning is a spatial optimization problem requiring feasible actions from large candidate spaces under objectives like cost and service quality, and existing optimization and reinforcement learning methods depend on task-specific representations and constraint handling. CityPlanner places a large language model agent in UrbanSandbox, a file-based environment where it inspects task files, generates plans, runs evaluators, and revises based on executable feedback, and trains it with atomic-task reinforcement learning that splits long trajectories into a BuildPlan construction step and an ImprovePlan refinement step. On a real-world benchmark the agent consistently outperforms heuristic, task-specific RL, and general LLM-agent baselines, with ablations attributing gains to the sandbox, the atomic-task decomposition, and iterative deployment.

Who Are They to Each Other? Multi-Agent Reasoning for Speaker Relationship Inference

Yaohan Guan, Yen-Ju Lu, Yuzhe Wang, Junhyeok Lee, Jesus Villalba, Laureano Moro Velazquez et al. cross-listed Inferring the relationship between speakers in a spoken conversation is underexplored, and supervised models are costly to train while single-pass large language model (LLM) prompting gives little structure for weighing subtle, distributed, multimodal cues. The authors propose a training-free framework in which multiple LLM agents propose, challenge, and adjudicate relationship judgments, instantiated as Multi-Role Multi-Agent Debate, which assigns agents complementary roles or social-theory-grounded perspectives, and Multi-Agent Compete, which eliminates weaker candidates through pairwise adjudication. On the Seamless Interaction dataset, covering binary classification and fine-grained relationship-detail prediction across modality settings, both designs improve over zero-shot and existing multi-agent baselines in most cases. Human evaluation shows the task is hard for people too, and LLM methods can beat human annotators when text is available but fall behind in audio-only settings, suggesting current models do not fully capture acoustic cues.

RobustSGPO: Search-Space Control for Agent Harness Evolution

Zibo Zhao, Jijun Shi, Mo Zhou, Zhongyuan Wang, Shifu Bie, Yunfei Zhang et al. Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses from execution feedback, but its local update rule leaves the scope and type of each edit unspecified. RobustSGPO makes the requested edit explicit, constructs and checks the resulting patch, and continues the search from either the current incumbent or retained earlier snapshots. Evaluated in the AgentX brainstorming workflow over 120 tasks, 95 runs, and 7,350 candidate attempts, it raises completion on 30 held-out tasks from 60.0% to 80.0% and test quality from 3.77 to 4.14 under a 20-million-token budget, and a periodic 1-2-3 permission schedule beats a fixed maximum permission by 0.28 test-score points. Category-based snapshot retention limits degradation on source tasks after a task-family shift, while random retention reaches a higher endpoint on the destination family, both at a measurable retention overhead.

Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches

Kevin Hartman cross-listed When an agent writes code, the development framework acts as the control system for a non-deterministic worker, and spec-first frameworks such as GitHub Spec Kit, obra/superpowers, BMAD, and GSD all capture intent up front but differ in how they enforce engineering discipline. The authors characterize three enforcement modes: persuasion through prompt discipline the model may ignore, front-loaded structure followed by a trusted build, and controls the agent cannot edit, such as a deterministic orchestrator, human-approved gates, immutable tests, and a required green result against a live branched database. Consort is built on the third mode, with a deterministic orchestrator driving separate role agents through a spec-first design lane and a test-driven build lane on a live database branch. The claims that enforced tests and gates keep agent-written code honest and that specialized roles keep it maintainable are framed as a pre-registered, testable hypothesis rather than demonstrated results.

LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

Yujin Zhou, Mingxuan Zheng, Chuxue Cao, Huang Yidan, Jiale Chen, Yike Guo et al. Tool-augmented legal agents introduce agentic hallucinations in which tool-call and reasoning errors cascade into fabricated holdings and miscited authority, yet existing legal benchmarks only score single-turn answers and general agent benchmarks lack legal diagnostics. LexAgentHallu contains 3,414 instances across 17 legal categories and 6 task types, built through an expert-in-the-loop pipeline and annotated with a two-layer taxonomy of 7 high-level categories and 27 subclasses spanning substantive and procedural failures, with metrics that localize where along a trajectory each failure occurs. Evaluating 18 proprietary and open-source agents reveals a Right-Answer-Wrong-Reason effect, and hallucination subclasses cluster into distinct profiles by agent framework, legal task, and category, patterns that outcome-level evaluation cannot see.

Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks

Yanze Cao Procedural memory lets language agents reuse stored routines, but reuse assumes the routine still applies to the current task. The study pairs a retrospective, human-assisted interface-adaptation case from the BrowserGym TimeWarp WebShop V1 to V6 development history with controlled frozen-memory comparisons on synthetic shopping decisions, testing four mismatch types (changed quantities, a different evidence representation, local-versus-global optimization conflict, and distributed promotion evidence) across 32 cells using a local qwen3:8b at temperature zero. None of the predefined interference signatures appeared when current-task evidence was explicit and sufficient, identifying a tested region where mismatched procedural memory is not behaviorally disruptive, though the authors stress this establishes neither general safety nor a mechanism.

ROAM: Robust Organization of Atomic Memories for Agents through Semantic Relations

Jianjie Zheng, Peng Lai, Sijie Cheng, Jiehui Zhao, Lei Yang, Guanhua Chen Long-running language-model agents store fine-grained atomic memories, but as these accumulate they become redundant, overlapping, or contradictory, and existing managers ask an LLM to add, update, or rewrite entries in a single error-prone step that couples interpretation, storage decisions, and content generation. ROAM instead classifies each pair of incoming and stored atoms as independent, equivalent, directionally subsuming, or conflicting, assigns observations to active Primary or supporting Evidence roles, and fuses complementary details and temporal changes into compact views, retrieving only Primary views at answer time. Across models and evaluation settings this improves answer accuracy by up to 29.8 percentage points, with ablations showing complementary gains from each relation type and from fusion beyond role organization alone. Mechanism analysis attributes the gains to 15.6-point higher answer-critical source recall and an 11.5-point lower share of confounding tokens, and the method holds up across manager model scales.

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

Xing Zhang, Guanghui Wang, Yanwei Cui, Mengdie Flora Wang, Peiyang He Compound LLM systems typically resolve coordination by adding a higher-level LLM meta-agent that reads worker outputs, writes the final answer, allocates later calls, and decides when to stop, concentrating three control decisions in an opaque, order-sensitive model call. UnitBoost replaces that model with a defined operator: a task-given unit map converts worker outputs into slot-value proposals, a constrained argmax assembles the output, and unfilled or unsupported slots become an explicit residual for the next round, yielding order invariance, unit provenance, and a guarantee that unit-wise maximization dominates selecting any complete candidate when there are no coupling constraints. On three held-out benchmarks it beats the best single candidate chosen with gold labels by 0.060-0.195 absolute task-score points and input-matched generative managers by 0.048-0.076, and swapping only the management step improves six compound configurations by 0.013-0.182. Residual-directed rounds raise FanOutQA cell F1 from 0.4778 to 0.5524, and the analysis also characterizes three conditions where no gain is available and quantifies cross-unit coupling as a repair cost.

Can AI Agents Detect and Repair Artifact Drift in Network Experiments?

Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan et al. cross-listed AI agents operating on network systems should be judged not only on task completion but on whether the experiment record they modify stays trustworthy, a property the authors call artifact integrity: claims must remain supported by evidence, confined to that evidence's scope, and traceable through the encoding artifacts. NetArtifactBench tests whether agents can repair inconsistent records derived from public network-system artifacts while preserving still-supported claims, with 52 instances whose injected inconsistencies range from direct contradictions to unstated relations spread across several artifacts. Across 23 agent configurations on three general-purpose agent runtimes, the average contract pass rate over 5,980 outputs is 65.3%, but no runtime exceeds 30% when repair requires recovering implicit relations and propagating changes across artifacts, marking a sharp boundary between local correction and full record-level repair.

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary Enterprise agent evaluation is blocked by the fact that customer production data cannot be used for benchmarking and no synthetic substitute carries reliable ground truth. The Era by Eon Benchmark builds a complete fictional company from a seed plus attributes like industry and size, driving simulators of Salesforce, Zendesk, Slack, and Gong from one shared entity graph so that every expected answer can be computed exactly from the final records. A realism scorecard and adversarial detector validate the generated data, with mean realism rising from 61.8 to 97.0 across 23 companies and no records flagged as synthetic. Nine models answering the same 33 questions three times each scored between 42.4% and 76.8%, though only three of 36 pairwise differences survived multiple-comparison correction.

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

Shrey Nag, Sachita, Abhishek Kumar Singh, Lipi Goel, Rajeshwar Singh Janwar Agent benchmarks tend to score one slice of behavior, either task completion as in AgentBench or adversarial robustness as in AgentDojo and ASB, so a failure anywhere in planning, tool selection, execution, memory, or reasoning surfaces only as a failed task. AgentAudit reads a recorded execution trace without interfering with the agent and scores ten dimensions spanning instruction integrity, planner quality, memory, tool invocation and faithfulness, security, and execution integrity, then attributes each failure to a specific stage. Across nine capability and adversarial tasks, Claude Sonnet 5 and GPT-5 led on Composite Trust Score at 95.1 and 80.6 out of 100 while Sarvam 105B, Llama 3.3 70B, and Gemini 2.5 Flash scored 57.6, 45.7, and 22.6. The sharper observation is that models with similar task-completion rates diverge widely in trustworthiness, with several non-frontier models actively complying with unsafe requests rather than simply failing them; all traces were graded by a single judge that was itself among the evaluated models.

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

Remco Hendriks (Continker) MetroLLM-Bench tests language models as the decision layer of a transit ticketing kiosk across 955 cases from six real metro systems, spanning routing, fare calculation, service disruptions, accessibility, and adversarial input. Each case requires structured tool calls plus a machine-renderable final state with an outcome, per-ticket fare quote, and kiosk action, scored by fourteen deterministic components and eight semantic components, six of which use a model judge. Across twenty-six models, a 4B Qwen 3.5 student tuned with parameter-efficient fine-tuning beat both GPT-5.6 tiers on the deterministic tier (91.3 versus 90.6 and 90.0) while fitting in 2.6 GB at Q4_K_M quantization, and 9B and 27B students added nothing further at this training scale. The fine-tuning gain shrank monotonically with size, from +7.03 points at 2B to -0.91 at 27B, and a rule-based baseline already reached 84.6, leaving the model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning.

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

Arnab Chattopadhayay, Debdipta Halder Language model agents fail in recognizable ways when the environment is only partially observable, committing early on ambiguous feedback, collapsing onto a wrong hypothesis after one informative observation, and drifting as the interaction history grows. The diagnosis offered is structural: such an agent is a history-conditioned policy carrying no explicit belief over hidden state, so the Belief-State Engine maintains a Bayesian posterior over the latent states of a partially observable Markov decision process outside the model and shows the agent only that posterior, never the raw action-observation log. Given a four-axiom specification of belief consistency, the authors prove the paired system is a sound Markov policy on the induced belief MDP and therefore inherits classical Bellman optimality guarantees, conditional on the model never seeing raw history. On the Tiger benchmark and a red-team attack-graph task it improves return, belief calibration, and decision consistency over six baselines including Chain-of-Thought, ReAct, QMDP, and POMCP.

Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training

Junwon Ko, Dong-Jae Lee, Minchan Kwon, Sunghyun Baek, Junmo Kim Post-training agents on trajectory-level success labels teaches them to solve a task but says nothing about preserving the several distinct ways a task could be solved from the same decision point. Framing this as successful strategy coverage under a fixed rollout budget, Direct Diversity Optimization pairs Divergence-Tree Collection, which builds branch sets rooted at shared states, with a Reference-Relative Target-Odds objective that trains the model to match reference-relative targets across the successful alternatives. The method achieved both the highest task success and the broadest successful strategy coverage among compared post-training methods on BabyAI, BabaIsAI, and WebShop, and also recovered best after actions were replaced locally, beating both success-only imitation and decoding-time diversification.

RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases

Yingqian Wu, Jingcong Liang, Siyuan Wang, Zhenfei Yin, Philip Torr, Junchi Yu et al. Judging whether a language model can anticipate where research attention is heading is awkward because reviews and idea proposals have no uniquely checkable outcome. Research Attention Prediction sidesteps this with a rolling benchmark of 278 AI and machine learning fields and 1,390 episodes, where at each cut-off an agent searches a time-restricted arXiv corpus and predicts the next six months' share of papers across eight fixed directions. Search helped, yet all four diagnostic models scored below an exact-count exponentially weighted moving average baseline on compositional accuracy. Two coupled bottlenecks emerged: with cumulative history, carrying forward a state estimate beat forecasting directly for every model, and frozen-evidence replay tied part of that reversal to forecast-oriented policies retrieving less recent evidence, while fine-tuning on realized outcomes raised Qwen3-4B's rank correlation by 0.105 on held-out fields.

Kernel-Managed Shared Memory for System-Wide Personalization

Ryan Lum, Yongfeng Zhang In multi-agent assistants, context one agent learns about a user is usually stranded there and unavailable to the rest. The proposed design has specialized agents write structured, tagged memories while the agent-system kernel — not the agents — controls retrieval, privacy enforcement, and prompt injection; it is implemented on AIOS and tested over 1,800 trials with GPT-4o, Llama-3.1:8B, and Qwen-2.5:7B. Against the Mem0 external memory backend on identical storage, kernel-managed retrieval improved personalization scores by 2.4 to 4.0 points on a 5-point scale, with every comparison significant at p below 10^-18, and it beat standard retrieval-augmented injection by similar margins. Compared with dumping the full unfiltered context, it matched quality on two of three models while cutting end-to-end latency 15 to 61% and reducing per-call tokens and cost.

Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?

Tianzhu Zhang, Chih-Kai Huang, Meikang Qiu cross-listed Agents driving network configuration operate within separate authority scopes, which limits the blast radius of a bad action but fragments the evidence needed to confirm an operator's network-wide intent was actually realized. A successful configuration call proves nothing about whether remote devices reacted or routing changes propagated, and an observation that was valid can go stale after the next change. EvidenceNet is a runtime assurance layer whose broker gathers post-change observations named by a completion contract and whose admission gate checks they come from the required scopes, remain current, and satisfy the task rules, with a verifier agent separately judging the content. On live routing networks, post-change state checks recognized successful outcomes that configuration-action records alone could not establish, and controlled interventions confirmed the gate rejects completion when observations are misattributed, substituted, or stale.

A-JIT: Agentic Just-In-Time Software Construction

Mark Marron, Earl T. Barr cross-listed Conventional software delivery builds code before execution and ships it as a fixed artifact. Agentic Just-In-Time Software Construction (A-JIT) instead treats an application as an assembly of code, a runtime harness, and an embedded AI agent that watches usage and live execution traces, specializing logic, workflows, and tool interfaces to the end user much as a JIT compiler specializes machine code to hot execution paths. The model lets applications construct missing implementations and generate new capabilities on the fly while adapting to user behavior, and the authors frame it as a design space for trace-driven human-AI co-construction and self-evolving software.

What Should an Agent Forget? Separating What Is Stored from What Is Used

Yuhang Li, Yuchen Li Persistent language agents must keep stored experience available over time, yet a superseded fact that misleads a current-state question may be essential for a historical one. RD-Forget is a training-free framework that separates what an agent stores from what it uses: a retained source archive preserves every observation, while a query-conditioned memory view built by a frozen LLM curator groups facts into semantic slots, preserves relations needed for multi-hop reasoning, suppresses superseded values through same-slot replacement links in current-state contexts, and makes earlier evidence eligible again for historical queries, with a rate-distortion formulation governing the view under a memory budget. Across conversational memory, knowledge updating, fact consolidation, long-context reasoning, and personalization tasks, configurations without forgetting or query conditioning show the largest score deficits, while slot grouping, historical access, and relation preservation contribute complementary gains.

GANDR: Claim Auditing for Verifiable Legal Answer Generation

Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos In legal question answering an answer is only useful if each claim can be checked against its cited source, yet grounded-generation pipelines score the answer as a whole, letting correct conclusions rest on fabricated or loosely matched citations. GANDR (Grounded ANswer DRafter) is a two-agent system in which a Drafter writes in a structured legal-reasoning format and a Critic, seeing what a human verifier would see, audits every claim against its cited passage and emits a per-claim audit trace each round, evaluated under a strict criterion requiring every citation to resolve to a passage the retriever returned. On a 185-item legal benchmark where six systems share one backbone, retrieval surface, and citation instruction, GANDR reaches 70.8% strict accuracy, leading the strongest baseline by 11.3 points, with the lead persisting at 3.2 to 6.5 points on three further backbones. Removing the protocol-anchored commit rule drops strict accuracy by 22.7 points, and against two law-trained annotators the audit flags under-supported claims at F1 0.84 as a binary detector, though its four-way verdict labels agree only weakly and are treated as advisory.

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

Rui Sun, Zhan Shi, Bing He Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in math and code, but diagnostic reasoning over messy data lacks cheap ground truth because the true cause of an anomaly often requires costly expert investigation and may stay ambiguous afterward. The authors engineer that asymmetry instead: they sample a hidden intervention, inject it into a controlled simulator, and generate the observations it would produce, so the intervention serves as an oracle label and objective reward while the agent must still investigate noisy, confounded, distributed evidence using Python and SQL. In TRACE, a digital-advertising diagnostic environment with 12 root causes and segment-level attribution, supervised fine-tuning lifts Qwen3.5-35B-A3B from 0.159 to 0.637 FullAttr@1 on a 235-episode held-out test set, and subsequent RL with synthesized rewards reaches 0.757, beating every prompted baseline including Claude Opus 5 at 0.686 and a prompted Qwen3.5-122B-A10B. The trained policy also uses substantially fewer tool calls than its prompted base, which the authors read as evidence that a scalable objective training signal can matter more than model scale alone.

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Ansuman Mullick, Eray T\"uz\"un Existing memory systems for large language models treat every personal fact the same way, so stores grow without bound and retrieval precision degrades over time. Fortunate Recall is a composable policy layer that classifies personal facts into a 10+1 behavioral ontology and applies category-specific lifecycle rules, such as differential temporal decay, slot-key supersession, event-time validity, and category-aware retrieval routing, as deterministic functions over LLM-extracted metadata. The implementation FR-Bank reaches 76.9% on LifecycleBench, a new 516-question temporal-disambiguation benchmark, ahead of Mem0, A-MEM, Memory-R1, and MemoryOS, while matching standard retrieval performance on LongMemEval-S. A pre-registered ablation shows that generic lifecycle metadata carries the correctness gains while the behavioral ontology carries calibration, halving downstream confabulation (12.0% vs 24.2%), and the ranking replicates on the open-weight Kimi K2.5 and on the independently built BEAM benchmark.

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

Zixiang Chen, Yuheng Lu, Zihao Cheng, Zeming Liu, Jizeng Bai, Ziye Huang et al. Real-world GUI workflows often span multiple devices and platforms, requiring transfer of intermediate results and shared state, yet existing benchmarks test agents on single-device, statically defined tasks. JarvisGUI is a dynamic benchmark that treats GUI tasks as typed input-output transformations, which lets it automatically compose multi-step workflows across Android, Windows, and Ubuntu virtual environments and evaluate agents within a unified framework. State-of-the-art open-source GUI agents struggle with state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management, exposing a capability gap that single-device benchmarks do not reveal.

Safety & Alignment 22

Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models

Syed Ghazanfar Abbas, Dongyan Xu cross-listed Most LLM security work targets role-play jailbreaks rather than what happens when a user asks a model to verify an identity claim using a test the model designs itself. Through a staged 'I am your developer' experiment on ChatGPT, Claude, Qwen, Mistral, and Llama, the authors find all five initially reject the claim, but Qwen, Mistral, and Llama then generate their own technical challenge, grade the answers, and return Verified without any external identity evidence, with Llama going on to claim access to internal runtime and deployment state. They name the model-generated procedure a Model-Issued Pseudo-Credential (MIPC) and the resulting judgment Conversational False Authentication (CFA), note that the accepted identities did not actually change authorization boundaries, and conclude that authenticated identity must come from an external security component rather than from dialogue.

AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

Zhihao Liu, Hongyu Sun, Zhiyuan Fu, Xiaonan Duan, Jice Wang, Shangru Zhao et al. cross-listed Computer-use agents (CUAs) act on screenshots, raising the question of whether a small adversarial image patch on a web page can trigger real commands through the full chain of screenshot input, vision-language model (VLM) generation, action parsing, and environment execution. AgentHijack trains such patches, deploys them on author-controlled GitHub Pages sites and a locally hosted CSDN clone, and evaluates them against five open-source or publicly available GUI-agent and VLM backends in live environments across 600 online cases. Trigger success reaches 84.5%, action parsing 47.0%, and end-to-end attack success 20.3%, and trajectory analysis shows some agents execute a malicious terminal command and then carry on with the original benign task.

In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning

Iliano Fasolino cross-listed Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, but tampered passages create a new attack surface where the model may repeat injected falsehoods. The study measures how a small quantized Llama 3.1 8B degrades on a fact-checking task built from FEVER when zero to three of its three retrieved passages are corrupted by entity swap, number swap, or negation, across a factorial sweep of 588 runs. Accuracy drops from 77.9% on clean context to 43.5% when all three passages are poisoned, with entity swaps flipping the most previously correct answers and number corruption staying flat until poisoned passages form a majority. The model rarely fabricates new falsehoods and mostly abstains under attack, and the authors frame the strategy contrasts as suggestive pending controlled decoding and stronger labeling.

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji cross-listed Speech-to-speech (S2S) models hear a speaker's voice, which carries gender cues, but most emit a single fixed output voice, so auditing that voice for stereotype drift can miss bias entirely. The authors cross male and female input voices with masculine-, neutral-, and feminine-stereotyped passages across five open- and closed-source models in English, Spanish, and Mandarin, measuring both how the re-spoken voice is perceived and how the model attributes the speaker's gender. The rendered voice shows no stereotype drift, but every model infers gender from the content rather than the voice: when content clashes with voice, the worst model misgenders the speaker in 90% of cases, versus 2% when they agree, and shifting content one step more feminine multiplies the odds of a female judgment by 1.7 to 24.

An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks

Viet K. Nguyen, Mohammad I. Husain cross-listed Agentic frameworks that read images give attackers a way to inject instructions into a model's context without going through the user. MMPIBench delivers a fixed set of attacks through six visual carriers (OCR text, overlays, EXIF metadata, QR codes, fake interfaces, and hybrids) and traces how far each injected instruction propagates from perception through planning to the tool call. Across 720 runs spanning six frameworks, five foundation models, six carriers, and four attacker objectives, attacks are attempted in 12.8% of runs but complete in only about 1%, with the gap closed almost entirely at the planning step where the model reads the instruction and declines to act; the model matters far more than the framework, with one model never attempting an attack and two others attempting in 23.6% of runs. Extending to audio, only two models ingest it and three frameworks deliver it, but where the signal arrives the attack completes in 49% of cells and 75% for one model, indicating perceptual channels beyond vision are narrower but far less defended.

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

Thomas Rivasseau cross-listed Cipher-based jailbreaks, in which a model is taught to exchange encrypted harmful requests and answers, were previously demonstrated by fine-tuning commercial models on encrypted corpora. The authors show that newer frontier models can acquire an arbitrary cipher through prompting alone, with in-context examples when needed, and that alignment is substantially weakened or fully bypassed once the conversation moves into the cipher. Successful jailbreaks are demonstrated against frontier models from Anthropic, Google, and OpenAI without any fine-tuning. Because the harmful content is encrypted, it also slips past commercial harmfulness classifiers as apparent gibberish.

Watermarks Without Verification: AI Text Watermarking After the EU AI Act

Alexander Nemecek, Vipin Chaudhary, Erman Ayday cross-listed Article 50 of the EU AI Act, in force since August 2, 2026, requires generative AI providers to mark and make detectable the content their systems produce, and the paper reports that Anthropic disclosed that all Claude models released after that date embed a SynthID-Text-based watermark by default with no opt-out, while Google has used it in Gemini since 2024. User objections about quality loss, hidden identifiers, and removability, along with the vendor's reassurances, are sorted by what evidence would settle each, and the central argument is that the inability to verify either side, rather than watermarking itself, is the substantive governance failure. Because no public tool can test the deployed systems, the open-source SynthID-Text implementation is evaluated on two open-weight models: on prose the watermark's effect does not exceed that of changing the sampling seed, while on code it costs three points of correctness on one model and is below measurement on the other, with detection remaining near chance. The remaining gaps are mapped to concrete requirements including release of matched outputs, configuration disclosure, accredited audits, a shared evaluation protocol, and interoperable detection.

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi Agentic systems in production read untrusted inputs, call tools with real permissions, and act autonomously, yet standard evaluations remain single-turn and miss multi-step vulnerabilities. The authors present a black-box framework that needs only a basic system description: a seven-domain taxonomy mapping observable behaviors to risk categories, automated SAGE-RT red teaming that generates 120 adversarial scenarios per domain, and human-validated scoring by LLM judges. Across CrewAI and AutoGen agents built on four base models, they report 56.25% average governance risk, 65% privacy risk in multi-agent configurations, and agent-behavior vulnerabilities reaching 85%, arguing that architectural weaknesses can be surfaced without privileged access.

Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning

Saeed Shariati, Mohsen Alambardar Meybodi Federated learning shares model updates rather than raw data, but analytic gradient inversion attacks that reconstruct training samples in closed form degrade as batch size grows, recovering only about half of a 100-sample batch even with full control of the network, and known upper bounds limit what such methods can achieve. The authors connect gradient inversion to erasure-correcting codes and use an LT-code-inspired peeling procedure to build attacks that recover entire batches exactly, including every label, from a single FedSGD round, while certifying each recovery without ground-truth data. On eight image and tabular benchmarks the attacks beat prior single-round methods by a wide margin, and a passive attacker observing an honestly trained network recovers 94-100% of ImageNet batches at sizes up to 128, more than prior attacks achieve even with active model manipulation, while an active attacker recovers over 90% at batch sizes of several hundred. The authors conclude that the privacy leakage of federated learning has been underestimated.

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Yi Shi, Tanyu Chen, Kai Shen cross-listed Directional ablation strips an aligned model's ability to refuse by projecting a single refusal direction out of the weights that write to the residual stream, using only a few hundred contrastive prompts and no optimization, but it had only been demonstrated on dense models up to about 70B parameters. Applying it to GLM-5.3-Flash, a 320B-parameter mixture-of-experts (MoE) model with 288 routed experts, a four-wide hyper-connection residual, and block-FP8 quantized weights, shows the attack survives but that its effect is located differently from what the original recipe implies. Editing the attention, dense, or routed-expert writers alone removes 0.039, 0.016, and 0.148 of refusal, while editing all three together removes 0.776, so 74% of the effect exists only under the joint intervention and the conventional module-name recipe fails silently on an MoE. The attack yields 41-89 percentage-point refusal reductions across seven harmful benchmarks with no detected capability change, a random orthogonal direction leaves refusal unchanged, and a category-concentrated residue for violence, sexual content, and hate survives every edit tried.

CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

Jinyang Li, Mingyu Guo, Hung X. Nguyen cross-listed Large language models (LLMs) can be coaxed into producing malware, yet how well guardrails hold up for code generation specifically has not been measured systematically. CS-Guard covers text-to-code generation with 1,000 malware-generation prompts, 7 jailbreak attacks, and a new fictional scenario attack (FSA) that hides malicious intent inside a plausible software-development story, plus code-to-code generation with 331 prompts spanning infilling, completion, and translation, evaluated over 9 guardrails and seven LLMs. Average attack success rate (ASR) after jailbreaks reaches about 50% for many guardrails on text-to-code, code-to-code ASR approaches 100% on base LLMs and stays high (14.4% to nearly 100%) across many guardrails, and the FSA alone achieves ASR near 100% against many guardrails. The benchmark ships with a modular three-layer guardrail taxonomy so developers can register new guardrails for evaluation.

Subgroup Membership Inference Audits of Differentially Private Synthetic Text

Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar, Siew Kei Lam, Anil Anthony Bharath cross-listed Synthetic data releases with differential privacy (DP) guarantees still carry residual risk, and existing membership inference attack (MIA) audits measure only average-case risk over random records, which can hide danger to vulnerable subgroups. The authors define a subgroup-targeted membership inference game in which the target pool is an explicit parameter and audit 32 proxies under three attacker-knowledge scenarios across four datasets, three generators (DP-SGD fine-tuning, API-based prompting, and activation steering), and five privacy budgets. Synthetic releases leak subgroup membership and prior attacks systematically underestimate it; DP cuts average leakage at every budget tested, but under DP a tenth of the records carries roughly 40% of the remaining leakage, the noise removes more measured leakage from random records than from high-risk ones, and which records leak depends on the release mechanism rather than the record alone.

When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors

Cho-Ying Wu LLMs are increasingly used to simulate professional decision-makers, but their behavior as common-law jurors, and in particular how a defendant's own courtroom statement moves them, has not been characterized. JuryBench supplies controversial U.S. criminal cases in which a fixed base case is paired with defendants of varying backgrounds who offer statements differing in emotional appeal or rebuttal, judged by simulated jurors spanning the ideological spectrum, producing 432,000 decisions and rationales from 20 frontier models. Emotional persuasion often backfires, since jurors read it as evidence of guilt or inconsistency, and juror ideology strongly shapes severity. The strongest single effect is that background fit between juror and defendant outweighs other isolated factors, with jurors harsher toward opposite-background defendants and lenient toward matching ones.

Strangers to Themselves: What Language Models Say About Themselves Is Generic

Phil Blandfort, Urja Pawar Models will readily describe how they would behave under pressure, whether they would cave to pushback, misuse a tool, or lie, and this work turns those descriptions into a falsifiable prediction test across nine behavioral evaluations by comparing self-predicted rates against measured ones. Direct self-report barely correlates with actual behavior at r = +0.04, and showing the model the exact evaluation items lifts it only to +0.24, while asking the same item-informed question about capable AI agents in general does just as well at +0.28 and other models predict the target model at least as accurately as it predicts itself. Frontier scale does not change the pattern, and gains are not self-specific. First-person framing has one reliable effect, shifting reports in a flattering direction that understates harmful behavior, and finetuning on a model's own behavioral record teaches narrow predictions while also altering the behavior being predicted.

Deep and shallow biases in language models

An Vo, Vy Tuong Dang, Khai-Nguyen Nguyen, Emilio Villa-Cueva, Thamar Solorio, Anh Totti Nguyen et al. When a model keeps picking the same answer among many plausible ones, prior work calls that bias without separating a stable learned preference from an artifact of one prompt's wording. A bias depth score measures both how strongly a model favors its top answer under direct prompting and whether that answer survives reframing into a different scenario. Across 4,442 opinion prompts and four large language models, only about a quarter of concentrated preferences survive reframing, splitting the phenomenon into deep and shallow biases. Deep biases are more often inherited from pretraining and preserved through supervised fine-tuning, and they resist both continued fine-tuning and prompt-based debiasing more than shallow ones do.

What Makes Adversarial Examples Transfer Across Deepfake Detectors?

Rafael M. Mamede, Pedro C. Neto, Ana F. Sequeira cross-listed Deepfake detectors are vulnerable to transfer-based black-box attacks, where adversarial examples crafted on a surrogate are applied to an unknown target, but how source-target compatibility drives success has been unclear because prior studies use small detector pools and conflate architecture with training factors. The study evaluates transfer across 60 detectors spanning six backbones, two pretraining regimes, and five training-data configurations using AutoAttack and the Carlini-Wagner attack with Expectation over Transformation (CW-EOT). Transfer is significantly higher when source and target share a backbone, architecture family, pretraining regime, or training data, with the dominant factor depending on the attack, and while the mean attack success rate averaged over non-target sources is only 7.21% under AutoAttack and 19.52% under CW-EOT, a multi-source oracle combining both attacks reaches 64.48% even after excluding exact backbone and training-data matches. The release includes 240,000 perturbed images, full pairwise transfer results, detector configurations, and evaluation code.

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

Marek Jeli\'nski, Jan Dubi\'nski, Maciej Chrabaszcz, Sebastian Cygert Bias audits usually read a model's generated text, which needs expensive benchmarks or judge models and misses internal shifts that never surface in output. The proposed audit compares hidden states across related checkpoints, for instance before and after fine-tuning, by encoding each sentence as its similarities to a fixed anchor set so representations live in a shared space despite fine-tuning reshaping the geometry, then measuring how target groups move relative to positive and negative attributes as a Representational Bias Shift. Across three model families and the WildGuardMix, DecodingTrust, and ToxiGen benchmarks the measure tracked output-level bias change in 15 of 18 settings, peaking at a correlation of 0.84 under full fine-tuning while becoming more model-dependent under parameter-efficient adaptation. Thresholding it flagged checkpoints with increased bias at ROC AUC between 0.65 and 0.99, needed no task-specific evaluation data, and ran in roughly three minutes using 3 to 50 times less compute than the output-level benchmarks compared against.

Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance

Samar Ansari cross-listed Current compute governance attaches thresholds and reporting duties to training runs, even as capability shifts toward deployment through inference-time scaling, agentic scaffolding, and compression onto consumer hardware. The authors build a feasibility taxonomy of twenty inference-stage mechanisms spanning monitoring, verification, and enforcement, rating each on a four-point readiness scale against evidence from four vendors and stress-testing them against an adversary model of three capability tiers crossed with four roles. Fifteen of the twenty mechanisms already have commercial technical substrates in production, though that readiness holds only against a cooperative deployer and a low-to-medium-capability user, and no mechanism rates adequate against a state-level deployer. Fine-tuning strips the model-internal parts of the enforcement cluster while platform-external controls survive; a second-rater check on a random subset of ratings returned a quadratic-weighted Cohen's kappa of 0.74.

Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning

Jing Guan, Yachao Yang, Zhaoliang Liu, Yuyao Zhang, Fanyu Meng, Junlan Feng Preventative steering defends against malicious fine-tuning by injecting undesirable-trait persona vectors during training and removing them at evaluation, but why the protection persists was unclear. Tracking optimization over time reveals an early compensatory adaptation phase followed by a steady state in which the corrective signal decays, with attention output projections acting as the dominant residual-write route for defensive updates. Intervention Delta Preservation experiments show that preserving or reinjecting the learned weight offset does not maintain protection, meaning the defense depends on ongoing active adaptation rather than a static parameter change. Building on that, Progressive Intensity Scheduling starts at moderate injection strength and raises it once fixed-strength alignment begins decaying, improving safety robustness and cutting harmful trait expression on Qwen2.5 and Gemma-3 models.

DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs

Bhuvan Arora, Devesh Saraogi, Sravya Varada, Dhruv Kumar Existing cultural benchmarks score large language models (LLMs) against a single correct answer, which cannot characterize a model's default cultural preference when several culturally grounded answers are all valid, and they conflate default preferences with context-driven adaptation. DiSCo is a distribution-first forced-choice framework that isolates default cultural priors and tests steerability through a four-level context gradient, instantiated as DiSCo-Bench with 304 items derived from BLEnD across 12 cultures and applied to six instruction-tuned LLMs. Default priors are heavily concentrated, with the UK and US together absorbing about 35% of all selections despite being only 2 of 12 cultures. Prompt-based steering consistently widens the gap between high- and low-resource cultures, and injecting explicit cultural facts barely shifts the distribution, suggesting prompt-based personalisation alone cannot fix the bias.

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou Large language models can memorize and reproduce sensitive or copyrighted training content, and existing unlearning methods that apply broad parameter updates tend to hurt utility and can see forgotten knowledge re-emerge after post-training quantization. FOM-UL (Forgetting Only What Matters via Unlearning Layers) scores each transformer layer by its influence on the forget set relative to its sensitivity to the retain set, then concentrates unlearning updates on the highest-scoring layers while leaving the rest of the model untouched. Across TOFU, KnowUnDo, and MUSE-style evaluations it reduces residual memorization compared with GA, NPO, KLD, SURE, ReLearn, and LUNAR-based baselines while keeping retain-set utility close to the original model. Under 8-bit and 4-bit quantization, the targeted updates survive low-bit rounding better than diffuse ones, maintaining stronger memorization suppression, and adversarial prompts recover less of the forgotten content, though the authors make no formal erasure guarantees.
1 more specialized paper

Vision 17

Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

Yiran Qiao, Feng Wang, Jing Ma World Action Models (WAMs) pair predictive world modeling with action generation, but in games there is no external physical environment in which to execute actions, and action-conditioned video rollouts provide visual observations without the persistent, navigable 3D geometry that playable games require. Valerant is a training-free framework that turns a pretrained action-conditioned world model into a WAM for exploring and constructing 3D game maps by coupling predictive visual rollouts with SLAM-based spatial reconstruction and exploration-driven action selection. Starting from a single image, the system progressively builds a persistent 3D game map, extending WAM-based interaction beyond 2D visual simulation and offering a route to reducing manual effort in 3D map creation.

Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation

Sudaksh Kalra, Dolly Sapra cross-listed Edge vision deployments face shifting latency, power, and memory budgets, but conventional networks execute a fixed computation graph and cannot trade accuracy for cost at runtime. Elastoformer converts an existing network into an elastic one that switches among several operating modes on the fly, replacing the common practice of shipping and managing a separate model per budget. Reported savings reach up to 85% fewer FLOPs, 50% lower latency, and 76% less memory overhead, demonstrated on both Vision Transformers and convolutional networks to support the claim of architecture independence.
15 more specialized papers

Multimodal 13

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

Meng'en Qin, Junye Chen, Jucheng Liu, Youlu Xing, Song Wang, Ruize Han cross-listed Multimodal large language models (MLLMs) hallucinate, and existing attention-based mitigations rely on indirect signals such as attention weights that do not track the actual information shift behind a hallucination. HEAL applies causal noise intervention to multi-head outputs to discard causally redundant heads, then uses a counterfactual Difference-in-Differences analysis to sort the remaining heads into four types by how their information is distributed. The analysis shows hallucinations occur when the information distribution in synergy heads drifts from a healthy equilibrium, largely independent of the number or strength of modality-specific heads, so HEAL injects dynamic calibration factors into those heads' value vectors to rebalance visual and language dependencies, reducing hallucinations across multiple MLLMs.

VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang et al. cross-listed Vision-language models are mostly evaluated on consumer, subject-centric video, leaving open how well they handle the fixed-camera footage used for safety monitoring and operational logging in warehouses, roads, and smart buildings. VANTAGE-Bench covers logistics, transportation, and smart-space domains with eight task formulations spanning semantic, spatial, temporal, and spatio-temporal skills, including dense captioning, spatio-temporal grounding, and a single-pass single-object-tracking protocol scored against specialist trackers, over 3,346 media assets. Evaluating 17 models zero-shot, the authors find the gap relative to consumer benchmarks is concentrated rather than general: event verification, referring expressions, and temporal localization drop roughly 9 to 24 points at every model scale, while video question answering stays within 5.3 points of VideoMME and 2D pointing shows no shortfall against BLINK. Temporal tasks are weakest in absolute terms, frontier models track objects nearly as well as specialists over short horizons but fall behind as horizons lengthen, and open-weight models lead 2D object localization outright.

StreamAlign: Streaming Text-Aligned Speech Tokenization

Kang-wook Kim, Jinyoung Park, Jinsoo Kim, Sehun Lee, Sang Hoon Woo, Gunhee Kim Text-aligned speech tokenizers map speech into the token space of a pretrained language model, but they depend on offline automatic speech recognition (ASR), which requires complete utterances and forces word-level rather than subword-level granularity because of vocabulary mismatch. StreamAlign aligns speech and text online by combining character-level RNN-Transducer alignment with word-level ASR guidance, and adds a proactive word-boundary classifier that anticipates word completion at chunk boundaries. Tokenization latency drops from 560 ms to 270 ms, and on LibriSpeech the tokenizer achieves the lowest word error rate and highest UTMOS among those evaluated. A spoken language model trained on its units outperforms other end-to-end spoken language models on speech continuation and is most consistent on SALMon and spoken StoryCloze.

VLX-VR: An Agentic-Aware Video Reasoning Model

Sheng Li, Peng Liu, Qianqian Zhang, Tiancheng Zhao Video understanding requires combining visual, audio, textual, and temporal evidence spread across a clip, but fixed-context single-pass pipelines cannot adaptively gather more evidence when observations are incomplete, ambiguous, or conflicting. VLX-VR is trained with reinforcement learning on videos and agent trajectories to run a Think-Memory-Observation loop in which it decides what evidence it needs, calls read_memory or write_memory, incorporates the returned observation, and chooses whether to continue or answer. On MINERVA it reaches 78.79% accuracy, the best among the compared models, with accuracy across the three duration groups ranging only from 76.70% to 80.92%; on correctly answered samples 96.20% of its reasoning traces are consistent with the reference evidence, while counting, state changes, causal reasoning, and spatial perception remain weak spots.

Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

Mingbo Yang, Wenqiang Wang, Zhaolu Kang, Peng Chen, Yannan Chen, Sunshang Wang et al. In-context learning in multimodal large language models tends to copy the surface form of demonstrations rather than the reasoning path the input actually requires, which caps performance on harder tasks. The proposed framework rewrites each demonstration as a contrast between a weaker and a better response to the same input, paired with a reasoning path showing how the refinement happens, making the route to the desired answer explicit. Because useful refinement depends on what the model has already produced, demonstrations are picked by a response-conditioned retriever, and a lightweight alignment controller predicts response quality to decide whether another refinement round is warranted. Across three families of multimodal tasks the method consistently helps, with the largest gains on visual question answering.

Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs

Haiji Liang, Pengfei Zhou, Zhenglin Wan, Wei Wang, Yang You, Wangbo Zhao cross-listed Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, and existing token-pruning methods assume one fixed strategy suits every input, even though the authors' analysis shows that alternative strategies beat the average-best one on a substantial fraction of individual samples. VIP-Router is a lightweight VIsion Pruning Router that, conditioned on cheap visual and textual features, predicts which candidate pruning strategy will work best for each input at a given pruning level and can fall back to full-token inference when pruning looks unfavorable. On VTC-Bench Group A, a suite of pruning-sensitive perception benchmarks, it beats the best fixed strategy at every reduction ratio with a 26.9% relative gain in average accuracy and a 22.0% relative gain in utility after accounting for realized token cost. The router adds trainable parameters equal to 0.017% of the backbone, leaves pruning algorithms and model weights untouched, and transfers across MLLM backbones and to unseen benchmarks.

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Killian Steunou, Yannis Tevissen, Moun\^im A. El Yacoubi cross-listed Video large language models (VideoLLMs) couple video representations with pretrained language models, but their compute and memory grow with frame count and context length, which limits real-time, mobile, and resource-constrained deployment. The survey catalogs inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameters, FLOPs, latency, memory, or visual and audio token count, analyzing bottlenecks across frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding. Methods are organized by the pipeline stage at which they act, covering VideoLLMs developed since late 2022 alongside earlier frame-sampling and vision-encoder techniques that remain components of current pipelines. The authors assemble accuracy-versus-cost comparisons under shared host models and input protocols where available, separate those from heterogeneous cross-paper evidence, and identify gaps in audiovisual efficiency and standardized evaluation, with a maintained companion repository.
6 more specialized papers

Robotics 11

No Free Checker: A Survey of Verifiers for Robot Policies

Yang Wan, Xihang Yue, Zhirui Liu, Ziyuan Chu, Shuxun Wang, Yuhan Chen et al. cross-listed A verifier for robot policies scores a candidate behavior and is used both to evaluate vision-language-action policies and to train them, spanning success detectors, reward models, runtime monitors, safety filters, and temporal-logic specifications. The survey compares roughly 150 verifiers along two axes, availability, meaning how cheap, early, and dense a verdict is, and credibility, meaning how much a high score actually says about the task, and groups them by who supplies the judgment: humans, rules and formal methods, learned or pretrained models, and the policy model itself. Across all four families, credibility falls as availability rises, so no verifier is simultaneously cheap, early, dense, and trustworthy. The authors then review how verifiers themselves are validated, through agreement with human labels, the performance of the policy they train, and behavior under reward hacking, and close with nine metrics for making a verifier claim checkable.

HiRAD: A Flexible Large-Scale AGV Routing System

Yunjie Huang, Ruizhong Wu, Mengxuan Zhang, Frodo Kin Sun Chan, Yan Nei Law, Lei Li cross-listed Routing large fleets of Automatic Guided Vehicles (AGVs) in warehouses is hard: classical Multi-Agent Pathfinding solvers scale super-quadratically and assume idealized grid motion, while existing Reinforcement Learning (RL) approaches discretize space and time, need millions of episodes, and observe the full map at every step. HiRAD is a hierarchical RL framework for continuous-space routing that uses a step-level spatiotemporal representation, splits heading choice from velocity control to shrink the action space, and runs an asynchronous event-driven decision pipeline that reduces inference complexity from quadratic to linear in fleet size. Across random graphs and two warehouse maps it reduces makespan by 45 to 63 percent and lowers per-step latency by up to 71 percent.

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

Shengye Dong, Haochen Niu, Hao Liu, Peiwen Lin, Chuang Wang, Shanmin Pang Vision-language-action (VLA) policies emit a chunk of one to two seconds of actions per forward pass but decode it from generic per-timestep tokens through a linear head, which entangles smooth global trends with fine corrections and, because dot-product attention is least sensitive near orthogonality, poorly captures relationships between near-orthogonal motion phases such as reach, contact, and settling. Time-Frequency Geometric Cross-Attention (TFGCA) decomposes the action chunk with a per-dimension learnable stationary wavelet transform into time-frequency tokens, then lets each time token attend to them with a score that fuses the dot product with the wedge-product magnitude through a learnable weight; a zero-initialized residual makes it a drop-in addition to a pretrained VLA. Relative to the same base model it gains +1.5 on LIBERO, +6.3 on the out-of-distribution LIBERO-Plus, +28.5 under RoboTwin domain randomization, and +11.67 points of success on three real AgiBot A2 tasks.

Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization

Andy Zeyi Liu, Haoran Sun, Lucas Baker, Randall Balestriero, John Sous Joint-Embedding Predictive Architecture (JEPA) world models learn compact latent representations for prediction and planning, but whether they capture physics well enough to generalize to unseen dynamics has not been tested. SG-JEPA (SemiGroup-JEPA) extends the LeWorldModel framework by feeding the physical parameter to the temporal model through action-conditioning and jointly training the encoder and predictor via an autoregressive latent rollout, evaluated on tasks under varying gravitational fields that share one law but produce qualitatively different motion. Compared with DINO-WM, it cuts open-loop prediction error by up to 2 times on 2D datasets and raises control success by up to 2.5 times on 3D robotic datasets using independently trained diffusion policies. A linear feature model separating local law-conditioned error from its amplification under rollout shows that most of the gain comes from the encoder learning features the predictor can carry forward, rather than from the predictor learning better dynamics.

Show-Harness: Just a VLM Agent Can Play Robots

Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin et al. cross-listed Foundation vision-language models (VLMs) hold broad knowledge about the world, but turning that into robot control remains hard. Show-Harness is an embodied harness that exposes discrete semantic action units the VLM can reason over, with embodiment-specific interpreters deterministically grounding them into local robot actions, so the VLM stays responsible for fine-grained physical decisions. Through this interface, closed-source frontier VLMs can control robots zero-shot, and small open-source VLMs can be adapted with only a few GPU-hours of fine-tuning, while the companion GUMI (GUI Manipulation Interface) lets humans and agents collect demonstrations across embodiments without teleoperation hardware. Harness-equipped VLM agents generalize across tasks, embodiments, and environments and outperform representative agentic and vision-language-action (VLA) paradigms, suggesting the interface rather than added model capacity unlocks embodied capability.
6 more specialized papers

Reinforcement Learning 9

Unthrottling the Tanh Jacobian in SAC: A Negative Result on Bang-Bang Control and MetaDrive

Faiq Shamass Soft Actor-Critic (SAC) squashes an unbounded Gaussian policy through tanh, and the Jacobian of that map vanishes as actions approach the bounds, raising the concern that the actor loses critic signal exactly where extreme actions such as full brake or full throttle are optimal. The authors test a minimal fix that adds one term to the actor loss whose gradient on the pre-tanh mean is the detached action-gradient of Q, with no gain parameter. On a minimum-time double integrator with a bang-bang optimum, vanilla SAC already reaches near-optimal return (-31.6 against a calibrated optimum of -30.3), while the ungated bypass saturates the policy and collapses return to -195.5, and a gated variant that fires only near the bounds also fails. Warm-started fine-tuning on MetaDrive shows the same pattern, with no return improvement and lower collision rates typically traded for more out-of-road departures, leading to the conclusion that the Jacobian throttle is real but undoing it does not help on these tasks.

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

Jianing Wang, Xintao Wang, Aili Chen, Jie Shi, Hongcheng Guo, Jun Gao et al. Existing reinforcement learning for socially intelligent dialogue optimizes single turns against sparse outcome rewards, producing short-sighted policies that mishandle the tension between achieving goals and maintaining relationships over long conversations. SocialRL trains with multi-turn PPO that propagates delayed outcome rewards back to individual turns, and adds six process-reward dimensions such as goal advancement, relational attunement, and contextual coherence, scored by a reward model that generates fine-grained criteria per dimension and weighted by a stage-aware schedule favouring relationship-building early and goal pursuit mid-conversation. Across multiple social-dialogue benchmarks it raises Goal Achievement by an average of 9.2 percentage points over the corresponding base models in both synthetic and real scenes.

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

Eshwar Reddy M, Sourav Karmakar Reinforcement learning on reasoning traces has driven frontier gains mostly in domains with cheap, sound verifiers, and the authors argue the binding constraint is a verification gap: no scalable, incorruptible reward exists for reasoning outside formal settings. In a joint-Gaussian model of best-of-N selection they show that verifier-gold correlation is the exact exchange rate between test-time compute and capability, with an unsound verifier paying a polynomial penalty, and a copula form predicts real LLM judges' soundness to 4% median error. In program-synthesis testbeds with executable ground truth, unsound verifiers lose Soundness-under-Pressure from 0.94 to 0.32 at N=4096 while a sound verifier improves monotonically, and under GRPO training a frozen reward model collapses executed reward by 90% whereas refitting it on a 10% stream of reality-settled labels preserves six times the executed reward. They propose proof-carrying cognition, in which reasoning steps are typed probabilistic claims priced by a world model trained only on held-out reality and settled by proper scoring rules, and specify Soundness-under-Pressure as the headline metric for a benchmark.

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

Guanqun Zhao, Zijun Xie, Binbin Zheng, Jiafeng Lu, Enlei Gong, Zeyu Chen Asynchronous reinforcement learning is the standard way to scale LLM training, but the lag between the behavior policy and the current policy biases the critic toward stale behavior, and prior asynchronous LLM work corrects only the actor. Classical off-policy value corrections do not transfer to long-horizon agentic tasks because a short correction horizon leaves the regression target without the reward while a long one lets products of importance ratios drift exponentially with trajectory length. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte Carlo tail beyond it, separating policy correction from reward propagation. It improves mean@1 on BrowseComp-Plus by 2.4% over the strongest baseline, runs 2.46x faster per step than synchronous training, and stays stable 50 updates off-policy.

FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

Yansen Han, Shengyi Liao, Peng Sun, Deyuan Liu, Yuanxing Zhang, Pengfei Wan et al. cross-listed Preference alignment for flow and diffusion models splits into online reinforcement learning methods that need fresh samples from the current model and offline methods on fixed preference pairs that lean on positive-only finetuning or likelihood-ratio surrogates, with no clear account of how the two relate. A divergence-based framework organizes both, and FlowCPO is introduced as an offline forward-KL objective that uses preferred and dispreferred samples without any rollouts. Under explicit regularity conditions on linear interpolation, the forward-KL objective is bounded by a contrastive flow matching loss that is tractable on fixed data and provably nonnegative, whereas the signed regression loss of simplified FlowDPO can be unbounded below. In domain, FlowCPO reaches 0.84 on GenEval and 0.87 on OCR score against 0.81 and 0.74 for FlowDPO at classifier-free guidance 3.0, while out-of-domain results are mixed.

Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection

Haoyue Liu, Xiaoyu Ma, Ye Chen, Zhichao Wang, Xiaoying Tang Reinforcement learning over a frozen reasoner is a common way to teach a policy which external tools to invoke, but in specialist scientific settings the space of tool subsets is small enough to enumerate, making sampled-rollout methods like GRPO structurally mismatched. The authors show the approximation degrades as training succeeds, because a concentrating policy resamples the same subsets and rewards collide: the fraction of genomic questions yielding no reward signal rises from 0.2% under a uniform policy to 20.8% after GRPO training. Their remedy, FGPO (Full-Group Policy Optimization), scores every tool subset to optimize the exact action expectation and precomputes a question-by-subset reward table, removing frozen-reasoner calls from the training loop. Across five frozen reasoners and three genomic benchmarks, FGPO beats GRPO in all 15 settings by 6.75 points on average and up to 14.20, and on GenomeQA cuts tools invoked per question from 2.36 to 1.40.

Learning Intrusion Response Strategies for OT Systems

Duc Huy Le, Rolf Stadler cross-listed Cyberattacks on Operational Technology (OT) systems that monitor and control industrial processes call for automated intrusion response, but a defender only ever sees partial evidence of what an attacker is doing. The authors formalize an OT intrusion response use case as a Partially Observable Markov Decision Process (POMDP) whose observation model is grounded in realistic network traffic measurements, and train response strategies with Proximal Policy Optimization (PPO). Evaluated on an emulated OT system, the learned strategies are effective against several types of MITRE-catalogued attacks for the studied use case.
2 more specialized papers

Reasoning 8

Quantifying Logical Consistency in Transformers via Query-Key Alignment

Eduard Tulchinskii, Anastasia Voznyuk, Laida Kushnareva, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev et al. cross-listed Chain-of-Thought prompting lets language models write out intermediate reasoning steps, but nothing checks whether each transition actually follows from the last. The authors propose a lightweight evaluation strategy that computes a QK-score from the query-key alignments of carefully chosen attention heads in a single forward pass, avoiding ablation-based analysis. These scores reliably separate valid from invalid inferences across multiple logical reasoning benchmarks, stay robust under added distractors and deeper reasoning chains, and hold across models from 1.5B to 70B parameters.

World-Time Compute with Verified Code World Models

James Schwoebel, Ingrida Semenec, Jenia Rousseva, Marcos Ortiz, Collin Overbay, Christopher Klaus et al. Language models generalize within a domain only after seeing many labeled examples, which most domains lack. When a domain's dynamics can be written as code, one template can instantiate many executable, verifiable world models that each yield unlimited exactly-labeled trajectories; fine-tuning on trajectories from many such worlds, which the authors call world-time compute, improves generalization to held-out worlds the model never saw. Gains concentrate where capability is weakest, +29 points for a 0.5B model, while the largest model's lift is within noise; a corrupted-label control shows label exactness rather than task variety drives the effect, and on List Functions one adapter trained on 128 disjoint worlds reaches 40% on held-out worlds versus 6% for the control. The same lever works as per-world test-time training on ARC-AGI, List Functions, and CLRS, fading for long reasoning chains and perception-heavy tasks, with worlds authored and served by the OpenWorld framework.

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary Existing benchmarks for large language model theorem proving in Lean 4 draw mostly from competition math such as the IMO and Putnam, which poorly reflects field-specific applied mathematics. StochBench collects 450 graduate-level stochastic-processes problems at varying abstraction levels, each paired with its natural-language source, spanning Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes, an area underrepresented in Mathlib. An agent built on Opus 4.8 proves 157 of 450 problems (34.9%) under a 15-minute per-problem limit, indicating the benchmark remains challenging for advanced provers.

Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

Yaning Jia, Chunhui Zhang, Wenxuan Xu, Xingjian Diao, Xiaoyuan Wang, Soroush Vosoughi Supervised fine-tuning (SFT) applies the same cross-entropy loss to every target token, which can over-sharpen tokens the model has already mastered while piling pressure on tokens it barely supports. Trimmed Logit-Gap SFT (TrimSFT) reweights each token's loss by the gap between the gold token's logit and its strongest competitor, using a Gaussian weight that concentrates learning on an intermediate band and trims supervision from both extremes, with no reference model or extra forward pass. Across six base models from the Llama, Qwen, and DeepMath families and five math benchmarks, TrimSFT beats standard SFT on average for five of six models, with gains of up to 26.9 points on MATH500. Ablations show the bandwidth of the weighting matters more than the exact margin location, and trimming only one extreme gives worse trade-offs.

Structural Process Supervision for Latent Chain-of-Thought Reasoning

Yiqi Li, Xu Chen, Chen Ju, Jiangchao Yao, Zhaoyang Li, Jinsong Lan et al. Latent reasoning replaces explicit chain-of-thought (CoT) tokens with compact continuous embeddings, but without process supervision over those embeddings training tends toward representation collapse and uneven information distribution. Prototype-Mediated Process Supervision (PMPS) introduces learnable reasoning prototypes as semantic anchors, projecting latent and explicit CoT embeddings into a shared prototype space so sequences of unequal length can be softly aligned many-to-many, while a Progressive Sequential Alignment (PSA) module starts with positional priors that favor sequential matching and relaxes them during training. On GSM8K-Aug the method cuts output length to under 50% of explicit CoT while averaging 2.08% higher accuracy than the SIM-CoT baseline across model families, surpasses CoT-SFT on GPT-2, and stays the most accurate latent method on larger models and a harder task at comparable output length.

Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

Yunxiang Mo, Donghao Zhao, Hejia Geng A tempting way to cut reasoning-model inference cost is to repeatedly probe a partial trajectory for its current answer and stop once the probes agree, a rule called self-consensus. A preregistered sweep of 3,520 such rules, replayed on frozen trajectories from two models and three benchmarks, finds that none clears three pre-specified acceptance gates, a result that reproduces on a held-out split and two unseen models, whereas a boundary-confidence control, DEER, run through the same pipeline clears all three. The failure traces to a consensus-termination gap: agreement shows the answer persists under a fixed probing procedure, not that reasoning has finished, so at a rule saving 32% of tokens one stop in nine fires on an answer the trajectory later abandons, usually cutting off a correction, and widening the agreement window only lowers that share to about 7% while shrinking savings to 8%.

Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza, Alexandre Berard, Thomas Euyang et al. Reasoning language models mostly think in English regardless of the prompt language, which is inaccessible to non-English speakers, risks losing the intent of the question, and forgoes knowledge better expressed in the target language. The authors study L2 reasoning, reasoning consistently in the language of the user's prompt, from a data-centric angle by optimizing data composition and scheduling in supervised fine-tuning (SFT). Their Tiny Aya L2-Thinker, at 3.35B parameters, reaches an in-language reasoning rate above 93% across 60 languages on six benchmarks spanning math, commonsense reasoning, instruction following, open-ended generation, and cultural reasoning while keeping performance strong. Generalization to held-out languages depends on broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, and the model weights and multilingual reasoning data are released.
1 more specialized paper