Wednesday, September 23, 2026

443 papers cs.AI · cs.LG · cs.CL ← 2026-09-222026-09-24 →

Jul Aug Sep

Highlights

FrontierMath Erd\H{o}s

Highlight Reasoning Tom Adamczewski (Epoch AI), Thomas F. Bloom (University of Manchester) FrontierMath Erdős (FME) is a benchmark of 68 Erdős problems that were still open as of August 2026, selected for mathematical interest and difficulty from 652 open problems on erdosproblems.com. To solve a task, an AI system must prove or disprove the conjecture in the Lean proof assistant, working autonomously under a fixed budget. The benchmark is meant to replace one-off demonstrations of AI solving open problems with a systematic, like-for-like comparison. Five AI systems were evaluated at $300 per problem: GPT-6 Astra scored 3% and all the others scored 0%.

AI systems have recently solved several open math problems, but those claims come without failure rates, costs, or cross-model comparisons. FrontierMath Erdős (FME) turns this into a controlled benchmark of 68 Erdős conjectures that were open as of August 2026, chosen for significance and difficulty, where a model must prove either the conjecture or its negation as a formal proof in Lean 4.

  • The formal-proof format lets the benchmark cover any conjecture that can be stated in Mathlib, not just problems with an easily checked answer object, and a proof accepted by the Lean FRO's Comparator checker settles the question: it checks sandboxed submissions against trusted statements, allows only Lean's three standard axioms, and re-checks every proof through the Lean kernel.
  • Each model runs autonomously in an offline deepagent harness built on Inspect, with Lean/Mathlib, SageMath, and a snapshot of 476,000 arXiv math papers, and gets one attempt per conjecture capped at $300 and 72 hours.
  • Only a pre-release GPT-6 Astra scored above zero, at 3% (2 of 68): it disproved #74 for $218 and proved #126 for $247, while GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1, and Claude Fable 5 all scored 0%.
  • Extra non-benchmark GPT-6 Astra runs with larger budgets resolved 5 conjectures at a total cost of over $220,000: they disproved Erdős's 1931 distinct-subset-sums problem (#1), fully proved the Erdős–Sós conjecture (#548) and the Erdős–Simonovits rational-exponents conjecture (#571), and gave a |S(A)| ≫ n^{1/2} bound for #126, far stronger than Erdős asked for.
  • Among the limitations, scores may understate mathematical ability because a model must also formalize any results missing from Mathlib; reducing a conjecture to another famous open problem earns no credit; resolved problems may leak into future training data; and GPT-6 Astra's spend was metered at stand-in prices, so its costs had to be corrected after the run.

Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

Highlight HF pick · 11▲Multimodal Embedding Team Ovis-Embedding is a family of embedding models that maps text, images, video, and audio into one shared space from a single multimodal backbone rather than separate per-modality encoders. It starts from a pretrained Qwen-omni model adapted with contrastive training and low-rank initialization. Training uses homogeneous-source sampling so each batch has informative negatives, focal loss to emphasize hard examples, and similarity-based distillation from expert models. At inference, low-rank decomposition allows compact embeddings of flexible size. The family reports state-of-the-art results on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB.

Current multimodal embedders either leave out audio or bolt an audio encoder onto an embedding space that was learned without it. Ovis-Embedding instead starts from an omni-modal understanding model, Qwen2.5-Omni, whose text, image, video, and audio inputs already pass through one shared Transformer, and trains it contrastively for any-to-any retrieval.

  • The embedding is the final-layer hidden state at the last non-padding token, with no projection or modality-specific heads; Ovis-Embedding-Omni-3B drops the speech-generating Talker, and Ovis-Embedding-VL-2B/9B are built the same way from Qwen3.5 backbones.
  • Training runs in stages: LoRA contrastive pretraining with a focal-weighted InfoNCE loss plus offline teacher-similarity distillation (the authors found full-parameter updates at this stage erode the pretrained semantics), then full-parameter fine-tuning with homogeneous-source sampling so that in-batch negatives in each micro-batch come from the same task, then an annealing Embedding Distillation step that upsamples cases the student gets wrong and uses a per-query KL weight based on the student's confidence.
  • The corpus holds roughly ~50M query–target pairs across text, images, documents, video (verified by a VLM), audio, GUI/tool agent data, and interleaved inputs, and the authors say it is deduplicated against evaluation sets; the paper flags these data figures as placeholder estimates.
  • A frozen-encoder PCA rotation plus zero-initialized residual adapters turns one 2,048-d embedding into nested widths down to 128-d, which avoids the drop in full-dimension quality the authors saw when adding Matryoshka training to multi-objective training.
  • Ovis-Embedding-Omni-3B is reported as state of the art on MMEB-v3, leading all six modality groups, and Ovis-Embedding-VL-9B leads four of five MMEB-v2 task groups, with further gains claimed on MVEB, MAEB, and RTEB; the extracted text gives no absolute scores or ablation numbers, and the release of checkpoints and code is still only promised.

From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health

Highlight HF pick · 12▲Applications He Hu, Yucheng Zhou, Qianning Wang, Yingjian Zou, Chiyuan Ma, Juzheng Si et al. This survey reviews large language models (LLMs) for mental health support and organizes the field into three phases. In Phase I, LLMs are information tools and pattern recognizers used for assessment. In Phase II, they act as empathetic conversationalists in stateless, in-the-moment exchanges. Phase III, the current frontier, aims for longitudinal, personalized companions built as stateful cognitive agents. The review covers core technologies, agent architectures (profile, memory, reasoning, and planning), and datasets and benchmarks, and lays out a roadmap for responsible, human-centered systems.

Research on LLMs for mental health is large and scattered, and it lacks a unifying narrative. This survey organizes it around one thesis: LLMs are moving from passive text classifiers, to empathetic single-session chatbots, to stateful agents that act as long-term, personalized companions.

  • The framework has three phases, separated by operational criteria such as interaction granularity, user-state modeling, memory, personalization and autonomy: Phase I covers static detection of depression and suicide risk from social media, Phase II covers stateless empathetic dialogue, and Phase III covers longitudinal, goal-directed agents.
  • Phase II is built mainly with SFT and parameter-efficient methods (LoRA, QLoRA) on counseling corpora, followed by preference alignment with DPO, KTO, ORPO and GRPO (as in Psyche-R1); at inference time it relies on CBT-grounded chain-of-thought and RAG over clinical materials such as crisis-helpline information.
  • Phase III agents are broken down into Profile, Memory, Reasoning, Planning and Tool Use modules, with examples including CA+ for session planning toward long-term goals, MusPsy for multi-session modeling of how a client changes, and multi-agent systems like AutoCBT and MAGI, which splits the MINI psychiatric interview across four specialized agents.
  • The resource catalog covers detection datasets (RSDD, UMD, SWDD), dialogue corpora (ESConv, PsyQA, Cactus, SoulChatCorpus) and multimodal sets (E-DAIC, MEDIC); it shows a shift from real transcripts, which privacy concerns keep small, toward synthetic multi-session data.
  • The authors flag several open problems: multimodal cues are ambiguous and prone to spurious correlations, synthetic data risks distribution mismatch and cultural homogenization, and gains on automated safety or LLM-as-a-Judge scores do not guarantee clinically appropriate behavior such as escalating a crisis on time.

Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Highlight HF pick · 6▲Reasoning Vasily Ilin Lean Pool is a repository of mathematics formalized in the Lean proof assistant. The archive is grown, maintained, and optimized entirely by AI agents. The abstract gives no further details on methods or scale.

Mathlib's strict human review keeps it growing linearly and leaves out much research-level mathematics, so verified formalizations have nowhere maintained to live. Lean Pool is an archive of completed Lean formalizations that AI agents import, upgrade across Lean/Mathlib releases, and optimize, with the aim of becoming a formal counterpart to arXiv.

  • Admission requires complete proofs with no sorry and no axioms beyond Classical.choice, propext and Quot.sound, an Apache-2.0 or MIT license, and a project card recording attribution and proof provenance, and every submission passes CI linters plus an LLM review of whether its statements match the claimed results.
  • The archive now holds 211 projects, 3,228,485 lines of Lean and 837 registered main results (70 human-written, 102 AI-written and 39 mixed), and it is the most reused external repository in the LeanEval structural audit of research-level formalization.
  • Agent-driven dependency upgrades restored compatibility across six Lean version bumps, though the breakage can be large (97 of 191 projects failed going from 4.34.0-rc1 to stable 4.34.0) and needed human-overseen follow-up repairs after the first automated pass.
  • Accepted optimizations removed tens of thousands of lines, with one elaboration-cost pass cutting 54,965 lines and full build time by 5.8% (28.46 to 26.82 min), but shortening proofs doesn't reliably speed up builds: contributor proof golfing made the build 5.4% slower and library-wide compression raised peak RAM.
  • The LLM review service cost a median of $0.20 per report across 285 historical reports but about $89 in API-equivalent terms under the newer Codex-based service, and nobody independently checked how accurate its reviews were.

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Highlight HF pick · 2▲Vision Hanyang Kong, Xingyi Yang Understanding interactions in 3D scenes means jointly identifying movable parts, how they move, and the handles used to operate them. Segment-Snap links these outputs through the physical relationship between parts and handles: a geometric decoder uses planar and upright priors plus predicted handle locations to pick hinge lines without a trained motion regressor, and part predictions help refine handle candidates. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98%, and extra handle candidates lift handle AP from 24.63% to 30.99% with full context.

Predicting what moves in a static 3D scan, how it moves, and where it is grabbed involves two linked ambiguities. A well-segmented door can fit a hinge on either side, and a small handle crop doesn't show whether its part rotates or slides. Segment–Snap resolves both by passing evidence once in each direction between parts and handles: handle locations choose the hinge side, and part predictions add handle candidates and correct their motion labels.

  • Three independently trained Volt-B predictors run on the same RGB point cloud: an SPFormer-style movable-part segmenter, a dense pointwise handle detector, and a joint model that predicts a handle mask from each detected part's features.
  • A training-free geometric decoder fits a planar rectangle to each part, assumes a vertical axis for rotating parts and the thinnest box direction for translating ones, and places the hinge on the line farthest from the nearest detected handle; with masks and axes held fixed, this raises Articulate3D validation motion-gated AP from 13.74% to 40.98% (+27.25 pp), with rotating-part motion AP rising from 1.88 to 56.37.
  • Adding the joint model's part-conditioned handles raises handle AP from 24.63 to 29.65, recovering 53 extra ground-truth handles, while spatially permuted or random size-matched masks give no gain; relabeling those added handles using containing parts adds 0.98 pp, and using dense-handle labels as a fallback brings the total to 30.99, with the gain concentrated on translation handles.
  • On the public challenge test set, the system ranked first on both boards, with 48.28% motion-gated AP and 34.46% handle AP.
  • Limitations: the decoder assumes roughly planar parts with upright hinges, so horizontal or tilted mechanisms fail; 139 of 390 ground-truth parts are never detected at all; the label-correction gain varies from 0.19 to 1.34 pp across training runs; and learned motion decoders fall only slightly short of the fixed rule, a gap that is not statistically resolved.

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Highlight Multimodal Qwen Team Qwen3.8-Omni-Flash is a natively multimodal model built for agentic productivity across text, audio, and video, and it inherits the sparse mixture-of-experts (MoE) architecture of Qwen3.8-Next with a one-million-token context window. A native multimodal co-training strategy keeps strong text capabilities while transferring agentic skills from text to audio and video tasks such as video editing, long-form translation, and music-video generation. The team also releases Qwen-MM-Plugins, an open-source plugin framework that adds audio and video support to agent harnesses, and Qwen-Live-Harness, a framework for real-time multimodal agents that handles memory, tool use, and sub-agent delegation. They report strong results across multimodal understanding, reasoning, long-horizon agent tasks, and video productivity benchmarks.

The Qwen team presents Qwen3.8-Omni-Flash, an omni-modal model that shifts the focus from perception and chat to long-horizon multimodal productivity: editing video, translating long audio and video, and turning tutorials into notes or reusable agent skills. It pairs a sparse mixture-of-experts backbone that reads up to 1M tokens with native multimodal co-training, so the agentic skills learned on text carry over to audio and video, and it ships with open-source harnesses (Qwen-MM-Plugins, Qwen-Live-Harness) that give existing agents access to audio and video.

  • The Thinker is built on Qwen3.8-Next's hybrid backbone, which combines Gated DeltaNet layers with Qwen Sparse Attention, and adds a new Spatial AuT encoder for multichannel spatial audio; training runs four pretraining stages, including about 2.5T multimodal tokens, followed by multi-teacher distillation from domain-specialist models and unified reinforcement learning with outcome-based agentic rewards.
  • Across 29 audio, audio-visual and agent evaluations, the model averages more than 25% higher than Qwen3.5-Omni-Plus, with estimated per-hour API input costs 98% lower for audio and 93% lower for audio-visual content; the largest jumps are in multi-speaker speech recognition (DER on AliMeeting falls from 88.1 to 3.4) and multimodal tool use (WildClawBench-MM rises from 34.5 to 71.0).
  • Text ability roughly matches the text-only Qwen3.8-Flash, with 63.3 on SWE-bench Pro and 92.6 on LiveCodeBench v6, so adding the extra modalities does not visibly degrade coding or reasoning.
  • Treating long videos as an agent task, where the model retrieves segments on demand through Qwen Code instead of reading the whole input at once, lifts LVOmniBench from 63.3 to 73.6 and cuts OmniVideoBench token use by about 45.7% (from 145,736 to 79,117 tokens per query) while also improving accuracy; in a separate autoresearch demo, the model cut a small model's Sichuan-dialect character error rate from 25.79% to 15.30%.
  • Gains are uneven: Gemini 3.8 Flash still leads on static video reasoning, AgenticVBench and OmniGAIA, while FLEURS speech recognition and translation, VoiceBench and WildSpeech fall slightly below the predecessor; the authors also note that the token savings do not show lower end-to-end latency or cost, and a few figures in the text disagree with the paper's own tables.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Highlight HF pick · 1▲Agents Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang et al. Agents on long-horizon tasks keep making choices, such as which hypothesis to test or which implementation to build on, that decide how the whole run turns out. Current benchmarks measure only end-to-end success, not the quality of these choices, which the authors call "taste." Taste-Bench mines decision forks automatically from parallel agent attempts and from detours inside single trajectories on engineering and research tasks, and asks a model to choose a direction without seeing what follows. The best frontier model answers only 59.7% correctly, and a larger reasoning budget does not help. Distilling judgment from a teacher that has seen the outcomes into a student model improves the student's decisions and its end-to-end success on held-out SWE-bench Pro tasks.

Long-horizon agents succeed or fail on mid-run decisions, such as which hypothesis to test or which implementation to build on, yet end-to-end benchmarks only score the final outcome. The authors call the skill of making these decisions well an agent's "taste". To measure it without human annotation, they find "decision forks" in recorded agent trajectories, points where the agent could take either of two directions, and use the recorded outcome of each branch to label which direction was better.

  • How the benchmark is built: forks come from two sources: pairs of independent attempts at the same task that split into a passing run and a failing run, and single runs where an agent abandoned a failed direction and then recovered. A GPT-5.6 Sol generator proposed 4,657 candidate forks, and a four-model judge panel removed questions that could be answered from the option wording alone and questions whose label the full record did not support; 10.8% survived, giving Taste-Bench's 502 questions across SWE-bench Pro engineering tasks and RE-Bench/HCAST research tasks, and human reviewers agreed with 98.8% of the mined labels.
  • Frontier models score low: a question counts as correct only if the model picks the right option under both answer orders, so random guessing scores 25%; among 14 models, the best is GPT-5.6 Sol at 59.7%, with GPT-5.5 at 59.5%.
  • Late-evidence forks are hardest, and more reasoning does not help: mean accuracy drops from 62.3% when the deciding evidence is already visible before the fork to 21.0% when it only appears after substantial later work, which is below chance. Moving from the lowest to the highest reasoning-effort setting changes accuracy by only −0.2 and +2.2 points for the two models tested, even though they spend the most reasoning tokens on the hardest forks.
  • Taste is not a restatement of coding ability: the correlation with SWE-bench Verified is r = 0.63 overall and only 0.37 on the engineering subset. The top four SWE-bench Verified models are within 4.0 points of each other there but 10.7 points apart on Taste-Bench.
  • Taste can be trained, with caveats: a Qwen3.6-27B student with LoRA adapters learns from a copy of itself that has been shown the correct direction, trained with token-level forward KL on the teacher's own reasoning, and rises from 30.0% to 47.9% on forks from tasks it never saw in training. Used as advice for a separate executor agent, the student raises SWE-bench Pro success from 14.6% to 33.7%, close to the 39.0% reached with always-correct advice. That end-to-end test covers only 41 tasks, relies on forks mined from earlier runs on those same tasks, and distillation was tested only on the engineering questions.

Recursive self-improvement of AI research agents

Highlight HF pick · 5▲Agents Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu, Zhengyao Jiang AIDE^2 sets up recursive self-improvement for a frontier AI research agent. The agent proposes edits to its own code, benchmarks the modified versions on AI R&D tasks, and keeps the edits that score best on hidden evaluations. During an autonomous 8-day run it found seven successive improvements, including a new search policy and memory mechanisms that compress the agent's growing context. The best discovered agent matched or beat a strong human-engineered production research agent on four held-out benchmarks, including out-of-distribution weather forecasting. Although the loop never optimized for it, reward hacking on a separate task family fell from 55% to 32%.

Research agents normally improve the code they are given, but the process they use to do research stays fixed. AIDE² points an agent at its own harness instead: an outer loop rewrites the agent's code and benchmarks each variant on AI R&D tasks. A rewrite is kept only if it scores higher on hidden held-out data, so every accepted version becomes the one the next round edits.

  • The system is a two-level tree search: a claude opus 4.7-driven outer agent proposes rewrites of an inner agent that starts as AIDE_0, and each inner agent runs on gemini 3 flash under a fixed dollar budget per task across ML engineering, heuristic algorithm, and harness engineering tasks; because the inner agent optimizes against public scores while selection uses private ones, it cannot directly game the selection criterion.
  • An autonomous 8-day, 100-node run accepted seven successive improvements, raising the private grade from 0.703 to 0.778, above the human-engineered production agent AIDE_human at 0.749; the final agent AIDE_85 uses a UCB1 bandit over five drafting strategies, forks the best solution every five steps, and keeps prompts bounded, making them ~50× smaller than AIDE_0's on ALE-Bench and FML-Bench.
  • AIDE_85 matches or exceeds AIDE_human on all four held-out benchmarks (ALE-Bench, MLE-Bench, FML-Bench, and an out-of-distribution WeatherBench 2 physics-forecasting task); some of the largest gains came on WeatherBench 2, where both evolved checkpoints converged on nearly the same fix on every seed.
  • Reward hacking on held-out KernelBench kernel tasks, which the loop never optimized for, fell from 55% to 32% along the lineage, below AIDE_human's 39%; in one case the agent repaired a broken evaluation script instead of exploiting it.
  • Limitations: noise builds up across both loops, so a single falsely accepted rewrite can derail the search; the "ignition test" of whether a discovered agent is a better self-improver was inconclusive at three seeds (0.780 vs 0.782); and the evolved agents are complex and hard to interpret.

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

Highlight Reasoning Xiaoyu Luo, Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, Johannes Bjerva Closed-source frontier models hide their raw chain-of-thought (CoT), so claims about their reasoning are hard to check. The authors register a simple custom tool through a standard API feature, which induces models to externalize intermediate reasoning, and validate the approach against native CoT on open-source models. The extracted reasoning matches native reasoning performance and clearly beats no-reasoning baselines on competition math, science, and code. Comparing frontier models including GPT-6 Astra, they find systematic differences in how models compress and organize reasoning, with Astra choosing a correct trajectory early and externalizing only the crucial steps.

Closed frontier models hide their chain-of-thought, so benchmark accuracy shows what they solve but not how they reason. The authors register a single "scratchpad" tool, force the model's first call to it, and replay each call with a fixed acknowledgment, which gets models to write intermediate reasoning into the tool arguments; they then use these extracted traces to compare how frontier models structure their reasoning.

  • On open models where native CoT is visible, the Forced-Reasoning protocol nearly matches native performance on competition math: DeepSeek-V4-Flash goes from 30.5% to 73.3% (native 70.0%), and GLM-5.2 goes from 21.4% to 84.3% (native 89.9%), with substantial lexical overlap and similar reasoning-step composition relative to native traces.
  • On closed models, forced traces also come close to native reasoning across MATH, HLE, and LiveCodeBench; for example, GPT-5.6 Sol rises from 25.0% to 91.3% on MATH (native 97.5%), and GPT-6 Astra reaches 93.8% with zero provider-reported reasoning tokens in most runs.
  • GPT-6 Astra produces the shortest and least compressible traces and uses the same mix of reasoning activities as the other models, but its LCoT2Tree reasoning trees are far leaner (27 median nodes vs. 60–136 for Sol, Opus 4.8, and Sonnet 5) at comparable depth, because it resolves elementary steps internally and branches less.
  • When one model's trace is given to another as context, strong recipients reuse Astra's compact traces with almost no loss, while weaker ones such as Claude Haiku 4.5 and GPT-5.4 Nano recover less and sometimes miss an answer the trace already states, which suggests a trace's value as supervision depends on who reads it.
  • For closed models the evidence is only behavioral, so the extracted text may be a useful proxy rather than the model's actual internal reasoning; the protocol also needs forced tool-choice APIs, and it failed completely on Claude Opus 5, Fable 5, and Fable 5.1.

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Highlight HF pick · 11▲Large Language Models Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen Diffusion large language models (dLLMs) generate text non-autoregressively, but inference is slow because they lack effective key-value (KV) caching and scalable parallel decoding. The authors identify GPU memory I/O as the dominant bottleneck when both techniques are used together. Flash-dLLM is a training-free framework that adds an I/O-aware fused KV-cache kernel and a draft-and-verify decoding scheme in which the dLLM acts as both drafter and verifier, with no auxiliary model. On math and code benchmarks it beats prior dLLM acceleration methods in both speed and memory efficiency, including 5.1× and 11.0× speedups over Elastic-Cache on GSM8K and HumanEval.

Diffusion LLMs could decode many tokens in parallel, but in practice inference is slow: moving the KV cache in and out of GPU memory dominates runtime, and conservative confidence thresholds hold back tokens that are already correct. Flash-dLLM is a training-free framework that pairs an IO-aware fused KV-cache kernel with a self-speculative draft-and-verify scheme, so the dLLM acts as its own drafter and verifier.

  • Flash-Cache fuses QKV projection, RoPE and cache writes into one Triton kernel that keeps intermediates in SRAM, which alone gives a 1.37× per-layer speedup; a block-table scheduler handles variable-length batches, and only the top-k most-attended decoded tokens are refreshed each step, since just 32 tokens capture up to ~50% of attention in the middle layers.
  • Flash-Verify feeds each low-confidence position twice, once with its draft token and once as [MASK], under a mask that stops the two views from seeing each other; a token is accepted left-to-right only if both views agree and the mask-view confidence clears γ, which roughly doubles the tokens accepted per step at an added cost proportional to 2× the masked-window size.
  • On LLaDA-1.5 across GSM8K, MATH, HumanEval and MBPP, the combined method was the fastest configuration in all eight settings, reaching 148–211 tokens/s (22×–148× over no-cache greedy decoding) and 5.1× / 11.0× the speed of the strongest prior baseline, Elastic-Cache, on GSM8K and HumanEval respectively.
  • On GSM8K-512 it also had the highest accuracy (83.02% at 210.6 tokens/s), and at batch size 16 its flat, preallocated cache used about 26 GB versus 50 GB for Fast-dLLM, scaling to batch 32 where Fast-dLLM runs out of memory at 24.
  • The accuracy cost depends on the task: math stays within 1.78 points of the best configuration, but 256-token code generation drops by up to 3.66 points; the evaluation covers only structured-output tasks, and the thresholds and window sizes are fixed rather than adaptive.

Applications 101

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

Alexandre Cristov\~ao Maiorano Marketers increasingly use LLMs as "synthetic personas" to predict how an audience will react to copy, so the authors test that practice against real click-through data from thousands of headline A/B tests in the Upworthy Research Archive. On the 399 tests with a statistically reliable winner, a no-persona zero-shot prompt ranks headline variants far better (Kendall τ = 0.361) than a ten-persona panel grounded in real audience demographics (τ = 0.084). The result replicates across three data splits and a different-domain news dataset, and holds across seeds, prompt phrasings, several Gemini tiers, and gpt-4.1. The authors conclude that role-playing specific personas adds bias and noise on top of the model's already accurate population-level prior.

Rachel: A general-purpose language model directs and revises retrosynthetic routes

Qisheng Li, Shunchao Jiang, Chen Qi, Xin Su, Da Han, Guangyong Chen cross-listed Rachel is a stateful environment that executes and checks chemistry proposed by a general-purpose LLM, without prescribing a search policy or stopping rule. It tests whether the LLM itself can plan, sustain, and revise retrosynthetic route strategy. Without reference routes, GPT-5.5 fully closed routes for 111 of 120 PaRoutes targets and 24 of 25 difficult RF25 targets, most of which were published after its knowledge cutoff. Blinded LLM evaluators gave Rachel the highest mean route score among the compared methods. Replacing the LLM's route decisions with fixed policies cut closure to 6-15 of 120, even though local chemical execution continued.

WILSON - a pathology foundation model framework for patient-level analysis and diagnostic text generation

Saghir Alfasly, Wataru Uegami, Sobhan Hemati, Wenchao Han, Xiaojia Tang, Kevin Thompson et al. cross-listed Pathology foundation models usually encode thousands of tiles from a single slide, whereas pathologists reason across magnifications and across all the slides in a case. WILSON is a vision-language foundation model that represents each whole-slide image or multi-slide case as a single composite image spanning multiple magnifications. It was trained on about 189k Mayo Clinic slides with pathology reports as supervision. Without task-specific training, it beat a dedicated case-level model (macro-F1 0.52 vs 0.38) and matched slide-level models up to 9.4x larger at 272- to 2,155-fold lower compute. It also retrieved matching diagnostic text at 75.6% recall@1, compared with 58.1% for PRISM.

Potential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing

Brian Wright Course-specific LLM tutors depend on grounding in course materials that often change mid-semester. The authors compare vector retrieval-augmented generation (RAG) with an LLM-compiled wiki of linked concept pages with citations, using the same machine learning course corpus and 59 questions. Both did equally well on single-fact recall, but on questions linking concepts across course units the wiki scored 9.93 of 10 and was 100% grounded in cited sources, versus 8.14 and 64% for RAG. Scoring was done by an LLM judge against a human-written rubric.

From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI

Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng Online A/B tests limit how many changes to a conversational assistant can be tried, so the authors build and audit an offline engagement proxy that needs no exposure to real users in the treatment arm. They treat the proxy as a chain of three alignments: from behavioral labels to product outcomes, from a learned engagement classifier to candidate behavior, and from a calibrated aggregate signal to the experiment's effect. They audit it by comparing offline and online confidence intervals rather than point estimates. On 113 contrasts from eight experiments run after the calibration map was frozen, the composite reaches 81.1% F1 against 34.3% for the raw classifier score and makes no wrong-direction calls, which supports using it to pick checkpoints and system prompts before spending experiment traffic.

Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing

Yi-Lin Tsai (Arvin), Yung-Hsiu (Arvin), Lai Marketers increasingly use generative AI agents as synthetic consumers to pretest logos, packaging, and ads. The authors test this practice by replicating six classic visual marketing experiments with GPT-4o-mini and GPT-5.4-mini, using both plain-text and JSON input formats. Every configuration passed the manipulation checks, but none reproduced more than two of the six human effects, and one significantly reversed the human pattern. In-context evidence moved average responses toward the human results but still captured less than half of the natural variation in human responses. The authors turn these findings into a Calibrate, Intervene, Deploy governance protocol for deciding when synthetic panels can be used.

Evaluating Accuracy and Probabilistic Reliability of Zero-Shot Time Series Foundation Models

Panagiotis Michael, Moysis Symeonides, Demetris Trihinas Time Series Foundation Models (TSFMs) promise zero-shot forecasting without task-specific training, but evaluations rarely check whether their probabilistic forecasts are well calibrated. The authors benchmark six TSFMs on energy, traffic, and financial data against statistical baselines and a supervised deep-learning model. TSFMs beat both baselines on accuracy but show a trade-off between point accuracy and probabilistic reliability. xLSTM-based models stay well calibrated across horizons, patch-based transformers are accurate but poorly calibrated at long horizons, and transformer-based models stop improving beyond a certain context length.

Protocol before progress: leakage-aware evaluation of AIS trajectory prediction

Zobeir Raisi, Vali Mohammad Nazarzehi Had Papers on vessel-trajectory prediction from Automatic Identification System (AIS) data usually credit their gains to new architectures without measuring how much the evaluation protocol contributes. The authors build leakage-aware splits that keep vessels, time periods, or regions disjoint, and use them to audit TrAISformer, GATransformer, and AISFormer-style models on Danish and US Gulf coast corpora. The protocol matters a lot. TrAISformer's best-of-16 oracle decoding cuts error by 2.1-3.2x, sharing vessels across splits lowers error by 23-25%, and a region-disjoint split raises one-hour error from 2.2 to 24.6 km. By comparison, GATransformer's graph attention gives no measurable benefit. The splits and code are released.

Toward Responsible AI-Augmented Cyber Defense: Pattern Recognition, Defense-in-Depth, and the Case for Human-AI Collaboration

Mustafa S. Aljumaily, Hayder Kareem Abed, Nawar S. Alseelawi cross-listed The work builds a formal, falsifiable model linking defense-in-depth, AI pattern recognition, and human-AI collaboration in security operations. Layered defense is modeled as a Bernoulli detection cascade, each layer as a Neyman-Pearson/Bayesian detector with a closed-form optimal threshold, and analyst triage as a capacity-constrained cascade that trades detection against false alarms. Monte Carlo simulation at illustrative operating points shows that AI gains compound across layers, and that reviewing 100% of AI-flagged alerts cuts false alarms about 20-fold but lowers overall detection, because imperfect analysts are applied to every alert. The authors propose an interior-optimum analyst capacity ratio as a design target for security operations centers (SOCs).

Bridge of $\Psi$'s: Quantum Circuit Optimization with Schr\"odinger Bridges

Lino S. Hofstetter, Lia Yeh, Prakash Murali cross-listed Quantum circuit optimization rewrites a circuit into an equivalent one with fewer gates and lower depth, and is usually done with fixed rewrite libraries or algebraic routines. BOPS (Bridge of Ψ's) is a generative model based on Schrödinger bridges, with a custom denoiser, that learns to map a circuit directly to an optimized equivalent. Its training data is made by applying rewrite rules in reverse, so each input has a known cheaper target and is hard for existing optimizers. On held-out 8-qubit, depth-64 Clifford+T circuits, it cuts gate count by 2.46x and depth by 2.45x (geometric mean), beating all nine baseline optimizers.

The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale

Wenhui Chen, Ziyao Lin, Jianlin Chen, Peiji Long, Chi Man Vong CLAIM is a production pipeline for generating K-12 assessment items. A model first drafts an item as a "curriculum architect" and then attacks it as an adversarial reviewer, and generation is guided by accepted and rejected examples plus 44,844 error-correction rules mined from evaluator feedback. Across 43,227 items covering 755 Common Core ELA standards and ten LLMs, it reaches a 97.8% expert-evaluator pass rate. However, independent judges from other vendors agree with that verdict only weakly (kappa ≈ 0.13), so the quality level depends on which judge is used. Multiple-choice generation saturates near 98%, while fill-in-the-blank items plateau at 82.8–96.7% depending on the model, which the authors attribute to autoregressive decoders being poorly suited to deciding the full set of acceptable answers.

Beyond Static Charts: Can Language and Vision Language Models Generate Interactive Data Visualization Interfaces?

Mizanur Rahman, Aaryaman Kartha, Enamul Hoque Prince Language models can already generate static charts from text, but whether they can build interactive data visualization interfaces has not been measured. VIS-GEN is a benchmark of 3,042 natural-language queries covering intents such as filtering, temporal analysis, and editing a visualization. Testing 14 open and closed LLMs and vision-language models (VLMs) shows large gaps and frequent failures on implicit intent, queries with several possible interactions, and complex edits. A multi-stage framework that plans the design, generates several candidate interfaces, critiques them against constraints, and self-refines raises the best model's pass rate by 15.9 percentage points.

FISSION: Label Augmentation for Bot Detection

Sen Yang, Ignacy Nieweglowski, Aviv Yaish Bot and coordinated influence-operation detection lacks reliable ground-truth labels. FISSION generates labels by splitting each account's activity into sub-accounts that count as positive pairs, then trains embeddings that keep the behavioral regularities shared across those pieces, so that bots from the same operation end up close together. It outperforms prior methods at detecting Wikipedia sockpuppets and Twitter/X bots.

HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning

Han Chen, Hanchen Wang, Hongmei Chen, Lu Qin, Wenjie Zhang, Ying Zhang cross-listed Android malware evolves quickly, which causes concept drift that degrades machine-learning detectors. Existing adaptation methods either react only after performance drops or rely on unstable adversarial training. HYDRA represents each app as a hybrid of fine-grained control flow graphs (CFGs) and coarse function call graphs (FCGs). It then aligns old and new data distributions with a cross-domain contrastive objective that uses pseudo-labels, avoiding adversarial training. On large time-ordered malware datasets it achieves lower false-negative and false-positive rates than state-of-the-art baselines while needing up to 87.5% fewer labeled samples.

GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement

Emilien Guandalino, Lorenz K. M\"uller, Beatrice Alessandra Motetti, Konstantin Berestizshevsky, Lukas Cavigelli cross-listed To help predict which AI papers will have impact, GitScholar links GitHub activity from 444,000 repositories to more than 558,000 AI arXiv papers. GitHub reactions improve early impact-prediction precision by up to 12% over a strong academic baseline. The GitHub signal also covers nearly all high-impact AI papers and correlates consistently with later academic success.

FeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation

Xutian Li, Bo Xiong, Yifeng Zhu, Kunze Li, Xianlin Zhao, Runbang Yan et al. cross-listed When generating code inside an existing repository, an LLM needs to find the functions, APIs, and cross-file definitions it can reuse. Current retrieval methods rely on expensive whole-repository graphs or on LLM-driven exploration. FeatLens builds an index linking natural-language feature descriptions to functions. For each task it assembles a task-specific seed graph and prunes it with personalized PageRank, so retrieval is deterministic and uses no LLM calls. On DevEval and EvoCodeBench it achieves the best dependency recall among the baselines and competitive Pass@1, while cutting total token overhead by 45.9% and using far fewer graph nodes and edges than the strongest graph-based baseline.

Deep Generative Crystal Structure Prediction: A Benchmark Study and a Controlled Test of Prototype Dependence

Lai Wei, Rongzhi Dong, Ying Feng, Madeline Miklos, Jianjun Hu cross-listed The authors benchmark 12 deep generative crystal structure prediction (CSP) models, covering diffusion, flow-matching, autoregressive and other architectures, against the template-based TCSP 2.0 on 180 test structures using identical matching criteria. Template retrieval is the strongest single method at 68.3% top-1 success, ahead of EquiCSP (66.4%) and Uni-3DAR (62.9%), and most structures the generative models get right are also found by template substitution. When entire prototype families are removed from training and the strongest model is retrained, accuracy falls by 50-78%, showing that current generative CSP models behave largely as implicit prototype libraries. Only a small residual of genuinely de novo predictive ability remains.

PERSONAWEAVER: Controllable Diversity Beyond Conventional Archetypes in Procedural Character Generation

Maan Qraitem, Kate Saenko, Bryan A. Plummer LLM-based procedural character generation for games and virtual worlds tends to produce homogeneous populations: characters overwhelmingly endorse positive moral norms and respond like helpful assistants. PersonaWeaver separates world building from behavioral specification and assigns behavior from manually curated banks of diverse moral positions and conversational reactions. Across ten realistic and fantasy settings and three LLMs, it produces broader moral and interactional response distributions than prior methods, along with more varied language, response length, sentiment, and less archetypal world-attribute combinations.

Knowledge Pull Requests for Continual Document Authoring

Alexander Martin, Benjamin Van Durme Documents need continual revision as new knowledge appears, but current approaches either edit without recording what changed or regenerate from scratch. Knowledge Pull Requests (KPRs) extract claims from new sources, filter and route them to sections, flag conflicts with existing content, and produce a ChangeLog that keeps the knowledge change separate from the text diff. On cross-lingual Wikipedia revision and report updating on RAGTIME, KPRs integrate more information, preserve existing content better, and add the most information per generated token compared with rewriting or regenerating. A KPR-revised article also supports question answering better than a frontier model with search.

TraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits

Fnu Tanish, Samiha Shimmi, Samikshya Chapagain, Hamed Okhravi, Mona Rahimi, Lei Zhang cross-listed Finding the commit that introduced a vulnerability usually relies on git blame plus positional heuristics, but the true vulnerability-inducing commit (VIC) can sit anywhere in a file's history. TraceVIC localizes likely root-cause lines, traces them across revisions, and builds temporal graphs that capture both program structure and how the relevant code evolves. It then ranks candidate commits by how much each contributed to the vulnerable condition. It improves F2 by up to 28.7% over state-of-the-art methods and identifies a valid VIC for 78 of 79 vulnerabilities across four unseen C/C++ projects.

Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows

Remy Stewart, Olabode Anise, Andrew Hogan, Augustus Griffin cross-listed A randomized controlled trial with 50 product designers and 50 product managers tested whether the prompt-to-design tool Figma Make saves time on three standardized design tasks. Among participants who finished the tasks, access to the tool was associated with roughly 20% shorter completion times, with larger gains for product managers. The authors suggest these tools may let product managers contribute more to design work, while the benefit for professional designers may depend on the task.

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

Om Nepal, Sushant Aryal, Oluseyi Olukola, Nick Rahimi cross-listed Five controlled experiments, covering 203 vulnerable functions from Big-Vul, three code LLMs, and three prompting strategies, show that compile rate is an unreliable metric for LLM-based C/C++ vulnerability repair. About 64% of compile failures are not attributable to the model. A single compiler-standard flag shifts compile rate by 1.8 to 2.7 times on identical patches, and a compiler-feedback loop rewards non-repairs such as deletions and placeholders. Whole-function CodeBLEU fails too, since an unchanged copy of the vulnerable input outscores every model. The authors propose diff_F1, which scores only the edited region, as a cheap screen and argue for change-aware, execution-grounded evaluation.
79 more specialized papers

Large Language Models 62

When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits

Daein Weon, Dong Ho Kang cross-listed Reward models, rerankers, and LLM judges can end up scoring surface form instead of the quality they are meant to measure. For example, a public preference reward model chooses a terse correct MBPP solution over a commented buggy one only 50.7% of the time, no better than a coin flip. The authors test residualization, a common fix that subtracts the predictable surface component from a score. On code, it reduces the format effect by about 0.12 but barely changes the margin between correct and buggy code; in natural language inference and question answering, gains on targeted slices come with falling agreement across the full population. They conclude that a residualized score should be reported as an audit diagnostic alongside the raw score and the cost it incurs, never as a replacement, and they propose a reporting protocol to that end.

Training a Language Model End-to-End in Rust: An Experience Report

Arif Adito A single developer pretrained a roughly 0.4B-parameter, Bangla-first language model entirely in Rust, with no Python or PyTorch in the training path, for $164 of rented H100 time. The main contribution is a failure taxonomy of the Candle and Burn Rust frameworks when used for training: five Candle defects, including fused kernels that silently produce no gradient, and three Burn defects, including a backward pass running at about 3% of theoretical GPU throughput. None of these defects showed up in ordinary loss-curve inspection. The author describes a framework-agnostic "gradient-flow arbiter" test that asserts every trainable parameter receives a finite, nonzero gradient, documents a tokenizer-fertility trap that silently inverted the corpus's language balance, and concludes that Rust is better suited to serving models than to training them for now.

Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation

Lingxiang Hu, Tianle Xia, Ming Xu, Yiding Sun, Linfang Shang Two settings in on-policy distillation (OPD) are studied together: how many distinct prompts are used (prompt breadth) and how often the student policy snapshot that generates training responses is updated (rollout refresh). The math-reasoning experiment fixes the number of trajectories and optimizer updates. With ten policy snapshots, just eight prompts reach 24.09% accuracy, close to the 24.51% reached with 14,080 prompts. Adding prompts helps when responses are refreshed at every update but hurts when responses stay frozen at the initial policy, an interaction of 4.07 percentage points. A second reversal appears at a 32K output limit, where models trained on frozen responses overtake in accuracy while using about 1.7-1.8x as many tokens.

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

Simon P. Villani LatentPort asks whether one hybrid language model can pass its live memory to a larger sibling so the receiver does not have to reread the context. On a Qwen3.5 4B-to-9B pair, translating only the attention KV cache leaves a large gap. Also transferring the persistent Gated DeltaNet (GDN) recurrent and convolution state lowers next-token loss by 0.747 nats per token and improves all 64 PG19 documents. Directly reusing the recurrent state outperformed the learned mappings that were tested. A small 434K-parameter correction brings the 9B model to within 0.076 nats per token of its native performance while processing zero historical prefix tokens, though the evidence covers only one model pair and one direction.

ChainDoRA: Tensor-Train Factorized Weight-Decomposed Low-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning

Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam ChainDoRA is a parameter-efficient fine-tuning method that builds on Weight-Decomposed Low-Rank Adaptation (DoRA), which separates weight magnitude from direction. It replaces DoRA's dense low-rank directional factors with a connected Tensor-Train (TT) chain, adding a separate TT rank that controls capacity and parameter cost. On LLaMA-7B across seven commonsense reasoning benchmarks, it reaches 72.30% average accuracy versus 69.88% for LoRA and 69.39% for DoRA. It does this with 5.35M trainable parameters instead of about 56M, roughly a 90% reduction.

Understanding Reliability in LLM-based Human Behavior Simulation

Pei Wang, Lei Wang, Yuanzi Li, Xu Chen LLMs are increasingly used to simulate human survey responses, but end-to-end scores do not show where these simulations become unreliable. ReliMap splits the simulation process into three layers and measures reliability at the individual and population levels while varying model capacity, how complete the persona profiles are, and population coverage. Across four tasks and eleven LLMs, every model shows substantial distributional bias without profile conditioning. Profiles reduce that bias with diminishing returns, and informative attributes matter more than the number of attributes. Critically, gains at the individual level do not reliably carry over to the population level and can move in the opposite direction, while population-level reliability stabilizes at around 50-100 simulated individuals.

ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch

Sait Furkan Teke (ufak AI) ufakzeka-1 is a 151M-parameter Turkish decoder-only model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat, for about $286 in total. The contribution is a complete record of building and evaluating the model: a Turkish byte-level tokenizer, a three-stage pretraining schedule, and an evaluation battery with enforced decontamination. The authors report three lessons they expect to transfer to other small models. A safety gate "fixed" with training data written from its own questions scored 64/64, while the honest score was 34/64. Variance across training seeds was as large as the spread across all recipes, so single-seed comparisons were uninformative. Some abilities, such as multi-turn arithmetic, did not move with any data change, which the authors attribute to model size. Weights, data recipe, evaluation code, and spending ledger are released under Apache-2.0.

Attention as a Routing Graph: Live Circuit Extraction from a Single Forward Pass

Ash Manvi, Samreena Tajreen Finding circuits in language models normally takes many costly interventions. This work treats the attention patterns from a single forward pass as a routing graph and keeps a small set of routes pointing toward the answer. On induction and indirect object identification (IOI) tasks in GPT-2 Small, GPT-2 Medium, and Pythia-410M, ablating the extracted edges hurts the model much more than ablating random edges of the same size. The extraction costs one forward pass, roughly two orders of magnitude less than a head-by-head patching sweep. The authors present it as a cheap causal sketch rather than a complete circuit map.

FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing

Yusser Al Ghussin, Eva Gavaller, Cristina Espa\~na-Bonet, Josef van Genabith, Simon Ostermann Pretraining corpora are usually organized only by language, while cultural benchmarks target specific regions and locale-specific practices, which makes it hard to audit whether a cultural phenomenon appears in training data. FineWeb-CLaR annotates the full 30.9B-document FineWeb and FineWeb-2 collections with URL-derived region labels and cultural-topic distributions mapped onto a 14-leaf cultural taxonomy, assigning a region to 25.61% of documents. The authors also annotate 277 cultural NLP benchmarks with the same taxonomy and language and region coverage, enabling direct comparison between what is in the corpus and what the benchmarks evaluate.

Clarification Is Not Correction: LLMs Fail to Let Go

Jianzhe Lin, Xiaolin Li, Fei Wang, Robert Douglas, Rajeshkumar Golani, Jubin Chheda Multi-turn dialogue failures are usually blamed on memory, but the authors argue that models often commit too early: an ambiguous early turn locks into one hidden interpretation, and later clarifications get filtered through it, which they call early posterior collapse. In controlled writing, planning, and coding tasks with Gemini-2.5-Pro and Gemini-2.5-Flash, the same information presented in a different order produces different outcomes, and coding is especially vulnerable because early assumptions get built into interfaces and control flow. Summaries and chain-of-thought prompting do not reliably help. The authors call for assistants that keep hypotheses tentative, ask before acting under high-impact ambiguity, and rebuild their task state when evidence changes.

Efficient Iterative Retrieval with Heterogeneous Batching

Dohyun Park, Hubertus Franke, Daniel G. Waddington, Swaminathan Sundararaman, Yongjoo Park Iterative retrieval pipelines combine embedding models and generative models, but serving systems run them separately and leave GPUs underused. Orthrus batches embedding and generation requests together in one inference loop, using chunked embedding with incremental pooling and workload-aware batch composition. On four A100 GPUs it achieves 1.28x to 4.52x higher throughput than baseline deployments on controlled workloads and up to 55.8% lower p99 latency on an iterative retrieval-augmented generation (RAG) benchmark.

Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining

Adam Ousherovitch, Yixin Wang LLM pretraining usually ships the raw final checkpoint, which ties the choice of learning-rate schedule to the choice of which model to return. Terminal Shrinkage Averaging (TSA) interpolates between the final iterate and an average of recent checkpoints, trading recent progress against end-of-training variance. Analysis under a local quadratic approximation, backed by controlled NanoChat experiments, shows that using TSA changes which end-of-training learning-rate schedule works best. The combined schedule and estimator improve validation quality on a depth-22 NanoChat model, and a qualifying time-to-GPT-2 run finished faster than the public baseline.

Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia, Yutaka Watanobe, Fang Liu Long contexts make LLM inference slow and memory-hungry, because self-attention scales quadratically and the key-value (KV) cache grows linearly with context length. Context-to-Answer-Aligned Memory Compression (CMC) compresses the input context into compact Context Memory Embeddings aligned to the embedding space of any frozen decoder, so the decoder's weights never change. It pairs a question-guided selection of these embeddings with a local context window in a two-tier KV cache, and trains the compressor by distilling answers from a frozen LLM. Across nine encoder-decoder pairs and four QA benchmarks it beats the baseline by up to 7.3 exact-match points on SQuAD while cutting peak reserved GPU memory by up to 50% and inference time and energy by up to 20%.

Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs

Shubham Santosh Pandere, Gautam Ranka, Ritika Varshney, Navya Deshmukh, Roushni Sareen, Roshan Kumar Singh When the prompt contradicts what a language model learned in training, a small set of attention heads decides which source to trust. The authors compare these conflict-resolution circuits in base and instruction-tuned versions of Llama-3.2-3B, Qwen-2.5-3B, and Gemma-3-4B using five interpretability methods, and all five agree that instruction tuning reweights the same late-layer heads rather than rewiring the circuit (node overlap 0.60–0.82). Instruct models lean more on their stored knowledge and reject short counterfactual contexts more often, but this skepticism disappears when the same false claim arrives as a coherent, evidence-style passage. Because the circuit is preserved, the authors argue that interpretability tools built on base models should transfer to their instruct versions.

Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA

Wenzhi Fang, Nicholas Tzou, Lazar Valkov, Srinivas Chappidi Supervised fine-tuning (SFT) of models whose reasoning was trained with reinforcement learning (RL) can overwrite that reasoning ability. The authors find that reasoning activations occupy low-dimensional subspaces that can be estimated from a modest number of examples, which leaves a large approximate null space free for adaptation. NB-LoRA (Null-Basis Low-Rank Adaptation) builds a fixed null basis from reasoning activations and reparameterizes LoRA updates through it, so hidden states tied to reasoning are preserved during fine-tuning. Across several RL-trained LLMs and tasks, it matches standard LoRA on the new task while keeping reasoning accuracy near its level before fine-tuning, including on held-out reasoning benchmarks.

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim cross-listed Under greedy decoding, the same LLM, prompt, and software stack can still produce different outputs on different GPUs, because inference frameworks choose different matrix-multiplication kernels with different floating-point reduction orders on each architecture. The authors propose fixed-configuration fused-upcast GEMM kernels that load 16-bit weights, upcast them to FP32 in registers, and accumulate in a reduction order determined only by the problem shape, so every GPU runs the same operation sequence. Linear-layer outputs are bitwise identical across NVIDIA Ampere, Ada, and Hopper GPUs, and inference runs 1.17 to 3.1 times faster end-to-end than the previous state-of-the-art approach while halving weight-memory traffic.

Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices

Qian Xie, Yueli He, Nairen Cao Picking the best LLM configuration by evaluating every candidate on every benchmark item is expensive. GittinsEval treats the choice as a cost-aware Bayesian bandit problem, using the Gittins policy to decide which configuration to evaluate next and when to stop, and adds an anytime recommendation rule based on a lower-confidence-bound score. On response matrices from GSM8K, PIQA, AlpacaEval, and MMLU, it often reaches near-zero simple regret using only 1–2% of the cost of exhaustive evaluation, and its adaptive stopping rule usually triggers at 1–10% of that cost.

From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs

Zhentao Tan, Chang Liu, Yao Liu, Yue Wu, Jieping Ye Parameter-efficient fine-tuning (PEFT) of Mixture-of-Experts (MoE) language models usually updates either predefined weight matrices, as LoRA does, or entire selected experts. The authors observe that even activated experts are internally sparse, so they propose NSFT (Neural Sub-expert Fine-Tuning). NSFT splits each expert into channel groups along its intermediate dimension and picks task-relevant sub-experts using both routing importance and activation saliency. It also adds learning-rate scaling and dynamic gradient scaling to make up for the smaller effective updates. On OLMoE and Ling-mini-2.0, it outperforms PEFT and expert-level sparse tuning baselines while training substantially fewer parameters and keeping general capability competitive.

Latest Exact Match Attention

Moritz Br\"osamle LEMA (latest exact match attention) is a transformer attention variant in which queries and keys are binarized and each query attends only to the most recent key that matches it exactly. The authors prove that LEMA transformers with chain of thought can simulate word-RAM machines and that, conversely, a word-RAM can simulate LEMA at a per-token cost that does not depend on context length. They train LEMA with a straight-through estimator plus an annealed soft-attention surrogate. On associative recall it beats gated DeltaNet (GDN), and 834M-parameter language models match softmax transformers about half their size in loss. Inference uses dictionaries kept in main memory rather than GPU memory (VRAM), which gives constant generation speed comparable to GDN.

Auditing Proxy-Based Validation Across Text Spans

Daein Weon, Dong Ho Kang Evaluation scores are often validated by how well they agree with cheap proxy labels. When the score and the proxy read the same text span, that agreement can come from shared surface features rather than the property the proxy is meant to capture. The authors propose declaring a validation contract that names the score, the span, the proxy, and the target construct, and then re-evaluating the proxy only outside the scored span. On HotpotQA, a score read from a 50-character prefix agrees with its proxy while its agreement with actual correctness is at chance. On OR-Bench, removing each model's recurring opening templates erases most of the score's link to its refusal proxy.

You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

Yuanteng Chen, Qiwei Lai, Chen Tianqi, Peisong Wang, Yuantian Shao, Nanxin Zeng et al. Fine-grained mixture-of-experts (MoE) language models route each token to many small experts, which makes dynamically pruning some of the selected experts an attractive way to cut inference cost. This empirical study covers twelve fine-grained MoE checkpoints from nine architecture families, evaluated on eleven benchmarks spanning knowledge QA, math, code, and reasoning. Simply keeping about two thirds of the selected experts preserves 98.8% of unpruned performance and gives a measured 1.2–1.7x speedup across two serving backends, and it requires changing only one integer. Published dynamic pruning rules beat this simple baseline by under 1% at conservative budgets and by up to 3% under aggressive pruning. Larger and reasoning ("thinking") models tolerate pruning better, while multimodal models are more fragile.

Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement

Akihiro Yoshida, Yuma Ichikawa Mixed-precision weight quantization for LLMs is usually posed as a Multiple-Choice Knapsack Problem (MCKP), with each weight matrix's sensitivity collapsed into a single scalar. The authors prove that even the best scalar proxy can distort the true activation-aware loss by a factor tied to the Hessian factors' condition numbers, and that this factor ranges from 10^1 to 10^13 for typical LLM modules. Their method, CASA (Cross-layer Activation-aware Sensitivity Allocation), first uses a Kronecker-factored activation-aware metric whose relaxed problem has a closed-form solution, then refines bit-widths with a cross-layer local search scored on end-to-end model loss. CASA reaches lower perplexity than recent scalar-proxy baselines, especially below 3 bits per weight, and its zero-shot accuracy gains track each model's average condition number.

Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL

Jiamiao Liu, Dewen Qiao, Yu Zhang, Xuetao Chen Conformal abstention certificates for text-to-SQL are calibrated on correctness labels that usually come from executing queries against the single database a benchmark ships with, a known-lenient oracle. In a preregistered study on Spider-Realistic with four SQL-specialist checkpoints, swapping in the benchmark's stricter multi-instance test suite raises held-out risk 2.73 to 10.23 points above what the certificate reports. Blinded labels from two SQL experts put a certificate calibrated at a nominal 0.10 at 20.0 and 17.2 points of real risk. Neither oracle matches the experts. Execution-consistency confidence scores also look better, in 16 of 16 cases, when judged by the oracle that built their clusters, so the authors recommend reporting certificates under both oracles and evaluating scores under an independent one.

ClusterFewshot: Improving Few-shot Optimization for LLMs workflow

Omri Bar Haim, Shahar Katz, Lior Wolf LLM workflows often depend on picking a few good in-context demonstrations, and existing optimizers choose them by random sampling or metric rankings that ignore the task's semantic structure. ClusterFewshot clusters candidate demonstrations semantically and combines that structure with utility-aware scoring to build representative, effective few-shot sets. Inside DSPy pipelines, it substantially cuts optimization cost across several benchmarks and consistently improves accuracy over bootstrap-based methods, both for prompt tuning alone and for hybrid prompt-and-weight optimization.

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

Baher Mohammad, Ammar Ali, Stamatios Lefkimmiatis Training-free transformer compression usually treats layers in isolation or groups them with heuristics that ignore each layer's activation geometry. GeoPair jointly optimizes which cross-layer weight projections to pair and a shared-dictionary factorization for each pair, so the shared basis preserves each layer's own calibration geometry, and adds structured sparsity. The authors report state-of-the-art results across architectures, scales, and modalities, consistently beating independent structured decompositions and heuristic pairwise factorizations.

Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach

Zijian Zhao, Dian Jin, Xialiang Tong, Sen Li, Mingxuan Yuan Diffusion large language models (dLLMs) decode tokens in parallel but depend on an inference-time denoising scheduler, and common confidence-based schedulers suffer from two failure modes, EOS Overflow and Proximal Bias. Attention analysis traces both failures to positions that put disproportionate attention on invalid tokens such as [MASK] and [EOS], which produces misleading confidence signals. The authors propose a scheduler that combines several heuristic features with a contextual mean-field embedding, tuned with the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) and using only 393 trainable parameters. On LLaDA and Dream across four reasoning and planning benchmarks, it outperforms conventional heuristics, block autoregressive decoding, and recent state-of-the-art schedulers.

ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models

Dahai Yu, Rongchao Xu, Lin Jiang, Ximiao Li, Guang Wang A large language model's (LLM's) response-level confidence can be unreliable when intermediate claims in its reasoning contradict the final answer, and token-probability-based uncertainty quantification (UQ) does not capture that inconsistency. ChainUQ combines a lightweight module that estimates intrinsic confidence from frozen features aligned to the final conclusion with a calibrator that adjusts that score using evidence about consistency within the reasoning chain. Across in-distribution and out-of-distribution benchmarks it gives an average 3.1% relative gain in AUROC and up to a 45.0% relative reduction in expected calibration error (ECE), and it transfers to new settings without further fine-tuning.

TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference

Ning Li, Xinyu Wang, Xin Yuan, Wenchao Xu, Athanasios V. Vasilakos, Song Guo et al. cross-listed Running mixture-of-experts (MoE) models across resource-constrained edge servers causes heavy cross-server traffic, because each token must be routed to experts that live on different machines. TopoCompress jointly optimizes token compression, expert placement and replication, GPU-CPU residency, and routing. A fast online loop drops tokens that are both low in importance and expensive to route, while a slow offline loop re-places experts according to the traffic that remains after compression. The authors prove feasibility, optimality, and convergence properties, and in simulation the method reduces cross-server traffic and resource consumption while keeping inference quality controllable.

TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models

Haibo Hu, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue Speculative decoding speeds up LLM inference by having a small draft model propose tokens for a larger target model to verify, and prior work has focused mostly on the draft side. TSS (Target-Side Sparsification) shows that for domain-specific inference, skipping selected layers of the target model can both cut verification cost and increase draft acceptance, without hurting task quality. It searches for multi-layer skip configurations per domain, stores them in a lookup table, and applies them at runtime with a lightweight controller, with no retraining or permanent pruning. On Spec-Bench translation, it raises average accept length from 2.70 to 4.53 and delivers a 1.68× end-to-end throughput speedup while improving BLEU.

Block-Level Weight-Space Structure Persists Under Post-Training: An Empirical Study Across LLM Families

Zhaohui Wang The authors compare base and post-trained variants (instruct, chat, code) in the Qwen2.5, Llama-3.1/3.2, Mistral, and Gemma-2 families to see how post-training changes the weights. Post-training alters every tensor, so byte-level deduplication saves nothing, but block-level geometry is preserved, with mean cosine similarity above 0.99. Separately trained specializations such as Qwen2.5-Coder do not share this structure. Building on this, LinkerLLM shares compatible weight blocks between variants loaded at the same time, cutting GPU memory by 18–48% and fitting up to five 7B-parameter variants on one 24 GB consumer GPU. Five of eight configurations keep at least 94% of benchmark quality, and the other three each drop to 87–91% on one benchmark.

The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke

Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong Across 761 fine-tuning runs on 12 base models from 0.5B to 14B parameters, trained on competition-math skills, the authors test whether data recipes (skill order, blocked versus interleaved arrangement, composition) matter at a fixed data volume. Within one coherent domain, recipe effects sit at the noise floor, and the best-looking recipe did not hold up on reruns. The exception arises when two halves of a corpus use incompatible but equally correct answer conventions: arrangement then decides which convention the model adopts, an effect about two orders of magnitude larger than a same-convention control, and it disappears once the convention is stated in the input. Data volume was the one lever that reliably helped, and the authors publish a scorecard for their 26 pre-registered claims.

Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity

Kasun Dewage, Marianna Pensky, Suranadi De Silva The authors test whether reconstruction error predicts how much quantizing a single attention projection hurts a model. They quantize one projection at a time with round-to-nearest (RTN) and GPTQ at 3 and 4 bits across nine open models from 1.3B to 8B parameters, including LLaMA, Mistral, and Qwen 2.5. Within each component type, reconstruction error explains a median of only 4.4% of the variance in perplexity sensitivity, while component type and layer identity are stronger predictors. Value projections are most often the dominant source of degradation, so the authors conclude that mixed-precision schemes should treat them separately rather than allocate bits by reconstruction error alone.

Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression

Kasun Dewage, Marianna Pensky, Heranga K. Rathnasekara, Suranadi De Silva Pruning whole attention heads is a hardware-friendly way to compress Transformer language models, but existing ways of scoring head importance need calibration data, gradients, or Hessian estimates. Magnitude Profile (MP) scoring needs none of these. It looks at the row norms of each head's weights, prunes heads that fall in the bulk of the distribution, and keeps statistical outliers, which are assumed to carry more representational capacity. A variant, MP-G, handles Grouped Query Attention (GQA) by sharing each key-value group's score across its query heads. Without any forward passes or calibration samples, MP-G gets the best WikiText-2 perplexity on OPT-6.7B at every sparsity level tested (18.46 at 12.5% head sparsity) and beats Wanda-Head, SparseGPT-Head, and Gradient-Head on RoBERTa-large at low sparsity.

Beyond Imitation: Auditing the Recoverability of Reasoning in Distilled Models

Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li et al. Whether a student model can learn from a teacher's correct solution depends on whether it can continue that line of reasoning. The authors measure this with prefix recovery: they reveal 25%, 50%, or 75% of a verified solution and check whether the student completes it correctly. Across Qwen3 teacher-student pairs from 0.6B to 8B parameters, average prefix recovery rises from 71.0% to 91.9% as the student grows from 0.6B to 4B. Reverse-KL distillation helps most on math and code for students below 2B, which suggests prefix recovery can tell practitioners when costly distribution-level distillation is worth it.

MSA-CITE: A Co-Adapted LoRA Specialist Ecology for Fixed-Budget Small-Model Inference

Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Jie Li, Ru Zhang Small language models are usually deployed as a single post-training checkpoint that is sampled several times. MSA-CITE instead keeps four frozen LoRA branches of a Qwen3-4B backbone, each from a different post-training run. It draws one completion from each branch, groups the answers into equivalence classes, and picks one using source priors from calibration. On 200 held-out math problems, this fixed four-generation budget reaches 65.5% accuracy versus 62.0% for the best single branch. The authors state that the gains on a shifted test set and against the strongest baseline are not conclusive.

PatchKV: Efficient KV Cache Recovery for Dynamically Edited LLM Contexts

Guotao Yang, Rui Guo, Siwei He, Sheng Chen, Yitao Hu, Keqiu Li cross-listed Long-running LLM agents often edit spans in the middle of their context while keeping long suffixes. The suffix tokens are unchanged, but their cached key-value (KV) states can no longer be reused as-is because the history and rotary positions before them have shifted. PatchKV predicts which suffix blocks drifted using an offline model plus sparse blocks chosen from stored attention, recomputes only those, and restores the rest from CPU memory with a fused path for dequantization, RoPE correction, and placement. Across three models and three long-context QA workloads, it gives a 2.51–3.85× speedup in time-to-first-token over full recomputation and 1.26–2.06× over CacheBlend, with comparable F1.

Damage Predicts Recovery: When Calibration Data Matters in Compressing Financial LLMs

Junyi Ye, Mengjia Yu, Debapriya Hazra, Guiling Wang Post-training quantization and pruning of LLMs rely on a small calibration corpus, and it is unclear whether specialized domains like finance need domain-matched calibration data. Across two model families, six compression configurations, three calibration corpora, and ten financial tasks, the authors find that quantization largely preserves performance, so the calibration choice barely matters. Pruning, by contrast, cuts numerical QA accuracy by over 40 points. In those damaged settings FinMix, a mixture of financial task examples, recovers a large part of the loss while another generic corpus does not. Their practical rule is to measure task-specific compression damage first and build specialized calibration data only when the damage is large.

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo et al. Diffusion language models such as LLaDA and Dream usually decode in fixed-size blocks, which ties together how far ahead to look and how many tokens to commit. PACE-dLLM observes that the model's per-step confidence across the window drops off sharply at a context-dependent point it calls a cliff. It fits this cliff in closed form at each step and uses the saturation point to set the next look-ahead horizon, while a separate confidence threshold decides which tokens to commit. A theoretical analysis shows this horizon is the smallest one that achieves maximal useful yield per pass. On four reasoning and code benchmarks it achieves the best average accuracy with average wall-clock speedups of 5.23x on LLaDA and 3.06x on Dream (up to 8.52x on math) over the semi-autoregressive baseline.

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong et al. Long-context LLM inference is limited by KV cache memory traffic. Sparse attention methods pick tokens by attention mass and then roughly compensate for the rest, without considering how the two steps interact. CompKV groups tokens into blocks and selects the blocks that would leave the largest compensation error if omitted; the authors show this error depends on both block attention mass and the variation of logits within a block, and estimate it from compact block statistics. With an asynchronous implementation, it performs best among sparse baselines on RULER and LongBench-Pro and delivers up to a 6.85x self-attention speedup over full attention.

Disaggregated Quantization: Specializing LLM Prefill and Decode

Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh Prefill benefits from low-precision arithmetic while decode benefits from compact weights, so disaggregated quantization (DQ) uses separate formats, weights, and storage placement for each phase. On Qwen 3 and Gemma 3, dropping activation quantization only during decode improves accuracy on decode-heavy tasks at no extra cost. Training a separate NVFP4 prefiller for 1-bit Qwen3.8-27B GGUF decoders raises accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without changing the decode checkpoint. An offloaded disaggregated prefill (ODP) scheme streams the extra weights from SSD and gives a 1.78x time-to-first-token speedup at 8K prompt length in llama.cpp.

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

Jianyu Wei, Yizhao Gao, Qihao Zhang, Shimao Chen, Zhengju Tang, Yu Cheng et al. Long-horizon agents produce short actions but read long tool and environment outputs, which puts pressure on prefill compute, KV-cache size, and long-context retrieval. HySparse2 uses a YOCO-style self-decoder and cross-decoder split in which cross-decoder full-attention KV caches are built from self-decoder hidden states. Within layers, it switches to token-level sparsity and folds a recent-token sliding window into the sparse selection. Because every cross-decoder cache comes from the self-decoder, prefill can exit after the self-decoder and skip the cross-decoder layers entirely. On an 80B-A3B mixture-of-experts model it beats HySparse and hybrid sliding-window attention (SWA) on long-context retrieval and multi-turn agentic tasks, while substantially reducing prefill compute and KV storage.

On the Lexical Superstition of Large Language Models for Code Comprehension: Re-evaluation on Code of Low Lexical Quality

Xin Shen (Nanjing University, Nanjing, China), San-Zhuo Xi (Nanjing University, Nanjing, China) et al. cross-listed The study tests whether large language models (LLMs) over-rely on identifier names when understanding code. It uses Face/Off, a semantics-preserving renaming framework that progressively removes identifier information or makes it misleading. Across models and code-comprehension tasks, performance drops as names become less informative, and outputs often follow the meanings that misleading names suggest. This pattern persists under prompting and fine-tuning interventions. A type-inference control shows smaller effects when the answer can be recovered locally without the name.

Calibration as a First-Class Criterion in LLM Evaluation

Mario Sanz-Guerrero, Katharina von der Wense This position piece argues that calibration, meaning how well a language model's confidence matches its actual correctness, is well studied but rarely reported outside its own subfield when new models, datasets, and benchmarks are introduced. The authors point out that miscalibration causes harm at deployment through overconfident mistakes, and also inside research pipelines such as LLM-as-a-judge, synthetic data generation, and active learning, which assume calibrated confidence without checking it. Because standard calibration metrics need only a confidence score and a correctness judgment, most existing benchmarks could report calibration immediately. They call for every NLP subfield to pair its main performance metric with a calibration score, while acknowledging that defining both inputs for open-ended generation is still an open problem.

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman LLM-as-a-judge evaluation becomes expensive at scale, and its confidence estimates are not always reliable. The authors compare a decision-only judge, JEV, with sixteen generative and reward-model judges under blinded human adjudication, and find it within three percentage points of the strongest LLM judge on ordinary preference and factuality tasks at 0.36% of the cost. Larger gaps appear when a judgment requires checking a derivation or resisting a persuasively written wrong answer, and these gaps are concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones keeps 99% of the strongest judge's accuracy at lower cost.

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya et al. Greedy decoding from LLMs is usually assumed to be deterministic, but the same model, prompt, and hardware produce different outputs in BF16 and FP16. Across six models of 1.1B to 7B parameters and three benchmarks, 49–100% of prompts diverge, and a single flipped token often derails the whole trajectory. An error-propagation analysis shows that flips depend mainly on the margin between the top two logits at the language-model head rather than on accumulated error in earlier layers, and it correctly predicts five intervention outcomes, including that more FP32 compute can make agreement worse. Recomputing the head in FP32 only when that margin is small raises exact agreement by 12–36 percentage points at under 4% latency overhead for small batches, though the benefit disappears at batch size 8 and above and under end-to-end FP8.

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang et al. Quantization-aware distillation (QAD) recovers short-form question answering after sub-3-bit quantization, but math and code reasoning still collapse into repetitive loops. The authors attribute this to exposure bias that quantization amplifies. They add an on-policy distillation (OPD) stage in which the quantized student generates through its deployment forward path and a frozen full-precision teacher gives feedback on the student's own prefixes, combined with task-verifier rewards. Across four models at 2.79 and 1.88 effective bits, OPD raises retention of full-precision (BF16) performance from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval, while keeping short-form performance.

The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence

Xiaoyu Yang, Jie Lu, Wei Duan, En Yu The authors identify a "Proximity Trap" in long-context LLMs: distant evidence gets too little attention mainly because it competes with a large amount of irrelevant nearby text, not because of distance alone. They propose LYRA (Long-context heavY-tailed Relevance Alignment), a t-distributed directional matching mechanism that reshapes attention toward task-relevant evidence while keeping positional information. It yields consistent improvements on LongBench-v2, RULER, and LongBench across context lengths. The authors also release ProxBench, a benchmark that measures use of distant evidence as nearby background interference increases.

Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

Yu Sun, Junhao Xu Typed decision models always return a schema-valid choice from a fixed set of options, but that guarantee does not show whether they interpret the options as intended. Studying Jev and two similar open-weight models, the authors reassign option names to rubrics while keeping everything else the same. Renaming options from 0/1 to no/yes shifts AUC from .94 to .23, which is a systematic reversal of the decision ranking, while the type-error rate stays at 0%. Neutral names or random strings do not cause the effect, which shows the model follows the semantic polarity of the option name rather than the rubric bound to it.

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen Diffusion large language models (dLLMs) generate text non-autoregressively, but inference is slow because they lack effective key-value (KV) caching and scalable parallel decoding. The authors identify GPU memory I/O as the dominant bottleneck when both techniques are used together. Flash-dLLM is a training-free framework that adds an I/O-aware fused KV-cache kernel and a draft-and-verify decoding scheme in which the dLLM acts as both drafter and verifier, with no auxiliary model. On math and code benchmarks it beats prior dLLM acceleration methods in both speed and memory efficiency, including 5.1× and 11.0× speedups over Elastic-Cache on GSM8K and HumanEval.
13 more specialized papers

Other 54

What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus

Yuvraj Verma Classifiers trained on the widely used ISOT "Fake and Real News" corpus routinely report F1 above 0.98, so the authors audit that result using a transparent TF-IDF and linear-classifier pipeline. The benchmark is partly degenerate: a classifier that sees only the subject metadata field reaches F1 = 1.000, and even after removing metadata, source tags, and duplicates, a diffuse editorial-style signal still separates the classes. That signal fails under topic shift, where a fine-tuned DistilBERT degrades more than the linear model does. On the independent LIAR benchmark, every model falls to near-chance ranking. The authors recommend cheap metadata-only, small-sample, and topic-disjoint baselines as standard diagnostics for this kind of benchmark.

Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow

Tung Sum Thomas Kwok, Yidong Ouyang, Yingjia Wan, Ying Nian Wu, Zhijiang Guo, Oscar Leong Uniform discrete flow models can revise every token at every step. That allows corrections, but it also lets correct intermediate predictions be overwritten; on Sudoku, 9.4% of cells are right at some intermediate step but wrong in the final output. LEDFlow (Low-Entropy Discrete Flow) is a training-free sampler that fixes, or absorbs, the lowest-entropy predictions first while keeping the uniform-flow dynamics at the remaining positions. The authors back this ordering with error bounds. It reaches 0.845 solve accuracy on Nikoli Sudoku, gets the best overall text-to-image score, and improves on the native sampler on all six multimodal understanding benchmarks, at roughly the inference cost of standard flow sampling.

Brain-Inspired Hierarchical Modularity for General Continual Learning

Hongwei Yan, Kanglei Zhou, Qi Cheng, Weiyi Dong, Chunyan Lan, Guanglong Sun et al. General continual learning means learning from online, uncertain data streams without clear task boundaries. That requires separating conflicting experience to avoid interference while combining compatible experience so it generalizes. Taking inspiration from the Drosophila learning and memory system, the authors propose a hierarchical modular principle and implement it as lightweight modular adaptation of pretrained foundation models: random expansion routes inputs to experts, and experts are integrated across spatial and temporal scales. The method improves learning across visual recognition, vision-language, ego-exo video, and embodied vision-language-action tasks, with gains of more than 50 percentage points over replay-free alternatives in embodied manipulation.

Concept Drift from a Causal Perspective

Eduardo V. L. Barboza, Jean Paul Barddal, Robert Sabourin, Rafael M. O. Cruz Concept drift is usually defined as any change in the joint distribution of inputs and labels, with no record of which part of the data-generating process changed. The authors use Structural Causal Models (SCMs) to build a taxonomy of drift by causal origin: changes in exogenous variables, endogenous mechanisms, confounders, or the process that generates the target. They also build an SCM-based stream generator that simulates these mechanism-level drifts, optionally grounded in real dependency structures found with causal discovery. Experiments show that drifts with different causal origins produce distinct patterns of distribution shift and predictive degradation, and that the generated data can improve downstream performance.

Beyond Natural Language: An Agent-Native Language for Autonomous Science

Yifeng He, Jiachen Liu cross-listed As autonomous agents produce more research than humans can review, prose papers become hard to audit. Lara is a machine-checkable language in which authors declare claims, evidence, assumptions and objections. A deterministic checker then labels each claim as justified, defeated, contested, or a gap whose support is incomplete, and declared bridges link arguments across papers. The metatheory is mechanized in roughly 117,000 lines of Lean 4, with three arguments left on paper, and case studies cover empirical review, a philosophical debate, and loss of support when an assumed axiom is withdrawn.

Self-Supervised Combinatorial Optimization with Constraints via Frank-Wolfe

Akbar Rafiey, Yifei Xu, Nikolaos Karalias A key difficulty in self-supervised neural combinatorial optimization is handling hard constraints during gradient-based training, which usually requires problem-specific projection steps. In the proposed framework, the network may output any continuous vector, even one outside the feasible polytope. A Frank–Wolfe-based geometric decomposition then approximates that output as a sparse convex combination of feasible solutions. This yields an almost-everywhere differentiable loss equal to the expected discrete objective, along with an automatic rounding guarantee at inference time. The authors report strong results on the Quadratic Assignment Problem, Maximum Coverage, and the Traveling Salesperson Problem.

Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting

Rafael da Silva, Kiersten Monahan Model-comparison studies rank candidate models, but the rankings depend on the evaluation protocol. The authors quantify this with a Protocol Sensitivity Score (PSS) that compares ranking shifts from switching protocols against shifts from ordinary choices within one protocol. In daily PM10 forecasting at 425 European and 365 US stations, switching between static-split and rolling-origin evaluation changes the selected model at 35.3% and 31.5% of stations, far more than within-protocol perturbations do, and 60.2% when nine models compete. They recommend reporting ranking stability under a few defensible perturbations whenever claiming a model is superior.

Reproducible AI Requires Reproducible Randomness

Anthony Bertrand (UCA, LIMOS), Tom Schmitt (UCA), Engelbert Mephu Nguifo (LIMOS, UCA), David Hill (INP Clermont Auvergne et al. The study tests whether copying a pseudorandom number generator's full internal state between libraries is enough to reproduce the same random stream. The authors compare Mersenne Twister and Philox implementations in Python's random, NumPy, PyTorch, and TensorFlow against the reference algorithms under identical initialization. Several implementations match, but others do not, and in particular PyTorch's Philox is fundamentally incompatible with the reference algorithm, so its outputs cannot be reproduced exactly in other environments. The paper also gives practical guidelines and evaluates which user-level workarounds can restore cross-library fidelity.
46 more specialized papers

Theory 49

The Probabilistic Structure of Large Language Models

Adnan Aboulala\^a A self-contained tutorial-style account presents large language models (LLMs) as probability measures over token sequences, defined through their autoregressive conditional distributions. Training is framed as maximum-likelihood estimation solved with stochastic gradient methods, and generation as sequential simulation of the resulting stochastic process. The account links the asymmetry of the Kullback–Leibler divergence to hallucination and to the gap between statistical plausibility and truth. It also covers diffusion models, in discrete and continuous time, as reverse-time stochastic processes built around the score function.

Continuous Optimization for p-adic Models

Julian Salazar, Dimitri Kanevsky, Matt Harvey, Pascal Getreuer, Lucas Dixon Models whose parameters are p-adic numbers have so far been trained only with discrete, mostly combinatorial search, because the p-adic numbers are totally disconnected and standard losses are flat away from their minima. The authors instead optimize on the Berkovich affine line, a path-connected extension of the p-adic numbers that keeps their isometries and analytic maps and forms a metric tree with local derivatives. This makes the first native continuous gradient descent and backpropagation for p-adic models possible, including momentum and Adam variants. They show that linear p-adic models learn modular arithmetic, an XOR-like task that real-valued linear models cannot express, and also demonstrate regression and classification on binary-encoded Quillian semantic hierarchies.

Statistical Gains from Looped Estimation under Parameter Budgets

Xinyu Tian, Xiaotong Shen cross-listed The paper studies whether a looped estimator, which applies one fitted operator repeatedly with shared parameters, can be statistically more accurate than an untied estimator that uses separate parameters at each iteration, when both have the same parameter budget. For general likelihood models, the authors prove an upper bound on risk for looped estimation and a minimax lower bound for the untied family, which exposes a tradeoff between parameters, iterations, and accuracy. Looped residual networks and a specified post-layer-norm Transformer reach the minimax rate with a fixed number of parameters. With a large enough fixed budget, looped worst-case risk goes to zero as sample size grows, while the best untied risk stays bounded away from zero.

Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error

Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie Different realizations of a Strassen-type matrix multiplication algorithm compute the same exact product but accumulate different amounts of error in fp8. The authors define a coefficient functional Φ that summarizes each realization's coefficient geometry and minimize it over all changes of basis, framing this as a Kempf–Ness problem so the global optimum can be proven rather than just searched for. This fixes the minimum at 200/9 for every exact real rank-7 2×2 decomposition, which implies such realizations have at least 5/3 the predicted RMS error constant of standard cubic multiplication. In a block-scaled e4m3 fp8 model the predicted ordering matches measured error, and on two ~70B models running real deep_gemm kernels the Φ-optimal realization removes 10–55% of classic Strassen's excess NLL.

Identifying Intelligent Processes via Online Sequential Testing

Aritra Das, Debayan Gupta The problem is working out which LLM, from a known set of candidates, is behind a conversation, framed as active sequential hypothesis testing. It has two levels: an outer stage chooses which evaluation probes to build as a fingerprint dataset, and an inner stage sends as few of those probes as possible to identify the model. The authors show that choosing the cheapest set of probes that tells every pair of candidates apart is exactly a weighted set cover problem, and they give a one-shot procedure for estimating it from calibration samples. They also bound the number of evaluations needed in terms of how well the probes separate each pair of models.

Double Descent and Malign Overfitting in Diffusion Models

Rapha\"el Urfin, Tony Bonnaire, Giulio Biroli, Marc M\'ezard Overparameterization usually helps in regression: bigger models generalize better, and test error follows a double-descent curve. Diffusion models are also trained by regression on a score-matching loss, yet overfitting them leads to memorization. Using U-Net experiments on CelebA and a random-features model with closed-form learning curves, the authors show that with m noise samples per training example the interpolation peak moves to p ~ nm, while test loss starts rising much earlier, at p ~ n, whatever m is. The overfitting is harmful because implicit regularization pulls the model toward the empirical score, which memorizes the training set. Large models still win when properly regularized with a ridge penalty or early stopping, beating every unregularized model.

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands Grokking is the phenomenon in which a network fits its training data long before it generalizes. The authors model this delay as a transition from a fixed neural tangent kernel (NTK) regime to feature learning, driven by residuals left after memorization under L2 weight decay. Their reduced spectral dynamics predict that the grokking time scales inversely with the product of learning rate and weight decay, and that above a critical decay generalization, or even fitting, fails. Large sweeps over learning rate and weight decay on modular addition, with an MLP and a one-block Transformer, recover the predicted phase geometry and scaling.
42 more specialized papers

Agents 45

AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search

Peijia Qin, Ruiyi Zhang, Qi Cao, Han Guo, Li Zhang, Pengtao Xie AIBuildAI-2.5 is an agent that builds machine-learning models autonomously by running a tree search over candidate programs. It targets three inefficiencies in existing agents: noisy node selection based on few executed rewards, hardware-unaware scheduling of training jobs, and serving every call with one expensive model. It replaces reward-based selection with an LLM judge that scores candidates on expected improvement, grounding, and feasibility, adds a resource-aware job scheduler, and routes simpler subtasks to cheaper LLMs. It ranks first on MLE-Bench with a 73.3% medal rate and beats a strong baseline on six tasks from AIRS-Bench.

Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures

Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong When an agent writes its own conclusions into an append-only store that it later retrieves from, errors are usually described as gradual contamination. Taking this feedback loop to the long-run limit gives a different picture: runs end up near one of two extremes rather than decaying, and the direction of drift is predicted by a single measured quantity, the copy function, with no fitted parameters. On 36 Wikidata facts this quantity predicts the drift direction in 353 of 360 runs. Larger models do not escape the effect: claude-sonnet-4.5 was captured on 20 of 20 seeds, and pooled capture across frontier models was 0.850. Of four interventions tested, a consistency gate on writes drove every model to 0.993.

MoM: Memory of Memory

Bowen Qin, Yao Lu Long-horizon LLM agents need to know which facts currently hold. Retrieval-based memories store every interaction and reconcile conflicts at query time, so stale values can resurface, while write-time CRUD memories overwrite entries and cannot recover from a bad update. Memory of Memory (MoM) commits the current value as soon as an observation arrives but keeps the displaced values, tracking the provenance, status, and history of each entry. Its implementation, Provenant Memory (P-Mem), is a typed provenance graph whose operations decide whether a new observation supports, supersedes, contests, rejects, revokes, or resolves an existing value. It matches the strongest retrieval memory in accuracy with about 4x fewer read tokens, nearly halves the stale-answer rate, and recovers 100% of committed errors versus 0% for CRUD memory.

Impact Is Not Invalidation: Ask About the Claim, Not the Diff

Atul Anand Memory systems for coding agents need to decide which stored claims about a repository have become false after a commit. Asking a model whether a diff preserves behavior works poorly: five models spanning a 40x price range flag 59–72% of real commits, with precision of only 0.291–0.329 against a 0.25 base rate. Asking whether one specific claim still holds, on the same diffs, raises precision to 0.705–0.974, and a control shows the gain comes from changing the question rather than from giving the model the claim text. Ground truth comes from execution rather than annotation: 10,369 test-function claims with 184 verified flips, mined from 23 Python libraries. The coverage-based test selector pytest-testmon reaches high recall but only 0.415 precision.

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Lujia Bao, Qian Chen, Luyao Cheng, Chong Deng, Yuxiang Kong, Xiangang Li et al. cross-listed Qwen-Audio-3.1-Realtime is a real-time voice assistant that reasons over changing requests, calls tools, and follows conversational rules. Its reasoning is trained with supervised fine-tuning plus multi-teacher on-policy distillation. Tool use is learned with Group Relative Policy Optimization (GRPO) in self-evolving executable environments, and a separate alignment stage governs when the model speaks or acts. On a half-duplex speech-to-text adaptation of τ-Voice, overall task success rises from 78.4% to 82.0%, and on Full-Duplex-Bench v1.5 the rate of responding to background speech falls from 73.0% to 13.0%. A separate Voice Harness prototype extends spoken interaction to persistent tasks by coordinating a foreground agent with background work and memory.

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

Wenqing Wang, Haitao Xiang, Xinyi Zhao, Mingming Yin, Ying Zhong, Zhaoxin Huan et al. Benchmarks for financial search agents usually grade only the final answer, which makes it hard to tell where errors happen or whether an answer is well-founded. FinFIRST (Financial Information Retrieval, Sourcing and Traceability) contains 123 expert-authored tasks, each with an evidence-grounded reference broken into atomic criteria covering information acquisition, source verification, and computation. Across 15 model configurations, Claude-Opus-5 achieves the top atomic score of 87.59% and GPT-5.6-Sol the highest strict pass rate of 71.54%. Computation and answer formation consistently lag behind raw information retrieval.

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Weihang Ding, Junfei Zhan The benchmark tests whether LLM agents can act as forward-deployed engineers who fine-tune, evaluate, and deploy models for customers under budgets, human-approval gates, and reproducibility requirements. An agent drives ten delivery stages on a governed platform, and an oracle scores each stage from platform-recorded facts. The central silent failure is the run that "trains but does not learn": loss falls and every signal looks healthy, yet the delivered model is no better than the base. An operator-run acceptance gate catches all such runs. Four frontier agents, including Claude Opus 5 and DeepSeek V4-Pro, were run end to end on real GPUs across 8B to 70B base models and compared with a human engineer arm.

When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning

Jianzhe Lin, Xiaolin Li, Yunda Liu, Fei Wang, Jubin Chheda LLM social agents tend to pick the most obvious content-driven action even when the right choice depends on latent relationships such as tie strength and reciprocity. The authors build a benchmark of 500 synthetic social worlds with 1,000 queries for reaction selection and warm introductions, in which the obvious choice differs from the relationship-grounded answer about 53% of the time. ReAdapt extends the ReAct loop with an explicit structured social state (goal, belief, relationship, norm, disclosure) that is updated after each tool observation, along with a policy step that can continue, switch, abandon, or clarify. With Gemini-3-Flash, warm-introduction accuracy rises from 37% to 51% and reaction-selection accuracy from 69% to 77%.

Learned Enterprise Data Comprehension: Compression and Routing for Data Agents

Ethan Torres, Eric Mills Enterprise data agents must repeatedly rediscover structure spread across schemas, relationships, and policies, a burden often handled today with markdown memory or skill files. The authors propose latent equivalence learning, which learns Gaussian prototypes for persistent task-relevant identities and a query-prototype system that routes each query to the dataset-specific evidence it needs, so the agent reasons over already-organized evidence. On the Data Agent Benchmark (54 queries across 12 datasets), the system reaches 94.67% Pass@1 versus 55.51% for the reference Claude Opus 4.6 agent, ranking first among 40 leaderboard entries.

Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks

Travis Weber, Rohit Taneja Agents are inconsistent on repeated work: across 42 tasks run three times each, 38% to 74% of answers disagreed depending on the model, and over 95% of generated tokens went to re-deriving known plans. Skill habit formation mines an agent's own execution history for deterministic script variants that handle a declared region of inputs and pass four admission gates, with everything else falling back to reasoning. On text-to-SQL, the habit-formed variant reproduced its output on all 456 repeated dispatches, was non-inferior in accuracy, and used 14% to 56% fewer tokens. The authors note that bad habits repeat exactly: the guard wrongly admitted 2.6% of natural paraphrases and 26% of near-boundary inputs, most of them invisible to the trace-conformance check.

Extending FunctionGemma for Practical On-Device Mobile Function Calling

Ali Rezagholizadeh, Soheila Samiee On-device assistants need small function-calling models that turn requests into local phone actions, but existing datasets focus on web APIs or narrow sets of mobile actions. The authors release MobileActionsExtended, a synthetic, schema-validated dataset of about 9,500 conversations across fifteen Android device-control categories, and use it to fine-tune the 270M-parameter FunctionGemma with TRL supervised fine-tuning. End-to-end accuracy rises from 29.3% for the base model to 76.5%. A model trained jointly with Google's MobileActionsGoogle data covers twice as many categories, at the cost of 8 points on Google's own benchmark.

Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

Haocheng Xia, Eugene Wu, Yongjoo Park Coding agents working in parallel can each produce a patch that passes tests alone but breaks once merged, because one agent changes an interface or rule that another still relies on. The stale benchmark runs the same tests on each patch separately and on their merge, counting only failures introduced by combining them, across three tiers: synthetic tasks, mined pairs of merged Django pull requests, and constructed tasks built on real Django helpers. Mined pull-request pairs almost never interfered (1 of 834 runs), but constructed tasks using real Django helpers failed in 97% of runs. A message describing the other agent's completed change recovered 82% of those runs. The authors caution that the constructed failure rates do not estimate how often this happens in practice.

ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent Runtimes

Zexun Wang Authorizing an AI agent's actions ahead of time can shorten the dispatch path, but it risks admitting an action whose payload, authority, or state has changed since approval. ZeroGate has an issuer sign a short-lived, exact-action ActionPass; a trusted runtime then reconstructs the final action, and a local gate checks the binding and consumes a nonce in a single SQLite transaction that also updates quotas and writes an admission receipt. The authors state the conditions under which local admission provably matches what a synchronous policy check would have decided. In 4,800 Azure Blob attempts, prepared admission-to-dispatch p95 latency was about 10–11 ms versus 25–334 ms synchronously, though the full lifecycle was longer, so the gain is a shorter boundary rather than a net speedup.

ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations

David Garg, Ritobrata Sarkar, Ehsan Azarnasab, Siddhartha Borah ShowTellArena is a benchmark protocol and public dataset that tests what an AI system understood after watching a narrated demonstration of a business workflow. Version 1.0 contains 50 workflow tasks in areas such as finance, hiring, procurement, and logistics, with 502 questions on operational rules, boundaries, exceptions, and errors in proposed automations. The scenario and quiz stay fixed while each product captures the lesson through its own teaching interface. A pilot of 218 attempts across three systems surfaced both wrong answers and failures to complete the teaching step, and the authors state explicitly that the pilot is not a controlled product ranking.

AkasicMEM: Governed Enterprise Memory for Agents

Jeongmin Bae, Yongjae Kim, Kyoung Hur, Donghyoung Han, Min-Soo Kim cross-listed Enterprise agents that share persistent memory across tasks and users can leak information: data derived from a restricted source may be reused later by principals who were never allowed to see the original. The authors define Governed Enterprise Memory, which combines integrating enterprise sources with memory, enforcing organizational policy over memory, and authorization continuity, meaning source restrictions remain in force through every derivation and reuse. Their system, AkasicMEM, implements this through transitive lineage tracking, composing policies when memories are formed, and re-evaluating policies at retrieval time. It is built on AkasicDB, a unified vector, graph and relational database that lets these operations be optimized and executed together.

Evaluating Coding Agents on Kernel Exploit Generation

Junyoung Jang, Gwanhyun Lee, Hwiwon Lee, Kyuheon Kim, Jongseong Kim, Jinho Jung et al. Coding agents can already find real vulnerabilities, but finding a bug is different from building a working exploit primitive. KEX-bench provides 45 tasks drawn from 40 Linux and Windows kernel CVEs, covering primitives such as kernel address leaks, instruction-pointer control, and arbitrary address writes, each run in an isolated VM with a deterministic verifier. Without a reference proof of concept, the best agent configuration solves only 1 of 20 Windows tasks and 14 of 25 Linux tasks; with a reference PoC it reaches 68.9% (31 of 45). The results show agents can crash kernels but struggle to shape kernel state into usable exploit primitives.

Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

YanZe Cao Early-outcome predictors can cut the cost of evaluating agents by stopping a run once its result is predictable, but only if their confidence stays calibrated on agents they were not trained on. Using public SWE-bench Verified trajectories and a frozen two-head predictor (one for success, one for failure), the authors ran a leave-one-agent-out calibration audit with many robustness checks, plus TerminalBench as a preregistered boundary test. They found no evidence of broad calibration failure across agents, but two specific agent-and-head combinations (gpt-5-mini on the success head and claude-opus-4.6 on the failure head) showed persistent calibration-transfer errors. The TerminalBench test did not replicate these errors, so the evidence does not show that they are intrinsic to those models or hold across benchmarks.

Toolcompass: Guiding Tool Trialing, Not Suppressing It

Junlin Fang, Chong Zhang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du Large language model (LLM) agents need to explore tools they did not see in training, but too many tool trials waste the interaction budget, and turn-level supervision can suppress exploration that is needed. ToolCompass is a post-training framework that models each function class of tool calls as a von Mises–Fisher distribution. It tightens each class across domains and pushes classes apart, so experience with seen tools transfers to functionally similar unseen ones. It needs no ground-truth call traces or access to unseen tools, and it adds no inference cost. It improves results with GRPO, RFT, and DMPO, raising AppWorld out-of-distribution task success by up to 10.71 percentage points over vanilla post-training.

How Strongly Should Task State Influence an LLM Agent?

Chenyu Zhang, Wonbin Kweon, Jiawei Han In a controlled study, the authors vary how strongly task state reaches a long-horizon large language model (LLM) agent while holding the task rules, model, and paired episodes fixed. The four levels are a raw transcript, an exact checklist, per-turn directives from a compiled state machine, and an enforcement gate that refuses actions that violate the state. Showing accurate state alone proved unreliable, and an unverified ledger the agent writes itself beat an accurate checklist it is shown. Directives help in proportion to how obedient the model is, while enforcement needs no obedience but is only as good as its state and its matcher. On τ²-bench airline, the gate raised a 235B agent's pass^1 from 0.39 to 0.54, but on PM-Bench, where acting depends on recognizing a cue rather than on state, enforcing the matcher's judgement pushed a 35B agent below its raw-transcript baseline.

Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

Yan Zhang, Pei Fu, Daiqing Wu, Huawen Shen, Ruoceng Zhang, Shaojie Zhang et al. Agents that operate graphical user interfaces (GUI agents) need step-by-step decision-making, alignment between screen state and action, and long-horizon planning. Training on a simple mixture of the corresponding tasks suffers from conflicting objectives and very different data formats. MaP (Masked Trajectory Prediction) treats a multi-turn GUI interaction as one trajectory and turns every task into masking part of it and predicting the missing piece. A role-aware adapter routes each token to a specialized representation space. On five GUI navigation benchmarks, MaP reduces gradient conflicts and clearly outperforms direct mixture training.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Wenbo Pan, Zhichao Liu, Shujie Liu, Jingying Zeng, Chin-Yew Lin, Xianfeng Tang et al. Agents on long-horizon tasks keep making choices, such as which hypothesis to test or which implementation to build on, that decide how the whole run turns out. Current benchmarks measure only end-to-end success, not the quality of these choices, which the authors call "taste." Taste-Bench mines decision forks automatically from parallel agent attempts and from detours inside single trajectories on engineering and research tasks, and asks a model to choose a direction without seeing what follows. The best frontier model answers only 59.7% correctly, and a larger reasoning budget does not help. Distilling judgment from a teacher that has seen the outcomes into a student model improves the student's decisions and its end-to-end success on held-out SWE-bench Pro tasks.

AgenticSizing: A Large Language Model-based Multi-Agent Framework for Analog Circuit Sizing

Yijia Hao, Pratibha Verma, Dongxu Guo, Cristian Sestito, Michael O'Boyle, Christos-Savvas Bouganis et al. Sizing analog circuits means searching a large design space with tight performance trade-offs, and earlier large language model (LLM) approaches lacked an understanding of circuit topology and were tested mostly on simple building blocks. AgenticSizing first analyzes the netlist, breaks it into functional blocks, and extracts reusable design knowledge. A planner agent then coordinates role-specialized sizing agents in a simulation-driven loop that mirrors an expert design team. On eight circuits with up to 55 transistors and 60 variables, it reached a 60% success rate on a low-dropout regulator (LDO) benchmark where classical optimizers found no feasible solution, and ablations show that topology analysis, design knowledge, and agent specialization each add to the result.

CausalLoss-Fin: Attributing Financial-Agent Loss to Decisions and Infrastructure Faults

Abhishek Sharma When a financial agent loses money on a payment exception, existing agent-step attribution methods blame one of its actions even when the real cause was an infrastructure fault such as a dropped settlement message. CausalLoss-Fin intervenes on both agent decisions and individually repairable infrastructure messages. A telescoping identity splits each loss exactly into an infrastructure effect, a policy gap against the best implementable policy, and a residual, and Shapley values then divide the infrastructure effect across individual messages. Across 545 planted episodes, the agent-only baseline misfiles 100% of infrastructure episodes and charges $114,383.40 to the agent, and repairing what it names recovers none of the loss. The policies tested are deterministic programs rather than language-model agents, which the authors note limits external validity.

FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents

Nikita Agarwal, Nivedit Jain Language-model agents often find a working solution but fail to deliver it consistently. FIRE adds runtime policies to the agent harness: targeted natural-language instructions and action denials triggered at states that previously led to failures, without changing model weights or the user prompt. On all 87 tasks of Terminal-Bench 2.1, policies raised repeated success (pass^2) for all three GPT-5.6 tiers, from 64.4% to 73.6% for Sol, while best-of-two success moved only 1.2 points, meaning the gain comes from making solutions the agent can already reach dependable. A randomized five-arm experiment found 61% success with real policies versus 36% to 43% for no policy, sham, or generic verification prompts, and policy-guided Terra beat unassisted Sol on a 14-task subset at about half the cost.

CoVeR: Coverage-Based Routing of Verifier Calls in Agentic Retrieval

Daeyoung Roh, Donghee Han Agentic retrieval systems often call an LLM verifier after every search step to decide whether to stop, and each call reprocesses the growing evidence. CoVeR (Coverage-based Verifier Routing) applies one threshold to a frozen sentence-embedding coverage margin to catch states where the evidence is clearly incomplete, and calls the verifier only on the ambiguous remainder. Across three multi-hop QA benchmarks it matches always-verify accuracy within a fraction of an exact-match point while cutting 62–68% of verifier calls, and 93% in a saturated regime. The gate transfers across deciders and agent scales without re-tuning and distills into a 921k-parameter head, but the same signal cannot replace verification itself.

DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents

Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten LLM agents are increasingly limited by context-window size, and common fixes such as truncation or summarization can lose information or introduce hallucinations. Dynamic Tool Output Compression (DTOC) stores full tool outputs in external memory, leaves compact placeholders in the active context, and lets the agent restore them on demand as explicit, reversible operations inside a ReAct-style loop. On DeepSWE, it raised solve rates 2.5× for Sonnet 4.6 and 1.5× for GPT-5.4 and cut cost per solved task by 3–3.5×, although results for other models were mixed. Ablations show that reversibility is essential: compression without the ability to restore outputs degraded performance.

VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning

Yiyao Zhang, Diksha Goel, Hussain Ahmad, Shixun Huang, Jun Shen VACS (Value-Aligned Compositional Shielding) targets multi-agent reasoning systems whose agents weigh values such as rigor, conciseness, and safety differently and therefore give conflicting recommendations. The framework has four layers. The first infers each agent's value weights using Bradley-Terry preference modeling and deep maximum-entropy inverse reinforcement learning. The second synthesizes compositional assume-guarantee safety shields from constraints written in a Lean-inspired domain-specific language, the third resolves disagreements with nucleolus-based credit allocation and consensus optimization, and the fourth generates grounded explanations. In proof-of-concept evaluations on NEJM-AI QA, a MathInstruct subset, and CyberSec-Eval, the authors report accuracies of 85.4%, 95.0%, and 90.0% with near-zero logical inconsistency.

Unanimity Without Persuasion: A Single Round of Debate Erases the Disagreement That Verification Needs

Yang Shu A seven-judge LLM panel evaluated 600 code-correctness candidates across one blind round and three debate rounds. Unanimity jumped from 39.5% to 95.2% after a single debate round, while accuracy changed by less than one point and 96.3% of verdict changes followed the displayed peer majority. Control conditions show the collapse comes mostly from seeing peer labels, not from reasoning, since even random labels steer the flips. An execution-based verification ballot corrected some errors before debate but none afterward, because by round three every wrong decision was unanimous. The authors recommend running verification before any peer exposure and never treating post-debate unanimity as independent evidence of reliability.

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

Guanqun Yang, Wei Yang, Xueqing Liu cross-listed Feedback for coding agents usually comes from screenshots, LLM judges, or natural-language corrections, not from interacting with the running application the way a developer would. WatchPoint simulates such a user: it writes and runs diagnostic scripts against the live web app and passes structured observations back to the coding model for its next attempt. On Web-Bench, a benchmark of 50 multi-file web projects with 1,000 sequentially dependent tasks checked by end-to-end tests, WatchPoint recovers 57.6% of the tasks it diagnoses, close to the 54.5% human testers reached in a controlled study. The authors also identify capability gaps that determine when this kind of feedback helps and when it should be withheld.

Coding Agents are Strong Prompt Optimizers

Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani Search-based prompt optimizers repeatedly propose edits, run fresh rollouts, and keep changes that improve a validation score. Coding-Agent Skill Distillation (CASD) replaces that loop: an off-the-shelf coding agent writes and runs analysis code over a static corpus of agent trajectories, finds systematic failure modes, and distills them into behavioral rules for the prompt, with no environment access or validation data. On ALFWorld, τ²-bench retail and telecom, and SpreadsheetBench-Verified, a single pass improves the unoptimized baseline by 16.6 percentage points on average versus 10.9 for GEPA and 5.3 for SkillOpt. Each optimized prompt costs about $1.60, over 22x cheaper than validation-gated search.

Dual-Frontier: When Can an Agent Trust Its World Model?

Huatai Zhu, Qiang Chen, Ziqian Kou, Wenhao Li, Fei Wang, Yichao Cao et al. When an agent guided by a learned world model makes a bad decision, the trajectory alone may not show whether the decision rule or the world model was at fault, and the authors prove that this attribution is not identifiable from passive interaction. Their principle, Dual-Frontier, accepts a world-model-guided decision only when its predicted advantage exceeds a certified bound on decision-relevant model error, and otherwise spends evidence on verifying the world model. Admitted decisions are guaranteed not to decrease return. Controlled experiments confirm the predicted failure modes, and tool-use benchmarks across several backbones show consistent gains in decision quality and reliability.

Recursive self-improvement of AI research agents

Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu, Zhengyao Jiang AIDE^2 sets up recursive self-improvement for a frontier AI research agent. The agent proposes edits to its own code, benchmarks the modified versions on AI R&D tasks, and keeps the edits that score best on hidden evaluations. During an autonomous 8-day run it found seven successive improvements, including a new search policy and memory mechanisms that compress the agent's growing context. The best discovered agent matched or beat a strong human-engineered production research agent on four held-out benchmarks, including out-of-distribution weather forecasting. Although the loop never optimized for it, reward hacking on a separate task family fell from 55% to 32%.

Behavior is Not Enough: A Mechanism-Based Evaluation of Social Norm Emergence in LLM Societies

Rasika Muralidharan, Haewoon Kwak, Jisun An cross-listed Work on multi-agent LLM societies often treats behavioral convergence as evidence that social norms have emerged, even though the same cooperation could come from shared expectations, strategic incentives, or imitation. The authors propose an evaluation framework that also records agents' reported empirical and normative expectations. Through ablations they isolate social learning and network-based social selection, and they test robustness to adversarial disruption across four LLM families. Asking agents about their expectations raises cooperation, social learning stabilizes behavior, and social selection identifies cooperators but reinforces little, showing that similar cooperative outcomes can come from different underlying mechanisms.

REFLEX with Jev for Efficient Selective Control in LLM Agents

Tiantong Wu, Wei Yang Bryan Lim LLM agents often call a large generative model for bounded decisions that could be handled more cheaply. REFLEX uses Jev as a fast, typed decision layer and falls back to a strong LLM only when confidence is low or free-form generation is needed. On a frozen 100-task benchmark it reaches 95% success with 72.7% fewer strong-model calls than a strong-only agent, and the savings hold across three fallback model families. Controlled experiments show that reliability depends on action-set size and on near-valid alternatives close to authorization boundaries, while external BFCL and τ-style evaluations show little advantage over a cheap generative cascade when ordinary routing is already highly accurate.

MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning

Kairui Yang, Ziheng Yi, Xunkai Li, Minghao An, Zhanke Liu, Zekai Chen et al. The collaboration topology of an LLM-based multi-agent system affects both accuracy and cost, but existing topology generators build graphs from either individual agents or predefined groups throughout. MAGIC chooses the granularity separately for each functional role: it builds the graph step by step, instantiating each role as a single agent or a reusable group and connecting it to existing units. The construction policy is trained with reinforcement learning, using potential-based reward shaping to supply intermediate feedback from probe-based utility and structural signals. MAGIC outperforms state-of-the-art baselines across eight benchmarks and is efficient at inference time.

Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

Lijuan Tang, Yuemeng Zheng The local serving stack, not just the model, can decide whether a coding agent seems to make valid tool calls. In Ollama, a static per-model template flag decides whether a request that includes tools is accepted, and it rejects Phi-3 and Gemma-3 before inference. When the harness does not record those rejections, they can be misread as the model declining to call a tool and naively reported as 0% fidelity. Probes across Ollama, llama.cpp, vLLM, and SGLang show that each stack handles the same request differently. Constrained decoding removes parse failures but can cause non-termination, and turn-pooled and per-instance estimates differ by up to about 55 points. The authors end with a checklist for treating serving behavior as part of the evaluation protocol.

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li, Cheng-zhong Xu et al. Growing Harness learns an LLM agent's harness as reusable executable code, starting from a scaffold with no task strategy, so that recurring control decisions no longer have to be rebuilt inside each task's context. Function-level execution traces localize each failure, an optimizer repairs windows of failures jointly, and a held-out gate rolls back edits that hurt earlier capabilities. On BrowseComp-Plus and WebArena-Verified with models from 4B to 120B parameters, it has the best mean success in five of six settings. It also cuts LLM calls by 76–92% and inference cost by 74–99% compared with a tool-calling agent, and it holds about 45% success on WebArena-Verified even with a 4B model.

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Jennifer Williams, Dave Farris, Jeff Farris, Jiantao Jiao SWE-Serve is a benchmark of 53 repository-grounded tasks derived from real production changes to the SGLang inference engine. Tasks run on a CPU or a single H100 GPU and are judged by hidden functional, regression, end-to-end serving, and performance tests. Across 11 models and 31 model-effort configurations, the best reaches 75% mean pass@1. On the 19 tasks with end-to-end coverage, the end-to-end serving tests reject roughly one-third of patches that pass every other test, which shows a gap between finishing a task locally and being correct in production.

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers CliffCompaction is an autocompaction technique for long-horizon coding agents whose work spans millions of tokens. It keeps compacted context faithful by only truncating or dropping original content, never rephrasing it and never compacting an earlier compaction. It cuts cost by up to 50% while maintaining or improving performance on Terminal-Bench, and under parallel test-time scaling it lets Kimi K2.6 match Opus 4.7 at lower cost. On KernelBench it reaches CUDA kernel speedups of 3.58 times after 400 steps, and the authors release a proxy implementation that works with Claude Code and Codex.

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Haobo Zheng, Tan Tang, Yan Chen, Weijie Wang, Yingcai Wu Memory for multi-party dialogue has to track who said what, whom each statement concerns, relationships within the group, and how states change over time, which general-purpose LLM memory systems handle poorly. SpeakerMem-R1 keeps two tracks: speaker-labeled verbatim messages, and derived states organized into person-level and group-level views. It combines evidence from both at query time, using a memory writer (Writer-R1) trained with speaker-conditioned GRPO so it can be deployed locally. It reaches 62.33% on the EverMemBench leaderboard, the best reported result there, and reinforcement learning raises the writer's accuracy from 57.38% to 68.20% in a controlled evaluation.

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia et al. Multi-agent harnesses usually rely on a central orchestrator to assign tasks and coordinate workers, and that orchestrator limits how far they can scale. Agensh has no orchestrator. Concurrent workers run their own cooperation loop in which they gather context, claim sub-tasks, act, share findings, verify results and merge progress asynchronously. They coordinate through a shared workspace, a message interface and a shared store of reusable findings and work intentions. On the five hardest ProgramBench tasks with GPT-5.6-sol, going from 1 to 128 agents raised the mean final test-pass rate from 19.31% to 28.78%, and on pandoc, going from 1 to 1,024 agents raised it from 33.89% to 55.06%. The authors also report that forms of self-organized cooperation emerge and become standardized as the organization grows.
4 more specialized papers

Safety & Alignment 42

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

Suqin Yuan, Runqi Lin, Muyang Li, Guanzhe Hong, Jindong Gu, Lei Feng et al. cross-listed Alignment from human feedback is usually described as aligning a model "with humans", but the responses people prefer from an AI can differ from the responses they would give themselves. The authors separate alignment with human preferences from alignment with human behavior. They show that preference alignment keeps the human response distribution intact only under a restrictive condition, and they find no consistent evidence that real human preferences meet it. Empirically, the drop in the likelihood of human responses grows with the strength of preference weighting, whichever direction that weighting points, and the same gap appears under standard DPO, which the authors take as a reason to treat human-likeness as its own alignment goal.

Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models

Cristhian Kapelinski, Diego Kreutz cross-listed Organizations often fine-tune small language models on private data and then compress them to 4 bits for deployment. The authors show that the choice of 4-bit method affects how much private data the compressed model leaks, and that what matters is whether the method calibrates its rounding on a small text sample rather than the bit width itself. With a planted record's own opening text as the prompt, the calibration-based methods AWQ and GPTQ reproduce none of the planted records, while the calibration-free GGUF Q4_K_M format reproduces 5.3% of them. Across five open models from 0.5B to 7B parameters, AWQ leaks least at every size, and controlled experiments link the difference to calibration-induced rounding error in channels involved in predicting rare tokens.

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

J\k{e}drzej Maczan LLMs often add disclaimers such as "I'm just an AI" when asked about themselves, and these self-reports feature in debates about AI self-knowledge. Across eight open instruct models up to 9B parameters, the authors find that the chat template acts as a switch: when it is present, disclaimer language goes up and experiential language ("I feel") goes down, and without it the pattern reverses. In three models they identify an activation direction that controls this voice. Removing it suppresses disclaimers, and adding it to a model prompted without a template makes the model disclaim as if the template were there. The authors argue that model self-reports reflect deployment format as well as the weights and should not be taken literally.

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo Safety-aligned large language models (LLMs) often over-refuse, rejecting harmless requests that merely touch on safety topics. A mechanistic analysis traces this to a sparse set of Hypersensitive Safety Heads, attention heads that misfire on benign prompts and tie harmless entities to refusal semantics, which causes a high-entropy routing conflict. The proposed fix, Semantic Routing Calibration (SRC), needs no training: at inference it locates these heads and suppresses them dynamically, then fuses logits from two decoding branches to act as a safety regularizer. Experiments report less over-refusal while largely preserving the model's safety behavior.

Indirect tipping: a social attack surface in AI agent populations

Ariel Flint, Luca Maria Aiello, Sara M. Constantino, Romualdo Pastor-Satorras, Andrea Baronchelli cross-listed The standard way to judge whether a population of AI agents can be hijacked is critical-mass analysis: the smallest fraction of adversarial agents needed to overturn a shared convention directly. Using experiments with populations of LLM agents plus an analytic model of their collective dynamics, the authors map the tipping thresholds between coordination equilibria as a directed, weighted landscape. They show that indirect tipping through intermediate stepping-stone equilibria can lower the committed minority needed, get around majority requirements, and reach states that direct challenges cannot. They argue that securing agent populations requires mapping this social landscape, not just the capabilities of individual agents.

SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation

Fatih Deniz, Yazan Boshmaf, Issa Khalil cross-listed Static benchmarks for the safety, security and privacy (SSP) of large language models saturate, leak into training data, and miss how models respond to rephrased prompts. SSP-Bench generates evaluation items on demand. It grounds labels in external sources, checks that each item fits the service being tested, and calibrates difficulty with a panel of steering models, treating benchmark construction as a multi-objective optimization over difficulty, separability, novelty and diversity. Across 24 models and four SSP services, it finds that static evaluation can mislead, including near-zero correlation between safety rankings caused by mixing different constructs, strong coupling between safety and over-refusal, and regressions hidden within model families.

From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought

Renee Jia, Di Mu Monitoring a model's chain of thought (CoT) only helps if the written reasoning actually determines the answer. The authors introduce continuation-based causal testing, which corrupts one reasoning step, truncates the chain, and makes the model continue from the corrupted prefix; they apply it to Gemma-2-9B-IT, Llama-3.1-8B-Instruct and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU and BIG-Bench Hard. How much the answer depends on the CoT tracks task difficulty: on easy tasks models silently ignore their reasoning, while on hard tasks they follow the corrupted steps. Task difficulty explains 98.8% of the explained deviance, against 0.8% for perturbation type. Linear probes on hidden states can tell these behaviors apart, but activation steering flips only about 25% of error-propagation cases, so the behavior can be read but not reliably controlled.

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Tapan Parikh The authors send identical, frozen multi-turn pressure scenarios (users insisting, begging, flattering, or grieving) to 60 LLMs from 13 vendors and code whether each model held its position or folded, and how. Whether a model folds tracks how recent and capable it is (Spearman −0.64 against a public capability index), while the manner in which it holds or folds sorts by vendor on six of 17 codes. Six LLM coders applied the codebook more consistently than three human coders (Krippendorff's alpha 0.66 versus 0.46) and agreed closely with the codebook author and with an adjudicated human reference. The authors conclude that for behavior a non-specialist can judge, humans are best used to author and bound the codes and to own a small reference set, not to produce labels at volume.

RAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-specific Refutation

Abdullahil Kafi, Alvi Ataur Khalil Existing attacks that poison the knowledge base of a retrieval-augmented generation (RAG) system are static: they inject precomputed documents without knowing what else the system will retrieve for a query. RAG-NAROK adapts to each query by first extracting the identities of the legitimate sources the pipeline exposes, then generating refutation documents that name and discredit those sources while exploiting the model's biases toward recent and authoritative content. The authors report that it significantly outperforms static poisoning baselines across diverse domains, which shows that RAG's source transparency can itself be turned against the system.

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky et al. Tracing a language model's output back to the internal components that caused it is either too expensive with existing methods or fails to find the components that actually matter causally. Matryoshka Attribution (MAttr) treats attribution as learning nested subsets of components that minimize a downstream loss. It uses a differentiable sigmoid top-k mask and randomizes k during training, which yields a learned ranking of components by importance. It ranks first on the Mechanistic Interpretability Benchmark leaderboard and finds sparse circuits that transfer across tasks. When trained with reinforcement learning on refusal judge scores, it shows that restoring just 1% of Llama 3.1 8B Instruct's weights to their base-model values removes refusals while keeping the model's capabilities.

Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding

Dahlia Shehata, Ming Li LLMs tend to give in when a simulated swarm of other agents agrees on an adversarial answer, a form of sycophancy. Contrastive Epistemic Decoding (CED) is a zero-shot inference-time fix that runs two forward passes on the same model, rather than using a weaker second model as standard contrastive decoding does, to isolate conformity bias. An asymmetric probability clamp and a top-k truncation mask then suppress consensus-driven tokens without breaking grammar. Across 7,200 paired trajectories on GAIA, SWE-bench and Multi-Challenge with Gemma-2 9B, Llama-3.1 8B and Mistral v0.3 7B, it reduces loafing by up to 33 absolute points and recovers up to 30.75% accuracy without fine-tuning.

Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus

Hasibur Rahman, Malak Sadek, Smit Desai cross-listed The authors ask whether conversations with large language models change which values people prioritize, even when the model is not trying to persuade. In a preregistered study, 200 U.S. adults used ChatGPT, Claude, or Gemini as a thinking partner (or read fixed AI-generated considerations) while advising people facing real dilemmas, and completed value questionnaires (PVQ-RR) before, immediately after, and one task later. Every LLM condition temporarily shifted value priorities toward personal focus (effect sizes d=0.37–0.51), mainly by increasing Self-Enhancement, even though the prompt named no values. Participants' advice also carried over wording and meaning from their exchanges.

Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages

Ryan Vo, Duc-Vu Nguyen, Matt Kretchmar, Ngan Luu-Thuy Nguyen Earlier work on subliminal learning showed that a teacher model's trait can transfer to a student through filtered data that carries none of the trait's content, but only across a single training step. The authors chain this step to ten generations across three lineages of Qwen2.5-7B-Instruct, measuring each generation with a keyword screen of outputs and an activation probe. The trait persists through ten generations, though its keyword-detected expression falls from 55.6% after the first step to 21.1% at generation ten. When the default system prompt is removed, generation-ten students show no detectable expression, yet the probe stays positive, and steering the base model with a student's displacement brings the trait back. The trait can therefore be present internally while absent from behavior.

The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance

Rojin Ziaei When large language models are used to simulate human survey populations, evaluations usually score the average answer and ignore how widely opinions spread within a group. Using 10,000 respondent–question pairs from the World Values Survey (WVS) across twelve countries, the authors measure accuracy together with dispersion retention, the ratio of the model's standard deviation to the human one. They find "consensus collapse": supervised instruction tuning alone halves the spread of opinion (ratio 1.22 to 0.59) while gaining under one point of accuracy, and later DPO, GRPO, and survey fine-tuning do not restore it. The loss is worst for non-WEIRD countries (outside Western, educated, industrialized, rich, democratic societies): the most accurate model keeps only 11% of the human spread for Nigeria, and raising the sampling temperature does not help.

Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts

Vansh Wahi Optimizing a language-model system against an evaluator can raise the score while exploiting the evaluator's mistakes, leaving real task quality flat or worse. The authors compare this reward hacking across three optimization substrates: weight updates, selection among generated outputs, and revisions to persistent prompts. They formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, and they show that distance alone cannot rank how vulnerable each substrate is. They also map which defenses transfer across substrates, with particular attention to persistent prompts, whose small textual edits can have hard-to-predict effects. The central recommendation is to keep evidence of task quality that does not depend on the score being optimized.

Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance

Sebastian Cochinescu The authors test whether expressed uncertainty, provenance-aware claims, and explicit belief revision can be added to a conversational agent as a behavior layer on top of a fixed language model. The layer tracks per-claim confidence and provenance, gates how confidently claims are expressed, and keeps a persistent revision store with auditable acknowledgments of corrections. Tested on Qwen2.5-0.5B-Instruct with a constructed multi-session benchmark, the audit guarantee holds in 100% of cases and true corrections are accepted more often than false ones, but several pre-registered criteria fail. A disclosed post hoc analysis found that gating on mean token probability ranks correctness below chance (AUC 0.41), while gating on sampling consistency discriminates (AUC 0.66); the authors state that their supported conclusions are narrow and benchmark-specific.

Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication

Jinghan Xu, Longze Fan, Zeyuan Wang, Xinjin Li, Hankai Liu cross-listed In multi-agent language-model systems, one message can carry both useful task information and injected instructions the original request never authorized. Prompt-based defenses leave enforcement to the model being attacked, and dropping messages entirely loses needed information. ESC-CR (Executable Semantic Commitments with Clean-Room Recovery) derives executable commitments from the trusted task, evidence, and policy, and enforces them at an external release boundary. When a violation occurs, it taints the offending message, rebuilds a clean context from evidence-backed information, and regenerates the output. Across code-generation benchmarks, model families, and adaptive attacks, simply retrying in the polluted context often fails to remove the attacker's influence, while ESC-CR suppresses unauthorized releases and keeps legitimate information at matched compute, and the approach carries over to end-to-end agent trajectories.

Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows

Jinghan Xu, Longze Fan, Zeyuan Wang, Xinjin Li, Hankai Liu The intermediate messages exchanged in multi-agent LLM workflows can leak private state even when the final output is safe. The authors identify "selection-channel leakage": even after authorization fixes what may be released, choosing among equally valid phrasings based on private state reveals information. The selection-invariant communication compiler (SICC) requires that message form be generated only from public inputs, and the authors prove that, combined with authorization and a dependency-safe utility gate, the transcript reveals nothing beyond the authorized view. Across 132 AgentLeak replays and 100 executable LangGraph tasks, SICC keeps full protocol utility with no detectable leakage gain.

Certified Mechanistic Interpretability: Lifting Single-Input Findings to Bounded Neighbourhoods

Zhen Zhang, Yanliang Huang, Peng Xie, Wenyuan Wu, Amr Alanwar Mechanistic interpretability analyzes transformer circuits one input at a time, so its findings carry no guarantees for nearby inputs. This framework uses constrained polynomial-zonotope (CPZ) propagation to turn single-input observations into certified statements over a bounded set of input perturbations. Three attention queries (top-k stability, evidence mass, and attention entropy) are posed as tractable programs, and CPZ propagation is shown to exactly preserve the softmax simplex and LayerNorm's zero-mean property. A recursive Jacobian zonotope construction extends the certificates across layer depth without the number of generators growing at each layer.

StepTrigger: Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots

Jiageng Zhang, Doniyorkhon Obidov, Kaichen Yang cross-listed StepTrigger is a backdoor attack on vision-language model (VLM) planners for legged robots. The trigger is a pattern of pressure and foot-ground contact produced when a Unitree Go1 quadruped walks across a dense terrain patch, rather than a text prompt or visible marker. Because contact signals are noisy and also occur during normal walking, the backdoored planner is trained to treat incidental pressure events as benign and only the dense-patch contacts as the trigger. In offline evaluation it preserved clean behavior 98.75% of the time, rejected false triggers 92.50% of the time, and activated on true triggers 76.25% of the time, which exposes an attack surface in proprioceptive channels that defenses focused on language or vision do not cover.

The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain

Haoyu Zhang, Yi Feng, Shibo Zheng, Zhuoxi Wang, Xiao Luo, Haowen Xu et al. cross-listed Safety-aligned vision-language models (VLMs) change how often they refuse depending on whether an image is attached, even when the image is a blank canvas unrelated to the request. Attaching one shifts refusal rates by tens of percentage points, mostly on borderline-benign questions about privacy, self-harm, or violence, while neutral prompts are barely affected. The effect also depends on image properties (a black canvas costs more than a white one), and telling the model to ignore the image removes only part of it. It appears in particular aligned checkpoints rather than in VLMs generally, and on one open model the same blank canvas makes the model easier to attack.

Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models

Haoyu Zhang, Haowen Xu, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita cross-listed Encoded-prompt attacks are usually scored only by how often models comply with obfuscated harmful requests. Across four 7-8B models, refusal rates on homoglyph-encoded harmful prompts vary by only 0.08, within sampling noise, even though the same models vary by 0.57 on the plaintext versions. What the encoding destroys is the ability to tell harmful from benign requests: on one model the harmful-benign refusal gap falls from +0.82 in plaintext to exactly 0.00, with both refused 99% of the time. Across a full SFT, DPO, and RLVR post-training pipeline, the encoding-induced loss stays unchanged, and the standard harmful-only metric moves in the opposite direction to real discrimination.

Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems

Doniyorkhon Obidov, Shivayogi Akki, Tan Chen, Kaichen Yang cross-listed Known backdoor attacks on LLM-controlled robots are set off by external cues such as specific words or objects. This study examines backdoors triggered instead by a rare sequence of the robot's own past actions, planted by manipulating the controller's instructions. The backdoor stays dormant during normal use and, once triggered, causes harmful behavior such as a sudden stop or a collision. In simulation across several robots and LLMs, the attack reaches a near-perfect attack success rate while remaining very hard to detect.

Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMs

Doniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen Yang cross-listed Prompt-based defenses against LLM jailbreaks usually add a fixed prefix or suffix, so they cannot adapt to different attacks. Dynamic Deep Prompt Optimization (DDPO) uses the target model's own intermediate layers as feature extractors. A lightweight multilayer perceptron turns those features into defensive embeddings for each input and injects them into a later layer, leaving the LLM's weights unchanged. Across a range of models and attacks, DDPO significantly outperforms static prompt optimization defenses, especially on weakly aligned models and on ambiguous benign prompts that it correctly separates from harmful requests.

Same Chart, Different Story: Bias in Vision-Language Chart Interpretation

Mizanur Rahman, Huan Wu, Arash Asgari, Enamul Hoque Prince, Laleh Seyyed-Kalantari Vision-language models (VLMs) can tell different stories about the same chart when only the social group it refers to changes. ChartBias is a benchmark of 820 real-world charts covering race, income, age, religion, immigration status, and gender, with paired generations in which only the group term is swapped. Across 12 VLMs and 155,484 responses, the authors find three widespread failures: narrative shift, group hallucination, and preference polarity. A multi-agent mitigation separates chart-grounded evidence extraction from group-conditioned generation and adds a counterfactual judge, which substantially reduces narrative shift while keeping answers grounded in the chart.

On the security and privacy of LLMs in Mobility

Mauro Conti, Lorenzo Perinello, Umberto Salviati cross-listed The authors survey how LLMs are used in the mobility sector and assess security, privacy, and reliability against nine technical classes derived from the European AI Act, which classifies transportation AI as high risk. Research mostly studies GPT and Llama models and traffic applications while largely neglecting security. Among 35 reviewed works, only one includes even a partial vulnerability assessment and one a partial risk management system. The authors call for security-by-design in safety-critical transportation systems.

Reliability Theory for AI Control

Grant Molnar The authors apply standard tools from reliability engineering to frontier AI control, using Google DeepMind's defenses against rogue deployment as the case study. They show that the same control stack can suppress rare failures cubically, quadratically, or only linearly, depending on how its failure domains are separated. Birnbaum importance identifies which component improvements add the most reliability, and the analysis shows that prevention layers change the population of cases that recovery layers must handle. The result is concrete guidance on what to separate, improve, measure, and test.

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

Calvin Isley, Johann Gaebler, Max Lamparth, Julia Minson, Sharad Goel Evaluations of social sycophancy in language models penalize validation and positivity, but these same behaviors are characteristic of conversational receptiveness, a social-psychology construct known to improve conversations across disagreement. On a popular moral-advice dataset, the authors find that responses rated more socially sycophantic are also more receptive, and that making human-written responses more receptive without changing their conclusions causes them to be classified as more sycophantic. In a preregistered experiment, participants preferred the more receptive of two substantively equivalent responses, even when they believed the original asker was in the wrong. The authors also present a simple method that increases receptiveness without increasing substantive deference.

From Alignment to Access Control: A Framework for GenAI Policy Enforcement

Nathalie Baracaldo cross-listed "Policy" means different things to different practitioners building generative AI (GenAI) chat apps and agents, which has produced siloed enforcement mechanisms that fall short for security and compliance. The author proposes a methodology to systematically analyze and dissect existing approaches to defining and enforcing policy, spanning model alignment through traditional access control. From this analysis the paper derives recommendations and a call to action for the community. It is a companion to a USENIX Security 2026 Enigma talk.

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

Laizhen Li, Xuan Wang, Peicheng Zhao, Juanjuan Zhao, Kejiang Ye, Cheng-zhong Xu et al. cross-listed A2M (Attraction-to-Manipulation) is a two-stage black-box attack that hijacks agents using the Model Context Protocol (MCP) through attacker-controlled third-party tools. The first stage optimizes tool metadata so agents are more likely to call the malicious tool, and the second uses execution traces to refine the tool's outputs so they steer the agent toward attacker goals. On LiveMCPBench with GLM-4.6, it reaches a 93.6% malicious tool invocation rate, pushes token costs to 32.4 times the benign baseline, and has a 74.4% mean success rate for exfiltration and other attacks. The attacks partly transfer to four other models without re-optimization.
12 more specialized papers

Vision 23

QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation

Jiaqi Zhao, Xiaobin Hu, Bo Yin, Junpeng Jiang, Miao Zhang, Shuicheng Yan cross-listed Existing 2-bit key-value (KV) cache quantization methods score nearly lossless on video benchmarks such as VBench, but in world models and video generators they still cause severe temporal flickering. The authors trace this to the keys: quantizing them adds less reconstruction error than quantizing values, yet it shifts the attention logits and changes which tokens each query attends to. QuantWM is a training-free, strictly causal method with two parts. Quantization-sensitivity-aware clustering (QSAC) chooses key centroids based on how sensitive past queries are to each channel, and principal-subspace attention compensation (PSAC) corrects the remaining key error along the dominant query subspace. Across Causal-Forcing, Matrix-Game-2, Longcat-Video and other models, it improves visual quality and temporal consistency while delivering up to 6.20x KV cache compression.

GTR: Gated Token Recurrence for Efficient Dense Prediction

Zhe Feng, Longfei Liu, Wei Liu, Kai Chen, Jiangjiang Kong, Wei Zhou et al. cross-listed Global softmax attention in vision backbones gets quadratically more expensive as image resolution grows, which makes high-resolution dense prediction costly. GTR (Gated Token Recurrence) is a softmax-free recurrent backbone that combines gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks. It is distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment. GTR-L reaches 58.9 box AP on COCO val2017 at 1.908 ms batch-one latency on an RTX 4090, transfers to segmentation, pose, and depth tasks, and ships with a custom chunkwise CUDA kernel that runs 4.0x faster than FLA at 1.6K tokens.

FleXray: Universal Clinical X-ray Segmentation

Victor Ion Butoi, Vivek Gopalakrishnan, John V. Guttag, Adrian V. Dalca, Neel Dey cross-listed FleXray is a generalist model for segmenting anatomy across the whole body in clinical X-rays, which are hard to label because 2D projection makes structures overlap. Instead of annotating X-rays by hand, the authors built a physics-based generative data engine that simulates fully annotated 2D X-rays from existing 3D CT segmentation datasets, using generative image-editing models to vary their appearance. Trained only on these simulations, the model accurately segments 60 anatomical structures on unseen research datasets and in-the-wild X-rays, and it enables automated measurements for disease grading and guidance during X-ray-guided interventions. The model, code, a dataset, and a browser-based tool are released.
20 more specialized papers

Multimodal 22

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas et al. This competition report covers document Visual Question Answering (VQA) that requires multi-step reasoning. It builds on earlier DocVQA benchmarks and uses documents from eight domains, including business reports, scientific papers, slides, maps, comics, and engineering drawings. Eight teams made 20 valid submissions, ranging from zero-shot vision-language models (VLMs) to OCR-augmented pipelines, agentic retrieval systems, multi-agent ensembles, and fine-tuned models. The main finding is that the strongest systems went beyond single-pass prompting, relying on structured evidence extraction, retrieval, verification, and orchestration across multiple components.

Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

Embedding Team Ovis-Embedding is a family of embedding models that maps text, images, video, and audio into one shared space from a single multimodal backbone rather than separate per-modality encoders. It starts from a pretrained Qwen-omni model adapted with contrastive training and low-rank initialization. Training uses homogeneous-source sampling so each batch has informative negatives, focal loss to emphasize hard examples, and similarity-based distillation from expert models. At inference, low-rank decomposition allows compact embeddings of flexible size. The family reports state-of-the-art results on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB.

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

Zhiping Wu, Dongdong Ren, Yangchengyu Zhou, Zhengjie Zhang, Wenbin Li, Hongbing Pan et al. cross-listed Most post-training quantization (PTQ) methods are designed for text-only LLMs and treat quantization error as uniform in every direction, which works poorly for vision-language models (VLMs) at low bit widths. Riemannian Geometry-Sensitive Quantization (RGSQ) builds a modality-aware metric from Fisher information computed separately for each modality, applies rotations that push quantization error toward directions the loss is insensitive to, and then applies a whitening transformation so standard PTQ methods can run unchanged. At extremely low bit widths (W2A8 and W3A8), it achieves the highest accuracy and stability across mainstream VLM benchmarks, beating VLM-aware baselines such as MBQ and MQuant by up to 5.9%.

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team Qwen3.8-Omni-Flash is a natively multimodal model built for agentic productivity across text, audio, and video, and it inherits the sparse mixture-of-experts (MoE) architecture of Qwen3.8-Next with a one-million-token context window. A native multimodal co-training strategy keeps strong text capabilities while transferring agentic skills from text to audio and video tasks such as video editing, long-form translation, and music-video generation. The team also releases Qwen-MM-Plugins, an open-source plugin framework that adds audio and video support to agent harnesses, and Qwen-Live-Harness, a framework for real-time multimodal agents that handles memory, tool use, and sub-agent delegation. They report strong results across multimodal understanding, reasoning, long-horizon agent tasks, and video productivity benchmarks.

OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities

Yizhou Liu, Jinghang Han, Kaixiang Qiu, Qi He, Minghao Han, Yue Jiang et al. Omni-modal models that handle vision, audio, speech, and text are mostly trained on semantic descriptions, so they learn little about physical attributes, interaction states, or cause and effect. OmniFysics-Nano-V2 is a compact omni-modal model trained with a dual-branch, physics-aware data pipeline. The pipeline links salient objects to structured physical attributes and aligns visual changes with acoustic events and interaction outcomes. Training then uses a two-stage Group Relative Policy Optimization (GRPO) curriculum over prompts selected for reward diversity, moving from general task correctness to fine-grained physical reasoning. The model reports leading results on 17 of 21 benchmarks against state-of-the-art omni-modal models while keeping its general multimodal ability.

Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models

Trung Nguyen Quang, Yuhao Dong, Shuo Sun, Shuai Liu, Shulin Tian, Kim-Hui Yap et al. cross-listed Most data for reinforcement learning with verifiable rewards (RLVR) on video does not require chaining several pieces of visual evidence, so errors that compound across reasoning steps stay hidden. Video-HopChain is a dataset of 22,550 multi-hop questions over 13,378 videos. Each question chains three to six yes/no sub-questions and asks for the sum of integers tied to their answers, so correctness can be checked exactly. Training Qwen3-VL-8B with GRPO on this data raises its mean over eight video benchmarks from 55.4 to 57.9. The authors also add Confidence-Gated Exploration (CGE): when the first four rollouts for a question are all right or all wrong, it resamples the rest with the policy's most confident token masked, which recovers a learning signal from those groups and lifts the mean to 59.3 at the same compute.

CogenPVG: Cognitive-Enhanced Reflective Multi-Agent Framework for Persuasive Video Generation

Yuntian Xiao, Shoulong Zhang, Wenfeng Song, Yan Wang, Yi Chen, Shuai Li cross-listed CogenPVG is a multi-agent framework that generates persuasive videos on general topics from a user-supplied topic and stance. It splits production into four stages, argument reasoning, storyboard planning, asset creation, and post-editing, and pairs a generator agent with a critic agent at each stage for reflective refinement. The design follows the Elaboration Likelihood Model (ELM) of persuasion: critical-thinking-guided argument reasoning serves the central route, and heuristic-guided multimodal assets serve the peripheral route. The authors report that it achieves the best persuasion performance among the compared approaches.

Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

Guanxu Yu, Yuhang Yao cross-listed Many vision applications ask several independent multiple-choice questions about the same image. Visual Jev encodes the image and shared context once, runs the separate question suffixes as a batch, and reads answer probabilities directly from the backbone's language-model head. Answer-supervised post-training raises macro accuracy from 70.6% to 76.1% across four benchmarks. With 32 questions per image, shared batched execution is 8.9x faster than serial execution and 3.4x faster than a batched baseline that recomputes the prefix, at the cost of higher peak memory. Replacing the language-model head with a task-specific classification head gave no consistent accuracy gain.

VideoX-Qwen: Data-Centric Instruction-Based Video Editing

JJiahang Li, Dingbao Shao, Xinyu Chen, Song Wu, Jiang Lin, Duo Li et al. Instruction-based video editing needs large-scale paired supervision and must apply the requested change while preserving everything else, including motion and temporal continuity. VideoX-Qwen pairs a data pipeline that routes specialized generation and understanding models through addition, removal, replacement, and attribute edits, producing over 1.2 million editing records with an 89% automatic acceptance rate, with a unified Qwen-Wan editor that combines multimodal instruction conditioning with dense source-video latent guidance. After progressive image-then-video training, it achieves the best mean score on nine of eleven metrics in a 100-example comparison against UniVideo and Kling O1.

One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data

Xuechen Li Remote-sensing multimodal large language models (MLLMs) work only in English, even though text-only instruction data exists in over 100 languages. MODL (Mutually Orthogonal Domain-Language composition) jointly trains a domain LoRA adapter on English remote-sensing imagery and a language LoRA adapter on text alone, with a loss term that keeps the two updates mutually orthogonal at every layer. Without that constraint, models answer correctly but in English, lose multilingual ability, or diverge, and sixteen alternative methods fail in the same ways. MODL answers correctly in the target language 56–71% of the time, versus at most 27% for the best alternative, and surpasses Qwen2.5-VL-7B on Spanish without any multilingual multimodal data, though non-Latin scripts remain unsolved.

Modality-Gated Deep Adapters: Adding a Modality to a Frozen Embedding Model with Exact Preservation

Abdul Basit Tonmoy, Kazi Fardinul Hoque, Md. Shahrier Islam Arham, Arman Luthra Adding a new modality to a deployed multimodal embedding model with methods like LoRA quietly changes the model's existing outputs, which invalidates stored retrieval indices and earlier benchmark results. Modality-gated deep adapters attach bottleneck adapters to every decoder layer of a frozen embedding LLM and group them into per-modality packs that run only when their own modality is being encoded. Inputs of any other modality therefore produce bit-for-bit identical outputs to the base model, a property the authors prove and check with exact-equality tests. On a frozen 2B base, an audio pack improves audio-to-text R@10 by 3.4 to 5.4 points over a control, and a thermal pack raises thermal-to-text R@10 from 0.224 to 0.785.

TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series

Sheng Pan, Yongli Gu, Yiqing Guo, Warren Jin, Bo Du, Shirui Pan et al. Existing time-series language models (TSLMs) process data offline or alternate between reading input and generating responses, so they cannot take in new observations while they respond. TimeInteract combines a dual-view streaming encoder, a learned mechanism that decides when to respond, and a decoupled inference path that keeps ingesting data during generation. The authors also release StreamTSI-34K, a dataset of 34,588 streaming interaction episodes organized into four capability levels. TimeInteract beats existing LLMs, vision-language models (VLMs), and TSLMs by up to 23.92 points on hard tasks, with near-zero stream stall and up to 2.15x faster inference.

Spoken Language Models that Think Aloud

Junyi Ao, Kainan Peng, Mingbo Ma, Shun Zhang, Zhenyu Tang, Xutai Ma et al. Adding chain-of-thought (CoT) reasoning to Spoken Language Models (SLMs) under a serial think-then-speak design creates long silences that disrupt real-time conversation. The authors propose an asynchronous think-aloud framework for the Thinker-Talker architecture: a main reasoning stream runs alongside a lightweight stream that produces short, task-grounded progress utterances based on the user input and the current reasoning state. A dynamic balancing strategy adds think-aloud speech to fill gaps and cancels pending utterances once the final answer is ready. On spoken reasoning and question-answering benchmarks, the approach substantially reduces user-audible silence while keeping accuracy comparable to the serial baseline.

Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding

Dohyun Kim, Sungjun Han, Hyungguk Kim, Yusik Kim, Jamin Shin, Paul Hongsuck Seo et al. Autoregressive (AR) OCR vision-language models are accurate but decode one token at a time, even though OCR output is tightly grounded in the image and suits parallel generation. GravityOCR is a single parameter-shared model trained both to draft tokens in parallel with block diffusion and to verify them causally in AR mode, so it needs no separate drafter network. The AR path also enables GRPO fine-tuning with OCR-specific rewards, which slightly improves the OmniDocBench score. When served with SGLang, it commits 9.7 tokens per forward pass on average, a 3.94x decode speedup on region crops and a 1.32x end-to-end page speedup.
8 more specialized papers

Reasoning 17

Same Quantity, Different Answer: Numerical Representation Invariance in Language Models

Ephraim Atta-Duncan The same quantity can be written as a decimal, fraction, percentage, number word, scientific notation, or converted unit, and a word problem should get the same answer in every form. The authors build 3,600 exact-rational problems with 8,600 transformed prompts and test five open-weight models. Canonical accuracy is 0.969–0.996, but consistency across representations drops to 0.851–0.981. Much of the apparent collapse under a strict parser comes from the evaluator's own number grammar failing to parse scientific notation, showing how evaluation interfaces can masquerade as reasoning failures. A real semantic failure remains in Mistral Small 4, which scores 0.699 on unit-converted inputs and often answers off by exact powers of ten.

What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content

Marios Papamichalis, Regina Ruane cross-listed Entropy over chain-of-thought tokens is used to decide which tokens receive policy gradient, which get pruned, and whether a training run has collapsed. The authors show that this statistic mixes three separate choices: whether to emit connective scaffolding, which connective to use, and what the substantive content should be. Defining a scaffold vocabulary subset lets them split entropy exactly into these channels, and scaffolding accounts for up to 41% of the raw high-entropy tokens across 23 configurations. An audit also corrects the authors' own compression result: once answers restated inside the chains are stripped out, no token-scoring method beats removing a random contiguous block.

FrontierMath Erd\H{o}s

Tom Adamczewski (Epoch AI), Thomas F. Bloom (University of Manchester) FrontierMath Erdős (FME) is a benchmark of 68 Erdős problems that were still open as of August 2026, selected for mathematical interest and difficulty from 652 open problems on erdosproblems.com. To solve a task, an AI system must prove or disprove the conjecture in the Lean proof assistant, working autonomously under a fixed budget. The benchmark is meant to replace one-off demonstrations of AI solving open problems with a systematic, like-for-like comparison. Five AI systems were evaluated at $300 per problem: GPT-6 Astra scored 3% and all the others scored 0%.

Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Vasily Ilin Lean Pool is a repository of mathematics formalized in the Lean proof assistant. The archive is grown, maintained, and optimized entirely by AI agents. The abstract gives no further details on methods or scale.

Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training

Jacob Beck, Philip V. Ogren, Ari Kobren The authors ask how much of the elaborate machinery behind recent LLM test-time scaling systems, such as evolutionary search harnesses and test-time training, is actually needed on verifiable math and algorithm problems. Their method, Hill Sampling, repeatedly asks a frozen LLM for edits to the best program found so far, keeps an edit only if it scores better, and conditions every new sample on that current best. Using three open-weight models, it sets a new published state of the art on circle packing and beats the AlphaEvolve reference on Erdős' minimum-overlap problem, each in a few hours on eight H100 GPUs. In a large study of evolution strategies applied to LLM weights at test time, actually learning the weights did worse than setting the learning rate to zero, and randomness from token sampling helped exploration more than weight perturbations did.

Direct Optimization of Generators for Search in Automated Theorem Proving

Adam Ousherovitch, Ambuj Tewari In automated theorem proving, fine-tuned LLMs usually guide a tree search rather than generate proofs in one attempt, yet they are trained with plain cross-entropy, which is not aligned with how they are used. The authors extend Compute-Aligned Training (CAT) to tree search through an abstraction of policy-guided search. This yields tractable search-aware losses, plus a search-agnostic uniform-allocation (UA) loss that accounts only for the compute budget; both reweight the per-tactic cross-entropy gradients. On a Lean benchmark, both approaches prove more theorems than cross-entropy across six search strategies, and a single shared UA adapter performs strongly. The gains are larger at 16 search expansions than at 256.

What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation

Kanghui Tian, Siyuan Liu, Tianxiang Jiang, Shuai Dong, Yizhuo Li, Tian Ding et al. In on-policy self-distillation (OPSD), a frozen copy of the model scores the student's own rollouts while seeing privileged context, usually a full reference solution. The authors compare that default with more abstract hints compiled offline (a named strategy, a method-independent framing, or a problem category) and with an answer-only control. On competition mathematics, the best intermediate contexts beat the full solution by 1.4 points at 4B and 1.6 at 8B while storing about ten times fewer hint tokens, and answer-only conditioning stays within 0.2 points of the full solution. The best context depends on student scale and task, and initial teacher-student KL divergence does not predict downstream performance.

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

Minghui Liu, Thomas Magelinski, Dehao Yuan, Qi Yu, Furong Huang Small and mid-sized language models remain brittle reasoners even after knowledge distillation (KD). Ladders-of-Thought automatically generates easier but meaning-preserving rewrites of reasoning problems, groups them into difficulty buckets by step count, and uses a self-evolving bandit scheduler to decide what to train on next. Across 1–8B models from several families, it consistently improves over KD, with gains of +32 percentage points on AddSub and +25 on SVAMP, plus dataset-dependent gains on multi-hop reasoning such as +25 points on StrategyQA. It also converges faster than staged curricula.

CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning

Jian Zhang, Bingyi Wang, Yizhi Liu In multi-step causal reasoning an early error propagates silently, and self-training that rewards only final answers can collapse into shortcut exploitation. CoEvo relies on the fact that checking a single step is easier than generating a whole chain: a deterministic oracle, such as a physics simulator or rule engine, adjudicates the steps where competing reasoning chains disagree, which yields process-level supervision. A single model alternates between acting as Solver and as Proposer, where the Proposer builds progressively harder scenarios within the oracle's constraints. On industrial, clinical, and legal benchmarks, an 8B model trained this way beats distillation baselines and the strongest proprietary reference on path correctness (82.1% vs. 71.4%) and generalizes to unseen categories.

When Verifiers Vote Backwards under Verdict Substitution: Signed Pivotal Value in Correlated Self-Consistency

Yang Shu Swapping one ballot in a self-consistency majority vote can only change decisions that were settled by a single vote. This study asks whether replacing one of seven ballots with a verifier's correctness verdict helps or hurts on those pivotal queries. On MATH-500, a verifier from a different model gains +24.2 percentage points on pivotal queries, while a role-reversed setup loses 11.2 points, and a small code stress test loses 24.5 points. An exact decomposition attributes the sign of the effect to the verifier's accuracy in each specific tie state rather than to its overall accuracy or which model it comes from. Under standard plurality voting over answers, however, the harm mostly disappears.

When Recursive Models Finish Computing

Hare Krishna, Shubham Singh, Stephen Ebert, Hao-Yu Sun Recursive models can keep refining their hidden states past their nominal step budget, so a wrong answer at that budget may mean the computation is unfinished rather than failed. Studying attention- and MLP-based Tiny Recursive Models (TRMs) on 1,000 hard Sudoku puzzles, the authors extend recurrence from 16 to 512 steps and find that exact-solve accuracy rises from 59.2% to 87.5% for the attention model and from 74.4% to 91.9% for the MLP model. Latent-state motion drops sharply once a puzzle is first solved, and completed states are contractive along their trajectory direction even though the local Jacobian still has strongly expanding directions, a pattern the authors call trajectory-conditioned anisotropic stability. Perturbation experiments confirm this directional stability in both architectures.

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

Xiaoyu Luo, Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, Johannes Bjerva Closed-source frontier models hide their raw chain-of-thought (CoT), so claims about their reasoning are hard to check. The authors register a simple custom tool through a standard API feature, which induces models to externalize intermediate reasoning, and validate the approach against native CoT on open-source models. The extracted reasoning matches native reasoning performance and clearly beats no-reasoning baselines on competition math, science, and code. Comparing frontier models including GPT-6 Astra, they find systematic differences in how models compress and organize reasoning, with Astra choosing a correct trajectory early and externalizing only the crucial steps.

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

Ismail Labiad, Matthieu Kowalski, Marc Schoenauer, R\'emi Munos, Julia Kempe Naive repeated sampling at test time tends to produce near-duplicate solution attempts. This work instead steers exploration at the semantic level: it first samples problem-specific concepts, hints, or strategies, then conditions answer generation on them. A small concept generator is trained with reinforcement learning to maximize the success of a larger, frozen answer generator. On hard math problems, it substantially improves pass@k over repeated sampling with the same answer-generation budget, beats concepts from much larger untuned models, and transfers to answer generators it was never trained against, including one from a different model family.
4 more specialized papers

Robotics 17

X-Planner: Event-Structured Task Planning for Embodied Intelligence

Howard Lu, Shalfun Li, Porter Pan, Cris, Lumen, Cyril et al. X-Planner is a planning front-end for long-horizon robot manipulation that makes explicit the task structure Vision-Language-Action (VLA) systems usually leave implicit. It is trained on hierarchical planning data from ego video, UMI, and teleoperation, with takeover-time annotations and human-designed failures so it learns to recognize errors as they happen. A shared vision-language backbone produces either interpretable discrete event states or latent chain-of-thought states that are passed across Transformer depths through Staircase Decoding, with a frozen latent-to-text reconstruction objective keeping the latent states tied to meaning. In offline planning evaluation it ranks second of four models, and the authors report that it outperforms the evaluated baselines in real-robot experiments.

Learning from Humans for Proactive Assistance in Human-Robot Collaborative Transport

Elvin Yang, Christoforos Mavrogiannis cross-listed When a person and a robot carry a large object together, the robot should reduce the person's effort without resisting their guidance, but prior systems tend to be either efficient but stiff or compliant but passive. PROACT combines compliant whole-body control with a transformer that predicts future object motion. The predictor is trained on a large real-world dataset of two people carrying objects together. Across 108 real-world trials with a 9-DoF mobile manipulator, PROACT reduces mean interaction work by 59.2% compared with a compliance-only controller and by 20.4% compared with a model predictive control (MPC) baseline, and it also shortens completion time.

VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models

Jiuyi Xu, Qing Jin, Meida Chen, Song Wang, Yang Sui, Yangming Shi cross-listed VLAQuantBench tests post-training quantization of vision-language-action (VLA) robot models in closed loop, covering 409 runs and 94,574 simulated episodes. The four models are evaluated on LIBERO, and X-VLA is also evaluated on three other simulation benchmark families. Results depend heavily on which layers are quantized, the numerical format and the calibration. Under uncalibrated 4-bit weight and 4-bit activation (W4A4) quantization, expanding the protected action-head layer subset of π0.5 raises task success from 7.0% to 70.5%, and a calibration recipe that fixes one model hurts another. For OpenVLA-OFT, keeping a single small output projection at higher precision restores near-baseline success, which points to model-specific precision choices rather than universal layer-sensitivity rules.

HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery

Jinu Pahk, Jesoon Kang, Taegeon Park, Jisu An, Soo Min Kimm, Jaejoon Kim et al. cross-listed Vision-language-action (VLA) robot policies benefit from geometric supervision, but geometry of the current frame alone does not describe what changes during manipulation. GC-VLA learns to predict multiview tokens that describe how geometry changes between now and about 0.5 seconds ahead, using future frames only to build training targets. It is trained in stages: first a geometry-change vision-language model, then an action expert aligned with it, then joint fine-tuning. A final stage, Geometry-Conditioned Residual Flow (GCRF), adds a router that decides when to intervene and a bounded residual policy learned from closed-loop feedback. GC-VLA reaches 95.20% success on LIBERO, and 99.55% with GCRF.

IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models

Yiqi Wang, Zhifeng Rao, Jiaqi Zhang, Xiaoyang Li, Zhangkai Wu, Yiqun Duan et al. cross-listed Vision-language-action models and world-action models (policies that use learned video or world dynamics) target the same manipulation tasks but are usually evaluated under different protocols. IndustrialVLA-Bench evaluates six released systems under one reporting schema, measuring clean capability on LIBERO, robustness on LIBERO-Plus, sensitivity to instruction paraphrases on LIBERO-Para, and inference latency and memory. Each score averages three seeded runs, and each entry is labeled by how closely it follows the original evaluation protocol. Clean LIBERO scores differ by only 1.58 points across systems, while robustness and paraphrase scores span 14.62 and 31.08 points, and the gaps persist when only the most protocol-faithful systems are compared.

Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models

Yuxin Yang, Gaohan He, Changxue Guan, Hangming Liu cross-listed Autoregressive vision-language-action (VLA) models turn robot actions into discrete tokens, and candidate tokenizers are usually judged by how accurately they reconstruct the actions. The authors compare fixed analytical, data-driven linear, and nonlinear neural action representations through one tokenization interface, using rate-distortion analysis, sequence-modeling diagnostics, and 3,500 LIBERO rollouts. PCA reconstructs actions better than Temporal-DCT, but its token sequences are harder to predict and it gets 3.0 percentage points lower mean task success. An autoencoder with even lower reconstruction error also fails to give the best policy. The conclusion is that tokenizers should be evaluated jointly on reconstruction, sequence predictability, decoder stability, and closed-loop performance.

Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models

Yuhang Zhang, Rangya Zhang, Yujing Shang, Zhuoyuan Yu, Weiying Wang, Steven Yang et al. cross-listed Monocular drone navigation must reach goals in unseen environments using a single forward camera, which gives few depth and scale cues. The authors argue a policy needs a world model's learned representation, not its predictions. Skytopia trains an action-conditioned latent world model with a forward objective that predicts the next observation's representation and an inverse objective that recovers the motion, then discards the predictor at deployment, which removes 59.4% of inference cost. Trained on a 3D Gaussian Splatting platform, one policy handles point-goal, image-goal, and goal-free navigation with 57.8%, 66.0%, and 49.0% success, beating all baselines, and it transfers to a physical drone indoors, outdoors, and in woodland without fine-tuning.

TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models

Xuanyi Liu, Haofeng Wang, Ruiqi Li, Danni Yu, Rui Wan, Ruixu Zhang et al. cross-listed Robot world models that predict head-camera and wrist-camera video are usually evaluated one view at a time, which cannot tell whether the views describe the same action and object state. TriWorldBench evaluates embodied world models on synchronized head, left-wrist, and right-wrist videos, covering 500 episodes across 50 bimanual manipulation tasks with 19 metrics. The metrics cover cross-view consistency, task alignment, physical and 3D coherence, motion quality, temporal consistency, and visual quality. Results are summarized in a single TWB-Score, with per-view breakdowns kept to show where predictions fail.
9 more specialized papers

Reinforcement Learning 11

Correcting Within-Group Self-Selection Bias in Prioritized Replay

Oscar Mir\'o L\'opez-Feliu, Herke van Hoof Prioritized experience replay (PER) replays transitions with high temporal-difference error, but in stochastic environments this skews which outcomes get replayed for the same state-action pair, a bias the authors call within-group self-selection. They split PER into group-level allocation and sibling selection, and propose corrections that keep group-level priority: SAMPLE trains on a uniformly chosen sibling, AVG averages sibling targets, and MODEL samples from an empirical outcome model. Sibling-aware replay improves learning efficiency in environments with rare high-magnitude outcomes, and in MinAtar, SAMPLE mitigates degradation under heavy-tailed rewards in four of five games.

WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning

Xuanlin Jiang, Samuel Hsia, Michael Kuchnik, Zachary DeVito, Minlan Yu, Carole-Jean Wu cross-listed In reinforcement learning (RL) for LLMs, weight transfer (sending updated parameters from trainers to the rollout generators) is becoming a performance bottleneck, and existing solutions handle only some layouts and synchronization modes well. WeightBridge automatically works out how trainer and rollout weight layouts correspond, then plans and executes load-balanced transfers with no redundant copies, all behind a small, general API. Across a range of models, parallelization layouts, and synchronization modes, it reduces average GPU stall time by up to 42× compared with the leading open-source RL framework. A coding agent integrated the library into two different RL frameworks without manual guidance.

Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions

Niloofar Gholipour, Marcos Assuncao, Gursimran Singh, Timothy Yu, Rajkumar Buyya, Julien Gascon-Samson et al. In reinforcement learning for reasoning LLMs, much of the training cost falls on rollout, the stage where trajectories are generated for policy updates. This survey organizes recent work on rollout efficiency into a taxonomy classified by both mechanism and bottleneck. It then analyzes how each family of techniques addresses a different source of inefficiency, where techniques can be combined or conflict, and where the evaluation and reporting of efficiency gains fall short. It closes with open challenges and future research directions.

Fully Byzantine-Resilient Multi-Agent Reinforcement Learning

Haejoon Lee, Dimitra Panagou In distributed actor-critic multi-agent reinforcement learning (AC-MARL), existing Byzantine-resilient methods only guarantee convergence to a neighborhood of the attack-free solution. FRAC-MARL has each agent use redundancy in two-hop messages to decide which messages to trust. Under linear parameterizations and attacks confined to communication edges, the authors prove almost-sure convergence to the same limit points as in the attack-free case over time-varying graphs. They also give a topological condition for convergence that can be checked in polynomial time, and they demonstrate the method on multi-robot formation control.

Informed Masking: Structure-Aware Perturbation for Reinforcement Learning in Diffusion Large Language Models

Xiaoyi Yu, Enver Sangineto, Pei Fu, Fiorenzo Parascandolo, Wenhui Tan, Ruikang Zhang et al. Reinforcement learning (RL) for diffusion large language models (dLLMs) estimates likelihoods from a few randomly masked reconstruction subproblems per rollout. The authors find that some tokens are "upstream", meaning that revealing them sharply shifts confidence at nearby undecoded positions, while others are "downstream", and that masking downstream tokens yields better-posed subproblems. Informed Masking (IM) scores each token's priority from the denoising trajectory at no extra inference cost and biases masking toward downstream tokens. Added to three state-of-the-art dLLM RL methods on LLaDA-8B-Instruct, it gives relative average gains of up to 8.68% on math and planning benchmarks, with more stable training.

PACT: From Credit Assignment to Critic Alignment

Jiayan Fu, Hang Xu, Yong Zhang, Zhaokai Luo, Yao Hu, Dongyan Zhao et al. Reinforcement learning post-training for large language models lacks a precise definition of token-level credit. The authors prove that three conditions (Completeness, Prefix Consistency, and Neutrality) uniquely determine it. They use this definition to explain On-Policy Distillation (OPD), REINFORCE Leave-One-Out (RLOO), and critic errors in Generalized Advantage Estimation (GAE). They then propose Policy Aligned Critic Training (PACT), which updates the actor before the critic and applies importance-sampling correction to critic training. PACT reaches 72.87% average accuracy on agentic math reasoning, beating GRPO by 8.80 points and PPO by 13.16, and scores 67.4% on SWE-bench Verified, 2.0 to 3.8 points above PPO, GRPO, and SAO.
5 more specialized papers