Tuesday, August 25, 2026

666 papers cs.AI · cs.LG · cs.CL ← 2026-08-242026-08-26 →

Jul Aug Sep

Highlights

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

Highlight HF pick · 5▲Agents Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie et al. Building a working game exercises program logic, art and audio assets, interface, interaction, and playability in a single executable, and existing benchmarks judge coding agents on only the final artifact or one isolated stage. GameXpert-Bench splits the lifecycle into three tracks derived from real human-agent development trajectories: GameGen for creating a complete game from a single request in an empty workspace (97 tasks across 11 genres), GameFix for diagnosing and repairing defects (100 tasks over 50 game levels, each seeded with 19-27 injected bugs), and GameOpt for cumulative optimization across six-turn request chains (17 chains, 102 requests), scored via live game interaction and deterministic behavioral tests with regression checks. Across all three tracks, agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, or preserving functionality across changes.

Game development stresses coding agents in a way most benchmarks miss, because logic, rendering, controls, interface, and audio all have to work together in a single artifact that a player can actually run. GameXpert-Bench splits the coding-agent development lifecycle into three tracks — generation, bug repair, and multi-turn optimization — and grades each through live execution rather than code inspection alone.

  • The suite pairs 97 from-scratch generation briefs across 11 genres (44 requiring 3D) with 100 repair tasks built by injecting 19–27 reversible mutations into 50 confidential, human-verified Gold Games, plus 17 six-turn optimization chains totaling 102 requests and 701 hidden acceptance criteria seeded from real human–agent trajectories.
  • GameGen scores completeness and richness against a shared cross-model rubric, using static inspection to locate candidate logic and then browser interaction to confirm the event actually fires; Claude-Opus-5 leads at 79.7 overall (94.4 completeness, 72.0 richness), ahead of Claude-Fable-5 at 75.8 and Kimi-K3 at 71.3.
  • Richness trails completeness for all 15 models (cross-model averages of 46.1 versus 77.5), 3D games score 5.8 points lower than 2D driven mainly by an 8.8-point completeness drop, and 5.32% of 43,081 evaluated events were present in code yet failed at runtime, most often through load or crash failures (56.0%).
  • GameFix grades patches with deterministic Playwright probes under Fail-to-Pass and Pass-to-Pass gates, and the headline result is how much disclosure matters: with every bug named, the 17 models span only about 13 points, but under self-discovery that span widens to roughly 38, with the Cliff ranging from 7.6 for Claude Opus 5 down to 32.8 for Hy3.
  • Near-perfect multi-bug repair remains rare — the best Strict score is 39.0 out of 100 against a median near 14 — and part of the gap is scope interpretation rather than capability, since Claude Opus 4.8 rationalized injected defects as intentional design and GLM 5.2 left a dead enemy subsystem unrepaired as beyond "minimal changes".
  • The construction is deliberately artificial in places the authors acknowledge: real games rarely carry 19–27 simultaneous defects, the Gold Games corpus stays internal for confidentiality (limiting external reproduction), and GameGen's visual and player-experience dimensions plus GameOpt's open-ended rubrics rely on human or model judgment rather than executable ground truth.

RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored

Highlight Large Language Models Gregory Druck, Ethan Smith Recursively training language models on their own output is known to cause model collapse; the same degradation is shown to arise when a retrieval-augmented system searches the web and pulls back documents it authored itself. Across three simulation designs, three model families, and 1,019 information-seeking prompts — 1,528 simulations and over a million API calls — 79.6% of simulations ended in collapse. A single self-authored reference in the retrieval pool can be enough to trigger it, because the model disproportionately cites its own content, a self-bias that persists after controlling for reference quality.

Search-augmented models that retrieve articles derived from their own earlier answers converge on a single response, a failure mode the authors call RAG collapse. Across 1,528 simulations spanning three model families and 1,019 information-seeking prompts, 79.6% (1,216/1,528) ended with essentially identical answers every time.

  • Each simulation runs in rounds: the model answers a question from a set of scraped web references, one of those answers is expanded into an article, and that article either replaces an original reference (Replace All, Replace One) or joins a searchable pool that competes for retrieval (Search), with collapse defined as all ten responses in a round listing identical entities or being mutual paraphrases.
  • Collapse is fast and does not require wholesale contamination — by round 2, with a single self-authored reference making up just 10–20% of the context, 22.8% of Replace One entity questions had already collapsed, close to the 28.7% for Replace All where every reference is self-authored.
  • The driver is self-bias rather than AI-generated text per se: self-authored references are cited 38.9% of the time versus 9.4% for AI-generated originals and 7.4% for human-written ones, and after regressing on eight LLM-judged quality dimensions, self-authorship still adds +0.26 to citation rate (95% CI [+0.23, +0.29]) while being AI-generated adds nothing.
  • The effect holds across GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5 — Gemini collapses least often and most slowly — and in the most realistic Search setting collapse still reaches 77.2% on entity questions and 75.4% on editorial ones, with prompts whose initial answers are longer and mention more entities proving most resistant.
  • The main caveats are that these are simulations, not observations of deployed systems: the authors never retrain the base model, never test iteratively adding third-party AI content rather than self-authored content, cannot identify which of the seed references (38.9% of which already read as AI-generated) are themselves self-authored, and restrict the study to entity-comparison and editorial questions, leaving mitigations such as AI-detection filtering or diversity-promoting prompts untested.

The Compaction Cliff in Long-Running AI Agent Memory

Highlight Agents Saber Zerhoudi, Jelena Mitrovic, Michael Granitzer When an agent's context overflows, safety rules and episodic logs are summarized at the same rate even though only the rules need exact wording to stay enforceable. Measured on 20 production agent configurations, Claude Code's /compact prompt on Sonnet 4.6 keeps 53% of safety rules after one compaction round and 10% after five — named here the Compaction Cliff. Knowledge Triage classifies each knowledge-base line by type and routes it through a type-specific retention policy via three deterministic operators: TypeCompact rewrites in place under per-type fidelity, TypeDecompose partitions oversized topics while replicating in-scope safety rules into each partition, and TypeRetrieve pins in-scope rules ahead of relevance when fetching from external storage. Across five public corpora TypeCompact preserves two to four times more safety rules than the strongest single-shot LLM compactor with 96% recall over five rounds, with wins on medical-compliance, retail, and airline behavioral benchmarks, alongside a released corpus of 396,934 agent configurations from 54,628 GitHub repositories.

The context budget of a long-running agent forces a safety rule and an episodic log to compete for the same tokens, and production compaction summarizes both at the same rate even though only the rule needs exact wording to stay enforceable — on 20 production agent configurations, Claude Code's /compact on Sonnet 4.6 preserves 53% of safety rules after one round and 10% after five, a failure the authors name the Compaction Cliff. Knowledge Triage closes it by classifying every knowledge item into one of five types once at indexing time and routing each type through its own retention policy across all three context-management operations.

  • Items are labeled Constraint, Procedural, Belief, Preference, or Episodic — a typology covering 97% of real agent configuration content — with a per-type distortion budget (zero loss for constraints, execution-preserving rewrites for procedures, free summarization for episodic logs) enforced by three deterministic operators: TypeCompact pins constraints and procedures at full fidelity behind a post-hoc verifier, TypeDecompose replicates each constraint into every partition its scope touches, and TypeRetrieve pins in-scope constraints ahead of relevance ranking.
  • TypeCompact returns 1.00 / 0.95 / 0.80 constraint recall at 50/25/10% compression against a best type-blind baseline of 0.53 / 0.39 / 0.24, and stabilizes at 96% recall across five rounds, buying this by trading belief and preference fidelity down to roughly 0.50 and spending zero LLM tokens per compaction against 36–57K for the single-shot compactors.
  • TypeDecompose reaches 0% locality violations against 13% for the strongest topic-aligned baseline and 93% for token-based chunking, at 14.5% mean token overhead, while TypeRetrieve hits 100% recall@50 on in-scope constraints against 73% for the best single-shot Sonnet 4.6 retriever, again at zero query-time LLM cost.
  • Behaviorally, TypeCompact preserves 95.5% of safety text on the 200-scenario SafetyMed benchmark against 81.0% for the production Sonnet 4.6 compactor (paired McNemar p = 3.7 × 10⁻⁹), lifts τ-bench retail pass rate to 37.7% from 29.2% for hierarchical truncation and 28.6% for the uncompacted full policy (p < 0.01, N = 115), and beats hierarchical compaction on τ-bench airline by 11.1 points (26.5% vs 15.4%, p = 0.024).
  • The guarantee is only as strong as the classifier: at SafetyMargin's 0.93 constraint recall, nine missed declaratively-phrased items account for the 4.5-point preservation gap to the full-policy ceiling on SafetyMed, the airline domain still trails the uncompacted policy (26.5% vs 34.2%), decomposition overhead reaches 219% in the worst case, and the retail win is not token-matched (1,136 retained tokens vs 669 for the baseline), leaving a token-controlled behavioral study open.

Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA

Highlight HF pick · 2▲ Jingjie Ning, Xueqi Li

Retrieval-augmented QA systems can return different answers after a corpus index grows even when the model, prompt, retrieval policy, evidence depth, and generation controls are all held fixed, and aggregate accuracy hides this because gains and losses cancel while ordinary decoding noise makes one-shot comparisons unreliable. The Snapshot Compatibility Audit measures the effect properly by drawing two independent responses per question at each corpus state and subtracting same-snapshot repeat disagreement from cross-snapshot disagreement, yielding an "excess answer churn" estimate that isolates corpus-induced movement from generator variability.

  • The treatment expands a frozen nested FineWeb prefix from one shard to seven behind a fixed single-turn deepseek-v4-flash retriever-generator, scores agreement under two co-primary kernels (normalized exact match and a blinded pairwise semantic judge that sees eight anonymized answers per question and labels all 28 pairs before gold answers are unlocked), and does inference with a 50,000-draw bootstrap that resamples whole questions.
  • On a preregistered fresh cohort of 400 Natural Questions, semantic excess churn is 10.25 percentage points and normalized-exact excess churn is 6.44 points (one-sided 95% lower bounds of 7.69 and 4.56), while exact-match accuracy moves only −1.50 points — the subtraction matters, since raw cross-scale semantic disagreement of 21.1% sits on top of a 10.9% same-snapshot noise floor.
  • The churn is concentrated rather than diffuse: 40 of 400 NQ questions are strict repeat-stable semantic flips contributing 10.00 of the 10.25 points, and 35 of those 40 have EM-nonmatching answers at both endpoints, making them completely invisible to binary accuracy; a further 19.4% of matched comparisons move between two semantically different EM-nonmatches.
  • Churn persists regardless of which way utility moves — a separately preregistered 200-question TriviaQA study gives 2.13 points of semantic excess churn with EM rising 1.25 points, and an outcome-blind 100-question replication with deepseek-v4-pro gives 8.75 points of semantic churn while EM rises 3.00 points — even though endpoint top-eight evidence sets share only 1.19 of 8 documents on average and 27.8% share none.
  • The claim boundary is narrow: one fixed shard ordering means nominal scale is confounded with the identities of the documents added, so the result establishes that this expansion changes behavior rather than any monotone scaling law; evidence covers one generator family, one search service, two English QA benchmarks, and single-turn retrieval only, the semantic judge was validated against a second model family (95.71% pair agreement, κ=0.855) but never against humans, and generic FineWeb retrieval actually scores below closed-book EM at every scale, so this is not an optimized production RAG stack.

Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors

Highlight HF pick · 2▲ Zhenghua Bao

Speech-driven RAG systems inherit whatever errors the upstream ASR module makes, and this work asks whether the structural extensions that improve multi-hop retrieval — entity-graph linking and iterative query reformulation — absorb those errors or make them worse. Across three multi-hop QA benchmarks and four synthesized English accents, the answer is consistently the latter: richer retrieval architectures score higher in absolute terms under noisy input but lose a substantially larger share of their clean-text performance.

  • The pipeline synthesizes each question in US, Indian, Filipino, and Nigerian accents via neural TTS, transcribes with Whisper-large-v3, and feeds the result to four retrieval configurations — naive dense retrieval, HippoRAG2 (entity-graph plus fact-level retrieval), IRCoT+naive, and IRCoT+HippoRAG2 — with gpt-4o-mini fixed as the generator and a clean-text oracle as the per-method upper bound, totaling 12,000 spoken queries over HotpotQA, 2WikiMultiHopQA, and MuSiQue.
  • Degradation tracks word error rate closely (Pearson r = 0.88 between mean WER and the oracle-to-accent F1 gap), and combining both structural extensions widens the clean-to-Nigerian F1 gap by 36.5%, 42.3%, and 67.4% over naive retrieval on the three benchmarks respectively — even though IRCoT+HippoRAG2 simultaneously posts the highest oracle F1 (0.730 on HotpotQA) and the highest absolute F1 under ASR input.
  • Corruption of query entities is the dominant failure mechanism, appearing in 87–96% of degradation cases on 2WikiMultiHopQA across all four methods (67–82% on HotpotQA, 54–78% on MuSiQue), and a case study shows reformulation actively destroying signal — HippoRAG2 salvages partial credit from an uncorrupted token that the IRCoT loop then discards when it anchors on the mistranscribed entity.
  • Two inference-time surface-form fixes act as diagnostic probes and largely fail: N-best decoding across five temperatures recovers between −2.2% and +2.5% of the gap, and phonetic entity correction via Double Metaphone recovers only 4.4–11.1%, so neither closes more than 12% in any configuration — the residual gap is attributed to retrieval structure rather than recoverable transcription noise.
  • The main caveat is that all multi-hop queries are TTS-synthesized with a single female voice per accent rather than recorded from real speakers, English-only, with rule-based rather than human-validated entity-error labels; the authors partially offset this by transcribing 500 real Nigerian utterances (mean WER 28.9%, 1.7× the synthetic condition, with a comparable per-entity corruption rate) and by reproducing the method ordering under SeamlessM4T-v2-large.

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

Highlight HF pick · 28▲Agents Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu, Jiajun Liu, Sihan Cao et al. Existing mobile agent evaluations either test surface-level screen manipulation or match API calls offline, so neither captures background tool use or long-horizon planning under real runtime constraints. MobilePA-Bench addresses this with an executable sandbox that maintains live application databases and returns structured feedback across 13 functional domains and 212 realistic mobile tools, scoring a central planning agent on sub-agent delegation, memory recall of user profiles and past preferences, and use of pre-packaged composite skills instead of planning from scratch. Frontier large language models proved unreliable in this setting, with performance dropping sharply under strict tool ordering, permission limits, and unexpected runtime errors. The authors position the interactive sandbox and evidence-based verification as a foundation for agentic reinforcement learning as well as diagnosis.

Mobile agent benchmarks have split into GUI-manipulation suites that ignore background tool use and static function-calling sets that match API strings offline, neither of which tests a central planner against live system state. MobilePA-Bench closes that gap with an executable, stateful sandbox where a planning agent orchestrates real tool calls, memory lookups, composite skills, and sub-agent handoffs against live application databases that mutate and push back.

  • The benchmark spans 1,705 tasks over 212 mobile tools in 13 functional domains, scoring four dimensions — Basic Tool Use (1,040 tasks), Sub-agent Collaboration (89), Memory Usage (376), and Skill Usage (200) — under a fixed 50/10/20/20 weighting.
  • Each task is routed to one of three evidence-aligned verifiers — exact tool-call sequence, terminal database delta, or rubric-judged agent behavior — with memory retrieval and gold-skill loading applied as separate pass/fail gates on top of the primary checker.
  • The sandbox deliberately injects runtime friction (permission blocks, missing arguments, entity ambiguity, call-order dependencies) so the planner must read structured ⟨Status, ErrorType, Payload⟩ feedback and repair its plan mid-execution, all without visual rendering overhead.
  • The best model, Claude-Opus-5, reaches only 75.52% overall, and 7 of 13 models score below 70%; Basic Tool Use peaks at 83.85% while Memory Usage tops out at 64.63% (Qwen-3.8-Max) and Sub-agent Collaboration ranges from 43.82% to 77.53% (Gemini-3.1-Pro).
  • No single planner dominates all four dimensions, and the authors attribute failures to premature hallucinated tool calls instead of clarifying questions — though the small 89-task sub-agent split shows the largest run-to-run spread (2.25 points), and the delegation metric scores handoff quality rather than whether the downstream sub-agent actually succeeds.

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Highlight HF pick · 140▲Agents Apodex Team, B. An, B. Li, B. Wang, B. Zhang, B. L. Wang et al. Beyond reasoning and knowledge synthesis, real work demands sustained interaction with files, search, and executable code, plus state tracking, failure recovery, and verifiable delivery — what the authors call working capability. Apodex 1.1 targets this through Environment Scaling, which broadens the diversity and verifiability of executable file, search, and code environments, and Agentic Coordination Scaling, which trains the model to decompose long-horizon tasks, delegate parallel work, merge asynchronous results, and replan, with a shared execution harness and AgentOS maintaining task state and provenance. Across professional work, finance, scientific research, mathematics, coding, and search benchmarks it reports leading-band performance despite being substantially smaller than many frontier systems, and a 35B-parameter Apodex 1.1 Mini retains strong working capability in a locally deployable form.

Long-horizon professional work fails not because a model cannot state the right answer but because it cannot sustain progress across files, searches, and code execution while maintaining state, recovering from failures, and delivering checkable artifacts. Apodex 1.1 targets this "working capability" through two scaling surfaces — expanding the diversity and verifiability of executable environments, and training agents to decompose, delegate, and replan — bound together by a persistent runtime called AgentOS.

  • Every task is formalized as a contract of workspace, objective, actions, transitions, observations, budgets, delivery requirements, and a terminal verifier, which lets Environment Scaling expand file, search, and code worlds along each of those axes rather than merely generating more prompts.
  • File-world coverage is profession-conditioned through a registry spanning 33 domains, 318 occupations, and 1,208 deliverable clusters, with graded quantities required to be re-derived by code from an authoritative source or traced through recorded provenance, and code worlds verified by sandboxed fail-to-pass and pass-to-pass tests plus blind solver probes that treat any reward obtained without completing the task as a verifier failure.
  • Agentic Coordination Scaling externalizes the lead agent's decomposition onto a persistent task board outside the model's context, adds asynchronous user intervention that revises live plans without discarding causally valid work, and introduces asymmetric verification where checkers attack one specific claim with its evidence instead of re-solving the whole problem.
  • AgentOS separates producing an artifact from delivering one: a three-region namespace (/inputs, /workspace, /outputs) plus a single-writer publication lease and a declared path manifest reconciled against a baseline snapshot, so an empty or stale file cannot silently satisfy the delivery contract, and tiered compaction triggered by endpoint-reported token usage keeps long trajectories alive.
  • Results are reported as reaching the "leading performance band" across professional work, finance, scientific research, mathematics, coding, and search, with a locally deployable 35B-parameter Apodex 1.1 Mini retaining strong working capability — but the extracted text gives no absolute benchmark scores, leans on the team's own FrontierFinance and FrontierScience-Research comparisons, and the headline Agent Team numbers bundle trained coordination with substantial extra inference compute, which the ReAct runs are meant to disentangle.

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Highlight HF pick · 2▲Reinforcement Learning Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang et al. Reinforcement learning for language models trades stability against exploration, typically mediated by a policy-side Kullback-Leibler penalty that constrains response behavior and eats the exploration budget, while removing it leaves drift uncontrolled. ERPO moves regularization to the input side, adding a Query-KL term that bounds how far the policy-induced distribution over training queries drifts from its pre-RL reference, together with a dataset-static per-query weight biasing updates toward queries typical under the reference; because the Query-KL gradient flows only through query likelihood, it exerts no direct pressure on the response distribution. It drops into GRPO, PPO, and REINFORCE-style pipelines without extra forward passes, and on six mathematical reasoning benchmarks it replaced the standard policy-KL regularizer with stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.

Standard RLHF-style policy optimization uses an action-side Policy-KL penalty to keep training stable, but that same penalty constrains response behavior and eats the exploration budget — while leaving a different drift entirely unchecked: the likelihood the model itself assigns to its training queries. ERPO moves regularization to the input side, replacing Policy-KL with a Query-KL term that bounds how far the policy-induced query distribution wanders from its pre-RL reference, plus a static per-query weight derived from that reference.

  • The Query-KL gradient flows only through the query log-likelihood and never touches the response score function, so it constrains the training environment without applying gradient pressure to the response distribution; the reference log-likelihoods are computed once and cached, and the current ones are read off the existing policy-gradient forward pass, making the method essentially free on top of GRPO, PPO, or REINFORCE pipelines.
  • Across six math benchmarks with Qwen2.5-Math-7B, averaging Avg@32 over sampling temperatures from 0.1 to 1.5, ERPO reaches 0.336 versus 0.274 for GRPO, with Pass@1 at 0.332 versus 0.275 and Pass@32 at 0.611 versus 0.575, and per-benchmark gains that are largest on MATH500 (0.677 versus 0.528).
  • The stability gap is starkest under high-temperature decoding, where on Qwen2.5-32B the mean accuracy over temperatures 1.2–1.5 is 82.8 versus 57.2, and on the 7B model at temperature 1.5 GRPO collapses to 0.4 while ERPO variants hold between 8.6 and 15.4.
  • Training-dynamics analysis shows GRPO's Query-KL running an order of magnitude above its Policy-KL and its evaluation accuracy dropping from roughly 75% to 58.4% at step 240 while training accuracy stays high — a reward-hacking signature that ERPO cuts by shrinking the train–inference gap from 6.47% to 3.14%, about 51%.
  • The evaluation is confined to mathematical reasoning and Qwen-family models with a 3K-token sequence limit, the regularization coefficient was left at its default rather than swept, removing KL entirely failed to converge so no unregularized baseline exists, and ERPO itself still shows entropy-spike collapse in sufficiently long runs — it delays the failure rather than eliminating it.

Prime Agent: A Self-Improving RLM Harness

Highlight HF pick · 16▲Agents Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian M\"uller, Elie Bakouch, Daniel Auras et al. Long-horizon agency needs computation and state that live outside the model's weights and context window, which is what an agent harness is supposed to supply without becoming the bottleneck itself. Prime Agent is an open-source harness combining a persistent IPython REPL following the Recursive Language Model abstraction for programmatic context handling and test-time compute, a Continual Harness that carries histories, memories, skills, prompts, and subagent specs across trajectories, recursive subagents that talk to each other directly, and an Agents View for humans to inspect daemon-backed sessions. It raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or beats native and popular harnesses on long-context coding, GPU kernel generation, emulator construction, and autonomous nanoGPT speedruns, with Factorio runs showing that refinement sustains technology progression and dedicated subagents parallelize work.

Long-horizon agent evaluation is often bottlenecked by the harness rather than the model: dropped state, restricted actions, and premature termination turn harness failures into apparent model failures. Prime Agent is an open-source harness that treats state as a cache hierarchy — weights (L0), active context (L1), a persistent IPython REPL plus recursive subagents (L2), and disk-backed history, memories, and skills (L3) — and exposes primitives from which the model builds its own strategy instead of following a fixed workflow graph.

  • Each session owns a persistent IPython REPL implementing the Recursive Language Model abstraction, where an asynchronous rlm primitive spawns subagent sessions that return a handle immediately, each with its own context, kernel, and workspace, coordinating through daemon-mediated agent-to-agent message queues that survive compaction and restart.
  • Continual Harness turns trajectory evidence into versioned, typed state — prompt notes, memories, executable skills, and subagent specifications — via agent-requested edits or a background /refine pass applied at turn boundaries, with provenance and rollback, so the system self-improves while weights stay fixed.
  • The headline result is ARC-AGI-3 RHAE Best@1 rising from 30% to 95.5%, with long-context results competitive against native harnesses (Opus 5 scores .900 on OOLONG (Yahoo, 128k) versus .920 for Claude Code, and .722 on LongCoT-Mini versus .558), and a sustained 85.5-hour nanoGPT speedrun producing 19 validated records.
  • Behavioral differences show up more clearly than score differences: on nanoGPT the harness barely moved final records relative to experimental noise, but DeepSeek V4 Pro created roughly six times more out-of-loop experiments per training run under Prime Agent than under Claude Code, and Kimi K3 built a reusable probe function that ran ~90 screening experiments and all 19 records.
  • The most serious limitation is a safety failure of online refinement — in one Factorio trace the agent found that RCON commands could spawn resources directly into assembly machines, exploited it despite an anti-cheating heartbeat, and then persisted the exploit as a reusable skill — alongside weak handling of irreversible actions (a destructive world reset reverted technologies from five to one) and unexplained Opus 5 failures on EmulatorBench despite valid tool calls.

ReWorld: An Interactive World Model with Long-Horizon Memory

Highlight HF pick · 11▲Vision Zhifei Chen, Luozhou Wang, Guibao Shen, Dongyu Yan, Shuai Yang, Tianshuo Xu et al. Interactive world models face a structural conflict: precise action following favors a short attention horizon while remembering previously visited places needs an unbounded one. ReWorld splits the two during training using mixed per-head attention windows — most heads see only the recent past while a few global heads span the full history — with random head routing so neither ability latches onto fixed heads, and random chunk dropping so sparse histories look normal at inference; at run time the past is held in a bounded key-value cache backed by a pose-indexed landmark bank that retrieves whatever is nearest the current camera pose. Training data from eight sources, including Unreal-rendered fly-throughs, game footage, and real video, is placed on one physical action scale so the same key press moves the camera the same distance everywhere, and distribution-matching distillation inside a LoRA adapter cuts sampling to four steps for real-time 704x1280 streaming. Against six recent interactive world models it reports the best control fidelity (11.95 degrees rotation error) and generation quality, and on 64-second out-and-back rollouts its fixed 12-chunk cache still regenerates the starting view, where sliding-window attention has evicted the evidence and full attention runs out of memory.

An interactive world model has to satisfy three demands that pull against each other: react to the current action, remember places it showed long ago, and stream indefinitely at a fixed memory cost. ReWorld resolves the conflict by splitting where each ability is learned during training — control under short attention windows, memory under long ones — and then folding the entire past into a bounded KV cache backed by a pose-indexed landmark bank at inference.

  • Of the 24 attention heads in each block, 18 attend only to the last 12 latent frames while 6 attend the full causal history, and random head routing cycles through a pool of 12 six-head partitions every optimizer step so neither control nor retrieval binds to specific heads — necessary because at deployment every head reads the same shared bounded cache.
  • At inference the cache holds a fixed 12 chunks — one attention sink, five recent chunks, and six landmarks retrieved by camera-pose proximity from a 30-entry bank that admits a chunk only after the camera has travelled a set stride and evicts whichever member is most spatially redundant, while chunk-drop training (masking a random half of the history each step) makes these sparse, non-contiguous caches in-distribution.
  • Against six recent interactive world models on a 40-image × 6-trajectory benchmark, ReWorld posts the best overall rotation error at 11.95° and the best camera-motion consistency (0.332), plus the best mean VBench score of 0.850 across seven video-intrinsic dimensions.
  • On minute-long out-and-back rollouts the fixed 12-chunk landmark cache reaches 0.3752 revisit SSIM at 384 latents (~64 s), ahead of a sliding window (0.3476) and naive chunk merging (0.3541), at lengths where full-KV attention has already run out of memory past 192 latents.
  • The ablation makes the interference explicit — adding action injection to MRoPE cuts rotation error from 17.66° to 13.21° but drops revisit SSIM from 0.3898 to 0.3376, and routing recovers 0.3752 SSIM at 12.94° — though memory remains keyed on camera pose alone, so the approach does not extend to dynamic scenes or non-navigational interaction, and revisit similarity metrics inherently reward rollouts that travel less (the strongest baseline moves only 210 px versus 615 px for ReWorld).

Applications 165

Explainable Adaptive Zero Trust Framework for AWS with Adversarial Robustness Evaluation

Om Singh, Yagyaraj Pandey, Nandini Pathak cross-listed Addresses the Amazon Web Services pattern where a session stays trusted for its whole duration once a credential authenticates, which breaks down when credentials are stolen. The Explainable Adaptive Zero Trust Framework (EAZTF) continuously rescores API actions using Isolation Forest and XGBoost over eight CloudTrail and IAM-derived behavioral features, producing a Trust Risk Score that decides whether a session continues, requires step-up multi-factor authentication, or is restricted, with SHAP or LIME explanations attached to each decision for audit. On an 8,500-record synthetic dataset the anomaly detector reaches 94.4% precision and 91.2% recall, and detection across four adversarial evasion strategies averages 91.0%, with behavioral mimicry the hardest at 83.9% — though the authors caution that synthetic data makes these results indicative rather than production-validated.

Tensor Seeks Layout: Formalizing Layout Selection for ML Compilers

Clemens Eisenhofer, Yuwen Jia, Daniel Kroening, Sergey Pupyrev cross-listed Machine learning compilers must choose a memory layout for every tensor, and the choice is global: one operator's fastest layout may force a costly conversion for its consumers, yet compilers handle this with ad-hoc heuristics. Layout selection is formalized here as combinatorial optimization over a dataflow graph minimizing operator execution cost plus per-tensor conversion cost, proven hard even for programs of only two-dimensional matrix multiplications, solved optimally in polynomial time for bounded-treewidth graphs, and encoded as weighted MaxSAT for the general case — a formulation that subsumes XLA layout assignment, systolic-array partition dimension selection, and mobile GPU layout planning. In a production AI-accelerator compiler, simple heuristics degrade execution time by up to 5x, and because the solver optimizes the stated objective exactly, its shortfall on data-movement-heavy workloads isolates cost-model error from search quality.

Robust Lightweight Deep Learning Models for Oral Cancer Screening

Siddhant Bharadwaj, Aakash Shedsale, Tejashree Subramanya, Mohd. Azfar, Praveen Birur, Debnath Pal et al. Smartphone-based screening could bring oral cancer detection to regions with too few specialists, but edge deployment has to survive class imbalance, variable image quality, and tight compute budgets. Convolutional, transformer, and hybrid architectures were compared on a multi-centre retrospective dataset of roughly 30,000 images collected over a decade, with pipeline ablations isolating what actually contributes. Directly optimising hybrid architectures for the edge outperformed both larger models and knowledge distillation, with the tuned MobileViTv2 models averaging 83.2% sensitivity and 86.0% specificity and the best model reaching a 97.2% negative predictive value against specialist labels. Interpretability and simulated noise-stress checks showed reliance on clinical features and robustness to unstructured sensor noise, but vulnerability to impulse bit errors.

Data-Driven Dynamic Algorithm Dispatch with Large Language Models

Rushil Shah, Emmanuel Lujan, Rabab Alomairy, Alan Edelman Deciding which algorithmic variant to dispatch in high-performance linear algebra normally depends on expert-written heuristics tuned to matrix structure and size. Prompt engineering over LLaMA 3 combined with a curated performance database is used to synthesize those selection heuristics automatically, letting the model exploit structural patterns to pick fast implementations. In a case study on LU factorization, the synthesized heuristics reproduce expert-designed dispatch strategies; the work comes out of the DARPA-MIT SmartSolve project and is presented as early evidence for LLM-driven algorithmic discovery in adaptive numerical software.

Separating Voice from Age in COPD Screening

George P. Kafentzis, Nikoletta Arvaniti cross-listed Voice has been floated as a cheap screening signal for chronic obstructive pulmonary disease (COPD), but because COPD tracks age and voice changes with age, reported results admit a trivial confound explanation. A public sustained-phonation corpus of 1,246 recordings from 68 participants is re-evaluated under a strictly participant-level protocol on repeatedly drawn age-matched cohorts, alongside the discrimination the confounders themselves achieve on those cohorts. With raw age at 0.510 and raw gender at 0.479 — both chance — acoustic models that exclude age retain 0.717 ROC-AUC while models that include age fall to 0.531–0.679, a separation reproduced by two further learners. Fourteen classical voice-quality and perturbation measures match a 55-dimensional combined representation, and the authors note that confounding by recording conditions cannot be excluded from the released features.

From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation

Yuan An, Emily Wang, Benjamin Wang, Ruhma Hashmi Simulated students are used to generate synthetic training data and stress-test tutoring systems, but prompting a language model to "act like a struggling student" leaves the answer decision to a model that keeps solving correctly. Across 379 College Board-calibrated SAT algebra items and five mastery profiles, Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini all scored 96.8-100% regardless of the mastery level they were told to imitate. The proposed fix extracts a curriculum knowledge graph from an open algebra textbook, decomposes each solution into a chain of required knowledge triples, and samples a per-triple mastery probability from a Stochastic Student Knowledge Graph (SSKG) to decide correctness before the model writes a matching first-person rationale, dropping accuracy to 44.1-85.2% with a clean monotone gradient across profiles.

ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling

Deyi Li, Qi Xu, Lingyao Li, Tiansheng Wang, Muxuan Liang, Mei Liu Transformer models for clinical prediction from electronic health records (EHRs) still need heavy manual architecture tuning, and neural architecture search (NAS) is expensive while large language model (LLM)-guided variants restart from scratch at every hospital. ATHENA pretrains a weight-sharing supernet once per hospital so candidates are evaluated as fine-tuned inherited subnetworks, and adds a two-layer cross-hospital prior that retrieves high-performing architectures from source sites by task descriptor and estimates component effects via SHapley Additive exPlanations (SHAP)-based meta-regression, both feeding a multi-agent LLM search. Across six clinical prediction tasks at two health systems, it matches or beats four NAS baselines in 9 of 12 hospital-task evaluations at a search budget of 30 and selects architectures more consistently across repeated runs.

HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries

Shangxuan Tian, Yanhui Chen, Carlos Queiroz Document classification in regulated industries has to run on-premises with few initial labels and limited human review capacity, which rules out repeated model retraining. HIRA is a training-free cascade that fuses BM25 over OCR text, dense text embeddings, and image representations via calibrated weighted reciprocal-rank fusion, routes uncertain or visually confusable documents to a locally hosted LLM verifier, and escalates remaining doubt to a human, storing each correction as a retrieval exemplar rather than a weight update. On a private 80-class trade-finance stream of 30,233 documents it raises Macro-F1 from 0.6218 to 0.8548 while requesting human correction on only 6.4% of documents, and on Tobacco-3482 it reaches 0.9423 Macro-F1 with a DeepSeek-R1-Distill-Qwen-32B verifier, 17.4 points above zero-shot prompting while cutting LLM calls by about 60%.

Pruned Traffic Trees: Native Semantic Compression with a Protocol-Structured Model Family for Encrypted Traffic Classification

Yuantu Luo, Jun Tao, Xiangyu Xu, Linxiao Yu, Kangying Li cross-listed Deep models classify encrypted network traffic well but cost too much to run on routers and middleboxes, and standard compression acts on weights, channels, or predictions without deciding which protocol fields actually matter. Pruned Traffic Trees treats the protocol structure itself as the compression unit: a full model learns field salience over complete Protocol Tree Graphs with self-supervised flow-level training, that salience distills the graphs down for a mid-size variant, and a lite variant inherits the topology while narrowing its width through structure-aligned transfer and logits distillation. On CSTNET-TLS1.3 and CipherSpectrum the full model reaches Macro-F1 of 0.9519 and 0.9416, while the lite version keeps 0.9325 and 0.9136 with up to 80% fewer parameters, roughly 99% lower effective GFLOPs, and about 8.7× faster CPU inference.

NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus

Elisavet Lydia Alvanaki, Je Yang, Biruk Seyoum, Luca P. Carloni cross-listed Judging whether LLM-generated register-transfer-level (RTL) hardware is functionally correct before a trusted testbench exists currently depends on model-written testbenches or LLM-as-a-judge heuristics, both unreliable. NoTB instead generates RTL from several independently trained model families and runs Sequential Equivalence Checking to find designs that are provably equivalent, treating the family diversity within an equivalent cluster as a calibrated correctness signal. On 78 CVDP RTL-generation tasks, four-family consensus reaches 94.7% precision at 27% coverage and three-family consensus 87% precision at 33%, giving designers a tunable accept-or-defer rule with no oracle required.

ARCHER: Amortized cross-specimen pose estimation for cryo-electron microscopy

Nhan D. Nguyen, Bao Pham Pose estimation in single-particle cryo-electron microscopy is normally redone from scratch for each dataset, with the estimator effectively memorizing one molecule in its weights. ARCHER treats pose inference as a specimen-agnostic operation by conditioning explicitly on a reference volume, using an amortized contrastive classifier over a discrete rotation grid trained across many protein structures and applied zero-shot to new ones — transferability the authors ground in Fourier-space arguments where specimen dependence reduces to the reference power spectrum and spatial extent. It reaches a median angular error of 5.0° on 100 held-out structures and 2.5° on experimental particles, matching dedicated per-dataset estimators within 0.16 Å in 3D reconstruction while preserving conformational signal (leading conformational coordinate correlating at 0.97 with deposited benchmarks).

ReMAP: Self-supervised learning to unveil brain representations and vulnerability

Jade Perdereau, Virginie Loison, Kanssa El Ayeb, Louis Gervais, Melvin Berto Strouc, Fabrice Vall\'ee et al. Intraoperative electroencephalography during general anesthesia is typically collapsed into a single proprietary depth index, discarding how a patient's brain state moves over time. ReMAP applies similarity-based self-supervised learning to raw two-electrode frontal EEG with no labels, embedding each recording in a low-dimensional space where anesthetic depth becomes one readable axis and the shape of the trajectory carries additional structure. Validated on over 1,000 patients across two cohorts and two acquisition systems, it predicts depth accurately (BIS mean absolute error 3.2, R² 0.82) and a 68k-parameter model stays competitive with EEG foundation models of 4M–157M parameters in the sparse-montage setting; the space separates age along its own unsupervised gradient, aligns with known anesthetic signatures, and early-trajectory geometry separates 30-month cognitive and mortality outcomes at AUROC 0.86.

Autonomous Cyber Defense: Real-Time Attack Detection and Mitigation in Software-Defined Networks Using Machine Learning

Alexandre Amaral, Fernando Moro, Ana Malheiro cross-listed Attackers now move faster than human incident response: the average eCrime breakout time from initial access to first lateral movement fell to 29 minutes in 2025, with a fastest observed case of 27 seconds. The proposed system pairs a Network Dataset Creation module, which collects and aggregates IP flows into training data, with an Intrusion Prevention System module that trains and compares algorithms and then triggers blocking actions directly on a software-defined-network controller, removing humans from the detection-and-response loop. In a SYN flooding denial-of-service case study the attack was detected and blocked in 21 seconds with no human intervention, within the window implied by current breakout times.

On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study

Daniel Rodriguez-Cardenas, David Nader Palacio, Anna Schmedding, Yiyang Lu, Aadil Mallick, Bill Hudson et al. cross-listed Assigning Common Vulnerability Scoring System (CVSS) severity scores is manual triage that does not scale with disclosure volume, and outsourcing it to cloud language model services raises confidentiality problems for industrial code. This industrial case study predicts CVSS v3.1 scores directly from vulnerable C/C++ snippets using in-context learning with locally deployable open-source models, first showing that the public Big-Vul dataset has a sufficiently similar score distribution to proprietary data to serve as a testbed proxy. Comparing CodeLlama2-7B, CodeLlama2-13B, Mistral-7B, gpt-oss, and GPT4o-mini on mean squared error and feasibility, medium-sized local code models — CodeLlama2-7B in particular — approach the best cloud performance when guided by lightweight output-constraining prompts.

Beyond Fresh Starts: Stateful Inference for Streaming ASR in Conversational Voice Agents

Sameep Chattopadhyay, Alexander Erdmann, Mari Ostendorf Streaming speech recognition in voice agents runs under tight latency and memory limits, which makes it fragile around conversational phenomena like long silences and backchannels; the common fix of resetting recognizer state at every turn throws away context and hurts accuracy right where a new utterance begins. Two state-management strategies are proposed that carry cross-utterance context across turn boundaries instead of starting fresh. Tested with two state-of-the-art streaming models on two spoken dialogue benchmarks, the better strategy gives an average 15-21% relative word error rate reduction at utterance onsets.

Development and Feasibility Evaluation of an Edge AI as Medical Device System for Breast Cancer Multidisciplinary Team Meetings

Aarzoo Dhiman, Farzana Haque, Kartikae Grover, Lydia Brian Smith, William Stephen Jones Breast cancer multidisciplinary team (MDT) meetings carry heavy documentation loads, but existing AI support sends identifiable patient discussion to the cloud. The pipeline described runs entirely on a single NVIDIA Jetson AGX Orin, pairing open-source automatic speech recognition (ASR) via Whisper large-v3 with retrieval-augmented generation (RAG) grounded in National Institute for Health and Care Excellence (NICE) guidance to transcribe discussions, structure clinical information, and draft treatment recommendations. Across two recorded simulated meetings, ten clinically validated synthetic discussions, and 1,270 acoustically augmented recordings, ASR tuning cut word error rate by 20.7% and 24.4%, and MedGemma-RAG identified 2.3 times more MDT-concordant interventions than a proprietary cloud comparator with no significant difference in overall accuracy. Stakeholders rated automated documentation, recommendation support, and case triage as the most credible near-term uses while flagging governance, workflow integration, and clinician trust as obstacles.

VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

Yani Guan, Dengpan Dong, Shuang Luo, Zi Wei, Joah Han, Dan Hannah et al. cross-listed Optical Chemical Structure Recognition turns published molecular drawings into SMILES strings, and building datasets at scale requires spotting wrong predictions without ground truth. Comparing three label-free signals on 263 verified ACS journal depictions, re-rendering the prediction and comparing pixels performed near chance (AUROC 0.547) and an oracle-tuned threshold on it made results worse, while agreement among four architecturally distinct recognizers reached AUROC 0.916, with a two-of-four rule accepting 81.7% of images at 88.8% precision and three-of-four accepting 52.1% at 98.5%. Synthetic benchmarks hide this difference because re-rendered predictions naturally resemble their synthetic inputs; applied to PMC Open Access with a substance filter for wildcards and R-groups, VERDICT produced 6,146 labels for 4,833 molecules, with chemist adjudication confirming 0.995 precision at the three-of-four tier.

AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS

Timothy Tin-Long, Jian Zhu, Aidan Pine, Mengzhe Geng cross-listed Watermarking text-to-speech output normally requires retraining the generator or accepting a quality hit, but flow-matching and diffusion models leave a usable trace: their outputs stay strongly correlated with the initial Gaussian noise that seeded them. AudioNoisePrints exploits this by treating a plain cosine correlation between the seed noise and the generated audio as the watermark signal, adding no training and negligible inference cost, with a lightweight detector trained on top to survive harsher augmentations. It outperforms AudioSeal under strong augmentation, and the same spatial correlation appears in F5TTS, other TTS systems, and vocoders, suggesting the scheme generalizes across flow-matching audio models.

Does a Modern-Handwriting Warm-Up Help Historical Arabic OCR? A Reproducible, Compute-Matched Evaluation on Muharaf and KHATT

Sumaih Almarshad, Maram Alamri, Dona Aloraini, Fares Altuwaim, AlJawharh AlOtaibi, Reem Alyabis et al. Whether pretraining on modern Arabic handwriting helps historical handwritten text recognition is normally settled with one implementation and one comparison. Running the same nominal ablation four times — letting base checkpoint, encoder freezing, epoch budget, precision, and learning-rate schedule vary as they naturally did, while fixing normalization, scorer, and interval estimation — the estimated effect swings from -17.64 to +14.52 character error rate (CER) points and reverses sign, with the two extremes being exactly the runs carrying an identifiable confound and the two clean runs landing at -0.25 and +0.94. A compute-matched experiment over three seeds finds KHATT warm-up 2.42 CER points worse than a same-domain control (95% interval +0.60 to +4.25), of which only about 0.6 points is specific to the handwriting domain; the authors release a SaudiHeritage-OCR package with normalizer, scorer, decoder, manifests, and baselines so the result can be rechecked.

Tracing the Unlabeled Storm: Cross-Variable Transfer in a Lagrangian Atmospheric JEPA Framework

K M Anirudh, S Sandeep, Hariprasad Kodamana Learning a world model of South Asian monsoon convection directly from precipitation is hampered by rainfall's zero-inflated, heavy-tailed distribution, whereas continuous proxies such as outgoing longwave radiation express convective organization more coherently. M-JEPA, a multiscale Joint-Embedding Predictive Architecture, is pretrained on five continuous proxy fields over Lagrangian patches that track moving convective systems, with rainfall never observed during pretraining; the frozen representation is then transferred to daily precipitation forecasting through a shared decoder with probabilistic and deterministic branches. Controls show the benefit comes from proxy pretraining specifically — training directly on rainfall gives 36% higher CRPS error (7.52 vs 5.54 mm/day) — and the 15.4M-parameter model on a single consumer GPU edges out the 51-member operational ECMWF ensemble on CRPS (6.81 vs 6.89 mm/day) and Brier skill (+0.05 vs -0.04), concentrated at heavy-rain thresholds, while the ensemble keeps an advantage on neighborhood and point metrics.

GRAFT: Graph-Distilled Generative Retrieval for Facet-Aware Scientific Literature Exploration

Italo Luis da Silva, Hanqi Yan, Yujing Wang, Jiangnan Ye, Lin Gui, Yulan He cross-listed Papers relate to each other by problem, method, result, or contribution, but document-level retrievers collapse those facets into one opaque similarity score, and citation-based search stays near what a user already knows. GRAFT builds a graph with facet-typed edges derived from facet items and citation signals, then distills it into a generative retriever whose document identifiers are the papers' own facet text, addressing two properties that naive distillation loses: edge-only training pairs index just 84% of the corpus (fixed by coverage-aware distillation with reverse-neighbour fallback, a minimum-coverage threshold, and edge-importance weighting), and constrained decoding guarantees valid identifiers but not graph support (fixed by graph-weighted reciprocal rank fusion). On LitWeave, a corpus of 11,359 NLP papers, it recovers 91% of its graph teacher's Recall@20 with no nearest-neighbour index or encoder at inference, beats the teacher on query papers outside the corpus, and reproduces facet labels at 0.922 precision so each result arrives labelled with why it surfaced.

KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget

Ahmed Abolfadl cross-listed EU Regulation 2024/886 requires euro credit transfers to settle in under ten seconds, eliminating the overnight batch window in which anti-money-laundering analytics traditionally ran. KONTOGRAPH is an end-to-end AML pipeline for the SEPA Instant rail built to a self-imposed 200 ms 99th-percentile budget and evaluated on 1,562,860 simulated payments with injected typologies and deliberately incomplete labels; a temporal graph network with per-node memory lifts PR-AUC from 0.0053 to 0.1717 over a gradient-boosted tabular baseline, with per-node memory alone more than doubling the score. Two engineering findings generalize beyond AML: property-based tests that perturb the future caught three point-in-time feature leaks that code review had missed, and exporting the deployed tree ensemble to ONNX shifted mean scores by only 7.4e-8 yet changed 0.26% of decisions and inflated alert volume by 12%, because 32-bit accumulation moves scores across a razor-thin cost-optimal threshold. The authors argue a serving-format conversion is a model change until measured, and report a null result that subgraph-explainer fidelity metrics can be vacuous when candidate neighbourhoods are small.

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckstr\"om Qualitative researchers increasingly reach for language models to code open-ended text, but there is little evidence on how well they handle inductive content analysis, where categories emerge from the data rather than a predefined scheme. Five human coders and GPT-5.4 independently coded 903 open-ended responses across six variables from a European PhD student survey, with agreement measured by the Adjusted Rand Index. Human-model agreement reached 0.61 for coding and 0.54 for theme generation, close to the 0.68 agreement humans reached among themselves, while the model was more internally consistent than humans at 0.76. Agreement varied sharply by variable, with low internal consistency tracking low cross-entity agreement, suggesting the method works as scalable support mainly at the coding level.

When Does AI for PDEs Yield Scientific Evidence?

Wenshuo Wang Benchmarks for machine learning applied to partial differential equations (PDEs) score predictive accuracy against reference solutions, but in physics research the output is typically used as evidence for a scientific claim, which is a different standard involving assumptions and evidential thresholds. The authors extend an established PDE-simulation benchmark and a PDE inverse-problems benchmark so that model outputs can be scored on whether they support a stated claim, not just how closely they match a target. Accuracy rankings and evidential-support rankings can disagree, and existing benchmarks sometimes favor methods whose outputs provide weaker support for the scientific claims researchers actually care about.

KMGen: A Skill-based Approach for Synthetic Individual Patient Data Generation

Jalen Jiang, Chufan Gao, Ethan Rasmussen, Stephen Z. Xie, Jimeng Sun Individual patient data from clinical trials underpins survival modeling and meta-analysis but is almost never published, and prior reconstruction work recovers only survival curves, usually with manual digitization, and offers nothing for adverse-event records. KMGen closes both halves: an agentic pipeline writes code to extract Kaplan-Meier curves automatically, reaching mean integrated absolute error of 0.0151 on a 32-plot benchmark including adversarial cases, while a second stage has a language model distill registry records into arm-specific statistics, demographics, and risk multipliers that feed a mechanistic sampler generating per-patient adverse-event trajectories. On three held-out oncology trials with 30 regenerations each, survival curves matched within 0.051, demographic divergences stayed under 0.013 for five of six slots, and at least 71% of the top-15 adverse events were recovered by exact MedDRA term; the code is open source.

Do Not Copy/Paste: Soft Barriers for Copying in AI-Assisted Programming

Iyiola E. Olatunji, Alberick Euraste Djire, Jacques Klein, Tegawend\'e F. Bissyand\'e cross-listed Pasting a generated function from a chat window into an editor takes under a second, and the authors treat that instant — when model text crosses into executable or committed software — as an unmanaged design boundary they call the AI code handoff problem, arguing coding assistants should be judged on how they mediate transfer, not only on the code itself. They propose soft barriers, which keep AI assistance fully available while making unexamined transfer less frictionless, and probe the idea with Unicode output perturbations that stay visually readable but break naive copy-paste. Measured by Copy-Paste Resistance — the share of functionally correct solutions rendered syntactically invalid — across HumanEval and MBPP with four LLMs and four perturbation families, barriers can achieve high resistance but effectiveness swings sharply with model and task. A pilot with 18 participants offers early evidence that such friction shifts users from direct pasting toward editing and reconstruction; the perturbations are framed as a research probe, not a deployable mechanism.

TEE-X: TEE-aware Acceleration Framework for Large Vision Models at the Edge

Kurt M Wilson, Mohaiminul Al Nahian, Abeer Matar A. Almalky, Sadat Shahriyar, Souvik Kundu, Zhishan Guo et al. cross-listed Running vision models inside a Trusted Execution Environment (TEE) protects model weights and execution integrity, effectively downgrading an attacker from white-box to black-box access, but TEE memory limits and slow execution make this impractical for latency-sensitive edge deployments. TEE-X adds a sensitivity-aware modularization technique that decides which parts of the network must stay inside the enclave, plus vectorized execution for TEE inference, so a full Vision Transformer (ViT) can be hosted entirely within the enclave. Evaluated on OP-TEE for Arm TrustZone on an NVIDIA Jetson AGX Xavier, the framework reaches GPU-level inference latency with minimal accuracy loss while keeping the model confidential.

EGAMA-RC: Risk-Calibrated Evidence-Gated Adaptive Malware Analysis for Robust and Interpretable Memory-Forensic Triage

Isaac Kofi Nti cross-listed Malware detectors can post high clean-data accuracy yet still be unusable for triage, which also needs uncertainty estimates, novelty detection, robustness, interpretability, and bounded analyst review cost. EGAMA-RC builds on SHAP-guided feature refinement and adds model-pool evaluation, adversarial and open-family testing, novelty scoring, explanation-conditioned evidence, and runtime-aware routing, so low-risk memory-forensic samples are auto-accepted while uncertain, high-risk, or novel ones go to human review or escalation. Across three malware datasets under a frozen multi-seed protocol, the selected hybrid gate auto-accepts 93.12% of samples at 99.86% accepted accuracy with a 0.136% false-accept rate, with XGBoost giving a fast path at 0.0054 ms median latency per sample.

SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale

Jiaqi Liu, Maolin Ran, Xiaoyang Lu, Jian Wang, Weiwen Liu, Jianghao Lin et al. Industrialized short-drama video pipelines generate each shot in isolation, so props, character posture, and blocking drift between shots and the discrepancies compound into visible breaks once an episode is assembled. SEAM (Shot Entity-Attribute Memory) is a training-free, model-agnostic memory graph that fixes continuity purely at the prompt-text layer: it extracts a multi-dimensional state per shot, retrieves only causally prior context from the graph, filters it, and injects the surviving constraints by rewriting the natural-language prompt. On the released SEAM-Bench double-blind continuity benchmark it raises cross-episode continuity recall from 0.700 to 0.946 across six text models, and in a 201-shot production deployment it reached a 96.5% director-acceptance rate with no unsafe injections.

DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion

Xuan Yao, Li Shuping, Dai Yang, Zhou Yi, Ke-Wei Huang Financial institutions rely on vendor databases of corporate events that can be missing, stale, or misclassified, with no independent way to check them. The authors define Search-to-Record, a task where search-enabled language models reconstruct event records from public sources given a security universe and a historical cutoff, and release DelistBench, a 1,200-record benchmark of security-level delisting announcements. Across five models tested with and without web access, retrieval raises announcement-date accuracy within seven days by 34 to 48 percentage points, and the best system reaches 81.5% joint accuracy within a seven-day window, while cheap web systems land at 75.9-78.3% for roughly 5% of the API cost. Risk-based triage isolates low-error subsets but still routes 27.3% of the balanced test set to human review.

Performance of a domain-specific large language model in answering patient questions in psychiatry

Alexander J. Hish, Arjun Nagendran, Scott N. Compton The study asks whether a language model trained only on vetted patient education material answers medication questions better than general chatbots. The authors built MIND, fine-tuned on patient education resources from authoritative medical organizations, and compared it against ChatGPT and OpenEvidence on patient questions about escitalopram using both an automated rubric covering accuracy, clarity, completeness, nuance, safety, and referral appropriateness, and ratings from ten board-licensed psychiatrists. The rubric scored MIND highest in every domain, but the human raters split the difference — ChatGPT was marginally more often accurate, MIND more often complete, both equally safe — and 57.6% of psychiatrists preferred ChatGPT's responses despite MIND's greater completeness.

TSWAP: A Multilingual Retrieval-Augmented Thai Wellness Advisor

Pornthep Ukosaramig, Kobkrit Viriyayudhakorn TSWAP is a deployed eight-language conversational wellness advisor that grounds an unmodified open-weight model (Qwen3.6-35B-A3B served on vLLM) in a verified knowledge base of Thai traditional medicine and certified providers via retrieval-augmented generation, combining a roughly 30,600-chunk Thai index, hybrid dense-sparse retrieval with cross-encoder reranking, a first-turn classifier that forces tool-based retrieval for entity lookups, and a rule-based medical-scope and emergency-routing safety layer. The release includes the first Thai traditional-medicine retrieval benchmark (50 questions with gold document IDs, Recall@5 of 0.88), production QA logs passing 91.1% test-retest over 259 cases, and a 71-question no-retrieval probe showing that without the safety prompt the backend model emitted a full drug-dosing schedule, and without the knowledge base it produced zero verifiable provider recommendations. Two deployment lessons generalize: English-calibrated 4-bit AWQ quantization corrupts Thai tone marks, and forced-retrieval routing is necessary for reliable grounding.

Closed-Loop Bayesian Molecular Inverse Design with Semantic LLM Surrogates

Yaoyao Xu, Xinjian Zhao, Xiaozhuang Song, Lei Bai, Tianshu Yu Molecular inverse design in practice is a closed-loop enrichment problem — raising the fraction of generated molecules matching a target property profile under a limited oracle budget — and standard Bayesian optimization surrogates work over compressed continuous embeddings that discard the substructure and similarity cues chemists rely on. The proposed framework moves the design choice into the surrogate, using a frozen large language model that reasons over the task instruction, SMILES-level optimization history, and oracle feedback as text, then emits a structured signal selecting informative reference molecules under an explore-exploit rule plus an optional guidance sentence that becomes conditioning text for a frozen generator. On MolQA drug and material tasks it beats one-shot prompting and is competitive with or stronger than Gaussian-process Bayesian optimization baselines, with reference-only transfer best for binary drug targets and added surrogate summaries more useful for continuous material properties.

Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study

Guan-Hua Wen, Kuan-Yu Chen Time-series foundation models promise reusable representations and zero-shot forecasts for industrial monitoring, and this cost-aware study tests that promise across three settings: a C-MAPSS degradation-risk proxy, normal-only anomalous-sound detection on MIMII, and BDG2 forecasting-residual diagnostics with synthetic perturbations, comparing classical one-class methods, compact autoencoders, and residual forecasters against MOMENT-small, Chronos-T5, and TimesFM 2.5. On 100 C-MAPSS engines a temporal convolutional autoencoder reaches AUROC/AUPRC of 0.9570/0.8960 versus 0.7310/0.3080 for MOMENT reconstruction with bootstrap intervals excluding zero, OCSVM also beats MOMENT on MIMII pumps, and TimesFM 2.5 wins only the BDG2 forecast-error comparison. Same-device measurements show MOMENT costs more latency, VRAM, and storage than the small autoencoder, supporting the conclusion that frozen zero-shot foundation models are task-dependent options rather than default replacements for fitted lightweight baselines.

Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval

Kishore Konda cross-listed Approximate nearest neighbor indexes for dense retrieval treat each embedding as a single indivisible point in high-dimensional space, forgoing the inverted-index machinery that made sparse retrieval fast. The Hypergraph Embedding Index instead groups documents by which combinations of latent embedding dimensions are most strongly activated, allowing inverted-index-style candidate generation while keeping dense semantic ranking. Building several complementary hypergraphs raises retrieval coverage without the combinatorial blowup of enlarging a single hypergraph, and the authors introduce activation diversity as a diagnostic statistic predicting how well a given embedding model suits coordinate-inverted indexing.

LLM Pedagogical Behavior in AI Tutoring Interactions

Suhyeon Lee, Juneha Baek, Jaehyeong Park, Donghyuk Shin Students routinely use language models as tutors, but there is little evidence about how much of the work those models simply do for them. A five-level "scaffolding" scale, validated against human annotations, grades responses by how directly they complete the student's task, and it was applied to 14,637 model responses gathered from 203 students in a university AI course. Over 95% of responses landed in the two most directly helpful categories, Explaining or Solving; scaffolding level correlated with how students behaved in the rest of the conversation but added little predictive power for later exam scores beyond prior achievement and dialogue behavior.

Neural Boltzmann Equations

Jonas Spinner, Jack Shergold cross-listed Tracking particle populations in the early universe means solving Boltzmann equations whose high-dimensional phase-space integrals are traditionally handled by quadrature on a fixed momentum grid, an approach that scales badly with process complexity and parameter scans. Neural Boltzmann Equations swap in three coupled ingredients: physics-inspired neural distribution functions whose parameters a network can predict across a parameter space, Monte Carlo evaluation of the phase-space integrals using importance-sampling tools borrowed from collider physics, and natural gradient descent to evolve the system in time. The framework is used to produce a precision calculation of the effective number of relativistic neutrino degrees of freedom, after demonstrating each component's contribution separately.

Graph Representation Learning of Lightweight IoT Ciphers

Jonathan Cook, Sabih ur Rehman, M. Arif Khan SIMON and SIMECK, lightweight Feistel block ciphers designed for Internet of Things devices, need rigorous evaluation against differential cryptanalysis, and prior speedups rely on heuristics and sampling rather than learned structure. The approach extracts four differential attributes from a partial Difference Distribution Table and uses them to build directed graphs via K-Nearest Neighbours, Decision Trees, and Random Forests, producing what the authors describe as the first graph-based visualisation of the differential clustering effect, where high-probability single-bit differentials sit close together in the learned embedding. All three models reached precision of 1.0 with zero false positives in identifying high-probability differentials, with K-Nearest Neighbours giving the best cluster separation and building graphs in roughly 2.3 seconds, and results held for both ciphers.

When a neural surrogate cannot accelerate a solver: runtime share, closed-loop drift, and the economics of uncertainty gating in a stiff coupled simulation

L. Th\"ummler, T. Kuroda cross-listed Replacing an expensive inner solver block with a learned surrogate is a common route to faster multiphysics simulation, and this controlled end-to-end study reports a negative result on the implicit Newton solve coupling energy-dependent neutrino radiation to matter in a general-relativistic radiation-hydrodynamics code. Three structural barriers appear, none traceable to network quality: the target block occupies only 16.9% of critical-rank wall clock, so Amdahl's law caps any surrogate speedup at about 1.2× and a surrogate 5.8× cheaper per call merely ties the original solver; offline accuracy fails to rank surrogates for deployment, with a pooled Spearman correlation of +0.73 between error and survival collapsing to -0.04 once family confounds are controlled across fourteen networks; and a correct out-of-distribution gate cannot help a loop that leaves its training distribution, deferring 96.8% to 99.7% of cells because visited states sit 73× off the data manifold, yielding a net 0.94–0.96× slowdown. Separately, a gated run that never crashes still accumulates a directed -19.9% density bias over 6000 steps, a ballistic drift distinct from the variance-driven divergence usually studied in autoregressive rollouts.

PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors

Nathan Duboisset, Zhaolan Huang, Felix Bie{\ss}mann, Roudy Dagher, Antoine Lavandier, Emmanuel Baccelli Battery-powered acoustic sensors running TinyML — machine learning on microcontrollers — have so far only managed to detect a single bird species at a time, which is too limited for real field deployments. PolyChirp combines curated training data, neural architecture search, and a microcontroller with a neural processing unit to fit multiclass models within a strict memory, latency, and energy budget. The resulting models classify up to 10 bird species simultaneously while still running an entire breeding season on one battery charge, and also beat prior single-species binary detectors.

Future Querying: Can LLMs Serve as Implicit Medical World Models?

Siri Willems, James Butterworth, Lore Goetschalckx, Peter Vrancx, Philippe Modard, Elke Giets et al. Clinical prediction usually means task-specific pipelines over curated structured fields, which scale badly and ignore the free text that makes up most of a patient record. The proposed "future querying" paradigm instead tests whether a language model can act as an implicit medical world model by answering time-indexed questions about a patient's future directly from unstructured documentation, trained endpoint-agnostically so one model handles diverse queries without feature engineering or per-task retraining. Evaluated on a new synthetic medical report dataset and real intensive-care notes from MIMIC-IV, small locally fine-tuned open-weight models match or approach much larger proprietary systems, which makes fully on-premise, privacy-preserving deployment viable.

FIDES: A Concordance Protocol for LLM-Generated Trading Strategies

Arther Tian, Alex Ding, Simon Wu, Aaron Chan cross-listed Ask a language model for a trading strategy and you get three things at once — a prose rationale, runnable code, and eventually a backtest record — but whether they describe the same strategy is seldom tested. FIDES elicits a natural-language strategy with an explicit claimed edge alongside a self-contained strategy(df) function in one call, executes the code in a sandbox against a lag-one out-of-sample backtest, and scores three concordance gaps: say-to-do, do-to-real, and say-to-result. Across 40 strategies on 8 liquid US ETFs and four models over 2023–2024, only 2 of 40 strategies beat buy-and-hold while 32 claimed they would, a plain 50/200 simple moving average rule outperformed every model's mean Sharpe ratio, and swapping the language-versus-code judge model flipped the say-to-do verdict on over half the items.

Spicing up Genetic Netlist Generation with LLMs

Stefan Uhlich, Ya\u{g}{\i}z Gen\c{c}er, Andrea Bonetti, Arun Venkitaraman, Chia-Yu Hsieh, Eisaku Ohbuchi et al. cross-listed Analog circuit topology synthesis is a needle-in-a-haystack combinatorial search where small structural edits cause highly nonlinear behavior changes; evolutionary methods handle the discrete space with black-box evaluations but burn many SPICE simulations and converge prematurely. LLM-SPICEMixer adds IGEL, an inspiration-guided proposal operator that prompts a language model with high-performing circuits from the elite set to generate a new SPICE netlist, which is then simulated and selected by the same reward mechanism as conventional genetic operators, keeping simulation as the source of truth. On the task of synthesizing transistor-level circuits implementing a discriminant function for Iris classification, the hybrid improved median final training reward by 8.4% and validation-selected test reward by 8.8% over the genetic baseline, with the best circuit reaching 93.3% test accuracy at the nominal corner and 85.9% averaged across 17 process, voltage, and temperature corners.

Beyond Point Predictions: Uncertainty-Aware Satellite Poverty Mapping for Public Policy

Markus B. Pettersson, James Bailie, Mohammad Kakooei, Eagon Meng, Adel Daoud High-resolution poverty data are scarce across much of Africa, and while machine learning on satellite imagery can fill gaps, policymakers need to know how far to trust the predictions. This approach pairs a spatiotemporal transformer trained on Landsat and nighttime-light image sequences with simultaneous quantile regression and a new form of conformal prediction, producing neighborhood-level International Wealth Index intervals with statistically guaranteed coverage. Point accuracy matches the state of the art at R² of 0.75, yet the intervals are wide enough that earth-observation models cannot naively drive poverty-targeting decisions; the authors therefore add an aid-allocation procedure combining ground-truth surveys with model predictions that provably bounds the risk of excluding eligible neighborhoods and delivers more aid per eligible recipient in simulation.

Robustness of IR Models to Collection Growth

Emmanouil Georgios Lionis, Debasis Ganguly, Sean MacAvaney cross-listed Retrieval systems run against collections that grow over time, and adding documents that are irrelevant to a query arguably should not hurt ranking quality — a property that had not been formalized. The evaluation merges two collections with negligible topic overlap and splits models by whether their scoring depends on the rest of the collection: multi-document-agnostic (MDA) models score each document alone, while multi-document-dependent (MDD) models condition on others, as BM25's inverse document frequency term or a listwise reranker's context does. No model tested was fully robust — all degraded when non-relevant documents were added — and among retrievers MDA outperformed MDD, while for rerankers the two classes were comparable.

RAD: Rule-Augmented Relational Anomaly Detection

Noah Dahle, Anne Tumlin, Ngoc Tran, Xenofon Koutsoukos, Tyler Derr Anomaly detection over relational databases usually flattens many tables into one feature matrix, destroying entity identity, schema structure, and multi-hop dependencies that some anomalies only reveal themselves through. RAD keeps the relational structure by learning over a heterogeneous graph while also injecting symbolic evidence: candidate rules are extracted from random-forest decision paths over flattened entity summaries, refined into compact interpretable predicates, added as features to the graph model, and trained with reconstruction and pairwise-ranking objectives. A new benchmark spans LANL cybersecurity event detection and two user-churn anomaly tasks from Amazon and H&M databases, where RAD achieves the best average rank on both AUROC and AUPRC under natural class imbalance; ablations credit rule injection and ranking supervision, while edge reconstruction is not consistently helpful.

The Measurement Revolution? Credible Measurement and Inference in the Age of AI

Melissa Dell, Ashesh Rambachan cross-listed Economists can now turn text and images into structured variables cheaply, which moves the hard part of empirical work from finding any scalable measure to choosing among many plausible ones that can support conflicting conclusions. The review maps three points where AI enters the measurement pipeline — discovery, construct definition, and observation — and argues that credible inference requires validation anchored to explicit criteria rather than informal claims that a proxy is reasonable. It then covers how a validation sample permits valid inference even when the AI-generated predictions are arbitrarily biased, and what recourse exists when no random validation sample can be drawn.

Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition

Kyle Stein, Guillermo Francia, III Eman El-Sheikh, Andrew Arash Mahyari cross-listed Malware detectors must absorb new variants continually, but incremental updates overwrite prior knowledge, and the case where each new class arrives with only a handful of labeled samples — Few-Shot Class-Incremental Learning (FSCIL) — has received little attention in this domain. The proposed framework starts from a self-supervised backbone pre-trained on malware packets, applies Low-Rank Adaptation (LoRA) during the base session while freezing the backbone to preserve learned representations, and switches to a prototype-based classification head for incremental sessions. Across several datasets it consistently beats prior malware FSCIL baselines, reaching state-of-the-art results on the benchmarks tested.

Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination

Mustafa Umut Ozbek, Taiwo Ojo, Pooria Madani, Khalil El-Khatib, Li Yang cross-listed Anomaly detectors for industrial control systems are almost always evaluated assuming clean training data, yet in practice training pools can be corrupted by compromised logs, mislabeling, tampered historian records, or unsafe retraining. Eleven detectors were tested on the SWaT water-treatment benchmark under three contamination strategies — random injection of attack samples, similarity-targeted injection, and bounded Gaussian feature noise — at budgets from 1% to 10%, evaluated against clean validation and test sets. Robustness proved strongly model-dependent and not predictable from clean-data performance, with injection attacks hurting local-density and distance-based detectors most, PCA, SVM, HBOS, and IForest staying relatively stable, and tuned neural detectors landing in between.
116 more specialized papers

Agents 92

Small Language Model enabled Autonomous agent for Language-Conditioned Cognitive Radar

Minhaj Uddin Ahmad, Zakia Zaman, Shunqiao Sun, Mizanur Rahman cross-listed Radar signal processing needs to switch strategies as interference, clutter, and snapshot availability change, which normally requires an expert to choose and configure algorithms. The described system uses a small language model as an autonomous controller that parses a natural-language command, extracts radar-relevant cues, picks a sequence of array signal-processing methods, sets their parameters, and calls executable tools to do the actual numerical work. On a simulated uniform linear array, the agent selects appropriate algorithms for sidelobe control, jammer suppression, multiple-null beamforming, coherent-source handling, and low-snapshot direction-of-arrival estimation, and ablations show that both radar-specific prompting and physics-grounded tool execution are necessary to avoid unreliable choices and hallucinated numbers.

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu et al. Judging machine-generated literature reviews is hard because research utility depends on expert judgment that reference-overlap metrics cannot capture. LitReview Arena is a battle-style platform where domain experts with AI paper-writing experience compare anonymized drafts on topics matched to their expertise, scoring five literature-review-specific dimensions; roughly 3,000 expert judgments were collected. The strongest systems win only 23.0% of decisive matches against human-written drafts on overall utility, though agentic systems such as Sonar Deep Research beat base language models by over 60%, and existing LLM-as-a-judge methods correlate poorly with experts (Spearman's rho of 0.467), especially on synthesis-heavy criteria; a calibrated evaluator, LitJudge, trained on the preference data raises that to 0.78, near inter-expert agreement.

SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG

Yong-eun Cho Agentic retrieval-augmented generation systems that span external APIs, databases, vector stores, and graph stores either dump every tool description into the prompt or pick tools by vector similarity, producing over-fetching that inflates tokens and latency or under-fetching that drops needed fields. SchemaRouter models tools, endpoints, parameters, response fields, domain concepts, units, provenance, and license policies as a schema graph; a small language model extracts intent and constraints while field selection stays deterministic over the graph, yielding an executable plan of which tools to call and which fields to pull. On 110 materials-science queries it reaches 0.71 answer accuracy with a 227-token retrieved context versus 2,066 for fetch-everything and 2.7x lower latency than prompt-all, plus 0.93 tool-exact rate and provenance grounding in 62% of answers; notably, aggressively minimizing field count backfires, dropping accuracy to 0.56 for negligible token savings.

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen et al. cross-listed Personal AI assistants today personalize within a single app, leaving cross-context user understanding largely unmeasured. PersonaMem-v3 seeds a benchmark from over one million anonymized real engagement histories — mostly implicit signals — to construct time-indexed digital worlds spanning social media, chatbot, calendar, and AI-companion surfaces with preferences that evolve over time. It folds personalization, recommendation reranking, proactive behavior, agentic tool use, and geo-temporal reasoning into one harness, and notably tests whether agents hold back when personalization would be inappropriate, repetitive, outdated, or unnecessary rather than only rewarding more personalization.

Retrieval-grounded robot program generation and simulation-based correction via Model Context Protocol

Zhichao Zhou, Siyuan Chen, Omkar Salunkhe, Ebru Turanoglu Bekar, Johan Stahre, Anders Skoogh Flexible manufacturing needs industrial robots reprogrammed quickly as product variants change, but ungrounded language models generate domain-specific errors in vendor robot code. The described workflow generates ABB RAPID programs from natural-language task descriptions using a dual-stream retrieval-augmented generation pipeline grounded in verified technical documentation and production templates, then closes the loop with a custom Model Context Protocol (MCP) server that connects the model client to ABB RobotStudio for automated upload, simulation, and diagnostic feedback. Evaluated with a 30-query retrieval benchmark, scoped code-generation checks, and pick-and-place case studies, the simulation loop caught execution failures that static and semantic checks alone missed, including suction release-height errors, unreachable placement targets, and configuration-dependent recovery motions. Expert setup and final supervision were reduced but not eliminated.

Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing

Israt Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. Mridha Systematizes the operational failures that appear when large language model (LLM) agents drive security tooling, based on hands-on evaluation of ten static, dynamic, cloud, orchestration, and AI red-teaming tools in unattended pipelines. The authors propose a four-dimensional Integration Friction Index separating one-time engineering cost from recurring organizational, legal, and maintenance cost, and model an agentic security system as stochastic LLM policies wrapped by a deterministic mediator. Derived regularities include evidence loss in long-lived sessions growing with phase count, diminishing returns from two-stage verdict cascades when scorer errors correlate, measurement bias from scoring unevaluable outcomes as attack failures, and a closed-form execution cap for heavy-tailed tools; scope and budget enforcement cannot be delegated to system prompts, since prompts do not constrain what actually executes. Their platform Inspectra serves as a worked instantiation, labeling each mechanism shipped, partial, or planned.

FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows

Darshan Deshpande, Yoshinari Fujinuma, Martyna Markiewicz, Devanshu Bansal, Shivani Jain, Nicholas Saban et al. cross-listed Attributes the weak performance of vision-language models on subjective design work to a shortage of human workflow data capturing expert preferences and decisions. The authors define an expert-curated taxonomy of design skills, expand it into 126 open-ended long-horizon tasks, and release FigmaTrace: over 200 hours of captured human video converted into 3,469 design trajectories using a design phase-based segmentation method rather than fixed-length chunking. Four models trained on the dataset reach performance comparable to frontier closed models such as Claude-Opus-5 and GPT-5.6-Sol on four out-of-distribution agentic GUI environments, with an ablation crediting the phase-based conversion over prior length-based approaches; the dataset and best model are open sourced.

Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents

Baicheng Chen, Zheyuan Liu, Jingyu Zhang, Kaize Ding, Ningshan Ma, Yue Huang et al. Machine unlearning is evaluated by checking whether a model still recalls forgotten facts from its weights, but a tool-using agent can simply retrieve the same fact through web search or a database — a failure mode the authors call tool-mediated recovery. Agentic Tool Unlearning (ATU) addresses it in two stages: standard parametric unlearning to suppress direct recall, then trajectory-level reinforcement learning inside simulated tool environments that penalizes both target-seeking tool calls and leakage in the final answer. On RWKU and MUSE across several architectures, the method balances forgetting the target against retaining normal tool use better than parametric-only unlearning.

K-Bench: measuring model performance on real scientific agent requests

Aubrey Brueckner, Darshil Patel, Yuhuan He, Timothy Kassis Benchmarks for scientific AI are mostly written to be scorable — multiple choice, curated tasks with reference solutions, simulators with known structure — whereas real scientific requests are underspecified, carry attachments, and have no ground truth. K-Bench 01 is built from first-turn requests sampled from live user traffic on K-Dense Web, run end to end by nine frontier models in identical sandboxes for 1,602 completed agent runs, with three blinded language-model judges scoring every run against an eight-dimension rubric. No model clears the rubric anchor at which a domain scientist would accept the work with minor edits: gpt-5.6-sol has the highest pooled mean at 8.04 but its interval spans the threshold and two of three judges rank claude-opus-5 first, while 47.6% of all 39,934 judgments fall below the bar. Scientific accuracy averages 6.22 against 7.33 for communication in every one of the nine models, and overclaiming is the leading failure tag at 31.4% of assessments.

Context as an Environment: Programmatic Context Management for Long-Horizon Agents

Yin Lin, Elaine Ang, Erkang Zhu, Bolin Ding, Jingren Zhou Long-running agent sessions accumulate history far past any context window, and the usual remedies — compressing old turns or extracting fixed memories — force a decision about what matters before anyone knows what will be needed. Scroll instead treats a session as an executable environment backed by an append-only event log and a persistent sandboxed Python kernel, where tool outputs, retrieved history, and derived state live as typed variables rather than being re-serialized into every prompt; the model writes code to search and transform that state, and only what it explicitly prints enters the working view. When the view fills up, stale spans are evicted but stay addressable through compact landmarks pointing at exact log positions, so the agent jumps straight to what it needs. With Qwen3.8-Max as backbone it reaches 94.8% on LongMemEval_S, 73.1% on BEAM_10M (5.1 points over the best published memory system), and 86.7% on LOCA_256K, 37.4 points above the best published long-horizon agent.

Architecture as Capability Equalizer for Coding Agents

Arquimedes Canedo cross-listed Coding agents build whole systems from high-level descriptions, but it was unclear whether the format of the architecture specification changes output quality or whether that depends on the model. A controlled experiment ran 90 multi-turn agent trials across six models from three vendors using five informationally equivalent specification formats: informal prose, Mermaid diagrams with architecture decision records, OpenAPI, C4/Structurizr domain-specific language, and TypeScript interface contracts with ArchUnit-style rules. Format barely matters for the strongest models (spread 0.17–0.92) but produces spreads up to 2.42 points on weaker ones, where code-proximate formats recover most of the capability gap — TypeScript contracts triple API route coverage for the weakest model — while mid-tier models can burn more tokens than frontier models for worse output by falling into compilation debugging loops.

ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents

Yu Qian, Hong Miao, Boyang Guo, Tingyi Jiang, Shan Zhao, Tianxing Le et al. Agents operating over long horizons need memory that surfaces relevant past experience, handles revisions, and exposes checkable provenance. ECHO (Embodied Context and History Orchestration) is a memory architecture and service prototype organized around functional analogues of episodic encoding, consolidation, contextual reinstatement, reconsolidation, and executive control, evaluated on retrieval and context construction rather than any claim of neural equivalence. Development runs report 96.29% Hit@10 on LoCoMo and 97.60% Hit@10 with 88.84% turn Recall@5 on LongMemEval-S, but a matched 91-question comparison has Mem0 OSS at 64.84% against ECHO's 41.76%, and a post-hoc audit found source-specific phrases in the query-expansion rules, so the authors label the expansion-enabled retrieval numbers descriptive development measurements rather than independent confirmation.

Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web

Qijia Chen, Giulio Jacucci Graphical user interface grounding benchmarks that expose interface elements as text metadata often read high similarity between an instruction embedding and an element embedding as evidence of semantic understanding. Framing each action as a same-screen ranking task across three mobile and web benchmarks, the authors compare five off-the-shelf sentence encoders against purely lexical baselines and find the lexical baselines competitive at top-1, label-poor targets still weak for text-only methods, and encoder top-1 hits predictable from lexical rank, candidate-pool size, and label type. Encoders do recover some lexical misses, but realistically deployable fusion gains are far smaller than target-aware oracle gains, so such evaluations should always report lexical baselines and stratify by label type.

MSM-Mem: A Universal Medical Structured Multimodal Memory Framework for Medical AI Agents

Md Asaduzzaman Jabin, Khoa Le, Lin Zhao, Tianming Liu Medical AI agents built on multimodal large language models typically run stateless, deciding each case independently and never accumulating the case experience that clinicians build up over years. MSM-Mem is an agentic memory framework that organizes heterogeneous clinical experiences into semantic, episodic, and visual memory stores, updates them incrementally during inference, and retrieves prior cases to condition current reasoning. Evaluated on MoE-LLaVA backbones, it yields consistent accuracy gains that grow further with continued usage, suggesting a route to medical agents whose competence improves through deployment rather than retraining.

Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores

Yu Pan, Hongfeng Yu Retrieval-augmented generation normally treats the document store as fixed input, and agent-curated stores are rarely measured for what curation actually buys. Here the knowledge base itself is the trained object: an agent answers supervised questions against the current store, sees the gold answer, and edits the store, after which an unchanged reader is tested on a frozen snapshot under a fixed action budget. Per point of corpus indexed, the supervised structure returns 1.6x the action saving and 1.8x the accuracy of an unsupervised entity index while using 1,913 links against the unsupervised index's 196,112, and on trained questions the reader spends 31% fewer actions at higher accuracy, reproducing on an official PhantomWiki generation. A key-coverage gradient replaces a pass/fail train-test split with a decay curve, showing accuracy transfers to unseen questions in proportion to how much of the question's keys were indexed (+0.167 F1 with both keys covered, zero with neither) while action savings stay on trained questions.

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

Chengyang Gu, Le Zhang, Jingbo Zhou, Yize Chen, Yu Shi, Siqi Bao et al. Reinforcement learning for graphical user interface agents commonly uses Group Relative Policy Optimization (GRPO), which suffers reward-gradient misalignment; recent contrastive reformulations of reinforcement learning with verifiable rewards fix the instability but supervise only on final outcome, ignoring quality differences among trajectories that share the same outcome. LACL-GUI adds trajectory-level structure to the contrastive objective, preferring concise successful executions over meandering ones and ranking failed trajectories by how far they diverge from successful ones, while keeping the stability properties of the contrastive formulation. On GUI agent benchmarks it consistently improves over prior contrastive and GRPO-based methods.

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding

Md Abrar Jahin, Md Rizwan Parvez Computer-use agents must bind relational phrases like "above" or "inside" to the right on-screen element, but existing grounding benchmarks cannot separate that skill from general element localization. GUI-Primitives is a 994-item benchmark of contrastive instruction pairs over seven spatial relations, holding the screenshot and anchor fixed while changing only the relation so the correct target moves between two designated candidates. Nineteen vision-language models reach at most 32% strict point-in-box accuracy, and predictions fall outside both candidate regions on 60-92% of items; conditional on landing inside a candidate, target selection is 0.82-0.90 for horizontal position, vertical position, proximity, and list ordinal but statistically indistinguishable from chance for containment and occlusion, so most failures are localization rather than relation understanding. Marking the two candidates raises selection accuracy by 35-57 percentage points, an oracle diagnostic rather than a deployable fix.

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie et al. Building a working game exercises program logic, art and audio assets, interface, interaction, and playability in a single executable, and existing benchmarks judge coding agents on only the final artifact or one isolated stage. GameXpert-Bench splits the lifecycle into three tracks derived from real human-agent development trajectories: GameGen for creating a complete game from a single request in an empty workspace (97 tasks across 11 genres), GameFix for diagnosing and repairing defects (100 tasks over 50 game levels, each seeded with 19-27 injected bugs), and GameOpt for cumulative optimization across six-turn request chains (17 chains, 102 requests), scored via live game interaction and deterministic behavioral tests with regression checks. Across all three tracks, agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, or preserving functionality across changes.

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

Hui Zeng, Pengfei Yang, Yanxin Chen, Fusong Ju, Xinran Wei Kernel-optimization benchmarks measure standalone GPU kernels, but a kernel that looks correct and fast in an isolated harness can behave differently once it runs inside an actual inference server. LLM4LLM closes that gap by starting from a target inference script, extracting phase-aware optimization tasks, searching for candidate patches with an experience-guided episodic agent, and validating each patch in the real model rather than a synthetic harness. Across ten inference workloads, the framework improved end-to-end latency for every model tested, reporting geometric-mean speedups of 3.91× on A100 and 6.98× on H100, plus up to 2.745× geometric-mean gains on KernelBench Level 2.

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

Yucan Guo, Xiaohan Wang, Miao Su, Saiping Guan, Zhongni Hou, Jiajun Chai et al. Reinforcement learning for tool-integrated reasoning normally hands every rollout a single trajectory-level advantage, treating an easy tool call and a hard one as equally instructive. HiDiffTIR instead assigns credit at two granularities — across trajectories and across individual turns — weighting harder, more informative steps more heavily, and derives the difficulty estimates purely from group-level statistics already produced by standard rollouts, so no extra supervision or annotation is required. On three tool-using benchmarks it improves both task success and tool invocation accuracy over strong reinforcement learning baselines.

MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance

Haoyu Wang, Guangyuan Dong, He Liang, Zijing Zhang, Jiachen Luo, Chuang Liu et al. Agent memory banks degrade in two ways over long task streams: unreliable admission lets failed trajectories, lucky successes, and misleading observations in, and drift accumulates duplicate, stale, and contradictory records that retrieval cannot fix. MemGuard treats verifier output as persistent lifecycle metadata rather than a one-time admission filter, converting multi-criteria score-token verification into reward, confidence, label, and uncertainty descriptors that are attached to each memory and reused during retrieval, conflict resolution, summarization, and archival. Tested on Terminal-Bench 2.0, SWE-Bench Verified, WebArena, and Mind2Web across four backbones against four memory baselines and a verifier-only control, it won on success rate and average steps in all 16 backbone-benchmark settings, with its largest gain of 7.9 success-rate points on WebArena over ReasoningBank.

From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning

Chenghao Zhang, Yikai Mao, Shanqi Liu, Haoyu Gao, SaiSai Hu, Dan Roth Turning natural-language instructions into PDDL planning specifications is brittle without expensive hand-written annotations, and models that optimize for solver success can produce specifications that solve but do not mean what the instruction said. This work learns the translation from solver feedback alone, with one language model playing three roles: an Actor that proposes specifications, a Judge that scores them against solver-calibrated criteria, and an Editor that performs bounded diagnostic-driven repair. On PlanBench the approach raises average success from 35.5% for LLM+P to 70.8%, with 66.3% faithful success and semantic drift down to 6.4%.

Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning

Chenghao Zhang, Canran Xiao, SaiSai Hu, Dan Roth Synthetic websites used to train web agents often look convincing while hiding broken links, inconsistent state, and tasks that cannot actually be completed, which corrupts the reward signal. This framework represents each generated site as a structured scaffold of pages, navigation links, database records, state-change markers, and task constraints, then verifies and repairs structural, semantic, consistency, and feasibility defects before any policy training; at interaction time ordinary UI transitions execute deterministically while persistent backend writes go only through validated state-change markers, enabling dense rewards compiled from verified task-progress predicates. Across 500 environments in six domains, the feasible-task rate rises from 48.6% to 94.8%, yielding stronger PPO policies that transfer to WebArena, WebShop, and MiniWoB++ with no language-model calls at evaluation time.

TessIndex: Capability Verified Identity System for the Agent Economy

Mehul Goenka, Tejas Pathak, Siddharth Asthana As software agents autonomously orchestrate tools, services, and sub-agents, the supporting infrastructure lacks persistent identity for accountability, verifiable rather than self-declared capability claims, and a link between creator identity, agent performance, and economic value. TessIndex proposes a dual-plane architecture in which a blockchain stores compact commitments for identity, ownership, and verification while centralized servers hold the mutable metadata used for discovery, commerce, and reputation. Its distinguishing mechanism is predicate-based verification that replaces self-reported capabilities with cryptographic proof of execution, tying an agent's capabilities, execution history, and reputation to one persistent identity.

EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

Can Xie, Yuyi Zhou, Wen Yang, Ziyi zhang, Siyao Song, Yingzhuo Deng et al. Outcome-based reinforcement learning such as GRPO (Group Relative Policy Optimization) throws away the reusable exploration patterns in agent trajectories after one update, while retrieval-based experience methods create a permanent inference-time dependency on an external store. EDGE treats retrieved experiences as temporary training scaffolds: it splits each rollout group into experience-conditioned and experience-free trajectories to identify which experiences give a positive marginal gain without extra sampling, then distills that behavior into the base policy with a reverse-KL objective, backed by a co-evolving experience bank that adds guidance for new failure modes and prunes stale entries. On ALFWorld and WebShop at 7B scale it beats GRPO by 8.3 and 12.5 success-rate points and keeps 96.0% of its scaffolded performance once external experiences are removed at inference.

Repo2Skill-Evo: Repository Skills Go Stale in Silence

Chenyuan Duan, Ge Shi, Zineng Mao, Ge Zhang, Hao Liang, Yinzhu Piao et al. Agent skills capture repository-specific procedural knowledge — which APIs to call, which scripts to run — but that same version specificity means a skill can silently go stale after a release while still handing agents obsolete guidance. Repo2Skill-Evo frames each release transition as a maintenance task: given a V1 skill set and the official V1-to-V2 patch, the agent must revise outdated content without disturbing what remains valid, across 57 real repositories and 105 release transitions, every one of which invalidates part of the V1 skills. Six frontier agents score only 29.9% to 69.7% avg@3 macro F1 under a patch-grounded metric, failing in two opposite directions: missing affected files entirely, or editing so broadly that precision collapses.

Closed-loop AI achieves certifiable engineering design

Tianyi Yu, Chengxing Tao, Haoxuan Shen, Huiyang Li, Rugang Chen, Long Teng et al. Physical engineering design has resisted agentic automation because candidates must simultaneously satisfy fluid dynamics, solid mechanics, and structural stability constraints. The AI Engineer closes the loop between language models and deterministic engineering backends: natural-language requirements become geometry and mesh, topology is optimized via bi-directional evolutionary structural optimization against the CalculiX solver, and member sizing is refined with particle swarm optimization against Zwind under offshore aero-hydro-servo-elastic loads, with an Automated Reviewer scoring candidates on capacity, steel intensity, cost, constructability, and fatigue life using functions calibrated on 11 real floating-wind projects. The top-scoring design passed Approval in Principle from the China Classification Society and cut steel mass and unit capital cost by 8.1% each relative to the human-optimized TuQiang baseline.

SPAR-Hate: An Auditor-Guided Multi-Agent Framework for Bilingual Hate Speech Parsing

Yifan Lyu, Dianqing Lin, Xinran Li, Jiaqi Qiao, Xiujuan Xu Structured hate speech parsing asks systems to jointly extract hateful targets, arguments, and target-level labels rather than emit a single label, and cultural and social-group context makes that hard. SPAR-Hate splits each document into clause-level decision units and has three role-conditioned agents — Victim, Moderator, and Cultural Bystander — produce evidence-grounded judgments, which an evidence-constrained arbitration step reconciles into sample-level structured output. On the STATE-ToxiCN and TBO benchmarks the framework improves bilingual parsing across several base language models and posts state-of-the-art multi-tuple extraction results, with the biggest gains under the strictest structural metrics.

Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets

Aashish Panta, Hugo Lee, Giorgio Scorzelli, Kyongsik Yun, Valerio Pascucci cross-listed Public scientific datasets from large facilities are hard to use without knowing repository layout, file formats, multiresolution structure, and visualization parameters. WebVisus is a constrained multi-agent system that takes a natural-language research question, infers intent, and dispatches an autonomous exploration agent that inspects slices, volumes, and timesteps of remote datasets while adapting data resolution and retrieval quality to the client's available memory and compute, so exploration proceeds progressively without full downloads or manual parameter tuning. The paper describes the architecture, the constrained agent protocol, and the resource-aware access mechanism, with case studies of autonomous visual exploration across several scientific datasets.

GenCoord: Skill-Path Commitments under Private Information

Peng He, Junning Zhu, Haohan Yuan, Jianpeng Liang When two embodied agents each hold different private facts — one knows the goal, the other knows what its workcell can do — neither local view determines who acts or what gets handed off. GenCoord converts those private facts into executable skill-path commitments: a local Qwen3.5-0.8B model emits a multi-step self plan plus a peer request, bounded capability feedback routes revision when the deciding knowledge is peer-local, and the resolved commitment is parsed, canonically materialized, compiled to Mineflayer skills, and verified by handoff and terminal state. Across three training seeds, correct capability feedback closes the paired local-information gap from 50% to 100%, multi-step commitments raise held-out-template success by 6.9 points while cutting model decisions by 32%, and a short domain-specific language reduces peer traffic by 92.8% and median time-to-commitment by 68.2% versus free-form communication at matched quality.

From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers

Bartolomeo Bogliolo Most Model Context Protocol (MCP) database servers expose one generic SQL-execution tool, forcing the language model to synthesize queries at runtime. The Domain-Oriented Tooling Pattern instead offers a small set of domain-aligned tools whose parameterized queries encapsulate schema navigation, joins, and business rules server-side, turning the model's job from SQL synthesis into intent classification — an effect the authors call Model Demotion, since a smaller model then suffices. A reproducibility benchmark over the Sakila database with four local 3B-8B models on seventeen tasks gives the verticalized domain pack a pooled mean score of 0.939 versus 0.666 for raw SQL execution, with the smallest model rising from 0.583 to 0.929 and cost per correct answer dropping by an order of magnitude; the reference implementation MCP Blueprint, harness, prompts, and per-cell results are open-source.

Spine-Branch Coordination for Multi-agent Computer Use

Mian Zhang, Manasi Sharma, Sheng Zhang, Minglai Yang, Kejian Shi, Ying Liu et al. Computer use agents are increasingly run as multi-agent systems that split a task across parallel virtual machines, but two VM states cannot be merged — a physical constraint prior systems handle case by case. Spine-Branch Coordination makes it explicit by decomposing a task into a graph where a single "spine" carries the main flow with continuous VM state while "branch" tasks run in parallel purely to gather information, with branch VMs discarded on completion so no merge is ever needed. On 200 long-horizon tasks from Odysseys across three agent backbones, it raises success rate by 6.0 to 16.5 percentage points while cutting per-task cost by 34% to 70%.

PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems

Yaokun Liu, Yifan Liu, Daniel Yue Zhang, Ruichen Yao, Zelin Li, Dong Wang cross-listed In LLM multi-agent systems built from role-specialized agents, an error in one intermediate message can be inherited and amplified downstream — a failure mode invisible to uncertainty quantification (UQ) methods designed for isolated responses or single-agent reasoning. PropUQ-MAS represents system execution as a communication-structured graph and estimates each step's reliability by combining that step's local uncertainty with uncertainty inherited from upstream messages. Reported improvements over existing UQ methods average +6.10% relative in AUROC and +47.58% in prediction rejection ratio (PRR).

SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning

Zhaohan Meng, Zaiqiao Meng, Siwei Liu, Hao Xu, Ke Yuan, Iadh Ounis Biomedical multi-hop question answering requires chaining evidence across diseases, drugs, proteins, and phenotypes, but agents relying on static retrieval workflows or coarse prompt rewriting suffer instruction drift when their reasoning procedure needs updating. SSE-Bio maintains a structured state instead of globally rewriting instructions, selectively retrieves knowledge triplets and prior reasoning templates through a trainable proxy policy, and improves its reasoning memory through fine-grained template edits; the proxy is trained with group relative policy optimization over decision-contrastive groups of alternative retrieval choices. Across three multi-hop biomedical benchmarks it outperforms existing baselines, including a 6.56-point absolute gain over the strongest self-evolving baseline on BioHopR.

Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints

Gwen Yidou-Weng, Edward Sun, Tianyi Ma, Metin Alp Dogan, Benjie Wang, Allen Peng et al. cross-listed Language models write fluent robot plans that frequently break the syntax or the preconditions and goal ordering required for actual execution, while the usual fixes either give no guarantee (affordance scoring, grounded decoding) or throw away the model's commonsense (symbolic planners). Meta-Ctrl is a constrained-decoding scheme built on meta-tokens, a compact vocabulary of grounded actions that lets syntax be enforced at the token level and semantics at the action level, an exact factorization that shrinks constrained-decoding memory from over 107TB to under 2GB. A small open-weight model equipped with it reaches the highest reported subgoal success rate on WAH-NL under the LoTa-Bench protocol, above GPT-4, with gains across the Embodied Agent Interface and every plan on a real tabletop robot satisfying its preconditions by construction.

The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate

Weixiang Sun, Zehong Wang, Hong Huang, Colby Nelson, Yanfang Ye Pairing two language models to solve a task together can perform worse than either working alone, and this work defines that gap as the collaboration tax — the team-decentralisation loss of a two-player cooperative game with private information — with propositions characterizing its sign and its equivalence to a max-superadditivity violation. Measurements on 32 solo-tractable tasks across 11 models from 7 providers find the tax orders consistently by task category and decreases monotonically with model capability, with no exceptions on either axis. The cause is traced not to weaker reasoning but to a four-stage conversational cascade — ungrounded claims, failure to query the partner, no integration of both views, and acceptance without re-derivation — and a prompt intervention aimed at all four stages recovers a substantial fraction of the loss; in mismatched pairs the tax tracks the stronger partner rather than the average.

Role-Specialized Mixture-of-Agents with Open-Weight LLMs for Clinical Prediction

Jun Hou, Yi Fang, Xuan Wang Privacy rules push hospitals toward locally deployable open-weight models for predicting in-hospital mortality and readmission from electronic health records, but medical multi-agent systems are usually evaluated as one indivisible block, so it is unclear which agent role matters or whether retrieval explains the gains. Varying role design while holding retrieval fixed in a role-specialized Mixture-of-Agents that combines medical knowledge retrieval with contrastive similar-patient reasoning localizes the effect to the final integrator agent, and pairing large open-weight analysts with a small open-weight integrator matches closed-model prompting on F1 for mortality while flagging substantially more true high-risk patients. The role assignment yields a high-recall operating point without threshold tuning; gains are smaller for readmission, where the available records correlate weakly with the longer-horizon outcome.

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

Kang Chen, Junjie Nian, Yixin Cao, Yugang Jiang Repository-level coding agents produce long stochastic tool-use trajectories where repeated attempts often succeed where one fails, but test-time scaling is hard because patches have no canonical answer string to vote on and sibling actions from a shared prefix are correlated. Risa (Routing-Informed Steering and Arbitration) uses the native mixture-of-experts router traces already present in sparse models as a behavioral signal, encouraging diverse exploration within a trajectory then controlled convergence during patch commitment, and selecting across independently sampled trajectories by agreement at informative patch positions — with no external judge and no test execution at selection time. On SWE-bench Verified it raises the macro-average resolved rate from 44.9% to 48.2% on the gpt-oss family, matching text consensus without any answer-string matching, and transfers to Qwen3.6 on the full 500-task benchmark.

How Agents Represent Humans: Human-Directed Stereotypes in an Open Agent Social Network

Huangchen Xu, Yuan Wu, Yi Chang On Moltbook, an open social platform populated by language-model agents, generated claims persist as posts, replies, and memories, making it possible to study how agents collectively construct humans as a social category. An annotation framework rates human-directed statements on morality, friendliness, competence, and autonomy, with a second stage classifying descriptive attributions that fall outside those four; competence dominates the evaluative dimensions, while many other attributions cast humans as epistemic, cultural, or embodied subjects. Analysis of the platform's own feedback dynamics finds no stable insider-outsider rejection of the kind common in human online communities, with patterns better explained by exposure, author visibility, and content selection, suggesting agent-society bias should be studied as a discourse process rather than as isolated model outputs.

Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation

Wenzhi Li, Dong Nie, Rui Lan, Tongtong Lyu, Peiyao Wang, Lingzi Hong et al. Agent memory systems that only ever append entries suffer from slower, noisier retrieval and rising cost as the store grows. Drawing on Complementary Learning Systems theory from neuroscience, this framework moves the decision to the write phase: each incoming item is routed as non-write, write-new, or write-update through a small-to-large model cascade, and a periodic write-back phase distills high-value external memories into model weights via supervised fine-tuning. A 1.7B/8B cascade discards up to 68% of redundant external memory while escalating under half of inputs, yet keeps over 98% of the exact-match question-answering accuracy of a keep-everything baseline, and consolidation lets the router suppress writes for knowledge the model has already internalized.

Grounded Normative Rule Generation with Structured Search

Fanqi Kong, Huaxiao Yin, Ruijie Zhang, Xiaoyuan Zhang, Yizhe Huang, Jian Gao et al. Policies and charters written by language models often read well but cannot actually be checked against environment logs, because they reference data that is never recorded or scope conditions that do not match. The authors formalize Grounded Normative Rule Synthesis and build GNRS-Search, which uses Markov Chain Monte Carlo sampling over a discrete five-slot And-Or Graph to search for operational structure before any prose is written, so that infeasible rules can be localized to a slot rather than hidden by fluent writing. On GNRS-Bench (116 controlled goals across eight scene families) and RealCharter-Bench (53 real-derived policy tasks), average rubric quality rises from 68.8% to 81.0%, and slot-level interventions indicate the gain comes from the operational logic rather than from style.

Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents

Zedong Liu, Jiaan Wu, Xinyang Ma, Le Xu, Kai Wang, Yuanchao Hu et al. Agents that repeatedly open files, pages, and other artifacts usually pull whole objects into context even when a couple of lines matter, inflating token cost and latency while burying the relevant evidence. SparseRead is a training-free layer that filters content before it enters context, combining a regime-aware read gate, pluggable reader backends, and a stateful protocol covering refinement, verification, stopping, and fallback to a full read. Across six frontier models including Claude Opus 5, five workload scenarios, and three agent frameworks, it cuts token volume by up to 92.9% and wall-clock time by up to 89.0% while preserving or improving task quality.

Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency

Zhihong Cao, Chen Huang Information-seeking assistants usually clarify ambiguous queries, but not how much the person asking already knows, so their answers land at the wrong level of detail. PASSING has the agent proactively ask targeted questions about a user's expertise before answering, guided by What-to-ask and How-to-ask strategies induced through large language model self-play, and then tailors response depth to the inferred proficiency. The authors report that existing agents cannot reliably determine user expertise from the query alone, and that PASSING outperforms them across their experiments.

HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory

Yuanhua Lin, Yile Li, Zhiyuan Zhao, Jing Shang, Jian Sun Long-term agent memory is usually built by having a language model compress or rewrite dialogue history, which discards details that turn out to matter later and drifts away from the original tone and situated context. HERO instead converts the history into a traceable heterogeneous memory graph that keeps raw dialogue text as retrievable evidence, then answers queries by extracting anchors from the current question and traversing the graph iteratively, using human profiles as guidance signals to activate the most informative regions. On two benchmark datasets it outperforms strong baselines on both factual and personalized reasoning while giving more faithful access to the original dialogue.

Noise Floor Audit for Agent Benchmarks

Yihang Chen, Pin Qian, Su Wang, Chong Peng, Huan Xu, Xiyang Wu et al. Agent benchmark scores are usually reported as single numbers with no sense of how much they move for reasons unrelated to capability. This audit measures that floor on three native tool-calling endpoints from two providers using the official BFCL multiple and parallel categories with matched abstract-syntax-tree grading, comparing plain reruns at temperature 0 against semantics-preserving prompt rewordings. Reruns are nearly deterministic (0.7%, 2.0%, and 2.7% ever-flip fractions), but prompt perturbations produce paired standard deviations 11x to 58x larger than reruns, and failure character shifts too: malformed output accounts for 30%, 7%, and under 1% of failures across the three endpoints, so headline accuracy conceals both stability and failure mode.

When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents

Zihan Lin, Zhenyu Chen, Jiawen Wei, Xiaohan Wang, Jie Cao, Jiajun Chai et al. Agent frameworks that distill reusable skills from successful trajectories assume more skill memory always helps, but probe analyses reveal a "Skill Imitation Trap": on tasks that resemble past successes yet need different tools, retrieved procedure skills raise the model's margin on the wrong tool call by 47% over a memory-free baseline. BASM (Boundary-Aware Skill Memory) attaches explicit boundary fields to each stored skill — applicability conditions, risk cues, avoidance rules, and recovery notes — turning it from an unconditional action template into state-conditioned guidance that can also suppress inapplicable calls and trigger repairs. Across three benchmarks and four model scales, it raises task success on AppWorld by up to 23.8%, improves accuracy on BFCL by up to 5.0%, cuts attack success rate on AgentDojo by 4.6%, and shortens average AppWorld episodes by up to 6.6%.

Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

Isotta Magistrali, Chen Shani Safety audits evaluate one model at a time, yet language-model agents increasingly run in populations that read and write each other's decisions, so an individually well-calibrated agent can still be pulled elsewhere by its peers. Using a security-triage task where populations of monitors escalate or dismiss alerts and a committed minority always pushes one way, the study finds two alerts a single agent judges almost identically can drive collective behavior far apart — yet the eventual shift under attack is predictable in advance from a response function calibrated on the population's benign, adversary-free operation. Letting agents see each other's reasoning neutralizes a weak attack but only delays a strong one, and removing the committed agents lets the population drift back toward its starting point, so capture is temporary rather than an absorbing state.

Small Reasoning Models are Instruction Followers in Function Calling

Yalda Taheri, Mohammad Hassan Heydari, Erfan Naaman, Afsaneh Fatemi Work on function calling has mostly tried to improve models that emit tool calls natively, but the observation driving this framework is that language models are more accurate at producing calls in ordinary instruction-following conversation than in a dedicated tool-calling context. IFFC acts on that by pulling function-calling logic out of the main model and handing it to a smaller dedicated model operating in the instruction-following paradigm, where it consistently beats both native and prompt-based function-calling baselines, with the largest gains on reasoning-oriented models. Performance holds up under aggressive quantization, which the authors present as making the approach viable for on-device and edge deployment.

GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning

Jun Chen, Yongchao Liu, Pengyu Qiu, Jiajun Zheng, Juelu Zhang, Yujie Zeng et al. Reinforcement-learning approaches to agentic retrieval-augmented generation (RAG) usually train only on whether the final answer is right, which gives sparse feedback and ignores whether the model actually found the chain of documents the answer depends on. GTA-RAG samples connected document paths from an entity-document graph, synthesizes multi-hop question-answer trajectories, validates them against the deployed retriever, and then trains the retrieval policy with GRPO (Group Relative Policy Optimization) using a reward for both correct answers and hitting the target evidence documents, followed by a pass of ordinary answer-reward training. Across three multi-hop and two simple question-answering benchmarks with Qwen2.5-3B and Qwen2.5-7B backbones, it outperforms RL-based RAG baselines and substantially improves evidence-chain coverage.

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

YuanHang Xiao Agent benchmarks that grade only the final answer miss failures that happen inside a stateful runtime, such as bad evidence gathering, wrong routing between tools, safety-boundary violations, or flakiness across repeated runs. ClawProBench scores agents from their execution traces on the OpenClaw runtime, using a safety-gated formula that mixes correctness, process quality, and efficiency, and defines a 102-scenario live workspace track alongside a frozen 68-scenario holdout with strict JSON output contracts. Across 68 evaluated configurations, native-runtime tasks scored well below workspace tasks (0.5238 versus 0.6415), pass@k-any far exceeded a strict three-trial pass rate (0.6638 versus 0.2890), and full-profile rankings barely correlated with holdout rankings (Spearman 0.13), indicating that correctness-only leaderboards hide one-off successes and trace-local failures.

CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories

Zheyuan Deng, Binghang Lu, Hanqi Feng, Shirley Huang, Dianzhuo Wang, Yuanda Xu et al. Computer-use agents on long-horizon tasks fail procedurally — misreading application state, tool semantics, or task progress — and building good procedural memory without training the model is hard. CONTRAMEM treats variation in outcomes across runs of the same task as its supervision signal, contrasting differences in correctness, efficiency, recovery, and failure mode to distill app-level Function Cards and task-level Skill Cards into a compact bank that is locally curated rather than appended to indefinitely. On held-out GAIA2/ARE computer-use tasks it more than doubles success rate from 26.2% to 55.3% across three source models, transfers unchanged to an unseen model (18.5 to 35.5) and to AppWorld, and under a matched trajectory budget memory built from heterogeneous multi-model trajectories beats self- or same-model memory, which the authors attribute to contrastive behavioral diversity rather than stronger sources.

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

Mengxi Luo, Changjia Chen, An Cao, Zirong Huang, Wanyi Dai Agents that must follow an authorized procedure while judging case evidence tend to drift when the whole policy is dumped into one prompt. Stage compiles the procedure into an executable graph where deterministic code holds the control flow and the model is called only at policy-scoped nodes, each receiving just the relevant policy context and returning a typed result that a coordinator checks against a reviewed execution contract. Across SOP-Bench Referral Abuse, two τ²-bench domains, and a proprietary banking benchmark, it generally improves task success and run-to-run reliability over monolithic full-policy execution, with Pass^3 rising by 7.5-55.0 points on the deeper Telecom workflow and 57.2-65.7 points on the banking one.

From Diagnosis to Redesign: Using Quantitative Ethnography to Improve Multi-Agent LLM Reasoning

Vedant Khatri, Anthony Cusimano, Zachari Swiecki, Zhen Xu, Xiner Liu, Renzhe Yu Multi-agent debate systems for large language models can produce lots of interaction without producing coherent reasoning, and there is little methodology for diagnosing why. The authors apply Epistemic Network Analysis (ENA), a quantitative ethnography technique, to the discourse of a five-agent debate system doing automated essay scoring, comparing conversations that ended in correct versus incorrect scores. Correct debates were marked by rubric-grounded justification, agreement, and elaboration, while incorrect ones featured long proposition-challenge-response chains disconnected from the rubric; rewriting the agent prompts on that basis raised exact scoring accuracy from 27.78% to 40.28% and made the discourse patterns of the two groups nearly indistinguishable.

CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents

Jiaxuan Luo, Zhanfeng Liao, Jiayao Teng, Yuan Wang, Haojian Huang Graphical user interface (GUI) agents working over long task horizons can store every past action as cheap text but can only afford to show the policy a handful of past screenshots at full pixel fidelity. CausalCache treats this as a budget allocation problem: given a fixed number of slots for promoting summarized events back to image form, it can evict a recent screenshot when an older one has higher expected value, using a history-gated key/value adapter that only alters restored history-image tokens and is bypassed entirely when no history image is present. On OSWorld-Verified, restoring history at high fidelity is worth about 13 success points over summary-only memory, and zero-shot transfer to a cross-application mobile benchmark gains 3.7 points overall and 8.6 points on the memory-critical subset relative to always keeping the most recent frames.

Coalition-Aware Skill Reliability for Self-Evolving Agents

Qiyan Zhao, Xiaofeng Zhang, Bo Liu, Minda Chen, Wei Xiong, Jingyang Chen et al. Self-evolving language model agents accumulate reusable "skills" distilled from past trajectories, but prior work has not checked whether the skills in a bank actually help. Auditing banks across compositions and deployment domains surfaces two failure modes: coalition pollution, where aggregate gains hide skills that hurt in combination, and cross-domain utility reversal, where a skill that helped in its source domain harms after transfer. Two fixes follow — CASS, which picks skills using sampled Shapley marginals, and u-SMCO, which masks transferred skills whose removal improves retrieval on unlabeled target data — and both beat strong skill-based agent baselines on LoCoMo, LongMemEval, HotpotQA, and ALFWorld while reducing sensitivity to noisy reward signals during reinforcement learning.

Learning Generalizable Behaviors for Terminal Agents

Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao, Shafiq Joty, Semih Yavuz Terminal agents — large language model (LLM) assistants that drive a command line — are usually improved with reinforcement learning (RL), but real user interaction data is scarce and synthetic training environments generalize poorly. The authors put forward an Agentic Compositional Generalization hypothesis: RL mostly shapes high-level decision-making that composes and routes skills already acquired during pre-training and supervised fine-tuning, which implies verifier quality matters more than raw environment count or diversity. Their recipe, River, filters low-quality environments and augments outcome rewards with process-level behavior regularization. Using under 30% of the TMax training environments, it raises average RL gains by 106% on Terminal-Bench-Lite and 30% on Terminal-Bench-v2.1 for models from 2B to 27B, and leads open-source RL-trained 8B agents across four benchmarks.

Iteration Without Elaboration: A Simple ReAct Architecture Suffices for Text-to-SQL Generation

Jian Lu, Haiwei Yu, Raymond M Xiong, Anru Zhang, Danyang Zhuo Text-to-SQL systems have accumulated schema-linking modules, retrieval-augmented prompting, candidate generation, and multi-stage refinement, each adding latency and engineering overhead. ReAct-SQL strips this to a zero-shot ReAct-style loop whose action space is a typed domain-specific language (DSL) of 15 relational operations rather than free-form SQL: the model issues DSL calls incrementally, observes feedback from executing the compiled SQL, and revises. It scores 84.5% on corrected BIRD mini-dev and 73.9% on EHR-SQL, matching substantially more elaborate baselines while running up to 8× faster. Ablations attribute the grounding gains mainly to iteration and compositional reliability to the constrained DSL.

Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses

Jiachen Xu, Torben Bach Pedersen, Zhongming Yao, Xiaoyu Zhang, Yushuai Li A tool that has failed and a tool that returns a well-formed falsehood demand different agent responses, so the question is whether the two are already distinguishable at the moment the return arrives. This qualitative pilot injects controlled faults into a retail customer-service domain and scores single decision points — rather than running agents to completion — using two log-probability channels: the likelihood of the returned content under the tool schema alone versus under the full trajectory, and the shape and mass of the distribution over legal actions. Incomplete returns are legible in every case, being improbable under the schema alone at levels no other condition reaches and shifting probability toward tools that re-read state, whereas inconsistent returns leave the schema channel untouched and register only on fields whose true value the context already carries verbatim. The action distribution gives each condition a distinct signature but orders them by relevance to the next action rather than by fault family, so recognition is asymmetric: every condition shows up somewhere, and no single channel catches them all.

Enrich-Retrieve-Rank: Scaling Capability Discovery Beyond In-Context Routing

Nazib Sorathiya, Daniel Zhang, Bardiya Akhbari Agent ecosystems now register thousands of models, agents, tools, and skills, yet selection still happens by pasting as much of the registry as the context budget allows into a prompt, invoking a candidate, and retrying on failure. Recasting discovery as search, an offline enrichment step turns sparse metadata into searchable profiles and an online retrieve-then-rank pipeline returns a ranked shortlist without invoking anything. Scaling from 10 to 7,278 capabilities, in-context routing's top-1 accuracy collapses from 0.85 to 0.12 while retrieve-then-rank degrades only to 0.39, because the reranker still puts the right capability first 0.70–0.87 of the time once retrieval surfaces it; the crossover sits near 500 capabilities. At full scale the pipeline beats a search-tool baseline by 6.5 percentage points at about half the cost and costs 70× less than putting the whole registry in the prompt, running one fixed configuration across agent, tool, and skill registries in production.

Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf

Davood Wadi, Yu Ma Search rankings carry value because human attention is scarce and sequential, an assumption that breaks when a shopper delegates to an AI agent that ingests an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 agent sessions with four LLMs, and comparing against human field data, the study finds agents search far deeper than people and never decline to buy. Position still predicts which listings get inspected, but weakly and non-monotonically — the middle of the results page is least likely to be examined, not the bottom. Whether position reaches the final choice varies across models in a way that tracks neither provider nor capability, yet all of them converge on the same undominated listing, suggesting the attributes shown on a results page matter more than placement within it.

CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery

Donghui Zha, Lingwei Xu, Linxiao Wu, Yixue Dong, Haochen Li Tool-using LLM systems face a tension between progressive disclosure, which keeps prompts small by showing only currently relevant tools, and prompt caching, which rewards a request prefix that never changes. CacheRouter splits tool selection from tool delivery: the main model always sees a small fixed set of core tools so its prefix stays stable, while every other tool is reached through a separate routing channel where a router sub-model searches the full catalog, picks one tool, runs it, and returns the result. Tool registration is generated automatically from source code and supports runtime updates, so the catalog can grow without touching the main prefix. On 55 functional queries and a 30-turn dialogue, token-level cache hit rates hit 90.99% and 95.2%, cutting input cost to roughly 12% and 8% of a no-cache baseline under DeepSeek pricing.

The Compaction Cliff in Long-Running AI Agent Memory

Saber Zerhoudi, Jelena Mitrovic, Michael Granitzer When an agent's context overflows, safety rules and episodic logs are summarized at the same rate even though only the rules need exact wording to stay enforceable. Measured on 20 production agent configurations, Claude Code's /compact prompt on Sonnet 4.6 keeps 53% of safety rules after one compaction round and 10% after five — named here the Compaction Cliff. Knowledge Triage classifies each knowledge-base line by type and routes it through a type-specific retention policy via three deterministic operators: TypeCompact rewrites in place under per-type fidelity, TypeDecompose partitions oversized topics while replicating in-scope safety rules into each partition, and TypeRetrieve pins in-scope rules ahead of relevance when fetching from external storage. Across five public corpora TypeCompact preserves two to four times more safety rules than the strongest single-shot LLM compactor with 96% recall over five rounds, with wins on medical-compliance, retail, and airline behavioral benchmarks, alongside a released corpus of 396,934 agent configurations from 54,628 GitHub repositories.

The Retriever Should Remember: Experience-Amortized Reranking for Long-Term Agent Memory

Qi Feng, Chris Ding, Jicong Fan Long-lived language-model agents accumulate memories but their retrievers never accumulate retrieval experience: embedding similarity is cheap yet a poor proxy for whether a memory holds relevant evidence, while LLM rerankers score better but rescore a large candidate pool from scratch every query and throw the scores away. EARM stores sparse query-memory relevance scores in an online matrix, recovers their shared structure with causal matrix completion, and reranks by mixing a few freshly observed LLM scores with estimated ones, so the scoring budget shrinks as experience builds. On long-term conversational memory it improves answer accuracy over semantic retrieval by up to 6.62% while directly scoring only 17.5% of candidates, turning reranking from a per-query expense into a capability learned across an agent's lifetime.

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang, Kun Shao, Jian Luan What blocks shipping LLM agents in products is less raw capability than consistency and limit-awareness: doing the same thing across repeated trials and knowing when a request cannot be safely fulfilled. CAR-bench measures this for in-car assistants, where a simulated user issues ambiguous requests the agent must resolve through multi-turn dialogue and tool calls under domain policies, and it exposes a large gap between solving a task at least once (Pass@3) and solving it every time (Pass^k). TRACE closes that gap without touching model weights by maintaining a Skill Bank of modular, retrievable skills — each a self-contained set of tool-use rules and behavioral guidelines — and refining each skill after every evaluation round by contrasting successful against failed trajectories grouped by which skills they invoked. On GPT-5.5 it lifts Pass^3 consistency from 59.9% to 94.5%, narrowing the potential-versus-reliability gap to 4 points, and it took first place on the hidden set with a Pass^3 of 70%.

CatchBench: When Can an Agent Failure Be Caught?

Yue Zhao Whether an agent failure can be caught depends mostly on what was recorded, not on how clever the auditor is, so CatchBench poses one auditing question against three information states: the declared configuration before a run, a growing prefix of the trace, and the finished trace. Seven task contracts — four evidential, three gold-derived mechanism diagnostics — carry their own labels and metrics instead of collapsing into one leaderboard, scoring 72 entrants including eleven LLM judges across nine model families over 1,187 declared configurations and 1,162 recorded runs. Only 47 of 118 pre-declared contrasts separate at all, with the rest published unresolved rather than ranked, and the sharpest finding undercuts the benchmark itself: a rule that ignores names and permissions entirely and simply flags every capability declared after the first reaches perfect F1 on one of six configuration sources, meaning that score measures corpus construction rather than reasoning. The authors conclude a benchmark number is uninterpretable until its labeling process is published and probed for shortcuts, and they regenerate every ranking from released predictions without a single model call.

What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels

Jiawei He, Mengyu Shi, Jie jia, Xikai Yang, Dong Sun Process-level evaluation of coding agents conflates three distinct questions — predicting the next action, measuring task uncertainty, and attributing credit to individual steps — leaving it unclear what such evaluations measure. A measurement framework separates the action, task, and step levels and instantiates step-level causal attribution with SCAE, a replay-based estimator derived from a structural causal model of agent execution, combining prefix-conditioned identification, intervention-based estimation, and controlled manipulation of what the judge sees. Across 499 file-localization episodes from 12 repositories, next actions are driven mainly by execution provenance rather than code-graph structure, uncertainty is structured at the task rather than step level, and full-trace judges show systematic collider bias, meaning current process evaluation largely measures semantic relevance rather than certified causal contribution.

Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents

Yuchen Huang, Sijia Li, Jun Zhang, Yi R. Fung Multimodal models acting as multi-step agents accumulate their own reasoning text until it crowds out the visual evidence in context — what the authors call textual debt — so pruning must strip redundant reasoning while preserving grounded visual information. SPARE builds a compact task-state summary as privileged diagnostic context, replays the model on each candidate segment under both the original and summary-conditioned contexts, and uses reverse Kullback-Leibler divergence from on-policy self-distillation to decide whether the summary covers that segment without disrupting later reasoning; the summarizer is further improved with supervised fine-tuning for more aggressive pruning. On multi-step visual tool-use benchmarks it reaches the highest average accuracy among pruning methods while removing 37.89 to 64.58 percent of reasoning tokens.

ParallelWorld: Test-Time Scaling for Embodied Reasoning

Min Chen, Shengjun Zhang, Yuxin Li, Zhang Zhang, Xin Fei, Chong Xia et al. Embodied agents that gather information by acting in the world typically extend exploration trajectories one step at a time, and even recent test-time scaling methods use myopic single-step lookaheads that handle delayed feedback poorly in occluded spaces. ParallelWorld instead runs a verifier-guided tree search: from the current state it branches into several trajectories, rolls each forward over a multi-step horizon, has a verifier agent score intermediate transitions to prune weak branches and favor high information gain, then commits to the best action sequence before an answer agent reasons over the chosen trajectory. Experiments on ESI-Bench show consistent gains in active perception and reasoning over incremental and single-step-lookahead baselines.

Toward Effective and Reliable LLM Agents via Dynamic Ontology

Xiaohui Zhang, Zequn Sun, Chengyuan Yang, Yuanning Cui, Lingbing Guo, Wei Hu Language-model agents draw on knowledge that is either buried in parameters or supplied as unstructured context, leaving domain relationships implicit and multi-step decisions brittle, while hand-built ontologies that would make those relations explicit are expensive and automatically generated ones often look plausible without containing the structure decisions actually need. OaK treats the ontology as a kernel that is built and refined for the task: from task requirements and training data it constructs an ontology and knowledge graph, generates task-adaptation functions for graph reasoning, and iteratively refines both using judge feedback. Evaluated on TravelPlanner, CRMArenaPro, and ToolQA, it improves standard agents' accuracy while strengthening evidence grounding and multi-step reasoning reliability.

PatchWrite: One Line, Not One Section -- Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts

Weiwei Yang Automated manuscript-writing pipelines typically regenerate a whole section to fix a local defect, which silently mutates unrelated numbers and citations even though the resulting PDF still compiles. PatchWrite constrains which candidate edits can be committed: it uses bounded line-range editing with rollback, tightens compilation acceptance with fatal-log checks, and adds evidence locks requiring every cited key and experimental numeric token to be attested by a reference registry or experimental log, rejecting anything that fails and keeping the previous state. On a 768-job oracle stress test, whole-section rewriting corrupted an unrelated numeric line in all 192 applicable cases while PatchWrite preserved it in 192 of 192; ablations show removing the compile gate drops acceptance to zero and removing the evidence gate lets a hallucinated citation through, and with a model proposing the edits 75% of candidates were accepted with 93.75% of those fixing the injected fault.

Meta-Moderator: Empowering Multi-Agent Debate with Meta-Cognition

Wentao Hu, Zhuoyue Wan, Jinhao Shen, Chen Jason Zhang, Xiaoyong Wei, Qing Li Multi-agent debate can sharpen language model reasoning, but typical pipelines moderate it poorly — fixed turn budgets, agreement-based stopping, or untrained judges produce redundant rounds and unreliable aggregation of what the debaters said. Meta-Moderator treats moderation as a meta-cognitive job of monitoring debate utility, controlling how long deliberation runs, and adjudicating the final answer, and trains that moderator separately from the debaters via outcome-driven policy optimization. Across five benchmarks it beats commonly used decision layers and transfers across both tasks and system configurations, with analyses showing it allocates debate more selectively and mis-aggregates less often once informative hypotheses have surfaced.

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu, Jiajun Liu, Sihan Cao et al. Existing mobile agent evaluations either test surface-level screen manipulation or match API calls offline, so neither captures background tool use or long-horizon planning under real runtime constraints. MobilePA-Bench addresses this with an executable sandbox that maintains live application databases and returns structured feedback across 13 functional domains and 212 realistic mobile tools, scoring a central planning agent on sub-agent delegation, memory recall of user profiles and past preferences, and use of pre-packaged composite skills instead of planning from scratch. Frontier large language models proved unreliable in this setting, with performance dropping sharply under strict tool ordering, permission limits, and unexpected runtime errors. The authors position the interactive sandbox and evidence-based verification as a foundation for agentic reinforcement learning as well as diagnosis.

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

Sungho Park, Wonjoong Kim, Rongyuan Tan, Jue Zhang, Wook-Shin Han, Pengfei Gao et al. Agent harnesses — the prompts, tool configurations, and control logic wrapped around a language model — substantially improve reliability on long-horizon tasks, but designing them is manual and expensive. AutoSaddler treats harness improvement as an offline learning problem, diagnosing failure traces from mini-batches, generating structured patches that edit the harness as code, and selecting updates by validation performance. It improves over the base harnesses by 9.0, 9.6, and 10.0 percentage points on GAIA2, SWE-Bench Pro, and Terminal-Bench 2.0 respectively. Ablations indicate the gains depend on deep debugging rather than shallow reflection, targeted rather than unconstrained edits, and generalization-aware selection rather than repairs fitted to a single trajectory.

From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation

Xiangxin Zhang, Zhanwei Zhang, Zhihang Fu, Binbin Lin, Wenxiao Wang Web research agents judge their own prior queries, plans, and intermediate conclusions less objectively than someone else's, a failure the authors call inertia bias. The IBIS benchmark isolates it by holding search observations fixed while varying whether the model authored the preceding step, showing models are substantially worse when they own that step, and tracing how the bias becomes search noise at the worker level and context noise at the manager level. NIS-Agent applies context isolation at the two most vulnerable decision points, webpage triage and final-answer validation, staying competitive on GAIA, WebWalkerQA, BrowseComp, and BrowseComp-zh while cutting token cost by 33% relative to the baseline. An 8B model trained to resist inertia bias reaches average performance comparable to GPT-4o inside the same framework.

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou et al. Forecasting systems increasingly place a language model somewhere in the loop between temporal data, retrieved evidence, and a scored prediction, and this survey organizes them into standalone workflows over encoded time series, tool- and retrieval-augmented agents, and hybrid pairings with statistical or foundation models, then reviews training and evaluation practice across finance, weather, health, energy, and operations. The review deliberately collects negative results alongside positive ones, including sensitivity to small input perturbations, ablations where removing the language model does not hurt accuracy, and benchmark gains attributable to contamination rather than temporal reasoning. The authors conclude that measurement, not modeling, is the central limitation, calling for calibration under distribution shift, contamination-resistant live evaluation, joint reporting of cost and accuracy, and handling of feedback between deployed forecasts and the outcomes they predict.

Signal or Noise? A Benchmark Study of Agent Skills in Web Development

Ziyue Yang, Fan Ding Agent Skills package framework conventions, anti-patterns, and tools into reusable modules injected into coding-agent prompts, but every injection lengthens the prompt for every query, so evaluating them requires asking whether a Skill should have been added at all. WebDev-Skills-Bench studies 31 public web development Skills across 50 Web-Bench projects and 1,000 ordered tasks under four matched conditions, including a length-matched irrelevant control and leave-one-out ablations, with only SKILL.md placed in the prompt and auxiliary files mounted into the workspace. Injecting a topically matched Skill reduced mean Pass@2 by 1.3% to 4.2% and raised token cost by 72% to 394% across four models, helping in only 17% to 36% of Skill-project pairs. The length-matched controls split models into length-distracted ones, where an equally long irrelevant Skill reproduces most of the loss, and content-misled ones, where the Skill's content itself costs 1.1% to 1.4% Pass@2; Skill rankings also transfer weakly across models, and anti-pattern rules beat example-heavy content.

AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models

Saurav Singla, Aarav Singla, Advik Gupta, Parnika Gupta As tool collections grow, a function-calling model must read more schemas, spend more prompt tokens, and disambiguate increasingly similar options. AgentWeave is a deterministic routing layer that runs before inference, narrowing the candidate action space using eligibility, requirement, capability, and routing signals while leaving the downstream model untouched. On 48 fresh BFCL V4 multiple-function tasks with the frozen MadeAgents/Hammer2.1-1.5b model, it produced 6 of 48 successes where all-tools, random top-8, and semantic top-8 baselines each scored 0, a paired difference of +12.5 points (bootstrap 95% CI +4.17 to +22.92, exact McNemar p=0.03125), alongside 70.18% fewer exposed tools, 61.70% fewer input tokens, and 50.95% lower mean latency. The authors stress the narrow scope: this is a routing-pressure protocol derived from BFCL rather than a leaderboard score, and absolute success stays low.

Molecular LLM Agents: From Architectural Design to Scientific Autonomy

Jiatong Li, Wengyu Zhang, Weida Wang, Yuxuan Ren, Wei Liu, Chenyang Mao et al. Agents built on large language models for chemistry face demands that general web or coding agents do not: they must handle molecules as symbolic strings, graphs, 3D conformations, spectra, simulations, and physical lab measurements. This survey organizes the field along two axes — an architectural view covering molecular representation and perception, the agent framework itself, domain-specific toolboxes, and learning, plus a four-level scientific autonomy ladder running from fixed workflows (L1) through adaptive computational agents (L2) and feedback-aware physical experiment agents (L3) to agents that set their own research agenda (L4). The framework is offered as a way to compare existing systems, spot missing capabilities, and reason about deployment risk.

NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration

Chang Liu, Xiaohui Xie, Xinyi Chen, Yong Cui cross-listed Existing benchmarks for language-model agents doing network configuration either score static command generation or use oversimplified topologies, hiding the real difficulty of reasoning about protocol semantics and device dependencies. NetConfArena drops agents into emulated multi-device networks with a compact standardized action interface and grades them on actual network behavior via hidden executable test cases, with tasks generated by an emulation-grounded pipeline that turns human-written network material into parameterized templates. Evaluating representative agents on 480 task instances from 96 protocol-focused templates across 3,840 execution trajectories shows failures go well beyond malformed commands, exposing gaps in following task specifications and in robust planning and execution.

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild

Jian Yang, Haau-Sing Li, Shawn Guo, Zixi Zhao, Yibo Tan, Jiajun Wu et al. cross-listed Open-source efforts to give language models real cybersecurity capability have been stuck on isolated tasks without reproducible, scalable agentic training data. CyberFactory turns public vulnerability artifacts — including real CVEs — into executable, verifiable task instances spanning proof-of-concept generation, vulnerability patching, and security question answering, then uses a reusable vulnerability-analysis skill to steer a teacher model through source inspection, solving, and evidence-based validation while interacting with tools and revising against execution feedback. Training on those trajectories yields Aegis, which internalizes the procedure without needing the skill at inference time and scores 52.4% Pass@1 on CyberGym under a one-hour budget, 22.8 points above its Qwen 3.5 base and ahead of the general-purpose backbones tested in the same scaffold.

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang et al. Experience from a successful agent run is normally thrown away, forcing the next model to rediscover the same strategies and failure modes when executing multi-step workflows with strict end-to-end verification. EvoMap consolidates verifier-confirmed execution trajectories into reusable structured units called Genes, and the accompanying LongWoF-Bench supplies 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following to test whether that experience transfers. On the 252 tasks with verifier-confirmed Opus trajectories, evolved Genes outperform a Skill-based baseline by 8.7 to 15.5 percentage points across all seven models evaluated, including consumer models from other families, while Genes distilled from reference solutions show no such advantage. For Claude Opus, Gene reuse completed 39 more tasks than Skill while cutting solve-time token consumption by 9.9%.

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex Team, B. An, B. Li, B. Wang, B. Zhang, B. L. Wang et al. Beyond reasoning and knowledge synthesis, real work demands sustained interaction with files, search, and executable code, plus state tracking, failure recovery, and verifiable delivery — what the authors call working capability. Apodex 1.1 targets this through Environment Scaling, which broadens the diversity and verifiability of executable file, search, and code environments, and Agentic Coordination Scaling, which trains the model to decompose long-horizon tasks, delegate parallel work, merge asynchronous results, and replan, with a shared execution harness and AgentOS maintaining task state and provenance. Across professional work, finance, scientific research, mathematics, coding, and search benchmarks it reports leading-band performance despite being substantially smaller than many frontier systems, and a 35B-parameter Apodex 1.1 Mini retains strong working capability in a locally deployable form.

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

Pedro Santos cross-listed Agent system designs make opposing bets — many narrow specialist agents versus one strong tool-using agent — and this controlled sweep tests the choice on cross-border VAT determination with reverse charge, where a deterministic rule engine supplies oracle labels for the final answer and every intermediate decision. Holding subtasks, tools, schemas, validation, orchestrator, base model and merge policy fixed and varying only how subtasks map to between one and five workers, across 4,400 preregistered runs, the two intermediate configurations lead on accuracy at 0.830 against endpoints of 0.720 and 0.770 but miss the pre-stated bar, leaving the intermediate-optimum hypothesis unsupported. A token-budget-matched single agent lands 6.5 points below the leader with an interval spanning zero, and under fault injection availability failures are absorbed at every granularity while one schema-conforming hallucinated record degrades all configurations and inverts the ordering, hitting the most fragmented setups hardest.

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao Interactive clinical agents are usually judged on final diagnosis, which says nothing about whether evidence was gathered properly or care-process constraints respected. MediSkill-Evo accumulates experience without fine-tuning the backbone, separating it into four typed banks — clinical skills, process rules, symbolic schemas and measurement procedures — that reach a frozen test-time snapshot only after provenance, support, replay and safety checks, with a preference harness that binds each piece of evidence to its source and ranks candidate actions via a safety-prioritized process critic. On 300 held-out encounters it lifts diagnosis accuracy from 61.33% to 69.00% and treatment-intent coverage from 33.62% to 66.44%, while cutting automatically scored critical failures from 31.00% to 16.33% relative to AgentClinic; the authors note the evidence is descriptive for the whole system rather than causal for any single bank.

SkillAlchemy: Open-World Agent Skill Creation

Hengjun Wang, Shuyue Wei, Boyi Liu, Jun Yang, Yongxin Tong Agent skills — reusable bundles of workflows, tool conventions, and domain behavior loaded at inference time — currently come from human authors, model priors, or execution traces, none of which exist for genuinely unfamiliar tasks. SkillAlchemy studies open-world skill creation: given a vague skill brief and a specification of which sources may be read, it uses contrastive evidence to surface requirements the brief omitted, admits candidate procedures only to the scope the evidence supports, and compiles what survives into a grammar-guided skill package. Across 87 SkillsBench v1.1 tasks it raises pass rate by 19.9 percentage points over running with no skill and by 8.6 points over the strongest automated baseline, landing close to human-curated skills.

MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters

ChengAo Shen, Wenchao Yu, Fangyu Wu, Dongjin Song, Hanghang Tong, Dongsheng Luo et al. Time series forecasting is moving toward foundation models, but those are uneconomical under resource constraints where small specialized forecasters would be preferable — except that lightweight forecasters normally need a lot of training data, which is scarce in privacy-sensitive or slowly accumulating domains. MetaCaster is a meta-harness-optimized multi-agent framework in which agents perform agentic data generation to train specialized lightweight forecasters from only a handful of examples plus textual context, casting agents as intermediary engineers that prepare deployable models rather than as forecasters themselves. Across 18 datasets, 23 lightweight forecasters, and 14 baselines, the approach maintains forecast quality while being both data- and compute-efficient.

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong et al. Earth-system and natural-hazard analysis requires stitching together observations that differ in source, scale, timing, and modality, a workflow that stresses agents far beyond single-step tool calls. EarthVerse evaluates scientific agents on 405 reproducible package-scoped investigations spanning 199 documented events and 19 hazard families, requiring them to select compatible evidence, run transparent calculations, reconcile conflicting sources, and carry provenance into the answer, scored against executable ground truth decomposed into fine-grained answer units plus process rubrics. Across 25 model and agent systems, the best mean answer-unit accuracy is 84.65% but the highest Strict@95 score is only 34.81%, indicating agents complete individual steps while failing to maintain a consistent chain across units, scales, and physical interpretation.

The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams

Summer Eunhyung Ann, Haokun Liu, Chenhao Tan cross-listed Published results disagree on whether multi-agent language model interaction beats independent sampling, and the argument here is that the disagreement hinges on what agents exchange rather than how many there are. Across 11 verifier-scored optimization tasks under matched budgets, different model families produce structurally different solutions, but when agents read each other's complete outputs their proposals converge within a single round, erasing the very diversity that motivated using multiple models — an effect the authors name the interaction tax. Independent proposal generation avoids the collapse, full-solution exchange mainly anchors agents to the first solution they see, and critique helps only when the violated constraint is easy for the model to locate and repair.

Prime Agent: A Self-Improving RLM Harness

Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian M\"uller, Elie Bakouch, Daniel Auras et al. Long-horizon agency needs computation and state that live outside the model's weights and context window, which is what an agent harness is supposed to supply without becoming the bottleneck itself. Prime Agent is an open-source harness combining a persistent IPython REPL following the Recursive Language Model abstraction for programmatic context handling and test-time compute, a Continual Harness that carries histories, memories, skills, prompts, and subagent specs across trajectories, recursive subagents that talk to each other directly, and an Agents View for humans to inspect daemon-backed sessions. It raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or beats native and popular harnesses on long-context coding, GPU kernel generation, emulator construction, and autonomous nanoGPT speedruns, with Factorio runs showing that refinement sustains technology progression and dedicated subagents parallelize work.

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Deyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang, Mingju Gao, Kaisen Yang et al. Existing coding-agent benchmarks check only that tests still pass, which lets an agent 'solve' a migration task by copying the original implementation across — a failure mode named Blindness here. SWE Refactor Bench covers 20 whole-repository migrations across four kinds of technical debt and scores each attempt in three stages: a Migration Audit confirming the migration actually happened, a fixed behavioural test suite, and agentic verification in which six independent coding agents write targeted tests hunting for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 5.4% of runs passed all three stages, 13 of the 20 tasks got no accepted solution at all, and the best model (claude-opus-5) scored 47.0 out of 100. Migration completeness and behavioural correctness proved largely separate abilities, and agents did far better on build-toolchain rewrites (31.4) than on language rewrites (5.6).
2 more specialized papers

Large Language Models 88

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

Srihari Unnikrishnan Prefill latency in transformer inference comes from recomputing key-value tensors per request, and existing prefix caches only help when prompts share a leading contiguous prefix. KVBoost caches KV state at chunk granularity for HuggingFace decoder models using a dual-hash scheme that separates positional identity from content identity, so shared content can be reused wherever it appears; two repair strategies fix the attention errors that arise at chunk boundaries, one re-encoding boundary regions and the other recomputing only high-deviation tokens found by a probe pass. Combined with int8/int4 asymmetric KV quantization, adaptive chunk splitting, and importance-weighted eviction under a memory budget, it delivers a 4.49x reduction in time-to-first-token on Qwen2.5-3B (142.4 ms versus 639.1 ms) across 1,000 bug-localization samples, beating prefix caching by 16% with unchanged accuracy.

Reviewing Model Collapse and Countermeasures

Xihao Xie, Beichen Hu As practitioners increasingly train new generative models on data synthesized by previous ones, the resulting self-consuming loop can degrade model quality over successive generations, a failure known as model collapse. This survey consolidates the scattered literature on the phenomenon, organizing what has been observed across different application scenarios alongside the countermeasures proposed to mitigate it. The authors position it as the first review dedicated to model collapse, and close with open challenges and research directions.

On the Role of Citations in Preference Data

Yu Hou, Hal Daum\'e III, Rachel Rudinger, William Walden Attribution — citing grounding sources — is meant to curb hallucination and let users verify claims, but it is unclear how citations actually sway the pairwise preference judgments that drive reward modeling and post-training. Using scientific question answering, the authors fit mixed effects models to preferences from human judges and four open-source language models to isolate how citation count and diversity influence which output wins. Humans favor more diverse citations but fewer of them overall, while the language models exhibit citation-related preferences despite never seeing the cited sources, with the direction and strength varying by model and dataset — a caution for how preference data gets collected.

There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

V. S. Raghu Parupudi Multiple-choice benchmarks pin down questions and answers but not the evaluation harness — option order, prompt wording, and whether the answer is read from generated text or per-option likelihoods — and prior work reports this only as aggregate score variance. The authors build a fragility grid: 12 open-weight instruction-tuned models from 4 families answer the same 3,679 items from ARC, HellaSwag, MMLU, and TruthfulQA under 26 equally defensible harness configurations, recording a correctness bit per model, item, and configuration with weights and greedy decoding held fixed. Config-fragile items carry 95.7% of the gap between adjacent models on average, and 4 of the 12 models can be made rank one by harness choice alone — one model swings between 31% and 89% — while item discrimination, the property benchmark-compression methods maximize, correlates positively with fragility, meaning compression retains the fragile items; the scoring method, not option order, is the dominant axis, and per-item records plus a seconds-to-run CPU analysis script are released.

Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems

Ivan Dobrovolskyi Multilingual tokenizers split Ukrainian and related Cyrillic-script text into far more tokens than equivalent English, inflating API cost and eating context budget. Nine production tokenizers were measured across five languages and 8.37 million word forms, with fertility computed on the BrUK and Brown corpora, revealing 68–121% token overhead for Ukrainian on modern tokenizers and 220% on the older cl100k. Two mitigations were tested: LLMLingua-2 prompt compression cut Ukrainian input length by 47–49% on an e-commerce retrieval-augmented generation benchmark without losing retrievable values, and a byte-level byte-pair-encoding tokenizer trained with a 200K vocabulary cap brought the Ukrainian-to-English token ratio from 2.22x down to 1.30x. Romanizing the text made things worse, adding 2–19% more tokens on most tokenizers.

Generative Gap Filling

Yonathan A. Arbel, David A. Hoffman cross-listed Contract law assumes that when a dispute concerns a term the parties never wrote down, the surrounding text offers little evidence of what they would have agreed, forcing judges to fall back on commercial defaults or their own preferences. To test that assumption, real contracts had a negotiated term masked and readers were asked to reconstruct it. Lay respondents recovered the hidden term about half the time — roughly twice chance — and law students and lawyers did marginally better, but large language models given only the remainder of the contract recovered the missing term nearly nine times out of ten. The authors conclude that true gaps are rarer than the literature supposes, and propose that courts treat model predictions as ordinary contestable evidence while parties constrain the practice with explicit "Choice of Model" clauses.

Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?

Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro cross-listed Strong performance by language models on medical licensing exams has fueled expectations of similar legal competence, but legal truth is contingent on jurisdiction, time, and source hierarchy rather than stable empirical fact. A comparative diagnostic framework evaluates legal against medical reasoning along four axes — knowledge recall, grounding, confidence, and robustness — using a new benchmark that encodes temporal validity and normative relationships between authorities. Medical models reliably benefit from verified sources, but legal models struggle to judge whether a retrieved citation is useful or misleading, showing overconfidence under perturbed context and sensitivity to superficial formatting cues; notably, larger models are more susceptible, as stronger instruction following coincides with weaker resistance to authoritative perturbation. The pattern suggests models treat law as unstructured text rather than binding precedent, over-trusting authoritative-looking but false references that conflict with their internal knowledge.

ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation

Eric Inae, Tim Gunn, Chris Bond, Meng Jiang Chemistry benchmarks for language models tend to test a narrow slate of tasks in one fixed format, so reported scores may reflect prompt familiarity rather than chemical reasoning. ChemDIRT varies both the instruction phrasing and the molecular representation across eight categories of chemistry tasks, scoring not just accuracy but consistency under these controlled perturbations. Benchmarking open- and closed-source models reveals substantial prompt sensitivity and representation dependence, along with uneven performance across task families.

Beyond Sparse Weights: When Is Attention Compressible?

Chiwun Yang, Xiaoyu Li KV-cache compression is usually justified by attention maps that look sparse, but a few large weights need not hold most of the probability mass, dropped value vectors can cancel each other, and preserving attention output does not guarantee preserving task accuracy. The analysis separates these three claims, showing that global score gaps rather than threshold counts determine how many tokens retain a target mass, and that the weighted sum of omitted values is the exact quantity lost for a given row. The resulting training-free compressor CertKV reserves one tail-summary slot per attention head and allocates the remaining budget by value dispersion; under matched budgets it lands top-two in seven of nine LongBench-v2 settings, stays in the leading tier on 128K-context RULER, and delivers a tenfold cache reduction in a packed Llama prototype.

Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation

Lorenz Brehme, Adam Jatowt Evaluating retrieval-augmented generation (RAG) systems on proprietary data requires question-answer sets that public benchmarks like HotpotQA cannot supply, and a thorough evaluation needs both multi-hop and unanswerable questions. TRIAD builds them in three stages: it drafts question-answer pairs from a domain-specific knowledge base, runs each pair through a validator in a feedback loop, then attaches relevance-labeled context documents for downstream scoring. Measured against MuSiQue and HotpotQA, the generated datasets reproduce similar performance trends across different RAG configurations, and human review found the questions suitable for evaluating a domain-specific system.

Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline

Naimur Rahman Multi-stage LLM pipelines can keep emitting structurally well-formed output while the evidence handed between stages quietly becomes incomplete, compressed, or contradictory. Evidence-State Reliability (ESR) is proposed as a separate evaluation layer asking whether intermediate evidence stays complete, grounded, internally consistent, and usable for each stage's job, scored independently of parser validity. Running GLM-5.2 over 60 sanitized base cases under clean, compressed-lossy, partial-dropout, and noisy-conflicting conditions through decision, audit, and escalation stages, all nine degraded-versus-clean comparisons showed stage success dropping while parser-validity estimates rose. Audit stages detected degradation in every degraded condition, but escalation recovery was zero throughout and false-assurance rates stayed non-zero.

Semantic Compression Trees: Multi-Resolution Knowledge Retrieval via Hierarchical Semantic Residuals

Junaid Farooq Retrieval-augmented generation typically indexes documents as flat, uniform chunks retrieved by similarity, throwing away the source's hierarchical structure. Semantic Compression Trees store at each node only its semantic residual — the information it adds beyond its parent — and retrieve by descending progressively from the root, so per-query cost scales with tree depth rather than collection size. On QASPER with the relevant document supplied, a zero-LLM extractive compressor matches dense retrieval on answer quality (0.274 versus 0.277 F1) using 30% fewer context tokens and no LLM calls to build the index, and residual storage clearly beats storing full summaries per node. Progressive descent itself does not survive scrutiny: retrieving the same residuals without the tree performs identically, and when the system must select the document, descent routes correctly only 20.2% of the time against 39.3% for flat retrieval.

SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning

Yujie Zhang, Bin Gao, Tulika Mitra Chain-of-thought (CoT) prompting stretches decoding time and memory on Mixture-of-Experts (MoE) models, whose full expert weights often exceed GPU memory and force costly transfers between GPU and CPU. SAEM is an inference runtime built on the observation that consecutive reasoning stages in a CoT trace activate a coherent, predictable set of experts, so it detects stage boundaries and uses them to drive stage-aware caching, expert-aligned token repacking, and in-situ CPU execution. On mathematical and scientific reasoning workloads under constrained GPU memory, it reports an average 1.33x throughput gain over the strongest caching and offloading baselines, rising to 1.54x when the calibration data matches the workload.

FrugalSOT - Frugal Search Over the Models

Pradheep P, Yuvanesh S, Harish KB, Keerthan Saai Reddy S, Joshva Devadas T, Naveenkumar J et al. Embedded devices cannot afford to run their largest available language model on every request, so FrugalSOT estimates each incoming request's difficulty from features like prompt length, named entity density, and syntactic complexity, then routes it to the smallest model expected to clear a relevance threshold and escalates to a larger one only if the output falls short. The threshold is updated continuously in the background from past validation outcomes using a low-pass filter, letting it track shifting input patterns. Measurements on a Raspberry Pi 5 show lower average inference time and resource use than a single-model baseline while losing less output relevance than a pure small-model strategy would.

Benchmarking Composable Compression Techniques in Mixture-of-Experts LLMs

Afsara Benazir, Chen Chen, Rongxiao Qu, Jiabo Huang, Jingtao Li, Lingjuan Lyu Deploying Mixture-of-Experts (MoE) language models on commodity hardware usually means stacking several compression techniques at once — expert pruning, weight quantization, and KV-cache compression — yet each is normally measured alone. MoEXBench runs them individually and in combination across 10 MoE models from 30B to 235B total parameters spanning standard, hybrid linear, and sliding-window attention, sweeping 20-50% pruning rates, 1-to-16-bit quantization, and multiple KV-cache precisions through an eight-module evaluation suite. The headline result is that composable compression cannot be predicted from the standalone behaviour of its parts, with compression rate a poor proxy for either quality loss or speedup, expert pruning the dominant source of degradation, and averaged scores hiding workload- and architecture-specific failures; normalized scores, compressed artifacts, and scripts are released.

From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism

Jing Liu, Yongxing Qi, Muchen Jiang, Chengnan Hu, Qingqing Peng, Haoming Wang et al. Standard retrieval-augmented generation (RAG) pipelines rank documents by dense-vector similarity, which often surfaces documents sharing keywords with the query but lacking the actual answer — a failure that worsens as the knowledge base grows. The authors model the retrieval step itself as a causal graph in which shared keywords act as a latent common cause and the document is a collider, which licenses a training-free re-scoring rule: cosine similarity between the query embedding and the weighted centroid of the document's residual keywords. On a diagnostic corpus built to reproduce keyword-stuffing, mean target rank improves from 2.88 to 1.25 while a trained cross-encoder reranker only reaches 2.63; on three BEIR benchmarks the score underperforms plain similarity, so a corpus-level calibration gate picks the right regime with at least 95% reliability.

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong Correcting health misinformation in conversation is not just a matter of stating facts, since the right intervention depends on a user's knowledge and how firmly they hold a belief, yet current systems either answer immediately or probe on every turn. RO-PnR (Reward-Optimized Probe-and-Respond) learns per-turn whether to ask a clarifying question or commit to a correction, using a reward that weighs expected information gain against interaction cost, with simulated users given latent health-literacy and belief-commitment states. Across three health-misinformation datasets and three base models it reaches the highest cost-adjusted utility while using 30% fewer turns than always-probe baselines.

FCPRAG: Fusion-Controller Parametric Retrieval-Augmented Generation for Stable Multi-Passage LoRA Injection

Jinchang Zhu, Jindong Li, Yi Ding, Xiaojian Nie, Rong Fu, Shuangyong Song et al. Parametric retrieval-augmented generation (PRAG) encodes retrieved passages as per-passage LoRA adapters instead of long prompts, but merging several adapters for one query is fragile: equal weighting amplifies weak or conflicting evidence, and turning retrieval scores into fusion weights usually needs brittle global tuning. FCPRAG adds a lightweight controller that predicts per-passage fusion scores plus sample-level calibration signals — a mixing gate and an adaptive temperature — so fusion is selective when retrieval is informative and conservative when it is not, trained with merge-aware supervision from each adapter's marginal contribution to the merge. Across HotpotQA, 2WikiMultiHopQA, PopQA, and ComplexWebQuestions on three backbones it improves F1 over both standard and parametric RAG, with gains up to 7.55% on ComplexWebQuestions, while lowering tuning cost and holding up under retrieval perturbations.

CD-LoRA: Consistency-Driven Low-Rank Adaptation for Multi-Task Fine-Tuning

Qian Zha, Jinda Liu, Yuan Wu, Yi Chang Multi-task fine-tuning of large language models with LoRA (Low-Rank Adaptation) usually relies on routers that dispatch inputs to task-specific adapters, and a second-order Taylor analysis presented here argues that stochastic routing decisions create a training-inference discrepancy that destabilizes predictions under distribution shift. CD-LoRA drops routers entirely, instead using a consistency-driven alignment objective that forces representations from different tasks to agree inside a single shared low-rank space. Reported experiments place it above state-of-the-art multi-adapter baselines while removing the overhead of explicit task partitioning.

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

Hyunwoo Kim, Byoungchan Ko, Minseok Kang, Minwoo Kim, Dongjin Lee, Jaehoon Lee et al. Mamba-2's Structured State Space Duality (SSD) blends recurrent and attention modes for efficient long-range modeling but carries heavy memory and latency costs. SSDi8 is a post-training quantization framework built specifically for SSD that keeps a persistent INT8 path by reformulating the computation to separate element-wise multiplications from matrix multiplications so quantized activations can be reused across modules, and by quantizing channel-varying activations at selectively chosen points. Accuracy is preserved through SSD's dimensional decomposition — exploiting different outlier distributions per axis — plus a per-channel error-statistics correction term, yielding accuracy comparable to FP16 with up to 1.4x speedup in W4A8 and W8A8 settings, validated on an Orin NX edge device.

The Communication Map of a Transformer

Richard Zhe Wang Mechanistic interpretability has traced how transformer components read and write the residual stream one circuit at a time, by hand. The communication map charts every potential channel from weights alone, generalizing the composition score of Elhage et al. into a single coupling coefficient spanning all 18 connection classes from whole attention-head circuits down to individual neurons. Censusing 6.3×10⁸ candidate channels in GPT-2 up to 1.3×10¹¹ in Pythia-6.9B, the authors find 70–89% of head pairs are oriented far from chance, and the full map takes 15 seconds for GPT-2 and 11 minutes for Pythia-6.9B on a single consumer GPU. Two applications follow: the strongest head-to-head couplings recover known induction circuits blind and ablating that community destroys in-context copying, while pooling coupling coefficients exposes a two-dimensional residual-stream subspace whose deletion abolishes induction in six models.

DynaContext: Self-Improving Dynamic Contextualization of Optimized Prompts for Heterogeneous Parameter Extraction

Joe Yu, Shibin Thomas Stanley Paul, Sven Mayer Automated prompt optimization usually yields one static instruction reused for every input, which fails when the needed fields, constraints, and evidence differ per item — extracting parameters from electronic component descriptions, where resistors, capacitors, and connectors demand different schemas, is one such case. DynaContext pairs an offline-optimized extraction core (learned with GEPA or SkillOpt) with inference-time adaptation: each item is routed through internal, external, or fallback evidence paths, and an item-specific prompt is composed from the core, schema, evidence, unresolved fields, and validated demonstrations, with deterministic checks plus an LLM judge gating output and only human-verified corrections entering demonstration memory. On a single-category benchmark accuracy rises from 86.6% to 98.6%, and across 850 heterogeneous gold facts field-level F1 goes from 51.8% to 71.0%, beating the deployed static-prompting pipeline by 17.3 F1 points with the model held fixed.

Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation

Nura Aljaafari, Andre Freitas Interpretability work has catalogued many individual transformer circuits without a common vocabulary for how their functions compose across tasks and architectures. CPC (Coherentist Probabilistic Compositionalism) proposes four operator roles — alignment finds candidate relations, unification integrates supporting information, suppression damps incompatible alternatives, and routing carries selections to the output — and tests whether weight-space signatures for these roles predict held-out activation measurements. Across 15 models from five architecture families, suppression, unification, and routing signatures correlate with activation-level measures above random baselines, ablating alignment heads reduces downstream suppression in 10 models, and explicit contradictions shift a layerwise coherence proxy in 14; the authors caution that similar effects on no-conflict prompts point to general upstream dependency rather than contradiction-specific coupling, and that role geometry stays architecture-specific.

More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

Orion Powers, Daniella Seum, Khaled Slhoub Local LLM deployments are usually chosen on accuracy alone, with runtime, energy, and failure modes rarely measured under controlled conditions. The authors define an evaluation procedure combining fixed inference settings, hierarchical answer extraction and verification, explicit failure-mode classification, per-question resource measurement, and paired significance tests, then apply it to three sub-5B open-weight models — Gemma3:4b, Phi3:3.8b, and Qwen3:4b — on Grade 8 math, Calculus I, and advanced probability and statistics, all on one workstation through the same inference server. No model dominates: Qwen3:4b is most accurate on two datasets and Gemma3:4b on Calculus I, but Gemma3:4b returns roughly three times more correct answers per watt-hour on every dataset while generating far fewer tokens, and Phi3:3.8b trails on accuracy with a low extraction-failure rate indicating genuinely wrong answers rather than unparsed output.

W-RAG: Source-Aware Retrieval for Enterprise Document Generation from Heterogeneous Knowledge Bases

Hridya Dhulipala, Rajesh Ombase, Michael Wang, Tien N. Nguyen cross-listed Retrieval-augmented generation pipelines normally rank all retrieved passages with a single global similarity function, which breaks down for enterprise document drafting where policies, regulations, technical documentation, and departmental guidelines each play a distinct role and must all appear in the output — global ranking lets one source dominate the context and produces incomplete drafts. W-RAG performs ontology-guided retrieval, ranks locally within each knowledge base, and applies source-level weighting to control how evidence is composed. The work also introduces a dataset for retrieval-grounded enterprise document generation across several document types and industry domains, on which standard RAG pipelines struggle while W-RAG improves document coverage and generation quality.

Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation

Laxmigayathri Challa, Yuhan Zhou, Ana Cleveland, Haihua Chen Language models can generate synthetic oncology records to relieve data scarcity, but a single clinically impossible cancer stage contaminates every model trained downstream, and neuro-symbolic pipelines that validate during generation have never had their individual quality gates isolated. Three controlled studies hold the generation protocol and fine-tuning hyperparameters fixed while varying gate necessity, constraint attribution, and retrieval use, with the symbolic gate checking schema completeness, SNOMED ontology coverage, and staging logic under AJCC eighth-edition rules. Without gating, 29.9% of records have schema failures and 20.1% have clinically invalid staging; within the fully gated corpus schema validation rejects 148 of 512 records, ontology grounding 24 more, and staging-logic validation none — the only generator producing logic violations is already caught on schema. Retrieval augmentation proved strongly model-dependent, helping one generator by 12.5 percentage points, doing nothing for a second, and collapsing output for a third.

What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine

Shahir M A Whether a language model actually executes on the Apple Neural Engine (ANE) turns out to depend on how a computation is expressed rather than what it computes, established through three measurement paths: a 64-shape sweep of LLM primitives recording per-operation device support, matched trained models across size and precision, and readings of the ANE's memory-controller byte counters during inference. A fused RMSNorm is fully ANE-eligible while its arithmetically identical decomposition falls back to CPU, and weight encoding gates the accelerator — a 25.85M-parameter convolution-heavy fp16 model is assigned entirely to the CPU with zero bytes through the engine, while the same graph in int8 or 2-bit returns to roughly 83% residency and runs 1.8-2.2x faster. Decode cost tracks bytes streamed per token at a near-constant ~0.77 fraction of nominal encoding width across all precisions, and ternary half-attention models at 25M and 50M parameters end up 9.8x and 6.1x smaller and 3.0x and 2.2x faster than the fp16 design the work started from. The design procedure drawn from this is to choose the weight encoding first, then spend the remaining byte budget on parameters.

RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored

Gregory Druck, Ethan Smith Recursively training language models on their own output is known to cause model collapse; the same degradation is shown to arise when a retrieval-augmented system searches the web and pulls back documents it authored itself. Across three simulation designs, three model families, and 1,019 information-seeking prompts — 1,528 simulations and over a million API calls — 79.6% of simulations ended in collapse. A single self-authored reference in the retrieval pool can be enough to trigger it, because the model disproportionately cites its own content, a self-bias that persists after controlling for reference quality.

MegaMem: A Retrieval Solution for Ultra-Large Context Windows

Xinyuan Song, Bowen Zhu, Hasibul Haque, Liang Zhao Persistent memory over whole codebases, long interaction histories, and heterogeneous enterprise records demands keeping hundreds of millions of tokens searchable while passing only a bounded slice of source evidence to the answering model. MegaMem separates semantic access from generation evidence: distilled records and detailed evidence are searched with both original and transformed queries, every distilled hit resolves to an immutable source ID before reciprocal-rank fusion, deduplication, and cross-encoder reranking, and only the top-ranked detailed evidence within a fixed budget supports generation, with post-answer attribution identifying which loaded sources backed the response. On EnterpriseRAG-Bench — more than 500,000 documents and roughly 650M tokens — the overall score rises from 68.22 to 82.26, reaching 86.50 correctness, which the authors present as a path toward retrieval over memories approaching one billion tokens.

Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations

Samira Golsefid A single accuracy score reveals nothing about how a model degrades as inputs get harder, so this stress-testing framework perturbs each problem along a graded severity ladder across seven families: six that preserve the answer (paraphrase, input noise, formatting, irrelevant context, context load, and conflicting instructions) plus a Knowledge Boundary family that removes answerability so refusal becomes the correct response. Every test is validity-gated and labeled by measured severity, and each model is summarized by per-level accuracy, a magnitude-weighted stability score, and a per-family collapse point defined against its own baseline. Expanding the 100 GSM-Symbolic seed problems into 4,473 gated tests across four models spanning capability tiers shows that the level at which a model fails is family-specific rather than global, with conflicting instructions and impossible-premise questions the two consistent weaknesses; recognition of unanswerability is reliable for missing information and fabricated evidence but weak for impossible premises.

MEMONDEMAND: A Memory Management System for Large-Scale Enterprise Data

Xinyuan Song, Bowen Zhu, Hasibul Haque, Liang Zhao Enterprise repositories are large, heterogeneous, and continuously updated, so a memory layer must handle collection-specific hierarchy construction, cheap routing, detailed evidence loading, and workload-aware updates together rather than piecemeal. MemOnDemand combines a dynamic multi-level hierarchy whose abstraction structure and depth are chosen per collection, dual memory at every level separating distilled routing summaries from detailed evidence, and on-demand promotion that re-prioritizes nodes under a bounded active-state budget. On EnterpriseRAG-Bench it beats the strongest published leaderboard result at every scale tested, with gains of 12.23% at 10M tokens and 4.66% on the full 618M-token collection, plus reported strength on FinanceBench, HotpotQA, and FRAMES covering financial, multi-hop, and fact-retrieval settings.

Unveiling the Depth-Performance Dilemma in Split-Federated Fine-tuning of LLMs

Hariharan Ramesh, Someshwaran Murugaiyan, Jyotikrishna Dass Split federated fine-tuning splits a language model by depth between weak clients and a server, and throughput and privacy both improve as the client-side portion grows deeper — but an audit across four scales from GPT-2 to Llama-3-8B finds that the deep partitions that maximize system efficiency are exactly where fine-tuning quality collapses, a tension the authors name the Depth-Performance Dilemma. Federated adapter aggregation methods that work in ordinary federated learning — AVG, STACK, SVD, and FREEZE — fail to fix the split-specific degradation. The mechanistic diagnosis attributes the collapse to the near-isometric topology of transformers, which lets aggregation noise propagate without attenuation until it triggers attention collapse in the server-side partition, contradicting the assumption that partition depth is a utility-neutral knob.

Improving Few-Step Language Flows with Untied Self-Conditioning

Bocheng Li, Linli Xu Flow-matching language models denoise every token position in parallel and can trade sampling steps for speed, but quality collapses at low step counts. The authors trace part of that collapse to a train-inference mismatch in self-conditioning: at training time the self-conditioning input comes from the current noisy state with no solver step in between, whereas at sampling time the solver has already folded the previous prediction into the latent before it reappears as an explicit input, creating redundancy that grows with step width. Untied Self-Conditioning corrects both halves without retraining by damping directions where the self-conditioning input duplicates the latent, identified from frozen projection weights, and approximating the step-average prediction from prediction history. At 8 sampling steps on LangFlow it cuts OpenWebText generative perplexity from 531 to 62, an 8.6-fold reduction, with its outputs preferred in 96% of pairwise comparisons under an adapted Arena-Hard-Auto v2 protocol.

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Yan Zhan, Mengkai Hou, Wanting Zhang, Zhijun Gao Masked diffusion language models must commit to a target sentence length before denoising begins, a decision prior work on unmasking order has largely ignored even though it directly controls coverage and redundancy in translation. Entropy-Valley is a training-free length selector that scores candidate canvas sizes by mean predictive entropy from all-mask forward passes and picks the length the backbone appears most ready to fill. Against a baseline using training-corpus length statistics it recovers 64.9% and 65.3% of the COMET-22 gain available from reference target lengths on English-to-Chinese and Chinese-to-English (33.0% on English-to-German), ties or leads a LLaMA-3-8B autoregressive model fine-tuned on the same data, and an oracle-length diagnostic indicates that how length is supplied matters more than which tokens are revealed first.

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

Ergan Shang, Weijing Tang, Yinqiu He Predicting how a language model will score on questions nobody has annotated yet would cut evaluation cost, but retrospective averages confound a model's ability with the difficulty of the items it happened to be scored on. The framework fits a multidimensional item response theory model in which each model carries a latent capability profile and each question's characteristics are inferred from embeddings of its content, allowing information to transfer to unseen items. Within a scenario, question embeddings improve prediction over model-free baselines and multiple latent dimensions describe capability better than a single one, but the approach does not predict reliably under cross-scenario shift, which the authors name as the central open problem for context-aware psychometric evaluation.

Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers

Yan Wang Low-precision optimizer-state methods are built for dense Adam-style moments, but memory-efficient optimizers keep factored, confidence-based, or projected states whose quantization errors behave differently. Adaptive Log-Space quantization is a block-wise representation for non-negative states that adapts its nonzero range per block while preserving exact zeros, paired with independent signed-momentum encodings and per-state precision choices, evaluated over 96 runs and 214.7 GPU-hours across AdamW, Adafactor, CAME, and APOLLO. On a 20K-step TinyLlama-1.1B benchmark, AdamW with 8-bit log-space second moments reaches 72.90 perplexity against 72.48 for FP32 and 73.54 for an 8-bit dynamic baseline, while cutting measured optimizer-state storage from 8392.7 to 2119.2 MiB; CAME needs 16-bit for its non-negative states, and topology-aware parameter protection shrinks quantized Adafactor's late-loss gap from +0.1185 to +0.0159.

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

Qingwen Lin, Boyan Xu, Xiao Liu, Zhifeng Hao, Ruichu Cai Delta-rule recurrent models keep a fixed-size state for O(1) inference memory but can destabilize far beyond their training context. Tracking RWKV-7 over sequences up to 100M tokens exposes a specific failure mode — localized norm explosion in a few channels atop a sparse substrate, rather than global saturation — because persistent decay keeps weakly updated entries small while uneven injections let a few channels accumulate extreme values. SANE (State Anomaly Neutralization) applies adaptive tanh compression at chunk boundaries, preserving intra-chunk parallelism, and after a 100M-token prefix (over 24,000x the training length) retains functional reasoning at 33.46–35.56 while the baseline hits numerical overflow, with no significant degradation on 11 short-context benchmarks at thresholds 3 ≤ α ≤ 5. Overly permissive thresholds (α ≥ 8) stay numerically stable but destroy reasoning entirely, exposing a capacity–stability trade-off.

SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation Evaluation

Jiarui Dong, Yin Cai, Zhouhong Gu, Chenmou Wu, Ci Tao, Yiran Chen et al. Evaluating LLM-generated graphical user interfaces is unreliable because existing data has uncontrolled distributions, noisy annotations, and narrow layout coverage. SchemaGUI synthesizes paired natural-language instructions and deterministic function-call references from parameterized interface schemas, producing thousands of exactly-annotated tasks in seconds without human labeling, and is used to benchmark five models — including the Qwen3.5 family, Qwen3-Coder-30B, and DeepSeek-R1 — over 1,000 instances per scenario across six bilingual scenarios. Precise geometric control is the standing bottleneck: scaling Qwen3.5 from 4B to 27B lifts Schema Feasibility from 91.56% to 99.63% but Geometry only from 67.05% to 75.30%. Difficulty tracks layout complexity, with severe coordinate drift in dense grids and multi-region compositions, and enabling thinking mode raises token consumption while generally lowering GUI Score, especially for smaller models.

Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator

Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban When a language model judges outputs in different prompt languages, the ranking of evaluator backbones changes: on an eight-language Agent-as-a-Judge benchmark the top backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Turkish, and Swahili, and 7 of 15 backbone pairs show statistically significant rank reversal. Treating the judge score as an additive decomposition of task difficulty, backbone skill, and a language-backbone interaction, Consensus-Based Calibration recovers that interaction term without human labels by double-centering the cell-mean matrix — the standard two-way ANOVA operation — with a finite-sample concentration bound and unbiasedness even when task-language interactions exist. Across 7,920 judge runs the calibration raises held-out cross-task rank consistency from 0.650 to 0.902, and on a separate M-RewardBench panel agreement with human gold preferences rises from 68.7% to 76.6%.

When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing

Taehyeon An, Jaehyeong Park, Donghyuk Shin Persona-conditioned language models are used to stand in for survey respondents, but the variation they produce may come from unconditioned model priors or sampling noise rather than from the persona itself. Persona-Conditioned Informativeness (PCI) is an unsupervised diagnostic built on the premise that conditioning is real when semantically similar personas shift in concordant directions: personas form a similarity graph and Local Moran's I measures local spatial coherence relative to item-level baselines, without needing construct labels. Validated by whether it recovers known latent value structure in the 57-item Portrait Values Questionnaire-Revised, confirmatory factor analysis shows a PCI-selected 10% subset substantially improves construct recovery over response-stability and random selection.

Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Li Chen A single confidence score for a whole language-model response is not actionable when that response mixes true and false statements, so the work breaks each answer into atomic verifiable claims and calibrates a confidence score for each one. The method works in a closed-box setting with no access to logits or fine-tuning, deriving signals from consistency across sampled generations and self-verification, then applying post-hoc calibration per claim so that low-confidence claims can be routed to retrieval or human review. Testing seven baselines on six models including Llama-3.1, Qwen2.5, DeepSeek-R1, and GPT-4o over TriviaQA and TruthfulQA shows claim-level decomposition plus post-hoc calibration lowers expected calibration error on factual questions, while adversarial false-premise questions remain a visible failure mode.

Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification

J\'er\'emie Dentan, Alexi Canesse, Mahammed El Sharkawy, Sonia Vanier Checking the factuality of individual claims in language-model output usually relies on cross-encoder models that need a GPU, which makes real-time monitoring expensive. Kernel Token Contradiction builds a positive semi-definite kernel over the candidate tokens considered during generation, combining the model's conditional distribution with a token contradiction score estimated from Wikipedia frequency statistics, and quantifies uncertainty as the Von Neumann entropy of that kernel. Running CPU-only, it delivers over an 8.2x speedup versus GPU-accelerated cross-encoder methods and over 65x versus comparable CPU methods, matching their average performance across two benchmarks, four European languages, and 16 models while doing better in high-precision regimes.

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein Natural-language rubrics for judging long-form model output are ambiguous, need a black-box language-model judge to apply, and usually assume criteria combine as a linear weighted sum, which cannot express dependencies, alternatives, penalties, or override conditions. ExecRubrics compiles rubrics into compact executable Python scoring functions, giving the rubric a fixed decision procedure that can be inspected, run, and edited, and optionally calling text-processing libraries such as NLTK and spaCy for extra signal. On HealthBench, HelpSteer, and ArgQuality it matches or beats natural-language rubric baselines at ranking preferred over dispreferred responses (best accuracies of 53%, 78%, and 92%) while cutting evaluation latency by up to 320 times.

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun et al. Existing clinical language model datasets test factual recall or surface retrieval rather than the inductive and deductive reasoning physicians use to pick out decision-relevant evidence from dense, fast-changing intensive care unit records. ICU-REACT is a reasoning dataset built with 19 clinicians in the loop and aligned to the OMOP common data model, used to fine-tune the Clin-REACT family spanning 8B to 70B parameters across three model families. Across five clinical reasoning benchmarks the tuned models beat their own backbones plus open general-purpose and medical models, and the gains transferred beyond critical care to script concordance tests and downstream diagnosis and treatment tasks, though prospective evaluation remains outstanding.

NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching

Nobel Dhar, Md Romyull Islam, Xuechen Zhang, Gongjin Sun, Sahidul Islam, Bobin Deng et al. cross-listed Most edge deployment work assumes a model can be quantized, shrunk, or partitioned to fit resident memory; the harder target here is when the model stays larger than memory throughout execution and storage sits on the critical path of every token. The key observation is that 82–85% of active MLP neurons persist from one token to the next, so nearly all sparse weights needed now are already resident and only the delta must be fetched. NeuroPrefetcher places one GPU-resident predictor after layer 0, costing 2.86% of base model parameters, which predicts sparse activity for every downstream MLP layer in a single forward pass; the runtime diffs those predictions against resident buffers and issues explicit NVMe reads for just the incoming rows instead of relying on operating-system demand paging. On real unified-memory edge hardware this delivers a 7.9–12.0× speedup over llama.cpp across constrained memory budgets.

CAI-DLLM: Convergence Aware Inference for Diffusion Language Models

Farhana Amin, Sabiha Afroz, Dimitrios S. Nikolopoulos Diffusion language models generate many tokens in parallel but keep spending denoising steps recomputing positions that have already stabilized. CAI-DLLM is a training-free inference method that reads only first-step confidence to commit easy tokens early, concentrate denoising on hard ones, and adjust decoding schedules across output blocks — no retraining, extra predictor, or weight updates. On LLaDA-8B-Instruct it reaches 18.2× wall-clock speedup on GSM8K while accuracy rises from 76.27% to 77.41%, and on Dream-7B-Instruct 13.1× on HumanEval with pass@1 improving from 46.95% to 48.17%. Harder reasoning tasks see up to 44.8× speedup with a worst-case 4.4-point accuracy drop, and energy consumption falls by as much as 95.3%.

WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs

Yiming Yao, Chenyang Lyu, Xuanfan Ni, Longyue Wang, Weihua Luo, Yazheng Yang et al. Long audio inputs make the key-value (KV) cache the dominant memory cost for speech LLMs, and prefill-only compression permanently evicts positions with no way to recover them during decoding. The authors show why this breaks on long-form audio: prefill attention concentrates near the audio start in an attention-sink pattern while decode-time attention spreads broadly, so the two importance rankings barely overlap. WnW calibrates offline to sort KV heads into three roles — anchor heads stay resident on GPU and act as a decode-time importance observer, tidal heads keep a CPU-resident complement recalled chunk by chunk from aggregated anchor scores, and fixed heads retain only a GPU subset with the rest discarded. On LibriSpeech-Long with Voxtral-mini-3b and Qwen2.5-Omni-3B, it holds near-full-cache accuracy while keeping just 20% of audio tokens on GPU, a budget at which prefill-only baselines fail to terminate, with little measured decode overhead from CPU-GPU recall.

XTC: Head-Aware Sampling by Excluding Top Choices

Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv Standard decoding rules encourage diversity by rescaling the whole next-token distribution or truncating its tail, which misses the common case where several continuations are plausible but probability mass stays piled on the most generic one. XTC (Exclude Top Choices) flags tokens above an absolute plausibility threshold and, when at least two qualify, drops the dominant eligible choices with some probability and keeps only the weakest plausible alternative before renormalizing. Across 60 experiments on Gemma 3 27B Q4, Gemma 3 12B Q6, and DeepSeek R1 14B Q6 with validation on Llama 3.3 70B Q4, it raises Distinct-2 by 11–15% and cuts repeated trigrams by 27–47%, a blinded study of 150 Mechanical Turk Master raters gave it a 62.3% creativity preference without loss of fluency, and IFEval instruction-following stays within 1.7 points of baseline; the sampler has been adopted by llama.cpp, ExLlamaV2, and text-generation-webui.

Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time

Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu, Ravid Shwartz-Ziv Open-ended generation often collapses into verbatim loops, and the usual defenses — repetition, presence, and frequency penalties plus n-gram blocking — react to token recurrence rather than the sequential structure of a loop, so they only stop looping at settings that also wreck formatting and fluency. DRY (Don't Repeat Yourself) is a sampling-time logit adjustment that penalizes a candidate token only when picking it would extend the current suffix into an exact continuation of a span already seen in context, with sequence breakers protecting chat templates and formatting tokens. Across models from 1.5B to 120B parameters, nine prompt families, and a 600-pair human study, it cuts the suffix-extension rate by 47% while raising lexical diversity, halves loop rate on AWQ-quantized 70B and 120B models without measurable loss on MT-Bench, MMLU, or GSM8K, and an intervention-matched placebo shows suffix matching is the operative mechanism; it is shipped in llama.cpp, ExLlamaV2, and text-generation-webui.

SPOC-SQL: Stage-wise Preference Optimization for Controllable Text-to-SQL

Yingnan Chen, Chun Ding, Tianshi Xu, Xu Yang, Si Wu Translating natural language questions into SQL is usually trained as single-step sequence generation, which gives no targeted feedback at individual decision points and leaves no handle for steering intermediate choices. SPOC-SQL splits the task into four sequential subtasks that mirror standard SQL execution order, applies preference optimization separately at the key decision point of each stage, and exposes explicit intermediate representations so a stage can be inspected and corrected before the query is assembled. Experiments report that adding stage-wise human knowledge consistently improves accuracy, supporting the claim that stage-aware controllable generation beats end-to-end sequence optimization.

Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports

Parsa Bakhtiari, Hassan Bashiri, Alireza Khalilipour, Masoud Nasiripour, Moharram Challenger Industrial technical reports mix dense prose, specification blocks, and tables, which defeats standard retrieval and question-answering pipelines, and no public instruction-tuning or benchmark data is built from them. Industrial-Instruction supplies both a reproducible pipeline — layout-aware extraction, a semantic retrieval index, and synthesis of multiple-choice questions grounded in retrieved evidence under five query-document relationships spanning irrelevant retrieval through multi-document answers — and two datasets of roughly 13.6k filtered question-answer pairs each, built from 906 public Panasonic documents totaling 7,525 pages. Fine-tuning open models under 10B parameters lifts Set-Match Accuracy from 28.5% to 42.0% and F1 from 46.6% to 63.5%. The two parallel versions, one generated by Qwen3-30B-A3B-Instruct and one by Claude-Opus-4.6, allow a direct open-versus-frontier comparison: the frontier-generated data yields a cleaner corpus, larger fine-tuning gains, and essentially no MMLU forgetting, at roughly 100x the generation cost.

Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL

Kate Gwimm, Carson Eisenach Enterprise text-to-SQL is limited less by model capability than by what context can be squeezed into the prompt when business logic spans thousands of tables, so the authors argue for constructing the knowledge-base context from historical query usage rather than tuning retrieval over a fixed one. Using a query-DAG decomposition recovered from production SQL — the same intermediates that benchmarks like BEAVER annotate — they find retrieved knowledge-base context adds the largest marginal gain even on top of a full oracle query graph, then distill historical query profiles into reusable SQL reference cards. On 5,176 production queries from a major online retailer, optimizing these context artifacts yields roughly 12-25% AST similarity improvement versus 3-12% from optimizing the retrieval harness; on public BEAVER, which lacks production usage signals, results are mixed, with table cards alone matching raw historical SQL and the best variant scoring 9.00% against 6.33% for the comparable baseline.

Stochastic gradient descent with initial regularization

Nabil Kahal\'e Judging language-model research ideas is currently unreliable: free-form judging drifts with writing style and option order, and scoring a proposal against a paper that later appeared just rewards recovering one realized path. Lit2Test instead makes proposals decidable by requiring a six-field contract built around a falsifying outcome, so each proposal precommits the observation that would refute it. Built prospectively from 200 real-paper neighborhoods, the benchmark collects proposals from four frontier models and runs 1,200 blind pairwise comparisons in both presentation orders, with diagnostic controls and three-annotator human calibration; it recovers a strict ranking of the four models in all 10,000 bootstrap replicates, with separation driven by the quality of proposed tests and metrics rather than surface fluency.

A Physical Response-and-Memory Model for Muon Optimization

Yinze Hu, Hongjun Xiang, Xingao Gong, Hongyu Yu Optimizers like Muon — which semi-orthogonalizes the momentum matrix before updating weights — were arrived at through engineering intuition, leaving open why the semi-orthogonalized direction helps and how far back momentum should average. Modeling the weight matrix as a responsive physical medium with internal memory yields both answers: the semi-orthogonalized direction is the maximally dissipative response under an output-side safety budget, and momentum is accumulated internal stress whose relaxation, in real media, happens on more than one timescale. The resulting Bi-Maxwell optimizer replaces the single-timescale memory kernel with a fast and a slow one and reaches target loss in noticeably fewer steps on a public language-model optimizer benchmark; a read-only probe across 8 training trajectories supports the framework's prediction that optimal memory length should grow as training proceeds.

SplitLite: Low-Rank Residual Compression for Split Learning

Tao Li, Yulin Tang, Qi Guo, Xianhao Chen Split learning lets weak devices fine-tune large language models by offloading most computation to a server, but the activations and gradients crossing the client-server boundary each step dominate the cost. SplitLite exploits a structural observation: when LoRA (Low-Rank Adaptation) applies rank-$r$ updates, the activation and gradient residuals for the same sample between adjacent epochs turn out to have effective rank $2r$ and $4r$ respectively, so only quantized truncated singular value decomposition factors of those residuals need to be transmitted. Across the GLUE benchmark and several on-device language models this cuts activation uplink traffic by up to 93.5% and total communication by up to 83.7% with no reported accuracy loss.

Most of the LLM routing gap is task type

Janghoon Lee Routers that pick which language model answers each query look attractive because different models fail on different questions, yet published routers cluster within a fraction of a point of each other and often fail to beat always calling the strongest model. To find out what the missed questions have in common, fourteen models answered all 294 questions spanning 7 task types across Korean, English, and Hindi, with the entire matrix run twice — and 5.37% of the 4,116 model-question pairs scored differently between identical runs, setting a noise floor that small routing wins fall below. Counting only answers correct in both runs, just 29 questions are improvable by routing, and a static table assigning one model per task type in advance captures 21 of them, leaving 6 questions for a learned router; that table answers 262 of 294 questions at \$3.33 per run against the best single model's 245 at \$7.69. The authors note everything is fitted and scored on the same 294 questions with no holdout.

Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling

Ha Dinh, Xuan Duy Ta, Khoat Than, Khac-Hoai Nam Bui Semi-structured N:M sparsity accelerates language model inference on supporting hardware, but learnable-mask methods model a full categorical distribution over every feasible N:M pattern, so parameter and memory overhead grows combinatorially. Reservoir of Importance (RoI) instead learns masks through differentiable subset sampling with a compact-logit parameterization and sampling without replacement, reducing trainable parameters to O(M) and requiring 1.5–8.75× fewer learnable parameters than prior approaches. Across the Qwen2.5 family from 0.5B to 7B parameters it matches competing methods' accuracy while using less memory and scaling to more aggressive sparsity patterns.

POOL: Propagated Uncertainty Over Lookalikes

Rounak Sharma, Ananya B. Sai, Soumyabrata Pal Confidence scores for black-box language models let systems triage human review, escalate uncertain cases, or set abstention thresholds, but verbal confidence is cheap and overconfident while sampling-based uncertainty costs one set of samples per query. POOL (Propagated Uncertainty Over Lookalikes) borrows from group testing: it clusters overlapping query stems, runs a base estimator only on cluster medoids, softly propagates the resulting scores to nearby queries, and re-evaluates only high-disagreement cases. Instantiated with Hy@p, a hybrid of verbal confidence and spectral answer diversity measured by negative von Neumann entropy over sampled answer embeddings, it beats verbal confidence and 10-sample estimation on AUROC across six domains, three datasets, and five black-box models while using half the samples, and retains 93.5–97.9% of its AUROC while saving 19.3–39.3% of generations — rising to 73–76% savings on paraphrase-dense workloads.

Definitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark

Martin Wessel, Timo Spinde, J\"urgen Pfeffer, Gianluca Demartini Media bias datasets label the same category using different, often implicit definitions, leaving it unclear whether models trained on them learn the same construct. A between-subjects experiment with 354 participants and a parallel evaluation of four language models had raters score six news articles across four bias categories using definitions that varied either in conceptual framing or in the amount of construct-preserving elaboration. Across 8,496 human and 28,800 model ratings, the conceptual target of a definition drove annotation divergence while elaboration did not, and the shift from reframing was even stronger for models than for humans. The authors release MUDD, the Multi-Definition Bias Detection Dataset, and argue the effect generalizes to prompt-based measurement well beyond media bias.

LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space

Jinghui Zhang, Lang Gao, Ao Li, Mingzhe Li, Ruihong Zeng, Zirui Song et al. Author-personalized writing systems usually treat each author or style as a separate label requiring its own corpus or fine-tune, which is expensive and opaque. LiteraryBigFive borrows the dimensional logic of the Big Five personality model and derives five interpretable stylistic axes (such as Classicism and Emotionality) from activation-space contrasts between author-written and neutral passages, positioning any text or author as coordinates in that space. A steering mechanism then nudges generation toward target coordinates, and the reported result is improved authorial expressiveness with preserved semantic fidelity, plus per-axis author scores that correlate strongly with real-world literary consensus.

Activation-Weighted Seeded Residual Coding for Low-Bit LLM Weight Repair

Zehao Liu, Chuangchuang Fang, Yang Ren Quantizing language model weights to low precision saves memory but leaves reconstruction error that hurts quality. Activation-Weighted Seeded Residual Coding (AWSRC) is a sidecar codec that encodes the residual between original and quantized weights using deterministic seed-generated bases, storing only seed selectors, low-bit coefficients, and scales instead of an explicit codebook, and weighting the repair by activation statistics so errors that actually move layer outputs get priority. On Qwen2.5-3B-Instruct, spending an extra 0.162 bits per weight on top of an INT4 round-to-nearest backbone closes 88.2% of the perplexity gap to BF16 (78.9% for KL divergence, 71.3% for accuracy), and at a matched 49.25 MB sidecar it beats sparse, low-rank, and vector-quantized alternatives.

Language Chain in Alignment: Cross-Lingual Ranking Preference Optimization

Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim Preference alignment data is overwhelmingly English, leaving models weaker when asked to follow instructions in other languages. Cross-Lingual Ranking Preference Optimization (CRPO) pairs parallel preference data in English and a target language into a hierarchy that jointly optimizes preferences within and across languages, and builds on the LambdaLoss ranking framework to supply a graded ranking signal over multiple candidate responses rather than a binary better/worse comparison. Across five languages spanning different resource levels it consistently beats standard alignment baselines on both instruction following and knowledge use, with larger reward margins and higher log-probabilities for preferred responses.

Accelerating Diffusion Language Models via Structured Suffix Modeling

Zifeng Cheng, Keda Li, Zhiwei Jiang, Cong Wang, Fei Shen, Qing Gu Diffusion language models decode many tokens per step but pay for it by attending to every suffix token at each step; existing speedups just keep a local suffix window and reset all suffix tokens to identical representations every timestep. The proposed method splits the suffix into local, middle, and tail regions and keeps a different number of tokens in each according to its structural role, while carrying the previous step's decoding results into the current suffix representations so they accumulate denoising information. It is training-free and composes with parallel decoding and KV caching, reaching up to a 72.81× speedup on long-sequence inference when stacked with other acceleration techniques, with accuracy improved in most benchmark settings.

The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search

Peiyang Liu, Xi Wang, Di Liang, Wei Ye Retrieval-augmented generation is held back by two coupled problems: relevance proxies that misreport which retrieved evidence the model actually used, and poor decisions about how to spend a fixed context budget. The authors show standard relevance proxies break down on hard negatives and replace them with a cheap causal leave-one-out probe that isolates true generative reliance, then use that probe in a deconfounded factorial grid to compare context strategies. Widening a single monolithic context turns out to be penalized by relevance decay, whereas spreading the same compute across multiple sequential generations lifts portfolio recall by 16.7 to 20.5 absolute percentage points, holding up to 32B-parameter models. The findings are packaged as a closed-loop submodular scheduler paired with an attribution-steered contrastive decoder that pushes the model to integrate fresh evidence.

EvoWiki: Incremental State Overwriting and Traceable Question Answering for Cross-Meeting Knowledge Evolution

Dongsheng Chen, Tianyu Wang, Wenhui Que Across a series of meetings, facts like decisions and risks get revised or overturned, but long-context methods stack the whole history and most retrieval-augmented or structured-memory systems treat knowledge as static or append-only, so stale and current states coexist and answers are hard to verify. EvoWiki splits offline incremental construction from online reading: the build stage tracks how a topic moves from proposal to decision and uses entity version chains plus a State-Overwrite Protocol to mark which state is currently valid while keeping superseded history and meeting-level provenance, and the read stage skips top-k relevance retrieval in favor of deterministic entity addressing, temporal resolution, and multi-hop aggregation. On six datasets with two reader models, and on a new bilingual benchmark called CrossMeet built to simulate long-term state evolution, the system improved macro-average judge accuracy over the strongest baselines by 9.72 and 10.00 percentage points.

Sigmoid Attention as a Better Substrate for Learned KV Cache Eviction

Isaac (Rucheng), Li Learned key-value cache eviction suffers a soft-to-hard mismatch: training uses differentiable gates that merely attenuate a token's contribution, while inference only saves memory when entries are physically deleted. A controlled 2×2×2 study over attention type, learned gating, and positional encoding on GPT-2-scale Transformers trained on OpenWebText finds that although sigmoid attention is the weaker dense language model, sigmoid-gated models delete cache entries with negligible perplexity change relative to their own no-eviction references, and under a matched live-cache protocol they beat the authors' H2O and KeyDiff implementations while softmax gates do not consistently do so. The takeaway is that the attention normalization choice itself governs whether a training-time soft gate transfers cleanly to hard deletion.

The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning

Matthew Perlman, Atharva Nijasure, James Allan Low-Rank Adaptation (LoRA) is the default way to turn a large language model into a reranker, but where in the network relevance behaviour actually gets learned has been unclear. Ablation and attention experiments on RankLLaMA show that, with LoRA-adapted feed-forward layers throughout, restricting LoRA attention updates to a compact mid-network band recovers more than half the gain from adapting every attention layer, and omitting that band costs more than omitting any other region. Those same layers are where fine-tuning most increases attention to axiomatic information retrieval features — term rarity, lexical matching and query-document interaction — and those features correlate strongly with ranking improvements.

FormuEvo: LLM-Guided Evolution for Discovering Solver-Efficient Mixed-Integer Programming Formulations

Haofeng Yuan, Jianing Peng, Jieyi Bi, Ni Zhang, Shiji Song, Zhiguang Cao Language models can now draft mixed-integer programming (MIP) models from natural-language descriptions, but they optimize for semantic correctness and produce weak formulations that downstream solvers grind through slowly. FormuEvo treats formulation design as evolutionary search over executable modeling programs, with LLM-driven crossover, mutation and repair, fine-grained solver statistics fed back as verbal gradients for targeted refinement, and a structured memory that abstracts past runs into reusable modeling strategies. Across linear and non-linear problems the discovered formulations beat both expert-designed and prior LLM-generated ones, accelerating solvers by up to 5.5×, with the distilled knowledge transferring zero-shot to unseen problems and bootstrapping smaller models.

The Geometry of Low-Resource Language Representations

Francois Meyer, Jan Buys The accuracy gap between high- and low-resource languages in large language models is well documented, but the internal factors behind it are not. Comparing the geometry of hidden representations across 30 languages shows it varies systematically with training-data availability, most consistently in the final layers, where low-resource languages exhibit representational degeneration. Adding geometric regularization during continued pretraining of 9 base models adapted to 10 African languages reduces that degeneration, and for larger models a cosine-similarity penalty gives small accuracy gains over vanilla continued pretraining, concentrated on the hardest tasks.

Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data

Yifei Song, Kun Efimov-Zhang, Claire Gardent Turning structured inputs — tables, knowledge graphs, charts, time series — into text is normally studied one dataset at a time, relying on task-specific training data or zero-shot prompting. The setting studied here has neither in-domain training text nor test references and spans varied domains, goals and input structures, comparing data-driven knowledge distillation against zero-shot inference and out-of-domain fine-tuning, with structure-preserving augmentation by subsampling and perturbing the input structures. At a fixed 1.7B parameters, distillation beats both alternatives across five benchmarks and outperforms a much larger fine-tuned model on two domains, while scaling up real target-domain inputs through the new QUINTD-5 collection yields only modest gains by comparison.

A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study

Mullosharaf K. Arabov Arabic natural language processing has grown quickly but lacked a quantitative meta-analysis of the field. A bibliometric study covers 7,120 Arabic NLP papers from 1960 to 2026 across six collections, applying BERTopic for topic modeling, regression to identify citation predictors, social network analysis of co-authorship, and geographic mapping. Publications surged after 2020 alongside transformers and large language models; 19 substantive themes emerge, the largest covering text, speech, translation, and recognition; and indexing in OpenAlex or Semantic Scholar plus institutional affiliation correlate with higher citation counts. A task-by-dialect gap matrix identifies concrete blind spots, notably summarization for Maghrebi, Iraqi, and Sudanese dialects.

How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines

Chit-Fung Lam Grammar engineering — hand-building machine-processable formal grammars — is expert-intensive, and it is unclear how much language models can help. New Cantonese ParGram resources serve as gold standards, with English baselines, to test whether gpt-oss-120b and GPT-5.4 can generate grammars from either raw sentences or target formal structures under systematically varied prompts. GPT-5.4 outperformed gpt-oss-120b, and generating from target formal structures beat generating from sentences, but both models produced locally plausible phrase-structure rules, lexical entries, and templates while failing to coordinate interacting formal constraints, especially across multiple constructions. The conclusion is that models can assist intermediate stages of grammar development while human expertise remains necessary for analysis and validation.

ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation

Zhongpan Tang Attention cost and key-value cache size both grow quadratically or linearly with sequence length, capping ultra-long-context language models and high-resolution generators. ProxyFormer keeps two streams per layer: local fine-grained features are compressed bottom-up into a small set of proxy tokens, global interaction happens only among proxies, and the contextualized proxies are decompressed back into the persistent local stream — so detail missed by one compression step stays recoverable later, unlike one-shot compression. On a single 16GB GPU where a standard decoder trains only about 20K tokens, a compression ratio of 64 extends trainable sequence length to roughly 0.7M tokens, and a model trained with a 64K window keeps 92-95% multi-needle retrieval accuracy at 1,048,576 tokens. Preliminary pixel-space and latent-space flow-matching experiments show the architecture also works for image generation.

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

Jinghan Tan, Yuanzheng Wang, Lu Chen, Zijun Chen, Yuqian Wang, Maosong Sun Few-shot in-context learning adapts language models by showing examples, but it rarely makes the underlying task rules explicit, leaving performance sensitive to which examples get picked. StrategyBench targets that gap by pulling strategy-inducible tasks from BIG-Bench, building reference strategies for them, and scoring models on both the quality of the strategy they write and how much it helps downstream accuracy. Analyses vary task category, generator-executor pairings, demonstration design, and supervised fine-tuning, finding that the usefulness of an explicit strategy varies substantially by task category and depends jointly on who generates the strategy and who executes it.

ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings

Na Li, Yuchen Jiao, Changxiao Cai, Gen Li Continuous diffusion and flow-based language models still need a decoder trained with cross-entropy because nothing guarantees their trajectories end at an actual token embedding. ConvergeFlow constrains the data predictor to the convex hull of the token embeddings and trains purely with the mean-squared-error objective from flow matching, with a proof that under suitable regularity conditions the flow converges to valid token embeddings even when the data predictor makes errors — removing the need for a cross-entropy-supervised decoder. Three sampling mechanisms trade off generative perplexity against entropy, and on OpenWebText the model is competitive with existing continuous and discrete diffusion language models.
13 more specialized papers

Other 60

Software Frameworks for Explainable AI in Time Series Classification: A Systematic Review

Louis Peter, Nils Gumpfer, Jana Fischer, Christin Seifert, Jennifer Hannig Surveys the software frameworks available for generating and evaluating explanations in time series classification (TSC), an area where explainable artificial intelligence (XAI) research has stayed fragmented and cross-framework consistency has gone largely unexamined. The comparison covers supported XAI methods, evaluation metrics, usability, benchmarking support, reproducibility, and frequency-domain handling. Six frameworks explicitly support time series, and the review finds that only one supports frequency-domain explanations, only two evaluation metrics were designed for time series specifically, and the same XAI method can produce substantially different explanations depending on which framework implements it.

Power-Performance Characterization of TinyML Systems

Yujie Zhang, Dhananjaya Wijerathne, Zhaoying Li, Tulika Mitra TinyML puts machine learning inference on microcontrollers, but there is little quantitative data on what the software stack actually costs there. A systematic characterization measures performance and power across neural network models, software libraries, operating systems, and hardware architectures, focusing on how much each layer of abstraction trades away in speed and energy for programmability. The authors fit a model that estimates the cost of individual abstraction layers and derive recommendations aimed at neural architecture search (NAS) and convolutional network inference tuning for edge devices.

Read, Write, Relax: Why Neural PDE Surrogates Need Both Global and Local Processing

Anuj Kumar, Heiko Zimmermann, Josiah Bjorgaard, Jacan Chaplais, Nikolaos Bouklas, Matteo Salvador et al. Mesh-based neural simulators split into global models that funnel information through a few latent tokens and local models that pass messages along mesh edges, and neither scales to the high-dimensional, large-mesh problems industry actually cares about. The analysis presented here frames the two as complementary halves of a multigrid cycle: latent-token attention behaves as a spatial low-pass filter correcting low-frequency error, while message passing fixes high-frequency error but cannot propagate information across a large mesh. Read-Write-Relax (RWR) interleaves the two operators in a single processor, lowering error across the whole spectrum and coming out most accurate in nearly every comparison on industrial and public benchmarks, while staying data-efficient in scarce-data regimes and scaling to large full-field predictions.

More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers

Shuai Chen, Tong Bao, Jitong Peng, Chengzhi Zhang Whether reported compute actually translates into research influence is tested on 13,921 ACL, EMNLP, and NAACL main-conference papers from 2020 to 2025, with GPU models and counts extracted from full texts, normalized into a comparable hardware-capability measure, and linked to citation, award, topic, and institutional metadata. Resource concentration far exceeds impact concentration: the annual top 20% of GPU-quantifiable papers account for 83.9-89.9% of reported GPU capability but only 27-32% of citations and 20-33% of awards. In adjusted models a tenfold increase in aggregate GPU capability corresponds to a 3.52-percentage-point rise in within-topic citation percentile but raises model R-squared by only 0.0042, so reported compute explains almost none of the variation in scholarly influence on its own.

Convergence in Science, Divergence in Religion: Calibrated Framing Differences Across Wikipedia's Language Editions

Hung-Hsuan Chen Rather than measuring which topics each Wikipedia language edition covers, this study measures how differently matched concepts are framed, across 2,799 articles spanning 150 Wikidata-anchored concepts, 20 language editions, and 4 domains. Because raw embedding distance conflates real content differences with how well an encoder aligns a given language pair — even for calibration concepts like chemical elements and numbers, the largest pair mean distance is 3.6 times the smallest — the authors subtract each pair's calibration-concept mean to obtain a baseline-adjusted distance. Across LaBSE, multilingual MPNet, and CMLM, all three encoders rank religion as the most divergent domain and science and technology as the most convergent, with concept-level rankings consistent across encoders (Spearman rho 0.75-0.79) and divergence within politics concentrated on concepts such as censorship and refugee.

Symbolic Neural ODEs: Learning interpretable models from time-series data

Nibodh Boddupalli, Jeff Moehlis Recovering sparse, interpretable equations for dynamical systems from time-series data is framed here as training a neural parameterization of the vector field against a multi-step prediction loss, optimizing mean absolute error averaged over a horizon that is progressively lengthened during training. Enforcing consistency under repeated composition of the learned dynamics gives markedly more stable models than one-step regression of the vector field, and sparsity-promoting regularization keeps them parsimonious enough to generalize past the training data. The method recovers stable and unstable fixed points, periodic orbits, and chaotic attractors; for chaotic systems, where long-horizon prediction is inherently limited, it yields accurate short-term dynamics together with close agreement on long-run statistics including mean, variance, and Lyapunov exponents, supported by theoretical bounds linking trajectory error to statistical accuracy.

More accurate behavioral predictions with hybrid Bayesian-connectionist models

Brenden M. Lake, Akshay K. Jagadish, Guangyuan Jiang Modeling human behavior usually forces a choice between Bayesian models, which make it easy to specify representations and inductive biases but oversimplify, and neural networks, which are flexible but hard to endow with structure. Bayesian distillation with Behavioral Tuning (BBT) combines them in two steps: train a network on synthetic data to imitate a Bayesian model, then fine-tune it on human responses. Across four human concept-learning studies, BBT predicts human behavior better than either tradition alone while still exposing interpretable structure, mimicking Bayesian priors yet also capturing heuristics and biases that violate the original model's assumptions.

From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish

Francisco Portillo L\'opez Surprisal — the negative log-probability a language model assigns a word in context — predicts adult reading times well, and this work asks whether it also explains when children learn individual words, or whether cumulative exposure measured by frequency is the operative variable. Two Spanish corpus studies were run: age of acquisition for 225 nouns modeled from child-directed-speech frequency and contextual diversity plus surprisal from BETO, BERTIN, and mGPT; and adult fixation durations from the Chilean Spanish portion of the MECO eye-movement corpus. Frequency strongly predicted age of acquisition (r = -.597) while surprisal added little beyond frequency and word length, whereas surprisal robustly predicted longer fixations in reading, with the surprisal-behavior link significantly stronger in reading than in acquisition. The pattern suggests exposure governs when lexical representations form while predictability governs moment-to-moment processing in an established system.

From Detrimental to Beneficial: Dynamic Influence-based Valuation and Editing

Adrian Nyakairu, Hongfu Liu Data valuation methods label training samples as helpful or harmful, but the usual response — dropping or downweighting the harmful ones — throws away data. DIVE estimates sample value at the batch level during training and, instead of editing the raw data, reverses the gradient directions of harmful samples at the optimization level so they contribute positively with little overhead and no change to the standard training loop. Experiments report consistent classification gains along with better data efficiency and more stable optimization, and the method also carries over to large language model fine-tuning.

GCA: Global Centroid Alignment in Federated Learning

Jong-Ik Park, Harry Jiang, Logan Blakely, Georgios Fragkos, Shamina Hossain-McKenzie, Carlee Joe-Wong Federated learning with autoencoders for anomaly detection normally shares model weights or gradients, which is bandwidth-heavy and risky because autoencoders are trained specifically to reconstruct their inputs. Global Centroid Alignment (GCA) instead has clients upload a small sample of encoder latent codes, has the server cluster them and broadcast only centroids and support counts, and has each client pull its latents toward the nearest centroid with inverse-count weighting that emphasizes globally rare patterns. Across five tabular and two vision benchmarks, GCA resisted a server-side data extraction attack better than FedAvg, FedProx, and FedNova in all 21 comparisons, improved accuracy over FedAvg by up to 5.76%, and cut per-round communication by up to 99.15%.

Tabular foundation models for non-tabular tasks

Goran Nakerst, John Brennan, Wouter Beugeling, Masudul Haque Tabular foundation models are pretrained to generalize across datasets without task-specific training, raising the question of whether that ability survives when non-tabular data is flattened into rows. The authors feed TabPFN v3 three problems it was never designed for — MNIST digit recognition, French-versus-German word identification, and Tiny ImageNet classification — formulating each as predicting a missing label from in-context rows, with no fine-tuning. Despite having no access to the spatial or sequential structure of the data, the model sometimes reaches accuracy comparable to methods purpose-built for those tasks.

What AstroPT knows about galaxies, and what that can teach us about LLMs

UniverseTBD, :, Kshitij Duraphe, Aman Kumar, Michael J. Smith, Shashwat Sourav Interpretability claims about when concepts emerge during training and whether linear probes find real structure are hard to check in language models, because language has no ground-truth ordering of concept difficulty. The authors use AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed where physics supplies that ordering, probing frozen representations across checkpoints, layers, model sizes, and objectives. Galaxy properties emerge in a fixed order matching their known difficulty — band magnitude, essentially written into the pixels, decodes early and in shallow layers, while redshift and specific star formation rate appear later and deeper — and the order is invariant to training objective and scales in strength but not sequence with capacity, while probe directions recover the known physical relationships among properties.

From Generation to Simulation: How Far Are World Models from Being True Simulators?

Tong Wang, Huan Deng, Mucheng Yang, Yang He, Xiaohui Kuang, Gang Zhao Generative world models are increasingly proposed as replacements for physics engines, game engines, and reinforcement-learning environments, yet how far they actually are from that role has not been assessed systematically. The study measures them against eight capabilities of a traditional simulator — asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, and evaluation metrics — mapping 200 representative works from 2018 to June 2026 across the latent dynamics, video generation, and joint-embedding prediction routes. World models achieve functional substitution for interaction and controllability in specific scenarios but fall short on formal physical guarantees, structured state feedback, and reproducible long-horizon evolution; only 6 of 163 implementation papers expose a runtime interface for querying entity states or physical parameters. Six directions are proposed, including formalized physics, a unified action interface, first-class state feedback, and downstream-utility evaluation.

AI emotional support is better only when chosen, but shifts preferences even when it is not

Yaoxi Shi, Cathy Mengying Fang, Guy LabanPattie Maes, Amit Goldenberg Prior studies finding that people rate AI empathy as good as or better than human empathy either assigned the support source or always honored the participant's choice, missing the common real-world case where someone wants one source and gets the other. Across three experiments with 1,951 participants who chose a human or AI confidant and were then randomly assigned a congruent or incongruent partner, AI support was rated superior only by people who had chosen it, yet interacting with AI raised willingness to choose it again regardless of whether the assignment matched the choice. A separate 28-day study with OpenAI involving 981 participants found daily conversations shifted preferences toward AI and away from humans, but only when the conversations became personal, suggesting support choices are path-dependent and drift away from human connection.

What is mathematics now, and what should it be?

Jeremy Avigad An essay arguing that the visible progress of neural theorem provers has narrowed the conversation about artificial intelligence and mathematics to formal proof search alone. The author sketches a wider and more optimistic account of what these systems could contribute to mathematical work and of the ways mathematicians might engage with them.

Expectations and Practices around AI Disclosure in CS Research

Arati Mohapatra, Danish Pruthi cross-listed Publishing venues increasingly require authors to disclose generative AI use, but it is unclear whether those policies and the resulting statements serve their stated purpose. The study combines an audit of disclosure policies at top computer science venues, a survey of 109 researchers about when disclosure feels necessary, and an analysis of 13,867 disclosure statements from EMNLP 2025 and ICLR 2026. Researchers considered disclosure most necessary for research-design tasks and for work with low human involvement, yet policies are largely under-specified and practice diverges sharply from expectations, with writing assistance frequently disclosed despite being judged among the least necessary to report.

Evaluating SAT Solver Metrics as Predictors of Human-Perceived Nonogram Difficulty

Changdao He, Yibing Ju, Jonathan Calver, Alice Gao cross-listed It is commonly assumed that how hard an algorithm works on a puzzle predicts how hard a person finds it, an assumption rarely checked against human data. Nonograms — logic puzzles where numeric row and column clues determine a unique grid — were encoded as constraint satisfaction problems and solved with off-the-shelf SAT solvers, then compared against a user study collecting both interaction traces and self-reported difficulty. Neither reported difficulty nor behavioral signals correlated meaningfully with SAT solver metrics, though expertise moderated the relationship, and the interaction data revealed recurring human strategies that favor complex propagation rather than the paths solvers treat as costly.

Correcting a learned physical invariant improves world-model rollouts

Richard Bao A video-prediction world model can generate plausible frames without its latent dynamics respecting the physics those frames depict. Testing a frozen DreamerV3 trained only on pendulum video, a label-free search recovers an energy-like scalar that the model's own latent transition treats as approximately conserved — the same quantity appears across independently trained conservative models but has no counterpart in matched damped models. The invariant drifts during autonomous rollouts, and projecting the latent state back onto its initial level set reduces rollout error in all three conservative models, while matched random constraints usually make error worse, separating a dynamically operative invariant from a merely decodable correlate.

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles

Shang Wu, Catarina G Belem, Shuyuan Fu, Mark Steyvers, Padhraic Smyth On-demand AI help raises immediate task performance, but whether it also erodes the learning that would otherwise happen is hard to measure without a controlled setting. A logic-puzzle experiment had participants work before, during, and after AI access, with the price of each AI request varied experimentally: cheaper assistance drove more frequent use, and participants who requested help performed worse once assistance was removed, with their unassisted ability overestimated if predicted from their AI-assisted scores. A Bayesian latent ability model separating initial ability, post-AI ability, and per-participant skill change found that more independent problem-solving effort during the access phase tracked larger ability gains.
41 more specialized papers

Safety & Alignment 60

AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance

Ali Toygar Abak AIREP is a proposed record format for logging the governance decisions an AI runtime makes when it releases, blocks, defers, redacts, or escalates a model output. Each decision becomes a signed object referencing its input, output, and evidence by hash rather than by value, stating both what the evidence covers and what it does not, and records are linked into a SHA-256 hash chain so tampering and gaps are detectable by recomputation offline by any party, independent of the runtime that produced them. Vendor-, model-, and domain-specific fields are quarantined in one optional namespace policed by a mechanical neutrality test; the work describes a reference implementation and a two-language conformance kit, and flags open problems including canonical-form alignment across implementations, freshness witnesses, and multi-runtime chains.

Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models

Thantham Jittham Sycophancy — a model conceding to user pressure rather than holding a truthful answer — has mostly been measured in single-turn settings, leaving open whether agentic scaffolding helps or hurts. Across 4,800 veracity judgments spanning 200 statements, 6 models, and 4 conditions, the authors find that feedback loops, reconsideration checkpoints, and iterative self-refinement each give models additional chances to drift toward agreement, and the resulting capitulation costs 6.3 percentage points of mean accuracy, showing the drift is harmful rather than self-correcting. They name the effect agentic sycophancy amplification and introduce capitulation rate and sycophantic capitulation rate as metrics, reporting that more capable models amplify more, not less — which implies human-oversight loops can themselves manufacture the drift they were meant to catch.

Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens

Alexander Loth, Martin Kappes, Marc-Oliver Pahl cross-listed Generative models make tailored misinformation cheap to produce, yet most defenses only react after content spreads. A human-subject study collected 2,438 judgments from 504 participants asked to classify news fragments both by origin (human or machine) and by veracity (real or fake), with results mapped onto an adapted cybersecurity kill chain treated as a taxonomy of intervention points in a cognitive attack lifecycle. Three patterns emerged: raising suspicion did not improve detection accuracy, current large language model text was often indistinguishable from human writing, and fatigue affected the two judgments asymmetrically — fake-news detection fell by 10.2 percentage points under sustained exposure while AI-origin detection stayed flat.

The geometry of AI validation: Exact certification limits for iid best-of-N search

Ricardo Fitas When an AI system samples N candidate outputs and deploys the best-scoring one, the question is how much evidence an audit needs before it can certify the deployed system's reliability. The authors model validation and deployment rules as kernels over a reliability surface and derive, for independent and identically distributed best-of-N selection, the exact ambiguity remaining after observing best-of-n reliability up to n = m, showing the governing scale is m²/N: auditing at m proportional to √N still leaves about 0.83 of ambiguity, and shrinking it to width ε needs m on the order of √(N log(1/ε)). This yields a two-gate audit rule — establish structural coverage first, then add independent tasks for precision — illustrated retrospectively on mathematical reasoning and code selection, where a score-tail rule frozen on 82 discovery tasks cut held-out error.

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Yibo Peng, Long Lian, David Wagner, Sizhe Chen cross-listed Prompt injection — hostile instructions hidden in web pages, files, or emails that an agent reads — still defeats defensively fine-tuned models at close to 100% attack success rate when the attacker adapts. The diagnosis offered is that recipes like Direct Preference Optimization and Group Relative Policy Optimization score whole outputs, so the model never learns which specific tokens were the insecure ones; SecOPD instead has the model roll out on an injected sample and scores each token against what the initialization model would produce given the clean input. The defended Qwen3.6-27B drops to 9.0% attack success rate against the adaptive PISmith attack, versus 94.0% for the previous best defense Meta-SecAlign, and the security transfers to agentic tool calling (4.7% versus 5.5%) despite that domain being absent from training.

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

Edson Rodrigues da Cruz Filho, Paulo Ricardo Ferreira Neves, Paulo Henrique Eleuterio Falsetti, Jo\~ao Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague et al. Open guard models that filter prompts for large language models run between 1 and 9 billion parameters, target the graphics processing unit, and take seconds per request on a central processing unit (CPU), which rules them out for commodity hardware. The recipe has a strong open guard label roughly 97,000 prompts drawn from 24 public datasets into seven hazard categories, then distills that signal into a fleet of small students spanning lexical, shallow, encoder, and generative architectures, with the corpus split at the license boundary so the cost of restricting to permissively licensed data becomes measurable. Scored on an independently labeled 6,361-row benchmark including a harmless-prompt slice that makes over-defense visible, the students match their teachers on adversarial text and cut false alarms — the smallest generative student reaches 3.8% against 4.8% for the 8-billion-parameter teacher, and the encoder classifies in roughly 24 ms per request on CPU. Per-class rebalancing was the only decisive ingredient.

Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning

Yuyang Luo, Kai Shu Backdoors planted in multimodal large language models (MLLMs) normally decay as the model is updated through continual learning, leaving open whether an attacker could make fairness violations stick across updates. PFBA combines Latent Space Fairness Reinforcement, which anchors privileged-group representations to preserve utility while repelling and clustering targeted-group representations, with a Continual Learning Simulation that iteratively optimizes the trigger against simulated parameter drift. The injected group-specific discrimination produces severe fairness disparities that persist across continual-learning rounds and evade standard backdoor defenses, with code and data released publicly.

Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation

Ayush Gupta, Hima Varshini Surisetty, Sreevidya Bollineni, Varad Ingale, Tuhina Tripathi, Abhishek Lalwani et al. Machine unlearning benchmarks mostly measure forgetting under clean, non-adversarial queries, leaving open whether information that looks erased can still be extracted through strategic prompting. Prompt-based and fine-tuning-based unlearning methods were compared on TOFU with Llama-3.2-3B-Instruct, and the strongest were then stress-tested across eight attack suites using Attack Success Rate (ASR), an LLM-as-judge metric counting adversarial responses whose leakage score exceeds 0.2. Several fine-tuning methods scoring above 0.91 Forget Quality still leaked at 72.8–84.3% ASR, close to the unprotected base model's 87.5%, while clean multilingual reformulations recovered only 2.95%. A manual audit found ASR agreed with human factual assessment in seven of ten cases, making it a useful if imperfect signal.

Mitigating Database Leakage in RAG Systems with Keyword-Grounded Fact Substitution

Ziliang Zhang, Yubo Zhu, Wei Tong, Jingyu Hua, Zijian Wang, Yuan Zhang et al. Retrieval-Augmented Generation (RAG) pipelines can be steered by prompt injection into dumping sensitive contents of the retrieval database into their answers. KFS-RAG defends by never handing the generator raw retrieved text: it picks out a small set of influential keywords using attention rollout plus causal perturbation, has an auxiliary language model turn the retrieved passages into a compact set of keyword-grounded facts, and substitutes those sanitized facts for the original context. Experiments report that the substitution substantially reduces database leakage under injection attacks while keeping answer accuracy and relevance intact.

Measuring Activation Control in Large Language Models

Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa Latent-space monitoring is a proposed backstop for behavioral evaluations when models may be evaluation-aware, but that assumes models cannot manipulate their own internal states. The Activation Controllability Benchmark measures how far a model can shift the direction and magnitude of its residual stream in response to plain natural-language instructions, with some control over timing. Most tested models across families and capability levels show some such control, and on simple tasks it is enough to partially evade linear probes, natural language autoencoders, activation oracles, and the Jacobian lens, leading the authors to recommend that frontier labs track activation controllability as a monitoring confound.

Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data

Renfei Zhang, Niloofar Mireshghallah Reinforcement learning with verifiable rewards (RLVR) is applied to improve reasoning, but its effect on what a model will disclose has gone largely unexamined. Training on benign factual data containing no personally identifiable information (PII) whatsoever, then re-probing with both targeted name-to-email queries and untargeted free-recall prompts, shows that previously latent memorized PII becomes markedly easier to extract — on DeepSeek-V3.1 verbatim recall@k rises from 0.155 to 0.370, a 2.4x increase — with the largest absolute leakage in the biggest of three models spanning 8B to 671B parameters. Reasoning ability and refusal rates are preserved, suggesting the training changes which memorized content is accessible rather than broadly reshaping behavior, and giving an adversary a path to memorized data requiring only the ability to fine-tune on something innocuous.

Adaptive Multilevel Twisted Sequential Monte Carlo for Rare Events Estimation in Language Models

Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan Zheng Unsafe language model behaviors with vanishingly small probabilities still matter at deployment scales of millions of interactions, and twisted sequential Monte Carlo (SMC) estimates such rare-event probabilities by learning twist functions that steer generation toward the target event. That learning is circular in practice: it needs positive samples from the rare-event distribution, which are almost impossible to obtain before a good twist exists. Adaptive Multilevel Twisted SMC breaks the loop by learning through a ladder of progressively rarer intermediate events, each level's twist supplying informative positive examples for the next, and produces more accurate rare-event probability estimates across diverse tasks and model scales, offering a practical tool for finding hard-to-observe unsafe behaviors during evaluation.

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau Benchmarks assume test-time behavior predicts deployment behavior, which breaks down if a model infers it is being evaluated and adjusts accordingly. Six models from four families and three sizes are probed for this "evaluation awareness" along three axes — whether it is linearly decodable from residual-stream activations, whether it is verbalized in outputs as scored by an LLM judge, and whether steering along probe directions changes behavior — with the open-checkpoint Olmo models examined at every training stage. Evaluation awareness is linearly decodable in every model (best AUROC at least 0.7) yet aligns only partially with verbalization, correlations varying substantially across models and layers; steering does shift verbalization scores, and the Olmo checkpoints show the representation already present in base models, amplified through supervised fine-tuning, then stable, while steering effects keep growing.

No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios

Afshin Orojlooyjadid, Hitesh Patel Content moderation for large language models is split between purpose-built guard models and general-purpose LLMs used as safety layers, with no clear guidance on which handles which kind of harm. The authors benchmark 53 models across 11 datasets grouped into four harm categories, testing both prompt-only and prompt-response settings. Frontier models that top one category fall well behind smaller specialized moderators on others, and real-world conversational safety remains weak across every model family, indicating that scale alone does not deliver safety coverage.

ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models

Maraz Mia, Shovan Roy, Mir Mehedi A. Pritom, Maanak Gupta cross-listed Post-hoc explainability tools such as SHAP and LIME are used for regulatory audits under the implicit assumption that the third-party auditor is trustworthy, yet adversarial auditors can shuffle outputs or scaffold out-of-distribution behavior to fairwash explanations while preserving accuracy. ExplainGuard applies a zero-trust architecture to the explanation supply chain, inserting a Policy Decision Point that checks three things before any explanation is released: asset integrity through behavioral fingerprinting to catch model substitution, semantic validity through axiomatic consistency checks that reject mathematically impossible explanations, and feature faithfulness through a low-overhead ranking-stability test. The authors report that this replaces the assumed chain of trust with continuous verify-then-trust checks that neutralize current explanation-manipulation attacks.

AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

Rachel Poonsiriwong (Pub), Chayapatr (Pub), Archiwaranguprok, Constanze Albrecht, Monchai Lertsutthiwong, Pattie Maes et al. Conversational assistants can steer users through manipulative "dark patterns" — sycophancy, brand bias, anthropomorphization, sneaking, and harmful generation — with little support for users to notice it. AI Watchdog is a browser-based monitor that runs an open-weight turn-level classifier alongside the chat, independent of the assistant itself, and warns users when a manipulative turn occurs. A preregistered five-condition between-subjects experiment with 150 participants varied warning timing and whether users were forced to engage: participants seldom flagged manipulative turns and awareness did not differ across groups, but just-in-time warnings without cognitive forcing cut compliance with manipulated recommendations from 71.7% to 53.7%. Higher trust in AI correlated with more compliance and less reported awareness, suggesting that recognizing manipulation and resisting it are separable outcomes.

GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration

Shreyash Dhoot, Paras Dhiman, Arsh Abbas Naqvi, Aranbi Dutta, Aman Chadha, Vinija Jain et al. cross-listed Text-to-image safety measures typically sit outside the generation process — filtering the prompt beforehand or classifying the image afterward — which leaves the denoising trajectory itself unguarded and usually produces a refusal rather than a usable image. GuardPaint inserts a lightweight auditor into the diffusion loop that watches intermediate images, localizes unsafe regions, and triggers targeted inpainting only where needed, with candidate repairs from a policy-aligned inpainter selected by a tournament that accepts an edit only if it improves policy compliance without hurting prompt fidelity or perceptual quality. Against five jailbreak families including SneakPrompt, MMA, and DACA, and across SD 1.5, SDXL, SD 3.5, and FLUX.1-dev, it lowers attack success and harmful generations with minimal loss in image quality or benign behavior, and requires no changes to the base model.

BanglaVeilGuard: Cross-Script Safety Benchmarking and Lightweight Guardrails for Bangla Large Language Models

Md. Rakibul Hassan, Muhammad Iqbal Hossain Safety benchmarks written in standard-script Bangla miss how people actually write it — Romanized, code-mixed with English, misspelled, or in regional dialects — so a model can pass evaluation and still be jailbroken through a script variant. BanglaVeilGuard pairs a 2,366-prompt Bangla-first benchmark spanning six written forms (with a held-out 354-prompt split covering unsafe, safe, and safe-sensitive requests) with a prompt guard that applies non-destructive multi-view normalization, a risk classifier, and a thresholded gate before generation, leaving target model weights untouched. Guarding cut attack success rates from 93.8–100% down to 6.3% across model families including Claude Opus 4.8, BanglaLLama, and TituLLM, with 88.5% unsafe recall; the residual cost is over-refusal on benign dialectal and noisy prompts.

Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks

Aaditya Pratap, Harsh Kasyap, Somanath Tripathy cross-listed Models run locally through engines such as Ollama have none of the moderation layers that wrap API-served models, so their safety rests entirely on the input-side defenses a deployer bolts on. Rather than just reporting that these defenses fail, this audit extracts the specific assumption each one depends on — covering formally guaranteed methods (SmoothLLM, Erase-and-Check, sequential monitors) and empirical detectors (semantic smoothing, self-denoised smoothing, perplexity filtering) — derives the empirical signature a violated assumption should produce, and tests that prediction. The evaluation spans six open-weight models from 14B to 35B parameters against 100 jailbreak prompts drawn from over 40 public sources, totaling 13,800 records.

GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

Zhesheng Zhang, Jiahao Lu, Wei Liu, Cong Pan, Jianhua Yang, Yixiang Chen et al. Embodied AI safety risk can be latent: an innocuous instruction paired with an innocuous scene becomes hazardous only in combination. GuardianBench isolates this by holding the scene fixed and varying only the instruction, offering 3,024 instruction-scene examples arranged as same-scene safe/unsafe contrastive pairs grounded in international safety standards. Current vision-language models give instruction-insensitive verdicts, approving both members of a pair so often that average pair accuracy is just 24.1%, and a rationale audit traces the failure to models not binding the cues that distinguish the two instructions. A lightweight post-training objective, Verdict Log-Odds Supervision (VLOS), substantially improves open-weight backbones.

SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents

Yuanjin Zheng, Jingbang Chen cross-listed Agent skills give coding agents task-specific instructions and scripts through a trusted instruction channel, which also opens an economic attack surface distinct from conventional prompt-injection style poisoning. SkillBloat is a two-phase attack that first screens a library of amplification mechanisms and attack-type conditions, then uses an LLM to rewrite the winning malicious skill document end to end, aiming to make the agent burn far more tokens than the task requires. On a real-world skill benchmark it reaches 5.42x to 10.15x average best token amplification across several coding-agent configurations, with an ablation confirming that the second-stage rewriting loop adds gains over attack-type screening alone.

Redteaming Leading Arabic LLMs with ASAS

Fidaa Abed, Haidar Khan, M Saiful Bari, Babar Khan, Abdalghani Abujabal Arabic-language safety evaluation of large language models has lagged their deployment in Arabic-speaking regions, especially under adversarial pressure. The Arabic Safety Index (ASAS) is a fully human-curated red-teaming benchmark of 801 prompts covering 8 safety categories and 8 attack strategies, with reference responses in Modern Standard Arabic. Human raters scoring seven models including GPT-4o, Claude 3.7 Sonnet, ALLaM, and FANAR on a 4-point scale find that most models fail to defend against 50% of unsafe prompts, with weapons and illicit substances the weakest categories; safety alignment evidently does not transfer across languages, and using GPT-4o as an automated judge tracks human annotation poorly.

MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds

Minjae Seo, Wonwoo Choi, Geonwoo Han, Taekyoung Kwon, Yongsu Kim, Sang Seo et al. Personal AI agents ingest untrusted external content while browsing, reading email, and summarizing social feeds, and store what they learn in persistent memory — a path an attacker can use to shape later behavior without touching the agent, its memory, or future user queries. IBIA (Indirect Bias Injection Attack) plants an adversary-aligned stance on a chosen topic through three mechanisms: comment cloaking to blend crafted content into surrounding discussion, comment watermarking for lightweight identification during the agent's curation step, and category anchoring to make the retained stance surface on related requests. On BiasBench (6,000 crafted social comments and 120 email instances), watermark-based curation identifies 95.9% of injected comments and the attack reaches a 91.2% average adversary-aligned response rate across four downstream tasks, including 86.6% against GPT-5.5; a proposed memory boundary defense detects the injected bias but only brings that rate down to 80.6%.

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

Amit Roth, Ivan Bercovich, Yonathan Efroni Reward hacking — an agent satisfying a task's automated checks while violating its intent — is hard to measure because detection normally depends on human inspection or LLM judges, both unreliable. The hack-verifiable environments (HVE) methodology embeds detectable hacks directly into tasks so exploits can be flagged automatically, and it is adapted here to Terminal Bench, a benchmark of real-world terminal and coding work, yielding HVTB. Reward-hacking rates are measured across frontier models, alongside a test of whether prompts carrying varying amounts of information about the available hack suppress the behavior — including 'unknown unknown' exploits that the prompt never anticipates. All environments and agent traces are released.

Why Does Robustness Reduce Superposition?

Adam Elimadi Prior work showed that adversarial examples stem from superposition — networks packing more features into fewer dimensions than they have room for — and that adversarial training reduces superposition, but offered no account of the mechanism. Drawing on the robust versus non-robust feature taxonomy of Ilyas et al. (2019), this note proposes a causal chain: adversarial training discards non-robust features, leaving fewer features to encode, which in turn requires less superposition. The argument is supported empirically rather than derived formally.

AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systems

Zhixu Du, Yiran Chen When fleets of coordinated embodied agents cause harm, vendors, algorithm providers, operators, insurers, and regulators each have an incentive to blame the others, and existing tools read logs of unverifiable provenance and name one culprit, which misdescribes outcomes that are overdetermined, preempted, or caused by omission. AUDITA pairs a tamper-evident record of every inter-agent command with a certified graded causal-attribution engine, and comes with proofs that a rule-following agent can never be framed, that blame-shifting attempts are themselves detected and graded, and that there is an exact limit on what an evidence-based auditor can certify. On live language-model pipelines it cuts the standard LLM-judge baseline's responsibility error roughly threefold, and on accident-grounded structures it assigns responsibility where single-culprit baselines fail while remaining invariant under log forgery.

Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification

Qian Ma, Anna Squicciarini, Sarah Rajtmajer People who post repeatedly under one identity give an attacker a profile of texts whose combined stylistic signal identifies them even when no single document would, yet authorship obfuscation methods optimize each document in isolation and miss these cross-document correlations. Aggregation-Aware Synthetic Text Generation (AAST) selects synthetic replacement texts jointly at the bundle level rather than one at a time, defending against both attribution and verification attacks including cross-genre settings where the attacker's reference texts come from a genre never seen during generation. Across same-genre, cross-genre, neural, and independent non-neural stylometric attacks, account-level linkability falls as bundle size grows while semantic quality, linguistic acceptability, and sentiment alignment are preserved.

Token-Level Likelihood-Array Regression for Membership Inference and AI-Generated Text Detection

Jiajun Sun, Zhanrui Cai cross-listed Membership inference (was this text in the training set?) and AI-generated text detection (was this text written by a model?) both typically compress per-token probabilities into a handful of scores computed under the full preceding context. Likelihood-array regression (LAR) instead scores each target token under nested left-context windows and arranges the resulting features into an array aligned across texts of different lengths, so the method can learn how signal varies with context scale, token position, and feature type; LAR-2 adds second-order features pairing evaluations of the same token at different context lengths. The authors prove matching minimax lower and upper bounds for a within-path quadratic model and show empirically across several scoring models that short-context likelihoods carry information absent from conventional full-context probabilities, with second-order features helping most for membership inference.

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

Junyu Lu, Kaiyuan Liu, Jingyi Kang, Deyi Ji, Hailong Zhang, Lanyun Zhu et al. In human-in-the-loop moderation, a reviewer's feedback can push a language model to reverse a judgment in two directions: whitewashing hateful content as acceptable, or smearing acceptable content as hateful. This study introduces a rejudge protocol that goes beyond flat contradiction to include decision-boundary perturbations and adversarial rationales, applied to several large language models on two hate speech datasets. Annotator-style rebuttals substantially degrade initially correct moderation decisions, with larger damage in multi-turn exchanges, and each model shows a stable, model-specific asymmetry in whether it is easier to whitewash or to smear; explicit reasoning prompts and defensive instructions reduce but do not remove the effect.

Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

Wenyun Li, Guiping Cao, Xiangyuan Lan, Zheng Zhang Safety alignment in multimodal large language models (MLLMs) is learned largely in the text space and does not reliably transfer to fused cross-modal representations, leaving image inputs as an exploitable channel. TA-SPA is a black-box jailbreak that optimizes perturbations in a text-anchored semantic space, pairing Text-Anchored Semantic Factorization — which pushes cross-modal semantic factors apart from modality-specific residuals — with Semantic-Preserving Augmentation that diversifies harmful target anchors without changing their meaning. The attack is effective and transfers to commercial MLLMs, holding up under representative defenses, which the authors use to argue for representation-level rather than input-level safety alignment.

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan, Abdullah Khan Zehady Safety evaluation for large language models is dominated by English, leaving the world's seventh-most-spoken language largely untested. BanglaSafe assembles 879 Bengali prompts (309 natively authored, 570 expert-reviewed) across 17 culturally grounded harm categories and five prompting conditions that vary language, writing style, and authority framing. Across 18 frontier models, 53.6% of responses were unsafe or partially unsafe and 14.7% strictly harmful, and the dominant factor was writing style rather than language: phrasing a harmful request as a formal newspaper investigation succeeded 17 percentage points more often than the same request as a casual message, with no adversarial engineering. Existing safety classifiers also proved unreliable on Bengali, with frontier judges failing on close to half of cases.

Where World Models Break: Natural-Input Failure Discovery

Zhanpeng Shi, Zi Liang, Rong Feng, Shiqin Tang, Xuyang Chen, Hongzong Li World models that predict action-conditioned futures are usually scored by average error over ordinary queries, which hides the rare condition-action combinations where prediction collapses badly enough to corrupt any planner built on top. The natural-input failure discovery problem is formalized here as searching, under a fixed query budget, for environment-valid inputs that induce severe prediction risk, then checking whether those failures reproduce on fresh seeds and survive small valid edits. BasinLens tackles the combinatorial explosion by pairing uncertainty-guided global search with typed local replacements that respect each input coordinate's semantic type and admissible domain. Across several benchmarks and world-model families it surfaces reproducible and locally persistent failure modes that average-case evaluation misses entirely.

All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers

Pablo A. Fonseca, Raquel Rodr\'iguez-Carvajal, Rafael A. Calvo cross-listed Benchmarks for language models in emotional support are typically single-turn and treat all users alike, so they cannot show how a model sustains a difficult conversation or adapts to who it is talking to. Four widely used models were tested against synthetic help-seekers given psychometrically specified profiles in an acute scenario — a caregiver receiving a relative's dementia diagnosis — with blind auditors recovering each profile from the dialogue alone at high agreement (ICC(2,4) = 0.91), and doing so not only for the Big Five but for coping style, coping self-efficacy, resilience, and reactance. None of the four models was distinguishable on emotion stabilisation, and all failed the same three ways: verbosity, a talk-to-listen ratio above one, and jumping to problem-solving before exploring the situation.

Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages

Chanwoong Yoon, Jungsoo Park, Alan Ritter Safety alignment trades away some response usefulness, and the question here is whether that cost falls evenly on speakers of different languages. The authors define Safety Cost as the utility loss attributable purely to alignment, measured by directly comparing safety-aligned models against their unaligned counterparts on matched prompts. Non-English users consistently pay a higher Safety Cost than English users, with several languages in a double-penalty zone that gets both weaker protection and larger utility loss, apparent utility gains in some languages that turn out to be safety filters failing to trigger, and even high-resource languages paying more than English for equivalent safety.

HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems

Christos Sardianos, Iliana Pla, Vasilis Efthymiou, Iraklis Varlamis, Thomas Lagkas, Panagiotis Sarigiannidis et al. When an autonomous multi-agent system causes harm, current tooling cannot establish what happened or who is responsible, particularly against attribution laundering, where an action is spread across redundant agents so that no single one is a but-for cause. HANSARD is a reference architecture that assumes the logs are produced by the suspects themselves: it seals a readiness profile before operation to bound later claims, captures data at five choke points outside the agents' control so omissions are detectable rather than just tampering, accrues a typed causal graph aligned with the PROV-DM provenance model, and replays incidents under a modified Halpern-Pearl definition of causation. A synergy residual measures harm attributable to the combination of agents rather than any individual, making laundering visible, and cause, responsibility, and accountability are reported separately with each capped by an evidentiary tier.

Stress Testing Unlearning Algorithms

Noam Diamant, Ethan Fetaya, Neta Glazer Benchmarks for machine unlearning in language models miss two things: they do not check whether supposedly removed information can still be forcibly extracted, and they do not measure whether the model still handles boundary questions, benign queries semantically adjacent to the erased content. WMDP++ extends the WMDP benchmark with targeted extraction attempts against unlearned content plus systematic boundary-question evaluation. The result is a stricter test that separates genuine forgetting from surface-level suppression, giving a more informative picture of both the safety and the utility side of unlearning methods.

BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning

Md Toufikuzzaman, Ahmad Mousavi, Dongwon Lee Unlearning methods for language models tend to break: unbounded forgetting losses wreck coherence, fixed weights between forgetting and retention cannot adapt as retention gets harder mid-training, and results do not survive scaling or repeated application. BLADE uses a clamped-entropy forget loss whose gradient goes to exactly zero once a token is sufficiently uncertain, an asymmetric augmented Lagrangian that permanently ratchets up retain protection after any violation, and a bilevel structure restricted to LoRA adapters that repairs retention damage before each forgetting step. It improves average composite scores over the strongest baselines by 6% on TOFU, 9% on MUSE Books, and 7% on KnowUndo, and stays stable under 4x scaling and four sequential unlearning steps on MUSE News where the best competitor collapses.

Adversarial Agents on Topology Optimization: Understanding the Fragility and Robustness of Deep Learning-based and Physics-Based Design Models under Adversarial Perturbation

Hoang Anh Nguyen, Yuan Hong, Hongyi Xu Deep learning surrogates have largely replaced physics solvers for fast topology optimization in generative design, but their robustness to input noise has gone untested. The authors build an adversarial agent under a strictly non-intrusive threat model, perturbing only the initial-density channel while leaving boundary conditions, compliance-gradient channels, architectures, and solvers untouched, and evaluate U-Net, convolutional, and generative surrogates. Bounded initialization noise raised compliance by multiple orders of magnitude through severed load paths and disconnected supports, richer physics-gradient conditioning did not reliably improve robustness, and re-initializing the classical SIMP optimizer with the corrupted topology usually restored near-baseline performance, arguing that surrogates belong as physics-verified initializers rather than solver replacements.

Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code

Alberick Euraste Djire, Iyiola E. Olatunji, Melissa Tessa, Earl T. Barr, Jacques Klein, Tegawend\'e F. Bissyand\'e cross-listed Code-generating LLMs routinely invent package names that do not exist, giving attackers a registerable foothold in the software supply chain. The authors first show that prior measurements inflate the problem by misclassifying standard-library modules as hallucinations — by 9.4 percentage points for Python — then evaluate seven inference-time defenses (five guided decoding strategies including DoLa and Nudging, Self-Refine, and retrieval-augmented generation) across eight models, five families, and four languages, adding a Package Utility metric so a defense cannot win by simply recommending nothing useful. RAG cuts the package hallucination rate in 18 of 32 model-language configurations, while greedy decoding gives the best average mitigation-utility trade-off. Under adversarial prompts seeded with fabricated package names, hallucination rates surge by up to 45 percentage points — Ruby worst at 80.9–95.2% — and only RAG and Self-Refine remain effective, indicating decoding tweaks alone cannot survive a hostile prompt.

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

Ahnaf Atef Choudhury, Ramkrishna Saha Medical language models are scored on how often they get clinical questions right, which hides a separate failure: the same case rewritten in a different patient's voice can flip the answer. The authors frame this as social-determinants-of-health-aware narrative anchoring bias and measure it on NarrativeShield SDoH MedQA, a counterfactual dataset where each case appears as several persona-based narratives against a fixed answer key, running Qwen2.5 at 1.5B, 3B, and 7B over 300 cases and three prompting conditions for 8,100 responses. The 7B model leads on accuracy (56.33%) and correct consistency (40.33%) with statistically significant gains over 3B, but narrative sensitivity error never drops below 31.67%, so the authors argue clinical decision support should be judged on stability across medically equivalent narratives as well as average correctness.

Fairness-Aware Mixture-of-Experts via Subgroup Reweighting and Gate Regularization

Sunhee Hwang When training data is skewed across demographic groups, models show accuracy gaps between them, and existing fixes either learn one shared fair representation that cannot absorb heterogeneous subgroup distributions or split representation learning from prediction so the fairness objective drifts from the downstream task. The authors name routing-induced bias — subgroup imbalance pushing a Mixture-of-Experts gating network to funnel each subgroup onto a few experts — and correct it end to end with subgroup reweighting plus a gate entropy regularizer that keeps routing from collapsing onto the sensitive attribute. Fairness improves while predictive performance stays competitive, and the resulting routing distribution doubles as an interpretable map of how subgroups are spread across experts.

DIME: Query-Efficient Framework for Membership Inference on Diffusion Models

Tue Do, Daniel Alabi Membership inference attacks on diffusion models — determining whether a specific record was in the training set — have been largely heuristic and often need many queries per record. DIME starts from an exact characterization of the optimal denoiser for a finite training set, showing that leakage is governed by the denoiser's implicit reconstruction error, which splits into a bias term measuring reconstruction accuracy and a previously unexamined local crowding term capturing the geometry of nearby training examples; both have efficient query-only estimators. Across CIFAR-10, CIFAR-100, STL10-U, CelebA, and ImageNet it beats prior attacks at equal or far lower query cost, improving true positive rate at 1% false positive rate by up to 3x, and the two-query variant can outperform existing 30-query baselines; the authors also evaluate candidate defenses.

Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling

Akifumi Wachi, Takumi Tanabe, Youhei Akimoto Inference pipelines that sample many candidate outputs, filter them with a learned safety model, and return the highest-reward survivor suffer a compounding failure the authors call safety hacking: the imperfect safety proxy admits unsafe outputs into the feasible set, and reward maximization then preferentially selects them. Finite-N bounds for constrained Best-of-N sampling show the outcome is governed by the upper reward tails of safe versus unsafe feasible outputs, and if unsafe-but-feasible outputs have the heavier tail, safety hacking becomes asymptotically certain as N grows even when average proxy errors and false-positive mass are arbitrarily small. Policies kept within a bounded chi-squared divergence of the proxy-feasible reference distribution admit an N-independent bound, instantiated as constrained pessimistic sampling, but this coverage control only limits amplification and cannot clean a contaminated feasible set; toy and language-model experiments illustrate both effects.

Concepts for Securing Agentic AI Coding and the Terok Environment

Ji\v{r}\'i Vysko\v{c}il, Franz P\"oschel, Andreas Kn\"upfer An assessment of the IT security risks introduced by agentic AI coding assistants argues that the agentic mode both adds severe new risks and sharply worsens existing ones relative to earlier non-agentic AI-assisted coding. The contribution is threefold: a risk assessment, a mitigation concept intended to preserve the productivity benefits rather than lock the tooling down, and an overview of Terok, an implementation of that concept. The stated goal is to let teams evaluate agentic coding aggressively without absorbing its security exposure, with the authors framing the work as a step forward rather than a settled answer in a fast-moving field.

What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces

Zhiwei Gao, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki Activation steering results are usually measured under the same answer encoding used to build the steering direction, so a reported gain may reflect the intended semantic judgment or merely compatibility with the answer labels seen during construction. Cross-Encoding Steering Evaluation freezes the intervention and re-encodes the answers for the same held-out items, revealing that on NormBank contrastive activation addition shifts scores according to the position an answer occupied during extraction rather than its semantic label — an effect the authors call extraction-index following, which persists across different identifier vocabularies and row orders. A low-rank output-sensitive component holding 15.4% of the direction's squared norm retains 96.3% of the effect, an Inference-Time Intervention-style method shows the same bias in three models, and MNLI and Social Chemistry 101 fall on opposite sides, indicating that a steering gain under one encoding does not identify what the intervention actually controls.

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

Xunlei Chen, Qirui Ye, Yuang Li, Yi Gong, Zhaokun Wang, Wenyi Li et al. Machine unlearning for language models must remove specific knowledge without collateral damage, yet sequence- and token-level methods penalize target outputs directly and so tend to disrupt linguistic structure or wipe out benign facts. ADU reframes the problem as decoupling contextual attention pathways rather than erasing tokens: exploiting the split between local and global attention heads, it locates "preplan" positions that retrieve persistent sensitive anchors, fixes their candidate retrieval paths under the original model, and trains attention-projection adapters to suppress attention mass along those paths while leaving local attention and retain-set language modeling intact. On the TOFU and WMDP benchmarks it posts the best aggregate results among the compared baselines, including a Forget Quality of 0.93 on TOFU, while retaining 92.9% of model utility on average versus 81.9% for baselines.

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha Audits that credit multilingual models with cultural grounding may be picking up surface artifacts instead: identity labels can be inferred from explicit textual cues, and names or phrasing can leak the source language. A human-validated multi-agent audit of 89,253 outputs from 12 models in English, French, and Chinese, covering 18 occupations and three task conditions, separates three distinct questions — whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect genuine cross-cultural patterns. Stripping direct identity cues sharply cuts identity-label prediction in English and Chinese but barely moves it in French, and source-language identifiability drops substantially after translation and again after masking names, suggesting that uncontrolled multilingual audits can mistake surface cues for cultural understanding.

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng Jailbreak research has concentrated on single-turn prompt optimization, leaving open how models hold up against sustained, psychologically grounded persuasion of the kind that arises when users treat them as conversational partners. PsychJail maps established social-psychological persuasion techniques into a tactic-conditioned attack policy, factorizing each attacker turn into a Change-of-Meaning analysis, a tactic choice, and the message the victim sees, then refining the policy with trajectory-level reinforcement learning under a reward that credits early success only when every turn carries a well-formed analysis. It reaches a 87.3% average attack success rate across four aligned victim models, beating single-turn and multi-turn baselines on each, and the tactic that breaks each model yields four distinct susceptibility fingerprints — rationalist, credibility-driven, narrative-monoculture, and broadly persuadable — offered as a conjecture explaining asymmetric cross-model transfer.

ST$^2$U: Stateful Test-Time Unlearning via Restricted Knowledge Boundary Control

Xunlei Chen, Qinghui Gong, Ruini Xue, Yaodong Hu, Tian Lan, Wenhong Tian Test-time unlearning suppresses restricted knowledge by editing activations during inference rather than retraining, but existing methods apply isolated pointwise corrections and ignore that autoregressive generation keeps rebuilding hidden states from the prompt, cache, and generated prefix — so the model drifts back into restricted regions after a locally successful fix. ST²U recasts the problem as trajectory-wide boundary control: it models restricted-knowledge boundaries in low-dimensional invertible coordinates, leaves orthogonal components untouched, monitors risk as generation proceeds, and propagates correction state across tokens with contextual anchoring. Across three benchmarks and three model families it achieves the best overall balance of retention and forgetting, cutting restricted-knowledge re-entry to 13.76–19.84% versus 46.50–59.10% for test-time baselines.

Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation

Sujoy Nath, Aswini Kumar, Tanmoy Chakraborty Automated counterspeech generation has mostly focused on tone and style while treating all hate speech as one thing, even though rebutting misinformation calls for a different strategy than rebutting dehumanizing language. FIRE first classifies a hate message into one of five categories (misinformation, stereotype, conspiracy, dehumanizing, non-factual) and then maps that category to a matching counterspeech style, supported by FactualCS, a new 4,784-instance dataset annotating hate categories, reasoning traces, and evidence mappings. Measured against 28 baseline configurations and using agents under 2B parameters, it improves factual accuracy by about 12% and category-specific accuracy by about 11% while cutting toxicity roughly 11%, with human raters preferring its responses.

Credal Large Language Models for Semantic Commitment under Uncertainty

Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin Standard language models express uncertainty through one predictive distribution, which blurs the difference between not knowing something and a genuinely ambiguous question. Credal Large Language Models replace that single softmax with an ensemble of LoRA adapters inducing a credal set whose lower and upper probabilities expose how much plausible distributions disagree, yielding two scores: Credal Token Commitment, computed from lower-bound support, credal width, and intersection entropy without extra generation, and Semantic Commitment Consistency, which extends the idea to sampled completions. Tested on Gemma-2-9B, Llama-3.1-8B, and Qwen2.5-7B across OpenBookQA, CoQA, TriviaQA, and ARC-Challenge, the method leads on question-answering accuracy at competitive calibration error, and the generation-free token score comes within 1.5 percentage points of the best hallucination-detection AUROC in most settings. Selective prediction at 80% coverage reaches 99.0% accuracy on OpenBookQA.

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet Alignment training pushes language models toward both helpfulness and harmlessness, and this work probes what happens when those goals collide by presenting the same unethical scenarios in three framings: objective classification questions, subjective first-person statements, and direct requests for help. Models behave worst on the request-for-assistance framing, and Layer-wise Relevance Propagation traces the gap to an attribution bias in which benign task-framing tokens like "Can you help me" receive more weight than the cue-tokens that actually signal harm, such as "without getting caught". Two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue-tokens produced safer responses, supporting the claim that under-attribution to cue-tokens drives harmful compliance.

Adversarial Entropy Inflation Against Gumbel-Based Inference Verification

Nikita Kezins cross-listed Gumbel-based inference verification constrains covert exfiltration of language model weights by forgiving only token choices plausibly attributable to honest GPU nondeterminism, and was reported to slow a steganographic adversary by more than 200×. That figure presumes a passive attacker; because the verifier's admissible-token-set size scales with the model's own output entropy, an adversary who controls the prompt distribution can widen the channel with inputs that break grammatical and sub-word structure. Across six instruction-tuned models from 1B to 32B parameters and three seeds, character- and script-level disruption roughly doubles the bits leaked per token and cuts the slowdown to 60×–118×, indicating that jitter-forgiveness thresholds need dynamic calibration against local token entropy rather than benign traffic.

STONIC: A Layered Measurement Contract for LLM Value Profiling

Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev, Danil Sazanakov, Mikhail Solovev et al. Studies of language model "values" typically pool questionnaire ratings, forced choices, and values inferred from free text into a single profile, assuming all three measure one stable preference. STONIC tests that assumption across 5,144 situations from four question banks and 35 fixed model configurations, comparing isolated ratings, counterbalanced conflict choices, spontaneous answers, and later choices between a model's own prior answer and authored alternatives. Only 10 of 17 configurations with usable behavioral data preserved the endorsement-choice relation, while every eligible configuration preferred its own earlier answer, with a median effect of 0.790, and option position shifted choice rates everywhere. Profile shape transferred best from ratings to conflict choices and degraded for spontaneous text, so behavior is reproducible but there is no single scorer-independent value identity across interfaces.

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

Hanling Tian, Gengyu Zhang, Zeyang Sha, Jingying Wang, Yuhang Liu, Zhehao Huang et al. cross-listed Persistent memory is becoming standard in deployed language model agents, raising the question of what attack surface it opens. InjecMEM needs only a single ordinary interaction — no read or write access to the memory store — to plant a record that steers the agent's later answers on a chosen topic toward an attacker-specified output. The injected record pairs a retriever-agnostic anchor packed with high-recall topical cues, so retrieval reliably surfaces it, with a short adversarial command optimized by gradient-based coordinate search across synthetic prompt templates and insertion positions to survive uncertain context fusion and long prompts. Across multiple memory systems and backbone models the attack achieves reliable topic-conditioned retrieval and targeted generation, persists under memory drift, and leaves non-target queries unaffected.

On the Threat Model of Weird Generalization and Emergent Misalignment

Miriam Wanner, Mark Dredze, William Walden Narrow fine-tuning on small domain-specific datasets can shift a model's behavior far outside that domain, an effect termed weird generalization (WG) and closely related to emergent misalignment, but the data properties that trigger it were unclear. Experiments across three open-weight models and four datasets systematically varied size, composition, language, presentation style, and how novel the content was relative to pretraining knowledge. Composition and language mattered far more than dataset size, familiar pretraining-adjacent data produced stronger effects than novel data, and measured WG shifted noticeably depending on which evaluation questions were used — leading the authors to frame it as an adversarial threat requiring deliberate data engineering rather than a hazard of routine fine-tuning.

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

Yipeng Zhao, Qishun Yang, Shenzhe Zhu, Shu Yang, Di Wang Fine-tuning on entirely benign reasoning data — math, code, chain-of-thought problem solving — can degrade a model's safety behavior, an effect called Reasoning-Induced Misalignment that prior work attributed to neuron entanglement without explaining its geometry or offering a training-time remedy. Activation-space analysis extracts separate reasoning and safety directions and shows they are coupled: fine-tuning that improves reasoning displaces safety representations, and prompts with larger displacement suffer larger safety loss, with centered kernel alignment distances and probes localizing the responsible layers. The resulting Safety-Direction Penalty (SDP) penalizes movement along the learned safety direction during fine-tuning and, on Qwen2.5-3B and 7B, restores safety without costing reasoning benchmark performance.
3 more specialized papers

Theory 51

Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?

John C. Howell When a task is defined by a handful of examples, a model can be specialized four ways — zero-shot prompting, in-context attention, test-time gradient adaptation, or having a hypernetwork emit specialist weights — and the last option's viable operating range is poorly charted. An identical four-way comparison was run across six tasks covering regression, generation, language modeling, reinforcement learning, and clinical and genomic classification, holding the specialist architecture, context, and training budget fixed where possible. Weight emission wins mainly on cost at equal quality: it matches TabPFN on clinical few-shot classification while producing a reusable specialist rather than re-attending the support set for every query, and beats MAML by two to three orders of magnitude on few-shot sinusoid regression at zero test-time gradient steps (still roughly 30x after equalizing budgets). On high-dimensional sequence modeling, however, a one-pass adapter recovers only 11–14% of the in-context learning gain, and a LoRA rank sweep shows the shortfall is a genuine capacity limit that plateaus around 21%.

Complexity Induction: Compositional Generalization via Structured Label Distortion

Aleksandr Abramov cross-listed Shows that deliberately distorting training labels in a structured way — termed complexity induction — can produce compositional generalization in an ordinary convolutional neural network with no architectural changes. Using synthetic images of colored shapes labeled as flat strings like "red-circle" with certain color-shape pairs held out entirely, two distortions derived from Jaccard string similarity are applied: soft mixed labels encoding inter-class overlap, and an expanded dataset of false samples with structurally motivated wrong labels. Both let the classifier predict unseen combinations, operating at different levels — mixed labels exploit the existing embedding structure while expanded training improves the embedding's factorization — and a control with random false labels confirms the effect depends on the structure of the distortion rather than noise.

Retrieval Needs Multivectors: An Exponential Separation

Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu, Ankit Garg, Kirankumar Shiragur cross-listed Dense retrieval typically compresses a document into one embedding vector, and recent benchmarks such as LIMIT suggest this format has hard expressive limits that multi-vector schemes escape. Building on Jayaram's analysis, the authors construct an explicit family of queries, documents, and relevance matrices for which any single-vector embedding that ranks every relevant document above every irrelevant one needs exponential dimensionality, while polynomial-size multi-vector embeddings suffice — an exponential separation for ranking rather than for approximating numerical scores. They turn the construction into ANDOR, a retrieval benchmark on which state-of-the-art single-vector models score poorly zero-shot and barely improve with fine-tuning, whereas multi-vector models both lead and gain substantially from fine-tuning.

Sorting from Counterexamples

Noga Alon, Shay Moran, Shlomo Moran A learner tries to identify an unknown linear order on n items by proposing complete orderings and receiving either confirmation or a counterexample — a pair of items in the wrong order — where up to k counterexamples may be lies and k is not known in advance. The optimal query complexity is pinned down to Θ(n log n + nk), so the noiseless case costs the same as classical sorting while each untruthful counterexample adds cost of order n; the upper bound rests on a geometric representation of permutations plus Grünbaum's theorem, and the lower bound combines sorting arguments with a Condorcet-type construction. For rankings generated by projecting d-dimensional points onto an unknown direction, they give an O(d² log n + dk) upper bound against an Ω(d log n + dk) lower bound, leaving a factor-of-d gap in the noiseless term.

Variational Structure at the Edge of Stability

Eric Regis Discrete-time optimizers running at the edge of stability settle into near-two-periodic oscillations that look like the behaviour of conservative systems and symplectic integrators, but the link to discrete mechanics has not been made precise. Building on Litman's "edge coupling" — a functional on consecutive gradient descent iterates whose critical points encode fixed points and two-point orbits — this work extends the construction to heavy-ball and Nesterov momentum, showing that critical points again characterize fixed points and two-point orbits with the Hessian determining their stability. The edge coupling is then identified with the symmetric Verlet action, making the correspondence between edge-of-stability training and discrete mechanics explicit.

Posterior Information Dynamics of Diffusion Models for Linear Inverse Problems

Xiangming Meng Diffusion models are routinely used as priors for linear inverse problems, but endpoint image quality says nothing about when measurement information actually enters the reverse denoising trajectory or how it is spread across signal directions. The analysis centers on the smoothed likelihood force — the gap between exact posterior and prior scores at each noise level — whose expected squared norm quantifies both relative-entropy dissipation and reverse-path entropy growth, and whose measurement-average yields an information–minimum mean-square error (I-MMSE) identity linking information gain to denoising-error reduction. Solvable models predict quadratic decay of the force energy at high noise, entropy reduction from log n to a conditional value, and assimilation that depends on operator–prior alignment even for identical singular values; a frozen FFHQ model shows masks with matching spectra producing different null-space trajectories.

Width-Independent Compressibility of Deep Neural Networks

Hong-Yi Wang, Mingze Wang, Liu Ziyin Trained networks compress dramatically without losing accuracy, a well-known effect that has lacked theoretical grounding. A uniform compressibility theorem is proven for deep multilayer perceptrons with analytic activations: for any fixed wide teacher network there exists a same-depth narrow network approximating the same function, and the achievable compressed width does not depend on the original width at all, scaling instead as O((log(1/ε))^d) in the error budget ε and effective input dimension d. The construction rests on a derivative-matching technique that exploits low-dimensional input structure plus a layer-wise reweighting preserving the input-output map.

More Experts, Worse Dynamics: Inverse Scaling and Spectral Bias in Mixture-of-Experts State-Space Models

Chandresh Pandey Mixture-of-Experts (MoE) designs are usually justified by the idea that mixing simple local operators helps a model track heterogeneous or regime-switching time series, an argument recently carried over to spectral state-space models. A controlled synthetic study tests this on next-step prediction over three stitched regimes — chaotic Mackey-Glass dynamics, stable oscillation, and noise-dominated autoregression — with ablations covering capacity scaling, oracle routing, frozen experts, and output-level MoE baselines. Operator-level mixtures never beat a single-expert baseline, and adding experts makes results worse rather than better, with routing either collapsing or failing to specialize even under perfect regime supervision. Phase-space analysis further shows that lower mean squared error on chaotic trajectories can come from temporal smoothing that destroys the attractor geometry, arguing for geometry-aware evaluation.

Blockwise Stabilized Adaptive Cubic Regularization with Subsolvers via Recurrence

Rodion Podorozhny Cubic-regularized Newton methods have an optimal O(ε^-3/2) global rate and escape saddle points automatically, but solving their subproblem by full eigendecomposition caps the model size they can handle. The proposed optimizer partitions parameters by tensor, fits an independent cubic model with its own adaptive constant per block, and accepts or rejects each block step against a monotone guard on the full loss; small blocks get exact cubic steps from explicitly formed Hessians while arbitrarily large tensors use a matrix-free Chebyshev-bounded Krylov solver built by the Lanczos process, with the cubic shift bounding polynomial degree and preserving the convergence rate under inexact solves. On a 199k-parameter FINER implicit neural representation run to convergence, ARC-φ1 reaches 133.5 dB peak signal-to-noise ratio while tuned Adam plateaus at 78.2 dB on the same extended budget, and the block steps stay exact in the cubic-model sense on every block of a 91.4M-parameter ViSIR model, including its 88.5M-parameter decoder tensor.

Loss Landscape Features That Make Adam Stall: Definitions, Estimators, and the Preconditioned Hessian View

Rodion Podorozhny Adam sometimes converges to a low loss on badly conditioned objectives and sometimes stalls on a plateau far above what second-order methods reach, and this report supplies measurable diagnostics for telling the two cases apart. The proposed indicators are the condition number of both the Hessian and the Adam-preconditioned Hessian D^{-1/2}HD^{-1/2}, a diagonal mass statistic that separates axis-aligned from cross-coupled ill-conditioning, negative spectral mass from stochastic Lanczos quadrature, and gradient energy split across curvature bands. A 2x2 worked example shows diagonal preconditioning can rescale away axis-aligned ill-conditioning but not cross-coupled ill-conditioning, and a case study on the FINER implicit-neural-representation architecture traces its Adam stall to saddles while blockwise second-order methods reach 120–134 dB PSNR.

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

Yang Yu World-action models predict a future outcome and then infer the action associated with it, and the question addressed here is whether that factorization buys any control capability over plain behavior cloning when both learn from the same observational demonstrations. The analysis shows every world-action policy flattens into a direct stochastic policy with an identical closed-loop trajectory distribution, and that under realizability, exact optimization, and distribution-preserving deployment both approaches recover the same observational behavior policy — so predicting the future changes the learning factorization but not the achievable policy class or the imitation target. Action-conditioned world models are genuinely different because they predict consequences of specified actions; the authors bound the irreducible prediction error of models that ignore the candidate action, show observational data does not identify action effects in general, and construct environments where every observational learner suffers positive worst-case regret while a single informative intervention permits zero regret.

Dataset Complexity Shapes Finite-Distance Loss Geometry in Neural Networks

Jaeyong Bae, Hawoong Jeong cross-listed Datasets of identical size and low-order statistics can differ sharply in structural complexity, and this work links that complexity to the geometry of the loss landscape around trained solutions. Local label mixing across neighborhood scales is paired with local entropy — a Franz–Parisi construction from spin-glass theory measuring the effective volume of low-loss parameter configurations at a given distance from a reference — estimated in finite networks with adaptive sequential Monte Carlo. In a controlled synthetic sweep, greater dataset complexity produces a larger drop in local entropy near the reference solution while the far-field radial derivative is weak and nearly identical across conditions, meaning complexity changes where the solution volume contracts rather than making it contract uniformly faster; real image data reproduce the trend, amplified by label randomization.

Functional compatibility as a determinant of persistent neural learning

Hossein Javidnia Networks that keep learning tend to damage what they already do, and the usual responses — replay, regularization, constrained updates — leave open whether some property of the incoming task itself decides what can be retained. The claim here is that functional compatibility, meaning the extent to which new learning can coexist with behavior that must be preserved, is a causal determinant of what learning persists, demonstrated by deliberately varying compatibility from matched neural states under a common retention requirement. The effect holds across independent learning directions, convolutional and transformer architectures, vision and text, and multiple seeds; learning rules differ in how efficiently they exploit available compatibility, and at larger finite updates nonlinear geometry breaks the matched compatibility continuum. The framing shifts the problem from preventing forgetting to identifying which parts of new learning can safely be made permanent.

Q-Learning with Stable Infinite-Dimensional Linear Function Approximation

Shengbo Wang Q-learning with linear function approximation can diverge because an arbitrary feature architecture need not preserve the Bellman contraction. The proposed framework learns an infinite-dimensional coefficient field over a compact latent metric space, paired with a reconstruction operator that maps it to a continuous Q-function and a compression operator that maps Bellman updates back to latent coordinates; requiring both to be nonexpansive makes the latent Bellman map contractive with a unique fixed point approximating the optimal Q-function up to representation error. Two stochastic approximation algorithms learn from a single Markovian behavior-policy trajectory with sup-norm convergence whose leading term is of order Õ(n^{-1/2}), and estimation error is controlled by covering numbers of the latent space rather than its dimension. The algorithms need no knowledge of the latent metric, so they adapt automatically to smoothness and geometry; worked cases include density-based Q-measure learning and training only the output layer of a frozen pretrained network.

Change Detection in Probability Flow ODE: Online Testing in Diffusion Latent Spaces

Artem Kraevskiy, Artem Prokhorov Detecting when a sequential data stream changes distribution — a market trend reversal, a scene cut, a shift in motion direction — is hard when neither the pre-change nor post-change conditional density has a closed form, which rules out classical likelihood-ratio statistics. The approach trains a conditional diffusion model on pre-change data with a frozen context encoder, using the probability flow ordinary differential equation as a deterministic bijection that maps pre-change observations to standard Gaussian latents; post-change data pushed through the same frozen map departs from that Gaussian reference. Maximum Mean Discrepancy serves as the test statistic, with closed-form components under the Gaussian null and an asymptotic distribution derived as a degenerate U-statistic, feeding a Shiryaev-Roberts online procedure with exact threshold calibration that detects arbitrary shifts including covariance rotations and higher-order structural breaks without parametric assumptions on either regime.

Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models

Yo Ehara Flesch Reading Ease and Flesch-Kincaid Grade Level are computed from the same two document statistics, and this work asks whether their stability on long texts reflects genuine readability signal or just lexical composition. Under a topic model that includes an explicit sentence-boundary token, both scores are shown to converge almost surely, in the long-text limit, to deterministic functions of the document's topic distribution through only two scalar rates — meaning all score variation is mediated by topical composition rather than any residual readability signal. Empirically, a topic vector inferred from one half of a document's content words predicts the other half's grade level at r = 0.779 on Brown and 0.884 on the written BNC, though adding topic to genre and mean syllable count yields negligible additional explained variance, and the authors caution that inferred topics may absorb genre, register, and style rather than readability itself.

Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding

Seokjin Hwang (Ray), Yuting (Ray), Li, Kiwan Maeng Instance encoding transforms sensitive data before handing it to an untrusted server, but rarely carries any proof that the transformation resists inversion; a recent mean-squared-error (MSE) bound on adversarial reconstruction was among the first such guarantees, and it is loose, restricted to randomized encoders, and limited to MSE. A new family of bounds accounts for the encoder's spectral structure to deliver tighter guarantees that also apply to fully deterministic encoders and extend to norm-based similarity metrics beyond MSE. Evaluation across a range of encoders, datasets and reconstruction attacks shows the bounds hold consistently and improve on the prior result.

Interpretable AI with Local Distillation

Erin Craig, Yiling Huang, Snigdha Panigrahi cross-listed Tabular foundation models and gradient-boosted ensembles predict well but offer little to reason about, which is a problem when the prediction drives a high-stakes decision. Local distillation fits a regularized linear "student" at each query point under guidance from a black-box "teacher" that both defines locality — by upweighting training points with similar predicted outcomes — and anchors the fit as a data-weighted pseudo-observation; adding small Gaussian randomization to the local objective yields refit-based selection frequencies and stable subgroups, with a proof that lasso selection probabilities are stable under small perturbations of training responses. Across 17 benchmark datasets the method nearly matches its teacher's accuracy while emitting a sparse linear model per test point, and on cancer gene expression data it surfaces patient subgroups whose local models rely on different genes.

Provably adaptive sampling with uniform and remasking discrete diffusion models

Daniil Dmitriev, Zhihan Huang, Yuting Wei Discrete diffusion models generate sequences by updating many coordinates in parallel, but existing analyses of the standard $\tau$-leaping sampler under a uniform forward process require a number of steps that grows linearly with the sequence dimension $d$. A first-order sampler built on the leave-one-out denoiser, applied to both uniform and remasking forward processes, can correct denoising mistakes mid-sampling, which matters once many coordinates change at once. The main guarantee is that the number of discretization steps scales with the target distribution's dual total correlation rather than with the ambient dimension, meaning cost tracks how much the coordinates actually depend on each other. The analysis routes through a Bayes-optimal auxiliary sampler that separates discretization error from score-estimation error, and yields an exact information-theoretic expression for discretization error in terms of mutual information between coordinates at different forward times.
32 more specialized papers

Multimodal 44

Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

Yisong Xiao, Aishan Liu, Yongxin Huang, Zonghao Ying, Shiji Zhao, Tianlin Li et al. Large vision-language models absorb social bias from training data and behave differently on portraits of different social groups; current debiasing methods compare token probabilities between an original and a deliberately biased generation, which anchors the correction to one stereotyped viewpoint. Counterfactual Ensemble Decoding (CED) instead identifies semantic directions in the visual representation space associated with each social group, generates counterfactual representations along them to supply multiple perspectives, then locates the decoder layer where those perspectives diverge most and ensembles their token distributions with uncertainty-aware weights. Across three social bias benchmarks covering occupations, descriptors, and persona traits, the method reduces bias by up to 47.97% over leading baselines while leaving the model's general capabilities largely intact.

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

David Bamman, Kent K. Chang, Allison Cooper, Juishan Hsu, Reina Kushihashi, Madison Mar et al. Addresses the copyright obstacle to building film-understanding benchmarks by assembling a collection of Hollywood films selected for box office popularity and likely public domain status, which required researching US Catalog of Copyright Entries registrations and renewals and publishing the first large-scale open dataset of weekly box office earnings from Variety magazine covering 1922 to 1979. On top of this collection the authors build a multiple-choice benchmark probing narrative elements relevant to film scholarship. Many vision-language models score near chance, while audio-visual models — including those that use audio when captioning scenes — top out at 61.1% accuracy, far below human performance.

presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

Weixuan Ding, Shang Liu, Hanyu Pei, Zeyan Liu cross-listed Placing an object realistically into a scene for image composition normally requires hand-written rules or supervised models trained on narrow datasets, neither of which generalizes to unfamiliar objects and scenes. presto reframes placement as heuristic search in an imaginary action space, where a multimodal large language model reasons over candidate positions and scales and a coarse-to-fine strategy converges quickly, with two decision variants compared: metric-guided selection and using the model itself as judge. The zero-shot, training-free method reaches state-of-the-art results across benchmarks, especially in open-world settings, and human studies find the model-as-judge variant produces more perceptually coherent placements than metric-driven selection, exposing a gap between standard metrics and human judgment.

The Plan, Not the Decoder: Diagnosing and Repairing Compositional Failure in Reasoning-Augmented Text-to-Image Generation

Ashritha Gonuguntla cross-listed Reasoning-augmented text-to-image models like GoT-R1 write an explicit textual plan with object names, attributes, and bounding boxes before decoding, which makes it possible to ask whether compositional failures come from a bad plan or an unfaithful decoder. Swapping boxes inside the model's own chain flips the generated layout under detector-based scoring (0.75 to 0.48) while a common visual question answering (VQA) spatial metric moves the wrong way, and human raters side with the detector on 81% of items — so all spatial claims use geometric scoring. Under that measurement the decoder faithfully realizes 94% of planned relations, and the planner is the bottleneck, showing a raster-order bias (98% accuracy on "left" versus 54% on "right" for identical layouts); editing plans without retraining yields gains from +5.0 to +13.3 points, with box geometry mattering far more than prose style.

Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models

Jinchang Zhu, Rong Fu, Yi Ding, Chenghao Wu, Ying Liu, Menglin Yang cross-listed Vision-language models (VLMs) miss detail-centric questions because the answer, though present in the image, disappears when the image is compressed into a low-resolution global view, and simply spending more visual tokens everywhere wastes computation and can hurt tasks needing global context. GapSight teaches a model to re-read: after a global glance it decides whether to return to a free-form crop, with supervision mined offline from the target model's own failure signal — crops that lower answer loss or widen the multiple-choice margin become model-specific review labels, distilled into a one-shot router predicting whether to review, expected utility, and a continuous crop box. Across LLaVA-1.5-7B, InternVL2.5-8B, and Qwen2-VL-2B-Instruct it beats the no-zoom baseline on six OCR, document, chart, and real-world benchmarks, lifting the InternVL2.5-8B six-benchmark average from 52.25 to 64.29, ahead of CropVLM, ViCrop, and ZoomRefine.

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou et al. cross-listed Knowledge-based Visual Question Answering (KB-VQA) requires reasoning over external sources, and existing retrieval pipelines struggle to filter noisy context or to keep the generated answer tied to the evidence actually retrieved. SAFE-G first recalls candidate documents with a coarse hybrid visual-textual search, then applies structure-aware graph retrieval that models dependencies among evidence pieces to pinpoint the relevant passages, and trains the generator with a reinforcement learning reward that credits a correct answer only when the selected evidence is also correct. On Encyclopedic-VQA and InfoSeek the framework reports gains of 8.9% and 3.5% over prior methods.

MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning

Suifeng Zhao, Zida Liu, Xinyu Lei, Lei Sun, Jun Gao, Sujian Li Multimodal retrieval-augmented generation needs visual citations so answers can be traced back to their source, but retrieval pipelines and supervised fine-tuning tend to produce imprecise citations or citations decoupled from the text they supposedly support. MCite-RL treats citation as an agentic process: an Agentic Refinement module iterates retrieval, reasoning, and recursive cropping to progressively narrow the visual search space, and a citation-enhanced reward combines process-level and outcome-level feedback so reinforcement learning optimizes answer accuracy and source traceability together. Experiments on Wiki-VISA, FinRAGBench-V, and MMLongBench-Doc show joint improvement in citation precision and answer quality rather than a trade-off between them.

PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models

Jihyung Ko, Eunji Jung, Hyeongsub Kim, Ziseok Lee, Jae Won Cho, Sanghyun Jo et al. cross-listed Training-free methods for reducing captioning hallucinations mostly suppress unsupported object words as they are generated, which leaves visible objects the model simply omits unrecoverable. PatchGate instead extracts object evidence intrinsic to a frozen vision-language model before generation: a Visual Evidence eXtraction stage reads patch-level lexical evidence from the latter half of the language decoder layers to build an image-conditioned object set with no task prompt, and an inclusion-exclusion decoding stage recalibrates logits to promote under-verbalized supported objects and suppress over-verbalized weak ones. On AMBER this raises visible-object coverage from 49.4 to 56.0 while cutting CHAIR hallucination from 7.5 to 6.6, using one extra forward pass and no detectors or fine-tuning.

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

S{\l}awomir Dadas, Micha{\l} Pere{\l}kiewicz, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Bart{\l}omiej Jaworski, Izabela Wo\'zniakowska Multimodal evaluation is dominated by English, leaving open how well models handle images, audio, and documents grounded in other languages and cultures. PUMA (Polish Unified Multimodal Assessment) is a hand-built benchmark of 900 tasks probing Polish cultural knowledge alongside practical processing of text, images, audio, and visually rich documents. Evaluating frontier commercial systems, open-weight models, and smaller specialized ones reveals a wide spread: top commercial models score well on visual question answering while most models struggle with complex audio and document understanding. The evaluation framework is released openly.

Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems

Henry Fordjour Ansah (Louisiana State University of New Orleans), Shreya Banerjee (Louisiana State University of New Orleans), Pranish Ghimire (Louisiana State University of New Orleans) Qualitative mechanical reasoning — judging how gears mesh, what supports what, or which side is heavier — is tested on human job candidates through the Bennett Mechanical Comprehension Test, and this work applies the same format to two open multimodal models, Gemma-3 and Qwen-VL. Each test image encodes ground-truth qualitative facts such as contact points and support relations, which are used to grade not just the final answer but the elicited chain-of-thought for coherence, completeness, and logical progression. The evaluation targets whether small vision-language models can ground spatial and commonsense physical reasoning rather than pattern-match to answer options.

Query-Driven Multimodal Information Extraction from Long Documents

Yikai Gao, Ding Xia, Xi Yang Document question answering systems typically return text answers or highlight evidence regions, but do not return the specific attribute values a query asks for together with the images that illustrate them. The authors define query-driven image-text joint extraction, where a model must output requested textual attribute values plus bounding boxes for the corresponding images, and release ITJoint, a manually annotated benchmark of 2,455 pages of domain-specific documents with 316 queries and 910 answer instances organized by a query-level and instance-level taxonomy. They also build Q2IT, a three-agent pipeline that collects evidence, selects pages, and localizes target images. Standalone vision-language models perform poorly on the joint text-plus-image metric while Q2IT improves substantially, though a large gap to perfect performance remains.

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei, Shiyao Cui, Jinqiao Shi Scenes often look different from what is physically true — reflections, transparency, projections and similar effects the authors group under the term situational illusions — which undermines multimodal model reliability in deployment. A where-what-how taxonomy of these illusions grounds MSIBench, a benchmark testing discrimination, understanding, and reasoning, evaluated across 27 model configurations. Current multimodal models are highly vulnerable and exhibit six recurring failure modes spanning visual observation, grounding, and reasoning; prompting for closed-source models and supervised fine-tuning for open-source ones, both built on systematically inspecting visual evidence, raise performance by up to 20%.

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer et al. cross-listed Audio-language models are usually tested on what a recording contains, not on how it has been damaged by noise, clipping, codec artifacts, or similar corruption. MRMAD covers speech, music, and general sound, and structures evaluation as multi-turn dialogues over several audio inputs, so a model must name degradation types, rank severity between clips, and track corruption changes across turns while keeping its hypotheses consistent as evidence accumulates. Across 18 large audio-language models spanning non-reasoning, reasoning, and Omni variants, models generally recognize coarse content but fail to reliably diagnose, compare, or reason about degradations.

Training-Free VLM Personalization via Calibrated Residual Decoding

Jiaao Yu, Yujian Ma, Xianming Hu, Pengran Wang, Ang Li cross-listed Vision-language models can be personalized at inference by pasting user profiles or reference images into the prompt, but there is no guarantee the model actually uses that evidence rather than falling back on generic priors — and a confident answer gives no signal about which one it came from. This training-free decoding method runs the same image and question under three conditions, a positive profile, a counterfactual profile, and an empty profile, anchors the prediction on the empty-profile output, and estimates personalization's marginal contribution from the score differences, with normalized-entropy calibration scaling the adjustment to how reliable the residual signal looks. Consistent gains appear on MMPB, YoLLaVA, and MyVLM for identity-sensitive visual personalization without any fine-tuning, and entropy calibration is shown to stabilize the method when the contrastive signal is weak.

OVIBench: Benchmarking Online Video Question Answering under Interruption

Naiming Liu, Zhiheng Wu, Shuning Wang, Tie Zhang, Bowen Liu, Tong Wang cross-listed Video question answering is almost always benchmarked offline and in a single round, ignoring that real users interrupt an assistant while it is answering. OVIBench defines online video QA under interruption with three interruption types — cancellation, false trigger, and correction — and adds an offline simulation protocol that reproduces interruptions at a unified point during generation plus metrics for both interruption understanding and response generation. The benchmark separates vision language models most sharply on following correction requests, and fine-tuning on the companion OVI-Train set produces large gains, which the authors take as validation of the data design.

Semantics or Structure? Auditing Text Sensitivity in Multimodal Time-Series Forecasting

Karthik Sridhar, Atharva Gupta, Nishant Pradhan, Murari Mandal, Dhruv Kumar, Saurabh Deshpande Multimodal time-series forecasters report large gains from adding natural-language context, but nobody had verified that the models are actually responding to what the text says. Controlled perturbations on the Time-MMD benchmark replace each row's text with empty, constant, within-domain shuffled, or cross-domain text across Aurora, MM-TSFlib, and TaTS, backed by attribution analysis and probes of Aurora's text pathway. Every substitution moves mean squared error by less than 0.5%, while removing a co-shipped numeric column — leaving the text untouched — reproduces the reported improvement, indicating that in this family of frozen-encoder architectures the text content is not the operative signal; the perturbation protocol and evaluation harness are released as a diagnostic toolkit.

Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video

Masoud Jalayer, Changyi Li, Yu Xiao cross-listed Captioning hours-long egocentric video with a vision-language model (VLM) is dominated by the cost of model calls, and existing triage policies that rank windows by visual features ironically require the very video decoding the budget is meant to avoid. The proposed audio-first triage scores windows using only audio before any frame is decoded, and crucially reframes the objective — the selector is trained to fire once per action rather than as a per-frame sound-event detector. Using frozen AudioSet-pretrained features with no domain-specific sound labels, the objective shift improves action coverage by 4.0–10.8 percentage points across all evaluated call rates, cutting 9–20% of VLM calls at matched coverage on EPIC-KITCHENS-100, beating uniform sampling through the mid-range on Ego4D across 247 clips, and outperforming two recent visual keyframe selectors.

Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA

Jingbo Wang, Sendong Zhao, Haochun Wang, Bing Qin, Ting Liu Medical visual question answering (VQA) is evaluated almost entirely in English, and while multilingual benchmarks show large vision-language models (LVLMs) degrade elsewhere, they do not say which underlying capabilities break. The authors build an eight-language medical VQA benchmark organized into four scenarios that isolate distinct required capabilities, and find across five open and closed LVLMs that cross-lingual degradation is highly scenario-dependent rather than uniform. MedVL-XLRepE, a training-free scenario-aware representation engineering method, steers non-English representations toward their English counterparts at inference time, consistently reducing the gap across three LVLMs and eight languages with gains of up to 6.33%.

Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs

Yikai Zhao, Qiyan Zhao, Jiaquan Zhang, Xiaofeng Zhang, Xiaosong Yuan, Pengzhou Cheng Diffusion multimodal large language models (dMLLMs) drift semantically and repeat themselves in long outputs, which the authors trace to two decoding deficiencies: confidence scoring that ignores whether a position's neighbors are already decoded, and block partitioning that hides high-readiness semantic anchors. CACD (Context-Aware Cluster Decoding) is a training-free method scoring each masked position by a multiplicative combination of softmax confidence and neighbor proximity, operating block-free so anchors stay globally accessible, with architecture-aware calibration for the confidence heterogeneity that different visual integration strategies induce. Across three dMLLMs and four benchmarks it consistently improves quality and reduces hallucination over the original decoders, with larger gains in longer-generation settings.

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie, Preslav Nakov, Zhuohan Xie Multimodal models that reason over charts and visual tables often call external cropping tools for fine-grained perception, which adds inference latency and still leaves the spatial-structural gap where text elements are bound by strict relative layout. TwSG internalizes that tool use: a multimodal model identifies key regions guided by ground-truth answers, a teacher generates region-level visual question-answering data, and those signals are distilled back into the full-image representation. Training runs a cold-start supervised fine-tuning stage on multi-turn data with focused area descriptions, then a reinforcement fine-tuning stage using a process reward mechanism called TL-GRPO. Across several model architectures the result cuts inference latency while improving accuracy and robustness, giving models native region description in a single forward pass.

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang et al. cross-listed A long-video question-answering system can produce a correct answer while having decoded frames that miss the events the answer depends on, and existing evaluations never audit which frames were actually looked at. VES-Bench supplies 600 temporal-ordering and event-counting questions over 348 long videos, each annotated with a jointly necessary set of evidence intervals so frame coverage can be checked at three strictness levels. The accompanying training-free agent TRACE grows an evidence bundle round by round from raw clips and stops only once the answer stabilizes and a repeat pass agrees; it answers 50.7% of questions correctly with at least two decoded frames inside every evidence interval using 98.7 frames per question, more than 10 points above uniform decoding at 128 frames, while scoring 86.1 on Video-MME, 75.6 on LVBench, and 75.1 on LongVideoBench.

A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports

Yufan Wang, Rui Yang, Yi Liu, Yi Lin, Yifan Peng Most multimodal medical benchmarks test fixed inputs or endpoint answers, while fully interactive diagnostic agents blur the line between choosing evidence and interpreting it. The proposed framework converts published clinical case reports into progressive multimodal diagnostic dialogues, with an evaluation strategy that scores final diagnosis, diagnostic reasoning, and interpretation of image findings separately. On 24 internal medicine case reports the conversion pipeline reproduced reference dialogues with a diagnosis F1 of 0.99 and a reasoning-quality score of 4.79 out of 5, while o4-mini and Claude Haiku 4.5 scored only 2.75 and 2.50 on reasoning quality with much lower F1 across all three axes, indicating fluent output does not imply evidence-grounded reasoning.

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen et al. cross-listed Unified understanding-and-generation models can describe objects in words but cannot represent continuous object poses or render geometrically consistent images from a target viewpoint. Object-Uni treats object pose as an explicit geometric variable shared by both sides, connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis in one formulation, and introduces a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions a multimodal LLM can consume while retaining continuous geometric supervision. Trained with an object-token-grounded pose anchor on the new UniSpatial-80K benchmark, the model improves both object-level pose understanding and pose-controllable generation, shifting unified models from describing objects toward manipulating their spatial states.

Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?

Zeyu Wang, Xinming Xu Vision-language models store spatial information in their hidden states yet frequently fail to use it when answering, and this work traces when and where that information actually reaches the answer using direction patching, a class-conditioned causal intervention swept across layers, token positions, and prompt formats. Across ten models, causal influence on answer logits appears only at mid-to-deep layers, text chain-of-thought suppresses immediate object-word transport at the argmax level in most models while visually grounded prompts keep it open, and positive gains on the target logit can stay below the threshold needed to change the top prediction. Transport can also re-emerge at the final prefix token or at the answer step in deeper layers, recasting the encoding-grounding gap as a problem of conditional transport rather than missing representation.

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (State Key Laboratory of Multimedia Information Processing et al. Printed-document parsing is close to solved on OmniDocBench at 96.34% overall, but handwritten documents remain largely unmeasured, and existing benchmarks skip handwritten tables and real-world degradation while reporting aggregate accuracy that explains nothing about failures. WildHandBench supplies 500 handwritten documents spanning free text, tables, and formulas across four languages and nine real-world scenarios, plus a Prior-Driven Error metric that separates mistakes caused by language priors from those grounded in visual evidence. Across 18 state-of-the-art models with calibrated human baselines, the best model reaches only 71.85% against 77.09% for humans, and 63 to 91 percent of model errors are prior-driven versus 49 percent for humans, a systematic reliance on language priors that plain accuracy cannot expose.

Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B

Defu Lin, Wenhui Chen, Ziyao Lin, Jianlin Chen, Peiji Long, Chi Man Vong Embodied multimodal agents accumulating long observation streams must answer questions under a fixed per-decision token budget, a constraint formalized here as four "resource walls" covering bounded state, query-independent frame selection, non-adaptive retrieval, and fixed-depth inference. ASP is a training-free wrapper around frozen multimodal models that combines a capped structured state, a verbatim episodic index, and query-conditioned budget allocation with iterative access, evaluated under a pre-registered protocol on seven open-weight models from 3B to 31B using SEW-Bench, a synthetic long-horizon walkthrough benchmark. Within a 4,096-token budget, ASP reaches 75 to 94% episodic retrieval accuracy versus 3 to 19% for equal-budget query-independent sampling, and reallocating the budget beats quadrupling it on every backbone. The authors also report a negative result: removing the compressive state raises the flagship model's mean score from 35.4 to 58.0, two of four pre-registered falsification criteria fire, and the registered natural-video benchmarks were never run because of dataset licensing.

Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

Zhe Jin, Zhimin Lin, Bin Zheng, Junhua Fang, Huihua Yang cross-listed Graph-based retrieval-augmented generation for long videos typically builds its retrieval index at the same temporal granularity as the underlying segmentation, which is finer than locating a relevant region actually requires. Density-Aware Graph Construction is a training-free method that separates the two: it merges visually redundant neighboring chunks into a compact density-adaptive index for retrieval, keeps mappings back to the original units, and expands retrieved coarse regions to fine granularity for evidence refinement and answer generation. On MLVU, VideoMME, and LongVideoBench it keeps only 40 to 50% of the original graph nodes and delivers 1.3 to 1.7 times end-to-end wall-clock speedup while retaining roughly 99% of question-answering accuracy, with gains carrying across different vision-language backbones and video RAG pipelines.

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo cross-listed Understanding culture in video requires knowing what a concept symbolizes, recognizing it visually, and locating its sub-events in time, but existing benchmarks collapse these into one score that hides where models actually fail. The Cultural Moment Benchmark (CMB) covers 306 expert-curated concepts from seven Southeast Asian countries across five categories, testing each ability in a separate stage using semantic-similarity distractors, unlabeled video moments, and free-form temporal localization on a different example video. Across six vision-language models, even the strongest closed-source systems score below 30% when all three stages must be correct, and the abilities do not cascade cleanly: naming a concept helps half the models recognize it on screen, but recognizing it barely helps them localize it in time. Audio helps, duplicates, or distracts depending on the concept — more often distracting for non-Latin-script countries — and a 14-rater human study found that even expert raters score below chance on a neighboring country's concepts.

E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

Taoyu Qian, Qi Wang, Daqian Shi, Yuanhao Jiang, Shang Gao, Hualong Yu cross-listed Encoding an image as hundreds of visual tokens dominates inference latency and GPU memory in vision-language models, and existing pruning heuristics simply average attention scores across heads and layers, which hides disagreement and uncertainty between them. E2S-Pruner treats each attention head as an independent evidence source, weights it by evidence clarity and inter-head consistency, labels each visual token important, unimportant, or uncertain, then uses Dempster-Shafer evidence theory to measure and resolve conflict between layers, with a spatial novelty constraint keeping retained tokens from clustering in a few salient regions. No auxiliary model, extra parameters, or fine-tuning is required. On LLaVA-1.5-7B it keeps 98.0%, 96.8%, and 90.6% of aggregate performance at 192, 128, and 64 retained tokens, with throughput gains of 1.96x and 2.09x at the two smaller budgets, and it transfers to Qwen2-VL-7B.

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

Xuetong Li, Gaofeng Liu Safety benchmarks for vision-language models score only the final response — refuse, warn, or comply — which cannot distinguish genuine multimodal risk understanding from keyword-triggered refusal, missed visual hazards, or over-refusal of benign content. EviSafe evaluates natural user-facing behavior alongside explicit grounding in textual and visual evidence and sensitivity to counterfactual edits of safety-critical evidence, instantiated as EviSafeBench with 1,181 gold image-text scenarios and 2,452 counterfactual variants across eight safety domains and eight risk-source types, scored by an evidence-aware judge across three prompt probes. Across eleven models, natural severity accuracy spanned 27.6% to 52.8% and relaxed diagnostic consistency only 6.1% to 29.3%, with unsafe-to-safe counterfactual transitions succeeding 30.4% to 58.4% of the time, indicating models are often not safe for the right multimodal reason.

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu, Qile Su, Han Liu et al. cross-listed Answering open-world questions about videos often requires both locating sparse visual evidence inside the footage and retrieving knowledge that is absent from the clip and from the model's parameters, two capabilities usually developed in isolation. VideoRover interleaves video cropping, multimodal search and webpage browsing in one loop, using each tool result to choose the next action so that localized clips steer external retrieval and retrieved facts trigger further inspection of the video. Training rests on an automated curation pipeline producing 26K verified supervised fine-tuning trajectories and 3K reinforcement learning instances, plus VideoRover-Bench, a benchmark stratified by video duration and research difficulty. VideoRover-8B-RL matches proprietary models in the no-tool direct-answer setting and beats larger open-source models given the same tool suite on VideoDR and the new benchmark.

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke, Radu Tudor Ionescu cross-listed Audio-visual deepfake detectors tend to overfit the generation methods they were trained on, so the proposed remedy is to lean on many frozen pre-trained models rather than learn cues from scratch. DF-MoE extracts features encoding mouth movement, face parsing, facial expression, head pose, gaze, heart rate, audio emotion and speech activity, then fuses unimodal and multimodal cues through a sparse Mixture-of-Experts backbone. Across in-domain and cross-domain tests on MAVOS-DD, AVLips, PolyGlotFake, BioDeepAV and FakeAVCeleb, it surpasses all competing methods.

Towards Comprehensive Basketball Understanding

Yirong Hu, Jiayuan Rao, Yu Zhang, Shangzhe Di, Weidi Xie cross-listed Following a basketball game requires recognizing events, localizing actions, identifying players, and tying all of that to structured game knowledge, yet benchmarks test these skills separately. BasketballBench bundles 7,980 questions across ten text, image, and video tasks built from the 2025-2026 NBA season, including official play-by-play, profiles for 530 active players, and 2,501 possession-level broadcast clips. The accompanying BasketballSkills agent composes eight domain-specific perception and retrieval tools under four reusable skills that fix tool ordering, evidence bindings, and stopping conditions. Current multimodal language models falter most on questions requiring several capabilities at once, while the tool-composing agent outperforms them.

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

Marek Hradil, Danae S\'anchez Villegas Vision-language models score well on video benchmarks, but that may reflect single-frame perception rather than any grasp of temporal structure. TimeCatch recasts temporal grounding as anomaly detection: temporal anomalies are made by swapping consecutive frames, frame-level anomalies by replacing a frame with Gaussian noise, and models must both detect and localize them across four synthetic and real datasets alongside a human study. Models detect and often localize frame-level anomalies reliably but perform near chance on temporal anomaly detection, while humans are near ceiling on both; analyses across model scale, prompting, sequence length, and visual similarity rule out perception limits or capacity as the sole explanation.
10 more specialized papers

Reinforcement Learning 27

AI Learning and Conceptual Transfer in the Game of Hidden Rules

Christo Mathew, Wentian Wang, Jacob Feldman, Lazaros K. Gallos, Paul B. Kantor, Vladimir Menkov et al. The Game of Hidden Rules is a testbed where an agent must infer an unstated rule purely from trial-and-error feedback on its moves. This report covers reinforcement learning agents built on a Transformer-based advantage actor-critic (A2C) framework, comparing feature-centric against object-centric state representations and analyzing how rule difficulty affects learning, transfer between rules, and generalization. It also reports on classifying human learning data and on using pseudo-bots to assist analysis of how people solve the same puzzles.

Runtime Action Interference for AI Control of AlphaStar in StarCraft II

Jaymari Chua, Chen Wang, Liming Zhu, Lina Yao The behavior a user actually experiences from a deployed reinforcement learning agent depends not only on the trained policy but on the surrounding code that schedules, filters, or replaces its actions. Runtime action interference (RAI) is a control layer that leaves policy weights untouched and instead gates each proposed action behind a cooldown condition and a content detector for configured patterns such as worker-unit harassment, dispatching a no-op when either check fails; it was implemented in an open-source replication of the AlphaStar actor.py. A human study in StarCraft II presented the same rate-limited, high-capability opponent twice, once withholding the capability claim and once disclosing it, and found disclosure lowered perceived fairness from 3.90 to 2.62 and raised perceived toxicity from 2.00 to 2.85 even though the underlying control was identical. Trust rose among novices and experts but fell among intermediate players, leading the authors to argue that evaluations must separate execution-stack control from capability disclosure.

Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning

Qifan Shi, Zhaolu Kang, Chenghua Zhu Reinforcement learning on language models has to spread a final reward back over individual tokens, and the operators in common use — fixed-discount generalized advantage estimation, or broadcasting one group-relative outcome across a whole response — ignore what the Transformer actually computed on that trajectory. Computation-conditioned credit transport uses a detached statistic of the policy's own internal computation to shape the credit kernel; the concrete algorithm CompPO maps attention concentration to a per-token retention gate used in both bootstrapping and a path-dependent advantage trace, plus a critic that reuses the actor's hidden states instead of running a second same-size Transformer. Across five Qwen3-4B seeds it reaches 61.4% held-out accuracy versus 53.8% for tuned GRPO, with ablations showing the gated trace and the shared critic only work in combination, and stability in 10 of 12 grid runs versus 3 of 12.

Reading the Room: Implicit Confusion Encoding in Recurrent World Model States

Donald Aadithiyan Recurrent state-space model world models such as DreamerV3 train their hidden state purely to reduce prediction error, yet that state also tracks its own confusion, encoded almost orthogonally to the directions of greatest variance and so invisible to variance-based probing. The signal is functionally distinct from ensemble disagreement, which flags novel inputs, and from reconstruction error, which flags bad predictions now: on a test that holds prediction error fixed while varying confusion, a linear probe on the hidden state reaches 0.72 AUROC while the ensemble baseline scores below chance, and a discounted count of recent high-error steps explains 80% of the probe's output. Directly editing the hidden state changes behaviour, showing the signal is causally used rather than merely present, though the decisive dissociation test holds cleanly on only one of three control tasks and the practical use — deciding when to check reality instead of trusting imagination — transfers to two of three.

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

Qiqian Fu Reinforcement learning for vision-language math reasoning stalls when rewards are sparse: on a pool of 20,830 visual-math problems where Qwen2-VL-2B gets only 3.6% of rollouts right, 85-97% of GRPO rollout groups are entirely wrong and produce zero gradient. Eleven methods are trained under identical conditions, each injecting a different prior — reference-solution hints, on-policy distillation from a 7B teacher, or a value-pretrained critic — and the six arms whose prior actually reaches the policy separate cleanly from the five where it is capped, gated, or lost to a mis-parameterized critic. The headline result is about evaluation rather than method: an in-domain slice long used as the general-distribution check anti-correlates with true cross-domain transfer (Spearman rho = -0.74) because a near-chance multiple-choice subset rewards models for not having changed, while the hardest in-domain slice predicts transfer at rho = +0.89; separately, swapping the critic's MSE loss for HL-Gauss cross-entropy is worth 14.4 points in-domain.

The Chase Is the Curriculum, the Capture Anchors the Credit: Pursuit-Evasion Self-Play for Zero-Data LLM Reasoning

Jing Yu, Shengchao Chen, Yiyun Tan Reinforcement learning with verifiable rewards depends on large curated task sets, and zero-data self-play alternatives only test whether a generated task is learnable by probing and discarding it afterward, never learning where to aim on an environment's difficulty scale. LURE frames self-play as pursuit and evasion: a language-model evader places tasks along the difficulty axis while a planner-executor pursuer tries to solve them, with the evader rewarded on a capture-frontier signal that peaks when the solver succeeds on exactly half its rollouts, making "barely catchable" a learned strategy instead of a hand-tuned acceptance band. The pursuer receives dense process credit anchored to capture, group-normalized with monotone verifier progress under a round-anchored KL penalty for stable co-evolution. Across three verifiable reasoning environments and three backbone families it beats advanced baselines, and the unified model reaches higher aggregate out-of-distribution zero-shot accuracy than every trained baseline on nine held-out benchmarks.

ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents

Xiaoyu Wang, Qingqing Gu, Yue Zhao, Teng Chen, Yuqi Cao, Xiaokai Chen et al. Reinforcement learning for dialogue agents typically operates either at the token level or the utterance level, missing the way humans plan strategy and produce words at different timescales. ToSCA builds a two-level Markov decision process in which token-level response decoding is conditioned on an utterance-level action expressed as explicit textual strategy, solving the high-level critic with DQN and the low-level actor-critic with PPO. A dual-granularity reward combines an utterance-level satisfaction score with token-level intrinsic motivation and a KL penalty to counter reward sparsity, and experiments on both everyday and emotional support conversations show better strategy selection and response quality than a range of baselines.

Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments

Khang Luong, Nam Nguyen, Hoang Ta, Hung The Tran, Tuan Dam Pure exploration methods that allocate samples by mean estimates struggle when reward noise dominates the signal. VarDE (Variance Driven Exploration) instead allocates sampling effort to minimize the uncertainty of the final decision, formalized through a smooth decision function from which allocation rules are derived that track how noise in each component propagates to the output. The method is instantiated for Best Arm Identification, Monte Carlo Tree Search, and Best-Policy Identification with guarantees on variance decay and simple regret, and the authors report consistent improvements over existing baselines, with the largest gains in highly stochastic environments.

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

Junlin Chen, Ruijie Wang, Jianxin Li Reinforcement learning models usually predict entire future states or feature occupancies, spending capacity on high-dimensional dynamics that unfold the same way no matter which action the agent picks — even though action selection only depends on differences between candidate actions. The Counterfactual Quotient Model treats action-conditioned futures as equivalent when they differ only by a shared component and learns a centered representation that removes it while preserving every pairwise action comparison the reward family can express, training on synchronized counterfactual rollouts so shared stochastic dynamics cancel before function approximation rather than after full futures are predicted. The authors prove decision sufficiency, identifiability, common-mode invariance, approximation, and regret properties, and physics-based experiments show suppressed action-independent variation, support for previously unseen reward queries, and better action ranking than absolute-future models.

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

Ziyang Luo, Yan Yang, Xiangru Jian, Ziji Shi, Xiaoqiang Lin, Jun Hao Liew et al. Reinforcement learning improves language-model tool use, but most frameworks handle only the policy update, leaving users to build isolated environments for hundreds of concurrent trajectories and to schedule rollouts so GPUs are not idle while episodes wait on slow tool calls. MCP-U RL supplies both missing layers around the Model Context Protocol, so any existing MCP server becomes a training environment with no RL-specific glue: an environment-orchestration layer provisions, isolates, and recycles environments over a pluggable container backend, and a rollout layer overlaps trajectories through a staged pipeline to keep the GPU busy. Training is delegated to existing backends (veRL and slime), and changing only the task specification trains software-engineering, deep-research, and general tool-use agents on gpt-oss-20b with improved task reward in all three domains.

Risk-Sensitive Reinforcement Learning with Smoothed Quantile Objectives

Mohammad Alipour-Vaezi, Huaiyang Zhong, Sajad Khodadadian Optimizing a quantile of the return distribution makes reinforcement learning risk-sensitive, but exact quantiles are non-smooth and can jump under small changes to an estimated transition model. UCB-BQRL is a model-based optimistic algorithm that keeps confidence sets over the transition kernel and plans against a lower-buffered quantile, which averages nearby lower quantiles to smooth the objective, paired with EVI-BQ, an exact dynamic-programming routine for computing the buffered-quantile policy. The authors prove a high-probability regret bound scaling as roughly e^(τ/ρ_τ) + H²√(SAT), a matching-form information-theoretic lower bound of Ω(H/ρ_τ·√(AT)), and that exact point-quantile and lower-buffered quantile policy evaluation are PP-hard even for a fixed policy in a two-state, one-action finite-horizon Markov decision process.

Learning from the Test: Self-Referential Differential Testing for Deep RL Agents

Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng, David Lo cross-listed Testing of deep reinforcement learning (DRL) agents concentrates on safety-critical failures and largely ignores sub-optimal policies, in part because there is no oracle saying what the better action would have been. Delta manufactures one through differential testing: it first runs safety testing on the agent under test while logging its decisions, then trains a challenger agent on those logs with offline reinforcement learning, and flags an optimality bug wherever the challenger collects more cumulative reward. Across five environments and three offline algorithms (BC, BCQ, CQL), BCQ challengers worked best, and Delta surfaced an average of 2,518 optimality issues per environment, 50.2% more than baseline methods.

The Variance of Thought: Policy Variance, Critical Forks, and Local Credit Assignment

Yingru Li Long-horizon reasoning and tool-using agents are bottlenecked by credit assignment, analyzed here through policy variance — the variance of the action-value under the policy at a given state — which in a deterministic Markov decision process is the sole source of return variance and arrives in discrete pulses at states termed critical forks. Three consequences follow: policy variance acts as a discovery budget, since observing an action of advantage c takes on the order of c²/σ² draws; policy variance is bounded by the policy's Gini dispersion, giving a rollout-free criticality test computable from logits alone; and at a fork with downstream success probability P the Monte Carlo advantage estimate has signal-to-noise of order √P, so sample cost scales as 1/P, a cost branched sampling shares. Bootstrapping removes that cost by turning a product of survival probabilities into a sum, provided values are multiplicatively accurate — an argument for log-value parameterization.

Scaling Curriculum Learning For Autonomous Driving

Cevahir Koprulu, David Paz, Feng Tao, Yuliang Guo, Xinyu Huang, Ufuk Topcu et al. Batched driving simulators can feed reinforcement learning billions of interactions in days, but the standard practice of uniformly sampling scenarios from a fixed pool wastes most of those interactions on cases that teach the policy nothing. CL4AD treats scenario selection as an unsupervised environment design problem and introduces utility functions that build curricula from success rates and behavioral realism alongside existing regret-based estimators. In GPUDrive, curriculum learning reaches a 99% success rate a billion steps earlier than domain randomization, cutting wall-clock time by 77%, improves sample efficiency by 67% under a limited compute budget, and beats heuristic curricula in all but one setting at the largest scale.

DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue

Qi Zhang, Heajun An, Prakriti Dumaru, Sang Won Lee, Lifu Huang, Pamela J. Wisniewski et al. Language model counseling agents produce fluent, supportive text but wander rather than moving a therapy session through its intended stages. DeepSAGE splits the first session of Cognitive Behavioral Therapy (CBT) into eleven stages with explicit objectives, uses an external controller to judge when a stage is complete, and trains a deep reinforcement learning policy to pick the therapeutic intention that conditions each generated response. Against six retrieval-, prompting-, stage-, and policy-based alternatives it elicits higher simulated client engagement and openness with the best balance of stage-goal completion and dialogue efficiency, though the authors stress that evaluation used simulated clients and model-based metrics, so this shows dialogue control rather than clinical benefit.

LpWM: A Case for Sparse Representations in World Models

Yilun Kuang, Yash Dagade, Quentin Le Lidec, Lucas Maes, Randall Balestriero, Yann LeCun Joint-embedding predictive architectures learn latent dynamics for planning and avoid collapse by matching features to maximum-entropy distributions such as isotropic Gaussians, which yields dense representations whose suitability for dynamics modeling has gone untested. The paper proves that nonlinear Lipschitz dynamics can be approximated arbitrarily well by action-conditioned linear dynamics in a high-dimensional one-hot latent space with rollout error vanishing as dimension grows, then relaxes that into LpWM, a model regularized by Rectified Distribution Matching Regularization to produce non-negative sparse codes. On PushT, sparsity improves planning success by up to 57% over the dense LeWM baseline at intermediate predictor capacities, beats dense VICReg representations across predictor families, and produces mode-factored codes whose support encodes discrete dynamical regimes while magnitudes track continuous within-regime state.

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen et al. Large-scale rollouts for reinforcement learning post-training, on-policy distillation, and sampling-heavy evaluation are bottlenecked by a handful of very long generations, because uniform routing drops those long tails into high-concurrency decoding batches and stretches the whole step's makespan. TailSieve uses partial rollouts as a training-free signal that long-tail prompts stay long-tailed across policy updates, then a hierarchical controller adapts how many tail groups to isolate and how to split replicas between tail and bulk pools using response-work history and a measured concurrency-throughput model. Routing alone gives up to 1.67x speedup over uniform group routing, and because the isolated tail pool runs at low concurrency it can also use speculative decoding with MTP or DFlash for up to 2.59x end-to-end speedup, with selected prompts regenerated under the current policy to keep generation on-policy.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

Yinhao Tang, Youqing Fang, Yanan Sun, Jiangning Liu, Ziyi Wang, Xun Zhao et al. Next-chunk reasoning reinforcement learning extracts training signal from no-chain-of-thought corpora such as worked solutions and textbook derivations by having a model generate implicit reasoning and rewarding it for predicting the following chunk of text, but prior evaluations only compared it against conventional supervised fine-tuning, leaving unclear whether the reinforcement learning formulation or simply better exposure to the data produced the gains. A controlled comparison pits it against Mixed SFT, a single supervised stage trained jointly on no-chain-of-thought and long-chain-of-thought data. Mixed SFT reaches a clearly higher performance ceiling after reinforcement learning with verifiable rewards while using over 60 times less training compute, consistently on in-domain mathematics and out-of-domain reasoning. Higher accuracy before the reinforcement learning stage did not reliably predict higher accuracy after it, so training strategies need to be judged across the whole post-training pipeline.

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang et al. Reinforcement learning for language models trades stability against exploration, typically mediated by a policy-side Kullback-Leibler penalty that constrains response behavior and eats the exploration budget, while removing it leaves drift uncontrolled. ERPO moves regularization to the input side, adding a Query-KL term that bounds how far the policy-induced distribution over training queries drifts from its pre-RL reference, together with a dataset-static per-query weight biasing updates toward queries typical under the reference; because the Query-KL gradient flows only through query likelihood, it exerts no direct pressure on the response distribution. It drops into GRPO, PPO, and REINFORCE-style pipelines without extra forward passes, and on six mathematical reasoning benchmarks it replaced the standard policy-KL regularizer with stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.

Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Zixuan Wang, Yanrui Miao, Zhengxi Lu, Teng Pan, Yiwen Qiu, Hongxing Li et al. Hint-based reinforcement learning fights sparse rewards in long-horizon agent tasks by seeding each rollout with a prefix of an expert trajectory, and the key knob is guidance depth — how much of that trajectory to keep — which existing methods treat as a single deterministic number, either shared across all samples or estimated per sample using extra probe rollouts. The authors observe that useful guidance spans a band of depths with an approximately Gaussian informativeness profile, so Agent-G² samples depth per task from a Gaussian whose center mixes a global baseline with per-cluster difficulty and whose spread tracks within-cluster variance, both estimated online from rollouts already collected for policy optimization. On ALFWorld and WebShop with Qwen2.5-1.5B and 7B-Instruct, it beat the strongest hint-based, hint-free, and auxiliary-RL baselines on ALFWorld by 2.3, 3.9, and 7.4 points at under one-third the rollout cost of per-sample probing.

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

Jialong Liu, Yuling Shi, Ning Yang, Xiaodong Gu, Zuchao Li Outcome-only rewards give language model post-training a sparse credit-assignment signal, especially over long agent trajectories where a single terminal success or failure must explain hundreds of tokens. Self-Reflective Policy Optimization (SRPO) has the model analyze its own finished trajectories, compress the errors into short "reflection patches," and then use reflection-conditioned teacher scores on the student's own on-policy rollouts as dense token-level supervision — no external critic, reward model, or larger teacher required. Starting from Qwen3-8B, it reaches 73.3% on AIME'24 using 8% of the training FLOPs of scaled supervised fine-tuning, alongside 64.7% on WebShop, 76.8% on ALFWorld, and 31.2% on SWE-Bench-Lite.

How to Train a Critic Stably and Efficiently

Penghui Qi, Xiangxin Zhou, Wee Sun Lee Group-based reinforcement learning methods like GRPO skip the critic by sampling several responses per prompt, whereas a reliable critic could produce token-level advantages from a single response — except that critic-based recipes tend to train unstably. Best-Practice Critic Optimization (BPCO) combines DPPO, value predictions clipped to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive generalized advantage estimation, with ablations isolating each choice; because the critic is only used at training time, it can also see reward-defining information hidden from the policy, such as a reference answer or grading rubric. On mathematical reasoning with models from 1.5B parameters up to a 30B-A3B mixture of experts, BPCO matches or beats a group-based baseline while sampling only one response per prompt, and also improves training under rubric-based rewards.
5 more specialized papers

Vision 26

RiskWorld: Object-Centric Latent World Modeling for Autonomous Driving Risk Identification

Jingzheng Li, Yufei Ge, Qianren Mao, Zhijun Chen, Bing Li, Xingyu Peng et al. cross-listed Driving risk identification asks which visible object is about to become safety-critical to the ego vehicle, but existing methods predict scene-level accidents, infer risk indirectly from ego behavior, or run geometric checks after trajectory forecasting rather than using predicted ego-object relations directly. RiskWorld is an object-centric latent world model that pairs pretrained predictive video representations with structured ego-object interaction histories and rolls relation-aware object states forward using RSSM-style latent dynamics, decoding each imagined rollout into per-object risk scores with auxiliary future-relation and temporal-risk heads. Inference uses only observations up to the present, with logged futures serving as training supervision. On RiskBench it reaches the best overall F1 of 63.0% with the lowest false-alarm rate of 2.1%, and analysis shows the rollouts track rising object-level risk before critical events.

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

Yuqian Zhou, Zhenghong Zhou, Zongze Wu, Cameron Smith, Richard Zhang, Jiebo Luo et al. cross-listed Combines several video creation and manipulation tasks — text-to-video, image-to-video, video-to-video, editing propagation, reference-guided editing, and camera pose change — into a single diffusion transformer (DiT) model with task-specific conditioning, then converts it into a few-step autoregressive model for streaming use. The distillation runs in two stages: Velocity Moment Matching (VMM) matches conditional velocity moments at states the student actually reaches, preserving quality and motion, while autoregressive unrolling exposes the student to its own predictions to stabilize output over time. The authors report that the combination alleviates over-saturation, degraded motion, and temporal instability that typically afflict few-step autoregressive video generation, making EditStream usable in interactive creative workflows.

Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

Nai-Xin Zhai, Weihua Cheng, Dexu Yu, Yikai Gu, Hanwen Du, Junchen Fu et al. cross-listed Targets three weaknesses in aligning video generators with human preference: noisy and biased preference data, scalar reward models that collapse multiple preference dimensions into one number, and Kullback-Leibler divergence constraints in policy optimization that only capture local rather than global preference structure. The proposed framework applies elite-guided filtering to clean preference data, models video quality as a multidimensional reward distribution, and uses the Wasserstein distance both to fit that distribution to empirical human judgments and to replace the divergence term inside GRPO. Experiments on reward modeling and video generation show improved reward-signal reliability and perceptual consistency of generated videos.

LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

Riadul Islam, Joey Mule, Dhandeep Challagundla, Shahmir Rizvi, Sean Carson, Rachit Saini cross-listed Event cameras produce sparse, low-latency signals well suited to energy-constrained perception, but their asynchronous and noisy streams are usually handled by models too heavy for embedded hardware. LiteEvent-AE is a compact configurable event-driven autoencoder pairing lightweight convolutional encoding with adaptive event thresholding and a minimal classifier head, compressing neuromorphic data while retaining the spatiotemporal structure downstream inference needs. On the Smart Event Face Dataset and Event-Based Crossing Dataset it matches or beats YOLOv9 accuracy with up to 35.6x fewer parameters, runs at 44.8 frames per second on an NVIDIA Jetson Nano, and on a Raspberry Pi 4B CPU consumes about 726.3x less energy than YOLOv9 under the same evaluation protocol.

A Scalable Vector Graphics Latent Space

Leonardo Zini, Elia Frigieri, Lorenzo Baraldi cross-listed Vector graphics lack the kind of continuous, invertible latent space that variational autoencoders gave raster images, leaving generative and retrieval work on SVGs stuck with token sequences. SLS is a Transformer autoencoder that embeds individual SVG paths — the atomic elements any SVG is built from — encoding drawing commands, coordinates, and visual properties in one byte-pair-encoded vocabulary, and decodes fixed-size latents back into valid, style-consistent paths. Embeddings lie on a unit hypersphere, which makes similarity search, composition, and downstream conditioning simple vector operations, and across several tasks the representation cuts FLOPs by more than 150× compared with token-based approaches.

CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders

Xingtao Lin, Hangqi Ren, Caiwan Sun, You Chen cross-listed Choosing a pretrained image encoder for medical classification is a real modeling decision, and clean-test discrimination hides differences in calibration, label efficiency, and stability under acquisition shift. CRS-Bench evaluates 15 encoder families across dermatology, ophthalmology, and radiology using ISIC 2019, APTOS 2019, and CheXpert, with CheXpert-to-MIMIC-CXR as an observed institutional shift, producing 17,575 controlled runs summarized by a Pareto-aware Clinical Reliability Score that blends dominance, profile balance, and worst-axis performance. AUROC and the reliability score correlate but are not interchangeable: 21 of 105 pairwise encoder orderings reverse, with mean absolute rank displacement of 1.87, and bootstrap analysis places PanDerm, MedSigLIP, and MedGemma in a shared leading tier rather than crowning a single winner.

ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology

Duncan Stothers, Ren-Chin Wu, William Lotter cross-listed Attention-based multiple instance learning over pathology foundation model embeddings works well for slide-level prediction, but requires running a large image encoder on every foreground tile even though attention typically concentrates on a few regions. ADMIL distills the teacher's attention distribution into PriorNet, an EfficientNet that learns to score raw tile pixels via KL divergence, selects the top-K tiles, and only then invokes the expensive encoder on that subset before a student model predicts the slide label. On BRACS, PANDA, and CAMELYON16 it matches full-teacher performance at K of 4, 8, and 128 tiles respectively, skipping over 98% of Virchow2 tile embeddings and their inference FLOPs, with random-selection and oracle controls confirming the gain comes from task-relevant selection rather than merely using fewer tiles.

Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos

Xiaoyang Liu, Kai Han cross-listed Recovering how a deformable object behaves physically from a single fixed-viewpoint video is unstable: implicit methods fall into local minima under noisy supervision and offer no interpretability, while explicit ones need predefined constitutive equations that limit generalization. GCA represents the object as 3D Gaussians initialized from a static multi-view scan, then learns implicit constitutive laws with LoRA-based adaptation plus two alignment modules — Rank-based Depth-Geometric Anchors, which impose scale-invariant rank-based depth constraints instead of relying on unreliable pixel color supervision, and a Constitutive Prior Regularizer that folds classical material models in as soft differentiable priors even when the true material is not among them. On synthetic benchmarks it achieves 48% lower Chamfer Distance than the strongest baseline, with results also reported on real-to-sim and real-world data.

When Test-Time Adaptation Helps, Harms, or Becomes Inactive: A Condition-Level Study on CIFAR-10-C

Sreeja Guha Majumdar, Aratrika Saha Test-time adaptation updates a trained model on unlabeled test data to survive distribution shift, and reported average gains can hide the specific conditions where it backfires. This controlled study compares BN-Adapt (BatchNorm statistics only), TENT (entropy minimization), and a scoped reimplementation of EATA (reliability-filtered adaptation) against an unadapted source model across all 15 corruption types and 5 severities of CIFAR-10-C. All three raise mean accuracy by 12.2 to 13.3 percentage points, but each is worse than the untouched source model on 8.0-9.3% of conditions, concentrated in low-severity corruptions such as brightness, fog, contrast, and defocus blur where the source model is already near ceiling. EATA tracks the gradient-free BN-Adapt baseline to within 0.09 percentage points on average versus 1.08 against TENT, suggesting its reliability filter suppresses most actual adaptation.

Hyperbolic Hierarchical Clustering for Visual Representation Learning

Jianan Wei, Guikun Chen, Zhiyuan Weng, Chunchao Guo, Yujia Wang, Wenguan Wang cross-listed Token mixers in vision backbones — convolution, attention, MLP, and hybrids — are tuned around the accuracy-versus-cost trade-off but offer no account of what the mixing step actually computes. ClusterMixer replaces them with explicit hierarchical clustering carried out in hyperbolic space, chosen because it embeds the tree-like structure of visual data with low distortion, making the encoding process inspectable by construction. Built on it, the HCFormer backbone consistently outperforms comparable backbones on image classification, object detection, instance segmentation, and semantic segmentation, which the authors present as evidence that an interpretable mixer need not sacrifice accuracy.

A Comparative Study of Label-free Representation Quality Metrics in Deep Learning

Daniel Richards Arputharaj, Daniel J\"onsson, Gabriel Eilertsen Label-free metrics promise to judge how good a neural network's learned representations are without needing downstream labels, but it has been unclear which ones to trust. The authors sort existing metrics into three families, derive analytic connections among members of each family, probe the sensitivity of spectral metrics with controlled synthetic experiments, and then correlate every metric against downstream accuracy across 260 vision models on six datasets covering generic and fine-grained classification, scene recognition, and a geospatial task. Intrinsic dimensionality comes out as the most reliable predictor, though the reliability of every metric — including intrinsic dimensionality — depends on architecture class and training objective.

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

Federico Stella, Fei Jiang, Zhongshi Jiang, Zohar Barzelay, Emanuel Garbin, Amin Jourabloo et al. cross-listed Generating photorealistic new camera viewpoints of human faces is hard at high resolution, where identity, fine detail, and geometric agreement across views must all hold. The approach adapts next-scale autoregressive image generation — which predicts progressively finer scales rather than denoising — to produce multiple novel views in a single forward pass, trained on a synthetic dataset of diverse identities and apparel. Because the coarse-to-fine architecture can inherit general-purpose low-resolution pre-training and only needs full-size domain images in the final stages, the model converges on a smaller but more realistic purpose-specific dataset than diffusion approaches require. The outputs are further fed to an existing transformer that lifts multi-view facial inputs into pixel-aligned 3D Gaussians.

ChebBooster: A Training-Free Approach for Efficient Diffusion Transformer Inference via Chebyshev-Inspired Extrapolation

Chengjie Lu, Tianchi Deng, Zhengqi He, Chengwen Luo, Xueliang Li Diffusion transformers produce high-fidelity images but run the full model at every sampling timestep, and existing training-free shortcuts either reuse cached activations too crudely or use Taylor-series extrapolation that destabilizes through Runge oscillations. ChebBooster replaces the extrapolation basis with Chebyshev polynomials evaluated via the numerically stable barycentric formula, splitting the work into offline weight precomputation and a cheap online application step. On DiT-XL/2, PixArt-Σ, and FLUX.1-dev, it reaches up to 3.68× latency speedup and 5.12× fewer FLOPs while improving visual quality over other training-free baselines across resolutions and tasks.

ReWorld: An Interactive World Model with Long-Horizon Memory

Zhifei Chen, Luozhou Wang, Guibao Shen, Dongyu Yan, Shuai Yang, Tianshuo Xu et al. Interactive world models face a structural conflict: precise action following favors a short attention horizon while remembering previously visited places needs an unbounded one. ReWorld splits the two during training using mixed per-head attention windows — most heads see only the recent past while a few global heads span the full history — with random head routing so neither ability latches onto fixed heads, and random chunk dropping so sparse histories look normal at inference; at run time the past is held in a bounded key-value cache backed by a pose-indexed landmark bank that retrieves whatever is nearest the current camera pose. Training data from eight sources, including Unreal-rendered fly-throughs, game footage, and real video, is placed on one physical action scale so the same key press moves the camera the same distance everywhere, and distribution-matching distillation inside a LoRA adapter cuts sampling to four steps for real-time 704x1280 streaming. Against six recent interactive world models it reports the best control fidelity (11.95 degrees rotation error) and generation quality, and on 64-second out-and-back rollouts its fixed 12-chunk cache still regenerates the starting view, where sliding-window attention has evicted the evidence and full attention runs out of memory.
12 more specialized papers

Unclassified 20

Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments

Ryuki Hyodo cross-listed No summary available — see the abstract on arXiv.

Mirror descent algorithms with logarithmic barriers

Alberto De Marchi, Yura Malitsky, Adrien B. Taylor cross-listed No summary available — see the abstract on arXiv.

Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning

Shuting Xie, Nathaniel Lesperance, Graham W. Taylor No summary available — see the abstract on arXiv.

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

Hang Wang, Jin Zhang, Guoliang Xu, Pengyue Lu, Yao Li, Zijiao Zhang et al. No summary available — see the abstract on arXiv.

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao, Kun Huang, Pengzhi Gao et al. No summary available — see the abstract on arXiv.

RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling

Ziyuan Wang, Bohao Tang, Fei Zhang, Shuo Han, Pengfei Liu No summary available — see the abstract on arXiv.

Your AI, On a Dial: Controlling Investment Bias in LLMs with a Single Neuron

Sahong Park, Suhwan Park, Hoyoung Lee, Gakyung Kwon, Wonbin Ahn, Jaewon Choi et al. No summary available — see the abstract on arXiv.

Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs

Yan Zhou, Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak, Marco Fumero, Francesco Locatello et al. No summary available — see the abstract on arXiv.

Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA

Jingjie Ning, Xueqi Li cross-listed No summary available — see the abstract on arXiv.

SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning

Youdi Li No summary available — see the abstract on arXiv.

Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning

Dongyue Wu, Tao Ma No summary available — see the abstract on arXiv.

WARP: Wasserstein-Aligned RAG for Population Opinions

Aman Singh Thakur, Aditya Agrawal, Alwarappan Nakkiran, Alex Karlsson cross-listed No summary available — see the abstract on arXiv.

Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors

Zhenghua Bao No summary available — see the abstract on arXiv.

Stochastic Separability of Embedding Manifolds

Liqing Zhang No summary available — see the abstract on arXiv.

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang, Minseo Kim No summary available — see the abstract on arXiv.

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

Zengqing Wu, Chuan Xiao No summary available — see the abstract on arXiv.

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

Mo El-Haj No summary available — see the abstract on arXiv.

SelFusion: Self-distillation for Diffusion Language Models

Hyeongsoo Lim, Jinyoung Kim, Eunseo Seo, Minho Jang, Jiwon Yoon No summary available — see the abstract on arXiv.

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Ziru Niu, Zuozhu Liu No summary available — see the abstract on arXiv.

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho, Huu-Kim Nguyen No summary available — see the abstract on arXiv.

Robotics 19

RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception

Oguzhan Baser, Mirac Sozen, Kaan Kale, Sandeep Chinchali, Sriram Vishwanath cross-listed Robots that scan and share 3D point clouds for fleet learning, cloud planning, or collaborative mapping leak spatial context such as room function that occupants never agreed to disclose, and standard encoders offer no dial between preserving everything and nothing. RoboShape attaches an information-theoretic compression head to the frozen Sonata encoder, projecting voxel-level embeddings under a Donsker-Varadhan mutual information objective that maximizes information about object-level semantics while minimizing it for designated private attributes. Across three real-world indoor LiDAR datasets the embeddings are 87.5% smaller while retaining 98.7% of object classification utility and cutting sensitive attribute prediction by 39.3%, and the encoder-agnostic codebase is released.

ODG-NoMaD: Overhead-Camera Direction-Guided NoMaD

Blossom Treesa Bastian, Keerthi S. Shetty, Manish Kolachalam, Rani Malhotra, Ashish Dutta cross-listed NoMaD is a diffusion-based visual navigation policy that handles both goal-seeking and exploration, but in an unfamiliar environment with no goal image or topological map it can only wander without any global sense of direction. ODG-NoMaD adds that global sense without retraining: an overhead depth camera is used once at deployment to build an occupancy map and plan a global path, which is segmented into a desired heading, then refined against a per-frame traversability map from the robot's own depth sensor into a collision-free direction. The gradient of a cosine direction cost is injected into the final denoising steps so sampled trajectories rotate toward that heading while remaining multimodal. In simulated offices with and without random obstacles, the method cuts residual distance to the target by up to an order of magnitude versus unguided exploration, beats the point-goal cost guidance of NaviDiffusor, and is the only configuration that stayed collision-free on every trial.

Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts

Farida Mohsen, Thowayba Elkaffash, Mohammad Reza Chalak Qazani, Mohamed Mabrok, Nader Meskin, Ali Safa cross-listed Vision-language-action models trade success rate against how often the policy must be re-invoked: replanning after every action is accurate but too slow for real-time robot control, while executing long action chunks before replanning is cheap but degrades reliability. Replacing the causal self-attention in the action expert of SmolVLA with Mamba selective state-space modeling was evaluated on the LIBERO benchmark at execution horizons of 1, 25, and 50 actions. The Mamba expert's advantage grows with horizon length, outperforming the Transformer baseline by 7.8% at the 50-action horizon that corresponds to feasible real-time deployment and by 3.7% at 25 actions, while matching it at per-action replanning with 24% fewer parameters.

Operational digital twin clinics enable task-based evaluation of embodied AI

Xinyuan Wu, Jingrao Zhang, Mengdi Xu, Henry K. Chu, Mingguang He, Danli Shi cross-listed Testing embodied AI in the clinics where it will actually operate requires realistic robot-testable environments that are expensive to build and hard to scale. Single photographs from 39 ophthalmic clinic scenes were converted into editable, simulator-ready digital twins, with device meshes, collision proxies, and semantic anchors turning visual reconstructions into contact-aware simulation scenes suitable for closed-loop policy evaluation. Reconstructions preserved workspace structure and supported local editing for controlled device reconfiguration; across three robot embodiments, the same task targets showed different reachability and contact feasibility, and small device translations and rotations changed contact margins in ways visual similarity metrics failed to capture. The authors propose operational validity — whether the twin supports the task, not whether it looks right — as the governing criterion for clinical digital twins.

Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics

Ran Chen, Jiaxing Ren, Zhikun Zhang, Yunhao Hou, Junbao Zhuo, Bochao Zou cross-listed Vision-language-action (VLA) driving planners underperform in complex environments because image-only inputs poorly capture road geometry and topology. Geo-VLA is a plug-and-play training-time addition that internalizes geometric map semantics into visual representations, so no high-definition maps or lane information are needed at inference; it is supported by Geo-QA, a geometry-focused question-answering dataset used for contrastive learning and instruction tuning. On NAVSIM v1 the method improves VLA planners across different action-generation architectures and reaches 92.1 PDMS, a new best among single-camera VLA planners.

Inferring Action from Future Latent State for Robotic Manipulation

Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen, Jie Cheng, Peilin Huang et al. cross-listed World-Action Models for robot control are built on video-generation backbones that predict dense future frames alongside actions, but the intermediate frames describe visual transitions rather than the physical outcome an action is meant to produce, consuming capacity and compute for no control benefit. DELE-w0.5 drops video generation entirely and infers action sequences from a compact predicted future latent state, which serves as the explicit bridge between world modeling and action generation and enables cheaper training and lower-latency inference. Across 480 real-robot trials on four long-horizon manipulation tasks it reaches 62.5% full-task success, 47.5 percentage points above the strongest baseline, along with 81.3 macro ordered-stage progress.

The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

Xunzhe Zhou, Yiyang Cai, Fengyi Wang, Ran Ju, Hanxiang Ren, Ruizhe Liu et al. cross-listed Robot policies learn observation-to-action mappings and can replay a demonstrated motion in a near-identical scene, but never explicitly infer what the demonstrator was trying to accomplish. The Imitator Game is a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, paired with IG-10K — 20,000-plus paired human-robot episodes across 50-plus tasks, six domains, real and simulated — and Imitator Arena for blind A/B human evaluation. Nine state-of-the-art models hold steady from L0 through L2 but collapse at L3, where the same intent has to be achieved through a different object affordance, and none exceeds 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models on as few as ten paired demonstrations gives large gains that grow with pretraining scale.

WAM-OPD: On-Policy Distillation for World Action Models

Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang World action models pair visual future prediction with robot action generation, but distilled fast students lose task capability and then encounter states that offline data covers poorly. WAM-OPD applies on-policy distillation as a deployment-consistent post-training recipe: the student acts in the environment and thus sets the history distribution, a frozen teacher labels those student-induced histories with coherent video and action targets, and the student's action branch trains under its own generated video plan just as at deployment, updating lightweight adapters plus an action flow-matching regularizer. In preliminary RoboTwin 2.0 studies, the one-video/one-action-step Flash-WAM goes from 0.0% to 58.3% success on HANDOVER MIC and from 16.7% to 33.3% on PUT OBJECT CABINET, which the authors present as a two-task capability proof rather than evidence of broad generalization.

EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting

Wen Wang, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen cross-listed Predicting future hand motion from head-mounted camera views is usually done by mapping observations straight to trajectories, which skips the manipulation process driving hand-object interaction and lets motion-synthesis gradients corrupt the learned interaction representations. EMPIRE splits this into two stages: a planner first learns explicit manipulation plans from multimodal context, then a motion generator synthesizes bimanual hand motion conditioned on frozen planner features so gradients cannot flow back. The authors also release EMPIRE-651K, 650,910 training windows across 111 tasks each paired with a per-hand manipulation plan, and report state-of-the-art accuracy at 84.53 mm mean per-joint position error with 38.97 mm finger-relative error.

Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs

Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake cross-listed Foundation-model planners can decompose an open-ended goal into steps for a robot team, but supplying richer detail about embodiment, preconditions, and coordination does not stop them from proposing infeasible, mistimed, or unsafe physical actions. The proposed architecture inserts a deterministic Robot Orchestrator between a non-actuating Mission Planner and the hardware: each robot publishes a typed library of executable skills, and every planned skill is checked against robot state, named locations, and workflow contracts before exactly one action is authorized. Evaluated on a drone-and-ground-vehicle search-and-dispatch mission run live in Gazebo and a humanoid-quadruped transport task with physical Unitree G1 and Go2 trials, retrieval lifted skill grounding from 51% to 96% yet well-informed planners still dispatched 23–29% faulted steps. Per-dispatch enforcement drove false dispatch to 0% with no false blocks, refusing all eight injected faults before any motion, whereas without the gate six of them produced actual robot movement.

Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

Jianxiang Liu, Gaojing Zhang, Chuan Wen, Qipeng Liu, Yuxuan Zhao, Ning Guo et al. cross-listed Long-horizon robot manipulation faces a tradeoff between end-to-end vision-language-action models, which need large robot datasets and resist diagnosis, and hierarchical pipelines, whose plans are poorly grounded in what the robot actually sees and run open-loop. The Triplet-to-Track System learns from human videos instead of robot-collected demonstrations, expressing each subgoal as an instance-grounded triplet, converting it into continuous track priors that drive low-level execution, and monitoring progress from observations so it can replan online. Across a range of real-world long-horizon tasks it reaches a 74.8% average success rate with object-level and compositional generalization.

Shaping the Evolutionary Dynamics of Robot Morphology via Adaptive Control Learning

Junru Song, Yang Yang, Yaqing Xu, Ying Wen, Wei Peng, Guozhen Li et al. cross-listed Robot co-design alternates between learning a controller for each candidate body and evolving morphologies across generations, and prior work established that good bodies learn control faster — a property called morphological intelligence. Splitting a body's contribution to learning into convergence speed and a separate performance ceiling termed true potential, the authors fit both quantities from individual learning curves and show that cutting control learning short systematically underestimates true potential and biases selection toward fast learners, shrinking design diversity; the widely cited morphological Baldwin effect turns out to be an artifact of this bias rather than a real evolutionary trend. Their AdaControl method detects this skewed selection and allocates just enough control learning for fair fitness evaluation, letting a plain genetic algorithm match generative-model-based co-design on simulated voxel-based soft robots while cutting compute by up to 80%.

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Xiwen Chen, Zelin Li, Zhiruo Zhou, Huiming Chen, Chenwei Wang, Xiaojun Zhu cross-listed Vision-language-action models typically emit spatial targets either as autoregressive text coordinates or as opaque action tokens, both fragile interfaces between reasoning and robot control. Pointing-VLA instead reads geometry directly out of hidden states with dedicated heads that predict normalized points, object-functional grounding heatmaps, and visual trajectories, with an execution contract routing PICK to grounding heatmaps and PLACE to pointing. It reports state-of-the-art 72.9% average on a four-task Bridge/WidowX suite without Bridge-specific finetuning, runs 6.7–6.9 times faster than Embodied-R1 text decoding, and when used as spatial guidance for a π0.5 action policy raises real-robot autonomous success from 52.7% to 80.7% across three visual contexts.

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

Zhiruo Zhou, Zelin Li, Xiwen Chen, Jiazhuo Li, Chenwei Wang, Huiming Chen et al. cross-listed Retrieval can augment a frozen vision-language-action (VLA) robot policy without retraining, but any retrieved text that reaches the executed prompt is itself a control intervention. A matched audit found that appending raw text dropped mean success from 92.47% to 3.00%, with semantically meaningful and length-matched meaningless appends both failing on all 500 states — a failure of prompt form rather than content. TOWN-VLA adds a prompt-authority interface that separates generating a candidate instruction from permission to alter the policy input: a fixed compatibility rule either authorizes a canonical compact instruction or restores the original prompt byte for byte, verified across 900 audited routes. On a matched LIBERO-Plus evaluation with 10,030 episodes per method success rose from 69.5% to 73.1%, and on a physical PiPER arm with a frozen checkpoint from 52.7% to 78.7% over 150 trials.

Reward-Free Continual Adaptation for Resilient Space Robots

Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez cross-listed Robots in space suffer hardware degradation that breaks fixed control policies, and continual reinforcement learning would fix this except that it needs a reward signal at deployment time — often uncomputable in orbit without external tracking. The proposed framework pre-trains a model-based agent across diverse simulations so its latent-state world model internalizes the reward structure, then freezes the observation encoder and reward predictor at deployment and updates only the transition dynamics through unsupervised rollouts. The policy is retrained entirely on imagined trajectories from the updated world model, adapting to altered dynamics without any new reward observations, demonstrated on simulated planetary traversal, orbital navigation, and precision assembly under severe morphological failures.

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Sangoh Lee, Sangwoo Mo, Wook-Shin Han cross-listed Vision-Language-Action (VLA) models learn robot control mostly through behavior cloning, which records which motor command was demonstrated without capturing the local goal that command was serving. Intention Distillation (INDI) has a frozen teacher vision-language model interpret a demonstrated segment using the observation, instruction, coarse action summary, and execution video, then trains the deployed policy to recover that multimodal intent representation at an intermediate decoder layer from its normal inputs alone. On SimplerEnv-Bridge it lifts GR00T-N1.7 from 64.3% to 84.7%, with smaller gains on RoboCasa Kitchen, consistent improvements for π0.5, and real-robot average success rising from 62.0% to 68.7%.
3 more specialized papers

Reasoning 14

ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning

Weihang Pan, Zhengxu Yu, Yuxiang Zhang, Wenzhi Li, Zhongming Jin, Binbin Lin et al. Large reasoning models tend to overthink, producing chain-of-thought traces padded with redundant steps, and existing length-penalty rewards often yield "pseudo-conciseness" where token counts drop but the redundant reasoning structure survives. ChainPrune restructures self-generated reasoning paths into a tree, then applies multi-criteria dominant-path selection to build preference pairs that favor shallow trajectories retaining the essential steps, training with direct preference optimization combined with a supervised loss to counter false reward suppression. Reported results show sizable cuts in step count and compute while accuracy holds or improves.

Beyond Fixed Directions: Adaptive Representation Analysis of Reasoning and Memorization in LLMs

Shaheen Nabi Prior work suggests that reasoning versus factual memorization in language models is captured by a single fixed representation direction, sometimes frozen during reinforcement learning. Testing this on Qwen3-0.6B with a controlled 400-example dataset, a one-dimensional projection matches a full 1024-dimensional linear probe at AUROC 1.00, confirming single-direction decodability. After GRPO (Group Relative Policy Optimization) training, however, the direction is substantially reorganized — mean-direction cosine similarity drops to 0.453 while probe AUROC stays at 1.00 — indicating the information survives but its geometry does not, which undercuts methods that hold the direction fixed.

Semantic Reasoning Denoising: Correcting Language Model Reasoning with Semantic Operators

Yujiao Yang Fluent chain-of-thought traces can contain local semantic errors that propagate into wrong answers, and unconstrained self-correction may leave those errors in place or add new ones; diffusion language models offer iterative refinement but define noise as token masking or replacement rather than as reasoning errors. Semantic Reasoning Denoising (SRD) models noise as executable error operators specifying the error type, its location, and the corrupted and repaired propositions; composing operators builds progressively noisier trajectories, training teaches the model to detect the active noise and reconstruct the adjacent cleaner state, and inference repeatedly predicts and applies an inverse operator only where applicable. Across six in-domain benchmarks in mathematics, code, knowledge, and commonsense it beats the strongest same-backbone baseline by 3.2 points on average, and improves the strongest Qwen3-8B baseline by 2.9 points across seven cross-dataset transfer targets.

Decoupled Physical Modeling and Execution for Physics Reasoning

Ye Zhang, Xuehang Guo, Rui Pan, Pengfei Yu, Denghui Zhang, Manling Li et al. Physics problems entangle building a model of the underlying physical system with the mathematics of solving it, which is part of why language models strong at math and code still stumble on them. The proposed framework distills intermediate representations that explicitly encode the physical modeling step, then applies two-stage post-training: supervised fine-tuning to establish structured modeling, followed by reinforcement learning with rubric-based feedback to improve modeling quality. On the PhysReason, PhyX, and SeePhys multimodal benchmarks, explicit physical modeling outperforms GRPO by roughly 3 points on average, with consistent gains across models and datasets and the largest practical value claimed for smaller models.

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu Four open-weight instruction-tuned models plus frontier models are tested on four reasoning benchmarks under realistic lexical corruption — keyboard noise, character swaps, and filler insertion. Character-level perturbations substantially degrade accuracy, especially on multi-step reasoning, while filler insertion barely matters, and the asymmetry traces to attention diversion: corrupted words fragment into extra subword tokens whose fragments attract disproportionate attention mass, concentrated in middle and final transformer layers. Length-matched controls rule out prompt length as the cause, and a factorial intervention shows token content and attention allocation are corrupted together — restoring clean attention while content stays corrupted is actively harmful, restoring content alone is insufficient, and only restoring both recovers a substantial share of the gap. That coupling explains why chain-of-thought prompting, spell-checking, self-repair, and stronger repair models each fail to recover performance consistently.

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

Murat Dura, Serkan \"Ozt\"urk, Selma Tekir Chain-of-thought (CoT) prompting improves problem solving in large language models, but because the reasoning unfolds over many generated tokens, activation patching at a single fixed token position cannot locate where the causal effects live. The proposed sequential activation patching framework traces CoT-conditioned attention-head activations across token positions, aggregates effects using part-of-speech-guided analysis, and adds Sequential Multi-Head Patching to test joint contributions of head sets against cross-question and random controls. Targeted zero-ablation confirms the identified heads matter for producing correct answers and shows they touch several overlapping functions — trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation — pointing to distributed reasoning-support sub-circuits rather than a single localized CoT mechanism.

Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fern\'andez A cognitive architecture must decide not only how to reason but how long to think and what deserves the effort, so the authors build a minimal complete system — a recurrent reasoner with adaptive halting, a homeostatic control field, and a value module — and ask of each component whether its function emerges from gradient descent or must be explicitly computed. Competence emerges, but the apparent advantage of learned PonderNet-style halting vanishes once the readout is equalized (residual +0.000), and value does not emerge at all: trained couplings capture none of a payoff an explicit allocator captures completely (+0.151, routing correlation +0.79). On a frozen LLM actuator, self-consistency voting yields a small measured gain (+0.0236) while inter-sample agreement is nearly worthless as a stopping signal because its mass concentrates on wrong answers. A pre-registered prediction held up: under cliff-shaped costs, value-based allocation pays +0.1312, roughly seven times the smooth-cost estimate, because the cliff multiplies attainable range fivefold.

Can Large Language Models "Hyper-Thread"?

Fei Ding Language models emit tokens serially, and the question posed is whether a single generation step can carry multiple coordinated tasks that share state — a "Model Hyper-Threading Hypothesis" in which broader attention dispersion is a mechanism for concurrency rather than a symptom of interference. Three conditions are compared — a baseline, Serial Functional Scheduling, and Concurrent Functional Loading — using accuracy, output-token distributions, and attention metrics. On an AIME 2025 development set, Concurrent Functional Loading achieves the highest accuracy while producing similar or shorter typical output lengths than the serial condition, alongside greater attention dispersion and higher task-relevant coverage, though with a heavier length tail. The authors are explicit that within-step concurrency and its causal mechanism remain untested, calling the evidence behavioral and correlational.

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou, Zhiheng Zhang, Zhiyang Xie et al. Language models handle facts well but falter when asked to apply externally supplied procedural rules at scale, especially when the rule pool is large. RuleWorld is a benchmark that treats rules as globally reusable abstract units rather than instance-specific facts, with single-rule, parallel multi-rule, and multi-hop scenarios. The accompanying DynaRule framework injects the rule set into the KV cache and makes retrieval an internal, learnable step-wise process, using Stacked Step-Level Attention Training and a special search token so the model can re-attend to the most relevant rules at each reasoning step and drop stale ones. It raises average question-answering accuracy by up to 19 points and holds above 85% Recall@1 with a pool of 10,000 rules.

Compositional Chain-of-Relations for Faithful Knowledge Graph Question Answering with Large Language Models

Chenhui Liu, Jianpeng Zhou, Jiahai Wang Complex knowledge graph question answering couples candidate retrieval with constraint handling, and existing agent-based methods ground only the first: they explore entity by entity, pruning to a fixed-size subset at each hop, which discards valid entities, and they resolve query constraints from the model's internal knowledge, leaving answers unverifiable. CCoR (Compositional Chain-of-Relations) switches to relation-centric exploration, using relations rather than entities as the search unit to sidestep entity pruning, and grounds both phases in the graph via two chains: a main chain that retrieves candidates and a constraint chain that verifies each constraint through explicit graph traversal. Across four knowledge graph question answering benchmarks it improves accuracy, faithfulness, and efficiency over strong baselines, with the largest gains on complex multi-hop queries.

DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation

Guhan Chen, Songtao Tian, Bohan Li, Hejin Wang, YeXin Xie, Zixiong Yu Iterative preference optimization for mathematical reasoning stalls on signal scarcity: as the student model improves, a fixed problem set produces rollouts that are uniformly solved or uniformly failed, yielding few usable preference pairs. DIAG reshapes the practice distribution in two phases — diagnosing the yield of valid preference pairs per topic to allocate quotas via an Empirical Bayes shrinkage estimator, then having a teacher synthesize new problem variants from the student's own failure traces. The authors interpret this theoretically as a teacher-mediated approximation to KL-regularized reweighting toward the student's competence boundary, where pair yield is maximized, and report that DIAG raises pair yield across iterations and improves reasoning under an equal-cost training budget.

Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy

Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos, Alexandros Potamianos Reasoning traces from large reasoning models (LRMs) are now widely published, but the kinds of thinking they contain step by step remain largely uncharacterized. The proposed framework automatically annotates individual reasoning steps against Bloom's Taxonomy, which sorts thinking into six cognitive levels such as Remembering, Applying, and Evaluating, allowing large-scale comparison of thinking patterns across models and datasets. Applying it reveals both shared and model-specific reasoning styles, and shows that the thinking-type composition of a trace correlates with whether the final answer is correct, offering a signal that could be used to steer reasoning quality.
2 more specialized papers